For Penetration Testing Firms

Run the test, not the document

Hexecution is built around the engagement itself. A port-scan ladder that goes deeper only on hosts that answer. Nessus configured, launched and imported without a human touching it. Findings that carry instances and locations, so a retest closes one URL rather than a whole report. And a peer-review gate that stops a report shipping unapproved.

It runs the commercial side too, pipeline through invoicing, and an MCP server and API let an agent do what a person does in the browser.

The Reality

Most firms run on a patchwork

A reporting tool for deliverables. A CRM for the pipeline. Spreadsheets for scoping. Word for templates. A couple of consultants maintaining scripts on the side. Every seam between those tools is a place where things get dropped, re-keyed, or forgotten. And underneath it all, tooling that treats a pentest like a generic document project. We got tired of it, so we built a platform where scope, scans, findings, instances and retests are first-class objects, and then we ran 850+ assessments through it.

It runs the test, not the document

Scans are orchestrated, not imported. A progressive port-scan ladder deepens only on hosts that answer, Nessus is configured and launched and imported without a human step, and results diff against the last run. Findings carry instances and locations, and a retest resolves one location at a time.

Content that compounds

500+ curated finding templates, Nessus plugin IDs mapped to them already, per-instance CVSS 3.1 and 4.0, field inheritance from a sibling instance, structured OWASP, CWE and CVE tags, and DOCX round-trip so your existing Word reports come with you.

It runs the practice too

Deal pipeline, a scoping engine with the hour math your spreadsheet is doing badly, e-signature that fires the won-side effects, Xero invoicing, commission runs, renewals and win-back. The parts of running a firm that reporting tools leave in your CRM.

Scope, Recon & Assets

Assessments you do not babysit

Orchestrating scans is the easy half. The half that saves your consultants real hours is that the results are the same objects everything else reads. Define scope once and it generates the scope files in each tool's own format, drives the scans, populates the asset inventory, seeds the service-coverage worksheet and the application-path list, attaches to findings as affected locations, becomes the host and service appendices in the report, and is the baseline next year's engagement diffs against.

Nobody retypes anything, and nothing quietly disagrees with anything else.

A port-scan ladder that spends time where it pays

Wave one hits every address at the top 100 ports. Waves two through four hit only the hosts that answered, at the next 900, the next 9,000, then everything. Port exclusion makes "the next N" literal, so nothing is rescanned. Timing is measured per host, so two slow firewalled boxes do not cost the other thirteen a deeper wave.

Nessus, configured and launched for you

Configure, launch, poll, export, import, with no human step. Hexecution builds the target list and the explicit open-port range from the nmap results. We audited two hand-run scans once: they agreed on two settings and diverged on five. A custom Nessus policy cannot fix that, because it exposes 11 editable settings out of 3,912 and the port range is not one of them.

Every run diffed against the last

Nuclei results are kept in a ledger keyed by template and host, with first seen, last seen and resolved dates, so a finding is brand new, returned, or gone. Listening services diff the same way. Schedule daily, weekly, monthly or quarterly, and route the changes to a Slack channel.

Three scans, one view: what is brand new, what returned after disappearing, and what dropped out since the last run
Service inventory with change detection: a new port on a host you scanned last week is flagged, not buried
Hundreds of open ports collapse into the handful of distinct applications actually worth testing, with screenshots attached

One scope definition, read by everything downstream. Scope in, tool-specific scope files out. Scan results in, asset inventory out. Domains in, breached credentials out, deduplicated across sources and exportable as an attack list your password spraying can read directly. Asset inventory into the coverage worksheet, the application-path list and the finding's affected locations. Those into the report appendices, and into the year-over-year comparison on the same scope next time. The integrations between the data are the product; the orchestration is just how the data arrives.

Findings

Structured records, not paragraphs

A finding has instances. An instance has locations. That model is the difference between “the report says this is still open” and “three of these four URLs are fixed and the fourth is not.”

500+ templates that already know your scanners

Curated writeups with Nessus plugin IDs mapped to them, so an imported scan lands on the right template instead of a blank editor. Structured OWASP, CWE, CVE and root-cause tags. Per-instance CVSS in both 3.1 and 4.0, and any section can be inherited from a sibling instance rather than retyped.

Retest resolves a location, not a report

Each affected location carries its own remediation status. A retest round is requested against the finding, tested against the original evidence, and recorded per location, so partial remediation is visible instead of rounded to open or closed.

AI review that does not leak your client

Six models available through one router (Claude, GPT and Gemini families), with zero-data-retention routing on by default: a model with no ZDR endpoint fails closed rather than quietly falling back. Screenshots never leave. Org names, URLs, IPs, emails and paths are replaced with reversible tokens before anything is sent. Bring your own API key.

The findings workspace: every row carries its instance count, location count and review state
Four affected URLs, four independent remediation states, and a retest that can close them one at a time

Coverage

Proof the work actually happened

The uncomfortable question on any engagement is what did not get tested. Three features answer it before a client asks.

A methodology checklist with honest waivers

Seeded per assessment type from an admin-editable template. Items are required, required-if-applicable, or optional. Skipping one takes a category and a written reason, so a reviewer can tell a deliberate decision from neglect six months later.

Application paths, diffed year over year

Every URL and API operation identified or tested, tagged with the auth roles that reach it and where it was discovered (crawl, Burp, OpenAPI, GraphQL introspection). New paths since the last engagement are flagged, so a v2 API that appeared in March does not quietly go untested.

Service coverage that collapses the noise

Four hundred open ports across sixty hosts becomes a short list of distinct web applications plus a tail you can mark reviewed in a few clicks, seeded straight from the scan results rather than retyped.

The methodology, itemized and tracked: 57 of 132 done, 38 of 81 required, every tick stamped with who did it and when
And the template it is seeded from: yours to edit, per assessment type, with guidance attached to each item and edits dated so you can see what changed
19 of 31 paths tested, 5 new since last year, each tagged with the roles that reach it and the tester's note on what was tried

Review & Delivery

Nothing ships unreviewed

Peer review is a state machine with a gate on the end, not a comment thread people mean to get to.

Review that survives a second round

Draft, ready, in progress, changes requested, approved. Each requested change is accepted, rejected or reopened against a frozen snapshot of the prose, so a reviewer sees exactly what moved between rounds. A delivery retro on scope, client readiness, timeliness and communication has to be completed before an assessment can be approved.

An AI pass before the human one

A 0-100 quality score per report, severity re-derived from the evidence and compared against what the tester picked, plus typo and gap detection. The prompt carries our severity rubric and finding-writing guide, so the human review is spent on judgment rather than proofreading.

A delivery queue that names the blocker

Per assessment: findings approved, retests completed, required checklist items resolved, peer review approved. The sections are unweighted, so 99% of findings approved cannot hide a peer review that never started. Unclaimed retests sit in a team-wide queue anyone can claim.

The queue a tech lead opens first: what is blocking each engagement, and which retests nobody has picked up
The review state machine, and the delivery retro that gates approval
One click to DOCX, PDF, CSV, executive summary or letter of attestation, with the peer-review state shown next to the export button

Delivery is a link, not an attachment. Reports go out as expiring, revocable share links rather than a PDF loose in someone's inbox. Multi-assessment projects export as one combined report. Findings push to Jira with severity mapped to priority and evidence images attached.

Seen enough of the delivery side? Tell us how many consultants you have and what you run today, and we will show you the rest live.

Request Early Access

AI-Native

Built to be driven by an agent

Not a chat box bolted onto a reporting tool. Most of the API has no user interface at all, because it exists so that agent skills, the MCP server and CI can do what a person does in the browser. An agent works through the same permissions, the same audit trail and the same review gates as your consultants.

Point Claude Code at your own deployment

18 read tools and 9 write tools over stdio: read scope, recon, coverage and findings, draft findings, attach evidence, add validation notes, launch scans. Each tool is a thin wrapper over exactly one API endpoint, so it inherits that endpoint's permission checks and cannot exceed them. It installs against your own deployment and runs on your consultants' own Claude subscriptions.

Your methodology, written as skills

Skills encode how your firm tests, and they run against your own data rather than a vendor's idea of a workflow. We run ours against this platform every day: severity scored against a calibrated rubric, findings triaged for false positives before a human reads them, coverage checked before delivery, timesheets and client status drafted from the week's activity.

Agent work is labelled, not laundered

Every instance an agent creates is stamped with the skill, method and model that produced it when the skill supplies them, and the API warns on every write that does not, separately from the consultant credited with identifying the finding. Nothing an agent writes reaches a client without passing the same peer-review gate, and the scorecard credits the human who directed the work rather than the tool that typed it.

We threat-modelled our own agent surface, and wrote it down. Scanner output, HTTP banners, TLS certificate fields, page titles and crawled paths all originate from the hosts you are testing. The moment a tool returns them they become instructions-eligible text in an agent's context, sitting alongside tools that change customer-visible data and spend money on paid APIs. Our internal documentation names every tool that carries third-party-influenced text and pairs it with the state-changing tools in the same window. Write tools are not registered at all unless you enable them (the hosted workspace turns them on, and every write still prompts), the API key is the real gate and is enforced server-side, and a client-scoped key cannot hold a write permission under any configuration. You test things for a living. You were going to ask.

Your Team

Who is free, and who is on what

Staffing, utilization and billability read the same records the engagements already produce, so the capacity picture and the delivery plan cannot disagree.

Capacity you can see a quarter out

A scheduling Gantt with per-consultant utilization, hours free per week, and over-allocation flagged in red. Tech-lead time is weighted, so a TL across six engagements does not read as 600% booked. Unbooked consultants surface on their own.

Timesheets people actually fill in

Rows seeded from the assignments already on the schedule, copy-from-last-week, and reminders that escalate: a direct message to whoever is late, then a team digest if it stays late. Hours land against the assessment, not a free-text project code.

Utilization, delivery team and whole firm

Logged against available hours, billable percentage, and the split between testing and reporting time, for the delivery team and for everyone. The number you want before deciding whether you can take the next engagement or need to hire.

Allocations, utilization and free capacity per week, with over-allocation flagged before it becomes a delivery problem
Delivery team and whole firm side by side, with the testing-versus-reporting split that tells you where the hours actually went

Measurement

Whether you are actually getting better

The obvious metrics all lie. Finding counts reward whoever books the most hours. Severity averages reward whoever rates hardest. Everything here is built to survive that.

A scorecard that corrects for luck

Impact per scoped hour, divided by the baseline for that assessment type, weighted by each person's mix. Scores shrink toward typical on small samples, thin-sample consultants are marked and dropped out of the ranking, and the page tells you in writing not to read a 0.88 against a 0.95 as a ranking. It is built to resist being used badly.

Novelty mix, not just finding counts

Every credited finding sorts into scanner output, standard AD tooling, hunting, or genuinely novel, quoted per 40 logged hours so booking more hours does not inflate it. Titles the taxonomy does not recognize stay in the denominator rather than being guessed into a tier, and the novel findings are listed by name so the number can be checked rather than trusted.

A trend with a built-in lie detector

Impact per hour and findings per hour are plotted together on purpose. Both rising is real improvement. Impact rising while findings stay flat is severity-rating drift, meaning you are rating the same work higher, and the chart says so rather than letting you take the win.

Type-normalized impact per scoped hour, ranked, with small samples shrunk toward typical and thin-sample rows pushed out of the ranking entirely
Are we improving? Impact per hour against findings per hour by quarter, per-type trend lines, and each tech lead’s value-add on engagements they did not test on themselves
Scanner, standard AD tooling, hunting or novel, per 40 logged hours, with the novel titles listed underneath
And per engagement: this one scored 0.76x a typical test of its type, against a cohort of fifteen

The Commercial Side

Lead to invoice, without leaving the platform

This is the part reporting platforms do not attempt, and the part that is currently three spreadsheets and a CRM you are not updating.

Scoping with the hour math built in

Hours are computed from a sublinear curve per assessment type, not multiplied out linearly, with project-manager and tech-lead overhead redistributed across the committed lines. Seventeen assessment types, three rate tiers, and optional lines priced as included, alternative or add-on so a client can see the trade.

Signature that fires the side effects

An interactive proposal portal with live pricing, then e-signature on the SOW. Completion marks the deal won, stamps the date, enrols it in renewal tracking, posts to Slack and drafts the invoice in Xero. Turning it into a project is one click, on your terms.

The money follows through

Xero invoicing, commission runs with tiered referral fees across several payment models, a renewals queue that knows when last year's engagement is due, and a win-back dashboard for the ones that went quiet.

Pipeline with the scoped hours and price attached to the card, so the deal and the delivery plan are the same object
Revenue against goal, hours sold, win rate and days to close, each against the same period in prior years
What you are actually selling, by assessment type, in hours and dollars against prior years. Read from the SOW scoping lines on won deals, so nobody has to tag anything.

The analytics are a by-product, not a data-entry job. Nobody tags a deal, fills in a category, or reconciles a spreadsheet at month end. Renewals know when last year's engagement is due because the project shape is stored rather than remembered. Commission runs read the same delivered assessments the invoices do. Retention, new versus existing business, and revenue by owner all fall out of records the pipeline already holds.

Seen how it runs the practice? Tell us how many consultants you have and what you run today, and we will show you the rest live.

Request Early Access

Getting In & Getting Out

Your content comes with you

Switching cost is the real objection. Here is how it is handled.

Bring your last five engagements

DOCX import turns delivered Word reports into structured assessments through a preview, diff and execute flow, so you see what it will create before it creates it. We have run it across whole multi-year backlogs. There is a CSV importer for platform exports too. If you want to try before you commit, send us five reports and we will have them imported before your first new engagement.

An API you can build against

Bearer tokens that live until you revoke them rather than expiring every fifteen minutes. Scopes per key, and a key can be restricted to specific assessments. Bulk-create up to 100 findings in a call. Live reference at /api/v1/docs. Client keys are clamped read-only. Every unscoped resource route was read end to end during an internal access-control review.

How it compares

Reporting platforms format the deliverable. Hexecution runs the test that produces it, then handles the practice around it too.

Swipe sideways to see each column.

How Hexecution compares with dedicated pentest reporting platforms and with a home-built stack of spreadsheets and Word templates.
CapabilityHexecutionReporting PlatformsSpreadsheets & Word
Orchestrates scans and diffs results between runsPartialManual
Automated Nessus config, launch, export and importManual
Findings library pre-mapped to scanner plugin IDsPartial
Per-instance CVSS and per-location remediationPartialManual
Retest tracked per location, not per reportPartial
AI review built in, with zero-data-retention routingAdd-on
Peer review gate that blocks deliveryPartial
Project to assessment hierarchy, combined reportsPartialManual
DOCX round-trip import and exportPartialManual
MCP server: run an engagement from Claude Code
API tokens scoped to a single assessment
Calibrated consultant scorecard and novelty mix
Deal pipeline, scoping and pricing engineSpreadsheet
E-signature, invoicing and commissionsSpreadsheet
Renewal automation and win-back queue
Vendor still runs its own engagements on itYou

“Reporting Platforms” covers dedicated pentest reporting and management tools as a category, and they differ from each other: at least one runs native scans, and at least one ships a far larger writeup library than ours. The rows are worded as what Hexecution does, not as what they cannot. Check any of them against your own shortlist.

Straight Answers

The questions you are going to ask anyway

You would ask these on the first call. Here they are before it.

“You compete with me.”

We do, and you should ask. The answer is a single-tenant deployment: your own database, your own domain, your own isolated cloud account. Your client data does not live alongside ours and we do not query it. We are happy to put non-solicitation terms in the agreement, and to talk about what happens if we ever bid against each other. If that answer is not good enough for you, it should not be, and we would rather find that out now.

“How far along is this really?”

We have run our own engagements on it since 2022, more than 850 assessments so far. The first external deployments are hand-run by us, and we will tell you before you sign which parts still assume our name and our integrations.

“Who is actually building this?”

A working pentest firm, funded by services revenue rather than by a runway. Every feature exists because we hit the problem on an engagement. Design partners get a weekly call, roadmap input, and a direct line to the people who test for a living.

Want to run your firm on Hexecution?

A single-tenant deployment, a weekly call, and real influence over what gets built. Tell us how many consultants you have and what you run today, and we will show you how we run every engagement, business side included.