The way most companies run internal pentest is a fixed point of failure. A team comes in for two or three weeks every twelve months. They build a domain map, they find some misconfigured shares, they get domain admin from an old service account, they write a report. Procurement signs off, executives feel reassured, and the report sits on a SharePoint while the environment it described drifts further from reality every day.
The external attack surface has at least the benefit of being scanned constantly by every bot on the internet. The internal one does not. It changes when an engineer joins, when a service account is created, when a new server is stood up in a region nobody is watching closely. By the time the next annual engagement comes around, the report you are commissioning is going to be about a network that no longer exists.
This is the part we built continuous pentest to fix.
What rolling actually means
The pipeline runs three classes of operation against your environment, on a rolling cadence, with no manual kickoff. After a few months of operation a typical customer is at 142 days uptime with zero manual triggers. The cadence is split across four triggers: on every push to a connected repository, nightly, weekly, and ad hoc when something interesting changes.
The three operation types live side by side in the same queue:
- diff scan runs against repositories the moment a change lands, looking at the surface area that was actually modified rather than the whole codebase. Fast, narrow, and cheap to run hundreds of times a day.
- full pentest runs against network targets on a longer cadence. Internal subnets, segmented environments, jump hosts, the production prod-eu-1 and prod-us-2 of the world. This is where the agentic internal pentest does its real work.
- redteam runs adversary emulation playbooks against the environment. APT29 today, Volt Typhoon tomorrow, whichever TTP profile is relevant to the threat model the customer actually cares about.
The numbers the pipeline reports next to each run, things like 47/120 or 84% coverage, are how far through the playbook the operation got and how much of the target surface it touched. They are the same numbers an internal red team lead would track on a manual engagement, except they accumulate every day instead of resetting every twelve months.
Why the internal side is the one that ages the fastest
External pentest output stays useful longer than internal pentest output. The shape of an externally exposed application does not change as fast as the shape of an Active Directory forest with thirty thousand objects in it.
Inside the perimeter, things move. A new server is provisioned with a default local admin password that nobody rotated. A helpdesk technician is added to a group that gets DCSync on a child domain. A misconfigured GPO pushes a registry change that re-enables a protocol that was disabled three quarters ago. A laptop is enrolled into a tenant with an old certificate. None of these will fail compliance on their own. All of them compound, and any one of them is enough to change the post-compromise path an attacker would take.
An annual internal pentest catches the state of those compounding mistakes on the day the engagement ends. The next day, the report starts losing accuracy. By month six it is a historical document. By month twelve, when you commission the next one, you have spent most of a year defending an environment you were not actually looking at.
Continuous pentest replaces that cadence. The agents sit inside the environment, mapping, probing, and re-validating the things that have changed since the last sweep. When somebody on your side wants to know whether a specific user could reach a specific resource as of this morning, the answer is in the platform.
What the agents actually do internally
The internal pentest piece is agentic, which is a word we are aware has been worn thin. In our case it means specifically that the operation is built around the same loop a human pentester runs: collect information, form a hypothesis, attempt a step, observe the result, adapt. The agents do not just run a vulnerability scanner and rebrand the output. They enumerate Active Directory the way an operator would, build relationship graphs the way BloodHound does, identify the shortest path to a high-value object, and attempt the lateral movement that path requires.
What changes from a human engagement is the cadence and the consistency. The agents do the unglamorous parts of the work every night without getting bored. They re-run the same checks against the same forest after every change. When something new appears in the network, like a freshly joined workstation or a newly granted permission, the next sweep includes it without anyone scheduling anything.
The output is a rolling record of what the post-compromise path looks like in your environment, today, against the configuration that exists today.
Adversary emulation is part of the pipeline, not a separate engagement
The redteam runs in the pipeline are not decorative. Adversary emulation lives in the same queue as the diff scans and the full pentests, against the same environment, with the same instrumentation. When a sweep runs the APT29 playbook against your tenant, the artifacts that get produced (the events, the logs, the alerts that fired, the ones that did not) are stitched together with the controls map you have in the platform.
The point is to find out, on a continuous basis, whether your detections actually fire on the TTPs that matter for your threat model. Not whether the SIEM has a rule for the technique. Whether the rule, the data source it depends on, the agent feeding it, and the analyst workflow on the other end of it all worked end to end when the technique was actually executed.
That answer changes constantly, because everything underneath it changes constantly. Continuous emulation is the only way to keep the answer current.
What "0 manual kickoffs" implies
The pipeline runs without anyone scheduling it. There is no security engineer on your side who needs to remember to start the quarterly internal engagement. There is no procurement cycle for the next sweep. The cadence is the product. The runs accumulate, the queue stays full, and the team you have can spend their time on the findings rather than on coordinating the engagement that produces them.
This is the part that compounds. A team running continuous pentest for twelve months ends the year with a year of evidence about how the environment changed and how the post-compromise picture shifted with it. A team that runs an annual pentest ends the year with one report.
Where this fits with everything else
Continuous pentest does not replace a human-led engagement on a piece of the environment that needs a specialist. If you are about to ship a new identity platform, or you want a focused two-week look at a specific high-value system, that is a different kind of work. We run those engagements too, and the platform's accumulated history makes them substantially more efficient because the operator does not start from zero.
Continuous pentest also does not replace the SOC. The detections are run against the SOC, not in place of it. The point of the pipeline is to feed the defensive side better evidence about what does and does not work.
What continuous pentest does replace is the assumption that an annual internal engagement plus a vulnerability scanner together constitute internal coverage. They do not. Everything important about the internal network is the part that changed since the last time anyone looked at it, and the last time anyone looked is the variable you want to drive to days, not months.
What the agents will do is run the unglamorous, repeatable, high-cadence part of internal pentest every night, on every push, and on every change, without the cost structure that makes that frequency impossible with a pure services model. The result is an internal posture that is measured continuously instead of sampled annually.
If you want to see what the pipeline looks like against your own environment, we can run a scoped sweep on a segmented test network and show you the kind of evidence the queue produces. It tends to be the most efficient way to decide whether this is the right shape of coverage for your team.