External Penetration Testing is now live in Pragma Core. You hand the platform a scope of CIDRs, IPs, ASNs, or domains, pick a depth, sign an authorization, and a lead orchestrator agent runs a full external engagement for you: reconnaissance, attack-surface discovery, vulnerability assessment, and, at the top tier, a single non-destructive proof of concept against a confirmed weakness. The part that changes the economics is that you can point this at your perimeter on a schedule instead of once a year, so a forgotten staging host or a management portal that leaked to the public internet last Tuesday shows up while it still matters.
The article closes by placing External Penetration Testing inside the wider Pragma Core model of continuous offensive security across three domains in one workspace.
Part I. Executive breakdown
What shipped
Pragma Core already ran an internal network engagement: drop a small agent on a host you control, and an AI operator walks the internal pentest playbook with full Active Directory and BloodHound coverage. External Penetration Testing is the outside-in counterpart. There is no agent to deploy and no VPN to a jump box. You describe your internet-facing footprint in a text box, and the platform does the rest.
The mental model is simple. Think of your organization's perimeter as the outside walls of a building: the doors, the windows, the service entrances that anyone on the street can walk up to and rattle. An external pentest is a professional walking that perimeter, writing down every door they find, noting which ones are unlocked, and, where you allow it, opening exactly one to prove it really was unlocked, without stepping inside or breaking anything. Pragma Core's version does that walk with an AI operator instead of a scheduled human consultant, and it does it against your live, current perimeter each time you run it.
The result is not a raw scanner dump. It is a finding pack: a reconciled inventory of hosts and exposed services, a deduplicated list of findings, and a full event timeline that shows what the operator did, in what order, and why it stopped where it stopped. A phase gate refuses to close the engagement until every required checkpoint has been addressed, so you never get a half-finished test dressed up as a clean report.
Who this is for
| Audience | Why it matters |
|---|---|
| Teams with a sprawling or shifting perimeter | Rebuild the live picture of what you actually expose every time you run, instead of trusting a stale asset inventory. |
| Security teams without a standing pentest budget | Validate that a finding is real at the Exploitation tier with a read-only proof of concept, without scheduling and paying for a separate consultancy engagement each time. |
| Compliance and audit owners | Run recurring external assessments on a cadence and keep an audited, phase-by-phase timeline of every engagement. |
| M&A and vendor-risk functions | Point an engagement at a target company's ASN or domain, with authorization in hand, and get a real read on their internet-facing risk in hours rather than weeks. |
The common thread is exposure that changes faster than an annual test can keep up with. A new service goes live, a certificate lapses, a default credential ships to production, an old VPN appliance never gets decommissioned. Each of those is a window that opened after your last review, and each is the kind of thing an attacker on the internet finds first.
Why external testing matters
The uncomfortable truth of perimeter security is that most high-impact external findings are not exotic exploit chains. They are forgotten assets. A staging host that was supposed to be temporary. A management portal that should never have been public. An admin interface still carrying the credentials it shipped with. The value of an external engagement is less about clever exploitation and more about honest, current visibility: knowing what you actually expose, not what your documentation says you expose.
That visibility decays quickly. The perimeter you tested in January is not the perimeter you run in July. This is exactly why continuous, on-demand external testing beats the once-a-year model. The change you care about is the one that appeared last week, and a yearly test will not see it for another eleven months.
Recommended actions
- Run a Recon-depth engagement against your own perimeter first. It is the fastest way to see whether your live attack surface matches your asset inventory. Most teams find at least one host they had forgotten about.
- Put an Assessment-depth engagement on a schedule. Weekly or monthly, depending on how fast you ship. This is the cadence that catches a lapsed certificate or a newly exposed service while it is still cheap to fix.
- Reserve the Exploitation tier for the findings you need to prove. Use it to turn a "this looks exposed" into a "this is confirmed exploitable, here is the read-only proof," which is the version that gets prioritized in an engineering sprint.
- Confirm authorization before you scope anything you do not own. For vendor or acquisition baselining, get written authorization for the target's ASN or domain first. The platform enforces an attestation, but the legal responsibility is yours.
- Feed the finding pack back into remediation, not into a drawer. The output is structured so hosts, services, and findings reconcile against a database. Treat each recurring engagement as a diff against the last one.
Part II. Technical breakdown
Architecture: an orchestrator and its phase subagents
The engine is not a single monolithic scanner. It is a lead orchestrator agent that owns the engagement and spawns a dedicated subagent for each phase of the methodology. Each subagent runs its own tools over SSH on the Kali worker, records what it finds, and hands control back to the orchestrator. The orchestrator's job is to keep the methodology honest from reconnaissance through wrap-up, so no phase gets skipped and no phase runs out of order.
This separation matters for two reasons. First, it maps cleanly onto how a real external engagement is structured, where discovery genuinely must precede assessment and assessment must precede any exploitation attempt. Second, it makes the whole run auditable. Because each subagent is bound to a phase, every command and every finding is attributable to a specific stage of the methodology, and the timeline reads like a methodology, not like a log file.
The tooling itself runs on a hardened Kali worker, the same shared worker pool that drives Pragma Core's DAST. There is nothing for you to stand up. The orchestrator picks up a worker, drives it over SSH, and tears the session down when the engagement closes.
Defining scope: four token types, expanded on the worker
The scope textarea accepts four kinds of token, and a whitelist grammar validates every one of them so free text never reaches the engine.
CIDR 203.0.113.0/24 a network range, enumerated to live hosts during discovery
IP 198.51.100.10 a single address, used exactly as given
ASN AS64500 an autonomous system number, resolved to its announced prefixes
Domain acme.com resolved to IPs and expanded to subdomains
The design decision worth calling out is that ASNs and domains are expanded on the Kali worker during recon, not inside the platform. The platform carries no whois or DNS dependency of its own. An ASN token resolves to its announced prefixes on the worker via whois and BGP lookups, so a single AS64500 can stand in for an entire organization's announced footprint. A domain is resolved to IPs and expanded to live subdomains on the worker with subfinder, amass, and dnsx, then folded back into the host inventory. You write one token; the operator turns it into the real list of hosts to assess.
Picking a depth: the tier is the contract
Depth is not a cosmetic setting. It sets the time budget, the set of phases that run, and how aggressive the operator is allowed to be. Because the phase set is gated to the depth you pick, the depth you select is the depth the operator executes, with no scope creep into more aggressive behavior than you authorized. An admin cap can clamp the budget down further for a whole workspace.
| Depth | Time budget | What runs |
|---|---|---|
| Recon | 90 minutes | Passive OSINT, DNS and subdomain enumeration, light live-host and port discovery. Maps the attack surface without intrusive scanning. |
| Assessment | 6 hours | Everything in Recon, plus a full port and service scan, nuclei vulnerability scanning, and TLS, web-exposure, and default-credential checks. Non-destructive throughout. |
| Exploitation | 18 hours | Everything in Assessment, plus active, non-destructive proof-of-concept exploitation of confirmed vulnerabilities, with an impact write-up against each finding. |
The tiering is deliberate. Recon answers "what do I expose." Assessment answers "which of those exposures are weak." Exploitation answers "is this specific weakness actually real," and it answers it with a single read-only proof rather than a spray of exploit attempts.
The phase-gated methodology
The methodology runs in five phases: recon, discovery, assessment, exploitation, and wrap-up. Each phase carries a required checklist, and the engagement cannot close until every checkpoint has been addressed. Two properties follow from that gate.
First, recorded findings are cross-checked against the database before the engagement is allowed to close, so the finding pack and the underlying inventory cannot silently disagree. Second, at the top tier the gate requires a real exploitation attempt to be present, which means an Exploitation-depth engagement cannot quietly downgrade itself into a glorified vulnerability scan.
Skipped checkpoints stay visible. If the operator stood down on a check, that decision is recorded on the timeline with a reason, rather than disappearing. You can see not only what the operator did, but where it deliberately chose not to act, and why.
Non-destructive by contract
This is the most important design constraint, so it is worth being precise about what it means. The engagement is non-destructive by contract: no denial of service, no data modification, no persistence, and no lateral movement beyond proving a single confirmed finding. That boundary is not a guideline in a runbook. It is written into the operator prompt and enforced by the depth-gated phase set, so only the Exploitation tier can run active checks at all, and even there the proof of concept is a single read-only demonstration.
Concretely, at the Exploitation tier the operator proves a finding with one read-only proof of concept and writes the impact up against that finding. It does not chain the foothold into deeper access. The point is to give you exploitability you can trust and a clear impact statement, not to actually breach and hold your environment.
Authorization and safety rails
Every engagement requires three things before it can launch: acceptance of the Rules of Engagement, an authorization attestation that you are permitted to test the targets, and a typed signature. This is the human checkpoint that keeps an external pentest from becoming an unauthorized one.
There are hard blocks on top of the attestation. Cloud metadata endpoints, link-local ranges, and the platform's own host are blocked outright, regardless of what you paste into the scope box. Combined with the whitelist grammar that validates every scope token, the effect is that both obvious footguns (scanning a cloud metadata service) and malformed input are refused before the engine ever runs.
Semantic finding dedup
Findings are deduplicated by category, host, target, and technique, rather than by raw text. That distinction matters more than it sounds. Text-based dedup fails the moment two tools describe the same exposed service with slightly different wording. Semantic dedup collapses "the same exposed service reported twice" into one row even when the descriptions differ. Critical and high findings raise a notification the moment they land, so you are not waiting for the wrap-up to learn that something serious turned up.
The toolkit the operator runs
The recon, scanning, and exploitation tooling all runs on the Kali worker. The operator reaches for a fairly standard, well-understood external toolkit, which is a feature, not a limitation: these are the tools a human external tester would use, driven by an agent.
Scope expansion whois + BGP lookups, subfinder, amass, dnsx
Live-host discovery nmap service detection, masscan
Fingerprinting httpx, whatweb
Vulnerability scan nuclei
TLS audit testssl
Exposure checks web exposure, default-credential, exposed-service review, info-leak detection
Exploitation non-destructive proof of concept, CVE confirmation
What an engagement actually records
Here is a lightly abbreviated excerpt from a sample Assessment-depth timeline, to make the phased structure concrete. Note how each event is tagged with its phase and an outcome, and how a finding is a distinct outcome type from a completed action.
00:02:31 recon ASN AS64500 resolved to two announced prefixes via whois completed
00:09:48 recon subdomain enum on acme.com, 38 names, 11 resolving in-scope completed
00:24:05 discovery port and service scan, 17 live hosts, 41 services fingerprinted completed
01:06:52 assessment exposed management portal on 8443 with default credentials finding
01:38:20 assessment testssl audit, legacy TLS still enabled on the mail gateway finding
02:55:14 assessment nuclei flagged a known CVE on the app load balancer, confirmed finding
03:21:40 assessment single non-destructive proof of concept, read-only response completed
03:44:09 wrap-up 6 findings recorded, host and service inventory reconciled completed
Two of those lines tell the whole story of the design. The default-credential finding at 01:06:52 is the classic external result: not a memory-corruption exploit, just a portal that should never have been reachable, carrying credentials it should never have kept. And the 03:21:40 proof of concept is read-only by contract, captured as evidence, and goes no further.
A note on why this is an AI operator, not a scanner
It would have been easier to ship a scheduled scanner with a nice report template. The reason to build an orchestrator-plus-subagents operator instead is that an external engagement is a sequence of decisions, not a fixed pipeline. What you scan depends on what discovery found. Whether a default-credential check is worth running depends on what services fingerprinting turned up. Whether a finding is worth a proof of concept depends on whether the operator can confirm the version and the exposure. Those are judgment calls, made in order, and the phase-gated agent structure is what lets the platform make them while still keeping the whole run inside an auditable, non-destructive boundary.
What this changes for security teams
- Perimeter visibility becomes a cadence, not an event. The moment external testing is a scheduled Recon or Assessment engagement rather than an annual consultancy booking, the forgotten-asset problem shrinks. You find the new exposure in the week it appears.
- "Is this finding real?" gets a cheap answer. The Exploitation tier turns a suspected exposure into a confirmed, read-only proof without a separate engagement. That single fact changes how findings get prioritized, because confirmed exploitability outranks a scanner's maybe.
- Auditability is a first-class output, not an afterthought. A phase-gated timeline with visible skipped checkpoints and reasons is exactly the artifact an auditor or an incident reviewer wants, and it is produced automatically.
- Authorization has to be designed in, not bolted on. Offensive automation without an enforced attestation and hard target blocks is a liability. The safer pattern is the one here: whitelist the scope grammar, require a signature, and block dangerous targets outright regardless of input.
- The best external findings are still boring, and that is the point. Default credentials, a lapsed certificate, a management portal that leaked to the public internet. Tooling that treats reconnaissance and exposure review as the main event, not as a warm-up to exploitation, is tooling aligned with where perimeter risk actually lives.
How this fits the rest of Pragma Core
Pragma Core's premise is continuous offensive security across three domains in one workspace: application security over your code, supply chain, and web apps; infrastructure security over your internal and external networks; and adversary emulation for breach and attack simulation. External Penetration Testing slots into the infrastructure domain as the outside-in counterpart to the internal network engagement, and it inherits the platform patterns that make the other domains coherent.
One finding lifecycle across every domain
The same login, the same role-based access control, and the same finding lifecycle apply whether the finding came from a SAST scan of your code, an internal network pentest, or an external engagement. An externally exposed service and a hardcoded secret in a repository live in the same triage workflow, which is what lets a team reason about risk across the boundary between "our code" and "our perimeter" instead of in two disconnected tools.
The same Kali worker pool that drives DAST
External engagements run on the same shared hardened Kali worker pool that powers Pragma Core's dynamic application security testing. That shared foundation is why there is nothing to deploy for an external test: the worker infrastructure already exists for DAST, and the external operator drives it over SSH the same way.
Internal and external, two halves of the network picture
Infrastructure Penetration Tests cover the inside: drop an agent on a host you control and an AI operator walks Active Directory recon, BloodHound attack paths, ADCS abuse, and safe credential attacks. External Penetration Testing covers the outside: hand over a scope and the operator maps and assesses your perimeter. Run both and you have the full network picture, the doors an attacker rattles from the street and what they could reach if one of those doors gives way, in one workspace with one finding model.
Feeds naturally into adversary emulation
Once an external engagement has told you what an attacker can reach and confirmed which exposures are real, the natural next question is whether your detections would fire against a real adversary using them. That is the adversary emulation domain: pick a MITRE-aligned adversary like APT29 or FIN7, run it dry to storyboard or live to exercise your SOC, and watch each technique land with action-level evidence. External testing finds the exposure; emulation tests whether you would notice it being used.
Closing thoughts
External Penetration Testing is not, at bottom, a feature about running nmap and nuclei on a schedule. Any team can do that. It is a feature about turning an inherently sequential, judgment-heavy engagement, the kind that has traditionally required a human consultant and a calendar slot, into something you can run against your live perimeter as often as your perimeter changes, while keeping the whole thing inside an auditable, non-destructive, authorization-gated boundary. The hard part was never the scanning. It was making the operator make good decisions in order, prove a finding without breaking anything, and record every step so you can trust the result. That is the part Pragma Core built.
The same perimeter-drift problem lives in almost every organization that ships to the internet, and the gap between "we scanned once this year" and "we know what we expose this week" is the gap most breaches walk through. If your team wants to move from an annual snapshot to a continuous, audited read on your external attack surface, you can define a scope, pick a depth, and run your first engagement at pragma-core.com, or reach out for a demo.
Sources
- Pragma Core, product homepage and platform overview, pragma-core.com.