Back to blog

Your perimeter, tested every week: External Penetration Testing is live in Pragma Core

· · 16 min read
Your perimeter, tested every week: External Penetration Testing is live in Pragma Core

External Penetration Testing is now live in Pragma Core. You hand the platform a scope of CIDRs, IPs, ASNs, or domains, pick a depth, sign an authorization, and a lead orchestrator agent runs a full external engagement for you: reconnaissance, attack-surface discovery, vulnerability assessment, and, at the top tier, a single non-destructive proof of concept against a confirmed weakness. The part that changes the economics is that you can point this at your perimeter on a schedule instead of once a year, so a forgotten staging host or a management portal that leaked to the public internet last Tuesday shows up while it still matters.

The article closes by placing External Penetration Testing inside the wider Pragma Core model of continuous offensive security across three domains in one workspace.


Part I. Executive breakdown

What shipped

Pragma Core already ran an internal network engagement: drop a small agent on a host you control, and an AI operator walks the internal pentest playbook with full Active Directory and BloodHound coverage. External Penetration Testing is the outside-in counterpart. There is no agent to deploy and no VPN to a jump box. You describe your internet-facing footprint in a text box, and the platform does the rest.

The mental model is simple. Think of your organization's perimeter as the outside walls of a building: the doors, the windows, the service entrances that anyone on the street can walk up to and rattle. An external pentest is a professional walking that perimeter, writing down every door they find, noting which ones are unlocked, and, where you allow it, opening exactly one to prove it really was unlocked, without stepping inside or breaking anything. Pragma Core's version does that walk with an AI operator instead of a scheduled human consultant, and it does it against your live, current perimeter each time you run it.

The result is not a raw scanner dump. It is a finding pack: a reconciled inventory of hosts and exposed services, a deduplicated list of findings, and a full event timeline that shows what the operator did, in what order, and why it stopped where it stopped. A phase gate refuses to close the engagement until every required checkpoint has been addressed, so you never get a half-finished test dressed up as a clean report.

Who this is for

Audience Why it matters
Teams with a sprawling or shifting perimeter Rebuild the live picture of what you actually expose every time you run, instead of trusting a stale asset inventory.
Security teams without a standing pentest budget Validate that a finding is real at the Exploitation tier with a read-only proof of concept, without scheduling and paying for a separate consultancy engagement each time.
Compliance and audit owners Run recurring external assessments on a cadence and keep an audited, phase-by-phase timeline of every engagement.
M&A and vendor-risk functions Point an engagement at a target company's ASN or domain, with authorization in hand, and get a real read on their internet-facing risk in hours rather than weeks.

The common thread is exposure that changes faster than an annual test can keep up with. A new service goes live, a certificate lapses, a default credential ships to production, an old VPN appliance never gets decommissioned. Each of those is a window that opened after your last review, and each is the kind of thing an attacker on the internet finds first.

Why external testing matters

The uncomfortable truth of perimeter security is that most high-impact external findings are not exotic exploit chains. They are forgotten assets. A staging host that was supposed to be temporary. A management portal that should never have been public. An admin interface still carrying the credentials it shipped with. The value of an external engagement is less about clever exploitation and more about honest, current visibility: knowing what you actually expose, not what your documentation says you expose.

That visibility decays quickly. The perimeter you tested in January is not the perimeter you run in July. This is exactly why continuous, on-demand external testing beats the once-a-year model. The change you care about is the one that appeared last week, and a yearly test will not see it for another eleven months.

Recommended actions

  1. Run a Recon-depth engagement against your own perimeter first. It is the fastest way to see whether your live attack surface matches your asset inventory. Most teams find at least one host they had forgotten about.
  2. Put an Assessment-depth engagement on a schedule. Weekly or monthly, depending on how fast you ship. This is the cadence that catches a lapsed certificate or a newly exposed service while it is still cheap to fix.
  3. Reserve the Exploitation tier for the findings you need to prove. Use it to turn a "this looks exposed" into a "this is confirmed exploitable, here is the read-only proof," which is the version that gets prioritized in an engineering sprint.
  4. Confirm authorization before you scope anything you do not own. For vendor or acquisition baselining, get written authorization for the target's ASN or domain first. The platform enforces an attestation, but the legal responsibility is yours.
  5. Feed the finding pack back into remediation, not into a drawer. The output is structured so hosts, services, and findings reconcile against a database. Treat each recurring engagement as a diff against the last one.

Part II. Technical breakdown

Architecture: an orchestrator and its phase subagents

The engine is not a single monolithic scanner. It is a lead orchestrator agent that owns the engagement and spawns a dedicated subagent for each phase of the methodology. Each subagent runs its own tools over SSH on the Kali worker, records what it finds, and hands control back to the orchestrator. The orchestrator's job is to keep the methodology honest from reconnaissance through wrap-up, so no phase gets skipped and no phase runs out of order.

This separation matters for two reasons. First, it maps cleanly onto how a real external engagement is structured, where discovery genuinely must precede assessment and assessment must precede any exploitation attempt. Second, it makes the whole run auditable. Because each subagent is bound to a phase, every command and every finding is attributable to a specific stage of the methodology, and the timeline reads like a methodology, not like a log file.

The tooling itself runs on a hardened Kali worker, the same shared worker pool that drives Pragma Core's DAST. There is nothing for you to stand up. The orchestrator picks up a worker, drives it over SSH, and tears the session down when the engagement closes.

Defining scope: four token types, expanded on the worker

The scope textarea accepts four kinds of token, and a whitelist grammar validates every one of them so free text never reaches the engine.

CIDR     203.0.113.0/24     a network range, enumerated to live hosts during discovery
IP       198.51.100.10      a single address, used exactly as given
ASN      AS64500            an autonomous system number, resolved to its announced prefixes
Domain   acme.com           resolved to IPs and expanded to subdomains

The design decision worth calling out is that ASNs and domains are expanded on the Kali worker during recon, not inside the platform. The platform carries no whois or DNS dependency of its own. An ASN token resolves to its announced prefixes on the worker via whois and BGP lookups, so a single AS64500 can stand in for an entire organization's announced footprint. A domain is resolved to IPs and expanded to live subdomains on the worker with subfinder, amass, and dnsx, then folded back into the host inventory. You write one token; the operator turns it into the real list of hosts to assess.

Picking a depth: the tier is the contract

Depth is not a cosmetic setting. It sets the time budget, the set of phases that run, and how aggressive the operator is allowed to be. Because the phase set is gated to the depth you pick, the depth you select is the depth the operator executes, with no scope creep into more aggressive behavior than you authorized. An admin cap can clamp the budget down further for a whole workspace.

Depth Time budget What runs
Recon 90 minutes Passive OSINT, DNS and subdomain enumeration, light live-host and port discovery. Maps the attack surface without intrusive scanning.
Assessment 6 hours Everything in Recon, plus a full port and service scan, nuclei vulnerability scanning, and TLS, web-exposure, and default-credential checks. Non-destructive throughout.
Exploitation 18 hours Everything in Assessment, plus active, non-destructive proof-of-concept exploitation of confirmed vulnerabilities, with an impact write-up against each finding.

The tiering is deliberate. Recon answers "what do I expose." Assessment answers "which of those exposures are weak." Exploitation answers "is this specific weakness actually real," and it answers it with a single read-only proof rather than a spray of exploit attempts.

The phase-gated methodology

The methodology runs in five phases: recon, discovery, assessment, exploitation, and wrap-up. Each phase carries a required checklist, and the engagement cannot close until every checkpoint has been addressed. Two properties follow from that gate.

First, recorded findings are cross-checked against the database before the engagement is allowed to close, so the finding pack and the underlying inventory cannot silently disagree. Second, at the top tier the gate requires a real exploitation attempt to be present, which means an Exploitation-depth engagement cannot quietly downgrade itself into a glorified vulnerability scan.

Skipped checkpoints stay visible. If the operator stood down on a check, that decision is recorded on the timeline with a reason, rather than disappearing. You can see not only what the operator did, but where it deliberately chose not to act, and why.

Non-destructive by contract

This is the most important design constraint, so it is worth being precise about what it means. The engagement is non-destructive by contract: no denial of service, no data modification, no persistence, and no lateral movement beyond proving a single confirmed finding. That boundary is not a guideline in a runbook. It is written into the operator prompt and enforced by the depth-gated phase set, so only the Exploitation tier can run active checks at all, and even there the proof of concept is a single read-only demonstration.

Concretely, at the Exploitation tier the operator proves a finding with one read-only proof of concept and writes the impact up against that finding. It does not chain the foothold into deeper access. The point is to give you exploitability you can trust and a clear impact statement, not to actually breach and hold your environment.

Authorization and safety rails

Every engagement requires three things before it can launch: acceptance of the Rules of Engagement, an authorization attestation that you are permitted to test the targets, and a typed signature. This is the human checkpoint that keeps an external pentest from becoming an unauthorized one.

There are hard blocks on top of the attestation. Cloud metadata endpoints, link-local ranges, and the platform's own host are blocked outright, regardless of what you paste into the scope box. Combined with the whitelist grammar that validates every scope token, the effect is that both obvious footguns (scanning a cloud metadata service) and malformed input are refused before the engine ever runs.

Semantic finding dedup

Findings are deduplicated by category, host, target, and technique, rather than by raw text. That distinction matters more than it sounds. Text-based dedup fails the moment two tools describe the same exposed service with slightly different wording. Semantic dedup collapses "the same exposed service reported twice" into one row even when the descriptions differ. Critical and high findings raise a notification the moment they land, so you are not waiting for the wrap-up to learn that something serious turned up.

The toolkit the operator runs

The recon, scanning, and exploitation tooling all runs on the Kali worker. The operator reaches for a fairly standard, well-understood external toolkit, which is a feature, not a limitation: these are the tools a human external tester would use, driven by an agent.

Scope expansion     whois + BGP lookups, subfinder, amass, dnsx
Live-host discovery nmap service detection, masscan
Fingerprinting      httpx, whatweb
Vulnerability scan  nuclei
TLS audit           testssl
Exposure checks     web exposure, default-credential, exposed-service review, info-leak detection
Exploitation        non-destructive proof of concept, CVE confirmation

What an engagement actually records

Here is a lightly abbreviated excerpt from a sample Assessment-depth timeline, to make the phased structure concrete. Note how each event is tagged with its phase and an outcome, and how a finding is a distinct outcome type from a completed action.

00:02:31  recon         ASN AS64500 resolved to two announced prefixes via whois        completed
00:09:48  recon         subdomain enum on acme.com, 38 names, 11 resolving in-scope      completed
00:24:05  discovery     port and service scan, 17 live hosts, 41 services fingerprinted  completed
01:06:52  assessment    exposed management portal on 8443 with default credentials       finding
01:38:20  assessment    testssl audit, legacy TLS still enabled on the mail gateway      finding
02:55:14  assessment    nuclei flagged a known CVE on the app load balancer, confirmed   finding
03:21:40  assessment    single non-destructive proof of concept, read-only response      completed
03:44:09  wrap-up       6 findings recorded, host and service inventory reconciled       completed

Two of those lines tell the whole story of the design. The default-credential finding at 01:06:52 is the classic external result: not a memory-corruption exploit, just a portal that should never have been reachable, carrying credentials it should never have kept. And the 03:21:40 proof of concept is read-only by contract, captured as evidence, and goes no further.

A note on why this is an AI operator, not a scanner

It would have been easier to ship a scheduled scanner with a nice report template. The reason to build an orchestrator-plus-subagents operator instead is that an external engagement is a sequence of decisions, not a fixed pipeline. What you scan depends on what discovery found. Whether a default-credential check is worth running depends on what services fingerprinting turned up. Whether a finding is worth a proof of concept depends on whether the operator can confirm the version and the exposure. Those are judgment calls, made in order, and the phase-gated agent structure is what lets the platform make them while still keeping the whole run inside an auditable, non-destructive boundary.


What this changes for security teams

  1. Perimeter visibility becomes a cadence, not an event. The moment external testing is a scheduled Recon or Assessment engagement rather than an annual consultancy booking, the forgotten-asset problem shrinks. You find the new exposure in the week it appears.
  2. "Is this finding real?" gets a cheap answer. The Exploitation tier turns a suspected exposure into a confirmed, read-only proof without a separate engagement. That single fact changes how findings get prioritized, because confirmed exploitability outranks a scanner's maybe.
  3. Auditability is a first-class output, not an afterthought. A phase-gated timeline with visible skipped checkpoints and reasons is exactly the artifact an auditor or an incident reviewer wants, and it is produced automatically.
  4. Authorization has to be designed in, not bolted on. Offensive automation without an enforced attestation and hard target blocks is a liability. The safer pattern is the one here: whitelist the scope grammar, require a signature, and block dangerous targets outright regardless of input.
  5. The best external findings are still boring, and that is the point. Default credentials, a lapsed certificate, a management portal that leaked to the public internet. Tooling that treats reconnaissance and exposure review as the main event, not as a warm-up to exploitation, is tooling aligned with where perimeter risk actually lives.

How this fits the rest of Pragma Core

Pragma Core's premise is continuous offensive security across three domains in one workspace: application security over your code, supply chain, and web apps; infrastructure security over your internal and external networks; and adversary emulation for breach and attack simulation. External Penetration Testing slots into the infrastructure domain as the outside-in counterpart to the internal network engagement, and it inherits the platform patterns that make the other domains coherent.

One finding lifecycle across every domain

The same login, the same role-based access control, and the same finding lifecycle apply whether the finding came from a SAST scan of your code, an internal network pentest, or an external engagement. An externally exposed service and a hardcoded secret in a repository live in the same triage workflow, which is what lets a team reason about risk across the boundary between "our code" and "our perimeter" instead of in two disconnected tools.

The same Kali worker pool that drives DAST

External engagements run on the same shared hardened Kali worker pool that powers Pragma Core's dynamic application security testing. That shared foundation is why there is nothing to deploy for an external test: the worker infrastructure already exists for DAST, and the external operator drives it over SSH the same way.

Internal and external, two halves of the network picture

Infrastructure Penetration Tests cover the inside: drop an agent on a host you control and an AI operator walks Active Directory recon, BloodHound attack paths, ADCS abuse, and safe credential attacks. External Penetration Testing covers the outside: hand over a scope and the operator maps and assesses your perimeter. Run both and you have the full network picture, the doors an attacker rattles from the street and what they could reach if one of those doors gives way, in one workspace with one finding model.

Feeds naturally into adversary emulation

Once an external engagement has told you what an attacker can reach and confirmed which exposures are real, the natural next question is whether your detections would fire against a real adversary using them. That is the adversary emulation domain: pick a MITRE-aligned adversary like APT29 or FIN7, run it dry to storyboard or live to exercise your SOC, and watch each technique land with action-level evidence. External testing finds the exposure; emulation tests whether you would notice it being used.


Closing thoughts

External Penetration Testing is not, at bottom, a feature about running nmap and nuclei on a schedule. Any team can do that. It is a feature about turning an inherently sequential, judgment-heavy engagement, the kind that has traditionally required a human consultant and a calendar slot, into something you can run against your live perimeter as often as your perimeter changes, while keeping the whole thing inside an auditable, non-destructive, authorization-gated boundary. The hard part was never the scanning. It was making the operator make good decisions in order, prove a finding without breaking anything, and record every step so you can trust the result. That is the part Pragma Core built.

The same perimeter-drift problem lives in almost every organization that ships to the internet, and the gap between "we scanned once this year" and "we know what we expose this week" is the gap most breaches walk through. If your team wants to move from an annual snapshot to a continuous, audited read on your external attack surface, you can define a scope, pick a depth, and run your first engagement at pragma-core.com, or reach out for a demo.


Sources

Related posts
CVE-2026-74820: how an unsanitized ORDER BY clause turned ServiceNow's AI Platform into an unauthenticated database backdoor
Sep 18, 2026
CVE-2026-85978: how one path normalization mismatch turned Akana's admin console into unauthenticated remote code execution
Sep 9, 2026
CVE-2026-78174: how an unredacted session token in a diagnostic log turned a low-privileged WatchGuard Dimension admin into super admin
Sep 1, 2026

Start securing your codebase today

Connect your repositories and let AI agents handle continuous scanning, research, and triage.

Have questions? Get in touch →