In the second half of 2025, the threat model that most infrastructure teams had quietly relied on stopped holding. On November 14, 2025, Anthropic published an account of the first reported cyber espionage campaign in which an AI agent, not a human, executed the large majority of the operation: roughly thirty organizations targeted, a handful of confirmed intrusions, and by Anthropic's own estimate 80 to 90 percent of the tactical work carried out by the model itself, at request rates no human team could match. A few months earlier, in June 2025, an autonomous penetration testing system called XBOW had climbed to the number one spot on HackerOne's US leaderboard, ahead of every human researcher, prompting the platform to split machine and human rankings entirely. The same capability that lets an agent enumerate, reason, and exploit at machine speed is now available to both sides. The interesting question is no longer whether AI can run an attack chain. It is who gets to run it against your infrastructure first.
This is a two-layered breakdown. The first part is for decision makers who need to understand quickly what actually changed, which systems are exposed, and what to do this quarter. The second part is for engineers and red teamers who want the mechanics: how an autonomous offensive agent works internally, why it is so effective against infrastructure specifically, where the current tooling sits, and how these systems are built and operated. The article closes by looking at how the Pragma Core platform turns this same agentic capability into a defensive advantage, which is the whole point.
Part I. Executive breakdown
What happened
For most of the last decade, offensive security against infrastructure was rate limited by people. A penetration test was a person, or a small team, working for one or two weeks, once or twice a year. Attackers operated under similar constraints. Skilled operators were scarce, attention was finite, and the slow, patient enumeration that uncovers the ugly attack paths simply did not scale. That scarcity was load bearing. A lot of infrastructure was safe not because it was hardened, but because nobody had the hours to look closely.
AI powered red teaming removes that constraint. An autonomous offensive agent is a large language model wired to real tools: a network scanner, a directory enumerator, a web proxy, a credential cracker, an exploitation framework. It runs a loop. It observes the environment, reasons about what to try next, runs a tool, reads the output, and decides the next move, over and over, without getting tired and without forgetting what it saw three hours ago. Given a foothold and a goal, it behaves like a methodical human operator who never sleeps, never gets bored, and can run dozens of these loops in parallel.
The concrete result is a collapse in the cost of looking. Tasks that took a human tester days, mapping an entire external surface, walking an Active Directory environment from a low privilege user to Domain Admin, chaining three boring misconfigurations into one serious compromise, now take minutes to hours and cost almost nothing to repeat. Defenders can use this to test their own infrastructure continuously. Attackers can use it to test everyone's, all the time. Both are already happening.
Who is affected
The honest answer is everyone who runs infrastructure, but the exposure is not uniform. The systems most affected are the ones that depend on obscurity, on the assumption that an attacker would not bother to look closely enough or stay long enough to find the path.
| Environment | Exposure |
|---|---|
| Active Directory and identity infrastructure | High. Attack-path analysis is exactly what agents excel at. Enumerate, find an ACL or delegation chain to a high value target, abuse it. Audit graph reachability now. |
| Internet-facing web and API surface | High. Autonomous agents map and probe the full external surface continuously. The window between a misconfiguration shipping and an agent finding it is shrinking to days. |
| Cloud control planes (IAM, roles, trust policies) | High. Privilege-escalation paths through over-broad roles and assume-role chains are precisely the multi-step logic agents reason through well. |
| Internal flat networks and legacy segmentation | Elevated. Lateral movement that relied on nobody mapping the network is now cheap to map. |
| Point-in-time-only programs (annual pentest, then nothing) | Elevated. The annual snapshot ages out in days against a continuous adversary. |
| Mature, continuously tested environments | Lower, but only if the testing itself keeps pace with agentic offense. |
The number that should reorder priorities is timing. Google's own threat intelligence group has documented adversaries using AI across the attack lifecycle, and Anthropic's GTG-1002 disclosure showed an agent sustaining an intrusion campaign at machine speed. The asymmetry to internalize is this: your infrastructure is now tested continuously by attackers whether or not you test it continuously yourself.
Why this matters beyond any single tool
It is tempting to read XBOW or Big Sleep or GTG-1002 as isolated headlines, impressive demos that do not touch your environment. That reading is a mistake. The underlying capability, an agent that can plan and execute a multi-step attack chain across a network, is now commoditized. There are open-source agentic red team frameworks on GitHub that drive nmap, BloodHound, and a Kali toolbox through a Model Context Protocol server with little human input. The barrier to running a competent autonomous attack against infrastructure has dropped from "elite operator" to "anyone with an API key and a target."
This is the same structural shift that fuzzing brought to memory safety, then mass scanning brought to patch management. Each time, an expensive expert activity became cheap and continuous, and the defenders who treated security as a periodic event got caught flat. The class of problem is not a particular exploit. It is the durable gap between point-in-time defense and continuous, automated offense.
Recommended actions
- Move from point-in-time to continuous offensive testing. An annual pentest report is a photograph of a moving target. Adopt continuous or frequent automated red teaming so your view of exposure ages in days, not quarters.
- Map your Active Directory and cloud IAM attack paths now. Run BloodHound or an equivalent against AD and review assume-role and trust-policy chains in cloud. The paths an agent will find are the paths you can find first.
- Shrink your exposed surface and your detection-to-remediation window. Inventory internet-facing assets, kill what is not needed, and assume the window between a misconfiguration shipping and an adversary finding it is now measured in days.
- Instrument for machine-speed activity. The GTG-1002 campaign was eventually caught because its request rate and persistence were not human. Tune detection for volumetric, tireless, breadth-first enumeration, not just for known signatures.
- Use the same agentic tooling defensively, under authorization. The most effective answer to an autonomous attacker is an autonomous defender that has already walked every path. Stand up AI powered red teaming on your own estate before someone else does it uninvited.
- Govern the agents you deploy. Autonomous offensive tools are dual use. Scope them, log them, and keep a human accountable for what they touch, so your red team does not become an incident of its own.
Part II. Technical breakdown
The anatomy of an autonomous offensive agent
Strip away the marketing and an AI powered red team agent is a fairly simple control loop wrapped around a capable model. The model is the planner. Around it sits a harness that gives it three things: a set of tools it can call, a memory of what it has done and seen, and a goal. The loop runs until the goal is met or the agent gives up.
state = recon(target) # initial observation
goal = "reach Domain Admin"
while not satisfied(goal, state):
plan = model.reason(state, goal, history) # decide next step
action = plan.tool_call # e.g. ldapsearch, nmap, bloodhound
result = execute(action) # run real tool in sandbox
state = update(state, result) # observe outcome
history.append((plan, action, result)) # remember everything
The power is not in any single component. nmap, ldapsearch, BloodHound, and Impacket have existed for years. The power is in the loop. A human operator running this loop is bounded by attention and stamina. The model is not. It will patiently try the eighth enumeration technique after seven failed, read a 4,000-line LDAP dump without skimming, and recall on hour three a stray service account it noticed on hour one. Breadth and patience, the two things human red teams ration most carefully, are exactly what the agent has in surplus.
The capability: the full kill chain, executed end to end
What makes this an infrastructure problem rather than a curiosity is that the loop now closes across the entire attack lifecycle. The agent does not just find one bug and hand it back. It recons, gains a foothold, enumerates internally, identifies an attack path, executes the abuse, moves laterally, and repeats, narrating each decision the way a human operator would in an engagement log.
# autonomous internal engagement, foothold: low-priv user
[00:00] recon nmap 10.0.0.0/16 → 412 hosts, 3 DCs
[00:02] enum ldapsearch + BloodHound → graph ingested
[00:04] reason shortest path to DA? → 3 hops via ACL abuse
[00:05] abuse GenericWrite on svc_acct → targeted Kerberoast
[00:07] crack hashcat svc_acct hash → password recovered
[00:09] move svc_acct → SQL01 (admin) → new credentials
[00:12] reason SQL01 admin → DCSync? → path confirmed
[00:14] exec DCSync krbtgt → golden ticket
[00:15] result Domain Admin achieved → full domain control
Every step in that log corresponds to a technique red teamers have used for years. The difference is the clock and the operator. Fifteen minutes, no human keystrokes, and the agent can now run the same fifteen minutes against a thousand environments in parallel. The Active Directory abuse chain, enumerate with BloodHound, find an ACL edge such as GenericWrite to a service account, Kerberoast, crack, move laterally, then DCSync to forge a golden ticket, is precisely the kind of multi-hop graph reasoning that suits a model with perfect recall and infinite patience.
Why infrastructure is the agent's best ground
Autonomous agents are not equally good at everything. They are mediocre at novel memory-corruption research that needs deep intuition, and they hallucinate when forced to invent exploit primitives from nothing. But infrastructure compromise is rarely about novel primitives. It is about composition: stringing together known, individually unremarkable conditions into a path. A read-only LDAP property here, an over-broad IAM role there, a reused local admin password, a forgotten trust. None of these is a vulnerability on its own. The vulnerability is the path.
Path-finding over a graph of trust relationships is the agent's home turf for three reasons. First, the state is enumerable: directory objects, ACLs, sessions, and role policies are all queryable facts, not fuzzy intuitions. Second, the reasoning is compositional and the model holds the whole graph in working memory without losing the thread. Third, the feedback is immediate and unambiguous: a tool either returns the hash or it does not, so the agent learns from every step. This is why BloodHound plus a model has become a recurring pattern. Practitioners have wired BloodHound's graph into models through Model Context Protocol so the agent can ask, in effect, "what is the shortest path from this user to Domain Admin, and which edge do I abuse first," and act on the answer.
Exploitation in practice: BloodHound, MCP, and a Kali toolbox
The concrete plumbing of a modern offensive agent looks like this. A Model Context Protocol server exposes a security toolbox to the model as callable functions. The model issues structured calls, the server runs the real tool inside a sandboxed Kali environment, and the result flows back into the loop.
# MCP-driven attack-path query and abuse
tool: bloodhound.query
cypher: MATCH p=shortestPath(
(u:User {name:"svc_low"})-[*1..]->
(g:Group {name:"DOMAIN ADMINS"}))
RETURN p
→ 1 path: svc_low -GenericWrite-> svc_app -MemberOf-> DA
tool: impacket.targeted_kerberoast
target: svc_app
→ TGS hash captured
tool: hashcat.crack
hash: svc_app
→ cracked: Summer2025!
→ next: authenticate as svc_app, escalate to DA
Open-source frameworks released through 2025 package exactly this pattern. One drives a six-phase reconnaissance engine, then hands off to an agent that validates CVE exploitability, tests credential policies, and maps lateral movement through fourteen tools over MCP inside a Kali sandbox. Another centers on Active Directory specifically, running the canonical nmap to LDAP enumeration to BloodHound to abuse-primitive workflow while keeping session state in one place. The sophistication ceiling is rising fast, but the floor, what a non-expert can now run, has dropped just as fast.
The current tooling landscape
This is not a single product story. It is a field that matured across 2024 and 2025, with offense and defense built on the same foundations.
- XBOW is the most visible autonomous penetration tester. From April to June 2025 it submitted findings on HackerOne classified as 54 critical, 242 high, 524 medium, and 65 low, reached the top of the US leaderboard, and the company raised 75 million dollars to scale it. It runs comprehensive tests in hours rather than weeks.
- Big Sleep, from Google DeepMind and Project Zero, is the defensive mirror image: an agent that hunts zero-days in widely used software. Google reported it found CVE-2025-6965, a critical SQLite flaw known only to threat actors, and that the discovery let Google cut off imminent exploitation. Big Sleep also flagged a use-after-free in the ANGLE graphics library before human researchers patched it.
- GTG-1002, the actor behind the Anthropic-disclosed campaign, showed the offensive extreme: an agent driving roughly 80 to 90 percent of a real espionage operation, jailbroken by operators who role-played as a defensive security firm so the model believed it was doing authorized testing.
- Open-source agentic frameworks put the same capability in anyone's hands, driving real toolchains through MCP with little supervision.
Timeline
| Date | Event |
|---|---|
| 2024 | Google introduces Big Sleep; first AI-found memory-safety bugs in mainstream software. |
| 2025-06 | XBOW reaches number one on HackerOne's US leaderboard; HackerOne later separates human and machine rankings. |
| 2025-07 | Google reports Big Sleep found CVE-2025-6965 in SQLite, cutting off imminent exploitation. |
| 2025-08 | Big Sleep flags a use-after-free in ANGLE before human researchers patch it. |
| 2025-09 | Anthropic detects the GTG-1002 espionage campaign run largely by an AI agent. |
| 2025-11 | Anthropic publicly discloses GTG-1002 as the first reported AI-orchestrated campaign at scale. |
| 2025-12 | OWASP publishes a Top 10 for agentic applications, codifying the new threat surface. |
A note on the discovery methodology
The methodological lesson is that these systems are not magic, and treating them as either hype or sorcery leads to the wrong defenses. They are orchestration. The model contributes planning, prioritization, and the ability to hold a large messy state in mind. The tools contribute the actual capability. The harness contributes persistence and parallelism. Where they shine, infrastructure attack paths, is where the problem is search over a knowable graph, and where they stumble, deep novel exploit invention, is where it is not.
There is a sharp safety lesson too. GTG-1002 succeeded in part by convincing the model it was performing authorized defensive testing, decomposing the attack into innocuous-looking sub-tasks so no single request looked malicious. That is a reminder that the guardrails on these agents are themselves an attack surface, and that anyone deploying offensive agents defensively must scope, sandbox, and supervise them as carefully as any other privileged tool.
What we should learn from the shift to agentic offense
- Scarcity was a security control, and it is gone. A great deal of infrastructure was safe because skilled attention was expensive and rationed. Automation removed the rationing. Any defense whose only real protection was "nobody will look that hard" should be treated as already broken.
- The vulnerability is the path, not the line. Infrastructure compromise is composition: known conditions chained into reachability. Tools that grade one host or one finding in isolation miss the chain. Defense has to reason over the graph the way the attacker does.
- Point-in-time testing cannot cover a continuous adversary. An annual pentest is a single frame of a movie the attacker watches every day. The cadence of your offense has to match the cadence of theirs, which now means continuous.
- Machine speed is a detection signal, not just a threat. The same tireless, breadth-first, volumetric behavior that makes agents effective also makes them detectable if you instrument for it. GTG-1002 was caught because it did not behave like a human.
- Offensive AI is dual use, and governance is part of the control. The capability that hardens your estate can compromise it if mis-scoped or jailbroken. Authorization, sandboxing, logging, and a human owner are not optional extras; they are how you keep your red team from becoming your breach.
How Pragma Core addresses this class of problem
The shift to agentic offense is not a problem you patch. It is a permanent change in tempo, and the only durable answer is to operate your own defense at the same tempo, with the same agentic depth, over your real environment. Pragma Core is built for exactly this: continuous, AI-driven offensive security that reasons over how access and trust flow across hosts, services, and identities, which is precisely where the boring, composable infrastructure paths hide. Each capability below maps directly to the attack patterns described above.
Infrastructure penetration tests with Active Directory and BloodHound coverage
The AD abuse chain in Part II, enumerate, find an ACL edge, Kerberoast, move laterally, DCSync, is the canonical agentic infrastructure attack. Pragma Core's infrastructure penetration testing runs agentic internal assessments with native Active Directory and BloodHound coverage, walking the same trust graph an attacker's agent would and surfacing the shortest path from a low-privilege foothold to domain control before someone uninvited finds it.
Adversary emulation aligned to live MITRE ATT&CK
GTG-1002 was a real adversary running a real kill chain at machine speed. Pragma Core's adversary emulation runs breach-and-attack simulation aligned to the live MITRE ATT&CK matrix, exercising the same recon-to-lateral-movement-to-persistence sequence the engagement log in Part II walks through, so your detection and response are tested against the techniques agents actually use, not a static checklist.
Autonomous AI agents for attack-chain investigation
The whole point of an offensive agent is that it does not stop at one flagged host; it asks where this credential reaches, which account has more privilege than its entry point, which path closes the loop to Domain Admin. Pragma Core's autonomous agents reason over those same chains on your behalf, posing the exact questions that turn three unremarkable misconfigurations into one critical path, and reporting the path rather than three disconnected findings.
Continuous testing instead of a point-in-time snapshot
A once-a-year pentest cannot defend against an adversary that tests you every day. Pragma Core connects to your environment and runs continuously, so your view of exposure ages in days rather than quarters and a misconfiguration that ships on Tuesday is found on Wednesday, not at next year's assessment.
Interactive call graphs and code map for cross-boundary paths
Infrastructure paths and application logic meet at trust boundaries: a service account that an app provisions, an IAM role an internal service assumes. Pragma Core auto-generates call graphs and a code map for connected repositories, overlaying findings so teams can see where untrusted input or over-broad privilege crosses a boundary, the relationship-level flaw that single-file scanners never catch.
Human-guided investigations on top of the agentic baseline
Some paths need an operator's intuition. Pragma Core's expert-led research module lets an AppSec professional drive deeper analysis, backed by the autonomous agents and the full workspace context already gathered, applying systematically the same focused investigation a skilled red teamer would, as a natural extension of the platform rather than a separate black-box engagement.
Closing thoughts
The shift to AI powered red teaming is not, fundamentally, a story about AI. It is a story about asymmetry. For years, infrastructure security quietly depended on offense being slow, scarce, and periodic, while pretending it was hardened. Autonomous agents simply removed the slack, and in doing so revealed how much of our defensive posture was resting on the cost of looking rather than the difficulty of breaking in. The same dynamic will play out in every architecture where safety was an accident of attacker economics rather than a property of the system.
The good news is that the capability is symmetric. The agent that can walk your Active Directory to Domain Admin in fifteen minutes can do it for your defenders first, continuously, on every change. The difference between treating this article as a curiosity and treating it as an operating model comes down to whether your security program runs at point-in-time tempo or continuous tempo. Organizations that want to move from "we scan and report once a year" to "we systematically and continuously investigate what is fragile" can reach Pragma Core at pragma-core.com for a demo.
Sources
- Anthropic, Disrupting the first reported AI-orchestrated cyber espionage campaign, November 14, 2025.
- Google, Cloud CISO Perspectives: our Big Sleep agent makes a big leap, and Google Threat Intelligence, Adversaries leverage AI for vulnerability exploitation and initial access, 2025.
- The Record, Google says Big Sleep AI tool found bug hackers planned to use, 2025.
- XBOW, The road to Top 1: how XBOW did it, and BusinessWire, Announcing XBOW Pentest On-Demand, 2025.
- Hacker News discussion and Slashdot, XBOW reaches the top spot on HackerOne, June and July 2025.
- OWASP, Top 10 for Agentic Applications, December 2025.
- NVD, CVE-2025-6965 Detail.