Back to blog

When the attacker has an AI too: AppSec scanning in an era where exploits arrive before patches

· · 19 min read
When the attacker has an AI too: AppSec scanning in an era where exploits arrive before patches

In April 2026, security researchers at Wiz published their analysis of an unreleased Anthropic model they referred to as Claude Mythos: a system that, running with little human guidance, discovered thousands of zero-day vulnerabilities across major operating systems and browsers and produced working exploits for many of them within hours. A year earlier, in June 2025, an autonomous offensive system called XBOW had quietly climbed to the top of HackerOne's US leaderboard, outscoring every human researcher on the platform. The verdict that ties these two events together is simple and uncomfortable: the cost of finding and weaponizing a software flaw has collapsed, and it has collapsed for the attacker first. The single fact worth sitting with is the one Mandiant put in its M-Trends 2026 report: time-to-exploit has effectively gone negative, with 28.3% of CVEs exploited within 24 hours of disclosure, down from a mean of over 700 days as recently as 2020.

This article is a two-layered breakdown. The first part is for readers who need to decide quickly whether this changes anything about how their organization handles application security, and what specifically they should do about it. The second part dissects the mechanics: what these AI offensive systems actually do, why traditional scanning was already losing the race before AI entered it, and where the new failure modes are. It closes by looking at how the Pragma Core platform addresses exactly the kind of problem this shift creates, the problem of finding what is fragile before an automated adversary does.


Part I. Executive breakdown

What happened

For most of the history of software security, there was a comfortable asymmetry working in the defender's favor: finding a serious vulnerability in real code was slow, expensive, and required rare expertise. A skilled researcher might spend weeks auditing a single component. That friction was, in practice, a security control. It meant that most bugs in most software were never found by anyone, and that the ones that were found arrived slowly enough for a patch cycle to absorb them.

That friction is dissolving. Large language models tuned for vulnerability research can now read code, reason about how data flows through it, propose a hypothesis about where a flaw might be, write a test to confirm it, and in the strongest demonstrations, write a working exploit. They do this at a speed and unit cost that no team of humans can match, and they do it in parallel across thousands of targets. Google's Big Sleep, built by DeepMind and Project Zero, reported 20 real vulnerabilities in open-source software including FFmpeg and ImageMagick in August 2025, each found and reproduced by the AI without human intervention. Anthropic's Mythos, per the Wiz writeup, went further: thousands of findings and exploits generated autonomously.

The concrete result for a defender is this. The window between a vulnerability becoming public and someone, or something, exploiting it has shrunk from years to hours. The skill barrier that used to protect obscure or internal applications is gone, because an AI does not need to be an expert in your specific framework to systematically probe it. The asymmetry has flipped: the attacker now scales cheaply, and the defender, still patching at human speed, does not.

Who is affected

This is not a single product with a version range. It is a shift in the economics of exploitation, so the "affected" question is really about who is exposed to faster, cheaper, more thorough offensive automation.

Who Exposure
Orgs relying on periodic (quarterly/annual) pentests Highly exposed. A point-in-time test cannot keep pace with continuous, automated discovery. The gap between tests is now the attacker's working window.
Teams using only off-the-shelf SAST/DAST in CI Exposed. Traditional scanners catch known sink patterns; AI attackers chain logic and state bugs those tools structurally miss.
Internal / low-profile applications Newly exposed. The "nobody will bother auditing this" assumption no longer holds when auditing is nearly free.
Orgs with strong continuous AppSec and fast patching Reduced exposure, but still racing a shrinking patch window.
Software with public CVEs and slow upgrade cycles Critically exposed. With 28.3% of CVEs hit within 24 hours, a multi-week upgrade cycle is an open door.

The number that should worry a security leader is from Zscaler's ThreatLabz 2026 VPN Risk Report: 79% of security leaders now believe attackers can exploit vulnerabilities faster than their teams can deploy patches. That is not a fear of a future threat. It is a description of the present operating condition.

Why this matters beyond any one vendor

The temptation is to read each AI bug-hunting headline as a story about a single impressive tool. That misses the pattern. What has actually changed is that vulnerability research, historically the most labor-constrained activity in all of security, is becoming an automated, parallel, and cheap process. Every activity downstream of "finding the bug" gets faster as a result: triage, weaponization, and exploitation.

The class of problem this creates is a timing problem. Defensive security was built around the assumption that there is meaningful time between disclosure and exploitation, time to assess, prioritize, test a patch, and roll it out. When that interval approaches zero, every control that depends on it degrades. Periodic scanning, scheduled pentests, monthly patch windows, and manual triage queues were all designed for a slower adversary. They do not fail loudly. They quietly stop being fast enough.

Recommended actions

  1. Move from periodic to continuous. Replace or supplement quarterly pentests and scheduled scans with continuous scanning and continuous testing that runs on every commit and against running systems, so the window between "introduced" and "found by you" stays small.
  2. Shorten the patch path, not just the scan path. Finding faster is useless if remediation still takes weeks. Pre-stage upgrade paths, automate dependency bumps where safe, and treat critical-CVE patching as an incident-grade workflow.
  3. Inventory what you actually run. Maintain a live SBOM per repository so that when a CVE lands, you can answer "do we run this, and which version" in minutes, not days.
  4. Test for logic and chained bugs, not just injection sinks. Adopt analysis that reasons across functions and trust boundaries, because that is the territory AI attackers are best at and classical scanners are weakest at.
  5. Use AI on defense, deliberately. The same capability that finds bugs for attackers finds them for you. Adopt AI-driven research and testing so your discovery speed is in the same order of magnitude as the adversary's.
  6. Tune for signal, not volume. AI-generated reports have already flooded bug-bounty programs with low-quality "AI slop." Whatever you adopt must produce confirmed, reproducible findings, or it just moves the bottleneck to triage.

Part II. Technical breakdown

Background: what AppSec scanning was built to do

To see why the ground has shifted, it helps to be precise about what the established tools do and the assumptions baked into them.

Static application security testing (SAST) parses source code and models how data flows through it. Its core technique is taint analysis: track values that originate from an untrusted source and flag the ones that reach a dangerous sink (a SQL query, a shell command, a file path) without passing through sanitization. This is genuinely useful and catches a large family of injection bugs. Its blind spot is structural: it is tuned for source-to-sink flows with recognizable sinks. Logic flaws, broken authorization, and state mismatches do not have a "sink" to flag. There is no dangerous function call at the end of "this code lets a user read another tenant's data because two functions disagree about who the current user is."

Dynamic application security testing (DAST) probes a running application from the outside, sending crafted inputs and observing responses. It finds things SAST cannot, but it is bounded by what it knows to send and what it can recognize as a finding. Software composition analysis (SCA) compares your dependency tree against databases of known-vulnerable versions. It is essential and largely mechanical: it tells you that a package you use has a published CVE. It says nothing about bugs that are not yet published, which is precisely the category AI research is now mass-producing.

Each of these tools encodes the slow-adversary assumption. SCA assumes the relevant bug is already a known CVE. Periodic SAST/DAST runs assume that finding a bug a few weeks after it is written is fast enough. None of them were designed for an opponent that can discover an unpublished, multi-step logic flaw in your specific codebase in an afternoon.

What the AI offensive systems actually do

The recent demonstrations are not magic, and understanding their method is the key to defending against it. An AI vulnerability research agent typically operates in a loop that mirrors how a strong human researcher works, but without the human's time and attention limits:

1. Ingest target        (source, binary, or live endpoint)
2. Build a mental model  (call graph, data flow, trust boundaries)
3. Hypothesize a flaw    ("this handler trusts a value set elsewhere")
4. Write a test/PoC      (input that should violate the invariant)
5. Execute and observe   (did the invariant break?)
6. If yes: refine into a working exploit; if no: revise hypothesis
7. Repeat, in parallel, across thousands of hypotheses

The two stages that used to be the expensive human bottleneck, step 3 (forming a good hypothesis about where a bug hides) and step 6 (turning a crash or anomaly into a reliable exploit), are exactly what frontier models have become competent at. Google's Big Sleep produced findings in which "each vulnerability was found and reproduced by the AI agent without human intervention," with a human only reviewing the final report for quality. Anthropic's Mythos, per the Wiz analysis, reportedly performs patch-diffing (comparing a patched binary against the unpatched one to locate the fix and therefore the bug) and chains multiple vulnerabilities into a single exploit, autonomously.

The strategically important capability is parallelism. A human researcher pursues one or two hypotheses at a time. An agent fleet pursues thousands. This is why the relevant metric is no longer "can an AI find a bug a human could have found" (it can) but "how many targets can be exhaustively probed per dollar" (a lot, and the cost keeps falling).

The root cause: a speed asymmetry, not a new bug class

It is worth being clear about what has and has not changed. The bugs themselves are not new. Big Sleep's FFmpeg and ImageMagick findings are the same families of memory-safety and parsing bugs researchers have hunted for decades. Mythos chains the same logic and memory primitives. AI did not invent a new vulnerability class. What it changed is the rate.

The root cause of the defender's problem is therefore an asymmetry in the rate of two competing processes. Call them discovery (the attacker finding an exploitable flaw) and remediation (the defender finding and fixing it first). For decades, remediation could afford to be slow because discovery was slow. The patch window, the time between a flaw becoming exploitable and an attacker actually exploiting it, was wide enough to absorb human-speed defense.

AI compresses the discovery process by orders of magnitude while leaving most organizations' remediation process untouched. Mandiant's data captures the result precisely: time-to-exploit fell from over 700 days in 2020 to 44 days in 2025, and into the negative in 2026, where exploits routinely arrive before patches are even available. The patch window did not shrink. It inverted.

Exploitation in practice: the inverted patch window

Concretely, here is the sequence that now plays out around a disclosed vulnerability:

T+0h   CVE published with advisory + affected versions
T+1h   Attacker AI ingests advisory, patch-diffs the fix
T+3h   Working exploit generated and validated
T+6h   Mass scanning of internet-facing instances begins
...
T+?d   Defender's next scheduled scan / patch window

The defender's controls fire on a calendar. The attacker's fire on an event. With 28.3% of CVEs exploited within 24 hours, any control whose cadence is measured in weeks is, for that fraction of vulnerabilities, decorative. This is the operational meaning of "exploits arrive before patches": for an unpublished (zero-day) flaw found by an AI, there is no patch to race at all, only the question of whether your side found it first.

There is also a defensive failure mode that is not about speed. The same generative capability that produces exploits produces noise. Bug-bounty programs in 2025 and 2026 have been flooded with AI-generated reports of low quality, what practitioners bluntly call "AI slop": plausible-sounding, unconfirmed, and often wrong. The lesson for defensive tooling is sharp. Volume is not value. A scanner that emits ten thousand unconfirmed findings has not helped you; it has relocated the bottleneck from discovery to triage and exhausted the humans who must clear the queue.

Affected practices and their failure modes

Practice Designed assumption Failure mode under AI-speed adversaries
Annual / quarterly pentest A point-in-time snapshot is representative Stale within days; misses everything introduced between tests
Scheduled SAST/DAST runs Finding bugs weeks after commit is fast enough Discovery cadence slower than attacker's
SCA against known CVEs The dangerous bugs are already published Blind to AI-found, unpublished zero-days
Manual triage of all findings Finding volume is human-manageable Drowned by AI-generated noise / slop
Monthly patch windows The patch window is wide Window inverted; 28.3% of CVEs hit within 24h

Timeline

Date Event
2020 Mandiant mean time-to-exploit measured at over 700 days
June 2025 XBOW tops HackerOne's US leaderboard, outscoring all human researchers
August 2025 Google Big Sleep reports 20 AI-found vulnerabilities (FFmpeg, ImageMagick, others)
2025 Mandiant time-to-exploit measured at 44 days
April 2026 Wiz publishes analysis of Anthropic's Claude Mythos: thousands of autonomous zero-day findings
2026 Mandiant M-Trends: time-to-exploit negative; 28.3% of CVEs exploited within 24h
2026 Zscaler: 79% of security leaders believe attackers out-pace their patching

A note on the discovery methodology

What unites Big Sleep, XBOW, and Mythos is not a single algorithm but a methodology: give a capable model the ability to read a target, form hypotheses, run experiments, and iterate, then run that loop cheaply and in parallel. The meta-lesson for the AppSec community is double-edged. The same loop is available to defenders, and it is far more valuable on the defensive side because defenders have something attackers do not: full source access, build context, and the authority to fix what is found. An AI that can patch-diff a binary from the outside is impressive; an AI that can read your source, map your call graph, and reason about your trust boundaries from the inside is strictly more powerful. The organizations that come out ahead will be the ones that point this capability at their own code first.


What we should learn from the AI-offense era

  1. Friction was a security control, and it is gone. The cost of finding a bug was an invisible defense. Now that it has collapsed, security posture has to be earned explicitly rather than inherited from the difficulty of the attacker's job.
  2. Cadence beats coverage. A perfect scan run once a quarter is worse than a good scan run continuously, because the adversary's clock runs on events, not calendars. Optimize for how fast you find, not just how much.
  3. The bugs that matter most have no sink. AI attackers excel at logic, authorization, and state-mismatch flaws that span multiple functions and services. Tooling that only chases injection sinks is fighting the last war.
  4. Volume without confirmation is a liability. Whether the findings come from your own AI or land in your bug-bounty inbox, unconfirmed output just moves the bottleneck. Insist on reproduced, confirmed findings.
  5. Defenders have the home-field advantage if they use it. Source access, build context, and the right to remediate make AI far more effective on defense. The losing move is to leave that advantage on the table and let the attacker be the only one with an AI.

How Pragma Core addresses this class of problem

The shift described here is not a bug you can patch. It is a change in tempo, and it punishes exactly the tools that were built for a slower adversary: the once-a-quarter pentest, the periodic scan, the SCA database that only knows about bugs after they are published. Pragma Core is a continuous, AI-driven application security platform built by zer0day Technologies and Expertware for precisely this environment. Its design premise is that defenders should be running the same kind of autonomous, parallel, source-aware research that the attacker now has, except with full code access and the authority to fix what is found. Below are the parts of the platform that map directly onto the failure modes above.

Continuous scanning instead of point-in-time tests

The core failure mode of the AI-offense era is cadence: scheduled tests leave windows, and windows are where automated adversaries live. Pragma Core connects to a team's repositories (GitHub, GitLab, Azure DevOps) and runs continuous scanning, research, and testing, contextualized to repository, branch, and commit. Discovery happens as code changes, not on a quarterly calendar, which is the only cadence that competes with an opponent whose clock runs on commits and disclosures.

Autonomous AI agents for attack-chain investigation

The bugs AI attackers find best are the chained, multi-step ones: a value trusted in one handler because it was set in another, a privilege boundary that holds in isolation but not in composition. Pragma Core's autonomous agents reason over attack chains rather than stopping at a single flagged line, asking the questions a human researcher would ask: where else does this value flow, which consumer runs with higher privilege than the entry point, under which inputs is the security check actually reached. This is the same hypothesize-test-refine loop the offensive systems run, pointed at your own code, with your source as ground truth.

SAST tuned for logic and state, not just injection sinks

Off-the-shelf static analysis is tuned for taint flows ending in SQL or shell sinks and structurally misses authorization and state-mismatch bugs, the category with no sink to flag. Pragma Core's AI-powered static analysis is built to surface exactly those patterns (a check that is skipped on one path, a last-write-wins assumption, two functions disagreeing about the current principal) as high-confidence findings, closing the gap that traditional SAST leaves open for AI attackers to walk through.

Interactive call graphs with vulnerability overlay

When the flaw is in the relationship between functions rather than in any single one, you have to see the relationship to find it. Pragma Core auto-generates call graphs for any connected repository, visualizing how classes, functions, and calls connect, and overlays findings so teams can see where untrusted input propagates and where it crosses a trust boundary. This is the inside view an external attacker's AI can only approximate, and it is where chained bugs become visible.

Continuous dependency tracking and full SBOM

With 28.3% of CVEs exploited within 24 hours, the question "do we run this package, and which version" has to be answerable in minutes. Pragma Core tracks every third-party package across all connected repositories, surfacing vulnerable versions with CVSS scores and fixed upgrade paths, and generates a complete per-repository SBOM exportable as CycloneDX JSON. When a CVE lands in the catalog, the affected repositories light up immediately rather than after a manual hunt, which is the difference between racing the inverted patch window and losing it by default.

Confirmed findings, not AI slop

The defensive counterpart to the bug-bounty "AI slop" problem is a scanner that floods you with unconfirmed noise. Pragma Core's DAST drives dynamic testing of live applications and reports confirmed findings, and its research agents are built to reproduce before they report. The goal is to keep the bottleneck at remediation, where it belongs, rather than relocating it to a triage queue no human team can clear.

Human-guided investigations on top of the autonomous layer

Some of what is fragile in a system is known only to the people who built it. Pragma Core's expert-led research module lets an AppSec operator drive deeper analysis, supported by the autonomous agents and the platform context already in the workspace, focused on the parts of the system the team itself suspects. It is the systematic, repeatable version of the one-off researcher investigation, available as part of the subscription rather than as a separate black-box engagement.


Closing thoughts

The AI-offense era is not, fundamentally, a story about clever new tools. It is a story about the disappearance of friction. For decades, the difficulty and cost of finding a real vulnerability was doing quiet defensive work for everyone, and the entire defensive playbook (periodic scans, scheduled pentests, monthly patch windows, manual triage) was tuned to an adversary slowed by that friction. Remove the friction, and the playbook does not break dramatically. It just stops being fast enough, silently, one inverted patch window at a time.

The same patterns live in nearly every modern architecture: code that trusts a value because some other function set it, a dependency buried three layers deep in a container, an internal app no one thought was worth attacking. What separates an organization that treats this article as a curiosity from one that treats it as an audit trigger is AppSec maturity and code visibility: whether you can see how input and state flow across your own functions, services, and trust boundaries, and whether you are finding what is fragile continuously rather than on a calendar. The decisive question of this era is not whether the attacker has an AI. They do. It is whether you have pointed one at your own code first.

Organizations that want to move from "we scan and report" to "we systematically investigate what is fragile" can reach Pragma Core at pragma-core.com for a demo.


Sources

Related posts
CVE-2026-74820: how an unsanitized ORDER BY clause turned ServiceNow's AI Platform into an unauthenticated database backdoor
Sep 18, 2026
CVE-2026-85978: how one path normalization mismatch turned Akana's admin console into unauthenticated remote code execution
Sep 9, 2026
CVE-2026-78174: how an unredacted session token in a diagnostic log turned a low-privileged WatchGuard Dimension admin into super admin
Sep 1, 2026

Start securing your codebase today

Connect your repositories and let AI agents handle continuous scanning, research, and triage.

Have questions? Get in touch →