Back to blog

CVE-2026-65321: how one backslash in PyAthena's quote escaper turned parameterized DELETE queries into SQL injection

· · 18 min read
CVE-2026-65321: how one backslash in PyAthena's quote escaper turned parameterized DELETE queries into SQL injection

The finding, reported by Rahul Karne and coordinated by VulnCheck, is the kind that makes developers wince, because it breaks a guarantee they were relying on without a second thought: that passing a value through the library's parameter substitution keeps it safely quoted. In a vulnerable configuration an unauthenticated attacker can inject arbitrary SQL through DELETE and CREATE TABLE ... AS SELECT statements that the application built with PyAthena's parameterization API, defeating the exact mechanism that was supposed to prevent injection. The issue is tracked as CVE-2026-65321 and carries a CVSS score of 9.8 (with a CVSS v4 score of 9.3). The striking detail is where the safety came from: the library chose how to escape quotes based on the first word of the query, kept an allowlist of statements it would escape correctly, and defaulted everything else, including the two statement types in this advisory, to the wrong escaper.

This is a two-layered breakdown. The first part is for readers who need to decide quickly whether they are exposed and what to do about it, in plain language and without the internals. The second part reconstructs the flaw the way the maintainer and the reporter saw it: the two escaping conventions PyAthena juggles, the prefix allowlist that routed the wrong statements to the wrong one, why a backslash that looks like an escape is not one, and the small allowlist inversion that fixed it. The article closes by looking at how the Pragma Core platform addresses exactly the kind of problem that made this possible, the kind that lives in a security-relevant branch keyed on a fragile signal, not in any single obviously wrong line.


Part I. Executive breakdown

What happened

PyAthena is the widely used Python client for Amazon Athena, the serverless query service that runs SQL over data in S3. Like most database clients, it offers parameter substitution: you write a query with placeholders, hand it your values separately, and the library is responsible for quoting and escaping those values so they cannot break out of the string they belong in. This is the standard defense against SQL injection, and developers reach for it precisely so they never have to think about escaping by hand.

Here is the flaw, in plain terms. Athena actually speaks two dialects under the hood. Most queries, including reads, inserts, and "create a table from a query," are handled by the Trino engine, which expects a single quote inside a string to be escaped by doubling it, turning ' into ''. A separate set of table-management commands are handled by an older Hive engine, which escapes quotes with a backslash, turning ' into \'. PyAthena decided which convention to use by looking at the first word of the query. It kept a short list of statement types it would escape the Trino way, and sent everything else the Hive way.

The problem is that "everything else" included DELETE and CREATE TABLE ... AS SELECT, both of which are actually run by the Trino engine. So PyAthena escaped their parameters with a backslash. But Trino does not treat a backslash as an escape character inside a string. It reads the backslash as a plain character and then sees the attacker's quote as the real end of the string. The value the developer thought was safely contained has just closed its own quotes, and whatever the attacker put after it is now live SQL. The concrete result is data exfiltration from other tables, mass deletion, and attacker-chosen output destinations, all through a query the developer believed was parameterized and safe.

Who is affected

Component Status
pyathena (pip), versions <= 3.35.3 Vulnerable. Upgrade to 3.35.4.
pyathena 3.35.4 and later Patched. Escaper is now selected by the engine, defaulting to the Trino-safe one.
Apps that only parameterize SELECT / WITH / INSERT / UPDATE / MERGE Not exposed through this specific path; those already used the safe escaper. Still upgrade.
Apps that never pass untrusted values as parameters to DELETE or CTAS Lower risk in practice, but the correct escaper selection still matters. Upgrade.

The exposure concentrates exactly where it hurts: data platforms and ETL code, which are the systems that most often run DELETE to prune partitions and CREATE TABLE ... AS SELECT to materialize results, and which increasingly build those statements from request parameters, job configs, or upstream records. An application that carefully used PyAthena's parameter API on a DELETE, believing that was the safe thing to do, was more exposed than one that sloppily used a SELECT, which is an uncomfortable inversion of the usual advice.

Why this matters beyond PyAthena

The bug class is CWE-89, SQL injection, but the interesting part is the mechanism: an escaping routine that was correct for one parser and wrong for another, selected by a heuristic that did not actually track which parser would run the statement. Any library that speaks to more than one backend dialect, or that supports more than one quoting convention, faces the same trap. The escaping has to be keyed to the authority that will parse the string, and if the selection is a guess based on a prefix, the guess will eventually be wrong for some input.

The second transferable lesson is about defaults. PyAthena kept an allowlist of statements to escape the safe way and defaulted everything else to the other way. When a safety decision falls through to a default, that default has to be the safe one. Here it was the dangerous one, so every statement type the authors did not explicitly enumerate, including ones added to SQL engines over time, silently inherited the injectable path.

Recommended actions

  1. Upgrade to PyAthena 3.35.4 or later. This is the fix and it should be first.
  2. As interim mitigation, do not pass untrusted values as parameters to DELETE or CREATE TABLE ... AS SELECT. Build those statements from trusted input only, or route the untrusted value through a SELECT or INSERT style path, which already used the safe escaper.
  3. Inventory every call site that formats a DELETE or CTAS with the parameter API and trace where those parameter values originate. The vulnerable pattern is untrusted input reaching one of those two statement types.
  4. Apply least privilege to the IAM role your Athena queries run as. The blast radius of this injection is bounded by what that role and workgroup can read, write, and drop, so a narrowly scoped role limits exfiltration and destruction.
  5. Review Athena query history and CloudTrail for anomalous UNION SELECT, unexpected DELETE volume, or CTAS statements writing to unfamiliar locations, over the window your vulnerable version was in use.

Part II. Technical breakdown

Background: two engines, two escaping conventions, one formatter

Athena runs SQL through two different engines depending on the statement. Data queries and data manipulation, SELECT, INSERT, UPDATE, MERGE, DELETE, and CTAS among them, are executed by the Trino engine. A set of Hive-style DDL statements, such as CREATE EXTERNAL TABLE ... LOCATION, ALTER TABLE ... SET LOCATION, MSCK REPAIR TABLE, and SHOW PARTITIONS, are executed by the Hive metastore layer. The two engines do not agree on how to escape a single quote inside a string literal. Trino wants the quote doubled (''). Hive accepts a backslash escape (\').

PyAthena implements parameter substitution in DefaultParameterFormatter inside pyathena/formatter.py. Two helper functions do the actual escaping: _escape_presto, which doubles single quotes and is the Trino-safe path, and _escape_hive, which backslash-escapes single quotes. The formatter's job, for any given statement, is to pick the escaper that matches the engine that will parse the statement. Everything hinges on that choice.

The vulnerability: an escaper chosen by statement prefix

Before 3.35.4, format() chose the escaper with a prefix check that looked, in essence, like this:

# pyathena/formatter.py (pre-3.35.4, simplified)
operation_upper = operation.upper()
if operation_upper.startswith(("SELECT", "WITH", "INSERT", "UPDATE", "MERGE")):
    escaper = _escape_presto      # Trino-safe: '  ->  ''
else:
    escaper = _escape_hive        # Hive-style: '  ->  \'

Read that else branch carefully, because it is the whole bug. The allowlist enumerates five statement types that get the Trino-safe escaper. Every other statement, no matter which engine actually runs it, falls through to _escape_hive. DELETE is not on the list, so it gets the Hive escaper. CREATE TABLE ... AS SELECT begins with CREATE, which is also not on the list, so it too gets the Hive escaper. Both statements are executed by Trino, and both were being escaped for the wrong parser.

The two escapers themselves are each internally consistent. _escape_presto doubles the quote, _escape_hive backslashes it. Neither is buggy in isolation. The defect is entirely in the selection: a security-relevant branch that decides how to neutralize attacker input, keyed on a signal (the leading keyword) that does not reliably indicate which engine will parse the result.

Root cause: a backslash that is not an escape, and a default that fails open

Take the classic probe value a' OR 1=1 -- and run it through _escape_hive, as a pre-patch DELETE would. The escaper turns the single quote into \', and the formatter wraps the whole thing in quotes, producing:

'a\' OR 1=1 --'

On an engine that honors backslash escapes, that \' would be a literal quote and the string would be safe. Trino and Athena do not honor it. To the Trino parser, the backslash is just an ordinary character inside the string, and the very next quote closes the literal. So the parser reads the string as 'a\', value a\, and then treats OR 1=1 -- as trailing SQL. The literal terminated early, exactly one character sooner than the developer intended, and the attacker's payload is now part of the query.

There are two root causes stacked here. The first is the parser mismatch: the escaping convention did not match the engine, so an escape that looked correct was inert. The second is the direction of the default: the allowlist enumerated the safe cases and let everything else fall to the unsafe escaper. Any statement type the authors did not list, including DELETE, CTAS, EXPLAIN, UNLOAD, and CREATE VIEW, all of which Trino runs, inherited the injectable path by omission. The patch notes make this explicit, observing that defaulting to the Trino escaper is the fail-safe direction, because over-doubling a quote in a genuine Hive statement is at worst a harmless formatting quirk, while backslash-escaping a quote in a Trino statement is an injection.

Exploitation: from a parameter to deletion and exfiltration

The impact depends on the statement the application builds and the IAM privileges of the query role, but two paths are direct.

Destructive DELETE. Suppose an application prunes a tenant's rows with a parameterized delete:

cursor.execute(
    "DELETE FROM events WHERE tenant = %(t)s",
    {"t": user_supplied},
)

With user_supplied = "x' OR 1=1 --", the pre-patch formatter renders:

DELETE FROM events WHERE tenant = 'x\' OR 1=1 --'

Trino closes the literal at the attacker's quote, so the predicate becomes tenant = 'x\' OR 1=1, with the rest commented out. OR 1=1 is always true, and the statement deletes every row in the table.

Exfiltration and attacker-controlled output via CTAS. A materialization step is an even richer target, because CTAS both reads and writes:

cursor.execute(
    "CREATE TABLE report WITH (format = 'PARQUET') "
    "AS SELECT id FROM orders WHERE region = %(r)s",
    {"r": user_supplied},
)

An injected UNION SELECT in the WHERE or projection can pull columns from tables the query role can read but the application never intended to expose, and because CTAS controls its own destination, an attacker who can influence the statement can steer results into a location or table of their choosing. The advisory names all three consequences: exfiltration via injected UNION SELECT, execution of destructive statements, and attacker-controlled CTAS destination and content. None of it requires authentication to PyAthena itself; it requires only that untrusted input reach the parameter of a DELETE or CTAS, which is the CVSS AV:N/AC:L/PR:N/UI:N picture behind the 9.8.

How the fix works

Version 3.35.4 inverts the selection logic. Instead of an allowlist of Trino statements with an unsafe default, it keeps an allowlist of genuine Hive DDL statements and defaults everything else to the Trino-safe escaper:

# pyathena/formatter.py (3.35.4)
def _get_escaper(operation: str) -> Callable[[str], str]:
    """Select the escaper matching the engine that will parse the statement."""
    if _HIVE_STATEMENT_PATTERN.match(_strip_leading_comments(operation)):
        return _escape_hive
    return _escape_presto

Three parts make it robust. The _HIVE_STATEMENT_PATTERN regex enumerates the real Hive DDL forms (ALTER/CREATE/DROP of DATABASE/SCHEMA/TABLE, MSCK REPAIR, SHOW, DESCRIBE), and crucially matches CREATE ... TABLE only when it is not followed by AS, using a negative lookahead so that CREATE TABLE ... AS SELECT is correctly recognized as a Trino statement rather than Hive DDL. The _strip_leading_comments helper removes any leading /* ... */ or -- comments before the check, so an attacker cannot dodge statement detection by prefixing a comment like /* generated by etl */ DELETE FROM .... And the default is now _escape_presto, so any statement not explicitly recognized as Hive DDL is escaped the Trino-safe way, which is the fail-safe direction.

The regression tests added with the fix pin the behavior down precisely. A hostile value a' OR 1=1 -- is run through a battery of Trino statements, including DELETE, CTAS, CREATE VIEW, EXPLAIN DELETE, UNLOAD, and comment-prefixed variants, asserting that no backslash escape survives and that quotes are doubled. A parallel set of true Hive DDL statements asserts that backslash escaping is retained. One test spells out the corrected DELETE rendering directly: the value comes back as 'a'' OR 1=1 --', with the quote doubled, so the literal the formatter opened stays open until its own closing quote. Another confirms that a table named as_of is not mistaken for the CTAS AS keyword, a nice guard against the lookahead overreaching.

Affected versions

Timeline

Date Event
2026 (private report) Rahul Karne reports the escaper-selection flaw, coordinated by VulnCheck.
2026-07-31 PyAthena 3.35.4 released and advisory GHSA-xwj5-g6cv-4r5c published by the maintainer.
2026-08-02 CVE-2026-65321 published, with a VulnCheck advisory the same day.

A note on the discovery methodology

This is a source-review find that rewards reading a security control against the thing it is supposed to protect. In isolation, _escape_hive and _escape_presto both look fine, and the prefix check reads like reasonable housekeeping. The bug only becomes visible when you ask which engine actually parses each statement type and then compare that against which escaper the prefix check assigns, at which point the mismatch for DELETE and CTAS jumps out. The tell for a reviewer is structural: a branch that selects how to neutralize untrusted input, decided by a heuristic (the leading keyword) that is not the same thing as the property it needs to track (the executing engine). Wherever a security decision is proxied through a fragile signal like that, it is worth checking every case the signal can misclassify, especially the default branch.


What we should learn from CVE-2026-65321

  1. Escaping must be keyed to the parser that will read the string. A correct escape for one dialect is an inert non-escape for another. When a library talks to more than one engine or quoting convention, the escaper has to track the actual downstream parser, not a proxy for it.
  2. Security defaults must fail safe. An allowlist that enumerates the safe cases and defaults the rest to the dangerous path will silently misclassify every case the authors did not anticipate. Enumerate the dangerous cases and default to the safe branch, so omissions fail closed.
  3. Parameterization is a promise the library must actually keep. Developers use the parameter API specifically so they do not have to think about escaping. A gap in that API is worse than a hand-built query, because it is invisible to the person who did the right thing.
  4. Backslash escaping is a classic false friend. \' looks universally safe but is engine-specific, and several SQL engines, Trino and Athena included, do not honor it inside string literals. Treat quote-doubling as the portable default and backslash escaping as the exception that needs justification.
  5. Watch the near-miss cases around a heuristic. The fix had to handle leading comments, CTAS versus plain CREATE TABLE, and a table named as_of. Whenever statement classification drives a security decision, the edge cases adjacent to the boundary are where the next bug hides.

How Pragma Core addresses this class of problem

Pragma Core is built for exactly this shape of issue: a vulnerability that is invisible when you read one function and only becomes clear when you follow how untrusted input reaches a branch that decides how it will be neutralized. Off-the-shelf scanners tend to flag string concatenation into a query and go quiet when the sink is a parameter API that is supposed to be safe. This bug lived one level up, in the selection of the escaper, which is precisely the kind of state-and-flow reasoning Pragma Core targets. Each capability below is tied back to the specific defect in DefaultParameterFormatter.

SAST tuned for the relevant pattern, not just injection sinks

Off-the-shelf static analysis is tuned for taint that ends in obvious string-built SQL and misses a case where the taint reaches a parameter formatter that then chooses the wrong escaper. Pragma Core's static analysis can be tuned to this pattern: a security-relevant branch that selects an escaping or encoding strategy from a prefix or type heuristic, with a default that falls through to a less strict path. Described that way, the pre-3.35.4 if operation.startswith(...) else _escape_hive is a high-confidence finding rather than a blind spot.

Autonomous AI agents for attack chain investigation

The autonomous agents reason over the whole path rather than stopping at the format call. For a codebase using PyAthena, the agent asks the question that surfaces this bug directly: does any untrusted value reach the parameter of a DELETE or CREATE TABLE ... AS SELECT statement, and does the formatter escape that statement type for the engine that will actually run it. That is a question about flow from an input boundary to an escaper selection, not about a single flagged line, which is where generic tooling stops.

Interactive call graphs with vulnerability overlay

Pragma Core auto-generates call graphs for any connected repository and overlays findings on them. Here the value is making visible the path from a request handler or job config, through the application's query-building helper, into DefaultParameterFormatter.format, and out to the two statement types that were misrouted. The flaw lived in that relationship, between where input enters and where the escaper is chosen, and seeing it drawn makes the risk legible.

Continuous tracking of third-party packages

Pragma Core tracks every third-party package across all connected repositories and surfaces vulnerable versions with their CVSS scores and fixed upgrade paths. A service pinned to pyathena <= 3.35.3 would light up the moment CVE-2026-65321 entered the catalog, with 3.35.4 as the exact target, so the team learns it is exposed from the dependency graph rather than from an incident in their data lake.

Full SBOM per repository

Pragma Core generates a complete component inventory per repository, exportable as CycloneDX JSON. When a data-access library ships an injection CVE, the first question is which services embed an affected version, often buried under a data-pipeline or analytics dependency, and most teams cannot answer it quickly. A current SBOM turns that from a scramble into a query.

Human-guided AppSec investigations

The expert-led research module lets an AppSec operator drive a deeper look at the areas a team already suspects are fragile, backed by the autonomous agents and the workspace context. Auditing every DELETE and CTAS call site across a data platform for untrusted parameters, exactly the sweep this advisory calls for, is the kind of systematic investigation this supports, applied across a portfolio rather than one file at a time.


Closing thoughts

CVE-2026-65321 is not, fundamentally, a bug about a backslash or about one Athena client. It is a bug about neutralizing untrusted input for the wrong authority, and about a safety decision that defaulted in the dangerous direction. Pick the escaper by a fragile proxy for the real parser, let unrecognized cases fall through to the unsafe branch, and the promise of parameterization quietly stops holding for exactly the statements no one thought to test. That pattern is not specific to PyAthena or to SQL. It recurs anywhere one component encodes or escapes data for another and has to guess which ruleset the consumer will apply, across template engines, serializers, shell builders, and multi-dialect database layers.

The difference between treating this write-up as a curiosity and using it as an audit trigger comes down to AppSec maturity and code visibility: whether you can see, across your repositories, where untrusted input reaches a query builder and whether the escaping is keyed to the engine that will parse it. Organizations that want to move from "we scan and report" to "we systematically investigate what is fragile" can reach Pragma Core at pragma-core.com for a demo.


Sources

Related posts
CVE-2026-74820: how an unsanitized ORDER BY clause turned ServiceNow's AI Platform into an unauthenticated database backdoor
Sep 18, 2026
CVE-2026-85978: how one path normalization mismatch turned Akana's admin console into unauthenticated remote code execution
Sep 9, 2026
CVE-2026-78174: how an unredacted session token in a diagnostic log turned a low-privileged WatchGuard Dimension admin into super admin
Sep 1, 2026

Start securing your codebase today

Connect your repositories and let AI agents handle continuous scanning, research, and triage.

Have questions? Get in touch →