Back to blog

CVE-2026-43499: how a 15-year-old cleanup shortcut in the Linux kernel's futex code became a one-user path to root and a container escape

· · 22 min read
CVE-2026-43499: how a 15-year-old cleanup shortcut in the Linux kernel's futex code became a one-user path to root and a container escape

On July 7, 2026, Nebula Security published the technical writeup for GhostLock, a Linux kernel vulnerability their VEGA research system had reported to the kernel maintainers back in April. The verdict is blunt: any local user with a shell on an unpatched machine can walk a chain of ordinary threading syscalls into full root, and from inside a container out onto the host. The flaw is tracked as CVE-2026-43499 and carries a CVSS 3.1 base score of 7.8 (High), with a CVSS 4.0 score of 8.5. The detail that should stop you mid-scroll: the bug was introduced in Linux 2.6.39 in April 2011 and shipped, untouched, in essentially every mainstream distribution for more than fifteen years. Nebula turned it into a privilege escalation that was 97% reliable in their testing, and Google paid out $92,337 for it through the kernelCTF program.

This is a two-layered breakdown. The first part is for people who need to decide, quickly, whether they have anything to do here and what exactly that is: what GhostLock is, who it touches, and what to do this week. The second part is for engineers and researchers who want the mechanism, and it follows the chain the way Nebula reconstructed it, from a single mislabeled pointer to a working root shell. At the end we look at how the Pragma Core platform addresses exactly the class of problem that let a bug like this hide in plain sight for fifteen years.


Part I. Executive breakdown

What happened

Deep inside the Linux kernel there is a small piece of machinery for handling lock priorities. When a high-priority task is stuck waiting on a lock held by a low-priority task, the kernel temporarily boosts the low-priority task so the important work is not blocked forever. This is called priority inheritance, and it runs through a subsystem called rtmutex. It is old, it is heavily used, and almost nobody reads it.

GhostLock is a mistake in how that subsystem cleans up after itself. A helper function was written years ago for one specific situation: a thread that put itself to sleep waiting on a lock, then had to tidy up its own bookkeeping. The helper always assumed the thread running the cleanup was the same thread that had gone to sleep. That assumption held for a decade and a half, until a second, newer code path started calling the same helper on behalf of a different, still-sleeping thread. Now the helper tidies up the wrong thread's records and leaves the sleeping thread holding a pointer to a chunk of kernel memory that has already been handed back and reused. That stale pointer is the whole ballgame.

The concrete result: an unprivileged local user, with nothing more than a normal login, can steer the kernel into following that stale pointer, spray their own data into the freed memory, and from there rewrite a table of function pointers the kernel trusts. That gives them code execution as the kernel itself, which is to say root. Because the same kernel is shared between a container and its host, the same trick also carries an attacker out of a container and onto the machine underneath it.

Who is affected

Component Status
Linux kernel 2.6.39-rc1 through 7.1-rc1 with CONFIG_FUTEX_PI=y Vulnerable. Upgrade to a patched LTS kernel now.
Linux kernel 7.1 and later (fix commit 3bfdc63936dd) Patched
Kernels built without CONFIG_FUTEX_PI Not affected (rare in mainstream distros)
Debian, Ubuntu, Red Hat, SUSE stable and LTSS kernels Vulnerable until the vendor backport is applied. Distro patches shipped from May 2026 onward.
Container hosts and multi-tenant Kubernetes nodes Vulnerable if the host kernel is unpatched. The escape crosses the container boundary.

The single build-time condition is CONFIG_FUTEX_PI, the priority-inheritance futex option, and it is enabled in the default configuration of virtually every mainstream distribution. There is no capability requirement, no user namespace requirement, and no network access needed to trigger the bug. Working exploit code is public, published alongside the writeup, so the barrier to weaponizing it has already fallen. No in-the-wild exploitation has been reported yet, but "yet" is doing a lot of work in that sentence once a reliable public PoC exists.

Why this matters beyond one kernel bug

Local privilege escalation is often dismissed as second-tier: the attacker already needs a foothold, so how bad can it be? That framing is wrong for modern infrastructure. In a world of containers, shared CI runners, multi-tenant clusters, and web apps that can be pushed into running attacker-controlled code, "local access" is not a high bar. It is the routine outcome of a phishing click, a leaked credential, a supply-chain package, or a sandbox escape in some other component. A dependable LPE is the second stage that turns any of those into total compromise.

The container escape is what elevates GhostLock from serious to strategic. Container isolation is the load-bearing assumption behind cloud multi-tenancy: your workload and someone else's run on the same host, kept apart by the kernel. A bug that lets code climb out of the container and onto the host does not just compromise one tenant, it undermines the boundary the whole platform is sold on. This is the same class of risk as the vsock and POSIX-timer kernel flaws that made multi-tenant hosts a priority target through 2025, and it belongs in the same triage bucket.

Recommended actions

  1. Upgrade to a patched kernel. Move to the latest LTS release for your distribution, or apply the vendor backport of commit 3bfdc63936dd ("rtmutex: Use waiter::task instead of current in remove_waiter()"). Patches have been shipping from Debian, Ubuntu, Red Hat, and SUSE since early May 2026.
  2. Prioritize container hosts and multi-tenant nodes. Any machine where untrusted or semi-trusted workloads share a kernel is your highest-value target for this patch, because the escape defeats the isolation those machines depend on.
  3. Inventory what you actually run. Enumerate every Linux kernel version across servers, VMs, container hosts, embedded devices, and appliances. Legacy and appliance systems are where a fifteen-year-old bug survives longest, precisely because they fall outside normal update cycles.
  4. Turn on defense-in-depth where patching lags. RANDOMIZE_KSTACK_OFFSET makes the memory-reuse step a roughly one-in-thirty-two gamble, and STATIC_USERMODE_HELPER blocks the specific final trick shown in the PoC. Neither is a complete fix, but both raise the cost while you roll out kernels.
  5. Watch for the behavior, not just the CVE. Add detection for unexpected privilege transitions and container-to-host activity, so an attempt shows up even if you cannot patch a given host immediately.
  6. Assign ownership. Make sure someone is explicitly responsible for the systems that never show up in the standard patch report. Those are the ones this bug will outlive.

Part II. Technical breakdown

Background: priority inheritance, requeue-PI, and the waiter object

To see GhostLock you need three pieces of the kernel's locking machinery.

The first is the priority-inheritance futex. A PI futex is a userspace lock that the kernel backs with an internal real-time mutex (rtmutex) so that priority inheritance works correctly across the userspace and kernel boundary. When a thread blocks on a PI futex, the kernel builds an rt_mutex_waiter object to represent that thread's place in the queue and its relationship to the lock owner.

The second is where that waiter object lives. For a thread blocking on its own behalf, the rt_mutex_waiter is allocated on that thread's own kernel stack, inside the stack frame of the syscall it is sleeping in. It stays valid exactly as long as the thread is parked in that syscall. The moment the thread returns to userspace, the stack frame is popped and that memory is free to be reused by the next syscall.

The third is requeue-PI, the awkward case that breaks everything. FUTEX_WAIT_REQUEUE_PI lets a thread block on one futex with the understanding that it may be moved onto a second PI futex by another thread. The move is performed with FUTEX_CMP_REQUEUE_PI, and internally it proxies the sleeping thread's waiter onto the target lock through rt_mutex_start_proxy_lock(). The critical property of this path is that the thread running the requeue is not the thread that owns the waiter. One thread is acting on another sleeping thread's behalf. Hold onto that, because it is the entire bug.

The invariant that rtmutex quietly relies on is: whoever is cleaning up a waiter is the task that owns it. Requeue-PI violates that invariant, and a helper that was never updated to notice is where GhostLock lives.

The vulnerability: remove_waiter() scrubs the wrong task

The flawed function is remove_waiter() in kernel/locking/rtmutex.c. It is called to roll a waiter back off a lock, and on the normal self-blocking slow path it correctly assumes that current, the running task, is the waiter's owner. On the proxy rollback path it is called on behalf of a sleeping task, and it never learned the difference.

Here is the caller. When the proxy attempt fails with a deadlock, it rolls back through the same buggy helper:

int __sched rt_mutex_start_proxy_lock(struct rt_mutex_base *lock,
                                      struct rt_mutex_waiter *waiter,
                                      struct task_struct *task)
{
  int ret;
  raw_spin_lock_irq(&lock->wait_lock);
  ret = __rt_mutex_start_proxy_lock(lock, waiter, task);
  if (unlikely(ret))
    remove_waiter(lock, waiter);          // ret == -EDEADLK
  raw_spin_unlock_irq(&lock->wait_lock);
  return ret;
}

And here is the helper doing the damage. The comment marks the exact line:

static void __sched remove_waiter(struct rt_mutex_base *lock,
                                  struct rt_mutex_waiter *waiter)
{
  ...
  raw_spin_lock(&current->pi_lock);
  rt_mutex_dequeue(lock, waiter);
  current->pi_blocked_on = NULL;            // should be waiter->task
  raw_spin_unlock(&current->pi_lock);
  ...
}

The waiter argument points at the object living on the sleeping thread's stack. But current here is the requeuer, a completely different task. So the function takes the requeuer's pi_lock, not the waiter task's, and clears the requeuer's pi_blocked_on, not the waiter task's. Three things go wrong at once: the red-black tree dequeue happens without the waiter task's pi_lock held, the waiter task's pi_blocked_on is never cleared, and a later priority-chain walk operates on the wrong top-priority task. The one that matters for exploitation is the second: the sleeping task keeps a pi_blocked_on pointing straight at its own stack frame, and that frame is freed the instant it returns to userspace. Any subsequent PI chain walk through that task follows a dangling pointer into freed kernel stack. It is a stack use-after-free.

A quiet aggravating factor is that this slips past lockdep, the kernel's runtime lock validator. Lockdep checks that a pi_lock is held, but not whose pi_lock it is, so the wrong-task locking raised no alarm for fifteen years.

Root cause: a helper reused by a caller it was never written for

Strip away the specifics and GhostLock has a shape you have seen before. A function was written for exactly one scenario and encoded that scenario as an unstated assumption. Years later a new caller reused the function in a scenario the assumption does not cover, and nothing in the type system, the review process, or the runtime checks flagged the mismatch.

remove_waiter() assumed current == the waiter's owning task. That was true for every caller that existed when it was written. Requeue-PI introduced a caller for which it is false, and the fix is simply to stop assuming: derive the task from waiter->task instead of from current. The bug is not a subtle race in the exotic sense, and it is not a memory-safety slip in a single line of arithmetic. It is a lifecycle and ownership error that lives in the relationship between two functions, one of which changed the rules the other depended on.

Triggering the dangling pointer

To reach the buggy -EDEADLK rollback you need a priority-inheritance cycle built from three futex words and three threads:

The sequence is:

  1. The waiter takes f_pi_chain, then blocks in FUTEX_WAIT_REQUEUE_PI(f_wait -> f_pi_target). Its rt_mutex_waiter is now sitting on its own stack.
  2. The owner takes f_pi_target, then blocks on f_pi_chain, which the waiter is holding.
  3. The main thread calls FUTEX_CMP_REQUEUE_PI(f_wait -> f_pi_target).

The requeue tries to proxy the waiter onto f_pi_target, but the owner of f_pi_target is already blocked behind the waiter through f_pi_chain. The chain walk closes the loop waiter -> f_pi_target -> owner -> f_pi_chain -> waiter, returns -EDEADLK, and takes the buggy rollback. The waiter wakes up with a dangling pi_blocked_on.

There is a pleasant property here for the attacker and an unpleasant one for defenders: once the cycle is staged, the only ordering that matters happens on its own, and after it resolves there is no time pressure at all. The waiter sits in userspace with a live dangling pointer, and the follow-up sched_setattr() that walks the chain can be fired whenever the attacker likes. Nebula note that although they describe it with three threads for clarity, a single CPU core is enough to win the race. The use-after-free window is not a narrow one to be threaded, it is wide open.

From dangling pointer to one controlled write

The freed object is the waiter's own stack rt_mutex_waiter:

struct rt_mutex_waiter {
  struct rt_waiter_node tree;     // rb node, lives in lock->waiters
  struct rt_waiter_node pi_tree;
  struct task_struct *task;
  struct rt_mutex_base *lock;
  unsigned int wake_state;
  struct ww_acquire_ctx *ww_ctx;
};

To reclaim that exact frame, the waiter thread returns from the futex syscall and immediately calls prctl(PR_SET_MM, PR_SET_MM_MAP, ...). Inside, prctl_set_mm_map() copies a user-supplied auxv into a fixed-size stack buffer that lands at roughly the same stack depth as the freed waiter. That gives the attacker a large, naturally aligned, namespace-free block of controlled 8-byte values laid directly over the old object. The auxv is backed by a memfd positioned so the copy straddles a page boundary, and a sibling thread races fallocate(PUNCH_HOLE) on the trailing page during the prctl to stretch the copy_from_user window. prctl is just convenient here; clone, setsockopt, pselect, keyctl, and other syscalls with large controlled stack locals work the same way.

The forged waiter is shaped so a consumer thread firing sched_setattr() walks the PI chain and does exactly one useful thing. The chain walk calls rt_mutex_dequeue(), which is a red-black tree erase, and erasing a single-child root writes that child pointer into the root slot. By pointing the fake lock at target - 8, the attacker lines the rt_mutex_base fields up over the memory around a chosen target pointer:

target - 8  ->  raw_spinlock_t wait_lock        // must read as "unlocked"
target      ->  waiters.rb_root.rb_node          // this slot gets written
target + 8  ->  waiters.rb_leftmost
target + 16 ->  owner

The result is a single constrained store: *(uint64_t *)target = W0_BASE. The constraints are strict. The qword before the target must read as an unlocked spinlock or the trylock fails silently, and the qwords after it must not steer the walk into an uncontrolled waiter or owner, or the kernel faults and panics. So this is not an arbitrary write, it is one carefully aimed pointer write into a location whose neighbors already satisfy the layout.

Exploitation: a function table, a loopback packet, and DirtyMode

The rest is a tidy sequence of using that one write well. Nebula's full chain looks like this:

On Google's LTS 6.12.80 target this whole chain wins the flag in about five seconds. The container escape falls out of the same primitive, because the kernel being corrupted is the one host and container share.

Affected versions

One wrinkle worth flagging for anyone tracking the fix: the upstream v1 patch introduced a separate null-pointer dereference corner case, where a non-top requeue waiter that already owns the target PI futex hits -EDEADLK before waiter->task is set, and the new code then dereferences a NULL waiter->task. That required follow-up hardening, so make sure your backport includes the corner-case fix, not just the original patch.

Timeline

Date Event
2011-04 Bug introduced with the rtmutex rework (commit 8161239a8bcc)
2026-04-18 Nebula's VEGA reports the bug and sends a draft patch to [email protected]
2026-04-20 Fixed upstream with commit 3bfdc63936dd
2026-05-04 Fix v1 backported to stable trees
2026-06-30 Google acknowledges the kernelCTF submission and awards $92,337
2026-07-07 Nebula publishes the technical writeup and open-source PoC

A note on the discovery methodology

GhostLock was found by VEGA, Nebula's automated vulnerability-research system, and it did not surface alone. In the same period, researchers disclosed Bad Epoll (CVE-2026-46242), a close cousin that also turns an unprivileged user into root and, unusually, works on Android, and Copy Fail (CVE-2026-31431), a cryptographic-template flaw already on CISA's Known Exploited Vulnerabilities list. The common thread is not a single subsystem, it is a category: old, heavily used kernel machinery that few humans had reread in years, now being combed systematically by automated tooling.

There is a meta-lesson here that is easy to over- or under-state, so let me be careful. Automated and AI-assisted analysis did not invent a new class of bug. remove_waiter() was reachable and wrong the entire time; anyone with the patience to trace requeue-PI through rtmutex could have found it. What changed is the economics of that patience. Tools that can walk call graphs, reason about which caller violates which assumption, and do it across millions of lines without getting bored have made fifteen-year-old lifecycle bugs findable at scale. The volume of high-quality kernel reports has climbed sharply for exactly this reason. Defenders who are still relying on "nobody has looked at that code in a decade" as an implicit control should assume attackers are now looking at all of it.


What we should learn from CVE-2026-43499

  1. A helper that encodes an unstated assumption is a latent bug waiting for a second caller. remove_waiter() was correct for every caller that existed when it was written, and wrong the moment requeue-PI reused it. Audit shared helpers for assumptions about identity and ownership ("this is always current", "the caller always holds this lock") and make those assumptions explicit parameters, not folklore.
  2. current is not a safe stand-in for "the object's owner". A large family of kernel bugs comes from code that reaches for current when it should have used the task attached to the object it is operating on. Any time a function operates on behalf of another task, current is a red flag.
  3. Runtime validators check shape, not intent. Lockdep confirmed a pi_lock was held and stayed silent, because it does not verify whose lock it is. Controls that check structural correctness will happily wave through a semantically wrong operation. Do not mistake a green lockdep run for proof of correctness.
  4. Local privilege escalation is a container escape in disguise. On shared-kernel infrastructure, "the attacker needs local access" and "the attacker can break tenant isolation" are the same sentence. Score and prioritize kernel LPEs on multi-tenant hosts as boundary-crossing bugs, not second-tier ones.
  5. Age is exposure, not safety. A bug that has survived fifteen years is not battle-tested, it is unexamined. The oldest, most trusted, least-touched code is now the most attractive hunting ground for systematic tooling, which inverts the old intuition that mature code is safe code.

How Pragma Core addresses this class of problem

GhostLock is not the kind of bug a single-function scanner is built to catch. There is no tainted input reaching a dangerous sink on one line. The flaw lives in the relationship between two functions, rt_mutex_start_proxy_lock() and remove_waiter(), where one caller silently violates an assumption the other depends on. That is precisely the category Pragma Core is designed for: issues about how state and ownership flow across functions and trust boundaries, not a single bad statement. Here is how the platform maps onto a problem like this.

Autonomous AI agents that reason about ownership across callers

Pragma Core's autonomous agents reason over chains rather than stopping at a flagged line. Pointed at rtmutex, the question that unlocks GhostLock is a concrete one: does remove_waiter() operate on current, and is there any caller for which current is not the task that owns waiter? Following rt_mutex_start_proxy_lock() back to futex_requeue() answers that in the affirmative. This is exactly the "which caller breaks the assumption" reasoning VEGA applied, made a routine, repeatable check rather than a once-a-decade stroke of attention.

Interactive call graphs that expose the dangerous relationship

The bug was invisible in either function alone and obvious in the edge between them. Pragma Core auto-generates call graphs for a connected repository and overlays findings on them, so the path from a proxy-lock caller into a cleanup helper that assumes current is highlighted as a relationship worth interrogating. When the flaw is in the wiring rather than the components, seeing the wiring is the whole battle.

SAST tuned for lifecycle and ownership patterns, not just injection sinks

Off-the-shelf static analysis is tuned for taint flows that end in SQL or shell execution and is effectively blind to "a helper clears state on current when it should clear it on waiter->task". Pragma Core's static analysis can be tuned to the pattern itself: a function that mutates per-task state via current while receiving an object that carries its own owning task. Framed that way, remove_waiter() is a high-confidence finding rather than a needle no rule was looking for.

Continuous tracking so the fix, and its follow-up, are not missed

The GhostLock story has a tail: the first upstream patch opened a null-pointer dereference corner case that needed a second fix. A team that backported v1 and moved on is not actually safe. Pragma Core tracks kernel and package versions continuously against advisory data, so both the original CVE and the follow-up hardening surface as required upgrades, with the fixed versions attached, rather than living in a maintainer thread nobody on your team is reading.

Full SBOM to answer "which kernel, where" before it slows you down

The first question in any kernel-CVE response is deceptively hard: exactly which kernel versions run where, including the container hosts, appliances, and embedded systems that never make it into the standard inventory. Pragma Core generates a complete, exportable component inventory per repository and environment, so "do we run an affected kernel, and which build" is a lookup instead of a fire drill. For a bug whose blast radius is every unpatched host, that speed is the difference between a scheduled patch and an incident.

Human-guided investigation for the code nobody has reread

The deepest value here mirrors what Nebula actually did: someone decided to seriously reread a fifteen-year-old subsystem. Pragma Core's expert-led research module lets an AppSec operator drive that kind of investigation into the parts of your stack you suspect are fragile and under-examined, backed by the platform's agents and the context already in your workspace. It turns "nobody has looked at that in years" from a standing risk into a scheduled activity.


Closing thoughts

CVE-2026-43499 is not, fundamentally, a bug about futexes, or about the red-black tree in rtmutex, or even about use-after-free. It is a bug about an assumption that was true when it was written and false when it was reused, and about the fact that nothing, not the compiler, not the review, not lockdep, not fifteen years of production, was checking whether it still held. The same shape lives in every large codebase: a helper with an implicit contract, a new caller that quietly breaks it, and a runtime that validates form without validating meaning. Kernels, hypervisors, browsers, and sprawling application backends are all full of these edges.

The difference between reading about GhostLock as a curiosity and using it as an audit trigger comes down to two things: how mature your AppSec practice is, and how much visibility you have into your own code. If your answer to "has anyone traced the ownership assumptions in our oldest, most-trusted modules?" is "not recently", that is the finding. Organizations that want to move from "we scan and report" to "we systematically investigate what is fragile" can reach Pragma Core at pragma-core.com for a demo.


Sources

Related posts
CVE-2026-74820: how an unsanitized ORDER BY clause turned ServiceNow's AI Platform into an unauthenticated database backdoor
Sep 18, 2026
CVE-2026-85978: how one path normalization mismatch turned Akana's admin console into unauthenticated remote code execution
Sep 9, 2026
CVE-2026-78174: how an unredacted session token in a diagnostic log turned a low-privileged WatchGuard Dimension admin into super admin
Sep 1, 2026

Start securing your codebase today

Connect your repositories and let AI agents handle continuous scanning, research, and triage.

Have questions? Get in touch →