modules: add bad_epoll (CVE-2026-46242, "Bad Epoll" epoll teardown race UAF)
release / build (arm64) (push) Waiting to run
release / build (x86_64) (push) Waiting to run
release / build (x86_64-static / musl) (push) Waiting to run
release / build (arm64-static / musl) (push) Waiting to run
release / release (push) Blocked by required conditions

Jaeyoung Chung's kernelCTF "Bad Epoll": a race-condition use-after-free in
fs/eventpoll.c. ep_remove() clears file->f_ep under f_lock but keeps using
the file (hlist_del_rcu + unlock) while a concurrent __fput() frees the
still-referenced struct eventpoll. Reachable by any unprivileged user with
no userns / CONFIG / capability; weaponised via cross-cache to a struct
file, /proc/self/fdinfo arb-read, and ROP. Introduced 58c9b016e128 (6.4),
fixed a6dc643c6931 (7.1-rc1), stable backport 7.0.13. CWE-416; not in KEV.

The corpus's first epoll / VFS-teardown module. Shipped as a reconstructed,
deliberately under-driven trigger on the stackrot / nft_catchall contract:
detect() is a pure version gate (>= 6.4 and below the on-branch fix; no
userns precondition -- epoll needs none); exploit() forks a CPU-pinned
child that exercises the ep_remove-vs-__fput close window a hard-bounded 48
attempts / 2s and returns EXPLOIT_FAIL. It does not grind the race to a
win, does not do the cross-cache reclaim, and does not bundle the fdinfo
R/W + ROP (a won race frees a live struct file and rarely trips KASAN ->
silent-corruption risk).

Wired: registry, Makefile, safety rank (12 -- lowest in the corpus; a
kernel race is the least predictable class), 5 detect() test rows (version
gating), CVE metadata (sorted insert, CWE-416 / T1068 / not-KEV), README +
CVES.md + website counts (44/39), RELEASE_NOTES v0.9.12, and a verify-vm
target (sweep pending). Detection rules are intentionally weak/structural
(epoll ubiquitous, rarely KASAN) -- post-exploitation euid-0 transition,
no yara. Also corrects pre-existing README drift in the "not yet verified"
count (8 -> 11). Not VM-verified, so the verified count stays 28 of 39.
Version 0.9.12. Credit: Jaeyoung Chung (J-jaeyoung).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KUq4DGXSPBmPkAyJ9WnM9n
This commit is contained in:
KaraZajac
2026-07-04 19:35:33 -04:00
parent 4d0a0e2443
commit 95589e26cb
17 changed files with 787 additions and 22 deletions
+106
View File
@@ -0,0 +1,106 @@
# bad_epoll — CVE-2026-46242
"Bad Epoll" — a race-condition use-after-free in the Linux kernel epoll
subsystem (`fs/eventpoll.c`) reachable by **any unprivileged local user**.
No user namespace, no capability, no special `CONFIG``epoll_create1(2)`,
`epoll_ctl(2)`, and `close(2)` are available to everyone, which is what
makes this bug unusually dangerous.
## The bug
On the file-teardown path, `ep_remove()` clears `file->f_ep` under
`file->f_lock` but keeps **using** the file inside the same critical
section — the `hlist_del_rcu()` walk over the eventpoll's `refs` list and
the trailing `spin_unlock()`. A concurrent `__fput()` of a linked epoll
file can observe the transient `NULL` `f_ep`, skip
`eventpoll_release_file()`, and jump straight to `f_op->release`, freeing
a `struct eventpoll` that the first path is still walking →
**use-after-free** on a live kernel object.
The public exploit (Jaeyoung Chung, submitted to Google's kernelCTF)
arranges four epoll objects in two pairs — one pair drives the race, the
other is the victim — and converts the 8-byte UAF write into control of a
`struct file` via a **cross-cache** attack (the freed `eventpoll` slab
page is drained to the buddy allocator and reclaimed as pipe backing
buffers). From there it reads arbitrary kernel memory through
`/proc/self/fdinfo` and ROPs to a root shell. Roughly **99% reliable**
despite a race window only ~6 instructions wide; the racer widens it with
`close(dup())` storms that induce false-sharing on the file's `f_count`
cache line. It **rarely trips KASAN**, which is why the bug survived three
years and why it is hard to detect at runtime.
## Affected range
| | |
|---|---|
| Vulnerable path introduced | commit `58c9b016e128` — Linux **6.4** (2023-04-08) |
| Fixed upstream | commit `a6dc643c69311677c574a0f17a3f4d66a5f3744b` — merged for **7.1-rc1** (2026-04-24) |
| Stable backport | **7.0.13** (Debian forky `7.0.13-1` / sid `7.0.14-1`) |
| Still vulnerable at time of writing | trixie **6.12.x** (no backport yet); 6.6 LTS pending |
| Not affected | 6.1 and older (predate the bug — Debian: "vulnerable code not present") |
| NVD class | CWE-416 (Use After Free) via CWE-362 (race) |
| CISA KEV | no (brand new) |
Table threshold is a single `{7,0,13}` entry — `kernel_range_is_patched()`
treats 7.1+ as patched-via-mainline and everything in `[6.4, 7.0.13)` as
vulnerable, matching the Debian tracker. Add 6.6.x / 6.12.x rows when
those LTS backports land (`tools/refresh-kernel-ranges.py` flags them).
## Trigger / detection
`detect()` is a **pure version gate** — no active probe, because there is
no cheap, safe way to distinguish a vulnerable kernel from a patched one
without actually winning the race (the dangerous part). It returns `OK`
below 6.4 or on a patched kernel, and `VULNERABLE` in range. There is **no
`PRECOND_FAIL` userns path** the way `nft_catchall` has — epoll needs no
namespace, so there is no unprivileged-userns stopgap to report or to
harden with.
`exploit()` forks a CPU-pinned child that builds the epoll race pair (a
waiter eventpoll watching a target eventpoll) and exercises the
`ep_remove`-vs-`__fput` concurrent-close window a **hard-bounded** number
of times (48 attempts / 2 s), widening it with `close(dup())`
false-sharing storms, snapshots the `eventpoll`/`kmalloc-192` slab, and
returns `EXPLOIT_FAIL`.
It is **deliberately under-driven**. A *won* race frees a live
`struct eventpoll` — genuine kernel memory corruption that rarely trips
KASAN, so on a vulnerable production host a completed race can silently
destabilise the box rather than cleanly oops. This module therefore does
**not** grind the race to a win, does **not** perform the cross-cache
reclaim, and does **not** bundle the per-kernel `fdinfo` arbitrary-read +
ROP that lands root (per-build offsets refused). The trigger is
**reconstructed from the public kernelCTF PoC and is not VM-verified**. It
never claims root it did not get.
Because a kernel race is the least predictable class in the corpus — and
this one can corrupt memory invisibly — `bad_epoll` carries the **lowest
`--auto` safety rank** (see `module_safety_rank()` in `skeletonkey.c`), so
`--auto` only ever reaches for it after every safer vulnerable module.
## Detection is hard — read this before shipping the rules
Unlike most modules, `bad_epoll` has **no high-fidelity signature**.
`epoll_create1` / `epoll_ctl` / `close` is the steady-state behaviour of
nginx, systemd, and every language runtime's event loop; the exploit
looks identical and rarely trips KASAN. The shipped auditd/sigma/falco
rules therefore key on the **post-exploitation** tell — an unprivileged
process transitioning to euid 0 without a setuid `execve` — plus a
recommendation to monitor kernel logs for oops/BUG lines. Expect false
positives from legitimate privilege-management daemons and tune per
environment. There is no yara rule (no file artifact). Treat this module
as much as a *blue-team teaching case* — "here is a root LPE your existing
stack is nearly blind to" — as an offensive one.
## Fix / mitigation
Upgrade the kernel (>= 7.0.13, or 7.1+). There is **no partial
mitigation**: epoll cannot be disabled in practice, and no
`unprivileged_userns_clone` / sysctl toggle closes this path the way it
does for the netfilter bugs. `mitigate()` is `NULL` for that reason.
## Credit
Discovery, exploitation, and the public kernelCTF PoC:
**Jaeyoung Chung** (`J-jaeyoung`). Upstream fix `a6dc643c6931`. See
`NOTICE.md`.