Files
SKELETONKEY/modules/refluxfs_cve_2026_64600/MODULE.md
T
KaraZajac e46a32f11e
build / build (clang / debug) (push) Waiting to run
build / build (clang / default) (push) Waiting to run
build / build (gcc / debug) (push) Waiting to run
build / build (gcc / default) (push) Waiting to run
build / sanitizers (ASan + UBSan) (push) Waiting to run
build / clang-tidy (push) Waiting to run
build / drift-check (CISA KEV + Debian tracker) (push) Waiting to run
build / static-build (push) Waiting to run
modules: add refluxfs (CVE-2026-64600, "RefluXFS" XFS reflink CoW ILOCK race)
Adds the corpus's first XFS module and its first data-oriented kernel bug —
every other kernel entry corrupts memory; this one corrupts file contents.

xfs_direct_write_iomap_begin() reads the data-fork extent map under ILOCK,
then xfs_reflink_fill_cow_hole() drops ILOCK to wait for transaction log
space. On reacquiring it, the code re-queries the refcount btree at the
ORIGINAL imap->br_startblock and never re-reads the data fork. A second
O_DIRECT writer holding only IOLOCK completes a whole CoW cycle in that
window, so the first writer's stale mapping sees refcount 1, treats a
still-shared block as private, and writes to it in place — landing its data
on the reflink source file's on-disk blocks.

The primitive is an arbitrary overwrite of the on-disk contents of any
readable file, which has three consequences that drive the design:
  - No offsets, no ROP, no KASLR/SMEP/SMAP; SELinux, containers and seccomp
    are all irrelevant.
  - The victim's inode is never written, so mtime/ctime/size never change
    and nothing is logged — FIM and `-w /etc/passwd -p wa` cannot see it.
  - The change persists across reboots.

Introduced 4.11 (3c68d44a2b49); fixed 2f4acd0fcd86 (mainline 7.2-rc4,
merged 2026-07-16), stable backports 7.1.4 / 6.18.39 / 6.12.96. Exposure is
distro-shaped: RHEL/CentOS/Rocky/Alma/Oracle/CloudLinux 8-10, Fedora Server
>= 31 and Amazon Linux 2023 ship XFS+reflink by default.

detect() is not a pure version gate — reachability here is safely
observable, so it pairs the backport table with a real storage precondition
(writable XFS via statfs XFS_SUPER_MAGIC, deliberately not via a successful
FICLONE since btrfs implements that too and is unaffected). --active
confirms reflink via FICLONE; SKELETONKEY_XFS_ASSUME_REFLINK=1/0 overrides.
On rpm-family hosts it warns that vendors backport without bumping the
upstream version, so the verdict speaks only to the upstream base.

exploit() forks a child that works only in a private mkdtemp scratch dir on
two files it owns: it establishes a shared extent (FICLONE, corroborated by
FIEMAP_EXTENT_SHARED) plus an O_DIRECT gate, then races a hard-bounded
8 writers / 2 helpers / 16 rounds / 2s and stops, reading the donor back
with O_DIRECT. Deliberately under-driven, and it never clones or targets a
file it does not own — the /etc/passwd overwrite -> su -> root step is
documented but NOT bundled. Always returns EXPLOIT_FAIL.

Safety rank 55, far above bad_epoll (12) and ghostlock (11): a won race
corrupts 4 KiB of our own scratch file and cannot touch kernel memory, so
there is no oops/KASAN/panic path.

Detection inverts the usual advice. auditd/sigma anchor on ioctl request
0x40049409 (FICLONE, matched exactly) and openat O_DIRECT; falco adds the
cross-uid reflink condition; and the yara rule is genuinely the right tool
here, matching the on-disk artifact because FIM is structurally blind.

VM-VERIFIED 2026-07-23 — the corpus's first rpm-family verification, taking
the empirical count to 29 of 41 CVEs. Rocky Linux 9.8 /
5.14.0-687.10.1.el9_8.0.1.x86_64 under qemu/KVM, stock GenericCloud layout
with no provisioner changes (root is XFS with reflink=1 out of the box).
detect() -> VULNERABLE, --active FICLONE witness confirmed reflink, phase A
observed FIEMAP_EXTENT_SHARED on a real shared extent, scratch self-cleaned,
clean build on el9 gcc. The underlying bug was separately confirmed winnable
on that kernel via tools/verify-vm/refluxfs_verify.c at the public PoC's
parameters (32 writers / 8 helpers, 60s): 4/4 runs won, first divergence
after 69/114/170/494 rounds. The shipped under-driven trigger did NOT win in
its 2s budget on that same vulnerable kernel — intended behaviour, and
exactly why a non-win must never be read as "patched".

14 new detect() unit rows (148 tests total, 0 failures). Bumps to v0.9.14.

Credit: Qualys Threat Research Unit (blog by Saeed Abbasi; the technical
advisory credits model-assisted kernel analysis performed with Anthropic),
and the upstream XFS maintainers who fixed it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0118iUgHY44hdRtANgyCmu7y
2026-07-23 17:49:40 -04:00

12 KiB

refluxfs — CVE-2026-64600

"RefluXFS" — a time-of-check/time-of-use race in the XFS reflink copy-on-write path that lets any unprivileged local user overwrite the on-disk contents of any file they can read, on any XFS volume mounted with reflink=1 that they can write to. No user namespace, no capability, no crafted filesystem image, no kernel offsets. It has been present since reflink direct-I/O CoW landed in 4.11 (2017) — a nine-year window.

This is the corpus's first XFS module, and its first data-oriented kernel bug: the primitive is an arbitrary file content overwrite, not memory corruption.

The bug

xfs_direct_write_iomap_begin() (fs/xfs/xfs_iomap.c) reads the data-fork extent map under ILOCK. To allocate a transaction it must wait for log space, so xfs_reflink_fill_cow_hole() (fs/xfs/xfs_reflink.c) drops ILOCK. On re-acquiring it, the code re-queries the refcount btree at the original physical block number (imap->br_startblock) — and never re-reads the data fork.

A second O_DIRECT writer, holding only the coarser IOLOCK, can complete an entire CoW cycle inside that window: allocate block Y, write it, and remap via xfs_reflink_end_cow(). The first writer's imap now points at a block owned solely by the reflink source. Its stale refcount lookup returns 1, it concludes the block is private, and writes to it in place — landing attacker data on the source file's on-disk blocks.

Three consequences follow, and they drive the whole module design:

  1. No offsets, no ROP, no KASLR/SMEP/SMAP. There is nothing to port per kernel build. Qualys is explicit that SELinux enforcing, container boundaries and seccomp are equally irrelevant.
  2. The victim's inode is never written. The data is applied to the shared physical block underneath it, so mtime/ctime/size do not change and there is no kernel log output. File-integrity monitoring does not fire.
  3. It persists across reboots, because the change is on disk.

The public demonstration (RHEL 10.2) reflink-clones /etc/passwd into /var/tmp, races concurrent direct-I/O writes against the clone, thereby rewriting /etc/passwd itself to strip root's password, and runs su.

Affected range

Introduced 4.11 (2017-02, commit 3c68d44a2b49, "xfs: allocate direct I/O COW blocks in iomap_begin")
Fixed upstream commit 2f4acd0fcd86 ("xfs: resample the data fork mapping after cycling ILOCK") — merged 2026-07-16, released 7.2-rc4
Stable backports 7.1.4 (e705d81a7193) · 6.18.39 (206c09b04dc5) · 6.12.96 (44f891bc0889)
Affected, no upstream fix 6.6 / 6.1 / 5.15 / 5.14 / 5.10 / 4.19 / 4.18 LTS lines (per the CNA record at time of writing)
Not affected < 4.11 — includes RHEL/CentOS 7 (3.10 predates reflink)
NVD class CWE-362 (race) → CWE-367 (TOCTOU). NVD had published no CWE and no CVSS vector at time of writing
CISA KEV no (disclosed 2026-07-22)

The exposure is distro-shaped, not kernel-shaped. What matters is whether XFS+reflink is the installer default:

Exploitable out of the box Not reachable by default
RHEL 8/9/10 · CentOS Stream 8/9/10 · Rocky/AlmaLinux 8/9/10 · Oracle Linux 8/9/10 (RHCK + UEK R6/R7/8) · CloudLinux 8/9/10 · Fedora Server ≥ 31 · Amazon Linux 2023 (and AL2 AMIs from 2022-12) Debian · Ubuntu · Fedora Workstation · SLES · openSUSE · Arch (ext4/btrfs defaults — unless an XFS volume was added deliberately)

⚠️ The version gate has a real blind spot here

The affected population is overwhelmingly RHEL-family, and those vendors backport fixes without bumping the upstream base version — a patched RHEL 8 kernel still reports 4.18.0-xxx.el8. An upstream-version gate cannot see that.

So on rpm-family hosts, a VULNERABLE verdict is a statement about the upstream base version only. detect() prints that warning explicitly rather than implying it checked the erratum. Confirm against the vendor advisory (RHSA / ELSA / ALSA / RLSA) before acting on it.

Trigger / detection

Unlike a pure kernel race, this bug's reachability can be established safely and deterministically, so detect() is a version gate plus a real precondition probe:

  • Passive — is there a writable directory on a mounted XFS filesystem? Identified via statfs(2) f_type == XFS_SUPER_MAGIC, not by a successful FICLONE, because btrfs implements FICLONE too and is unaffected. No such directory → PRECOND_FAIL, the correct verdict on a stock Debian/Ubuntu host.
  • Active (--active / --auto) — confirms reflink=1 empirically by cloning and removing two 4 KiB files, rather than assuming the mkfs.xfs default. reflink=0 → no shared extents can exist → PRECOND_FAIL.
  • OverrideSKELETONKEY_XFS_ASSUME_REFLINK=1 (force reachable) / 0 (force unreachable), for when you know the fleet's storage layout better than a local probe can. This also drives the unit tests.

exploit() forks an isolated child that creates a private mkdtemp scratch directory on the XFS mount and works only on two files it owns:

  • (A) deterministic + safe — writes a donor file, FICLONE-clones it, and confirms via FIEMAP_EXTENT_SHARED that the clone's extent really is shared (refcount > 1), plus that O_DIRECT opens succeed. That is a read-only observation that the exact filesystem state the bug misjudges exists here. Reflink cloning is an ordinary supported operation, so this phase is safe on any kernel.
  • (B) hard-bounded window exercise — races 8 concurrent O_DIRECT 4 KiB writes against the clone while 2 helper threads cycle ftruncate/fdatasync to keep the transaction allocator dropping ILOCK to wait for log space, for at most 16 rounds / 2 s. Then it stops and reads the donor back with O_DIRECT — a buffered read would be served from the page cache that the corruption bypasses, and would hide a win.

It is deliberately under-driven (the public PoC uses 32 writers and 8 helpers and grinds far longer) and, more importantly, it never clones or targets a file it does not own. The step that makes this root — reflink-cloning a root-owned file such as /etc/passwd and racing writes onto its shared blocks — persistently rewrites a system file on disk with no undo, so it is documented here and not bundled. exploit() always returns EXPLOIT_FAIL.

If the race is won, the module says so loudly: that is CVE-2026-64600 confirmed present, empirically, with the damage contained to 4 KiB of the operator's own scratch file.

VM verification (2026-07-23)

Confirmed on Rocky Linux 9.8 / 5.14.0-687.10.1.el9_8.0.1.x86_64 under qemu/KVM with 6 vCPUs — the stock GenericCloud installer layout, root on /dev/vda4 XFS with reflink=1, no provisioner changes:

Check Result
detect() on real XFS VULNERABLE (found writable XFS at /var/tmp)
rpm-family backport caveat fired correctly
--active FICLONE witness reflink CONFIRMED
Phase A shared extent FIEMAP_EXTENT_SHARED set (btrfs never reported it; XFS does)
Phase A O_DIRECT gate available
Shipped trigger (8 writers / 2 helpers / 2 s) ran 16 rounds, did not winby design
Scratch cleanup no artifacts left
Build on el9 gcc clean

The underlying bug was separately confirmed winnable on that kernel, using a VM-only harness driven at the public PoC's parameters (32 writers / 8 helpers, 60 s budget — tools/verify-vm/refluxfs_verify.c): 4 out of 4 runs won, with the first divergence after 69, 114, 170 and 494 rounds. A racing O_DIRECT write landed on a still-shared block and rewrote the donor's on-disk bytes — the arbitrary-overwrite primitive, observed directly, contained to files the test user owned.

Note carefully what this does and does not say. The shipped trigger not winning in 2 s on a kernel that is provably vulnerable is exactly the designed behaviour, and is the concrete reason a non-win must never be read as "patched" — trust the version gate and the vendor erratum instead.

Why this ranks above the other reconstructed race triggers

bad_epoll (12) and ghostlock (11) sit at the bottom of the --auto safety ranking because a won race frees a live struct file or corrupts the kernel stack — silent destabilisation or near-certain panic. Neither applies here. RefluXFS corrupts file data, not kernel memory: there is no oops, no KASAN report, no panic risk, and the blast radius of a win is one 4 KiB scratch file we created and delete. That is why refluxfs carries safety rank 55 — it is genuinely safe to run, and the ranking should say so. The VM run above bears this out: the bug was won 4/4 times on a vulnerable kernel with no oops, no dmesg output and no instability.

Detection — the obvious rule does not work

Do not rely on -w /etc/passwd -p wa, AIDE, or Tripwire for this CVE. The attacker never issues a write(2) against the victim inode; XFS applies their data to the shared physical block beneath it. Size, mtime and ctime are unchanged and nothing is logged. Anyone relying on FIM to catch a passwd modification is blind to this bug by construction.

What does work, in descending order of fidelity:

  1. The reflink itselfioctl(fd, FICLONE, srcfd) where FICLONE is 0x40049409. auditd can match the request number exactly, so it does not flood, and the attack cannot avoid it. Tune out cp --reflink=auto, podman and systemd-nspawn image work.
  2. O_DIRECT opensopenat flags & 0x4000. Also on the critical path, and rare outside databases and backup agents.
  3. Content-vs-metadata drift — because the bytes change while mtime does not, hashing /etc/passwd, /etc/shadow and the setuid binaries on a schedule and alerting when the content hash moves without a corresponding mtime change is a near-zero-false-positive detector for this whole bug class.

The shipped rules cover all three: auditd/sigma anchor on the FICLONE request number and O_DIRECT opens (correlated per-pid, plus the post-exploitation euid-0 transition), falco adds the high-fidelity "reflinked a file owned by another user" condition, and — unusually for a kernel bug — the yara rule is genuinely the right tool, matching the on-disk artifact (/etc/passwd with a password-less root entry or an added uid-0 account) precisely because there is no metadata trace for FIM to find.

Fix / mitigation

Upgrade the kernel (≥ 7.1.4 / 6.18.39 / 6.12.96 on-branch, or 7.2+; on RHEL-family, the vendor erratum) and reboot.

There is no partial mitigation, which is why mitigate() is NULL: reflink is a superblock feature that cannot be disabled on a live filesystem, O_DIRECT cannot be turned off, and — because this is a data-oriented bug — SELinux enforcing, container boundaries, KASLR, SMEP, SMAP and seccomp are all irrelevant. Qualys puts it plainly: "This isn't a vulnerability you can harden around, isolate, or live-patch."

cleanup() sweeps any skeletonkey-refluxfs-* scratch directories left behind if a trigger run was killed mid-round; normal runs remove their own.

Credit

Discovery and research: Qualys Threat Research Unit (TRU); the blog post is authored by Saeed Abbasi, and the technical advisory credits model-assisted kernel analysis performed with Anthropic. Upstream fix 2f4acd0fcd86. See NOTICE.md.