Re: Fwd: BadBunny: UFFDIO_COPY shmem killpriv bypass leading to local privilege escalation
Pedro Falcato <[email protected]> Sat, 8 Aug 2026 13:17:09 +0100
| Newsgroups | gmane.linux.file-systems,gmane.linux.kernel.mm |
|---|---|
| Message-ID | <[email protected]> |
I'm adding a bunch of fs people that might have opinions about this. Please see the rest of the thread below. On Sat, Aug 08, 2026 at 12:13:06PM +0300, vova tokarev wrote: > Pedro, > > Thanks for the review. Let me address both points. > > On the page fault comparison: > > UFFDIO_COPY is not comparable to a page fault. Page faults serve > existing page cache content (read-like). UFFDIO_COPY adds new, A page fault can also write to the page cache. On MAP_SHARED mappings. > > user-controlled content to the shmem page cache (write-like). > > The correct comparison is to write(), which calls file_modified() > > -> __remove_privs() to strip SUID/SGID. > > > This is exactly the reasoning used when fallocate() was fixed > > across XFS (commit fbe7e5200365), ext4, and f2fs to call > > file_modified(). The XFS commit message says: > > "as various fallocate modes can change the file contents [...] > > we should drop file privileges like suid just like we do for a > > regular write()" I can't speak for fallocate, or other system calls. As far as I'm aware, this is a best-effort kind of thing. As I said, writing to a MAP_SHARED mapping does not clear the setuid bit. It's a super trivial thing to do, too. But it's not a problem because setuid executables are not world-writable. I simply don't think this can feasibly be a security boundary, considering how much it has been historically screwed up, and how the second-most basic way to write to a file Just Bypasses It. I also don't know a single setuid program that's packaged as 04777. Do you? > > UFFDIO_COPY on shmem changes file contents. > > It should drop file privileges like suid, just like write(). > > > On the PoC setup and exploitability: > > The PoC uses mode 04777 for simplicity of demonstration, > > but the underlying bug is a killpriv invariant violation: > > every VFS write path calls file_modified() to strip SUID on > > content modification, but shmem_mfill_filemap_add() does not. > > The killpriv mechanism is defense-in-depth - if file permissions > > alone were sufficient to protect SUID, the kernel wouldn't bother > > stripping SUID on write(). But it does, because writable SUID files > > do occur in practice (group-writable SUID binaries, POSIX ACLs, > > container shared mounts, chained with a separate write-access bug). > > > For precedent: CVE-2023-0386 (overlayfs copy-up preserving SUID > > across namespaces) is the same class of bug - a kernel code path > > that modifies or copies file content without stripping SUID - > > and was scored CVSS 7.8 and added to CISA's KEV catalog. > > The fallocate killpriv fixes were backported to all stable trees. > > > Additionally, this path is available even with > > vm.unprivileged_userfaultfd=0, since UFFD_USER_MODE_ONLY > > bypasses the privilege check (userfaultfd_syscall_allowed() > > returns true unconditionally for USER_MODE_ONLY). > > The kernel considers this path safe for unprivileged use, > > yet it skips killpriv. > > > *On the fix:* > > Regardless of how we classify severity, the fix is trivial and makes > > UFFDIO_COPY consistent with every other write path - > > add file_modified() to shmem_mfill_filemap_add(). > > I'm happy to submit a patch if you'd like. > > > Best regards, > > Vladimir > > On Fri, Aug 7, 2026 at 5:14 PM Pedro Falcato <[email protected]> wrote: > > > On Fri, Aug 07, 2026 at 01:40:41PM +0300, vova tokarev wrote: > > > Hi, > > > > > > It's been almost two months since I sent this report, and I haven't > > > heard back. I'd really appreciate any feedback when you get a chance. > > > > > > I've rechecked both mainline master and stable 6.12.95 -- the > > > vulnerability remains unfixed in both trees: > > > > > > 1. mm/shmem.c: shmem_mfill_atomic_pte() (6.12) / > > shmem_mfill_filemap_add() > > > (7.x) still adds pages to the page cache without calling > > file_modified() > > > or __remove_privs(). Writing to a SUID binary on tmpfs via UFFDIO_COPY > > > preserves the setuid bit. > > > > > > 2. mm/userfaultfd.c: I noticed commit 85668fda932a added retry state > > > tracking on master, but MFILL_RETRY_STATE_VMA_FLAGS still does not > > > include VMA_WRITE_BIT -- the mprotect TOCTOU remains exploitable. > > > > > > This is a deterministic local privilege escalation (no race timing > > > needed for the killpriv bypass), affects every kernel since 4.11 > > > (8+ years), and works on any system with userfaultfd + tmpfs (the > > > default on virtually all distributions). > > > > > > I have a full working PoC that gets uid=0 from uid=1000 reliably. > > > Happy to provide any additional information if needed. > > > > > > Thanks, > > > Vladimir > > > > > > > > > ---------- Forwarded message --------- > > > From: vova tokarev <[email protected]> > > > Date: Tue, Jun 16, 2026 at 12:37 PM > > > Subject: BadBunny: UFFDIO_COPY shmem killpriv bypass leading to local > > > privilege escalation > > > To: <[email protected]> > > > > > > > > > Hi, > > > > > > I found a local privilege escalation (BadBunny) in the userfaultfd + > > > shmem subsystems that affects all Linux kernels from 4.11 to 7.1 > > > (every major distribution: Ubuntu, Debian, Fedora, RHEL, SUSE, Arch, > > > Android, ChromeOS, and any system with CONFIG_USERFAULTFD=y and a > > > tmpfs/shmem mount). > > > > > > Two bugs are chained: > > > > > > 1. TOCTOU in UFFDIO_COPY retry path (mm/userfaultfd.c): > > > mfill_retry_state_changed() does not re-validate VM_WRITE after > > > dropping and re-acquiring the mmap lock in mfill_copy_folio_retry(). > > > A concurrent mprotect(PROT_READ) installs a writable PTE into a > > > now-read-only VMA. > > > > This sounds like a bug, but not really exploitable. > > > > > > > > 2. Missing killpriv in shmem UFFDIO_COPY (mm/shmem.c): > > > shmem_mfill_filemap_add() adds pages to the shmem page cache > > > without calling file_modified()/killpriv. This preserves SUID/SGID > > > bits when file content is replaced via UFFDIO_COPY, unlike normal > > > write() which strips them. > > > > > > NOTE: The killpriv bypass (Bug 2) does not require > > > unprivileged userfaultfd and works even with > > vm.unprivileged_userfaultfd=0, > > > since UFFDIO_COPY on shmem is available to any process that can open > > > a tmpfs file O_RDWR and call userfaultfd with UFFD_USER_MODE_ONLY. > > > > Who made the suid file world-writable? Note that this is not a bug, page > > fault paths don't clear the suid bit either. > > > > > > > > An unprivileged user can replace the content of a SUID-root binary on > > > tmpfs via UFFDIO_COPY while preserving its setuid permission, then > > > execute it to obtain root. > > > > > > The attack is deterministic (no timing dependency for the killpriv > > > bypass), requires no heap spraying, and bypasses all modern kernel > > > mitigations (KASLR, SMEP, SMAP, CFI, PAC, heap hardening). > > > > > > Affected versions: Linux 4.11+ (since shmem UFFDIO_COPY support, > > > commit 4c27fe4c4c84 "userfaultfd: shmem: add shmem_mcopy_atomic_pte") > > > > This sounds like a bug, but not really exploitable. > > > Confirmed on: 7.1.0 (aarch64) > > > Affected distros: All major distributions (Ubuntu, Debian, Fedora, > > > RHEL, SUSE, Arch, Android, ChromeOS) that have CONFIG_USERFAULTFD=y > > > and tmpfs mounted (virtually all Linux systems) > > > > > > Attached files: > > > - bad_bunny.c Full LPE exploit (uid=1000 to > > > uid=0) Build: gcc -static -O2 -pthread > > > - suidhelper.c Standalone SUID payload > > binary > > > Build: gcc -static -O2 > > > - uffdio_copy_lpe_report.md Detailed writeup with root cause, > > > reproduction steps, and suggested fix > > > - badbunny_demo.mp4 PoC demo clip > > > > > > The PoC (bad_bunny.c) sets up a SUID target on tmpfs, drops to > > > > Since the exploit isn't public, I assume you set up the tmpfs file as root > > and world writable. This is not an LPE. > > > > > > -- > > Pedro > > -- Pedro