[PATCH V11 0/9] famfs: port into fuse
John Groves <[email protected]> Mon, 20 Jul 2026 03:44:08 +0000
| Newsgroups | dev.linux.lists.nvdimm,dev.linux.lists.fuse-devel,org.kernel.vger.linux-cxl,org.kernel.vger.linux-doc,org.kernel.vger.linux-fsdevel,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <0100019f7d9fbe81-6cb16662-2522-47ea-a152-fab0ee3d9b35-000000@email.amazonses.com> |
From: John Groves <[email protected]> This has been quite a gauntlet so far. Some requests were made but eventually ruled out (e.g. using BPF for fault handling). Concrete requests (demands?) from the fuse community included the following: - Pass daxdevs into the kernel via an ioctl rather than a new fuse message/response - Lose the famfs interleaved extent format. I have complied with this request. It makes some things worse, but we'll discuss that later. There has also been a desire from the fuse community for famfs to use an experimental implementation of virtual backing-devs from Darrick. I'm leaving this out as a potential future optimization - although Darrick has announced that he (and $employer) intend to abandon that work, so it's not clear whether that will become viable. NOTE: the dax support that famfs requires (devdax fs-dax / fsdev) is now in upstream Linux, so this series no longer depends on a bundled or out-of-tree dax patch set - it applies on mainline. (It currently uses the upstream dax_dev_get(); a follow-on will switch to dax_dev_find() once that helper is upstream too - see "Changes v10 -> v11" below.) There are a few orthogonal fix patches for the devdax stack, and one that will trigger a patch to this code (changing from dax_get_get() to dax_dev_find()). I'll patch this code when that fix lands. Changes v10 -> v11 - The dax support famfs depends on is now upstream, so this series no longer carries or depends on a bundled dax patch set (v10 was based on Ira's for-7.1/dax-famfs branch [0]); v11 applies on mainline. - Dropped the controversial famfs interleaved extent format structures from the ABI. - Daxdev acquisition changed from a fuse message to an ioctl. The GET_DAXDEV message/response (kernel pulls a daxdev by pathname) has been replaced by FUSE_DEV_IOC_DAXDEV_OPEN: the fuse server now pushes each daxdev to the kernel by fd, and the kernel resolves it and acquires it exclusively via fs_dax_get(). This removes the GET_DAXDEV message plumbing (patch "famfs_fuse: GET_DAXDEV message and daxdev_table" becomes "fuse: register fs-dax daxdevs via FUSE_DEV_IOC_DAXDEV_OPEN"). - New patch "famfs_fuse: fail I/O on invalid or errored daxdevs": the iomap resolution path now gates on the state of the backing daxdev (slot never installed, exclusive acquire failed, or a memory error was reported via notify_failure) and fails the access (-EIO / -EHWPOISON) rather than mapping it. - Removed FUSE_FAMFS_MAX_EXTENTS. A GET_FMAP reply is now bounded by the reply buffer size rather than a separate extent-count cap. - Code-review pass over the whole series (self-review of every patch). Description: This patch series introduces famfs into the fuse file system framework. The dax support that famfs relies on is now upstream. The famfs user space code can be found at [1]. Fuse Overview: Famfs started as a standalone file system, but this series is intended to permanently supersede that implementation. At a high level, famfs adds one new fuse server message plus one device ioctl: GET_FMAP - Retrieves a famfs fmap (the file-to-dax map for a famfs file) FUSE_DEV_IOC_DAXDEV_OPEN - The fuse server pushes a daxdev (by fd) that was referenced by an fmap, so the kernel can acquire it Famfs Overview Famfs exposes shared memory as a file system. Famfs consumes shared memory from dax devices, and provides memory-mappable files that map directly to the memory - no page cache involvement. Famfs differs from conventional file systems in fs-dax mode, in that it handles in-memory metadata in a sharable way (which begins with never caching dirty shared metadata). Famfs started as a standalone file system [2,3], but the consensus at LSFMM was that it should be ported into fuse [4,5]. The key performance requirement is that famfs must resolve mapping faults without upcalls. This is achieved by fully caching the file-to-devdax metadata for all active files. The fmap is retrieved via the GET_FMAP message/response; the daxdevs an fmap references are registered by the server via the FUSE_DEV_IOC_DAXDEV_OPEN ioctl. Famfs remains the first fs-dax file system that is backed by devdax rather than pmem in fs-dax mode (hence the need for the new dax mode). Notes - When a file is opened in a famfs mount, the OPEN is followed by a GET_FMAP message and response. The "fmap" is the full file-to-dax mapping, allowing the fuse/famfs kernel code to handle read/write/fault without any upcalls. - Each fmap is checked for extents that reference previously-unknown daxdevs. The fuse server registers each such daxdev with the kernel via the FUSE_DEV_IOC_DAXDEV_OPEN ioctl (passing an fd to the devdax device), rather than the kernel pulling it by name. - Daxdevs are stored in a table (which might become an xarray at some point). When entries are added to the table, we acquire exclusive access to the daxdev via the fs_dax_get() call (modeled after how fs-dax handles this with pmem devices). Famfs provides holder_operations to devdax, providing a notification path in the event of memory errors or forced reconfiguration. - If devdax notifies famfs of memory errors on a dax device, famfs currently blocks all subsequent accesses to data on that device. The recovery is to re-initialize the memory and file system. Famfs is memory, not storage... - Because famfs uses backing (devdax) devices, only privileged mounts are supported (i.e. the fuse server requires CAP_SYS_RAWIO). - The famfs kernel code never accesses the memory directly - it only facilitates read, write and mmap on behalf of user processes, using fmap metadata provided by its privileged fuse server. As such, the RAS of the shared memory affects applications, but not the kernel. - Famfs has backing device(s), but they are devdax (char) rather than block. Right now there is no way to tell the vfs layer that famfs has a char backing device (unless we say it's block, but it's not). Currently we use the standard anonymous fuse fs_type - but I'm not sure that's ultimately optimal (thoughts?) Changes v9 -> v10 - Rebased to Ira's for-7.1/dax-famfs branch [0], which contains the required dax patches - Add parentheses to FUSE_IS_VIRTIO_DAX() macro, in case something bad is passed in as fuse_inode (thanks Jonathan's AI) Changes v8 -> v9 - Kconfig: fs/fuse/Kconfig:CONFIG_FUSE_FAMFS_DAX now depends on the new CONFIG_DEV_DAX_FSDEV (from drivers/dax/Kconfig) rather than just CONFIG_DEV_DAX and CONFIG_FS_DAX. (CONFIG_FUSE_FAMFS_DAX depends on those...) Changes v7 -> v8 - Moved to inline __free declaration in fuse_get_fmap() and famfs_fuse_meta_alloc(), famfs_teardown() - Adopted FIELD_PREP() macro rather than manual bitfield manipulation - Minor doc edits - I dropped adding magic numbers to include/uapi/linux/magic.h. That can be done later if appropriate Changes v6 -> v7 - Fixed a regression in famfs_interleave_fileofs_to_daxofs() that was reported by Intel's kernel test robot - Added a check in __fsdev_dax_direct_access() for negative return from pgoff_to_phys(), which would indicate an out-of-range offset - Fixed a bug in __famfs_meta_free(), where not all interleaved extents were freed - Added chunksize alignment checks in famfs_fuse_meta_alloc() and famfs_interleave_fileofs_to_daxofs() as interleaved chunks must be PTE or PMD aligned - Simplified famfs_file_init_dax() a bit - Re-ran CM's kernel code review prompts on the entire series and fixed several minor issues Changes v4 -> v5 -> v6 - None. Re-sending due to technical difficulties Changes v3 [9] -> v4 - The patch "dax: prevent driver unbind while filesystem holds device" has been dropped. Dan Williams indicated that the favored behavior is for a file system to stop working if an underlying driver is unbound, rather than preventing the unbind. - The patch "famfs_fuse: Famfs mount opt: -o shadow=<shadowpath>" has been dropped. Found a way for the famfs user space to do without the -o opt (via getxattr). - Squashed the fs/fuse/Kconfig patch into the first subsequent patch that needed the change ("famfs_fuse: Basic fuse kernel ABI enablement for famfs") - Many review comments addressed. - Addressed minor kerneldoc infractions reported by test robot. Changes v2 [7] -> v3 - Dax: Completely new fsdev driver (drivers/dax/fsdev.c) replaces the dev_dax_iomap modifications to bus.c/device.c. Devdax devices can now be switched among 'devdax', 'famfs' and 'system-ram' modes via daxctl or sysfs. - Dax: fsdev uses MEMORY_DEVICE_FS_DAX type and leaves folios at order-0 (no vmemmap_shift), allowing fs-dax to manage folio lifecycles dynamically like pmem does. - Dax: The "poisoned page" problem is properly fixed via fsdev_clear_folio_state(), which clears stale mapping/compound state when fsdev binds. The temporary WARN_ON_ONCE workaround in fs/dax.c has been removed. - Dax: Added dax_set_ops() so fsdev can set dax_operations at bind time (and clear them on unbind), since the dax_device is created before we know which driver will bind. - Dax: Added custom bind/unbind sysfs handlers; unbind return -EBUSY if a filesystem holds the device, preventing unbind while famfs is mounted. - Fuse: Famfs mounts now require that the fuse server/daemon has CAP_SYS_RAWIO because they expose raw memory devices. - Fuse: Added DAX address_space_operations with noop_dirty_folio since famfs is memory-backed with no writeback required. - Rebased to latest kernels, fully compatible with Alistair Popple et. al's recent dax refactoring. - Ran this series through Chris Mason's code review AI prompts to check for issues - several subtle problems found and fixed. - Dropped RFC status - this version is intended to be mergeable. Changes v1 [8] -> v2: - The GET_FMAP message/response has been moved from LOOKUP to OPEN, as was the pretty much unanimous consensus. - Made the response payload to GET_FMAP variable sized (patch 12) - Dodgy kerneldoc comments cleaned up or removed. - Fixed memory leak of fc->shadow in patch 11 (thanks Joanne) - Dropped many pr_debug and pr_notice calls References [0] - https://git.kernel.org/pub/scm/linux/kernel/git/nvdimm/nvdimm.git/ [1] - https://famfs.org (famfs user space) [2] - https://lore.kernel.org/linux-cxl/[email protected]/ [3] - https://lore.kernel.org/linux-cxl/[email protected]/ [4] - https://lwn.net/Articles/983105/ (lsfmm 2024) [5] - https://lwn.net/Articles/1020170/ (lsfmm 2025) [6] - https://lore.kernel.org/linux-cxl/cover.8068ad144a7eea4a813670301f4d2a86a8e68ec4.1740713401.git-series.apopple@nvidia.com/ [7] - https://lore.kernel.org/linux-fsdevel/[email protected]/ (famfs fuse v2) [8] - https://lore.kernel.org/linux-fsdevel/[email protected]/ (famfs fuse v1) [9] - https://lore.kernel.org/linux-fsdevel/[email protected]/T/#mb2c868801be16eca82dab239a1d201628534aea7 (famfs fuse v3) [10] - [TODO: John - famfs fuse v10 lore link] John Groves (9): famfs_fuse: Update macro s/FUSE_IS_DAX/FUSE_IS_VIRTIO_DAX/ famfs_fuse: Basic fuse kernel ABI enablement for famfs famfs_fuse: Plumb the GET_FMAP message/response famfs_fuse: Create files with famfs fmaps famfs_fuse: register fs-dax daxdevs via FUSE_DEV_IOC_DAXDEV_OPEN famfs_fuse: Plumb dax iomap and fuse read/write/mmap famfs_fuse: fail I/O on invalid or errored daxdevs famfs_fuse: Add DAX address_space_operations with noop_dirty_folio famfs_fuse: Add documentation Documentation/filesystems/famfs.rst | 143 ++++ Documentation/filesystems/index.rst | 1 + MAINTAINERS | 9 + fs/fuse/Kconfig | 13 + fs/fuse/Makefile | 1 + fs/fuse/dev.c | 31 + fs/fuse/dev.h | 1 + fs/fuse/dir.c | 2 +- fs/fuse/famfs.c | 1066 +++++++++++++++++++++++++++ fs/fuse/famfs_kfmap.h | 91 +++ fs/fuse/file.c | 48 +- fs/fuse/fuse_i.h | 125 +++- fs/fuse/inode.c | 35 +- fs/fuse/iomode.c | 2 +- include/uapi/linux/fuse.h | 55 +- 15 files changed, 1608 insertions(+), 15 deletions(-) create mode 100644 Documentation/filesystems/famfs.rst create mode 100644 fs/fuse/famfs.c create mode 100644 fs/fuse/famfs_kfmap.h -- 2.53.0