Re: [PATCH v4 3/6] xfs: implement write-stream management support
Kanchan Joshi <[email protected]> Fri, 31 Jul 2026 13:33:04 +0530
| Newsgroups | org.kernel.vger.linux-xfs,org.kernel.vger.linux-block,org.kernel.vger.linux-fsdevel |
|---|---|
| Message-ID | <[email protected]> |
On 7/21/2026 8:38 AM, Darrick J. Wong wrote: >> Write streams, filestreams, and write-life-time hints are mutually exclusive; >> combining any two of them returns -EINVAL: >> - GET_MAX reports 0 whenever xfs_inode_is_filestream() is true, >> covering both mount-wide filestreams and the per-inode chattr >> flag. Also when the file is on the realtime device. >> - SET refuses to bind a stream to a file that already has a >> write-life-time hint (fcntl F_SET_RW_HINT), is filestream, or is >> on the realtime device. >> - chattr refuses to set the filestream or realtime flag on a file >> that already has a write stream set. > These special "files" that represent stream ids could be generic code > instaed of in xfs. AFAICT the only thing you need from xfs is a pointer > from struct xfs_inode to struct (xfs_)write_stream, right? Right. Is it fine if we come to it when everything else is settled. I was hoping to lift common things up when write-stream is applied on another FS. > >> Suggested-by: Christoph Hellwig<[email protected]> >> Co-developed-by: Kanchan Joshi<[email protected]> >> Signed-off-by: Anuj Gupta<[email protected]> >> Signed-off-by: Kanchan Joshi<[email protected]> >> --- >> fs/xfs/xfs_icache.c | 1 + >> fs/xfs/xfs_inode.c | 155 ++++++++++++++++++++++++++++++++++++++++++++ >> fs/xfs/xfs_inode.h | 8 +++ >> fs/xfs/xfs_ioctl.c | 69 ++++++++++++++++++++ >> fs/xfs/xfs_iomap.c | 1 + >> fs/xfs/xfs_mount.h | 3 + >> fs/xfs/xfs_super.c | 12 ++++ >> 7 files changed, 249 insertions(+) >> >> diff --git a/fs/xfs/xfs_icache.c b/fs/xfs/xfs_icache.c >> index 9d8dd30bd927..7b9dda74122f 100644 >> --- a/fs/xfs/xfs_icache.c >> +++ b/fs/xfs/xfs_icache.c >> @@ -129,6 +129,7 @@ xfs_inode_alloc( >> spin_lock_init(&ip->i_ioend_lock); >> ip->i_next_unlinked = NULLAGINO; >> ip->i_prev_unlinked = 0; >> + ip->i_write_stream = 0; >> >> return ip; >> } >> diff --git a/fs/xfs/xfs_inode.c b/fs/xfs/xfs_inode.c >> index 15279d22a894..aafc3ffa6e0a 100644 >> --- a/fs/xfs/xfs_inode.c >> +++ b/fs/xfs/xfs_inode.c >> @@ -4,6 +4,7 @@ >> * All Rights Reserved. >> */ >> #include <linux/iversion.h> >> +#include <linux/anon_inodes.h> >> >> #include "xfs_platform.h" >> #include "xfs_fs.h" >> @@ -47,6 +48,160 @@ >> >> struct kmem_cache *xfs_inode_cache; >> >> +int >> +xfs_inode_max_write_streams( >> + struct xfs_inode *ip) >> +{ >> + struct block_device *bdev; >> + bool is_filestream, is_realtime; >> + >> + xfs_ilock(ip, XFS_ILOCK_SHARED); >> + is_filestream = xfs_inode_is_filestream(ip); >> + is_realtime = XFS_IS_REALTIME_INODE(ip); >> + bdev = xfs_inode_buftarg(ip)->bt_bdev; >> + xfs_iunlock(ip, XFS_ILOCK_SHARED); >> + >> + if (!bdev || is_filestream || is_realtime) > Would be nice if realtime worked, or someone at least adds a comment > about why it isn't (e.g. "we have something more exciting for rt/zoned > filesystems") etc. Will add, and that possibility is there. Since each allocator (default,filestream, rt, zoned) has a different logical placement, a single write-hint based scheme does not fit all. For non-zoned RT device: write-stream based physical placement can just work, it's mostly about lifting the checks that we have now. And adding logical placement (use write-stream to choose RTG similar to how we pick AG) is straight. But would you prefer that over default round-robin rotor? For zoned RT device: two options - Do nothing as write-hint based zone selection is already in place. - Or publish N write-streams (corresponds to N isolation buckets over >N open zones), and use file's write-stream value for zone selction. Current write-hint based scheme, based on xfs_zoned_hint_score[][] matrix, supports 4 isolation buckets. An application may need more, say 6. And two such applications will require 12. Zoned-RT can care more isolation buckets, but write-hints don't allow to publish that. Also, two unrelated applications that both pick WRITE_LIFE_SHORT bucket may land in the same zone. This goes back to FD based interface for write-steam. Not having enough buckets, and having multi-tenant mixing - both can improved with the write-stream. But I might be missing many zoned-xfs details; Maybe Christoph could weigh in?