Re: [PATCH v4 3/6] xfs: implement write-stream management support

Kanchan Joshi <[email protected]>
Newsgroups gmane.linux.block,gmane.linux.file-systems
Message-ID <[email protected]>
On 7/21/2026 8:38 AM, Darrick J. Wong wrote:
>> Write streams, filestreams, and write-life-time hints are mutually exclusive;
>> combining any two of them returns -EINVAL:
>>    - GET_MAX reports 0 whenever xfs_inode_is_filestream() is true,
>>      covering both mount-wide filestreams and the per-inode chattr
>>      flag. Also when the file is on the realtime device.
>>    - SET refuses to bind a stream to a file that already has a
>>      write-life-time hint (fcntl F_SET_RW_HINT), is filestream, or is
>>      on the realtime device.
>>    - chattr refuses to set the filestream or realtime flag on a file
>>      that already has a write stream set.
> These special "files" that represent stream ids could be generic code
> instaed of in xfs.  AFAICT the only thing you need from xfs is a pointer
> from struct xfs_inode to struct (xfs_)write_stream, right?

Right. Is it fine if we come to it when everything else is settled.
I was hoping to lift common things up when write-stream is applied on 
another FS.
> 
>> Suggested-by: Christoph Hellwig<[email protected]>
>> Co-developed-by: Kanchan Joshi<[email protected]>
>> Signed-off-by: Anuj Gupta<[email protected]>
>> Signed-off-by: Kanchan Joshi<[email protected]>
>> ---
>>   fs/xfs/xfs_icache.c |   1 +
>>   fs/xfs/xfs_inode.c  | 155 ++++++++++++++++++++++++++++++++++++++++++++
>>   fs/xfs/xfs_inode.h  |   8 +++
>>   fs/xfs/xfs_ioctl.c  |  69 ++++++++++++++++++++
>>   fs/xfs/xfs_iomap.c  |   1 +
>>   fs/xfs/xfs_mount.h  |   3 +
>>   fs/xfs/xfs_super.c  |  12 ++++
>>   7 files changed, 249 insertions(+)
>>
>> diff --git a/fs/xfs/xfs_icache.c b/fs/xfs/xfs_icache.c
>> index 9d8dd30bd927..7b9dda74122f 100644
>> --- a/fs/xfs/xfs_icache.c
>> +++ b/fs/xfs/xfs_icache.c
>> @@ -129,6 +129,7 @@ xfs_inode_alloc(
>>   	spin_lock_init(&ip->i_ioend_lock);
>>   	ip->i_next_unlinked = NULLAGINO;
>>   	ip->i_prev_unlinked = 0;
>> +	ip->i_write_stream = 0;
>>   
>>   	return ip;
>>   }
>> diff --git a/fs/xfs/xfs_inode.c b/fs/xfs/xfs_inode.c
>> index 15279d22a894..aafc3ffa6e0a 100644
>> --- a/fs/xfs/xfs_inode.c
>> +++ b/fs/xfs/xfs_inode.c
>> @@ -4,6 +4,7 @@
>>    * All Rights Reserved.
>>    */
>>   #include <linux/iversion.h>
>> +#include <linux/anon_inodes.h>
>>   
>>   #include "xfs_platform.h"
>>   #include "xfs_fs.h"
>> @@ -47,6 +48,160 @@
>>   
>>   struct kmem_cache *xfs_inode_cache;
>>   
>> +int
>> +xfs_inode_max_write_streams(
>> +	struct xfs_inode	*ip)
>> +{
>> +	struct block_device	*bdev;
>> +	bool			is_filestream, is_realtime;
>> +
>> +	xfs_ilock(ip, XFS_ILOCK_SHARED);
>> +	is_filestream = xfs_inode_is_filestream(ip);
>> +	is_realtime = XFS_IS_REALTIME_INODE(ip);
>> +	bdev = xfs_inode_buftarg(ip)->bt_bdev;
>> +	xfs_iunlock(ip, XFS_ILOCK_SHARED);
>> +
>> +	if (!bdev || is_filestream || is_realtime)
> Would be nice if realtime worked, or someone at least adds a comment
> about why it isn't (e.g. "we have something more exciting for rt/zoned
> filesystems") etc.

Will add, and that possibility is there.

Since each allocator (default,filestream, rt, zoned) has a different 
logical placement, a single write-hint based scheme does not fit all.

For non-zoned RT device: write-stream based physical placement can just 
work, it's mostly about lifting the checks that we have now. And adding 
logical placement (use write-stream to choose RTG similar to how we pick 
AG) is straight. But would you prefer that over default round-robin rotor?

For zoned RT device: two options
- Do nothing as write-hint based zone selection is already in place.
- Or publish N write-streams (corresponds to N isolation buckets over >N 
open zones), and use file's write-stream value for zone selction.

Current write-hint based scheme, based on xfs_zoned_hint_score[][] 
matrix, supports 4 isolation buckets. An application may need more, say 
6. And two such applications will require 12. Zoned-RT can care more 
isolation buckets, but write-hints don't allow to publish that.

Also, two unrelated applications that both pick WRITE_LIFE_SHORT bucket 
may land in the same zone. This goes back to FD based interface for 
write-steam.

Not having enough buckets, and having multi-tenant mixing - both can 
improved with the write-stream.
But I might be missing many zoned-xfs details; Maybe Christoph could 
weigh in?
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.