Re: [PATCH v4 4/6] xfs: generic AG set based steering

"Darrick J. Wong" <[email protected]>
Newsgroups org.kernel.vger.linux-xfs,org.kernel.vger.linux-block,org.kernel.vger.linux-fsdevel
Message-ID <20260721032027.GW7380@frogsfrogsfrogs>
On Fri, Jul 17, 2026 at 06:25:36PM +0530, Kanchan Joshi wrote:
> Improve allocator concurrency and reduce interleaving by introducing
> fixed sized AG set.
> Use low bits of the inode as a hash to select AG within the AG set.
> Overall, a file will try to use the same AG (and contiguity is maintained),
> but multiple files will be spread across all AGs in the target AG set.
> 
> Suggested-by: Dave Chinner <[email protected]>
> Signed-off-by: Kanchan Joshi <[email protected]>
> Signed-off-by: Anuj Gupta <[email protected]>
> ---
>  fs/xfs/libxfs/xfs_bmap.c | 38 ++++++++++++++++++++++++++++++++++++++
>  1 file changed, 38 insertions(+)
> 
> diff --git a/fs/xfs/libxfs/xfs_bmap.c b/fs/xfs/libxfs/xfs_bmap.c
> index d64defeda645..fd1a3aa4ad3f 100644
> --- a/fs/xfs/libxfs/xfs_bmap.c
> +++ b/fs/xfs/libxfs/xfs_bmap.c
> @@ -3192,6 +3192,36 @@ xfs_bmap_select_minlen(
>  	return args->maxlen;
>  }
>  
> +#define	GENERIC_AG_SET_SZ	(2)

What does this define?

> +
> +static inline xfs_agnumber_t
> +xfs_default_ag_set_size(
> +	struct xfs_inode	*ip)
> +{
> +	struct xfs_mount	*mp = ip->i_mount;
> +
> +	return min_t(xfs_agnumber_t, GENERIC_AG_SET_SZ, mp->m_sb.sb_agcount);

Because I'm not sure what it means on a single-AG filesystem.

> +}
> +
> +static xfs_agnumber_t
> +xfs_ag_to_ag_set(
> +	struct xfs_bmalloca	*ap,
> +	xfs_agnumber_t		base_agno)
> +{
> +	struct xfs_inode	*ip = ap->ip;
> +	struct xfs_mount	*mp = ip->i_mount;
> +	xfs_agnumber_t		set_size;
> +
> +	/* Apply fanning only for regular file data */
> +	if (!(ap->datatype & XFS_ALLOC_USERDATA))
> +		return base_agno;
> +
> +	set_size = xfs_default_ag_set_size(ip);
> +	/* Fan out within the AG set using low bits of the inode */
> +	return (base_agno + (XFS_INO_TO_AGINO(mp, I_INO(ip)) % set_size)) %
> +		mp->m_sb.sb_agcount;
> +}
> +
>  static int
>  xfs_bmap_btalloc_select_lengths(
>  	struct xfs_bmalloca	*ap,
> @@ -3587,8 +3617,16 @@ xfs_bmap_btalloc_best_length(
>  {
>  	xfs_extlen_t		blen = 0;
>  	int			error;
> +	xfs_agnumber_t		target_ag, start_ag;
>  
>  	ap->blkno = XFS_INODE_TO_FSB(ap->ip);
> +
> +	/* fan out initial AG across the generic AG set */
> +	start_ag = XFS_FSB_TO_AGNO(args->mp, ap->blkno);
> +	target_ag = xfs_ag_to_ag_set(ap, start_ag);
> +	if (target_ag != start_ag)
> +		ap->blkno = XFS_AGB_TO_FSB(args->mp, target_ag, 0);

/me wonders, if xfs_bmap_rtalloc looked at ap->blkno for a hint the way
that the data device allocator does, then would it be trivial to have
write streams on the rt device too?

I guess the tricky part would be figuring out what to do if you ever
want to switch a file between rt and data devices -- presumably you'd
just reset the write stream id to the default, but I guess you could
reject such a switch if the id had been set explicitly.

--D

> +
>  	if (!xfs_bmap_adjacent(ap))
>  		ap->eof = false;
>  
> -- 
> 2.25.1
> 
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.