Re: [PATCH 2/5] xfs: fix racy open zone caching

"Darrick J. Wong" <[email protected]>
Newsgroups org.kernel.vger.linux-xfs
Message-ID <20260810182142.GW3556460@frogsfrogsfrogs>
On Mon, Aug 10, 2026 at 08:37:43AM -0700, Christoph Hellwig wrote:
> When testing on very fast storage devices, I've observed writers using
> io_uring creating many open zones with just a few kiB written to it,
> which then don't get used.  I tracked this down to multiple io_uring
> helper threads finding a full zone in i_private, and then going on to
> select a one, with the final one winning the race and leaving it in
> i_private.
> 
> Fix this by dropping full zones from i_private as soon we find them,
> checking cached for a cached zoned when a single writes needs a new zone,
> and by keeping an existing cached zone in xfs_set_cached_zone when it
> still has space available, dropping the newly found/allocated one
> instead.  This uses i_flags_lock as a low-level spinlock for short
> hold times to avoid interactions with the ilock, which is used for
> completions.
> 
> Signed-off-by: Christoph Hellwig <[email protected]>
> ---
>  fs/xfs/xfs_zone_alloc.c | 56 +++++++++++++++++++++++++++++++----------
>  1 file changed, 43 insertions(+), 13 deletions(-)
> 
> diff --git a/fs/xfs/xfs_zone_alloc.c b/fs/xfs/xfs_zone_alloc.c
> index 7d13fa7ab30a..dee21f65f7b7 100644
> --- a/fs/xfs/xfs_zone_alloc.c
> +++ b/fs/xfs/xfs_zone_alloc.c
> @@ -793,17 +793,35 @@ xfs_get_cached_zone(
>  
>  	rcu_read_lock();
>  	oz = VFS_I(ip)->i_private;
> -	if (oz) {
> -		/*
> -		 * GC only steals open zones at mount time, so no GC zones
> -		 * should end up in the cache.
> -		 */
> -		ASSERT(!oz->oz_is_gc);
> -		if (!atomic_inc_not_zero(&oz->oz_ref))
> +	if (!oz)
> +		goto out_unlock;
> +
> +	/*
> +	 * GC only steals open zones at mount time, so no GC zones should end up
> +	 * in the cache.
> +	 */
> +	ASSERT(!oz->oz_is_gc);
> +
> +	/*
> +	 * Drop the old cached open zone if it is full.
> +	 */
> +	if (oz->oz_allocated == rtg_blocks(oz->oz_rtg)) {
> +		spin_lock(&ip->i_flags_lock);
> +		oz = VFS_I(ip)->i_private;
> +		if (oz && oz->oz_allocated == rtg_blocks(oz->oz_rtg)) {
> +			VFS_I(ip)->i_private = NULL;
> +			spin_unlock(&ip->i_flags_lock);
> +			xfs_open_zone_put(oz);
>  			oz = NULL;
> +			goto out_unlock;
> +		}
> +		spin_unlock(&ip->i_flags_lock);
>  	}
> -	rcu_read_unlock();
>  
> +	if (oz && !atomic_inc_not_zero(&oz->oz_ref))
> +		oz = NULL;

Do we still need to test oz for null-ness here?  AFAICT we've already
handled those cases here.

> +out_unlock:
> +	rcu_read_unlock();
>  	return oz;
>  }
>  
> @@ -819,17 +837,30 @@ xfs_get_cached_zone(
>   * lookup.  Because the open_zone is clearly marked as full when all data
>   * in the underlying RTG was written, the caching is always safe.
>   */
> -static void
> +static struct xfs_open_zone *
>  xfs_set_cached_zone(
>  	struct xfs_inode	*ip,
>  	struct xfs_open_zone	*oz)
>  {
>  	struct xfs_open_zone	*old_oz;
>  
> +	/*
> +	 * If the open zone cached in the inode still has free space, use that
> +	 * instead of the inode we just selected.  This can happen when multiple
> +	 * threads race to perform zone selection for an inode.  io_uring worker
> +	 * threads seem to be good at triggering this.
> +	 */
> +	spin_lock(&ip->i_flags_lock);
> +	old_oz = VFS_I(ip)->i_private;
> +	if (old_oz && old_oz->oz_allocated < rtg_blocks(old_oz->oz_rtg))
> +		swap(oz, old_oz);
>  	atomic_inc(&oz->oz_ref);

Hmm, I'm confused about oz_ref handling here.

If old_oz still has space, we swap oz and old_oz, after which oz alias
i_private and old_oz is the zone that the caller passed in.  The above
line then increments oz->oz_ref and puts the zone that the caller passed
in.

Doesn't that cause oz->oz_ref to be too high?  We already had a ref
via i_private, and now we have another one.

--D

> -	old_oz = xchg(&VFS_I(ip)->i_private, oz);
> +	VFS_I(ip)->i_private = oz;
> +	spin_unlock(&ip->i_flags_lock);
> +
>  	if (old_oz)
>  		xfs_open_zone_put(old_oz);
> +	return oz;
>  }
>  
>  static void
> @@ -873,14 +904,13 @@ xfs_zone_alloc_and_submit(
>  	 * the inode is still associated with a zone and use that if so.
>  	 */
>  	if (!*oz)
> +select_zone:
>  		*oz = xfs_get_cached_zone(ip);
> -
>  	if (!*oz) {
> -select_zone:
>  		*oz = xfs_select_zone(mp, write_hint, pack_tight);
>  		if (!*oz)
>  			goto out_error;
> -		xfs_set_cached_zone(ip, *oz);
> +		*oz = xfs_set_cached_zone(ip, *oz);
>  	}
>  
>  	alloc_len = xfs_zone_alloc_blocks(*oz, XFS_B_TO_FSB(mp, ioend->io_size),
> -- 
> 2.53.0
> 
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.