Re: [PATCH 2/5] xfs: fix racy open zone caching
"Darrick J. Wong" <[email protected]>
| Newsgroups | org.kernel.vger.linux-xfs |
|---|---|
| Message-ID | <20260810182142.GW3556460@frogsfrogsfrogs> |
On Mon, Aug 10, 2026 at 08:37:43AM -0700, Christoph Hellwig wrote: > When testing on very fast storage devices, I've observed writers using > io_uring creating many open zones with just a few kiB written to it, > which then don't get used. I tracked this down to multiple io_uring > helper threads finding a full zone in i_private, and then going on to > select a one, with the final one winning the race and leaving it in > i_private. > > Fix this by dropping full zones from i_private as soon we find them, > checking cached for a cached zoned when a single writes needs a new zone, > and by keeping an existing cached zone in xfs_set_cached_zone when it > still has space available, dropping the newly found/allocated one > instead. This uses i_flags_lock as a low-level spinlock for short > hold times to avoid interactions with the ilock, which is used for > completions. > > Signed-off-by: Christoph Hellwig <[email protected]> > --- > fs/xfs/xfs_zone_alloc.c | 56 +++++++++++++++++++++++++++++++---------- > 1 file changed, 43 insertions(+), 13 deletions(-) > > diff --git a/fs/xfs/xfs_zone_alloc.c b/fs/xfs/xfs_zone_alloc.c > index 7d13fa7ab30a..dee21f65f7b7 100644 > --- a/fs/xfs/xfs_zone_alloc.c > +++ b/fs/xfs/xfs_zone_alloc.c > @@ -793,17 +793,35 @@ xfs_get_cached_zone( > > rcu_read_lock(); > oz = VFS_I(ip)->i_private; > - if (oz) { > - /* > - * GC only steals open zones at mount time, so no GC zones > - * should end up in the cache. > - */ > - ASSERT(!oz->oz_is_gc); > - if (!atomic_inc_not_zero(&oz->oz_ref)) > + if (!oz) > + goto out_unlock; > + > + /* > + * GC only steals open zones at mount time, so no GC zones should end up > + * in the cache. > + */ > + ASSERT(!oz->oz_is_gc); > + > + /* > + * Drop the old cached open zone if it is full. > + */ > + if (oz->oz_allocated == rtg_blocks(oz->oz_rtg)) { > + spin_lock(&ip->i_flags_lock); > + oz = VFS_I(ip)->i_private; > + if (oz && oz->oz_allocated == rtg_blocks(oz->oz_rtg)) { > + VFS_I(ip)->i_private = NULL; > + spin_unlock(&ip->i_flags_lock); > + xfs_open_zone_put(oz); > oz = NULL; > + goto out_unlock; > + } > + spin_unlock(&ip->i_flags_lock); > } > - rcu_read_unlock(); > > + if (oz && !atomic_inc_not_zero(&oz->oz_ref)) > + oz = NULL; Do we still need to test oz for null-ness here? AFAICT we've already handled those cases here. > +out_unlock: > + rcu_read_unlock(); > return oz; > } > > @@ -819,17 +837,30 @@ xfs_get_cached_zone( > * lookup. Because the open_zone is clearly marked as full when all data > * in the underlying RTG was written, the caching is always safe. > */ > -static void > +static struct xfs_open_zone * > xfs_set_cached_zone( > struct xfs_inode *ip, > struct xfs_open_zone *oz) > { > struct xfs_open_zone *old_oz; > > + /* > + * If the open zone cached in the inode still has free space, use that > + * instead of the inode we just selected. This can happen when multiple > + * threads race to perform zone selection for an inode. io_uring worker > + * threads seem to be good at triggering this. > + */ > + spin_lock(&ip->i_flags_lock); > + old_oz = VFS_I(ip)->i_private; > + if (old_oz && old_oz->oz_allocated < rtg_blocks(old_oz->oz_rtg)) > + swap(oz, old_oz); > atomic_inc(&oz->oz_ref); Hmm, I'm confused about oz_ref handling here. If old_oz still has space, we swap oz and old_oz, after which oz alias i_private and old_oz is the zone that the caller passed in. The above line then increments oz->oz_ref and puts the zone that the caller passed in. Doesn't that cause oz->oz_ref to be too high? We already had a ref via i_private, and now we have another one. --D > - old_oz = xchg(&VFS_I(ip)->i_private, oz); > + VFS_I(ip)->i_private = oz; > + spin_unlock(&ip->i_flags_lock); > + > if (old_oz) > xfs_open_zone_put(old_oz); > + return oz; > } > > static void > @@ -873,14 +904,13 @@ xfs_zone_alloc_and_submit( > * the inode is still associated with a zone and use that if so. > */ > if (!*oz) > +select_zone: > *oz = xfs_get_cached_zone(ip); > - > if (!*oz) { > -select_zone: > *oz = xfs_select_zone(mp, write_hint, pack_tight); > if (!*oz) > goto out_error; > - xfs_set_cached_zone(ip, *oz); > + *oz = xfs_set_cached_zone(ip, *oz); > } > > alloc_len = xfs_zone_alloc_blocks(*oz, XFS_B_TO_FSB(mp, ioend->io_size), > -- > 2.53.0 > >