Re: [PATCH v2] btrfs: replace writeback inhibition xarray with a fixed inline buffer
David Sterba <[email protected]> Tue, 21 Jul 2026 14:27:07 +0200
| Newsgroups | org.kernel.vger.linux-btrfs |
|---|---|
| Message-ID | <[email protected]> |
On Wed, Jul 01, 2026 at 04:47:10PM -0700, Leo Martins wrote:
> Commit f9a48549a15a ("btrfs: inhibit extent buffer writeback to prevent
> COW amplification") tracks the extent buffers a transaction handle has
> inhibited in a per-handle xarray. Keying the tracking to the transaction
> handle is correct, but using an xarray for it causes two problems in
> production.
>
> First, a write_iops regression. Every COW calls
> btrfs_inhibit_eb_writeback() from btrfs_force_cow_block() and
> should_cow_block(), which does an xa_store() keyed by eb->start. The
> kernel test robot reported a 22.6% fio.write_iops regression on a
> single-task 4k randwrite workload (ftruncate ioengine, buffered IO) on
> btrfs. The cost is the per-COW xarray store done on every COW'd block.
> Replacing it with a non-allocating fixed buffer recovers the lost
> throughput, and that buffer does more per-COW bookkeeping yet still
> recovers, so the cost is the xarray operation itself rather than the
> extra tracking work.
>
> Second, an unbounded cleanup walk. btrfs_uninhibit_all_eb_writeback()
> iterates every eb the handle inhibited with xa_for_each(). A single
> handle that COWs a very large number of blocks (inode eviction, or
> truncate of a file with many extents, where btrfs_truncate_inode_items()
> loops over many search_again descents under one handle) makes that walk
> arbitrarily long. It runs in __btrfs_end_transaction() before
> num_writers is dropped, so it blocks the committing thread; this shows up
> as multi-second stalls and RCU stall reports.
>
> Replace the xarray with a fixed inline array on btrfs_trans_handle,
> managed with a CLOCK (second-chance) eviction policy. Inhibiting a buffer
> becomes an array append with no allocation and no tree walk, and the
> end-of-handle cleanup is bounded by the array size.
>
> The set that actually needs protection is the working set the handle
> revisits across search_again descents, the search path frontier, which is
> on the order of the tree height. It is not every block the handle ever
> COWs. should_cow_block() re-inhibiting an already tracked buffer marks it
> referenced, so revisited buffers survive eviction while write-once buffers
> are reclaimed first. A small fixed buffer is therefore enough where a
> non-evicting array would either overflow or have to grow without bound.
> BTRFS_INHIBITED_EBS_SLOTS is 8 and the reference bits pack into a u32.
>
> The CLOCK eviction is what justifies the extra complexity over a plain
> non-evicting array. The test workload stresses amplification: it removes
> 16 heavily fragmented 64 MiB files in one transaction while background
> writeback keeps writing out in-use metadata. A re-COW event is a buffer
> already COWed in the running transaction that was written back and then
> COWed again; the figure below is the ratio of re-COW events to first-COW
> events summed across the eviction (n=5, lower is better):
>
> tracking re-COW per first-COW
> no inhibition 6.1
> non-evicting array, 32 slots 3.8
> CLOCK array, 8 slots (this patch) 1.6
> unbounded xarray (reverted) 1.4
>
> The non-evicting array fills with write-once buffers and stops covering
> the buffers the handle keeps revisiting, so even at four times the slots
> it leaves most of the amplification. CLOCK evicts the cold buffers and
> keeps the revisited ones, recovering almost all of the unbounded benefit.
> The eviction policy, not the buffer size, is what closes the gap.
>
> eb->writeback_inhibitors and the WB_SYNC_ALL bypass in
> lock_extent_buffer_for_io() are unchanged, so fsync and commit behavior
> are unaffected. A reference is taken on each tracked buffer so it cannot
> be freed while the array points at it; eviction drops that reference and
> the inhibitor count.
>
> Fixes: f9a48549a15a ("btrfs: inhibit extent buffer writeback to prevent COW amplification")
> Reported-by: kernel test robot <[email protected]>
> Closes: https://lore.kernel.org/oe-lkp/[email protected]
> Signed-off-by: Leo Martins <[email protected]>
> Reviewed-by: Sun YangKai <[email protected]>
> ---
> v2:
> - Present the amplification numbers as a table instead of prose (David Sterba).
> - Use int for the loop indices and the slot local instead of u32 (David Sterba).
> - Replace the BTRFS_INHIBITED_EBS_SLOTS comment with static_assert checks for
> the <= 32 bound and the power-of-two size (David Sterba).
> - Widen inhibited_ebs_hand from u8 to u32; the handle stays in the same slab
> bucket and the u8 only left an alignment hole (David Sterba).
> - Factor slot selection and eviction into btrfs_inhibit_claim_slot()
> (Sun YangKai).
> - Add Reviewed-by from Sun YangKai.
I've fixed up the thing Filipe pointed out and added the patch to
for-next, also with a reference to the testing report from Chengfeng Lin.