Re: [PATCH v2] btrfs: replace writeback inhibition xarray with a fixed inline buffer

David Sterba <[email protected]> Tue, 21 Jul 2026 14:32:02 +0200
Newsgroups org.kernel.vger.linux-btrfs
Message-ID <[email protected]>
On Fri, Jul 17, 2026 at 01:11:11AM +0800, Chengfeng Lin wrote:
> Hi Leo,
> 
> I independently tested the exact v2 code diff on a physical x86-64 machine.
> In this narrow synthetic Btrfs remap-range microbenchmark, it was consistently
> faster than the unpatched xarray control.  This is not an application
> workload.  It also differs from the LKP fio 4 KiB random-write case: this test
> exercises the same Btrfs writeback-inhibition functions through 4 KiB clone
> and dedupe ioctls.
> 
> I applied the code diff from this v2 message to a frozen Linux 7.1.0 source
> snapshot containing the Btrfs change at
> f9a48549a15aa369d42cebc08a6a72b71a53d547.  The exact saved diff used for the
> build has SHA-256
> 5ec741be5a89d6dae0c0608cc036512770b55d6a49e9b576b4aa3115ebfdffd3.
> Aside from the v2 diff, the control and patched kernels used the same base
> source snapshot, normalized config, GCC 15.2.0 toolchain and Kbuild metadata,
> with `CONFIG_PREEMPT_DYNAMIC=y` and `preempt=none`.
> 
> The machine was an Intel Core i7-12700KF system with 20 logical CPUs and
> 32 GiB RAM.  For each kernel point I created a fresh Btrfs filesystem on a
> 1 GiB RAM-backed brd device.  The timed process was pinned to P-core logical
> CPU 2 with the `intel_pstate` `performance` governor.  Each point ran 15
> rounds; each round performed 10,000 4 KiB FICLONERANGE operations and 10,000
> 4 KiB FIDEDUPERANGE operations.
> 
> In a control -> v2 -> control sandwich, the means across all 15 rounds were in
> ns/op (lower is better):
> 
>   operation          control A   v2         control B   v2 vs control midpoint
>   FICLONERANGE 4 KiB  2943.790    2159.123   2970.446    -26.986%
>   FIDEDUPERANGE 4 KiB 3535.456    2762.835   3546.942    -21.980%
> 
> The two control points drifted by only 0.905% for clone and 0.325% for dedupe.
> Dropping the first round from each point produced deltas of -27.178% and
> -22.128%.
> 
> All 90 timing rows passed the semantic checks: expected_match_ratio=100 and
> unexpected_results=0.  A separate ftrace direct-hit run on the v2 kernel with
> 1,000 clone plus 1,000 dedupe operations observed 2,000 calls to
> btrfs_remap_file_range, 7,473 to btrfs_inhibit_eb_writeback, and 2,033 to
> btrfs_uninhibit_all_eb_writeback.
> 
> For this specific brd-backed 4 KiB clone/dedupe workload, v2 reduced the
> measured per-operation latency relative to the unpatched xarray control by
> about 27% for clone and 22% for dedupe.  I have not tested application-level
> impact, physical-storage impact, or a broad filesystem matrix.  The Tested-by
> below applies to the v2 code diff identified by this message and the SHA-256
> above.
> 
> Evidence bundle:
> 
>   https://github.com/lcf0399/linux-regression-evidence/tree/7e47ea94a4d151610f16e98429efc3c3b8865265/btrfs-remap-writeback-inhibition-v2
> 
> Standalone workload:
> 
>   https://github.com/lcf0399/linux-regression-evidence/tree/7e47ea94a4d151610f16e98429efc3c3b8865265/btrfs-remap-writeback-inhibition-v2/reproducer
> 
> Tested-by: Chengfeng Lin <[email protected]>

Thanks for the report, I've added link to the patch.