Re: [REGRESSION] io_uring/futex: scalar wait/wake slowdown after 079afb081c42
Chengfeng Lin <[email protected]> Fri, 31 Jul 2026 16:31:30 +0800
| Newsgroups | org.kernel.vger.io-uring,dev.linux.lists.regressions,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <CANGjgdkhQZWntcnpa2zBshGn_E7yaKDnPcSbH-HBBfXGWAw1+g@mail.gmail.com> |
Hi Jens,
Thanks for your reply, and sorry that I forgot to include the reproducer
parameters in my previous message. The exact standalone parameters are:
ring entries: 64
futex words: 32 private u32 words, one per cache line
futex flags: FUTEX2_SIZE_U32 | FUTEX2_PRIVATE
per cycle: submit 32 scalar WAITs for value 0, then 32 scalar WAKEs
with nr_wake=3D1, then drain exactly 64 CQEs
per round: 512 cycles, or 16,384 WAIT/WAKE pairs
run: 3 untimed warm-up rounds followed by 15 measured rounds
machine: Intel Core i7-12700KF, 20 logical CPUs, 32 GiB RAM
CPU: one process pinned to P-core CPU 2
Each WAIT CQE must return 0, each WAKE CQE must return 1, and the program
fails on a missing or duplicate CQE, an unexpected flag or result, timeout,
CQ overflow, outstanding request, or CPU migration. The normal command has
no arguments:
./io_uring_futex_wait_wake_standalone
I also tested the exact two patches you attached on frozen Linus master
48a5a7ab8d6a. I used the same normalized config, GCC 15.2.0 toolchain, and
Kbuild metadata, with a fresh boot for every measured point:
baseline A -> patch 1 A -> full series -> patch 1 B -> baseline B
The means were:
point ns/pair
baseline A 196.271
patch 1 A 191.071
full series 189.900
patch 1 B 188.824
baseline B 196.962
Using the surrounding-control midpoints, patch 1 was 3.392% faster than the
baseline and the full series was 3.416% faster. Full versus the patch-1
midpoint changed by only -0.025%. This is consistent with patch 2 retaining
inflight tracking for private waits.
Baseline drift was 0.352% and patch-1 drift was -1.176%. Dropping the first
measured round gave -3.323%, -3.345%, and -0.022% for the same three
comparisons.
I then added a matched --shared mode. It keeps the workload and timing the
same. The only changes are that the words use a shared anonymous mapping an=
d
the SQEs use FUTEX2_SIZE_U32 without FUTEX2_PRIVATE. This directly exercise=
s
patch 2's scalar shared-WAIT path. I compared:
patch 1 A -> full series -> patch 1 B
The means were 271.107, 267.254, and 271.971 ns/pair. The full series was
1.578% faster than the patch-1 midpoint. Patch-1 drift was 0.319%, the
drop-first result was also -1.578%, and the maximum CV was 0.200%.
All 120 measured rows across the private and shared tests passed their
checks. All eight measured boots actually ran with preempt=3Dfull; the proc=
ess
used the performance governor and EPP, with Turbo disabled.
On this frozen-master baseline, patch 1 reduced latency by 3.392% in the
private-futex workload. Adding patch 2 did not materially change the privat=
e
result. In the matched scalar shared-futex test, patch 2 reduced latency by
1.578%. I have not tested shared WAITV.
This is a different source baseline from my earlier direct-parent test of
079afb081c42, so I am not subtracting the 3.4% here from the earlier 9.27%
result.
The compact results and exact patch identities are available here:
https://github.com/lcf0399/linux-regression-evidence/tree/87da380476eb56c=
b3642d6ed15baf22c59496094/io-uring-futex-inflight-wait-wake
Thanks,
Chengfeng
Jens Axboe <[email protected]> =E4=BA=8E2026=E5=B9=B47=E6=9C=8830=E6=97=A5=E5=
=91=A8=E5=9B=9B 23:49=E5=86=99=E9=81=93=EF=BC=9A
>
> On 7/30/26 8:34 AM, Jens Axboe wrote:
> > On 7/30/26 4:42 AM, Chengfeng Lin wrote:
> >> Hi Jens,
> >>
> >> I tested 079afb081c42 against its direct parent on bare metal. In a na=
rrow
> >> scalar io_uring futex wait/wake workload, the child was 9.27% slower. =
A
> >> separate 338-line standalone reproduced the result at 10.02%. All comp=
ared
> >> kernels actually ran with preempt=3Dfull.
> >>
> >> #regzbot introduced: 079afb081c4288e94d5e4223d3eb6306d853c68b
> >> #regzbot title: io_uring scalar futex wait/wake slowdown
> >>
> >> This is a focused synthetic microbenchmark, not an application benchma=
rk. It
> >> uses one raw-UAPI ring and 32 cacheline-separated private futex words =
on one
> >> pinned P-core. Each timed cycle submits 32 scalar IORING_OP_FUTEX_WAIT
> >> requests, then 32 scalar IORING_OP_FUTEX_WAKE requests, and drains exa=
ctly 64
> >> CQEs. Every wait must return 0 and every wake must return 1.
> >>
> >> I used a fresh boot for each point:
> >>
> >> 6a8118a77eec parent A -> 079afb081c42 child -> 6a8118a77eec parent B
> >>
> >> Each point had 3 warm-up rounds and 15 measured rounds. Every measured=
round
> >> ran 512 cycles, or 16,384 wait/wake pairs. The results in ns/pair were=
:
> >>
> >> implementation parent A child parent B child vs midpoint
> >> formal 180.079 196.647 179.856 +9.268%
> >> standalone 180.171 197.856 179.502 +10.020%
> >>
> >> For the formal source, dropping the first measured round gave +9.269%.=
Parent
> >> drift was -0.124%, and the maximum CV was 0.169%. The standalone drop-=
first
> >> result was +9.996%, with -0.371% parent drift. All 90 scalar timing ro=
ws
> >> passed the CQE, result, timeout, overflow, outstanding-request, and CP=
U
> >> checks.
> >>
> >> An untimed child trace also hit io_futex_prep(), io_futex_wait(),
> >> io_futex_wake(), and io_futex_complete() with the expected request cou=
nts.
> >>
> >> A matched WAITV -> WAKE profile changed by only +1.385%, below my prer=
egistered
> >> 5% signal gate, so my claim is limited to scalar wait/wake.
> >>
> >> I understand that 079afb fixes the exit-time use-after-free by keeping=
pending
> >> private futex waits visible to cancellation before their mm state disa=
ppears.
> >> Scalar WAIT and WAKE both use io_futex_prep(), so in the child both si=
des of
> >> each measured pair execute the added tracking call. I am not suggestin=
g a
> >> revert.
> >>
> >> Is this per-request cost an expected trade-off for the lifetime fix, o=
r could
> >> the same exit/mm-lifetime guarantee be retained with cheaper tracking?
> >>
> >> Evidence bundle:
> >>
> >> https://github.com/lcf0399/linux-regression-evidence/tree/65c8cbf86f=
40cbe759e3f7db1d29d152ba03a8f2/io-uring-futex-inflight-wait-wake
> >>
> >> Standalone reproducer:
> >>
> >> https://github.com/lcf0399/linux-regression-evidence/tree/65c8cbf86f=
40cbe759e3f7db1d29d152ba03a8f2/io-uring-futex-inflight-wait-wake/reproducer
> >
> > Great report, thanks for that! I'll take a look at this. The inflight
> > tracking is a bit of a big hammer for sure for this, and it isn't even
> > needed on the wake side. Can you tell me what parameters you're using
> > for the reproducers?
>
> Try with these two patches.
>
> --
> Jens Axboe