[REGRESSION] io_uring/futex: scalar wait/wake slowdown after 079afb081c42
Chengfeng Lin <[email protected]> Thu, 30 Jul 2026 18:42:34 +0800
| Newsgroups | dev.linux.lists.regressions,org.kernel.vger.io-uring,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <CANGjgdn=R_qyUdE=j9za+vkmqcxacbP-84OHXF4nZ4ho9qRyVg@mail.gmail.com> |
Hi Jens, I tested 079afb081c42 against its direct parent on bare metal. In a narrow scalar io_uring futex wait/wake workload, the child was 9.27% slower. A separate 338-line standalone reproduced the result at 10.02%. All compared kernels actually ran with preempt=full. #regzbot introduced: 079afb081c4288e94d5e4223d3eb6306d853c68b #regzbot title: io_uring scalar futex wait/wake slowdown This is a focused synthetic microbenchmark, not an application benchmark. It uses one raw-UAPI ring and 32 cacheline-separated private futex words on one pinned P-core. Each timed cycle submits 32 scalar IORING_OP_FUTEX_WAIT requests, then 32 scalar IORING_OP_FUTEX_WAKE requests, and drains exactly 64 CQEs. Every wait must return 0 and every wake must return 1. I used a fresh boot for each point: 6a8118a77eec parent A -> 079afb081c42 child -> 6a8118a77eec parent B Each point had 3 warm-up rounds and 15 measured rounds. Every measured round ran 512 cycles, or 16,384 wait/wake pairs. The results in ns/pair were: implementation parent A child parent B child vs midpoint formal 180.079 196.647 179.856 +9.268% standalone 180.171 197.856 179.502 +10.020% For the formal source, dropping the first measured round gave +9.269%. Parent drift was -0.124%, and the maximum CV was 0.169%. The standalone drop-first result was +9.996%, with -0.371% parent drift. All 90 scalar timing rows passed the CQE, result, timeout, overflow, outstanding-request, and CPU checks. An untimed child trace also hit io_futex_prep(), io_futex_wait(), io_futex_wake(), and io_futex_complete() with the expected request counts. A matched WAITV -> WAKE profile changed by only +1.385%, below my preregistered 5% signal gate, so my claim is limited to scalar wait/wake. I understand that 079afb fixes the exit-time use-after-free by keeping pending private futex waits visible to cancellation before their mm state disappears. Scalar WAIT and WAKE both use io_futex_prep(), so in the child both sides of each measured pair execute the added tracking call. I am not suggesting a revert. Is this per-request cost an expected trade-off for the lifetime fix, or could the same exit/mm-lifetime guarantee be retained with cheaper tracking? Evidence bundle: https://github.com/lcf0399/linux-regression-evidence/tree/65c8cbf86f40cbe759e3f7db1d29d152ba03a8f2/io-uring-futex-inflight-wait-wake Standalone reproducer: https://github.com/lcf0399/linux-regression-evidence/tree/65c8cbf86f40cbe759e3f7db1d29d152ba03a8f2/io-uring-futex-inflight-wait-wake/reproducer Thanks, Chengfeng