Re: [REGRESSION] sched/idle: Sysbench threads regression after f4c31b07b136
Joseph Salisbury <[email protected]> Fri, 24 Jul 2026 13:20:41 -0400
| Newsgroups | dev.linux.lists.regressions,org.kernel.vger.linux-kernel,org.kernel.vger.linux-pm |
|---|---|
| Message-ID | <[email protected]> |
Hi Rafael, Christian, On 7/6/26 10:29 AM, Christian Loehle wrote: > On 7/2/26 19:47, Rafael J. Wysocki (Intel) wrote: >> Hi, >> >> On Thu, Jul 2, 2026 at 6:30 PM Joseph Salisbury >> <[email protected]> wrote: >>> Hi Rafael, >>> >>> We are seeing a reproducible MySQL Sysbench threads regression. A >>> bisect indicated the following commit as the first bad commit: >>> f4c31b07b136 ("sched: idle: Consolidate the handling of two special cases") >>> >>> The regression was found in Oracle kernel performance testing on OCI VM >>> shapes: >>> >>> VM Details: >>> * VM.Standard2.1: >>> x86 OCI VM shape, 1 OCPU / 2 hardware threads, about 14.5 GB RAM >>> >>> * VM.Standard.A1.Flex.2: >>> Arm/Ampere A1 flexible VM shape, 2 vCPU threads, about 10.9 GB RAM >>> >>> >>> The ResultsDB runs show the regression in the Sysbench threads metric: >>> >>> - VM.Standard2.1: 333 -> 236 (-29.1%) >>> - VM.Standard.A1.Flex.2: 1286 -> 1152 (-10.4%) >>> >>> A test kernel was built with f4c31b07b136 reverted and the performance >>> regression was recovered. >>> >>> From the code, it is possible the regression is due to the new >>> previous-wakeup heuristic in the special idle cases. Before the commit: >>> >>> - no cpuidle driver: >>> tick_nohz_idle_stop_tick() >>> default_idle_call() >>> >>> - one idle state: >>> tick_nohz_idle_retain_tick() >>> cpuidle state 0 >> I think that this is your case and the tick stops for you sometimes >> now while it had never stopped before. >> >> Can you confirm? The guest-visible data does not show the single-idle-state cpuidle case. Both affected guests report: /sys/devices/system/cpu/cpuidle/current_driver = none /sys/devices/system/cpu/cpuidle/current_governor = menu There are also no /sys/devices/system/cpu/cpu*/cpuidle entries on either guest. So from the guest data, this looks like the no-cpuidle-driver special case rather than the one-idle-state case. The full revert of f4c31b07b136 recovered the regression. I also tested Rafael's suggested one-line change, applied as: - idle_call_stop_or_retain_tick(stop_tick); + idle_call_stop_or_retain_tick(false); That test kernel still showed regressed performance on VM.Standard2.1. >> >> Overall, it would be good to know the idle state lists for both the VM >> and the host. For the VMs, there are no guest cpuidle state lists exposed because the cpuidle driver is "none". I do not currently have the host-side idle-state lists from the OCI hosts. I can try to get that data if it would still be useful. >> > +1, but also which HZ are you using? VM.Standard2.1, x86_64: CONFIG_HZ_1000=y CONFIG_HZ=1000 VM.Standard.A1.Flex.2, aarch64: CONFIG_HZ_250=y CONFIG_HZ=250 > Both systems reported have 2 logical CPUs then? Yes: VM.Standard2.1: CPU(s): 2 Thread(s) per core: 2 Core(s) per socket: 1 Socket(s): 1 VM.Standard.A1.Flex.2: CPU(s): 2 Thread(s) per core: 1 Core(s) per socket: 2 Socket(s): 1 > Were higher core counts also > tested and how does it affect them? Yes. Higher-core-count runs were checked. The regression appears limited to the smaller core-count shapes. The current data shows regressions on: - VM.Standard2.1 - VM.Standard.A1.Flex.2 - VM.Standard.E4.Flex.1 The larger tested shapes did not show the same regression pattern. The test runs use one sysbench thread per online CPU/core count as encoded in the metric name.