Re: [REGRESSION] sched/idle: Sysbench threads regression after f4c31b07b136
Joseph Salisbury <[email protected]> Wed, 8 Jul 2026 11:25:02 -0400
| Newsgroups | dev.linux.lists.regressions,org.kernel.vger.linux-kernel,org.kernel.vger.linux-pm |
|---|---|
| Message-ID | <[email protected]> |
On 7/6/26 10:29 AM, Christian Loehle wrote: > On 7/2/26 19:47, Rafael J. Wysocki (Intel) wrote: >> Hi, >> >> On Thu, Jul 2, 2026 at 6:30 PM Joseph Salisbury >> <[email protected]> wrote: >>> Hi Rafael, >>> >>> We are seeing a reproducible MySQL Sysbench threads regression. A >>> bisect indicated the following commit as the first bad commit: >>> f4c31b07b136 ("sched: idle: Consolidate the handling of two special cases") >>> >>> The regression was found in Oracle kernel performance testing on OCI VM >>> shapes: >>> >>> VM Details: >>> * VM.Standard2.1: >>> x86 OCI VM shape, 1 OCPU / 2 hardware threads, about 14.5 GB RAM >>> >>> * VM.Standard.A1.Flex.2: >>> Arm/Ampere A1 flexible VM shape, 2 vCPU threads, about 10.9 GB RAM >>> >>> >>> The ResultsDB runs show the regression in the Sysbench threads metric: >>> >>> - VM.Standard2.1: 333 -> 236 (-29.1%) >>> - VM.Standard.A1.Flex.2: 1286 -> 1152 (-10.4%) >>> >>> A test kernel was built with f4c31b07b136 reverted and the performance >>> regression was recovered. >>> >>> From the code, it is possible the regression is due to the new >>> previous-wakeup heuristic in the special idle cases. Before the commit: >>> >>> - no cpuidle driver: >>> tick_nohz_idle_stop_tick() >>> default_idle_call() >>> >>> - one idle state: >>> tick_nohz_idle_retain_tick() >>> cpuidle state 0 >> I think that this is your case and the tick stops for you sometimes >> now while it had never stopped before. >> >> Can you confirm? >> >> Overall, it would be good to know the idle state lists for both the VM >> and the host. >> > +1, but also which HZ are you using? > Both systems reported have 2 logical CPUs then? Were higher core counts also > tested and how does it affect them? Hi Rafael, Christian, Thanks for the feedback! We will collect the requested runtime data from the affected systems before testing the proposed one-line change. For Christian’s HZ question, I can answer from the UEK kernel configs. The relevant configs for the affected kernel are: - x86_64 / VM.Standard2.1: - CONFIG_HZ=1000 - CONFIG_HZ_1000=y - CONFIG_NO_HZ_FULL=y - CONFIG_NO_HZ=y - CONFIG_CPU_IDLE=y - CONFIG_INTEL_IDLE=y - CONFIG_ACPI_PROCESSOR_IDLE=y - arm64 / VM.Standard.A1.Flex.2: - CONFIG_HZ=250 - CONFIG_HZ_250=y - CONFIG_NO_HZ_FULL=y - CONFIG_NO_HZ=y - CONFIG_CPU_IDLE=y - CONFIG_ACPI_PROCESSOR_IDLE=y - CONFIG_ARM_PSCI_CPUIDLE is not set The ResultsDB metadata for the two reported systems shows both had 2 logical CPUs: - VM.Standard2.1: 1 OCPU / 2 hardware threads - VM.Standard.A1.Flex.2: 2 vCPU threads I do not yet have higher-core-count results for this same comparison. I will check whether those runs exist, and if not, whether we can schedule them. Thanks, Joe