Re: [REGRESSION] sched/idle: Sysbench threads regression after f4c31b07b136

Joseph Salisbury <[email protected]> Wed, 8 Jul 2026 11:25:02 -0400
Newsgroups dev.linux.lists.regressions,org.kernel.vger.linux-kernel,org.kernel.vger.linux-pm
Message-ID <[email protected]>

On 7/6/26 10:29 AM, Christian Loehle wrote:
> On 7/2/26 19:47, Rafael J. Wysocki (Intel) wrote:
>> Hi,
>>
>> On Thu, Jul 2, 2026 at 6:30 PM Joseph Salisbury
>> <[email protected]> wrote:
>>> Hi Rafael,
>>>
>>> We are seeing a reproducible MySQL Sysbench threads regression.  A
>>> bisect indicated the following commit as the first bad commit:
>>> f4c31b07b136 ("sched: idle: Consolidate the handling of two special cases")
>>>
>>> The regression was found in Oracle kernel performance testing on OCI VM
>>> shapes:
>>>
>>> VM Details:
>>> * VM.Standard2.1:
>>>         x86 OCI VM shape, 1 OCPU / 2 hardware threads, about 14.5 GB RAM
>>>
>>>    * VM.Standard.A1.Flex.2:
>>>         Arm/Ampere A1 flexible VM shape, 2 vCPU threads, about 10.9 GB RAM
>>>
>>>
>>> The ResultsDB runs show the regression in the Sysbench threads metric:
>>>
>>>     - VM.Standard2.1:       333 -> 236  (-29.1%)
>>>     - VM.Standard.A1.Flex.2: 1286 -> 1152 (-10.4%)
>>>
>>> A test kernel was built with f4c31b07b136 reverted and the performance
>>> regression was recovered.
>>>
>>>   From the code, it is possible the regression is due to the new
>>> previous-wakeup heuristic in the special idle cases.  Before the commit:
>>>
>>>     - no cpuidle driver:
>>>         tick_nohz_idle_stop_tick()
>>>         default_idle_call()
>>>
>>>     - one idle state:
>>>         tick_nohz_idle_retain_tick()
>>>         cpuidle state 0
>> I think that this is your case and the tick stops for you sometimes
>> now while it had never stopped before.
>>
>> Can you confirm?
>>
>> Overall, it would be good to know the idle state lists for both the VM
>> and the host.
>>
> +1, but also which HZ are you using?
> Both systems reported have 2 logical CPUs then? Were higher core counts also
> tested and how does it affect them?
Hi Rafael, Christian,

Thanks for the feedback!

We will collect the requested runtime data from the affected systems 
before testing the proposed one-line change.

For Christian’s HZ question, I can answer from the UEK kernel configs. 
The relevant configs for the affected kernel are:

- x86_64 / VM.Standard2.1:
     - CONFIG_HZ=1000
     - CONFIG_HZ_1000=y
     - CONFIG_NO_HZ_FULL=y
     - CONFIG_NO_HZ=y
     - CONFIG_CPU_IDLE=y
     - CONFIG_INTEL_IDLE=y
     - CONFIG_ACPI_PROCESSOR_IDLE=y

- arm64 / VM.Standard.A1.Flex.2:
     - CONFIG_HZ=250
     - CONFIG_HZ_250=y
     - CONFIG_NO_HZ_FULL=y
     - CONFIG_NO_HZ=y
     - CONFIG_CPU_IDLE=y
     - CONFIG_ACPI_PROCESSOR_IDLE=y
     - CONFIG_ARM_PSCI_CPUIDLE is not set

The ResultsDB metadata for the two reported systems shows both had 2 
logical CPUs:

  - VM.Standard2.1: 1 OCPU / 2 hardware threads
  - VM.Standard.A1.Flex.2: 2 vCPU threads

I do not yet have higher-core-count results for this same comparison.  I 
will check whether those runs exist, and if not, whether we can schedule 
them.

Thanks,

Joe