Re: [PATCH v11 05/12] sched/core: Try to use a preferred CPU in is_cpu_allowed

Shrikanth Hegde <[email protected]>
Newsgroups dev.linux.lists.virtualization,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
Hi Dietmar, Vincent

On 8/28/26 4:08 PM, Vincent Guittot wrote:
> On Fri, 28 Aug 2026 at 09:51, Dietmar Eggemann <[email protected]> wrote:
>>
>> On 25.08.26 12:38, Shrikanth Hegde wrote:
>>> When possible, try to choose a preferred CPU.
>>>
>>> This is essential to maintain user affinities when preferred
>>> CPUs change. A task pinned on a non-preferred CPU should continue
>>> to run there, since this is a non-user triggered event.
>>>
>>> If a CPU is non-preferred and the task can run on other CPUs which are
>>> currently preferred, then choose a preferred CPU instead.
>>> This is decided by checking if cpus_ptr and cpu_preferred_mask
>>> intersect or not. If yes, then the task has other preferred CPUs.
>>>
>>> The push task mechanism uses a stopper thread which calls
>>> select_fallback_rq() and uses this mechanism to pick a preferred CPU.
>>>
>>> This takes care of the wakeup path for FAIR tasks too.
>>> is_cpu_allowed() is called to ensure wakeups happen on preferred CPUs.
>>> With that, additional checks in available_idle_cpu() are not necessary.
>>>
>>> Ignore the preferred CPU state if a task's affinity is changing and
>>> its new mask no longer includes the CPU it is currently running on.
>>> This ensures migration_cpu_stop() does not abort, preventing the task
>>> from being stranded outside its allowed affinity.
>>>
>>> Account for tasks with architecture-specific CPU masks
>>> (e.g., 32-bit tasks on arm64). For such tasks, explicitly check against
>>> the arch-allowed CPUs to determine if any of the preferred CPUs are
>>> actually valid.
>>>
>>> For the majority of cases, this would still keep select_fallback_rq()
>>> as O(N). cpumask_intersects(), which is O(N), is called only if
>>> !cpu_preferred. The task running there is expected to move out.
>>> Subsequently, it should run on a preferred CPU. This becomes O(N**2)
>>> only for tasks pinned solely to non-preferred CPUs. That is a rare case.
>>>
>>> Overhead is minimal when the CPU is preferred.
>>>
>>> Signed-off-by: Shrikanth Hegde <[email protected]>
>>> ---
>>>   kernel/sched/core.c | 41 +++++++++++++++++++++++++++++++++++++++--
>>>   1 file changed, 39 insertions(+), 2 deletions(-)
>>>
>>> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
>>> index a45f7c308329..f71317fe281d 100644
>>> --- a/kernel/sched/core.c
>>> +++ b/kernel/sched/core.c
>>> @@ -2494,6 +2494,35 @@ static inline bool rq_has_pinned_tasks(struct rq *rq)
>>>        return rq->nr_pinned;
>>>   }
>>>
>>> +static inline bool task_can_sched_on_preferred(int cpu, struct task_struct *p)
>>> +{
>>> +     const struct cpumask *valid_mask;
>>> +     int i;
>>> +
>>> +     if (cpu_preferred(cpu))
>>> +             return false;
>>> +
>>> +     /* Only FAIR tasks honor preferred CPU state */
>>> +     if (unlikely(p->sched_class != &fair_sched_class))
>>> +             return false;
>>> +
>>> +     /* Ignore preferred state if task affinity is changing */
>>> +     if (unlikely(!cpumask_test_cpu(task_cpu(p), p->cpus_ptr)))
>>> +             return false;
>>> +
>>> +     valid_mask = task_cpu_possible_mask(p);
>>> +     if (likely(valid_mask == cpu_possible_mask))
>>> +             return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
>>> +
>>> +     /* Tasks with arch-specific CPU masks. e.g. 32-bit tasks on arm64. */
>>> +     for_each_cpu_and(i, p->cpus_ptr, cpu_preferred_mask) {
>>> +             if (cpumask_test_cpu(i, valid_mask))
>>> +                     return true;
>>> +     }
>>
>> Looking more into this, there might be a window in 64-32-bit execve()
>> for 32bit EL0 tasks on Arm64 (w/ allow_mismatched_32bit_el0 command line
>> option).
>>
>> The time before arch_setup_new_exec() calls
>> force_compatible_cpus_allowed_ptr() to restrict CPU affinity for those
>> tasks.
>>
>> Let me run more test on this ...
>>
>> Why not simply:
>>
>> - return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
>> + return cpumask_first_and_and(p->cpus_ptr, cpu_preferred_mask,
>> +                              task_cpu_possible_mask(p)) < nr_cpu_ids;
> 
> +1
> This is the best way to check that there is a valid cpu
> 
>>

Thanks Dietmar, Vincent for checking it further.

It is good to know that there might be a very narrow window.
The three-way cpumask check you suggested will handle that case safely.


>> IMHO, you want to know whether there is at least one CPU that belongs to
>> all three CPU masks?
>>
>> [...]

With that task_can_sched_on_preferred essentially becomes this.
I will put that in v12 and probably send it out after 7.3-rc1 lands.


static inline bool task_can_sched_on_preferred(int cpu, struct task_struct *p)
{
         if (cpu_preferred(cpu))
                 return false;

         /* Only FAIR tasks honor preferred CPU state */
         if (unlikely(p->sched_class != &fair_sched_class))
                 return false;

         /* Ignore preferred state if task affinity is changing */
         if (unlikely(!cpumask_test_cpu(task_cpu(p), p->cpus_ptr)))
                 return false;

         return cpumask_first_and_and(p->cpus_ptr, cpu_preferred_mask,
                              task_cpu_possible_mask(p)) < nr_cpu_ids;
}


Thanks again for checking and for the suggested code!
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.