Re: [PATCH v6 5/6] sched/fair: Allow load balancing between CPUs of identical capacity
Vincent Guittot <[email protected]>
| Newsgroups | gmane.linux.kernel |
|---|---|
| Message-ID | <CAKfTPtBHNwuxua8F7G_U8mFd_kwOxJ9Q8SPFNgDEcwAkapgd2Q@mail.gmail.com> |
On Tue, 21 Jul 2026 at 04:33, Ricardo Neri <[email protected]> wrote: > > sched_balance_find_src_rq() avoids selecting a runqueue with a single > running task as busiest if doing so results in migrating the task to a > CPU with less than ~5% of extra capacity. It also unintentionally > prevents migrations between CPUs of identical capacity. > > When CONFIG_SCHED_CLUSTER is enabled, load should be balanced across > clusters of CPUs with the same capacity. Allowing migration between CPUs > of identical capacity is necessary to meet this goal. > > Use get_actual_cpu_capacity() to reflect architectural capacity as well > as diminished capacity due to hardware or cpufreq pressure. Guard this > check with the sched_cluster_active static key so that systems without > cluster topology are unaffected. > > Tested-by: Christian Loehle <[email protected]> > Tested-by: Andrea Righi <[email protected]> > Signed-off-by: Ricardo Neri <[email protected]> Reviewed-by: Vincent Guittot <[email protected]> > --- > Changes in v6: > * Switched to use get_actual_cpu_capacity() instead of > arch_scale_cpu_capacity(). The former considers rq->avg_hw.load_avg and > cpufreq_pressure and their impact on CPU capacity. (Vincent) > * Renamed the variable same_arch_cluster as cluster_equal_cap for > clarity. (Andrea) > * Added Tested-by tag from Andrea. Thanks! > > Changes in v5: > * Optimized logic to identify same-arch clusters only when needed. > * Added Tested-by tag from Christian. Thanks! > > Changes in v4: > * Implemented the check for cluster with a local variable for improved > readability. > > Changes in v3: > * Reverted the inverted capacity check; the inverted form incorrectly > allows migrations to CPUs of slightly less capacity. > * Guarded the check for architectural capacity with the > sched_cluster_active static key. > > Changes in v2: > * Used arch_scale_cpu_capacity() instead of capacity_of() to ignore > runtime variability. > * Inverted the check for runtime capacity. (Christian) > * Reworded patch description for clarity. > --- > kernel/sched/fair.c | 9 ++++++++- > 1 file changed, 8 insertions(+), 1 deletion(-) > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > index feea47e6abea..de4189b562ac 100644 > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -13104,13 +13104,20 @@ static struct rq *sched_balance_find_src_rq(struct lb_env *env, > */ > if (env->sd->flags & SD_ASYM_CPUCAPACITY && > nr_running == 1) { > + bool cluster_equal_cap = static_branch_unlikely(&sched_cluster_active) && > + (get_actual_cpu_capacity(env->dst_cpu) == > + get_actual_cpu_capacity(i)); > bool smt_degraded_cap = sched_smt_active() && !is_core_idle(i); > > /* > * Busy SMT siblings reduce the capacity of CPU @i. Do > * not skip it in this case. > + * > + * CONFIG_SCHED_CLUSTER requires balancing load across > + * clusters of identical capacity, accounting for > + * hardware and cpufreq pressure. > */ > - if (!smt_degraded_cap && > + if (!smt_degraded_cap && !cluster_equal_cap && > !capacity_greater(capacity_of(env->dst_cpu), capacity)) > continue; > } > > -- > 2.43.0 >