Re: [PATCH] sched/cache: honor migrate_llc_task semantics in active load balance
"Chen, Yu C" <[email protected]>
| Newsgroups | org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On 8/7/2026 1:22 AM, Tim Chen wrote:
> On Thu, 2026-08-06 at 23:35 +0800, Chen, Yu C wrote:
>> Hi Lu Wang, Tim,
>>
>> On 8/1/2026 8:22 PM, Lu Wang wrote:
>>> A passive load-balance pass marks group_llc_balance as migrate_llc_task
>>> and queues active balance when it cannot move a task. The CPU stopper
>>> callback constructs a fresh lb_env, so preserve the migration type on
>>> the runqueue across the asynchronous boundary.
>>>
>>> For CAS-directed active balance, reject a candidate whose preferred LLC
>>> does not match the destination LLC. This keeps the fallback from moving
>>> a task away from its preferred LLC.
>>>
>>
>> It looks like this proposal provides fine-grain control on per-task base
>> migration strategy is promising.
>>
>>> +static inline bool
>>> +migrate_llc_task_wrong_dst(struct task_struct *p, struct lb_env *env)
>>> +{
>>> + return sched_cache_enabled() &&
>>> + env->migration_type == migrate_llc_task &&
>>> + READ_ONCE(p->preferred_llc) != llc_id(env->dst_cpu);
>>> +}
>>> +
>>
>> [ ... ]
>>
>>> @@ -13654,6 +13672,7 @@ static int active_load_balance_cpu_stop(void *data)
>>> .src_rq = busiest_rq,
>>> .idle = CPU_IDLE,
>>> .flags = LBF_ACTIVE_LB,
>>> + .migration_type = (enum migration_type)busiest_rq->active_balance_type,
>>
>> If we overwrite migration_type for ALB (default is 0, i.e. migrate_load),
>> then in can_migrate_task() a delayed task might not be migrated in ALB:
>>
>> if ((p->se.sched_delayed) && (env->migration_type != migrate_load))
>
> This is a good catch.
>
> It may be easier to create a migrate_llc_task_alb type and pass that
> in migration type. Then modify the above as
>
> if ((p->se.sched_delayed) && env->migration_type != migrate_load
> && env->migration_type != migrate_llc_task_alb)
> return 0
>
We still need a channel to carry migrate_llc_task/migrate_llc_task_alb
into the alb, since the stopper builds a fresh lb_env - hence
rq->active_balance_type was introduced in Lu Wang's proposal. We can
reuse the callback slot in active_balance_work instead, pick
active_load_balance_llc_cpu_stop() at kick time, thus no new rq field
is needed.
thanks,
Chenyu