Re: [PATCH] sched/cache: honor migrate_llc_task semantics in active load balance

wanglu15 <[email protected]>
Newsgroups org.kernel.vger.linux-kernel
Message-ID <[email protected]>
From: Lu Wang <[email protected]>

Thanks, Tim.

On Tue, 2026-08-04 at 12:42 -0700, Tim Chen wrote:
> On Tue, 2026-08-04 at 23:07 +0800, Lu Wang wrote:
> > Thanks, Chenyu.
> >
> > On Tue, 2026-08-04 at 16:17 +0800, Chen, Yu C wrote:
> > > Yes. Besides, if I understand correctly, I suppose Lu Wang was
> > > referring to the following scenario:
> > >
> > > src_rq has 2 runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc),
> > > while p2 prefers src_rq (src_llc). In this case, migrate_llc_task is
> > > set because src_rq has at least one task, p1, that wants to migrate
> > > to dst_rq. In ALB, can_migrate_task() found p2 and returns true for p2
> > > thus moves p2 out of its preferred LLC.
> >
> > That's exactly the scenario I had in mind.
> >
> > > Firstly, before ALB is triggered, the generic (passive) load balance is
> > > triggered. It iterates over p1 and p2 on src_rq to see if it can move any
> > > one of them to dst_rq, and in most cases it succeeds in moving p1 to
> > > dst_cpu. As a result, ALB will not be triggered.
> >
> > My question is whether p1 is guaranteed to be moved out in passive
> > LB. can_migrate_task()/migrate_degrades_llc() can reject p1 for
> > several independent reasons — p1 pinned by cpus_ptr, p1 cache-hot
> > with nr_balance_failed still below cache_nice_tries, or
> > can_migrate_llc_task() returning something other than mig_forbid due
> > to capacity constraints on dst_llc at that instant. If passive LB
> > rejects p1 for any of these, ALB is still triggered with p1 and p2
> > both present on src_rq.
> >
> > Can we conclude that p1 and p2 never end up on src_rq together when
> > ALB fires? Or would it help to set up a simple experiment and trace
> > this path to see whether it actually occurs in practice?
>
> With 2 tasks on rq with different preference, active load balance could
> pick the wrong task as can_migrate_task() checked in active load balance
> will not consult migrate_degrades_llc().  How about the following patch
> to fix this issue.
>
> [...]
>
> + if (env->migration_type == migrate_llc_task &&
> +     env->src_rq->cfs.h_nr_runnable > 1)
> +         return true;
> +
>   return false;
> }

Your approach is simpler than mine — it avoids threading
migration_type across the CPU stopper boundary and doesn't need any
new rq field.

One thing I'd like to flag, IMO: this approach skips the ALB path
entirely for migrate_llc_task whenever more than one task is
runnable, deferring the fix to the next passive LB pass. So it
trades "delay" for a simpler implementation.

Wang
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.