Re: [PATCH] sched/cache: honor migrate_llc_task semantics in active load balance
wanglu15 <[email protected]>
| Newsgroups | org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
From: Lu Wang <[email protected]> Thanks, Tim. On Tue, 2026-08-04 at 12:42 -0700, Tim Chen wrote: > On Tue, 2026-08-04 at 23:07 +0800, Lu Wang wrote: > > Thanks, Chenyu. > > > > On Tue, 2026-08-04 at 16:17 +0800, Chen, Yu C wrote: > > > Yes. Besides, if I understand correctly, I suppose Lu Wang was > > > referring to the following scenario: > > > > > > src_rq has 2 runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc), > > > while p2 prefers src_rq (src_llc). In this case, migrate_llc_task is > > > set because src_rq has at least one task, p1, that wants to migrate > > > to dst_rq. In ALB, can_migrate_task() found p2 and returns true for p2 > > > thus moves p2 out of its preferred LLC. > > > > That's exactly the scenario I had in mind. > > > > > Firstly, before ALB is triggered, the generic (passive) load balance is > > > triggered. It iterates over p1 and p2 on src_rq to see if it can move any > > > one of them to dst_rq, and in most cases it succeeds in moving p1 to > > > dst_cpu. As a result, ALB will not be triggered. > > > > My question is whether p1 is guaranteed to be moved out in passive > > LB. can_migrate_task()/migrate_degrades_llc() can reject p1 for > > several independent reasons — p1 pinned by cpus_ptr, p1 cache-hot > > with nr_balance_failed still below cache_nice_tries, or > > can_migrate_llc_task() returning something other than mig_forbid due > > to capacity constraints on dst_llc at that instant. If passive LB > > rejects p1 for any of these, ALB is still triggered with p1 and p2 > > both present on src_rq. > > > > Can we conclude that p1 and p2 never end up on src_rq together when > > ALB fires? Or would it help to set up a simple experiment and trace > > this path to see whether it actually occurs in practice? > > With 2 tasks on rq with different preference, active load balance could > pick the wrong task as can_migrate_task() checked in active load balance > will not consult migrate_degrades_llc(). How about the following patch > to fix this issue. > > [...] > > + if (env->migration_type == migrate_llc_task && > + env->src_rq->cfs.h_nr_runnable > 1) > + return true; > + > return false; > } Your approach is simpler than mine — it avoids threading migration_type across the CPU stopper boundary and doesn't need any new rq field. One thing I'd like to flag, IMO: this approach skips the ALB path entirely for migrate_llc_task whenever more than one task is runnable, deferring the fix to the next passive LB pass. So it trades "delay" for a simpler implementation. Wang