Re: [PATCH v5 sched_ext/for-7.3 28/33] sched_ext: Route ops.update_idle() to sub-schedulers and re-notify owed scheds
Andrea Righi <[email protected]> Tue, 14 Jul 2026 08:18:20 +0200
| Newsgroups | dev.linux.lists.sched-ext,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <alXUrLY3WMdAyXIc@gpd4> |
Hi Tejun, On Thu, Jul 09, 2026 at 12:50:36PM -1000, Tejun Heo wrote: > __scx_update_idle() notified only the root scheduler. A sub-scheduler that > holds a cid needs that cid's idle state to place and kick on it. > > Deliver ops.update_idle() to every scheduler that holds SCX_CAP_BASE on the > transitioning cid. The root holds every cap, so a real transition always > reaches it. > > Real transitions are not enough on their own. A cid that is already idle > when a sub-sched gains baseline access produces no transition, so the new > holder would never learn it is idle. The ecaps sync arms a re-notify on the > gain, and the next idle pick delivers ops.update_idle() to just that sched, > leaving holders that already track the cpu untouched. A matching loss of > baseline access drops any pending re-notify. > > Bypass suppresses ops.update_idle() too, so a cpu that goes idle during a > bypass window and stays idle yields no transition to re-deliver on > un-bypass. Arm the same re-notify for every sched leaving bypass. The acute > case is a child granted cids during its own ops.sub_attach(). The grant > lands while the child is bypassed and the notify walk skips it, so on > un-bypass it holds cids it never saw go idle. The root is owed the same and > is armed through a separate per-rq flag, which keeps this working when > sub-schedulers are compiled out. > > v2: Gate the idle catch-up in pick_task_idle() to avoid a double ops.update_idle(). (sashiko AI) > > Signed-off-by: Tejun Heo <[email protected]> > --- ... > +/* > + * Notify schedulers of an idle transition on @cpu's cid, delivering to every > + * sched that holds %SCX_CAP_BASE on the cid (the root holds every cap). A real > + * transition (@do_notify) reaches all holders. A forced one (@root_renotify for > + * the root, a sub-sched's idle_renotify marker for a sub) reaches only the owed > + * scheds. > + */ > +static void scx_idle_notify(struct rq *rq, bool idle, bool do_notify, bool root_renotify) > +{ > + s32 cpu = cpu_of(rq); > + s32 cid = scx_cpu_arg(cpu); > + struct scx_sched *pos; > + > + lockdep_assert_rq_held(rq); > + > + pos = scx_next_descendant_pre(NULL, scx_root); > + while (pos) { > + bool forced = false; > + > + if (unlikely(scx_missing_caps(pos, cpu, SCX_CAP_BASE))) { > + pos = scx_skip_subtree_pre(pos, scx_root); > + continue; > + } > + > + if (pos == scx_root) { > + forced = root_renotify; > + } > +#ifdef CONFIG_EXT_SUB_SCHED > + else if (per_cpu_ptr(pos->pcpu, cpu)->idle_renotify) { > + per_cpu_ptr(pos->pcpu, cpu)->idle_renotify = false; > + forced = true; > + } > +#endif > + if ((do_notify || forced) && SCX_HAS_OP(pos, update_idle) && > + !scx_bypassing(pos, cpu)) > + SCX_CALL_OP(pos, update_idle, rq, cid, idle); > + pos = scx_next_descendant_pre(pos, scx_root); > + } > +} So, this makes every real idle/busy transition walk the whole scheduler hierarchy and potentially invoke ops.update_idle() for every cap-holding scheduler while the rq lock is held and IRQs are disabled? This becomes O(number of cap-holding schedulers) BPF callbacks per idle transition. I haven't benchmarked this, so I'm not sure if it's a valid performance concern. We don't have to fix this now, it can be a future improvement. In that case, if we prove that we have a real bottleneck here, would it make sense to maintain a per-rq list of schedulers that both hold effective SCX_CAP_BASE and implement ops.update_idle()? Also, for the root-only case, would it be worth keeping the old direct ops.update_idle() path behind a static key which is enabled while any sub-scheduler is attached? Thanks, -Andrea