Re: [PATCH v5 sched_ext/for-7.3 28/33] sched_ext: Route ops.update_idle() to sub-schedulers and re-notify owed scheds

Andrea Righi <[email protected]> Tue, 14 Jul 2026 08:18:20 +0200
Newsgroups dev.linux.lists.sched-ext,org.kernel.vger.linux-kernel
Message-ID <alXUrLY3WMdAyXIc@gpd4>
Hi Tejun,

On Thu, Jul 09, 2026 at 12:50:36PM -1000, Tejun Heo wrote:
> __scx_update_idle() notified only the root scheduler. A sub-scheduler that
> holds a cid needs that cid's idle state to place and kick on it.
> 
> Deliver ops.update_idle() to every scheduler that holds SCX_CAP_BASE on the
> transitioning cid. The root holds every cap, so a real transition always
> reaches it.
> 
> Real transitions are not enough on their own. A cid that is already idle
> when a sub-sched gains baseline access produces no transition, so the new
> holder would never learn it is idle. The ecaps sync arms a re-notify on the
> gain, and the next idle pick delivers ops.update_idle() to just that sched,
> leaving holders that already track the cpu untouched. A matching loss of
> baseline access drops any pending re-notify.
> 
> Bypass suppresses ops.update_idle() too, so a cpu that goes idle during a
> bypass window and stays idle yields no transition to re-deliver on
> un-bypass. Arm the same re-notify for every sched leaving bypass. The acute
> case is a child granted cids during its own ops.sub_attach(). The grant
> lands while the child is bypassed and the notify walk skips it, so on
> un-bypass it holds cids it never saw go idle. The root is owed the same and
> is armed through a separate per-rq flag, which keeps this working when
> sub-schedulers are compiled out.
> 
> v2: Gate the idle catch-up in pick_task_idle() to avoid a double ops.update_idle(). (sashiko AI)
> 
> Signed-off-by: Tejun Heo <[email protected]>
> ---

...

> +/*
> + * Notify schedulers of an idle transition on @cpu's cid, delivering to every
> + * sched that holds %SCX_CAP_BASE on the cid (the root holds every cap). A real
> + * transition (@do_notify) reaches all holders. A forced one (@root_renotify for
> + * the root, a sub-sched's idle_renotify marker for a sub) reaches only the owed
> + * scheds.
> + */
> +static void scx_idle_notify(struct rq *rq, bool idle, bool do_notify, bool root_renotify)
> +{
> +	s32 cpu = cpu_of(rq);
> +	s32 cid = scx_cpu_arg(cpu);
> +	struct scx_sched *pos;
> +
> +	lockdep_assert_rq_held(rq);
> +
> +	pos = scx_next_descendant_pre(NULL, scx_root);
> +	while (pos) {
> +		bool forced = false;
> +
> +		if (unlikely(scx_missing_caps(pos, cpu, SCX_CAP_BASE))) {
> +			pos = scx_skip_subtree_pre(pos, scx_root);
> +			continue;
> +		}
> +
> +		if (pos == scx_root) {
> +			forced = root_renotify;
> +		}
> +#ifdef CONFIG_EXT_SUB_SCHED
> +		else if (per_cpu_ptr(pos->pcpu, cpu)->idle_renotify) {
> +			per_cpu_ptr(pos->pcpu, cpu)->idle_renotify = false;
> +			forced = true;
> +		}
> +#endif
> +		if ((do_notify || forced) && SCX_HAS_OP(pos, update_idle) &&
> +		    !scx_bypassing(pos, cpu))
> +			SCX_CALL_OP(pos, update_idle, rq, cid, idle);
> +		pos = scx_next_descendant_pre(pos, scx_root);
> +	}
> +}

So, this makes every real idle/busy transition walk the whole scheduler
hierarchy and potentially invoke ops.update_idle() for every cap-holding
scheduler while the rq lock is held and IRQs are disabled?

This becomes O(number of cap-holding schedulers) BPF callbacks per idle
transition.

I haven't benchmarked this, so I'm not sure if it's a valid performance concern.
We don't have to fix this now, it can be a future improvement. In that case, if
we prove that we have a real bottleneck here, would it make sense to maintain a
per-rq list of schedulers that both hold effective SCX_CAP_BASE and implement
ops.update_idle()?

Also, for the root-only case, would it be worth keeping the old direct
ops.update_idle() path behind a static key which is enabled while any
sub-scheduler is attached?

Thanks,
-Andrea