Re: [PATCH 1/5] sched_ext: Make exit claiming lock-free
Andrea Righi <[email protected]> Mon, 27 Jul 2026 21:35:12 +0200
| Newsgroups | dev.linux.lists.sched-ext,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <amey8FunQb8M7qwI@gpd4> |
Hi Tejun, On Fri, Jul 24, 2026 at 02:50:15PM -1000, Tejun Heo wrote: > scx_claim_exit() claims descendants' exits by walking the subtree under > scx_sched_lock, making exit claiming, and thus scx_error(), unusable from > NMI and from under scx_sched_lock. However, kfuncs raising errors can run > from NMI-attached BPF progs, the hardlockup handler runs in NMI, and > scx_link_sched() wants to report failures under the lock. > > The walk does two things with different urgencies: ->aborting must be > asserted synchronously to break IRQs-off dispatch-path live-locks, while the > descendants' exit_kind claims can happen later. Split them: sweep ->aborting > locklessly under RCU to unwedge the system and defer the locked > SCX_EXIT_PARENT walk to a new irq_work, both of which are NMI-safe. > > The sweep stores each node's ->aborting and then reads its children list > while scx_link_sched() inserts and then checks the parent's ->aborting, the > two sides paired by full barriers - one side always sees the other. A link > that sees ->aborting undoes its insert and fails. As the undo's > list_del_rcu() leaves ->sibling non-empty, list_empty() can no longer > identify a never-linked sched during teardown - add sch->linked instead. > > trace_sched_ext_exit can now fire from NMI. The exit backtrace is skipped > for NMI exits as stack_trace_save()'s NMI-safety is arch-dependent and > undocumented. > > Signed-off-by: Tejun Heo <[email protected]> > --- ... > diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c > index aca8d2380509..30ce4c9428cf 100644 > --- a/kernel/sched/ext/ext.c > +++ b/kernel/sched/ext/ext.c > @@ -5027,6 +5027,7 @@ static void scx_sched_free_rcu_work(struct work_struct *work) ... > /* > * Claim the exit on @sch. The caller must ensure that the helper kthread work > * is kicked before the current task can be preempted. Once exit_kind is > * claimed, scx_error() can no longer trigger, so if the current task gets > * preempted and the BPF scheduler fails to schedule it back, the helper work > * will never be kicked and the whole system can wedge. > + * > + * Lock-free and safe to call from any context including NMI. > */ > static bool scx_claim_exit(struct scx_sched *sch, enum scx_exit_kind kind) > { > @@ -6279,35 +6310,28 @@ static bool scx_claim_exit(struct scx_sched *sch, enum scx_exit_kind kind) > if (!atomic_try_cmpxchg(&sch->exit_kind, &none, kind)) > return false; > > - /* > - * Some CPUs may be trapped in the dispatch paths. Set the aborting > - * flag to break potential live-lock scenarios, ensuring we can > - * successfully reach scx_bypass(). > - */ > - WRITE_ONCE(sch->aborting, true); > - > trace_sched_ext_exit(sch, kind); > > - /* > - * Propagate exits to descendants immediately. Each has a dedicated > - * helper kthread and can run in parallel. While most of disabling is > - * serialized, running them in separate threads allows parallelizing > - * ops.exit(), which can take arbitrarily long prolonging bypass mode. > - * > - * To guarantee forward progress, this propagation must be in-line so > - * that ->aborting is synchronously asserted for all sub-scheds. The > - * propagation is also the interlocking point against sub-sched > - * attachment. See scx_link_sched(). > - * > - * This doesn't cause recursions as propagation only takes place for > - * non-propagation exits. > - */ > - if (kind != SCX_EXIT_PARENT) { > - scoped_guard (raw_spinlock_irqsave, &scx_sched_lock) { > - struct scx_sched *pos; > + if (kind == SCX_EXIT_PARENT) { > + /* an ancestor is already sweeping the subtree */ > + WRITE_ONCE(sch->aborting, true); > + } else { > + struct scx_sched *pos; > + > + /* > + * CPUs may be live-locked in the dispatch paths of @sch or its > + * descendants, which ->aborting breaks. Sweep the subtree > + * locklessly so that this works from NMI. smp_store_mb() orders > + * each node's ->aborting store before its children are walked - > + * either we see a racing scx_link_sched() on ->children or it > + * sees ->aborting. > + */ > + scoped_guard (rcu) { > scx_for_each_descendant_pre(pos, sch) > - scx_disable(pos, SCX_EXIT_PARENT); > + smp_store_mb(pos->aborting, true); > } > + > + irq_work_queue(&sch->propagate_exit_irq_work); > } Should we move the trace_sched_ext_exit() after this block? Before this patch, sch->aborting was set before the tracepoint. Now the tracepoint callbacks run before any scheduler in the subtree is marked aborting. Patch 4 also makes this path callable directly from the hardlockup NMI. While the tracepoint is running, other CPUs may remain in the live-lock-prone dispatch paths instead of observing ->aborting and backing out. Probably it doesn't really matter much, but this should help reduce the recovery window a bit. Thanks, -Andrea