Re: [PATCH] sched/core: Don't pin the idle task in migrate_disable_switch()

Peter Zijlstra <[email protected]>
Newsgroups org.kernel.vger.stable,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
On Thu, Aug 06, 2026 at 11:01:32AM +0200, Krystian Slowik wrote:
> [Resend as plain text, the lists rejected my first attempt for HTML.]
> 
> > What is actually doing migrate_disable() here? Why would the idle thread
> > ever hit this...
> 
> Nothing, as far as I can tell. I went back into the vmcore and the
> counter didn't come from migrate_disable() at all:
> 
>   swapper/3->migration_disabled == 8      /* 0b1000 */
>   cpu3 rq->nr_pinned == 0
>   the other 15 idle tasks: 0 and 0
> 
> That pair can't come from the API. Surrounding bytes are clean, the
> whole anomaly is one set bit:
> 
>   swapper/2 +2416: 00 00 00 00 00 00 00 00  00 00 00 00 00 00 00 00
>   swapper/3 +2416: 00 00 00 00 00 00 00 00  08 00 00 00 00 00 00 00
>                    migration_pending        ^^ migration_disabled
> 
> A second dump from the same box two days later (rcu_cblist_dequeue,
> unrelated) has all 16 idle tasks at 0.
> 
> Correction to my patch description while I'm here: "two separate
> machines" was wrong, both idle-loop oopses are the same box.
> 
> Fleet context, since it changes how much you should care: three
> Ryzen 7 7840HS boxes, no ECC, DDR5 at JEDEC 5600, no EXPO. Same
> board, BIOS and microcode, identical apart from DIMM vendor and SSD.
> Taint 0 on all three, no out-of-tree modules ever, BPF is systemd
> cgroup boilerplate only. Two of the three have produced corruption
> oopses over five weeks (css_rstat_flush GPF twice at the same RIP
> plus a slab freelist on one, the rcu_cblist fault plus this bit flip
> on the other), and they were on different kernel builds at the time.
> So I can't separate a scribbler from a platform fault yet. DDR5
> on-die ECC makes marginal cells the less likely half, though it's
> SEC per burst so not excluded. Memtest next, with low expectations
> at 11 events in five weeks.

Oof, memory corruption is a pain.

> Which makes this not a fix but an assertion, and the stable Cc
> should go. If you want it at all:
> 
>     if (likely(!p->migration_disabled))
>         return;
> 
>     if (WARN_ON_ONCE(p == rq->idle))
>         return;
> 
> behind the existing early return, so the hot path doesn't pay for it
> and the state screams instead of oopsing under the rq lock.
> 
> Placement is a guess though. Anything reaching do_set_cpus_allowed()
> with p == rq->idle hits the same missing enqueue_task(), so
> sched_change_begin() might be the better home. And if you'd rather
> not paper over corruption in the scheduler at all, that's a fine
> answer too. The thread already answered the question I actually had.

Well, do_set_cpus_allowed() or any of the other sched_change users. I
think I'm leaning towards no change here, as you found we already get a
'nice' splat if this happens. Additionally, tripping a WARN inside
schedule() is commonly fatal all on its own (printk likes to do a
wakeup, which doesn't really work all that great from inside the
scheduler -- this is being worked on on the prink side).

Thanks!
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.