Re: [BUG] sched/fair: divide error in __calc_prop_weight() from the enqueue path (flat-hierarchy series)

Jake Steinman <[email protected]>
Newsgroups org.kernel.vger.cgroups,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
The diagnostic WARN fired. tg_cpus() does return 0 in practice, so the
cpuset route is real and not just theoretically reachable.

21.4 h into the boot, on the kernel carrying the three guards:

  sched: tg_cpus() == 0, empty cpuset
  WARNING: kernel/sched/fair.c:4904 at calc_concur_shares+0xf8/0x130, CPU#14: swapper/14/0
  CPU: 14 UID: 0 PID: 0 Comm: swapper/14
  Tainted: G     U   C OE       7.2.0-1-cachyos-rc #1 PREEMPT(full)
  Hardware name: Dell Inc. XPS 16 DA16260/0RMV2Y, BIOS 1.5.1 04/01/2026
  RIP: 0010:calc_concur_shares+0xfe/0x130
  RAX: ffffffff88eec4b0 RBX: 0000000000000000 RCX: 0000000000000000
  RDX: ffff8aedcc2379a0 RSI: 0000000000000030 RDI: ffffffff88eec4b0
  RBP: fffff3f0ff990cc0 R08: 0000000000000000 R09: 0000000000000000
  R10: 000000000000b74f R11: 0000000000000001 R12: fffff3f0ff98c5c0
  R13: 0000000000100000 R14: ffff8aed89796c00 R15: 0000000000000001
  Call Trace:
   <IRQ>
   enqueue_task_fair+0x201/0x970
   enqueue_task+0x8e/0x250
   ttwu_do_activate+0x89/0x1d0
   sched_ttwu_pending+0xce/0x1f0
   __flush_smp_call_function_queue+0x2b8/0x3e0
   __sysvec_call_function_single+0x22/0xb0
   fred_sysvec_call_function_single+0x54/0x70
   </IRQ>
   <TASK>
   fred_extint+0x4e/0x80
   asm_fred_entrypoint_kernel+0x41/0x70

Two things worth noting.

**The context is not the one that crashed.** The original oops was
bash in fork(), wake_up_new_task() -> enqueue_task_fair(). This is a
remote wakeup: an idle CPU (swapper/14) draining its pending-wakeup
queue from a CALL_FUNCTION_SINGLE IPI, in IRQ context. Same
calc_concur_shares() -> tg_cpus() == 0, reached a different way. So the
zero is a property of the task_group's cpuset at that moment, not of
either call path.

**IRQ context may matter for your RCU-race question.** You wrote:

> Now, cpuset_num_cpus() uses RCU, so perhaps there is a race somewhere.

This fired inside a hardirq, which is at least consistent with that --
though I can't distinguish "raced with an update" from "legitimately
empty at that instant" from the WARN alone. Waiman, if there is a
specific place you'd want instrumented to tell those apart, I have a
machine that reproduces it in about a day of ordinary use and I'm happy
to run whatever you want.

For completeness on what was and wasn't happening:

- Fired 21.4 h into the boot (wall clock 2026-08-19 13:24), during
  ordinary desktop use, no reproducer.
- One s2idle suspend/resume in that boot, which ended 4h20m earlier.
- No CPU hotplug events logged in the boot at all.
- No isolcpus, no domain isolation, no partition roots -- every
  cpuset.cpus.partition on the system reads "member", and no cgroup
  showed an empty cpuset.cpus.effective when I sampled afterwards.
- intel_lpmd is running and does rewrite AllowedCPUs on the top-level
  systemd slices, but logged nothing near the event, so I still can't
  pin it as the trigger. I mention it only because I raised it earlier
  and want to be clear it remains unproven.

The guard did its job: WARN_ONCE, floor to 1, machine still up 22 h
later with no divide error and no panic. That is 21 h longer than the
unguarded kernel managed on the day it crashed, though obviously one
data point.

Nothing here changes the patch I sent -- flooring tg_cpus() is still
right on symmetry grounds. It does mean the floor is load-bearing
rather than defensive, and that whatever lets a cpuset read as empty
here is a real thing worth understanding separately.

Thanks,
Jake
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.