Re: [PATCH RFC] sched_ext: warn when cpu.max is set but the BPF scheduler doesn't implement bandwidth control
Tao Cui <[email protected]>
| Newsgroups | dev.linux.lists.sched-ext,org.kernel.vger.bpf,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
Hi, tejun 在 2026/8/18 21:53, Tao Cui 写道: > From: Tao Cui <[email protected]> > > The kernel stores cpu.max bandwidth parameters in the task_group and > passes them to the BPF scheduler via ops.cgroup_set_bandwidth() and > scx_cgroup_init_args, but does not enforce the quota itself. If the > loaded BPF scheduler doesn't implement the callback, cpu.max is > silently ignored -- the cgroup gets unlimited CPU regardless of the > configured quota. > > Of the example schedulers, only scx_qmap implements the callback -- > and only to bpf_printk() the parameters, so no in-tree scheduler > actually enforces the quota. Measured with scx_simple: a > cgroup with cpu.max = "50000 100000" (50% of one CPU) and one > busy task used 9946ms of CPU in 10 seconds with nr_throttled > remaining 0. > > Print a one-time warning when a finite quota is configured on a > cgroup while the active scheduler lacks the callback, so users and > container orchestrators know the quota is not enforced. > Some background on how I found this: I was testing how sched_ext interacts with cgroup CPU controls in a VM, and the cpu.max case stood out: cgroup with cpu.max = "50000 100000" (50% of one CPU), one busy task, scx_simple loaded sched_ext : 9946ms of CPU in 10s, nr_throttled = 0 CFS : ~5000ms in 10s, throttling as expected The same happens with scx_flatcg and scx_central -- neither they nor scx_simple implement ops.cgroup_set_bandwidth(), so the quota is silently ignored. grep shows scx_qmap is the only in-tree user of the callback -- and its implementation just bpf_printk()s the parameters, so even there the quota is not enforced. Outside the tree, lavd implements its own bandwidth accounting, but as far as I can tell rusty and bpfland don't, which suggests users on those schedulers are running containers with cpu.max that does nothing. That's why I drafted the warning patch -- but I'm not sure a warning is the right approach. Some alternatives I can think of: 1. pr_warn_once() as in this patch (minimal, but the quota is still not enforced) 2. refuse to enable sched_ext (or the cgroup support) when a finite quota exists and the callback is missing 3. kernel-side fallback enforcement, e.g. throttle in scx_next_task_picked()/dispatch path based on tg->scx.bw_quota_us Is the missing enforcement considered the BPF scheduler's responsibility by design (and just under-documented), or would a kernel fallback be welcome? If it's the former, maybe sched-ext.rst should mention that cpu.max requires ops.cgroup_set_bandwidth() from the loaded scheduler. Happy to work on whichever direction you prefer. Thanks, Tao > Signed-off-by: Tao Cui <[email protected]> > --- > kernel/sched/ext/ext.c | 6 ++++++ > 1 file changed, 6 insertions(+) > > diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c > index 10af28a9f2c0..de786d0b928a 100644 > --- a/kernel/sched/ext/ext.c > +++ b/kernel/sched/ext/ext.c > @@ -4953,6 +4953,12 @@ void scx_group_set_bandwidth(struct task_group *tg, > tg->scx.bw_burst_us != burst_us)) > SCX_CALL_OP(sch, cgroup_set_bandwidth, NULL, > tg_cgrp(tg), period_us, quota_us, burst_us); > + else if (scx_cgroup_enabled && sch && > + !SCX_HAS_OP(sch, cgroup_set_bandwidth) && > + quota_us != RUNTIME_INF) > + pr_warn_once("sched_ext: BPF scheduler \"%s\" does not implement " > + "ops.cgroup_set_bandwidth(); cpu.max will not be enforced\n", > + sch->ops.name); > > tg->scx.bw_period_us = period_us; > tg->scx.bw_quota_us = quota_us;