Re: [PATCH v3] docs/sched_ext: document that cgroup CPU knobs are scheduler-dependent
Andrea Righi <[email protected]>
| Newsgroups | dev.linux.lists.sched-ext,org.kernel.vger.bpf,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <aowxJK_KgBLLnyiN@gpd4> |
Hi Tao, On Mon, Aug 24, 2026 at 05:15:01PM +0800, Tao Cui wrote: > From: Tao Cui <[email protected]> > > The fair class enforces cpu controller knobs such as cpu.max, > cpu.weight and cpu.idle in the kernel. sched_ext only passes them to > the BPF scheduler through the ops.cgroup_set_*() callbacks. Whether > and how a knob takes effect is up to the loaded scheduler: it may > implement the corresponding callback partially or not at all. The > same applies to other knobs like nice levels. > > Document this in the basics section so users and container > orchestrators know what to expect from a BPF scheduler. > > Signed-off-by: Tao Cui <[email protected]> > --- > v2 -> v3: Drop the scheduler list and the nr_throttled example, keep > the section concise and generic, per review. > > v2: https://lore.kernel.org/r/[email protected] > > Documentation/scheduler/sched-ext.rst | 12 ++++++++++++ > 1 file changed, 12 insertions(+) > > diff --git a/Documentation/scheduler/sched-ext.rst b/Documentation/scheduler/sched-ext.rst > index 35b550671ca7..b594d93dd6aa 100644 > --- a/Documentation/scheduler/sched-ext.rst > +++ b/Documentation/scheduler/sched-ext.rst > @@ -242,6 +242,18 @@ optional. The following modified excerpt is from > .name = "simple", > }; > > +Scheduler-Dependent Knobs > +------------------------- > + > +The fair class enforces cpu controller knobs such as ``cpu.max``, > +``cpu.weight`` and ``cpu.idle`` in the kernel. sched_ext only passes > +them to the BPF scheduler through ``ops.cgroup_set_weight()``, > +``ops.cgroup_set_idle()``, ``ops.cgroup_set_bandwidth()`` and friends. Existing cpu.weight and bandwidth values are delivered via ops.cgroup_init(), ops.cgroup_set_*() callbacks handle later changes. Moreover, cpu.idle state is not included in scx_cgroup_init_args, so apparently BPF schedulers don't receive an initial value (only subsequent writes). This should be probably fixed separately. > +Whether and how a knob takes effect is up to the loaded scheduler: it > +may implement the corresponding callback partially or not at all. The > +same applies to other knobs like nice levels. When relying on these > +knobs, check the documentation or source of the loaded scheduler. > + "nice levels" can be a bit ambiguous, per-task nice changes are converted to weights and reported through ops.set_weight(), writes to the cgroup file cpu.weight.nice are reported through ops.cgroup_set_weight(). Maybe we can rephrase the whole paragraph as following, or something along these lines: The fair-class scheduler enforces CPU controller settings such as cpu.max, cpu.weight, and cpu.idle. For sched_ext tasks, the scheduler core communicates these settings to the BPF scheduler through ops.cgroup_init() and reports subsequent changes through the corresponding ops.cgroup_set_*() callbacks. Similarly, per-task nice changes are converted to weights and reported through ops.set_weight(). Each BPF scheduler is responsible for implementing the scheduling semantics of these settings and may choose to ignore them. Consult the loaded scheduler's documentation before relying on these controls. Thanks, -Andrea