Re: [PATCH] sched_ext/scx_flatcg: expire cached hweights on weight changes

Tao Cui <[email protected]>
Newsgroups dev.linux.lists.sched-ext,org.kernel.vger.bpf,org.kernel.vger.linux-kernel
Message-ID <[email protected]>

在 2026/8/15 06:23, Tejun Heo 写道:
> Hello,
> 
> On Fri, Aug 14, 2026 at 10:48:35PM +0800, Tao Cui wrote:
>> Is this the intended behavior of the budget clamping, or should the
>> steady-state distribution converge to the compounded shares over time?
>> The header comment's model doesn't seem to hold in this scenario.
> 
> The clamp isn't the culprit. I reproduced your setup (3 busy tasks per
> leaf, 4 CPUs) and dug in.
> 
> The dominant factor is that D is runnable-task limited. At weight 800
> its compounded share is 3.55 CPUs but it only has 3 tasks, so 75% is
> the best it can do. Below that cap, the window granting loses more:
> whenever all three of D's tasks are already running and another CPU
> picks D, the pop from D's DSQ comes up empty, the cgv_node gets
> stashed and that CPU grants a full cgrp_slice_ns window to another
> cgroup. The loss scales with the window size. Measured D shares at
> weight 800:
> 
> 	3 tasks/leaf			55-59%
> 	3 tasks/leaf, 100ms slices	52%
> 	5 tasks/leaf			84-86%
> 
> Disabling the clamp entirely is the 59% above, so it barely matters.
> With enough runnable tasks per cgroup, the distribution converges to
> the documented compounding. The model holds, but only when no cgroup
> is runnable-task limited, and the window granting degrades sooner than
> per-task fair queueing would as that limit is approached. That's an
> inherent simplification of this example scheduler.
> 
> Separately, while digging into this, I found that the true-up in
> fcg_dispatch() is broken:
> 
> 	__sync_fetch_and_add(&cgc->cvtime_delta,
> 			     (cpuc->cur_at + cgrp_slice_ns - now) *
> 			     FCG_HWEIGHT_ONE / (cgc->hweight ?: 1));
> 
> In the CNS_EXPIRE case, now is past cur_at + cgrp_slice_ns, so the u64
> subexpression wraps. The multiplication preserves the two's complement
> encoding but the unsigned division by hweight destroys it, adding about
> 2^64 / hweight per expiry instead of a small correction. The sign is
> also inverted. The true-up should be actual minus charged, so the
> expiry overrun should be added and the CNS_EMPTY unused portion
> subtracted. Under saturation the budget clamp mostly masks the garbage,
> which is why the numbers above barely move with it fixed (the 86% in
> the 5 tasks/leaf row), but the accounting is broken all the same. The
> following fixes it and tests fine (BPF division is unsigned, keep the
> dividends positive):
> 
> 	s64 delta = now - cpuc->cur_at - cgrp_slice_ns;
> 
> 	if (delta >= 0)
> 		__sync_fetch_and_add(&cgc->cvtime_delta,
> 				     (u64)delta * FCG_HWEIGHT_ONE /
> 				     (cgc->hweight ?: 1));
> 	else
> 		__sync_fetch_and_sub(&cgc->cvtime_delta,
> 				     (u64)-delta * FCG_HWEIGHT_ONE /
> 				     (cgc->hweight ?: 1));
> 
> Care to send a patch?
> 
Hi, tejun

Thanks for the detailed walkthrough. I was chasing the clamp, but the
runnable-task limit and the window granting behavior explain it much
better, learned a lot from your analysis.

I sent a fix based on your suggestion, and saw it's been applied to
your tree:

  sched_ext/scx_flatcg: Fix cvtime true-up on slice overrun

  https://lore.kernel.org/all/[email protected]/

Kept your approach (s64 delta, keep dividends positive) and added
A/B numbers from the same VM setup in the commit message.

Looking forward to more of this while digging through sched_ext.

Thanks,
Tao

> Thanks.
> 
> --
> tejun
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.