Re: [PATCH v9 02/11] cpumask: Introduce cpu_preferred_mask
Shrikanth Hegde <[email protected]> Sat, 25 Jul 2026 08:46:23 +0530
| Newsgroups | dev.linux.lists.virtualization,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On 7/25/26 1:30 AM, Yury Norov wrote: > On Fri, Jul 24, 2026 at 07:37:23PM +0530, Shrikanth Hegde wrote: >> Provide cpu_preferred_mask infrastructure. Define get/set macros >> which could be used to get/set CPU state as preferred. >> >> PREFERRED_CPU config will be selected by the driver which handles >> steal time values. It is going to set/clear preferred CPU state. >> This driver will be called steal_governor and it is introduced in >> subsequent patches. It periodically samples the steal time and >> decides on preferred CPU state. >> >> A CPU is set to preferred when it becomes active. Later it may be >> marked as non-preferred depending on steal time values with >> steal_governor being enabled. >> >> Always maintain design construct of preferred is subset of active. >> i.e. preferred ⊆ active ⊆ online ⊆ present ⊆ possible >> >> With PREFERRED_CPU=n, ensure set_cpu_preferred is a nop and get >> method returns the active state in that case. >> >> Signed-off-by: Shrikanth Hegde <[email protected]> >> --- >> include/linux/cpumask.h | 24 ++++++++++++++++++++++++ >> kernel/Kconfig.preempt | 4 ++++ >> kernel/cpu.c | 6 ++++++ >> kernel/sched/core.c | 5 +++++ >> 4 files changed, 39 insertions(+) >> >> diff --git a/include/linux/cpumask.h b/include/linux/cpumask.h >> index d3cda0544954..34d08a3d80e1 100644 >> --- a/include/linux/cpumask.h >> +++ b/include/linux/cpumask.h >> @@ -122,12 +122,20 @@ extern struct cpumask __cpu_enabled_mask; >> extern struct cpumask __cpu_present_mask; >> extern struct cpumask __cpu_active_mask; >> extern struct cpumask __cpu_dying_mask; >> + >> +#ifdef CONFIG_PREFERRED_CPU >> +extern struct cpumask __cpu_preferred_mask; >> +#else >> +#define __cpu_preferred_mask __cpu_active_mask >> +#endif >> + >> #define cpu_possible_mask ((const struct cpumask *)&__cpu_possible_mask) >> #define cpu_online_mask ((const struct cpumask *)&__cpu_online_mask) >> #define cpu_enabled_mask ((const struct cpumask *)&__cpu_enabled_mask) >> #define cpu_present_mask ((const struct cpumask *)&__cpu_present_mask) >> #define cpu_active_mask ((const struct cpumask *)&__cpu_active_mask) >> #define cpu_dying_mask ((const struct cpumask *)&__cpu_dying_mask) >> +#define cpu_preferred_mask ((const struct cpumask *)&__cpu_preferred_mask) >> >> extern atomic_t __num_online_cpus; >> extern unsigned int __num_possible_cpus; >> @@ -1164,6 +1172,12 @@ void init_cpu_possible(const struct cpumask *src); >> #define set_cpu_active(cpu, active) assign_cpu((cpu), &__cpu_active_mask, (active)) >> #define set_cpu_dying(cpu, dying) assign_cpu((cpu), &__cpu_dying_mask, (dying)) >> >> +#ifdef CONFIG_PREFERRED_CPU >> +#define set_cpu_preferred(cpu, preferred) assign_cpu((cpu), &__cpu_preferred_mask, (preferred)) >> +#else >> +#define set_cpu_preferred(cpu, preferred) do { } while (0) >> +#endif >> + >> void set_cpu_online(unsigned int cpu, bool online); >> void set_cpu_possible(unsigned int cpu, bool possible); >> >> @@ -1258,6 +1272,11 @@ static __always_inline bool cpu_dying(unsigned int cpu) >> return cpumask_test_cpu(cpu, cpu_dying_mask); >> } >> >> +static __always_inline bool cpu_preferred(unsigned int cpu) >> +{ >> + return cpumask_test_cpu(cpu, cpu_preferred_mask); >> +} >> + >> #else >> >> #define num_online_cpus() 1U >> @@ -1296,6 +1315,11 @@ static __always_inline bool cpu_dying(unsigned int cpu) >> return false; >> } >> >> +static __always_inline bool cpu_preferred(unsigned int cpu) >> +{ >> + return cpu == 0; >> +} >> + >> #endif /* NR_CPUS > 1 */ >> >> #define cpu_is_offline(cpu) unlikely(!cpu_online(cpu)) >> diff --git a/kernel/Kconfig.preempt b/kernel/Kconfig.preempt >> index 88c594c6d7fc..de789b274ba3 100644 >> --- a/kernel/Kconfig.preempt >> +++ b/kernel/Kconfig.preempt >> @@ -192,3 +192,7 @@ config SCHED_CLASS_EXT >> For more information: >> Documentation/scheduler/sched-ext.rst >> https://github.com/sched-ext/scx >> + >> +config PREFERRED_CPU >> + bool >> + depends on SMP && PARAVIRT >> diff --git a/kernel/cpu.c b/kernel/cpu.c >> index b3c8553d7bd6..376d297a6292 100644 >> --- a/kernel/cpu.c >> +++ b/kernel/cpu.c >> @@ -3103,6 +3103,11 @@ EXPORT_SYMBOL(__cpu_dying_mask); >> atomic_t __num_online_cpus __read_mostly; >> EXPORT_SYMBOL(__num_online_cpus); >> >> +#ifdef CONFIG_PREFERRED_CPU >> +struct cpumask __cpu_preferred_mask __read_mostly; >> +EXPORT_SYMBOL_GPL(__cpu_preferred_mask); >> +#endif >> + >> void init_cpu_present(const struct cpumask *src) >> { >> cpumask_copy(&__cpu_present_mask, src); >> @@ -3160,6 +3165,7 @@ void __init boot_cpu_init(void) >> /* Mark the boot cpu "present", "online" etc for SMP and UP case */ >> set_cpu_online(cpu, true); >> set_cpu_active(cpu, true); >> + set_cpu_preferred(cpu, true); >> set_cpu_present(cpu, true); >> set_cpu_possible(cpu, true); >> >> diff --git a/kernel/sched/core.c b/kernel/sched/core.c >> index 2e7cde033a31..a45f7c308329 100644 >> --- a/kernel/sched/core.c >> +++ b/kernel/sched/core.c >> @@ -8690,6 +8690,9 @@ int sched_cpu_activate(unsigned int cpu) >> */ >> sched_set_rq_online(rq, cpu); >> >> + /* preferred is subset of active and follows its state */ >> + set_cpu_preferred(cpu, true); >> + >> return 0; >> } >> >> @@ -8703,6 +8706,8 @@ int sched_cpu_deactivate(unsigned int cpu) >> if (ret) >> return ret; >> >> + set_cpu_preferred(cpu, false); >> + > > Is it possible that this CPU would be the last preferred CPU in the > system? If so, you'll make the preferred mask empty. > Possible case is, say there are 80 CPUs and all CPUs are part of housekeeping. driver marked 40-80 as non-preferred and before driver gets a chance to run again, user disabled 0-39. Now preferred mask is empty. if steal time is low, it might recover without check broken in the next sampling, but it stays in between or high, then that check is broken. I don't think there is any side effect in core mechanism since is_cpu_allowed will pass due to empty preferred mask. In driver, further reduction will not happen. But yes, it will break the design checks. > In v9 you disabled integrity check while the steal time is withing the > threshold, so this condition may stay undetected quite a long. > I think simplest solution is do the design checks always and restore the preferred state if such case happens. I.e drop the optimization that was done in v9 compared to v8. > Can you add another integrity check here? If you're going to remove > the last preferred CPU, you need to force-enable some alternative. > Something like: > > if (cpumask_nth(1, cpu_preferred_mask) >= nr_cpu_ids) { > new_cpu = cpumask_any_andnot_but(cpu_active_mask, cpu_preferred_mask, cpu); > if (!WARN_ON(new_cpu >= nr_cpu_ids)) > set_cpu_preferred(new_cpu); > } > > set_cpu_preferred(cpu, false); > I think we shouldn't do such change. The design constraints are of driver. Hotplug mechanism just ensure to set preferred after setting active and clear preferred before clearing the active. That's all. Driver runs only once in 100ms at the very least and enforcing design checks of driver into core hotplug/scheduler mechanism is not right IMHO. It should be the role of driver to either actively recover or gracefully shut. That is user triggered edge case, i think simplest solution is gracefully shut the driver and let user to load the driver again. Always run the design checks. I can add this corner case details to the driver change log. What do you think?