Re: [PATCH v2 13/20] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops
Pedro Falcato <[email protected]> Wed, 5 Aug 2026 13:50:58 +0100
| Newsgroups | gmane.linux.kernel.stable,gmane.linux.ports.arm.kernel |
|---|---|
| Message-ID | <[email protected]> |
On Wed, Aug 05, 2026 at 11:27:07AM +0100, David Laight wrote: > On Tue, 4 Aug 2026 23:45:56 +0100 > Pedro Falcato <[email protected]> wrote: > > > On Tue, Aug 04, 2026 at 06:04:56PM +0100, Mark Rutland wrote: > > > Currently arm64's this_cpu_*() ops transiently disable preemption in > > > order to guarantee that the address generation and memory access(es) > > > occur on the same CPU. > > > > > > Transiently disabling preemption can be expensive. When re-enabling > > > preemption it is necessary to make a conditional function call to > > > preempt_schedule[_notrace]() in order to handle the rare case that the > > > task needs to be rescheduled. The potential function call has a number > > > of negative effects on code generation (e.g. due to the need to create a > > > stack frame and spill registers), and the conditionality can result in > > > poor code generation and/or poor branch prediction. > > > > > > This patch adds infrastructure for a scheme where this_cpu_*() ops do > > > not need to transiently disable preemption, avoiding the negative > > > impacts described above. Individual operations will be converted in > > > subsequent patches. > > > > > > Each operation registers a critical section during which the exception > > > return code will adjust the offset and addresses if preemption occurs > > > mid-sequence. The critical section is registered/unregistered with a > > > small prologue and epilogue which encodes three distinct GPRRs (<pcp>, > > > <off>, <addr>) into a new thread_info::pcp_gprs field: > > > > > > // Prologue. Enable fixups for <off> and <addr>. > > > mrs <tsk>, sp_el0 > > > mov <tmp>, #__VAL_PCPU_GPRS(<pcp>, <off>, <addr>) > > > strh <tmp>, [<tsk>, #TSK_TI_PCPU_GPRS] > > > > > > // Generate cpu-specific address > > > mrs <off>, TPIDR_ELx > > > add <addr>, <pcp>, <off> > > > > > > // Perform access sequence > > > ldr <val>, [<addr>] > > > > > > // Epilogue. Disable fixups > > > strh wzr, [<tsk>, #TSK_TI_PCPU_GPRS] > > > > > > If an exception is taken from within the critical section, the exception > > > return code will adjust <off> to be the current CPU's offset, and will > > > adjust <addr> to be (<pcp> + <off>). Distinct registers are used for > > > <pcp>, <off>, and <addr>, so that the fixup can be applied safely at any > > > point during the critical section. > > > > > > To ensure that this_cpu_*() operations within exception handlers work > > > correctly and do not corrupt state, thread_info::pcpu_gprs is saved > > > into a new pt_regs::pcpu_gprs field upon exception entry, and restored > > > upon exception return. > > > > > > Looking at a simple this_cpu_operation: > > > > > > | void outline_this_cpu_add_u64(u64 __percpu *p, u64 v) > > > | { > > > | this_cpu_add(*p, v); > > > | } > ... > > I think I had an Interesting Idea(tm) while reading the per-cpu discussion > > in linux-mm. In case the 3 instruction preamble is too expensive: > > > > 1) Pass -ffixed-x18 (this natively conflicts with SHADOW_CALL_STACK. > > SHADOW_CALL_STACK is already not-optimal codegen wise, so maybe not a big deal). > > 2) arm64 kernel bits will use x18 as a cheap task flags register > > 3) #define TASK_KRSEQ (1 << 0) > > 4) Switching into the krseq mode is just a matter of toggling the bit in x18, so > > orr x18, x18, #TASK_KRSEQ > > a single instruction. > > 5) Switching off is just a matter of clearing the bit in x18, so: > > and x18, x18, #~TASK_KRSEQ > > 6) On the preempt side we keep the krseq tables in memory, and do a sort of lookup > > (binary search sounds easiest?) on them. But _only_ if x18 TASK_KRSEQ is set. > > This penalises unlucky preempts but keeps fast paths maximally fast. > > 7) entry points of course get to clear it after saving it > > > > The end result would look something like: > > | <outline_this_cpu_add_u64>: > > | orr x18, x18, #TASK_KRSEQ > > | mrs x4, tpidr_el1 > > | add x3, x0, x4 > > | 1: ldxr x6, [x3] > > | add x6, x6, x1 > > | stxr w5, x6, [x3] > > | cbnz w5, 1b > > | 2: > > | and x18, x18, #~TASK_KRSEQ > > | ret > > | .pushsection .data.krseq > > | .word 1b > > | .word 2b > > | .word whateverelse > > | .popsection ... Something I overlooked is that, since the solution shifted from "literally kernel rseq" to "funky kernel rseq but pcpu-specific" you don't probably don't need the table at all, as long as you dedicate a good few of those x18 bits to stash the pcp_gprs value. > > > > This of course precludes the use of x18 for the compiler, so it would > > require careful benchmarking in case it negatively affects codegen too much. > > But it avoids any sort of extraneous stores in the fast path. > > The extra stores are independent of the main instruction flow. > On a multi-issue (and especially out-of-order) cpu they are pretty much > likely to be noise. Oh, I agree, it is probably in the noise, modern uarchs are awesome :) > The biggest cost is likely to be in the I-cache and instruction decoders. > Put a memory read in the 'main' path and the few clocks needed for the > D-cache read are likely to dominate - so the writes to the pcp_gprs > are actually likely to be free. While I do like theorycrafting, I think we need numbers to know (numbers which I do not have, and seemingly no one seems to have for now). > > OTOH stealing a gpr for some flags will cost everwhere. Perhaps, but arm64 isn't exactly short on registers :) In any case, I generally agree with your take that it quite possibly doesn't matter, I was just throwing this out there in case it can help. -- Pedro