Re: [PATCH v2 13/20] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops

Pedro Falcato <[email protected]>
Newsgroups org.kernel.vger.stable,org.infradead.lists.linux-arm-kernel
Message-ID <[email protected]>
On Wed, Aug 05, 2026 at 11:27:07AM +0100, David Laight wrote:
> On Tue, 4 Aug 2026 23:45:56 +0100
> Pedro Falcato <[email protected]> wrote:
> 
> > On Tue, Aug 04, 2026 at 06:04:56PM +0100, Mark Rutland wrote:
> > > Currently arm64's this_cpu_*() ops transiently disable preemption in
> > > order to guarantee that the address generation and memory access(es)
> > > occur on the same CPU.
> > > 
> > > Transiently disabling preemption can be  expensive. When re-enabling
> > > preemption it is necessary to make a conditional function call to
> > > preempt_schedule[_notrace]() in order to handle the rare case that the
> > > task needs to be rescheduled. The potential function call has a number
> > > of negative effects on code generation (e.g. due to the need to create a
> > > stack frame and spill registers), and the conditionality can result in
> > > poor code generation and/or poor branch prediction.
> > > 
> > > This patch adds infrastructure for a scheme where this_cpu_*() ops do
> > > not need to transiently disable preemption, avoiding the negative
> > > impacts described above. Individual operations will be converted in
> > > subsequent patches.
> > > 
> > > Each operation registers a critical section during which the exception
> > > return code will adjust the offset and addresses if preemption occurs
> > > mid-sequence. The critical section is registered/unregistered with a
> > > small prologue and epilogue which encodes three distinct GPRRs (<pcp>,
> > > <off>, <addr>) into a new thread_info::pcp_gprs field:
> > > 
> > >          // Prologue. Enable fixups for <off> and <addr>.
> > >          mrs	<tsk>, sp_el0
> > >          mov	<tmp>, #__VAL_PCPU_GPRS(<pcp>, <off>, <addr>)
> > >          strh	<tmp>, [<tsk>, #TSK_TI_PCPU_GPRS]
> > > 
> > >          // Generate cpu-specific address
> > >          mrs	<off>, TPIDR_ELx
> > >          add	<addr>, <pcp>, <off>
> > > 
> > >          // Perform access sequence
> > >          ldr	<val>, [<addr>]
> > > 
> > >          // Epilogue. Disable fixups
> > >          strh	wzr, [<tsk>, #TSK_TI_PCPU_GPRS]
> > > 
> > > If an exception is taken from within the critical section, the exception
> > > return code will adjust <off> to be the current CPU's offset, and will
> > > adjust <addr> to be (<pcp> + <off>). Distinct registers are used for
> > > <pcp>, <off>, and <addr>, so that the fixup can be applied safely at any
> > > point during the critical section.
> > > 
> > > To ensure that this_cpu_*() operations within exception handlers work
> > > correctly and do not corrupt state, thread_info::pcpu_gprs is saved
> > > into a new pt_regs::pcpu_gprs field upon exception entry, and restored
> > > upon exception return.
> > > 
> > > Looking at a simple this_cpu_operation:
> > > 
> > > | void outline_this_cpu_add_u64(u64 __percpu *p, u64 v)
> > > | {
> > > | 	this_cpu_add(*p, v);
> > > | }
> ...
> > I think I had an Interesting Idea(tm) while reading the per-cpu discussion
> > in linux-mm. In case the 3 instruction preamble is too expensive:
> > 
> > 1) Pass -ffixed-x18 (this natively conflicts with SHADOW_CALL_STACK.
> > SHADOW_CALL_STACK is already not-optimal codegen wise, so maybe not a big deal).
> > 2) arm64 kernel bits will use x18 as a cheap task flags register
> > 3) #define TASK_KRSEQ     (1 << 0)
> > 4) Switching into the krseq mode is just a matter of toggling the bit in x18, so
> > 	orr x18, x18, #TASK_KRSEQ
> > a single instruction.
> > 5) Switching off is just a matter of clearing the bit in x18, so:
> > 	and x18, x18, #~TASK_KRSEQ
> > 6) On the preempt side we keep the krseq tables in memory, and do a sort of lookup
> > (binary search sounds easiest?) on them. But _only_ if x18 TASK_KRSEQ is set.
> > This penalises unlucky preempts but keeps fast paths maximally fast.
> > 7) entry points of course get to clear it after saving it
> > 
> > The end result would look something like:
> > | <outline_this_cpu_add_u64>:
> > |	 orr x18, x18, #TASK_KRSEQ
> > |        mrs     x4, tpidr_el1
> > |        add     x3, x0, x4
> > | 1:     ldxr    x6, [x3]
> > |        add     x6, x6, x1
> > |        stxr    w5, x6, [x3]
> > |        cbnz    w5, 1b
> > | 2:
> > |	 and x18, x18, #~TASK_KRSEQ
> > |        ret
> > | .pushsection .data.krseq
> > | .word 1b
> > | .word 2b
> > | .word whateverelse
> > | .popsection

... Something I overlooked is that, since the solution shifted from
"literally kernel rseq" to "funky kernel rseq but pcpu-specific" you don't
probably don't need the table at all, as long as you dedicate a good few of
those x18 bits to stash the pcp_gprs value.

> > 
> > This of course precludes the use of x18 for the compiler, so it would
> > require careful benchmarking in case it negatively affects codegen too much.
> > But it avoids any sort of extraneous stores in the fast path.
> 
> The extra stores are independent of the main instruction flow.
> On a multi-issue (and especially out-of-order) cpu they are pretty much
> likely to be noise.

Oh, I agree, it is probably in the noise, modern uarchs are awesome :)

> The biggest cost is likely to be in the I-cache and instruction decoders.
> Put a memory read in the 'main' path and the few clocks needed for the
> D-cache read are likely to dominate - so the writes to the pcp_gprs
> are actually likely to be free.

While I do like theorycrafting, I think we need numbers to know (numbers
which I do not have, and seemingly no one seems to have for now).

> 
> OTOH stealing a gpr for some flags will cost everwhere.

Perhaps, but arm64 isn't exactly short on registers :) In any case,
I generally agree with your take that it quite possibly doesn't matter,
I was just throwing this out there in case it can help.

-- 
Pedro
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.