Re: [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series)

David Laight <[email protected]> Wed, 5 Aug 2026 12:10:39 +0100
Newsgroups gmane.linux.kernel,gmane.linux.kernel.mm,gmane.linux.ports.arm.kernel
Message-ID <20260805121039.471c7340@pumpkin>
On Tue, 4 Aug 2026 18:47:26 +0200
"David Hildenbrand (Arm)" <[email protected]> wrote:

> On 8/4/26 18:19, Christoph Lameter (Ampere) wrote:
> > On Tue, 4 Aug 2026, Lorenzo Stoakes (ARM) wrote:
> >   
> >> Since this work seems to be very much arm64-focused, perhaps it's therefore
> >> worth looking at an alterative solution that's specific to the arch, like the
> >> one suggested by Mark ([1])?
> >>
> >> [0]:https://lore.kernel.org/all/CAHk-=wire3dzhHx=KiL_f5Rj0=1u9ustsa33QoR-F9-v-NU9Ng@mail.gmail.com/
> >> [1]:https://lore.kernel.org/linux-arm-kernel/al_DpFJFcmVhxpvW@J2N7QTR9R3/  
> > 
> > Mark's solution does replace the preempt_enable/disable sections with a
> > rather hacky restart logic. It relies on a long preemable and postscript
> > to each per cpu operations.  
> 
> Okay, so 3 simple instructions of preemable is "long preemable"? In which universe?
> 
> But I am sure you did you homework and have data to back up your claims. Please
> share that data, because I am very curious.

The proposed sequence is:
>           // Prologue. Enable fixups for <off> and <addr>.
>  1        mrs	<tsk>, sp_el0
>  2        mov	<tmp>, #__VAL_PCPU_GPRS(<pcp>, <off>, <addr>)
>  3        strh	<tmp>, [<tsk>, #TSK_TI_PCPU_GPRS]
> 
>           // Generate cpu-specific address
>  4        mrs	<off>, TPIDR_ELx
>  5        add	<addr>, <pcp>, <off>
> 
>           // Perform access sequence
>  6        ldr	<val>, [<addr>]
> 
>           // Epilogue. Disable fixups
>  7        strh	wzr, [<tsk>, #TSK_TI_PCPU_GPRS]

Think about how that actually gets execute by a real cpu.
Instructions will be read from the I-cache in 'chunks' (maybe half a cache line).
They are then fed to multiple decoders that generate u-ops for the execution units.
The decoder is unlikely to be a bottleneck.
I've numbered the instructions:
First clock can run instructions 1, 2 and 4.
Assuming the mrs have no extra latency the second runs 3 and 5.
The third will then run 6 and 7.
The cpu then probably has to wait for the result of the ldr.

If the access is a write then there may be a stall waiting for the value
to be written to be available.

The only real effect of the extra instructions is likely to be code size.

	David