Re: [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series)

David Laight <[email protected]>
Newsgroups org.infradead.lists.linux-arm-kernel,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <20260805121039.471c7340@pumpkin>
On Tue, 4 Aug 2026 18:47:26 +0200
"David Hildenbrand (Arm)" <[email protected]> wrote:

> On 8/4/26 18:19, Christoph Lameter (Ampere) wrote:
> > On Tue, 4 Aug 2026, Lorenzo Stoakes (ARM) wrote:
> >   
> >> Since this work seems to be very much arm64-focused, perhaps it's therefore
> >> worth looking at an alterative solution that's specific to the arch, like the
> >> one suggested by Mark ([1])?
> >>
> >> [0]:https://lore.kernel.org/all/CAHk-=wire3dzhHx=KiL_f5Rj0=1u9ustsa33QoR-F9-v-NU9Ng@mail.gmail.com/
> >> [1]:https://lore.kernel.org/linux-arm-kernel/al_DpFJFcmVhxpvW@J2N7QTR9R3/  
> > 
> > Mark's solution does replace the preempt_enable/disable sections with a
> > rather hacky restart logic. It relies on a long preemable and postscript
> > to each per cpu operations.  
> 
> Okay, so 3 simple instructions of preemable is "long preemable"? In which universe?
> 
> But I am sure you did you homework and have data to back up your claims. Please
> share that data, because I am very curious.

The proposed sequence is:
>           // Prologue. Enable fixups for <off> and <addr>.
>  1        mrs	<tsk>, sp_el0
>  2        mov	<tmp>, #__VAL_PCPU_GPRS(<pcp>, <off>, <addr>)
>  3        strh	<tmp>, [<tsk>, #TSK_TI_PCPU_GPRS]
> 
>           // Generate cpu-specific address
>  4        mrs	<off>, TPIDR_ELx
>  5        add	<addr>, <pcp>, <off>
> 
>           // Perform access sequence
>  6        ldr	<val>, [<addr>]
> 
>           // Epilogue. Disable fixups
>  7        strh	wzr, [<tsk>, #TSK_TI_PCPU_GPRS]

Think about how that actually gets execute by a real cpu.
Instructions will be read from the I-cache in 'chunks' (maybe half a cache line).
They are then fed to multiple decoders that generate u-ops for the execution units.
The decoder is unlikely to be a bottleneck.
I've numbered the instructions:
First clock can run instructions 1, 2 and 4.
Assuming the mrs have no extra latency the second runs 3 and 5.
The third will then run 6 and 7.
The cpu then probably has to wait for the result of the ldr.

If the access is a write then there may be a stall waiting for the value
to be written to be available.

The only real effect of the extra instructions is likely to be code size.

	David
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.