Re: [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series)
"David Hildenbrand (Arm)" <[email protected]> Tue, 4 Aug 2026 23:47:03 +0200
| Newsgroups | gmane.linux.kernel,gmane.linux.kernel.mm,gmane.linux.ports.arm.kernel |
|---|---|
| Message-ID | <[email protected]> |
On 8/4/26 23:25, Christoph Lameter (Ampere) wrote: > On Tue, 4 Aug 2026, David Hildenbrand (Arm) wrote: > >>> rather hacky restart logic. It relies on a long preemable and postscript >>> to each per cpu operations. > >> But I am sure you did you homework and have data to back up your claims. Please >> share that data, because I am very curious. > > The comparison here is a single instruction on one side and abnd the per > cpu operation prefixed with a preeamble and another instruction / set of > instructions on the other side. > > Even if its only one before and one after then we are still looking at 3 > instructions instead of one. This is a frequently used operation in the > VM. > > You need data to show that one instruction is faster than 3??? I want you to show that the difference *matters* in *practice* enough that we should *care* about accepting the additional *complexity*. No hand-waving about "3 > 1". IIRC people were asking that previously but received no results back. So to be clear: claiming that the additional instructions are unacceptable I'd expect something minimal like: "removing these 3 instructions speeds up micro-benchmark (a, b, c) by X% and macro-benchmarks (d, e, f) by Y% compared to Mark's approach". Maybe there really is a big difference. > >> Mark's solution is the obvious improvement to the problem, doing it just like >> s390 already does. > > Well there is interest by the S390 folks to move to what we proposed from > what I can tell. Heiko said that as reply to v1, bit he's been replying a lot on Mark's patches. Only Heiko can tell. (on CC) > > >> If the proposed solution has a significant benefit over Mark's solution it could >> be discussed as a next step; and it must be backed with sufficient data to >> warrant all this complexity here. > > You cannot accept that one instruction is faster than the same > instruction prefixed with another and something else at its tail? I am pretty sure that you got exactly what I meant, you likely just don't have the data to back it up. Proof me wrong! -- Cheers, David