Re: [PATCH] libcpp: add an AArch64 SVE2 search_line_fast

Kyrylo Tkachov <[email protected]>
Newsgroups gmane.comp.gcc.patches
Message-ID <[email protected]>
Hi Wilco,

> On 7 Aug 2026, at 14:10, Wilco Dijkstra <[email protected]> wrote:
> 
> Hi Kyrill,
> 
>> The AArch64 helper that finds the next '\n', '\r', '\\' or '?' scans a
>> fixed 16 bytes per iteration and spends more than half of its per-call
>> cost outside the loop, turning a vector comparison into a byte index.
>> Each iteration needs four CMEQs, three ORRs and an ADDP/FMOV pair to get
>> the comparison into a general register for the loop branch, and the exit
>> path builds a 16-bit mask with AND/UADDLP/USHL/ADDV/FMOV before it can
>> count trailing zeros.  Both sequences also transfer from SIMD to GP
>> regs, which can be expensive.
> 
> This sounds like an ancient sequence - we know how to do this fast by
> using SHRN or ADDHN to narrow the mask so that it can both be tested
> and immediately used in CTZ to get the bit position. We should fix this in
> another patch.
> 
> From a high level perspective, I'm wondering why SVE2 would use a pagecross
> check when you have a start and end pointer? The aligning and the pagecross
> check add extra overhead, so looping without aligning would be faster if we
> check that there is at least 1 whole vector left to do (then only the last partial
> vector needs special treatment).
> 
> Alternatively, if a pagecross check is useful (eg. if input buffer is frequently small)
> then it's better to process at least 2 and possibly 4 vectors before either aligning or
> falling back to the above loop.
> 
> It would be good to list the statistics that resulted in this design. Lines are short on
> average, but there are also lots of lines that need 4 or more vectors.
> 
> As for dispatch, I'm assuming HAVE_SVE2 would be true in almost all cases, so
> we get overhead from the indirect call (and no more inlining) in all cases. I think it
> would be better if we wrote:
> 
> static bool has_sve2;
> 
> static const uchar *
> search_line_fast (const uchar *s, const uchar *end ATTRIBUTE_UNUSED)
> {
>  if (has_sve2)
>    return search_line_sve2 (s, end);
>  // Neon version
> }
> 
> Although this likely can't inline the SVE2 version, the dispatch is very cheap and the
> Neon version can still be inlined as before - so you either get existing performance or
> a speedup from SVE2.

Thanks for having a look!
Indeed the page cross checks were not necessary, I overengineered it by looking at the old Neon version.
The new version as I think you described it is simpler and a bit faster (90% gain vs 80% gain). I’ve sent a v2 with your suggestions incorporated.
Kyrill

> 
> Cheers,
> Wilco
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.