Re: [V2] riscv: Implement Zbb based strlen and prefer it over the RVV based strlen implementation when Zbb is available
Jeffrey Law <[email protected]>
| Newsgroups | gmane.comp.lib.glibc.alpha |
|---|---|
| Message-ID | <[email protected]> |
On 7/5/2026 6:20 AM, Florian Weimer wrote: > * Jeffrey Law: > >> I'm seeing some really weird results and I wish camel-cdr had data on >> a Zbb implemenation to correlate against -- his data has been >> incredibly helpful in confirming the vector behavior as well as the >> generic implementation behavior. But he didn't have a zbb variant. >> >> Basically it appears yours is meaningfully slower for small strings. >> It starts off about 50% slower, but closes the gap for each 2^n length >> bucket. The crossover point is around 128 bytes after which it blows >> away the other zbb variant -- it was still widening the gap at the 8k >> bucket where it was about 60% faster. *BUT* the loop itself is >> unchanged and while it appears there's one less architectural register >> in use, I'd be amazed if that one register was enough to make this >> kind of difference at the uarch level. > Hmm, maybe it's not working correctly and incorrectly returning early? That was my suspicion. I'll need to see what make check reports when I'm back home and I don't have to fight internet connection drops. I certainly don't see anything obvious when I look at the code. > As I said, it's untested. Or maybe it's related to the .p2align? That's a secondary possibility as well. I recall the K1 having a loop buffer of some sort, so the K3 might as well. The details were quite scarce, but I could imagine the possibility that if we the main loop's location within a icache line changed that we may or may not be running out of the loop buffer. > > The dependency chains should be equal or shorter even for short strings, > so the regression is surprising. And the loop really should not get any > faster. That's my read of the code as well. Hence my confusion. jeff