Re: [PATCH v2 2/3] target/arm: Fix SVE2 WHILEWR/WHILERW zero diff boundary case

Peter Maydell <[email protected]>
Newsgroups org.nongnu.qemu-arm,org.nongnu.qemu-devel
Message-ID <CAFEAcA91Gc7-n2SBvA-+vF2PemB6gqWFKKsoS9EPZ4Ks+N7awg@mail.gmail.com>
On Mon, 17 Aug 2026 at 23:34, Richard Henderson
<[email protected]> wrote:
>
> On 8/17/26 09:37, Peter Maydell wrote:
> >> +    tcg_gen_addi_i64(diff, diff, -1);
> >> +    tcg_gen_umin_i64(diff, diff, tcg_constant_i64((vsz >> a->esz) - 1));
> >> +
> >> +    /*
> >> +     * Since we're bounded, pass as a 32-bit type.
> >> +     * Sink the diff += 1 from above into the 32-bit type.
> >
> > Is it worth doing that? Hosts are all 64-bits now, and surely there's
> > not going to be a perf or codesize difference between "add 1"
> > on a 32-bit type vs a 64-bit type ?
> Technically one byte difference on x86_64.
> But we could drop the whole truncation thing it you like.

I think mostly my feeling was that since the rearrangement of the
umin() expressions above was already quite complicated, moving
the last part of that down into the truncate-to-32-bits was
adding extra complexity on top. Just doing the subtract-one in
64 bits and keeping the truncate would be fine, as then the
two steps aren't interleaved with each other.

-- PMM
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.