Re: [PATCH 19/20] target/arm: Implement PMULL, PMLAL
Peter Maydell <[email protected]> Tue, 28 Jul 2026 17:37:28 +0100
| Newsgroups | org.nongnu.qemu-arm,org.nongnu.qemu-devel |
|---|---|
| Message-ID | <CAFEAcA8SinMnnYyWxzbNsE-B9tB3TQ8KiXQem1KZ9Vt9sH5EEQ@mail.gmail.com> |
On Tue, 21 Jul 2026 at 19:49, Richard Henderson <[email protected]> wrote: > > Signed-off-by: Richard Henderson <[email protected]> > --- > +void HELPER(sve_pmull_q)(void *vd, void *vn, void *vm, uint32_t desc) > +{ > + intptr_t opr_sz = simd_oprsz(desc); > + uint64_t *n = vn, *m = vm; > + Int128 *d0 = vd; > + Int128 *d1 = vd + sizeof(ARMVectorReg); > + > + for (intptr_t i = 0; i < opr_sz / 16; ++i) { > + Int128 rl = clmul_64(n[2 * i + 0], m[2 * i + 0]); > + Int128 rh = clmul_64(n[2 * i + 1], m[2 * i + 1]); > + d0[i] = rl; > + d1[i] = rh; Doesn't this put the two 64-bit halves of the results in the wrong order on big-endian hosts? Compare how we handle writing back the Int128 result in sve2_adcl_d. > + } > +} > + > +void HELPER(sve_pmlal_q)(void *vd, void *vn, void *vm, uint32_t desc) > +{ > + intptr_t opr_sz = simd_oprsz(desc); > + uint64_t *n = vn, *m = vm; > + Int128 *d0 = vd; > + Int128 *d1 = vd + sizeof(ARMVectorReg); > + > + for (intptr_t i = 0; i < opr_sz / 16; ++i) { > + Int128 rl = clmul_64(n[2 * i + 0], m[2 * i + 0]); > + Int128 rh = clmul_64(n[2 * i + 1], m[2 * i + 1]); > + d0[i] = int128_xor(d0[i], rl); > + d1[i] = int128_xor(d1[i], rh); > + } > +} Similarly here. (The other places where we use Int128 in helpers -- ZIP, UZP, TRN, and some moves -- don't need endianness handling as they are just copying 128 bits of data from one vector to another without interpreting it.) thanks -- PMM