[Bug target/126597] [15/16/17 Regression] Wrong aarch64 code with -O0 and aarch64_expand_vec_perm_const_1

"cvs-commit at gcc dot gnu.org via Gcc-bugs" <[email protected]>
Newsgroups gmane.comp.gcc.bugs
Message-ID <[email protected]/bugzilla/>
https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126597

--- Comment #2 from GCC Commits <cvs-commit at gcc dot gnu.org> ---
The master branch has been updated by Kyrylo Tkachov <[email protected]>:

https://gcc.gnu.org/g:75b259d068e1db2c8ac2246e3118dd5d24724674

commit r17-3401-g75b259d068e1db2c8ac2246e3118dd5d24724674
Author: Kyrylo Tkachov <[email protected]>
Date:   Tue Aug 18 14:19:57 2026 +0200

    aarch64: Swap the zeroness flags when swapping vec_perm operands [PR126597]

    aarch64_expand_vec_perm_const_1 normalizes a permutation whose first index
    selects the second operand by rotating the indices and swapping op0 and
op1.
    It left zero_op0_p and zero_op1_p pointing at the old operands, so the
later
    recognizers that consult them, aarch64_evpc_and and aarch64_evpc_tbl, read
    the wrong vector.

    Swap the two flags together with the operands.

    For

            typedef int v4si __attribute__ ((vector_size (16)));
            v4si f (v4si x)
            {
              const v4si m = { 4, 1, 2, 3 };
              return __builtin_shuffle (x, (v4si) { 0, 0, 0, 0 }, m);
            }

    at -O0 the AND was applied to the all-zero operand, so the function
returned
    {0,0,0,0} instead of {0,x1,x2,x3}:

            sub     sp, sp, #32
            str     q0, [sp]
            adrp    x0, .LC0
            add     x0, x0, :lo12:.LC0
            ldr     q31, [x0]
            str     q31, [sp, 16]
            movi    v31.4s, 0
            fmov    s31, s31
            mov     v0.16b, v31.16b
            add     sp, sp, 32
            ret

    With the fix the AND is applied to the incoming vector:

            sub     sp, sp, #32
            str     q0, [sp]
            adrp    x0, .LC0
            add     x0, x0, :lo12:.LC0
            ldr     q31, [x0]
            str     q31, [sp, 16]
            ldr     q30, [sp]
            adrp    x0, .LC1
            add     x0, x0, :lo12:.LC1
            ldr     q31, [x0]
            and     v31.16b, v30.16b, v31.16b
            mov     v0.16b, v31.16b
            add     sp, sp, 32
            ret

    Bootstrapped and tested on aarch64-none-linux-gnu.

    gcc/ChangeLog:

            PR target/126597
            * config/aarch64/aarch64.cc (aarch64_expand_vec_perm_const_1): Swap
            zero_op0_p and zero_op1_p along with the operands.

    gcc/testsuite/ChangeLog:

            PR target/126597
            * gcc.target/aarch64/pr126597.c: New test.

    Signed-off-by: Kyrylo Tkachov <[email protected]>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.