RE: [PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum reductions
Tamar Christina <[email protected]> Tue, 4 Aug 2026 12:52:17 +0000
| Newsgroups | gmane.comp.gcc.patches |
|---|---|
| Message-ID | <VI0PR08MB10392E606F946FBCA47CED357FFD42@VI0PR08MB10392.eurprd08.prod.outlook.com> |
Hi Kyrill, > -----Original Message----- > From: [email protected] <[email protected]> > Sent: 04 August 2026 13:39 > To: [email protected] > Cc: Tamar Christina <[email protected]>; [email protected]; Kyrylo > Tkachov <[email protected]> > Subject: [PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum > reductions >=20 > From: Kyrylo Tkachov <[email protected]> >=20 > The Advanced SIMD widen_[su]sum optabs only cover a single widening step, > expanded as a dependent <su>addw + <su>addw2 pair, plus a 4x form that > requires dot product. A reduction into an accumulator that is more than > twice as wide as the data therefore has to extend the input explicitly an= d > then issue one widening add per half vector. Summing bytes into a 64-bit > accumulator costs fifteen SIMD operations per 16 bytes of input. >=20 > [SU]ADDLP and [SU]ADALP add adjacent lane pairs into the next wider > element, so a chain of them expresses any power-of-two widening sum > reduction in one operation per step. The regrouping is exact because the > sum of two elements always fits in the doubled element width, and the > grouping of lanes inside a reduction accumulator is already unconstrained > for WIDEN_SUM_EXPR, which the existing dot product based 4x expander also > relies on. >=20 > Expand the 2x forms as a single [SU]ADALP, add the missing V4SI <- V16QI > and V2SI <- V8QI forms for !TARGET_DOTPROD, and add the V2DI <- V8HI and > V2DI <- V16QI forms that no expander covered. All of them are built by > aarch64_expand_widen_sum, which halves the lane count with [SU]ADDLP > until > one pairwise step remains and then accumulates with [SU]ADALP. >=20 > For a sum of unsigned char into long the inner loop changes from >=20 > ldr q30, [x1], 16 > zip1 v28.16b, v30.16b, v29.16b > zip2 v30.16b, v30.16b, v29.16b > zip1 v26.8h, v28.8h, v29.8h > zip2 v28.8h, v28.8h, v29.8h > zip1 v27.8h, v30.8h, v29.8h > zip2 v30.8h, v30.8h, v29.8h > uaddw v31.2d, v31.2d, v26.2s > uaddw2 v31.2d, v31.2d, v26.4s > ... (six more uaddw/uaddw2) >=20 > to >=20 > ldr q31, [x1], 16 > uaddlp v31.8h, v31.16b > uaddlp v31.4s, v31.8h > uadalp v30.2d, v31.4s >=20 Have you considered considered even for the b to d case to use dotprod for the b -> s and then uadalp for the final 4 to d? or is the above codege= n the fallback for when !dotprod? wasn't quite clear.. That should still shave off 1 cycle. Thanks, Tamar > and for a sum of int into long the saddw/saddw2 pair becomes one sadalp. > On a Neoverse V2 core with an L1 resident working set this cuts the time = of > the > byte loop by about 88% and of the int loop by about 68%. >=20 > Bootstrapped and tested on aarch64-none-linux-gnu. > Ok for trunk? > Thanks, > Kyrill >=20 > gcc/ChangeLog: >=20 > * config/aarch64/aarch64-protos.h (aarch64_expand_widen_sum): > Declare. > * config/aarch64/aarch64.cc (aarch64_expand_widen_sum): New > function. > * config/aarch64/aarch64-simd.md (aarch64_<su>adalp<mode>): > Rename > to ... > (@aarch64_<su>adalp<mode>): ... this. > (widen_ssum<Vdblw><mode>3, widen_usum<Vdblw><mode>3): > Replace by ... > (widen_<su>sum<Vdblw><mode>3): ... this. Expand to [SU]ADALP. > (widen_ssum<mode><vsi2qi>3, widen_usum<mode><vsi2qi>3): > Replace > by ... > (widen_<su>sum<mode><vsi2qi>3): ... this. Handle > !TARGET_DOTPROD. > (widen_<su>sumv2di<mode>3): New expander. > * config/aarch64/iterators.md (VQ_BH): New mode iterator. >=20 > gcc/testsuite/ChangeLog: >=20 > * gcc.target/aarch64/pr122069_1.c: Update expected output. > * gcc.target/aarch64/pr122069_3.c: Likewise. > * gcc.target/aarch64/saddw-1.c: Renamed to... > * gcc.target/aarch64/sadalp-1.c: ...this. Update expected output. > * gcc.target/aarch64/saddw-2.c: Renamed to... > * gcc.target/aarch64/sadalp-2.c: ...this. Update expected output. > * gcc.target/aarch64/uaddw-1.c: Renamed to... > * gcc.target/aarch64/uadalp-1.c: ...this. Update expected output. > * gcc.target/aarch64/uaddw-2.c: Renamed to... > * gcc.target/aarch64/uadalp-2.c: ...this. Update expected output. > * gcc.target/aarch64/uaddw-3.c: Renamed to... > * gcc.target/aarch64/uadalp-3.c: ...this. Update expected output. > * gcc.target/aarch64/widen_sum_pairwise_1.c: New test. > * gcc.target/aarch64/widen_sum_pairwise_2.c: New test. >=20 > Signed-off-by: Kyrylo Tkachov <[email protected]> > --- > gcc/config/aarch64/aarch64-protos.h | 1 + > gcc/config/aarch64/aarch64-simd.md | 78 ++++++++----------- > gcc/config/aarch64/aarch64.cc | 27 +++++++ > gcc/config/aarch64/iterators.md | 4 + > gcc/testsuite/gcc.target/aarch64/pr122069_1.c | 11 +-- > gcc/testsuite/gcc.target/aarch64/pr122069_3.c | 3 +- > .../aarch64/{saddw-1.c =3D> sadalp-1.c} | 3 +- > .../aarch64/{saddw-2.c =3D> sadalp-2.c} | 3 +- > .../aarch64/{uaddw-1.c =3D> uadalp-1.c} | 3 +- > .../aarch64/{uaddw-2.c =3D> uadalp-2.c} | 3 +- > .../aarch64/{uaddw-3.c =3D> uadalp-3.c} | 3 +- > .../gcc.target/aarch64/widen_sum_pairwise_1.c | 39 ++++++++++ > .../gcc.target/aarch64/widen_sum_pairwise_2.c | 29 +++++++ > 13 files changed, 141 insertions(+), 66 deletions(-) > rename gcc/testsuite/gcc.target/aarch64/{saddw-1.c =3D> sadalp-1.c} (74%= ) > rename gcc/testsuite/gcc.target/aarch64/{saddw-2.c =3D> sadalp-2.c} (74%= ) > rename gcc/testsuite/gcc.target/aarch64/{uaddw-1.c =3D> uadalp-1.c} (75%= ) > rename gcc/testsuite/gcc.target/aarch64/{uaddw-2.c =3D> uadalp-2.c} (75%= ) > rename gcc/testsuite/gcc.target/aarch64/{uaddw-3.c =3D> uadalp-3.c} (74%= ) > create mode 100644 > gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c > create mode 100644 > gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >=20 > diff --git a/gcc/config/aarch64/aarch64-protos.h > b/gcc/config/aarch64/aarch64-protos.h > index bcc833cfaa1..9303f12f80c 100644 > --- a/gcc/config/aarch64/aarch64-protos.h > +++ b/gcc/config/aarch64/aarch64-protos.h > @@ -1066,6 +1066,7 @@ void aarch64_emit_sve_pred_vec_duplicate > (machine_mode, rtx, rtx); > void aarch64_expand_prologue (void); > void aarch64_decompose_vec_struct_index (machine_mode, rtx *, rtx *, > bool); > void aarch64_expand_vector_init (rtx, rtx); > +void aarch64_expand_widen_sum (rtx, rtx, rtx, rtx_code); > void aarch64_sve_expand_vector_init_subvector (rtx, rtx); > void aarch64_sve_expand_vector_init (rtx, rtx); > void aarch64_init_cumulative_args (CUMULATIVE_ARGS *, const_tree, rtx, > diff --git a/gcc/config/aarch64/aarch64-simd.md > b/gcc/config/aarch64/aarch64-simd.md > index 433f16052bf..d119ac17352 100644 > --- a/gcc/config/aarch64/aarch64-simd.md > +++ b/gcc/config/aarch64/aarch64-simd.md > @@ -1182,7 +1182,7 @@ > } > ) >=20 > -(define_expand "aarch64_<su>adalp<mode>" > +(define_expand "@aarch64_<su>adalp<mode>" > [(set (match_operand:<VDBLW> 0 "register_operand") > (plus:<VDBLW> > (plus:<VDBLW> > @@ -5283,19 +5283,17 @@ >=20 > ;; <su><addsub>w<q>. >=20 > -(define_expand "widen_ssum<Vdblw><mode>3" > +;; A widening sum reduction that halves the lane count is a single pairw= ise > +;; widening accumulate. > +(define_expand "widen_<su>sum<Vdblw><mode>3" > [(set (match_operand:<VDBLW> 0 "register_operand") > - (plus:<VDBLW> (sign_extend:<VDBLW> > - (match_operand:VQW 1 "register_operand")) > + (plus:<VDBLW> (ANY_EXTEND:<VDBLW> > + (match_operand:VQW 1 "register_operand")) > (match_operand:<VDBLW> 2 "register_operand")))] > "TARGET_SIMD" > { > - rtx p =3D aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, fal= se); > - rtx temp =3D gen_reg_rtx (GET_MODE (operands[0])); > - > - emit_insn (gen_aarch64_saddw<mode>_internal (temp, operands[2], > - operands[1], p)); > - emit_insn (gen_aarch64_saddw2<mode> (operands[0], temp, > operands[1])); > + emit_insn (gen_aarch64_<su>adalp<mode> (operands[0], operands[2], > + operands[1])); > DONE; > } > ) > @@ -5311,23 +5309,6 @@ > DONE; > }) >=20 > -(define_expand "widen_usum<Vdblw><mode>3" > - [(set (match_operand:<VDBLW> 0 "register_operand") > - (plus:<VDBLW> (zero_extend:<VDBLW> > - (match_operand:VQW 1 "register_operand")) > - (match_operand:<VDBLW> 2 "register_operand")))] > - "TARGET_SIMD" > - { > - rtx p =3D aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, fal= se); > - rtx temp =3D gen_reg_rtx (GET_MODE (operands[0])); > - > - emit_insn (gen_aarch64_uaddw<mode>_internal (temp, operands[2], > - operands[1], p)); > - emit_insn (gen_aarch64_uaddw2<mode> (operands[0], temp, > operands[1])); > - DONE; > - } > -) > - > (define_expand "widen_usum<Vwide><mode>3" > [(set (match_operand:<VWIDE> 0 "register_operand") > (plus:<VWIDE> (zero_extend:<VWIDE> > @@ -5339,38 +5320,43 @@ > DONE; > }) >=20 > -(define_expand "widen_ssum<mode><vsi2qi>3" > +;; A widening sum reduction that quarters the lane count. With dot prod= uct > +;; this is one [SU]DOT with a vector of ones, i.e. +=3D a becomes +=3D (= a * 1). > +;; Otherwise it is a pairwise widening add feeding a pairwise widening > +;; accumulate. > +(define_expand "widen_<su>sum<mode><vsi2qi>3" > [(set (match_operand:VS 0 "register_operand") > - (plus:VS (sign_extend:VS > + (plus:VS (ANY_EXTEND:VS > (match_operand:<VSI2QI> 1 "register_operand")) > (match_operand:VS 2 "register_operand")))] > - "TARGET_DOTPROD" > + "TARGET_SIMD" > { > - rtx ones =3D force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); > - emit_insn (gen_sdot_prod<mode><vsi2qi> (operands[0], operands[1], > ones, > - operands[2])); > + if (TARGET_DOTPROD) > + { > + rtx ones =3D force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); > + emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0], > operands[1], > + ones, operands[2])); > + } > + else > + aarch64_expand_widen_sum (operands[0], operands[2], operands[1], > <CODE>); > DONE; > } > ) >=20 > -;; Use dot product to perform double widening sum reductions by > -;; changing +=3D a into +=3D (a * 1). i.e. we seed the multiplication w= ith 1. > -(define_expand "widen_usum<mode><vsi2qi>3" > - [(set (match_operand:VS 0 "register_operand") > - (plus:VS (zero_extend:VS > - (match_operand:<VSI2QI> 1 "register_operand")) > - (match_operand:VS 2 "register_operand")))] > - "TARGET_DOTPROD" > +;; Widening sum reductions into 64-bit elements. These need two or thre= e > +;; pairwise widening steps. > +(define_expand "widen_<su>sumv2di<mode>3" > + [(set (match_operand:V2DI 0 "register_operand") > + (plus:V2DI (ANY_EXTEND:V2DI > + (match_operand:VQ_BH 1 "register_operand")) > + (match_operand:V2DI 2 "register_operand")))] > + "TARGET_SIMD" > { > - rtx ones =3D force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); > - emit_insn (gen_udot_prod<mode><vsi2qi> (operands[0], operands[1], > ones, > - operands[2])); > + aarch64_expand_widen_sum (operands[0], operands[2], operands[1], > <CODE>); > DONE; > } > ) >=20 > -;; Use dot product to perform double widening sum reductions by > -;; changing +=3D a into +=3D (a * 1). i.e. we seed the multiplication w= ith 1. > (define_insn "aarch64_<ANY_EXTEND:su>subw<mode>" > [(set (match_operand:<VWIDE> 0 "register_operand" "=3Dw") > (minus:<VWIDE> (match_operand:<VWIDE> 1 "register_operand" > "w") > diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.c= c > index d19ca305d82..628e94e8c40 100644 > --- a/gcc/config/aarch64/aarch64.cc > +++ b/gcc/config/aarch64/aarch64.cc > @@ -26327,6 +26327,33 @@ aarch64_expand_vector_init (rtx target, rtx > vals) > emit_insn (seq_total_cost < fallback_seq_cost ? seq : fallback_seq); > } >=20 > +/* Expand the widening sum reduction DEST =3D ACC + (WIDE) SRC, where th= e > + Advanced SIMD vector SRC holds an even multiple of the number of lane= s > + of the accumulator ACC and of the result DEST. EXTEND_CODE is > + SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form. > + Halve the lane count with [SU]ADDLP until a single pairwise step is > + left, then accumulate into ACC with [SU]ADALP. */ > + > +void > +aarch64_expand_widen_sum (rtx dest, rtx acc, rtx src, rtx_code extend_co= de) > +{ > + unsigned int dest_nunits =3D GET_MODE_NUNITS (GET_MODE > (dest)).to_constant (); > + machine_mode mode =3D GET_MODE (src); > + gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * 2) > =3D=3D 0); > + > + while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2) > + { > + insn_code icode =3D code_for_aarch64_addlp (extend_code, mode); > + mode =3D insn_data[icode].operand[0].mode; > + rtx tmp =3D gen_reg_rtx (mode); > + emit_insn (GEN_FCN (icode) (tmp, src)); > + src =3D tmp; > + } > + > + emit_insn (GEN_FCN (code_for_aarch64_adalp (extend_code, mode)) > (dest, acc, > + src)); > +} > + > /* Emit RTL corresponding to: > insr TARGET, ELEM. */ >=20 > diff --git a/gcc/config/aarch64/iterators.md > b/gcc/config/aarch64/iterators.md > index 0d319751430..6a8c93cce37 100644 > --- a/gcc/config/aarch64/iterators.md > +++ b/gcc/config/aarch64/iterators.md > @@ -313,6 +313,10 @@ > ;; All quad integer widen-able modes. > (define_mode_iterator VQW [V16QI V8HI V4SI]) >=20 > +;; Quad integer modes that reach 64-bit elements through more than one > +;; pairwise widening step. > +(define_mode_iterator VQ_BH [V16QI V8HI]) > + > ;; Double vector modes for combines. > (define_mode_iterator VDC [V8QI V4HI V4BF V4HF V2SI V2SF DI DF]) >=20 > diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c > b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c > index b2f973261ea..d99b5493ade 100644 > --- a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c > +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c > @@ -10,12 +10,8 @@ inline char char_abs(char i) { > ** foo_int: > ** ... > ** sub v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b > -** zip1 v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b > -** zip2 v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b > -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h > -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h > -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h > -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h > +** uaddlp v[0-9]+.8h, v[0-9]+.16b > +** uadalp v[0-9]+.4s, v[0-9]+.8h > ** ... > */ > int foo_int(unsigned char *x, unsigned char * restrict y) { > @@ -29,8 +25,7 @@ int foo_int(unsigned char *x, unsigned char * restrict = y) { > ** foo2_int: > ** ... > ** add v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h > -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h > -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h > +** uadalp v[0-9]+.4s, v[0-9]+.8h > ** ... > */ > int foo2_int(unsigned short *x, unsigned short * restrict y) { > diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c > b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c > index 0e832c43032..f29fc2b2ed4 100644 > --- a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c > +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c > @@ -24,8 +24,7 @@ int foo_int(unsigned char *x, unsigned char * restrict = y) { > ** foo2_int: > ** ... > ** add v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h > -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h > -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h > +** uadalp v[0-9]+.4s, v[0-9]+.8h > ** ... > */ > int foo2_int(unsigned short *x, unsigned short * restrict y) { > diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-1.c > b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c > similarity index 74% > rename from gcc/testsuite/gcc.target/aarch64/saddw-1.c > rename to gcc/testsuite/gcc.target/aarch64/sadalp-1.c > index f8871209b8a..61f9633f1a0 100644 > --- a/gcc/testsuite/gcc.target/aarch64/saddw-1.c > +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c > @@ -14,5 +14,4 @@ t6(int len, void * dummy, short * __restrict x) > return result; > } >=20 > -/* { dg-final { scan-assembler "saddw" } } */ > -/* { dg-final { scan-assembler "saddw2" } } */ > +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } = */ > diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-2.c > b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c > similarity index 74% > rename from gcc/testsuite/gcc.target/aarch64/saddw-2.c > rename to gcc/testsuite/gcc.target/aarch64/sadalp-2.c > index b9fc442a2f7..873fda2e1ea 100644 > --- a/gcc/testsuite/gcc.target/aarch64/saddw-2.c > +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c > @@ -14,5 +14,4 @@ t6(int len, void * dummy, int * __restrict x) > return result; > } >=20 > -/* { dg-final { scan-assembler "saddw" } } */ > -/* { dg-final { scan-assembler "saddw2" } } */ > +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s} } } = */ > diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c > b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c > similarity index 75% > rename from gcc/testsuite/gcc.target/aarch64/uaddw-1.c > rename to gcc/testsuite/gcc.target/aarch64/uadalp-1.c > index 14dff87d7f0..c4034384aae 100644 > --- a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c > +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c > @@ -14,5 +14,4 @@ t6(int len, void * dummy, unsigned short * __restrict x= ) > return result; > } >=20 > -/* { dg-final { scan-assembler "uaddw" } } */ > -/* { dg-final { scan-assembler "uaddw2" } } */ > +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } = */ > diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c > b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c > similarity index 75% > rename from gcc/testsuite/gcc.target/aarch64/uaddw-2.c > rename to gcc/testsuite/gcc.target/aarch64/uadalp-2.c > index 79d0d094fc3..395d36c7c00 100644 > --- a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c > +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c > @@ -14,6 +14,5 @@ t6(int len, void * dummy, unsigned short * __restrict x= ) > return result; > } >=20 > -/* { dg-final { scan-assembler "uaddw" } } */ > -/* { dg-final { scan-assembler "uaddw2" } } */ > +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } = */ >=20 > diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c > b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c > similarity index 74% > rename from gcc/testsuite/gcc.target/aarch64/uaddw-3.c > rename to gcc/testsuite/gcc.target/aarch64/uadalp-3.c > index 39cbd6b6cc2..5fdb1639ab8 100644 > --- a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c > +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c > @@ -14,5 +14,4 @@ t6(int len, void * dummy, char * __restrict x) > return result; > } >=20 > -/* { dg-final { scan-assembler "uaddw" } } */ > -/* { dg-final { scan-assembler "uaddw2" } } */ > +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.8h, v[0-9]+\.16b} } }= */ > diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c > b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c > new file mode 100644 > index 00000000000..0aec0bf81c8 > --- /dev/null > +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c > @@ -0,0 +1,39 @@ > +/* { dg-do compile } */ > +/* { dg-options "-O3 -march=3Darmv8-a -mautovec-preference=3Dasimd-only = -- > param vect-epilogues-nomask=3D0" } */ > + > +/* Widening sum reductions should use the pairwise widening add and > + accumulate instructions rather than a chain of extensions feeding > + [SU]ADDW pairs. */ > + > +#define DEF(NAME, ITYPE, OTYPE) \ > + OTYPE NAME (const ITYPE *a, long n) \ > + { \ > + OTYPE s =3D 0; \ > + for (long i =3D 0; i < n; i++) \ > + s +=3D a[i]; \ > + return s; \ > + } > + > +DEF (sum_u8_l, unsigned char, long) > +DEF (sum_i8_l, signed char, long) > +DEF (sum_u16_l, unsigned short, long) > +DEF (sum_i16_l, short, long) > +DEF (sum_u32_l, unsigned int, long) > +DEF (sum_i32_l, int, long) > +DEF (sum_u8_i, unsigned char, int) > +DEF (sum_i8_i, signed char, int) > +DEF (sum_u16_i, unsigned short, int) > +DEF (sum_i16_i, short, int) > + > +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16= b\n} > 2 } } */ > +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.8h, v[0-9]+\.16= b\n} > 2 } } */ > +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h= \n} 2 > } } */ > +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0-9]+\.8h= \n} 2 > } } */ > +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s= \n} 3 > } } */ > +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s= \n} 3 > } } */ > +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h= \n} 2 > } } */ > +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h= \n} 2 > } } */ > + > +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */ > +/* { dg-final { scan-assembler-not {\tsaddw2?\t} } } */ > +/* { dg-final { scan-assembler-not {\tzip1\t} } } */ > diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c > b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c > new file mode 100644 > index 00000000000..01537deeb9f > --- /dev/null > +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c > @@ -0,0 +1,29 @@ > +/* { dg-do compile } */ > +/* { dg-options "-O3 -march=3Darmv8.2-a+dotprod -mautovec- > preference=3Dasimd-only --param vect-epilogues-nomask=3D0" } */ > + > +/* With dot product a 4x widening sum stays a single [SU]DOT, while a > + sum into 64-bit elements uses the pairwise widening instructions. */ > + > +int > +sum_u8_i (const unsigned char *a, long n) > +{ > + int s =3D 0; > + for (long i =3D 0; i < n; i++) > + s +=3D a[i]; > + return s; > +} > + > +long > +sum_u8_l (const unsigned char *a, long n) > +{ > + long s =3D 0; > + for (long i =3D 0; i < n; i++) > + s +=3D a[i]; > + return s; > +} > + > +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b,= v[0- > 9]+\.16b\n} 1 } } */ > +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16= b\n} > 1 } } */ > +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h= \n} 1 > } } */ > +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s= \n} 1 > } } */ > +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */ > -- > 2.50.1 (Apple Git-155)