RE: [PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum reductions

Tamar Christina <[email protected]> Tue, 4 Aug 2026 12:52:17 +0000
Newsgroups gmane.comp.gcc.patches
Message-ID <VI0PR08MB10392E606F946FBCA47CED357FFD42@VI0PR08MB10392.eurprd08.prod.outlook.com>
Hi Kyrill,

> -----Original Message-----
> From: [email protected] <[email protected]>
> Sent: 04 August 2026 13:39
> To: [email protected]
> Cc: Tamar Christina <[email protected]>; [email protected]; Kyrylo
> Tkachov <[email protected]>
> Subject: [PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum
> reductions
>=20
> From: Kyrylo Tkachov <[email protected]>
>=20
> The Advanced SIMD widen_[su]sum optabs only cover a single widening step,
> expanded as a dependent <su>addw + <su>addw2 pair, plus a 4x form that
> requires dot product.  A reduction into an accumulator that is more than
> twice as wide as the data therefore has to extend the input explicitly an=
d
> then issue one widening add per half vector.  Summing bytes into a 64-bit
> accumulator costs fifteen SIMD operations per 16 bytes of input.
>=20
> [SU]ADDLP and [SU]ADALP add adjacent lane pairs into the next wider
> element, so a chain of them expresses any power-of-two widening sum
> reduction in one operation per step.  The regrouping is exact because the
> sum of two elements always fits in the doubled element width, and the
> grouping of lanes inside a reduction accumulator is already unconstrained
> for WIDEN_SUM_EXPR, which the existing dot product based 4x expander also
> relies on.
>=20
> Expand the 2x forms as a single [SU]ADALP, add the missing V4SI <- V16QI
> and V2SI <- V8QI forms for !TARGET_DOTPROD, and add the V2DI <- V8HI and
> V2DI <- V16QI forms that no expander covered.  All of them are built by
> aarch64_expand_widen_sum, which halves the lane count with [SU]ADDLP
> until
> one pairwise step remains and then accumulates with [SU]ADALP.
>=20
> For a sum of unsigned char into long the inner loop changes from
>=20
> 	ldr	q30, [x1], 16
> 	zip1	v28.16b, v30.16b, v29.16b
> 	zip2	v30.16b, v30.16b, v29.16b
> 	zip1	v26.8h, v28.8h, v29.8h
> 	zip2	v28.8h, v28.8h, v29.8h
> 	zip1	v27.8h, v30.8h, v29.8h
> 	zip2	v30.8h, v30.8h, v29.8h
> 	uaddw	v31.2d, v31.2d, v26.2s
> 	uaddw2	v31.2d, v31.2d, v26.4s
> 	...  (six more uaddw/uaddw2)
>=20
> to
>=20
> 	ldr	q31, [x1], 16
> 	uaddlp	v31.8h, v31.16b
> 	uaddlp	v31.4s, v31.8h
> 	uadalp	v30.2d, v31.4s
>=20

Have you considered considered even for the b to d case to use dotprod
for the b -> s and then uadalp for the final 4 to d? or is the above codege=
n
the fallback for when !dotprod? wasn't quite clear..

That should still shave off 1 cycle.

Thanks,
Tamar

> and for a sum of int into long the saddw/saddw2 pair becomes one sadalp.
> On a Neoverse V2 core with an L1 resident working set this cuts the time =
of
> the
> byte loop by about 88% and of the int loop by about 68%.
>=20
> Bootstrapped and tested on aarch64-none-linux-gnu.
> Ok for trunk?
> Thanks,
> Kyrill
>=20
> gcc/ChangeLog:
>=20
> 	* config/aarch64/aarch64-protos.h (aarch64_expand_widen_sum):
> Declare.
> 	* config/aarch64/aarch64.cc (aarch64_expand_widen_sum): New
> function.
> 	* config/aarch64/aarch64-simd.md (aarch64_<su>adalp<mode>):
> Rename
> 	to ...
> 	(@aarch64_<su>adalp<mode>): ... this.
> 	(widen_ssum<Vdblw><mode>3, widen_usum<Vdblw><mode>3):
> Replace by ...
> 	(widen_<su>sum<Vdblw><mode>3): ... this.  Expand to [SU]ADALP.
> 	(widen_ssum<mode><vsi2qi>3, widen_usum<mode><vsi2qi>3):
> Replace
> 	by ...
> 	(widen_<su>sum<mode><vsi2qi>3): ... this.  Handle
> !TARGET_DOTPROD.
> 	(widen_<su>sumv2di<mode>3): New expander.
> 	* config/aarch64/iterators.md (VQ_BH): New mode iterator.
>=20
> gcc/testsuite/ChangeLog:
>=20
> 	* gcc.target/aarch64/pr122069_1.c: Update expected output.
> 	* gcc.target/aarch64/pr122069_3.c: Likewise.
> 	* gcc.target/aarch64/saddw-1.c: Renamed to...
> 	* gcc.target/aarch64/sadalp-1.c: ...this.  Update expected output.
> 	* gcc.target/aarch64/saddw-2.c: Renamed to...
> 	* gcc.target/aarch64/sadalp-2.c: ...this.  Update expected output.
> 	* gcc.target/aarch64/uaddw-1.c: Renamed to...
> 	* gcc.target/aarch64/uadalp-1.c: ...this.  Update expected output.
> 	* gcc.target/aarch64/uaddw-2.c: Renamed to...
> 	* gcc.target/aarch64/uadalp-2.c: ...this.  Update expected output.
> 	* gcc.target/aarch64/uaddw-3.c: Renamed to...
> 	* gcc.target/aarch64/uadalp-3.c: ...this.  Update expected output.
> 	* gcc.target/aarch64/widen_sum_pairwise_1.c: New test.
> 	* gcc.target/aarch64/widen_sum_pairwise_2.c: New test.
>=20
> Signed-off-by: Kyrylo Tkachov <[email protected]>
> ---
>  gcc/config/aarch64/aarch64-protos.h           |  1 +
>  gcc/config/aarch64/aarch64-simd.md            | 78 ++++++++-----------
>  gcc/config/aarch64/aarch64.cc                 | 27 +++++++
>  gcc/config/aarch64/iterators.md               |  4 +
>  gcc/testsuite/gcc.target/aarch64/pr122069_1.c | 11 +--
>  gcc/testsuite/gcc.target/aarch64/pr122069_3.c |  3 +-
>  .../aarch64/{saddw-1.c =3D> sadalp-1.c}         |  3 +-
>  .../aarch64/{saddw-2.c =3D> sadalp-2.c}         |  3 +-
>  .../aarch64/{uaddw-1.c =3D> uadalp-1.c}         |  3 +-
>  .../aarch64/{uaddw-2.c =3D> uadalp-2.c}         |  3 +-
>  .../aarch64/{uaddw-3.c =3D> uadalp-3.c}         |  3 +-
>  .../gcc.target/aarch64/widen_sum_pairwise_1.c | 39 ++++++++++
>  .../gcc.target/aarch64/widen_sum_pairwise_2.c | 29 +++++++
>  13 files changed, 141 insertions(+), 66 deletions(-)
>  rename gcc/testsuite/gcc.target/aarch64/{saddw-1.c =3D> sadalp-1.c} (74%=
)
>  rename gcc/testsuite/gcc.target/aarch64/{saddw-2.c =3D> sadalp-2.c} (74%=
)
>  rename gcc/testsuite/gcc.target/aarch64/{uaddw-1.c =3D> uadalp-1.c} (75%=
)
>  rename gcc/testsuite/gcc.target/aarch64/{uaddw-2.c =3D> uadalp-2.c} (75%=
)
>  rename gcc/testsuite/gcc.target/aarch64/{uaddw-3.c =3D> uadalp-3.c} (74%=
)
>  create mode 100644
> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
>  create mode 100644
> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
>=20
> diff --git a/gcc/config/aarch64/aarch64-protos.h
> b/gcc/config/aarch64/aarch64-protos.h
> index bcc833cfaa1..9303f12f80c 100644
> --- a/gcc/config/aarch64/aarch64-protos.h
> +++ b/gcc/config/aarch64/aarch64-protos.h
> @@ -1066,6 +1066,7 @@ void aarch64_emit_sve_pred_vec_duplicate
> (machine_mode, rtx, rtx);
>  void aarch64_expand_prologue (void);
>  void aarch64_decompose_vec_struct_index (machine_mode, rtx *, rtx *,
> bool);
>  void aarch64_expand_vector_init (rtx, rtx);
> +void aarch64_expand_widen_sum (rtx, rtx, rtx, rtx_code);
>  void aarch64_sve_expand_vector_init_subvector (rtx, rtx);
>  void aarch64_sve_expand_vector_init (rtx, rtx);
>  void aarch64_init_cumulative_args (CUMULATIVE_ARGS *, const_tree, rtx,
> diff --git a/gcc/config/aarch64/aarch64-simd.md
> b/gcc/config/aarch64/aarch64-simd.md
> index 433f16052bf..d119ac17352 100644
> --- a/gcc/config/aarch64/aarch64-simd.md
> +++ b/gcc/config/aarch64/aarch64-simd.md
> @@ -1182,7 +1182,7 @@
>    }
>  )
>=20
> -(define_expand "aarch64_<su>adalp<mode>"
> +(define_expand "@aarch64_<su>adalp<mode>"
>    [(set (match_operand:<VDBLW> 0 "register_operand")
>  	(plus:<VDBLW>
>  	  (plus:<VDBLW>
> @@ -5283,19 +5283,17 @@
>=20
>  ;; <su><addsub>w<q>.
>=20
> -(define_expand "widen_ssum<Vdblw><mode>3"
> +;; A widening sum reduction that halves the lane count is a single pairw=
ise
> +;; widening accumulate.
> +(define_expand "widen_<su>sum<Vdblw><mode>3"
>    [(set (match_operand:<VDBLW> 0 "register_operand")
> -	(plus:<VDBLW> (sign_extend:<VDBLW>
> -		        (match_operand:VQW 1 "register_operand"))
> +	(plus:<VDBLW> (ANY_EXTEND:<VDBLW>
> +			(match_operand:VQW 1 "register_operand"))
>  		      (match_operand:<VDBLW> 2 "register_operand")))]
>    "TARGET_SIMD"
>    {
> -    rtx p =3D aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, fal=
se);
> -    rtx temp =3D gen_reg_rtx (GET_MODE (operands[0]));
> -
> -    emit_insn (gen_aarch64_saddw<mode>_internal (temp, operands[2],
> -						operands[1], p));
> -    emit_insn (gen_aarch64_saddw2<mode> (operands[0], temp,
> operands[1]));
> +    emit_insn (gen_aarch64_<su>adalp<mode> (operands[0], operands[2],
> +					    operands[1]));
>      DONE;
>    }
>  )
> @@ -5311,23 +5309,6 @@
>    DONE;
>  })
>=20
> -(define_expand "widen_usum<Vdblw><mode>3"
> -  [(set (match_operand:<VDBLW> 0 "register_operand")
> -	(plus:<VDBLW> (zero_extend:<VDBLW>
> -		        (match_operand:VQW 1 "register_operand"))
> -		      (match_operand:<VDBLW> 2 "register_operand")))]
> -  "TARGET_SIMD"
> -  {
> -    rtx p =3D aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, fal=
se);
> -    rtx temp =3D gen_reg_rtx (GET_MODE (operands[0]));
> -
> -    emit_insn (gen_aarch64_uaddw<mode>_internal (temp, operands[2],
> -						 operands[1], p));
> -    emit_insn (gen_aarch64_uaddw2<mode> (operands[0], temp,
> operands[1]));
> -    DONE;
> -  }
> -)
> -
>  (define_expand "widen_usum<Vwide><mode>3"
>    [(set (match_operand:<VWIDE> 0 "register_operand")
>  	(plus:<VWIDE> (zero_extend:<VWIDE>
> @@ -5339,38 +5320,43 @@
>    DONE;
>  })
>=20
> -(define_expand "widen_ssum<mode><vsi2qi>3"
> +;; A widening sum reduction that quarters the lane count.  With dot prod=
uct
> +;; this is one [SU]DOT with a vector of ones, i.e. +=3D a becomes +=3D (=
a * 1).
> +;; Otherwise it is a pairwise widening add feeding a pairwise widening
> +;; accumulate.
> +(define_expand "widen_<su>sum<mode><vsi2qi>3"
>    [(set (match_operand:VS 0 "register_operand")
> -	(plus:VS (sign_extend:VS
> +	(plus:VS (ANY_EXTEND:VS
>  		   (match_operand:<VSI2QI> 1 "register_operand"))
>  		 (match_operand:VS 2 "register_operand")))]
> -  "TARGET_DOTPROD"
> +  "TARGET_SIMD"
>    {
> -    rtx ones =3D force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
> -    emit_insn (gen_sdot_prod<mode><vsi2qi> (operands[0], operands[1],
> ones,
> -					    operands[2]));
> +    if (TARGET_DOTPROD)
> +      {
> +	rtx ones =3D force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
> +	emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0],
> operands[1],
> +						   ones, operands[2]));
> +      }
> +    else
> +      aarch64_expand_widen_sum (operands[0], operands[2], operands[1],
> <CODE>);
>      DONE;
>    }
>  )
>=20
> -;; Use dot product to perform double widening sum reductions by
> -;; changing +=3D a into +=3D (a * 1).  i.e. we seed the multiplication w=
ith 1.
> -(define_expand "widen_usum<mode><vsi2qi>3"
> -  [(set (match_operand:VS 0 "register_operand")
> -	(plus:VS (zero_extend:VS
> -		        (match_operand:<VSI2QI> 1 "register_operand"))
> -		      (match_operand:VS 2 "register_operand")))]
> -  "TARGET_DOTPROD"
> +;; Widening sum reductions into 64-bit elements.  These need two or thre=
e
> +;; pairwise widening steps.
> +(define_expand "widen_<su>sumv2di<mode>3"
> +  [(set (match_operand:V2DI 0 "register_operand")
> +	(plus:V2DI (ANY_EXTEND:V2DI
> +		     (match_operand:VQ_BH 1 "register_operand"))
> +		   (match_operand:V2DI 2 "register_operand")))]
> +  "TARGET_SIMD"
>    {
> -    rtx ones =3D force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
> -    emit_insn (gen_udot_prod<mode><vsi2qi> (operands[0], operands[1],
> ones,
> -					    operands[2]));
> +    aarch64_expand_widen_sum (operands[0], operands[2], operands[1],
> <CODE>);
>      DONE;
>    }
>  )
>=20
> -;; Use dot product to perform double widening sum reductions by
> -;; changing +=3D a into +=3D (a * 1).  i.e. we seed the multiplication w=
ith 1.
>  (define_insn "aarch64_<ANY_EXTEND:su>subw<mode>"
>    [(set (match_operand:<VWIDE> 0 "register_operand" "=3Dw")
>  	(minus:<VWIDE> (match_operand:<VWIDE> 1 "register_operand"
> "w")
> diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.c=
c
> index d19ca305d82..628e94e8c40 100644
> --- a/gcc/config/aarch64/aarch64.cc
> +++ b/gcc/config/aarch64/aarch64.cc
> @@ -26327,6 +26327,33 @@ aarch64_expand_vector_init (rtx target, rtx
> vals)
>    emit_insn (seq_total_cost < fallback_seq_cost ? seq : fallback_seq);
>  }
>=20
> +/* Expand the widening sum reduction DEST =3D ACC + (WIDE) SRC, where th=
e
> +   Advanced SIMD vector SRC holds an even multiple of the number of lane=
s
> +   of the accumulator ACC and of the result DEST.  EXTEND_CODE is
> +   SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form.
> +   Halve the lane count with [SU]ADDLP until a single pairwise step is
> +   left, then accumulate into ACC with [SU]ADALP.  */
> +
> +void
> +aarch64_expand_widen_sum (rtx dest, rtx acc, rtx src, rtx_code extend_co=
de)
> +{
> +  unsigned int dest_nunits =3D GET_MODE_NUNITS (GET_MODE
> (dest)).to_constant ();
> +  machine_mode mode =3D GET_MODE (src);
> +  gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * 2)
> =3D=3D 0);
> +
> +  while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2)
> +    {
> +      insn_code icode =3D code_for_aarch64_addlp (extend_code, mode);
> +      mode =3D insn_data[icode].operand[0].mode;
> +      rtx tmp =3D gen_reg_rtx (mode);
> +      emit_insn (GEN_FCN (icode) (tmp, src));
> +      src =3D tmp;
> +    }
> +
> +  emit_insn (GEN_FCN (code_for_aarch64_adalp (extend_code, mode))
> (dest, acc,
> +								   src));
> +}
> +
>  /* Emit RTL corresponding to:
>     insr TARGET, ELEM.  */
>=20
> diff --git a/gcc/config/aarch64/iterators.md
> b/gcc/config/aarch64/iterators.md
> index 0d319751430..6a8c93cce37 100644
> --- a/gcc/config/aarch64/iterators.md
> +++ b/gcc/config/aarch64/iterators.md
> @@ -313,6 +313,10 @@
>  ;; All quad integer widen-able modes.
>  (define_mode_iterator VQW [V16QI V8HI V4SI])
>=20
> +;; Quad integer modes that reach 64-bit elements through more than one
> +;; pairwise widening step.
> +(define_mode_iterator VQ_BH [V16QI V8HI])
> +
>  ;; Double vector modes for combines.
>  (define_mode_iterator VDC [V8QI V4HI V4BF V4HF V2SI V2SF DI DF])
>=20
> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
> b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
> index b2f973261ea..d99b5493ade 100644
> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
> @@ -10,12 +10,8 @@ inline char char_abs(char i) {
>  ** foo_int:
>  ** 	...
>  ** 	sub	v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
> -** 	zip1	v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
> -** 	zip2	v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
> -** 	uaddw	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
> -** 	uaddw2	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
> -** 	uaddw	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
> -** 	uaddw2	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
> +** 	uaddlp	v[0-9]+.8h, v[0-9]+.16b
> +** 	uadalp	v[0-9]+.4s, v[0-9]+.8h
>  ** 	...
>  */
>  int foo_int(unsigned char *x, unsigned char * restrict y) {
> @@ -29,8 +25,7 @@ int foo_int(unsigned char *x, unsigned char * restrict =
y) {
>  ** foo2_int:
>  ** 	...
>  ** 	add	v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h
> -** 	uaddw	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
> -** 	uaddw2	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
> +** 	uadalp	v[0-9]+.4s, v[0-9]+.8h
>  ** 	...
>  */
>  int foo2_int(unsigned short *x, unsigned short * restrict y) {
> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
> b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
> index 0e832c43032..f29fc2b2ed4 100644
> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
> @@ -24,8 +24,7 @@ int foo_int(unsigned char *x, unsigned char * restrict =
y) {
>  ** foo2_int:
>  ** 	...
>  ** 	add	v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h
> -** 	uaddw	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
> -** 	uaddw2	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
> +** 	uadalp	v[0-9]+.4s, v[0-9]+.8h
>  ** 	...
>  */
>  int foo2_int(unsigned short *x, unsigned short * restrict y) {
> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-1.c
> b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c
> similarity index 74%
> rename from gcc/testsuite/gcc.target/aarch64/saddw-1.c
> rename to gcc/testsuite/gcc.target/aarch64/sadalp-1.c
> index f8871209b8a..61f9633f1a0 100644
> --- a/gcc/testsuite/gcc.target/aarch64/saddw-1.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c
> @@ -14,5 +14,4 @@ t6(int len, void * dummy, short * __restrict x)
>    return result;
>  }
>=20
> -/* { dg-final { scan-assembler "saddw" } } */
> -/* { dg-final { scan-assembler "saddw2" } } */
> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } =
*/
> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-2.c
> b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c
> similarity index 74%
> rename from gcc/testsuite/gcc.target/aarch64/saddw-2.c
> rename to gcc/testsuite/gcc.target/aarch64/sadalp-2.c
> index b9fc442a2f7..873fda2e1ea 100644
> --- a/gcc/testsuite/gcc.target/aarch64/saddw-2.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c
> @@ -14,5 +14,4 @@ t6(int len, void * dummy, int * __restrict x)
>    return result;
>  }
>=20
> -/* { dg-final { scan-assembler "saddw" } } */
> -/* { dg-final { scan-assembler "saddw2" } } */
> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s} } } =
*/
> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c
> b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c
> similarity index 75%
> rename from gcc/testsuite/gcc.target/aarch64/uaddw-1.c
> rename to gcc/testsuite/gcc.target/aarch64/uadalp-1.c
> index 14dff87d7f0..c4034384aae 100644
> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c
> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c
> @@ -14,5 +14,4 @@ t6(int len, void * dummy, unsigned short * __restrict x=
)
>    return result;
>  }
>=20
> -/* { dg-final { scan-assembler "uaddw" } } */
> -/* { dg-final { scan-assembler "uaddw2" } } */
> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } =
*/
> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c
> b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c
> similarity index 75%
> rename from gcc/testsuite/gcc.target/aarch64/uaddw-2.c
> rename to gcc/testsuite/gcc.target/aarch64/uadalp-2.c
> index 79d0d094fc3..395d36c7c00 100644
> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c
> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c
> @@ -14,6 +14,5 @@ t6(int len, void * dummy, unsigned short * __restrict x=
)
>    return result;
>  }
>=20
> -/* { dg-final { scan-assembler "uaddw" } } */
> -/* { dg-final { scan-assembler "uaddw2" } } */
> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } =
*/
>=20
> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c
> b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c
> similarity index 74%
> rename from gcc/testsuite/gcc.target/aarch64/uaddw-3.c
> rename to gcc/testsuite/gcc.target/aarch64/uadalp-3.c
> index 39cbd6b6cc2..5fdb1639ab8 100644
> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c
> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c
> @@ -14,5 +14,4 @@ t6(int len, void * dummy, char * __restrict x)
>    return result;
>  }
>=20
> -/* { dg-final { scan-assembler "uaddw" } } */
> -/* { dg-final { scan-assembler "uaddw2" } } */
> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.8h, v[0-9]+\.16b} } }=
 */
> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
> new file mode 100644
> index 00000000000..0aec0bf81c8
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
> @@ -0,0 +1,39 @@
> +/* { dg-do compile } */
> +/* { dg-options "-O3 -march=3Darmv8-a -mautovec-preference=3Dasimd-only =
--
> param vect-epilogues-nomask=3D0" } */
> +
> +/* Widening sum reductions should use the pairwise widening add and
> +   accumulate instructions rather than a chain of extensions feeding
> +   [SU]ADDW pairs.  */
> +
> +#define DEF(NAME, ITYPE, OTYPE)				\
> +  OTYPE NAME (const ITYPE *a, long n)			\
> +  {							\
> +    OTYPE s =3D 0;					\
> +    for (long i =3D 0; i < n; i++)			\
> +      s +=3D a[i];					\
> +    return s;						\
> +  }
> +
> +DEF (sum_u8_l, unsigned char, long)
> +DEF (sum_i8_l, signed char, long)
> +DEF (sum_u16_l, unsigned short, long)
> +DEF (sum_i16_l, short, long)
> +DEF (sum_u32_l, unsigned int, long)
> +DEF (sum_i32_l, int, long)
> +DEF (sum_u8_i, unsigned char, int)
> +DEF (sum_i8_i, signed char, int)
> +DEF (sum_u16_i, unsigned short, int)
> +DEF (sum_i16_i, short, int)
> +
> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16=
b\n}
> 2 } } */
> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.8h, v[0-9]+\.16=
b\n}
> 2 } } */
> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h=
\n} 2
> } } */
> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0-9]+\.8h=
\n} 2
> } } */
> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s=
\n} 3
> } } */
> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s=
\n} 3
> } } */
> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h=
\n} 2
> } } */
> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h=
\n} 2
> } } */
> +
> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
> +/* { dg-final { scan-assembler-not {\tsaddw2?\t} } } */
> +/* { dg-final { scan-assembler-not {\tzip1\t} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> new file mode 100644
> index 00000000000..01537deeb9f
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> @@ -0,0 +1,29 @@
> +/* { dg-do compile } */
> +/* { dg-options "-O3 -march=3Darmv8.2-a+dotprod -mautovec-
> preference=3Dasimd-only --param vect-epilogues-nomask=3D0" } */
> +
> +/* With dot product a 4x widening sum stays a single [SU]DOT, while a
> +   sum into 64-bit elements uses the pairwise widening instructions.  */
> +
> +int
> +sum_u8_i (const unsigned char *a, long n)
> +{
> +  int s =3D 0;
> +  for (long i =3D 0; i < n; i++)
> +    s +=3D a[i];
> +  return s;
> +}
> +
> +long
> +sum_u8_l (const unsigned char *a, long n)
> +{
> +  long s =3D 0;
> +  for (long i =3D 0; i < n; i++)
> +    s +=3D a[i];
> +  return s;
> +}
> +
> +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b,=
 v[0-
> 9]+\.16b\n} 1 } } */
> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16=
b\n}
> 1 } } */
> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h=
\n} 1
> } } */
> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s=
\n} 1
> } } */
> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
> --
> 2.50.1 (Apple Git-155)