[PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum reductions

<[email protected]> Tue, 4 Aug 2026 14:39:25 +0200
Newsgroups gmane.comp.gcc.patches
Message-ID <[email protected]>
From: Kyrylo Tkachov <[email protected]>

The Advanced SIMD widen_[su]sum optabs only cover a single widening step,
expanded as a dependent <su>addw + <su>addw2 pair, plus a 4x form that
requires dot product.  A reduction into an accumulator that is more than
twice as wide as the data therefore has to extend the input explicitly and
then issue one widening add per half vector.  Summing bytes into a 64-bit
accumulator costs fifteen SIMD operations per 16 bytes of input.

[SU]ADDLP and [SU]ADALP add adjacent lane pairs into the next wider
element, so a chain of them expresses any power-of-two widening sum
reduction in one operation per step.  The regrouping is exact because the
sum of two elements always fits in the doubled element width, and the
grouping of lanes inside a reduction accumulator is already unconstrained
for WIDEN_SUM_EXPR, which the existing dot product based 4x expander also
relies on.

Expand the 2x forms as a single [SU]ADALP, add the missing V4SI <- V16QI
and V2SI <- V8QI forms for !TARGET_DOTPROD, and add the V2DI <- V8HI and
V2DI <- V16QI forms that no expander covered.  All of them are built by
aarch64_expand_widen_sum, which halves the lane count with [SU]ADDLP until
one pairwise step remains and then accumulates with [SU]ADALP.

For a sum of unsigned char into long the inner loop changes from

	ldr	q30, [x1], 16
	zip1	v28.16b, v30.16b, v29.16b
	zip2	v30.16b, v30.16b, v29.16b
	zip1	v26.8h, v28.8h, v29.8h
	zip2	v28.8h, v28.8h, v29.8h
	zip1	v27.8h, v30.8h, v29.8h
	zip2	v30.8h, v30.8h, v29.8h
	uaddw	v31.2d, v31.2d, v26.2s
	uaddw2	v31.2d, v31.2d, v26.4s
	...  (six more uaddw/uaddw2)

to

	ldr	q31, [x1], 16
	uaddlp	v31.8h, v31.16b
	uaddlp	v31.4s, v31.8h
	uadalp	v30.2d, v31.4s

and for a sum of int into long the saddw/saddw2 pair becomes one sadalp.
On a Neoverse V2 core with an L1 resident working set this cuts the time of the
byte loop by about 88% and of the int loop by about 68%.

Bootstrapped and tested on aarch64-none-linux-gnu.
Ok for trunk?
Thanks,
Kyrill

gcc/ChangeLog:

	* config/aarch64/aarch64-protos.h (aarch64_expand_widen_sum): Declare.
	* config/aarch64/aarch64.cc (aarch64_expand_widen_sum): New function.
	* config/aarch64/aarch64-simd.md (aarch64_<su>adalp<mode>): Rename
	to ...
	(@aarch64_<su>adalp<mode>): ... this.
	(widen_ssum<Vdblw><mode>3, widen_usum<Vdblw><mode>3): Replace by ...
	(widen_<su>sum<Vdblw><mode>3): ... this.  Expand to [SU]ADALP.
	(widen_ssum<mode><vsi2qi>3, widen_usum<mode><vsi2qi>3): Replace
	by ...
	(widen_<su>sum<mode><vsi2qi>3): ... this.  Handle !TARGET_DOTPROD.
	(widen_<su>sumv2di<mode>3): New expander.
	* config/aarch64/iterators.md (VQ_BH): New mode iterator.

gcc/testsuite/ChangeLog:

	* gcc.target/aarch64/pr122069_1.c: Update expected output.
	* gcc.target/aarch64/pr122069_3.c: Likewise.
	* gcc.target/aarch64/saddw-1.c: Renamed to...
	* gcc.target/aarch64/sadalp-1.c: ...this.  Update expected output.
	* gcc.target/aarch64/saddw-2.c: Renamed to...
	* gcc.target/aarch64/sadalp-2.c: ...this.  Update expected output.
	* gcc.target/aarch64/uaddw-1.c: Renamed to...
	* gcc.target/aarch64/uadalp-1.c: ...this.  Update expected output.
	* gcc.target/aarch64/uaddw-2.c: Renamed to...
	* gcc.target/aarch64/uadalp-2.c: ...this.  Update expected output.
	* gcc.target/aarch64/uaddw-3.c: Renamed to...
	* gcc.target/aarch64/uadalp-3.c: ...this.  Update expected output.
	* gcc.target/aarch64/widen_sum_pairwise_1.c: New test.
	* gcc.target/aarch64/widen_sum_pairwise_2.c: New test.

Signed-off-by: Kyrylo Tkachov <[email protected]>
---
 gcc/config/aarch64/aarch64-protos.h           |  1 +
 gcc/config/aarch64/aarch64-simd.md            | 78 ++++++++-----------
 gcc/config/aarch64/aarch64.cc                 | 27 +++++++
 gcc/config/aarch64/iterators.md               |  4 +
 gcc/testsuite/gcc.target/aarch64/pr122069_1.c | 11 +--
 gcc/testsuite/gcc.target/aarch64/pr122069_3.c |  3 +-
 .../aarch64/{saddw-1.c => sadalp-1.c}         |  3 +-
 .../aarch64/{saddw-2.c => sadalp-2.c}         |  3 +-
 .../aarch64/{uaddw-1.c => uadalp-1.c}         |  3 +-
 .../aarch64/{uaddw-2.c => uadalp-2.c}         |  3 +-
 .../aarch64/{uaddw-3.c => uadalp-3.c}         |  3 +-
 .../gcc.target/aarch64/widen_sum_pairwise_1.c | 39 ++++++++++
 .../gcc.target/aarch64/widen_sum_pairwise_2.c | 29 +++++++
 13 files changed, 141 insertions(+), 66 deletions(-)
 rename gcc/testsuite/gcc.target/aarch64/{saddw-1.c => sadalp-1.c} (74%)
 rename gcc/testsuite/gcc.target/aarch64/{saddw-2.c => sadalp-2.c} (74%)
 rename gcc/testsuite/gcc.target/aarch64/{uaddw-1.c => uadalp-1.c} (75%)
 rename gcc/testsuite/gcc.target/aarch64/{uaddw-2.c => uadalp-2.c} (75%)
 rename gcc/testsuite/gcc.target/aarch64/{uaddw-3.c => uadalp-3.c} (74%)
 create mode 100644 gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
 create mode 100644 gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c

diff --git a/gcc/config/aarch64/aarch64-protos.h b/gcc/config/aarch64/aarch64-protos.h
index bcc833cfaa1..9303f12f80c 100644
--- a/gcc/config/aarch64/aarch64-protos.h
+++ b/gcc/config/aarch64/aarch64-protos.h
@@ -1066,6 +1066,7 @@ void aarch64_emit_sve_pred_vec_duplicate (machine_mode, rtx, rtx);
 void aarch64_expand_prologue (void);
 void aarch64_decompose_vec_struct_index (machine_mode, rtx *, rtx *, bool);
 void aarch64_expand_vector_init (rtx, rtx);
+void aarch64_expand_widen_sum (rtx, rtx, rtx, rtx_code);
 void aarch64_sve_expand_vector_init_subvector (rtx, rtx);
 void aarch64_sve_expand_vector_init (rtx, rtx);
 void aarch64_init_cumulative_args (CUMULATIVE_ARGS *, const_tree, rtx,
diff --git a/gcc/config/aarch64/aarch64-simd.md b/gcc/config/aarch64/aarch64-simd.md
index 433f16052bf..d119ac17352 100644
--- a/gcc/config/aarch64/aarch64-simd.md
+++ b/gcc/config/aarch64/aarch64-simd.md
@@ -1182,7 +1182,7 @@
   }
 )
 
-(define_expand "aarch64_<su>adalp<mode>"
+(define_expand "@aarch64_<su>adalp<mode>"
   [(set (match_operand:<VDBLW> 0 "register_operand")
 	(plus:<VDBLW>
 	  (plus:<VDBLW>
@@ -5283,19 +5283,17 @@
 
 ;; <su><addsub>w<q>.
 
-(define_expand "widen_ssum<Vdblw><mode>3"
+;; A widening sum reduction that halves the lane count is a single pairwise
+;; widening accumulate.
+(define_expand "widen_<su>sum<Vdblw><mode>3"
   [(set (match_operand:<VDBLW> 0 "register_operand")
-	(plus:<VDBLW> (sign_extend:<VDBLW> 
-		        (match_operand:VQW 1 "register_operand"))
+	(plus:<VDBLW> (ANY_EXTEND:<VDBLW>
+			(match_operand:VQW 1 "register_operand"))
 		      (match_operand:<VDBLW> 2 "register_operand")))]
   "TARGET_SIMD"
   {
-    rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false);
-    rtx temp = gen_reg_rtx (GET_MODE (operands[0]));
-
-    emit_insn (gen_aarch64_saddw<mode>_internal (temp, operands[2],
-						operands[1], p));
-    emit_insn (gen_aarch64_saddw2<mode> (operands[0], temp, operands[1]));
+    emit_insn (gen_aarch64_<su>adalp<mode> (operands[0], operands[2],
+					    operands[1]));
     DONE;
   }
 )
@@ -5311,23 +5309,6 @@
   DONE;
 })
 
-(define_expand "widen_usum<Vdblw><mode>3"
-  [(set (match_operand:<VDBLW> 0 "register_operand")
-	(plus:<VDBLW> (zero_extend:<VDBLW> 
-		        (match_operand:VQW 1 "register_operand"))
-		      (match_operand:<VDBLW> 2 "register_operand")))]
-  "TARGET_SIMD"
-  {
-    rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false);
-    rtx temp = gen_reg_rtx (GET_MODE (operands[0]));
-
-    emit_insn (gen_aarch64_uaddw<mode>_internal (temp, operands[2],
-						 operands[1], p));
-    emit_insn (gen_aarch64_uaddw2<mode> (operands[0], temp, operands[1]));
-    DONE;
-  }
-)
-
 (define_expand "widen_usum<Vwide><mode>3"
   [(set (match_operand:<VWIDE> 0 "register_operand")
 	(plus:<VWIDE> (zero_extend:<VWIDE>
@@ -5339,38 +5320,43 @@
   DONE;
 })
 
-(define_expand "widen_ssum<mode><vsi2qi>3"
+;; A widening sum reduction that quarters the lane count.  With dot product
+;; this is one [SU]DOT with a vector of ones, i.e. += a becomes += (a * 1).
+;; Otherwise it is a pairwise widening add feeding a pairwise widening
+;; accumulate.
+(define_expand "widen_<su>sum<mode><vsi2qi>3"
   [(set (match_operand:VS 0 "register_operand")
-	(plus:VS (sign_extend:VS
+	(plus:VS (ANY_EXTEND:VS
 		   (match_operand:<VSI2QI> 1 "register_operand"))
 		 (match_operand:VS 2 "register_operand")))]
-  "TARGET_DOTPROD"
+  "TARGET_SIMD"
   {
-    rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
-    emit_insn (gen_sdot_prod<mode><vsi2qi> (operands[0], operands[1], ones,
-					    operands[2]));
+    if (TARGET_DOTPROD)
+      {
+	rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
+	emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0], operands[1],
+						   ones, operands[2]));
+      }
+    else
+      aarch64_expand_widen_sum (operands[0], operands[2], operands[1], <CODE>);
     DONE;
   }
 )
 
-;; Use dot product to perform double widening sum reductions by
-;; changing += a into += (a * 1).  i.e. we seed the multiplication with 1.
-(define_expand "widen_usum<mode><vsi2qi>3"
-  [(set (match_operand:VS 0 "register_operand")
-	(plus:VS (zero_extend:VS
-		        (match_operand:<VSI2QI> 1 "register_operand"))
-		      (match_operand:VS 2 "register_operand")))]
-  "TARGET_DOTPROD"
+;; Widening sum reductions into 64-bit elements.  These need two or three
+;; pairwise widening steps.
+(define_expand "widen_<su>sumv2di<mode>3"
+  [(set (match_operand:V2DI 0 "register_operand")
+	(plus:V2DI (ANY_EXTEND:V2DI
+		     (match_operand:VQ_BH 1 "register_operand"))
+		   (match_operand:V2DI 2 "register_operand")))]
+  "TARGET_SIMD"
   {
-    rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
-    emit_insn (gen_udot_prod<mode><vsi2qi> (operands[0], operands[1], ones,
-					    operands[2]));
+    aarch64_expand_widen_sum (operands[0], operands[2], operands[1], <CODE>);
     DONE;
   }
 )
 
-;; Use dot product to perform double widening sum reductions by
-;; changing += a into += (a * 1).  i.e. we seed the multiplication with 1.
 (define_insn "aarch64_<ANY_EXTEND:su>subw<mode>"
   [(set (match_operand:<VWIDE> 0 "register_operand" "=w")
 	(minus:<VWIDE> (match_operand:<VWIDE> 1 "register_operand" "w")
diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc
index d19ca305d82..628e94e8c40 100644
--- a/gcc/config/aarch64/aarch64.cc
+++ b/gcc/config/aarch64/aarch64.cc
@@ -26327,6 +26327,33 @@ aarch64_expand_vector_init (rtx target, rtx vals)
   emit_insn (seq_total_cost < fallback_seq_cost ? seq : fallback_seq);
 }
 
+/* Expand the widening sum reduction DEST = ACC + (WIDE) SRC, where the
+   Advanced SIMD vector SRC holds an even multiple of the number of lanes
+   of the accumulator ACC and of the result DEST.  EXTEND_CODE is
+   SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form.
+   Halve the lane count with [SU]ADDLP until a single pairwise step is
+   left, then accumulate into ACC with [SU]ADALP.  */
+
+void
+aarch64_expand_widen_sum (rtx dest, rtx acc, rtx src, rtx_code extend_code)
+{
+  unsigned int dest_nunits = GET_MODE_NUNITS (GET_MODE (dest)).to_constant ();
+  machine_mode mode = GET_MODE (src);
+  gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * 2) == 0);
+
+  while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2)
+    {
+      insn_code icode = code_for_aarch64_addlp (extend_code, mode);
+      mode = insn_data[icode].operand[0].mode;
+      rtx tmp = gen_reg_rtx (mode);
+      emit_insn (GEN_FCN (icode) (tmp, src));
+      src = tmp;
+    }
+
+  emit_insn (GEN_FCN (code_for_aarch64_adalp (extend_code, mode)) (dest, acc,
+								   src));
+}
+
 /* Emit RTL corresponding to:
    insr TARGET, ELEM.  */
 
diff --git a/gcc/config/aarch64/iterators.md b/gcc/config/aarch64/iterators.md
index 0d319751430..6a8c93cce37 100644
--- a/gcc/config/aarch64/iterators.md
+++ b/gcc/config/aarch64/iterators.md
@@ -313,6 +313,10 @@
 ;; All quad integer widen-able modes.
 (define_mode_iterator VQW [V16QI V8HI V4SI])
 
+;; Quad integer modes that reach 64-bit elements through more than one
+;; pairwise widening step.
+(define_mode_iterator VQ_BH [V16QI V8HI])
+
 ;; Double vector modes for combines.
 (define_mode_iterator VDC [V8QI V4HI V4BF V4HF V2SI V2SF DI DF])
 
diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
index b2f973261ea..d99b5493ade 100644
--- a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
+++ b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
@@ -10,12 +10,8 @@ inline char char_abs(char i) {
 ** foo_int:
 ** 	...
 ** 	sub	v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
-** 	zip1	v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
-** 	zip2	v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
-** 	uaddw	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
-** 	uaddw2	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
-** 	uaddw	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
-** 	uaddw2	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
+** 	uaddlp	v[0-9]+.8h, v[0-9]+.16b
+** 	uadalp	v[0-9]+.4s, v[0-9]+.8h
 ** 	...
 */
 int foo_int(unsigned char *x, unsigned char * restrict y) {
@@ -29,8 +25,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) {
 ** foo2_int:
 ** 	...
 ** 	add	v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h
-** 	uaddw	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
-** 	uaddw2	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
+** 	uadalp	v[0-9]+.4s, v[0-9]+.8h
 ** 	...
 */
 int foo2_int(unsigned short *x, unsigned short * restrict y) {
diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
index 0e832c43032..f29fc2b2ed4 100644
--- a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
+++ b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
@@ -24,8 +24,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) {
 ** foo2_int:
 ** 	...
 ** 	add	v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h
-** 	uaddw	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
-** 	uaddw2	v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
+** 	uadalp	v[0-9]+.4s, v[0-9]+.8h
 ** 	...
 */
 int foo2_int(unsigned short *x, unsigned short * restrict y) {
diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-1.c b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c
similarity index 74%
rename from gcc/testsuite/gcc.target/aarch64/saddw-1.c
rename to gcc/testsuite/gcc.target/aarch64/sadalp-1.c
index f8871209b8a..61f9633f1a0 100644
--- a/gcc/testsuite/gcc.target/aarch64/saddw-1.c
+++ b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c
@@ -14,5 +14,4 @@ t6(int len, void * dummy, short * __restrict x)
   return result;
 }
 
-/* { dg-final { scan-assembler "saddw" } } */
-/* { dg-final { scan-assembler "saddw2" } } */
+/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-2.c b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c
similarity index 74%
rename from gcc/testsuite/gcc.target/aarch64/saddw-2.c
rename to gcc/testsuite/gcc.target/aarch64/sadalp-2.c
index b9fc442a2f7..873fda2e1ea 100644
--- a/gcc/testsuite/gcc.target/aarch64/saddw-2.c
+++ b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c
@@ -14,5 +14,4 @@ t6(int len, void * dummy, int * __restrict x)
   return result;
 }
 
-/* { dg-final { scan-assembler "saddw" } } */
-/* { dg-final { scan-assembler "saddw2" } } */
+/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s} } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c
similarity index 75%
rename from gcc/testsuite/gcc.target/aarch64/uaddw-1.c
rename to gcc/testsuite/gcc.target/aarch64/uadalp-1.c
index 14dff87d7f0..c4034384aae 100644
--- a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c
+++ b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c
@@ -14,5 +14,4 @@ t6(int len, void * dummy, unsigned short * __restrict x)
   return result;
 }
 
-/* { dg-final { scan-assembler "uaddw" } } */
-/* { dg-final { scan-assembler "uaddw2" } } */
+/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c
similarity index 75%
rename from gcc/testsuite/gcc.target/aarch64/uaddw-2.c
rename to gcc/testsuite/gcc.target/aarch64/uadalp-2.c
index 79d0d094fc3..395d36c7c00 100644
--- a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c
+++ b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c
@@ -14,6 +14,5 @@ t6(int len, void * dummy, unsigned short * __restrict x)
   return result;
 }
 
-/* { dg-final { scan-assembler "uaddw" } } */
-/* { dg-final { scan-assembler "uaddw2" } } */
+/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
 
diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c
similarity index 74%
rename from gcc/testsuite/gcc.target/aarch64/uaddw-3.c
rename to gcc/testsuite/gcc.target/aarch64/uadalp-3.c
index 39cbd6b6cc2..5fdb1639ab8 100644
--- a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c
+++ b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c
@@ -14,5 +14,4 @@ t6(int len, void * dummy, char * __restrict x)
   return result;
 }
 
-/* { dg-final { scan-assembler "uaddw" } } */
-/* { dg-final { scan-assembler "uaddw2" } } */
+/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.8h, v[0-9]+\.16b} } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
new file mode 100644
index 00000000000..0aec0bf81c8
--- /dev/null
+++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
@@ -0,0 +1,39 @@
+/* { dg-do compile } */
+/* { dg-options "-O3 -march=armv8-a -mautovec-preference=asimd-only --param vect-epilogues-nomask=0" } */
+
+/* Widening sum reductions should use the pairwise widening add and
+   accumulate instructions rather than a chain of extensions feeding
+   [SU]ADDW pairs.  */
+
+#define DEF(NAME, ITYPE, OTYPE)				\
+  OTYPE NAME (const ITYPE *a, long n)			\
+  {							\
+    OTYPE s = 0;					\
+    for (long i = 0; i < n; i++)			\
+      s += a[i];					\
+    return s;						\
+  }
+
+DEF (sum_u8_l, unsigned char, long)
+DEF (sum_i8_l, signed char, long)
+DEF (sum_u16_l, unsigned short, long)
+DEF (sum_i16_l, short, long)
+DEF (sum_u32_l, unsigned int, long)
+DEF (sum_i32_l, int, long)
+DEF (sum_u8_i, unsigned char, int)
+DEF (sum_i8_i, signed char, int)
+DEF (sum_u16_i, unsigned short, int)
+DEF (sum_i16_i, short, int)
+
+/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n} 2 } } */
+/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n} 2 } } */
+/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 2 } } */
+/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 2 } } */
+/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 3 } } */
+/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 3 } } */
+/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 2 } } */
+/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 2 } } */
+
+/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
+/* { dg-final { scan-assembler-not {\tsaddw2?\t} } } */
+/* { dg-final { scan-assembler-not {\tzip1\t} } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
new file mode 100644
index 00000000000..01537deeb9f
--- /dev/null
+++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
@@ -0,0 +1,29 @@
+/* { dg-do compile } */
+/* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec-preference=asimd-only --param vect-epilogues-nomask=0" } */
+
+/* With dot product a 4x widening sum stays a single [SU]DOT, while a
+   sum into 64-bit elements uses the pairwise widening instructions.  */
+
+int
+sum_u8_i (const unsigned char *a, long n)
+{
+  int s = 0;
+  for (long i = 0; i < n; i++)
+    s += a[i];
+  return s;
+}
+
+long
+sum_u8_l (const unsigned char *a, long n)
+{
+  long s = 0;
+  for (long i = 0; i < n; i++)
+    s += a[i];
+  return s;
+}
+
+/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, v[0-9]+\.16b\n} 1 } } */
+/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n} 1 } } */
+/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 1 } } */
+/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 1 } } */
+/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
-- 
2.50.1 (Apple Git-155)