RE: [PATCH] aarch64: complex multiply with FCMLA without -fno-signed-zeros [PR126589]

Tamar Christina <[email protected]>
Newsgroups gmane.comp.gcc.patches
Message-ID <GVXPR08MB1040589C99BE2CC0BA85A0960FFDC2@GVXPR08MB10405.eurprd08.prod.outlook.com>
> -----Original Message-----
> From: [email protected] <[email protected]>
> Sent: 12 August 2026 14:11
> To: [email protected]
> Cc: Tamar Christina <[email protected]>; Kyrylo Tkachov
> <[email protected]>
> Subject: [PATCH] aarch64: complex multiply with FCMLA without -fno-signed-
> zeros [PR126589]
> 
> From: Kyrylo Tkachov <[email protected]>
> 
> FCMLA only accumulates, so a plain complex multiply was built by seeding the
> accumulator with +0.0.  That turns a -0.0 product into +0.0, so the fix for
> PR126589 gave the cmul optabs a !HONOR_SIGNED_ZEROS requirement,
> which means
> FCMLA is only reached with -ffast-math and not with -fcx-limited-range alone.
> 
> The first rotation of the pair computes
> 
>   Vd.re = Vd.re + Vn.re * Vm.re
>   Vd.im = Vd.im + Vn.re * Vm.im
> 
> that is, a multiply of Vm by the real part of Vn broadcast over its own
> pair.  TRN1 of Vn with itself is that broadcast, so the accumulator is not
> needed at all:
> 
>   trn1  vt.4s, vn.4s, vn.4s
>   fmul  vd.4s, vt.4s, vm.4s
>   fcmla vd.4s, vn.4s, vm.4s, #90
> 
> The sign of a zero now survives, and there is one fewer accumulate.  Where
> signed zeros are not honoured the previous sequence is kept: MOVI with a
> zero
> immediate is a zeroing idiom, so seeding the accumulator there is free, while
> materialising any other constant is not.
> 
> The same applies to the SVE expander.
> 
> On Neoverse V2, over arrays of 4096 _Complex float at -O3 -mcpu=neoverse-
> v2
> -fcx-limited-range, c[i] = a[i] * b[i] drops 5.9% and y[i] += alpha * x[i]
> drops 21.8% in run time.  -ffast-math is unchanged.
> 
> Bootstrapped and tested on aarch64-none-linux-gnu.
> Ok for trunk?

Sure, though this should be reverted once PR121925 lands because it
hides the permutations from the vectorizer.

For now this is fine though.

Thanks,
Tamar

> Thanks,
> Kyrill
> 
> gcc/ChangeLog:
> 
> 	PR target/126589
> 	* config/aarch64/aarch64-protos.h (aarch64_emit_cmul_first_half):
> 	Declare.
> 	* config/aarch64/aarch64.cc (aarch64_emit_cmul_first_half): New
> 	function.
> 	* config/aarch64/aarch64-simd.md (cmul<conj_op><mode>3): Use it
> when
> 	signed zeros are honoured, and drop the requirement.
> 	* config/aarch64/aarch64-sve.md (cmul<conj_op><mode>3):
> Likewise.
> 
> gcc/testsuite/ChangeLog:
> 
> 	PR target/126589
> 	* gcc.target/aarch64/complex-mul-signed-zeros.c: New test.
> 	* gcc.target/aarch64/sve/complex-mul-signed-zeros.c: New test.
> 	* gcc.dg/vect/complex/complex-mul-signed-zero-run.c: New test.
> 
> Signed-off-by: Kyrylo Tkachov <[email protected]>
> ---
>  gcc/config/aarch64/aarch64-protos.h           |   1 +
>  gcc/config/aarch64/aarch64-simd.md            |  19 +++-
>  gcc/config/aarch64/aarch64-sve.md             |  19 +++-
>  gcc/config/aarch64/aarch64.cc                 |  24 ++++
>  .../complex/complex-mul-signed-zero-run.c     | 105 ++++++++++++++++++
>  .../aarch64/complex-mul-signed-zeros.c        |  39 +++++++
>  .../aarch64/sve/complex-mul-signed-zeros.c    |  27 +++++
>  7 files changed, 224 insertions(+), 10 deletions(-)
>  create mode 100644 gcc/testsuite/gcc.dg/vect/complex/complex-mul-
> signed-zero-run.c
>  create mode 100644 gcc/testsuite/gcc.target/aarch64/complex-mul-signed-
> zeros.c
>  create mode 100644 gcc/testsuite/gcc.target/aarch64/sve/complex-mul-
> signed-zeros.c
> 
> diff --git a/gcc/config/aarch64/aarch64-protos.h
> b/gcc/config/aarch64/aarch64-protos.h
> index bcc833cfaa1..24668aca125 100644
> --- a/gcc/config/aarch64/aarch64-protos.h
> +++ b/gcc/config/aarch64/aarch64-protos.h
> @@ -1047,6 +1047,7 @@ rtx aarch64_convert_sve_data_to_pred (rtx, rtx);
>  rtx aarch64_expand_sve_dupq (rtx, machine_mode, rtx);
>  void aarch64_expand_mov_immediate (rtx, rtx);
>  rtx aarch64_stack_protect_canary_mem (machine_mode, rtx,
> aarch64_salt_type);
> +void aarch64_emit_cmul_first_half (rtx, rtx, rtx);
>  rtx aarch64_ptrue_reg (machine_mode);
>  rtx aarch64_ptrue_reg (machine_mode, unsigned int);
>  rtx aarch64_ptrue_reg (machine_mode, machine_mode);
> diff --git a/gcc/config/aarch64/aarch64-simd.md
> b/gcc/config/aarch64/aarch64-simd.md
> index 8ad1ec20f46..48ca6b10d51 100644
> --- a/gcc/config/aarch64/aarch64-simd.md
> +++ b/gcc/config/aarch64/aarch64-simd.md
> @@ -681,12 +681,23 @@
>  	(unspec:VHSDF [(match_operand:VHSDF 1 "register_operand")
>  		       (match_operand:VHSDF 2 "register_operand")]
>  		       FCMUL_OP))]
> -  "TARGET_COMPLEX && !BYTES_BIG_ENDIAN && !HONOR_SIGNED_ZEROS
> (<MODE>mode)"
> +  "TARGET_COMPLEX && !BYTES_BIG_ENDIAN"
>  {
> -  rtx tmp = force_reg (<MODE>mode, CONST0_RTX (<MODE>mode));
>    rtx res1 = gen_reg_rtx (<MODE>mode);
> -  emit_insn (gen_aarch64_fcmla<rotsplit1><mode> (res1, tmp,
> -						 operands[2], operands[1]));
> +  if (HONOR_SIGNED_ZEROS (<MODE>mode))
> +    /* The first rotation multiplies the real part of operand 2 by the whole
> +       of operand 1, so express it as a multiply rather than as an accumulate
> +       into zero.  Accumulating into +0.0 would turn a -0.0 product into
> +       +0.0.  */
> +    aarch64_emit_cmul_first_half (res1, operands[2], operands[1]);
> +  else
> +    {
> +      /* MOVI with a zero immediate is a zeroing idiom, so seeding the
> +	 accumulator here is free.  */
> +      rtx zero = force_reg (<MODE>mode, CONST0_RTX (<MODE>mode));
> +      emit_insn (gen_aarch64_fcmla<rotsplit1><mode> (res1, zero,
> +						     operands[2],
> operands[1]));
> +    }
>    emit_insn (gen_aarch64_fcmla<rotsplit2><mode> (operands[0], res1,
>  						 operands[2], operands[1]));
>    DONE;
> diff --git a/gcc/config/aarch64/aarch64-sve.md
> b/gcc/config/aarch64/aarch64-sve.md
> index 5f9a19c42e9..539030ab1bc 100644
> --- a/gcc/config/aarch64/aarch64-sve.md
> +++ b/gcc/config/aarch64/aarch64-sve.md
> @@ -8300,16 +8300,23 @@
>  	   [(match_operand:SVE_FULL_F 1 "register_operand")
>  	    (match_operand:SVE_FULL_F 2 "register_operand")]
>  	  FCMUL_OP))]
> -  "TARGET_SVE && !HONOR_SIGNED_ZEROS (<MODE>mode)"
> +  "TARGET_SVE"
>  {
>    rtx pred_reg = aarch64_ptrue_reg (<VPRED>mode);
>    rtx gp_mode = gen_int_mode (SVE_RELAXED_GP, SImode);
> -  rtx accum = force_reg (<MODE>mode, CONST0_RTX (<MODE>mode));
>    rtx tmp = gen_reg_rtx (<MODE>mode);
> -  emit_insn
> -    (gen_aarch64_pred_fcmla<sve_rot1><mode> (tmp, pred_reg,
> -					     operands[2], operands[1],
> -					     accum, gp_mode));
> +  if (HONOR_SIGNED_ZEROS (<MODE>mode))
> +    /* Accumulating into +0.0 would turn a -0.0 product into +0.0, so do the
> +       first rotation as the equivalent explicit multiply.  */
> +    aarch64_emit_cmul_first_half (tmp, operands[2], operands[1]);
> +  else
> +    {
> +      rtx accum = force_reg (<MODE>mode, CONST0_RTX (<MODE>mode));
> +      emit_insn
> +	(gen_aarch64_pred_fcmla<sve_rot1><mode> (tmp, pred_reg,
> +						 operands[2], operands[1],
> +						 accum, gp_mode));
> +    }
>    emit_insn
>      (gen_aarch64_pred_fcmla<sve_rot2><mode> (operands[0], pred_reg,
>  					     operands[2], operands[1],
> diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc
> index 3041a6ee62a..1a2c578381c 100644
> --- a/gcc/config/aarch64/aarch64.cc
> +++ b/gcc/config/aarch64/aarch64.cc
> @@ -4229,6 +4229,30 @@ aarch64_ptrue_reg (machine_mode pred_mode,
> machine_mode data_mode)
>    return aarch64_ptrue_reg (pred_mode, size);
>  }
> 
> +/* Emit into TARGET the first half of the complex multiplication of A by B,
> +   that is the pair (Are * Bre, Are * Bim) for each complex element.  This is
> +   what FCMLA with a rotation of #0 computes, but as a plain multiply, so no
> +   additive identity is needed and the sign of a zero product survives.
> +   TARGET, A and B all have the same floating-point vector mode.  */
> +
> +void
> +aarch64_emit_cmul_first_half (rtx target, rtx a, rtx b)
> +{
> +  machine_mode mode = GET_MODE (target);
> +  rtx dup_real = gen_reg_rtx (mode);
> +  /* TRN1 of a vector with itself broadcasts each even element over the
> +     odd one, which for complex data is the real part over its own pair.
> +     The two inputs are the same register, so no big-endian correction is
> +     needed here.  */
> +  emit_set_insn (dup_real,
> +		 gen_rtx_UNSPEC (mode, gen_rtvec (2, a, a), UNSPEC_TRN1));
> +  rtx res = expand_binop (mode, smul_optab, dup_real, b, target, 0,
> +			  OPTAB_DIRECT);
> +  gcc_assert (res);
> +  if (res != target)
> +    emit_move_insn (target, res);
> +}
> +
>  /* Return an all-false predicate register of mode MODE.  */
> 
>  rtx
> diff --git a/gcc/testsuite/gcc.dg/vect/complex/complex-mul-signed-zero-run.c
> b/gcc/testsuite/gcc.dg/vect/complex/complex-mul-signed-zero-run.c
> new file mode 100644
> index 00000000000..0e5f98144c6
> --- /dev/null
> +++ b/gcc/testsuite/gcc.dg/vect/complex/complex-mul-signed-zero-run.c
> @@ -0,0 +1,105 @@
> +/* { dg-do run } */
> +/* { dg-require-effective-target vect_complex_add_double } */
> +/* { dg-additional-options "-fcx-limited-range" } */
> +/* { dg-add-options arm_v8_3a_complex_neon } */
> +
> +/* Vectorised complex multiplication must keep the sign of a zero result
> +   when -fno-signed-zeros is not in effect.  */
> +
> +#include <math.h>
> +
> +#define N 64
> +
> +_Complex float af[N], bf[N], cf[N], reff[N];
> +_Complex double ad[N], bd[N], cd[N], refd[N];
> +
> +__attribute__((noipa)) void
> +mulf (_Complex float *__restrict c, _Complex float *__restrict a,
> +      _Complex float *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * b[i];
> +}
> +
> +__attribute__((noipa, optimize ("no-tree-vectorize"))) void
> +mulf_ref (_Complex float *__restrict c, _Complex float *__restrict a,
> +	  _Complex float *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * b[i];
> +}
> +
> +__attribute__((noipa)) void
> +muld (_Complex double *__restrict c, _Complex double *__restrict a,
> +      _Complex double *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * b[i];
> +}
> +
> +__attribute__((noipa, optimize ("no-tree-vectorize"))) void
> +muld_ref (_Complex double *__restrict c, _Complex double *__restrict a,
> +	  _Complex double *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * b[i];
> +}
> +
> +/* Compare bit patterns so that the sign of a zero matters.  Any two NaNs
> +   are equal, their sign and payload are unspecified.  */
> +
> +static int
> +eqf (float x, float y)
> +{
> +  if (isnan (x) && isnan (y))
> +    return 1;
> +  return __builtin_memcmp (&x, &y, sizeof x) == 0;
> +}
> +
> +static int
> +eqd (double x, double y)
> +{
> +  if (isnan (x) && isnan (y))
> +    return 1;
> +  return __builtin_memcmp (&x, &y, sizeof x) == 0;
> +}
> +
> +static const float vals[]
> +  = { 0.0f, -0.0f, 1.0f, -1.0f, 3.5f, -2.25f,
> +      __builtin_inff (), -__builtin_inff (), __builtin_nanf ("") };
> +#define NV ((int) (sizeof (vals) / sizeof (vals[0])))
> +
> +int
> +main (void)
> +{
> +  for (int i0 = 0; i0 < NV; i0++)
> +    for (int i1 = 0; i1 < NV; i1++)
> +      for (int i2 = 0; i2 < NV; i2++)
> +	for (int i3 = 0; i3 < NV; i3++)
> +	  {
> +	    for (int k = 0; k < N; k++)
> +	      {
> +		__real__ af[k] = vals[i0];
> +		__imag__ af[k] = vals[i1];
> +		__real__ bf[k] = vals[i2];
> +		__imag__ bf[k] = vals[i3];
> +		__real__ ad[k] = vals[i0];
> +		__imag__ ad[k] = vals[i1];
> +		__real__ bd[k] = vals[i2];
> +		__imag__ bd[k] = vals[i3];
> +	      }
> +
> +	    mulf (cf, af, bf, N);
> +	    mulf_ref (reff, af, bf, N);
> +	    muld (cd, ad, bd, N);
> +	    muld_ref (refd, ad, bd, N);
> +
> +	    for (int k = 0; k < N; k++)
> +	      if (!eqf (__real__ cf[k], __real__ reff[k])
> +		  || !eqf (__imag__ cf[k], __imag__ reff[k])
> +		  || !eqd (__real__ cd[k], __real__ refd[k])
> +		  || !eqd (__imag__ cd[k], __imag__ refd[k]))
> +		__builtin_abort ();
> +	  }
> +  return 0;
> +}
> diff --git a/gcc/testsuite/gcc.target/aarch64/complex-mul-signed-zeros.c
> b/gcc/testsuite/gcc.target/aarch64/complex-mul-signed-zeros.c
> new file mode 100644
> index 00000000000..a15accf8e07
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/aarch64/complex-mul-signed-zeros.c
> @@ -0,0 +1,39 @@
> +/* { dg-do compile } */
> +/* { dg-options "-O3 -march=armv8.3-a+nosve -fcx-limited-range" } */
> +
> +/* FCMLA only accumulates, so seeding it with +0.0 to get a plain complex
> +   multiply would turn a -0.0 product into +0.0.  Doing the first rotation as
> +   an explicit multiply keeps the sign of a zero and needs no accumulator, so
> +   this works without -fno-signed-zeros.  */
> +
> +void
> +mulf (_Complex float *__restrict c, _Complex float *__restrict a,
> +      _Complex float *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * b[i];
> +}
> +
> +void
> +muld (_Complex double *__restrict c, _Complex double *__restrict a,
> +      _Complex double *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * b[i];
> +}
> +
> +void
> +mulconjf (_Complex float *__restrict c, _Complex float *__restrict a,
> +	  _Complex float *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * ~b[i];
> +}
> +
> +/* { dg-final { scan-assembler {trn1\tv[0-9]+\.4s, v[0-9]+\.4s, v[0-9]+\.4s} } }
> */
> +/* { dg-final { scan-assembler {trn1\tv[0-9]+\.2d, v[0-9]+\.2d, v[0-9]+\.2d} }
> } */
> +/* { dg-final { scan-assembler {fcmla\tv[0-9]+\.4s, v[0-9]+\.4s, v[0-9]+\.4s,
> #90} } } */
> +/* { dg-final { scan-assembler {fcmla\tv[0-9]+\.2d, v[0-9]+\.2d, v[0-9]+\.2d,
> #90} } } */
> +/* { dg-final { scan-assembler {fcmla\tv[0-9]+\.4s, v[0-9]+\.4s, v[0-9]+\.4s,
> #270} } } */
> +/* No accumulator is needed, so no zero constant is materialised.  */
> +/* { dg-final { scan-assembler-not {fcmla\tv[0-9]+\.[24][sd], v[0-
> 9]+\.[24][sd], v[0-9]+\.[24][sd], #0} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/sve/complex-mul-signed-zeros.c
> b/gcc/testsuite/gcc.target/aarch64/sve/complex-mul-signed-zeros.c
> new file mode 100644
> index 00000000000..a79332460f5
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/aarch64/sve/complex-mul-signed-zeros.c
> @@ -0,0 +1,27 @@
> +/* { dg-do compile } */
> +/* { dg-options "-O3 -march=armv8.3-a+sve -msve-vector-bits=scalable -fcx-
> limited-range" } */
> +
> +/* The SVE complex multiply also does the first rotation as an explicit
> +   multiply, so it does not need -fno-signed-zeros either.  */
> +
> +void
> +mulf (_Complex float *__restrict c, _Complex float *__restrict a,
> +      _Complex float *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * b[i];
> +}
> +
> +void
> +muld (_Complex double *__restrict c, _Complex double *__restrict a,
> +      _Complex double *__restrict b, int n)
> +{
> +  for (int i = 0; i < n; i++)
> +    c[i] = a[i] * b[i];
> +}
> +
> +/* { dg-final { scan-assembler {trn1\tz[0-9]+\.s, z[0-9]+\.s, z[0-9]+\.s} } } */
> +/* { dg-final { scan-assembler {trn1\tz[0-9]+\.d, z[0-9]+\.d, z[0-9]+\.d} } } */
> +/* { dg-final { scan-assembler {fcmla\tz[0-9]+\.s, p[0-9]+/m, z[0-9]+\.s, z[0-
> 9]+\.s, #90} } } */
> +/* { dg-final { scan-assembler {fcmla\tz[0-9]+\.d, p[0-9]+/m, z[0-9]+\.d, z[0-
> 9]+\.d, #90} } } */
> +/* { dg-final { scan-assembler-not {fcmla\tz[0-9]+\.[sd], p[0-9]+/m, z[0-
> 9]+\.[sd], z[0-9]+\.[sd], #0} } } */
> --
> 2.50.1 (Apple Git-155)
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.