[Bug target/126770] New: inefficient reduction of comparison sum

"rguenth at gcc dot gnu.org via Gcc-bugs" <[email protected]>
Newsgroups gmane.comp.gcc.bugs
Message-ID <[email protected]/bugzilla/>
https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126770

            Bug ID: 126770
           Summary: inefficient reduction of comparison sum
           Product: gcc
           Version: 17.0
            Status: UNCONFIRMED
          Keywords: missed-optimization
          Severity: normal
          Priority: P3
         Component: target
          Assignee: unassigned at gcc dot gnu.org
          Reporter: rguenth at gcc dot gnu.org
  Target Milestone: ---
            Target: x86_64-*-*

For the tail reduction in PR126028,

        bool test0 = (r2_0 < groupplcutoff2);
        bool test1 = (r2_1 < groupplcutoff2);
        hu += test0 + test1;

we generate awful code with -O3 -march=znver5 which can be seen with the
following reduced testcase (which needs -fno-vect-cost-model):

int foo (double g, double *r, int n)
{
  int hu = 0;
  bool test0 = r[0] < g;
  bool test1 = r[1] < g;
  hu += test0 + test1;
  return hu;
}

we generate

foo:
.LFB0:
        .cfi_startproc
        vmovupd (%rdi), %xmm1
        vmovddup        %xmm0, %xmm0
        movl    $1, %eax
        vcmppd  $1, %xmm0, %xmm1, %k1
        vpbroadcastq    %rax, %xmm0{%k1}{z}
        vpmovqd %xmm0, %xmm0
        vmovd   %xmm0, %eax
        vpextrd $1, %xmm0, %edx
        addl    %edx, %eax

which is in .optimized

  _17 = {g_8(D), g_8(D)};
  vect__1.5_16 = MEM <vector(2) double> [(double *)r_7(D)];
  mask_test0_9.6_18 = vect__1.5_16 < _17;
  vect_patt_11.7_19 = .VCOND_MASK (mask_test0_9.6_18, { 1, 1 }, { 0, 0 });
  vect_patt_12.8_20 = (vector(2) int) vect_patt_11.7_19;
  _22 = BIT_FIELD_REF <vect_patt_12.8_20, 32, 0>;
  _23 = BIT_FIELD_REF <vect_patt_12.8_20, 32, 32>;
  _24 = _22 + _23;

on a high-level the reduction of the mask as 1/0 could be matched as
population count.  But on a low-level STV could have figured that
doing the addl in SSE and the SSE->GPR move only once should be prefered
and similar the movl $1,%eax; vpbroadcastq %rax, %xmm0{%k1}{z}
sequence looks like there must be a more optimal form, not to name
the vpmovqd.

I realize most of this should probably happen on the GIMPLE level.

The first mistake is from the vectorizer and the bool pattern recognition of

  _1 = *r_7(D);
  test0_9 = _1 < g_8(D);
  _2 = MEM[(double *)r_7(D) + 8B];
  test1_10 = _2 < g_8(D);
  _3 = (int) test0_9;
  _4 = (int) test1_10;

where it choses a COND_EXPR with a unsigned long.  I'll open a separate PR
for that.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.