[Bug tree-optimization/126946] New: [14/15/16/17 Regression] Phiopt lengthening dependency chain in loop for MIN/MAX operations

"ktkachov at gcc dot gnu.org via Gcc-bugs" <[email protected]>
Newsgroups gmane.comp.gcc.bugs
Message-ID <[email protected]/bugzilla/>
https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126946

            Bug ID: 126946
           Summary: [14/15/16/17 Regression] Phiopt lengthening dependency
                    chain in loop for MIN/MAX operations
           Product: gcc
           Version: 17.0
            Status: UNCONFIRMED
          Keywords: missed-optimization
          Severity: normal
          Priority: P3
         Component: tree-optimization
          Assignee: unassigned at gcc dot gnu.org
          Reporter: ktkachov at gcc dot gnu.org
                CC: rguenth at gcc dot gnu.org
  Target Milestone: ---
            Target: aarch64

We got an internal customer reporting a large performance regression with GCC
14 on a tight min/max loop on aarch64.

The crux of the matter is:
float
loop_max (long n, const float *__restrict x)
{
  float m = 0.0f;
  for (long i = 0; i < n; i++)
    {
      float a = __builtin_fabsf (x[i]);
      m = m < a ? a : m;
    }
  return m;
}

GCC at -O2 GCC 13 on aarch64 produced:
loop_max:
        movi    v0.2s, #0
        cmp     x0, 0
        ble     .L1
        add     x0, x1, x0, lsl 2
.L4:
        ldr     s1, [x1]
        fabs    s1, s1
        fcmpe   s0, s1
        bmi     .L6
.L3:
        add     x1, x1, 4
        cmp     x0, x1
        bne     .L4
.L1:
        ret
.L6:
        fmov    s0, s1
        b       .L3

whereas GCC 14 produces a conditional select inside the loop:
loop_max:
        cmp     x0, 0
        ble     .L4
        movi    v0.2s, #0
        mov     w2, 0
.L3:
        ldr     s31, [x1, x2, lsl 2]
        add     x2, x2, 1
        fabs    s31, s31
        fcmpe   s31, s0
        fcsel   s0, s31, s0, gt
        cmp     x0, x2
        bne     .L3
        ret
.L4:
        movi    v0.2s, #0
        ret

This costs a 4x drop in the particular workload.
The issue is g:9f8f37f5490076 from PR88540 produces the COND_EXPR when
MIN/MAX_EXPR formation is blocked rather than leaving the decision to RTL,
which has a more fine-grained logic for target specifics.

Maybe we can track whether the potential COND_EXPR operands would come from a
loop recurrence and avoid the COND_EXPR formation then?
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.