Re: [PATCH] x86: Double unaligned move cost for short data
Richard Biener <[email protected]>
| Newsgroups | gmane.comp.gcc.patches |
|---|---|
| Message-ID | <CAFiYyc2z4OBPWc_80gM2MKFEjuzLQcyHE2N=aNjwidTuyHnK-w@mail.gmail.com> |
On Fri, Aug 14, 2026 at 10:44 AM H.J. Lu <[email protected]> wrote: > > On Fri, Aug 14, 2026 at 4:13 PM Richard Biener > <[email protected]> wrote: > > > > On Fri, Aug 14, 2026 at 9:28 AM H.J. Lu <[email protected]> wrote: > > > > > > On Fri, Aug 14, 2026 at 2:49 PM Richard Biener > > > <[email protected]> wrote: > > > > > > > > On Fri, Aug 14, 2026 at 8:24 AM H.J. Lu <[email protected]> wrote: > > > > > > > > > > On Fri, Aug 14, 2026 at 2:21 PM Richard Biener > > > > > <[email protected]> wrote: > > > > > > > > > > > > On Thu, Aug 13, 2026 at 10:32 AM H.J. Lu <[email protected]> wrote: > > > > > > > > > > > > > > On Thu, Aug 13, 2026 at 3:59 PM Richard Biener > > > > > > > <[email protected]> wrote: > > > > > > > > > > > > > > > > On Wed, Aug 12, 2026 at 4:25 PM H.J. Lu <[email protected]> wrote: > > > > > > > > > > > > > > > > > > Double unaligned load and store cost if the vector mode size is less > > > > > > > > > than 8 bytes and the number of vector element is less than 4 when SSE > > > > > > > > > is enabled to avoid using vector instructions on unaligned short data > > > > > > > > > with 2 vector elements so that for > > > > > > > > > > > > > > > > > > extern char *var1; > > > > > > > > > extern int var2; > > > > > > > > > > > > > > > > > > void > > > > > > > > > func (void) > > > > > > > > > { > > > > > > > > > var2 = var1[1] + var1[0]; > > > > > > > > > } > > > > > > > > > > > > > > > > > > we generate > > > > > > > > > > > > > > > > > > movq var1(%rip), %rdx > > > > > > > > > movsbl 1(%rdx), %eax > > > > > > > > > movsbl (%rdx), %edx > > > > > > > > > addl %edx, %eax > > > > > > > > > movl %eax, var2(%rip) > > > > > > > > > > > > > > > > > > instead of > > > > > > > > > > > > > > > > > > movq var1(%rip), %rax > > > > > > > > > pxor %xmm1, %xmm1 > > > > > > > > > pinsrw $0, (%rax), %xmm0 > > > > > > > > > pcmpgtb %xmm0, %xmm1 > > > > > > > > > punpcklbw %xmm1, %xmm0 > > > > > > > > > movdqa %xmm0, %xmm1 > > > > > > > > > psraw $15, %xmm1 > > > > > > > > > punpcklwd %xmm1, %xmm0 > > > > > > > > > movd %xmm0, %edx > > > > > > > > > pshufd $0xe5, %xmm0, %xmm2 > > > > > > > > > movd %xmm2, %eax > > > > > > > > > addl %edx, %eax > > > > > > > > > movl %eax, var2(%rip) > > > > > > > > > > > > > > > > > > with -O2 -march=x86-64. > > > > > > > > > > > > > > > > > > Compile PR 125100 tests with -mno-sse since unaligned V2QImode load is no > > > > > > > > > longer generated when SSE is enabled. > > > > > > > > > > > > > > > > Instead of just doubling the load/store cost can we try to more accurately > > > > > > > > model the cost of loading of 1, 2 or 4 byte vectors to SSE registers? > > > > > > > > For example with SSE4 we get > > > > > > > > > > > > > > > > movq var1(%rip), %rax > > > > > > > > pinsrw $0, (%rax), %xmm0 > > > > > > > > pmovsxbd %xmm0, %xmm0 > > > > > > > > movd %xmm0, %edx > > > > > > > > pextrd $1, %xmm0, %eax > > > > > > > > addl %edx, %eax > > > > > > > > movl %eax, var2(%rip) > > > > > > > > ret > > > > > > > > > > > > > > Is this really better than > > > > > > > > > > > > > > movq var1(%rip), %rdx > > > > > > > movswl 2(%rdx), %eax > > > > > > > movswl (%rdx), %edx > > > > > > > addl %edx, %eax > > > > > > > movl %eax, var2(%rip) > > > > > > > > > > > > No, but it's better than the SSE2 version ;) If you just want to prevent > > > > > > these small vectorizations you can deny any scalar_cost < 30 vectorizations > > > > > > or so. I thought the point is to make the costing more precise - you > > > > > > correctly identified at least HImode -> XMM loads/stores to be not accurately > > > > > > costed. > > > > > > > > > > > > > > > > The problems are V2QImode and V2HImode. Others seem OK. > > > > > > > > But I see movd (%rax), %xmm0 for V2HImode, so that seems fine > > > > for loads and stores. > > > > > > > > movd (%rdi), %xmm0 > > > > paddw %xmm0, %xmm0 > > > > movd %xmm0, x(%rip) > > > > > > > > For V2QImode: > > > > > > > > pinsrw $0, (%rdi), %xmm0 > > > > paddb %xmm0, %xmm0 > > > > movd %xmm0, %eax > > > > movw %ax, x(%rip) > > > > > > > > so even SSE2 has pinsrw for the load but nothing for the store (SSE4 > > > > has pextrw there). > > > > > > > > So what am I missing? > > > > > > > > > > For > > > > > > extern short *var1; > > > extern int var2; > > > > > > void > > > func (void) > > > { > > > var2 = var1[1] + var1[0]; > > > } > > > > > > -O2 generates: > > > > > > movq var1(%rip), %rax > > > movd (%rax), %xmm0 > > > movdqa %xmm0, %xmm1 > > > psraw $15, %xmm1 > > > punpcklwd %xmm1, %xmm0 > > > movd %xmm0, %edx > > > pshufd $0xe5, %xmm0, %xmm2 > > > movd %xmm2, %eax > > > addl %edx, %eax > > > movl %eax, var2(%rip) > > > > > > With SSE4, we get > > > > > > movq var1(%rip), %rax > > > movd (%rax), %xmm0 > > > pmovsxwd %xmm0, %xmm0 > > > movd %xmm0, %edx > > > pextrd $1, %xmm0, %eax > > > addl %edx, %eax > > > movl %eax, var2(%rip) > > > > > > It isn't much better than > > > > > > movq var1(%rip), %rdx > > > movswl 2(%rdx), %eax > > > movswl (%rdx), %edx > > > addl %edx, %eax > > > movl %eax, var2(%rip) > > > > I'm not arguing about the profitability of the vectorization. I'm arguing > > of wheter the costing of the V2HImode load is wrong. That doesn't seem > > to be the case? > > > > Load is one part of the computation. The total cost of V2HImode > computation is too low. I think the total cost of the HImode calculation is too high ;) Or rather, load and store costs tend to dominate and for scalar we fail to realize we can issue the two loads in parallel and for stores we fail to realize they are irrelevant for the computation of latency. I'll note that x86 cost tables do not tell us the number of loads (which might depend on load width) that can issue in parallel, that is, we only have latency information, not throughput. We can possibly win most of the two-lane BB reduction cases by doing a DIV_CEIL (.., 2) on the total number of loads. For 2 scalar vs 1 vector load that would cancel out. That said, we make no attempt at identifying dependence chains so we add individual operation latencies as if we had a single linear dependence chain. That's the most fundamental issue with computing more accurate profitability. Richard. > > > -- > H.J.