Re: [PATCH] x86: Double unaligned move cost for short data

"H.J. Lu" <[email protected]>
Newsgroups gmane.comp.gcc.patches
Message-ID <CAMe9rOoaaPVfdBH4w0So32LHqH2tYFbuQwzzFbjkxcZLEHNgsw@mail.gmail.com>
On Thu, Aug 13, 2026 at 3:59 PM Richard Biener
<[email protected]> wrote:
>
> On Wed, Aug 12, 2026 at 4:25 PM H.J. Lu <[email protected]> wrote:
> >
> > Double unaligned load and store cost if the vector mode size is less
> > than 8 bytes and the number of vector element is less than 4 when SSE
> > is enabled to avoid using vector instructions on unaligned short data
> > with 2 vector elements so that for
> >
> > extern char *var1;
> > extern int var2;
> >
> > void
> > func (void)
> > {
> >   var2 = var1[1] + var1[0];
> > }
> >
> > we generate
> >
> > movq var1(%rip), %rdx
> > movsbl 1(%rdx), %eax
> > movsbl (%rdx), %edx
> > addl %edx, %eax
> > movl %eax, var2(%rip)
> >
> > instead of
> >
> > movq var1(%rip), %rax
> > pxor %xmm1, %xmm1
> > pinsrw $0, (%rax), %xmm0
> > pcmpgtb %xmm0, %xmm1
> > punpcklbw %xmm1, %xmm0
> > movdqa %xmm0, %xmm1
> > psraw $15, %xmm1
> > punpcklwd %xmm1, %xmm0
> > movd %xmm0, %edx
> > pshufd $0xe5, %xmm0, %xmm2
> > movd %xmm2, %eax
> > addl %edx, %eax
> > movl %eax, var2(%rip)
> >
> > with -O2 -march=x86-64.
> >
> > Compile PR 125100 tests with -mno-sse since unaligned V2QImode load is no
> > longer generated when SSE is enabled.
>
> Instead of just doubling the load/store cost can we try to more accurately
> model the cost of loading of 1, 2 or 4 byte vectors to SSE registers?
> For example with SSE4 we get
>
>         movq    var1(%rip), %rax
>         pinsrw  $0, (%rax), %xmm0
>         pmovsxbd        %xmm0, %xmm0
>         movd    %xmm0, %edx
>         pextrd  $1, %xmm0, %eax
>         addl    %edx, %eax
>         movl    %eax, var2(%rip)
>         ret

Is this really better than

 movq var1(%rip), %rdx
movswl 2(%rdx), %eax
movswl (%rdx), %edx
addl %edx, %eax
movl %eax, var2(%rip)

> and 4 bytes with SSE can be loaded via
>
>         movq    var1(%rip), %rax
>         movd    (%rax), %xmm0
>
> (why the indirection through rax?).  So it's only 2 byte vectors that

var1 is a pointer.   2 shorts have the same issue.

> are problematic
> when SSE4 is not available?
>
> In general we're missing to model the fact that modern x86 uarchs can happily
> execute two streams of same scalar integer operations in parallel and using
> vector intructions for them a) increases code size due to larger insn encodings,
> b) for some uarchs has GPR <-> XMM move penalties.  So I was thinking
> that while the vectorizer accepts scalar-cost == vector-cost as being profitable
> to vectorize, at least for integer code with only two lanes, there should be a
> larger profitability gap.  We can see to adjust the finish_cost hook in case
> of BB vectorization to implement such heuristic, like requiring 66% of the
> scalar cost?  Benefits/costs of the surrounding code are of course not
> and difficult to evaluate (like register pressure on both sides, or EU
> occupancy).
>
> So I think you want to be just testing the vector mode size, not element size,
> and size 4 looks OK to me.  You also want to handle pinsrw availability and
> not double cost (load cost is target dependent) but instead cost a GPR load
> plus the relevant vector instruction cost.
>
> Richard.
>
> > gcc/
> >
> > PR target/126802
> > * config/i386/i386.cc (sse_adjust_unaligned_cost): New.
> > (ix86_default_vector_cost): Call sse_adjust_unaligned_cost for
> > unaligned load and store to adjust unaligned move cost.
> >
> > testsuite/
> >
> > PR target/126802
> > * gcc.target/i386/pr125100-1.c: Add -mno-sse.
> > * gcc.target/i386/pr125100-2.c: Likewise.
> > * gcc.target/i386/pr125100-3.c: Likewise.
> > * gcc.target/i386/pr126802-1a.c: New test.
> > * gcc.target/i386/pr126802-1b.c: Likewise.
> > * gcc.target/i386/pr126802-2a.c: Likewise.
> > * gcc.target/i386/pr126802-2b.c: Likewise.
> > * gcc.target/i386/pr126802-3a.c: Likewise.
> > * gcc.target/i386/pr126802-3b.c: Likewise.
> > * gcc.target/i386/pr126802-4a.c: Likewise.
> > * gcc.target/i386/pr126802-4b.c: Likewise.
> >
> > Tested on Linux/x86-64 without any regressions.
> >
> > --
> > H.J.



-- 
H.J.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.