Re: [PATCH 1/2] x86: optimize XCHG to MOV for same-register forms

Jan Beulich via Valgrind-developers <valgrind-developers-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org>
Newsgroups gmane.comp.debugging.valgrind.devel,gmane.comp.gnu.binutils
Message-ID <[email protected]>
On 07.07.2026 23:18, H.J. Lu wrote:
> On Tue, Jul 7, 2026 at 10:35 PM Jan Beulich <[email protected]> wrote:
>> On 07.07.2026 16:21, Sam James wrote:
>>> Just what you keep saying?
>>>
>>> "Assembler optimization is intended for hand-written code, not to be
>>> used on compiler output." which then gives us licence to dispose with
>>> bugs reports like this, as opposed to the ambiguity right now.
>>
>> I fear I wouldn't be happy with making such a statement in doc. For one,
>> "intended" is weak enough that people may still think the options are
>> worthwhile to use on compiler output. Plus there's the issue with inline
>> assembly, which imo can plausibly be subject to optimization. Yet at
>> this time we have no way to have optimization "kick in" only on those
>> portions.
>>
>> I could perhaps live with a yet weaker version of what you suggest:
>>
>> "Assembler optimization is intended primarily for hand-written code.  If
> 
> I disagree.  I added -O to assembler for
> 
> On x86, some instructions have alternate shorter encodings:
> 
> 1. When the upper 32 bits of destination registers of
> 
> andq $imm31, %r64
> testq $imm31, %r64
> xorq %r64, %r64
> subq %r64, %r64
> 
> known to be zero, we can encode them without the REX_W bit:
> 
> andl $imm31, %r32
> testl $imm31, %r32
> xorl %r32, %r32
> subl %r32, %r32
> 
> This optimization is enabled with -O, -O2 and -Os.
> 2. Since 0xb0 mov with 32-bit destination registers zero-extends 32-bit
> immediate to 64-bit destination register, we can use it to encode 64-bit
> mov with 32-bit immediates.  This optimization is enabled with -O, -O2
> and -Os.
> 3. Since the upper bits of destination registers of VEX128 and EVEX128
> instructions are extended to zero, if all bits of destination registers
> of AVX256 or AVX512 instructions are zero, we can use VEX128 or EVEX128
> encoding to encode AVX256 or AVX512 instructions.  When 2 source
> registers are identical, AVX256 and AVX512 andn and xor instructions:
> 
> VOP %reg, %reg, %dest_reg
> 
> can be encoded with
> 
> VOP128 %reg, %reg, %dest_reg
> 
> This optimization is enabled with -O2 and -Os.
> 4. 16-bit, 32-bit and 64-bit register tests with immediate may be
> encoded as 8-bit register test with immediate.  This optimization is
> enabled with -Os.
> 
> These optimizations were intended for compiler generated assembly
> codes.

Why would that be? I.e. why would the compiler emit sub-optimal code,
for the assembler to tidy after it?

Jan


_______________________________________________
Valgrind-developers mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/valgrind-developers
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.