Re: [PATCH 1/2] x86: optimize XCHG to MOV for same-register forms
Jan Beulich via Valgrind-developers <valgrind-developers-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org>
| Newsgroups | gmane.comp.debugging.valgrind.devel,gmane.comp.gnu.binutils |
|---|---|
| Message-ID | <[email protected]> |
On 07.07.2026 23:18, H.J. Lu wrote: > On Tue, Jul 7, 2026 at 10:35 PM Jan Beulich <[email protected]> wrote: >> On 07.07.2026 16:21, Sam James wrote: >>> Just what you keep saying? >>> >>> "Assembler optimization is intended for hand-written code, not to be >>> used on compiler output." which then gives us licence to dispose with >>> bugs reports like this, as opposed to the ambiguity right now. >> >> I fear I wouldn't be happy with making such a statement in doc. For one, >> "intended" is weak enough that people may still think the options are >> worthwhile to use on compiler output. Plus there's the issue with inline >> assembly, which imo can plausibly be subject to optimization. Yet at >> this time we have no way to have optimization "kick in" only on those >> portions. >> >> I could perhaps live with a yet weaker version of what you suggest: >> >> "Assembler optimization is intended primarily for hand-written code. If > > I disagree. I added -O to assembler for > > On x86, some instructions have alternate shorter encodings: > > 1. When the upper 32 bits of destination registers of > > andq $imm31, %r64 > testq $imm31, %r64 > xorq %r64, %r64 > subq %r64, %r64 > > known to be zero, we can encode them without the REX_W bit: > > andl $imm31, %r32 > testl $imm31, %r32 > xorl %r32, %r32 > subl %r32, %r32 > > This optimization is enabled with -O, -O2 and -Os. > 2. Since 0xb0 mov with 32-bit destination registers zero-extends 32-bit > immediate to 64-bit destination register, we can use it to encode 64-bit > mov with 32-bit immediates. This optimization is enabled with -O, -O2 > and -Os. > 3. Since the upper bits of destination registers of VEX128 and EVEX128 > instructions are extended to zero, if all bits of destination registers > of AVX256 or AVX512 instructions are zero, we can use VEX128 or EVEX128 > encoding to encode AVX256 or AVX512 instructions. When 2 source > registers are identical, AVX256 and AVX512 andn and xor instructions: > > VOP %reg, %reg, %dest_reg > > can be encoded with > > VOP128 %reg, %reg, %dest_reg > > This optimization is enabled with -O2 and -Os. > 4. 16-bit, 32-bit and 64-bit register tests with immediate may be > encoded as 8-bit register test with immediate. This optimization is > enabled with -Os. > > These optimizations were intended for compiler generated assembly > codes. Why would that be? I.e. why would the compiler emit sub-optimal code, for the assembler to tidy after it? Jan _______________________________________________ Valgrind-developers mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/valgrind-developers