Re: [PATCH 2/2] x86: fold F16C templates with their AVX512VL counterparts

Jan Beulich <[email protected]> Tue, 21 Jul 2026 08:37:42 +0200
Newsgroups gmane.comp.gnu.binutils
Message-ID <[email protected]>
On 21.07.2026 08:18, Jiang, Haochen wrote:
>> From: Jan Beulich <[email protected]>
>> Sent: Tuesday, July 21, 2026 2:07 PM
>>
>> On 21.07.2026 05:16, Jiang, Haochen wrote:
>>>> From: Jan Beulich <[email protected]>
>>>> Sent: Friday, July 17, 2026 9:50 PM
>>>>
>>>> This actually points out another issue with FMA and F16C: Code created
>>>> prior to the introduction of AVX512 might have used ".arch .no{fma,f16c}"
>>>> to have uses of the respective insns properly rejected. That protection
>>>> suddenly ceased to work properly when AVX512 support was added to the
>>>> assembler: Those insns would now assemble to their AVX512VL EVEX forms.
>>>> The least bad option to deal with this looks to be to reject the AVX512VL
>>>> forms also when FMA/F16C are explicitly disabled while AVX512VL isn't
>>>> explicitly enabled. Are there any better suggestions?
>>>
>>> When enabling AVX512VL, those EVEX forms should be enabled since it
>>> is an "or" claimed in SDM. It is a bug but just not reported.
>>
>> Here pretty prominently, but also elsewhere below I can't resolve ambiguities
>> in your reply. On one hand I'm getting the impression that you want us to
>> follow what the SDM mandates, while otoh ...
> 
> To be clear, I mean it is not SDM bug. .nofma disabling EVEX FMA is buggy
> to me. 
> 
>>
>>> The behavior
>>> is the same in GCC, where -mno-fma will not control EVEX version. The
>>> actual flag to totally disable FMA is -ffp-contract=off.
>>>
>>> So, I believe we should just enable those insts under this scenario. The
>>> thing we might need to do is to add a mandatory EVEX prefix here, but
>>> maybe tricky.
>>>
>>> Ref from GCC repo in gcc/config/i386/sse.md, which actually applied
>>> to all FMA, but it is mentioned in scalar part:
>>>
>>> ;; The standard names for scalar FMA are only available with SSE math enabled.
>>> ;; CPUID bit AVX512F enables evex encoded scalar and 512-bit fma.  It doesn't
>>> ;; care about FMA bit, so we enable fma for TARGET_AVX512F even when TARGET_FMA
>>> ;; and TARGET_FMA4 are both false.
>>> ;; TODO: In theory AVX512F does not automatically imply FMA, and without FMA
>>> ;; one must force the EVEX encoding of the fma insns.  Ideally we'd improve
>>> ;; GAS to allow proper prefix selection.  However, for the moment all hardware
>>> ;; that supports AVX512F also supports FMA so we can ignore this for now.
>>
>> ... you quote this. If we followed this (and then also for F16C, which in gcc
>> doesn't have a similar comment but looks to be treated similarly).
>>
>> So - are you suggesting to keep dependencies as we have them, or are you
>> suggesting to introduce "artificial" dependencies of AVX512F (or AVX512VL)
>> on FMA and F16C (thus allowing simplification here and perhaps elsewhere)? Or
>> yet something else?
> 
> I suppose .nofma should only disable VEX version FMA, leaving EVEX version there
> (better with {evex} prefix). Similar for F16C.

Ah, you want to _demand_ presence of the prefix, for disambiguation. Let me
think about this; at the first glance it feels like this may be a little
intrusive, but maybe I'm wrong.

Jan