Re: [PATCH 1/2] genopinit: Distribute generated code across multiple files

Kyrylo Tkachov <[email protected]>
Newsgroups gmane.comp.gcc.patches
Message-ID <[email protected]>

> On 5 Aug 2026, at 16:37, Kyrylo Tkachov <[email protected]> wrote:
> 
> 
> 
>> On 5 Aug 2026, at 10:11, Tamar Christina <[email protected]> wrote:
>> 
>>> -----Original Message-----
>>> From: Robin Dapp <[email protected]>
>>> Sent: 04 August 2026 20:21
>>> To: Tamar Christina <[email protected]>; Jeffrey Law
>>> <[email protected]>; [email protected]; gcc-
>>> [email protected]
>>> Cc: [email protected]; Robin Dapp <[email protected]>
>>> Subject: Re: [PATCH 1/2] genopinit: Distribute generated code across multiple
>>> files
>>> 
>>>> But in my opinion we don't want to keep adding split code to every
>>>> gen* file individually
>>>> but concentrate that code in gensupport so ever gen* can use it.
>>> 
>>> The individual splits are there because I started out with one
>>> (genemit), then later realized genrecog is large as well.  It was
>>> certainly not the initial intention to split everything individually.
>>> But I agree that by now we have reached a tipping point.
>>> 
>>> BTW I also touched opinit before, splitting one function into several,
>>> because it would trigger pathological compiler behavior.
>>> So we already perform some kind of splitting.  Are we sure we're not
>>> actually seeing pathological behavior again that could be helped by
>>> something else than file-level splitting?
>> 
>> geninit is also one that I think can benefit from a different way of generating
>> the values.
>> 
>> If you look at the large arrays it generates for like init_all_optabs
>> 
>> These are huge, and also result in at runtime just a series of branches and calls.
>> 
>> i.e.
>> 
>> static void
>> init_optabs_01 (struct target_optabs *optabs)
>> {
>> bool *ena = optabs->pat_enable;
>> ena[1000] = HAVE_udot_prodv2siv8qi;
>> ena[1001] = HAVE_udot_prodv4siv16qi;
>> ena[1002] = HAVE_udot_prodvnx8hivnx16qi;
>> ena[1003] = HAVE_udot_prodvnx4sivnx16qi;
>> ena[1004] = HAVE_udot_prodvnx4sivnx8hi;
>> ena[1005] = HAVE_udot_prodvnx2divnx8hi;
>> ena[1006] = HAVE_usdot_prodv2siv8qi;
>> ena[1007] = HAVE_usdot_prodv4siv16qi;
>> ena[1008] = HAVE_usdot_prodvnx4sivnx16qi;
>> ena[1009] = HAVE_while_ultsivnx16bi;
>> ena[1010] = HAVE_while_ultdivnx16bi;
>> ena[1011] = HAVE_while_ultsivnx8bi;
>> ena[1012] = HAVE_while_ultdivnx8bi;
>> ena[1013] = HAVE_while_ultsivnx4bi;
>> ena[1014] = HAVE_while_ultdivnx4bi;
>> ena[1015] = HAVE_while_ultsivnx2bi;
>> ena[1016] = HAVE_while_ultdivnx2bi;
>> ena[1017] = HAVE_addsi3;
>> ena[1018] = HAVE_adddi3;
>> ena[1019] = HAVE_addti3;
>> ena[1020] = HAVE_addhf3;
>> ena[1021] = HAVE_addsf3;
>> ena[1022] = HAVE_adddf3;
>> ena[1023] = HAVE_addv8qi3;
>> ena[1024] = HAVE_addv4hi3;
>> ena[1025] = HAVE_addv2si3;
>> ena[1026] = HAVE_addv16qi3;
>> ena[1027] = HAVE_addv8hi3;
>> ena[1028] = HAVE_addv4si3;
>> ena[1029] = HAVE_addv2di3;
>> ena[1030] = HAVE_addvnx16qi3;
>> ena[1031] = HAVE_addvnx8hi3;
>> ena[1032] = HAVE_addvnx4si3;
>> ena[1033] = HAVE_addvnx2di3;
>> ena[1034] = HAVE_addvnx2qi3;
>> ena[1035] = HAVE_addvnx4qi3;
>> ena[1036] = HAVE_addvnx2hi3;
>> ena[1037] = HAVE_addvnx8qi3;
>> ena[1038] = HAVE_addvnx4hi3;
>> ena[1039] = HAVE_addvnx2si3;
>> ena[1040] = HAVE_addv4hf3;
>> ena[1041] = HAVE_addv2sf3;
>> ena[1042] = HAVE_addv8hf3;
>> ena[1043] = HAVE_addv4sf3;
>> ena[1044] = HAVE_addv2df3;
>> 
>> etc.
>> 
>> However if you look at insn-flags.h you'll see that
>> These flags mostly have the same definitions
>> 
>> For instance 
>> 
>>> grep "HAVE_while_ult" ../../build-aarch64-none-elf/obj/gcc2/gcc/insn-flags.h
>> 
>> #define HAVE_while_ultsivnx16bi (TARGET_SVE)
>> #define HAVE_while_ultdivnx16bi (TARGET_SVE)
>> #define HAVE_while_ultsivnx8bi (TARGET_SVE)
>> #define HAVE_while_ultdivnx8bi (TARGET_SVE)
>> #define HAVE_while_ultsivnx4bi (TARGET_SVE)
>> #define HAVE_while_ultdivnx4bi (TARGET_SVE)
>> #define HAVE_while_ultsivnx2bi (TARGET_SVE)
>> #define HAVE_while_ultdivnx2bi (TARGET_SVE)
>> #define HAVE_while_ultsivnx16bi_ptest (TARGET_SVE)
>> #define HAVE_while_ultdivnx16bi_ptest (TARGET_SVE)
>> #define HAVE_while_ultsivnx8bi_ptest (TARGET_SVE)
>> #define HAVE_while_ultdivnx8bi_ptest (TARGET_SVE)
>> #define HAVE_while_ultsivnx4bi_ptest (TARGET_SVE)
>> #define HAVE_while_ultdivnx4bi_ptest (TARGET_SVE)
>> #define HAVE_while_ultsivnx2bi_ptest (TARGET_SVE)
>> #define HAVE_while_ultdivnx2bi_ptest (TARGET_SVE)
>> 
>> So we're spending quite a bit of time reading and compiling these files to just set
>> the same flag over and over again.
>> 
>> If we were to group the generated code by the defines instead, we wouldn't
>> need giant array assignments.
>> 
>> Init_all_optabs would just become:
>> 
>> bool *ena = optabs->pat_enable
>> 
>> memset (ena, 0, sizeof optabs->pat_enable);
>> 
>> if (TARGET_SVE)
>>  memset (ena + offset, 1, count);
>> 
>> which I think would be both a runtime and a compile time win because
>> quite a few insn have either no condition, or are just TARGET_SVE or
>> TARGET_SVE2 etc.  And I assume this is common across many other
>> targets.
>> 
>> Though I have only experimented with a small version years ago.
>> 
>> Other files have similar things that interesting to evaluate wrt to the
>> scale the generators are at now.
>> 
> 
> This turned out to not be too hard after all. I’ve tried to factor common code out to gensupport.cc and made this change genopinit. Sent out as a separate series.

I ended up exploring the rest of the gen* generators to find areas where their output could be improved, before we tackle the splitting. I’ve sent out a separate patch series that I think tackles the RTL side of the generators more comprehensively.
Thanks,
Kyrill


> Thanks,
> Kyrill
> 
>> Thanks,
>> Tamar
>> 
>>> 
>>>> Lastly I also don't think splitting on an iterative way is going to
>>>> give you the best compile
>>>> time increase.  Because that tends to bias the output.
>>>> 
>>>> In genmatch I use ftell to tell how big the files is so far an pick the smallest
>>> file. This allows
>>>> me to balance out the compile time over all files. I believe Robin eventually
>>> respun his
>>>> gen<something> match to do this a swell as that allows you to amortize the
>>> compile time
>>>> over cores.
>>> 
>>> Yeah, I added your suggestion to the initial genemit split.  And at
>>> least the "which file to write to next"/ftell functionality is not local
>>> to genemit but already in gensupport.  The function is choose_output and
>>> Kyryll's patch even uses it already so it shouldn't do iterative
>>> splitting if I'm not mistaken.
>>> 
>>> Each gen* still keep track of its own files, though, and that part could
>>> be unified still.  I don't think it's terribly difficult but will
>>> require some plumbing.
>>> 
>>> 
>>> --
>>> Regards
>>> Robin
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.