Re: [PATCH 1/2] genopinit: Distribute generated code across multiple files
Kyrylo Tkachov <[email protected]> Wed, 5 Aug 2026 14:38:06 +0000
| Newsgroups | gmane.comp.gcc.patches |
|---|---|
| Message-ID | <[email protected]> |
> On 5 Aug 2026, at 10:11, Tamar Christina <[email protected]> wrote: > >> -----Original Message----- >> From: Robin Dapp <[email protected]> >> Sent: 04 August 2026 20:21 >> To: Tamar Christina <[email protected]>; Jeffrey Law >> <[email protected]>; [email protected]; gcc- >> [email protected] >> Cc: [email protected]; Robin Dapp <[email protected]> >> Subject: Re: [PATCH 1/2] genopinit: Distribute generated code across multiple >> files >> >>> But in my opinion we don't want to keep adding split code to every >>> gen* file individually >>> but concentrate that code in gensupport so ever gen* can use it. >> >> The individual splits are there because I started out with one >> (genemit), then later realized genrecog is large as well. It was >> certainly not the initial intention to split everything individually. >> But I agree that by now we have reached a tipping point. >> >> BTW I also touched opinit before, splitting one function into several, >> because it would trigger pathological compiler behavior. >> So we already perform some kind of splitting. Are we sure we're not >> actually seeing pathological behavior again that could be helped by >> something else than file-level splitting? > > geninit is also one that I think can benefit from a different way of generating > the values. > > If you look at the large arrays it generates for like init_all_optabs > > These are huge, and also result in at runtime just a series of branches and calls. > > i.e. > > static void > init_optabs_01 (struct target_optabs *optabs) > { > bool *ena = optabs->pat_enable; > ena[1000] = HAVE_udot_prodv2siv8qi; > ena[1001] = HAVE_udot_prodv4siv16qi; > ena[1002] = HAVE_udot_prodvnx8hivnx16qi; > ena[1003] = HAVE_udot_prodvnx4sivnx16qi; > ena[1004] = HAVE_udot_prodvnx4sivnx8hi; > ena[1005] = HAVE_udot_prodvnx2divnx8hi; > ena[1006] = HAVE_usdot_prodv2siv8qi; > ena[1007] = HAVE_usdot_prodv4siv16qi; > ena[1008] = HAVE_usdot_prodvnx4sivnx16qi; > ena[1009] = HAVE_while_ultsivnx16bi; > ena[1010] = HAVE_while_ultdivnx16bi; > ena[1011] = HAVE_while_ultsivnx8bi; > ena[1012] = HAVE_while_ultdivnx8bi; > ena[1013] = HAVE_while_ultsivnx4bi; > ena[1014] = HAVE_while_ultdivnx4bi; > ena[1015] = HAVE_while_ultsivnx2bi; > ena[1016] = HAVE_while_ultdivnx2bi; > ena[1017] = HAVE_addsi3; > ena[1018] = HAVE_adddi3; > ena[1019] = HAVE_addti3; > ena[1020] = HAVE_addhf3; > ena[1021] = HAVE_addsf3; > ena[1022] = HAVE_adddf3; > ena[1023] = HAVE_addv8qi3; > ena[1024] = HAVE_addv4hi3; > ena[1025] = HAVE_addv2si3; > ena[1026] = HAVE_addv16qi3; > ena[1027] = HAVE_addv8hi3; > ena[1028] = HAVE_addv4si3; > ena[1029] = HAVE_addv2di3; > ena[1030] = HAVE_addvnx16qi3; > ena[1031] = HAVE_addvnx8hi3; > ena[1032] = HAVE_addvnx4si3; > ena[1033] = HAVE_addvnx2di3; > ena[1034] = HAVE_addvnx2qi3; > ena[1035] = HAVE_addvnx4qi3; > ena[1036] = HAVE_addvnx2hi3; > ena[1037] = HAVE_addvnx8qi3; > ena[1038] = HAVE_addvnx4hi3; > ena[1039] = HAVE_addvnx2si3; > ena[1040] = HAVE_addv4hf3; > ena[1041] = HAVE_addv2sf3; > ena[1042] = HAVE_addv8hf3; > ena[1043] = HAVE_addv4sf3; > ena[1044] = HAVE_addv2df3; > > etc. > > However if you look at insn-flags.h you'll see that > These flags mostly have the same definitions > > For instance > >> grep "HAVE_while_ult" ../../build-aarch64-none-elf/obj/gcc2/gcc/insn-flags.h > > #define HAVE_while_ultsivnx16bi (TARGET_SVE) > #define HAVE_while_ultdivnx16bi (TARGET_SVE) > #define HAVE_while_ultsivnx8bi (TARGET_SVE) > #define HAVE_while_ultdivnx8bi (TARGET_SVE) > #define HAVE_while_ultsivnx4bi (TARGET_SVE) > #define HAVE_while_ultdivnx4bi (TARGET_SVE) > #define HAVE_while_ultsivnx2bi (TARGET_SVE) > #define HAVE_while_ultdivnx2bi (TARGET_SVE) > #define HAVE_while_ultsivnx16bi_ptest (TARGET_SVE) > #define HAVE_while_ultdivnx16bi_ptest (TARGET_SVE) > #define HAVE_while_ultsivnx8bi_ptest (TARGET_SVE) > #define HAVE_while_ultdivnx8bi_ptest (TARGET_SVE) > #define HAVE_while_ultsivnx4bi_ptest (TARGET_SVE) > #define HAVE_while_ultdivnx4bi_ptest (TARGET_SVE) > #define HAVE_while_ultsivnx2bi_ptest (TARGET_SVE) > #define HAVE_while_ultdivnx2bi_ptest (TARGET_SVE) > > So we're spending quite a bit of time reading and compiling these files to just set > the same flag over and over again. > > If we were to group the generated code by the defines instead, we wouldn't > need giant array assignments. > > Init_all_optabs would just become: > > bool *ena = optabs->pat_enable > > memset (ena, 0, sizeof optabs->pat_enable); > > if (TARGET_SVE) > memset (ena + offset, 1, count); > > which I think would be both a runtime and a compile time win because > quite a few insn have either no condition, or are just TARGET_SVE or > TARGET_SVE2 etc. And I assume this is common across many other > targets. > > Though I have only experimented with a small version years ago. > > Other files have similar things that interesting to evaluate wrt to the > scale the generators are at now. > This turned out to not be too hard after all. I’ve tried to factor common code out to gensupport.cc and made this change genopinit. Sent out as a separate series. Thanks, Kyrill > Thanks, > Tamar > >> >>> Lastly I also don't think splitting on an iterative way is going to >>> give you the best compile >>> time increase. Because that tends to bias the output. >>> >>> In genmatch I use ftell to tell how big the files is so far an pick the smallest >> file. This allows >>> me to balance out the compile time over all files. I believe Robin eventually >> respun his >>> gen<something> match to do this a swell as that allows you to amortize the >> compile time >>> over cores. >> >> Yeah, I added your suggestion to the initial genemit split. And at >> least the "which file to write to next"/ftell functionality is not local >> to genemit but already in gensupport. The function is choose_output and >> Kyryll's patch even uses it already so it shouldn't do iterative >> splitting if I'm not mistaken. >> >> Each gen* still keep track of its own files, though, and that part could >> be unified still. I don't think it's terribly difficult but will >> require some plumbing. >> >> >> -- >> Regards >> Robin