Re: 回复:[RFC] RISC-V: Raise noce if-c onversion cost limit

Jeffrey Law <[email protected]>
Newsgroups gmane.comp.gcc.patches
Message-ID <[email protected]>

On 8/7/2026 4:44 PM, wangjue wrote:
> Thank you for the suggestion. You're right that raising BRANCH_COST sufficiently makes this case if-convert. My concern is that changing it globally also affects other expansion and RTL profitability decisions, rather than just this particular noce transformation.
>
> Current threshold
> I compared the current TARGET_MAX_NOCE_IFCVT_SEQ_COST implementations:
> /* Generic default.  */
> return BRANCH_COST (true, predictable_p) * COSTS_N_INSNS (3);
> /* i386.  */
> return BRANCH_COST (true, predictable_p) * COSTS_N_INSNS (2);
> /* RISC-V.  */
> return COSTS_N_INSNS (BRANCH_COST (true, predictable_p));
>
> With BRANCH_COST == 4, these correspond to thresholds of 48, 32 and 16. This made the current RISC-V threshold appear relatively conservative, particularly when Zicond is available.
Yea.  It's also worth noting that BRANCH_COST gets used in two very 
different places.

First way up in the generic/gimple/tree optimizers it's used to drive 
decisions for if combination, conversion and related transformations.  
  Then it's used down in the RTL space to drive decisions around if 
conversion.    I've had a nagging suspicion that we need to break it 
into two different values.  On our vt1 and vt2 designs we generally 
found that very aggressive branch costing early tended to hurt 
performance while very aggressive branch costing in the RTL tended to 
improve performance.  But we had enough other things to chase down that 
were more important so we never really did the analysis to confirm what 
we were seeing in limited testing and draw meaningful conclusions on 
where to go.


>
> Increasing it globally may therefore generate more branchless code in unrelated places, including cases where the branch is actually predicted well by the hardware.
> Instead, I was considering restricting the adjustment to the RISC-V noce threshold:
See above.  Yes, increasing can have unintended consequences.  That 
would be a secondary benefit of splitting BRANCH_COST into independent 
knobs for gimple vs RTL -- fewer unintended consequences as you work 
through issues with adjusting one value for your uarch.

Note that it may still be advantageous to convert things that are well 
predicted as if-converting will create larger blocks for the compiler's 
various optimizers to chew on.  It's a behavior I've always been curious 
to investigate, but it hasn't bubbled up high enough yet on the priority 
list.  Essentially one behavior we've seen is that LLVM consistently has 
fewer conditional branches (from a dynamic standpoint), but consistently 
higher branch mispredicts than GCC compiled code.  We investigated this 
a few years back (before focusing on vector) -- in x264 LLVM 
aggressively if-converted the clamp idioms with essentially zero 
benefit, but missed converting a tough to predict branch elsewhere that 
was far more important.   I've meant to go back and see if if-converting 
those clamps in GCC would in turn allow additional downstream 
optimization opportunities (GCC was if-converting the tougher, more 
valuable case).
>> It might also help if you passed along a testcase. I've found many of
>> the failure to if-convert problems are due to inefficiencies in the
>> sequences we generate; we can often get if-conversion to fire by
>> generating better generic sequences and fixing costing goofs elsewhere.
> The case is in spec_qsort from 505.mcf_r. It occurs after med3 and its comparator have both been inlined into spec_qsort. The relevant call is:
> pn = med3 (pn - 2 * d, pn - d, pn, cmp);
> After inlining, the same ID condition controls three pointer selections. RTL if-conversion initially costs the candidate sequence at 48 cost units, equivalent to 12 instructions. A maximum cost of 47 does not trigger the conversion, while 48 does.
> After later RTL optimization, the final sequence is reduced to:
> 1 slt
> 6 czero
> 3 add
> This is 10 instructions, or 40 cost units. It appears that redundant condition calculations included in the initial candidate are removed by later passes.
> The two preceding abs_cost comparisons remain branches. Only the final ID tie-breaker and the three dependent pointer selections are converted.
Yea, that could well be the one I was looking at fairly recently.  At 
the point where you're considering if-conversion, what values are you 
selecting across?

The one I remember from mcf was a select across 1, -1.  That was a 4 
instruction sequence (slt+li+czero+addi), so cost 16 and thus a 
BRANCH_COST of 4 to trigger.  The sequence could be improved to a 3 
instruction sequence (slt+slli+addi) and thus would start to trigger at 
a BRANCH_COST of 3.  The testcase ultimately reduced down to something 
like this:

     int foo(int x, int y) { return (y < x) ? 1 : -1; }

The solution was to improve selection across 2^n and 0, once that was 
done generic code would subtract 1 from that result giving us an 
efficient select across 2^n-1 and -1 which obviously includes a select 
across 1, -1.


Andrea pointed out during that work that defining a spaceship pattern 
would likely help in meaningful way and is almost certainly correct.  
That might be another avenue to explore and would likely help numerous 
things without doing costing adjustments.

And yes, the if-converted sequence initially tends to have unnecessary 
copies, extensions, etc that make some sequences look unprofitable to 
use, but later optimizations would clean things up and they would have 
been profitable.  I've fixed several bugs in that area recently, though 
the most serious remains and I haven't found a good solution.  I touched 
on the problems with this commit:

So what I'd suggest would be to take the case you've found in MCF, 
reduce it down to a small testcase if possible and we'll start working 
with that.  We'd probably want that for the final patch anyway as a 
testcase, so even if we go with your approach on costing, this work 
would still be desirable.

> commit 3ab2199392445f2deb3855dcfe68ecf0fed8466d
> Author: Jeff Law <[email protected]>
> Date:   Thu Jul 16 06:37:53 2026 -0600
>
>     [RISC-V] Improve costing for if-converted sequences
>
>     So I'm not happy with this entire routine, but after trying to fix 
> things
>     right, I'm back to adjusting the cost routine instead.
>
>     The fundamental problem as I see it is the cstore patterns on 
> RISC-V are
>     suboptimal.
[ ... ]

The full commit message describes the issues.  I suspect finding a good 
solution to the problems in that full commit message would significantly 
help the RTL if-conversion phase make better choices.

Jeff
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.