Re: [PATCH] LoongArch: improve 64-bit bitwise AND operation
"Ben Shi" <[email protected]>
| Newsgroups | gmane.comp.gcc.patches |
|---|---|
| Message-ID | <[email protected]> |
Shall we still use `define_insn_and_split` at first ? At least it does emit more efficient assembly. The solution you suggested involved changes to immediate materializing, which seems tricky, and I need more time to calculate a best form. ------------------ Original ------------------ From: "Xi Ruoyao" <[email protected]>; Date: Sun, Aug 2, 2026 05:11 AM To: "Jeffrey Law"<[email protected]>;"Ben Shi"<[email protected]>;"gcc-patches"<[email protected]>; Cc: "chenglulu"<[email protected]>;"郭杰"<[email protected]>; Subject: Re: [PATCH] LoongArch: improve 64-bit bitwise AND operation On Sun, 2026-08-02 at 04:49 +0800, Xi Ruoyao wrote: > On Sat, 2026-08-01 at 05:54 -0600, Jeffrey Law wrote: > > > > > > On 7/31/2026 3:47 AM, Ben Shi wrote: > > > It usually costs 4-5 instructions for 64-bit bitwise AND with a large > > > immediate. For some immediates the operation can be simplified to two > > > 'BSTRINS.D' instructions if the immediate satisfies: > > > 1. its top bit is not zero. > > > 2. has two sections of consecutive zero bits. > > > > > > gcc/ChangeLog: > > > * config/loongarch/loongarch.md: Add a new RTL expression > > > (define_insn_and_split "bstrins_bstrins_for_and_imm"). > > > > > > * config/loongarch/loongarch.cc: Add a helper function > > > 'loongarch_use_bstrins_bstrins_for_and' for the above RTL expression. > > > > > > * config/loongarch/loongarch-protos.h: Add prototype of function > > > 'loongarch_use_bstrins_bstrins_for_and'. > > > > > > * testsuite/gcc.target/loongarch/la64/and-large-immediate-opt-2.c: > > > Add a new test. > > You might want ot look at how RISC-V handles this. There's all kinds of > > primitives you can use to clear bits and you're usually better using > > those primitives to synthesize the logical operation during initial > > expansion without constructing the constant. For example you can use > > shifts in pairs or triplets, bit clear style instructions, zero > > extensions, and-immediate if you have them, rotates in combination with > > and-immediate, and so-on. These are composable. > > > > In general define_insn_and_split is not a great way to solve these > > problems because it effectively lies about the cost of the patterns it > > matches and it will tend to inhibit further optimizations because of > > those lies about the cost (and to be clear, I'm not talking about > > rtx_cost or insn cost, but instead the internal costing done by combine > > which relies solely on insn counts, particularly with regards to insn > > splitting). > > > > I'm not at all familiar with the details of the loongarch port and this > > should not be considered a review. Just some advice based on working in > > this space on the RISC-V port. > > I tried doing this in define_expand but then combine insists to > transform the bstrins pair back to materializing the mask. The reason > is we have > > /* When not optimizing for size, we care more about the cost > of hot code, and hot code is often in a loop. If a constant > operand needs to be forced into a register, we will often be > able to hoist the constant load out of the loop, so the load > should not contribute to the cost. */ > if (speed || loongarch_immediate_operand_p (outer_code, INTVAL (x))) > { > *total = 0; > return true; > } > > in loongarch_rtx_costs for CONST_INT, i.e. the cost for materializing > the immediate is considered 0 unless optimizing for size... > > Perhaps this loongarch_rtx_costs logic is incorrect but I'm not sure. Tweaking loongarch_rtx_costs prevents the combine: diff --git a/gcc/config/loongarch/loongarch.cc b/gcc/config/loongarch/loongarch.cc index a0f91ab2843..7a86b904264 100644 --- a/gcc/config/loongarch/loongarch.cc +++ b/gcc/config/loongarch/loongarch.cc @@ -4090,18 +4090,13 @@ loongarch_rtx_costs (rtx x, machine_mode mode, int outer_code, switch (code) { case CONST_INT: - if (TARGET_64BIT && outer_code == AND && UINTVAL (x) == 0xffffffff) + if (outer_code == AND && and_operand (x, mode)) { *total = 0; return true; } - /* When not optimizing for size, we care more about the cost - of hot code, and hot code is often in a loop. If a constant - operand needs to be forced into a register, we will often be - able to hoist the constant load out of the loop, so the load - should not contribute to the cost. */ - if (speed || loongarch_immediate_operand_p (outer_code, INTVAL (x))) + if (loongarch_immediate_operand_p (outer_code, INTVAL (x))) { *total = 0; return true; @@ -4118,8 +4113,6 @@ loongarch_rtx_costs (rtx x, machine_mode mode, int outer_code, if (cost == 1 && outer_code == SET && !(float_mode_p && TARGET_HARD_FLOAT)) cost = 0; - else if ((outer_code == SET || GET_MODE (x) == VOIDmode)) - cost = 1; *total = COSTS_N_INSNS (cost); return true; } The first diff is obviously correct, while the following diffs remove some special cases for const materialization. I don't know if removing them will cause a net win, or lose. P.S. even if the logic in the comment removed above is incorrect, there might be still a valid reason not to penalty const_int materialization: the micro architecture likely has some optimization for it. -- Xi Ruoyao <[email protected]>