Re: [PATCH][committed][v2] match: fold two idioms built from the negation of a value

Kyrylo Tkachov <[email protected]>
Newsgroups gmane.comp.gcc.patches
Message-ID <[email protected]>

> On 8 Aug 2026, at 00:51, H.J. Lu <[email protected]> wrote:
> 
> On Thu, Aug 6, 2026 at 10:32 PM <[email protected]> wrote:
>> 
>> From: Kyrylo Tkachov <[email protected]>
>> 
>> X | -X has every bit from the lowest set bit of X upwards, so adding X to
>> it clears that bit, and masking with it isolates the padding needed to
>> round X up:
>> 
>>  X + (X | -X) -> X & (X - 1)
>>  X + ((-X) & (C - 1)) -> (X + C - 1) & -C   for a power of two C
>> 
>> The second is the alignment round up written with the padding computed
>> first, which is how allocators tend to spell it.
>> 
>> Neither needs a wrapping type.  X - 1 overflows only for the most
>> negative value, where the source already does, and rounding X up is
>> representable exactly when X + C - 1 is, because the largest multiple of
>> C below the maximum leaves room for C - 1.  Restrict both rules to integral
>> types.  The bitwise operations also accept fixed-point types, whose
>> saturating arithmetic does not have these integer semantics.
>> 
>>  int f (int x) { return x + ((-x) & 15); }
>> 
>> aarch64 -O2:
>> 
>>  before                          after
>>    neg   w1, w0                    add   w0, w0, 15
>>    and   w1, w1, 15                and   w0, w0, -16
>>    add   w0, w1, w0
>> 
>> The vector spelling folds too, a uniform vector constant is matched with
>> uniform_integer_cst_p.
>> 
>> Keep trapping and sanitized negations.  Also keep the source addition
>> of X + (X | -X) under overflow sanitization because the replacement
>> would remove one recoverable diagnostic.  Use :s on the consumed bitwise
>> expression.  This prevents new work when it remains live, but allows the
>> fold when the replacement arithmetic is already available.
>> 
>> Reuse the matched uniform constant and test its wide value directly.
>> This avoids rebuilding the same constant only to inspect it.
>> 
>> Bootstrapped and tested on aarch64-none-linux-gnu.
>> Committed as per Richard's approval.
>> 
>> gcc/ChangeLog:
>> 
>>        * match.pd (X + (X | -X)): New simplification.
>>        (X + ((-X) & (C - 1))): Likewise.
>> 
>> gcc/testsuite/ChangeLog:
>> 
>>        * gcc.dg/tree-ssa/signbit-1.c: New test.
>>        * gcc.dg/tree-ssa/alignup-2.c: New test.
>>        * gcc.dg/tree-ssa/vector-alignup-1.c: New test.
>>        * gcc.dg/tree-ssa/alignup-overflow-1.c: New test.
>>        * gcc.dg/tree-ssa/alignup-overflow-2.c: New test.
>>        * gcc.dg/tree-ssa/signbit-4.c: New test.
>> 
> 
> On Linux/x86-64, I got
> 
> FAIL: gcc.dg/tree-ssa/signbit-1.c scan-tree-dump-times optimized " \\| " 1
> FAIL: gcc.dg/tree-ssa/signbit-1.c scan-tree-dump-times optimized " != 0" 3

Sorry, I had some other match.pd fold patterns in my tree and there was a phase mismatch in my commit order.
These scans and functions should not be in this patch.
I’ve removed them with g:1c37c631abe1de02e9b843a264b9923c12e9a68d
Thanks for catching it.
Kyrill

> 
> -- 
> H.J.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.