Re: [PATCH 1/2] doc: describe the lane freedom of the widen_ssum and widen_usum patterns

Richard Biener <[email protected]> Wed, 5 Aug 2026 14:41:27 +0200 (CEST)
Newsgroups gmane.comp.gcc.patches
Message-ID <[email protected]>
On Wed, 5 Aug 2026, Kyrylo Tkachov wrote:

> 
> 
> > On 4 Aug 2026, at 16:19, Tamar Christina <[email protected]> wrote:
> > 
> >> -----Original Message-----
> >> From: Robin Dapp <[email protected]>
> >> Sent: 04 August 2026 15:07
> >> To: Richard Biener <[email protected]>; Kyrylo Tkachov
> >> <[email protected]>
> >> Cc: [email protected]; Tamar Christina <[email protected]>;
> >> Robin Dapp <[email protected]>
> >> Subject: Re: [PATCH 1/2] doc: describe the lane freedom of the widen_ssum
> >> and widen_usum patterns
> >> 
> >>> I wonder if this is a step in the correct direction, see PR67612 where
> >>> we want to apply a SLP pattern to exercise the actual lane mapping
> >>> done.
> >> 
> >>> Do we want to have separate optabs for this?  IMO what you document
> >>> would be better named reduc_widen_ssum_optab (as opposed to
> >>> reduc_widen_ssum_scal_optab or widen_ssum_optab).  Where the _scal
> >>> variant would be expected to accumulate to a scalar (or lane zero),
> >>> the reduc_*_optab would be free in how to accumulate lanes and
> >>> the widen_ssum_optab would lay out exactly which source lanes are
> >>> summed to which destination lanes (I hope ISAs have matching behavior
> >>> here).
> >> 
> >> I would be very much in favor of a separate optab rather than narrowing
> >> widen_[us]sum's scope.  You could argue that sum vs plus is already an
> >> inconsistency, though.
> >> Last year I experimented with making widen_[us]sum available to RVV.
> >> That necessitated a few changes in the vectorizer (we don't expect
> >> "SLP-style" "number of lanes stays the same" everywhere) but it's not
> >> that big of a change.
> >> 
> >> Of course the general question remains where the widening should
> >> actually be recognized.  For riscv it works quite well doing it "late".
> >> 
> > 
> > I think the widening and other similar optimizations need to stay in the vectorizer
> > because for masked based targets like SVE if we generate an unpacking then
> > the predicate needs to be unpacked too.
> > 
> > That means any "late" detection has to clean up the predicate operations too
> > which usually brings the chain out of what combine supports.  But also by
> > recognizing it In the vectorizer it allows us to cost alternate strategies like
> > unpacking or using sparse vectors.
> > 
> > So I don't think we can move this past costing as it has a significant effect on
> > Codegen.
> > 
> 
> So as I’m not dealing day-to-day with the vectorizer, what would you recommend as the way forward here?
> Rename the optabs currently used for lane-agnostic reduction to reduc_widen_[us]sum_optab and implement those for aarch64 as in patch 2/2 and document the widen_[us]sum optabs as lane-preserving?

naming consistency would be nice indeed.  Note WIDEN_SUM_EXPR
and related should also be unambiguously documented (I'm fine with
chosing lane-agnostic interpretations for them - in fact they should
all cease to exist and be replaced with internal functions only).

> Thanks,
> Kyrill
> 
> > Thanks,
> > Tamar
> > 
> >> --
> >> Regards
> >> Robin
> 
> 
> 

-- 
Richard Biener <[email protected]>
SUSE Software Solutions Germany GmbH,
Frankenstrasse 146, 90461 Nuernberg, Germany;
GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG Nuernberg)