[Bug target/126516] [17 Regression] 9% slowdown of calculix on Zen3 since r17-2723-g3f2b3eced0f827

"rguenth at gcc dot gnu.org via Gcc-bugs" <[email protected]>
Newsgroups gmane.comp.gcc.bugs
Message-ID <[email protected]/bugzilla/>
https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126516

--- Comment #5 from Richard Biener <rguenth at gcc dot gnu.org> ---
I'll note the change will aggressively prefer SSE vectorization width,
irrespective of the cost of the operations leading up to the fold-left
reduction.

Also (but unlikely here), for slp_inst_kind_reduc_group we only
actually implement fold-left reductions for SLP_TREE_LANES == 1
unless there's no actual fold-left reduction to be performed
(VF == 1).  Meaning, slp_inst_kind_reduc_group with VF == 1
is free, which isn't reflected in the change.  But unlikely to be
the issue here.

So you might want to check whether the following fixes the regression:

diff --git a/gcc/config/i386/i386.cc b/gcc/config/i386/i386.cc
index dcfe4531f11..b2a96820ad3 100644
--- a/gcc/config/i386/i386.cc
+++ b/gcc/config/i386/i386.cc
@@ -27212,7 +27212,7 @@ ix86_vector_costs::finish_cost (const vector_costs
*scalar_costs)
     {
       unsigned int vf = vect_vf_for_cost (loop_vinfo);
       for (auto inst : LOOP_VINFO_SLP_INSTANCES (loop_vinfo))
-       if ((SLP_INSTANCE_KIND (inst) == slp_inst_kind_reduc_group
+       if (((SLP_INSTANCE_KIND (inst) == slp_inst_kind_reduc_group && vf != 1)
             || SLP_INSTANCE_KIND (inst) == slp_inst_kind_reduc_chain)
            && (vect_reduc_type (loop_vinfo, SLP_INSTANCE_TREE (inst))
                == FOLD_LEFT_REDUCTION))
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.