Re: [PATCH v14 12/19] vect: Extend BB SLP vectorization to use predicated tails

Pengfei Li <[email protected]>
Newsgroups gmane.comp.gcc.patches
Message-ID <[email protected]>
On 30/07/2026 15:56, Christopher Bazley wrote:
> This enables use of a predicate mask or length limit for
> vectorization of basic blocks in cases where previously only the
> equivalent rolled (i.e. loop) form of some source code would have
> been vectorized. Predication is used for groups whose size
> is not neatly divisible into vectors of lengths that can be
> supported directly by the target.
>
> The initial vector mode for an SLP region is "autodetected" by calling
> aarch64_preferred_simd_mode, which prefers SVE modes if supported
> unless configured otherwise (e.g. VNx4SI for int). If at least one
> profitable subgraph can be scheduled then GCC does not try to vectorize
> the region using any other modes, even though their estimated costs
> might otherwise have been lower.
>
> For example, if analysis of a 24-byte group succeeds with vector mode
> V16QI (using types vector(16) and vector(8) char) then the estimated
> cost of the vectorized code is 11+11=22. If analysis of the same group
> succeeds with vector mode VNx16QI (using type vector([16,16]) char for
> both subtrees) then the estimated cost is 15+15=30. In both cases, the
> estimated vectorized cost would beat the estimated scalar cost of
> 96+48=144, but vector([16,16]) is chosen because VNx16QI is tried
> first.
>
> This is mitigated by the fact that a sequence of GIMPLE stmts such as:
>
> vectp.14_86 = x_50(D) + 16;
> slp_mask_87 = .WHILE_ULT (0, 8, { 0, ... });
> .MASK_STORE (vectp.14_86, 8B, slp_mask_87, vect__34.12_85);
>
> are lowered to a fixed-length vector store (e.g., str d30, [x0, 16]) if
> possible, instead of a more literal interpretation such as:
>
> add   x0, x0, 16
> ptrue p7.b, vl8
> st1b  z30.b, p7, [x0]
>
> The vect_record_max_nunits function used during building of an SLP
> tree is updated to prevent it returning failure for BB SLP if the
> group size is not an integral multiple of the number of lanes in the
> vector type; it now allows such cases if the group size is known to
> be less than or equal to the minimum number of lanes in the vector type.
>
> For BB SLP, vect_analyze_slp_instance previously gave up after
> building an SLP tree if it could not prove that the group size was
> at least the maximum lane count across all of the vector types in
> the SLP tree (which is unprovable for scalable vector types), or
> attempted to split the group if it could prove that the group size
> was greater than this maximum but not exactly divisible by it
> (which is also unprovable for scalable vector types).
>
> This function will now provisionally create a new SLP instance if the
> group size definitely does not exceed the minimum number of lanes,
> even if the group size otherwise satisfies conditions that would
> require a loop to be unrolled (e.g., a group of size 3 that uses a
> mixture of V4SI and V8HI types). If the group size lies between the
> minimum and maximum number of lanes then vectorization is still
> abandoned (e.g., a group of size 3 that uses a mixture of
> V2DI and V4SI types).
>
> The vect_prologue_cost_for_slp and vect_create_constant_vectors
> functions expect to be able to get the number of vectors needed for
> an SLP node by dividing the number of scalar elements needed for the
> vectorized statements by the number of subparts in the vector type.
> This relies on the scalar elements occupying a whole number of vectors
> or else one partial vector, which is true of nodes created during
> building of the SLP tree but not always true when a vector type is
> assigned to an external definition node during analysis of statements.
> Update vect_maybe_update_slp_op_vectype to reject late attempts to
> assign a vector type that would result in an external definition
> that occupies "one and half" vectors or similar.
>
> Update test expectations for gcc.dg/vect/vect-over-widen-*.c,
> gcc.target/aarch64/sve/slp_6.c and
> gcc.target/aarch64/sve/vec_construct_*.c.
>
> The vec_construct_*.c tests previously expected their output
> to use Advanced SIMD instead of SVE despite their use of
> vector length agnostic types such as svint16_t and despite
> the fact that they are in the aarch64/sve directory. Since
> BB SLP can now vectorize these tests using VLA types such
> as 'vector([8,8]) char', and because (with one exception) the
> resultant code is deemed profitable relative to scalar code,
> GCC no longer considers vectorizing using non-VLA types such
> as 'vector(8) char' (although the estimated cost with non-VLA
> types might have been lower, had it been calculated).
> Instruction selection is not the focus of these tests, therefore
> I updated them to expect SVE instead (e.g. st1b instead of str)
> and added --param=aarch64-autovec-preference=sve-only to reduce
> future churn.
>
> Because the cost model takes into account predicate mask
> generation for BB SLP with VLA types, the threshold at which
> vectorized code wins against scalar code is higher than
> before. The number of elements stored by vec_construct_3.c was
> increased just enough to allow for that.
>
> gcc/ChangeLog:
>
> 	* tree-vect-slp.cc (vect_record_max_nunits): For BB SLP
> 	vectorization, allow group sizes that are indivisible by the
> 	vector length if they are known not to exceed it.
> 	(vect_build_slp_tree_1): If building SLP failed because of an
> 	unsupported data type or because unrolling was required then use the
> 	known minimum number of subparts when choosing a splitting point for
> 	a VLA vector type.
> 	(vect_build_slp_tree_2): Don't call
> 	can_duplicate_and_interleave_p when doing basic block SLP
> 	vectorization.
> 	(vect_update_slp_min_nunits_for_node): New recursive function.
> 	Update min_nunits to reflect the minimum number of subparts for
> 	all of the vector types used by an SLP subgraph.
> 	(vect_slp_tree_min_nunits): New function. Initialize min_nunits
> 	then call vect_update_slp_min_nunits_for_node.
> 	(vect_analyze_slp_instance): For BB SLP vectorization, create
> 	a new SLP instance if the group size definitely does not exceed
> 	the minimum number of subparts for all of the vector types used
> 	in the SLP tree, even if the group size otherwise satisfies
> 	conditions that would require a loop to be unrolled.
> 	(vectorizable_slp_permutation_1): Instead of asserting that an
> 	SLP tree node's number of lanes is compatible with the chosen
> 	vector width, return a failure indication if incompatible.
> 	* tree-vect-stmts.cc (get_vectype_for_scalar_type): Execute the
> 	old algorithm to find a narrower vector type that can be used to
> 	carve an SLP group into pieces only if partial vector support is
> 	unavailable for either loads or stores, and only if the number of
> 	subparts may be greater than the group size rather than also if
> 	the number of subparts may be equal to the group size.
> 	Otherwise, assume that using a partial vector is more efficient
> 	than carving the group into smaller pieces.  If the natural
> 	vector type is not VLA and it has more subparts than the group
> 	size then find the smallest vector type that contains the group.
> 	(vect_maybe_update_slp_op_vectype): Reject external definitions
> 	when the number of scalar elements represented by the vectorized
> 	operation is not divisible by the number of subparts in a vector
> 	type inferred from the scalar type.  This takes into account the
> 	vectorization factor like vect_get_num_copies although it should
> 	be 1 for BB SLP vectorization.
> 	(vect_get_vector_types_for_stmt): Update the description.
>
> gcc/testsuite/ChangeLog:
>
> 	* gcc.dg/vect/vect-over-widen-10.c: Update test expectations to
> 	avoid spurious matching of scan-tree-dump-not pattern.
> 	* gcc.dg/vect/vect-over-widen-13.c: As above.
> 	* gcc.dg/vect/vect-over-widen-14.c: As above.
> 	* gcc.dg/vect/vect-over-widen-17.c: As above.
> 	* gcc.dg/vect/vect-over-widen-18.c: As above.
> 	* gcc.dg/vect/vect-over-widen-5.c: As above.
> 	* gcc.dg/vect/vect-over-widen-6.c: As above.
> 	* gcc.dg/vect/vect-over-widen-7.c: As above.
> 	* gcc.dg/vect/vect-over-widen-8.c: As above.
> 	* gcc.dg/vect/vect-over-widen-9.c: As above.
> 	* gcc.target/aarch64/sve/slp_6.c: Require ld1b, ld1h and ld1w
> 	instructions in the output instead of forbidding them because
> 	discovery of reduction groups now succeeds.  Also require
> 	the corresponding st1b, st1h and st1w instructions.
> 	* gcc.target/aarch64/sve/vec_construct_1.c:
> 	Expect SVE instead of ASIMD instructions and add
> 	--param=aarch64-autovec-preference=sve-only to stop
> 	flip-flopping.
> 	* gcc.target/aarch64/sve/vec_construct_2.c: Forbid SVE
> 	instructions as well as ASIMD instructions and add
> 	--param=aarch64-autovec-preference=sve-only to stop
> 	flip-flopping.
> 	* gcc.target/aarch64/sve/vec_construct_3.c: Expect SVE instead
> 	of ASIMD instructions and add
> 	--param=aarch64-autovec-preference=sve-only to avoid
> 	flip-flopping. Increase the number of elements stored to
> 	ensure vectorization using SVE is deemed profitable despite
> 	predicate mask costs.
> 	* gcc.target/aarch64/sve/vec_construct_4.c:
> 	Forbid SVE instructions as well as ASIMD instructions and add
> 	--param=aarch64-autovec-preference=sve-only to stop
> 	flip-flopping.  Fix a missing end of comment delimiter.
> 	* gcc.target/aarch64/sve/vec_construct_5.c: As above.
> ---
>   .../gcc.dg/vect/vect-over-widen-10.c          |   2 +-
>   .../gcc.dg/vect/vect-over-widen-13.c          |   2 +-
>   .../gcc.dg/vect/vect-over-widen-14.c          |   2 +-
>   .../gcc.dg/vect/vect-over-widen-17.c          |   2 +-
>   .../gcc.dg/vect/vect-over-widen-18.c          |   2 +-
>   gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c |   2 +-
>   gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c |   2 +-
>   gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c |   2 +-
>   gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c |   2 +-
>   gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c |   2 +-
>   gcc/testsuite/gcc.target/aarch64/sve/slp_6.c  |  18 ++-
>   .../gcc.target/aarch64/sve/vec_construct_1.c  |   6 +-
>   .../gcc.target/aarch64/sve/vec_construct_2.c  |   4 +-
>   .../gcc.target/aarch64/sve/vec_construct_3.c  |  20 +++-
>   .../gcc.target/aarch64/sve/vec_construct_4.c  |   5 +-
>   .../gcc.target/aarch64/sve/vec_construct_5.c  |   6 +-
>   gcc/tree-vect-slp.cc                          | 113 +++++++++++++++---
>   gcc/tree-vect-stmts.cc                        |  97 ++++++++++++---
>   18 files changed, 228 insertions(+), 61 deletions(-)
>
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-10.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-10.c
> index f0140e4ef6d..6efcf739db9 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-10.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-10.c
> @@ -16,5 +16,5 @@
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 1} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 2} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: detected:[^\n]* \(unsigned char\)} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-13.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-13.c
> index 08a65ea5518..720353716cf 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-13.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-13.c
> @@ -48,5 +48,5 @@ main (void)
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* \+} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* / 2} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: detected:[^\n]* = \(signed char\)} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-14.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-14.c
> index dfa09f5d2ca..f1d5f95c543 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-14.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-14.c
> @@ -15,5 +15,5 @@
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* \+} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 1} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: detected:[^\n]* = \(unsigned char\)} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-17.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-17.c
> index 53fcfd0c06c..ac1a0f86727 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-17.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-17.c
> @@ -46,5 +46,5 @@ main (void)
>      adopts realign_load scheme.  It requires rs6000_builtin_mask_for_load to
>      generate mask whose return type is vector char.  */
>   /* { dg-final { scan-tree-dump-not {vector[^\n]*char} "vect" { target vect_hw_misalign } } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-18.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-18.c
> index aa58cd1c957..3ebfaa78270 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-18.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-18.c
> @@ -47,5 +47,5 @@ main (void)
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* |} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* <<} "vect" } } */
>   /* { dg-final { scan-tree-dump {vector[^\n]*char} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c
> index c2ab11a9d32..1d89789a86d 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c
> @@ -49,5 +49,5 @@ main (void)
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* \+ } "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 1} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: detected:[^\n]* \(signed char\)} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c
> index bda92c965e0..62d5a52587e 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c
> @@ -13,5 +13,5 @@
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* \+ } "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 1} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: detected:[^\n]* \(unsigned char\)} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c
> index 1d55e13fb1f..6e09631009a 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c
> @@ -51,5 +51,5 @@ main (void)
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* \+ } "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 2} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: detected:[^\n]* \(signed char\)} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c
> index 553c0712a79..b6d650beab4 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c
> @@ -16,5 +16,5 @@
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* \+ } "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 2} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: detected:[^\n]* \(unsigned char\)} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c b/gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c
> index 36bfc68e053..e82f8a571da 100644
> --- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c
> +++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c
> @@ -56,5 +56,5 @@ main (void)
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 1} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: detected:[^\n]* >> 2} "vect" } } */
>   /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: detected:[^\n]* \(signed char\)} "vect" } } */
> -/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
> +/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
>   /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/sve/slp_6.c b/gcc/testsuite/gcc.target/aarch64/sve/slp_6.c
> index 44d128477d2..0ac3f09cbdb 100644
> --- a/gcc/testsuite/gcc.target/aarch64/sve/slp_6.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sve/slp_6.c
> @@ -36,11 +36,21 @@ vec_slp_##TYPE (TYPE *restrict a, TYPE *restrict b, int n)	\
>   
>   TEST_ALL (VEC_PERM)
>   
> -/* These loops can't use SLP.  */
> -/* { dg-final { scan-assembler-not {\tld1b\t} } } */
> -/* { dg-final { scan-assembler-not {\tld1h\t} } } */
> -/* { dg-final { scan-assembler-not {\tld1w\t} } } */
> +/* When this test was written, these loops couldn't use SLP.  Since changes to
> +   enable tail predication, SLP discovery of size 3 reduction groups instead
> +   succeeds, producing predicated loads such as ld1b z25.b, p7/z, [x1] for
> +   {b[0], b[1], b[2]}.  */
> +
> +/* { dg-final { scan-assembler {\tld1b\t} } } */
> +/* { dg-final { scan-assembler {\tld1h\t} } } */
> +/* { dg-final { scan-assembler {\tld1w\t} } } */
>   /* { dg-final { scan-assembler-not {\tld1d\t} } } */
> +
> +/* { dg-final { scan-assembler {\tst1b\t} } } */
> +/* { dg-final { scan-assembler {\tst1h\t} } } */
> +/* { dg-final { scan-assembler {\tst1w\t} } } */
> +/* { dg-final { scan-assembler-not {\tst1d\t} } } */
> +
>   /* { dg-final { scan-assembler {\tld3b\t} } } */
>   /* { dg-final { scan-assembler {\tld3h\t} } } */
>   /* { dg-final { scan-assembler {\tld3w\t} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_1.c b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_1.c
> index 2f8ce6808a9..eea13c28e49 100644
> --- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_1.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_1.c
> @@ -1,5 +1,5 @@
>   /* { dg-do compile } */
> -/* { dg-options "-O2 -ftree-slp-vectorize" } */
> +/* { dg-options "-O2 -ftree-slp-vectorize --param=aarch64-autovec-preference=sve-only" } */
>   
>   /* Test that a group of stores of 8 elements derived from a horizontal
>      reduction is vectorized by constructing a vector and storing it.
> @@ -30,8 +30,8 @@ foo (svint8_t src0, svint8_t src1, svint8_t src2, svint8_t src3, svint8_t src4,
>     s.h = svaddv_s8 (all, src7);
>   }
>   
> -/* { dg-final { scan-assembler-times {\tins\tv[0-9]+\.b\[[0-9]+\], v[0-9]+\.b\[[0-9]+\]\n} 7 } } */
> -/* { dg-final { scan-assembler-times {\tstr\td[0-9]+, } 1 } } */
> +/* { dg-final { scan-assembler-times {\tinsr\tz[0-9]+\.h, h[0-9]+\n} 7 } } */
> +/* { dg-final { scan-assembler-times {\tst1b\tz[0-9]+\.h, p[0-9]+, \[x[0-9]+\]\n} 1 } } */
>   
>   /* { dg-final { scan-assembler-not {\tstr\tb[0-9]+, } } } */
>   /* { dg-final { scan-assembler-not {\tstrb\tw[0-9]+, } } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_2.c b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_2.c
> index 6715118d7b0..2bf537e13e2 100644
> --- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_2.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_2.c
> @@ -1,5 +1,5 @@
>   /* { dg-do compile } */
> -/* { dg-options "-O2 -ftree-slp-vectorize" } */
> +/* { dg-options "-O2 -ftree-slp-vectorize --param=aarch64-autovec-preference=sve-only" } */
>   
>   /* Test that a group of stores of 8 elements derived from the results of calls
>      to a function that has only vector parameters and returns a scalar result is
> @@ -40,3 +40,5 @@ foo (svint8_t src0, svint8_t src1, svint8_t src2, svint8_t src3, svint8_t src4,
>   
>   /* { dg-final { scan-assembler-not {\tins\tv[0-9]+\.b\[[0-9]+\], w[0-9]+\n} } } */
>   /* { dg-final { scan-assembler-not {\tstr\td[0-9]+, } } } */
> +/* { dg-final { scan-assembler-not {\tfmov\th[0-9]+, h[0-9]+\n} } } */
> +/* { dg-final { scan-assembler-not {\tinsr\tz[0-9]+\.b, w[0-9]+\n} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_3.c b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_3.c
> index 8143d0050ad..ccadaccbcb4 100644
> --- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_3.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_3.c
> @@ -1,7 +1,7 @@
>   /* { dg-do compile } */
> -/* { dg-options "-O2 -ftree-slp-vectorize" } */
> +/* { dg-options "-O2 -ftree-slp-vectorize --param=aarch64-autovec-preference=sve-only" } */
>   
> -/* Test that a group of stores of 8 elements derived from a horizontal
> +/* Test that a group of stores of 14 elements derived from a horizontal
>      reduction is vectorized by constructing a vector and storing it
>      even if the results of the reductions are narrowed.
>      Since there are no GPR-to-SIMD register transfers, there is no
> @@ -13,12 +13,14 @@
>   
>   struct S
>   {
> -  char a, b, c, d, e, f, g, h;
> +  char a, b, c, d, e, f, g, h, i, j, k, l, m, n;
>   } s;
>   
>   void
>   foo (svint16_t src0, svint32_t src1, svint16_t src2, svint32_t src3,
> -     svint32_t src4, svint16_t src5, svint32_t src6, svint16_t src7)
> +     svint32_t src4, svint16_t src5, svint32_t src6, svint16_t src7,
> +     svint16_t src8, svint32_t src9, svint16_t src10, svint32_t src11,
> +     svint32_t src12, svint16_t src13)
>   {
>     svbool_t all16 = svptrue_b16 ();
>     svbool_t all32 = svptrue_b32 ();
> @@ -30,10 +32,16 @@ foo (svint16_t src0, svint32_t src1, svint16_t src2, svint32_t src3,
>     s.f = svminv_s16 (all16, src5);
>     s.g = svlastb_s32 (svptrue_pat_b32 (SV_VL1), src6);
>     s.h = svaddv_s16 (all16, src7);
> +  s.i = svmaxv_s16 (all16, src8);
> +  s.j = svminv_s32 (all32, src9);
> +  s.k = svlastb_s16 (svptrue_pat_b16 (SV_VL1), src10);
> +  s.l = svaddv_s32 (all32, src11);
> +  s.m = svmaxv_s32 (all32, src12);
> +  s.n = svminv_s16 (all16, src13);
>   }
>   
> -/* { dg-final { scan-assembler-times {\tins\tv[0-9]+\.b\[[0-9]+\], v[0-9]+\.b\[[0-9]+\]\n} 7 } } */
> -/* { dg-final { scan-assembler-times {\tstr\td[0-9]+, } 1 } } */
> +/* { dg-final { scan-assembler-times {\tinsr\tz[0-9]+\.b, b[0-9]+\n} 13 } } */
> +/* { dg-final { scan-assembler-times {\tst1b\tz[0-9]+\.b, p[0-9]+, \[x[0-9]\]\n} 1 } } */
>   
>   /* { dg-final { scan-assembler-not {\tstr\tb[0-9]+, } } } */
>   /* { dg-final { scan-assembler-not {\tstrb\tw[0-9]+, } } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_4.c b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_4.c
> index 49f8114b64c..3d41af684a3 100644
> --- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_4.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_4.c
> @@ -1,5 +1,5 @@
>   /* { dg-do compile } */
> -/* { dg-options "-O2 -ftree-slp-vectorize" } */
> +/* { dg-options "-O2 -ftree-slp-vectorize --param=aarch64-autovec-preference=sve-only" } */
>   
>   /* Test that a group of stores of 8 elements derived from a horizontal
>      reduction is not vectorized by constructing a vector and storing it
> @@ -33,5 +33,6 @@ foo (svint16_t src0, svint8_t src1, svint16_t src2, svint8_t src3,
>   /* { dg-final { scan-assembler-times {\tstp\tw[0-9]+, w[0-9]+,} 4 } } */
>   
>   /* { dg-final { scan-assembler-not {\tins\tv[0-9]+\.s\[[0-9]+\], w[0-9]+\n} } } */
> -/* { dg-final { scan-assembler-not {\tfmov\ts[0-9]+, w[0-9]+\n} } }
> +/* { dg-final { scan-assembler-not {\tfmov\ts[0-9]+, w[0-9]+\n} } } */
>   /* { dg-final { scan-assembler-not {\tstp\tq[0-9]+, q[0-9]+,} } } */
> +/* { dg-final { scan-assembler-not {\tinsr\tz[0-9]+.s, w[0-9]+\n} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_5.c b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_5.c
> index 983d6c69ebc..89e57406c0e 100644
> --- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_5.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_5.c
> @@ -1,5 +1,5 @@
>   /* { dg-do compile } */
> -/* { dg-options "-O2 -ftree-slp-vectorize" } */
> +/* { dg-options "-O2 -ftree-slp-vectorize --param=aarch64-autovec-preference=sve-only" } */
>   
>   /* Test that a group of stores of 8 elements derived from lane extractions is
>      vectorized by constructing a vector and storing it.  Since there are no
> @@ -30,8 +30,8 @@ foo (svint8_t src0, svint8_t src1, svint8_t src2, svint8_t src3, svint8_t src4,
>     s.h = svlastb_s8 (p, src7);
>   }
>   
> -/* { dg-final { scan-assembler-times {\tins\tv[0-9]+\.b\[[0-9]+\], v[0-9]+\.b\[[0-9]+\]\n} 7 } } */
> -/* { dg-final { scan-assembler-times {\tstr\td[0-9]+, } 1 } } */
> +/* { dg-final { scan-assembler-times {\tinsr\tz[0-9]+\.h, h[0-9]+\n} 7 } } */
> +/* { dg-final { scan-assembler-times {\tst1b\tz[0-9]+\.h, p[0-9]+, \[x[0-9]+\]\n} 1 } } */
>   
>   /* { dg-final { scan-assembler-not {\tstr\tb[0-9]+, } } } */
>   /* { dg-final { scan-assembler-not {\tstrb\tw[0-9]+, } } } */
> diff --git a/gcc/tree-vect-slp.cc b/gcc/tree-vect-slp.cc
> index 7bb3d400ca0..ca9bae97e7e 100644
> --- a/gcc/tree-vect-slp.cc
> +++ b/gcc/tree-vect-slp.cc
> @@ -1114,8 +1114,12 @@ vect_record_max_nunits (vec_info *vinfo, stmt_vec_info stmt_info,
>       }
>   
>     /* If populating the vector type requires unrolling then fail
> -     before adjusting *max_nunits for basic-block vectorization.  */
> +     before adjusting *max_nunits for basic-block vectorization.
> +     Allow group sizes that are indivisible by the vector length only if they
> +     are known not to exceed the vector length.  We may be able to support such
> +     cases by generating constant masks.  */
>     if (is_a <bb_vec_info> (vinfo)
> +      && maybe_gt (group_size, TYPE_VECTOR_SUBPARTS (vectype))
>         && !multiple_p (group_size, TYPE_VECTOR_SUBPARTS (vectype)))
>       {
>         if (dump_enabled_p ())
> @@ -1693,16 +1697,22 @@ vect_build_slp_tree_1 (vec_info *vinfo, unsigned char *swap,
>   
>     if (maybe_soft_fail)
>       {
> -      unsigned HOST_WIDE_INT const_nunits;
> -      if (!TYPE_VECTOR_SUBPARTS
> -	    (soft_fail_nunits_vectype).is_constant (&const_nunits)
> -	  || const_nunits > group_size)
> +      /* Use the known minimum number of subparts for VLA because we still need
> +	 to choose a splitting point although the choice is more arbitrary.  */
> +      unsigned HOST_WIDE_INT const_nunits = constant_lower_bound (
> +	  TYPE_VECTOR_SUBPARTS (soft_fail_nunits_vectype));
> +
> +      if (const_nunits > group_size)
>   	matches[0] = false;
>         else
>   	{
>   	  /* With constant vector elements simulate a mismatch at the
>   	     point we need to split.  */
> +	  gcc_assert ((const_nunits & (const_nunits - 1)) == 0);
>   	  unsigned tail = group_size & (const_nunits - 1);
> +	  if (tail == 0)
> +	    tail = const_nunits;
> +	  gcc_assert (group_size >= tail);
>   	  memset (&matches[group_size - tail], 0, sizeof (bool) * tail);
>   	}
>         return false;
> @@ -2442,13 +2452,21 @@ vect_build_slp_tree_2 (vec_info *vinfo, slp_tree node,
>   		  /* Check whether we can build the invariant.  If we can't
>   		     we never will be able to.  */
>   		  tree type = TREE_TYPE (chains[0][n].op);
> -		  if (!GET_MODE_SIZE (vinfo->vector_mode).is_constant ()
> -		      && (TREE_CODE (type) == BOOLEAN_TYPE
> -			  || !can_duplicate_and_interleave_p (vinfo, group_size,
> -							      type)))
> +		  if (!GET_MODE_SIZE (vinfo->vector_mode).is_constant ())
>   		    {
> -		      matches[0] = false;
> -		      goto out;
> +		      if (TREE_CODE (type) == BOOLEAN_TYPE)
> +			{
> +			  matches[0] = false;
> +			  goto out;
> +			}
> +
> +		      if (!is_a<bb_vec_info> (vinfo)
> +			  && !can_duplicate_and_interleave_p (vinfo, group_size,
> +							      type))
> +			{
> +			  matches[0] = false;
> +			  goto out;
> +			}
>   		    }
>   		}
>   	      else if (dt != vect_internal_def)
> @@ -2881,7 +2899,7 @@ out:
>   		    uniform_val = NULL_TREE;
>   		    break;
>   		  }
> -	      if (!uniform_val
> +	      if (!uniform_val && !is_a<bb_vec_info> (vinfo)
>   		  && !can_duplicate_and_interleave_p (vinfo,
>   						      oprnd_info->ops.length (),
>   						      TREE_TYPE (op0)))
> @@ -5079,6 +5097,53 @@ vect_analyze_slp_reductions (loop_vec_info loop_vinfo,
>     return true;
>   }
>   
> +/* Update MIN_NUNITS to reflect the minimum number of subparts for all of the
> +   vector types used by the SLP subgraph rooted at NODE.  VISITED is used to
> +   avoid reevaluating any node in the subgraph; it thereby prevents infinite
> +   recursion should a cycle be encountered. The value of MIN_NUNITS will only be
> +   updated if any node in the subgraph has a vector type with a number of
> +   subparts that is smaller than the passed-in value of MIN_NUNITS. Before
> +   calling this function for the first time, initialize MIN_NUNITS to
> +   UINT64_MAX.  */
> +
> +static void
> +vect_update_slp_min_nunits_for_node (slp_tree node, poly_uint64 &min_nunits,
> +				     hash_set<slp_tree> &visited)
> +{
> +  if (!node || SLP_TREE_DEF_TYPE (node) != vect_internal_def)
> +    return;
> +
> +  if (visited.add (node))
> +    return;
> +
> +  for (slp_tree child : SLP_TREE_CHILDREN (node))
> +    vect_update_slp_min_nunits_for_node (child, min_nunits, visited);
> +
> +  tree vectype = SLP_TREE_VECTYPE (node);
> +  if (!vectype)
> +    return;
> +
> +  /* All unit counts have the form vec_info::vector_size * X for some
> +     rational X, therefore we know the values are ordered.  */
> +  poly_uint64 nunits = TYPE_VECTOR_SUBPARTS (vectype);
> +  min_nunits = known_eq (min_nunits, UINT64_MAX)
> +		 ? nunits
> +		 : ordered_min (min_nunits, nunits);
> +}
> +
> +/* For NODE, return the minimum number of subparts for all of the vector
> +   types used in the given SLP subgraph.  */
> +
> +static poly_uint64
> +vect_slp_tree_min_nunits (slp_tree node)
> +{
> +  poly_uint64 min_nunits = UINT64_MAX;
> +  hash_set<slp_tree> visited;
> +  vect_update_slp_min_nunits_for_node (node, min_nunits, visited);
> +  gcc_checking_assert (known_ne (min_nunits, UINT64_MAX));
> +  return min_nunits;
> +}
> +
>   /* Analyze an SLP instance starting from a group of grouped stores.  Call
>      vect_build_slp_tree to build a tree of packed stmts if possible.
>      Return FALSE if it's impossible to SLP any stmt in the group.  */
> @@ -5148,8 +5213,8 @@ vect_analyze_slp_instance (vec_info *vinfo,
>         poly_uint64 unrolling_factor
>   	= calculate_unrolling_factor (max_nunits, group_size);
>   
> -      if (maybe_ne (unrolling_factor, 1U)
> -	  && is_a <bb_vec_info> (vinfo))
> +      if (maybe_ne (unrolling_factor, 1U) && is_a<bb_vec_info> (vinfo)
> +	  && !known_ge (vect_slp_tree_min_nunits (node), group_size))

Could the new condition use maybe_lt instead? Avoiding the negated 
known_ge reads clearer.

You may also want to update the dump message in this if block.


Thanks,
Pengfei


>   	{
>   	  unsigned HOST_WIDE_INT const_max_nunits;
>   	  if (!max_nunits.is_constant (&const_max_nunits)
> @@ -5235,8 +5300,10 @@ vect_analyze_slp_instance (vec_info *vinfo,
>   	  tree vectype = get_vectype_for_scalar_type (vinfo, scalar_type,
>   						      1 << floor_log2 (i));
>   	  unsigned HOST_WIDE_INT const_nunits;
> +	  poly_uint64 nunits = TYPE_VECTOR_SUBPARTS (vectype);
>   	  if (vectype
> -	      && TYPE_VECTOR_SUBPARTS (vectype).is_constant (&const_nunits))
> +	      && (const_nunits = constant_lower_bound (nunits)) > 1
> +	      && (i % const_nunits) == 0)
>   	    {
>   	      /* Split into two groups at the first vector boundary.  */
>   	      gcc_assert ((const_nunits & (const_nunits - 1)) == 0);
> @@ -11863,7 +11930,21 @@ vectorizable_slp_permutation_1 (vec_info *vinfo, gimple_stmt_iterator *gsi,
>         unpack_factor = 1;
>       }
>     unsigned olanes = unpack_factor * ncopies * SLP_TREE_LANES (node);
> -  gcc_assert (repeating_p || multiple_p (olanes, nunits));
> +
> +  /* With fully-predicated BB-SLP, an external node's number of lanes can be
> +     incompatible with the chosen vector width (e.g., lane packs of 3 with a
> +     natural 2-lane vector type).  */
> +  if (!repeating_p && !multiple_p (olanes, nunits))
> +    {
> +      if (dump_p)
> +	dump_printf_loc (MSG_MISSED_OPTIMIZATION, vect_location,
> +			 "unsupported permutation %p: vector type %T,"
> +			 " nunits=" HOST_WIDE_INT_PRINT_UNSIGNED
> +			 " ncopies=%" PRIu64 ", lanes=%u and unpack=%u\n",
> +			 (void *) node, vectype, estimated_poly_value (nunits),
> +			 ncopies, SLP_TREE_LANES (node), unpack_factor);
> +      return -1;
> +    }
>   
>     /* Compute the { { SLP operand, vector index}, lane } permutation sequence
>        from the { SLP operand, scalar lane } permutation as recorded in the
> diff --git a/gcc/tree-vect-stmts.cc b/gcc/tree-vect-stmts.cc
> index bcc8581b4cc..168d79c91d0 100644
> --- a/gcc/tree-vect-stmts.cc
> +++ b/gcc/tree-vect-stmts.cc
> @@ -13826,9 +13826,11 @@ get_related_vectype_for_scalar_type (machine_mode prevailing_mode,
>   /* Function get_vectype_for_scalar_type.
>   
>      Returns the vector type corresponding to SCALAR_TYPE as supported
> -   by the target.  If GROUP_SIZE is nonzero and we're performing BB
> -   vectorization, make sure that the number of elements in the vector
> -   is no bigger than GROUP_SIZE.  */
> +   by the target, or NULL_TREE if none.  If GROUP_SIZE is nonzero and we're
> +   performing BB vectorization, ensure that the number of subparts in the
> +   returned type is no bigger than GROUP_SIZE, or no smaller than GROUP_SIZE if
> +   a vector mask or length can be used to limit the number of lanes loaded and
> +   stored.  */
>   
>   tree
>   get_vectype_for_scalar_type (vec_info *vinfo, tree scalar_type,
> @@ -13853,11 +13855,23 @@ get_vectype_for_scalar_type (vec_info *vinfo, tree scalar_type,
>     if (vectype)
>       vinfo->used_vector_modes.add (TYPE_MODE (vectype));
>   
> -  /* If the natural choice of vector type doesn't satisfy GROUP_SIZE,
> -     try again with an explicit number of elements.  */
> -  if (vectype
> -      && group_size
> -      && maybe_ge (TYPE_VECTOR_SUBPARTS (vectype), group_size))
> +  if (!vectype || !group_size)
> +    return vectype;
> +
> +  /* The natural choice of vector type satisfies GROUP_SIZE if it can be used to
> +     store the whole group without a mask or length limit (known_eq) or if it
> +     can be used to carve the group into pieces (known_lt).  */
> +  const poly_uint64 natural_nelts = TYPE_VECTOR_SUBPARTS (vectype);
> +  if (known_le (natural_nelts, group_size))
> +    return vectype;
> +
> +  /* We might be able to use a partial vector to handle a GROUP_SIZE that is not
> +     satisfied by the natural choice of vector type.  Assume this is more
> +     efficient than carving the group into pieces.  */
> +  if (vect_get_load_store_partial_vector_style (vectype, true)
> +	== vect_load_store_partial_vectors_none
> +      || vect_get_load_store_partial_vector_style (vectype, false)
> +	   == vect_load_store_partial_vectors_none)
>       {
>         /* Start with the biggest number of units that fits within
>   	 GROUP_SIZE and halve it until we find a valid vector type.
> @@ -13879,8 +13893,32 @@ get_vectype_for_scalar_type (vec_info *vinfo, tree scalar_type,
>   	  nunits /= 2;
>   	}
>         while (nunits > 1 && !vectype);
> +
> +      return vectype;
> +    }
> +
> +  /* Interoperability between modes requires one to be a constant multiple of
> +     the other, so do not try to substitute a narrower fixed-length vector type
> +     for a variable-length vector type.  Do not try to minimize the number of
> +     elements in a variable-length type either because it can increase the size
> +     of vector constants (e.g., if QImode values are stored in HImode).  */
> +  if (!natural_nelts.is_constant ())
> +    return vectype;
> +
> +  /* Try to minimize the number of subparts in the partial vector type.  Start
> +     with the smallest number of units that contains GROUP_SIZE and double it
> +     until we find a valid vector type.  */
> +  unsigned int nunits = 1 << ceil_log2 (group_size);
> +  do
> +    {
> +      vectype = get_related_vectype_for_scalar_type (vinfo->vector_mode,
> +						     scalar_type, nunits);
> +      nunits *= 2;
>       }
> +  while (known_le (nunits, natural_nelts) && !vectype);
>   
> +  /* We should have rediscovered the natural vector type if nothing else.  */
> +  gcc_assert (vectype != NULL);
>     return vectype;
>   }
>   
> @@ -13900,10 +13938,10 @@ get_vectype_for_scalar_type (vec_info *vinfo, tree scalar_type, slp_tree node)
>   /* Function get_mask_type_for_scalar_type.
>   
>      Returns the mask type corresponding to a result of comparison
> -   of vectors of specified SCALAR_TYPE as supported by target.
> -   If GROUP_SIZE is nonzero and we're performing BB vectorization,
> -   make sure that the number of elements in the vector is no bigger
> -   than GROUP_SIZE.  */
> +   of vectors of specified SCALAR_TYPE as supported by target or NULL_TREE if
> +   none.  If GROUP_SIZE is nonzero and we're performing BB vectorization, ensure
> +   that the returned mask type is suitable for use with the vector type that
> +   would be chosen by get_vectype_for_scalar_type.  */
>   
>   tree
>   get_mask_type_for_scalar_type (vec_info *vinfo, tree scalar_type,
> @@ -14173,7 +14211,33 @@ vect_maybe_update_slp_op_vectype (vec_info *vinfo, slp_tree op, tree vectype)
>         && SLP_TREE_DEF_TYPE (op) == vect_external_def
>         && SLP_TREE_LANES (op) > 1)
>       return false;
> -  (void) vinfo; /* FORNOW */
> +
> +  /* When the vectorizer falls back to building vector operands from scalars,
> +     it can create SLP trees with external defs that have a number of lanes not
> +     divisible by the number of subparts in a vector type naively inferred from
> +     the scalar type.  Reject such types to avoid ICE when later computing the
> +     prologue cost for invariant operands.  */
> +  if (SLP_TREE_DEF_TYPE (op) == vect_external_def)
> +    {
> +      poly_uint64 vf = vinfo->vectorization_factor;
> +
> +      vf *= SLP_TREE_LANES (op);
> +
> +      if (maybe_lt (TYPE_VECTOR_SUBPARTS (vectype), vf)
> +	  && !multiple_p (vf, TYPE_VECTOR_SUBPARTS (vectype)))
> +	{
> +	  if (dump_enabled_p ())
> +	    dump_printf_loc (MSG_MISSED_OPTIMIZATION, vect_location,
> +			     "lanes=" HOST_WIDE_INT_PRINT_UNSIGNED
> +			     " is not divisible by "
> +			     "subparts=" HOST_WIDE_INT_PRINT_UNSIGNED ".\n",
> +			     estimated_poly_value (vf),
> +			     estimated_poly_value (
> +			       TYPE_VECTOR_SUBPARTS (vectype)));
> +	  return false;
> +	}
> +    }
> +
>     SLP_TREE_VECTYPE (op) = vectype;
>     return true;
>   }
> @@ -14858,9 +14922,10 @@ vect_gen_while_not (gimple_seq *seq, tree mask_type, tree start_index,
>   
>   /* Try to compute the vector types required to vectorize STMT_INFO,
>      returning true on success and false if vectorization isn't possible.
> -   If GROUP_SIZE is nonzero and we're performing BB vectorization,
> -   take sure that the number of elements in the vectors is no bigger
> -   than GROUP_SIZE.
> +   If GROUP_SIZE is nonzero and we're performing BB vectorization, ensure that
> +   the number of subparts in *STMT_VECTYPE_OUT is no bigger than GROUP_SIZE, or
> +   no smaller than GROUP_SIZE if a vector mask or length can be used to limit
> +   the number of lanes loaded and stored.
>   
>      On success:
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.