[gcc(refs/users/meissner/heads/work252-bugs)] PR target/992493: Optimize splat of a V2DF/V2DI extract with constant element

Michael Meissner via Gcc-cvs <[email protected]>
Newsgroups gmane.comp.gcc.cvs
Message-ID <[email protected]>
https://gcc.gnu.org/g:930a96583b3e1081054719af06ed349294658326

commit 930a96583b3e1081054719af06ed349294658326
Author: Michael Meissner <[email protected]>
Date:   Fri Jul 24 00:26:35 2026 -0400

    PR target/992493: Optimize splat of a V2DF/V2DI extract with constant element
    
    We had optimizations for splat of a vector extract for the other vector
    types, but we missed having one for V2DI and V2DF.  This patch adds a
    combiner insn to do this optimization.
    
    In looking at the source, we had similar optimizations for V4SI and V4SF
    extract and splats, but we missed doing V2DI/V2DF.
    
    Without the patch for the code:
    
            vector long long splat_dup_l_0 (vector long long v)
            {
              return __builtin_vec_splats (__builtin_vec_extract (v, 0));
            }
    
    the compiler generates (on a little endian power9):
    
            splat_dup_l_0:
                    mfvsrld 9,34
                    mtvsrdd 34,9,9
                    blr
    
    Now it generates:
    
            splat_dup_l_0:
                    xxpermdi 34,34,34,3
                    blr
    
    I have committed all of the patches in my backlog (dense math registers, other
    -mcpu=future instructions, random bug fixes, support for _Float16 and
    __bfloat16, and optimizations for vector logical operations on power10/power11)
    into the IBM vendor branch:
    
            vendors/ibm/gcc-17-future
    
    2026-07-24  Michael Meissner  <[email protected]>
    
    gcc/
    
            PR target/99293
            * config/rs6000/vsx.md (vsx_splat_extract_<mode>): New insn.
    
    gcc/testsuite/
    
            PR target/99293
            * gcc.target/powerpc/pr99293.c: New test.

Diff:
---
 gcc/config/rs6000/vsx.md                   | 18 ++++++++++++++++++
 gcc/testsuite/gcc.target/powerpc/pr99293.c | 22 ++++++++++++++++++++++
 2 files changed, 40 insertions(+)

diff --git a/gcc/config/rs6000/vsx.md b/gcc/config/rs6000/vsx.md
index ee89107528a3..58318f5bb5e2 100644
--- a/gcc/config/rs6000/vsx.md
+++ b/gcc/config/rs6000/vsx.md
@@ -4829,6 +4829,24 @@
   "lxvdsx %x0,%y1"
   [(set_attr "type" "vecload")])
 
+;; Optimize SPLAT of an extract from a V2DF/V2DI vector with a constant element
+(define_insn "*vsx_splat_extract_<mode>"
+  [(set (match_operand:VSX_D 0 "vsx_register_operand" "=wa")
+	(vec_duplicate:VSX_D
+	 (vec_select:<VEC_base>
+	  (match_operand:VSX_D 1 "vsx_register_operand" "wa")
+	  (parallel [(match_operand 2 "const_0_to_1_operand" "n")]))))]
+  "VECTOR_MEM_VSX_P (<MODE>mode)"
+{
+  int which_word = INTVAL (operands[2]);
+  if (!BYTES_BIG_ENDIAN)
+    which_word = 1 - which_word;
+
+  operands[3] = GEN_INT (which_word ? 3 : 0);
+  return "xxpermdi %x0,%x1,%x1,%3";
+}
+  [(set_attr "type" "vecperm")])
+
 ;; V4SI splat support
 (define_insn "vsx_splat_v4si"
   [(set (match_operand:V4SI 0 "vsx_register_operand" "=wa,wa")
diff --git a/gcc/testsuite/gcc.target/powerpc/pr99293.c b/gcc/testsuite/gcc.target/powerpc/pr99293.c
new file mode 100644
index 000000000000..20adc1f27f65
--- /dev/null
+++ b/gcc/testsuite/gcc.target/powerpc/pr99293.c
@@ -0,0 +1,22 @@
+/* { dg-do compile { target powerpc*-*-* } } */
+/* { dg-require-effective-target powerpc_vsx_ok } */
+/* { dg-options "-O2 -mvsx" } */
+
+/* Test for PR 99263, which wants to do:
+	__builtin_vec_splats (__builtin_vec_extract (v, n))
+
+   where v is a V2DF or V2DI vector and n is either 0 or 1.  Previously the
+   compiler would do a direct move to the GPR registers to select the item and a
+   direct move from the GPR registers to do the splat.  */
+
+vector long long splat_dup_l_0 (vector long long v)
+{
+  return __builtin_vec_splats (__builtin_vec_extract (v, 0));
+}
+
+vector long long splat_dup_l_1 (vector long long v)
+{
+  return __builtin_vec_splats (__builtin_vec_extract (v, 1));
+}
+
+/* { dg-final { scan-assembler-times "xxpermdi" 2 } } */
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.