From: Kyrylo Tkachov <[email protected]>
combine_simplify_rtx runs make_compound_operation over the address of a MEM
so that the shifts and masks combine has introduced are turned back into the
forms the target's address predicates recognize. A PREFETCH holds an address
too, and aarch64 classifies it with the same routine it uses for a MEM, but
that address never gets the treatment, so a zero extend that combine has
rewritten as an AND with 0xffffffff is never rewritten back:
Failed to match this instruction:
(prefetch (plus:DI (and:DI (subreg:DI (reg:SI 108 [ i ]) 0)
(const_int 4294967295 [0xffffffff]))
(reg/f:DI 107 [ b ]))
(const_int 0 [0])
(const_int 3 [0x3]))
On aarch64 the result is that __builtin_prefetch (b + i, 0, 3) with a 32-bit
i needs a separate address computation while a load of b[i] at the same
address does not:
before after
add x1, x0, w1, uxtw prfm PLDL1KEEP, [x0, w1, uxtw]
prfm PLDL1KEEP, [x1]
Handle PREFETCH alongside MEM.
zstd 1.5.7 prefetches its match candidate through a 32-bit index in
ZSTD_RowFindBestMatch, which is 85% of the work at compression level 9.
Compressing the Silesia corpus on Grace with -mcpu=grace -O3, this saves
~1.65% dynamic instructions.
Bootstrapped and tested on aarch64-none-linux-gnu and x86_64-linux.
Ok for trunk?
Thanks,
Kyrill
gcc/ChangeLog:
* combine.cc (combine_simplify_rtx): Handle PREFETCH like MEM when
simplifying the address.
gcc/testsuite/ChangeLog:
* gcc.target/aarch64/prefetch-extend-1.c: New test.
Signed-off-by: Kyrylo Tkachov <[email protected]>
---
gcc/combine.cc | 4 ++++
.../gcc.target/aarch64/prefetch-extend-1.c | 20 +++++++++++++++++++
2 files changed, 24 insertions(+)
create mode 100644 gcc/testsuite/gcc.target/aarch64/prefetch-extend-1.c
diff --git a/gcc/combine.cc b/gcc/combine.cc
index 743026317ce..35c6a529950 100644
--- a/gcc/combine.cc
+++ b/gcc/combine.cc
@@ -5954,6 +5954,10 @@ combine_simplify_rtx (rtx x, machine_mode op0_mode, bool in_dest, bool in_cond)
/* A little bit of algebraic simplification here. */
switch (code)
{
+ case PREFETCH:
+ /* A prefetch reaches memory through an address, and targets recognize
+ that address with the same predicates they use for a MEM, so it
+ needs the same treatment. */
case MEM:
/* Ensure that our address has any ASHIFTs converted to MULT in case
address-recognizing predicates are called later. */
diff --git a/gcc/testsuite/gcc.target/aarch64/prefetch-extend-1.c b/gcc/testsuite/gcc.target/aarch64/prefetch-extend-1.c
new file mode 100644
index 00000000000..0037dfd43c3
--- /dev/null
+++ b/gcc/testsuite/gcc.target/aarch64/prefetch-extend-1.c
@@ -0,0 +1,20 @@
+/* A prefetch reaches memory through an address that PRFM can hold in an
+ extended-register form, just like a load of the same address does. */
+/* { dg-do compile } */
+/* { dg-options "-O2" } */
+
+void
+pf_uxtw (const unsigned char *b, unsigned int i)
+{
+ __builtin_prefetch (b + i, 0, 3);
+}
+
+void
+pf_sxtw (const unsigned char *b, int i)
+{
+ __builtin_prefetch (b + i, 0, 3);
+}
+
+/* { dg-final { scan-assembler-times {prfm\tPLDL1KEEP, \[x[0-9]+, w[0-9]+, uxtw\]} 1 } } */
+/* { dg-final { scan-assembler-times {prfm\tPLDL1KEEP, \[x[0-9]+, w[0-9]+, sxtw\]} 1 } } */
+/* { dg-final { scan-assembler-not {add\tx[0-9]+, x[0-9]+, w[0-9]+, [us]xtw\n} } } */
--
2.50.1 (Apple Git-155)
lmpx.com only provides a reader for public news (NNTP) servers. It is not
affiliated with the servers or forums shown here and is not responsible for
the content of articles, which is written by their respective authors.