[PATCH v2 0/6] RISC-V: Optimize memmove() for speed
m fally <[email protected]>
| Newsgroups | gmane.comp.lib.newlib |
|---|---|
| Message-ID | <[email protected]> |
This is version 2 of a patch series that optimizes the RISC-V port of
memmove() for speed. The implementation is based on the generic port of
the function, since that is what is currently used when compiling newlib
for RISC-V.
Changes in v2 include redirection to memcpy() if the memory areas of
source and destination do not overlap, as well as small corrections of
comments in the code.
Link to v1: https://sourceware.org/pipermail/newlib/2025/021810.html
In the stock implementation, an unroll-factor of 4 is used for
the word-copy-loop in the case where both source and destination
addresses are aligned on a long-boundary, and the memory areas
overlap non-destructively or not at all. No unrolling is done in
the destructive-overlap case. The proposed implementation uses an
unroll-factor of 9 for both overlap-cases when both addresses are
aligned to xlen. The unroll-factor was chosen to match memcpy() and
speeds up the copying-process for lengths >= 9*SZREG, while almost
not at all degrading performance for shorter lengths.
If at least one address is unaligned, misaligned accesses are slow
or prohibited, and there are >= 2*SZREG bytes left to copy, the
proposed implementation first aligns the source address. Then, one
whole word (or doubleword for rv64) is loaded at a time and individual
bytes are stored back to the destination. The threshold of 2*SZREG was
chosen in order to keep the negative effect on shorter copies caused
by the additional overhead of the alignment operation low.
If there is no overlap between the two memory regions, the function
redirects to memcpy(). This is only done if length > SZREG in order
to reduce overhead on very short copies.
Furthermore, the function now only uses fixed-width types.
Macros from the generic port are replaced with RISC-V-specific macros
and static inline functions.
The proposed implementation was tested on spike with pk for each of the
following configurations (compiled with gcc):
rv32ic -mtune=thead-c906 -mstrict-align -O3
rv32ic -mtune=thead-c906 -mno-strict-align -O3
rv64ic -mtune=thead-c906 -mstrict-align -O3
rv64ic -mtune=thead-c906 -mno-strict-align -O3
For each configuration, the following cases were considered when comparing
the old and new implementations:
both addresses are xlen-aligned
both addresses are unaligned and have the same alignment
source address is xlen-aligned, destination address is not
destination address is xlen-aligned, source address is not
For each configuration, 98280 tests were run. In total, there were 13027 cases
where the new implementation was slower than the original, with a maximum
difference of 342 retired instructions. In 324 cases the implementations were
equally fast. In all other cases, the new implementation was faster than the
original.
Please see here for graphical comparisons between the two implementations:
https://cloud.servus.at/s/jnHnN9gEoNMkLLc
Below tables show the number of cases where the new implementation was
faster, slower, or equally fast as the current implementation, as well as
the max. differences in instructions retired between the implementations.
The graphs and tables show the differences for copied lengths between
2 and 8191 bytes. For lengths < 2, old and new implementations were
equally fast.
--------- BENCHMARKING RESULTS ---------
+--------------------------------------+
| rv32ic -mstrict-align |
+--------------------------------------+
| |Number |Max. difference in |
| |of cases |instructions retired |
+------+---------+---------------------+
|Faster| 98082 | -35899 |
+------+---------+---------------------+
|Slower| 172 | 24 |
+------+---------+---------------------+
|Tied | 26 | |
+------+---------+---------------------+
+--------------------------------------+
| rv64ic -mstrict-align |
+--------------------------------------+
| |Number |Max. difference in |
| |of cases |instructions retired |
+------+---------+---------------------+
|Faster| 98026 | -38382 |
+------+---------+---------------------+
|Slower| 229 | 30 |
+------+---------+---------------------+
|Tied | 25 | |
+------+---------+---------------------+
+--------------------------------------+
| rv32ic -mno-strict-align |
+--------------------------------------+
| |Number |Max. difference in |
| |of cases |instructions retired |
+------+---------+---------------------+
|Faster| 95568 | -5231 |
+------+---------+---------------------+
|Slower| 2626 | 160 |
+------+---------+---------------------+
|Tied | 86 | |
+------+---------+---------------------+
+--------------------------------------+
| rv64ic -mno-strict-align |
+--------------------------------------+
| |Number |Max. difference in |
| |of cases |instructions retired |
+------+---------+---------------------+
|Faster| 88093 | -2609 |
+------+---------+---------------------+
|Slower| 10000 | 342 |
+------+---------+---------------------+
|Tied | 187 | |
+------+---------+---------------------+
m fally (6):
RISC-V: memmove() speed optimized: Add implementation
RISC-V: memmove() speed optimized: Replace macros and use fixed-width
types
RISC-V: memmove() speed optimized: Add loop-unrolling
RISC-V: memmove() speed optimized: Align source address
RISC-V: memmove() speed optimized: Call memcpy()
newlib: Regenerate configuration files
newlib/Makefile.in | 52 ++--
newlib/libc/machine/riscv/Makefile.inc | 2 +-
.../riscv/{memmove.S => memmove-asm.S} | 0
newlib/libc/machine/riscv/memmove-stub.c | 14 -
newlib/libc/machine/riscv/memmove.c | 259 ++++++++++++++++++
5 files changed, 286 insertions(+), 41 deletions(-)
rename newlib/libc/machine/riscv/{memmove.S => memmove-asm.S} (100%)
delete mode 100644 newlib/libc/machine/riscv/memmove-stub.c
create mode 100644 newlib/libc/machine/riscv/memmove.c
--
2.49.0