Re: [PATCH v2 0/6] RISC-V: Optimize memmove() for speed

Kito Cheng <[email protected]>
Newsgroups gmane.comp.lib.newlib
Message-ID <CA+yXCZABFgYz0hvn_6Y5j+p5nE8Lba1pR0b-=6LK+yA2JoRXBQ@mail.gmail.com>
Thanks! committed to trunk after passing GCC regression :)

On Tue, Jun 17, 2025 at 11:10 PM m fally <[email protected]> wrote:
>
> This is version 2 of a patch series that  optimizes the RISC-V port of
> memmove() for speed. The implementation is based on the generic port of
> the function, since that is what is currently used when compiling newlib
> for RISC-V.
>
> Changes in v2 include redirection to memcpy() if the memory areas of
> source and destination do not overlap, as well as small corrections of
> comments in the code.
>
> Link to v1: https://sourceware.org/pipermail/newlib/2025/021810.html
>
> In the stock implementation, an unroll-factor of 4 is used for
> the word-copy-loop in the case where both source and destination
> addresses are aligned on a long-boundary, and the memory areas
> overlap non-destructively or not at all. No unrolling is done in
> the destructive-overlap case. The proposed implementation uses an
> unroll-factor of 9 for both overlap-cases when both addresses are
> aligned to xlen. The unroll-factor was chosen to match memcpy() and
> speeds up the copying-process for lengths >= 9*SZREG, while almost
> not at all degrading performance for shorter lengths.
>
> If at least one address is unaligned, misaligned accesses are slow
> or prohibited, and there are >= 2*SZREG bytes left to copy, the
> proposed implementation first aligns the source address. Then, one
> whole word (or doubleword for rv64) is loaded at a time and individual
> bytes are stored back to the destination. The threshold of 2*SZREG was
> chosen in order to keep the negative effect on shorter copies caused
> by the additional overhead of the alignment operation low.
>
> If there is no overlap between the two memory regions, the function
> redirects to memcpy(). This is only done if length > SZREG in order
> to reduce overhead on very short copies.
>
> Furthermore, the function now only uses fixed-width types.
> Macros from the generic port are replaced with RISC-V-specific macros
> and static inline functions.
>
>
> The proposed implementation was tested on spike with pk for each of the
> following configurations (compiled with gcc):
>
> rv32ic -mtune=thead-c906 -mstrict-align -O3
> rv32ic -mtune=thead-c906 -mno-strict-align -O3
> rv64ic -mtune=thead-c906 -mstrict-align -O3
> rv64ic -mtune=thead-c906 -mno-strict-align -O3
>
>
> For each configuration, the following cases were considered when comparing
> the old and new implementations:
>
> both addresses are xlen-aligned
> both addresses are unaligned and have the same alignment
> source address is xlen-aligned, destination address is not
> destination address is xlen-aligned, source address is not
>
> For each configuration, 98280 tests were run. In total, there were 13027 cases
> where the new implementation was slower than the original, with a maximum
> difference of 342 retired instructions. In 324 cases the implementations were
> equally fast. In all other cases, the new implementation was faster than the
> original.
>
> Please see here for graphical comparisons between the two implementations:
> https://cloud.servus.at/s/jnHnN9gEoNMkLLc
>
> Below tables show the number of cases where the new implementation was
> faster, slower, or equally fast as the current implementation, as well as
> the max. differences in instructions retired between the implementations.
> The graphs and tables show the differences for copied lengths between
> 2 and 8191 bytes. For lengths < 2, old and new implementations were
> equally fast.
>
> --------- BENCHMARKING RESULTS ---------
>
> +--------------------------------------+
> |        rv32ic -mstrict-align         |
> +--------------------------------------+
> |      |Number   |Max. difference in   |
> |      |of cases |instructions retired |
> +------+---------+---------------------+
> |Faster|   98082 |              -35899 |
> +------+---------+---------------------+
> |Slower|     172 |                  24 |
> +------+---------+---------------------+
> |Tied  |      26 |                     |
> +------+---------+---------------------+
>
> +--------------------------------------+
> |        rv64ic -mstrict-align         |
> +--------------------------------------+
> |      |Number   |Max. difference in   |
> |      |of cases |instructions retired |
> +------+---------+---------------------+
> |Faster|   98026 |              -38382 |
> +------+---------+---------------------+
> |Slower|     229 |                 30  |
> +------+---------+---------------------+
> |Tied  |      25 |                     |
> +------+---------+---------------------+
>
> +--------------------------------------+
> |        rv32ic -mno-strict-align      |
> +--------------------------------------+
> |      |Number   |Max. difference in   |
> |      |of cases |instructions retired |
> +------+---------+---------------------+
> |Faster|   95568 |               -5231 |
> +------+---------+---------------------+
> |Slower|    2626 |                 160 |
> +------+---------+---------------------+
> |Tied  |      86 |                     |
> +------+---------+---------------------+
>
> +--------------------------------------+
> |        rv64ic -mno-strict-align      |
> +--------------------------------------+
> |      |Number   |Max. difference in   |
> |      |of cases |instructions retired |
> +------+---------+---------------------+
> |Faster|   88093 |               -2609 |
> +------+---------+---------------------+
> |Slower|   10000 |                 342 |
> +------+---------+---------------------+
> |Tied  |     187 |                     |
> +------+---------+---------------------+
>
>
> m fally (6):
>   RISC-V: memmove() speed optimized: Add implementation
>   RISC-V: memmove() speed optimized: Replace macros and use fixed-width
>     types
>   RISC-V: memmove() speed optimized: Add loop-unrolling
>   RISC-V: memmove() speed optimized: Align source address
>   RISC-V: memmove() speed optimized: Call memcpy()
>   newlib: Regenerate configuration files
>
>  newlib/Makefile.in                            |  52 ++--
>  newlib/libc/machine/riscv/Makefile.inc        |   2 +-
>  .../riscv/{memmove.S => memmove-asm.S}        |   0
>  newlib/libc/machine/riscv/memmove-stub.c      |  14 -
>  newlib/libc/machine/riscv/memmove.c           | 259 ++++++++++++++++++
>  5 files changed, 286 insertions(+), 41 deletions(-)
>  rename newlib/libc/machine/riscv/{memmove.S => memmove-asm.S} (100%)
>  delete mode 100644 newlib/libc/machine/riscv/memmove-stub.c
>  create mode 100644 newlib/libc/machine/riscv/memmove.c
>
> --
> 2.49.0
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.