RE: [PATCH] eal/x86: optimize memcpy of small 64-byte blocks

Konstantin Ananyev <[email protected]>
Newsgroups org.dpdk.dev
Message-ID <[email protected]>

> -----Original Message-----
> From: Morten Brørup <[email protected]>
> Sent: Wednesday, August 5, 2026 6:56 AM
> To: Konstantin Ananyev <[email protected]>; Stephen
> Hemminger <[email protected]>
> Cc: [email protected]; Bruce Richardson <[email protected]>
> Subject: RE: [PATCH] eal/x86: optimize memcpy of small 64-byte blocks
> 
> > From: Konstantin Ananyev [mailto:[email protected]]
> > Sent: Wednesday, 5 August 2026 07.46
> >
> > > > From: Stephen Hemminger [mailto:[email protected]]
> > > > Sent: Tuesday, 4 August 2026 17.52
> > > >
> > > > On Tue,  4 Aug 2026 14:33:04 +0000
> > > > Morten Brørup <[email protected]> wrote:
> > > >
> > > > > +	/* Common way for small copy size of 64-byte blocks.
> > Unlikely, so
> > > > constant size only */
> > > > > +	if (__rte_constant(n) && (n & 63) == 0 && n <=
> > > > RTE_MEMCPY_BLOCK_64_MAX) {
> > > > > +		void *ret = dst;
> > > > > +
> > > >
> > > > Maybe just let compiler decide, it will generate vector
> > instructions in
> > > > most cases.
> > > >
> > > > 	if (__rte_constant(n))
> > > > 		return mempcpy(dst, src, n);
> > >
> > > Maybe in most, but not in all:
> > > https://godbolt.org/z/KvdKqT5rY
> >
> > With '-mavx' or '-mavx512f' it looks like it does for your sample code.
> 
> It also does with -msse4.2 when SZ is reduced to 256 bytes.
> Clang switches to inline when SZ is reduced to 128 bytes.
> 
> It seems the compiler has a threshold for when to inline and when to call the C
> library's memcpy subroutine.
> The threshold depends on both copy size and vector register size.
> And it is compiler dependent.

I think there are compiler options to specify desired threshold values.
Let say for gcc there is  ' -mmemcpy-strategy=strategy'.
For that example in that particular case
-mmemcpy-strategy=vector_loop:512:align,loop:-1:align
generates sse loads/stores.
Might be we can exploit it somehow?
I am not really happy that our home-brewed memcpy code-block keeps growing,
while we keep talking that it would be good to eliminate it completely.




> With rte_memcpy() it is always inline.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.