Re: [RFC 12/32] stack: always use C11 memory model implementation

Stephen Hemminger <[email protected]> Sat, 1 Aug 2026 10:01:50 -0700
Newsgroups org.dpdk.dev
Message-ID <[email protected]>
On Fri, 31 Jul 2026 16:53:45 +0200
Morten Br=C3=B8rup <[email protected]> wrote:

> +TO: x86 maintainers, ThunderX maintainers
>=20
> > From: Stephen Hemminger [mailto:[email protected]]
> > Sent: Wednesday, 29 July 2026 19.54
> >=20
> > The generic and C11 lock-free stack implementations differ only in
> > memory ordering. The generic version uses a full barrier where its
> > own comments state an acquire fence is sufficient, and seq_cst for
> > all length counter operations.
> >=20
> > Only x86 and ThunderX still used the generic version. On x86 the
> > switch removes a locked add per CAS attempt in push and pop; TSO
> > provides the acquire semantics. On ThunderX the pop fence weakens
> > from dmb ish to dmb ishld and the push fence goes away. Unlike the
> > ring, no platform selected the generic stack for measured
> > performance reasons.
> >=20
> > Remove it and use the C11 implementation everywhere. =20
>=20
> The lack of measured performance difference documentation is not a valid =
reason to remove the generic version!
>=20
> It would be reasonable to assume that x86 (and ThunderX) use the generic =
version for non-insignificant performance reasons.
>=20
> If there is no performance difference, I agree with this patch. Otherwise=
 not.
> This could be verified by providing the missing measurements.
>=20

Surprisingly, the performance of the C11 version is better than the old gen=
eric
version that had smp_mb.  That is because C11 code generates no locked pref=
ixes.
Gets speedup of upto 60%.

Between main (with rte_smp_mb) and the unified C11 version on the 32-core x=
86 machine:

Test	main (n=3D9)	unified C11 (n=3D9)	delta
single push/pop	46.62 =C2=B10.30	33.41 =C2=B10.10	-28%
empty pop	1.47 =C2=B10.01	0.98 =C2=B10.01	-33%
1 lcore, bulk 8	9.06 =C2=B10.05	8.20 =C2=B10.08	-10%
1 lcore, bulk 32	6.09 =C2=B10.02	6.15 =C2=B10.03	+1%
2 HT, bulk 8	42.05 =C2=B10.31	39.24 =C2=B10.52	-7%
2 HT, bulk 32	11.92 =C2=B10.13	11.89 =C2=B10.10	0
2 cores, bulk 8	78.90 =C2=B10.60	72.96 =C2=B11.11	-7%
2 cores, bulk 32	20.74 =C2=B11.56	7.70 =C2=B10.13	-63%
32 cores, bulk 8	6126 =C2=B172	6121 =C2=B189	0
32 cores, bulk 32	1953.9 =C2=B12.9	1984.6 =C2=B113.3	+1.6%