Re: [RFC 12/32] stack: always use C11 memory model implementation
Stephen Hemminger <[email protected]> Sat, 1 Aug 2026 10:01:50 -0700
| Newsgroups | org.dpdk.dev |
|---|---|
| Message-ID | <[email protected]> |
On Fri, 31 Jul 2026 16:53:45 +0200 Morten Br=C3=B8rup <[email protected]> wrote: > +TO: x86 maintainers, ThunderX maintainers >=20 > > From: Stephen Hemminger [mailto:[email protected]] > > Sent: Wednesday, 29 July 2026 19.54 > >=20 > > The generic and C11 lock-free stack implementations differ only in > > memory ordering. The generic version uses a full barrier where its > > own comments state an acquire fence is sufficient, and seq_cst for > > all length counter operations. > >=20 > > Only x86 and ThunderX still used the generic version. On x86 the > > switch removes a locked add per CAS attempt in push and pop; TSO > > provides the acquire semantics. On ThunderX the pop fence weakens > > from dmb ish to dmb ishld and the push fence goes away. Unlike the > > ring, no platform selected the generic stack for measured > > performance reasons. > >=20 > > Remove it and use the C11 implementation everywhere. =20 >=20 > The lack of measured performance difference documentation is not a valid = reason to remove the generic version! >=20 > It would be reasonable to assume that x86 (and ThunderX) use the generic = version for non-insignificant performance reasons. >=20 > If there is no performance difference, I agree with this patch. Otherwise= not. > This could be verified by providing the missing measurements. >=20 Surprisingly, the performance of the C11 version is better than the old gen= eric version that had smp_mb. That is because C11 code generates no locked pref= ixes. Gets speedup of upto 60%. Between main (with rte_smp_mb) and the unified C11 version on the 32-core x= 86 machine: Test main (n=3D9) unified C11 (n=3D9) delta single push/pop 46.62 =C2=B10.30 33.41 =C2=B10.10 -28% empty pop 1.47 =C2=B10.01 0.98 =C2=B10.01 -33% 1 lcore, bulk 8 9.06 =C2=B10.05 8.20 =C2=B10.08 -10% 1 lcore, bulk 32 6.09 =C2=B10.02 6.15 =C2=B10.03 +1% 2 HT, bulk 8 42.05 =C2=B10.31 39.24 =C2=B10.52 -7% 2 HT, bulk 32 11.92 =C2=B10.13 11.89 =C2=B10.10 0 2 cores, bulk 8 78.90 =C2=B10.60 72.96 =C2=B11.11 -7% 2 cores, bulk 32 20.74 =C2=B11.56 7.70 =C2=B10.13 -63% 32 cores, bulk 8 6126 =C2=B172 6121 =C2=B189 0 32 cores, bulk 32 1953.9 =C2=B12.9 1984.6 =C2=B113.3 +1.6%