RE: [RFC 01/32] bpf: replace deprecated SMP barriers with C11 fences
Konstantin Ananyev <[email protected]> Thu, 30 Jul 2026 09:23:35 +0000
| Newsgroups | org.dpdk.dev |
|---|---|
| Message-ID | <[email protected]> |
> The use counter handshake is a store-load pattern, so rte_smp_mb() > becomes a seq_cst thread fence; the read barrier in > bpf_eth_cbi_unuse() becomes an acquire fence. Same code generated > on x86 and arm64. > > Use relaxed loads and stores for the counter itself. With > enable_stdatomic, the plain increment of an RTE_ATOMIC() field > compiled to a seq_cst add, i.e. two locked operations per burst. > > Signed-off-by: Stephen Hemminger <[email protected]> > --- > lib/bpf/bpf_pkt.c | 18 +++++++++++------- > 1 file changed, 11 insertions(+), 7 deletions(-) > > diff --git a/lib/bpf/bpf_pkt.c b/lib/bpf/bpf_pkt.c > index f072fdaaed..a831b5ad86 100644 > --- a/lib/bpf/bpf_pkt.c > +++ b/lib/bpf/bpf_pkt.c > @@ -80,9 +80,11 @@ static struct bpf_eth_cbh tx_cbh = { > static __rte_always_inline void > bpf_eth_cbi_inuse(struct bpf_eth_cbi *cbi) > { > - cbi->use++; > + rte_atomic_store_explicit(&cbi->use, > + rte_atomic_load_explicit(&cbi->use, rte_memory_order_relaxed) > + 1, > + rte_memory_order_relaxed); > /* make sure no store/load reordering could happen */ > - rte_smp_mb(); > + rte_atomic_thread_fence(rte_memory_order_seq_cst); > } > > /* > @@ -92,8 +94,10 @@ static __rte_always_inline void > bpf_eth_cbi_unuse(struct bpf_eth_cbi *cbi) > { > /* make sure all previous loads are completed */ > - rte_smp_rmb(); > - cbi->use++; > + rte_atomic_thread_fence(rte_memory_order_acquire); Probably safer to use 'acq_release' order here? We do want that all previous loads to be completed at that point. > + rte_atomic_store_explicit(&cbi->use, > + rte_atomic_load_explicit(&cbi->use, rte_memory_order_relaxed) > + 1, > + rte_memory_order_relaxed); > } > > /* > @@ -105,9 +109,9 @@ bpf_eth_cbi_wait(const struct bpf_eth_cbi *cbi) > uint32_t puse; > > /* make sure all previous loads and stores are completed */ > - rte_smp_mb(); > + rte_atomic_thread_fence(rte_memory_order_seq_cst); > > - puse = cbi->use; > + puse = rte_atomic_load_explicit(&cbi->use, rte_memory_order_relaxed); > > /* in use, busy wait till current RX/TX iteration is finished */ > if ((puse & BPF_ETH_CBI_INUSE) != 0) { > @@ -439,7 +443,7 @@ bpf_eth_cbi_unload(struct bpf_eth_cbi *bc) > { > /* mark this cbi as empty */ > bc->cb = NULL; > - rte_smp_mb(); > + rte_atomic_thread_fence(rte_memory_order_seq_cst); > > /* make sure datapath doesn't use bpf anymore, then destroy bpf */ > bpf_eth_cbi_wait(bc); > -- > 2.53.0