Re: [RFC PATCH 1/1] hpref: Hazard Pointers with Reference Counter

Mathieu Desnoyers <[email protected]> Wed, 7 Jan 2026 15:04:11 -0500
Newsgroups dev.linux.lists.lkmm
Message-ID <[email protected]>
On 2026-01-07 12:48, Paul E. McKenney wrote:
> On Wed, Jan 07, 2026 at 09:40:56AM -0500, Mathieu Desnoyers wrote:
[...]
> 
>> * RCU grace period guarantee:
>>
>> For the read-side critical section (rcu_read_{lock,unlock}) pairing with
>> synchronize_rcu, we've updated the liburcu read lock/unlock
>> implementation to use either atomic thread fence as full barrier or
>> SEQ_CST stores. We've also introduced a urcu/annotate.h header to help
>> annotating memory access groups. This helps threadsanitizer track the
>> dependency ordering inherent to RCU grace periods.
> 
> The atomic thread fence is for threadsanitizer's benefit?  In the MB
> rcu_read_lock() and rcu_read_unlock() case, is barrier() or a light-weight
> asymmetric fence still used?

For urcu-memb and urcu-bp, we use a lightweight compiler cmm_barrier()
in rcu_read_{lock,unlock}() paired with membarrier(2) within
synchronize_rcu() when membarrier is available, else we fallback
on cmm_smp_mb() everywhere.

For urcu-mb and urcu-qsbr, we previously (before 0.15) had the more
heavyweight cmm_smp_mb(). When compiling with the atomic builtins
configure flag, this just maps to atomic thread fence. AFAIK
TSAN does not model atomic thread fence, so we need to rely on
explicit TSAN annotations of happens-before relationships of
relaxed memory accesses ordered by atomic thread fence (or
barrier+membarrier pairing trick) in those scenarios.

Now that I review the urcu-mb and urcu-qsbr 0.15 code again, I suspect
we're missing ordering between the seq-cst store to rcu_reader.ctr
and relaxed load of the futex state. I have prepared the following
commits to address this:

https://review.lttng.org/c/userspace-rcu/+/16151 urcu-mb: Use CMM_SEQ_CST_FENCE for _urcu_mb_read_unlock_update_and_wakeup
https://review.lttng.org/c/userspace-rcu/+/16152 urcu-qsbr: Use CMM_SEQ_CST_FENCE for quiescent state update and offline

We're also missing a TSAN acquire annotation for the urcu_mb_gp.ctr load
in _urcu_mb_read_lock_update. This is only relevant for TSAN analyses:

https://review.lttng.org/c/userspace-rcu/+/16156 urcu-mb: Add missing TSAN annotation to _urcu_mb_read_lock_update

And for the sake of combining the store+cmm_smp_mb() pattern into
a CMM_SEQ_CST_FENCE which can be optimized by architectures,
the following commits:

https://review.lttng.org/c/userspace-rcu/+/16153 urcu-mb: Use CMM_SEQ_CST_FENCE for _urcu_mb_read_lock_update
https://review.lttng.org/c/userspace-rcu/+/16154 urcu-qsbr: Use CMM_SEQ_CST_FENCE for _urcu_qsbr_thread_online

Thoughts ?

Thanks,

Mathieu

-- 
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com