Re: [PATCH v14 05/15] arm64: rqspinlock: Remove private copy of smp_cond_load_acquire_timewait()

Ankur Arora <[email protected]> Thu, 30 Jul 2026 16:31:55 -0700
Newsgroups org.kernel.vger.linux-arch,dev.linux.lists.sashiko-reviews,org.infradead.lists.linux-arm-kernel,org.kernel.vger.bpf,org.kernel.vger.linux-kernel,org.kernel.vger.linux-pm
Message-ID <[email protected]>
[email protected] writes:

> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] Removing the ARM64-specific `smp_cond_load_acquire_timewait()` fallback before replacing it with a working alternative breaks BPF `rqspinlock` timeout functionality, causing indefinite hangs on platforms without a timer event stream.

A private copy of this interface was committed to the tree prematurely.

Removal (while temporarily breaking bisection) was acked by the author
here:
  https://lore.kernel.org/lkml/CAP01T75VaNGk+adULuX3q_8uxc0z78J5B_5ECLy2VR=i7pBi0g@mail.gmail.com/

Ankur

>
> --- Patch [5]: [PATCH v14 05/15] arm64: rqspinlock: Remove private copy of smp_cond_load_acquire_timewait() ---
> commit e150f6dc634a9b6af6666a67d9bc5ea262bedad7
> Author: Ankur Arora <[email protected]>
>
> arm64: rqspinlock: Remove private copy of smp_cond_load_acquire_timewait()
>
> This commit removes the private definition of smp_cond_load_acquire_timewait()
> and its helper macros from arch/arm64/include/asm/rqspinlock.h. The commit
> message notes this is in preparation for defining
> smp_cond_load_acquire_timeout(), allowing rqspinlock to fall back to using
> smp_cond_load_acquire() in the interim.
>
>> diff --git a/arch/arm64/include/asm/rqspinlock.h b/arch/arm64/include/asm/rqspinlock.h
>> index 9ea0a74e58927..a385603436e98 100644
>> --- a/arch/arm64/include/asm/rqspinlock.h
>> +++ b/arch/arm64/include/asm/rqspinlock.h
> [ ... ]
>> -#define smp_cond_load_acquire_timewait(ptr, cond_expr,			\
>> -				      time_expr_ns, time_limit_ns)	\
>> -({									\
>> -	__unqual_scalar_typeof(*ptr) _val;				\
>> -	int __wfe = arch_timer_evtstrm_available();			\
>> -									\
>> -	if (likely(__wfe)) {						\
>> -		_val = __smp_cond_load_acquire_timewait(ptr, cond_expr,	\
>> -							time_expr_ns,	\
>> -							time_limit_ns);	\
>> -	} else {							\
>> -		_val = __smp_cond_load_relaxed_spinwait(ptr, cond_expr,	\
>> -							time_expr_ns,	\
>> -							time_limit_ns);	\
>> -		smp_acquire__after_ctrl_dep();				\
>> -	}								\
>> -	(typeof(*ptr))_val;						\
>> -})
>> -
>> -#endif
>> -
>> -#define res_smp_cond_load_acquire(v, c) smp_cond_load_acquire_timewait(v, c, 0, 1)
>
> [Severity: High]
> With this custom fallback removed, the BPF rqspinlock code will use the generic
> smp_cond_load_acquire(). Does this mean on platforms without a timer event
> stream (arch_timer_evtstrm_available() is false), a CPU could get permanently
> stuck in a WFE instruction during a lock timeout?
>
> In kernel/bpf/rqspinlock.c:resilient_queued_spin_lock_slowpath(), the timeout is
> evaluated as part of the condition expression:
>
>     res_smp_cond_load_acquire(&lock->locked, !VAL || RES_CHECK_TIMEOUT(ts, ret, _Q_LOCKED_MASK));
>
> But smp_cond_load_acquire() calls __cmpwait_relaxed() which executes WFE:
>
>     __cmpwait_relaxed(__PTR, VAL);
>
> If no stores occur to the locked byte, WFE will wait indefinitely, and the
> embedded timeout check won't be evaluated while the CPU is asleep.