Re: [PATCH v2 1/2] futex/requeue: Fix rtmutex schedule preparation for requeue PI
Yao Kai <[email protected]>
| Newsgroups | org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On 8/5/2026 4:00 PM, Yao Kai wrote:
>
>
> On 8/4/2026 8:21 PM, Peter Zijlstra wrote:
>> On Wed, Jul 22, 2026 at 04:51:39PM +0800, Yao Kai wrote:
>>> A waiter requeued onto a PI futex can reach rt_mutex_wait_proxy_lock()
>>> without rtmutex schedule preparation:
>>>
>>> WARNING: CPU: 0 PID: 293 at kernel/sched/core.c:7606
>>> RIP: rt_mutex_schedule+0x43/0x50
>>> Call Trace:
>>> rt_mutex_slowlock_block.constprop.0+0x5b/0x320
>>> rt_mutex_wait_proxy_lock+0x3e/0x80
>>> futex_wait_requeue_pi+0x3ba/0x590
>>> do_futex+0x171/0x1f0
>>>
>>> rt_mutex_schedule() requires current->sched_rt_mutex to be set. Normally,
>>> rt_mutex_pre_schedule() sets it before an rtmutex waiter can schedule. With
>>> requeue PI, another task can enqueue the waiter after its futex_q becomes
>>> visible:
>>>
>>> waiter requeue task
>>> ------ ------------
>>> futex_wait_requeue_pi()
>>> futex_wait_setup()
>>> futex_queue(&q)
>>> futex_requeue()
>>> rt_mutex_start_proxy_lock()
>>> enqueue rt_waiter
>>> install pi_blocked_on
>>> requeue_futex()
>>> plist_del(&q->list)
>>> futex_do_wait()
>>> plist_node_empty(&q->list)
>>> skip schedule()
>>> plist_add(&q->list)
>>> futex_requeue_pi_complete()
>>> IN_PROGRESS -> DONE
>>> futex_requeue_pi_wakeup_sync() // DONE
>>> rt_mutex_wait_proxy_lock()
>>> rt_mutex_schedule()
>>>
>>> futex_do_wait() mistakes the temporary removal for a wakeup and skips
>>> schedule(). The proxy waiter can nevertheless remain blocked on the target
>>> rtmutex and subsequently enter rt_mutex_schedule() with
>>> current->sched_rt_mutex clear.
>>>
>>> Call rt_mutex_pre_schedule() and rt_mutex_post_schedule() directly around
>>> rt_mutex_wait_proxy_lock() so this second blocking point has the required
>>> scheduler preparation.
>>>
>>> Fixes: d14f9e930b90 ("locking/rtmutex: Use rt_mutex specific scheduler helpers")
>>> Suggested-by: Sebastian Andrzej Siewior <[email protected]>
>>> Cc: [email protected]
>>> Signed-off-by: Yao Kai <[email protected]>
>>> ---
>>> kernel/futex/requeue.c | 8 ++++++++
>>> 1 file changed, 8 insertions(+)
>>>
>>> diff --git a/kernel/futex/requeue.c b/kernel/futex/requeue.c
>>> index 79823ad13683..f7889fb2fce4 100644
>>> --- a/kernel/futex/requeue.c
>>> +++ b/kernel/futex/requeue.c
>>> @@ -1,6 +1,7 @@
>>> // SPDX-License-Identifier: GPL-2.0-or-later
>>> #include <linux/plist.h>
>>> +#include <linux/sched/rt.h>
>>> #include <linux/sched/signal.h>
>>> #include "futex.h"
>>> @@ -865,7 +866,14 @@ int futex_wait_requeue_pi(u32 __user *uaddr, unsigned int flags,
>>> case Q_REQUEUE_PI_DONE:
>>> /* Requeue completed. Current is 'pi_blocked_on' the rtmutex */
>>> pi_mutex = &q.pi_state->pi_mutex;
>>> + /*
>>> + * Requeue temporarily removes q from the hash bucket, so
>>> + * futex_do_wait() may skip schedule() even though the proxy
>>> + * waiter still has to block on the rtmutex.
>>> + */
>>> + rt_mutex_pre_schedule();
>>> ret = rt_mutex_wait_proxy_lock(pi_mutex, to, &rt_waiter);
>>> + rt_mutex_post_schedule();
>>
>> As per always, I'm totally confused about everything. I mean, futexes
>> suck, but requeue sucks worse.
>>
>> So the purpose of rt_mutex_pre_schedule() was to avoid the double waiter
>> enqueue for rt_mutex on RT, where sched_submit_work() will hit a
>> spinlock-nee-rtlock.
>>
>> So rt_mutex_pre_schedule() must happen before the rt_mutex is added as a
>> waiter. However, AFAICT we're already a waiter at the above spot, no? So
>> this cannot be right.
>>
>> The changelogs doesn't at all explain why this is correct. Please help?
>
> You are right that rt_mutex_pre_schedule() is normally required before the
> rtmutex waiter is enqueued.
>
> The v2 placement relies on a property of this particular call path that I
> failed to explain in the changelog: futex_wait_requeue_pi() is reachable
> only via the FUTEX_WAIT_REQUEUE_PI syscall from userspace. Therefore,
> 'current' cannot be a workqueue/io-wq worker, nor can it carry a live
> block plug across the syscall boundary.
>
> Consequently, sched_submit_work() inside rt_mutex_pre_schedule() has no
> pending plugged I/O or worker notifications to process here, so it cannot
> recurse into an rtlock after the proxy waiter has been installed. In this
> specific path, calling rt_mutex_pre_schedule() here effectively only sets
> current->sched_rt_mutex before rt_mutex_wait_proxy_lock() calls
> rt_mutex_schedule(), without executing any blocking submit_work after enqueue.
>
> This was the rationale for adopting Sebastian's simpler suggestion from the
> v1 discussion, but this essential invariant was omitted from the v2
> changelog and code comments.
>
> Thanks,
> Yao Kai
>
>
Gentle ping. If the v2 approach is not acceptable, would the v1 approach be
preferable instead?
https://lore.kernel.org/r/[email protected]
Thanks,
Yao Kai