Re: Hang during smokey posix_cond test w/ cobalt and PREEMPT_RT enabled
Jan Kiszka <[email protected]> Tue, 21 Jul 2026 10:03:03 +0200
| Newsgroups | dev.linux.lists.xenomai |
|---|---|
| Message-ID | <[email protected]> |
On 21.07.26 09:21, Dominic Rath wrote: > Dear Xenomai team, > > we're seeing a hang during the execution of the posix_cond test in > smokey. When breaking into the stuck test, I can see that the main > thread is waiting on a semaphore in thread_spawn, and the newly spawned > thread appears to be stuck in cobalt_thread_harden, before it had a > chance to post the semaphore. > > Apart from that single test, all other tests run successfully, and our > application was running fine, too. > > We're using Xenomai 3.3.3 on an ARM64 running Linux kernel 6.12 with > PREEMPT_RT enabled. The kernel is a heavily patched vendor tree based on > v6.12.y-cip merged with v6.12.y-cip-dovetail, and I can't rule out that > we've made some mistake there. > > When running the tests without PREEMPT_RT enabled, the tests complete > successfully. > > Looking at some tracepoints and function traces we noticed that the > newly spawned thread actually wanted to relax again, but instead the > kernel is in an endless loop calling lostage_task_wakeup on the main > thread. > > What happened is that the main thread wanted to relax, hardened again, > and called xnthread_relax again. On that second relax the > relax_work.inband_work was queued via irq_work_queue again, while it was > already on the lazy_list, resulting in the endless loop. > > It seems that the main thread is continuing execution in the Linux > domain after the first call to xnthread_relax without > lostage_task_wakeup running first. That gives it a chance to harden and > call relax the second time, where the endless loop on the lazy_list gets > created. Only afterwards is irq_work_run_list executing again and calls > cobalt_lostage_wakeup, eventually in an endless loop. > > This is possible, because relax_work.inband_work gets completely > initialized on each call to xnthread_relax via IRQ_WORK_INIT. > This causes the check for pending irq_work in irq_work_queue to not > detect that this item was already queued. > > I was able to fix this by moving the initialization of thread-> > relax_work.inband_work to __xnthread_init. With that changed, the test > successfully completes even with PREEMPT_RT enabled. > > I believe the reason why this works without PREEMPT_RT, or when using > IRQ_WORK_INIT_HARD (tried that, too), is that the irq_work is called > immediately when the core is back in secondary mode, whereas with > PREEMPT_RT it is deferred (to the next tick?). > > What I'm not so sure about is the following: > > - is running Xenomai 3 with cobalt on a kernel with PREEMPT_RT > supported, i.e. should this work? > Conceptually, yes. Practically, this combination is very likely undertested (e.g. not part of our CI). > - is it conceivable that the main thread continues execution in the > Linux domain before lostage_task_wakeup got called? Or does this > indicate that there's something else wrong with our setup? > > I've enabled all of the Xenomai, dovetail and PREEMPT_RT debug checks > that made sense to me and I'm not getting any errors, but of course I > might have missed some useful CONFIG option. > Did you already compare the event flow you are seeing with a reference flow you could generated from qemu-arm64 image of xenomai-images? That may help looking for the essential differences. Furthermore, you could patch xenomai-image to turn on PREEMPT_RT for 6.12 on qemu-arm64 as well and possibly run your test in that setup (provided its portable enough). Jan -- Siemens AG, Foundational Technologies Linux Expert Center