Re: Fw: Random system freezes with Linux 6.12 + Xenomai 3.3.1 on Intel Ultra 5 125H

Jan Kiszka <[email protected]> Wed, 29 Jul 2026 07:30:10 +0200
Newsgroups dev.linux.lists.xenomai
Message-ID <[email protected]>
On 29.07.26 01:38, 刘杨 wrote:
> Hi Xenomai team,
> I am using Linux 6.12 with Xenomai 3.3.1 on an Intel Ultra 5 125H platform. The kernel randomly freezes, causing the entire system to hang. Below is my environment.
> • Platform: Intel Ultra 5 125H + NVIDIA 4060 GPU
> • Linux version: 6.12.55
> • Xenomai version: 3.3.1

Already tried to update to latest versions? Also if varying the major
kernel version makes any difference?

> • Linux grub parameters:
> isolcpus=0,1 nohz=0,1 nohz_full=0,1 rcu_nocbs=0,1 rcu_nocb_poll acpi_irq_nobalance noirqbalance processor.max_cstate=0 processor_idle.max_cstate=0 intel_idle.max_cstate=0 clocksource=tsc tsc=reliable hpet=disable nmi_watchdog=0 nowatchdog nosoftlockup softlockup_panic=0 selinux=0 enforcing=0 intel_pstate=disable idle=poll cpuidle.off=1 nohalt nosmap pcie_port_pm=off pcie_aspm.policy=performance ahci.mobile_lpm_policy=1 mce=off numa_balancing=disable apparmor=1 mem_sleep_default=deep console=tty console=ttyS0,115200n8 splash ibt=off usbcore.usbfs_memory_mb=512 vt.handoff=7
> • Use case: Humanoid robot
> • Applications: ROS2, CUDA 13

Okay... CUDA. Implies you are using an Nvidia driver - which one, binary
or OSS?

> The following is the log captured via serial port before the system froze:
>     fiveages@W2-0000:~$ [34822.953313] i915 0000:00:02.0: [drm] *ERROR* GT0: GUC: TLB invalidation response timed out for seqno 68639
> [34824.808351] xhci_hcd 0000:00:0d.0: xHCI host controller not responding, assume dead
> [34824.816041] xhci_hcd 0000:00:0d.0: HC died; cleaning up
> [34825.000349] i915 0000:00:02.0: [drm] *ERROR* GT0: GUC: TLB invalidation response timed out for seqno 68640
> [34870.898076] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
> [34870.904192] rcu: 7-...0: (1 GPs behind) idle=2684/1/0x4000000000000000 softirq=12697480/12697481 fqs=15002
> [34870.913954] rcu: (detected by 1, t=60017 jiffies, g=24464761, q=13288 ncpus=12)

What is the systems's workload at this point? Is there possibly some RT
thread running 100% of the time on some core? Do you have
CONFIG_XENO_OPT_WATCHDOG enabled?

Also, can you only reproduce this when the CUDA stack is in use?

Jan

-- 
Siemens AG, Foundational Technologies
Linux Expert Center