Re: [REGRESSION] 6.12.36+: usb: hub: post-resume delayed work triggers > uncorrected MCE / data fabric sync flood on Threadripper 7970X > (bisected to aec11e5f9c45)

Mario Limonciello <[email protected]>
Newsgroups dev.linux.lists.regressions,org.kernel.vger.linux-kernel,org.kernel.vger.linux-usb,org.kernel.vger.stable
Message-ID <[email protected]>

On 8/24/26 14:37, Mathieu Fluhr wrote:
>> The issue is obviously a severe HW malfunction (you mentioned MCEs, the
>> Ryzen CPUs simply totally locked up), triggered by poking certain xHCI
>> controllers on the I/O die of these CPUs in some wrong way.
> 
> Yes. As mentioned, I first thought that the emulator itself triggered that by
> doing something that the CPU did not like.To be honest, I barely play with old
> Android versions anymore, but seeing that I could reproduce it even with
> Android 13 or 14 made me suspicious.
> 
> I _guess_ Google implemented a workaround inside adb for version Android
> 15 since using this version, it remains stable for more than 2 hours.
> 
> But, in the end, the situation is that from a simple user account having access
> to some usb plugged in devices (I usually add my user account to the plugdev
> group and use some known udev rules to access my Android tests devices),
> you have a way to crash the complete system.
> 
>> Opinions seem to vary on whether CPU load must be present or absent.
> 
> On my side (and I am here only speaking about my TR. I don't know about
> other Ryzen CPUs), I can't reproduce it under load, and one condition to
> reproduce it is my CPU going in C2 state.
> -> I did a 2:30 hour test using several youtube videos playing at the same
> time on my desktop, also stressing the emulator with some 3D Mark runs
> (as mentioned, I first suspected the nvidia driver to be faulty). As long as my
> computer was busy everything went fine. But then I let it stand still for a few
> minutes, and it just crashed.
> 
> If you need me to do some further tests or experiments, let me know. I will
> be more than happy to play the guinea pig here.
> 
The behavior that is described here sounds like a platform firmware bug 
to me.  Are you on the latest BIOS available from your OEM?

Can you please confirm:

1. Your CPU model number/codename
2. Your OEM (from /sys/class/dmi/id)
3. OEM BIOS version (from /sys/class/dmi/id)
4. AGESA version (see commit bc91133e260c8113c1119073c03b93c12aa41738 if 
the OEM didn't tear it out.  Otherwise look in BIOS menus)?
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.