Re: [PATCH] arm64: smp: distinguish secondary CPUs that hang after reaching head.S
Naman Jain <[email protected]> Wed, 5 Aug 2026 10:12:03 +0530
| Newsgroups | gmane.linux.kernel,gmane.linux.ports.arm.kernel |
|---|---|
| Message-ID | <[email protected]> |
On 8/4/2026 7:13 PM, Will Deacon wrote: > On Mon, Jul 27, 2026 at 11:08:05AM +0530, Naman Jain wrote: >> On 7/26/2026 7:16 PM, Will Deacon wrote: >>> On Fri, Jul 24, 2026 at 10:16:03AM +0530, Naman Jain wrote: >>>> On 7/23/2026 11:26 AM, Anshuman Khandual wrote: >>>>> On 23/07/26 8:19 AM, Jinjie Ruan wrote: >>>>>> 在 2026/7/22 19:30, Naman Jain 写道: >>>>>>> When a secondary CPU fails to come online, __cpu_up() falls back to >>>>>>> __early_cpu_boot_status, but boot status 0x0 is ambiguous: it cannot >>>>>>> distinguish a CPU that never executed head.S (firmware/hypervisor never >>>>>>> dispatched it, so it never ran a single instruction) from one that >>>>>>> entered head.S, started executing, and then got stuck somewhere in kernel >>>>>>> bring-up. Add a change to let us tell those two cases apart, which >>>>>>> narrows down where to look when a CPU goes missing during boot. >>>>>> I previously encountered this issue when debugging the parallel startup >>>>>> of ARM64 secondary cores. It is difficult for the kernel to determine >>>>>> whether the secondary core is hung in the firmware or whether it has not >>>>>> executed a single instruction. So I think this motive is reasonable. >>>>> >>>>> Why should kernel determine the difference here ? Would not the firmware >>>>> know if it has started any secondary CPU for the kernel which must have >>>>> come inside head.S ? If the cpu gets hung inside firmware while starting >>>>> up then the debug responsibilities belong there instead. >>>>> >>>>> Still wondering what's the rationale for this change. >>>> >>>> Hello Anshuman, >>>> This sounds fair to me. Let me elaborate the problem, beyond the scope of >>>> this patch. In production, we occasionally see these crashes where one of >>>> the CPU fails to bring up online, with 0x0 status code. Hypervisor may be >>>> missing the telemetry, but the problem is that we don't know if the >>>> secondary CPU ever started executing the instructions or is stuck somewhere >>>> between the start of head.S and marking itself online at the end of >>>> secondary_start_kernel(). >>>> There are couple of places, where we get those other status codes, but not >>>> everywhere. If the issue is not easily reproducible, experiments on local >>>> setups do not yield anything. That's where I am attempting to add some more >>>> information in kernel to debug these issues. >>> >>> I think this is a game of diminishing returns. There's a lot of stuff >>> that the firmware/hypervisor can get wrong here and trying to detect or >>> handle that in Linux is going to be a real mess. For example, if it >>> enters the kernel at the wrong address, or in the wrong mode, or with >>> the MMU enabled etc. It sounds like you don't have much idea about >>> what happens in the failure case, so it might not even execute the code >>> that you're adding correctly. >> >> That is true. I agree. While we cannot and should not worry about adding >> logs for each of these firmware failure points in kernel, do you see any >> merit in adding any of this information to the kernel to at least narrow >> down the problem? > > Possibly, but I'd prefer to land the parallel bringup work first and then > see whether this still makes sense. > > Will Acked. Regards, Naman