Re: [PATCH v2] ACPI: APEI: Handle repeated SEA error storms

hejunhao <[email protected]>
Newsgroups org.kernel.vger.linux-edac,org.infradead.lists.linux-arm-kernel,org.kernel.vger.linux-acpi,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
Hi All,


a gentle ping on this.
All comments are welcomed!


On 2026/7/23 16:32, hejunhao wrote:
>
> On 2026/7/9 1:14, Luck, Tony wrote:
>> On Wed, May 27, 2026 at 04:27:07PM +0800, Junhao He wrote:
>>> When hardware memory corruption occurs and a user process accesses the
>>> corrupted page, the CPU triggers a Synchronous External Abort (SEA).
>>> The kernel invokes do_sea() to handle the exception, which calls
>>> memory_failure() to handle the faulty page.
>>>
>>> Scenario 1: Memory Error Interrupt First, then SEA
>>> The page is already poisoned by the memory error interrupt path. The
>>> subsequent SEA handler sends a SIGBUS to the task, which accesses the
>>> poisoned page. This flow is correct.
>>>
>>> Scenario 2: SEA first, then memory error interrupt (problematic scenario)
>>> If a user task directly accesses corrupted memory through a PFNMAP-style
>>> mapping (e.g., devmem), the page may still be in the free-buddy state when
>>> SEA is handled. In this case, memory_failure() will poison the page without
>>> invoking kill_accessing_process(), and then takes the free-buddy recovery
>>> path.
>>>
>>> After the CPU returns to the task context, the task re-enters the SEA
>>> handler due to the same access. However, ghes_estatus_cached() suppresses
>>> all subsequent entries during the 10-second window, preventing
>>> ghes_do_proc() from being called. This suppression blocks the
>>> MF_ACTION_REQUIRED-based SIGBUS delivery, causing the kernel to fail to
>>> kill the task immediately. Consequently, the process keeps re-entering
>>> the SEA handler, leading to an SEA storm. Later, the memory error
>>> interrupt path also cannot kill the task, leaving the system stuck in
>>> this repeated loop.
>>>
>>> The following error logs are explained using the devmem process:
>>>   NOTICE:  SEA Handle
>>>   [Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 9
>>>   [Hardware Error]: event severity: recoverable
>>>   [Hardware Error]:   section_type: ARM processor error
>>>   [Hardware Error]:    physical fault address: 0x0000001000093c00
>>>   [T54990] Memory failure: 0x1000093: recovery action for free buddy page: Recovered
>>>   [ T9955] EDAC MC0: 1 UE Multi-bit ECC on unknown memory
>>>            (page:0x1000093 offset:0xc00 grain:1 - APEI location: ...)
>>>   NOTICE:  SEA Handle
>>>   NOTICE:  SEA Handle
>>>   ...
>>>   ...        ---> SEA storm
>>>   ...
>>>   NOTICE:  SEA Handle
>>>   [ T9955] Memory failure: 0x1000093: already hardware poisoned
>>>   ghes_print_estatus: 1 callbacks suppressed
>>>   [Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 9
>>>   [Hardware Error]: event severity: recoverable
>>>   [Hardware Error]:   section_type: ARM processor error
>>>   [Hardware Error]:    physical fault address: 0x0000001000093c00
>>>   [T54990] Memory failure: 0x1000093: already hardware poisoned
>>>   [T54990] 0x1000093: Sending SIGBUS to devmem:54990 due to hardware memory corruption
>>>
>>> To resolve this, return an error when encountering the same SEA again.
>>> The subsequent SEA handler invocation uses arm64_notify_die() to send a
>>> SIGBUS signal to the task, which terminates the process and prevents it
>>> from re-entering the handler loop.
>>>
>>> Signed-off-by: Junhao He <[email protected]>
>>> ---
>>>  drivers/acpi/apei/ghes.c | 10 +++++++++-
>>>  1 file changed, 9 insertions(+), 1 deletion(-)
>>>
>>> Changes in V2:
>>> 1. update the commit message per suggestion from Xueshuai
>>> 2. Add a check to only return failure on the ghes_notify_sea() path,
>>>    avoiding impact on other NMI-type GHES handlers.
>>> Link to V1 - https://lore.kernel.org/all/[email protected]/
>>>
>>> diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c
>>> index 3236a3ce79d6..787664740150 100644
>>> --- a/drivers/acpi/apei/ghes.c
>>> +++ b/drivers/acpi/apei/ghes.c
>>> @@ -1383,8 +1383,16 @@ static int ghes_in_nmi_queue_one_entry(struct ghes *ghes,
>> ghes_in_nmi_queue_one_entry() is called from two places:
>>
>> __ghes_sdei_callback()
>>
>> I think, but I'm not sure, that you are addressing call from here. By
>> returning -ECANCELLED this code skips the irq_work_queue() call and
>> returns -ENOENT (instead of zero).
>>
>> ghes_in_nmi_spool_from_list()
>>
>> Effect here is more complex. Depends on whether there are multiple
>> ghes entries on the rcu_list. If there was just one, and you return
>> -ECANCELLED, then irq_work_queue() is skipped. But if there are multiple
>> entries on the list and at least one of them gets a zero return from
>> ghes_in_nmi_queue_one_entry() then irq_work_queue() is called.
> Hi Luck Tony,
>
> Thanks for the comments.
>
> This fix is only for the call to ghes_in_nmi_spool_from_list() from ghes_notify_sea() on
> ARM platforms (ACPI_HEST_NOTIFY_SEA); other paths are unchanged.
>
> If multiple SEA errors occur simultaneously, multiple entries in ghes_sea are active
> and one of them returns 0, then -ECANCELED is discarded, but the unhandled
> SEA will simply trigger again and re-enter the SEA handler to be processed.
>
> The worst case is when multiple valid SEA entries persist within the 10-second
> ghes_estatus_cached() window — the fact is that a great many SEA errors are
> produced, which actually constitutes a real SEA storm, not one triggered by a single
> SEA, and in this situation not delivering SIGBUS in time to kill the process is acceptable.
>
> Best regards,
> Junhao.
>
>>
>> So which of these callers is the one that needs this fix? And is the
>> other called OK with getting -ECANCELLED error?
>>
>>>  	ghes_clear_estatus(ghes, &tmp_header, buf_paddr, fixmap_idx);
>>>  
>>>  	/* This error has been reported before, don't process it again. */
>>> -	if (ghes_estatus_cached(estatus))
>>> +	if (ghes_estatus_cached(estatus)) {
>>> +		/*
>>> +		 * Return failure on duplicate SEA entries so that the
>>> +		 * subsequent SEA handler invocation sends a SIGBUS signal to
>>> +		 * the task to prevent it from re-entering the handler loop.
>>> +		 */
>>> +		if (is_hest_sync_notify(ghes))
>>> +			rc = -ECANCELED;
>>>  		goto no_work;
>>> +	}
>>>  
>>>  	llist_add(&estatus_node->llnode, &ghes_estatus_llist);
>>>  
>>> -- 
>>> 2.33.0
>>>
>> -Tony
>>
>> .
>>
> .
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.