Re: [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood
Borislav Petkov <[email protected]>
| Newsgroups | org.kernel.vger.linux-doc,org.kernel.vger.linux-edac,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <20260825191555.GCao3p63v3fXzwa7E7@fat_crate.local> |
On Tue, Aug 25, 2026 at 04:26:26PM +0000, Luck, Tony wrote:
> > Why isn't the strategy here page offlining and shutting up the source of the
> > error instead of doing silly counting and not doing anything to contain the
> > errors in the first place?
>
> While Breno wasn't specific about the source of the errors in these messages, I'm
> guessing that they might be coming from cache errors rather than DDR memory.
>
> There isn't a good way for software[1] to suppress these errors. Taking memory
> pages offline when the problem is the cache will just deplete available memory
> without solving the problem.
I have been thinking about this *years* ago. If it is cache errors, we should
simply offline the core or cores using that cache. We have a lot of cores
nowadays :)
In general, us being a lot more resilient and applying automatic containment
and recovery actions should be the goal IMO.
Thx.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette