Re: [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood

Borislav Petkov <[email protected]>
Newsgroups org.kernel.vger.linux-doc,org.kernel.vger.linux-edac,org.kernel.vger.linux-kernel
Message-ID <20260825191555.GCao3p63v3fXzwa7E7@fat_crate.local>
On Tue, Aug 25, 2026 at 04:26:26PM +0000, Luck, Tony wrote:
> > Why isn't the strategy here page offlining and shutting up the source of the
> > error instead of doing silly counting and not doing anything to contain the
> > errors in the first place?
> 
> While Breno wasn't specific about the source of the errors in these messages, I'm
> guessing that they might be coming from cache errors rather than DDR memory.
> 
> There isn't a good way for software[1] to suppress these errors. Taking memory
> pages offline when the problem is the cache will just deplete available memory
> without solving the problem.

I have been thinking about this *years* ago. If it is cache errors, we should
simply offline the core or cores using that cache. We have a lot of cores
nowadays :)

In general, us being a lot more resilient and applying automatic containment
and recovery actions should be the goal IMO.

Thx.

-- 
Regards/Gruss,
    Boris.

https://people.kernel.org/tglx/notes-about-netiquette
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.