Re: [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood

Borislav Petkov <[email protected]>
Newsgroups org.kernel.vger.linux-doc,org.kernel.vger.linux-edac,org.kernel.vger.linux-kernel
Message-ID <20260825225202.GEao4ckknyv8nRc8Dh@fat_crate.local>
On Tue, Aug 25, 2026 at 08:01:23PM +0000, Luck, Tony wrote:
> L3 on Intel is shared by the whole socket. So you'd lose 50% of cores for an L3 cache
> issue on a typical two socket system

Would panicking the whole system be better?

> (plus we'd have to bring back offline of CPU 0 if you want this to work for
> socket 0).

As long as you offline whatever you can and cordon off the accesses to the
faulty area as much as possible...

> Taking cores offline likely needs a bunch more plumbing outside of the kernel.
> Bare metal systems sometime isolate critical workloads on specific cores. VMM
> systems may bind guests to specific cores to provide consistent performance.

Well, nothing's free, right? If you want to run with degraded performance, you
should put that into the list of testing scenarios.

Thx.

-- 
Regards/Gruss,
    Boris.

https://people.kernel.org/tglx/notes-about-netiquette
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.