Re: [PATCH RFC 0/3] efi: mm/memory-failure: keep hardware-poisoned pages out of the next kexec
Kiryl Shutsemau <[email protected]> Wed, 22 Jul 2026 14:50:18 +0100
| Newsgroups | org.kernel.vger.linux-efi,org.infradead.lists.kexec,org.kernel.vger.linux-kernel,org.kvack.linux-mm |
|---|---|
| Message-ID | <amDHmNWfQ8eid9jH@thinkstation> |
On Fri, Jul 17, 2026 at 07:03:02AM -0700, Breno Leitao wrote: > Problem: > ======== > > When a page is hard-offlined due to an uncorrectable memory error (multi > bit ECC), memory_failure() sets PG_hwpoison, unmaps it, removes it from > the buddy allocator. This information is not carried to the next kernel > that is kexeced. The new kernel kexecs and trip over that bad memory > bank _again_. > > Why now: > ======== > > Several industry trends make this increasingly important: > > 1) DRAM is getting more expensive > 2) soldered / on-package memory (LPDDR, HBM) is becoming more common, so a > failing part can no longer simply be swapped; > 3) memory is kept in service far longer (at Meta, DRAM lifetime is being > drastically extended). > 4) It is more and more common to kexec instead of full reboot > 5) Increase of memory per system with CXL > > Proposed Solution: > ================== > > Carry the poisoned frames to the next kernel in a new EFI configuration > table, LINUX_EFI_POISONED_MEMORY, modeled on the existing > LINUX_EFI_MEMRESERVE table. > > EFI configuration tables already survive kexec: firmware hands the EFI > system table to every kernel in the chain, so a table installed once is > seen by all successors without a new handover channel. > > The mechanism is architecture independent, so x86 and arm64 use the same > code. > > The EFI stub installs an empty table while EFI boot service is up. A > configuration table can only be created there; the running kernel can only > append to it. > > Each hard-offlined frame is appended; an unpoison "removes" its entry > so a frame that is good again is not carried forward. > > The next kernel walks the inherited table early in > efi_config_parse_tables(), before memblock and the buddy allocator are > up, and memblock_reserve()s every recorded frame. The bad RAM is never > handed out. The allocator is only half of the problem. kexec segment placement doesn't know about any of this: kexec_file picks destinations from System RAM resources (memblock on arm64), and poisoned frames are only ever removed from buddy. So the next kernel image, initrd or purgatory can be placed on top of a poisoned frame -- the relocation memcpy then consumes the poison and the machine checks during the very kexec this series is supposed to protect. Same for the table's own list pages: overwrite one at placement time and the next kernel parses garbage. Hooking num_poisoned_pages_inc() also means soft-offlined pages are recorded. Those are functional pages, offlined predictively. Turning them into permanent losses for every kernel down the kexec chain does not seem right. I would record hard failures only. On the data structure: I am not an expert on DRAM failure modes, so I went reading. The field studies [1] say roughly a quarter of DRAM faults are multi-bit structures (row/column/bank), and the address interleaving means one such fault shows up to the OS as many separate 4K pages. A row fault lands in a window of tens of KB to about a MB; a column fault is one bad line per row, strided across the bank's entire footprint -- potentially thousands of pages scattered over gigabytes, reported one MCE at a time as they get touched. DDR5 on-die ECC hides most single-bit faults from the host, so the visible mix shifts toward these multi-bit modes over time. If that is accurate, per-4K entries have no natural bound, and every entry becomes a separate scattered memblock_reserve() in every future boot. I would consider a bitmap with one bit per 2M instead, modeled on struct efi_unaccepted_memory: the stub sizes it from the EFI memory map at cold boot, the running kernel sets a bit on hard offline, clears it on unpoison, and the next kernel reserves set units. It is one fixed-size allocation, so the whole grow-and-link machinery and the chain-parsing trust problem go away, and a row fault collapses into one or two bits. [1] https://arxiv.org/abs/2408.15302 -- Kiryl Shutsemau / Kirill A. Shutemov