Re: [PATCH] kexec: keep the next kernel off hardware-poisoned pages
Miaohe Lin <[email protected]> Wed, 29 Jul 2026 17:33:17 +0800
| Newsgroups | org.infradead.lists.kexec,org.kernel.vger.linux-kernel,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
On 2026/7/29 0:22, Breno Leitao wrote: > Memory failures (such as unrecoverable ECCs errors) are getting more and > more common. The kernel knows how to handle it while running, marking it > as poisoned (and SIGBUS user tasks). > > Poisoned memory is removed from the buddy allocator, but, not from > other places. A current problem is that kexec will load new kernel > on top of a bad/poisoned memory, which is undesirable. > > If the next kernel's image, initrd or purgatory lands on poisoned frame, > the relocation copy writes to the bad memory and the machine checks > during the kexec. > > Skip hardware-poisoned frames when placing segments: check them in the > kexec_file hole finder so it lays the next kernel down on good memory, > and reject a poisoned destination in sanity_check_segment_list() for > the kexec_load path, which cannot relocate. > > Suggested-by: Kiryl Shutsemau <[email protected]> > Signed-off-by: Breno Leitao <[email protected]> > --- > include/linux/mm.h | 6 ++++++ > kernel/kexec_core.c | 10 ++++++++++ > kernel/kexec_file.c | 14 ++++++++++++++ > mm/memory-failure.c | 18 ++++++++++++++++++ > 4 files changed, 48 insertions(+) > > diff --git a/include/linux/mm.h b/include/linux/mm.h > index 7fabe6c66b4b7..a89108bcc3f90 100644 > --- a/include/linux/mm.h > +++ b/include/linux/mm.h > @@ -5192,6 +5192,7 @@ extern const struct attribute_group memory_failure_attr_group; > extern void memory_failure_queue(unsigned long pfn, int flags); > void num_poisoned_pages_inc(unsigned long pfn); > void num_poisoned_pages_sub(unsigned long pfn, long i); > +bool range_contains_hwpoison(phys_addr_t start, unsigned long size); > #else > static inline void memory_failure_queue(unsigned long pfn, int flags) > { > @@ -5204,6 +5205,11 @@ static inline void num_poisoned_pages_inc(unsigned long pfn) > static inline void num_poisoned_pages_sub(unsigned long pfn, long i) > { > } > + > +static inline bool range_contains_hwpoison(phys_addr_t start, unsigned long size) > +{ > + return false; > +} > #endif > > #if defined(CONFIG_MEMORY_FAILURE) && defined(CONFIG_MEMORY_HOTPLUG) > diff --git a/kernel/kexec_core.c b/kernel/kexec_core.c > index dc770b9a6d053..18793f925835f 100644 > --- a/kernel/kexec_core.c > +++ b/kernel/kexec_core.c > @@ -212,6 +212,16 @@ int sanity_check_segment_list(struct kimage *image) > } > #endif > > + /* > + * Reject destinations that land on hardware-poisoned memory: the > + * relocation copy would machine-check on the bad frame. > + */ > + for (i = 0; i < nr_segments; i++) { > + if (range_contains_hwpoison(image->segment[i].mem, > + image->segment[i].memsz)) > + return -EADDRNOTAVAIL; > + } > + > /* > * The destination addresses are searched from system RAM rather than > * being allocated from the buddy allocator, so they are not guaranteed > diff --git a/kernel/kexec_file.c b/kernel/kexec_file.c > index 59fb9d71e9d86..cd77350f12f85 100644 > --- a/kernel/kexec_file.c > +++ b/kernel/kexec_file.c > @@ -504,6 +504,13 @@ static int locate_mem_hole_top_down(unsigned long start, unsigned long end, > continue; > } > > + /* Avoid placing the next kernel on hardware-poisoned memory */ > + if (range_contains_hwpoison(temp_start, > + temp_end - temp_start + 1)) { > + temp_start = temp_start - PAGE_SIZE; > + continue; > + } > + > /* We found a suitable memory range */ > break; > } while (1); > @@ -546,6 +553,13 @@ static int locate_mem_hole_bottom_up(unsigned long start, unsigned long end, > continue; > } > > + /* Avoid placing the next kernel on hardware-poisoned memory */ > + if (range_contains_hwpoison(temp_start, > + temp_end - temp_start + 1)) { > + temp_start = temp_start + PAGE_SIZE; > + continue; > + } > + > /* We found a suitable memory range */ > break; > } while (1); > diff --git a/mm/memory-failure.c b/mm/memory-failure.c > index a8b03e2920ba8..032bb23db7566 100644 > --- a/mm/memory-failure.c > +++ b/mm/memory-failure.c > @@ -96,6 +96,24 @@ void num_poisoned_pages_sub(unsigned long pfn, long i) > memblk_nr_poison_sub(pfn, i); > } > > +/* > + * Check if any page in [start, start + size) has hardware-poisoned segments > + */ > +bool range_contains_hwpoison(phys_addr_t start, unsigned long size) > +{ > + unsigned long pfn, end_pfn; > + > + if (!size || !atomic_long_read(&num_poisoned_pages)) > + return false; > + > + end_pfn = PHYS_PFN(start + size - 1); > + for (pfn = PHYS_PFN(start); pfn <= end_pfn; pfn++) { > + if (pfn_valid(pfn) && PageHWPoison(pfn_to_page(pfn))) Should we use pfn_to_online_page here? I doubt we might meet the same problem in [1] if pfn_to_page is used here. [1]: https://git.kernel.org/pub/scm/linux/kernel/git/next/linux-next.git/commit/?id=d613f53c83ec47089c4e25859d5e8e0359f6f8da Thanks. .