Re: [PATCH v5] kexec: keep the next kernel off hardware-poisoned pages
Mike Rapoport <[email protected]>
| Newsgroups | org.infradead.lists.kexec,org.kernel.vger.linux-kernel,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
Hi Breno, On Mon, Aug 10, 2026 at 06:32:04AM -0700, Breno Leitao wrote: > Memory failures (such as unrecoverable ECCs errors) are getting more and > more common. The kernel knows how to handle it while running, marking it > as poisoned (and SIGBUS user tasks). > > Poisoned memory is removed from the buddy allocator, but, not from > other places. A current problem is that kexec will load new kernel > on top of a bad/poisoned memory, which is undesirable. > > If the next kernel's image, initrd or purgatory lands on poisoned frame, > the relocation copy writes to the bad memory and the machine checks What does the machine check here? ;-) > during the kexec. > > Skip hardware-poisoned frames when placing segments: check them in the > kexec_file hole finder so it lays the next kernel down on good memory, > and reject a poisoned destination in sanity_check_segment_list() for > the kexec_load path, which cannot relocate. > > The two hole finders walk in opposite directions, so each asks for the > end of the poison it has to clear: the top-down walk for the first > poisoned page in the window, the bottom-up walk for the last. A poisoned > hugetlb folio counts in full, as hugetlb keeps the flag on the folio and > the poisoned subpages on its raw hwpoison list. I had hard time parsing these two paragraphs. Can you please add more human touch to them? > Suggested-by: Kiryl Shutsemau <[email protected]> > Signed-off-by: Breno Leitao <[email protected]> > > diff --git a/kernel/kexec_core.c b/kernel/kexec_core.c > index dc770b9a6d053..7ee8c9f078f6b 100644 > --- a/kernel/kexec_core.c > +++ b/kernel/kexec_core.c > @@ -212,6 +212,16 @@ int sanity_check_segment_list(struct kimage *image) > } > #endif > > + /* > + * Reject destinations that land on hardware-poisoned memory: the > + * relocation copy would machine-check on the bad frame. Would cause machine-check exception? > + */ > + for (i = 0; i < nr_segments; i++) { > + if (range_first_hwpoison(image->segment[i].mem, > + image->segment[i].memsz) != PHYS_ADDR_MAX) > + return -EHWPOISON; > + } > + > /* > * The destination addresses are searched from system RAM rather than > * being allocated from the buddy allocator, so they are not guaranteed > diff --git a/kernel/kexec_file.c b/kernel/kexec_file.c > index 59fb9d71e9d86..9ba6cc01af929 100644 > --- a/kernel/kexec_file.c > +++ b/kernel/kexec_file.c > @@ -475,6 +475,7 @@ static int locate_mem_hole_top_down(unsigned long start, unsigned long end, > { > struct kimage *image = kbuf->image; > unsigned long temp_start, temp_end; > + phys_addr_t poison; > > temp_end = min(end, kbuf->buf_max); > temp_start = temp_end - kbuf->memsz + 1; > @@ -504,6 +505,15 @@ static int locate_mem_hole_top_down(unsigned long start, unsigned long end, > continue; > } > > + poison = range_first_hwpoison(temp_start, kbuf->memsz); > + if (poison != PHYS_ADDR_MAX) { > + /* we hit a poisoned page */ > + if (poison < kbuf->memsz) > + return 0; Won't we break out on the next iteration boundaries check? I.e. if (temp_start < start || temp_start < kbuf->buf_min) return 0; > + temp_start = poison - kbuf->memsz; > + continue; > + } > + > /* We found a suitable memory range */ > break; > } while (1); > @@ -520,6 +530,7 @@ static int locate_mem_hole_bottom_up(unsigned long start, unsigned long end, > { > struct kimage *image = kbuf->image; > unsigned long temp_start, temp_end; > + phys_addr_t poison; > > temp_start = max(start, kbuf->buf_min); > > @@ -546,6 +557,13 @@ static int locate_mem_hole_bottom_up(unsigned long start, unsigned long end, > continue; > } > > + poison = range_last_hwpoison(temp_start, kbuf->memsz); > + if (poison != PHYS_ADDR_MAX) { > + /* we hit a poisoned page */ > + temp_start = poison + PAGE_SIZE; > + continue; > + } > + > /* We found a suitable memory range */ > break; > } while (1); > diff --git a/mm/memory-failure.c b/mm/memory-failure.c > index a8b03e2920ba8..f3875680e7955 100644 > --- a/mm/memory-failure.c > +++ b/mm/memory-failure.c > @@ -96,6 +96,47 @@ void num_poisoned_pages_sub(unsigned long pfn, long i) > memblk_nr_poison_sub(pfn, i); > } > > +/* > + * Return the first or the last hardware-poisoned online page in [start, > + * start + size), or PHYS_ADDR_MAX if the range is clean. > + */ > +static phys_addr_t range_hwpoison(phys_addr_t start, unsigned long size, > + bool first) > +{ > + phys_addr_t poison = PHYS_ADDR_MAX; > + unsigned long pfn, end_pfn; > + > + if (!size || !atomic_long_read(&num_poisoned_pages)) > + return poison; > + > + end_pfn = PHYS_PFN(start + size - 1); > + for (pfn = PHYS_PFN(start); pfn <= end_pfn; pfn++) { > + struct page *page = pfn_to_online_page(pfn); > + > + cond_resched(); cond_resched() for every pfn is too much, isn't it? > + > + if (!page || !is_page_hwpoison(page)) > + continue; > + > + if (first) > + return PFN_PHYS(pfn); > + > + poison = PFN_PHYS(pfn); > + } > + > + return poison; > +} > + > +phys_addr_t range_first_hwpoison(phys_addr_t start, unsigned long size) > +{ > + return range_hwpoison(start, size, true); > +} > + > +phys_addr_t range_last_hwpoison(phys_addr_t start, unsigned long size) > +{ > + return range_hwpoison(start, size, false); > +} > + > /** > * MF_ATTR_RO - Create sysfs entry for each memory failure statistics. > * @_name: name of the file in the per NUMA sysfs directory. > > --- > base-commit: c5e32e86ca02b003f86e095d379b38148999293d > change-id: 20260727-kexec_posioned-72bb0a4143a0 > > Best regards, > -- > Breno Leitao <[email protected]> > -- Sincerely yours, Mike.