Re: [PATCH v5] kexec: keep the next kernel off hardware-poisoned pages

Mike Rapoport <[email protected]>
Newsgroups org.infradead.lists.kexec,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <[email protected]>
Hi Breno,

On Mon, Aug 10, 2026 at 06:32:04AM -0700, Breno Leitao wrote:
> Memory failures (such as unrecoverable ECCs errors) are getting more and
> more common. The kernel knows how to handle it while running, marking it
> as poisoned (and SIGBUS user tasks).
> 
> Poisoned memory is removed from the buddy allocator, but, not from
> other places. A current problem is that kexec will load new kernel
> on top of a bad/poisoned memory, which is undesirable.
> 
> If the next kernel's image, initrd or purgatory lands on poisoned frame,
> the relocation copy writes to the bad memory and the machine checks

What does the machine check here? ;-)

> during the kexec.
> 
> Skip hardware-poisoned frames when placing segments: check them in the
> kexec_file hole finder so it lays the next kernel down on good memory,
> and reject a poisoned destination in sanity_check_segment_list() for
> the kexec_load path, which cannot relocate.
> 
> The two hole finders walk in opposite directions, so each asks for the
> end of the poison it has to clear: the top-down walk for the first
> poisoned page in the window, the bottom-up walk for the last. A poisoned
> hugetlb folio counts in full, as hugetlb keeps the flag on the folio and
> the poisoned subpages on its raw hwpoison list.

I had hard time parsing these two paragraphs. Can you please add more human
touch to them?
 
> Suggested-by: Kiryl Shutsemau <[email protected]>
> Signed-off-by: Breno Leitao <[email protected]>
>
> diff --git a/kernel/kexec_core.c b/kernel/kexec_core.c
> index dc770b9a6d053..7ee8c9f078f6b 100644
> --- a/kernel/kexec_core.c
> +++ b/kernel/kexec_core.c
> @@ -212,6 +212,16 @@ int sanity_check_segment_list(struct kimage *image)
>  	}
>  #endif
>  
> +	/*
> +	 * Reject destinations that land on hardware-poisoned memory: the
> +	 * relocation copy would machine-check on the bad frame.

Would cause machine-check exception?

> +	 */
> +	for (i = 0; i < nr_segments; i++) {
> +		if (range_first_hwpoison(image->segment[i].mem,
> +					 image->segment[i].memsz) != PHYS_ADDR_MAX)
> +			return -EHWPOISON;
> +	}
> +
>  	/*
>  	 * The destination addresses are searched from system RAM rather than
>  	 * being allocated from the buddy allocator, so they are not guaranteed
> diff --git a/kernel/kexec_file.c b/kernel/kexec_file.c
> index 59fb9d71e9d86..9ba6cc01af929 100644
> --- a/kernel/kexec_file.c
> +++ b/kernel/kexec_file.c
> @@ -475,6 +475,7 @@ static int locate_mem_hole_top_down(unsigned long start, unsigned long end,
>  {
>  	struct kimage *image = kbuf->image;
>  	unsigned long temp_start, temp_end;
> +	phys_addr_t poison;
>  
>  	temp_end = min(end, kbuf->buf_max);
>  	temp_start = temp_end - kbuf->memsz + 1;
> @@ -504,6 +505,15 @@ static int locate_mem_hole_top_down(unsigned long start, unsigned long end,
>  			continue;
>  		}
>  
> +		poison = range_first_hwpoison(temp_start, kbuf->memsz);
> +		if (poison != PHYS_ADDR_MAX) {
> +			/* we hit a poisoned page */
> +			if (poison < kbuf->memsz)
> +				return 0;

Won't we break out on the next iteration boundaries check? I.e.

		if (temp_start < start || temp_start < kbuf->buf_min)
			return 0;

> +			temp_start = poison - kbuf->memsz;
> +			continue;
> +		}
> +
>  		/* We found a suitable memory range */
>  		break;
>  	} while (1);
> @@ -520,6 +530,7 @@ static int locate_mem_hole_bottom_up(unsigned long start, unsigned long end,
>  {
>  	struct kimage *image = kbuf->image;
>  	unsigned long temp_start, temp_end;
> +	phys_addr_t poison;
>  
>  	temp_start = max(start, kbuf->buf_min);
>  
> @@ -546,6 +557,13 @@ static int locate_mem_hole_bottom_up(unsigned long start, unsigned long end,
>  			continue;
>  		}
>  
> +		poison = range_last_hwpoison(temp_start, kbuf->memsz);
> +		if (poison != PHYS_ADDR_MAX) {
> +			/* we hit a poisoned page */
> +			temp_start = poison + PAGE_SIZE;
> +			continue;
> +		}
> +
>  		/* We found a suitable memory range */
>  		break;
>  	} while (1);
> diff --git a/mm/memory-failure.c b/mm/memory-failure.c
> index a8b03e2920ba8..f3875680e7955 100644
> --- a/mm/memory-failure.c
> +++ b/mm/memory-failure.c
> @@ -96,6 +96,47 @@ void num_poisoned_pages_sub(unsigned long pfn, long i)
>  		memblk_nr_poison_sub(pfn, i);
>  }
>  
> +/*
> + * Return the first or the last hardware-poisoned online page in [start,
> + * start + size), or PHYS_ADDR_MAX if the range is clean.
> + */
> +static phys_addr_t range_hwpoison(phys_addr_t start, unsigned long size,
> +				  bool first)
> +{
> +	phys_addr_t poison = PHYS_ADDR_MAX;
> +	unsigned long pfn, end_pfn;
> +
> +	if (!size || !atomic_long_read(&num_poisoned_pages))
> +		return poison;
> +
> +	end_pfn = PHYS_PFN(start + size - 1);
> +	for (pfn = PHYS_PFN(start); pfn <= end_pfn; pfn++) {
> +		struct page *page = pfn_to_online_page(pfn);
> +
> +		cond_resched();

cond_resched() for every pfn is too much, isn't it?

> +
> +		if (!page || !is_page_hwpoison(page))
> +			continue;
> +
> +		if (first)
> +			return PFN_PHYS(pfn);
> +
> +		poison = PFN_PHYS(pfn);
> +	}
> +
> +	return poison;
> +}
> +
> +phys_addr_t range_first_hwpoison(phys_addr_t start, unsigned long size)
> +{
> +	return range_hwpoison(start, size, true);
> +}
> +
> +phys_addr_t range_last_hwpoison(phys_addr_t start, unsigned long size)
> +{
> +	return range_hwpoison(start, size, false);
> +}
> +
>  /**
>   * MF_ATTR_RO - Create sysfs entry for each memory failure statistics.
>   * @_name: name of the file in the per NUMA sysfs directory.
> 
> ---
> base-commit: c5e32e86ca02b003f86e095d379b38148999293d
> change-id: 20260727-kexec_posioned-72bb0a4143a0
> 
> Best regards,
> --  
> Breno Leitao <[email protected]>
> 

-- 
Sincerely yours,
Mike.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.