On 7/31/2026 1:07 PM, Matthew Wilcox (Oracle) wrote:
> If we fail to get a reference on the hugetlb folio, then it might be
> freed as we operate on it. Prevent the freeing and the attendant races
> around manipulation of the raw_hwp list by holding the hugetlb_lock,
> which is also held by the hugetlb
>
>
> If we fail to get a reference on the hugetlb folio, then it might be
> freed as we operate on it. Prevent the freeing and the attendant races
> around manipulation of the raw_hwp list by holding the hugetlb_lock,
> which is also held by the hugetlb code when freeing hugetlb folios.
>
> Fixes: ac5fcde0a96a ("mm, hwpoison: make unpoison aware of raw error info in hwpoisoned hugepage")
> Signed-off-by: Matthew Wilcox (Oracle) <[email protected]>
> Reviewed-by: Gregory Price (Meta) <[email protected]>
> ---
> include/linux/hugetlb.h | 19 +++++++++++++++++++
> mm/memory-failure.c | 6 +++++-
> 2 files changed, 24 insertions(+), 1 deletion(-)
>
> diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
> index 2abaf99321e9..50eab2c23299 100644
> --- a/include/linux/hugetlb.h
> +++ b/include/linux/hugetlb.h
> @@ -110,6 +110,17 @@ extern struct resv_map *resv_map_alloc(void);
> void resv_map_release(struct kref *ref);
>
> extern spinlock_t hugetlb_lock;
> +
> +static inline void hugetlb_lock_irq(void)
> +{
> + spin_lock_irq(&hugetlb_lock);
> +}
> +
> +static inline void hugetlb_unlock_irq(void)
> +{
> + spin_unlock_irq(&hugetlb_lock);
> +}
> +
> extern int hugetlb_max_hstate __read_mostly;
> #define for_each_hstate(h) \
> for ((h) = hstates; (h) < &hstates[hugetlb_max_hstate]; (h)++)
> @@ -279,6 +290,14 @@ unsigned int arch_hugetlb_cma_order(void);
>
> #else /* !CONFIG_HUGETLB_PAGE */
>
> +static inline void hugetlb_lock_irq(void)
> +{
> +}
> +
> +static inline void hugetlb_unlock_irq(void)
> +{
> +}
> +
> static inline void hugetlb_dup_vma_private(struct vm_area_struct *vma)
> {
> }
> diff --git a/mm/memory-failure.c b/mm/memory-failure.c
> index 944e6e1d4971..1dd0e7b99bb1 100644
> --- a/mm/memory-failure.c
> +++ b/mm/memory-failure.c
> @@ -2725,13 +2725,17 @@ int unpoison_memory(unsigned long pfn)
>
> ghp = get_hwpoison_page(p, MF_UNPOISON);
> if (!ghp) {
> + hugetlb_lock_irq();
> if (folio_test_hugetlb(folio)) {
> huge = true;
> count = folio_free_raw_hwp(folio, false);
> - if (count == 0)
> + if (count == 0) {
> + hugetlb_unlock_irq();
> goto unlock_mutex;
> + }
> }
> ret = folio_test_clear_hwpoison(folio) ? 0 : -EBUSY;
> + hugetlb_unlock_irq();
> } else if (ghp < 0) {
> if (ghp == -EHWPOISON) {
> ret = put_page_back_buddy(p) ? 0 : -EBUSY;
> --
> 2.47.3
>
Patch itself looks good, so Reviewed-by: Jane Chu <[email protected]>
That said, there is a pre-existing issue:
folio_free_raw_hwp() should check HPG_raw_hwp_unreliable, and fail the
act of unpoison just like what __update_and_free_hugetlb_folio() does -
static void __update_and_free_hugetlb_folio(struct hstate *h,
struct folio *folio)
{
bool clear_flag = folio_test_hugetlb_vmemmap_optimized(folio);
if (hstate_is_gigantic_no_runtime(h))
return;
/*
* If we don't know which subpages are hwpoisoned, we can't free
* the hugepage, so it's leaked intentionally.
*/
if (folio_test_hugetlb_raw_hwp_unreliable(folio))
return;
thanks,
-jane
lmpx.com only provides a reader for public news (NNTP) servers. It is not
affiliated with the servers or forums shown here and is not responsible for
the content of articles, which is written by their respective authors.