Re: [PATCH v2] mm/madvise: avoid skipping pages after splitting large folios

"David Hildenbrand (Arm)" <[email protected]>
Newsgroups org.kernel.vger.stable,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <[email protected]>
On 8/6/26 07:55, Yunhui Cui wrote:
> madvise_inject_error() advances through the requested range using the
> size of the page returned by get_user_pages_fast(). Saving the size
> before error injection is required for hugetlb pages because successful
> soft offlining can dissolve the source huge page.
> 
> That stride is incorrect for non-hugetlb large folios in system memory.
> The memory failure handlers split such a folio and handle only the base
> page for the supplied PFN. Advancing by the pre-split folio size then
> skips the remaining pages in the requested range while madvise() still
> reports success.
> 
> Advance by PAGE_SIZE for non-hugetlb folios in system memory. Retain
> folio_size() for hugetlb and ZONE_DEVICE folios, as compound Device DAX
> folios are handled as a whole.
> 
> Fixes: 19bfbe22f59a ("mm, hugetlb, soft_offline: save compound page order before page migration")
> Cc: [email protected]
> Signed-off-by: Yunhui Cui <[email protected]>
> ---
>  mm/madvise.c | 13 +++++++++----
>  1 file changed, 9 insertions(+), 4 deletions(-)
> 
> diff --git a/mm/madvise.c b/mm/madvise.c
> index 5a09cc24f04a0..e9d4c3bbc5290 100644
> --- a/mm/madvise.c
> +++ b/mm/madvise.c
> @@ -1455,20 +1455,25 @@ static int madvise_inject_error(struct madvise_behavior *madv_behavior)
>  
>  	for (; start < end; start += size) {
>  		unsigned long pfn;
> +		struct folio *folio;
>  		struct page *page;
>  		int ret;
>  
>  		ret = get_user_pages_fast(start, 1, 0, &page);
>  		if (ret != 1)
>  			return ret;
> +		folio = page_folio(page);
>  		pfn = page_to_pfn(page);
>  
>  		/*
> -		 * When soft offlining hugepages, after migrating the page
> -		 * we dissolve it, therefore in the second loop "page" will
> -		 * no longer be a compound page.
> +		 * Non-hugetlb large folios in system memory are split and only
> +		 * the addressed base page is handled. Hugetlb folios may be
> +		 * dissolved and ZONE_DEVICE folios may be handled as a whole,
> +		 * so save their size before error injection.
>  		 */
> -		size = page_size(compound_head(page));
> +		size = PAGE_SIZE;
> +		if (folio_test_hugetlb(folio) || folio_is_zone_device(folio))
> +			size = folio_size(folio);


We should never ever try deferring "how much has been mapped" from a single PTE.

While this currently works for hugetlb, it's just an anti-pattern to throw
hugetlb checks and similar around.

So this is not the way to fix it.

-- 
Cheers,

David
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.