Re: [PATCH v2] mm/madvise: avoid skipping pages after splitting large folios

"Lorenzo Stoakes (ARM)" <[email protected]>
Newsgroups org.kernel.vger.stable,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <anRPzJiraQUnsHa5@lucifer>
On Thu, Aug 06, 2026 at 10:41:06AM +0200, David Hildenbrand (Arm) wrote:
> On 8/6/26 07:55, Yunhui Cui wrote:
> > madvise_inject_error() advances through the requested range using the
> > size of the page returned by get_user_pages_fast(). Saving the size
> > before error injection is required for hugetlb pages because successful
> > soft offlining can dissolve the source huge page.
> >
> > That stride is incorrect for non-hugetlb large folios in system memory.
> > The memory failure handlers split such a folio and handle only the base
> > page for the supplied PFN. Advancing by the pre-split folio size then
> > skips the remaining pages in the requested range while madvise() still
> > reports success.
> >
> > Advance by PAGE_SIZE for non-hugetlb folios in system memory. Retain
> > folio_size() for hugetlb and ZONE_DEVICE folios, as compound Device DAX
> > folios are handled as a whole.
> >
> > Fixes: 19bfbe22f59a ("mm, hugetlb, soft_offline: save compound page order before page migration")
> > Cc: [email protected]
> > Signed-off-by: Yunhui Cui <[email protected]>
> > ---
> >  mm/madvise.c | 13 +++++++++----
> >  1 file changed, 9 insertions(+), 4 deletions(-)
> >
> > diff --git a/mm/madvise.c b/mm/madvise.c
> > index 5a09cc24f04a0..e9d4c3bbc5290 100644
> > --- a/mm/madvise.c
> > +++ b/mm/madvise.c
> > @@ -1455,20 +1455,25 @@ static int madvise_inject_error(struct madvise_behavior *madv_behavior)
> >
> >  	for (; start < end; start += size) {
> >  		unsigned long pfn;
> > +		struct folio *folio;
> >  		struct page *page;
> >  		int ret;
> >
> >  		ret = get_user_pages_fast(start, 1, 0, &page);
> >  		if (ret != 1)
> >  			return ret;
> > +		folio = page_folio(page);
> >  		pfn = page_to_pfn(page);
> >
> >  		/*
> > -		 * When soft offlining hugepages, after migrating the page
> > -		 * we dissolve it, therefore in the second loop "page" will
> > -		 * no longer be a compound page.
> > +		 * Non-hugetlb large folios in system memory are split and only
> > +		 * the addressed base page is handled. Hugetlb folios may be
> > +		 * dissolved and ZONE_DEVICE folios may be handled as a whole,
> > +		 * so save their size before error injection.
> >  		 */
> > -		size = page_size(compound_head(page));
> > +		size = PAGE_SIZE;
> > +		if (folio_test_hugetlb(folio) || folio_is_zone_device(folio))
> > +			size = folio_size(folio);
>
>
> We should never ever try deferring "how much has been mapped" from a single PTE.

Inferring? :)

>
> While this currently works for hugetlb, it's just an anti-pattern to throw
> hugetlb checks and similar around.
>
> So this is not the way to fix it.

All of this code is disgusting. Also what about an address range that is
partially inside a folio...? Then presumably the whole folio is discarded? Or
does it figure it out somehow and does the split for the rest of the range?

And what if end < the end of the hugetlb size?

Ugh god I hate all of this, it's so so bad. And hugetlb being the special
snowflake is the cherry on the s*** cake...

>
> --
> Cheers,
>
> David

--
Cheers, Lorenzo
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.