Re: [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file()

Vernon Yang <[email protected]>
Newsgroups org.kernel.vger.linux-kernel,org.kernel.vger.stable,org.kvack.linux-mm
Message-ID <[email protected]>
On Tue, Aug 11, 2026 at 12:12:25PM -0700, Andrew Morton wrote:
> On Tue, 11 Aug 2026 21:36:55 +0800 Vernon Yang <[email protected]> wrote:
>
> > When the swap entries found exceed max_ptes_swap, the loop is left via
> > break with folio still holding the xarray value that encodes the swap
> > entry, not valid folio pointer.
> >
> > That value is passed to trace_mm_khugepaged_scan_file(), which feeds it
> > to folio_pfn(). On FLATMEM and SPARSEMEM_VMEMMAP, the page_to_pfn() is
> > plain pointer arithmetic, so the trace event merely prints bogus
> > scan_pfn. On classic SPARSEMEM, the page_to_pfn() reads page->flags,
> > dereferencing the tiny encoded integer and oopsing khugepaged whenever
> > the trace event is enabled.
> >
> > So set folio to NULL before breaking out, the tracepoint maps NULL to
> > scan_pfn of -1, just like exhausted scan naturally.
> >
> > Fixes: d41fd2016ed0 ("mm/khugepaged: add tracepoint to hpage_collapse_scan_file()")
>
> Added in 2022.  Why so long - do people not use tracing?

For most architectures, SPARSEMEM_VMEMMAP is the default, where
page_to_pfn() is plain pointer arithmetic, merely printing bogus
scan_pfn. Only on classic SPARSEMEM, the swap entry value cause
khugepaged to oops.

> Sashiko might have a found a couple of other tracing bugs in this code,
> which I suggest are on-topic for your patch:
>
> 	https://sashiko.dev/#/patchset/[email protected]
>

David also mentioned "Using the folio after dropping the reference",
and the same issue exists in
trace_mm_khugepaged_collapse_file()/trace_mm_khugepaged_scan_pmd(),
which merely prints a overdue scan_pfn without oopsing khugepaged.
However, this approach does have potential issues, just that they
haven't been triggered yet. I can fix them together.

I will add PATCH#2 and PATCH#3 in the next version of the patchset to
fix the other two trace_xxx() functions, rather than mixing them
together into one large patch.

> Also a possible bug mapping large folios which straddle i_size, which
> is a separate thing.

Yes, I will submit a separate fix patch later to resolve this issue.

--
Cheers,
Vernon
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.