[RFC PATCH v3 0/8] batch lookups in follow_page_mask()

Rik van Riel <[email protected]>
Newsgroups gmane.linux.kernel,gmane.linux.kernel.mm
Message-ID <[email protected]>
follow_page_mask() walks the page tables one page at a time, even when the
caller asked for a whole run of contiguous pages. Every page of a large folio
re-enters the pmd/pud/pte walk and re-takes the page table lock.

This series changes follow_page_mask() to return a page count and a fill an
array of pages, instead of a single struct page, so a walker can hand back
more than one page per call.

Patches 1 to 4 are preparation, no functional change:

  1: move __get_user_pages()'s open-coded pages[] fill and cache flush into a
     gup_fill_pages() helper, which the rest of the series reuses.
  2: convert the follow_page_mask()/follow_p4d_mask()/follow_pud_mask()/
     follow_pmd_mask()/follow_page_pte() call chain to return a long instead of
     a struct page pointer or ERR_PTR(). Every path still handles one page.
  3: split the "commit to a resolved page" tail of follow_page_pte() into
     follow_page_pte_commit().
  4: split the "work out which page this PTE maps" half of follow_page_pte()
     into follow_one_pte(), leaving one unlock and one exit.

Patch 5 has the huge page paths store the page and leave the array fill to
follow_pud_mask()/follow_pmd_mask() after they unlock, so the cache flushes
happen outside the pud/pmd critical section.

Patch 6 has follow_huge_pud()/follow_huge_pmd() report the huge page's real
subpage count instead of a separate *page_mask output, and retires *page_mask
and __get_user_pages()'s try_grab_folio() call and subpage loop.

Patch 7 walks every PTE in a page table in one follow_page_pte() call instead
of one per page.

Patch 8 adds follow_pte_batch() so a contiguous same-folio run is committed
with one refcount grab. This is the only patch whose benefit depends on folio
size; patch 7 alone covers plain base pages.

Patches 7 and 8 carry their own benchmark tables, both measured against the
base of the series, so the split between the two mechanisms is visible: the
single-call walk is worth 2.3x on base pages and 2.2x on 64 kB mTHP, and
refcount batching adds a further 5.9x on the mTHP case.

v3:
 - split up the series into 8 much smaller patches (David & Lorenzo)
 - shorten changelogs where they were too long (Lorenzo)
 - fix FOLL_WRITE folio dirtying by gathering dirty bits from all PTEs

Link: https://lore.kernel.org/r/20260730035350.1fc95dd8@fangorn/ [RFC]
Link: https://lore.kernel.org/r/[email protected]/ [RFC v2]
Suggested-by: David Hildenbrand <[email protected]>

Rik van Riel (8):
  mm/gup: break out gup_fill_pages() helper
  mm/gup: convert follow_page_mask() to return a long
  mm/gup: split follow_page_pte_commit() out of follow_page_pte()
  mm/gup: break out follow_one_pte() helper
  mm/gup: fill the pages array outside the pud/pmd lock
  mm/gup: return a huge page's full count from follow_page_mask()
  mm/gup: walk multiple PTEs per follow_page_pte() call
  mm/gup: batch contiguous same-folio PTEs into one refcount grab

 mm/gup.c | 515 +++++++++++++++++++++++++++++++++----------------------
 1 file changed, 308 insertions(+), 207 deletions(-)

-- 
2.55.0
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.