Re: [PATCH v5 10/11] mm: install PMD swap entries on swap-out
Luiz Capitulino <[email protected]>
| Newsgroups | gmane.linux.kernel,gmane.linux.kernel.mm |
|---|---|
| Message-ID | <[email protected]> |
On 2026-07-22 11:19, Usama Arif wrote: > Reclaim today splits a PMD-mapped anonymous THP into 512 PTE swap > entries before unmap, losing the huge mapping across the swap > round-trip and forcing khugepaged to rebuild it later. The contiguous > swap range was already secured when the folio was added to the swap > cache (a non-contiguous allocation would have split the folio earlier), > so the PMD can be replaced by a single PMD-level swap entry instead. > > This patch mirrors the existing PTE swap-out path at PMD granularity: > - shrink_folio_list() drops TTU_SPLIT_HUGE_PMD for PMD-mappable > swapcache folios. zswap is handled by the PMD swap-in users: if any > covered slot currently has a zswap entry, they split the PMD swap > entry and fall back to the per-PTE path. > - try_to_unmap_one() now has a PMD branch that calls > set_pmd_swap_entry() and adjusts MM_ANONPAGES / MM_SWAPENTS by > HPAGE_PMD_NR before walk_done. TTU_SPLIT_HUGE_PMD remains the > fallback. > - set_pmd_swap_entry() is the installer. Mirroring the PTE swap-out > sequence at PMD granularity, it clears the present mapping (keeping > the original for rollback), bumps the swap_map refcount for the > folio's 512 slots, transfers the exclusive state in the swap entry, > propagates the dirty bit to the folio so writeback is not lost, > and installs a swap PMD that preserves the original > soft-dirty / uffd-wp / exclusive bits. Any failing step rolls back > the present mapping. > > The swap entry value matches what 512 PTE swap entries would encode, so > swap_map refcounting is unchanged: each of the 512 slots carries a > count of 1, released individually on later split or together on swap-in. > > Add thp_swpout_pmd to count each PMD mapping replaced by a PMD-level > swap entry. Unlike the folio-level thp_swpout counter, a fork-shared THP > can increment this counter once for each mapping; document that > distinction. > > Signed-off-by: Usama Arif <[email protected]> > --- > Documentation/admin-guide/mm/transhuge.rst | 5 ++ > include/linux/huge_mm.h | 2 + > include/linux/vm_event_item.h | 1 + > mm/huge_memory.c | 80 ++++++++++++++++++++++ > mm/rmap.c | 19 +++++ > mm/vmscan.c | 9 ++- > mm/vmstat.c | 1 + > 7 files changed, 116 insertions(+), 1 deletion(-) > > diff --git a/Documentation/admin-guide/mm/transhuge.rst b/Documentation/admin-guide/mm/transhuge.rst > index 16f37135ed80..b421d7982db6 100644 > --- a/Documentation/admin-guide/mm/transhuge.rst > +++ b/Documentation/admin-guide/mm/transhuge.rst > @@ -630,6 +630,11 @@ thp_swpout > is incremented every time a huge page is swapout in one > piece without splitting. > > +thp_swpout_pmd > + is incremented every time a PMD mapping is replaced by a PMD-level > + swap entry. A fork-shared THP can increment this counter once for each > + PMD mapping that is swapped out. > + > thp_swpout_fallback > is incremented if a huge page has to be split before swapout. > Usually because failed to allocate some continuous swap space > diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h > index e7107e0991ad..41cf643a3f55 100644 > --- a/include/linux/huge_mm.h > +++ b/include/linux/huge_mm.h > @@ -554,6 +554,8 @@ vm_fault_t do_huge_pmd_device_private(struct vm_fault *vmf); > > #ifdef CONFIG_THP_SWAP > vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf); > +int set_pmd_swap_entry(struct page_vma_mapped_walk *pvmw, > + struct folio *folio); > #else > static inline vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf) > { > diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h > index 2628ccda076a..f8fd4e13698c 100644 > --- a/include/linux/vm_event_item.h > +++ b/include/linux/vm_event_item.h > @@ -108,6 +108,7 @@ enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT, > THP_ZERO_PAGE_ALLOC_FAILED, > THP_SWPOUT, > THP_SWPOUT_FALLBACK, > + THP_SWPOUT_PMD, > #endif > #ifdef CONFIG_BALLOON > BALLOON_INFLATE, > diff --git a/mm/huge_memory.c b/mm/huge_memory.c > index 6ef56936ced4..c014631e5a26 100644 > --- a/mm/huge_memory.c > +++ b/mm/huge_memory.c > @@ -5561,3 +5561,83 @@ void remove_migration_pmd(struct page_vma_mapped_walk *pvmw, struct page *new) > trace_remove_migration_pmd(address, pmd_val(pmde)); > } > #endif > + > +#ifdef CONFIG_THP_SWAP > +/** > + * set_pmd_swap_entry() - Replace a PMD mapping with a PMD-level swap entry. > + * @pvmw: Page vma mapped walk context, must have pvmw->pmd set and > + * pvmw->pte NULL (i.e. PMD-mapped). > + * @folio: The folio being swapped out. Must be in the swap cache. > + * > + * This installs a PMD-level swap entry in place of a present PMD mapping, > + * avoiding the need to split the PMD into PTE-level swap entries. > + * > + * Return: 0 on success, negative error code on failure. > + */ > +int set_pmd_swap_entry(struct page_vma_mapped_walk *pvmw, > + struct folio *folio) > +{ > + struct vm_area_struct *vma = pvmw->vma; > + struct mm_struct *mm = vma->vm_mm; > + unsigned long address = pvmw->address; > + unsigned long haddr = address & HPAGE_PMD_MASK; > + struct page *page = folio_page(folio, 0); > + bool anon_exclusive; > + pmd_t pmdval; > + swp_entry_t entry; > + pmd_t pmdswp; > + > + if (!(pvmw->pmd && !pvmw->pte)) > + return 0; Should we call VM_WARN_ON_ONCE() and return an error instead? This function is only called by try_to_unmap_one() under this condition. In addition, returning zero would cause try_to_unmap_one() to assume that the PMD swap entry was installed, which is not the case. > + > + VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio); > + VM_BUG_ON_FOLIO(!folio_test_anon(folio), folio); > + > + if (unlikely(folio_test_swapbacked(folio) != > + folio_test_swapcache(folio))) { > + WARN_ON_ONCE(1); > + return -EBUSY; > + } > + > + flush_cache_range(vma, haddr, haddr + HPAGE_PMD_SIZE); > + > + pmdval = pmdp_invalidate(vma, haddr, pvmw->pmd); > + > + /* Update high watermark before we lower rss */ > + update_hiwater_rss(mm); > + > + if (folio_dup_swap(folio, NULL) < 0) { > + set_pmd_at(mm, haddr, pvmw->pmd, pmdval); > + return -ENOMEM; > + } > + > + /* See folio_try_share_anon_rmap_pmd(): invalidate PMD first. */ > + anon_exclusive = PageAnonExclusive(page); > + if (anon_exclusive && folio_try_share_anon_rmap_pmd(folio, page)) { > + folio_put_swap(folio, NULL); > + set_pmd_at(mm, haddr, pvmw->pmd, pmdval); > + return -EBUSY; > + } > + > + mm_prepare_for_swap_entries(mm); > + > + if (pmd_dirty(pmdval)) > + folio_mark_dirty(folio); > + > + entry = folio->swap; > + pmdswp = softleaf_to_pmd(entry); > + if (pmd_soft_dirty(pmdval)) > + pmdswp = pmd_swp_mksoft_dirty(pmdswp); > + if (pmd_uffd(pmdval)) > + pmdswp = pmd_swp_mkuffd(pmdswp); > + if (anon_exclusive) > + pmdswp = pmd_swp_mkexclusive(pmdswp); > + set_pmd_at(mm, haddr, pvmw->pmd, pmdswp); > + > + folio_remove_rmap_pmd(folio, page, vma); > + folio_put(folio); > + > + count_vm_event(THP_SWPOUT_PMD); > + return 0; > +} > +#endif /* CONFIG_THP_SWAP */ > diff --git a/mm/rmap.c b/mm/rmap.c > index b7ead3e9f064..4574f7b969b6 100644 > --- a/mm/rmap.c > +++ b/mm/rmap.c > @@ -2282,6 +2282,25 @@ static bool try_to_unmap_one(struct folio *folio, struct vm_area_struct *vma, > goto walk_abort; > } > > +#ifdef CONFIG_THP_SWAP > + /* > + * If the folio is in the swap cache and we're not > + * asked to split, install a PMD-level swap entry. > + */ > + if (!(flags & TTU_SPLIT_HUGE_PMD) && > + folio_test_anon(folio) && > + folio_test_swapcache(folio)) { > + if (set_pmd_swap_entry(&pvmw, folio)) > + goto walk_abort; > + > + add_mm_counter(mm, MM_ANONPAGES, > + -HPAGE_PMD_NR); > + add_mm_counter(mm, MM_SWAPENTS, > + HPAGE_PMD_NR); > + goto walk_done; > + } > +#endif > + > if (flags & TTU_SPLIT_HUGE_PMD) { > /* > * We temporarily have to drop the PTL and > diff --git a/mm/vmscan.c b/mm/vmscan.c > index 9b8ee902f972..83bb34403e42 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -1333,7 +1333,14 @@ static unsigned int shrink_folio_list(struct list_head *folio_list, > enum ttu_flags flags = TTU_BATCH_FLUSH; > bool was_swapbacked = folio_test_swapbacked(folio); > > - if (folio_test_pmd_mappable(folio)) > + /* > + * With THP_SWAP, PMD-mappable folios already in the > + * swap cache can be unmapped with a PMD-level swap > + * entry, avoiding the cost of splitting the PMD. > + */ > + if (folio_test_pmd_mappable(folio) && > + !(IS_ENABLED(CONFIG_THP_SWAP) && > + folio_test_swapcache(folio))) > flags |= TTU_SPLIT_HUGE_PMD; > /* > * Without TTU_SYNC, try_to_unmap will only begin to > diff --git a/mm/vmstat.c b/mm/vmstat.c > index 4e26e5fd6666..68e6efc0dabd 100644 > --- a/mm/vmstat.c > +++ b/mm/vmstat.c > @@ -1422,6 +1422,7 @@ const char * const vmstat_text[] = { > [I(THP_ZERO_PAGE_ALLOC_FAILED)] = "thp_zero_page_alloc_failed", > [I(THP_SWPOUT)] = "thp_swpout", > [I(THP_SWPOUT_FALLBACK)] = "thp_swpout_fallback", > + [I(THP_SWPOUT_PMD)] = "thp_swpout_pmd", > #endif > #ifdef CONFIG_BALLOON > [I(BALLOON_INFLATE)] = "balloon_inflate",