Re: [PATCH v5 10/11] mm: install PMD swap entries on swap-out
Usama Arif <[email protected]>
| Newsgroups | org.kvack.linux-mm,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On 06/08/2026 03:29, Luiz Capitulino wrote: > On 2026-07-22 11:19, Usama Arif wrote: >> Reclaim today splits a PMD-mapped anonymous THP into 512 PTE swap >> entries before unmap, losing the huge mapping across the swap >> round-trip and forcing khugepaged to rebuild it later. The contiguous >> swap range was already secured when the folio was added to the swap >> cache (a non-contiguous allocation would have split the folio earlier), >> so the PMD can be replaced by a single PMD-level swap entry instead. >> >> This patch mirrors the existing PTE swap-out path at PMD granularity: >> - shrink_folio_list() drops TTU_SPLIT_HUGE_PMD for PMD-mappable >> swapcache folios. zswap is handled by the PMD swap-in users: if any >> covered slot currently has a zswap entry, they split the PMD swap >> entry and fall back to the per-PTE path. >> - try_to_unmap_one() now has a PMD branch that calls >> set_pmd_swap_entry() and adjusts MM_ANONPAGES / MM_SWAPENTS by >> HPAGE_PMD_NR before walk_done. TTU_SPLIT_HUGE_PMD remains the >> fallback. >> - set_pmd_swap_entry() is the installer. Mirroring the PTE swap-out >> sequence at PMD granularity, it clears the present mapping (keeping >> the original for rollback), bumps the swap_map refcount for the >> folio's 512 slots, transfers the exclusive state in the swap entry, >> propagates the dirty bit to the folio so writeback is not lost, >> and installs a swap PMD that preserves the original >> soft-dirty / uffd-wp / exclusive bits. Any failing step rolls back >> the present mapping. >> >> The swap entry value matches what 512 PTE swap entries would encode, so >> swap_map refcounting is unchanged: each of the 512 slots carries a >> count of 1, released individually on later split or together on swap-in. >> >> Add thp_swpout_pmd to count each PMD mapping replaced by a PMD-level >> swap entry. Unlike the folio-level thp_swpout counter, a fork-shared THP >> can increment this counter once for each mapping; document that >> distinction. >> >> Signed-off-by: Usama Arif <[email protected]> >> --- >> Documentation/admin-guide/mm/transhuge.rst | 5 ++ >> include/linux/huge_mm.h | 2 + >> include/linux/vm_event_item.h | 1 + >> mm/huge_memory.c | 80 ++++++++++++++++++++++ >> mm/rmap.c | 19 +++++ >> mm/vmscan.c | 9 ++- >> mm/vmstat.c | 1 + >> 7 files changed, 116 insertions(+), 1 deletion(-) >> >> diff --git a/Documentation/admin-guide/mm/transhuge.rst b/Documentation/admin-guide/mm/transhuge.rst >> index 16f37135ed80..b421d7982db6 100644 >> --- a/Documentation/admin-guide/mm/transhuge.rst >> +++ b/Documentation/admin-guide/mm/transhuge.rst >> @@ -630,6 +630,11 @@ thp_swpout >> is incremented every time a huge page is swapout in one >> piece without splitting. >> +thp_swpout_pmd >> + is incremented every time a PMD mapping is replaced by a PMD-level >> + swap entry. A fork-shared THP can increment this counter once for each >> + PMD mapping that is swapped out. >> + >> thp_swpout_fallback >> is incremented if a huge page has to be split before swapout. >> Usually because failed to allocate some continuous swap space >> diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h >> index e7107e0991ad..41cf643a3f55 100644 >> --- a/include/linux/huge_mm.h >> +++ b/include/linux/huge_mm.h >> @@ -554,6 +554,8 @@ vm_fault_t do_huge_pmd_device_private(struct vm_fault *vmf); >> #ifdef CONFIG_THP_SWAP >> vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf); >> +int set_pmd_swap_entry(struct page_vma_mapped_walk *pvmw, >> + struct folio *folio); >> #else >> static inline vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf) >> { >> diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h >> index 2628ccda076a..f8fd4e13698c 100644 >> --- a/include/linux/vm_event_item.h >> +++ b/include/linux/vm_event_item.h >> @@ -108,6 +108,7 @@ enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT, >> THP_ZERO_PAGE_ALLOC_FAILED, >> THP_SWPOUT, >> THP_SWPOUT_FALLBACK, >> + THP_SWPOUT_PMD, >> #endif >> #ifdef CONFIG_BALLOON >> BALLOON_INFLATE, >> diff --git a/mm/huge_memory.c b/mm/huge_memory.c >> index 6ef56936ced4..c014631e5a26 100644 >> --- a/mm/huge_memory.c >> +++ b/mm/huge_memory.c >> @@ -5561,3 +5561,83 @@ void remove_migration_pmd(struct page_vma_mapped_walk *pvmw, struct page *new) >> trace_remove_migration_pmd(address, pmd_val(pmde)); >> } >> #endif >> + >> +#ifdef CONFIG_THP_SWAP >> +/** >> + * set_pmd_swap_entry() - Replace a PMD mapping with a PMD-level swap entry. >> + * @pvmw: Page vma mapped walk context, must have pvmw->pmd set and >> + * pvmw->pte NULL (i.e. PMD-mapped). >> + * @folio: The folio being swapped out. Must be in the swap cache. >> + * >> + * This installs a PMD-level swap entry in place of a present PMD mapping, >> + * avoiding the need to split the PMD into PTE-level swap entries. >> + * >> + * Return: 0 on success, negative error code on failure. >> + */ >> +int set_pmd_swap_entry(struct page_vma_mapped_walk *pvmw, >> + struct folio *folio) >> +{ >> + struct vm_area_struct *vma = pvmw->vma; >> + struct mm_struct *mm = vma->vm_mm; >> + unsigned long address = pvmw->address; >> + unsigned long haddr = address & HPAGE_PMD_MASK; >> + struct page *page = folio_page(folio, 0); >> + bool anon_exclusive; >> + pmd_t pmdval; >> + swp_entry_t entry; >> + pmd_t pmdswp; >> + >> + if (!(pvmw->pmd && !pvmw->pte)) >> + return 0; > > Should we call VM_WARN_ON_ONCE() and return an error instead? This > function is only called by try_to_unmap_one() under this condition. In > addition, returning zero would cause try_to_unmap_one() to assume that > the PMD swap entry was installed, which is not the case. > Thanks for review! It makes sense, I will add a warning and return -EINVAL.