Re: [PATCH v5 10/11] mm: install PMD swap entries on swap-out

Usama Arif <[email protected]>
Newsgroups org.kvack.linux-mm,org.kernel.vger.linux-kernel
Message-ID <[email protected]>

On 06/08/2026 03:29, Luiz Capitulino wrote:
> On 2026-07-22 11:19, Usama Arif wrote:
>> Reclaim today splits a PMD-mapped anonymous THP into 512 PTE swap
>> entries before unmap, losing the huge mapping across the swap
>> round-trip and forcing khugepaged to rebuild it later. The contiguous
>> swap range was already secured when the folio was added to the swap
>> cache (a non-contiguous allocation would have split the folio earlier),
>> so the PMD can be replaced by a single PMD-level swap entry instead.
>>
>> This patch mirrors the existing PTE swap-out path at PMD granularity:
>> - shrink_folio_list() drops TTU_SPLIT_HUGE_PMD for PMD-mappable
>>    swapcache folios. zswap is handled by the PMD swap-in users: if any
>>    covered slot currently has a zswap entry, they split the PMD swap
>>    entry and fall back to the per-PTE path.
>> - try_to_unmap_one() now has a PMD branch that calls
>>    set_pmd_swap_entry() and adjusts MM_ANONPAGES / MM_SWAPENTS by
>>    HPAGE_PMD_NR before walk_done. TTU_SPLIT_HUGE_PMD remains the
>>    fallback.
>> - set_pmd_swap_entry() is the installer. Mirroring the PTE swap-out
>>    sequence at PMD granularity, it clears the present mapping (keeping
>>    the original for rollback), bumps the swap_map refcount for the
>>    folio's 512 slots, transfers the exclusive state in the swap entry,
>>    propagates the dirty bit to the folio so writeback is not lost,
>>    and installs a swap PMD that preserves the original
>>    soft-dirty / uffd-wp / exclusive bits. Any failing step rolls back
>>    the present mapping.
>>
>> The swap entry value matches what 512 PTE swap entries would encode, so
>> swap_map refcounting is unchanged: each of the 512 slots carries a
>> count of 1, released individually on later split or together on swap-in.
>>
>> Add thp_swpout_pmd to count each PMD mapping replaced by a PMD-level
>> swap entry. Unlike the folio-level thp_swpout counter, a fork-shared THP
>> can increment this counter once for each mapping; document that
>> distinction.
>>
>> Signed-off-by: Usama Arif <[email protected]>
>> ---
>>   Documentation/admin-guide/mm/transhuge.rst |  5 ++
>>   include/linux/huge_mm.h                    |  2 +
>>   include/linux/vm_event_item.h              |  1 +
>>   mm/huge_memory.c                           | 80 ++++++++++++++++++++++
>>   mm/rmap.c                                  | 19 +++++
>>   mm/vmscan.c                                |  9 ++-
>>   mm/vmstat.c                                |  1 +
>>   7 files changed, 116 insertions(+), 1 deletion(-)
>>
>> diff --git a/Documentation/admin-guide/mm/transhuge.rst b/Documentation/admin-guide/mm/transhuge.rst
>> index 16f37135ed80..b421d7982db6 100644
>> --- a/Documentation/admin-guide/mm/transhuge.rst
>> +++ b/Documentation/admin-guide/mm/transhuge.rst
>> @@ -630,6 +630,11 @@ thp_swpout
>>       is incremented every time a huge page is swapout in one
>>       piece without splitting.
>>   +thp_swpout_pmd
>> +    is incremented every time a PMD mapping is replaced by a PMD-level
>> +    swap entry. A fork-shared THP can increment this counter once for each
>> +    PMD mapping that is swapped out.
>> +
>>   thp_swpout_fallback
>>       is incremented if a huge page has to be split before swapout.
>>       Usually because failed to allocate some continuous swap space
>> diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h
>> index e7107e0991ad..41cf643a3f55 100644
>> --- a/include/linux/huge_mm.h
>> +++ b/include/linux/huge_mm.h
>> @@ -554,6 +554,8 @@ vm_fault_t do_huge_pmd_device_private(struct vm_fault *vmf);
>>     #ifdef CONFIG_THP_SWAP
>>   vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf);
>> +int set_pmd_swap_entry(struct page_vma_mapped_walk *pvmw,
>> +               struct folio *folio);
>>   #else
>>   static inline vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf)
>>   {
>> diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h
>> index 2628ccda076a..f8fd4e13698c 100644
>> --- a/include/linux/vm_event_item.h
>> +++ b/include/linux/vm_event_item.h
>> @@ -108,6 +108,7 @@ enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT,
>>           THP_ZERO_PAGE_ALLOC_FAILED,
>>           THP_SWPOUT,
>>           THP_SWPOUT_FALLBACK,
>> +        THP_SWPOUT_PMD,
>>   #endif
>>   #ifdef CONFIG_BALLOON
>>           BALLOON_INFLATE,
>> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
>> index 6ef56936ced4..c014631e5a26 100644
>> --- a/mm/huge_memory.c
>> +++ b/mm/huge_memory.c
>> @@ -5561,3 +5561,83 @@ void remove_migration_pmd(struct page_vma_mapped_walk *pvmw, struct page *new)
>>       trace_remove_migration_pmd(address, pmd_val(pmde));
>>   }
>>   #endif
>> +
>> +#ifdef CONFIG_THP_SWAP
>> +/**
>> + * set_pmd_swap_entry() - Replace a PMD mapping with a PMD-level swap entry.
>> + * @pvmw: Page vma mapped walk context, must have pvmw->pmd set and
>> + *        pvmw->pte NULL (i.e. PMD-mapped).
>> + * @folio: The folio being swapped out. Must be in the swap cache.
>> + *
>> + * This installs a PMD-level swap entry in place of a present PMD mapping,
>> + * avoiding the need to split the PMD into PTE-level swap entries.
>> + *
>> + * Return: 0 on success, negative error code on failure.
>> + */
>> +int set_pmd_swap_entry(struct page_vma_mapped_walk *pvmw,
>> +               struct folio *folio)
>> +{
>> +    struct vm_area_struct *vma = pvmw->vma;
>> +    struct mm_struct *mm = vma->vm_mm;
>> +    unsigned long address = pvmw->address;
>> +    unsigned long haddr = address & HPAGE_PMD_MASK;
>> +    struct page *page = folio_page(folio, 0);
>> +    bool anon_exclusive;
>> +    pmd_t pmdval;
>> +    swp_entry_t entry;
>> +    pmd_t pmdswp;
>> +
>> +    if (!(pvmw->pmd && !pvmw->pte))
>> +        return 0;
> 
> Should we call VM_WARN_ON_ONCE() and return an error instead? This
> function is only called by try_to_unmap_one() under this condition. In
> addition, returning zero would cause try_to_unmap_one() to assume that
> the PMD swap entry was installed, which is not the case.
> 
Thanks for review!

It makes sense, I will add a warning and return -EINVAL.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.