Re: [PATCH v3 08/17] mm/sparse-vmemmap: support section-based vmemmap optimization

Muchun Song <[email protected]>
Newsgroups org.kvack.linux-mm,org.kernel.vger.linux-kernel
Message-ID <[email protected]>

On 2026/8/6 11:04, Muchun Song wrote:
>
>
> On 2026/8/4 11:55, Muchun Song wrote:
>> Teach sparse-vmemmap population code to use the compound page order
>> when deciding whether a vmemmap page can be optimized.
>>
>> With this information, the common sparse-vmemmap population path can
>> allocate or reuse shared tail vmemmap pages directly instead of relying
>> on HugeTLB-specific handling.
>>
>> This centralizes vmemmap optimization logic in the sparse-vmemmap code,
>> based on section metadata, and prepares for sharing the same mechanism
>> across different users of vmemmap optimization, including HugeTLB and
>> DAX.
>>
>> Signed-off-by: Muchun Song <[email protected]>
>> ---
>> v2:
>> - Keep vmemmap accounting and population logic in sparse-vmemmap.c
>>    (suggested by Mike Rapoport)
>> - Move vmemmap_get_tail() before its first use instead of adding only a
>>    forward declaration in the previous patch (suggested by Mike 
>> Rapoport)
>> - Simplify the PMD path handling for HVO-covered sections
>> ---
>>   mm/sparse-vmemmap.c | 36 ++++++++++++++++++++++++++++++------
>>   mm/sparse.c         |  4 ++--
>>   mm/sparse.h         |  7 +++++++
>>   3 files changed, 39 insertions(+), 8 deletions(-)
>>
>> diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
>> index b770fe2428fd..b69a7af76858 100644
>> --- a/mm/sparse-vmemmap.c
>> +++ b/mm/sparse-vmemmap.c
>> @@ -186,6 +186,11 @@ static __meminit struct page 
>> *vmemmap_get_tail(unsigned int order, struct zone *
>>         return tail;
>>   }
>> +#else
>> +static inline struct page *vmemmap_get_tail(unsigned int order, 
>> struct zone *zone)
>> +{
>> +    return NULL;
>> +}
>>   #endif
>>     static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, 
>> unsigned long addr, int node,
>> @@ -193,12 +198,24 @@ static pte_t * __meminit 
>> vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
>>                          unsigned long ptpfn, unsigned long flags)
>>   {
>>       pte_t *pte = pte_offset_kernel(pmd, addr);
>> +    unsigned long pfn = page_to_pfn((struct page *)addr);
>> +
>>       if (pte_none(ptep_get(pte))) {
>>           pte_t entry;
>> -        void *p;
>> +
>> +        if (pfn_vmemmap_optimizable(pfn) && ptpfn == (unsigned 
>> long)-1) {
>> +            unsigned int order = pfn_to_section_order(pfn);
>> +            struct zone *zone = pfn_to_zone(pfn, node);
>> +            struct page *page = vmemmap_get_tail(order, zone);
>> +
>> +            if (!page)
>> +                return NULL;
>> +            ptpfn = page_to_pfn(page);
>> +        }
>>             if (ptpfn == (unsigned long)-1) {
>> -            p = vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
>> +            void *p = vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
>> +
>>               if (!p)
>>                   return NULL;
>>               ptpfn = PHYS_PFN(__pa(p));
>> @@ -217,7 +234,8 @@ static pte_t * __meminit 
>> vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
>>           }
>>           entry = pfn_pte(ptpfn, PAGE_KERNEL);
>>           set_pte_at(&init_mm, addr, pte, entry);
>> -    }
>> +    } else if (WARN_ON_ONCE(pfn_vmemmap_optimizable(pfn)))
>> +        return NULL;
>>       return pte;
>>   }
>>   @@ -406,6 +424,9 @@ int __meminit 
>> vmemmap_populate_hugepages(unsigned long start, unsigned long end,
>>       pmd_t *pmd;
>>         for (addr = start; addr < end; addr = next) {
>> +        unsigned long pfn = page_to_pfn((struct page *)addr);
>> +        const struct mem_section *ms = __pfn_to_section(pfn);
>> +
>>           next = pmd_addr_end(addr, end);
>>             pgd = vmemmap_pgd_populate(addr, node);
>> @@ -421,7 +442,7 @@ int __meminit vmemmap_populate_hugepages(unsigned 
>> long start, unsigned long end,
>>               return -ENOMEM;
>>             pmd = pmd_offset(pud, addr);
>> -        if (pmd_none(pmdp_get(pmd))) {
>> +        if (pmd_none(pmdp_get(pmd)) && 
>> !section_vmemmap_optimizable(ms)) {
>>               void *p;
>>                 p = vmemmap_alloc_block_buf(PMD_SIZE, node, altmap);
>> @@ -439,8 +460,11 @@ int __meminit 
>> vmemmap_populate_hugepages(unsigned long start, unsigned long end,
>>                    */
>>                   return -ENOMEM;
>>               }
>> -        } else if (vmemmap_check_pmd(pmd, node, addr, next))
>> +        } else if (vmemmap_check_pmd(pmd, node, addr, next)) {
>> +            if (WARN_ON_ONCE(section_vmemmap_optimizable(ms)))
>> +                return -ENOTSUPP;
>>               continue;
>> +        }
>>           if (vmemmap_populate_basepages(addr, next, node, altmap))
>>               return -ENOMEM;
>>       }
>> @@ -648,7 +672,7 @@ void offline_mem_sections(unsigned long 
>> start_pfn, unsigned long end_pfn)
>>       }
>>   }
>>   -static int __meminit section_nr_vmemmap_pages(unsigned long pfn, 
>> unsigned long nr_pages,
>> +int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned 
>> long nr_pages,
>
> The kernel test robot reported a compilation issue: when
> CONFIG_MEMORY_HOTPLUG = n && CONFIG_SPARSEMEM_VMEMMAP=y,
> section_nr_vmemmap_pages is undefined. This problem is easy to fix, and
> I will move the entire function outside the CONFIG_MEMORY_HOTPLUG guard
> in the next version. 

Hi,

I will wait a few more days to see if there are any reviews. If there is
no further input during this period, I will update a new version to resolve
this compilation issue.

Muchun,
Thanks
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.