Re: [PATCH v4 03/17] mm/mm_init: skip initializing shared vmemmap tail pages
Muchun Song <[email protected]>
| Newsgroups | org.kvack.linux-mm,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On 2026/8/24 16:59, Mike Rapoport wrote: >> memmap_init_range() initializes every struct page in the target range. >> For compound pages with vmemmap optimization, the tail struct pages are >> backed by a shared vmemmap page. >> >> Initializing those tail struct pages would overwrite the shared >> vmemmap page contents, requiring users such as HugeTLB to restore the >> metadata afterwards. >> >> Track the compound order for HVO-backed sections and use that metadata >> to detect struct pages that fall into the shared tail vmemmap range. >> Skip those shared tail pages in memmap_init_range(), then initialize >> pageblock migratetypes for the processed range with a helper after the >> per-page initialization loop. >> >> Keep direct mem_section access inside sparse helpers by exposing >> pfn_to_section_order() to users that only need the order associated with >> a PFN. This lets memmap_init_range() skip shared tail vmemmap pages >> without exposing __pfn_to_section() to !SPARSEMEM builds. >> >> This is a preparatory change for consolidating handling across users of >> vmemmap optimization, and it also avoids redundant initialization of >> shared tail vmemmap pages during early boot. >> >> Signed-off-by: Muchun Song <[email protected]> >> >> diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h >> index 5fb9b37819d5..df31cac12311 100644 >> --- a/include/linux/mmzone.h >> +++ b/include/linux/mmzone.h >> @@ -2022,6 +2022,14 @@ struct mem_section { >> unsigned long section_mem_map; >> >> struct mem_section_usage *usage; >> +#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP >> + /* >> + * Normally, sections hold regular (order-0) pages. However, for >> + * sections with HVO enabled, this tracks the compound page order >> + * to enable deduplication of redundant vmemmap pages. >> + */ >> + unsigned int order; >> +#endif >> #ifdef CONFIG_PAGE_EXTENSION >> /* >> * If SPARSEMEM, pgdat doesn't have page_ext pointer. We use >> diff --git a/mm/mm_init.c b/mm/mm_init.c >> index 1533aebafb68..05c09e755e0b 100644 >> --- a/mm/mm_init.c >> +++ b/mm/mm_init.c >> @@ -29,6 +29,7 @@ >> #include <linux/cma.h> >> #include <linux/crash_dump.h> >> #include <linux/execmem.h> >> +#include <linux/sizes.h> >> #include <linux/vmstat.h> >> #include <linux/kexec_handover.h> >> #include <linux/hugetlb.h> >> @@ -677,21 +678,19 @@ static inline void fixup_hashdist(void) >> static inline void fixup_hashdist(void) {} >> #endif /* CONFIG_NUMA */ >> >> -#if defined(CONFIG_ZONE_DEVICE) || defined(CONFIG_DEFERRED_STRUCT_PAGE_INIT) >> static __meminit void pageblock_migratetype_init_range(unsigned long pfn, >> - unsigned long nr_pages, int migratetype, bool atomic) >> + unsigned long nr_pages, int migratetype, bool isolate, bool atomic) > Growing boolean flags makes the callsites harder to read. > One way to deal with it is to add comments to the callers saying what > each true and false mean. Make sense. > >> { >> const unsigned long end = pfn + nr_pages; >> >> for (pfn = pageblock_align(pfn); pfn < end; pfn += pageblock_nr_pages) { >> enum migratetype mt = kho_scratch_migratetype(pfn, migratetype); >> >> - init_pageblock_migratetype(pfn_to_page(pfn), mt, false); >> - if (!atomic && IS_ALIGNED(pfn, PAGES_PER_SECTION)) >> + init_pageblock_migratetype(pfn_to_page(pfn), mt, isolate); >> + if (!atomic && IS_ALIGNED(pfn, PFN_DOWN(SZ_1G))) >> cond_resched(); >> } >> } >> -#endif >> >> #ifdef CONFIG_DEFERRED_STRUCT_PAGE_INIT >> static inline void pgdat_set_deferred_range(pg_data_t *pgdat) >> @@ -886,6 +885,13 @@ void __meminit memmap_init_range(unsigned long size, int nid, unsigned long zone >> } >> } >> >> + if (vmemmap_optimizable_pfn(pfn)) { > A short comment above would be nice :) No problem. Thanks for your review. Muchun, Thanks.