Re: [PATCH] mm/huge_memory: let special huge VMAs bypass the THP policy check
Cédric Le Goater <[email protected]>
| Newsgroups | gmane.linux.kernel,gmane.linux.kernel.mm,gmane.linux.kernel.stable |
|---|---|
| Message-ID | <[email protected]> |
On 8/6/26 16:28, Lorenzo Stoakes (ARM) wrote: > On Thu, Aug 06, 2026 at 04:26:20PM +0200, David Hildenbrand (Arm) wrote: >> On 8/5/26 07:55, Cédric Le Goater wrote: >>> From: Cedric Le Goater <[email protected]> >>> >>> The global THP sysfs policy (transparent_hugepage=never/madvise/always) >>> gates the huge fault dispatch path in __thp_vma_allowable_orders() for >>> all non-anonymous VMAs, including PFN-mapped device BARs (VM_PFNMAP). >>> >>> DAX VMAs already bypass this check via an early return: >>> >>> if (vma_is_dax(vma)) >>> return in_pf ? orders : 0; >>> >>> But "special huge" VMAs -- identified by vma_is_special_huge() -- do not >>> get this early return, even though they share the same fundamental >>> property: they map physical addresses directly into page tables and >>> involve no memory allocation, no compaction, no splitting, and no >>> reclaim. The THP policy has no meaningful effect on them. >>> >>> This matters for VFIO PCI passthrough of large-BAR devices such as >>> NVIDIA H200 NVL GPUs (256 GB BAR each). The VFIO driver registers a >>> .huge_fault handler (vfio_pci_mmap_huge_fault) that dispatches to >>> vmf_insert_pfn_pmd/pud, and QEMU's vfio_region_mmap() aligns the BAR >>> mappings for huge page table entries. Both prerequisites are met, but >>> with THP=never or THP=madvise, __thp_vma_allowable_orders() returns 0 >>> before reaching the "trust huge_fault handlers" code. >>> >>> The result: each 256 GB BAR is mapped at 4 KiB granularity -- 67 million >>> page faults per GPU instead of a few thousand PMD/PUD faults. On hosts >>> with 8 GPUs (2 TB of BAR space), this causes VM boot times to degrade >>> severely, with 99.98% of CPU time spent in the VFIO BAR mapping path. >>> >>> Configurations that trigger this: >>> - transparent_hugepage=never on the kernel command line >>> - The tuned cpu-partitioning profile (inherits network-latency, which >>> sets transparent_hugepages=never via sysfs) >>> - transparent_hugepage=madvise (the RHEL default), since VFIO VMAs >>> lack VM_HUGEPAGE and QEMU does not call madvise(MADV_HUGEPAGE) on >>> BAR mmap regions >>> >>> Extend the existing DAX early return to also cover vma_is_special_huge() >>> VMAs. This is consistent with how vma_is_special_huge() is already >>> treated for supported_orders (grouped with DAX). The mm/Kconfig TODO >>> comment "Allow to be enabled without THP" also acknowledges this >>> coupling is wrong. >>> >>> Cc: Peter Xu <[email protected]> >>> Cc: Andrew Morton <[email protected]> >>> Cc: Lorenzo Stoakes <[email protected]> >>> Cc: David Hildenbrand <[email protected]> >>> Cc: Alex Williamson <[email protected]> >>> Cc: Jason Gunthorpe <[email protected]> >>> Cc: Zi Yan <[email protected]> >>> Fixes: 5dd40721f147 ("mm: allow THP orders for PFNMAPs") >>> Cc: [email protected] >>> Assisted-by: Claude:claude-opus-4 >>> Signed-off-by: Cedric Le Goater <[email protected]> >>> --- >>> mm/huge_memory.c | 8 ++++++-- >>> 1 file changed, 6 insertions(+), 2 deletions(-) >>> >>> diff --git a/mm/huge_memory.c b/mm/huge_memory.c >>> index 58cabe6af33d031e48250e21db51506bc46c97b2..6dfef5500a054f09f9ece6df8bf7a0194624350f 100644 >>> --- a/mm/huge_memory.c >>> +++ b/mm/huge_memory.c >>> @@ -139,8 +139,12 @@ unsigned long __thp_vma_allowable_orders(struct vm_area_struct *vma, >>> if (thp_disabled_by_hw() || vma_thp_disabled(vma, vm_flags, forced_collapse)) >>> return 0; >>> >>> - /* khugepaged doesn't collapse DAX vma, but page fault is fine. */ >>> - if (vma_is_dax(vma)) >>> + /* >>> + * khugepaged doesn't collapse DAX or special huge VMAs, but page >>> + * fault is fine. These map physical addresses directly — the THP >>> + * policy is irrelevant for them. >> >> emdash in a code comment? There are quite a few of these in the code in fact. >> Then I spot >> >> Assisted-by: Claude:claude-opus-4 >> >> and really have to shake my head. It's a one liner. Good enough to raise the discussion no ? > Ha, I missed that! Me too. But, you have more to add to it anyway. It should be dropped. > Well all the more reason for me to take over this patch... :) > > At least the AI is acked here (appreciate that at least Cedric). Yeah. Let's be honest. Even if I understand what is going on, Claude was faster at digging through the code and connecting the dots than I would have been on my own. The AI behemoth dropped my accent though. Cédric it should have been. I wonder why. > I think the _actual issue_ is valid at least. The huge pfn stuff did seem to > completely miss this aspect of things. It's a real problem indeed and to cover mix of workloads, it it difficult to address without a kernel patch. Cheers, C.