Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio

Anshuman Khandual <[email protected]>
Newsgroups gmane.linux.kernel,gmane.linux.kernel.mm
Message-ID <2sipzzq7amhd5kn4xbqmtzc25iud2zp76dhpefcujqqpcnpqrt@czyz3dag6hlp>
On Thu, Aug 06, 2026 at 09:11:29PM -0700, Sourav Panda wrote:
> On Thu, Aug 6, 2026 at 4:07 AM Usama Arif <[email protected]> wrote:
> >
> >
> >
> > On 06/08/2026 06:33, Sourav Panda wrote:
> > > On Wed, Aug 5, 2026 at 7:47 PM Anshuman Khandual
> > > <[email protected]> wrote:
> > >>
> > >> On Wed, Aug 05, 2026 at 01:24:42PM -0700, Usama Arif wrote:
> > >>> On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <[email protected]> wrote:
> > >>>
> > >>>> alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> > >>>> alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> > >>>> nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> > >>>> the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> > >>>> alloc_gigantic_frozen_folio().
> > >>>>
> > >>>> hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> > >>>> node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> > >>>> leading to a null pointer dereference kernel panic.
> > >>>>
> > >>>> Fix this by checking if nodemask is NULL in
> > >>>> hugetlb_cma_alloc_frozen_folio() and defaulting it to
> > >>>> node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> > >>>> back to any node with memory, keeping behavior consistent with
> > >>>> alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
> > >>>>
> > >>>> >From a userspace perspective, this bug allows an unprivileged user to
> > >>>> crash the kernel (trigger a panic) by requesting a gigantic hugepage
> > >>>> allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> > >>>> configured on a subset of NUMA nodes.
> > >>>>
> > >>>> This can be reproduced by booting a VM with two NUMA nodes, restricting
> > >>>> CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> > >>>> hugepagesz=1G hugepages=0), and running a program that allocates a
> > >>>> 1GB hugepage area without reserving, restricts allocation to Node 0
> > >>>> using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
> > >>>>
> > >>>>   void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> > >>>>                    MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> > >>>>                    MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> > >>>>   unsigned long nodemask = 1; /* Node 0 */
> > >>>>   mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> > >>>>         sizeof(nodemask) * 8, 0);
> > >>>>   memset(ptr, 0, 1UL << 30); /* Trigger fault */
> > >>>>
> > >>>> This results in a NULL pointer dereference:
> > >>>>
> > >>>>   BUG: kernel NULL pointer dereference, address: 0000000000000000
> > >>>>   #PF: supervisor read access in kernel mode
> > >>>>   #PF: error_code(0x0000) - not-present page
> > >>>>   Oops: Oops: 0000 [#1] SMP NOPTI
> > >>>>   RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> > >>>>   Call Trace:
> > >>>>    <TASK>
> > >>>>    only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> > >>>>    alloc_surplus_hugetlb_folio+0x6d/0x100
> > >>>>    alloc_hugetlb_folio+0x3c5/0x660
> > >>>>    hugetlb_no_page+0x3d9/0x650
> > >>>>
> > >>>> Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
> > >>>> Cc: [email protected]
> > >>>> Signed-off-by: Sourav Panda <[email protected]>
> > >>>> ---
> > >>>> Changes in v4:
> > >>>> - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
> > >>>>   As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
> > >>>>   not prevent false-positive allocation failures without complex retry loops,
> > >>>>   and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
> > >>>>   NULL nodemasks safely internally.
> > >>>> - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
> > >>>>   by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
> > >>>>   HugeTLB allocators clean and consistent.
> > >>>> - v3: https://lore.kernel.org/linux-mm/[email protected]/
> > >>>> - v2: https://lore.kernel.org/linux-mm/[email protected]/
> > >>>> - v1: https://lore.kernel.org/linux-mm/[email protected]/
> > >>>>
> > >>>>  mm/hugetlb_cma.c | 5 ++++-
> > >>>>  1 file changed, 4 insertions(+), 1 deletion(-)
> > >>>>
> > >>>> diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
> > >>>> index 39344d6c78d8..5744de0ceeb7 100644
> > >>>> --- a/mm/hugetlb_cma.c
> > >>>> +++ b/mm/hugetlb_cma.c
> > >>>> @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> > >>>>     if (!hugetlb_cma_size)
> > >>>>             return NULL;
> > >>>>
> > >>>> -   if (hugetlb_cma[nid])
> > >>>> +   if (!nodemask)
> > >>>> +           nodemask = &node_states[N_MEMORY];
> > >>>
> > >>> Hi Sourav,
> > >>>
> > >>> hmm what if cpuset.mems only allows allocation from node 0, and hugetlb CMA
> > >>> is only available on node 1. This will no now allow allocating a gigantic page
> > >>> on node 1 when its explicitly not allowed?
> > >>
> > >> That's fair point but should not the user be also responsible in provding a right
> > >> nodemask containing CMA memory if it prefers avoiding node_states[N_MEMORY] based
> > >> fallback mechanism in kernel ?
> > >>
> >
> > I don't think the task can be made responsible for providing a CMA-aware fallback mask
> > here.  Userspace supplies the MPOL_PREFERRED_MANY preference, but the kernel itself
> > changes the second attempt to a NULL nodemask.  That permits fallback outside the
> > preferred policy nodes; it does not permit fallback outside the task's hardwall cpuset.
> >
> > >
> > > We have precedence for three options, which makes selecting one difficult :)
> > >
> > > 1) nodemask = &cpuset_current_mems_allowed is currently being used in:
> > >        only_alloc_fresh_hugetlb_folio --> alloc_buddy_frozen_folio -->
> > > __alloc_frozen_pages_noprof --> prepare_alloc_pages.
> >
> > Here I think using &cpuset_current_mems_allowed is part of the page allocator's wider cpuset handling.
> >
> > > 2) Some places use cookies with cpuset_current_mems_allowed to prevent
> > > torn writes. But this adds complexity.
> > >        An example would be dequeue_hugetlb_folio_nodemask()
> > > 3) node_states[N_MEMORY] is sprinkled all over the kernel (simplest).
> > >
> >
> > I would prefer option 2.  A NULL mempolicy mask should fall back to
> > cpuset_current_mems_allowed, not all N_MEMORY nodes.
> >
> 
> Sounds good Usama :)
> 
> I will send a patch that replaces nodemask = &node_states[N_MEMORY];
> with the below:
> 
> + unsigned int cpuset_mems_cookie;
> +
> + do {
> + cpuset_mems_cookie = read_mems_allowed_begin();
> + local_node_mask = cpuset_current_mems_allowed;
> + } while (read_mems_allowed_retry(cpuset_mems_cookie));
> +
> + nodemask = &local_node_mask;

Although complexity increases with the above mechanism but probably
it might be better in terms of being compliant with task's hardwall
cpuset as explained by Usama earlier.

> 
> Will send it out tomorrow! Let me know if anyone has any objections.

I would say let's wait for a day or two before sending the respin.

> 
> >
> > >>>
> > >>> Would it not be better to check cpuset as well?
> > >>>
> > >>> Thanks,
> > >>> Usama
> > >>>
> > >>>> +
> > >>>> +   if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> > >>>>             page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
> > >>>>
> > >>>>     if (!page && !(gfp_mask & __GFP_THISNODE)) {
> > >>>> --
> > >>>> 2.55.0
> > >>>>
> >
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.