Re: [PATCH v3 11/26] x86/mm: introduce the mermap

Mike Rapoport <[email protected]>
Newsgroups gmane.linux.kernel,gmane.linux.kernel.mm
Message-ID <[email protected]>
On Sun, Jul 26, 2026 at 10:22:44PM +0000, Brendan Jackman wrote:
> The mermap provides a fast way to create ephemeral mm-local mappings of
> physical pages. The purpose of this is to access pages that have been
> removed from the direct map. Potential use cases are:
> 
> 1. For zeroing direct-map-nonpresent pages (added in a later patch).
> 
> 2. For populating guest_memfd pages that are protected by the
>    GUEST_MEMFD_NO_DIRECT_MAP feature [0].
> 
> 3. For efficient access of pages protected by Address Space Isolation
>    [1].
> 
> [0] https://lore.kernel.org/all/[email protected]/
> [1] https://linuxasi.dev
> 
> The details of this mechanism are described in the API comments. However
> the key idea is to use CPU-local virtual regions to avoid a need for
> synchronizing. On x86, this can also be used to prevent TLB shootdowns.
> 
> Because the virtual region is CPU-local, allocating from the mermap
> disables migration. The caller is forbidden to use the returned value
> from any other context, and migration is re-enabled when it's freed.
> 
> One might notice that mermap_get() bears a strong similarity to
> kmap_local_page(). The most important differences between mermap_get()
> and kmap_local_page() are:
> 
> 1. mermap_get() allows mapping variable sizes while kmap_local_page()
>    specifically maps a single order-0 page.
> 2. As a consequence of 1 (combined with the need for mermap_get() to be
>    an extremely simple allocator), mermap_get() should be expected to
>    fail, while kmap_local_page() is guaranteed to work up to a certain
>    degree of nesting.
> 3. While the mappings provided by kmap_local_page() are _logically_
>    local to the calling context (it's a bug for software to access them
>    from elsewhere), they are _physically_ installed into the shared
>    kernel pagetables. This means their locality doesn't provide any
>    protection from hardware attacks. In contrast, the mermap is
>    physically local to the creating mm, taking advantage of the new
>    mm-local kernel address region.
> 
> So that the mermap is available even in contexts where failure is not
> tolerable there is also a _reserved() variant, which is fixed at
> allocating a single base page. This is useful, for example, for zeroing
> unmapped pages, where handling failure would be extremely inconvenient.
> The _reserved() variant is simply implemented by leaving one base-page
> space unavailable for non-_reserved allocations, and requiring an atomic
> context.
> 
> Note for Sashiko: Yes, the data mapped by the mermap is exposed to
> Meltdown-style attacks by the current process. This is completely
> intentional. Data is only supposed to be mapped there that the current
> process is allowed to read anyway.
> 
> Signed-off-by: Brendan Jackman <[email protected]>
> ---
>  arch/x86/Kconfig                        |   1 +
>  arch/x86/include/asm/mermap.h           |  23 +++
>  arch/x86/include/asm/pgtable_64_types.h |   8 +-
>  arch/x86/include/asm/pgtable_types.h    |   2 +
>  include/linux/mermap.h                  |  63 ++++++
>  include/linux/mermap_types.h            |  41 ++++
>  include/linux/mm_types.h                |   4 +
>  kernel/fork.c                           |   5 +
>  mm/Kconfig                              |   9 +
>  mm/Makefile                             |   1 +
>  mm/mermap.c                             | 338 ++++++++++++++++++++++++++++++++
>  11 files changed, 494 insertions(+), 1 deletion(-)
> 
> diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig
> index 33c1282bfbf93..6b4d81a280d3b 100644
> --- a/arch/x86/Kconfig
> +++ b/arch/x86/Kconfig
> @@ -37,6 +37,7 @@ config X86_64
>  	select ZONE_DMA32
>  	select EXECMEM if DYNAMIC_FTRACE
>  	select ACPI_MRRM if ACPI
> +	select ARCH_SUPPORTS_MERMAP
>  
>  config FORCE_DYNAMIC_FTRACE
>  	def_bool y
> diff --git a/arch/x86/include/asm/mermap.h b/arch/x86/include/asm/mermap.h
> new file mode 100644
> index 0000000000000..9d7614716b718
> --- /dev/null
> +++ b/arch/x86/include/asm/mermap.h
> @@ -0,0 +1,23 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef _ASM_X86_MERMAP_H
> +#define _ASM_X86_MERMAP_H
> +
> +#include <asm/tlbflush.h>
> +
> +static inline void arch_mermap_flush_tlb(void)
> +{
> +	/*
> +	 * No shootdown allowed, IRQs may be off. Luckily other CPUs are not
> +	 * allowed to access our region so the stale mappings are harmless, as
> +	 * long as they still point to data belonging to this process.
> +	 */
> +	__flush_tlb_all();
> +}
> +
> +static inline bool arch_mermap_pgprot_allowed(pgprot_t prot)
> +{
> +	/* Mermap is mm-local so global mappings would be a bug. */
> +	return !(pgprot_val(prot) & _PAGE_GLOBAL);
> +}
> +
> +#endif /* _ASM_X86_MERMAP_H */
> diff --git a/arch/x86/include/asm/pgtable_64_types.h b/arch/x86/include/asm/pgtable_64_types.h
> index 1181565966405..fb6c3daacfeb8 100644
> --- a/arch/x86/include/asm/pgtable_64_types.h
> +++ b/arch/x86/include/asm/pgtable_64_types.h
> @@ -105,11 +105,17 @@ extern unsigned int ptrs_per_p4d;
>  
>  #define MM_LOCAL_PGD_ENTRY	-240UL
>  #define MM_LOCAL_BASE_ADDR	(MM_LOCAL_PGD_ENTRY << PGDIR_SHIFT)
> -#define MM_LOCAL_END_ADDR	((MM_LOCAL_PGD_ENTRY + 1) << PGDIR_SHIFT)
> +#define MM_LOCAL_START_ADDR	((MM_LOCAL_PGD_ENTRY) << PGDIR_SHIFT)
> +#define MM_LOCAL_END_ADDR	(MM_LOCAL_START_ADDR + (1UL << PGDIR_SHIFT))
>  
>  #define LDT_BASE_ADDR		MM_LOCAL_BASE_ADDR
>  #define LDT_END_ADDR		(LDT_BASE_ADDR + PMD_SIZE)
>  
> +#define MERMAP_BASE_ADDR	LDT_END_ADDR
> +#define MERMAP_CPU_REGION_SIZE	PMD_SIZE
> +#define MERMAP_SIZE		(MERMAP_CPU_REGION_SIZE * NR_CPUS)
> +#define MERMAP_END_ADDR		(MERMAP_BASE_ADDR + (NR_CPUS * MERMAP_CPU_REGION_SIZE))
> +
>  #define __VMALLOC_BASE_L4	0xffffc90000000000UL
>  #define __VMALLOC_BASE_L5 	0xffa0000000000000UL
>  
> diff --git a/arch/x86/include/asm/pgtable_types.h b/arch/x86/include/asm/pgtable_types.h
> index af08d98be9309..f397e4311cf66 100644
> --- a/arch/x86/include/asm/pgtable_types.h
> +++ b/arch/x86/include/asm/pgtable_types.h
> @@ -223,6 +223,7 @@ enum page_cache_mode {
>  #define __PAGE_KERNEL_RO	 (__PP|   0|   0|___A|__NX|   0|   0|___G)
>  #define __PAGE_KERNEL_ROX	 (__PP|   0|   0|___A|   0|   0|   0|___G)
>  #define __PAGE_KERNEL		 (__PP|__RW|   0|___A|__NX|___D|   0|___G)
> +#define __PAGE_KERNEL_NOGLOBAL	 (__PP|__RW|   0|___A|__NX|___D|   0|   0)
>  #define __PAGE_KERNEL_EXEC	 (__PP|__RW|   0|___A|   0|___D|   0|___G)
>  #define __PAGE_KERNEL_NOCACHE	 (__PP|__RW|   0|___A|__NX|___D|   0|___G| __NC)
>  #define __PAGE_KERNEL_VVAR	 (__PP|   0|_USR|___A|__NX|   0|   0|___G)
> @@ -245,6 +246,7 @@ enum page_cache_mode {
>  #define __pgprot_mask(x)	__pgprot((x) & __default_kernel_pte_mask)
>  
>  #define PAGE_KERNEL		__pgprot_mask(__PAGE_KERNEL            | _ENC)
> +#define PAGE_KERNEL_NOGLOBAL	__pgprot_mask(__PAGE_KERNEL_NOGLOBAL   | _ENC)
>  #define PAGE_KERNEL_NOENC	__pgprot_mask(__PAGE_KERNEL            |    0)
>  #define PAGE_KERNEL_RO		__pgprot_mask(__PAGE_KERNEL_RO         | _ENC)
>  #define PAGE_KERNEL_EXEC	__pgprot_mask(__PAGE_KERNEL_EXEC       | _ENC)
> diff --git a/include/linux/mermap.h b/include/linux/mermap.h
> new file mode 100644
> index 0000000000000..5457dcb8c9789
> --- /dev/null
> +++ b/include/linux/mermap.h
> @@ -0,0 +1,63 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef _LINUX_MERMAP_H
> +#define _LINUX_MERMAP_H
> +
> +#include <linux/mermap_types.h>
> +#include <linux/mm.h>
> +
> +#ifdef CONFIG_MERMAP
> +
> +#include <asm/mermap.h>
> +
> +int mermap_mm_prepare(struct mm_struct *mm);
> +void mermap_mm_init(struct mm_struct *mm);
> +void mermap_mm_teardown(struct mm_struct *mm);
> +
> +/* Can the mermap be called from this context? */
> +static inline bool mermap_ready(void)
> +{
> +	return in_task() && current->mm && current->mm->mermap.cpu;
> +}
> +
> +struct mermap_alloc *mermap_get(struct page *page, unsigned long size, pgprot_t prot);
> +void *mermap_get_reserved(struct page *page, pgprot_t prot);
> +void mermap_put(struct mermap_alloc *alloc);
> +
> +static inline void *mermap_addr(struct mermap_alloc *alloc)
> +{
> +	return (void *)alloc->base;
> +}
> +
> +/*
> + * arch_mermap_flush_tlb() is called before a part of the local CPU's mermap
> + * region is remapped to a new address. No other CPU is allowed to _access_ that
> + * region, but the region was mapped there.
> + *
> + * This may be called with IRQs off.
> + *
> + * On arm64, this will need to be a broadcast TLB flush. Although the other CPUs
> + * are forbidden to access the region, they can leak the data that was mapped
> + * there via CPU exploits. Violating break-before-make would mean the data
> + * available to these CPU exploits is unpredictable.
> + */
> +extern void arch_mermap_flush_tlb(void);
> +extern bool arch_mermap_pgprot_allowed(pgprot_t prot);
> +
> +#if IS_ENABLED(CONFIG_KUNIT)
> +struct mermap_alloc *__mermap_get(struct mm_struct *mm, struct page *page,
> +			unsigned long size, pgprot_t prot, bool use_reserve);
> +void __mermap_put(struct mm_struct *mm, struct mermap_alloc *alloc);
> +unsigned long mermap_cpu_base(int cpu);
> +unsigned long mermap_cpu_end(int cpu);
> +#endif
> +
> +#else /* CONFIG_MERMAP */
> +
> +static inline int mermap_mm_prepare(struct mm_struct *mm) { return 0; }
> +static inline void mermap_mm_init(struct mm_struct *mm) { }
> +static inline void mermap_mm_teardown(struct mm_struct *mm) { }
> +static inline bool mermap_ready(void) { return false; }
> +
> +#endif /* CONFIG_MERMAP */
> +
> +#endif /* _LINUX_MERMAP_H */
> diff --git a/include/linux/mermap_types.h b/include/linux/mermap_types.h
> new file mode 100644
> index 0000000000000..c1c83b223c28d
> --- /dev/null
> +++ b/include/linux/mermap_types.h
> @@ -0,0 +1,41 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef _LINUX_MERMAP_TYPES_H
> +#define _LINUX_MERMAP_TYPES_H
> +
> +#include <linux/mutex.h>
> +#include <linux/percpu.h>
> +#include <linux/types.h>
> +
> +#ifdef CONFIG_MERMAP
> +
> +/* Tracks an individual allocation in the mermap. */
> +struct mermap_alloc {
> +	/* Currently allocated. */
> +	bool in_use;
> +	/* Requires flush before reallocating. */
> +	bool need_flush;
> +	unsigned long base;
> +	/* Non-inclusive. */
> +	unsigned long end;
> +};
> +
> +struct mermap_cpu {
> +	/* Next address immediately available for alloc (no TLB flush needed). */
> +	unsigned long next_addr;
> +	struct mermap_alloc normal_allocs[3];
> +	struct mermap_alloc reserve_alloc;
> +};
> +
> +struct mermap {
> +	struct mutex init_lock;
> +	struct mermap_cpu __percpu *cpu;
> +};
> +
> +#else /* CONFIG_MERMAP */
> +
> +struct mermap {};
> +
> +#endif /* CONFIG_MERMAP */
> +
> +#endif /* _LINUX_MERMAP_TYPES_H */
> +
> diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
> index d39fddf57edc8..bb80d60cf3498 100644
> --- a/include/linux/mm_types.h
> +++ b/include/linux/mm_types.h
> @@ -7,6 +7,7 @@
>  #include <linux/auxvec.h>
>  #include <linux/kref.h>
>  #include <linux/list.h>
> +#include <linux/mermap_types.h>
>  #include <linux/spinlock.h>
>  #include <linux/rbtree.h>
>  #include <linux/maple_tree.h>
> @@ -35,6 +36,7 @@
>  struct address_space;
>  struct futex_private_hash;
>  struct mem_cgroup;
> +struct mermap;
>  
>  typedef struct {
>  	unsigned long f;
> @@ -1211,6 +1213,8 @@ struct mm_struct {
>  		atomic_t membarrier_state;
>  #endif
>  
> +		struct mermap mermap;
> +
>  		/**
>  		 * @mm_users: The number of users including userspace.
>  		 *
> diff --git a/kernel/fork.c b/kernel/fork.c
> index 4fb23ea33b7da..7c2050e76d1bc 100644
> --- a/kernel/fork.c
> +++ b/kernel/fork.c
> @@ -13,6 +13,7 @@
>   */
>  
>  #include <linux/anon_inodes.h>
> +#include <linux/mermap.h>
>  #include <linux/slab.h>
>  #include <linux/sched/autogroup.h>
>  #include <linux/sched/mm.h>
> @@ -1143,6 +1144,9 @@ static struct mm_struct *mm_init(struct mm_struct *mm, struct task_struct *p)
>  		goto fail_pcpu;
>  
>  	lru_gen_init_mm(mm);
> +
> +	mermap_mm_init(mm);
> +
>  	return mm;
>  
>  fail_pcpu:
> @@ -1186,6 +1190,7 @@ static inline void __mmput(struct mm_struct *mm)
>  	ksm_exit(mm);
>  	khugepaged_exit(mm); /* must run before exit_mmap */
>  	exit_mmap(mm);
> +	mermap_mm_teardown(mm);
>  	mm_put_huge_zero_folio(mm);
>  	set_mm_exe_file(mm, NULL);
>  	if (!list_empty(&mm->mmlist)) {
> diff --git a/mm/Kconfig b/mm/Kconfig
> index cb531c1436f77..bf8c4c6264c73 100644
> --- a/mm/Kconfig
> +++ b/mm/Kconfig
> @@ -1507,6 +1507,15 @@ config MM_LOCAL_REGION
>  	bool
>  	depends on ARCH_SUPPORTS_MM_LOCAL_REGION
>  
> +config ARCH_SUPPORTS_MERMAP
> +	bool
> +	select ARCH_SUPPORTS_MM_LOCAL_REGION
> +
> +config MERMAP
> +	bool
> +	depends on ARCH_SUPPORTS_MERMAP
> +	select MM_LOCAL_REGION
> +
>  source "mm/damon/Kconfig"
>  
>  endmenu
> diff --git a/mm/Makefile b/mm/Makefile
> index ab37ef428d98d..9cf282c154104 100644
> --- a/mm/Makefile
> +++ b/mm/Makefile
> @@ -147,3 +147,4 @@ obj-$(CONFIG_EXECMEM) += execmem.o
>  obj-$(CONFIG_TMPFS_QUOTA) += shmem_quota.o
>  obj-$(CONFIG_LAZY_MMU_MODE_KUNIT_TEST) += tests/lazy_mmu_mode_kunit.o
>  obj-$(CONFIG_MEM_ALLOC_PROFILING) += alloc_tag.o
> +obj-$(CONFIG_MERMAP) += mermap.o
> diff --git a/mm/mermap.c b/mm/mermap.c
> new file mode 100644
> index 0000000000000..2bead38eadfe8
> --- /dev/null
> +++ b/mm/mermap.c
> @@ -0,0 +1,338 @@
> +// SPDX-License-Identifier: GPL-2.0
> +#include <linux/io.h>
> +#include <linux/error-injection.h>
> +#include <linux/mermap.h>
> +#include <linux/mm.h>
> +#include <linux/mmu_context.h>
> +#include <linux/mutex.h>
> +#include <linux/pagemap.h>
> +#include <linux/pgtable.h>
> +#include <linux/sched.h>
> +
> +#include <kunit/visibility.h>
> +
> +#include "internal.h"
> +
> +static inline int set_unmapped_pte(pte_t *ptep, unsigned long addr, void *data)
> +{
> +	set_pte(ptep, __pte(0));
> +	return 0;
> +}
> +
> +VISIBLE_IF_KUNIT void __mermap_put(struct mm_struct *mm, struct mermap_alloc *alloc)
> +{
> +	unsigned long size = PAGE_ALIGN(alloc->end - alloc->base);
> +
> +	__apply_to_page_range(mm, alloc->base, size, set_unmapped_pte,
> +			      NULL, PGRANGE_CREATE | PGRANGE_NOLOCK);
> +

Sorry if I missed that in previous discussions.

__apply_to_page_range() acts only on PTE mappings, and looking forward I
presume we'd want PMD and maybe event PUD mappings in guest_memfd and
subsequently in mermap.

We anyway have a ton of page table walkers, so maybe it'll make sense to
add yet another one rather than adjust __apply_to_page_range() to the
mermap needs?

Or maybe there's a suitable walk_ API in mm/pagewalk.c?

> +	WRITE_ONCE(alloc->in_use, false);
> +}
> +EXPORT_SYMBOL_IF_KUNIT(__mermap_put);

-- 
Sincerely yours,
Mike.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.