Re: [PATCH v7 bpf-next] bpf: Populate mmap-able array map memory lazily

Andrii Nakryiko <[email protected]>
Newsgroups org.kernel.vger.bpf
Message-ID <CAEf4Bzb33ah9dk1nSPYUH1XhWvSPw3sKR_KjUCJ2Nb-3stTYAA@mail.gmail.com>
On Fri, Aug 14, 2026 at 8:56 AM Song Liu <[email protected]> wrote:
>
> An mmap-able BPF array map (BPF_F_MMAPABLE) has its backing memory
> vmalloc'ed up front at map creation time. array_map_mmap() then wired up
> the whole mapping eagerly via remap_vmalloc_range(), which calls
> vm_insert_page() for every page of the map. For large maps this makes
> every mmap() O(number of pages): an 8MiB map inserts 2048 PTEs per
> mmap() and tears them all down again on munmap(), even when user space
> only touches a few pages (or none at all).
>
> Populate the mapping lazily instead, the same way the arena map already
> does. array_map_mmap() now only performs the bounds check and returns,
> leaving the PTEs unpopulated; pages are inserted on demand by a new
> array_map_mmap_fault() handler. Because the memory is already resident,
> the fault handler simply resolves the vmalloc page and hands it to the
> fault path. This makes mmap() O(1), and munmap() proportional to the
> number of pages that were actually faulted in rather than to the size of
> the map.
>
> The handler is reached through a new optional ->map_mmap_fault callback.
> Maps that provide it get a vm_operations_struct with a .fault handler;
> maps that populate their mapping eagerly keep the one they had. Both
> share the same open/close callbacks, so the existing VMA accounting
> (VM_MAYWRITE write-active tracking, freeze handling) stays centralized
> rather than each map installing its own vm_operations_struct.
>
> Callers that want the pages populated up front can still request that
> explicitly with MAP_POPULATE. Kernel-side access to the map (via the
> vmalloc address) is unaffected.
>
> Time for one mmap()+munmap() of an 8MiB mmap-able array map:
>
>                                        before     after
>   no MAP_POPULATE, no access            226us     1.1us
>   no MAP_POPULATE, access all pages     236us    1341us
>   MAP_POPULATE, no access               312us     493us
>   MAP_POPULATE, access all pages        318us     519us
>
> Mapping without touching the data, which is what this change targets,
> gets ~160x cheaper. Faulting in the whole mapping one page at a time is
> more expensive than the eager remap_vmalloc_range() loop, so users that
> do touch every page should ask for MAP_POPULATE. Note that MAP_POPULATE
> is not free before this change either: it adds ~85us (226us => 312us)
> for no benefit, as the mapping is already fully populated.
>
> Signed-off-by: Song Liu <[email protected]>
> Assisted-by: Claude:claude-opus-4-8
>
> ---
> Changes in v7:
> - Set VM_MIXEDMAP, which the eager remap_vmalloc_range() path set via
>   vm_insert_page(), so that e.g. NUMA balancing keeps skipping these
>   VMAs. (BPF CI AI review)
> - Only install the .fault handler for maps that populate their mapping
>   lazily, so that eagerly populated maps (ringbuf) keep exactly the
>   vm_operations_struct they had before. (BPF CI AI review)
> v6: https://lore.kernel.org/bpf/[email protected]/
>
> Changes in v6:
> - Drop the !CONFIG_MMU branch. nommu cannot mmap() a BPF map fd in the
>   first place: the fd is an anon inode and bpf_map_fops has no
>   ->mmap_capabilities, so validate_mmap_request() returns -EINVAL before
>   ->mmap() runs. (BPF CI AI review)
> - Scope the O(1) claim to mmap(); munmap() still tears down whatever was
>   faulted in. (BPF CI AI review)
> v5: https://lore.kernel.org/bpf/[email protected]/
>
> Changes in v5:
> - Drop the ->map_pages (fault-around) handler, it pulls in too much mm
>   internal API for the gain. (Andrii)
> - Drop the fault path overflow and bounds checks, and the verbose
>   comments; VM_DONTEXPAND plus the mmap() time check already bound the
>   faulting offset. (Andrii)
> - No cover letter for a single patch. (Andrii)
> v4: https://lore.kernel.org/bpf/[email protected]/
>
> Changes in v4:
> - Flush the D-cache before exposing a page at a new user address, as the
>   eager vm_insert_page() path did. (Sashiko AI review)
> - Fix the build on !CONFIG_MMU: keep populating the mapping eagerly
>   there, as there are no page faults. (kernel test robot)
> v3: https://lore.kernel.org/bpf/[email protected]/
>
> Changes in v3:
> - Add a ->map_pages (fault-around) handler so mmap(MAP_POPULATE) and
>   linear access populate PTEs in batches instead of one fault per page.
> - Harden the fault path with check_shl_overflow() and explicit bounds
>   checks instead of a plain (u64) cast. (Andrii)
> - Drop selftests (2/2 in v2). (Andrii)
> v2: https://lore.kernel.org/bpf/[email protected]/
>
> Changes in v2:
> - Use 64-bit arithmetic for the mmap offset and bounds check to avoid a
>   potential overflow on 32-bit architectures.
> v1: https://lore.kernel.org/bpf/[email protected]/
> ---
>  include/linux/bpf.h   |  1 +
>  kernel/bpf/arraymap.c | 34 ++++++++++++++++++++++++++++++----
>  kernel/bpf/syscall.c  | 22 +++++++++++++++++++++-
>  3 files changed, 52 insertions(+), 5 deletions(-)
>
> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
> index f4e8d372253a..04cadd987169 100644
> --- a/include/linux/bpf.h
> +++ b/include/linux/bpf.h
> @@ -145,6 +145,7 @@ struct bpf_map_ops {
>         int (*map_direct_value_meta)(const struct bpf_map *map,
>                                      u64 imm, u32 *off);
>         int (*map_mmap)(struct bpf_map *map, struct vm_area_struct *vma);
> +       vm_fault_t (*map_mmap_fault)(struct bpf_map *map, struct vm_fault *vmf);
>         __poll_t (*map_poll)(struct bpf_map *map, struct file *filp,
>                              struct poll_table_struct *pts);
>         unsigned long (*map_get_unmapped_area)(struct file *filep, unsigned long addr,
> diff --git a/kernel/bpf/arraymap.c b/kernel/bpf/arraymap.c
> index 34865701f7f7..ef315b168b29 100644
> --- a/kernel/bpf/arraymap.c
> +++ b/kernel/bpf/arraymap.c
> @@ -608,17 +608,42 @@ static int array_map_check_btf(struct bpf_map *map,
>  static int array_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
>  {
>         struct bpf_array *array = container_of(map, struct bpf_array, map);
> -       pgoff_t pgoff = PAGE_ALIGN(sizeof(*array)) >> PAGE_SHIFT;
>
>         if (!(map->map_flags & BPF_F_MMAPABLE))
>                 return -EINVAL;
>
> -       if (vma->vm_pgoff * PAGE_SIZE + (vma->vm_end - vma->vm_start) >
> +       /* use u64 math so the offset cannot overflow on 32-bit archs */
> +       if ((u64)vma->vm_pgoff * PAGE_SIZE + (vma->vm_end - vma->vm_start) >
>             PAGE_ALIGN((u64)array->map.max_entries * array->elem_size))
>                 return -EINVAL;
>
> -       return remap_vmalloc_range(vma, array_map_vmalloc_addr(array),
> -                                  vma->vm_pgoff + pgoff);
> +       /*
> +        * Pages are faulted in on demand by array_map_mmap_fault(). Set the
> +        * same flags that the eager remap_vmalloc_range() path used to set
> +        * via vm_insert_page(), so that e.g. NUMA balancing keeps skipping
> +        * these VMAs.
> +        */
> +       vm_flags_set(vma, VM_DONTEXPAND | VM_DONTDUMP | VM_MIXEDMAP);
> +
> +       return 0;
> +}
> +
> +static vm_fault_t array_map_mmap_fault(struct bpf_map *map,
> +                                      struct vm_fault *vmf)
> +{
> +       struct bpf_array *array = container_of(map, struct bpf_array, map);
> +       struct page *page;
> +
> +       page = vmalloc_to_page(array->value + ((u64)vmf->pgoff << PAGE_SHIFT));
> +       if (!page)
> +               return VM_FAULT_SIGBUS;
> +
> +       /* the eager remap_vmalloc_range() flushed via vm_insert_page() */
> +       flush_dcache_folio(page_folio(page));
> +       get_page(page);
> +       vmf->page = page;
> +
> +       return 0;
>  }
>
>  static bool array_map_meta_equal(const struct bpf_map *meta0,
> @@ -844,6 +869,7 @@ const struct bpf_map_ops array_map_ops = {
>         .map_direct_value_addr = array_map_direct_value_addr,
>         .map_direct_value_meta = array_map_direct_value_meta,
>         .map_mmap = array_map_mmap,
> +       .map_mmap_fault = array_map_mmap_fault,
>         .map_seq_show_elem = array_map_seq_show_elem,
>         .map_check_btf = array_map_check_btf,
>         .map_lookup_batch = generic_map_lookup_batch,
> diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
> index 7d8c3e8e6d62..caf3e202ad6c 100644
> --- a/kernel/bpf/syscall.c
> +++ b/kernel/bpf/syscall.c
> @@ -1076,11 +1076,30 @@ static void bpf_map_mmap_close(struct vm_area_struct *vma)
>                 bpf_map_write_active_dec(map);
>  }
>
> +static vm_fault_t bpf_map_mmap_fault(struct vm_fault *vmf)
> +{
> +       struct bpf_map *map = vmf->vma->vm_private_data;
> +
> +       return map->ops->map_mmap_fault(map, vmf);
> +}
> +
>  static const struct vm_operations_struct bpf_map_default_vmops = {
>         .open           = bpf_map_mmap_open,
>         .close          = bpf_map_mmap_close,
>  };
>
> +/*
> + * Only for maps that populate their mapping lazily. Maps that populate it
> + * eagerly keep the default vm_ops, so that a fault racing the temporary
> + * pte clearing of a read/modify/write update still gets the pte re-checked
> + * under the ptl by do_fault() instead of a SIGBUS.
> + */

removed this comment altogether

> +static const struct vm_operations_struct bpf_map_lazy_vmops = {
> +       .open           = bpf_map_mmap_open,
> +       .close          = bpf_map_mmap_close,
> +       .fault          = bpf_map_mmap_fault,
> +};
> +
>  static int bpf_map_mmap(struct file *filp, struct vm_area_struct *vma)
>  {
>         struct bpf_map *map = filp->private_data;
> @@ -1116,7 +1135,8 @@ static int bpf_map_mmap(struct file *filp, struct vm_area_struct *vma)
>                 return err;
>
>         /* set default open/close callbacks */
> -       vma->vm_ops = &bpf_map_default_vmops;
> +       vma->vm_ops = map->ops->map_mmap_fault ? &bpf_map_lazy_vmops
> +                                              : &bpf_map_default_vmops;

made this into a single line statement, applied to bpf-next, thanks!

>         vma->vm_private_data = map;
>         vm_flags_clear(vma, VM_MAYEXEC);
>         /* If mapping is read-only, then disallow potentially re-mapping with
>
> base-commit: 259d60f5bfa41056fe01cbf2ba3f6f0331865a16
> --
> 2.53.0-Meta
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.