Re: [External] Re: [LSF/MM/BPF TOPIC] 64k (or 16k) base page size on x86

yunhui cui <[email protected]> Tue, 4 Aug 2026 20:24:44 +0800
Newsgroups org.kvack.linux-mm,org.kernel.vger.linux-kernel
Message-ID <CAEEQ3wmam=maXfXL+oyw=h4UsQ4hkd0gG1MytnuG0zV1_LdWcw@mail.gmail.com>
Hi Kiryl,


On Sat, May 2, 2026 at 2:01=E2=80=AFAM Kiryl Shutsemau <[email protected]> wro=
te:
>
> On Thu, Feb 19, 2026 at 03:08:58PM +0000, Kiryl Shutsemau wrote:
> > =3D=3D Status =3D=3D
> >
> > I have a POC implementation on top of v6.17:
> >
> > git://git.kernel.org/pub/scm/linux/kernel/git/kas/linux.git pte_size
>
> Just a heads up.
>
> I've updated the tree. It is on top of v7.0 now.

  Thank you for proposing the PG_SIZE/PTE_SIZE split (PGCL) and for
  sharing the implementation:

    https://git.kernel.org/pub/scm/linux/kernel/git/kas/linux.git/log/?h=3D=
pte_size

  Based on the pte_size branch, I developed a RISC-V PoC and used it for
  this evaluation on a RISC-V server. The branch includes the RISC-V
  adaptation, Svnapot PTE folding, split-PG correctness fixes and
  regression tests:

    https://github.com/cuiyunhui/linux/commits/pgcl_riscv

  The complete benchmark results, methodology, run-to-run variation,
  UnixBench subtests, fio results and investigation details are available
  here:

    https://github.com/cuiyunhui/linux/blob/pgcl_riscv/Documentation/mm/pgc=
l-riscv-evaluation.rst

  The RISC-V Svnapot PTE-folding implementation was also posted separately:

    https://lore.kernel.org/all/[email protected]=
m/

  For this evaluation, PGCL means:

    PG_SIZE  =3D 64 KiB
    PTE_SIZE =3D 4 KiB

  A 64 KiB allocator page is represented by sixteen 4 KiB PTEs. When the
  virtual address, physical layout and PTE attributes satisfy the folding
  requirements, these PTEs can form a 64 KiB Svnapot translation.

  I compared three configurations on the same RISC-V server:

    A. Conventional 4 KiB kernel
    B. 64 KiB PG_SIZE / 4 KiB PTE_SIZE (PGCL) + Svnapot PTE folding
    C. 4 KiB base pages + 64 KiB anonymous mTHP + the same PTE folding

  Both B and C use the same PTE-folding implementation. The main
  difference is whether 64 KiB is the kernel allocator-page granularity
  or a best-effort anonymous mTHP size.

  The same userspace binaries, CPU placement and datasets were used.
  No cpufreq driver was active during these measurements.

  PTE-folding benefit
  =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D

  I first isolated the effect of PTE folding by comparing PGCL with
  Svnapot disabled and enabled.

  For lat_mem_rd, DTLB miss/access decreased from 7.202% to 0.910%, while
  access latency decreased from 35.633 ns to 30.145 ns, a 15.40%
  improvement.

  For SPEC CPU2017 505.mcf_r, DTLB miss/access decreased from 6.062% to
  1.021%, while elapsed time decreased from 282.68 seconds to 273.61
  seconds, a 3.21% improvement.

  In the complete A/B/C comparison, PGCL and 64 KiB mTHP produced nearly
  identical DTLB results. Their lat_mem_rd DTLB miss/access ratios were
  0.9180% and 0.9200%, respectively, compared with 5.6768% for A. Their
  505.mcf_r ratios were 1.0189% and 1.0234%, compared with 6.0635% for A.

  Application benchmarks
  =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D

  Most Redis, MariaDB and small-response Nginx results remained within
  approximately 4% of the conventional 4 KiB kernel.

  The largest difference appeared in Nginx with large static responses.
  PGCL improved 64 KiB and 1 MiB response throughput by 18.33% and
  19.26%, while 64 KiB anonymous mTHP improved them by 0.70% and 0.67%.

  The Nginx test used a warmed page cache, sendfile and loopback TCP. Its
  large-response path includes page-cache, sendfile and TCP processing
  that anonymous mTHP does not change.

  A separate physical-NIC iPerf comparison showed a difference of a
  similar magnitude with one unidirectional stream, while the PGCL/mTHP
  difference became small with multiple concurrent streams.

  One possible contributor on a physical NIC path is page_pool allocation
  geometry. With a 64 KiB allocator page, an order-0 page_pool allocation
  naturally obtains one complete 64 KiB page. A 4 KiB kernel would require
  an explicit order-4 allocation to obtain the same physically contiguous
  size, which drivers may avoid because of fragmentation risk. This is an
  allocator and buffer-backing property, not Svnapot folding of the kernel
  linear map.

  UnixBench regression
  =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D

  The 16-CPU concurrent UnixBench system index was:

    A: 20265.5
    B: 14513.1 (-28.39%)
    C: 20186.3 (-0.39%)

  The PGCL loss was concentrated in Execl, Process Creation and Shell.
  Arithmetic, pipe throughput and ordinary system-call throughput remained
  close to A.

  One issue found during the investigation was that file-backed
  fault-around had been disabled when PTES_PER_PAGE > 1 as a correctness
  workaround for sub-PG mappings. I implemented PTE-granular file
  fault-around with mapping and file-size bounds.

  After the change, filemap_map_pages() increased from zero to 21.28 calls
  per exec, close to A at 21.66. filemap_fault() decreased from 123.03 to
  4.31 calls per exec, close to A at 3.10.

  The single-CPU UnixBench index improved by 2.38%, and single-CPU Execl
  throughput improved by 16.52%. However, the 16-CPU system index remained
  at 14588.2, compared with 14513.1 before the fix and 20265.5 for A.

  Restoring file-backed fault-around therefore fixes a real single-CPU
  issue but does not recover the 16-CPU regression. The concurrent Execl,
  Process Creation and Shell paths have another remaining bottleneck.

  Direct-I/O results
  =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D

  The final fio direct-read comparison used psync, direct=3D1, iodepth=3D1,
  numjobs=3D1 and three measured runs per workload.

  Random-read throughput was close to A. The remaining regression appeared
  in 64 KiB sequential reads:

    A: 3279.804 MiB/s
    B: 3036.649 MiB/s (-7.41%)
    C: 3049.891 MiB/s (-7.01%)

  An A/C kprobe comparison over 20,000 64 KiB I/Os showed that A performed
  20,000 batched fast-GUP operations. The tested C path instead performed
  320,000 PTE-granular page pins and 320,031 VMA lookups, corresponding to
  approximately sixteen operations per I/O.

  The tested code processes one PTE at a time to preserve the correct
  4 KiB sub-page offset. This is an implementation overhead rather than a
  fundamental limit of PGCL or mTHP. Possible optimizations include
  retaining batched fast-GUP when PG_SIZE =3D=3D PTE_SIZE and providing
  batched PTE-granular pinning that preserves sub-page offsets for split-PG
  configurations.

  Userspace ABI, PTE folding and TLB coverage
  =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D

  A foldable 64 KiB mapping requires at least:

    - a 64 KiB-aligned virtual address;
    - a 64 KiB-aligned and physically contiguous range;
    - sixteen PTEs with compatible permissions and attributes.

  These requirements apply to both data and instruction mappings and
  therefore affect both DTLB and iTLB coverage.

  Some processors, including AMD implementations with PTE Coalescing, can
  coalesce adjacent 4 KiB translations within the microarchitecture
  without an architectural encoding such as RISC-V Svnapot. Although the
  implementation details are architecture-specific, these mechanisms
  similarly depend on a suitably aligned virtual-address range and
  compatible adjacent translations.

  For instruction translations, the virtual-address layout of executable
  code is especially important because it determines whether code mappings
  can satisfy the alignment requirements and obtain iTLB benefits.

  PGCL currently exposes a 4 KiB userspace page size:

    getpagesize()         =3D 4096
    sysconf(_SC_PAGESIZE) =3D 4096
    AT_PAGESZ             =3D 4096

  This preserves the existing 4 KiB userspace ABI, but applications and
  runtimes using the reported page size for reservation, alignment,
  mapping and permission transitions are not informed that 64 KiB is the
  preferred folding granularity.

  This is relevant to JVM-based workloads such as Flink. The OpenJDK JVM
  uses the operating-system page size when reserving and aligning virtual
  memory, including executable JIT CodeCache mappings. A reported 64 KiB
  page size makes the runtime more likely to preserve 64 KiB alignment and
  obtain better iTLB coverage.

  Another example of userspace behavior that changes according to
  getpagesize() is discussed here:

    https://github.com/intel/lmbench/pull/55

  Should the operating-system ecosystem also provide an optional,
  end-to-end 64 KiB userspace environment?

  Such support would require coordination across the kernel userspace ABI,
  distribution build rules, toolchains, ELF binaries, shared libraries,
  loaders, memory allocators and language runtimes. In such an environment,
  getpagesize() and AT_PAGESZ would report 64 KiB, allowing userspace
  mappings and executable code to preserve the alignment needed for
  large-page mappings, PTE folding and potentially better DTLB and iTLB
  coverage.

  Crash and kdump tooling
  =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D

  PGCL may also require changes to crash, makedumpfile and related tooling,
  because allocator pages use a 64 KiB PG_SIZE while page-table translation
  uses a 4 KiB PTE_SIZE, but existing VMCOREINFO and tools may assume a
  single page size.

>
> I can boot my VM fully (not only to shell as before) on all three page
> size 4k, 16k and 64k. And it doesn't crash immediately.
>
> Page fault handler is the most challenging part of the design. I wanted
> to flash it out to RFC quality before the summit, but it is not there
> yet. But it started getting a shape.
>
> Volume of the code changes makes it hard to deal with the patchset:
>
>  1530 files changed, 12675 insertions(+), 8669 deletions(-)
>
> Most of it mechanical, but still.
>
> --
>   Kiryl Shutsemau / Kirill A. Shutemov
>

Thanks,
Yunhui