Re: [RFC] KVM: x86/mmu: Prefetch forward run of pages on TDP page faults

James Houghton <[email protected]>
Newsgroups org.kernel.vger.kvm,org.kernel.vger.linux-kernel
Message-ID <CADrL8HUcr5Fei1PX179bkFmMJ_S+eD267SVgc85cm-rpcahwmA@mail.gmail.com>
On Tue, Aug 25, 2026 at 7:27 AM Marco Marangoni <[email protected]> wrote:
>
> The TDP MMU installs one SPTE per fault, so faulting in guest memory via
> userfaultfd (e.g. snapshot restore) without hugetlbfs costs a VM-exit per
> 4KiB page. Mirror the shadow MMU's prefetch: after a 4KiB fault, resolve
> the forward run of host-present pages in the faulting leaf table (one
> guest 2MiB region) with one non-blocking GUP and fill the empty SPTEs.
> When userfaultfd populates in large chunks this maps up to 511 neighbours
> per fault, cutting EPT violations up to 512x.
>
> Touching 128MiB backed by userfaultfd with 2MiB UFFD_COPY chunks:
> c8i.metal-48xl 129.1 -> 100.3 ms (-22%), nested 495.8 -> 121.5 ms (-75%).
> With no batching (one copy per fault) there is a ~2-4% regression.

Hi Marco!

I think part of the problem in this case is that UFFDIO_COPY will
install 4K pages (IIRC), I think a more natural way to fix this
problem is to:

1. MADV_COLLAPSE after doing UFFDIO_COPY.
2. Make UFFDIO_COPY install PMDs when it is able to do so.

These don't solve the exact same problem, but really userfaultfd
should already try to install PMDs when it can (#2). If we have #2, #1
is mostly a no-op.

What do you think?
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.