Re: [RFC PATCH 0/3] KVM: Dirty page logging for guest_memfd-only memslots

Sean Christopherson <[email protected]>
Newsgroups dev.linux.lists.kvmarm,org.infradead.lists.linux-arm-kernel,org.kernel.vger.kvm
Message-ID <[email protected]>
On Thu, Aug 13, 2026, David Hildenbrand wrote:
> On 8/13/26 18:09, Alexandru Elisei wrote:
> > On Mon, Jul 13, 2026 at 04:11:57PM +0200, David Hildenbrand wrote:
> >>> Yeah. I agree (with the caveats you mention). Arm folk need to go figure
> >>> out if it's worthwhile to support with all those caveats.
> >>
> >> Right, disabling migration (once gmem supports it) was also what I discussed
> >> with Alexandru when that topic comes up.
> >>
> >> How to communicate to gmem that it wants these fixed mappings is a good question.
> > 
> > There's already a proposal for how to do this in the migratable guest_memfd
> > series [1] - it's a new guest_memfd creation flag that disables migration.
> 
> In the guest_memfd call I was arguing against the flag in the first version, and
> instead adding it when actually required.

Ya, right now a flag is meaningless.  Telling guest_memfd not to do something it
doesn't ever do...

> I was also raising whether KVM couldn't tell guest_memfd (e.g., at creation
> time?) that it supports a CPU feature that requires S2 to be always mapped to
> disable migration.
> 
> It would then be a contract between KVM and guest_memfd without user space
> having to be involved on that level.
> 
> > With Sean's comment that he expects swap/reclaim to be fully userspace
> > driven, I believe that would be enough to guarantee on the _kernel_ side
> > that SPE will work as intended for a guest.
> 
> That's my understanding.
> 
> > 
> > I'm a slightly concerned though that all of this will work by chance, and
> > not by design, and in the future the behaviour might change to allow
> > guest_memfd memory to be unmapped from stage 2 without the VMM or KVM
> > explicitly allowing it or initiating it.

Meh, TDX on x86 already has the same requirement.  Unmapping a page from the S-EPT
kills the VM unless the VM was expecting the page to be lost. 

> Thus my idea of the explicit contract between KVM and guest_memfd. Instead of
> being a "this doesn't support migration" it would be a "S2 always mapped"
> kind-of contract.

Who would that contract be between though?  KVM can tell a guest_memfd instance
that page migration is/isn't supported, but telling guest_memfd that the VM will
always keep the entirety of the guest_memfd mapped in stage-2 is nonsensical.
guest_memfd simply doesn't care if the page is mapped or not, it only cares if
migration is supported.

> > My understanding from the conversation so far is that the plan for the
> > future of guest_memfd is to support an option/mode where the memory is
> > effectively "pinned" at stage 2 (but which allows userspace to explicitly
> > free/unmap it, of course). Is that correct, or am I being overly optimistic
> > in my interpretation?
> 
> We could then even disallow fallocate() to punch holes if that contract is
> negotiated.

Why?  If userspace pulls a stupid and kills its guest, that's userspace's problem.

KVM would also have to block memslot changes, and probably other things in the
future that would unmap stage-2 in response to userspace syscalls/ioctls.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.