Re: [PATCH v4 08/12] hostmem: Support fully shared guest memfd to back a VM
Daniel P. Berrangé <[email protected]>
| Newsgroups | org.nongnu.qemu-devel |
|---|---|
| Message-ID | <[email protected]> |
On Thu, Aug 13, 2026 at 05:10:45PM -0500, Michael Roth wrote: > On Thu, Aug 13, 2026 at 01:48:21PM +0100, Daniel P. Berrangé wrote: > > On Thu, Aug 13, 2026 at 08:28:24AM -0400, Peter Xu wrote: > > > On Thu, Aug 13, 2026 at 09:24:22AM +0100, Daniel P. Berrangé wrote: > > > > On Wed, Aug 12, 2026 at 03:16:46PM -0500, Michael Roth wrote: > > > > > From: Peter Xu <[email protected]> > > > > > > > > > > Host backends supports guest-memfd now by detecting whether it's a > > > > > confidential VM. There's no way to choose it yet from the memory level to > > > > > use it fully shared. If we use guest-memfd, it so far always implies we > > > > > need two layers of memory backends, while the guest-memfd only provides the > > > > > private set of pages. > > > > > > > > > > This patch introduces a way so that QEMU can consume guest memfd as the > > > > > only source of memory to back the object (aka, fully shared). > > > > > > > > > > To use the fully shared guest-memfd, one can add a memfd object with: > > > > > > > > > > -object memory-backend-memfd,guest-memfd=on,share=on > > > > > > > > > > Note that share=on is required with fully shared guest_memfd. > > > > > > > > > > PS: there's a trivial touch-up on fd<0 check, because the stub to create > > > > > guest-memfd may return negative but not -1. > > > > > > > > > > Signed-off-by: Peter Xu <[email protected]> > > > > > Reviewed-by: Xiaoyao Li <[email protected]> > > > > > Reviewed-by: Fabiano Rosas <[email protected]> > > > > > Signed-off-by: Michael Roth <[email protected]> > > > > > --- > > > > > backends/hostmem-memfd.c | 56 ++++++++++++++++++++++++++++++++++++---- > > > > > qapi/qom.json | 6 ++++- > > > > > 2 files changed, 56 insertions(+), 6 deletions(-) snip > > > > > diff --git a/qapi/qom.json b/qapi/qom.json > > > > > index c55776af7d..ee981fc44c 100644 > > > > > --- a/qapi/qom.json > > > > > +++ b/qapi/qom.json > > > > > @@ -771,13 +771,17 @@ > > > > > # @seal: if true, create a sealed-file, which will block further > > > > > # resizing of the memory (default: true) > > > > > # > > > > > +# @guest-memfd: if true, use guest-memfd to back the memory region. > > > > > +# (default: false, since: 11.2) > > > > > +# > > > > > # Since: 2.12 > > > > > ## > > > > > { 'struct': 'MemoryBackendMemfdProperties', > > > > > 'base': 'MemoryBackendProperties', > > > > > 'data': { '*hugetlb': 'bool', > > > > > '*hugetlbsize': 'size', > > > > > - '*seal': 'bool' }, > > > > > + '*seal': 'bool', > > > > > + '*guest-memfd': 'bool' }, > > > > > 'if': 'CONFIG_LINUX' } > > > > > > > > We're reusing the 'memory-backend-memfd' class, and then at runtime > > > > refusing allow the user to control any of properties in > > > > MemoryBackendProperties. > > > > > > gmemfd should be able to use all ultimately. > > > > > > For seal, IMHO it's already implied, kind of forced seal=on but it doesn't > > > matter, gmemfd was introduced with sealing, at least what QEMU implies with > > > "F_SEAL_GROW | F_SEAL_SHRINK | F_SEAL_SEAL". So IMHO we could ignore what > > > user selected and assume it's ON. > > The naming threw me off, but looking at the actual schema description for > 'seal', it basically implies 'fixed-size=on', which guest_memfd does > enforce. So, aside from the question of whether to reuse the memfd backend, > I do think it's more correct to imply seal=on for guest_memfd. > > > > > Then we should not have a 'seal' property defined for guest memfd > > at all. Defining a property and then ignoring it, or only ever > > allowing 1 value to be set is a design mistake. The property should > > not exist if it can't ever be changed by the user/app. > > We have a couple examples of what seems like something similar with > memory-backend-file,rom=on,readonly=off and > memory-backend-*,prealloc=on,reserve=off. Granted, those seem a little > more obvious to recognize as mutually-exclusive, but as far as the > discussion around libvirt/mgmt/introspection: are there mechanisms in > place already to handle cases like that? Or are these special-cases that > would be problematic to try to handle similarly? The difference only matters if what QEMU supports changes over time. eg if current QEMU required 'reserve=off', but later changed to allow either reserve=on or reserve=off, as libvirt can't introspect the built-in runtime restriction. > Regarding hugepage options, there will undoubtedly be kernels that support > guest_memfd but not hugetlb, so even if we introduce > memory-backend-guest-memfd now so that we can add options when/where it > makes sense, libvirt/mgmt. would still need to eventually handle a > 'hugetlb' option existing, but not necessarilly implying that the guest can > actually use them. It seems like the issue exists in either case. Yes, the varying kernel dependancy makes the QEMU introspection less valuable than it would otherwise be :-( > FWIW, I do anticipate that we will need something like > memory-backend-guest-memfd for some of the use-cases coming down the > pipeline, but for some stuff like CXL/HBM memory support where folks are > talking about stuff like custom guest_memfd allocators or special-purpose > NUMA nodes I'm not reasonably confident that memory-backend-guest-memfd won't > itself end up being too generic of a construct and be immediately relegated > to only handling the exact same set of options as memory-backend-memfd (give > or take a 'seal'). > > Peter's suggested approach allows us to assume less about how things will > eventually look by reusing existing options/command-lines and handling things > underneath the covers for the more basic use-cases in the meantime. Once we > hit cases that clearly have no business in memory-backend-memfd, we won't > really have lost anything as far as our options to introduce > memory-backend-guest-memfd at that point or maybe some subclass or something > else entirely. Hmm, if we think that even a separate memory-backend-guest-memfd is not going to suitable / sufficiently flexible for future enhancements, that reduces its value somewhat :-( > > > This is indeed what Michael used to suggest, and we were discussing in > > > previous version on which is better, > > > > > > https://lore.kernel.org/r/rjqfiwh57gip3u3psqg33jhmo7ixaj2qwzupc7zdk7f3d26qnu@tglactz67ogk > > > > > > The hope is this is also easier for either libvirt or most users, but > > > please correct me if it's not the case, especially for libvirt. The plan > > > is when CoCo flags are provided, all things will automatically switch to a > > > CoCo-friendly implementation within QEMU. > > > > > > It also means here the guest-memfd= parameter shouldn't be needed in real > > > CoCo contexts because they'll simply be implied (no cmdline change needed > > > for the same "-object memory-backend-memfd" one used to use without CoCo). > > > It's only needed for only special use of guest-memfd, in this case > > > init-shared is the special case where CoCo doesn't use. > > > > Reading all this, IMHO reusing memory-backend-memfd for the current > > Coco support was a design mistake, it should have have a > > memory-backend-guest-memfd object from the start. > > I think your point still stands, but the current CoCo support doesn't rely > on memory-backend-memfd, the private guest_memfd instance is handled by > QEMU completely separately as a function of whether or not we are running > a CoCo VM. The backends are only for shared memory, which aren't relevant > to guest_memfd (without in-place conversion support), so users can select > whatever they'd like. Ah, I missed that the guest_memfd stuff for Coco is magically handled with QEMU able to "do the right thing" there. > However, backends like memory-backend-file tend to get used for special > cases like persisting memory, where there is a high chance of surprises > for users trying to make use of these sorts of things for a CoCo VM. > That's why memory-backend-memfd is the configuration we normally suggest > for CoCo VMs, but it's not required, since nothing stops users from > persisting shared memory ranges if that's actually their intent. > > The need to enforce the backend more explicitly becomes much more apparent > once in-place conversion however, where even shared memory goes through > guest-memfd and so the backend necessarily needs to be > guest_memfd-aware. The in-place conversion series tracks this via a > RAMBlock flag that can be set by whatever backends we decide would be > appropriate for managing guest_memfd instances based on these > discussions. > > So, at least for users following the recommended configuration: > > qemu -object sev-snp-guest,... > -object memory-backend-memfd,... > > they can then switch on in-place conversion via, e.g.: > > qemu -object sev-snp-guest,...,convert-in-place=on > -object memory-backend-memfd,... > > and we can flip guest_memfd=on automatically underneath the covers > so that both shared/private memory go through guest_memfd as > convert-in-place would necessarily imply. Yes, if we think we can make QEMU "do the right thing" internally that makes a separate object much less desirable. That would be saying that we have a tri-state, guest_memfd=auto|on|off, and such tri-state's would not invite a separate sub-class design. The need for "auto" rather forces the approach you have in this patch. With regards, Daniel -- |: https://berrange.com ~~ https://hachyderm.io/@berrange :| |: https://libvirt.org ~~ https://entangle-photo.org :| |: https://pixelfed.art/berrange ~~ https://fstop138.berrange.com :|