Re: [PATCH v10 12/41] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion
Binbin Wu <[email protected]>
| Newsgroups | org.kernel.vger.linux-trace-kernel,dev.linux.lists.linux-coco,org.kernel.vger.kvm,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel,org.kernel.vger.linux-kselftest,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
On 8/8/2026 5:52 AM, Ackerley Tng via B4 Relay wrote: > From: Ackerley Tng <[email protected]> > > When memory in guest_memfd is converted from private to shared, the > platform-specific state associated with the guest-private pages must be > invalidated or cleaned up. > > Iterate over the folios in the affected range and call the > kvm_arch_gmem_make_shared() hook for each PFN range. This allows > architectures to update hardware metadata or encryption states to > transition pages to the shared state. > > Invoke this helper after indicating to KVM's mmu code that an invalidation > is in progress to stop in-flight page faults from succeeding. > > Omit support for calling the arch hook to make private, since SNP, the only > implementer of the arch make-private hook today, would actually prefer > making private only just before faulting memory into the NPTs. > > Calling the make-private arch hook would require iterating both bindings > and the filemap to find the intersection of bindings and allocated > folios. On top of that, SNP would need to figure out whether to actually > make private based on whether the memory is about the be faulted, or ^ the -> to? > whether it is a conversion. > [...] > > +#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT > +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) > +{ > + struct folio_batch fbatch; > + pgoff_t next = start; > + int i; > + > + folio_batch_init(&fbatch); > + while (filemap_get_folios(inode->i_mapping, &next, end - 1, &fbatch)) { > + for (i = 0; i < folio_batch_count(&fbatch); ++i) { > + struct folio *folio = fbatch.folios[i]; > + pgoff_t start_index, end_index; > + kvm_pfn_t start_pfn; > + kvm_pfn_t nr_pages; > + > + start_index = max(start, folio->index); > + end_index = min(end, folio_next_index(folio)); > + /* > + * end_index is either in folio or points to > + * the first page of the next folio. Hence, > + * all pages in range [start_index, end_index) > + * are contiguous. > + */ > + start_pfn = folio_file_pfn(folio, start_index); > + nr_pages = end_index - start_index; > + > + kvm_arch_gmem_make_shared(start_pfn, nr_pages); > + } > + > + folio_batch_release(&fbatch); > + cond_resched(); > + } > +} > +#else > +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) {} > +#endif > + > static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start, > size_t nr_pages, uint64_t attrs, > pgoff_t *err_index) > @@ -599,7 +636,12 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start, > > filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE; > kvm_gmem_invalidate_start(inode, start, end, filter); > + > + if (!to_private) > + kvm_gmem_make_shared(inode, start, end); If both KVM_AMD_SEV and KVM_INTEL_TDX are enabled, HAVE_KVM_ARCH_GMEM_CONVERT will be enabled and the logic in kvm_gmem_make_shared() introduces unnecessary overhead for TDX. Not sure about CSPs, but in a standard distribution kernel, it's very likely that both are enabled, right? Should kvm_gmem_make_shared() do some optimization or the overhead is relative small in the conversion to shared path so that the optimization is not worth it? > + > mas_store_prealloc(&mas, xa_mk_value(attrs)); > + > kvm_gmem_invalidate_end(inode, start, end); > out: > filemap_invalidate_unlock(mapping); >