Re: [PATCH v10 12/41] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion

Binbin Wu <[email protected]>
Newsgroups org.kernel.vger.linux-trace-kernel,dev.linux.lists.linux-coco,org.kernel.vger.kvm,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel,org.kernel.vger.linux-kselftest,org.kvack.linux-mm
Message-ID <[email protected]>
On 8/8/2026 5:52 AM, Ackerley Tng via B4 Relay wrote:
> From: Ackerley Tng <[email protected]>
> 
> When memory in guest_memfd is converted from private to shared, the
> platform-specific state associated with the guest-private pages must be
> invalidated or cleaned up.
> 
> Iterate over the folios in the affected range and call the
> kvm_arch_gmem_make_shared() hook for each PFN range. This allows
> architectures to update hardware metadata or encryption states to
> transition pages to the shared state.
> 
> Invoke this helper after indicating to KVM's mmu code that an invalidation
> is in progress to stop in-flight page faults from succeeding.
> 
> Omit support for calling the arch hook to make private, since SNP, the only
> implementer of the arch make-private hook today, would actually prefer
> making private only just before faulting memory into the NPTs.
> 
> Calling the make-private arch hook would require iterating both bindings
> and the filemap to find the intersection of bindings and allocated
> folios. On top of that, SNP would need to figure out whether to actually
> make private based on whether the memory is about the be faulted, or
                                                     ^
the -> to?

> whether it is a conversion.
> 

[...]

>  
> +#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT
> +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end)
> +{
> +	struct folio_batch fbatch;
> +	pgoff_t next = start;
> +	int i;
> +
> +	folio_batch_init(&fbatch);
> +	while (filemap_get_folios(inode->i_mapping, &next, end - 1, &fbatch)) {
> +		for (i = 0; i < folio_batch_count(&fbatch); ++i) {
> +			struct folio *folio = fbatch.folios[i];
> +			pgoff_t start_index, end_index;
> +			kvm_pfn_t start_pfn;
> +			kvm_pfn_t nr_pages;
> +
> +			start_index = max(start, folio->index);
> +			end_index = min(end, folio_next_index(folio));
> +			/*
> +			 * end_index is either in folio or points to
> +			 * the first page of the next folio. Hence,
> +			 * all pages in range [start_index, end_index)
> +			 * are contiguous.
> +			 */
> +			start_pfn = folio_file_pfn(folio, start_index);
> +			nr_pages = end_index - start_index;
> +
> +			kvm_arch_gmem_make_shared(start_pfn, nr_pages);
> +		}
> +
> +		folio_batch_release(&fbatch);
> +		cond_resched();
> +	}
> +}
> +#else
> +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) {}
> +#endif
> +
>  static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
>  				     size_t nr_pages, uint64_t attrs,
>  				     pgoff_t *err_index)
> @@ -599,7 +636,12 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
>  
>  	filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE;
>  	kvm_gmem_invalidate_start(inode, start, end, filter);
> +
> +	if (!to_private)
> +		kvm_gmem_make_shared(inode, start, end);

If both KVM_AMD_SEV and KVM_INTEL_TDX are enabled, HAVE_KVM_ARCH_GMEM_CONVERT
will be enabled and the logic in kvm_gmem_make_shared() introduces unnecessary
overhead for TDX.
Not sure about CSPs, but in a standard distribution kernel, it's very likely
that both are enabled, right? Should kvm_gmem_make_shared() do some
optimization or the overhead is relative small in the conversion to shared
path so that the optimization is not worth it? 


> +
>  	mas_store_prealloc(&mas, xa_mk_value(attrs));
> +
>  	kvm_gmem_invalidate_end(inode, start, end);
>  out:
>  	filemap_invalidate_unlock(mapping);
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.