Re: [PATCH v10 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion

Yan Zhao <[email protected]>
Newsgroups org.kernel.vger.linux-kselftest,dev.linux.lists.linux-coco,org.kernel.vger.kvm,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel,org.kernel.vger.linux-trace-kernel,org.kvack.linux-mm
Message-ID <[email protected]>
On Sat, Aug 08, 2026 at 08:29:15AM +0800, Yan Zhao wrote:
> On Fri, Aug 07, 2026 at 02:52:50PM -0700, Ackerley Tng via B4 Relay wrote:
> > +static bool kvm_gmem_is_safe_for_conversion(struct inode *inode, pgoff_t start,
> > +					    size_t nr_pages, pgoff_t *err_index)
> > +{
> > +	struct address_space *mapping = inode->i_mapping;
> > +	const int filemap_get_folios_refcount = 1;
> > +	pgoff_t last = start + nr_pages - 1;
> > +	struct folio_batch fbatch;
> > +	bool safe = true;
> > +	pgoff_t next;
> > +	int i;
> > +
> > +	folio_batch_init(&fbatch);
> > +
> > +	next = start;
> > +	while (safe && filemap_get_folios(mapping, &next, last, &fbatch)) {
> > +		for (i = 0; i < folio_batch_count(&fbatch); ++i) {
> > +			struct folio *folio = fbatch.folios[i];
> > +
> > +			if (folio_ref_count(folio) !=
> > +			    folio_nr_pages(folio) + filemap_get_folios_refcount) {
> > +				safe = false;
> > +				*err_index = max(start, folio->index);
> > +				break;
> > +			}
> > +		}
> > +
> > +		folio_batch_release(&fbatch);
> > +		cond_resched();
> > +	}
> > +
> > +	return safe;
> > +}
> > +
> >  static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> > -				     size_t nr_pages, uint64_t attrs)
> > +				     size_t nr_pages, uint64_t attrs,
> > +				     pgoff_t *err_index)
> >  {
> >  	bool to_private = attrs & KVM_MEMORY_ATTRIBUTE_PRIVATE;
> >  	struct address_space *mapping = inode->i_mapping;
> > @@ -542,8 +576,21 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> >  
> >  	mas_init(&mas, mt, start);
> >  	r = kvm_gmem_mas_preallocate(&mas, attrs, start, nr_pages);
> > -	if (r)
> > +	if (r) {
> > +		*err_index = start;
> >  		goto out;
> > +	}
> > +
> > +	if (to_private) {
> > +		unmap_mapping_pages(mapping, start, nr_pages, false);
> > +
> > +		if (!kvm_gmem_is_safe_for_conversion(inode, start, nr_pages,
> > +						     err_index)) {
> Note: conversion failures could occur if another vCPU is attempting to map a GFN
> within this range.
> 
> CPU 0 (setting attributes)          CPU 1 (attempting to map)
> --------------------------          --------------------
>                                  A: mmu_invalidate_retry_gfn_unsafe
>                                     filemap_invalidate_lock_shared
>                                     __kvm_gmem_get_pfn ==> folio refcount++
>                                     filemap_invalidate_unlock_shared
> 
> filemap_invalidate_lock
> filemap_get_folios
> check folio_ref_count(folio) ==> Not match !!
> filemap_invalidate_unlock
> 
>                                  B: read_lock(&vcpu->kvm->mmu_lock);
>                                     is_page_fault_stale
>                                     kvm_mmu_finish_page_fault ==>folio recount--
> 				    read_unlock(&vcpu->kvm->mmu_lock);
> 
> 
> Retrying in kvm_gmem_is_safe_for_conversion() or moving the invocation of
> kvm_mmu_invalidate_start() + kvm_mmu_invalidate_range_add() to an earlier
> position does not help as long as CPU 1 stays at stage A.
> 
> So, should we avoid this failure?
> e.g., by moving filemap_invalidate_unlock_shared() from stage A to after
> stage B?

Or what about having KVM always treat gmem page as non-refcounted, and have
kvm_gmem_get_pfn() put folio refcount before releasing the filemap invalidate
lock?
Below patch is applied and tested at the end of this series.

From 8c2f29bc15bceb6a8fa103cf2585ec11354fd74e Mon Sep 17 00:00:00 2001
From: Yan Zhao <[email protected]>
Date: Mon, 10 Aug 2026 06:24:52 +0800
Subject: [PATCH] KVM: guest_memfd: Return gmem page as non-refcounted

Have kvm_gmem_get_pfn() put gmem page refcount before releasing filemap
invalidate lock and return the gmem page as non-refcounted. This avoids
gmem memory attribute conversion failure caused by temporarily holding gmem
page after faulting and before completing mapping.

guest_memfd always holds gmem page in filemap cache. TDX does not increment
gmem page refcount when having gmem pages mapped in S-EPT. Additionally,
as gmem pages are not swappable, setting dirty or accessed bit is not
necessary. Therefore, there's no need to treat gmem pages as refcounted
pages.

Signed-off-by: Yan Zhao <[email protected]>
---
 virt/kvm/guest_memfd.c | 5 ++---
 1 file changed, 2 insertions(+), 3 deletions(-)

diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 2115e73e455a..e357b4ffa777 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -1332,11 +1332,10 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot,
 #endif
 
 	folio_unlock(folio);
+	folio_put(folio);
 
 	if (!r)
-		*page = folio_file_page(folio, index);
-	else
-		folio_put(folio);
+		*page = NULL;
 
 out:
 	filemap_invalidate_unlock_shared(file_inode(file)->i_mapping);
-- 
2.43.2
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.