Re: [PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure)

Itaru Kitayama <[email protected]>
Newsgroups dev.linux.lists.kvmarm,org.infradead.lists.linux-arm-kernel,org.kernel.vger.linux-kernel
Message-ID <anvWfkxd1b97xxFw@sm-arm-grace07>
On Mon, Aug 10, 2026 at 09:50:32PM +0100, Wei-Lin Chang wrote:
> Hi,
> 
> This is v5 of optimizing the shadow s2 mmu unmapping during MMU
> notifiers.

I've tested your series v5 on a Grace system. L2 booted into prompt
with the Ubuntu filesystem image, your two kvm selftest for nested 
virtualization ran fine in L1, and also did stress-ng in L1:

projects $ sudo stress-ng --kvm 16 --cpu 16 --vm 8 --vma 8 --fork 8 --timeout 10m --verify    --metrics-brief
stress-ng: info:  [1053] setting to a 10 mins run per stressor
stress-ng: info:  [1053] dispatching hogs: 16 kvm, 16 cpu, 8 vm, 8 vma, 8 fork
stress-ng: info:  [1087] vm: using 32MB per stressor instance (total 256MB of 2.75GB available memory)
stress-ng: metrc: [1053] stressor       bogo ops real time  usr time  sys time   bogo ops/s     bogo ops/s
stress-ng: metrc: [1053]                           (secs)    (secs)    (secs)   (real time) (usr+sys time)
stress-ng: metrc: [1053] kvm                  81    602.31    250.98    357.55         0.13           0.13
stress-ng: metrc: [1053] cpu               63310    595.70    166.40      0.60       106.28         379.12
stress-ng: metrc: [1053] vm              4252063    600.90     38.98     56.59      7076.16       44491.45
stress-ng: metrc: [1053] vma               56633    601.76     12.64    157.96        94.11         331.97
stress-ng: metrc: [1053] fork                149    600.98      0.03      0.45         0.25         314.59
stress-ng: info:  [1053] skipped: 0
stress-ng: info:  [1053] passed: 56: kvm (16) cpu (16) vm (8) vma (8) fork (8)
stress-ng: info:  [1053] failed: 0
stress-ng: info:  [1053] metrics untrustworthy: 0
stress-ng: info:  [1053] successful run completed in 10 mins 8.62 secs

Tested-by: Itaru Kitayama <[email protected]>

Thanks,
Itaru.

> 
> This time, a major overhaul is done to the implementation. After
> receiving some suggestions from Marc, I have identified that using the
> interval tree to store the guest stage-2 mappings solves many problems
> compared to using the maple tree.
> 
> Interval Tree vs Maple Tree
> ===========================
> 
> First of all, interval trees are capable of storing overlapping ranges,
> which is helpful when the L1 hypervisor maps something like:
> 
> nested IPA [x, x+4K) -> canonical IPA [a, a+4K)
> nested IPA [y, y+2M) -> canonical IPA [a, a+2M)
> 
> No problems with storing that in the interval tree with different nodes.
> We can avoid the maple tree UNKNOWN_IPA mechanism as a compromise.
> 
> Second, ideally we would want to save the canonical IPA <-> nested IPA
> mapping in both directions to allow MMU notifier unmap speed up, and
> stale shadow mapping removals. If we use the maple tree, we'll have to
> have 2 separate trees, and make sure they store the same mappings, which
> isn't simple given the first point.
> 
> On the other hand, by using this pattern:
> 
> /* Record of a guest stage-2 mapping. */
> struct kvm_guest_s2_mapping {
>        struct interval_tree_node canonical; // CIPA range of the mapping
>        struct interval_tree_node nested;    // NIPA range of the mapping
>        struct kvm_s2_mmu *nested_mmu;       // mmu of the NIPA space
> };
> 
> and equip each mmu with an interval tree storing mapping records
> corresponding to the IPA space it represents, we can insert the
> respective nodes into the canonical IPA tree, and the corresponding
> nested IPA tree. This makes it trivial to find the range of the other
> IPA space from a range in one IPA space.
> 
> Diagram to help understanding:
> 
> struct kvm_guest_s2_mapping mapping1, mapping2;
> 
>     ---------------------> mapping2.canonical
>     |                      mapping1.canonical
>     |                          ^   (both stored in canonical mmu's tree)
>     |                          |
> --*****-----------------------*****----------- CIPA
>    \\\\\                      |||||                  mapping1.nested_mmu
>     \\\\\                     \\\\\                            |
>      \\\\\                     \\\\\                           v
> ------\\\\\---------------------*****--------- NIPA #1 (nested mmu #1)
>        \\\\\                      |
>         \\\\\                     -> mapping1.nested
>          \\\\\                       (stored in nested mmu #1's tree)
>           \\\\\
> -----------*****------------------------------ NIPA #2 (nested mmu #2)
>              |                                                 ^
>              -> mapping2.nested                                |
>                 (stored in nested mmu #2's tree)   mapping2.nested_mmu
> 
> Third, maple tree does its own memory allocation. In the KVM stage-2
> fault path we only find out what the mapping ranges are after taking the
> KVM MMU lock, and the maple tree has to know the range and entry to be
> stored to preallocate, therefore in our case the maple tree is forced to
> only use GFP_NOWAIT, which isn't the best. With the interval tree the
> user does the memory management, and we can just allocate before taking
> the locks.
> 
> Locking
> =======
> 
> The guest_s2_tracking_lock serializes accesses to the tracking interval
> trees. It is taken after the mmu_lock. However in reality it is only
> taken after we take the read mmu_lock in the stage-2 fault path, as
> other accesses have the write mmu_lock already. This saves us some
> manual lock/unlocks.
> 
> vCPU Stage-2 Fault Scalability Reduction
> ========================================
> 
> KVM/arm64 is able to handle stage-2 faults from multiple vCPUs in
> parallel, thanks to the engineering done to the s2 pgtable code. However
> to safely insert mappings into the interval trees we have to serialize
> using the guest_s2_tracking_lock. We trade some performance in stage-2
> fault for faster MMU notifier unmaps, and keeping the unaffected shadow
> mappings.
> 
> Memory Usage
> ============
> 
> Each interval tree node is 48 bytes, and a kvm_guest_s2_mapping is 104
> bytes, residing in 128-byte slab objects. Each shadow stage-2 fault
> requires one kvm_guest_s2_mapping instance. This is 32MB for a fully 4KB
> mapped 1GB region, and 64KB for a 2MB mapped 1GB region.
> 
> Series Structure
> ================
> 
> Patch 1:   Preparatory refactoring.
> Patch 2:   Introduce data structures for guest stage-2 tracking.
> Patch 3-4: Guest stage-2 tracking addition and removal
> Patch 5:   Avoid full unmap during MMU notifier unmap using the tracked
>            guest stage-2 mapping information.
> Patch 6:   Minor clean up.
> 
> As this is a complete rework, I will omit the change log this time.
> Series is based on v7.2-rc5.
> 
> Thanks!
> 
> Link to v4: https://lore.kernel.org/kvmarm/[email protected]/
> 
> Wei-Lin Chang (6):
>   KVM: arm64: Use a variable for the canonical IPA in kvm_s2_fault_map()
>   KVM: arm64: nv: Introduce guest stage-2 tracking structures
>   KVM: arm64: nv: Track guest stage-2 mapping creation
>   KVM: arm64: nv: Track guest stage-2 mapping removal
>   KVM: arm64: nv: Avoid full shadow stage-2 unmap
>   KVM: arm64: Refactor kvm_unmap_gfn_range() with common variables
> 
>  arch/arm64/include/asm/kvm_host.h   |  20 ++++++
>  arch/arm64/include/asm/kvm_nested.h |   7 ++
>  arch/arm64/kvm/mmu.c                | 105 ++++++++++++++++++++++++----
>  arch/arm64/kvm/nested.c             |  95 +++++++++++++++++++++++++
>  4 files changed, 215 insertions(+), 12 deletions(-)
> 
> -- 
> 2.43.0
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.