[PATCH v2 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race
"Lorenzo Stoakes (ARM)" <[email protected]>
| Newsgroups | dev.linux.lists.kvmarm,org.infradead.lists.linux-arm-kernel,org.kernel.vger.linux-kernel,org.kernel.vger.stable |
|---|---|
| Message-ID | <[email protected]> |
When GFNs are invalidated in L0 an MMU notifier triggers kvm_unmap_gfn_range() which tears down all of the stage 2 shadow page tables for nested guests via kvm_nested_s2_unmap(). To avoid lockup, the kvm->mmu_lock is dropped while doing this and the task rescheduled once for each block of physical address space (32 MiB for 16 KiB page size), with the lock being reacquired once the task is scheduled again. This results in a potential race between this L0 tear down and tear down of the guest itself in kvm_flush_shadow_all(), a race which has been observed on real hardware. When this race occurs it causes an invalid kernel warning when the PGT of a nested MMU is cleared by kvm_flush_shadow_all() -> kvm_arch_flush_shadow_all() -> kvm_free_stage2_pgd(). Patch 1 fixes this by having stage2_apply_range() no longer return an error when it has experienced a benign race with pgt teardown when it drops the lock. Patch 2 addresses something more serious - bad timing can turn this spurious warning into a NULL pointer dereference. kvm_arch_flush_shadow_all() calls kvm_uninit_stage2_mmu() which calls kvm_free_stage2_pgd() on the canonical kvm->arch.mmu for that guest's S2 mappings, making it NULL. This is problematic if it happens before stage2_apply_range() reacquires the kvm->mmu_lock, as it ultimately returns to kvm_nested_s2_unmap() which dereferences kvm->arch.mmu.pgt with the mmu lock held under the incorrect assumption that it means it's valid, resulting in a NULL pointer dereference. Fix that by checking if kvm->arch.mmu.pgt is NULL before dereferencing it in kvm_nested_s2_unmap() and kvm_nested_s2_wp(). v2: * Rebased onto next * Updated 1/2's commit message to say that it was all of the kvmtool hosts that were stopped, as per discussion with Wei-Lin and Yao Yuan. * Updated 2/2 to remove the !may_block WARN_ON() as duplicative, as per Marc. * Updated 2/2 to abstract the VNCR IPA invalidation in kvm_invalidate_vncr_ipa_all() and perform the same check for kvm_nested_s2_wp(), as per discussion with Marc and sashiko report. v1: https://lore.kernel.org/r/[email protected] Signed-off-by: Lorenzo Stoakes (ARM) <[email protected]> --- Lorenzo Stoakes (ARM) (2): KVM: arm64: Fix spurious warning for benign stage 2 teardown race KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race arch/arm64/kvm/mmu.c | 10 ++++++++-- arch/arm64/kvm/nested.c | 15 +++++++++++++-- 2 files changed, 21 insertions(+), 4 deletions(-) --- base-commit: aa8e5dc6a7a2a1141ab40706a51010adcd0e57d2 change-id: 20260811-kvm-arm-nested-virt-fix-031e9ab1be87 Best regards, -- Lorenzo Stoakes (ARM) <[email protected]>