Re: [PATCH v2 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race
Marc Zyngier <[email protected]>
| Newsgroups | org.kernel.vger.stable,dev.linux.lists.kvmarm,org.infradead.lists.linux-arm-kernel,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On Mon, 24 Aug 2026 14:15:29 +0100, "Lorenzo Stoakes (ARM)" <[email protected]> wrote: > > On Sun, Aug 23, 2026 at 08:53:31AM +0100, Marc Zyngier wrote: > > On Sat, 22 Aug 2026 18:46:52 +0100, > > "Lorenzo Stoakes (ARM)" <[email protected]> wrote: > > > > > > When GFNs are invalidated in L0 an MMU notifier triggers > > > kvm_unmap_gfn_range() which tears down all of the stage 2 shadow page > > > tables for nested guests via kvm_nested_s2_unmap(). > > > > > > To avoid lockup, the kvm->mmu_lock is dropped while doing this and the task > > > rescheduled once for each block of physical address space (32 MiB for 16 > > > KiB page size), with the lock being reacquired once the task is scheduled > > > again. > > > > > > This results in a potential race between this L0 tear down and tear down of > > > the guest itself in kvm_flush_shadow_all(), a race which has been observed > > > on real hardware. > > > > > > When this race occurs it causes an invalid kernel warning when the PGT of a > > > nested MMU is cleared by kvm_flush_shadow_all() -> > > > kvm_arch_flush_shadow_all() -> kvm_free_stage2_pgd(). > > > > > > Patch 1 fixes this by having stage2_apply_range() no longer return an error > > > when it has experienced a benign race with pgt teardown when it drops the > > > lock. > > > > > > Patch 2 addresses something more serious - bad timing can turn this spurious > > > warning into a NULL pointer dereference. > > > > > > kvm_arch_flush_shadow_all() calls kvm_uninit_stage2_mmu() which calls > > > kvm_free_stage2_pgd() on the canonical kvm->arch.mmu for that guest's S2 > > > mappings, making it NULL. > > > > > > This is problematic if it happens before stage2_apply_range() reacquires > > > the kvm->mmu_lock, as it ultimately returns to kvm_nested_s2_unmap() which > > > dereferences kvm->arch.mmu.pgt with the mmu lock held under the incorrect > > > assumption that it means it's valid, resulting in a NULL pointer > > > dereference. > > > > > > Fix that by checking if kvm->arch.mmu.pgt is NULL before dereferencing it > > > in kvm_nested_s2_unmap() and kvm_nested_s2_wp(). > > > > With the commit message for patch #1 trimmed to something that fits on > > a couple of standard terminal screens ;-) : > > Haha sure will put it on a diet and respin :) Oliver can probably do so when applying the series. Thanks, M. -- Jazz isn't dead. It just smells funny.