Re: [PATCH v2 1/8] KVM: arm64: Remove VM-wide VNCR mapping counter

[email protected] Thu, 06 Aug 2026 09:35:41 +0000
Newsgroups dev.linux.lists.kvmarm,org.kernel.vger.kvm
Message-ID <[email protected]>
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

Pre-existing issues:
- [High] S1E2 TLBI instructions fail to invalidate the VNCR pseudo-TLB when TGE=0 or E2H=0.
- [High] Hardcoded `TR_EL20` in `kvm_translate_vncr` misinterprets `TCR_EL2` for non-VHE guests.
--

--- Patch [1]: [PATCH v2 1/8] KVM: arm64: Remove VM-wide VNCR mapping counter ---
commit 9a09630623c9341a08c8ecea540015307bde098c
Author: Marc Zyngier <[email protected]>
Subject: KVM: arm64: Remove VM-wide VNCR mapping counter

The global VNCR mapping counter is used to decide whether an L1
provided VNCR page is mapped in L0 on any CPU at the point of
dealing with a TLB invalidation. It is incremented when a mapping
is made in the fixmap, and decremented when unmapped.

> diff --git a/arch/arm64/kvm/hyp/vhe/switch.c b/arch/arm64/kvm/hyp/vhe/switch.c
> index bbe9cebd3d9d5..c09b1d411c584 100644
> --- a/arch/arm64/kvm/hyp/vhe/switch.c
> +++ b/arch/arm64/kvm/hyp/vhe/switch.c
> @@ -427,8 +427,7 @@ static bool kvm_hyp_handle_tlbi_el2(struct kvm_vcpu *vcpu, u64 *exit_code)
>  	 * If we have to check for any VNCR mapping being invalidated,
>  	 * go back to the slow path for further processing.
>  	 */
> -	if (vcpu_el2_e2h_is_set(vcpu) && vcpu_el2_tge_is_set(vcpu) &&
> -	    atomic_read(&vcpu->kvm->arch.vncr_map_count))
> +	if (vcpu_el2_e2h_is_set(vcpu) && vcpu_el2_tge_is_set(vcpu))

[Severity: High]
This is a pre-existing issue, but does this condition incorrectly limit the
slow-path fallback to only E2H=1 and TGE=1 contexts?

If a non-VHE guest (E2H=0) or a VHE guest running a VM (TGE=0) executes an
S1E2 TLBI instruction (e.g. TLBI VAE2), the fast-path handler
kvm_hyp_handle_tlbi_el2() performs the hardware TLBI but will evaluate the
above condition as false.

It will then execute __kvm_skip_instr() and return true, bypassing the EL1
slow-path handler entirely.

Since the EL1 handler is exclusively responsible for invalidating the
software VNCR pseudo-TLB, won't this leave stale mappings in the pseudo-TLB
when a guest hypervisor invalidates its EL2 mappings?

>  		return false;
>  
>  	__kvm_skip_instr(vcpu);

> diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c
> index d9486010835a7..983ba0adaf4e6 100644
> --- a/arch/arm64/kvm/nested.c
> +++ b/arch/arm64/kvm/nested.c

[ ... ]

[Severity: High]
This is also a pre-existing issue in this file within kvm_translate_vncr(),
where the TR_EL20 translation regime is hardcoded:

kvm_translate_vncr() {
    ...
    vt->wi = (struct s1_walk_info) {
        .regime = TR_EL20,
        .as_el0 = false,
        .pan    = false,
    };
    ...
}

Could this misinterpret TCR_EL2 for non-VHE guests?

When setup_s1_walk() parses TCR_EL2 for the TR_EL20 regime, it assumes the
E2H=1 register layout (reading TBI from bit 37 and PS from bits 34:32).

However, an untrusted nested guest hypervisor with E2H=0 programs TCR_EL2
using the standard E2H=0 layout, where TBI is at bit 20 and PS is at bits
18:16.

Won't this cause KVM to use garbage bits to configure the walk, leading to
spurious translation faults or using the wrong physical address size?

-- 
Sashiko AI review ยท https://sashiko.dev/#/patchset/[email protected]?part=1