Re: [PATCH v4 11/48] KVM: arm64: gic-v5: Implement VMT/vIST IRS MMIO Ops

Sascha Bischoff <[email protected]> Fri, 7 Aug 2026 08:12:55 +0000
Newsgroups dev.linux.lists.kvmarm,dev.linux.lists.sashiko-reviews,org.kernel.vger.kvm
Message-ID <[email protected]>
On Fri, 2026-07-24 at 11:20 +0000, [email protected] wrote:
> Thank you for your contribution! Sashiko AI review found 6 potential
> issue(s) to consider:
> - [Critical] Timeout of IRS MMIO commands leads to hardware Use-
> After-Free for IST/VMTE memory.
> - [Critical] Cache line false sharing on non-coherent GICv5 VMT
> causes silent memory corruption.
> - [Critical] Unsafe use of kzalloc for hardware data structures
> requiring strict alignment.
> - [High] Missing wait for idle before writing to IRS VMT
> configuration registers.
> - [High] GICv5 probe fails systematically across kexec because it
> rejects an already valid VMT base register.
> - [Medium] The newly allocated VMT is leaked if
> vgic_v5_irs_assign_vmt() returns -EBUSY and vgic_v5_irs_clear_vmt()
> fails.
> --
> 
> --- Patch [11]: [PATCH v4 11/48] KVM: arm64: gic-v5: Implement
> VMT/vIST IRS MMIO Ops ---
> commit ef6f1e2217f9b0055cd57ec9bca6826104eed2cf
> Author: Sascha Bischoff <[email protected]>
> 
> KVM: arm64: gic-v5: Implement VMT/vIST IRS MMIO Ops
> 
> This commit adds support for making level 1 VMTs valid, allowing
> dynamic
> level 2 array allocation. It also implements IRS MMIO operations for
> making VMs and ISTs valid or invalid.
> 
> > --- a/arch/arm64/kvm/vgic/vgic-v5-tables.c
> > +++ b/arch/arm64/kvm/vgic/vgic-v5-tables.c
> 
> [ ... ]
> 
> > @@ -730,8 +738,13 @@ static int vgic_v5_vmte_assign_ist(struct kvm
> > *kvm, phys_addr_t ist_base,
> >  	/* Finally, mark the entry as valid */
> >  	cmd = spi_ist ? SPI_VIST_MAKE_VALID : LPI_VIST_MAKE_VALID;
> >  	ret = irq_set_vcpu_affinity(vgic_v5_vpe_db(vcpu0), &cmd);
> > +	if (ret) {
> > +		WRITE_ONCE(vmte->val[section], 0ULL);
> > +		vgic_v5_clean_inval(vmte, sizeof(*vmte));
> > +		return ret;
> > +	}
> >  
> > -	return ret;
> > +	return 0;
> >  }
> 
> [Severity: Critical]
> If the IRS MMIO command times out and returns an error here, KVM
> zeroes the
> VMTE and the caller frees the IST memory. Is it possible that the
> hardware
> IRS is still processing the delayed command?

This would only realistically happen on very broken hardware. We wait
for up to 1 second, which should be more than enough time for the
hardware to complete such a request.

> 
> If the hardware eventually completes the command, it might use the
> physical
> address KVM just freed, potentially causing a Use-After-Free DMA by
> writing
> to the freed memory.

Again, this would be very broken hardware if it was going this wrong.
If the MMIO command times out, then we're already in the situation
where KVM's view might differ from the hardware's view. Things will
only get worse from here, I'm afraid.

> 
> [Severity: Critical]
> Could there be a false-sharing issue leading to memory corruption on
> non-coherent GICv5 VMT systems?
> 
> The vmtl2_entry structure is 32 bytes:
> 
> arch/arm64/kvm/vgic/vgic-v5-tables.h:
>     struct vmtl2_entry {
>         __le64 val[4];
>     };
> 
> When modifying a VMTE, the cache line is flushed:
> 
> arch/arm64/kvm/vgic/vgic-v5-tables.c:vgic_v5_clean_inval() {
>     ...
>     if (kvm_vgic_global_state.vgic_v5_irs_caps.non_coherent)
>         dcache_clean_inval_poc(base, base + size);
> }
> 
> Since ARM64 cache lines are at least 64 bytes, multiple VMTEs share a
> single
> cache line. If the CPU writes back this dirty cache line, won't it
> overwrite
> any concurrent updates the IRS made directly to adjacent VMTEs in
> main memory?

I believe that the way that this was written it could have happened.
I've extended the locking around these to ensure that there is never a
case where more than one VMTE is being operated on at any one time.
This should avoid any false sharing issues.

> 
> [Severity: Critical]
> Are the alignment requirements for hardware data structures
> guaranteed when
> using kzalloc()?
> 
> In vgic_v5_vmte_init(), vmd is allocated with kzalloc() and its
> physical
> address is bit-shifted:
> 
> arch/arm64/kvm/vgic/vgic-v5-tables.c:vgic_v5_vmte_init() {
>     ...
>     vmd = kzalloc(vmt_info->vmd_size, GFP_KERNEL);
>     ...
>     tmp = FIELD_PREP(GICV5_VMTEL2E_VMD_ADDR,
>                      virt_to_phys(vmd) >>
> GICV5_VMTEL2E_VMD_ADDR_SHIFT);
>     ...
> }
> 
> Since kzalloc() only guarantees natural alignment, shifting the
> address right
> by GICV5_VMTEL2E_VMD_ADDR_SHIFT could truncate the lower bits if the
> allocation isn't perfectly aligned to 1 << SHIFT. Could this lead to
> the
> hardware accessing the wrong memory location?

kzalloc() will provide naturally aligned memory, with alignment being
at least ARCH_KMALLOC_MINALIGN, which is 8 for arm64. We already ensure
that the allocation is large enough that those lower bits are not set,
so that shift should never lose bits. FWIW, the shift is there because
we are using FIELD_PREP() which both masks and shifts, whereas for the
address we only really want to zero the lower bits. We're effectively
undoing the shift that FIELD_PREP() is implicitly applying.

> 
> > --- a/arch/arm64/kvm/vgic/vgic-v5.c
> > +++ b/arch/arm64/kvm/vgic/vgic-v5.c
> 
> [ ... ]
> 
> > @@ -105,13 +131,55 @@ int vgic_v5_probe(const struct gic_kvm_info
> > *info)
> 
> [ ... ]
> 
> > +	ret =
> > vgic_v5_irs_assign_vmt(irs_caps.two_level_vmt_support,
> > +				     ilog2(irs_caps.max_vms),
> > +				     vgic_v5_get_vmt_base());
> > +	if (ret) {
> > +		kvm_err("Failed to assign the GICv5 VM tables to
> > the IRS; no GICv5 support\n");
> > +		if (!vgic_v5_irs_clear_vmt())
> > +			vgic_v5_vmt_free();
> > +		return -ENODEV;
> > +	}
> 
> [Severity: Medium]
> If vgic_v5_irs_assign_vmt() returns -EBUSY and
> vgic_v5_irs_clear_vmt() fails,
> this skips calling vgic_v5_vmt_free().
> 
> Since vgic_v5_irs_assign_vmt() returns -EBUSY when the hardware
> VMT_BASER is
> already valid (holding a pointer to the old VMT), doesn't skipping
> the free
> here leak the newly allocated VMT?

It would only do that if we also failed to clear the IRS_VMT_BASER. The
only way we realistically hit this case is if we've kexec'd or similar
and have inherited an IRS with a pre-set IRS_VMT_BASER. We'd fail to
write the new VMT to it, would then clear it, and then free the newly
allocated VMT. The clear would only fail if we timed out here, which
would imply dodgy hardware.

> 
> [ ... ]
> 
> > @@ -153,18 +222,136 @@ int vgic_v5_probe(const struct gic_kvm_info
> > *info)
> 
> [ ... ]
> 
> > +static int vgic_v5_irs_assign_vmt(bool two_level, u8 vm_id_bits,
> > +				  phys_addr_t vmt_base)
> > +{
> > +	u64 vmt_baser;
> > +	u32 vmt_cfgr;
> > +
> > +	guard(raw_spinlock_irqsave)(&global_irs_lock);
> > +
> > +	vmt_baser = irs_readq_relaxed(GICV5_IRS_VMT_BASER);
> > +	if (!!FIELD_GET(GICV5_IRS_VMT_BASER_VALID, vmt_baser))
> > +		return -EBUSY;
> 
> [Severity: High]
> During a kexec, KVM does not clear this register on shutdown, so the
> GICV5_IRS_VMT_BASER_VALID bit may already be set by the previous
> kernel.
> 
> Will vgic_v5_probe() systematically fail and permanently disable
> GICv5 on
> kexec because it aborts instead of clearing the old VMT?

As the code is currently written, yes. That's something that should be
addressed in the future, but right now kexec will result in coming back
without support for VMs.

> 
> > +	vmt_cfgr = FIELD_PREP(GICV5_IRS_VMT_CFGR_VM_ID_BITS,
> > vm_id_bits);
> > +	if (two_level)
> > +		vmt_cfgr |=
> > FIELD_PREP(GICV5_IRS_VMT_CFGR_STRUCTURE,
> > +				      
> > GICV5_IRS_VMT_CFGR_STRUCTURE_TWO_LEVEL);
> > +
> > +	irs_writel_relaxed(vmt_cfgr, GICV5_IRS_VMT_CFGR);
> 
> [Severity: High]
> Should this function call vgic_v5_irs_wait_for_vm_op() to poll for
> the IRS
> to be idle before writing directly to GICV5_IRS_VMT_CFGR?
> 
> Other IRS writer functions enforce an idle check before writing. If
> the IRS
> is busy processing an operation, could this direct write corrupt the
> hardware
> state?
> 

I'd previously dismissed this one, but I think it does make sense to
add in that additional vgic_v5_irs_wait_for_vm_op() here. While it is
very unlikely for the IDLE bit to be non-zero, the GICv5 spec is
explicit that the IRS_VMT_CFGR and IRS_VMT_BASER are read only while
IDLE is 0.

Thanks,
Sascha