Re: [RFC] RISC-V KVM: guest local sfence.vma may miss stale VS-stage TLB after vCPU migration

Anup Patel <[email protected]> Mon, 3 Aug 2026 19:52:14 +0530
Newsgroups org.kernel.vger.kvm,org.infradead.lists.kvm-riscv,org.infradead.lists.linux-riscv,org.kernel.vger.linux-kernel
Message-ID <CAAhSdy3QE6Fe_fanKOLHLjDCb+vsVax=wy18UxzR23XN7ggJGw@mail.gmail.com>
On Sun, Aug 2, 2026 at 10:30=E2=80=AFAM Yaxing Guo <[email protected]> w=
rote:
>
> Hi all,
>
> I would like to ask whether RISC-V KVM has a gap around guest sfence.vma
> and VS-stage TLB invalidation when a vCPU migrates across host CPUs.
>
> On bare metal, Linux uses mm_cpumask in __flush_tlb_range() to decide
> whether to send a local or remote sfence.vma. That works because the cpum=
ask
> reflects physical CPUs. In a guest, however, the kernel only sees vCPUs, =
so
> from the guest point of view a task may never "migrate" even though the
> underlying vCPU moved to a different host CPU. In other words, guest
> mm_cpumask is effectively a vCPU mask, not a host CPU mask.
>
> There is also a related switch_mm()/set_mm_asid() aspect. With the ASID
> allocator enabled, RISC-V Linux does not flush on every context switch. A
> local_flush_tlb_all() is only done when a deferred ASID rollover flush is
> pending for the current CPU. That is fine on bare metal, because previous
> shootdowns are also based on physical CPUs. In a guest, however, the earl=
ier
> shootdown may have been only a guest-local sfence.vma and may have missed=
 an
> old host CPU where the same vCPU ran previously.
>
> A simplified sequence is:
>
> 1. A guest task with mm A runs on vcpu0 while vcpu0 is scheduled on host
>    CPU1. CPU1 fills a VS-stage TLB entry for that mm/ASID.
> 2. vcpu0 migrates to host CPU0. KVM's vCPU migration sanitization is loca=
l
>    to the current host CPU, so it does not invalidate the VS-stage TLB en=
try
>    that was left on host CPU1.

When vcpu0 migrates to host CPU0, the host CPU0 may or may not have
VS-stage TLB entry so VCPU migration sanitization does the right thing.

In the future, vcpu0 may again move back to host CPU1 containing stale
VS-stage TLB entries left in step1 above so the VCPU migration sanitization
will cleanup these stale TLB enteries.

> 3. The guest later updates mm A's page table from vcpu0, for example due =
to
>    COW. Since the guest still sees mm A as having run only on guest CPU0,
>    __flush_tlb_range() can choose a local sfence.vma. That local sfence.v=
ma
>    only affects host CPU0, where vcpu0 is currently running, so the stale
>    VS-stage entry on host CPU1 remains.

No need to worry about stale VS-stage entry on host CPU1 because these
will age-out or same guest may again move to host CPU1 resulting in
VCPU migration sanitization.

> 4. Later, the task is scheduled again while vcpu0 is running on host CPU1=
.
>    If the ASID is still valid and no deferred ASID rollover flush is pend=
ing,
>    set_mm_asid() may not issue any flush. The access on host CPU1 can the=
n
>    hit the old VS-stage TLB entry.

See above comments.

>
> We hit this on a XiangShan RISC-V system. A bash task in the guest was
> scheduled on vcpu0 while that vCPU was running on host CPU1. A read-only
> page was prefetched and cached in CPU1's VS-stage TLB. Later the same vCP=
U
> migrated to host CPU0, the guest triggered COW and updated the mapping, a=
nd
> the guest still issued only a local sfence.vma. The stale VS-stage entry =
on
> CPU1 was not invalidated. After running for some time, the task later
> accessed that mapping while running on host CPU1 again and hit the stale
> translation. The data in that page was a pointer from the GOT, so it
> dereferenced to NULL and eventually caused a segmentation fault.

The above example is already taken care by TLB flushes in KVM RISC-V.
This smells like some HW errata to me.

>
> I checked current upstream RISC-V KVM (v7.0.2) and I do not see HSTATUS.V=
TVM
> being enabled, so guest sfence.vma is not trapped. I found
> kvm_riscv_local_tlb_sanitize() on host CPU migration, which handles G-sta=
ge
> VMID entries on the current host CPU, but I do not think it generally sol=
ves
> the stale VS-stage TLB case described above.
>
> I also noticed the vendor-specific kvm_riscv_vsstage_tlb_no_gpa path in
> current upstream. It looks like a workaround for a particular implementat=
ion
> issue on Andes AX66, where the VS-stage TLB does not cache guest physical
> address and VMID, so the normal VMID/GPA-based invalidation model is
> insufficient. In that sense, it does not seem to be a generic solution fo=
r
> guest sfence.vma coherence across vCPU migration.
>
> I am wondering what the right direction should be. Enabling HSTATUS.VTVM =
and
> trapping guest sfence.vma would allow KVM to translate guest local sfence=
.vma
> into the required host-side hfence.vvma, but that may be too heavy for th=
e
> common case. Another possible direction may be to make the current
> Andes-specific vCPU migration flush logic more generic, so implementation=
s
> that need VS-stage TLB sanitization on vCPU migration can opt in through =
a
> common mechanism. A third possibility may be to define a lightweight SBI
> interface for guest local sfence.vma, so Linux could use it in
> __flush_tlb_range() for the local case when running under KVM, and KVM co=
uld
> then decide whether a local sfence.vma is sufficient or whether host-side
> hfence.vvma is also needed.

Enabling HSTATUS.VTVM has a big impact on performance so certainly
NACK from myside.

>
> My questions are:
> 1. Is the behavior above expected on current RISC-V KVM?

Yes, it should work fine unless there is some HW bug.

> 2. If not, what would be the preferred fix direction?
> 3. Should this be handled by trapping guest sfence.vma with VTVM, by
>    generalizing the existing vCPU-migration VS-stage TLB flush logic, or =
by
>    adding a lighter SBI-mediated path for guest local sfence.vma?
> 4. Or is there another intended mechanism to keep VS-stage TLB state cohe=
rent
>    across vCPU migration?
>
> If needed, I can send a reproducer and the exact hardware/kernel details.
>

Even if you share some reporducing code sequence, we still have don't
have access to your HW so it won't help much.

Regards,
Anup