Re: [RFC] RISC-V KVM: guest local sfence.vma may miss stale VS-stage TLB after vCPU migration
guoyaxing <[email protected]> Tue, 4 Aug 2026 10:56:01 +0800
| Newsgroups | org.kernel.vger.kvm,org.infradead.lists.kvm-riscv,org.infradead.lists.linux-riscv,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
+cc wangzhizun<[email protected]> 在 2026/8/4 10:32, guoyaxing 写道: > > > 在 2026/8/3 22:22, Anup Patel 写道: >> On Sun, Aug 2, 2026 at 10:30 AM Yaxing Guo <[email protected]> wrote: >>> >>> Hi all, >>> >>> I would like to ask whether RISC-V KVM has a gap around guest sfence.vma >>> and VS-stage TLB invalidation when a vCPU migrates across host CPUs. >>> >>> On bare metal, Linux uses mm_cpumask in __flush_tlb_range() to decide >>> whether to send a local or remote sfence.vma. That works because the >>> cpumask >>> reflects physical CPUs. In a guest, however, the kernel only sees >>> vCPUs, so >>> from the guest point of view a task may never "migrate" even though the >>> underlying vCPU moved to a different host CPU. In other words, guest >>> mm_cpumask is effectively a vCPU mask, not a host CPU mask. >>> >>> There is also a related switch_mm()/set_mm_asid() aspect. With the ASID >>> allocator enabled, RISC-V Linux does not flush on every context >>> switch. A >>> local_flush_tlb_all() is only done when a deferred ASID rollover >>> flush is >>> pending for the current CPU. That is fine on bare metal, because >>> previous >>> shootdowns are also based on physical CPUs. In a guest, however, the >>> earlier >>> shootdown may have been only a guest-local sfence.vma and may have >>> missed an >>> old host CPU where the same vCPU ran previously. >>> >>> A simplified sequence is: >>> >>> 1. A guest task with mm A runs on vcpu0 while vcpu0 is scheduled on host >>> CPU1. CPU1 fills a VS-stage TLB entry for that mm/ASID. >>> 2. vcpu0 migrates to host CPU0. KVM's vCPU migration sanitization is >>> local >>> to the current host CPU, so it does not invalidate the VS-stage >>> TLB entry >>> that was left on host CPU1. >> >> When vcpu0 migrates to host CPU0, the host CPU0 may or may not have >> VS-stage TLB entry so VCPU migration sanitization does the right thing. >> >> In the future, vcpu0 may again move back to host CPU1 containing stale >> VS-stage TLB entries left in step1 above so the VCPU migration >> sanitization >> will cleanup these stale TLB enteries. >> > > Sorry, perhaps I didn't describe it clearly enough. In short, the issue > is: because vcpu0 migrated away, cpu1 still has stale TLB entries while > cpu0 has the new mapping. After that, neither vcpu migrated again, but > the process was scheduled from vcpu0 to vcpu1 (which is on cpu1), > causing it to hit the stale TLB. Consider the following scenario: > > 1. vcpu0 and vcpu1 are both running on cpu1. At this point, vcpu0 is > on cpu1, and a process on it reads a page, which gets cached in the TLB > (let's call this the "stale TLB entry"). > 2. vcpu0 migrates to cpu0, and then the process (on vcpu0, now on > cpu0) does a CoW (copy-on-write) and establishes a new mapping. > 3. Neither vcpu0 nor vcpu1 migrates again. At this point, the process > gets scheduled inside the VM onto vcpu1 (which is on cpu1), reads the > same page, and hits the stale TLB entry. > >>> 3. The guest later updates mm A's page table from vcpu0, for example >>> due to >>> COW. Since the guest still sees mm A as having run only on guest >>> CPU0, >>> __flush_tlb_range() can choose a local sfence.vma. That local >>> sfence.vma >>> only affects host CPU0, where vcpu0 is currently running, so the >>> stale >>> VS-stage entry on host CPU1 remains. >> >> No need to worry about stale VS-stage entry on host CPU1 because these >> will age-out or same guest may again move to host CPU1 resulting in >> VCPU migration sanitization. >> >>> 4. Later, the task is scheduled again while vcpu0 is running on host >>> CPU1. > > Sorry, that was a typo on my part. 'Later, the task is scheduled again > while vcpu0 is running on host CPU1' should say vcpu1 instead of vcpu0. > >>> If the ASID is still valid and no deferred ASID rollover flush is >>> pending, >>> set_mm_asid() may not issue any flush. The access on host CPU1 >>> can then >>> hit the old VS-stage TLB entry. >> >> See above comments. >> >>> >>> We hit this on a XiangShan RISC-V system. A bash task in the guest was >>> scheduled on vcpu0 while that vCPU was running on host CPU1. A read-only >>> page was prefetched and cached in CPU1's VS-stage TLB. Later the same >>> vCPU >>> migrated to host CPU0, the guest triggered COW and updated the >>> mapping, and >>> the guest still issued only a local sfence.vma. The stale VS-stage >>> entry on >>> CPU1 was not invalidated. After running for some time, the task later >>> accessed that mapping while running on host CPU1 again and hit the stale >>> translation. The data in that page was a pointer from the GOT, so it >>> dereferenced to NULL and eventually caused a segmentation fault. >> >> The above example is already taken care by TLB flushes in KVM RISC-V. >> This smells like some HW errata to me. >> >>> >>> I checked current upstream RISC-V KVM (v7.0.2) and I do not see >>> HSTATUS.VTVM >>> being enabled, so guest sfence.vma is not trapped. I found >>> kvm_riscv_local_tlb_sanitize() on host CPU migration, which handles >>> G-stage >>> VMID entries on the current host CPU, but I do not think it generally >>> solves >>> the stale VS-stage TLB case described above. >>> >>> I also noticed the vendor-specific kvm_riscv_vsstage_tlb_no_gpa path in >>> current upstream. It looks like a workaround for a particular >>> implementation >>> issue on Andes AX66, where the VS-stage TLB does not cache guest >>> physical >>> address and VMID, so the normal VMID/GPA-based invalidation model is >>> insufficient. In that sense, it does not seem to be a generic >>> solution for >>> guest sfence.vma coherence across vCPU migration. >>> >>> I am wondering what the right direction should be. Enabling >>> HSTATUS.VTVM and >>> trapping guest sfence.vma would allow KVM to translate guest local >>> sfence.vma >>> into the required host-side hfence.vvma, but that may be too heavy >>> for the >>> common case. Another possible direction may be to make the current >>> Andes-specific vCPU migration flush logic more generic, so >>> implementations >>> that need VS-stage TLB sanitization on vCPU migration can opt in >>> through a >>> common mechanism. A third possibility may be to define a lightweight SBI >>> interface for guest local sfence.vma, so Linux could use it in >>> __flush_tlb_range() for the local case when running under KVM, and >>> KVM could >>> then decide whether a local sfence.vma is sufficient or whether host- >>> side >>> hfence.vvma is also needed. >> >> Enabling HSTATUS.VTVM has a big impact on performance so certainly >> NACK from myside. >> >>> >>> My questions are: >>> 1. Is the behavior above expected on current RISC-V KVM? >> >> Yes, it should work fine unless there is some HW bug. >> >>> 2. If not, what would be the preferred fix direction? >>> 3. Should this be handled by trapping guest sfence.vma with VTVM, by >>> generalizing the existing vCPU-migration VS-stage TLB flush >>> logic, or by >>> adding a lighter SBI-mediated path for guest local sfence.vma? >>> 4. Or is there another intended mechanism to keep VS-stage TLB state >>> coherent >>> across vCPU migration? >>> >>> If needed, I can send a reproducer and the exact hardware/kernel >>> details. >>> >> >> Even if you share some reporducing code sequence, we still have don't >> have access to your HW so it won't help much. >> >> Regards, >> Anup