Re: [PATCH v7 07/36] KVM: x86: Restructure get_kvmclock()

Sean Christopherson <[email protected]>
Newsgroups dev.linux.lists.sashiko-reviews,org.kernel.vger.kvm
Message-ID <[email protected]>
On Tue, Jul 28, 2026, [email protected] wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] Removal of `get_cpu()` reintroduces a race condition between `rdtsc` and `cpu_tsc_khz`, and triggers a kernel warning for preemptible `__this_cpu_read`.
> --
> 
> --- Patch [7]: [PATCH v7 07/36] KVM: x86: Restructure get_kvmclock() ---
> commit c5c131ac2510f7641714a694b1cf0192a9da9749
> Author: David Woodhouse <[email protected]>
> 
> KVM: x86: Restructure get_kvmclock()
> 
> This commit wraps the use_master_clock block in an ifdef for CONFIG_X86_64
> to resolve 32-bit build issues with unused variables and host_tsc_clocksource.
> It also refactors the master-clock success path to use continue, allowing
> the fallback computation to become the common tail.
> 
> > diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
> > index ee115dc6df6e1..d68d00cdbea73 100644
> > --- a/arch/x86/kvm/x86.c
> > +++ b/arch/x86/kvm/x86.c
> [ ... ]
> > @@ -1636,34 +1636,38 @@ static unsigned long get_cpu_tsc_khz(void)
> >  static void get_kvmclock(struct kvm *kvm, struct kvm_clock_data *data)
> >  {
> >  	struct kvm_arch *ka = &kvm->arch;
> > -	struct pvclock_vcpu_time_info hv_clock;
> >  	unsigned int seq;
> >  
> >  	do {
> >  		seq = read_seqcount_begin(&ka->pvclock_sc);
> >  
> >  		data->flags = 0;
> > -		if (ka->use_master_clock) {
> >  #ifdef CONFIG_X86_64
> > +		if (ka->use_master_clock) {
> > +			struct pvclock_vcpu_time_info hv_clock;
> >  			struct timespec64 ts;
> >  
> >  			if (kvm_get_walltime_and_clockread(&ts, &data->host_tsc)) {
> >  				data->realtime = ts.tv_nsec + NSEC_PER_SEC * ts.tv_sec;
> > -				data->flags |= KVM_CLOCK_REALTIME | KVM_CLOCK_HOST_TSC;
> > -			} else
> > -#endif
> > -			data->host_tsc = rdtsc();
> > -
> > -			data->flags |= KVM_CLOCK_TSC_STABLE;
> > -			hv_clock.tsc_timestamp = ka->master_cycle_now;
> > -			hv_clock.system_time = ka->master_kernel_ns + ka->kvmclock_offset;
> > -			kvm_get_time_scale(NSEC_PER_SEC, get_cpu_tsc_khz() * 1000LL,
> > -					   &hv_clock.tsc_shift,
> > -					   &hv_clock.tsc_to_system_mul);
> > +				data->flags |= KVM_CLOCK_REALTIME | KVM_CLOCK_HOST_TSC | KVM_CLOCK_TSC_STABLE;
> > +
> > +				hv_clock.tsc_timestamp = ka->master_cycle_now;
> > +				hv_clock.system_time = ka->master_kernel_ns + ka->kvmclock_offset;
> > +				kvm_get_time_scale(NSEC_PER_SEC, get_cpu_tsc_khz() * 1000LL,
> > +						   &hv_clock.tsc_shift,
> > +						   &hv_clock.tsc_to_system_mul);
> 
> [Severity: High]
> Does this code trigger a preemptible __this_cpu_read() warning?

Yes.  And my suggestion to isolate the change was "bad", but also "good".  "Bad"
in the sense that we can't simply drop the pinning, "good" in that it forced me
to dig into why the code is the way it is.

Happily, I ended up with this, which segues very well into "KVM: x86: Fix KVM clock
precision in get_kvmclock() with TSC scaling".

---
From: Sean Christopherson <[email protected]>
Date: Tue, 28 Jul 2026 14:57:58 -0700
Subject: [PATCH] KVM: x86: Drop unnecessary CPU pinning when computing/getting
 kvmclock

When computing the current kvmclock value, don't pin the task to the
current CPU for the entire duration of the master clock path, as the
CPU pinning was never about ensuring rdtsc() and cpu_tsc_khz would
agree.  As pointed out by David, ka->use_master_clock can only be true
when the host clocksource is TSC based, which in turn requires a stable,
constant and synchronised TSC across all CPUs.

The CPU pinning was added in commit e2c2206a1899 ("KVM: x86: Fix potential
preemption when get the current kvmclock timestamp") purely in response to
a CONFIG_DEBUG_PREEMPT=y bug due.  Despite what the comment would suggest,
including rdtsc() in the {get,put}_cpu() section was opportunistic.  In
fact, Paolo even said exactly that when suggesting that KVM guarantee the
rdtsc() would execute on the same CPU[*].

 : Also, rdtsc() should really be on the same CPU as __this_cpu_read.  We
 : know it's not really really necessary because the master clock is
 : active, but since we need a get_cpu/put_cpu pair, better be clean.

Nothing has changed in the last ~9 years, i.e. the rdtsc() still *should*
be on the same CPU, but super strictly speaking, all will be fine if the
task is migrated between grabbing the frequency and doing rdtsc().
Dropping the CPU pinning will allow dropping the rdtsc() entirely without
having to resort to a large "rewrite get_kvmclock()" patch.

Opportunistically add a comment to explain why KVM needs to snapshot the
frequency, because that _is_ a hard requirement to avoid reintroducing the
bug fixed by commit e70b57a6ce4e ("KVM: X86: Fix softlockup when get the
current kvmclock")

Link: https://lore.kernel.org/all/[email protected] [*]

Suggested-by: David Woodhouse <[email protected]>
Signed-off-by: Sean Christopherson <[email protected]>
---
 arch/x86/kvm/x86.c | 15 +++++++++------
 1 file changed, 9 insertions(+), 6 deletions(-)

diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index d5a05acf20ad..f64726e08f00 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -1638,13 +1638,18 @@ static void __get_kvmclock(struct kvm *kvm, struct kvm_clock_data *data)
 {
 	struct kvm_arch *ka = &kvm->arch;
 	struct pvclock_vcpu_time_info hv_clock;
+	u64 tsc_hz;
 
-	/* both __this_cpu_read() and rdtsc() should be on the same cpu */
+	/*
+	 * Snapshot and validate the TSC frequency as kvmclock_cpu_down_prep()
+	 * zeros the per-CPU value when a CPU is going offline.
+	 */
 	get_cpu();
+	tsc_hz = (u64)get_cpu_tsc_khz() * HZ_PER_KHZ;
+	put_cpu();
 
 	data->flags = 0;
-	if (ka->use_master_clock &&
-	    (static_cpu_has(X86_FEATURE_CONSTANT_TSC) || __this_cpu_read(cpu_tsc_khz))) {
+	if (ka->use_master_clock && tsc_hz) {
 #ifdef CONFIG_X86_64
 		struct timespec64 ts;
 
@@ -1658,15 +1663,13 @@ static void __get_kvmclock(struct kvm *kvm, struct kvm_clock_data *data)
 		data->flags |= KVM_CLOCK_TSC_STABLE;
 		hv_clock.tsc_timestamp = ka->master_cycle_now;
 		hv_clock.system_time = ka->master_kernel_ns + ka->kvmclock_offset;
-		kvm_get_time_scale(NSEC_PER_SEC, get_cpu_tsc_khz() * 1000LL,
+		kvm_get_time_scale(NSEC_PER_SEC,  tsc_hz,
 				   &hv_clock.tsc_shift,
 				   &hv_clock.tsc_to_system_mul);
 		data->clock = __pvclock_read_cycles(&hv_clock, data->host_tsc);
 	} else {
 		data->clock = get_kvmclock_base_ns() + ka->kvmclock_offset;
 	}
-
-	put_cpu();
 }
 
 static void get_kvmclock(struct kvm *kvm, struct kvm_clock_data *data)

base-commit: c9ed17959337edfc904f59e7ed1f569918f260ca
--
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.