Re: [PATCH 3/9] KVM: selftests: Affine child tasks to other pCPUs in hardware disable test

Sean Christopherson <[email protected]> Fri, 31 Jul 2026 09:03:12 -0700
Newsgroups dev.linux.lists.sashiko-reviews,dev.linux.lists.kvmarm,org.kernel.vger.kvm
Message-ID <[email protected]>
On Fri, Jul 31, 2026, Sean Christopherson wrote:
> On Thu, Jul 30, 2026, [email protected] wrote:
> > > diff --git a/tools/testing/selftests/kvm/hardware_disable_test.c b/tools/testing/selftests/kvm/hardware_disable_test.c
> > > index 3147f5c97e943..1bcbd13e99947 100644
> > > --- a/tools/testing/selftests/kvm/hardware_disable_test.c
> > > +++ b/tools/testing/selftests/kvm/hardware_disable_test.c
> > [ ... ]
> > > @@ -84,16 +85,11 @@ static void run_test(u32 run)
> > [ ... ]
> > >  		check_create_thread(&threads[i], NULL, run_vcpu, vcpu);
> > > -		check_set_affinity(threads[i], &cpu_set);
> > > +		check_set_affinity(threads[i], &child_cpu_set);
> > >  
> > >  		for (j = 0; j < SLEEPING_THREAD_NUM; ++j) {
> > >  			check_create_thread(&throw_away, NULL, sleeping_thread,
> > >  					    (void *)NULL);
> > > -			check_set_affinity(throw_away, &cpu_set);
> > > +			check_set_affinity(throw_away, &child_cpu_set);
> > >  		}
> > >  	}
> > [ ... ]
> > > @@ -147,12 +143,30 @@ void wait_for_child_setup(pid_t pid)
> > [ ... ]
> > > +static void setup_child_cpu_set(void)
> > > +{
> > > +	int cpu;
> > > +
> > > +	kvm_sched_getaffinity(0, sizeof(child_cpu_set), &child_cpu_set);
> > > +
> > > +	if (CPU_COUNT(&child_cpu_set) < 2)
> > > +		return;
> > > +
> > > +	cpu = pin_task_to_random_cpu(pthread_self(), &child_cpu_set);
> > > +	CPU_CLR(cpu, &child_cpu_set);
> > > +
> > > +	while (CPU_COUNT(&child_cpu_set) > VCPU_NUM)
> > > +		CPU_CLR(kvm_pick_random_cpu(&child_cpu_set), &child_cpu_set);
> > > +}
> > 
> > [Severity: Medium]
> > Since setup_child_cpu_set() pins the parent thread to a single random CPU
> > prior to spawning children, won't the newly created threads in run_test()
> > inherit this single-CPU affinity upon creation?
> > 
> > Because check_create_thread() spawns the threads before their affinity is
> > explicitly updated to child_cpu_set, they might immediately begin executing
> > on the exact same CPU as the parent.
> 
> Yeah, I was trying to avoid this since pthread_attr_setaffinity_np() is only
> available with __USE_GNU.
> 
> > Could this result in the child threads starving the parent on that single CPU
> > before the parent gets a chance to execute check_set_affinity()? This seems
> > like it might heavily delay the migration and trigger the same CPU contention
> > that this patch aims to fix.
> 
> After far too much testing, the problem isn't direct CPU contention, AFAICT it's
> due to incurring cross-node NUMA penalties, and presumably scheduler contention
> when the pool of possible CPUs to run on is huge.
> 
> I tried a bunch of fancy approaches, e.g. making parent, child, and grandchildren
> all mutually exclusive, and while the fanciness can provide marginal benefits when
> pthread_attr_setaffinity_np() isn't available, all that really matters is that the
> worker threads get affined before they start doing work.

Oh, and I also ruled out wakeup latency, i.e. this isn't the same underlying cause
that prompted 0297cdc12a87 ("KVM: selftests: Add option to rseq test to override
/dev/cpu_dma_latency").