Re: [PATCH v2] KVM: PPC: Book3S HV: Don't drop pending doorbell across L2 entry
Gautam Menghani <[email protected]>
| Newsgroups | org.kernel.vger.kvm,org.kernel.vger.kvm-ppc,org.ozlabs.lists.linuxppc-dev |
|---|---|
| Message-ID | <[email protected]> |
> Hi Vaibhav, > > I have tested this v2 patch and it is still giving me the issue I reported, > so here is my analysis: > > a) Without applying the patch : > > 1) Start the guest and run stress-ng as below for sometime > localhost:~ # stress-ng --cpu 4 --vm 2 --vm-bytes 1G --hdd 2 --hdd-bytes 1G > --sched other --timeout 3600000s > stress-ng: info: [1464] setting to a 41 days, 16 hours, 0 secs run per > stressor > stress-ng: info: [1464] dispatching hogs: 4 cpu, 2 vm, 2 hdd > > > > 2) Start the migration from H1 to H2: > > ltc-lp7:~ # virsh migrate --live --domain sles16_anu > qemu+ssh://10.xx.xx.xx/system --verbose --undefinesource --persistent > --auto-converge --postcopy > ([email protected]) Password: > Migration: [100.00 %] > > 3) Migration got completed but guest is not getting recovered from > continuous softlockups > > [ 1336.003836][  C1] watchdog: BUG: soft lockup - CPU#1 stuck for 977s! > [htxd_monitor:1337] > [ 1336.006834][  C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1002s! > [rcu_exp_par_gp_:19] > [ 1346.015839][  C0] BUG: workqueue lockup - pool cpus=1 node=0 flags=0x0 > nice=0 stuck for 1090s! > [ 1346.016355][  C0] BUG: workqueue lockup - pool cpus=3 node=0 flags=0x0 > nice=0 stuck for 1107s! > [ 1346.016874][  C0] BUG: workqueue lockup - pool cpus=7 node=0 flags=0x0 > nice=0 stuck for 1093s! > [ 1356.007835][  C6] watchdog: BUG: soft lockup - CPU#6 stuck for 912s! > [systemd:1353] > [ 1356.008835][  C7] watchdog: BUG: soft lockup - CPU#7 stuck for 998s! > [systemd-journal:570] > [ 1360.003836][  C1] watchdog: BUG: soft lockup - CPU#1 stuck for 999s! > [htxd_monitor:1337] > [ 1360.006834][  C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1024s! > [rcu_exp_par_gp_:19] > [ 1368.933835][  C4] rcu: INFO: rcu_preempt self-detected stall on CPU > [ 1368.933973][  C4] rcu:   4-....: (1129830 ticks this GP) > idle=afc4/1/0x4000000000000002 softirq=3694/428556 fqs=259639 > [ 1368.934106][  C4] rcu:       hardirqs  softirqs  csw/system > [ 1368.934188][  C4] rcu:   number:    1   444039 0 > [ 1368.934271][  C4] rcu:   cputime:    3     8 1096165  ==> > 1110021(ms) > [ 1368.934373][  C4] rcu:   (t=1140022 jiffies g=6177 q=1684 ncpus=8) > [ 1376.224839][  C0] BUG: workqueue lockup - pool cpus=1 node=0 flags=0x0 > nice=0 stuck for 1120s! > [ 1376.225307][  C0] BUG: workqueue lockup - pool cpus=3 node=0 flags=0x0 > nice=0 stuck for 1138s! > [ 1376.225428][  C0] BUG: workqueue lockup - pool cpus=5 node=0 flags=0x0 > nice=0 stuck for 715s! > [ 1376.225548][  C0] BUG: workqueue lockup - pool cpus=6 node=0 flags=0x0 > nice=0 stuck for 1027s! > [ 1376.225667][  C0] BUG: workqueue lockup - pool cpus=7 node=0 flags=0x0 > nice=0 stuck for 1123s! > [ 1444.006835][  C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1100s! > [rcu_exp_par_gp_:19] > > > b) Even after applying the patch also it is giving same softlockup issue as > mentioned above: > Though I have enough vcpus and memory on the guest (16 vcpus , 13Gi of > memory) and ample amount of memory and cpus > present on host still these softlockups are happening after applying the > patch too. I tried reducing stress also on the guest > but still this issue is seen. > > stress-ng --cpu 4 --vm 2 --vm-bytes 1G --hdd 2 --hdd-bytes 1G --sched other > --timeout 3600000s > > If you are planning to send next version of this patch, > Please do add my reported-by: > Reported-by: Anushree Mathur <[email protected]> > Thanks for testing this Anushree. Upon further debugging, I found that this issue is likely due to a bug in QEMU, for which I've sent a fix[1]. @maddy: Please don't pull in this patch for now, will ping here if this is needed. [1] : lore.kernel.org/qemu-devel/[email protected]/ Thanks, Gautam