Re: [PATCH v2] KVM: PPC: Book3S HV: Don't drop pending doorbell across L2 entry

Gautam Menghani <[email protected]>
Newsgroups org.kernel.vger.kvm,org.kernel.vger.kvm-ppc,org.ozlabs.lists.linuxppc-dev
Message-ID <[email protected]>
> Hi Vaibhav,
> 
> I have tested this v2 patch and it is still giving me the issue I reported,
> so here is my analysis:
> 
> a) Without applying the patch :
> 
> 1) Start the guest and run stress-ng as below for sometime
> localhost:~ # stress-ng --cpu 4 --vm 2 --vm-bytes 1G --hdd 2 --hdd-bytes 1G
> --sched other --timeout 3600000s
> stress-ng: info:  [1464] setting to a 41 days, 16 hours, 0 secs run per
> stressor
> stress-ng: info:  [1464] dispatching hogs: 4 cpu, 2 vm, 2 hdd
> 
> 
> 
> 2) Start the migration from H1 to H2:
> 
> ltc-lp7:~ # virsh migrate --live --domain sles16_anu
> qemu+ssh://10.xx.xx.xx/system --verbose --undefinesource --persistent
> --auto-converge --postcopy
> ([email protected]) Password:
> Migration: [100.00 %]
> 
> 3) Migration got completed but guest is not getting recovered from
> continuous softlockups
> 
> [ 1336.003836][    C1] watchdog: BUG: soft lockup - CPU#1 stuck for 977s!
> [htxd_monitor:1337]
> [ 1336.006834][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1002s!
> [rcu_exp_par_gp_:19]
> [ 1346.015839][    C0] BUG: workqueue lockup - pool cpus=1 node=0 flags=0x0
> nice=0 stuck for 1090s!
> [ 1346.016355][    C0] BUG: workqueue lockup - pool cpus=3 node=0 flags=0x0
> nice=0 stuck for 1107s!
> [ 1346.016874][    C0] BUG: workqueue lockup - pool cpus=7 node=0 flags=0x0
> nice=0 stuck for 1093s!
> [ 1356.007835][    C6] watchdog: BUG: soft lockup - CPU#6 stuck for 912s!
> [systemd:1353]
> [ 1356.008835][    C7] watchdog: BUG: soft lockup - CPU#7 stuck for 998s!
> [systemd-journal:570]
> [ 1360.003836][    C1] watchdog: BUG: soft lockup - CPU#1 stuck for 999s!
> [htxd_monitor:1337]
> [ 1360.006834][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1024s!
> [rcu_exp_par_gp_:19]
> [ 1368.933835][    C4] rcu: INFO: rcu_preempt self-detected stall on CPU
> [ 1368.933973][    C4] rcu:     4-....: (1129830 ticks this GP)
> idle=afc4/1/0x4000000000000002 softirq=3694/428556 fqs=259639
> [ 1368.934106][    C4] rcu:              hardirqs   softirqs  csw/system
> [ 1368.934188][    C4] rcu:      number:        1     444039 0
> [ 1368.934271][    C4] rcu:     cputime:        3          8 1096165   ==>
> 1110021(ms)
> [ 1368.934373][    C4] rcu:     (t=1140022 jiffies g=6177 q=1684 ncpus=8)
> [ 1376.224839][    C0] BUG: workqueue lockup - pool cpus=1 node=0 flags=0x0
> nice=0 stuck for 1120s!
> [ 1376.225307][    C0] BUG: workqueue lockup - pool cpus=3 node=0 flags=0x0
> nice=0 stuck for 1138s!
> [ 1376.225428][    C0] BUG: workqueue lockup - pool cpus=5 node=0 flags=0x0
> nice=0 stuck for 715s!
> [ 1376.225548][    C0] BUG: workqueue lockup - pool cpus=6 node=0 flags=0x0
> nice=0 stuck for 1027s!
> [ 1376.225667][    C0] BUG: workqueue lockup - pool cpus=7 node=0 flags=0x0
> nice=0 stuck for 1123s!
> [ 1444.006835][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1100s!
> [rcu_exp_par_gp_:19]
> 
> 
> b) Even after applying the patch also it is giving same softlockup issue as
> mentioned above:
> Though I have enough vcpus and memory on the guest (16 vcpus , 13Gi of
> memory)  and ample amount of memory and cpus
> present on host still these softlockups are happening after applying the
> patch too. I tried reducing stress also on the guest
> but still this issue is seen.
> 
> stress-ng --cpu 4 --vm 2 --vm-bytes 1G --hdd 2 --hdd-bytes 1G --sched other
> --timeout 3600000s
> 
> If you are planning to send next version of this patch,
> Please do add my reported-by:
> Reported-by: Anushree Mathur <[email protected]>
> 

Thanks for testing this Anushree. Upon further debugging, I found that
this issue is likely due to a bug in QEMU, for which I've sent a fix[1].

@maddy: Please don't pull in this patch for now, will ping here if this
is needed.

[1] : lore.kernel.org/qemu-devel/[email protected]/

Thanks,
Gautam
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.