Re: [PATCH] srcu: Queue sdp->work when the delay timer is successfully deleted
"Zqiang" <[email protected]> Mon, 03 Aug 2026 14:37:22 +0000
| Newsgroups | org.kernel.vger.rcu,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
>=20 >=20>=20 >=20> On Sat, Aug 01, 2026 at 05:19:25AM +0000, Zqiang wrote: > >=20=20 >=20>=20=20 >=20> Hello Zqiang, > >=20=20 >=20> On Thu, Jul 09, 2026 at 06:06:02PM +0800, Zqiang wrote: > >=20=20 >=20> >=20 >=20> > In the cleanup_srcu_struct(), when iterating over per-cpu's srcu= _data, > > > the timer_delete_sync(&sdp->delay_work) is called to cancel the de= lay > > > timer before flush_work(&sdp->work). > > >=20 >=20> > However, if the timer_delete_sync() returns 1 means that it succ= essfully > > > deleted an pending timer before it had a chance to fire, also mean= s that > > > the sdp->work cannot be queued, the subsequent flush_work(&sdp->wo= rk) > > > will returns immediately without waiting for anything, this causes= SRCU > > > callbacks to not be processed. > > >=20 >=20> > Fix this by checking the return value of timer_delete_sync(), if= it > > > returns 1, explicitly queue sdp->work so that the following flush_= work() > > > can correctly wait for the work to complete. > > >=20 >=20> > Signed-off-by: Zqiang <[email protected]> > > > --- > > > kernel/rcu/srcutree.c | 6 +++++- > > > 1 file changed, 5 insertions(+), 1 deletion(-) > > >=20 >=20> > diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c > > > index 7c2f7cc131f7..02c322b7c6f1 100644 > > > --- a/kernel/rcu/srcutree.c > > > +++ b/kernel/rcu/srcutree.c > > > @@ -725,7 +725,11 @@ void cleanup_srcu_struct(struct srcu_struct *= ssp) > > > for_each_possible_cpu(cpu) { > > > struct srcu_data *sdp =3D per_cpu_ptr(ssp->sda, cpu); > > >=20 >=20> > - timer_delete_sync(&sdp->delay_work); > > > + //In most scenarios, calling srcu_barrier before cleanup > > > + //will not trigger WARN_ON(). > > > + if (WARN_ON(timer_delete_sync(&sdp->delay_work)) && > > > + rcu_cpu_beenfullyonline(sdp->cpu)) > > > + queue_work_on(sdp->cpu, rcu_gp_wq, &sdp->work); > > >=20 >=20> I started seeing this on my tests, it is not trivial to decode thi= s one, > > but, I can try harder if _really_ needed. > >=20=20 >=20>=20=20 >=20> The scenario I can think of is that we missed the call to srcu_bar= rier() > > before cleanup_srcu_struct(): > >=20=20 >=20> loop_add() > > ->blk_mq_alloc_tag_set > > init_srcu_struct(&set->tags_srcu) > >=20=20 >=20> blk_mq_alloc_set_map_and_rqs() { > > ->__blk_mq_alloc_rq_maps() > > ->__blk_mq_alloc_map_and_rqs() return error > > goto out_unwind: __blk_mq_free_map_and_rqs() > > ->blk_mq_free_rq_map() > > ->blk_mq_free_tags() > > ->call_srcu(&set->tags_srcu, &tags->rcu_head, blk_mq_free_tags_callb= ack); > > } return error=20 >=20>=20=20 >=20> goto out_free_mq_map: > > .... > > cleanup_srcu_struct(&set->tags_srcu) > > -> trigger WARN_ON(timer_delete_sync(&sdp->delay_work) > >=20=20 >=20> Can you try the following patch? > >=20=20 >=20> I tried it, and it does not help -- the WARN still fires at the sa= me rate. I think the analysis points at the wrong call site. > >=20=20 >=20> Setup: linux-next-20260731 (arm64), 32 vCPU VM, HZ=3D1000, PROVE_L= OCKING > > and DEBUG_OBJECTS_TIMERS enabled, reproducer stress-ng --loop 32 > > --timeout 60s. > >=20 >=20Thanks for provide testing methods, I will also testing it. >=20 >=20>=20 >=20> baseline 3 x WARN srcutree.c:706 in 60s > > + your blk-mq patch 4 x WARN srcutree.c:706 in 60s > >=20=20 >=20> (3 vs 4 is just jitter on a one-jiffy race, not a regression.) > >=20=20 >=20> The reason it cannot help is that the splat comes from > > blk_mq_free_tag_set(), not from the blk_mq_alloc_tag_set() error pat= h > > your patch touches. All four splats in the patched run have the same > > trace: > >=20=20 >=20> cleanup_srcu_struct+0x274/0x450 (P) > > blk_mq_free_tag_set+0x1a4/0x1e0 > > loop_remove+0x2c/0x78 > > loop_control_ioctl+0x248/0x2a0 > > __arm64_sys_ioctl+0x9c0/0xb00 > >=20 >=20This may trigger a new srcu grace period again during the window peri= od between > srcu-barrier() and cleanup_srcu_struct(). >=20 >=20>=20 >=20> I also put a pr_warn() at out_cleanup_tags_srcu: to be sure -- it f= ired > > zero times over the whole run, so loop_add() never takes that error = path > > in this workload. > >=20=20 >=20> Why do you wangt to have this > > WARN_ON(timer_delete_sync(&sdp->delay_work)) ? > >=20=20 >=20> timer_delete_sync() !=3D 0 "a timer was armed", not "callbacks are > > pending". > >=20 >=20There are only two types of return values for timer_delete_sync(), > return 0 or 1, the timer_delete_sync() !=3D 0 means that this timer > was pending and has been deactivated, right? How about this? diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c index b9fe57ff9100..62da31ec2bd0 100644 --- a/kernel/rcu/srcutree.c +++ b/kernel/rcu/srcutree.c @@ -751,7 +751,8 @@ void cleanup_srcu_struct(struct srcu_struct *ssp) // Call srcu_barrier() before this cleanup_srcu_struct() // to avoid triggering this WARN_ON(). - if (WARN_ON(timer_delete_sync(&sdp->delay_work)) && + if (WARN_ON(rcu_segcblist_n_cbs(&sdp->srcu_cblist) && + timer_delete_sync(&sdp->delay_work)) && rcu_cpu_beenfullyonline(sdp->cpu)= ) queue_work_on(sdp->cpu, rcu_gp_wq, &sdp->work); flush_work(&sdp->work); Thanks Zqiang >=20 >=20Thanks > Zqiang >=20 >=20>=20 >=20>=20 >=20> >