Re: [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)

"Paul E. McKenney" <[email protected]>
Newsgroups gmane.linux.kernel,gmane.linux.kernel.mm
Message-ID <23c8020c-3d2c-4c89-9a55-83d64a7133d9@paulmck-laptop>
On Thu, Aug 06, 2026 at 10:18:57AM +0200, Vlastimil Babka (SUSE) wrote:
> On 8/5/26 22:28, Paul E. McKenney wrote:
> > On Wed, Aug 05, 2026 at 12:29:52PM -0700, Andrew Morton wrote:
> >> On Tue, 04 Aug 2026 17:01:48 -0700 syzbot <[email protected]> wrote:
> >> 
> >> > Hello,
> >> > 
> >> > syzbot found the following issue on:
> >> > 
> >> > HEAD commit:    3708dd948844 Merge tag 'pm-7.2-rc6' of git://git.kernel.or..
> >> > git tree:       upstream
> >> > console output: https://syzkaller.appspot.com/x/log.txt?x=11ac703e580000
> >> > kernel config:  https://syzkaller.appspot.com/x/.config?x=4e38b15c29e6a1d9
> >> > dashboard link: https://syzkaller.appspot.com/bug?extid=d2401aeb74cc84adba04
> >> > compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
> >> > 
> >> > Unfortunately, I don't have any reproducer for this issue yet.
> >> 
> >> Thanks.
> >> 
> >> Lazy optimists (ahem) paste this gunk into Gemini and ask "what the
> >> heck just happened".  The results are often useful, but should be
> >> treated with skepticism.  In this case I think it came usably close.
> >> 
> >> 	https://share.gemini.google/vq4TLhTiLBih
> >> 
> >> 
> >> tl;dr: khugepaged's collapse_scan_file() is taking too long and RCU got
> >> starved.  I don't think khugepaged is doing anything wrong here,
> >> per-se.  There's a lot of work to do and we're doing it.
> >> 
> >> An appropriate fix would be to take a break, let RCU do its thing then
> >> get back to work.  But I don't think RCU offers interfaces for that?
> >> 
> >> collapse_scan_file()'s main loop has
> >> 
> >> 		if (need_resched()) {
> >> 			xas_pause(&xas);
> >> 			cond_resched_rcu();
> >> 		}
> >> 
> >> but that won't help with the RCU stall detector(?).
> >> 
> >> I suggest that a suitable fix here would be to add the analogous
> >> 
> >> 	if (rcu_i_need_to_take_a_break()) {
> >> 		rcu_read_unlock();
> >> 		rcu_take_a_break())	
> >> 		rcu_read_lock();
> >> 	}
> >> 
> >> (iirc rcu_read_unlock() does an rcu run, so rcu_take_a_break() isn't
> >> needed here)
> >> 
> >> Paul, wdyt?
> > 
> > Let's see...
> > 
> > The console log says "rcu_preempt detected stalls on CPUs/tasks",
> > which means that cond_resched() is a no-op, but it also means that
> > the rcu_read_unlock() in cond_resched_rcu() will directly take care of
> > informing RCU of the pause.
> > 
> > But that is clearly not happening.  Why?
> > 
> > Well, we have this:
> > 
> > rcu: 	Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l
> > 
> > This means that the task whose RCU read-side critical section is blocking
> > the current RCU grace period isn't even running, and thus cannot invoke
> > cond_resched_rcu(), let alone the rcu_read_unlock() within that function.
> > So an RCU CPU stall warning is expected behavior.  Or at least it is not
> > in any way ruled out.
> > 
> > What we need is RCU priority boosting.  Except that the .config file
> > does not enable this.  Not only is there no CONFIG_RCU_BOOST=y, there
> > is also no CONFIG_RCU_EXPERT=y and no CONFIG_PREEMPT_RT=y.  But there
> > is CONFIG_RT_MUTEX=y and CONFIG_RCU_EXPERT=y.
> 
> It comes from syzbot so might be likely a randconfig and there's no point in
> trying to find any sense in that combination :)

;-) ;-) ;-)

> > Because we don't have RCU priority boosting, if the load on the system
> > is heavy enough to prevent our poor preempted RCU reader (PID 37) from
> > running, the grace period cannot end.
> > 
> > I am not sure why this task is saving its stack, but maybe that is normal
> > for this code path?
> 
> That's because page_owner is also enabled so it's saving the freeing stack
> for the page it's freeing. That's not a normal production config, only when
> debugging.

Ah, OK, I feel much better now.

> > My bemusement aside, I recommend running this test either with
> > non-preemptible RCU (CONFIG_PREEMPT_LAZY=y these days) or enabling RCU
> > priority boosting (CONFIG_RCU_EXPERT=y and CONFIG_RCU_BOOST=y).
> > 
> > Maybe RCU_BOOST should no longer depend on RCU_EXPERT?  I would of
> > course need ot remove the prompt ("Enable RCU priority boosting") to
> > avoid annoying Linus.  Maybe as shown below.
> 
> The "no longer depend" part alone would make no difference with randconfigs.
> Removing the prompt too should help indeed.

Agreed!

> Maybe a possible strategy in general would be indeed to unconditionally
> select what's the expected config, like you did below, and only make it
> possible to override that with RCU_EXPERT. So here with RCU_EXPERT you could
> disable RCU_BOOST even if it was automatically enabled - assuming this is
> useful for development or internal rcu testing by people who know what they
> are doing (not syzbot randconfig) or whatnot.
> 
> But then RCU_EXPERT should be excluded from (impossible to be enabled by)
> randconfig to indicate it's not valid for this kind of testing.
> I don't know if there's any precedent for such a strategy.

Good point!  For the first cut, I will just force it, so that someone
wanting to do that sort of testing gets to edit the Kconfig file, but
it is good to have a trick like that in my back pocket, so thank you!

> Specifically for the proposal below, could the problem still happen with
> PREEMPT_RCU without RT_MUTEXES? If yes, it wouldn't be enough?

Quite true!  Maybe I should make PREEMPT_RCU select RT_MUTEXES?

But that might need a bit of discussion, so if I take that approach
it needs to be a separate patch.  ;-)

							Thanx, Paul

> > ------------------------------------------------------------------------
> > 
> > diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig
> > index 1a5fb3156c062a..5141ad8d1cd029 100644
> > --- a/kernel/rcu/Kconfig
> > +++ b/kernel/rcu/Kconfig
> > @@ -237,17 +237,16 @@ config RCU_FANOUT_LEAF
> >  	  Take the default if unsure.
> >  
> >  config RCU_BOOST
> > -	bool "Enable RCU priority boosting"
> > -	depends on (RT_MUTEXES && PREEMPT_RCU && RCU_EXPERT) || PREEMPT_RT
> > +	bool
> > +	depends on (RT_MUTEXES && PREEMPT_RCU) || PREEMPT_RT
> >  	default y if PREEMPT_RT
> >  	help
> >  	  This option boosts the priority of preempted RCU readers that
> >  	  block the current preemptible RCU grace period for too long.
> >  	  This option also prevents heavy loads from blocking RCU
> > -	  callback invocation.
> > +	  callback invocation.  It is now automatically enabled in
> > +	  any kernel that can benefit from it and that can support it.
> >  
> > -	  Say Y here if you are working with real-time apps or heavy loads
> > -	  Say N here if you are unsure.
> >  
> >  config RCU_BOOST_DELAY
> >  	int "Milliseconds to delay boosting after RCU grace-period start"
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.