Re: [PATCH] mm/mmu_notifier: Remove non_block_start/end() from notifier invocation
Jason Gunthorpe <[email protected]>
| Newsgroups | dev.linux.lists.linux-rt-devel,org.kernel.vger.kvm,org.kernel.vger.linux-kernel,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
On Tue, Aug 11, 2026 at 03:21:35PM +0100, David Woodhouse wrote: > > > - On PREEMPT_RT, spinning locks become sleeping locks, and perfectly > > > legitimate spinlock/rwlock usage in notifier implementations (e.g. > > > KVM's mn_invalidate_lock and gfn_to_pfn_cache locks) triggers the > > > splat despite having no allocator dependency whatsoever. This is > > > reproducible today on a PREEMPT_RT kernel: KVM takes > > > kvm->mn_invalidate_lock in kvm_mmu_notifier_invalidate_range_start(), > > > and if the OOM reaper reaps a KVM process the result is a "BUG: > > > sleeping function called from invalid context" from > > > rt_spin_lock(). > > > > I don't know anything about PREEEMPT_RT, but this seems like an issue > > with RT if a traditionally atomic safe functions are now triggering > > might sleep failures? > > I can sympathise with that point of view. In fact I've spent the last > couple of years mostly ignoring this "problem" and just blaming RT for > doing exactly that, but I don't think we can really get away with it > any more. > > cf. https://lore.kernel.org/all/[email protected]/ If might_sleep doesn't work sanely at all in preempt_rt then just globally turn it off? > > > - A notifier implementation may legitimately need to wait for an RCU > > > grace period before allowing the caller to proceed with unmapping > > > > That's not allowed. We really want to forbid that, it is not an > > acceptable way to implement a driver using these APIs due to > > performance. > > Speak for yourself. For the KVM gfn-to-pfn-cache the performance scales > *much* better with RCU than with explicit locking: > https://lore.kernel.org/all/[email protected]/ At the cost of completely destroying the mm shootdown performance with 1s RCU grace period waits every mm operation. No thanks. The unstated secondary purprose of the atomic context is to force the driver implementors to make sane choices that don't degrade the MM spectacularly. > Perhaps we could find a way to push down an *accurate* sanity check > into the code paths where what you say is *true*? I guess it could be > done with a flag on each notifier? Or *into* the notifier callback > function(s)? I think it is right and correct the way it is. Jason