Re: [PATCH v2 0/4] mm/hmm: Clarify notifier retry state and scope HMM timeouts

Stanislav Kinsburskii <[email protected]> Wed, 15 Jul 2026 07:42:48 -0700
Newsgroups org.freedesktop.lists.nouveau,org.freedesktop.lists.dri-devel,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel,org.kernel.vger.linux-kselftest,org.kernel.vger.linux-rdma,org.kvack.linux-mm
Message-ID <alecaJ1FUso39D9k@skinsburskii>
On Wed, Jul 15, 2026 at 02:41:43PM +0200, David Hildenbrand (Arm) wrote:
> On 7/15/26 00:21, Stanislav Kinsburskii wrote:
> > This small fixup series applies on top of:
> > 
> >   [PATCH v8 0/8] mm/hmm: Add mmap lock-drop support for userfaultfd-backed mappings
> > 
> > The first patch updates the HMM documentation example to make the
> > mmu_interval_read_retry() state explicit: callers should use the notifier and
> > notifier_seq stored in the same hmm_range that was passed to
> > hmm_range_fault_unlocked_timeout().
> > 
> > The remaining patches adjust nouveau, amdxdna, and drm_gpusvm users so the
> > timeout passed to hmm_range_fault_unlocked_timeout() is treated as a relative
> > HMM retry budget. These callers no longer keep an absolute deadline around
> > their outer driver retry loops or pass a computed remaining time into HMM.
> > 
> > This keeps the timeout scoped to HMM's internal mmu-notifier retry handling. If
> > HMM succeeds and the driver later observes an invalidation through
> > mmu_interval_read_retry(), the driver retries the operation with a fresh HMM
> > retry budget.
> > 
> > Changes in v2:
> >   - Kept the nouveau outer absolute timeout around the
> >     mmu_interval_read_retry() loop. hmm_range_fault_unlocked_timeout() only
> >     bounds HMM’s internal retries, while nouveau faults are handled from a GPU
> >     fault worker, so userspace fatal signals cannot break an endless stream of
> >     invalidations there.
> >   - Updated nouveau to use time_after_eq() before calling HMM, so the remaining
> >     timeout passed to hmm_range_fault_unlocked_timeout() is always positive and
> >     never 0, which would mean retry indefinitely.
> >   - Updated the nouveau fixup commit message to explain the worker-thread
> >     timeout issue and the time_after_eq() boundary behavior.
> >   - Fixed the amdxdna fixup commit message. It now describes
> >     aie2_populate_range() correctly instead of carrying stale nouveau prose,
> >     and notes that command submission still keeps its broader timeout while HMM
> >     gets a fresh relative retry budget.
> > 
> > 
> > ---
> > 
> > Stanislav Kinsburskii (4):
> >       fixup! mm/hmm: add hmm_range_fault_unlocked_timeout() for mmap lock-drop support
> >       fixup! drm/nouveau: use hmm_range_fault_unlocked_timeout() for SVM faults
> >       fixup! accel/amdxdna: use hmm_range_fault_unlocked_timeout() for range population
> >       fixup! drm/gpusvm: use hmm_range_fault_unlocked_timeout() for range faults
> 
> Why a fixup series instead of properly resending the full thing?
> 

The goal was to get a Sashiko review, and v8 has already been applied to
both `mm-new` and `linux-next`.

You can find more details here:

  https://sashiko.dev/#/message/alaWmUEeIBeSkmO0%40skinsburskii

Thanks, Stanislav

> -- 
> Cheers,
> 
> David