Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL

Tales A. Mendonça <[email protected]> Tue, 4 Aug 2026 20:00:06 -0300
Newsgroups org.freedesktop.lists.intel-xe,org.freedesktop.lists.dri-devel
Message-ID <CAHBRX4ErWF0e4X9DCXyycyqJUHaVcv36J=X0PRrNANqZxrzidg@mail.gmail.com>
> Are we seeing this on i915?

I have not tested i915 on the affected machines yet - the
instrumentation that measured the stalls (late-ack logging, kick
results) is xe-only, so I have no comparable i915 data. I can boot one
of the ARL machines with i915 for a few days and watch for TLB
invalidation timeouts there, if that data helps.

More generally: both machines here reproduce reliably (~1 stall/hour
on a desktop workload, much more under memory pressure), so I am happy
to test anything on them - including any GuC build the firmware team
would like data on.

On Stuart's masking concern: fully agreed, that is why patch 3 is
marked RFC. Patches 1-2 are pure diagnostics and stand on their own; I
am fine holding patch 3 until the firmware side has been looked at.
The data point it adds is that a doorbell ring unblocks the ack in the
majority of episodes, while the severe ones ignore 8-9 consecutive
rings - hopefully that narrows where to look inside the GuC.

I will send a v2 addressing Matt's review comments (the
__xe_devcoredump unification and the fixes on patch 2).

Thanks,
Tales

Em ter., 4 de ago. de 2026 às 19:08, Daniele Ceraolo Spurio
<[email protected]> escreveu:
>
>
>
> On 8/4/2026 2:33 PM, Summers, Stuart wrote:
> > On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
> >> On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
> >>> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> >>>> Hi,
> >>>>
> >>>> This series is a follow-up to the TLB invalidation ack stall I
> >>>> have
> >>>> been debugging on ARL, tracked in:
> >>>>
> >>>>    https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> >>>>
> >>>> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
> >>>> and
> >>>> 7dd1 machines here, plus an independent Arc Pro 130T report on
> >>>> the
> >>>> issue above), TLB invalidation acks intermittently stall for
> >>>> ~2.3s.
> >>>> The H2G request is consumed from the CTB immediately and the G2H
> >>>> CTB
> >>>> is empty the whole time - the firmware simply does not send the
> >>>> ack
> >>>> until much later. The fence timeout fires at 2.25s and the ack
> >>>> lands
> >>>> tens of ms after it. Userspace blocked on the invalidation
> >>>> (compositor
> >>>> buffer unmaps etc.) hitches for the full window.
> >>> Firstly, thanks for the patch!
> >>>
> >>> I haven't looked in to all the details of the sighting you were
> >>> debugging, but we have had similar issues that were fixed in a
> >>> later
> >>> GuC version. I think around 70.60.0? It might be worth trying on
> >>> something later than that to see if that helps... (+Daniele)
> >>>
> >> I think this would require an AR on our end to make a new firmware
> >> version available.
> >>
> >> The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
> >> aliases to MTL for firmware). (+Julia too).
> >>
> >> Presumably, the GuC changelogs should indicate whether an issue
> >> related
> >> this has been fixed. If so, we need to update all GuC versions across
> >> both i915 and Xe.
> > Right... I guess I'd still like to see if we can test this in GuC (or
> > get confirmation we can't for some reason) before committing something.
> > My worry is we will prevent bug reports like this by working around it
> > and miss critical bugs that need to be fixed in the right component.
> >
> >> [1]
> >> https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
> >>
> >>>> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
> >>>> no
> >>> Is there a reason we don't just re-use the main xe_devcoredump()?
> >>>
> >> This is my suggestion: the main devcoredump infrastructure is job-
> >> based,
> >> so it cannot be used for hangs that are not associated with a job.
> >>
> >> In my opinion, this is a gap on our end. Introducing something like
> >> `xe_devcoredump_gt()`, which can be used for non-job-based hangs
> >> (e.g.,
> >> TLB invalidation timeouts like those addressed in this series, or
> >> more
> >> generally any GuC protocol hang), makes sense to me.
> > Ok makes sense. We can do that here. It would be nice to have a more
> > inclusive implementation that lets us call this from anywhere so we
> > aren't duplicating things around for different use cases. But not a
> > blocker here.
> >
> >> I haven't looked at the patch yet, but at a high level, adding
> >> `xe_devcoredump_gt()` seems like a reasonable approach.
> >>
> >>>> exec queue or job to blame - leaves a devcoredump with the GuC
> >>>> log
> >>>> and
> >>>> CT state behind (Matt suggested capturing devcoredumps when we
> >>>> discussed the issue; devcoredumps from both machines are attached
> >>>> to
> >>>> the issue above).
> >>>>
> >>>> Patch 2 logs when the ack for a timed out invalidation finally
> >>>> arrives. This is what established that the acks are late rather
> >>>> than
> >>>> lost.
> >>>>
> >>>> Patch 3 is the RFC part: a delayed work that pokes the GuC
> >>>> (status
> >>>> register read, CT flush, doorbell ring) every 250ms while an ack
> >>>> is
> >>>> overdue. On my machines this converts the guaranteed 2.3s stall
> >>>> into
> >>> I'm a little worried we're just papering over something here that
> >>> needs
> >>> to be addressed in GuC, particularly around GT going to sleep or
> >>> something around the time we're expecting a response, so the pings
> >>> on
> >>> registers might be prematurely waking things up which is something
> >>> we'd
> >>> want to happen in GuC, not the KMD.
> >>>
> >> In general, I agree with this. We should avoid papering over the
> >> issue
> >> and instead fix it properly in the GuC. That said, this workaround
> >> provides a pretty strong data point, since it appears to get the TLB
> >> invalidation unstuck.
> > So if we hit this issue I guess we're already going to have some
> > performance degredation and the workaround makes that better. I need to
> > look at the implementation, but we could be potentially introducing
> > performance penalties in other areas doing these pings.
> >
> > Again, I'd like to see if we can fix this in the right place before
> > implementing a workaround for it. Hopefully Daniele or Julia can give
> > some direction there.
>
> Are we seeing this on i915 at all? Given that Xe does not officially
> support MTL/ARL and is missing several critical WAs for those platforms,
> the approach so far has been to only update the GuC FW if it is required
> for i915.
> Looking at the GuC release notes, there have been a couple of
> TLB-related fixes after 70.53, but they're both marked as only affecting
> PVC and Xe2+ platforms, so no fixes seem to be available for ARL (or at
> least they're not listed in the release notes).
>
> Daniele
>
> >
> > Thanks,
> > Stuart
> >
> >> Matt
> >>
> >>> Thanks,
> >>> Stuart
> >>>
> >>>> a
> >>>> sub-500ms hiccup for the majority of occurrences; a minority of
> >>>> severe
> >>>> episodes ignore 8-9 consecutive doorbells, which points at the
> >>>> GuC
> >>>> firmware being internally blocked for the whole window. Full data
> >>>> on
> >>>> the issue. I am happy to rework the approach (different delay,
> >>>> tying it to the G2H handler, dropping the status read, etc.) -
> >>>> mainly
> >>>> I would like the firmware side investigated, since no host-side
> >>>> poke
> >>>> can fix the severe cases.
> >>>>
> >>>> Based on drm-tip. Tested for several days on both ARL machines
> >>>> under
> >>>> desktop and VM-heavy workloads.
> >>>>
> >>>> Thanks,
> >>>> Tales
> >>>>
> >>>> Tales A. Mendonça (3):
> >>>>    drm/xe: Capture devcoredump on TLB invalidation timeout
> >>>>    drm/xe: Log when a timed out TLB invalidation ack finally
> >>>> arrives
> >>>>    drm/xe: Kick GuC while TLB invalidation acks are overdue
> >>>>
> >>>>   drivers/gpu/drm/xe/xe_devcoredump.c     |  68 ++++++++++++
> >>>>   drivers/gpu/drm/xe/xe_devcoredump.h     |   6 ++
> >>>>   drivers/gpu/drm/xe/xe_tlb_inval.c       | 131
> >>>> +++++++++++++++++++++++-
> >>>>   drivers/gpu/drm/xe/xe_tlb_inval_types.h |  42 ++++++++
> >>>>   4 files changed, 243 insertions(+), 4 deletions(-)
> >>>>
>


-- 
Com os cumprimentos,

Tales A. Mendonça
talesam.org
communitybig.org