Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
Matthew Brost <[email protected]> Tue, 4 Aug 2026 14:29:27 -0700
| Newsgroups | org.freedesktop.lists.intel-xe,org.freedesktop.lists.dri-devel |
|---|---|
| Message-ID | <[email protected]> |
On Tue, Aug 04, 2026 at 01:34:19PM -0300, Tales A. Mendonça wrote: > Patchwork reports my address is not on the CI allowlist, so CI was not > triggered for this series: > > Series author address '[email protected]' is not on the allowlist, > which prevents CI from being automatically triggered. > > Could one of the project owners click 'retest' on the series (and/or > add me to the allowlist)? Series URL: > > https://patchwork.freedesktop.org/series/171539/ > We'd have to resend this ourselves. I can do this, but I've requested for you to be on our allow list as well. I'll ping here once that goes through. Matt > Thanks! > Tales > > > Em seg., 3 de ago. de 2026 às 23:14, Tales A. Mendonça > <[email protected]> escreveu: > > > > Hi, > > > > This series is a follow-up to the TLB invalidation ack stall I have > > been debugging on ARL, tracked in: > > > > https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678 > > > > Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and > > 7dd1 machines here, plus an independent Arc Pro 130T report on the > > issue above), TLB invalidation acks intermittently stall for ~2.3s. > > The H2G request is consumed from the CTB immediately and the G2H CTB > > is empty the whole time - the firmware simply does not send the ack > > until much later. The fence timeout fires at 2.25s and the ack lands > > tens of ms after it. Userspace blocked on the invalidation (compositor > > buffer unmaps etc.) hitches for the full window. > > > > Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no > > exec queue or job to blame - leaves a devcoredump with the GuC log and > > CT state behind (Matt suggested capturing devcoredumps when we > > discussed the issue; devcoredumps from both machines are attached to > > the issue above). > > > > Patch 2 logs when the ack for a timed out invalidation finally > > arrives. This is what established that the acks are late rather than > > lost. > > > > Patch 3 is the RFC part: a delayed work that pokes the GuC (status > > register read, CT flush, doorbell ring) every 250ms while an ack is > > overdue. On my machines this converts the guaranteed 2.3s stall into a > > sub-500ms hiccup for the majority of occurrences; a minority of severe > > episodes ignore 8-9 consecutive doorbells, which points at the GuC > > firmware being internally blocked for the whole window. Full data on > > the issue. I am happy to rework the approach (different delay, > > tying it to the G2H handler, dropping the status read, etc.) - mainly > > I would like the firmware side investigated, since no host-side poke > > can fix the severe cases. > > > > Based on drm-tip. Tested for several days on both ARL machines under > > desktop and VM-heavy workloads. > > > > Thanks, > > Tales > > > > Tales A. Mendonça (3): > > drm/xe: Capture devcoredump on TLB invalidation timeout > > drm/xe: Log when a timed out TLB invalidation ack finally arrives > > drm/xe: Kick GuC while TLB invalidation acks are overdue > > > > drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++ > > drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++ > > drivers/gpu/drm/xe/xe_tlb_inval.c | 131 +++++++++++++++++++++++- > > drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++ > > 4 files changed, 243 insertions(+), 4 deletions(-) > > > > -- > > 2.55.0 > > > > > -- > Com os cumprimentos, > > Tales A. Mendonça > talesam.org > communitybig.org