Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL

Tales A. Mendonça <[email protected]>
Newsgroups org.freedesktop.lists.dri-devel,org.freedesktop.lists.intel-xe
Message-ID <CAHBRX4GYEKniGb8jw-h=QYPi=FpbWRKwHHnfw-K5wYiy6AWfJA@mail.gmail.com>
Patchwork reports my address is not on the CI allowlist, so CI was not
triggered for this series:

    Series author address '[email protected]' is not on the allowlist,
which prevents CI from being automatically triggered.

  Could one of the project owners click 'retest' on the series (and/or
add me to the allowlist)? Series URL:

    https://patchwork.freedesktop.org/series/171539/

  Thanks!
  Tales


Em seg., 3 de ago. de 2026 às 23:14, Tales A. Mendonça
<[email protected]> escreveu:
>
> Hi,
>
> This series is a follow-up to the TLB invalidation ack stall I have
> been debugging on ARL, tracked in:
>
>   https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>
> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> 7dd1 machines here, plus an independent Arc Pro 130T report on the
> issue above), TLB invalidation acks intermittently stall for ~2.3s.
> The H2G request is consumed from the CTB immediately and the G2H CTB
> is empty the whole time - the firmware simply does not send the ack
> until much later. The fence timeout fires at 2.25s and the ack lands
> tens of ms after it. Userspace blocked on the invalidation (compositor
> buffer unmaps etc.) hitches for the full window.
>
> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
> exec queue or job to blame - leaves a devcoredump with the GuC log and
> CT state behind (Matt suggested capturing devcoredumps when we
> discussed the issue; devcoredumps from both machines are attached to
> the issue above).
>
> Patch 2 logs when the ack for a timed out invalidation finally
> arrives. This is what established that the acks are late rather than
> lost.
>
> Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> register read, CT flush, doorbell ring) every 250ms while an ack is
> overdue. On my machines this converts the guaranteed 2.3s stall into a
> sub-500ms hiccup for the majority of occurrences; a minority of severe
> episodes ignore 8-9 consecutive doorbells, which points at the GuC
> firmware being internally blocked for the whole window. Full data on
> the issue. I am happy to rework the approach (different delay,
> tying it to the G2H handler, dropping the status read, etc.) - mainly
> I would like the firmware side investigated, since no host-side poke
> can fix the severe cases.
>
> Based on drm-tip. Tested for several days on both ARL machines under
> desktop and VM-heavy workloads.
>
> Thanks,
> Tales
>
> Tales A. Mendonça (3):
>   drm/xe: Capture devcoredump on TLB invalidation timeout
>   drm/xe: Log when a timed out TLB invalidation ack finally arrives
>   drm/xe: Kick GuC while TLB invalidation acks are overdue
>
>  drivers/gpu/drm/xe/xe_devcoredump.c     |  68 ++++++++++++
>  drivers/gpu/drm/xe/xe_devcoredump.h     |   6 ++
>  drivers/gpu/drm/xe/xe_tlb_inval.c       | 131 +++++++++++++++++++++++-
>  drivers/gpu/drm/xe/xe_tlb_inval_types.h |  42 ++++++++
>  4 files changed, 243 insertions(+), 4 deletions(-)
>
> --
> 2.55.0
>


-- 
Com os cumprimentos,

Tales A. Mendonça
talesam.org
communitybig.org
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.