[PATCH 0/1] drm/amdkfd: Fix SVM page migration hang on MES-based GPUs

Priya Hosur <[email protected]>
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <[email protected]>
This patch addresses a GPU queue hang issue during SVM (Shared Virtual
Memory) page migration on MES-based GPUs, specifically observed on
gfx1151 (Ryzen AI MAX).

Problem:
--------
KFDSVMRangeTest.MultiThreadMigrationTest/1 fails with a queue consumption
timeout when XNACK mode 1 is enabled. The GPU compute queue hangs with
packets submitted but never consumed.

Root Cause:
-----------
MES (Micro Engine Scheduler) does not perform heavy-weight TLB invalidation
after unmapping queues, unlike HWS which does this automatically. This
causes in-flight SDMA DMA descriptors to access memory that has been
unmapped, leading to page faults and queue hangs.

Fix:
----
Add kfd_flush_tlb() with TLB_FLUSH_HEAVYWEIGHT after MES removes/suspends
queues in evict_process_queues_cpsch() and suspend_queues().

Testing:
--------
Platform: gfx1151, ROCm 7.15.0, kernel 7.0.0-28-generic

Results:
  - Baseline (no fix):     100% failure (50 runs)
  - With TLB flush:        7-10% failure (100 runs)

The TLB flush significantly reduces failures from 100% to ~10%. The
residual failures require further investigation and may be addressed
in a follow-up patch.

Priya Hosur (1):
  drm/amdkfd: Add TLB flush after MES queue eviction/suspension

 .../gpu/drm/amd/amdkfd/kfd_device_queue_manager.c   | 13 ++++++++++++-
 1 file changed, 12 insertions(+), 1 deletion(-)

-- 
2.43.0
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.