[PATCH] drm/amdgpu: force complete the MES scheduler ring fence on reset

Jesse Zhang <[email protected]> Thu, 6 Aug 2026 15:56:47 +0800
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <[email protected]>
The MES scheduler ring has no drm scheduler (no_scheduler = true), so it is
skipped by the force-completion loop in amdgpu_device_pre_asic_reset(). It
uses a polling fence whose hw value lives in wb (GTT) memory and survives a
MODE1 reset, while fence_drv.sync_seq keeps advancing for every packet.

When the reset is triggered because MES itself stopped responding, the
timed-out packets advance sync_seq past the last hw fence value MES wrote.
After resume the first MES submission polls forever on a seq that is never
written back, failing the resume and wedging the box on a second reset:

  amdgpu: MES ring buffer is full.
  amdgpu: *ERROR* ring gfx_0.0.0 test failed (-110)
  amdgpu: resume of IP block <gfx_v11_0> failed -110
  amdgpu: GPU reset end with ret = -110

Force complete the MES scheduler ring fence together with the scheduler
rings so its hw fence is realigned to sync_seq.

Signed-off-by: Jesse Zhang <[email protected]>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
index 638dfaf127fd..5ad13afa87c8 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
@@ -4988,6 +4988,18 @@ int amdgpu_device_pre_asic_reset(struct amdgpu_device *adev,
 		amdgpu_fence_driver_force_completion(ring, fence);
 	}
 
+	/*
+	 * The MES scheduler ring has no drm scheduler (no_scheduler), so it is
+	 * skipped by the loop above. It uses a polling fence whose hw value lives
+	 * in wb memory and survives the reset, while its sync_seq keeps advancing
+	 * for every (possibly timed-out) packet. Force it too, otherwise the first
+	 * post-reset submission polls forever on a seq that will never be written
+	 * back ("MES ring buffer is full") and the resume fails.
+	 */
+	if (adev->mes.ring[0].fence_drv.initialized &&
+	    adev->mes.ring[0].sched.ready)
+		amdgpu_fence_driver_force_completion(&adev->mes.ring[0], fence);
+
 	amdgpu_fence_driver_isr_toggle(adev, false);
 
 	r = amdgpu_reset_prepare_hwcontext(adev, reset_context);
-- 
2.49.0