[PATCH] drm/amdkfd: mark queues as reset after a full GPU reset

YuanShang Mao <[email protected]>
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <[email protected]>
From: Yuansheng Mao <[email protected]>

A full GPU reset reloads the MES firmware and wipes its queue
table, but KFD keeps the queue objects with is_reset left false.
Pre-reset cleanup cannot fix this: it runs with the reset domain
write lock held, so remove_queue_mes() fails with -EIO and the
cleanup loops bail out early.

When the process later exits, process_termination_cpsch() sends
REMOVE_QUEUE for doorbells MES no longer knows about. MES rejects
them, remove_queue_mes_on_reset_option() mistakes the rejection
for a hung queue and ends up in kfd_hws_hang(), requesting another
GPU reset. On SR-IOV that resets the whole card, so one injected
hang costs two resets and disturbs the other VFs.

Observed on Navi32 SR-IOV with a compute workload running:

  ring sdma1 timeout, signaled seq=239, emitted seq=240
  GPU reset begin!. Source:  1
  Failed to evict queue 2
  GPU reset(1) succeeded!
  [ ~34s later, when the killed process is reaped ]
  failed to remove queue from MES, doorbell=0x1002
  MES might be in unrecoverable state, issue a GPU reset
  GPU reset begin!. Source:  3

remove_queue_mes_on_reset_option() already skips queues with
is_reset set, but only the per-queue reset paths ever set it. Set
it in kgd2kfd_post_reset() as well. Suspend/resume is unaffected:
it goes through kgd2kfd_resume(), where queues are re-added to MES.

Signed-off-by: Yuansheng Mao <[email protected]>
---
 drivers/gpu/drm/amd/amdkfd/kfd_device.c | 21 +++++++++++++++++++++
 1 file changed, 21 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_device.c b/drivers/gpu/drm/amd/amdkfd/kfd_device.c
index 586e640f13dc..d326e41a8656 100644
--- a/drivers/gpu/drm/amd/amdkfd/kfd_device.c
+++ b/drivers/gpu/drm/amd/amdkfd/kfd_device.c
@@ -1033,6 +1033,27 @@ int kgd2kfd_post_reset(struct kfd_dev *kfd)
 	if (!kfd->init_complete)
 		return 0;
 
+	/*
+	 * The reset wiped the MES queue table. Mark the queues as reset so
+	 * that teardown does not try to remove them from MES again.
+	 */
+	if (kfd->shared_resources.enable_mes) {
+		for (i = 0; i < kfd->num_nodes; i++) {
+			struct device_queue_manager *dqm = kfd->nodes[i]->dqm;
+			struct device_process_node *cur;
+			struct qcm_process_device *qpd;
+			struct queue *q;
+
+			dqm_lock(dqm);
+			list_for_each_entry(cur, &dqm->queues, list) {
+				qpd = cur->qpd;
+				list_for_each_entry(q, &qpd->queues_list, list)
+					q->properties.is_reset = true;
+			}
+			dqm_unlock(dqm);
+		}
+	}
+
 	for (i = 0; i < kfd->num_nodes; i++) {
 		ret = kfd_resume(kfd->nodes[i]);
 		if (ret)
-- 
2.25.1
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.