Re: [PATCH] drm/amdkfd: mark queues as reset after a full GPU reset

Alex Deucher <[email protected]>
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <CADnq5_P8Vyzysy0SBncfK0AkxJ9t5A3p_0ccBfBMeF7SVv_Jtw@mail.gmail.com>
On Thu, Aug 20, 2026 at 8:40 AM YuanShang Mao <[email protected]> wrote:
>
> From: Yuansheng Mao <[email protected]>
>
> A full GPU reset reloads the MES firmware and wipes its queue
> table, but KFD keeps the queue objects with is_reset left false.
> Pre-reset cleanup cannot fix this: it runs with the reset domain
> write lock held, so remove_queue_mes() fails with -EIO and the
> cleanup loops bail out early.
>
> When the process later exits, process_termination_cpsch() sends
> REMOVE_QUEUE for doorbells MES no longer knows about. MES rejects
> them, remove_queue_mes_on_reset_option() mistakes the rejection
> for a hung queue and ends up in kfd_hws_hang(), requesting another
> GPU reset. On SR-IOV that resets the whole card, so one injected
> hang costs two resets and disturbs the other VFs.
>
> Observed on Navi32 SR-IOV with a compute workload running:
>
>   ring sdma1 timeout, signaled seq=239, emitted seq=240
>   GPU reset begin!. Source:  1
>   Failed to evict queue 2
>   GPU reset(1) succeeded!
>   [ ~34s later, when the killed process is reaped ]
>   failed to remove queue from MES, doorbell=0x1002
>   MES might be in unrecoverable state, issue a GPU reset
>   GPU reset begin!. Source:  3
>
> remove_queue_mes_on_reset_option() already skips queues with
> is_reset set, but only the per-queue reset paths ever set it. Set
> it in kgd2kfd_post_reset() as well. Suspend/resume is unaffected:
> it goes through kgd2kfd_resume(), where queues are re-added to MES.
>
> Signed-off-by: Yuansheng Mao <[email protected]>
> ---
>  drivers/gpu/drm/amd/amdkfd/kfd_device.c | 21 +++++++++++++++++++++
>  1 file changed, 21 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_device.c b/drivers/gpu/drm/amd/amdkfd/kfd_device.c
> index 586e640f13dc..d326e41a8656 100644
> --- a/drivers/gpu/drm/amd/amdkfd/kfd_device.c
> +++ b/drivers/gpu/drm/amd/amdkfd/kfd_device.c
> @@ -1033,6 +1033,27 @@ int kgd2kfd_post_reset(struct kfd_dev *kfd)
>         if (!kfd->init_complete)
>                 return 0;
>
> +       /*
> +        * The reset wiped the MES queue table. Mark the queues as reset so
> +        * that teardown does not try to remove them from MES again.
> +        */
> +       if (kfd->shared_resources.enable_mes) {

Why is this MES specific?  Shouldn't this be true for KIQ as well?

Alex

> +               for (i = 0; i < kfd->num_nodes; i++) {
> +                       struct device_queue_manager *dqm = kfd->nodes[i]->dqm;
> +                       struct device_process_node *cur;
> +                       struct qcm_process_device *qpd;
> +                       struct queue *q;
> +
> +                       dqm_lock(dqm);
> +                       list_for_each_entry(cur, &dqm->queues, list) {
> +                               qpd = cur->qpd;
> +                               list_for_each_entry(q, &qpd->queues_list, list)
> +                                       q->properties.is_reset = true;
> +                       }
> +                       dqm_unlock(dqm);
> +               }
> +       }
> +
>         for (i = 0; i < kfd->num_nodes; i++) {
>                 ret = kfd_resume(kfd->nodes[i]);
>                 if (ret)
> --
> 2.25.1
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.