Re: [Re] GPU reset when running the ROCm hsa runtime tests on gfx12 and next-20260701

Bert Karwatzki <[email protected]>
Newsgroups org.kernel.vger.linux-next,org.freedesktop.lists.amd-gfx,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
Am Montag, dem 10.08.2026 um 10:15 -0400 schrieb Alex Deucher:
> @Amber Lin
> , @Zhang, Jesse(Jie), @Shaoyun Liu  Can you take a look at the firmware version?
> 
> Alex
> 
> On Mon, Jul 6, 2026 at 9:34 AM Bert Karwatzki <[email protected]> wrote:
> > 
> > I found the real cause of my problems:
> > 
> > 
> > bool amdgpu_mes_queue_reset_by_mes_supported(struct amdgpu_device *adev)
> > {
> >  u32 ip_maj = IP_VERSION_MAJ(amdgpu_ip_version(adev, GC_HWIP, 0));
> >  u32 ip_min = IP_VERSION_MIN(amdgpu_ip_version(adev, GC_HWIP, 0));
> >  u32 mes_sched = adev->mes.sched_version & AMDGPU_MES_VERSION_MASK;
> >  printk(KERN_INFO "%s: ip_maj = %u ip_min = %u mes_sched = 0x%x", __func__, ip_maj, ip_min, mes_sched);
> > 
> >  return (ip_maj == 11 && mes_sched >= 0x8c) ||
> >  ((ip_maj == 12 && ip_min == 0) && mes_sched >= 0x8d) ||
> >  ((ip_maj == 12 && ip_min == 1) && mes_sched >= 0x73);
> > }
> > 
> > returns false on my machine (because mes_sched is not large enough)
> > 
> > [ T8549] amdgpu_mes_queue_reset_by_mes_supported: ip_maj = 12 ip_min = 0 mes_sched = 0x76
> > 
> > So I skipped the call to amdgpu_mes_queue_reset_by_mes_supported()
> > 
> > diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> > index 5c9dfb0c424f..462f20aeb681 100644
> > --- a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> > +++ b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> > @@ -448,17 +448,22 @@ static int reset_queues_mes(struct device_queue_manager *dqm, struct queue *q)
> >         unsigned int num_hung = 0;
> >         int r = 0;
> >         struct mes_remove_queue_input queue_input;
> > +       printk(KERN_INFO "entering %s", __func__);
> > 
> > +       /*
> >         if (!amdgpu_mes_queue_reset_by_mes_supported(adev)) {
> >                 r = -ENOTRECOVERABLE;
> > +               printk(KERN_INFO "%s: reset by mes not supported", __func__);
> >                 goto fail;
> > -       }
> > +       }*/
> > +       printk(KERN_INFO "%s: skip calling amdgpu_mes_queue_reset_by_mes_supported()", __func__);
> > 
> >         /* reset should be used only in dqm locked queue reset */
> >         if (WARN_ON(dqm->detect_hang_count > 0))
> >                 return 0;
> > 
> >         if (!amdgpu_gpu_recovery) {
> > +               printk(KERN_INFO "%s: gpu recovery not enabled", __func__);
> >                 r = -ENOTRECOVERABLE;
> >                 goto fail;
> >         }
> > @@ -470,6 +475,7 @@ static int reset_queues_mes(struct device_queue_manager *dqm, struct queue *q)
> >         queue_input.xcc_id = ffs(dqm->dev->xcc_mask) - 1;
> >         /* pass the known bad queue info to the reset function */
> >         r = amdgpu_gfx_reset_mes_compute(adev, NULL, NULL, NULL, &num_hung, &queue_input);
> > +       printk(KERN_INFO "%s: amdgpu_gfx_reset_mes_compute() returned %d", __func__, r);
> >         if (r)
> >                 goto fail;
> > 
> > @@ -3231,6 +3237,7 @@ int kfd_dqm_suspend_bad_queue_mes(struct kfd_node *knode, u32 pasid, u32 doorbel
> >         struct qcm_process_device *qpd;
> >         struct queue *q = NULL;
> >         int ret = 0;
> > +       printk(KERN_INFO "entering %s", __func__);
> > 
> >         if (!pdd)
> >                 return -EINVAL;
> > @@ -3242,6 +3249,7 @@ int kfd_dqm_suspend_bad_queue_mes(struct kfd_node *knode, u32 pasid, u32 doorbel
> > 
> >                 list_for_each_entry(q, &qpd->queues_list, list) {
> >                         if (q->doorbell_id == doorbell_id && q->properties.is_active) {
> > +                               printk(KERN_INFO "%s: calling recover_bad_queue_mes() for queue %px", __func__, q);
> > 
> > and got this output when running the hsaruntime (which inject illegal opcodes into the
> > command stream) test:
> > 
> > [  113.811612] [    T645] [drm:gfx_v12_0_bad_op_irq [amdgpu]] *ERROR* Illegal opcode in command stream
> > [  113.811675] [   T2558] entering kfd_dqm_suspend_bad_queue_mes
> > [  113.811676] [   T2558] kfd_dqm_suspend_bad_queue_mes: calling recover_bad_queue_mes() for queue ffffa03160950400
> > [  113.811676] [   T2558] entering reset_queues_mes
> > [  113.811677] [   T2558] reset_queues_mes: skip calling amdgpu_mes_queue_reset_by_mes_supported()
> > [  113.811887] [   T2558] reset_queues_mes: amdgpu_gfx_reset_mes_compute() returned 0
> > 
> > So even though  amdgpu_mes_queue_reset_by_mes_supported() reported false,
> > amdgpu_gfx_reset_mes_compute() returns 0, suggesting that resetting actually works here!
> > 
> > Perhaps the minimum required mes_sched version for 12.0 can be relaxed to 0x76 to solve this:
> > 
> > From 6d7af652177063963012eb4df228e99caeb03b31 Mon Sep 17 00:00:00 2001
> > From: Bert Karwatzki <[email protected]>
> > Date: Mon, 6 Jul 2026 14:36:22 +0200
> > Subject: [PATCH] amdgpu: relax required mes_sched version
> > 
> > This mes_sched version is actually enough on this hardware:
> > 03:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 44 [Radeon RX 9060 XT] [1002:7590] (rev c0)
> > 
> > [    3.200538] [    T537] amdgpu 0000:03:00.0: initializing kernel modesetting (IP DISCOVERY 0x1002:0x7590 0x148C:0x2437 0xC0).
> > [    3.200546] [    T537] amdgpu 0000:03:00.0: register mmio base: 0xDFC00000
> > [    3.200547] [    T537] amdgpu 0000:03:00.0: register mmio size: 524288
> > [    3.203739] [    T537] amdgpu 0000:03:00.0: detected ip block number 0 <common_v1_0_0> (soc24_common)
> > [    3.203740] [    T537] amdgpu 0000:03:00.0: detected ip block number 1 <gmc_v12_0_0> (gmc_v12_0)
> > [    3.203741] [    T537] amdgpu 0000:03:00.0: detected ip block number 2 <ih_v7_0_0> (ih_v7_0)
> > [    3.203741] [    T537] amdgpu 0000:03:00.0: detected ip block number 3 <psp_v14_0_0> (psp)
> > [    3.203742] [    T537] amdgpu 0000:03:00.0: detected ip block number 4 <smu_v14_0_0> (smu)
> > [    3.203742] [    T537] amdgpu 0000:03:00.0: detected ip block number 5 <dce_v1_0_0> (dm)
> > [    3.203743] [    T537] amdgpu 0000:03:00.0: detected ip block number 6 <gfx_v12_0_0> (gfx_v12_0)
> > [    3.203743] [    T537] amdgpu 0000:03:00.0: detected ip block number 7 <sdma_v7_0_0> (sdma_v7_0)
> > [    3.203744] [    T537] amdgpu 0000:03:00.0: detected ip block number 8 <vcn_v5_0_0> (vcn_v5_0_0)
> > [    3.203744] [    T537] amdgpu 0000:03:00.0: detected ip block number 9 <jpeg_v5_0_0> (jpeg_v5_0_0)
> > [    3.203745] [    T537] amdgpu 0000:03:00.0: detected ip block number 10 <mes_v12_0_0> (mes_v12_0)
> > 
> > Signed-off-by: Bert Karwatzki <[email protected]>
> > ---
> >  drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c | 2 +-
> >  1 file changed, 1 insertion(+), 1 deletion(-)
> > 
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c
> > index 6c0dde3786e3..c88fdc8a187d 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c
> > @@ -869,7 +869,7 @@ bool amdgpu_mes_queue_reset_by_mes_supported(struct amdgpu_device *adev)
> >  u32 mes_sched = adev->mes.sched_version & AMDGPU_MES_VERSION_MASK;
> > 
> >  return (ip_maj == 11 && mes_sched >= 0x8c) ||
> > - ((ip_maj == 12 && ip_min == 0) && mes_sched >= 0x8d) ||
> > + ((ip_maj == 12 && ip_min == 0) && mes_sched >= 0x76) ||
> >  ((ip_maj == 12 && ip_min == 1) && mes_sched >= 0x73);
> >  }
> > 
> > --
> > 2.53.0
> > 
> > 
> > Bert Karwatzki

These are the firmware version reported in sysfs:

# find /sys/class/drm/card0/device/fw_version/ -print -type f -exec cat {} \;
/sys/class/drm/card0/device/fw_version/
/sys/class/drm/card0/device/fw_version/pfp_fw_version
0x00000c12
/sys/class/drm/card0/device/fw_version/vcn_fw_version
0x0910c01a
/sys/class/drm/card0/device/fw_version/mec_fw_version
0x00000d0c
/sys/class/drm/card0/device/fw_version/mes_fw_version
0x0000008b
/sys/class/drm/card0/device/fw_version/dmcub_fw_version
0x0a000c00
/sys/class/drm/card0/device/fw_version/me_fw_version
0x00000bc2
/sys/class/drm/card0/device/fw_version/asd_fw_version
0x21000110
/sys/class/drm/card0/device/fw_version/sdma2_fw_version
0x00798e96
/sys/class/drm/card0/device/fw_version/sdma_fw_version
0x00798e96
/sys/class/drm/card0/device/fw_version/smc_fw_version
0x00664600
/sys/class/drm/card0/device/fw_version/sos_fw_version
0x003b0f0d
/sys/class/drm/card0/device/fw_version/imu_fw_version
0x0c2c2500
/sys/class/drm/card0/device/fw_version/rlc_fw_version
0x00bde160
/sys/class/drm/card0/device/fw_version/mes_kiq_fw_version
0x0000008b

In addition dmesg reports a problem with the mes version:

[    4.353074] [    T489] amdgpu 0000:03:00.0: MES firmware reports incorrect version in ucode binary (0x1 vs 0x8b)

Both these result use the firmware package shipped by debian sid:
ii  firmware-amd-graphics         20260622-1      all     Binary firmware for AMD/ATI graphics and NPU chips

The earlier result of mes_sched = 0x76 is probably from the old firmware (debian stable version 20250410), but
the gpu resets also occur with the new firmware from debian sid.

Bert Karwatzki
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.