[PATCH] drm/amdgpu/mes12: use basic+ unmapped-doorbell handling for concurrent AQL queues
Jesse Zhang <[email protected]> Thu, 30 Jul 2026 17:37:30 +0800
| Newsgroups | org.freedesktop.lists.amd-gfx |
|---|---|
| Message-ID | <[email protected]> |
On gfx12, concurrent dispatch from multiple AQL compute user queues in one process can intermittently stall one queue's first dispatch: its completion never signals, while queue creation succeeds and no obvious dmesg error is reported. With unmapped-doorbell mode 1 (basic), a doorbell hit only sets a coarse level-wide hasReadyQueues flag. Queue-level selection can still be filtered by work_done/is_queue_ready, so under concurrency the ringing queue may be skipped and its new dispatch not fetched in time. Switch mes12 to mode 2 (basic+). In this mode, MES consumes per-queue CP_UNMAPPED_QUEUE bitmaps, identifies the ringing queue, and clears that queue's work_done in the unmapped-doorbell path so it is reconsidered promptly by scheduling. Tested on the amd-unified-interface branch on gfx1201: multiple concurrent AQL compute queues now pass repeatedly. Signed-off-by: Jesse Zhang <[email protected]> --- drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 15 +++++++++++++-- 1 file changed, 13 insertions(+), 2 deletions(-) diff --git a/drivers/gpu/drm/amd/amdgpu/mes_v12_0.c b/drivers/gpu/drm/amd/amdgpu/mes_v12_0.c index 04465804c254..5f9e28f8f007 100644 --- a/drivers/gpu/drm/amd/amdgpu/mes_v12_0.c +++ b/drivers/gpu/drm/amd/amdgpu/mes_v12_0.c @@ -990,10 +990,21 @@ static int mes_v12_0_set_hw_resources(struct amdgpu_mes *mes, int pipe) /* * Keep oversubscribe timer for sdma . When we have unmapped doorbell * handling support, other queue will not use the oversubscribe timer. - * handling mode - 0: disabled; 1: basic version; 2: basic+ version + * handling mode - 0: disabled; 1: basic; 2: basic+; 3: basic++ + * + * Use basic+ (2): on an unmapped-queue doorbell ring MES reads the + * per-queue CP_UNMAPPED_QUEUE bitmap to identify exactly which queue(s) + * rang and clears their work_done so the scheduler re-maps them. The + * basic (1) mode only sets a coarse level-wide "ready" flag without + * clearing the ringing queue's work_done, so a concurrent AQL compute + * queue whose doorbell rings after work_done was set gets skipped by the + * scheduler and its dispatch is silently lost (intermittent hang with + * multiple concurrent queues). basic++ (3) additionally map_gang()s in + * the interrupt handler for lower latency, but basic+ is sufficient and + * touches less state. */ mes_set_hw_res_pkt.oversubscription_timer = mes_rev < 0x8b ? 0 : 50; - mes_set_hw_res_pkt.unmapped_doorbell_handling = 1; + mes_set_hw_res_pkt.unmapped_doorbell_handling = 2; if (amdgpu_mes_log_enable) { mes_set_hw_res_pkt.enable_mes_event_int_logging = 1; -- 2.49.0