Re: [PATCH 1/1] drm/amdkfd: Add TLB flush after MES queue eviction/suspension

"Kuehling, Felix" <[email protected]>
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <[email protected]>
On 2026-08-12 02:11, Priya Hosur wrote:
> MES (Micro Engine Scheduler) does not perform heavy-weight TLB invalidation
> after unmapping queues, unlike HWS which does this automatically. This causes
> a race condition where in-flight SDMA DMA descriptors can access memory that
> has been unmapped, leading to page faults and GPU queue hangs during SVM
> page migration.
>
> The issue manifests as KFDSVMRangeTest.MultiThreadMigrationTest/1 failures
> on gfx1151 (Ryzen AI MAX) with XNACK mode 1 enabled - the GPU compute queue
> hangs with packets submitted but never consumed.
>
> Add kfd_flush_tlb() with TLB_FLUSH_HEAVYWEIGHT in two MES code paths:
> 1. evict_process_queues_cpsch() - after MES removes queues
> 2. suspend_queues() - after MES suspends queues and mem_fence completes
>
> This ensures all in-flight memory accesses from unmapped queues are flushed
> before memory is freed or migrated.
>
> Testing on gfx1151 shows this reduces failure rate from 100% to approximately
> 7-10%. The residual failures require further investigation.
>
> Signed-off-by: Priya Hosur <[email protected]>
> ---
>   .../gpu/drm/amd/amdkfd/kfd_device_queue_manager.c   | 13 ++++++++++++-
>   1 file changed, 12 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> index a23384571193..5eb85290126e 100644
> --- a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> +++ b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> @@ -1450,6 +1450,14 @@ static int evict_process_queues_cpsch(struct device_queue_manager *dqm,
>   		dqm_evict_mqd_bo(dqm, q);
>   	}
>   
> +	/*
> +	 * Heavy-weight TLB flush after MES removes queues to ensure
> +	 * in-flight SDMA accesses complete before memory is freed/migrated.
> +	 * HWS does this automatically, MES does not.

I'm not sure why you call out SDMA specifically here. This affects 
in-flight memory accesses from compute jobs as well. Just remove "SDMA" 
from the comment. With that fixed, the patch is

Reviewed-by: Felix Kuehling <[email protected]>


> +	 */
> +	if (dqm->dev->kfd->shared_resources.enable_mes)
> +		kfd_flush_tlb(pdd, TLB_FLUSH_HEAVYWEIGHT);
> +
>   	if (!dqm->dev->kfd->shared_resources.enable_mes) {
>   		pdd->last_evict_timestamp = get_jiffies_64();
>   		retval = execute_queues_cpsch(dqm,
> @@ -3736,8 +3744,11 @@ int suspend_queues(struct kfd_process *p,
>   		if (!per_device_suspended) {
>   			dqm_unlock(dqm);
>   			mutex_unlock(&p->event_mutex);
> -			if (total_suspended)
> +			if (total_suspended) {
>   				amdgpu_amdkfd_debug_mem_fence(dqm->dev->adev);
> +				/* Heavy-weight TLB flush after MES suspends queues */
> +				kfd_flush_tlb(pdd, TLB_FLUSH_HEAVYWEIGHT);
> +			}
>   			continue;
>   		}
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.