Re: [PATCH] drm/amdgpu: clamp the isolation index for rings outside a partition

Alex Deucher <[email protected]>
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <CADnq5_M_8azMxrWBNukehd4ytrXmM_JVN5crHq7_sTSsUSrmzw@mail.gmail.com>
On Fri, Aug 21, 2026 at 6:20 AM Xiang Liu <[email protected]> wrote:
>
> adev->isolation[] has one slot per partition, but a ring that is not
> assigned to one keeps AMDGPU_XCP_NO_PARTITION, which is ~0, so indexing
> the array with it is out of bounds. SDMA submissions hit this on both
> the isolation enforcement and the VM flush path and trip UBSAN.
>
> Fall back to the first slot the way the cleaner shader path already
> does, and stop taking the address before the ring type check that makes
> it relevant.
>
> Signed-off-by: Xiang Liu <[email protected]>

Reviewed-by: Alex Deucher <[email protected]>

> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 5 ++++-
>  drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c     | 4 +++-
>  2 files changed, 7 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> index baadb676f76b..d8d182599fc6 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> @@ -6756,8 +6756,8 @@ struct dma_fence *amdgpu_device_enforce_isolation(struct amdgpu_device *adev,
>                                                   struct amdgpu_ring *ring,
>                                                   struct amdgpu_job *job)
>  {
> -       struct amdgpu_isolation *isolation = &adev->isolation[ring->xcp_id];
>         struct drm_sched_fence *f = job->base.s_fence;
> +       struct amdgpu_isolation *isolation;
>         struct dma_fence *dep;
>         void *owner;
>         int r;
> @@ -6770,6 +6770,9 @@ struct dma_fence *amdgpu_device_enforce_isolation(struct amdgpu_device *adev,
>             ring->funcs->type != AMDGPU_RING_TYPE_COMPUTE)
>                 return NULL;
>
> +       isolation = &adev->isolation[ring->xcp_id == AMDGPU_XCP_NO_PARTITION ?
> +                                    0 : ring->xcp_id];
> +
>         /*
>          * All submissions where enforce isolation is false are handled as if
>          * they come from a single client. Use ~0l as the owner to distinct it
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c
> index 5a6f5151c3ca..743d307c71fd 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c
> @@ -774,7 +774,9 @@ void amdgpu_vm_flush(struct amdgpu_ring *ring, struct amdgpu_job *job,
>                      bool *emit_gds_needed)
>  {
>         struct amdgpu_device *adev = ring->adev;
> -       struct amdgpu_isolation *isolation = &adev->isolation[ring->xcp_id];
> +       struct amdgpu_isolation *isolation =
> +               &adev->isolation[ring->xcp_id == AMDGPU_XCP_NO_PARTITION ?
> +                                0 : ring->xcp_id];
>         unsigned vmhub = ring->vm_hub;
>         struct amdgpu_vmid_mgr *id_mgr = &adev->vm_manager.id_mgr[vmhub];
>         struct amdgpu_vmid *id = &id_mgr->ids[job->vmid];
> --
> 2.34.1
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.