RE: [PATCH] drm/amdgpu: skip BOs being torn down during GTT recovery

"Zhang, Yifan" <[email protected]> Thu, 30 Jul 2026 03:14:29 +0000
Newsgroups org.freedesktop.lists.amd-gfx
Message-ID <CY5PR12MB636940C9F7CFED80EFE97DC4C1C92@CY5PR12MB6369.namprd12.prod.outlook.com>
Public

> +#define GART_ENTRY_BO_COLOR            0
> #define GART_ENTRY_WITHOUT_BO_COLOR     1
> +#define GART_ENTRY_BO_TEARDOWN_COLOR   2
These names are confusing - a GART entry can have three colors
- color, without color, and teardown color. Could these names be made more descriptive - especially the first one?

Agreed. We can rename them to GART_ENTRY_COLOR_LIVE_BO, GART_ENTRY_COLOR_NO_BO and GART_ENTRY_COLOR_STALE_BO.

> kernel: amdgpu 0000:a3:00.0: GPU reset succeeded, trying to resume
> kernel: BUG: kernel NULL pointer dereference, address:
> 0000000000000000
> kernel: #PF: supervisor read access in kernel mode
> kernel: #PF: error_code(0x0000) - not-present page
> kernel: PGD 62256dd067 P4D 0
Could this be summarized instead of having the entire kernel dump in the commit message?

I will trim it down, but since the log shows the race scenario, I would like to keep part of it in the commit description.

-----Original Message-----
From: Francis, David <[email protected]>
Sent: Tuesday, July 28, 2026 10:20 PM
To: Zhang, Yifan <[email protected]>; [email protected]
Cc: Deucher, Alexander <[email protected]>; Koenig, Christian <[email protected]>; Yuan, Perry <[email protected]>
Subject: Re: [PATCH] drm/amdgpu: skip BOs being torn down during GTT recovery

Design seems reasonable, some minor comments:

> kernel: amdgpu 0000:a3:00.0: GPU reset succeeded, trying to resume
> kernel: BUG: kernel NULL pointer dereference, address:
> 0000000000000000
> kernel: #PF: supervisor read access in kernel mode
> kernel: #PF: error_code(0x0000) - not-present page
> kernel: PGD 62256dd067 P4D 0
Could this be summarized instead of having the entire kernel dump in the commit message?

> +#define GART_ENTRY_BO_COLOR            0
> #define GART_ENTRY_WITHOUT_BO_COLOR     1
> +#define GART_ENTRY_BO_TEARDOWN_COLOR   2
These names are confusing - a GART entry can have three colors
- color, without color, and teardown color. Could these names be made more descriptive - especially the first one?

________________________________________
From: amd-gfx <[email protected]> on behalf of Yifan Zhang <[email protected]>
Sent: Tuesday, July 28, 2026 1:40 AM
To: [email protected]
Cc: Deucher, Alexander; Koenig, Christian; Yuan, Perry; Zhang, Yifan
Subject: [PATCH] drm/amdgpu: skip BOs being torn down during GTT recovery

A GPU reset can race with BO teardown after the BO's GTT resource has been marked for deletion but before its drm_mm node is removed. In this window, amdgpu_gtt_mgr_recover() can treat the node as a live BO and try to restore its GART mapping while its TT backing is being destroyed.

Mark the GTT node with a dedicated teardown color from amdgpu_bo_delete_mem_notify(). During recovery, process only nodes colored as live BOs, leaving both teardown nodes and nodes without BOs untouched.

This prevents reset recovery from accessing a BO whose backing storage is no longer valid.

kernel: amdgpu 0000:a3:00.0: GPU reset succeeded, trying to resume
kernel: BUG: kernel NULL pointer dereference, address: 0000000000000000
kernel: #PF: supervisor read access in kernel mode
kernel: #PF: error_code(0x0000) - not-present page
kernel: PGD 62256dd067 P4D 0
kernel: Oops: Oops: 0000 [#1] SMP NOPTI
kernel: CPU: 378 UID: 0 PID: 3775143 Comm: kworker/u1536:4 Tainted: G        W  OE       6.17.0-35-generic #35~24.04.1-Ubuntu PREEMPT(voluntary)
kernel: Tainted: [W]=WARN, [O]=OOT_MODULE, [E]=UNSIGNED_MODULE
kernel: Hardware name: Supermicro AS -4125GS-TNRT/H13DSG-O-CPU, BIOS 3.8a 11/01/2025
kernel: Workqueue: amdgpu-reset-dev amdgpu_debugfs_reset_work [amdgpu]
kernel: RIP: 0010:amdgpu_gart_map+0x60/0xc0 [amdgpu]
kernel: Code: 71 98 c3 48 89 45 d0 31 c0 c7 45 cc 00 00 00 00 e8 25 4b 9e c1 84 c0 74 39 48 89 d8 48 c1 e8 0c 89 c3 45 85 e4 74 23 41 01 c4 <49> 8b 0e 4c 8b 45 c0 89 da 4c 89 fe 4c 89 ef 83 c3 01 49 83 c6 08
kernel: RSP: 0018:ff4f7e3af8a1fb18 EFLAGS: 00010206
kernel: RAX: 00000000000002b6 RBX: 00000000000002b6 RCX: 0000000000000000
kernel: RDX: 0000000000000001 RSI: 0000000000000000 RDI: 0000000000000000
kernel: RBP: ff4f7e3af8a1fb58 R08: 80c0000000000073 R09: ff4f7e4ad6b00000
kernel: R10: 0000000000000000 R11: 0000000000000000 R12: 00000000000002b7
kernel: R13: ff2209225a000000 R14: 0000000000000000 R15: ff4f7e4ad6b00000
kernel: FS:  0000000000000000(0000) GS:ff22098108562000(0000) knlGS:0000000000000000
kernel: CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
kernel: CR2: 0000000000000000 CR3: 00000064b7011002 CR4: 0000000000f71ef0
kernel: PKRU: 55555554
kernel: Call Trace:
kernel:  <TASK>
kernel:  amdgpu_gart_bind+0x1a/0x50 [amdgpu]
kernel:  amdgpu_ttm_gart_bind+0xd8/0xe0 [amdgpu]
kernel:  amdgpu_ttm_recover_gart+0x5f/0x80 [amdgpu]
kernel:  amdgpu_gtt_mgr_recover+0x43/0x70 [amdgpu]
kernel:  gmc_v12_0_hw_init+0x46/0x160 [amdgpu]
kernel:  gmc_v12_0_resume+0x14/0x40 [amdgpu]
kernel:  amdgpu_ip_block_resume+0x24/0x80 [amdgpu]
kernel:  amdgpu_device_ip_resume_phase1+0xce/0x180 [amdgpu]
kernel:  amdgpu_device_reinit_after_reset+0x13d/0x340 [amdgpu]
kernel:  amdgpu_do_asic_reset.part.0+0x56/0x1b0 [amdgpu]
kernel:  amdgpu_device_asic_reset+0x426/0x640 [amdgpu]
kernel:  ? srso_alias_return_thunk+0x5/0xfbef5
kernel:  amdgpu_device_gpu_recover+0x260/0x410 [amdgpu]
kernel:  amdgpu_debugfs_reset_work+0x68/0x90 [amdgpu]
kernel:  process_one_work+0x18e/0x3e0
kernel:  worker_thread+0x2e3/0x420
kernel:  ? _raw_spin_lock_irqsave+0xe/0x20
kernel:  ? srso_alias_return_thunk+0x5/0xfbef5
kernel:  ? __pfx_worker_thread+0x10/0x10
kernel:  kthread+0x10a/0x230
kernel:  ? __pfx_kthread+0x10/0x10
kernel:  ret_from_fork+0x121/0x140
kernel:  ? __pfx_kthread+0x10/0x10
kernel:  ret_from_fork_asm+0x1a/0x30
kernel:  </TASK>

Signed-off-by: Yifan Zhang <[email protected]>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c | 36 +++++++++++++++++++--
 drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c     |  1 +
 drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h     |  1 +
 3 files changed, 36 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c
index 0ea32561c4bc..1f9063b30526 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c
@@ -26,7 +26,9 @@

 #include "amdgpu.h"

+#define GART_ENTRY_BO_COLOR            0
 #define GART_ENTRY_WITHOUT_BO_COLOR    1
+#define GART_ENTRY_BO_TEARDOWN_COLOR   2

 static inline struct amdgpu_gtt_mgr *
 to_gtt_mgr(struct ttm_resource_manager *man) @@ -102,6 +104,35 @@ bool amdgpu_gtt_mgr_has_gart_addr(struct ttm_resource *res)
        return drm_mm_node_allocated(&node->mm_nodes[0]);
 }

+/**
+ * amdgpu_gtt_mgr_mark_bo_teardown - exclude a BO from GART recovery
+ *
+ * @tbo: TTM BO whose TT backing is about to be destroyed
+ *
+ * Keep the GART range allocated until the resource is freed, but
+prevent
+ * reset recovery from using the BO after TT teardown has started.
+ */
+void amdgpu_gtt_mgr_mark_bo_teardown(struct ttm_buffer_object *tbo) {
+       struct amdgpu_device *adev = amdgpu_ttm_adev(tbo->bdev);
+       struct ttm_resource *res = tbo->resource;
+       struct ttm_range_mgr_node *node;
+       struct amdgpu_gtt_mgr *mgr;
+
+       dma_resv_assert_held(tbo->base.resv);
+
+       if (!res || res->mem_type != TTM_PL_TT)
+               return;
+
+       node = to_ttm_range_mgr_node(res);
+       mgr = &adev->mman.gtt_mgr;
+
+       spin_lock(&mgr->lock);
+       if (drm_mm_node_allocated(&node->mm_nodes[0]))
+               node->mm_nodes[0].color = GART_ENTRY_BO_TEARDOWN_COLOR;
+       spin_unlock(&mgr->lock);
+}
+
 /**
  * amdgpu_gtt_mgr_new - allocate a new node
  *
@@ -137,7 +168,8 @@ static int amdgpu_gtt_mgr_new(struct ttm_resource_manager *man,
                spin_lock(&mgr->lock);
                r = drm_mm_insert_node_in_range(&mgr->mm, &node->mm_nodes[0],
                                                num_pages, tbo->page_alignment,
-                                               0, place->fpfn, place->lpfn,
+                                               GART_ENTRY_BO_COLOR,
+                                               place->fpfn,
+ place->lpfn,
                                                DRM_MM_INSERT_BEST);
                spin_unlock(&mgr->lock);
                if (unlikely(r))
@@ -248,7 +280,7 @@ void amdgpu_gtt_mgr_recover(struct amdgpu_gtt_mgr *mgr)
        adev = container_of(mgr, typeof(*adev), mman.gtt_mgr);
        spin_lock(&mgr->lock);
        drm_mm_for_each_node(mm_node, &mgr->mm) {
-               if (mm_node->color == GART_ENTRY_WITHOUT_BO_COLOR)
+               if (mm_node->color != GART_ENTRY_BO_COLOR)
                        continue;

                node = container_of(mm_node, typeof(*node), mm_nodes[0]); diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
index 5fe29e4972d8..6fa8301b8ddc 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
@@ -1676,6 +1676,7 @@ static int amdgpu_ttm_access_memory(struct ttm_buffer_object *bo,  static void  amdgpu_bo_delete_mem_notify(struct ttm_buffer_object *bo)  {
+       amdgpu_gtt_mgr_mark_bo_teardown(bo);
        amdgpu_bo_move_notify(bo, false, NULL);  }

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h
index ff9e2e346609..af1e7fcc7175 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h
@@ -145,6 +145,7 @@ int amdgpu_vram_mgr_init(struct amdgpu_device *adev);  void amdgpu_vram_mgr_fini(struct amdgpu_device *adev);

 bool amdgpu_gtt_mgr_has_gart_addr(struct ttm_resource *mem);
+void amdgpu_gtt_mgr_mark_bo_teardown(struct ttm_buffer_object *tbo);
 void amdgpu_gtt_mgr_recover(struct amdgpu_gtt_mgr *mgr);

 int amdgpu_gtt_mgr_alloc_entries(struct amdgpu_gtt_mgr *mgr,
--
2.43.0