RE: [PATCH] drm/amdgpu: skip BOs being torn down during GTT recovery
"Zhang, Yifan" <[email protected]> Thu, 30 Jul 2026 03:14:29 +0000
| Newsgroups | org.freedesktop.lists.amd-gfx |
|---|---|
| Message-ID | <CY5PR12MB636940C9F7CFED80EFE97DC4C1C92@CY5PR12MB6369.namprd12.prod.outlook.com> |
Public > +#define GART_ENTRY_BO_COLOR 0 > #define GART_ENTRY_WITHOUT_BO_COLOR 1 > +#define GART_ENTRY_BO_TEARDOWN_COLOR 2 These names are confusing - a GART entry can have three colors - color, without color, and teardown color. Could these names be made more descriptive - especially the first one? Agreed. We can rename them to GART_ENTRY_COLOR_LIVE_BO, GART_ENTRY_COLOR_NO_BO and GART_ENTRY_COLOR_STALE_BO. > kernel: amdgpu 0000:a3:00.0: GPU reset succeeded, trying to resume > kernel: BUG: kernel NULL pointer dereference, address: > 0000000000000000 > kernel: #PF: supervisor read access in kernel mode > kernel: #PF: error_code(0x0000) - not-present page > kernel: PGD 62256dd067 P4D 0 Could this be summarized instead of having the entire kernel dump in the commit message? I will trim it down, but since the log shows the race scenario, I would like to keep part of it in the commit description. -----Original Message----- From: Francis, David <[email protected]> Sent: Tuesday, July 28, 2026 10:20 PM To: Zhang, Yifan <[email protected]>; [email protected] Cc: Deucher, Alexander <[email protected]>; Koenig, Christian <[email protected]>; Yuan, Perry <[email protected]> Subject: Re: [PATCH] drm/amdgpu: skip BOs being torn down during GTT recovery Design seems reasonable, some minor comments: > kernel: amdgpu 0000:a3:00.0: GPU reset succeeded, trying to resume > kernel: BUG: kernel NULL pointer dereference, address: > 0000000000000000 > kernel: #PF: supervisor read access in kernel mode > kernel: #PF: error_code(0x0000) - not-present page > kernel: PGD 62256dd067 P4D 0 Could this be summarized instead of having the entire kernel dump in the commit message? > +#define GART_ENTRY_BO_COLOR 0 > #define GART_ENTRY_WITHOUT_BO_COLOR 1 > +#define GART_ENTRY_BO_TEARDOWN_COLOR 2 These names are confusing - a GART entry can have three colors - color, without color, and teardown color. Could these names be made more descriptive - especially the first one? ________________________________________ From: amd-gfx <[email protected]> on behalf of Yifan Zhang <[email protected]> Sent: Tuesday, July 28, 2026 1:40 AM To: [email protected] Cc: Deucher, Alexander; Koenig, Christian; Yuan, Perry; Zhang, Yifan Subject: [PATCH] drm/amdgpu: skip BOs being torn down during GTT recovery A GPU reset can race with BO teardown after the BO's GTT resource has been marked for deletion but before its drm_mm node is removed. In this window, amdgpu_gtt_mgr_recover() can treat the node as a live BO and try to restore its GART mapping while its TT backing is being destroyed. Mark the GTT node with a dedicated teardown color from amdgpu_bo_delete_mem_notify(). During recovery, process only nodes colored as live BOs, leaving both teardown nodes and nodes without BOs untouched. This prevents reset recovery from accessing a BO whose backing storage is no longer valid. kernel: amdgpu 0000:a3:00.0: GPU reset succeeded, trying to resume kernel: BUG: kernel NULL pointer dereference, address: 0000000000000000 kernel: #PF: supervisor read access in kernel mode kernel: #PF: error_code(0x0000) - not-present page kernel: PGD 62256dd067 P4D 0 kernel: Oops: Oops: 0000 [#1] SMP NOPTI kernel: CPU: 378 UID: 0 PID: 3775143 Comm: kworker/u1536:4 Tainted: G W OE 6.17.0-35-generic #35~24.04.1-Ubuntu PREEMPT(voluntary) kernel: Tainted: [W]=WARN, [O]=OOT_MODULE, [E]=UNSIGNED_MODULE kernel: Hardware name: Supermicro AS -4125GS-TNRT/H13DSG-O-CPU, BIOS 3.8a 11/01/2025 kernel: Workqueue: amdgpu-reset-dev amdgpu_debugfs_reset_work [amdgpu] kernel: RIP: 0010:amdgpu_gart_map+0x60/0xc0 [amdgpu] kernel: Code: 71 98 c3 48 89 45 d0 31 c0 c7 45 cc 00 00 00 00 e8 25 4b 9e c1 84 c0 74 39 48 89 d8 48 c1 e8 0c 89 c3 45 85 e4 74 23 41 01 c4 <49> 8b 0e 4c 8b 45 c0 89 da 4c 89 fe 4c 89 ef 83 c3 01 49 83 c6 08 kernel: RSP: 0018:ff4f7e3af8a1fb18 EFLAGS: 00010206 kernel: RAX: 00000000000002b6 RBX: 00000000000002b6 RCX: 0000000000000000 kernel: RDX: 0000000000000001 RSI: 0000000000000000 RDI: 0000000000000000 kernel: RBP: ff4f7e3af8a1fb58 R08: 80c0000000000073 R09: ff4f7e4ad6b00000 kernel: R10: 0000000000000000 R11: 0000000000000000 R12: 00000000000002b7 kernel: R13: ff2209225a000000 R14: 0000000000000000 R15: ff4f7e4ad6b00000 kernel: FS: 0000000000000000(0000) GS:ff22098108562000(0000) knlGS:0000000000000000 kernel: CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 kernel: CR2: 0000000000000000 CR3: 00000064b7011002 CR4: 0000000000f71ef0 kernel: PKRU: 55555554 kernel: Call Trace: kernel: <TASK> kernel: amdgpu_gart_bind+0x1a/0x50 [amdgpu] kernel: amdgpu_ttm_gart_bind+0xd8/0xe0 [amdgpu] kernel: amdgpu_ttm_recover_gart+0x5f/0x80 [amdgpu] kernel: amdgpu_gtt_mgr_recover+0x43/0x70 [amdgpu] kernel: gmc_v12_0_hw_init+0x46/0x160 [amdgpu] kernel: gmc_v12_0_resume+0x14/0x40 [amdgpu] kernel: amdgpu_ip_block_resume+0x24/0x80 [amdgpu] kernel: amdgpu_device_ip_resume_phase1+0xce/0x180 [amdgpu] kernel: amdgpu_device_reinit_after_reset+0x13d/0x340 [amdgpu] kernel: amdgpu_do_asic_reset.part.0+0x56/0x1b0 [amdgpu] kernel: amdgpu_device_asic_reset+0x426/0x640 [amdgpu] kernel: ? srso_alias_return_thunk+0x5/0xfbef5 kernel: amdgpu_device_gpu_recover+0x260/0x410 [amdgpu] kernel: amdgpu_debugfs_reset_work+0x68/0x90 [amdgpu] kernel: process_one_work+0x18e/0x3e0 kernel: worker_thread+0x2e3/0x420 kernel: ? _raw_spin_lock_irqsave+0xe/0x20 kernel: ? srso_alias_return_thunk+0x5/0xfbef5 kernel: ? __pfx_worker_thread+0x10/0x10 kernel: kthread+0x10a/0x230 kernel: ? __pfx_kthread+0x10/0x10 kernel: ret_from_fork+0x121/0x140 kernel: ? __pfx_kthread+0x10/0x10 kernel: ret_from_fork_asm+0x1a/0x30 kernel: </TASK> Signed-off-by: Yifan Zhang <[email protected]> --- drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c | 36 +++++++++++++++++++-- drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c | 1 + drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h | 1 + 3 files changed, 36 insertions(+), 2 deletions(-) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c index 0ea32561c4bc..1f9063b30526 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gtt_mgr.c @@ -26,7 +26,9 @@ #include "amdgpu.h" +#define GART_ENTRY_BO_COLOR 0 #define GART_ENTRY_WITHOUT_BO_COLOR 1 +#define GART_ENTRY_BO_TEARDOWN_COLOR 2 static inline struct amdgpu_gtt_mgr * to_gtt_mgr(struct ttm_resource_manager *man) @@ -102,6 +104,35 @@ bool amdgpu_gtt_mgr_has_gart_addr(struct ttm_resource *res) return drm_mm_node_allocated(&node->mm_nodes[0]); } +/** + * amdgpu_gtt_mgr_mark_bo_teardown - exclude a BO from GART recovery + * + * @tbo: TTM BO whose TT backing is about to be destroyed + * + * Keep the GART range allocated until the resource is freed, but +prevent + * reset recovery from using the BO after TT teardown has started. + */ +void amdgpu_gtt_mgr_mark_bo_teardown(struct ttm_buffer_object *tbo) { + struct amdgpu_device *adev = amdgpu_ttm_adev(tbo->bdev); + struct ttm_resource *res = tbo->resource; + struct ttm_range_mgr_node *node; + struct amdgpu_gtt_mgr *mgr; + + dma_resv_assert_held(tbo->base.resv); + + if (!res || res->mem_type != TTM_PL_TT) + return; + + node = to_ttm_range_mgr_node(res); + mgr = &adev->mman.gtt_mgr; + + spin_lock(&mgr->lock); + if (drm_mm_node_allocated(&node->mm_nodes[0])) + node->mm_nodes[0].color = GART_ENTRY_BO_TEARDOWN_COLOR; + spin_unlock(&mgr->lock); +} + /** * amdgpu_gtt_mgr_new - allocate a new node * @@ -137,7 +168,8 @@ static int amdgpu_gtt_mgr_new(struct ttm_resource_manager *man, spin_lock(&mgr->lock); r = drm_mm_insert_node_in_range(&mgr->mm, &node->mm_nodes[0], num_pages, tbo->page_alignment, - 0, place->fpfn, place->lpfn, + GART_ENTRY_BO_COLOR, + place->fpfn, + place->lpfn, DRM_MM_INSERT_BEST); spin_unlock(&mgr->lock); if (unlikely(r)) @@ -248,7 +280,7 @@ void amdgpu_gtt_mgr_recover(struct amdgpu_gtt_mgr *mgr) adev = container_of(mgr, typeof(*adev), mman.gtt_mgr); spin_lock(&mgr->lock); drm_mm_for_each_node(mm_node, &mgr->mm) { - if (mm_node->color == GART_ENTRY_WITHOUT_BO_COLOR) + if (mm_node->color != GART_ENTRY_BO_COLOR) continue; node = container_of(mm_node, typeof(*node), mm_nodes[0]); diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c index 5fe29e4972d8..6fa8301b8ddc 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c @@ -1676,6 +1676,7 @@ static int amdgpu_ttm_access_memory(struct ttm_buffer_object *bo, static void amdgpu_bo_delete_mem_notify(struct ttm_buffer_object *bo) { + amdgpu_gtt_mgr_mark_bo_teardown(bo); amdgpu_bo_move_notify(bo, false, NULL); } diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h index ff9e2e346609..af1e7fcc7175 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.h @@ -145,6 +145,7 @@ int amdgpu_vram_mgr_init(struct amdgpu_device *adev); void amdgpu_vram_mgr_fini(struct amdgpu_device *adev); bool amdgpu_gtt_mgr_has_gart_addr(struct ttm_resource *mem); +void amdgpu_gtt_mgr_mark_bo_teardown(struct ttm_buffer_object *tbo); void amdgpu_gtt_mgr_recover(struct amdgpu_gtt_mgr *mgr); int amdgpu_gtt_mgr_alloc_entries(struct amdgpu_gtt_mgr *mgr, -- 2.43.0