Re: [PATCH v4 0/4] drm/tyr: GPU reset infrastructure
Onur Özkan <[email protected]>
| Newsgroups | org.kernel.vger.rust-for-linux,org.freedesktop.lists.dri-devel,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On Sat, 15 Aug 2026 13:23:56 +0300 Onur Özkan <[email protected]> wrote: > Add support for scheduling GPU resets on a dedicated workqueue. Track > the reset state to avoid queueing another reset while one is already > pending or in progress. > > Use an SRCU based gate with mutex-protected reader admission to block > hardware accesses while reset work runs and wait for current users > before resetting. > > Stop new reset requests during teardown and drain any queued or running > reset work before releasing the device resources. > > This is the initial reset infrastructure only. It is not wired to a reset > source yet as those will follow in separate work. > > Based on 'commit 53441a9cae3c ("drm/tyr: program CSF global interface")' > from tyr-for-upstream with the following patch series on the ML: > - rust: workqueue: add cancel_sync support [1] > - rust: add SRCU abstraction [2] > - rust: workqueue: add ScopedQueue for lifetime bound items [3] > - Creation of workqueues in Rust [4] > > TODOs: > - On reset failure, we don't do anything for now. We should unplug > the GPU (that's what Panthor does) but we don't have the infrastructure > for that yet [5]. > - In schedule(), similar to panthor_device_schedule_reset(), we should > have a PM check but similar to the note above, we don't have the > infrastructure for that yet. > > Changes since v4: > - Moved iomem behind the HwGate. > - Renamed the guard names in hw_gate (after the iomem patch, I realized > the old names were confusing). > > Changes since v3: > - Dropped Resettable, ActiveHwState, pre/post-reset hooks and typestate. > - Removed ResetState::Enqueueing and teardown spinning. > - Replaced fail-fast try_access() and epoch tracking with a blocking > hardware-access gate. > - Used b4 for the prerequisites. > > Changes since v2: > - Replaced Work::disable_sync with Work::cancel_sync. > - Used type state pattern for reset stages. > - Using ScopedQueue with device lifetime instead of a static workqueue. > - Removed SRCU patch from this series. > > Changes since v1: > - Removed OrderedQueue and using Alice's workqueue implementation [1] > instead. > - Added Resettable trait with pre_reset and post_reset hooks to be > implemented by reset-managed hardwares. > - Added SRCU abstraction and used it to synchronize the reset work and > hardware access. > > v1: https://lore.kernel.org/rust-for-linux/[email protected] > v2: https://lore.kernel.org/rust-for-linux/[email protected] > v3: https://lore.kernel.org/all/[email protected] > v4: https://lore.kernel.org/all/[email protected] > > Link: https://lore.kernel.org/all/[email protected] [1] > Link: https://lore.kernel.org/all/[email protected] [2] > Link: https://lore.kernel.org/all/[email protected] [3] > Link: https://lore.kernel.org/all/[email protected] [4] > Link: https://gitlab.freedesktop.org/panfrost/linux/-/work_items/29#note_3391826 [5] > Link: https://gitlab.freedesktop.org/panfrost/linux/-/issues/28 > Signed-off-by: Onur Özkan <[email protected]> > --- > Onur Özkan (4): > rust: workqueue: impl Send and Sync for OwnedQueue > drm/tyr: clear stale IRQ state before soft reset > drm/tyr: add GPU reset infrastructure > drm/tyr: put iomem behind the hardware gate > > drivers/gpu/drm/tyr/driver.rs | 61 +++----- > drivers/gpu/drm/tyr/fw.rs | 16 +- > drivers/gpu/drm/tyr/mmu.rs | 9 +- > drivers/gpu/drm/tyr/mmu/address_space.rs | 49 +++--- > drivers/gpu/drm/tyr/reset.rs | 252 +++++++++++++++++++++++++++++++ > drivers/gpu/drm/tyr/reset/hw_gate.rs | 102 +++++++++++++ > drivers/gpu/drm/tyr/tyr.rs | 1 + > rust/kernel/workqueue/mod.rs | 8 + > 8 files changed, 421 insertions(+), 77 deletions(-) > --- > base-commit: 53441a9cae3c4be552fa8aefa6e61b0f1bceb8a5 > change-id: 20260813-tyr-reset-impl-93e951f996b4 > prerequisite-message-id: <[email protected]> > prerequisite-patch-id: f5e24f7b3717f2ab0445b5395c7f12b554861d78 > prerequisite-patch-id: 0bf6ae4abcef7090d0e48921d6ec627a59dc6a7a > prerequisite-patch-id: f2bd17b7ba9f4626ddec56fec8e108872ddc6af1 > prerequisite-message-id: <[email protected]> > prerequisite-patch-id: 5b966d09cfd455dd09fb0911d174b123efa02a16 > prerequisite-message-id: <[email protected]> > prerequisite-patch-id: 9e1efee190d212ba1b01cd0acb5a4357e0b4da42 > prerequisite-patch-id: 26aba035f4d1e212fa6ea7078095febefff5c5ba > prerequisite-patch-id: ffa25d5aadec4c04589af2bdd59a9c022f0fc9b0 > prerequisite-patch-id: c2e05a4ac9d665d331952b291a622df6be673e41 > prerequisite-patch-id: c14a4bd8a68b045356d61f7fa94690bd037ad08b > prerequisite-message-id: <[email protected]> > prerequisite-patch-id: b40e7a218c4e0467933e83c45b09eb78c7ebace2 > The version prefix is wrong on this series. It's not V4; it's V5. I am not sure if I should re-send the whole series with the correct prefix. Sorry, Onur