Re: [v9,09/10] drm/xe/pci: Introduce PCIe FLR
"Laguna, Lukasz" <[email protected]>
| Newsgroups | org.freedesktop.lists.intel-xe |
|---|---|
| Message-ID | <[email protected]> |
On 7/1/2026 10:29, Raag Jadav wrote: > With bare minimum pieces in place, we can finally introduce PCIe Function > Level Reset (FLR) support which re-initializes hardware state without the > need for reloading the driver from userspace. All VRAM contents are lost > along with hardware state and driver takes care of recreating the required > kernel bos as part of re-initialization, but user still needs to recreate > user bos and reload context after PCIe FLR. > > Signed-off-by: Raag Jadav <[email protected]> > Tested-by: Lukasz Laguna <[email protected]> > Acked-by: Rodrigo Vivi <[email protected]> > --- > v2: Spell out Function Level Reset (Jani) > v5: Prevent PM ref leak for wedged device (Matthew Brost) > v6: Add PCIe FLR documentation (Daniele) > v7: Refine PCIe FLR documentation (Daniele) > Introduce xe_pci_reset_skip() helper (Lukasz) > v9: Add 'Xe' prefix to document title (Rodrigo) > --- > drivers/gpu/drm/xe/Makefile | 1 + > drivers/gpu/drm/xe/xe_device_types.h | 3 + > drivers/gpu/drm/xe/xe_pci.c | 1 + > drivers/gpu/drm/xe/xe_pci.h | 2 + > drivers/gpu/drm/xe/xe_pci_error.c | 129 +++++++++++++++++++++++++++ > 5 files changed, 136 insertions(+) > create mode 100644 drivers/gpu/drm/xe/xe_pci_error.c > > diff --git a/drivers/gpu/drm/xe/Makefile b/drivers/gpu/drm/xe/Makefile > index 8e7b146880f4..3c001b2a4aec 100644 > --- a/drivers/gpu/drm/xe/Makefile > +++ b/drivers/gpu/drm/xe/Makefile > @@ -101,6 +101,7 @@ xe-y += xe_bb.o \ > xe_page_reclaim.o \ > xe_pat.o \ > xe_pci.o \ > + xe_pci_error.o \ > xe_pci_rebar.o \ > xe_pcode.o \ > xe_pm.o \ > diff --git a/drivers/gpu/drm/xe/xe_device_types.h b/drivers/gpu/drm/xe/xe_device_types.h > index 80fa60821a47..3ccfa8af482c 100644 > --- a/drivers/gpu/drm/xe/xe_device_types.h > +++ b/drivers/gpu/drm/xe/xe_device_types.h > @@ -480,6 +480,9 @@ struct xe_device { > /** @pxp: Encapsulate Protected Xe Path support */ > struct xe_pxp *pxp; > > + /** @flr_prepared: Prepared for function-reset */ > + bool flr_prepared; > + > /** @needs_flr_on_fini: requests function-reset on fini */ > bool needs_flr_on_fini; > > diff --git a/drivers/gpu/drm/xe/xe_pci.c b/drivers/gpu/drm/xe/xe_pci.c > index 03362480e3e0..d11188790eca 100644 > --- a/drivers/gpu/drm/xe/xe_pci.c > +++ b/drivers/gpu/drm/xe/xe_pci.c > @@ -1353,6 +1353,7 @@ static struct pci_driver xe_pci_driver = { > #ifdef CONFIG_PM_SLEEP > .driver.pm = &xe_pm_ops, > #endif > + .err_handler = &xe_pci_error_handlers, > }; > > /** > diff --git a/drivers/gpu/drm/xe/xe_pci.h b/drivers/gpu/drm/xe/xe_pci.h > index 11bcc5fe2c5b..24e51a71a959 100644 > --- a/drivers/gpu/drm/xe/xe_pci.h > +++ b/drivers/gpu/drm/xe/xe_pci.h > @@ -8,6 +8,8 @@ > > struct pci_dev; > > +extern const struct pci_error_handlers xe_pci_error_handlers; > + > int xe_register_pci_driver(void); > void xe_unregister_pci_driver(void); > struct xe_device *xe_pci_to_pf_device(struct pci_dev *pdev); > diff --git a/drivers/gpu/drm/xe/xe_pci_error.c b/drivers/gpu/drm/xe/xe_pci_error.c > new file mode 100644 > index 000000000000..544005af4d3f > --- /dev/null > +++ b/drivers/gpu/drm/xe/xe_pci_error.c > @@ -0,0 +1,129 @@ > +// SPDX-License-Identifier: MIT > +/* > + * Copyright © 2026 Intel Corporation > + */ > + > +#include "xe_device.h" > +#include "xe_printk.h" > +#include "xe_sriov_pf_helpers.h" > + > +/** > + * DOC: Xe PCI Error Handling > + * > + * Xe driver registers PCI callbacks which are called by PCI core in case of > + * bus errors or resets. > + * > + * Currently only Function Level Reset (FLR) callbacks are supported. PCIe FLR This is outdated and needs to be updated. > + * wipes the VRAM and resets the state of all the hardware units. Therefore, the > + * contents of all exec queues and BOs in VRAM are lost and the hardware needs a > + * full re-initialization. The way Xe driver handles it, is pretty much similar > + * to system suspend/resume flow with a few notable exceptions. > + * > + * Prepare phase: > + * > + * - Temporarily wedge the device to prevent userspace access > + * - Kill exec queues which signals all fences and frees in-flight jobs > + * - Stop the scheduler and all submissions to GuC > + * - The fact that FLR is needed is because hardware could be in corrupted state > + * and access unreliable, so skip memory eviction due to untrustworthy VRAM > + * contents > + * - Remove all memory mappings since VRAM contents will be lost > + * > + * Re-initialization phase: > + * > + * - Recreate kernel BOs due to skipped memory eviction in prepare phase > + * - Restore kernel queues which were killed in prepare phase > + * - Reload all uC firmwares > + * - Bring up all hardware units > + * - Unwedge the device to allow userspace access > + * > + * Since VRAM contents are lost, the user is expected to recreate user memory > + * and reload context. > + * > + * TODO: Add PCIe error handling callbacks using similar flow. This is outdated and needs to be updated. > + * > + * Current implementation is only limited to re-initializing GT. This needs to > + * be extended for a lot of components listed below. > + * > + * - Proper re-initialization of GSC and PXP for integrated platforms > + * - SR-IOV cases which need PF and VF synchronization > + * - Re-initialization of all child devices registered by Xe > + * - MM corner cases > + * - Display > + */ > + > +static inline bool xe_pci_reset_skip(struct xe_device *xe) > +{ > + return !IS_DGFX(xe) || IS_SRIOV_VF(xe) || xe_sriov_pf_num_vfs(xe) || xe->info.probe_display; > +} > + > +static void xe_pci_reset_prepare(struct pci_dev *pdev) > +{ > + struct xe_device *xe = pdev_to_xe_device(pdev); > + > + if (xe_pci_reset_skip(xe)) { > + xe_err(xe, "PCIe FLR not supported\n"); > + return; > + } > + > + if (xe_device_wedged(xe)) { > + xe_err(xe, "PCIe FLR aborted, device in unexpected state\n"); > + return; > + } > + > + /* Wedge the device to prevent userspace access but don't send the event yet */ > + xe_device_wedged_get(xe); > + > + /* > + * The hardware could be in corrupted state and access unreliable, but we try to > + * update data structures and cleanup any pending work to avoid side effects during > + * PCIe FLR. This will be similar to system suspend flow but without eviction. > + */ > + if (xe_device_suspend(xe, true)) { > + xe_err(xe, "Failed to prepare for PCIe FLR\n"); > + xe_device_wedged_put(xe); > + return; > + } > + > + xe->flr_prepared = true; > + xe_info(xe, "Prepared for PCIe FLR\n"); > +} > + > +static void xe_pci_reset_done(struct pci_dev *pdev) > +{ > + struct xe_device *xe = pdev_to_xe_device(pdev); > + > + if (xe_pci_reset_skip(xe)) > + goto out; > + > + if (!xe_device_wedged(xe) || !xe->flr_prepared) > + goto out; > + > + /* Unprepare early in case we fail */ > + xe->flr_prepared = false; > + > + /* > + * We already have the data structures intact, so try to re-initialize the device. > + * This will be similar to system resume flow, except we'll also need to recreate > + * kernel bos and restore kernel queues. > + */ > + if (xe_device_resume(xe, true)) { > + xe_err(xe, "Re-initialization failed\n"); > + xe_device_wedged_put(xe); I'm not sure if we should bother, but it introduces a small window where the device is unwedged while still broken. Lukasz > + goto out; > + } > + > + /* Unwedge to allow userspace access */ > + xe_device_wedged_put(xe); > + xe_info(xe, "Re-initialization success\n"); > + > + return; > +out: > + /* Most likely the device is unusable and there's nothing we can do about it */ > + xe_device_declare_wedged(xe); > +} > + > +const struct pci_error_handlers xe_pci_error_handlers = { > + .reset_prepare = xe_pci_reset_prepare, > + .reset_done = xe_pci_reset_done, > +};