Re: [PATCH V15 12/14] drm/xe: Add sysfs interface for bad gpu vram pages
"Ghimiray, Himal Prasad" <[email protected]>
| Newsgroups | org.freedesktop.lists.intel-xe |
|---|---|
| Message-ID | <[email protected]> |
On 11-08-2026 18:10, Tejas Upadhyay wrote: > Include a sysfs interface designed to expose information about bad > VRAM pages — those identified as having hardware faults (e.g., ECC > errors). This interface allows userspace tools and administrators to > monitor the health of the GPU's local memory and track the status of > page retirement. Details on bad gpu vram pages can be found under > /sys/bus/pci/devices/<bdf>/vram_bad_pages. > > The format is: pfn : gpu_page_size : flags > > flags: > R: reserved, this gpu page is reserved. > P: pending for reserve, this gpu page is marked as bad, will be > reserved in next window of page_reserve. > F: unable to reserve, this gpu page can't be reserved due to some > reasons. > > For example, cat /sys/bus/pci/devices/<bdf>/vram_bad_pages: > max_pages : 10000 > 0x0000000000000000 : 0x0000000000001000 : R > 0x0000000000001234 : 0x0000000000001000 : P > > The sysfs binary attribute is created under the PCI device kobject > when the platform supports it and the configfs bad_page_reservation > policy is enabled. Uses RCU-protected list traversal so reads never > block normal VRAM allocation operations. > > Signed-off-by: Tejas Upadhyay <[email protected]> > --- > drivers/gpu/drm/xe/xe_device_sysfs.c | 7 +++++++ > drivers/gpu/drm/xe/xe_ttm_vram_mgr.h | 1 + > 2 files changed, 8 insertions(+) > > diff --git a/drivers/gpu/drm/xe/xe_device_sysfs.c b/drivers/gpu/drm/xe/xe_device_sysfs.c > index a73e0e957cb0..47c5be4180fe 100644 > --- a/drivers/gpu/drm/xe/xe_device_sysfs.c > +++ b/drivers/gpu/drm/xe/xe_device_sysfs.c > @@ -8,12 +8,14 @@ > #include <linux/pci.h> > #include <linux/sysfs.h> > > +#include "xe_configfs.h" > #include "xe_device.h" > #include "xe_device_sysfs.h" > #include "xe_mmio.h" > #include "xe_pcode_api.h" > #include "xe_pcode.h" > #include "xe_pm.h" > +#include "xe_ttm_vram_mgr.h" > > /** > * DOC: Xe device sysfs > @@ -267,6 +269,7 @@ static const struct attribute_group auto_link_downgrade_attr_group = { > int xe_device_sysfs_init(struct xe_device *xe) > { > struct device *dev = xe->drm.dev; > + bool policy; > int ret; > > if (xe->d3cold.capable) { > @@ -285,5 +288,9 @@ int xe_device_sysfs_init(struct xe_device *xe) > return ret; > } > > + policy = xe_configfs_get_bad_page_reservation(to_pci_dev(dev)); > + if (xe->info.platform == XE_CRESCENTISLAND && policy) > + xe_ttm_vram_sysfs_init(xe); > + > return 0; > } > diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h > index d5392beff30c..eb55b0f74ef3 100644 > --- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h > +++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h > @@ -32,6 +32,7 @@ void xe_ttm_vram_get_used(struct ttm_resource_manager *man, > u64 *used, u64 *used_visible); > > int xe_ttm_vram_handle_addr_fault(struct xe_device *xe, u64 addr); > +int xe_ttm_vram_sysfs_init(struct xe_device *xe); Move implementation to this patch. > static inline struct xe_ttm_vram_mgr_resource * > to_xe_ttm_vram_mgr_resource(struct ttm_resource *res) > {