[PATCH V15 12/14] drm/xe: Add sysfs interface for bad gpu vram pages

Tejas Upadhyay <[email protected]>
Newsgroups org.freedesktop.lists.intel-xe
Message-ID <[email protected]>
Include a sysfs interface designed to expose information about bad
VRAM pages — those identified as having hardware faults (e.g., ECC
errors). This interface allows userspace tools and administrators to
monitor the health of the GPU's local memory and track the status of
page retirement. Details on bad gpu vram pages can be found under
/sys/bus/pci/devices/<bdf>/vram_bad_pages.

The format is: pfn : gpu_page_size : flags

flags:
  R: reserved, this gpu page is reserved.
  P: pending for reserve, this gpu page is marked as bad, will be
     reserved in next window of page_reserve.
  F: unable to reserve, this gpu page can't be reserved due to some
     reasons.

For example, cat /sys/bus/pci/devices/<bdf>/vram_bad_pages:
  max_pages : 10000
  0x0000000000000000 : 0x0000000000001000 : R
  0x0000000000001234 : 0x0000000000001000 : P

The sysfs binary attribute is created under the PCI device kobject
when the platform supports it and the configfs bad_page_reservation
policy is enabled. Uses RCU-protected list traversal so reads never
block normal VRAM allocation operations.

Signed-off-by: Tejas Upadhyay <[email protected]>
---
 drivers/gpu/drm/xe/xe_device_sysfs.c | 7 +++++++
 drivers/gpu/drm/xe/xe_ttm_vram_mgr.h | 1 +
 2 files changed, 8 insertions(+)

diff --git a/drivers/gpu/drm/xe/xe_device_sysfs.c b/drivers/gpu/drm/xe/xe_device_sysfs.c
index a73e0e957cb0..47c5be4180fe 100644
--- a/drivers/gpu/drm/xe/xe_device_sysfs.c
+++ b/drivers/gpu/drm/xe/xe_device_sysfs.c
@@ -8,12 +8,14 @@
 #include <linux/pci.h>
 #include <linux/sysfs.h>
 
+#include "xe_configfs.h"
 #include "xe_device.h"
 #include "xe_device_sysfs.h"
 #include "xe_mmio.h"
 #include "xe_pcode_api.h"
 #include "xe_pcode.h"
 #include "xe_pm.h"
+#include "xe_ttm_vram_mgr.h"
 
 /**
  * DOC: Xe device sysfs
@@ -267,6 +269,7 @@ static const struct attribute_group auto_link_downgrade_attr_group = {
 int xe_device_sysfs_init(struct xe_device *xe)
 {
 	struct device *dev = xe->drm.dev;
+	bool policy;
 	int ret;
 
 	if (xe->d3cold.capable) {
@@ -285,5 +288,9 @@ int xe_device_sysfs_init(struct xe_device *xe)
 			return ret;
 	}
 
+	policy = xe_configfs_get_bad_page_reservation(to_pci_dev(dev));
+	if (xe->info.platform == XE_CRESCENTISLAND && policy)
+		xe_ttm_vram_sysfs_init(xe);
+
 	return 0;
 }
diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h
index d5392beff30c..eb55b0f74ef3 100644
--- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h
+++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h
@@ -32,6 +32,7 @@ void xe_ttm_vram_get_used(struct ttm_resource_manager *man,
 			  u64 *used, u64 *used_visible);
 
 int xe_ttm_vram_handle_addr_fault(struct xe_device *xe, u64 addr);
+int xe_ttm_vram_sysfs_init(struct xe_device *xe);
 static inline struct xe_ttm_vram_mgr_resource *
 to_xe_ttm_vram_mgr_resource(struct ttm_resource *res)
 {
-- 
2.52.0
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.