Re: [PATCH 4/4] drm/xe/ras: Report CSC errors using SIGID
Umesh Nerlige Ramappa <[email protected]>
| Newsgroups | org.freedesktop.lists.intel-xe |
|---|---|
| Message-ID | <[email protected]> |
On Sat, Aug 15, 2026 at 09:59:24AM +0200, Michal Wajdeczko wrote: > > >On 8/14/2026 8:27 PM, Umesh Nerlige Ramappa wrote: >> On Wed, Aug 12, 2026 at 09:52:22PM +0200, Michal Wajdeczko wrote: >>> >>> >>> On 8/12/2026 1:52 AM, Umesh Nerlige Ramappa wrote: >>>> Use xe_log_err() to report CSC errors using SIGID. >>>> >>>> Signed-off-by: Umesh Nerlige Ramappa <[email protected]> >>>> --- >>>> drivers/gpu/drm/xe/xe_ras.c | 5 ++--- >>>> 1 file changed, 2 insertions(+), 3 deletions(-) >>>> >>>> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c >>>> index d1aa3794f2e0..e50ab49bce5a 100644 >>>> --- a/drivers/gpu/drm/xe/xe_ras.c >>>> +++ b/drivers/gpu/drm/xe/xe_ras.c >>>> @@ -283,9 +283,8 @@ static u8 handle_soc_internal_errors(struct xe_device *xe, struct xe_ras_error_a >>>> * is required. >>>> */ >>>> if (csc_error->hec_fw_error) { >>>> - xe_err(xe, "[RAS]: CSC %s detected: 0x%x\n", >>>> - sev_to_str(counter->common.severity), >>>> - csc_error->hec_fw_error); >>>> + xe_log_err(xe, SOC, 0, "[RAS]: CSC %s detected: 0x%x\n", >>>> + sev_to_str(counter->common.severity), csc_error->hec_fw_error); >>> >>> hmm, it was assumed that for HW errors we will use 12 byte data >>> while xe_log_err is mostly for the SW errors where we use errno >> >> Not sure I understand what the 12 bytes data is. > >IIUC it would be struct xe_ras_error_class > >> >>> >>> do we still need to use "RAS" prefix ? >>> >>> maybe we should add component XE_LOG_COMPONENT_CSC with SIGID_SOC_INTERNAL ? >>> >>> shouldn't we use common.severity to select right xe_log/CPER severity ? >> >> Not sure how to use this macro to do that though. I thought severity was automatically handled in the backend. > >if you plan to report errno-like SW errors from the CSC FW, >then IMO we should add new CSC component in abi/xe_log_abi.h >with XE_SIGID_DEVICE_FW: > >+ define(DRIVER_FIRMWARE, 18, CSC, DEVICE_FW, "CSC") \ > >then you can use > > xe_log_err(xe, CSC, err, "detected: %#x\n", csc_error->hec_fw_error); > xe_log_err_fatal(xe, CSC, err, "detected: %#x\n", csc_error->hec_fw_error); > >or if you can match counter severity with CPER severity: > > cper_sev = to_cper_sev(counter->common.severity); > > xe_log_comp(xe, cper_sev, CSC, ERR_PTR(err), 0, "detected: %#x\n", csc_error->hec_fw_error); > >but if that error is representing HW SIGID - SOC_INTERNAL >then I guess it is expected that xe_ras_error_class should >be passed to xe_log() using: > > struct xe_ras_error_class data; > > xe_log_comp(xe, cper_sev, SOC_INTERNAL, &data, sizeof(data), "...\n"); It's just going to be a CPER_SEV_FATAL since it triggers survivability. I think I am leaning into the xe_log_comp (SOC_INTERNAL) recommendation here since it also matches Aravind/Mallesh's recommendation offline. I will post that as the next rev. Thanks, Umesh > >to allow proper CPER record generation > >> >> Thanks, >> Umesh >>> >>>> xe_survivability_mode_runtime_enable(xe); >>>> return XE_RAS_RECOVERY_ACTION_DISCONNECT; >>>> } >>>