Re: [PATCH 4/4] drm/xe/ras: Report CSC errors using SIGID

Umesh Nerlige Ramappa <[email protected]>
Newsgroups org.freedesktop.lists.intel-xe
Message-ID <[email protected]>
On Sat, Aug 15, 2026 at 09:59:24AM +0200, Michal Wajdeczko wrote:
>
>
>On 8/14/2026 8:27 PM, Umesh Nerlige Ramappa wrote:
>> On Wed, Aug 12, 2026 at 09:52:22PM +0200, Michal Wajdeczko wrote:
>>>
>>>
>>> On 8/12/2026 1:52 AM, Umesh Nerlige Ramappa wrote:
>>>> Use xe_log_err() to report CSC errors using SIGID.
>>>>
>>>> Signed-off-by: Umesh Nerlige Ramappa <[email protected]>
>>>> ---
>>>>  drivers/gpu/drm/xe/xe_ras.c | 5 ++---
>>>>  1 file changed, 2 insertions(+), 3 deletions(-)
>>>>
>>>> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
>>>> index d1aa3794f2e0..e50ab49bce5a 100644
>>>> --- a/drivers/gpu/drm/xe/xe_ras.c
>>>> +++ b/drivers/gpu/drm/xe/xe_ras.c
>>>> @@ -283,9 +283,8 @@ static u8 handle_soc_internal_errors(struct xe_device *xe, struct xe_ras_error_a
>>>>           * is required.
>>>>           */
>>>>          if (csc_error->hec_fw_error) {
>>>> -            xe_err(xe, "[RAS]: CSC %s detected: 0x%x\n",
>>>> -                   sev_to_str(counter->common.severity),
>>>> -                   csc_error->hec_fw_error);
>>>> +            xe_log_err(xe, SOC, 0, "[RAS]: CSC %s detected: 0x%x\n",
>>>> +                   sev_to_str(counter->common.severity), csc_error->hec_fw_error);
>>>
>>> hmm, it was assumed that for HW errors we will use 12 byte data
>>> while xe_log_err is mostly for the SW errors where we use errno
>>
>> Not sure I understand what the 12 bytes data is.
>
>IIUC it would be struct xe_ras_error_class
>
>>
>>>
>>> do we still need to use "RAS" prefix ?
>>>
>>> maybe we should add component XE_LOG_COMPONENT_CSC with SIGID_SOC_INTERNAL ?
>>>
>>> shouldn't we use common.severity to select right xe_log/CPER severity ?
>>
>> Not sure how to use this macro to do that though. I thought severity was automatically handled in the backend.
>
>if you plan to report errno-like SW errors from the CSC FW,
>then IMO we should add new CSC component in abi/xe_log_abi.h
>with XE_SIGID_DEVICE_FW:
>
>+	define(DRIVER_FIRMWARE, 18, CSC, DEVICE_FW, "CSC")			\
>
>then you can use
>
>	xe_log_err(xe, CSC, err, "detected: %#x\n", csc_error->hec_fw_error);
>	xe_log_err_fatal(xe, CSC, err, "detected: %#x\n", csc_error->hec_fw_error);
>
>or if you can match counter severity with CPER severity:
>
>	cper_sev = to_cper_sev(counter->common.severity);
>
>	xe_log_comp(xe, cper_sev, CSC, ERR_PTR(err), 0, "detected: %#x\n", csc_error->hec_fw_error);
>
>but if that error is representing HW SIGID - SOC_INTERNAL
>then I guess it is expected that xe_ras_error_class should
>be passed to xe_log() using:
>
>	struct xe_ras_error_class data;
>
>	xe_log_comp(xe, cper_sev, SOC_INTERNAL, &data, sizeof(data), "...\n");

It's just going to be a CPER_SEV_FATAL since it triggers survivability.  
I think I am leaning into the xe_log_comp (SOC_INTERNAL) recommendation 
here since it also matches Aravind/Mallesh's recommendation offline. I 
will post that as the next rev.

Thanks,
Umesh

>
>to allow proper CPER record generation
>
>>
>> Thanks,
>> Umesh
>>>
>>>>              xe_survivability_mode_runtime_enable(xe);
>>>>              return XE_RAS_RECOVERY_ACTION_DISCONNECT;
>>>>          }
>>>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.