[PATCH v20 5/9] cxl: Update CXL Endpoint AER handler
Terry Bowman <[email protected]>
| Newsgroups | org.kernel.vger.linux-acpi,org.kernel.vger.linux-cxl,org.kernel.vger.linux-doc,org.kernel.vger.linux-kernel,org.kernel.vger.linux-pci |
|---|---|
| Message-ID | <[email protected]> |
Rename cxl_error_detected() to cxl_pci_error_detected() and rename the struct pci_error_handlers instance from cxl_error_handlers to cxl_pci_error_handlers for consistency with the cxl_pci_ prefix used by the renamed handler and the rest of the PCI-facing entry points. Document the unconditional CXL RAS read policy: on a dead link, readl() returns 0xFFFFFFFF which is interpreted as UCE bits set and triggers a panic. If RAS registers are not mapped the read is skipped and the frozen/perm_failure switch cases defer to AER recovery for devices without active CXL.mem traffic. Signed-off-by: Terry Bowman <[email protected]> Reviewed-by: Dave Jiang <[email protected]> Reviewed-by: Jonathan Cameron <[email protected]> Reviewed-by: Alison Schofield <[email protected]> --- Changes in v19->v20: - Expand the comment block in cxl_pci_error_detected(). - Clarify commit message's first paragraph. cxl_error_handlers is a variable - Clarify AER handling in cxl_pci_error_detected() block comment Changes in v18->v19: - Add review-by for DaveJ and Jonathan - Remove reindent introduced in v18 at cxl_error_handlers definition. Changes in v17->v18: - Fix cxl_pci_error_detected() to use find_cxl_port_by_uport() and port->uport_dev - Read CXL RAS unconditionally; panic on UCE regardless of channel state - Document unconditional read policy and 0xFFFFFFFF behavior in comment - Drop guard removal paragraph from commit message (not in this diff) - Drop Reviewed-by tags pending re-review after message change Changes in v16->v17: - Rename pci_error_handlers struct instance to cxl_pci_error_handlers for cxl_pci_ naming consistency. - Restore scoped_guard(device) and dev->driver check around AER read. - NULL-check find_cxl_port_by_dev() before deref of port->uport_dev. - Updated commit message. (Terry) - Add scope cleanup for port variable in cxl_pci_error_detected() (Terry) - Drop cxl_uncor_aer_present(), rely on AER state Changes in v15->v16: - Update commit message (DaveJ) - s/cxl_handle_aer()/cxl_uncor_aer_present()/g (Jonathan) - cxl_uncor_aer_present(): Leave original result calculation based on if a UCE is present and the provided state (Terry) - Add call to pci_print_aer(). AER fails to log because is upstream link (Terry) Changes in v14->v15: - Update commit message and title. Added Bjorn's ack. - Move CE and UCE handling logic here Changes in v13->v14: - Add Dave Jiang's review-by - Update commit message & headline (Bjorn) - Refactor cxl_port_error_detected()/cxl_port_cor_error_detected() to one line (Jonathan) - Remove cxl_walk_port() (Dan) - Remove cxl_pci_drv_bound(). Check for 'is_cxl' parent port is sufficient (Dan) - Remove device_lock_if() - Combined CE and UCE here (Terry) Changes in v12->v13: - Move get_pci_cxl_host_dev() and cxl_handle_proto_error() to Dequeue patch (Terry) - Remove EP case in cxl_get_ras_base(), not used. (Terry) - Remove check for dport->dport_dev (Dave) - Remove whitespace (Terry) Changes in v11->v12: - Add call to cxl_pci_drv_bound() in cxl_handle_proto_error() and pci_to_cxl_dev() - Change cxl_error_detected() -> cxl_cor_error_detected() - Remove NULL variable assignments - Replace bus_find_device() with find_cxl_port_by_uport() for upstream port searches. Changes in v10->v11: - None --- drivers/cxl/core/ras.c | 25 ++++++++++++++++--------- drivers/cxl/cxlpci.h | 8 ++++---- drivers/cxl/pci.c | 6 +++--- 3 files changed, 23 insertions(+), 16 deletions(-) diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c index e99dcd028b738..bf479fac08565 100644 --- a/drivers/cxl/core/ras.c +++ b/drivers/cxl/core/ras.c @@ -61,7 +61,7 @@ cxl_cper_trace_uncorr_prot_err(struct cxl_memdev *cxlmd, /* * ras_cap.header_log[] holds CXL_HEADERLOG_SIZE_U32 (16) hardware - * dwords. Copy them into the front of a zero-filled + * dwords. Copy them into the front of a zero-filled * CXL_HEADERLOG_TRACE_SIZE_U32 (128) u32 staging buffer so the trace * event memcpy sees a full 512-byte source and the userspace ABI * (rasdaemon) is preserved. @@ -320,8 +320,8 @@ bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem return true; } -pci_ers_result_t cxl_error_detected(struct pci_dev *pdev, - pci_channel_state_t state) +pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev, + pci_channel_state_t state) { struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_uport(&pdev->dev); bool ue = false; @@ -341,11 +341,18 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev, } /* - * The CXL RAS read is unconditional regardless of channel - * state. Any uncorrectable error bit set in the CXL RAS - * status register triggers a panic below because CXL.mem - * cache coherency is already lost; continuing risks silent - * data corruption. + * The CXL RAS uncorrectable status is the only signal here + * that the error is a CXL internal (protocol) error. A set + * UCE bit confirms it and triggers the panic below. On a dead + * link readl() returns 0xFFFFFFFF, which sets all UCE bits and + * triggers the panic intentionally. + * + * If RAS is not mapped the read is skipped. Unlike + * cxl_do_recovery(), which is reached only after + * is_aer_internal_error() has already confirmed a CXL internal + * UCE, this path has no such confirmation, so an unmapped RAS + * block cannot attribute the error to CXL and must not panic. + * The switch cases below then handle AER recovery. */ ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL)); } @@ -373,7 +380,7 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev, } return PCI_ERS_RESULT_NEED_RESET; } -EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL"); +EXPORT_SYMBOL_NS_GPL(cxl_pci_error_detected, "CXL"); static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port, struct cxl_dport *dport, int severity) diff --git a/drivers/cxl/cxlpci.h b/drivers/cxl/cxlpci.h index 606fadd2476f3..7421132899fc4 100644 --- a/drivers/cxl/cxlpci.h +++ b/drivers/cxl/cxlpci.h @@ -79,13 +79,13 @@ struct cxl_dev_state; void read_cdat_data(struct cxl_port *port); #ifdef CONFIG_CXL_RAS -pci_ers_result_t cxl_error_detected(struct pci_dev *pdev, - pci_channel_state_t state); +pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev, + pci_channel_state_t state); void devm_cxl_dport_rch_ras_setup(struct cxl_dport *dport); void devm_cxl_port_ras_setup(struct cxl_port *port); #else -static inline pci_ers_result_t cxl_error_detected(struct pci_dev *pdev, - pci_channel_state_t state) +static inline pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev, + pci_channel_state_t state) { return PCI_ERS_RESULT_NONE; } diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c index fe11be7fd9aab..1e7be77ded634 100644 --- a/drivers/cxl/pci.c +++ b/drivers/cxl/pci.c @@ -998,8 +998,8 @@ static void cxl_reset_done(struct pci_dev *pdev) } } -static const struct pci_error_handlers cxl_error_handlers = { - .error_detected = cxl_error_detected, +static const struct pci_error_handlers cxl_pci_error_handlers = { + .error_detected = cxl_pci_error_detected, .slot_reset = cxl_slot_reset, .resume = cxl_error_resume, .reset_done = cxl_reset_done, @@ -1009,7 +1009,7 @@ static struct pci_driver cxl_pci_driver = { .name = KBUILD_MODNAME, .id_table = cxl_mem_pci_tbl, .probe = cxl_pci_probe, - .err_handler = &cxl_error_handlers, + .err_handler = &cxl_pci_error_handlers, .dev_groups = cxl_rcd_groups, .driver = { .probe_type = PROBE_PREFER_ASYNCHRONOUS, -- 2.34.1