[PATCH v3 9/9] perf/cxl: Don't log through pmu.dev in the overflow interrupt handler
Dave Jiang <[email protected]> Fri, 31 Jul 2026 16:28:27 -0700
| Newsgroups | org.kernel.vger.linux-cxl,org.kernel.vger.linux-perf-users |
|---|---|
| Message-ID | <[email protected]> |
perf_pmu_unregister() frees pmu->dev without clearing the pointer, and
cxl_pmu_probe() registers its devm actions so that teardown runs
perf_pmu_unregister() first, then the hotplug instance removal, then
free_irq(). Nothing before free_irq() masks the interrupt, so the handler
stays live across a window where info->pmu.dev is freed and its dev_dbg()
walks that pointer.
Unsharing the interrupt does not close that window. cxl_pmu_event_stop()
leaves the counter's overflow status bit set, and only the handler clears
it, so an overflow taken just before teardown is still delivered and still
gets past the "did anything overflow" early-out. It lands in the !event
branch, where the dev_dbg() is.
Log through info->pmu.parent instead, which is devm-managed and outlives
every teardown action.
Clear the overflow status before requesting the interrupt too, since the
driver never touched it at probe and a counter left enabled with
INT_ON_OVRFLW by firmware or a previous kernel can raise an interrupt at
any point.
Fixes: 5d7107c72796 ("perf: CXL Performance Monitoring Unit driver")
Reported-by: [email protected]
Closes: https://sashiko.dev/#/patchset/[email protected]?part=1
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Dave Jiang <[email protected]>
---
v3:
- Clear CXL_PMU_OVERFLOW_REG at probe, before the handler is armed. A stale
status bit from firmware can cause overflow interrupt. (Robin)
- Ack dropped as the patch grew a hunk.
- Drop the cxl_pmu_offline_cpu() hunk. That dev_err() cannot be reached.
---
drivers/perf/cxl_pmu.c | 11 ++++++++++-
1 file changed, 10 insertions(+), 1 deletion(-)
diff --git a/drivers/perf/cxl_pmu.c b/drivers/perf/cxl_pmu.c
index 37742ce43d9f..3683427fbb7e 100644
--- a/drivers/perf/cxl_pmu.c
+++ b/drivers/perf/cxl_pmu.c
@@ -806,7 +806,7 @@ static irqreturn_t cxl_pmu_irq(int irq, void *data)
struct perf_event *event = info->hw_events[i];
if (!event) {
- dev_dbg(info->pmu.dev,
+ dev_dbg(info->pmu.parent,
"overflow but on non enabled counter %d\n", i);
continue;
}
@@ -903,6 +903,15 @@ static int cxl_pmu_probe(struct device *dev)
if (!irq_name)
return -ENOMEM;
+ /*
+ * Clear any overflow status left set by firmware or a previous kernel
+ * before the handler goes live, so it cannot mistake a stale bit for an
+ * overflow on a counter no event owns yet. The register is RW1C, and
+ * bits above the implemented counters are RsvdZ, so only write those.
+ */
+ writeq(GENMASK_ULL(info->num_counters - 1, 0),
+ info->base + CXL_PMU_OVERFLOW_REG);
+
/*
* The handler must run on info->on_cpu, so the interrupt cannot be
* shared - IRQF_NOBALANCING is only honoured for the first action on a
--
2.55.0