Re: [PATCH v2 8/9] perf/cxl: Don't use pmu.dev in IRQ and hotplug callbacks after unregister
Jonathan Cameron <[email protected]> Wed, 29 Jul 2026 20:17:32 +0100
| Newsgroups | org.kernel.vger.linux-cxl,org.kernel.vger.linux-perf-users |
|---|---|
| Message-ID | <20260729201732.493818b4@jic23-huawei> |
On Wed, 29 Jul 2026 07:55:54 -0700 Dave Jiang <[email protected]> wrote: > On device removal the devm actions unwind LIFO, so cxl_pmu_perf_unregister() > runs first and perf_pmu_unregister() frees info->pmu.dev (device_del() + > put_device() -> kfree()). The overflow IRQ (freed last) and the CPU-hotplug > instance (removed next) are still live at that point, and both > cxl_pmu_irq() and cxl_pmu_offline_cpu() log via dev_dbg()/dev_err() on > info->pmu.dev, dereferencing freed memory. The shared IRQ can be entered > for a co-function on the same MSI vector, and a CPU can go offline in the > window before the hotplug instance is removed. > > Log through info->pmu.parent instead, the cxl_pmu device passed to probe, > which is devm-managed and outlives every teardown action. > > Fixes: 5d7107c72796 ("perf: CXL Performance Monitoring Unit driver") > Reported-by: [email protected] > Closes: https://sashiko.dev/#/patchset/[email protected]?part=1 > Assisted-by: Claude:claude-opus-4-8 > Signed-off-by: Dave Jiang <[email protected]> In general logging on subsystem devices is flakey so Acked-by: Jonathan Cameron <[email protected]> > --- > drivers/perf/cxl_pmu.c | 4 ++-- > 1 file changed, 2 insertions(+), 2 deletions(-) > > diff --git a/drivers/perf/cxl_pmu.c b/drivers/perf/cxl_pmu.c > index 2e817a52ff1e..f42238b2b6b0 100644 > --- a/drivers/perf/cxl_pmu.c > +++ b/drivers/perf/cxl_pmu.c > @@ -803,7 +803,7 @@ static irqreturn_t cxl_pmu_irq(int irq, void *data) > struct perf_event *event = info->hw_events[i]; > > if (!event) { > - dev_dbg(info->pmu.dev, > + dev_dbg(info->pmu.parent, > "overflow but on non enabled counter %d\n", i); > continue; > } > @@ -966,7 +966,7 @@ static int cxl_pmu_offline_cpu(unsigned int cpu, struct hlist_node *node) > info->on_cpu = -1; > target = cpumask_any_but(cpu_online_mask, cpu); > if (target >= nr_cpu_ids) { > - dev_err(info->pmu.dev, "Unable to find a suitable CPU\n"); > + dev_err(info->pmu.parent, "Unable to find a suitable CPU\n"); > return 0; > } >