Re: [PATCH v6 0/3] perf/core: sched_task() dispatch and branch entry fixes
James Clark <[email protected]>
| Newsgroups | org.kernel.vger.linux-perf-users,org.infradead.lists.linux-arm-kernel,org.kernel.vger.bpf,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On 06/08/2026 14:52, Puranjay Mohan wrote: > These three fixes were found while adding BRBE support for > bpf_get_branch_snapshot() on arm64 and have been carried in that series > since v1 [1]. They do not depend on it, so they go on their own from > here; the version number continues from that series to avoid two > numbering schemes for the same patches. > > Patch 1 stops __perf_pmu_sched_task() passing a NULL pmu_ctx to > pmu->sched_task(). armv8pmu_sched_task() is the only implementation that > dereferences the argument, so the oops needs BRBE. > > Patch 2 makes perf_pmu_sched_task() visit PMUs whose events are all > CPU-wide. They are skipped today whenever the scheduled task has a perf > event of its own, so branch records leak across task boundaries with > perf record -b -a. intel_pmu_lbr_add() calls perf_sched_cb_inc() > unconditionally, so x86 LBR is affected the same way. > > Dropping the early return alone leaves both dispatch paths running for > one case, so perf_ctx_sched_task_cb() gains a matching gate. The two > could instead be collapsed into perf_pmu_sched_task() alone, since > __perf_pmu_sched_task() already passes the same epc, but that would move > the callback out of the perf_ctx_disable() window for every PMU rather > than just that one case, which seemed like too much for a fix tagged for > stable. > > Patch 3 clears struct perf_branch_entry with a single struct assignment. > perf_clear_branch_entry_bitfields() had drifted from the struct: new_type > and priv were never cleared, and arm_pmuv3.c allocates the per-CPU branch > stack with kmalloc(). > > Tested on a 128 CPU arm64 machine with BRBE. A WARN_ON_ONCE() at the gate > patch 2 adds to perf_ctx_sched_task_cb() fires within seconds of running > perf record -b -a alongside a task-bound event pinned to a different CPU. > > Changes in v6: > - Split the sched_task() fix into patches 1 and 2; the NULL dereference > and the missed dispatch are separate bugs with different reachability. > - Gate perf_ctx_sched_task_cb() on cpc->task_epc. v5 removed the early > return in perf_pmu_sched_task() without it, so both paths ran for a > task whose event for that PMU is pinned to another CPU. Caught by the > WARN_ON_ONCE() described above. > - Tag patches 1 and 2 for stable. > - Send separately from the BRBE series, rebased onto tip perf/core. > > [1] https://lore.kernel.org/all/[email protected]/ > > Based on tip perf/core (f4dfab174244). > > Puranjay Mohan (3): > perf/core: Fix NULL pmu_ctx passed to pmu->sched_task() > perf/core: Run sched_task() for PMUs with only CPU-wide events > perf/core: Clear the whole branch entry in perf_clear_branch_entry() > > arch/x86/events/amd/brs.c | 2 +- > arch/x86/events/amd/lbr.c | 2 +- > arch/x86/events/intel/lbr.c | 6 +++--- > drivers/perf/arm_brbe.c | 2 +- > include/linux/perf_event.h | 16 ++-------------- > kernel/events/core.c | 15 +++++++++++---- > 6 files changed, 19 insertions(+), 24 deletions(-) > Reviewed-by: James Clark <[email protected]>