Re: [PATCH v7 00/11] arm_mpam: Add MPAM-Fb firmware support

Gavin Shan <[email protected]>
Newsgroups gmane.linux.kernel,gmane.linux.acpi.devel,gmane.linux.ports.arm.kernel
Message-ID <[email protected]>
On 8/1/26 3:03 AM, Andre Przywara wrote:
> v7 is another quick respin of the MPAM-Fb code, for firmware based MSC
> accesses. This propagates MSC access errors in a function where it was
> missing before, covering more functions that pass their return value
> to userspace (new patch 07/11). Also some bug fixes, formatting
> adjustments and adding the accumulated tags - many thanks to the
> reviewers. Find the changelog below. Based on v7.2-rc1.
> 
> =======================
> The Arm MPAM specification defines Memory System Components (MSCs),
> which are devices that are programmed through an MMIO register frame. In
> some occasions this turned out to be too limiting: the MSC might be
> located behind a separate bus system (for instance inside an on-board
> controller), it might be mapped secure-only, or in a different processor
> socket without direct MMIO mapping. Also the MMIO access might be too slow
> or it would need to be filtered or otherwise access controlled. Finally
> there might be bugs in the MSC integration, which require a mediating
> firmware to be accessible.
> 
> To accommodate all those different use cases, the MPAM-Fb specification
> [1] describes an alternative way to access MSCs. Accesses to an MSC
> would be wrapped in a message and communicated to the system using a
> shared-memory/mailbox system mostly mimicking the Arm SCMI spec.
> For ACPI systems, this would be abstracted through an ACPI PCC channel,
> which provides the shared-memory region and the mailbox trigger. We can
> lean on existing ACPI parsing code to register with these two
> subsystems, but cannot rely on the existing SCMI code in the kernel.
> This means we somewhat need to open code a very simplified SCMI handler,
> which just provides enough functionality for the very basic subset of
> SCMI that the MPAM-Fb spec requires.
> 
> The first seven patches rework all MSC access wrappers to propagate error
> information. Pure MMIO based MSC accesses would never fail, but the
> MPAM-Fb access can go wrong in multiple ways. The patches have been split
> up purely for reviewing reasons, if the number is a problem, we could as
> well squash them. Please note that until the very last patch of this series
> any MSC accesses would always only return 0, it's only the final enablement
> of MPAM-Fb that could possibly introduce errors. Hence all former patches
> can add error handling gradually, those code paths wouldn't be triggered
> before patch 11/11.
> Patch 8/11 solves a nasty problem: At the moment we protect stateful MSC
> register accesses (mon_sel) through a spinlock. Unfortunately the mailbox
> subsystem and the slow nature of the communication through this channel
> forbid MPAM-Fb access in atomic context. So this patch keeps using a
> spinlock for MMIO based accesses, but reverts to a mutex otherwise.
> We just deny taking the lock for MPAM-Fb in atomic context, ideally we
> wouldn't need that (no need to IPI another core when the MSC access does
> not need to be local to one particular core), or we simply deny that part
> of the functionality (access through perf).
> Patch 9/11 adds the code to redirect MSC accesses through the
> PCC shmem/mailbox system.
> Patch 10/11 reworks the error interrupt handler to use a threaded IRQ for
> MPAM-Fb, to be able to do MPAM-Fb MSC accesses inside (which might sleep).
> The final patch 11/11 then adds the code to detect and store the PCC
> channel information from the ACPI tables, and eventually enables
> MPAM-Fb accesses.
> 
> This would enable systems where some MSCs are not accessible via MMIO to
> use those components anyway.
> 
> Please have a look and test!
> 
> Cheers,
> Andre
> 
> [1] https://developer.arm.com/documentation/den0144/latest
> 
> Changes in v7:
> - add tags
> - add new patch to propagate errors in mpam_reprogram_ris_partid()
> - prevent loop when mpam_diable() tries MSC accesses again
> - register MMIO error IRQ handler without IRQF_ONESHOT
> - re-use existing "err" variable instead of declaring "ret"
> - initialise mon_sel_lock later, to wait for interface decision
> - convert timeout units for nominal latency, between us and ms
>   

[...]

> 
> Andre Przywara (11):
>    arm_mpam: let low level MSC accessors return an error
>    arm_mpam: propagate MSC access errors for hw_probe functions
>    arm_mpam: propagate MSC access errors for MBWU counters
>    arm_mpam: propagate MSC access errors for msmon helpers
>    arm_mpam: propagate MSC access errors for __ris_msmon_read()
>    arm_mpam: propagate MSC access errors for state saving function
>    arm_mpam: propagate MSC access errors for mpam_reprogram_ris_partid()
>    arm_mpam: prepare mon_sel locking for MPAM-Fb
>    arm_mpam: add MPAM-Fb MSC firmware access support
>    arm_mpam: change MPAM-Fb error IRQ to use a threaded IRQ handler
>    arm_mpam: detect and enable MPAM-Fb PCC support
> 
>   drivers/resctrl/Makefile        |   2 +-
>   drivers/resctrl/mpam_devices.c  | 820 +++++++++++++++++++++++---------
>   drivers/resctrl/mpam_fb.c       | 250 ++++++++++
>   drivers/resctrl/mpam_internal.h |  62 ++-
>   include/linux/arm_mpam.h        |   2 +-
>   5 files changed, 915 insertions(+), 221 deletions(-)
>   create mode 100644 drivers/resctrl/mpam_fb.c
> 

With this series (manually) applied to v7.2.rc6, the tests for the existing functions
like kunit-tests, L3 cache partitioning, MBW (soft) limiting, llc_occupancy monitor
look fine on NVidia's grace-hopper machine.

Tested-by: Gavin Shan <[email protected]>

Thanks,
Gavin
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.