Re: [PATCH 0/3] ACPI: CPPC: Reduce .target() callback overhead

Christian Loehle <[email protected]> Mon, 3 Aug 2026 22:01:30 +0100
Newsgroups org.kernel.vger.linux-pm,org.kernel.vger.linux-acpi,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
On 8/3/26 21:46, Christian Loehle wrote:
> On 8/3/26 15:54, Rafael J. Wysocki (Intel) wrote:
>> On Mon, Aug 3, 2026 at 4:07 PM Christian Loehle
>> <[email protected]> wrote:
>>>
>>> On 7/24/26 14:42, Christian Loehle wrote:
>>>> cppc-cpufreq reaches cppc_set_perf() from every target callback.  For
>>>> direct SystemMemory controls, that path currently does several steps
>>>> which are unnecessary once the immutable _CPC layout is known:
>>>>
>>>>   - a full-width write first reads the access unit and merges the value.
>>>>   - every write takes the descriptor's RMW lock, even when its access unit
>>>>     is not shared with another _CPC entry.
>>>>   - cppc_set_perf() evaluates the same three PCC predicates at each phase
>>>>     of the transaction.
>>>>
>>>> Remove those costs while retaining the existing conservative paths for
>>>> partial fields, overlapping or malformed access units, and PCC controls.
>>>>
>>>> The series was tested on Arm Power-Orion O6 and AmpereOne systems using
>>>> cppc-cpufreq and schedutil.  An rt-app task pinned to one CPU ran for
>>>> 500 us every 2 ms with uclamp.min=512, for 5000 periods per run.  schedutil
>>>> rate_limit_us was 1000.  cppc_cpufreq_fast_switch() latency was measured
>>>> for 100 ten-second runs.  Each sample is the mean callback latency within
>>>> one run, and the table reports the median and sample standard deviation
>>>> of those samples.
>>>>
>>>> Orion O6               median       stdev       callbacks
>>>> baseline               5703 ns      307 ns       188017
>>>> complete series        5023 ns      240 ns       188209
>>>> ==> 680 ns (11.9%) reduction
>>>>
>>>> AmpereOne              median       stdev
>>>> baseline               2090 ns      157 ns
>>>> complete series      1907.5 ns      140 ns
>>>> ==> 182.5 ns (8.7%) reduction
>>>>
>>>> The cumulative intermediate results on the Orion O6 attribute roughly half
>>>> of the gain to each of the first two patches: avoiding the read reduced
>>>> the median by 284 ns (5.0%), and avoiding the lock reduced it by another
>>>> 336 ns (6.2%).
>>>> Together they account for 620 ns of the 680 ns total reduction.
>>>>
>>>> With the same arm64 configuration and GCC 11.4, caching the PCC predicate
>>>> also reduces cppc_set_perf() from 1124 to 884 bytes.  The generated
>>>> function has 60 fewer instructions, 17 fewer loads and 20 fewer branches.
>>>>
>>>> This series is based on the CPPC fixes posted at (already queued):
>>>> https://lore.kernel.org/lkml/[email protected]/
>>>>
>>>> and the separately posted fix still under review:
>>>> https://lore.kernel.org/lkml/[email protected]/
>>>>
>>>> PS:
>>>> There's a final optimization that I actually wanted to make but decided to
>>>> split it out for now as it somewhat replicated Sumit's series:
>>>> Skipping redundant perf ctrl writes in cpufreq-cppc if registers are non-PCC
>>>> and !shared (because the values are unchanged, the
>>>> !autonomous-common-case), but that requires the driver to have a more
>>>> complex caching- and atomic-updating machinery in place, similar to
>>>> hotplug. As opposed to this series the optimization would be for the
>>>> microcontroller handling the CPPC requests, which may be shared across
>>>> many CPUs and therefore redundant requests can increase the dvfs
>>>> transition latency.
>>>> That patch will follow once Sumit's is queued:
>>>> https://lore.kernel.org/lkml/[email protected]/
>>>>
>>>> Christian Loehle (3):
>>>>   ACPI: CPPC: Avoid unnecessary reads for full-width writes
>>>>   ACPI: CPPC: Avoid locking standalone full-width registers
>>>>   ACPI: CPPC: Evaluate performance-control PCC use once
>>>>
>>>>  drivers/acpi/cppc_acpi.c | 135 +++++++++++++++++++++++++++++++--------
>>>>  include/acpi/cppc_acpi.h |   5 +-
>>>>  2 files changed, 113 insertions(+), 27 deletions(-)
>>>>
>>>>
>>>> base-commit: 1590cf0329716306e948a8fc29f1d3ee87d3989f
>>>> prerequisite-patch-id: 4c1f4063800e3658717eb77e43af13dad52d55cc
>>>> prerequisite-patch-id: 5c40d9f099543123c9b904d18ce651e740fbefb8
>>>> prerequisite-patch-id: 7202391dca8f31f9eb8f6d59fc753b3ce665e103
>>>
>>>
>>> I was about the rebase and resend this series but noticed it still applies cleanly
>>> on 7.2-rc6 + ACPI: CPPC: Avoid Desired Performance reads on ACPI 6.6+
>>> https://lore.kernel.org/all/[email protected]/
>>>
>>> Any comments on this?
>>
>> Sashiko says that it couldn't apply this series:
>>
>> https://sashiko.dev/#/patchset/20260724134251.1632824-1-christian.loehle%40arm.com
>>
>> Would it apply on top of plain 7.2-rc6?  If so, can you please rebase
>> and resend it?  Or is it linux-next-only?
> 
> Strange, it still works for me, anything I'm missing?
> Anyway I can rebase on v7.2-rc6 and rebase, the desired_perf changes are unrelated,
> but here's my log FWIW:
> 
> # git checkout v7.2-rc6
> # git switch -c cloehle/cppc-target-optimizations-7.2-rc6-test
> # b4 am -c -o - [email protected] | git am
> Grabbing thread from lore.kernel.org/all/[email protected]/t.mbox.gz
> Checking for newer revisions
> Grabbing search results from lore.kernel.org
> Analyzing 6 messages in the thread
> Looking for additional code-review trailers on lore.kernel.org
> Analyzing 0 code-review messages
> Checking attestation on all messages, may take a moment...
> ---
>   ✓ [PATCH 1/3] ACPI: CPPC: Avoid unnecessary reads for full-width writes
>   ✓ [PATCH 2/3] ACPI: CPPC: Avoid locking standalone full-width registers
>   ✓ [PATCH 3/3] ACPI: CPPC: Evaluate performance-control PCC use once
>   ---
>   ✓ Signed: DKIM/arm.com
> ---
> Total patches: 3
> ---
>  Link: https://lore.kernel.org/r/[email protected]
>  Base: using specified base-commit 1590cf0329716306e948a8fc29f1d3ee87d3989f
>        git checkout -b 20260724_christian_loehle_arm_com 1590cf0329716306e948a8fc29f1d3ee87d3989f
> Applying: ACPI: CPPC: Avoid unnecessary reads for full-width writes
> Applying: ACPI: CPPC: Avoid locking standalone full-width registers
> Applying: ACPI: CPPC: Evaluate performance-control PCC use once
> 
> # git log --oneline
> 86f53dfeddb6 (HEAD -> cloehle/cppc-target-optimizations-7.2-rc6-test) ACPI: CPPC: Evaluate performance-control PCC use once
> 61378f475006 ACPI: CPPC: Avoid locking standalone full-width registers
> 0c5b12480412 ACPI: CPPC: Avoid unnecessary reads for full-width writes
> 075b74841bd0 (tag: v7.2-rc6, origin/master, origin/HEAD) Linux 7.2-rc6
> 

Ah looks like sashiko doesn't understand prerequisite-patch-id yet:
https://github.com/sashiko-dev/sashiko/issues/49