Re: [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression
Usama Arif <[email protected]>
| Newsgroups | dev.linux.lists.oe-lkp,org.kernel.vger.cgroups,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
On 13/08/2026 18:29, Shakeel Butt wrote:
> On Thu, Aug 13, 2026 at 06:14:50PM +0100, Usama Arif wrote:
>>
>>
>> On 13/08/2026 18:01, Shakeel Butt wrote:
>>> On Thu, Aug 13, 2026 at 09:16:22PM +0800, kernel test robot wrote:
>>>>
>>>>
>>>> Hello,
>>>>
>>>> kernel test robot noticed a 36.2% regression of stress-ng.mremap.ops_per_sec on:
>>>>
>>>> commit: ea928e9e18da682e9a5bc40aa862bff7ce5ae42e ("mm/vmpressure: move v1 userspace eventfd code into memcontrol-v1.c") https://git.kernel.org/cgit/linux/kernel/git/next/linux-next.git master
>>>>
>>>> in testcase: stress-ng
>>>> version: stress-ng-x86_64-29ce10a2c-1_20260712
>>>> with following parameters:
>>>>
>>>> nr_threads: 100%
>>>> testtime: 60s
>>>> test: mremap
>>>> cpufreq_governor: performance
>>>>
>>>>
>>>>
>>>> config: x86_64-rhel-9.4 (CONFIG_MEMCG=y and CONFIG_MEMCG_V1 is not set)
>>>> compiler: gcc-14
>>>> test machine: 256 threads 4 sockets INTEL(R) XEON(R) PLATINUM 8592+ (Emerald Rapids) with 256G memory
>>>>
>>>> (please refer to attached dmesg/kmsg for entire log/backtrace)
>>>>
>>>
>>> Hi there,
>>>
>>> Can you please test the following patch and see if it fixes the regression?
>>>
>>>
>>> From 84c0b05b3bc5cf73ee66ead75aafb1ad684462c3 Mon Sep 17 00:00:00 2001
>>> From: Shakeel Butt <[email protected]>
>>> Date: Thu, 13 Aug 2026 09:38:28 -0700
>>> Subject: [PATCH] memcg: keep vmstats_percpu off the memory_events[] cacheline
>>>
>>> Signed-off-by: Shakeel Butt <[email protected]>
>>> ---
>>> include/linux/memcontrol.h | 10 ++++++----
>>> 1 file changed, 6 insertions(+), 4 deletions(-)
>>>
>>> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
>>> index e78bc98ab229..e25d5b9a1db8 100644
>>> --- a/include/linux/memcontrol.h
>>> +++ b/include/linux/memcontrol.h
>>> @@ -246,8 +246,13 @@ struct mem_cgroup {
>>> /* handle for "memory.swap.events" */
>>> struct cgroup_file swap_events_file;
>>>
>>> - /* memory.stat */
>>> + /* Read-mostly. */
>>> struct memcg_vmstats *vmstats;
>>> + struct memcg_vmstats_percpu __percpu *vmstats_percpu;
>>> + int kmemcg_id;
>>> +
>>> + /* Write-hot from here on; do not let it share with the above. */
>>> + CACHELINE_PADDING(_pad_);
>>>
>>> /* memory.events */
>>> atomic_long_t memory_events[MEMCG_NR_MEMORY_EVENTS];
>>> @@ -266,9 +271,6 @@ struct mem_cgroup {
>>> #if BITS_PER_LONG < 64
>>> seqlock_t socket_pressure_seqlock;
>>> #endif
>>> - int kmemcg_id;
>>> -
>>> - struct memcg_vmstats_percpu __percpu *vmstats_percpu;
>>>
>>> #ifdef CONFIG_CGROUP_WRITEBACK
>>> struct list_head cgwb_list;
>>
>>
>> I was currently testing this diff, not sure which one would be better.
>
> I was just checking if false sharing of vmstats_percpu is the cause. If your
> patch does not increase the struct size, we can go with that as a backportable
> fix. I am planning to rearrange fields of struct mem_cgroup more drastically and
> have it more stable as future work as we continuously see these regressions keep
> popping up.
>
Yes this makes sense. I did not expect such a big change in a benchmark
with my patch, although I feel like the microbenchmark is probably not
that realistic.
I think another issue is that its a 4 socket system.
I only have access to a single socket system, and I see a 4.38% regression.
Do you know if there a way for kernel test robot to test the below patch
on its host?
From b862e84e7bd6a54b1546b7f21a6cf991253def18 Mon Sep 17 00:00:00 2001
From: Usama Arif <[email protected]>
Date: Thu, 13 Aug 2026 11:42:05 -0700
Subject: [PATCH] mm/memcontrol: avoid false sharing between vmstats and events
Moving v1 userspace eventfd handling into memcontrol-v1.c shrank
struct vmpressure from 112 to 24 bytes when CONFIG_MEMCG_V1 is disabled.
This moved memory_events_local[MEMCG_SWAP_FAIL] and the hot
vmstats_percpu pointer onto the same cacheline.
The stress-ng mremap stressor exercises MADV_PAGEOUT with swap
disabled, generating about 20 million MEMCG_SWAP_FAIL updates per
60-second run on a 176-CPU test system. Those writes bounce the line
while memcg statistics paths load vmstats_percpu.
Move cgwb_list into the existing alignment gap and cacheline-align
vmstats_percpu. This separates the pointer from the event counters
without increasing the size of struct mem_cgroup in the tested
configuration.
The blamed commit reduced median mremap throughput by 4.38% on the
test system with one socket. The patched kernel brings the performance
to within 0.5% of the parent which is within the observed boot-to-boot
spread (up to 1.2%).
Fixes: ea928e9e18da ("mm/vmpressure: move v1 userspace eventfd code into memcontrol-v1.c")
Reported-by: kernel test robot <[email protected]>
Closes: https://lore.kernel.org/oe-lkp/[email protected]
Signed-off-by: Usama Arif <[email protected]>
---
include/linux/memcontrol.h | 9 +++++++--
1 file changed, 7 insertions(+), 2 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index e78bc98ab229b..215e2e87f42b2 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -268,10 +268,15 @@ struct mem_cgroup {
#endif
int kmemcg_id;
- struct memcg_vmstats_percpu __percpu *vmstats_percpu;
-
#ifdef CONFIG_CGROUP_WRITEBACK
struct list_head cgwb_list;
+#endif
+
+ /* Keep the hot per-CPU stats pointer away from memory event counters. */
+ struct memcg_vmstats_percpu __percpu *vmstats_percpu
+ ____cacheline_aligned_in_smp;
+
+#ifdef CONFIG_CGROUP_WRITEBACK
struct wb_domain cgwb_domain;
struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
#endif
--
2.53.0-Meta
>>
>> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
>> index e78bc98ab229b..215e2e87f42b2 100644
>> --- a/include/linux/memcontrol.h
>> +++ b/include/linux/memcontrol.h
>> @@ -268,10 +268,15 @@ struct mem_cgroup {
>> #endif
>> int kmemcg_id;
>>
>> - struct memcg_vmstats_percpu __percpu *vmstats_percpu;
>> -
>> #ifdef CONFIG_CGROUP_WRITEBACK
>> struct list_head cgwb_list;
>> +#endif
>> +
>> + /* Keep the hot per-CPU stats pointer away from memory event counters. */
>> + struct memcg_vmstats_percpu __percpu *vmstats_percpu
>> + ____cacheline_aligned_in_smp;
>> +
>> +#ifdef CONFIG_CGROUP_WRITEBACK
>> struct wb_domain cgwb_domain;
>> struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
>> #endif
>>