Re: [PATCH v20 0/6] Hierarchical Percpu Counters for RSS
Mathieu Desnoyers <[email protected]>
| Newsgroups | gmane.linux.kernel,gmane.linux.kernel.mm |
|---|---|
| Message-ID | <[email protected]> |
On 2026-08-05 14:57, Andrew Morton wrote: > On Wed, 5 Aug 2026 14:37:08 -0400 Mathieu Desnoyers <[email protected]> wrote: > >> On 2026-07-23 13:53, Mathieu Desnoyers wrote: >>> On 2026-07-07 09:15, Mathieu Desnoyers wrote: >>>> Hi Andrew, >>>> >>>> Here is the hierarchical percpu counters series rebased on top of >>>> v7.2-rc2. It includes small bootup fixes which were needed to fix >>>> bootup sequence on specific architectures, and a rename of the >>>> test config option to include "KUNIT_". >>>> >>>> This aims at replacing the prior version of the series you had >>>> in mm. >>>> >>>> As a reminder, the goal here is to provide more precise RSS counters >>>> through /proc. >>> >>> Hello, >>> >>> Just a gentle ping for feedback whenever it's convenient. >> >> Hi Andrew, >> >> Can you pick this up after the upcoming merge window ? > > Resending after -rc1 would be appropriate. Will do, thanks! > >> That's of course assuming this is solving an issue which is >> still relevant. >> >> If there is no interest in solving this issue anymore, kindly >> let me know and I'll drop this series. > > This was 1000-2000 patches ago so I've forgotten what the issue was! > Please ensure that changelogging describes the issue in glorious > detail and I'm sure it'll all come back. Certainly. For the immediate records, the issue addressed by this series is RSS stats inaccuracy when reading /proc files on large SMP boxes. It gets increasingly worse for long-lifetime processes which jump around various cores while allocating/freeing memory over their lifetime. And once we get this in, it would be relevant to use this infrastructure to address the tail latency of memcg oom killer, which AFAIU is now a scan in the order of num_possible_cpus * nb processes. Although not so bad when the OOM killer fires due to a machine memory use ballooning out of proportions, it's rather inconvenient to cause large latencies when it's used to apply the mem cgroup limits, which is a rather more common routine scenario nowadays. Thanks, Mathieu -- Mathieu Desnoyers EfficiOS Inc. https://www.efficios.com