Re: [PATCH] memcg: trim the per-cpu charge stock instead of draining it
Shakeel Butt <[email protected]>
| Newsgroups | org.kernel.vger.cgroups,org.kernel.vger.linux-kernel,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
On Tue, Aug 18, 2026 at 12:03:07PM +0200, Michal Hocko wrote: > On Mon 17-08-26 16:46:51, Shakeel Butt wrote: > > Joy reported that an application generating a request/response traffic > > pattern spends 44.6% to 57.0% of CPU in the memcg charge/uncharge path > > for a range of message sizes, against 0.27% to 0.71% outside that range. > > Running from the root memcg, where socket memory accounting is skipped, > > recovers the performance. > > > > Tracing the charge path showed that the application generates a pattern > > where the write syscall charges one page and the read syscall uncharges > > two pages on the same CPU. This hits a corner case in the memcg percpu > > stock code that thrashes the stock continuously. > > > > In the memcg percpu stock code, MEMCG_CHARGE_BATCH (64) is both the high > > watermark and the emptying target, i.e. on a request to charge one page > > the kernel charges MEMCG_CHARGE_BATCH pages and caches > > (MEMCG_CHARGE_BATCH - 1) of them in the percpu stock. The following > > uncharge of 2 pages takes the cached count to (MEMCG_CHARGE_BATCH + 1), > > and refill_stock() then empties the cache completely. With such a > > pattern the percpu stock becomes completely ineffective. > > > > Instead of a single boundary point for charges, use the technique the > > page allocator uses for its own percpu caches, which keeps the watermark > > and the emptying target apart: nr_pcp_free() frees between batch and > > high - batch pages, leaving at least pcp->batch on the list. Add a high > > watermark MEMCG_STOCK_HIGH and, once the cached count goes over it, > > return only the pages above MEMCG_STOCK_LOW. The watermarks are > > MEMCG_CHARGE_BATCH apart, so a page_counter update still covers a full > > batch. Peak cached pages per memcg grows from 64 to 96, the same > > high-versus-batch tradeoff the page allocator makes. > > The idea is sound. I would just not increase the overall stock size in > the same patch. Fine tuning can be done independently and ideally with > some numbers. > Would it make sense to start with MEMCG_STOCK_HIGH := MEMCG_CHARGE_BATCH > and MEMCG_CHARGE_BATCH := MEMCG_CHARGE_BATCH / 2. That would preserve > the maximum stock size while preventing all or nothing behavior which is > indeed suboptimal and pushing charging path to a slower path way too > aggressively. > > WDYT? Yes, this makes sense. Let me run the experiment with that workload to make sure the newer number works and resend the patch. Thanks for the review.