Re: [PATCH] mm: memcg: flush empty per-cpu stock slots on memcg offlining

Song Hu <[email protected]>
Newsgroups org.kernel.vger.cgroups,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <[email protected]>
Hi, 

在 2026/8/17 12:30, Joshua Hahn 写道:
> On Mon, 17 Aug 2026 10:59:17 +0800 Song Hu <[email protected]> wrote:
> 
>> refill_stock() takes a css reference for each cached memcg slot and
>> mem_cgroup_css_offline() relies on drain_all_stock() to release those
>> references before the cgroup can finish dying.
>>
>> consume_stock() can drive a slot's nr_pages to zero while its cached[]
>> pointer stays set.  is_memcg_drain_needed() skips such empty slots, so
>> the offlining drain is a no-op for them: the pinned css reference is
>> released only if the slot happens to be evicted by later unrelated
>> charges on the same CPU, or by CPU hotplug.  The memcg then lingers
>> in the dying state indefinitely.  On a CPU with container churn, a
>> per-cpu stock can pin up to NR_MEMCG_STOCK (7) zombie memcgs.
>>
>> The sibling obj_stock_flush_required() has no such gate: a cached
>> objcgs slot is flushed regardless of its byte count.  Drop the
>> nr_pages gate from is_memcg_drain_needed() accordingly.  Draining an
>> empty slot is a single css_put() and a NULL store, and it happens at
>> most once per slot per drain, so the extra work is negligible.
> 
> Hi Song,
> 
> I feel that this above statement is quite misleading.
> Scheduling a drain on a memcg that doesn't have stock eventually
> performs a drain operation on every memcg cached in that CPU.
> 
Right, I missed that drain_local_memcg_stock() flushes the whole
stock, so an unrelated charge below a dying memcg would throw away the
other slots on that CPU.  Will drop that claim and the objcg
comparison from the changelog; the objcg side has no reference to
release from the stock, so there is no symmetry to argue from.

> drain_local_memcg_stock
>   drain_stock_fully
>   {
>     for (i = 0; i < NR_MEMCG_STOCK; ++i)
>       drain_stock(stock, i);
>   {
> 
> Sure, for any empty stock during this iteration drain_stock is trivial,
> but for nonempty stock that happen to co-exist on the CPU, it's forcing
> a cache miss on the next charge.
> 
> The problem that you note does seem quite real though. I think a less
> invasive solution could be something like
> 
> if (mem_cgroup_is_descendant(memcg, root_memct) &&
>     (READ_ONCE(stock->nr_pages[i]) || css_is_dying(&memcg->css))
> 
Will do.  The ordering works out: kill_css_sync() sets CSS_DYING
before css_offline is queued, so css_is_dying() is already true when
mem_cgroup_css_offline() calls drain_all_stock().
> By the way, I think the argument that "obj_stock_flush_required has no
> such gate so is_memcg_drain_needed doesn't need one too" is also not
> really correct since a dying objcg doesn't flush anyways. 
> 
Thanks,
Song> I hope you have a great day,
> Joshua
> 
>> This is easily reproduced with short-lived cgroups pinned to one CPU:
>> 12 charge/exit/rmdir cycles leave nr_dying_subsys_memory at +2, stable
>> across later slot displacement.
>>
>> Fixes: d1a05b6973c7 ("memcg: do not try to drain per-cpu caches without pages")
>> Signed-off-by: Song Hu <[email protected]>
>> ---
>>  mm/memcontrol.c | 8 ++++++--
>>  1 file changed, 6 insertions(+), 2 deletions(-)
>>
>> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
>> index 17da1f43b7d3..b931ec16bb82 100644
>> --- a/mm/memcontrol.c
>> +++ b/mm/memcontrol.c
>> @@ -2291,8 +2291,12 @@ static bool is_memcg_drain_needed(struct memcg_stock_pcp *stock,
>>  		if (!memcg)
>>  			continue;
>>  
>> -		if (READ_ONCE(stock->nr_pages[i]) &&
>> -		    mem_cgroup_is_descendant(memcg, root_memcg)) {
>> +		/*
>> +		 * An empty slot still pins a css reference which
>> +		 * mem_cgroup_css_offline() relies on drain_all_stock()
>> +		 * to release.
>> +		 */
>> +		if (mem_cgroup_is_descendant(memcg, root_memcg)) {
>>  			flush = true;
>>  			break;
>>  		}
>> -- 
>> 2.43.0
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.