Re: [PATCH] Fix unbounded loop within try_charge_memcg

Shakeel Butt <[email protected]> Fri, 7 Aug 2026 12:10:20 -0700
Newsgroups org.kernel.vger.cgroups,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <[email protected]>
On Fri, Aug 07, 2026 at 08:47:10PM +0200, Michal Hocko wrote:
> On Fri 07-08-26 11:40:06, Audra Mitchell wrote:
> > On Fri, Aug 07, 2026 at 10:06:53AM +0200, Michal Hocko wrote:
> > > On Thu 06-08-26 11:10:03, Audra Mitchell wrote:
> > > > Originally nr_retries was actually nr_oom_retries and we used it to track (and
> > > > limit) the number of times we entered the mem_cgroup_oom path and then attempted
> > > > a retry. The purpose of nr_retries counter changed with the introduction of
> > > > 9b1306192d33 ("mm: memcontrol: retry reclaim for oom-disabled and __GFP_NOFAIL
> > > > charges") so that the oom-disabled and __GFP_NOFAIL charges would also continue
> > > > to retry within the desired nr_retries threshold. Later d977aa939fca
> > > > ("mm, memcg: unify reclaim retry limits with page allocator") changed the
> > > > nr_retries counter from 5 to 16.
> > > > 
> > > > As the function has evolved we now have multiple paths that have a goto retry
> > > > path and we have lost the original purpose of the nr_retries counter, allowing
> > > > us to take a goto retry path an unbounded number of times.
> > > > 
> > > > Fix the unbounded retries by nesting the code in a loop and decrementing the
> > > > nr_retries counter correctly.
> > > 
> > > Are you trying to fix a theoretical problem spotted by the code review
> > > or is there any actual problem that you are trying to fix?
> > 
> > We have had some customer complaints that performance has slowed to a crawl when
> > the cgroup memory limit has come close to the maximum limit. In those cases, we
> > have noticed that each process is spending a large amount of time in the direct
> > reclaim path acquiring just enough memory for their specific allocation, thus
> > by-passing the oom condition yet degrading the system's overall performance.
> 
> Yes, this is entirely possible scenario.
> 
> > In the global case, direct reclaim is bounded by DEF_PRIORITY,
> 
> Well, both global and memcg reclaim share the reclaim logic. Both of
> them try to exercise all reclaim priorities (i.e. check whole eligible
> LRU lists) and they fall back to OOM killer only if there is no other
> option left. For the global case should_reclaim_retry is the gate keeper
> around direct reclaim retries while for the memcg we have more or less
> fixed number of retries.
> 
> > however, a cgroup
> > will go through the try_charge_memcg path which will call
> > try_to_free_mem_cgroup_pages->do_try_to_free_pages each time it does a retry (16
> > times).
> 
> > If we bound the loop in try_charge_memcg, the worst case is 16*12 passes
> > attempting to reclaim. This patch is meant to address the unbound case, limiting
> > the loops to 16 attempts at following the direct reclaim path. An argument could
> > be made to reduce nr_retries as well, but given that the nr_retries has been set
> > to 16 for sometime, it seemed unlikely such a change would be considered.
> 
> As Shakeel said in other reply, this is a deliberate implementation
> decision. The OOM killer is the very last resort and we are giving
> chance to userspace to handle close to OOM situation much more
> gracefully and also workload aware. Keep in mind that what might be seen
> as a slow progress for one workload might be acceptable for others where
> OOM killer could mean a lot of work being lost.
> 
> From what you are describing above those users might be hitting reclaim trashing.
> I.e. last small portion of a reclaimable memory is bounced back and
> forth for the workload to make tiny but steady forward progress. While
> OOM killer might help to stop the suffering and restart the workload
> sooner I would generally recommend revisiting limits set for the
> particular workload.

+1 to this. Beside memory.pressure, we also have refault and reclaim metrics
in memory.stat which can further help in debugging if the workload is thrashing
due to workingset larger than the limits.

> Especially if restarting it might lead to the same
> state sooner or later. Watching PSI metric would be a good start to see
> how the workload behaves wrt memory stalling. User space oom handlers
> might be a proper measure as well but that will always be safeguard
> rather than a solution.
> -- 
> Michal Hocko
> SUSE Labs