Re: [PATCH] Fix unbounded loop within try_charge_memcg

Audra Mitchell <[email protected]>
Newsgroups gmane.linux.kernel.cgroups,gmane.linux.kernel.mm,gmane.linux.kernel
Message-ID <anYuUpgvsUGZcGMQ@fedora>
On Fri, Aug 07, 2026 at 08:47:10PM +0200, Michal Hocko wrote:
> On Fri 07-08-26 11:40:06, Audra Mitchell wrote:
> > On Fri, Aug 07, 2026 at 10:06:53AM +0200, Michal Hocko wrote:
> > > On Thu 06-08-26 11:10:03, Audra Mitchell wrote:
> 
> Well, both global and memcg reclaim share the reclaim logic. Both of
> them try to exercise all reclaim priorities (i.e. check whole eligible
> LRU lists) and they fall back to OOM killer only if there is no other
> option left. For the global case should_reclaim_retry is the gate keeper
> around direct reclaim retries while for the memcg we have more or less
> fixed number of retries.

> From what you are describing above those users might be hitting reclaim trashing.
> I.e. last small portion of a reclaimable memory is bounced back and
> forth for the workload to make tiny but steady forward progress. While
> OOM killer might help to stop the suffering and restart the workload
> sooner I would generally recommend revisiting limits set for the
> particular workload. Especially if restarting it might lead to the same
> state sooner or later. Watching PSI metric would be a good start to see
> how the workload behaves wrt memory stalling. User space oom handlers
> might be a proper measure as well but that will always be safeguard
> rather than a solution.

Both of these recommendations were made to the end customer, with a
significant emphasis on reviewing the workload and the limits in place.

However, while reviewing the code, I was truly surprised to find retry
pathways in try_charge_memcg that do not decrement the counter, such as this:

        nr_reclaimed = try_to_free_mem_cgroup_pages(mem_over_limit, nr_pages,
                                                    gfp_mask, reclaim_options, NULL);
        psi_memstall_leave(&pflags);

        if (mem_cgroup_margin(mem_over_limit) >= nr_pages)
                goto retry;

If the intent is to have a counter of nr_retries, I am genuinely surprised
we are comfortable not adhering to said counter. The patch is not meant to
change any heuristic, but only enforce a counter that is already in place.

I have no real dog in the fight of pushing this change to upstream, so if the
patch is not desired I will let this lie. Thank you for taking a look and
the dialouge. Happy Weekend!
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.