Re: [PATCH v2 2/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark
Ridong Chen <[email protected]>
| Newsgroups | org.kernel.vger.cgroups,org.kernel.vger.linux-kernel,org.kvack.linux-mm |
|---|---|
| Message-ID | <[email protected]> |
On 8/13/2026 12:48 AM, Johannes Weiner wrote: > On Fri, Aug 07, 2026 at 05:00:00PM +0800, Ridong wrote: >> From: Ridong Chen <[email protected]> >> >> Writing to memory.peak resets the peak for that fd only. Each fd is a >> watcher and reads back max(its own value, the shared local_watermark). >> >> peak_write() resets by lowering local_watermark to the current usage. >> To keep the other watchers' peaks it then walks the watcher list, but it >> stores the current usage into them instead of the old watermark. So once >> usage has dropped from a peak, a reset on one fd wrongly drags every >> other fd's peak down too, even fds that never reset. >> >> Reproduced on 7.2.0-rc5-next under QEMU, two fds A and B on one cgroup: >> B sees the peak (410624 KB), usage drops, then A resets -- and B's peak >> collapses to 1060 KB although B never reset. With this patch B keeps >> reading 410624 KB. >> >> Fix: save the old watermark before lowering it and use that to floor the >> other watchers, so a reset only affects the fd that issued it. >> >> Fixes: c6f53ed8f213 ("mm, memcg: cg2 memory{.swap,}.peak write handlers") >> Assisted-by: Claude:claude-opus-4-8 >> Signed-off-by: Ridong Chen <[email protected]> >> --- >> mm/memcontrol.c | 7 ++++--- >> 1 file changed, 4 insertions(+), 3 deletions(-) >> >> diff --git a/mm/memcontrol.c b/mm/memcontrol.c >> index 2da55b778ae3..28577beeb3d0 100644 >> --- a/mm/memcontrol.c >> +++ b/mm/memcontrol.c >> @@ -4746,7 +4746,7 @@ static ssize_t peak_write(struct kernfs_open_file *of, char *buf, size_t nbytes, >> loff_t off, struct page_counter *pc, >> struct list_head *watchers) >> { >> - unsigned long usage; >> + unsigned long usage, peer_watermark; >> struct cgroup_of_peak *peer_ctx; >> struct mem_cgroup *memcg = mem_cgroup_from_css(of_css(of)); >> struct cgroup_of_peak *ofp = of_peak(of); >> @@ -4754,11 +4754,12 @@ static ssize_t peak_write(struct kernfs_open_file *of, char *buf, size_t nbytes, >> spin_lock(&memcg->peaks_lock); >> >> usage = page_counter_read(pc); >> + peer_watermark = max(usage, READ_ONCE(pc->local_watermark)); >> WRITE_ONCE(pc->local_watermark, usage); >> >> list_for_each_entry(peer_ctx, watchers, list) >> - if (usage > peer_ctx->value) >> - WRITE_ONCE(peer_ctx->value, usage); >> + if (peer_ctx != ofp && peer_watermark > peer_ctx->value) >> + WRITE_ONCE(peer_ctx->value, peer_watermark); > > Sorry for letting your previous reply sit unanswered. You made a good > point on the peer_watermark = max(usage, local_watermark) being > pointless because that's how local_watermark moves to begin with. > > So what you had before was indeed better. It was just me missing that > detail. Could you please go back to your original? Feel free to > include: > > Acked-by: Johannes Weiner <[email protected]> Sure, thank you for your review. -- Best regards Ridong