Re: [RFC PATCH 3/6] mm/mglru: skip empty PUD subtrees during aging
Barry Song <[email protected]>
| Newsgroups | org.kvack.linux-mm |
|---|---|
| Message-ID | <CAGsJ_4yF=NN896RqhGWGtkhQPFr2ZSsFC2izhA456J_fnRv78A@mail.gmail.com> |
On Thu, Aug 6, 2026 at 6:30 PM Baoquan He <[email protected]> wrote: > > The aging walk descends every present PUD and iterates all 512 of its > PMDs, testing the PMD-level Bloom filter on each. For a process whose > memory lives only on other NUMA nodes, every PUD of this lruvec fails > the PMD test, so the whole PMD iteration is pure waste - and on > multi-socket systems these cross-node walks are common because > lru_gen_use_mm() marks an mm for all nodes at every context switch. > > Add a PUD-level Bloom filter (pud_filters) one level up. walk_pmd_range() > now reports whether it found any young leaf entries; walk_pud_range() > records that in the PUD filter and, on subsequent generations, skips the > whole 1GB subtree when the filter says it had none last generation. > The double-buffered filter flips with each new iteration, and the > existing eviction feedback (lru_gen_look_around()) keeps hot regions > marked, so newly hot or migrated-in pages are re-examined promptly > rather than suppressed indefinitely. > > force_scan walks bypass the PUD test, so manual aging and newly added > mm's always rescan and re-populate the filter. > > Signed-off-by: Baoquan He <[email protected]> > --- > mm/vmscan.c | 50 +++++++++++++++++++++++++++++++++++++++++++++++--- > 1 file changed, 47 insertions(+), 3 deletions(-) > > diff --git a/mm/vmscan.c b/mm/vmscan.c > index a397c62b2e5d..74edfe2a747d 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -2816,6 +2816,13 @@ static bool __maybe_unused seq_is_valid(struct lruvec *lruvec) > * walk_pmd_range(); the eviction also report them when walking the rmap > * in lru_gen_look_around(). > * > + * A second, coarser pair of filters (pud_filters) sits one level up. It > + * remembers which 1GB PUD subtrees had young leaf entries, so walk_pud_range() > + * can skip whole subtrees whose 512 PMDs would all fail the PMD-level test — > + * e.g. the page tables of a process whose memory lives only on other NUMA > + * nodes (cross-node empty walks). It mirrors the PMD-level filters: populated > + * by walk_pmd_range()/lru_gen_look_around(), flipped by reset_pud_bloom_filter(). > + * > * For future optimizations: > * 1. It's not necessary to keep both filters all the time. The spare one can be > * freed after the RCU grace period and reallocated if needed again. > @@ -2907,6 +2914,23 @@ static void reset_bloom_filter(struct lru_gen_mm_state *mm_state, unsigned long > __reset_bloom_filter(mm_state->filters, seq); > } > > +static bool test_pud_bloom_filter(struct lru_gen_mm_state *mm_state, unsigned long seq, > + void *item) > +{ > + return __test_bloom_filter(mm_state->pud_filters, seq, item); > +} > + > +static void update_pud_bloom_filter(struct lru_gen_mm_state *mm_state, unsigned long seq, > + void *item) > +{ > + __update_bloom_filter(mm_state->pud_filters, seq, item); > +} > + > +static void reset_pud_bloom_filter(struct lru_gen_mm_state *mm_state, unsigned long seq) > +{ > + __reset_bloom_filter(mm_state->pud_filters, seq); > +} I'd rather have symmetric names such as update_pmd_bloom_filter() and update_pud_bloom_filter(), rather than update_bloom_filter() and update_pud_bloom_filter(). it could also be: update_bloom_filter(mm_state, seq, pmd + i, PMD); update_bloom_filter(mm_state, seq, pud + i, PUD); I think either approach would make the intent clearer than the current naming. > + > /****************************************************************************** > * mm_struct list > ******************************************************************************/ > @@ -3146,8 +3170,10 @@ static bool iterate_mm_list(struct lru_gen_mm_walk *walk, struct mm_struct **ite > > spin_unlock(&mm_list->lock); > > - if (mm && first) > + if (mm && first) { > reset_bloom_filter(mm_state, walk->seq + 1); > + reset_pud_bloom_filter(mm_state, walk->seq + 1); > + } > > if (*iter) > mmdrop(*iter); > @@ -3728,10 +3754,11 @@ static void walk_pmd_range_locked(pud_t *pud, unsigned long addr, struct vm_area > *first = -1; > } > > -static void walk_pmd_range(pud_t *pud, unsigned long start, unsigned long end, > +static bool walk_pmd_range(pud_t *pud, unsigned long start, unsigned long end, > struct mm_walk *args) > { > int i; > + bool young = false; > pmd_t *pmd; > unsigned long next; > unsigned long addr; > @@ -3790,6 +3817,7 @@ static void walk_pmd_range(pud_t *pud, unsigned long start, unsigned long end, > continue; > > walk->mm_stats[MM_NONLEAF_ADDED]++; > + young = true; > > /* carry over to the next generation */ > update_bloom_filter(mm_state, walk->seq + 1, pmd + i); > @@ -3799,6 +3827,8 @@ static void walk_pmd_range(pud_t *pud, unsigned long start, unsigned long end, > > if (i < PTRS_PER_PMD && get_next_vma(PUD_MASK, PMD_SIZE, args, &start, &end)) > goto restart; > + > + return young; > } > > static int walk_pud_range(p4d_t *p4d, unsigned long start, unsigned long end, > @@ -3809,6 +3839,7 @@ static int walk_pud_range(p4d_t *p4d, unsigned long start, unsigned long end, > unsigned long addr; > unsigned long next; > struct lru_gen_mm_walk *walk = args->private; > + struct lru_gen_mm_state *mm_state = get_mm_state(walk->lruvec); > > VM_WARN_ON_ONCE(p4d_leaf(*p4d)); > > @@ -3822,7 +3853,20 @@ static int walk_pud_range(p4d_t *p4d, unsigned long start, unsigned long end, > if (!pud_present(val) || WARN_ON_ONCE(pud_leaf(val))) > continue; > > - walk_pmd_range(&val, addr, next, args); > + /* > + * Cross-node empty walk suppression. A 1GB PUD subtree whose > + * 512 PMDs all failed the PMD-level Bloom filter last generation > + * found no young leaf entries for this lruvec, so skip the whole > + * PMD iteration instead of re-checking every entry. This mirrors > + * the PMD-level filter one level up and mainly cuts the cost of > + * walking page tables of processes whose memory lives only on > + * other NUMA nodes. > + */ Is this a common case? It seems a bit odd to me that NUMA balancing doesn't keep the process and its memory on the same NUMA node. Or is it because users don't pin processes and memory properly? > + if (!walk->force_scan && !test_pud_bloom_filter(mm_state, walk->seq, pud + i)) > + continue; > + > + if (walk_pmd_range(&val, addr, next, args)) > + update_pud_bloom_filter(mm_state, walk->seq + 1, pud + i); > > if (need_resched() || walk->batched >= MAX_LRU_BATCH) { > end = (addr | ~PUD_MASK) + 1; Thanks Barry