Re: [PATCH 0/1] mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance cost

Usama Arif <[email protected]>
Newsgroups org.kernel.vger.cgroups,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <[email protected]>
On Mon,  6 Jul 2026 05:28:25 -0700 Usama Arif <[email protected]> wrote:

> The anon/file scan balance heuristic in get_scan_count() is fed by two
> scalars in struct lruvec (anon_cost, file_cost) that every reclaim
> producer updates under lruvec->lru_lock. The cost-recording work
> itself is trivial, but it both contends for and contributes to
> contention on lru_lock - which is often a contention point on
> memory-pressured workloads. Specifically:
> 
> - shrink_inactive_list() re-acquires lru_lock at function exit just
>   to call lru_note_cost_unlock_irq().
> - shrink_active_list() does the same after rotation accounting.
> - workingset_refault() takes folio_lruvec_lock_irq() purely to
>   record the refault cost.
> - prepare_scan_control() snapshots anon_cost/file_cost under
>   lru_lock.
> - lru_note_cost_unlock_irq() itself walks parent_lruvec() and
>   re-acquires lru_lock on every ancestor, multiplying the cost
>   of every update by memcg-hierarchy depth.
> 
> This patch removes those producer-side acquisitions entirely. The
> producer-local inputs (PGROTATE_*, PGRECLAIM_PAGEOUT_*) become
> per-LRU vmstat counters; WORKINGSET_RESTORE_* already captures the
> refault input. prepare_scan_control() reads the raw cost signal
> lock-free from those vmstats and folds the delta into a per-lruvec
> accumulator. A dedicated per-lruvec cost_lock, not touched by
> isolate_lru_folios(), move_folios_to_lru(), or folio_add_lru(),
> serialises the accumulator RMW and the lrusize/4 halving check.
> Hierarchy aggregation is implicit in rstat propagation, so the
> parent_lruvec() walk and the lru_reparent_memcg() cost-splice both
> disappear.

Another update on this, I was profiling another Meta workload that suffers
from very heavy reclaim and has a lot of cgroups. I ran:

bpftrace -q -e 'profile:hz:99 { @[kstack, comm] = count(); }'

The biggest entry was below:
@[
        lru_note_cost_unlock_irq+146
        shrink_lruvec+1913
        shrink_node+869
        do_try_to_free_pages+197
        try_to_free_mem_cgroup_pages+311
        __mem_cgroup_charge+1832
        filemap_add_folio+127
        page_cache_ra_unbounded+347
        filemap_fault+956
        __do_fault+40
        handle_mm_fault+4034
        do_user_addr_fault+406
        exc_page_fault+105
        asm_exc_page_fault+34
, Func___]: 5354

5354 represents 51.42% of all kernel samples collected.

A very signifcant amount of time was being spent just in lru_note_cost_unlock_irq
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.