Re: [PATCH v4 0/2] mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance cost

Usama Arif <[email protected]> Fri, 24 Jul 2026 14:31:12 +0100
Newsgroups org.kernel.vger.cgroups,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <[email protected]>

On 23/07/2026 21:35, Andrew Morton wrote:
> On Mon, 20 Jul 2026 09:41:21 -0700 Usama Arif <[email protected]> wrote:
> 
>> The anon/file scan balance heuristic in get_scan_count() is fed by two
>> scalars in struct lruvec (anon_cost, file_cost) that every reclaim
>> producer updates under lruvec->lru_lock. The cost-recording work
>> itself is trivial, but it both contends for and contributes to
>> contention on lru_lock - which is often a contention point on
>> memory-pressured workloads. Specifically:
>>
>> - shrink_inactive_list() re-acquires lru_lock at function exit just
>>   to call lru_note_cost_unlock_irq().
>> - shrink_active_list() does the same after rotation accounting.
>> - workingset_refault() takes folio_lruvec_lock_irq() purely to
>>   record the refault cost.
>> - prepare_scan_control() snapshots anon_cost/file_cost under
>>   lru_lock.
>> - lru_note_cost_unlock_irq() itself walks parent_lruvec() and
>>   re-acquires lru_lock on every ancestor, multiplying the cost
>>   of every update by memcg-hierarchy depth.
>>
>> This series removes those producer-side acquisitions entirely.
> 
> Thanks.  We're still awaiting review of [2/2].  I'll get this under
> test in order to parallelize things.

Thanks Andrew! Yes I think we have the reviews, but just 1 outstanding
comment from Shakeel and need to shorten commit message as Johannes
mentioned. Hopefully we will have a final version next week.

> 
>> Measurement is a 30 s `perf lock record -a` window
>> over otherwise-idle hardware.
>>
>> Workload rates are identical on both kernels (the bench drives the
>> same memory pressure):
>>
>>                           baseline    patched      delta
>>   pgscan_direct  / s      172,662     171,817      ~0%
>>   pgsteal_direct / s       67,162      66,306      ~0%
>>   workingset_refault_anon / s
>>                            40,696      39,830      ~0%
>>
>> perf lock contention (total wait per 30 s window):
>>
>>   Lock Name                Before      After     % change
>>   shrink_lruvec+0x770     722.84 ms    0         -100% (eliminated)
>>         (= lru_note_cost_unlock_irq)
>>   workingset_refault+0x167 385.26 ms   0         -100% (eliminated)
>>         (= lru_note_cost_refault)
>>   shrink_node+0x4ad       689.43 ms    26.95 ms  -96%
>>   shrink_active_list      208.34 ms    15.97 ms  -92%
>>   lru_add_drain_cpu+0x34    1.96 s    917.71 ms  -53%
>>
>>   Total LRU lock wait      ~4.23 s     ~1.66 s   -61%
> 
> So it saves 2.5 seconds out of 30 seconds?

Yes 2.5 seconds total CPU time in the 30 second perf record.