[RFC PATCH v4 2/3] mm/filemap: allow filemap faults to retry under the VMA lock
Hongru Zhang <[email protected]> Tue, 4 Aug 2026 17:52:47 +0800
| Newsgroups | gmane.linux.kernel,gmane.linux.kernel.mm |
|---|---|
| Message-ID | <[email protected]> |
From: Hongru Zhang <[email protected]> filemap_fault() currently returns VM_FAULT_RETRY_MMAP_LOCK on every retry, which forces faults taken under the per-VMA lock to fall back to mmap_lock even when the retry could complete under that lock alone. Add vmf_retry_with_fault_lock() and use it in filemap_fault() so that faults taken under the per-VMA lock return VM_FAULT_RETRY_VMA_LOCK. This lets the bounded VMA-lock retry handle the second attempt before falling back to mmap_lock. All retry paths in filemap_fault() already use a common return path. Use vmf_retry_with_fault_lock() there instead of hard-coding VM_FAULT_RETRY_MMAP_LOCK. try_handle_fault_under_vma_lock() bounds a VMA-lock retry by setting FAULT_FLAG_TRIED before retrying under the per-VMA lock, so a retried VMA-lock filemap fault may reach synchronous I/O while still holding that lock. In stress testing with the workload [1], there were over 140 million filemap_fault() calls, but the synchronous I/O path under retried VMA-lock faults was reached only 115 times, less than 0.0001% of all filemap_fault() calls; all but one completed within 1 ms, and none exceeded 4 ms. This did not show a problematic latency tail from the bounded VMA-lock retry. Based on the stress model from Kunwu Chan and Wang Lian in v2, we adapted a benchmark [2] to a 20-core Intel i7-12700 desktop by reducing the thread count and adjusting the memcg limits. The benchmark uses concurrent page faults under memcg pressure with parallel munmap to amplify mmap_lock read-write contention. Throughput (higher is better): +---------+------------+------------+-------------+ | Threads | Vanilla | Patched | Improvement | +---------+------------+------------+-------------+ | 40 | 1069.34 /s | 1404.47 /s | +31.3% | +---------+------------+------------+-------------+ | 60 | 1038.12 /s | 1682.88 /s | +62.1% | +---------+------------+------------+-------------+ | 80 | 1042.62 /s | 1766.72 /s | +69.5% | +---------+------------+------------+-------------+ mmap_lock contention count (lower is better): +---------+-----------+---------+-----------+ | Threads | Vanilla | Patched | Reduction | +---------+-----------+---------+-----------+ | 40 | 3,187,336 | 68,490 | -97.9% | +---------+-----------+---------+-----------+ | 60 | 4,385,154 | 96,234 | -97.8% | +---------+-----------+---------+-----------+ | 80 | 5,337,890 | 116,184 | -97.8% | +---------+-----------+---------+-----------+ Since filemap_fault() can now return either VM_FAULT_RETRY_* value, update the generic fault handling and fault-signal/accounting checks to recognize VM_FAULT_RETRY_MASK as incomplete faults. Update the filemap_fault() and helper comments for the new fault-lock retry path. [1] https://gist.github.com/zhr250/4d53483a6e91aa2f7ded091f6019e491 [2] https://gist.github.com/zhr250/c36c2c54d9351df37e12fd072d4926ef Signed-off-by: Hongru Zhang <[email protected]> Suggested-by: Barry Song <[email protected]> Suggested-by: Suren Baghdasaryan <[email protected]> --- mm/filemap.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/mm/filemap.c b/mm/filemap.c index 6afec636881f..a86355c5f1c7 100644 --- a/mm/filemap.c +++ b/mm/filemap.c @@ -3709,7 +3709,7 @@ vm_fault_t filemap_fault(struct vm_fault *vmf) filemap_invalidate_unlock_shared(mapping); if (fpin) fput(fpin); - return ret | VM_FAULT_RETRY; + return ret | VM_FAULT_RETRY | VM_FAULT_MAY_USE_VMA_LOCK; } EXPORT_SYMBOL(filemap_fault); -- 2.43.0