Re: [PATCH v2 00/17] mm/huge_memory: clean up folio split and lift swapcache split limits
"David Hildenbrand (Arm)" <[email protected]>
| Newsgroups | gmane.linux.kernel,gmane.linux.kernel.mm |
|---|---|
| Message-ID | <[email protected]> |
On 8/12/26 20:48, Kairui Song via B4 Relay wrote: > This series clean up the split code, add better swap cache split support > for mappingless, large order, uniform and non-uniform split. Generic > performance is on par or slightly better, and stack usage is reduced. > > The swap cache infrastructure can handle non-uniform or high order folio > replace, so there is no reason for either restriction from the THP side. > What stands in the way is the mixed anon/file folio split routine, > which makes lifting the restrictions hard to follow, and it already > carries some buggy or redundant checks. > > So this series cleans up the split path and separates anon and file > splitting into two helpers. The file split path never sees a swap > cache folio, and that is now enforced up front: a folio that is both > in the page cache and the swap cache can only be a shmem folio, which > remains unsupported and is rejected early. That helps to rule out swap > cache handling in that part completely. Only the anon split path > handles swap cache folios, with an anon mapping or mappingless: > either way the splitting is similar, and non-uniform split is > supported as well. > > Order-1 is still forbidden for swap cache splitting. In theory it is > doable for shmem swap cache folios, but a mappingless swap cache > folio cannot currently be told apart from a shmem one, so forbid it > for all swap cache folios for now. > > Testing: > > The in-tree split_huge_page_test selftest (uniform, non-uniform and > in-folio-offset splits of anon and pagecache folios) passes 62/62 on > the patched kernel. > > ftrace function_graph tracing filtered on __folio_split() was used to > compare per-call durations between the base and the patched kernel on > the same x86-64 box (interleaved runs across alternating reboots; > mean +- stddev of the per-run averages, 135 split calls per run): > > base: 24 runs, 69.6 +- 0.7 us per __folio_split() > patched: 26 runs, 68.8 +- 1.3 us per __folio_split() > > The patched kernel is consistently ~1% faster; with this sample > count the difference is outside run-to-run noise. > > On x86-64 with gcc 12 (-fstack-usage), the stack frame of > __folio_split() shrinks from 240 to 96 bytes, and the worst-case > split call chain from ~544 to ~384 (anon) or ~464 (file) bytes. > > Bloat-o-meter shows a tiny growth of huge_memory.o: > before=58419 after=58446, chg +0.05% (+27 bytes). > > Signed-off-by: Kairui Song <[email protected]> > --- I might need a bit to get to this; but the merge window is about to open either way so, so this is material for the one afterwards. -- Cheers, David