[PATCH 6.6.y] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios
Kiryl Shutsemau <[email protected]>
| Newsgroups | org.kernel.vger.stable |
|---|---|
| Message-ID | <[email protected]> |
From: "Kiryl Shutsemau (Meta)" <[email protected]> __folio_split() keeps dereferencing the mapping after the split: shmem_uncharge(mapping->host) and remap_page() while the folios are still frozen/locked, and i_mmap_unlock_read(mapping) at the very end, after the after-split folios have been unlocked and freed. Nothing holds an inode reference across that. The split relies on @folio -- which the beyond-EOF drop loop never removes, as it starts at folio_next(folio) -- staying locked and in the page cache to hold off eviction. But the unlock loop unlocks @folio before i_mmap_unlock_read() runs. If the caller's @lock_at is a tail beyond EOF, as memory_failure() passes when splitting a poisoned tail of a shmem THP that reaches past i_size during truncation, it too is gone from the page cache; so once @folio is unlocked no locked, in-cache folio pins the inode, and a concurrent final iput() can evict and RCU-free it before i_mmap_unlock_read() touches i_mmap_rwsem: BUG: KASAN: slab-use-after-free in __up_read+0x634/0x790 i_mmap_unlock_read include/linux/fs.h:537 [inline] __folio_split+0x732/0x1640 mm/huge_memory.c:4100 try_to_split_thp_page+0xab/0x390 mm/memory-failure.c:1675 memory_failure+0x1394/0x26e0 mm/memory-failure.c:2470 Freed by task 4601: shmem_free_in_core_inode+0x54/0xb0 mm/shmem.c:5177 evict+0x57f/0xac0 fs/inode.c:870 Do every mapping dereference while @folio still pins the inode: drop i_mmap_rwsem right after remap_page(), before the loop that unlocks and frees the after-split folios, and clear @mapping so the exit path does not unlock it again. shmem_uncharge() and remap_page() already run before that point, so after this nothing past the unlock loop touches the inode or the mapping. This is now a rule the split depends on, alongside keeping @folio frozen until the page cache is updated: no inode or mapping dereference once the after-split folios start being unlocked. Link: https://lore.kernel.org/[email protected] Fixes: baa355fd3314 ("thp: file pages support for split_huge_page()") Signed-off-by: Kiryl Shutsemau (Meta) <[email protected]> Reported-by: Hao Zhang <[email protected]> Closes: https://lore.kernel.org/linux-mm/20260710071344.GA106129@zh-pc Co-developed-by: Hao Zhang <[email protected]> Signed-off-by: Hao Zhang <[email protected]> Acked-by: David Hildenbrand (Arm) <[email protected]> Reviewed-by: Zi Yan <[email protected]> Reviewed-by: Baolin Wang <[email protected]> Reviewed-by: Miaohe Lin <[email protected]> Cc: Baolin Wang <[email protected]> Cc: Barry Song <[email protected]> Cc: Dev Jain <[email protected]> Cc: Lance Yang <[email protected]> Cc: Liam R. Howlett <[email protected]> Cc: Lorenzo Stoakes <[email protected]> Cc: Naoya Horiguchi <[email protected]> Cc: Nico Pache <[email protected]> Cc: Ryan Roberts <[email protected]> Cc: <[email protected]> Signed-off-by: Andrew Morton <[email protected]> (cherry picked from commit e923bd21058ea02fd0dcd3549d151d143fd036e5) [ kas: adapt to the __split_huge_page()/split_huge_page_to_list() two-function split: pass @mapping into __split_huge_page() and drop it there, before the loop that frees the after-split subpages while the head is still locked; the caller then skips its own i_mmap unlock ] Signed-off-by: Kiryl Shutsemau (Meta) <[email protected]> --- mm/huge_memory.c | 16 ++++++++++++++-- 1 file changed, 14 insertions(+), 2 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 4443cc44cbf9..ff95a802d158 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2489,7 +2489,7 @@ static void __split_huge_page_tail(struct folio *folio, int tail, } static void __split_huge_page(struct page *page, struct list_head *list, - pgoff_t end) + pgoff_t end, struct address_space *mapping) { struct folio *folio = page_folio(page); struct page *head = &folio->page; @@ -2564,6 +2564,16 @@ static void __split_huge_page(struct page *page, struct list_head *list, if (folio_test_swapcache(folio)) split_swap_cluster(folio->swap); + /* + * Drop the mapping while the head page is still locked and thus pins + * the inode. The loop below may free the after-split subpages -- + * including the head, when @page is a tail beyond EOF that the split + * dropped from the page cache -- which could otherwise let the inode, + * and @mapping, be freed before this unlock. + */ + if (mapping) + i_mmap_unlock_read(mapping); + for (i = 0; i < nr; i++) { struct page *subpage = head + i; if (subpage == page) @@ -2745,7 +2755,9 @@ int split_huge_page_to_list(struct page *page, struct list_head *list) } } - __split_huge_page(page, list, end); + __split_huge_page(page, list, end, mapping); + /* __split_huge_page() dropped the i_mmap lock */ + mapping = NULL; ret = 0; } else { spin_unlock(&ds_queue->split_queue_lock); -- 2.54.0