[PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry
Breno Leitao <[email protected]>
| Newsgroups | org.kvack.linux-mm,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
I've seen some machines at Meta fleet that show the following type of
problem:
1) It gets some weird warning:
BUG: Bad page map in process khugepaged pte:f000eef300000017 pmd:00000067
addr:00007f57c0a01000 vm_flags:20200073 anon_vma:ffff88829af7c340 mapping:0000000000000000 index:7f57c0a01
The corruption is most likely the collapse/PT_RECLAIM race fixed by
commit 366a4532d96f ("mm: fix the race between collapse and PT_RECLAIM
under per-vma lock"). But this series is not about this one.
2) Then it floods all the monitoring of the fleet, sending the same
message in the loop, crashing the our fleet kernel monitoring
subsystem (which is the part that I am interested in protecting)
get_swap_device: Bad swap offset entry 3ffffffc043c5
For instance, in a host today it logged 6M in a few hours, and it is still
going forever. Two things go wrong.
1) get_swap_device() prints unconditionally, unlike print_bad_pte() next
door which suppresses itself with is_bad_page_map_ratelimited().
1) do_swap_page() returns 0 when get_swap_device() fails, so the
fault is retried, reads the same entry and faults again.
Nothing in the round trip changes the PTE.
Trying to fix it in a naive way:
Patch 1 is super simple, and rate limits the two prints.
Patch 2 makes get_swap_device() return ERR_PTR(-EINVAL) for an entry
that can never name a slot on any device, keeping NULL for a device
swapoff is taking away, and converts the callers. No functional change
expected.
Patch 3 uses that to return VM_FAULT_SIGBUS instead of retrying.
PS: Sashiko flagged several pre-existing issues, and get_swap_device()
returning an error opens the door to fixing some of them. For this
series, I am focused in landing the basic cases first and build on top,
if needed.
---
Changes in v2:
- Rate limit swap_dup_entry_direct()'s print too (Andrew)
- Drop "in get_swap_device()" from patch 1's subject, it now covers all
three prints
- Return ERR_PTR(-EIO) rather than ERR_PTR(-EINVAL) for a malformed
entry; -EINVAL is too soft for a corrupt page table (David)
- Document the malformed entry case in get_swap_device()'s kerneldoc,
in patch 2 instead of patch 3 (David)
- Reword patch 2's changelog, "an entry that can never name a slot on
any device" was unclear (David)
- Link to v1: https://patch.msgid.link/[email protected]
To: Andrew Morton <[email protected]>
To: Chris Li <[email protected]>
To: Kairui Song <[email protected]>
To: Kemeng Shi <[email protected]>
To: Nhat Pham <[email protected]>
To: Baoquan He <[email protected]>
To: Barry Song <[email protected]>
To: Youngjun Park <[email protected]>
To: David Hildenbrand <[email protected]>
To: Lorenzo Stoakes <[email protected]>
To: "Liam R. Howlett" <[email protected]>
To: Vlastimil Babka <[email protected]>
To: Mike Rapoport <[email protected]>
To: Suren Baghdasaryan <[email protected]>
To: Michal Hocko <[email protected]>
To: Jann Horn <[email protected]>
To: Pedro Falcato <[email protected]>
To: Hugh Dickins <[email protected]>
To: Baolin Wang <[email protected]>
To: Peter Xu <[email protected]>
To: Johannes Weiner <[email protected]>
To: Yosry Ahmed <[email protected]>
To: Chengming Zhou <[email protected]>
Cc: [email protected]
Cc: [email protected]
---
Breno Leitao (3):
mm, swap: ratelimit bad swap entry reports
mm, swap: distinguish a malformed swap entry from a dying device
mm: fail the fault on a malformed swap entry instead of retrying it
mm/memory.c | 9 +++++++--
mm/mincore.c | 2 +-
mm/shmem.c | 2 +-
mm/swap_state.c | 4 ++--
mm/swapfile.c | 20 ++++++++++++--------
mm/userfaultfd.c | 3 ++-
mm/zswap.c | 2 +-
7 files changed, 26 insertions(+), 16 deletions(-)
---
base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
change-id: 20260810-swap-25420f9c8ba9
Best regards,
--
Breno Leitao <[email protected]>