Re: [PATCH] mm: swap_cgroup: fix NULL deref in lookup_swap_cgroup_id on swapless host

Barry Song <[email protected]>
Newsgroups gmane.linux.kernel.stable,gmane.linux.kernel.cgroups,gmane.linux.kernel.mm,gmane.linux.kernel
Message-ID <CAGsJ_4y+kGUH7TjKEGU+hbwCotHeubNdbJtoBExP+Ya-Wf3fWA@mail.gmail.com>
On Fri, Aug 14, 2026 at 3:48 PM <[email protected]> wrote:
>
> From: "Jose Fernandez (Anthropic)" <[email protected]>
>
> [ Upstream commit 63b02a9409cb5180398491b093e48bcb5315f5fb ]
>
> lookup_swap_cgroup_id() passes swap_cgroup_ctrl[type].map to
> __swap_cgroup_id_lookup() without checking that the type was ever
> registered via swap_cgroup_swapon().  On a swapless host every ctrl->map
> is NULL, so __swap_cgroup_id_lookup() dereferences NULL + a scaled
> swp_offset().
>
> Since commit bea67dcc5eea ("mm: attempt to batch free swap entries for
> zap_pte_range()"), zap_pte_range() -> swap_pte_batch() calls
> lookup_swap_cgroup_id() on any non-present, non-none PTE that decodes as a
> real swap entry, without first validating it against swap_info[].  A
> single PTE corrupted into a type-0 swap entry takes the host down at
> process exit.

Thanks for the patch. However, we have a strict check to ensure that
this is only done for valid swap entries:

static inline int swap_pte_batch(pte_t *start_ptep, int max_nr, pte_t pte)
{
        pte_t expected_pte = pte_next_swp_offset(pte);
        const pte_t *end_ptep = start_ptep + max_nr;
        pte_t *ptep = start_ptep + 1;

        VM_WARN_ON(max_nr < 1);
        VM_WARN_ON(!softleaf_is_swap(softleaf_from_pte(pte)));

        while (ptep < end_ptep) {
                pte = ptep_get(ptep);

                if (!pte_same(pte, expected_pte))
                        break;
                expected_pte = pte_next_swp_offset(expected_pte);
                ptep++;
        }

        return ptep - start_ptep;
}

I don't know why this can happen on a swapless system.

>
> We hit this in production on a swapless 6.12.58 host: ~1s of
> "get_swap_device: Bad swap file entry 3f800204222bb" (do_swap_page() being
> correctly defensive about the same entry) followed by
>
>   BUG: unable to handle page fault for address: 000003f800204220
>   RIP: 0010:lookup_swap_cgroup_id+0x2b/0x60
>   Call Trace:
>    swap_pte_batch+0xbf/0x230
>    zap_pte_range+0x4c8/0x780
>    unmap_page_range+0x190/0x3e0
>    exit_mmap+0xd9/0x3c0
>    do_exit+0x20c/0x4b0
>
> syzbot has reported the identical stack.
>
> The source of the PTE corruption is a separate bug; this change makes the
> teardown path as robust as the fault path already is.  Every other caller
> of lookup_swap_cgroup_id() is downstream of a get_swap_device() that has
> already validated the entry, so the new branch is cold.

If the source is PTE corruption, I think we should fix the corruption
itself rather than work around it here.

Best Regards
Barry
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.