Re: [RFC PATCH v2 01/10] mm: xswap support for zswap
Baoquan He <[email protected]>
| Newsgroups | org.kvack.linux-mm,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <anWhKU8NPzosQR8d@MiWiFi-R3L-srv> |
Hi Johannes, On 08/05/26 at 10:17am, Johannes Weiner wrote: > On Wed, Aug 05, 2026 at 03:53:24PM +0800, Baoquan He wrote: > > From: Chris Li <[email protected]> > > > > Introduce extendable (virtual) swap device support ??? xswap. > > > > The current zswap requires a backing swapfile. The swap slot used > > by zswap is not able to be used by the swapfile, wasting swapfile > > space. > > > > An xswap device is a swapfile that only contains the swap header, > > with the header indicating the size of the virtual swap space. There > > is no swap data section, therefore no waste of swapfile space. Any > > write to an xswap device will fail. To prevent accidental read or > > write, bdev of swap_info_struct is set to NULL. Xswap devices set > > the SSD flag because there is no rotational disk access when using > > zswap. > > > > Zswap writeback is disabled if all swapfiles in the system are > > xswap devices (tracked via nr_real_swapfiles). > > > > How to create an xswap device: > > touch swap.1G > > truncate -s 1G swap.1G > > mkswap swap.1G > > dd if=swap.1G of=xswap.1G bs=4096 count=1 > > # xswap.1G is 4K on disk but reports 1G capacity > > swapon xswap.1G > > Sigh. > > Why does the user have to go through this dance? > > Why does the user have to decide in advance what size the space needs > to be? > > You point out no inherent limit to how much can be compressed, so > there is no reason to make userspace decide on an arbitrary one. > > There is no reason to tie an address space that can be managed > transparently inside the kernel to TWO named files on disk. Thanks for looking into this. The file-based creation dance is there only because this is RFC — I wanted to reuse the existing swapon path so the core grow/shrink machinery could be measured and tested without also designing a new userspace interface. I agree it's not the right final interface. The direction I'm thinking for the next revision: - Drop the file requirement entirely. An xswap device has no backing store, so there is no reason it needs a file. - Use totalram_pages as the initial per-device size. Chris suggested this, and it's a natural bound: if all anonymous memory is swapped out, that is the maximum number of swap entries zswap will ever need, assuming a reasonable compression ratio. The hard upper limit could be 2 times of system RAM, or the max system RAM memory hotplug can add to. Doing this because we need consider swap.tier support. A single global xswap device in swap.tier would mean all memcgs compress into the same device — there is only one swap entry namespace. With per-device xswap instances, swap.tier can bind different memcgs to different xswap devices, giving each its own swap slot namespace. Total isolation on slot usage, no cross-memcg interference. ----- Hi Chris, Joungjun, Please correct me if I misunderstood the swap.tier concept and xswap use case in there.) ----- For creation, something like: 1. echo $((4 * 1024 * 1024 * 1024)) > /sys/kernel/mm/xswap/create or 2. even simpler, with automatic sizing: echo 1 > /sys/kernel/mm/xswap/enable (use totalram_pages as the default size.) 3. swapon -t xswap xswap0 (use totalram_pages as the default size.) I'm open to other ideas. If anyone have a preference for the interface, I'd like to hear it. > > There are plenty of past discussions on this very topic. I don't see > the point in resubmitting the same thing under different names, > without even a reference to previous discussions. > > As a side note: if you have to add "(virtual)" after every instance of > "extendable", then maybe "extendable" is a terrible name and you > should just call it "virtual". "extendable (virtual)" wasn't meant to explain one with the other. Chris prefers "extendable", you prefer "virtual" — I put both in the cover letter so the community could weigh in. I don't have a strong preference to either. I'll address all of this in RFC v3 — drop the file requirement, auto-size to totalram_pages (with grow/shrink for dynamic adjustment on top), and settle the name. The goal of RFC v2 was to get the core mechanics reviewed; I think that part is in reasonable shape, and the interface is exactly the kind of thing I was hoping to get feedback on. Thanks for providing it. Thanks Baoquan