Re: [RFC PATCH v2 01/10] mm: xswap support for zswap

Baoquan He <[email protected]>
Newsgroups org.kvack.linux-mm,org.kernel.vger.linux-kernel
Message-ID <anWhKU8NPzosQR8d@MiWiFi-R3L-srv>
Hi Johannes,

On 08/05/26 at 10:17am, Johannes Weiner wrote:
> On Wed, Aug 05, 2026 at 03:53:24PM +0800, Baoquan He wrote:
> > From: Chris Li <[email protected]>
> > 
> > Introduce extendable (virtual) swap device support ??? xswap.
> > 
> > The current zswap requires a backing swapfile. The swap slot used
> > by zswap is not able to be used by the swapfile, wasting swapfile
> > space.
> > 
> > An xswap device is a swapfile that only contains the swap header,
> > with the header indicating the size of the virtual swap space. There
> > is no swap data section, therefore no waste of swapfile space. Any
> > write to an xswap device will fail. To prevent accidental read or
> > write, bdev of swap_info_struct is set to NULL. Xswap devices set
> > the SSD flag because there is no rotational disk access when using
> > zswap.
> > 
> > Zswap writeback is disabled if all swapfiles in the system are
> > xswap devices (tracked via nr_real_swapfiles).
> > 
> > How to create an xswap device:
> >   touch swap.1G
> >   truncate -s 1G swap.1G
> >   mkswap swap.1G
> >   dd if=swap.1G of=xswap.1G bs=4096 count=1
> >   # xswap.1G is 4K on disk but reports 1G capacity
> >   swapon xswap.1G
> 
> Sigh.
> 
> Why does the user have to go through this dance?
> 
> Why does the user have to decide in advance what size the space needs
> to be?
> 
> You point out no inherent limit to how much can be compressed, so
> there is no reason to make userspace decide on an arbitrary one.
> 
> There is no reason to tie an address space that can be managed
> transparently inside the kernel to TWO named files on disk.

Thanks for looking into this.

The file-based creation dance is there only because this is RFC —
I wanted to reuse the existing swapon path so the core grow/shrink
machinery could be measured and tested without also designing a new
userspace interface. I agree it's not the right final interface.

The direction I'm thinking for the next revision:

- Drop the file requirement entirely.  An xswap device has no backing
  store, so there is no reason it needs a file.

- Use totalram_pages as the initial per-device size.  Chris suggested
  this, and it's a natural bound: if all anonymous memory is swapped
  out, that is the maximum number of swap entries zswap will ever need,
  assuming a reasonable compression ratio. The hard upper limit could
  be 2 times of system RAM, or the max system RAM memory hotplug can
  add to.

  Doing this because we need consider swap.tier support. A single global
  xswap device in swap.tier would mean all memcgs compress into the
  same device — there is only one swap entry namespace. With per-device
  xswap instances, swap.tier can bind different memcgs to different xswap
  devices, giving each its own swap slot namespace. Total isolation on slot
  usage, no cross-memcg interference.

   -----
   Hi Chris, Joungjun,
   Please correct me if I misunderstood the swap.tier concept and xswap
   use case in there.)
   -----

For creation, something like:

1. 
  echo $((4 * 1024 * 1024 * 1024)) > /sys/kernel/mm/xswap/create

or
2.
  even simpler, with automatic sizing:

  echo 1 > /sys/kernel/mm/xswap/enable

  (use totalram_pages as the default size.)

3. 
swapon -t xswap xswap0

(use totalram_pages as the default size.)

I'm open to other ideas.  If anyone have a preference for the interface,
I'd like to hear it.

> 
> There are plenty of past discussions on this very topic. I don't see
> the point in resubmitting the same thing under different names,
> without even a reference to previous discussions.
> 
> As a side note: if you have to add "(virtual)" after every instance of
> "extendable", then maybe "extendable" is a terrible name and you
> should just call it "virtual".

"extendable (virtual)" wasn't meant to explain one with the other.
Chris prefers "extendable", you prefer "virtual" — I put both in the
cover letter so the community could weigh in. I don't have a strong
preference to either. 

I'll address all of this in RFC v3 — drop the file requirement,
auto-size to totalram_pages (with grow/shrink for dynamic adjustment
on top), and settle the name. The goal of RFC v2 was to get the
core mechanics reviewed; I think that part is in reasonable shape,
and the interface is exactly the kind of thing I was hoping to get
feedback on.  Thanks for providing it.

Thanks
Baoquan
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.