Re: Hard system lock-ups when using encrypted swap and RAM is exhausted

Aaron Rainbolt <[email protected]> Thu, 27 Nov 2025 17:24:16 -0600
Newsgroups dev.linux.lists.cryptsetup,dev.linux.lists.dm-devel,org.kernel.vger.linux-kernel,org.kvack.linux-mm
Message-ID <20251127172323.7913c99f@kf-m2g5>
--Sig_/r1E/HN.et+YE1IVFD_tEODW
Content-Type: text/plain; charset=US-ASCII
Content-Transfer-Encoding: quoted-printable

On Thu, 27 Nov 2025 18:54:04 +0100 (CET)
Mikulas Patocka <[email protected]> wrote:

> On Thu, 27 Nov 2025, Milan Broz wrote:
>=20
> > Hi,
> >=20
> > On 11/12/25 6:18 AM, Aaron Rainbolt wrote: =20
> > > Not sure if this is a memory management issue, a LUKS issue, or
> > > both, so I wrote both mailing lists. =20
> >=20
> > It is not a LUKS issue; cryptsetup/LUKS activates the encrypted
> > device, so it is only the kernel/dm-crypt handling IOs.
> >=20
> > Adding cc to dm-devel as this would be another combination
> > device-mapper and encrypted swap that could cause issues...
> >=20
> > However, could you please specify exactly your storage
> > configuration?
> >=20
> > From the subject, I expected you to have an encrypted swap, but it
> > is not clear if there are other encrypted devices.
> >=20
> > Please paste at least lsblk, lsblk -f output, and also luksDump
> > (or crypttab if it is not LUKS) for LUKS/dm-crypt configuration.
> >=20
> > Thanks,
> > Milan
> >=20
> >  =20
> > >=20
> > > I'm seeing an issue with both the latest mainline kernel
> > > (6.18-rc5) and Debian 13's 6.12 kernel package. When physical
> > > memory fills up, the entire system locks up hard, as if it hit
> > > rather severe thrashing, despite the fact that there appears to
> > > be disk cache that can still be evicted, and there is ample
> > > amounts of swap space remaining (gigabytes of it). This issue did
> > > not occur with the 6.1 kernel in Debian 12. I'm seeing this occur
> > > in very low-memory Debian VMs, with between 512 and 900 MB RAM,
> > > running under VirtualBox and KVM. (I suspect, but have not
> > > verified, that I'm seeing similar behavior under Xen as well.)
> > > These VMs generally use a swappiness of 1, though I have seen a
> > > lockup occur even with a swappiness of 60. The filesystem in use,
> > > in case it matters, is ext4.
> > >=20
> > > To reproduce on a system running Linux 6.18-rc5, with :
> > >=20
> > > * Follow the steps from
> > >    https://gitlab.com/cryptsetup/cryptsetup/-/wikis/FrequentlyAskedQu=
estions,
> > >    section "2.3 How do I set up encrypted swap?", but creating a
> > >    swapfile rather than a swap partition. =20
>=20
> Hi
>=20
> Encrypted swap file is not supposed to work. It uses the loop device
> that routes the requests to a filesystem and the filesystem needs to
> allocate memory to process requests.
>=20
> So, this is what happened to you - the machine runs out of memory, it=20
> needs to swap out some pages, dm-crypt encrypts the pages and
> generates write bios, the write bios are directed to the loop device,
> the loop device directs them to the filesystem, the filesystem
> attempts to allocate more memory =3D> deadlock.
>=20
> I got the deadlock with 6.18-rc4 when I used dm-crypt on a file and I=20
> didn't get the deadlock when I used dm-crypt on a SCSI block device.
> That is expected behavior.

Is it only expected behavior since some time after kernel 6.1, or has
it always been expected behavior and encrypted swapfiles simply worked
by accident with kernel 6.1? Is there any reasonable way to reserve
some memory for in-kernel filesystem code (at least for some
filesystems like ext4 in the event it's not feasible for all of them)
that will ensure it has enough memory to handle I/O operations even if
the system is completely out of memory from userspace's perspective?
I'd be happy to try to contribute a fix if possible.

With kernel 6.1, this was working reliably, and Kicksecure (a
security-focused Debian derivative with a relatively sizable userbase)
was using encrypted swapfiles by default. We never got any reports of
lockups like we're seeing now with kernel 6.12. Whether this was
intended to work or not before, it seemed to work very well, and this
seems like a regression from our standpoint.

--
Aaron

> Note that the in-kernel OOM killer sometimes doesn't kill the
> application and discards read-only program pages (which generates big
> I/O churn and general system slowdown) - if you are hitting this
> problem, I recommend installing userspace OOM killer, such as
> earlyoom.
>=20
> Mikulas
>=20


--Sig_/r1E/HN.et+YE1IVFD_tEODW
Content-Type: application/pgp-signature
Content-Description: OpenPGP digital signature

-----BEGIN PGP SIGNATURE-----

iQIzBAEBCgAdFiEEudh48PFXwyPDa0wGpwkWDXPHkQkFAmko3aAACgkQpwkWDXPH
kQntQhAAxJBsISFb0MTNGHlS2xtdPK9zFeL3vVCaWdSpBfMH8BQFUZVbMYrBXEBQ
JMalBPkjor4pOFJiDNIKWrir/Er67jbtlsY/nFvK5g+AOQ0CaiPcLiokv6Jfx1oR
IFy8UNm82uoMamzv61zWwzv7BdEuVctYaGBpabJX876PqxrATeYLamEOE3Rx04h8
7uFSdc3Oz3WK99/G09lAHXibZCD053Lo++JtumG7dN3p0CQ1YGWYYR9ApssYIeZq
bFYjH2SIJf2V+4D8x5NUZAVVXfvv0QoyKk8WAyivo8jKP25wSbmOiUId8npo6JQM
fsTNO5NOba5K6TZPzjfxEsHOT54qlQs/OKu2tzC37u2DVNUWwlZHs4Iqav8NlIZY
ikYQcARj+6kvEJRgjhNaBxt40zEoT1uu5maBRLqHuM8alFmmr1IwJh+QuR5WRoXL
qXFoRZoywlCZuVwrFUh/g3N+yfZXDZ04OIWxPMqkNaPwIyNOn6JXpm/tG1B/95Nn
VU/2ZRadxceBuPrzbHnLnzSLD33wFpl8181IpRBMDTSHwCA5TMtvcJhuV6+aH12x
3O1/CZ7YqyOHEHS2moKcwves7XCBqCeION5+nJoDEIgkmjOKkZx9CEBmc8v7T17v
SXayP7eYNKLbBQfhV1pSJnt8j/n9zS7lfGTzZVN+cnWiNclF1Y4=
=yBJC
-----END PGP SIGNATURE-----

--Sig_/r1E/HN.et+YE1IVFD_tEODW--