[PATCH RFC 4/7] fs: support FD_FAILFS_ROOT in fchroot()
Christian Brauner <[email protected]> Thu, 23 Jul 2026 13:30:23 +0200
| Newsgroups | org.kernel.vger.linux-api,org.kernel.vger.linux-doc,org.kernel.vger.linux-fsdevel,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
Allow a process to move its root directory into failfs via
fchroot(FD_FAILFS_ROOT). From that point on every absolute path lookup
and every absolute symlink fails with EOPNOTSUPP. Combined with
fchdir(FD_FAILFS_ROOT) this leaves the process with lookups anchored
at explicit directory file descriptors only. It is the fs_struct
equivalent of RESOLVE_BENEATH. This allows taks to drop their filesystem
state completely.
Callers with CAP_SYS_CHROOT in their user namespace may always do
this, mirroring chroot(2). Unprivileged callers are subject to two
requirements (which may be loosened later):
(1) no_new_privs must be set
After entering failfs suid binaries on regular mounts remain
reachable via inherited directory file descriptors or the working
directory. A setuid program executing with an unusable root
directory might be tricked by this. I'm not 100% convinced that this
is needed but it feels more secure initially and it also forces more
no_new_privs on userspace. So win-win imo.
(2) The caller must not already be chrooted.
The root directory is what confines .. resolution. The failfs root
can never be reached by walking up a real mount tree. A task whose
root is failfs has no .. barrier left below the top of its mount
tree. A .. walk from any real directory fd it still holds climbs
to the mount-namespace root. Which is kinda the point if you want to
do fd-based lookup only. If failfs prevented you from doing that
then it doesn't make a lot of sense.
A task that a privileged manager chrooted into a subtree could use
chroot()ing into failfs as a way to allow for an inherited fd to
resolve it again.
So reject already-chrooted callers closing that issue without losing
anything for the intended self-sandboxing use case.
There's also some thought needed around shared fs_struct state. It's
obviously possible to chroot into failfs with a shared fs_struct if the
caller has CAP_SYS_CHROOT and shares the fs_struct or if the caller is
no_new_privs and shares the fs_struct. The non-chrooted-currently
requirement still applies.
Once entered, failfs is a throw-away-the-key moment. The task is
considered chrooted so it cannot create user namespaces to regain
CAP_SYS_CHROOT. chroot()/fchroot() back out require CAP_SYS_CHROOT.
The only other exit is setns() to a mount namespace file descriptor
which requires CAP_SYS_ADMIN over the target namespace plus
CAP_SYS_CHROOT and CAP_SYS_ADMIN in the caller's user namespace and
resets both root and working directory. A process that closes or never
had such file descriptors and restricts *chdir()/*chroot()/setns() via
seccomp has thrown away the key.
Signed-off-by: Christian Brauner (Amutable) <[email protected]>
---
fs/open.c | 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
diff --git a/fs/open.c b/fs/open.c
index c57f641f2e29..8adc9f00889a 100644
--- a/fs/open.c
+++ b/fs/open.c
@@ -618,6 +618,27 @@ SYSCALL_DEFINE1(chroot, const char __user *, filename)
return error;
}
+static int fchroot_failfs(void)
+{
+ struct path path;
+ int error;
+
+ if (!ns_capable(current_user_ns(), CAP_SYS_CHROOT)) {
+ if (!task_no_new_privs(current))
+ return -EPERM;
+ /* Moving the root to failfs lifts the old root's ".." barrier. */
+ if (current_chrooted())
+ return -EPERM;
+ }
+
+ failfs_get_root(&path);
+ error = security_path_chroot(&path);
+ if (!error)
+ set_fs_root(current->fs, &path);
+ path_put(&path);
+ return error;
+}
+
SYSCALL_DEFINE2(fchroot, int, fd, unsigned int, flags)
{
int error;
@@ -625,6 +646,9 @@ SYSCALL_DEFINE2(fchroot, int, fd, unsigned int, flags)
if (flags)
return -EINVAL;
+ if (fd == FD_FAILFS_ROOT)
+ return fchroot_failfs();
+
CLASS(fd_raw, f)(fd);
if (fd_empty(f))
return -EBADF;
--
2.53.0