Re: [PATCH] ovl: add ioctls to retrieve layer file descriptors

Giuseppe Scrivano <[email protected]> Wed, 08 Jul 2026 22:45:32 +0200
Newsgroups org.kernel.vger.linux-unionfs,org.kernel.vger.linux-api,org.kernel.vger.linux-fsdevel
Message-ID <[email protected]>
Miklos Szeredi <[email protected]> writes:

> On Wed, 8 Jul 2026 at 15:44, Giuseppe Scrivano <[email protected]> wrot=
e:
>>
>> Amir Goldstein <[email protected]> writes:
>>
>> > On Wed, Jul 8, 2026 at 2:31=E2=80=AFPM Giuseppe Scrivano <gscrivan@red=
hat.com> wrote:
>> >>
>> >> Miklos Szeredi <[email protected]> writes:
>> >>
>> >> > On Wed, 8 Jul 2026 at 12:00, Giuseppe Scrivano <[email protected]=
> wrote:
>> >> >>
>> >> >> Add two ioctls to overlay filesystem to allow userspace to retrieve
>> >> >> information about the overlay layers:
>> >> >>
>> >> >> OVL_IOC_OPEN_LAYER:     return an O_PATH fd to the root of a layer.
>> >> >>                         arg =3D=3D 0 returns the upper layer (-ENO=
ENT if
>> >> >>                         no upper is configured), arg >=3D 1 returns
>> >> >>                         lower layers (-ENOENT if index is out of
>> >> >>                         range).
>> >> >
>> >> > We could do this with a plain open() call.  Something like the magic
>> >> > symlinks we have under /proc/PID/fd/.   Question is where could the=
se
>> >> > live...
>> >>
>> >> is there any existing user of such a mechanism?  I don't see any mount
>> >> specific info under /proc or /sys.
>> >>
>> >> >
>> >> >> OVL_IOC_GET_LAYERS_INFO: copy a struct ovl_layers_info to userspace
>> >> >>                         with numlower, numlowerdata, and has_upper.
>> >> >
>> >> > Isn't this info obtainable via statmount(2) already?  If not, it
>> >> > should be there, instead of a specialized ioctl.
>> >>
>> >> no that is not exposed by statmount and I don't see any way to export
>> >> file system specific data through it.  Do you've anything in mind?
>> >>
>> >> >> --- a/fs/overlayfs/ovl_entry.h
>> >> >> +++ b/fs/overlayfs/ovl_entry.h
>> >> >> @@ -35,6 +35,8 @@ struct ovl_layer {
>> >> >>         struct vfsmount *mnt;
>> >> >>         /* Trap in ovl inode cache */
>> >> >>         struct inode *trap;
>> >> >> +       /* Keeps the original fsmount file alive for OVL_IOC_OPEN_=
LAYER */
>> >> >> +       struct file *origin;
>> >> >
>> >> > Don't need to keep the file open: the only info missing is the
>> >> > original vfsmount, everything else is already there to reconstruct =
the
>> >> > file.
>> >>
>> >> I didn't manage to get that to work.  As soon as the userspace process
>> >> closes the mount fd that was passed to fsconfig, the anonymous mount
>> >> namespace is destroyed and dissolve_on_fput sets mnt->mnt_ns to NULL.
>> >>
>> >> So whenever I try to use this mount again from userspace, it is not
>> >> usable because the mount namespace is empty, causing check_mnt() to
>> >> fail.
>> >>
>> >> Do you have any suggestions on how to solve this problem?
>> >
>> > The suggestion was to store origin->f_path->mnt instead of storing ori=
gin file,
>> > because you only end up using the origin vfsmount.
>>
>> the reason I am keeping the file and not just the vfsmount is that the
>> file is what keeps the mount namespace alive (preventing
>> dissolve_on_fput from fire).
>>
>> Should we export open_detached_copy from fs/namespace.c?  I've not
>> tested it, but it might work.  Are there other ways to solve it that I
>> am not seeing?
>
> Using an anon namespace sounds good to me, that means the original
> vfsmount isn't needed at all.
>
> Not sure if it's okay for the case where the original ns is not anon,
> but we can save the vfsmount in that case if it turns out to be a
> problem.
>
> Thanks,
> Miklos

for this to work, I need something like:

diff --git a/fs/namespace.c b/fs/namespace.c
index 3d5cd5bf3b05..138d15ab37ef 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -3091,6 +3091,52 @@ static struct file *open_detached_copy(struct path *=
path, unsigned int flags)
 	return file;
 }
=20
+struct file *open_detached_copy_internal(struct path *path)
+{
+	struct mnt_namespace *ns, *mnt_ns =3D current->nsproxy->mnt_ns;
+	struct mount *mnt;
+	struct file *file;
+
+	ns =3D alloc_mnt_ns(mnt_ns->user_ns, true);
+	if (IS_ERR(ns))
+		return ERR_CAST(ns);
+
+	guard(namespace_excl)();
+
+	mnt =3D clone_mnt(real_mount(path->mnt), path->dentry,
+			CL_COPY_MNT_NS_FILE);
+	if (IS_ERR(mnt)) {
+		free_mnt_ns(ns);
+		return ERR_CAST(mnt);
+	}
+
+	mnt_add_to_ns(ns, mnt);
+	ns->nr_mounts++;
+	ns->root =3D mnt;
+
+	mntput(path->mnt);
+	path->mnt =3D mntget(&mnt->mnt);
+	file =3D dentry_open(path, O_PATH, current_cred());
+	if (IS_ERR(file))
+		dissolve_on_fput(path->mnt);
+	else
+		file->f_mode |=3D FMODE_NEED_UNMOUNT;
+	return file;
+}
+EXPORT_SYMBOL_GPL(open_detached_copy_internal);
+
 enum mount_copy_flags_t {
 	MOUNT_COPY_RECURSIVE    =3D (1 << 0),
 	MOUNT_COPY_NEW		=3D (1 << 1),

is this acceptable?

Regards,
Giuseppe