Re: [PATCH 2/3] binfmt_misc: don't warn when the mount is completed from another user namespace

Jan Kara <[email protected]> Mon, 3 Aug 2026 14:05:50 +0200
Newsgroups org.kernel.vger.linux-unionfs,org.kernel.vger.linux-fsdevel,org.kernel.vger.stable,org.kvack.linux-mm
Message-ID <lp42anrqt3khwf4fyp56odomifxeuitf4ejxahhxns7ze7nyha@dblztvmz4ocj>
On Sun 02-08-26 20:00:44, Christian Brauner wrote:
> fsopen() records the caller's user namespace in fc->user_ns and hands
> back an ordinary file descriptor. Nothing ties the task that calls
> fsconfig(FSCONFIG_CMD_CREATE) to the task that created the context. The
> fd is inherited across fork() and exec() and it can be passed over a
> unix socket.
> 
> Completing a context from another user namespace is allowed on purpose.
> vfs_cmd_create() authorizes the create with mount_capable(), which for
> FS_USERNS_MOUNT checks ns_capable(fc->user_ns, CAP_SYS_ADMIN), and that
> succeeds for a task holding CAP_SYS_ADMIN in an ancestor of fc->user_ns.
> So an unprivileged task can reach the WARN_ON() in bm_fill_super():
> create a user and a mount namespace in a child, call
> fsopen("binfmt_misc") there, send the fscontext fd to the parent and let
> the parent issue FSCONFIG_CMD_CREATE. Both namespaces come from a plain
> unshare(1) and no capability is needed anywhere:
> 
>   WARNING: fs/binfmt_misc.c:938 at bm_fill_super+0xa2/0xc0 [binfmt_misc]
>   CPU: 15 UID: 1000 PID: 3243382 Comm: fswarn
>   Call Trace:
>    get_tree_keyed+0x7d/0xb0
>    bm_get_tree+0x34/0x90 [binfmt_misc]
>    vfs_get_tree+0x2a/0x100
>    vfs_cmd_create+0x60/0xf0
>    __do_sys_fsconfig+0x4b2/0x500
> 
> The child needs the mount namespace because fsopen() itself gates on
> may_mount(), which asks for CAP_SYS_ADMIN in the user namespace owning
> the caller's mount namespace. fsconfig() doesn't repeat that check.
> 
> It is a WARN_ON() and not a WARN_ON_ONCE(), so the condition can be
> raised in a loop to taint the kernel and flood the log, and it panics a
> kernel booted with panic_on_warn.
> 
> Keep refusing the mount and stop warning about it. Nothing in
> bm_fill_super() depends on the two namespaces matching, it derives
> everything from sb->s_user_ns.
> 
> Fixes: 21ca59b365c0 ("binfmt_misc: enable sandboxed mounts")
> Cc: [email protected] # v6.7+
> Signed-off-by: Christian Brauner (Amutable) <[email protected]>

This one as well. Feel free to add:

Reviewed-by: Jan Kara <[email protected]>

								Honza

> ---
>  fs/binfmt_misc.c | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
> 
> diff --git a/fs/binfmt_misc.c b/fs/binfmt_misc.c
> index c97f10b48b5b..613dd28e3f1a 100644
> --- a/fs/binfmt_misc.c
> +++ b/fs/binfmt_misc.c
> @@ -937,7 +937,8 @@ static int bm_fill_super(struct super_block *sb, struct fs_context *fc)
>  		/* last one */ {""}
>  	};
>  
> -	if (WARN_ON(user_ns != current_user_ns()))
> +	/* The fscontext fd may have been passed to another user namespace. */
> +	if (user_ns != current_user_ns())
>  		return -EINVAL;
>  
>  	/* Never exec off this instance and never let anything stack on it. */
> 
> -- 
> 2.53.0
> 
-- 
Jan Kara <[email protected]>
SUSE Labs, CR