Re: [PATCH 0/2] mount: add OPEN_TREE_NAMESPACE
Jeff Layton <[email protected]> Mon, 19 Jan 2026 17:21:30 -0500
| Newsgroups | org.kernel.vger.initramfs,dev.linux.lists.containers,org.kernel.vger.linux-api,org.kernel.vger.linux-fsdevel,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On Mon, 2026-01-19 at 11:05 -0800, Andy Lutomirski wrote: > On Mon, Jan 19, 2026 at 10:56=E2=80=AFAM Askar Safin <[email protected]= m> wrote: > >=20 > > Christian Brauner <[email protected]>: > > > Extend open_tree() with a new OPEN_TREE_NAMESPACE flag. Similar to > > > OPEN_TREE_CLONE only the indicated mount tree is copied. Instead of > > > returning a file descriptor referring to that mount tree > > > OPEN_TREE_NAMESPACE will cause open_tree() to return a file descripto= r > > > to a new mount namespace. In that new mount namespace the copied moun= t > > > tree has been mounted on top of a copy of the real rootfs. > >=20 > > I want to point at security benefits of this. > >=20 > > [[ TL;DR: [1] and [2] are very big changes to how mount namespaces work= . > > I like them, and I think they should get wider exposure. ]] > >=20 > > If this patchset ([1]) and [2] both land (they are both in "next" now a= nd > > likely will be submitted to mainline soon) and "nullfs_rootfs" is passe= d on > > command line, then mount namespace created by open_tree(OPEN_TREE_NAMES= PACE) will > > usually contain exactly 2 mounts: nullfs and whatever was passed to > > open_tree(OPEN_TREE_NAMESPACE). > >=20 > > This means that even if attacker somehow is able to unmount its root an= d > > get access to underlying mounts, then the only underlying thing they wi= ll > > get is nullfs. > >=20 > > Also this means that other mounts are not only hidden in new namespace,= they > > are fully absent. This prevents attacks discussed here: [3], [4]. > >=20 > > Also this means that (assuming we have both [1] and [2] and "nullfs_roo= tfs" > > is passed), there is no anymore hidden writable mount shared by all con= tainers, > > potentially available to attackers. This is concern raised in [5]: > >=20 > > > You want rootfs to be a NULLFS instead of ramfs. You don't seem to wa= nt it to > > > actually _be_ a filesystem. Even with your "fix", containers could co= mmunicate > > > with each _other_ through it if it becomes accessible. If a container= can get > > > access to an empty initramfs and write into it, it can ask/answer the= question > > > "Are there any other containers on this machine running stux24" and t= hen coordinate. >=20 > I think this new OPEN_TREE_NAMESPACE is nifty, but I don't think the > path that gives it sensible behavior should be conditional like this. > Either make it *always* mount on top of nullfs (regardless of boot > options) or find some way to have it actually be the root. I assume > the latter is challenging for some reason. >=20 I think that's the plan. I suggested the same to Christian last week, and he was amenable to removing the option and just always doing a nullfs_rootfs mount. We think that older runtimes should still "just work" with this scheme. Out of an abundance of caution, we _might_ want a command-line option to make it go back to old way, in case we find some userland stuff that doesn't like this for some reason, but hopefully we won't even need that. --=20 Jeff Layton <[email protected]>