From e2c6d326216642011da318b9edfd5d57ebd5a56f Mon Sep 17 00:00:00 2001 From: Christian Brauner Date: Sat, 26 Sep 2026 15:44:26 +0200 Subject: namespace: don't inherit MNT_UMOUNT in clone_mnt() MNT_UMOUNT gets raised on a mount that umount_tree() or umount_one() has taken out of its namespace. It's never cleared. __detach_mounts() unhooks a mount that carries it instead of unmounting it. disconnect_mount() lets an unmounted child stay attached only if its parent carries it. propagate_umount() takes it as the sign that a mount is going to be unmounted when it decides what else can be unmounted, where a chain of locked mounts ends and where a for reparenting an overmount. clone_mnt() copies the flags of the source and only masks MNT_INTERNAL_FLAGS which doesn't include MNT_UMOUNT. A mount that was lazily unmounted but is kept alive by a file descriptor can still be copied when its root is an nsfs or pidfs file because may_copy_tree() accepts those before checking whether the source is mounted at all: mount --bind /proc/self/ns/net /tmp/x exec 3< /tmp/x umount -l /tmp/x mount --bind /proc/self/fd/3 /tmp/y The mount on /tmp/y is attached to the caller's namespace with MNT_UMOUNT already set and every copy of it inherits the flag again. That causes a number of bugs: - Unlink its mountpoint from a mount namespace where it isn't one. __detach_mounts() takes it for a mount that was unmounted earlier, unhooks it from its parent and drops the reference the namespace holds on it. It stays in the namespace with one reference too few and as its own parent, so it can't be unmounted through the file descriptor anymore and the namespace never finds it again. When the namespace goes away it's leaked and points to a freed struct mnt_namespace that may_change_propagation() and anon_ns_root() read on the next propagation change or mount_setattr() through the file descriptor. - Make it shared and unmount a peer. trace_transfers() follows the flag from the peer into it and strips it from its peer group. - Let a propagated umount reach it. gather_candidates() marks it with T_UMOUNT_CANDIDATE but never puts it on the candidate list, so the mark is never cleared and every later propagated umount skips it. If it's the parent of a locked mount handle_locked() takes it as the cutoff and commits the locked child that should have stayed and disconnect_mount() leaves that child attached to it, unmounted under a mounted parent, which is what disconnect_mount() exists to prevent. When the namespace goes away umount_tree() visits the child a second time. And once __detach_mounts() has made it its own parent reparent() never finds the end of the chain of unmounted parents. A clone is a new mount that has never been unmounted. Add MNT_UMOUNT to MNT_INTERNAL_FLAGS so clone_mnt() never copies it. And warn when umount_tree() reaches a mount that already has it set. Fixes: 590ce4bcbfb4 ("mnt: Add MNT_UMOUNT flag") Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260926-work-mount-fixes-2-v1-3-f357abf3d17b@kernel.org Signed-off-by: Christian Brauner (Amutable) --- include/linux/mount.h | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) (limited to 'include') diff --git a/include/linux/mount.h b/include/linux/mount.h index acfe7ef86a1b..b16ba5d45605 100644 --- a/include/linux/mount.h +++ b/include/linux/mount.h @@ -52,7 +52,7 @@ enum mount_flags { MNT_ATIME_MASK = MNT_NOATIME | MNT_NODIRATIME | MNT_RELATIME, MNT_INTERNAL_FLAGS = MNT_INTERNAL | MNT_DOOMED | - MNT_SYNC_UMOUNT | MNT_LOCKED + MNT_SYNC_UMOUNT | MNT_LOCKED | MNT_UMOUNT }; struct vfsmount { -- cgit v1.2.3 From 0b6f31a3c0b16066f812e94b52a704174a6f23cd Mon Sep 17 00:00:00 2001 From: Christian Brauner Date: Wed, 30 Sep 2026 15:32:08 +0200 Subject: unshare: don't drop active namespace references that were never taken Active references on the namespaces of an nsproxy are taken when the nsproxy is installed into a task in switch_task_namespaces() and copy_namespaces() and dropped again by put_nsproxy() through deactivate_nsproxy(). But ksys_unshare() calls put_nsproxy() on an nsproxy that was never installed when set_cred_ucounts() fails. The new namespaces of that nsproxy go from zero to minus one and the namespaces shared with the caller lose a reference that belongs to the nsproxy the caller keeps using: WARNING: kernel/nscommon.c:171 at __ns_ref_active_put+0x1cd/0x230 nsproxy_ns_active_put deactivate_nsproxy ksys_unshare Afterwards the namespaces the caller lives in aren't listed by listns() anymore. set_cred_ucounts() only fails when alloc_ucounts() can't allocate and it only allocates when the real uid of the caller differs from its effective uid. Commit cefd55bd2159 ("nsproxy: fix free_nsproxy() and simplify create_new_namespaces()") separated the two cases on purpose. nsproxy_free() frees an nsproxy that was prepared but never installed and that's what a failed setns() uses in put_nsset(). Export it and use it for a failed unshare() as well. Fixes: a98621a0f187 ("unshare: fix nsproxy leak in ksys_unshare() on set_cred_ucounts() failure") Cc: stable@vger.kernel.org # v7.1+ Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-16-be34c83956ae@kernel.org Signed-off-by: Christian Brauner (Amutable) --- include/linux/nsproxy.h | 1 + 1 file changed, 1 insertion(+) (limited to 'include') diff --git a/include/linux/nsproxy.h b/include/linux/nsproxy.h index 5a67648721c7..dc2447f5f092 100644 --- a/include/linux/nsproxy.h +++ b/include/linux/nsproxy.h @@ -100,6 +100,7 @@ void exit_cred_namespaces(struct task_struct *tsk); void switch_task_namespaces(struct task_struct *tsk, struct nsproxy *new); int exec_task_namespaces(void); void deactivate_nsproxy(struct nsproxy *ns); +void nsproxy_free(struct nsproxy *ns); int unshare_nsproxy_namespaces(unsigned long, struct nsproxy **, struct cred *, struct fs_struct *); int __init nsproxy_cache_init(void); -- cgit v1.2.3 From c9ec8607b4d1ce4e862f5b6ba2f0eb727387988a Mon Sep 17 00:00:00 2001 From: Christian Brauner Date: Fri, 2 Oct 2026 15:52:47 +0200 Subject: fsnotify: let a filesystem refuse marks on its objects Don't let nullfs be watched. fanotify refuses mount and filesystem marks on SB_NOUSER superblocks but inode marks of inotify, fanotify and dnotify go through. The one inode of knullfs is the root of every kernel thread and the following patches make it reachable from userspace as the directory that stands in for an unmounted mount. A watch placed through one such directory would report the opens through all the others, across users. Add FS_DISALLOW_NOTIFY next to FS_DISALLOW_NOTIFY_PERM, refuse a mark on any object of such a filesystem in fsnotify_add_mark_list() where every backend ends up and set it for nullfs. There's nothing to watch on a permanently empty and immutable filesystem. Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-16-dd44b89d44ce@kernel.org Signed-off-by: Christian Brauner (Amutable) --- include/linux/fs.h | 1 + 1 file changed, 1 insertion(+) (limited to 'include') diff --git a/include/linux/fs.h b/include/linux/fs.h index f9d1e05e8ae6..784fa20217c4 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -2296,6 +2296,7 @@ struct file_system_type { #define FS_POWER_FREEZE 256 /* Always freeze on suspend/hibernate */ #define FS_USERNS_MOUNT_RESTRICTED 512 /* Restrict mount in userns if not already visible */ #define FS_USERNS_DELEGATABLE 1024 /* Can be mounted inside userns from outside */ +#define FS_DISALLOW_NOTIFY 2048 /* No fsnotify marks on its objects */ #define FS_RENAME_DOES_D_MOVE 32768 /* FS will handle d_move() during rename() internally. */ int (*init_fs_context)(struct fs_context *); const struct fs_parameter_spec *parameters; -- cgit v1.2.3 From e08f57a7dcb83818c058e415988c23beb6c5a47f Mon Sep 17 00:00:00 2001 From: Christian Brauner Date: Fri, 2 Oct 2026 15:52:50 +0200 Subject: readdir: take no inode lock on an immutable directory iterate_dir() takes the directory's i_rwsem shared and holds it across ->iterate_shared(). For a directory that never has an entry and is never removed the lock keeps nothing still, it only orders every reader and every writer of that inode behind each other. For the directory of a nullfs instance that matters. The instance of the initial mount namespace is the root of every empty mount namespace and the private instance is the root of every kernel thread, so one inode is shared across users who have nothing else in common. And a reader can hold the lock for as long as it likes: back the getdents() buffer with a mapping of a file on a FUSE mount of your own, let the copy of "." and ".." fault and let the server wait. Queue an exclusive taker behind it, a mkdir() in that directory goes through start_dirop() before the read-only mount is reported, and from then on every lookup that misses the dcache in that directory, every create and every mount on it waits until the server answers. One user of an empty mount namespace stalls all the others. Add FOP_IMMUTABLE for the file operations of a directory that never changes and is never removed and let iterate_dir() skip the lock for it. The flag never changes for a file, ->f_pos is protected by f_pos_lock since directories are FMODE_ATOMIC_POS, IS_DEADDIR can't be set on such a directory and neither touch_atime() nor fsnotify take i_rwsem. Set it on the nullfs directory. The placeholder directories of libfs never have an entry either but their owners remove them, so they keep the lock. Fixes: 9d4e752a24f7 ("namespace: allow creating empty mount namespaces") Cc: stable@vger.kernel.org # v7.1+ Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-19-dd44b89d44ce@kernel.org Signed-off-by: Christian Brauner (Amutable) --- include/linux/fs.h | 2 ++ 1 file changed, 2 insertions(+) (limited to 'include') diff --git a/include/linux/fs.h b/include/linux/fs.h index 784fa20217c4..deb411e86661 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -1978,6 +1978,8 @@ struct file_operations { #define FOP_ASYNC_LOCK ((__force fop_flags_t)(1 << 6)) /* File system supports uncached read/write buffered IO */ #define FOP_DONTCACHE ((__force fop_flags_t)(1 << 7)) +/* Never changes and is never removed, readdir of a directory takes no lock */ +#define FOP_IMMUTABLE ((__force fop_flags_t)(1 << 8)) /* Wrap a directory iterator that needs exclusive inode access */ int wrap_directory_iterator(struct file *, struct dir_context *, -- cgit v1.2.3