diff options
| author | Linus Torvalds <torvalds@linux-foundation.org> | 2026-10-09 23:37:44 +0200 |
|---|---|---|
| committer | Linus Torvalds <torvalds@linux-foundation.org> | 2026-10-09 23:37:44 +0200 |
| commit | f99005fbed4e8071110faa350cc5fa453afa1455 (patch) | |
| tree | 2e5e41e0cadba5b7c961a82a7c17d00439bf0751 /include/linux/mount.h | |
| parent | 9515f634dbed2bb8928fbb3b01102dd4e64f4aeb (diff) | |
| parent | 3d399224425573875b6f6f1181bd8cf9a28cc7d1 (diff) | |
Merge tag 'vfs-7.3-rc7.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
"This contains fixes for the current development cycle.
All of them came out of a review of the mount code that started with a
bug report. The review modeled the corner cases of mount propagation,
unmounting and mount reference counting and turned up a lot of bugs.
Most of them years old. Most fixes come with a selftest.
- Rework connected mounts.
A mount that is unmounted together with its parent can stay
attached to the parent to keep its mountpoint covered. That happens
when the mountpoint is removed with rmdir(), unlink() or rename(),
when a detached tree is dissolved, and for locked mounts in any
umount that isn't synchronous, including the teardown of their
mount namespace. The parent then owns the child and drops it on its
own final mntput(). So any reference from the child's superblock
back to one of its ancestors becomes a cycle that is never freed.
A loop device backed by an image on a tmpfs and mounted on that
same tmpfs is enough. Remove the directory the tmpfs is mounted on
from the host, let the container's mount namespace exit, and the
loop device, the tmpfs and the filesystem on the loop device are
leaked for good. The same works with autofs, zram, ecryptfs,
binfmt_misc, fuse passthrough, zloop, a mass storage gadget and md,
and the selftests have reproducers for them. This has been possible
since v4.1. It is also why "put_mnt_ns(): leave mounts connected"
was reverted in -rc5. Keeping every mount of a dying mount
namespace connected made these cycles trivial to create.
Every unmounted mount is now detached from its parent. Where the
mountpoint has to stay covered the mount leaves a cover on the
parent instead, allocated together with the mount. A lookup on the
unmounted parent that hits a cover finds an empty immutable
directory or file on the private nullfs instance. Nothing leads
from a cover to another mount, so no unmounted mount owns another
one and no cycle can form.
This is visible to userspace. A formerly connected mount can no
longer be reached through its unmounted parent and ".." inside it
leads nowhere, as for every other lazily unmounted mount.
With that the private nullfs instance becomes reachable from
userspace, so it now refuses mounts on top, is mounted read-only,
and refuses fsnotify marks and file locks. Its inodes are shared by
every holder and a watch or a lock would otherwise reach across
users. may_decode_fh() now also decides its subtree check under a
single mount_lock hold, as a racing umount could otherwise let it
decode into what a locked child covered.
- umount:
* Don't silently unmount busy mounts.
Since v4.13 propagate_umount() takes down propagated copies of
the victim with children as long as each child is an overmount or
another copy of the victim, but propagate_mount_busy() only ever
checked copies without children or with just an overmount. A
container that moved a tree beneath its copy of a host mount lost
that tree from under its open file descriptor to a plain umount()
on the host. propagate_mount_busy() now applies the same rules,
walking each chain of copies once.
* Don't let a migrating task hide its reference from umount().
mnt_get_count() sums the per-cpu counters under mount_lock but
the mntget() and mntput() fast paths don't take it. A task that
takes a reference on a cpu the sum has already passed and drops
it after migrating to one the sum hasn't reached yet hides the
reference it held to begin with, and umount() succeeds with the
file still open. Gets and puts now live in separate per-cpu
counters and all puts are summed before all gets with a full
barrier in between, the way srcu_readers_active_idx_check() does
it. mntget() is unchanged and mntput() gains an smp_wmb().
* Check each submount for references right before unmounting it.
shrink_submounts() and mark_mounts_for_expiry() checked all their
victims up front. Unmounting the first could move a busy
overmount to where the next victim's propagated copy is looked up
and it was then unmounted without a check.
* Never expire a locked mount.
A shrinkable mount moved beneath a locked mount with
MOVE_MOUNT_BENEATH takes over the lock, and umount() of an
unlocked ancestor expired it and revealed what it covered. That
umount() now fails with EBUSY as it does for any other locked
child. A lazy umount still takes the whole tree.
- Overmounts and locked mounts:
* Unhash a dentry before detaching the mounts on it.
unlink(), rmdir() and rename() detach the mounts on the victim
but only d_delete() it once its inode is unlocked, a window that
includes an expedited RCU grace period. In between, a lookup from
a mount namespace in which the dentry is a mountpoint found it
hashed, positive and uncovered. Drop the dentry first, as
d_invalidate() already does.
* Don't reveal overmounted entries in refwalk.
A refwalk that had grabbed the dentry before the unlink never
rechecked it the way rcuwalk does with d_seq and mount_lock.
Without any artificial widening three walkers read the covered
file 27 times in a minute. step_into() now fails an unhashed
dentry marked DCACHE_CANT_MOUNT with -ESTALE and the walk is
retried.
* Keep covered mounts covered in OPEN_TREE_NAMESPACE.
Creating such a mount namespace only takes a user namespace and
the copy followed bind mount rules: no children without
AT_RECURSIVE and no unbindable mounts with it. An unprivileged
user could see what mounts covered in the source, such as the
parts of /proc and /sys that container runtimes mask. If the
caller doesn't own the source mount namespace a non-recursive
copy of a mount with something mounted below the requested
directory is now refused and a recursive copy includes unbindable
mounts, the way unshare() copies.
* Keep the lock on a mount that a propagated copy is moved beneath.
MNT_LOCKED moved to any mount that ended up beneath a locked
mount, propagated copies included. A host mount and umount on a
directory covered by a locked mount in a less privileged mount
namespace left that cover unlocked for the namespace's owner to
remove. Only mounts the caller places beneath take over the lock
now.
* Handle mount locking for automounts correctly. Which copies to
lock was decided by the mount namespace of the task that
triggered the automount. A task in a user namespace that
triggered one on a host mount through a file descriptor got the
host's own automount locked while its own copy stayed unlocked
and could have nosuid, nodev and noexec cleared. Use the owner of
the mount namespace the mount lands in.
- Use-after-free and crashes:
* Refuse an automount below a mount that is in no namespace.
The private clones overlayfs uses for its layers have the
MNT_NS_INTERNAL error pointer as their namespace, which
finish_automount() let through and count_mounts() dereferenced. A
fanotify filesystem mark on an overlayfs lower layer hands out
file descriptors on such a clone. With debugfs as the lower layer
opening "tracing" oopses with namespace_sem held for writing and
every mount operation on the system blocks from then on.
* Reset the old parent's ->overmount in mnt_change_mountpoint().
When propagate_umount() moved an overmount off a mount that a
file descriptor kept alive, MOVE_MOUNT_BENEATH through that
descriptor later followed the stale pointer into the freed
overmount.
* statmount() with STATMOUNT_BY_FD and pivot_root() read the parent
of a mount that may be unmounted and only held by a file
descriptor, while the parent's final mntput() can free it.
statmount() now reads it under mount_lock and pivot_root() first
checks that both mounts are in the caller's mount namespace.
* Queue a mount only once for mount notifications. A mount
reparented by one umount_tree() and taken down by the next under
the same namespace_sem hold, as in shrink_submounts(), was queued
twice. That cut the mounts queued in between out of notify_list
while it still pointed at them, and once they were freed every
later mount operation walked freed memory.
* Don't let a pseudo dentry become the root of a mount. A bind
mount of a bpf token file did that with a DCACHE_NORCU dentry,
which is freed without an RCU grace period while lockless path
walks may still look at it. Refuse to clone such a mount.
* Don't inherit MNT_UMOUNT in clone_mnt(). A bind mount of a lazily
unmounted nsfs or pidfs mount through its file descriptor started
out flagged as unmounted. Among other things __detach_mounts()
then dropped the namespace's reference on it, the mount outlived
its namespace and mount_setattr() through the descriptor read the
freed namespace. A recursive bind mount of such a mount also
copied the unmounted stack still attached to it. That now fails
with EINVAL, copying the mount itself still works.
* Remove the fsnotify marks of a mount namespace in free_mnt_ns()
instead of the RCU callback that frees the namespace, where
taking the group mutexes meant sleeping in softirq context.
- Propagation and copies:
* Keep a copied mount unbindable. Since v6.17 clone_mnt() didn't
copy the unbindable flag, so every mount namespace created with
CLONE_NEWNS had bindable copies of all unbindable mounts. This
had been fixed once before.
* Refuse MOVE_MOUNT_SET_GROUP on an unbindable mount. It made the
mount an unbindable slave, a state nothing else can produce, or
silently dropped the unbindable flag. CRIU applies MS_UNBINDABLE
after restoring sharing and isn't affected.
* Check a recursive bind mount for mount namespace loops.
Recursively bind mounting a tree from another mount namespace
could put a mount of a namespace's file inside that same
namespace, which then pins itself and all its mounts. Repeating
it leaks without limit, the reproducer took Shmem from 380 kB to
65916 kB. The copy is now checked with check_for_nsfs_mounts()
before it is grafted, as move_mount() does.
* Look at the topmost mount for a mount namespace file.
attach_recursive_mnt() never looked at the topmost mount of the
source's chain of overmounts. If that was the chain's only mount
namespace file an existing mount at a propagated destination got
buried below the root of the nsfs file where no path walk reaches
it.
* Don't put a mountpoint on a dentry that's being removed.
attach_recursive_mnt() makes a mountpoint of the source's root
without its inode lock, so a racing rmdir() of that directory
could leave a mount on it that nothing ever detaches.
d_set_mounted() now checks cant_mount() as well.
- nullfs:
* Take no inode lock for readdir of an immutable directory.
The root of every empty mount namespace is the same nullfs
directory and iterate_dir() held its i_rwsem across
->iterate_shared(). A reader whose buffer faults on a FUSE mount
of its own holds it for as long as its server wants, and with an
exclusive locker queued behind it every lookup that misses the
dcache, every create and every mount in that directory waits. One
user of an empty mount namespace stalls all others. Directories
with the new FOP_IMMUTABLE flag skip the lock.
* Refuse to reconfigure internal superblocks through fspick(),
MS_REMOUNT or the read-only remount that a synchronous umount()
of the root does. For nullfs only root in the initial user
namespace could do it, but the superblock is shared by every
mount namespace and the flags showed up in statfs() for all of
them.
* Don't update the access time on nullfs and refuse F_SET_RW_HINT
on an immutable inode.
- unshare: Free an nsproxy that was never installed with
nsproxy_free() when set_cred_ucounts() fails. put_nsproxy() dropped
active references that were never taken, which triggered a warning
and hid the caller's own namespaces from listns().
- Smaller changes: mount_setattr() checks the target before it walks
the tree to allocate peer group ids, unshare() puts the old
fs_struct before the old namespaces, dissolve_on_fput() drops the
file's reference to the tree itself, disconnect_mount() is
simplified and the documentation of the propagated unmount rule is
brought up to date"
* tag 'vfs-7.3-rc7.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (58 commits)
namespace: simplify disconnect_mount()
selftests/filesystems: test covered mounts
namespace: rework connected mounts
nullfs: add an empty immutable regular file
selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody
selftests/filesystems: add a helper that holds a readdir in a page fault
readdir: take no inode lock on an immutable directory
nullfs: refuse file locks
fsnotify: let a filesystem refuse marks on its objects
namespace: nothing is mounted on or written through knullfs
namespace: keep the private nullfs instance in knullfs
fhandle: decide the subtree check under mount_lock
selftests/filesystems: check that an automount below an overlay layer is refused
selftests/filesystems: check the atime of the empty mount namespace root
selftests/filesystems: check that a lock lands on the right mount and stays
namespace: keep the lock on a mount that a propagated copy is moved beneath
namespace: never expire a locked mount
nullfs: don't update the access time
namespace: handle mount locking for automounts correctly
namespace: refuse an automount below a mount that is in no namespace
...
Diffstat (limited to 'include/linux/mount.h')
| -rw-r--r-- | include/linux/mount.h | 2 |
1 files changed, 1 insertions, 1 deletions
diff --git a/include/linux/mount.h b/include/linux/mount.h index acfe7ef86a1b..b16ba5d45605 100644 --- a/include/linux/mount.h +++ b/include/linux/mount.h @@ -52,7 +52,7 @@ enum mount_flags { MNT_ATIME_MASK = MNT_NOATIME | MNT_NODIRATIME | MNT_RELATIME, MNT_INTERNAL_FLAGS = MNT_INTERNAL | MNT_DOOMED | - MNT_SYNC_UMOUNT | MNT_LOCKED + MNT_SYNC_UMOUNT | MNT_LOCKED | MNT_UMOUNT }; struct vfsmount { |
