aboutsummaryrefslogtreecommitdiff
path: root/kernel
AgeCommit message (Collapse)AuthorFilesLines
20 hoursMerge tag 'dma-mapping-7.3-2026-10-09' of ↵Linus Torvalds1-3/+3
git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux Pull dma-mapping fixes from Marek Szyprowski: "Two more fixes for the corner cases in the DMA-mapping SWIOTLB code (Peng Fan and Marek Szyprowski)" * tag 'dma-mapping-7.3-2026-10-09' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux: swiotlb: fix default_swiotlb_limit() for non-growable default pool iommu/dma: skip swiotlb bounce for DMA_ATTR_MMIO in iommu_dma_map_phys
21 hoursMerge tag 'vfs-7.3-rc7.fixes' of ↵Linus Torvalds2-5/+9
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs Pull vfs fixes from Christian Brauner: "This contains fixes for the current development cycle. All of them came out of a review of the mount code that started with a bug report. The review modeled the corner cases of mount propagation, unmounting and mount reference counting and turned up a lot of bugs. Most of them years old. Most fixes come with a selftest. - Rework connected mounts. A mount that is unmounted together with its parent can stay attached to the parent to keep its mountpoint covered. That happens when the mountpoint is removed with rmdir(), unlink() or rename(), when a detached tree is dissolved, and for locked mounts in any umount that isn't synchronous, including the teardown of their mount namespace. The parent then owns the child and drops it on its own final mntput(). So any reference from the child's superblock back to one of its ancestors becomes a cycle that is never freed. A loop device backed by an image on a tmpfs and mounted on that same tmpfs is enough. Remove the directory the tmpfs is mounted on from the host, let the container's mount namespace exit, and the loop device, the tmpfs and the filesystem on the loop device are leaked for good. The same works with autofs, zram, ecryptfs, binfmt_misc, fuse passthrough, zloop, a mass storage gadget and md, and the selftests have reproducers for them. This has been possible since v4.1. It is also why "put_mnt_ns(): leave mounts connected" was reverted in -rc5. Keeping every mount of a dying mount namespace connected made these cycles trivial to create. Every unmounted mount is now detached from its parent. Where the mountpoint has to stay covered the mount leaves a cover on the parent instead, allocated together with the mount. A lookup on the unmounted parent that hits a cover finds an empty immutable directory or file on the private nullfs instance. Nothing leads from a cover to another mount, so no unmounted mount owns another one and no cycle can form. This is visible to userspace. A formerly connected mount can no longer be reached through its unmounted parent and ".." inside it leads nowhere, as for every other lazily unmounted mount. With that the private nullfs instance becomes reachable from userspace, so it now refuses mounts on top, is mounted read-only, and refuses fsnotify marks and file locks. Its inodes are shared by every holder and a watch or a lock would otherwise reach across users. may_decode_fh() now also decides its subtree check under a single mount_lock hold, as a racing umount could otherwise let it decode into what a locked child covered. - umount: * Don't silently unmount busy mounts. Since v4.13 propagate_umount() takes down propagated copies of the victim with children as long as each child is an overmount or another copy of the victim, but propagate_mount_busy() only ever checked copies without children or with just an overmount. A container that moved a tree beneath its copy of a host mount lost that tree from under its open file descriptor to a plain umount() on the host. propagate_mount_busy() now applies the same rules, walking each chain of copies once. * Don't let a migrating task hide its reference from umount(). mnt_get_count() sums the per-cpu counters under mount_lock but the mntget() and mntput() fast paths don't take it. A task that takes a reference on a cpu the sum has already passed and drops it after migrating to one the sum hasn't reached yet hides the reference it held to begin with, and umount() succeeds with the file still open. Gets and puts now live in separate per-cpu counters and all puts are summed before all gets with a full barrier in between, the way srcu_readers_active_idx_check() does it. mntget() is unchanged and mntput() gains an smp_wmb(). * Check each submount for references right before unmounting it. shrink_submounts() and mark_mounts_for_expiry() checked all their victims up front. Unmounting the first could move a busy overmount to where the next victim's propagated copy is looked up and it was then unmounted without a check. * Never expire a locked mount. A shrinkable mount moved beneath a locked mount with MOVE_MOUNT_BENEATH takes over the lock, and umount() of an unlocked ancestor expired it and revealed what it covered. That umount() now fails with EBUSY as it does for any other locked child. A lazy umount still takes the whole tree. - Overmounts and locked mounts: * Unhash a dentry before detaching the mounts on it. unlink(), rmdir() and rename() detach the mounts on the victim but only d_delete() it once its inode is unlocked, a window that includes an expedited RCU grace period. In between, a lookup from a mount namespace in which the dentry is a mountpoint found it hashed, positive and uncovered. Drop the dentry first, as d_invalidate() already does. * Don't reveal overmounted entries in refwalk. A refwalk that had grabbed the dentry before the unlink never rechecked it the way rcuwalk does with d_seq and mount_lock. Without any artificial widening three walkers read the covered file 27 times in a minute. step_into() now fails an unhashed dentry marked DCACHE_CANT_MOUNT with -ESTALE and the walk is retried. * Keep covered mounts covered in OPEN_TREE_NAMESPACE. Creating such a mount namespace only takes a user namespace and the copy followed bind mount rules: no children without AT_RECURSIVE and no unbindable mounts with it. An unprivileged user could see what mounts covered in the source, such as the parts of /proc and /sys that container runtimes mask. If the caller doesn't own the source mount namespace a non-recursive copy of a mount with something mounted below the requested directory is now refused and a recursive copy includes unbindable mounts, the way unshare() copies. * Keep the lock on a mount that a propagated copy is moved beneath. MNT_LOCKED moved to any mount that ended up beneath a locked mount, propagated copies included. A host mount and umount on a directory covered by a locked mount in a less privileged mount namespace left that cover unlocked for the namespace's owner to remove. Only mounts the caller places beneath take over the lock now. * Handle mount locking for automounts correctly. Which copies to lock was decided by the mount namespace of the task that triggered the automount. A task in a user namespace that triggered one on a host mount through a file descriptor got the host's own automount locked while its own copy stayed unlocked and could have nosuid, nodev and noexec cleared. Use the owner of the mount namespace the mount lands in. - Use-after-free and crashes: * Refuse an automount below a mount that is in no namespace. The private clones overlayfs uses for its layers have the MNT_NS_INTERNAL error pointer as their namespace, which finish_automount() let through and count_mounts() dereferenced. A fanotify filesystem mark on an overlayfs lower layer hands out file descriptors on such a clone. With debugfs as the lower layer opening "tracing" oopses with namespace_sem held for writing and every mount operation on the system blocks from then on. * Reset the old parent's ->overmount in mnt_change_mountpoint(). When propagate_umount() moved an overmount off a mount that a file descriptor kept alive, MOVE_MOUNT_BENEATH through that descriptor later followed the stale pointer into the freed overmount. * statmount() with STATMOUNT_BY_FD and pivot_root() read the parent of a mount that may be unmounted and only held by a file descriptor, while the parent's final mntput() can free it. statmount() now reads it under mount_lock and pivot_root() first checks that both mounts are in the caller's mount namespace. * Queue a mount only once for mount notifications. A mount reparented by one umount_tree() and taken down by the next under the same namespace_sem hold, as in shrink_submounts(), was queued twice. That cut the mounts queued in between out of notify_list while it still pointed at them, and once they were freed every later mount operation walked freed memory. * Don't let a pseudo dentry become the root of a mount. A bind mount of a bpf token file did that with a DCACHE_NORCU dentry, which is freed without an RCU grace period while lockless path walks may still look at it. Refuse to clone such a mount. * Don't inherit MNT_UMOUNT in clone_mnt(). A bind mount of a lazily unmounted nsfs or pidfs mount through its file descriptor started out flagged as unmounted. Among other things __detach_mounts() then dropped the namespace's reference on it, the mount outlived its namespace and mount_setattr() through the descriptor read the freed namespace. A recursive bind mount of such a mount also copied the unmounted stack still attached to it. That now fails with EINVAL, copying the mount itself still works. * Remove the fsnotify marks of a mount namespace in free_mnt_ns() instead of the RCU callback that frees the namespace, where taking the group mutexes meant sleeping in softirq context. - Propagation and copies: * Keep a copied mount unbindable. Since v6.17 clone_mnt() didn't copy the unbindable flag, so every mount namespace created with CLONE_NEWNS had bindable copies of all unbindable mounts. This had been fixed once before. * Refuse MOVE_MOUNT_SET_GROUP on an unbindable mount. It made the mount an unbindable slave, a state nothing else can produce, or silently dropped the unbindable flag. CRIU applies MS_UNBINDABLE after restoring sharing and isn't affected. * Check a recursive bind mount for mount namespace loops. Recursively bind mounting a tree from another mount namespace could put a mount of a namespace's file inside that same namespace, which then pins itself and all its mounts. Repeating it leaks without limit, the reproducer took Shmem from 380 kB to 65916 kB. The copy is now checked with check_for_nsfs_mounts() before it is grafted, as move_mount() does. * Look at the topmost mount for a mount namespace file. attach_recursive_mnt() never looked at the topmost mount of the source's chain of overmounts. If that was the chain's only mount namespace file an existing mount at a propagated destination got buried below the root of the nsfs file where no path walk reaches it. * Don't put a mountpoint on a dentry that's being removed. attach_recursive_mnt() makes a mountpoint of the source's root without its inode lock, so a racing rmdir() of that directory could leave a mount on it that nothing ever detaches. d_set_mounted() now checks cant_mount() as well. - nullfs: * Take no inode lock for readdir of an immutable directory. The root of every empty mount namespace is the same nullfs directory and iterate_dir() held its i_rwsem across ->iterate_shared(). A reader whose buffer faults on a FUSE mount of its own holds it for as long as its server wants, and with an exclusive locker queued behind it every lookup that misses the dcache, every create and every mount in that directory waits. One user of an empty mount namespace stalls all others. Directories with the new FOP_IMMUTABLE flag skip the lock. * Refuse to reconfigure internal superblocks through fspick(), MS_REMOUNT or the read-only remount that a synchronous umount() of the root does. For nullfs only root in the initial user namespace could do it, but the superblock is shared by every mount namespace and the flags showed up in statfs() for all of them. * Don't update the access time on nullfs and refuse F_SET_RW_HINT on an immutable inode. - unshare: Free an nsproxy that was never installed with nsproxy_free() when set_cred_ucounts() fails. put_nsproxy() dropped active references that were never taken, which triggered a warning and hid the caller's own namespaces from listns(). - Smaller changes: mount_setattr() checks the target before it walks the tree to allocate peer group ids, unshare() puts the old fs_struct before the old namespaces, dissolve_on_fput() drops the file's reference to the tree itself, disconnect_mount() is simplified and the documentation of the propagated unmount rule is brought up to date" * tag 'vfs-7.3-rc7.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (58 commits) namespace: simplify disconnect_mount() selftests/filesystems: test covered mounts namespace: rework connected mounts nullfs: add an empty immutable regular file selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody selftests/filesystems: add a helper that holds a readdir in a page fault readdir: take no inode lock on an immutable directory nullfs: refuse file locks fsnotify: let a filesystem refuse marks on its objects namespace: nothing is mounted on or written through knullfs namespace: keep the private nullfs instance in knullfs fhandle: decide the subtree check under mount_lock selftests/filesystems: check that an automount below an overlay layer is refused selftests/filesystems: check the atime of the empty mount namespace root selftests/filesystems: check that a lock lands on the right mount and stays namespace: keep the lock on a mount that a propagated copy is moved beneath namespace: never expire a locked mount nullfs: don't update the access time namespace: handle mount locking for automounts correctly namespace: refuse an automount below a mount that is in no namespace ...
2 daysswiotlb: fix default_swiotlb_limit() for non-growable default poolMarek Szyprowski1-3/+3
Commit ad96ce3252db ("swiotlb: determine potential physical address limit") changed default_swiotlb_limit() to return the maximum physical address that could ever be used by dynamically allocated pools, so that the value stays constant when CONFIG_SWIOTLB_DYNAMIC=y. However, io_tlb_default_mem.phys_limit is returned even when the default pool cannot grow, e.g. when swiotlb_init_remap() is called with a remap callback. This breaks Xen dom0 on x86: pci_xen_swiotlb_init() passes SWIOTLB_ANY and xen_swiotlb_fixup(), so phys_limit is set to the top of dom0 memory, although xen_swiotlb_fixup() makes the default pool below the 32-bit boundary. xen_swiotlb_dma_supported() then rejects 32-bit DMA masks, even though the bounce buffer is perfectly usable by such devices. Return phys_limit only if the default pool can actually grow, otherwise report the real end of the default pool, like in the !CONFIG_SWIOTLB_DYNAMIC case. Reported-by: Andreas Greve <andreas.greve@a-greve.de> Reported-by: James Dingwall <james@dingwall.me.uk> Fixes: ad96ce3252db ("swiotlb: determine potential physical address limit") Closes: https://lore.kernel.org/xen-devel/f74668db-52fd-4575-8372-7bfdf10d62ac@a-greve.de/ Cc: stable@vger.kernel.org Assisted-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Tested-by: James Dingwall <james@dingwall.me.uk> Link: https://lore.kernel.org/all/20261005070518.1590097-1-m.szyprowski@samsung.com/ Signed-off-by: Marek Szyprowski <m.szyprowski@samsung.com>
3 daysMerge tag 'urgent.2026.10.01a' of ↵Linus Torvalds1-1/+1
git://git.kernel.org/pub/scm/linux/kernel/git/rcu/linux Pull RCU fix from Paul McKenney: "Fix spurious WARN_ON() for rcu_segcblist_n_cbs() in cleanup_srcu_struct() This issue was introduced by 78a38cbf6f20 ("srcu: Queue sdp->work when the delay timer is successfully deleted") during this merge window. Enough people are hitting this that I am sending it now rather than waiting for the next merge window. Especially given that it is a simple one-liner" * tag 'urgent.2026.10.01a' of git://git.kernel.org/pub/scm/linux/kernel/git/rcu/linux: srcu: Fix WARN_ON() for rcu_segcblist_n_cbs() in cleanup_srcu_struct()
4 daysMerge tag 'wq-for-7.3-rc6-fixes' of ↵Linus Torvalds1-1/+1
git://git.kernel.org/pub/scm/linux/kernel/git/tj/wq Pull workqueue fixes from Tejun Heo: - Fix a NULL dereference in the chained work check when a kworker queues work on a draining or destroying workqueue outside work item execution. * tag 'wq-for-7.3-rc6-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/wq: workqueue: Fix NULL current_pwq deref in chained work check
4 daysMerge tag 'cgroup-for-7.3-rc6-fixes' of ↵Linus Torvalds1-8/+32
git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup Pull cgroup fixes from Tejun Heo: - During CPU offline, the active mask drops the CPU before cpuset updates the effective CPUs, so a task placement in that window could find no active CPU in the top cpuset and dereference NULL. Restore the NULL check. - The cpuset v2-mode test read the subsystem's root pointer, which is stale during a cgroup filesystem rebind, and the hotplug handler evaluated it before taking the cpuset mutex. Record the mode in a flag and test it under the mutex. * tag 'cgroup-for-7.3-rc6-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup: cgroup/cpuset: Call is_in_v2_mode() after acquiring cpuset_mutex in cpuset_handle_hotplug() cgroup/cpuset: Handle cpu hotplug race in guarantee_active_cpus() cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode()
4 daysMerge tag 'sched_ext-for-7.3-rc6-fixes' of ↵Linus Torvalds5-17/+52
git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext Pull sched_ext fixes from Tejun Heo: - Taking a CPU offline could hang, or stall until the watchdog ejected the BPF scheduler, when tasks on the dying CPU were still held by the scheduler or sitting on a user dispatch queue. Re-enqueue them onto the local queue when the runqueue goes offline so that the CPU pushes them off like the other sched classes. - The sequence number guarding against stale dispatches was per runqueue, so a task re-enqueued on another CPU could get the same number and a dispatch meant for its earlier instance was applied to the new one. Use a per-task counter. - A task dispatched to another CPU's local queue got its ops.dequeue() only when picked to run and flagged as a core-sched pick. Call it at insertion like for same-CPU dispatches. * tag 'sched_ext-for-7.3-rc6-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext: sched_ext: Generate qseq from a per-task counter selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ moves sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ sched_ext: Fix CPU hotplug hang when a dying CPU's tasks sit in the BPF scheduler
4 dayscgroup/cpuset: Call is_in_v2_mode() after acquiring cpuset_mutex in ↵Waiman Long1-4/+5
cpuset_handle_hotplug() It is reported by sashiko that calling is_in_v2_mode() outside of cpuset_mutex critical section in cpuset_handle_hotplug() can introduce a TOCTOU race where cgroup hierarchy may have changed from v1 to v2 or vice versa after is_in_v2_mode() is called leading to erroneously skip the allocation of tmpmasks or incorrectly modify cpus_allowed masks, resulting in cpuset state corruption. Fix that by calling is_in_v2_mode() after acquiring the cpuset_mutex. Fixes: b8d1b8ee93df ("cpuset: Allow v2 behavior in v1 cgroup") Signed-off-by: Waiman Long <longman@redhat.com> Signed-off-by: Tejun Heo <tj@kernel.org>
4 daysMerge tag 'printk-for-7.3-rc7' of ↵Linus Torvalds1-0/+79
git://git.kernel.org/pub/scm/linux/kernel/git/printk/linux Pull printk fix from Petr Mladek: - Allow using Braille console with a serial console driver converted to NBCON API * tag 'printk-for-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/printk/linux: braille: nbcon: Allow to use a serial console with NBCON API as Braille console
6 daysMerge tag 'perf-urgent-2026-10-04' of ↵Linus Torvalds2-185/+137
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull perf events fixes from Ingo Molnar: - Fix race between perf_event_exit_task() and perf_pending_task() (Luo Gengkun) - Fix perf header output management regressions (Ian Rogers) - Require kernel access for text poke events (Zhengchuan Liang) * tag 'perf-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: perf: Require kernel access for text poke events perf: Replace perf_event_header__init_id with full header init perf: Fix race between perf_event_exit_task() and perf_pending_task()
6 daysMerge tag 'locking-urgent-2026-10-04' of ↵Linus Torvalds3-5/+19
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull locking fixes from Ingo Molnar: - Don't run refcount kunit self-test when !CONFIG_KUNIT_ALL_TESTS (Kuan-Wei Chiu) - Fix futex private hash use-after-free on resize (Chris Mason) * tag 'locking-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: futex: Fix private hash use-after-free on resize irq: Make refcount_interrupt kunit test selectable
7 dayssched_ext: Generate qseq from a per-task counterKuba Piecuch3-5/+15
finish_dispatch() uses the qseq embedded in p->scx.ops_state to tell whether the QUEUED instance of a task it's about to claim is the one scx_bpf_dsq_insert() saw. qseq is generated from rq->scx.ops_qseq, but the counters of different rqs are independent, so if a task is dequeued and re-enqueued on a different rq between scx_bpf_dsq_insert() and finish_dispatch(), the new QUEUED instance can end up with the same qseq as the old one: CPU X CPU Z ----- ----- enqueue p on rq A, qseq = N ops.dispatch() scx_bpf_dsq_insert(p) records qseq N sched_setaffinity(p) dequeue p from rq A enqueue p on rq B, qseq = N finish_dispatch(p, N) qseq matches, p is claimed The claim itself is still atomic so the core stays consistent, but an insert issued for a previous QUEUED instance gets applied to a new one which the BPF scheduler has just received through ops.enqueue(). This breaks the guarantee that dispatches targeting a stale instance are ignored. Generate qseq from a per-task counter, p->scx.ops_qseq, instead so that consecutive QUEUED instances of a task never share a qseq regardless of which rq they're on. The counter is only updated in scx_do_enqueue_task() with the task's rq locked, so no additional synchronization is needed, and it fits in an existing hole in struct sched_ext_entity on 64bit. Remove the now unused rq->scx.ops_qseq. Never generate qseq 0. NONE and DISPATCHING don't carry a qseq, so scx_bpf_dsq_insert() on a task in either state records 0. With a per-task counter, every task's first QUEUED instance would otherwise get qseq 0 and could be claimed by such an insert. Wrap the counter where the QSEQ field wraps so that it can't reach a value that shifts to 0 on 32bit either. Fixes: f0e1a0643a59 ("sched_ext: Implement BPF extensible scheduler class") Cc: stable@vger.kernel.org # v6.12+ Assisted-by: Claude:claude-opus-5.5 Signed-off-by: Kuba Piecuch <jpiecuch@google.com> Signed-off-by: Tejun Heo <tj@kernel.org>
8 daysMerge tag 'probes-fixes-v7.3-rc5' of ↵Linus Torvalds2-13/+21
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace Pull probes fixes from Masami Hiramatsu: - fprobe: Use guard(rcu_sched_notrace) and check rcu_is_watching() Exit handlers early if !rcu_is_watching() to prevent potential use-after-free during unregistration in idle/quiescent states. Switch to guard(rcu_sched_notrace) to avoid fast-path lockdep overhead and recursion while ensuring safe grace period synchronization. - kprobes: Skip disarmed probes when checking optkprobe overlap Continue past disarmed or unprepared probes in get_optimized_kprobe() to find active optimized probes. This avoids overwriting active jump displacements which can lead to panic. * tag 'probes-fixes-v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: kprobes: Skip disarmed probes when checking optkprobe overlap fprobe: Use guard(rcu_sched_notrace) and check rcu_is_watching()
8 daysMerge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpfLinus Torvalds6-32/+168
Pull bpf fixes from Alexei Starovoitov: - Fix overflow of backward jump offset in constant blinding (Alexei Starovoitov) - Fix packet range of packet pointers sharing an id when var_off tightens umax of one pointer and not the other (Alexei Starovoitov) - Fix objects stuck in free_by_rcu_ttrace list of bpf memalloc (Alexei Starovoitov) - Fix use-after-free of progs detached from busy trampolines: wait for an RCU tasks grace period before freeing trampoline progs, and patch detached progs out of trampoline images that are still in use (Florent Revest) - Hold map BTF for the memory allocator destructor record to fix UAF in deferred bpf_mem_alloc destruction (Kumar Kartikeya Dwivedi) - Fix missing migration protection in resizable hashtab lookup_and_delete batch operation (Ömer Mete Kaya) * tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch() selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace bpf: Fix objects stuck in free_by_rcu_ttrace bpf: Factor out __do_call_rcu_ttrace() selftests/bpf: Test packet range of pointers sharing an id bpf: Fix packet range of pointers sharing an id selftests/bpf: Detach a trampoline prog while a task sleeps before it bpf: Skip detached progs in trampoline images that are still in use bpf: Wait for an RCU tasks grace period before freeing trampoline progs bpf: Hold map BTF for the memory allocator destructor record bpf: Fix overflow of jump offset in constant blinding
8 dayscgroup/cpuset: Handle cpu hotplug race in guarantee_active_cpus()Waiman Long1-2/+18
With commit 2125c0034c5d ("cgroup/cpuset: Make cpuset hotplug processing synchronous"), the cpuset hotplug operation becomes synchronous. That commit also removes the code that handles race between cpuset_hotplug_work and cpu hotplug notifier with the assumption that race is now gone. Later commit 7a0aabd9ce69 ("cgroup/cpuset: Always use cpu_active_mask") updates the cpuset code to always use cpu_active_mask instead of cpu_online_mask in some places including guarantee_online_cpus() which is renamed to guarantee_active_cpus() in that commit. In the case of CPU offline operation, cpu_active_mask is updated first in sched_cpu_deactivate() to remove the offline CPU before cpuset_handle_hotplug() is called to update the effective_cpus of the affected cpusets. The cpu_online_mask is updated after that near the end of the offline operation to remove the offline CPU. As a result, the race comes back and the top cpuset may not have any active CPU leading to NULL pointer dereference during the race window when guarantee_active_cpus() is called after cpu_active_mask is updated to remove the CPU to be torn down but before cpuset_handle_hotplug() is able to properly update the effective_cpus of the top cpuset. Fix this by adding back the NULL cs check to avoid this problem. Fixes: 7a0aabd9ce69 ("cgroup/cpuset: Always use cpu_active_mask") Cc: stable@vger.kernel.org # v6.16+ Reported-by: Farhad Alemi <farhad.alemi@berkeley.edu> Link: https://lore.kernel.org/lkml/CA+0ovChh3VjsKN1g+ZGjwwY2fGTpP7uD+aCCByLj5Qbymw=bfQ@mail.gmail.com Signed-off-by: Waiman Long <longman@redhat.com> Tested-by: Farhad Alemi <farhad.alemi@berkeley.edu> Signed-off-by: Tejun Heo <tj@kernel.org>
8 daysbpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch()Ömer Mete Kaya1-0/+2
bpf_mem_cache_free_rcu() uses this_cpu_ptr() which requires migration to be disabled. All callers of rhtab_delete_elem() disable migration except __rhtab_map_lookup_and_delete_batch(), which calls it under rcu_read_lock() only. On CONFIG_PREEMPT_RCU, rcu_read_lock() does not disable preemption or migration, so the task can migrate between CPUs during the delete loop, causing this_cpu_ptr() to trigger: BUG: using smp_processor_id() in preemptible [00000000] code Fix by wrapping the delete loop in migrate_disable()/migrate_enable() in __rhtab_map_lookup_and_delete_batch(), matching the migration protection that the other callers already provide. Fixes: 818e00848227 ("bpf: Implement iteration ops for resizable hashtab") Reported-by: syzbot+fd7e415d891073b83e1f@syzkaller.appspotmail.com Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Mykyta Yatsenko <yatsenko@meta.com> Link: https://patch.msgid.link/20260929081609.557899-1-omermetekaya0@gmail.com Closes: https://syzkaller.appspot.com/bug?extid=fd7e415d891073b83e1f
8 daysfutex: Fix private hash use-after-free on resizeChris Mason1-4/+6
poll_state_synchronize_rcu(mm->futex.phash.batches) is used by futex_ref_drop() to check that a grace period has passed since the current hash was published. This relies on batches referencing a grace period which started after the hash pointer was assigned. __futex_pivot_hash() sets mmph->batches before it replaces mmph->hash: scoped_guard(rcu) { mmph->batches = get_state_synchronize_rcu(); rcu_assign_pointer(mmph->hash, new); } The scoped_guard(rcu) doesn't stop new grace periods from starting, and if one starts between those two assignments, futex_ref_drop() can move forward while a reader still holds a pointer to the old hash. Fix things by setting mmph->batches after assigning mmph->hash. The scoped_guard(rcu) isn't needed, so let's drop that as well. Fixes: 56180dd20c19 ("futex: Use RCU-based per-CPU reference counting instead of rcuref_t") Assisted-by: kres Signed-off-by: Chris Mason <mason@kernel.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: Paul E. McKenney <paulmck@kernel.org> Link: https://patch.msgid.link/20261001135022.2220288-1-mason@kernel.org
8 daysirq: Make refcount_interrupt kunit test selectableKuan-Wei Chiu2-1/+13
Currently, refcount_interrupt_test is built unconditionally when CONFIG_KUNIT is enabled, causing it to run unexpectedly during boot. Fix this by introducing CONFIG_REFCOUNT_INTERRUPT_KUNIT_TEST so the test can be configured independently, following standard kunit practices. Fixes: 07a88e2bcd5b ("irq: Add KUnit test for refcounted interrupt enable/disable") Signed-off-by: Kuan-Wei Chiu <visitorckw@gmail.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: Lyude Paul <lyude@redhat.com> Reviewed-by: Radu Rendec <radu@rendec.net> Reviewed-by: Boqun Feng <boqun@kernel.org> Link: https://patch.msgid.link/20261001082519.16195-2-boqun@kernel.org
8 daysbraille: nbcon: Allow to use a serial console with NBCON API as Braille consolePetr Mladek1-0/+79
The Braille console is not registered in console_list. Instead, it is integrated with the virtual terminal (VT) and shows what is displayed on the terminal. It writes the data using con->write*() callback of the associated serial console driver. braille_write() is called from the VT code under console_lock(). The associated serial console driver can be converted to the NBCON API though. The situation is similar to flushing nbcon consoles in the legacy loop when some boot consoles are still registered, see nbcon_legacy_emit_next_record(). But there is a big difference though. braille_write() is not directly called from the code paths flushing registered consoles. The VT code expects that braille_write() succeeds. It does not replay the message when it can't acquire the ownership. As a result, braille_write(): + must try harder to get the ownership. + has to be synchronized only against non-printk serial console which depend nbcon_device_try_acquire() using NBCON_PRIO_NORMAL. Let's look at it from another side and try to simulate the original locking using the NBCON API: 1. Take con->device_lock(), aka the port->lock in the legacy serial console driver. 2. Acquire nbcon context to provide some synchronization for a panic() context. Use NBCON_PRIO_NORMAL because it contends only with nbcon_device_try_acquire() users. Do it in a busy loop. It should always succeed when con->device_lock() succeeded because all other users do the same. The only exception is when the context get acquired by a CPU handling panic. 3. In panic, disable interrupts and try to acquire the nbcon context. Use NBCON_PRIO_PANIC. And try even an unsafe takeover because otherwise the Braille console won't see the text shown during panic(). It is similar to the oops_in_progress/trylock handling in the legacy serial console driver. Finally, avoid the newline prepending logic in the existing serial console drivers when they are used as a Braille console. As explained above, the Braille console shows the last modified line on the terminal (VT). braille_write() is called when single characters are added. Most messages are not ended by newline. Anyway, the VT code does not have logic to reply partially printed messages. Fixes: 13189fa73afa ("printk: nbcon: Rely on kthreads for normal operation") Reviewed-by: John Ogness <john.ogness@linutronix.de> [pmladek@suse.com: Remove white space changes in braille_register_console(). Typo fix.] Link: https://patch.msgid.link/20261001140727.124398-2-pmladek@suse.com Signed-off-by: Petr Mladek <pmladek@suse.com>
8 dayskprobes: Skip disarmed probes when checking optkprobe overlapLeon Hwang1-3/+5
On x86, an optkprobe at A replaces five bytes with a jump. If a disabled probe B is at A+2, get_optimized_kprobe() stops at B when arming a new probe C at A+4. It leaves A optimized: A A+1 A+2 A+3 A+4 A's jump | e9 | d0 | d1 | d2 | d3 | after C | e9 | d0 | d1 | d2 | cc | The INT3 for C overwrites the last byte of A's jump displacement, so execution can jump to the wrong address. B can have prepared optinsns while disarmed, but has no jump to unoptimize. Continue past disarmed and unprepared probes to find the active optimized probe before arming a probe in its jump. Link: https://lore.kernel.org/all/20260930053618.104498-1-leon.hwang@linux.dev/ Fixes: afd66255b9a4 ("kprobes: Introduce kprobes jump optimization") Cc: stable@vger.kernel.org Signed-off-by: Leon Hwang <leon.hwang@linux.dev> Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
9 dayssrcu: Fix WARN_ON() for rcu_segcblist_n_cbs() in cleanup_srcu_struct()Sunho Park1-1/+1
The WARN_ON() added by commit 78a38cbf6f20 ("srcu: Queue sdp->work when the delay timer is successfully deleted") uses rcu_segcblist_n_cbs() to detect callbacks that srcu_barrier() failed to wait for. However, the ->len counter is decremented only at the end of srcu_invoke_callbacks(), after the invoking loop has finished. Since srcu_barrier() can return right after the barrier callback is invoked, cleanup_srcu_struct() can see a non-zero n_cbs even though the cblist is already physically empty, falsely triggering the WARN_ON() together with a still-pending delay_work timer. This can be triggered as follows, as seen in the syzbot report against kvm_destroy_vm() -> cleanup_srcu_struct(&kvm->srcu): 1. call_srcu(&kvm->srcu, &bus->rcu, __free_bus) starts SRCU grace period GP1. 2. GP1 ends: a delay timer is armed and sdp->work is queued, but sdp->work has not run yet. 3. Another call_srcu(&kvm->srcu, &bus->rcu, __free_bus) call invokes srcu_segcblist_advance(), which moves the GP1 callback to RCU_DONE_TAIL, and starts SRCU grace period GP2. 4. srcu_barrier() is called. It queues its barrier callback after the GP2 callback and waits for srcu_invoke_callbacks() to invoke it. 5. GP2 ends: another delay timer is armed, and the sdp->work queued in step 2 begins to run. Its srcu_invoke_callbacks() call invokes srcu_segcblist_advance() again, moving the GP2 and barrier callbacks to RCU_DONE_TAIL as well, and then invokes all of them. However, rcu_segcblist_add_len(), which updates srcu_cblist's ->len, has not run yet at this point. 6. srcu_barrier() returns once its callback has been invoked, and cleanup_srcu_struct() starts running. It finds the delay timer armed in step 5 still pending and srcu_cblist's ->len still non-zero (because step 5 has not reached rcu_segcblist_add_len() yet), and WARN_ON() fires even though every callback has actually been invoked. Use rcu_segcblist_empty(), which checks the actual head of the cblist, instead of rcu_segcblist_n_cbs(), which checks the racy ->len counter. Callbacks that have genuinely not been invoked yet still leave the list non-empty, so the WARN_ON() still catches callers that skip srcu_barrier() or queue callbacks after it. Link: https://lore.kernel.org/rcu/e6350377085ddd85d6ef00d8e9a67bd50c762d3c@linux.dev/T/#t Reported-by: syzbot+d4faf7db59e11f6fd1ab@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=d4faf7db59e11f6fd1ab Fixes: 78a38cbf6f20 ("srcu: Queue sdp->work when the delay timer is successfully deleted") Suggested-by: Zqiang <qiang.zhang@linux.dev> Signed-off-by: Sunho Park <shpark061104@gmail.com> Signed-off-by: Paul E. McKenney <paulmck@kernel.org> Tested-by: Sean Christopherson <seanjc@google.com>
9 daysbpf: Fix objects stuck in free_by_rcu_ttraceAlexei Starovoitov1-2/+32
do_call_rcu_ttrace() returns early when call_rcu_ttrace_in_progress is set and leaves the objects in free_by_rcu_ttrace. __free_rcu() frees waiting_for_gp_ttrace only and clears the flag. Hence the objects that free_bulk() or __free_by_rcu() added while RCU tasks trace GP was in flight stay in free_by_rcu_ttrace until free_bulk() or alloc_bulk() is called for the same bpf_mem_cache again, which may never happen. The number of such objects is not bounded. Turn call_rcu_ttrace_in_progress into three states: 0 - idle 1 - __free_rcu() is queued 2 - __free_rcu() is queued and free_by_rcu_ttrace got more objects since do_call_rcu_ttrace() sets 2. __free_rcu() does cmpxchg(1 -> 0) and starts the next GP when it fails. It cannot clear the flag first and check free_by_rcu_ttrace later, since bpf_mem_alloc_destroy() frees bpf_mem_cache without waiting for RCU callbacks when the flag is zero. Now __free_rcu() queues itself, so the one that didn't see 'draining' may do call_rcu_tasks_trace() after rcu_barrier_tasks_trace() in free_mem_alloc(). Queue it under rcu_read_lock() and do synchronize_rcu() before the barriers. Calling rcu_barrier_tasks_trace() twice works too, but creating and destroying hash maps in a loop on many cpus slows down to one free_mem_alloc() per GP and kworkers pile up. Fixes: 8d5a8011b35d ("bpf: Batch call_rcu callbacks instead of SLAB_TYPESAFE_BY_RCU.") Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20260930095920.601738-3-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
9 daysbpf: Factor out __do_call_rcu_ttrace()Alexei Starovoitov1-10/+17
Move the part of do_call_rcu_ttrace() that runs after call_rcu_ttrace_in_progress is set into __do_call_rcu_ttrace(). The next patch will call it from __free_rcu(). No functional change. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20260930095920.601738-2-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
9 daysbpf: Fix packet range of pointers sharing an idAlexei Starovoitov1-1/+9
Since commit 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers"), find_good_pkt_pointers() sets the range of all packet pointers sharing an id from the umax of the compared pointer, and check_packet_access() requires umax + off + size <= range. That assumes the umax of two such pointers differ by exactly their constant distance. reg_bounds_sync() breaks it when var_off tightens one umax and not the other: r4 &= 0x38 if r4 > 50 goto exit ; umax 50, var_off (0x0; 0x38) r5 = pkt + r4 ; umax 50 r6 = r5 r6 += 8 ; umax 56, not 58 Comparing r6 with pkt_end sets the range to 56, and the valid 8-byte load at r5 is rejected (50 + 8 > 56). Comparing r5 sets it to 50, and the out-of-bounds 1-byte load at r6 - 7, i.e. r5 + 1, is accepted (56 - 7 + 1 <= 50). Don't call reg_bounds_sync() on a packet pointer that keeps its id (a constant was added or subtracted) or its range (an unknown non-negative value was subtracted), so that var_off cannot tighten its umax. Only update the 32-bit bounds from var_off: reg_bounds_sanity_check() wants them constant when the lower half of var_off is, e.g. for pkt + 8. This relies on nothing else changing the 64-bit bounds of a packet pointer, which holds today. var_off of such a pointer is no longer narrowed by its bounds. Adjust three verifier_align expectations; the low bits, which the alignment checks use, don't change. veristat on the selftests shows no verdict changes and +0.8% insns in test_cls_redirect_subprogs. Fixes: 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers") Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20261001145255.855630-1-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
9 daysMerge tag 'sysctl-7.03-fixes-rc6' of ↵Linus Torvalds2-1/+3
git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl Pull sysctl fixes from Joel Granados: "Fix sysctl jiffies conversions errors introduced in 2dc164a48e6f ("sysctl: Create converter functions with two new macros") - Ensure that mult_hz does *not* wrap - Ensure we pass just the magnitude for the negative branch in proc_int_k2u_conv_kop" * tag 'sysctl-7.03-fixes-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl: time/jiffies: Saturate in mult_hz() instead of wrapping sysctl: Negate before converting in the int read path
9 daysunshare: don't drop active namespace references that were never takenChristian Brauner2-2/+3
Active references on the namespaces of an nsproxy are taken when the nsproxy is installed into a task in switch_task_namespaces() and copy_namespaces() and dropped again by put_nsproxy() through deactivate_nsproxy(). But ksys_unshare() calls put_nsproxy() on an nsproxy that was never installed when set_cred_ucounts() fails. The new namespaces of that nsproxy go from zero to minus one and the namespaces shared with the caller lose a reference that belongs to the nsproxy the caller keeps using: WARNING: kernel/nscommon.c:171 at __ns_ref_active_put+0x1cd/0x230 nsproxy_ns_active_put deactivate_nsproxy ksys_unshare Afterwards the namespaces the caller lives in aren't listed by listns() anymore. set_cred_ucounts() only fails when alloc_ucounts() can't allocate and it only allocates when the real uid of the caller differs from its effective uid. Commit cefd55bd2159 ("nsproxy: fix free_nsproxy() and simplify create_new_namespaces()") separated the two cases on purpose. nsproxy_free() frees an nsproxy that was prepared but never installed and that's what a failed setns() uses in put_nsset(). Export it and use it for a failed unshare() as well. Fixes: a98621a0f187 ("unshare: fix nsproxy leak in ksys_unshare() on set_cred_ucounts() failure") Cc: stable@vger.kernel.org # v7.1+ Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-16-be34c83956ae@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
9 daysperf: Require kernel access for text poke eventsZhengchuan Liang1-1/+1
Perf events with exclude_kernel=1 can be opened without kernel perf access. However, exclude_kernel does not suppress text-poke sideband records. Every PERF_RECORD_TEXT_POKE is marked PERF_RECORD_MISC_KERNEL and contains a raw kernel instruction address. An unprivileged task can therefore open and mmap a task-local software event with text_poke=1. Both opening a count-only tracepoint event and configuring UDP GRO for ESP-in-UDP cause updates to inline static calls; the observer receives the relocated addresses of the modified instructions. For a known kernel image, any such address reveals the runtime kernel text base despite KASLR. Call perf_allow_kernel() whenever attr.text_poke is set, regardless of exclude_kernel. Events that neither monitor kernel execution nor request text-poke records retain their existing permissions. Fixes: e17d43b93e54 ("perf: Add perf text poke event") Assisted-by: LLM Signed-off-by: Zhengchuan Liang <zcliangcn@gmail.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/99131354c41e23188f778b92f90363775b482395.1790573390.git.zcliangcn@gmail.com
9 daysperf: Replace perf_event_header__init_id with full header initIan Rogers2-182/+132
perf_iterate_sb() invokes its callback for each matching perf_event on the CPU and task context, passing a shared caller-allocated event structure. perf_event_header__init_id() mutated header->size in place by adding event->id_header_size, requiring sideband output callbacks to save and restore header fields across iterations. Three sideband callbacks failed to save and restore header.size around perf_event_header__init_id(): - perf_event_ksymbol_output() - perf_event_bpf_output() - perf_event_text_poke_output() When multiple events with attr.ksymbol, attr.bpf_event, or attr.text_poke and sample_id_all are active on the same CPU, each subsequent event receives a record whose header.size is inflated by all preceding events' id_header_size values while only a single id_sample is written, leaving uninitialized ring-buffer bytes at the end of the record and causing userspace perf to fail with -EFAULT ("Bad address") when parsing the sample_id trailer. Similarly, perf_event_mmap_output() set PERF_RECORD_MISC_MMAP_BUILD_ID in mmap_event->event_id.header.misc when event->attr.build_id was enabled, but only saved and restored header.size and header.type. If an event with attr.build_id was followed by an event with attr.mmap2 and !attr.build_id, the second event received PERF_RECORD_MISC_MMAP_BUILD_ID in header.misc while its payload contained maj/min/ino/ino_generation instead of a build ID. Rather than splitting header initialization between callers and output callbacks and saving/restoring mutated header fields, replace perf_event_header__init_id() with perf_event_header__init(), which initializes header->type, header->misc, and header->size alongside the sample_id fields on each invocation. Fixes: 76193a94522f ("perf, bpf: Introduce PERF_RECORD_KSYMBOL") Fixes: 6ee52e2a3fe4 ("perf, bpf: Introduce PERF_RECORD_BPF_EVENT") Fixes: e17d43b93e54 ("perf: Add perf text poke event") Fixes: 88a16a130933 ("perf: Add build id data in mmap2 event") Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260929222332.973435-1-irogers@google.com Cc: stable@vger.kernel.org
9 daysperf: Fix race between perf_event_exit_task() and perf_pending_task()Luo Gengkun1-2/+4
A race condition exists between perf_event_exit_task() and perf_pending_task() during begin_new_exec(). During begin_new_exec(), perf_event_exit_task() may be called, and the PF_EXITING flag is not set on task. So perf_sigtrap() continues to execute and triggers WARN_ON_ONCE(event->ctx->task != current). Since both task exit and exec paths can call perf_event_exit_task() which sets ctx->task to TASK_TOMBSTONE, fix this by explicitly checking if event->ctx->task equals TASK_TOMBSTONE and dropping the redundant PF_EXITING check. Fixes: 97ba62b27867 ("perf: Add support for SIGTRAP on perf events") Signed-off-by: Luo Gengkun <luogengkun2@huawei.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260920075026.990582-1-luogengkun2@huawei.com
10 daysMerge tag 'audit-pr-20260930' of ↵Linus Torvalds1-4/+20