aboutsummaryrefslogtreecommitdiff
AgeCommit message (Collapse)AuthorFilesLines
14 hoursMerge tag 'pci-v7.3-fixes-4' of ↵HEADmasterLinus Torvalds1-6/+16
git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci Pull PCI fix from Bjorn Helgaas: "This fixes some GPU initialization regressions caused by eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors"), which appeared in v7.3-rc1. That commit also caused a MacBookPro16,1 spontaneous power-off regression; I expect a fix for that next week" * tag 'pci-v7.3-fixes-4' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci: PCI/AER: Skip error recovery on false alarms
14 hoursMerge tag 'mmc-v7.3-rc1-2' of ↵Linus Torvalds6-28/+26
git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/mmc Pull MMC/MEMSTICK fixes from Ulf Hansson: "MMC host: - cavium-octeon|thunderx: Destroy slot platform devices on remove - mtk-sd: Cancel request timeout work on remove - sdhci-sprd: Disable runtime PM on remove MEMSTICK: - Wait for request completion before freeing card - rtsx_usb_ms: Complete requests after eject instead of dropping them" * tag 'mmc-v7.3-rc1-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/mmc: memstick: rtsx_usb_ms: complete requests after eject instead of dropping them memstick: core: wait for request completion before freeing card mmc: cavium-thunderx: destroy slot platform devices on remove mmc: cavium-octeon: destroy slot platform devices on remove mmc: sdhci-sprd: disable runtime PM on remove mmc: mtk-sd: Cancel request timeout work on remove
14 hoursMerge tag 'media/v7.3-3' of ↵Linus Torvalds1-2/+2
git://git.kernel.org/pub/scm/linux/kernel/git/mchehab/linux-media Pull media fix from Mauro Carvalho Chehab: "A fix for em28xx unregister code affecting devices with FM radio support" * tag 'media/v7.3-3' of git://git.kernel.org/pub/scm/linux/kernel/git/mchehab/linux-media: media: em28xx: use video_unregister_device for radio_dev
14 hoursMerge tag 'sound-7.3-rc7' of ↵Linus Torvalds13-14/+161
git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound Pull sound fixes from Takashi Iwai: "A dozen of small fixes. All device-specific quirks or fixes, and nothing exciting is expected. - USB-audio and HD-audio quirks - HD-audio TAS2781 codec fix - ctxfi driver memory leak fix - ASoC AMD quirks - ASoC cs35l56 and adau1372 codec fixes" * tag 'sound-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound: ASoC: amd: acp: Add more ACP7.0 match entries for Cirrus Logic parts ASoC: amd: acp: Add DMI override for ASUS EXPERTBOOK AM7406CKA ASoC: sdw_utils: Switch CS35L56/57/62/63 to use OT25 DAI ASoC: cs35l56: Add DAI for SDCA OT25 stream ASoC: samsung: i2s: enable op_clk in probe to balance runtime PM ASoC: adau1372: fix invalid default state of output ASRC and DAC muxes ASoC: adau1372: allow sleepy powerdown GPIOs ASoC: amd: yc: Add DMI quirk for Acer Aspire Go 15 (AG15-21P) ALSA: hda/realtek: Add mute LED quirk for HP Pavilion 15-eg3xxx (103c:8bfd) ALSA: hda/tas2781: Accept V1 calibration data without CRC ALSA: ctxfi: Fix memory leakage when clearing DAO input mappers ALSA: hda/realtek: Add quirk for ASRock NUC BOX-358H ALSA: usb-audio: Add native DSD support for Cambridge Audio streamers
14 hoursMerge tag 'pmdomain-v7.3-rc2' of ↵Linus Torvalds2-3/+32
git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm Pull pmdomain provider fixes from Ulf Hansson: - imx: Serialize power on/off across sibling domains for imx8m-blk-ctrl - rockchip: Fix a couple of errors during probe * tag 'pmdomain-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm: pmdomain: rockchip: don't ignore clock lookup errors on attach pmdomain: rockchip: fix clock leak on domain probe failure pmdomain: rockchip: propagate subdomain add errors pmdomain: imx8m-blk-ctrl: Serialize power on/off across sibling domains
14 hoursMerge tag 'dma-mapping-7.3-2026-10-09' of ↵Linus Torvalds2-5/+5
git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux Pull dma-mapping fixes from Marek Szyprowski: "Two more fixes for the corner cases in the DMA-mapping SWIOTLB code (Peng Fan and Marek Szyprowski)" * tag 'dma-mapping-7.3-2026-10-09' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux: swiotlb: fix default_swiotlb_limit() for non-growable default pool iommu/dma: skip swiotlb bounce for DMA_ATTR_MMIO in iommu_dma_map_phys
15 hoursMerge tag 'fsverity-for-linus' of git://git.kernel.org/pub/scm/fs/fsverity/linuxLinus Torvalds2-2/+26
Pull fsverity fix from Eric Biggers: "Fix a regression from commit f77f281b6118 ("fsverity: use a hashtable to find the fsverity_info")" * tag 'fsverity-for-linus' of git://git.kernel.org/pub/scm/fs/fsverity/linux: fsverity: RCU-delay the freeing of struct fsverity_info
16 hoursMerge tag 'vfs-7.3-rc7.fixes' of ↵Linus Torvalds53-195/+6462
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs Pull vfs fixes from Christian Brauner: "This contains fixes for the current development cycle. All of them came out of a review of the mount code that started with a bug report. The review modeled the corner cases of mount propagation, unmounting and mount reference counting and turned up a lot of bugs. Most of them years old. Most fixes come with a selftest. - Rework connected mounts. A mount that is unmounted together with its parent can stay attached to the parent to keep its mountpoint covered. That happens when the mountpoint is removed with rmdir(), unlink() or rename(), when a detached tree is dissolved, and for locked mounts in any umount that isn't synchronous, including the teardown of their mount namespace. The parent then owns the child and drops it on its own final mntput(). So any reference from the child's superblock back to one of its ancestors becomes a cycle that is never freed. A loop device backed by an image on a tmpfs and mounted on that same tmpfs is enough. Remove the directory the tmpfs is mounted on from the host, let the container's mount namespace exit, and the loop device, the tmpfs and the filesystem on the loop device are leaked for good. The same works with autofs, zram, ecryptfs, binfmt_misc, fuse passthrough, zloop, a mass storage gadget and md, and the selftests have reproducers for them. This has been possible since v4.1. It is also why "put_mnt_ns(): leave mounts connected" was reverted in -rc5. Keeping every mount of a dying mount namespace connected made these cycles trivial to create. Every unmounted mount is now detached from its parent. Where the mountpoint has to stay covered the mount leaves a cover on the parent instead, allocated together with the mount. A lookup on the unmounted parent that hits a cover finds an empty immutable directory or file on the private nullfs instance. Nothing leads from a cover to another mount, so no unmounted mount owns another one and no cycle can form. This is visible to userspace. A formerly connected mount can no longer be reached through its unmounted parent and ".." inside it leads nowhere, as for every other lazily unmounted mount. With that the private nullfs instance becomes reachable from userspace, so it now refuses mounts on top, is mounted read-only, and refuses fsnotify marks and file locks. Its inodes are shared by every holder and a watch or a lock would otherwise reach across users. may_decode_fh() now also decides its subtree check under a single mount_lock hold, as a racing umount could otherwise let it decode into what a locked child covered. - umount: * Don't silently unmount busy mounts. Since v4.13 propagate_umount() takes down propagated copies of the victim with children as long as each child is an overmount or another copy of the victim, but propagate_mount_busy() only ever checked copies without children or with just an overmount. A container that moved a tree beneath its copy of a host mount lost that tree from under its open file descriptor to a plain umount() on the host. propagate_mount_busy() now applies the same rules, walking each chain of copies once. * Don't let a migrating task hide its reference from umount(). mnt_get_count() sums the per-cpu counters under mount_lock but the mntget() and mntput() fast paths don't take it. A task that takes a reference on a cpu the sum has already passed and drops it after migrating to one the sum hasn't reached yet hides the reference it held to begin with, and umount() succeeds with the file still open. Gets and puts now live in separate per-cpu counters and all puts are summed before all gets with a full barrier in between, the way srcu_readers_active_idx_check() does it. mntget() is unchanged and mntput() gains an smp_wmb(). * Check each submount for references right before unmounting it. shrink_submounts() and mark_mounts_for_expiry() checked all their victims up front. Unmounting the first could move a busy overmount to where the next victim's propagated copy is looked up and it was then unmounted without a check. * Never expire a locked mount. A shrinkable mount moved beneath a locked mount with MOVE_MOUNT_BENEATH takes over the lock, and umount() of an unlocked ancestor expired it and revealed what it covered. That umount() now fails with EBUSY as it does for any other locked child. A lazy umount still takes the whole tree. - Overmounts and locked mounts: * Unhash a dentry before detaching the mounts on it. unlink(), rmdir() and rename() detach the mounts on the victim but only d_delete() it once its inode is unlocked, a window that includes an expedited RCU grace period. In between, a lookup from a mount namespace in which the dentry is a mountpoint found it hashed, positive and uncovered. Drop the dentry first, as d_invalidate() already does. * Don't reveal overmounted entries in refwalk. A refwalk that had grabbed the dentry before the unlink never rechecked it the way rcuwalk does with d_seq and mount_lock. Without any artificial widening three walkers read the covered file 27 times in a minute. step_into() now fails an unhashed dentry marked DCACHE_CANT_MOUNT with -ESTALE and the walk is retried. * Keep covered mounts covered in OPEN_TREE_NAMESPACE. Creating such a mount namespace only takes a user namespace and the copy followed bind mount rules: no children without AT_RECURSIVE and no unbindable mounts with it. An unprivileged user could see what mounts covered in the source, such as the parts of /proc and /sys that container runtimes mask. If the caller doesn't own the source mount namespace a non-recursive copy of a mount with something mounted below the requested directory is now refused and a recursive copy includes unbindable mounts, the way unshare() copies. * Keep the lock on a mount that a propagated copy is moved beneath. MNT_LOCKED moved to any mount that ended up beneath a locked mount, propagated copies included. A host mount and umount on a directory covered by a locked mount in a less privileged mount namespace left that cover unlocked for the namespace's owner to remove. Only mounts the caller places beneath take over the lock now. * Handle mount locking for automounts correctly. Which copies to lock was decided by the mount namespace of the task that triggered the automount. A task in a user namespace that triggered one on a host mount through a file descriptor got the host's own automount locked while its own copy stayed unlocked and could have nosuid, nodev and noexec cleared. Use the owner of the mount namespace the mount lands in. - Use-after-free and crashes: * Refuse an automount below a mount that is in no namespace. The private clones overlayfs uses for its layers have the MNT_NS_INTERNAL error pointer as their namespace, which finish_automount() let through and count_mounts() dereferenced. A fanotify filesystem mark on an overlayfs lower layer hands out file descriptors on such a clone. With debugfs as the lower layer opening "tracing" oopses with namespace_sem held for writing and every mount operation on the system blocks from then on. * Reset the old parent's ->overmount in mnt_change_mountpoint(). When propagate_umount() moved an overmount off a mount that a file descriptor kept alive, MOVE_MOUNT_BENEATH through that descriptor later followed the stale pointer into the freed overmount. * statmount() with STATMOUNT_BY_FD and pivot_root() read the parent of a mount that may be unmounted and only held by a file descriptor, while the parent's final mntput() can free it. statmount() now reads it under mount_lock and pivot_root() first checks that both mounts are in the caller's mount namespace. * Queue a mount only once for mount notifications. A mount reparented by one umount_tree() and taken down by the next under the same namespace_sem hold, as in shrink_submounts(), was queued twice. That cut the mounts queued in between out of notify_list while it still pointed at them, and once they were freed every later mount operation walked freed memory. * Don't let a pseudo dentry become the root of a mount. A bind mount of a bpf token file did that with a DCACHE_NORCU dentry, which is freed without an RCU grace period while lockless path walks may still look at it. Refuse to clone such a mount. * Don't inherit MNT_UMOUNT in clone_mnt(). A bind mount of a lazily unmounted nsfs or pidfs mount through its file descriptor started out flagged as unmounted. Among other things __detach_mounts() then dropped the namespace's reference on it, the mount outlived its namespace and mount_setattr() through the descriptor read the freed namespace. A recursive bind mount of such a mount also copied the unmounted stack still attached to it. That now fails with EINVAL, copying the mount itself still works. * Remove the fsnotify marks of a mount namespace in free_mnt_ns() instead of the RCU callback that frees the namespace, where taking the group mutexes meant sleeping in softirq context. - Propagation and copies: * Keep a copied mount unbindable. Since v6.17 clone_mnt() didn't copy the unbindable flag, so every mount namespace created with CLONE_NEWNS had bindable copies of all unbindable mounts. This had been fixed once before. * Refuse MOVE_MOUNT_SET_GROUP on an unbindable mount. It made the mount an unbindable slave, a state nothing else can produce, or silently dropped the unbindable flag. CRIU applies MS_UNBINDABLE after restoring sharing and isn't affected. * Check a recursive bind mount for mount namespace loops. Recursively bind mounting a tree from another mount namespace could put a mount of a namespace's file inside that same namespace, which then pins itself and all its mounts. Repeating it leaks without limit, the reproducer took Shmem from 380 kB to 65916 kB. The copy is now checked with check_for_nsfs_mounts() before it is grafted, as move_mount() does. * Look at the topmost mount for a mount namespace file. attach_recursive_mnt() never looked at the topmost mount of the source's chain of overmounts. If that was the chain's only mount namespace file an existing mount at a propagated destination got buried below the root of the nsfs file where no path walk reaches it. * Don't put a mountpoint on a dentry that's being removed. attach_recursive_mnt() makes a mountpoint of the source's root without its inode lock, so a racing rmdir() of that directory could leave a mount on it that nothing ever detaches. d_set_mounted() now checks cant_mount() as well. - nullfs: * Take no inode lock for readdir of an immutable directory. The root of every empty mount namespace is the same nullfs directory and iterate_dir() held its i_rwsem across ->iterate_shared(). A reader whose buffer faults on a FUSE mount of its own holds it for as long as its server wants, and with an exclusive locker queued behind it every lookup that misses the dcache, every create and every mount in that directory waits. One user of an empty mount namespace stalls all others. Directories with the new FOP_IMMUTABLE flag skip the lock. * Refuse to reconfigure internal superblocks through fspick(), MS_REMOUNT or the read-only remount that a synchronous umount() of the root does. For nullfs only root in the initial user namespace could do it, but the superblock is shared by every mount namespace and the flags showed up in statfs() for all of them. * Don't update the access time on nullfs and refuse F_SET_RW_HINT on an immutable inode. - unshare: Free an nsproxy that was never installed with nsproxy_free() when set_cred_ucounts() fails. put_nsproxy() dropped active references that were never taken, which triggered a warning and hid the caller's own namespaces from listns(). - Smaller changes: mount_setattr() checks the target before it walks the tree to allocate peer group ids, unshare() puts the old fs_struct before the old namespaces, dissolve_on_fput() drops the file's reference to the tree itself, disconnect_mount() is simplified and the documentation of the propagated unmount rule is brought up to date" * tag 'vfs-7.3-rc7.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (58 commits) namespace: simplify disconnect_mount() selftests/filesystems: test covered mounts namespace: rework connected mounts nullfs: add an empty immutable regular file selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody selftests/filesystems: add a helper that holds a readdir in a page fault readdir: take no inode lock on an immutable directory nullfs: refuse file locks fsnotify: let a filesystem refuse marks on its objects namespace: nothing is mounted on or written through knullfs namespace: keep the private nullfs instance in knullfs fhandle: decide the subtree check under mount_lock selftests/filesystems: check that an automount below an overlay layer is refused selftests/filesystems: check the atime of the empty mount namespace root selftests/filesystems: check that a lock lands on the right mount and stays namespace: keep the lock on a mount that a propagated copy is moved beneath namespace: never expire a locked mount nullfs: don't update the access time namespace: handle mount locking for automounts correctly namespace: refuse an automount below a mount that is in no namespace ...
16 hoursMerge tag 'arm64-fixes' of ↵Linus Torvalds2-2/+6
git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux Pull arm64 fixes from Will Deacon: "It finally seems to have calmed down on the arm64 fixes front, so please pull these two straightforward fixes for -rc7. One fixes the EL2 trap configuration for implementation-defined CPU PMU hardware during boot and the other fixes a kcov selftest failure by excluding our softirq early entry code: - Fix PMU EL2 trap configuration for CPUs with an IMPDEF PMU - Fix kcov boot selftest failure by excluding our early IRQ entry code" * tag 'arm64-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux: arm64: irq: exclude the softirq stack switch from KCOV arm64/boot: Don't set PMUv3p9 FGT2 bits without PMUv3
32 hoursMerge tag 'net-7.3-rc7' of ↵Linus Torvalds149-813/+2063
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Including fixes from wireless, wireguard, CAN and Bluetooth. We have one known regression to wrap up in VLAN handling. Current release - regressions: - Bluetooth: RFCOMM: fix deadlock on rfcomm_mutex Previous releases - regressions: - can: fix regression in handling RPS after migrating metadata to skb_ext - eth: - iavf: fix regressions in reconfig impacting bonding - mana: fix packet forwarding performance regression - stmmac: remove buggy VLAN acceleration support Previous releases - always broken: - a few high prio fixes for tun, and af_packet - amt: fix a UaF on tunnel teardown - eth: - bnxt: fix PCIe AER recovery and FLR handling issues - macb: don't modify Tx skbs before taking ownership - axienet: don't leak Tx skbs on interface stop - wifi: - nxpwifi: number of LLM-ish fixes - assorted mt76 fixes" * tag 'net-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (128 commits) net: macb: copy shared skbs before appending the FCS net: macb: check TX ring before modifying skb vsock: Fix memory leak in vmci_transport_recv_dgram_cb() wireguard: noise: reject response consumption after intermediate initiation wireguard: queueing: preserve tstamp_type when encapsulating packet net: openvswitch: validate transport header presence in set_ipv6_addr net/smc: protect clcsock lifetime in smc_getname ipv6: do not warn on route notification size race ipv4: do not warn on route notification size race ipv4: validate checksum_start before completing checksum ptp: ocp: fix PCIe delay estimation calculation xen/netfront: don't leak the skb when xennet_fill_frags() fails net/packet: call packet_parse_headers after virtio_net_hdr_to_skb xen/netfront: drop RX packets with a short Ethernet header net: skbuff: don't leave stale bytes in skb_copy_and_csum_bits() net: sparx5: free the matchall entry on destroy selftests: mlxsw: Test port range occupancy on template create mlxsw: spectrum_flower: Fix port range register leak in tmplt_create() net: dsa: microchip: fix KSZ8765 fiber detection net/mlx5e: Order ICOSQ cc update after CQ doorbell ...
42 hoursPCI/AER: Skip error recovery on false alarmsLukas Wunner1-6/+16
Alex is seeing a probe failure of the amdgpu driver after the Root Port above an AMD Navi10 GPU has been reset. The reset was performed to recover from a Firmware First reported Fatal Error. However all status registers in the Root Port's AER Extended Capability are blank, so apparently the platform firmware raised a false alarm. The issue is only occurring since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors"). It looks like enabling Advisory Non-Fatal Errors causes code paths to be exercised in platform firmware which were never validated before. Skip error recovery on false alarms, i.e. if no unmasked errors were actually signaled. Note that this will also skip recovery if both the Status and Mask registers are "all ones", as would be the case for inaccessible devices. However that seems justified because it would imply either a hot-unplug event or a Surprise Down Error further up in the hierarchy. Interfering with recovery from that seems uncalled for. Fixes: eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors") Reported-by: Alex Deucher <alexander.deucher@amd.com> Closes: https://bugzilla.kernel.org/show_bug.cgi?id=222095 Signed-off-by: Lukas Wunner <lukas@wunner.de> Signed-off-by: Bjorn Helgaas <bhelgaas@google.com> Tested-by: Alex Deucher <alexander.deucher@amd.com> Acked-by: Alex Deucher <alexander.deucher@amd.com> Link: https://patch.msgid.link/0552ed277e40a288e0157af799257ee6ec722534.1791460615.git.lukas@wunner.de
42 hoursMerge branch 'net-macb-fix-software-fcs-handling-of-shared-and-requeued-skbs'Jakub Kicinski1-31/+49
Nicolai Buchwitz says: ==================== net: macb: fix software FCS handling of shared and requeued skbs While testing the genet MTU series I used a Raspberry Pi CM5 (RP1 GEM) as pktgen source for the CM4. With clone_skb the CM5 rebooted after a few seconds. Further investigation showed that macb_pad_and_fcs() appends the FCS in place, so the shared skb grows with every transmit until BQL completes more than was queued and dql_completed() hits its BUG_ON. The same code also modifies the skb before the TX ring check, so a NETDEV_TX_BUSY retry gets an skb that was already replaced or grown. Patch 1 checks the ring first, patch 2 copies shared skbs. Tested on Raspberry CM5 with pktgen at 60/20000 bytes, clone_skb 0/ 1000, burst 1/32. ==================== Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-0-a80641479041@tipi-net.de Signed-off-by: Jakub Kicinski <kuba@kernel.org>
42 hoursnet: macb: copy shared skbs before appending the FCSNicolai Buchwitz1-2/+4
macb_pad_and_fcs() appends the FCS in place when the skb has tailroom. A shared skb, as pktgen sends in clone_skb mode, grows by one FCS per transmit. BQL then completes more bytes than were queued and dql_completed() hits its BUG_ON. On a Raspberry Pi CM5 (RP1 GEM) pktgen with clone_skb 1000 burst 32 at 60 bytes kills the box within seconds. Copy shared skbs before appending the FCS. Clearing IFF_TX_SKB_SHARING would also fix it but makes pktgen refuse clone_skb on macb. Fixes: 653e92a9175e ("net: macb: add support for padding and fcs computation") Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de> Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-2-a80641479041@tipi-net.de Signed-off-by: Jakub Kicinski <kuba@kernel.org>
42 hoursnet: macb: check TX ring before modifying skbNicolai Buchwitz1-30/+46
macb_pad_and_fcs() replaces or extends the skb before the ring space check. On NETDEV_TX_BUSY the stack requeues an skb that is already freed or grown. Check the ring first, using the padded length for the descriptor count. Nonlinear skbs always take the copy path so the count can assume a linear skb. Fixes: 653e92a9175e ("net: macb: add support for padding and fcs computation") Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de> Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-1-a80641479041@tipi-net.de Signed-off-by: Jakub Kicinski <kuba@kernel.org>
42 hoursvsock: Fix memory leak in vmci_transport_recv_dgram_cb()Ilia Gavrilov1-0/+2
During the closure of a datagram socket, the vmci_transport_recv_dgram_cb() function may be called, which will add the packets to the socket's backlog; then, after the receive queue is cleared, the __release_sock() function will move the packet back from the socket's backlog to the receive queue, which will lead to a memory leak. sock_close __sock_release __vsock_release // take ownership by user-space lock_sock_nested sock_set_flag(sk, SOCK_DEAD) vmci_transport_release vmci_dispatch_dgs vmci_datagram_invoke_guest_handler vmci_transport_recv_dgram_cb sk_receive_skb if (!sock_owned_by_user(sk)) [...] (!) else if sk_add_backlog() [...] (!) skb_queue_purge(&sk->sk_receive_queue) release_sock __release_sock // move all packets from the backlog // to the receive queue (!) sk_backlog_rcv vsock_queue_rcv_skb [...] __sock_queue_rcv_skb if (!sock_flag(sk, SOCK_DEAD)) // Since the socket is "dead", the function will not // be called sk->sk_data_ready(sk) sock_release_ownership(sk); sock_put Add a receive queue cleanup in the socket destructor to fix this. syzkaller report: unreferenced object 0xffff8880117bfb80 (size 240): comm "irq/56-vmw_vmci", pid 205, jiffies 4295064448 (age 73.992s) hex dump (first 32 bytes): 78 dc 26 1e 80 88 ff ff 78 dc 26 1e 80 88 ff ff x.&.....x.&..... 00 00 00 00 00 00 00 00 c0 da 26 1e 80 88 ff ff ..........&..... backtrace: [<ffffffff82be8d97>] __alloc_skb+0x287/0x330 net/core/skbuff.c:505 [<ffffffffa114bcbd>] vmci_transport_recv_dgram_cb+0xbd/0x210 [vmw_vsock_vmci_transport] [<ffffffffa05072c6>] vmci_datagram_invoke_guest_handler+0x366/0x450 [vmw_vmci] [<ffffffffa050ad75>] vmci_dispatch_dgs+0x235/0x490 [vmw_vmci] [<ffffffffa050b025>] vmci_interrupt+0x55/0x290 [vmw_vmci] [<ffffffff8137ad2b>] irq_thread_fn+0x8b/0x1a0 kernel/irq/manage.c:1205 [<ffffffff8137c7ae>] irq_thread+0x28e/0x530 kernel/irq/manage.c:1314 [<ffffffff81253dd6>] kthread+0x2e6/0x3a0 kernel/kthread.c:376 [<ffffffff81004392>] ret_from_fork+0x22/0x30 arch/x86/entry/entry_64.S:295 Found by InfoTeCS on behalf of Linux Verification Center (linuxtesting.org) with Syzkaller. Fixes: d021c344051a ("VSOCK: Introduce VM Sockets") Reviewed-by: Stefano Garzarella <sgarzare@redhat.com> Signed-off-by: Ilia Gavrilov <Ilia.Gavrilov@infotecs.ru> Link: https://patch.msgid.link/20261008144046.90853-1-Ilia.Gavrilov@infotecs.ru Signed-off-by: Jakub Kicinski <kuba@kernel.org>
42 hoursMerge branch 'wireguard-fixes-for-7-3-rc7'Jakub Kicinski2-3/+6
Jason A. Donenfeld says: ==================== WireGuard fixes for 7.3-rc7 This series contains two important WireGuard fixes 1) Stop zeroing out skb->tstamp_type when encapsulating packets, so that fq behaves correctly, from Ramses de Norre. 2) Make sure handshake state isn't swapped out while locks are released, reported by Jérémy Jean. ==================== Link: https://patch.msgid.link/20261008130124.724119-1-Jason@zx2c4.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
42 hourswireguard: noise: reject response consumption after intermediate initiationJason A. Donenfeld1-3/+4
Two threads begin processing the identical response message, received twice. The first thread, A, runs. While it's running, the second one, B, gets partway through, and during that slow calculation, or even while blocking on down_write(), A completes and then also a handshake initiation that's already been queued up runs in thread C, which itself takes that same down_write(). The handshake initiation creation succeeds, and sets the state back to waiting-for-response, and calls up_write(), at which point thread B resumes, because either its finished its calculations or was finally allowed to acquire down_write(). Thread B then copies the state back to the peer, and begins a new session, using that state, which is the same session as the one made in thread A. Thread A Thread B Thread C down_read() sA = handshake->state memcpy(cA, handshake->crypto) up_read() if (sA != 1) goto fail slow_crypto(cA) down_read() sB = handshake->state memcpy(cB, handshake->crypto) up_read() if (sB != 1) goto fail slow_crypto(cB) down_write() if (sA != handshake->state) goto fail memcpy(handshake->crypto, cA) handshake->state = 2 up_write() down_write() if (handshake->state != 2) goto fail derive_session(handshake->crypto) up_write() down_write() slow_crypto(handshake->crypto) handshake->state = 1 up_write() down_write() if (sB != handshake->state) goto fail memcpy(handshake->crypto, cB) handshake->state = 2 up_write() down_write() if (handshake->state != 2) goto fail derive_session(handshake->crypto) up_write() This seems basically impossible to hit in a meaningful way in practice, but ensure that it absolutely cannot happen by comparing the ephemeral private key that's on the stack with the latest one that the peer's handshake state has. Cc: stable@vger.kernel.org Fixes: e7096c131e51 ("net: WireGuard secure network tunnel") Reported-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr> Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com> Link: https://patch.msgid.link/20261008130124.724119-4-Jason@zx2c4.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
42 hourswireguard: queueing: preserve tstamp_type when encapsulating packetRamses de Norre1-0/+2
Sending traffic through a wireguard tunnel on a host using the fq qdisc fills the log with: fq: likely mono tstamp with tstamp_type 0 An skb carries a timestamp in skb->tstamp and, separately, a skb->tstamp_type field recording which clock that timestamp came from. The two have to agree. When wireguard encapsulates a packet it calls wg_reset_packet(), which clears the fields that must not leak from the inner packet into the tunnel packet. It does so in two steps: skb_scrub_packet(skb, true); memset(&skb->headers, 0, sizeof(skb->headers)); skb_scrub_packet() deliberately keeps skb->tstamp when it holds a monotonic timestamp: that value is the time the packet is scheduled to be sent, and the qdisc still needs it. The memset then zeroes skb->tstamp_type, because that field sits inside the headers group while skb->tstamp does not. The packet therefore leaves wireguard carrying a monotonic timestamp labelled as a realtime one. Nothing noticed until commit c4f796c4f16b ("net_sched: sch_fq: convert skb->tstamp if not monotonic"): fq used to assume every timestamp was monotonic. It now consults tstamp_type, spots the mismatch, warns, and falls back to treating the value as monotonic. Pacing still ends up correct, so the log spam is the actual problem. Save tstamp_type before the memset and restore it when encapsulating, next to the hash fields that are already carried over this way. When decapsulating it stays zeroed, which is right: an incoming packet's timestamp is a realtime receive timestamp. Fixes: d98d58a00261 ("net: Set skb->mono_delivery_time and clear it after sch_handle_ingress()") Signed-off-by: Ramses de Norre <ramses@well-founded.dev> Reviewed-by: Toke Høiland-Jørgensen <toke@kernel.org> Reviewed-by: Eric Dumazet <edumazet@google.com> Cc: stable@vger.kernel.org Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com> Link: https://patch.msgid.link/20261008130124.724119-3-Jason@zx2c4.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursnet: openvswitch: validate transport header presence in set_ipv6_addrFernando Fernandez Mancera1-1/+6
When executing IPv6 address rewrite actions on IPv6 fragments, set_ipv6_addr() calls update_ipv6_checksum(). If parse_ipv6hdr() processes a non-first IPv6 fragment, it sets key->ip.proto to NEXTHDR_FRAGMENT and returns early without calling skb_set_transport_header(). update_ipv6_checksum() unconditionally evaluates skb_transport_offset() on entry before checking l4_proto. Because skb->transport_header is uninitialized, this triggers a warning under CONFIG_DEBUG_NET=y although it is completely harmless. Fix this by returning early in update_ipv6_checksum() if l4_proto is NEXTHDR_FRAGMENT. This avoids reading the uninitialized transport offset for fragments while preserving the debug warning for any other protocol where the transport header is unexpectedly missing. See the syzbot trace: !skb_transport_header_was_set(skb) WARNING: ./include/linux/skbuff.h:3075 at skb_transport_header include/linux/skbuff.h:3075 [inline], CPU#1: syz-executor463/5635 WARNING: ./include/linux/skbuff.h:3075 at skb_transport_offset include/linux/skbuff.h:3250 [inline], CPU#1: syz-executor463/5635 WARNING: ./include/linux/skbuff.h:3075 at update_ipv6_checksum net/openvswitch/actions.c:361 [inline], CPU#1: syz-executor463/5635 WARNING: ./include/linux/skbuff.h:3075 at set_ipv6_addr+0x462/0x660 net/openvswitch/actions.c:399, CPU#1: syz-executor463/5635 [...] RIP: 0010:skb_transport_header include/linux/skbuff.h:3075 [inline] RIP: 0010:skb_transport_offset include/linux/skbuff.h:3250 [inline] RIP: 0010:update_ipv6_checksum net/openvswitch/actions.c:361 [inline] RIP: 0010:set_ipv6_addr+0x462/0x660 net/openvswitch/actions.c:399 [...] Call Trace: <TASK> set_ipv6 net/openvswitch/actions.c:531 [inline] do_execute_actions+0x557e/0x8600 net/openvswitch/actions.c:1366 ovs_execute_actions+0xde/0x520 net/openvswitch/actions.c:1592 ovs_packet_cmd_execute+0xb4f/0xf10 net/openvswitch/datapath.c:705 genl_family_rcv_msg_doit+0x233/0x340 net/netlink/genetlink.c:1114 genl_rcv_msg+0x614/0x7a0 net/netlink/genetlink.c:1209 netlink_rcv_skb+0x226/0x4a0 net/netlink/af_netlink.c:2572 genl_rcv+0x28/0x40 net/netlink/genetlink.c:1218 netlink_unicast+0x7bd/0x940 net/netlink/af_netlink.c:1361 netlink_sendmsg+0x813/0xb40 net/netlink/af_netlink.c:1916 sock_sendmsg_nosec+0x14e/0x190 net/socket.c:800 Reported-by: syzbot+4cc63fcfb3845e149969@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=4cc63fcfb3845e149969 Fixes: 3fdbd1ce11e5 ("openvswitch: add ipv6 'set' action") Signed-off-by: Fernando Fernandez Mancera <fmancera@suse.de> Reviewed-by: Ilya Maximets <i.maximets@ovn.org> Link: https://patch.msgid.link/20261008104135.4722-1-fmancera@suse.de Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursnet/smc: protect clcsock lifetime in smc_getnameChengfeng Ye1-1/+6
smc_getname() dereferences smc->clcsock without holding clcsock_release_lock. Link-group termination can release the CLC socket through smc_close_active_abort() while the SMC socket is still open, for example after shutdown(SHUT_WR). A getsockname() caller can load smc->clcsock, then the termination worker can clear the pointer and call sock_release() before the caller accesses clcsock->ops or invokes getname(). This causes a use-after-free; if the worker clears the pointer before the load, it causes a NULL dereference. The syscall's file reference keeps the SMC socket alive but does not prevent asynchronous release of its CLC socket. KASAN reported the following with test-only timing instrumentation: BUG: KASAN: slab-use-after-free in smc_getname+0x19e/0x1b0 Read of size 8 at addr ffff888109abb4e0 by task poc/103 Call Trace: smc_getname+0x19e/0x1b0 do_getsockname+0xe5/0x170 __sys_getsockname+0x8c/0x100 Allocated by task 95: sock_alloc_inode+0x1e/0x280 sock_alloc+0x3d/0x240 __sock_create+0x7e/0x430 smc_create+0x121/0x240 Freed by task 0: kmem_cache_free+0xcc/0x340 rcu_core+0x50a/0x1850 Last potentially related work creation: evict+0x446/0x6c0 smc_clcsock_release+0xa8/0xd0 smc_close_active_abort+0x26a/0x3a0 __smc_lgr_terminate.part.0+0x137/0x2e0 Hold clcsock_release_lock across the pointer check and the getname() callback to serialize with smc_clcsock_release(). Return -EBADF if the CLC socket has already been released, preserving the existing peer state check and the callback's return value otherwise. Fixes: b03faa1fafc8 ("net/smc: postpone release of clcsock") Cc: stable@vger.kernel.org Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Reviewed-by: Mahanta Jambigi <mjambigi@linux.ibm.com> Reviewed-by: Dust Li <dust.li@linux.alibaba.com> Link: https://patch.msgid.link/20261008084444.787449-1-nicoyip.dev@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursMerge branch 'ipv4-ipv6-do-not-warn-on-route-notification-size-races'Jakub Kicinski2-4/+0
Daehyeon Ko says: ==================== ipv4/ipv6: do not warn on route notification size races A nexthop group can grow between route notification sizing and filling. Both IPv4 and IPv6 can then legitimately return -EMSGSIZE, so remove the stale warnings while preserving their existing error paths. Patch 1 is unchanged from v1. Patch 2 adds the IPv6 counterpart requested by Ido. No new build or runtime test was run; both changes only delete the stale comment and WARN_ON(). ==================== Link: https://patch.msgid.link/cover.1791423190.git.4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursipv6: do not warn on route notification size raceDaehyeon Ko1-2/+0
fib6_info_hw_flags_set() sizes the skb with rt6_nlmsg_size() and later fills it with rt6_fill_node(). With nexthop compatibility mode enabled, a concurrent replacement can grow the group between these independent snapshots. rt6_fill_node() can then legitimately return -EMSGSIZE, so the warning does not prove a sizing bug. Remove the warning. The existing error path still frees the skb and reports the error to listeners. Fixes: 907eea486888 ("net: ipv6: Emit notification when fib hardware flags are changed") Reported-by: Ido Schimmel <idosch@nvidia.com> Link: https://lore.kernel.org/netdev/20261007114003.GA1011260@shredder/ Suggested-by: Ido Schimmel <idosch@nvidia.com> Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Link: https://patch.msgid.link/4012604727e5ba80b69ccd4d08796511397ddaef.1791423190.git.4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursipv4: do not warn on route notification size raceDaehyeon Ko1-2/+0
Commit 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed") added asynchronous route notifications for hardware flag changes. fib_alias_hw_flags_set() sizes the skb with fib_nlmsg_size() and later fills it with fib_dump_info() while holding only RCU. With nexthop compatibility mode enabled, a concurrent replacement can grow the group between these independent snapshots. fib_dump_info() can then legitimately return -EMSGSIZE, so the warning does not prove a sizing bug. Remove the warning. The existing error path still frees the skb and reports the error to listeners. Fixes: 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed") Reported-by: Ido Schimmel <idosch@nvidia.com> Link: https://lore.kernel.org/netdev/20261004082043.GB92032@shredder/ Suggested-by: Ido Schimmel <idosch@nvidia.com> Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Link: https://patch.msgid.link/b7e203fc5867651cb67ef84aa03e082f0f76008e.1791423190.git.4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursipv4: validate checksum_start before completing checksumMichael S. Tsirkin6-12/+44
If a packet with bad checksum metadata gets into the ipv4 stack, skb_checksum_help can corrupt the network header and cause a bunch of mischief. This was discovered and reported by Paulos, and has been reporoduced by others independently since. We really shouldn't allow such packets in, but as a defence in depth measure, let's also check before we complete the checksum. A more complete validation at input is forthcoming, but needs more work. Cc: stable@vger.kernel.org Fixes: f43798c27684 ("tun: Allow GSO using virtio_net_hdr") Fixes: bfd5f4a3d605 ("packet: Add GSO/csum offload support.") Reported-by: Paulos Yibelo <habte.yibelo@gmail.com> Closes: https://lore.kernel.org/netdev/20260922030310.8684-2-habte.yibelo@gmail.com/ Signed-off-by: Michael S. Tsirkin <mst@redhat.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/0a012b4923e189c4c593ef4f471e5ff0edbe9030.1791412497.git.mst@redhat.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursptp: ocp: fix PCIe delay estimation calculationVadim Fedorenko1-1/+1
The commit in fixes introduced a high cap for delayas U64_MAX value while ktime_t is actually s64. This is wrong cap as it becomes negative value and any comparison to a real delay will fail to update delay value. Use KTIME_MAX constant as correct max cap for PCIe delay. The issue was hit in production (a negative value is observed): # cat /sys/class/timecard/ocp0/ts_window_adjust -3 Fixes: aa05fe67bcd64 ("ptp: ocp: Improve PCIe delay estimation") Signed-off-by: Vadim Fedorenko <vadim.fedorenko@linux.dev> Reviewed-by: Daniel Machon <daniel.machon@microchip.com> Link: https://patch.msgid.link/20261007203359.417270-1-vadim.fedorenko@linux.dev Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursxen/netfront: don't leak the skb when xennet_fill_frags() failsJosef Bacik1-1/+3
When a response chain has more slots than fit in the skb's frags, xennet_fill_frags() returns an error and xennet_poll() jumps to its error path. That path moves what's left on tmpq to errq to be freed, but the skb being filled was already dequeued from tmpq, so it's never freed. Each chain that overflows leaks the skb and the pages attached to it as frags, and the backend decides how many slots it sends. Put the skb back on tmpq before taking the error path, like the xennet_set_skb_gso() failure just above it does. Fixes: ad4f15dc2c70 ("xen/netfront: don't bug in case of too many frags") Cc: stable@vger.kernel.org Signed-off-by: Josef Bacik <josef@toxicpanda.com> Reviewed-by: Juergen Gross <jgross@suse.com> Link: https://patch.msgid.link/20261007-b4-xen-netfront-fill-frags-leak-v1-1-a8a01ff9cd52@toxicpanda.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursnet/packet: call packet_parse_headers after virtio_net_hdr_to_skbWillem de Bruijn1-10/+15
SOCK_RAW packet sockets incorrectly drop VLAN-tagged GSO packets without VIRTIO_NET_HDR_F_NEEDS_CSUM. packet_parse_headers() sets skb->protocol to VLAN but advances skb->network_header past the VLAN tags. virtio_net_hdr_to_skb() then flow dissects these packets, which parses the network header as a VLAN tag, and fails. Call packet_parse_headers() after virtio_net_hdr_to_skb() instead. This matches tun_get_user() and tap_get_user() and thus simplifies overall complexity. Commit 01fdecc0480d ("net: packet: fix wrong transport_header when sending VLAN-tagged frame") fixed the same issue for the transport header probe in packet_parse_headers(). That probe is now skipped if virtio_net_hdr_to_skb() already set the transport header, avoiding a redundant dissection. Implementation details: - Do not set skb->protocol to the RX wildcard ETH_P_ALL. SOCK_RAW then derives it from the link layer header, as before. SOCK_DGRAM now leaves it 0, or derives it from gso_type for GSO with NEEDS_CSUM. - Keep virtio_net_hdr_set_proto() last, so that the link layer protocol takes priority over gso_type. Reported-by: Junnan Zhang <zhangjn_dev@163.com> Closes: https://lore.kernel.org/netdev/20260821085722.24036-1-zhangjn_dev@163.com/ Fixes: dfed913e8b55 ("net/af_packet: add VLAN support for AF_PACKET SOCK_RAW GSO") Signed-off-by: Willem de Bruijn <willemb@google.com> Acked-by: Michael S. Tsirkin <mst@redhat.com> Link: https://patch.msgid.link/20261007175816.3138556-1-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursxen/netfront: drop RX packets with a short Ethernet headerJosef Bacik1-2/+10
handle_incoming_queue() pulls pull_to bytes into the head before calling eth_type_trans(). pull_to is the length of the first RX slot, capped at RX_COPY_THRESHOLD, and that length comes from the backend. Nothing checks it against ETH_HLEN. If the first slot is shorter than ETH_HLEN and more slots follow, the head ends up shorter than an Ethernet header while skb->len is longer, and eth_type_trans() BUG()s in __skb_pull(). If the whole packet is shorter than ETH_HLEN, eth_type_trans() reads the header past the end of the data instead. Pull at least ETH_HLEN, and drop the packet if that fails, which also drops packets too short to hold an Ethernet header. This also checks the return value of the pull, which was ignored. Fixes: 0d160211965b ("xen: add virtual network device driver") Cc: stable@vger.kernel.org Signed-off-by: Josef Bacik <josef@toxicpanda.com> Link: https://patch.msgid.link/20261007-b4-xen-netfront-short-head-v1-1-12d7113a7e4e@toxicpanda.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursnet: skbuff: don't leave stale bytes in skb_copy_and_csum_bits()Josef Bacik1-1/+5
When skb_copy_and_csum_bits() reaches unreadable frags it returns 0 after copying only the linear part, and the rest of the caller's buffer is left as it was. The callers copy into a buffer that is about to go out on the wire: an ICMP error quoting the offending packet, or a driver's TX bounce buffer in skb_copy_and_csum_dev(). Neither buffer is zeroed beforehand, so whatever was in memory there gets sent. Zero the part of the buffer we didn't fill. The checksum usually won't match the data any more, so the receiver will usually drop the packet, but either way it no longer carries anything it shouldn't. Only zero for a positive @len, a negative one from a broken caller must not turn into a huge memset(). Fixes: 65249feb6b3d ("net: add support for skbs with unreadable frags") Cc: stable@vger.kernel.org Signed-off-by: Josef Bacik <josef@toxicpanda.com> Reviewed-by: Mina Almasry <almasrymina@google.com> Link: https://patch.msgid.link/20261007-b4-skb-copy-csum-stale-bytes-v1-1-adbbde033fb3@toxicpanda.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
43 hoursnet: sparx5: free the matchall entry on destroyDaniel Machon1-0/+1
sparx5_tc_matchall_replace() allocates a struct sparx5_mall_entry for every offloaded matchall filter and adds it to sparx5->mall_entries. sparx5_tc_matchall_destroy() removes the entry from the list, but never frees it, so the entry of every deleted mirror and goto matchall filter is leaked. Free the entry after unlinking it. The leak was discovered by an AI code review agent, and reproduced with kmemleak on a lan969x EV board (EV23X71A) by rep