| Age | Commit message (Collapse) | Author | Files | Lines |
|
git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci
Pull PCI fix from Bjorn Helgaas:
"This fixes some GPU initialization regressions caused by eddba19b8b5f
("PCI/AER: Support Advisory Non-Fatal Errors"), which appeared in
v7.3-rc1.
That commit also caused a MacBookPro16,1 spontaneous power-off
regression; I expect a fix for that next week"
* tag 'pci-v7.3-fixes-4' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci:
PCI/AER: Skip error recovery on false alarms
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/mmc
Pull MMC/MEMSTICK fixes from Ulf Hansson:
"MMC host:
- cavium-octeon|thunderx: Destroy slot platform devices on remove
- mtk-sd: Cancel request timeout work on remove
- sdhci-sprd: Disable runtime PM on remove
MEMSTICK:
- Wait for request completion before freeing card
- rtsx_usb_ms: Complete requests after eject instead of dropping
them"
* tag 'mmc-v7.3-rc1-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/mmc:
memstick: rtsx_usb_ms: complete requests after eject instead of dropping them
memstick: core: wait for request completion before freeing card
mmc: cavium-thunderx: destroy slot platform devices on remove
mmc: cavium-octeon: destroy slot platform devices on remove
mmc: sdhci-sprd: disable runtime PM on remove
mmc: mtk-sd: Cancel request timeout work on remove
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mchehab/linux-media
Pull media fix from Mauro Carvalho Chehab:
"A fix for em28xx unregister code affecting devices with FM radio
support"
* tag 'media/v7.3-3' of git://git.kernel.org/pub/scm/linux/kernel/git/mchehab/linux-media:
media: em28xx: use video_unregister_device for radio_dev
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound
Pull sound fixes from Takashi Iwai:
"A dozen of small fixes. All device-specific quirks or fixes, and
nothing exciting is expected.
- USB-audio and HD-audio quirks
- HD-audio TAS2781 codec fix
- ctxfi driver memory leak fix
- ASoC AMD quirks
- ASoC cs35l56 and adau1372 codec fixes"
* tag 'sound-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound:
ASoC: amd: acp: Add more ACP7.0 match entries for Cirrus Logic parts
ASoC: amd: acp: Add DMI override for ASUS EXPERTBOOK AM7406CKA
ASoC: sdw_utils: Switch CS35L56/57/62/63 to use OT25 DAI
ASoC: cs35l56: Add DAI for SDCA OT25 stream
ASoC: samsung: i2s: enable op_clk in probe to balance runtime PM
ASoC: adau1372: fix invalid default state of output ASRC and DAC muxes
ASoC: adau1372: allow sleepy powerdown GPIOs
ASoC: amd: yc: Add DMI quirk for Acer Aspire Go 15 (AG15-21P)
ALSA: hda/realtek: Add mute LED quirk for HP Pavilion 15-eg3xxx (103c:8bfd)
ALSA: hda/tas2781: Accept V1 calibration data without CRC
ALSA: ctxfi: Fix memory leakage when clearing DAO input mappers
ALSA: hda/realtek: Add quirk for ASRock NUC BOX-358H
ALSA: usb-audio: Add native DSD support for Cambridge Audio streamers
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm
Pull pmdomain provider fixes from Ulf Hansson:
- imx: Serialize power on/off across sibling domains for imx8m-blk-ctrl
- rockchip: Fix a couple of errors during probe
* tag 'pmdomain-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm:
pmdomain: rockchip: don't ignore clock lookup errors on attach
pmdomain: rockchip: fix clock leak on domain probe failure
pmdomain: rockchip: propagate subdomain add errors
pmdomain: imx8m-blk-ctrl: Serialize power on/off across sibling domains
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux
Pull dma-mapping fixes from Marek Szyprowski:
"Two more fixes for the corner cases in the DMA-mapping SWIOTLB code
(Peng Fan and Marek Szyprowski)"
* tag 'dma-mapping-7.3-2026-10-09' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux:
swiotlb: fix default_swiotlb_limit() for non-growable default pool
iommu/dma: skip swiotlb bounce for DMA_ATTR_MMIO in iommu_dma_map_phys
|
|
Pull fsverity fix from Eric Biggers:
"Fix a regression from commit f77f281b6118 ("fsverity: use a hashtable
to find the fsverity_info")"
* tag 'fsverity-for-linus' of git://git.kernel.org/pub/scm/fs/fsverity/linux:
fsverity: RCU-delay the freeing of struct fsverity_info
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
"This contains fixes for the current development cycle.
All of them came out of a review of the mount code that started with a
bug report. The review modeled the corner cases of mount propagation,
unmounting and mount reference counting and turned up a lot of bugs.
Most of them years old. Most fixes come with a selftest.
- Rework connected mounts.
A mount that is unmounted together with its parent can stay
attached to the parent to keep its mountpoint covered. That happens
when the mountpoint is removed with rmdir(), unlink() or rename(),
when a detached tree is dissolved, and for locked mounts in any
umount that isn't synchronous, including the teardown of their
mount namespace. The parent then owns the child and drops it on its
own final mntput(). So any reference from the child's superblock
back to one of its ancestors becomes a cycle that is never freed.
A loop device backed by an image on a tmpfs and mounted on that
same tmpfs is enough. Remove the directory the tmpfs is mounted on
from the host, let the container's mount namespace exit, and the
loop device, the tmpfs and the filesystem on the loop device are
leaked for good. The same works with autofs, zram, ecryptfs,
binfmt_misc, fuse passthrough, zloop, a mass storage gadget and md,
and the selftests have reproducers for them. This has been possible
since v4.1. It is also why "put_mnt_ns(): leave mounts connected"
was reverted in -rc5. Keeping every mount of a dying mount
namespace connected made these cycles trivial to create.
Every unmounted mount is now detached from its parent. Where the
mountpoint has to stay covered the mount leaves a cover on the
parent instead, allocated together with the mount. A lookup on the
unmounted parent that hits a cover finds an empty immutable
directory or file on the private nullfs instance. Nothing leads
from a cover to another mount, so no unmounted mount owns another
one and no cycle can form.
This is visible to userspace. A formerly connected mount can no
longer be reached through its unmounted parent and ".." inside it
leads nowhere, as for every other lazily unmounted mount.
With that the private nullfs instance becomes reachable from
userspace, so it now refuses mounts on top, is mounted read-only,
and refuses fsnotify marks and file locks. Its inodes are shared by
every holder and a watch or a lock would otherwise reach across
users. may_decode_fh() now also decides its subtree check under a
single mount_lock hold, as a racing umount could otherwise let it
decode into what a locked child covered.
- umount:
* Don't silently unmount busy mounts.
Since v4.13 propagate_umount() takes down propagated copies of
the victim with children as long as each child is an overmount or
another copy of the victim, but propagate_mount_busy() only ever
checked copies without children or with just an overmount. A
container that moved a tree beneath its copy of a host mount lost
that tree from under its open file descriptor to a plain umount()
on the host. propagate_mount_busy() now applies the same rules,
walking each chain of copies once.
* Don't let a migrating task hide its reference from umount().
mnt_get_count() sums the per-cpu counters under mount_lock but
the mntget() and mntput() fast paths don't take it. A task that
takes a reference on a cpu the sum has already passed and drops
it after migrating to one the sum hasn't reached yet hides the
reference it held to begin with, and umount() succeeds with the
file still open. Gets and puts now live in separate per-cpu
counters and all puts are summed before all gets with a full
barrier in between, the way srcu_readers_active_idx_check() does
it. mntget() is unchanged and mntput() gains an smp_wmb().
* Check each submount for references right before unmounting it.
shrink_submounts() and mark_mounts_for_expiry() checked all their
victims up front. Unmounting the first could move a busy
overmount to where the next victim's propagated copy is looked up
and it was then unmounted without a check.
* Never expire a locked mount.
A shrinkable mount moved beneath a locked mount with
MOVE_MOUNT_BENEATH takes over the lock, and umount() of an
unlocked ancestor expired it and revealed what it covered. That
umount() now fails with EBUSY as it does for any other locked
child. A lazy umount still takes the whole tree.
- Overmounts and locked mounts:
* Unhash a dentry before detaching the mounts on it.
unlink(), rmdir() and rename() detach the mounts on the victim
but only d_delete() it once its inode is unlocked, a window that
includes an expedited RCU grace period. In between, a lookup from
a mount namespace in which the dentry is a mountpoint found it
hashed, positive and uncovered. Drop the dentry first, as
d_invalidate() already does.
* Don't reveal overmounted entries in refwalk.
A refwalk that had grabbed the dentry before the unlink never
rechecked it the way rcuwalk does with d_seq and mount_lock.
Without any artificial widening three walkers read the covered
file 27 times in a minute. step_into() now fails an unhashed
dentry marked DCACHE_CANT_MOUNT with -ESTALE and the walk is
retried.
* Keep covered mounts covered in OPEN_TREE_NAMESPACE.
Creating such a mount namespace only takes a user namespace and
the copy followed bind mount rules: no children without
AT_RECURSIVE and no unbindable mounts with it. An unprivileged
user could see what mounts covered in the source, such as the
parts of /proc and /sys that container runtimes mask. If the
caller doesn't own the source mount namespace a non-recursive
copy of a mount with something mounted below the requested
directory is now refused and a recursive copy includes unbindable
mounts, the way unshare() copies.
* Keep the lock on a mount that a propagated copy is moved beneath.
MNT_LOCKED moved to any mount that ended up beneath a locked
mount, propagated copies included. A host mount and umount on a
directory covered by a locked mount in a less privileged mount
namespace left that cover unlocked for the namespace's owner to
remove. Only mounts the caller places beneath take over the lock
now.
* Handle mount locking for automounts correctly. Which copies to
lock was decided by the mount namespace of the task that
triggered the automount. A task in a user namespace that
triggered one on a host mount through a file descriptor got the
host's own automount locked while its own copy stayed unlocked
and could have nosuid, nodev and noexec cleared. Use the owner of
the mount namespace the mount lands in.
- Use-after-free and crashes:
* Refuse an automount below a mount that is in no namespace.
The private clones overlayfs uses for its layers have the
MNT_NS_INTERNAL error pointer as their namespace, which
finish_automount() let through and count_mounts() dereferenced. A
fanotify filesystem mark on an overlayfs lower layer hands out
file descriptors on such a clone. With debugfs as the lower layer
opening "tracing" oopses with namespace_sem held for writing and
every mount operation on the system blocks from then on.
* Reset the old parent's ->overmount in mnt_change_mountpoint().
When propagate_umount() moved an overmount off a mount that a
file descriptor kept alive, MOVE_MOUNT_BENEATH through that
descriptor later followed the stale pointer into the freed
overmount.
* statmount() with STATMOUNT_BY_FD and pivot_root() read the parent
of a mount that may be unmounted and only held by a file
descriptor, while the parent's final mntput() can free it.
statmount() now reads it under mount_lock and pivot_root() first
checks that both mounts are in the caller's mount namespace.
* Queue a mount only once for mount notifications. A mount
reparented by one umount_tree() and taken down by the next under
the same namespace_sem hold, as in shrink_submounts(), was queued
twice. That cut the mounts queued in between out of notify_list
while it still pointed at them, and once they were freed every
later mount operation walked freed memory.
* Don't let a pseudo dentry become the root of a mount. A bind
mount of a bpf token file did that with a DCACHE_NORCU dentry,
which is freed without an RCU grace period while lockless path
walks may still look at it. Refuse to clone such a mount.
* Don't inherit MNT_UMOUNT in clone_mnt(). A bind mount of a lazily
unmounted nsfs or pidfs mount through its file descriptor started
out flagged as unmounted. Among other things __detach_mounts()
then dropped the namespace's reference on it, the mount outlived
its namespace and mount_setattr() through the descriptor read the
freed namespace. A recursive bind mount of such a mount also
copied the unmounted stack still attached to it. That now fails
with EINVAL, copying the mount itself still works.
* Remove the fsnotify marks of a mount namespace in free_mnt_ns()
instead of the RCU callback that frees the namespace, where
taking the group mutexes meant sleeping in softirq context.
- Propagation and copies:
* Keep a copied mount unbindable. Since v6.17 clone_mnt() didn't
copy the unbindable flag, so every mount namespace created with
CLONE_NEWNS had bindable copies of all unbindable mounts. This
had been fixed once before.
* Refuse MOVE_MOUNT_SET_GROUP on an unbindable mount. It made the
mount an unbindable slave, a state nothing else can produce, or
silently dropped the unbindable flag. CRIU applies MS_UNBINDABLE
after restoring sharing and isn't affected.
* Check a recursive bind mount for mount namespace loops.
Recursively bind mounting a tree from another mount namespace
could put a mount of a namespace's file inside that same
namespace, which then pins itself and all its mounts. Repeating
it leaks without limit, the reproducer took Shmem from 380 kB to
65916 kB. The copy is now checked with check_for_nsfs_mounts()
before it is grafted, as move_mount() does.
* Look at the topmost mount for a mount namespace file.
attach_recursive_mnt() never looked at the topmost mount of the
source's chain of overmounts. If that was the chain's only mount
namespace file an existing mount at a propagated destination got
buried below the root of the nsfs file where no path walk reaches
it.
* Don't put a mountpoint on a dentry that's being removed.
attach_recursive_mnt() makes a mountpoint of the source's root
without its inode lock, so a racing rmdir() of that directory
could leave a mount on it that nothing ever detaches.
d_set_mounted() now checks cant_mount() as well.
- nullfs:
* Take no inode lock for readdir of an immutable directory.
The root of every empty mount namespace is the same nullfs
directory and iterate_dir() held its i_rwsem across
->iterate_shared(). A reader whose buffer faults on a FUSE mount
of its own holds it for as long as its server wants, and with an
exclusive locker queued behind it every lookup that misses the
dcache, every create and every mount in that directory waits. One
user of an empty mount namespace stalls all others. Directories
with the new FOP_IMMUTABLE flag skip the lock.
* Refuse to reconfigure internal superblocks through fspick(),
MS_REMOUNT or the read-only remount that a synchronous umount()
of the root does. For nullfs only root in the initial user
namespace could do it, but the superblock is shared by every
mount namespace and the flags showed up in statfs() for all of
them.
* Don't update the access time on nullfs and refuse F_SET_RW_HINT
on an immutable inode.
- unshare: Free an nsproxy that was never installed with
nsproxy_free() when set_cred_ucounts() fails. put_nsproxy() dropped
active references that were never taken, which triggered a warning
and hid the caller's own namespaces from listns().
- Smaller changes: mount_setattr() checks the target before it walks
the tree to allocate peer group ids, unshare() puts the old
fs_struct before the old namespaces, dissolve_on_fput() drops the
file's reference to the tree itself, disconnect_mount() is
simplified and the documentation of the propagated unmount rule is
brought up to date"
* tag 'vfs-7.3-rc7.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (58 commits)
namespace: simplify disconnect_mount()
selftests/filesystems: test covered mounts
namespace: rework connected mounts
nullfs: add an empty immutable regular file
selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody
selftests/filesystems: add a helper that holds a readdir in a page fault
readdir: take no inode lock on an immutable directory
nullfs: refuse file locks
fsnotify: let a filesystem refuse marks on its objects
namespace: nothing is mounted on or written through knullfs
namespace: keep the private nullfs instance in knullfs
fhandle: decide the subtree check under mount_lock
selftests/filesystems: check that an automount below an overlay layer is refused
selftests/filesystems: check the atime of the empty mount namespace root
selftests/filesystems: check that a lock lands on the right mount and stays
namespace: keep the lock on a mount that a propagated copy is moved beneath
namespace: never expire a locked mount
nullfs: don't update the access time
namespace: handle mount locking for automounts correctly
namespace: refuse an automount below a mount that is in no namespace
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux
Pull arm64 fixes from Will Deacon:
"It finally seems to have calmed down on the arm64 fixes front, so
please pull these two straightforward fixes for -rc7. One fixes the
EL2 trap configuration for implementation-defined CPU PMU hardware
during boot and the other fixes a kcov selftest failure by excluding
our softirq early entry code:
- Fix PMU EL2 trap configuration for CPUs with an IMPDEF PMU
- Fix kcov boot selftest failure by excluding our early IRQ entry
code"
* tag 'arm64-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux:
arm64: irq: exclude the softirq stack switch from KCOV
arm64/boot: Don't set PMUv3p9 FGT2 bits without PMUv3
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from wireless, wireguard, CAN and Bluetooth.
We have one known regression to wrap up in VLAN handling.
Current release - regressions:
- Bluetooth: RFCOMM: fix deadlock on rfcomm_mutex
Previous releases - regressions:
- can: fix regression in handling RPS after migrating metadata to skb_ext
- eth:
- iavf: fix regressions in reconfig impacting bonding
- mana: fix packet forwarding performance regression
- stmmac: remove buggy VLAN acceleration support
Previous releases - always broken:
- a few high prio fixes for tun, and af_packet
- amt: fix a UaF on tunnel teardown
- eth:
- bnxt: fix PCIe AER recovery and FLR handling issues
- macb: don't modify Tx skbs before taking ownership
- axienet: don't leak Tx skbs on interface stop
- wifi:
- nxpwifi: number of LLM-ish fixes
- assorted mt76 fixes"
* tag 'net-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (128 commits)
net: macb: copy shared skbs before appending the FCS
net: macb: check TX ring before modifying skb
vsock: Fix memory leak in vmci_transport_recv_dgram_cb()
wireguard: noise: reject response consumption after intermediate initiation
wireguard: queueing: preserve tstamp_type when encapsulating packet
net: openvswitch: validate transport header presence in set_ipv6_addr
net/smc: protect clcsock lifetime in smc_getname
ipv6: do not warn on route notification size race
ipv4: do not warn on route notification size race
ipv4: validate checksum_start before completing checksum
ptp: ocp: fix PCIe delay estimation calculation
xen/netfront: don't leak the skb when xennet_fill_frags() fails
net/packet: call packet_parse_headers after virtio_net_hdr_to_skb
xen/netfront: drop RX packets with a short Ethernet header
net: skbuff: don't leave stale bytes in skb_copy_and_csum_bits()
net: sparx5: free the matchall entry on destroy
selftests: mlxsw: Test port range occupancy on template create
mlxsw: spectrum_flower: Fix port range register leak in tmplt_create()
net: dsa: microchip: fix KSZ8765 fiber detection
net/mlx5e: Order ICOSQ cc update after CQ doorbell
...
|
|
Alex is seeing a probe failure of the amdgpu driver after the Root Port
above an AMD Navi10 GPU has been reset. The reset was performed to recover
from a Firmware First reported Fatal Error.
However all status registers in the Root Port's AER Extended Capability are
blank, so apparently the platform firmware raised a false alarm.
The issue is only occurring since commit eddba19b8b5f ("PCI/AER: Support
Advisory Non-Fatal Errors"). It looks like enabling Advisory Non-Fatal
Errors causes code paths to be exercised in platform firmware which were
never validated before.
Skip error recovery on false alarms, i.e. if no unmasked errors were
actually signaled.
Note that this will also skip recovery if both the Status and Mask
registers are "all ones", as would be the case for inaccessible devices.
However that seems justified because it would imply either a hot-unplug
event or a Surprise Down Error further up in the hierarchy. Interfering
with recovery from that seems uncalled for.
Fixes: eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors")
Reported-by: Alex Deucher <alexander.deucher@amd.com>
Closes: https://bugzilla.kernel.org/show_bug.cgi?id=222095
Signed-off-by: Lukas Wunner <lukas@wunner.de>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Tested-by: Alex Deucher <alexander.deucher@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/0552ed277e40a288e0157af799257ee6ec722534.1791460615.git.lukas@wunner.de
|
|
Nicolai Buchwitz says:
====================
net: macb: fix software FCS handling of shared and requeued skbs
While testing the genet MTU series I used a Raspberry Pi CM5 (RP1 GEM)
as pktgen source for the CM4. With clone_skb the CM5 rebooted after a
few seconds. Further investigation showed that macb_pad_and_fcs()
appends the FCS in place, so the shared skb grows with every transmit
until BQL completes more than was queued and dql_completed() hits its
BUG_ON.
The same code also modifies the skb before the TX ring check, so a
NETDEV_TX_BUSY retry gets an skb that was already replaced or grown.
Patch 1 checks the ring first, patch 2 copies shared skbs.
Tested on Raspberry CM5 with pktgen at 60/20000 bytes, clone_skb 0/
1000, burst 1/32.
====================
Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-0-a80641479041@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
macb_pad_and_fcs() appends the FCS in place when the skb has tailroom.
A shared skb, as pktgen sends in clone_skb mode, grows by one FCS per
transmit. BQL then completes more bytes than were queued and
dql_completed() hits its BUG_ON.
On a Raspberry Pi CM5 (RP1 GEM) pktgen with clone_skb 1000 burst 32 at
60 bytes kills the box within seconds.
Copy shared skbs before appending the FCS. Clearing IFF_TX_SKB_SHARING
would also fix it but makes pktgen refuse clone_skb on macb.
Fixes: 653e92a9175e ("net: macb: add support for padding and fcs computation")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-2-a80641479041@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
macb_pad_and_fcs() replaces or extends the skb before the ring space
check. On NETDEV_TX_BUSY the stack requeues an skb that is already freed
or grown.
Check the ring first, using the padded length for the descriptor count.
Nonlinear skbs always take the copy path so the count can assume a
linear skb.
Fixes: 653e92a9175e ("net: macb: add support for padding and fcs computation")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-1-a80641479041@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
During the closure of a datagram socket, the vmci_transport_recv_dgram_cb()
function may be called, which will add the packets
to the socket's backlog; then, after the receive queue is cleared,
the __release_sock() function will move the packet back from the socket's
backlog to the receive queue, which will lead to a memory leak.
sock_close
__sock_release
__vsock_release
// take ownership by user-space
lock_sock_nested
sock_set_flag(sk, SOCK_DEAD)
vmci_transport_release
vmci_dispatch_dgs
vmci_datagram_invoke_guest_handler
vmci_transport_recv_dgram_cb
sk_receive_skb
if (!sock_owned_by_user(sk))
[...]
(!) else if sk_add_backlog()
[...]
(!) skb_queue_purge(&sk->sk_receive_queue)
release_sock
__release_sock
// move all packets from the backlog
// to the receive queue
(!) sk_backlog_rcv
vsock_queue_rcv_skb
[...]
__sock_queue_rcv_skb
if (!sock_flag(sk, SOCK_DEAD))
// Since the socket is "dead", the function will not
// be called
sk->sk_data_ready(sk)
sock_release_ownership(sk);
sock_put
Add a receive queue cleanup in the socket destructor to fix this.
syzkaller report:
unreferenced object 0xffff8880117bfb80 (size 240):
comm "irq/56-vmw_vmci", pid 205, jiffies 4295064448 (age 73.992s)
hex dump (first 32 bytes):
78 dc 26 1e 80 88 ff ff 78 dc 26 1e 80 88 ff ff x.&.....x.&.....
00 00 00 00 00 00 00 00 c0 da 26 1e 80 88 ff ff ..........&.....
backtrace:
[<ffffffff82be8d97>] __alloc_skb+0x287/0x330 net/core/skbuff.c:505
[<ffffffffa114bcbd>] vmci_transport_recv_dgram_cb+0xbd/0x210 [vmw_vsock_vmci_transport]
[<ffffffffa05072c6>] vmci_datagram_invoke_guest_handler+0x366/0x450 [vmw_vmci]
[<ffffffffa050ad75>] vmci_dispatch_dgs+0x235/0x490 [vmw_vmci]
[<ffffffffa050b025>] vmci_interrupt+0x55/0x290 [vmw_vmci]
[<ffffffff8137ad2b>] irq_thread_fn+0x8b/0x1a0 kernel/irq/manage.c:1205
[<ffffffff8137c7ae>] irq_thread+0x28e/0x530 kernel/irq/manage.c:1314
[<ffffffff81253dd6>] kthread+0x2e6/0x3a0 kernel/kthread.c:376
[<ffffffff81004392>] ret_from_fork+0x22/0x30 arch/x86/entry/entry_64.S:295
Found by InfoTeCS on behalf of Linux Verification Center
(linuxtesting.org) with Syzkaller.
Fixes: d021c344051a ("VSOCK: Introduce VM Sockets")
Reviewed-by: Stefano Garzarella <sgarzare@redhat.com>
Signed-off-by: Ilia Gavrilov <Ilia.Gavrilov@infotecs.ru>
Link: https://patch.msgid.link/20261008144046.90853-1-Ilia.Gavrilov@infotecs.ru
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Jason A. Donenfeld says:
====================
WireGuard fixes for 7.3-rc7
This series contains two important WireGuard fixes
1) Stop zeroing out skb->tstamp_type when encapsulating packets, so that
fq behaves correctly, from Ramses de Norre.
2) Make sure handshake state isn't swapped out while locks are
released, reported by Jérémy Jean.
====================
Link: https://patch.msgid.link/20261008130124.724119-1-Jason@zx2c4.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Two threads begin processing the identical response message, received
twice. The first thread, A, runs. While it's running, the second one,
B, gets partway through, and during that slow calculation, or even while
blocking on down_write(), A completes and then also a handshake
initiation that's already been queued up runs in thread C, which itself
takes that same down_write(). The handshake initiation creation
succeeds, and sets the state back to waiting-for-response, and calls
up_write(), at which point thread B resumes, because either its finished
its calculations or was finally allowed to acquire down_write(). Thread
B then copies the state back to the peer, and begins a new session,
using that state, which is the same session as the one made in thread A.
Thread A Thread B Thread C
down_read()
sA = handshake->state
memcpy(cA, handshake->crypto)
up_read()
if (sA != 1)
goto fail
slow_crypto(cA)
down_read()
sB = handshake->state
memcpy(cB, handshake->crypto)
up_read()
if (sB != 1)
goto fail
slow_crypto(cB)
down_write()
if (sA != handshake->state)
goto fail
memcpy(handshake->crypto, cA)
handshake->state = 2
up_write()
down_write()
if (handshake->state != 2)
goto fail
derive_session(handshake->crypto)
up_write()
down_write()
slow_crypto(handshake->crypto)
handshake->state = 1
up_write()
down_write()
if (sB != handshake->state)
goto fail
memcpy(handshake->crypto, cB)
handshake->state = 2
up_write()
down_write()
if (handshake->state != 2)
goto fail
derive_session(handshake->crypto)
up_write()
This seems basically impossible to hit in a meaningful way in practice,
but ensure that it absolutely cannot happen by comparing the ephemeral
private key that's on the stack with the latest one that the peer's
handshake state has.
Cc: stable@vger.kernel.org
Fixes: e7096c131e51 ("net: WireGuard secure network tunnel")
Reported-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
Link: https://patch.msgid.link/20261008130124.724119-4-Jason@zx2c4.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Sending traffic through a wireguard tunnel on a host using the fq
qdisc fills the log with:
fq: likely mono tstamp with tstamp_type 0
An skb carries a timestamp in skb->tstamp and, separately, a
skb->tstamp_type field recording which clock that timestamp came from.
The two have to agree.
When wireguard encapsulates a packet it calls wg_reset_packet(), which
clears the fields that must not leak from the inner packet into the
tunnel packet. It does so in two steps:
skb_scrub_packet(skb, true);
memset(&skb->headers, 0, sizeof(skb->headers));
skb_scrub_packet() deliberately keeps skb->tstamp when it holds a
monotonic timestamp: that value is the time the packet is scheduled to
be sent, and the qdisc still needs it. The memset then zeroes
skb->tstamp_type, because that field sits inside the headers group
while skb->tstamp does not. The packet therefore leaves wireguard
carrying a monotonic timestamp labelled as a realtime one.
Nothing noticed until commit c4f796c4f16b ("net_sched: sch_fq: convert
skb->tstamp if not monotonic"): fq used to assume every timestamp was
monotonic. It now consults tstamp_type, spots the mismatch, warns, and
falls back to treating the value as monotonic. Pacing still ends up
correct, so the log spam is the actual problem.
Save tstamp_type before the memset and restore it when encapsulating,
next to the hash fields that are already carried over this way. When
decapsulating it stays zeroed, which is right: an incoming packet's
timestamp is a realtime receive timestamp.
Fixes: d98d58a00261 ("net: Set skb->mono_delivery_time and clear it after sch_handle_ingress()")
Signed-off-by: Ramses de Norre <ramses@well-founded.dev>
Reviewed-by: Toke Høiland-Jørgensen <toke@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
Link: https://patch.msgid.link/20261008130124.724119-3-Jason@zx2c4.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When executing IPv6 address rewrite actions on IPv6 fragments,
set_ipv6_addr() calls update_ipv6_checksum(). If parse_ipv6hdr()
processes a non-first IPv6 fragment, it sets key->ip.proto to
NEXTHDR_FRAGMENT and returns early without calling
skb_set_transport_header().
update_ipv6_checksum() unconditionally evaluates skb_transport_offset()
on entry before checking l4_proto. Because skb->transport_header is
uninitialized, this triggers a warning under CONFIG_DEBUG_NET=y although
it is completely harmless.
Fix this by returning early in update_ipv6_checksum() if l4_proto is
NEXTHDR_FRAGMENT. This avoids reading the uninitialized transport offset
for fragments while preserving the debug warning for any other protocol
where the transport header is unexpectedly missing.
See the syzbot trace:
!skb_transport_header_was_set(skb)
WARNING: ./include/linux/skbuff.h:3075 at skb_transport_header include/linux/skbuff.h:3075 [inline], CPU#1: syz-executor463/5635
WARNING: ./include/linux/skbuff.h:3075 at skb_transport_offset include/linux/skbuff.h:3250 [inline], CPU#1: syz-executor463/5635
WARNING: ./include/linux/skbuff.h:3075 at update_ipv6_checksum net/openvswitch/actions.c:361 [inline], CPU#1: syz-executor463/5635
WARNING: ./include/linux/skbuff.h:3075 at set_ipv6_addr+0x462/0x660 net/openvswitch/actions.c:399, CPU#1: syz-executor463/5635
[...]
RIP: 0010:skb_transport_header include/linux/skbuff.h:3075 [inline]
RIP: 0010:skb_transport_offset include/linux/skbuff.h:3250 [inline]
RIP: 0010:update_ipv6_checksum net/openvswitch/actions.c:361 [inline]
RIP: 0010:set_ipv6_addr+0x462/0x660 net/openvswitch/actions.c:399
[...]
Call Trace:
<TASK>
set_ipv6 net/openvswitch/actions.c:531 [inline]
do_execute_actions+0x557e/0x8600 net/openvswitch/actions.c:1366
ovs_execute_actions+0xde/0x520 net/openvswitch/actions.c:1592
ovs_packet_cmd_execute+0xb4f/0xf10 net/openvswitch/datapath.c:705
genl_family_rcv_msg_doit+0x233/0x340 net/netlink/genetlink.c:1114
genl_rcv_msg+0x614/0x7a0 net/netlink/genetlink.c:1209
netlink_rcv_skb+0x226/0x4a0 net/netlink/af_netlink.c:2572
genl_rcv+0x28/0x40 net/netlink/genetlink.c:1218
netlink_unicast+0x7bd/0x940 net/netlink/af_netlink.c:1361
netlink_sendmsg+0x813/0xb40 net/netlink/af_netlink.c:1916
sock_sendmsg_nosec+0x14e/0x190 net/socket.c:800
Reported-by: syzbot+4cc63fcfb3845e149969@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=4cc63fcfb3845e149969
Fixes: 3fdbd1ce11e5 ("openvswitch: add ipv6 'set' action")
Signed-off-by: Fernando Fernandez Mancera <fmancera@suse.de>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Link: https://patch.msgid.link/20261008104135.4722-1-fmancera@suse.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
smc_getname() dereferences smc->clcsock without holding
clcsock_release_lock. Link-group termination can release the CLC socket
through smc_close_active_abort() while the SMC socket is still open,
for example after shutdown(SHUT_WR).
A getsockname() caller can load smc->clcsock, then the termination worker
can clear the pointer and call sock_release() before the caller accesses
clcsock->ops or invokes getname(). This causes a use-after-free; if the
worker clears the pointer before the load, it causes a NULL dereference.
The syscall's file reference keeps the SMC socket alive but does not
prevent asynchronous release of its CLC socket.
KASAN reported the following with test-only timing instrumentation:
BUG: KASAN: slab-use-after-free in smc_getname+0x19e/0x1b0
Read of size 8 at addr ffff888109abb4e0 by task poc/103
Call Trace:
smc_getname+0x19e/0x1b0
do_getsockname+0xe5/0x170
__sys_getsockname+0x8c/0x100
Allocated by task 95:
sock_alloc_inode+0x1e/0x280
sock_alloc+0x3d/0x240
__sock_create+0x7e/0x430
smc_create+0x121/0x240
Freed by task 0:
kmem_cache_free+0xcc/0x340
rcu_core+0x50a/0x1850
Last potentially related work creation:
evict+0x446/0x6c0
smc_clcsock_release+0xa8/0xd0
smc_close_active_abort+0x26a/0x3a0
__smc_lgr_terminate.part.0+0x137/0x2e0
Hold clcsock_release_lock across the pointer check and the getname()
callback to serialize with smc_clcsock_release(). Return -EBADF if the
CLC socket has already been released, preserving the existing peer
state check and the callback's return value otherwise.
Fixes: b03faa1fafc8 ("net/smc: postpone release of clcsock")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Reviewed-by: Mahanta Jambigi <mjambigi@linux.ibm.com>
Reviewed-by: Dust Li <dust.li@linux.alibaba.com>
Link: https://patch.msgid.link/20261008084444.787449-1-nicoyip.dev@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Daehyeon Ko says:
====================
ipv4/ipv6: do not warn on route notification size races
A nexthop group can grow between route notification sizing and filling.
Both IPv4 and IPv6 can then legitimately return -EMSGSIZE, so remove the
stale warnings while preserving their existing error paths.
Patch 1 is unchanged from v1. Patch 2 adds the IPv6 counterpart requested
by Ido. No new build or runtime test was run; both changes only delete the
stale comment and WARN_ON().
====================
Link: https://patch.msgid.link/cover.1791423190.git.4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
fib6_info_hw_flags_set() sizes the skb with rt6_nlmsg_size() and later
fills it with rt6_fill_node(). With nexthop compatibility mode enabled,
a concurrent replacement can grow the group between these independent
snapshots. rt6_fill_node() can then legitimately return -EMSGSIZE, so
the warning does not prove a sizing bug.
Remove the warning. The existing error path still frees the skb and
reports the error to listeners.
Fixes: 907eea486888 ("net: ipv6: Emit notification when fib hardware flags are changed")
Reported-by: Ido Schimmel <idosch@nvidia.com>
Link: https://lore.kernel.org/netdev/20261007114003.GA1011260@shredder/
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Link: https://patch.msgid.link/4012604727e5ba80b69ccd4d08796511397ddaef.1791423190.git.4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 680aea08e78c ("net: ipv4: Emit notification when fib hardware
flags are changed") added asynchronous route notifications for hardware
flag changes.
fib_alias_hw_flags_set() sizes the skb with fib_nlmsg_size() and later
fills it with fib_dump_info() while holding only RCU. With nexthop
compatibility mode enabled, a concurrent replacement can grow the group
between these independent snapshots. fib_dump_info() can then
legitimately return -EMSGSIZE, so the warning does not prove a sizing bug.
Remove the warning. The existing error path still frees the skb and
reports the error to listeners.
Fixes: 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed")
Reported-by: Ido Schimmel <idosch@nvidia.com>
Link: https://lore.kernel.org/netdev/20261004082043.GB92032@shredder/
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Link: https://patch.msgid.link/b7e203fc5867651cb67ef84aa03e082f0f76008e.1791423190.git.4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
If a packet with bad checksum metadata gets into the ipv4 stack,
skb_checksum_help can corrupt the network header and cause a bunch of
mischief.
This was discovered and reported by Paulos, and has been reporoduced
by others independently since.
We really shouldn't allow such packets in, but as a defence
in depth measure, let's also check before we complete the checksum.
A more complete validation at input is forthcoming, but needs more work.
Cc: stable@vger.kernel.org
Fixes: f43798c27684 ("tun: Allow GSO using virtio_net_hdr")
Fixes: bfd5f4a3d605 ("packet: Add GSO/csum offload support.")
Reported-by: Paulos Yibelo <habte.yibelo@gmail.com>
Closes: https://lore.kernel.org/netdev/20260922030310.8684-2-habte.yibelo@gmail.com/
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/0a012b4923e189c4c593ef4f471e5ff0edbe9030.1791412497.git.mst@redhat.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The commit in fixes introduced a high cap for delayas U64_MAX value
while ktime_t is actually s64. This is wrong cap as it becomes negative
value and any comparison to a real delay will fail to update delay
value. Use KTIME_MAX constant as correct max cap for PCIe delay.
The issue was hit in production (a negative value is observed):
# cat /sys/class/timecard/ocp0/ts_window_adjust
-3
Fixes: aa05fe67bcd64 ("ptp: ocp: Improve PCIe delay estimation")
Signed-off-by: Vadim Fedorenko <vadim.fedorenko@linux.dev>
Reviewed-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20261007203359.417270-1-vadim.fedorenko@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When a response chain has more slots than fit in the skb's frags,
xennet_fill_frags() returns an error and xennet_poll() jumps to its
error path. That path moves what's left on tmpq to errq to be freed,
but the skb being filled was already dequeued from tmpq, so it's never
freed. Each chain that overflows leaks the skb and the pages attached
to it as frags, and the backend decides how many slots it sends.
Put the skb back on tmpq before taking the error path, like the
xennet_set_skb_gso() failure just above it does.
Fixes: ad4f15dc2c70 ("xen/netfront: don't bug in case of too many frags")
Cc: stable@vger.kernel.org
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
Reviewed-by: Juergen Gross <jgross@suse.com>
Link: https://patch.msgid.link/20261007-b4-xen-netfront-fill-frags-leak-v1-1-a8a01ff9cd52@toxicpanda.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
SOCK_RAW packet sockets incorrectly drop VLAN-tagged GSO packets
without VIRTIO_NET_HDR_F_NEEDS_CSUM.
packet_parse_headers() sets skb->protocol to VLAN but advances
skb->network_header past the VLAN tags. virtio_net_hdr_to_skb()
then flow dissects these packets, which parses the network header as a
VLAN tag, and fails.
Call packet_parse_headers() after virtio_net_hdr_to_skb() instead.
This matches tun_get_user() and tap_get_user() and thus simplifies
overall complexity.
Commit 01fdecc0480d ("net: packet: fix wrong transport_header when
sending VLAN-tagged frame") fixed the same issue for the transport
header probe in packet_parse_headers(). That probe is now skipped if
virtio_net_hdr_to_skb() already set the transport header, avoiding a
redundant dissection.
Implementation details:
- Do not set skb->protocol to the RX wildcard ETH_P_ALL. SOCK_RAW then
derives it from the link layer header, as before. SOCK_DGRAM now
leaves it 0, or derives it from gso_type for GSO with NEEDS_CSUM.
- Keep virtio_net_hdr_set_proto() last, so that the link layer protocol
takes priority over gso_type.
Reported-by: Junnan Zhang <zhangjn_dev@163.com>
Closes: https://lore.kernel.org/netdev/20260821085722.24036-1-zhangjn_dev@163.com/
Fixes: dfed913e8b55 ("net/af_packet: add VLAN support for AF_PACKET SOCK_RAW GSO")
Signed-off-by: Willem de Bruijn <willemb@google.com>
Acked-by: Michael S. Tsirkin <mst@redhat.com>
Link: https://patch.msgid.link/20261007175816.3138556-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
handle_incoming_queue() pulls pull_to bytes into the head before
calling eth_type_trans(). pull_to is the length of the first RX slot,
capped at RX_COPY_THRESHOLD, and that length comes from the backend.
Nothing checks it against ETH_HLEN.
If the first slot is shorter than ETH_HLEN and more slots follow, the
head ends up shorter than an Ethernet header while skb->len is longer,
and eth_type_trans() BUG()s in __skb_pull(). If the whole packet is
shorter than ETH_HLEN, eth_type_trans() reads the header past the end
of the data instead.
Pull at least ETH_HLEN, and drop the packet if that fails, which also
drops packets too short to hold an Ethernet header. This also checks
the return value of the pull, which was ignored.
Fixes: 0d160211965b ("xen: add virtual network device driver")
Cc: stable@vger.kernel.org
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
Link: https://patch.msgid.link/20261007-b4-xen-netfront-short-head-v1-1-12d7113a7e4e@toxicpanda.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When skb_copy_and_csum_bits() reaches unreadable frags it returns 0
after copying only the linear part, and the rest of the caller's buffer
is left as it was. The callers copy into a buffer that is about to go
out on the wire: an ICMP error quoting the offending packet, or a
driver's TX bounce buffer in skb_copy_and_csum_dev(). Neither buffer
is zeroed beforehand, so whatever was in memory there gets sent.
Zero the part of the buffer we didn't fill. The checksum usually
won't match the data any more, so the receiver will usually drop the
packet, but either way it no longer carries anything it shouldn't.
Only zero for a positive @len, a negative one from a broken caller must
not turn into a huge memset().
Fixes: 65249feb6b3d ("net: add support for skbs with unreadable frags")
Cc: stable@vger.kernel.org
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
Reviewed-by: Mina Almasry <almasrymina@google.com>
Link: https://patch.msgid.link/20261007-b4-skb-copy-csum-stale-bytes-v1-1-adbbde033fb3@toxicpanda.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
sparx5_tc_matchall_replace() allocates a struct sparx5_mall_entry for
every offloaded matchall filter and adds it to sparx5->mall_entries.
sparx5_tc_matchall_destroy() removes the entry from the list, but never
frees it, so the entry of every deleted mirror and goto matchall filter
is leaked.
Free the entry after unlinking it.
The leak was discovered by an AI code review agent, and reproduced with
kmemleak on a lan969x EV board (EV23X71A) by rep |