| Age | Commit message (Collapse) | Author | Files | Lines |
|
git://git.kernel.org/pub/scm/linux/kernel/git/ras/ras
Pull EDAC fix from Borislav Petkov:
- amd64_edac: Shorten the ErrorInformation field read from the MCA_SYND
MSR to only two bits. It is perfectly fine to do so because no system
ever supported more than 2 bits of information (the Chip Selects used
were only 4 maximum) and newer hardware will use only 2 bits anyway
* tag 'edac_urgent_for_v7.3_rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/ras/ras:
EDAC/amd64: Mask UMC chip select to the four implemented selects
|
|
Pull drm fixes from Dave Ailie:
"Live from Dublin Airport, it's Saturday Night drm fixes.
This week has the missing misc fixes from last week which I tracked
down and seemed to be a race/bug in my lei setup somehow, once I asked
lei to ignore it's cache I got the missing email. But there are more
misc fixes this week and amd and intel ones.
The main ones in this are amdgpu and xe, with vc4, vmwgfx, nouveau and
imagination in the middle, with a bunch of small single fixes.
Bit busier than I'd like, but the missing misc might explain it,
anyways time for me to fly home.
fb:
- defer setup when fbdev_probe() fails, not just on -EAGAIN
bridge:
- Fix unlocked list_del in drm_bridge_add()
- Fix unlocked list access in drm_bridge_attach()
- th1520-dw-hdmi:
- Fix error check on dw_hdmi_probe() return value
- Fix remove() callback
aspeed:
- Balance the display clock enable on teardown
i915:
- Fix a GT park race that could trigger spurious RPM wakelock
warnings
- Fix execbuffer relocation cleanup to avoid out-of-bounds free
xe:
- i2c removal
- system Controller Maibox header handling
- two bo pin/unpin accounting bugs
- not emitting a w/a twice
- xe_mmio_wait32() to honor delay/sleep maximums
amdgpu:
- Fix direct scanout alpha on some DRM_FORMATs
- UserQ fixes
- Fix tearing flips with PSR
- GPU reset vblank fix
- Debugfs register interface fix
- Fix display mode patching
- Expand Mac reserved memory workaround
amdkfd:
- SDMA doorbell fix
- Fix dma-buf reference leak in error path
imagination:
- fix reference and vm_bo handling in remap()
- size page table preallocation by device address
gud:
- don't keep a connector without a CRTC as connector_state
panthor:
- validate userspace queue count against initialized firmware
slot count
nouveau:
- fix null ptr regression
- skip turing CE workaround for non-GR
vmwgfx:
- validate pitch value
- add blend mode property
vc4:
- drain hangcheck timer on unbind
- free BO cache on teardown
- use kvmalloc_objs for bo cache size list
- disable V3D interrupt across runtime suspend
- fix binner slot allocation
- fix fw refcount leak"
* tag 'drm-fixes-2026-10-11' of https://gitlab.freedesktop.org/drm/kernel: (39 commits)
drm/nouveau: Skip the Turing CE workaround object for non-GR channels
drm/vc4: Fix binner slot allocation failing on an idle GPU
drm/vc4: Disable the V3D interrupt across runtime suspend
drm/imagination: Size page table preallocation by device address
drm/imagination: Fix reference and vm_bo handling in remap()
drm/amdgpu: make Mac FB workaround generic
drm/amdkfd: fix dma_buf reference leak in get_dmabuf_info
drm/amd/display: Don't replace a sink mode that shares the native totals
drm/amdkfd: align SDMA doorbell base for shader doorbell writes
drm/amdgpu: copy debugfs register data outside the GRBM/SRBM locks
drm/amdgpu: serialize vblank counter reads against GPU reset
drm/amd/display: disable self-refresh on tearing flips
drm/amdgpu/userq: return the memdup_user() error for the user MQD
drm/amdgpu/userq: only accept doorbell BOs as queue doorbell
drm/amd/display: Fix direct scanout alpha on some DRM_FORMATs
drm/xe/i2c: cancel the client work on remove
drm/xe/sysctrl: Fix mailbox header handling
drm/i915/gt: Unmask interrupts only when ACTIVE is true
drm/i915/gem: Prevent overstepping exec array boundary
drm/gud: don't keep a connector without a CRTC as connector_state
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/andi.shyti/linux
Pull i2c fix from Andi Shyti:
"Just one qcom-geni fix for a runtime PM reference leak
during transfer setup"
* tag 'i2c-fixes-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/andi.shyti/linux:
i2c: qcom-geni: release runtime PM reference when set_rate fails
|
|
https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes
A pointer assignment fix for gud, a reference fix and a preallocation
size fix for imagination, a suspend fix and an allocation fix when idle
for vc4, and a hardware variant fix for nouveau.
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Maxime Ripard <self@mripard.dev>
Link: https://patch.msgid.link/asjkbknTHg2t8JGy@houat
|
|
https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes
amd-drm-fixes-7.3-2026-10-08:
amdgpu:
- Fix direct scanout alpha on some DRM_FORMATs
- UserQ fixes
- Fix tearing flips with PSR
- GPU reset vblank fix
- Debugfs register interface fix
- Fix display mode patching
- Expand Mac reserved memory workaround
amdkfd:
- SDMA doorbell fix
- Fix dma-buf reference leak in error path
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20261008174658.2506431-1-alexander.deucher@amd.com
|
|
https://gitlab.freedesktop.org/drm/xe/kernel into drm-fixes
Fixes on:
- i2c removal (Fan)
- system Controller Maibox header handling (Mallesh)
- two bo pin/unpin accounting bugs (Thomas)
- not emitting a w/a twice (Tvrtko)
- xe_mmio_wait32() to honor delay/sleep maximums (Alan)
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Rodrigo Vivi <rodrigo.vivi@intel.com>
Link: https://patch.msgid.link/asfEfgbIjs0GJ-RF@intel.com
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci
Pull PCI fix from Bjorn Helgaas:
"This fixes some GPU initialization regressions caused by eddba19b8b5f
("PCI/AER: Support Advisory Non-Fatal Errors"), which appeared in
v7.3-rc1.
That commit also caused a MacBookPro16,1 spontaneous power-off
regression; I expect a fix for that next week"
* tag 'pci-v7.3-fixes-4' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci:
PCI/AER: Skip error recovery on false alarms
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/mmc
Pull MMC/MEMSTICK fixes from Ulf Hansson:
"MMC host:
- cavium-octeon|thunderx: Destroy slot platform devices on remove
- mtk-sd: Cancel request timeout work on remove
- sdhci-sprd: Disable runtime PM on remove
MEMSTICK:
- Wait for request completion before freeing card
- rtsx_usb_ms: Complete requests after eject instead of dropping
them"
* tag 'mmc-v7.3-rc1-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/mmc:
memstick: rtsx_usb_ms: complete requests after eject instead of dropping them
memstick: core: wait for request completion before freeing card
mmc: cavium-thunderx: destroy slot platform devices on remove
mmc: cavium-octeon: destroy slot platform devices on remove
mmc: sdhci-sprd: disable runtime PM on remove
mmc: mtk-sd: Cancel request timeout work on remove
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mchehab/linux-media
Pull media fix from Mauro Carvalho Chehab:
"A fix for em28xx unregister code affecting devices with FM radio
support"
* tag 'media/v7.3-3' of git://git.kernel.org/pub/scm/linux/kernel/git/mchehab/linux-media:
media: em28xx: use video_unregister_device for radio_dev
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm
Pull pmdomain provider fixes from Ulf Hansson:
- imx: Serialize power on/off across sibling domains for imx8m-blk-ctrl
- rockchip: Fix a couple of errors during probe
* tag 'pmdomain-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm:
pmdomain: rockchip: don't ignore clock lookup errors on attach
pmdomain: rockchip: fix clock leak on domain probe failure
pmdomain: rockchip: propagate subdomain add errors
pmdomain: imx8m-blk-ctrl: Serialize power on/off across sibling domains
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux
Pull dma-mapping fixes from Marek Szyprowski:
"Two more fixes for the corner cases in the DMA-mapping SWIOTLB code
(Peng Fan and Marek Szyprowski)"
* tag 'dma-mapping-7.3-2026-10-09' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux:
swiotlb: fix default_swiotlb_limit() for non-growable default pool
iommu/dma: skip swiotlb bounce for DMA_ATTR_MMIO in iommu_dma_map_phys
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from wireless, wireguard, CAN and Bluetooth.
We have one known regression to wrap up in VLAN handling.
Current release - regressions:
- Bluetooth: RFCOMM: fix deadlock on rfcomm_mutex
Previous releases - regressions:
- can: fix regression in handling RPS after migrating metadata to skb_ext
- eth:
- iavf: fix regressions in reconfig impacting bonding
- mana: fix packet forwarding performance regression
- stmmac: remove buggy VLAN acceleration support
Previous releases - always broken:
- a few high prio fixes for tun, and af_packet
- amt: fix a UaF on tunnel teardown
- eth:
- bnxt: fix PCIe AER recovery and FLR handling issues
- macb: don't modify Tx skbs before taking ownership
- axienet: don't leak Tx skbs on interface stop
- wifi:
- nxpwifi: number of LLM-ish fixes
- assorted mt76 fixes"
* tag 'net-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (128 commits)
net: macb: copy shared skbs before appending the FCS
net: macb: check TX ring before modifying skb
vsock: Fix memory leak in vmci_transport_recv_dgram_cb()
wireguard: noise: reject response consumption after intermediate initiation
wireguard: queueing: preserve tstamp_type when encapsulating packet
net: openvswitch: validate transport header presence in set_ipv6_addr
net/smc: protect clcsock lifetime in smc_getname
ipv6: do not warn on route notification size race
ipv4: do not warn on route notification size race
ipv4: validate checksum_start before completing checksum
ptp: ocp: fix PCIe delay estimation calculation
xen/netfront: don't leak the skb when xennet_fill_frags() fails
net/packet: call packet_parse_headers after virtio_net_hdr_to_skb
xen/netfront: drop RX packets with a short Ethernet header
net: skbuff: don't leave stale bytes in skb_copy_and_csum_bits()
net: sparx5: free the matchall entry on destroy
selftests: mlxsw: Test port range occupancy on template create
mlxsw: spectrum_flower: Fix port range register leak in tmplt_create()
net: dsa: microchip: fix KSZ8765 fiber detection
net/mlx5e: Order ICOSQ cc update after CQ doorbell
...
|
|
This workaround is here for the old Gallium driver and predate async CE
and NVDEC. Instead of only skipping it for NVDEC, let's only apply it
in case of GR channels.
This should make async CE work properly on Turing by stopping the assign
of a GRCE. (here CE0)
Signed-off-by: Mary Guillemard <mary@mary.zone>
Reviewed-by: Lyude Paul <lyude@redhat.com>
Reviewed-by: Daniel Almeida <daniel.almeida@collabora.com>
Signed-off-by: Lyude Paul <lyude@redhat.com>
Link: https://patch.msgid.link/20261007-turing-async-ce-fix-v1-1-12ee6ccd1a3c@mary.zone
|
|
Alex is seeing a probe failure of the amdgpu driver after the Root Port
above an AMD Navi10 GPU has been reset. The reset was performed to recover
from a Firmware First reported Fatal Error.
However all status registers in the Root Port's AER Extended Capability are
blank, so apparently the platform firmware raised a false alarm.
The issue is only occurring since commit eddba19b8b5f ("PCI/AER: Support
Advisory Non-Fatal Errors"). It looks like enabling Advisory Non-Fatal
Errors causes code paths to be exercised in platform firmware which were
never validated before.
Skip error recovery on false alarms, i.e. if no unmasked errors were
actually signaled.
Note that this will also skip recovery if both the Status and Mask
registers are "all ones", as would be the case for inaccessible devices.
However that seems justified because it would imply either a hot-unplug
event or a Surprise Down Error further up in the hierarchy. Interfering
with recovery from that seems uncalled for.
Fixes: eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors")
Reported-by: Alex Deucher <alexander.deucher@amd.com>
Closes: https://bugzilla.kernel.org/show_bug.cgi?id=222095
Signed-off-by: Lukas Wunner <lukas@wunner.de>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Tested-by: Alex Deucher <alexander.deucher@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/0552ed277e40a288e0157af799257ee6ec722534.1791460615.git.lukas@wunner.de
|
|
A job's binner slots are returned to bin_alloc_used by vc4_complete_exec(),
which runs from the job_done workqueue, but the seqno
vc4_v3d_get_bin_slot() waits on is incremented earlier, in the function
vc4_irq_finish_render_job(). Therefore, a waiter can wake, retry, and still
find the pool full because the worker has not run. By then, the job might
have left the render_job_list, so no seqno remains to wait on and the
allocation fails returning -ENOMEM.
This scenario can be reproduced on a Raspberry Pi 3 by using a burst of
small jobs (e.g. Piglit's `quick_gl` suite), with userspace seeing a
rejected submit for a pool about to become free.
Note that the slots are held from validation until the render completes,
so the pool can also be exhausted by jobs still queued on bin_job_list,
leaving render_job_list empty and nothing to wait on.
Release the slots from the FRDONE handler and wait on the pool instead of
on one job's seqno. This bounds the wait to the actual resource. Also, by
claiming the slot inside the wait condition, another submit is unable to
take the slot between the wakeup and the retry. vc4_complete_exec() keeps
clearing the slots for jobs that never completed, and now wakes the
waiters too.
Note that the wait takes a deadline, since the two callers need different
ones. The submit path is a user task and can be interrupted, so it waits
without one. On the other hand, vc4_overflow_mem_work() runs on a worker
and never sees a signal, so it needs a deadline. During a reset,
vc4_irq_disable() masks V3D_DRIVER_IRQS before draining the worker, and
with those masked nothing can release a slot. This deadlocks: the worker
would wait for slots that only the reset can free, while the reset waits
in cancel_work_sync() for the worker.
Fixes: 553c942f8b2c ("drm/vc4: Allow using more than 256MB of CMA memory.")
Reviewed-by: Iago Toral Quiroga <itoral@igalia.com>
Link: https://patch.msgid.link/20260930182204.1877361-1-mcanal@igalia.com
Signed-off-by: Maíra Canal <mcanal@igalia.com>
|
|
vc4_irq_disable() masks the V3D interrupt sources and then calls
synchronize_irq() before the V3D is powered down. However, by itself,
this is not enough to quiesce the interrupt handler.
synchronize_irq() waits for handlers that have already set
IRQD_IRQ_INPROGRESS, and for irqchips reporting IRQCHIP_STATE_ACTIVE.
A GIC interrupt chip is able to mark an interrupt active as soon as a
CPU acknowledges it, so there these checks cover the whole dispatch path.
On RPi 0-3, however, the interrupt controller is ARMCTRL, which has no
active state. A CPU that has read the hwirq out of the pending register
but has not yet reached handle_level_irq() stays invisible to
synchronize_irq().
During a power transition, vc4_irq() may therefore run after the power
domain is off, where every V3D register read will return 0xdeadbeef.
0xdeadbeef has FLDONE, FRDONE and OUTOMEM set. This problem doesn't
trigger NULL pointer dereference issues only because the functions
vc4_irq_finish_bin_job() and vc4_irq_finish_render_job() return early
when there is no job pointer, but OUTOMEM is still able to schedule
vc4_overflow_mem_work() after vc4_irq_disable() has already cancelled it.
Therefore, fix the spurious interrupts by disabling the interrupt line
across the PM transition. To preserve the enable/disable balance, move
vc4_irq_install() and request the IRQ line disabled before the first
runtime resume.
Fixes: 9b6f461582e6 ("drm/vc4: v3d: Stop disabling interrupts")
Reviewed-by: Iago Toral Quiroga <itoral@igalia.com>
Link: https://patch.msgid.link/20260925191318.541938-1-mcanal@igalia.com
Signed-off-by: Maíra Canal <mcanal@igalia.com>
|
|
macb_pad_and_fcs() appends the FCS in place when the skb has tailroom.
A shared skb, as pktgen sends in clone_skb mode, grows by one FCS per
transmit. BQL then completes more bytes than were queued and
dql_completed() hits its BUG_ON.
On a Raspberry Pi CM5 (RP1 GEM) pktgen with clone_skb 1000 burst 32 at
60 bytes kills the box within seconds.
Copy shared skbs before appending the FCS. Clearing IFF_TX_SKB_SHARING
would also fix it but makes pktgen refuse clone_skb on macb.
Fixes: 653e92a9175e ("net: macb: add support for padding and fcs computation")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-2-a80641479041@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
macb_pad_and_fcs() replaces or extends the skb before the ring space
check. On NETDEV_TX_BUSY the stack requeues an skb that is already freed
or grown.
Check the ring first, using the padded length for the descriptor count.
Nonlinear skbs always take the copy path so the count can assume a
linear skb.
Fixes: 653e92a9175e ("net: macb: add support for padding and fcs computation")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-1-a80641479041@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Two threads begin processing the identical response message, received
twice. The first thread, A, runs. While it's running, the second one,
B, gets partway through, and during that slow calculation, or even while
blocking on down_write(), A completes and then also a handshake
initiation that's already been queued up runs in thread C, which itself
takes that same down_write(). The handshake initiation creation
succeeds, and sets the state back to waiting-for-response, and calls
up_write(), at which point thread B resumes, because either its finished
its calculations or was finally allowed to acquire down_write(). Thread
B then copies the state back to the peer, and begins a new session,
using that state, which is the same session as the one made in thread A.
Thread A Thread B Thread C
down_read()
sA = handshake->state
memcpy(cA, handshake->crypto)
up_read()
if (sA != 1)
goto fail
slow_crypto(cA)
down_read()
sB = handshake->state
memcpy(cB, handshake->crypto)
up_read()
if (sB != 1)
goto fail
slow_crypto(cB)
down_write()
if (sA != handshake->state)
goto fail
memcpy(handshake->crypto, cA)
handshake->state = 2
up_write()
down_write()
if (handshake->state != 2)
goto fail
derive_session(handshake->crypto)
up_write()
down_write()
slow_crypto(handshake->crypto)
handshake->state = 1
up_write()
down_write()
if (sB != handshake->state)
goto fail
memcpy(handshake->crypto, cB)
handshake->state = 2
up_write()
down_write()
if (handshake->state != 2)
goto fail
derive_session(handshake->crypto)
up_write()
This seems basically impossible to hit in a meaningful way in practice,
but ensure that it absolutely cannot happen by comparing the ephemeral
private key that's on the stack with the latest one that the peer's
handshake state has.
Cc: stable@vger.kernel.org
Fixes: e7096c131e51 ("net: WireGuard secure network tunnel")
Reported-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
Link: https://patch.msgid.link/20261008130124.724119-4-Jason@zx2c4.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Sending traffic through a wireguard tunnel on a host using the fq
qdisc fills the log with:
fq: likely mono tstamp with tstamp_type 0
An skb carries a timestamp in skb->tstamp and, separately, a
skb->tstamp_type field recording which clock that timestamp came from.
The two have to agree.
When wireguard encapsulates a packet it calls wg_reset_packet(), which
clears the fields that must not leak from the inner packet into the
tunnel packet. It does so in two steps:
skb_scrub_packet(skb, true);
memset(&skb->headers, 0, sizeof(skb->headers));
skb_scrub_packet() deliberately keeps skb->tstamp when it holds a
monotonic timestamp: that value is the time the packet is scheduled to
be sent, and the qdisc still needs it. The memset then zeroes
skb->tstamp_type, because that field sits inside the headers group
while skb->tstamp does not. The packet therefore leaves wireguard
carrying a monotonic timestamp labelled as a realtime one.
Nothing noticed until commit c4f796c4f16b ("net_sched: sch_fq: convert
skb->tstamp if not monotonic"): fq used to assume every timestamp was
monotonic. It now consults tstamp_type, spots the mismatch, warns, and
falls back to treating the value as monotonic. Pacing still ends up
correct, so the log spam is the actual problem.
Save tstamp_type before the memset and restore it when encapsulating,
next to the hash fields that are already carried over this way. When
decapsulating it stays zeroed, which is right: an incoming packet's
timestamp is a realtime receive timestamp.
Fixes: d98d58a00261 ("net: Set skb->mono_delivery_time and clear it after sch_handle_ingress()")
Signed-off-by: Ramses de Norre <ramses@well-founded.dev>
Reviewed-by: Toke Høiland-Jørgensen <toke@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
Link: https://patch.msgid.link/20261008130124.724119-3-Jason@zx2c4.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The commit in fixes introduced a high cap for delayas U64_MAX value
while ktime_t is actually s64. This is wrong cap as it becomes negative
value and any comparison to a real delay will fail to update delay
value. Use KTIME_MAX constant as correct max cap for PCIe delay.
The issue was hit in production (a negative value is observed):
# cat /sys/class/timecard/ocp0/ts_window_adjust
-3
Fixes: aa05fe67bcd64 ("ptp: ocp: Improve PCIe delay estimation")
Signed-off-by: Vadim Fedorenko <vadim.fedorenko@linux.dev>
Reviewed-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20261007203359.417270-1-vadim.fedorenko@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When a response chain has more slots than fit in the skb's frags,
xennet_fill_frags() returns an error and xennet_poll() jumps to its
error path. That path moves what's left on tmpq to errq to be freed,
but the skb being filled was already dequeued from tmpq, so it's never
freed. Each chain that overflows leaks the skb and the pages attached
to it as frags, and the backend decides how many slots it sends.
Put the skb back on tmpq before taking the error path, like the
xennet_set_skb_gso() failure just above it does.
Fixes: ad4f15dc2c70 ("xen/netfront: don't bug in case of too many frags")
Cc: stable@vger.kernel.org
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
Reviewed-by: Juergen Gross <jgross@suse.com>
Link: https://patch.msgid.link/20261007-b4-xen-netfront-fill-frags-leak-v1-1-a8a01ff9cd52@toxicpanda.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
handle_incoming_queue() pulls pull_to bytes into the head before
calling eth_type_trans(). pull_to is the length of the first RX slot,
capped at RX_COPY_THRESHOLD, and that length comes from the backend.
Nothing checks it against ETH_HLEN.
If the first slot is shorter than ETH_HLEN and more slots follow, the
head ends up shorter than an Ethernet header while skb->len is longer,
and eth_type_trans() BUG()s in __skb_pull(). If the whole packet is
shorter than ETH_HLEN, eth_type_trans() reads the header past the end
of the data instead.
Pull at least ETH_HLEN, and drop the packet if that fails, which also
drops packets too short to hold an Ethernet header. This also checks
the return value of the pull, which was ignored.
Fixes: 0d160211965b ("xen: add virtual network device driver")
Cc: stable@vger.kernel.org
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
Link: https://patch.msgid.link/20261007-b4-xen-netfront-short-head-v1-1-12d7113a7e4e@toxicpanda.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
sparx5_tc_matchall_replace() allocates a struct sparx5_mall_entry for
every offloaded matchall filter and adds it to sparx5->mall_entries.
sparx5_tc_matchall_destroy() removes the entry from the list, but never
frees it, so the entry of every deleted mirror and goto matchall filter
is leaked.
Free the entry after unlinking it.
The leak was discovered by an AI code review agent, and reproduced with
kmemleak on a lan969x EV board (EV23X71A) by repeatedly adding and
deleting matchall mirror and goto filters. With the fix, kmemleak no
longer reports the leak.
Cc: stable@vger.kernel.org
Fixes: 1ede4acf045c ("net: sparx5: add bookkeeping code for matchall rules")
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20261007-sparx5-matchall-kfree-net-v1-1-c8918685b337@microchip.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
mlxsw_sp_flower_tmplt_create() parses a flow_cls_offload template
into a stack-local struct mlxsw_sp_acl_rule_info purely to compute
rulei.values.elusage. Parsing can acquire port range registers via
mlxsw_sp_flower_parse_ports_range(), but since this rulei never goes
through mlxsw_sp_acl_rulei_destroy(), those registers were never
released, including through chain template deletion.
Factor out of mlxsw_sp_acl_rulei_destroy() the code to actually
release the necessary resources and call from
mlxsw_sp_flower_tmplt_create() to plug the leak.
The issue was found during a review of Wentao Liang's patch referenced
below.
Fixes: fe22f7410527 ("mlxsw: spectrum_flower: Add ability to match on port ranges")
Reported-by: Wentao Liang <vulab@iscas.ac.cn>
Closes: https://lore.kernel.org/netdev/20260917113236.2149095-1-vulab@iscas.ac.cn/
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Petr Machata <petrm@nvidia.com>
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
Link: https://patch.msgid.link/95339cf970ce78e1a74aa86ffcc104ecf7c7762c.1791294384.git.petrm@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The KSZ8765 is similar to the KSZ8795 but has two fiber ports.
KSZ8_PORT_STATUS_0 is used to detect fiber mode. It is currently
defined as 0x08, which is the Global Control 6 MIB Control register
and has bit 7 defined as Flush Counter.
Set KSZ8_PORT_STATUS_0 to 0x18, which is the Port 1 Status 0 register
where bit 7 is the Fiber Mode bit.
This issue was discovered when upgrading an embedded device from kernel
5.10 to 6.16. The device was using a KSZ8765 device tree configuration
and the corresponding hardware, but during boot the kernel incorrectly
detected it as a KSZ8795. The relevant boot messages were:
ksz-switch spi0.0: found switch: KSZ8795, rev 0
ksz-switch spi0.0: Device tree specifies chip KSZ8765 but found KSZ8795, please fix it!
Cc: stable@vger.kernel.org
Fixes: 91a98917a883 ("net: dsa: microchip: move switch chip_id detection to ksz_common")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Sebastien Royen <sebastien.royen@armadeus.com>
Link: https://patch.msgid.link/20261007080557.15719-1-sebastien.royen@armadeus.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
mlx5e_poll_ico_cq() requires sq->cc to be updated only after
mlx5_cqwq_update_db_record(), otherwise a CQ overrun may occur.
The current implementation updates sq->cc before the CQ doorbell
record, violating this ordering requirement.
Update the CQ doorbell record first and use dma_wmb() before updating
sq->cc. This ensures that the CQ space is released to the device
before the corresponding ICOSQ consumer index is updated by software.
Fixes: fd9b4be8002c ("net/mlx5e: RX, Support multiple outstanding UMR posts")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20261006105820.257208-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A veth device advertises NETDEV_XDP_ACT_NDO_XMIT only if its peer has an
XDP program attached or GRO enabled, that is, only if the peer will have
NAPI to receive the frames.
veth_set_features() updates the peer's flag when GRO is toggled, but
returns early if the device is down, and veth_open() only refreshes the
flags of the device being opened. Toggling GRO while the device is down
therefore leaves the peer's flag stale after the device comes up.
If GRO was enabled while down, the device comes up with NAPI but the
peer does not advertise NDO_XMIT, and devmap rejects redirects to the
peer with -EOPNOTSUPP. If GRO was disabled while down, the device comes
up without NAPI but the peer still advertises NDO_XMIT, so redirects are
accepted and then dropped in veth_xdp_xmit() with -ENXIO.
Commit 7a6102aa6df0 ("veth: Update XDP feature set when bringing up
device") made veth_open() refresh the device's own flags. Refresh the
peer's flags there too. The peer's flag depends on this device's XDP
program and GRO setting, not on whether the peer is up, so it is
correct to set it even if the peer is down.
Fixes: 8267fc71abb2 ("veth: take into account peer device for NETDEV_XDP_ACT_NDO_XMIT xdp_features flag")
Signed-off-by: Tianyi Gao <tianyi@cloudflare.com>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Acked-by: Jesper Dangaard Brouer <hawk@kernel.org>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20261006173241.65945-2-tianyi@cloudflare.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers
instead of full pages to improve memory efficiency.") started handing
out RX buffers with zero headroom so that two buffers fit into one page
at the default MTU.
The MANA TX path, however, stores the per scatter-gather entry DMA
mappings in `struct mana_skb_head` at skb->head, and mana_start_xmit()
therefore calls skb_cow_head(skb, MANA_HEADROOM). The port advertises
this requirement as ndev->needed_headroom = MANA_HEADROOM.
As a result every packet that is received and then forwarded out of a
MANA port fails the skb_cow() in ip_forward() and gets reallocated and
copied by pskb_expand_head(). This is invisible to a plain RX or TX
workload, but it puts a full skb reallocation plus memcpy on the hot
path of every single forwarded packet, which is exactly what a
router/NVA workload does.
Restore the headroom. Note that reserving MANA_HEADROOM (232) is not
enough: ip_forward() asks for LL_RESERVED_SPACE(dev), which rounds
hard_header_len + needed_headroom up to HH_DATA_MOD and is 256 bytes on
ethernet. Use LL_RESERVED_SPACE() directly so the value keeps tracking
both constants. Also, since LL_RESERVED_SPACE() tracks MANA_HEADROOM,
it grows with MAX_SKB_FRAGS and for MAX_SKB_FRAGS >= 19 it is greater
than 256, so we have to account for that by using the headroom the RX
queue actually uses (instead of assuming XDP_PACKET_HEADROOM) and
turning MANA_XDP_MTU_MAX into MANA_XDP_MTU_MAX(ndev) (note that at the
default CONFIG_MAX_SKB_FRAGS=17 they are equivalent).
At the default MTU on a 4K page this means a buffer no longer fits twice
into a page (SKB_DATA_ALIGN(1500 + MANA_RXBUF_PAD + 256) = 2112), so the
frag-vs-single decision is now made by computing the real buffer size
instead of comparing the MTU against PAGE_SIZE / 2. The page_pool
fragment path is still used wherever at least two buffers genuinely fit,
e.g. on 16K and 64K page sizes.
Measured on an Azure VM with a MANA NIC acting as a forwarding NVA (UDP,
1400 byte payload, 4 streams, 8 Gbps offered, only the forwarding
node's kernel differs), 8 runs each, median:
forwarded pps throughput
before 272,830 3.06 Gbps
after 390,560 4.37 Gbps (+43%)
perf on the forwarding node, same workload:
memset_orig __pi_memcpy pskb_expand_head
before 10.07% 3.96% present
after 0.94% 0.64% gone
Cc: stable@vger.kernel.org
Fixes: 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers instead of full pages to improve memory efficiency.")
Signed-off-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20261003013647.2051416-1-hamzamahfooz@linux.microsoft.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
https://gitlab.freedesktop.org/drm/i915/kernel into drm-fixes
drm/i915 fixes for v7.3-rc7:
- Fix a GT park race that could trigger spurious RPM wakelock warnings
- Fix execbuffer relocation cleanup to avoid out-of-bounds free
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Jani Nikula <jani.nikula@intel.com>
Link: https://patch.msgid.link/49ac4a9f496e1dadbf71740eb0d1f443598b2632@intel.com
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux
Pull power sequencing fixes from Bartosz Golaszewski:
- add missing PCI device IDs for Thinkpad T14s gen6 which should have
been part of commit a39ac4651e3b ("power: sequencing: pcie-m2: Match
WCN6855 and WCN7851 UART BT variants by subdevice ID") in
pwrseq-pcie-m2
- fix memory leak in pwrseq-pcie-m2 (leaking the array returned by
of_regulator_bulk_get_all())
* tag 'pwrseq-fixes-for-v7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
power: sequencing: pcie-m2: Fix leaking array from of_regulator_bulk_get_all()
power: sequencing: pcie-m2: Add Lenovo ThinkPad T14s gen6 WCN7850 subsystem PCI ids
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux
Pull gpio fixes from Bartosz Golaszewski:
- fix runtime PM leak in error path in gpio-xilinx
- fix race when arming the IRQ poll worker in gpio-mpsse
- fix devres cleanup path on probe error in gpio-exar
* tag 'gpio-fixes-for-v7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
gpio: mpsse: fix race when arming the IRQ poll worker
gpio: exar: initialize the ID before registering its cleanup
gpio: xilinx: fix runtime PM leak on request error path
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull MM fixes from Andrew Morton:
- Update .mailmap entries for Andy Yan and John Garry
- Fix read-only MAP_SHARED /dev/zero mappings so they retain
shared-file semantics instead of being treated as anonymous memory,
also avoiding a CONFIG_DEBUG_VM assertion
- Fix 32-bit build warnings in the hugetlb-mmap selftest caused by
using the wrong printf format for size_t values
- Fix a boot-time crash when early function tracing causes CPA to free
kernel page tables before the workqueues used for deferred freeing
are available
- Fix two MREMAP_DONTUNMAP locked_vm accounting leaks: one caused by
an mlock-on-fault VMA self-merging, and one caused by partially
remapping a locked VMA
* tag 'mm-hotfixes-stable-2026-10-07-21-48' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
mailmap: update entry for Andy Yan
drivers/char/mem: mmap readonly MAP_SHARED-/dev/zero correctly
selftests/mm: cleanup -Wformat issues in hugetlb-mmap
mm: don't schedule deferred kernel page table freeing while booting
mailmap: update addresses for John Garry
mm/mremap: fix locked_vm leak by splitting VMA for MREMAP_DONTUNMAP
mm/mremap: fix locked_vm leak from MREMAP_DONTUNMAP self-merge
|
|
Commit 0a8224058a58 ("drm/imagination: Fix page count for page table for
map() interface") passed the device address to
pvr_mmu_op_context_create(), but the preallocation is still sized from
device_addr + sgt_offset with an exclusive end. The offset into the
object's pages has no place in a device-virtual range, and the exclusive
end allocates one table too many when the range ends on a table
boundary.
Count the tables from device_addr and the inclusive end of the range.
The new calculation does not take page table indices into account. This
fixes another issue with the previous implementation, where the
difference between start and end indices could underflow when a mapping
crossed a page table boundary. This would lead to the driver
preallocating a huge number of MMU pages, exhausting system memory and
triggering an out of memory panic.
Fixes: 0a8224058a58 ("drm/imagination: Fix page count for page table for map() interface")
Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code")
Signed-off-by: Gyeyoung Baek <gye976@gmail.com>
Reviewed-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
Reviewed-by: Alessio Belle <alessio.belle@imgtec.com>
Cc: stable@vger.kernel.org
[Alessio: add paragraph to point out OOM bug being fixed, cc stable]
Link: https://patch.msgid.link/20261001-pvr-fixes-a-v2-2-f58254e5dfb8@gmail.com
Signed-off-by: Alessio Belle <alessio.belle@imgtec.com>
|
|
When a map overlaps part of an existing mapping, pvr_vm_gpuva_remap()
splits the mapping into prev/next parts covering what the request did
not take, instead of creating something new. It gets two things wrong.
- A GEM reference is taken for each part but it's never needed and never
dropped, so it leaks one reference per split (remap-next-in-2m in
tests/imagination/pvr_vm_map.c):
CRITICAL: 33558528 bytes of shmem still held after close
- The parts still belong to the original object being split, but
pvr_vm_gpuva_remap() links them to ctx->gpuvm_bo, the new object's
vm_bo:
prev_va --obj--> BO_A BO_B <--obj-- vm_bo (ctx->gpuvm_bo)
\___________link___________/ (mismatch: BO_A != BO_B)
drm_gpuva_link() catches the mismatch:
WARNING: drivers/gpu/drm/drm_gpuvm.c:2108 at drm_gpuva_link+0x2ec/0x310
drm_WARN_ON(obj != vm_bo->obj)
Call trace:
drm_gpuva_link
pvr_vm_gpuva_remap
__drm_gpuvm_sm_map
pvr_vm_map
pvr_ioctl_vm_map
Link them to op->remap.unmap->va->vm_bo instead.
The locking of the GPUVA lists of the other objects touched by a split is
not addressed here; it is handled by switching the GPUVM to immediate
mode.
Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code")
Signed-off-by: Gyeyoung Baek <gye976@gmail.com>
Reviewed-by: Brajesh Gupta <brajesh.gupta@imgtec.com>
Link: https://patch.msgid.link/20261001-pvr-fixes-a-v2-1-f58254e5dfb8@gmail.com
Signed-off-by: Alessio Belle <alessio.belle@imgtec.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/soc/soc
Pull SoC fixes from Arnd Bergmann:
- Four distinct issues in TEE firmware, all fairly minor
- Five devicetree mistakes on NXP i.MX8, lx2160a and Qualcomm
based machines, one of these may cause file system corruption
from an incorrect SD card supply voltage
- Five fixes for clk drivers on new Qualcomm platforms,
addressing issues with incorrect enable states
* tag 'soc-fixes-7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/soc/soc:
MAINTAINERS: update my email address
tee: optee: ffa: support shared memory offsets on large-page kernels
optee: register TEE devices only once fully initialized
tee: shm: reject zero-sized allocations in tee_dyn_shm_alloc_helper()
arm64: dts: imx8mp-var-dart-sonata: Fix Sonata SD I/O supply
arm64: dts: lx2160a: fix the iic5 spi3 pinmux offset and value
arm64: dts: lx2160a: fix IIC1 pinmux submask rejected by pinctrl-single
arm64: dts: lx2160a: fix incorrect pinmux
clk: qcom: gpucc-kaanapali: Mark the GPU CX GDSC as votable
clk: qcom: gpucc-glymur: Mark the GPU CX GDSC as votable
clk: qcom: gcc-kaanapali: Fix always-enabling PCIE_RSCC clocks
clk: qcom: gcc-hawi: Fix always-enabling PCIE_RSCC clocks
clk: qcom: gcc-eliza: Fix always-enabling PCIE_RSCC clocks
arm64: dts: qcom: x1-denali: Fix microphone distortion
tee: qcomtee: fix kernel-doc warnings
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/paulmck/linux-rcu
Pull tracing fix from Paul McKenney:
"Fix a double-dereference splat in TP_printk in usb mtu3. This was a
pre-existing bug that can corrupt tracing output, but which was
exposed this cycle by additional checking that was added to the
tracing subsystem.
This splat is reporting a double-dereference that can result in
garbage traces being dumped due to the possibility of the TP_printk()
being executed without the benefit of main memory being present.
The fix is to move the extra dereference from TP_printk() time to
TP_fast_assign() time"
* tag 'usb-mtu3.2026.10.06a' of git://git.kernel.org/pub/scm/linux/kernel/git/paulmck/linux-rcu:
usb: mtu3: Fix double dereference in TP_printk
|
|
bareudp_xmit_skb() sets the inner protocol of the skb to the configured
ethertype, and bareudp6_xmit_skb() does not set it at all. The inner
protocol is used by skb_udp_tunnel_segment() to segment the inner packet
when GSO has to be done in software, for instance when the lower device
does not offload the checksum.
With multiproto, an IPv4 device also carries IPv6, so the IPv6 header
is parsed as an IPv4 one. With IPv6 underlay, the inner protocol is
whatever the skb had before. In both cases segmentation fails and the
packets are dropped. With veth tx checksum offload disabled, iperf3 TCP
over bareudp gets about 10 Mbit/s with thousands of retransmits instead
of about 700 Mbit/s for IPv6 over IPv4 and for both families over IPv6,
while IPv4 over IPv4 works.
At this point skb->protocol is the protocol of the inner packet, which
bareudp_xmit() has already checked against the configuration, so use it
on both paths.
Fixes: 571912c69f0e ("net: UDP tunnel encapsulation module for tunnelling different protocols like MPLS, IP, NSH etc.")
Fixes: 4b5f67232d95 ("net: Special handling for IP & MPLS.")
Signed-off-by: Haishuang Yan <yanhaishuang@cmss.chinamobile.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20261002153039.462663-1-yanhaishuang@cmss.chinamobile.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Javier reports seeing packet loss issues related to the re-enablement
of K1; disabling it makes the issues go away. Add the reported system
to have K1 disabled by default.
Reproduction steps (from the Link):
Plug the cable and ping the default gateway on the LAN. No suspend/re |