| Age | Commit message (Collapse) | Author | Files | Lines |
|
Commit 680aea08e78c ("net: ipv4: Emit notification when fib hardware
flags are changed") added asynchronous route notifications for hardware
flag changes.
fib_alias_hw_flags_set() sizes the skb with fib_nlmsg_size() and later
fills it with fib_dump_info() while holding only RCU. With nexthop
compatibility mode enabled, a concurrent replacement can grow the group
between these independent snapshots. fib_dump_info() can then
legitimately return -EMSGSIZE, so the warning does not prove a sizing bug.
Remove the warning. The existing error path still frees the skb and
reports the error to listeners.
Fixes: 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed")
Reported-by: Ido Schimmel <idosch@nvidia.com>
Link: https://lore.kernel.org/netdev/20261004082043.GB92032@shredder/
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Link: https://patch.msgid.link/b7e203fc5867651cb67ef84aa03e082f0f76008e.1791423190.git.4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
If a packet with bad checksum metadata gets into the ipv4 stack,
skb_checksum_help can corrupt the network header and cause a bunch of
mischief.
This was discovered and reported by Paulos, and has been reporoduced
by others independently since.
We really shouldn't allow such packets in, but as a defence
in depth measure, let's also check before we complete the checksum.
A more complete validation at input is forthcoming, but needs more work.
Cc: stable@vger.kernel.org
Fixes: f43798c27684 ("tun: Allow GSO using virtio_net_hdr")
Fixes: bfd5f4a3d605 ("packet: Add GSO/csum offload support.")
Reported-by: Paulos Yibelo <habte.yibelo@gmail.com>
Closes: https://lore.kernel.org/netdev/20260922030310.8684-2-habte.yibelo@gmail.com/
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/0a012b4923e189c4c593ef4f471e5ff0edbe9030.1791412497.git.mst@redhat.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
cipso_v4_skbuff_delattr() removes the CIPSO bytes but does not adjust
cached offsets for options that follow them.
For example, with a valid 10-byte CIPSO option followed by a seven-byte
RR option, parsing records rr at offset 30. Removing CIPSO moves RR to
offset 20, while the cached offset remains 30. Consumers such as
ip_forward_options() and __ip_options_echo() then access the wrong bytes;
the latter may interpret packet data as the option length and copy it into
fixed-size option storage.
Mirror cipso_v4_delopt() and subtract cipso_len from the srr, rr, ts and
router_alert offsets when they follow CIPSO. cipso_len is the distance
the first memmove() shifts those options. The later header move and
network-header reset relocate the bytes and their offset base together.
Fixes: 89aa3619d141 ("cipso: make cipso_v4_skbuff_delattr() fully remove the CIPSO options")
Cc: stable@vger.kernel.org
Reviewed-by: Ondrej Mosnáček <omosnacek@gmail.com>
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Acked-by: Paul Moore <paul@paul-moore.com>
Link: https://patch.msgid.link/20261006035108.3101440-1-4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 680aea08e78c ("net: ipv4: Emit notification when fib
hardware flags are changed") added an RCU-only fib_nlmsg_size() call
for asynchronous hardware flag notifications.
fib_nlmsg_size() checks fib_info_num_path() before each iteration,
then fib_info_nhc() independently reloads nh->nh_grp. If
RTM_NEWNEXTHOP replaces a group with fewer paths between those loads,
fib_info_nhc() returns NULL and fib_nexthop_nlmsg_size() dereferences
it.
Stop sizing when the indexed path is absent. Nexthop groups are
dense, so the current snapshot has no later path.
Fixes: 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed")
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/1054fa46165bba1b473782fa41367804aa908e86.1790915964.git.4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 4ce5dc9316de ("inet: switch inet_dump_fib() to RCU
protection") allowed IPv4 FIB dumps to run concurrently with nexthop
replacement.
fib_dump_info_fnhe() uses fib_info_num_path() as the loop bound,
then fib_info_nhc() independently reloads nh->nh_grp. If
RTM_NEWNEXTHOP replaces a group with fewer paths between those loads,
fib_info_nhc() returns NULL and the dump dereferences it through
nhc_flags.
Stop the walk when the indexed path is absent. Nexthop groups are
dense, so the current snapshot has no later path.
Fixes: 4ce5dc9316de ("inet: switch inet_dump_fib() to RCU protection")
Reported-by: Ido Schimmel <idosch@nvidia.com>
Link: https://lore.kernel.org/netdev/20261001170513.GA1657889@shredder/
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/5bb7bd97411eb9db454464bf449107395e760f26.1790915964.git.4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 7d3f3b4367f3 ("net: ipv4: Cache pmtu for all packet paths if
multipath enabled") made __ip_rt_update_pmtu() update every path.
For nexthop objects, fib_info_num_path() and fib_info_nhc()
independently load nh->nh_grp. RCU protects each group's lifetime but
does not make the loads observe the same group. If replacement shrinks
the group after the loop accepts an index, fib_info_nhc() returns NULL
and update_or_create_fnhe() dereferences it.
A deterministic probe only widened the existing window between the real
operations. On v7.2, an RTM_NEWNEXTHOP replacement published a
one-member group after the PMTU reader accepted index 1 from a two-member
group, producing:
KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
RIP: update_or_create_fnhe+0x45/0x15b0
Stop when the indexed path is absent. Groups are dense, so no later
index exists in that snapshot. The same real-writer window completed
114,423 PMTU iterations without an oops with this check. A reproducer
is available on request.
Replacement requires CAP_NET_ADMIN in the network namespace. Where
unprivileged user namespaces are permitted, a local user can obtain it
in a new user and network namespace.
Fixes: 7d3f3b4367f3 ("net: ipv4: Cache pmtu for all packet paths if multipath enabled")
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/4fb87ed994dfc067178fc845cdd8bc9b5446b039.1790915964.git.4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
tcp_v4_syn_recv_sock() transfers ownership of ireq->ireq_opt to the
child socket (newinet->inet_opt) without copying it.
Another cpu can concurrently process a retransmitted SYN for the same
request socket, and send a SYNACK from tcp_check_req().
tcp_v4_send_synack() and inet_csk_route_req() read ireq->ireq_opt
under rcu_read_lock() only, and ip_build_and_send_pkt() and
ip_options_build() then read opt->optlen twice.
Note that the SYNACK timer itself is not an issue: it holds a
reference on its request socket, and inet_csk_reqsk_queue_drop()
calls timer_delete_sync() before the child can be freed.
Since commit 079096f103fa ("tcp/dccp: install syn_recv requests
into ehash table"), request sockets are processed without holding
the listener lock, so nothing prevents the child socket from being
freed while the SYNACK is still being built. TCP child sockets do
not have SOCK_RCU_FREE, and inet_sock_destruct() frees inet_opt
with a plain kfree(), leading to a use-after-free in
ip_options_build().
A similar issue exists with request socket migration
(net.ipv4.tcp_migrate_req=1, or a BPF_SK_REUSEPORT_SELECT_OR_MIGRATE
program). reqsk_timer_handler() clones the request socket with
inet_reqsk_clone(), so that the old request socket and its clone
share the same ireq_opt, then reqsk_migrate_reset() clears the
pointer in the old one. Another cpu holding a reference on the old
request socket can still be using these options (sending a SYNACK,
or creating a child in tcp_v4_syn_recv_sock()) when the clone is
freed, and tcp_v4_reqsk_destructor() also uses a plain kfree().
Readers already use RCU, and other paths replacing inet_opt
(do_ip_setsockopt(), cipso_v4_sock_setattr()...) already use
kfree_rcu(). Use kfree_rcu() in inet_sock_destruct() and
tcp_v4_reqsk_destructor() as well.
IPv6 is not affected by the first issue, because tcp_v6_syn_recv_sock()
duplicates the options. tcp_v6_reqsk_destructor() has the same
migration issue with ipv6_opt, which is only set by CALIPSO. This
will be addressed in a separate patch.
Fixes: 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table")
Fixes: c905dee62232 ("tcp: Migrate TCP_NEW_SYN_RECV requests at retransmitting SYN+ACKs.")
Reported-by: Xinyang Ge <xinyang@anthropic.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20261001221253.2964024-1-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
After copying a readable prefix, receive_fallback_to_copy() asks
tcp_zerocopy_set_hint_for_skb() where page mapping can resume. If the
next skb is unreadable, find_next_mappable_frag() passes its net_iov
fragment to can_map_frag().
skb_frag_page() returns NULL for a net_iov, but can_map_frag()
dereferences it in PageCompound(). A v7.2 KASAN run on a connected TCP
socket with 64 readable bytes followed by a 4096-byte NET_IOV_DMABUF
fragment reported:
BUG: KASAN: null-ptr-deref in can_map_frag
tcp_zerocopy_receive -> can_map_frag
Kernel panic - not syncing: KASAN: panic_on_warn set
The diagnostic inserted the net_iov directly because the test host has
no devmem-capable NIC. Hardware end-to-end reachability remains untested
and requires CONFIG_NET_DEVMEM plus a supported DMA-buf-bound RX queue.
Reject all net_iov fragments before skb_frag_page(). This covers both
DMABUF and IOURING net_iov types while leaving page-backed checks
unchanged. With the guard, the same queue copied the readable prefix,
returned a 4096-byte skip hint, and completed without a fault.
Fixes: 9f6b619edf2e ("net: support non paged skb frags")
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Reviewed-by: Mina Almasry <almasrymina@google.com>
Link: https://patch.msgid.link/20260930102524.1659847-1-4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A TCP packet can carry new data while acknowledging traffic in the
opposite direction. With overlapping traffic in both directions, a
delayed packet's acknowledgment can be older than one Linux has already
accepted, even when that packet fills a gap in the received data.
Linux accepts the data, but tcp_ack() takes the old_ack path and skips
updating TS.Recent, the timestamp saved for outgoing acknowledgments.
The reply therefore echoes an older timestamp. If the sender uses this
echo to measure round-trip time after a long idle period, its estimate
includes the idle time and can reduce its sending rate.
Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK
processing can trigger a transmission. This reuses the existing timestamp
and sequence checks, including PAWS protection against old duplicate
packets. ACK validation already rejects old ACKs in SYN_RECV before this
path, so no additional state check is needed.
Echoing the timestamp of the packet that fills the receive gap follows
RFC 7323 section 4.3. In a socket reproduction with 300 seconds idle,
controlled reordering and retransmission to exercise timestamp-based RTT
sampling, the sender's smoothed round-trip time was 37.5 seconds without
the fix and 15.5 ms with it.
Fixes: 12fb3dd9dc3c ("tcp: call tcp_replace_ts_recent() from tcp_ack()")
Assisted-by: LLM sparse
Signed-off-by: Jeff Jo <jeffjo@openai.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
IP-in-IP GSO can re-enter inet_gso_segment() or ipv6_gso_segment()
for each nested IP header. encap_level tracks header bytes, not callback
depth, so a deep chain can exhaust the kernel stack. Making
inet_gso_segment() stackable introduced unbounded IPv4 nesting; IPIP
GSO/TSO later made the path reachable. The IPv6 stackable path was
introduced separately and uses the same guard.
Count IPv4 and IPv6 GSO handler entries in skb_gso_cb, initialized once
per top-level GSO operation and preserved across GRE/UDP context changes.
Use the existing IP_TUNNEL_RECURSION_LIMIT for both handlers. The first
five entries pass, and the sixth returns -EINVAL before dispatching
another GSO callback.
Fixes: 3347c9602955 ("ipv4: gso: make inet_gso_segment() stackable")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Closes: https://lore.kernel.org/all/cover.1790157745.git.zihanx@nebusec.ai/
Assisted-by: LLM
Co-developed-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Zihan Xi <zihanx@nebusec.ai>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260924051521.32568-2-zihanx@nebusec.ai
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
patches. It doesn't seem like we're creating any regressions with all
these fixes, three 'Fixes' tags here point to 7.2 commits but none are
true regression fixes. We're trying to keep the count down,
nonetheless.
Previous releases - regressions:
- net: don't require the hwtstamp NDOs when a PHY provides
timestamping
- ipv6: fix dst leak for uncached routes
- vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
packets
Previous releases - always broken:
- packet: use ubuf_info completion for TX_RING packets
- arp: terminate device name before lookup
- ipv6: do not let ipv6_find_hdr() return an offset past the packet
end
- udp: remove a disconnected socket from the 4-tuple hash table
- sctp: discard the rest of the packet on a stale-cookie error
- eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
eswitch"
[ And lots of other random network driver fixes ]
* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
tcp: prevent collapsing skbs across boundary in rtx queue
vlan: ensure sufficient headroom in vlan_dev_hard_header()
net/sched: sch_teql: fix shadowed err in __teql_resolve()
bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
llc: reserve device headroom for allocated frames
gve: DQO: reject TSO packets with an out of range MSS
gve: fix TX drop when GSO MSS is too small for hw
gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
net: flush skb_defer_nodes in dev_cpu_dead()
net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
tipc: Fix a data race on mon->peer_cnt in mon_timeout()
net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
net/smc: fix UAF on lgr list traversal in smcr_port_err()
net/rds: size a connection's path set by the transport it ends up with
nfp: hold IPsec RX state under the XArray lock
net: ena: fix MMIO read buffer leak on probe failure
net: ena: fix PHC cleanup on probe failure
net/sched: act_ct: fix helper UAF due to extensions realloc
...
|
|
When tcp_send_synack() replaces the cloned SYN skb at the head of the
retransmit queue with a copy, it frees the original with
tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack.
tp->retransmit_skb_hint keeps pointing at the freed
skbuff_fclone_cache object.
The dangling hint is read in tcp_verify_retransmit_hint() and used as
the root of the rbtree walk in tcp_xmit_retransmit_queue(). An
unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with
an attacker-supplied ICMP fragmentation-needed message, after which a
simultaneous open frees the armed SYN skb:
BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
Read of size 4 at addr ffff88800604d928 by task swapper/1/0
Call Trace:
tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
tcp_simple_retransmit (net/ipv4/tcp_input.c:3158)
tcp_v4_err (net/ipv4/tcp_ipv4.c:587)
Sync the hint to the copy.
Fixes: c31b70c9968f ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit")
Reported-by: Kimi Security Team <bug-report@moonshot.ai>
Tested-by: Weiming Shi <shiweiming@moonshot.ai>
Signed-off-by: Yilin Zhang <yilinzhang@moonshot.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix bpf_skb_change_tail() to drop the checksum offload instead of
rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)
- Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
read arbitrary memory and for untrusted read-only memory reads
(Daniel Borkmann)
- Clear scalar delta on narrowing stack spill (Daniel Borkmann)
- Set up the frame pointer for the exception callback in arm64 JIT, and
zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
element (Donggeun Yoo)
- Various fixes (Emil Tsalapatis):
- Fix bounds check underflow for skb-backed dynptrs
- Fix rx_queue_mapping context access code generation in bpf_sock
- Reject packet pointer arguments to subprogs that may mutate the
packet
- Reject ALU instructions that see arena and non-arena operands on
different code paths
- Fix copied_seq double-counting on sockmap self-redirect
(Geliang Tang)
- Fix divide-by-zero in btf_struct_walk() on a flexible array of
zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
(Jiayuan Chen)
- Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
and request socks, and sleeping under RCU when destroying a listener
with pending children (Jiayuan Chen)
- Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)
- Avoid soft lockup in htab lookup[_and_delete] batch operations on
large maps (Jose Fernandez)
- Various fixes (Kumar Kartikeya Dwivedi):
- Verify global subprogs in each sleepability context they are
called from
- Make post-verification instruction rewrites killable
- Preserve packet pointer displacement in regsafe()
- Apply CO-RE relocations before subprogram validation, restrict
CO-RE poisoning to relocatable instructions, and reject truncated
ldimm64 CO-RE relocations in libbpf
- Assign lock identity to callback map values
- Compare stack frames in regs_exact()
- Bound ownership depth through local kptrs and graph roots
- Fix u32 overflow in map batch operations when the map size exceeds
4GB (Masoud Aghasi)
- Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
in alloc_bulk() (Pu Lehui)
- Allow gotox as the terminal instruction of a program or a subprogram
(Siddharth Chintamaneni)
- Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
in link iterator, and reject dev-bound-only programs on other devices
(Weiming Shi)
- Reject non-negative stack offsets in stack_slot_obj_get_spi()
(Xu Yunxiang)
- Check params size before reading reserved fields in
bpf_crypto_ctx_create() (Yuqi Xu)
- Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)
- Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)
- Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
bpf: Fix BSWAP 32 and 16 on MIPS64
bpf: Fix immediate JMP JEQ/JNE on MIPS32
bpf: Reject dev-bound-only programs on other devices
bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
selftests/bpf: Test for mixed arena/nonarena code paths
bpf: Prevent variable arena/non-arena register contents
selftests/bpf: Test rejection of pkt args to mutating subprogs
bpf: Reject pkt arguments in mutating subprogs
selftests/bpf: Add selftests for rx_queue_mapping context access
bpf: Fix bpf_sock context code generation
selftests/bpf: Test dynptr slices past end of skb
bpf: Fix bounds check for skb-backed dynptrs
selftests/bpf: Reject iterator destruction through fp+0
bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
bpf: Check params size before reading reserved fields
selftests/bpf: Check local object ownership depth
bpf: Bound ownership depth through local kptrs and graph roots
selftests/bpf: Cover frame changes in bounded loops
...
|
|
ic_dhcp_init_options() appends the hostname (option 12), vendor-class
(option 60) and client-ID (option 61) options into the fixed 312-byte
bootp_pkt.exten[] buffer. Only the client-ID branch checked the
remaining space; the hostname and vendor-class writes were unbounded.
A 64-byte hostname together with the maximum 252-byte dhcpclass=
identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available
bytes even before the terminating END marker, so the vendor-class memcpy
runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is
reported as a field-spanning write and, when the kernel is booted with
panic_on_warn=1, aborts boot with a panic.
Route the optional options through a common helper that makes sure the
option, its 2-byte header and the END marker all fit and drops an option
that would not. Configurations with short options keep sending exactly
the same bytes as before.
Fixes: 130c0f47fdf9 ("ipconfig: send host-name in DHCP requests")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP
or ERSPAN device. Unlike newlink, changelink does not enforce metadata
tunnel uniqueness. Converting a non-metadata device can therefore
replace the metadata receive entry for another device of the same type
in the same netns. Deleting either device then clears the shared entry,
breaking metadata receive lookup for the surviving device.
If parameter validation fails after collect_md is set, deleting the
modified device can also clear an entry it never owned.
Reject enabling metadata mode in both changelink callbacks before any
encapsulation or tunnel parameters are modified. Allow requests that
repeat the metadata attribute on an existing metadata device.
Fixes: 2e15ea390e6f ("ip_gre: Add support to collect tunnel metadata.")
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The ARP ioctl copies a user-provided struct arpreq into a stack object. Its
arp_dev field may contain IFNAMSIZ bytes without a NUL terminator.
Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where
strcmp() can read past the end of the stack object when a matching
alternative interface name exists.
Terminate the field before the lookup to prevent the out-of-bounds read.
Fixes: 36fbf1e52bd3 ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Zijie Huang <milkory@outlook.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added
NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which
rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with
-ERANGE.
However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user
sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits
FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and
parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0,
sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with
fou->protocol == 0.
In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb()
triggers IP protocol resubmission when fou->protocol > 0, whereas
returning 0 tells the UDP tunnel layer that the skb was consumed without
freeing it. When fou->protocol == 0, every packet received on the socket
returns 0 from fou_udp_recv() and leaks the sk_buff.
Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that
creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with
-EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share
parse_nl_config()) unaffected.
Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic
Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE =
FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel,
FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and
fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks
all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to
59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected
with -EINVAL (-22).
Fixes: 23461551c006 ("fou: Support for foo-over-udp RX path")
Fixes: 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.")
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A UDP socket bound to a specific address and port keeps its entry in the
4-tuple hash table after it is disconnected:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed in the 4-tuple table
sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0
__udp_disconnect() takes a socket out of that table only as a side effect
of ->rehash() or ->unhash(), and it skips ->rehash() when
SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set.
commit 6996a2d2d0a6 ("udp: Unhash auto-bound connected sk from 4-tuple hash
table when disconnected.") fixed the same end state for a wildcard-bound
socket, by a path this one does not take.
The entry is counted whether or not anything hits it. hash4_cnt on the
hash2 slot stays raised for as long as the socket lives, so udp_has_hash4()
keeps sending every packet for that address and port through the 4-tuple
lookup first.
On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr,
so udp_v6_rehash() files the entry under the peer the socket was connected
to with a zero dport, and inet6_match() compares that same
field: a datagram from the former peer with a zero source port matches,
and source port zero is accepted on receive. On IPv4 the peer is cleared,
so a match would need a zero source address as well, which the routing
layer rejects as martian. The stale sk_v6_daddr is a separate defect, not
addressed here; removing the entry closes this path either way.
The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if,
so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive
address is still specific udp_lib_rehash() moves the entry instead of
removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a
pure function of the address and port, so every socket reaching this state
on one address and port collects in one bucket. The bucket cannot be chosen
from outside, as udp_ehashfn() is seeded with a per-boot secret. This last
one became reachable only with commit 644f9108f3a5 ("udp: Make rehash4
independent in udp_lib_rehash()"), which moved the hash4 handling out of a
branch a disconnected socket does not take; the stale entry itself dates
from the commit in Fixes.
Take the socket out of the table before __udp_disconnect() runs, while it
still matches how it was filed. This also reaches the wildcard case ahead
of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable
from udp_disconnect(); removing it belongs in net-next. udp_disconnect()
and udp_abort() are the only UDP entries into __udp_disconnect(), which is
shared with raw, ping and l2tp sockets that are not struct udp_sock:
ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one
would read past the allocation.
Fixes: 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
A connected UDP socket that connects again to a different peer is not
re-filed in the 4-tuple hash table:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1)
sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1)
packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the
// lookup falls back to scoring the
// hash2 chain for this address
// and port
udp_lib_hash4() returns early when the socket is already hashed, assuming
->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect()
only while the receive address is unset, which a second connect never is:
the first connect assigns it, whether the socket was bound to a specific
address or to the wildcard. commit 644f9108f3a5 ("udp: Make rehash4
independent in udp_lib_rehash()") added that early return and named
connect(AF_UNSPEC) as the way around it. That workaround does not help a
socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because
__udp_disconnect() skips ->rehash() for the first and ->unhash() for the
second.
Delivery is correct either way.
Relocate the socket when the hash it is filed under differs from the one
requested, which is what commit 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash
for connected socket") did before the early return became unconditional. It
is done here under hslot->lock, which that version did not take, to match
udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt
needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected,
and IPv6 shares the code.
With 500 sockets on the port, a re-connected socket measured 522,553 pps
without this change and 2,055,078 with it. The UDP side was noted as
remaining work in [1].
Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1]
Fixes: 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Commit fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly")
added validation of IFLA_INET_CONF attributes, and in the process
changed the call of nla_for_each_nested() to nla_parse_nested(). A
side effect of this change is that the IFLA_INET_CONF option is now
tested for NLA_F_NESTED being set, and fails if it is not. Prior to the
commit there was no check of NLA_F_NESTED.
Change nla_parse_nested() to nla_parse(). This restores the previous
functionality of not checking NLA_F_NESTED, thereby allowing code that
(incorrectly) doesn't set NLA_F_NESTED to continue to work.
This issue was identified because keepalived started logging errors when
it was configuring macvlans that it created.
Fixes: fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly")
Signed-off-by: Quentin Armitage <quentin@armitage.org.uk>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
fib_select_multipath() compares nexthop_nh->nh_saddr against the flow
source address with no lock held, while fib_info_update_nhc_saddr()
stores a new value from another CPU as soon as the preferred source
address of the egress device changes.
Commit 195374d89368 ("ipv4: fib: annotate races around nh->nh_saddr_genid
and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE()
in fib_result_prefsrc() after syzbot reported
BUG: KCSAN: data-race in fib_select_path / fib_select_path
but it only covered that reader. fib_select_multipath(), reached from
fib_select_path(), is a second lockless reader of nh->nh_saddr and was
left bare.
Moreover, nh_saddr is only meaningful when nh_saddr_genid matches
dev_addr_genid, as established by commit 436c3b66ec98 ("ipv4: Invalidate
nexthop cache nh_saddr more correctly."). fib_select_multipath()
skips that validation, so it can score a nexthop using a stale source
address and skew the ECMP selection.
Annotate both reads with READ_ONCE() and refresh the cached source
address via fib_info_update_nhc_saddr() when the genid does not match,
mirroring fib_result_prefsrc().
Fixes: 32607a332cfe ("ipv4: prefer multipath nexthop that matches source address")
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260916125316.988044-1-xiaolinkui@126.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Exclude old ACKs before SND.UNA from the tcp fast path
as well as ACKs after SND.NXT.
Such ACKs will fall through to the slow path, where tcp_ack()
performs the appropriate validation and challenge ACK handling
according to RFC5961 and Commit 3d501dd326fb1c7 ("tcp: do not
accept ACK of bytes we never sent").
This prevents old ACKs from being accepted
or modifying connection state as part of the fast path before
appropriate ACK validation is applied.
In particular, this prevents payload carried by a segment with
an excessively old ACK from advancing RCV.NXT before the ACK
is rejected.
Fixes: 31770e34e43d ("tcp: Revert "tcp: remove header prediction"")
Reported-by: Amit Klein <amit.klein@mail.huji.ac.il>
Reported-by: Tamir Shahar <tamir.shahar1@mail.huji.ac.il>
Reported-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260914090408.1435080-2-inbal.lipshtat@mail.huji.ac.il
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
PSP conflicts with TLS ULP in its usage of both skb->decrypted and
sk->sk_validate_xmit_skb().
Make PSP mutually exclusive with TLS ULP, the only other user of either
of these. As other users of skb->decrypted come along, they can be added
to sk_has_decrypt_user(). It would make sense to also assert that
sk->sk_validate_xmit_skb() is also NULL in both of these setup paths for
similar future proofing, but the PSP listener/sk_clone() path is still
broken and it could be seen as a regression to not allow rx assoc to run
on a child of a listener socket with PSP tx assoc state.
Include all TCP ULPs in the sk_has_decrypt_user() check, even though TLS
is the only one that conflicts with PSP via the decrypted bit. This is
intentional because PSP was not designed to be used with ULPs. It is
best to close off surface area that may make bugs reachable, until
someone wishes to design and test an actual user of PSP with ULPs.
Fixes: 6b46ca260e22 ("net: psp: add socket security association code")
Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260915-psp-ktls-fix-v2-1-0eedc3b148ec@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec
Steffen Klassert says:
====================
pull request (net): ipsec 2026-09-16
1) xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk()
Add the up-front nr_frags guard iptfs_skb_add_frags() already has,
so an out-of-range offset can't walk past the on-stack frags[] array.
2) xfrm: serialize state GC with device state flush
Serialize xfrm_state destruction against the deferred-device pass
with a dedicated mutex, since the device GC list doesn't hold a state
reference and the two paths could free the same state.
3) xfrm: add missing RCU read lock in xfrm_send_migrate_state()
Hold the RCU read lock around xfrm_nlmsg_multicast() so the
rcu_dereference() of net->xfrm.nlsk doesn't warn.
4) xfrm: iptfs: fix runt reassembly panic from short inner tot_len
Require the runt length to cover at least the minimum IP header,
so a tot_len in [6, 19] (IPv4) can't write past the declared length
and trip skb_over_panic().
5) ipv6: xfrm: use full sockets in local error paths
Use skb_to_full_sk() in xfrm6_local_rxpmtu() and xfrm6_local_error()
and bail out without a full socket, so a TCP_NEW_SYN_RECV request_sock
isn't miscast as a full inet/IPv6 socket.
6) xfrm: fix compat ALLOCSPI request use-after-free
Drop the redundant alloc_compat() in xfrm_alloc_userspi() so the
compat translator no longer reads past the payload and publishes a
child a multicast clone can still see after xfrm_user_rcv_msg() frees.
7) xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
Force the dst before queuing, hold dev across the workqueue deferral,
and take rcu_read_lock() around the finish() loop, so transport-mode
reinjection doesn't deref non-refcounted dst/dev under workqueue.
8) xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
Switch to hlist_del_init_rcu() so a second __xfrm_state_delete() is
a no-op instead of writing through LIST_POISON2, closing the UAFs.
9) esp: downgrade zerocopy managed frags before mutating skb frags
Call skb_zcopy_downgrade_managed() before ESP rewrites the skb frag
array, so per-frag unrefs in esp_ssg_unref() and skb_release_data()
stay balanced for ubuf-owned managed frags.
10) xfrm: hold net_device reference under RCU in bundle creation
Read dst->dev via dst_dev_rcu() and keep RCU active through
xfrm_fill_dst(), so a concurrent RTM_DELLINK can't free dev
under bundle creation.
11) xfrm: save input state data before secpath resets
Save the state protocol on the stack while it's still valid and
use the saved address family for transport_finish(), so post-reset
dereferences (VTI, XFRM if, MAX_DEPTH error) can't UAF the state.
12) net: xfrm: reject unrepresentable espintcp transport headers
Use the careful transport-header helper and drop the skb through
the XFRM error path when the offset can't be represented, instead
of silently truncating it.
* tag 'ipsec-2026-09-16' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec:
net: xfrm: reject unrepresentable espintcp transport headers
xfrm: save input state data before secpath resets
xfrm: hold net_device reference under RCU in bundle creation
esp: downgrade zerocopy managed frags before mutating skb frags
xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
xfrm: fix compat ALLOCSPI request use-after-free
ipv6: xfrm: use full sockets in local error paths
xfrm: iptfs: fix runt reassembly panic from short inner tot_len
xfrm: add missing RCU read lock in xfrm_send_migrate_state()
xfrm: serialize state GC with device state flush
xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk()
====================
Link: https://patch.msgid.link/20260916101938.118628-1-steffen.klassert@secunet.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The following command triggers a kernel panic:
ip link add d0 type dummy; ip link set d0 up
ip route add 10.30.0.0/16 \
encap ip id 300 geneve_opts 4660:66:11223344 dev d0
memcpy: detected buffer overflow: 4 byte write of buffer size 0
kernel BUG at lib/string_helpers.c:1044!
...
ip_tun_parse_opts.part.0.cold+0x10/0x10
ip_tun_build_state+0x116/0x2a0
On kernels built with GCC 15+ and `CONFIG_FORTIFY_SOURCE`, the fortified
`memcpy()` got 0 sized destination with request of 4 bytes length:
static int ip_tun_parse_opts_geneve(...)
{
...
attr = tb[LWTUNNEL_IP_OPT_GENEVE_DATA];
data_len = nla_len(attr); /* == 4 */
struct geneve_opt *opt = ip_tunnel_info_opts(info) + opts_len;
memcpy(opt->opt_data, nla_data(attr), data_len);
/* ^^^^^^^^^^^^^ 0 since options_len is assigned afterwards */
Fixed by initializing the counter before the options are referenced.
Matching what `tunnel_key_opts_set()` already does.
Fixes: bb5e62f2d547 ("net: Add options as a flexible array to struct ip_tunnel_info")
Cc: stable@vger.kernel.org
Signed-off-by: Gris Ge <cnfourt@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Gustavo A. R. Silva <gustavoars@kernel.org>
Link: https://patch.msgid.link/20260913090851.468216-1-cnfourt@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
We can hit a division by zero crash in tcp_rcvbuf_grow()
and tcp_rcv_space_adjust():
divide error: 0000 [#1] PREEMPT SMP
RIP: 0010:tcp_rcvbuf_grow+0x187/0x450 net/ipv4/tcp_input.c:939
...
grow = div_u64(((u64)rcvwin << 1) * (newval - oldval), oldval);
The division uses oldval = tp->rcvq_space.space as divisor.
When tp->rcvq_space.space is zero, this leads to a divide-by-zero
exception.
tp->rcvq_space.space is initialized in tcp_init_buffer_space():
tp->rcvq_space.space = min3(tp->rcv_ssthresh, tp->rcv_wnd,
(u32)TCP_INIT_CWND * tp->advmss);
If tcp_rmem[1] is configured to very small values (such as 1),
sk->sk_rcvbuf is initialized to 1. Then tcp_full_space(sk), which
computes (sk->sk_rcvbuf * scaling_ratio) >> 8, truncates to 0.
This sets tp->window_clamp = 0, tp->rcv_ssthresh = 0, and
tp->rcvq_space.space = 0. Later, when data arrives and DRS is invoked,
tcp_rcvbuf_grow() divides by oldval == 0.
Back in 2015, commit b1cb59cf2efe ("net: sysctl_net_core: check SNDBUF
and RCVBUF for min length") ensured that net.core.rmem_default and
net.core.rmem_max cannot be set below SOCK_MIN_RCVBUF. Similarly,
SO_RCVBUF setsockopt enforces max_t(int, val * 2, SOCK_MIN_RCVBUF).
However, net.ipv4.tcp_rmem still had .extra1 = SYSCTL_ONE, allowing
arbitrarily small values.
Because SOCK_MIN_RCVBUF depends on sizeof(struct sk_buff) and cacheline
alignment, its value varies across architectures and configuration options.
Using a fixed constant of 4096 ensures a predictable, architecture-
independent lower bound that is safely above SOCK_MIN_RCVBUF everywhere
and matches the documented 4K default.
Fix this by setting tcp_rmem.extra1 to 4096 and updating the documentation.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260912144848.3448026-1-edumazet@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
When the forward output route cannot be used in icmp_route_lookup(),
it enters the "reverse path" and calls ip_route_input() on fl4_dec.daddr,
the original packet's source address.
ip_route_input() only returns an error for truly invalid packets. For
unreachable addresses it will succeed and return an input route whose
dst.output is set to ip_rt_bug(). The existing check only rejects
RTN_LOCAL routes, so the RTN_UNREACHABLE route types can still be returned
and later used for output, syzkaller triggering a WARN_ON_ONCE()
in ip_rt_bug() as bellow:
------------[ cut here ]------------
WARNING: net/ipv4/route.c:1273 at ip_rt_bug+0x14/0x20
RIP: 0010:ip_rt_bug+0x14/0x20
Call Trace:
ip_push_pending_frames+0xfa/0x100
__icmp_send+0x905/0xf10
ip_options_compile+0xc0/0xd0
ip_rcv_finish_core+0x321/0xae0
ip_rcv+0x1de/0x260
__netif_receive_skb_one_core+0x11a/0x130
netif_receive_skb+0x7b/0x260
tun_get_user+0x11bf/0x1c10
------------[ cut here ]------------
Reject input route that is RTN_UNREACHABLE to fix it. The net warning
is only printed for RTN_LOCAL, as RTN_UNREACHABLE is not the result of
a race condition.
Fixes: 8b7817f3a959 ("[IPSEC]: Add ICMP host relookup support")
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Dong Chenchen <dongchenchen2@huawei.com>
Link: https://patch.msgid.link/20260910140042.1880242-1-dongchenchen2@huawei.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
bpf_sock_destroy() runs from the tcp iterator, under rcu_read_lock(). If
the sock is a listener that still has children in its accept queue,
tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched()
there trips the debug check:
BUG: sleeping function called from invalid context at net/ipv4/inet_connection_sock.c:1523
in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name: test_progs
preempt_count: 0, expected: 0
RCU nest depth: 1, expected: 0
locks held by test_progs/628: 3, last CPU#3:
#0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at: bpf_seq_read+0x56/0x1210
#1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at: bpf_iter_tcp_seq_show+0x32b/0x4b0
#2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at: bpf_iter_run_prog+0x46b/0xde0
CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G W 7.2.0+ #65 PREEMPT
Tainted: [W]=WARN
Call Trace:
<TASK>
dump_stack_lvl+0xc1/0xf0
dump_stack+0x10/0x20
__might_resched+0x3d2/0x610
inet_csk_listen_stop+0x7b/0xbf0
tcp_abort+0x23b/0x3b0
bpf_sock_destroy+0xfc/0x140
bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a
bpf_iter_run_prog+0x538/0xde0
bpf_iter_tcp_seq_show+0x26b/0x4b0
bpf_seq_read+0x424/0x1210
vfs_read+0x197/0xe40
ksys_read+0x119/0x240
__x64_sys_read+0x72/0xc0
x64_sys_call+0x647/0x27e0
do_syscall_64+0xe5/0x610
entry_SYSCALL_64_after_hwframe+0x76/0x7e
RIP: 0033:0x7fad39b28aca
RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000
RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca
RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014
RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003
R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000
</TASK>
The commit that added the kfunc already guards lock_sock() in tcp_abort()
and udp_abort() with has_current_bpf_ctx(), but missed the listener path.
Do the same for the cond_resched(). The loop runs inside the iterator's
rcu_read_lock(), it must not reschedule or report a quiescent state there.
Fixes: 4ddbcb886268 ("bpf: A |