aboutsummaryrefslogtreecommitdiff
path: root/net/ipv4
AgeCommit message (Collapse)AuthorFilesLines
2 daysipv4: do not warn on route notification size raceDaehyeon Ko1-2/+0
Commit 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed") added asynchronous route notifications for hardware flag changes. fib_alias_hw_flags_set() sizes the skb with fib_nlmsg_size() and later fills it with fib_dump_info() while holding only RCU. With nexthop compatibility mode enabled, a concurrent replacement can grow the group between these independent snapshots. fib_dump_info() can then legitimately return -EMSGSIZE, so the warning does not prove a sizing bug. Remove the warning. The existing error path still frees the skb and reports the error to listeners. Fixes: 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed") Reported-by: Ido Schimmel <idosch@nvidia.com> Link: https://lore.kernel.org/netdev/20261004082043.GB92032@shredder/ Suggested-by: Ido Schimmel <idosch@nvidia.com> Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Link: https://patch.msgid.link/b7e203fc5867651cb67ef84aa03e082f0f76008e.1791423190.git.4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2 daysipv4: validate checksum_start before completing checksumMichael S. Tsirkin1-3/+7
If a packet with bad checksum metadata gets into the ipv4 stack, skb_checksum_help can corrupt the network header and cause a bunch of mischief. This was discovered and reported by Paulos, and has been reporoduced by others independently since. We really shouldn't allow such packets in, but as a defence in depth measure, let's also check before we complete the checksum. A more complete validation at input is forthcoming, but needs more work. Cc: stable@vger.kernel.org Fixes: f43798c27684 ("tun: Allow GSO using virtio_net_hdr") Fixes: bfd5f4a3d605 ("packet: Add GSO/csum offload support.") Reported-by: Paulos Yibelo <habte.yibelo@gmail.com> Closes: https://lore.kernel.org/netdev/20260922030310.8684-2-habte.yibelo@gmail.com/ Signed-off-by: Michael S. Tsirkin <mst@redhat.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/0a012b4923e189c4c593ef4f471e5ff0edbe9030.1791412497.git.mst@redhat.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 dayscipso: adjust cached option offsets when removing CIPSODaehyeon Ko1-0/+8
cipso_v4_skbuff_delattr() removes the CIPSO bytes but does not adjust cached offsets for options that follow them. For example, with a valid 10-byte CIPSO option followed by a seven-byte RR option, parsing records rr at offset 30. Removing CIPSO moves RR to offset 20, while the cached offset remains 30. Consumers such as ip_forward_options() and __ip_options_echo() then access the wrong bytes; the latter may interpret packet data as the option length and copy it into fixed-size option storage. Mirror cipso_v4_delopt() and subtract cipso_len from the srr, rr, ts and router_alert offsets when they follow CIPSO. cipso_len is the distance the first memmove() shifts those options. The later header move and network-header reset relocate the bytes and their offset base together. Fixes: 89aa3619d141 ("cipso: make cipso_v4_skbuff_delattr() fully remove the CIPSO options") Cc: stable@vger.kernel.org Reviewed-by: Ondrej Mosnáček <omosnacek@gmail.com> Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Acked-by: Paul Moore <paul@paul-moore.com> Link: https://patch.msgid.link/20261006035108.3101440-1-4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
4 daysipv4: stop route notification sizing when nexthop group shrinksDaehyeon Ko1-0/+3
Commit 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed") added an RCU-only fib_nlmsg_size() call for asynchronous hardware flag notifications. fib_nlmsg_size() checks fib_info_num_path() before each iteration, then fib_info_nhc() independently reloads nh->nh_grp. If RTM_NEWNEXTHOP replaces a group with fewer paths between those loads, fib_info_nhc() returns NULL and fib_nexthop_nlmsg_size() dereferences it. Stop sizing when the indexed path is absent. Nexthop groups are dense, so the current snapshot has no later path. Fixes: 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed") Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/1054fa46165bba1b473782fa41367804aa908e86.1790915964.git.4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
4 daysipv4: stop exception dump when nexthop group shrinksDaehyeon Ko1-0/+3
Commit 4ce5dc9316de ("inet: switch inet_dump_fib() to RCU protection") allowed IPv4 FIB dumps to run concurrently with nexthop replacement. fib_dump_info_fnhe() uses fib_info_num_path() as the loop bound, then fib_info_nhc() independently reloads nh->nh_grp. If RTM_NEWNEXTHOP replaces a group with fewer paths between those loads, fib_info_nhc() returns NULL and the dump dereferences it through nhc_flags. Stop the walk when the indexed path is absent. Nexthop groups are dense, so the current snapshot has no later path. Fixes: 4ce5dc9316de ("inet: switch inet_dump_fib() to RCU protection") Reported-by: Ido Schimmel <idosch@nvidia.com> Link: https://lore.kernel.org/netdev/20261001170513.GA1657889@shredder/ Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/5bb7bd97411eb9db454464bf449107395e760f26.1790915964.git.4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
4 daysipv4: stop PMTU walk when nexthop group shrinksDaehyeon Ko1-0/+2
Commit 7d3f3b4367f3 ("net: ipv4: Cache pmtu for all packet paths if multipath enabled") made __ip_rt_update_pmtu() update every path. For nexthop objects, fib_info_num_path() and fib_info_nhc() independently load nh->nh_grp. RCU protects each group's lifetime but does not make the loads observe the same group. If replacement shrinks the group after the loop accepts an index, fib_info_nhc() returns NULL and update_or_create_fnhe() dereferences it. A deterministic probe only widened the existing window between the real operations. On v7.2, an RTM_NEWNEXTHOP replacement published a one-member group after the PMTU reader accepted index 1 from a two-member group, producing: KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007] RIP: update_or_create_fnhe+0x45/0x15b0 Stop when the indexed path is absent. Groups are dense, so no later index exists in that snapshot. The same real-writer window completed 114,423 PMTU iterations without an oops with this check. A reproducer is available on request. Replacement requires CAP_NET_ADMIN in the network namespace. Where unprivileged user namespaces are permitted, a local user can obtain it in a new user and network namespace. Fixes: 7d3f3b4367f3 ("net: ipv4: Cache pmtu for all packet paths if multipath enabled") Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/4fb87ed994dfc067178fc845cdd8bc9b5446b039.1790915964.git.4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
5 daysipv4: free inet_opt and ireq_opt after an RCU grace periodEric Dumazet2-2/+2
tcp_v4_syn_recv_sock() transfers ownership of ireq->ireq_opt to the child socket (newinet->inet_opt) without copying it. Another cpu can concurrently process a retransmitted SYN for the same request socket, and send a SYNACK from tcp_check_req(). tcp_v4_send_synack() and inet_csk_route_req() read ireq->ireq_opt under rcu_read_lock() only, and ip_build_and_send_pkt() and ip_options_build() then read opt->optlen twice. Note that the SYNACK timer itself is not an issue: it holds a reference on its request socket, and inet_csk_reqsk_queue_drop() calls timer_delete_sync() before the child can be freed. Since commit 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table"), request sockets are processed without holding the listener lock, so nothing prevents the child socket from being freed while the SYNACK is still being built. TCP child sockets do not have SOCK_RCU_FREE, and inet_sock_destruct() frees inet_opt with a plain kfree(), leading to a use-after-free in ip_options_build(). A similar issue exists with request socket migration (net.ipv4.tcp_migrate_req=1, or a BPF_SK_REUSEPORT_SELECT_OR_MIGRATE program). reqsk_timer_handler() clones the request socket with inet_reqsk_clone(), so that the old request socket and its clone share the same ireq_opt, then reqsk_migrate_reset() clears the pointer in the old one. Another cpu holding a reference on the old request socket can still be using these options (sending a SYNACK, or creating a child in tcp_v4_syn_recv_sock()) when the clone is freed, and tcp_v4_reqsk_destructor() also uses a plain kfree(). Readers already use RCU, and other paths replacing inet_opt (do_ip_setsockopt(), cipso_v4_sock_setattr()...) already use kfree_rcu(). Use kfree_rcu() in inet_sock_destruct() and tcp_v4_reqsk_destructor() as well. IPv6 is not affected by the first issue, because tcp_v6_syn_recv_sock() duplicates the options. tcp_v6_reqsk_destructor() has the same migration issue with ipv6_opt, which is only set by CALIPSO. This will be addressed in a separate patch. Fixes: 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table") Fixes: c905dee62232 ("tcp: Migrate TCP_NEW_SYN_RECV requests at retransmitting SYN+ACKs.") Reported-by: Xinyang Ge <xinyang@anthropic.com> Signed-off-by: Eric Dumazet <edumazet@kernel.org> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20261001221253.2964024-1-edumazet@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daystcp: reject net_iov in zerocopy receive mapping hintsDaehyeon Ko1-0/+2
After copying a readable prefix, receive_fallback_to_copy() asks tcp_zerocopy_set_hint_for_skb() where page mapping can resume. If the next skb is unreadable, find_next_mappable_frag() passes its net_iov fragment to can_map_frag(). skb_frag_page() returns NULL for a net_iov, but can_map_frag() dereferences it in PageCompound(). A v7.2 KASAN run on a connected TCP socket with 64 readable bytes followed by a 4096-byte NET_IOV_DMABUF fragment reported: BUG: KASAN: null-ptr-deref in can_map_frag tcp_zerocopy_receive -> can_map_frag Kernel panic - not syncing: KASAN: panic_on_warn set The diagnostic inserted the net_iov directly because the test host has no devmem-capable NIC. Hardware end-to-end reachability remains untested and requires CONFIG_NET_DEVMEM plus a supported DMA-buf-bound RX queue. Reject all net_iov fragments before skb_frag_page(). This covers both DMABUF and IOURING net_iov types while leaving page-backed checks unchanged. With the guard, the same queue copied the readable prefix, returned a 4096-byte skip hint, and completed without a fault. Fixes: 9f6b619edf2e ("net: support non paged skb frags") Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Reviewed-by: Mina Almasry <almasrymina@google.com> Link: https://patch.msgid.link/20260930102524.1659847-1-4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
11 daystcp: refresh TS.Recent for accepted old ACKsJeff Jo1-0/+6
A TCP packet can carry new data while acknowledging traffic in the opposite direction. With overlapping traffic in both directions, a delayed packet's acknowledgment can be older than one Linux has already accepted, even when that packet fills a gap in the received data. Linux accepts the data, but tcp_ack() takes the old_ack path and skips updating TS.Recent, the timestamp saved for outgoing acknowledgments. The reply therefore echoes an older timestamp. If the sender uses this echo to measure round-trip time after a long idle period, its estimate includes the idle time and can reduce its sending rate. Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK processing can trigger a transmission. This reuses the existing timestamp and sequence checks, including PAWS protection against old duplicate packets. ACK validation already rejects old ACKs in SYN_RECV before this path, so no additional state check is needed. Echoing the timestamp of the packet that fills the receive gap follows RFC 7323 section 4.3. In a socket reproduction with 300 seconds idle, controlled reordering and retransmission to exercise timestamp-based RTT sampling, the sender's smoothed round-trip time was 37.5 seconds without the fix and 15.5 ms with it. Fixes: 12fb3dd9dc3c ("tcp: call tcp_replace_ts_recent() from tcp_ack()") Assisted-by: LLM sparse Signed-off-by: Jeff Jo <jeffjo@openai.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Signed-off-by: David S. Miller <davem@davemloft.net>
11 daysnet: gso: limit recursive IP-in-IP segmentationZihan Xi1-0/+3
IP-in-IP GSO can re-enter inet_gso_segment() or ipv6_gso_segment() for each nested IP header. encap_level tracks header bytes, not callback depth, so a deep chain can exhaust the kernel stack. Making inet_gso_segment() stackable introduced unbounded IPv4 nesting; IPIP GSO/TSO later made the path reachable. The IPv6 stackable path was introduced separately and uses the same guard. Count IPv4 and IPv6 GSO handler entries in skb_gso_cb, initialized once per top-level GSO operation and preserved across GRE/UDP context changes. Use the existing IP_TUNNEL_RECURSION_LIMIT for both handlers. The first five entries pass, and the sixth returns -EINVAL before dispatching another GSO callback. Fixes: 3347c9602955 ("ipv4: gso: make inet_gso_segment() stackable") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Closes: https://lore.kernel.org/all/cover.1790157745.git.zihanx@nebusec.ai/ Assisted-by: LLM Co-developed-by: Luxing Yin <root@tr0jan.top> Signed-off-by: Luxing Yin <root@tr0jan.top> Signed-off-by: Zihan Xi <zihanx@nebusec.ai> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260924051521.32568-2-zihanx@nebusec.ai Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-24Merge tag 'net-7.3-rc5' of ↵Linus Torvalds8-28/+99
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Including fixes from Bluetooth, NFC and Netfilter. Every week in this release is record-setting for number of posted patches. It doesn't seem like we're creating any regressions with all these fixes, three 'Fixes' tags here point to 7.2 commits but none are true regression fixes. We're trying to keep the count down, nonetheless. Previous releases - regressions: - net: don't require the hwtstamp NDOs when a PHY provides timestamping - ipv6: fix dst leak for uncached routes - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets Previous releases - always broken: - packet: use ubuf_info completion for TX_RING packets - arp: terminate device name before lookup - ipv6: do not let ipv6_find_hdr() return an offset past the packet end - udp: remove a disconnected socket from the 4-tuple hash table - sctp: discard the rest of the packet on a stale-cookie error - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged eswitch" [ And lots of other random network driver fixes ] * tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits) tcp: prevent collapsing skbs across boundary in rtx queue vlan: ensure sufficient headroom in vlan_dev_hard_header() net/sched: sch_teql: fix shadowed err in __teql_resolve() bridge: check llc_mac_hdr_init() return value in br_send_bpdu() llc: fix skb UAF and leaks on llc_mac_hdr_init() failure llc: reserve device headroom for allocated frames gve: DQO: reject TSO packets with an out of range MSS gve: fix TX drop when GSO MSS is too small for hw gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO net: flush skb_defer_nodes in dev_cpu_dead() net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails af_packet: fix integer overflow in prb_calc_retire_blk_tmo() tipc: Fix a data race on mon->peer_cnt in mon_timeout() net: phy: intel-xway: workaround 100BASE-TX Link-Up issue net/smc: fix UAF on lgr list traversal in smcr_port_err() net/rds: size a connection's path set by the transport it ends up with nfp: hold IPsec RX state under the XArray lock net: ena: fix MMIO read buffer leak on probe failure net: ena: fix PHC cleanup on probe failure net/sched: act_ct: fix helper UAF due to extensions realloc ...
2026-09-24tcp: fix use-after-free of retransmit_skb_hint in tcp_send_synack()Yilin Zhang1-0/+3
When tcp_send_synack() replaces the cloned SYN skb at the head of the retransmit queue with a copy, it frees the original with tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack. tp->retransmit_skb_hint keeps pointing at the freed skbuff_fclone_cache object. The dangling hint is read in tcp_verify_retransmit_hint() and used as the root of the rbtree walk in tcp_xmit_retransmit_queue(). An unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with an attacker-supplied ICMP fragmentation-needed message, after which a simultaneous open frees the armed SYN skb: BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316) Read of size 4 at addr ffff88800604d928 by task swapper/1/0 Call Trace: tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316) tcp_simple_retransmit (net/ipv4/tcp_input.c:3158) tcp_v4_err (net/ipv4/tcp_ipv4.c:587) Sync the hint to the copy. Fixes: c31b70c9968f ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit") Reported-by: Kimi Security Team <bug-report@moonshot.ai> Tested-by: Weiming Shi <shiweiming@moonshot.ai> Signed-off-by: Yilin Zhang <yilinzhang@moonshot.ai> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24Merge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpfLinus Torvalds1-1/+2
Pull bpf fixes from Alexei Starovoitov: - Fix bpf_skb_change_tail() to drop the checksum offload instead of rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann) - Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that read arbitrary memory and for untrusted read-only memory reads (Daniel Borkmann) - Clear scalar delta on narrowing stack spill (Daniel Borkmann) - Set up the frame pointer for the exception callback in arm64 JIT, and zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash element (Donggeun Yoo) - Various fixes (Emil Tsalapatis): - Fix bounds check underflow for skb-backed dynptrs - Fix rx_queue_mapping context access code generation in bpf_sock - Reject packet pointer arguments to subprogs that may mutate the packet - Reject ALU instructions that see arena and non-arena operands on different code paths - Fix copied_seq double-counting on sockmap self-redirect (Geliang Tang) - Fix divide-by-zero in btf_struct_walk() on a flexible array of zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops (Jiayuan Chen) - Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT and request socks, and sleeping under RCU when destroying a listener with pending children (Jiayuan Chen) - Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh) - Avoid soft lockup in htab lookup[_and_delete] batch operations on large maps (Jose Fernandez) - Various fixes (Kumar Kartikeya Dwivedi): - Verify global subprogs in each sleepability context they are called from - Make post-verification instruction rewrites killable - Preserve packet pointer displacement in regsafe() - Apply CO-RE relocations before subprogram validation, restrict CO-RE poisoning to relocatable instructions, and reject truncated ldimm64 CO-RE relocations in libbpf - Assign lock identity to callback map values - Compare stack frames in regs_exact() - Bound ownership depth through local kptrs and graph roots - Fix u32 overflow in map batch operations when the map size exceeds 4GB (Masoud Aghasi) - Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists in alloc_bulk() (Pu Lehui) - Allow gotox as the terminal instruction of a program or a subprogram (Siddharth Chintamaneni) - Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links in link iterator, and reject dev-bound-only programs on other devices (Weiming Shi) - Reject non-negative stack offsets in stack_slot_obj_get_spi() (Xu Yunxiang) - Check params size before reading reserved fields in bpf_crypto_ctx_create() (Yuqi Xu) - Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi) - Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou) - Use kvfree() in xdp_test_run_teardown() (Zhixing Chen) * tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits) selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element bpf: Fix BSWAP 32 and 16 on MIPS64 bpf: Fix immediate JMP JEQ/JNE on MIPS32 bpf: Reject dev-bound-only programs on other devices bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc selftests/bpf: Test for mixed arena/nonarena code paths bpf: Prevent variable arena/non-arena register contents selftests/bpf: Test rejection of pkt args to mutating subprogs bpf: Reject pkt arguments in mutating subprogs selftests/bpf: Add selftests for rx_queue_mapping context access bpf: Fix bpf_sock context code generation selftests/bpf: Test dynptr slices past end of skb bpf: Fix bounds check for skb-backed dynptrs selftests/bpf: Reject iterator destruction through fp+0 bpf: Reject non-negative offsets in stack_slot_obj_get_spi() bpf: Check params size before reading reserved fields selftests/bpf: Check local object ownership depth bpf: Bound ownership depth through local kptrs and graph roots selftests/bpf: Cover frame changes in bounded loops ...
2026-09-24net: ipconfig: bound DHCP option constructionYuqi Xu1-19/+26
ic_dhcp_init_options() appends the hostname (option 12), vendor-class (option 60) and client-ID (option 61) options into the fixed 312-byte bootp_pkt.exten[] buffer. Only the client-ID branch checked the remaining space; the hostname and vendor-class writes were unbounded. A 64-byte hostname together with the maximum 252-byte dhcpclass= identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available bytes even before the terminating END marker, so the vendor-class memcpy runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is reported as a field-spanning write and, when the kernel is booted with panic_on_warn=1, aborts boot with a panic. Route the optional options through a common helper that makes sure the option, its 2-byte header and the END marker all fit and drops an option that would not. Configurations with short options keep sending exactly the same bytes as before. Fixes: 130c0f47fdf9 ("ipconfig: send host-name in DHCP requests") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Assisted-by: LLM Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com> Reviewed-by: Ren Wei <weir@nebusec.ai> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-23ip_gre: Reject enabling collect metadata through changelinkXuanqiang Luo1-0/+12
ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP or ERSPAN device. Unlike newlink, changelink does not enforce metadata tunnel uniqueness. Converting a non-metadata device can therefore replace the metadata receive entry for another device of the same type in the same netns. Deleting either device then clears the shared entry, breaking metadata receive lookup for the surviving device. If parameter validation fails after collect_md is set, deleting the modified device can also clear an entry it never owned. Reject enabling metadata mode in both changelink callbacks before any encapsulation or tunnel parameters are modified. Allow requests that repeat the metadata attribute on an existing metadata device. Fixes: 2e15ea390e6f ("ip_gre: Add support to collect tunnel metadata.") Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23net: arp: terminate device name before lookupZijie Huang1-0/+1
The ARP ioctl copies a user-provided struct arpreq into a stack object. Its arp_dev field may contain IFNAMSIZ bytes without a NUL terminator. Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where strcmp() can read past the end of the stack object when a matching alternative interface name exists. Terminate the field before the lookup to prevent the out-of-bounds read. Fixes: 36fbf1e52bd3 ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Signed-off-by: Zijie Huang <milkory@outlook.com> Signed-off-by: Ren Wei <weir@nebusec.ai> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23fou: reject omitted FOU_ATTR_IPPROTO on FOU_ENCAP_DIRECTHui Peng1-0/+4
Commit 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with -ERANGE. However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0, sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with fou->protocol == 0. In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb() triggers IP protocol resubmission when fou->protocol > 0, whereas returning 0 tells the UDP tunnel layer that the skb was consumed without freeing it. When fou->protocol == 0, every packet received on the socket returns 0 from fou_udp_recv() and leaks the sk_buff. Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with -EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share parse_nl_config()) unaffected. Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE = FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel, FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to 59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected with -EINVAL (-22). Fixes: 23461551c006 ("fou: Support for foo-over-udp RX path") Fixes: 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") Cc: stable@vger.kernel.org Signed-off-by: Hui Peng <benquike@gmail.com> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22udp: remove a disconnected socket from the 4-tuple hash tableShardul Bankar1-0/+23
A UDP socket bound to a specific address and port keeps its entry in the 4-tuple hash table after it is disconnected: sk binds to 127.0.0.1:21001 sk connects to 127.0.0.2:20001 // filed in the 4-tuple table sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0 __udp_disconnect() takes a socket out of that table only as a side effect of ->rehash() or ->unhash(), and it skips ->rehash() when SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set. commit 6996a2d2d0a6 ("udp: Unhash auto-bound connected sk from 4-tuple hash table when disconnected.") fixed the same end state for a wildcard-bound socket, by a path this one does not take. The entry is counted whether or not anything hits it. hash4_cnt on the hash2 slot stays raised for as long as the socket lives, so udp_has_hash4() keeps sending every packet for that address and port through the 4-tuple lookup first. On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr, so udp_v6_rehash() files the entry under the peer the socket was connected to with a zero dport, and inet6_match() compares that same field: a datagram from the former peer with a zero source port matches, and source port zero is accepted on receive. On IPv4 the peer is cleared, so a match would need a zero source address as well, which the routing layer rejects as martian. The stale sk_v6_daddr is a separate defect, not addressed here; removing the entry closes this path either way. The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if, so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive address is still specific udp_lib_rehash() moves the entry instead of removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a pure function of the address and port, so every socket reaching this state on one address and port collects in one bucket. The bucket cannot be chosen from outside, as udp_ehashfn() is seeded with a per-boot secret. This last one became reachable only with commit 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()"), which moved the hash4 handling out of a branch a disconnected socket does not take; the stale entry itself dates from the commit in Fixes. Take the socket out of the table before __udp_disconnect() runs, while it still matches how it was filed. This also reaches the wildcard case ahead of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable from udp_disconnect(); removing it belongs in net-next. udp_disconnect() and udp_abort() are the only UDP entries into __udp_disconnect(), which is shared with raw, ping and l2tp sockets that are not struct udp_sock: ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one would read past the allocation. Fixes: 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket") Assisted-by: LLM Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22udp: relocate a connected socket in the 4-tuple hash table on re-connectShardul Bankar1-5/+14
A connected UDP socket that connects again to a different peer is not re-filed in the 4-tuple hash table: sk binds to 127.0.0.1:21001 sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1) sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1) packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the // lookup falls back to scoring the // hash2 chain for this address // and port udp_lib_hash4() returns early when the socket is already hashed, assuming ->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect() only while the receive address is unset, which a second connect never is: the first connect assigns it, whether the socket was bound to a specific address or to the wildcard. commit 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()") added that early return and named connect(AF_UNSPEC) as the way around it. That workaround does not help a socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because __udp_disconnect() skips ->rehash() for the first and ->unhash() for the second. Delivery is correct either way. Relocate the socket when the hash it is filed under differs from the one requested, which is what commit 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket") did before the early return became unconditional. It is done here under hslot->lock, which that version did not take, to match udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected, and IPv6 shares the code. With 500 sockets on the port, a re-connected socket measured 522,553 pps without this change and 2,055,078 with it. The UDP side was noted as remaining work in [1]. Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1] Fixes: 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()") Assisted-by: LLM Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-18net: allow IFLA_INET_CONF messages when NLA_F_NESTED unsetQuentin Armitage1-3/+4
Commit fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly") added validation of IFLA_INET_CONF attributes, and in the process changed the call of nla_for_each_nested() to nla_parse_nested(). A side effect of this change is that the IFLA_INET_CONF option is now tested for NLA_F_NESTED being set, and fails if it is not. Prior to the commit there was no check of NLA_F_NESTED. Change nla_parse_nested() to nla_parse(). This restores the previous functionality of not checking NLA_F_NESTED, thereby allowing code that (incorrectly) doesn't set NLA_F_NESTED to continue to work. This issue was identified because keepalived started logging errors when it was configuring macvlans that it created. Fixes: fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly") Signed-off-by: Quentin Armitage <quentin@armitage.org.uk> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17ipv4: fib: fix data-race and stale genid check around nh->nh_saddrLinkui Xiao1-1/+12
fib_select_multipath() compares nexthop_nh->nh_saddr against the flow source address with no lock held, while fib_info_update_nhc_saddr() stores a new value from another CPU as soon as the preferred source address of the egress device changes. Commit 195374d89368 ("ipv4: fib: annotate races around nh->nh_saddr_genid and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE() in fib_result_prefsrc() after syzbot reported BUG: KCSAN: data-race in fib_select_path / fib_select_path but it only covered that reader. fib_select_multipath(), reached from fib_select_path(), is a second lockless reader of nh->nh_saddr and was left bare. Moreover, nh_saddr is only meaningful when nh_saddr_genid matches dev_addr_genid, as established by commit 436c3b66ec98 ("ipv4: Invalidate nexthop cache nh_saddr more correctly."). fib_select_multipath() skips that validation, so it can score a nexthop using a stale source address and skew the ECMP selection. Annotate both reads with READ_ONCE() and refresh the cached source address via fib_info_update_nhc_saddr() when the genid does not match, mirroring fib_result_prefsrc(). Fixes: 32607a332cfe ("ipv4: prefer multipath nexthop that matches source address") Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/20260916125316.988044-1-xiaolinkui@126.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17tcp: exclude old ACKs from tcp fast pathInbal Schussheim1-1/+2
Exclude old ACKs before SND.UNA from the tcp fast path as well as ACKs after SND.NXT. Such ACKs will fall through to the slow path, where tcp_ack() performs the appropriate validation and challenge ACK handling according to RFC5961 and Commit 3d501dd326fb1c7 ("tcp: do not accept ACK of bytes we never sent"). This prevents old ACKs from being accepted or modifying connection state as part of the fast path before appropriate ACK validation is applied. In particular, this prevents payload carried by a segment with an excessively old ACK from advancing RCV.NXT before the ACK is rejected. Fixes: 31770e34e43d ("tcp: Revert "tcp: remove header prediction"") Reported-by: Amit Klein <amit.klein@mail.huji.ac.il> Reported-by: Tamir Shahar <tamir.shahar1@mail.huji.ac.il> Reported-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il> Suggested-by: Eric Dumazet <edumazet@google.com> Cc: stable@vger.kernel.org Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/20260914090408.1435080-2-inbal.lipshtat@mail.huji.ac.il Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-16net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()Daniel Zahka1-0/+4
PSP conflicts with TLS ULP in its usage of both skb->decrypted and sk->sk_validate_xmit_skb(). Make PSP mutually exclusive with TLS ULP, the only other user of either of these. As other users of skb->decrypted come along, they can be added to sk_has_decrypt_user(). It would make sense to also assert that sk->sk_validate_xmit_skb() is also NULL in both of these setup paths for similar future proofing, but the PSP listener/sk_clone() path is still broken and it could be seen as a regression to not allow rx assoc to run on a child of a listener socket with PSP tx assoc state. Include all TCP ULPs in the sk_has_decrypt_user() check, even though TLS is the only one that conflicts with PSP via the decrypted bit. This is intentional because PSP was not designed to be used with ULPs. It is best to close off surface area that may make bugs reachable, until someone wishes to design and test an actual user of PSP with ULPs. Fixes: 6b46ca260e22 ("net: psp: add socket security association code") Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260915-psp-ktls-fix-v2-1-0eedc3b148ec@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16Merge tag 'ipsec-2026-09-16' of ↵Jakub Kicinski1-0/+6
git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec Steffen Klassert says: ==================== pull request (net): ipsec 2026-09-16 1) xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk() Add the up-front nr_frags guard iptfs_skb_add_frags() already has, so an out-of-range offset can't walk past the on-stack frags[] array. 2) xfrm: serialize state GC with device state flush Serialize xfrm_state destruction against the deferred-device pass with a dedicated mutex, since the device GC list doesn't hold a state reference and the two paths could free the same state. 3) xfrm: add missing RCU read lock in xfrm_send_migrate_state() Hold the RCU read lock around xfrm_nlmsg_multicast() so the rcu_dereference() of net->xfrm.nlsk doesn't warn. 4) xfrm: iptfs: fix runt reassembly panic from short inner tot_len Require the runt length to cover at least the minimum IP header, so a tot_len in [6, 19] (IPv4) can't write past the declared length and trip skb_over_panic(). 5) ipv6: xfrm: use full sockets in local error paths Use skb_to_full_sk() in xfrm6_local_rxpmtu() and xfrm6_local_error() and bail out without a full socket, so a TCP_NEW_SYN_RECV request_sock isn't miscast as a full inet/IPv6 socket. 6) xfrm: fix compat ALLOCSPI request use-after-free Drop the redundant alloc_compat() in xfrm_alloc_userspi() so the compat translator no longer reads past the payload and publishes a child a multicast clone can still see after xfrm_user_rcv_msg() frees. 7) xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject() Force the dst before queuing, hold dev across the workqueue deferral, and take rcu_read_lock() around the finish() loop, so transport-mode reinjection doesn't deref non-refcounted dst/dev under workqueue. 8) xfrm: use hlist_del_init_rcu for state_cache and state_cache_input Switch to hlist_del_init_rcu() so a second __xfrm_state_delete() is a no-op instead of writing through LIST_POISON2, closing the UAFs. 9) esp: downgrade zerocopy managed frags before mutating skb frags Call skb_zcopy_downgrade_managed() before ESP rewrites the skb frag array, so per-frag unrefs in esp_ssg_unref() and skb_release_data() stay balanced for ubuf-owned managed frags. 10) xfrm: hold net_device reference under RCU in bundle creation Read dst->dev via dst_dev_rcu() and keep RCU active through xfrm_fill_dst(), so a concurrent RTM_DELLINK can't free dev under bundle creation. 11) xfrm: save input state data before secpath resets Save the state protocol on the stack while it's still valid and use the saved address family for transport_finish(), so post-reset dereferences (VTI, XFRM if, MAX_DEPTH error) can't UAF the state. 12) net: xfrm: reject unrepresentable espintcp transport headers Use the careful transport-header helper and drop the skb through the XFRM error path when the offset can't be represented, instead of silently truncating it. * tag 'ipsec-2026-09-16' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec: net: xfrm: reject unrepresentable espintcp transport headers xfrm: save input state data before secpath resets xfrm: hold net_device reference under RCU in bundle creation esp: downgrade zerocopy managed frags before mutating skb frags xfrm: use hlist_del_init_rcu for state_cache and state_cache_input xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject() xfrm: fix compat ALLOCSPI request use-after-free ipv6: xfrm: use full sockets in local error paths xfrm: iptfs: fix runt reassembly panic from short inner tot_len xfrm: add missing RCU read lock in xfrm_send_migrate_state() xfrm: serialize state GC with device state flush xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk() ==================== Link: https://patch.msgid.link/20260916101938.118628-1-steffen.klassert@secunet.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15net: ip_tunnel: initialize `options_len` before referencing optionsGris Ge1-5/+11
The following command triggers a kernel panic: ip link add d0 type dummy; ip link set d0 up ip route add 10.30.0.0/16 \ encap ip id 300 geneve_opts 4660:66:11223344 dev d0 memcpy: detected buffer overflow: 4 byte write of buffer size 0 kernel BUG at lib/string_helpers.c:1044! ... ip_tun_parse_opts.part.0.cold+0x10/0x10 ip_tun_build_state+0x116/0x2a0 On kernels built with GCC 15+ and `CONFIG_FORTIFY_SOURCE`, the fortified `memcpy()` got 0 sized destination with request of 4 bytes length: static int ip_tun_parse_opts_geneve(...) { ... attr = tb[LWTUNNEL_IP_OPT_GENEVE_DATA]; data_len = nla_len(attr); /* == 4 */ struct geneve_opt *opt = ip_tunnel_info_opts(info) + opts_len; memcpy(opt->opt_data, nla_data(attr), data_len); /* ^^^^^^^^^^^^^ 0 since options_len is assigned afterwards */ Fixed by initializing the counter before the options are referenced. Matching what `tunnel_key_opts_set()` already does. Fixes: bb5e62f2d547 ("net: Add options as a flexible array to struct ip_tunnel_info") Cc: stable@vger.kernel.org Signed-off-by: Gris Ge <cnfourt@gmail.com> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Reviewed-by: Gustavo A. R. Silva <gustavoars@kernel.org> Link: https://patch.msgid.link/20260913090851.468216-1-cnfourt@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15tcp: do not let tcp_rmem be set below 4096Eric Dumazet1-1/+3
We can hit a division by zero crash in tcp_rcvbuf_grow() and tcp_rcv_space_adjust(): divide error: 0000 [#1] PREEMPT SMP RIP: 0010:tcp_rcvbuf_grow+0x187/0x450 net/ipv4/tcp_input.c:939 ... grow = div_u64(((u64)rcvwin << 1) * (newval - oldval), oldval); The division uses oldval = tp->rcvq_space.space as divisor. When tp->rcvq_space.space is zero, this leads to a divide-by-zero exception. tp->rcvq_space.space is initialized in tcp_init_buffer_space(): tp->rcvq_space.space = min3(tp->rcv_ssthresh, tp->rcv_wnd, (u32)TCP_INIT_CWND * tp->advmss); If tcp_rmem[1] is configured to very small values (such as 1), sk->sk_rcvbuf is initialized to 1. Then tcp_full_space(sk), which computes (sk->sk_rcvbuf * scaling_ratio) >> 8, truncates to 0. This sets tp->window_clamp = 0, tp->rcv_ssthresh = 0, and tp->rcvq_space.space = 0. Later, when data arrives and DRS is invoked, tcp_rcvbuf_grow() divides by oldval == 0. Back in 2015, commit b1cb59cf2efe ("net: sysctl_net_core: check SNDBUF and RCVBUF for min length") ensured that net.core.rmem_default and net.core.rmem_max cannot be set below SOCK_MIN_RCVBUF. Similarly, SO_RCVBUF setsockopt enforces max_t(int, val * 2, SOCK_MIN_RCVBUF). However, net.ipv4.tcp_rmem still had .extra1 = SYSCTL_ONE, allowing arbitrarily small values. Because SOCK_MIN_RCVBUF depends on sizeof(struct sk_buff) and cacheline alignment, its value varies across architectures and configuration options. Using a fixed constant of 4096 ensures a predictable, architecture- independent lower bound that is safely above SOCK_MIN_RCVBUF everywhere and matches the documented 4K default. Fix this by setting tcp_rmem.extra1 to 4096 and updating the documentation. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Signed-off-by: Eric Dumazet <edumazet@google.com> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/20260912144848.3448026-1-edumazet@google.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15ipv4: icmp: reject RTN_UNREACHABLE input routes in icmp_route_lookupDong Chenchen1-7/+10
When the forward output route cannot be used in icmp_route_lookup(), it enters the "reverse path" and calls ip_route_input() on fl4_dec.daddr, the original packet's source address. ip_route_input() only returns an error for truly invalid packets. For unreachable addresses it will succeed and return an input route whose dst.output is set to ip_rt_bug(). The existing check only rejects RTN_LOCAL routes, so the RTN_UNREACHABLE route types can still be returned and later used for output, syzkaller triggering a WARN_ON_ONCE() in ip_rt_bug() as bellow: ------------[ cut here ]------------ WARNING: net/ipv4/route.c:1273 at ip_rt_bug+0x14/0x20 RIP: 0010:ip_rt_bug+0x14/0x20 Call Trace: ip_push_pending_frames+0xfa/0x100 __icmp_send+0x905/0xf10 ip_options_compile+0xc0/0xd0 ip_rcv_finish_core+0x321/0xae0 ip_rcv+0x1de/0x260 __netif_receive_skb_one_core+0x11a/0x130 netif_receive_skb+0x7b/0x260 tun_get_user+0x11bf/0x1c10 ------------[ cut here ]------------ Reject input route that is RTN_UNREACHABLE to fix it. The net warning is only printed for RTN_LOCAL, as RTN_UNREACHABLE is not the result of a race condition. Fixes: 8b7817f3a959 ("[IPSEC]: Add ICMP host relookup support") Suggested-by: Ido Schimmel <idosch@nvidia.com> Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Dong Chenchen <dongchenchen2@huawei.com> Link: https://patch.msgid.link/20260910140042.1880242-1-dongchenchen2@huawei.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-10tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF contextJiayuan Chen1-1/+2
bpf_sock_destroy() runs from the tcp iterator, under rcu_read_lock(). If the sock is a listener that still has children in its accept queue, tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched() there trips the debug check: BUG: sleeping function called from invalid context at net/ipv4/inet_connection_sock.c:1523 in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name: test_progs preempt_count: 0, expected: 0 RCU nest depth: 1, expected: 0 locks held by test_progs/628: 3, last CPU#3: #0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at: bpf_seq_read+0x56/0x1210 #1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at: bpf_iter_tcp_seq_show+0x32b/0x4b0 #2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at: bpf_iter_run_prog+0x46b/0xde0 CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G W 7.2.0+ #65 PREEMPT Tainted: [W]=WARN Call Trace: <TASK> dump_stack_lvl+0xc1/0xf0 dump_stack+0x10/0x20 __might_resched+0x3d2/0x610 inet_csk_listen_stop+0x7b/0xbf0 tcp_abort+0x23b/0x3b0 bpf_sock_destroy+0xfc/0x140 bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a bpf_iter_run_prog+0x538/0xde0 bpf_iter_tcp_seq_show+0x26b/0x4b0 bpf_seq_read+0x424/0x1210 vfs_read+0x197/0xe40 ksys_read+0x119/0x240 __x64_sys_read+0x72/0xc0 x64_sys_call+0x647/0x27e0 do_syscall_64+0xe5/0x610 entry_SYSCALL_64_after_hwframe+0x76/0x7e RIP: 0033:0x7fad39b28aca RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000 RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014 RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003 R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000 </TASK> The commit that added the kfunc already guards lock_sock() in tcp_abort() and udp_abort() with has_current_bpf_ctx(), but missed the listener path. Do the same for the cond_resched(). The loop runs inside the iterator's rcu_read_lock(), it must not reschedule or report a quiescent state there. Fixes: 4ddbcb886268 ("bpf: A