aboutsummaryrefslogtreecommitdiff
path: root/include/trace
AgeCommit message (Collapse)AuthorFilesLines
9 daysMerge tag 'landlock-7.3-rc5' of ↵Linus Torvalds1-121/+180
git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux Pull Landlock fixes from Mickaël Salaün: "This mainly fixes the Landlock tracepoint support merged this cycle so that denial and rule events report the intended policy context, whether through tracefs or BTF-visible callbacks. The size of this all is mainly from propagating the corrected contract through event definitions and producers, adding new tests for the reported context, and updating the documentation. Also improve annotation and fix a GCC 16 build warning" * tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux: landlock: Widen ruleset versions to 64 bits landlock: Add counted_by in landlock_domain landlock: Fix tracepoint contract documentation selftests/landlock: Test network denial context selftests/landlock: Test filesystem denial blockers landlock: Report the effective signal number landlock: Report the actual ptrace tracer landlock: Fix network denial trace context landlock: Fix rule tracepoint context landlock: Fix filesystem denial blocker reporting landlock: Fix tracepoint fixed-width type names landlock: Work around gcc-16 -Wuninitialized warning
11 dayslandlock: Widen ruleset versions to 64 bitsMickaël Salaün1-10/+10
Tracepoint consumers use a ruleset ID and version to identify the successful landlock_add_rule(2) call prefix used to create a domain. LANDLOCK_MAX_NUM_RULES bounds distinct stored rules, not successful calls: re-adding already-present rights for an object or port succeeds without increasing num_rules. Because every successful call increments the version, these calls can wrap the 32-bit counter and give different prefixes the same trace identity. Widen the counter and its trace fields to 64 bits so the counter cannot wrap in practice, while preserving the successful-call semantics. Saturating would alias all subsequent histories, while rejecting a call at the limit would change otherwise valid syscall behavior solely for trace metadata. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: 63747c94774d ("landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints") Reviewed-by: Günther Noack <gnoack@google.com> Link: https://patch.msgid.link/20260922132615.1025945-1-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
13 dayslandlock: Fix tracepoint contract documentationMickaël Salaün1-18/+18
The tracepoint documentation claims that denial and lifecycle events expose every input needed to reproduce a verdict. Instead document how denial, ruleset, and domain events identify the denying policy, checked operation and object, and reason for denial. Direct consumers to generic tracepoints for additional operational context. State the reconstruction limits: IDs are boot-local, rule checks have no request ID, and exported records may be lost or cross-CPU reordered. Also replace the incorrect BPF_RAW_TRACEPOINT guidance with libbpf SEC("tp_btf/...") attachment and refer consumers to the event prototypes for callback argument layouts. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Link: https://patch.msgid.link/20260918185036.608651-10-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
13 dayslandlock: Report the effective signal numberMickaël Salaün1-3/+5
The signal-scope denial callback identifies its target but not the effective signal. This loses permission-probe signal zero and makes the file-owner hook's zero sentinel ambiguous. Append an int signal argument to the typed-BPF callback. Preserve sig, including zero, in hook_task_kill(). In hook_file_send_sigiotask(), translate signum zero to SIGIO at the producer, where its meaning is known. Carry the effective signal and target domain ID in a private, stack-backed context consumed synchronously. This requires no allocation or task reference in the interrupt-capable file-owner path. Gate this context and the remaining scope-only domain IDs with CONFIG_TRACEPOINTS. Keep the tracefs record and audit output unchanged. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: bb91730f16c0 ("landlock: Add tracepoints for ptrace and scope denials") Link: https://patch.msgid.link/20260918185036.608651-7-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
13 dayslandlock: Report the actual ptrace tracerMickaël Salaün1-3/+8
The ptrace denial callback identifies only the tracee. Current is the tracer during hook_ptrace_access_check(), but it is the tracee during PTRACE_TRACEME, where the parent is the actual tracer. A consumer therefore cannot infer both parties from the existing arguments. Append the actual tracer task to the typed-BPF callback: current for hook_ptrace_access_check() and parent for hook_ptrace_traceme(). Carry it with the tracee domain ID in a private ptrace context. Both hooks keep the selected tasks alive through synchronous dispatch, so no extra task reference is needed. Keep the tracefs record unchanged. The new context is available only to typed BPF, while same_exec continues to describe the tracer that owns the denying policy. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: bb91730f16c0 ("landlock: Add tracepoints for ptrace and scope denials") Link: https://patch.msgid.link/20260918185036.608651-6-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
13 dayslandlock: Fix network denial trace contextMickaël Salaün1-27/+46
Network denial events report source and destination ports reconstructed from audit data. Their zero values are ambiguous, and neither identifies the complete endpoint that Landlock checked. Carry the checked sockaddr and its signed length in a private trace-only context. For an enabled event, validate the length and copy only the initialized prefix into zeroed local storage. This prevents a typed BPF program from reading uninitialized bytes while exposing the socket family, socket, address, and length. Replace the source and destination trace-record fields with one signed port derived from the checked address. A value of -1 means that no port was checked, zero is a valid port, and positive values use host endianness. Bind blockers select the bind address; connect and send blockers select the destination. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: 01ce260f5ccf ("landlock: Add landlock_deny_access_fs and landlock_deny_access_net") Link: https://patch.msgid.link/20260918185036.608651-5-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
13 dayslandlock: Fix rule tracepoint contextMickaël Salaün1-27/+35
Name each event after the identity it reports. Add-rule events describe UAPI rule insertion, so rename them after LANDLOCK_RULE_PATH_BENEATH and LANDLOCK_RULE_NET_PORT. Check-rule events describe matches in internal rule trees, so rename them after LANDLOCK_KEY_INODE and LANDLOCK_KEY_NET_PORT. This remains accurate if multiple UAPI rule types share one lookup and stored rule. Keep denial event names based on filesystem and network families because they describe final access decisions. Use u64 for growable access masks passed by value to add-rule and check-rule typed BTF callbacks. CO-RE can relocate pointer-reached fields, but it cannot widen a scalar callback slot declared by a BPF program. Keep native access_mask_t for internal state and trace records. For add-rule callbacks, report the normalized per-call contribution passed to landlock_insert_rule() and expose the complete validated flags value. Put the ruleset and flags first as a common invocation prefix. This distinguishes duplicate and effective-zero additions without recovering arguments from saved syscall registers. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: 63747c94774d ("landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints") Fixes: 3f1f106e4c14 ("landlock: Add tracepoints for rule checking") Link: https://patch.msgid.link/20260918185036.608651-4-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
13 dayslandlock: Fix filesystem denial blocker reportingMickaël Salaün1-13/+38
Filesystem topology denials are rendered with an empty blockers value because their blocker is identified by the request type instead of an access mask. Introduce the private struct landlock_blockers to carry the request type and final missing access mask to filesystem and network denial tracepoints. Copy both members into named trace-record fields, then use the type to print change_topology for topology denials while preserving symbolic access masks for ordinary denials. The request type lets typed BPF consumers distinguish topology denials from access denials. Keeping the native access mask in a pointer-reached field also lets CO-RE adjust existing programs' load width if access_mask_t grows. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Fixes: 01ce260f5ccf ("landlock: Add landlock_deny_access_fs and landlock_deny_access_net") Link: https://patch.msgid.link/20260918185036.608651-3-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
13 dayslandlock: Fix tracepoint fixed-width type namesMickaël Salaün1-34/+34
The new Landlock tracepoints use UAPI-prefixed __u32 and __u64 names for callback arguments and record fields, including internal IDs that are not Landlock UAPI values. Typed BPF consumers see callback typedef names through BTF. Use the kernel u32 and u64 aliases before release so the tracepoint contract does not present internal values as Landlock UAPI types. This changes BTF-visible typedef spelling but not integer widths, calling conventions, tracefs formats, or record layouts. Cc: Günther Noack <gnoack@google.com> Cc: Steven Rostedt <rostedt@goodmis.org> Link: https://patch.msgid.link/20260918185036.608651-2-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-17Merge tag 'dma-mapping-7.3-2026-09-17' of ↵Linus Torvalds1-1/+1
git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux Pull dma-mapping fixes from Marek Szyprowski: "A few fixes for the DMA-mapping code: - resolved regression in accessing encrypted memory by IOMMU-backed devices (Aneesh Kumar K.V) - improved failure handling and removed rare bug in swiotlb/highmem (Donggeun Yoo)" * tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux: x86/mm: Don't force unencrypted DMA for IOMMU-backed devices dma-mapping: don't trace the DMA address when the allocation fails swiotlb: use the adjusted address for the highmem page lookup dma-coherent: report a failed reserved memory assignment
2026-09-13Merge tag 'trace-v7.3-rc2' of ↵Linus Torvalds1-1/+1
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace Pull tracing fixes from Steven Rostedt: - Don't destroy user event fields when removal fails User event fields are destroyed before the event is removed from visibility. But that can fail leaving the still visible event with no fields. Move the destroying of the fields to after the event is successfully removed from visibility. - Initialize function graph state is fork before calling copy_exec_state() For non-CLONE_VM forks, copy_exec_state() allocates a new task_exec_state. If that allocation fails, ftrace_graph_exit_task() will free the tasks ret_stack pointer. Since that pointer is still using the parent's ret_stack, it mistakenly frees the parent's pointer too. Call ftrace_graph_init() on the task first which will NULL out the new tasks's ret_stack and if the copy fails, it will not free anything. - Remove FGRAPH_MAX_INDEX The macro FGRAPH_MAX_INDEX was added but never used. Remove it. - Save ent_size in function graph printing of nested functions The function graph tracer needs to look at the next event to see if the next event is the return of the current function entry. If it is, it prints a single line: ktime_get(); Otherwise it prints it like a nested function: tick_nohz_irq_exit() { ktime_get(); kcpustat_irq_exit(); } In order to look at the next event, it must save the current event so that it has the information to print from it. It saves the event in the iterator descriptor called "ent". What it doesn't save is the ent_size of the event which is now used to know if the function graph arguments are to be printed. The peek doesn't save the size so the size used happens to be that of the size of the last event that was seen. Save the entry event size in the iterator descriptor so that the correct size is used. - Fix several errors with freeing data in the histogram code The histogram code had a lot of leaked or or incorrect accounting when failures happen. Correct them. - Fix histogram regression of .percent and .graph modifiers Up until 6.3 histogram values could have "percent" or "graph" modifiers that changed how they were printed. But a change that added restricting histograms values from being strings, stack traces and other modifiers inadvertently prevented them from using the percent and graph modifiers, which were legal use cases for values. Put back the percent and graph modifiers. - Fix various typos in the comments - Set the trace_clock before initializing a histogram with clock argument The histogram API allows the user to specific which trace clock to use via a "clock=" string. The histogram is set up first before the clock is checked. If the passed in clock is not valid, it exits without fully fixing up the histogram leaving it on the list and a use-after-free can trigger. Update the clock argument first and if it fails then exit gracefully before the histogram trigger is placed on any lists. - Restore :mod: trailer after parsing in ftrace_set_clr_event The function ftrace_set_clr_event() modifies the parse string and needs to put it back to what was passed in. It searches for ":mod:" via a strsep() but fails to put back the first ':' in the string. Add back the ':' in the passed in string. - Take trace_array reference when opening a tracer options file The options files are dynamically created and some tracers add their own options. When a tracer adds their own list of options, the trace_array holding them has an array to hold the list of options for each tracer. This array increases in size via a krealloc(), and the new entry gets a newly allocated array to hold the options of the new tracer being added. The element in each entry of the tracer's option array holds a pointer back to the trace_array, a pointer to the tracer it is associated to, a pointer to the flags of the option. The issue is that these arrays are freed when the trace_array is freed when its instance it represents is removed from the instances directory. There's a race that an open of one of these options files can happen when the instance is being removed. Add a new helper function to be called by the open function of the options file to iterate all existing trace_arrays under a lock and find the one that has the given option element in one of it's tracer arrays. If found, then update the associated trace_array's reference counter to keep it from being freed. If not found, have the open call return -ENODEV. - Disable interrupts when acquiring the lock in rb_wake_up_waiters() The function rb_wake_up_waiters() assumes it will be called in interrupt context and does not disable irqs when taking cpu_buffer->reader_lock, which can be called in hard interrupt context. The issue is in PREEMPT_RT, this function is called in thread context leaving this lock open to a deadlock. Take the lock with interrupts disabled. - Use rcu_assign_pointer() for tmp_ops filter hash The tmp_ops used in update_ftrace_direct_mod() assigns its filter_hash field directly, but that field is annotated as __rcu and sparse complains. Assign it with rcu_assign_pointer() - Fix use-after-free in enable_trigger_private_data_free() The trace_event_call is accessed through the event_trigger_data's trace_event_file pointer to put the trace_event_call on freeing. The issue is that the trace_event_file data may have been freed already causing a use-after-free. Add a field to the event_trigger_data that points directly to the trace_event_call so that it can decrement its reference directly without needing to go through the trace_event_file. - Fix accounting of buffer data remote headers trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount the number of pages is needed for the asked for size as it doesn't take into account the meta data on each page. Add a helper function to do the calculation properly and use that in these functions. - Catch nr_page_va overflow in ring_buffer_desc sizing The number of pages per remote ring buffer is capped by ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to overflow that field would silently allocate a descriptor smaller than what was asked for. - Do not resize the subbuf order if any per_cpu buffer is disabled The mmapping of ring buffers disables resizing the subbuffers, but it is done per-cpu whereas the subbuf size change is done for all the per_cpu buffers under the buffer->mutex. It could change the size of some while the mapping is happening on others. Have the resize of the subbuf order check all the per_cpu buffers under the lock to see if any of them is disabled before starting and causing an inconsistency between buffers that are being mapped. * tag 'trace-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: (25 commits) ring-buffer: Check resize_disabled before publishing the new subbuf order tracing/remotes: Catch nr_page_va overflow in ring_buffer_desc sizing tracing/remotes: Account for ring buffer page header in size calculation tracing: Don't dereference trace_event_file in deferred trigger free ftrace: Use rcu_assign_pointer() for tmp_ops filter hash ring-buffer: Acquire the lock with irqsave in rb_wake_up_waiters() tracing: Take trace_array reference when opening a tracer options file tracing: Fix ring_buffer_read_page_size() kernel-doc tracing: Restore :mod: trailer after parsing in ftrace_set_clr_event() tracing: Fix memory corruption from a "STACKTRACE" histogram key tracing: Fix memory corruption from the histogram stacktrace modifier tracing: Undo the registration when enabling the histogram trigger fails tracing: Take the reference before publishing the named histogram trigger tracing: Set the trace clock before registering the histogram trigger tracing: Fix typo "preceeded" in comment tracing: Fix typo "availabe" in comment tracing: Let histogram values keep the percent and graph modifiers tracing: Keep the entry count when the histogram stats allocation fails tracing: Free histogram the field rejected for a bad modifier tracing: Free histogram the var ref when its initialization fails ...
2026-09-11tracing: Fix typo "preceeded" in commentHemanth Selam1-1/+1
Correct "preceeded" to "Preceded", reported by scripts/checkpatch.pl using the misspelling list in scripts/spelling.txt. Only touches comments, no code changes. Link: https://patch.msgid.link/20260907065607.36615-1-hemanth.selam@gmail.com Assisted-by: Cursor:claude-opus-5 Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com> Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2026-09-09Merge tag 'landlock-7.3-rc3' of ↵Linus Torvalds1-17/+53
git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux Pull Landlock fixes from Mickaël Salaün: "This fixes a use-after-free and a lockdep assert NULL dereferencing, and properly truncates too-long strings printed by a Landlock tracepoint. Most of the changes are brought by new tests" * tag 'landlock-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux: landlock: Test trace path output boundaries landlock: Bound escaped trace path output landlock: Clean up ruleset validation checks selftests/landlock: Test abstract socket trace name limits landlock: Fix use-after-free of the source's parent directory
2026-09-09Merge tag 'vfs-7.3-rc3.fixes' of ↵Linus Torvalds2-4/+45
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs Pull vfs fixes from Christian Brauner: - netfs: - Fix an uninitialized return value in netfs_unbuffered_write() when preparing the first subrequest fails - For partial unbuffered/DIO writes return the amount transferred rather than an error - Update i_size with the amount actually written when a partial transfer ends in an error - Fix a subrequest reference leak when the io_iter ends up empty - Handle netfs_alloc_subrequest() failure during unbuffered writes - Load all readahead folios into the rolling buffer upfront and drop the readahead references once the first subrequest is dispatched - Mark folios for copy-to-cache while issuing subrequests - Fix read progress reporting - afs: - Add the missing kunmap in the error path of afs_dir_search_bucket() - Fix a double kunmap in afs_edit_dir_remove() - Don't free an existing server's endpoint state when cleaning up a candidate server in afs_lookup_server() - Unbind peers removed from a server's address list - ufs: - Load the cylinder group metadata before creating the root dentry - Validate the cylinder group index and rotor positions before caching them - Treat an unreadable directory block as not empty - exec: - Close the close-on-exec files before taking exec_update_lock Closing a file can block on the filesystem, so a hung filesystem blocked everything that takes exec_update_lock and a FUSE server inspecting the calling process could deadlock - Drop the bprm loader before closing bprm->file in free_bprm() - exit: Hold a reference to thread_pid across proc_flush_pid() - reboot: Fix a use-after-free on cad_pid - nsfs: Keep the namespace tree fields out of the rcu_head used by kfree_rcu() - nstree: Check listing permission before taking a namespace reference in listns() - super: Return 0 when a nested thaw drops its hold while other freezers remain - ext4: Don't set I_METADATA_WRITEBACK during fastcommit replay - adfs: Free s_fs_info in ->kill_sb() - autofs: Free the inode info allocated in autofs_fill_super() when the root inode allocation fails - ovl: Return EINVAL instead of EIO on a user namespace mismatch now that it's a plain refusal and not an internal error - cachefiles: Don't cast the variable-length coherency data to a __be64 in the coherency tracepoint * tag 'vfs-7.3-rc3.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (28 commits) nstree: check listing permission before taking a namespace reference exec: do_close_on_exec() before taking exec_update_lock exit: hold a reference to thread_pid across proc_flush_pid fs: autofs: fix memory leak in autofs_fill_super() exec: Drop bprm loader before closing bprm->file afs: Clear stale peer app data after address list changes afs: Fix incorrect free in candidate cleanup in afs_lookup_server() afs: Fix double-unmap of directory block afs: Fix missing kunmap in afs_dir_search_bucket() ovl: return EINVAL instead of EIO in case of mismatched user_ns reboot: fix cad_pid use-after-free race cachefiles: Fix potential UAF/KASAN warning netfs: Fix read progress reporting netfs: Mark folios with COPY_TO_CACHE whilst issuing subreqs netfs: Fix readahead synchronisation issues by loading all folios upfront netfs: break unbuffered write when netfs_alloc_subrequest() fails netfs: Fix subreq ref leak netfs: Fix i_size update for partial transfer netfs: Fix error vs transferred passed to ->ki_complete() netfs: Fix unbuffered/DIO write partial transfer error return ...
2026-09-09dma-mapping: don't trace the DMA address when the allocation failsDonggeun Yoo1-1/+1
dma_alloc_attrs() passes *dma_handle to trace_dma_alloc() without checking whether the allocation succeeded. No backend writes it on failure: dma_direct_alloc(), iommu_dma_alloc() and the dma_map_ops instances assign it only on the path that returns a buffer. Callers usually pass an uninitialized automatic variable, so a failed allocation records whatever the stack held, next to the virt_addr=(null) that marks the record as an error: dma_alloc: dmatrace dir=BIDIRECTIONAL dma_addr=deadbeefdeadbeef size=1099511627776 virt_addr=0000000000000000 The device coherent pool path reaches the same call: a non-zero return from dma_alloc_from_dev_coherent() means the request was handled, not that it succeeded, so cpu_addr is NULL and dma_handle is untouched once the pool runs out. For an allocation event a NULL virt_addr already means the request failed, so the address field carries nothing. Report 0 for it in the event class rather than at each call site, which covers dma_alloc_pages() and dma_alloc_sgt_err() as well. Fixes: 038eb433dc14 ("dma-mapping: add tracing for dma-mapping API calls") Fixes: 68b6dbf1f441 ("dma-mapping: trace more error paths") Suggested-by: Marek Szyprowski <m.szyprowski@samsung.com> Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com> Link: https://lore.kernel.org/r/20260907120124.603373-1-donggeunyoo.kernel@gmail.com Reviewed-by: Sean Anderson <sean.anderson@linux.dev> Signed-off-by: Marek Szyprowski <m.szyprowski@samsung.com>
2026-09-08landlock: Bound escaped trace path outputMickaël Salaün1-17/+53
Filesystem paths may expand fourfold when trace text escapes spaces and other untrusted bytes. A sufficiently long representation can exhaust the shared scratch sequence. A sibling __print_flags() helper may then return an unterminated one-past pointer because TP_printk() argument ordering is unspecified. Use a fixed budget rather than the scratch space available at call time, so output does not vary with sibling evaluation order. Limit an untrusted string to three quarters of the trace sequence, leaving the rest for sibling helpers and final event metadata. Compute and commit complete escaped output transactionally so an exact fill cannot consume the terminating NUL or poison the scratch sequence. For strings that exceed the limit, retain the largest prefix ending at a complete escape unit, then append a raw UTF-8 ellipsis. Keep the helper's existing octal fallback so complete values remain unchanged. Hex fallback would consume the same four bytes per escaped byte without increasing the prefix or strengthening the marker. ESCAPE_NAP renders every non-ASCII input byte in octal, so legitimate data cannot reproduce the marker without being escaped. Cc: Günther Noack <gnoack@google.com> Link: https://patch.msgid.link/20260907154401.124362-1-mic@digikod.net Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-03Merge tag 'net-7.3-rc2' of ↵Linus Torvalds1-5/+8
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Paolo Abeni: "Including fixes from bluetooth. Previous releases - regressions: - page_pool: keep frag_offset aligned for odd-sized requests - sched: fix u32 duplicate handle when node ID pool is exhausted - udp: create exceptions before socket matching - igmp: convert struct ip_sf_list to RCU - ip6_gre: check tunnel info before xmit in ip6gre_tunnel_xmit - rds: acquire the fastpath locks in rds_conn_shutdown() - tipc: - protect node reset trace dump with node lock - fix NULL deref in tipc_named_node_up() on empty publication list - bluetooth: - L2CAP: fix out-of-bounds write in l2cap_ecred_connect - hci_core: fix race condition during device registration - eth: - mlx5e: prevent stale XSK buffer release on refill retries - bridge: don't truncate the port group walk on teardown Previous releases - always broken: - gro: fix nesting of TCP GSO SKBs in skb_gro_receive_list() - sched: fix skb sizing and action leak on reoffload delete - tcp: fix use-after-free in do_tcp_getsockopt() - af_packet: don't cast tpacket_hdr.tp_len to int in tpacket_parse_header() - sctp: fix soft lockup from unpadded ASCONF-ACK parameter iteration - iptunnel: fix stale transport header during tunnel decapsulation - eth: - vxlan: fix use-after-free in vxlan_mdb_remote_src_del() - bonding: fix uninitialized transport header access in alb_determine_nd()" * tag 'net-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (83 commits) net: gro: Fix nesting of TCP GSO SKBs in skb_gro_receive_list() net: stmmac: reconfigure RX packet parser table in stmmac_hw_setup() after reset net: airoha: enable RX_DONE interrupt for RX queue 31 net/rds: don't let rds_conn_shutdown() consume a concurrent drop net/rds: acquire the fastpath locks in rds_conn_shutdown() net/rds: acquire RDS_IN_XMIT in rds_tcp_reset_callbacks() net/rds: tcp: don't force RDS_CONN_RESETTING over a concurrent shutdown net/rds: clear cp_flags bits individually in rds_conn_path_reset() net/rds: use clear_bit_unlock() in release_refill() net/rds: use wq_has_sleeper() in release_in_xmit() net: usb: qmi_wwan: add Compal EXM-G1x support net: macb: exclude software FCS from TX byte statistics net: Remove conflicting altnames for dying netns in __dev_change_net_namespace(). net: bridge: mcast: don't truncate the port group walk on teardown bonding: do not clear curr_active_slave prematurely when releasing all slaves net: qrtr: Send HELLO message on endpoint register octeontx2-af: Fix limiting SRIOV VF count logic bonding: alb: fix uninitialized transport header access in alb_determine_nd() s390/ctcm: Prevent XID null dereference net: psp: do not inherit the Rx association on clone ...
2026-08-31cachefiles: Fix potential UAF/KASAN warningDavid Howells1-2/+17
Currently, trace_cachefiles_coherency() is being passed a pointer to a __be64 lain over the coherency data in struct cachefiles_xattr so that it can display the first 8 bytes. However, the data is of variable length and could even be 0 bytes. This could lead to a UAF or KASAN warning. Fix this by making sure the buffer has room for at least 8 bytes and that those 8 bytes are pre-cleared. Further, those bytes are not 8-byte aligned, so fix the tracepoint to extract the data as four 2-byte words (they are 2-byte aligned) and reassemble the __be64. The compiler will convert this into a single 8-byte load where the CPU supports it. Fixes: 229105e5cfd9 ("cachefiles: Add auxiliary data trace") Link: https://sashiko.dev/#/patchset/20260810144746.574036-1-dhowells%40redhat.com Signed-off-by: David Howells <dhowells@redhat.com> Link: https://patch.msgid.link/20260827134304.2075713-11-dhowells@redhat.com Acked-by: Paulo Alcantara <pc@manguebit.org> cc: Paulo Alcantara <pc@manguebit.org> cc: netfs@lists.linux.dev cc: linux-fsdevel@vger.kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-31netfs: Fix read progress reportingDavid Howells1-0/+21
For really big read RPC ops that span multiple folios, netfslib allows the filesystem to give progress notifications to wake up the collector thread to do a collection of folios that have now been fetched, even if the RPC is still ongoing, thereby allowing the application to make progress. This works by taking the current rreq->cleaned_to value (which indicates which folios have been unlocked) and adding the stashed size of the next folio to it. cleaned_to, however, is subject to 64-bit tearing on a 32-bit arch. Fix this by stashing the next progress notification point as a size_t (which won't tear) to be added to rreq->start (which won't change), with the collector thread calculating that from cleaned_to plus the next folio size. Further, however, if the folios are small, the collector thread gets constantly woken up - which has a negative performance impact on the system. Fix that too by setting a minimum trigger of 256KiB or the size of the folio at the front of the queue, whichever is larger. Note that this has an issue that different subreqs have different need-to-be-cached properties; this is solved by a preceding patch that marks the property on the folios whilst issuing subreqs rather than when collecting them. Also, make sure rreq->cleaned_to is initialised up front, along with rreq->collected_to and stream->collected_to. Fixes: e2d46f2ec332 ("netfs: Change the read result collector to only use one work item") Link: https://sashiko.dev/#/patchset/20260804100224.2748935-1-dhowells%40redhat.com Signed-off-by: David Howells <dhowells@redhat.com> Link: https://patch.msgid.link/20260827134304.2075713-10-dhowells@redhat.com Acked-by: Paulo Alcantara <pc@manguebit.org> cc: Paulo Alcantara <pc@manguebit.org> cc: netfs@lists.linux.dev cc: linux-fsdevel@vger.kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-31netfs: Mark folios with COPY_TO_CACHE whilst issuing subreqsDavid Howells1-2/+4
Mark folios with NETFS_FOLIO_COPY_TO_CACHE whilst issuing subreqs rather than when collecting them. This means that the collector thread doesn't have to try and keep track of which subreqs contribute to which folios - and thus which folios will need to be copied to the cache because at least one byte wasn't in the cache. Instead, this is marked on the folios up front and the collector need only consider the folios. For PG_private_2-using filesystems, PG_private_2 is set instead of NETFS_FOLIO_COPY_TO_CACHE, but otherwise it works the same. The NETFS_RREQ_COPY_TO_CACHE is replaced with NETFS_RREQ_CANCEL_CACHING, which is now set if caching fails somewhere, thereby causing the collection thread to cancel the copy-to-cache marks on the remaining folios. Signed-off-by: David Howells <dhowells@redhat.com> Link: https://patch.msgid.link/20260827134304.2075713-9-dhowells@redhat.com Acked-by: Paulo Alcantara <pc@manguebit.org> cc: Paulo Alcantara (Red Hat) <pc@manguebit.org> cc: Matthew Wilcox <willy@infradead.org> cc: netfs@lists.linux.dev cc: linux-mm@kvack.org cc: linux-fsdevel@vger.kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-31netfs: Fix readahead synchronisation issues by loading all folios upfrontDavid Howells1-0/+3
There are some synchronisation issues that derive from the app thread adding more folios to the rolling buffer whilst the collector thread is looking at them or trying to clear them, such as determining the setting of front_folio_order when the next folio hasn't been added yet, The reason for the rolling buffer approach is that loading the buffer upfront and then dropping all the refs just acquired is quite a slow operation, and loading progressively allows some of the cost to be deferred until after at least some of the I/O is started. Instead, a better way is to load all the folios into the rolling buffer upfront - and then drop the refs later, once the I/O is in progress. (Even better would be for the refs not to be there at all.) Fix this by changing the rolling buffer loader to load all the folios selected by the VM for readahead upfront into the folio queue. The folio queue is allocated a batch worth at a time as we don't know how many folios are involved (the readahead_control struct, alas, has a page count, not a folio count). The folio refs acquired from readahead are then dropped in bulk once the first subrequest is dispatched as it's quite a slow operation. The collector waits for NETFS_RREQ_NEED_PUT_RA_REFS to be cleared so that it doesn't unlock folios before the xarray has been scanned for them. This simplifies the buffer handling later and isn't noticeably slower as the xarray doesn't need to be modified and the folios are all already pre-locked. Fixes: ee4cdf7ba857 ("netfs: Speed up buffered reading") Link: https://sashiko.dev/#/patchset/20260824120224.504575-1-dhowells%40redhat.com Signed-off-by: David Howells <dhowells@redhat.com> Link: https://patch.msgid.link/20260827134304.2075713-8-dhowells@redhat.com Acked-by: Paulo Alcantara <pc@manguebit.org> cc: Paulo Alcantara (Red Hat) <pc@manguebit.org> cc: Matthew Wilcox <willy@infradead.org> cc: netfs@lists.linux.dev cc: linux-mm@kvack.org cc: linux-fsdevel@vger.kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-28net: icmp: avoid invalid transport header access in icmp_send tracepointEric Dumazet1-5/+8
syzbot reported a WARNING triggered by DEBUG_NET_WARN_ON_ONCE(): WARNING: at skb_transport_header include/linux/skbuff.h:3087 [inline] WARNING: at udp_hdr include/linux/udp.h:23 [inline] WARNING: at do_trace_event_raw_event_icmp_send include/trace/events/icmp.h:30 [inline] WARNING: at trace_event_raw_event_icmp_send+0x48c/0x6ec include/trace/events/icmp.h:11 Call trace: skb_transport_header include/linux/skbuff.h:3087 [inline] udp_hdr include/linux/udp.h:23 [inline] do_trace_event_raw_event_icmp_send include/trace/events/icmp.h:30 [inline] trace_event_raw_event_icmp_send+0x48c/0x6ec include/trace/events/icmp.h:11 __traceiter_icmp_send include/trace/events/icmp.h:11 [inline] __do_trace_icmp_send include/trace/events/icmp.h:11 [inline] trace_icmp_send+0x320/0x49c include/trace/events/icmp.h:11 __icmp_send+0xcfc/0x11d8 net/ipv4/icmp.c:1013 ipv4_send_dest_unreach net/ipv4/route.c:1280 [inline] ipv4_link_failure+0x57c/0x8dc net/ipv4/route.c:1287 dst_link_failure include/net/dst.h:438 [inline] vti_tunnel_xmit+0xe40/0x17a4 net/ipv4/ip_vti.c:307 TP_fast_assign() unconditionally calls udp_hdr(skb) before checking whether the packet is UDP. Furthermore, __icmp_send() can be invoked from paths (e.g., link failures, ARP errors, forwarding, AF_PACKET) where skb->transport_header was never initialized (~0U). Under CONFIG_DEBUG_NET=y, calling skb_transport_header(skb) triggers DEBUG_NET_WARN_ON_ONCE(!skb_transport_header_was_set(skb)). Fix this by: 1. Only parsing transport info when iph->protocol == IPPROTO_UDP. 2. Using skb_header_pointer() at skb_network_offset(skb) + (iph->ihl << 2) to safely fetch the UDP header without assuming transport_header is set. Fixes: db3efdcf70c7 ("net/ipv4: add tracepoint for icmp_send") Reported-by: syzbot+6d2762674103618994b0@syzkaller.appspotmail.com Closes: https://lore.kernel.org/netdev/6a8d5538.91706f20.ef82.0009.GAE@google.com/T/#u Signed-off-by: Eric Dumazet <edumazet@google.com> Cc: Peilin He <he.peilin@zte.com.cn> Cc: xu xin <xu.xin16@zte.com.cn> Cc: Steven Rostedt <rostedt@goodmis.org> Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev> Reviewed-by: David Ahern <dsahern@kernel.org> Link: https://patch.msgid.link/20260825084551.1562967-1-edumazet@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-28Merge tag 'f2fs-for-7.3-rc1' of ↵Linus Torvalds1-5/+18
git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs Pull f2fs updates from Jaegeuk Kim: "In this round, key enhancements focus on reducing inode management memory overhead, introducing resizable tail sections with unified pinned allocation, and boosting I/O throughput via parallel multi-device flushes and asynchronous f2fs_write_end_io() execution. We also add dynamic device alias reservations to allow on-the-fly space donation from user partitions. Alongside these features, critical bug fixes resolve folio race conditions, lingering dirty flags, dentry and block counter leaks, and potential deadloops in f2fs_fsync_node_pages(). Additional stability patches address error-path handling across symlink, sync, and rename/unlink operations, prevent pinned file fragmentation, and correct segment migration and free section accounting in free_segment_range. Enhancements: - reduce memory footprint of ino management - support dynamic reserve/release for device aliasing - issue multi-device flushes in parallel - add a way to run f2fs_write_end_io() asynchronously - support resizable tail section and unify pinned allocation Bug fixes: - fix to pass folio->index to f2fs_sanity_check_node_footer() - fix folio_nr_pages() race after put in large folio invalidate - fix to clear dirty flag on folio in error path - accurately adjust free_sections during free_segment_range - fix to avoid potential deadloop in f2fs_fsync_node_pages() - fix the error path in symlink, device alias in rename/unlink, f2fs_sync_fs - fix to migrate all curseg types during free_segment_range - fix to avoid pinfile fragment on fragment:{block, segment} mode - fix valid block count leak on data block allocation failure - fix dentry folio leak in find_in_level - reject overlapping move range after len expansion - fix some bugs related to file pinning, GC functions, i_size And, the series includes a number of minor bug fixes" * tag 'f2fs-for-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs: (51 commits) f2fs: support resizable tail section and unify pinned allocation f2fs: don't leave the hashed inode while it's unlinked f2fs: accurately adjust free_sections during free_segment_range f2fs: fix to avoid potential deadloop in f2fs_fsync_node_pages() f2fs: use adjusted write range after f2fs_write_checks() f2fs: fix to propagate error from f2fs_sync_fs() f2fs: return symlink writeback errors f2fs: fix error handling on device alias check in rename and unlink f2fs: fix to reset all pinned status during fggc f2fs: use f2fs_{down, up}_(read, write}_trace() for nat_tree_lock f2fs: reduce memory footprint of ino management f2fs: fix i_size when pinned fallocate partially fails f2fs: fix to migrate all curseg types during free_segment_range f2fs: avoid setting SBI_NEED_FSCK on transient resize failure f2fs: fix to avoid pinfile fragment on fragment:{block, segment} mode f2fs: cleanup w/ f2fs_need_rand_{blk, seg, seg_blk} f2fs: fix to shrink gc_lock coverage in f2fs_gc_range() f2fs: fix to reclaim space in f2fs_allocate_pinning_section() f2fs: unify add/remove ino entry API for all ino types f2fs: fix to zero post-EOF data when extending file size ...
2026-08-25Merge tag 'tty-7.3-rc1' of ↵Linus Torvalds1-4/+3
git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/tty Pull TTY / serial driver updates from Greg KH: "Here is the "big" set of tty and serial driver updates for 7.3-rc1. Not really all that much happened this development cycle for this subsystem, changes in here are: - removal of the ipwireless driver as it's no longer used or needed - new 8250_mxpcie driver added - qcom serial driver updates and additions - vt mode validation addition - lots of other small serial driver updates and additions All of these have been in linux-next for weeks with no reported issues" * tag 'tty-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/tty: (97 commits) serial: imx: serialize imx_uart_ports[] lifetime tty: clear cdev pointer after cdev_add() failure tty: skip cdev_del() when no cdev is registered serial: core: clear freed pointers on uart_register_driver() failure serial: core: do fallible allocations before the console can be registered serial: 8250_mxpcie: implement rx_trig_bytes callbacks via MUEx50 RTL serial: 8250_mxpcie: introduce per-port private data structure serial: 8250: allow UART drivers to override rx_trig_bytes handling serial: 8250_mxpcie: add break support for RS485 using MUEx50 features serial: 8250: allow low-level drivers to override break control serial: 8250_mxpcie: support serial interface mode switching serial: 8250_mxpcie: speed up TX using memory-mapped FIFO window serial: 8250_mxpcie: speed up RX using memory-mapped FIFO window serial: 8250_mxpcie: add custom handle_irq callback serial: 8250_mxpcie: offload XON/XOFF flow control to MUEx50 hardware serial: 8250_mxpcie: enable automatic RTS/CTS flow control serial: 8250_mxpcie: enable enhanced mode and program FIFO trigger levels serial: 8250: add Moxa MUEx50 UART port type serial: 8250: split Moxa PCIe serial board support out of 8250_pci serial: qcom-geni: Use geni_se_set_perf_level() for baud rate perf level ...
2026-08-24Merge tags 'dma-mapping-7.3-2026-08-24' and 'dma-mapping-7.3-2026-08-24-2' ↵Linus Torvalds1-1/+2
of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux Pull dma-mapping updates from Marek Szyprowski: - swiotlb: - new configuration option for the default pool size (Jagadeesh Pagadala) - reduce overhead for high watermark tracking (chenhuguanshen) - minor code cleanups and improvements (Vova Sharaienko, Honglei Huang and Marek Szyprowski) - add proper tracking of the shared DMA state through direct, pool and swiotlb paths (Aneesh Kumar K.V) This is important for confidential-computing * tag 'dma-mapping-7.3-2026-08-24' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux: dma/swiotlb: decouple high watermark tracking from CONFIG_DEBUG_FS MAINTAINERS: update tree for DMA MAPPING HELPERS dma/swiotlb: introduce Kconfig option for compile-time default pool size dma-direct: Improve readability of the dma_direct_map_sg() for P2PDMA case iommu/dma: simplify dma_iova_destroy() and drop the free_iova helper dma-coherent: use KiB in DMA allocation logs dma-coherent: fix spacing coding style issue * tag 'dma-mapping-7.3-2026-08-24-2' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux: (23 commits) swiotlb: remove unused SWIOTLB_FORCE flag dma: swiotlb: handle set_memory_decrypted() failures dma: swiotlb: free dynamic pools from process context dma-direct: rename ret to cpu_addr in alloc helpers dma-direct: select DMA address encoding from __DMA_ATTR_ALLOC_CC_SHARED dma-direct: set decrypted flag for remapped DMA allocations dma-direct: make dma_direct_map_phys() honor DMA_ATTR_CC_SHARED dma-direct: Move dma_direct_map_phys() to dma/direct.c dma-direct: pass attrs to dma_capable() for DMA_ATTR_CC_SHARED checks dma-mapping: make dma_pgprot() honor __DMA_ATTR_ALLOC_CC_SHARED dma: swiotlb: track pool encryption state and honor DMA_ATTR_CC_SHARED dma: swiotlb: pass mapping attributes by reference dma-pool: track decrypted atomic pools and select them via attrs dma-direct: use __DMA_ATTR_ALLOC_CC_SHARED in alloc/free paths dma-mapping: Add internal shared allocation attribute coco: arm64: s390: powerpc: Mark secure guests with CC_ATTR_GUEST_MEM_ENCRYPT dma-direct: swiotlb: handle swiotlb alloc/free outside __dma_direct_alloc_pages s390: Expose protected virtualization through cc_platform_has() swiotlb: Preserve allocation virtual address for dynamic pools dma: free atomic pool pages by physical address ...
2026-08-24Merge tag 'slab-for-7.3' of ↵Linus Torvalds1-1/+1
git://git.kernel.org/pub/scm/linux/kernel/git/vbabka/slab Pull slab updates from Vlastimil Babka: - Add kfree_rcu_nolock() that can be used from contexts where spinning on a lock might be unsafe, such as a BPF program attached to an arbitrary function, or in NMI context. This complements the existing kfree_nolock() support (Harry Yoo) - Runtime instead of compile-time slabobj_ext sizing. Avoid wasting memory when memory allocation profiling is compiled but not enabled, with initial partial support to also avoid wasting memory for objcg pointers when those are not needed, while profiling is enabled (Vlastimil Babka) - Various non-urgent fixes, cleanups and optimizations (Hao Li, Hongling Zeng, Li RongQing, Li Xiasong, Seongjun Hong, Shengming Hu) * tag 'slab-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/vbabka/slab: (31 commits) mm/slab, kfence, memcg: completely remove obj_ext for kfence objects mm/slab: stop allocating objcg pointers when unnecessary mm/slab: add cache_ and slab_needs_objcg() helpers mm/slab: stop exporting kvfree_rcu_barrier[_on_cache]() slub_kunit: extend the test for kfree_rcu_nolock() mm/slab: introduce kfree_rcu_nolock() mm/slab: introduce struct kvfree_rcu_head for kvfree_rcu batching mm/slab: reduce slabobj_ext memory with allocation profiling disabled mm/slab: introduce slab_obj_ext_has_codetag() mm/slab: allow kfree_rcu_sheaf() on PREEMPT_RT mm/slab: extend deferred free mechanism to handle rcu sheaves mm/slab: use call_rcu() in unknown context if irqs are enabled mm/slab: handle the !allow_spin case in kfree_rcu_sheaf() mm/slab: change struct slabobj_ext to a union mm/slab: replace slab.stride with obj_exts_in_object mm/slab: abstract slabobj_ext.ref access mm/slab: abstract slabobj_ext.objcg access mm/slab: make slab_obj_ext() determine object index mm: move struct slabobj_ext to mm/slab.h mm/slab: remove objs_per_slab() ...
2026-08-23Merge tag 'rcu.2026.08.18a' of ↵Linus Torvalds1-2/+3
git://git.kernel.org/pub/scm/linux/kernel/git/rcu/linux Pull RCU updates from Paul McKenney: "Make expedited grace periods expedite normal RCU callbacks Miscellaneous fixes: - Improve diagnostic output with character task states - Mark accesses to inform KCSAN of concurrency design - Move from kmalloc() to kmalloc_obj() - Documentation updates - Improve handling of RCU deferred quiescent states - Clean up unused function arguments and structure fields - Reduce show_rcu_gp_kthreads() stack space Tasks RCU updates: - Clean up after SRCU re-implementation of Tasks Trace RCU - Mark accesses to inform KCSAN of concurrency design - Add ->lazy_timer status to diagnostic output - Remove an unnecessary memory barrier - Fix a data r