aboutsummaryrefslogtreecommitdiffstatshomepage
path: root/net
AgeCommit message (Collapse)AuthorFilesLines
4 daysipv4: free inet_opt and ireq_opt after an RCU grace periodEric Dumazet2-2/+2
tcp_v4_syn_recv_sock() transfers ownership of ireq->ireq_opt to the child socket (newinet->inet_opt) without copying it. Another cpu can concurrently process a retransmitted SYN for the same request socket, and send a SYNACK from tcp_check_req(). tcp_v4_send_synack() and inet_csk_route_req() read ireq->ireq_opt under rcu_read_lock() only, and ip_build_and_send_pkt() and ip_options_build() then read opt->optlen twice. Note that the SYNACK timer itself is not an issue: it holds a reference on its request socket, and inet_csk_reqsk_queue_drop() calls timer_delete_sync() before the child can be freed. Since commit 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table"), request sockets are processed without holding the listener lock, so nothing prevents the child socket from being freed while the SYNACK is still being built. TCP child sockets do not have SOCK_RCU_FREE, and inet_sock_destruct() frees inet_opt with a plain kfree(), leading to a use-after-free in ip_options_build(). A similar issue exists with request socket migration (net.ipv4.tcp_migrate_req=1, or a BPF_SK_REUSEPORT_SELECT_OR_MIGRATE program). reqsk_timer_handler() clones the request socket with inet_reqsk_clone(), so that the old request socket and its clone share the same ireq_opt, then reqsk_migrate_reset() clears the pointer in the old one. Another cpu holding a reference on the old request socket can still be using these options (sending a SYNACK, or creating a child in tcp_v4_syn_recv_sock()) when the clone is freed, and tcp_v4_reqsk_destructor() also uses a plain kfree(). Readers already use RCU, and other paths replacing inet_opt (do_ip_setsockopt(), cipso_v4_sock_setattr()...) already use kfree_rcu(). Use kfree_rcu() in inet_sock_destruct() and tcp_v4_reqsk_destructor() as well. IPv6 is not affected by the first issue, because tcp_v6_syn_recv_sock() duplicates the options. tcp_v6_reqsk_destructor() has the same migration issue with ipv6_opt, which is only set by CALIPSO. This will be addressed in a separate patch. Fixes: 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table") Fixes: c905dee62232 ("tcp: Migrate TCP_NEW_SYN_RECV requests at retransmitting SYN+ACKs.") Reported-by: Xinyang Ge <xinyang@anthropic.com> Signed-off-by: Eric Dumazet <edumazet@kernel.org> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20261001221253.2964024-1-edumazet@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
4 daysnet: make dev_xdp_sb_prog_count() see programs attached through a linkJakub Kicinski1-4/+7
ethtool refuses to enable tcp-data-split on a device running a single-buffer XDP program (which can't deal with the payload landing in a separate buffer). dev_xdp_sb_prog_count() only looks at xdp_state[].prog though, ignoring bpf_link integration. Get the prog pointer from dev_xdp_prog(), which will consult both. bpf_xdp_link_update() needs a bit of a touch-up, too, since ethtool calls dev_xdp_sb_prog_count() with just the instance lock (for ops-locked devices) - we now have to make sure the link updates happen under that lock. Spotted by AI while reviewing the XDP propagation series, verified and tested on netdevsim with a C test along these lines: link = bpf_program__attach_xdp(prog, ifindex); system("ethtool -G eth0 tcp-data-split on") Fixes: 197258f0ef68 ("net: ethtool: add hds_config member in ethtool_netdev_state") Acked-by: Stanislav Fomichev <sdf@fomichev.me> Link: https://patch.msgid.link/20261001012559.297751-1-kuba@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
5 daysopenvswitch: fix soft lockup in the netlink flow dumpDenis V. Lunev1-0/+8
A production compute node carrying a few thousand datapath flows hit a soft lockup inside a single netlink flow dump and panicked. ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it emits, and that releases stats->lock with spin_unlock_bh() once per CPU that has touched the flow. Every release is a local_bh_enable(), and each one runs the pending softirq backlog in the dumping thread's own context. The skb bounds the flows one callback emits, but not the softirq work it absorbs. On a CPU that carries the box's packet load the backlog refills as fast as it drains, so the dumping thread becomes that CPU's softirq engine. It never sleeps and it has no reschedule point, so under voluntary preemption nothing can take the CPU away from it: neither the ksoftirqd the kernel woke to take the work over, nor the stopper thread the softlockup detector dispatches to refresh its timestamp. Hold BH off across the whole callback instead, the way ctnetlink_dump_table() does, so the nested spin_unlock_bh() stop draining softirqs. The loop already runs under rcu_read_lock() and cannot sleep. What it gives up is preemption under CONFIG_PREEMPT, since a BH-off region is not preemptible outside PREEMPT_RT. The region stays short: the skb caps the flows one callback emits, and empty buckets cost no skb space but are each visited once per dump, as the cursor only moves forward. The table grows on insert and shrinks only on flush, so the walk is bounded by the largest flow count the datapath has held. The softirq backlog the callback used to absorb has no bound at all. Fixes: 63e7959c4b9b ("openvswitch: Per NUMA node flow stats.") Cc: stable@vger.kernel.org Signed-off-by: Denis V. Lunev <den@openvz.org> Reviewed-by: Ilya Maximets <i.maximets@ovn.org> Signed-off-by: David S. Miller <davem@davemloft.net>
5 daysnet: bridge: avoid recursive multicast port cleanupZixuan Chai1-5/+18
br_multicast_sg_del_exclude_ports() removes automatically added STAR_EXCL port groups by calling br_multicast_del_pg(). For S,G entries, that deletion calls back into br_multicast_sg_del_exclude_ports(), so a large number of STAR_EXCL ports consumes one kernel stack frame per port and can hit the stack guard page. Keep the complete port-group deletion path for cleanup, but suppress the S,G exclude-port cleanup while that helper is already walking the same entry. Normal deletion callers retain the existing behavior. Fixes: 8266a0491e92 ("net: bridge: mcast: handle port group filter modes") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Assisted-by: LLM Signed-off-by: Zixuan Chai <petalzu987@gmail.com> Signed-off-by: Ren Wei <weir@nebusec.ai> Signed-off-by: David S. Miller <davem@davemloft.net>
7 daystcp: reject net_iov in zerocopy receive mapping hintsDaehyeon Ko1-0/+2
After copying a readable prefix, receive_fallback_to_copy() asks tcp_zerocopy_set_hint_for_skb() where page mapping can resume. If the next skb is unreadable, find_next_mappable_frag() passes its net_iov fragment to can_map_frag(). skb_frag_page() returns NULL for a net_iov, but can_map_frag() dereferences it in PageCompound(). A v7.2 KASAN run on a connected TCP socket with 64 readable bytes followed by a 4096-byte NET_IOV_DMABUF fragment reported: BUG: KASAN: null-ptr-deref in can_map_frag tcp_zerocopy_receive -> can_map_frag Kernel panic - not syncing: KASAN: panic_on_warn set The diagnostic inserted the net_iov directly because the test host has no devmem-capable NIC. Hardware end-to-end reachability remains untested and requires CONFIG_NET_DEVMEM plus a supported DMA-buf-bound RX queue. Reject all net_iov fragments before skb_frag_page(). This covers both DMABUF and IOURING net_iov types while leaving page-backed checks unchanged. With the guard, the same queue copied the readable prefix, returned a 4096-byte skip hint, and completed without a fault. Fixes: 9f6b619edf2e ("net: support non paged skb frags") Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Reviewed-by: Mina Almasry <almasrymina@google.com> Link: https://patch.msgid.link/20260930102524.1659847-1-4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daystipc: destroy topsrv workqueues before closing connectionsYuqi Xu1-15/+65
tipc_topsrv_stop() closed subscriber connections while the topology server's send and receive workqueues were still running. Socket callbacks and subscription events could then queue more work, and in-flight send/recv work could drop the last connection reference during the conn_idr walk. That race produced several teardown failures: refcount_t addition on 0 from conn_get() on a connection whose release was blocked on idr_lock, a subsequent use-after-free in tipc_conn_close(), queue_work() on an already destroyed workqueue from the listener data-ready callback, and an RCU stall in tipc_topsrv_exit_net() while the walk spun under idr_lock. Clear srv->listener under idr_lock so it acts as a shutdown flag, skip queue_work() once it is NULL, destroy the workqueues to flush in-flight work, and only then close the remaining connections. Refuse tipc_conn_lookup() after that flag is cleared, and drop idr_lock when the teardown walk finds no connection, so an in-flight subscription event cannot pin idr_in_use while the walk holds the lock. v3 supersedes the narrower in-thread diff Tung Quang Nguyen posted on 2026-09-23. It keeps his teardown order and adds the lookup refusal and the empty-idr unlock. Fixes: c5fa7b3cf3cb ("tipc: introduce new TIPC server infrastructure") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Assisted-by: LLM Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com> Reviewed-by: Ren Wei <weir@nebusec.ai> Signed-off-by: David S. Miller <davem@davemloft.net>
8 daysnet/packet: guard the ll header push in packet_rcv_spkt()Quchaosheng1-1/+2
packet_rcv_spkt() restores the link layer header with skb_push(skb, skb->data - skb_mac_header(skb)); That subtraction is only meaningful when the device actually has a link layer header. packet_rcv() and tpacket_rcv() both wrap it in dev_has_header(), which is also the predicate the block comment at the top of the file states the restore in terms of; packet_rcv_spkt() does not. commit d549699048b4 ("net/packet: fix packet receive on L3 devices without visible hard header") introduced the helper and changed the two call sites, and this one stayed behind. A device without a visible ll header can leave skb->mac_header at the 0xFFFF sentinel that __alloc_skb() initialises it to. A CAN skb does: init_can_skb() sets pkt_type and ip_summed but does not reset the headers, and commit 9f10374bb024 ("can: remove private CAN skb headroom infrastructure") dropped the skb_reset_*_header() calls that used to be there. skb_mac_header() is then 0xFFFF, the length becomes a large negative number and skb_push() reports it through skb_under_panic() -- from softirq context, so it is a full system panic even with panic_on_oops=0: skbuff: skb_under_panic: text:ffffffff8bd21bc1 len:-65455 put:-65471 head:... data:... tail:0x50 end:0x180 dev:can0 kernel BUG at net/core/skbuff.c:214! RIP: 0010:skb_panic+0x50/0x60 Call Trace: <IRQ> skb_push+0x38/0x40 packet_rcv_spkt+0xe1/0x170 __netif_receive_skb_core.constprop.0+0x7e8/0xd30 ... Kernel panic - not syncing: Fatal exception in interrupt The socket type is reachable: packet_create() accepts SOCK_PACKET alongside SOCK_RAW and SOCK_DGRAM behind the same CAP_NET_RAW check, and neither the socket length nor a capability check keeps it away from a CAN interface. The missing skb_reset_*_header() calls in init_can_skb() are a regression in their own right and are being fixed separately, but a packet socket should not turn a link layer that did not initialise its mac header into a kernel panic. Guard the push the way the other two receive paths do. Tested on v7.3-rc5 under QEMU with a slcan device on a pty, which is the driver RX path: vcan does not reproduce it, because can_send() resets the headers on the way out. One SOCK_PACKET socket bound to can0 and one frame written into the line discipline panics an unpatched kernel with the trace above; the same image with this patch prints no panic and powers off normally. Both kernels are this tree, defconfig plus CONFIG_CAN_SLCAN=y, differing only in this hunk. Fixes: d549699048b4 ("net/packet: fix packet receive on L3 devices without visible hard header") Cc: stable@vger.kernel.org Signed-off-by: Quchaosheng <quchaosheng000406@163.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260928113108.2127215-1-quchaosheng000406@163.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daysnet/sched: em_text: reject unterminated algo before autoloadJamal Hadi Salim1-0/+3
This is a follow-up to commit 9c572a83037a ("net/sched: fix potential stack infoleak in em_text_dump()"), which zero-initialised the dump-side struct. Sashiko pointed at a pre-existing bug that exposed the change-side input-validation gap open. em_text_change() casts the raw netlink attribute payload to struct tcf_em_text and passes conf->algo, a 16-byte array with no enforced NUL terminator, to textsearch_prepare(). On the TS_AUTOLOAD retry textsearch_prepare() calls request_module("ts_%s", algo); vsnprintf() walks %s until it finds a NUL, so an algo[] with no NUL reads past the array into the adjacent struct fields and payload and feeds those bytes to the usermode helper command line. This is an out-of-bounds read and a kernel memory disclosure. Reject the attribute when no NUL exists within TC_EM_TEXT_ALGOSIZ bytes, matching the check xt_string has had since 3ab720881b6e. Conditions to recreate the bug: - CAP_NET_ADMIN on the target netns (namespace-local via unshare -Urn is enough) - install clsact on an interface, then add a basic filter with one text ematch whose algo[] is 16 bytes with no NUL (e.g. "A"*16) and pattern_len at least 1 - the add fails with -ENOENT and the kernel invokes modprobe with a module name made of "ts_" plus the 16 bytes and the adjacent payload Fixes: d675c989ed2d4 ("[PKT_SCHED]: Packet classification based on textsearch (ematch)") Reported-by: Sashiko (gemini) <sashiko-bot@kernel.org> Closes: https://sashiko.dev/#/patchset/20260918133953.12494-1-bernard.ladenthin%40gmail.com Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Link: https://patch.msgid.link/QDISC-2DIU.v1.20260929183844@mojatatu.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daystipc: hold a reference to nodes found by link nameChengfeng Ye1-1/+9
tipc_node_find_by_name() returns a node after dropping its RCU read lock without taking a reference. The LINK_SET, LINK_GET and LINK_RESET_STATS handlers then lock and access the node, racing with timer-driven cleanup of a down peer. Generic netlink serialization does not exclude the node timer. The following interleaving can leave a handler using a freed node: CPU 0: find the node under RCU and release the node read lock CPU 1: tipc_node_timeout() clears the links and unlinks the down node CPU 1: drop the list and timer references, queuing tipc_node_free() CPU 0: leave the RCU read-side critical section CPU 1: complete the grace period and free the node CPU 0: acquire the node lock through the stale pointer LINK_SET also uses the node's media address after releasing the node lock, when passing queued packets to tipc_bearer_xmit(). KASAN on v7.3-rc5 reported: BUG: KASAN: slab-use-after-free in _raw_read_lock_bh Write of size 4 at addr ffff88807e06f008 by task poc/92 Call Trace: _raw_read_lock_bh kernel/locking/spinlock.c:287 tipc_nl_node_set_link net/tipc/node.c:2475 genl_family_rcv_msg_doit net/netlink/genetlink.c:1114 genl_rcv_msg net/netlink/genetlink.c:1209 netlink_rcv_skb net/netlink/af_netlink.c:2575 Allocated by task 0: tipc_node_create net/tipc/node.c:539 tipc_node_check_dest net/tipc/node.c:1196 tipc_disc_rcv net/tipc/discover.c:252 Freed by task 92: kfree mm/slub.c:6801 rcu_core kernel/rcu/tree.c:2919 Last potentially related work creation: __call_rcu_common.constprop.0 kernel/rcu/tree.c:3181 tipc_node_timeout net/tipc/node.c:814 Acquire a reference to the selected node with kref_get_unless_zero() before leaving RCU, returning NULL if the node has already been released. Release that reference on every caller exit after the last node access, including transmission in LINK_SET. Keep the existing link lookup order and locking so concurrent link removal still takes the existing error paths. Fixes: 6a939f365bdb ("tipc: Auto removal of peer down node instance") Cc: stable@vger.kernel.org Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Link: https://patch.msgid.link/20260928161246.1895273-1-nicoyip.dev@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
9 daysMerge tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nfPaolo Abeni14-18/+77
Pablo Neira Ayuso says: ==================== Netfilter/IPVS fixes for net The following batch contains Netfilter fixes for net. This batch fixes crashes as recent feature regression, one of the due to a dependency that has been pulled into -stable: 1) Expand existing ipset fix for bitmap sets to disallow comments updates from kernel-side adds, from Florian Westphal. 2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise flowtable cannot ever be removed, from Aohan Mei. 3) nft_rbtree GC should collect end elements that contained in this transaction batch, new or deleted elements are never expired. From Weiming Shi. 4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_* values, from Fernando F. Mancera. 5) Flowtable GC must skip flows that are pending hardware updates, generalize the PENDING flag and use it to inhibit GC. 6) Restore flowtable with ieee80211 which broke due to a relatively recent commit, which was pulled in by -stable, causing a regression in 6.18 kernels. And the following IPVS fixes: 1) Fix accounting of cache entries in IPVS LBLC for destinations, which eventually fills up the table and trigger recurrent resizing, from Julian Anastasov. 2) Limit IPVS cache growth for LBLCR and LBLC schedulers, from Zhiling Zou. 3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections, do not allow to use it with templates. Also from Julian. 4) Sanitize flags in IPVS sync messages received in the backup. From Julian Anastasov. netfilter pull request 26-09-30 * tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf: netfilter: flowtable: restore ieee80211 forward path netfilter: flowtable: generalize pending status bit netfilter: bpf: reject invalid NAT manipulation types netfilter: nft_set_rbtree: skip transaction elements during GC ipvs: filter some flags received in the backup server ipvs: do not create invisible templates ipvs: bound LBLCR and LBLC cache growth ipvs: fix missing counter decrement in lblc netfilter: nft_flow_offload: drop flowtable reference on init error path netfilter: ipset: do not update comments from kernel-side adds ==================== Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org Signed-off-by: Paolo Abeni <pabeni@redhat.com>
9 daysipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel()Eric Dumazet1-1/+1
Since commit d58e468b1112 ("flow_dissector: implements flow dissector BPF hook") __skb_flow_dissect() needs a net pointer, either from skb->dev, skb->sk, or since commit 3cbf4ffba5ee ("net: plumb network namespace into __skb_flow_dissect") a caller provided pointer. syzbot was able to reach seg6_make_flowlabel() with an skb having neither skb->dev nor skb->sk set: a TIPC UDP bearer sends a discovery message through an IPv4 route using seg6 encap, while net.ipv6.seg6_flowlabel is set to 1. seg6_make_flowlabel() already has a net pointer, use skb_get_hash_net(). WARNING: net/core/flow_dissector.c:1131 at __skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126, CPU#0: syz.0.17/4930 Call trace: __skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126 (P) __skb_get_hash_net+0xe0/0x29c net/core/flow_dissector.c:1903 skb_get_hash include/linux/skbuff.h:1663 [inline] seg6_make_flowlabel+0xcc/0x1ec net/ipv6/seg6_iptunnel.c:132 __seg6_do_srh_encap+0x320/0xbf4 net/ipv6/seg6_iptunnel.c:161 seg6_do_srh+0x4b4/0xa44 net/ipv6/seg6_iptunnel.c:431 seg6_output_core+0x164/0x688 net/ipv6/seg6_iptunnel.c:681 seg6_output+0x44/0x1ac net/ipv6/seg6_iptunnel.c:744 lwtunnel_output+0x3d0/0x664 net/core/lwtunnel.c:356 dst_output include/net/dst.h:470 [inline] ip_local_out+0x110/0x148 net/ipv4/ip_output.c:131 iptunnel_xmit+0x50c/0xd38 net/ipv4/ip_tunnel_core.c:97 udp_tunnel_xmit_skb+0x220/0x348 net/ipv4/udp_tunnel_core.c:187 tipc_udp_xmit+0x75c/0x9c0 net/tipc/udp_media.c:202 tipc_udp_send_msg+0x214/0x374 net/tipc/udp_media.c:274 tipc_bearer_xmit_skb+0x260/0x3b0 net/tipc/bearer.c:576 tipc_enable_bearer net/tipc/bearer.c:366 [inline] __tipc_nl_bearer_enable+0xc90/0xfb0 net/tipc/bearer.c:1048 tipc_nl_bearer_enable+0x2c/0x48 net/tipc/bearer.c:1057 genl_family_rcv_msg_doit+0x1e4/0x2d4 net/netlink/genetlink.c:1114 Fixes: d58e468b1112 ("flow_dissector: implements flow dissector BPF hook") Reported-by: syzbot+9408fbe0e6452a12e9ab@syzkaller.appspotmail.com Closes: https://lore.kernel.org/netdev/6abac2eb.3654fce1.1bec97.0000.GAE@google.com/ Signed-off-by: Eric Dumazet <edumazet@kernel.org> Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/20260928194524.3617299-1-edumazet@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
9 dayssctp: check RCV_SHUTDOWN after the sendmsg connect waitJun Yang1-1/+1
sctp_wait_for_connect() drops the socket lock while it sleeps. An out-of-the-blue ABORT can then be processed from the socket backlog and unlink the association. If a concurrent shutdown(fd, SHUT_RD) sets RCV_SHUTDOWN, the waiter breaks with err == 0 before checking asoc->base.dead. Its final sctp_association_put() can then free the association, leaving sctp_sendmsg_to_asoc() to continue with a dangling pointer. Check RCV_SHUTDOWN along with the wait error in sctp_sendmsg_to_asoc() before using the association again. The check only accesses the socket, so it needs no additional association reference. Return the existing -ESRCH so that sctp_sendmsg() skips freeing a new association that may already have been destroyed. Keep sctp_wait_for_connect() unchanged to preserve its behavior for the connect() caller. Fixes: 668c9beb9020 ("sctp: implement assign_number for sctp_stream_interleave") Cc: stable@vger.kernel.org Reported-by: TencentOS Corvus AI <corvus@tencent.com> Signed-off-by: Jun Yang <junvyyang@tencent.com> Acked-by: Xin Long <lucien.xin@gmail.com> Link: https://patch.msgid.link/20260926095606.68601-1-juny24602@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
9 daysMerge tag 'wireless-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wirelessJakub Kicinski16-61/+293
Johannes Berg says: ==================== Still more fixes coming in, notably: - ath11k: avoid running out of stations on HW restart - mac80211: - drop too large fragmented MPDUs - mesh path handling fixes - validation improvements - reject CSA with bad 320 MHz bandwidth - cfg80211: fix RTS for single radio devices * tag 'wireless-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (27 commits) wifi: mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue() wifi: mac80211: reject invalid 320 MHz CSA bandwidth wifi: mac80211: set info->band for 802.3 encap offload frames wifi: mac80211: prevent AP VLAN tx from other interfaces wifi: cfg80211: preserve hidden-group beacon IE ownership wifi: mac80211: keep fallback association elements alive wifi: mac80211: shut down RX BA session timer on teardown wifi: mac80211: validate TX status rate metadata wifi: ath9k_htc: bound TX aggregation to MAX_TX_BUF_SIZE wifi: ath9k: reject short WMI command responses wifi: ath9k: Clean up device initialisation guards wifi: ath11k: reset ar->num_stations on hardware start wifi: cfg80211: fix RTS threshold setting for single-radio PHY wifi: mac80211: handle empty FILS association request payload wifi: mac80211: minstrel_ht: validate fixed rate index wifi: p54: validate firmware record lengths wifi: mac80211: fix mesh fast xmit path deletion UAF wifi: mac80211: drop oversized fragments to avoid extra_len overflow wifi: wlcore: Fix runtime PM leak in wlcore_remove() wifi: mac80211: drain PS delivery work during station teardown ... ==================== Link: https://patch.msgid.link/20260930124440.224799-3-johannes@sipsolutions.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
10 daysnetfilter: flowtable: restore ieee80211 forward pathPablo Neira Ayuso2-0/+10
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered"), there was a fallback to set up a forward path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used. Such fallback was used by commit d787a3e38f01 ("mac80211: add support for .ndo_fill_forward_path"). One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable forward path discovery. However, this is only used internally by drivers to retrieve mtk_wdma information to set up hardware offload. Felix decided to use the .fill_forward_path interface for this purpose due to the lack of a better interface at that time. Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211 flag is set on in the struct net_device_path_ctx to restore the flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211 path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last netdevice in the stack. This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path. Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered") Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daysnetfilter: flowtable: generalize pending status bitPablo Neira Ayuso3-11/+12
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the flowtable GC worker until pending hw offload work has been completed. Apparently, nf_flow_offload_stats() can schedule work to retrieve stats while the flow is being removed by GC. And this bit can also be used in a follow up patch to disable GC until the flow has been fully added in both directions. Revert the reordering done in commit d644b23afe1e ("netfilter: flowtable: publish HW_DEAD after worker is done") to prevent a race between GC and hw offload handler. Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work") Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daysnetfilter: bpf: reject invalid NAT manipulation typesFernando Fernandez Mancera2-2/+6
As bpf_ct_set_nat_info() is not validating the NAT manipulation type a wrong value can be passed directly to nf_nat_setup_info(). This triggers the WARN_ON() at nf_nat_setup_info() and if panic_on_warn isn't set, then IPS_SRC_NAT_DONE is set without adding nat_bysource and conntrack cleanup tries to unlink an uninitialized hlist node. Fix this by checking that NAT manipulation type is correct before calling nf_nat_setup_info(). In addition, if the WARN_ON is hit, return NF_DROP instead of continuing with the processing to avoid similar situations in the future. Reported-by: VEGA <vega@nebusec.ai> Fixes: 0fabd2aa199f ("net: netfilter: add bpf_ct_set_nat_info kfunc helper") Signed-off-by: Fernando Fernandez Mancera <fmancera@suse.de> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daysnetfilter: nft_set_rbtree: skip transaction elements during GCWeiming Shi1-0/+2
Since nft_set_commit_update() runs set commit callbacks before processing NEWSETELEM transactions, nft_rbtree_gc_scan() can observe elements added by the transaction being committed. The scan records an interval end in rbe_end without checking the element's transaction state. A later, unrelated expired start then moves both elements to the expired list. The synchronous GC queue can free the new end element before the transaction subsequently activates it, causing a use-after-free. Only consider elements that are fully active in both generations. This keeps transaction-state elements out of the GC scan and preserves interval pairing across skipped elements. KASAN reports: BUG: KASAN: slab-use-after-free in nft_setelem_activate nft_setelem_activate net/netfilter/nf_tables_api.c:7047 nf_tables_commit net/netfilter/nf_tables_api.c:11137 Allocated by task 130: nft_set_elem_init net/netfilter/nf_tables_api.c:6794 nft_add_set_elem net/netfilter/nf_tables_api.c:7523 Freed by task 130: nft_trans_gc_trans_free net/netfilter/nf_tables_api.c:10506 rcu_core kernel/rcu/tree.c:2919 Fixes: 1e3b9e1c77fe ("netfilter: nf_tables: call set ops .commit when building new ruleset blob") Reported-by: <co+ee5e50ef2670e5f4@bugs.sh> Assisted-by: LLM Signed-off-by: Weiming Shi <bestswngs@gmail.com> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daysipvs: filter some flags received in the backup serverJulian Anastasov1-3/+30
While the IPVS SYNC protocol is not secure by design we can still protect the backup server from messages that can wreak havoc. This commit addresses problems from received connection flags or their combinations. We now drop messages as follows: 1. the NO_CPORT+TEMPLATE combination allows lookups for normal connections to hit template which can break in many ways. While the master does not sync connections with NO_CPORT flag, i.e. before they are established, we still accept NO_CPORT without TEMPLATE. 2. ONE_PACKET: it is not sent by master, so we do not expect it in backup. Before now it was ignored by IP_VS_CONN_F_BACKUP_MASK for protocol v1 while protocol v0 created connections that are not hashed and dropped immediately. Better to apply the IP_VS_CONN_F_BACKUP_MASK also to the flags from v0 messages for consistency with v1. Fixes: 87375ab47cd0 ("[IPVS]: ip_vs_ftp breaks connections using persistence") Signed-off-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daysipvs: do not create invisible templatesJulian Anastasov1-0/+3
The IP_VS_CONN_F_ONE_PACKET flag was implemented for normal connections. When conn template inherits this flag from dest->conn_flags it will not be hashed. As result, we will create new template for every new normal connection. Fix it to allow one template to be used by many normal connections. Fixes: 26ec037f9841 ("IPVS: one-packet scheduling") Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260916231652.127456-1-pablo%40netfilter.org Signed-off-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daysipvs: bound LBLCR and LBLC cache growthZhiling Zou2-0/+6
ip_vs_lblcr_new() and ip_vs_lblc_new() create cache entries for every previously unseen destination address. The table max_size only tells the periodic collector to reclaim entries after the cache has already exceeded the limit. It does not reclaim entries that the attacker continues to use. Reject new cache entries once either table reaches max_size * 3 / 2. The extra headroom lets the periodic collector catch up while the existing scheduler fallback continues to use the selected destination when cache creation fails. New traffic therefore stays serviceable without growing the tables further. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Suggested-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai> Acked-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daysipvs: fix missing counter decrement in lblcJulian Anastasov1-0/+1
LBLC may delete cache entries for destinations that are removed or overloaded and replace them with available ones. But ip_vs_lblc_new() forgets to decrement the tbl->entries counter after calling ip_vs_lblc_del(). This can lead to increased shrinking of the cache with every new garbage collection. Fixes: 2f3d771a35fe ("ipvs: do not use dest after ip_vs_dest_put in LBLC") Link: https://sashiko.dev/#/patchset/0bdd5abe9968ded7ca2b9cb6844ba83d94cc8d53.1787318053.git.zhilinz%40nebusec.ai Signed-off-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daysnetfilter: nft_flow_offload: drop flowtable reference on init error pathAohan Mei1-1/+6
nft_flow_offload_init() bumps the flowtable use count with nft_use_inc() before calling nf_ct_netns_get(). When the latter fails, the error is returned as-is and the reference is leaked. The upper layers do not balance it either: nf_tables_newexpr() clears expr->ops when the expression init callback fails, so the nft_expr_more() iteration in nft_rule_expr_deactivate() and nf_tables_rule_destroy() stops right before the failed expression and its ->destroy callback, which would drop the reference, never runs. Each failed rule addition therefore leaks one flowtable reference and the flowtable can no longer be removed: NFT_MSG_DELFLOWTABLE keeps reporting -EBUSY even though no rule references it. Save the nf_ct_netns_get() return value and undo the nft_use_inc() when it fails, restoring the inc/dec pairing within nft_flow_offload_init() itself. Fixes: a3c90f7a2323 ("netfilter: nf_tables: flow offload expression") Reported-by: TencentOS Corvus AI <corvus@tencent.com> Cc: stable@vger.kernel.org Assisted-by: CodeBuddy:Kimi-K3 Signed-off-by: Aohan Mei <henrymei@tencent.com> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
10 daystcp: refresh TS.Recent for accepted old ACKsJeff Jo1-0/+6
A TCP packet can carry new data while acknowledging traffic in the opposite direction. With overlapping traffic in both directions, a delayed packet's acknowledgment can be older than one Linux has already accepted, even when that packet fills a gap in the received data. Linux accepts the data, but tcp_ack() takes the old_ack path and skips updating TS.Recent, the timestamp saved for outgoing acknowledgments. The reply therefore echoes an older timestamp. If the sender uses this echo to measure round-trip time after a long idle period, its estimate includes the idle time and can reduce its sending rate. Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK processing can trigger a transmission. This reuses the existing timestamp and sequence checks, including PAWS protection against old duplicate packets. ACK validation already rejects old ACKs in SYN_RECV before this path, so no additional state check is needed. Echoing the timestamp of the packet that fills the receive gap follows RFC 7323 section 4.3. In a socket reproduction with 300 seconds idle, controlled reordering and retransmission to exercise timestamp-based RTT sampling, the sender's smoothed round-trip time was 37.5 seconds without the fix and 15.5 ms with it. Fixes: 12fb3dd9dc3c ("tcp: call tcp_replace_ts_recent() from tcp_ack()") Assisted-by: LLM sparse Signed-off-by: Jeff Jo <jeffjo@openai.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Signed-off-by: David S. Miller <davem@davemloft.net>
10 daysnet: extend IPv6 exthdr detection of tunneled packetsWillem de Bruijn1-16/+31
Commit c4336a07eb6b ("net: correctly handle tunneled traffic on IPV6_CSUM GSO fallback") split skb_gso_has_extension_hdr() into mutually exclusive branches on skb->encapsulation. This did not yet address all paths: 1. With skb->encapsulation set, the outer header is not checked. IPv6 tunnels such as ip6_gre and ip6_tunnel add an outer Destination Options header by default (encap_limit). Their GSO packets skip software GSO, then skb_csum_hwoffload_help() sees the outer extension header and calls skb_checksum_help() on the GSO skb, which warns and drops it. 2. Tunnels over IPv6 without an outer transport header, such as ip6_tunnel, leave skb->transport_header at the inner transport header. skb_network_header_len() then spans the outer IPv6, tunnel and inner IP headers, a false positive. 3. UDP tunnels without an inner network header, such as SCTP-in-UDP or PSP, have no inner IPv6 header to check. Decide on the outer header alone. 4. Directly dereferencing inner_ip_hdr(skb)->version without skb_header_pointer() is unsafe. Instead, check ipv6_ext_hdr(nexthdr) on the outer IPv6 header and, if set, on the inner IPv6 header. Read the headers with skb_header_pointer(). Keep the skb_network_header_len() check when !skb->encapsulation to also catch encapsulated packets without skb->encapsulation (e.g., virtio). Use the same helper in skb_csum_hwoffload_help(). Its open coded test has the false positive of (2) and ignores the inner header. Background: checksum offload of tunneled packets invariants: Non-GSO skb: - If the inner packet is CHECKSUM_PARTIAL, Local Checksum Offload computes the outer checksum in software and the device offloads only the inner L4 checksum. - If the inner packet is CHECKSUM_NONE (e.g., SCTP-in-UDP, ESP-in-UDP, Remote Checksum Offload), the device offloads the outer UDP or GRE checksum instead. GSO skb: - A device with NETIF_F_GSO_UDP_TUNNEL_CSUM or NETIF_F_GSO_GRE_CSUM computes both inner and outer checksums per segment. The outer checksum is seeded with only the pseudo-header checksum. A NETIF_F_IPV6_CSUM device must parse through the outer headers to reach the inner ones. Fixes: c4336a07eb6b ("net: correctly handle tunneled traffic on IPV6_CSUM GSO fallback") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260926140506.2335137-1-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
10 daysnet: cap skb->queue_mapping when the tx queue is pickedJamal Hadi Salim2-10/+26
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit() cleared the flag before sch_handle_egress() and only read it afterwards, so the flag was not confined to the xmit that set it: a nested xmit (mirred redirect or mirror, or a drop after skbedit) could set the flag and the outer xmit would consume it for an skb that never went through skbedit. A forwarded packet still carries the ingress NIC's rx_queue + 1 in skb->queue_mapping, so the outer device then indexes its tx queue state with that stale value. Taprio's child array q->qdiscs[] is sized to the device's queue count, so taprio_enqueue() indexes past its allocation and dereferences the result as a struct Qdisc *. We (ab)use the skb->nf_skip_egress which means "skip netfilter egress for this packet" to tag to "am I in tc egress?". Despite the overload I dont see it as a conflict since the marker is set only around the single sch_handle_egress() call and ingress path is guarded by tc_at_ingress. I will send a followup(net-next) patch once this hits net-next to rename the skb->nf_skip_egress bit/flag to skb->skip_egress Arm the flag only from the egress classifier that can use it: raise skip_txqueue from tcf_skbedit_act() only when it runs inside sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc classifier runs in q->enqueue(), after the tx queue has been picked, so a mapping it sets cannot affect the current packet; arming the flag there only pollutes it for a later xmit. Then own the flag for the xmit frame the egress hook runs in: save the incoming value and clear it just before sch_handle_egress(), and restore it after the hook - on the consumed (drop) path, or, in the same call that reads it, on the surviving path. The save and the restores stay inside the egress_needed_key static branch, so a packet pays for them only when egress hooks are active (2f1e85b1aee4). Store the value netdev_cap_txqueue() selected back into skb->queue_mapping in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so the skip_txqueue path never hands a later reader on the xmit path a mapping the device cannot serve. A store made still later in the same frame, by a tc BPF program attached to a transmit qdisc, is outside this path and is not re-capped; a separate followup will resolve that path. netdev_xmit_skip_txqueue() returns the previous flag value so the save-and-clear is one call, and a no-op stub is provided when CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so skb_at_tc_egress() is valid whenever the egress path is built. A local user in a network namespace can redirect a packet from a device with more TX queues to one with fewer after setting a mapping valid only on the larger device. That reaches these reads and, under KASAN, faults with "slab-out-of-bounds in taprio_enqueue". Conditions to recreate the bug: the report's own trigger is a local user with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed. With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y, CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y, CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb (2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb (num_tc 1, queues 2@0) and clsact on all three; then add an egress matchall filter on every device. On qa: "action skbedit queue_mapping 2 pipe action mirred egress redirect dev qb". On qb: "action mirred egress mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one packet out qa. qc's skbedit sets the flag while qb's outer xmit is in flight; without the fix qb consumes it and reads its two-entry taprio child array with the forwarded packet's stale mapping. A qc whose skbedit is instead installed in a transmit-qdisc classifier (a matchall filter on the qc root qdisc) reaches the same code path the same way without the fix. Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes past a 16-byte taprio_init() allocation, for the clsact-setter and the transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF store; the fixed kernel runs all three with no report, and the BPF store variant additionally shows the expected "selects TX queue" clamp notice from the write-back. Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue") Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com> Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/ Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/ Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/ Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/ Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/ Suggested-by: Eric Dumazet <edumazet@google.com> Suggested-by: Jakub Kicinski <kuba@kernel.org> Tested-by: hybris <hybris@mojatatu.ai> Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
10 daysnet: fix checksum offsets in skb_splice_from_iter()Alireza Asgari1-3/+5
skb_splice_from_iter() appends multiple page fragments before updating skb->len. However, skb_splice_csum_page() uses that unchanged length as the offset for every fragment checksum. If bytes already appended during the call have an odd total length, a later fragment's checksum is combined with the wrong parity. For example, prime a UDP socket with sendto(MSG_MORE) and splice two distinct pipe buffers containing "abc" and "DEFGH". Uncorking reports success, but the receiver discards the packet for a bad checksum. This also affects IPv6 and fragments beginning near a page boundary. Pass the initial skb length plus the bytes already spliced to the checksum helper. Keep the existing final length update and partial-progress error handling unchanged. CHECKSUM_PARTIAL does not use this helper and remains unaffected. Fixes: 2e910b95329c ("net: Add a function to splice pages into an skbuff for MSG_SPLICE_PAGES") Cc: stable@vger.kernel.org Signed-off-by: Alireza Asgari <alireza@asgari.net> Link: https://patch.msgid.link/20260924-fix-udp-splice-checksum-v1-1-fe61d65447a8@asgari.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
10 daysnet/packet: prevent TX_RING byte count overflowIlkka Lappeteläinen1-0/+3
tpacket_snd() drains every SEND_REQUEST frame in a TX ring and adds each valid packet length to the signed int len_sum. Userspace can recycle completed frames concurrently, so a single blocking send call is not bounded by the ring size. After enough successful transmissions len_sum wraps into the negative range. In particular, 65536 65535-byte frames followed by a 65007-byte frame produce 0xfffffdef, or -EIOCBQUEUED. sock_sendmsg_nosec() treats that as an impossible return and triggers a BUG. On systems configured to panic on oops, a CAP_NET_RAW holder can panic the host kernel. Stop before the next frame would make the byte count exceed INT_MAX. The positive short count leaves that frame pending for a later send call. A two-thread TPACKET_V2 reproducer triggered the BUG on an arm64 build of Linux 7.3-rc4. With this change, the same reproducer returned the expected positive short count without an oops. Fixes: 69e3c75f4d54 ("net: TX_RING and packet mmap") Cc: stable@vger.kernel.org Signed-off-by: Ilkka Lappeteläinen <ilkka.lappetelainen@gmail.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260925124829.6595-1-ilkka.lappetelainen@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
10 daysnet/sched: sch_codel: match the no-drop threshold to the packet sizeJamal Hadi Salim1-1/+1
commit 6439461f1618 ("net/sched: sch_codel: clamp default mtu to avoid disabling CoDel") clamped q->params.mtu to [256, 1 << 20]. The value is the CoDel no-drop threshold (codel_impl.h "*backlog <= params->mtu"), so on a link whose maximum transmitted packet size is below 256 the floor extends CoDel's minimum-backlog exemption beyond one packet and delays drop or mark eligibility by several small packets. Keep the upper bound that guards the original overflow (psched_mtu() wrapping to ~2 GiB on a huge-MTU device) but drop the 256 floor, so the threshold tracks the real device packet size. Conditions to recreate the bug: attach a codel qdisc on a link whose MTU plus hard_header_len is below 256 (e.g. a CAN interface). At that MTU the no-drop threshold must equal the device MTU plus its hard-header length; before this patch it was forced to 256. Requires CAP_NET_ADMIN in a user namespace. Fixes: 6439461f1618 ("net/sched: sch_codel: clamp default mtu to avoid disabling CoDel") Reported-by: Sashiko (nipa) <sashiko-bot@kernel.org> Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260819143136.57350-1-jhs@mojatatu.com Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Link: https://patch.msgid.link/QDISC-34MS.v1.20260925165535@mojatatu.com.2 Signed-off-by: Paolo Abeni <pabeni@redhat.com>
10 daysnet/sched: fq_codel: match the no-drop threshold to the packet sizeJamal Hadi Salim1-2/+2
commit d9ebd8f9aa8b ("net/sched: fq_codel: clamp default quantum and mtu") clamped both q->quantum and q->cparams.mtu to [256, FQ_CODEL_QUANTUM_MAX]. The two fields mean different things: quantum is a DRR credit that wants the 256 floor, but cparams.mtu is the CoDel no-drop threshold (codel_impl.h "*backlog <= params->mtu"). On a link whose maximum transmitted packet size is below 256, the floor extends CoDel's minimum-backlog exemption beyond one packet and delays drop or mark eligibility by several small packets. Split the clamp. quantum keeps [256, FQ_CODEL_QUANTUM_MAX]; cparams.mtu tracks psched_mtu() (the device MTU plus its hard-header length) with only the upper bound that guards the original overflow (psched_mtu() wrapping to ~2 GiB on a huge-MTU device). Conditions to recreate the bug: attach an fq_codel qdisc on a link whose MTU plus hard_header_len is below 256 (e.g. a CAN interface). At that MTU the no-drop threshold must equal the device MTU plus its hard-header length; before this patch it was forced to 256. Basic Testing done: with dev->mtu=100 and hard_header_len=14, a return probe on fq_codel_init() observed cparams.mtu change from 256 to 114 Fixes: d9ebd8f9aa8b ("net/sched: fq_codel: clamp default quantum and mtu") Reported-by: Sashiko (nipa) <sashiko-bot@kernel.org> Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260819143136.57350-1-jhs@mojatatu.com Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Link: https://patch.msgid.link/QDISC-34MS.v1.20260925165535@mojatatu.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
11 daysnet: gso: limit recursive IP-in-IP segmentationZihan Xi3-0/+8
IP-in-IP GSO can re-enter inet_gso_segment() or ipv6_gso_segment() for each nested IP header. encap_level tracks header bytes, not callback depth, so a deep chain can exhaust the kernel stack. Making inet_gso_segment() stackable introduced unbounded IPv4 nesting; IPIP GSO/TSO later made the path reachable. The IPv6 stackable path was introduced separately and uses the same guard. Count IPv4 and IPv6 GSO handler entries in skb_gso_cb, initialized once per top-level GSO operation and preserved across GRE/UDP context changes. Use the existing IP_TUNNEL_RECURSION_LIMIT for both handlers. The first five entries pass, and the sixth returns -EINVAL before dispatching another GSO callback. Fixes: 3347c9602955 ("ipv4: gso: make inet_gso_segment() stackable") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Closes: https://lore.kernel.org/all/cover.1790157745.git.zihanx@nebusec.ai/ Assisted-by: LLM Co-developed-by: Luxing Yin <root@tr0jan.top> Signed-off-by: Luxing Yin <root@tr0jan.top> Signed-off-by: Zihan Xi <zihanx@nebusec.ai> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260924051521.32568-2-zihanx@nebusec.ai Signed-off-by: Paolo Abeni <pabeni@redhat.com>
11 daysnet: restrict SO_RESERVE_MEM to TCP sockets and cap max valueEric Dumazet1-2/+2
Commit 2bb2f5fb21b0 ("net: add new socket option SO_RESERVE_MEM") and commit d00c8ee31729 ("net: fix possible NULL deref in sock_reserve_memory") only checked sk_has_account(sk), which is true for both TCP and UDP sockets. However, SO_RESERVE_MEM and sk_unused_reserved_mem() are currently only supported by TCP: - On UDP sockets, sk->sk_forward_alloc is protected by sk->sk_receive_queue.lock, whereas sock_reserve_memory() and sock_release_reserved_memory() only acquire lock_sock(sk). Concurrent UDP packet reception/release and setsockopt(SO_RESERVE_MEM) corrupt sk_forward_alloc and memcg accounting. - udp_rmem_release() reclaims excess sk_forward_alloc without accounting for sk_unused_reserved_mem(sk). Restrict sock_reserve_memory() to TCP sockets (sk_is_tcp(sk)) for now. Supporting SO_RESERVE_MEM for UDP (acquiring sk_receive_queue.lock and honoring sk_unused_reserved_mem() in udp_rmem_release()) can be done in a future net-next series if needed. In addition, reject val > SZ_1G with -EINVAL in sk_setsockopt(SO_RESERVE_MEM). Without an upper bound, values near INT_MAX cause sk_mem_pages(delta) and (pages << PAGE_SHIFT) to overflow 32-bit signed int, corrupting sk->sk_forward_alloc and sk->sk_reserved_mem. Using a page-aligned cap (SZ_1G) ensures that the page-rounded sk->sk_reserved_mem reported by getsockopt(SO_RESERVE_MEM) can always be passed back to setsockopt(SO_RESERVE_MEM). Fixes: 2bb2f5fb21b0 ("net: add new socket option SO_RESERVE_MEM") Reported-by: Cai Xinchen <caixinchen1@huawei.com> Closes: https://lore.kernel.org/netdev/5a88421d-10ef-4fca-9acb-85a27a3c1173@huawei.com/ Signed-off-by: Eric Dumazet <edumazet@google.com> Reviewed-by: Wei Wang <weiwan@google.com> Link: https://patch.msgid.link/20260925135244.3715196-2-edumazet@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
11 daysnet: ethtool: let tsconfig reach a PHY-only timestamp providerNicolai Buchwitz2-15/+12
TSCONFIG_GET and TSCONFIG_SET reject a device that implements neither hwtstamp NDO, even when its PHY can serve the request. The ioctls they meant to replace handle it, so the two interfaces disagree on the same hardware and user space has to pick one. Accept the default timestamping PHY and an already installed provider on both sides. On the set side the test moves into ethnl_set_tsconfig() as the validate callback runs without rtnl. On the get side it stays ahead of ethnl_ops_begin(), so a device that can serve nothing keeps failing with EOPNOTSUPP and a dump still skips it rather than stopping there. A netdev provider needs ndo_hwtstamp_set to be programmed at all, so don't pick that source without it, and test for ndo_hwtstamp_get before calling it. Fixes: 6e9e2eed4f39 ("net: ethtool: Add support for tsconfig command to get/set hwtstamp config") Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de> Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Reviewed-by: Kory Maincent <kory.maincent@bootlin.com> Link: https://patch.msgid.link/20260925135237.3432266-3-nb@tipi-net.de Signed-off-by: Jakub Kicinski <kuba@kernel.org>
11 daysseg6: ensure packet data is writable before modifying SRH and IPv6 DAAndrea Mayer1-23/+117
advance_nextseg() modifies the SRH Segments Left field and the IPv6 destination address without ensuring the packet data is writable. seg6_next_csid_advance_arg() has the same problem when it advances the NEXT-C-SID argument in the destination address. The skb may be cloned, for example by an AF_PACKET socket receiving on the ingress device. advance_nextseg() and seg6_next_csid_advance_arg() then write into the packet data shared with the clone. A read from that socket can return the modified packet instead of the received one. The simplified path below shows this for advance_nextseg(): __netif_receive_skb_one_core __netif_receive_skb_core deliver_skb [orig: users=2, cloned=0] packet_rcv skb_clone clone queued to the AF_PACKET socket consume_skb(orig) [orig: users=1, cloned=1] ipv6_rcv ip6_rcv_core skb_share_check: no-op [orig: users=1] [...] input_action_end_core advance_nextseg writes into the data shared with the clone Call skb_ensure_writable() in advance_nextseg() and in seg6_next_csid_advance_arg() before they modify the packet data. skb_ensure_writable() may reallocate skb->head, which invalidates the pointers into the packet data taken before the call. advance_nextseg() now returns a valid SRH pointer, or NULL if skb_ensure_writable() fails. seg6_next_csid_advance_arg() takes the pointer to the destination address from the skb after skb_ensure_writable(). On failure, the callers drop the packet with SKB_DROP_REASON_NOMEM. Fixes: 140f04c33bbc ("ipv6: sr: implement several seg6local actions") Fixes: 848f3c0d4769 ("seg6: add NEXT-C-SID support for SRv6 End behavior") Signed-off-by: Andrea Mayer <andrea.mayer@uniroma2.it> Link: https://patch.msgid.link/20260925133807.32-1-andrea.mayer@uniroma2.it Signed-off-by: Jakub Kicinski <kuba@kernel.org>
11 daysnet: page_pool: fix use-after-free in page_pool_recycle_ring_bulk()Mina Almasry1-1/+1
With CONFIG_PAGE_POOL_STATS=y, page_pool_recycle_ring_bulk() updates the recycle stats after dropping the producer lock. If it just recycled the last inflight netmems of a pool being destroyed, page_pool_release() can pass its producer-lock barrier and free the pool before recycle_stat_add() runs. Commit fcc680a647ba7 ("page_pool: allow mixing PPs within one bulk") moved this stat update after the unlock. Commit 271683bb2cf32 ("page_pool: Fix use-after-free in page_pool_recycle_in_ring") later added the barrier, but only fixed page_pool_recycle_in_ring(). Move the update back under the lock. Fixes: fcc680a647ba7 ("page_pool: allow mixing PPs within one bulk") Link: https://lore.kernel.org/r/179028939171.2160803.706522228583664641@kernel.org Cc: Kaifeng Wang <kaifengw@google.com> Cc: Dong Chenchen <dongchenchen2@huawei.com> Signed-off-by: Mina Almasry <almasrymina@google.com> Reviewed-by: Simon Horman <horms@kernel.org> Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com> Link: https://patch.msgid.link/20260925144127.1445667-1-almasrymina@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
11 daystipc: prevent GCM nonce reuse on peer key changesJérémy Jean1-15/+5
TIPC can encrypt traffic between nodes using a different transmit key for each node. In this mode, the AES-GCM nonce for a packet sent to a known peer consists of a 32-bit prefix (a per-key salt XOR the peer's address) followed by a 64-bit counter. That counter is stored in the peer's RX crypto object. When the peer reports a change in which key it uses to receive packets, TIPC resets this counter. The sender can still be using the same TX key and salt, so subsequent packets reuse earlier nonces. This nonce reuse breaks confidentiality and exposes GCM's authentication key. This makes forgeries trivial: an attacker can exploit CTR malleability to alter captured ciphertexts and use the recovered authentication key to compute a valid tag for the modified ciphertext, under the same key and nonce. Use the TX key's existing aead->seqno counter instead. All encryptions using that key object share the same atomic counter, so concurrent encryptions get distinct nonce counter values. The counter survives key activation and peer reconnection, and peer key-status reports cannot reset it. This prevents those transitions from causing nonce reuse while the same TX key remains installed. The nonce format is unchanged, and receivers do not require consecutive counter values, so sharing the counter across peers remains compatible with existing receivers. A pre-existing check still invokes key revocation in the unlikely event that the counter wraps to zero. Fixes: fc1b6d6de220 ("tipc: introduce TIPC encryption & authentication") Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr> Reviewed-by: Tung Nguyen <tung.quang.nguyen@est.tech> Link: https://patch.msgid.link/20260924202105.3722778-1-Jeremy.Jean@oss.cyber.gouv.fr Signed-off-by: Jakub Kicinski <kuba@kernel.org>
11 daysMerge tag 'for-net-2026-09-28' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetoothJakub Kicinski6-22/+80
Luiz Augusto von Dentz says: ==================== bluetooth pull request for net: Core: - hci_core: Serialize ACL scheduling with channel deletion - hci_core: Serialize SCO and ISO scheduling with teardown - hci_core: Serialize fragmented ISO packet queueing - hci_core: Fix inquiry cache timestamps on 64-bit systems - hci_core: free the HCI ID if naming fails - hci_conn: Lock parent access during enhanced SCO setup - hci_sync: Fix inquiry cache use-after-free - hci_sync: don't drain cmd_sync backlog on unregister - RFCOMM: Fix initial port reference race - SMP: Serialize SMP remote OOB data access Drivers: - btintel: fix buffer over-read in btintel_hw_error() - btintel: validate DDC record lengths - btintel_pcie: fix plen overflow in btintel_pcie_recv_frame() - btintel_pcie: reject oversized TX packets in send_frame() * tag 'for-net-2026-09-28' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth: Bluetooth: hci_core: Serialize SCO and ISO scheduling with teardown Bluetooth: hci_core: Serialize ACL scheduling with channel deletion Bluetooth: Serialize SMP remote OOB data access Bluetooth: RFCOMM: Fix initial port reference race Bluetooth: hci_sync: Fix inquiry cache use-after-free Bluetooth: hci_sync: don't drain cmd_sync backlog on unregister Bluetooth: hci_core: Serialize fragmented ISO packet queueing Bluetooth: hci_conn: Lock parent access during enhanced SCO setup Bluetooth: btintel: validate DDC record lengths Bluetooth: hci_core: free the HCI ID if naming fails Bluetooth: hci_core: Fix inquiry cache timestamps on 64-bit systems Bluetooth: btintel_pcie: reject oversized TX packets in send_frame() Bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame() Bluetooth: btintel: fix buffer over-read in btintel_hw_error() ==================== Link: https://patch.msgid.link/20260928152919.942973-1-luiz.dentz@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
11 daysBluetooth: hci_core: Serialize SCO and ISO scheduling with teardownChengfeng Ye1-8/+18
hci_low_sent() selects a connection under RCU but drops the read lock before calculating its quota and returning it to the scheduler. The connection is then used without lifetime protection by hci_sched_iso() and hci_sched_sco(). After the TX worker drops the RCU read lock, hci_abort_conn_sync() on hdev->req_workqueue can remove the connection from the hash, complete synchronize_rcu(), purge its queues and release it. The TX worker on hdev->workqueue can then access the freed connection in hci_quote_sent() or while dequeuing packets and updating conn->sent. KASAN reported: BUG: KASAN: slab-use-after-free in hci_low_sent+0x730/0x840 Workqueue: hci0 hci_tx_work Call Trace: hci_low_sent+0x730/0x840 hci_sched_iso+0x25e/0x4d0 hci_tx_work+0x239/0xcb0 Allocated by task 93: __hci_conn_add+0x16f/0x1b40 hci_bind_bis+0x782/0x17b0 hci_connect_bis+0xa0/0x510 iso_sock_connect+0x589/0x1050 Freed by task 88: kfree+0x131/0x3c0 device_release+0xc8/0x240 kobject_put+0x14d/0x280 hci_conn_del+0x55a/0xe80 hci_disconnect_sync+0x156/0x180 hci_abort_conn_sync+0x3e7/0x940 hci_cmd_sync_work+0x13c/0x290 Hold hci_dev_lock() across connection selection and transmission in both SCO and ISO scheduling, serializing them with connection teardown. This also prevents queuing completion timestamps after the connection queues have been purged. Extending RCU across transmission would be unsafe because the transmit path can sleep. Keep the ISO timeout check outside the mutex since hci_link_tx_to() takes it itself. The preceding channel fix already holds this mutex in the ACL and LE schedulers, which call the SCO scheduler between packets. Move the SCO body to __hci_sched_sco(), assert that its caller holds the mutex, and use it directly from these locked paths. Keep a locking hci_sched_sco() wrapper for the direct calls from hci_tx_work(). This avoids recursively acquiring the device mutex while preserving the scheduling order. Remove the obsolete claim that connection removal disables TX. Fixes: bf4c63252490 ("Bluetooth: convert conn hash to RCU") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
11 daysBluetooth: hci_core: Serialize ACL scheduling with channel deletionChengfeng Ye1-0/+8
hci_chan_sent() selects a channel under RCU but drops the read lock before accessing chan->conn and returning the channel. The ACL and LE schedulers then use its packet queue and update its transmit counters without any protection against channel deletion. After TX selects a channel and releases RCU, a disconnect command timeout on the separate request workqueue can run hci_conn_failed() and l2cap_conn_del(). hci_chan_del() can then unlink the channel, complete synchronize_rcu(), purge its queue and free it before TX resumes. This causes use-after-free both in hci_chan_sent() and in its callers. KASAN reported: BUG: KASAN: slab-use-after-free in hci_chan_sent+0x892/0x9b0 Workqueue: hci0 hci_tx_work Call Trace: hci_chan_sent+0x892/0x9b0 hci_tx_work+0x5e6/0xb70 Allocated by task 91: hci_chan_create+0xe3/0x350 l2cap_conn_add.part.0+0x12/0xa30 l2cap_chan_connect+0x110d/0x1b60 l2cap_sock_connect+0x310/0x530 Freed by task 99: hci_chan_del+0x11f/0x170 l2cap_conn_del+0x4f1/0x800 l2cap_connect_cfm+0x88c/0xd30 hci_conn_failed+0x150/0x250 hci_abort_conn_sync+0x3e3/0x800 hci_cmd_sync_run+0x7e/0xc0 hci_abort_conn+0x105/0x1f0 disconnect_sync+0x157/0x290 hci_cmd_sync_work+0x13c/0x290 Hold the existing device mutex across channel selection and transmission in both schedulers to serialize them with channel teardown. Keep timeout handling outside the critical sections because hci_link_tx_to() acquires the same mutex. This also permits the transmit path to sleep, unlike extending the RCU read-side critical section across packet submission. Fixes: 3eff45eaf817 ("Bluetooth: convert tx_task to workqueue") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
11 daysBluetooth: Serialize SMP remote OOB data accessChengfeng Ye2-1/+13
build_pairing_cmd() looks up remote OOB data and copies its contents without holding hdev->lock, which serializes the list's writers. After SMP finds an entry, a concurrent management Remove Remote OOB Data command can unlink and free it before SMP reads its present flag or copies its random and confirmation values. Removal can also invalidate an entry while the lookup is still traversing the list. KASAN reported: BUG: KASAN: slab-use-after-free in build_pairing_cmd+0x948/0x9b0 Call Trace: build_pairing_cmd+0x948/0x9b0 smp_recv_cb+0x459f/0x8110 l2cap_recv_frame+0xf14/0x9190 l2cap_recv_acldata+0xa64/0xd40 hci_rx_work+0x4ca/0x730 Allocated by task 87: hci_add_remote_oob_data+0x11d/0x530 add_remote_oob_data+0x282/0x400 hci_sock_sendmsg+0x1033/0x1ea0 Freed by task 93: hci_remote_oob_data_clear+0x108/0x1c0 remove_remote_oob_data+0x198/0x220 hci_sock_sendmsg+0x1033/0x1ea0 Taking hdev->lock in build_pairing_cmd() would recurse for callers that already hold it and invert the device-to-L2CAP lock order on the receive path. Add a per-device remote_oob_lock instead, held across the SMP lookup and copies and by the add, remove and clear helpers. Cover initialization and in-place updates as well, so SMP cannot read partially initialized or updated OOB values. Release the mutex on allocation failure, preserving the existing error return. The new critical sections acquire no device, connection or channel locks. Writers retain their existing hdev->lock protection, which continues to serialize the other readers without changing their locking or behavior. Link: https://lore.kernel.org/r/00660cd3-7d71-13a4-f617-229e6defb701@gmail.com Fixes: 02b05bd8b0a6 ("Bluetooth: Set SMP OOB flag if OOB data is available") Cc: stable@vger.kernel.org Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
11 daysBluetooth: RFCOMM: Fix initial port reference raceChengfeng Ye1-4/+3
For RFCOMM_RELEASE_ONHUP devices, rfcomm_tty_install() and __rfcomm_release_dev() both use RFCOMM_TTY_OWNED to decide who drops the initial tty_port reference. The release path tests the bit separately from the install path setting it, so both can drop that reference. The synchronous hangup does not prevent this race: port->tty is not assigned until tty_port_open(), after installation. The following interleaving is possible: 1. Install and release each obtain a reference with rfcomm_dev_get(). 2. Release observes RFCOMM_TTY_OWNED clear. 3. Install sets RFCOMM_TTY_OWNED and drops the initial reference. 4. Release drops the same initial reference again. 5. TTY cleanup drops its reference and frees the device. 6. Release performs its final tty_port_put() on the freed port. KASAN reported: BUG: KASAN: slab-use-after-free in tty_port_put+0x22/0x190 Write of size 4 at addr ffff8881001d3d5c by task poc/92 Call Trace: tty_port_put+0x22/0x190 rfcomm_dev_ioctl+0x1d4/0x1930 sock_do_ioctl+0x110/0x260 sock_ioctl+0x380/0x590 __x64_sys_ioctl+0x134/0x1c0 Allocated by task 88: rfcomm_dev_ioctl+0x8f6/0x1930 sock_do_ioctl+0x110/0x260 sock_ioctl+0x380/0x590 __x64_sys_ioctl+0x134/0x1c0 Freed by task 65: kfree+0x131/0x3c0 rfcomm_dev_destruct+0x23b/0x2f0 release_one_tty+0xc1/0x370 process_one_work+0x661/0x1090 Use test_and_set_bit() in both paths so that only the caller that changes RFCOMM_TTY_OWNED from clear to set drops the initial reference. Each path keeps its own lookup reference until release or TTY cleanup, preserving the existing callback ordering and error handling. Fixes: 80ea73378af4 ("Bluetooth: Fix unreleased rfcomm_dev reference") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
11 daysBluetooth: hci_sync: Fix inquiry cache use-after-freeChengfeng Ye1-2/+11
The sync command worker holds hdev->req_lock, but inquiry-cache updates and flushes use hdev->lock. Both hci_acl_create_conn_sync() and hci_stop_discovery_sync() look up entries and read their fields without taking hdev->lock. After either lookup returns, a concurrent HCIINQUIRY ioctl can acquire hdev->lock and flush the cache, freeing the entry. The worker then reads the freed entry while preparing a create-connection or remote-name-cancel command. The list traversal also races with cache updates and removal. KASAN reported these accesses: BUG: KASAN: slab-use-after-free in hci_acl_create_conn_sync+0x5f1/0x650 Workqueue: hci0 hci_cmd_sync_work Call Trace: hci_acl_create_conn_sync+0x5f1/0x650 hci_cmd_sync_work+0x13c/0x290 Allocated by task 90: hci_inquiry_cache_update+0x3e6/0x7d0 hci_inquiry_result_evt+0x3cb/0x560 Freed by task 93: hci_inquiry_cache_flush+0x111/0x2b0 hci_inquiry+0x2f2/0x780 hci_sock_ioctl+0x269/0x5f0 BUG: KASAN: slab-use-after-free in hci_stop_discovery_sync+0x3b1/0x3c0 Workqueue: hci0 hci_cmd_sync_work Call Trace: hci_stop_discovery_sync+0x3b1/0x3c0 hci_cmd_sync_work+0x173/0x300 Allocated by task 86: hci_inquiry_cache_update+0x483/0x940 hci_inquiry_result_evt+0x3cb/0x560 Freed by task 91: hci_inquiry_cache_flush+0x13e/0x2f0 hci_inquiry+0x2f2/0x780 hci_sock_ioctl+0x269/0x5f0 Hold hdev->lock across each lookup and all reads from its result. Copy the remote address before unlocking so discovery cancellation does not retain a cache entry pointer. Release the lock before sending synchronous HCI commands, since their completion handlers may need the same lock. Fixes: cf75ad8b41d2 ("Bluetooth: hci_sync: Convert MGMT_SET_POWERED") Fixes: 45340097ce6e ("Bluetooth: hci_conn: Only do ACL connections sequentially") Cc: stable@vger.kernel.org Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
11 daysBluetooth: hci_sync: don't drain cmd_sync backlog on unregisterNguyen Ngoc Thang1-0/+6
hci_unregister_dev() disables cmd_work and cmd_timer, then calls hci_cmd_sync_clear(), whose cancel_work_sync() waits for hci_cmd_sync_work() to return. That worker dequeues and runs every entry on cmd_sync_work_list. With cmd_work disabled nothing reaches the controller, so each queued HCI command waits the full HCI_CMD_TIMEOUT (2s). The backlog has no bound. Userspace can keep queuing MGMT commands such as MGMT_OP_GET_CLOCK_INFO against a controller that doesn't answer. Closing /dev/vhci then blocks in vhci_release() for backlog * 2s: INFO: task syz-executor:5749 blocked for more than 143 seconds. cancel_work_sync hci_cmd_sync_clear hci_unregister_dev vhci_release In a local reproduction the backlog held more than 5000 entries, which comes to hours of hang. Stop the worker from taking new entries once HCI_UNREGISTER is set, and wake any request still waiting with -ENODEV before cancelling the work. The entries left over are destroyed with -ECANCELED by the existing sweep in hci_cmd_sync_clear(). They stay on the list until the work has stopped, so hci_cmd_sync_dequeue() and friends still see them. A callback already running may still issue another command, which delays unregister by at most one timeout per remaining command, not by the whole backlog. Fixes: 008ee9eb8a11 ("Bluetooth: hci_sync: Fix not processing all entries on cmd_sync_work") Reported-by: syzbot+217e3f1283cafe80586e@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=217e3f1283cafe80586e Signed-off-by: Nguyen Ngoc Thang <ngocthang2710.1999@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
11 daysBluetooth: hci_core: Serialize fragmented ISO packet queueingChengfeng Ye1-0/+4
The fragmented path in hci_queue_iso() calls __skb_queue_tail() without holding conn->data_q.lock. The socket lock serializes senders, but the transmit worker can concurrently remove packets from this queue using skb_dequeue(). After an insertion saves the old tail pointer, hci_sched_iso() can dequeue that packet and pass it to the driver. The driver can free the packet before the insertion resumes and writes to the old tail's next pointer, causing a use-after-free. Concurrent updates can also corrupt the queue. KASAN reported: BUG: KASAN: slab-use-after-free in hci_send_iso+0xc7d/0xe10 Call Trace: hci_send_iso+0xc7d/0xe10 iso_sock_sendmsg+0x710/0x900 __sys_sendto+0x34a/0x3a0 Allocated by task 86: __alloc_skb+0xdd/0x820 alloc_skb_with_frags+0x7e/0x770 sock_alloc_send_pskb+0x67a/0x820 bt_skb_sendmsg.constprop.0+0xc0/0x6c0 iso_sock_sendmsg+0x44e/0x900 Freed by task 90: kmem_cache_free+0xcb/0x3d0 vhci_read+0x33f/0x4d0 vfs_read+0x177/0xa20 Hold the queue lock across the entire fragment batch, as hci_queue_acl() does, to serialize insertion against dequeue while preserving fragment ordering. Fixes: 26afbd826ee3 ("Bluetooth: Add initial implementation of CIS connections") Cc: stable@vger.kernel.org Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
11 daysBluetooth: hci_conn: Lock parent access during enhanced SCO setupChengfeng Ye1-5/+13
Bluetooth: hci_conn: Lock parent access during enhanced SCO setup hci_enhanced_setup_sync() runs on the request workqueue without the hci_dev_lock held by its caller when setup was queued. Its CVSD capability check and find_next_esco_param() dereference conn->parent while the receive workqueue can unlink and release that parent. The following interleaving can cause a use-after-free: hci_enhanced_setup_sync() hci_disconn_complete_evt() load conn->parent hci_dev_lock() hci_conn_del(ACL parent) unlink SCO child drop link's parent reference clear child->parent release ACL parent bt_link_release() kfree(parent) hci_dev_unlock() read parent->features[0][3] The reference held for the queued SCO child does not keep its ACL parent alive after unlinking. Commit 42de40abe25d ("Bluetooth: hci_conn: fix the SCO setup context lifetime") protects the child stored in the queued context, but leaves these parent accesses unprotected. With a 40 ms diagnostic delay after loading conn->parent, an instrumented kernel based on fd179f8a05be, which already contains 42de40abe25d, reported: BUG: KASAN: slab-use-after-free in hci_enhanced_setup_sync+0xda5/0xdf0 Read of size 1 at addr ffff888102204047 by task kworker/u17:0/93 Call Trace: hci_enhanced_setup_sync+0xda5/0xdf0 hci_cmd_sync_work+0x13c/0x290 process_one_work+0x6b4/0x10e0 Allocated by task 92: __hci_conn_add+0x304/0x1df0 hci_connect_acl+0x349/0x3e0 hci_connect_sco+0x3b/0x9a0 sco_sock_connect+0x475/0xca0 Freed by task 94: kfree+0x121/0x3c0 bt_link_release+0x79/0xa0 device_release+0xc8/0x240 kobject_put+0x14d/0x280 hci_conn_del+0x524/0xe30 hci_disconn_complete_evt+0x403/0x8c0 hci_event_packet+0x71b/0xb20 hci_rx_work+0x293/0x730 The accessed address is 71 bytes into the freed ACL parent, at its features[0][3] byte; the queued SCO child is a different object. The diagnostic preserves the loaded parent across the delay, matching the unmodified compiled capability check, and does not change the parent references or teardown path. Hold hci_dev_lock() across the codec switch, including every call to find_next_esco_param(), and release it on all selection errors. Keep configure_datapath_sync() outside the critical section because it waits for HCI events. Preserve parameter selection and existing return values. Fixes: e07a06b4eb41 ("Bluetooth: Convert SCO configure_datapath to hci_sync") Cc: stable@vger.kernel.org Assisted-by: GPT-6-Astra Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
12 dayswifi: mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()Cen Zhang (Microsoft)2-7/+7
ieee80211_monitor_select_queue() validates the skb length using skb->len before dereferencing pointers into skb->data to access the 802.11 header. However, skb->len includes data in both the linear head buffer and non-linear fragments/pages. When AF_PACKET sends a packet large enough to become non-linear, the 802.11 header at skb->data + len_rthdr may extend past the linear head buffer even though skb->len appears sufficient. This leads to a slab-out-of-bounds read when accessing hdr->frame_control, as the kernel reads beyond the allocated skb head buffer: BUG: KASAN: slab-out-of-bounds in ieee80211_monitor_select_queue+0x1ed/0x220 syzbot has hit the same issues. Fix this by checking skb_headlen(skb) (which gives the length of the linear data region) instead of skb->len before dereferencing skb->data pointers. This ensures the required bytes are actually present in the linear portion of the skb. Apply the same fix to ieee80211_validate_radiotap_len() which has the same class of bug: it uses skb->len to validate accesses to skb->data, and to the corresponding checks in ieee80211_monitor_start_xmit(): for drivers advertising NETIF_F_SG the skb is not linearized before ndo_start_xmit, so the 802.11 header can be read past the linear buffer there as well. Fixes: cf0277e714a0 ("mac80211: fix skb buffering issue") Fixes: 9b8a74e3482f ("[MAC80211]: Improve sanity checks on injected packets") Reported-by: AutonomousCodeSecurity@microsoft.com Reported-by: syzbot+610e40369bc02181bad0@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=610e40369bc02181bad0 Reported-by: syzbot+878643e0580bc580f883@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=878643e0580bc580f883 Assisted-by: GitHub-Copilot:claude-opus-4.6 Signed-off-by: Cen Zhang (Microsoft) <cenzhang@linux.microsoft.com> Reviewed-by: Francis Perron <francis@akrites.dev> Link: https://patch.msgid.link/20260925024408.32143-1-cenzhang@linux.microsoft.com Signed-off-by: Johannes Berg <johannes.berg@intel.com>
12 dayswifi: mac80211: reject invalid 320 MHz CSA bandwidthRuide Cao1-2/+2
The HT/VHT channel definition validator can receive a 320 MHz bandwidth indication from a received CSA frame on a non-6 GHz link. It warns for this unsupported width but continues with an uninitialized vht_operation.chan_width, which is then read by ieee80211_chandef_vht_oper(). With panic_on_warn enabled, this lets a received frame panic the kernel. Reject the channel definition before entering the VHT operation conversion. This preserves the existing CSA fallback while avoiding both the warning and the uninitialized read. Fixes: 21c3f8f95554 ("wifi: mac80211: refactor STA CSA parsing flows") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Assisted-by: LLM Signed-off-by: Ruide Cao <caoruide123@gmail.com> Signed-off-by: Ren Wei <weir@nebusec.ai> Link: https://patch.msgid.link/bd68e360e06135c750397cb6be003a01a834a179.1787222001.git.vega.cover-letter@nebusec.ai Signed-off-by: Johannes Berg <johannes.berg@intel.com>
12 dayswifi: mac80211: set info->band for 802.3 encap offload framesFelix Fietkau1-0/+11
Unlike the 802.11 TX paths, ieee80211_8023_xmit() does not set info->band from the channel context on non-MLD interfaces. It stays at 0, so ieee80211_get_tx_rates() uses the wrong band for these frames. Set it the same way as ieee80211_build_hdr(): drop the frame if there is no channel context, and leave the band at 0 on MLD interfaces. Reported-by: Andrea Pesaresi <andreapesaresi82@gmail.com> Fixes: 50ff477a8639 ("mac80211: add 802.11 encapsulation offloading support") Cc: stable@vger.kernel.org Signed-off-by: Felix Fietkau <nbd@nbd.name> Link: https://patch.msgid.link/20260926120216.4054712-1-nbd@nbd.name Signed-off-by: Johannes Berg <johannes.berg@intel.com>
12 daysnetfilter: ipset: do not update comments from kernel-side addsFlorian Westphal1-1/+1
'Fixes' commit stopped calling ip_set_init_comment() for hash types from kernel-side-adds (xtables .. -j SET). ip_set_init_comment() says: "The kadt functions don't use the comment extensions in any way." But bitmap set type calls the function from kadt cb too. While this appears to be safe (serialized via the set spinlock), it seems better to not call the init function either, least of all to keep behaviour consistent. ip_set_list calls ip_set_init_comment() only from uadt cb, it can be kept as-is. This was triggered by yet another LLM review, hinting that the existing rcu_dereference_protected() cannot be downgraded to only check if the nfnl mutex is held. Fixes: f30415929be8 ("netfilter: ipset: do not update comments from kernel-side hash adds") Signed-off-by: Florian Westphal <fw@strlen.de> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
14 daysipv6: fix prefix route expiry in modify_prefix_route()Qishuai Liu1-1/+1
modify_prefix_route() is given the lifetime in clock_t relative to now, but fib6_set_expires() wants an absolute jiffies value. So when a permanent address is changed to a finite valid_lft, the prefix route ends up already expired and GC removes it. Steps to reproduce: ip link add dummy9 type dummy ip link set dummy9 up ip -6 addr add 2001:db8:9::1/64 dev dummy9 ip -6 addr change 2001:db8:9::1/64 dev dummy9 valid_lft 3600 preferred_lft 3600 ip -6 route show dev dummy9 # expires is negative Fixes: 8308f3ff1753 ("net/ipv6: Add support for specifying metric of connected routes") Signed-off-by: Qishuai Liu <lqs@lqs.me> Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/20260923132048.3272878-1-lqs@lqs.me Signed-off-by: Jakub Kicinski <kuba@kernel.org>
14 daysnet/sched: cls_api: reclaim an empty proto on the error pathJamal Hadi Salim1-1/+14
Two racing tc filter add requests on the same chain/prio of an unlocked classifier both run change() on the shared proto and both can fail: the winner's tcf_chain_tp_delete_empty() attempt gives up because the loser's handle is still in the idr, and the loser's error path drops only its own reference without a second reclamation attempt. The empty proto stays linked in the chain, holding the chain reference, a block reference and the classifier module reference until the chain or block is torn down. Reclaim the proto on the error path of any failed request that holds a proto reference. The reclamation is emptiness-gated: tcf_chain_tp_delete_empty() unlinks the proto only when delete_empty() admits it is empty, so a live shared proto is never unlinked. A proto the request created is reclaimed unconditionally - it is the only owner, so marking it for deletion is safe even without a delete_empty callback. Classifiers without one (the check marks the proto unconditionally) are rtnl-serialized, so the raced window this guard closes cannot arise for them. This is a follow-up to commit d4e359b3608a ("net/sched: cls_api: fix teardown of an adopted proto on insert-race loss"), which stopped the loser of the insert race from unlinking the winner's live proto but left the empty-proto residual in place. Conditions to recreate: - CONFIG_NET_CLS_FLOWER=y; veth pair - tc qdisc add dev veth0 ingress - two concurrent `tc filter add dev veth0 ingress protocol ip pref 1 flower skip_sw ... action drop` (both fail in fl_hw_replace_filter after publishing their handle in the idr); repeat in a loop - an empty flower tp stays linked after both requests fail; visible as a bare `filter protocol ip pref 1 flower chain 0` header in `tc filter show` with no filter entries - CAP_NET_ADMIN (namespace-local via unshare -Urn suffices) Fixes: 8b64678e0af8 ("net: sched: refactor tp insert/delete for concurrent execution") Reported-by: Sashiko (nipa) <sashiko-bot@kernel.org> Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260805134049.927864-1-victor@mojatatu.com Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/QDISC-JCOT.v1.20260910090924@mojatatu.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>