| Age | Commit message (Collapse) | Author | Files | Lines |
|
tcp_v4_syn_recv_sock() transfers ownership of ireq->ireq_opt to the
child socket (newinet->inet_opt) without copying it.
Another cpu can concurrently process a retransmitted SYN for the same
request socket, and send a SYNACK from tcp_check_req().
tcp_v4_send_synack() and inet_csk_route_req() read ireq->ireq_opt
under rcu_read_lock() only, and ip_build_and_send_pkt() and
ip_options_build() then read opt->optlen twice.
Note that the SYNACK timer itself is not an issue: it holds a
reference on its request socket, and inet_csk_reqsk_queue_drop()
calls timer_delete_sync() before the child can be freed.
Since commit 079096f103fa ("tcp/dccp: install syn_recv requests
into ehash table"), request sockets are processed without holding
the listener lock, so nothing prevents the child socket from being
freed while the SYNACK is still being built. TCP child sockets do
not have SOCK_RCU_FREE, and inet_sock_destruct() frees inet_opt
with a plain kfree(), leading to a use-after-free in
ip_options_build().
A similar issue exists with request socket migration
(net.ipv4.tcp_migrate_req=1, or a BPF_SK_REUSEPORT_SELECT_OR_MIGRATE
program). reqsk_timer_handler() clones the request socket with
inet_reqsk_clone(), so that the old request socket and its clone
share the same ireq_opt, then reqsk_migrate_reset() clears the
pointer in the old one. Another cpu holding a reference on the old
request socket can still be using these options (sending a SYNACK,
or creating a child in tcp_v4_syn_recv_sock()) when the clone is
freed, and tcp_v4_reqsk_destructor() also uses a plain kfree().
Readers already use RCU, and other paths replacing inet_opt
(do_ip_setsockopt(), cipso_v4_sock_setattr()...) already use
kfree_rcu(). Use kfree_rcu() in inet_sock_destruct() and
tcp_v4_reqsk_destructor() as well.
IPv6 is not affected by the first issue, because tcp_v6_syn_recv_sock()
duplicates the options. tcp_v6_reqsk_destructor() has the same
migration issue with ipv6_opt, which is only set by CALIPSO. This
will be addressed in a separate patch.
Fixes: 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table")
Fixes: c905dee62232 ("tcp: Migrate TCP_NEW_SYN_RECV requests at retransmitting SYN+ACKs.")
Reported-by: Xinyang Ge <xinyang@anthropic.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20261001221253.2964024-1-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
ethtool refuses to enable tcp-data-split on a device running
a single-buffer XDP program (which can't deal with the payload
landing in a separate buffer). dev_xdp_sb_prog_count() only looks
at xdp_state[].prog though, ignoring bpf_link integration.
Get the prog pointer from dev_xdp_prog(), which will consult both.
bpf_xdp_link_update() needs a bit of a touch-up, too, since
ethtool calls dev_xdp_sb_prog_count() with just the instance
lock (for ops-locked devices) - we now have to make sure the
link updates happen under that lock.
Spotted by AI while reviewing the XDP propagation series,
verified and tested on netdevsim with a C test along these lines:
link = bpf_program__attach_xdp(prog, ifindex);
system("ethtool -G eth0 tcp-data-split on")
Fixes: 197258f0ef68 ("net: ethtool: add hds_config member in ethtool_netdev_state")
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Link: https://patch.msgid.link/20261001012559.297751-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A production compute node carrying a few thousand datapath flows hit a
soft lockup inside a single netlink flow dump and panicked.
ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it
emits, and that releases stats->lock with spin_unlock_bh() once per
CPU that has touched the flow. Every release is a local_bh_enable(),
and each one runs the pending softirq backlog in the dumping thread's
own context.
The skb bounds the flows one callback emits, but not the softirq work
it absorbs. On a CPU that carries the box's packet load the backlog
refills as fast as it drains, so the dumping thread becomes that CPU's
softirq engine. It never sleeps and it has no reschedule point, so
under voluntary preemption nothing can take the CPU away from it:
neither the ksoftirqd the kernel woke to take the work over, nor the
stopper thread the softlockup detector dispatches to refresh its
timestamp.
Hold BH off across the whole callback instead, the way
ctnetlink_dump_table() does, so the nested spin_unlock_bh() stop
draining softirqs. The loop already runs under rcu_read_lock() and
cannot sleep. What it gives up is preemption under CONFIG_PREEMPT,
since a BH-off region is not preemptible outside PREEMPT_RT. The
region stays short: the skb caps the flows one callback emits, and
empty buckets cost no skb space but are each visited once per dump, as
the cursor only moves forward. The table grows on insert and shrinks
only on flush, so the walk is bounded by the largest flow count the
datapath has held. The softirq backlog the callback used to absorb has
no bound at all.
Fixes: 63e7959c4b9b ("openvswitch: Per NUMA node flow stats.")
Cc: stable@vger.kernel.org
Signed-off-by: Denis V. Lunev <den@openvz.org>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
br_multicast_sg_del_exclude_ports() removes automatically added STAR_EXCL
port groups by calling br_multicast_del_pg(). For S,G entries, that
deletion calls back into br_multicast_sg_del_exclude_ports(), so a large
number of STAR_EXCL ports consumes one kernel stack frame per port and can
hit the stack guard page.
Keep the complete port-group deletion path for cleanup, but suppress the
S,G exclude-port cleanup while that helper is already walking the same
entry. Normal deletion callers retain the existing behavior.
Fixes: 8266a0491e92 ("net: bridge: mcast: handle port group filter modes")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Zixuan Chai <petalzu987@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
After copying a readable prefix, receive_fallback_to_copy() asks
tcp_zerocopy_set_hint_for_skb() where page mapping can resume. If the
next skb is unreadable, find_next_mappable_frag() passes its net_iov
fragment to can_map_frag().
skb_frag_page() returns NULL for a net_iov, but can_map_frag()
dereferences it in PageCompound(). A v7.2 KASAN run on a connected TCP
socket with 64 readable bytes followed by a 4096-byte NET_IOV_DMABUF
fragment reported:
BUG: KASAN: null-ptr-deref in can_map_frag
tcp_zerocopy_receive -> can_map_frag
Kernel panic - not syncing: KASAN: panic_on_warn set
The diagnostic inserted the net_iov directly because the test host has
no devmem-capable NIC. Hardware end-to-end reachability remains untested
and requires CONFIG_NET_DEVMEM plus a supported DMA-buf-bound RX queue.
Reject all net_iov fragments before skb_frag_page(). This covers both
DMABUF and IOURING net_iov types while leaving page-backed checks
unchanged. With the guard, the same queue copied the readable prefix,
returned a 4096-byte skip hint, and completed without a fault.
Fixes: 9f6b619edf2e ("net: support non paged skb frags")
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Reviewed-by: Mina Almasry <almasrymina@google.com>
Link: https://patch.msgid.link/20260930102524.1659847-1-4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
tipc_topsrv_stop() closed subscriber connections while the topology
server's send and receive workqueues were still running. Socket
callbacks and subscription events could then queue more work, and
in-flight send/recv work could drop the last connection reference
during the conn_idr walk.
That race produced several teardown failures: refcount_t addition on
0 from conn_get() on a connection whose release was blocked on
idr_lock, a subsequent use-after-free in tipc_conn_close(),
queue_work() on an already destroyed workqueue from the listener
data-ready callback, and an RCU stall in tipc_topsrv_exit_net()
while the walk spun under idr_lock.
Clear srv->listener under idr_lock so it acts as a shutdown flag,
skip queue_work() once it is NULL, destroy the workqueues to flush
in-flight work, and only then close the remaining connections.
Refuse tipc_conn_lookup() after that flag is cleared, and drop
idr_lock when the teardown walk finds no connection, so an
in-flight subscription event cannot pin idr_in_use while the
walk holds the lock.
v3 supersedes the narrower in-thread diff Tung Quang Nguyen
posted on 2026-09-23. It keeps his teardown order and adds the
lookup refusal and the empty-idr unlock.
Fixes: c5fa7b3cf3cb ("tipc: introduce new TIPC server infrastructure")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <weir@nebusec.ai>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
packet_rcv_spkt() restores the link layer header with
skb_push(skb, skb->data - skb_mac_header(skb));
That subtraction is only meaningful when the device actually has a link
layer header. packet_rcv() and tpacket_rcv() both wrap it in
dev_has_header(), which is also the predicate the block comment at the
top of the file states the restore in terms of; packet_rcv_spkt() does
not. commit d549699048b4 ("net/packet: fix packet receive on L3
devices without visible hard header") introduced the helper and changed
the two call sites, and this one stayed behind.
A device without a visible ll header can leave skb->mac_header at the
0xFFFF sentinel that __alloc_skb() initialises it to. A CAN skb does:
init_can_skb() sets pkt_type and ip_summed but does not reset the
headers, and commit 9f10374bb024 ("can: remove private CAN skb
headroom infrastructure") dropped the skb_reset_*_header() calls that
used to be there. skb_mac_header() is then 0xFFFF, the length becomes
a large negative number and skb_push() reports it through
skb_under_panic() -- from softirq context, so it is a full system panic
even with panic_on_oops=0:
skbuff: skb_under_panic: text:ffffffff8bd21bc1 len:-65455 put:-65471 head:... data:... tail:0x50 end:0x180 dev:can0
kernel BUG at net/core/skbuff.c:214!
RIP: 0010:skb_panic+0x50/0x60
Call Trace:
<IRQ>
skb_push+0x38/0x40
packet_rcv_spkt+0xe1/0x170
__netif_receive_skb_core.constprop.0+0x7e8/0xd30
...
Kernel panic - not syncing: Fatal exception in interrupt
The socket type is reachable: packet_create() accepts SOCK_PACKET
alongside SOCK_RAW and SOCK_DGRAM behind the same CAP_NET_RAW check,
and neither the socket length nor a capability check keeps it away
from a CAN interface.
The missing skb_reset_*_header() calls in init_can_skb() are a
regression in their own right and are being fixed separately, but a
packet socket should not turn a link layer that did not initialise its
mac header into a kernel panic. Guard the push the way the other two
receive paths do.
Tested on v7.3-rc5 under QEMU with a slcan device on a pty, which is
the driver RX path: vcan does not reproduce it, because can_send()
resets the headers on the way out. One SOCK_PACKET socket bound to
can0 and one frame written into the line discipline panics an
unpatched kernel with the trace above; the same image with this patch
prints no panic and powers off normally. Both kernels are this tree,
defconfig plus CONFIG_CAN_SLCAN=y, differing only in this hunk.
Fixes: d549699048b4 ("net/packet: fix packet receive on L3 devices without visible hard header")
Cc: stable@vger.kernel.org
Signed-off-by: Quchaosheng <quchaosheng000406@163.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260928113108.2127215-1-quchaosheng000406@163.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
This is a follow-up to commit 9c572a83037a ("net/sched: fix potential
stack infoleak in em_text_dump()"), which zero-initialised the dump-side
struct. Sashiko pointed at a pre-existing bug that exposed the change-side
input-validation gap open.
em_text_change() casts the raw netlink attribute payload to struct
tcf_em_text and passes conf->algo, a 16-byte array with no enforced NUL
terminator, to textsearch_prepare(). On the TS_AUTOLOAD retry
textsearch_prepare() calls request_module("ts_%s", algo); vsnprintf()
walks %s until it finds a NUL, so an algo[] with no NUL reads past the
array into the adjacent struct fields and payload and feeds those bytes
to the usermode helper command line. This is an out-of-bounds read and
a kernel memory disclosure.
Reject the attribute when no NUL exists within TC_EM_TEXT_ALGOSIZ bytes,
matching the check xt_string has had since 3ab720881b6e.
Conditions to recreate the bug:
- CAP_NET_ADMIN on the target netns (namespace-local via unshare -Urn
is enough)
- install clsact on an interface, then add a basic filter with one text
ematch whose algo[] is 16 bytes with no NUL (e.g. "A"*16) and
pattern_len at least 1
- the add fails with -ENOENT and the kernel invokes modprobe with a
module name made of "ts_" plus the 16 bytes and the adjacent payload
Fixes: d675c989ed2d4 ("[PKT_SCHED]: Packet classification based on textsearch (ematch)")
Reported-by: Sashiko (gemini) <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260918133953.12494-1-bernard.ladenthin%40gmail.com
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2DIU.v1.20260929183844@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
tipc_node_find_by_name() returns a node after dropping its RCU read lock
without taking a reference. The LINK_SET, LINK_GET and LINK_RESET_STATS
handlers then lock and access the node, racing with timer-driven cleanup of
a down peer. Generic netlink serialization does not exclude the node
timer.
The following interleaving can leave a handler using a freed node:
CPU 0: find the node under RCU and release the node read lock
CPU 1: tipc_node_timeout() clears the links and unlinks the down node
CPU 1: drop the list and timer references, queuing tipc_node_free()
CPU 0: leave the RCU read-side critical section
CPU 1: complete the grace period and free the node
CPU 0: acquire the node lock through the stale pointer
LINK_SET also uses the node's media address after releasing the node lock,
when passing queued packets to tipc_bearer_xmit().
KASAN on v7.3-rc5 reported:
BUG: KASAN: slab-use-after-free in _raw_read_lock_bh
Write of size 4 at addr ffff88807e06f008 by task poc/92
Call Trace:
_raw_read_lock_bh kernel/locking/spinlock.c:287
tipc_nl_node_set_link net/tipc/node.c:2475
genl_family_rcv_msg_doit net/netlink/genetlink.c:1114
genl_rcv_msg net/netlink/genetlink.c:1209
netlink_rcv_skb net/netlink/af_netlink.c:2575
Allocated by task 0:
tipc_node_create net/tipc/node.c:539
tipc_node_check_dest net/tipc/node.c:1196
tipc_disc_rcv net/tipc/discover.c:252
Freed by task 92:
kfree mm/slub.c:6801
rcu_core kernel/rcu/tree.c:2919
Last potentially related work creation:
__call_rcu_common.constprop.0 kernel/rcu/tree.c:3181
tipc_node_timeout net/tipc/node.c:814
Acquire a reference to the selected node with kref_get_unless_zero() before
leaving RCU, returning NULL if the node has already been released. Release
that reference on every caller exit after the last node access, including
transmission in LINK_SET. Keep the existing link lookup order and locking
so concurrent link removal still takes the existing error paths.
Fixes: 6a939f365bdb ("tipc: Auto removal of peer down node instance")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Link: https://patch.msgid.link/20260928161246.1895273-1-nicoyip.dev@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:
1) Expand existing ipset fix for bitmap sets to disallow comments
updates from kernel-side adds, from Florian Westphal.
2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
flowtable cannot ever be removed, from Aohan Mei.
3) nft_rbtree GC should collect end elements that contained in
this transaction batch, new or deleted elements are never
expired. From Weiming Shi.
4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_*
values, from Fernando F. Mancera.
5) Flowtable GC must skip flows that are pending hardware updates,
generalize the PENDING flag and use it to inhibit GC.
6) Restore flowtable with ieee80211 which broke due to a relatively
recent commit, which was pulled in by -stable, causing a regression
in 6.18 kernels.
And the following IPVS fixes:
1) Fix accounting of cache entries in IPVS LBLC for destinations,
which eventually fills up the table and trigger recurrent
resizing, from Julian Anastasov.
2) Limit IPVS cache growth for LBLCR and LBLC schedulers,
from Zhiling Zou.
3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections,
do not allow to use it with templates. Also from Julian.
4) Sanitize flags in IPVS sync messages received in the backup.
From Julian Anastasov.
netfilter pull request 26-09-30
* tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
netfilter: ipset: do not update comments from kernel-side adds
====================
Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Since commit d58e468b1112 ("flow_dissector: implements flow dissector
BPF hook") __skb_flow_dissect() needs a net pointer, either from
skb->dev, skb->sk, or since commit 3cbf4ffba5ee ("net: plumb network
namespace into __skb_flow_dissect") a caller provided pointer.
syzbot was able to reach seg6_make_flowlabel() with an skb having
neither skb->dev nor skb->sk set: a TIPC UDP bearer sends a discovery
message through an IPv4 route using seg6 encap, while
net.ipv6.seg6_flowlabel is set to 1.
seg6_make_flowlabel() already has a net pointer, use skb_get_hash_net().
WARNING: net/core/flow_dissector.c:1131 at __skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126, CPU#0: syz.0.17/4930
Call trace:
__skb_flow_dissect+0x910/0x5368 net/core/flow_dissector.c:1126 (P)
__skb_get_hash_net+0xe0/0x29c net/core/flow_dissector.c:1903
skb_get_hash include/linux/skbuff.h:1663 [inline]
seg6_make_flowlabel+0xcc/0x1ec net/ipv6/seg6_iptunnel.c:132
__seg6_do_srh_encap+0x320/0xbf4 net/ipv6/seg6_iptunnel.c:161
seg6_do_srh+0x4b4/0xa44 net/ipv6/seg6_iptunnel.c:431
seg6_output_core+0x164/0x688 net/ipv6/seg6_iptunnel.c:681
seg6_output+0x44/0x1ac net/ipv6/seg6_iptunnel.c:744
lwtunnel_output+0x3d0/0x664 net/core/lwtunnel.c:356
dst_output include/net/dst.h:470 [inline]
ip_local_out+0x110/0x148 net/ipv4/ip_output.c:131
iptunnel_xmit+0x50c/0xd38 net/ipv4/ip_tunnel_core.c:97
udp_tunnel_xmit_skb+0x220/0x348 net/ipv4/udp_tunnel_core.c:187
tipc_udp_xmit+0x75c/0x9c0 net/tipc/udp_media.c:202
tipc_udp_send_msg+0x214/0x374 net/tipc/udp_media.c:274
tipc_bearer_xmit_skb+0x260/0x3b0 net/tipc/bearer.c:576
tipc_enable_bearer net/tipc/bearer.c:366 [inline]
__tipc_nl_bearer_enable+0xc90/0xfb0 net/tipc/bearer.c:1048
tipc_nl_bearer_enable+0x2c/0x48 net/tipc/bearer.c:1057
genl_family_rcv_msg_doit+0x1e4/0x2d4 net/netlink/genetlink.c:1114
Fixes: d58e468b1112 ("flow_dissector: implements flow dissector BPF hook")
Reported-by: syzbot+9408fbe0e6452a12e9ab@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6abac2eb.3654fce1.1bec97.0000.GAE@google.com/
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260928194524.3617299-1-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
sctp_wait_for_connect() drops the socket lock while it sleeps. An
out-of-the-blue ABORT can then be processed from the socket backlog and
unlink the association. If a concurrent shutdown(fd, SHUT_RD) sets
RCV_SHUTDOWN, the waiter breaks with err == 0 before checking
asoc->base.dead. Its final sctp_association_put() can then free the
association, leaving sctp_sendmsg_to_asoc() to continue with a dangling
pointer.
Check RCV_SHUTDOWN along with the wait error in sctp_sendmsg_to_asoc()
before using the association again. The check only accesses the socket,
so it needs no additional association reference. Return the existing
-ESRCH so that sctp_sendmsg() skips freeing a new association that may
already have been destroyed.
Keep sctp_wait_for_connect() unchanged to preserve its behavior for the
connect() caller.
Fixes: 668c9beb9020 ("sctp: implement assign_number for sctp_stream_interleave")
Cc: stable@vger.kernel.org
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Signed-off-by: Jun Yang <junvyyang@tencent.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260926095606.68601-1-juny24602@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Johannes Berg says:
====================
Still more fixes coming in, notably:
- ath11k: avoid running out of stations on HW restart
- mac80211:
- drop too large fragmented MPDUs
- mesh path handling fixes
- validation improvements
- reject CSA with bad 320 MHz bandwidth
- cfg80211: fix RTS for single radio devices
* tag 'wireless-2026-09-30' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (27 commits)
wifi: mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()
wifi: mac80211: reject invalid 320 MHz CSA bandwidth
wifi: mac80211: set info->band for 802.3 encap offload frames
wifi: mac80211: prevent AP VLAN tx from other interfaces
wifi: cfg80211: preserve hidden-group beacon IE ownership
wifi: mac80211: keep fallback association elements alive
wifi: mac80211: shut down RX BA session timer on teardown
wifi: mac80211: validate TX status rate metadata
wifi: ath9k_htc: bound TX aggregation to MAX_TX_BUF_SIZE
wifi: ath9k: reject short WMI command responses
wifi: ath9k: Clean up device initialisation guards
wifi: ath11k: reset ar->num_stations on hardware start
wifi: cfg80211: fix RTS threshold setting for single-radio PHY
wifi: mac80211: handle empty FILS association request payload
wifi: mac80211: minstrel_ht: validate fixed rate index
wifi: p54: validate firmware record lengths
wifi: mac80211: fix mesh fast xmit path deletion UAF
wifi: mac80211: drop oversized fragments to avoid extra_len overflow
wifi: wlcore: Fix runtime PM leak in wlcore_remove()
wifi: mac80211: drain PS delivery work during station teardown
...
====================
Link: https://patch.msgid.link/20260930124440.224799-3-johannes@sipsolutions.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward
path cannot be discovered"), there was a fallback to set up a forward
path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used.
Such fallback was used by commit d787a3e38f01 ("mac80211: add support
for .ndo_fill_forward_path").
One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable
forward path discovery. However, this is only used internally by drivers
to retrieve mtk_wdma information to set up hardware offload. Felix
decided to use the .fill_forward_path interface for this purpose due to
the lack of a better interface at that time.
Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211
flag is set on in the struct net_device_path_ctx to restore the
flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211
path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last
netdevice in the stack.
This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which
call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path.
Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the
flowtable GC worker until pending hw offload work has been completed.
Apparently, nf_flow_offload_stats() can schedule work to retrieve stats
while the flow is being removed by GC.
And this bit can also be used in a follow up patch to disable GC until
the flow has been fully added in both directions.
Revert the reordering done in commit d644b23afe1e ("netfilter:
flowtable: publish HW_DEAD after worker is done") to prevent a race
between GC and hw offload handler.
Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
As bpf_ct_set_nat_info() is not validating the NAT manipulation type a
wrong value can be passed directly to nf_nat_setup_info(). This triggers
the WARN_ON() at nf_nat_setup_info() and if panic_on_warn isn't set,
then IPS_SRC_NAT_DONE is set without adding nat_bysource and conntrack
cleanup tries to unlink an uninitialized hlist node.
Fix this by checking that NAT manipulation type is correct before
calling nf_nat_setup_info(). In addition, if the WARN_ON is hit, return
NF_DROP instead of continuing with the processing to avoid similar
situations in the future.
Reported-by: VEGA <vega@nebusec.ai>
Fixes: 0fabd2aa199f ("net: netfilter: add bpf_ct_set_nat_info kfunc helper")
Signed-off-by: Fernando Fernandez Mancera <fmancera@suse.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
Since nft_set_commit_update() runs set commit callbacks before processing
NEWSETELEM transactions, nft_rbtree_gc_scan() can observe elements added by
the transaction being committed.
The scan records an interval end in rbe_end without checking the element's
transaction state. A later, unrelated expired start then moves both
elements to the expired list. The synchronous GC queue can free the new end
element before the transaction subsequently activates it, causing a
use-after-free.
Only consider elements that are fully active in both generations. This
keeps transaction-state elements out of the GC scan and preserves interval
pairing across skipped elements.
KASAN reports:
BUG: KASAN: slab-use-after-free in nft_setelem_activate
nft_setelem_activate net/netfilter/nf_tables_api.c:7047
nf_tables_commit net/netfilter/nf_tables_api.c:11137
Allocated by task 130:
nft_set_elem_init net/netfilter/nf_tables_api.c:6794
nft_add_set_elem net/netfilter/nf_tables_api.c:7523
Freed by task 130:
nft_trans_gc_trans_free net/netfilter/nf_tables_api.c:10506
rcu_core kernel/rcu/tree.c:2919
Fixes: 1e3b9e1c77fe ("netfilter: nf_tables: call set ops .commit when building new ruleset blob")
Reported-by: <co+ee5e50ef2670e5f4@bugs.sh>
Assisted-by: LLM
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
While the IPVS SYNC protocol is not secure by design
we can still protect the backup server from messages that
can wreak havoc.
This commit addresses problems from received connection flags
or their combinations. We now drop messages as follows:
1. the NO_CPORT+TEMPLATE combination allows lookups for normal
connections to hit template which can break in many ways.
While the master does not sync connections with NO_CPORT flag,
i.e. before they are established, we still accept NO_CPORT
without TEMPLATE.
2. ONE_PACKET: it is not sent by master, so we do not
expect it in backup. Before now it was ignored by
IP_VS_CONN_F_BACKUP_MASK for protocol v1 while protocol
v0 created connections that are not hashed and dropped
immediately. Better to apply the IP_VS_CONN_F_BACKUP_MASK
also to the flags from v0 messages for consistency with v1.
Fixes: 87375ab47cd0 ("[IPVS]: ip_vs_ftp breaks connections using persistence")
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
The IP_VS_CONN_F_ONE_PACKET flag was implemented for normal
connections. When conn template inherits this flag from
dest->conn_flags it will not be hashed. As result, we will
create new template for every new normal connection.
Fix it to allow one template to be used by many normal
connections.
Fixes: 26ec037f9841 ("IPVS: one-packet scheduling")
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260916231652.127456-1-pablo%40netfilter.org
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
ip_vs_lblcr_new() and ip_vs_lblc_new() create cache entries for
every previously unseen destination address. The table max_size only
tells the periodic collector to reclaim entries after the cache has
already exceeded the limit. It does not reclaim entries that the
attacker continues to use.
Reject new cache entries once either table reaches max_size * 3 / 2.
The extra headroom lets the periodic collector catch up while the
existing scheduler fallback continues to use the selected destination
when cache creation fails. New traffic therefore stays serviceable
without growing the tables further.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Suggested-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai>
Acked-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
LBLC may delete cache entries for destinations that are
removed or overloaded and replace them with available ones.
But ip_vs_lblc_new() forgets to decrement the tbl->entries
counter after calling ip_vs_lblc_del(). This can lead to
increased shrinking of the cache with every new garbage
collection.
Fixes: 2f3d771a35fe ("ipvs: do not use dest after ip_vs_dest_put in LBLC")
Link: https://sashiko.dev/#/patchset/0bdd5abe9968ded7ca2b9cb6844ba83d94cc8d53.1787318053.git.zhilinz%40nebusec.ai
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
nft_flow_offload_init() bumps the flowtable use count with
nft_use_inc() before calling nf_ct_netns_get(). When the latter
fails, the error is returned as-is and the reference is leaked.
The upper layers do not balance it either: nf_tables_newexpr()
clears expr->ops when the expression init callback fails, so the
nft_expr_more() iteration in nft_rule_expr_deactivate() and
nf_tables_rule_destroy() stops right before the failed expression
and its ->destroy callback, which would drop the reference, never
runs.
Each failed rule addition therefore leaks one flowtable reference
and the flowtable can no longer be removed: NFT_MSG_DELFLOWTABLE
keeps reporting -EBUSY even though no rule references it.
Save the nf_ct_netns_get() return value and undo the nft_use_inc()
when it fails, restoring the inc/dec pairing within
nft_flow_offload_init() itself.
Fixes: a3c90f7a2323 ("netfilter: nf_tables: flow offload expression")
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Assisted-by: CodeBuddy:Kimi-K3
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
A TCP packet can carry new data while acknowledging traffic in the
opposite direction. With overlapping traffic in both directions, a
delayed packet's acknowledgment can be older than one Linux has already
accepted, even when that packet fills a gap in the received data.
Linux accepts the data, but tcp_ack() takes the old_ack path and skips
updating TS.Recent, the timestamp saved for outgoing acknowledgments.
The reply therefore echoes an older timestamp. If the sender uses this
echo to measure round-trip time after a long idle period, its estimate
includes the idle time and can reduce its sending rate.
Update TS.Recent in old_ack using tcp_replace_ts_recent(), before SACK
processing can trigger a transmission. This reuses the existing timestamp
and sequence checks, including PAWS protection against old duplicate
packets. ACK validation already rejects old ACKs in SYN_RECV before this
path, so no additional state check is needed.
Echoing the timestamp of the packet that fills the receive gap follows
RFC 7323 section 4.3. In a socket reproduction with 300 seconds idle,
controlled reordering and retransmission to exercise timestamp-based RTT
sampling, the sender's smoothed round-trip time was 37.5 seconds without
the fix and 15.5 ms with it.
Fixes: 12fb3dd9dc3c ("tcp: call tcp_replace_ts_recent() from tcp_ack()")
Assisted-by: LLM sparse
Signed-off-by: Jeff Jo <jeffjo@openai.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
Commit c4336a07eb6b ("net: correctly handle tunneled traffic on IPV6_CSUM
GSO fallback") split skb_gso_has_extension_hdr() into mutually exclusive
branches on skb->encapsulation. This did not yet address all paths:
1. With skb->encapsulation set, the outer header is not checked. IPv6
tunnels such as ip6_gre and ip6_tunnel add an outer Destination
Options header by default (encap_limit). Their GSO packets skip
software GSO, then skb_csum_hwoffload_help() sees the outer extension
header and calls skb_checksum_help() on the GSO skb, which warns and
drops it.
2. Tunnels over IPv6 without an outer transport header, such as
ip6_tunnel, leave skb->transport_header at the inner transport
header. skb_network_header_len() then spans the outer IPv6, tunnel
and inner IP headers, a false positive.
3. UDP tunnels without an inner network header, such as SCTP-in-UDP or
PSP, have no inner IPv6 header to check. Decide on the outer header
alone.
4. Directly dereferencing inner_ip_hdr(skb)->version without
skb_header_pointer() is unsafe.
Instead, check ipv6_ext_hdr(nexthdr) on the outer IPv6 header and, if
set, on the inner IPv6 header. Read the headers with skb_header_pointer().
Keep the skb_network_header_len() check when !skb->encapsulation to also
catch encapsulated packets without skb->encapsulation (e.g., virtio).
Use the same helper in skb_csum_hwoffload_help(). Its open coded test
has the false positive of (2) and ignores the inner header.
Background: checksum offload of tunneled packets invariants:
Non-GSO skb:
- If the inner packet is CHECKSUM_PARTIAL, Local Checksum Offload computes
the outer checksum in software and the device offloads only the inner L4
checksum.
- If the inner packet is CHECKSUM_NONE (e.g., SCTP-in-UDP, ESP-in-UDP,
Remote Checksum Offload), the device offloads the outer UDP or GRE
checksum instead.
GSO skb:
- A device with NETIF_F_GSO_UDP_TUNNEL_CSUM or NETIF_F_GSO_GRE_CSUM
computes both inner and outer checksums per segment.
The outer checksum is seeded with only the pseudo-header checksum.
A NETIF_F_IPV6_CSUM device must parse through the outer headers to
reach the inner ones.
Fixes: c4336a07eb6b ("net: correctly handle tunneled traffic on IPV6_CSUM GSO fallback")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260926140506.2335137-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue
flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
cleared the flag before sch_handle_egress() and only read it afterwards,
so the flag was not confined to the xmit that set it: a nested xmit
(mirred redirect or mirror, or a drop after skbedit) could set the flag
and the outer xmit would consume it for an skb that never went through
skbedit.
A forwarded packet still carries the ingress NIC's rx_queue + 1 in
skb->queue_mapping, so the outer device then indexes its tx queue state
with that stale value. Taprio's child array q->qdiscs[] is sized to the
device's queue count, so taprio_enqueue() indexes past its allocation
and dereferences the result as a struct Qdisc *.
We (ab)use the skb->nf_skip_egress which means "skip netfilter egress
for this packet" to tag to "am I in tc egress?". Despite the overload
I dont see it as a conflict since the marker is set only around the
single sch_handle_egress() call and ingress path is guarded by
tc_at_ingress.
I will send a followup(net-next) patch once this hits net-next to
rename the skb->nf_skip_egress bit/flag to skb->skip_egress
Arm the flag only from the egress classifier that can use it: raise
skip_txqueue from tcf_skbedit_act() only when it runs inside
sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc
classifier runs in q->enqueue(), after the tx queue has been picked,
so a mapping it sets cannot affect the current packet; arming the flag
there only pollutes it for a later xmit. Then own the flag for the xmit
frame the egress hook runs in: save the incoming value and clear it just
before sch_handle_egress(), and restore it after the hook - on the
consumed (drop) path, or, in the same call that reads it, on the
surviving path. The save and the restores stay inside the
egress_needed_key static branch, so a packet pays for them only when
egress hooks are active (2f1e85b1aee4).
Store the value netdev_cap_txqueue() selected back into skb->queue_mapping
in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
the skip_txqueue path never hands a later reader on the xmit path a
mapping the device cannot serve. A store made still later in the same
frame, by a tc BPF program attached to a transmit qdisc, is outside this
path and is not re-capped; a separate followup will resolve that path.
netdev_xmit_skip_txqueue() returns the previous flag value so the
save-and-clear is one call, and a no-op stub is provided when
CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under
CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
skb_at_tc_egress() is valid whenever the egress path is built.
A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer after setting a mapping valid only
on the larger device. That reaches these reads and, under KASAN, faults
with "slab-out-of-bounds in taprio_enqueue".
Conditions to recreate the bug: the report's own trigger is a local user
with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
(2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
(num_tc 1, queues 2@0) and clsact on all three; then add an egress
matchall filter on every device. On qa: "action skbedit queue_mapping 2
pipe action mirred egress redirect dev qb". On qb: "action mirred egress
mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
flight; without the fix qb consumes it and reads its two-entry taprio
child array with the forwarded packet's stale mapping. A qc whose
skbedit is instead installed in a transmit-qdisc classifier (a matchall
filter on the qc root qdisc) reaches the same code path the same way
without the fix.
Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
past a 16-byte taprio_init() allocation, for the clsact-setter and the
transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
store; the fixed kernel runs all three with no report, and the BPF store
variant additionally shows the expected "selects TX queue" clamp notice
from the write-back.
Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
Suggested-by: Eric Dumazet <edumazet@google.com>
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
skb_splice_from_iter() appends multiple page fragments before updating
skb->len. However, skb_splice_csum_page() uses that unchanged length as the
offset for every fragment checksum. If bytes already appended during the
call have an odd total length, a later fragment's checksum is combined
with the wrong parity.
For example, prime a UDP socket with sendto(MSG_MORE) and splice two
distinct pipe buffers containing "abc" and "DEFGH". Uncorking reports
success, but the receiver discards the packet for a bad checksum.
This also affects IPv6 and fragments beginning near a page boundary.
Pass the initial skb length plus the bytes already spliced to the checksum
helper. Keep the existing final length update and partial-progress error
handling unchanged. CHECKSUM_PARTIAL does not use this helper and remains
unaffected.
Fixes: 2e910b95329c ("net: Add a function to splice pages into an skbuff for MSG_SPLICE_PAGES")
Cc: stable@vger.kernel.org
Signed-off-by: Alireza Asgari <alireza@asgari.net>
Link: https://patch.msgid.link/20260924-fix-udp-splice-checksum-v1-1-fe61d65447a8@asgari.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
tpacket_snd() drains every SEND_REQUEST frame in a TX ring and adds
each valid packet length to the signed int len_sum. Userspace can
recycle completed frames concurrently, so a single blocking send call
is not bounded by the ring size.
After enough successful transmissions len_sum wraps into the negative
range. In particular, 65536 65535-byte frames followed by a 65007-byte
frame produce 0xfffffdef, or -EIOCBQUEUED. sock_sendmsg_nosec() treats
that as an impossible return and triggers a BUG. On systems configured
to panic on oops, a CAP_NET_RAW holder can panic the host kernel.
Stop before the next frame would make the byte count exceed INT_MAX.
The positive short count leaves that frame pending for a later send
call.
A two-thread TPACKET_V2 reproducer triggered the BUG on an arm64 build
of Linux 7.3-rc4. With this change, the same reproducer returned the
expected positive short count without an oops.
Fixes: 69e3c75f4d54 ("net: TX_RING and packet mmap")
Cc: stable@vger.kernel.org
Signed-off-by: Ilkka Lappeteläinen <ilkka.lappetelainen@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260925124829.6595-1-ilkka.lappetelainen@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
commit 6439461f1618 ("net/sched: sch_codel: clamp default mtu to avoid
disabling CoDel") clamped q->params.mtu to [256, 1 << 20]. The value is
the CoDel no-drop threshold (codel_impl.h "*backlog <= params->mtu"), so
on a link whose maximum transmitted packet size is below 256 the floor
extends CoDel's minimum-backlog exemption beyond one packet and delays
drop or mark eligibility by several small packets.
Keep the upper bound that guards the original overflow (psched_mtu()
wrapping to ~2 GiB on a huge-MTU device) but drop the 256 floor, so the
threshold tracks the real device packet size.
Conditions to recreate the bug: attach a codel qdisc on a link whose
MTU plus hard_header_len is below 256 (e.g. a CAN interface). At that
MTU the no-drop threshold must equal the device MTU plus its
hard-header length; before this patch it was forced to 256.
Requires CAP_NET_ADMIN in a user namespace.
Fixes: 6439461f1618 ("net/sched: sch_codel: clamp default mtu to avoid disabling CoDel")
Reported-by: Sashiko (nipa) <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260819143136.57350-1-jhs@mojatatu.com
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-34MS.v1.20260925165535@mojatatu.com.2
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
commit d9ebd8f9aa8b ("net/sched: fq_codel: clamp default quantum and mtu")
clamped both q->quantum and q->cparams.mtu to [256, FQ_CODEL_QUANTUM_MAX].
The two fields mean different things: quantum is a DRR credit that wants
the 256 floor, but cparams.mtu is the CoDel no-drop threshold
(codel_impl.h "*backlog <= params->mtu"). On a link whose maximum
transmitted packet size is below 256, the floor extends CoDel's
minimum-backlog exemption beyond one packet and delays drop or mark
eligibility by several small packets.
Split the clamp. quantum keeps [256, FQ_CODEL_QUANTUM_MAX]; cparams.mtu
tracks psched_mtu() (the device MTU plus its hard-header length) with
only the upper bound that guards the original overflow (psched_mtu()
wrapping to ~2 GiB on a huge-MTU device).
Conditions to recreate the bug: attach an fq_codel qdisc on a link
whose MTU plus hard_header_len is below 256 (e.g. a CAN interface). At
that MTU the no-drop threshold must equal the device MTU plus its
hard-header length; before this patch it was forced to 256.
Basic Testing done: with dev->mtu=100 and hard_header_len=14, a
return probe on fq_codel_init() observed cparams.mtu change from 256 to 114
Fixes: d9ebd8f9aa8b ("net/sched: fq_codel: clamp default quantum and mtu")
Reported-by: Sashiko (nipa) <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260819143136.57350-1-jhs@mojatatu.com
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-34MS.v1.20260925165535@mojatatu.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
IP-in-IP GSO can re-enter inet_gso_segment() or ipv6_gso_segment()
for each nested IP header. encap_level tracks header bytes, not callback
depth, so a deep chain can exhaust the kernel stack. Making
inet_gso_segment() stackable introduced unbounded IPv4 nesting; IPIP
GSO/TSO later made the path reachable. The IPv6 stackable path was
introduced separately and uses the same guard.
Count IPv4 and IPv6 GSO handler entries in skb_gso_cb, initialized once
per top-level GSO operation and preserved across GRE/UDP context changes.
Use the existing IP_TUNNEL_RECURSION_LIMIT for both handlers. The first
five entries pass, and the sixth returns -EINVAL before dispatching
another GSO callback.
Fixes: 3347c9602955 ("ipv4: gso: make inet_gso_segment() stackable")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Closes: https://lore.kernel.org/all/cover.1790157745.git.zihanx@nebusec.ai/
Assisted-by: LLM
Co-developed-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Luxing Yin <root@tr0jan.top>
Signed-off-by: Zihan Xi <zihanx@nebusec.ai>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260924051521.32568-2-zihanx@nebusec.ai
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Commit 2bb2f5fb21b0 ("net: add new socket option SO_RESERVE_MEM") and
commit d00c8ee31729 ("net: fix possible NULL deref in sock_reserve_memory")
only checked sk_has_account(sk), which is true for both TCP and UDP
sockets.
However, SO_RESERVE_MEM and sk_unused_reserved_mem() are currently only
supported by TCP:
- On UDP sockets, sk->sk_forward_alloc is protected by
sk->sk_receive_queue.lock, whereas sock_reserve_memory() and
sock_release_reserved_memory() only acquire lock_sock(sk). Concurrent
UDP packet reception/release and setsockopt(SO_RESERVE_MEM) corrupt
sk_forward_alloc and memcg accounting.
- udp_rmem_release() reclaims excess sk_forward_alloc without
accounting for sk_unused_reserved_mem(sk).
Restrict sock_reserve_memory() to TCP sockets (sk_is_tcp(sk)) for now.
Supporting SO_RESERVE_MEM for UDP (acquiring sk_receive_queue.lock and
honoring sk_unused_reserved_mem() in udp_rmem_release()) can be done in
a future net-next series if needed.
In addition, reject val > SZ_1G with -EINVAL in
sk_setsockopt(SO_RESERVE_MEM). Without an upper bound, values near
INT_MAX cause sk_mem_pages(delta) and (pages << PAGE_SHIFT) to overflow
32-bit signed int, corrupting sk->sk_forward_alloc and
sk->sk_reserved_mem. Using a page-aligned cap (SZ_1G) ensures that the
page-rounded sk->sk_reserved_mem reported by getsockopt(SO_RESERVE_MEM)
can always be passed back to setsockopt(SO_RESERVE_MEM).
Fixes: 2bb2f5fb21b0 ("net: add new socket option SO_RESERVE_MEM")
Reported-by: Cai Xinchen <caixinchen1@huawei.com>
Closes: https://lore.kernel.org/netdev/5a88421d-10ef-4fca-9acb-85a27a3c1173@huawei.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Wei Wang <weiwan@google.com>
Link: https://patch.msgid.link/20260925135244.3715196-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
TSCONFIG_GET and TSCONFIG_SET reject a device that implements neither
hwtstamp NDO, even when its PHY can serve the request. The ioctls they
meant to replace handle it, so the two interfaces disagree on the same
hardware and user space has to pick one.
Accept the default timestamping PHY and an already installed provider on
both sides. On the set side the test moves into ethnl_set_tsconfig() as
the validate callback runs without rtnl. On the get side it stays ahead
of ethnl_ops_begin(), so a device that can serve nothing keeps failing
with EOPNOTSUPP and a dump still skips it rather than stopping there.
A netdev provider needs ndo_hwtstamp_set to be programmed at all, so
don't pick that source without it, and test for ndo_hwtstamp_get before
calling it.
Fixes: 6e9e2eed4f39 ("net: ethtool: Add support for tsconfig command to get/set hwtstamp config")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Reviewed-by: Kory Maincent <kory.maincent@bootlin.com>
Link: https://patch.msgid.link/20260925135237.3432266-3-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
advance_nextseg() modifies the SRH Segments Left field and the IPv6
destination address without ensuring the packet data is writable.
seg6_next_csid_advance_arg() has the same problem when it advances the
NEXT-C-SID argument in the destination address.
The skb may be cloned, for example by an AF_PACKET socket receiving on
the ingress device. advance_nextseg() and seg6_next_csid_advance_arg()
then write into the packet data shared with the clone. A read from that
socket can return the modified packet instead of the received one. The
simplified path below shows this for advance_nextseg():
__netif_receive_skb_one_core
__netif_receive_skb_core
deliver_skb [orig: users=2, cloned=0]
packet_rcv
skb_clone clone queued to the AF_PACKET socket
consume_skb(orig) [orig: users=1, cloned=1]
ipv6_rcv
ip6_rcv_core skb_share_check: no-op [orig: users=1]
[...]
input_action_end_core
advance_nextseg writes into the data shared with the clone
Call skb_ensure_writable() in advance_nextseg() and in
seg6_next_csid_advance_arg() before they modify the packet data.
skb_ensure_writable() may reallocate skb->head, which invalidates the
pointers into the packet data taken before the call.
advance_nextseg() now returns a valid SRH pointer, or NULL if
skb_ensure_writable() fails.
seg6_next_csid_advance_arg() takes the pointer to the destination
address from the skb after skb_ensure_writable().
On failure, the callers drop the packet with SKB_DROP_REASON_NOMEM.
Fixes: 140f04c33bbc ("ipv6: sr: implement several seg6local actions")
Fixes: 848f3c0d4769 ("seg6: add NEXT-C-SID support for SRv6 End behavior")
Signed-off-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260925133807.32-1-andrea.mayer@uniroma2.it
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
With CONFIG_PAGE_POOL_STATS=y, page_pool_recycle_ring_bulk() updates
the recycle stats after dropping the producer lock. If it just recycled
the last inflight netmems of a pool being destroyed, page_pool_release()
can pass its producer-lock barrier and free the pool before
recycle_stat_add() runs.
Commit fcc680a647ba7 ("page_pool: allow mixing PPs within one bulk")
moved this stat update after the unlock. Commit 271683bb2cf32
("page_pool: Fix use-after-free in page_pool_recycle_in_ring") later
added the barrier, but only fixed page_pool_recycle_in_ring(). Move
the update back under the lock.
Fixes: fcc680a647ba7 ("page_pool: allow mixing PPs within one bulk")
Link: https://lore.kernel.org/r/179028939171.2160803.706522228583664641@kernel.org
Cc: Kaifeng Wang <kaifengw@google.com>
Cc: Dong Chenchen <dongchenchen2@huawei.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Link: https://patch.msgid.link/20260925144127.1445667-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
TIPC can encrypt traffic between nodes using a different transmit key
for each node. In this mode, the AES-GCM nonce for a packet sent to a
known peer consists of a 32-bit prefix (a per-key salt XOR the peer's
address) followed by a 64-bit counter. That counter is stored in the
peer's RX crypto object. When the peer reports a change in which key
it uses to receive packets, TIPC resets this counter. The sender can
still be using the same TX key and salt, so subsequent packets reuse
earlier nonces. This nonce reuse breaks confidentiality and exposes
GCM's authentication key. This makes forgeries trivial: an attacker
can exploit CTR malleability to alter captured ciphertexts and use the
recovered authentication key to compute a valid tag for the modified
ciphertext, under the same key and nonce.
Use the TX key's existing aead->seqno counter instead. All encryptions
using that key object share the same atomic counter, so concurrent
encryptions get distinct nonce counter values. The counter survives
key activation and peer reconnection, and peer key-status reports
cannot reset it. This prevents those transitions from causing nonce
reuse while the same TX key remains installed.
The nonce format is unchanged, and receivers do not require consecutive
counter values, so sharing the counter across peers remains compatible
with existing receivers. A pre-existing check still invokes key
revocation in the unlikely event that the counter wraps to zero.
Fixes: fc1b6d6de220 ("tipc: introduce TIPC encryption & authentication")
Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Reviewed-by: Tung Nguyen <tung.quang.nguyen@est.tech>
Link: https://patch.msgid.link/20260924202105.3722778-1-Jeremy.Jean@oss.cyber.gouv.fr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Luiz Augusto von Dentz says:
====================
bluetooth pull request for net:
Core:
- hci_core: Serialize ACL scheduling with channel deletion
- hci_core: Serialize SCO and ISO scheduling with teardown
- hci_core: Serialize fragmented ISO packet queueing
- hci_core: Fix inquiry cache timestamps on 64-bit systems
- hci_core: free the HCI ID if naming fails
- hci_conn: Lock parent access during enhanced SCO setup
- hci_sync: Fix inquiry cache use-after-free
- hci_sync: don't drain cmd_sync backlog on unregister
- RFCOMM: Fix initial port reference race
- SMP: Serialize SMP remote OOB data access
Drivers:
- btintel: fix buffer over-read in btintel_hw_error()
- btintel: validate DDC record lengths
- btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
- btintel_pcie: reject oversized TX packets in send_frame()
* tag 'for-net-2026-09-28' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
Bluetooth: hci_core: Serialize SCO and ISO scheduling with teardown
Bluetooth: hci_core: Serialize ACL scheduling with channel deletion
Bluetooth: Serialize SMP remote OOB data access
Bluetooth: RFCOMM: Fix initial port reference race
Bluetooth: hci_sync: Fix inquiry cache use-after-free
Bluetooth: hci_sync: don't drain cmd_sync backlog on unregister
Bluetooth: hci_core: Serialize fragmented ISO packet queueing
Bluetooth: hci_conn: Lock parent access during enhanced SCO setup
Bluetooth: btintel: validate DDC record lengths
Bluetooth: hci_core: free the HCI ID if naming fails
Bluetooth: hci_core: Fix inquiry cache timestamps on 64-bit systems
Bluetooth: btintel_pcie: reject oversized TX packets in send_frame()
Bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
Bluetooth: btintel: fix buffer over-read in btintel_hw_error()
====================
Link: https://patch.msgid.link/20260928152919.942973-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
hci_low_sent() selects a connection under RCU but drops the read lock
before calculating its quota and returning it to the scheduler. The
connection is then used without lifetime protection by hci_sched_iso()
and hci_sched_sco().
After the TX worker drops the RCU read lock, hci_abort_conn_sync() on
hdev->req_workqueue can remove the connection from the hash, complete
synchronize_rcu(), purge its queues and release it. The TX worker on
hdev->workqueue can then access the freed connection in hci_quote_sent()
or while dequeuing packets and updating conn->sent.
KASAN reported:
BUG: KASAN: slab-use-after-free in hci_low_sent+0x730/0x840
Workqueue: hci0 hci_tx_work
Call Trace:
hci_low_sent+0x730/0x840
hci_sched_iso+0x25e/0x4d0
hci_tx_work+0x239/0xcb0
Allocated by task 93:
__hci_conn_add+0x16f/0x1b40
hci_bind_bis+0x782/0x17b0
hci_connect_bis+0xa0/0x510
iso_sock_connect+0x589/0x1050
Freed by task 88:
kfree+0x131/0x3c0
device_release+0xc8/0x240
kobject_put+0x14d/0x280
hci_conn_del+0x55a/0xe80
hci_disconnect_sync+0x156/0x180
hci_abort_conn_sync+0x3e7/0x940
hci_cmd_sync_work+0x13c/0x290
Hold hci_dev_lock() across connection selection and transmission in both
SCO and ISO scheduling, serializing them with connection teardown. This
also prevents queuing completion timestamps after the connection queues
have been purged. Extending RCU across transmission would be unsafe
because the transmit path can sleep. Keep the ISO timeout check outside
the mutex since hci_link_tx_to() takes it itself.
The preceding channel fix already holds this mutex in the ACL and LE
schedulers, which call the SCO scheduler between packets. Move the SCO
body to __hci_sched_sco(), assert that its caller holds the mutex, and
use it directly from these locked paths. Keep a locking hci_sched_sco()
wrapper for the direct calls from hci_tx_work(). This avoids recursively
acquiring the device mutex while preserving the scheduling order.
Remove the obsolete claim that connection removal disables TX.
Fixes: bf4c63252490 ("Bluetooth: convert conn hash to RCU")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
hci_chan_sent() selects a channel under RCU but drops the read lock before
accessing chan->conn and returning the channel. The ACL and LE schedulers
then use its packet queue and update its transmit counters without any
protection against channel deletion.
After TX selects a channel and releases RCU, a disconnect command timeout
on the separate request workqueue can run hci_conn_failed() and
l2cap_conn_del(). hci_chan_del() can then unlink the channel, complete
synchronize_rcu(), purge its queue and free it before TX resumes. This
causes use-after-free both in hci_chan_sent() and in its callers.
KASAN reported:
BUG: KASAN: slab-use-after-free in hci_chan_sent+0x892/0x9b0
Workqueue: hci0 hci_tx_work
Call Trace:
hci_chan_sent+0x892/0x9b0
hci_tx_work+0x5e6/0xb70
Allocated by task 91:
hci_chan_create+0xe3/0x350
l2cap_conn_add.part.0+0x12/0xa30
l2cap_chan_connect+0x110d/0x1b60
l2cap_sock_connect+0x310/0x530
Freed by task 99:
hci_chan_del+0x11f/0x170
l2cap_conn_del+0x4f1/0x800
l2cap_connect_cfm+0x88c/0xd30
hci_conn_failed+0x150/0x250
hci_abort_conn_sync+0x3e3/0x800
hci_cmd_sync_run+0x7e/0xc0
hci_abort_conn+0x105/0x1f0
disconnect_sync+0x157/0x290
hci_cmd_sync_work+0x13c/0x290
Hold the existing device mutex across channel selection and transmission
in both schedulers to serialize them with channel teardown. Keep timeout
handling outside the critical sections because hci_link_tx_to() acquires
the same mutex. This also permits the transmit path to sleep, unlike
extending the RCU read-side critical section across packet submission.
Fixes: 3eff45eaf817 ("Bluetooth: convert tx_task to workqueue")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
build_pairing_cmd() looks up remote OOB data and copies its contents
without holding hdev->lock, which serializes the list's writers. After
SMP finds an entry, a concurrent management Remove Remote OOB Data
command can unlink and free it before SMP reads its present flag or
copies its random and confirmation values. Removal can also invalidate
an entry while the lookup is still traversing the list.
KASAN reported:
BUG: KASAN: slab-use-after-free in build_pairing_cmd+0x948/0x9b0
Call Trace:
build_pairing_cmd+0x948/0x9b0
smp_recv_cb+0x459f/0x8110
l2cap_recv_frame+0xf14/0x9190
l2cap_recv_acldata+0xa64/0xd40
hci_rx_work+0x4ca/0x730
Allocated by task 87:
hci_add_remote_oob_data+0x11d/0x530
add_remote_oob_data+0x282/0x400
hci_sock_sendmsg+0x1033/0x1ea0
Freed by task 93:
hci_remote_oob_data_clear+0x108/0x1c0
remove_remote_oob_data+0x198/0x220
hci_sock_sendmsg+0x1033/0x1ea0
Taking hdev->lock in build_pairing_cmd() would recurse for callers that
already hold it and invert the device-to-L2CAP lock order on the receive
path. Add a per-device remote_oob_lock instead, held across the SMP
lookup and copies and by the add, remove and clear helpers. Cover
initialization and in-place updates as well, so SMP cannot read partially
initialized or updated OOB values. Release the mutex on allocation
failure, preserving the existing error return.
The new critical sections acquire no device, connection or channel locks.
Writers retain their existing hdev->lock protection, which continues to
serialize the other readers without changing their locking or behavior.
Link: https://lore.kernel.org/r/00660cd3-7d71-13a4-f617-229e6defb701@gmail.com
Fixes: 02b05bd8b0a6 ("Bluetooth: Set SMP OOB flag if OOB data is available")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
For RFCOMM_RELEASE_ONHUP devices, rfcomm_tty_install() and
__rfcomm_release_dev() both use RFCOMM_TTY_OWNED to decide who drops the
initial tty_port reference. The release path tests the bit separately
from the install path setting it, so both can drop that reference.
The synchronous hangup does not prevent this race: port->tty is not
assigned until tty_port_open(), after installation. The following
interleaving is possible:
1. Install and release each obtain a reference with rfcomm_dev_get().
2. Release observes RFCOMM_TTY_OWNED clear.
3. Install sets RFCOMM_TTY_OWNED and drops the initial reference.
4. Release drops the same initial reference again.
5. TTY cleanup drops its reference and frees the device.
6. Release performs its final tty_port_put() on the freed port.
KASAN reported:
BUG: KASAN: slab-use-after-free in tty_port_put+0x22/0x190
Write of size 4 at addr ffff8881001d3d5c by task poc/92
Call Trace:
tty_port_put+0x22/0x190
rfcomm_dev_ioctl+0x1d4/0x1930
sock_do_ioctl+0x110/0x260
sock_ioctl+0x380/0x590
__x64_sys_ioctl+0x134/0x1c0
Allocated by task 88:
rfcomm_dev_ioctl+0x8f6/0x1930
sock_do_ioctl+0x110/0x260
sock_ioctl+0x380/0x590
__x64_sys_ioctl+0x134/0x1c0
Freed by task 65:
kfree+0x131/0x3c0
rfcomm_dev_destruct+0x23b/0x2f0
release_one_tty+0xc1/0x370
process_one_work+0x661/0x1090
Use test_and_set_bit() in both paths so that only the caller that changes
RFCOMM_TTY_OWNED from clear to set drops the initial reference. Each path
keeps its own lookup reference until release or TTY cleanup, preserving
the existing callback ordering and error handling.
Fixes: 80ea73378af4 ("Bluetooth: Fix unreleased rfcomm_dev reference")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
The sync command worker holds hdev->req_lock, but inquiry-cache updates
and flushes use hdev->lock. Both hci_acl_create_conn_sync() and
hci_stop_discovery_sync() look up entries and read their fields without
taking hdev->lock.
After either lookup returns, a concurrent HCIINQUIRY ioctl can acquire
hdev->lock and flush the cache, freeing the entry. The worker then reads
the freed entry while preparing a create-connection or remote-name-cancel
command. The list traversal also races with cache updates and removal.
KASAN reported these accesses:
BUG: KASAN: slab-use-after-free in hci_acl_create_conn_sync+0x5f1/0x650
Workqueue: hci0 hci_cmd_sync_work
Call Trace:
hci_acl_create_conn_sync+0x5f1/0x650
hci_cmd_sync_work+0x13c/0x290
Allocated by task 90:
hci_inquiry_cache_update+0x3e6/0x7d0
hci_inquiry_result_evt+0x3cb/0x560
Freed by task 93:
hci_inquiry_cache_flush+0x111/0x2b0
hci_inquiry+0x2f2/0x780
hci_sock_ioctl+0x269/0x5f0
BUG: KASAN: slab-use-after-free in hci_stop_discovery_sync+0x3b1/0x3c0
Workqueue: hci0 hci_cmd_sync_work
Call Trace:
hci_stop_discovery_sync+0x3b1/0x3c0
hci_cmd_sync_work+0x173/0x300
Allocated by task 86:
hci_inquiry_cache_update+0x483/0x940
hci_inquiry_result_evt+0x3cb/0x560
Freed by task 91:
hci_inquiry_cache_flush+0x13e/0x2f0
hci_inquiry+0x2f2/0x780
hci_sock_ioctl+0x269/0x5f0
Hold hdev->lock across each lookup and all reads from its result. Copy the
remote address before unlocking so discovery cancellation does not retain
a cache entry pointer. Release the lock before sending synchronous HCI
commands, since their completion handlers may need the same lock.
Fixes: cf75ad8b41d2 ("Bluetooth: hci_sync: Convert MGMT_SET_POWERED")
Fixes: 45340097ce6e ("Bluetooth: hci_conn: Only do ACL connections sequentially")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
hci_unregister_dev() disables cmd_work and cmd_timer, then calls
hci_cmd_sync_clear(), whose cancel_work_sync() waits for
hci_cmd_sync_work() to return. That worker dequeues and runs every
entry on cmd_sync_work_list. With cmd_work disabled nothing reaches
the controller, so each queued HCI command waits the full
HCI_CMD_TIMEOUT (2s).
The backlog has no bound. Userspace can keep queuing MGMT commands
such as MGMT_OP_GET_CLOCK_INFO against a controller that doesn't
answer. Closing /dev/vhci then blocks in vhci_release() for
backlog * 2s:
INFO: task syz-executor:5749 blocked for more than 143 seconds.
cancel_work_sync
hci_cmd_sync_clear
hci_unregister_dev
vhci_release
In a local reproduction the backlog held more than 5000 entries,
which comes to hours of hang.
Stop the worker from taking new entries once HCI_UNREGISTER is set,
and wake any request still waiting with -ENODEV before cancelling
the work. The entries left over are destroyed with -ECANCELED by the
existing sweep in hci_cmd_sync_clear(). They stay on the list until
the work has stopped, so hci_cmd_sync_dequeue() and friends still see
them. A callback already running may still issue another command,
which delays unregister by at most one timeout per remaining command,
not by the whole backlog.
Fixes: 008ee9eb8a11 ("Bluetooth: hci_sync: Fix not processing all entries on cmd_sync_work")
Reported-by: syzbot+217e3f1283cafe80586e@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=217e3f1283cafe80586e
Signed-off-by: Nguyen Ngoc Thang <ngocthang2710.1999@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
The fragmented path in hci_queue_iso() calls __skb_queue_tail() without
holding conn->data_q.lock. The socket lock serializes senders, but the
transmit worker can concurrently remove packets from this queue using
skb_dequeue().
After an insertion saves the old tail pointer, hci_sched_iso() can dequeue
that packet and pass it to the driver. The driver can free the packet
before the insertion resumes and writes to the old tail's next pointer,
causing a use-after-free. Concurrent updates can also corrupt the queue.
KASAN reported:
BUG: KASAN: slab-use-after-free in hci_send_iso+0xc7d/0xe10
Call Trace:
hci_send_iso+0xc7d/0xe10
iso_sock_sendmsg+0x710/0x900
__sys_sendto+0x34a/0x3a0
Allocated by task 86:
__alloc_skb+0xdd/0x820
alloc_skb_with_frags+0x7e/0x770
sock_alloc_send_pskb+0x67a/0x820
bt_skb_sendmsg.constprop.0+0xc0/0x6c0
iso_sock_sendmsg+0x44e/0x900
Freed by task 90:
kmem_cache_free+0xcb/0x3d0
vhci_read+0x33f/0x4d0
vfs_read+0x177/0xa20
Hold the queue lock across the entire fragment batch, as hci_queue_acl()
does, to serialize insertion against dequeue while preserving fragment
ordering.
Fixes: 26afbd826ee3 ("Bluetooth: Add initial implementation of CIS connections")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
Bluetooth: hci_conn: Lock parent access during enhanced SCO setup
hci_enhanced_setup_sync() runs on the request workqueue without the
hci_dev_lock held by its caller when setup was queued. Its CVSD
capability check and find_next_esco_param() dereference conn->parent
while the receive workqueue can unlink and release that parent.
The following interleaving can cause a use-after-free:
hci_enhanced_setup_sync() hci_disconn_complete_evt()
load conn->parent
hci_dev_lock()
hci_conn_del(ACL parent)
unlink SCO child
drop link's parent reference
clear child->parent
release ACL parent
bt_link_release()
kfree(parent)
hci_dev_unlock()
read parent->features[0][3]
The reference held for the queued SCO child does not keep its ACL parent
alive after unlinking. Commit 42de40abe25d ("Bluetooth: hci_conn: fix the
SCO setup context lifetime") protects the child stored in the queued
context, but leaves these parent accesses unprotected.
With a 40 ms diagnostic delay after loading conn->parent, an instrumented
kernel based on fd179f8a05be, which already contains 42de40abe25d,
reported:
BUG: KASAN: slab-use-after-free in hci_enhanced_setup_sync+0xda5/0xdf0
Read of size 1 at addr ffff888102204047 by task kworker/u17:0/93
Call Trace:
hci_enhanced_setup_sync+0xda5/0xdf0
hci_cmd_sync_work+0x13c/0x290
process_one_work+0x6b4/0x10e0
Allocated by task 92:
__hci_conn_add+0x304/0x1df0
hci_connect_acl+0x349/0x3e0
hci_connect_sco+0x3b/0x9a0
sco_sock_connect+0x475/0xca0
Freed by task 94:
kfree+0x121/0x3c0
bt_link_release+0x79/0xa0
device_release+0xc8/0x240
kobject_put+0x14d/0x280
hci_conn_del+0x524/0xe30
hci_disconn_complete_evt+0x403/0x8c0
hci_event_packet+0x71b/0xb20
hci_rx_work+0x293/0x730
The accessed address is 71 bytes into the freed ACL parent, at its
features[0][3] byte; the queued SCO child is a different object. The
diagnostic preserves the loaded parent across the delay, matching the
unmodified compiled capability check, and does not change the parent
references or teardown path.
Hold hci_dev_lock() across the codec switch, including every call to
find_next_esco_param(), and release it on all selection errors. Keep
configure_datapath_sync() outside the critical section because it waits
for HCI events. Preserve parameter selection and existing return values.
Fixes: e07a06b4eb41 ("Bluetooth: Convert SCO configure_datapath to hci_sync")
Cc: stable@vger.kernel.org
Assisted-by: GPT-6-Astra
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
|
|
ieee80211_monitor_select_queue() validates the skb length using skb->len
before dereferencing pointers into skb->data to access the 802.11 header.
However, skb->len includes data in both the linear head buffer and
non-linear fragments/pages. When AF_PACKET sends a packet large enough to
become non-linear, the 802.11 header at skb->data + len_rthdr may extend
past the linear head buffer even though skb->len appears sufficient.
This leads to a slab-out-of-bounds read when accessing hdr->frame_control,
as the kernel reads beyond the allocated skb head buffer:
BUG: KASAN: slab-out-of-bounds in ieee80211_monitor_select_queue+0x1ed/0x220
syzbot has hit the same issues.
Fix this by checking skb_headlen(skb) (which gives the length of the
linear data region) instead of skb->len before dereferencing skb->data
pointers. This ensures the required bytes are actually present in the
linear portion of the skb.
Apply the same fix to ieee80211_validate_radiotap_len() which has the
same class of bug: it uses skb->len to validate accesses to skb->data,
and to the corresponding checks in ieee80211_monitor_start_xmit(): for
drivers advertising NETIF_F_SG the skb is not linearized before
ndo_start_xmit, so the 802.11 header can be read past the linear
buffer there as well.
Fixes: cf0277e714a0 ("mac80211: fix skb buffering issue")
Fixes: 9b8a74e3482f ("[MAC80211]: Improve sanity checks on injected packets")
Reported-by: AutonomousCodeSecurity@microsoft.com
Reported-by: syzbot+610e40369bc02181bad0@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=610e40369bc02181bad0
Reported-by: syzbot+878643e0580bc580f883@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=878643e0580bc580f883
Assisted-by: GitHub-Copilot:claude-opus-4.6
Signed-off-by: Cen Zhang (Microsoft) <cenzhang@linux.microsoft.com>
Reviewed-by: Francis Perron <francis@akrites.dev>
Link: https://patch.msgid.link/20260925024408.32143-1-cenzhang@linux.microsoft.com
Signed-off-by: Johannes Berg <johannes.berg@intel.com>
|
|
The HT/VHT channel definition validator can receive a 320 MHz
bandwidth indication from a received CSA frame on a non-6 GHz link.
It warns for this unsupported width but continues with an uninitialized
vht_operation.chan_width, which is then read by
ieee80211_chandef_vht_oper(). With panic_on_warn enabled, this lets a
received frame panic the kernel.
Reject the channel definition before entering the VHT operation
conversion. This preserves the existing CSA fallback while avoiding
both the warning and the uninitialized read.
Fixes: 21c3f8f95554 ("wifi: mac80211: refactor STA CSA parsing flows")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Ruide Cao <caoruide123@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Link: https://patch.msgid.link/bd68e360e06135c750397cb6be003a01a834a179.1787222001.git.vega.cover-letter@nebusec.ai
Signed-off-by: Johannes Berg <johannes.berg@intel.com>
|
|
Unlike the 802.11 TX paths, ieee80211_8023_xmit() does not set
info->band from the channel context on non-MLD interfaces. It stays at
0, so ieee80211_get_tx_rates() uses the wrong band for these frames.
Set it the same way as ieee80211_build_hdr(): drop the frame if there is
no channel context, and leave the band at 0 on MLD interfaces.
Reported-by: Andrea Pesaresi <andreapesaresi82@gmail.com>
Fixes: 50ff477a8639 ("mac80211: add 802.11 encapsulation offloading support")
Cc: stable@vger.kernel.org
Signed-off-by: Felix Fietkau <nbd@nbd.name>
Link: https://patch.msgid.link/20260926120216.4054712-1-nbd@nbd.name
Signed-off-by: Johannes Berg <johannes.berg@intel.com>
|
|
'Fixes' commit stopped calling ip_set_init_comment() for hash types
from kernel-side-adds (xtables .. -j SET). ip_set_init_comment() says:
"The kadt functions don't use the comment extensions in any way."
But bitmap set type calls the function from kadt cb too.
While this appears to be safe (serialized via the set spinlock), it seems
better to not call the init function either, least of all to keep
behaviour consistent.
ip_set_list calls ip_set_init_comment() only from uadt cb, it can be
kept as-is.
This was triggered by yet another LLM review, hinting that the existing
rcu_dereference_protected() cannot be downgraded to only check if the
nfnl mutex is held.
Fixes: f30415929be8 ("netfilter: ipset: do not update comments from kernel-side hash adds")
Signed-off-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
|
|
modify_prefix_route() is given the lifetime in clock_t relative to now,
but fib6_set_expires() wants an absolute jiffies value. So when a
permanent address is changed to a finite valid_lft, the prefix route
ends up already expired and GC removes it.
Steps to reproduce:
ip link add dummy9 type dummy
ip link set dummy9 up
ip -6 addr add 2001:db8:9::1/64 dev dummy9
ip -6 addr change 2001:db8:9::1/64 dev dummy9 valid_lft 3600 preferred_lft 3600
ip -6 route show dev dummy9 # expires is negative
Fixes: 8308f3ff1753 ("net/ipv6: Add support for specifying metric of connected routes")
Signed-off-by: Qishuai Liu <lqs@lqs.me>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260923132048.3272878-1-lqs@lqs.me
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Two racing tc filter add requests on the same chain/prio of an
unlocked classifier both run change() on the shared proto and both
can fail: the winner's tcf_chain_tp_delete_empty() attempt gives up
because the loser's handle is still in the idr, and the loser's
error path drops only its own reference without a second reclamation
attempt. The empty proto stays linked in the chain, holding the
chain reference, a block reference and the classifier module
reference until the chain or block is torn down.
Reclaim the proto on the error path of any failed request that holds
a proto reference. The reclamation is emptiness-gated:
tcf_chain_tp_delete_empty() unlinks the proto only when
delete_empty() admits it is empty, so a live shared proto is never
unlinked. A proto the request created is reclaimed unconditionally -
it is the only owner, so marking it for deletion is safe even without
a delete_empty callback. Classifiers without one (the check marks the
proto unconditionally) are rtnl-serialized, so the raced window this
guard closes cannot arise for them.
This is a follow-up to commit d4e359b3608a ("net/sched: cls_api: fix
teardown of an adopted proto on insert-race loss"), which stopped the
loser of the insert race from unlinking the winner's live proto but
left the empty-proto residual in place.
Conditions to recreate:
- CONFIG_NET_CLS_FLOWER=y; veth pair
- tc qdisc add dev veth0 ingress
- two concurrent `tc filter add dev veth0 ingress protocol ip pref 1
flower skip_sw ... action drop` (both fail in fl_hw_replace_filter
after publishing their handle in the idr); repeat in a loop
- an empty flower tp stays linked after both requests fail; visible
as a bare `filter protocol ip pref 1 flower chain 0` header in
`tc filter show` with no filter entries
- CAP_NET_ADMIN (namespace-local via unshare -Urn suffices)
Fixes: 8b64678e0af8 ("net: sched: refactor tp insert/delete for concurrent execution")
Reported-by: Sashiko (nipa) <sashiko-bot@kernel.org>
Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260805134049.927864-1-victor@mojatatu.com
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-JCOT.v1.20260910090924@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|