aboutsummaryrefslogtreecommitdiffstatshomepage
AgeCommit message (Collapse)AuthorFilesLines
65 min.wireguard: noise: reject response consumption after intermediate initiationHEADstableJason A. Donenfeld1-3/+4
Two threads begin processing the identical response message, received twice. The first thread, A, runs. While it's running, the second one, B, gets partway through, and during that slow calculation, or even while blocking on down_write(), A completes and then also a handshake initiation that's already been queued up runs in thread C, which itself takes that same down_write(). The handshake initiation creation succeeds, and sets the state back to waiting-for-response, and calls up_write(), at which point thread B resumes, because either its finished its calculations or was finally allowed to acquire down_write(). Thread B then copies the state back to the peer, and begins a new session, using that state, which is the same session as the one made in thread A. Thread A Thread B Thread C down_read() sA = handshake->state memcpy(cA, handshake->crypto) up_read() if (sA != 1) goto fail slow_crypto(cA) down_read() sB = handshake->state memcpy(cB, handshake->crypto) up_read() if (sB != 1) goto fail slow_crypto(cB) down_write() if (sA != handshake->state) goto fail memcpy(handshake->crypto, cA) handshake->state = 2 up_write() down_write() if (handshake->state != 2) goto fail derive_session(handshake->crypto) up_write() down_write() slow_crypto(handshake->crypto) handshake->state = 1 up_write() down_write() if (sB != handshake->state) goto fail memcpy(handshake->crypto, cB) handshake->state = 2 up_write() down_write() if (handshake->state != 2) goto fail derive_session(handshake->crypto) up_write() This seems basically impossible to hit in a meaningful way in practice, but ensure that it absolutely cannot happen by comparing the ephemeral private key that's on the stack with the latest one that the peer's handshake state has. Cc: stable@vger.kernel.org Fixes: e7096c131e51 ("net: WireGuard secure network tunnel") Reported-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr> Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
47 hourswireguard: queueing: preserve tstamp_type when encapsulating packetRamses de Norre1-0/+2
Sending traffic through a wireguard tunnel on a host using the fq qdisc fills the log with: fq: likely mono tstamp with tstamp_type 0 An skb carries a timestamp in skb->tstamp and, separately, a skb->tstamp_type field recording which clock that timestamp came from. The two have to agree. When wireguard encapsulates a packet it calls wg_reset_packet(), which clears the fields that must not leak from the inner packet into the tunnel packet. It does so in two steps: skb_scrub_packet(skb, true); memset(&skb->headers, 0, sizeof(skb->headers)); skb_scrub_packet() deliberately keeps skb->tstamp when it holds a monotonic timestamp: that value is the time the packet is scheduled to be sent, and the qdisc still needs it. The memset then zeroes skb->tstamp_type, because that field sits inside the headers group while skb->tstamp does not. The packet therefore leaves wireguard carrying a monotonic timestamp labelled as a realtime one. Nothing noticed until commit c4f796c4f16b ("net_sched: sch_fq: convert skb->tstamp if not monotonic"): fq used to assume every timestamp was monotonic. It now consults tstamp_type, spots the mismatch, warns, and falls back to treating the value as monotonic. Pacing still ends up correct, so the log spam is the actual problem. Save tstamp_type before the memset and restore it when encapsulating, next to the hash fields that are already carried over this way. When decapsulating it stays zeroed, which is right: an incoming packet's timestamp is a realtime receive timestamp. Fixes: d98d58a00261 ("net: Set skb->mono_delivery_time and clear it after sch_handle_ingress()") Signed-off-by: Ramses de Norre <ramses@well-founded.dev> Reviewed-by: Toke Høiland-Jørgensen <toke@kernel.org> Reviewed-by: Eric Dumazet <edumazet@google.com> Cc: stable@vger.kernel.org Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
47 hourswireguard: noise: remove unused variableJason A. Donenfeld1-2/+0
This is a harmless artifact from early pre-release wireguard development. These days, static_private is read from a shared data structure under a read lock directly, and it doesn't need to be copied to the stack. So remove the unused variable. Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
3 daysRevert "net: stmmac: propagate PTP addend and system time programming errors"davem/netJakub Kicinski2-84/+32
This reverts commit 232d49dd4b40a666283de9e722899f088ed581b2. Anirudh reports a regression from this change. It is not appropriate for net this late in the release cycle in the first place, so let's revert and revisit in net-next. Reported-by: Anirudh Srinivasan <asrinivasan@oss.tenstorrent.com> Link: https://lore.kernel.org/CAEev2e8+fr1tJnSErD+2TtLDw5fGJbQOMcM5eefMnLyFUR+m7w@mail.gmail.com Fixes: 232d49dd4b40 ("net: stmmac: propagate PTP addend and system time programming errors") Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet: airoha: Add retry mechanism to airoha_qdma_set_trtcm_param()Leto Liu (刘涛)2-6/+28
CPU accesses QDMA via the bus. When multiple modules are using the bus simultaneously, CPU access to QDMA may encounter bus timeouts and fails, resulting in QDMA configuration failures and potentially causing packet transmission issues. In order to mitigate the issue, introduce a retry mechanism to airoha_qdma_set_trtcm_param routine in order to ensure the configuration is correctly applied to the hardware. It's enough to try a second time for the TRTCM config to be actually applied to make sure we are not in the middle of bucket handling as it does operate on fixed time slot internally. Fixes: ef1ca9271313b ("net: airoha: Add sched HTB offload support") Signed-off-by: Leto Liu (刘涛) <Leto.Liu@airoha.com> [ improve commit description, drop memory block, apply better loop logic] Signed-off-by: Christian Marangi <ansuelsmth@gmail.com> Link: https://patch.msgid.link/20260930065855.46171-1-ansuelsmth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysMerge branch 'gve-various-xdp-fixes'Jakub Kicinski2-5/+16
Joshua Washington says: ==================== gve: various XDP fixes This patch series consists of a number of small XDP-related fixes. These changes are split off from https://lore.kernel.org/20260922194533.631387-1-joshwash@google.com in an effort to prevent blocking relatively simple fixes on more complex fixes that will need more feedback to merge. A summary of the changes: 1) fix an issue where XSK buffers that have not been processed are leaked when disabling XSK pools 2) fix an XSK buffer leak when an RX error descriptor comes back from the hardware 3) fix a deadlock introduced by attempting to acquire the netdev lock after it has already been acquired ==================== Link: https://patch.msgid.link/20260930214735.3288780-1-joshwash@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysgve: fix napi_disable deadlock when attempting to disable XSK poolsJoshua Washington1-4/+4
When disabling XSK pools, GVE calls the unlocked versions of napi_disable and napi_enable. However, the netdev lock has already been acquired before ndo_bpf is called because GVE supports queue management ops. Calling the unlocked versions of napi_disable/enable results in a deadlock when attempting to disable XSK pools, as the thread attempts to re-acquire a lock it already holds. Update the NAPI calls to use the locked versions. This issue was found in production. Fixes: 606048cbd834 ("net: designate XSK pool pointers in queues as "ops protected"") Cc: stable@vger.kernel.org Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com> Reviewed-by: Jordan Rhee <jordanrhee@google.com> Signed-off-by: Joshua Washington <joshwash@google.com> Link: https://patch.msgid.link/20260930214735.3288780-4-joshwash@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysgve: fix XSK buffer leak on error descriptorJoshua Washington1-1/+6
When the error bit is set in the RX completion descriptor, the buf_state and its attached buffer should be freed. In the case of AF_XDP ZC, the XSK buffer was not freed, leading to a leak. This issue was caught by LLM. Fixes: c1fffc5d66a7 ("gve: implement DQO RX datapath and control path for AF_XDP zero-copy") Cc: stable@vger.kernel.org Reviewed-by: Jordan Rhee <jordanrhee@google.com> Reviewed-by: Tim Hostetler <thostet@google.com> Signed-off-by: Joshua Washington <joshwash@google.com> Link: https://patch.msgid.link/20260930214735.3288780-3-joshwash@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysgve: fix XSK buffer leak when rings are stoppedJoshua Washington1-0/+6
GVE does not free XSK buffers when resetting ring state as a part of stopping queues. This causes all XSK buffers which are posted to the NIC to be leaked. Free XSK buffers attached to an allocated buf_state when stopping rings. This issue was hit in production. Fixes: c1fffc5d66a7 ("gve: implement DQO RX datapath and control path for AF_XDP zero-copy") Cc: stable@vger.kernel.org Reviewed-by: Tim Hostetler <thostet@google.com> Reviewed-by: Jordan Rhee <jordanrhee@google.com> Signed-off-by: Joshua Washington <joshwash@google.com> Link: https://patch.msgid.link/20260930214735.3288780-2-joshwash@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysMerge branch 'net-stmmac-fix-double-vlan-802-1ad-tag-handling'Jakub Kicinski5-96/+63
Ovidiu Panait says: ==================== net: stmmac: Fix double VLAN 802.1ad tag handling Currently, hardware VLAN stripping is broken for 802.1ad tags. vlan_rx_hw() hardcodes ETH_P_8021Q when putting the hardware tag into the skb, rather than using the actual protocol from the packet. Because of this, packets that contain an 802.1ad outer tag are incorrectly passed up the stack as having an 802.1Q tag. This issue was observed on the Renesas RZ/V2H platform (which has a dwmac4 IP), when testing QinQ ping: # DUT ip link add link end0 name end0.100 type vlan proto 802.1ad id 100 ip link add link end0.100 name end0.100.200 type vlan proto 802.1q id 200 ip addr add 172.16.3.2/24 dev end0.100.200 ip link set end0 up ip link set end0.100 up ip link set end0.100.200 up # Peer ip link add link eth0 name eth0.100 type vlan proto 802.1ad id 100 ip link add link eth0.100 name eth0.100.200 type vlan proto 802.1q id 200 ip addr add 172.16.3.1/24 dev eth0.100.200 ip link set eth0 up ip link set eth0.100 up ip link set eth0.100.200 up ping 172.16.3.2 -- FAIL -- Note that this series only fixes the issue on dwmac4. dwxgmac2 has the same issue but I do not have access to hw to test on. Since dwmac4 does not expose the tag type in the RDES3 descriptor, it cannot support hardware S-Tag stripping correctly. This series disables S-tag stripping for it, so the 802.1ad tags are left in place and are handled by the software VLAN path. ==================== Link: https://patch.msgid.link/20260928203441.34876-1-ovidiu.panait.rb@renesas.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet: stmmac: Disable S-Tag processing on dwmac4Ovidiu Panait1-3/+7
Currently, hardware VLAN stripping is broken for 802.1ad tags. vlan_rx_hw() hardcodes ETH_P_8021Q when putting the hardware tag into the skb, rather than using the actual protocol from the packet. Because of this, packets that contain a 802.1ad outer tag are incorrectly passed up the stack as having an 802.1Q tag. This causes QinQ ping between two hosts to fail. vlan_rx_hw() is shared by dwxgmac2 and dwmac4: on dwxgmac2 the tag type is available in the RDES3 write-back descriptor (the ET_LT field), so the outer tag type can be determined based on that info. However, dwmac4 doesn't seem to provide the tag type. The Length/Type field in RDES3 only indicates whether the packet is single or double-tagged, not which tag type was stripped. Since dwmac4 cannot report the stripped tag type, it cannot support hardware S-Tag stripping correctly. Therefore, restrict the NETIF_F_HW_VLAN_STAG_RX and NETIF_F_HW_VLAN_STAG_FILTER advertisement to dwxgmac2 only. With this, 802.1ad tags are left in place and handled by the software VLAN path. Fixes: 750011e239a5 ("net: stmmac: Add support for HW-accelerated VLAN stripping") Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com> Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Link: https://patch.msgid.link/20260928203441.34876-6-ovidiu.panait.rb@renesas.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet: stmmac: Do not advertise S-VLAN stripping when it is disabledOvidiu Panait1-0/+7
C-VLAN and S-VLAN tag stripping are both controlled by the EVLS bit, so disabling rx-vlan-offload also disables S-VLAN tag stripping. However, rx-vlan-stag-hw-parse keeps being advertised as enabled: root@rzv2h-evk:~# ethtool -K end1 rx-vlan-offload off root@rzv2h-evk:~# ethtool -k end1 | grep -i vlan rx-vlan-offload: off tx-vlan-offload: off [fixed] rx-vlan-filter: on [fixed] vlan-challenged: off [fixed] tx-vlan-stag-hw-insert: off [fixed] rx-vlan-stag-hw-parse: on [fixed] rx-vlan-stag-filter: on [fixed] Fix this inconsistency by making NETIF_F_HW_VLAN_STAG_RX follow NETIF_F_HW_VLAN_CTAG_RX. Fixes: 750011e239a5 ("net: stmmac: Add support for HW-accelerated VLAN stripping") Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com> Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Link: https://patch.msgid.link/20260928203441.34876-5-ovidiu.panait.rb@renesas.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet: stmmac: Rename double VLAN references to svlanOvidiu Panait5-38/+38
The ESVL and DOVLTC bits control S-VLAN tag processing and have nothing to do with the double VLAN feature, which only provides a way to process an additional inner VLAN tag. However, the driver code that handles them always refers to "double VLAN", which is unrelated and makes the implementation confusing. The driver does not use any of the inner VLAN tag features, and the networking core does not support offloads for the inner tag anyway. To reduce the confusion regarding S-Tag vs double VLAN handling, rename double -> svlan. No functional change intended. Suggested-by: Joseph Steel <recv.jo@gmail.com> Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com> Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Link: https://patch.msgid.link/20260928203441.34876-4-ovidiu.panait.rb@renesas.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet: stmmac: Stop toggling the EDVLP bitOvidiu Panait1-8/+12
Currently, the double VLAN EDVLP bit is toggled whenever an 802.1ad VLAN is registered. This bit enables processing of the inner VLAN tag, which is completely unrelated to S-Tag VLAN handling. Move EDVLP handling into vlan_set_hw_mode() instead, and keep it always enabled, so that COE can work for packets with an inner VLAN header. Add a dedicated callback for dwxlgmac2, as it doesn't implement the set_hw_vlan_mode callback, like the other cores. Suggested-by: Joseph Steel <recv.jo@gmail.com> Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com> Link: https://patch.msgid.link/20260928203441.34876-3-ovidiu.panait.rb@renesas.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet: stmmac: Remove VLAN perfect matching dead codeOvidiu Panait3-52/+4
stmmac_vlan_update() falls back to "perfect matching" when the VLAN hash filter is unavailable (!priv->dma_cap.vlhash). This fallback has been unreachable in normal operation since its introduction in commit c7ab0b8088d7 ("net: stmmac: Fallback to VLAN Perfect filtering if HASH is not available") because the NETIF_F_HW_VLAN_{CTAG,STAG}_FILTER features are advertised only when priv->dma_cap.vlhash is true. The fallback is also duplicating the code in vlan_add_hw_rx_fltr(), which is always available since stmmac_get_num_vlan() returns at least 1. Therefore, remove it. Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com> Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Link: https://patch.msgid.link/20260928203441.34876-2-ovidiu.panait.rb@renesas.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysipv4: free inet_opt and ireq_opt after an RCU grace periodEric Dumazet2-2/+2
tcp_v4_syn_recv_sock() transfers ownership of ireq->ireq_opt to the child socket (newinet->inet_opt) without copying it. Another cpu can concurrently process a retransmitted SYN for the same request socket, and send a SYNACK from tcp_check_req(). tcp_v4_send_synack() and inet_csk_route_req() read ireq->ireq_opt under rcu_read_lock() only, and ip_build_and_send_pkt() and ip_options_build() then read opt->optlen twice. Note that the SYNACK timer itself is not an issue: it holds a reference on its request socket, and inet_csk_reqsk_queue_drop() calls timer_delete_sync() before the child can be freed. Since commit 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table"), request sockets are processed without holding the listener lock, so nothing prevents the child socket from being freed while the SYNACK is still being built. TCP child sockets do not have SOCK_RCU_FREE, and inet_sock_destruct() frees inet_opt with a plain kfree(), leading to a use-after-free in ip_options_build(). A similar issue exists with request socket migration (net.ipv4.tcp_migrate_req=1, or a BPF_SK_REUSEPORT_SELECT_OR_MIGRATE program). reqsk_timer_handler() clones the request socket with inet_reqsk_clone(), so that the old request socket and its clone share the same ireq_opt, then reqsk_migrate_reset() clears the pointer in the old one. Another cpu holding a reference on the old request socket can still be using these options (sending a SYNACK, or creating a child in tcp_v4_syn_recv_sock()) when the clone is freed, and tcp_v4_reqsk_destructor() also uses a plain kfree(). Readers already use RCU, and other paths replacing inet_opt (do_ip_setsockopt(), cipso_v4_sock_setattr()...) already use kfree_rcu(). Use kfree_rcu() in inet_sock_destruct() and tcp_v4_reqsk_destructor() as well. IPv6 is not affected by the first issue, because tcp_v6_syn_recv_sock() duplicates the options. tcp_v6_reqsk_destructor() has the same migration issue with ipv6_opt, which is only set by CALIPSO. This will be addressed in a separate patch. Fixes: 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table") Fixes: c905dee62232 ("tcp: Migrate TCP_NEW_SYN_RECV requests at retransmitting SYN+ACKs.") Reported-by: Xinyang Ge <xinyang@anthropic.com> Signed-off-by: Eric Dumazet <edumazet@kernel.org> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com> Link: https://patch.msgid.link/20261001221253.2964024-1-edumazet@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet/mlx5: Lag, split aggregate speed into oper and max helpersOr Har-Toov1-22/+57
mlx5_lag_sum_devices_speed computes the LAG aggregate by summing oper speeds across all ports. This has two bugs. First, it relies on the assumption that a port whose carrier is down will report an oper speed of zero and therefore not contribute to the sum. This assumption does not always hold: when the link partner disconnects the port transitions to DOWN state but firmware may still report a non-zero oper speed, causing the aggregate to include a port that is not actively carrying traffic. Second, in active-backup mode only one port transmits at a time, so the aggregate should reflect a single port speed rather than the sum of all ports. Fix this by splitting mlx5_lag_sum_devices_speed into two helpers. mlx5_lag_get_devices_oper_speed reflects the speed currently available: it queries the vport state of each port and skips any port that is not UP, rather than relying on oper speed being zero. mlx5_lag_get_devices_max_speed is state-independent and returns the maximum achievable speed used as a fallback when speed is 0; for active-backup it takes the maximum single-port speed instead of the sum. The bugs were discovered during integration and reproduced on a back-to-back ConnectX-8 setup. The first bug (stale oper speed when link goes down) was observed: after the link partner disconnected, the aggregate TX speed reported to vports remained full speed because firmware still returned the last oper speed, causing vports to be rate-limited to the old aggregate speed. The second bug (active-backup summing all port speeds instead of the active port's speed) was also triggered: in active-backup mode the computed aggregate was double the expected value since both port speeds were summed even though only one port was active. Fixes: 28ea6036dad2 ("net/mlx5: Handle port and vport speed change events in MPESW") Signed-off-by: Or Har-Toov <ohartoov@nvidia.com> Reviewed-by: Shay Drori <shayd@nvidia.com> Reviewed-by: Mark Bloch <mbloch@nvidia.com> Signed-off-by: Tariq Toukan <tariqt@nvidia.com> Link: https://patch.msgid.link/20260930170406.148548-1-tariqt@nvidia.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysMerge branch '40GbE' of git://git.kernel.org/pub/scm/linux/kernel/git/tnguy/net-queueJakub Kicinski5-45/+217
Tony Nguyen says: ==================== Fix i40e/ice/iavf VF bonding after netdev lock changes Jose Ignacio Tornos Martinez says: This series fixes VF bonding failures introduced by commit ad7c7b2172c3 ("net: hold netdev instance lock during sysfs operations"). When adding VFs to a bond immediately after setting trust mode, MAC address changes fail with -EAGAIN, preventing bonding setup. This affects both i40e (700-series) and ice (800-series) Intel NICs. The core issue is lock contention: iavf_set_mac() is now called with the netdev lock held and waits for MAC change completion while holding it. However, both the watchdog task that sends the request and the adminq_task that processes PF responses also need this lock, creating a deadlock where neither can run, causing timeouts. Additionally, setting VF trust triggers an unnecessary ~10 second VF reset in i40e driver that delays bonding setup, even though filter synchronization happens naturally during normal VF operation. For ice driver, the delay is not so big, but in the same way the operation is not necessary. This series: 1. Eliminates unnecessary VF reset when setting trust in i40e (reset only if revoking trust and VF has advanced features configured). 2. Fixes lock contention by polling admin queue synchronously 3. Eliminates unnecessary VF reset when setting trust in ice, (reset only if revoking trust and VF has advanced features configured). The key fix (patch 2/3) implements a synchronous MAC change operation similar to the approach used for ndo_change_mtu deadlock fix: https://lore.kernel.org/20260211191855.1532226-1-poros@redhat.com Instead of scheduling work and waiting, it: - Sends the virtchnl message directly (not via watchdog) - Polls the admin queue hardware directly for responses - Processes all messages inline (including non-MAC messages) - Returns when complete or times out This allows the operation to complete synchronously while holding netdev_lock, without relying on watchdog or adminq_task. The function can sleep for up to 2.5 seconds polling hardware, but this is acceptable since netdev_lock is per-device and only serializes operations on the same interface. Testing shows VF bonding now works reliably in ~5 seconds vs 15+ seconds before (i40e), without timeouts or errors (i40e and ice). Tested on Intel 700-series (i40e) and 800-series (ice) dual-port NICs with iavf driver. Thanks to Jan Tluka <jtluka@redhat.com> and Yuying Ma <yuma@redhat.com> for reporting the issues. * '40GbE' of git://git.kernel.org/pub/scm/linux/kernel/git/tnguy/net-queue: ice: skip unnecessary VF reset when setting trust iavf: send MAC change request synchronously i40e: skip unnecessary VF reset when setting trust ==================== Link: https://patch.msgid.link/20260928224454.483072-1-anthony.l.nguyen@intel.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysnet: make dev_xdp_sb_prog_count() see programs attached through a linkJakub Kicinski1-4/+7
ethtool refuses to enable tcp-data-split on a device running a single-buffer XDP program (which can't deal with the payload landing in a separate buffer). dev_xdp_sb_prog_count() only looks at xdp_state[].prog though, ignoring bpf_link integration. Get the prog pointer from dev_xdp_prog(), which will consult both. bpf_xdp_link_update() needs a bit of a touch-up, too, since ethtool calls dev_xdp_sb_prog_count() with just the instance lock (for ops-locked devices) - we now have to make sure the link updates happen under that lock. Spotted by AI while reviewing the XDP propagation series, verified and tested on netdevsim with a C test along these lines: link = bpf_program__attach_xdp(prog, ifindex); system("ethtool -G eth0 tcp-data-split on") Fixes: 197258f0ef68 ("net: ethtool: add hds_config member in ethtool_netdev_state") Acked-by: Stanislav Fomichev <sdf@fomichev.me> Link: https://patch.msgid.link/20261001012559.297751-1-kuba@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysopenvswitch: fix soft lockup in the netlink flow dumpDenis V. Lunev1-0/+8
A production compute node carrying a few thousand datapath flows hit a soft lockup inside a single netlink flow dump and panicked. ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it emits, and that releases stats->lock with spin_unlock_bh() once per CPU that has touched the flow. Every release is a local_bh_enable(), and each one runs the pending softirq backlog in the dumping thread's own context. The skb bounds the flows one callback emits, but not the softirq work it absorbs. On a CPU that carries the box's packet load the backlog refills as fast as it drains, so the dumping thread becomes that CPU's softirq engine. It never sleeps and it has no reschedule point, so under voluntary preemption nothing can take the CPU away from it: neither the ksoftirqd the kernel woke to take the work over, nor the stopper thread the softlockup detector dispatches to refresh its timestamp. Hold BH off across the whole callback instead, the way ctnetlink_dump_table() does, so the nested spin_unlock_bh() stop draining softirqs. The loop already runs under rcu_read_lock() and cannot sleep. What it gives up is preemption under CONFIG_PREEMPT, since a BH-off region is not preemptible outside PREEMPT_RT. The region stays short: the skb caps the flows one callback emits, and empty buckets cost no skb space but are each visited once per dump, as the cursor only moves forward. The table grows on insert and shrinks only on flush, so the walk is bounded by the largest flow count the datapath has held. The softirq backlog the callback used to absorb has no bound at all. Fixes: 63e7959c4b9b ("openvswitch: Per NUMA node flow stats.") Cc: stable@vger.kernel.org Signed-off-by: Denis V. Lunev <den@openvz.org> Reviewed-by: Ilya Maximets <i.maximets@ovn.org> Signed-off-by: David S. Miller <davem@davemloft.net>
3 daysnet: bridge: avoid recursive multicast port cleanupZixuan Chai1-5/+18
br_multicast_sg_del_exclude_ports() removes automatically added STAR_EXCL port groups by calling br_multicast_del_pg(). For S,G entries, that deletion calls back into br_multicast_sg_del_exclude_ports(), so a large number of STAR_EXCL ports consumes one kernel stack frame per port and can hit the stack guard page. Keep the complete port-group deletion path for cleanup, but suppress the S,G exclude-port cleanup while that helper is already walking the same entry. Normal deletion callers retain the existing behavior. Fixes: 8266a0491e92 ("net: bridge: mcast: handle port group filter modes") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Assisted-by: LLM Signed-off-by: Zixuan Chai <petalzu987@gmail.com> Signed-off-by: Ren Wei <weir@nebusec.ai> Signed-off-by: David S. Miller <davem@davemloft.net>
6 daysMerge branch 'intel-wired-lan-driver-updates-2026-09-28-idpf-ice-iavf'Jakub Kicinski3-4/+4
Tony Nguyen says: ==================== Intel Wired LAN Driver Updates 2026-09-28 (ice) [part] For ice: Xuanqiang Luo resolves a use-after-free issue by changing order of operations so index is used before being freed. Bryan Fraschetti restores ordered MMIO writes for ice Tx doorbells by replacing writel_relaxed() call with writel(). Tristan Madani fixes representor use-after-free by releasing metadata_dst through dst_release() to ensure it is not freed until all references are dropped. ==================== Link: https://patch.msgid.link/20260928230429.495442-1-anthony.l.nguyen@intel.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
6 daysice: fix metadata_dst refcount handling on representor teardownTristan Madani1-1/+1
ice_eswitch_release_repr() uses metadata_dst_free() to release the representor's metadata_dst. metadata_dst_free() directly frees the underlying memory without checking the dst_entry refcount. When ice_eswitch_port_start_xmit() processes a packet, it takes a reference via dst_hold() and attaches the metadata_dst to the skb. If the representor is torn down while packets are still queued on the lower device (e.g. in a qdisc), the metadata_dst is freed while references are still held. Use dst_release() instead, which correctly decrements the refcount and only frees the object when all references are dropped. The dst subsystem already handles metadata_dst cleanup in dst_destroy() when DST_METADATA is set. Other drivers sharing this pattern (nfp, airoha, bnxt) already use dst_release() for their metadata_dst lifecycle. Fixes: f5396b8a663f7 ("ice: switchdev slow path") Cc: stable@vger.kernel.org Signed-off-by: Tristan Madani <tristan@talencesecurity.com> Reviewed-by: Simon Horman <horms@kernel.org> Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com> Link: https://patch.msgid.link/20260928230429.495442-5-anthony.l.nguyen@intel.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
6 daysice: Restore Ordered MMIO Writes for Tx DoorbellsBryan Fraschetti1-2/+2
The DQL accounting state which tracks the number of bytes queued for the NIC to transmit, namely dql->num_queued, is updated by the ICE driver when it invokes __netdev_tx_sent_queue(). Subsequently the driver updates the hardware queue tail allowing the NIC to begin processing the work queue. The queue accounting update should be ordered before the NIC begins transmitting descriptors. The transmit path currently updates the hardware doorbell using writel_relaxed(), which (on arm64) does not provide the same ordering guarantees between writes to normal memory and writes to MMIO registers that writel() does. This introduces a potential race where the NIC begins transmitting descriptors before the dql->num_queued update is globally visible. If the NIC finishes before the update is observed by dql_completed(), it detects an invalid state where more bytes have been completed than have been queued. When this happens the following BUG_ON is triggered. BUG_ON(count > num_queued - dql->num_completed); This has been observed and manifests as the following crash (note that the trace has been trimmed) in an environment with sustained network load that uses an Intel Corporation Ethernet Controller E810-XXV for SFP (rev 02) on an arm64 machine. Replacing writel_relaxed() with writel() in a test kernel eliminated the crash in the user's workload, which previously reproduced the issue reliably. kernel BUG at lib/dynamic_queue_limits.c:99 Internal error: Oops - BUG: 00000000f2000800 [#1] SMP pc : dql_completed+0x268/0x2a0 lr : ice_clean_tx_irq+0x1d4/0x620 [ice] Call trace: dql_completed+0x268/0x2a0 (P) ice_napi_poll+0x94/0x520 [ice] __napi_poll+0x48/0x3f0 net_rx_action+0x194/0x420 This restores the behaviour prior to commit ccde82e90946 ("ice: add E830 Earliest TxTime First Offload support"), which changed the notification mechanism from writel() to writel_relaxed(). Link: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2161572 Fixes: ccde82e90946 ("ice: add E830 Earliest TxTime First Offload support") Signed-off-by: Bryan Fraschetti <bryan.fraschetti@canonical.com> Tested-by: Alexander Nowlin <alexander.nowlin@intel.com> Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com> Link: https://patch.msgid.link/20260928230429.495442-4-anthony.l.nguyen@intel.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
6 daysice: fix use-after-free in dynamic port cleanupXuanqiang Luo1-1/+1
ice_dealloc_dynamic_port() uses dyn_port->vsi->idx to erase the dynamic port from pf->dyn_ports. However, it frees the VSI before reading the index for the erase, resulting in a use-after-free. Follow the reverse of the allocation order in ice_alloc_dynamic_port() by erasing the xarray entry before freeing the VSI. Fixes: eda69d654c7e ("ice: add basic devlink subfunctions support") Cc: stable@vger.kernel.org Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn> Reviewed-by: Marcin Szycik <marcin.szycik@linux.intel.com> Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com> Tested-by: Patryk Holda <patryk.holda@intel.com> Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com> Link: https://patch.msgid.link/20260928230429.495442-3-anthony.l.nguyen@intel.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
6 daystcp: reject net_iov in zerocopy receive mapping hintsDaehyeon Ko1-0/+2
After copying a readable prefix, receive_fallback_to_copy() asks tcp_zerocopy_set_hint_for_skb() where page mapping can resume. If the next skb is unreadable, find_next_mappable_frag() passes its net_iov fragment to can_map_frag(). skb_frag_page() returns NULL for a net_iov, but can_map_frag() dereferences it in PageCompound(). A v7.2 KASAN run on a connected TCP socket with 64 readable bytes followed by a 4096-byte NET_IOV_DMABUF fragment reported: BUG: KASAN: null-ptr-deref in can_map_frag tcp_zerocopy_receive -> can_map_frag Kernel panic - not syncing: KASAN: panic_on_warn set The diagnostic inserted the net_iov directly because the test host has no devmem-capable NIC. Hardware end-to-end reachability remains untested and requires CONFIG_NET_DEVMEM plus a supported DMA-buf-bound RX queue. Reject all net_iov fragments before skb_frag_page(). This covers both DMABUF and IOURING net_iov types while leaving page-backed checks unchanged. With the guard, the same queue copied the readable prefix, returned a 4096-byte skip hint, and completed without a fault. Fixes: 9f6b619edf2e ("net: support non paged skb frags") Cc: stable@vger.kernel.org Signed-off-by: Daehyeon Ko <4ncienth@gmail.com> Reviewed-by: Mina Almasry <almasrymina@google.com> Link: https://patch.msgid.link/20260930102524.1659847-1-4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
6 daysnet: xilinx: axienet: Free outstanding DMA buffers on dmaengine stopSuraj Gupta2-11/+47
In the dmaengine path the driver pre-submits RX buffers and holds in-flight TX buffers whose SKBs are DMA-mapped by the driver and freed only in the completion callbacks. On ndo_stop() dmaengine_terminate_sync() aborts these descriptors without running their callbacks, and the driver then frees only the ring shells, leaking every SKB still owned by the engine and its DMA mapping on each ifdown. With 128 RX buffers pre-posted per channel, the mapping leak can eventually exhaust a limited IOMMU aperture. Clear the slot's skb in the TX and RX callbacks so a non-NULL skb marks a slot that still owns a live, DMA-mapped buffer, and on stop unmap and free every such buffer. axienet_dma_rx_cb() runs from the DMA tasklet and re-arms the RX ring on each completion, so it can race axienet_stop(): a completion may submit a fresh buffer after dmaengine_terminate_sync() has returned, leaving the channel armed with a buffer the teardown then frees while the engine may still write into it (dma_release_channel() does not stop it either). Add a lock that axienet_dma_rx_cb() holds across the @stopping check and the resubmit, and axienet_stop() holds to set @stopping before terminating. Once @stopping is set no callback can arm a new buffer, and any armed just before is aborted by the terminate, so teardown only frees buffers the engine no longer owns. Fixes: 6a91b846af85 ("net: axienet: Introduce dmaengine support") Cc: stable@vger.kernel.org Signed-off-by: Suraj Gupta <suraj.gupta2@amd.com> Link: https://patch.msgid.link/20260928184207.2361931-1-suraj.gupta2@amd.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
6 daysbonding: fix the program leak when XDP is replacedJakub Kicinski1-4/+4
We need to release the reference on the old BPF prog, whether we are going to no-prog, or replacing the prog with another. Spotted by AI while running tests for the propagation series (leaked programs were accumulating). Selftest will come in the net-next series to avoid conflicts. Fixes: 9e2ee5c7e7c3 ("net, bonding: Add XDP support to the bonding driver") Fixes: 6d5f1ef83868 ("bonding: Fix negative jump label count on nested bonding") Acked-by: Daniel Borkmann <daniel@iogearbox.net> Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev> Link: https://patch.msgid.link/20261001012639.298690-1-kuba@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
6 daysnet: allwinner: remove dmaengine_desc_free()Frank Li1-4/+1
dmaengine_desc_free() is designed for DMA_CTRL_REUSE. If DMA_CTRL_REUSE is not set, dmaengine_desc_free() will return -EPERM. if (!dmaengine_desc_test_reuse(desc)) return -EPERM; So calling dmaengine_desc_free() does nothing. dmaengine_terminate_(a)sync() will free all already allocated descriptors. Signed-off-by: Frank Li <Frank.Li@nxp.com> Reviewed-by: Joe Damato <joe@dama.to> Signed-off-by: David S. Miller <davem@davemloft.net>
6 daysnet: stmmac: propagate PTP addend and system time programming errorsLorenzo Bianconi2-32/+84
stmmac_update_subsecond_increment() ignores the error returned by stmmac_config_addend(), and stmmac_init_tstamp_counter() discards the addend and system time programming errors, always returning success. A failure to program the addend (PTP_TCR_TSADDREG) or to initialize the system time counter (PTP_TCR_TSINIT) is therefore silently swallowed, leaving the hardware timestamp counter in a non-running or partially configured state while the driver keeps operating as if timestamping were up. This matters for TAPRIO/EST offloading, which derives the gate base time from the hardware timestamp counter. The same hooks are also called from the PHC callbacks: settime64 and adjfine drop the error and report success to clock_settime() and clock_adjtime(), so a dead PTP reference clock goes unnoticed by ptp4l/phc2sys. Return error codes from stmmac_update_subsecond_increment(), stmmac_init_tstamp_counter(), stmmac_dl_ts_coarse_set() and the settime64/adjfine callbacks instead of silently returning success. On failure, roll back the partially applied configuration so the hardware and the driver bookkeeping stay consistent, and report the reason through the devlink extack. Also guard against a zero sub-second increment, which would otherwise divide by zero when computing the addend. Reset the persistent timestamping state (hwts_tx_en, hwts_rx_en, tstamp_config, systime_flags and tsfupdt_coarse) when (re)initializing timestamping, so a failed init does not leave TX/RX timestamping enabled on a counter that never started. Fixes: cc4c9001ce31 ("net: stmmac: Switch stmmac_hwtimestamp to generic HW Interface Helpers") Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com> Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Signed-off-by: David S. Miller <davem@davemloft.net>
6 dayscxgb4/ch_ktls: disable softirqs around tid_list erasesSang-Hoon Choi1-2/+2
chcr_ktls_cpl_act_open_rpl() inserts into tid_list from the receive path using xa_insert_bh(). The array is shared by the adapter's connections and initialized with XA_FLAGS_LOCK_BH. However, chcr_ktls_dev_del() and the chcr_ktls_dev_add() error path use xa_erase(), which takes the XArray lock without disabling softirqs. When cleanup runs with softirqs enabled, a receive softirq for another connection on the same CPU can interrupt the erase and spin on the lock held by the interrupted task. Use xa_erase_bh() at both sites to match the insertion path. Fixes: 65e302a9bd57 ("cxgb4/ch_ktls: Clear resources when pf4 device is removed") Reported-by: Changyul Lee <lcy8047@gmail.com> Assisted-by: LLM Signed-off-by: Sang-Hoon Choi <csh0052@gmail.com> Reviewed-by: Ayush Sawal <ayush.sawal@chelsio.com> Signed-off-by: David S. Miller <davem@davemloft.net>
6 daystipc: destroy topsrv workqueues before closing connectionsYuqi Xu1-15/+65
tipc_topsrv_stop() closed subscriber connections while the topology server's send and receive workqueues were still running. Socket callbacks and subscription events could then queue more work, and in-flight send/recv work could drop the last connection reference during the conn_idr walk. That race produced several teardown failures: refcount_t addition on 0 from conn_get() on a connection whose release was blocked on idr_lock, a subsequent use-after-free in tipc_conn_close(), queue_work() on an already destroyed workqueue from the listener data-ready callback, and an RCU stall in tipc_topsrv_exit_net() while the walk spun under idr_lock. Clear srv->listener under idr_lock so it acts as a shutdown flag, skip queue_work() once it is NULL, destroy the workqueues to flush in-flight work, and only then close the remaining connections. Refuse tipc_conn_lookup() after that flag is cleared, and drop idr_lock when the teardown walk finds no connection, so an in-flight subscription event cannot pin idr_in_use while the walk holds the lock. v3 supersedes the narrower in-thread diff Tung Quang Nguyen posted on 2026-09-23. It keeps his teardown order and adds the lookup refusal and the empty-idr unlock. Fixes: c5fa7b3cf3cb ("tipc: introduce new TIPC server infrastructure") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Assisted-by: LLM Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com> Reviewed-by: Ren Wei <weir@nebusec.ai> Signed-off-by: David S. Miller <davem@davemloft.net>
6 daysMerge branch 'dim-overflow'David S. Miller4-3/+150
Shashank Mohan Jain says: ==================== lib/dim: fix 32-bit overflow in dim_calc_stats() On 32-bit kernels dim_calc_stats() multiplies the u32 byte, packet and completion counts of a DIM window by USEC_PER_MSEC (1000L) in 32-bit long arithmetic. Once a window carries more than about 4.3 MB, bpms wraps, and net_dim steers interrupt moderation on a meaningless throughput value. At 1 Gbit/s line rate a 64-event window passes that size when there are fewer than about 1,860 DIM events per second, which is common while NAPI keeps the interrupt masked under load. 32-bit users of the library include mtk_eth_soc (MT7621, MT7623), bcmgenet and bcmsysport on 32-bit ARM, xilinx_axienet on Zynq-7000 and MicroBlaze, and virtio_net in 32-bit guests. The overflow goes back to the mlx5e code the library was moved from. Patch 1 does the multiplications in 64 bits and divides with DIV_ROUND_UP_ULL(); the results on 64-bit are unchanged. Patch 2 adds a KUnit suite for dim_calc_stats() whose large-window cases fail on 32-bit without patch 1. It is part of this series as described under "Co-posting selftests" in maintainer-netdev.rst. The series is based on net (a7bfaba4823e) and has no dependencies; both patches also apply to mainline (fd179f8a05be) and net-next. The bug was found and the patches were prepared with Claude Code (Anthropic), model Claude Opus 5.5 (claude-opus-5-5). Tested: - KUnit (CONFIG_DIMLIB_KUNIT_TEST=y) on UML i386 (SUBARCH=i386): without patch 1, 4 of the 7 dim_calc_stats cases fail (many_bytes, bytes_32bit_limit, gigabit, many_packets); with it all pass. On UML x86_64 all cases pass with and without patch 1. Both were run on mainline and on net. - W=1 builds of lib/dim/ for UML x86_64 and i386 without warnings; dim.o references no libgcc 64-bit division helpers. A native i386 defconfig build (vmlinux and modules, DIMLIB=y) succeeds. - checkpatch --strict. Its "does MAINTAINERS need updating?" warning on patch 2 does not apply: lib/dim/ is already covered by the DYNAMIC INTERRUPT MODERATION entry. Not tested: 32-bit ARM or MIPS builds (no cross compiler was available), and no run on a 32-bit NIC. The traffic levels at which the drivers hit the overflow are derived from how they count DIM events, not measured. Signed-off-by: David S. Miller <davem@davemloft.net>
6 dayslib/dim: add KUnit test for dim_calc_stats()Shashank Mohan Jain3-0/+143
Add a KUnit suite for the DIM library that checks the packet, byte, event and completion rates computed by dim_calc_stats(), including counter wraparound and windows whose byte or packet count times USEC_PER_MSEC does not fit in 32 bits. The latter cases fail on 32-bit architectures without the previous commit. Assisted-by: LLM Signed-off-by: Shashank Mohan Jain <jain.sm@gmail.com> Signed-off-by: David S. Miller <davem@davemloft.net>
6 dayslib/dim: fix 32-bit overflow in dim_calc_stats() ratesShashank Mohan Jain1-3/+7
dim_calc_stats() computes the per-millisecond rates as DIV_ROUND_UP(nbytes * USEC_PER_MSEC, delta_us) where nbytes is a u32 and USEC_PER_MSEC is 1000L. On 64-bit the product is done in 64-bit long arithmetic, but on 32-bit architectures long is 32 bits wide and the product wraps as soon as a measurement window carries more than 4294967 bytes (about 4.3 MB). The same applies to the packet and completion counts, although those need more than 4.29 million packets or completions per window. A DIM window spans DIM_NEVENTS (64) events. Drivers count events per interrupt or per NAPI poll, so under sustained load a window can easily carry more than 4.3 MB: 64 full NAPI polls of 64 MTU-sized frames are already 6.2 MB, and drivers such as mtk_eth_soc count one event per interrupt while NAPI keeps polling with the interrupt masked. On 32-bit users of the library (for example mtk_eth_soc on MT7621, bcmgenet and bcmsysport on 32-bit ARM, or virtio_net in a 32-bit guest) bpms then becomes the product modulo 2^32 divided by the window length, and net_dim_stats_compare() makes its BETTER/WORSE decisions on a value that has little to do with the real throughput. For example, a 1 Gbit/s link at line rate that moves 5 MB in a 40 ms window gives bpms = 125000 on 64-bit but 17626 on 32-bit, and 5 million packets in 2 s gives ppms = 353 instead of 2500. Widen the products to 64 bits and divide with DIV_ROUND_UP_ULL(). The results are unchanged on 64-bit. Fixes: cb3c7fd4f839 ("net/mlx5e: Support adaptive RX coalescing") Fixes: 4c4dbb4a7363 ("net/mlx5e: Move dynamic interrupt coalescing code to include/linux") Assisted-by: LLM Signed-off-by: Shashank Mohan Jain <jain.sm@gmail.com> Signed-off-by: David S. Miller <davem@davemloft.net>
7 daysnet/packet: guard the ll header push in packet_rcv_spkt()Quchaosheng1-1/+2
packet_rcv_spkt() restores the link layer header with skb_push(skb, skb->data - skb_mac_header(skb)); That subtraction is only meaningful when the device actually has a link layer header. packet_rcv() and tpacket_rcv() both wrap it in dev_has_header(), which is also the predicate the block comment at the top of the file states the restore in terms of; packet_rcv_spkt() does not. commit d549699048b4 ("net/packet: fix packet receive on L3 devices without visible hard header") introduced the helper and changed the two call sites, and this one stayed behind. A device without a visible ll header can leave skb->mac_header at the 0xFFFF sentinel that __alloc_skb() initialises it to. A CAN skb does: init_can_skb() sets pkt_type and ip_summed but does not reset the headers, and commit 9f10374bb024 ("can: remove private CAN skb headroom infrastructure") dropped the skb_reset_*_header() calls that used to be there. skb_mac_header() is then 0xFFFF, the length becomes a large negative number and skb_push() reports it through skb_under_panic() -- from softirq context, so it is a full system panic even with panic_on_oops=0: skbuff: skb_under_panic: text:ffffffff8bd21bc1 len:-65455 put:-65471 head:... data:... tail:0x50 end:0x180 dev:can0 kernel BUG at net/core/skbuff.c:214! RIP: 0010:skb_panic+0x50/0x60 Call Trace: <IRQ> skb_push+0x38/0x40 packet_rcv_spkt+0xe1/0x170 __netif_receive_skb_core.constprop.0+0x7e8/0xd30 ... Kernel panic - not syncing: Fatal exception in interrupt The socket type is reachable: packet_create() accepts SOCK_PACKET alongside SOCK_RAW and SOCK_DGRAM behind the same CAP_NET_RAW check, and neither the socket length nor a capability check keeps it away from a CAN interface. The missing skb_reset_*_header() calls in init_can_skb() are a regression in their own right and are being fixed separately, but a packet socket should not turn a link layer that did not initialise its mac header into a kernel panic. Guard the push the way the other two receive paths do. Tested on v7.3-rc5 under QEMU with a slcan device on a pty, which is the driver RX path: vcan does not reproduce it, because can_send() resets the headers on the way out. One SOCK_PACKET socket bound to can0 and one frame written into the line discipline panics an unpatched kernel with the trace above; the same image with this patch prints no panic and powers off normally. Both kernels are this tree, defconfig plus CONFIG_CAN_SLCAN=y, differing only in this hunk. Fixes: d549699048b4 ("net/packet: fix packet receive on L3 devices without visible hard header") Cc: stable@vger.kernel.org Signed-off-by: Quchaosheng <quchaosheng000406@163.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260928113108.2127215-1-quchaosheng000406@163.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysMerge branch 'amt-send-the-relay-s-general-query-directly-from-the-receive-path'Jakub Kicinski3-37/+44
Omar Ramadan says: ==================== amt: send the relay's General Query directly from the receive path The relay queues its General Query on the amt device with a raw tunnel pointer in skb->cb, and a query that waits in a qdisc can outlive its tunnel: a use-after-free in amt_dev_xmit(), reported by Microsoft with a KASAN reproducer. Patch 1 sends the query directly from amt_request_handler(), inside the RCU section that found or created the tunnel, so it never waits in a qdisc and nothing is stored in skb->cb. Patch 2 adds the selftest Taehee asked for. It counts the queries that leave the relay through its amt device and expects none. It fails without patch 1 and passes with it. ==================== Link: https://patch.msgid.link/20260928202312.74574-1-omar@blockcast.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysselftests: net: amt: check that the relay's queries bypass the amt deviceOmar Ramadan1-0/+29
The relay used to hand its General Queries to dev_queue_xmit() on the amt device, where a query could wait in a qdisc and outlive the tunnel it pointed to. The previous patch sends them directly from the receive path instead. Count the IGMP and MLD queries that leave the relay through amtr with tc flower filters on its egress, installed before the gateway comes up, and check that there are none. The forwarding tests before it already show that the gateway received its queries, since it cannot join without one. Without the previous patch the new test fails (one run counted 7 IGMP and 6 MLD queries); with it, all of amt.sh passes. Signed-off-by: Omar Ramadan <omar@blockcast.net> Link: https://patch.msgid.link/20260928202312.74574-3-omar@blockcast.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysamt: send the relay's General Query directly from the receive pathOmar Ramadan2-37/+15
An skb queued in a qdisc can outlive the tunnel it references through a raw pointer in skb->cb. For example, with igmp_qrv set to 1 on the relay a tunnel lives for 135s, so a netem delay of 180s on the amt device outlives it; when the tunnel expires and is freed, the subsequent dequeue triggers a use-after-free in amt_dev_xmit(). BUG: KASAN: slab-use-after-free in amt_dev_xmit+0x2763/0x2e20 Call Trace: amt_dev_xmit+0x2763/0x2e20 [drivers/net/amt.c:1262] dev_hard_start_xmit+0x22f/0x620 sch_direct_xmit+0x12e/0xac0 netem_dequeue+0x333/0xc50 net_tx_action+0x35c/0xa60 amt_send_igmp_gq() and amt_send_mld_gq() are only called from amt_request_handler(), inside the rcu_read_lock_bh() section of amt_rcv(). amt_request_handler() already has the tunnel the query is for: it found or created it inside that section. Queuing the query with dev_queue_xmit() only leads back into amt_dev_xmit(), which strips the Ethernet header and calls amt_send_membership_query() for that tunnel. Make that call directly from the two senders instead, the same way amt_send_advertisement() transmits from the receive path. The query never waits in a qdisc, the tunnel is only dereferenced inside the RCU section that found or created it, and nothing is stored in skb->cb, so no lookup or refcount is needed. Remove the query branch of amt_dev_xmit(), amt_skb_cb() and struct amt_skb_cb, which have no users left. Behaviour changes: - The relay's own General Queries no longer pass through the amt device's egress path: its qdisc, tc egress (clsact/tcx), the netfilter egress hook and packet taps. They are still visible as UDP on the underlay. - A query that is sent successfully is no longer counted as tx_dropped. The old query branch left through the unlock label, which counted every sent query as dropped. - A query that reaches amt_dev_xmit() on a relay from elsewhere, such as a userspace querier, is now dropped at the IGMP/MLD type switch. Before, it trusted whatever skb->cb held, and a NULL tunnel hit the WARN_ON(1). Fixes: cbc21dc1cfe9 ("amt: add data plane of amt interface") Reported-by: AutonomousCodeSecurity@microsoft.com Reported-by: Xiang Mei (Microsoft) <xmei5@asu.edu> Reported-by: Cen Zhang (Microsoft Security FORGE Labs) <cenzhang@linux.microsoft.com> Signed-off-by: Omar Ramadan <omar@blockcast.net> Link: https://patch.msgid.link/20260928202312.74574-2-omar@blockcast.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysnet/sched: em_text: reject unterminated algo before autoloadJamal Hadi Salim1-0/+3
This is a follow-up to commit 9c572a83037a ("net/sched: fix potential stack infoleak in em_text_dump()"), which zero-initialised the dump-side struct. Sashiko pointed at a pre-existing bug that exposed the change-side input-validation gap open. em_text_change() casts the raw netlink attribute payload to struct tcf_em_text and passes conf->algo, a 16-byte array with no enforced NUL terminator, to textsearch_prepare(). On the TS_AUTOLOAD retry textsearch_prepare() calls request_module("ts_%s", algo); vsnprintf() walks %s until it finds a NUL, so an algo[] with no NUL reads past the array into the adjacent struct fields and payload and feeds those bytes to the usermode helper command line. This is an out-of-bounds read and a kernel memory disclosure. Reject the attribute when no NUL exists within TC_EM_TEXT_ALGOSIZ bytes, matching the check xt_string has had since 3ab720881b6e. Conditions to recreate the bug: - CAP_NET_ADMIN on the target netns (namespace-local via unshare -Urn is enough) - install clsact on an interface, then add a basic filter with one text ematch whose algo[] is 16 bytes with no NUL (e.g. "A"*16) and pattern_len at least 1 - the add fails with -ENOENT and the kernel invokes modprobe with a module name made of "ts_" plus the 16 bytes and the adjacent payload Fixes: d675c989ed2d4 ("[PKT_SCHED]: Packet classification based on textsearch (ematch)") Reported-by: Sashiko (gemini) <sashiko-bot@kernel.org> Closes: https://sashiko.dev/#/patchset/20260918133953.12494-1-bernard.ladenthin%40gmail.com Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Link: https://patch.msgid.link/QDISC-2DIU.v1.20260929183844@mojatatu.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daystipc: hold a reference to nodes found by link nameChengfeng Ye1-1/+9
tipc_node_find_by_name() returns a node after dropping its RCU read lock without taking a reference. The LINK_SET, LINK_GET and LINK_RESET_STATS handlers then lock and access the node, racing with timer-driven cleanup of a down peer. Generic netlink serialization does not exclude the node timer. The following interleaving can leave a handler using a freed node: CPU 0: find the node under RCU and release the node read lock CPU 1: tipc_node_timeout() clears the links and unlinks the down node CPU 1: drop the list and timer references, queuing tipc_node_free() CPU 0: leave the RCU read-side critical section CPU 1: complete the grace period and free the node CPU 0: acquire the node lock through the stale pointer LINK_SET also uses the node's media address after releasing the node lock, when passing queued packets to tipc_bearer_xmit(). KASAN on v7.3-rc5 reported: BUG: KASAN: slab-use-after-free in _raw_read_lock_bh Write of size 4 at addr ffff88807e06f008 by task poc/92 Call Trace: _raw_read_lock_bh kernel/locking/spinlock.c:287 tipc_nl_node_set_link net/tipc/node.c:2475 genl_family_rcv_msg_doit net/netlink/genetlink.c:1114 genl_rcv_msg net/netlink/genetlink.c:1209 netlink_rcv_skb net/netlink/af_netlink.c:2575 Allocated by task 0: tipc_node_create net/tipc/node.c:539 tipc_node_check_dest net/tipc/node.c:1196 tipc_disc_rcv net/tipc/discover.c:252 Freed by task 92: kfree mm/slub.c:6801 rcu_core kernel/rcu/tree.c:2919 Last potentially related work creation: __call_rcu_common.constprop.0 kernel/rcu/tree.c:3181 tipc_node_timeout net/tipc/node.c:814 Acquire a reference to the selected node with kref_get_unless_zero() before leaving RCU, returning NULL if the node has already been released. Release that reference on every caller exit after the last node access, including transmission in LINK_SET. Keep the existing link lookup order and locking so concurrent link removal still takes the existing error paths. Fixes: 6a939f365bdb ("tipc: Auto removal of peer down node instance") Cc: stable@vger.kernel.org Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com> Link: https://patch.msgid.link/20260928161246.1895273-1-nicoyip.dev@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysMerge branch 'ibmveth-fix-open-fail-unwind-after-lan-registration'Jakub Kicinski1-10/+12
Mingming Cao says: ==================== ibmveth: fix open-fail unwind after LAN registration ibmveth_open() has two independent unwind holes after the logical LAN is set up. Neither needs the MQ RX series. This posting is against current net.git. Patch 1 issues h_free_logical_lan() on every post-register failure. A pool allocation failure jumped to out_free_buffer_pools without the hypercall, so PHYP still owned the buffer-list page when it was unmapped. request_irq() failure already issued H_FREE before taking the same label. Fixes: d43732ce021f ("ibmveth: properly unwind on init errors") Patch 2 gives TX LTBs their own walk. out_free_buffer_pools reuses open()'s loop index, so after the pool unwind i is -1 and the TX buffers leak. request_irq() failure has the same leak. The TX-fail goto on that walk also skips dma_unmap of the filter list (pre-existing since the LTB was added). Fixes: d926793c1de9 ("ibmveth: Implement multi queue on xmit") The MQ RX series on net-next keeps the helper versions of these guards and does not depend on this pair. If both land, keep the helpers in that series; these two patches are the current single-queue open-fail path only. ==================== Link: https://patch.msgid.link/cover.1790357373.git.mmc@linux.ibm.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysibmveth: fix TX LTB and filter unwind on open-failMingming Cao1-6/+9
out_free_buffer_pools reuses open()'s loop index for the TX LTB walk. After the pool unwind i is -1, so the TX buffers allocated earlier are leaked. request_irq() failure has the same leak. The TX-fail goto on that walk also skips dma_unmap of the filter list (pre-existing since the LTB was added). Rewrite the TX walk to real_num_tx_queues with a pointer check, and unmap the filter list before those frees. Fixes: d926793c1de9 ("ibmveth: Implement multi queue on xmit") Cc: stable@vger.kernel.org Signed-off-by: Mingming Cao <mmc@linux.ibm.com> Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com> Link: https://patch.msgid.link/d14314c2b10f125f5e578782484ba3cc50564d53.1790357373.git.mmc@linux.ibm.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysibmveth: h_free logical LAN on open-fail after registerMingming Cao1-4/+3
ibmveth_open() registers the logical LAN, then allocates the RX buffer pools. A pool failure jumps to out_free_buffer_pools without h_free_logical_lan(), so PHYP still owns the buffer-list page when it is unmapped. request_irq() failure already issued the hypercall before taking the same label. Move that h_free to the shared post-register unwind so both paths deregister before the buffer list is unmapped. Fixes: d43732ce021f ("ibmveth: properly unwind on init errors") Cc: stable@vger.kernel.org Signed-off-by: Mingming Cao <mmc@linux.ibm.com> Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com> Link: https://patch.msgid.link/7a56fbfb9e0674f011f4e310fbd7dfd8e67c30aa.1790357373.git.mmc@linux.ibm.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
7 daysMerge tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/netLinus Torvalds111-344/+1441
Pull networking fixes from Paolo Abeni: "Including fixes from Bluetooth, WiFi and netfilter. We are actively retargeting several non-urgent fixes towards next, but the traffic on the ML looks ever-increasing, and propagating the push-back towards subsystems is not immediate. No known outstanding regressions. Current release - regressions: - netfilter: nft_set_rbtree: skip transaction elements during GC Previous releases - regressions: - sched: cls_api: reclaim an empty proto on the error path - core: - fix checksum offsets in skb_splice_from_iter() - cap skb->queue_mapping when the tx queue is picked - page_pool: fix use-after-free in page_pool_recycle_ring_bulk() - wifi: - mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue() - mac80211: drop oversized fragments to avoid extra_len overflow - netfilter: - flowtable: restore ieee80211 forward path - bluetooth: hci_conn: Lock parent access during enhanced SCO setup - eth: - bcmgenet: allocate RX buffers as page fragments - stmmac: fix rx Scatter-Gather support - octeontx2-pf: fix aura BPID assignment when CONFIG_DCB is enabled - gve: DQO: accept TSO packets with non-protocol gso_type bits - r8169: disable EEE on RTL8168h/8111h Previous releases - always broken: - tcp: refresh TS.Recent for accepted old ACKs - wifi: - ath11k: reset ar->num_stations on hardware start - cfg80211: fix RTS threshold setting for single-radio PHY - bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame() - eth: bcmgenet: fix NULL dereference in set_coalesce before first open Misc: - Eric is retiring from google and updating his contact info" * tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (96 commits) net: phy: aquantia: fix system interface type not updated in forced mode net: usb: qmi_wwan: add Rolling Wireless RN947R net: mvneta: clear XDP pfmemalloc flag between frames ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel() net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference r8169: disable EEE on RTL8168h/8111h octeontx2-pf: Fix RSS indirection table size sctp: check RCV_SHUTDOWN after the sendmsg connect wait net: sparx5: make ports inherit the switch base mac address type net: microchip: vcap: stop scanning after deleting key field netfilter: flowtable: restore ieee80211 forward path netfilter: flowtable: generalize pending status bit netfilter: bpf: reject invalid NAT manipulation types netfilter: nft_set_rbtree: skip transaction elements during GC ipvs: filter some flags received in the backup server ipvs: do not create invisible templates ipvs: bound LBLCR and LBLC cache growth ipvs: fix missing counter decrement in lblc netfilter: nft_flow_offload: drop flowtable reference on init error path selftests: net: check timestamp echo after an old ACK ...
7 daysMerge tag 'sysctl-7.03-fixes-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctlLinus Torvalds2-1/+3
Pull sysctl fixes from Joel Granados: "Fix sysctl jiffies conversions errors introduced in 2dc164a48e6f ("sysctl: Create converter functions with two new macros") - Ensure that mult_hz does *not* wrap - Ensure we pass just the magnitude for the negative branch in proc_int_k2u_conv_kop" * tag 'sysctl-7.03-fixes-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl: time/jiffies: Saturate in mult_hz() instead of wrapping sysctl: Negate before converting in the int read path
7 daysMerge tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nfPaolo Abeni16-19/+81
Pablo Neira Ayuso says: ==================== Netfilter/IPVS fixes for net The following batch contains Netfilter fixes for net. This batch fixes crashes as recent feature regression, one of the due to a dependency that has been pulled into -stable: 1) Expand existing ipset fix for bitmap sets to disallow comments updates from kernel-side adds, from Florian Westphal. 2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise flowtable cannot ever be removed, from Aohan Mei. 3) nft_rbtree GC should collect end elements that contained in this transaction batch, new or deleted elements are never expired. From Weiming Shi. 4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_* values, from Fernando F. Mancera. 5) Flowtable GC must skip flows that are pending hardware updates, generalize the PENDING flag and use it to inhibit GC. 6) Restore flowtable with ieee80211 which broke due to a relatively recent commit, which was pulled in by -stable, causing a regression in 6.18 kernels. And the following IPVS fixes: 1) Fix accounting of cache entries in IPVS LBLC for destinations, which eventually fills up the table and trigger recurrent resizing, from Julian Anastasov. 2) Limit IPVS cache growth for LBLCR and LBLC schedulers, from Zhiling Zou. 3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections, do not allow to use it with templates. Also from Julian. 4) Sanitize flags in IPVS sync messages received in the backup. From Julian Anastasov. netfilter pull request 26-09-30 * tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf: netfilter: flowtable: restore ieee80211 forward path netfilter: flowtable: generalize pending status bit netfilter: bpf: reject invalid NAT manipulation types netfilter: nft_set_rbtree: skip transaction elements during GC ipvs: filter some flags received in the backup server ipvs: do not create invisible templates ipvs: bound LBLCR and LBLC cache growth ipvs: fix missing counter decrement in lblc netfilter: nft_flow_offload: drop flowtable reference on init error path netfilter: ipset: do not update comments from kernel-side adds ==================== Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org Signed-off-by: Paolo Abeni <pabeni@redhat.com>
7 daysnet: phy: aquantia: fix system interface type not updated in forced modeBartosz Golaszewski1-1/+4
aqr_gen1_read_status() decodes the MDIO_PHYXS_VEND_IF_STATUS register to determine which SerDes interface the PHY is currently using on its system side and stores the result in phydev->interface. phylink relies on this value to configure the MAC. The autoneg == AUTONEG_DISABLE check is not correct: MDIO_PHYXS_VEND_IF_STATUS is set by the PHY firmware based on the negotiated link speed, not based on whether autoneg was used to reach it. When the link comes up at 1G in forced mode, the register correctly reads SGMII, but the early return prevents phydev->interface from being updated. It stays at whatever value it held before (typically 2500BASE-X from the initial autoneg run), so phylink configures the MAC for the wrong interface and the link cannot come up. Remove the autoneg guard so that the system interface type is always decoded when the link is up. Cc: stable@vger.kernel.org Fixes: 110a2432c520 ("net: phy: aquantia: add downshift support") Signed-off-by: Bartosz Golaszewski <bartosz.golaszewski@oss.qualcomm.com> Link: https://patch.msgid.link/20260923-qcom-sa8255p-emac-v15-1-e82f33720737@oss.qualcomm.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
7 daysMerge tag 'audit-pr-20260930' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/auditLinus Torvalds1-4/+20
Pull audit fix from Paul Moore: "A single audit fix for a potential UAF error in some audit filter configurations" * tag 'audit-pr-20260930' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/audit: audit: fix exe mark UAF in kill_rules()
8 daysnet: usb: qmi_wwan: add Rolling Wireless RN947RReinhard Speyerer1-0/+1
Add Rolling Wireless RN947R 0x9300 composition: DIAG + ADB + NMEA + MODEM + RMNET + QDSS + ADPL T: Bus=01 Lev=01 Prnt=01 Port=02 Cnt=02 Dev#= 6 Spd=480 MxCh= 0 D: Ver= 2.10 Cls=ef(misc ) Sub=02 Prot=01 MxPS=64 #Cfgs= 1 P: Vendor=33f8 ProdID=9300 Rev= 0.00 S: Manufacturer=Rolling Wireless S: Product=RN947R S: SerialNumber=xxxxxxxxxxx C:* #Ifs= 7 Cfg#= 1 Atr=a0 MxPwr=500mA I:* If#= 0 Alt= 0 #EPs= 2 Cls=ff(vend.) Sub=ff Prot=30 Driver=option E: Ad=01(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms E: Ad=81(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms I:* If#= 1 Alt= 0 #EPs= 2 Cls=ff(vend.) Sub=42 Prot=01 Driver=usbfs E: Ad=02(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms E: Ad=82(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms I:* If#= 2 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=40 Driver=option E: Ad=84(I) Atr=03(Int.) MxPS= 10 Ivl=32ms E: Ad=83(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms E: Ad=03(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms I:* If#= 3 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=40 Driver=option E: Ad=86(I) Atr=03(Int.) MxPS= 10 Ivl=32ms E: Ad=85(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms E: Ad=04(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms I:* If#= 8 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=50 Driver=qmi_wwan E: Ad=88(I) Atr=03(Int.) MxPS= 8 Ivl=32ms E: Ad=87(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms E: Ad=05(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms I:* If#=21 Alt= 0 #EPs= 1 Cls=ff(vend.) Sub=ff Prot=70 Driver=(none) E: Ad=89(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms I:* If#=22 Alt= 0 #EPs= 1 Cls=ff(vend.) Sub=ff Prot=80 Driver=(none) E: Ad=8a(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms Signed-off-by: Reinhard Speyerer <rspmn@arcor.de> Link: https://patch.msgid.link/arZzuOOvLUZtdOlv@arcor.de Signed-off-by: Jakub Kicinski <kuba@kernel.org>