| Age | Commit message (Collapse) | Author | Files | Lines |
|
Two threads begin processing the identical response message, received
twice. The first thread, A, runs. While it's running, the second one,
B, gets partway through, and during that slow calculation, or even while
blocking on down_write(), A completes and then also a handshake
initiation that's already been queued up runs in thread C, which itself
takes that same down_write(). The handshake initiation creation
succeeds, and sets the state back to waiting-for-response, and calls
up_write(), at which point thread B resumes, because either its finished
its calculations or was finally allowed to acquire down_write(). Thread
B then copies the state back to the peer, and begins a new session,
using that state, which is the same session as the one made in thread A.
Thread A Thread B Thread C
down_read()
sA = handshake->state
memcpy(cA, handshake->crypto)
up_read()
if (sA != 1)
goto fail
slow_crypto(cA)
down_read()
sB = handshake->state
memcpy(cB, handshake->crypto)
up_read()
if (sB != 1)
goto fail
slow_crypto(cB)
down_write()
if (sA != handshake->state)
goto fail
memcpy(handshake->crypto, cA)
handshake->state = 2
up_write()
down_write()
if (handshake->state != 2)
goto fail
derive_session(handshake->crypto)
up_write()
down_write()
slow_crypto(handshake->crypto)
handshake->state = 1
up_write()
down_write()
if (sB != handshake->state)
goto fail
memcpy(handshake->crypto, cB)
handshake->state = 2
up_write()
down_write()
if (handshake->state != 2)
goto fail
derive_session(handshake->crypto)
up_write()
This seems basically impossible to hit in a meaningful way in practice,
but ensure that it absolutely cannot happen by comparing the ephemeral
private key that's on the stack with the latest one that the peer's
handshake state has.
Cc: stable@vger.kernel.org
Fixes: e7096c131e51 ("net: WireGuard secure network tunnel")
Reported-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
|
|
Sending traffic through a wireguard tunnel on a host using the fq
qdisc fills the log with:
fq: likely mono tstamp with tstamp_type 0
An skb carries a timestamp in skb->tstamp and, separately, a
skb->tstamp_type field recording which clock that timestamp came from.
The two have to agree.
When wireguard encapsulates a packet it calls wg_reset_packet(), which
clears the fields that must not leak from the inner packet into the
tunnel packet. It does so in two steps:
skb_scrub_packet(skb, true);
memset(&skb->headers, 0, sizeof(skb->headers));
skb_scrub_packet() deliberately keeps skb->tstamp when it holds a
monotonic timestamp: that value is the time the packet is scheduled to
be sent, and the qdisc still needs it. The memset then zeroes
skb->tstamp_type, because that field sits inside the headers group
while skb->tstamp does not. The packet therefore leaves wireguard
carrying a monotonic timestamp labelled as a realtime one.
Nothing noticed until commit c4f796c4f16b ("net_sched: sch_fq: convert
skb->tstamp if not monotonic"): fq used to assume every timestamp was
monotonic. It now consults tstamp_type, spots the mismatch, warns, and
falls back to treating the value as monotonic. Pacing still ends up
correct, so the log spam is the actual problem.
Save tstamp_type before the memset and restore it when encapsulating,
next to the hash fields that are already carried over this way. When
decapsulating it stays zeroed, which is right: an incoming packet's
timestamp is a realtime receive timestamp.
Fixes: d98d58a00261 ("net: Set skb->mono_delivery_time and clear it after sch_handle_ingress()")
Signed-off-by: Ramses de Norre <ramses@well-founded.dev>
Reviewed-by: Toke Høiland-Jørgensen <toke@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
|
|
This is a harmless artifact from early pre-release wireguard
development. These days, static_private is read from a shared
data structure under a read lock directly, and it doesn't need to be
copied to the stack. So remove the unused variable.
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
|
|
This reverts commit 232d49dd4b40a666283de9e722899f088ed581b2.
Anirudh reports a regression from this change.
It is not appropriate for net this late in the release cycle
in the first place, so let's revert and revisit in net-next.
Reported-by: Anirudh Srinivasan <asrinivasan@oss.tenstorrent.com>
Link: https://lore.kernel.org/CAEev2e8+fr1tJnSErD+2TtLDw5fGJbQOMcM5eefMnLyFUR+m7w@mail.gmail.com
Fixes: 232d49dd4b40 ("net: stmmac: propagate PTP addend and system time programming errors")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
CPU accesses QDMA via the bus. When multiple modules are using the bus
simultaneously, CPU access to QDMA may encounter bus timeouts and fails,
resulting in QDMA configuration failures and potentially causing packet
transmission issues. In order to mitigate the issue, introduce a retry
mechanism to airoha_qdma_set_trtcm_param routine in order to ensure the
configuration is correctly applied to the hardware.
It's enough to try a second time for the TRTCM config to be actually
applied to make sure we are not in the middle of bucket handling as it does
operate on fixed time slot internally.
Fixes: ef1ca9271313b ("net: airoha: Add sched HTB offload support")
Signed-off-by: Leto Liu (刘涛) <Leto.Liu@airoha.com>
[ improve commit description, drop memory block, apply better loop logic]
Signed-off-by: Christian Marangi <ansuelsmth@gmail.com>
Link: https://patch.msgid.link/20260930065855.46171-1-ansuelsmth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Joshua Washington says:
====================
gve: various XDP fixes
This patch series consists of a number of small XDP-related fixes.
These changes are split off from
https://lore.kernel.org/20260922194533.631387-1-joshwash@google.com
in an effort to prevent blocking relatively simple fixes on more complex
fixes that will need more feedback to merge.
A summary of the changes:
1) fix an issue where XSK buffers that have not been processed are
leaked when disabling XSK pools
2) fix an XSK buffer leak when an RX error descriptor comes back from
the hardware
3) fix a deadlock introduced by attempting to acquire the netdev lock
after it has already been acquired
====================
Link: https://patch.msgid.link/20260930214735.3288780-1-joshwash@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When disabling XSK pools, GVE calls the unlocked versions of
napi_disable and napi_enable. However, the netdev lock has already been
acquired before ndo_bpf is called because GVE supports queue management
ops. Calling the unlocked versions of napi_disable/enable results in a
deadlock when attempting to disable XSK pools, as the thread attempts to
re-acquire a lock it already holds.
Update the NAPI calls to use the locked versions.
This issue was found in production.
Fixes: 606048cbd834 ("net: designate XSK pool pointers in queues as "ops protected"")
Cc: stable@vger.kernel.org
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Reviewed-by: Jordan Rhee <jordanrhee@google.com>
Signed-off-by: Joshua Washington <joshwash@google.com>
Link: https://patch.msgid.link/20260930214735.3288780-4-joshwash@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When the error bit is set in the RX completion descriptor, the buf_state
and its attached buffer should be freed. In the case of AF_XDP ZC, the
XSK buffer was not freed, leading to a leak.
This issue was caught by LLM.
Fixes: c1fffc5d66a7 ("gve: implement DQO RX datapath and control path for AF_XDP zero-copy")
Cc: stable@vger.kernel.org
Reviewed-by: Jordan Rhee <jordanrhee@google.com>
Reviewed-by: Tim Hostetler <thostet@google.com>
Signed-off-by: Joshua Washington <joshwash@google.com>
Link: https://patch.msgid.link/20260930214735.3288780-3-joshwash@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
GVE does not free XSK buffers when resetting ring state as a part of
stopping queues. This causes all XSK buffers which are posted to the
NIC to be leaked.
Free XSK buffers attached to an allocated buf_state when stopping rings.
This issue was hit in production.
Fixes: c1fffc5d66a7 ("gve: implement DQO RX datapath and control path for AF_XDP zero-copy")
Cc: stable@vger.kernel.org
Reviewed-by: Tim Hostetler <thostet@google.com>
Reviewed-by: Jordan Rhee <jordanrhee@google.com>
Signed-off-by: Joshua Washington <joshwash@google.com>
Link: https://patch.msgid.link/20260930214735.3288780-2-joshwash@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Ovidiu Panait says:
====================
net: stmmac: Fix double VLAN 802.1ad tag handling
Currently, hardware VLAN stripping is broken for 802.1ad tags. vlan_rx_hw()
hardcodes ETH_P_8021Q when putting the hardware tag into the skb, rather
than using the actual protocol from the packet. Because of this, packets
that contain an 802.1ad outer tag are incorrectly passed up the stack as
having an 802.1Q tag.
This issue was observed on the Renesas RZ/V2H platform (which has a dwmac4
IP), when testing QinQ ping:
# DUT
ip link add link end0 name end0.100 type vlan proto 802.1ad id 100
ip link add link end0.100 name end0.100.200 type vlan proto 802.1q id 200
ip addr add 172.16.3.2/24 dev end0.100.200
ip link set end0 up
ip link set end0.100 up
ip link set end0.100.200 up
# Peer
ip link add link eth0 name eth0.100 type vlan proto 802.1ad id 100
ip link add link eth0.100 name eth0.100.200 type vlan proto 802.1q id 200
ip addr add 172.16.3.1/24 dev eth0.100.200
ip link set eth0 up
ip link set eth0.100 up
ip link set eth0.100.200 up
ping 172.16.3.2
-- FAIL --
Note that this series only fixes the issue on dwmac4. dwxgmac2 has the same
issue but I do not have access to hw to test on.
Since dwmac4 does not expose the tag type in the RDES3 descriptor, it
cannot support hardware S-Tag stripping correctly. This series
disables S-tag stripping for it, so the 802.1ad tags are left in
place and are handled by the software VLAN path.
====================
Link: https://patch.msgid.link/20260928203441.34876-1-ovidiu.panait.rb@renesas.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Currently, hardware VLAN stripping is broken for 802.1ad tags. vlan_rx_hw()
hardcodes ETH_P_8021Q when putting the hardware tag into the skb, rather
than using the actual protocol from the packet. Because of this, packets
that contain a 802.1ad outer tag are incorrectly passed up the stack as
having an 802.1Q tag. This causes QinQ ping between two hosts to fail.
vlan_rx_hw() is shared by dwxgmac2 and dwmac4: on dwxgmac2 the tag type
is available in the RDES3 write-back descriptor (the ET_LT field), so the
outer tag type can be determined based on that info. However, dwmac4
doesn't seem to provide the tag type. The Length/Type field in RDES3 only
indicates whether the packet is single or double-tagged, not which tag
type was stripped.
Since dwmac4 cannot report the stripped tag type, it cannot support
hardware S-Tag stripping correctly. Therefore, restrict the
NETIF_F_HW_VLAN_STAG_RX and NETIF_F_HW_VLAN_STAG_FILTER advertisement
to dwxgmac2 only.
With this, 802.1ad tags are left in place and handled by the software
VLAN path.
Fixes: 750011e239a5 ("net: stmmac: Add support for HW-accelerated VLAN stripping")
Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260928203441.34876-6-ovidiu.panait.rb@renesas.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
C-VLAN and S-VLAN tag stripping are both controlled by the EVLS bit,
so disabling rx-vlan-offload also disables S-VLAN tag stripping.
However, rx-vlan-stag-hw-parse keeps being advertised as enabled:
root@rzv2h-evk:~# ethtool -K end1 rx-vlan-offload off
root@rzv2h-evk:~# ethtool -k end1 | grep -i vlan
rx-vlan-offload: off
tx-vlan-offload: off [fixed]
rx-vlan-filter: on [fixed]
vlan-challenged: off [fixed]
tx-vlan-stag-hw-insert: off [fixed]
rx-vlan-stag-hw-parse: on [fixed]
rx-vlan-stag-filter: on [fixed]
Fix this inconsistency by making NETIF_F_HW_VLAN_STAG_RX follow
NETIF_F_HW_VLAN_CTAG_RX.
Fixes: 750011e239a5 ("net: stmmac: Add support for HW-accelerated VLAN stripping")
Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260928203441.34876-5-ovidiu.panait.rb@renesas.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The ESVL and DOVLTC bits control S-VLAN tag processing and have
nothing to do with the double VLAN feature, which only provides a way
to process an additional inner VLAN tag. However, the driver code
that handles them always refers to "double VLAN", which is unrelated
and makes the implementation confusing. The driver does not use any
of the inner VLAN tag features, and the networking core does not
support offloads for the inner tag anyway.
To reduce the confusion regarding S-Tag vs double VLAN handling,
rename double -> svlan.
No functional change intended.
Suggested-by: Joseph Steel <recv.jo@gmail.com>
Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260928203441.34876-4-ovidiu.panait.rb@renesas.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Currently, the double VLAN EDVLP bit is toggled whenever an 802.1ad VLAN
is registered. This bit enables processing of the inner VLAN tag, which
is completely unrelated to S-Tag VLAN handling.
Move EDVLP handling into vlan_set_hw_mode() instead, and keep it always
enabled, so that COE can work for packets with an inner VLAN header.
Add a dedicated callback for dwxlgmac2, as it doesn't implement the
set_hw_vlan_mode callback, like the other cores.
Suggested-by: Joseph Steel <recv.jo@gmail.com>
Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com>
Link: https://patch.msgid.link/20260928203441.34876-3-ovidiu.panait.rb@renesas.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
stmmac_vlan_update() falls back to "perfect matching" when the VLAN hash
filter is unavailable (!priv->dma_cap.vlhash). This fallback has been
unreachable in normal operation since its introduction in
commit c7ab0b8088d7 ("net: stmmac: Fallback to VLAN Perfect filtering if
HASH is not available") because the NETIF_F_HW_VLAN_{CTAG,STAG}_FILTER
features are advertised only when priv->dma_cap.vlhash is true.
The fallback is also duplicating the code in vlan_add_hw_rx_fltr(), which
is always available since stmmac_get_num_vlan() returns at least 1.
Therefore, remove it.
Signed-off-by: Ovidiu Panait <ovidiu.panait.rb@renesas.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260928203441.34876-2-ovidiu.panait.rb@renesas.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
tcp_v4_syn_recv_sock() transfers ownership of ireq->ireq_opt to the
child socket (newinet->inet_opt) without copying it.
Another cpu can concurrently process a retransmitted SYN for the same
request socket, and send a SYNACK from tcp_check_req().
tcp_v4_send_synack() and inet_csk_route_req() read ireq->ireq_opt
under rcu_read_lock() only, and ip_build_and_send_pkt() and
ip_options_build() then read opt->optlen twice.
Note that the SYNACK timer itself is not an issue: it holds a
reference on its request socket, and inet_csk_reqsk_queue_drop()
calls timer_delete_sync() before the child can be freed.
Since commit 079096f103fa ("tcp/dccp: install syn_recv requests
into ehash table"), request sockets are processed without holding
the listener lock, so nothing prevents the child socket from being
freed while the SYNACK is still being built. TCP child sockets do
not have SOCK_RCU_FREE, and inet_sock_destruct() frees inet_opt
with a plain kfree(), leading to a use-after-free in
ip_options_build().
A similar issue exists with request socket migration
(net.ipv4.tcp_migrate_req=1, or a BPF_SK_REUSEPORT_SELECT_OR_MIGRATE
program). reqsk_timer_handler() clones the request socket with
inet_reqsk_clone(), so that the old request socket and its clone
share the same ireq_opt, then reqsk_migrate_reset() clears the
pointer in the old one. Another cpu holding a reference on the old
request socket can still be using these options (sending a SYNACK,
or creating a child in tcp_v4_syn_recv_sock()) when the clone is
freed, and tcp_v4_reqsk_destructor() also uses a plain kfree().
Readers already use RCU, and other paths replacing inet_opt
(do_ip_setsockopt(), cipso_v4_sock_setattr()...) already use
kfree_rcu(). Use kfree_rcu() in inet_sock_destruct() and
tcp_v4_reqsk_destructor() as well.
IPv6 is not affected by the first issue, because tcp_v6_syn_recv_sock()
duplicates the options. tcp_v6_reqsk_destructor() has the same
migration issue with ipv6_opt, which is only set by CALIPSO. This
will be addressed in a separate patch.
Fixes: 079096f103fa ("tcp/dccp: install syn_recv requests into ehash table")
Fixes: c905dee62232 ("tcp: Migrate TCP_NEW_SYN_RECV requests at retransmitting SYN+ACKs.")
Reported-by: Xinyang Ge <xinyang@anthropic.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20261001221253.2964024-1-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
mlx5_lag_sum_devices_speed computes the LAG aggregate by summing oper
speeds across all ports. This has two bugs.
First, it relies on the assumption that a port whose carrier is down
will report an oper speed of zero and therefore not contribute to the
sum. This assumption does not always hold: when the link partner
disconnects the port transitions to DOWN state but firmware may still
report a non-zero oper speed, causing the aggregate to include a port
that is not actively carrying traffic.
Second, in active-backup mode only one port transmits at a time, so
the aggregate should reflect a single port speed rather than the sum
of all ports.
Fix this by splitting mlx5_lag_sum_devices_speed into two helpers.
mlx5_lag_get_devices_oper_speed reflects the speed currently available:
it queries the vport state of each port and skips any port that is not
UP, rather than relying on oper speed being zero.
mlx5_lag_get_devices_max_speed is state-independent and returns the
maximum achievable speed used as a fallback when speed is 0; for
active-backup it takes the maximum single-port speed instead of the sum.
The bugs were discovered during integration and reproduced on a
back-to-back ConnectX-8 setup.
The first bug (stale oper speed when link goes down) was observed: after
the link partner disconnected, the aggregate TX speed reported to vports
remained full speed because firmware still returned the last oper speed,
causing vports to be rate-limited to the old aggregate speed.
The second bug (active-backup summing all port speeds instead of the
active port's speed) was also triggered: in active-backup mode the
computed aggregate was double the expected value since both port speeds
were summed even though only one port was active.
Fixes: 28ea6036dad2 ("net/mlx5: Handle port and vport speed change events in MPESW")
Signed-off-by: Or Har-Toov <ohartoov@nvidia.com>
Reviewed-by: Shay Drori <shayd@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260930170406.148548-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Tony Nguyen says:
====================
Fix i40e/ice/iavf VF bonding after netdev lock changes
Jose Ignacio Tornos Martinez says:
This series fixes VF bonding failures introduced by commit ad7c7b2172c3
("net: hold netdev instance lock during sysfs operations").
When adding VFs to a bond immediately after setting trust mode, MAC
address changes fail with -EAGAIN, preventing bonding setup. This
affects both i40e (700-series) and ice (800-series) Intel NICs.
The core issue is lock contention: iavf_set_mac() is now called with the
netdev lock held and waits for MAC change completion while holding it.
However, both the watchdog task that sends the request and the adminq_task
that processes PF responses also need this lock, creating a deadlock where
neither can run, causing timeouts.
Additionally, setting VF trust triggers an unnecessary ~10 second VF reset
in i40e driver that delays bonding setup, even though filter
synchronization happens naturally during normal VF operation. For ice
driver, the delay is not so big, but in the same way the operation is not
necessary.
This series:
1. Eliminates unnecessary VF reset when setting trust in i40e (reset only
if revoking trust and VF has advanced features configured).
2. Fixes lock contention by polling admin queue synchronously
3. Eliminates unnecessary VF reset when setting trust in ice, (reset only
if revoking trust and VF has advanced features configured).
The key fix (patch 2/3) implements a synchronous MAC change operation
similar to the approach used for ndo_change_mtu deadlock fix:
https://lore.kernel.org/20260211191855.1532226-1-poros@redhat.com
Instead of scheduling work and waiting, it:
- Sends the virtchnl message directly (not via watchdog)
- Polls the admin queue hardware directly for responses
- Processes all messages inline (including non-MAC messages)
- Returns when complete or times out
This allows the operation to complete synchronously while holding
netdev_lock, without relying on watchdog or adminq_task.
The function can sleep for up to 2.5 seconds polling hardware, but this
is acceptable since netdev_lock is per-device and only serializes
operations on the same interface.
Testing shows VF bonding now works reliably in ~5 seconds vs 15+ seconds
before (i40e), without timeouts or errors (i40e and ice).
Tested on Intel 700-series (i40e) and 800-series (ice) dual-port NICs
with iavf driver.
Thanks to Jan Tluka <jtluka@redhat.com> and Yuying Ma <yuma@redhat.com> for
reporting the issues.
* '40GbE' of git://git.kernel.org/pub/scm/linux/kernel/git/tnguy/net-queue:
ice: skip unnecessary VF reset when setting trust
iavf: send MAC change request synchronously
i40e: skip unnecessary VF reset when setting trust
====================
Link: https://patch.msgid.link/20260928224454.483072-1-anthony.l.nguyen@intel.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
ethtool refuses to enable tcp-data-split on a device running
a single-buffer XDP program (which can't deal with the payload
landing in a separate buffer). dev_xdp_sb_prog_count() only looks
at xdp_state[].prog though, ignoring bpf_link integration.
Get the prog pointer from dev_xdp_prog(), which will consult both.
bpf_xdp_link_update() needs a bit of a touch-up, too, since
ethtool calls dev_xdp_sb_prog_count() with just the instance
lock (for ops-locked devices) - we now have to make sure the
link updates happen under that lock.
Spotted by AI while reviewing the XDP propagation series,
verified and tested on netdevsim with a C test along these lines:
link = bpf_program__attach_xdp(prog, ifindex);
system("ethtool -G eth0 tcp-data-split on")
Fixes: 197258f0ef68 ("net: ethtool: add hds_config member in ethtool_netdev_state")
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Link: https://patch.msgid.link/20261001012559.297751-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A production compute node carrying a few thousand datapath flows hit a
soft lockup inside a single netlink flow dump and panicked.
ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it
emits, and that releases stats->lock with spin_unlock_bh() once per
CPU that has touched the flow. Every release is a local_bh_enable(),
and each one runs the pending softirq backlog in the dumping thread's
own context.
The skb bounds the flows one callback emits, but not the softirq work
it absorbs. On a CPU that carries the box's packet load the backlog
refills as fast as it drains, so the dumping thread becomes that CPU's
softirq engine. It never sleeps and it has no reschedule point, so
under voluntary preemption nothing can take the CPU away from it:
neither the ksoftirqd the kernel woke to take the work over, nor the
stopper thread the softlockup detector dispatches to refresh its
timestamp.
Hold BH off across the whole callback instead, the way
ctnetlink_dump_table() does, so the nested spin_unlock_bh() stop
draining softirqs. The loop already runs under rcu_read_lock() and
cannot sleep. What it gives up is preemption under CONFIG_PREEMPT,
since a BH-off region is not preemptible outside PREEMPT_RT. The
region stays short: the skb caps the flows one callback emits, and
empty buckets cost no skb space but are each visited once per dump, as
the cursor only moves forward. The table grows on insert and shrinks
only on flush, so the walk is bounded by the largest flow count the
datapath has held. The softirq backlog the callback used to absorb has
no bound at all.
Fixes: 63e7959c4b9b ("openvswitch: Per NUMA node flow stats.")
Cc: stable@vger.kernel.org
Signed-off-by: Denis V. Lunev <den@openvz.org>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
br_multicast_sg_del_exclude_ports() removes automatically added STAR_EXCL
port groups by calling br_multicast_del_pg(). For S,G entries, that
deletion calls back into br_multicast_sg_del_exclude_ports(), so a large
number of STAR_EXCL ports consumes one kernel stack frame per port and can
hit the stack guard page.
Keep the complete port-group deletion path for cleanup, but suppress the
S,G exclude-port cleanup while that helper is already walking the same
entry. Normal deletion callers retain the existing behavior.
Fixes: 8266a0491e92 ("net: bridge: mcast: handle port group filter modes")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Zixuan Chai <petalzu987@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
Tony Nguyen says:
====================
Intel Wired LAN Driver Updates 2026-09-28 (ice) [part]
For ice:
Xuanqiang Luo resolves a use-after-free issue by changing order of
operations so index is used before being freed.
Bryan Fraschetti restores ordered MMIO writes for ice Tx doorbells
by replacing writel_relaxed() call with writel().
Tristan Madani fixes representor use-after-free by releasing
metadata_dst through dst_release() to ensure it is not freed until
all references are dropped.
====================
Link: https://patch.msgid.link/20260928230429.495442-1-anthony.l.nguyen@intel.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
ice_eswitch_release_repr() uses metadata_dst_free() to release the
representor's metadata_dst. metadata_dst_free() directly frees the
underlying memory without checking the dst_entry refcount.
When ice_eswitch_port_start_xmit() processes a packet, it takes a
reference via dst_hold() and attaches the metadata_dst to the skb.
If the representor is torn down while packets are still queued on
the lower device (e.g. in a qdisc), the metadata_dst is freed while
references are still held.
Use dst_release() instead, which correctly decrements the refcount
and only frees the object when all references are dropped. The dst
subsystem already handles metadata_dst cleanup in dst_destroy() when
DST_METADATA is set.
Other drivers sharing this pattern (nfp, airoha, bnxt) already use
dst_release() for their metadata_dst lifecycle.
Fixes: f5396b8a663f7 ("ice: switchdev slow path")
Cc: stable@vger.kernel.org
Signed-off-by: Tristan Madani <tristan@talencesecurity.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
Link: https://patch.msgid.link/20260928230429.495442-5-anthony.l.nguyen@intel.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The DQL accounting state which tracks the number of bytes queued for the
NIC to transmit, namely dql->num_queued, is updated by the ICE driver
when it invokes __netdev_tx_sent_queue(). Subsequently the driver
updates the hardware queue tail allowing the NIC to begin processing the
work queue. The queue accounting update should be ordered before the NIC
begins transmitting descriptors.
The transmit path currently updates the hardware doorbell using
writel_relaxed(), which (on arm64) does not provide the same ordering
guarantees between writes to normal memory and writes to MMIO registers
that writel() does. This introduces a potential race where the NIC
begins transmitting descriptors before the dql->num_queued update is
globally visible. If the NIC finishes before the update is observed by
dql_completed(), it detects an invalid state where more bytes have been
completed than have been queued. When this happens the following BUG_ON
is triggered.
BUG_ON(count > num_queued - dql->num_completed);
This has been observed and manifests as the following crash (note that
the trace has been trimmed) in an environment with sustained network
load that uses an Intel Corporation Ethernet Controller E810-XXV for SFP
(rev 02) on an arm64 machine. Replacing writel_relaxed() with writel()
in a test kernel eliminated the crash in the user's workload, which
previously reproduced the issue reliably.
kernel BUG at lib/dynamic_queue_limits.c:99
Internal error: Oops - BUG: 00000000f2000800 [#1] SMP
pc : dql_completed+0x268/0x2a0
lr : ice_clean_tx_irq+0x1d4/0x620 [ice]
Call trace:
dql_completed+0x268/0x2a0 (P)
ice_napi_poll+0x94/0x520 [ice]
__napi_poll+0x48/0x3f0
net_rx_action+0x194/0x420
This restores the behaviour prior to commit ccde82e90946 ("ice: add E830
Earliest TxTime First Offload support"), which changed the notification
mechanism from writel() to writel_relaxed().
Link: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2161572
Fixes: ccde82e90946 ("ice: add E830 Earliest TxTime First Offload support")
Signed-off-by: Bryan Fraschetti <bryan.fraschetti@canonical.com>
Tested-by: Alexander Nowlin <alexander.nowlin@intel.com>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
Link: https://patch.msgid.link/20260928230429.495442-4-anthony.l.nguyen@intel.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
ice_dealloc_dynamic_port() uses dyn_port->vsi->idx to erase the dynamic
port from pf->dyn_ports. However, it frees the VSI before reading the
index for the erase, resulting in a use-after-free.
Follow the reverse of the allocation order in ice_alloc_dynamic_port()
by erasing the xarray entry before freeing the VSI.
Fixes: eda69d654c7e ("ice: add basic devlink subfunctions support")
Cc: stable@vger.kernel.org
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Marcin Szycik <marcin.szycik@linux.intel.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Tested-by: Patryk Holda <patryk.holda@intel.com>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
Link: https://patch.msgid.link/20260928230429.495442-3-anthony.l.nguyen@intel.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
After copying a readable prefix, receive_fallback_to_copy() asks
tcp_zerocopy_set_hint_for_skb() where page mapping can resume. If the
next skb is unreadable, find_next_mappable_frag() passes its net_iov
fragment to can_map_frag().
skb_frag_page() returns NULL for a net_iov, but can_map_frag()
dereferences it in PageCompound(). A v7.2 KASAN run on a connected TCP
socket with 64 readable bytes followed by a 4096-byte NET_IOV_DMABUF
fragment reported:
BUG: KASAN: null-ptr-deref in can_map_frag
tcp_zerocopy_receive -> can_map_frag
Kernel panic - not syncing: KASAN: panic_on_warn set
The diagnostic inserted the net_iov directly because the test host has
no devmem-capable NIC. Hardware end-to-end reachability remains untested
and requires CONFIG_NET_DEVMEM plus a supported DMA-buf-bound RX queue.
Reject all net_iov fragments before skb_frag_page(). This covers both
DMABUF and IOURING net_iov types while leaving page-backed checks
unchanged. With the guard, the same queue copied the readable prefix,
returned a 4096-byte skip hint, and completed without a fault.
Fixes: 9f6b619edf2e ("net: support non paged skb frags")
Cc: stable@vger.kernel.org
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Reviewed-by: Mina Almasry <almasrymina@google.com>
Link: https://patch.msgid.link/20260930102524.1659847-1-4ncienth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
In the dmaengine path the driver pre-submits RX buffers and holds
in-flight TX buffers whose SKBs are DMA-mapped by the driver and freed
only in the completion callbacks. On ndo_stop() dmaengine_terminate_sync()
aborts these descriptors without running their callbacks, and the driver
then frees only the ring shells, leaking every SKB still owned by the
engine and its DMA mapping on each ifdown. With 128 RX buffers pre-posted
per channel, the mapping leak can eventually exhaust a limited IOMMU
aperture.
Clear the slot's skb in the TX and RX callbacks so a non-NULL skb marks a
slot that still owns a live, DMA-mapped buffer, and on stop unmap and free
every such buffer.
axienet_dma_rx_cb() runs from the DMA tasklet and re-arms the RX ring on
each completion, so it can race axienet_stop(): a completion may submit a
fresh buffer after dmaengine_terminate_sync() has returned, leaving the
channel armed with a buffer the teardown then frees while the engine may
still write into it (dma_release_channel() does not stop it either). Add a
lock that axienet_dma_rx_cb() holds across the @stopping check and the
resubmit, and axienet_stop() holds to set @stopping before terminating.
Once @stopping is set no callback can arm a new buffer, and any armed just
before is aborted by the terminate, so teardown only frees buffers the
engine no longer owns.
Fixes: 6a91b846af85 ("net: axienet: Introduce dmaengine support")
Cc: stable@vger.kernel.org
Signed-off-by: Suraj Gupta <suraj.gupta2@amd.com>
Link: https://patch.msgid.link/20260928184207.2361931-1-suraj.gupta2@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
We need to release the reference on the old BPF prog, whether
we are going to no-prog, or replacing the prog with another.
Spotted by AI while running tests for the propagation series
(leaked programs were accumulating). Selftest will come in
the net-next series to avoid conflicts.
Fixes: 9e2ee5c7e7c3 ("net, bonding: Add XDP support to the bonding driver")
Fixes: 6d5f1ef83868 ("bonding: Fix negative jump label count on nested bonding")
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://patch.msgid.link/20261001012639.298690-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
dmaengine_desc_free() is designed for DMA_CTRL_REUSE. If DMA_CTRL_REUSE
is not set, dmaengine_desc_free() will return -EPERM.
if (!dmaengine_desc_test_reuse(desc))
return -EPERM;
So calling dmaengine_desc_free() does nothing.
dmaengine_terminate_(a)sync() will free all already allocated descriptors.
Signed-off-by: Frank Li <Frank.Li@nxp.com>
Reviewed-by: Joe Damato <joe@dama.to>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
stmmac_update_subsecond_increment() ignores the error returned by
stmmac_config_addend(), and stmmac_init_tstamp_counter() discards the
addend and system time programming errors, always returning success. A
failure to program the addend (PTP_TCR_TSADDREG) or to initialize the
system time counter (PTP_TCR_TSINIT) is therefore silently swallowed,
leaving the hardware timestamp counter in a non-running or partially
configured state while the driver keeps operating as if timestamping
were up. This matters for TAPRIO/EST offloading, which derives the gate
base time from the hardware timestamp counter.
The same hooks are also called from the PHC callbacks: settime64 and
adjfine drop the error and report success to clock_settime() and
clock_adjtime(), so a dead PTP reference clock goes unnoticed by
ptp4l/phc2sys.
Return error codes from stmmac_update_subsecond_increment(),
stmmac_init_tstamp_counter(), stmmac_dl_ts_coarse_set() and the
settime64/adjfine callbacks instead of silently returning success. On
failure, roll back the partially applied configuration so the hardware
and the driver bookkeeping stay consistent, and report the reason
through the devlink extack. Also guard against a zero sub-second
increment, which would otherwise divide by zero when computing the
addend.
Reset the persistent timestamping state (hwts_tx_en, hwts_rx_en,
tstamp_config, systime_flags and tsfupdt_coarse) when (re)initializing
timestamping, so a failed init does not leave TX/RX timestamping
enabled on a counter that never started.
Fixes: cc4c9001ce31 ("net: stmmac: Switch stmmac_hwtimestamp to generic HW Interface Helpers")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
chcr_ktls_cpl_act_open_rpl() inserts into tid_list from the receive path
using xa_insert_bh(). The array is shared by the adapter's connections
and initialized with XA_FLAGS_LOCK_BH. However, chcr_ktls_dev_del() and
the chcr_ktls_dev_add() error path use xa_erase(), which takes the XArray
lock without disabling softirqs.
When cleanup runs with softirqs enabled, a receive softirq for another
connection on the same CPU can interrupt the erase and spin on the lock
held by the interrupted task. Use xa_erase_bh() at both sites to match
the insertion path.
Fixes: 65e302a9bd57 ("cxgb4/ch_ktls: Clear resources when pf4 device is removed")
Reported-by: Changyul Lee <lcy8047@gmail.com>
Assisted-by: LLM
Signed-off-by: Sang-Hoon Choi <csh0052@gmail.com>
Reviewed-by: Ayush Sawal <ayush.sawal@chelsio.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
tipc_topsrv_stop() closed subscriber connections while the topology
server's send and receive workqueues were still running. Socket
callbacks and subscription events could then queue more work, and
in-flight send/recv work could drop the last connection reference
during the conn_idr walk.
That race produced several teardown failures: refcount_t addition on
0 from conn_get() on a connection whose release was blocked on
idr_lock, a subsequent use-after-free in tipc_conn_close(),
queue_work() on an already destroyed workqueue from the listener
data-ready callback, and an RCU stall in tipc_topsrv_exit_net()
while the walk spun under idr_lock.
Clear srv->listener under idr_lock so it acts as a shutdown flag,
skip queue_work() once it is NULL, destroy the workqueues to flush
in-flight work, and only then close the remaining connections.
Refuse tipc_conn_lookup() after that flag is cleared, and drop
idr_lock when the teardown walk finds no connection, so an
in-flight subscription event cannot pin idr_in_use while the
walk holds the lock.
v3 supersedes the narrower in-thread diff Tung Quang Nguyen
posted on 2026-09-23. It keeps his teardown order and adds the
lookup refusal and the empty-idr unlock.
Fixes: c5fa7b3cf3cb ("tipc: introduce new TIPC server infrastructure")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <weir@nebusec.ai>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
Shashank Mohan Jain says:
====================
lib/dim: fix 32-bit overflow in dim_calc_stats()
On 32-bit kernels dim_calc_stats() multiplies the u32 byte, packet and
completion counts of a DIM window by USEC_PER_MSEC (1000L) in 32-bit
long arithmetic. Once a window carries more than about 4.3 MB, bpms
wraps, and net_dim steers interrupt moderation on a meaningless
throughput value. At 1 Gbit/s line rate a 64-event window passes that
size when there are fewer than about 1,860 DIM events per second,
which is common while NAPI keeps the interrupt masked under load.
32-bit users of the library include mtk_eth_soc (MT7621, MT7623),
bcmgenet and bcmsysport on 32-bit ARM, xilinx_axienet on Zynq-7000 and
MicroBlaze, and virtio_net in 32-bit guests. The overflow goes back to
the mlx5e code the library was moved from.
Patch 1 does the multiplications in 64 bits and divides with
DIV_ROUND_UP_ULL(); the results on 64-bit are unchanged. Patch 2 adds a
KUnit suite for dim_calc_stats() whose large-window cases fail on
32-bit without patch 1. It is part of this series as described under
"Co-posting selftests" in maintainer-netdev.rst.
The series is based on net (a7bfaba4823e) and has no dependencies; both
patches also apply to mainline (fd179f8a05be) and net-next.
The bug was found and the patches were prepared with Claude Code
(Anthropic), model Claude Opus 5.5 (claude-opus-5-5).
Tested:
- KUnit (CONFIG_DIMLIB_KUNIT_TEST=y) on UML i386 (SUBARCH=i386):
without patch 1, 4 of the 7 dim_calc_stats cases fail (many_bytes,
bytes_32bit_limit, gigabit, many_packets); with it all pass. On UML
x86_64 all cases pass with and without patch 1. Both were run on
mainline and on net.
- W=1 builds of lib/dim/ for UML x86_64 and i386 without warnings;
dim.o references no libgcc 64-bit division helpers. A native i386
defconfig build (vmlinux and modules, DIMLIB=y) succeeds.
- checkpatch --strict. Its "does MAINTAINERS need updating?" warning
on patch 2 does not apply: lib/dim/ is already covered by the
DYNAMIC INTERRUPT MODERATION entry.
Not tested: 32-bit ARM or MIPS builds (no cross compiler was
available), and no run on a 32-bit NIC. The traffic levels at which
the drivers hit the overflow are derived from how they count DIM
events, not measured.
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
Add a KUnit suite for the DIM library that checks the packet, byte,
event and completion rates computed by dim_calc_stats(), including
counter wraparound and windows whose byte or packet count times
USEC_PER_MSEC does not fit in 32 bits. The latter cases fail on 32-bit
architectures without the previous commit.
Assisted-by: LLM
Signed-off-by: Shashank Mohan Jain <jain.sm@gmail.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
dim_calc_stats() computes the per-millisecond rates as
DIV_ROUND_UP(nbytes * USEC_PER_MSEC, delta_us)
where nbytes is a u32 and USEC_PER_MSEC is 1000L. On 64-bit the product
is done in 64-bit long arithmetic, but on 32-bit architectures long is
32 bits wide and the product wraps as soon as a measurement window
carries more than 4294967 bytes (about 4.3 MB). The same applies to the
packet and completion counts, although those need more than 4.29
million packets or completions per window.
A DIM window spans DIM_NEVENTS (64) events. Drivers count events per
interrupt or per NAPI poll, so under sustained load a window can easily
carry more than 4.3 MB: 64 full NAPI polls of 64 MTU-sized frames are
already 6.2 MB, and drivers such as mtk_eth_soc count one event per
interrupt while NAPI keeps polling with the interrupt masked. On 32-bit
users of the library (for example mtk_eth_soc on MT7621, bcmgenet and
bcmsysport on 32-bit ARM, or virtio_net in a 32-bit guest) bpms then
becomes the product modulo 2^32 divided by the window length, and
net_dim_stats_compare() makes its BETTER/WORSE decisions on a value
that has little to do with the real throughput.
For example, a 1 Gbit/s link at line rate that moves 5 MB in a 40 ms
window gives bpms = 125000 on 64-bit but 17626 on 32-bit, and 5 million
packets in 2 s gives ppms = 353 instead of 2500.
Widen the products to 64 bits and divide with DIV_ROUND_UP_ULL(). The
results are unchanged on 64-bit.
Fixes: cb3c7fd4f839 ("net/mlx5e: Support adaptive RX coalescing")
Fixes: 4c4dbb4a7363 ("net/mlx5e: Move dynamic interrupt coalescing code to include/linux")
Assisted-by: LLM
Signed-off-by: Shashank Mohan Jain <jain.sm@gmail.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
packet_rcv_spkt() restores the link layer header with
skb_push(skb, skb->data - skb_mac_header(skb));
That subtraction is only meaningful when the device actually has a link
layer header. packet_rcv() and tpacket_rcv() both wrap it in
dev_has_header(), which is also the predicate the block comment at the
top of the file states the restore in terms of; packet_rcv_spkt() does
not. commit d549699048b4 ("net/packet: fix packet receive on L3
devices without visible hard header") introduced the helper and changed
the two call sites, and this one stayed behind.
A device without a visible ll header can leave skb->mac_header at the
0xFFFF sentinel that __alloc_skb() initialises it to. A CAN skb does:
init_can_skb() sets pkt_type and ip_summed but does not reset the
headers, and commit 9f10374bb024 ("can: remove private CAN skb
headroom infrastructure") dropped the skb_reset_*_header() calls that
used to be there. skb_mac_header() is then 0xFFFF, the length becomes
a large negative number and skb_push() reports it through
skb_under_panic() -- from softirq context, so it is a full system panic
even with panic_on_oops=0:
skbuff: skb_under_panic: text:ffffffff8bd21bc1 len:-65455 put:-65471 head:... data:... tail:0x50 end:0x180 dev:can0
kernel BUG at net/core/skbuff.c:214!
RIP: 0010:skb_panic+0x50/0x60
Call Trace:
<IRQ>
skb_push+0x38/0x40
packet_rcv_spkt+0xe1/0x170
__netif_receive_skb_core.constprop.0+0x7e8/0xd30
...
Kernel panic - not syncing: Fatal exception in interrupt
The socket type is reachable: packet_create() accepts SOCK_PACKET
alongside SOCK_RAW and SOCK_DGRAM behind the same CAP_NET_RAW check,
and neither the socket length nor a capability check keeps it away
from a CAN interface.
The missing skb_reset_*_header() calls in init_can_skb() are a
regression in their own right and are being fixed separately, but a
packet socket should not turn a link layer that did not initialise its
mac header into a kernel panic. Guard the push the way the other two
receive paths do.
Tested on v7.3-rc5 under QEMU with a slcan device on a pty, which is
the driver RX path: vcan does not reproduce it, because can_send()
resets the headers on the way out. One SOCK_PACKET socket bound to
can0 and one frame written into the line discipline panics an
unpatched kernel with the trace above; the same image with this patch
prints no panic and powers off normally. Both kernels are this tree,
defconfig plus CONFIG_CAN_SLCAN=y, differing only in this hunk.
Fixes: d549699048b4 ("net/packet: fix packet receive on L3 devices without visible hard header")
Cc: stable@vger.kernel.org
Signed-off-by: Quchaosheng <quchaosheng000406@163.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260928113108.2127215-1-quchaosheng000406@163.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Omar Ramadan says:
====================
amt: send the relay's General Query directly from the receive path
The relay queues its General Query on the amt device with a raw tunnel
pointer in skb->cb, and a query that waits in a qdisc can outlive its
tunnel: a use-after-free in amt_dev_xmit(), reported by Microsoft with
a KASAN reproducer. Patch 1 sends the query directly from
amt_request_handler(), inside the RCU section that found or created
the tunnel, so it never waits in a qdisc and nothing is stored in
skb->cb.
Patch 2 adds the selftest Taehee asked for. It counts the queries that
leave the relay through its amt device and expects none. It fails
without patch 1 and passes with it.
====================
Link: https://patch.msgid.link/20260928202312.74574-1-omar@blockcast.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The relay used to hand its General Queries to dev_queue_xmit() on the
amt device, where a query could wait in a qdisc and outlive the tunnel
it pointed to. The previous patch sends them directly from the receive
path instead.
Count the IGMP and MLD queries that leave the relay through amtr with
tc flower filters on its egress, installed before the gateway comes
up, and check that there are none. The forwarding tests before it
already show that the gateway received its queries, since it cannot
join without one.
Without the previous patch the new test fails (one run counted 7 IGMP
and 6 MLD queries); with it, all of amt.sh passes.
Signed-off-by: Omar Ramadan <omar@blockcast.net>
Link: https://patch.msgid.link/20260928202312.74574-3-omar@blockcast.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
An skb queued in a qdisc can outlive the tunnel it references
through a raw pointer in skb->cb. For example, with igmp_qrv set
to 1 on the relay a tunnel lives for 135s, so a netem delay of
180s on the amt device outlives it; when the tunnel expires and is
freed, the subsequent dequeue triggers a use-after-free in
amt_dev_xmit().
BUG: KASAN: slab-use-after-free in amt_dev_xmit+0x2763/0x2e20
Call Trace:
amt_dev_xmit+0x2763/0x2e20 [drivers/net/amt.c:1262]
dev_hard_start_xmit+0x22f/0x620
sch_direct_xmit+0x12e/0xac0
netem_dequeue+0x333/0xc50
net_tx_action+0x35c/0xa60
amt_send_igmp_gq() and amt_send_mld_gq() are only called from
amt_request_handler(), inside the rcu_read_lock_bh() section of
amt_rcv(). amt_request_handler() already has the tunnel the query is
for: it found or created it inside that section. Queuing the query with
dev_queue_xmit() only leads back into amt_dev_xmit(), which strips the
Ethernet header and calls amt_send_membership_query() for that tunnel.
Make that call directly from the two senders instead, the same way
amt_send_advertisement() transmits from the receive path. The query
never waits in a qdisc, the tunnel is only dereferenced inside the
RCU section that found or created it, and nothing is stored in
skb->cb, so no lookup or refcount is needed. Remove the query branch
of amt_dev_xmit(), amt_skb_cb() and struct amt_skb_cb, which have no
users left.
Behaviour changes:
- The relay's own General Queries no longer pass through the amt
device's egress path: its qdisc, tc egress (clsact/tcx), the
netfilter egress hook and packet taps. They are still visible as
UDP on the underlay.
- A query that is sent successfully is no longer counted as
tx_dropped. The old query branch left through the unlock label,
which counted every sent query as dropped.
- A query that reaches amt_dev_xmit() on a relay from elsewhere,
such as a userspace querier, is now dropped at the IGMP/MLD type
switch. Before, it trusted whatever skb->cb held, and a NULL
tunnel hit the WARN_ON(1).
Fixes: cbc21dc1cfe9 ("amt: add data plane of amt interface")
Reported-by: AutonomousCodeSecurity@microsoft.com
Reported-by: Xiang Mei (Microsoft) <xmei5@asu.edu>
Reported-by: Cen Zhang (Microsoft Security FORGE Labs) <cenzhang@linux.microsoft.com>
Signed-off-by: Omar Ramadan <omar@blockcast.net>
Link: https://patch.msgid.link/20260928202312.74574-2-omar@blockcast.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
This is a follow-up to commit 9c572a83037a ("net/sched: fix potential
stack infoleak in em_text_dump()"), which zero-initialised the dump-side
struct. Sashiko pointed at a pre-existing bug that exposed the change-side
input-validation gap open.
em_text_change() casts the raw netlink attribute payload to struct
tcf_em_text and passes conf->algo, a 16-byte array with no enforced NUL
terminator, to textsearch_prepare(). On the TS_AUTOLOAD retry
textsearch_prepare() calls request_module("ts_%s", algo); vsnprintf()
walks %s until it finds a NUL, so an algo[] with no NUL reads past the
array into the adjacent struct fields and payload and feeds those bytes
to the usermode helper command line. This is an out-of-bounds read and
a kernel memory disclosure.
Reject the attribute when no NUL exists within TC_EM_TEXT_ALGOSIZ bytes,
matching the check xt_string has had since 3ab720881b6e.
Conditions to recreate the bug:
- CAP_NET_ADMIN on the target netns (namespace-local via unshare -Urn
is enough)
- install clsact on an interface, then add a basic filter with one text
ematch whose algo[] is 16 bytes with no NUL (e.g. "A"*16) and
pattern_len at least 1
- the add fails with -ENOENT and the kernel invokes modprobe with a
module name made of "ts_" plus the 16 bytes and the adjacent payload
Fixes: d675c989ed2d4 ("[PKT_SCHED]: Packet classification based on textsearch (ematch)")
Reported-by: Sashiko (gemini) <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260918133953.12494-1-bernard.ladenthin%40gmail.com
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2DIU.v1.20260929183844@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
tipc_node_find_by_name() returns a node after dropping its RCU read lock
without taking a reference. The LINK_SET, LINK_GET and LINK_RESET_STATS
handlers then lock and access the node, racing with timer-driven cleanup of
a down peer. Generic netlink serialization does not exclude the node
timer.
The following interleaving can leave a handler using a freed node:
CPU 0: find the node under RCU and release the node read lock
CPU 1: tipc_node_timeout() clears the links and unlinks the down node
CPU 1: drop the list and timer references, queuing tipc_node_free()
CPU 0: leave the RCU read-side critical section
CPU 1: complete the grace period and free the node
CPU 0: acquire the node lock through the stale pointer
LINK_SET also uses the node's media address after releasing the node lock,
when passing queued packets to tipc_bearer_xmit().
KASAN on v7.3-rc5 reported:
BUG: KASAN: slab-use-after-free in _raw_read_lock_bh
Write of size 4 at addr ffff88807e06f008 by task poc/92
Call Trace:
_raw_read_lock_bh kernel/locking/spinlock.c:287
tipc_nl_node_set_link net/tipc/node.c:2475
genl_family_rcv_msg_doit net/netlink/genetlink.c:1114
genl_rcv_msg net/netlink/genetlink.c:1209
netlink_rcv_skb net/netlink/af_netlink.c:2575
Allocated by task 0:
tipc_node_create net/tipc/node.c:539
tipc_node_check_dest net/tipc/node.c:1196
tipc_disc_rcv net/tipc/discover.c:252
Freed by task 92:
kfree mm/slub.c:6801
rcu_core kernel/rcu/tree.c:2919
Last potentially related work creation:
__call_rcu_common.constprop.0 kernel/rcu/tree.c:3181
tipc_node_timeout net/tipc/node.c:814
Acquire a reference to the selected node with kref_get_unless_zero() before
leaving RCU, returning NULL if the node has already been released. Release
that reference on every caller exit after the last node access, including
transmission in LINK_SET. Keep the existing link lookup order and locking
so concurrent link removal still takes the existing error paths.
Fixes: 6a939f365bdb ("tipc: Auto removal of peer down node instance")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Link: https://patch.msgid.link/20260928161246.1895273-1-nicoyip.dev@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Mingming Cao says:
====================
ibmveth: fix open-fail unwind after LAN registration
ibmveth_open() has two independent unwind holes after the logical LAN
is set up. Neither needs the MQ RX series. This posting is against
current net.git.
Patch 1 issues h_free_logical_lan() on every post-register failure.
A pool allocation failure jumped to out_free_buffer_pools without
the hypercall, so PHYP still owned the buffer-list page when it was
unmapped. request_irq() failure already issued H_FREE before taking
the same label.
Fixes: d43732ce021f ("ibmveth: properly unwind on init errors")
Patch 2 gives TX LTBs their own walk. out_free_buffer_pools reuses
open()'s loop index, so after the pool unwind i is -1 and the TX
buffers leak. request_irq() failure has the same leak. The TX-fail
goto on that walk also skips dma_unmap of the filter list
(pre-existing since the LTB was added).
Fixes: d926793c1de9 ("ibmveth: Implement multi queue on xmit")
The MQ RX series on net-next keeps the helper versions of these
guards and does not depend on this pair. If both land, keep the
helpers in that series; these two patches are the current
single-queue open-fail path only.
====================
Link: https://patch.msgid.link/cover.1790357373.git.mmc@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
out_free_buffer_pools reuses open()'s loop index for the TX LTB
walk. After the pool unwind i is -1, so the TX buffers allocated
earlier are leaked. request_irq() failure has the same leak.
The TX-fail goto on that walk also skips dma_unmap of the filter
list (pre-existing since the LTB was added). Rewrite the TX walk
to real_num_tx_queues with a pointer check, and unmap the filter
list before those frees.
Fixes: d926793c1de9 ("ibmveth: Implement multi queue on xmit")
Cc: stable@vger.kernel.org
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Link: https://patch.msgid.link/d14314c2b10f125f5e578782484ba3cc50564d53.1790357373.git.mmc@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
ibmveth_open() registers the logical LAN, then allocates the RX
buffer pools. A pool failure jumps to out_free_buffer_pools without
h_free_logical_lan(), so PHYP still owns the buffer-list page when
it is unmapped. request_irq() failure already issued the hypercall
before taking the same label.
Move that h_free to the shared post-register unwind so both paths
deregister before the buffer list is unmapped.
Fixes: d43732ce021f ("ibmveth: properly unwind on init errors")
Cc: stable@vger.kernel.org
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Link: https://patch.msgid.link/7a56fbfb9e0674f011f4e310fbd7dfd8e67c30aa.1790357373.git.mmc@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Pull networking fixes from Paolo Abeni:
"Including fixes from Bluetooth, WiFi and netfilter.
We are actively retargeting several non-urgent fixes towards next,
but the traffic on the ML looks ever-increasing, and propagating the
push-back towards subsystems is not immediate.
No known outstanding regressions.
Current release - regressions:
- netfilter: nft_set_rbtree: skip transaction elements during GC
Previous releases - regressions:
- sched: cls_api: reclaim an empty proto on the error path
- core:
- fix checksum offsets in skb_splice_from_iter()
- cap skb->queue_mapping when the tx queue is picked
- page_pool: fix use-after-free in page_pool_recycle_ring_bulk()
- wifi:
- mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()
- mac80211: drop oversized fragments to avoid extra_len overflow
- netfilter:
- flowtable: restore ieee80211 forward path
- bluetooth: hci_conn: Lock parent access during enhanced SCO setup
- eth:
- bcmgenet: allocate RX buffers as page fragments
- stmmac: fix rx Scatter-Gather support
- octeontx2-pf: fix aura BPID assignment when CONFIG_DCB is enabled
- gve: DQO: accept TSO packets with non-protocol gso_type bits
- r8169: disable EEE on RTL8168h/8111h
Previous releases - always broken:
- tcp: refresh TS.Recent for accepted old ACKs
- wifi:
- ath11k: reset ar->num_stations on hardware start
- cfg80211: fix RTS threshold setting for single-radio PHY
- bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
- eth: bcmgenet: fix NULL dereference in set_coalesce before first open
Misc:
- Eric is retiring from google and updating his contact info"
* tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (96 commits)
net: phy: aquantia: fix system interface type not updated in forced mode
net: usb: qmi_wwan: add Rolling Wireless RN947R
net: mvneta: clear XDP pfmemalloc flag between frames
ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel()
net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference
r8169: disable EEE on RTL8168h/8111h
octeontx2-pf: Fix RSS indirection table size
sctp: check RCV_SHUTDOWN after the sendmsg connect wait
net: sparx5: make ports inherit the switch base mac address type
net: microchip: vcap: stop scanning after deleting key field
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
selftests: net: check timestamp echo after an old ACK
...
|
|
Pull sysctl fixes from Joel Granados:
"Fix sysctl jiffies conversions errors introduced in 2dc164a48e6f
("sysctl: Create converter functions with two new macros")
- Ensure that mult_hz does *not* wrap
- Ensure we pass just the magnitude for the negative branch in
proc_int_k2u_conv_kop"
* tag 'sysctl-7.03-fixes-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl:
time/jiffies: Saturate in mult_hz() instead of wrapping
sysctl: Negate before converting in the int read path
|
|
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:
1) Expand existing ipset fix for bitmap sets to disallow comments
updates from kernel-side adds, from Florian Westphal.
2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
flowtable cannot ever be removed, from Aohan Mei.
3) nft_rbtree GC should collect end elements that contained in
this transaction batch, new or deleted elements are never
expired. From Weiming Shi.
4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_*
values, from Fernando F. Mancera.
5) Flowtable GC must skip flows that are pending hardware updates,
generalize the PENDING flag and use it to inhibit GC.
6) Restore flowtable with ieee80211 which broke due to a relatively
recent commit, which was pulled in by -stable, causing a regression
in 6.18 kernels.
And the following IPVS fixes:
1) Fix accounting of cache entries in IPVS LBLC for destinations,
which eventually fills up the table and trigger recurrent
resizing, from Julian Anastasov.
2) Limit IPVS cache growth for LBLCR and LBLC schedulers,
from Zhiling Zou.
3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections,
do not allow to use it with templates. Also from Julian.
4) Sanitize flags in IPVS sync messages received in the backup.
From Julian Anastasov.
netfilter pull request 26-09-30
* tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
netfilter: ipset: do not update comments from kernel-side adds
====================
Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
aqr_gen1_read_status() decodes the MDIO_PHYXS_VEND_IF_STATUS register
to determine which SerDes interface the PHY is currently using on its
system side and stores the result in phydev->interface. phylink relies
on this value to configure the MAC.
The autoneg == AUTONEG_DISABLE check is not correct:
MDIO_PHYXS_VEND_IF_STATUS is set by the PHY firmware based on the
negotiated link speed, not based on whether autoneg was used to reach
it. When the link comes up at 1G in forced mode, the register correctly
reads SGMII, but the early return prevents phydev->interface from being
updated. It stays at whatever value it held before (typically 2500BASE-X
from the initial autoneg run), so phylink configures the MAC for the
wrong interface and the link cannot come up.
Remove the autoneg guard so that the system interface type is always
decoded when the link is up.
Cc: stable@vger.kernel.org
Fixes: 110a2432c520 ("net: phy: aquantia: add downshift support")
Signed-off-by: Bartosz Golaszewski <bartosz.golaszewski@oss.qualcomm.com>
Link: https://patch.msgid.link/20260923-qcom-sa8255p-emac-v15-1-e82f33720737@oss.qualcomm.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
|
|
Pull audit fix from Paul Moore:
"A single audit fix for a potential UAF error in some audit filter
configurations"
* tag 'audit-pr-20260930' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/audit:
audit: fix exe mark UAF in kill_rules()
|
|
Add Rolling Wireless RN947R 0x9300 composition:
DIAG + ADB + NMEA + MODEM + RMNET + QDSS + ADPL
T: Bus=01 Lev=01 Prnt=01 Port=02 Cnt=02 Dev#= 6 Spd=480 MxCh= 0
D: Ver= 2.10 Cls=ef(misc ) Sub=02 Prot=01 MxPS=64 #Cfgs= 1
P: Vendor=33f8 ProdID=9300 Rev= 0.00
S: Manufacturer=Rolling Wireless
S: Product=RN947R
S: SerialNumber=xxxxxxxxxxx
C:* #Ifs= 7 Cfg#= 1 Atr=a0 MxPwr=500mA
I:* If#= 0 Alt= 0 #EPs= 2 Cls=ff(vend.) Sub=ff Prot=30 Driver=option
E: Ad=01(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms
E: Ad=81(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms
I:* If#= 1 Alt= 0 #EPs= 2 Cls=ff(vend.) Sub=42 Prot=01 Driver=usbfs
E: Ad=02(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms
E: Ad=82(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms
I:* If#= 2 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=40 Driver=option
E: Ad=84(I) Atr=03(Int.) MxPS= 10 Ivl=32ms
E: Ad=83(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms
E: Ad=03(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms
I:* If#= 3 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=40 Driver=option
E: Ad=86(I) Atr=03(Int.) MxPS= 10 Ivl=32ms
E: Ad=85(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms
E: Ad=04(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms
I:* If#= 8 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=50 Driver=qmi_wwan
E: Ad=88(I) Atr=03(Int.) MxPS= 8 Ivl=32ms
E: Ad=87(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms
E: Ad=05(O) Atr=02(Bulk) MxPS= 512 Ivl=0ms
I:* If#=21 Alt= 0 #EPs= 1 Cls=ff(vend.) Sub=ff Prot=70 Driver=(none)
E: Ad=89(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms
I:* If#=22 Alt= 0 #EPs= 1 Cls=ff(vend.) Sub=ff Prot=80 Driver=(none)
E: Ad=8a(I) Atr=02(Bulk) MxPS= 512 Ivl=0ms
Signed-off-by: Reinhard Speyerer <rspmn@arcor.de>
Link: https://patch.msgid.link/arZzuOOvLUZtdOlv@arcor.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|