Thread (4 messages) flat view 4 messages, 4 authors, 5d ago

Re: [PATCH net] bonding: avoid ARP flood on RTNL contention in active-backup mode

From: Jay Vosburgh <jv@jvosburgh.net>
Date: 2026-09-01 23:44:32

Paolo Abeni [off-list ref] wrote:
On 8/31/26 11:09 AM, Eric Dumazet wrote:
quoted
Commit f1986b3a9f2e ("net: bonding: skip the 2nd trylock when first one
fail") changed bond_activebackup_arp_mon() to reschedule arp_work in
1 tick if the second rtnl_trylock() fails (for sending peer/slave
notifications).

However, by the time bond_activebackup_arp_mon() reaches this second lock
check, bond_ab_arp_probe() has already been executed and sent an ARP probe.
If RTNL remains contended, rescheduling every 1 tick causes
bond_activebackup_arp_mon() to re-execute bond_ab_arp_probe() every jiffy,
flooding the network with ARP probes at HZ frequency (e.g. 1000 pkts/sec)
instead of respecting the configured arp_interval.

If rtnl_trylock() fails at the second check, do not change delta_in_ticks
to 1 so that the next ARP monitor execution is scheduled according to the
configured arp_interval, matching the behavior in
bond_loadbalance_arp_mon().

Fixes: f1986b3a9f2e ("net: bonding: skip the 2nd trylock when first one fail")
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
Cc: Tonghao Zhang <redacted>
Cc: Hangbin Liu <redacted>
Cc: Jay Vosburgh <jv@jvosburgh.net>
---
 drivers/net/bonding/bond_main.c | 4 +---
 1 file changed, 1 insertion(+), 3 deletions(-)
diff --git a/drivers/net/bonding/bond_main.c b/drivers/net/bonding/bond_main.c
index ef9eb0c53c66..c23cf18a996a 100644
--- a/drivers/net/bonding/bond_main.c
+++ b/drivers/net/bonding/bond_main.c
@@ -3871,10 +3871,8 @@ static void bond_activebackup_arp_mon(struct bonding *bond)
 	rcu_read_unlock();
 
 	if (READ_ONCE(bond->send_peer_notif) || should_notify_rtnl) {
-		if (!rtnl_trylock()) {
-			delta_in_ticks = 1;
+		if (!rtnl_trylock())
 			goto re_arm;
Sashiko noted this should cause a regression, with notifications
potentially delayed for an unbounded time:

https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260831090937.3342052-1-edumazet%40google.com

That was also the behavior prior to f1986b3a9f2e, so I guess is 
a reasonable trade-off, but a 2nd opinion would help :)
	Yeah, without reworking all of this so it's just one round trip
on RTNL, it's a choice between possible ARP spam or an unlikely
possibility of egregiously delayed probes.  At the default missed_max of
2, with the rearm interval set to delta_in_ticks (i.e., this patch
applied), the ARP mon will fail over if it misses RTNL twice, with
caveat that the first miss needs to be the second RTNL acquisition in
bond_activebackup_arp_mon.

	I suppose another possibility would be to set delta_in_ticks to
something larger than 1, on the theory that RTNL shouldn't generally be
held for very long, so a sufficiently large value would be likely to
miss the contention but not wait too long.  Choosing a value is going to
have voodoo in there, and would likely have to be some fraction of
delta_in_ticks.

	Regardless of the rearm interval (1, delta_in_ticks, or
somewhere in between), the notification can be delayed for unbounded
time if we are sufficiently unlucky, although it's more likely with the
larger value from delta_in_ticks.

	That said, I don't have a major objection to changing this back.

Acked-by: Jay Vosburgh <jv@jvosburgh.net>

	-J

---
	-Jay Vosburgh, jv@jvosburgh.net
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help