Re: [PATCH net] bonding: avoid ARP flood on RTNL contention in active-backup mode
From: Jay Vosburgh <jv@jvosburgh.net>
Date: 2026-09-01 23:44:32
Paolo Abeni [off-list ref] wrote:
On 8/31/26 11:09 AM, Eric Dumazet wrote:quoted
Commit f1986b3a9f2e ("net: bonding: skip the 2nd trylock when first one fail") changed bond_activebackup_arp_mon() to reschedule arp_work in 1 tick if the second rtnl_trylock() fails (for sending peer/slave notifications). However, by the time bond_activebackup_arp_mon() reaches this second lock check, bond_ab_arp_probe() has already been executed and sent an ARP probe. If RTNL remains contended, rescheduling every 1 tick causes bond_activebackup_arp_mon() to re-execute bond_ab_arp_probe() every jiffy, flooding the network with ARP probes at HZ frequency (e.g. 1000 pkts/sec) instead of respecting the configured arp_interval. If rtnl_trylock() fails at the second check, do not change delta_in_ticks to 1 so that the next ARP monitor execution is scheduled according to the configured arp_interval, matching the behavior in bond_loadbalance_arp_mon(). Fixes: f1986b3a9f2e ("net: bonding: skip the 2nd trylock when first one fail") Signed-off-by: Eric Dumazet <edumazet@google.com> --- Cc: Tonghao Zhang <redacted> Cc: Hangbin Liu <redacted> Cc: Jay Vosburgh <jv@jvosburgh.net> --- drivers/net/bonding/bond_main.c | 4 +--- 1 file changed, 1 insertion(+), 3 deletions(-)diff --git a/drivers/net/bonding/bond_main.c b/drivers/net/bonding/bond_main.c index ef9eb0c53c66..c23cf18a996a 100644 --- a/drivers/net/bonding/bond_main.c +++ b/drivers/net/bonding/bond_main.c@@ -3871,10 +3871,8 @@ static void bond_activebackup_arp_mon(struct bonding *bond) rcu_read_unlock(); if (READ_ONCE(bond->send_peer_notif) || should_notify_rtnl) { - if (!rtnl_trylock()) { - delta_in_ticks = 1; + if (!rtnl_trylock()) goto re_arm;Sashiko noted this should cause a regression, with notifications potentially delayed for an unbounded time: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260831090937.3342052-1-edumazet%40google.com That was also the behavior prior to f1986b3a9f2e, so I guess is a reasonable trade-off, but a 2nd opinion would help :)
Yeah, without reworking all of this so it's just one round trip on RTNL, it's a choice between possible ARP spam or an unlikely possibility of egregiously delayed probes. At the default missed_max of 2, with the rearm interval set to delta_in_ticks (i.e., this patch applied), the ARP mon will fail over if it misses RTNL twice, with caveat that the first miss needs to be the second RTNL acquisition in bond_activebackup_arp_mon. I suppose another possibility would be to set delta_in_ticks to something larger than 1, on the theory that RTNL shouldn't generally be held for very long, so a sufficiently large value would be likely to miss the contention but not wait too long. Choosing a value is going to have voodoo in there, and would likely have to be some fraction of delta_in_ticks. Regardless of the rearm interval (1, delta_in_ticks, or somewhere in between), the notification can be delayed for unbounded time if we are sufficiently unlucky, although it's more likely with the larger value from delta_in_ticks. That said, I don't have a major objection to changing this back. Acked-by: Jay Vosburgh <jv@jvosburgh.net> -J --- -Jay Vosburgh, jv@jvosburgh.net