Thread (2 messages) flat view 2 messages, 2 authors, 5d ago

Re: [PATCH net v2] bonding: fix initial last_rx vs ARP-monitor slack window

From: Hangbin Liu <hidden>
Date: 2026-08-28 02:50:11
Also in: lkml

On Thu, Aug 27, 2026 at 01:44:42PM +0200, Ramses de Norre via B4 Relay wrote:
From: Ramses de Norre <redacted>

Commit f31c7937c254 ("bonding: start slaves with link down for ARP
monitor") initialises a freshly enslaved port's last_rx to
jiffies - (arp_interval + 1) so that it does not "immediately cause
fake detection of 'up' state". At the time, the comparison was a plain
<= arp_interval and the value was just stale enough.

Commit da210f559019 ("bonding: add some slack to arp monitoring time
limits"), four months later, added a +arp_interval/2 slack term to
every comparison (now bond_time_in_interval()) but did not widen the
init to match. Since then, bond_time_in_interval(bond, last_rx, 1) is
true for the first ~arp_interval/2 after enslavement even though no
packet has been received: the upper bound is last_rx + 1.5*delta and
last_rx was set to jiffies - delta - 1.

If the ARP monitor tick lands in that window, bond_ab_arp_inspect()
proposes the slave UP. If the slave is the configured primary,
bond_ab_arp_commit() sets do_failover and the still-armed
force_primary in bond_choose_primary_or_current() makes it the active
slave regardless of primary_reselect. ARP validation as the active
slave then fails (the link has not actually received anything; on
SFP+ ports the PHY is often still negotiating) and the bond falls back
to the backup. With primary_reselect=failure, force_primary has now
been spent and the bond stays on the backup until something else
triggers a reselect.
Can we set primary_reselect to always or better to avoid this? If you prefer
to using the primary slave.
Reproducer:

  ip link add bond0 type bond mode active-backup arp_interval 1000 \
      arp_validate all arp_ip_target 192.0.2.1 \
      primary eth0 primary_reselect failure
  # eth0: SFP+ (slow link-up), eth1: RJ45 (fast link-up)
  ip link set eth0 master bond0
  ip link set eth1 master bond0
  ip link set bond0 up
  # bond0 lands on eth0 via force_primary, ARP-fails it before the
  # SFP+ has carrier, falls to eth1, and stays there.

Initialise last_rx (and the per-target array, and last_tx) to two full
intervals in the past so it is outside the slack window from the
start.
Is this trying to init the backup slave down by default?

Thanks
Hangbin
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help