[REGRESSION] 6.12.111: "net: advertise TCP MSS from the configured MTU, not the learned PMTU" breaks TCP through relays when the peer's network drops ICMP

From: Echoo Wall <hidden>
Date: 2026-10-09 16:27:48
Also in: regressions, stable

Hi,

After a routine stable update from 6.12.107 to 6.12.111 (Debian trixie-security,
6.12.111-1), TCP connections to our servers started timing out for a subset of
clients. Booting back to 6.12.107 / 6.12.101 fixed it immediately. I believe the
cause is:

  2640e64195948a601430d230c9864f5426574cde
  net: advertise TCP MSS from the configured MTU, not the learned PMTU
  (in stable 6.12.111 and 6.18.51)

#regzbot introduced: 2640e64195948a601430d230c9864f5426574cde

Topology
--------

  client ---(internet)---> relay (Linux, nftables DNAT + masquerade) ---> server
                                                               ^
                                                 path MTU 1456 on this hop

- Servers have a 1500-byte interface MTU, but the path between the relay and the
  server has a PMTU of 1456 (encapsulation upstream). The server learns this
  correctly: `ip route get <relay>` shows `cache mtu 1456`.
- Clients reach the server only through the relay (port-based DNAT, SNAT to the
  relay address).
- Some client networks (typical corporate firewalls) drop ICMP
  "fragmentation needed".

Behaviour
---------

Before (<= 6.12.110): the server's SYN-ACK advertised MSS 1416, derived from the
learned PMTU. Clients never sent segments that were too large. Everything worked,
including from networks that filter ICMP.

After (6.12.111): the SYN-ACK advertises MSS 1460 from the configured MTU. The
client sends 1500-byte packets. They are dropped at the 1456 hop, and the ICMP
error never reaches the client because its network filters it. The handshake
completes, then the connection stalls: a classic PMTU black hole.

Direct connections to the same servers, without the relay hop, keep working, and
so do servers still on <= 6.12.110 behind the same relays. That made the problem
very hard to attribute: it only appears from certain client networks, only on
relayed paths, and only after a kernel update that nothing points to.

What I verified
---------------

- Reproducible per host by switching only the kernel. Rebooting without changing
  the kernel does not help. Same sysctls on both kernels (309 TCP parameters
  compared), same userspace.
- On 6.12.107 the server has `cache mtu 1456` towards the relay, so the
  SYN-ACK MSS is derived from it (1416). With this commit the advertised
  MSS comes from the 1500-byte device MTU (1460) instead. That matches the
  symptom exactly, but I have not yet captured a 6.12.111 SYN-ACK on the
  affected path.
- As a workaround we now clamp MSS on the relay
  (`tcp flags syn tcp option maxseg size > 1416 tcp option maxseg size set 1416`
  in the forward chain). I have not yet re-tested 6.12.111 behind it.

Why I think this matters
------------------------

The commit message argues that on symmetric paths nothing is lost because "the
peer usually already knows the real path MTU". That assumption does not hold when:

  1. the peer sits behind a firewall that drops ICMP fragmentation-needed, which
     is common in enterprise networks; and/or
  2. the narrow hop is only on the server side (relay, tunnel or overlay egress),
     so the peer has no way to learn it other than via ICMP.

In both cases the receiver-side PMTU knowledge was the only thing that kept the
connection working. The old behaviour was effectively a built-in safeguard for
these black holes, and it has been in place for a very long time.

This is now shipping as a security update in LTS/stable kernels and
distributions. It silently breaks previously working deployments, and only for
some users, so most operators will not connect the symptom to the kernel. The
6.6 / 6.1 / 5.15 backports appear to be still pending, so there may still be
time to reconsider before it spreads further.

Possible ways forward (just suggestions):

  - revert from stable and keep it in mainline until there is a safer variant;
  - make it opt-in (sysctl or per-route flag) for the asymmetric / DSR case it
    targets;
  - or only ignore the learned PMTU when there is evidence of asymmetry, instead
    of unconditionally.

I am happy to test patches on the affected setup.

Thanks.
Esko Mobius
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help