Thread (8 messages) 8 messages, 2 authors, 20m ago
HOTtoday

Revision v3 of 3 in this series.

Revisions (3)
  1. v1 [diff vs current]
  2. v2 [diff vs current]
  3. v3 current

[PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs

From: Wang Zhan <hidden>
Date: 2026-09-28 04:41:17

BIG TCP is negotiated per netdevice, so BIG TCP and non-BIG TCP ports can
coexist in one path.  When an skb exceeds the GSO limits of the port it is
sent to, it loses its GSO feature mask and is segmented into individual MSS
sized packets, which that port then sends without TSO.

This series cuts such a TCP GSO skb into GSO skbs which fit the device
limits instead, so the rest of the path keeps using TSO.  On a
veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP enabled on the
veth endpoints and left off in the guest, a single iperf3 TCP flow, six
alternating runs per state (-t 15 -O 5, fixed CPU affinity and port tuple):

  protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
  TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
  TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps

The middle column is the same tree with the bounded path disabled.  A BIG
TCP hop which feeds a 64 KiB hop loses 69% of the throughput of a path
which never enables BIG TCP; bounded resegmentation recovers it.

1/5 stands on its own as a fix: the size limit is picked from
skb->protocol, which is the VLAN ethertype for a frame which already
carries its tag in the packet, so an IPv6 frame was measured against the
IPv4 limit.  2/5 is the preparation which lets the limit tests be skipped
for one caller, and carries no functional change.

The new path is taken only when the skb is an unencapsulated TCP GSO skb
without a frag_list, it exceeds gso_max_size or gso_max_segs, and the
device offloads that GSO type.  Everything else keeps today's
segmentation.

The output obeys the GSO feature and limit contract the device already
advertises, so this needs no new UAPI, no device state and no driver
change, and it applies automatically.  The bound only says how the output
is grouped; an over-limit skb pays one extra ndo_features_check() in
exchange for staying a GSO skb.

Alternatives considered:

  - The caller could set skb_shinfo(skb)->gso_size to ~64K and adjust the
    gso bits in the shared info afterwards, which would need
    skb_unclone(), a repeat of the grouping logic skb_segment() already
    has, and a recomputed IPv4 ID for the DF=0 case.

Patch layout:

  [1/5] the GSO size limit follows the packet's L3 protocol
  [2/5] factor the device limit check out of gso_features_check()
  [3/5] let the GSO engine bound the MSS segments per output skb
  [4/5] apply that bound to oversized TCP GSO skbs in the TX path
  [5/5] KUnit coverage for the bound, the device limits and the TCP path

---
v3:
- patch 1 is new: the size limit follows the packet's L3 protocol
- patch 2: move the check_gso_limits flag and the wrapper split in from patch 4
- patch 3: drop the tcp_gso_segment() exception, the caller keeps the features
- patch 3: cap the bound with GSO_MAX_SEGS and drop the output reset
- patch 3: document the bound as TCP only and note the frag_list gate
- patch 4: enter from the limit predicate gso_features_check() uses
- patch 4: keep the features, so the bound shapes the output only
- patch 4: drop the SG and checksum tests, fold the MSS minimum
- patch 4: cap each output at GSO_LEGACY_MAX_SIZE, not the BIG TCP size
- patch 5: cover the IPv4 and the IPv6 limit, with and without the tag
- patch 5: skip the TCP cases without CONFIG_INET, reserve headroom
- patch 5: free on the failure paths, check the ungrouped single MSS
v2: https://lore.kernel.org/20260918084651.3022878-1-wang.zhan@smartx.com/ (local)
v1: https://lore.kernel.org/20260917063854.2011613-1-wang.zhan@smartx.com/ (local)

Wang Zhan (5):
  net: core: use the packet's L3 protocol for the GSO size limit
  net: core: factor out the GSO device limit check
  net: gso: support bounded TCP segmentation
  net: core: resegment oversized TCP GSO skbs
  net: net_test: add tests for bounded GSO segmentation

 drivers/net/tap.c          |   3 +-
 include/linux/netdevice.h  |   4 +-
 include/net/gso.h          |   6 +-
 include/net/udp.h          |   2 +-
 net/core/dev.c             |  91 ++++++++--
 net/core/gso.c             |   6 +-
 net/core/net_test.c        | 347 +++++++++++++++++++++++++++++++++++++
 net/core/skbuff.c          |  14 +-
 net/openvswitch/datapath.c |   2 +-
 9 files changed, 455 insertions(+), 20 deletions(-)


base-commit: 014d795c73837ea2339a4ea8e8f82c6e959b845d
-- 
2.47.3
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help