Thread (17 messages) 17 messages, 5 authors, 14h ago

[PATCH v5 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT.

flat view
HOTtoday IN BPF-NEXT

From: Kuniyuki Iwashima <kuniyu@google.com>
Date: 2026-10-08 03:16:07
Also in: bpf

Revision v5 of 5 in this series; queued in bpf-next as 5ce194eeb35a on 2026-10-10.

Revisions (5)
  1. v1 [diff vs current]
  2. v2 [diff vs current]
  3. v3 [diff vs current]
  4. v4 [diff vs current]
  5. v5 current
This series introduces two callbacks for bpf_tcp_ops:

  .enqueue_rcvq(): invoked when TCP stack enqueues skb to
                   sk->sk_receive_queue

  .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
                   is dequeued from sk->sk_receive_queue

Those callbacks can be enabled on a per-socket basis by
a new kfunc bpf_tcp_ops_set_flags():

  bpf_tcp_ops_set_flags((struct tcp_sock *)sk,
                        BPF_TCP_OPS_FLAG_RCVQ, 0);

This allows the BPF prog to dynamically adjust sk->sk_rcvlowat,
suppressing unnecessary EPOLLIN wakeups until sufficient data
is available in the receive queue.

This functionality, which we call "TCP AutoLOWAT", was originally
developed in 2020 by Tenzin Ukyab with the help of Soheil Hassas
Yeganeh, Arjun Roy, and Eric Dumazet.  It has served Google RPC
workloads for more than 5 years.

Combined with TCP RX zerocopy, this typically allows us to read an
entire RPC frame with just a single wakeup and a single system call.

While the original implementation was specialised for our
internal RPC format, this series introduces a more flexible
version by leveraging BPF.

The bpf prog in the last selftest patch closely mirrors the core
logic of the original implementation to provide a real-world
example.

Note that the new callbacks are not supported on legacy SOCK_OPS.

Changes:
  v5:
    * Patch 4 & 5
      * Always inline bpf_tcp_ops_hdr_opt_len() and
        bpf_skops_hdr_opt_len() to tcp_established_options().

  v4: https://lore.kernel.org/bpf/20261006192601.1875100-1-kuniyu@google.com/ (local)
    * Patch 2
      * Allow-listed BPF_PROG_TYPE_CGROUP_SOCKOPT and _SKB
        in bpf_tcp_ops_set_flags_kfunc_filter()
      * Add doc for flag enum and bpf_tcp_ops members
    * Patch 4
      * Check static key first and then flags
    * Patch 5 (new)
      * Bubble up static key check for SOCK_OPS similar to bpf_tcp_ops
    * Patch 6
      * Use bpf_tcp_ops_call_flag()

  v3: https://lore.kernel.org/bpf/20261005154533.4147685-1-kuniyu@google.com/ (local)
    * Add patch 2 ~ 4 (guard bpf_tcp_ops w/ a new flag)
    * Patch 5 & 9 : Adapt to a new flag and kfunc

  v2: https://lore.kernel.org/netdev/20260923213719.224838-1-kuniyu@google.com/ (local)
    * Add patch 1 not to allow setsockopt() from new callbacks
    * Patch 8 (selftest)
      * Avoid address comparison for a specific version of gcc.
      * Make rpc_test_cases[] static.
      * Update comment in rpc_test_case[].

  v1: https://lore.kernel.org/bpf/20260920195633.3033620-1-kuniyu@google.com/ (local)

Legacy SOCK_OPS version:
  v3: https://lore.kernel.org/bpf/20260523083001.2911931-1-kuniyu@google.com/ (local)
  v2: https://lore.kernel.org/bpf/20260522074601.1658705-1-kuniyu@google.com/ (local)
  v1: https://lore.kernel.org/bpf/20260508073355.3916746-1-kuniyu@google.com/ (local)


Kuniyuki Iwashima (10):
  bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
  bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops.
  selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c.
  bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag.
  bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after
    cgroup_bpf_enabled(CGROUP_SOCK_OPS).
  bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  tcp: Split out __tcp_set_rcvlowat().
  bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ.
  bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
  selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().

 .../networking/net_cachelines/tcp_sock.rst    |   1 +
 include/linux/bpf-cgroup.h                    |  36 +-
 include/linux/tcp.h                           |   7 +
 include/net/tcp.h                             | 122 +++---
 include/uapi/linux/bpf.h                      |  25 ++
 net/ipv4/af_inet.c                            |   2 +-
 net/ipv4/bpf_tcp_ops.c                        | 159 +++++++-
 net/ipv4/tcp.c                                |  20 +-
 net/ipv4/tcp_fastopen.c                       |   2 +
 net/ipv4/tcp_input.c                          |  37 +-
 net/ipv4/tcp_nv.c                             |   2 +-
 net/ipv4/tcp_output.c                         |  89 ++---
 net/ipv4/tcp_timer.c                          |   6 +-
 tools/include/uapi/linux/bpf.h                |  25 ++
 .../selftests/bpf/prog_tests/tcp_autolowat.c  | 350 ++++++++++++++++++
 .../selftests/bpf/progs/bpf_tcp_ops_hdr.c     |  20 +
 .../selftests/bpf/progs/bpf_tracing_net.h     |   2 +
 .../selftests/bpf/progs/tcp_autolowat.c       | 294 +++++++++++++++
 18 files changed, 1052 insertions(+), 147 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
 create mode 100644 tools/testing/selftests/bpf/progs/tcp_autolowat.c

-- 
2.56.0.360.g66cac248cb-goog
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help