Thread (21 messages) read the whole thread 21 messages, 5 authors, 26d ago

Re: [PATCH bpf-next 2/6] bpf: Add ksock kfuncs

From: Mahe Tardy <hidden>
Date: 2026-07-07 09:41:16
Also in: bpf

On Mon, Jul 06, 2026 at 10:50:33PM +0000, Kuniyuki Iwashima wrote:
From: Stanislav Fomichev <redacted>
Date: Mon, 6 Jul 2026 14:02:39 -0700
quoted
On 07/06, Mahe Tardy wrote:
quoted
On Mon, Jul 06, 2026 at 09:58:09AM -0700, Stanislav Fomichev wrote:
quoted
On 07/06, Mahe Tardy wrote:
quoted
Add BPF kfuncs that allow BPF LSM programs to create and use sockets for
sending data. This provides a mechanism for BPF programs to emit
telemetry. For this first patch set, it's restricted to SOCK_DGRAM
socket types with IPPROTO_UDP protocol but could be easily extended to
SOCK_STREAM and IPPROTO_TCP in the future.

The API consists of six kfuncs:

  bpf_ksock_create()   - Create a socket (sleepable)
[..]
quoted
  bpf_ksock_bind()     - Bind socket to local address (sleepable)
  bpf_ksock_connect()  - Connect socket to remote address (sleepable)
Since you're doing only UDP for now, maybe you don't need bind/connect? The
kernel should autobind (by default) when you sendmsg over UDP socket (IIRC).
Yep indeed, I kinda overlooked that as I started with UDP & TCP supports
and mostly added the args checks. Another thing is that send is simpler
since only used on connected sockets, so you just pass the struct
bpf_ksock and data. So on one side it would simplify the current
UDP-only API for now by removing the kfuncs but we might need a more
complex send kfunc (something like sendto).
Since you were targeting bpf_netpoll_send_udp originally, maybe sendto
is a better fit? You get the payload and the destination and you
bpf_sendto() it? We can later move to stateful bind/connect if needed.

(mostly coming from the pow of minimizing api exposure initially, but
not a strong preference)
+1, small start would be better.
Okay I feel at least we can confidently remove bind for now. Indeed for
netpoll it was kinda obvious that it was going to be sendto-style but
now with the sockets I'm unsure what's best:

As Amery Hung wrote in a parallel thread:
[...] but connect() + send() make sense to me. In the stated use case,
the dst addr probably doesn't change often and I think avoiding route
lookup everytime should be a good thing.
Looks like there's benefit to both approaches, I'll experiment.
quoted
quoted
quoted
quoted
  bpf_ksock_send()     - Send data through the socket (sleepable)
  bpf_ksock_acquire()  - Acquire a reference to a socket context
  bpf_ksock_release()  - Release a reference (cleanup via
                         queue_rcu_work since sock_release sleeps)
[..]
quoted
A bpf_ksock_max sysctl is added to limit the maximum number of BPF
kernel sockets that may exist in each network namespace. Out of
simplicity for now, the settings is host wide but the counters are per
network namespace.
What is this guarding against? Rogue bpf programs creating too many sockets?
Yes. AI review raised this because users are prevented from creating too
many sockets by bumping against the max number of fd and this would
allow them to create way more sockets. I kind of agreed that having "a
limit" on resource creation would make sense but maybe it doesn't and we
can simplify this!
I believe even the kernel sockets go via lsm layer, so this enforcement
can be done in an lsm bpf program. Seems like that should be enough?
Right, I don't think the per-netns limit is useful for CAP_BPF users.
Ok, agree, let's remove all this then for the next version.
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help