Re: [PATCH bpf-next 2/6] bpf: Add ksock kfuncs
From: Mahe Tardy <hidden>
Date: 2026-07-07 09:41:16
Also in:
bpf
On Mon, Jul 06, 2026 at 10:50:33PM +0000, Kuniyuki Iwashima wrote:
From: Stanislav Fomichev <redacted> Date: Mon, 6 Jul 2026 14:02:39 -0700quoted
On 07/06, Mahe Tardy wrote:quoted
On Mon, Jul 06, 2026 at 09:58:09AM -0700, Stanislav Fomichev wrote:quoted
On 07/06, Mahe Tardy wrote:quoted
Add BPF kfuncs that allow BPF LSM programs to create and use sockets for sending data. This provides a mechanism for BPF programs to emit telemetry. For this first patch set, it's restricted to SOCK_DGRAM socket types with IPPROTO_UDP protocol but could be easily extended to SOCK_STREAM and IPPROTO_TCP in the future. The API consists of six kfuncs: bpf_ksock_create() - Create a socket (sleepable)[..]quoted
bpf_ksock_bind() - Bind socket to local address (sleepable) bpf_ksock_connect() - Connect socket to remote address (sleepable)Since you're doing only UDP for now, maybe you don't need bind/connect? The kernel should autobind (by default) when you sendmsg over UDP socket (IIRC).Yep indeed, I kinda overlooked that as I started with UDP & TCP supports and mostly added the args checks. Another thing is that send is simpler since only used on connected sockets, so you just pass the struct bpf_ksock and data. So on one side it would simplify the current UDP-only API for now by removing the kfuncs but we might need a more complex send kfunc (something like sendto).Since you were targeting bpf_netpoll_send_udp originally, maybe sendto is a better fit? You get the payload and the destination and you bpf_sendto() it? We can later move to stateful bind/connect if needed. (mostly coming from the pow of minimizing api exposure initially, but not a strong preference)+1, small start would be better.
Okay I feel at least we can confidently remove bind for now. Indeed for netpoll it was kinda obvious that it was going to be sendto-style but now with the sockets I'm unsure what's best: As Amery Hung wrote in a parallel thread:
[...] but connect() + send() make sense to me. In the stated use case, the dst addr probably doesn't change often and I think avoiding route lookup everytime should be a good thing.
Looks like there's benefit to both approaches, I'll experiment.
quoted
quoted
quoted
quoted
bpf_ksock_send() - Send data through the socket (sleepable) bpf_ksock_acquire() - Acquire a reference to a socket context bpf_ksock_release() - Release a reference (cleanup via queue_rcu_work since sock_release sleeps)[..]quoted
A bpf_ksock_max sysctl is added to limit the maximum number of BPF kernel sockets that may exist in each network namespace. Out of simplicity for now, the settings is host wide but the counters are per network namespace.What is this guarding against? Rogue bpf programs creating too many sockets?Yes. AI review raised this because users are prevented from creating too many sockets by bumping against the max number of fd and this would allow them to create way more sockets. I kind of agreed that having "a limit" on resource creation would make sense but maybe it doesn't and we can simplify this!I believe even the kernel sockets go via lsm layer, so this enforcement can be done in an lsm bpf program. Seems like that should be enough?Right, I don't think the per-netns limit is useful for CAP_BPF users.
Ok, agree, let's remove all this then for the next version.