Thread (41 messages) 41 messages, 7 authors, 2d ago

Re: [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory

flat view

From: Ihor Solodrai <ihor.solodrai@linux.dev>
Date: 2026-10-05 19:01:33
Also in: bpf, linux-doc, linux-input, linux-kbuild, linux-kselftest, linux-modules, linux-perf-users, lkml, rust-for-linux, sched-ext

On 10/4/26 3:21 PM, Jay Wang wrote:
quoted
   - If I know I don't need BPF on the workload, and these 5.4 MB
     bite me, I can just turn off BPF
They cannot, even when they know they do not need it.  As the cover
letter says, the cloud provider usually ships one kernel build to all of
its customers. [...]
Ah, I get it now. Thanks.

So the constraint is that the cloud provider ships a single kernel
revision to all users, and there is currently no mechanism to on/off
BPF other than a kernel build flag.

I'd be tempted to say "just ship two kernels", it's simpler than the
upstream change. But I understand the downsides of that as well.

And obviously a good upstream feature may benefit many more users.
[...]
quoted
     Keep zstd-compressed blob in the kernel image and decompress and
     parse it synchronously on first use.
Thanks for putting this forward.  This compression approach is simpler 
than loading a module.  I expect the overall code complexity to stay similar
to v3, since both need to find the first user and defer the module BTF parsing
until then, but it should be easier to merge into mainline because it avoids
the request_module() deadlocks.

Would you like to prepare it for a formal submission, or would you
like me to do so?  Either works for me.  Ideally it lands in time for
the next LTS kernel, so that we can adopt it there.
Please go ahead and submit the compression approach. Looks like we are
converging on that. Feel free to use the prototype code if it's
useful, it uses parts of your series anyway.
From what I can see, most of the work left is in that deferral.
Whoever needs the BTF first also parses the waiting module BTFs and
replays the queued registrations, struct_ops ->init() included, under
whatever locks it holds at that point, e.g. event_mutex or
bpf_verifier_lock, or inside the module notifier.  So the
cand_cache_mutex ABBA you found may not be the only deadlock.  Besides
that, only x86_64 was touched so far, so it needs extending to the
other architectures as well.
Yes, I think you are right that the hard part here is the deferral,
because everything expects vmlinux BTF to just be there.  I don't
think this is avoidable whatever memory saving approach we choose.
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help