Re: [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module
From: bot+bpf-ci@kernel.org
Date: 2026-10-01 23:45:57
Also in:
bpf, linux-doc, linux-kbuild, linux-kselftest, linux-modules, linux-perf-users, linux-trace-kernel, lkml, rust-for-linux, sched-ext
bpf: take the vmlinux BTF from the btf_vmlinux module
TL;DR: Retrying __sys_bpf() in bpf() may break BPF_PROG_LOAD: the failed first run can write the kernel's record size into uattr, so the retry may fail with -EINVAL instead of -E2BIG and libbpf won't recover.
quoted hunk ↗ jump to hunk
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c --- a/kernel/bpf/syscall.c +++ b/kernel/bpf/syscall.c
[ ... ]
quoted hunk ↗ jump to hunk
@@ -6525,10 +6544,40 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size, return err; } +/* + * With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF is loaded on demand, but never + * from within a command: loading waits for user space, and a command may hold + * locks or run from a BPF program (bpf_sys_bpf()). A command that needs the + * BTF while it is not loaded fails as it would without BTF. If the command + * is one whose failure leaves nothing behind, load the BTF here, on entry + * from user space with nothing held, and run the command once more. + */
[ ... ]
quoted hunk ↗ jump to hunk
SYSCALL_DEFINE5(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size, struct bpf_common_attr __user *, uattr_common, unsigned int, size_common) { - return __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common); + unsigned int misses = bpf_btf_vmlinux_misses(); + int err; + + err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common); + if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) && err < 0 && bpf_btf_vmlinux_retry(cmd, misses)) + err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), + size_common); + return err; }
Is it safe to run __sys_bpf() a second time on the same uattr? The comment
and the commit message say a failed BPF_PROG_LOAD "leaves nothing behind" and
"fails again the same way", but the first run can write into the caller's
attr.
When the func_info, line_info or core_relo records are larger than the
kernel's struct and have a non-zero tail, the failing run writes the
kernel's record size back into the attr, which is an input field:
kernel/bpf/check_btf.c:
if (ret == -E2BIG) {
verbose(env, "nonzero tailing record in func info");
/* set the size kernel expects so loader can zero
* out the rest of the record.
*/
if (copy_to_bpfptr_offset(uattr,
offsetof(union bpf_attr, func_info_rec_size),
&min_size, sizeof(min_size)))
The same happens for line_info_rec_size in check_btf_line() and for
core_relo_rec_size in bpf_check_core_relo().
With CONFIG_DEBUG_INFO_BTF=m, suppose the global miss counter moves while
this load runs. That happens if any bpf_get_btf_vmlinux() call misses, for
example another process loading BPF, or btf_module_notify() for a module that
is loading. Suppose the carrier module can also be loaded. Then
bpf_btf_vmlinux_retry() returns true and the second __sys_bpf() copies the
attr in again.
func_info_rec_size (or line_info_rec_size) now holds the kernel's size, but
the user's func_info/line_info array is still laid out with the original,
larger stride. prepare_btf_func() and check_btf_line() then read record i
from the wrong offset, landing in the previous record's non-zero tail, and
fail with -EINVAL ("same or smaller insn offset", "Invalid
line_info[%u].insn_off") instead of -E2BIG.
libbpf's bpf_prog_load() recovers from E2BIG only here:
while (errno == E2BIG && (!finfo || !linfo))
which rebuilds the records with the size the kernel wrote back. With EINVAL
that path is skipped, so a load that works without this patch (E2BIG, libbpf
trims the records, then success) now fails, and the verifier log shows a
misleading message.
A miss in the same run does not trigger this, because it fails earlier
(attach_btf, CO-RE candidates, kfunc). Only a concurrent miss before the
vmlinux BTF is first loaded does, for example during boot.
Should the retry be skipped when err == -E2BIG, or more generally when the
first run may have written into uattr? kernel/bpf/syscall.c is not touched
by later commits in the series, so nothing there changes this.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36938681172