Thread (16 messages) 16 messages, 2 authors, 1h ago

[PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load

HOTtoday

From: Jay Wang <hidden>
Date: 2026-10-01 22:54:27
Also in: bpf, linux-doc, linux-kbuild, linux-kselftest, linux-modules, linux-perf-users, linux-trace-kernel, lkml, rust-for-linux, sched-ext
Subsystem: bpf [general] (safe dynamic programs and tools), the rest · Maintainers: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Linus Torvalds

Revision v4 of 3 in this series.

Revisions (3)
  1. v2 [diff vs current]
  2. v3 [diff vs current]
  3. v4 current
A module with a .BTF.base section (built out of tree) that is loaded
before the vmlinux BTF only gets its /sys/kernel/btf file once the
vmlinux BTF has been loaded and its BTF relocated, because the raw .BTF
is only valid against the distilled base and relocation rewrites it in
place.  Until then the module is missing from /sys/kernel/btf, unlike
with =y, and reading its file cannot trigger the load.

Create the file at module load instead, with its final size: relocation
only rewrites type ids and string offsets, never the length.  Its reader,
btf_module_sysfs_read_deferred(), has the vmlinux BTF loaded, which parses
and relocates the kept modules, then waits until this module's BTF is
published (btf_mod->ready) before serving it, so no unrelocated or
half-relocated data is ever visible.  If the module goes away or its BTF
turns out unusable first (btf_mod->gone), or the vmlinux BTF cannot be
loaded, the read fails with -ENODEV.

The reader does not load the vmlinux BTF itself but queues a work item
that does: it holds the file's kernfs active reference, which
MODULE_STATE_GOING waits for when it removes the file, with the module
notifier chain held; loading btf_vmlinux takes that chain, so with a
writer queued on it the two would wait for each other.  The reader's own
wait ends when the module goes.

Because that reader may be waiting for btf_parse_deferred_modules(), and
removing a sysfs file waits for its readers, nothing removes a sysfs file
from that path (a failed entry
stays dead until the module goes, which the previous patch already
arranged), and MODULE_STATE_GOING removes the file after dropping
btf_module_mutex, which the reader may need to get there.  A dead entry
keeps its raw data only for a file that serves it as is, the deferred
reader never does.

Modules without .BTF.base are unchanged: their data is final and is
served as is.

Signed-off-by: Jay Wang <redacted>
---
 kernel/bpf/btf.c | 144 ++++++++++++++++++++++++++++++++++++++++-------
 1 file changed, 124 insertions(+), 20 deletions(-)
diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 4d51fb212218..3c3aba0fc4cb 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -9208,17 +9208,27 @@ struct btf_module {
 	u32 data_size;
 	u32 base_data_size;
 	struct list_head deferred_regs;
-	/* the kept BTF turned out unusable; the entry stays until the module goes */
+	/*
+	 * For the sysfs reader of a module whose data is only final once its
+	 * BTF is relocated: @ready once @btf is published, @gone once the
+	 * entry is dead (parse failed or module going).  Both only ever go
+	 * from false to true; waiters sleep on btf_module_wq.
+	 */
+	bool ready;
 	bool gone;
 };
 
 static LIST_HEAD(btf_modules);
 static DEFINE_MUTEX(btf_module_mutex);
+static DECLARE_WAIT_QUEUE_HEAD(btf_module_wq);
 
 static void purge_cand_cache(struct btf *btf);
 
 static int btf_module_sysfs_add(struct btf_module *btf_mod, const char *name,
-				void *data, size_t data_size)
+				void *private, size_t size,
+				ssize_t (*read)(struct file *, struct kobject *,
+						const struct bin_attribute *,
+						char *, loff_t, size_t))
 {
 	struct bin_attribute *attr;
 	int err;
@@ -9233,9 +9243,9 @@ static int btf_module_sysfs_add(struct btf_module *btf_mod, const char *name,
 	sysfs_bin_attr_init(attr);
 	attr->attr.name = name;
 	attr->attr.mode = 0444;
-	attr->size = data_size;
-	attr->private = data;
-	attr->read = sysfs_bin_attr_simple_read;
+	attr->size = size;
+	attr->private = private;
+	attr->read = read;
 
 	err = sysfs_create_bin_file(btf_kobj, attr);
 	if (err) {
@@ -9249,8 +9259,14 @@ static int btf_module_sysfs_add(struct btf_module *btf_mod, const char *name,
 	return 0;
 }
 
+/*
+ * Called with btf_module_mutex NOT held: removing the sysfs file waits for
+ * readers to leave, and a deferred reader may need the mutex to get there.
+ */
 static void btf_module_free(struct btf_module *btf_mod)
 {
+	WRITE_ONCE(btf_mod->gone, true);
+	wake_up_all(&btf_module_wq);
 	if (btf_mod->sysfs_attr)
 		sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr);
 	if (btf_mod->btf) {
@@ -9311,15 +9327,87 @@ static bool btf_is_vmlinux_carrier(const struct module *mod)
 	return !strcmp(mod->name, btf_vmlinux_link.module_name);
 }
 
+/*
+ * sysfs reader for a module kept aside with a .BTF.base section: its .BTF is
+ * split against the distilled base and only becomes valid split BTF against
+ * the vmlinux BTF once relocated, which rewrites the buffer in place.  So
+ * have the vmlinux BTF loaded (which parses and relocates the kept modules),
+ * then wait until this module's BTF is published.  The size does not change:
+ * relocation only rewrites ids and string offsets.
+ *
+ * The reader does not load the vmlinux BTF itself: it holds the file's
+ * kernfs active reference, which MODULE_STATE_GOING waits for when it
+ * removes the file with the module notifier chain held, and loading
+ * btf_vmlinux needs that chain.  A work item loads it, and the reader waits
+ * in a way that the module going away (@gone) ends.
+ */
+static bool btf_module_published(struct btf_module *btf_mod)
+{
+	/* Pairs with the smp_store_release() of @ready after btf_mod->btf is set */
+	return smp_load_acquire(&btf_mod->ready);
+}
+
+/* Bumped after each load attempt by btf_vmlinux_load_work */
+static atomic_t btf_vmlinux_load_seq = ATOMIC_INIT(0);
+
+static void btf_vmlinux_load_workfn(struct work_struct *work)
+{
+	bpf_load_btf_vmlinux();
+	atomic_inc(&btf_vmlinux_load_seq);
+	wake_up_all(&btf_module_wq);
+}
+
+static DECLARE_WORK(btf_vmlinux_load_work, btf_vmlinux_load_workfn);
+
+/* The vmlinux BTF could not be had: a load attempt ended without it. */
+static bool btf_vmlinux_load_failed(int seq)
+{
+	return atomic_read(&btf_vmlinux_load_seq) != seq &&
+	       IS_ERR_OR_NULL(bpf_peek_btf_vmlinux());
+}
+
+static ssize_t btf_module_sysfs_read_deferred(struct file *filp, struct kobject *kobj,
+					      const struct bin_attribute *attr,
+					      char *buf, loff_t off, size_t count)
+{
+	struct btf_module *btf_mod = attr->private;
+	int seq = atomic_read(&btf_vmlinux_load_seq);
+	struct btf *vmlinux_btf = bpf_peek_btf_vmlinux();
+	int err;
+
+	if (IS_ERR(vmlinux_btf))
+		return -ENODEV;
+	if (!vmlinux_btf)
+		queue_work(system_unbound_wq, &btf_vmlinux_load_work);
+
+	/*
+	 * Another thread may still be relocating and publishing it; if the
+	 * module goes away or its BTF turns out unusable, btf_module_free()
+	 * or btf_parse_deferred_modules() set @gone and wake us.
+	 */
+	err = wait_event_interruptible(btf_module_wq,
+				       btf_module_published(btf_mod) ||
+				       READ_ONCE(btf_mod->gone) ||
+				       btf_vmlinux_load_failed(seq));
+	if (err)
+		return err;
+	if (!btf_module_published(btf_mod))
+		return -ENODEV;
+
+	/* sysfs clamps @off and @count to attr->size == btf->data_size */
+	memcpy(buf, btf_mod->btf->data + off, count);
+	return count;
+}
+
 /*
  * The vmlinux BTF is not available yet and must not be loaded from the
  * module notifier (that would nest a module load into a module load).  Keep
  * the module's BTF for btf_parse_deferred_modules().
  *
- * Without a .BTF.base section the .BTF data is final and can be exposed in
- * sysfs right away, it needs no parsing.  With one, parsing relocates the
- * data in place against the vmlinux BTF, so the file is created afterwards,
- * as with =y where it also only appears once the BTF is parsed.
+ * The sysfs file is created right away with its final size, as with =y.
+ * Without a .BTF.base section the .BTF data is final and is served as is;
+ * with one, it is only valid once relocated, so its reader waits for that
+ * (btf_module_sysfs_read_deferred()).
  */
 static int btf_module_defer(struct btf_module *btf_mod, struct module *mod)
 {
@@ -9338,10 +9426,12 @@ static int btf_module_defer(struct btf_module *btf_mod, struct module *mod)
 			return -ENOMEM;
 		}
 		btf_mod->base_data_size = mod->btf_base_data_size;
-	} else {
 		/* not fatal, the module BTF is usable without the sysfs file */
+		btf_module_sysfs_add(btf_mod, mod->name, btf_mod, btf_mod->data_size,
+				     btf_module_sysfs_read_deferred);
+	} else {
 		btf_module_sysfs_add(btf_mod, mod->name, btf_mod->data,
-				     btf_mod->data_size);
+				     btf_mod->data_size, sysfs_bin_attr_simple_read);
 	}
 
 	list_add(&btf_mod->list, &btf_modules);
@@ -9469,7 +9559,8 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 		mutex_unlock(&btf_module_mutex);
 
 		/* not fatal, the module BTF is usable without the sysfs file */
-		btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size);
+		btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size,
+				     sysfs_bin_attr_simple_read);
 		break;
 	case MODULE_STATE_LIVE:
 		mutex_lock(&btf_module_mutex);
@@ -9507,8 +9598,10 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			if (btf_mod->btf)
 				btf_free_id(btf_mod->btf);
 			list_del(&btf_mod->list);
+			mutex_unlock(&btf_module_mutex);
+			/* off the list, nobody else can find it now */
 			btf_module_free(btf_mod);
-			break;
+			goto out;
 		}
 		mutex_unlock(&btf_module_mutex);
 		break;
@@ -9536,23 +9629,34 @@ static bool btf_deferred_modules_done;
 /*
  * A kept module whose BTF cannot be used after all.  The module is loaded
  * and stays, so there is no way to reject it: the entry stays on the list,
- * dead, until the module goes.  A sysfs file it has keeps serving the raw
- * data, which is kept for that.
+ * dead, until the module goes.  Its sysfs file stays too, its reader, or
+ * the caller of this function, may be inside it right now: a .BTF.base
+ * reader wakes up and fails without touching the data, which can go; a
+ * plain one keeps serving the raw data, which is kept for that.
  */
 static void btf_module_dead(struct btf_module *btf_mod, const char *what, int err)
 {
 	pr_warn("failed to %s module [%s] BTF: %d\n", what, btf_mod->module->name, err);
+	if (!btf_mod->sysfs_attr ||
+	    btf_mod->sysfs_attr->read != sysfs_bin_attr_simple_read) {
+		kvfree(btf_mod->data);
+		btf_mod->data = NULL;
+	}
 	kvfree(btf_mod->base_data);
 	btf_mod->base_data = NULL;
 	btf_free_deferred_regs(&btf_mod->deferred_regs);
-	btf_mod->gone = true;
+	WRITE_ONCE(btf_mod->gone, true);
+	wake_up_all(&btf_module_wq);
 }
 
 /*
  * CONFIG_DEBUG_INFO_BTF=m: the vmlinux BTF has just become available.  Parse
  * the BTF of the modules that were loaded before it, and apply the
  * registrations that waited for them.  Called from bpf_load_btf_vmlinux()
- * once btf_vmlinux is published, serialized by it, with no locks held.
+ * once btf_vmlinux is published, serialized by it, with no locks held.  The
+ * sysfs reader of one of these modules may be waiting for it (see
+ * btf_module_sysfs_read_deferred()), which is why no sysfs file is removed
+ * here.
  *
  * A module's BTF is published (btf_mod->btf set, id installed) only after
  * its queued registrations are applied, so nobody sees a module BTF without
@@ -9617,12 +9721,12 @@ void btf_parse_deferred_modules(void)
 		if (btf_mod->flags & BTF_MODULE_F_LIVE)
 			btf_module_apply_regs(btf_mod, btf);
 
-		/* modules with .BTF.base get their sysfs file now, the data is relocated */
-		if (!btf_mod->sysfs_attr)
-			btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size);
 		btf_mod->data = NULL;
 		btf_mod->btf = btf;
 		btf_install_id(btf);
+		/* Pairs with the smp_load_acquire() in btf_module_sysfs_read_deferred() */
+		smp_store_release(&btf_mod->ready, true);
+		wake_up_all(&btf_module_wq);
 		parsed = true;
 		/* the list may have changed while the mutex was dropped */
 		goto restart;
-- 
2.47.3
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help