[PATCH net v3] net: iterate online nodes in skb_defer_free_flush()
From: Kris Pan <hidden>
Date: 2026-09-16 00:37:43
Also in:
lkml
Subsystem:
networking [general], the rest · Maintainers:
"David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, Linus Torvalds
skb_attempt_defer_free() only queues skbs on the current CPU's node
(numa_node_id() of a running CPU), which is always online, so the
flush loop never needs to visit nodes that are merely possible.
for_each_node() walks node_possible_map. On machines where the
possible map is much larger than the online map -- e.g. a POWER10
LPAR with 32 possible but 1 online node -- the flush loop touches 31
cold, always-empty per-node lists on every softirq pass.
Use for_each_online_node() to iterate only node_online_map.
Since the flush now skips offline nodes, also drain the per-node
lists in dev_cpu_dead(). A skb is queued on the alloc CPU's list at
the freeing CPU's node index; if that node is offlined before the
alloc CPU flushes, the skb is stranded. Likewise, skbs queued on a
CPU that later goes offline are never freed, since its softirq no
longer runs.
Loopback UDP throughput in a QEMU guest with 32 possible / 1 online
nodes (bench_udp, 8 senders, interleaved runs) improves by ~5%.
Fixes: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()")
Reported-by: kernel test robot <redacted>
Suggested-by: Adrian Tomasov <redacted>
Suggested-by: Eric Dumazet <edumazet@google.com>
Closes: https://lore.kernel.org/oe-lkp/202512112119.5b9829a-lkp@intel.com (local)
Signed-off-by: Kris Pan <redacted>
---
net/core/dev.c | 27 +++++++++++++++++++++++++--
1 file changed, 25 insertions(+), 2 deletions(-)
diff --git a/net/core/dev.c b/net/core/dev.c
index ecfbd72d5d1a4..896498b349500 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c@@ -6907,7 +6907,7 @@ static void skb_defer_free_flush(void) struct skb_defer_node *sdn; int node; - for_each_node(node) { + for_each_online_node(node) { sdn = this_cpu_ptr(net_hotdata.skb_defer_nodes) + node; if (llist_empty(&sdn->defer_list))
@@ -12895,9 +12895,12 @@ int __dev_change_net_namespace(struct net_device *dev, struct net *net, static int dev_cpu_dead(unsigned int oldcpu) { struct sk_buff **list_skb; - struct sk_buff *skb; + struct llist_node *free_list; + struct sk_buff *skb, *next; + struct skb_defer_node *sdn; unsigned int cpu; struct softnet_data *sd, *oldsd, *remsd = NULL; + int node; local_irq_disable(); cpu = smp_processor_id();
@@ -12957,6 +12960,26 @@ static int dev_cpu_dead(unsigned int oldcpu) rps_input_queue_head_incr(oldsd); } + /* Drain skbs deferred to the offlining CPU: its softirq no longer + * runs, so skb_defer_free_flush() would never free them. Also drain + * skbs deferred to this CPU's node, which may leave node_online_map + * and thus be skipped by the for_each_online_node() loop in + * skb_defer_free_flush(). + */ + for_each_node(node) { + sdn = per_cpu_ptr(net_hotdata.skb_defer_nodes, oldcpu) + node; + free_list = llist_del_all(&sdn->defer_list); + llist_for_each_entry_safe(skb, next, free_list, ll_node) + dev_consume_skb_any(skb); + } + node = cpu_to_node(oldcpu); + for_each_online_cpu(cpu) { + sdn = per_cpu_ptr(net_hotdata.skb_defer_nodes, cpu) + node; + free_list = llist_del_all(&sdn->defer_list); + llist_for_each_entry_safe(skb, next, free_list, ll_node) + dev_consume_skb_any(skb); + } + return 0; }
--
2.43.0