Thread (12 messages) 12 messages, 3 authors, 6d ago

[PATCH v1 net-next 3/5] net: Track state in ops_undo_list().

COOLING6d

From: Kuniyuki Iwashima <kuniyu@google.com>
Date: 2026-09-27 20:24:41
Subsystem: networking [general], the rest · Maintainers: "David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, Linus Torvalds

We will call rt_flush_dev() and rt6_uncached_list_flush_dev()
from ->pre_exit_batch().

Then, we want them to return early when called again from
 ->exit_rtnl() or default_device_exit_batch().

However, we cannot simply return early when !check_net(net).

In the following cases, even if check_net(net) is false,
we cannot skip rt_flush_dev() / rt6_uncached_list_flush_dev():

  1. some ->pre_exit() call unregister_netdevice() before
     fib_net_ops (e.g. ovs_pre_exit_net(), l2tp_pre_exit_net()).

  2. ->dellink() could call unregister_netdevice() for another
     netdev in a dying netns queued for the next cleanup_net()
     batch, for which ->pre_exit_batch() has not been called
     yet (e.g. veth).

Thus, we need a clear flag to indicate that ->pre_exit_batch()
has already been called.

Let's add net->undo_state and update it only for dying netns.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 include/net/net_namespace.h | 9 +++++++++
 net/core/net_namespace.c    | 7 +++++++
 2 files changed, 16 insertions(+)
diff --git a/include/net/net_namespace.h b/include/net/net_namespace.h
index 58b2601bb869..6a817326fbb8 100644
--- a/include/net/net_namespace.h
+++ b/include/net/net_namespace.h
@@ -57,6 +57,9 @@ struct uevent_sock;
 struct netns_ipvs;
 struct bpf_prog;
 
+enum {
+	NET_PRE_EXIT_DONE = 1,
+};
 
 #define NETDEV_HASHBITS    8
 #define NETDEV_HASHENTRIES (1 << NETDEV_HASHBITS)
@@ -125,6 +128,7 @@ struct net {
 	 * it is critical that it is on a read_mostly cache line.
 	 */
 	u32			hash_mix;
+	u8			undo_state;
 
 	struct net_device       *loopback_dev;          /* The loopback */
 
@@ -364,6 +368,11 @@ static inline bool net_initialized(const struct net *net)
 	return READ_ONCE(net->list.next);
 }
 
+static inline bool net_pre_exit_done(const struct net *net)
+{
+	return READ_ONCE(net->undo_state) >= NET_PRE_EXIT_DONE;
+}
+
 static inline void __netns_tracker_alloc(struct net *net,
 					 netns_tracker *tracker,
 					 bool refcounted,
diff --git a/net/core/net_namespace.c b/net/core/net_namespace.c
index 476fbf913bad..9e2462bfa478 100644
--- a/net/core/net_namespace.c
+++ b/net/core/net_namespace.c
@@ -226,7 +226,9 @@ static void ops_undo_list(const struct list_head *ops_list,
 			  bool expedite_rcu)
 {
 	const struct pernet_operations *saved_ops;
+	bool dying = ops_list == &pernet_list;
 	bool hold_rtnl = false;
+	struct net *net;
 
 	if (!ops)
 		ops = list_entry(ops_list, typeof(*ops), list);
@@ -248,6 +250,11 @@ static void ops_undo_list(const struct list_head *ops_list,
 	else
 		synchronize_rcu();
 
+	if (dying) {
+		list_for_each_entry(net, net_exit_list, exit_list)
+			WRITE_ONCE(net->undo_state, NET_PRE_EXIT_DONE);
+	}
+
 	if (hold_rtnl)
 		ops_exit_rtnl_list(ops_list, saved_ops, net_exit_list);
 
-- 
2.56.0.rc1.315.gc6ed9934b7-goog
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help