Thread (18 messages) 18 messages, 3 authors, 2026-04-02

Re: "Dead loop on virtual device" error without softirq-BKL on PREEMPT_RT

From: Bert Karwatzki <hidden>
Date: 2026-02-17 11:24:44
Also in: linux-rt-devel, lkml
Subsystem: networking [general], the rest · Maintainers: "David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, Linus Torvalds

Am Dienstag, dem 17.02.2026 um 11:42 +0100 schrieb Bert Karwatzki:
I just wondered if we can completely skip the

	if (READ_ONCE(txq->xmit_lock_owner) != cpu) {
		[...]
	} else 
	{
		/* "Recursion" alert */
	}
 
check, as the synchronization will we provided by HARD_TX_{LOCK,UNLOCK}.
I thought about that again, and it seems like a bad idea as (in the non-preempt) case
other threads trying to access the queue would wait for the spinlock to be freed, perhaps
one can just change the code like this:

commit 05026868843a4eea51d45811d87706f36896e828
Author: Bert Karwatzki [off-list ref]
Date:   Tue Feb 17 12:08:35 2026 +0100

    net: core: dev: don't warn about recursion when on same CPU
    
    This prints a message if we're on the same CPU and the lock is
    already taken, in a production use we would of course skip this
    message.
    
    Signed-off-by: Bert Karwatzki [off-list ref]
diff --git a/net/core/dev.c b/net/core/dev.c
index 5b536860138d..cac5588640b3 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -4784,15 +4784,18 @@ int __dev_queue_xmit(struct sk_buff *skb, struct net_device *sb_dev)
 			net_crit_ratelimited("Virtual device %s asks to queue packet!\n",
 					     dev->name);
 		} else {
-			/* Recursion is detected! It is possible,
-			 * unfortunately
-			 */
-recursion_alert:
-			net_crit_ratelimited("Dead loop on virtual device %s, fix it urgently!\n",
-					     dev->name);
+			net_crit_ratelimited("Lock taken already on %s!\n", dev->name);
+			goto lock_taken;
 		}
 	}
 
+	/* Recursion is detected! It is possible,
+	 * unfortunately
+	 */
+recursion_alert:
+	net_crit_ratelimited("Dead loop on virtual device %s, fix it urgently!\n",
+					     dev->name);
+lock_taken:
 	rc = -ENETDOWN;
 	rcu_read_unlock_bh();
 
With this I get these messages (be skipped in production code)
instead of the "Dead loop on virtual device":

[   51.994435] [   T1525] Lock taken already on wlp4s0!
[   51.994546] [   T1525] Lock taken already on wlp4s0!
[   51.994650] [   T1525] Lock taken already on wlp4s0!
[   51.994746] [   T1525] Lock taken already on wlp4s0!
[   51.994845] [   T1525] Lock taken already on wlp4s0!
[   51.994948] [   T1525] Lock taken already on wlp4s0!
[   51.995037] [   T1525] Lock taken already on wlp4s0!
[   51.995128] [   T1525] Lock taken already on wlp4s0!
[   51.995220] [   T1525] Lock taken already on wlp4s0!


Bert Karwatzki
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help