Thread (4 messages) flat view 4 messages, 2 authors, 13h ago
HOTtoday

[PATCH net v2] net: yield the CPU on every exit of the threaded NAPI poll loop

From: Vitaliy Sochnev <hidden>
Date: 2026-09-02 19:01:34
Also in: lkml
Subsystem: networking [general], the rest · Maintainers: "David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, Linus Torvalds

napi_threaded_poll_loop() reaches cond_resched() only when it is about to
iterate. When __napi_poll() clears repoll the loop breaks first, and
napi_thread_wait() returns without scheduling if work is already pending.
Under a receive load arriving as fast as it is drained the kthread never
yields, and on CONFIG_PREEMPT_NONE nothing else on that CPU runs.

Everything waiting for deferred work on that CPU then blocks. Deleting a
netdev hits three such waits - synchronize_net(), flush_all_backlogs() ->
flush_work() and rcu_barrier() from netdev_run_todo() - which is how this
was found. run_backlog_napi() runs the same loop, so the backlog kthread
can be held off the same way.

rcu_softirq_qs_periodic() does not cover it: it reports a quiescent state
but does not schedule, so the work items and the callbacks still wait.
Moving only that call is not enough either - the delay then migrates from
synchronize_net() to flush_work() and rcu_barrier().

Without the patch the kernel reports the thread holding the CPU:

  rcu: INFO: rcu_sched self-detected stall on CPU
  rcu:  0-....: (5999 ticks this GP) ... (t=6000 jiffies g=913 q=1218)
  CPU: 0 UID: 0 PID: 203 Comm: napi/qdma_eth-0 Not tainted 6.18.44 #0
  Hardware name: Nokia XG-040G-MD (UBI) (DT)
  pc : __dma_sync_single_for_device+0x8/0xfc

Measured on that board (Airoha AN7581, quad core Cortex-A53, PREEMPT_NONE,
HZ=100, airoha_eth with threaded NAPI) while it terminates a 950 Mbit/s TCP
receive load. 30 minute runs, timing "ip link del" of a dummy interface:

                          before        after
  mean                    31.09 s       0.20 s
  worst                  151.92 s       1.00 s
  over 1 s              13 of 37      0 of 89
  RCU stalls, classic          14            0
  RCU stalls, expedited        43            0
  packet rate           79363 p/s    79520 p/s

The packet rate is the control: the same work is done in both runs, so the
difference is not a lighter load. Three of the four CPUs sat around 65%
idle throughout the first run and did not help - the deferred work the
delete waits for is tied to the CPU the poll loop holds.

On preemption, raised in v1: same board and load, unpatched, two halves
differing only in that choice - worst "ip link del" 248.48 s with 8 stalls
under PREEMPT_NONE against 0.39 s and none under PREEMPT_LAZY. So lazy
preemption does hide the symptom, and PREEMPT_NONE and PREEMPT_VOLUNTARY
builds are what is left. It is not the NAPI thread being preempted more -
nonvoluntary_ctxt_switches on it is 4.6/s under LAZY against 13.9/s under
PREEMPT_NONE - so that is a measurement, not a mechanism. The missing yield
is there under either model.

The loop is unchanged in Linus's tree - net/core/dev.c at v7.3-rc1 is
identical here to net/main. The numbers come from 6.18 because that is the
only kernel this board runs: mainline carries en7581-evb alone, while the
SoC dtsi, the board DTS and the airoha_eth changes it needs are still out
of tree. On 6.18 the loop has no busy_poll_last_qs parameter, but with that
pointer NULL the two are the same code, so the change under test is this
one. The busy-poll path is unaffected; cond_resched() was already reached
there.

Reproducing needs the loop re-entered tens of thousands of times a second,
which depends on the driver and the shape of the load rather than on this
board: threaded NAPI can be turned on for any driver through
/sys/class/net/<dev>/threaded, and what it takes on top is little or no
interrupt coalescing. It did not reproduce on mtk_eth_soc, whose net_dim
moderation folds the same packet rate into far fewer interrupts.

Fixes: 29863d41bb6e ("net: implement threaded-able napi poll loop support")
Signed-off-by: Vitaliy Sochnev <redacted>
---
v2:
  - answered the tree question: the loop is unchanged in Linus's tree at
    v7.3-rc1, and said why the numbers have to come from 6.18
  - re-ran the A/B on a kernel with no out-of-tree module, so the splat and
    the numbers now come from an untainted 6.18.44 build
  - added the preemption-model measurement, and why PREEMPT_NONE and
    PREEMPT_VOLUNTARY are still the exposed configs
  - noted the repro is not board-specific
  - shortened the comment; no other code change
v1: https://lore.kernel.org/netdev/20260814220427.623427-1-sochnev.v.74@gmail.com/ (local)

 net/core/dev.c | 10 ++++++----
 1 file changed, 6 insertions(+), 4 deletions(-)
diff --git a/net/core/dev.c b/net/core/dev.c
index 38336858c168..5c7f8cdf8443 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -7924,11 +7924,13 @@ static void napi_threaded_poll_loop(struct napi_struct *napi,
 			gro_flush_normal(&napi->gro, HZ >= 1000);
 		local_bh_enable();
 
-		/* Call cond_resched here to avoid watchdog warnings. */
-		if (repoll || busy_poll_last_qs) {
+		if (repoll || busy_poll_last_qs)
 			rcu_softirq_qs_periodic(last_qs);
-			cond_resched();
-		}
+
+		/* napi_thread_wait() can return without scheduling, so yield on
+		 * every exit, not only when the loop iterates.
+		 */
+		cond_resched();
 
 		if (!repoll)
 			break;
-- 
2.55.0
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help