Thread (12 messages) 12 messages, 5 authors, 13d ago

Re: [PATCH net-next] net/sched: fq_codel, cake: widen backlogs to u64

flat view

From: Jamal Hadi Salim <jhs@mojatatu.com>
Date: 2026-09-25 11:28:08

On Fri, Sep 25, 2026 at 5:13 AM Eric Dumazet [off-list ref] wrote:
On Fri, Sep 25, 2026 at 10:54 AM Jamal Hadi Salim [off-list ref] wrote:
quoted
This is a follow-up to commit 8f735d64382d ("net/sched: bound
qdisc_pkt_len to prevent qdisc soft lockup"), which capped
qdisc_pkt_len() at QDISC_PKT_LEN_MAX (1 MiB). That cap bounds the stab
amplifier but leaves the per-flow backlog counter u32:
fq_codel_enqueue() accumulates qdisc_pkt_len(skb) into q->backlogs[idx],
so a flow can still accumulate 4096 packets of 1 MiB each and wrap the
counter mod 2^32. After a wrap, fq_codel_drop() sees a tiny maxbacklog
and drops from an almost-empty flow, and the dequeue-side subtractions
corrupt the counter further.

Widen the fq_codel backlogs table, the fat-flow scan (maxbacklog/len) and
the drop threshold to u64. fq_codel is not lockless: every writer runs
under the root qdisc lock, so plain u64 arithmetic keeps the WRITE_ONCE
publish / READ_ONCE-consume pattern. The dump path
(fq_codel_dump_class_stats) stays a lockless stat-only read.

CAKE accumulates the same generic qdisc_pkt_len(skb) into its per-flow
b->backlogs[] and per-tin b->tin_backlog and consumes the values for
longest-flow pruning (cake_heapify/cake_heapify_up) and for the shaper
staleness check, so it shares the bug. Widen those counters and the heap
comparison locals to u64; the class/tin stats keep exporting the low 32
bits through the unchanged uAPI fields.

Conditions to recreate the bug: CAP_NET_ADMIN in a user namespace;
CONFIG_NET_SCH_FQ_CODEL=y.

  ip tuntap add tun0 mode tun
  ip link set tun0 txqueuelen 32 up
  ip addr add 10.99.0.1/24 dev tun0
  tc qdisc add dev tun0 root handle 1: stab overhead 2000000000 \
      fq_codel flows 1 limit 4200 ecn drop_batch 4096
  # hold the tun fd open without reading (IFF_BACKPRESSURE) so the qdisc
  # backlog persists, then send at least 4300 packets (the wrap starts
  # at 4096 resident; the over-limit drop that reads the wrapped
  # threshold fires past the 4200 limit): backlogs[0] wraps at 4096 x
  # 1 MiB and the fat-flow threshold reads the wrapped value.

With a 1 MiB qdisc_pkt_len cap the counter wraps at 4096 resident
packets. At limit 4200 the first over-limit enqueue (the 4201st) sees a
wrapped 105 MiB (half-backlog threshold 52 MiB, a ~52 packet drop burst),
where the u64 counter sees 4201 MiB (threshold 2100 MiB, a ~2100 packet
drop burst).

Reported-by: Sashiko (nipa) <sashiko-bot@kernel.org>
Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260818101130.16203-1-jhs@mojatatu.com
Link: https://lore.kernel.org/netdev/20260818101130.16203-1-jhs@mojatatu.com/ (local)
Tested-by: hybris <redacted>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
---
This makes no sense.
As absurd as it looks that code is reachable ;->
These qdisc have been developped to address bufferbloat issues.

Storing 4GB in a qdisc is absolutely insane.
The counter wrap is not because we stored 4GB, rather it is because
"tc .. stab overhead ..." inflates qdisc_pkt_len() for a 64B pkt to
1MB. So ~4K packets (put in other words a few "real" KB) makes that
backlog[0] cross 2^32.
Result is pruning the wrong flow..
Let's drop at enqueue if the current backlog is approaching 4GB (this
can later be a new config/attribute in net-next)
Note, this is net-next already. Enqueue is fast path - are you ok with that?

cheers,
jamal
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help