Thread (2 messages) flat view 2 messages, 2 authors, 18h ago
HOTtoday REVIEWED: 4 (0M)

4 review trailers.

[PATCH net v3] net/sched: taprio: catch up in bounded time when the schedule falls behind

From: Junjie Cao <hidden>
Date: 2026-09-01 09:33:56
Also in: lkml
Subsystem: cbs/etf/taprio qdiscs, networking [general], tc subsystem, the rest · Maintainers: Vinicius Costa Gomes, "David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, Jamal Hadi Salim, Jiri Pirko, Linus Torvalds

advance_sched() advances exactly one entry per hrtimer expiry. When the
operational schedule falls behind - the timer was delayed, the CPU was
starved, or the reference clock stepped forward - every elapsed entry is
replayed back to back from hrtimer context with current_entry_lock held,
and each replay rearms the timer with an expiry in the past. Once the
backlog is large enough the CPU never leaves timer processing and RCU
stalls follow. syzbot triggers this with schedules whose intervals are
shorter than the cost of servicing one expiry, so the backlog only ever
grows.

Skip whole periods arithmetically and walk at most one more to the entry
covering the current time. The software schedule restarts the list after
its last entry even when that is before cycle_time, so its period is
min(cycle_time, sum of intervals); record it at parse time. Advance
cycle_end_time by the period as well: with cycle_time it runs ahead of
the entries by the difference every lap, and skipping whole laps at once
would overflow it within minutes for a schedule with nanosecond
intervals and a cycle_time of seconds. Gate close times and budgets are
still only computed for the entry landed on, and an admin schedule
crossed by the jump is picked up by the existing
should_change_schedules() check on the recomputed end time. The walk is
capped at twice the entry count; a leftover is handled by the next
expiry as today.

Fixes: 5a781ccbd19e ("tc: Add support for configuring the taprio scheduler")
Reported-by: syzbot+19d01f6082ec61dd45b2@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=19d01f6082ec61dd45b2
Tested-by: syzbot+19d01f6082ec61dd45b2@syzkaller.appspotmail.com
Reported-by: syzbot+8785aaf121cfb2141e0d@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=8785aaf121cfb2141e0d
Tested-by: syzbot+8785aaf121cfb2141e0d@syzkaller.appspotmail.com
Reported-by: syzbot+2642f347f7309b4880dc@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=2642f347f7309b4880dc
Tested-by: syzbot+2642f347f7309b4880dc@syzkaller.appspotmail.com
Reported-by: syzbot+e4aa91d7f20c34417d4e@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=e4aa91d7f20c34417d4e
Tested-by: syzbot+e4aa91d7f20c34417d4e@syzkaller.appspotmail.com
Signed-off-by: Junjie Cao <redacted>
---
v3:
- jump in whole periods of min(cycle_time, sum of intervals) rather
  than cycle_time, and fold the walk into advance_sched() instead of a
  helper duplicating its step (Jakub)
- treat end == now as behind; comment on the cap and on rewriting the
  published entry (Jakub)
- advance cycle_end_time by the period as well, not left alone as I
  said in the v2 thread: with cycle_time it runs ahead of the entries
  by (cycle_time - sum) per lap, and skipping whole laps turns that
  into an s64 overflow within minutes for a schedule with nanosecond
  intervals and a cycle_time of seconds
- the walk finishes within num_entries steps once the first expiry has
  been serviced; a schedule whose cycle_time is shorter than its first
  entry starts with cycle_end_time behind that entry and takes one more
  expiry to line up, where today it replays until it does
- drop the minimum-interval patch and its selftest (Jakub)
- tags for a fourth syzbot bucket, tested by akpm with the v2 patch
  alone; all four re-tested with this version
v2: https://lore.kernel.org/all/20260820062715.278124-1-junjie.cao@intel.com/ (local)
- take now from the timer's clock base (Hillf Danton)
v1: https://lore.kernel.org/all/20260818071706.251035-1-junjie.cao@intel.com/ (local)

 net/sched/sch_taprio.c | 72 ++++++++++++++++++++++++++++++------------
 1 file changed, 52 insertions(+), 20 deletions(-)
diff --git a/net/sched/sch_taprio.c b/net/sched/sch_taprio.c
index 39ac5b97aa3a..901dfd2484e1 100644
--- a/net/sched/sch_taprio.c
+++ b/net/sched/sch_taprio.c
@@ -83,6 +83,10 @@ struct sched_gate_list {
 	s64 cycle_time;
 	s64 cycle_time_extension;
 	s64 base_time;
+	/* min(cycle_time, sum of intervals): the software schedule restarts
+	 * the list after the last entry even when cycle_time is not up yet.
+	 */
+	s64 period;
 };
 
 struct taprio_sched {
@@ -871,12 +875,13 @@ static struct sk_buff *taprio_dequeue(struct Qdisc *sch)
 }
 
 static bool should_restart_cycle(const struct sched_gate_list *oper,
-				 const struct sched_entry *entry)
+				 const struct sched_entry *entry,
+				 ktime_t end_time)
 {
 	if (list_is_last(&entry->list, &oper->entries))
 		return true;
 
-	if (ktime_compare(entry->end_time, oper->cycle_end_time) == 0)
+	if (ktime_compare(end_time, oper->cycle_end_time) == 0)
 		return true;
 
 	return false;
@@ -925,8 +930,9 @@ static enum hrtimer_restart advance_sched(struct hrtimer *timer)
 	int num_tc = netdev_get_num_tc(dev);
 	struct sched_entry *entry, *next;
 	struct Qdisc *sch = q->root;
-	ktime_t end_time;
-	int tc;
+	ktime_t end_time, next_start, now;
+	int budget, tc;
+	s64 behind;
 
 	spin_lock(&q->current_entry_lock);
 	entry = rcu_dereference_protected(q->current_entry,
@@ -952,23 +958,49 @@ static enum hrtimer_restart advance_sched(struct hrtimer *timer)
 		goto first_run;
 	}
 
-	if (should_restart_cycle(oper, entry)) {
-		next = list_first_entry(&oper->entries, struct sched_entry,
-					list);
-		oper->cycle_end_time = ktime_add_ns(oper->cycle_end_time,
-						    oper->cycle_time);
-	} else {
-		next = list_next_entry(entry, list);
+	now = hrtimer_cb_get_time(timer);
+	end_time = entry->end_time;
+	behind = ktime_sub(now, end_time);
+
+	/* Behind, e.g. delayed timer or stepped clock: skip whole periods
+	 * arithmetically and walk at most one more to the entry covering
+	 * now, instead of replaying the backlog one expiry at a time. The
+	 * cap bounds the walk; a leftover is picked up by the next expiry.
+	 */
+	if (unlikely(behind >= oper->period)) {
+		s64 jump = div64_s64(behind, oper->period) * oper->period;
+
+		end_time = ktime_add_ns(end_time, jump);
+		oper->cycle_end_time = ktime_add_ns(oper->cycle_end_time, jump);
 	}
 
-	end_time = ktime_add_ns(entry->end_time, next->interval);
-	end_time = min_t(ktime_t, end_time, oper->cycle_end_time);
+	budget = 2 * oper->num_entries;
+	do {
+		if (should_restart_cycle(oper, entry, end_time)) {
+			next = list_first_entry(&oper->entries,
+						struct sched_entry, list);
+			oper->cycle_end_time = ktime_add_ns(oper->cycle_end_time,
+							    oper->period);
+		} else {
+			next = list_next_entry(entry, list);
+		}
+
+		next_start = end_time;
+		end_time = ktime_add_ns(next_start, next->interval);
+		end_time = min_t(ktime_t, end_time, oper->cycle_end_time);
+		entry = next;
+	} while (unlikely(ktime_compare(end_time, now) <= 0) && budget--);
 
+	/* next can be the entry already published as q->current_entry (a
+	 * single-entry schedule, or a catch-up of whole periods), so the
+	 * close times and budgets below are rewritten in place while
+	 * taprio_dequeue_from_txq() may be reading them.
+	 */
 	for (tc = 0; tc < num_tc; tc++) {
 		if (next->gate_duration[tc] == oper->cycle_time)
 			next->gate_close_time[tc] = KTIME_MAX;
 		else
-			next->gate_close_time[tc] = ktime_add_ns(entry->end_time,
+			next->gate_close_time[tc] = ktime_add_ns(next_start,
 								 next->gate_duration[tc]);
 	}
 
@@ -1130,6 +1162,8 @@ static int parse_taprio_schedule(struct taprio_sched *q, struct nlattr **tb,
 				 struct sched_gate_list *new,
 				 struct netlink_ext_ack *extack)
 {
+	struct sched_entry *entry;
+	ktime_t cycle = 0;
 	int err = 0;
 
 	if (tb[TCA_TAPRIO_ATTR_SCHED_SINGLE_ENTRY]) {
@@ -1152,13 +1186,10 @@ static int parse_taprio_schedule(struct taprio_sched *q, struct nlattr **tb,
 	if (err < 0)
 		return err;
 
-	if (!new->cycle_time) {
-		struct sched_entry *entry;
-		ktime_t cycle = 0;
-
-		list_for_each_entry(entry, &new->entries, list)
-			cycle = ktime_add_ns(cycle, entry->interval);
+	list_for_each_entry(entry, &new->entries, list)
+		cycle = ktime_add_ns(cycle, entry->interval);
 
+	if (!new->cycle_time) {
 		if (cycle < 0 || cycle > INT_MAX) {
 			NL_SET_ERR_MSG(extack, "'cycle_time' is too big");
 			return -EINVAL;
@@ -1172,6 +1203,7 @@ static int parse_taprio_schedule(struct taprio_sched *q, struct nlattr **tb,
 		return -EINVAL;
 	}
 
+	new->period = min(new->cycle_time, cycle);
 	taprio_calculate_gate_durations(q, new);
 
 	return 0;
-- 
2.43.0
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help