Thread (29 messages) 29 messages, 7 authors, 11d ago

Re: [PATCH v5 02/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines

flat view

From: Frederic Weisbecker <frederic@kernel.org>
Date: 2026-09-22 09:26:03
Also in: bpf, linux-trace-kernel, lkml, rcu

Le Tue, Sep 22, 2026 at 02:23:21AM +0000, Josef Bacik a écrit :
Tasks RCU waits for every task to pass through a voluntary context
switch, usermode or idle, because a preempted task might be sitting in a
trampoline that is about to be freed and nothing marks it as such.  With
PREEMPT_LAZY that is a poor fit for servers: cond_resched() is a no-op,
so a CPU-bound kthread only ever leaves the CPU by preemption, and one
such kthread holds every synchronize_rcu_tasks() caller -- ftrace and
BPF trampoline teardown under their mutexes, the kprobe jump optimizer
under text_mutex and cpus_read_lock() -- hostage for as long as it runs.

Following the discussion on v2, take the other road: let the
architecture make its trampolines Tasks Trace RCU readers.  When an
architecture selects HAVE_RCU_TRAMPOLINE_READERS it promises that every
trampoline whose lifetime Tasks RCU guards enters rcu_read_lock_trace()
(or its assembly equivalent) before calling out and leaves it before
returning, so a task anywhere inside such a call-out, preempted or not,
is an ordinary Tasks Trace reader.

That leaves the few instructions of trampoline text before the reader is
entered and after it is left (plus, in a later patch, the bytes a kprobe
jump optimization is about to overwrite). A task can only linger there
by being interrupted there, and such text never calls anything that
schedules, so instead of tracking tasks we track CPUs: every pass
through __schedule() is a per-CPU quiescent event, except that the one
context switch that can catch a task at an arbitrary instruction -- a
preemption from irq exit -- first records the interrupted IP in the task
and parks it on a per-CPU list for the duration (reusing the fields and
lists the classic flavor keeps for its exit-path bookkeeping), and, if
the IP is inside such "unmarked" text, puts the task on a short holdout
list; the task takes itself off at its next context switch outside such
a preemption or irq-exit check that finds it elsewhere. Usermode (the
existing tick hook, or a nohz_full CPU in an RCU extended quiescent
state) and idle count as well. rcu_tasks_trampoline_text() does the
classification: anything outside core and module text, plus an arch hook
for things like static ftrace stubs and return thunks.

The grace period, run by the existing rcu_tasks kthread so that
call_rcu_tasks(), synchronize_rcu_tasks() and rcu_barrier_tasks() keep
their names and callers, is: wait for every online CPU to context switch
or be seen in an RCU extended quiescent state (nudging stragglers with
resched_cpu() after a jiffy), drain the holdout list as it stood,
synchronize_rcu_tasks_trace() for everything inside the readers, then
one more CPU pass and drain for tasks that have since left the reader
into the trailing instructions. That is bounded by a few jiffies,
preempt-off latency and an SRCU grace period rather than by the longest
stretch any task runs without sleeping, needs no per-task scan, and
makes cond_resched_tasks_rcu_qs() unnecessary on such architectures.
Unlike the classic flavor it also waits for an idle task caught in a
trampoline, since an idle CPU only counts while RCU is not watching it.
rcu_tasks_wait_irq_preempted() walks the parked lists for the one caller
(the kprobe jump optimizer, later in the series) that makes ordinary
text unsafe to be parked in and so has to wait out tasks that were
preempted there before it said so.

The classic implementation is untouched and remains the default; the
new one is built only as CONFIG_TASKS_RCU_TRAMPOLINE_READERS when the
architecture opts in and uses the generic irq entry code, whose
reschedule check gains the rcu_tasks_irq_resched() call.  Nothing
selects it yet.

Suggested-by: Paul E. McKenney <paulmck@kernel.org>
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
One review might have fell into the cracks: https://lore.kernel.org/lkml/aqxLgT41UyA-bV5J@pavilion.home/ (local)

-- 
Frederic Weisbecker
SUSE Labs
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help