Re: [PATCH v1 00/11] RCU: Enable callbacks to benefit from expedited grace periods
From: "Paul E. McKenney" <paulmck@kernel.org>
Date: 2026-07-24 00:31:17
Also in:
lkml, rcu
On Wed, Jun 24, 2026 at 06:23:42AM -0700, Puranjay Mohan wrote:
This series lets call_rcu() callbacks be reclaimed as soon as either a normal or an expedited grace period that covers them has elapsed, rather than always waiting for a normal grace period. Motivation ========== Today there is an asymmetry: synchronize_rcu_expedited() callers get fast reclaim, but call_rcu() callers never benefit from those same expedited grace periods, even though an expedited GP proves exactly the same thing as a normal one -- all pre-existing readers are done. When expedited GPs are running on the system (driven by other subsystems), call_rcu() callbacks that could already be freed instead sit in RCU_WAIT_TAIL until the next normal GP. This series treats a grace period as a grace period regardless of how it was driven, so memory is reclaimed sooner. Design ====== Callback segments now record both the normal and expedited grace-period sequence in struct rcu_gp_seq, and rcu_segcblist_advance() releases a segment as soon as poll_state_synchronize_rcu_full() reports that either has completed. Three notification paths are taught about expedited completion so the advance actually happens: the NOCB rcuog kthreads, the rcu_pending() tick gate, and rcu_core(). Changelog: RFC: https://lore.kernel.org/all/20260417231203.785172-1-puranjay@kernel.org/ (local) Changes in v1: - New prep patch 1 renames struct rcu_gp_oldstate to struct rcu_gp_seq and its fields rgos_norm/rgos_exp to norm/exp tree-wide (Frederic). - The rcu_segcblist segment field stays named gp_seq; only its type changes (Frederic). - Patch 8 (NOCB wake) is reworked. v1 woke the wrong waitqueue (rdp_gp->nocb_gp_wq via wake_nocb_gp() rather than the leaf rnp->nocb_gp_wq[] that an rcuog kthread waiting for a GP sleeps on), and the wait condition only checked the normal ->gp_seq. The rcuog grace-period wait now tracks a struct rcu_gp_seq and is released via poll_state_synchronize_rcu_full(); rcu_exp_wait_wake() wakes the leaf node through the new rcu_nocb_exp_cleanup() (Frederic). - rcu_pending() uses a new memory-ordering-free poll_state_synchronize_rcu_full_unordered() to avoid memory barriers on every tick, leaving the ordering duty to rcu_core() (Frederic). Still open: Frederic asked whether the first smp_mb() in poll_state_synchronize_rcu_full() is needed on the callback-advance path (patch 6). That path still uses the fully ordered helper; only rcu_pending() was switched to the unordered variant. Happy to revisit. Puranjay Mohan (11): rcu: Rename struct rcu_gp_oldstate to rcu_gp_seq rcu/segcblist: Add SRCU and Tasks RCU wrapper functions rcu/segcblist: Factor out rcu_segcblist_advance_compact() helper rcu/segcblist: Track segment grace periods with struct rcu_gp_seq rcu: Add RCU_GET_STATE_NOT_TRACKED for subsystems without expedited GPs rcu: Enable RCU callbacks to benefit from expedited grace periods rcu: Update comments for gp_seq and expedited GP tracking rcu: Wake NOCB rcuog kthreads on expedited grace period completion rcu: Detect expedited grace period completion in rcu_pending() rcu: Advance callbacks for expedited GP completion in rcu_core() rcuscale: Add concurrent expedited GP threads for callback scaling tests
I do see the occasional failure when running TREE01, as in a failure or four every couple hundred hours of testing. This is the only rcutorture scenario that exercises callback (de)offloading, which might (or might not) be a clue. The console log of a failing run is attached, or you can find it here: https://drive.google.com/file/d/1EH2P4SDy2a-SJndkN3GbCTGAXxI2xevm/view?usp=sharing This is a "??? Writer stall state" failure, where rcutorture never learns of grace-period completion. Possibly because no grace period completed. But there are no RCU CPU stall warnings. In addition, these two lines [ 631.032657] ??? Writer stall state RTWS_EXP_SYNC(4) g152813 f0x0 ->state D cpu 1 . . . [ 3610.863114] ??? Writer stall state RTWS_EXP_SYNC(4) g736160 f0x0 ->state D cpu 1 show that the grace-period sequence number advanced from 152813 to 736160 in about 50 minutes. So grace periods really are completing, lots of them. This rcutorture scenario has rcu_nocbs CPUs, and CPU 5 is in trouble: [ 631.066109] rcu: CB 5^4->4 KblSW F21044 L21044 C116 .W120(148380/168068).N295B101 q516 S CPU 16 . . . [ 3610.869914] rcu: CB 5^4->4 KblSW F3000848 L3000848 C116 .W120(148380/168068).N295B101 q516 S CPU 16 There are 120 callbacks on the WAIT segment, which are waiting on grace-period sequence number 148380, which was long gone even back at 631.032657 seconds in (g152813). There are 295 more callbacks in the NEXT segment, which long ago should have been assigned a grace period and moved into the currently empty NEXT_READY segment (which is the "." right before the "N295"). There are yet another 101 callbacks on the bypass list. This suggests that the rcuog kthread isn't being awakened when it should be. And that something is failing to advance callbacks out of NEXT in a timely fashion. Thoughts? Thanx, Paul
Attachments
- console.log [text/plain] 782639 bytes · preview