Re: [PATCH 09/19] arm64: smp: Defer RCU registration during secondary CPU bringup
From: Jinjie Ruan <hidden>
Date: 2026-09-10 02:48:20
Also in:
lkml
在 2026/9/9 20:36, Will Deacon 写道:
On Tue, Sep 08, 2026 at 07:25:54PM +0800, Jinjie Ruan wrote:quoted
在 2026/9/8 18:19, Will Deacon 写道:quoted
On Tue, Sep 08, 2026 at 04:55:36PM +0800, Jinjie Ruan wrote:quoted
I think we need to handle the printk problem before this patch as we discussed earlier. Otherwise defer the rcutree_report_cpu_starting() will trigger a false-positive lockdep"suspicious RCU usage" splat during early lock acquisitions as commit ce3d31ad3cac ("arm64/smp: Move rcu_cpu_starting() earlier") pointed out.Sorry, I meant to mention this in the cover letter but forgot about it. I'm not sure that ce3d31ad3cac ("arm64/smp: Move rcu_cpu_starting() earlier") is still relevant with the latest printk/console/lockdep code. I tried quite hard to trigger lockdep splats manually, but the only way I could do it was by using the "%pS" specifier to print the name of a symbol in a module, which would cause an RCU walk of the module symbols in the kallsyms code! Manually calling WARN() or even rcu_read_lock() / spin_lock() did _not_ trigger a splat.Add "dyndbg="+p"" in cmdline, CONFIG_DEBUG_LOCK_ALLOC=y, CONFIG_PROVE_RCU_LIST=y, we can reproduce the warning as below: I believe there is also a problem in the RISC-V code itself here as store_cpu_topology() is common for RISC-V. [ 0.335162] smp: Bringing up secondary CPUs ... [ 0.345495] [ 0.345513] ============================= [ 0.345523] WARNING: suspicious RCU usage [ 0.345621] 7.3.0-rc2-00010-g2311ba2cd56f #500 Tainted: G W [ 0.345637] ----------------------------- [ 0.345646] kernel/locking/lockdep.c:3845 RCU-list traversed in non-reader section!! [ 0.345659] [ 0.345659] other info that might help us debug this: [ 0.345659] [ 0.345680] [ 0.345680] RCU used illegally from offline CPU! [ 0.345680] rcu_scheduler_active = 1, debug_locks = 1 [ 0.345725] locks held by swapper/1/0: 0, last CPU#1 [ 0.345743] [ 0.345743] stack backtrace: [ 0.345834] CPU: 1 UID: 0 PID: 0 Comm: swapper/1 Tainted: G W 7.3.0-rc2-00010-g2311ba2cd56f #500 PREEMPT(full) [ 0.345885] Tainted: [W]=WARN [ 0.346077] Call trace: [ 0.346102] show_stack+0x20/0x38 (C) [ 0.346153] dump_stack_lvl+0xc4/0x150 [ 0.346176] dump_stack+0x18/0x28 [ 0.346194] lockdep_rcu_suspicious+0x170/0x238 [ 0.346217] __lock_acquire+0xf08/0x1818 [ 0.346237] lock_acquire+0x1e0/0x450 [ 0.346256] _raw_spin_lock_irqsave+0x70/0xc0 [ 0.346277] down_trylock+0x20/0x60 [ 0.346293] __down_trylock_console_sem+0x4c/0x118 [ 0.346316] vprintk_emit+0x2d8/0x3f8 [ 0.346333] vprintk_default+0x40/0x58 [ 0.346350] vprintk+0x3c/0x80 [ 0.346366] _printk+0x64/0x98 [ 0.346386] __dynamic_pr_debug+0x90/0xd8 [ 0.346406] acpi_get_cache_info+0x140/0x1a0 [ 0.346430] init_cache_level+0xec/0x110 [ 0.346450] detect_cache_attributes+0x74/0x7c0 [ 0.346473] update_siblings_masks+0x30/0x300 [ 0.346495] store_cpu_topology+0x70/0xf0 [ 0.346515] secondary_start_kernel+0xe0/0x178 [ 0.346535] __secondary_switched+0xc0/0xc8I was about to say "don't do this" but then I realised two things: 1. update_siblings_masks() can trigger lockdep splats outside of pr_debug() if RCU isn't up and running, e.g.: [ 0.524042] show_stack+0x18/0x24 (C) [ 0.524519] __dump_stack+0x28/0x38 [ 0.524546] dump_stack_lvl+0x64/0x84 [ 0.524562] dump_stack+0x18/0x24 [ 0.524576] lockdep_rcu_suspicious+0x134/0x1cc [ 0.524591] __lock_acquire+0xee8/0x2cb0 [ 0.524606] lock_acquire+0x11c/0x2fc [ 0.524621] _raw_spin_lock_irqsave+0x64/0x84 [ 0.524641] of_find_property+0x2c/0x8c [ 0.524659] detect_cache_attributes+0x1c0/0x6d0 [ 0.524676] update_siblings_masks+0x38/0x288 [ 0.524692] store_cpu_topology+0x4c/0x58 [ 0.524706] secondary_start_kernel+0xdc/0x1c8 [ 0.524722] __secondary_switched+0x120/0x124 2. This code is running _after_ cpuhp_ap_sync_alive(). So for the next version, I'll reintroduce the call to rcutree_report_cpu_starting(), but move it immediately after the call to cpuhp_ap_sync_alive(). I think that will solve these issues, without
pr_crit() and pr_warn() (such as vec_verify_vq_map()) in
check_local_cpu_capabilities() can also trigger lockdep splats as below.
But I think this is not common on the failure path, so it seems to have
little impact..
[ 0.158619] smp: Bringing up secondary CPUs ...
[ 0.173958] CPU1: missing HWCAP.
[ 0.174071]
[ 0.174088] =============================
[ 0.174099] WARNING: suspicious RCU usage
[ 0.174197] 7.3.0-rc2-00020-gef0bd63bdd97-dirty #504 Tainted: G W
[ 0.174217] -----------------------------
[ 0.174226] kernel/locking/lockdep.c:3845 RCU-list traversed in
non-reader section!!
[ 0.174240]
[ 0.174240] other info that might help us debug this:
[ 0.174240]
[ 0.174262]
[ 0.174262] RCU used illegally from offline CPU!
[ 0.174262] rcu_scheduler_active = 1, debug_locks = 1
[ 0.174306] locks held by swapper/1/0: 0, last CPU#1
[ 0.174326]
[ 0.174326] stack backtrace:
[ 0.174412] CPU: 1 UID: 0 PID: 0 Comm: swapper/1 Tainted: G W
7.3.0-rc2-00020-gef0bd63bdd97-dirty #504 PREEMPT(full)
[ 0.174462] Tainted: [W]=WARN
[ 0.174488] Call trace:
[ 0.174513] show_stack+0x20/0x38 (C)
[ 0.174568] dump_stack_lvl+0xc4/0x150
[ 0.174594] dump_stack+0x18/0x28
[ 0.174615] lockdep_rcu_suspicious+0x170/0x238
[ 0.174641] __lock_acquire+0xf08/0x1818
[ 0.174664] lock_acquire+0x1e0/0x450
[ 0.174687] _raw_spin_lock_irqsave+0x70/0xc0
[ 0.174709] down_trylock+0x20/0x60
[ 0.174728] __down_trylock_console_sem+0x4c/0x118
[ 0.174748] vprintk_emit+0x2d8/0x3f8
[ 0.174769] vprintk_default+0x40/0x58
[ 0.174788] vprintk+0x3c/0x80
[ 0.174808] _printk+0x64/0x98
[ 0.174832] secondary_start_kernel+0xc8/0x190
[ 0.174857] __secondary_switched+0x120/0x128
causing issues with the concurrent part of early boot and also without reintroducing the early call to rcutree_report_cpu_dead(). Cheers, Will