Thread (56 messages) flat view 56 messages, 10 authors, 14h ago

Re: [Linux PPC] Disable PREEMPT

From: Shrikanth Hegde <sshegde@linux.ibm.com>
Date: 2026-09-15 09:09:34

Hi.

On 9/15/26 1:05 PM, Michal Suchánek wrote:
On Tue, Sep 15, 2026 at 11:17:23AM +0530, Narayana Murty N wrote:
quoted

On 14/09/26 4:28 PM, Michal Suchánek wrote:
quoted
On Fri, Sep 11, 2026 at 10:39:10PM +0530, Shrikanth Hegde wrote:
quoted
Hi Michal,

On 9/10/26 5:46 PM, Michal Suchánek wrote:
quoted
On Thu, Sep 10, 2026 at 01:29:08PM +0200, Michal Suchánek wrote:
quoted
On Thu, Sep 10, 2026 at 04:11:20PM +0530, Shrikanth Hegde wrote:
quoted

On 9/10/26 2:29 PM, Michal Suchánek wrote:
quoted
On Fri, Sep 04, 2026 at 02:34:21PM +0530, Shrikanth Hegde wrote:
quoted

On 9/4/26 1:05 PM, Michal Suchánek wrote:
quoted
On Thu, Sep 03, 2026 at 10:52:59PM +0530, Shrikanth Hegde wrote:
quoted
quoted
quoted
If possible run against current upstream and share the results.
https://github.com/openSUSE/kernel-source/blob/6824496d1801f73def615dca8794202eeb7b0d86/config/ppc64le/default

[  472.091531][ T6181] Kernel panic - not syncing: stack-protector: Kernel stack is corrupted in: kvmhv_run_single_vcpu+0x19d4/0x1b50 [kvm_hv]
[  472.091598][ T6181] CPU: 29 UID: 107 PID: 6181 Comm: CPU 112/KVM Not tainted 7.2.2-5.g6824496-default #1 PREEMPT(full) openSUSE Tumbleweed (unreleased)  61871a5f06863b4006f5ef27cd9f18e8a7a2edad
[  472.091612][ T6181] Hardware name: IBM,9824-42A Power11 (architected) 0x820200 0xf000007 of:IBM,FW1110.20 (OB1110_130) hv:phyp pSeries
[  472.091624][ T6181] Call Trace:
[  472.091630][ T6181] [c00000000fdfb680] [c00000000134ce90] dump_stack_lvl+0x84/0xc0 (unreliable)
[  472.091653][ T6181] [c00000000fdfb6b0] [c00000000022e7c8] vpanic+0x324/0x5e4
[  472.091666][ T6181] [c00000000fdfb760] [c00000000022eac4] do_panic_on_target_cpu+0x0/0x2c
[  472.091677][ T6181] [c00000000fdfb780] [c0000000013c1ff8] __stack_chk_fail+0x48/0x60
[  472.091689][ T6181] [c00000000fdfb7f0] [c00800001aae219c] kvmhv_run_single_vcpu+0x19d4/0x1b50 [kvm_hv]
[  472.091712][ T6181] [c00000000fdfb940] [c00800001aae24b4] kvmppc_vcpu_run_hv+0x19c/0x12f0 [kvm_hv]
[  472.091732][ T6181] [c00000000fdfba10] [c00800001aeeed18] kvmppc_vcpu_run+0x30/0x48 [kvm]
[  472.091779][ T6181] [c00000000fdfba30] [c00800001aee9ef4] kvm_arch_vcpu_ioctl_run+0x35c/0x4a0 [kvm]
[  472.091813][ T6181] [c00000000fdfbac0] [c00800001aedaac4] kvm_vcpu_ioctl+0x1ac/0xad8 [kvm]
[  472.091844][ T6181] [c00000000fdfbca0] [c0000000007f1244] sys_ioctl+0x374/0x1060
[  472.091857][ T6181] [c00000000fdfbdb0] [c00000000002f7f8] system_call_exception+0x188/0x430
[  472.091871][ T6181] [c00000000fdfbe50] [c00000000000cfdc] system_call_vectored_common+0x15c/0x2ec
[  472.091886][ T6181] ---- interrupt: 3000 at 0x7fffb5565fac
[  472.091896][ T6181] NIP:  00007fffb5565fac LR: 00007fffb5565fac CTR: 0000000000000000
[  472.091904][ T6181] REGS: c00000000fdfbe80 TRAP: 3000   Not tainted  (7.2.2-5.g6824496-default)
[  472.091911][ T6181] MSR:  800000000280f033 <SF,VEC,VSX,EE,PR,FP,ME,IR,DR,RI,LE>  CR: 42044402  XER: 00000000
[  472.091938][ T6181] IRQMASK: 0
[  472.091938][ T6181] GPR00: 0000000000000036 00007fbfa77ed7a0 00007fffb5677100 00000000000000fa
[  472.091938][ T6181] GPR04: 000000002000ae80 0000000000000000 0000000000000000 0000000000000000
[  472.091938][ T6181] GPR08: 00000000000000fa 0000000000000000 0000000000000000 0000000000000000
[  472.091938][ T6181] GPR12: 0000000000000000 00007fbfa77f5ec0 000000014676f000 00007fbfa77ee7c0
[  472.091938][ T6181] GPR16: 000000014674e8d0 00007fbfa77eeec0 00007fbfa77eeec0 fffffffffffffff7
[  472.091938][ T6181] GPR20: 00007fffb71210d0 0000000000000001 00007fbfa77eeec0 0000000000000000
[  472.091938][ T6181] GPR24: 00007fbfa77ed8e8 0000000105971428 000000002000ae80 0000000105f77a70
[  472.091938][ T6181] GPR28: 0000000000000000 0000000000000000 000000002000ae80 000000014674f000
[  472.092020][ T6181] NIP [00007fffb5565fac] 0x7fffb5565fac
[  472.092027][ T6181] LR [00007fffb5565fac] 0x7fffb5565fac
[  472.092033][ T6181] ---- interrupt: 3000
[  472.098256][ T6181] pstore: backend (nvram) writing error (-1)

This is the host, cannot run the kernel as guest because it fails to boot most
of the time inside KVM.
quoted
quoted
Nonethless, there are quite a few platforms. Originally no preemption
was the only option, and that's the reason why many people run that.
It's the conservative, known working option. And that's the reason a lot
of platfrom code does not get tested with more aggressive preemtion
models, and never gets fixed to work with them.
Full preemption has been there for many years!.
Possible for years, forced only recently.
quoted
Lazy is not that aggressive compared to that.
quoted
Simply disabling the no preemtion option does not make the platform code
ready.
Let's understand your crash case. Let's see where it is going wrong. I am suspecting
it is some wrong usage of preemption api rather than arch can't support preemption.
Very likely some wrong use of the preemption API by the arch code, or no
use where it should have been used. It did not matter so long as people
could run their no preempt configs and ignore the problem.

Thanks

Friendly LLM analysis says preemption is enabled too early and before
completing kvmppc_handle_exit_hv() / kvmppc_handle_nested_exit()

Below is ONLY a speculation and completely UNTESTED.
Maybe worth a try.
The patch is munged by the e-mail client, and it causes immediate
voluntary preemprion in rcu critical section and hard lockup on starting
a KVM VM.

Also it would be sort of bad news if it worked because that would be
specific to book3s KVM HV and would not help with the KVM HV from the
original report which likely is not book3s, nor with KVM PR.
Thanks for trying. We will try a local repro and look into it why stack is
getting corrupted.
As we discussed offlist, samir helped to run a similar test on his machine, and he didn't
run into issue so far. we will try more.

I am just wondering what different in your case?
By any chance we are running into below one? Can you check your gcc version?

https://lore.kernel.org/all/CAABZP2z=xu+07-y5fqFLidZz1VpSgrSwXa1mFHPb=b3Ezr3OtA@mail.gmail.com/ (local)

Maybe CONFIG_DEBUG_PREEMPT worth a try to if it shows up anything.
Does not change anything, the kernel crashes all the same.

Nonetheless, I noticed that a tool that does some BPF tracing is
running, and without it the problem is not reproducible.

Thanks

Michal
Thanks Michal for the update.

Interesting finding! Since the crash is only reproducible when the BPF
tracing tool is running, could you share the following details?

Which BPF tracing tool are you using (e.g., bpftrace, perf, bcc-based tool,
a custom tool)?
A custom tool, likely a version of this:
https://build.opensuse.org/package/show/security:sensor/velociraptor
quoted
What is the specific BPF program or filter being attached —
particularly which kernel function(s) or tracepoints it hooks into (e.g., is
it attaching a kprobe/uprobe/tracepoint anywhere inside
kvmhv_run_single_vcpu or related KVM HV paths)?
What is the BPF program type — kprobe, tracepoint, perf_event, fentry/fexit,
etc.?
It would be likely these programs:
https://github.com/SUSE/linux-security-sensor/tree/sensor-base-0.76.7/vql/linux/bpf

Thanks

Michal
It seems easy to re-create.
The key is running bcc/bpf program in parallel to kernel build.

For example, these two in parallal, leads to stack corruption panic.
Stacktrace does differ sometimes.

./funccount  rcu* -d 100
make -j 32

Boom. Will try to debug this further.

[ 3979.381003][    T0] Kernel panic - not syncing: stack-protector: Kernel stack is corrupted in: prb_reserve+0x498/0x4a0
[ 3979.385530][  C114] pstore: backend (nvram) writing error (-1)
[ 3980.385590][  C114] Oops: System Reset, sig: 6 [#1]
[ 3980.385605][  C114] LE PAGE_SIZE=64K MMU=Radix  SMP NR_CPUS=2048 NUMA pSeries
[ 3980.385611][  C114] Modules linked in: bonding pseries_rng rng_core vmx_crypto fuse ibmvfc ibmveth dm_mirror dm_region_hash dm_log autofs4
[ 3980.385630][  C114] CPU: 114 UID: 0 PID: 0 Comm: swapper/114 Not tainted 7.3.0-rc1-00026-gd3d9a20b4eb9 #23 PREEMPT(full)
[ 3980.385634][  C114] Hardware name: IBM,9043-MRU Power11 (architected) 0x820200 0xf000007 of:IBM,FW1120.00 (RF1120_183) hv:phyp pSeries
[ 3980.385636][  C114] NIP:  c0000000001ae714 LR: c000000001552b08 CTR: 0000000000000000
[ 3980.385638][  C114] REGS: c0000007ffc2fd60 TRAP: 0100   Not tainted  (7.3.0-rc1-00026-gd3d9a20b4eb9)
[ 3980.385641][  C114] MSR:  800000000298b033 <SF,VEC,VSX,EE,FP,ME,IR,DR,RI,LE>  CR: 22000202  XER: 20040006
[ 3980.385649][  C114] CFAR: 000000000000011c IRQMASK: 3
[ 3980.385649][  C114] GPR00: 0000000000000000 c00000040030fd50 c000000001be8100 0000000000000000
[ 3980.385649][  C114] GPR04: 0000000000000001 000000000000003a 0000000000000000 0000010000000000
[ 3980.385649][  C114] GPR08: ffffffffffffffbf 0000000000000000 ffffffffffffffff 0000000000000000
[ 3980.385649][  C114] GPR12: 0000000000000000 c0000007fffdeb00 0000000000000000 000000002ef01f40
[ 3980.385649][  C114] GPR16: 0000000000000000 0000000000000000 0000000000000000 0000000000000000
[ 3980.385649][  C114] GPR20: 0000000000000000 0000000000000000 0000000000000000 0000000000000001
[ 3980.385649][  C114] GPR24: 0000000000000001 0000000000000000 0000039e85974a0e c0000000029ab5a8
[ 3980.385649][  C114] GPR28: c0000007fe8f9a40 0000000000000001 c0000000022020a8 c0000000022020b0
[ 3980.385675][  C114] NIP [c0000000001ae714] plpar_hcall_norets_notrace+0x18/0x2c
[ 3980.385684][  C114] LR [c000000001552b08] check_and_cede_processor+0x48/0x60
[ 3980.385692][  C114] Call Trace:
[ 3980.385693][  C114] [c00000040030fd50] [0000000042000202] 0x42000202 (unreliable)
[ 3980.385703][  C114] [c00000040030fdb0] [c000000001552d30] shared_cede_loop+0x70/0x170
[ 3980.385708][  C114] [c00000040030fdf0] [c000000001552170] cpuidle_enter_state+0x300/0x748
[ 3980.385712][  C114] [c00000040030fe90] [c0000000011394f0] cpuidle_enter+0x50/0x80
[ 3980.385718][  C114] [c00000040030fed0] [c0000000002a5a78] call_cpuidle+0x48/0x90
[ 3980.385722][  C114] [c00000040030fef0] [c0000000002acb0c] do_idle+0x2cc/0x470
[ 3980.385725][  C114] [c00000040030ff60] [c0000000002acf7c] cpu_startup_entry+0x4c/0x50
[ 3980.385729][  C114] [c00000040030ff90] [c00000000005be80] start_secondary+0x290/0x2a0
[ 3980.385733][  C114] [c00000040030ffe0] [c00000000000e258] start_secondary_prolog+0x10/0x14
[ 3980.385737][  C114] Code: 3d220063 39295780 7c634a78 7c630074 7863d182 4e800020 3c4c01a4 38429a04 7c421378 7c000026 90010008 44000022 <38800000> 988d0931 80010008 7c0ff120

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help