[PATCH 0/2] KVM: PPC: BookE: Fix boot hang on preemptible kernels

WARM1d

8 messages, 3 authors, 1d ago · open the first message on its own page

[PATCH 0/2] KVM: PPC: BookE: Fix boot hang on preemptible kernels

From: Shrikanth Hegde <sshegde@linux.ibm.com>
Date: 2026-09-28 11:04:45

Christian Zigotzky had reported that preemptible kernels don't boot on
his FSL Cyrus+ board. The board freezes on running either full/lazy
preemption and has been an issue for a while now.

https://lore.kernel.org/all/b897b0fd-90f2-4215-bcd4-3714e497d773@xenosoft.de/#t

After 7.0, there is only preempt=full/lazy which removed the
preempt=none/voluntary workaround used to mask the issue. There was
challenge is getting the console logs which made the fixes difficult since
one couldn't know where the issue is.

Thanks to Michal for helping in getting the logs. That pointed at few
places where the issue could be. This series is an attempt at fixing
those.

Christian, KVM team, 
Please *test* the patches.

Shrikanth Hegde (2):
  KVM: PPC: BookE: Disable preemption before loading guest FP and
    Altivec
  KVM: PPC: Replay pending interrupts before entering the guest

 arch/powerpc/kvm/booke.c   |  9 ++++++++-
 arch/powerpc/kvm/powerpc.c | 13 +++++++++++++
 2 files changed, 21 insertions(+), 1 deletion(-)

-- 
2.52.0

[PATCH 1/2] KVM: PPC: BookE: Disable preemption before loading guest FP and Altivec

From: Shrikanth Hegde <sshegde@linux.ibm.com>
Date: 2026-09-28 11:04:51

Christian reported that booting preemptible kernel on FSL Cyrus+ board
causes boot hang.

The logs pointed that system was busy in printing below warning.

WARNING: at .enable_kernel_fp+0x30/0x78, CPU#3: qemu-system-ppc/4884
Modules linked in:
CPU: 3 UID: 1000 PID: 4884 Comm: qemu-system-ppc Not tainted 7.3.0-rc1-powerpc64-smp-preempt #1 PREEMPT
NIP [c000000000003338] .enable_kernel_fp+0x30/0x78
LR [c00000000005de84] .kvmppc_load_guest_fp+0x30/0x80
Call Trace:
[c000000085ca7700] [c00000000005de84] .kvmppc_load_guest_fp+0x30/0x80
[c000000085ca7780] [c00000000005f2a0] .kvmppc_handle_exit+0x5bc/0x5cc
[c000000085ca7830] [c00000000006204c] .kvmppc_resume_host+0xb8/0x10c

Which is...

void enable_kernel_fp(void)
{
        unsigned long cpumsr;
        WARN_ON(preemptible());

And...

Though irq's are hard disabled after kvmppc_prepare_to_enter, but
kvmppc_fix_ee_before_entry enables the softmask's IRQ state.
That causes the irqs_disabled to return false.
Hence leading to the warnings.

Fix it by disabling the preemption using the preempt disable.
Note, it is calling noresched variant of preempt enable, since hard
irq are disabled. It is likely not a good idea to call schedule.

Fixes: 3efc7da61f6c ("KVM: PPC: Book3E: Increase FPU laziness")
Reported-by: Christian Zigotzky <redacted>
Closes: https://lore.kernel.org/all/33342fbf-eb7b-bde6-2c8c-254fe8bfb993@xenosoft.de/
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
 arch/powerpc/kvm/booke.c | 9 ++++++++-
 1 file changed, 8 insertions(+), 1 deletion(-)
diff --git a/arch/powerpc/kvm/booke.c b/arch/powerpc/kvm/booke.c
index 13ad4cf5fa71..5b9118eefe1d 100644
--- a/arch/powerpc/kvm/booke.c
+++ b/arch/powerpc/kvm/booke.c
@@ -1404,10 +1404,17 @@ int kvmppc_handle_exit(struct kvm_vcpu *vcpu, unsigned int exit_nr)
 		if (s <= 0)
 			r = (s << 2) | RESUME_HOST | (r & RESUME_FLAG_NV);
 		else {
-			/* interrupts now hard-disabled */
+			/*
+			 * kvmppc_fix_ee_before_entry() marks the software
+			 * IRQ state enabled while interrupts are still
+			 * hard-disabled. So disable preemption while loading
+			 * guest FP and Altivec.
+			 */
 			kvmppc_fix_ee_before_entry();
+			preempt_disable();
 			kvmppc_load_guest_fp(vcpu);
 			kvmppc_load_guest_altivec(vcpu);
+			preempt_enable_no_resched();
 		}
 	}
 
-- 
2.52.0

[PATCH 2/2] KVM: PPC: Replay pending interrupts before entering the guest

From: Shrikanth Hegde <sshegde@linux.ibm.com>
Date: 2026-09-28 11:04:59

After applying preempt disable patch, i.e PATCH 1/2, Christian reported
a subsequent warning stopping his board to boot properly.

WARNING: at .kvmppc_fix_ee_before_entry+0x10/0x28, CPU#0: qemu-system-ppc/4667
CPU: 0 UID: 1000 PID: 4667 Comm: qemu-system-ppc Tainted: G        W           7.3.0-rc4-2-powerpc64-smp #1 PREEMPT
Tainted: [W]=WARN
Hardware name: varisys,CYRUS5040 e5500 0x80240012 CoreNet Generic
NIP [c00000000005dc3c] .kvmppc_fix_ee_before_entry+0x10/0x28
LR [c00000000005f298] .kvmppc_handle_exit+0x5b4/0x5e8
Call Trace:
[c000000086587780] [c00000000005ee10] .kvmppc_handle_exit+0x12c/0x5e8 (unreliable)
[c000000086587830] [c000000000062068] .kvmppc_resume_host+0xb8/0x10c

It triggers below warning...

static inline void kvmppc_fix_ee_before_entry(void)
{
        trace_hardirqs_on();

        /*
         * To avoid races, the caller must have gone directly from having
         * interrupts fully-enabled to hard-disabled.
         */
        WARN_ON(local_paca->irq_happened != PACA_IRQ_HARD_DIS);

This happens since kvmppc_prepare_to_enter does first local_irq_disable
followed by hard_irq_disable. This leaves a small window where interrupt
may occur and it could set the irq pending bit in PACA. When that
happens replay that interrupt before entering the guest.

Fixes: 12013e3d4695 ("KVM: powerpc: Use generic xfer to guest work function")
Reported-by: Christian Zigotzky <redacted>
Closes: https://lore.kernel.org/all/b8f82519-9ae5-c247-e020-e79c627e57df@xenosoft.de/
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
 arch/powerpc/kvm/powerpc.c | 13 +++++++++++++
 1 file changed, 13 insertions(+)
diff --git a/arch/powerpc/kvm/powerpc.c b/arch/powerpc/kvm/powerpc.c
index 9194cf492d1c..847ff07c364b 100644
--- a/arch/powerpc/kvm/powerpc.c
+++ b/arch/powerpc/kvm/powerpc.c
@@ -147,6 +147,19 @@ int kvmppc_prepare_to_enter(struct kvm_vcpu *vcpu)
 			continue;
 		}
 
+#ifdef CONFIG_PPC64
+		/*
+		 * Interrupt arrived between the soft and hard
+		 * disable. Replay it and retry guest entry.
+		 */
+		if (unlikely(local_paca->irq_happened != PACA_IRQ_HARD_DIS)) {
+			local_irq_enable();
+			local_irq_disable();
+			hard_irq_disable();
+			continue;
+		}
+#endif
+
 		guest_enter_irqoff();
 		return 1;
 	}
-- 
2.52.0

Re: [PATCH 0/2] KVM: PPC: BookE: Fix boot hang on preemptible kernels

From: Christian Zigotzky <hidden>
Date: 2026-09-29 15:51:40

On 28/09/26 13:04, Shrikanth Hegde wrote:
Christian Zigotzky had reported that preemptible kernels don't boot on
his FSL Cyrus+ board. The board freezes on running either full/lazy
preemption and has been an issue for a while now.

https://lore.kernel.org/all/b897b0fd-90f2-4215-bcd4-3714e497d773@xenosoft.de/#t

After 7.0, there is only preempt=full/lazy which removed the
preempt=none/voluntary workaround used to mask the issue. There was
challenge is getting the console logs which made the fixes difficult since
one couldn't know where the issue is.

Thanks to Michal for helping in getting the logs. That pointed at few
places where the issue could be. This series is an attempt at fixing
those.

Christian, KVM team,
Please *test* the patches.

Shrikanth Hegde (2):
   KVM: PPC: BookE: Disable preemption before loading guest FP and
     Altivec
   KVM: PPC: Replay pending interrupts before entering the guest

  arch/powerpc/kvm/booke.c   |  9 ++++++++-
  arch/powerpc/kvm/powerpc.c | 13 +++++++++++++
  2 files changed, 21 insertions(+), 1 deletion(-)
Hi,

I was able to compile the RC5 of kernel 7.3 with your patches today. :-) [1]

After that I successfully tested it with KVM PR and KVM HV on my PowerPC machines.

Tested-by: Christian Zigotzky <redacted>

Many thanks for your help,

Christian


[1] https://github.com/chzigotzky/kernels/releases/tag/v7.3.0-rc5-2


-- 
Sent with BrassMonkey 34.3.2.1 (https://github.com/chzigotzky/Web-Browsers-and-Suites-for-Linux-PPC/releases/tag/BrassMonkey_34.3.2.1)

Re: [PATCH 1/2] KVM: PPC: BookE: Disable preemption before loading guest FP and Altivec

From: Narayana Murty N <hidden>
Date: 2026-10-01 11:27:13

Hi Shrikanth,
Thanks for the fixes. I had one question on patch 1.

On 28/09/26 4:34 PM, Shrikanth Hegde wrote:
quoted hunk
Christian reported that booting preemptible kernel on FSL Cyrus+ board
causes boot hang.

The logs pointed that system was busy in printing below warning.

WARNING: at .enable_kernel_fp+0x30/0x78, CPU#3: qemu-system-ppc/4884
Modules linked in:
CPU: 3 UID: 1000 PID: 4884 Comm: qemu-system-ppc Not tainted 7.3.0-rc1-powerpc64-smp-preempt #1 PREEMPT
NIP [c000000000003338] .enable_kernel_fp+0x30/0x78
LR [c00000000005de84] .kvmppc_load_guest_fp+0x30/0x80
Call Trace:
[c000000085ca7700] [c00000000005de84] .kvmppc_load_guest_fp+0x30/0x80
[c000000085ca7780] [c00000000005f2a0] .kvmppc_handle_exit+0x5bc/0x5cc
[c000000085ca7830] [c00000000006204c] .kvmppc_resume_host+0xb8/0x10c

Which is...

void enable_kernel_fp(void)
{
         unsigned long cpumsr;
         WARN_ON(preemptible());

And...

Though irq's are hard disabled after kvmppc_prepare_to_enter, but
kvmppc_fix_ee_before_entry enables the softmask's IRQ state.
That causes the irqs_disabled to return false.
Hence leading to the warnings.

Fix it by disabling the preemption using the preempt disable.
Note, it is calling noresched variant of preempt enable, since hard
irq are disabled. It is likely not a good idea to call schedule.

Fixes: 3efc7da61f6c ("KVM: PPC: Book3E: Increase FPU laziness")
Reported-by: Christian Zigotzky <redacted>
Closes: https://lore.kernel.org/all/33342fbf-eb7b-bde6-2c8c-254fe8bfb993@xenosoft.de/
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
  arch/powerpc/kvm/booke.c | 9 ++++++++-
  1 file changed, 8 insertions(+), 1 deletion(-)
diff --git a/arch/powerpc/kvm/booke.c b/arch/powerpc/kvm/booke.c
index 13ad4cf5fa71..5b9118eefe1d 100644
--- a/arch/powerpc/kvm/booke.c
+++ b/arch/powerpc/kvm/booke.c
@@ -1404,10 +1404,17 @@ int kvmppc_handle_exit(struct kvm_vcpu *vcpu, unsigned int exit_nr)
  		if (s <= 0)
  			r = (s << 2) | RESUME_HOST | (r & RESUME_FLAG_NV);
  		else {
-			/* interrupts now hard-disabled */
+			/*
+			 * kvmppc_fix_ee_before_entry() marks the software
+			 * IRQ state enabled while interrupts are still
+			 * hard-disabled. So disable preemption while loading
+			 * guest FP and Altivec.
+			 */
  			kvmppc_fix_ee_before_entry();
+			preempt_disable();
  			kvmppc_load_guest_fp(vcpu);
  			kvmppc_load_guest_altivec(vcpu);
+			preempt_enable_no_resched();
Would it be simpler to move kvmppc_fix_ee_before_entry() after the
FP/Altivec loads instead?

The normal kvmppc_vcpu_run() entry path already loads the guest
FP/Altivec state while interrupts are still disabled and calls
kvmppc_fix_ee_before_entry() immediately before entering the guest.

So could this path follow the same ordering:

kvmppc_load_guest_fp(vcpu);
kvmppc_load_guest_altivec(vcpu);
kvmppc_fix_ee_before_entry();

That would avoid making the software IRQ state enabled before loading
the guest FP/Altivec state, and also avoid the additional
preempt_disable()/preempt_enable_no_resched() pair.

Thanks,
Narayana Murty.
  		}
  	}
  

Re: [PATCH 2/2] KVM: PPC: Replay pending interrupts before entering the guest

From: Narayana Murty N <hidden>
Date: 2026-10-01 11:29:20

Hi Shrikanth,

On 28/09/26 4:34 PM, Shrikanth Hegde wrote:
quoted hunk
After applying preempt disable patch, i.e PATCH 1/2, Christian reported
a subsequent warning stopping his board to boot properly.

WARNING: at .kvmppc_fix_ee_before_entry+0x10/0x28, CPU#0: qemu-system-ppc/4667
CPU: 0 UID: 1000 PID: 4667 Comm: qemu-system-ppc Tainted: G        W           7.3.0-rc4-2-powerpc64-smp #1 PREEMPT
Tainted: [W]=WARN
Hardware name: varisys,CYRUS5040 e5500 0x80240012 CoreNet Generic
NIP [c00000000005dc3c] .kvmppc_fix_ee_before_entry+0x10/0x28
LR [c00000000005f298] .kvmppc_handle_exit+0x5b4/0x5e8
Call Trace:
[c000000086587780] [c00000000005ee10] .kvmppc_handle_exit+0x12c/0x5e8 (unreliable)
[c000000086587830] [c000000000062068] .kvmppc_resume_host+0xb8/0x10c

It triggers below warning...

static inline void kvmppc_fix_ee_before_entry(void)
{
         trace_hardirqs_on();

         /*
          * To avoid races, the caller must have gone directly from having
          * interrupts fully-enabled to hard-disabled.
          */
         WARN_ON(local_paca->irq_happened != PACA_IRQ_HARD_DIS);

This happens since kvmppc_prepare_to_enter does first local_irq_disable
followed by hard_irq_disable. This leaves a small window where interrupt
may occur and it could set the irq pending bit in PACA. When that
happens replay that interrupt before entering the guest.

Fixes: 12013e3d4695 ("KVM: powerpc: Use generic xfer to guest work function")
Reported-by: Christian Zigotzky <redacted>
Closes: https://lore.kernel.org/all/b8f82519-9ae5-c247-e020-e79c627e57df@xenosoft.de/
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
  arch/powerpc/kvm/powerpc.c | 13 +++++++++++++
  1 file changed, 13 insertions(+)
diff --git a/arch/powerpc/kvm/powerpc.c b/arch/powerpc/kvm/powerpc.c
index 9194cf492d1c..847ff07c364b 100644
--- a/arch/powerpc/kvm/powerpc.c
+++ b/arch/powerpc/kvm/powerpc.c
@@ -147,6 +147,19 @@ int kvmppc_prepare_to_enter(struct kvm_vcpu *vcpu)
  			continue;
  		}
  
+#ifdef CONFIG_PPC64
+		/*
+		 * Interrupt arrived between the soft and hard
+		 * disable. Replay it and retry guest entry.
+		 */
+		if (unlikely(local_paca->irq_happened != PACA_IRQ_HARD_DIS)) {
+			local_irq_enable();
+			local_irq_disable();
+			hard_irq_disable();
+			continue;
+		}
+#endif
+
  		guest_enter_irqoff();
  		return 1;
  	}
It look good to me.

Thanks,
Narayana

Re: [PATCH 1/2] KVM: PPC: BookE: Disable preemption before loading guest FP and Altivec

From: Shrikanth Hegde <sshegde@linux.ibm.com>
Date: 2026-10-01 11:54:13

Hi Narayana.

On 10/1/26 4:56 PM, Narayana Murty N wrote:
Hi Shrikanth,
Thanks for the fixes. I had one question on patch 1.
See response below.
quoted
diff --git a/arch/powerpc/kvm/booke.c b/arch/powerpc/kvm/booke.c
index 13ad4cf5fa71..5b9118eefe1d 100644
--- a/arch/powerpc/kvm/booke.c
+++ b/arch/powerpc/kvm/booke.c
@@ -1404,10 +1404,17 @@ int kvmppc_handle_exit(struct kvm_vcpu *vcpu, unsigned int exit_nr)
          if (s <= 0)
              r = (s << 2) | RESUME_HOST | (r & RESUME_FLAG_NV);
          else {
-            /* interrupts now hard-disabled */
+            /*
+             * kvmppc_fix_ee_before_entry() marks the software
+             * IRQ state enabled while interrupts are still
+             * hard-disabled. So disable preemption while loading
+             * guest FP and Altivec.
+             */
              kvmppc_fix_ee_before_entry();
+            preempt_disable();
              kvmppc_load_guest_fp(vcpu);
              kvmppc_load_guest_altivec(vcpu);
+            preempt_enable_no_resched();
Would it be simpler to move kvmppc_fix_ee_before_entry() after the
FP/Altivec loads instead?

The normal kvmppc_vcpu_run() entry path already loads the guest
FP/Altivec state while interrupts are still disabled and calls
kvmppc_fix_ee_before_entry() immediately before entering the guest.

So could this path follow the same ordering:

kvmppc_load_guest_fp(vcpu);
kvmppc_load_guest_altivec(vcpu);
kvmppc_fix_ee_before_entry();

That would avoid making the software IRQ state enabled before loading
the guest FP/Altivec state, and also avoid the additional
preempt_disable()/preempt_enable_no_resched() pair.

Thanks,
Narayana Murty.
quoted
          }
      }
I thought I had put that for discussion after ---, but looks like I forgot.

I don't mind the above too.  but I didn't have a way to test it.
So kept it as is based on what Christian said works for him.

If you have a way to test the patches, please let me know.
We can try that too.

Re: [PATCH 2/2] KVM: PPC: Replay pending interrupts before entering the guest

From: Shrikanth Hegde <sshegde@linux.ibm.com>
Date: 2026-10-01 11:57:26

Hi Narayana,
quoted
diff --git a/arch/powerpc/kvm/powerpc.c b/arch/powerpc/kvm/powerpc.c
index 9194cf492d1c..847ff07c364b 100644
--- a/arch/powerpc/kvm/powerpc.c
+++ b/arch/powerpc/kvm/powerpc.c
@@ -147,6 +147,19 @@ int kvmppc_prepare_to_enter(struct kvm_vcpu *vcpu)
              continue;
          }
+#ifdef CONFIG_PPC64
+        /*
+         * Interrupt arrived between the soft and hard
+         * disable. Replay it and retry guest entry.
+         */
+        if (unlikely(local_paca->irq_happened != PACA_IRQ_HARD_DIS)) {
+            local_irq_enable();
+            local_irq_disable();
+            hard_irq_disable();
+            continue;
+        }
+#endif
+
          guest_enter_irqoff();
          return 1;
      }
It look good to me.
Thanks.
Thanks,
Narayana
Do you want me to consider the above as rwb tag? If yes, I prefer to see explicit tag.
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help