From: Amit Machhiwal <hidden> Date: 2026-05-22 15:28:22
On POWER systems, newer processor generations can operate in compatibility
modes corresponding to earlier generations (e.g., a Power11 system running
in Power10 compatibility mode). In such cases, the effective CPU level
exposed to guests differs from the physical processor generation.
This creates a problem for nested virtualization. When booting a nested KVM
guest (L2) inside a host KVM guest (L1) running in a compatibility mode,
userspace (e.g., QEMU) may derive the CPU model from the raw hardware PVR
and attempt to configure the nested guest accordingly. However, the L1
partition is constrained by the compatibility level negotiated with the
hypervisor (L0), and requests exceeding that level are rejected, leading to
guest boot failures such as:
KVM-NESTEDv2: couldn't set guest wide elements
This series addresses the issue in two steps:
1. Detect and reject invalid compatibility requests early in KVM to avoid
late failures.
2. Provide a mechanism for userspace to query the effective CPU
compatibility modes supported by the host, so it can select an
appropriate CPU model for nested guests.
To achieve this, the series introduces a new KVM capability and ioctl
(KVM_CAP_PPC_COMPAT_CAPS / KVM_PPC_GET_COMPAT_CAPS) that expose the
compatibility modes supported by the host.
Why a new UAPI?
---------------
While cpu-version is available in /proc/device-tree/cpus/<cpu#>/cpu-version
on both L1 booted on PowerNV and PowerVM LPARs, the UAPI approach is
preferable for several reasons:
1. pHYP (L0) capabilities: On PowerVM, we need to rely on capabilities
negotiated with pHYP in KVM, not just device tree properties. The
cpu-version property depicts the current compat mode but doesn't point
to what all compat modes are supported for the nested guest.
2. procfs dependency: Not all systems run with procfs enabled (CONFIG_PROC_FS
is optional). Minimal configurations like buildroot might disable it, but
KVM ioctl works regardless since it accesses kernel data structures
directly.
3. Kernel validation: The kernel validates and normalizes the compatibility
information. Patch 1 adds validation logic that rejects invalid
compatibility requests early, ensuring userspace gets validated,
consistent data.
4. Abstraction & stability: /proc/device-tree is an implementation detail.
The UAPI provides a stable interface that won't break if the underlying
mechanism changes.
5. Semantic clarity: KVM_PPC_GET_COMPAT_CAPS clearly expresses what
compatibility modes can be used for KVM guests, vs. parsing device tree
which requires understanding the semantic meaning of cpu-version.
The implementation supports both:
- PowerVM (nested API v2), where compatibility information is obtained
via the H_GUEST_GET_CAPABILITIES hypercall.
- PowerNV (nested API v1), where compatibility is derived from the device
tree ("cpu-version") representing the effective processor compatibility
level.
This allows userspace (e.g., QEMU) to select a CPU model consistent with
the host compatibility mode, avoiding mismatches and enabling successful
nested guest boot.
Changes in v3:
- Added "Why a new UAPI?" section to cover letter addressing questions
about the need for a new UAPI vs. using existing mechanisms like
/proc/device-tree
- Fixed initialization of 'r' in KVM_PPC_GET_COMPAT_CAPS ioctl handler
from 0 to -ENOTTY for proper error handling when the operation is not
supported
- Added Vaibhav's "Suggested-by" tags
- Have retained Anushree's "Tested-by" tags as no major code changes
- Fixed documentation build warning reported by kernel test robot and
added "Reported-by" and "Closes" tags to patch 5
Changes in v2:
- Squashed patches 2 and 3 from v1 (capability introduction and ioctl
wiring) into a single patch for better logical grouping
- Changed kvm_ppc_compat_caps.flags from __u32 to __u64 for consistency
and future extensibility
- Addressed other review comments
- Improved commit messages with clearer explanations of the changes
Patch summary:
[1/5] Validate arch_compat against host compatibility mode
[2/5] Introduce KVM_CAP_PPC_COMPAT_CAPS and wire up ioctl
[3/5] Implement capability retrieval for PowerVM (API v2)
[4/5] Add PowerNV support (API v1)
[5/5] Document the new ioctl
Tested on:
- Power11 pSeries LPAR in Power10 compatibility mode (nested API v2)
- Power10 PowerNV system (and QEMU TCG PowerNV 11) with nested
virtualization (API v1) with various combinations of KVM L1/L2 guests
in various supported compatibility modes.
With this series, nested guests boot successfully in configurations where
they previously failed due to compatibility mismatches.
Related QEMU series:
A corresponding QEMU series adds support for querying and using these
compatibility capabilities when configuring nested KVM guests:
https://lore.kernel.org/all/20260502140021.69712-1-amachhiw@linux.ibm.com/
v2: https://lore.kernel.org/linuxppc-dev/20260513100755.83215-1-amachhiw@linux.ibm.com/
v1: https://lore.kernel.org/linuxppc-dev/20260430054906.94431-1-amachhiw@linux.ibm.com/
Amit Machhiwal (5):
KVM: PPC: Book3S HV: Validate arch_compat against host compatibility
mode
KVM: PPC: Introduce KVM_CAP_PPC_COMPAT_CAPS and wire up ioctl
KVM: PPC: Book3S HV: Implement compat CPU capability retrieval for KVM
on PowerVM
KVM: PPC: Book3S HV: Add support for compat CPU capabilities for KVM
on PowerNV
KVM: PPC: Document KVM_PPC_GET_COMPAT_CAPS ioctl
Documentation/virt/kvm/api.rst | 35 ++++++++++++++++
arch/powerpc/include/asm/kvm_ppc.h | 1 +
arch/powerpc/include/uapi/asm/kvm.h | 6 +++
arch/powerpc/kvm/book3s_hv.c | 63 +++++++++++++++++++++++++++++
arch/powerpc/kvm/powerpc.c | 21 ++++++++++
include/uapi/linux/kvm.h | 4 ++
6 files changed, 130 insertions(+)
base-commit: 1d5dcaa3bd65f2e8c9baa14a393d3a2dc5db7524
--
2.50.1 (Apple Git-155)
From: Amit Machhiwal <hidden> Date: 2026-05-22 15:28:24
On POWER systems, the host CPU may run in a compatibility mode (e.g., a
Power11 processor operating in Power10 compatibility mode). In such
cases, the effective CPU level exposed to guests differs from the
physical processor generation.
When running nested KVM guests, QEMU derives the host CPU type using
mfpvr(), which reflects the physical processor version. This can result
in a mismatch between the CPU model selected by QEMU and the
compatibility mode enforced by the host, leading to guest boot failures.
For example, booting a nested guest on a Power11 LPAR configured in
Power10 compatibility mode fails with:
KVM-NESTEDv2: couldn't set guest wide elements
[..KVM reg dump..]
This occurs because QEMU selects a CPU model corresponding to the
physical processor (via mfpvr()), while the host operates in a lower
compatibility mode. As a result, KVM rejects the requested compatibility
level during guest initialization.
Add support for retrieving host CPU compatibility capabilities for
nested guests on PowerVM (PAPR nested API v2). The hypervisor provides
the effective compatibility levels via the H_GUEST_GET_CAPABILITIES
hcall, which reflects the processor modes negotiated between the Power
hypervisor (L0) and the host partition (L1).
On pseries systems, obtain the capability bitmap using
plpar_guest_get_capabilities() and return it via struct
kvm_ppc_compat_caps. This information is then exposed to userspace
through the KVM_PPC_GET_COMPAT_CAPS ioctl.
Hook the implementation into the Book3S HV kvmppc_ops so that it can be
invoked by the generic KVM ioctl handling code.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
From: Amit Machhiwal <hidden> Date: 2026-05-22 15:28:24
On IBM POWER systems, newer processor generations can operate in
compatibility modes corresponding to earlier generations. This becomes
relevant for nested virtualization, where nested KVM guests may need to
run with a specific processor compatibility level.
Currently, when running a nested KVM guest (L2) inside a Power11 pSeries
logical partition (L1) booted in Power10 compatibility mode, the guest
fails to boot while setting 'arch_compat'. This happens because the CPU
class is derived from the hardware PVR (via mfspr()), which reflects the
physical processor generation (Power11), rather than the effective
compatibility mode (Power10).
As a result, userspace may request a Power11 arch_compat for the L2
guest. However, the L1 partition, running in Power10 compatibility, has
only negotiated support up to Power10 with the Power Hypervisor (L0).
When H_SET_STATE is invoked with a Power11 Logical PVR, the hypervisor
rejects the request, leading to a late guest boot failure:
KVM-NESTEDv2: couldn't set guest wide elements
[..KVM reg dump..]
This situation should be detected earlier. Rejecting unsupported
'arch_compat' values in 'kvmppc_set_arch_compat()' avoids issuing an
invalid H_SET_STATE hcall and provides a clearer failure mode.
Add a check to reject Power11 'arch_compat' requests when the host is
running in Power10 compatibility mode, returning -EINVAL early instead
of deferring the failure to the hypervisor.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 12 ++++++++++++
1 file changed, 12 insertions(+)
From: Amit Machhiwal <hidden> Date: 2026-05-22 15:28:25
Introduce a new capability and ioctl to expose CPU compatibility modes
supported by the host processor for nested guests.
On IBM POWER systems, newer processor generations (N) can operate in
compatibility modes corresponding to earlier generations, like (N-1) and
(N-2). This is particularly relevant for nested virtualization, where
nested KVM guests may need to run with a specific processor compatibility
level.
Introduce KVM_CAP_PPC_COMPAT_CAPS capability and the corresponding
KVM_PPC_GET_COMPAT_CAPS vm ioctl. The ioctl returns a bitmap describing
the compatibility modes supported by the host in respective bit numbers,
allowing userspace (e.g., QEMU) to select an appropriate compatibility
level when configuring nested KVM guests.
The ioctl handling is added in kvm_arch_vm_ioctl() and retrieves host
CPU compatibility capabilities via a PowerPC-specific backend
implementation when available. If the capability is not supported, the
ioctl returns success with no capabilities set, allowing userspace to
fall back gracefully.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/include/asm/kvm_ppc.h | 1 +
arch/powerpc/include/uapi/asm/kvm.h | 6 ++++++
arch/powerpc/kvm/powerpc.c | 21 +++++++++++++++++++++
include/uapi/linux/kvm.h | 4 ++++
4 files changed, 32 insertions(+)
@@ -437,6 +437,12 @@ struct kvm_ppc_cpu_char {__u64behaviour_mask;/* valid bits in behaviour */};+/* For KVM_PPC_GET_COMPAT_CAPS */+structkvm_ppc_compat_caps{+__u64flags;/* Reserved for future use */+__u64compat_capabilities;/* Capabilities supported by the host */+};+/**Valuesforcharacterandcharacter_mask.*TheseareidenticaltothevaluesusedbyH_GET_CPU_CHARACTERISTICS.
From: Amit Machhiwal <hidden> Date: 2026-05-22 15:28:30
Currently, when booting a compatibility-mode KVM guest (L1) on a PowerNV
hypervisor (L0), the guest runs with the expected processor
compatibility level. However, when booting a nested KVM guest (L2)
inside the L1, QEMU derives the CPU model from the raw host PVR and
attempts to run the nested guest at that level, instead of honoring the
compatibility mode of the L1.
Extend host CPU compatibility capability reporting to support nested
virtualization on PowerNV systems (PAPR nested API v1).
For nested API v2 (PowerVM), compatibility capabilities are obtained
from the hypervisor via the H_GUEST_GET_CAPABILITIES hcall. This
information is not available on PowerNV systems.
For nested API v1, derive the compatibility capabilities from the L1
guest by reading the "cpu-version" property from the device tree, which
reflects the effective (logical) processor compatibility level. Map this
value to the corresponding compatibility capability bitmap.
Introduce a helper to translate CPU version values into compatibility
capability bits and integrate it into kvmppc_get_compat_cpu_caps().
This allows userspace to query host CPU compatibility modes on both
PowerVM and PowerNV platforms via the KVM_PPC_GET_COMPAT_CAPS ioctl.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 37 +++++++++++++++++++++++++++++++++++-
1 file changed, 36 insertions(+), 1 deletion(-)
@@ -6555,6 +6555,41 @@ KVM_S390_KEYOP_SSKE.._kvm_run:+4.145 KVM_PPC_GET_COMPAT_CAPS+-----------------------------+:Capability: KVM_CAP_PPC_COMPAT_CAPS+:Architectures: powerpc+:Type: vm ioctl+:Parameters: struct kvm_ppc_compat_caps (out)+:Returns:+ 0 on successful completion,+ -EFAULT if ``struct kvm_ppc_compat_caps`` cannot be written++IBM POWER system server-based processors provide a compatibility mode feature+where an Nth generation processor can operate in modes consistent with earlier+generations such as (N-1) and (N-2).++This ioctl provides userspace with information about the CPU compatibility modes+supported by the current host processor for booting the nested KVM guests on+PowerNV (KVM nested APIv1) and PowerVM (KVM nested APIv2) platforms.++::++ struct kvm_ppc_compat_caps {+ __u64 flags; /* Reserved for future use */+ __u64 compat_capabilities; /* Capabilities supported by the host */+ };++The ``compat_capabilities`` bit field describes the processor compatibility+modes supported by the host. For example, the following bits indicate support+for specific processor modes.++::++ H_GUEST_CAP_POWER9 (bit 1): KVM guests can run in Power9 processor mode+ H_GUEST_CAP_POWER10 (bit 2): KVM guests can run in Power10 processor mode+ H_GUEST_CAP_POWER11 (bit 3): KVM guests can run in Power11 processor mode+5. The kvm_run structure ========================
On IBM POWER systems, newer processor generations can operate in
compatibility modes corresponding to earlier generations. This becomes
relevant for nested virtualization, where nested KVM guests may need to
run with a specific processor compatibility level.
Currently, when running a nested KVM guest (L2) inside a Power11 pSeries
logical partition (L1) booted in Power10 compatibility mode, the guest
fails to boot while setting 'arch_compat'. This happens because the CPU
class is derived from the hardware PVR (via mfspr()), which reflects the
physical processor generation (Power11), rather than the effective
compatibility mode (Power10).
As a result, userspace may request a Power11 arch_compat for the L2
guest. However, the L1 partition, running in Power10 compatibility, has
only negotiated support up to Power10 with the Power Hypervisor (L0).
When H_SET_STATE is invoked with a Power11 Logical PVR, the hypervisor
s/H_SET_STATE/H_GUEST_SET_STATE
rejects the request, leading to a late guest boot failure:
KVM-NESTEDv2: couldn't set guest wide elements
[..KVM reg dump..]
I think irrespective of the other UAPI changes, we should still get this
fixed - so that we don't see a late KVM guest boot failure msgs.
So, in this review, I would like to mainly look at fixing this issue
first and would request if we can defer the UAPI changes as a separate
patch series please.
This situation should be detected earlier. Rejecting unsupported
'arch_compat' values in 'kvmppc_set_arch_compat()' avoids issuing an
invalid H_SET_STATE hcall and provides a clearer failure mode.
s/H_SET_STATE/H_GUEST_SET_STATE
quoted hunk
Add a check to reject Power11 'arch_compat' requests when the host is
running in Power10 compatibility mode, returning -EINVAL early instead
of deferring the failure to the hypervisor.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 12 ++++++++++++
1 file changed, 12 insertions(+)
Instead of the complicated check can we simply do this?
if (!cpu_has_feature(CPU_FTR_P11_PVR))
return -EINVAL;
which means that if the Qemu is trying to set the arch_compat with P11
PVR (arch_compat) and if the host cpu FTR doesn't support P11 PVR, then
simply return -EINVAL
-ritesh
From: Amit Machhiwal <hidden> Date: 2026-05-29 10:28:48
Hi Ritesh,
Thanks for taking a look at this patch. Please find my response inline
below:
On 2026/05/28 08:43 AM, Ritesh Harjani wrote:
Amit Machhiwal [off-list ref] writes:
quoted
On IBM POWER systems, newer processor generations can operate in
compatibility modes corresponding to earlier generations. This becomes
relevant for nested virtualization, where nested KVM guests may need to
run with a specific processor compatibility level.
Currently, when running a nested KVM guest (L2) inside a Power11 pSeries
logical partition (L1) booted in Power10 compatibility mode, the guest
fails to boot while setting 'arch_compat'. This happens because the CPU
class is derived from the hardware PVR (via mfspr()), which reflects the
physical processor generation (Power11), rather than the effective
compatibility mode (Power10).
As a result, userspace may request a Power11 arch_compat for the L2
guest. However, the L1 partition, running in Power10 compatibility, has
only negotiated support up to Power10 with the Power Hypervisor (L0).
When H_SET_STATE is invoked with a Power11 Logical PVR, the hypervisor
s/H_SET_STATE/H_GUEST_SET_STATE
Good catch! I'll fix this in the next version.
quoted
rejects the request, leading to a late guest boot failure:
KVM-NESTEDv2: couldn't set guest wide elements
[..KVM reg dump..]
I think irrespective of the other UAPI changes, we should still get this
fixed - so that we don't see a late KVM guest boot failure msgs.
So, in this review, I would like to mainly look at fixing this issue
first and would request if we can defer the UAPI changes as a separate
patch series please.
This patch only enables the guest boot to bail out early instead of going upto
making an H_GUEST_SET_STATE hcall with a non supported compatibility mode. In
addition to that, it does not fix the real problem where guest fails to boot on
Power11 LPAR booted in a Power10 compatibility mode.
I understand your point and to make the segregation explicitly clear, I shall
update the cover letter to mention that patch 1 only takes care of failing
earlier as soon as an invalid compatibility mode is detected and Patch 2-5
introduce a new uAPI for evaluating the right compatibility mode.
The actual L2 guest boot fix with L1 booted in a compatibility mode is done via
the next 4 patches which enables correct compatibility mode detection using the
newly introduced uAPI. So, we would still want to prioritize the whole series
instead of just this one patch.
quoted
This situation should be detected earlier. Rejecting unsupported
'arch_compat' values in 'kvmppc_set_arch_compat()' avoids issuing an
invalid H_SET_STATE hcall and provides a clearer failure mode.
s/H_SET_STATE/H_GUEST_SET_STATE
Will rectify in the next version.
quoted
Add a check to reject Power11 'arch_compat' requests when the host is
running in Power10 compatibility mode, returning -EINVAL early instead
of deferring the failure to the hypervisor.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 12 ++++++++++++
1 file changed, 12 insertions(+)
Instead of the complicated check can we simply do this?
if (!cpu_has_feature(CPU_FTR_P11_PVR))
return -EINVAL;
Sure, I can base this check on CPU features.
Thanks,
Amit
which means that if the Qemu is trying to set the arch_compat with P11
PVR (arch_compat) and if the host cpu FTR doesn't support P11 PVR, then
simply return -EINVAL
-ritesh
So, we would still want to prioritize the whole series
instead of just this one patch.
Patch-1 could go as a bug fix even in 7.1-rc6 (or maybe with 7.2
bug fixes). - Maddy?
So, you may want to add a fixes tag and maybe even cc stable if you are
seeing this issue from older kernels maybe when nestedv2 got introduced?
However the new UAPI discussion might still require more discussion with
the community and I don't think it is ready for 7.2 yet ;)
-ritesh
Hi Ritesh, thanks for looking into this patch. My responses to your
review comments inline below.
Ritesh Harjani (IBM) [off-list ref] writes:
Amit Machhiwal [off-list ref] writes:
quoted
So, we would still want to prioritize the whole series
instead of just this one patch.
Patch-1 could go as a bug fix even in 7.1-rc6 (or maybe with 7.2
bug fixes). - Maddy?
So, you may want to add a fixes tag and maybe even cc stable if you are
seeing this issue from older kernels maybe when nestedv2 got introduced?
This isnt a 'bug fix' per-se but rather strengthening of compat mode
checks so that any non compatible PVR being used by the VMM can be
caught early. The hypervisor anyway ultimately prevents non-compatible
PVRs from being used by the VMM. So there isnt a bug thats being fixed
in this patch.
The rest of the patch series builds on top of this patch to advertise
the available compatible PVRs to the VMM so that it can further
preemptively prevent users from forcibly using a non-compatible PVR.
Hence IMHO, this patch can be marked for stable tree and potential
candidate for 7.2 merge window. But dont see applicability of a 'fixes'
tag to this patch
However the new UAPI discussion might still require more discussion with
the community and I don't think it is ready for 7.2 yet ;)
Hi Amit,
Thanks for this patch. Few review comments below:
Amit Machhiwal [off-list ref] writes:
quoted hunk
Introduce a new capability and ioctl to expose CPU compatibility modes
supported by the host processor for nested guests.
On IBM POWER systems, newer processor generations (N) can operate in
compatibility modes corresponding to earlier generations, like (N-1) and
(N-2). This is particularly relevant for nested virtualization, where
nested KVM guests may need to run with a specific processor compatibility
level.
Introduce KVM_CAP_PPC_COMPAT_CAPS capability and the corresponding
KVM_PPC_GET_COMPAT_CAPS vm ioctl. The ioctl returns a bitmap describing
the compatibility modes supported by the host in respective bit numbers,
allowing userspace (e.g., QEMU) to select an appropriate compatibility
level when configuring nested KVM guests.
The ioctl handling is added in kvm_arch_vm_ioctl() and retrieves host
CPU compatibility capabilities via a PowerPC-specific backend
implementation when available. If the capability is not supported, the
ioctl returns success with no capabilities set, allowing userspace to
fall back gracefully.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/include/asm/kvm_ppc.h | 1 +
arch/powerpc/include/uapi/asm/kvm.h | 6 ++++++
arch/powerpc/kvm/powerpc.c | 21 +++++++++++++++++++++
include/uapi/linux/kvm.h | 4 ++++
4 files changed, 32 insertions(+)
@@ -437,6 +437,12 @@ struct kvm_ppc_cpu_char {__u64behaviour_mask;/* valid bits in behaviour */};+/* For KVM_PPC_GET_COMPAT_CAPS */+structkvm_ppc_compat_caps{+__u64flags;/* Reserved for future use */
Please introduce a size field also for the UAPI so that in this
structure can evolve in future without breaking kernel ABI.
quoted hunk
+ __u64 compat_capabilities; /* Capabilities supported by the host */+};+ /* * Values for character and character_mask. * These are identical to the values used by H_GET_CPU_CHARACTERISTICS.
As mentioned above please introduce a size field in the structure thats
being copied to the userspace and use the size field to copy the
apporiate structure to the userspace. Otherwise a future kernel may
unintentionally overwrite unintended userspace memory if it happens to
be using a larger structure size then what VMM knows about.
quoted hunk
+ r = -EFAULT;+ break;+ } default: { struct kvm *kvm = filp->private_data; r = kvm->arch.kvm_ops->arch_vm_ioctl(filp, ioctl, arg);
Hi Amit,
Thanks for the patch. My review comments inline below:
Amit Machhiwal [off-list ref] writes:
quoted hunk
On POWER systems, the host CPU may run in a compatibility mode (e.g., a
Power11 processor operating in Power10 compatibility mode). In such
cases, the effective CPU level exposed to guests differs from the
physical processor generation.
When running nested KVM guests, QEMU derives the host CPU type using
mfpvr(), which reflects the physical processor version. This can result
in a mismatch between the CPU model selected by QEMU and the
compatibility mode enforced by the host, leading to guest boot failures.
For example, booting a nested guest on a Power11 LPAR configured in
Power10 compatibility mode fails with:
KVM-NESTEDv2: couldn't set guest wide elements
[..KVM reg dump..]
This occurs because QEMU selects a CPU model corresponding to the
physical processor (via mfpvr()), while the host operates in a lower
compatibility mode. As a result, KVM rejects the requested compatibility
level during guest initialization.
Add support for retrieving host CPU compatibility capabilities for
nested guests on PowerVM (PAPR nested API v2). The hypervisor provides
the effective compatibility levels via the H_GUEST_GET_CAPABILITIES
hcall, which reflects the processor modes negotiated between the Power
hypervisor (L0) and the host partition (L1).
On pseries systems, obtain the capability bitmap using
plpar_guest_get_capabilities() and return it via struct
kvm_ppc_compat_caps. This information is then exposed to userspace
through the KVM_PPC_GET_COMPAT_CAPS ioctl.
Hook the implementation into the Book3S HV kvmppc_ops so that it can be
invoked by the generic KVM ioctl handling code.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
since this value will trikle back to userspace please apply a mask on
the hcall return value so that any reserved and non-PVR related bits
doesnt leak back to userspace.
Hi Amit,
Thanks for the patch. My review comments inline:
Amit Machhiwal [off-list ref] writes:
quoted hunk
Currently, when booting a compatibility-mode KVM guest (L1) on a PowerNV
hypervisor (L0), the guest runs with the expected processor
compatibility level. However, when booting a nested KVM guest (L2)
inside the L1, QEMU derives the CPU model from the raw host PVR and
attempts to run the nested guest at that level, instead of honoring the
compatibility mode of the L1.
Extend host CPU compatibility capability reporting to support nested
virtualization on PowerNV systems (PAPR nested API v1).
For nested API v2 (PowerVM), compatibility capabilities are obtained
from the hypervisor via the H_GUEST_GET_CAPABILITIES hcall. This
information is not available on PowerNV systems.
For nested API v1, derive the compatibility capabilities from the L1
guest by reading the "cpu-version" property from the device tree, which
reflects the effective (logical) processor compatibility level. Map this
value to the corresponding compatibility capability bitmap.
Introduce a helper to translate CPU version values into compatibility
capability bits and integrate it into kvmppc_get_compat_cpu_caps().
This allows userspace to query host CPU compatibility modes on both
PowerVM and PowerNV platforms via the KVM_PPC_GET_COMPAT_CAPS ioctl.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 37 +++++++++++++++++++++++++++++++++++-
1 file changed, 36 insertions(+), 1 deletion(-)
Need to mask capabilities as mentioned in the review comments for
previous patch. I would suggest creating a helper that performs the
hcall and applies the mask which can then be used at
plpar_guest_get_capabilities() call sites.
Hi Ritesh, thanks for looking into this patch. My responses to your
review comments inline below.
Ritesh Harjani (IBM) [off-list ref] writes:
quoted
Amit Machhiwal [off-list ref] writes:
quoted
So, we would still want to prioritize the whole series
instead of just this one patch.
Patch-1 could go as a bug fix even in 7.1-rc6 (or maybe with 7.2
bug fixes). - Maddy?
So, you may want to add a fixes tag and maybe even cc stable if you are
seeing this issue from older kernels maybe when nestedv2 got introduced?
This isnt a 'bug fix' per-se but rather strengthening of compat mode
checks so that any non compatible PVR being used by the VMM can be
caught early. The hypervisor anyway ultimately prevents non-compatible
PVRs from being used by the VMM. So there isnt a bug thats being fixed
in this patch.
The rest of the patch series builds on top of this patch to advertise
the available compatible PVRs to the VMM so that it can further
preemptively prevent users from forcibly using a non-compatible PVR.
Hence IMHO, this patch can be marked for stable tree and potential
candidate for 7.2 merge window. But dont see applicability of a 'fixes'
tag to this patch
amit, can you just post this alone as a separate patch, so that we could
pull it for 7.2 merge?
quoted
However the new UAPI discussion might still require more discussion with
the community and I don't think it is ready for 7.2 yet ;)
From: Harsh Prateek Bora <hidden> Date: 2026-06-03 05:11:08
On 03/06/26 10:03 am, Madhavan Srinivasan wrote:
On 6/3/26 9:03 AM, Vaibhav Jain wrote:
quoted
Hi Ritesh, thanks for looking into this patch. My responses to your
review comments inline below.
Ritesh Harjani (IBM) [off-list ref] writes:
quoted
Amit Machhiwal [off-list ref] writes:
quoted
So, we would still want to prioritize the whole series
instead of just this one patch.
Patch-1 could go as a bug fix even in 7.1-rc6 (or maybe with 7.2
bug fixes). - Maddy?
So, you may want to add a fixes tag and maybe even cc stable if you are
seeing this issue from older kernels maybe when nestedv2 got introduced?
This isnt a 'bug fix' per-se but rather strengthening of compat mode
checks so that any non compatible PVR being used by the VMM can be
caught early. The hypervisor anyway ultimately prevents non-compatible
PVRs from being used by the VMM. So there isnt a bug thats being fixed
in this patch.
The rest of the patch series builds on top of this patch to advertise
the available compatible PVRs to the VMM so that it can further
preemptively prevent users from forcibly using a non-compatible PVR.
Hence IMHO, this patch can be marked for stable tree and potential
candidate for 7.2 merge window. But dont see applicability of a 'fixes'
tag to this patch
amit, can you just post this alone as a separate patch, so that we could
pull it for 7.2 merge?
FWIW, b4 am -P1 <mbox> should fetch this patch alone (and not the entire series), See b4 am --help for more options to select a subset of patches.
regards,
Harsh>
quoted
quoted
However the new UAPI discussion might still require more discussion with
the community and I don't think it is ready for 7.2 yet ;)
amit, can you just post this alone as a separate patch, so that we could
pull it for 7.2 merge?
FWIW, b4 am -P1 <mbox> should fetch this patch alone (and not the entire
series), See b4 am --help for more options to select a subset of patches.
I agree, however as an FYI in this case -
I had few review comments on PATCH-1 here [1] - which along with the
commit msg changes, also had a code change involved, so IMO, it's still
a good idea if Amit can test and send an updated patch separately for this -
to be pulled in for 7.2.
[1]: https://lore.kernel.org/linuxppc-dev/pl2g6xbz.ritesh.list@gmail.com/
Replying to Vaibhav comment here so that we can reach to the conclusion
at one place.
Hence IMHO, this patch can be marked for stable tree and potential
candidate for 7.2 merge window. But dont see applicability of a 'fixes'
tag to this patch
I agree, we need not use a fixes tag then. So, we shall mark this
with v6.10 tag then.
Cc: stable@vger.kernel.org # v6.10+
(I calculated this based on when Power11 was added:
git tag --contains c2ed087ed35ca | grep -E "^v" |head -1
v6.10
)
-ritesh
From: Harsh Prateek Bora <hidden> Date: 2026-06-03 06:31:43
On 03/06/26 11:35 am, Ritesh Harjani (IBM) wrote:
Harsh Prateek Bora [off-list ref] writes:
quoted
quoted
amit, can you just post this alone as a separate patch, so that we could
pull it for 7.2 merge?
FWIW, b4 am -P1 <mbox> should fetch this patch alone (and not the entire
series), See b4 am --help for more options to select a subset of patches.
I agree, however as an FYI in this case -
I had few review comments on PATCH-1 here [1] - which along with the
commit msg changes, also had a code change involved, so IMO, it's still
a good idea if Amit can test and send an updated patch separately for this -
to be pulled in for 7.2.
[1]: https://lore.kernel.org/linuxppc-dev/pl2g6xbz.ritesh.list@gmail.com/
Replying to Vaibhav comment here so that we can reach to the conclusion
at one place.
quoted
Hence IMHO, this patch can be marked for stable tree and potential
candidate for 7.2 merge window. But dont see applicability of a 'fixes'
tag to this patch
I agree, we need not use a fixes tag then. So, we shall mark this
with v6.10 tag then.
Cc: stable@vger.kernel.org # v6.10+
(I calculated this based on when Power11 was added:
git tag --contains c2ed087ed35ca | grep -E "^v" |head -1
v6.10
)
Hence IMHO, this patch can be marked for stable tree and potential
candidate for 7.2 merge window. But dont see applicability of a 'fixes'
tag to this patch
I agree, we need not use a fixes tag then. So, we shall mark this
with v6.10 tag then.
Cc: stable@vger.kernel.org # v6.10+
Please note that I have marked the patch for stable v6.13+ as the KVM
support for Power11 was added via 96e266e3bcd6 ("KVM: PPC: Book3S HV:
Add Power11 capability support for Nested PAPR guests"). Also, this
commit had introduced CPU_FTR_P11_PVR on which the compat PVR check in
the patch is based on.
Thanks,
Amit
From: Amit Machhiwal <hidden> Date: 2026-06-10 15:47:43
Hi Vaibhav,
Thanks for reviewing the patches. Please find my response inline.
On 2026/06/03 09:16 AM, Vaibhav Jain wrote:
Hi Amit,
Thanks for this patch. Few review comments below:
Amit Machhiwal [off-list ref] writes:
quoted
Introduce a new capability and ioctl to expose CPU compatibility modes
supported by the host processor for nested guests.
On IBM POWER systems, newer processor generations (N) can operate in
compatibility modes corresponding to earlier generations, like (N-1) and
(N-2). This is particularly relevant for nested virtualization, where
nested KVM guests may need to run with a specific processor compatibility
level.
Introduce KVM_CAP_PPC_COMPAT_CAPS capability and the corresponding
KVM_PPC_GET_COMPAT_CAPS vm ioctl. The ioctl returns a bitmap describing
the compatibility modes supported by the host in respective bit numbers,
allowing userspace (e.g., QEMU) to select an appropriate compatibility
level when configuring nested KVM guests.
The ioctl handling is added in kvm_arch_vm_ioctl() and retrieves host
CPU compatibility capabilities via a PowerPC-specific backend
implementation when available. If the capability is not supported, the
ioctl returns success with no capabilities set, allowing userspace to
fall back gracefully.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/include/asm/kvm_ppc.h | 1 +
arch/powerpc/include/uapi/asm/kvm.h | 6 ++++++
arch/powerpc/kvm/powerpc.c | 21 +++++++++++++++++++++
include/uapi/linux/kvm.h | 4 ++++
4 files changed, 32 insertions(+)
@@ -437,6 +437,12 @@ struct kvm_ppc_cpu_char {__u64behaviour_mask;/* valid bits in behaviour */};+/* For KVM_PPC_GET_COMPAT_CAPS */+structkvm_ppc_compat_caps{+__u64flags;/* Reserved for future use */
Please introduce a size field also for the UAPI so that in this
structure can evolve in future without breaking kernel ABI.
Sure, adding a size field for validations and future extensibility makes
sense. I'll add it in the next version.
quoted
+ __u64 compat_capabilities; /* Capabilities supported by the host */+};+ /* * Values for character and character_mask. * These are identical to the values used by H_GET_CPU_CHARACTERISTICS.
As mentioned above please introduce a size field in the structure thats
being copied to the userspace and use the size field to copy the
apporiate structure to the userspace. Otherwise a future kernel may
unintentionally overwrite unintended userspace memory if it happens to
be using a larger structure size then what VMM knows about.
Sure, I'll add a validation around the structure size.
quoted
+ r = -EFAULT;+ break;+ } default: { struct kvm *kvm = filp->private_data; r = kvm->arch.kvm_ops->arch_vm_ioctl(filp, ioctl, arg);
From: Amit Machhiwal <hidden> Date: 2026-06-10 15:52:11
Hi Vaibhav,
Thanks for taking a look at this patch. My response is inline.
On 2026/06/03 09:31 AM, Vaibhav Jain wrote:
Hi Amit,
Thanks for the patch. My review comments inline below:
Amit Machhiwal [off-list ref] writes:
quoted
On POWER systems, the host CPU may run in a compatibility mode (e.g., a
Power11 processor operating in Power10 compatibility mode). In such
cases, the effective CPU level exposed to guests differs from the
physical processor generation.
When running nested KVM guests, QEMU derives the host CPU type using
mfpvr(), which reflects the physical processor version. This can result
in a mismatch between the CPU model selected by QEMU and the
compatibility mode enforced by the host, leading to guest boot failures.
For example, booting a nested guest on a Power11 LPAR configured in
Power10 compatibility mode fails with:
KVM-NESTEDv2: couldn't set guest wide elements
[..KVM reg dump..]
This occurs because QEMU selects a CPU model corresponding to the
physical processor (via mfpvr()), while the host operates in a lower
compatibility mode. As a result, KVM rejects the requested compatibility
level during guest initialization.
Add support for retrieving host CPU compatibility capabilities for
nested guests on PowerVM (PAPR nested API v2). The hypervisor provides
the effective compatibility levels via the H_GUEST_GET_CAPABILITIES
hcall, which reflects the processor modes negotiated between the Power
hypervisor (L0) and the host partition (L1).
On pseries systems, obtain the capability bitmap using
plpar_guest_get_capabilities() and return it via struct
kvm_ppc_compat_caps. This information is then exposed to userspace
through the KVM_PPC_GET_COMPAT_CAPS ioctl.
Hook the implementation into the Book3S HV kvmppc_ops so that it can be
invoked by the generic KVM ioctl handling code.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
since this value will trikle back to userspace please apply a mask on
the hcall return value so that any reserved and non-PVR related bits
doesnt leak back to userspace.
Though currently we only supply the bits corresponding to supported
processor versions, it makes sense to mask out unrelated bits so that
they don't unnecesarily passed on to the userspace. I'll make the
changes in v4.
Thanks,
Amit
From: Amit Machhiwal <hidden> Date: 2026-06-10 15:53:37
On 2026/06/03 09:47 AM, Vaibhav Jain wrote:
Hi Amit,
Thanks for the patch. My review comments inline:
Amit Machhiwal [off-list ref] writes:
quoted
Currently, when booting a compatibility-mode KVM guest (L1) on a PowerNV
hypervisor (L0), the guest runs with the expected processor
compatibility level. However, when booting a nested KVM guest (L2)
inside the L1, QEMU derives the CPU model from the raw host PVR and
attempts to run the nested guest at that level, instead of honoring the
compatibility mode of the L1.
Extend host CPU compatibility capability reporting to support nested
virtualization on PowerNV systems (PAPR nested API v1).
For nested API v2 (PowerVM), compatibility capabilities are obtained
from the hypervisor via the H_GUEST_GET_CAPABILITIES hcall. This
information is not available on PowerNV systems.
For nested API v1, derive the compatibility capabilities from the L1
guest by reading the "cpu-version" property from the device tree, which
reflects the effective (logical) processor compatibility level. Map this
value to the corresponding compatibility capability bitmap.
Introduce a helper to translate CPU version values into compatibility
capability bits and integrate it into kvmppc_get_compat_cpu_caps().
This allows userspace to query host CPU compatibility modes on both
PowerVM and PowerNV platforms via the KVM_PPC_GET_COMPAT_CAPS ioctl.
Suggested-by: Vaibhav Jain <redacted>
Tested-by: Anushree Mathur <redacted>
Signed-off-by: Amit Machhiwal <redacted>
---
arch/powerpc/kvm/book3s_hv.c | 37 +++++++++++++++++++++++++++++++++++-
1 file changed, 36 insertions(+), 1 deletion(-)
Need to mask capabilities as mentioned in the review comments for
previous patch. I would suggest creating a helper that performs the
hcall and applies the mask which can then be used at
plpar_guest_get_capabilities() call sites.