Thread (10 messages) flat view 10 messages, 3 authors, 3d ago

Re: [PATCH v6 0/4] KVM: PPC: Expose CPU compatibility modes for nested guests

From: Amit Machhiwal <hidden>
Date: 2026-08-06 05:32:47
Also in: kvm, linux-doc, lkml

Hi Ritesh,

Thanks for taking a look. Please find my response inline.

On 2026/08/06 12:09 AM, Ritesh Harjani wrote:
Hi Amit,

Amit Machhiwal [off-list ref] writes:
quoted
On POWER systems, newer processor generations can operate in compatibility
modes corresponding to earlier generations (e.g., a Power11 system running
in Power10 compatibility mode). In such cases, the effective CPU level
exposed to guests differs from the physical processor generation.

This creates a problem for nested virtualization. When booting a nested KVM
guest (L2) inside a host KVM guest (L1) running in a compatibility mode,
userspace (e.g., QEMU) may derive the CPU model from the raw hardware PVR
and attempt to configure the nested guest accordingly. However, the L1
partition is constrained by the compatibility level negotiated with the
hypervisor (L0), and requests exceeding that level are rejected, leading to
guest boot failures such as:

  KVM-NESTEDv2: couldn't set guest wide elements

This series provides a mechanism for userspace to query the effective CPU
compatibility modes supported by the host, so it can select an appropriate
CPU model for nested guests.

To achieve this, the series introduces a new KVM capability and ioctl
(KVM_CAP_PPC_COMPAT_CAPS / KVM_PPC_GET_COMPAT_CAPS) that expose the
compatibility modes supported by the host.
Sorry, but I am somehow not convinced on whether we need all of this
machinary just to get these 3 bits of information, which we are
returning today.

Since KVM_CHECK_EXTENSION can already return an int, so why can't we use
KVM_CAP_PPC_COMPAT_CAPS itself and return the bitmap of supported compat
modes to the user?
Say if the cap is not supported, we can return 0, otherwise we can
return the bitmap of supported compat modes. This will easily allow us
to use 31-bits which as I see would be hardly a problem in the near
future. In the future if it grows - we can always use KVM_CAP_PPC_COMPAT_CAPS2.

This should reduce the code complexity both in the kernel and
userspace and we don't even need a new ioctl then.
Thanks for the suggestion. I considered this approach but would like to
go with a dedicated ioctl for the following reasons:

1. Intended semantics: The KVM API documentation states:

   ..kvm defines extension identifiers and a facility to query
   whether a particular extension identifier is available.  If it is, a
   set of ioctls is available for application use.

   [...]

   KVM defines many constants of the form KVM_CAP_*, each corresponding
   to a set of functionality provided by one or more ioctls. Availability
   of these capabilities can be checked with KVM_CHECK_EXTENSION.

   The intended role of KVM_CAP_* is to signal ioctl availability, not
   to serve as a data retrieval mechanism itself.

   You may take a look at KVM_CAP_PPC_GET_CPU_CHAR for instance.

2. Return type constraint: KVM_CHECK_EXTENSION returns a signed 32-bit int. The
   capability bits are defined as (1ULL << 62), (1ULL << 61), and (1ULL << 60) —
   64-bit values that cannot fit in a 32-bit return. Renumbering them to small
   integers would be a UAPI change and would lose alignment with the
   H_GUEST_CAP_* values from the hypervisor ABI.

3. Extensibility: The struct-based approach with the size field provides clean
   forward and backward ABI versioning via copy_struct_from/to_user(), without
   needing a KVM_CAP_PPC_COMPAT_CAPS2 in the future.

Thanks,
Amit
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help