Thread (112 messages) flat view 112 messages, 12 authors, 7d ago

Re: [PATCH v16 44/45] KVM: arm64: CCA: Require ICH_HCR_EL2.TDIR for realms

From: Kohei Enju <hidden>
Date: 2026-08-27 12:09:40
Also in: kvm, kvmarm, linux-coco, lkml

On 08/10 10:41, Marc Zyngier wrote:
On Mon, 10 Aug 2026 05:58:10 +0100,
Kohei Enju [off-list ref] wrote:
quoted
On 08/03 14:44, Steven Price wrote:
quoted
KVM advertises realm support when the RMM is available, and allows
userspace to create a VM with KVM_VM_TYPE_ARM_REALM on that basis.

On CPUs that lack ICH_HCR_EL2.TDIR, KVM uses ICH_HCR_EL2.TC for
normal guests so that ICC_DIR_EL1 is still trapped via the common GICv3
CPU interface trap. Realms cannot rely on the normal hyp-side trap
handling for that fallback, so advertising RMI support on such systems
lets userspace create a realm that cannot safely run.

Require the finalized ARM64_HAS_ICH_HCR_EL2_TDIR capability when
reporting KVM_CAP_ARM_RMI and when accepting KVM_VM_TYPE_ARM_REALM.
This leaves normal VM creation unchanged on systems that need the TC
workaround.
Hi Steven,

Thanks for your work on upstreaming CCA.

In the v15 discussion [0], you asked whether the system I was testing was a
"hacked up test system" or closer to "production hardware", and I said I would
share more when the time came. I can now say that this is not a hacked-up test
system. At Fujitsu, we have real hardware (FUJITSU-MONAKA) which implements CCA
(FEAT_RME) but does not implement FEAT_GICv3_TDIR. The hardware details are as
follows:

  - GICv4.2 compliant implementation
  - Supports FEAT_GICv3, FEAT_GICv3p1, FEAT_GICv4, FEAT_GICv4p1, and FEAT_GICv3_NMI
  - Does not support FEAT_GICv3_LEGACY (deprecated)
  - Does not support FEAT_GICv3_TDIR (ICH_VTR_EL2.TDS == 0)

For reference, compared with Arm Neoverse V3, the virtual GIC configuration is
largely equivalent. The only missing non-deprecated architectural feature is
FEAT_GICv3_TDIR.
A *very* significant difference. Given the cost of trapping between
R-EL1 and NS-EL2, something as simple as accesses to ICV_PMR_EL1
result in an extremely expensive trap.
quoted
The issue I see is that the CCA KVM code currently does not support a
configuration (non-TDIR/common-trap) that normal KVM already supports. For
normal guests, KVM handles systems without TDIR by using ICH_HCR_EL2.TC and the
existing GICv3 CPU interface emulation path. However, Realm guests currently
fail because the CCA path bypasses that existing emulation path, as Marc also
pointed out in [1].
Plugging CCA in the emulation code will solve the *functional* aspect.
The performance aspect is still there, unfortunately, and there isn't
much KVM can do about that.
quoted
Also, this is not limited to systems that actually lack TDIR. The same failure
can be reproduced on a TDIR-capable system by booting with:
  kvm-arm.vgic_v3_common_trap=1
This is a *debug* option for broken hardware. ThunderX, for
example. You really are in good company when this bit is set.
quoted
So it seems that the current CCA KVM implementation does not yet cover a
configuration that normal KVM already supports today, rather than this being a
limitation of the RMM specification or the underlying hardware.

I've included a patch below which reuses the existing GICv3 early emulation
path for Realm sysreg exits. This patch does not add any new vGIC emulation
code, and leaves the existing vGIC emulation code unchanged. So I believe this
is in line with Marc's request in [1]. With this patch, Realm guests can run
when the common CPU interface trap path is enabled.

I tested the exact patch both on our real silicon and on QEMU, and
confirmed that all Realm-related tests in kvm-unit-tests-cca passed.
KUTs are unfortunately not something that people run in production. I
wonder why...

Please run a Linux guest compiled with CONFIG_ARM64_PSEUDO_NMI=y and
irqchip.gicv3_pseudo_nmi=1 on the command line. Run any significant
workload (hackbench, for example), and report the overhead. This will
give you the expected impact introduced by the lack of TDIR.
Sorry for the delayed response.

Understood. I will run a Linux guest with CONFIG_ARM64_PSEUDO_NMI=y and
irqchip.gicv3_pseudo_nmi=1, measure a representative workload such as
hackbench, and report the overhead.

Thanks,
Kohei
	M.

-- 
Without deviation from the norm, progress is not possible.
  
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help