Re: [PATCH 3/8] KVM: arm64: Propagate and use kvm_s2_fault_result on S2 fault
flat view
From: Marc Zyngier <maz@kernel.org>
Date: 2026-09-10 08:49:11
Also in:
kvm, kvmarm, linux-doc, linux-kselftest, lkml
On Tue, 25 Aug 2026 17:00:37 +0100, "Lorenzo Stoakes (ARM)" [off-list ref] wrote:
quoted hunk ↗ jump to hunk
When stage 2 page tables fault the net result may either be that a page is mapped, an error occurred or the fault should be retried (-EAGAIN). When a fault succeeds it may be upgraded to a PMD size via transparent_hugepage_adjust(). In order to support KVM pre-faulting the outcome of the fault and the mapping size must be recorded. Track this in the new kvm_s2_fault_result struct, which is threaded through gmem_abort(), user_mem_abort() and kvm_s2_fault_map(). PKVM and SEA aren't relevant to synthetic pre-faulting so neither kvm_inject_sea() nor pkvm_mem_abort() are altered. Actual hardware faulting doesn't require this information, so kvm_handle_guest_abort() simply passes NULL kvm_s2_fault_result to gmem_abort() and user_mem_abort(). Faults are necessarily ephemeral and pre-faulting can't guarantee what may happen in parallel, so do not store the GFN or PFN in kvm_s2_fault_result. Pre-faulting only needs to know what was mapped in at the point of the fault. This struct could be replaced with a pointer to an unsigned long, however it's clearer to separate out the mapped flag and having a struct allows us to easily add additional fields in future as needed. No functional change intended. Suggested-by: Vincent Donnefort <redacted> Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> --- arch/arm64/kvm/mmu.c | 33 +++++++++++++++++++++++++++------ 1 file changed, 27 insertions(+), 6 deletions(-)diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c index 80cb520e25b9..da15da4e40e6 100644 --- a/arch/arm64/kvm/mmu.c +++ b/arch/arm64/kvm/mmu.c@@ -1607,6 +1607,11 @@ struct kvm_s2_fault_desc { struct kvm_s2_mmu *mmu; }; +struct kvm_s2_fault_result { + unsigned long mapping_size; + bool mapped; +}; + static bool kvm_s2_fault_is_perm(const struct kvm_s2_fault_desc *s2fd) { return esr_fsc_is_permission_fault(s2fd->esr);@@ -1632,7 +1637,17 @@ static u64 kvm_s2_perm_fault_granule(const struct kvm_s2_fault_desc *s2fd) return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(level)); } -static int gmem_abort(const struct kvm_s2_fault_desc *s2fd) +static void populate_fault_result(struct kvm_s2_fault_result *result, + unsigned long mapping_size) +{ + /* A THP upgrade may have altered mapping size. */ + result->mapping_size = mapping_size; + /* -EAGAIN is swallowed so be explicit when we actually map. */ + result->mapped = true;
I'm not sold on this boolean. I'd rather we use the fact that the fault handler has passed a result pointer to return -EAGAIN rather than turning into a 0, because that's a clear sign that the fault hasn't been generated by a vcpu. M. -- Without deviation from the norm, progress is not possible.