From: Marc Zyngier <maz@kernel.org> Date: 2021-07-26 15:36:07
A while ago, Willy and Sean pointed out[0] that arm64 is the last user
of kvm_is_transparent_hugepage(), and that there would actually be
some benefit in looking at the userspace mapping directly instead.
This small series does exactly that, although it doesn't try to
support more than a PMD-sized mapping yet for THPs. We could probably
look into unifying this with the huge PUD code, and there is still
some potential use of the contiguous hint.
As a consequence, it removes kvm_is_transparent_hugepage(),
PageTransCompoundMap() and kvm_get_pfn(), all of which have no user
left after this rework.
This has been lightly tested on an Altra box (VHE) and on a SC2A11
system (nVHE). Although nothing caught fire, it requires some careful
reviewing on the arm64 side.
* From v1 [1]:
- Move the PT helper into its own function, as both Quentin and I
need it for other developments
- Fixed stupid bug introduced by a bad conflict resolution, spotted
by Alexandru
- Collected Acks from Paolo, with thanks
[0] https://lore.kernel.org/r/YLpLvFPXrIp8nAK4@google.com
[1] https://lore.kernel.org/r/20210717095541.1486210-1-maz@kernel.org
Marc Zyngier (6):
KVM: arm64: Introduce helper to retrieve a PTE and its level
KVM: arm64: Walk userspace page tables to compute the THP mapping size
KVM: arm64: Avoid mapping size adjustment on permission fault
KVM: Remove kvm_is_transparent_hugepage() and PageTransCompoundMap()
KVM: arm64: Use get_page() instead of kvm_get_pfn()
KVM: Get rid of kvm_get_pfn()
arch/arm64/include/asm/kvm_pgtable.h | 19 ++++++++++++
arch/arm64/kvm/hyp/pgtable.c | 39 ++++++++++++++++++++++++
arch/arm64/kvm/mmu.c | 45 +++++++++++++++++++++++-----
include/linux/kvm_host.h | 1 -
include/linux/page-flags.h | 37 -----------------------
virt/kvm/kvm_main.c | 19 +-----------
6 files changed, 97 insertions(+), 63 deletions(-)
--
2.30.2
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel
From: Marc Zyngier <maz@kernel.org> Date: 2021-07-26 15:36:00
We currently rely on the kvm_is_transparent_hugepage() helper to
discover whether a given page has the potential to be mapped as
a block mapping.
However, this API doesn't really give un everything we want:
- we don't get the size: this is not crucial today as we only
support PMD-sized THPs, but we'd like to have larger sizes
in the future
- we're the only user left of the API, and there is a will
to remove it altogether
To address the above, implement a simple walker using the existing
page table infrastructure, and plumb it into transparent_hugepage_adjust().
No new page sizes are supported in the process.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/kvm/mmu.c | 34 ++++++++++++++++++++++++++++++----
1 file changed, 30 insertions(+), 4 deletions(-)
@@ -433,6 +433,32 @@ int create_hyp_exec_mappings(phys_addr_t phys_addr, size_t size,return0;}+staticstructkvm_pgtable_mm_opskvm_user_mm_ops={+/* We shouldn't need any other callback to walk the PT */+.phys_to_virt=kvm_host_va,+};++staticintget_user_mapping_size(structkvm*kvm,u64addr)+{+structkvm_pgtablepgt={+.pgd=(kvm_pte_t*)kvm->mm->pgd,+.ia_bits=VA_BITS,+.start_level=(KVM_PGTABLE_MAX_LEVELS-+CONFIG_PGTABLE_LEVELS),+.mm_ops=&kvm_user_mm_ops,+};+kvm_pte_tpte=0;/* Keep GCC quiet... */+u32level=~0;+intret;++ret=kvm_pgtable_get_leaf(&pgt,addr,&pte,&level);+VM_BUG_ON(ret);+VM_BUG_ON(level>=KVM_PGTABLE_MAX_LEVELS);+VM_BUG_ON(!(pte&PTE_VALID));++returnBIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(level));+}+staticstructkvm_pgtable_mm_opskvm_s2_mm_ops={.zalloc_page=stage2_memcache_zalloc_page,.zalloc_pages_exact=kvm_host_zalloc_pages_exact,
From: Marc Zyngier <maz@kernel.org> Date: 2021-07-26 15:36:01
It is becoming a common need to fetch the PTE for a given address
together with its level. Add such a helper.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/include/asm/kvm_pgtable.h | 19 ++++++++++++++
arch/arm64/kvm/hyp/pgtable.c | 39 ++++++++++++++++++++++++++++
2 files changed, 58 insertions(+)
From: Marc Zyngier <maz@kernel.org> Date: 2021-07-26 15:36:03
When mapping a THP, we are guaranteed that the page isn't reserved,
and we can safely avoid the kvm_is_reserved_pfn() call.
Replace kvm_get_pfn() with get_page(pfn_to_page()).
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/kvm/mmu.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Marc Zyngier <maz@kernel.org> Date: 2021-07-26 15:36:05
Since we only support PMD-sized mappings for THP, getting
a permission fault on a level that results in a mapping
being larger than PAGE_SIZE is a sure indication that we have
already upgraded our mapping to a PMD.
In this case, there is no need to try and parse userspace page
tables, as the fault information already tells us everything.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/kvm/mmu.c | 11 ++++++++---
1 file changed, 8 insertions(+), 3 deletions(-)
From: Marc Zyngier <maz@kernel.org> Date: 2021-07-26 15:36:10
Nobody is using kvm_get_pfn() anymore. Get rid of it.
Acked-by: Paolo Bonzini <pbonzini@redhat.com>
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
include/linux/kvm_host.h | 1 -
virt/kvm/kvm_main.c | 9 +--------
2 files changed, 1 insertion(+), 9 deletions(-)
From: Marc Zyngier <maz@kernel.org> Date: 2021-07-26 15:36:13
Now that arm64 has stopped using kvm_is_transparent_hugepage(),
we can remove it, as well as PageTransCompoundMap() which was
only used by the former.
Acked-by: Paolo Bonzini <pbonzini@redhat.com>
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
include/linux/page-flags.h | 37 -------------------------------------
virt/kvm/kvm_main.c | 10 ----------
2 files changed, 47 deletions(-)
@@ -632,43 +632,6 @@ static inline int PageTransCompound(struct page *page)returnPageCompound(page);}-/*-*PageTransCompoundMapisthesameasPageTransCompound,butitalso-*guaranteestheprimaryMMUhastheentirecompoundpagemapped-*throughpmd_trans_huge,whichinturnguaranteesthesecondaryMMUs-*canalsomaptheentirecompoundpage.Thisallowsthesecondary-*MMUstocallget_user_pages()onlyonceforeachcompoundpageand-*toimmediatelymaptheentirecompoundpagewithasinglesecondary-*MMUfault.Iftherewillbeapmdsplitlater,thesecondaryMMUs-*willgetanupdatethroughtheMMUnotifierinvalidationthrough-*split_huge_pmd().-*-*UnlikePageTransCompound,thisissafetobecalledonlywhile-*split_huge_pmd()cannotrunfromunderus,likeifprotectedbythe-*MMUnotifier,otherwiseitmayresultinpage->_mapcountcheckfalse-*positives.-*-*WehavetotreatpagecacheTHPdifferentlysinceeverysubpageofit-*wouldget_mapcountinc'edonceitisPMDmapped.But,itmaybePTE-*mappedinthecurrentprocesssocomparingsubpage's_mapcountto-*compound_mapcounttofilteroutPTEmappedcase.-*/-staticinlineintPageTransCompoundMap(structpage*page)-{-structpage*head;--if(!PageTransCompound(page))-return0;--if(PageAnon(page))-returnatomic_read(&page->_mapcount)<0;--head=compound_head(page);-/* File THP is PMD mapped and not PTE mapped */-returnatomic_read(&page->_mapcount)==-atomic_read(compound_mapcount_ptr(head));-}-/**PageTransTailreturnstrueforbothtransparenthugepages*andhugetlbfspages,soitshouldonlybecalledwhenit'sknown
From: Alexandru Elisei <hidden> Date: 2021-07-27 15:24:32
Hi Marc,
On 7/26/21 4:35 PM, Marc Zyngier wrote:
quoted hunk
It is becoming a common need to fetch the PTE for a given address
together with its level. Add such a helper.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/include/asm/kvm_pgtable.h | 19 ++++++++++++++
arch/arm64/kvm/hyp/pgtable.c | 39 ++++++++++++++++++++++++++++
2 files changed, 58 insertions(+)
Yet in the next patch you use a struct kvm_pgtable_pgt not initialized by any of
the kvm_pgtable_*_init() functions. It doesn't hurt correctness, but it might
confuse potential users of this function.
quoted hunk
+ * @addr: Input address for the start of the walk.
+ * @ptep: Pointer to storage for the retrieved PTE.
+ * @level: Pointer to storage for the level of the retrieved PTE.
+ *
+ * The offset of @addr within a page is ignored.
+ *
+ * The walker will walk the page-table entries corresponding to the input
+ * address specified, retrieving the leaf corresponding to this address.
+ * Invalid entries are treated as leaf entries.
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,
+ kvm_pte_t *ptep, u32 *level);
+
/**
* kvm_pgtable_stage2_find_range() - Find a range of Intermediate Physical
* Addresses with compatible permission
kvm_pgtable_walk() already aligns addr down to PAGE_SIZE, I don't think that's
needed here. But not harmful either.
Otherwise, the patch looks good to me:
Reviewed-by: Alexandru Elisei <redacted>
Thanks,
Alex
From: Alexandru Elisei <hidden> Date: 2021-07-27 15:54:37
Hi Marc,
On 7/26/21 4:35 PM, Marc Zyngier wrote:
quoted hunk
We currently rely on the kvm_is_transparent_hugepage() helper to
discover whether a given page has the potential to be mapped as
a block mapping.
However, this API doesn't really give un everything we want:
- we don't get the size: this is not crucial today as we only
support PMD-sized THPs, but we'd like to have larger sizes
in the future
- we're the only user left of the API, and there is a will
to remove it altogether
To address the above, implement a simple walker using the existing
page table infrastructure, and plumb it into transparent_hugepage_adjust().
No new page sizes are supported in the process.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/kvm/mmu.c | 34 ++++++++++++++++++++++++++++++----
1 file changed, 30 insertions(+), 4 deletions(-)
@@ -433,6 +433,32 @@ int create_hyp_exec_mappings(phys_addr_t phys_addr, size_t size,return0;}+staticstructkvm_pgtable_mm_opskvm_user_mm_ops={+/* We shouldn't need any other callback to walk the PT */
That looks correct to me, mm_ops is used in __kvm_pgtable_visit(), and then only
the phys_to_virt field callback is used. kvm_host_va() is also the callback used
by kvm_s2_mm_ops, which looks right to me.
@@ -780,7 +806,7 @@ static bool fault_supports_stage2_huge_mapping(struct kvm_memory_slot *memslot, * Returns the size of the mapping. */ static unsigned long-transparent_hugepage_adjust(struct kvm_memory_slot *memslot,+transparent_hugepage_adjust(struct kvm *kvm, struct kvm_memory_slot *memslot, unsigned long hva, kvm_pfn_t *pfnp, phys_addr_t *ipap) {
@@ -791,8 +817,8 @@ transparent_hugepage_adjust(struct kvm_memory_slot *memslot, * sure that the HVA and IPA are sufficiently aligned and that the * block map is contained within the memslot. */- if (kvm_is_transparent_hugepage(pfn) &&- fault_supports_stage2_huge_mapping(memslot, hva, PMD_SIZE)) {+ if (fault_supports_stage2_huge_mapping(memslot, hva, PMD_SIZE) &&+ get_user_mapping_size(kvm, hva) >= PMD_SIZE) { /* * The address we faulted on is backed by a transparent huge * page. However, because we map the compound huge page and
@@ -1051,7 +1077,7 @@ static int user_mem_abort(struct kvm_vcpu *vcpu, phys_addr_t fault_ipa, * backed by a THP and thus use block mapping if possible. */ if (vma_pagesize == PAGE_SIZE && !(force_pte || device))- vma_pagesize = transparent_hugepage_adjust(memslot, hva,+ vma_pagesize = transparent_hugepage_adjust(kvm, memslot, hva, &pfn, &fault_ipa); if (fault_status != FSC_PERM && !device && kvm_has_mte(kvm)) {
Sean explained well why holding the mmap lock isn't needed here. The patch looks
correct to me:
Reviewed-by: Alexandru Elisei <redacted>
Thanks,
Alex
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel
From: Alexandru Elisei <hidden> Date: 2021-07-27 15:59:31
Hi Marc,
On 7/26/21 4:35 PM, Marc Zyngier wrote:
quoted hunk
Since we only support PMD-sized mappings for THP, getting
a permission fault on a level that results in a mapping
being larger than PAGE_SIZE is a sure indication that we have
already upgraded our mapping to a PMD.
In this case, there is no need to try and parse userspace page
tables, as the fault information already tells us everything.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/kvm/mmu.c | 11 ++++++++---
1 file changed, 8 insertions(+), 3 deletions(-)
From: Alexandru Elisei <hidden> Date: 2021-07-27 17:45:28
Hi Marc,
On 7/26/21 4:35 PM, Marc Zyngier wrote:
quoted hunk
When mapping a THP, we are guaranteed that the page isn't reserved,
and we can safely avoid the kvm_is_reserved_pfn() call.
Replace kvm_get_pfn() with get_page(pfn_to_page()).
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/kvm/mmu.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
I am not very familiar with the mm subsystem, but I did my best to review this change.
kvm_get_pfn() uses get_page(pfn) if !PageReserved(pfn_to_page(pfn)). I looked at
the documentation for the PG_reserved page flag, and for normal memory, what
looked to me like the most probable situation where that can be set for a
transparent hugepage was for the zero page. Looked at mm/huge_memory.c, and
huge_zero_pfn is allocated via alloc_pages(__GFP_ZERO) (and other flags), which
doesn't call SetPageReserved().
I looked at how a huge page can be mapped from handle_mm_fault and from
khugepaged, and it also looks to like both are using using alloc_pages() to
allocate a new hugepage.
I also did a grep for SetPageReserved(), and there are very few places where that
is called, and none looked like they have anything to do with hugepages.
As far as I can tell, this change is correct, but I think someone who is familiar
with mm would be better suited for reviewing this patch.
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel
From: Marc Zyngier <maz@kernel.org> Date: 2021-07-28 12:17:16
Hi Alex,
On Tue, 27 Jul 2021 16:25:34 +0100,
Alexandru Elisei [off-list ref] wrote:
Hi Marc,
On 7/26/21 4:35 PM, Marc Zyngier wrote:
quoted
It is becoming a common need to fetch the PTE for a given address
together with its level. Add such a helper.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/include/asm/kvm_pgtable.h | 19 ++++++++++++++
arch/arm64/kvm/hyp/pgtable.c | 39 ++++++++++++++++++++++++++++
2 files changed, 58 insertions(+)
Yet in the next patch you use a struct kvm_pgtable_pgt not
initialized by any of the kvm_pgtable_*_init() functions. It doesn't
hurt correctness, but it might confuse potential users of this
function.
Fair enough. I'll add something like "[...] or any similar initialisation".
quoted
+ * @addr: Input address for the start of the walk.
+ * @ptep: Pointer to storage for the retrieved PTE.
+ * @level: Pointer to storage for the level of the retrieved PTE.
+ *
+ * The offset of @addr within a page is ignored.
+ *
+ * The walker will walk the page-table entries corresponding to the input
+ * address specified, retrieving the leaf corresponding to this address.
+ * Invalid entries are treated as leaf entries.
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,
+ kvm_pte_t *ptep, u32 *level);
+
/**
* kvm_pgtable_stage2_find_range() - Find a range of Intermediate Physical
* Addresses with compatible permission
kvm_pgtable_walk() already aligns addr down to PAGE_SIZE, I don't
think that's needed here. But not harmful either.
It is more that if you don't align it down, the size becomes awkward
to express. Masking is both cheap and readable.
Otherwise, the patch looks good to me:
Reviewed-by: Alexandru Elisei <redacted>
Thanks!
M.
--
Without deviation from the norm, progress is not possible.
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel
From: Marc Zyngier <maz@kernel.org> Date: 2021-08-02 13:39:59
On Mon, 26 Jul 2021 16:35:46 +0100, Marc Zyngier wrote:
A while ago, Willy and Sean pointed out[0] that arm64 is the last user
of kvm_is_transparent_hugepage(), and that there would actually be
some benefit in looking at the userspace mapping directly instead.
This small series does exactly that, although it doesn't try to
support more than a PMD-sized mapping yet for THPs. We could probably
look into unifying this with the huge PUD code, and there is still
some potential use of the contiguous hint.
[...]
Applied to next, thanks!
[1/6] KVM: arm64: Introduce helper to retrieve a PTE and its level
commit: 63db506e07622c344a3c748a1c06293d48780f83
[2/6] KVM: arm64: Walk userspace page tables to compute the THP mapping size
commit: 6011cf68c88545e16cb32039c2cecfdae6a32315
[3/6] KVM: arm64: Avoid mapping size adjustment on permission fault
commit: f2cc327303b13a70311e823bd52aa0bca8c7ddbc
[4/6] KVM: Remove kvm_is_transparent_hugepage() and PageTransCompoundMap()
commit: 205d76ff0684a0b4fe3ff3a283d143a47439d191
[5/6] KVM: arm64: Use get_page() instead of kvm_get_pfn()
commit: 0fe49630101b3ce23bd21a2788440ac719ec868a
[6/6] KVM: Get rid of kvm_get_pfn()
commit: 36c3ce6c0d03a6c9992c3359f879cdc70fde836a
Cheers,
M.
--
Without deviation from the norm, progress is not possible.
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel