Re: [PATCH v10 12/41] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion
From: Binbin Wu <hidden>
Date: 2026-08-13 06:56:29
Also in:
kvm, linux-coco, linux-doc, linux-kselftest, linux-mm, lkml
On 8/8/2026 5:52 AM, Ackerley Tng via B4 Relay wrote:
From: Ackerley Tng <redacted> When memory in guest_memfd is converted from private to shared, the platform-specific state associated with the guest-private pages must be invalidated or cleaned up. Iterate over the folios in the affected range and call the kvm_arch_gmem_make_shared() hook for each PFN range. This allows architectures to update hardware metadata or encryption states to transition pages to the shared state. Invoke this helper after indicating to KVM's mmu code that an invalidation is in progress to stop in-flight page faults from succeeding. Omit support for calling the arch hook to make private, since SNP, the only implementer of the arch make-private hook today, would actually prefer making private only just before faulting memory into the NPTs. Calling the make-private arch hook would require iterating both bindings and the filemap to find the intersection of bindings and allocated folios. On top of that, SNP would need to figure out whether to actually make private based on whether the memory is about the be faulted, or
^ the -> to?
whether it is a conversion.
[...]
quoted hunk ↗ jump to hunk
+#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) +{ + struct folio_batch fbatch; + pgoff_t next = start; + int i; + + folio_batch_init(&fbatch); + while (filemap_get_folios(inode->i_mapping, &next, end - 1, &fbatch)) { + for (i = 0; i < folio_batch_count(&fbatch); ++i) { + struct folio *folio = fbatch.folios[i]; + pgoff_t start_index, end_index; + kvm_pfn_t start_pfn; + kvm_pfn_t nr_pages; + + start_index = max(start, folio->index); + end_index = min(end, folio_next_index(folio)); + /* + * end_index is either in folio or points to + * the first page of the next folio. Hence, + * all pages in range [start_index, end_index) + * are contiguous. + */ + start_pfn = folio_file_pfn(folio, start_index); + nr_pages = end_index - start_index; + + kvm_arch_gmem_make_shared(start_pfn, nr_pages); + } + + folio_batch_release(&fbatch); + cond_resched(); + } +} +#else +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) {} +#endif + static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start, size_t nr_pages, uint64_t attrs, pgoff_t *err_index)@@ -599,7 +636,12 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start, filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE; kvm_gmem_invalidate_start(inode, start, end, filter); + + if (!to_private) + kvm_gmem_make_shared(inode, start, end);
If both KVM_AMD_SEV and KVM_INTEL_TDX are enabled, HAVE_KVM_ARCH_GMEM_CONVERT will be enabled and the logic in kvm_gmem_make_shared() introduces unnecessary overhead for TDX. Not sure about CSPs, but in a standard distribution kernel, it's very likely that both are enabled, right? Should kvm_gmem_make_shared() do some optimization or the overhead is relative small in the conversion to shared path so that the optimization is not worth it?
+ mas_store_prealloc(&mas, xa_mk_value(attrs)); + kvm_gmem_invalidate_end(inode, start, end); out: filemap_invalidate_unlock(mapping);