From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:36
Here's v13. Thanks everyone for the comments and fast responses!
We're now at ~3 weeks to soft-close at 7.3-rc5.
v13 is based on 7.3-rc2 and:
+ Hugh's patch to export lru_cache_drain_for_folio()
+ Another series [1], which makes kvm_gmem_get_pfn() NOT return a
refcounted page to KVM.
Here's everything stitched together for your convenience:
https://github.com/googleprodkernel/linux-cc/commits/guest_memfd-inplace-conversion-v13
Changes in v13:
+ Picked up Reviewed-bys
+ Fixed bug Fuad pointed out in "Ensure pages are not in use before
conversion", and squashed the lru draining in, adopted Sean's
suggestion.
+ Fixed documentation bug that Fuad pointed out in
kernel-parameters.txt.
Here's v13 with tests:
https://github.com/googleprodkernel/linux-cc/commits/guest_memfd-inplace-conversion-coco-selftests-v13
Tested with both CONFIG_KVM_VM_MEMORY_ATTRIBUTES enabled and disabled:
+ tools/testing/selftests/kvm/guest_memfd_test.c
+ tools/testing/selftests/kvm/pre_fault_memory_test.c
+ tools/testing/selftests/kvm/x86/guest_memfd_conversions_test.c
+ tools/testing/selftests/kvm/x86/private_mem_conversions_test.c
+ tools/testing/selftests/kvm/x86/private_mem_kvm_exits_test.c
[1] https://lore.kernel.org/all/20260826-gmem-no-return-page-v4-0-3bb9c1ddb4e3@google.com/
v12: https://lore.kernel.org/r/20260830-gmem-inplace-conversion-v12-0-85e5fd25252a@google.com
v11: https://lore.kernel.org/r/20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com
v10: https://lore.kernel.org/r/20260807-gmem-inplace-conversion-v10-0-2fc18ee6d3ba@google.com
v9: https://lore.kernel.org/r/20260728-gmem-inplace-conversion-v9-0-35f9aec2aed2@google.com
v8: https://lore.kernel.org/r/20260618-gmem-inplace-conversion-v8-0-9d2959357853@google.com
v7: https://lore.kernel.org/r/20260522-gmem-inplace-conversion-v7-0-2f0fae496530@google.com
v6: https://lore.kernel.org/r/20260507-gmem-inplace-conversion-v6-0-91ab5a8b19a4@google.com
RFC v5: https://lore.kernel.org/r/20260428-gmem-inplace-conversion-v5-0-d8608ccfca22@google.com
RFC v4: https://lore.kernel.org/r/20260326-gmem-inplace-conversion-v4-0-e202fe950ffd@google.com
RFC v3: https://lore.kernel.org/r/20260313-gmem-inplace-conversion-v3-0-5fc12a70ec89@google.com
RFC v2: https://lore.kernel.org/r/cover.1770071243.git.ackerleytng@google.com
RFC v1: https://lore.kernel.org/r/cover.1760731772.git.ackerleytng@google.com
Previous versions of this feature, part of other series, are available at:
+ https://lore.kernel.org/all/bd163de3118b626d1005aa88e71ef2fb72f0be0f.1726009989.git.ackerleytng@google.com/
+ https://lore.kernel.org/all/20250117163001.2326672-6-tabba@google.com/
+ https://lore.kernel.org/all/b784326e9ccae6a08388f1bf39db70a2204bdc51.1747264138.git.ackerleytng@google.com/
Signed-off-by: Ackerley Tng <redacted>
---
Ackerley Tng (22):
KVM: Rename kvm_mem_is_private() to kvm_is_private_gfn()
KVM: guest_memfd: Always fault from guest_memfd if in-place conversion is enabled
KVM: guest_memfd: Pass mapping type filter to invalidation helper
KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2
KVM: guest_memfd: Ensure pages are not in use before conversion
KVM: guest_memfd: Call arch make_shared callback for to-shared conversion
KVM: guest_memfd: Return early if range already has requested attributes
KVM: guest_memfd: Zero page while getting pfn
KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION
KVM: selftests: Test basic single-page conversion flow
KVM: selftests: Test conversion flow when INIT_SHARED
KVM: selftests: Test conversion precision in guest_memfd
KVM: selftests: Test conversion before allocation
KVM: selftests: Convert with allocated folios in different layouts
KVM: selftests: Test that truncation does not change shared/private status
KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST
KVM: selftests: Test conversion with elevated page refcount
KVM: selftests: Reset shared memory after hole-punching
KVM: selftests: Provide function to look up guest_memfd details from gpa
KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe
KVM: selftests: Set up page size and alignment independently for guest_memfd
KVM: selftests: Update private_mem_conversions_test for in-place conversions
Michael Roth (1):
KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE
Sean Christopherson (21):
KVM: guest_memfd: Optimize away conversion overheads via dead-code elimination
KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory
KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings
KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES
KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined
KVM: Rename memory attribute APIs to prepare for in-place gmem conversion
KVM: Provide generic interface for checking memory private/shared status
KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion
KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h
KVM: guest_memfd: Invalidate both SHARED and PRIVATE mappings for in-place conversions
KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86
KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes
KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs
KVM: selftests: Create gmem fd before "regular" fd when adding memslot
KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset}
KVM: selftests: Add support for mmap() on guest_memfd in core library
KVM: selftests: Add selftests global for guest memory attributes capability
KVM: selftests: Add helpers for calling ioctls on guest_memfd
KVM: selftests: Test that shared/private status is consistent across processes
KVM: selftests: Provide common function to set memory attributes
KVM: selftests: Update private memory exits test to work with per-gmem attributes
Documentation/admin-guide/kernel-parameters.txt | 25 +
Documentation/virt/kvm/api.rst | 110 ++++-
.../virt/kvm/x86/amd-memory-encryption.rst | 14 +-
Documentation/virt/kvm/x86/intel-tdx.rst | 4 +
arch/x86/include/asm/kvm-x86-ops.h | 2 +-
arch/x86/include/asm/kvm_host.h | 9 +-
arch/x86/kvm/Kconfig | 15 +-
arch/x86/kvm/mmu/mmu.c | 28 +-
arch/x86/kvm/svm/sev.c | 13 +-
arch/x86/kvm/vmx/tdx.c | 8 +-
arch/x86/kvm/x86.c | 20 +-
include/linux/kvm_host.h | 83 ++--
include/trace/events/kvm.h | 6 +-
include/uapi/linux/kvm.h | 16 +
tools/testing/selftests/kvm/Makefile.kvm | 1 +
tools/testing/selftests/kvm/include/kvm_util.h | 139 +++++-
tools/testing/selftests/kvm/include/test_util.h | 34 +-
tools/testing/selftests/kvm/lib/kvm_util.c | 222 +++++----
tools/testing/selftests/kvm/lib/test_util.c | 7 -
.../kvm/x86/guest_memfd_conversions_test.c | 512 +++++++++++++++++++++
.../kvm/x86/private_mem_conversions_test.c | 66 ++-
.../selftests/kvm/x86/private_mem_kvm_exits_test.c | 36 +-
virt/kvm/Kconfig | 3 -
virt/kvm/guest_memfd.c | 468 +++++++++++++++++--
virt/kvm/kvm_main.c | 92 ++--
25 files changed, 1642 insertions(+), 291 deletions(-)
---
base-commit: 5e036ce12de91c6fd674dad33b169c6150be2a7a
change-id: 20260225-gmem-inplace-conversion-bd0dbd39753a
prerequisite-message-id: 02876cea-5727-2ca4-bead-73659ea6fec4@google.com
prerequisite-patch-id: 28922a76dee80792da6e445a55cdbb4bedc85374
prerequisite-change-id: 20260818-gmem-no-return-page-614927a29f97:v4
prerequisite-patch-id: 38553271bcb80c17dbae29cef5c20bcf95e804f9
prerequisite-patch-id: 0d4b50926073ba4f5e5d8b2c0cdae8869a278863
prerequisite-patch-id: bf39277ddc12818fe562a006f31d7992e5c7d42a
prerequisite-patch-id: 74fe849c04c187074bd380e664e536a07076292d
prerequisite-patch-id: f92bdb4870dd7c1a2026fca1083e4400c2be11d8
Best regards,
--
Ackerley Tng [off-list ref]
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Add and use kvm_arch_has_gmem_convert() to guard guest_memfd's invocation
of arch hooks related to converting memory between private and shared, as
only one half of the x86 CoCo duo needs the runtime hooks (any pre-work is
pure overhead for TDX). At this exact moment, the overhead is negligible,
but that will change when in-place conversion comes along, at which point
to-shared conversions will "need" to find all affected folios prior to
calling into arch code. In quotes because very technically that work could
be pushed to arch code, but that would bleed guest_memfd details into arch
code and would be far worse than adding yet another kvm_arch_has... hook.
Opportunistically provide the kvm_arch_gmem_make_private() declaration, and
rely on dead-code elimination to eliminate the call to non-existent code
when CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT=n.
Reported-by: Binbin Wu <redacted>
Closes: https://lore.kernel.org/all/1ec08cd8-3072-4753-ad5e-cd34956647f8@linux.intel.com
Suggested-by: Ackerley Tng <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
arch/x86/include/asm/kvm_host.h | 3 +++
include/linux/kvm_host.h | 3 ++-
virt/kvm/guest_memfd.c | 5 ++---
3 files changed, 7 insertions(+), 4 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Use kvm_mem_is_private() when populating guest_memfd instead of using an
open coded equivalent. In addition to simplifying the populate code *now*,
this avoids the need to provide a range-based gmem lookup API in the future
as well.
No functional change intended.
Suggested-by: Xiaoyao Li <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
virt/kvm/guest_memfd.c | 4 +---
1 file changed, 1 insertion(+), 3 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Start plumbing in guest_memfd support for in-place private<=>shared
conversions by tracking attributes via a maple tree. KVM currently tracks
private vs. shared attributes on a per-VM basis, which made sense when a
guest_memfd _only_ supported private memory, but tracking per-VM simply
can't work for in-place conversions as the shared/private status of a given
page needs to be per-gmem_inode, not per-VM.
Use the filemap invalidation lock to protect the maple tree, as taking the
lock for read when faulting in memory (for userspace or the guest) isn't
expected to result in meaningful contention, and using a separate lock
would add significant complexity (avoiding deadlock is quite difficult).
In kvm_gmem_get_pfn(), drop the folio refcount before releasing
filemap_invalidate_lock(). This ensures that a competing conversion request
from userspace (to be added in a later patch), which also takes the
filemap_invalidate_lock(), will never see an elevated refcount due to
kvm_gmem_get_pfn().
Co-developed-by: Vishal Annapurve <redacted>
Signed-off-by: Vishal Annapurve <redacted>
Co-developed-by: Fuad Tabba <redacted>
Signed-off-by: Fuad Tabba <redacted>
Co-developed-by: Ackerley Tng <redacted>
Signed-off-by: Ackerley Tng <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Binbin Wu <redacted>
---
virt/kvm/guest_memfd.c | 136 ++++++++++++++++++++++++++++++++++++++++++-------
1 file changed, 119 insertions(+), 17 deletions(-)
@@ -550,16 +625,9 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)gotoerr_fops;}-inode->i_op=&kvm_gmem_iops;-inode->i_mapping->a_ops=&kvm_gmem_aops;-inode->i_mode|=S_IFREG;-inode->i_size=size;-mapping_set_gfp_mask(inode->i_mapping,GFP_HIGHUSER);-mapping_set_inaccessible(inode->i_mapping);-/* Unmovable mappings are supposed to be marked unevictable as well. */-WARN_ON_ONCE(!mapping_unevictable(inode->i_mapping));--GMEM_I(inode)->flags=flags;+err=kvm_gmem_init_inode(inode,size,flags);+if(err)+gotoerr_inode;file=alloc_file_pseudo(inode,kvm_gmem_mnt,name,O_RDWR,&kvm_gmem_fops);if(IS_ERR(file)){
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Rename the per-VM memory attributes Kconfig to make it explicitly about
per-VM attributes in anticipation of adding memory attributes support to
guest_memfd, at which point it will be possible (and desirable) to have
memory attributes without the per-VM support, even in x86.
No functional change intended.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Binbin Wu <redacted>
---
arch/x86/include/asm/kvm_host.h | 2 +-
arch/x86/kvm/Kconfig | 6 +++---
arch/x86/kvm/mmu/mmu.c | 2 +-
arch/x86/kvm/x86.c | 2 +-
include/linux/kvm_host.h | 8 ++++----
include/trace/events/kvm.h | 4 ++--
virt/kvm/Kconfig | 2 +-
virt/kvm/kvm_main.c | 14 +++++++-------
8 files changed, 20 insertions(+), 20 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Explicitly guard reporting support for KVM_MEMORY_ATTRIBUTE_PRIVATE based
on kvm_arch_has_private_mem being #defined in anticipation of tracking
PRIVATE vs. SHARED state per-guest_memfd, not per-VM (to allow in-place
conversion).
guest_memfd support for memory attributes is expected to be unconditional
to avoid yet more macros (all architectures that support guest_memfd are
expected to use per-gmem attributes at some point), at which point
enumerating support KVM_MEMORY_ATTRIBUTE_PRIVATE based solely on memory
attributes being supported by KVM at-large would result in a system-scope
check (NULL @kvm) over-reporting support on arm64.
Give architectures full control over overriding the default definition of
kvm_arch_has_private_mem() by removing the coupling with
CONFIG_KVM_VM_MEMORY_ATTRIBUTES.
In a later patch, kvm_arch_has_private_mem() will be defined based on
whether architectural features are compiled in, and made orthogonal to
CONFIG_KVM_VM_MEMORY_ATTRIBUTES.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: Xiaoyao Li <redacted>
---
include/linux/kvm_host.h | 2 +-
virt/kvm/kvm_main.c | 2 ++
2 files changed, 3 insertions(+), 1 deletion(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Ackerley Tng <redacted>
Rename kvm_mem_is_private() to kvm_is_private_gfn() to prepare for in-place
conversion, where there will be two lookup functions,
kvm_vm_is_private_gfn() and kvm_gmem_is_private_gfn().
This renaming allows consistent prefixing of "vm" vs "gmem" for
kvm_*_is_private_gfn(), as opposed to kvm_gmem_mem_is_private(), which
looks like a typo.
Suggested-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <redacted>
---
arch/x86/kvm/mmu/mmu.c | 12 ++++++------
arch/x86/kvm/svm/sev.c | 2 +-
include/linux/kvm_host.h | 4 ++--
virt/kvm/guest_memfd.c | 2 +-
4 files changed, 10 insertions(+), 10 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Rename memory attribute APIs to add a "vm_" in the name in anticipation of
moving PRIVATE tracking into guest_memfd, to allow in-place conversion
between SHARED and PRIVATE. At that point, there will effectively be two
(potential) sources of memory attributes: the VM and guest_memfd.
kvm_vm_set_mem_attributes() already has "vm" in the name to indicate that
it is a VM ioctl. Rename it to kvm_set_vm_mem_attributes() to show that it
is setting the VM's memory attributes. (Drop the VM-ioctl scoping since it
is a helper local to the file.) Update the accompanying trace function to
match.
No functional change intended.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Binbin Wu <redacted>
---
arch/x86/kvm/mmu/mmu.c | 14 +++++++-------
include/linux/kvm_host.h | 16 ++++++++--------
include/trace/events/kvm.h | 2 +-
virt/kvm/kvm_main.c | 32 ++++++++++++++++----------------
4 files changed, 32 insertions(+), 32 deletions(-)
@@ -2533,18 +2533,18 @@ static bool kvm_pre_set_memory_attributes(struct kvm *kvm,*/kvm_mmu_invalidate_range_add(kvm,range->start,range->end);-returnkvm_arch_pre_set_memory_attributes(kvm,range);+returnkvm_arch_pre_set_vm_memory_attributes(kvm,range);}/* Set @attributes for the gfn range [@start, @end). */-staticintkvm_vm_set_mem_attributes(structkvm*kvm,gfn_tstart,gfn_tend,+staticintkvm_set_vm_mem_attributes(structkvm*kvm,gfn_tstart,gfn_tend,unsignedlongattributes){structkvm_mmu_notifier_rangepre_set_range={.start=start,.end=end,.arg.attributes=attributes,-.handler=kvm_pre_set_memory_attributes,+.handler=kvm_pre_set_vm_memory_attributes,.on_lock=kvm_mmu_invalidate_start,.flush_on_ret=true,.may_block=true,
@@ -2563,12 +2563,12 @@ static int kvm_vm_set_mem_attributes(struct kvm *kvm, gfn_t start, gfn_t end,entry=attributes?xa_mk_value(attributes):NULL;-trace_kvm_vm_set_mem_attributes(start,end,attributes);+trace_kvm_set_vm_mem_attributes(start,end,attributes);mutex_lock(&kvm->slots_lock);/* Nothing to do if the entire range has the desired attributes. */-if(kvm_range_has_memory_attributes(kvm,start,end,~0,attributes))+if(kvm_range_has_vm_memory_attributes(kvm,start,end,~0,attributes))gotoout_unlock;/*
@@ -2607,7 +2607,7 @@ static int kvm_vm_ioctl_set_mem_attributes(struct kvm *kvm,/* flags is currently not used. */if(attrs->flags)return-EINVAL;-if(attrs->attributes&~kvm_supported_mem_attributes(kvm))+if(attrs->attributes&~kvm_supported_vm_mem_attributes(kvm))return-EINVAL;if(attrs->size==0||attrs->address+attrs->size<attrs->address)return-EINVAL;
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Introduce a generic kvm_is_private_gfn() interface using a static call to
determine if a GFN is private. This allows the implementation for checking
a GFN's private/shared status to be set at runtime.
In preparation for choosing implementations between a guest_memfd lookup
and the existing VM attribute lookup, rename the existing
VM-attribute-based check to kvm_vm_is_private_gfn() to emphasize that it
looks up VM attributes.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Binbin Wu <redacted>
---
include/linux/kvm_host.h | 14 ++++++++++++--
virt/kvm/kvm_main.c | 15 +++++++++++++++
2 files changed, 27 insertions(+), 2 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Stub in global variable to enable in-place guest_memfd private<=>shared
memory conversion, which will eventually be exposed to userspace via a
module param, and wire up the __kvm_is_private_gfn() static call to the
guest_memfd version when in-place conversion is enabled, i.e. when gmem is
the sole authority on private vs. shared memory.
Co-developed-by: Ackerley Tng <redacted>
Signed-off-by: Ackerley Tng <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
Cc: Fuad Tabba <redacted>
Cc: Xiaoyao Li <redacted>
---
Documentation/virt/kvm/api.rst | 17 ++++++++++++-----
include/linux/kvm_host.h | 6 ++++++
virt/kvm/guest_memfd.c | 26 ++++++++++++++++++++++++++
virt/kvm/kvm_main.c | 12 +++++++++++-
4 files changed, 55 insertions(+), 6 deletions(-)
@@ -6381,11 +6381,16 @@ mapping for userspace_addr is not required to be valid/populated at the time of KVM_SET_USER_MEMORY_REGION2, e.g. shared memory can be lazily mapped/allocated on-demand.-When mapping a gfn into the guest, KVM selects shared vs. private, i.e consumes-userspace_addr vs. guest_memfd, based on the gfn's KVM_MEMORY_ATTRIBUTE_PRIVATE-state. At VM creation time, all memory is shared, i.e. the PRIVATE attribute-is '0' for all gfns. Userspace can control whether memory is shared/private by-toggling KVM_MEMORY_ATTRIBUTE_PRIVATE via KVM_SET_MEMORY_ATTRIBUTES as needed.+When mapping a gfn into the guest, KVM selects shared vs. private, i.e. consumes+userspace_addr vs. guest_memfd, based on the state in guest_memfd, which is the+sole authority on private vs. shared memory. See :ref:`KVM_CREATE_GUEST_MEMFD`+to find out more about the creation-time shared/private status.++If in-place conversion is disabled, KVM selects shared vs. private based on the+gfn's KVM_MEMORY_ATTRIBUTE_PRIVATE state. At VM creation time, all memory is+shared, i.e. the PRIVATE attribute is '0' for all gfns. Userspace can control+whether memory is shared/private by toggling KVM_MEMORY_ATTRIBUTE_PRIVATE via+KVM_SET_MEMORY_ATTRIBUTES as needed. S390: ^^^^^
@@ -6429,6 +6434,8 @@ the state of a gfn/page as needed. The "flags" field is reserved for future extensions and must be '0'.+.._KVM_CREATE_GUEST_MEMFD:+ 4.142 KVM_CREATE_GUEST_MEMFD ----------------------------
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Ackerley Tng <redacted>
If a guest_memfd memslot is created but the guest_memfd does not have the
GUEST_MEMFD_FLAG_MMAP set, KVM still fulfils guest faults by looking up the
memslot's userspace_addr.
Set KVM_MEMSLOT_GMEM_ONLY if in-place conversion is enabled so that the
guest_memfd's memory will be used for both shared and private memory. With
in-place conversion, guest_memfd will be the only backing memory for the
memslot.
No validation is performed to require userspace_addr to be a mapping from
the associated guest_memfd because even after validation, userspace is free
to remap something else at the provided userspace_addr.
userspace_addr will still be used by functions like kvm_read_guest(), and
if userspace_addr does not match up with the corresponding memory in the
memslot's guest_memfd (whether userspace_addr points to the wrong offset or
some non-guest_memfd memory, etc), that is a user error.
Requiring both shared and private memory to come from the only associated
guest_memfd simplifies invalidation in stage 2 page tables. On a PUNCH_HOLE
operation on a guest_memfd, the invalidation is now guaranteed to be
invalidating only memory mapped from the given guest_memfd.
Suggested-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
Documentation/virt/kvm/api.rst | 22 ++++++++++++++--------
virt/kvm/guest_memfd.c | 2 +-
2 files changed, 15 insertions(+), 9 deletions(-)
@@ -6381,10 +6381,16 @@ mapping for userspace_addr is not required to be valid/populated at the time of KVM_SET_USER_MEMORY_REGION2, e.g. shared memory can be lazily mapped/allocated on-demand.-When mapping a gfn into the guest, KVM selects shared vs. private, i.e. consumes-userspace_addr vs. guest_memfd, based on the state in guest_memfd, which is the-sole authority on private vs. shared memory. See :ref:`KVM_CREATE_GUEST_MEMFD`-to find out more about the creation-time shared/private status.+When mapping a gfn into the guest, guest faults are always serviced from+guest_memfd regardless of whether memory is shared or private. KVM determines+shared vs. private based on the state in guest_memfd, which is the sole+authority on private vs. shared memory. See :ref:`KVM_CREATE_GUEST_MEMFD` to+find out more about the creation-time shared/private status.++userspace_addr is expected to be the mmap()-ed address corresponding to the+right offset within the guest_memfd. Any mismatch between userspace_addr and+guest_memfd is not validated and is a user error. userspace_addr is only used+for host-side guest accesses such as kvm_read_guest(). If in-place conversion is disabled, KVM selects shared vs. private based on the gfn's KVM_MEMORY_ATTRIBUTE_PRIVATE state. At VM creation time, all memory is
@@ -6490,10 +6496,10 @@ specified via KVM_CREATE_GUEST_MEMFD. Currently defined flags: page tables. Private memory cannot. ============================ ================================================-When the KVM MMU performs a PFN lookup to service a guest fault and the backing-guest_memfd has the GUEST_MEMFD_FLAG_MMAP set, then the fault will always be-consumed from guest_memfd, regardless of whether it is a shared or a private-fault.+When the KVM MMU performs a PFN lookup to service a guest fault, the fault will+always be consumed from guest_memfd, regardless of whether it is a shared or a+private fault (unless in-place conversion is disabled and the backing+guest_memfd does not have the GUEST_MEMFD_FLAG_MMAP flag set). See KVM_SET_USER_MEMORY_REGION2 for additional details.
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
Move the kvm_arch_has_private_mem() stub and a few guest_memfd function
definitions/declarations "down" in kvm_host.h to utilize existing #ifdefs,
and so that related code is clustered together.
No functional change intended.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
include/linux/kvm_host.h | 37 ++++++++++++++++---------------------
1 file changed, 16 insertions(+), 21 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:37
From: Sean Christopherson <seanjc@google.com>
When removing one or more folios from a guest_memfd instance, invalidate
both SHARED and PRIVATE mappings if in-place conversion is enabled, because
stating the obvious, KVM needs to ensure that all mappings to the folio(s)
are dropped.
Opportunistically rename the helper to capture that it returns a filter for
all gfns in anticipation of zapping only the previous mapping types on
conversion. I.e. when doing in-place conversion to PRIVATE, only SHARED
mappings need to be zapped (ignoring that KVM would ideally not invalidate
ranges whose attributes aren't changing in the first place).
Note, precisely zapping only the possible mapping types when in-place
conversion is disabled is important for functional correctness, not just
for performance. Specifically, if KVM zaps both when SHARED vs. PRIVATE is
tracked per-VM, then a PUNCH_HOLE operation on a PRIVATE guest_memfd will
incorrectly zap SHARED mappings that have nothing to do with that gmem
instance (because they're mapped via a VMA, not a gmem fd).
The incorrect over-zapping of SHARED memory that doesn't belong to the gmem
fd requesting the zapping will be resolved in a later patch, where, if
in-place conversion is enabled, KVM will use both shared and private memory
from the guest_memfd. If both shared and private memory are from the
guest_memfd, invalidation will only zap memory belonging to the given gmem
instance.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
virt/kvm/guest_memfd.c | 11 ++++++-----
1 file changed, 6 insertions(+), 5 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:38
From: Ackerley Tng <redacted>
Accept the mapping type filter as a parameter in the invalidation start
helper instead of querying it internally. This allows callers to specify
which mappings (shared, private, or both) should be invalidated.
In the next patch, the conversion process will use this new parameter to
invalidate mappings only when they're different from the target state of
the conversion, i.e. invalidate only shared mappings on a shared to private
conversion and not both.
No functional change intended.
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
virt/kvm/guest_memfd.c | 14 +++++++++-----
1 file changed, 9 insertions(+), 5 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:38
From: Ackerley Tng <redacted>
Add a new ioctl (and matching struct), KVM_SET_MEMORY_ATTRIBUTES2, using
the same base ioctl number (0xd2), but with R/W semantics for the kernel
instead of just read semantics. "Officially" documenting that KVM writes
to the payload will allow KVM to support partial/incremental conversions,
instead of all-or-nothing updates (which requires complex unwinding), by
recording the failing offset if an error occurs.
Opportunistically add a new struct as well, even though KVM could squeeze
the error offset into "struct kvm_memory_attributes", as there's no cost to
doing so in practice. Pad the struct with a pile of extra space to try and
avoid ending up with "struct kvm_memory_attributes3" in the future. Use
the same layout for the fields common to version 1 of the struct, e.g. to
ease upgrading userspace, and to provide flexibility if KVM ever adds
support for KVM_SET_MEMORY_ATTRIBUTES2 at VM scope.
Introduce KVM_CAP_GUEST_MEMFD_MEMORY_ATTRIBUTES to advertise the
availability of the KVM_SET_MEMORY_ATTRIBUTES2 ioctl.
Update the KVM API documentation to define the new ioctl and its behavior,
and add the necessary UAPI definitions and capability checks.
The process of setting memory attributes has a clear point of no return
because, for CoCo VMs, zapping stage 2 page tables is a destructive
operation. Unlike regular VMs, where re-faulting pages into the stage 2
page tables merely incurs a performance penalty, CoCo guests must
(re-):accept pages after every fault. To preserve CoCo security guarantees,
guests will not accept pages they did not explicitly request faults
for. Consequently, during memory conversions, any operation that could
cause the process to abort must be completed before the stage 2 page tables
are zapped.
Zap only the ranges that are not already in the requested state to avoid
inadvertently destroying (CoCo) data. ARM CCA guests will try to mark the
entire DRAM as private at boot. If there are no shared pages at all, the
to-private conversion can be skipped, but the existence of a single shared
page would require the conversion process to proceed, and if it proceeds,
zapping both shared and private pages would destroy data and break the
guest.
Suggested-by: Michael Roth <redacted>
Suggested-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Co-developed-by: Vishal Annapurve <redacted>
Signed-off-by: Vishal Annapurve <redacted>
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Documentation/virt/kvm/api.rst | 71 +++++++++++++++++++++++-
include/uapi/linux/kvm.h | 15 ++++++
virt/kvm/guest_memfd.c | 119 +++++++++++++++++++++++++++++++++++++++++
virt/kvm/kvm_main.c | 23 +++++---
4 files changed, 219 insertions(+), 9 deletions(-)
@@ -117,7 +117,7 @@ description: x86 includes both i386 and x86_64. Type:- system, vm, or vcpu.+ system, vm, vcpu or guest_memfd. Parameters: what parameters are accepted by the ioctl.
@@ -6385,7 +6385,9 @@ When mapping a gfn into the guest, guest faults are always serviced from guest_memfd regardless of whether memory is shared or private. KVM determines shared vs. private based on the state in guest_memfd, which is the sole authority on private vs. shared memory. See :ref:`KVM_CREATE_GUEST_MEMFD` to-find out more about the creation-time shared/private status.+find out more about the creation-time shared/private status. Userspace can+control whether memory is shared/private by toggling+KVM_MEMORY_ATTRIBUTE_PRIVATE via :ref:`KVM_SET_MEMORY_ATTRIBUTES2` as needed. userspace_addr is expected to be the mmap()-ed address corresponding to the right offset within the guest_memfd. Any mismatch between userspace_addr and
@@ -6404,6 +6406,8 @@ S390: Returns -EINVAL if the VM has the KVM_VM_S390_UCONTROL flag set. Returns -EINVAL if called on a protected VM.+.._KVM_SET_MEMORY_ATTRIBUTES:+ 4.141 KVM_SET_MEMORY_ATTRIBUTES -------------------------------
@@ -6440,6 +6444,8 @@ the state of a gfn/page as needed. The "flags" field is reserved for future extensions and must be '0'.+See also: :ref:`KVM_SET_MEMORY_ATTRIBUTES2`.+.._KVM_CREATE_GUEST_MEMFD: 4.142 KVM_CREATE_GUEST_MEMFD
@@ -6676,6 +6682,67 @@ significant bit): Userspace should use the defined constants from ``<linux/kvm.h>`` rather than hardcoding bit positions.+.._KVM_SET_MEMORY_ATTRIBUTES2:++4.146 KVM_SET_MEMORY_ATTRIBUTES2+---------------------------------++:Capability: KVM_CAP_GUEST_MEMFD_MEMORY_ATTRIBUTES+:Architectures: all+:Type: guest_memfd ioctl+:Parameters: struct kvm_memory_attributes2 (in)+:Returns: 0 on success, <0 on error++Errors:++ ========== ===============================================================+ EINVAL The specified `offset` or `size` was invalid (e.g. not+ page aligned, causes an overflow, or size is zero).+ EFAULT The parameter address was invalid.+ ENOMEM Ran out of memory trying to track private/shared state+ ========== ===============================================================++KVM_SET_MEMORY_ATTRIBUTES2 is an extension to+KVM_SET_MEMORY_ATTRIBUTES that supports returning (writing) values to+userspace. The original (pre-extension) fields are shared with+KVM_SET_MEMORY_ATTRIBUTES identically.++Attribute values are shared with KVM_SET_MEMORY_ATTRIBUTES.++::++ struct kvm_memory_attributes2 {+ union {+ __u64 address;+ __u64 offset;+ };+ __u64 size;+ __u64 attributes;+ __u64 flags;+ __u64 reserved[12];+ };++ #define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3)++Set attributes for a range of offsets within a guest_memfd to+KVM_MEMORY_ATTRIBUTE_PRIVATE to limit the specified guest_memfd backed+memory range for guest use. Even if KVM_CAP_GUEST_MEMFD_MMAP is+supported, after a successful call to set+KVM_MEMORY_ATTRIBUTE_PRIVATE, the requested range will not be mappable+into host userspace and will only be mappable by the guest.++To allow the range to be mappable into host userspace again, call+KVM_SET_MEMORY_ATTRIBUTES2 on the guest_memfd again with+KVM_MEMORY_ATTRIBUTE_PRIVATE unset.++KVM does not directly manipulate the memory contents of pages during+attribute updates. However, the process of setting these attributes,+which includes operations such as unmapping pages from the host or+stage-2 page tables, may result in side effects on memory contents+that vary across different trusted firmware implementations.++See also: :ref:`KVM_SET_MEMORY_ATTRIBUTES`.+.._kvm_run:5. The kvm_run structure
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:38
From: Ackerley Tng <redacted>
When converting memory to private in guest_memfd, it is necessary to ensure
that the pages are not currently being accessed by any other part of the
kernel or userspace to avoid any current user writing to guest private
memory.
guest_memfd checks for any outstanding references to determine whether a
page is still in use. The only expected references after unmapping the
range requested for conversion are those that are held by guest_memfd
itself.
A folio will have outstanding references if it is present in a per-CPU
lru_add fbatch. guest_memfd does not actually participate in LRU, but
freshly-allocated folios are still added to the lru_add fbatch for batch
LRU statistics processing.
A folio may also have extra refcounts if it is on the mlock fbatch.
These two known usages of the folio are handled by draining both the
lru_add and mlock fbatches for the folio. After draining, if the refcount
is still elevated, then there are truly outstanding references.
If the page may be DMA-pinned, DMA is using it and hence there are
outstanding references. Checking if the page may be DMA-pinned can have
false positives, but that is only with a significant number of refcounts,
at which point draining LRU is not going to move the needle - it can still
be concluded that the folio has outstanding references.
If the page is still mapped after guest_memfd tried to unmap it earlier in
the conversion process, it also has outstanding references.
Exit early to avoid unnecessary draining in these two cases. If an
outstanding reference is detected, stop scanning, record the failing
offset, and return immediately.
Track the drain status to avoid repeatedly draining LRU caches across
multiple folios while scanning the requested range.
Update the kvm_memory_attributes2 structure to include an error_offset
field. This allows KVM to report the exact offset where a conversion
failed. If the safety check fails, return -EAGAIN and copy the error_offset
back to userspace so that it can potentially retry the operation or handle
the failure gracefully.
Report error_offset if preallocating memory to track attributes fails with
-ENOMEM as well, in which case the start offset of the range is returned.
Update documentation to document the error_offset field and the possible
-EAGAIN error.
Suggested-by: David Hildenbrand <david@kernel.org>
Co-developed-by: Vishal Annapurve <redacted>
Signed-off-by: Vishal Annapurve <redacted>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
---
Documentation/virt/kvm/api.rst | 10 +++++
include/uapi/linux/kvm.h | 3 +-
virt/kvm/guest_memfd.c | 90 +++++++++++++++++++++++++++++++++++++++---
3 files changed, 97 insertions(+), 6 deletions(-)
@@ -6741,6 +6741,16 @@ which includes operations such as unmapping pages from the host or stage-2 page tables, may result in side effects on memory contents that vary across different trusted firmware implementations.+If this ioctl returns -EAGAIN, the offset of the page with unexpected+refcounts will be returned in ``error_offset``. This can occur if+there are transient refcounts on the pages, taken by other parts of+the kernel.++Userspace is expected to figure out how to remove all known refcounts+on the shared pages, such as refcounts taken by get_user_pages(), and+try the ioctl again. A possible source of these long term refcounts is+if the guest_memfd memory was pinned in IOMMU page tables.+ See also: :ref:`KVM_SET_MEMORY_ATTRIBUTES`..._kvm_run:
@@ -538,8 +539,58 @@ static int kvm_gmem_mas_preallocate(struct ma_state *mas, u64 attributes,returnmas_preallocate(mas,xa_mk_value(attributes),GFP_KERNEL);}+staticbool__folio_has_outstanding_references(structfolio*folio,+enumlru_cache_drained*drained)+{+if(folio_maybe_dma_pinned(folio)||folio_mapped(folio))+returntrue;++/* 1 reference held by filemap_get_folios() in the folio batch. */+lru_cache_drain_for_folio(folio,1,drained);++/*+*Outstandingreferencesareanythingotherthanthosefromthepage+*cache,plus1temporaryreferenceheldbyfilemap_get_folios()inthe+*foliobatch.+*/+returnfolio_ref_count(folio)!=folio_nr_pages(folio)+1;+}++staticboolkvm_gmem_has_outstanding_references(structinode*inode,+pgoff_tstart,size_tnr_pages,+pgoff_t*err_index)+{+enumlru_cache_draineddrained=LRU_CACHE_NOT_DRAINED;+structaddress_space*mapping=inode->i_mapping;+pgoff_tlast=start+nr_pages-1;+structfolio_batchfbatch;+pgoff_tnext;+inti;++folio_batch_init(&fbatch);++next=start;+while(filemap_get_folios(mapping,&next,last,&fbatch)){+for(i=0;i<folio_batch_count(&fbatch);++i){+structfolio*folio=fbatch.folios[i];++if(__folio_has_outstanding_references(folio,&drained)){+*err_index=max(start,folio->index);+folio_batch_release(&fbatch);+returntrue;+}+}++folio_batch_release(&fbatch);+cond_resched();+}++returnfalse;+}+staticint__kvm_gmem_set_attributes(structinode*inode,pgoff_tstart,-size_tnr_pages,uint64_tattrs)+size_tnr_pages,uint64_tattrs,+pgoff_t*err_index){boolto_private=attrs&KVM_MEMORY_ATTRIBUTE_PRIVATE;structaddress_space*mapping=inode->i_mapping;
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:38
From: Ackerley Tng <redacted>
When doing in-place conversion from PRIVATE to SHARED, immediately inform
arch code of the conversion for all allocated pages/folios, e.g. so that
arch code can put hardware metadata tables in the correct state. Eagerly
updating the table for to SHARED conversions avoids having to implement
on-demand updates, e.g. when faulting in host userspace mappings. Skip the
entire flow if the arch doesn't implement conversion callbacks, as getting
folios from the filemap is noticeably expensive, especially when converting
large chunks of memory.
Deliberately don't eagerly update the metadata table on conversions from
SHARED to PRIVATE, because assigning a page to a VM (versus "returning" it
to the host) requires the exact GFN associated with the page, i.e would
require walking the memslot bindings. And because KVM *must* do on-demand
metadata updates when getting a PFN for KVM-internal usage, as that's the
only time a relevant memslot binding is guaranteed to exist.
Note! Inform arch code of the conversion within the protection of the
invalidation sequence, to ensure that any existing mappings are dropped
before hardware is updated, and to ensure that new mappings can't be
established until after the conversion is complete.
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
---
arch/x86/include/asm/kvm-x86-ops.h | 2 +-
arch/x86/include/asm/kvm_host.h | 2 +-
arch/x86/kvm/x86.c | 5 +++++
include/linux/kvm_host.h | 1 +
virt/kvm/guest_memfd.c | 42 ++++++++++++++++++++++++++++++++++++++
5 files changed, 50 insertions(+), 2 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:38
From: Ackerley Tng <redacted>
Move the folio initialization logic from kvm_gmem_get_pfn() into
__kvm_gmem_get_pfn() to also zero pages if the page is to be used in
kvm_gmem_populate().
With in-place conversion, the existing data in a guest_memfd page can be
populated into guest memory through platform-specific ioctls.
Without first zeroing the page obtained using __kvm_gmem_get_pfn(), it
might contain uninitialized host memory, which would leak to the guest if
the populate completes.
guest_memfd pages are zeroed at most once in the page's entire lifetime
with guest_memfd, and that is tracked using the uptodate flag.
Zeroing the page in __kvm_gmem_get_pfn() is chosen over zeroing in
kvm_gmem_get_folio() since other flows, such as a future write() syscall,
can get a page, write to the page and then set page uptodate without
zeroing.
There may be some performance penalty due to redundant zeroing, but this
would pale in comparison to the cost of actually assigning the page to the
VM.
This aligns with the concept of zeroing before first use - the other place
where zeroing happens is in kvm_gmem_fault_user_mapping().
On populate failure, the page is not re-zeroed, since on SNP, if firmware
rejects a CPUID page, the expected CPUID values provided by firmware are
returned to userspace via page contents. More generally, page contents may
be modified on populate failure.
Don't mark the page uptodate again after populating, since the page would
already be marked uptodate before the post_populate() call.
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
virt/kvm/guest_memfd.c | 12 +++++-------
1 file changed, 5 insertions(+), 7 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:38
From: Ackerley Tng <redacted>
Provide a function to check that a range has given attributes.
Optimize setting memory attributes by returning early if all pages in the
requested range already have the requested attributes.
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
virt/kvm/guest_memfd.c | 23 ++++++++++++++++++++++-
1 file changed, 22 insertions(+), 1 deletion(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Sean Christopherson <seanjc@google.com>
Accept gmem_flags in vm_mem_add() to be able to create a guest_memfd within
vm_mem_add().
When vm_mem_add() is used to set up a guest_memfd for a memslot, set up the
provided (or created) gmem_fd as the fd for the user memory region. This
makes it available to be mmap()-ed from just like fds from other memory
sources.
For guest_memfds, mmap() using gmem_offset instead of 0 all the time.
Always use MAP_SHARED if mmap-ing from guest_memfd instead of reading flag
from the configured src_type, which doesn't include guest_memfd.
Add a kvm_slot_to_fd() helper to provide convenient access to the file
descriptor of a memslot.
Update existing callers of vm_mem_add() to pass 0 for gmem_flags to
preserve existing behavior.
Co-developed-by: Ackerley Tng <redacted>
Signed-off-by: Ackerley Tng <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/include/kvm_util.h | 7 +++++-
tools/testing/selftests/kvm/lib/kvm_util.c | 28 ++++++++++++----------
.../kvm/x86/private_mem_conversions_test.c | 2 +-
3 files changed, 23 insertions(+), 14 deletions(-)
@@ -1002,12 +1002,14 @@ void vm_set_user_memory_region2(struct kvm_vm *vm, u32 slot, u32 flags,/* FIXME: This thing needs to be ripped apart and rewritten. */voidvm_mem_add(structkvm_vm*vm,enumvm_mem_backing_src_typesrc_type,gpa_tgpa,u32slot,u64npages,u32flags,-intgmem_fd,u64gmem_offset)+intgmem_fd,u64gmem_offset,u64gmem_flags){intret;structuserspace_mem_region*region;size_tbacking_src_pagesz=get_backing_src_pagesz(src_type);+intmmap_flags=vm_mem_backing_src_alias(src_type)->flag;size_tmem_size=npages*vm->page_size;+off_tmmap_offset=0;size_talignment=1;TEST_REQUIRE_SET_USER_MEMORY_REGION2();
@@ -1079,8 +1081,6 @@ void vm_mem_add(struct kvm_vm *vm, enum vm_mem_backing_src_type src_type,if(flags&KVM_MEM_GUEST_MEMFD){if(gmem_fd<0){-u32gmem_flags=0;-TEST_ASSERT(!gmem_offset,"Offset must be zero when creating new guest_memfd");gmem_fd=vm_create_guest_memfd(vm,mem_size,gmem_flags);
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Sean Christopherson <seanjc@google.com>
Now that guest_memfd supports tracking private vs. shared within gmem
itself, allow userspace to specify INIT_SHARED on a guest_memfd instance
for x86 Confidential Computing (CoCo) VMs, so long as in-place conversion
is enabled, i.e. when it's actually possible for a guest_memfd instance to
contain shared memory.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Binbin Wu <redacted>
---
arch/x86/kvm/x86.c | 13 +++++++------
1 file changed, 7 insertions(+), 6 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Ackerley Tng <redacted>
Update tdx_gmem_post_populate() to handle cases where userspace requests
"no source page". To handle "no source page", populate (perform
TDH.MEM.PAGE.ADD) using memory in-place at the target PFN.
Allow "no source page" only when gmem_in_place_conversion is enabled,
because retroactively adding support for out-of-place conversion would mean
requiring a userspace update for a feature that's being deprecated.
Also, KVM supporting "no source page" without in-place conversion would
effectively be an obscure zero-page optimization that relies on the page
being zeroed when it is allocated by guest_memfd.
Rejecting "no source page" without in-place conversion scenario is valuable
for KVM developers since it helps newcomers understand what exactly is and
isn't possible.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Tested-by: Yan Zhao <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: Yan Zhao <redacted>
Reviewed-by: Xiaoyao Li <redacted>
---
Documentation/virt/kvm/x86/intel-tdx.rst | 4 ++++
arch/x86/kvm/vmx/tdx.c | 8 +++++---
2 files changed, 9 insertions(+), 3 deletions(-)
@@ -158,6 +158,10 @@ KVM_TDX_INIT_MEM_REGION Initialize @nr_pages TDX guest private memory starting from @gpa with userspace provided data from @source_addr. @source_addr must be PAGE_SIZE-aligned.+If guest_memfd in-place conversion is enabled, pass 0 for @source_addr+to represent "no source page". A source page is required if in-place+conversion is not enabled or not supported.+ Note, before calling this sub command, memory attribute of the range [gpa, gpa + nr_pages] needs to be private. Userspace can use KVM_SET_MEMORY_ATTRIBUTES to set the attribute.
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Sean Christopherson <seanjc@google.com>
Rename local variables and function parameters for the guest memory file
descriptor and its offset to use a "gmem_" prefix instead of
"guest_memfd_".
No functional change intended.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/include/kvm_util.h | 6 +++---
tools/testing/selftests/kvm/lib/kvm_util.c | 26 +++++++++++++-------------
2 files changed, 16 insertions(+), 16 deletions(-)
@@ -1002,7 +1002,7 @@ void vm_set_user_memory_region2(struct kvm_vm *vm, u32 slot, u32 flags,/* FIXME: This thing needs to be ripped apart and rewritten. */voidvm_mem_add(structkvm_vm*vm,enumvm_mem_backing_src_typesrc_type,gpa_tgpa,u32slot,u64npages,u32flags,-intguest_memfd,u64guest_memfd_offset)+intgmem_fd,u64gmem_offset){intret;structuserspace_mem_region*region;
@@ -1078,12 +1078,12 @@ void vm_mem_add(struct kvm_vm *vm, enum vm_mem_backing_src_type src_type,region->mmap_size+=alignment;if(flags&KVM_MEM_GUEST_MEMFD){-if(guest_memfd<0){-u32guest_memfd_flags=0;+if(gmem_fd<0){+u32gmem_flags=0;-TEST_ASSERT(!guest_memfd_offset,+TEST_ASSERT(!gmem_offset,"Offset must be zero when creating new guest_memfd");-guest_memfd=vm_create_guest_memfd(vm,mem_size,guest_memfd_flags);+gmem_fd=vm_create_guest_memfd(vm,mem_size,gmem_flags);}else{/**Installauniquefdforeachmemslotsothatthefd
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Sean Christopherson <seanjc@google.com>
Add a global variable, kvm_has_gmem_attributes, to make the result of
checking for KVM_CAP_GUEST_MEMFD_MEMORY_ATTRIBUTES available to all tests.
kvm_has_gmem_attributes is true if guest_memfd tracks memory attributes, as
opposed to VM-level tracking.
This global variable is synced to the guest for testing convenience, to
avoid introducing subtle bugs when host/guest state is desynced.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/include/test_util.h | 2 ++
tools/testing/selftests/kvm/lib/kvm_util.c | 5 +++++
2 files changed, 7 insertions(+)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Sean Christopherson <seanjc@google.com>
When adding a memslot associated with a guest_memfd instance, create/dup
the guest_memfd before creating the "normal" backing file. This will allow
dup'ing the gmem fd as the normal fd when guest_memfd supports mmap(),
i.e. to make guest_memfd the _only_ backing source for the memslot.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/lib/kvm_util.c | 45 +++++++++++++++---------------
1 file changed, 23 insertions(+), 22 deletions(-)
@@ -1077,6 +1077,29 @@ void vm_mem_add(struct kvm_vm *vm, enum vm_mem_backing_src_type src_type,if(alignment>1)region->mmap_size+=alignment;+if(flags&KVM_MEM_GUEST_MEMFD){+if(guest_memfd<0){+u32guest_memfd_flags=0;++TEST_ASSERT(!guest_memfd_offset,+"Offset must be zero when creating new guest_memfd");+guest_memfd=vm_create_guest_memfd(vm,mem_size,guest_memfd_flags);+}else{+/*+*Installauniquefdforeachmemslotsothatthefd+*canbeclosedwhentheregionisdeletedwithout+*needingtotrackifthefdisownedbytheframework+*orbythecaller.+*/+guest_memfd=kvm_dup(guest_memfd);+}++region->region.guest_memfd=guest_memfd;+region->region.guest_memfd_offset=guest_memfd_offset;+}else{+region->region.guest_memfd=-1;+}+region->fd=-1;if(backing_src_is_shared(src_type))region->fd=kvm_memfd_alloc(region->mmap_size,
@@ -1106,28 +1129,6 @@ void vm_mem_add(struct kvm_vm *vm, enum vm_mem_backing_src_type src_type,region->backing_src_type=src_type;-if(flags&KVM_MEM_GUEST_MEMFD){-if(guest_memfd<0){-u32guest_memfd_flags=0;-TEST_ASSERT(!guest_memfd_offset,-"Offset must be zero when creating new guest_memfd");-guest_memfd=vm_create_guest_memfd(vm,mem_size,guest_memfd_flags);-}else{-/*-*Installauniquefdforeachmemslotsothatthefd-*canbeclosedwhentheregionisdeletedwithout-*needingtotrackifthefdisownedbytheframework-*orbythecaller.-*/-guest_memfd=kvm_dup(guest_memfd);-}--region->region.guest_memfd=guest_memfd;-region->region.guest_memfd_offset=guest_memfd_offset;-}else{-region->region.guest_memfd=-1;-}-region->unused_phy_pages=sparsebit_alloc();if(vm_arch_has_protected_memory(vm))region->protected_phy_pages=sparsebit_alloc();
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Ackerley Tng <redacted>
Add a selftest for the guest_memfd memory attribute conversion ioctls.
The test starts the guest_memfd as all-private (the default state), and
verifies the basic flow of converting a single page to shared and then back
to private.
Add infrastructure that supports extensions to other conversion flow
tests. This infrastructure will be used in upcoming patches for other
conversion tests.
Add test as an x86-specific test since guest_memfd's testing
vehicle (KVM_X86_SW_PROTECTED_VM) is x86-specific.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/Makefile.kvm | 1 +
.../kvm/x86/guest_memfd_conversions_test.c | 200 +++++++++++++++++++++
2 files changed, 201 insertions(+)
@@ -0,0 +1,200 @@+// SPDX-License-Identifier: GPL-2.0-only+/*+*Copyright(c)2024,GoogleLLC.+*/+#include<sys/mman.h>+#include<unistd.h>++#include<linux/align.h>+#include<linux/kvm.h>+#include<linux/sizes.h>++#include"kvm_util.h"+#include"kselftest_harness.h"+#include"test_util.h"+#include"ucall_common.h"++FIXTURE(gmem_conversions){+structkvm_vcpu*vcpu;+intgmem_fd;+/* HVA of the first byte of the memory mmap()-ed from gmem_fd. */+char*mem;+};++typedefFIXTURE_DATA(gmem_conversions)test_data_t;++FIXTURE_SETUP(gmem_conversions){}++staticsize_tpage_size;++staticvoidguest_do_rmw(void);+#define GUEST_MEMFD_SHARING_TEST_GVA 0x90000000ULL++/*+*Defersetupuntiltheindividualtestisinvokedsothattestscanspecify+*thenumberofpagesandflagsfortheguest_memfdinstance.+*/+staticvoidgmem_conversions_do_setup(test_data_t*t,intnr_pages,+intgmem_flags)+{+conststructvm_shapeshape={+.mode=VM_MODE_DEFAULT,+.type=KVM_X86_SW_PROTECTED_VM,+};+/*+*UsehighGPAaboveAPIC_DEFAULT_PHYS_BASEtoavoidclashingwith+*APIC_DEFAULT_PHYS_BASE.+*/+constgpa_tgpa=SZ_4G;+constu32slot=1;+structkvm_vm*vm;++vm=__vm_create_shape_with_one_vcpu(shape,&t->vcpu,nr_pages,guest_do_rmw);++vm_mem_add(vm,VM_MEM_SRC_SHMEM,gpa,slot,nr_pages,+KVM_MEM_GUEST_MEMFD,-1,0,gmem_flags);++t->gmem_fd=kvm_slot_to_fd(vm,slot);+t->mem=addr_gpa2hva(vm,gpa);+virt_map(vm,GUEST_MEMFD_SHARING_TEST_GVA,gpa,nr_pages);+}++staticvoidgmem_conversions_do_teardown(test_data_t*t)+{+/* No need to close gmem_fd, it's owned by the VM structure. */+kvm_vm_free(t->vcpu->vm);+}++FIXTURE_TEARDOWN(gmem_conversions)+{+gmem_conversions_do_teardown(self);+}++/*+*Inthesetestdefinitionmacros,__nr_pagesandnr_pagesisusedtosetup+*thetotalnumberofpagesintheguest_memfdundertest.Thiswillbe+*availableinthetestdefinitionsasnr_pages.+*/++#define __GMEM_CONVERSION_TEST(test, __nr_pages, flags) \+staticvoid__gmem_conversions_##test(test_data_t*t,intnr_pages);\+\+TEST_F(gmem_conversions,test)\+{\+constintnr=(__nr_pages);\+\+gmem_conversions_do_setup(self,nr,flags);\+__gmem_conversions_##test(self,nr);\+}\+staticvoid__gmem_conversions_##test(test_data_t*t,intnr_pages)\++#define GMEM_CONVERSION_TEST(test, __nr_pages, flags) \+__GMEM_CONVERSION_TEST(test,__nr_pages,(flags)|GUEST_MEMFD_FLAG_MMAP)++#define __GMEM_CONVERSION_TEST_INIT_PRIVATE(test, __nr_pages) \+GMEM_CONVERSION_TEST(test,__nr_pages,0)++#define GMEM_CONVERSION_TEST_INIT_PRIVATE(test) \+__GMEM_CONVERSION_TEST_INIT_PRIVATE(test,1)++structguest_check_data{+void*mem;+charexpected_val;+charwrite_val;+};+staticstructguest_check_dataguest_data;++staticvoidguest_do_rmw(void)+{+for(;;){+char*mem=READ_ONCE(guest_data.mem);++GUEST_ASSERT_EQ(READ_ONCE(*mem),READ_ONCE(guest_data.expected_val));+WRITE_ONCE(*mem,READ_ONCE(guest_data.write_val));++GUEST_SYNC(0);+}+}++staticvoidrun_guest_do_rmw(structkvm_vcpu*vcpu,u64pgoff,+charexpected_val,charwrite_val)+{+structucalluc;+intr;++guest_data.mem=(void*)GUEST_MEMFD_SHARING_TEST_GVA+pgoff*page_size;+guest_data.expected_val=expected_val;+guest_data.write_val=write_val;+sync_global_to_guest(vcpu->vm,guest_data);++do{+r=__vcpu_run(vcpu);+}while(r==-1&&errno==EINTR);++TEST_ASSERT_EQ(r,0);++switch(get_ucall(vcpu,&uc)){+caseUCALL_ABORT:+REPORT_GUEST_ASSERT(uc);+caseUCALL_SYNC:+break;+default:+TEST_FAIL("Unexpected ucall %lu",uc.cmd);+}+}++staticvoidhost_do_rmw(char*mem,u64pgoff,charexpected_val,+charwrite_val)+{+TEST_ASSERT_EQ(READ_ONCE(mem[pgoff*page_size]),expected_val);+WRITE_ONCE(mem[pgoff*page_size],write_val);+}++staticvoidtest_private(test_data_t*t,u64pgoff,charstarting_val,+charwrite_val)+{+TEST_EXPECT_SIGBUS(WRITE_ONCE(t->mem[pgoff*page_size],write_val));+run_guest_do_rmw(t->vcpu,pgoff,starting_val,write_val);+TEST_EXPECT_SIGBUS(READ_ONCE(t->mem[pgoff*page_size]));+}++staticvoidtest_convert_to_private(test_data_t*t,u64pgoff,+charstarting_val,charwrite_val)+{+gmem_set_private(t->gmem_fd,pgoff*page_size,page_size);+test_private(t,pgoff,starting_val,write_val);+}++staticvoidtest_shared(test_data_t*t,u64pgoff,charstarting_val,+charhost_write_val,charwrite_val)+{+host_do_rmw(t->mem,pgoff,starting_val,host_write_val);+run_guest_do_rmw(t->vcpu,pgoff,host_write_val,write_val);+TEST_ASSERT_EQ(READ_ONCE(t->mem[pgoff*page_size]),write_val);+}++staticvoidtest_convert_to_shared(test_data_t*t,u64pgoff,+charstarting_val,charhost_write_val,+charwrite_val)+{+gmem_set_shared(t->gmem_fd,pgoff*page_size,page_size);+test_shared(t,pgoff,starting_val,host_write_val,write_val);+}++GMEM_CONVERSION_TEST_INIT_PRIVATE(init_private)+{+test_private(t,0,0,'A');+test_convert_to_shared(t,0,'A','B','C');+test_convert_to_private(t,0,'C','E');+}++intmain(intargc,char*argv[])+{+TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES)&BIT(KVM_X86_SW_PROTECTED_VM));+TEST_REQUIRE(kvm_check_cap(KVM_CAP_GUEST_MEMFD_MEMORY_ATTRIBUTES)&+KVM_MEMORY_ATTRIBUTE_PRIVATE);++page_size=getpagesize();++returntest_harness_run(argc,argv);+}
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Sean Christopherson <seanjc@google.com>
Add helper functions to kvm_util.h to support calling ioctls, specifically
KVM_SET_MEMORY_ATTRIBUTES2, on a guest_memfd file descriptor.
Introduce gmem_ioctl() and __gmem_ioctl() macros, modeled after the
existing vm_ioctl() helpers, to provide a standard way to call ioctls
on a guest_memfd.
Add gmem_set_memory_attributes() and its derivatives (gmem_set_private(),
gmem_set_shared()) to set memory attributes on a guest_memfd region.
Also provide "__" variants that return the ioctl error code instead of
aborting the test. These helpers will be used by upcoming guest_memfd
tests.
To avoid code duplication, factor out the check for supported memory
attributes into a new macro, TEST_ASSERT_SUPPORTED_ATTRIBUTES, and use
it in both the existing vm_set_memory_attributes() and the new
gmem_set_memory_attributes() helpers.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/include/kvm_util.h | 94 +++++++++++++++++++++++---
1 file changed, 86 insertions(+), 8 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Ackerley Tng <redacted>
Add a test case to verify that conversions between private and shared
memory work correctly when the memory is initially created as shared.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
.../selftests/kvm/x86/guest_memfd_conversions_test.c | 13 +++++++++++++
1 file changed, 13 insertions(+)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Sean Christopherson <seanjc@google.com>
Bury KVM_VM_MEMORY_ATTRIBUTES in x86 to discourage other architectures
from adding support for per-VM memory attributes, because tracking private
vs. shared memory on a per-VM basis is now deprecated in favor of tracking
on a per-guest_memfd basis, and while RWX memory attributes are on the
horizon, they too are expected to be x86-only.
This will also allow modifying KVM_VM_MEMORY_ATTRIBUTES to be
user-selectable (in x86) without creating weirdness in KVM's Kconfigs.
Now that guest_memfd supports in-place conversions, it's entirely possible
to run x86 CoCo VMs without support for KVM_VM_MEMORY_ATTRIBUTES.
Leave the code itself in common KVM so that it's trivial to undo this
change if new per-VM attributes do come along.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
---
arch/x86/kvm/Kconfig | 3 +++
virt/kvm/Kconfig | 3 ---
2 files changed, 3 insertions(+), 3 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Sean Christopherson <seanjc@google.com>
Allow the user to disable KVM_VM_MEMORY_ATTRIBUTES even when KVM supports
PRIVATE and SHARED attributes, and expose gmem_in_place_conversion as a
module parameter when per-VM attributes are supported. I.e. let userspace
enable in-place PRIVATE<=>SHARED conversion of guest_memfd pages.
Provide both a Kconfig option and a (conditional) module param so that
deployments that use a custom kernel can fully disable per-VM tracking,
while not forcing distros to ship two separate kernels in order to provide
backwards compatibility for downstream users.
Don't allow running VMs with mixed tracking for a given instance of KVM,
i.e. disallow toggling the module param after KVM is loaded, as the extra
complexity needed to handle per-VM behavior far outweighs any potential
benefit. E.g. neither TDX nor SNP supports live migration, so in effect
the requirement is that existing deployments that want to support both the
old and the new models would need to tell their VMM which flavor of
tracking to use.
[ Xiaoyao: Define module_param only if CONFIG_KVM_VM_MEMORY_ATTRIBUTES is enabled ]
Suggested-by: Xiaoyao Li <redacted>
Co-developed-by: Ackerley Tng <redacted>
Signed-off-by: Ackerley Tng <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
Reviewed-by: Xiaoyao Li <redacted>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
---
Documentation/admin-guide/kernel-parameters.txt | 25 +++++++++++++++++++++++++
arch/x86/include/asm/kvm_host.h | 4 +++-
arch/x86/kvm/Kconfig | 14 ++++++++++----
virt/kvm/kvm_main.c | 5 ++++-
4 files changed, 42 insertions(+), 6 deletions(-)
@@ -3156,6 +3156,31 @@ Kernel parameters kvm.enable_vmware_backdoor=[KVM] Support VMware backdoor PV interface. Default is false (don't support).+ kvm.gmem_in_place_conversion=+ [KVM] Controls whether KVM enables in-place conversion+ support for guest_memfd and tracks the private/shared+ state of memory per guest_memfd instead of per VM.++ If enabled, KVM enables the KVM_SET_MEMORY_ATTRIBUTES2+ ioctl on guest_memfd file descriptors and disables the+ legacy VM-scoped KVM_SET_MEMORY_ATTRIBUTES ioctl for+ private memory state tracking. Only the+ KVM_MEMORY_ATTRIBUTE_PRIVATE attribute moves to+ per-guest_memfd tracking; other attributes remain+ per-VM.++ This parameter toggles KVM's in-place conversion+ capability support. Whether a VMM uses separate backends+ or out-of-place memory management is determined by+ userspace VMM design.++ Note, this parameter is only available when+ CONFIG_KVM_VM_MEMORY_ATTRIBUTES=y. When+ CONFIG_KVM_VM_MEMORY_ATTRIBUTES is not set, in-place+ conversion is unconditionally enabled.++ Default is N (off).+ kvm.nx_huge_pages= [KVM] Controls the software workaround for the X86_BUG_ITLB_MULTIHIT bug.
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:39
From: Michael Roth <redacted>
Make the source page for populating an SNP guest_memfd instance optional
if in-place conversion/population is enabled. If KVM can convert the page
in-place, then it's possible for guest memory to be initialized directly
from userspace by mmap()'ing the guest_memfd and writing to it while the
corresponding GPA ranges are in a 'shared' state, before converting them
to the 'private' state expected by KVM_SEV_SNP_LAUNCH_UPDATE.
Update the handling/documentation for KVM_SEV_SNP_LAUNCH_UPDATE to allow
for 'uaddr' to be set to NULL when in-place conversion is enabled, which
SNP_LAUNCH_UPDATE will then use to determine when it should/shouldn't
copy in data from a separate memory location. Continue to enforce
non-NULL when PRIVATE is tracked per-VM, not per-guest_memfd.
[ Sean: Moved condition to snp_launch_update ]
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
---
Documentation/virt/kvm/x86/amd-memory-encryption.rst | 14 ++++++++++----
arch/x86/kvm/svm/sev.c | 11 ++++++-----
virt/kvm/kvm_main.c | 1 +
3 files changed, 17 insertions(+), 9 deletions(-)
@@ -503,7 +503,8 @@ secrets. It is required that the GPA ranges initialized by this command have had the KVM_MEMORY_ATTRIBUTE_PRIVATE attribute set in advance. See the documentation-for KVM_SET_MEMORY_ATTRIBUTES for more details on this aspect.+for KVM_SET_MEMORY_ATTRIBUTES/KVM_SET_MEMORY_ATTRIBUTES2 for more details on+this aspect. Upon success, this command is not guaranteed to have processed the entire range requested. Instead, the ``gfn_start``, ``uaddr``, and ``len`` fields of
@@ -511,9 +512,14 @@ range requested. Instead, the ``gfn_start``, ``uaddr``, and ``len`` fields of remaining range that has yet to be processed. The caller should continue calling this command until those fields indicate the entire range has been processed, e.g. ``len`` is 0, ``gfn_start`` is equal to the last GFN in the-range plus 1, and ``uaddr`` is the last byte of the userspace-provided source-buffer address plus 1. In the case where ``type`` is KVM_SEV_SNP_PAGE_TYPE_ZERO,-``uaddr`` will be ignored completely.+range plus 1, and ``uaddr`` (if specified) is the last byte of the+userspace-provided source buffer address plus 1.++In the case where ``type`` is KVM_SEV_SNP_PAGE_TYPE_ZERO, ``uaddr`` will be+ignored completely. For all other page types, ``uaddr`` is optional if in-place+conversion is enabled (i.e. when the data had been written directly to+guest_memfd while the page was in the shared state) and is required if in-place+conversion is disabled. Parameters (in): struct kvm_sev_snp_launch_update
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
The TEST_EXPECT_SIGBUS macro is not thread-safe as it uses a global
sigjmp_buf and installs a global SIGBUS signal handler. If multiple threads
execute the macro concurrently, they will race on installing the signal
handler and stomp on other threads' jump buffers, leading to incorrect test
behavior.
Make TEST_EXPECT_SIGBUS thread-safe with the following changes:
Share the KVM tests' global signal handler. sigaction() applies to all
threads; without sharing a global signal handler, one thread may have
removed the signal handler that another thread added, hence leading to
unexpected signals.
The alternative of layering signal handlers was considered, but calling
sigaction() within TEST_EXPECT_SIGBUS() necessarily creates a race. To
avoid adding new setup and teardown routines to do sigaction() and keep
usage of TEST_EXPECT_SIGBUS() simple, share the KVM tests' global signal
handler.
Opportunistically rename report_unexpected_signal to
catchall_signal_handler.
To continue to only expect SIGBUS within specific regions of code, use a
thread-specific variable, expecting_sigbus, to replace installing and
removing signal handlers.
Make the execution environment for the thread, sigjmp_buf, a
thread-specific variable.
As part of TEST_EXPECT_SIGBUS(), assert the prerequisite for this setup,
that the current signal handler is the catchall_signal_handler.
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/include/test_util.h | 32 +++++++++++++------------
tools/testing/selftests/kvm/lib/kvm_util.c | 18 ++++++++++----
tools/testing/selftests/kvm/lib/test_util.c | 7 ------
3 files changed, 30 insertions(+), 27 deletions(-)
@@ -83,21 +83,23 @@ do { \__builtin_unreachable();\}while(0)-externsigjmp_bufexpect_sigbus_jmpbuf;-voidexpect_sigbus_handler(intsignum);--#define TEST_EXPECT_SIGBUS(action) \-do{\-structsigactionsa_old,sa_new={\-.sa_handler=expect_sigbus_handler,\-};\-\-sigaction(SIGBUS,&sa_new,&sa_old);\-if(sigsetjmp(expect_sigbus_jmpbuf,1)==0){\-action;\-TEST_FAIL("'%s' should have triggered SIGBUS",#action);\-}\-sigaction(SIGBUS,&sa_old,NULL);\+extern__threadsigjmp_bufexpect_sigbus_jmpbuf;+extern__threadvolatilesig_atomic_texpecting_sigbus;+voidcatchall_signal_handler(intsignum);++#define TEST_EXPECT_SIGBUS(action) \+do{\+structsigaction__sa={};\+\+TEST_ASSERT_EQ(sigaction(SIGBUS,NULL,&__sa),0);\+TEST_ASSERT_EQ(__sa.sa_handler,&catchall_signal_handler);\+\+expecting_sigbus=true;\+if(sigsetjmp(expect_sigbus_jmpbuf,1)==0){\+action;\+TEST_FAIL("'%s' should have triggered SIGBUS",#action);\+}\+expecting_sigbus=false;\}while(0)size_tparse_size(constchar*size);
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
Add a guest_memfd selftest to verify that memory conversions work
correctly with allocated folios in different layouts.
By iterating through which pages are initially faulted, the test covers
various layouts of contiguous allocated and unallocated regions, exercising
conversion with different range layouts.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
.../kvm/x86/guest_memfd_conversions_test.c | 30 ++++++++++++++++++++++
1 file changed, 30 insertions(+)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
The existing guest_memfd conversion tests only use single-page memory
regions. This provides no coverage for multi-page guest_memfd objects,
specifically whether KVM correctly handles the page index for conversion
operations. An incorrect implementation could, for example, always operate
on the first page regardless of the index provided.
Add a new test case to verify that conversions between private and shared
memory correctly target the specified page within a multi-page guest_memfd.
This test also verifies the precision of memory conversions by converting a
single page and then iterating through all other pages to ensure they
remain in their original state.
To support this test, add a new GMEM_CONVERSION_MULTIPAGE_TEST_INIT_SHARED
macro that handles setting up and tearing down the VM for each page
iteration. The teardown logic is adjusted to prevent a double-free in this
new scenario.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
.../kvm/x86/guest_memfd_conversions_test.c | 67 ++++++++++++++++++++++
1 file changed, 67 insertions(+)
@@ -61,8 +61,13 @@ static void gmem_conversions_do_setup(test_data_t *t, int nr_pages,staticvoidgmem_conversions_do_teardown(test_data_t*t){+/* Use NULL to avoid second free in FIXTURE_TEARDOWN (multipage tests). */+if(!t->vcpu)+return;+/* No need to close gmem_fd, it's owned by the VM structure. */kvm_vm_free(t->vcpu->vm);+t->vcpu=NULL;}FIXTURE_TEARDOWN(gmem_conversions)
@@ -201,6 +230,44 @@ GMEM_CONVERSION_TEST_INIT_SHARED(init_shared)test_convert_to_shared(t,0,'C','D','E');}+GMEM_CONVERSION_MULTIPAGE_TEST_INIT_SHARED(indexing,4)+{+inti;++/* Get a char that varies with both i and n. */+#define combine(x, n) (((x) << 4) + (n))+#define i_(n) (combine(i, n))+#define t_(n) (combine(test_page, n))++/*+*Startwiththehighestindex,tocatchanyerrorswhen,perhaps,the+*firstpageisreturnedevenforthelastindex.+*/+for(i=nr_pages-1;i>=0;--i)+test_shared(t,i,0,i_(0),i_(2));++test_convert_to_private(t,test_page,t_(2),t_(3));++for(i=0;i<nr_pages;++i){+if(i==test_page)+test_private(t,test_page,t_(3),t_(4));+else+test_shared(t,i,i_(2),i_(3),i_(4));+}++test_convert_to_shared(t,test_page,t_(4),t_(5),t_(6));++for(i=0;i<nr_pages;++i){+charexpected=i==test_page?t_(6):i_(4);++test_shared(t,i,expected,i_(7),i_(8));+}++#undef t_+#undef i_+#undef combine+}+intmain(intargc,char*argv[]){TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES)&BIT(KVM_X86_SW_PROTECTED_VM));
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
Add two test cases to the guest_memfd conversions selftest to cover
the scenario where a conversion is requested before any memory has been
allocated in the guest_memfd region.
The KVM_SET_MEMORY_ATTRIBUTES2 ioctl can be called on a memory region at
any time. If the guest had not yet faulted in any pages for that region,
the kernel must record the conversion request and apply the requested state
when the pages are eventually allocated.
The new tests cover both conversion directions.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
.../selftests/kvm/x86/guest_memfd_conversions_test.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Sean Christopherson <seanjc@google.com>
Add a test to verify that a guest_memfd's shared/private status is
consistent across processes, and that any shared pages previously mapped in
any process are unmapped from all processes.
The test forks a child process after creating the shared guest_memfd
region so that the second process exists alongside the main process for the
entire test.
The processes then take turns to access memory to check that the
shared/private status is consistent across processes.
Co-developed-by: Ackerley Tng <redacted>
Signed-off-by: Ackerley Tng <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
.../kvm/x86/guest_memfd_conversions_test.c | 118 +++++++++++++++++++++
1 file changed, 118 insertions(+)
@@ -326,6 +328,122 @@ GMEM_CONVERSION_TEST_INIT_SHARED(truncate)test_private(t,0,0,'A');}+/* Test that shared/private memory protections work and are seen from any process. */+GMEM_CONVERSION_TEST_INIT_SHARED(forked_accesses)+{+enumtest_state{+STATE_INIT,+STATE_CHECK_SHARED,+STATE_DONE_CHECKING_SHARED,+STATE_CHECK_PRIVATE,+STATE_DONE_CHECKING_PRIVATE,+};++structsync_state{+pthread_mutex_tmutex;+pthread_cond_tcond;+enumtest_statestep;+}*sync;++pthread_mutexattr_tmattr;+pthread_condattr_tcattr;+pid_tchild_pid,parent_pid;+intstatus;++sync=kvm_mmap(sizeof(*sync),PROT_READ|PROT_WRITE,+MAP_SHARED|MAP_ANONYMOUS,-1);++pthread_mutexattr_init(&mattr);+pthread_mutexattr_setpshared(&mattr,PTHREAD_PROCESS_SHARED);+pthread_mutex_init(&sync->mutex,&mattr);+pthread_mutexattr_destroy(&mattr);++pthread_condattr_init(&cattr);+pthread_condattr_setpshared(&cattr,PTHREAD_PROCESS_SHARED);+pthread_cond_init(&sync->cond,&cattr);+pthread_condattr_destroy(&cattr);++sync->step=STATE_INIT;++#define TEST_STATE_AWAIT(__state) \+do{\+pthread_mutex_lock(&sync->mutex);\+while(sync->step!=(__state)){\+structtimespects,stop;\+intret;\+\+clock_gettime(CLOCK_REALTIME,&ts);\+stop=timespec_add_ns(ts,100*1000000UL);\+\+ret=pthread_cond_timedwait(&sync->cond,&sync->mutex,&stop);\+if(ret==ETIMEDOUT){\+boolalive=(child_pid==0)?\+(getppid()==parent_pid):\+(waitpid(child_pid,NULL,WNOHANG)==0);\+TEST_ASSERT(alive,"Other process exited prematurely");\+}else{\+TEST_ASSERT(!ret,"pthread_cond_timedwait failed");\+}\+}\+pthread_mutex_unlock(&sync->mutex);\+}while(0)++#define TEST_STATE_SET(__state) \+do{\+pthread_mutex_lock(&sync->mutex);\+sync->step=(__state);\+pthread_cond_broadcast(&sync->cond);\+pthread_mutex_unlock(&sync->mutex);\+}while(0)++parent_pid=getpid();+child_pid=fork();+TEST_ASSERT(child_pid!=-1,"fork failed");++if(child_pid==0){+constcharinconsequential=0xdd;++TEST_STATE_AWAIT(STATE_CHECK_SHARED);++/*+*Thismapsthepagesintothechildprocessaswell,andtests+*thattheconversionprocesswillunmaptheguest_memfdmemory+*fromallprocesses.+*/+host_do_rmw(t->mem,0,0xB,0xC);++TEST_STATE_SET(STATE_DONE_CHECKING_SHARED);+TEST_STATE_AWAIT(STATE_CHECK_PRIVATE);++TEST_EXPECT_SIGBUS(READ_ONCE(t->mem[0]));+TEST_EXPECT_SIGBUS(WRITE_ONCE(t->mem[0],inconsequential));++TEST_STATE_SET(STATE_DONE_CHECKING_PRIVATE);+exit(0);+}++test_shared(t,0,0,0xA,0xB);++TEST_STATE_SET(STATE_CHECK_SHARED);+TEST_STATE_AWAIT(STATE_DONE_CHECKING_SHARED);++test_convert_to_private(t,0,0xC,0xD);++TEST_STATE_SET(STATE_CHECK_PRIVATE);+TEST_STATE_AWAIT(STATE_DONE_CHECKING_PRIVATE);++TEST_ASSERT_EQ(waitpid(child_pid,&status,0),child_pid);+TEST_ASSERT(WIFEXITED(status)&&WEXITSTATUS(status)==0,+"Child exited with unexpected status");++pthread_mutex_destroy(&sync->mutex);+pthread_cond_destroy(&sync->cond);+kvm_munmap(sync,sizeof(*sync));++#undef TEST_STATE_SET+#undef TEST_STATE_AWAIT+}+intmain(intargc,char*argv[]){TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES)&BIT(KVM_X86_SW_PROTECTED_VM));
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
Add helper functions to allow KVM selftests to pin memory using
CONFIG_GUP_TEST. This is useful for creating test scenarios where some page
has an increased refcount, such as when testing guest_memfd in-place
conversion.
The helpers open /sys/kernel/debug/gup_test and invoke the
PIN_LONGTERM_TEST_START and PIN_LONGTERM_TEST_STOP ioctls. Since this
functionality depends on the kernel being built with CONFIG_GUP_TEST,
provide stub implementations that trigger a test failure if the
configuration is missing.
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/include/kvm_util.h | 3 +++
tools/testing/selftests/kvm/lib/kvm_util.c | 25 +++++++++++++++++++++++++
2 files changed, 28 insertions(+)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
Add a selftest to verify that converting a shared guest_memfd page to a
private page fails if the page has an elevated reference count.
When KVM converts a shared page to a private one, it expects the page to
have a reference count equal to the reference counts taken by the
filemap. If another kernel subsystem holds a reference to the page, the
conversion must be aborted.
The test asserts that both bulk and single-page conversion attempts
correctly fail with EAGAIN for the pinned page. After the page is unpinned,
the test verifies that subsequent conversions succeed.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
.../kvm/x86/guest_memfd_conversions_test.c | 56 ++++++++++++++++++++++
1 file changed, 56 insertions(+)
@@ -444,6 +444,62 @@ GMEM_CONVERSION_TEST_INIT_SHARED(forked_accesses)#undef TEST_STATE_AWAIT}+staticvoidtest_convert_to_private_fails(test_data_t*t,u64pgoff,+size_tnr_pages,+u64expected_error_offset)+{+/* +1 to make it anything but expected_error_offset. */+u64error_offset=expected_error_offset+1;+u64offset=pgoff*page_size;+intret;++do{+ret=__gmem_set_private(t->gmem_fd,offset,+nr_pages*page_size,&error_offset);+}while(ret==-1&&errno==EINTR);+TEST_ASSERT(ret==-1&&errno==EAGAIN,+"Wanted EAGAIN on page %lu, got %d (ret = %d)",pgoff,+errno,ret);+TEST_ASSERT_EQ(error_offset,expected_error_offset);+}++GMEM_CONVERSION_MULTIPAGE_TEST_INIT_SHARED(elevated_refcount,4)+{+inti;++pin_pages(t->mem+test_page*page_size,page_size);++for(i=0;i<nr_pages;i++)+test_shared(t,i,0,'A','B');++/*+*Convertinginbulkshouldfailaslonganypageintherangehas+*unexpectedrefcounts.+*/+test_convert_to_private_fails(t,0,nr_pages,test_page*page_size);++for(i=0;i<nr_pages;i++){+/*+*Convertingpage-wiseshouldalsofailaslonganypageinthe+*rangehasunexpectedrefcounts.+*/+if(i==test_page)+test_convert_to_private_fails(t,i,1,test_page*page_size);+else+test_convert_to_private(t,i,'B','C');+}++unpin_pages();++gmem_set_private(t->gmem_fd,0,nr_pages*page_size);++for(i=0;i<nr_pages;i++){+charexpected=i==test_page?'B':'C';++test_private(t,i,expected,'D');+}+}+intmain(intargc,char*argv[]){TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES)&BIT(KVM_X86_SW_PROTECTED_VM));
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
Introduce a new helper, kvm_gpa_to_guest_memfd(), to find the
guest_memfd-related details of a memory region that contains a given guest
physical address (GPA).
The function returns the file descriptor for the memfd, the offset into
the file that corresponds to the GPA, and the number of bytes remaining
in the region from that GPA.
kvm_gpa_to_guest_memfd() was factored out from vm_guest_mem_fallocate();
refactor vm_guest_mem_fallocate() to use the new helper.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/include/kvm_util.h | 3 +++
tools/testing/selftests/kvm/lib/kvm_util.c | 37 ++++++++++++++++----------
2 files changed, 26 insertions(+), 14 deletions(-)
@@ -1332,27 +1332,20 @@ void vm_guest_mem_fallocate(struct kvm_vm *vm, u64 base, u64 size,boolpunch_hole){constintmode=FALLOC_FL_KEEP_SIZE|(punch_hole?FALLOC_FL_PUNCH_HOLE:0);-structuserspace_mem_region*region;u64end=base+size;-gpa_tgpa,len;off_tfd_offset;-intret;+intfd,ret;+size_tlen;+gpa_tgpa;for(gpa=base;gpa<end;gpa+=len){-u64offset;--region=userspace_mem_region_find(vm,gpa,gpa);-TEST_ASSERT(region&®ion->region.flags&KVM_MEM_GUEST_MEMFD,-"Private memory region not found for GPA 0x%lx",gpa);+fd=kvm_gpa_to_guest_memfd(vm,gpa,&fd_offset,&len);+len=min(end-gpa,len);-offset=gpa-region->region.guest_phys_addr;-fd_offset=region->region.guest_memfd_offset+offset;-len=min_t(u64,end-gpa,region->region.memory_size-offset);--ret=fallocate(region->region.guest_memfd,mode,fd_offset,len);+ret=fallocate(fd,mode,fd_offset,len);TEST_ASSERT(!ret,"fallocate() failed to %s at %lx (len = %lu), fd = %d, mode = %x, offset = %lx",punch_hole?"punch hole":"allocate",gpa,len,-region->region.guest_memfd,mode,fd_offset);+fd,mode,fd_offset);}}
@@ -1689,6 +1682,22 @@ void *addr_gpa2alias(struct kvm_vm *vm, gpa_t gpa)return(void*)((uintptr_t)region->host_alias+offset);}+intkvm_gpa_to_guest_memfd(structkvm_vm*vm,gpa_tgpa,off_t*fd_offset,+size_t*nr_bytes)+{+structuserspace_mem_region*region;+gpa_tgpa_offset;++region=userspace_mem_region_find(vm,gpa,gpa);+TEST_ASSERT(region&®ion->region.flags&KVM_MEM_GUEST_MEMFD,+"guest_memfd memory region not found for GPA 0x%lx",gpa);++gpa_offset=gpa-region->region.guest_phys_addr;+*fd_offset=region->region.guest_memfd_offset+gpa_offset;+*nr_bytes=region->region.memory_size-gpa_offset;+returnregion->region.guest_memfd;+}+/* Create an interrupt controller chip for the specified VM. */voidvm_create_irqchip(structkvm_vm*vm){
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Sean Christopherson <seanjc@google.com>
Introduce vm_mem_set_memory_attributes(), which handles setting of memory
attributes for a range of guest physical addresses, regardless of whether
the attributes should be set via guest_memfd or via the memory attributes
at the VM level.
Refactor existing vm_mem_set_{shared,private} functions to use the new
function. Opportunistically update the size parameter to use size_t instead
of u64.
Co-developed-by: Ackerley Tng <redacted>
Signed-off-by: Ackerley Tng <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/include/kvm_util.h | 46 +++++++++++++++++++-------
1 file changed, 34 insertions(+), 12 deletions(-)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
Add a test to verify that deallocating a page in a guest memfd region via
fallocate() with FALLOC_FL_PUNCH_HOLE does not alter the shared or private
status of the corresponding memory range.
When a page backing a guest memfd mapping is deallocated, e.g., by punching
a hole or truncating the file, and then subsequently faulted back in, the
new page must inherit the correct shared/private status tracked by
guest_memfd.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
.../selftests/kvm/x86/guest_memfd_conversions_test.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:40
From: Ackerley Tng <redacted>
private_mem_conversions_test used to reset the shared memory that was used
for the test to an initial pattern at the end of each test iteration. Then,
it would punch out the pages, which would zero memory.
Without in-place conversion, the resetting would write shared memory, and
hole-punching will zero private memory, hence resetting the test to the
state at the beginning of the for loop.
With in-place conversion, resetting writes memory as shared, and
hole-punching zeroes the same physical memory, hence undoing the reset
done before the hole punch.
Move the resetting after the hole-punching, and reset the entire
PER_CPU_DATA_SIZE instead of just the tested range.
With in-place conversion, this zeroes and then resets the same physical
memory. Without in-place conversion, the private memory is zeroed, and the
shared memory is reset to init_p.
This is sufficient since at each test stage, the memory is assumed to start
as shared, and private memory is always assumed to start zeroed. Conversion
zeroes memory, so the future test stages will work as expected.
Fixes: 43f623f350ce1 ("KVM: selftests: Add x86-only selftest for private memory conversions")
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
tools/testing/selftests/kvm/x86/private_mem_conversions_test.c | 9 ++++++---
1 file changed, 6 insertions(+), 3 deletions(-)
@@ -202,15 +202,18 @@ static void guest_test_explicit_conversion(u64 base_gpa, bool do_fallocate)guest_sync_shared(gpa,size,p3,p4);memcmp_g(gpa,p4,size);-/* Reset the shared memory back to the initial pattern. */-memset((void*)gpa,init_p,size);-/**Free(viaPUNCH_HOLE)*all*privatememorysothatthenext*iterationstartsfromacleanslate,e.g.withrespectto*whetherornottherearepages/foliosinguest_mem.*/guest_map_shared(base_gpa,PER_CPU_DATA_SIZE,true);++/*+*Hole-punchingabovezeroedprivatememory.Resetshared+*memoryinpreparationforthenextGUEST_STAGE.+*/+memset((void*)base_gpa,init_p,PER_CPU_DATA_SIZE);}}
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:41
From: Ackerley Tng <redacted>
Update private_mem_conversions_test for in-place conversions. In-place
conversions support is detected in selftests with kvm_has_gmem_attributes.
With in-place conversions, specifying userspace_addr from some other memory
provider that is not the guest_memfd associated with the memslot is a user
error, since KVM will only use both shared and private memory from the
guest_memfd.
Hence, when kvm_has_gmem_attributes, only test in-place conversions with
single backing, where guest_memfd provides both shared and private memory.
For single backing, guest_memfd must be created with
GUEST_MEMFD_FLAG_MMAP. Initialize the guest_memfd as shared to align with
how memory would default to shared when shared/private state was tracked at
the VM level (the test expects this, it was written for
non-in-place-conversions).
When handling a hypercall to set attributes, use
vm_mem_set_memory_attributes() to send the ioctl to the guest_memfd instead
of the VM.
When testing in-place conversions (single-backing), don't allow the user to
configure src_type, since src_type will be ignored. Don't use src_type to
determine alignment for rounding up per-cpu test memory size, since
guest_memfd's backing page size is always the system PAGE_SIZE.
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
---
.../kvm/x86/private_mem_conversions_test.c | 57 +++++++++++++++++-----
1 file changed, 44 insertions(+), 13 deletions(-)
@@ -352,8 +355,21 @@ static void *__test_mem_conversions(void *__vcpu)size_tnr_bytes=min_t(size_t,vm->page_size,size-i);u8*hva=addr_gpa2hva(vm,gpa+i);-/* In all cases, the host should observe the shared data. */-memcmp_h(hva,gpa+i,uc.args[3],nr_bytes);+if(kvm_has_gmem_attributes&&+uc.args[0]==SYNC_PRIVATE){+TEST_EXPECT_SIGBUS(READ_ONCE(*hva));+}else{+/*+*Ifnottestingin-placeconversion+*(dualbacking),thehostshould+*alwaysobserveshareddata,sincethe+*sharedmemoryisaseparatepage.+*+*ForSYNC_SHARED,testthatthehost+*canseesharedmemory.+*/+memcmp_h(hva,gpa+i,uc.args[3],nr_bytes);+}/* For shared, write the new pattern to guest memory. */if(uc.args[0]==SYNC_SHARED)
@@ -369,20 +385,29 @@ static void *__test_mem_conversions(void *__vcpu)}}+/* Align each vCPU's chunk of memory naturally to the size of the backing store. */+staticsize_tcompute_per_cpu_size(enumvm_mem_backing_src_typesrc_type)+{+size_talignment;++if(kvm_has_gmem_attributes)+alignment=getpagesize();+else+alignment=get_backing_src_pagesz(src_type);++returnalign_up(PER_CPU_DATA_SIZE,max_t(size_t,SZ_2M,alignment));+}+staticvoidtest_mem_conversions(enumvm_mem_backing_src_typesrc_type,u32nr_vcpus,u32nr_memslots){-/*-*AllocateenoughmemorysothateachvCPU'schunkofmemorycanbe-*naturallyalignedwithrespecttothesizeofthebackingstore.-*/-constsize_talignment=max_t(size_t,SZ_2M,get_backing_src_pagesz(src_type));-constsize_tper_cpu_size=align_up(PER_CPU_DATA_SIZE,alignment);+constsize_tper_cpu_size=compute_per_cpu_size(src_type);constsize_tmemfd_size=per_cpu_size*nr_vcpus;constsize_tslot_size=memfd_size/nr_memslots;structkvm_vcpu*vcpus[KVM_MAX_VCPUS];pthread_tthreads[KVM_MAX_VCPUS];structkvm_vm*vm;+u64gmem_flags;intmemfd,i;conststructvm_shapeshape={
@@ -462,6 +491,8 @@ int main(int argc, char *argv[])while((opt=getopt(argc,argv,"hm:s:n:"))!=-1){switch(opt){case's':+TEST_ASSERT(!kvm_has_gmem_attributes,+"src_type is only configurable when testing without in-place conversion");src_type=parse_backing_src_type(optarg);break;case'n':
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:41
From: Ackerley Tng <redacted>
Currently, vm_mem_add derives the backing source page size, alignment
padding, and mmap size from the backing source type upfront before checking
if guest_memfd is being mmapped.
With shared memory also mmap()-ed from guest_memfd, the alignment of the
mmap-ed address needs to respect guest_memfd's backing page size.
Refactor the backing store setup to configure the backing source page
size, alignment, mmap flags, and mmap offset directly for guest_memfd
when it is mmapped, ignoring the backing source type.
Skip hugepage validation and anonymous memory madvise calls when mmapping
from guest_memfd, since those are not applicable when mmapping guest_memfd.
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
---
tools/testing/selftests/kvm/lib/kvm_util.c | 74 ++++++++++++++++++------------
1 file changed, 45 insertions(+), 29 deletions(-)
@@ -1090,19 +1091,31 @@ void vm_mem_add(struct kvm_vm *vm, enum vm_mem_backing_src_type src_type,/* Allocate and initialize new mem region structure. */region=calloc(1,sizeof(*region));TEST_ASSERT(region!=NULL,"Insufficient Memory");-region->mmap_size=mem_size;-/*-*WhenusingTHPmmapisnotguaranteedtoreturnedahugepagealigned-*addresssowehavetopadthemmap.PaddingisnotneededforHugeTLB-*becausemmapwillalwaysreturnanaddressalignedtotheHugeTLB-*pagesize.-*/-if(src_type==VM_MEM_SRC_ANONYMOUS_THP)-alignment=max(backing_src_pagesz,alignment);+is_gmem_mmap=(flags&KVM_MEM_GUEST_MEMFD)&&+(gmem_flags&GUEST_MEMFD_FLAG_MMAP);++if(is_gmem_mmap){+backing_src_pagesz=getpagesize();+alignment=1;+mmap_flags=MAP_SHARED;+mmap_offset=gmem_offset;+}else{+backing_src_pagesz=get_backing_src_pagesz(src_type);+/*+*WhenusingTHPmmapisnotguaranteedtoreturnedahugepagealigned+*addresssowehavetopadthemmap.PaddingisnotneededforHugeTLB+*becausemmapwillalwaysreturnanaddressalignedtotheHugeTLB+*pagesize.+*/+alignment=src_type==VM_MEM_SRC_ANONYMOUS_THP?backing_src_pagesz:1;+mmap_flags=vm_mem_backing_src_alias(src_type)->flag;+mmap_offset=0;+}TEST_ASSERT_EQ(gpa,align_up(gpa,backing_src_pagesz));+region->mmap_size=mem_size;/* Add enough memory to align up if necessary */if(alignment>1)region->mmap_size+=alignment;
From: Ackerley Tng via B4 Relay <devnull+ackerleytng.google.com@kernel.org> Date: 2026-09-10 23:55:41
From: Sean Christopherson <seanjc@google.com>
Skip setting memory to private in the private memory exits test when using
per-gmem memory attributes, as memory is initialized to private by default
for guest_memfd, and using vm_mem_set_private() on a guest_memfd instance
requires creating guest_memfd with GUEST_MEMFD_FLAG_MMAP (which is totally
doable, but would need to be conditional and is ultimately unnecessary).
Expect an emulated MMIO instead of a memory fault exit when attributes are
per-gmem, as deleting the memslot effectively drops the private status,
i.e. the GPA becomes shared and thus supports emulated MMIO.
Skip the "memslot not private" test entirely, as private vs. shared state
for x86 software-protected VMs comes from the memory attributes themselves,
and so when doing in-place conversions there can never be a disconnect
between the expected and actual states.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Tested-by: Shivank Garg <redacted>
Reviewed-by: Fuad Tabba <redacted>
---
.../selftests/kvm/x86/private_mem_kvm_exits_test.c | 36 ++++++++++++++++++----
1 file changed, 30 insertions(+), 6 deletions(-)
From: Xiaoyao Li <hidden> Date: 2026-09-11 09:18:11
On 9/11/2026 7:55 AM, Ackerley Tng via B4 Relay wrote:
From: Sean Christopherson <seanjc@google.com>
Add and use kvm_arch_has_gmem_convert() to guard guest_memfd's invocation
of arch hooks related to converting memory between private and shared, as
only one half of the x86 CoCo duo needs the runtime hooks (any pre-work is
pure overhead for TDX). At this exact moment, the overhead is negligible,
but that will change when in-place conversion comes along, at which point
to-shared conversions will "need" to find all affected folios prior to
calling into arch code. In quotes because very technically that work could
be pushed to arch code, but that would bleed guest_memfd details into arch
code and would be far worse than adding yet another kvm_arch_has... hook.
Opportunistically provide the kvm_arch_gmem_make_private() declaration, and
rely on dead-code elimination to eliminate the call to non-existent code
when CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT=n.
Reported-by: Binbin Wu <redacted>
Closes: https://lore.kernel.org/all/1ec08cd8-3072-4753-ad5e-cd34956647f8@linux.intel.com
Suggested-by: Ackerley Tng <redacted>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <redacted>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
From: Xiaoyao Li <hidden> Date: 2026-09-11 09:19:54
On 9/11/2026 7:55 AM, Ackerley Tng via B4 Relay wrote:
From: Ackerley Tng <redacted>
Rename kvm_mem_is_private() to kvm_is_private_gfn() to prepare for in-place
conversion, where there will be two lookup functions,
kvm_vm_is_private_gfn() and kvm_gmem_is_private_gfn().
This renaming allows consistent prefixing of "vm" vs "gmem" for
kvm_*_is_private_gfn(), as opposed to kvm_gmem_mem_is_private(), which
looks like a typo.
Suggested-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Ackerley Tng <redacted>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <redacted>
On Thu, Sep 10, 2026 at 04:55:26PM -0700, Ackerley Tng via B4 Relay wrote:
Here's v13. Thanks everyone for the comments and fast responses!
We're now at ~3 weeks to soft-close at 7.3-rc5.
v13 is based on 7.3-rc2 and:
+ Hugh's patch to export lru_cache_drain_for_folio()
+ Another series [1], which makes kvm_gmem_get_pfn() NOT return a
refcounted page to KVM.
Here's everything stitched together for your convenience:
https://github.com/googleprodkernel/linux-cc/commits/guest_memfd-inplace-conversion-v13
The following configurations were tested and all booted TDs successfully:
(1) gmem_in_place_conversion=1, in-place conversions + in-place-add init mem
(2) gmem_in_place_conversion=1, in-place conversions + out-of-place-add init mem
(3) gmem_in_place_conversion=0, out-of-place conversions + out-of-place-add init mem
Tested-by: Yan Zhao <redacted>