Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create... | linux-api

[PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Chao Peng <hidden> · 2022-10-25
[PATCH v9 2/8] KVM: Extend the memslot to support fd-based private memory · Chao Peng <hidden> · 2022-10-25
Re: [PATCH v9 2/8] KVM: Extend the memslot to support fd-based private memory · Fuad Tabba <hidden> · 2022-10-27
Re: [PATCH v9 2/8] KVM: Extend the memslot to support fd-based private memory · Xiaoyao Li <hidden> · 2022-10-28
Re: [PATCH v9 2/8] KVM: Extend the memslot to support fd-based private memory · Chao Peng <hidden> · 2022-10-31
Re: [PATCH v9 2/8] KVM: Extend the memslot to support fd-based private memory · Alex Bennée <hidden> · 2022-11-14
Re: [PATCH v9 2/8] KVM: Extend the memslot to support fd-based private memory · Chao Peng <hidden> · 2022-11-15
[PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-10-25
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Isaku Yamahata <hidden> · 2022-10-26
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-10-28
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Fuad Tabba <hidden> · 2022-10-27
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-10-31
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-11-01
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-01
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-01
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-11-02
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-02
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Vlastimil Babka <hidden> · 2022-11-14
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Kirill A. Shutemov <hidden> · 2022-11-14
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-14
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-11-15
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-14
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Kirill A. Shutemov <hidden> · 2022-11-02
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-02
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-02
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Kirill A. Shutemov <hidden> · 2022-11-03
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Kirill A. Shutemov <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · David Hildenbrand <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-11-30
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Michael Roth <hidden> · 2022-11-30
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Vishal Annapurve <hidden> · 2022-11-29
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Vishal Annapurve <hidden> · 2022-12-02
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Chao Peng <hidden> · 2022-12-02
Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory · Kirill A . Shutemov <hidden> · 2022-12-02
[PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Chao Peng <hidden> · 2022-10-25
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Peter Maydell <hidden> · 2022-10-25
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Sean Christopherson <seanjc@google.com> · 2022-10-25
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Fuad Tabba <hidden> · 2022-10-27
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Chao Peng <hidden> · 2022-10-28
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Alex Bennée <hidden> · 2022-11-15
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Chao Peng <hidden> · 2022-11-16
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Alex Bennée <hidden> · 2022-11-16
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Chao Peng <hidden> · 2022-11-17
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Alex Bennée <hidden> · 2022-11-17
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Chao Peng <hidden> · 2022-11-18
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Alex Bennée <hidden> · 2022-11-18
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Sean Christopherson <seanjc@google.com> · 2022-11-18
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Chao Peng <hidden> · 2022-11-22
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Sean Christopherson <seanjc@google.com> · 2022-11-23
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · "Andy Lutomirski" <luto@kernel.org> · 2022-11-16
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Sean Christopherson <seanjc@google.com> · 2022-11-16
Re: [PATCH v9 3/8] KVM: Add KVM_EXIT_MEMORY_FAULT exit · Chao Peng <hidden> · 2022-11-17
[PATCH v9 4/8] KVM: Use gfn instead of hva for mmu_notifier_retry · Chao Peng <hidden> · 2022-10-25
Re: [PATCH v9 4/8] KVM: Use gfn instead of hva for mmu_notifier_retry · Fuad Tabba <hidden> · 2022-10-27
Re: [PATCH v9 4/8] KVM: Use gfn instead of hva for mmu_notifier_retry · Chao Peng <hidden> · 2022-11-04
Re: [PATCH v9 4/8] KVM: Use gfn instead of hva for mmu_notifier_retry · Sean Christopherson <seanjc@google.com> · 2022-11-04
Re: [PATCH v9 4/8] KVM: Use gfn instead of hva for mmu_notifier_retry · Chao Peng <hidden> · 2022-11-08
Re: [PATCH v9 4/8] KVM: Use gfn instead of hva for mmu_notifier_retry · Sean Christopherson <seanjc@google.com> · 2022-11-10
Re: [PATCH v9 4/8] KVM: Use gfn instead of hva for mmu_notifier_retry · Sean Christopherson <seanjc@google.com> · 2022-11-10
Re: [PATCH v9 4/8] KVM: Use gfn instead of hva for mmu_notifier_retry · Chao Peng <hidden> · 2022-11-11
[PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Chao Peng <hidden> · 2022-10-25
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Fuad Tabba <hidden> · 2022-10-27
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Sean Christopherson <seanjc@google.com> · 2022-11-03
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Chao Peng <hidden> · 2022-11-04
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Sean Christopherson <seanjc@google.com> · 2022-11-04
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Chao Peng <hidden> · 2022-11-08
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Yuan Yao <hidden> · 2022-11-08
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Chao Peng <hidden> · 2022-11-08
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Yuan Yao <hidden> · 2022-11-09
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Sean Christopherson <seanjc@google.com> · 2022-11-16
Re: [PATCH v9 5/8] KVM: Register/unregister the guest private memory regions · Chao Peng <hidden> · 2022-11-17
[PATCH v9 6/8] KVM: Update lpage info when private/shared memory are mixed · Chao Peng <hidden> · 2022-10-25
Re: [PATCH v9 6/8] KVM: Update lpage info when private/shared memory are mixed · Isaku Yamahata <hidden> · 2022-10-26
Re: [PATCH v9 6/8] KVM: Update lpage info when private/shared memory are mixed · Chao Peng <hidden> · 2022-10-28
Re: [PATCH v9 6/8] KVM: Update lpage info when private/shared memory are mixed · Yuan Yao <hidden> · 2022-11-08
Re: [PATCH v9 6/8] KVM: Update lpage info when private/shared memory are mixed · Chao Peng <hidden> · 2022-11-09
[PATCH v9 7/8] KVM: Handle page fault for private memory · Chao Peng <hidden> · 2022-10-25
Re: [PATCH v9 7/8] KVM: Handle page fault for private memory · Isaku Yamahata <hidden> · 2022-10-26
Re: [PATCH v9 7/8] KVM: Handle page fault for private memory · Chao Peng <hidden> · 2022-10-28
Re: [PATCH v9 7/8] KVM: Handle page fault for private memory · Isaku Yamahata <hidden> · 2022-11-01
Re: [PATCH v9 7/8] KVM: Handle page fault for private memory · Chao Peng <hidden> · 2022-11-01
Re: [PATCH v9 7/8] KVM: Handle page fault for private memory · Ackerley Tng <hidden> · 2022-11-16
Re: [PATCH v9 7/8] KVM: Handle page fault for private memory · Sean Christopherson <seanjc@google.com> · 2022-11-16
Re: [PATCH v9 7/8] KVM: Handle page fault for private memory · Chao Peng <hidden> · 2022-11-17
[PATCH v9 8/8] KVM: Enable and expose KVM_MEM_PRIVATE · Chao Peng <hidden> · 2022-10-25
Re: [PATCH v9 8/8] KVM: Enable and expose KVM_MEM_PRIVATE · Fuad Tabba <hidden> · 2022-10-27
Re: [PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Vishal Annapurve <hidden> · 2022-11-03
Re: [PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Isaku Yamahata <hidden> · 2022-11-08
Re: [PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Kirill A. Shutemov <hidden> · 2022-11-09
Re: [PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Kirill A. Shutemov <hidden> · 2022-11-15
Re: [PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Alex Bennée <hidden> · 2022-11-14
Re: [PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Chao Peng <hidden> · 2022-11-16
Re: [PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Alex Bennée <hidden> · 2022-11-16
Re: [PATCH v9 0/8] KVM: mm: fd-based approach for supporting KVM · Chao Peng <hidden> · 2022-11-17

Re: [PATCH v9 1/8] mm: Introduce memfd_restricted system call to create restricted user memory

From: Michael Roth <hidden>
Date: 2022-11-30 14:32:50
Also in: kvm, linux-arch, linux-doc, linux-fsdevel, linux-mm, lkml, qemu-devel

On Wed, Nov 30, 2022 at 05:39:31PM +0800, Chao Peng wrote:

On Tue, Nov 29, 2022 at 01:18:15PM -0600, Michael Roth wrote:

quoted

On Tue, Nov 29, 2022 at 01:06:58PM -0600, Michael Roth wrote:

quoted

On Tue, Nov 29, 2022 at 10:06:15PM +0800, Chao Peng wrote:

quoted

On Mon, Nov 28, 2022 at 06:37:25PM -0600, Michael Roth wrote:

quoted

On Tue, Oct 25, 2022 at 11:13:37PM +0800, Chao Peng wrote:

...

quoted

+static long restrictedmem_fallocate(struct file *file, int mode,
+				    loff_t offset, loff_t len)
+{
+	struct restrictedmem_data *data = file->f_mapping->private_data;
+	struct file *memfd = data->memfd;
+	int ret;
+
+	if (mode & FALLOC_FL_PUNCH_HOLE) {
+		if (!PAGE_ALIGNED(offset) || !PAGE_ALIGNED(len))
+			return -EINVAL;
+	}
+
+	restrictedmem_notifier_invalidate(data, offset, offset + len, true);

The KVM restrictedmem ops seem to expect pgoff_t, but here we pass
loff_t. For SNP we've made this strange as part of the following patch
and it seems to produce the expected behavior:

That's correct. Thanks.

quoted

  https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fgithub.com%2Fmdroth%2Flinux%2Fcommit%2Fd669c7d3003ff7a7a47e73e8c3b4eeadbd2c4eb6&amp;data=05%7C01%7Cmichael.roth%40amd.com%7Cf3ad9d505bec4006028308dad2b76bc5%7C3dd8961fe4884e608e11a82d994e183d%7C0%7C0%7C638053982483658905%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=ipHjTVNhiRmaa%2BKTJiodbxHS7TOaYbBhAPD0VZ%2FFU2k%3D&amp;reserved=0

quoted

+	ret = memfd->f_op->fallocate(memfd, mode, offset, len);
+	restrictedmem_notifier_invalidate(data, offset, offset + len, false);
+	return ret;
+}
+

<snip>

quoted

+int restrictedmem_get_page(struct file *file, pgoff_t offset,
+			   struct page **pagep, int *order)
+{
+	struct restrictedmem_data *data = file->f_mapping->private_data;
+	struct file *memfd = data->memfd;
+	struct page *page;
+	int ret;
+
+	ret = shmem_getpage(file_inode(memfd), offset, &page, SGP_WRITE);

This will result in KVM allocating pages that userspace hasn't necessary
fallocate()'d. In the case of SNP we need to get the PFN so we can clean
up the RMP entries when restrictedmem invalidations are issued for a GFN
range.

Yes fallocate() is unnecessary unless someone wants to reserve some
space (e.g. for determination or performance purpose), this matches its
semantics perfectly at:
https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.man7.org%2Flinux%2Fman-pages%2Fman2%2Ffallocate.2.html&amp;data=05%7C01%7Cmichael.roth%40amd.com%7Cf3ad9d505bec4006028308dad2b76bc5%7C3dd8961fe4884e608e11a82d994e183d%7C0%7C0%7C638053982483658905%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=NJXs0bvvqb3oU%2FGhcvgHSvh8r1DouskOY5CreP1Q5OU%3D&amp;reserved=0

quoted

If the guest supports lazy-acceptance however, these pages may not have
been faulted in yet, and if the VMM defers actually fallocate()'ing space
until the guest actually tries to issue a shared->private for that GFN
(to support lazy-pinning), then there may never be a need to allocate
pages for these backends.

However, the restrictedmem invalidations are for GFN ranges so there's
no way to know inadvance whether it's been allocated yet or not. The
xarray is one option but currently it defaults to 'private' so that
doesn't help us here. It might if we introduced a 'uninitialized' state
or something along that line instead of just the binary
'shared'/'private' though...

How about if we change the default to 'shared' as we discussed at
https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Flore.kernel.org%2Fall%2FY35gI0L8GMt9%2BOkK%40google.com%2F&amp;data=05%7C01%7Cmichael.roth%40amd.com%7Cf3ad9d505bec4006028308dad2b76bc5%7C3dd8961fe4884e608e11a82d994e183d%7C0%7C0%7C638053982483658905%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=%2F1g3NdU0iLO6rWVgSm42UYlfHGG2EJ1Wp0r%2FGEznUoo%3D&amp;reserved=0?

Need to look at this a bit more, but I think that could work as well.

quoted

But for now we added a restrictedmem_get_page_noalloc() that uses
SGP_NONE instead of SGP_WRITE to avoid accidentally allocating a bunch
of memory as part of guest shutdown, and a
kvm_restrictedmem_get_pfn_noalloc() variant to go along with that. But
maybe a boolean param is better? Or maybe SGP_NOALLOC is the better
default, and we just propagate an error to userspace if they didn't
fallocate() in advance?

This (making fallocate() a hard requirement) not only complicates the
userspace but also forces the lazy-faulting going through a long path of
exiting to userspace. Unless we don't have other options I would not go
this way.

Unless I'm missing something, it's already the case that userspace is
responsible for handling all the shared->private transitions in response
to KVM_EXIT_MEMORY_FAULT or (in our case) KVM_EXIT_VMGEXIT. So it only
places the additional requirements on the VMM that if they *don't*
preallocate, then they'll need to issue the fallocate() prior to issuing
the KVM_MEM_ENCRYPT_REG_REGION ioctl in response to these events.

Preallocating and memory conversion between shared<->private are two
different things. No double fallocate() and conversion can be called

I just mean that we don't actually have additional userspace exits for
doing lazy-faulting in this manner, because prior to mapping restricted
page into the TDP, we will have gotten a KVM_EXIT_MEMORY_FAULT anyway so
that userspace can handle the conversion, so if you do the fallocate()
prior to KVM_MEM_ENCRYPT_REG_REGION, there's no additional KVM exits
(unless you count the fallocate() syscall itself but that seems
negligable compared to memory allocation).

For instance on QEMU side we do the fallocate() as part of
kvm_convert_memory() helper.

But thinking about it more, the main upside to this approach (giving VMM
control/accounting over restrictedmem allocations), doesn't actually
work out. For instance if VMM fallocate()'s memory for a single 4K page
prior to shared->private conversion, shmem might still allocate a THP for
that whole 2M range, and userspace doesn't have a good way to account
for this. So what I'm proposing probably isn't feasible anyway.

different things. No double fallocate() and conversion can be called
together in response to KVM_EXIT_MEMORY_FAULT, but they don't have to be
paired. And the fallocate() does not have to operate on the same memory
range as memory conversion does.

quoted

QEMU for example already has a separate 'prealloc' option for cases
where they want to prefault all the guest memory, so it makes sense to
continue making that an optional thing with regard to UPM.

Making 'prealloc' work for UPM in QEMU does sound reasonable. Anyway,
it's just an option so not change the assumption here.

quoted

Although I guess what you're suggesting doesn't stop userspace from
deciding whether they want to prefault or not. I know the Google folks
had some concerns over unexpected allocations causing 2x memory usage
though so giving userspace full control of what is/isn't allocated in
the restrictedmem backend seems to make it easier to guard against this,
but I think checking the xarray and defaulting to 'shared' would work
for us if that's the direction we end up going.

Yeah, that looks very likely the direction satisfying all people here.

Ok, yah after some more thought this probably is the more feasible
approach. Thanks for your input on this.

-Mike

Chao

quoted

-Mike

quoted

-Mike

quoted

Chao

quoted

-Mike

quoted

+	if (ret)
+		return ret;
+
+	*pagep = page;
+	if (order)
+		*order = thp_order(compound_head(page));
+
+	SetPageUptodate(page);
+	unlock_page(page);
+
+	return 0;
+}
+EXPORT_SYMBOL_GPL(restrictedmem_get_page);
-- 
2.25.1

`h`	back out one level
`j`	next message in thread
`k`	previous message in thread
`l`	drill in
`Esc`	close help / fold thread tree
`?`	toggle this help