This part of Secure Encrypted Paging (SEV-SNP) series focuses on the changes
required in a guest OS for SEV-SNP support.
SEV-SNP builds upon existing SEV and SEV-ES functionality while adding
new hardware-based memory protections. SEV-SNP adds strong memory integrity
protection to help prevent malicious hypervisor-based attacks like data
replay, memory re-mapping and more in order to create an isolated memory
encryption environment.
This series provides the basic building blocks to support booting the SEV-SNP
VMs, it does not cover all the security enhancement introduced by the SEV-SNP
such as interrupt protection.
Many of the integrity guarantees of SEV-SNP are enforced through a new
structure called the Reverse Map Table (RMP). Adding a new page to SEV-SNP
VM requires a 2-step process. First, the hypervisor assigns a page to the
guest using the new RMPUPDATE instruction. This transitions the page to
guest-invalid. Second, the guest validates the page using the new PVALIDATE
instruction. The SEV-SNP VMs can use the new "Page State Change Request NAE"
defined in the GHCB specification to ask hypervisor to add or remove page
from the RMP table.
Each page assigned to the SEV-SNP VM can either be validated or unvalidated,
as indicated by the Validated flag in the page's RMP entry. There are two
approaches that can be taken for the page validation: Pre-validation and
Lazy Validation.
Under pre-validation, the pages are validated prior to first use. And under
lazy validation, pages are validated when first accessed. An access to a
unvalidated page results in a #VC exception, at which time the exception
handler may validate the page. Lazy validation requires careful tracking of
the validated pages to avoid validating the same GPA more than once. The
recently introduced "Unaccepted" memory type can be used to communicate the
unvalidated memory ranges to the Guest OS.
At this time we only sypport the pre-validation, the OVMF guest BIOS
validates the entire RAM before the control is handed over to the guest kernel.
The early_set_memory_{encrypt,decrypt} and set_memory_{encrypt,decrypt} are
enlightened to perform the page validation or invalidation while setting or
clearing the encryption attribute from the page table.
This series does not provide support for the Interrupt security yet which will
be added after the base support.
The series is based on tip/master
a6d06ef25c4e (origin/master, origin/HEAD, master) Merge branch 'irq/core
Additional resources
---------------------
SEV-SNP whitepaper
https://www.amd.com/system/files/TechDocs/SEV-SNP-strengthening-vm-isolation-with-integrity-protection-and-more.pdf
APM 2: https://www.amd.com/system/files/TechDocs/24593.pdf
(section 15.36)
GHCB spec:
https://developer.amd.com/wp-content/resources/56421.pdf
SEV-SNP firmware specification:
https://developer.amd.com/sev/
v5: https://lore.kernel.org/lkml/20210820151933.22401-1-brijesh.singh@amd.com/
Changes since v5:
* move the seqno allocation in the sevguest driver.
* extend snp_issue_guest_request() to accept the exit_info to simplify the logic.
* use smaller structure names based on feedback.
* explicitly clear the memory after the SNP guest request is completed.
* cpuid validation: use a local copy of cpuid table instead of keeping
firmware table mapped throughout boot.
* cpuid validation: coding style fix-ups and refactor cpuid-related helpers
as suggested.
* cpuid validation: drop a number of BOOT_COMPRESSED-guarded defs/declarations
by moving things like snp_cpuid_init*() out of sev-shared.c and keeping only
the common bits there.
* Break up EFI config table helpers and related acpi.c changes into separate
patches.
* re-enable stack protection for 32-bit kernels as well, not just 64-bit
Changes since v4:
* Address the cpuid specific review comment
* Simplified the macro based on the review feedback
* Move macro definition to the patch that needs it
* Fix the issues reported by the checkpath
* Address the AP creation specific review comment
Changes since v3:
* Add support to use the PSP filtered CPUID.
* Add support for the extended guest request.
* Move sevguest driver in driver/virt/coco.
* Add documentation for sevguest ioctl.
* Add support to check the vmpl0.
* Pass the VM encryption key and id to be used for encrypting guest messages
through the platform drv data.
* Multiple cleanup and fixes to address the review feedbacks.
Changes since v2:
* Add support for AP startup using SNP specific vmgexit.
* Add snp_prep_memory() helper.
* Drop sev_snp_active() helper.
* Add sev_feature_enabled() helper to check which SEV feature is active.
* Sync the SNP guest message request header with latest SNP FW spec.
* Multiple cleanup and fixes to address the review feedbacks.
Changes since v1:
* Integerate the SNP support in sev.{ch}.
* Add support to query the hypervisor feature and detect whether SNP is supported.
* Define Linux specific reason code for the SNP guest termination.
* Extend the setup_header provide a way for hypervisor to pass secret and cpuid page.
* Add support to create a platform device and driver to query the attestation report
and the derive a key.
* Multiple cleanup and fixes to address Boris's review fedback.
Borislav Petkov (3):
x86/sev: Get rid of excessive use of defines
x86/head64: Carve out the guest encryption postprocessing into a
helper
x86/sev: Remove do_early_exception() forward declarations
Brijesh Singh (22):
x86/mm: Extend cc_attr to include AMD SEV-SNP
x86/sev: Shorten GHCB terminate macro names
x86/sev: Define the Linux specific guest termination reasons
x86/sev: Save the negotiated GHCB version
x86/sev: Add support for hypervisor feature VMGEXIT
x86/sev: Check SEV-SNP features support
x86/sev: Add a helper for the PVALIDATE instruction
x86/sev: Check the vmpl level
x86/compressed: Add helper for validating pages in the decompression
stage
x86/compressed: Register GHCB memory when SEV-SNP is active
x86/sev: Register GHCB memory when SEV-SNP is active
x86/sev: Add helper for validating pages in early enc attribute
changes
x86/kernel: Make the bss.decrypted section shared in RMP table
x86/kernel: Validate rom memory before accessing when SEV-SNP is
active
x86/mm: Add support to validate memory when changing C-bit
KVM: SVM: Define sev_features and vmpl field in the VMSA
x86/boot: Add Confidential Computing type to setup_data
x86/sev: Provide support for SNP guest request NAEs
x86/sev: Register SNP guest request platform device
virt: Add SEV-SNP guest driver
virt: sevguest: Add support to derive key
virt: sevguest: Add support to get extended report
Michael Roth (13):
x86/sev-es: initialize sev_status/features within #VC handler
x86/head: re-enable stack protection for 32/64-bit builds
x86/sev: move MSR-based VMGEXITs for CPUID to helper
KVM: x86: move lookup of indexed CPUID leafs to helper
x86/compressed/acpi: move EFI system table lookup to helper
x86/compressed/acpi: move EFI config table lookup to helper
x86/compressed/acpi: move EFI vendor table lookup to helper
x86/compressed/64: add support for SEV-SNP CPUID table in #VC handlers
boot/compressed/64: use firmware-validated CPUID for SEV-SNP guests
x86/boot: add a pointer to Confidential Computing blob in bootparams
x86/compressed/64: store Confidential Computing blob address in
bootparams
x86/compressed/64: add identity mapping for Confidential Computing
blob
x86/sev: use firmware-validated CPUID for SEV-SNP guests
Tom Lendacky (4):
KVM: SVM: Create a separate mapping for the SEV-ES save area
KVM: SVM: Create a separate mapping for the GHCB save area
KVM: SVM: Update the SEV-ES save area mapping
x86/sev: Use SEV-SNP AP creation to start secondary CPUs
Documentation/virt/coco/sevguest.rst | 117 ++++
arch/x86/boot/compressed/Makefile | 1 +
arch/x86/boot/compressed/acpi.c | 120 +---
arch/x86/boot/compressed/efi.c | 171 +++++
arch/x86/boot/compressed/head_64.S | 1 +
arch/x86/boot/compressed/ident_map_64.c | 44 +-
arch/x86/boot/compressed/idt_64.c | 5 +-
arch/x86/boot/compressed/misc.h | 42 ++
arch/x86/boot/compressed/sev.c | 189 +++++-
arch/x86/include/asm/bootparam_utils.h | 1 +
arch/x86/include/asm/cpuid.h | 26 +
arch/x86/include/asm/msr-index.h | 2 +
arch/x86/include/asm/setup.h | 2 +-
arch/x86/include/asm/sev-common.h | 137 +++-
arch/x86/include/asm/sev.h | 80 ++-
arch/x86/include/asm/svm.h | 167 ++++-
arch/x86/include/uapi/asm/bootparam.h | 4 +-
arch/x86/include/uapi/asm/svm.h | 13 +
arch/x86/kernel/Makefile | 1 -
arch/x86/kernel/cc_platform.c | 2 +
arch/x86/kernel/head64.c | 79 ++-
arch/x86/kernel/head_64.S | 24 +
arch/x86/kernel/probe_roms.c | 13 +-
arch/x86/kernel/sev-shared.c | 569 +++++++++++++++-
arch/x86/kernel/sev.c | 860 ++++++++++++++++++++++--
arch/x86/kernel/smpboot.c | 3 +
arch/x86/kvm/cpuid.c | 17 +-
arch/x86/kvm/svm/sev.c | 24 +-
arch/x86/kvm/svm/svm.c | 4 +-
arch/x86/kvm/svm/svm.h | 2 +-
arch/x86/mm/mem_encrypt.c | 55 +-
arch/x86/mm/pat/set_memory.c | 15 +
drivers/virt/Kconfig | 3 +
drivers/virt/Makefile | 1 +
drivers/virt/coco/sevguest/Kconfig | 9 +
drivers/virt/coco/sevguest/Makefile | 2 +
drivers/virt/coco/sevguest/sevguest.c | 703 +++++++++++++++++++
drivers/virt/coco/sevguest/sevguest.h | 98 +++
include/linux/cc_platform.h | 8 +
include/linux/efi.h | 1 +
include/uapi/linux/sev-guest.h | 81 +++
41 files changed, 3389 insertions(+), 307 deletions(-)
create mode 100644 Documentation/virt/coco/sevguest.rst
create mode 100644 arch/x86/boot/compressed/efi.c
create mode 100644 arch/x86/include/asm/cpuid.h
create mode 100644 drivers/virt/coco/sevguest/Kconfig
create mode 100644 drivers/virt/coco/sevguest/Makefile
create mode 100644 drivers/virt/coco/sevguest/sevguest.c
create mode 100644 drivers/virt/coco/sevguest/sevguest.h
create mode 100644 include/uapi/linux/sev-guest.h
--
2.25.1
The CC_ATTR_SEV_SNP can be used by the guest to query whether the SNP -
Secure Nested Paging feature is active.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/msr-index.h | 2 ++
arch/x86/kernel/cc_platform.c | 2 ++
arch/x86/mm/mem_encrypt.c | 4 ++++
include/linux/cc_platform.h | 8 ++++++++
4 files changed, 16 insertions(+)
@@ -1397,7 +1397,7 @@ DEFINE_IDTENTRY_VC_KERNEL(exc_vmm_communication)show_regs(regs);/* Ask hypervisor to sev_es_terminate */-sev_es_terminate(GHCB_SEV_ES_REASON_GENERAL_REQUEST);+sev_es_terminate(GHCB_SEV_ES_GEN_REQ);/* If that fails and we get here - just panic */panic("Returned from Terminate-Request to Hypervisor\n");
@@ -1445,7 +1445,7 @@ bool __init handle_vc_boot_ghcb(struct pt_regs *regs)/* Do initial setup or terminate the guest */if(unlikely(boot_ghcb==NULL&&!sev_es_setup_ghcb()))-sev_es_terminate(GHCB_SEV_ES_REASON_GENERAL_REQUEST);+sev_es_terminate(GHCB_SEV_ES_GEN_REQ);vc_ghcb_invalidate(boot_ghcb);
From: Borislav Petkov <redacted>
Remove all the defines of masks and bit positions for the GHCB MSR
protocol and use comments instead which correspond directly to the spec
so that following those can be a lot easier and straightforward with the
spec opened in parallel to the code.
Aligh vertically while at it.
No functional changes.
Signed-off-by: Borislav Petkov <redacted>
---
arch/x86/include/asm/sev-common.h | 51 +++++++++++++++++--------------
1 file changed, 28 insertions(+), 23 deletions(-)
GHCB specification defines the reason code for reason set 0. The reason
codes defined in the set 0 do not cover all possible causes for a guest
to request termination.
The reason set 1 to 255 is reserved for the vendor-specific codes.
Reseve the reason set 1 for the Linux guest. Define an error codes for
reason set 1.
While at it, change the sev_es_terminate() to accept the reason set
parameter.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/sev.c | 6 +++---
arch/x86/include/asm/sev-common.h | 8 ++++++++
arch/x86/kernel/sev-shared.c | 11 ++++-------
arch/x86/kernel/sev.c | 4 ++--
4 files changed, 17 insertions(+), 12 deletions(-)
@@ -1397,7 +1397,7 @@ DEFINE_IDTENTRY_VC_KERNEL(exc_vmm_communication)show_regs(regs);/* Ask hypervisor to sev_es_terminate */-sev_es_terminate(GHCB_SEV_ES_GEN_REQ);+sev_es_terminate(SEV_TERM_SET_GEN,GHCB_SEV_ES_GEN_REQ);/* If that fails and we get here - just panic */panic("Returned from Terminate-Request to Hypervisor\n");
@@ -1445,7 +1445,7 @@ bool __init handle_vc_boot_ghcb(struct pt_regs *regs)/* Do initial setup or terminate the guest */if(unlikely(boot_ghcb==NULL&&!sev_es_setup_ghcb()))-sev_es_terminate(GHCB_SEV_ES_GEN_REQ);+sev_es_terminate(SEV_TERM_SET_GEN,GHCB_SEV_ES_GEN_REQ);vc_ghcb_invalidate(boot_ghcb);
The SEV-ES guest calls the sev_es_negotiate_protocol() to negotiate the
GHCB protocol version before establishing the GHCB. Cache the negotiated
GHCB version so that it can be used later.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev.h | 2 +-
arch/x86/kernel/sev-shared.c | 17 ++++++++++++++---
2 files changed, 15 insertions(+), 4 deletions(-)
@@ -99,7 +110,7 @@ static enum es_result sev_es_ghcb_hv_call(struct ghcb *ghcb,enumes_resultret;/* Fill in protocol and format specifiers */-ghcb->protocol_version=GHCB_PROTOCOL_MAX;+ghcb->protocol_version=ghcb_version;ghcb->ghcb_usage=GHCB_DEFAULT_USAGE;ghcb_set_sw_exit_code(ghcb,exit_code);
From: Borislav Petkov <redacted>
Carve it out so that it is abstracted out of the main boot path. All
other encrypted guest-relevant processing should be placed in there.
No functional changes.
Signed-off-by: Borislav Petkov <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/head64.c | 60 +++++++++++++++++++++-------------------
1 file changed, 31 insertions(+), 29 deletions(-)
@@ -126,6 +126,36 @@ static bool __head check_la57_support(unsigned long physaddr)}#endif+staticunsignedlongsme_postprocess_startup(structboot_params*bp,pmdval_t*pmd)+{+unsignedlongvaddr,vaddr_end;+inti;++/* Encrypt the kernel and related (if SME is active) */+sme_encrypt_kernel(bp);++/*+*Clearthememoryencryptionmaskfromthe.bss..decryptedsection.+*Thebsssectionwillbememsettozerolaterintheinitializationso+*thereisnoneedtozeroitafterchangingthememoryencryption+*attribute.+*/+if(sme_get_me_mask()){+vaddr=(unsignedlong)__start_bss_decrypted;+vaddr_end=(unsignedlong)__end_bss_decrypted;+for(;vaddr<vaddr_end;vaddr+=PMD_SIZE){+i=pmd_index(vaddr);+pmd[i]-=sme_get_me_mask();+}+}++/*+*ReturntheSMEencryptionmask(ifSMEisactive)tobeusedasa+*modifierfortheinitialpgdirentryprogrammedintoCR3.+*/+returnsme_get_me_mask();+}+/* Code in __startup_64() can be relocated during execution, but the compiler*doesn'thavetogeneratePC-relativerelocationswhenaccessingglobalsfrom*thatfunction.Clangactuallydoesnotgeneratethem,whichleadsto
@@ -135,7 +165,6 @@ static bool __head check_la57_support(unsigned long physaddr)unsignedlong__head__startup_64(unsignedlongphysaddr,structboot_params*bp){-unsignedlongvaddr,vaddr_end;unsignedlongload_delta,*p;unsignedlongpgtable_flags;pgdval_t*pgd;
@@ -276,34 +305,7 @@ unsigned long __head __startup_64(unsigned long physaddr,*/*fixup_long(&phys_base,physaddr)+=load_delta-sme_get_me_mask();-/* Encrypt the kernel and related (if SME is active) */-sme_encrypt_kernel(bp);--/*-*Clearthememoryencryptionmaskfromthe.bss..decryptedsection.-*Thebsssectionwillbememsettozerolaterintheinitializationso-*thereisnoneedtozeroitafterchangingthememoryencryption-*attribute.-*-*Thisisearlycode,useanopencodedcheckforSMEinsteadof-*usingcc_platform_has().Thiseliminatesworriesaboutremoving-*instrumentationorcheckingboot_cpu_datainthecc_platform_has()-*function.-*/-if(sme_get_me_mask()){-vaddr=(unsignedlong)__start_bss_decrypted;-vaddr_end=(unsignedlong)__end_bss_decrypted;-for(;vaddr<vaddr_end;vaddr+=PMD_SIZE){-i=pmd_index(vaddr);-pmd[i]-=sme_get_me_mask();-}-}--/*-*ReturntheSMEencryptionmask(ifSMEisactive)tobeusedasa-*modifierfortheinitialpgdirentryprogrammedintoCR3.-*/-returnsme_get_me_mask();+returnsme_postprocess_startup(bp,pmd);}unsignedlong__startup_secondary_64(void)
From: Michael Roth <redacted>
Generally access to MSR_AMD64_SEV is only safe if the 0x8000001F CPUID
leaf indicates SEV support. With SEV-SNP, CPUID responses from the
hypervisor are not considered trustworthy, particularly for 0x8000001F.
SEV-SNP provides a firmware-validated CPUID table to use as an
alternative, but prior to checking MSR_AMD64_SEV there are no
guarantees that this is even an SEV-SNP guest.
Rather than relying on these CPUID values early on, allow SEV-ES and
SEV-SNP guests to instead use a cpuid instruction to trigger a #VC and
have it cache MSR_AMD64_SEV in sev_status, since it is known to be safe
to access MSR_AMD64_SEV if a #VC has triggered.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/sev-shared.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
Version 2 of GHCB specification introduced advertisement of a features
that are supported by the hypervisor. Add support to query the HV
features on boot.
Version 2 of GHCB specification adds several new NAEs, most of them are
optional except the hypervisor feature. Now that hypervisor feature NAE
is implemented, so bump the GHCB maximum support protocol version.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev-common.h | 3 +++
arch/x86/include/asm/sev.h | 2 +-
arch/x86/include/uapi/asm/svm.h | 2 ++
arch/x86/kernel/sev-shared.c | 30 ++++++++++++++++++++++++++++++
4 files changed, 36 insertions(+), 1 deletion(-)
@@ -23,6 +23,9 @@*/staticu16__ro_after_initghcb_version;+/* Bitmap of SEV features supported by the hypervisor */+staticu64__ro_after_initsev_hv_features;+staticbool__initsev_es_check_cpu_features(void){if(!has_cpuflag(X86_FEATURE_RDRAND)){
@@ -48,6 +51,30 @@ static void __noreturn sev_es_terminate(unsigned int set, unsigned int reason)asmvolatile("hlt\n":::"memory");}+/*+*ThehypervisorfeaturesareavailablefromGHCBversion2onward.+*/+staticboolget_hv_features(void)+{+u64val;++sev_hv_features=0;++if(ghcb_version<2)+returnfalse;++sev_es_wr_ghcb_msr(GHCB_MSR_HV_FT_REQ);+VMGEXIT();++val=sev_es_rd_ghcb_msr();+if(GHCB_RESP_CODE(val)!=GHCB_MSR_HV_FT_RESP)+returnfalse;++sev_hv_features=GHCB_MSR_HV_FT_RESP_VAL(val);++returntrue;+}+staticboolsev_es_negotiate_protocol(void){u64val;
Version 2 of the GHCB specification added the advertisement of features
that are supported by the hypervisor. If hypervisor supports the SEV-SNP
then it must set the SEV-SNP features bit to indicate that the base
SEV-SNP is supported.
Check the SEV-SNP feature while establishing the GHCB, if failed,
terminate the guest.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/sev.c | 16 ++++++++++++++--
arch/x86/include/asm/sev-common.h | 3 +++
arch/x86/kernel/sev.c | 8 ++++++--
3 files changed, 23 insertions(+), 4 deletions(-)
@@ -631,12 +631,16 @@ static enum es_result vc_handle_msr(struct ghcb *ghcb, struct es_em_ctxt *ctxt)*Thisfunctionrunsonthefirst#VCexceptionafterthekernel*switchedtovirtualaddresses.*/-staticbool__initsev_es_setup_ghcb(void)+staticbool__initsetup_ghcb(void){/* First make sure the hypervisor talks a supported protocol. */if(!sev_es_negotiate_protocol())returnfalse;+/* If SNP is active, make sure that hypervisor supports the feature. */+if(cc_platform_has(CC_ATTR_SEV_SNP)&&!(sev_hv_features&GHCB_HV_FT_SNP))+sev_es_terminate(SEV_TERM_SET_GEN,GHCB_SNP_UNSUPPORTED);+/**Cleartheboot_ghcb.Thefirstexceptioncomesinbeforethebss*sectioniscleared.
@@ -1444,7 +1448,7 @@ bool __init handle_vc_boot_ghcb(struct pt_regs *regs)enumes_resultresult;/* Do initial setup or terminate the guest */-if(unlikely(boot_ghcb==NULL&&!sev_es_setup_ghcb()))+if(unlikely(!boot_ghcb&&!setup_ghcb()))sev_es_terminate(SEV_TERM_SET_GEN,GHCB_SEV_ES_GEN_REQ);vc_ghcb_invalidate(boot_ghcb);
An SNP-active guest uses the PVALIDATE instruction to validate or
rescind the validation of a guest page’s RMP entry. Upon completion,
a return code is stored in EAX and rFLAGS bits are set based on the
return code. If the instruction completed successfully, the CF
indicates if the content of the RMP were changed or not.
See AMD APM Volume 3 for additional details.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev.h | 21 +++++++++++++++++++++
1 file changed, 21 insertions(+)
Many of the integrity guarantees of SEV-SNP are enforced through the
Reverse Map Table (RMP). Each RMP entry contains the GPA at which a
particular page of DRAM should be mapped. The VMs can request the
hypervisor to add pages in the RMP table via the Page State Change VMGEXIT
defined in the GHCB specification. Inside each RMP entry is a Validated
flag; this flag is automatically cleared to 0 by the CPU hardware when a
new RMP entry is created for a guest. Each VM page can be either
validated or invalidated, as indicated by the Validated flag in the RMP
entry. Memory access to a private page that is not validated generates
a #VC. A VM must use PVALIDATE instruction to validate the private page
before using it.
To maintain the security guarantee of SEV-SNP guests, when transitioning
pages from private to shared, the guest must invalidate the pages before
asking the hypervisor to change the page state to shared in the RMP table.
After the pages are mapped private in the page table, the guest must issue
a page state change VMGEXIT to make the pages private in the RMP table and
validate it.
On boot, BIOS should have validated the entire system memory. During
the kernel decompression stage, the VC handler uses the
set_memory_decrypted() to make the GHCB page shared (i.e clear encryption
attribute). And while exiting from the decompression, it calls the
set_page_encrypted() to make the page private.
Add sev_snp_set_page_{private,shared}() helper that is used by the
set_memory_{decrypt,encrypt}() to change the page state in the RMP table.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/ident_map_64.c | 18 ++++++++++-
arch/x86/boot/compressed/misc.h | 6 ++++
arch/x86/boot/compressed/sev.c | 41 +++++++++++++++++++++++++
arch/x86/include/asm/sev-common.h | 26 ++++++++++++++++
4 files changed, 90 insertions(+), 1 deletion(-)
@@ -154,6 +154,47 @@ static bool is_vmpl0(void)returntrue;}+staticvoid__page_state_change(unsignedlongpaddr,enumpsc_opop)+{+u64val;++if(!sev_snp_enabled())+return;++/*+*Ifprivate->sharedtheninvalidatethepagebeforerequestingthe+*statechangeintheRMPtable.+*/+if(op==SNP_PAGE_STATE_SHARED&&pvalidate(paddr,RMP_PG_SIZE_4K,0))+sev_es_terminate(SEV_TERM_SET_LINUX,GHCB_TERM_PVALIDATE);++/* Issue VMGEXIT to change the page state in RMP table. */+sev_es_wr_ghcb_msr(GHCB_MSR_PSC_REQ_GFN(paddr>>PAGE_SHIFT,op));+VMGEXIT();++/* Read the response of the VMGEXIT. */+val=sev_es_rd_ghcb_msr();+if((GHCB_RESP_CODE(val)!=GHCB_MSR_PSC_RESP)||GHCB_MSR_PSC_RESP_VAL(val))+sev_es_terminate(SEV_TERM_SET_LINUX,GHCB_TERM_PSC);++/*+*NowthatpageisaddedintheRMPtable,validateitsothatitis+*consistentwiththeRMPentry.+*/+if(op==SNP_PAGE_STATE_PRIVATE&&pvalidate(paddr,RMP_PG_SIZE_4K,1))+sev_es_terminate(SEV_TERM_SET_LINUX,GHCB_TERM_PVALIDATE);+}++voidsnp_set_page_private(unsignedlongpaddr)+{+__page_state_change(paddr,SNP_PAGE_STATE_PRIVATE);+}++voidsnp_set_page_shared(unsignedlongpaddr)+{+__page_state_change(paddr,SNP_PAGE_STATE_SHARED);+}+staticbooldo_early_sev_setup(void){if(!sev_es_negotiate_protocol())
The SEV-SNP guest is required to perform GHCB GPA registration. This is
because the hypervisor may prefer that a guest use a consistent and/or
specific GPA for the GHCB associated with a vCPU. For more information,
see the GHCB specification.
If hypervisor can not work with the guest provided GPA then terminate the
guest boot.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/sev.c | 4 ++++
arch/x86/include/asm/sev-common.h | 13 +++++++++++++
arch/x86/kernel/sev-shared.c | 16 ++++++++++++++++
3 files changed, 33 insertions(+)
@@ -223,6 +223,10 @@ static bool do_early_sev_setup(void)/* Initialize lookup tables for the instruction decoder */inat_init_tables();+/* SEV-SNP guest requires the GHCB GPA must be registered */+if(sev_snp_enabled())+snp_register_ghcb_early(__pa(&boot_ghcb_page));+returntrue;}
@@ -75,6 +75,22 @@ static bool get_hv_features(void)returntrue;}+staticvoidsnp_register_ghcb_early(unsignedlongpaddr)+{+unsignedlongpfn=paddr>>PAGE_SHIFT;+u64val;++sev_es_wr_ghcb_msr(GHCB_MSR_REG_GPA_REQ_VAL(pfn));+VMGEXIT();++val=sev_es_rd_ghcb_msr();++/* If the response GPA is not ours then abort the guest */+if((GHCB_RESP_CODE(val)!=GHCB_MSR_REG_GPA_RESP)||+(GHCB_MSR_REG_GPA_RESP_VAL(val)!=pfn))+sev_es_terminate(SEV_TERM_SET_LINUX,GHCB_TERM_REGISTER);+}+staticboolsev_es_negotiate_protocol(void){u64val;
Virtual Machine Privilege Level (VMPL) is an optional feature in the
SEV-SNP architecture, which allows a guest VM to divide its address space
into four levels. The level can be used to provide the hardware isolated
abstraction layers with a VM. The VMPL0 is the highest privilege, and
VMPL3 is the least privilege. Certain operations must be done by the VMPL0
software, such as:
* Validate or invalidate memory range (PVALIDATE instruction)
* Allocate VMSA page (RMPADJUST instruction when VMSA=1)
The initial SEV-SNP support assumes that the guest kernel is running on
VMPL0. Let's add a check to make sure that kernel is running at VMPL0
before continuing the boot. There is no easy method to query the current
VMPL level, so use the RMPADJUST instruction to determine whether its
booted at the VMPL0.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/sev.c | 41 ++++++++++++++++++++++++++++---
arch/x86/include/asm/sev-common.h | 1 +
arch/x86/include/asm/sev.h | 3 +++
3 files changed, 42 insertions(+), 3 deletions(-)
The encryption attribute for the bss.decrypted region is cleared in the
initial page table build. This is because the section contains the data
that need to be shared between the guest and the hypervisor.
When SEV-SNP is active, just clearing the encryption attribute in the
page table is not enough. The page state need to be updated in the RMP
table.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/head64.c | 7 +++++++
1 file changed, 7 insertions(+)
The hypervisor uses the sev_features field (offset 3B0h) in the Save State
Area to control the SEV-SNP guest features such as SNPActive, vTOM,
ReflectVC etc. An SEV-SNP guest can read the SEV_FEATURES fields through
the SEV_STATUS MSR.
While at it, update the dump_vmcb() to log the VMPL level.
See APM2 Table 15-34 and B-4 for more details.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/svm.h | 6 ++++--
arch/x86/kvm/svm/svm.c | 4 ++--
2 files changed, 6 insertions(+), 4 deletions(-)
The probe_roms() access the memory range (0xc0000 - 0x10000) to probe
various ROMs. The memory range is not part of the E820 system RAM
range. The memory range is mapped as private (i.e encrypted) in page
table.
When SEV-SNP is active, all the private memory must be validated before
the access. The ROM range was not part of E820 map, so the guest BIOS
did not validate it. An access to invalidated memory will cause a VC
exception. The guest does not support handling not-validated VC exception
yet, so validate the ROM memory regions before it is accessed.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/probe_roms.c | 13 ++++++++++++-
1 file changed, 12 insertions(+), 1 deletion(-)
@@ -197,11 +198,21 @@ static int __init romchecksum(const unsigned char *rom, unsigned long length)void__initprobe_roms(void){-constunsignedchar*rom;unsignedlongstart,length,upper;+constunsignedchar*rom;unsignedcharc;inti;+/*+*TheROMmemoryisnotpartoftheE820systemRAMandisnotpre-validated+*bytheBIOS.ThekernelpagetablemapstheROMregionasencryptedmemory,+*theSEV-SNPrequirestheencryptedmemorymustbevalidatedbeforethe+*access.ValidatetheROMbeforeaccessingit.+*/+snp_prep_memory(video_rom_resource.start,+((system_rom_resource.end+1)-video_rom_resource.start),+SNP_PAGE_STATE_PRIVATE);+/* video rom */upper=adapter_rom_resources[0].start;for(start=video_rom_resource.start;start<upper;start+=2048){
The set_memory_{encrypt,decrypt}() are used for changing the pages
from decrypted (shared) to encrypted (private) and vice versa.
When SEV-SNP is active, the page state transition needs to go through
additional steps.
If the page is transitioned from shared to private, then perform the
following after the encryption attribute is set in the page table:
1. Issue the page state change VMGEXIT to add the memory region in
the RMP table.
2. Validate the memory region after the RMP entry is added.
To maintain the security guarantees, if the page is transitioned from
private to shared, then perform the following before encryption attribute
is removed from the page table:
1. Invalidate the page.
2. Issue the page state change VMGEXIT to remove the page from RMP table.
To change the page state in the RMP table, use the Page State Change
VMGEXIT defined in the GHCB specification.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev-common.h | 22 ++++
arch/x86/include/asm/sev.h | 4 +
arch/x86/include/uapi/asm/svm.h | 2 +
arch/x86/kernel/sev.c | 165 ++++++++++++++++++++++++++++++
arch/x86/mm/pat/set_memory.c | 15 +++
5 files changed, 208 insertions(+)
@@ -655,6 +655,171 @@ void __init snp_prep_memory(unsigned long paddr, unsigned int sz, enum psc_op opWARN(1,"invalid memory op %d\n",op);}+staticintvmgexit_psc(structsnp_psc_desc*desc)+{+intcur_entry,end_entry,ret;+structsnp_psc_desc*data;+structghcb_statestate;+structghcb*ghcb;+structpsc_hdr*hdr;+unsignedlongflags;++local_irq_save(flags);++ghcb=__sev_get_ghcb(&state);+if(unlikely(!ghcb))+panic("SEV-SNP: Failed to get GHCB\n");++/* Copy the input desc into GHCB shared buffer */+data=(structsnp_psc_desc*)ghcb->shared_buffer;+memcpy(ghcb->shared_buffer,desc,sizeof(*desc));++hdr=&data->hdr;+cur_entry=hdr->cur_entry;+end_entry=hdr->end_entry;++/*+*AspertheGHCBspecification,thehypervisorcanresumetheguest+*beforeprocessingalltheentries.Checkswhetheralltheentries+*areprocessed.Ifnot,thenkeepretrying.+*+*Thestragtegyhereistowaitforthehypervisortochangethepage+*stateintheRMPtablebeforeguestaccessthememorypages.Ifthe+*pagestatewasnotsuccessful,thenlatermemoryaccesswillresult+*inthecrash.+*/+while(hdr->cur_entry<=hdr->end_entry){+ghcb_set_sw_scratch(ghcb,(u64)__pa(data));++ret=sev_es_ghcb_hv_call(ghcb,NULL,SVM_VMGEXIT_PSC,0,0);++/*+*PageStateChangeVMGEXITcanpasserrorcodethrough+*exit_info_2.+*/+if(WARN(ret||ghcb->save.sw_exit_info_2,+"SEV-SNP: PSC failed ret=%d exit_info_2=%llx\n",+ret,ghcb->save.sw_exit_info_2)){+ret=1;+gotoout;+}++/*+*Sanitycheckthatentryprocessingisnotgoingbackward.+*Thiswillhappenonlyifhypervisoristrickingus.+*/+if(WARN(hdr->end_entry>end_entry||cur_entry>hdr->cur_entry,+"SEV-SNP: PSC processing going backward, end_entry %d (got %d) cur_entry %d (got %d)\n",+end_entry,hdr->end_entry,cur_entry,hdr->cur_entry)){+ret=1;+gotoout;+}++/* Verify that reserved bit is not set */+if(WARN(hdr->reserved,"Reserved bit is set in the PSC header\n")){+ret=1;+gotoout;+}+}++out:+__sev_put_ghcb(&state);+local_irq_restore(flags);++return0;+}++staticvoid__set_page_state(structsnp_psc_desc*data,unsignedlongvaddr,+unsignedlongvaddr_end,intop)+{+structpsc_hdr*hdr;+structpsc_entry*e;+unsignedlongpfn;+inti;++hdr=&data->hdr;+e=data->entries;++memset(data,0,sizeof(*data));+i=0;++while(vaddr<vaddr_end){+if(is_vmalloc_addr((void*)vaddr))+pfn=vmalloc_to_pfn((void*)vaddr);+else+pfn=__pa(vaddr)>>PAGE_SHIFT;++e->gfn=pfn;+e->operation=op;+hdr->end_entry=i;++/*+*TheGHCBspecificationprovidestheflexibilityto+*useeither4Kor2MBpagesizeintheRMPtable.+*ThecurrentSNPsupportdoesnotkeeptrackofthe+*pagesizeusedintheRMPtable.Toavoidthe+*overlaprequest,usethe4KpagesizeintheRMP+*table.+*/+e->pagesize=RMP_PG_SIZE_4K;++vaddr=vaddr+PAGE_SIZE;+e++;+i++;+}++if(vmgexit_psc(data))+sev_es_terminate(SEV_TERM_SET_LINUX,GHCB_TERM_PSC);+}++staticvoidset_page_state(unsignedlongvaddr,unsignedintnpages,intop)+{+unsignedlongvaddr_end,next_vaddr;+structsnp_psc_desc*desc;++vaddr=vaddr&PAGE_MASK;+vaddr_end=vaddr+(npages<<PAGE_SHIFT);++desc=kmalloc(sizeof(*desc),GFP_KERNEL_ACCOUNT);+if(!desc)+panic("SEV-SNP: failed to allocate memory for PSC descriptor\n");++while(vaddr<vaddr_end){+/*+*Calculatethelastvaddrthatcanbefitinone+*structsnp_psc_desc.+*/+next_vaddr=min_t(unsignedlong,vaddr_end,+(VMGEXIT_PSC_MAX_ENTRY*PAGE_SIZE)+vaddr);++__set_page_state(desc,vaddr,next_vaddr,op);++vaddr=next_vaddr;+}++kfree(desc);+}++voidsnp_set_memory_shared(unsignedlongvaddr,unsignedintnpages)+{+if(!cc_platform_has(CC_ATTR_SEV_SNP))+return;++pvalidate_pages(vaddr,npages,0);++set_page_state(vaddr,npages,SNP_PAGE_STATE_SHARED);+}++voidsnp_set_memory_private(unsignedlongvaddr,unsignedintnpages)+{+if(!cc_platform_has(CC_ATTR_SEV_SNP))+return;++set_page_state(vaddr,npages,SNP_PAGE_STATE_PRIVATE);++pvalidate_pages(vaddr,npages,1);+}+intsev_es_setup_ap_jump_table(structreal_mode_header*rmh){u16startup_cs,startup_ip;
@@ -2010,8 +2011,22 @@ static int __set_memory_enc_dec(unsigned long addr, int numpages, bool enc)*/cpa_flush(&cpa,!this_cpu_has(X86_FEATURE_SME_COHERENT));+/*+*TomaintainthesecurityguranteesofSEV-SNPguestinvalidatethememory+*beforeclearingtheencryptionattribute.+*/+if(!enc)+snp_set_memory_shared(addr,numpages);+ret=__change_page_attr_set_clr(&cpa,1);+/*+*Nowthatmemoryismappedencryptedinthepagetable,validateit+*sothatisconsistentwiththeabovepagestate.+*/+if(!ret&&enc)+snp_set_memory_private(addr,numpages);+/**Afterchangingtheencryptionattribute,weneedtoflushTLBsagain*incaseanyspeculativeTLBcachingoccurred(butnoneedtoflush
The SEV-SNP guest is required to perform GHCB GPA registration. This is
because the hypervisor may prefer that a guest use a consistent and/or
specific GPA for the GHCB associated with a vCPU. For more information,
see the GHCB specification section GHCB GPA Registration.
During the boot, init_ghcb() allocates a per-cpu GHCB page. On very first
VC exception, the exception handler switch to using the per-cpu GHCB page
allocated during the init_ghcb(). The GHCB page must be registered in
the current vcpu context.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/sev.c | 124 +++++++++++++++++++++++++-----------------
1 file changed, 75 insertions(+), 49 deletions(-)
@@ -160,55 +167,6 @@ void noinstr __sev_es_ist_exit(void)this_cpu_write(cpu_tss_rw.x86_tss.ist[IST_INDEX_VC],*(unsignedlong*)ist);}-/*-*Nothingshallinterruptthiscodepathwhileholdingtheper-CPU-*GHCB.ThebackupGHCBisonlyforNMIsinterruptingthispath.-*-*Callersmustdisablelocalinterruptsaroundit.-*/-staticnoinstrstructghcb*__sev_get_ghcb(structghcb_state*state)-{-structsev_es_runtime_data*data;-structghcb*ghcb;--WARN_ON(!irqs_disabled());--data=this_cpu_read(runtime_data);-ghcb=&data->ghcb_page;--if(unlikely(data->ghcb_active)){-/* GHCB is already in use - save its contents */--if(unlikely(data->backup_ghcb_active)){-/*-*Backup-GHCBisalsoalreadyinuse.Thereisnoway-*tocontinueheresojustkillthemachine.Tomake-*panic()work,markGHCBsinactivesothatmessages-*canbeprintedout.-*/-data->ghcb_active=false;-data->backup_ghcb_active=false;--instrumentation_begin();-panic("Unable to handle #VC exception! GHCB and Backup GHCB are already in use");-instrumentation_end();-}--/* Mark backup_ghcb active before writing to it */-data->backup_ghcb_active=true;--state->ghcb=&data->backup_ghcb;--/* Backup GHCB content */-*state->ghcb=*ghcb;-}else{-state->ghcb=NULL;-data->ghcb_active=true;-}--returnghcb;-}-/* Needed in vc_early_forward_exception */voiddo_early_exception(structpt_regs*regs,inttrapnr);
@@ -464,6 +422,69 @@ static enum es_result vc_slow_virt_to_phys(struct ghcb *ghcb, struct es_em_ctxt/* Include code shared with pre-decompression boot stage */#include"sev-shared.c"+staticvoidsnp_register_ghcb(structsev_es_runtime_data*data,unsignedlongpaddr)+{+if(data->snp_ghcb_registered)+return;++snp_register_ghcb_early(paddr);++data->snp_ghcb_registered=true;+}++/*+*Nothingshallinterruptthiscodepathwhileholdingtheper-CPU+*GHCB.ThebackupGHCBisonlyforNMIsinterruptingthispath.+*+*Callersmustdisablelocalinterruptsaroundit.+*/+staticnoinstrstructghcb*__sev_get_ghcb(structghcb_state*state)+{+structsev_es_runtime_data*data;+structghcb*ghcb;++WARN_ON(!irqs_disabled());++data=this_cpu_read(runtime_data);+ghcb=&data->ghcb_page;++if(unlikely(data->ghcb_active)){+/* GHCB is already in use - save its contents */++if(unlikely(data->backup_ghcb_active)){+/*+*Backup-GHCBisalsoalreadyinuse.Thereisnoway+*tocontinueheresojustkillthemachine.Tomake+*panic()work,markGHCBsinactivesothatmessages+*canbeprintedout.+*/+data->ghcb_active=false;+data->backup_ghcb_active=false;++instrumentation_begin();+panic("Unable to handle #VC exception! GHCB and Backup GHCB are already in use");+instrumentation_end();+}++/* Mark backup_ghcb active before writing to it */+data->backup_ghcb_active=true;++state->ghcb=&data->backup_ghcb;++/* Backup GHCB content */+*state->ghcb=*ghcb;+}else{+state->ghcb=NULL;+data->ghcb_active=true;+}++/* SEV-SNP guest requires that GHCB must be registered. */+if(cc_platform_has(CC_ATTR_SEV_SNP))+snp_register_ghcb(data,__pa(ghcb));++returnghcb;+}+staticnoinstrvoid__sev_put_ghcb(structghcb_state*state){structsev_es_runtime_data*data;
@@ -650,6 +671,10 @@ static bool __init setup_ghcb(void)/* Alright - Make the boot-ghcb public */boot_ghcb=&boot_ghcb_page;+/* SEV-SNP guest requires that GHCB GPA must be registered. */+if(cc_platform_has(CC_ATTR_SEV_SNP))+snp_register_ghcb_early(__pa(&boot_ghcb_page));+returntrue;}
From: Tom Lendacky <thomas.lendacky@amd.com>
The save area for SEV-ES/SEV-SNP guests, as used by the hardware, is
different from the save area of a non SEV-ES/SEV-SNP guest.
This is the first step in defining the multiple save areas to keep them
separate and ensuring proper operation amongst the different types of
guests. Create an SEV-ES/SEV-SNP save area and adjust usage to the new
save area definition where needed.
Signed-off-by: Tom Lendacky <thomas.lendacky@amd.com>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/svm.h | 83 +++++++++++++++++++++++++++++---------
arch/x86/kvm/svm/sev.c | 24 +++++------
arch/x86/kvm/svm/svm.h | 2 +-
3 files changed, 77 insertions(+), 32 deletions(-)
@@ -227,6 +227,7 @@ struct vmcb_seg {u64base;}__packed;+/* Save area definition for legacy and SEV-MEM guests */structvmcb_save_area{structvmcb_seges;structvmcb_segcs;
@@ -243,8 +244,58 @@ struct vmcb_save_area {u8cpl;u8reserved_2[4];u64efer;+u8reserved_3[112];+u64cr4;+u64cr3;+u64cr0;+u64dr7;+u64dr6;+u64rflags;+u64rip;+u8reserved_4[88];+u64rsp;+u64s_cet;+u64ssp;+u64isst_addr;+u64rax;+u64star;+u64lstar;+u64cstar;+u64sfmask;+u64kernel_gs_base;+u64sysenter_cs;+u64sysenter_esp;+u64sysenter_eip;+u64cr2;+u8reserved_5[32];+u64g_pat;+u64dbgctl;+u64br_from;+u64br_to;+u64last_excp_from;+u64last_excp_to;+u8reserved_6[72];+u32spec_ctrl;/* Guest version of SPEC_CTRL at 0x2E0 */+}__packed;++/* Save area definition for SEV-ES and SEV-SNP guests */+structsev_es_save_area{+structvmcb_seges;+structvmcb_segcs;+structvmcb_segss;+structvmcb_segds;+structvmcb_segfs;+structvmcb_seggs;+structvmcb_seggdtr;+structvmcb_segldtr;+structvmcb_segidtr;+structvmcb_segtr;+u8reserved_1[43];+u8cpl;+u8reserved_2[4];+u64efer;u8reserved_3[104];-u64xss;/* Valid for SEV-ES only */+u64xss;u64cr4;u64cr3;u64cr0;
@@ -272,22 +323,14 @@ struct vmcb_save_area {u64br_to;u64last_excp_from;u64last_excp_to;--/*-*Thefollowingpartofthesaveareaisvalidonlyfor-*SEV-ESguestswhenreferencedthroughtheGHCBorfor-*savingtothehostsavearea.-*/-u8reserved_7[72];-u32spec_ctrl;/* Guest version of SPEC_CTRL at 0x2E0 */-u8reserved_7b[4];+u8reserved_7[80];u32pkru;-u8reserved_7a[20];-u64reserved_8;/* rax already available at 0x01f8 */+u8reserved_9[20];+u64reserved_10;/* rax already available at 0x01f8 */u64rcx;u64rdx;u64rbx;-u64reserved_9;/* rsp already available at 0x01d8 */+u64reserved_11;/* rsp already available at 0x01d8 */u64rbp;u64rsi;u64rdi;
@@ -551,12 +551,20 @@ static int sev_launch_update_data(struct kvm *kvm, struct kvm_sev_cmd *argp)staticintsev_es_sync_vmsa(structvcpu_svm*svm){-structvmcb_save_area*save=&svm->vmcb->save;+structsev_es_save_area*save=svm->vmsa;/* Check some debug related fields before encrypting the VMSA */-if(svm->vcpu.guest_debug||(save->dr7&~DR7_FIXED_1))+if(svm->vcpu.guest_debug||(svm->vmcb->save.dr7&~DR7_FIXED_1))return-EINVAL;+/*+*SEV-ESwilluseaVMSAthatispointedtobytheVMCB,not+*thetraditionalVMSAthatispartoftheVMCB.Copythe+*traditionalVMSAasithasbeenbuiltsofar(inprep+*forLAUNCH_UPDATE_VMSA)tobetheinitialSEV-ESstate.+*/+memcpy(save,&svm->vmcb->save,sizeof(svm->vmcb->save));+/* Sync registgers */save->rax=svm->vcpu.arch.regs[VCPU_REGS_RAX];save->rbx=svm->vcpu.arch.regs[VCPU_REGS_RBX];
@@ -584,14 +592,6 @@ static int sev_es_sync_vmsa(struct vcpu_svm *svm)save->xss=svm->vcpu.arch.ia32_xss;save->dr6=svm->vcpu.arch.dr6;-/*-*SEV-ESwilluseaVMSAthatispointedtobytheVMCB,not-*thetraditionalVMSAthatispartoftheVMCB.Copythe-*traditionalVMSAasithasbeenbuiltsofar(inprep-*forLAUNCH_UPDATE_VMSA)tobetheinitialSEV-ESstate.-*/-memcpy(svm->vmsa,save,sizeof(*save));-return0;}
@@ -2655,7 +2655,7 @@ void sev_es_prepare_guest_switch(struct vcpu_svm *svm, unsigned int cpu)vmsave(__sme_page_pa(sd->save_area));/* XCR0 is restored on VMEXIT, save the current host value */-hostsa=(structvmcb_save_area*)(page_address(sd->save_area)+0x400);+hostsa=(structsev_es_save_area*)(page_address(sd->save_area)+0x400);hostsa->xcr0=xgetbv(XCR_XFEATURE_ENABLED_MASK);/* PKRU is restored on VMEXIT, save the current host value */
From: Tom Lendacky <thomas.lendacky@amd.com>
The initial implementation of the GHCB spec was based on trying to keep
the register state offsets the same relative to the VM save area. However,
the save area for SEV-ES has changed within the hardware causing the
relation between the SEV-ES save area to change relative to the GHCB save
area.
This is the second step in defining the multiple save areas to keep them
separate and ensuring proper operation amongst the different types of
guests. Create a GHCB save area that matches the GHCB specification.
Signed-off-by: Tom Lendacky <thomas.lendacky@amd.com>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/svm.h | 48 +++++++++++++++++++++++++++++++++++---
1 file changed, 45 insertions(+), 3 deletions(-)
From: Tom Lendacky <thomas.lendacky@amd.com>
This is the final step in defining the multiple save areas to keep them
separate and ensuring proper operation amongst the different types of
guests. Update the SEV-ES/SEV-SNP save area to match the APM. This save
area will be used for the upcoming SEV-SNP AP Creation NAE event support.
Signed-off-by: Tom Lendacky <thomas.lendacky@amd.com>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/svm.h | 66 +++++++++++++++++++++++++++++---------
1 file changed, 50 insertions(+), 16 deletions(-)
@@ -325,12 +341,12 @@ struct sev_es_save_area {u64last_excp_to;u8reserved_7[80];u32pkru;-u8reserved_9[20];-u64reserved_10;/* rax already available at 0x01f8 */+u8reserved_8[20];+u64reserved_9;/* rax already available at 0x01f8 */u64rcx;u64rdx;u64rbx;-u64reserved_11;/* rsp already available at 0x01d8 */+u64reserved_10;/* rsp already available at 0x01d8 */u64rbp;u64rsi;u64rdi;
@@ -342,16 +358,34 @@ struct sev_es_save_area {u64r13;u64r14;u64r15;-u8reserved_12[16];-u64sw_exit_code;-u64sw_exit_info_1;-u64sw_exit_info_2;-u64sw_scratch;+u8reserved_11[16];+u64guest_exit_info_1;+u64guest_exit_info_2;+u64guest_exit_int_info;+u64guest_nrip;u64sev_features;-u8reserved_13[48];+u64vintr_ctrl;+u64guest_exit_code;+u64virtual_tom;+u64tlb_id;+u64pcpu_id;+u64event_inj;u64xcr0;-u8valid_bitmap[16];-u64x87_state_gpa;+u8reserved_12[16];++/* Floating point area */+u64x87_dp;+u32mxcsr;+u16x87_ftw;+u16x87_fsw;+u16x87_fcw;+u16x87_fop;+u16x87_ds;+u16x87_cs;+u64x87_rip;+u8fpreg_x87[80];+u8fpreg_xmm[256];+u8fpreg_ymm[256];}__packed;structghcb_save_area{
From: Tom Lendacky <thomas.lendacky@amd.com>
To provide a more secure way to start APs under SEV-SNP, use the SEV-SNP
AP Creation NAE event. This allows for guest control over the AP register
state rather than trusting the hypervisor with the SEV-ES Jump Table
address.
During native_smp_prepare_cpus(), invoke an SEV-SNP function that, if
SEV-SNP is active, will set/override apic->wakeup_secondary_cpu. This
will allow the SEV-SNP AP Creation NAE event method to be used to boot
the APs. As a result of installing the override when SEV-SNP is active,
this method of starting the APs becomes the required method. The override
function will fail to start the AP if the hypervisor does not have
support for AP creation.
Signed-off-by: Tom Lendacky <thomas.lendacky@amd.com>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev-common.h | 1 +
arch/x86/include/asm/sev.h | 4 +
arch/x86/include/uapi/asm/svm.h | 5 +
arch/x86/kernel/sev.c | 205 ++++++++++++++++++++++++++++++
arch/x86/kernel/smpboot.c | 3 +
5 files changed, 218 insertions(+)
@@ -820,6 +824,207 @@ void snp_set_memory_private(unsigned long vaddr, unsigned int npages)pvalidate_pages(vaddr,npages,1);}+staticintrmpadjust(void*va,boolvmsa)+{+u64attrs;+interr;++/*+*TheRMPADJUSTinstructionisusedtosetorcleartheVMSAbitfor+*apage.AchangetotheVMSAbitisonlyperformedwhenrunning+*atVMPL0andisignoredatotherVMPLlevels.Iftoolowofatarget+*VMPLlevelisspecified,theinstructioncansucceedwithoutchanging+*theVMSAbitshouldthekernelnotbeinVMPL0.UsingatargetVMPL+*levelof1willreturnaFAIL_PERMISSIONerrorifthekernelisnot+*atVMPL0,thusensuringthattheVMSAbithasbeenproperlysetwhen+*noerrorisreturned.+*/+attrs=1;+if(vmsa)+attrs|=RMPADJUST_VMSA_PAGE_BIT;++/* Instruction mnemonic supported in binutils versions v2.36 and later */+asmvolatile(".byte 0xf3,0x0f,0x01,0xfe\n\t"+:"=a"(err)+:"a"(va),"c"(RMP_PG_SIZE_4K),"d"(attrs)+:"memory","cc");++returnerr;+}++#define __ATTR_BASE (SVM_SELECTOR_P_MASK | SVM_SELECTOR_S_MASK)+#define INIT_CS_ATTRIBS (__ATTR_BASE | SVM_SELECTOR_READ_MASK | SVM_SELECTOR_CODE_MASK)+#define INIT_DS_ATTRIBS (__ATTR_BASE | SVM_SELECTOR_WRITE_MASK)++#define INIT_LDTR_ATTRIBS (SVM_SELECTOR_P_MASK | 2)+#define INIT_TR_ATTRIBS (SVM_SELECTOR_P_MASK | 3)++staticintwakeup_cpu_via_vmgexit(intapic_id,unsignedlongstart_ip)+{+structsev_es_save_area*cur_vmsa,*vmsa;+structghcb_statestate;+unsignedlongflags;+structghcb*ghcb;+intcpu,err,ret;+u8sipi_vector;+u64cr4;++if((sev_hv_features&GHCB_HV_FT_SNP_AP_CREATION)!=GHCB_HV_FT_SNP_AP_CREATION)+return-EOPNOTSUPP;++/*+*VerifythedesiredstartIPagainsttheknowntrampolinestartIP+*tocatchanyfuturenewtrampolinesthatmaybeintroducedthat+*wouldrequireanewprotectedguestentrypoint.+*/+if(WARN_ONCE(start_ip!=real_mode_header->trampoline_start,+"Unsupported SEV-SNP start_ip: %lx\n",start_ip))+return-EINVAL;++/* Override start_ip with known protected guest start IP */+start_ip=real_mode_header->sev_es_trampoline_start;++/* Find the logical CPU for the APIC ID */+for_each_present_cpu(cpu){+if(arch_match_cpu_phys_id(cpu,apic_id))+break;+}+if(cpu>=nr_cpu_ids)+return-EINVAL;++cur_vmsa=per_cpu(snp_vmsa,cpu);++/*+*AnewVMSAiscreatedeachtimebecausethereisnoguaranteethat+*thecurrentVMSAisthekernelsorthatthevCPUisnotrunning.If+*anattemptwasdonetousethecurrentVMSAwitharunningvCPU,a+*#VMEXITofthatvCPUwouldwipeoutallofthesettingsbeingdone+*here.+*/+vmsa=(structsev_es_save_area*)get_zeroed_page(GFP_KERNEL);+if(!vmsa)+return-ENOMEM;++/* CR4 should maintain the MCE value */+cr4=native_read_cr4()&X86_CR4_MCE;++/* Set the CS value based on the start_ip converted to a SIPI vector */+sipi_vector=(start_ip>>12);+vmsa->cs.base=sipi_vector<<12;+vmsa->cs.limit=0xffff;+vmsa->cs.attrib=INIT_CS_ATTRIBS;+vmsa->cs.selector=sipi_vector<<8;++/* Set the RIP value based on start_ip */+vmsa->rip=start_ip&0xfff;++/* Set VMSA entries to the INIT values as documented in the APM */+vmsa->ds.limit=0xffff;+vmsa->ds.attrib=INIT_DS_ATTRIBS;+vmsa->es=vmsa->ds;+vmsa->fs=vmsa->ds;+vmsa->gs=vmsa->ds;+vmsa->ss=vmsa->ds;++vmsa->gdtr.limit=0xffff;+vmsa->ldtr.limit=0xffff;+vmsa->ldtr.attrib=INIT_LDTR_ATTRIBS;+vmsa->idtr.limit=0xffff;+vmsa->tr.limit=0xffff;+vmsa->tr.attrib=INIT_TR_ATTRIBS;++vmsa->efer=0x1000;/* Must set SVME bit */+vmsa->cr4=cr4;+vmsa->cr0=0x60000010;+vmsa->dr7=0x400;+vmsa->dr6=0xffff0ff0;+vmsa->rflags=0x2;+vmsa->g_pat=0x0007040600070406ULL;+vmsa->xcr0=0x1;+vmsa->mxcsr=0x1f80;+vmsa->x87_ftw=0x5555;+vmsa->x87_fcw=0x0040;++/*+*SettheSNP-specificfieldsforthisVMSA:+*VMPLlevel+*SEV_FEATURES(matchestheSEVSTATUSMSRrightshifted2bits)+*/+vmsa->vmpl=0;+vmsa->sev_features=sev_status>>2;++/* Switch the page over to a VMSA page now that it is initialized */+ret=rmpadjust(vmsa,true);+if(ret){+pr_err("set VMSA page failed (%u)\n",ret);+free_page((unsignedlong)vmsa);++return-EINVAL;+}++/* Issue VMGEXIT AP Creation NAE event */+local_irq_save(flags);++ghcb=__sev_get_ghcb(&state);++vc_ghcb_invalidate(ghcb);+ghcb_set_rax(ghcb,vmsa->sev_features);+ghcb_set_sw_exit_code(ghcb,SVM_VMGEXIT_AP_CREATION);+ghcb_set_sw_exit_info_1(ghcb,((u64)apic_id<<32)|SVM_VMGEXIT_AP_CREATE);+ghcb_set_sw_exit_info_2(ghcb,__pa(vmsa));++sev_es_wr_ghcb_msr(__pa(ghcb));+VMGEXIT();++if(!ghcb_sw_exit_info_1_is_valid(ghcb)||+lower_32_bits(ghcb->save.sw_exit_info_1)){+pr_alert("SNP AP Creation error\n");+ret=-EINVAL;+}++__sev_put_ghcb(&state);++local_irq_restore(flags);++/* Perform cleanup if there was an error */+if(ret){+err=rmpadjust(vmsa,false);+if(err)+pr_err("clear VMSA page failed (%u), leaking page\n",err);+else+free_page((unsignedlong)vmsa);++vmsa=NULL;+}++/* Free up any previous VMSA page */+if(cur_vmsa){+err=rmpadjust(cur_vmsa,false);+if(err)+pr_err("clear VMSA page failed (%u), leaking page\n",err);+else+free_page((unsignedlong)cur_vmsa);+}++/* Record the current VMSA page */+per_cpu(snp_vmsa,cpu)=vmsa;++returnret;+}++voidsnp_set_wakeup_secondary_cpu(void)+{+if(!cc_platform_has(CC_ATTR_SEV_SNP))+return;++/*+*AlwayssetthisoverrideifSEV-SNPisenabled.Thismakesitthe+*requiredmethodtostartAPsunderSEV-SNP.Ifthehypervisordoes+*notsupportAPcreation,thennoAPswillbestarted.+*/+apic->wakeup_secondary_cpu=wakeup_cpu_via_vmgexit;+}+intsev_es_setup_ap_jump_table(structreal_mode_header*rmh){u16startup_cs,startup_ip;
From: Michael Roth <redacted>
As of commit 103a4908ad4d ("x86/head/64: Disable stack protection for
head$(BITS).o") kernel/head64.c is compiled with -fno-stack-protector
to allow a call to set_bringup_idt_handler(), which would otherwise
have stack protection enabled with CONFIG_STACKPROTECTOR_STRONG. While
sufficient for that case, there may still be issues with calls to any
external functions that were compiled with stack protection enabled that
in-turn make stack-protected calls, or if the exception handlers set up
by set_bringup_idt_handler() make calls to stack-protected functions.
As part of 103a4908ad4d, stack protection was also disabled for
kernel/head32.c as a precaution.
Subsequent patches for SEV-SNP CPUID validation support will introduce
both such cases. Attempting to disable stack protection for everything
in scope to address that is prohibitive since much of the code, like
SEV-ES #VC handler, is shared code that remains in use after boot and
could benefit from having stack protection enabled. Attempting to inline
calls is brittle and can quickly balloon out to library/helper code
where that's not really an option.
Instead, re-enable stack protection for head32.c/head64.c and make the
appropriate changes to ensure the segment used for the stack canary is
initialized in advance of any stack-protected C calls.
for head64.c:
- The BSP will enter from startup_64 and call into C code
(startup_64_setup_env) shortly after setting up the stack, which may
result in calls to stack-protected code. Set up %gs early to allow
for this safely.
- APs will enter from secondary_startup_64*, and %gs will be set up
soon after. There is one call to C code prior to this
(__startup_secondary_64), but it is only to fetch sme_me_mask, and
unlikely to be stack-protected, so leave things as they are, but add
a note about this in case things change in the future.
for head32.c:
- BSPs/APs will set %fs to __BOOT_DS prior to any C calls. In recent
kernels, the compiler is configured to access the stack canary at
%fs:__stack_chk_guard, which overlaps with the initial per-cpu
__stack_chk_guard variable in the initial/'master' .data..percpu
area. This is sufficient to allow access to the canary for use
during initial startup, so no changes are needed there.
Suggested-by: Joerg Roedel <redacted> #for 64-bit %gs set up
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/Makefile | 1 -
arch/x86/kernel/head_64.S | 24 ++++++++++++++++++++++++
2 files changed, 24 insertions(+), 1 deletion(-)
From: Michael Roth <redacted>
Future patches for SEV-SNP-validated CPUID will also require early
parsing of the EFI configuration. Incrementally move the related code
into a set of helpers that can be re-used for that purpose.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/Makefile | 1 +
arch/x86/boot/compressed/acpi.c | 18 ++++-----
arch/x86/boot/compressed/efi.c | 64 +++++++++++++++++++++++++++++++
arch/x86/boot/compressed/misc.h | 14 +++++++
4 files changed, 87 insertions(+), 10 deletions(-)
create mode 100644 arch/x86/boot/compressed/efi.c
@@ -98,18 +98,16 @@ static acpi_physical_address kexec_get_rsdp_addr(void)return0;}-ei=&boot_params->efi_info;-sig=(char*)&ei->efi_loader_signature;-if(strncmp(sig,EFI64_LOADER_SIGNATURE,4)){+/* Get systab from boot params. */+ret=efi_get_system_table(boot_params,(unsignedlong*)&systab,&efi_64);+if(ret)+error("EFI system table not found in kexec boot_params.");++if(!efi_64){debug_putstr("Wrong kexec EFI loader signature.\n");return0;}-/* Get systab from boot params. */-systab=(efi_system_table_64_t*)(ei->efi_systab|((__u64)ei->efi_systab_hi<<32));-if(!systab)-error("EFI system table not found in kexec boot_params.");-return__efi_get_rsdp_addr((unsignedlong)esd->tables,systab->nr_tables,true);}#else
@@ -0,0 +1,64 @@+// SPDX-License-Identifier: GPL-2.0+/*+*HelpersforearlyaccesstoEFIconfigurationtable+*+*Copyright(C)2021AdvancedMicroDevices,Inc.+*+*Author:MichaelRoth<michael.roth@amd.com>+*/++#include"misc.h"+#include<linux/efi.h>+#include<asm/efi.h>++/**+*Givenboot_params,retrievethephysicaladdressofEFIsystemtable.+*+*@boot_params:pointertoboot_params+*@sys_tbl_pa:locationtostorephysicaladdressofsystemtable+*@is_efi_64:locationtostorewhetherusing64-bitEFIornot+*+*Returns0onsuccess.Onerror,returnparamsareleftunchanged.+*/+intefi_get_system_table(structboot_params*boot_params,unsignedlong*sys_tbl_pa,+bool*is_efi_64)+{+unsignedlongsys_tbl;+structefi_info*ei;+boolefi_64;+char*sig;++if(!sys_tbl_pa||!is_efi_64)+return-EINVAL;++ei=&boot_params->efi_info;+sig=(char*)&ei->efi_loader_signature;++if(!strncmp(sig,EFI64_LOADER_SIGNATURE,4)){+efi_64=true;+}elseif(!strncmp(sig,EFI32_LOADER_SIGNATURE,4)){+efi_64=false;+}else{+debug_putstr("Wrong EFI loader signature.\n");+return-ENOENT;+}++/* Get systab from boot params. */+#ifdef CONFIG_X86_64+sys_tbl=ei->efi_systab|((__u64)ei->efi_systab_hi<<32);+#else+if(ei->efi_systab_hi||ei->efi_memmap_hi){+debug_putstr("Error: EFI system table located above 4GB.\n");+return-EINVAL;+}+sys_tbl=ei->efi_systab;+#endif+if(!sys_tbl){+debug_putstr("EFI system table not found.");+return-ENOENT;+}++*sys_tbl_pa=sys_tbl;+*is_efi_64=efi_64;+return0;+}
The early_set_memory_{encrypt,decrypt}() are used for changing the
page from decrypted (shared) to encrypted (private) and vice versa.
When SEV-SNP is active, the page state transition needs to go through
additional steps.
If the page is transitioned from shared to private, then perform the
following after the encryption attribute is set in the page table:
1. Issue the page state change VMGEXIT to add the page as a private
in the RMP table.
2. Validate the page after its successfully added in the RMP table.
To maintain the security guarantees, if the page is transitioned from
private to shared, then perform the following before clearing the
encryption attribute from the page table.
1. Invalidate the page.
2. Issue the page state change VMGEXIT to make the page shared in the
RMP table.
The early_set_memory_{encrypt,decrypt} can be called before the GHCB
is setup, use the SNP page state MSR protocol VMGEXIT defined in the GHCB
specification to request the page state change in the RMP table.
While at it, add a helper snp_prep_memory() that can be used outside
the sev specific files to change the page state for a specified memory
range.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev.h | 10 ++++
arch/x86/kernel/sev.c | 102 +++++++++++++++++++++++++++++++++++++
arch/x86/mm/mem_encrypt.c | 51 +++++++++++++++++--
3 files changed, 159 insertions(+), 4 deletions(-)
@@ -553,6 +553,108 @@ static u64 get_jump_table_addr(void)returnret;}+staticvoidpvalidate_pages(unsignedlongvaddr,unsignedintnpages,boolvalidate)+{+unsignedlongvaddr_end;+intrc;++vaddr=vaddr&PAGE_MASK;+vaddr_end=vaddr+(npages<<PAGE_SHIFT);++while(vaddr<vaddr_end){+rc=pvalidate(vaddr,RMP_PG_SIZE_4K,validate);+if(WARN(rc,"Failed to validate address 0x%lx ret %d",vaddr,rc))+sev_es_terminate(SEV_TERM_SET_LINUX,GHCB_TERM_PVALIDATE);++vaddr=vaddr+PAGE_SIZE;+}+}++staticvoid__initearly_set_page_state(unsignedlongpaddr,unsignedintnpages,enumpsc_opop)+{+unsignedlongpaddr_end;+u64val;++paddr=paddr&PAGE_MASK;+paddr_end=paddr+(npages<<PAGE_SHIFT);++while(paddr<paddr_end){+/*+*UsetheMSRprotocolbecausethisfunctioncanbecalledbeforetheGHCB+*isestablished.+*/+sev_es_wr_ghcb_msr(GHCB_MSR_PSC_REQ_GFN(paddr>>PAGE_SHIFT,op));+VMGEXIT();++val=sev_es_rd_ghcb_msr();++if(WARN(GHCB_RESP_CODE(val)!=GHCB_MSR_PSC_RESP,+"Wrong PSC response code: 0x%x\n",+(unsignedint)GHCB_RESP_CODE(val)))+gotoe_term;++if(WARN(GHCB_MSR_PSC_RESP_VAL(val),+"Failed to change page state to '%s' paddr 0x%lx error 0x%llx\n",+op==SNP_PAGE_STATE_PRIVATE?"private":"shared",+paddr,GHCB_MSR_PSC_RESP_VAL(val)))+gotoe_term;++paddr=paddr+PAGE_SIZE;+}++return;++e_term:+sev_es_terminate(SEV_TERM_SET_LINUX,GHCB_TERM_PSC);+}++void__initearly_snp_set_memory_private(unsignedlongvaddr,unsignedlongpaddr,+unsignedintnpages)+{+if(!cc_platform_has(CC_ATTR_SEV_SNP))+return;++/*+*AskthehypervisortomarkthememorypagesasprivateintheRMP+*table.+*/+early_set_page_state(paddr,npages,SNP_PAGE_STATE_PRIVATE);++/* Validate the memory pages after they've been added in the RMP table. */+pvalidate_pages(vaddr,npages,1);+}++void__initearly_snp_set_memory_shared(unsignedlongvaddr,unsignedlongpaddr,+unsignedintnpages)+{+if(!cc_platform_has(CC_ATTR_SEV_SNP))+return;++/*+*Invalidatethememorypagesbeforetheyaremarkedsharedinthe+*RMPtable.+*/+pvalidate_pages(vaddr,npages,0);++/* Ask hypervisor to mark the memory pages shared in the RMP table. */+early_set_page_state(paddr,npages,SNP_PAGE_STATE_SHARED);+}++void__initsnp_prep_memory(unsignedlongpaddr,unsignedintsz,enumpsc_opop)+{+unsignedlongvaddr,npages;++vaddr=(unsignedlong)__va(paddr);+npages=PAGE_ALIGN(sz)>>PAGE_SHIFT;++if(op==SNP_PAGE_STATE_PRIVATE)+early_snp_set_memory_private(vaddr,paddr,npages);+elseif(op==SNP_PAGE_STATE_SHARED)+early_snp_set_memory_shared(vaddr,paddr,npages);+else+WARN(1,"invalid memory op %d\n",op);+}+intsev_es_setup_ap_jump_table(structreal_mode_header*rmh){u16startup_cs,startup_ip;
@@ -49,6 +50,34 @@ EXPORT_SYMBOL_GPL(sev_enable_key);/* Buffer used for early in-place encryption by BSP, no locking needed */staticcharsme_early_buffer[PAGE_SIZE]__initdata__aligned(PAGE_SIZE);+/*+*WhenSNPisactive,changethepagestatefromprivatetosharedbefore+*copyingthedatafromthesourcetodestinationandrestoreafterthecopy.+*Thisisrequiredbecausethesourceaddressismappedasdecryptedbythe+*calleroftheroutine.+*/+staticinlinevoid__initsnp_memcpy(void*dst,void*src,size_tsz,+unsignedlongpaddr,booldecrypt)+{+unsignedlongnpages=PAGE_ALIGN(sz)>>PAGE_SHIFT;++if(!cc_platform_has(CC_ATTR_SEV_SNP)||!decrypt){+memcpy(dst,src,sz);+return;+}++/*+*WithSNP,thepaddrneedstobeaccesseddecrypted,markthepage+*sharedintheRMPtablebeforecopyingit.+*/+early_snp_set_memory_shared((unsignedlong)__va(paddr),paddr,npages);++memcpy(dst,src,sz);++/* Restore the page state after the memcpy. */+early_snp_set_memory_private((unsignedlong)__va(paddr),paddr,npages);+}+/**Thisroutinedoesnotchangetheunderlyingencryptionsettingofthe*page(s)thatmapthismemory.Itassumesthateventuallythememoryis
From: Michael Roth <redacted>
This code will also be used later for SEV-SNP-validated CPUID code in
some cases, so move it to a common helper.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/sev-shared.c | 84 +++++++++++++++++++++++++-----------
1 file changed, 58 insertions(+), 26 deletions(-)
From: Michael Roth <redacted>
Determining which CPUID leafs have significant ECX/index values is
also needed by guest kernel code when doing SEV-SNP-validated CPUID
lookups. Move this to common code to keep future updates in sync.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/cpuid.h | 26 ++++++++++++++++++++++++++
arch/x86/kvm/cpuid.c | 17 ++---------------
2 files changed, 28 insertions(+), 15 deletions(-)
create mode 100644 arch/x86/include/asm/cpuid.h
From: Michael Roth <redacted>
Future patches for SEV-SNP-validated CPUID will also require early
parsing of the EFI configuration. Incrementally move the related code
into a set of helpers that can be re-used for that purpose.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/acpi.c | 50 ++++++++-----------------
arch/x86/boot/compressed/efi.c | 65 +++++++++++++++++++++++++++++++++
arch/x86/boot/compressed/misc.h | 9 +++++
3 files changed, 90 insertions(+), 34 deletions(-)
@@ -104,3 +104,68 @@ int efi_get_conf_table(struct boot_params *boot_params, unsigned long *cfg_tbl_preturn0;}++/* Get vendor table address/guid from EFI config table at the given index */+staticintget_vendor_table(void*cfg_tbl,unsignedintidx,+unsignedlong*vendor_tbl_pa,+efi_guid_t*vendor_tbl_guid,+boolefi_64)+{+if(efi_64){+efi_config_table_64_t*tbl_entry=+(efi_config_table_64_t*)cfg_tbl+idx;++if(!IS_ENABLED(CONFIG_X86_64)&&tbl_entry->table>>32){+debug_putstr("Error: EFI config table entry located above 4GB.\n");+return-EINVAL;+}++*vendor_tbl_pa=tbl_entry->table;+*vendor_tbl_guid=tbl_entry->guid;++}else{+efi_config_table_32_t*tbl_entry=+(efi_config_table_32_t*)cfg_tbl+idx;++*vendor_tbl_pa=tbl_entry->table;+*vendor_tbl_guid=tbl_entry->guid;+}++return0;+}++/**+*GivenEFIconfigtable,searchitforthephysicaladdressofthevendor+*tableassociatedwithGUID.+*+*@cfg_tbl_pa:pointertoEFIconfigurationtable+*@cfg_tbl_len:numberofentriesinEFIconfigurationtable+*@guid:GUIDofvendortable+*@efi_64:trueifusing64-bitEFI+*@vendor_tbl_pa:locationtostorephysicaladdressofvendortable+*+*Returns0onsuccess.Onerror,returnparamsareleftunchanged.+*/+intefi_find_vendor_table(unsignedlongcfg_tbl_pa,unsignedintcfg_tbl_len,+efi_guid_tguid,boolefi_64,unsignedlong*vendor_tbl_pa)+{+unsignedinti;++for(i=0;i<cfg_tbl_len;i++){+unsignedlongvendor_tbl_pa_tmp;+efi_guid_tvendor_tbl_guid;+intret;++if(get_vendor_table((void*)cfg_tbl_pa,i,+&vendor_tbl_pa_tmp,+&vendor_tbl_guid,efi_64))+return-EINVAL;++if(!efi_guidcmp(guid,vendor_tbl_guid)){+*vendor_tbl_pa=vendor_tbl_pa_tmp;+return0;+}+}++return-ENOENT;+}
From: Michael Roth <redacted>
Future patches for SEV-SNP-validated CPUID will also require early
parsing of the EFI configuration. Incrementally move the related code
into a set of helpers that can be re-used for that purpose.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/acpi.c | 52 +++++----------------------------
arch/x86/boot/compressed/efi.c | 42 ++++++++++++++++++++++++++
arch/x86/boot/compressed/misc.h | 9 ++++++
3 files changed, 58 insertions(+), 45 deletions(-)
@@ -179,6 +179,8 @@ unsigned long sev_verify_cbit(unsigned long cr3);/* helpers for early EFI config table access */intefi_get_system_table(structboot_params*boot_params,unsignedlong*sys_tbl_pa,bool*is_efi_64);+intefi_get_conf_table(structboot_params*boot_params,unsignedlong*cfg_tbl_pa,+unsignedint*cfg_tbl_len,bool*is_efi_64);#elsestaticinlineintefi_get_system_table(structboot_params*boot_params,
While launching the encrypted guests, the hypervisor may need to provide
some additional information during the guest boot. When booting under the
EFI based BIOS, the EFI configuration table contains an entry for the
confidential computing blob that contains the required information.
To support booting encrypted guests on non-EFI VM, the hypervisor needs to
pass this additional information to the kernel with a different method.
For this purpose, introduce SETUP_CC_BLOB type in setup_data to hold the
physical address of the confidential computing blob location. The boot
loader or hypervisor may choose to use this method instead of EFI
configuration table. The CC blob location scanning should give preference
to setup_data data over the EFI configuration table.
In AMD SEV-SNP, the CC blob contains the address of the secrets and CPUID
pages. The secrets page includes information such as a VM to PSP
communication key and CPUID page contains PSP filtered CPUID values.
Define the AMD SEV confidential computing blob structure.
While at it, define the EFI GUID for the confidential computing blob.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev.h | 12 ++++++++++++
arch/x86/include/uapi/asm/bootparam.h | 1 +
include/linux/efi.h | 1 +
3 files changed, 14 insertions(+)
The SNP_GET_DERIVED_KEY ioctl interface can be used by the SNP guest to
ask the firmware to provide a key derived from a root key. The derived
key may be used by the guest for any purposes it choose, such as a
sealing key or communicating with the external entities.
See SEV-SNP firmware spec for more information.
Signed-off-by: Brijesh Singh <redacted>
---
Documentation/virt/coco/sevguest.rst | 19 ++++++++++-
drivers/virt/coco/sevguest/sevguest.c | 49 +++++++++++++++++++++++++++
include/uapi/linux/sev-guest.h | 24 +++++++++++++
3 files changed, 91 insertions(+), 1 deletion(-)
@@ -64,10 +64,27 @@ The SNP_GET_REPORT ioctl can be used to query the attestation report from the SEV-SNP firmware. The ioctl uses the SNP_GUEST_REQUEST (MSG_REPORT_REQ) command provided by the SEV-SNP firmware to query the attestation report.-On success, the snp_report_resp.data will contains the report. The report+On success, the snp_report_resp.data will contain the report. The report will contain the format described in the SEV-SNP specification. See the SEV-SNP specification for further details.+2.2 SNP_GET_DERIVED_KEY+-----------------------+:Technology: sev-snp+:Type: guest ioctl+:Parameters (in): struct snp_derived_key_req+:Returns (out): struct snp_derived_key_req on success, -negative on error++The SNP_GET_DERIVED_KEY ioctl can be used to get a key derive from a root key.+The derived key can be used by the guest for any purpose, such as sealing keys+or communicating with external entities.++The ioctl uses the SNP_GUEST_REQUEST (MSG_KEY_REQ) command provided by the+SEV-SNP firmware to derive the key. See SEV-SNP specification for further details+on the various fileds passed in the key derivation request.++On success, the snp_derived_key_resp.data will contains the derived key value. See+the SEV-SNP specification for further details. Reference ---------
@@ -364,6 +364,52 @@ static int get_report(struct snp_guest_dev *snp_dev, struct snp_guest_request_ioreturnrc;}+staticintget_derived_key(structsnp_guest_dev*snp_dev,structsnp_guest_request_ioctl*arg)+{+structsnp_guest_crypto*crypto=snp_dev->crypto;+structsnp_derived_key_respresp={0};+structsnp_derived_key_reqreq;+intrc,resp_len;+u8buf[89];++if(!arg->req_data||!arg->resp_data)+return-EINVAL;++/* Copy the request payload from userspace */+if(copy_from_user(&req,(void__user*)arg->req_data,sizeof(req)))+return-EFAULT;++/* Message version must be non-zero */+if(!req.msg_version)+return-EINVAL;++/*+*Theintermediateresponsebufferisusedwhiledecryptingthe+*responsepayload.Makesurethatithasenoughspacetocoverthe+*authtag.+*/+resp_len=sizeof(resp.data)+crypto->a_len;+if(sizeof(buf)<resp_len)+return-ENOMEM;++/* Issue the command to get the attestation report */+rc=handle_guest_request(snp_dev,SVM_VMGEXIT_GUEST_REQUEST,req.msg_version,+SNP_MSG_KEY_REQ,&req.data,sizeof(req.data),buf,resp_len,+&arg->fw_err);+if(rc)+gotoe_free;++/* Copy the response payload to userspace */+memcpy(resp.data,buf,sizeof(resp.data));+if(copy_to_user((void__user*)arg->resp_data,&resp,sizeof(resp)))+rc=-EFAULT;++e_free:+memzero_explicit(buf,sizeof(buf));+memzero_explicit(&resp,sizeof(resp));+returnrc;+}+staticlongsnp_guest_ioctl(structfile*file,unsignedintioctl,unsignedlongarg){structsnp_guest_dev*snp_dev=to_snp_dev(file);
@@ -382,6 +428,9 @@ static long snp_guest_ioctl(struct file *file, unsigned int ioctl, unsigned longcaseSNP_GET_REPORT:ret=get_report(snp_dev,&input);break;+caseSNP_GET_DERIVED_KEY:+ret=get_derived_key(snp_dev,&input);+break;default:break;}
@@ -36,9 +36,33 @@ struct snp_guest_request_ioctl {__u64fw_err;};+struct__snp_derived_key_req{+__u32root_key_select;+__u32rsvd;+__u64guest_field_select;+__u32vmpl;+__u32guest_svn;+__u64tcb_version;+};++structsnp_derived_key_req{+/* message version number (must be non-zero) */+__u8msg_version;++struct__snp_derived_key_reqdata;+};++structsnp_derived_key_resp{+/* response data, see SEV-SNP spec for the format */+__u8data[64];+};+#define SNP_GUEST_REQ_IOC_TYPE 'S'/* Get SNP attestation report */#define SNP_GET_REPORT _IOWR(SNP_GUEST_REQ_IOC_TYPE, 0x0, struct snp_guest_request_ioctl)+/* Get a derived key from the root */+#define SNP_GET_DERIVED_KEY _IOWR(SNP_GUEST_REQ_IOC_TYPE, 0x1, struct snp_guest_request_ioctl)+#endif /* __UAPI_LINUX_SEV_GUEST_H_ */
Version 2 of GHCB specification provides Non Automatic Exit (NAE) that can
be used by the SNP guest to communicate with the PSP without risk from a
malicious hypervisor who wishes to read, alter, drop or replay the messages
sent.
SNP_LAUNCH_UPDATE can insert two special pages into the guest’s memory:
the secrets page and the CPUID page. The PSP firmware populate the contents
of the secrets page. The secrets page contains encryption keys used by the
guest to interact with the firmware. Because the secrets page is encrypted
with the guest’s memory encryption key, the hypervisor cannot read the keys.
See SNP FW ABI spec for further details about the secrets page.
Create a platform device that the SNP guest driver can bind to get the
platform resources such as encryption key and message id to use to
communicate with the PSP. The SNP guest driver provides a userspace
interface to get the attestation report, key derivation, extended
attestation report etc.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev.h | 4 +++
arch/x86/kernel/sev.c | 61 ++++++++++++++++++++++++++++++++++++++
2 files changed, 65 insertions(+)
Version 2 of GHCB specification provides SNP_GUEST_REQUEST and
SNP_EXT_GUEST_REQUEST NAE that can be used by the SNP guest to communicate
with the PSP.
While at it, add a snp_issue_guest_request() helper that can be used by
driver or other subsystem to issue the request to PSP.
See SEV-SNP and GHCB spec for more details.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev-common.h | 3 ++
arch/x86/include/asm/sev.h | 13 ++++++++
arch/x86/include/uapi/asm/svm.h | 4 +++
arch/x86/kernel/sev.c | 50 +++++++++++++++++++++++++++++++
4 files changed, 70 insertions(+)
@@ -2121,3 +2121,53 @@ static int __init snp_cpuid_check_status(void)}arch_initcall(snp_cpuid_check_status);++intsnp_issue_guest_request(u64exit_code,structsnp_req_data*input,unsignedlong*fw_err)+{+structghcb_statestate;+unsignedlongflags;+structghcb*ghcb;+intret;++if(!cc_platform_has(CC_ATTR_SEV_SNP))+return-ENODEV;++local_irq_save(flags);++ghcb=__sev_get_ghcb(&state);+if(!ghcb){+ret=-EIO;+gotoe_restore_irq;+}++vc_ghcb_invalidate(ghcb);++if(exit_code==SVM_VMGEXIT_EXT_GUEST_REQUEST){+ghcb_set_rax(ghcb,input->data_gpa);+ghcb_set_rbx(ghcb,input->data_npages);+}++ret=sev_es_ghcb_hv_call(ghcb,NULL,exit_code,input->req_gpa,input->resp_gpa);+if(ret)+gotoe_put;++if(ghcb->save.sw_exit_info_2){+/* Number of expected pages are returned in RBX */+if(exit_code==SVM_VMGEXIT_EXT_GUEST_REQUEST&&+ghcb->save.sw_exit_info_2==SNP_GUEST_REQ_INVALID_LEN)+input->data_npages=ghcb_get_rbx(ghcb);++if(fw_err)+*fw_err=ghcb->save.sw_exit_info_2;++ret=-EIO;+}++e_put:+__sev_put_ghcb(&state);+e_restore_irq:+local_irq_restore(flags);++returnret;+}+EXPORT_SYMBOL_GPL(snp_issue_guest_request);
SEV-SNP specification provides the guest a mechanisum to communicate with
the PSP without risk from a malicious hypervisor who wishes to read, alter,
drop or replay the messages sent. The driver uses snp_issue_guest_request()
to issue GHCB SNP_GUEST_REQUEST or SNP_EXT_GUEST_REQUEST NAE events to
submit the request to PSP.
The PSP requires that all communication should be encrypted using key
specified through the platform_data.
The userspace can use SNP_GET_REPORT ioctl() to query the guest
attestation report.
See SEV-SNP spec section Guest Messages for more details.
Signed-off-by: Brijesh Singh <redacted>
---
Documentation/virt/coco/sevguest.rst | 77 ++++
drivers/virt/Kconfig | 3 +
drivers/virt/Makefile | 1 +
drivers/virt/coco/sevguest/Kconfig | 9 +
drivers/virt/coco/sevguest/Makefile | 2 +
drivers/virt/coco/sevguest/sevguest.c | 561 ++++++++++++++++++++++++++
drivers/virt/coco/sevguest/sevguest.h | 98 +++++
include/uapi/linux/sev-guest.h | 44 ++
8 files changed, 795 insertions(+)
create mode 100644 Documentation/virt/coco/sevguest.rst
create mode 100644 drivers/virt/coco/sevguest/Kconfig
create mode 100644 drivers/virt/coco/sevguest/Makefile
create mode 100644 drivers/virt/coco/sevguest/sevguest.c
create mode 100644 drivers/virt/coco/sevguest/sevguest.h
create mode 100644 include/uapi/linux/sev-guest.h
@@ -0,0 +1,77 @@+.. SPDX-License-Identifier: GPL-2.0++===================================================================+The Definitive SEV Guest API Documentation+===================================================================++1. General description+======================++The SEV API is a set of ioctls that are used by the guest or hypervisor+to get or set certain aspect of the SEV virtual machine. The ioctls belong+to the following classes:++- Hypervisor ioctls: These query and set global attributes which affect the+ whole SEV firmware. These ioctl are used by platform provision tools.++- Guest ioctls: These query and set attributes of the SEV virtual machine.++2. API description+==================++This section describes ioctls that can be used to query or set SEV guests.+For each ioctl, the following information is provided along with a+description:++ Technology:+ which SEV techology provides this ioctl. sev, sev-es, sev-snp or all.++ Type:+ hypervisor or guest. The ioctl can be used inside the guest or the+ hypervisor.++ Parameters:+ what parameters are accepted by the ioctl.++ Returns:+ the return value. General error numbers (ENOMEM, EINVAL)+ are not detailed, but errors with specific meanings are.++The guest ioctl should be issued on a file descriptor of the /dev/sev-guest device.+The ioctl accepts struct snp_user_guest_request. The input and output structure is+specified through the req_data and resp_data field respectively. If the ioctl fails+to execute due to a firmware error, then fw_err code will be set.++::+ struct snp_guest_request_ioctl {+ /* Request and response structure address */+ __u64 req_data;+ __u64 resp_data;++ /* firmware error code on failure (see psp-sev.h) */+ __u64 fw_err;+ };++2.1 SNP_GET_REPORT+------------------++:Technology: sev-snp+:Type: guest ioctl+:Parameters (in): struct snp_report_req+:Returns (out): struct snp_report_resp on success, -negative on error++The SNP_GET_REPORT ioctl can be used to query the attestation report from the+SEV-SNP firmware. The ioctl uses the SNP_GUEST_REQUEST (MSG_REPORT_REQ) command+provided by the SEV-SNP firmware to query the attestation report.++On success, the snp_report_resp.data will contains the report. The report+will contain the format described in the SEV-SNP specification. See the SEV-SNP+specification for further details.+++Reference+---------++SEV-SNP and GHCB specification: developer.amd.com/sev++The driver is based on SEV-SNP firmware spec 0.9 and GHCB spec version 2.0.
@@ -0,0 +1,98 @@+/* SPDX-License-Identifier: GPL-2.0-only */+/*+*Copyright(C)2021AdvancedMicroDevices,Inc.+*+*Author:BrijeshSingh<brijesh.singh@amd.com>+*+*SEV-SNPAPIspecisavailableathttps://developer.amd.com/sev+*/++#ifndef __LINUX_SEVGUEST_H_+#define __LINUX_SEVGUEST_H_++#include<linux/types.h>++#define MAX_AUTHTAG_LEN 32++/* See SNP spec SNP_GUEST_REQUEST section for the structure */+enummsg_type{+SNP_MSG_TYPE_INVALID=0,+SNP_MSG_CPUID_REQ,+SNP_MSG_CPUID_RSP,+SNP_MSG_KEY_REQ,+SNP_MSG_KEY_RSP,+SNP_MSG_REPORT_REQ,+SNP_MSG_REPORT_RSP,+SNP_MSG_EXPORT_REQ,+SNP_MSG_EXPORT_RSP,+SNP_MSG_IMPORT_REQ,+SNP_MSG_IMPORT_RSP,+SNP_MSG_ABSORB_REQ,+SNP_MSG_ABSORB_RSP,+SNP_MSG_VMRK_REQ,+SNP_MSG_VMRK_RSP,++SNP_MSG_TYPE_MAX+};++enumaead_algo{+SNP_AEAD_INVALID,+SNP_AEAD_AES_256_GCM,+};++structsnp_guest_msg_hdr{+u8authtag[MAX_AUTHTAG_LEN];+u64msg_seqno;+u8rsvd1[8];+u8algo;+u8hdr_version;+u16hdr_sz;+u8msg_type;+u8msg_version;+u16msg_sz;+u32rsvd2;+u8msg_vmpck;+u8rsvd3[35];+}__packed;++structsnp_guest_msg{+structsnp_guest_msg_hdrhdr;+u8payload[4000];+}__packed;++/*+*Thesecretspagecontains96-bytesofreservedfieldthatcanbeusedby+*theguestOS.TheguestOSusestheareatosavethemessagesequence+*numberforeachVMPCK.+*+*SeetheGHCBspecsectionSecretpagelayoutfortheformatforthisarea.+*/+structsecrets_os_area{+u32msg_seqno_0;+u32msg_seqno_1;+u32msg_seqno_2;+u32msg_seqno_3;+u64ap_jump_table_pa;+u8rsvd[40];+u8guest_usage[32];+}__packed;++#define VMPCK_KEY_LEN 32++/* See the SNP spec version 0.9 for secrets page format */+structsnp_secrets_page_layout{+u32version;+u32imien:1,+rsvd1:31;+u32fms;+u32rsvd2;+u8gosvw[16];+u8vmpck0[VMPCK_KEY_LEN];+u8vmpck1[VMPCK_KEY_LEN];+u8vmpck2[VMPCK_KEY_LEN];+u8vmpck3[VMPCK_KEY_LEN];+structsecrets_os_areaos_area;+u8rsvd3[3840];+}__packed;++#endif /* __LINUX_SNP_GUEST_H__ */
@@ -0,0 +1,44 @@+/* SPDX-License-Identifier: GPL-2.0-only WITH Linux-syscall-note */+/*+*UserspaceinterfaceforAMDSEVandSEV-SNPguestdriver.+*+*Copyright(C)2021AdvancedMicroDevices,Inc.+*+*Author:BrijeshSingh<brijesh.singh@amd.com>+*+*SEVAPIspecificationisavailableat:https://developer.amd.com/sev/+*/++#ifndef __UAPI_LINUX_SEV_GUEST_H_+#define __UAPI_LINUX_SEV_GUEST_H_++#include<linux/types.h>++structsnp_report_req{+/* message version number (must be non-zero) */+__u8msg_version;++/* user data that should be included in the report */+__u8user_data[64];+};++structsnp_report_resp{+/* response data, see SEV-SNP spec for the format */+__u8data[4000];+};++structsnp_guest_request_ioctl{+/* Request and response structure address */+__u64req_data;+__u64resp_data;++/* firmware error code on failure (see psp-sev.h) */+__u64fw_err;+};++#define SNP_GUEST_REQ_IOC_TYPE 'S'++/* Get SNP attestation report */+#define SNP_GET_REPORT _IOWR(SNP_GUEST_REQ_IOC_TYPE, 0x0, struct snp_guest_request_ioctl)++#endif /* __UAPI_LINUX_SEV_GUEST_H_ */
From: Michael Roth <redacted>
The run-time kernel will need to access the Confidential Computing
blob very early in boot to access the CPUID table it points to. At
that stage of boot it will be relying on the identity-mapped page table
set up by boot/compressed kernel, so make sure the blob and the CPUID
table it points to are mapped in advance.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/ident_map_64.c | 26 ++++++++++++++++++++++++-
arch/x86/boot/compressed/misc.h | 2 ++
arch/x86/boot/compressed/sev.c | 2 +-
3 files changed, 28 insertions(+), 2 deletions(-)
@@ -37,6 +37,8 @@#include<asm/setup.h>/* For COMMAND_LINE_SIZE */#undef _SETUP+#include<asm/sev.h>/* For ConfidentialComputing blob */+externunsignedlongget_cmd_line_ptr(void);/* Used by PAGE_KERN* macros: */
@@ -106,6 +108,27 @@ static void add_identity_map(unsigned long start, unsigned long end)error("Error: kernel_ident_mapping_init() failed\n");}+voidsev_prep_identity_maps(void)+{+/*+*TheConfidentialComputingblobisusedveryearlyinuncompressed+*kerneltofindthein-memorycpuidtabletohandlecpuid+*instructions.Makesureanidentity-mappingexistssoitcanbe+*accessedafterswitchover.+*/+if(sev_snp_enabled()){+structcc_blob_sev_info*cc_info=+(void*)(unsignedlong)boot_params->cc_blob_address;++add_identity_map((unsignedlong)cc_info,+(unsignedlong)cc_info+sizeof(*cc_info));+add_identity_map((unsignedlong)cc_info->cpuid_phys,+(unsignedlong)cc_info->cpuid_phys+cc_info->cpuid_len);+}++sev_verify_cbit(top_level_pgt);+}+/* Locates and clears a region for a new top level page table. */voidinitialize_identity_maps(void*rmode){
@@ -163,8 +186,9 @@ void initialize_identity_maps(void *rmode)cmdline=get_cmd_line_ptr();add_identity_map(cmdline,cmdline+COMMAND_LINE_SIZE);+sev_prep_identity_maps();+/* Load the new page-table. */-sev_verify_cbit(top_level_pgt);write_cr3(top_level_pgt);}
From: Michael Roth <redacted>
SEV-SNP guests will be provided the location of special 'secrets' and
'CPUID' pages via the Confidential Computing blob. This blob is
provided to the run-time kernel either through bootparams field that
was initialized by the boot/compressed kernel, or via a setup_data
structure as defined by the Linux Boot Protocol.
Locate the Confidential Computing from these sources and, if found,
use the provided CPUID page/table address to create a copy that the
run-time kernel will use when servicing cpuid instructions via a #VC
handler.
This must be set up during early startup before any cpuid instructions
are issued. As result, some pointer fixups are needed early on that
must be adjusted later in boot, which is why there are 2 init routines.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/sev.c | 2 +-
arch/x86/include/asm/setup.h | 2 +-
arch/x86/include/asm/sev.h | 17 +----
arch/x86/kernel/head64.c | 12 ++-
arch/x86/kernel/sev-shared.c | 23 +++++-
arch/x86/kernel/sev.c | 135 +++++++++++++++++++++++++++++++++
6 files changed, 170 insertions(+), 21 deletions(-)
@@ -361,7 +361,7 @@ void snp_cpuid_init_boot(struct boot_params *bp)if(!cc_info)return;-snp_cpuid_info_create(cc_info);+snp_cpuid_info_create(cc_info,0);/* SEV-SNP CPUID table is set up now. Do some sanity checks. */if(!snp_cpuid_active())
@@ -571,7 +571,7 @@ static void set_bringup_idt_handler(gate_desc *idt, int n, void *handler)}/* This runs while still in the direct mapping */-staticvoidstartup_64_load_idt(unsignedlongphysbase)+staticvoidstartup_64_load_idt(unsignedlongphysbase,structboot_params*bp){structdesc_ptr*desc=fixup_pointer(&bringup_idt_descr,physbase);gate_desc*idt=fixup_pointer(bringup_idt_table,physbase);
@@ -587,6 +587,9 @@ static void startup_64_load_idt(unsigned long physbase)desc->address=(unsignedlong)idt;native_load_idt(desc);++if(IS_ENABLED(CONFIG_AMD_MEM_ENCRYPT))+snp_cpuid_init_startup(bp,physbase);}/* This is used when running on kernel addresses */
@@ -988,6 +988,22 @@ snp_find_cc_blob_setup_data(struct boot_params *bp)return(structcc_blob_sev_info*)(unsignedlong)sd->cc_blob_address;}+staticconststructsnp_cpuid_info*+snp_cpuid_info_get_ptr(unsignedlongphysbase)+{+void*ptr=&cpuid_info_copy;++/* physbase is only 0 when the caller doesn't need adjustments */+if(!physbase)+returnptr;++/*+*Handlerelocationadjustmentsforglobalpointers,asdoneby+*fixup_pointer()in__startup64().+*/+returnptr-(void*)_text+(void*)physbase;+}+/**Initializethekernel'scopyoftheSEV-SNPCPUIDtable,andsetupthe*pointerthatwillbeusedtoaccessit.
@@ -1986,3 +1986,138 @@ bool __init handle_vc_boot_ghcb(struct pt_regs *regs)while(true)halt();}++/*+*InitialsetupofSEV-SNPCPUIDtablereliesoninformationprovided+*bytheConfidentialComputingblob,whichcanbepassedtothekernel+*inthefollowingways,dependingonhowitisbooted:+*+*-whenbootedviatheboot/decompresskernel:+*-viaboot_params+*+*-whenbooteddirectlybyfirmware/bootloader(e.g.CONFIG_PVH):+*-viaasetup_dataentry,asdefinedbytheLinuxBootProtocol+*+*Scanfortheblobinthatorder.+*/+structcc_blob_sev_info*snp_find_cc_blob(structboot_params*bp)+{+structcc_blob_sev_info*cc_info;++/* Boot kernel would have passed the CC blob via boot_params. */+if(bp->cc_blob_address){+cc_info=(structcc_blob_sev_info*)+(unsignedlong)bp->cc_blob_address;+gotofound_cc_info;+}++/*+*Ifkernelwasbooteddirectly,withouttheuseofthe+*boot/decompressionkernel,theCCblobmayhavebeenpassedvia+*setup_datainstead.+*/+cc_info=snp_find_cc_blob_setup_data(bp);+if(!cc_info)+returnNULL;++found_cc_info:+if(cc_info->magic!=CC_BLOB_SEV_HDR_MAGIC)+sev_es_terminate(1,GHCB_SNP_UNSUPPORTED);++returncc_info;+}++/*+*InitialsetupofSEV-SNPCPUIDtableduringearlystartupwhenstill+*usingidentity-mappedaddresses.+*+*Sincethisisduringearlystartup,physbaseisneededtogeneratethe+*correctpointertotheinitializedCPUIDtable.Thispointerwillbe+*adjustedagainlaterviasnp_cpuid_init()afterthekernelswitchesover+*tovirtualaddressesandpointerfixupsarenolongerneeded.+*/+void__initsnp_cpuid_init_startup(structboot_params*bp,+unsignedlongphysbase)+{+structcc_blob_sev_info*cc_info;+u32eax;++if(!bp)+return;++cc_info=snp_find_cc_blob(bp);+if(!cc_info)+return;++snp_cpuid_info_create(cc_info,physbase);++/* SEV-SNP CPUID table is set up now. Do some sanity checks. */+if(!snp_cpuid_active())+sev_es_terminate(1,GHCB_TERM_CPUID);++/* SEV (bit 1) and SEV-SNP (bit 4) should be enabled in CPUID. */+eax=native_cpuid_eax(0x8000001f);+if(!(eax&(BIT(4)|BIT(1))))+sev_es_terminate(1,GHCB_TERM_CPUID);++/* #VC generated by CPUID above will set sev_status based on SEV MSR. */+if(!(sev_status&MSR_AMD64_SEV_SNP_ENABLED))+sev_es_terminate(1,GHCB_TERM_CPUID);++/*+*TheCCblobwillbeusedlatertoaccessthesecretspage.Cache+*itherelikethebootkerneldoes.+*/+bp->cc_blob_address=(u32)(unsignedlong)cc_info;+}++/*+*Thisiscalledafterthekernelswitchesovertovirtualaddresses.Fixup+*offsetsarenolongerneededatthispoint,soupdatetheCPUIDtable+*pointeraccordingly.+*/+voidsnp_cpuid_init(void)+{+if(!cc_platform_has(CC_ATTR_SEV_SNP)){+/* Firmware should not have advertised the feature. */+if(snp_cpuid_active())+panic("Invalid use of SEV-SNP CPUID table.");+return;+}++/* CPUID table should always be available when SEV-SNP is enabled. */+if(!snp_cpuid_active())+sev_es_terminate(1,GHCB_TERM_CPUID);++/* Remove the fixup offset from the cpuid_info pointer. */+cpuid_info=snp_cpuid_info_get_ptr(0);+}++/*+*Itisusefulfromanauditing/testingperspectivetoprovideaneasyway+*fortheguestownertoknowthattheCPUIDtablehasbeeninitializedas+*expected,butthatinitializationhappenstooearlyinboottoprintany+*sortofindicator,andthere'snotreallyanyothergoodplacetodoit.So+*doithere,andwhileatit,goaheadandre-verifythatnothingstrangehas+*happenedbetweenearlybootandnow.+*/+staticint__initsnp_cpuid_check_status(void)+{+if(!cc_platform_has(CC_ATTR_SEV_SNP)){+/* Firmware should not have advertised the feature. */+if(snp_cpuid_active())+panic("Invalid use of SEV-SNP CPUID table.");+return0;+}++/* CPUID table should always be available when SEV-SNP is enabled. */+if(!snp_cpuid_active())+sev_es_terminate(1,GHCB_TERM_CPUID);++pr_info("Using SEV-SNP CPUID table, %d entries present.\n",+cpuid_info->count);++return0;+}++arch_initcall(snp_cpuid_check_status);
Version 2 of GHCB specification defines Non-Automatic-Exit(NAE) to get
the extended guest report. It is similar to the SNP_GET_REPORT ioctl.
The main difference is related to the additional data that will be
returned. The additional data returned is a certificate blob that can
be used by the SNP guest user. The certificate blob layout is defined
in the GHCB specification. The driver simply treats the blob as a opaque
data and copies it to userspace.
Signed-off-by: Brijesh Singh <redacted>
---
Documentation/virt/coco/sevguest.rst | 23 +++++++
drivers/virt/coco/sevguest/sevguest.c | 97 ++++++++++++++++++++++++++-
include/uapi/linux/sev-guest.h | 13 ++++
3 files changed, 131 insertions(+), 2 deletions(-)
@@ -86,6 +86,29 @@ on the various fileds passed in the key derivation request. On success, the snp_derived_key_resp.data will contains the derived key value. See the SEV-SNP specification for further details.++2.3 SNP_GET_EXT_REPORT+----------------------+:Technology: sev-snp+:Type: guest ioctl+:Parameters (in/out): struct snp_ext_report_req+:Returns (out): struct snp_report_resp on success, -negative on error++The SNP_GET_EXT_REPORT ioctl is similar to the SNP_GET_REPORT. The difference is+related to the additional certificate data that is returned with the report.+The certificate data returned is being provided by the hypervisor through the+SNP_SET_EXT_CONFIG.++The ioctl uses the SNP_GUEST_REQUEST (MSG_REPORT_REQ) command provided by the SEV-SNP+firmware to get the attestation report.++On success, the snp_ext_report_resp.data will contain the attestation report+and snp_ext_report_req.certs_address will contain the certificate blob. If the+length of the blob is smaller than expected then snp_ext_report_req.certs_len will+be updated with the expected value.++See GHCB specification for further detail on how to parse the certificate blob.+ Reference ---------
@@ -410,6 +411,88 @@ static int get_derived_key(struct snp_guest_dev *snp_dev, struct snp_guest_requereturnrc;}+staticintget_ext_report(structsnp_guest_dev*snp_dev,structsnp_guest_request_ioctl*arg)+{+structsnp_guest_crypto*crypto=snp_dev->crypto;+structsnp_ext_report_reqreq;+structsnp_report_resp*resp;+intret,npages=0,resp_len;++if(!arg->req_data||!arg->resp_data)+return-EINVAL;++/* Copy the request payload from userspace */+if(copy_from_user(&req,(void__user*)arg->req_data,sizeof(req)))+return-EFAULT;++/* Message version must be non-zero */+if(!req.data.msg_version)+return-EINVAL;++if(req.certs_len){+if(req.certs_len>SEV_FW_BLOB_MAX_SIZE||+!IS_ALIGNED(req.certs_len,PAGE_SIZE))+return-EINVAL;+}++if(req.certs_address&&req.certs_len){+if(!access_ok(req.certs_address,req.certs_len))+return-EFAULT;++/*+*Initializetheintermediatebufferwithallzero's.Thisbuffer+*isusedintheguestrequestmessagetogetthecertsblobfrom+*thehost.Ifhostdoesnotsupplyanycertsinit,thencopy+*zerostoindicatethatcertificatedatawasnotprovided.+*/+memset(snp_dev->certs_data,0,req.certs_len);++npages=req.certs_len>>PAGE_SHIFT;+}++/*+*Theintermediateresponsebufferisusedwhiledecryptingthe+*responsepayload.Makesurethatithasenoughspacetocoverthe+*authtag.+*/+resp_len=sizeof(resp->data)+crypto->a_len;+resp=kzalloc(resp_len,GFP_KERNEL_ACCOUNT);+if(!resp)+return-ENOMEM;++snp_dev->input.data_npages=npages;+ret=handle_guest_request(snp_dev,SVM_VMGEXIT_EXT_GUEST_REQUEST,req.data.msg_version,+SNP_MSG_REPORT_REQ,&req.data.user_data,+sizeof(req.data.user_data),resp->data,resp_len,&arg->fw_err);++/* If certs length is invalid then copy the returned length */+if(arg->fw_err==SNP_GUEST_REQ_INVALID_LEN){+req.certs_len=snp_dev->input.data_npages<<PAGE_SHIFT;++if(copy_to_user((void__user*)arg->req_data,&req,sizeof(req)))+ret=-EFAULT;+}++if(ret)+gotoe_free;++/* Copy the certificate data blob to userspace */+if(req.certs_address&&req.certs_len&&+copy_to_user((void__user*)req.certs_address,snp_dev->certs_data,+req.certs_len)){+ret=-EFAULT;+gotoe_free;+}++/* Copy the response payload to userspace */+if(copy_to_user((void__user*)arg->resp_data,resp,sizeof(*resp)))+ret=-EFAULT;++e_free:+kfree(resp);+returnret;+}+staticlongsnp_guest_ioctl(structfile*file,unsignedintioctl,unsignedlongarg){structsnp_guest_dev*snp_dev=to_snp_dev(file);
@@ -431,6 +514,9 @@ static long snp_guest_ioctl(struct file *file, unsigned int ioctl, unsigned longcaseSNP_GET_DERIVED_KEY:ret=get_derived_key(snp_dev,&input);break;+caseSNP_GET_EXT_REPORT:+ret=get_ext_report(snp_dev,&input);+break;default:break;}
@@ -57,6 +57,16 @@ struct snp_derived_key_resp {__u8data[64];};+structsnp_ext_report_req{+structsnp_report_reqdata;++/* where to copy the certificate blob */+__u64certs_address;++/* length of the certificate blob */+__u32certs_len;+};+#define SNP_GUEST_REQ_IOC_TYPE 'S'/* Get SNP attestation report */
@@ -65,4 +75,7 @@ struct snp_derived_key_resp {/* Get a derived key from the root */#define SNP_GET_DERIVED_KEY _IOWR(SNP_GUEST_REQ_IOC_TYPE, 0x1, struct snp_guest_request_ioctl)+/* Get SNP extended report as defined in the GHCB specification version 2. */+#define SNP_GET_EXT_REPORT _IOWR(SNP_GUEST_REQ_IOC_TYPE, 0x2, struct snp_guest_request_ioctl)+#endif /* __UAPI_LINUX_SEV_GUEST_H_ */
From: Michael Roth <redacted>
When the Confidential Computing blob is located by the boot/compressed
kernel, store a pointer to it in bootparams->cc_blob_address to avoid
the need for the run-time kernel to rescan the EFI config table to find
it again.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/sev.c | 7 +++++++
1 file changed, 7 insertions(+)
@@ -375,4 +375,11 @@ void snp_cpuid_init_boot(struct boot_params *bp)/* It should be safe to read SEV MSR and check features now. */if(!sev_snp_enabled())sev_es_terminate(1,GHCB_TERM_CPUID);++/*+*Passrun-timekernelapointertoCCinfoviaboot_paramssoEFI+*configtabledoesn'tneedtobesearchedagainduringearlystartup+*phase.+*/+bp->cc_blob_address=(u32)(unsignedlong)cc_info;}
From: Michael Roth <redacted>
CPUID instructions generate a #VC exception for SEV-ES/SEV-SNP guests,
for which early handlers are currently set up to handle. In the case
of SEV-SNP, guests can use a configurable location in guest memory
that has been pre-populated with a firmware-validated CPUID table to
look up the relevant CPUID values rather than requesting them from
hypervisor via a VMGEXIT. Add the various hooks in the #VC handlers to
allow CPUID instructions to be handled via the table. The code to
actually configure/enable the table will be added in a subsequent
commit.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/sev.c | 1 +
arch/x86/include/asm/sev-common.h | 2 +
arch/x86/kernel/sev-shared.c | 308 ++++++++++++++++++++++++++++++
arch/x86/kernel/sev.c | 1 +
4 files changed, 312 insertions(+)
@@ -26,6 +61,28 @@ static u16 __ro_after_init ghcb_version;/* Bitmap of SEV features supported by the hypervisor */staticu64__ro_after_initsev_hv_features;+/*+*Thesearestoredin.datasectiontoavoidtheneedtore-parseboot_params+*andregeneratetheCPUIDtable/pointerwhen.bssiscleared.+*/++/*+*TheCPUIDinfocan'talwaysbereferenceddirectlyduetotheneedfor+*pointerfixupsduringinitialstartupphaseofkernelproper,soaccessmust+*bedonethroughthispointer,whichwillbefixedupas-neededduringboot.+*/+staticconststructsnp_cpuid_info*cpuid_info__ro_after_init;++/*+*ThesewillbeinitializedbasedonCPUIDtablesothatnon-present+*all-zeroleaves(forsparsetables)canbedifferentiatedfrom+*invalid/out-of-rangeleaves.Thisisneededsinceall-zeroleaves+*stillneedtobepost-processed.+*/+u32cpuid_std_range_max__ro_after_init;+u32cpuid_hyp_range_max__ro_after_init;+u32cpuid_ext_range_max__ro_after_init;+staticbool__initsev_es_check_cpu_features(void){if(!has_cpuflag(X86_FEATURE_RDRAND)){
@@ -245,6 +302,224 @@ static int sev_cpuid_hv(u32 func, u32 subfunc, u32 *eax, u32 *ebx,return0;}+staticinlineboolsnp_cpuid_active(void)+{+return!!cpuid_info;+}++staticintsnp_cpuid_calc_xsave_size(u64xfeatures_en,u32base_size,+u32*xsave_size,boolcompacted)+{+u32xsave_size_total=base_size;+u64xfeatures_found=0;+inti;++for(i=0;i<cpuid_info->count;i++){+conststructsnp_cpuid_fn*fn=&cpuid_info->fn[i];++if(!(fn->eax_in==0xD&&fn->ecx_in>1&&fn->ecx_in<64))+continue;+if(!(xfeatures_en&(BIT_ULL(fn->ecx_in))))+continue;+if(xfeatures_found&(BIT_ULL(fn->ecx_in)))+continue;++xfeatures_found|=(BIT_ULL(fn->ecx_in));++if(compacted)+xsave_size_total+=fn->eax;+else+xsave_size_total=max(xsave_size_total,+fn->eax+fn->ebx);+}++/*+*EithertheguestsetunsupportedXCR0/XSSbits,orthecorresponding+*entriesintheCPUIDtablewerenotpresent.Thisisnotavalid+*statetobein.+*/+if(xfeatures_found!=(xfeatures_en&GENMASK_ULL(63,2)))+return-EINVAL;++*xsave_size=xsave_size_total;++return0;+}++staticvoidsnp_cpuid_hv(u32func,u32subfunc,u32*eax,u32*ebx,u32*ecx,+u32*edx)+{+/*+*MSRprotocoldoesnotsupportfetchingindexedsubfunction,butis+*sufficienttohandlecurrentfallbackcases.Shouldthatchange,+*makesuretoterminateratherthanignoringtheindexandgrabbing+*randomvalues.Ifthisissuearisesinthefuture,handlingcanbe+*addedheretouseGHCB-pageprotocolforcasesthatoccurlate+*enoughinbootthatGHCBpageisavailable.+*/+if(cpuid_function_is_indexed(func)&&subfunc)+sev_es_terminate(1,GHCB_TERM_CPUID_HV);++if(sev_cpuid_hv(func,0,eax,ebx,ecx,edx))+sev_es_terminate(1,GHCB_TERM_CPUID_HV);+}++staticbool+snp_cpuid_find_validated_func(u32func,u32subfunc,u32*eax,u32*ebx,+u32*ecx,u32*edx)+{+inti;++for(i=0;i<cpuid_info->count;i++){+conststructsnp_cpuid_fn*fn=&cpuid_info->fn[i];++if(fn->eax_in!=func)+continue;++if(cpuid_function_is_indexed(func)&&fn->ecx_in!=subfunc)+continue;++*eax=fn->eax;+*ebx=fn->ebx;+*ecx=fn->ecx;+*edx=fn->edx;++returntrue;+}++returnfalse;+}++staticboolsnp_cpuid_check_range(u32func)+{+if(func<=cpuid_std_range_max||+(func>=0x40000000&&func<=cpuid_hyp_range_max)||+(func>=0x80000000&&func<=cpuid_ext_range_max))+returntrue;++returnfalse;+}++staticintsnp_cpuid_postprocess(u32func,u32subfunc,u32*eax,u32*ebx,+u32*ecx,u32*edx)+{+u32ebx2,ecx2,edx2;++switch(func){+case0x1:+snp_cpuid_hv(func,subfunc,NULL,&ebx2,NULL,&edx2);++/* initial APIC ID */+*ebx=(ebx2&GENMASK(31,24))|(*ebx&GENMASK(23,0));+/* APIC enabled bit */+*edx=(edx2&BIT(9))|(*edx&~BIT(9));++/* OSXSAVE enabled bit */+if(native_read_cr4()&X86_CR4_OSXSAVE)+*ecx|=BIT(27);+break;+case0x7:+/* OSPKE enabled bit */+*ecx&=~BIT(4);+if(native_read_cr4()&X86_CR4_PKE)+*ecx|=BIT(4);+break;+case0xB:+/* extended APIC ID */+snp_cpuid_hv(func,0,NULL,NULL,NULL,edx);+break;+case0xD:{+boolcompacted=false;+u64xcr0=1,xss=0;+u32xsave_size;++if(subfunc!=0&&subfunc!=1)+return0;++if(native_read_cr4()&X86_CR4_OSXSAVE)+xcr0=xgetbv(XCR_XFEATURE_ENABLED_MASK);+if(subfunc==1){+/* Get XSS value if XSAVES is enabled. */+if(*eax&BIT(3)){+unsignedlonglo,hi;++asmvolatile("rdmsr":"=a"(lo),"=d"(hi)+:"c"(MSR_IA32_XSS));+xss=(hi<<32)|lo;+}++/*+*ThePPRandAPMaren'tclearonwhatsizeshouldbe+*encodedin0xD:0x1:EBXwhencompactionisnotenabled+*byeitherXSAVEC(featurebit1)orXSAVES(feature+*bit3)sinceSNP-capablehardwarehasthesefeature+*bitsfixedas1.KVMsetsitto0inthiscase,but+*toavoidthisbecominganissueit'ssafertosimply+*treatthisasunsupportedforSEV-SNPguests.+*/+if(!(*eax&(BIT(1)|BIT(3))))+return-EINVAL;++compacted=true;+}++if(snp_cpuid_calc_xsave_size(xcr0|xss,*ebx,&xsave_size,+compacted))+return-EINVAL;++*ebx=xsave_size;+}+break;+case0x8000001E:+/* extended APIC ID */+snp_cpuid_hv(func,subfunc,eax,&ebx2,&ecx2,NULL);+/* compute ID */+*ebx=(*ebx&GENMASK(31,8))|(ebx2&GENMASK(7,0));+/* node ID */+*ecx=(*ecx&GENMASK(31,8))|(ecx2&GENMASK(7,0));+break;+default:+/* No fix-ups needed, use values as-is. */+break;+}++return0;+}++/*+*Returns-EOPNOTSUPPiffeaturenotenabled.Anyotherreturnvalueshouldbe+*treatedasfatalbycaller.+*/+staticintsnp_cpuid(u32func,u32subfunc,u32*eax,u32*ebx,u32*ecx,+u32*edx)+{+if(!snp_cpuid_active())+return-EOPNOTSUPP;++if(!snp_cpuid_find_validated_func(func,subfunc,eax,ebx,ecx,edx)){+/*+*SomehypervisorswillavoidkeepingtrackofCPUIDentries+*whereallvaluesarezero,sincetheycanbehandledthe+*sameasout-of-rangevalues(all-zero).Thisisusefulhere+*aswellasitallowsvirtuallyallguestconfigurationsto+*workusingasingleSEV-SNPCPUIDtable.+*+*Toallowforthis,thereisaneedtodistinguishbetween+*out-of-rangeentriesandin-rangezeroentries,sincethe+*CPUIDtableentriesareonlyatemplatethatmayneedtobe+*augmentedwithadditionalvaluesforthingslike+*CPU-specificinformationduringpost-processing.Soifit's+*notinthetable,butisstillinthevalidrange,proceed+*withthepost-processing.Otherwise,justreturnzeros.+*/+*eax=*ebx=*ecx=*edx=0;+if(!snp_cpuid_check_range(func))+return0;+}++returnsnp_cpuid_postprocess(func,subfunc,eax,ebx,ecx,edx);+}+/**BootVCHandler-ThisisthefirstVChandlerduringboot,thereisnoGHCB*pageyet,soitonlysupportstheMSRbasedcommunicationwiththe
@@ -252,8 +527,10 @@ static int sev_cpuid_hv(u32 func, u32 subfunc, u32 *eax, u32 *ebx,*/void__initdo_vc_no_ghcb(structpt_regs*regs,unsignedlongexit_code){+unsignedintsubfn=lower_bits(regs->cx,32);unsignedintfn=lower_bits(regs->ax,32);u32eax,ebx,ecx,edx;+intret;/* Only CPUID is supported via MSR protocol */if(exit_code!=SVM_EXIT_CPUID)
From: Michael Roth <redacted>
SEV-SNP guests will be provided the location of special 'secrets' and
'CPUID' pages via the Confidential Computing blob. This blob is
provided to the boot kernel either through an EFI config table entry,
or via a setup_data structure as defined by the Linux Boot Protocol.
Locate the Confidential Computing from these sources and, if found,
use the provided CPUID page/table address to create a copy that the
boot kernel will use when servicing cpuid instructions via a #VC
handler.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/boot/compressed/head_64.S | 1 +
arch/x86/boot/compressed/idt_64.c | 5 +-
arch/x86/boot/compressed/misc.h | 2 +
arch/x86/boot/compressed/sev.c | 79 ++++++++++++++++++++++++++++++
arch/x86/include/asm/sev.h | 14 ++++++
arch/x86/kernel/sev-shared.c | 78 +++++++++++++++++++++++++++++
6 files changed, 178 insertions(+), 1 deletion(-)
@@ -297,3 +297,82 @@ void do_boot_stage2_vc(struct pt_regs *regs, unsigned long exit_code)elseif(result!=ES_RETRY)sev_es_terminate(SEV_TERM_SET_GEN,GHCB_SEV_ES_GEN_REQ);}++/* Search for Confidential Computing blob in the EFI config table. */+staticstructcc_blob_sev_info*snp_find_cc_blob_efi(structboot_params*bp)+{+structcc_blob_sev_info*cc_info;+unsignedlongconf_table_pa;+unsignedintconf_table_len;+boolefi_64;+intret;++ret=efi_get_conf_table(bp,&conf_table_pa,&conf_table_len,&efi_64);+if(ret)+returnNULL;++ret=efi_find_vendor_table(conf_table_pa,conf_table_len,+EFI_CC_BLOB_GUID,efi_64,+(unsignedlong*)&cc_info);+if(ret)+returnNULL;++returncc_info;+}++/*+*InitialsetupofSEV-SNPCPUIDtablereliesoninformationprovided+*bytheConfidentialComputingblob,whichcanbepassedtothebootkernel+*byfirmware/bootloaderinthefollowingways:+*+*-viaanentryintheEFIconfigtable+*-viaasetup_datastructure,asdefinedbytheLinuxBootProtocol+*+*Scanfortheblobinthatorder.+*/+structcc_blob_sev_info*snp_find_cc_blob(structboot_params*bp)+{+structcc_blob_sev_info*cc_info;++cc_info=snp_find_cc_blob_efi(bp);+if(cc_info)+gotofound_cc_info;++cc_info=snp_find_cc_blob_setup_data(bp);+if(!cc_info)+returnNULL;++found_cc_info:+if(cc_info->magic!=CC_BLOB_SEV_HDR_MAGIC)+sev_es_terminate(0,GHCB_SNP_UNSUPPORTED);++returncc_info;+}++voidsnp_cpuid_init_boot(structboot_params*bp)+{+structcc_blob_sev_info*cc_info;+u32eax;++if(!bp)+return;++cc_info=snp_find_cc_blob(bp);+if(!cc_info)+return;++snp_cpuid_info_create(cc_info);++/* SEV-SNP CPUID table is set up now. Do some sanity checks. */+if(!snp_cpuid_active())+sev_es_terminate(1,GHCB_TERM_CPUID);++/* CPUID bits for SEV (bit 1) and SEV-SNP (bit 4) should be enabled. */+eax=native_cpuid_eax(0x8000001f);+if(!(eax&(BIT(4)|BIT(1))))+sev_es_terminate(1,GHCB_TERM_CPUID);++/* It should be safe to read SEV MSR and check features now. */+if(!sev_snp_enabled())+sev_es_terminate(1,GHCB_TERM_CPUID);+}
From: Michael Roth <redacted>
The previously defined Confidential Computing blob is provided to the
kernel via a setup_data structure or EFI config table entry. Currently
these are both checked for by boot/compressed kernel to access the
CPUID table address within it for use with SEV-SNP CPUID enforcement.
To also enable SEV-SNP CPUID enforcement for the run-time kernel,
similar early access to the CPUID table is needed early on while it's
still using the identity-mapped page table set up by boot/compressed,
where global pointers need to be accessed via fixup_pointer().
This isn't much of an issue for accessing setup_data, and the EFI
config table helper code currently used in boot/compressed *could* be
used in this case as well since they both rely on identity-mapping.
However, it has some reliance on EFI helpers/string constants that
would need to be accessed via fixup_pointer(), and fixing it up while
making it shareable between boot/compressed and run-time kernel is
fragile and introduces a good bit of uglyness.
Instead, add a boot_params->cc_blob_address pointer that the
boot/compressed kernel can initialize so that the run-time kernel can
access the CC blob from there instead of re-scanning the EFI config
table.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/bootparam_utils.h | 1 +
arch/x86/include/uapi/asm/bootparam.h | 3 ++-
2 files changed, 3 insertions(+), 1 deletion(-)
Hi Brijesh,
On 08/10/2021 21:04, Brijesh Singh wrote:
SEV-SNP specification provides the guest a mechanisum to communicate with
the PSP without risk from a malicious hypervisor who wishes to read, alter,
drop or replay the messages sent. The driver uses snp_issue_guest_request()
to issue GHCB SNP_GUEST_REQUEST or SNP_EXT_GUEST_REQUEST NAE events to
submit the request to PSP.
The PSP requires that all communication should be encrypted using key
specified through the platform_data.
The userspace can use SNP_GET_REPORT ioctl() to query the guest
attestation report.
See SEV-SNP spec section Guest Messages for more details.
Signed-off-by: Brijesh Singh <redacted>
---
Documentation/virt/coco/sevguest.rst | 77 ++++
drivers/virt/Kconfig | 3 +
drivers/virt/Makefile | 1 +
drivers/virt/coco/sevguest/Kconfig | 9 +
drivers/virt/coco/sevguest/Makefile | 2 +
drivers/virt/coco/sevguest/sevguest.c | 561 ++++++++++++++++++++++++++
drivers/virt/coco/sevguest/sevguest.h | 98 +++++
include/uapi/linux/sev-guest.h | 44 ++
8 files changed, 795 insertions(+)
create mode 100644 Documentation/virt/coco/sevguest.rst
create mode 100644 drivers/virt/coco/sevguest/Kconfig
create mode 100644 drivers/virt/coco/sevguest/Makefile
create mode 100644 drivers/virt/coco/sevguest/sevguest.c
create mode 100644 drivers/virt/coco/sevguest/sevguest.h
create mode 100644 include/uapi/linux/sev-guest.h
On Fri, Oct 08, 2021 at 01:04:14PM -0500, Brijesh Singh wrote:
From: Borislav Petkov <redacted>
Remove all the defines of masks and bit positions for the GHCB MSR
protocol and use comments instead which correspond directly to the spec
so that following those can be a lot easier and straightforward with the
spec opened in parallel to the code.
Aligh vertically while at it.
No functional changes.
Signed-off-by: Borislav Petkov <redacted>
When you handle someone else's patch, you need to add your SOB
underneath to state that fact. I'll add it now but don't forget rule as
it is important to be able to show how a patch found its way upstream.
Like you've done for the next patch. :)
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
On Fri, Oct 08, 2021 at 01:04:13PM -0500, Brijesh Singh wrote:
Yeah, let's add a trivial commit message here anyway:
<-- "Shorten macro names for improved readability."
Hi Brijesh,
On 08/10/2021 21:04, Brijesh Singh wrote:
quoted
SEV-SNP specification provides the guest a mechanisum to communicate with
the PSP without risk from a malicious hypervisor who wishes to read, alter,
drop or replay the messages sent. The driver uses snp_issue_guest_request()
to issue GHCB SNP_GUEST_REQUEST or SNP_EXT_GUEST_REQUEST NAE events to
submit the request to PSP.
The PSP requires that all communication should be encrypted using key
specified through the platform_data.
The userspace can use SNP_GET_REPORT ioctl() to query the guest
attestation report.
See SEV-SNP spec section Guest Messages for more details.
Signed-off-by: Brijesh Singh <redacted>
---
Documentation/virt/coco/sevguest.rst | 77 ++++
drivers/virt/Kconfig | 3 +
drivers/virt/Makefile | 1 +
drivers/virt/coco/sevguest/Kconfig | 9 +
drivers/virt/coco/sevguest/Makefile | 2 +
drivers/virt/coco/sevguest/sevguest.c | 561 ++++++++++++++++++++++++++
drivers/virt/coco/sevguest/sevguest.h | 98 +++++
include/uapi/linux/sev-guest.h | 44 ++
8 files changed, 795 insertions(+)
create mode 100644 Documentation/virt/coco/sevguest.rst
create mode 100644 drivers/virt/coco/sevguest/Kconfig
create mode 100644 drivers/virt/coco/sevguest/Makefile
create mode 100644 drivers/virt/coco/sevguest/sevguest.c
create mode 100644 drivers/virt/coco/sevguest/sevguest.h
create mode 100644 include/uapi/linux/sev-guest.h
Yes, I did caught that during my testing and the hunk to fix it is in
42/42. I missed merging the hunk in this patch and will take care in
next rev. thanks
On Fri, Oct 08, 2021 at 01:04:18PM -0500, Brijesh Singh wrote:
Version 2 of GHCB specification introduced advertisement of a features
that are supported by the hypervisor. Add support to query the HV
features on boot.
Version 2 of GHCB specification adds several new NAEs, most of them are
optional except the hypervisor feature. Now that hypervisor feature NAE
is implemented, so bump the GHCB maximum support protocol version.
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/include/asm/sev-common.h | 3 +++
arch/x86/include/asm/sev.h | 2 +-
arch/x86/include/uapi/asm/svm.h | 2 ++
arch/x86/kernel/sev-shared.c | 30 ++++++++++++++++++++++++++++++
4 files changed, 36 insertions(+), 1 deletion(-)
For the next version, when you add those variables, do this too pls:
@@ -21,10 +21,10 @@**GHCBprotocolversionnegotiatedwiththehypervisor.*/-staticu16__ro_after_initghcb_version;+staticu16ghcb_version__ro_after_init;/* Bitmap of SEV features supported by the hypervisor */-staticu64__ro_after_initsev_hv_features;+staticu64sev_hv_features__ro_after_init;staticbool__initsev_es_check_cpu_features(void){
On Fri, Oct 08, 2021 at 01:04:19PM -0500, Brijesh Singh wrote:
quoted hunk
From: Michael Roth <redacted>
Generally access to MSR_AMD64_SEV is only safe if the 0x8000001F CPUID
leaf indicates SEV support. With SEV-SNP, CPUID responses from the
hypervisor are not considered trustworthy, particularly for 0x8000001F.
SEV-SNP provides a firmware-validated CPUID table to use as an
alternative, but prior to checking MSR_AMD64_SEV there are no
guarantees that this is even an SEV-SNP guest.
Rather than relying on these CPUID values early on, allow SEV-ES and
SEV-SNP guests to instead use a cpuid instruction to trigger a #VC and
have it cache MSR_AMD64_SEV in sev_status, since it is known to be safe
to access MSR_AMD64_SEV if a #VC has triggered.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/sev-shared.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
Ok, you guys are killing me. ;-\
How is bolting some pretty much unrelated code into the early #VC
handler not a hack? Do you not see it?
So sme_enable() is reading MSR_AMD64_SEV and setting up everything
there, including sev_status. If a SNP guest does not trust CPUID, why
can't you attempt to read that MSR there, even if CPUID has lied to the
guest?
And not just slap it somewhere just because it works?
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Michael Roth <hidden> Date: 2021-10-18 18:40:30
On Mon, Oct 18, 2021 at 04:29:07PM +0200, Borislav Petkov wrote:
On Fri, Oct 08, 2021 at 01:04:19PM -0500, Brijesh Singh wrote:
quoted
From: Michael Roth <redacted>
Generally access to MSR_AMD64_SEV is only safe if the 0x8000001F CPUID
leaf indicates SEV support. With SEV-SNP, CPUID responses from the
hypervisor are not considered trustworthy, particularly for 0x8000001F.
SEV-SNP provides a firmware-validated CPUID table to use as an
alternative, but prior to checking MSR_AMD64_SEV there are no
guarantees that this is even an SEV-SNP guest.
Rather than relying on these CPUID values early on, allow SEV-ES and
SEV-SNP guests to instead use a cpuid instruction to trigger a #VC and
have it cache MSR_AMD64_SEV in sev_status, since it is known to be safe
to access MSR_AMD64_SEV if a #VC has triggered.
Signed-off-by: Michael Roth <redacted>
Signed-off-by: Brijesh Singh <redacted>
---
arch/x86/kernel/sev-shared.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
Ok, you guys are killing me. ;-\
How is bolting some pretty much unrelated code into the early #VC
handler not a hack? Do you not see it?
This was the result of my proposal in v5:
> More specifically, the general protocol to determine SNP is enabled
> seems
> to be:
>
> 1) check cpuid 0x8000001f to determine if SEV bit is enabled and SEV
> MSR is available
> 2) check the SEV MSR to see if SEV-SNP bit is set
>
> but the conundrum here is the CPUID page is only valid if SNP is
> enabled, otherwise it can be garbage. So the code to set up the page
> skips those checks initially, and relies on the expectation that UEFI,
> or whatever the initial guest blob was, will only provide a CC_BLOB if
> it already determined SNP is enabled.
>
> It's still possible something goes awry and the kernel gets handed a
> bogus CC_BLOB even though SNP isn't actually enabled. In this case the
> cpuid values could be bogus as well, but the guest will fail
> attestation then and no secrets should be exposed.
>
> There is one thing that could tighten up the check a bit though. Some
> bits of SEV-ES code will use the generation of a #VC as an indicator
> of SEV-ES support, which implies SEV MSR is available without relying
> on hypervisor-provided CPUID bits. I could add a one-time check in
> the cpuid #VC to check SEV MSR for SNP bit, but it would likely
> involve another static __ro_after_init variable store state. If that
> seems worthwhile I can look into that more as well.
Yes, the skipping of checks above sounds weird: why don't you simply
keep the checks order: SEV, -ES, -SNP and then parse CPUID. It'll fail
at attestation eventually, but you'll have the usual flow like with the
rest of the SEV- feature picking apart.
https://lore.kernel.org/lkml/YS3+saDefHwkYwny@zn.tnic/
I'd thought you didn't like the previous approach of having snp_cpuid_init()
defer the CPUID/MSR checks until sme_enable() sets up sev_status later on,
then failing the boot retroactively if SNP bit isn't set but CPUID table
was advertised. So I added those checks in snp_cpuid_init(), along with the
additional #VC-based indicator of SEV-ES/SEV-SNP support as an additional
sanity check of what EFI firmware was providing, since I thought that was
the key concern here.
Now I'm realizing that perhaps your suggestion was to actually defer the
entire CPUID page setup until after sme_enable(). Is that correct?
So sme_enable() is reading MSR_AMD64_SEV and setting up everything
there, including sev_status. If a SNP guest does not trust CPUID, why
can't you attempt to read that MSR there, even if CPUID has lied to the
guest?
If CPUID has lied, that would result in a #GP, rather than a controlled
termination in the various checkers/callers. The latter is easier to
debug.
Additionally, #VC is arguably a better indicator of SEV MSR availability
for SEV-ES/SEV-SNP guests, since it is only generated by ES/SNP hardware
and doesn't rely directly on hypervisor/EFI-provided CPUID values. It
doesn't work for SEV guests, but I don't think it's a bad idea to allow
SEV-ES/SEV-SNP guests to initialize sev_status in #VC handler to make
use of the added assurance.
Is it just the way it's currently implemented as something
cpuid-table-specific that's at issue, or are you opposed to doing so in
general?
Thanks,
Mike
On Mon, Oct 18, 2021 at 01:40:03PM -0500, Michael Roth wrote:
If CPUID has lied, that would result in a #GP, rather than a controlled
termination in the various checkers/callers. The latter is easier to
debug.
Additionally, #VC is arguably a better indicator of SEV MSR availability
for SEV-ES/SEV-SNP guests, since it is only generated by ES/SNP hardware
and doesn't rely directly on hypervisor/EFI-provided CPUID values. It
doesn't work for SEV guests, but I don't think it's a bad idea to allow
SEV-ES/SEV-SNP guests to initialize sev_status in #VC handler to make
use of the added assurance.
Ok, let's take a step back and analyze what we're trying to solve first.
So I'm looking at sme_enable():
1. Code checks SME/SEV support leaf. HV lies and says there's none. So
guest doesn't boot encrypted. Oh well, not a big deal, the cloud vendor
won't be able to give confidentiality to its users => users go away or
do unencrypted like now.
Problem is solved by political and economical pressure.
2. Check SEV and SME bit. HV lies here. Oh well, same as the above.
3. HV lies about 1. and 2. but says that SME/SEV is supported.
Guest attempts to read the MSR Guest explodes due to the #GP. The same
political/economical pressure thing happens.
If the MSR is really there, we've landed at the place where we read the
SEV MSR. Moment of truth - SEV/SNP guests have a communication protocol
which is independent from the HV and all good.
Now, which case am I missing here which justifies the need to do those
acrobatics of causing #VCs just to detect the SEV MSR?
Thx.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
On Fri, Oct 08, 2021 at 01:04:20PM -0500, Brijesh Singh wrote:
quoted hunk
+static bool do_early_sev_setup(void) { if (!sev_es_negotiate_protocol()) sev_es_terminate(SEV_TERM_SET_GEN, GHCB_SEV_ES_PROT_UNSUPPORTED);+ /*+ * If SEV-SNP is enabled, then check if the hypervisor supports the SEV-SNP+ * features.
This and the other comment should say something along the lines of:
"SNP is supported in v2 of the GHCB spec which mandates support for HV
features."
because it wasn't clear to me why we're enforcing that support here.
Thx.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Michael Roth <hidden> Date: 2021-10-20 16:10:53
On Mon, Oct 18, 2021 at 09:18:13PM +0200, Borislav Petkov wrote:
On Mon, Oct 18, 2021 at 01:40:03PM -0500, Michael Roth wrote:
quoted
If CPUID has lied, that would result in a #GP, rather than a controlled
termination in the various checkers/callers. The latter is easier to
debug.
Additionally, #VC is arguably a better indicator of SEV MSR availability
for SEV-ES/SEV-SNP guests, since it is only generated by ES/SNP hardware
and doesn't rely directly on hypervisor/EFI-provided CPUID values. It
doesn't work for SEV guests, but I don't think it's a bad idea to allow
SEV-ES/SEV-SNP guests to initialize sev_status in #VC handler to make
use of the added assurance.
[Sorry for the wall of text, just trying to work through everything.]
Ok, let's take a step back and analyze what we're trying to solve first.
So I'm looking at sme_enable():
I'm not sure if this is pertaining to using the CPUID table prior to
sme_enable(), or just the #VC-based SEV MSR read. The following comments
assume the former. If that assumption is wrong you can basically ignore
the rest of this email :)
[The #VC-based SEV MSR read is not necessary for anything in sme_enable(),
it's simply a way to determine whether the guest is an SNP guest, without
any reliance on CPUID, which seemed useful in the context of doing some
additional sanity checks against the SNP CPUID table and determining that
it's appropriate to use it early on (rather than just trust that this is an
SNP guest by virtue of the CC blob being present, and then failing later
once sme_enable() checks for the SNP feature bits through the normal
mechanism, as was done in v5).]
1. Code checks SME/SEV support leaf. HV lies and says there's none. So
guest doesn't boot encrypted. Oh well, not a big deal, the cloud vendor
won't be able to give confidentiality to its users => users go away or
do unencrypted like now.
Problem is solved by political and economical pressure.
2. Check SEV and SME bit. HV lies here. Oh well, same as the above.
I'd be worried about the possibility that, through some additional exploits
or failures in the attestation flow, a guest owner was tricked into booting
unencrypted on a compromised host and exposing their secrets. Their
attestation process might even do some additional CPUID sanity checks, which
would at the point be via the SNP CPUID table and look legitimate, unaware
that the kernel didn't actually use the SNP CPUID table until after
0x8000001F was parsed (if we were to only initialize it after/as-part-of
sme_enable()).
Fortunately in this scenario I think the guest kernel actually would fail to
boot due to the SNP hardware unconditionally treating code/page tables as
encrypted pages. I tested some of these scenarios just to check, but not
all, and I still don't feel confident enough about it to say that there's
not some way to exploit this by someone who is more clever/persistant than
me.
3. HV lies about 1. and 2. but says that SME/SEV is supported.
Guest attempts to read the MSR Guest explodes due to the #GP. The same
political/economical pressure thing happens.
That's seems likely, but maybe some future hardware bug, or some other
exploit, makes it possible to intercept that MSR read? I don't know, but
if that particular branch of execution can be made less likely by utilizing
SNP CPUID validation I think it makes sense to make use of it.
If the MSR is really there, we've landed at the place where we read the
SEV MSR. Moment of truth - SEV/SNP guests have a communication protocol
which is independent from the HV and all good.
At which point we then switch to using the CPUID table? But at that
point all the previous CPUID checks, both SEV-related/non-SEV-related,
are now possibly not consistent with what's in the CPUID table. Do we
then revalidate? Even a non-malicious hypervisor might provide
inconsistent values between the two sources due to bugs, or SNP
validation suppressing certain feature bits that hypervisor otherwise
exposes, etc. Now all the code after sme_enable() can potentially take
unexpected execution paths, where post-sme_enable() code makes
assumptions about pre-sme_enable() checks that may no longer hold true.
Also, it would be useful from an attestation perspective that the CPUID
bits visible to userspace correspond to what the kernel used during boot,
which wouldn't necessarily be the case if hypervisor-provided values were
used during early boot and potentially put the kernel into some unexpected
state that could persist beyond the point of attestation.
Code-wise, thanks in large part to your suggestions, it really isn't all
that much more complicated to hook in the CPUID table lookup in the #VC
handlers (which are already needed anyway for SEV-ES) early on so all
these checks are against the same trusted (or more-trusted at least)
CPUID source.
Now, which case am I missing here which justifies the need to do those
acrobatics of causing #VCs just to detect the SEV MSR?
There are a few more places where cpuid is utilized prior to
sme_enable():
# In boot/compressed
paging_prepare():
ecx = cpuid(7, 0) # SNP-verified against host values
# check ecx for LA57
# In boot/compressed and kernel proper
verify_cpu():
eax, ebx, ecx, edx = cpuid(0, 0) # SNP-verified against host values
# check eax for range > 0
# check ebx, ecx, edx for "AuthenticAMD" or "GenuineIntel"
if_amd:
edx = cpuid(1, 0) # SNP-verified against host values
# check edx feature bits against REQUIRED_MASK0 (PAE|FPU|PSE|etc.)
eax = cpuid(0x80000001, 0) # SNP-verified against host values
# check eax against REQUIRED_MASK1 (LM|3DNOW)
edx = cpuid(1, 0) # SNP-verified against host values
# check eax against SSE_MASK
# if not set, try to force it on via MSR_K7_HWCR if this is an AMD CPU
# if forcing fails, report no_longmode available
if_intel:
# completely different stuff
It's possible that various lies about the values checked for in
REQUIRED_MASK0/REQUIRED_MASK1, LA57 enablement, etc., can be audited in
similar fashion as you've done above to find nothing concerning, but
what about 5 years from now? And these are all checks/configuration that
can put the kernel in unexpected states that persist beyond the point of
attestation, where we really need to care about the possible effects. If
SNP CPUID validation isn't utilized until after-the-fact, we'd end up
not utilizing it for some of the more 'interesting' CPUID bits.
It's also worth noting that TDX guards against most of this through
CPUID virtualization, where hardware/microcode provides similar
validation for these sorts of CPUID bits in early boot. It's only because
the SEV-SNP CPUID 'virtualization' lives in the guest code that we have to
deal with the additional complexity of initializing the CPUID table early
on. But if both platforms are capable of providing similar assurances then
it seems worthwhile to pursue that.
acrobatics of causing #VCs just to detect the SEV MSR?
The CPUID calls in snp_cpuid_init() weren't added specifically to induce
the #VC-based SEV MSR read, they were added only because I thought the
gist of your earlier suggestions were to do more validation against the
CPUID table advertised by EFI rather than the v5 approach of deferring
them till later after sev_status gets set by sme_enable(), and then
failing later if turned out that SNP CPUID feature bit wasn't set by
sme_enable(). I thought this was motivated by a desire to be more
paranoid about what EFI provides, so once I had the cpuid checks added
in snp_cpuid_init() it seemed like a logical step to further
sanity-check the SNP CPUID bit using a mechanism that was completely
independent of the CPUID table.
I'm not dead set on that at all however, it was only added based on me
[mis-]interpreting your comments as a desire to be less trusting of EFI
as to whether this is an SNP guest or not. But I also don't think it's
a good idea to not utilize the CPUID table until after sme_enable(),
based on the above reasons.
What if we simply add a check in sme_enable() that terminates the guest
if the cc_blob/cpuid table is provided, but the CPUID/MSR checks in
sme_enable() determine that this isn't an SNP guest? That would be similar
to the v5 approach, but in a less roundabout way, and then the cpuid/MSR
checks could be dropped from snp_cpuid_init().
If we did decide that it is useful to use #VC-based initialization of
sev_status there, it would all be self-contained there (but again, that's
a separate thing that I don't have a strong opinion on).
Thanks,
Mike
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
[Sorry for the wall of text, just trying to work through everything.]
And I'm going to respond in a couple of mails just for my own sanity.
I'm not sure if this is pertaining to using the CPUID table prior to
sme_enable(), or just the #VC-based SEV MSR read. The following comments
assume the former. If that assumption is wrong you can basically ignore
the rest of this email :)
This is pertaining to me wanting to show you that the design of this SNP
support needs to be sane and maintainable and every function needs to
make sense not only now but in the future.
In this particular example, we should set sev_status *once*, *before*
anything accesses it so that it is prepared when something needs it. Not
do a #VC and go, "oh, btw, is sev_status set? No? Ok, lemme set it."
which basically means our design is seriously lacking.
And I had suggested a similar thing for TDX and tglx was 100% right in
shooting it down because we do properly designed things - not, get stuff
in so that vendor is happy and then, once the vendor programmers have
disappeared to do their next enablement task, the maintainers get to mop
up and maintain it forever.
Because this mopping up doesn't scale - trust me.
[The #VC-based SEV MSR read is not necessary for anything in sme_enable(),
it's simply a way to determine whether the guest is an SNP guest, without
any reliance on CPUID, which seemed useful in the context of doing some
additional sanity checks against the SNP CPUID table and determining that
it's appropriate to use it early on (rather than just trust that this is an
SNP guest by virtue of the CC blob being present, and then failing later
once sme_enable() checks for the SNP feature bits through the normal
mechanism, as was done in v5).]
So you need to make up your mind here design-wise, what you wanna do.
The proper thing to do would be, to detect *everything*, detect whether
this is an SNP guest, yadda yadda, everything your code is going to need
later on, and then be done with it.
Then you continue with the boot and now your other code queries
everything that has been detected up til now and uses it.
End of mail 1.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
quoted
1. Code checks SME/SEV support leaf. HV lies and says there's none. So
guest doesn't boot encrypted. Oh well, not a big deal, the cloud vendor
won't be able to give confidentiality to its users => users go away or
do unencrypted like now.
Problem is solved by political and economical pressure.
2. Check SEV and SME bit. HV lies here. Oh well, same as the above.
I'd be worried about the possibility that, through some additional exploits
or failures in the attestation flow,
Well, that puts forward an important question: how do you verify
*reliably* that this is an SNP guest?
- attestation?
- CPUID?
- anything else?
I don't see this written down anywhere. Because this assumption will
guide the design in the kernel.
a guest owner was tricked into booting unencrypted on a compromised
host and exposing their secrets. Their attestation process might even
do some additional CPUID sanity checks, which would at the point
be via the SNP CPUID table and look legitimate, unaware that the
kernel didn't actually use the SNP CPUID table until after 0x8000001F
was parsed (if we were to only initialize it after/as-part-of
sme_enable()).
So what happens with that guest owner later?
How is she to notice that she booted unencrypted?
Fortunately in this scenario I think the guest kernel actually would fail to
boot due to the SNP hardware unconditionally treating code/page tables as
encrypted pages. I tested some of these scenarios just to check, but not
all, and I still don't feel confident enough about it to say that there's
not some way to exploit this by someone who is more clever/persistant than
me.
All this design needs to be preceded with: "We protect against cases A,
B and C and not against D, E, etc."
So that it is clear to all parties involved what we're working with and
what we're protecting against and what we're *not* protecting against.
End of mail 2, more later.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Peter Gonda <hidden> Date: 2021-10-20 21:34:01
On Fri, Oct 8, 2021 at 12:06 PM Brijesh Singh [off-list ref] wrote:
quoted hunk
SEV-SNP specification provides the guest a mechanisum to communicate with
the PSP without risk from a malicious hypervisor who wishes to read, alter,
drop or replay the messages sent. The driver uses snp_issue_guest_request()
to issue GHCB SNP_GUEST_REQUEST or SNP_EXT_GUEST_REQUEST NAE events to
submit the request to PSP.
The PSP requires that all communication should be encrypted using key
specified through the platform_data.
The userspace can use SNP_GET_REPORT ioctl() to query the guest
attestation report.
See SEV-SNP spec section Guest Messages for more details.
Signed-off-by: Brijesh Singh <redacted>
---
Documentation/virt/coco/sevguest.rst | 77 ++++
drivers/virt/Kconfig | 3 +
drivers/virt/Makefile | 1 +
drivers/virt/coco/sevguest/Kconfig | 9 +
drivers/virt/coco/sevguest/Makefile | 2 +
drivers/virt/coco/sevguest/sevguest.c | 561 ++++++++++++++++++++++++++
drivers/virt/coco/sevguest/sevguest.h | 98 +++++
include/uapi/linux/sev-guest.h | 44 ++
8 files changed, 795 insertions(+)
create mode 100644 Documentation/virt/coco/sevguest.rst
create mode 100644 drivers/virt/coco/sevguest/Kconfig
create mode 100644 drivers/virt/coco/sevguest/Makefile
create mode 100644 drivers/virt/coco/sevguest/sevguest.c
create mode 100644 drivers/virt/coco/sevguest/sevguest.h
create mode 100644 include/uapi/linux/sev-guest.h
@@ -0,0 +1,77 @@+.. SPDX-License-Identifier: GPL-2.0++===================================================================+The Definitive SEV Guest API Documentation+===================================================================++1. General description+======================++The SEV API is a set of ioctls that are used by the guest or hypervisor+to get or set certain aspect of the SEV virtual machine. The ioctls belong+to the following classes:++- Hypervisor ioctls: These query and set global attributes which affect the+ whole SEV firmware. These ioctl are used by platform provision tools.++- Guest ioctls: These query and set attributes of the SEV virtual machine.++2. API description+==================++This section describes ioctls that can be used to query or set SEV guests.+For each ioctl, the following information is provided along with a+description:++ Technology:+ which SEV techology provides this ioctl. sev, sev-es, sev-snp or all.++ Type:+ hypervisor or guest. The ioctl can be used inside the guest or the+ hypervisor.++ Parameters:+ what parameters are accepted by the ioctl.++ Returns:+ the return value. General error numbers (ENOMEM, EINVAL)+ are not detailed, but errors with specific meanings are.++The guest ioctl should be issued on a file descriptor of the /dev/sev-guest device.+The ioctl accepts struct snp_user_guest_request. The input and output structure is+specified through the req_data and resp_data field respectively. If the ioctl fails+to execute due to a firmware error, then fw_err code will be set.++::+ struct snp_guest_request_ioctl {+ /* Request and response structure address */+ __u64 req_data;+ __u64 resp_data;++ /* firmware error code on failure (see psp-sev.h) */+ __u64 fw_err;+ };++2.1 SNP_GET_REPORT+------------------++:Technology: sev-snp+:Type: guest ioctl+:Parameters (in): struct snp_report_req+:Returns (out): struct snp_report_resp on success, -negative on error++The SNP_GET_REPORT ioctl can be used to query the attestation report from the+SEV-SNP firmware. The ioctl uses the SNP_GUEST_REQUEST (MSG_REPORT_REQ) command+provided by the SEV-SNP firmware to query the attestation report.++On success, the snp_report_resp.data will contains the report. The report+will contain the format described in the SEV-SNP specification. See the SEV-SNP+specification for further details.+++Reference+---------++SEV-SNP and GHCB specification: developer.amd.com/sev++The driver is based on SEV-SNP firmware spec 0.9 and GHCB spec version 2.0.
@@ -0,0 +1,561 @@+// SPDX-License-Identifier: GPL-2.0-only+/*+*AMDSecureEncryptedVirtualizationNestedPaging(SEV-SNP)guestrequestinterface+*+*Copyright(C)2021AdvancedMicroDevices,Inc.+*+*Author:BrijeshSingh<brijesh.singh@amd.com>+*/++#include<linux/module.h>+#include<linux/kernel.h>+#include<linux/types.h>+#include<linux/mutex.h>+#include<linux/io.h>+#include<linux/platform_device.h>+#include<linux/miscdevice.h>+#include<linux/set_memory.h>+#include<linux/fs.h>+#include<crypto/aead.h>+#include<linux/scatterlist.h>+#include<linux/psp-sev.h>+#include<uapi/linux/sev-guest.h>+#include<uapi/linux/psp-sev.h>++#include<asm/svm.h>+#include<asm/sev.h>++#include"sevguest.h"++#define DEVICE_NAME "sev-guest"+#define AAD_LEN 48+#define MSG_HDR_VER 1++structsnp_guest_crypto{+structcrypto_aead*tfm;+u8*iv,*authtag;+intiv_len,a_len;+};++structsnp_guest_dev{+structdevice*dev;+structmiscdevicemisc;++structsnp_guest_crypto*crypto;+structsnp_guest_msg*request,*response;+structsnp_secrets_page_layout*layout;+structsnp_req_datainput;+u32*os_area_msg_seqno;+};++staticu32vmpck_id;+module_param(vmpck_id,uint,0444);+MODULE_PARM_DESC(vmpck_id,"The VMPCK ID to use when communicating with the PSP.");++staticDEFINE_MUTEX(snp_cmd_mutex);++staticinlineu64__snp_get_msg_seqno(structsnp_guest_dev*snp_dev)+{+u64count;++/* Read the current message sequence counter from secrets pages */+count=*snp_dev->os_area_msg_seqno;++returncount+1;+}++/* Return a non-zero on success */+staticu64snp_get_msg_seqno(structsnp_guest_dev*snp_dev)+{+u64count=__snp_get_msg_seqno(snp_dev);++/*+*ThemessagesequencecounterfortheSNPguestrequestisa64-bit+*valuebuttheversion2ofGHCBspecificationdefinesa32-bitstorage+*fortheit.Ifthecounterexceedsthe32-bitvaluethenreturnzero.+*Thecallershouldcheckthereturnvalue,butifthecallerhappento+*notcheckthevalueanduseit,thenthefirmwaretreatszeroasan+*invalidnumberandwillfailthemessagerequest.+*/+if(count>=UINT_MAX){+pr_err_ratelimited("SNP guest request message sequence counter overflow\n");+return0;+}++returncount;+}++staticvoidsnp_inc_msg_seqno(structsnp_guest_dev*snp_dev)+{+/*+*ThecounterisalsoincrementedbythePSP,soincrementitby2+*andsaveinsecretspage.+*/+*snp_dev->os_area_msg_seqno+=2;+}++staticinlinestructsnp_guest_dev*to_snp_dev(structfile*file)+{+structmiscdevice*dev=file->private_data;++returncontainer_of(dev,structsnp_guest_dev,misc);+}++staticstructsnp_guest_crypto*init_crypto(structsnp_guest_dev*snp_dev,u8*key,size_tkeylen)+{+structsnp_guest_crypto*crypto;++crypto=kzalloc(sizeof(*crypto),GFP_KERNEL_ACCOUNT);+if(!crypto)+returnNULL;++crypto->tfm=crypto_alloc_aead("gcm(aes)",0,0);+if(IS_ERR(crypto->tfm))+gotoe_free;++if(crypto_aead_setkey(crypto->tfm,key,keylen))+gotoe_free_crypto;++crypto->iv_len=crypto_aead_ivsize(crypto->tfm);+if(crypto->iv_len<12){+dev_err(snp_dev->dev,"IV length is less than 12.\n");+gotoe_free_crypto;+}++crypto->iv=kmalloc(crypto->iv_len,GFP_KERNEL_ACCOUNT);+if(!crypto->iv)+gotoe_free_crypto;++if(crypto_aead_authsize(crypto->tfm)>MAX_AUTHTAG_LEN){+if(crypto_aead_setauthsize(crypto->tfm,MAX_AUTHTAG_LEN)){+dev_err(snp_dev->dev,"failed to set authsize to %d\n",MAX_AUTHTAG_LEN);+gotoe_free_crypto;+}+}++crypto->a_len=crypto_aead_authsize(crypto->tfm);+crypto->authtag=kmalloc(crypto->a_len,GFP_KERNEL_ACCOUNT);+if(!crypto->authtag)+gotoe_free_crypto;++returncrypto;++e_free_crypto:+crypto_free_aead(crypto->tfm);+e_free:+kfree(crypto->iv);+kfree(crypto->authtag);+kfree(crypto);++returnNULL;+}++staticvoiddeinit_crypto(structsnp_guest_crypto*crypto)+{+crypto_free_aead(crypto->tfm);+kfree(crypto->iv);+kfree(crypto->authtag);+kfree(crypto);+}++staticintenc_dec_message(structsnp_guest_crypto*crypto,structsnp_guest_msg*msg,+u8*src_buf,u8*dst_buf,size_tlen,boolenc)+{+structsnp_guest_msg_hdr*hdr=&msg->hdr;+structscatterlistsrc[3],dst[3];+DECLARE_CRYPTO_WAIT(wait);+structaead_request*req;+intret;++req=aead_request_alloc(crypto->tfm,GFP_KERNEL);+if(!req)+return-ENOMEM;++/*+*AEADmemoryoperations:+*+------AAD-------+-------DATA-----+----AUTHTAG----++*|msgheader|plaintext|hdr->authtag|+*|bytes30h-5Fh|or||+*||cipher||+*+------------------+------------------+----------------++*/+sg_init_table(src,3);+sg_set_buf(&src[0],&hdr->algo,AAD_LEN);+sg_set_buf(&src[1],src_buf,hdr->msg_sz);+sg_set_buf(&src[2],hdr->authtag,crypto->a_len);++sg_init_table(dst,3);+sg_set_buf(&dst[0],&hdr->algo,AAD_LEN);+sg_set_buf(&dst[1],dst_buf,hdr->msg_sz);+sg_set_buf(&dst[2],hdr->authtag,crypto->a_len);++aead_request_set_ad(req,AAD_LEN);+aead_request_set_tfm(req,crypto->tfm);+aead_request_set_callback(req,0,crypto_req_done,&wait);++aead_request_set_crypt(req,src,dst,len,crypto->iv);+ret=crypto_wait_req(enc?crypto_aead_encrypt(req):crypto_aead_decrypt(req),&wait);++aead_request_free(req);+returnret;+}++staticint__enc_payload(structsnp_guest_dev*snp_dev,structsnp_guest_msg*msg,+void*plaintext,size_tlen)+{+structsnp_guest_crypto*crypto=snp_dev->crypto;+structsnp_guest_msg_hdr*hdr=&msg->hdr;++memset(crypto->iv,0,crypto->iv_len);+memcpy(crypto->iv,&hdr->msg_seqno,sizeof(hdr->msg_seqno));++returnenc_dec_message(crypto,msg,plaintext,msg->payload,len,true);+}++staticintdec_payload(structsnp_guest_dev*snp_dev,structsnp_guest_msg*msg,+void*plaintext,size_tlen)+{+structsnp_guest_crypto*crypto=snp_dev->crypto;+structsnp_guest_msg_hdr*hdr=&msg->hdr;++/* Build IV with response buffer sequence number */+memset(crypto->iv,0,crypto->iv_len);+memcpy(crypto->iv,&hdr->msg_seqno,sizeof(hdr->msg_seqno));++returnenc_dec_message(crypto,msg,msg->payload,plaintext,len,false);+}++staticintverify_and_dec_payload(structsnp_guest_dev*snp_dev,void*payload,u32sz)+{+structsnp_guest_crypto*crypto=snp_dev->crypto;+structsnp_guest_msg*resp=snp_dev->response;+structsnp_guest_msg*req=snp_dev->request;+structsnp_guest_msg_hdr*req_hdr=&req->hdr;+structsnp_guest_msg_hdr*resp_hdr=&resp->hdr;++dev_dbg(snp_dev->dev,"response [seqno %lld type %d version %d sz %d]\n",+resp_hdr->msg_seqno,resp_hdr->msg_type,resp_hdr->msg_version,resp_hdr->msg_sz);++/* Verify that the sequence counter is incremented by 1 */+if(unlikely(resp_hdr->msg_seqno!=(req_hdr->msg_seqno+1)))+return-EBADMSG;++/* Verify response message type and version number. */+if(resp_hdr->msg_type!=(req_hdr->msg_type+1)||+resp_hdr->msg_version!=req_hdr->msg_version)+return-EBADMSG;++/*+*Ifthemessagesizeisgreaterthanourbufferlengththenreturn+*anerror.+*/+if(unlikely((resp_hdr->msg_sz+crypto->a_len)>sz))+return-EBADMSG;++returndec_payload(snp_dev,resp,payload,resp_hdr->msg_sz+crypto->a_len);+}++staticboolenc_payload(structsnp_guest_dev*snp_dev,u64seqno,intversion,u8type,+void*payload,size_tsz)+{+structsnp_guest_msg*req=snp_dev->request;+structsnp_guest_msg_hdr*hdr=&req->hdr;++memset(req,0,sizeof(*req));++hdr->algo=SNP_AEAD_AES_256_GCM;+hdr->hdr_version=MSG_HDR_VER;+hdr->hdr_sz=sizeof(*hdr);+hdr->msg_type=type;+hdr->msg_version=version;+hdr->msg_seqno=seqno;+hdr->msg_vmpck=vmpck_id;+hdr->msg_sz=sz;++/* Verify the sequence number is non-zero */+if(!hdr->msg_seqno)+return-ENOSR;++dev_dbg(snp_dev->dev,"request [seqno %lld type %d version %d sz %d]\n",+hdr->msg_seqno,hdr->msg_type,hdr->msg_version,hdr->msg_sz);++return__enc_payload(snp_dev,req,payload,sz);+}++staticinthandle_guest_request(structsnp_guest_dev*snp_dev,u64exit_code,intmsg_ver,+u8type,void*req_buf,size_treq_sz,void*resp_buf,+u32resp_sz,__u64*fw_err)+{+unsignedlongerr;+u64seqno;+intrc;++/* Get message sequence and verify that its a non-zero */+seqno=snp_get_msg_seqno(snp_dev);+if(!seqno)+return-EIO;++memset(snp_dev->response,0,sizeof(*snp_dev->response));++/* Encrypt the userspace provided payload */+rc=enc_payload(snp_dev,seqno,msg_ver,type,req_buf,req_sz);+if(rc)+returnrc;++/* Call firmware to process the request */+rc=snp_issue_guest_request(exit_code,&snp_dev->input,&err);+if(fw_err)+*fw_err=err;++if(rc)+returnrc;++rc=verify_and_dec_payload(snp_dev,resp_buf,resp_sz);+if(rc)+returnrc;++/* Increment to new message sequence after the command is successful. */+snp_inc_msg_seqno(snp_dev);
Thanks for updating this sequence number logic. But I still have some
concerns. In verify_and_dec_payload() we check the encryption header
but all these fields are accessible to the hypervisor, meaning it can
change the header and cause this sequence number to not get
incremented. We then will reuse the sequence number for the next
command, which isn't great for AES GCM. It seems very hard to tell if
the FW actually got our request and created a response there by
incrementing the sequence number by 2, or if the hypervisor is acting
in bad faith. It seems like to be safe we need to completely stop
using this vmpck if we cannot confirm the PSP has gotten our request
and created a response. Thoughts?
quoted hunk
++ return 0;+}++static int get_report(struct snp_guest_dev *snp_dev, struct snp_guest_request_ioctl *arg)+{+ struct snp_guest_crypto *crypto = snp_dev->crypto;+ struct snp_report_resp *resp;+ struct snp_report_req req;+ int rc, resp_len;++ if (!arg->req_data || !arg->resp_data)+ return -EINVAL;++ /* Copy the request payload from userspace */+ if (copy_from_user(&req, (void __user *)arg->req_data, sizeof(req)))+ return -EFAULT;++ /* Message version must be non-zero */+ if (!req.msg_version)+ return -EINVAL;++ /*+ * The intermediate response buffer is used while decrypting the+ * response payload. Make sure that it has enough space to cover the+ * authtag.+ */+ resp_len = sizeof(resp->data) + crypto->a_len;+ resp = kzalloc(resp_len, GFP_KERNEL_ACCOUNT);+ if (!resp)+ return -ENOMEM;++ /* Issue the command to get the attestation report */+ rc = handle_guest_request(snp_dev, SVM_VMGEXIT_GUEST_REQUEST, req.msg_version,+ SNP_MSG_REPORT_REQ, &req.user_data, sizeof(req.user_data),+ resp->data, resp_len, &arg->fw_err);+ if (rc)+ goto e_free;++ /* Copy the response payload to userspace */+ if (copy_to_user((void __user *)arg->resp_data, resp, sizeof(*resp)))+ rc = -EFAULT;++e_free:+ kfree(resp);+ return rc;+}++static long snp_guest_ioctl(struct file *file, unsigned int ioctl, unsigned long arg)+{+ struct snp_guest_dev *snp_dev = to_snp_dev(file);+ void __user *argp = (void __user *)arg;+ struct snp_guest_request_ioctl input;+ int ret = -ENOTTY;++ if (copy_from_user(&input, argp, sizeof(input)))+ return -EFAULT;++ input.fw_err = 0;++ mutex_lock(&snp_cmd_mutex);++ switch (ioctl) {+ case SNP_GET_REPORT:+ ret = get_report(snp_dev, &input);+ break;+ default:+ break;+ }++ mutex_unlock(&snp_cmd_mutex);++ if (input.fw_err && copy_to_user(argp, &input, sizeof(input)))+ return -EFAULT;++ return ret;+}++static void free_shared_pages(void *buf, size_t sz)+{+ unsigned int npages = PAGE_ALIGN(sz) >> PAGE_SHIFT;++ if (!buf)+ return;++ /* If fail to restore the encryption mask then leak it. */+ if (WARN_ONCE(set_memory_encrypted((unsigned long)buf, npages),+ "Failed to restore encryption mask (leak it)\n"))+ return;++ __free_pages(virt_to_page(buf), get_order(sz));+}++static void *alloc_shared_pages(size_t sz)+{+ unsigned int npages = PAGE_ALIGN(sz) >> PAGE_SHIFT;+ struct page *page;+ int ret;++ page = alloc_pages(GFP_KERNEL_ACCOUNT, get_order(sz));+ if (IS_ERR(page))+ return NULL;++ ret = set_memory_decrypted((unsigned long)page_address(page), npages);+ if (ret) {+ pr_err("SEV-SNP: failed to mark page shared, ret=%d\n", ret);+ __free_pages(page, get_order(sz));+ return NULL;+ }++ return page_address(page);+}++static const struct file_operations snp_guest_fops = {+ .owner = THIS_MODULE,+ .unlocked_ioctl = snp_guest_ioctl,+};++static u8 *get_vmpck(int id, struct snp_secrets_page_layout *layout, u32 **seqno)+{+ u8 *key = NULL;++ switch (id) {+ case 0:+ *seqno = &layout->os_area.msg_seqno_0;+ key = layout->vmpck0;+ break;+ case 1:+ *seqno = &layout->os_area.msg_seqno_1;+ key = layout->vmpck1;+ break;+ case 2:+ *seqno = &layout->os_area.msg_seqno_2;+ key = layout->vmpck2;+ break;+ case 3:+ *seqno = &layout->os_area.msg_seqno_3;+ key = layout->vmpck3;+ break;+ default:+ break;+ }++ return NULL;+}++static int __init snp_guest_probe(struct platform_device *pdev)+{+ struct snp_secrets_page_layout *layout;+ struct snp_guest_platform_data *data;+ struct device *dev = &pdev->dev;+ struct snp_guest_dev *snp_dev;+ struct miscdevice *misc;+ u8 *vmpck;+ int ret;++ if (!dev->platform_data)+ return -ENODEV;++ data = (struct snp_guest_platform_data *)dev->platform_data;+ layout = (__force void *)ioremap_encrypted(data->secrets_gpa, PAGE_SIZE);+ if (!layout)+ return -ENODEV;++ ret = -ENOMEM;+ snp_dev = devm_kzalloc(&pdev->dev, sizeof(struct snp_guest_dev), GFP_KERNEL);+ if (!snp_dev)+ goto e_fail;++ ret = -EINVAL;+ vmpck = get_vmpck(vmpck_id, layout, &snp_dev->os_area_msg_seqno);+ if (!vmpck) {+ dev_err(dev, "invalid vmpck id %d\n", vmpck_id);+ goto e_fail;+ }++ platform_set_drvdata(pdev, snp_dev);+ snp_dev->dev = dev;+ snp_dev->layout = layout;++ /* Allocate the shared page used for the request and response message. */+ snp_dev->request = alloc_shared_pages(sizeof(struct snp_guest_msg));+ if (!snp_dev->request)+ goto e_fail;++ snp_dev->response = alloc_shared_pages(sizeof(struct snp_guest_msg));+ if (!snp_dev->response)+ goto e_fail;++ ret = -EIO;+ snp_dev->crypto = init_crypto(snp_dev, vmpck, VMPCK_KEY_LEN);+ if (!snp_dev->crypto)+ goto e_fail;++ misc = &snp_dev->misc;+ misc->minor = MISC_DYNAMIC_MINOR;+ misc->name = DEVICE_NAME;+ misc->fops = &snp_guest_fops;++ /* initial the input address for guest request */+ snp_dev->input.req_gpa = __pa(snp_dev->request);+ snp_dev->input.resp_gpa = __pa(snp_dev->response);++ ret = misc_register(misc);+ if (ret)+ goto e_fail;++ dev_dbg(dev, "Initialized SNP guest driver (using vmpck_id %d)\n", vmpck_id);+ return 0;++e_fail:+ iounmap(layout);+ free_shared_pages(snp_dev->request, sizeof(struct snp_guest_msg));+ free_shared_pages(snp_dev->response, sizeof(struct snp_guest_msg));++ return ret;+}++static int __exit snp_guest_remove(struct platform_device *pdev)+{+ struct snp_guest_dev *snp_dev = platform_get_drvdata(pdev);++ free_shared_pages(snp_dev->request, sizeof(struct snp_guest_msg));+ free_shared_pages(snp_dev->response, sizeof(struct snp_guest_msg));+ deinit_crypto(snp_dev->crypto);+ misc_deregister(&snp_dev->misc);++ return 0;+}++static struct platform_driver snp_guest_driver = {+ .remove = __exit_p(snp_guest_remove),+ .driver = {+ .name = "snp-guest",+ },+};++module_platform_driver_probe(snp_guest_driver, snp_guest_probe);++MODULE_AUTHOR("Brijesh Singh <brijesh.singh@amd.com>");+MODULE_LICENSE("GPL");+MODULE_VERSION("1.0.0");+MODULE_DESCRIPTION("AMD SNP Guest Driver");
@@ -0,0 +1,98 @@+/* SPDX-License-Identifier: GPL-2.0-only */+/*+*Copyright(C)2021AdvancedMicroDevices,Inc.+*+*Author:BrijeshSingh<brijesh.singh@amd.com>+*+*SEV-SNPAPIspecisavailableathttps://developer.amd.com/sev+*/++#ifndef __LINUX_SEVGUEST_H_+#define __LINUX_SEVGUEST_H_++#include<linux/types.h>++#define MAX_AUTHTAG_LEN 32++/* See SNP spec SNP_GUEST_REQUEST section for the structure */+enummsg_type{+SNP_MSG_TYPE_INVALID=0,+SNP_MSG_CPUID_REQ,+SNP_MSG_CPUID_RSP,+SNP_MSG_KEY_REQ,+SNP_MSG_KEY_RSP,+SNP_MSG_REPORT_REQ,+SNP_MSG_REPORT_RSP,+SNP_MSG_EXPORT_REQ,+SNP_MSG_EXPORT_RSP,+SNP_MSG_IMPORT_REQ,+SNP_MSG_IMPORT_RSP,+SNP_MSG_ABSORB_REQ,+SNP_MSG_ABSORB_RSP,+SNP_MSG_VMRK_REQ,+SNP_MSG_VMRK_RSP,++SNP_MSG_TYPE_MAX+};++enumaead_algo{+SNP_AEAD_INVALID,+SNP_AEAD_AES_256_GCM,+};++structsnp_guest_msg_hdr{+u8authtag[MAX_AUTHTAG_LEN];+u64msg_seqno;+u8rsvd1[8];+u8algo;+u8hdr_version;+u16hdr_sz;+u8msg_type;+u8msg_version;+u16msg_sz;+u32rsvd2;+u8msg_vmpck;+u8rsvd3[35];+}__packed;++structsnp_guest_msg{+structsnp_guest_msg_hdrhdr;+u8payload[4000];+}__packed;++/*+*Thesecretspagecontains96-bytesofreservedfieldthatcanbeusedby+*theguestOS.TheguestOSusestheareatosavethemessagesequence+*numberforeachVMPCK.+*+*SeetheGHCBspecsectionSecretpagelayoutfortheformatforthisarea.+*/+structsecrets_os_area{+u32msg_seqno_0;+u32msg_seqno_1;+u32msg_seqno_2;+u32msg_seqno_3;+u64ap_jump_table_pa;+u8rsvd[40];+u8guest_usage[32];+}__packed;++#define VMPCK_KEY_LEN 32++/* See the SNP spec version 0.9 for secrets page format */+structsnp_secrets_page_layout{+u32version;+u32imien:1,+rsvd1:31;+u32fms;+u32rsvd2;+u8gosvw[16];+u8vmpck0[VMPCK_KEY_LEN];+u8vmpck1[VMPCK_KEY_LEN];+u8vmpck2[VMPCK_KEY_LEN];+u8vmpck3[VMPCK_KEY_LEN];+structsecrets_os_areaos_area;+u8rsvd3[3840];+}__packed;++#endif /* __LINUX_SNP_GUEST_H__ */
@@ -0,0 +1,44 @@+/* SPDX-License-Identifier: GPL-2.0-only WITH Linux-syscall-note */+/*+*UserspaceinterfaceforAMDSEVandSEV-SNPguestdriver.+*+*Copyright(C)2021AdvancedMicroDevices,Inc.+*+*Author:BrijeshSingh<brijesh.singh@amd.com>+*+*SEVAPIspecificationisavailableat:https://developer.amd.com/sev/+*/++#ifndef __UAPI_LINUX_SEV_GUEST_H_+#define __UAPI_LINUX_SEV_GUEST_H_++#include<linux/types.h>++structsnp_report_req{+/* message version number (must be non-zero) */+__u8msg_version;++/* user data that should be included in the report */+__u8user_data[64];+};++structsnp_report_resp{+/* response data, see SEV-SNP spec for the format */+__u8data[4000];+};++structsnp_guest_request_ioctl{+/* Request and response structure address */+__u64req_data;+__u64resp_data;++/* firmware error code on failure (see psp-sev.h) */+__u64fw_err;+};++#define SNP_GUEST_REQ_IOC_TYPE 'S'++/* Get SNP attestation report */+#define SNP_GET_REPORT _IOWR(SNP_GUEST_REQ_IOC_TYPE, 0x0, struct snp_guest_request_ioctl)++#endif /* __UAPI_LINUX_SEV_GUEST_H_ */--
From: Michael Roth <hidden> Date: 2021-10-21 02:08:39
On Wed, Oct 20, 2021 at 08:01:07PM +0200, Borislav Petkov wrote:
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
quoted
[Sorry for the wall of text, just trying to work through everything.]
And I'm going to respond in a couple of mails just for my own sanity.
quoted
I'm not sure if this is pertaining to using the CPUID table prior to
sme_enable(), or just the #VC-based SEV MSR read. The following comments
assume the former. If that assumption is wrong you can basically ignore
the rest of this email :)
This is pertaining to me wanting to show you that the design of this SNP
support needs to be sane and maintainable and every function needs to
make sense not only now but in the future.
Absolutely.
In this particular example, we should set sev_status *once*, *before*
anything accesses it so that it is prepared when something needs it. Not
do a #VC and go, "oh, btw, is sev_status set? No? Ok, lemme set it."
which basically means our design is seriously lacking.
Yes, taking a step back there are some things that could probably be
improved upon there.
currently:
- boot kernel initializes sev_status in set_sev_encryption_mask()
- run-time kernel initializes sev_status in sme_enable()
with this series the following are introduced:
- boot kernel initializes sev_status on-demand in sev_snp_enabled()
- initially used by snp_cpuid_init_boot(), which happens before
set_sev_encryption_mask()
- run-time kernel initializes sev_status on-demand via #VC handler
- initially used by snp_cpuid_init(), which happens before
sme_enable()
Fortunately, all the code makes use of sev_status to get at the SEV MSR
bits, so breaking the appropriate bits out of sme_enable() into an earlier
sev_init() routine that's the exclusive writer of sev_status sounds like a
promising approach.
It makes sense to do it immediately after the first #VC handler is set
up, so CPUID is available, and since that's where SNP CPUID table
initialization would need to happen if it's to be made available in
#VC handler.
It may even be similar enough between boot/compressed and run-time kernel
that it could be a shared routine in sev-shared.c. But then again it also
sounds like the appropriate place to move the snp_cpuid_init*() calls,
and locating the cc_blob, and since there's differences there it might make
sense to keep the boot/compressed and kernel proper sev_init() routines
separate to avoid #ifdeffery).
Not to get ahead of myself though. Just seems like a good starting point
for how to consolidate the various users.
And I had suggested a similar thing for TDX and tglx was 100% right in
shooting it down because we do properly designed things - not, get stuff
in so that vendor is happy and then, once the vendor programmers have
disappeared to do their next enablement task, the maintainers get to mop
up and maintain it forever.
Because this mopping up doesn't scale - trust me.
Got it, and my apologies if I've given you that impression as it's
certainly not my intent. (though I'm sure you've heard that before.)
quoted
[The #VC-based SEV MSR read is not necessary for anything in sme_enable(),
it's simply a way to determine whether the guest is an SNP guest, without
any reliance on CPUID, which seemed useful in the context of doing some
additional sanity checks against the SNP CPUID table and determining that
it's appropriate to use it early on (rather than just trust that this is an
SNP guest by virtue of the CC blob being present, and then failing later
once sme_enable() checks for the SNP feature bits through the normal
mechanism, as was done in v5).]
So you need to make up your mind here design-wise, what you wanna do.
The proper thing to do would be, to detect *everything*, detect whether
this is an SNP guest, yadda yadda, everything your code is going to need
later on, and then be done with it.
Then you continue with the boot and now your other code queries
everything that has been detected up til now and uses it.
Agreed, if we need to check SEV MSR early for the purposes of SNP it makes
sense to move the overall SEV feature detection code earlier as well. I
should have looked into that aspect more closely before introducing the
changes.
From: Michael Roth <hidden> Date: 2021-10-21 02:08:59
On Wed, Oct 20, 2021 at 08:08:39PM +0200, Borislav Petkov wrote:
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
quoted
quoted
1. Code checks SME/SEV support leaf. HV lies and says there's none. So
guest doesn't boot encrypted. Oh well, not a big deal, the cloud vendor
won't be able to give confidentiality to its users => users go away or
do unencrypted like now.
Problem is solved by political and economical pressure.
2. Check SEV and SME bit. HV lies here. Oh well, same as the above.
I'd be worried about the possibility that, through some additional exploits
or failures in the attestation flow,
Well, that puts forward an important question: how do you verify
*reliably* that this is an SNP guest?
- attestation?
- CPUID?
- anything else?
I don't see this written down anywhere. Because this assumption will
guide the design in the kernel.
According to the APM at least, (Rev 3.37, 15.34.10, "SEV_STATUS MSR"), the
SEV MSR is the appropriate source for guests to use. This is what is used
in the EFI code as well. So that seems to be the right way to make the
initial determination.
There's a dependency there on the SEV CPUID bit however, since setting the
bit to 0 would generally result in a guest skipping the SEV MSR read and
assuming 0. So for SNP it would be more reliable to make use of the CPUID
table at that point, since it's less-susceptible to manipulation, or do the
#VC-based SEV MSR read (or both).
quoted
a guest owner was tricked into booting unencrypted on a compromised
host and exposing their secrets. Their attestation process might even
do some additional CPUID sanity checks, which would at the point
be via the SNP CPUID table and look legitimate, unaware that the
kernel didn't actually use the SNP CPUID table until after 0x8000001F
was parsed (if we were to only initialize it after/as-part-of
sme_enable()).
So what happens with that guest owner later?
How is she to notice that she booted unencrypted?
Fully-unencrypted should result in a crash due to the reasons below.
But there may exist some carefully crafted outside influences that could
goad the guest into, perhaps, not marking certain pages as private. The
best that can be done to prevent that is to audit/harden all the code in the
boot stack so that it is less susceptible to that kind of outside
manipulation (via mechanisms like SEV-ES, SNP page validation, SNP CPUID
table, SNP restricted injection, etc.)
Then of course that boot stack needs to be part of the attestation process
to provide any meaningful assurances about the resulting guest state.
Outside of the boot stack the guest owner might take some extra precautions.
Perhaps custom some kernel driver to verify encryption/validated status of
guest pages, some checks against the CPUID table to verify it contains sane
values, but not really worth speculating on that aspect as it will be
ultimately dependent on how the cloud vendor decides to handle things after
boot.
quoted
Fortunately in this scenario I think the guest kernel actually would fail to
boot due to the SNP hardware unconditionally treating code/page tables as
encrypted pages. I tested some of these scenarios just to check, but not
all, and I still don't feel confident enough about it to say that there's
not some way to exploit this by someone who is more clever/persistant than
me.
All this design needs to be preceded with: "We protect against cases A,
B and C and not against D, E, etc."
So that it is clear to all parties involved what we're working with and
what we're protecting against and what we're *not* protecting against.
That would indeed be useful. Perhaps as a nice big comment in sme_enable()
and/or the proposed sev_init() so that those invariants can be maintained,
or updated in sync with future changes. I'll look into that for the next
spin and check with Brijesh on the details.
On Wed, Oct 20, 2021 at 07:35:35PM -0500, Michael Roth wrote:
Fortunately, all the code makes use of sev_status to get at the SEV MSR
bits, so breaking the appropriate bits out of sme_enable() into an earlier
sev_init() routine that's the exclusive writer of sev_status sounds like a
promising approach.
Ack.
It makes sense to do it immediately after the first #VC handler is set
up, so CPUID is available, and since that's where SNP CPUID table
initialization would need to happen if it's to be made available in
#VC handler.
Right, and you can do all your init/CPUID prep there.
It may even be similar enough between boot/compressed and run-time kernel
that it could be a shared routine in sev-shared.c.
Uuh, bonus points! :-)
But then again it also sounds like the appropriate place to move the
snp_cpuid_init*() calls, and locating the cc_blob, and since there's
differences there it might make sense to keep the boot/compressed and
kernel proper sev_init() routines separate to avoid #ifdeffery).
Not to get ahead of myself though. Just seems like a good starting point
for how to consolidate the various users.
I like how you're thinking. :)
Got it, and my apologies if I've given you that impression as it's
certainly not my intent. (though I'm sure you've heard that before.)
Nothing to apologize - all good.
Agreed, if we need to check SEV MSR early for the purposes of SNP it makes
sense to move the overall SEV feature detection code earlier as well. I
should have looked into that aspect more closely before introducing the
changes.
On Wed, Oct 20, 2021 at 09:05:42PM -0500, Michael Roth wrote:
According to the APM at least, (Rev 3.37, 15.34.10, "SEV_STATUS MSR"), the
SEV MSR is the appropriate source for guests to use. This is what is used
in the EFI code as well. So that seems to be the right way to make the
initial determination.
Yap.
There's a dependency there on the SEV CPUID bit however, since setting the
bit to 0 would generally result in a guest skipping the SEV MSR read and
assuming 0. So for SNP it would be more reliable to make use of the CPUID
table at that point, since it's less-susceptible to manipulation, or do the
#VC-based SEV MSR read (or both).
So the CPUID page is supplied by the firmware, right?
Then, you parse it and see that the CPUID bit is 1, then you start using
the SEV_STATUS MSR and all good.
If there *is* a CPUID page but that bit is 0, then you can safely assume
that something is playing tricks on ya so you simply refuse booting.
Fully-unencrypted should result in a crash due to the reasons below.
Crash is a good thing in confidential computing. :)
But there may exist some carefully crafted outside influences that could
goad the guest into, perhaps, not marking certain pages as private. The
best that can be done to prevent that is to audit/harden all the code in the
boot stack so that it is less susceptible to that kind of outside
manipulation (via mechanisms like SEV-ES, SNP page validation, SNP CPUID
table, SNP restricted injection, etc.)
So to me I wonder why would one use anything *else* but an SNP guest. We
all know that those previous technologies were just the stepping stones
towards SNP.
Then of course that boot stack needs to be part of the attestation process
to provide any meaningful assurances about the resulting guest state.
Outside of the boot stack the guest owner might take some extra precautions.
Perhaps custom some kernel driver to verify encryption/validated status of
guest pages, some checks against the CPUID table to verify it contains sane
values, but not really worth speculating on that aspect as it will be
ultimately dependent on how the cloud vendor decides to handle things after
boot.
Well, I've always advocated having a best-practices writeup somewhere
goes a long way to explain this technology to people and how to get
their feet wet. And there you can give hints how such verification could
look like in detail...
That would indeed be useful. Perhaps as a nice big comment in sme_enable()
and/or the proposed sev_init() so that those invariants can be maintained,
or updated in sync with future changes. I'll look into that for the next
spin and check with Brijesh on the details.
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
At which point we then switch to using the CPUID table? But at that
point all the previous CPUID checks, both SEV-related/non-SEV-related,
are now possibly not consistent with what's in the CPUID table. Do we
then revalidate?
Well, that's a tough question. That's basically the same question as,
does Linux support heterogeneous cores and can it handle hardware
features which get enabled after boot. The perfect example is, late
microcode loading which changes CPUID bits and adds new functionality.
And the answer to that is, well, hard. You need to decide this on a
case-by-case basis.
But isn't it that the SNP CPUID page will be parsed early enough anyway
so that kernel proper will see only SNP CPUID info and init properly
using that?
Even a non-malicious hypervisor might provide inconsistent values
between the two sources due to bugs, or SNP validation suppressing
certain feature bits that hypervisor otherwise exposes, etc.
Now all the code after sme_enable() can potentially take unexpected
execution paths, where post-sme_enable() code makes assumptions about
pre-sme_enable() checks that may no longer hold true.
So as I said above, if you parse SNP CPUID page early enough, you don't
have to worry about feature rediscovery. Early enough means, before
identify_boot_cpu().
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
The CPUID calls in snp_cpuid_init() weren't added specifically to induce
the #VC-based SEV MSR read, they were added only because I thought the
gist of your earlier suggestions were to do more validation against the
CPUID table advertised by EFI
Well, if EFI is providing us with the CPUID table, who verified it? The
attestation process? Is it signed with the AMD platform key?
Because if we can verify the firmware is ok, then we can trust the CPUID
page, right?
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Dr. David Alan Gilbert <hidden> Date: 2021-10-21 15:56:21
* Borislav Petkov (bp@alien8.de) wrote:
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
quoted
At which point we then switch to using the CPUID table? But at that
point all the previous CPUID checks, both SEV-related/non-SEV-related,
are now possibly not consistent with what's in the CPUID table. Do we
then revalidate?
Well, that's a tough question. That's basically the same question as,
does Linux support heterogeneous cores and can it handle hardware
features which get enabled after boot. The perfect example is, late
microcode loading which changes CPUID bits and adds new functionality.
And the answer to that is, well, hard. You need to decide this on a
case-by-case basis.
I can imagine a malicious hypervisor trying to return different cpuid
answers to different threads or even the same thread at different times.
But isn't it that the SNP CPUID page will be parsed early enough anyway
so that kernel proper will see only SNP CPUID info and init properly
using that?
quoted
Even a non-malicious hypervisor might provide inconsistent values
between the two sources due to bugs, or SNP validation suppressing
certain feature bits that hypervisor otherwise exposes, etc.
which is exactly what you say - a non-malicious HV taking care of its
migration pool. So how do you handle that?
Well, the spec (AMD 56860 SEV spec) says:
'If firmware encounters a CPUID function that is in the standard or extended ranges, then the
firmware performs a check to ensure that the provided output would not lead to an insecure guest
state'
so I take that 'firmware' to be the PSP; that wording doesn't say that
it checks that the CPUID is identical, just that it 'would not lead to
an insecure guest' - so a hypervisor could hide any 'no longer affected
by' flag for all the CPUs in it's migration pool and the firmware
shouldn't complain; so it should be OK to pessimise.
Dave
quoted
Now all the code after sme_enable() can potentially take unexpected
execution paths, where post-sme_enable() code makes assumptions about
pre-sme_enable() checks that may no longer hold true.
So as I said above, if you parse SNP CPUID page early enough, you don't
have to worry about feature rediscovery. Early enough means, before
identify_boot_cpu().
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
--
Dr. David Alan Gilbert / dgilbert@redhat.com / Manchester, UK
On Thu, Oct 21, 2021 at 04:56:09PM +0100, Dr. David Alan Gilbert wrote:
I can imagine a malicious hypervisor trying to return different cpuid
answers to different threads or even the same thread at different times.
Haha, I guess that will fail not because of SEV* but because of the
kernel not really being able to handle heterogeneous CPUIDs.
Well, the spec (AMD 56860 SEV spec) says:
'If firmware encounters a CPUID function that is in the standard or extended ranges, then the
firmware performs a check to ensure that the provided output would not lead to an insecure guest
state'
so I take that 'firmware' to be the PSP; that wording doesn't say that
it checks that the CPUID is identical, just that it 'would not lead to
an insecure guest' - so a hypervisor could hide any 'no longer affected
by' flag for all the CPUs in it's migration pool and the firmware
shouldn't complain; so it should be OK to pessimise.
AFAIU this, I think this would depend on "[t]he policy used by the
firmware to assess CPUID function output can be found in [PPR]."
So if the HV sets the "no longer affected by" flag but the firmware
deems this set flag as insecure, I'm assuming the firmare will clear
it when it returns the CPUID leafs. I guess I need to go find that
policy...
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Dr. David Alan Gilbert <hidden> Date: 2021-10-21 17:13:03
* Borislav Petkov (bp@alien8.de) wrote:
On Thu, Oct 21, 2021 at 04:56:09PM +0100, Dr. David Alan Gilbert wrote:
quoted
I can imagine a malicious hypervisor trying to return different cpuid
answers to different threads or even the same thread at different times.
Haha, I guess that will fail not because of SEV* but because of the
kernel not really being able to handle heterogeneous CPUIDs.
My worry is if it fails cleanly or fails in a way an evil hypervisor can
exploit.
quoted
Well, the spec (AMD 56860 SEV spec) says:
'If firmware encounters a CPUID function that is in the standard or extended ranges, then the
firmware performs a check to ensure that the provided output would not lead to an insecure guest
state'
so I take that 'firmware' to be the PSP; that wording doesn't say that
it checks that the CPUID is identical, just that it 'would not lead to
an insecure guest' - so a hypervisor could hide any 'no longer affected
by' flag for all the CPUs in it's migration pool and the firmware
shouldn't complain; so it should be OK to pessimise.
AFAIU this, I think this would depend on "[t]he policy used by the
firmware to assess CPUID function output can be found in [PPR]."
So if the HV sets the "no longer affected by" flag but the firmware
deems this set flag as insecure, I'm assuming the firmare will clear
it when it returns the CPUID leafs. I guess I need to go find that
policy...
<digs - ppr_B1_pub_1 55898 rev 0.50 >
OK, so that bit is 8...21 Eax ext2eax bit 6 page 1-109
then 2.1.5.3 CPUID policy enforcement shows 8...21 EAX as
'bitmask'
'bits set in the GuestVal must also be set in HostVal.
This is often applied to feature fields where each bit indicates
support for a feature'
So that's right isn't it?
Dave
On Thu, Oct 21, 2021 at 06:12:53PM +0100, Dr. David Alan Gilbert wrote:
OK, so that bit is 8...21 Eax ext2eax bit 6 page 1-109
then 2.1.5.3 CPUID policy enforcement shows 8...21 EAX as
'bitmask'
'bits set in the GuestVal must also be set in HostVal.
This is often applied to feature fields where each bit indicates
support for a feature'
So that's right isn't it?
Yap, AFAIRC, it would fail the check if:
(GuestVal & HostVal) != GuestVal
and GuestVal is "the CPUID result value created by the hypervisor that
it wants to give to the guest". Let's say it clears bit 6 there.
Then HostVal comes in which is "the actual CPUID result value specified
in this PPR" and there the guest catches the HV lying its *ss off.
:-)
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Dr. David Alan Gilbert <hidden> Date: 2021-10-21 17:48:07
* Borislav Petkov (bp@alien8.de) wrote:
On Thu, Oct 21, 2021 at 06:12:53PM +0100, Dr. David Alan Gilbert wrote:
quoted
OK, so that bit is 8...21 Eax ext2eax bit 6 page 1-109
then 2.1.5.3 CPUID policy enforcement shows 8...21 EAX as
'bitmask'
'bits set in the GuestVal must also be set in HostVal.
This is often applied to feature fields where each bit indicates
support for a feature'
So that's right isn't it?
Yap, AFAIRC, it would fail the check if:
(GuestVal & HostVal) != GuestVal
and GuestVal is "the CPUID result value created by the hypervisor that
it wants to give to the guest". Let's say it clears bit 6 there.
^^^^^^^
Then HostVal comes in which is "the actual CPUID result value specified
in this PPR" and there the guest catches the HV lying its *ss off.
:-)
Hang on, I think it's perfectly fine for it to clear that bit - it just
gets caught if it *sets* it (i.e. claims to be a chip unaffected by the
bug).
i.e. if guestval=0 then (GustVal & whatever) == GuestVal
fine
?
Dave
On Thu, Oct 21, 2021 at 06:47:50PM +0100, Dr. David Alan Gilbert wrote:
Hang on, I think it's perfectly fine for it to clear that bit - it just
gets caught if it *sets* it (i.e. claims to be a chip unaffected by the
bug).
i.e. if guestval=0 then (GustVal & whatever) == GuestVal
fine
?
Bah, ofc. The name of the bit is NullSelectorClearsBase - so when it is
clear, we will note we're affected, as that patch does:
+ /*
+ * CPUID bit above wasn't set. If this kernel is still running
+ * as a HV guest, then the HV has decided not to advertize
+ * that CPUID bit for whatever reason. For example, one
+ * member of the migration pool might be vulnerable. Which
+ * means, the bug is present: set the BUG flag and return.
+ */
+ if (cpu_has(c, X86_FEATURE_HYPERVISOR)) {
+ set_cpu_bug(c, X86_BUG_NULL_SEG);
+ return;
+ }
I have managed to flip the meaning in my mind.
Ok, that makes more sense.
Thx.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Michael Roth <hidden> Date: 2021-10-21 21:35:02
On Thu, Oct 21, 2021 at 04:51:06PM +0200, Borislav Petkov wrote:
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
quoted
The CPUID calls in snp_cpuid_init() weren't added specifically to induce
the #VC-based SEV MSR read, they were added only because I thought the
gist of your earlier suggestions were to do more validation against the
CPUID table advertised by EFI
Well, if EFI is providing us with the CPUID table, who verified it? The
attestation process? Is it signed with the AMD platform key?
For CPUID table pages, the only thing that's assured/attested to by firmware
is that:
1) it is present at the expected guest physical address (that address
is generally baked into the EFI firmware, which *is* attested to)
2) its contents have been validated by the PSP against the current host
CPUID capabilities as defined by the AMD PPR (Publication #55898),
Section 2.1.5.3, "CPUID Policy Enforcement"
3) it is encrypted with the guest key
4) it is in a validated state at launch
The actual contents of the CPUID table are *not* attested to, so in theory
it can still be manipulated by a malicious hypervisor as part of the initial
SNP_LAUNCH_UPDATE firmware commands that provides the initial plain-text
encoding of the CPUID table that is provided to the PSP via
SNP_LAUNCH_UPDATE. It's also not signed in any way (apparently there were
some security reasons for that decision, though I don't know the full
details).
[A guest owner can still validate their CPUID values against known good
ones as part of their attestation flow, but that is not part of the
attestation report as reported by SNP firmware. (So long as there is some
care taken to ensure the source of the CPUID values visible to
userspace/guest attestion process are the same as what was used by the boot
stack: i.e. EFI/bootloader/kernel all use the CPUID page at that same
initial address, or in cases where a copy is used, that copy is placed in
encrypted/private/validated guest memory so it can't be tampered with during
boot.]
So, while it's more difficult to do, and the scope of influence is reduced,
there are still some games that can be played to mess with boot via
manipulation of the initial CPUID table values, so long as they are within
the constraints set by the CPUID enforcement policy defined in the PPR.
Unfortunately, the presence of the SEV/SEV-ES/SEV-SNP bits in 0x8000001F,
EAX, are not enforced by PSP. The only thing enforced there is that the
hypervisor cannot advertise bits that aren't supported by hardware. So
no matter how much the boot stack is trusted, the CPUID table does not
inherit that trust, and even values that we *know* should be true should be
verified rather than assumed.
But I think there are a couple approaches for verifying this is an SNP
guest that are robust against this sort of scenario. You've touched on
some of them in your other replies, so I'll respond there.
From: Michael Roth <hidden> Date: 2021-10-21 21:35:20
On Thu, Oct 21, 2021 at 04:48:16PM +0200, Borislav Petkov wrote:
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
quoted
At which point we then switch to using the CPUID table? But at that
point all the previous CPUID checks, both SEV-related/non-SEV-related,
are now possibly not consistent with what's in the CPUID table. Do we
then revalidate?
Well, that's a tough question. That's basically the same question as,
does Linux support heterogeneous cores and can it handle hardware
features which get enabled after boot. The perfect example is, late
microcode loading which changes CPUID bits and adds new functionality.
And the answer to that is, well, hard. You need to decide this on a
case-by-case basis.
But isn't it that the SNP CPUID page will be parsed early enough anyway
so that kernel proper will see only SNP CPUID info and init properly
using that?
At the time I wrote that I thought you were suggesting moving the SNP CPUID
table initialization to where sme_enable() is in current upstream, so it
seemed worth mentioning, but since the idea was actually to move all the
sev_status initialization in sme_enable() earlier in the code to where
SNP CPUID table init needs to happen (before first cpuid calls are made), I
this scenario is avoided.
quoted
Even a non-malicious hypervisor might provide inconsistent values
between the two sources due to bugs, or SNP validation suppressing
certain feature bits that hypervisor otherwise exposes, etc.
I concur with David's assessment on that solution being compatible with
CPUID enforcement policy. But it's certainly something to consider more
generally.
Fortunately I think I misspoke earlier, I thought there was a case or 2
where bits were suppressed, rather than causing a validation failure,
but looking back through the PPR I doesn't seem like that's actually the
case. Which is good, since that would indeed be painful to deal with in
the context of migration.
From: Michael Roth <hidden> Date: 2021-10-21 23:00:44
On Thu, Oct 21, 2021 at 04:39:31PM +0200, Borislav Petkov wrote:
On Wed, Oct 20, 2021 at 09:05:42PM -0500, Michael Roth wrote:
quoted
According to the APM at least, (Rev 3.37, 15.34.10, "SEV_STATUS MSR"), the
SEV MSR is the appropriate source for guests to use. This is what is used
in the EFI code as well. So that seems to be the right way to make the
initial determination.
Yap.
quoted
There's a dependency there on the SEV CPUID bit however, since setting the
bit to 0 would generally result in a guest skipping the SEV MSR read and
assuming 0. So for SNP it would be more reliable to make use of the CPUID
table at that point, since it's less-susceptible to manipulation, or do the
#VC-based SEV MSR read (or both).
So the CPUID page is supplied by the firmware, right?
Yes.
Then, you parse it and see that the CPUID bit is 1, then you start using
the SEV_STATUS MSR and all good.
If there *is* a CPUID page but that bit is 0, then you can safely assume
that something is playing tricks on ya so you simply refuse booting.
I think that's a good way to deal with this.
I was going to suggest we could assume the presence of SEV status MSR by
virtue of EFI/bootloader/etc having provided a cc_blob, and just read it
right away to confirm this is SNP. But with your approach we could basically
just set up the table early, based on the presence of the cc_blob, and do all
the checks in sme_enable() in the same order as with SEV/SEV-ES, then just
have additional sanity checks against the CPUID/MSR response values to
ensure the SNP bits are present for the cases where a cpuid table / cc_blob
are provided.
I'll work on implementing things in this way and see how it goes.
quoted
Fully-unencrypted should result in a crash due to the reasons below.
Crash is a good thing in confidential computing. :)
quoted
But there may exist some carefully crafted outside influences that could
goad the guest into, perhaps, not marking certain pages as private. The
best that can be done to prevent that is to audit/harden all the code in the
boot stack so that it is less susceptible to that kind of outside
manipulation (via mechanisms like SEV-ES, SNP page validation, SNP CPUID
table, SNP restricted injection, etc.)
So to me I wonder why would one use anything *else* but an SNP guest. We
all know that those previous technologies were just the stepping stones
towards SNP.
Yah, I think ultimately that's where things are headed.
quoted
Then of course that boot stack needs to be part of the attestation process
to provide any meaningful assurances about the resulting guest state.
Outside of the boot stack the guest owner might take some extra precautions.
Perhaps custom some kernel driver to verify encryption/validated status of
guest pages, some checks against the CPUID table to verify it contains sane
values, but not really worth speculating on that aspect as it will be
ultimately dependent on how the cloud vendor decides to handle things after
boot.
Well, I've always advocated having a best-practices writeup somewhere
goes a long way to explain this technology to people and how to get
their feet wet. And there you can give hints how such verification could
look like in detail...
Our security team is working on some initial reference designs / tooling
for attestation. It'll eventually make it's way to here:
https://github.com/AMDESE/sev-guest
but it's still mostly an internal effort so nothing there ATM. But hopefully
that will fill in some of these gaps. But I agree some an accompanying best
practices document to highlight some of these considerations is also
something that should be considered, I'll need to check to see if there's
anything like that in the works already.
quoted
That would indeed be useful. Perhaps as a nice big comment in sme_enable()
and/or the proposed sev_init() so that those invariants can be maintained,
or updated in sync with future changes. I'll look into that for the next
spin and check with Brijesh on the details.
There is Documentation/x86/amd-memory-encryption.rst, for example.
Makes sense, will work with Brijesh on this.
Thanks!
-Mike
On Thu, Oct 21, 2021 at 03:41:49PM -0500, Michael Roth wrote:
On Thu, Oct 21, 2021 at 04:51:06PM +0200, Borislav Petkov wrote:
quoted
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
quoted
The CPUID calls in snp_cpuid_init() weren't added specifically to induce
the #VC-based SEV MSR read, they were added only because I thought the
gist of your earlier suggestions were to do more validation against the
CPUID table advertised by EFI
Well, if EFI is providing us with the CPUID table, who verified it? The
attestation process? Is it signed with the AMD platform key?
For CPUID table pages, the only thing that's assured/attested to by firmware
is that:
1) it is present at the expected guest physical address (that address
is generally baked into the EFI firmware, which *is* attested to)
2) its contents have been validated by the PSP against the current host
CPUID capabilities as defined by the AMD PPR (Publication #55898),
Section 2.1.5.3, "CPUID Policy Enforcement"
3) it is encrypted with the guest key
4) it is in a validated state at launch
The actual contents of the CPUID table are *not* attested to,
Why?
so in theory it can still be manipulated by a malicious hypervisor as
part of the initial SNP_LAUNCH_UPDATE firmware commands that provides
the initial plain-text encoding of the CPUID table that is provided
to the PSP via SNP_LAUNCH_UPDATE. It's also not signed in any way
(apparently there were some security reasons for that decision, though
I don't know the full details).
So this sounds like an unnecessary complication. I'm sure there are
reasons to do it this way but my simple thinking would simply want the
CPUID page to be read-only and signed so that the guest can trust it
unconditionally.
[A guest owner can still validate their CPUID values against known good
ones as part of their attestation flow, but that is not part of the
attestation report as reported by SNP firmware. (So long as there is some
care taken to ensure the source of the CPUID values visible to
userspace/guest attestion process are the same as what was used by the boot
stack: i.e. EFI/bootloader/kernel all use the CPUID page at that same
initial address, or in cases where a copy is used, that copy is placed in
encrypted/private/validated guest memory so it can't be tampered with during
boot.]
This sounds like the good practices advice to guest owners would be,
"Hey, I just booted your SNP guest but for full trust, you should go and
verify the CPUID page's contents."
"And if I were you, I wouldn't want to run any verification of CPUID
pages' contents on the same guest because it itself hasn't been verified
yet."
It all sounds weird.
So, while it's more difficult to do, and the scope of influence is reduced,
there are still some games that can be played to mess with boot via
manipulation of the initial CPUID table values, so long as they are within
the constraints set by the CPUID enforcement policy defined in the PPR.
Unfortunately, the presence of the SEV/SEV-ES/SEV-SNP bits in 0x8000001F,
EAX, are not enforced by PSP. The only thing enforced there is that the
hypervisor cannot advertise bits that aren't supported by hardware. So
no matter how much the boot stack is trusted, the CPUID table does not
inherit that trust, and even values that we *know* should be true should be
verified rather than assumed.
But I think there are a couple approaches for verifying this is an SNP
guest that are robust against this sort of scenario. You've touched on
some of them in your other replies, so I'll respond there.
Yah, I guess the kernel can do good enough verification and then the
full thing needs to be done by the guest owner and in *some* userspace
- not necessarily on the currently booted, unverified guest - but
somewhere, where you have maximal flexibility.
IMHO.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Michael Roth <hidden> Date: 2021-10-25 16:36:03
On Mon, Oct 25, 2021 at 01:04:10PM +0200, Borislav Petkov wrote:
On Thu, Oct 21, 2021 at 03:41:49PM -0500, Michael Roth wrote:
quoted
On Thu, Oct 21, 2021 at 04:51:06PM +0200, Borislav Petkov wrote:
quoted
On Wed, Oct 20, 2021 at 11:10:23AM -0500, Michael Roth wrote:
quoted
The CPUID calls in snp_cpuid_init() weren't added specifically to induce
the #VC-based SEV MSR read, they were added only because I thought the
gist of your earlier suggestions were to do more validation against the
CPUID table advertised by EFI
Well, if EFI is providing us with the CPUID table, who verified it? The
attestation process? Is it signed with the AMD platform key?
For CPUID table pages, the only thing that's assured/attested to by firmware
is that:
1) it is present at the expected guest physical address (that address
is generally baked into the EFI firmware, which *is* attested to)
2) its contents have been validated by the PSP against the current host
CPUID capabilities as defined by the AMD PPR (Publication #55898),
Section 2.1.5.3, "CPUID Policy Enforcement"
3) it is encrypted with the guest key
4) it is in a validated state at launch
The actual contents of the CPUID table are *not* attested to,
Why?
As counter-intuitive as it sounds, it actually doesn't buy us if the CPUID
table is part of the PSP attestation report, since:
- the boot stack is attested to, and if the boot stack isn't careful to
use the CPUID table at all times, then attesting CPUID table after
boot doesn't provide any assurance that the boot wasn't manipulated
by CPUID
- given the boot stack must take these precautions, guest-specific
attestation code is just as capable of attesting the CPUID table
contents/values, since it has the same view of the CPUID values that
were used during boot.
So leaving it to the guest owner to attest it provides some flexibility
to guest owners to implement it as they see fit, whereas making it part
of the attestation report means that the guest needs the exact contents
of the CPUID page for a particular guest configuration so it can be
incorporated into the measurement they are expecting, which would likely
require some tooling provided by the cloud vendor, since every different
guest configuration, or even changes like the ordering in which entries
are placed in the table, would affect measurement, so it's not something
that could be easily surmised separately with minimal involvement from a
cloud vendor.
And even if the cloud vendor provided a simple way to export the table
contents for measurement, can you really trust it? If you have to audit
individual entries to be sure there's nothing fishy, why not just
incorporate those checks into the guest owner's attestation flow and
leave the vendor out of it completely?
So not including it in the measurement meshes well with the overall
SEV-SNP approach of reducing the cloud vendor's involvement in the
overall attestation process.
quoted
so in theory it can still be manipulated by a malicious hypervisor as
part of the initial SNP_LAUNCH_UPDATE firmware commands that provides
the initial plain-text encoding of the CPUID table that is provided
to the PSP via SNP_LAUNCH_UPDATE. It's also not signed in any way
(apparently there were some security reasons for that decision, though
I don't know the full details).
So this sounds like an unnecessary complication. I'm sure there are
reasons to do it this way but my simple thinking would simply want the
CPUID page to be read-only and signed so that the guest can trust it
unconditionally.
The thing here is that it's not just a specific CPUID page that's valid
for all guests for a particular host. Booting a guest with additional
vCPUs changes the contents, different CPU models/flags changes the
contents, etc. So it needs to be generated for each specific guest
configuration, and can't just be a read-only page.
Some sort of signature that indicates the PSP's stamp of approval on a
particular CPUID page would be nice, but we do sort of have this in the
sense that CPUID page 'address' is part of measurement, and can only
contain values that were blessed by the PSP. The problem then becomes
ensuring that only that address it used for CPUID lookups, and that it's
contents weren't manipulated in a way where it's 'valid' as far as the
PSP is concerned, but still not the 'expected' values for a particular
guest (which is where the attestation mentioned above would come into
play).
quoted
[A guest owner can still validate their CPUID values against known good
ones as part of their attestation flow, but that is not part of the
attestation report as reported by SNP firmware. (So long as there is some
care taken to ensure the source of the CPUID values visible to
userspace/guest attestion process are the same as what was used by the boot
stack: i.e. EFI/bootloader/kernel all use the CPUID page at that same
initial address, or in cases where a copy is used, that copy is placed in
encrypted/private/validated guest memory so it can't be tampered with during
boot.]
This sounds like the good practices advice to guest owners would be,
"Hey, I just booted your SNP guest but for full trust, you should go and
verify the CPUID page's contents."
"And if I were you, I wouldn't want to run any verification of CPUID
pages' contents on the same guest because it itself hasn't been verified
yet."
It all sounds weird.
Yes, understandably so. But the only way to avoid that sort of weirdness
in general is for *all* guest state to be measured, all pages, all
registers, etc. Baking that directly into the SEV-SNP attestation report
would be a non-starter for most since computing the measurement for all
that state independently would require lots of additional inputs from
cloud vendor (who we don't necessarily trust in the first place), and
constant updates of measurement values since they would change with
every guest configuration change, every different starting TSC offset,
different, maybe the order in which vCPUs were onlined, stuff that the
kernel prints to log buffers, etc.
But, if a guest owner wants to attempt clever ways to account for some/all
of that in their attestation flow, they are welcome to try. That's sort of
the idea behind SNP attestation vs. SEV. Things like page
validation/encryption, cpuid enforcement, etc., reduce some of the
variables/possibilities guest owners need to account for during attestation
to make the process more secure/tenable, but they don't rule out all
possibilities, just as Trusted Boot doesn't necessarily mean you can fully
trust your OS state immediately afer boot; there are still outside
influences at play, and the boot stack should guard against them
wherever possible.
quoted
So, while it's more difficult to do, and the scope of influence is reduced,
there are still some games that can be played to mess with boot via
manipulation of the initial CPUID table values, so long as they are within
the constraints set by the CPUID enforcement policy defined in the PPR.
Unfortunately, the presence of the SEV/SEV-ES/SEV-SNP bits in 0x8000001F,
EAX, are not enforced by PSP. The only thing enforced there is that the
hypervisor cannot advertise bits that aren't supported by hardware. So
no matter how much the boot stack is trusted, the CPUID table does not
inherit that trust, and even values that we *know* should be true should be
verified rather than assumed.
But I think there are a couple approaches for verifying this is an SNP
guest that are robust against this sort of scenario. You've touched on
some of them in your other replies, so I'll respond there.
Yah, I guess the kernel can do good enough verification and then the
full thing needs to be done by the guest owner and in *some* userspace
- not necessarily on the currently booted, unverified guest - but
somewhere, where you have maximal flexibility.
Exactly, moving attestation into the guest allows for more of the these
unexpected states to be accounted for at whatever level of paranoia a guest
owner sees fit, while still allowing firmware to provide some basic
assurances via attestation report and various features to reduce common
attack vectors during/after boot.
On Mon, Oct 25, 2021 at 11:35:18AM -0500, Michael Roth wrote:
As counter-intuitive as it sounds, it actually doesn't buy us if the CPUID
table is part of the PSP attestation report, since:
Thanks for taking the time to explain in detail - I think I know now
what's going on, and David explained some additional stuff to me
yesterday.
So, to cut to the chase:
- yeah, ok, I guess guest owner attestation is what should happen.
- as to the boot detection, I think you should do in sme_enable(), in
pseudo:
bool snp_guest_detected;
if (CPUID page address) {
read SEV_STATUS;
snp_guest_detected = SEV_STATUS & MSR_AMD64_SEV_SNP_ENABLED;
}
/* old SME/SEV detection path */
read 0x8000_001F_EAX and look at bits SME and SEV, yadda yadda.
if (snp_guest_detected && (!SME || !SEV))
/*
* HV is lying to me, do something there, dunno what. I guess we can
* continue booting unencrypted so that the guest owner knows that
* detection has failed and maybe the HV didn't want us to force SNP.
* This way, attestation will fail and the user will know why.
* Or something like that.
*/
/* normal feature detection continues. */
How does that sound?
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
From: Michael Roth <hidden> Date: 2021-10-27 15:14:27
On Wed, Oct 27, 2021 at 01:17:11PM +0200, Borislav Petkov wrote:
On Mon, Oct 25, 2021 at 11:35:18AM -0500, Michael Roth wrote:
quoted
As counter-intuitive as it sounds, it actually doesn't buy us if the CPUID
table is part of the PSP attestation report, since:
Thanks for taking the time to explain in detail - I think I know now
what's going on, and David explained some additional stuff to me
yesterday.
So, to cut to the chase:
- yeah, ok, I guess guest owner attestation is what should happen.
- as to the boot detection, I think you should do in sme_enable(), in
pseudo:
bool snp_guest_detected;
if (CPUID page address) {
read SEV_STATUS;
snp_guest_detected = SEV_STATUS & MSR_AMD64_SEV_SNP_ENABLED;
}
/* old SME/SEV detection path */
read 0x8000_001F_EAX and look at bits SME and SEV, yadda yadda.
if (snp_guest_detected && (!SME || !SEV))
/*
* HV is lying to me, do something there, dunno what. I guess we can
* continue booting unencrypted so that the guest owner knows that
* detection has failed and maybe the HV didn't want us to force SNP.
* This way, attestation will fail and the user will know why.
* Or something like that.
*/
/* normal feature detection continues. */
How does that sound?
That seems promising. I've been testing a similar approach in conjunction with
moving sme_enable() to after the initial #VC handler is set up and things seem
to work out pretty nicely.
boot/compressed is a little less straightforward since the sme_enable()
equivalent is set_sev_encryption_mask() which sets sev_status and is written
in assembly, whereas the SNP-specific bits we're adding relies on C code
that handles stuff like scanning EFI config table are in C, so probably
worthwhile to see if everything can be redone in C. But then there's
get_sev_encryption_bit(), which needs to be in assembly since it needs
to be called from 32-bit entry path as well, but that doesn't actually
rely on anything set by set_sev_encryption_mask(), so it seems like it
should be okay to split set_sev_encryption_mask() out into a separate C
routine.
Will work on implementing/testing that approach, but if you or Joerg are
aware of any showstoppers there just let me know.
Thanks!
-Mike