From: Kai Huang <hidden> Date: 2022-11-21 00:27:36
Intel Trusted Domain Extensions (TDX) protects guest VMs from malicious
host and certain physical attacks. TDX specs are available in [1].
This series is the initial support to enable TDX with minimal code to
allow KVM to create and run TDX guests. KVM support for TDX is being
developed separately[2]. A new "userspace inaccessible memfd" approach
to support TDX private memory is also being developed[3]. The KVM will
only support the new "userspace inaccessible memfd" as TDX guest memory.
This series doesn't aim to support all functionalities (i.e. exposing TDX
module via /sysfs), and doesn't aim to resolve all things perfectly.
Especially, the implementation to how to choose "TDX-usable" memory and
memory hotplug handling is simple, that this series just makes sure all
pages in the page allocator are TDX memory.
A better solution, suggested by Kirill, is similar to the per-node memory
encryption flag in this series [4]. Similarly, a per-node TDX flag can
be added so both "TDX-capable" and "non-TDX-capable" nodes can co-exist.
With exposing the TDX flag to userspace via /sysfs, the userspace can
then use NUMA APIs to bind TDX guests to those "TDX-capable" nodes.
For more information please refer to "Kernel policy on TDX memory" and
"Memory hotplug" sections below. Huang, Ying is working on this
"per-node TDX flag" support and will post another series independently.
(For memory hotplug, sorry for broadcasting widely but I cc'ed the
linux-mm@kvack.org following Kirill's suggestion so MM experts can also
help to provide comments.)
Also, other optimizations will be posted as follow-up once this initial
TDX support is upstreamed.
Hi Dave, Dan, Kirill, Ying (and Intel reviewers),
Please kindly help to review, and I would appreciate reviewed-by or
acked-by tags if the patches look good to you.
This series has been reviewed by Isaku who is developing KVM TDX patches.
Kirill also has reviewed couple of patches as well.
Also, I highly appreciate if anyone else can help to review this series.
----- Changelog history: ------
- v6 -> v7:
- Added memory hotplug support.
- Changed how to choose the list of "TDX-usable" memory regions from at
kernel boot time to TDX module initialization time.
- Addressed comments received in previous versions. (Andi/Dave).
- Improved the commit message and the comments of kexec() support patch,
and the patch handles returnning PAMTs back to the kernel when TDX
module initialization fails. Please also see "kexec()" section below.
- Changed the documentation patch accordingly.
- For all others please see individual patch changelog history.
- v5 -> v6:
- Removed ACPI CPU/memory hotplug patches. (Intel internal discussion)
- Removed patch to disable driver-managed memory hotplug (Intel
internal discussion).
- Added one patch to introduce enum type for TDX supported page size
level to replace the hard-coded values in TDX guest code (Dave).
- Added one patch to make TDX depends on X2APIC being enabled (Dave).
- Added one patch to build all boot-time present memory regions as TDX
memory during kernel boot.
- Added Reviewed-by from others to some patches.
- For all others please see individual patch changelog history.
- v4 -> v5:
This is essentially a resent of v4. Sorry I forgot to consult
get_maintainer.pl when sending out v4, so I forgot to add linux-acpi
and linux-mm mailing list and the relevant people for 4 new patches.
There are also very minor code and commit message update from v4:
- Rebased to latest tip/x86/tdx.
- Fixed a checkpatch issue that I missed in v4.
- Removed an obsoleted comment that I missed in patch 6.
- Very minor update to the commit message of patch 12.
For other changes to individual patches since v3, please refer to the
changelog histroy of individual patches (I just used v3 -> v5 since
there's basically no code change to v4).
- v3 -> v4 (addressed Dave's comments, and other comments from others):
- Simplified SEAMRR and TDX keyID detection.
- Added patches to handle ACPI CPU hotplug.
- Added patches to handle ACPI memory hotplug and driver managed memory
hotplug.
- Removed tdx_detect() but only use single tdx_init().
- Removed detecting TDX module via P-SEAMLDR.
- Changed from using e820 to using memblock to convert system RAM to TDX
memory.
- Excluded legacy PMEM from TDX memory.
- Removed the boot-time command line to disable TDX patch.
- Addressed comments for other individual patches (please see individual
patches).
- Improved the documentation patch based on the new implementation.
- V2 -> v3:
- Addressed comments from Isaku.
- Fixed memory leak and unnecessary function argument in the patch to
configure the key for the global keyid (patch 17).
- Enhanced a little bit to the patch to get TDX module and CMR
information (patch 09).
- Fixed an unintended change in the patch to allocate PAMT (patch 13).
- Addressed comments from Kevin:
- Slightly improvement on commit message to patch 03.
- Removed WARN_ON_ONCE() in the check of cpus_booted_once_mask in
seamrr_enabled() (patch 04).
- Changed documentation patch to add TDX host kernel support materials
to Documentation/x86/tdx.rst together with TDX guest staff, instead
of a standalone file (patch 21)
- Very minor improvement in commit messages.
- RFC (v1) -> v2:
- Rebased to Kirill's latest TDX guest code.
- Fixed two issues that are related to finding all RAM memory regions
based on e820.
- Minor improvement on comments and commit messages.
v6:
https://lore.kernel.org/linux-mm/cover.1666824663.git.kai.huang@intel.com/T/
v5:
https://lore.kernel.org/lkml/cover.1655894131.git.kai.huang@intel.com/T/
v3:
https://lore.kernel.org/lkml/68484e168226037c3a25b6fb983b052b26ab3ec1.camel@intel.com/T/
V2:
https://lore.kernel.org/lkml/cover.1647167475.git.kai.huang@intel.com/T/
RFC (v1):
https://lore.kernel.org/all/e0ff030a49b252d91c789a89c303bb4206f85e3d.1646007267.git.kai.huang@intel.com/T/
== Background ==
TDX introduces a new CPU mode called Secure Arbitration Mode (SEAM)
and a new isolated range pointed by the SEAM Ranger Register (SEAMRR).
A CPU-attested software module called 'the TDX module' runs in the new
isolated region as a trusted hypervisor to create/run protected VMs.
TDX also leverages Intel Multi-Key Total Memory Encryption (MKTME) to
provide crypto-protection to the VMs. TDX reserves part of MKTME KeyIDs
as TDX private KeyIDs, which are only accessible within the SEAM mode.
TDX is different from AMD SEV/SEV-ES/SEV-SNP, which uses a dedicated
secure processor to provide crypto-protection. The firmware runs on the
secure processor acts a similar role as the TDX module.
The host kernel communicates with SEAM software via a new SEAMCALL
instruction. This is conceptually similar to a guest->host hypercall,
except it is made from the host to SEAM software instead.
Before being able to manage TD guests, the TDX module must be loaded
and properly initialized. This series assumes the TDX module is loaded
by BIOS before the kernel boots.
How to initialize the TDX module is described at TDX module 1.0
specification, chapter "13.Intel TDX Module Lifecycle: Enumeration,
Initialization and Shutdown".
== Design Considerations ==
1. Initialize the TDX module at runtime
There are basically two ways the TDX module could be initialized: either
in early boot, or at runtime before the first TDX guest is run. This
series implements the runtime initialization.
This series adds a function tdx_enable() to allow the caller to initialize
TDX at runtime:
if (tdx_enable())
goto no_tdx;
// TDX is ready to create TD guests.
This approach has below pros:
1) Initializing the TDX module requires to reserve ~1/256th system RAM as
metadata. Enabling TDX on demand allows only to consume this memory when
TDX is truly needed (i.e. when KVM wants to create TD guests).
2) SEAMCALL requires CPU being already in VMX operation (VMXON has been
done). So far, KVM is the only user of TDX, and it already handles VMXON.
Letting KVM to initialize TDX avoids handling VMXON in the core kernel.
3) It is more flexible to support "TDX module runtime update" (not in
this series). After updating to the new module at runtime, kernel needs
to go through the initialization process again.
2. CPU hotplug
TDX doesn't support physical (ACPI) CPU hotplug. A non-buggy BIOS should
never support hotpluggable CPU devicee and/or deliver ACPI CPU hotplug
event to the kernel. This series doesn't handle physical (ACPI) CPU
hotplug at all but depends on the BIOS to behave correctly.
Note TDX works with CPU logical online/offline, thus this series still
allows to do logical CPU online/offline.
3. Kernel policy on TDX memory
The TDX module reports a list of "Convertible Memory Region" (CMR) to
indicate which memory regions are TDX-capable. The TDX architecture
allows the VMM to designate specific convertible memory regions as usable
for TDX private memory.
The initial support of TDX guests will only allocate TDX private memory
from the global page allocator. This series chooses to designate _all_
system RAM in the core-mm at the time of initializing TDX module as TDX
memory to guarantee all pages in the page allocator are TDX pages.
4. Memory Hotplug
After the kernel passes all "TDX-usable" memory regions to the TDX
module, the set of "TDX-usable" memory regions are fixed during module's
runtime. No more "TDX-usable" memory can be added to the TDX module
after that.
To achieve above "to guarantee all pages in the page allocator are TDX
pages", this series simply choose to reject any non-TDX-usable memory in
memory hotplug.
This _will_ be enhanced in the future after first submission. The
direction we are heading is to allow adding/onlining non-TDX memory to
separate NUMA nodes so that both "TDX-capable" nodes and "TDX-capable"
nodes can co-exist. The TDX flag can be exposed to userspace via /sysfs
so userspace can bind TDX guests to "TDX-capable" nodes via NUMA ABIs.
Note TDX assumes convertible memory is always physically present during
machine's runtime. A non-buggy BIOS should never support hot-removal of
any convertible memory. This implementation doesn't handle ACPI memory
removal but depends on the BIOS to behave correctly.
5. Kexec()
There are two problems in terms of using kexec() to boot to a new kernel
when the old kernel has enabled TDX: 1) Part of the memory pages are
still TDX private pages (i.e. metadata used by the TDX module, and any
TDX guest memory if kexec() happens when there's any TDX guest alive).
2) There might be dirty cachelines associated with TDX private pages.
Just like SME, TDX hosts require special cache flushing before kexec().
Similar to SME handling, the kernel uses wbinvd() to flush cache in
stop_this_cpu() when TDX is enabled.
This series doesn't convert all TDX private pages back to normal due to
below considerations:
1) The kernel doesn't have existing infrastructure to track which pages
are TDX private pages.
2) The number of TDX private pages can be large, and converting all of
them (cache flush + using MOVDIR64B to clear the page) in kexec() can
be time consuming.
3) The new kernel will almost only use KeyID 0 to access memory. KeyID
0 doesn't support integrity-check, so it's OK.
4) The kernel doesn't (and may never) support MKTME. If any 3rd party
kernel ever supports MKTME, it should do MOVDIR64B to clear the page
with the new MKTME KeyID (just like TDX does) before using it.
Also, if the old kernel ever enables TDX, the new kernel cannot use TDX
again. When the new kernel goes through the TDX module initialization
process it will fail immediately at the first step.
Ideally, it's better to shutdown the TDX module in kexec(), but there's
no guarantee that CPUs are in VMX operation in kexec() so just leave the
module open.
== Reference ==
[1]: TDX specs
https://software.intel.com/content/www/us/en/develop/articles/intel-trust-domain-extensions.html
[2]: KVM TDX basic feature support
https://lore.kernel.org/lkml/CAAhR5DFrwP+5K8MOxz5YK7jYShhaK4A+2h1Pi31U_9+Z+cz-0A@mail.gmail.com/T/
[3]: KVM: mm: fd-based approach for supporting KVM
https://lore.kernel.org/lkml/20220915142913.2213336-1-chao.p.peng@linux.intel.com/T/
[4]: per-node memory encryption flag
https://lore.kernel.org/linux-mm/20221007155323.ue4cdthkilfy4lbd@box.shutemov.name/t/
Kai Huang (20):
x86/tdx: Define TDX supported page sizes as macros
x86/virt/tdx: Detect TDX during kernel boot
x86/virt/tdx: Disable TDX if X2APIC is not enabled
x86/virt/tdx: Add skeleton to initialize TDX on demand
x86/virt/tdx: Implement functions to make SEAMCALL
x86/virt/tdx: Shut down TDX module in case of error
x86/virt/tdx: Do TDX module global initialization
x86/virt/tdx: Do logical-cpu scope TDX module initialization
x86/virt/tdx: Get information about TDX module and TDX-capable memory
x86/virt/tdx: Use all system memory when initializing TDX module as
TDX memory
x86/virt/tdx: Add placeholder to construct TDMRs to cover all TDX
memory regions
x86/virt/tdx: Create TDMRs to cover all TDX memory regions
x86/virt/tdx: Allocate and set up PAMTs for TDMRs
x86/virt/tdx: Set up reserved areas for all TDMRs
x86/virt/tdx: Reserve TDX module global KeyID
x86/virt/tdx: Configure TDX module with TDMRs and global KeyID
x86/virt/tdx: Configure global KeyID on all packages
x86/virt/tdx: Initialize all TDMRs
x86/virt/tdx: Flush cache in kexec() when TDX is enabled
Documentation/x86: Add documentation for TDX host support
Documentation/x86/tdx.rst | 181 +++-
arch/x86/Kconfig | 15 +
arch/x86/Makefile | 2 +
arch/x86/coco/tdx/tdx.c | 6 +-
arch/x86/include/asm/tdx.h | 30 +
arch/x86/kernel/process.c | 8 +-
arch/x86/mm/init_64.c | 10 +
arch/x86/virt/Makefile | 2 +
arch/x86/virt/vmx/Makefile | 2 +
arch/x86/virt/vmx/tdx/Makefile | 2 +
arch/x86/virt/vmx/tdx/seamcall.S | 52 ++
arch/x86/virt/vmx/tdx/tdx.c | 1422 ++++++++++++++++++++++++++++++
arch/x86/virt/vmx/tdx/tdx.h | 118 +++
arch/x86/virt/vmx/tdx/tdxcall.S | 19 +-
14 files changed, 1852 insertions(+), 17 deletions(-)
create mode 100644 arch/x86/virt/Makefile
create mode 100644 arch/x86/virt/vmx/Makefile
create mode 100644 arch/x86/virt/vmx/tdx/Makefile
create mode 100644 arch/x86/virt/vmx/tdx/seamcall.S
create mode 100644 arch/x86/virt/vmx/tdx/tdx.c
create mode 100644 arch/x86/virt/vmx/tdx/tdx.h
base-commit: 00e07cfbdf0b232f7553f0175f8f4e8d792f7e90
--
2.38.1
From: Kai Huang <hidden> Date: 2022-11-21 00:27:39
TDX supports 4K, 2M and 1G page sizes. The corresponding values are
defined by the TDX module spec and used as TDX module ABI. Currently,
they are used in try_accept_one() when the TDX guest tries to accept a
page. However currently try_accept_one() uses hard-coded magic values.
Define TDX supported page sizes as macros and get rid of the hard-coded
values in try_accept_one(). TDX host support will need to use them too.
Reviewed-by: Kirill A. Shutemov <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Removed the helper to convert kernel page level to TDX page level.
- Changed to use macro to define TDX supported page sizes.
---
arch/x86/coco/tdx/tdx.c | 6 +++---
arch/x86/include/asm/tdx.h | 9 +++++++++
2 files changed, 12 insertions(+), 3 deletions(-)
From: Kai Huang <hidden> Date: 2022-11-21 00:27:44
Intel Trust Domain Extensions (TDX) protects guest VMs from malicious
host and certain physical attacks. A CPU-attested software module
called 'the TDX module' runs inside a new isolated memory range as a
trusted hypervisor to manage and run protected VMs.
Pre-TDX Intel hardware has support for a memory encryption architecture
called MKTME. The memory encryption hardware underpinning MKTME is also
used for Intel TDX. TDX ends up "stealing" some of the physical address
space from the MKTME architecture for crypto-protection to VMs. The
BIOS is responsible for partitioning the "KeyID" space between legacy
MKTME and TDX. The KeyIDs reserved for TDX are called 'TDX private
KeyIDs' or 'TDX KeyIDs' for short.
TDX doesn't trust the BIOS. During machine boot, TDX verifies the TDX
private KeyIDs are consistently and correctly programmed by the BIOS
across all CPU packages before it enables TDX on any CPU core. A valid
TDX private KeyID range on BSP indicates TDX has been enabled by the
BIOS, otherwise the BIOS is buggy.
The TDX module is expected to be loaded by the BIOS when it enables TDX,
but the kernel needs to properly initialize it before it can be used to
create and run any TDX guests. The TDX module will be initialized at
runtime by the user (i.e. KVM) on demand.
Add a new early_initcall(tdx_init) to do TDX early boot initialization.
Only detect TDX private KeyIDs for now. Some other early checks will
follow up. Also add a new function to report whether TDX has been
enabled by BIOS (TDX private KeyID range is valid). Kexec() will also
need it to determine whether need to flush dirty cachelines that are
associated with any TDX private KeyIDs before booting to the new kernel.
To start to support TDX, create a new arch/x86/virt/vmx/tdx/tdx.c for
TDX host kernel support. Add a new Kconfig option CONFIG_INTEL_TDX_HOST
to opt-in TDX host kernel support (to distinguish with TDX guest kernel
support). So far only KVM is the only user of TDX. Make the new config
option depend on KVM_INTEL.
Reviewed-by: Kirill A. Shutemov <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- No change.
v5 -> v6:
- Removed SEAMRR detection to make code simpler.
- Removed the 'default N' in the KVM_TDX_HOST Kconfig (Kirill).
- Changed to use 'obj-y' in arch/x86/virt/vmx/tdx/Makefile (Kirill).
---
arch/x86/Kconfig | 12 +++++
arch/x86/Makefile | 2 +
arch/x86/include/asm/tdx.h | 7 +++
arch/x86/virt/Makefile | 2 +
arch/x86/virt/vmx/Makefile | 2 +
arch/x86/virt/vmx/tdx/Makefile | 2 +
arch/x86/virt/vmx/tdx/tdx.c | 95 ++++++++++++++++++++++++++++++++++
arch/x86/virt/vmx/tdx/tdx.h | 15 ++++++
8 files changed, 137 insertions(+)
create mode 100644 arch/x86/virt/Makefile
create mode 100644 arch/x86/virt/vmx/Makefile
create mode 100644 arch/x86/virt/vmx/tdx/Makefile
create mode 100644 arch/x86/virt/vmx/tdx/tdx.c
create mode 100644 arch/x86/virt/vmx/tdx/tdx.h
@@ -246,6 +246,8 @@ archheaders:libs-y+=arch/x86/lib/+core-y+=arch/x86/virt/+# drivers-y are linked after core-ydrivers-$(CONFIG_MATH_EMULATION)+=arch/x86/math-emu/drivers-$(CONFIG_PCI)+=arch/x86/pci/
From: Kai Huang <hidden> Date: 2022-11-21 00:27:48
Before the TDX module can be used to create and run TDX guests, it must
be loaded and properly initialized. The TDX module is expected to be
loaded by the BIOS, and to be initialized by the kernel.
TDX introduces a new CPU mode: Secure Arbitration Mode (SEAM). The host
kernel communicates with the TDX module via a new SEAMCALL instruction.
The TDX module implements a set of SEAMCALL leaf functions to allow the
host kernel to initialize it.
The TDX module can be initialized only once in its lifetime. Instead
of always initializing it at boot time, this implementation chooses an
"on demand" approach to initialize TDX until there is a real need (e.g
when requested by KVM). This approach has below pros:
1) It avoids consuming the memory that must be allocated by kernel and
given to the TDX module as metadata (~1/256th of the TDX-usable memory),
and also saves the CPU cycles of initializing the TDX module (and the
metadata) when TDX is not used at all.
2) It is more flexible to support TDX module runtime updating in the
future (after updating the TDX module, it needs to be initialized
again).
3) It avoids having to do a "temporary" solution to handle VMXON in the
core (non-KVM) kernel for now. This is because SEAMCALL requires CPU
being in VMX operation (VMXON is done), but currently only KVM handles
VMXON. Adding VMXON support to the core kernel isn't trivial. More
importantly, from long-term a reference-based approach is likely needed
in the core kernel as more kernel components are likely needed to
support TDX as well. Allow KVM to initialize the TDX module avoids
having to handle VMXON during kernel boot for now.
Add a placeholder tdx_enable() to detect and initialize the TDX module
on demand, with a state machine protected by mutex to support concurrent
calls from multiple callers.
The TDX module will be initialized in multi-steps defined by the TDX
module:
1) Global initialization;
2) Logical-CPU scope initialization;
3) Enumerate the TDX module capabilities and platform configuration;
4) Configure the TDX module about TDX usable memory ranges and global
KeyID information;
5) Package-scope configuration for the global KeyID;
6) Initialize usable memory ranges based on 4).
The TDX module can also be shut down at any time during its lifetime.
In case of any error during the initialization process, shut down the
module. It's pointless to leave the module in any intermediate state
during the initialization.
Both logical CPU scope initialization and shutting down the TDX module
require calling SEAMCALL on all boot-time present CPUs. For simplicity
just temporarily disable CPU hotplug during the module initialization.
Note TDX architecturally doesn't support physical CPU hot-add/removal.
A non-buggy BIOS should never support ACPI CPU hot-add/removal. This
implementation doesn't explicitly handle ACPI CPU hot-add/removal but
depends on the BIOS to do the right thing.
Reviewed-by: Chao Gao <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- No change.
v5 -> v6:
- Added code to set status to TDX_MODULE_NONE if TDX module is not
loaded (Chao)
- Added Chao's Reviewed-by.
- Improved comments around cpus_read_lock().
- v3->v5 (no feedback on v4):
- Removed the check that SEAMRR and TDX KeyID have been detected on
all present cpus.
- Removed tdx_detect().
- Added num_online_cpus() to MADT-enabled CPUs check within the CPU
hotplug lock and return early with error message.
- Improved dmesg printing for TDX module detection and initialization.
---
arch/x86/include/asm/tdx.h | 2 +
arch/x86/virt/vmx/tdx/tdx.c | 150 ++++++++++++++++++++++++++++++++++++
2 files changed, 152 insertions(+)
@@ -10,15 +10,34 @@#include<linux/types.h>#include<linux/init.h>#include<linux/printk.h>+#include<linux/mutex.h>+#include<linux/cpu.h>+#include<linux/cpumask.h>#include<asm/msr-index.h>#include<asm/msr.h>#include<asm/apic.h>#include<asm/tdx.h>#include"tdx.h"+/* TDX module status during initialization */+enumtdx_module_status_t{+/* TDX module hasn't been detected and initialized */+TDX_MODULE_UNKNOWN,+/* TDX module is not loaded */+TDX_MODULE_NONE,+/* TDX module is initialized */+TDX_MODULE_INITIALIZED,+/* TDX module is shut down due to initialization error */+TDX_MODULE_SHUTDOWN,+};+staticu32tdx_keyid_start__ro_after_init;staticu32tdx_keyid_num__ro_after_init;+staticenumtdx_module_status_ttdx_module_status;+/* Prevent concurrent attempts on TDX detection and initialization */+staticDEFINE_MUTEX(tdx_module_lock);+/**DetectTDXprivateKeyIDstoseewhetherTDXhasbeenenabledbythe*BIOS.BothinitializingtheTDXmoduleandrunningTDXguestrequire
@@ -104,3 +123,134 @@ bool platform_tdx_enabled(void){return!!tdx_keyid_num;}++/*+*DetectandinitializetheTDXmodule.+*+*Return-ENODEVwhentheTDXmoduleisnotloaded,0whenit+*issuccessfullyinitialized,orothererrorwhenitfailsto+*initialize.+*/+staticintinit_tdx_module(void)+{+/* The TDX module hasn't been detected */+return-ENODEV;+}++staticvoidshutdown_tdx_module(void)+{+/* TODO: Shut down the TDX module */+}++staticint__tdx_enable(void)+{+intret;++/*+*InitializingtheTDXmodulerequiresdoingSEAMCALLonall+*boot-timepresentCPUs.Forsimplicitytemporarilydisable+*CPUhotplugtopreventanyCPUfromgoingofflineduring+*theinitialization.+*/+cpus_read_lock();++/*+*Checkwhetherallboot-timepresentCPUsareonlineand+*returnearlywithamessagesotheusercanbeaware.+*+*Noteanon-buggyBIOSshouldneversupportphysical(ACPI)+*CPUhotplugwhenTDXisenabled,andallboot-timepresent+*CPUshouldbeenabledinMADT,sothereshouldbeno+*disabled_cpusandnum_processorswon'tchangeatruntime+*either.+*/+if(disabled_cpus||num_online_cpus()!=num_processors){+pr_err("Unable to initialize the TDX module when there's offline CPU(s).\n");+ret=-EINVAL;+gotoout;+}++ret=init_tdx_module();+if(ret==-ENODEV){+pr_info("TDX module is not loaded.\n");+tdx_module_status=TDX_MODULE_NONE;+gotoout;+}++/*+*ShutdowntheTDXmoduleincaseofanyerrorduringthe+*initializationprocess.It'smeaninglesstoleavetheTDX+*moduleinanymiddlestateoftheinitializationprocess.+*+*ShuttingdownthemodulealsorequiresdoingSEAMCALLonall+*MADT-enabledCPUs.DoitwhileCPUhotplugisdisabled.+*+*Returnallerrorsduringtheinitializationas-EFAULTasthe+*moduleisalwaysshutdown.+*/+if(ret){+pr_info("Failed to initialize TDX module. Shut it down.\n");+shutdown_tdx_module();+tdx_module_status=TDX_MODULE_SHUTDOWN;+ret=-EFAULT;+gotoout;+}++pr_info("TDX module initialized.\n");+tdx_module_status=TDX_MODULE_INITIALIZED;+out:+cpus_read_unlock();++returnret;+}++/**+*tdx_enable-EnableTDXbyinitializingtheTDXmodule+*+*CallertomakesureallCPUsareonlineandinVMXoperationbefore+*callingthisfunction.CPUhotplugistemporarilydisabledinternally+*topreventanycpufromgoingoffline.+*+*Thisfunctioncanbecalledinparallelbymultiplecallers.+*+*Return:+*+**0:TheTDXmodulehasbeensuccessfullyinitialized.+**-ENODEV:TheTDXmoduleisnotloaded,orTDXisnotsupported.+**-EINVAL:TheTDXmodulecannotbeinitializedduetocertain+*conditionsarenotmet(i.e.whennotallMADT-enabled+*CPUsarenotonline).+**-EFAULT:Otherinternalfatalerrors,ortheTDXmoduleisin+*shutdownmodeduetoitfailedtoinitializeinprevious+*attempts.+*/+inttdx_enable(void)+{+intret;++if(!platform_tdx_enabled())+return-ENODEV;++mutex_lock(&tdx_module_lock);++switch(tdx_module_status){+caseTDX_MODULE_UNKNOWN:+ret=__tdx_enable();+break;+caseTDX_MODULE_NONE:+ret=-ENODEV;+break;+caseTDX_MODULE_INITIALIZED:+ret=0;+break;+default:+WARN_ON_ONCE(tdx_module_status!=TDX_MODULE_SHUTDOWN);+ret=-EFAULT;+break;+}++mutex_unlock(&tdx_module_lock);++returnret;+}+EXPORT_SYMBOL_GPL(tdx_enable);
From: Kai Huang <hidden> Date: 2022-11-21 00:27:58
The MMIO/xAPIC interface has some problems, most notably the APIC LEAK
[1]. This bug allows an attacker to use the APIC MMIO interface to
extract data from the SGX enclave.
TDX is not immune from this either. Early check X2APIC and disable TDX
if X2APIC is not enabled, and make INTEL_TDX_HOST depend on X86_X2APIC.
[1]: https://aepicleak.com/aepicleak.pdf
Link: https://lore.kernel.org/lkml/d6ffb489-7024-ff74-bd2f-d1e06573bb82@intel.com/
Link: https://lore.kernel.org/lkml/ba80b303-31bf-d44a-b05d-5c0f83038798@intel.com/
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Changed to use "Link" for the two lore links to get rid of checkpatch
warning.
---
arch/x86/Kconfig | 1 +
arch/x86/virt/vmx/tdx/tdx.c | 11 +++++++++++
2 files changed, 12 insertions(+)
@@ -81,6 +82,16 @@ static int __init tdx_init(void)gotono_tdx;}+/*+*TDXrequiresX2APICbeingenabledtopreventpotentialdata+*leakviaAPICMMIOregisters.JustdisableTDXifnotusing+*X2APIC.+*/+if(!x2apic_enabled()){+pr_info("Disable TDX as X2APIC is not enabled.\n");+gotono_tdx;+}+return0;no_tdx:clear_tdx();
From: Kai Huang <hidden> Date: 2022-11-21 00:28:09
TDX introduces a new CPU mode: Secure Arbitration Mode (SEAM). This
mode runs only the TDX module itself or other code to load the TDX
module.
The host kernel communicates with SEAM software via a new SEAMCALL
instruction. This is conceptually similar to a guest->host hypercall,
except it is made from the host to SEAM software instead.
The TDX module defines a set of SEAMCALL leaf functions to allow the
host to initialize it, and to create and run protected VMs. SEAMCALL
leaf functions use an ABI different from the x86-64 system-v ABI.
Instead, they share the same ABI with the TDCALL leaf functions.
Implement a function __seamcall() to allow the host to make SEAMCALL
to SEAM software using the TDX_MODULE_CALL macro which is the common
assembly for both SEAMCALL and TDCALL.
SEAMCALL instruction causes #GP when SEAMRR isn't enabled, and #UD when
CPU is not in VMX operation. The current TDX_MODULE_CALL macro doesn't
handle any of them. There's no way to check whether the CPU is in VMX
operation or not.
Initializing the TDX module is done at runtime on demand, and it depends
on the caller to ensure CPU is in VMX operation before making SEAMCALL.
To avoid getting Oops when the caller mistakenly tries to initialize the
TDX module when CPU is not in VMX operation, extend the TDX_MODULE_CALL
macro to handle #UD (and also #GP, which can theoretically still happen
when TDX isn't actually enabled by the BIOS, i.e. due to BIOS bug).
Introduce two new TDX error codes for #UD and #GP respectively so the
caller can distinguish. Also, Opportunistically put the new TDX error
codes and the existing TDX_SEAMCALL_VMFAILINVALID into INTEL_TDX_HOST
Kconfig option as they are only used when it is on.
As __seamcall() can potentially return multiple error codes, besides the
actual SEAMCALL leaf function return code, also introduce a wrapper
function seamcall() to convert the __seamcall() error code to the kernel
error code, so the caller doesn't need to duplicate the code to check
return value of __seamcall() and return kernel error code accordingly.
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- No change.
v5 -> v6:
- Added code to handle #UD and #GP (Dave).
- Moved the seamcall() wrapper function to this patch, and used a
temporary __always_unused to avoid compile warning (Dave).
- v3 -> v5 (no feedback on v4):
- Explicitly tell TDX_SEAMCALL_VMFAILINVALID is returned if the
SEAMCALL itself fails.
- Improve the changelog.
---
arch/x86/include/asm/tdx.h | 9 ++++++
arch/x86/virt/vmx/tdx/Makefile | 2 +-
arch/x86/virt/vmx/tdx/seamcall.S | 52 ++++++++++++++++++++++++++++++++
arch/x86/virt/vmx/tdx/tdx.c | 42 ++++++++++++++++++++++++++
arch/x86/virt/vmx/tdx/tdx.h | 8 +++++
arch/x86/virt/vmx/tdx/tdxcall.S | 19 ++++++++++--
6 files changed, 129 insertions(+), 3 deletions(-)
create mode 100644 arch/x86/virt/vmx/tdx/seamcall.S
@@ -124,6 +124,48 @@ bool platform_tdx_enabled(void)return!!tdx_keyid_num;}+/*+*Wrapperof__seamcall()toconvertSEAMCALLleaffunctionerrorcode+*tokernelerrorcode.@seamcall_retand@outcontaintheSEAMCALL+*leaffunctionreturncodeandtheadditionaloutputrespectivelyif+*notNULL.+*/+staticint__always_unusedseamcall(u64fn,u64rcx,u64rdx,u64r8,u64r9,+u64*seamcall_ret,+structtdx_module_output*out)+{+u64sret;++sret=__seamcall(fn,rcx,rdx,r8,r9,out);++/* Save SEAMCALL return code if caller wants it */+if(seamcall_ret)+*seamcall_ret=sret;++/* SEAMCALL was successful */+if(!sret)+return0;++switch(sret){+caseTDX_SEAMCALL_GP:+/*+*platform_tdx_enabled()ischeckedtobetrue+*beforemakinganySEAMCALL.+*/+WARN_ON_ONCE(1);+fallthrough;+caseTDX_SEAMCALL_VMFAILINVALID:+/* Return -ENODEV if the TDX module is not loaded. */+return-ENODEV;+caseTDX_SEAMCALL_UD:+/* Return -EINVAL if CPU isn't in VMX operation. */+return-EINVAL;+default:+/* Return -EIO if the actual SEAMCALL leaf failed. */+return-EIO;+}+}+/**DetectandinitializetheTDXmodule.*
@@ -57,10 +59,23 @@*ThisvaluewillneverbeusedasactualSEAMCALLerrorcodeas*itisfromtheReservedstatuscodeclass.*/-jnc.Lno_vmfailinvalid+jnc.Lseamcall_outmov$TDX_SEAMCALL_VMFAILINVALID,%rax-.Lno_vmfailinvalid:+jmp.Lseamcall_out+2:+/*+*SEAMCALLcaused#GP or #UD. By reaching here %eax contains+*thetrapnumber.ConvertthetrapnumbertotheTDXerror+*codebysettingTDX_SW_ERRORtothehigh32-bitsof%rax.+*+*NotecannotORTDX_SW_ERRORdirectlyto%raxasORinstruction+*onlyaccepts32-bitimmediateatmost.+*/+mov$TDX_SW_ERROR,%r12+orq%r12, %rax+_ASM_EXTABLE_FAULT(1b,2b)+.Lseamcall_out:.elsetdcall.endif
From: Kai Huang <hidden> Date: 2022-11-21 00:28:15
TDX supports shutting down the TDX module at any time during its
lifetime. After the module is shut down, no further TDX module SEAMCALL
leaf functions can be made to the module on any logical cpu.
Shut down the TDX module in case of any error during the initialization
process. It's pointless to leave the TDX module in some middle state.
Shutting down the TDX module requires calling TDH.SYS.LP.SHUTDOWN on all
BIOS-enabled CPUs, and the SEMACALL can run concurrently on different
CPUs. Implement a mechanism to run SEAMCALL concurrently on all online
CPUs and use it to shut down the module. Later logical-cpu scope module
initialization will use it too.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- No change.
v5 -> v6:
- Removed the seamcall() wrapper to previous patch (Dave).
- v3 -> v5 (no feedback on v4):
- Added a wrapper of __seamcall() to print error code if SEAMCALL fails.
- Made the seamcall_on_each_cpu() void.
- Removed 'seamcall_ret' and 'tdx_module_out' from
'struct seamcall_ctx', as they must be local variable.
- Added the comments to tdx_init() and one paragraph to changelog to
explain the caller should handle VMXON.
- Called out after shut down, no "TDX module" SEAMCALL can be made.
---
arch/x86/virt/vmx/tdx/tdx.c | 43 +++++++++++++++++++++++++++++++++----
arch/x86/virt/vmx/tdx/tdx.h | 5 +++++
2 files changed, 44 insertions(+), 4 deletions(-)
@@ -181,7 +214,9 @@ static int init_tdx_module(void)staticvoidshutdown_tdx_module(void){-/* TODO: Shut down the TDX module */+structseamcall_ctxsc={.fn=TDH_SYS_LP_SHUTDOWN};++seamcall_on_each_cpu(&sc);}staticint__tdx_enable(void)
From: Kai Huang <hidden> Date: 2022-11-21 00:28:21
After the global module initialization, the next step is logical-cpu
scope module initialization. Logical-cpu initialization requires
calling TDH.SYS.LP.INIT on all BIOS-enabled CPUs. This SEAMCALL can run
concurrently on all CPUs.
Use the helper introduced for shutting down the module to do logical-cpu
scope initialization.
Signed-off-by: Kai Huang <redacted>
---
arch/x86/virt/vmx/tdx/tdx.c | 14 ++++++++++++++
arch/x86/virt/vmx/tdx/tdx.h | 1 +
2 files changed, 15 insertions(+)
From: Kai Huang <hidden> Date: 2022-11-21 00:28:26
The first step of initializing the module is to call TDH.SYS.INIT once
on any logical cpu to do module global initialization. Do the module
global initialization.
It also detects the TDX module, as seamcall() returns -ENODEV when the
module is not loaded.
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Improved changelog.
---
arch/x86/virt/vmx/tdx/tdx.c | 19 +++++++++++++++++--
arch/x86/virt/vmx/tdx/tdx.h | 1 +
2 files changed, 18 insertions(+), 2 deletions(-)
From: Kai Huang <hidden> Date: 2022-11-21 00:28:49
TDX provides increased levels of memory confidentiality and integrity.
This requires special hardware support for features like memory
encryption and storage of memory integrity checksums. Not all memory
satisfies these requirements.
As a result, TDX introduced the concept of a "Convertible Memory Region"
(CMR). During boot, the firmware builds a list of all of the memory
ranges which can provide the TDX security guarantees. The list of these
ranges, along with TDX module information, is available to the kernel by
querying the TDX module via TDH.SYS.INFO SEAMCALL.
The host kernel can choose whether or not to use all convertible memory
regions as TDX-usable memory. Before the TDX module is ready to create
any TDX guests, the kernel needs to configure the TDX-usable memory
regions by passing an array of "TD Memory Regions" (TDMRs) to the TDX
module. Constructing the TDMR array requires information of both the
TDX module (TDSYSINFO_STRUCT) and the Convertible Memory Regions. Call
TDH.SYS.INFO to get this information as a preparation.
Use static variables for both TDSYSINFO_STRUCT and CMR array to avoid
having to pass them as function arguments when constructing the TDMR
array. And they are too big to be put to the stack anyway. Also, KVM
needs to use the TDSYSINFO_STRUCT to create TDX guests.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Simplified the check of CMRs due to the fact that TDX actually
verifies CMRs (that are passed by the BIOS) before enabling TDX.
- Changed the function name from check_cmrs() -> trim_empty_cmrs().
- Added CMR page aligned check so that later patch can just get the PFN
using ">> PAGE_SHIFT".
v5 -> v6:
- Added to also print TDX module's attribute (Isaku).
- Removed all arguments in tdx_gete_sysinfo() to use static variables
of 'tdx_sysinfo' and 'tdx_cmr_array' directly as they are all used
directly in other functions in later patches.
- Added Isaku's Reviewed-by.
- v3 -> v5 (no feedback on v4):
- Renamed sanitize_cmrs() to check_cmrs().
- Removed unnecessary sanity check against tdx_sysinfo and tdx_cmr_array
actual size returned by TDH.SYS.INFO.
- Changed -EFAULT to -EINVAL in couple places.
- Added comments around tdx_sysinfo and tdx_cmr_array saying they are
used by TDH.SYS.INFO ABI.
- Changed to pass 'tdx_sysinfo' and 'tdx_cmr_array' as function
arguments in tdx_get_sysinfo().
- Changed to only print BIOS-CMR when check_cmrs() fails.
---
arch/x86/virt/vmx/tdx/tdx.c | 125 ++++++++++++++++++++++++++++++++++++
arch/x86/virt/vmx/tdx/tdx.h | 61 ++++++++++++++++++
2 files changed, 186 insertions(+)
@@ -40,6 +41,11 @@ static enum tdx_module_status_t tdx_module_status;/* Prevent concurrent attempts on TDX detection and initialization */staticDEFINE_MUTEX(tdx_module_lock);+/* Below two are used in TDH.SYS.INFO SEAMCALL ABI */+staticstructtdsysinfo_structtdx_sysinfo;+staticstructcmr_infotdx_cmr_array[MAX_CMRS]__aligned(CMR_INFO_ARRAY_ALIGNMENT);+staticinttdx_cmr_num;+/**DetectTDXprivateKeyIDstoseewhetherTDXhasbeenenabledbythe*BIOS.BothinitializingtheTDXmoduleandrunningTDXguestrequire
@@ -208,6 +214,121 @@ static int tdx_module_init_cpus(void)returnatomic_read(&sc.err);}+staticinlineboolis_cmr_empty(structcmr_info*cmr)+{+return!cmr->size;+}++staticinlineboolis_cmr_ok(structcmr_info*cmr)+{+/* CMR must be page aligned */+returnIS_ALIGNED(cmr->base,PAGE_SIZE)&&+IS_ALIGNED(cmr->size,PAGE_SIZE);+}++staticvoidprint_cmrs(structcmr_info*cmr_array,intcmr_num,+constchar*name)+{+inti;++for(i=0;i<cmr_num;i++){+structcmr_info*cmr=&cmr_array[i];++pr_info("%s : [0x%llx, 0x%llx)\n",name,+cmr->base,cmr->base+cmr->size);+}+}++/* Check CMRs reported by TDH.SYS.INFO, and trim tail empty CMRs. */+staticinttrim_empty_cmrs(structcmr_info*cmr_array,int*actual_cmr_num)+{+structcmr_info*cmr;+inti,cmr_num;++/*+*IntelTDXmodulespec,20.7.3CMR_INFO:+*+*TDH.SYS.INFOleaffunctionreturnsaMAX_CMRS(32)entry+*arrayofCMR_INFOentries.TheCMRsaresortedfromthe+*lowestbaseaddresstothehighestbaseaddress,andthey+*arenon-overlapping.+*+*ThisimpliesthatBIOSmaygenerateinvalidemptyentries+*iftotalCMRsarelessthan32.Needtoskipthemmanually.+*+*CMRalsomustbe4Kaligned.TDXdoesn'ttrustBIOS.TDX+*actuallyverifiesCMRsbeforeitgetsenabled,soanything+*doesn'tmeetabovemeanskernelbug(orTDXisbroken).+*/+cmr=&cmr_array[0];+/* There must be at least one valid CMR */+if(WARN_ON_ONCE(is_cmr_empty(cmr)||!is_cmr_ok(cmr)))+gotoerr;++cmr_num=*actual_cmr_num;+for(i=1;i<cmr_num;i++){+structcmr_info*cmr=&cmr_array[i];+structcmr_info*prev_cmr=NULL;++/* Skip further empty CMRs */+if(is_cmr_empty(cmr))+break;++/*+*DosanitycheckanywaytomakesureCMRs:+*-are4Kaligned+*-don'toverlap+*-areinaddressascendingorder.+*/+if(WARN_ON_ONCE(!is_cmr_ok(cmr)))+gotoerr;++prev_cmr=&cmr_array[i-1];+if(WARN_ON_ONCE((prev_cmr->base+prev_cmr->size)>+cmr->base))+gotoerr;+}++/* Update the actual number of CMRs */+*actual_cmr_num=i;++/* Print kernel checked CMRs */+print_cmrs(cmr_array,*actual_cmr_num,"Kernel-checked-CMR");++return0;+err:+pr_info("[TDX broken ?]: Invalid CMRs detected\n");+print_cmrs(cmr_array,cmr_num,"BIOS-CMR");+return-EINVAL;+}++staticinttdx_get_sysinfo(void)+{+structtdx_module_outputout;+intret;++BUILD_BUG_ON(sizeof(structtdsysinfo_struct)!=TDSYSINFO_STRUCT_SIZE);++ret=seamcall(TDH_SYS_INFO,__pa(&tdx_sysinfo),TDSYSINFO_STRUCT_SIZE,+__pa(tdx_cmr_array),MAX_CMRS,NULL,&out);+if(ret)+returnret;++/* R9 contains the actual entries written the CMR array. */+tdx_cmr_num=out.r9;++pr_info("TDX module: atributes 0x%x, vendor_id 0x%x, major_version %u, minor_version %u, build_date %u, build_num %u",+tdx_sysinfo.attributes,tdx_sysinfo.vendor_id,+tdx_sysinfo.major_version,tdx_sysinfo.minor_version,+tdx_sysinfo.build_date,tdx_sysinfo.build_num);++/*+*trim_empty_cmrs()updatestheactualnumberofCMRsby+*droppingalltailemptyCMRs.+*/+returntrim_empty_cmrs(tdx_cmr_array,&tdx_cmr_num);+}+/**DetectandinitializetheTDXmodule.*
@@ -232,6 +353,10 @@ static int init_tdx_module(void)if(ret)gotoout;+ret=tdx_get_sysinfo();+if(ret)+gotoout;+/**Return-EINVALuntilallstepsofTDXmoduleinitialization*processaredone.
From: Kai Huang <hidden> Date: 2022-11-21 00:28:55
TDX reports a list of "Convertible Memory Region" (CMR) to indicate all
memory regions that can possibly be used by the TDX module, but they are
not automatically usable to the TDX module. As a step of initializing
the TDX module, the kernel needs to choose a list of memory regions (out
from convertible memory regions) that the TDX module can use and pass
those regions to the TDX module. Once this is done, those "TDX-usable"
memory regions are fixed during module's lifetime. No more TDX-usable
memory can be added to the TDX module after that.
The initial support of TDX guests will only allocate TDX guest memory
from the global page allocator. To keep things simple, this initial
implementation simply guarantees all pages in the page allocator are TDX
memory. To achieve this, use all system memory in the core-mm at the
time of initializing the TDX module as TDX memory, and at the meantime,
refuse to add any non-TDX-memory in the memory hotplug.
Specifically, walk through all memory regions managed by memblock and
add them to a global list of "TDX-usable" memory regions, which is a
fixed list after the module initialization (or empty if initialization
fails). To reject non-TDX-memory in memory hotplug, add an additional
check in arch_add_memory() to check whether the new region is covered by
any region in the "TDX-usable" memory region list.
Note this requires all memory regions in memblock are TDX convertible
memory when initializing the TDX module. This is true in practice if no
new memory has been hot-added before initializing the TDX module, since
in practice all boot-time present DIMM is TDX convertible memory. If
any new memory has been hot-added, then initializing the TDX module will
fail due to that memory region is not covered by CMR.
This can be enhanced in the future, i.e. by allowing adding non-TDX
memory to a separate NUMA node. In this case, the "TDX-capable" nodes
and the "non-TDX-capable" nodes can co-exist, but the kernel/userspace
needs to guarantee memory pages for TDX guests are always allocated from
the "TDX-capable" nodes.
Note TDX assumes convertible memory is always physically present during
machine's runtime. A non-buggy BIOS should never support hot-removal of
any convertible memory. This implementation doesn't handle ACPI memory
removal but depends on the BIOS to behave correctly.
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Changed to use all system memory in memblock at the time of
initializing the TDX module as TDX memory
- Added memory hotplug support
---
arch/x86/Kconfig | 1 +
arch/x86/include/asm/tdx.h | 3 +
arch/x86/mm/init_64.c | 10 ++
arch/x86/virt/vmx/tdx/tdx.c | 183 ++++++++++++++++++++++++++++++++++++
4 files changed, 197 insertions(+)
@@ -46,6 +58,9 @@ static struct tdsysinfo_struct tdx_sysinfo;staticstructcmr_infotdx_cmr_array[MAX_CMRS]__aligned(CMR_INFO_ARRAY_ALIGNMENT);staticinttdx_cmr_num;+/* All TDX-usable memory regions */+staticLIST_HEAD(tdx_memlist);+/**DetectTDXprivateKeyIDstoseewhetherTDXhasbeenenabledbythe*BIOS.BothinitializingtheTDXmoduleandrunningTDXguestrequire
@@ -329,6 +344,107 @@ static int tdx_get_sysinfo(void)returntrim_empty_cmrs(tdx_cmr_array,&tdx_cmr_num);}+/* Check whether the given pfn range is covered by any CMR or not. */+staticboolpfn_range_covered_by_cmr(unsignedlongstart_pfn,+unsignedlongend_pfn)+{+inti;++for(i=0;i<tdx_cmr_num;i++){+structcmr_info*cmr=&tdx_cmr_array[i];+unsignedlongcmr_start_pfn;+unsignedlongcmr_end_pfn;++cmr_start_pfn=cmr->base>>PAGE_SHIFT;+cmr_end_pfn=(cmr->base+cmr->size)>>PAGE_SHIFT;++if(start_pfn>=cmr_start_pfn&&end_pfn<=cmr_end_pfn)+returntrue;+}++returnfalse;+}++/*+*AddamemoryregiononagivennodeasaTDXmemoryblock.Thecaller+*tomakesureallmemoryregionsareaddedinaddressascendingorder+*anddon'toverlap.+*/+staticintadd_tdx_memblock(unsignedlongstart_pfn,unsignedlongend_pfn,+intnid)+{+structtdx_memblock*tmb;++tmb=kmalloc(sizeof(*tmb),GFP_KERNEL);+if(!tmb)+return-ENOMEM;++INIT_LIST_HEAD(&tmb->list);+tmb->start_pfn=start_pfn;+tmb->end_pfn=end_pfn;+tmb->nid=nid;++list_add_tail(&tmb->list,&tdx_memlist);+return0;+}++staticvoidfree_tdx_memory(void)+{+while(!list_empty(&tdx_memlist)){+structtdx_memblock*tmb=list_first_entry(&tdx_memlist,+structtdx_memblock,list);++list_del(&tmb->list);+kfree(tmb);+}+}++/*+*Addallmemblockmemoryregionstothe@tdx_memlistasTDXmemory.+*Mustbecalledwhenget_online_mems()iscalledbythecaller.+*/+staticintbuild_tdx_memory(void)+{+unsignedlongstart_pfn,end_pfn;+inti,nid,ret;++for_each_mem_pfn_range(i,MAX_NUMNODES,&start_pfn,&end_pfn,&nid){+/*+*Thefirst1MBmaynotbereportedasTDXconvertible+*memory.ManuallyexcludethemasTDXmemory.+*+*Thisisfineasthefirst1MBisalreadyreservedin+*reserve_real_mode()andwon'tenduptoZONE_DMAas+*freepageanyway.+*/+start_pfn=max(start_pfn,(unsignedlong)SZ_1M>>PAGE_SHIFT);+if(start_pfn>=end_pfn)+continue;++/* Verify memory is truly TDX convertible memory */+if(!pfn_range_covered_by_cmr(start_pfn,end_pfn)){+pr_info("Memory region [0x%lx, 0x%lx) is not TDX convertible memorry.\n",+start_pfn<<PAGE_SHIFT,+end_pfn<<PAGE_SHIFT);+return-EINVAL;+}++/*+*AddthememoryregionsasTDXmemory.Theregionsin+*memblockhasalreadyguaranteedtheyareinaddress+*ascendingorderanddon'toverlap.+*/+ret=add_tdx_memblock(start_pfn,end_pfn,nid);+if(ret)+gotoerr;+}++return0;+err:+free_tdx_memory();+returnret;+}+/**DetectandinitializetheTDXmodule.*
@@ -357,12 +473,56 @@ static int init_tdx_module(void)if(ret)gotoout;+/*+*AllmemoryregionsthatcanbeusedbytheTDXmodulemustbe+*passedtotheTDXmoduleduringthemoduleinitialization.+*Oncethisisdone,all"TDX-usable"memoryregionsarefixed+*duringmodule'sruntime.+*+*TheinitialsupportofTDXguestsonlyallocatesmemoryfrom+*theglobalpageallocator.Tokeepthingssimple,fornow+*justmakesureallpagesinthepageallocatorareTDXmemory.+*+*Toachievethis,useallsystemmemoryinthecore-mmatthe+*timeofinitializingtheTDXmoduleasTDXmemory,andatthe+*meantime,rejectanynewmemoryinmemoryhot-add.+*+*Thisworksasinpractice,allboot-timepresentDIMMisTDX+*convertiblememory.Howeverifanynewmemoryishot-added+*beforeinitializingtheTDXmodule,theinitializationwill+*failduetothatmemoryisnotcoveredbyCMR.+*+*Thiscanbeenhancedinthefuture,i.e.byallowingaddingor+*onliningnon-TDXmemorytoaseparatenode,inwhichcasethe+*"TDX-capable"nodesandthe"non-TDX-capable"nodescanexist+*together--theuserspace/kerneljustneedstomakesurepages+*forTDXguestsmustcomefromthose"TDX-capable"nodes.+*+*BuildthelistofTDXmemoryregionsasmentionedaboveso+*theycanbepassedtotheTDXmodulelater.+*/+get_online_mems();++ret=build_tdx_memory();+if(ret)+gotoout;/**Return-EINVALuntilallstepsofTDXmoduleinitialization*processaredone.*/ret=-EINVAL;out:+/*+*Memoryhotplugchecksthehot-addedmemoryregionagainstthe+*@tdx_memlisttoseeiftheregionisTDXmemory.+*+*Doput_online_mems()heretomakesureanymodificationto+*@tdx_memlistisdonewhileholdingthememoryhotplugread+*lock,sothatthememoryhotplugpathcanjustcheckthe+*@tdx_memlistw/oholdingthe@tdx_module_lockwhichmaycause+*deadlock.+*/+put_online_mems();returnret;}
@@ -485,3 +645,26 @@ int tdx_enable(void)returnret;}EXPORT_SYMBOL_GPL(tdx_enable);++/*+*CheckwhetherthegivenrangeisTDXmemory.Mustbecalledbetween+*mem_hotplug_begin()/mem_hotplug_done().+*/+booltdx_cc_memory_compatible(unsignedlongstart_pfn,unsignedlongend_pfn)+{+structtdx_memblock*tmb;++/* Empty list means TDX isn't enabled successfully */+if(list_empty(&tdx_memlist))+returntrue;++list_for_each_entry(tmb,&tdx_memlist,list){+/*+*ThenewrangeisTDXmemoryifitisfullycovered+*byanyTDXmemoryblock.+*/+if(start_pfn>=tmb->start_pfn&&end_pfn<=tmb->end_pfn)+returntrue;+}+returnfalse;+}
From: Kai Huang <hidden> Date: 2022-11-21 00:28:57
TDX provides increased levels of memory confidentiality and integrity.
This requires special hardware support for features like memory
encryption and storage of memory integrity checksums. Not all memory
satisfies these requirements.
As a result, the TDX introduced the concept of a "Convertible Memory
Region" (CMR). During boot, the firmware builds a list of all of the
memory ranges which can provide the TDX security guarantees. The list
of these ranges is available to the kernel by querying the TDX module.
The TDX architecture needs additional metadata to record things like
which TD guest "owns" a given page of memory. This metadata essentially
serves as the 'struct page' for the TDX module. The space for this
metadata is not reserved by the hardware up front and must be allocated
by the kernel and given to the TDX module.
Since this metadata consumes space, the VMM can choose whether or not to
allocate it for a given area of convertible memory. If it chooses not
to, the memory cannot receive TDX protections and can not be used by TDX
guests as private memory.
For every memory region that the VMM wants to use as TDX memory, it sets
up a "TD Memory Region" (TDMR). Each TDMR represents a physically
contiguous convertible range and must also have its own physically
contiguous metadata table, referred to as a Physical Address Metadata
Table (PAMT), to track status for each page in the TDMR range.
Unlike a CMR, each TDMR requires 1G granularity and alignment. To
support physical RAM areas that don't meet those strict requirements,
each TDMR permits a number of internal "reserved areas" which can be
placed over memory holes. If PAMT metadata is placed within a TDMR it
must be covered by one of these reserved areas.
Let's summarize the concepts:
CMR - Firmware-enumerated physical ranges that support TDX. CMRs are
4K aligned.
TDMR - Physical address range which is chosen by the kernel to support
TDX. 1G granularity and alignment required. Each TDMR has
reserved areas where TDX memory holes and overlapping PAMTs can
be put into.
PAMT - Physically contiguous TDX metadata. One table for each page size
per TDMR. Roughly 1/256th of TDMR in size. 256G TDMR = ~1G
PAMT.
As one step of initializing the TDX module, the kernel configures
TDX-usable memory regions by passing an array of TDMRs to the TDX module.
Constructing the array of TDMRs consists below steps:
1) Create TDMRs to cover all memory regions that the TDX module can use;
2) Allocate and set up PAMT for each TDMR;
3) Set up reserved areas for each TDMR.
Add a placeholder to construct TDMRs to do the above steps after all
TDX memory regions are verified to be truly convertible. Always free
TDMRs at the end of the initialization (no matter successful or not)
as TDMRs are only used during the initialization.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Improved commit message to explain 'int' overflow cannot happen
in cal_tdmr_size() and alloc_tdmr_array(). -- Andy/Dave.
v5 -> v6:
- construct_tdmrs_memblock() -> construct_tdmrs() as 'tdx_memblock' is
used instead of memblock.
- Added Isaku's Reviewed-by.
- v3 -> v5 (no feedback on v4):
- Moved calculating TDMR size to this patch.
- Changed to use alloc_pages_exact() to allocate buffer for all TDMRs
once, instead of allocating each TDMR individually.
- Removed "crypto protection" in the changelog.
- -EFAULT -> -EINVAL in couple of places.
---
arch/x86/virt/vmx/tdx/tdx.c | 83 +++++++++++++++++++++++++++++++++++++
arch/x86/virt/vmx/tdx/tdx.h | 23 ++++++++++
2 files changed, 106 insertions(+)
@@ -445,6 +445,63 @@ static int build_tdx_memory(void)returnret;}+/* Calculate the actual TDMR_INFO size */+staticinlineintcal_tdmr_size(void)+{+inttdmr_sz;++/*+*TheactualsizeofTDMR_INFOdependsonthemaximumnumber+*ofreservedareas.+*+*Note:forTDX1.0themax_reserved_per_tdmris16,and+*TDMR_INFOsizeisalignedupto512-byte.Evenitis+*extendedinthefuture,itwouldbeinsaneifTDMR_INFO+*becomeslargerthan4K.Thetdmr_szhereshouldnever+*overflow.+*/+tdmr_sz=sizeof(structtdmr_info);+tdmr_sz+=sizeof(structtdmr_reserved_area)*+tdx_sysinfo.max_reserved_per_tdmr;++/*+*TDXrequireseachTDMR_INFOtobe512-bytealigned.Always+*roundupTDMR_INFOsizetothe512-byteboundary.+*/+returnALIGN(tdmr_sz,TDMR_INFO_ALIGNMENT);+}++staticstructtdmr_info*alloc_tdmr_array(int*array_sz)+{+/*+*TDXrequireseachTDMR_INFOtobe512-bytealigned.+*Usealloc_pages_exact()toallocateallTDMRsatonce.+*EachTDMR_INFOwillstillbe512-bytealignedsince+*cal_tdmr_size()alwaysreturns512-bytealignedsize.+*/+*array_sz=cal_tdmr_size()*tdx_sysinfo.max_tdmrs;++/*+*Zerothebufferso'structtdmr_info::size'canbe+*usedtodeterminewhetheraTDMRisvalid.+*+*Note:forTDX1.0themax_tdmrsis64andTDMR_INFOsize+*is512-byte.Eventheyareextendedinthefuture,it+*wouldbeinsaneifthetotalsizeexceeds4MB.+*/+returnalloc_pages_exact(*array_sz,GFP_KERNEL|__GFP_ZERO);+}++/*+*ConstructanarrayofTDMRstocoverallTDXmemoryranges.+*TheactualnumberofTDMRsiskeptto@tdmr_num.+*/+staticintconstruct_tdmrs(structtdmr_info*tdmr_array,int*tdmr_num)+{+/* Return -EINVAL until constructing TDMRs is done */+return-EINVAL;+}+/**DetectandinitializetheTDXmodule.*
@@ -454,6 +511,9 @@ static int build_tdx_memory(void)*/staticintinit_tdx_module(void){+structtdmr_info*tdmr_array;+inttdmr_array_sz;+inttdmr_num;intret;/*
@@ -506,11 +566,34 @@ static int init_tdx_module(void)ret=build_tdx_memory();if(ret)gotoout;++/* Prepare enough space to construct TDMRs */+tdmr_array=alloc_tdmr_array(&tdmr_array_sz);+if(!tdmr_array){+ret=-ENOMEM;+gotoout_free_tdx_mem;+}++/* Construct TDMRs to cover all TDX memory ranges */+ret=construct_tdmrs(tdmr_array,&tdmr_num);+if(ret)+gotoout_free_tdmrs;+/**Return-EINVALuntilallstepsofTDXmoduleinitialization*processaredone.*/ret=-EINVAL;+out_free_tdmrs:+/*+*ThearrayofTDMRsisfreednomattertheinitializationis+*successfulornot.Theyarenotneededanymoreafterthe+*moduleinitialization.+*/+free_pages_exact(tdmr_array,tdmr_array_sz);+out_free_tdx_mem:+if(ret)+free_tdx_memory();out:/**Memoryhotplugchecksthehot-addedmemoryregionagainstthe
From: Kai Huang <hidden> Date: 2022-11-21 00:29:14
The kernel configures TDX-usable memory regions by passing an array of
"TD Memory Regions" (TDMRs) to the TDX module. Each TDMR contains the
information of the base/size of a memory region, the base/size of the
associated Physical Address Metadata Table (PAMT) and a list of reserved
areas in the region.
Create a number of TDMRs to cover all TDX memory regions. To keep it
simple, always try to create one TDMR for each memory region. As the
first step only set up the base/size for each TDMR.
Each TDMR must be 1G aligned and the size must be in 1G granularity.
This implies that one TDMR could cover multiple memory regions. If a
memory region spans the 1GB boundary and the former part is already
covered by the previous TDMR, just create a new TDMR for the remaining
part.
TDX only supports a limited number of TDMRs. Disable TDX if all TDMRs
are consumed but there is more memory region to cover.
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- No change.
v5 -> v6:
- Rebase due to using 'tdx_memblock' instead of memblock.
- v3 -> v5 (no feedback on v4):
- Removed allocating TDMR individually.
- Improved changelog by using Dave's words.
- Made TDMR_START() and TDMR_END() as static inline function.
---
arch/x86/virt/vmx/tdx/tdx.c | 104 +++++++++++++++++++++++++++++++++++-
1 file changed, 103 insertions(+), 1 deletion(-)
@@ -445,6 +445,24 @@ static int build_tdx_memory(void)returnret;}+/* TDMR must be 1gb aligned */+#define TDMR_ALIGNMENT BIT_ULL(30)+#define TDMR_PFN_ALIGNMENT (TDMR_ALIGNMENT >> PAGE_SHIFT)++/* Align up and down the address to TDMR boundary */+#define TDMR_ALIGN_DOWN(_addr) ALIGN_DOWN((_addr), TDMR_ALIGNMENT)+#define TDMR_ALIGN_UP(_addr) ALIGN((_addr), TDMR_ALIGNMENT)++staticinlineu64tdmr_start(structtdmr_info*tdmr)+{+returntdmr->base;+}++staticinlineu64tdmr_end(structtdmr_info*tdmr)+{+returntdmr->base+tdmr->size;+}+/* Calculate the actual TDMR_INFO size */staticinlineintcal_tdmr_size(void){
@@ -492,14 +510,98 @@ static struct tdmr_info *alloc_tdmr_array(int *array_sz)returnalloc_pages_exact(*array_sz,GFP_KERNEL|__GFP_ZERO);}+staticstructtdmr_info*tdmr_array_entry(structtdmr_info*tdmr_array,+intidx)+{+return(structtdmr_info*)((unsignedlong)tdmr_array++cal_tdmr_size()*idx);+}++/*+*CreateTDMRstocoverallTDXmemoryregions.Theactualnumber+*ofTDMRsissetto@tdmr_num.+*/+staticintcreate_tdmrs(structtdmr_info*tdmr_array,int*tdmr_num)+{+structtdx_memblock*tmb;+inttdmr_idx=0;++/*+*LoopoverTDXmemoryregionsandcreateTDMRstocoverthem.+*Tokeepitsimple,alwaystrytouseoneTDMRtocover+*onememoryregion.+*/+list_for_each_entry(tmb,&tdx_memlist,list){+structtdmr_info*tdmr;+u64start,end;++tdmr=tdmr_array_entry(tdmr_array,tdmr_idx);+start=TDMR_ALIGN_DOWN(tmb->start_pfn<<PAGE_SHIFT);+end=TDMR_ALIGN_UP(tmb->end_pfn<<PAGE_SHIFT);++/*+*IfthecurrentTDMR'ssizehasn'tbeeninitialized,+*itisanewTDMRtocoverthenewmemoryregion.+*Otherwise,thecurrentTDMRhasalreadycoveredthe+*previousmemoryregion.Inthelattercase,check+*whetherthecurrentmemoryregionhasbeenfullyor+*partiallycoveredbythecurrentTDMR,sinceTDMRis+*1Galigned.+*/+if(tdmr->size){+/*+*Looptothenextmemoryregionifthecurrent+*blockhasalreadybeenfullycoveredbythe+*currentTDMR.+*/+if(end<=tdmr_end(tdmr))+continue;++/*+*Ifpartofthecurrentmemoryregionhas+*alreadybeencoveredbythecurrentTDMR,+*skipthealreadycoveredpart.+*/+if(start<tdmr_end(tdmr))+start=tdmr_end(tdmr);++/*+*CreateanewTDMRtocoverthecurrentmemory+*region,ortheremainingpartofit.+*/+tdmr_idx++;+if(tdmr_idx>=tdx_sysinfo.max_tdmrs)+return-E2BIG;++tdmr=tdmr_array_entry(tdmr_array,tdmr_idx);+}++tdmr->base=start;+tdmr->size=end-start;+}++/* @tdmr_idx is always the index of last valid TDMR. */+*tdmr_num=tdmr_idx+1;++return0;+}+/**ConstructanarrayofTDMRstocoverallTDXmemoryranges.*TheactualnumberofTDMRsiskeptto@tdmr_num.*/staticintconstruct_tdmrs(structtdmr_info*tdmr_array,int*tdmr_num){+intret;++ret=create_tdmrs(tdmr_array,tdmr_num);+if(ret)+gotoerr;+/* Return -EINVAL until constructing TDMRs is done */-return-EINVAL;+ret=-EINVAL;+err:+returnret;}/*
From: Kai Huang <hidden> Date: 2022-11-21 00:29:19
The TDX module uses additional metadata to record things like which
guest "owns" a given page of memory. This metadata, referred as
Physical Address Metadata Table (PAMT), essentially serves as the
'struct page' for the TDX module. PAMTs are not reserved by hardware
up front. They must be allocated by the kernel and then given to the
TDX module.
TDX supports 3 page sizes: 4K, 2M, and 1G. Each "TD Memory Region"
(TDMR) has 3 PAMTs to track the 3 supported page sizes. Each PAMT must
be a physically contiguous area from a Convertible Memory Region (CMR).
However, the PAMTs which track pages in one TDMR do not need to reside
within that TDMR but can be anywhere in CMRs. If one PAMT overlaps with
any TDMR, the overlapping part must be reported as a reserved area in
that particular TDMR.
Use alloc_contig_pages() since PAMT must be a physically contiguous area
and it may be potentially large (~1/256th of the size of the given TDMR).
The downside is alloc_contig_pages() may fail at runtime. One (bad)
mitigation is to launch a TD guest early during system boot to get those
PAMTs allocated at early time, but the only way to fix is to add a boot
option to allocate or reserve PAMTs during kernel boot.
TDX only supports a limited number of reserved areas per TDMR to cover
both PAMTs and memory holes within the given TDMR. If many PAMTs are
allocated within a single TDMR, the reserved areas may not be sufficient
to cover all of them.
Adopt the following policies when allocating PAMTs for a given TDMR:
- Allocate three PAMTs of the TDMR in one contiguous chunk to minimize
the total number of reserved areas consumed for PAMTs.
- Try to first allocate PAMT from the local node of the TDMR for better
NUMA locality.
Also dump out how many pages are allocated for PAMTs when the TDX module
is initialized successfully.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Changes due to using macros instead of 'enum' for TDX supported page
sizes.
v5 -> v6:
- Rebase due to using 'tdx_memblock' instead of memblock.
- 'int pamt_entry_nr' -> 'unsigned long nr_pamt_entries' (Dave/Sagis).
- Improved comment around tdmr_get_nid() (Dave).
- Improved comment in tdmr_set_up_pamt() around breaking the PAMT
into PAMTs for 4K/2M/1G (Dave).
- tdmrs_get_pamt_pages() -> tdmrs_count_pamt_pages() (Dave).
- v3 -> v5 (no feedback on v4):
- Used memblock to get the NUMA node for given TDMR.
- Removed tdmr_get_pamt_sz() helper but use open-code instead.
- Changed to use 'switch .. case..' for each TDX supported page size in
tdmr_get_pamt_sz() (the original __tdmr_get_pamt_sz()).
- Added printing out memory used for PAMT allocation when TDX module is
initialized successfully.
- Explained downside of alloc_contig_pages() in changelog.
- Addressed other minor comments.
---
arch/x86/Kconfig | 1 +
arch/x86/virt/vmx/tdx/tdx.c | 191 ++++++++++++++++++++++++++++++++++++
2 files changed, 192 insertions(+)
@@ -586,6 +586,187 @@ static int create_tdmrs(struct tdmr_info *tdmr_array, int *tdmr_num)return0;}+/*+*CalculatePAMTsizegivenaTDMRandapagesize.Thereturned+*PAMTsizeisalwaysalignedupto4Kpageboundary.+*/+staticunsignedlongtdmr_get_pamt_sz(structtdmr_info*tdmr,intpgsz)+{+unsignedlongpamt_sz,nr_pamt_entries;++switch(pgsz){+caseTDX_PS_4K:+nr_pamt_entries=tdmr->size>>PAGE_SHIFT;+break;+caseTDX_PS_2M:+nr_pamt_entries=tdmr->size>>PMD_SHIFT;+break;+caseTDX_PS_1G:+nr_pamt_entries=tdmr->size>>PUD_SHIFT;+break;+default:+WARN_ON_ONCE(1);+return0;+}++pamt_sz=nr_pamt_entries*tdx_sysinfo.pamt_entry_size;+/* TDX requires PAMT size must be 4K aligned */+pamt_sz=ALIGN(pamt_sz,PAGE_SIZE);++returnpamt_sz;+}++/*+*PickaNUMAnodeonwhichtoallocatethisTDMR'smetadata.+*+*ThisisimprecisesinceTDMRsare1GalignedandNUMAnodesmight+*notbe.IftheTDMRcoversmorethanonenode,justusethe_first_+*one.Thiscanleadtosmallareasofoff-nodemetadataforsome+*memory.+*/+staticinttdmr_get_nid(structtdmr_info*tdmr)+{+structtdx_memblock*tmb;++/* Find the first memory region covered by the TDMR */+list_for_each_entry(tmb,&tdx_memlist,list){+if(tmb->end_pfn>(tdmr_start(tdmr)>>PAGE_SHIFT))+returntmb->nid;+}++/*+*FallbacktoallocatingtheTDMR'smetadatafromnode0when+*noTDXmemoryblockcanbefound.Thisshouldneverhappen+*sinceTDMRsoriginatefromTDXmemoryblocks.+*/+WARN_ON_ONCE(1);+return0;+}++staticinttdmr_set_up_pamt(structtdmr_info*tdmr)+{+unsignedlongpamt_base[TDX_PS_1G+1];+unsignedlongpamt_size[TDX_PS_1G+1];+unsignedlongtdmr_pamt_base;+unsignedlongtdmr_pamt_size;+structpage*pamt;+intpgsz,nid;++nid=tdmr_get_nid(tdmr);++/*+*CalculatethePAMTsizeforeachTDXsupportedpagesize+*andthetotalPAMTsize.+*/+tdmr_pamt_size=0;+for(pgsz=TDX_PS_4K;pgsz<=TDX_PS_1G;pgsz++){+pamt_size[pgsz]=tdmr_get_pamt_sz(tdmr,pgsz);+tdmr_pamt_size+=pamt_size[pgsz];+}++/*+*Allocateonechunkofphysicallycontiguousmemoryforall+*PAMTs.ThishelpsminimizethePAMT'suseofreservedareas+*inoverlappedTDMRs.+*/+pamt=alloc_contig_pages(tdmr_pamt_size>>PAGE_SHIFT,GFP_KERNEL,+nid,&node_online_map);+if(!pamt)+return-ENOMEM;++/*+*Breakthecontiguousallocationbackupintothe+*individualPAMTsforeachpagesize.+*/+tdmr_pamt_base=page_to_pfn(pamt)<<PAGE_SHIFT;+for(pgsz=TDX_PS_4K;pgsz<=TDX_PS_1G;pgsz++){+pamt_base[pgsz]=tdmr_pamt_base;+tdmr_pamt_base+=pamt_size[pgsz];+}++tdmr->pamt_4k_base=pamt_base[TDX_PS_4K];+tdmr->pamt_4k_size=pamt_size[TDX_PS_4K];+tdmr->pamt_2m_base=pamt_base[TDX_PS_2M];+tdmr->pamt_2m_size=pamt_size[TDX_PS_2M];+tdmr->pamt_1g_base=pamt_base[TDX_PS_1G];+tdmr->pamt_1g_size=pamt_size[TDX_PS_1G];++return0;+}++staticvoidtdmr_get_pamt(structtdmr_info*tdmr,unsignedlong*pamt_pfn,+unsignedlong*pamt_npages)+{+unsignedlongpamt_base,pamt_sz;++/*+*ThePAMTwasallocatedinonecontiguousunit.The4KPAMT+*shouldalwayspointtothebeginningofthatallocation.+*/+pamt_base=tdmr->pamt_4k_base;+pamt_sz=tdmr->pamt_4k_size+tdmr->pamt_2m_size+tdmr->pamt_1g_size;++*pamt_pfn=pamt_base>>PAGE_SHIFT;+*pamt_npages=pamt_sz>>PAGE_SHIFT;+}++staticvoidtdmr_free_pamt(structtdmr_info*tdmr)+{+unsignedlongpamt_pfn,pamt_npages;++tdmr_get_pamt(tdmr,&pamt_pfn,&pamt_npages);++/* Do nothing if PAMT hasn't been allocated for this TDMR */+if(!pamt_npages)+return;++if(WARN_ON_ONCE(!pamt_pfn))+return;++free_contig_range(pamt_pfn,pamt_npages);+}++staticvoidtdmrs_free_pamt_all(structtdmr_info*tdmr_array,inttdmr_num)+{+inti;++for(i=0;i<tdmr_num;i++)+tdmr_free_pamt(tdmr_array_entry(tdmr_array,i));+}++/* Allocate and set up PAMTs for all TDMRs */+staticinttdmrs_set_up_pamt_all(structtdmr_info*tdmr_array,inttdmr_num)+{+inti,ret=0;++for(i=0;i<tdmr_num;i++){+ret=tdmr_set_up_pamt(tdmr_array_entry(tdmr_array,i));+if(ret)+gotoerr;+}++return0;+err:+tdmrs_free_pamt_all(tdmr_array,tdmr_num);+returnret;+}++staticunsignedlongtdmrs_count_pamt_pages(structtdmr_info*tdmr_array,+inttdmr_num)+{+unsignedlongpamt_npages=0;+inti;++for(i=0;i<tdmr_num;i++){+unsignedlongpfn,npages;++tdmr_get_pamt(tdmr_array_entry(tdmr_array,i),&pfn,&npages);+pamt_npages+=npages;+}++returnpamt_npages;+}+/**ConstructanarrayofTDMRstocoverallTDXmemoryranges.*TheactualnumberofTDMRsiskeptto@tdmr_num.
@@ -598,8 +779,13 @@ static int construct_tdmrs(struct tdmr_info *tdmr_array, int *tdmr_num)if(ret)gotoerr;+ret=tdmrs_set_up_pamt_all(tdmr_array,*tdmr_num);+if(ret)+gotoerr;+/* Return -EINVAL until constructing TDMRs is done */ret=-EINVAL;+tdmrs_free_pamt_all(tdmr_array,*tdmr_num);err:returnret;}
@@ -686,6 +872,11 @@ static int init_tdx_module(void)*processaredone.*/ret=-EINVAL;+if(ret)+tdmrs_free_pamt_all(tdmr_array,tdmr_num);+else+pr_info("%lu pages allocated for PAMT.\n",+tdmrs_count_pamt_pages(tdmr_array,tdmr_num));out_free_tdmrs:/**ThearrayofTDMRsisfreednomattertheinitializationis
From: Kai Huang <hidden> Date: 2022-11-21 00:30:35
As the last step of constructing TDMRs, set up reserved areas for all
TDMRs. For each TDMR, put all memory holes within this TDMR to the
reserved areas. And for all PAMTs which overlap with this TDMR, put
all the overlapping parts to reserved areas too.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- No change.
v5 -> v6:
- Rebase due to using 'tdx_memblock' instead of memblock.
- Split tdmr_set_up_rsvd_areas() into two functions to handle memory
hole and PAMT respectively.
- Added Isaku's Reviewed-by.
---
arch/x86/virt/vmx/tdx/tdx.c | 190 +++++++++++++++++++++++++++++++++++-
1 file changed, 188 insertions(+), 2 deletions(-)
@@ -767,6 +768,187 @@ static unsigned long tdmrs_count_pamt_pages(struct tdmr_info *tdmr_array,returnpamt_npages;}+staticinttdmr_add_rsvd_area(structtdmr_info*tdmr,int*p_idx,+u64addr,u64size)+{+structtdmr_reserved_area*rsvd_areas=tdmr->reserved_areas;+intidx=*p_idx;++/* Reserved area must be 4K aligned in offset and size */+if(WARN_ON(addr&~PAGE_MASK||size&~PAGE_MASK))+return-EINVAL;++/* Cannot exceed maximum reserved areas supported by TDX */+if(idx>=tdx_sysinfo.max_reserved_per_tdmr)+return-E2BIG;++rsvd_areas[idx].offset=addr-tdmr->base;+rsvd_areas[idx].size=size;++*p_idx=idx+1;++return0;+}++staticinttdmr_set_up_memory_hole_rsvd_areas(structtdmr_info*tdmr,+int*rsvd_idx)+{+structtdx_memblock*tmb;+u64prev_end;+intret;++/* Mark holes between memory regions as reserved */+prev_end=tdmr_start(tdmr);+list_for_each_entry(tmb,&tdx_memlist,list){+u64start,end;++start=tmb->start_pfn<<PAGE_SHIFT;+end=tmb->end_pfn<<PAGE_SHIFT;++/* Break if this region is after the TDMR */+if(start>=tdmr_end(tdmr))+break;++/* Exclude regions before this TDMR */+if(end<tdmr_start(tdmr))+continue;++/*+*Skipifnoholeexistsbeforethisregion."<="is+*usedbecauseonememoryregionmightspantwoTDMRs+*(whenthepreviousTDMRcoverspartofthisregion).+*Inthiscasethestartaddressofthisregionis+*smallerthanthestartaddressofthesecondTDMR.+*+*Updatetheprev_endtotheendofthisregionwhere+*thepossiblememoryholestarts.+*/+if(start<=prev_end){+prev_end=end;+continue;+}++/* Add the hole before this region */+ret=tdmr_add_rsvd_area(tdmr,rsvd_idx,prev_end,+start-prev_end);+if(ret)+returnret;++prev_end=end;+}++/* Add the hole after the last region if it exists. */+if(prev_end<tdmr_end(tdmr)){+ret=tdmr_add_rsvd_area(tdmr,rsvd_idx,prev_end,+tdmr_end(tdmr)-prev_end);+if(ret)+returnret;+}++return0;+}++staticinttdmr_set_up_pamt_rsvd_areas(structtdmr_info*tdmr,int*rsvd_idx,+structtdmr_info*tdmr_array,+inttdmr_num)+{+inti,ret;++/*+*IfanyPAMToverlapswiththisTDMR,theoverlappingpart+*mustalsobeputtothereservedareatoo.Walkoverall+*TDMRstofindoutthoseoverlappingPAMTsandputthemto+*reservedareas.+*/+for(i=0;i<tdmr_num;i++){+structtdmr_info*tmp=tdmr_array_entry(tdmr_array,i);+unsignedlongpamt_start_pfn,pamt_npages;+u64pamt_start,pamt_end;++tdmr_get_pamt(tmp,&pamt_start_pfn,&pamt_npages);+/* Each TDMR must already have PAMT allocated */+WARN_ON_ONCE(!pamt_npages||!pamt_start_pfn);++pamt_start=pamt_start_pfn<<PAGE_SHIFT;+pamt_end=pamt_start+(pamt_npages<<PAGE_SHIFT);++/* Skip PAMTs outside of the given TDMR */+if((pamt_end<=tdmr_start(tdmr))||+(pamt_start>=tdmr_end(tdmr)))+continue;++/* Only mark the part within the TDMR as reserved */+if(pamt_start<tdmr_start(tdmr))+pamt_start=tdmr_start(tdmr);+if(pamt_end>tdmr_end(tdmr))+pamt_end=tdmr_end(tdmr);++ret=tdmr_add_rsvd_area(tdmr,rsvd_idx,pamt_start,+pamt_end-pamt_start);+if(ret)+returnret;+}++return0;+}++/* Compare function called by sort() for TDMR reserved areas */+staticintrsvd_area_cmp_func(constvoid*a,constvoid*b)+{+structtdmr_reserved_area*r1=(structtdmr_reserved_area*)a;+structtdmr_reserved_area*r2=(structtdmr_reserved_area*)b;++if(r1->offset+r1->size<=r2->offset)+return-1;+if(r1->offset>=r2->offset+r2->size)+return1;++/* Reserved areas cannot overlap. The caller should guarantee. */+WARN_ON_ONCE(1);+return-1;+}++/* Set up reserved areas for a TDMR, including memory holes and PAMTs */+staticinttdmr_set_up_rsvd_areas(structtdmr_info*tdmr,+structtdmr_info*tdmr_array,+inttdmr_num)+{+intret,rsvd_idx=0;++/* Put all memory holes within the TDMR into reserved areas */+ret=tdmr_set_up_memory_hole_rsvd_areas(tdmr,&rsvd_idx);+if(ret)+returnret;++/* Put all (overlapping) PAMTs within the TDMR into reserved areas */+ret=tdmr_set_up_pamt_rsvd_areas(tdmr,&rsvd_idx,tdmr_array,tdmr_num);+if(ret)+returnret;++/* TDX requires reserved areas listed in address ascending order */+sort(tdmr->reserved_areas,rsvd_idx,sizeof(structtdmr_reserved_area),+rsvd_area_cmp_func,NULL);++return0;+}++staticinttdmrs_set_up_rsvd_areas_all(structtdmr_info*tdmr_array,+inttdmr_num)+{+inti;++for(i=0;i<tdmr_num;i++){+intret;++ret=tdmr_set_up_rsvd_areas(tdmr_array_entry(tdmr_array,i),+tdmr_array,tdmr_num);+if(ret)+returnret;+}++return0;+}+/**ConstructanarrayofTDMRstocoverallTDXmemoryranges.*TheactualnumberofTDMRsiskeptto@tdmr_num.
@@ -783,8 +965,12 @@ static int construct_tdmrs(struct tdmr_info *tdmr_array, int *tdmr_num)if(ret)gotoerr;-/* Return -EINVAL until constructing TDMRs is done */-ret=-EINVAL;+ret=tdmrs_set_up_rsvd_areas_all(tdmr_array,*tdmr_num);+if(ret)+gotoerr_free_pamts;++return0;+err_free_pamts:tdmrs_free_pamt_all(tdmr_array,*tdmr_num);err:returnret;
From: Kai Huang <hidden> Date: 2022-11-21 00:30:39
TDX module initialization requires to use one TDX private KeyID as the
global KeyID to protect the TDX module metadata. The global KeyID is
configured to the TDX module along with TDMRs.
Just reserve the first TDX private KeyID as the global KeyID. Keep the
global KeyID as a static variable as KVM will need to use it too.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
arch/x86/virt/vmx/tdx/tdx.c | 9 +++++++++
1 file changed, 9 insertions(+)
@@ -62,6 +62,9 @@ static int tdx_cmr_num;/* All TDX-usable memory regions */staticLIST_HEAD(tdx_memlist);+/* TDX module global KeyID. Used in TDH.SYS.CONFIG ABI. */+staticu32tdx_global_keyid;+/**DetectTDXprivateKeyIDstoseewhetherTDXhasbeenenabledbythe*BIOS.BothinitializingtheTDXmoduleandrunningTDXguestrequire
@@ -1053,6 +1056,12 @@ static int init_tdx_module(void)if(ret)gotoout_free_tdmrs;+/*+*ReservethefirstTDXKeyIDasglobalKeyIDtoprotect+*TDXmodulemetadata.+*/+tdx_global_keyid=tdx_keyid_start;+/**Return-EINVALuntilallstepsofTDXmoduleinitialization*processaredone.
From: Kai Huang <hidden> Date: 2022-11-21 00:30:42
After the TDX-usable memory regions are constructed in an array of TDMRs
and the global KeyID is reserved, configure them to the TDX module using
TDH.SYS.CONFIG SEAMCALL. TDH.SYS.CONFIG can only be called once and can
be done on any logical cpu.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
arch/x86/virt/vmx/tdx/tdx.c | 37 +++++++++++++++++++++++++++++++++++++
arch/x86/virt/vmx/tdx/tdx.h | 2 ++
2 files changed, 39 insertions(+)
@@ -979,6 +979,37 @@ static int construct_tdmrs(struct tdmr_info *tdmr_array, int *tdmr_num)returnret;}+staticintconfig_tdx_module(structtdmr_info*tdmr_array,inttdmr_num,+u64global_keyid)+{+u64*tdmr_pa_array;+inti,array_sz;+u64ret;++/*+*TDMR_INFOentriesareconfiguredtotheTDXmoduleviaan+*arrayofthephysicaladdressofeachTDMR_INFO.TDXmodule+*requiresthearrayitselftobe512-bytealigned.Roundup+*thearraysizeto512-bytealignedsothebufferallocated+*bykzalloc()willmeetthealignmentrequirement.+*/+array_sz=ALIGN(tdmr_num*sizeof(u64),TDMR_INFO_PA_ARRAY_ALIGNMENT);+tdmr_pa_array=kzalloc(array_sz,GFP_KERNEL);+if(!tdmr_pa_array)+return-ENOMEM;++for(i=0;i<tdmr_num;i++)+tdmr_pa_array[i]=__pa(tdmr_array_entry(tdmr_array,i));++ret=seamcall(TDH_SYS_CONFIG,__pa(tdmr_pa_array),tdmr_num,+global_keyid,0,NULL,NULL);++/* Free the array as it is not required anymore. */+kfree(tdmr_pa_array);++returnret;+}+/**DetectandinitializetheTDXmodule.*
@@ -1062,11 +1093,17 @@ static int init_tdx_module(void)*/tdx_global_keyid=tdx_keyid_start;+/* Pass the TDMRs and the global KeyID to the TDX module */+ret=config_tdx_module(tdmr_array,tdmr_num,tdx_global_keyid);+if(ret)+gotoout_free_pamts;+/**Return-EINVALuntilallstepsofTDXmoduleinitialization*processaredone.*/ret=-EINVAL;+out_free_pamts:if(ret)tdmrs_free_pamt_all(tdmr_array,tdmr_num);else
From: Kai Huang <hidden> Date: 2022-11-21 00:31:06
After the array of TDMRs and the global KeyID are configured to the TDX
module, use TDH.SYS.KEY.CONFIG to configure the key of the global KeyID
on all packages.
TDH.SYS.KEY.CONFIG must be done on one (any) cpu for each package. And
it cannot run concurrently on different CPUs. Implement a helper to
run SEAMCALL on one cpu for each package one by one, and use it to
configure the global KeyID on all packages.
Intel hardware doesn't guarantee cache coherency across different
KeyIDs. The kernel needs to flush PAMT's dirty cachelines (associated
with KeyID 0) before the TDX module uses the global KeyID to access the
PAMT. Following the TDX module specification, flush cache before
configuring the global KeyID on all packages.
Given the PAMT size can be large (~1/256th of system RAM), just use
WBINVD on all CPUs to flush.
Note if any TDH.SYS.KEY.CONFIG fails, the TDX module may already have
used the global KeyID to write any PAMT. Therefore, need to use WBINVD
to flush cache before freeing the PAMTs back to the kernel. Note using
MOVDIR64B (which changes the page's associated KeyID from the old TDX
private KeyID back to KeyID 0, which is used by the kernel) to clear
PMATs isn't needed, as the KeyID 0 doesn't support integrity check.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Improved changelong and comment to explain why MOVDIR64B isn't used
when returning PAMTs back to the kernel.
---
arch/x86/virt/vmx/tdx/tdx.c | 89 ++++++++++++++++++++++++++++++++++++-
arch/x86/virt/vmx/tdx/tdx.h | 1 +
2 files changed, 88 insertions(+), 2 deletions(-)
@@ -1010,6 +1050,22 @@ static int config_tdx_module(struct tdmr_info *tdmr_array, int tdmr_num,returnret;}+staticintconfig_global_keyid(void)+{+structseamcall_ctxsc={.fn=TDH_SYS_KEY_CONFIG};++/*+*ConfigurethekeyoftheglobalKeyIDonallpackagesby+*callingTDH.SYS.KEY.CONFIGonallpackagesinaserialized+*wayasitcannotrunconcurrentlyondifferentCPUs.+*+*TDH.SYS.KEY.CONFIGmayfailwithentropyerror(whichis+*arecoverableerror).Assumethisisexceedinglyrareand+*justreturnerrorifencounteredinsteadofretrying.+*/+returnseamcall_on_each_package_serialized(&sc);+}+/**DetectandinitializetheTDXmodule.*
@@ -1098,15 +1154,44 @@ static int init_tdx_module(void)if(ret)gotoout_free_pamts;+/*+*Hardwaredoesn'tguaranteecachecoherencyacrossdifferent+*KeyIDs.ThekernelneedstoflushPAMT'sdirtycachelines+*(associatedwithKeyID0)beforetheTDXmodulecanusethe+*globalKeyIDtoaccessthePAMT.GivenPAMTsarepotentially+*large(~1/256thofsystemRAM),justuseWBINVDonallcpus+*toflushthecache.+*+*FollowtheTDXspectoflushcachebeforeconfiguringthe+*globalKeyIDonallpackages.+*/+wbinvd_on_all_cpus();++/* Config the key of global KeyID on all packages */+ret=config_global_keyid();+if(ret)+gotoout_free_pamts;+/**Return-EINVALuntilallstepsofTDXmoduleinitialization*processaredone.*/ret=-EINVAL;out_free_pamts:-if(ret)+if(ret){+/*+*PartofPAMTmayalreadyhavebeeninitializedby+*TDXmodule.FlushcachebeforereturningPAMTback+*tothekernel.+*+*Notethere'snoneedtodoMOVDIR64B(whichchanges+*thepage'sassociatedKeyIDfromtheoldTDXprivate+*KeyIDbacktoKeyID0,whichisusedbythekernel),+*asKeyID0doesn'tsupportintegritycheck.+*/+wbinvd_on_all_cpus();tdmrs_free_pamt_all(tdmr_array,tdmr_num);-else+}elsepr_info("%lu pages allocated for PAMT.\n",tdmrs_count_pamt_pages(tdmr_array,tdmr_num));out_free_tdmrs:
From: Kai Huang <hidden> Date: 2022-11-21 00:31:12
There are two problems in terms of using kexec() to boot to a new kernel
when the old kernel has enabled TDX: 1) Part of the memory pages are
still TDX private pages (i.e. metadata used by the TDX module, and any
TDX guest memory if kexec() happens when there's any TDX guest alive).
2) There might be dirty cachelines associated with TDX private pages.
Because the hardware doesn't guarantee cache coherency among different
KeyIDs, the old kernel needs to flush cache (of those TDX private pages)
before booting to the new kernel. Also, reading TDX private page using
any shared non-TDX KeyID with integrity-check enabled can trigger #MC.
Therefore ideally, the kernel should convert all TDX private pages back
to normal before booting to the new kernel.
However, this implementation doesn't convert TDX private pages back to
normal in kexec() because of below considerations:
1) The kernel doesn't have existing infrastructure to track which pages
are TDX private pages.
2) The number of TDX private pages can be large, and converting all of
them (cache flush + using MOVDIR64B to clear the page) in kexec() can
be time consuming.
3) The new kernel will almost only use KeyID 0 to access memory. KeyID
0 doesn't support integrity-check, so it's OK.
4) The kernel doesn't (and may never) support MKTME. If any 3rd party
kernel ever supports MKTME, it should do MOVDIR64B to clear the page
with the new MKTME KeyID (just like TDX does) before using it.
Therefore, this implementation just flushes cache to make sure there are
no stale dirty cachelines associated with any TDX private KeyIDs before
booting to the new kernel, otherwise they may silently corrupt the new
kernel.
Following SME support, use wbinvd() to flush cache in stop_this_cpu().
Theoretically, cache flush is only needed when the TDX module has been
initialized. However initializing the TDX module is done on demand at
runtime, and it takes a mutex to read the module status. Just check
whether TDX is enabled by BIOS instead to flush cache.
Also, the current TDX module doesn't play nicely with kexec(). The TDX
module can only be initialized once during its lifetime, and there is no
ABI to reset the module to give a new clean slate to the new kernel.
Therefore ideally, if the TDX module is ever initialized, it's better
to shut it down. The new kernel won't be able to use TDX anyway (as it
needs to go through the TDX module initialization process which will
fail immediately at the first step).
However, shutting down the TDX module requires all CPUs being in VMX
operation, but there's no such guarantee as kexec() can happen at any
time (i.e. when KVM is not even loaded). So just do nothing but leave
leave the TDX module open.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Improved changelog to explain why don't convert TDX private pages back
to normal.
---
arch/x86/kernel/process.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
From: Kai Huang <hidden> Date: 2022-11-21 00:31:15
Initialize TDMRs via TDH.SYS.TDMR.INIT as the last step to complete the
TDX initialization.
All TDMRs need to be initialized using TDH.SYS.TDMR.INIT SEAMCALL before
the memory pages can be used by the TDX module. The time to initialize
TDMR is proportional to the size of the TDMR because TDH.SYS.TDMR.INIT
internally initializes the PAMT entries using the global KeyID.
To avoid long latency caused in one SEAMCALL, TDH.SYS.TDMR.INIT only
initializes an (implementation-specific) subset of PAMT entries of one
TDMR in one invocation. The caller needs to call TDH.SYS.TDMR.INIT
iteratively until all PAMT entries of the given TDMR are initialized.
TDH.SYS.TDMR.INITs can run concurrently on multiple CPUs as long as they
are initializing different TDMRs. To keep it simple, just initialize
all TDMRs one by one. On a 2-socket machine with 2.2G CPUs and 64GB
memory, each TDH.SYS.TDMR.INIT roughly takes couple of microseconds on
average, and it takes roughly dozens of milliseconds to complete the
initialization of all TDMRs while system is idle.
Reviewed-by: Isaku Yamahata <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Removed need_resched() check. -- Andi.
---
arch/x86/virt/vmx/tdx/tdx.c | 69 ++++++++++++++++++++++++++++++++++---
arch/x86/virt/vmx/tdx/tdx.h | 1 +
2 files changed, 65 insertions(+), 5 deletions(-)
@@ -1066,6 +1066,65 @@ static int config_global_keyid(void)returnseamcall_on_each_package_serialized(&sc);}+/* Initialize one TDMR */+staticintinit_tdmr(structtdmr_info*tdmr)+{+u64next;++/*+*InitializingPAMTentriesmightbetime-consuming(in+*proportiontothesizeoftherequestedTDMR).Toavoidlong+*latencyinoneSEAMCALL,TDH.SYS.TDMR.INITonlyinitializes+*an(implementation-defined)subsetofPAMTentriesinone+*invocation.+*+*CallTDH.SYS.TDMR.INITiterativelyuntilallPAMTentries+*oftherequestedTDMRareinitialized(ifnext-to-initialize+*addressmatchestheendaddressoftheTDMR).+*/+do{+structtdx_module_outputout;+intret;++ret=seamcall(TDH_SYS_TDMR_INIT,tdmr->base,0,0,0,NULL,+&out);+if(ret)+returnret;+/*+*RDXcontains'next-to-initialize'addressif+*TDH.SYS.TDMR.INTsucceeded.+*/+next=out.rdx;+/* Allow scheduling when needed */+cond_resched();+}while(next<tdmr->base+tdmr->size);++return0;+}++/* Initialize all TDMRs */+staticintinit_tdmrs(structtdmr_info*tdmr_array,inttdmr_num)+{+inti;++/*+*InitializeTDMRsone-by-oneforsimplicity,thoughtheTDX+*architecturedoesallowdifferentTDMRstobeinitializedin+*parallelonmultipleCPUs.Parallelinitializationcould+*beaddedlaterwhenthetimespentintheserializedscheme+*becomesarealconcern.+*/+for(i=0;i<tdmr_num;i++){+intret;++ret=init_tdmr(tdmr_array_entry(tdmr_array,i));+if(ret)+returnret;+}++return0;+}+/**DetectandinitializetheTDXmodule.*
@@ -1172,11 +1231,11 @@ static int init_tdx_module(void)if(ret)gotoout_free_pamts;-/*-*Return-EINVALuntilallstepsofTDXmoduleinitialization-*processaredone.-*/-ret=-EINVAL;+/* Initialize TDMRs to complete the TDX module initialization */+ret=init_tdmrs(tdmr_array,tdmr_num);+if(ret)+gotoout_free_pamts;+out_free_pamts:if(ret){/*
From: Kai Huang <hidden> Date: 2022-11-21 00:31:32
Add documentation for TDX host kernel support. There is already one
file Documentation/x86/tdx.rst containing documentation for TDX guest
internals. Also reuse it for TDX host kernel support.
Introduce a new level menu "TDX Guest Support" and move existing
materials under it, and add a new menu for TDX host kernel support.
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- Changed "TDX Memory Policy" and "Kexec()" sections.
---
Documentation/x86/tdx.rst | 181 +++++++++++++++++++++++++++++++++++---
1 file changed, 170 insertions(+), 11 deletions(-)
@@ -10,6 +10,165 @@ encrypting the guest memory. In TDX, a special module running in a special mode sits between the host and the guest and manages the guest/host separation.+TDX Host Kernel Support+=======================++TDX introduces a new CPU mode called Secure Arbitration Mode (SEAM) and+a new isolated range pointed by the SEAM Ranger Register (SEAMRR). A+CPU-attested software module called 'the TDX module' runs inside the new+isolated range to provide the functionalities to manage and run protected+VMs.++TDX also leverages Intel Multi-Key Total Memory Encryption (MKTME) to+provide crypto-protection to the VMs. TDX reserves part of MKTME KeyIDs+as TDX private KeyIDs, which are only accessible within the SEAM mode.+BIOS is responsible for partitioning legacy MKTME KeyIDs and TDX KeyIDs.++Before the TDX module can be used to create and run protected VMs, it+must be loaded into the isolated range and properly initialized. The TDX+architecture doesn't require the BIOS to load the TDX module, but the+kernel assumes it is loaded by the BIOS.++TDX boot-time detection+-----------------------++The kernel detects TDX by detecting TDX private KeyIDs during kernel+boot. Below dmesg shows when TDX is enabled by BIOS::++ [..] tdx: TDX enabled by BIOS. TDX private KeyID range: [16, 64).++TDX module detection and initialization+---------------------------------------++There is no CPUID or MSR to detect the TDX module. The kernel detects it+by initializing it.++The kernel talks to the TDX module via the new SEAMCALL instruction. The+TDX module implements SEAMCALL leaf functions to allow the kernel to+initialize it.++Initializing the TDX module consumes roughly ~1/256th system RAM size to+use it as 'metadata' for the TDX memory. It also takes additional CPU+time to initialize those metadata along with the TDX module itself. Both+are not trivial. The kernel initializes the TDX module at runtime on+demand. The caller to call tdx_enable() to initialize the TDX module::++ ret = tdx_enable();+ if (ret)+ goto no_tdx;+ // TDX is ready to use++Initializing the TDX module requires all logical CPUs being online.+tdx_enable() internally temporarily disables CPU hotplug to prevent any+CPU from going offline, but the caller still needs to guarantee all+present CPUs are online before calling tdx_enable().++Also, tdx_enable() requires all CPUs are already in VMX operation+(requirement of making SEAMCALL). Currently, tdx_enable() doesn't handle+VMXON internally, but depends on the caller to guarantee that. So far+KVM is the only user of TDX and KVM already handles VMXON.++User can consult dmesg to see the presence of the TDX module, and whether+it has been initialized.++If the TDX module is not loaded, dmesg shows below::++ [..] tdx: TDX module is not loaded.++If the TDX module is initialized successfully, dmesg shows something+like below::++ [..] tdx: TDX module: attributes 0x0, vendor_id 0x8086, major_version 1, minor_version 0, build_date 20211209, build_num 160+ [..] tdx: 65667 pages allocated for PAMT.+ [..] tdx: TDX module initialized.++If the TDX module failed to initialize, dmesg shows below::++ [..] tdx: Failed to initialize TDX module. Shut it down.++TDX Interaction to Other Kernel Components+------------------------------------------++TDX Memory Policy+~~~~~~~~~~~~~~~~~++TDX reports a list of "Convertible Memory Region" (CMR) to indicate all+memory regions that can possibly be used by the TDX module, but they are+not automatically usable to the TDX module. As a step of initializing+the TDX module, the kernel needs to choose a list of memory regions (out+from convertible memory regions) that the TDX module can use and pass+those regions to the TDX module. Once this is done, those "TDX-usable"+memory regions are fixed during module's lifetime. No more TDX-usable+memory can be added to the TDX module after that.++To keep things simple, currently the kernel simply guarantees all pages+in the page allocator are TDX memory. Specifically, the kernel uses all+system memory in the core-mm at the time of initializing the TDX module+as TDX memory, and at the meantime, refuses to add any non-TDX-memory in+the memory hotplug.++This can be enhanced in the future, i.e. by allowing adding non-TDX+memory to a separate NUMA node. In this case, the "TDX-capable" nodes+and the "non-TDX-capable" nodes can co-exist, but the kernel/userspace+needs to guarantee memory pages for TDX guests are always allocated from+the "TDX-capable" nodes.++Note TDX assumes convertible memory is always physically present during+machine's runtime. A non-buggy BIOS should never support hot-removal of+any convertible memory. This implementation doesn't handle ACPI memory+removal but depends on the BIOS to behave correctly.++CPU Hotplug+~~~~~~~~~~~++TDX doesn't support physical (ACPI) CPU hotplug. During machine boot,+TDX verifies all boot-time present logical CPUs are TDX compatible before+enabling TDX. A non-buggy BIOS should never support hot-add/removal of+physical CPU. Currently the kernel doesn't handle physical CPU hotplug,+but depends on the BIOS to behave correctly.++Note TDX works with CPU logical online/offline, thus the kernel still+allows to offline logical CPU and online it again.++Kexec()+~~~~~~~++There are two problems in terms of using kexec() to boot to a new kernel+when the old kernel has enabled TDX: 1) Part of the memory pages are+still TDX private pages (i.e. metadata used by the TDX module, and any+TDX guest memory if kexec() is executed when there's live TDX guests).+2) There might be dirty cachelines associated with TDX private pages.++Because the hardware doesn't guarantee cache coherency among different+KeyIDs, the old kernel needs to flush cache (of TDX private pages)+before booting to the new kernel. Also, the kernel doesn't convert all+TDX private pages back to normal because of below considerations:++1) The kernel doesn't have existing infrastructure to track which pages+ are TDX private page.+2) The number of TDX private pages can be large, and converting all of+ them (cache flush + using MOVDIR64B to clear the page) can be time+ consuming.+3) The new kernel will almost only use KeyID 0 to access memory. KeyID+ 0 doesn't support integrity-check, so it's OK.+4) The kernel doesn't (and may never) support MKTME. If any 3rd party+ kernel ever supports MKTME, it should do MOVDIR64B to clear the page+ with the new MKTME KeyID (just like TDX does) before using it.++The current TDX module architecture doesn't play nicely with kexec().+The TDX module can only be initialized once during its lifetime, and+there is no SEAMCALL to reset the module to give a new clean slate to+the new kernel. Therefore, ideally, if the module is ever initialized,+it's better to shut down the module. The new kernel won't be able to+use TDX anyway (as it needs to go through the TDX module initialization+process which will fail immediately at the first step).++However, there's no guarantee CPU is in VMX operation during kexec(), so+it's impractical to shut down the module. Currently, the kernel just+leaves the module in open state.++TDX Guest Support+================= Since the host cannot directly access guest registers or memory, much normal functionality of a hypervisor must be moved into the guest. This is implemented using a Virtualization Exception (#VE) that is handled by the
@@ -20,7 +179,7 @@ TDX includes new hypercall-like mechanisms for communicating from the guest to the hypervisor or the TDX module. New TDX Exceptions-==================+------------------ TDX guests behave differently from bare-metal and traditional VMX guests. In TDX guests, otherwise normal instructions or memory accesses can cause
@@ -30,7 +189,7 @@ Instructions marked with an '*' conditionally cause exceptions. The details for these instructions are discussed below. Instruction-based #VE----------------------+~~~~~~~~~~~~~~~~~~~~~- Port I/O (INS, OUTS, IN, OUT)- HLT
@@ -52,7 +211,7 @@ Instruction-based #GP- RDMSR*,WRMSR* RDMSR/WRMSR Behavior---------------------+~~~~~~~~~~~~~~~~~~~~ MSR access behavior falls into three categories:
@@ -73,7 +232,7 @@ trapping and handling in the TDX module. Other than possibly being slow, these MSRs appear to function just as they would on bare metal. CPUID Behavior---------------+~~~~~~~~~~~~~~ For some CPUID leaves and sub-leaves, the virtualized bit fields of CPUID return values (in guest EAX/EBX/ECX/EDX) are configurable by the
@@ -93,7 +252,7 @@ not know how to handle. The guest kernel may ask the hypervisor for the value with a hypercall. #VE on Memory Accesses-======================+---------------------- There are essentially two classes of TDX memory: private and shared. Private memory receives full TDX protections. Its content is protected
@@ -107,7 +266,7 @@ entries. This helps ensure that a guest does not place sensitive information in shared memory, exposing it to the untrusted hypervisor. #VE on Shared Memory---------------------+~~~~~~~~~~~~~~~~~~~~ Access to shared mappings can cause a #VE. The hypervisor ultimately controls whether a shared memory access causes a #VE, so the guest must be
@@ -127,7 +286,7 @@ be careful not to access device MMIO regions unless it is also prepared to handle a #VE. #VE on Private Pages---------------------+~~~~~~~~~~~~~~~~~~~~ An access to private mappings can also cause a #VE. Since all kernel memory is also private memory, the kernel might theoretically need to
@@ -145,7 +304,7 @@ The hypervisor is permitted to unilaterally move accepted pages to a to handle the exception. Linux #VE handler-=================+----------------- Just like page faults or #GP's, #VE exceptions can be either handled or be fatal. Typically, an unhandled userspace #VE results in a SIGSEGV.
@@ -167,7 +326,7 @@ While the block is in place, any #VE is elevated to a double fault (#DF) which is not recoverable. MMIO handling-=============+------------- In non-TDX VMs, MMIO is usually implemented by giving a guest access to a mapping which will cause a VMEXIT on access, and then the hypervisor
@@ -189,7 +348,7 @@ MMIO access via other means (like structure overlays) may result in an oops. Shared Memory Conversions-=========================+------------------------- All TDX guest memory starts out as private at boot. This memory can not be accessed by the hypervisor. However, some kernel users like device
Intel Trust Domain Extensions (TDX) protects guest VMs from malicious
host and certain physical attacks. A CPU-attested software module
called 'the TDX module' runs inside a new isolated memory range as a
trusted hypervisor to manage and run protected VMs.
Pre-TDX Intel hardware has support for a memory encryption architecture
called MKTME. The memory encryption hardware underpinning MKTME is also
used for Intel TDX. TDX ends up "stealing" some of the physical address
space from the MKTME architecture for crypto-protection to VMs. The
BIOS is responsible for partitioning the "KeyID" space between legacy
MKTME and TDX. The KeyIDs reserved for TDX are called 'TDX private
KeyIDs' or 'TDX KeyIDs' for short.
TDX doesn't trust the BIOS. During machine boot, TDX verifies the TDX
private KeyIDs are consistently and correctly programmed by the BIOS
across all CPU packages before it enables TDX on any CPU core. A valid
TDX private KeyID range on BSP indicates TDX has been enabled by the
BIOS, otherwise the BIOS is buggy.
The TDX module is expected to be loaded by the BIOS when it enables TDX,
but the kernel needs to properly initialize it before it can be used to
create and run any TDX guests. The TDX module will be initialized at
runtime by the user (i.e. KVM) on demand.
Add a new early_initcall(tdx_init) to do TDX early boot initialization.
Only detect TDX private KeyIDs for now. Some other early checks will
follow up. Also add a new function to report whether TDX has been
enabled by BIOS (TDX private KeyID range is valid). Kexec() will also
need it to determine whether need to flush dirty cachelines that are
associated with any TDX private KeyIDs before booting to the new kernel.
To start to support TDX, create a new arch/x86/virt/vmx/tdx/tdx.c for
TDX host kernel support. Add a new Kconfig option CONFIG_INTEL_TDX_HOST
to opt-in TDX host kernel support (to distinguish with TDX guest kernel
support). So far only KVM is the only user of TDX. Make the new config
option depend on KVM_INTEL.
Reviewed-by: Kirill A. Shutemov <redacted>
Signed-off-by: Kai Huang <redacted>
---
v6 -> v7:
- No change.
v5 -> v6:
- Removed SEAMRR detection to make code simpler.
- Removed the 'default N' in the KVM_TDX_HOST Kconfig (Kirill).
- Changed to use 'obj-y' in arch/x86/virt/vmx/tdx/Makefile (Kirill).
---
arch/x86/Kconfig | 12 +++++
arch/x86/Makefile | 2 +
arch/x86/include/asm/tdx.h | 7 +++
arch/x86/virt/Makefile | 2 +
arch/x86/virt/vmx/Makefile | 2 +
arch/x86/virt/vmx/tdx/Makefile | 2 +
arch/x86/virt/vmx/tdx/tdx.c | 95 ++++++++++++++++++++++++++++++++++
arch/x86/virt/vmx/tdx/tdx.h | 15 ++++++
8 files changed, 137 insertions(+)
create mode 100644 arch/x86/virt/Makefile
create mode 100644 arch/x86/virt/vmx/Makefile
create mode 100644 arch/x86/virt/vmx/tdx/Makefile
create mode 100644 arch/x86/virt/vmx/tdx/tdx.c
create mode 100644 arch/x86/virt/vmx/tdx/tdx.h
@@ -246,6 +246,8 @@ archheaders:libs-y+=arch/x86/lib/+core-y+=arch/x86/virt/+# drivers-y are linked after core-ydrivers-$(CONFIG_MATH_EMULATION)+=arch/x86/math-emu/drivers-$(CONFIG_PCI)+=arch/x86/pci/
It is not very clear why you increment tdx_keyid_start. What we read from
MSR_IA32_MKTME_KEYID_PARTITIONING is not the correct start keyid?
Also why is this global variable? At least in this patch, there seems to
be no use case.
quoted hunk
++ pr_info("TDX enabled by BIOS. TDX private KeyID range: [%u, %u)\n",+ tdx_keyid_start, tdx_keyid_start + tdx_keyid_num);++ return 0;+}++static void __init clear_tdx(void)+{+ tdx_keyid_start = tdx_keyid_num = 0;+}++static int __init tdx_init(void)+{+ if (detect_tdx())+ return -ENODEV;++ /*+ * Initializing the TDX module requires one TDX private KeyID.+ * If there's only one TDX KeyID then after module initialization+ * KVM won't be able to run any TDX guest, which makes the whole+ * thing worthless. Just disable TDX in this case.+ */+ if (tdx_keyid_num < 2) {+ pr_info("Disable TDX as there's only one TDX private KeyID available.\n");+ goto no_tdx;+ }++ return 0;+no_tdx:+ clear_tdx();+ return -ENODEV;+}+early_initcall(tdx_init);++/* Return whether the BIOS has enabled TDX */+bool platform_tdx_enabled(void)+{+ return !!tdx_keyid_num;+}
From: Huang, Kai <hidden> Date: 2022-11-21 09:16:02
On Sun, 2022-11-20 at 18:52 -0800, Sathyanarayanan Kuppuswamy wrote:
On 11/20/22 4:26 PM, Kai Huang wrote:
quoted
+/*+ * TDX supported page sizes (4K/2M/1G).+ *+ * Those values are part of the TDX module ABI. Do not change them.
It would be better if you include specification version and section
title.
Such as below?
"Those values are part of the TDX module ABI (section "Physical Page Size", TDX
module 1.0 spec). Do not change them."
Btw, Dave mentioned we should not put the "section numbers" to the comment:
https://lore.kernel.org/lkml/2a1886e7-fa5d-99e2-b1da-55ed7c0d024b@intel.com/
I was trying to follow.
From: Huang, Kai <hidden> Date: 2022-11-21 09:37:26
quoted
+static u32 tdx_keyid_start __ro_after_init;+static u32 tdx_keyid_num __ro_after_init;++/*+ * Detect TDX private KeyIDs to see whether TDX has been enabled by the+ * BIOS. Both initializing the TDX module and running TDX guest require+ * TDX private KeyID.+ *+ * TDX doesn't trust BIOS. TDX verifies all configurations from BIOS+ * are correct before enabling TDX on any core. TDX requires the BIOS+ * to correctly and consistently program TDX private KeyIDs on all CPU+ * packages. Unless there is a BIOS bug, detecting a valid TDX private+ * KeyID range on BSP indicates TDX has been enabled by the BIOS. If+ * there's such BIOS bug, it will be caught later when initializing the+ * TDX module.+ */+static int __init detect_tdx(void)+{+ int ret;++ /*+ * IA32_MKTME_KEYID_PARTIONING:+ * Bit [31:0]: Number of MKTME KeyIDs.+ * Bit [63:32]: Number of TDX private KeyIDs.+ */+ ret = rdmsr_safe(MSR_IA32_MKTME_KEYID_PARTITIONING, &tdx_keyid_start,+ &tdx_keyid_num);+ if (ret)+ return -ENODEV;++ if (!tdx_keyid_num)+ return -ENODEV;++ /*+ * KeyID 0 is for TME. MKTME KeyIDs start from 1. TDX private+ * KeyIDs start after the last MKTME KeyID.+ */+ tdx_keyid_start++;
It is not very clear why you increment tdx_keyid_start.
Please see above comment around rdmsr_safe():
/*
* IA32_MKTME_KEYID_PARTIONING:
* Bit [31:0]: Number of MKTME KeyIDs.
* Bit [63:32]: Number of TDX private KeyIDs.
*/
And the comment right above 'tdx_keyid_start++':
/*
* KeyID 0 is for TME. MKTME KeyIDs start from 1. TDX private
* KeyIDs start after the last MKTME KeyID.
*/
Do the above two comments answer the question?
Or I can be more explicit, such as below?
"
Now tdx_start_keyid is the last MKTME KeyID. TDX private KeyIDs start after the
last MKTME KeyID. Increase tdx_start_keyid by 1 to set it to the first TDX
private KeyID.
"
What we read from
MSR_IA32_MKTME_KEYID_PARTITIONING is not the correct start keyid?
TDX verifies TDX private KeyID range is configured consistently across all
packages. Any wrong KeyID range means BIOS bug, and such bug will cause TDX
being not enabled -- later TDX module initialization will catch this.
Also why is this global variable? At least in this patch, there seems to
be no use case.
Platform_tdx_enabled() uses tdx_keyid_num to determine whether TDX is enabled by
BIOS.
Also, in the changlog I can add "both initializing the TDX module and creating
TDX guest will need to use TDX private KeyID".
But I also have a comment saying something similar around ...
quoted
+ /*+ * Initializing the TDX module requires one TDX private KeyID.+ * If there's only one TDX KeyID then after module initialization+ * KVM won't be able to run any TDX guest, which makes the whole+ * thing worthless. Just disable TDX in this case.+ */+ if (tdx_keyid_num < 2) {+ pr_info("Disable TDX as there's only one TDX private KeyID available.\n");+ goto no_tdx;+ }+
From: Huang, Kai <hidden> Date: 2022-11-21 09:45:00
On Sun, 2022-11-20 at 19:51 -0800, Sathyanarayanan Kuppuswamy wrote:
On 11/20/22 4:26 PM, Kai Huang wrote:
quoted
The MMIO/xAPIC interface has some problems, most notably the APIC LEAK
"some problems" looks more generic. May be we can be specific here. Like
it has security issues?
It was quoted from below upstream commit id (I only kept the one that I quoted
to save space):
commit b8d1d163604bd1e600b062fb00de5dc42baa355f (tag: x86_apic_for_v6.1_rc1,
tip/x86/apic)
Author: Daniel Sneddon [off-list ref]
Date: Tue Aug 16 16:19:42 2022 -0700
x86/apic: Don't disable x2APIC if locked
....
The MMIO/xAPIC interface has some problems, most notably the APIC LEAK
[1]. This bug allows an attacker to use the APIC MMIO interface to
extract data from the SGX enclave.
....
[1]: https://aepicleak.com/aepicleak.pdf
Signed-off-by: Daniel Sneddon [off-list ref]
Signed-off-by: Dave Hansen [off-list ref]
Acked-by: Dave Hansen [off-list ref]
Tested-by: Neelima Krishnan [off-list ref]
Link:
https://lkml.kernel.org/r/20220816231943.1152579-1-daniel.sneddon@linux.intel.com
From: Dave Hansen <hidden> Date: 2022-11-21 18:14:46
On 11/20/22 18:52, Sathyanarayanan Kuppuswamy wrote:
On 11/20/22 4:26 PM, Kai Huang wrote:
quoted
+/*+ * TDX supported page sizes (4K/2M/1G).+ *+ * Those values are part of the TDX module ABI. Do not change them.
It would be better if you include specification version and section
title.
I actually think TDX code, in general, spends way too much time quoting
and referring to the spec.
Also, why quote the version? Do we quote the SDM version when we add
new SDM-defined architecture?
It's just busywork that bloats the kernel and adds noise. Please focus
on adding value to the comments that came from your brain and not just
pasting boilerplate gunk over and over.
On Sun, 2022-11-20 at 19:51 -0800, Sathyanarayanan Kuppuswamy wrote:
quoted
On 11/20/22 4:26 PM, Kai Huang wrote:
quoted
The MMIO/xAPIC interface has some problems, most notably the APIC LEAK
"some problems" looks more generic. May be we can be specific here. Like
it has security issues?
It was quoted from below upstream commit id (I only kept the one that I quoted
to save space):
Ok.
commit b8d1d163604bd1e600b062fb00de5dc42baa355f (tag: x86_apic_for_v6.1_rc1,
tip/x86/apic)
Author: Daniel Sneddon [off-list ref]
Date: Tue Aug 16 16:19:42 2022 -0700
x86/apic: Don't disable x2APIC if locked
....
The MMIO/xAPIC interface has some problems, most notably the APIC LEAK
[1]. This bug allows an attacker to use the APIC MMIO interface to
extract data from the SGX enclave.
....
[1]: https://aepicleak.com/aepicleak.pdf
Signed-off-by: Daniel Sneddon [off-list ref]
Signed-off-by: Dave Hansen [off-list ref]
Acked-by: Dave Hansen [off-list ref]
Tested-by: Neelima Krishnan [off-list ref]
Link:
https://lkml.kernel.org/r/20220816231943.1152579-1-daniel.sneddon@linux.intel.com
From: Dave Hansen <hidden> Date: 2022-11-21 23:47:21
On 11/20/22 16:26, Kai Huang wrote:
The MMIO/xAPIC interface has some problems, most notably the APIC LEAK
[1]. This bug allows an attacker to use the APIC MMIO interface to
extract data from the SGX enclave.
TDX is not immune from this either. Early check X2APIC and disable TDX
if X2APIC is not enabled, and make INTEL_TDX_HOST depend on X86_X2APIC.
This makes no sense.
This is TDX host code. TDX hosts are untrusted. Zero of the TDX
security guarantees are provided by the host.
What is the benefit of disabling TDX from the host if x2APIC is not
enabled? It can't be for security reasons since the host does not help
provide TDX security guarantees. It also can't be for SGX because SGX
doesn't depend on the OS doing anything in order to be secure.
So, this boils down to the most fundamental of questions you need to
answer about every patch:
What does this code do?
What end-user-visible effect is there if this code is not present?
From: Dave Hansen <hidden> Date: 2022-11-21 23:48:54
On 11/20/22 16:26, Kai Huang wrote:
quoted hunk
+/*+ * TDX supported page sizes (4K/2M/1G).+ *+ * Those values are part of the TDX module ABI. Do not change them.+ */+#define TDX_PS_4K 0+#define TDX_PS_2M 1+#define TDX_PS_1G 2
That comment can just be:
/* TDX supported page sizes from the TDX module ABI. */
I think folks understand that the kernel can't willy nilly change ABI
values.
Also why is this global variable? At least in this patch, there seems to
be no use case.
Platform_tdx_enabled() uses tdx_keyid_num to determine whether TDX is enabled by
BIOS.
Also, in the changlog I can add "both initializing the TDX module and creating
TDX guest will need to use TDX private KeyID".
But I also have a comment saying something similar around ...
I am asking about the tdx_keyid_start. It mainly used in detect_tdx(). Maybe you
declared it as global as a preparation for next patches. But it is not explained
in change log.
--
Sathyanarayanan Kuppuswamy
Linux Kernel Developer
From: Huang, Kai <hidden> Date: 2022-11-22 00:01:37
On Mon, 2022-11-21 at 15:48 -0800, Hansen, Dave wrote:
On 11/20/22 16:26, Kai Huang wrote:
quoted
+/*+ * TDX supported page sizes (4K/2M/1G).+ *+ * Those values are part of the TDX module ABI. Do not change them.+ */+#define TDX_PS_4K 0+#define TDX_PS_2M 1+#define TDX_PS_1G 2
That comment can just be:
/* TDX supported page sizes from the TDX module ABI. */
I think folks understand that the kernel can't willy nilly change ABI
values.
From: Dave Hansen <hidden> Date: 2022-11-22 00:10:50
On 11/20/22 16:26, Kai Huang wrote:
Intel Trust Domain Extensions (TDX) protects guest VMs from malicious
host and certain physical attacks. A CPU-attested software module
called 'the TDX module' runs inside a new isolated memory range as a
trusted hypervisor to manage and run protected VMs.
Pre-TDX Intel hardware has support for a memory encryption architecture
called MKTME. The memory encryption hardware underpinning MKTME is also
used for Intel TDX. TDX ends up "stealing" some of the physical address
space from the MKTME architecture for crypto-protection to VMs. The
BIOS is responsible for partitioning the "KeyID" space between legacy
MKTME and TDX. The KeyIDs reserved for TDX are called 'TDX private
KeyIDs' or 'TDX KeyIDs' for short.
TDX doesn't trust the BIOS. During machine boot, TDX verifies the TDX
private KeyIDs are consistently and correctly programmed by the BIOS
across all CPU packages before it enables TDX on any CPU core. A valid
TDX private KeyID range on BSP indicates TDX has been enabled by the
BIOS, otherwise the BIOS is buggy.
The TDX module is expected to be loaded by the BIOS when it enables TDX,
but the kernel needs to properly initialize it before it can be used to
create and run any TDX guests. The TDX module will be initialized at
runtime by the user (i.e. KVM) on demand.
Calling KVM "the user" is a stretch. How about we give actual user
facts instead of filling this with i.e.'s when there's only one actual
way it happens?
The TDX module will be initialized by the KVM subsystem when
<insert actual trigger description here>.
Add a new early_initcall(tdx_init) to do TDX early boot initialization.
Only detect TDX private KeyIDs for now. Some other early checks will
follow up.
Just say what this patch is doing. Don't try to
Also add a new function to report whether TDX has been
enabled by BIOS (TDX private KeyID range is valid). Kexec() will also
need it to determine whether need to flush dirty cachelines that are
associated with any TDX private KeyIDs before booting to the new kernel.
That last sentence doesn't parse correctly.
To start to support TDX, create a new arch/x86/virt/vmx/tdx/tdx.c for
TDX host kernel support. Add a new Kconfig option CONFIG_INTEL_TDX_HOST
to opt-in TDX host kernel support (to distinguish with TDX guest kernel
support). So far only KVM is the only user of TDX. Make the new config
option depend on KVM_INTEL.
..
quoted hunk
+config INTEL_TDX_HOST+ bool "Intel Trust Domain Extensions (TDX) host support"+ depends on CPU_SUP_INTEL+ depends on X86_64+ depends on KVM_INTEL+ help+ Intel Trust Domain Extensions (TDX) protects guest VMs from malicious+ host and certain physical attacks. This option enables necessary TDX+ support in host kernel to run protected VMs.++ If unsure, say N.+ config EFI bool "EFI runtime service support" depends on ACPI
@@ -246,6 +246,8 @@ archheaders:libs-y+=arch/x86/lib/+core-y+=arch/x86/virt/+# drivers-y are linked after core-ydrivers-$(CONFIG_MATH_EMULATION)+=arch/x86/math-emu/drivers-$(CONFIG_PCI)+=arch/x86/pci/
This comment is not right, sorry.
Talk about the function at a *HIGH* level. Don't talk about every
little detailed facet of the function. That's what the code is there for.
quoted hunk
+ * TDX doesn't trust BIOS. TDX verifies all configurations from BIOS+ * are correct before enabling TDX on any core. TDX requires the BIOS+ * to correctly and consistently program TDX private KeyIDs on all CPU+ * packages. Unless there is a BIOS bug, detecting a valid TDX private+ * KeyID range on BSP indicates TDX has been enabled by the BIOS. If+ * there's such BIOS bug, it will be caught later when initializing the+ * TDX module.+ */
I have no idea what that comment is doing. Can it just be removed?
quoted hunk
+static int __init detect_tdx(void)+{+ int ret;++ /*+ * IA32_MKTME_KEYID_PARTIONING:+ * Bit [31:0]: Number of MKTME KeyIDs.+ * Bit [63:32]: Number of TDX private KeyIDs.+ */+ ret = rdmsr_safe(MSR_IA32_MKTME_KEYID_PARTITIONING, &tdx_keyid_start,+ &tdx_keyid_num);
'tdx_keyid_start' appears to be named wrong.
quoted hunk
+ if (ret)+ return -ENODEV;++ if (!tdx_keyid_num)+ return -ENODEV;++ /*+ * KeyID 0 is for TME. MKTME KeyIDs start from 1. TDX private+ * KeyIDs start after the last MKTME KeyID.+ */
Is the TME key a "MKTME KeyID"?
+ tdx_keyid_start++;
... and this confirms it.
This probably should be:
u32 nr_mktme_keyids;
ret = rdmsr_safe(MSR_IA32_MKTME_KEYID_PARTITIONING,
&nr_mktme_keyids,
&tdx_keyid_num);
...
/* TDX KeyIDs start after the last MKTME KeyID */
tdx_keyid_start = nr_mktme_keyids + 1;
See how that makes actual logical sense and barely even needs the comment?
This is where a comment is needed and can actually help.
/*
* tdx_keyid_start/num indicate that TDX is uninitialized. This
* is used in TDX initialization error paths to take it from
* initialized -> uninitialized.
*/
quoted hunk
+static int __init tdx_init(void)+{+ if (detect_tdx())+ return -ENODEV;
This reads as:
if tdx is detected:
return error
So, first, why bother having detect_tdx() return fancy -ERRNO codes if
they're going to be throw away? You could at *least* do:
int err;
err = tdx_record_keyid_partioning();
if (err)
return err;
Note how tdx_record_keyid_partioning() actually talks about what the
function does. There's also a recent trend in x86 land not to put
obvious prefixes on functions. That would make the naming more or less
record_keyid_partioning().
I kinda like the consistent prefixes but Boris doesn't.
quoted hunk
+ /*+ * Initializing the TDX module requires one TDX private KeyID.+ * If there's only one TDX KeyID then after module initialization+ * KVM won't be able to run any TDX guest, which makes the whole+ * thing worthless. Just disable TDX in this case.+ */+ if (tdx_keyid_num < 2) {+ pr_info("Disable TDX as there's only one TDX private KeyID available.\n");+ goto no_tdx;+ }
'tdx_keyid_num' is a crummy name. Here it reads like, "if the tdx keyid
number is < 2'. Which is wrong. A better name would be: nr_tdx_keyids
That's also a horrible error message. You have:
+#define pr_fmt(fmt) "tdx: " fmt
so that message will look like:
tdx: Disable TDX as there's only one TDX private KeyID available.
How many 'TDX' strings do we need in one message. How about:
pr_info("initialization failed: too few private KeyIDs available
(%d).\n", nr_tdx_keyids;
That gives a lot more information and removes the two redundant TDX strings.
+no_tdx:
quoted hunk
+ clear_tdx();+ return -ENODEV;+}+early_initcall(tdx_init);++/* Return whether the BIOS has enabled TDX */+bool platform_tdx_enabled(void)+{+ return !!tdx_keyid_num;+}
From: Huang, Kai <hidden> Date: 2022-11-22 00:30:44
On Mon, 2022-11-21 at 15:46 -0800, Dave Hansen wrote:
On 11/20/22 16:26, Kai Huang wrote:
quoted
The MMIO/xAPIC interface has some problems, most notably the APIC LEAK
[1]. This bug allows an attacker to use the APIC MMIO interface to
extract data from the SGX enclave.
TDX is not immune from this either. Early check X2APIC and disable TDX
if X2APIC is not enabled, and make INTEL_TDX_HOST depend on X86_X2APIC.
This makes no sense.
This is TDX host code. TDX hosts are untrusted. Zero of the TDX
security guarantees are provided by the host.
What is the benefit of disabling TDX from the host if x2APIC is not
enabled? It can't be for security reasons since the host does not help
provide TDX security guarantees. It also can't be for SGX because SGX
doesn't depend on the OS doing anything in order to be secure.
Agreed. Although in practice I think if we do some hardening in the kernel, it
would raise some attack bar.
So, this boils down to the most fundamental of questions you need to
answer about every patch:
What does this code do?
What end-user-visible effect is there if this code is not present?
Considering TDX host cannot be trusted (i.e. can be attacked/modified), I agree
the check isn't needed.
I was following your suggestion in the patch which handles "x2apic locked" case:
https://lore.kernel.org/lkml/ba80b303-31bf-d44a-b05d-5c0f83038798@intel.com/
I guess I misunderstood your point.
Reading that discussion again, if I understand correctly, you just want to make
INTEL_TDX_HOST depend on X86_X2APIC?
How about still having a patch to make INTEL_TDX_HOST depend on X86_X2APIC but
with something below in the changelog?
"
TDX capable platforms are locked to X2APIC mode and cannot fall back to the
legacy xAPIC mode when TDX is enabled by the BIOS. It doesn't make sense to
turn on INTEL_TDX_HOST while X86_X2APIC is not enabled. Make INTEL_TDX_HOST
depend on X86_X2APIC.
"
From: Dave Hansen <hidden> Date: 2022-11-22 00:44:59
On 11/21/22 16:30, Huang, Kai wrote:
How about still having a patch to make INTEL_TDX_HOST depend on X86_X2APIC but
with something below in the changelog?
"
TDX capable platforms are locked to X2APIC mode and cannot fall back to the
legacy xAPIC mode when TDX is enabled by the BIOS. It doesn't make sense to
turn on INTEL_TDX_HOST while X86_X2APIC is not enabled. Make INTEL_TDX_HOST
depend on X86_X2APIC.
That's fine and it makes logical sense as a dependency. TDX host
support requires x2APIC. Period.
From: Huang, Kai <hidden> Date: 2022-11-22 00:58:26
On Mon, 2022-11-21 at 16:44 -0800, Dave Hansen wrote:
On 11/21/22 16:30, Huang, Kai wrote:
quoted
How about still having a patch to make INTEL_TDX_HOST depend on X86_X2APIC but
with something below in the changelog?
"
TDX capable platforms are locked to X2APIC mode and cannot fall back to the
legacy xAPIC mode when TDX is enabled by the BIOS. It doesn't make sense to
turn on INTEL_TDX_HOST while X86_X2APIC is not enabled. Make INTEL_TDX_HOST
depend on X86_X2APIC.
That's fine and it makes logical sense as a dependency. TDX host
support requires x2APIC. Period.
Thanks. Perhaps I can reuse your second sentence in the changelog:
"
TDX capable platforms are locked to X2APIC mode and cannot fall back to the
legacy xAPIC mode when TDX is enabled by the BIOS. TDX host support requires
x2APIC. Make INTEL_TDX_HOST depend on X86_X2APIC.
"
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 09:03:00
On Mon, Nov 21, 2022 at 01:26:26PM +1300, Kai Huang wrote:
quoted hunk
+static int __tdx_enable(void)+{+ int ret;++ /*+ * Initializing the TDX module requires doing SEAMCALL on all+ * boot-time present CPUs. For simplicity temporarily disable+ * CPU hotplug to prevent any CPU from going offline during+ * the initialization.+ */+ cpus_read_lock();++ /*+ * Check whether all boot-time present CPUs are online and+ * return early with a message so the user can be aware.+ *+ * Note a non-buggy BIOS should never support physical (ACPI)+ * CPU hotplug when TDX is enabled, and all boot-time present+ * CPU should be enabled in MADT, so there should be no+ * disabled_cpus and num_processors won't change at runtime+ * either.+ */+ if (disabled_cpus || num_online_cpus() != num_processors) {+ pr_err("Unable to initialize the TDX module when there's offline CPU(s).\n");+ ret = -EINVAL;+ goto out;+ }++ ret = init_tdx_module();+ if (ret == -ENODEV) {+ pr_info("TDX module is not loaded.\n");+ tdx_module_status = TDX_MODULE_NONE;+ goto out;+ }++ /*+ * Shut down the TDX module in case of any error during the+ * initialization process. It's meaningless to leave the TDX+ * module in any middle state of the initialization process.+ *+ * Shutting down the module also requires doing SEAMCALL on all+ * MADT-enabled CPUs. Do it while CPU hotplug is disabled.+ *+ * Return all errors during the initialization as -EFAULT as the+ * module is always shut down.+ */+ if (ret) {+ pr_info("Failed to initialize TDX module. Shut it down.\n");+ shutdown_tdx_module();+ tdx_module_status = TDX_MODULE_SHUTDOWN;+ ret = -EFAULT;+ goto out;+ }++ pr_info("TDX module initialized.\n");+ tdx_module_status = TDX_MODULE_INITIALIZED;+out:+ cpus_read_unlock();++ return ret;+}
Uhm.. so if we've offlined all the SMT siblings because of some
speculation fail or other, this TDX thing will fail to initialize?
Because as I understand it; this TDX initialization happens some random
time after boot, when the first (TDX using) KVM instance gets created,
long after the speculation mitigations are enforced.
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 09:07:02
On Mon, Nov 21, 2022 at 01:26:27PM +1300, Kai Huang wrote:
quoted hunk
+/*+ * Wrapper of __seamcall() to convert SEAMCALL leaf function error code+ * to kernel error code. @seamcall_ret and @out contain the SEAMCALL+ * leaf function return code and the additional output respectively if+ * not NULL.+ */+static int __always_unused seamcall(u64 fn, u64 rcx, u64 rdx, u64 r8, u64 r9,+ u64 *seamcall_ret,+ struct tdx_module_output *out)+{
What's the point of a 'static __always_unused' function again? Other
than to test the DCE pass of a linker, that is?
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 09:10:58
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
quoted hunk
+/*+ * Data structure to make SEAMCALL on multiple CPUs concurrently.+ * @err is set to -EFAULT when SEAMCALL fails on any cpu.+ */+struct seamcall_ctx {+ u64 fn;+ u64 rcx;+ u64 rdx;+ u64 r8;+ u64 r9;+ atomic_t err;+};
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 09:13:58
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
quoted hunk
+/*+ * Call the SEAMCALL on all online CPUs concurrently. Caller to check+ * @sc->err to determine whether any SEAMCALL failed on any cpu.+ */+static void seamcall_on_each_cpu(struct seamcall_ctx *sc)+{+ on_each_cpu(seamcall_smp_call_function, sc, true);+}
Suppose the user has NOHZ_FULL configured, and is already running
userspace that will terminate on interrupt (this is desired feature for
NOHZ_FULL), guess how happy they'll be if someone, on another parition,
manages to tickle this TDX gunk?
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 09:21:12
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
Shutting down the TDX module requires calling TDH.SYS.LP.SHUTDOWN on all
BIOS-enabled CPUs, and the SEMACALL can run concurrently on different
CPUs. Implement a mechanism to run SEAMCALL concurrently on all online
CPUs and use it to shut down the module. Later logical-cpu scope module
initialization will use it too.
Consider:
CPU0 CPU1 CPU2
local_irq_disable()
...
seamcall_on_each_cpu()
send-IPIs to 0 and 2
<IPI>
runs local seamcall
(seamcall done)
waits for 0 and 2
<has an NMI delay things>
runs seamcall
clears CSD_LOCK
</IPI>
... spinning ...
local_irq_enable()
<IPI>
runs seamcall
clears CSD_LOCK
*FINALLY* observes CSD_LOCK cleared on
all CPU and continues
</IPI>
IOW, they all 3 run seamcall at different times.
Either the Changelog is broken or this TDX crud is worse crap than I
thought possible, because the only way to actually meet that requirement
as stated is stop_machine().
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 10:11:23
On Mon, Nov 21, 2022 at 01:26:32PM +1300, Kai Huang wrote:
quoted hunk
+static int build_tdx_memory(void)+{+ unsigned long start_pfn, end_pfn;+ int i, nid, ret;++ for_each_mem_pfn_range(i, MAX_NUMNODES, &start_pfn, &end_pfn, &nid) {+ /*+ * The first 1MB may not be reported as TDX convertible+ * memory. Manually exclude them as TDX memory.+ *+ * This is fine as the first 1MB is already reserved in+ * reserve_real_mode() and won't end up to ZONE_DMA as+ * free page anyway.+ */+ start_pfn = max(start_pfn, (unsigned long)SZ_1M >> PAGE_SHIFT);+ if (start_pfn >= end_pfn)+ continue;++ /* Verify memory is truly TDX convertible memory */+ if (!pfn_range_covered_by_cmr(start_pfn, end_pfn)) {+ pr_info("Memory region [0x%lx, 0x%lx) is not TDX convertible memorry.\n",+ start_pfn << PAGE_SHIFT,+ end_pfn << PAGE_SHIFT);+ return -EINVAL;
Given how tdx_cc_memory_compatible() below relies on tdx_memlist being
empty; this error patch is wrong and should goto err.
quoted hunk
+ }++ /*+ * Add the memory regions as TDX memory. The regions in+ * memblock has already guaranteed they are in address+ * ascending order and don't overlap.+ */+ ret = add_tdx_memblock(start_pfn, end_pfn, nid);+ if (ret)+ goto err;+ }++ return 0;+err:+ free_tdx_memory();+ return ret;+}
quoted hunk
+bool tdx_cc_memory_compatible(unsigned long start_pfn, unsigned long end_pfn)+{+ struct tdx_memblock *tmb;++ /* Empty list means TDX isn't enabled successfully */+ if (list_empty(&tdx_memlist))+ return true;++ list_for_each_entry(tmb, &tdx_memlist, list) {+ /*+ * The new range is TDX memory if it is fully covered+ * by any TDX memory block.+ */+ if (start_pfn >= tmb->start_pfn && end_pfn <= tmb->end_pfn)+ return true;+ }+ return false;+}
From: Thomas Gleixner <hidden> Date: 2022-11-22 10:35:28
On Tue, Nov 22 2022 at 10:02, Peter Zijlstra wrote:
On Mon, Nov 21, 2022 at 01:26:26PM +1300, Kai Huang wrote:
quoted
+ cpus_read_unlock();++ return ret;+}
Uhm.. so if we've offlined all the SMT siblings because of some
speculation fail or other, this TDX thing will fail to initialize?
Because as I understand it; this TDX initialization happens some random
time after boot, when the first (TDX using) KVM instance gets created,
long after the speculation mitigations are enforced.
Correct. Aside of that it's completely unclear from the changelog why
TDX needs to run the seamcall on _all_ present CPUs and why it cannot
handle CPU being hotplugged later.
It's pretty much obvious that a TDX guest can only run on CPUs where
the seam module has been initialized, but where does the requirement
come from that _ALL_ CPUs must be initialized and _ALL_ CPUs must be
able to run TDX guests?
I just went and read through the documentation again.
"1. After loading the Intel TDX module, the host VMM should call the
TDH.SYS.INIT function to globally initialize the module.
2. The host VMM should then call the TDH.SYS.LP.INIT function on each
logical processor. TDH.SYS.LP.INIT is intended to initialize the
module within the scope of the Logical Processor (LP)."
This clearly tells me, that:
1) TDX must be globally initialized (once)
2) TDX must be initialized on each logical processor on which TDX
root/non-root operation should be executed
But it does not define any requirement for doing this on all logical
processors and for preventing physical hotplug (Neither for CPUs nor for
memory).
Nothing in the TDX specs and docs mentions physical hotplug or a
requirement for invoking seamcall on the world.
Thanks,
tglx
From: Huang, Kai <hidden> Date: 2022-11-22 11:35:39
quoted
The TDX module is expected to be loaded by the BIOS when it enables TDX,
but the kernel needs to properly initialize it before it can be used to
create and run any TDX guests. The TDX module will be initialized at
runtime by the user (i.e. KVM) on demand.
Calling KVM "the user" is a stretch. How about we give actual user
facts instead of filling this with i.e.'s when there's only one actual
way it happens?
The TDX module will be initialized by the KVM subsystem when
<insert actual trigger description here>.
quoted
Add a new early_initcall(tdx_init) to do TDX early boot initialization.
Only detect TDX private KeyIDs for now. Some other early checks will
follow up.
Just say what this patch is doing. Don't try to
quoted
Also add a new function to report whether TDX has been
enabled by BIOS (TDX private KeyID range is valid). Kexec() will also
need it to determine whether need to flush dirty cachelines that are
associated with any TDX private KeyIDs before booting to the new kernel.
That last sentence doesn't parse correctly.
Will do all above. Please see updated patch at the bottom.
[...]
quoted
+/*+ * Detect TDX private KeyIDs to see whether TDX has been enabled by the+ * BIOS. Both initializing the TDX module and running TDX guest require+ * TDX private KeyID.
This comment is not right, sorry.
Talk about the function at a *HIGH* level. Don't talk about every
little detailed facet of the function. That's what the code is there for.
quoted
+ * TDX doesn't trust BIOS. TDX verifies all configurations from BIOS+ * are correct before enabling TDX on any core. TDX requires the BIOS+ * to correctly and consistently program TDX private KeyIDs on all CPU+ * packages. Unless there is a BIOS bug, detecting a valid TDX private+ * KeyID range on BSP indicates TDX has been enabled by the BIOS. If+ * there's such BIOS bug, it will be caught later when initializing the+ * TDX module.+ */
I have no idea what that comment is doing. Can it just be removed?
Will remove this part and update the entire comment.
Also will address all your other comments. Please see the updated patch.
[...]
quoted
+ /*+ * KeyID 0 is for TME. MKTME KeyIDs start from 1. TDX private+ * KeyIDs start after the last MKTME KeyID.+ */
Is the TME key a "MKTME KeyID"?
I don't think so. Hardware handles TME KeyID 0 differently from non-0 MKTME
KeyIDs. And PCONFIG only accept non-0 KeyIDs.
This is where a comment is needed and can actually help.
/*
* tdx_keyid_start/num indicate that TDX is uninitialized. This
* is used in TDX initialization error paths to take it from
* initialized -> uninitialized.
*/
Just want to point out after removing the !x2apic_enabled() check, the only
thing need to do here is to detect/record the TDX KeyIDs.
And the purpose of this TDX boot-time initialization code is to provide
platform_tdx_enabled() function so that kexec() can use.
To distinguish boot-time TDX initialization from runtime TDX module
initialization, how about change the comment to below?
static void __init clear_tdx(void)
{
/*
* tdx_keyid_start and nr_tdx_keyids indicate that TDX is not
* enabled by the BIOS. This is used in TDX boot-time
* initializatiton error paths to take it from enabled to not
* enabled.
*/
tdx_keyid_start = nr_tdx_keyids = 0;
}
[...]
And below is the updated patch. How does it look to you?
==========================================================
x86/virt/tdx: Detect TDX during kernel boot
Intel Trust Domain Extensions (TDX) protects guest VMs from malicious
host and certain physical attacks. A CPU-attested software module
called 'the TDX module' runs inside a new isolated memory range as a
trusted hypervisor to manage and run protected VMs.
Pre-TDX Intel hardware has support for a memory encryption architecture
called MKTME. The memory encryption hardware underpinning MKTME is also
used for Intel TDX. TDX ends up "stealing" some of the physical address
space from the MKTME architecture for crypto-protection to VMs. The
BIOS is responsible for partitioning the "KeyID" space between legacy
MKTME and TDX. The KeyIDs reserved for TDX are called 'TDX private
KeyIDs' or 'TDX KeyIDs' for short.
TDX doesn't trust the BIOS. During machine boot, TDX verifies the TDX
private KeyIDs are consistently and correctly programmed by the BIOS
across all CPU packages before it enables TDX on any CPU core. A valid
TDX private KeyID range on BSP indicates TDX has been enabled by the
BIOS, otherwise the BIOS is buggy.
The TDX module is expected to be loaded by the BIOS when it enables TDX,
but the kernel needs to properly initialize it before it can be used to
create and run any TDX guests. The TDX module will be initialized by
the KVM subsystem when the KVM module is loaded.
Add a new early_initcall(tdx_init) to detect the TDX private KeyIDs.
Both TDX module initialization and creating TDX guest require to use TDX
private KeyID. Also add a function to report whether TDX is enabled by
the BIOS (TDX KeyID range is valid). Similar to AMD SME, kexec() will
use it to determine whether cache flush is needed.
To start to support TDX, create a new arch/x86/virt/vmx/tdx/tdx.c for
TDX host kernel support. Add a new Kconfig option CONFIG_INTEL_TDX_HOST
to opt-in TDX host kernel support (to distinguish with TDX guest kernel
support). So far only KVM uses TDX. Make the new config option depend
on KVM_INTEL.
Reviewed-by: Kirill A. Shutemov [off-list ref]
Signed-off-by: Kai Huang [off-list ref]
@@ -246,6 +246,8 @@ archheaders:libs-y+=arch/x86/lib/+core-y+=arch/x86/virt/+# drivers-y are linked after core-ydrivers-$(CONFIG_MATH_EMULATION)+=arch/x86/math-emu/drivers-$(CONFIG_PCI)+=arch/x86/pci/
@@ -0,0 +1,90 @@+// SPDX-License-Identifier: GPL-2.0+/*+*Copyright(c)2022IntelCorporation.+*+*IntelTrustedDomainExtensions(TDX)support+*/++#define pr_fmt(fmt) "tdx: " fmt++#include<linux/types.h>+#include<linux/cache.h>+#include<linux/init.h>+#include<linux/printk.h>+#include<asm/msr.h>+#include<asm/tdx.h>+#include"tdx.h"++staticu32tdx_keyid_start__ro_after_init;+staticu32nr_tdx_keyids__ro_after_init;++staticint__initrecord_keyid_partitioning(void)+{+u32nr_mktme_keyids;+intret;++/*+*IA32_MKTME_KEYID_PARTIONING:+*Bit[31:0]:NumberofMKTMEKeyIDs.+*Bit[63:32]:NumberofTDXprivateKeyIDs.+*/+ret=rdmsr_safe(MSR_IA32_MKTME_KEYID_PARTITIONING,&nr_mktme_keyids,+&nr_tdx_keyids);+if(ret)+return-ENODEV;++if(!nr_tdx_keyids)+return-ENODEV;++/* TDX KeyIDs start after the last MKTME KeyID. */+tdx_keyid_start++;++pr_info("enabled: private KeyID range [%u, %u)\n",+tdx_keyid_start,tdx_keyid_start+nr_tdx_keyids);++return0;+}++staticvoid__initclear_tdx(void)+{+/*+*tdx_keyid_startandnr_tdx_keyidsindicatethatTDXisnot+*enabledbytheBIOS.ThisisusedinTDXboot-time+*initializatitonerrorpathstotakeitfromenabledtonot+*enabled.+*/+tdx_keyid_start=nr_tdx_keyids=0;+}++staticint__inittdx_init(void)+{+interr;++err=record_keyid_partitioning();+if(err)+returnerr;++/*+*InitializingtheTDXmodulerequiresoneTDXprivateKeyID.+*Ifthere'sonlyoneTDXKeyIDthenaftermoduleinitialization+*KVMwon'tbeabletorunanyTDXguest,whichmakesthewhole+*thingworthless.JustdisableTDXinthiscase.+*/+if(nr_tdx_keyids<2){+pr_info("initialization failed: too few private KeyIDs available
From: Huang, Kai <hidden> Date: 2022-11-22 11:43:54
On Tue, 2022-11-22 at 11:10 +0100, Peter Zijlstra wrote:
On Mon, Nov 21, 2022 at 01:26:32PM +1300, Kai Huang wrote:
quoted
+static int build_tdx_memory(void)+{+ unsigned long start_pfn, end_pfn;+ int i, nid, ret;++ for_each_mem_pfn_range(i, MAX_NUMNODES, &start_pfn, &end_pfn, &nid) {+ /*+ * The first 1MB may not be reported as TDX convertible+ * memory. Manually exclude them as TDX memory.+ *+ * This is fine as the first 1MB is already reserved in+ * reserve_real_mode() and won't end up to ZONE_DMA as+ * free page anyway.+ */+ start_pfn = max(start_pfn, (unsigned long)SZ_1M >> PAGE_SHIFT);+ if (start_pfn >= end_pfn)+ continue;++ /* Verify memory is truly TDX convertible memory */+ if (!pfn_range_covered_by_cmr(start_pfn, end_pfn)) {+ pr_info("Memory region [0x%lx, 0x%lx) is not TDX convertible memorry.\n",+ start_pfn << PAGE_SHIFT,+ end_pfn << PAGE_SHIFT);+ return -EINVAL;
Given how tdx_cc_memory_compatible() below relies on tdx_memlist being
empty; this error patch is wrong and should goto err.
Oops. Thanks for catching.
Also thanks for review! Today is too late for me and I'll catch up with others
tomorrow.
quoted
+ }++ /*+ * Add the memory regions as TDX memory. The regions in+ * memblock has already guaranteed they are in address+ * ascending order and don't overlap.+ */+ ret = add_tdx_memblock(start_pfn, end_pfn, nid);+ if (ret)+ goto err;+ }++ return 0;+err:+ free_tdx_memory();+ return ret;+}
quoted
+bool tdx_cc_memory_compatible(unsigned long start_pfn, unsigned long end_pfn)+{+ struct tdx_memblock *tmb;++ /* Empty list means TDX isn't enabled successfully */+ if (list_empty(&tdx_memlist))+ return true;++ list_for_each_entry(tmb, &tdx_memlist, list) {+ /*+ * The new range is TDX memory if it is fully covered+ * by any TDX memory block.+ */+ if (start_pfn >= tmb->start_pfn && end_pfn <= tmb->end_pfn)+ return true;+ }+ return false;+}
From: Thomas Gleixner <hidden> Date: 2022-11-22 15:06:32
On Tue, Nov 22 2022 at 10:20, Peter Zijlstra wrote:
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
quoted
Shutting down the TDX module requires calling TDH.SYS.LP.SHUTDOWN on all
BIOS-enabled CPUs, and the SEMACALL can run concurrently on different
CPUs. Implement a mechanism to run SEAMCALL concurrently on all online
CPUs and use it to shut down the module. Later logical-cpu scope module
initialization will use it too.
Uhh, those requirements ^ are not met by this:
Can run concurrently != Must run concurrently
The documentation clearly says "can run concurrently" as quoted above.
Thanks,
tglx
From: Dave Hansen <hidden> Date: 2022-11-22 15:14:26
On 11/22/22 01:13, Peter Zijlstra wrote:
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
quoted
+/*+ * Call the SEAMCALL on all online CPUs concurrently. Caller to check+ * @sc->err to determine whether any SEAMCALL failed on any cpu.+ */+static void seamcall_on_each_cpu(struct seamcall_ctx *sc)+{+ on_each_cpu(seamcall_smp_call_function, sc, true);+}
Suppose the user has NOHZ_FULL configured, and is already running
userspace that will terminate on interrupt (this is desired feature for
NOHZ_FULL), guess how happy they'll be if someone, on another parition,
manages to tickle this TDX gunk?
Yeah, they'll be none too happy.
But, what do we do?
There are technical solutions like detecting if NOHZ_FULL is in play and
refusing to initialize TDX. There are also non-technical solutions like
telling folks in the documentation that they better modprobe kvm early
if they want to do TDX, or their NOHZ_FULL apps will pay.
We could also force the TDX module to be loaded early in boot before
NOHZ_FULL is in play, but that would waste memory on TDX metadata even
if TDX is never used.
How do NOHZ_FULL folks deal with late microcode updates, for example?
Those are roughly equally disruptive to all CPUs.
From: Dave Hansen <hidden> Date: 2022-11-22 15:21:07
On 11/22/22 01:20, Peter Zijlstra wrote:
Either the Changelog is broken or this TDX crud is worse crap than I
thought possible, because the only way to actually meet that requirement
as stated is stop_machine().
I think the changelog is broken. I don't see anything in the TDX module
spec about "the SEMACALL can run concurrently on different CPUs".
Shutdown, as far as I can tell, just requires that the shutdown seamcall
be run once on each CPU. Concurrency and ordering don't seem to matter
at all.
From: Dave Hansen <hidden> Date: 2022-11-22 15:36:13
On 11/22/22 02:31, Thomas Gleixner wrote:
Nothing in the TDX specs and docs mentions physical hotplug or a
requirement for invoking seamcall on the world.
The TDX module source is actually out there[1] for us to look at. It's
in a lovely, convenient zip file, but you can read it if sufficiently
motivated.
It has this lovely nugget in it:
WARNING!!! Proprietary License!! Avert your virgin eyes!!!
if (tdx_global_data_ptr->num_of_init_lps < tdx_global_data_ptr->num_of_lps)
{
TDX_ERROR("Num of initialized lps %d is smaller than total num of lps %d\n",
tdx_global_data_ptr->num_of_init_lps, tdx_global_data_ptr->num_of_lps);
retval = TDX_SYS_CONFIG_NOT_PENDING;
goto EXIT;
}
From: Dave Hansen <hidden> Date: 2022-11-22 16:50:45
On 11/22/22 03:28, Huang, Kai wrote:
quoted
quoted
+ /*+ * KeyID 0 is for TME. MKTME KeyIDs start from 1. TDX private+ * KeyIDs start after the last MKTME KeyID.+ */
Is the TME key a "MKTME KeyID"?
I don't think so. Hardware handles TME KeyID 0 differently from non-0 MKTME
KeyIDs. And PCONFIG only accept non-0 KeyIDs.
Let's say we have 4 MKTME hardware bits, we'd have:
0: TME Key
1->3: MKTME Keys
4->7: TDX Private Keys
First, the MSR values:
quoted hunk
+ * IA32_MKTME_KEYID_PARTIONING:+ * Bit [31:0]: Number of MKTME KeyIDs.+ * Bit [63:32]: Number of TDX private KeyIDs.
These would be:
Bit [ 31:0] = 3
Bit [63:22] = 4
And in the end the variables:
tdx_keyid_start would be 4 and tdx_keyid_num would be 4.
Right?
That's a bit wonky for my brain because I guess I know too much about
the internal implementation and how the key space is split up. I guess
I (wrongly) expected Bit[31:0]==Bit[63:22].
This is where a comment is needed and can actually help.
/*
* tdx_keyid_start/num indicate that TDX is uninitialized. This
* is used in TDX initialization error paths to take it from
* initialized -> uninitialized.
*/
Just want to point out after removing the !x2apic_enabled() check, the only
thing need to do here is to detect/record the TDX KeyIDs.
And the purpose of this TDX boot-time initialization code is to provide
platform_tdx_enabled() function so that kexec() can use.
To distinguish boot-time TDX initialization from runtime TDX module
initialization, how about change the comment to below?
static void __init clear_tdx(void)
{
/*
* tdx_keyid_start and nr_tdx_keyids indicate that TDX is not
* enabled by the BIOS. This is used in TDX boot-time
* initializatiton error paths to take it from enabled to not
* enabled.
*/
tdx_keyid_start = nr_tdx_keyids = 0;
}
[...]
I honestly have no idea what "boot-time TDX initialization" is versus
"runtime TDX module initialization". This doesn't hel.
And below is the updated patch. How does it look to you?
Let's see...
...
quoted hunk
+static u32 tdx_keyid_start __ro_after_init;+static u32 nr_tdx_keyids __ro_after_init;++static int __init record_keyid_partitioning(void)+{+ u32 nr_mktme_keyids;+ int ret;++ /*+ * IA32_MKTME_KEYID_PARTIONING:+ * Bit [31:0]: Number of MKTME KeyIDs.+ * Bit [63:32]: Number of TDX private KeyIDs.+ */+ ret = rdmsr_safe(MSR_IA32_MKTME_KEYID_PARTITIONING, &nr_mktme_keyids,+ &nr_tdx_keyids);+ if (ret)+ return -ENODEV;++ if (!nr_tdx_keyids)+ return -ENODEV;++ /* TDX KeyIDs start after the last MKTME KeyID. */+ tdx_keyid_start++;
tdx_keyid_start is uniniitalized here. So, it'd be 0, then ++'d.
Kai, please take a moment and slow down. This isn't a race. I offered
some replacement code here, which you've discarded, missed or ignored
and in the process broken this code.
This approach just wastes reviewer time. It's not working for me.
I'm going to make a suggestion (aka. a demand): You can post these
patches at most once a week. You get a whole week to (carefully)
incorporate reviewer feedback, make the patch better, and post a new
version. Need more time? Go ahead and take it. Take as much time as
you want.
From: Thomas Gleixner <hidden> Date: 2022-11-22 16:52:27
On Tue, Nov 22 2022 at 07:20, Dave Hansen wrote:
On 11/22/22 01:20, Peter Zijlstra wrote:
quoted
Either the Changelog is broken or this TDX crud is worse crap than I
thought possible, because the only way to actually meet that requirement
as stated is stop_machine().
I think the changelog is broken. I don't see anything in the TDX module
spec about "the SEMACALL can run concurrently on different CPUs".
Shutdown, as far as I can tell, just requires that the shutdown seamcall
be run once on each CPU. Concurrency and ordering don't seem to matter
at all.
You're right. The 'can concurrently run' thing is for LP.INIT:
4.2.2.
LP-Scope Initialization: TDH.SYS.LP.INIT
TDH.SYS.LP.INIT is intended to perform LP-scope, core-scope and
package-scope initialization of the Intel TDX module. It can be called
only after TDH.SYS.INIT completes successfully, and it can run
concurrently on multiple LPs.
From: Dave Hansen <hidden> Date: 2022-11-22 18:06:09
On 11/20/22 16:26, Kai Huang wrote:
2) It is more flexible to support TDX module runtime updating in the
future (after updating the TDX module, it needs to be initialized
again).
I hate this generic blabber about "more flexible". There's a *REASON*
it's more flexible, so let's talk about the reasons, please.
It's really something like this, right?
The TDX module design allows it to be updated while the system
is running. The update procedure shares quite a few steps with
this "on demand" loading mechanism. The hope is that much of
this "on demand" mechanism can be shared with a future "update"
mechanism. A boot-time TDX module implementation would not be
able to share much code with the update mechanism.
3) It avoids having to do a "temporary" solution to handle VMXON in the
core (non-KVM) kernel for now. This is because SEAMCALL requires CPU
being in VMX operation (VMXON is done), but currently only KVM handles
VMXON. Adding VMXON support to the core kernel isn't trivial. More
importantly, from long-term a reference-based approach is likely needed
in the core kernel as more kernel components are likely needed to
support TDX as well. Allow KVM to initialize the TDX module avoids
having to handle VMXON during kernel boot for now.
There are a lot of words in there.
3) Loading the TDX module requires VMX to be enabled. Currently, only
the kernel KVM code mucks with VMX enabling. If the TDX module were
to be initialized separately from KVM (like at boot), the boot code
would need to be taught how to muck with VMX enabling and KVM would
need to be taught how to cope with that. Making KVM itself
responsible for TDX initialization lets the rest of the kernel stay
blissfully unaware of VMX.
Add a placeholder tdx_enable() to detect and initialize the TDX module
on demand, with a state machine protected by mutex to support concurrent
calls from multiple callers.
As opposed to concurrent calls from one caller? ;)
The TDX module will be initialized in multi-steps defined by the TDX
module:
1) Global initialization;
2) Logical-CPU scope initialization;
3) Enumerate the TDX module capabilities and platform configuration;
4) Configure the TDX module about TDX usable memory ranges and global
KeyID information;
5) Package-scope configuration for the global KeyID;
6) Initialize usable memory ranges based on 4).
This would actually be a nice place to call out the SEAMCALL names and
mention that each of these steps involves a set of SEAMCALLs.
The TDX module can also be shut down at any time during its lifetime.
In case of any error during the initialization process, shut down the
module. It's pointless to leave the module in any intermediate state
during the initialization.
Both logical CPU scope initialization and shutting down the TDX module
require calling SEAMCALL on all boot-time present CPUs. For simplicity
just temporarily disable CPU hotplug during the module initialization.
You might want to more precisely define "boot-time present CPUs". The
boot of *what*?
@@ -10,15 +10,34 @@#include<linux/types.h>#include<linux/init.h>#include<linux/printk.h>+#include<linux/mutex.h>+#include<linux/cpu.h>+#include<linux/cpumask.h>#include<asm/msr-index.h>#include<asm/msr.h>#include<asm/apic.h>#include<asm/tdx.h>#include"tdx.h"+/* TDX module status during initialization */+enumtdx_module_status_t{+/* TDX module hasn't been detected and initialized */+TDX_MODULE_UNKNOWN,+/* TDX module is not loaded */+TDX_MODULE_NONE,+/* TDX module is initialized */+TDX_MODULE_INITIALIZED,+/* TDX module is shut down due to initialization error */+TDX_MODULE_SHUTDOWN,+};
Are these part of the ABI or just a purely OS-side construct?
quoted hunk
static u32 tdx_keyid_start __ro_after_init; static u32 tdx_keyid_num __ro_after_init;+static enum tdx_module_status_t tdx_module_status;+/* Prevent concurrent attempts on TDX detection and initialization */+static DEFINE_MUTEX(tdx_module_lock);+ /* * Detect TDX private KeyIDs to see whether TDX has been enabled by the * BIOS. Both initializing the TDX module and running TDX guest require
@@ -104,3 +123,134 @@ bool platform_tdx_enabled(void) { return !!tdx_keyid_num; }++/*+ * Detect and initialize the TDX module.+ *+ * Return -ENODEV when the TDX module is not loaded, 0 when it+ * is successfully initialized, or other error when it fails to+ * initialize.+ */+static int init_tdx_module(void)+{+ /* The TDX module hasn't been detected */+ return -ENODEV;+}++static void shutdown_tdx_module(void)+{+ /* TODO: Shut down the TDX module */+}++static int __tdx_enable(void)+{+ int ret;++ /*+ * Initializing the TDX module requires doing SEAMCALL on all+ * boot-time present CPUs. For simplicity temporarily disable+ * CPU hotplug to prevent any CPU from going offline during+ * the initialization.+ */+ cpus_read_lock();++ /*+ * Check whether all boot-time present CPUs are online and+ * return early with a message so the user can be aware.+ *+ * Note a non-buggy BIOS should never support physical (ACPI)+ * CPU hotplug when TDX is enabled, and all boot-time present+ * CPU should be enabled in MADT, so there should be no+ * disabled_cpus and num_processors won't change at runtime+ * either.+ */
Again, there are a lot of words in that comment, but I'm not sure why
it's here. Despite all the whinging about ACPI, doesn't it boil down to:
The TDX module itself establishes its own concept of how many
logical CPUs there are in the system when it is loaded. The
module will reject initialization attempts unless the kernel
runs TDX initialization code on every last CPU.
Ensure that the kernel is able to run code on all known logical
CPUs.
and these checks are just to see if the kernel has shot itself in the
foot and is *KNOWS* that it is currently unable to run code on some
logical CPU?
quoted hunk
+ if (disabled_cpus || num_online_cpus() != num_processors) {+ pr_err("Unable to initialize the TDX module when there's offline CPU(s).\n");+ ret = -EINVAL;+ goto out;+ }++ ret = init_tdx_module();+ if (ret == -ENODEV) {
Why check for -ENODEV exclusively? Is there some other error nonzero
code that indicates success?
quoted hunk
+ pr_info("TDX module is not loaded.\n");+ tdx_module_status = TDX_MODULE_NONE;+ goto out;+ }++ /*+ * Shut down the TDX module in case of any error during the+ * initialization process. It's meaningless to leave the TDX+ * module in any middle state of the initialization process.+ *+ * Shutting down the module also requires doing SEAMCALL on all+ * MADT-enabled CPUs. Do it while CPU hotplug is disabled.+ *+ * Return all errors during the initialization as -EFAULT as the+ * module is always shut down.+ */+ if (ret) {+ pr_info("Failed to initialize TDX module. Shut it down.\n");
"Shut it down" seems wrong here. That could be interpreted as "I have
already shut it down". "Shutting down" seems better.
quoted hunk
+ shutdown_tdx_module();+ tdx_module_status = TDX_MODULE_SHUTDOWN;+ ret = -EFAULT;+ goto out;+ }++ pr_info("TDX module initialized.\n");+ tdx_module_status = TDX_MODULE_INITIALIZED;+out:+ cpus_read_unlock();++ return ret;+}++/**+ * tdx_enable - Enable TDX by initializing the TDX module+ *+ * Caller to make sure all CPUs are online and in VMX operation before+ * calling this function. CPU hotplug is temporarily disabled internally+ * to prevent any cpu from going offline.
"cpu" or "CPU"?
quoted hunk
+ * This function can be called in parallel by multiple callers.+ *+ * Return:+ *+ * * 0: The TDX module has been successfully initialized.+ * * -ENODEV: The TDX module is not loaded, or TDX is not supported.+ * * -EINVAL: The TDX module cannot be initialized due to certain+ * conditions are not met (i.e. when not all MADT-enabled+ * CPUs are not online).+ * * -EFAULT: Other internal fatal errors, or the TDX module is in+ * shutdown mode due to it failed to initialize in previous+ * attempts.+ */
I honestly don't think all these error codes mean anything. They're
plumbed nowhere and the use of -EFAULT is just plain wrong.
Nobody can *DO* anything with these anyway.
Just give one error code and make sure that you have pr_info()'s around
to make it clear what went wrong. Then just do -EINVAL universally.
Remove all the nonsense comments.
quoted hunk
+int tdx_enable(void)+{+ int ret;++ if (!platform_tdx_enabled())+ return -ENODEV;++ mutex_lock(&tdx_module_lock);++ switch (tdx_module_status) {+ case TDX_MODULE_UNKNOWN:+ ret = __tdx_enable();+ break;+ case TDX_MODULE_NONE:+ ret = -ENODEV;+ break;
TDX_MODULE_NONE should probably be called TDX_MODULE_NOT_LOADED. A
comment would also be nice:
/* The BIOS did not load the module. No way to fix that. */
+ case TDX_MODULE_INITIALIZED:
/* Already initialized, great, tell the caller: */
quoted hunk
+ ret = 0;+ break;+ default:+ WARN_ON_ONCE(tdx_module_status != TDX_MODULE_SHUTDOWN);+ ret = -EFAULT;+ break;+ }
I don't get what that default: is for or what it has to do with
TDX_MODULE_SHUTDOWN.
From: Dave Hansen <hidden> Date: 2022-11-22 18:20:52
On 11/20/22 16:26, Kai Huang wrote:
TDX introduces a new CPU mode: Secure Arbitration Mode (SEAM). This
mode runs only the TDX module itself or other code to load the TDX
module.
The host kernel communicates with SEAM software via a new SEAMCALL
instruction. This is conceptually similar to a guest->host hypercall,
except it is made from the host to SEAM software instead.
The TDX module defines a set of SEAMCALL leaf functions to allow the
host to initialize it, and to create and run protected VMs. SEAMCALL
leaf functions use an ABI different from the x86-64 system-v ABI.
Instead, they share the same ABI with the TDCALL leaf functions.
I may have suggested this along the way, but the mention of the sysv ABI
is just confusing here. This is enough for a changelog:
The TDX module establishes a new SEAMCALL ABI which allows the
host to initialize the module and to and to manage VMs.
Kill the rest.
Implement a function __seamcall() to allow the host to make SEAMCALL
to SEAM software using the TDX_MODULE_CALL macro which is the common
assembly for both SEAMCALL and TDCALL.
In general, I dislike mentioning function names in changelogs. Keep
this high-level, like:
Add infrastructure to make SEAMCALLs. The SEAMCALL ABI is very
similar to the TDCALL ABI and leverages much TDCALL
infrastructure.
SEAMCALL instruction causes #GP when SEAMRR isn't enabled, and #UD when
CPU is not in VMX operation. The current TDX_MODULE_CALL macro doesn't
handle any of them. There's no way to check whether the CPU is in VMX
operation or not.
What is SEAMRR?
Why even mention this behavior in the changelog. Is this a problem?
Does it have a solution?
Initializing the TDX module is done at runtime on demand, and it depends
on the caller to ensure CPU is in VMX operation before making SEAMCALL.
To avoid getting Oops when the caller mistakenly tries to initialize the
TDX module when CPU is not in VMX operation, extend the TDX_MODULE_CALL
macro to handle #UD (and also #GP, which can theoretically still happen
when TDX isn't actually enabled by the BIOS, i.e. due to BIOS bug).
I'm not completely sure this is worth it. If the BIOS lies, we oops.
There are lots of ways that the BIOS lying can make the kernel oops.
What's one more?
Introduce two new TDX error codes for #UD and #GP respectively so the
caller can distinguish. Also, Opportunistically put the new TDX error
codes and the existing TDX_SEAMCALL_VMFAILINVALID into INTEL_TDX_HOST
Kconfig option as they are only used when it is on.
As __seamcall() can potentially return multiple error codes, besides the
actual SEAMCALL leaf function return code, also introduce a wrapper
function seamcall() to convert the __seamcall() error code to the kernel
error code, so the caller doesn't need to duplicate the code to check
return value of __seamcall() and return kernel error code accordingly.
@@ -124,6 +124,48 @@ bool platform_tdx_enabled(void)return!!tdx_keyid_num;}+/*+*Wrapperof__seamcall()toconvertSEAMCALLleaffunctionerrorcode+*tokernelerrorcode.@seamcall_retand@outcontaintheSEAMCALL+*leaffunctionreturncodeandtheadditionaloutputrespectivelyif+*notNULL.+*/+staticint__always_unusedseamcall(u64fn,u64rcx,u64rdx,u64r8,u64r9,+u64*seamcall_ret,+structtdx_module_output*out)+{+u64sret;++sret=__seamcall(fn,rcx,rdx,r8,r9,out);++/* Save SEAMCALL return code if caller wants it */+if(seamcall_ret)+*seamcall_ret=sret;++/* SEAMCALL was successful */+if(!sret)+return0;++switch(sret){+caseTDX_SEAMCALL_GP:+/*+*platform_tdx_enabled()ischeckedtobetrue+*beforemakinganySEAMCALL.+*/
This doesn't make any sense. "platform_tdx_enabled() is checked"???
Do you mean that it *should* be checked and probably wasn't which is
what caused the error?
quoted hunk
+ WARN_ON_ONCE(1);+ fallthrough;+ case TDX_SEAMCALL_VMFAILINVALID:+ /* Return -ENODEV if the TDX module is not loaded. */+ return -ENODEV;
Pro tip: you don't need to rewrite code in comments. If the code
literally says, "return -ENODEV", there is very little value in writing
virtually identical bytes "Return -ENODEV" in the comment.
From: Dave Hansen <hidden> Date: 2022-11-22 18:58:51
On 11/20/22 16:26, Kai Huang wrote:
TDX supports shutting down the TDX module at any time during its
lifetime. After the module is shut down, no further TDX module SEAMCALL
leaf functions can be made to the module on any logical cpu.
Shut down the TDX module in case of any error during the initialization
process. It's pointless to leave the TDX module in some middle state.
Shutting down the TDX module requires calling TDH.SYS.LP.SHUTDOWN on all
BIOS-enabled CPUs, and the SEMACALL can run concurrently on different
CPUs. Implement a mechanism to run SEAMCALL concurrently on all online
CPUs and use it to shut down the module. Later logical-cpu scope module
initialization will use it too.
To me, this starts to veer way too far into internal implementation details.
Issue the TDH.SYS.LP.SHUTDOWN SEAMCALL on all BIOS-enabled CPUs
to shut down the TDX module.
This is also the point where you should talk about the new
infrastructure. Why do you need a new 'struct seamcall_something'?
What makes it special?
The atomic_t is kinda silly. I guess it's not *that* wasteful though.
I think it would have actually been a lot more clear if instead of
containing an errno it was a *count* of the number of encountered errors.
An "atomic_set()" where everyone is overwriting each other is a bit
counterintuitive. It's OK here, of course, but it still looks goofy.
If this were:
atomic_inc(&sc->nr_errors);
it would be a lot more clear that *anyone* can increment and that it
truly is shared.
quoted hunk
+/*+ * Call the SEAMCALL on all online CPUs concurrently. Caller to check+ * @sc->err to determine whether any SEAMCALL failed on any cpu.+ */+static void seamcall_on_each_cpu(struct seamcall_ctx *sc)+{+ on_each_cpu(seamcall_smp_call_function, sc, true);+}+ /* * Detect and initialize the TDX module. *
@@ -181,7 +214,9 @@ static int init_tdx_module(void) static void shutdown_tdx_module(void) {- /* TODO: Shut down the TDX module */+ struct seamcall_ctx sc = { .fn = TDH_SYS_LP_SHUTDOWN };++ seamcall_on_each_cpu(&sc); }
The seamcall_on_each_cpu() function is silly as-is. Either collapse the
functions or note in the changelog why this is not as silly as it looks.
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 19:07:28
On Tue, Nov 22, 2022 at 04:06:25PM +0100, Thomas Gleixner wrote:
On Tue, Nov 22 2022 at 10:20, Peter Zijlstra wrote:
quoted
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
quoted
Shutting down the TDX module requires calling TDH.SYS.LP.SHUTDOWN on all
BIOS-enabled CPUs, and the SEMACALL can run concurrently on different
CPUs. Implement a mechanism to run SEAMCALL concurrently on all online
CPUs and use it to shut down the module. Later logical-cpu scope module
initialization will use it too.
Uhh, those requirements ^ are not met by this:
Can run concurrently != Must run concurrently
The documentation clearly says "can run concurrently" as quoted above.
The next sentense says: "Implement a mechanism to run SEAMCALL
concurrently" -- it does not.
Anyway, since we're all in agreement there is no such requirement at
all, a schedule_on_each_cpu() might be more appropriate, there is no
reason to use IPIs and spin-waiting for any of this.
That said; perhaps we should grow:
schedule_on_cpu(struct cpumask *cpus, work_func_t func);
to only disturb a given mask of CPUs.
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 19:13:43
On Tue, Nov 22, 2022 at 07:14:14AM -0800, Dave Hansen wrote:
On 11/22/22 01:13, Peter Zijlstra wrote:
quoted
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
quoted
+/*+ * Call the SEAMCALL on all online CPUs concurrently. Caller to check+ * @sc->err to determine whether any SEAMCALL failed on any cpu.+ */+static void seamcall_on_each_cpu(struct seamcall_ctx *sc)+{+ on_each_cpu(seamcall_smp_call_function, sc, true);+}
Suppose the user has NOHZ_FULL configured, and is already running
userspace that will terminate on interrupt (this is desired feature for
NOHZ_FULL), guess how happy they'll be if someone, on another parition,
manages to tickle this TDX gunk?
Yeah, they'll be none too happy.
But, what do we do?
Not intialize TDX on busy NOHZ_FULL cpus and hard-limit the cpumask of
all TDX using tasks.
There are technical solutions like detecting if NOHZ_FULL is in play and
refusing to initialize TDX. There are also non-technical solutions like
telling folks in the documentation that they better modprobe kvm early
if they want to do TDX, or their NOHZ_FULL apps will pay.
Surely modprobe kvm isn't the point where TDX gets loaded? Because
that's on boot for everybody due to all the auto-probing nonsense.
I was expecting TDX to not get initialized until the first TDX using KVM
instance is created. Am I wrong?
We could also force the TDX module to be loaded early in boot before
NOHZ_FULL is in play, but that would waste memory on TDX metadata even
if TDX is never used.
I'm thikning it makes sense to have a tdx={off,on-demand,force} toggle
anyway.
How do NOHZ_FULL folks deal with late microcode updates, for example?
Those are roughly equally disruptive to all CPUs.
I imagine they don't do that -- in fact I would recommend we make the
whole late loading thing mutually exclusive with nohz_full; can't have
both.
From: Dave Hansen <hidden> Date: 2022-11-22 19:14:24
On 11/20/22 16:26, Kai Huang wrote:
The first step of initializing the module is to call TDH.SYS.INIT once
on any logical cpu to do module global initialization. Do the module
global initialization.
It also detects the TDX module, as seamcall() returns -ENODEV when the
module is not loaded.
Part of making a good patch set is telling a bit of a story. In patch
4, you laid out 6 steps necessary to initialize TDX. On top of that,
there is infrastructure It would be great to lay that out in a way that
folks can actually follow along.
For instance, it would be great to tell the reader here that this patch
is an inflection point. It is transitioning out of the infrastructure
(patches 1->6) and into the actual "multi-steps" of initialization that
the module spec requires.
This patch is *TOTALLY* different from the one before it because it
actually _starts_ to do something useful.
But, you wouldn't know it from the changelog.
@@ -208,8 +208,23 @@ static void seamcall_on_each_cpu(struct seamcall_ctx *sc)*/staticintinit_tdx_module(void){-/* The TDX module hasn't been detected */-return-ENODEV;+intret;++/*+*CallTDH.SYS.INITtodotheglobalinitializationof+*theTDXmodule.Italsodetectsthemodule.+*/+ret=seamcall(TDH_SYS_INIT,0,0,0,0,NULL,NULL);+if(ret)+gotoout;
Please also note that the 0's are all just unused parameters. They mean
nothing.
quoted hunk
++ /*+ * Return -EINVAL until all steps of TDX module initialization+ * process are done.+ */+ ret = -EINVAL;+out:+ return ret; }
It might be a bit unconventional, but can you imagine how well it would
tell the story if this comment said:
/*
* TODO:
* - Logical-CPU scope initialization (TDH_SYS_INIT_LP)
* - Enumerate capabilities and platform configuration
(TDH_SYS_CONFIG)
...
*/
and then each of the following patches that *did* those things removed
the TODO line from the list.
That TODO list could have been added in patch 4.
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 19:14:44
On Tue, Nov 22, 2022 at 10:57:52AM -0800, Dave Hansen wrote:
To me, this starts to veer way too far into internal implementation details.
Issue the TDH.SYS.LP.SHUTDOWN SEAMCALL on all BIOS-enabled CPUs
to shut down the TDX module.
We really need to let go of the whole 'all BIOS-enabled CPUs' thing.
From: Dave Hansen <hidden> Date: 2022-11-22 19:25:00
On 11/22/22 11:13, Peter Zijlstra wrote:
On Tue, Nov 22, 2022 at 07:14:14AM -0800, Dave Hansen wrote:
quoted
On 11/22/22 01:13, Peter Zijlstra wrote:
quoted
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
quoted
+/*+ * Call the SEAMCALL on all online CPUs concurrently. Caller to check+ * @sc->err to determine whether any SEAMCALL failed on any cpu.+ */+static void seamcall_on_each_cpu(struct seamcall_ctx *sc)+{+ on_each_cpu(seamcall_smp_call_function, sc, true);+}
Suppose the user has NOHZ_FULL configured, and is already running
userspace that will terminate on interrupt (this is desired feature for
NOHZ_FULL), guess how happy they'll be if someone, on another parition,
manages to tickle this TDX gunk?
Yeah, they'll be none too happy.
But, what do we do?
Not intialize TDX on busy NOHZ_FULL cpus and hard-limit the cpumask of
all TDX using tasks.
I don't think that works. As I mentioned to Thomas elsewhere, you don't
just need to initialize TDX on the CPUs where it is used. Before the
module will start working you need to initialize it on *all* the CPUs it
knows about. The module itself has a little counter where it tracks
this and will refuse to start being useful until it gets called
thoroughly enough.
quoted
There are technical solutions like detecting if NOHZ_FULL is in play and
refusing to initialize TDX. There are also non-technical solutions like
telling folks in the documentation that they better modprobe kvm early
if they want to do TDX, or their NOHZ_FULL apps will pay.
Surely modprobe kvm isn't the point where TDX gets loaded? Because
that's on boot for everybody due to all the auto-probing nonsense.
I was expecting TDX to not get initialized until the first TDX using KVM
instance is created. Am I wrong?
I went looking for it in this series to prove you wrong. I failed. :)
tdx_enable() is buried in here somewhere:
I don't have the patience to dig it out today, so I guess we'll have Kai
tell us.
quoted
We could also force the TDX module to be loaded early in boot before
NOHZ_FULL is in play, but that would waste memory on TDX metadata even
if TDX is never used.
I'm thikning it makes sense to have a tdx={off,on-demand,force} toggle
anyway.
Yep, that makes total sense. Kai had one in an earlier version but I
made him throw it out because it wasn't *strictly* required and this set
is fat enough.
quoted
How do NOHZ_FULL folks deal with late microcode updates, for example?
Those are roughly equally disruptive to all CPUs.
I imagine they don't do that -- in fact I would recommend we make the
whole late loading thing mutually exclusive with nohz_full; can't have
both.
So, if we just use schedule_on_cpu() for now and have the TDX code wait,
will a NOHZ_FULL task just block the schedule_on_cpu() indefinitely?
That doesn't seem like _horrible_ behavior to start off with for a
minimal series.
From: Sean Christopherson <seanjc@google.com> Date: 2022-11-22 19:32:51
On Tue, Nov 22, 2022, Peter Zijlstra wrote:
On Tue, Nov 22, 2022 at 04:06:25PM +0100, Thomas Gleixner wrote:
quoted
On Tue, Nov 22 2022 at 10:20, Peter Zijlstra wrote:
quoted
On Mon, Nov 21, 2022 at 01:26:28PM +1300, Kai Huang wrote:
quoted
Shutting down the TDX module requires calling TDH.SYS.LP.SHUTDOWN on all
BIOS-enabled CPUs, and the SEMACALL can run concurrently on different
CPUs. Implement a mechanism to run SEAMCALL concurrently on all online
CPUs and use it to shut down the module. Later logical-cpu scope module
initialization will use it too.
Uhh, those requirements ^ are not met by this:
Can run concurrently != Must run concurrently
The documentation clearly says "can run concurrently" as quoted above.
The next sentense says: "Implement a mechanism to run SEAMCALL
concurrently" -- it does not.
Anyway, since we're all in agreement there is no such requirement at
all, a schedule_on_each_cpu() might be more appropriate, there is no
reason to use IPIs and spin-waiting for any of this.
Backing up a bit, what's the reason for _any_ of this? The changelog says
It's pointless to leave the TDX module in some middle state.
but IMO it's just as pointless to do a shutdown unless the kernel benefits in
some meaningful way. And IIUC, TDH.SYS.LP.SHUTDOWN does nothing more than change
the SEAM VMCS.HOST_RIP to point to an error trampoline. E.g. it's not like doing
a shutdown lets the kernel reclaim memory that was gifted to the TDX module.
In other words, this is just a really expensive way of changing a function pointer,
and the only way it would ever benefit the kernel is if there is a kernel bug that
leads to trying to use TDX after a fatal error. And even then, the only difference
seems to be that subsequent bogus SEAMCALLs would get a more unique error message.
From: Peter Zijlstra <peterz@infradead.org> Date: 2022-11-22 19:33:47
On Tue, Nov 22, 2022 at 11:24:48AM -0800, Dave Hansen wrote:
quoted
Not intialize TDX on busy NOHZ_FULL cpus and hard-limit the cpumask of
all TDX using tasks.
I don't think that works. As I mentioned to Thomas elsewhere, you don't
just need to initialize TDX on the CPUs where it is used. Before the
module will start working you need to initialize it on *all* the CPUs it
knows about. The module itself has a little counter where it tracks
this and will refuse to start being useful until it gets called
thoroughly enough.
That's bloody terrible, that is. How are we going to make that work with
the SMT mitigation crud that forces the SMT sibilng offline?
Then the counters don't match and TDX won't work.
Can we get this limitiation removed and simply let the module throw a
wobbly (error) when someone tries and use TDX without that logical CPU
having been properly initialized?
From: Thomas Gleixner <hidden> Date: 2022-11-22 20:03:47
On Tue, Nov 22 2022 at 07:35, Dave Hansen wrote:
On 11/22/22 02:31, Thomas Gleixner wrote:
quoted
Nothing in the TDX specs and docs mentions physical hotplug or a
requirement for invoking seamcall on the world.
The TDX module source is actually out there[1] for us to look at. It's
in a lovely, convenient zip file, but you can read it if sufficiently
motivated.
zip file? Version control from the last millenium?
The whole thing wants to be @github with a proper change history if
Intel wants anyone to trust this and take it serious.
/me refrains from ranting about the outrageous license choice.
It has this lovely nugget in it:
WARNING!!! Proprietary License!! Avert your virgin eyes!!!
It's probably not the only reasons to avert the eyes.
quoted
if (tdx_global_data_ptr->num_of_init_lps < tdx_global_data_ptr->num_of_lps)
{
TDX_ERROR("Num of initialized lps %d is smaller than total num of lps %d\n",
tdx_global_data_ptr->num_of_init_lps, tdx_global_data_ptr->num_of_lps);
retval = TDX_SYS_CONFIG_NOT_PENDING;
goto EXIT;
}
tdx_global_data_ptr->num_of_init_lps is incremented at TDH.SYS.INIT
time. That if() is called at TDH.SYS.CONFIG time to help bring the
module up.
So, I think you're right. I don't see the docs that actually *explain*
this "you must seamcall all the things" requirement.
The code actually enforces this.
At TDH.SYS.INIT which is the first operation it gets the total number
of LPs from the sysinfo table:
src/vmm_dispatcher/api_calls/tdh_sys_init.c:
tdx_global_data_ptr->num_of_lps = sysinfo_table_ptr->mcheck_fields.tot_num_lps;
Then TDH.SYS.LP.INIT increments the count of initialized LPs.
src/vmm_dispatcher/api_calls/tdh_sys_lp_init.c:
increment_num_of_lps(tdx_global_data_ptr)
_lock_xadd_32b(&tdx_global_data_ptr->num_of_init_lps, 1);
Finally TDH.SYS.CONFIG checks whether _ALL_ LPs have been initialized.
src/vmm_dispatcher/api_calls/tdh_sys_config.c:
if (tdx_global_data_ptr->num_of_init_lps < tdx_global_data_ptr->num_of_lps)
Clearly that's nowhere spelled out in the documentation, but I don't
buy the 'architecturaly required' argument not at all. It's an
implementation detail of the TDX module.
Technically there is IMO ZERO requirement to do so.
1) The TDX module is global
2) Seam-root and Seam-non-root operation are strictly a LP property.
The only architectural prerequisite for using Seam on a LP is that
obviously the encryption/decryption mechanics have been initialized
on the package to which the LP belongs.
I can see why it might be complicated to add/remove an LP after
initialization fact, but technically it should be possible.
TDX/Seam is not that special.
But what's absolutely annoying is that the documentation lacks any
information about the choice of enforcement which has been hardcoded
into the Seam module for whatever reasons.
Maybe I overlooked it, but then it's definitely well hidden.
Thanks,
tglx
From: Sean Christopherson <seanjc@google.com> Date: 2022-11-22 20:11:26
On Tue, Nov 22, 2022, Thomas Gleixner wrote:
On Tue, Nov 22 2022 at 07:35, Dave Hansen wrote:
quoted
On 11/22/22 02:31, Thomas Gleixner wrote:
quoted
Nothing in the TDX specs and docs mentions physical hotplug or a
requirement for invoking seamcall on the world.
The TDX module source is actually out there[1] for us to look at. It's
in a lovely, convenient zip file, but you can read it if sufficiently
motivated.
zip file? Version control from the last millenium?
The whole thing wants to be @github with a proper change history if
Intel wants anyone to trust this and take it serious.
/me refrains from ranting about the outrageous license choice.
Let me know if you grab pitchforcks and torches, I'll join the mob :-)
From: Huang, Kai <hidden> Date: 2022-11-22 23:21:51
On Tue, 2022-11-22 at 08:50 -0800, Dave Hansen wrote:
On 11/22/22 03:28, Huang, Kai wrote:
quoted
quoted
quoted
+ /*+ * KeyID 0 is for TME. MKTME KeyIDs start from 1. TDX private+ * KeyIDs start after the last MKTME KeyID.+ */
Is the TME key a "MKTME KeyID"?
I don't think so. Hardware handles TME KeyID 0 differently from non-0 MKTME
KeyIDs. And PCONFIG only accept non-0 KeyIDs.
Let's say we have 4 MKTME hardware bits, we'd have:
0: TME Key
1->3: MKTME Keys
4->7: TDX Private Keys
First, the MSR values:
quoted
+ * IA32_MKTME_KEYID_PARTIONING:+ * Bit [31:0]: Number of MKTME KeyIDs.+ * Bit [63:32]: Number of TDX private KeyIDs.
These would be:
Bit [ 31:0] = 3
Bit [63:22] = 4
And in the end the variables:
tdx_keyid_start would be 4 and tdx_keyid_num would be 4.
Right?
Yes.
That's a bit wonky for my brain because I guess I know too much about
the internal implementation and how the key space is split up. I guess
I (wrongly) expected Bit[31:0]==Bit[63:22].
The spec says the The Bit[31:0] only reports the number of MKTME KeyIDs, and it
does exclude KeyID 0.
My machine has 6 hardware bits in total (that is KeyID 0 ~ 63), and the upper 48
KeyIDs are reserved to TDX. In my case:
[Bit 31:0] = 15
[Bit 63:32] = 48
And tdx_keyid_start and nr_tdx_keyids are 16 and 48.
The TDX KeyID range: [16, 63], or [16, 64).
So [Bit 31:0] reports only "NUM_MKTME_KIDS", which excludes KeyID 0.
This is where a comment is needed and can actually help.
/*
* tdx_keyid_start/num indicate that TDX is uninitialized. This
* is used in TDX initialization error paths to take it from
* initialized -> uninitialized.
*/
Just want to point out after removing the !x2apic_enabled() check, the only
thing need to do here is to detect/record the TDX KeyIDs.
And the purpose of this TDX boot-time initialization code is to provide
platform_tdx_enabled() function so that kexec() can use.
To distinguish boot-time TDX initialization from runtime TDX module
initialization, how about change the comment to below?
static void __init clear_tdx(void)
{
/*
* tdx_keyid_start and nr_tdx_keyids indicate that TDX is not
* enabled by the BIOS. This is used in TDX boot-time
* initializatiton error paths to take it from enabled to not
* enabled.
*/
tdx_keyid_start = nr_tdx_keyids = 0;
}
[...]
I honestly have no idea what "boot-time TDX initialization" is versus
"runtime TDX module initialization". This doesn't hel.
I'll use your original comment.
quoted
And below is the updated patch. How does it look to you?
Let's see...
...
quoted
+static u32 tdx_keyid_start __ro_after_init;+static u32 nr_tdx_keyids __ro_after_init;++static int __init record_keyid_partitioning(void)+{+ u32 nr_mktme_keyids;+ int ret;++ /*+ * IA32_MKTME_KEYID_PARTIONING:+ * Bit [31:0]: Number of MKTME KeyIDs.+ * Bit [63:32]: Number of TDX private KeyIDs.+ */+ ret = rdmsr_safe(MSR_IA32_MKTME_KEYID_PARTITIONING, &nr_mktme_keyids,+ &nr_tdx_keyids);+ if (ret)+ return -ENODEV;++ if (!nr_tdx_keyids)+ return -ENODEV;++ /* TDX KeyIDs start after the last MKTME KeyID. */+ tdx_keyid_start++;
tdx_keyid_start is uniniitalized here. So, it'd be 0, then ++'d.
Kai, please take a moment and slow down. This isn't a race. I offered
some replacement code here, which you've discarded, missed or ignored
and in the process broken this code.
This approach just wastes reviewer time. It's not working for me.
Apology. I missed it this time.
I'm going to make a suggestion (aka. a demand): You can post these
patches at most once a week. You get a whole week to (carefully)
incorporate reviewer feedback, make the patch better, and post a new
version. Need more time? Go ahead and take it. Take as much time as
you want.
From: Dave Hansen <hidden> Date: 2022-11-22 23:39:58
On 11/20/22 16:26, Kai Huang wrote:
TDX provides increased levels of memory confidentiality and integrity.
This requires special hardware support for features like memory
encryption and storage of memory integrity checksums. Not all memory
satisfies these requirements.
As a result, TDX introduced the concept of a "Convertible Memory Region"
(CMR). During boot, the firmware builds a list of all of the memory
ranges which can provide the TDX security guarantees. The list of these
ranges, along with TDX module information, is available to the kernel by
querying the TDX module via TDH.SYS.INFO SEAMCALL.
I think the last sentence goes too far. What does it matter what the
name of the SEAMCALL is? Who cares at this point? It's in the patch.
Scroll down two pages if you really care.
The host kernel can choose whether or not to use all convertible memory
regions as TDX-usable memory. Before the TDX module is ready to create
any TDX guests, the kernel needs to configure the TDX-usable memory
regions by passing an array of "TD Memory Regions" (TDMRs) to the TDX
module. Constructing the TDMR array requires information of both the
TDX module (TDSYSINFO_STRUCT) and the Convertible Memory Regions. Call
TDH.SYS.INFO to get this information as a preparation.
That last sentece is kinda goofy. I think there's a way to distill this
whole thing down more effecively.
CMRs tell the kernel which memory is TDX compatible. The kernel
takes CMRs and constructs "TD Memory Regions" (TDMRs). TDMRs
let the kernel grante TDX protections to some or all of the CMR
areas.
Use static variables for both TDSYSINFO_STRUCT and CMR array to avoid
I find it very useful to be precise when referring to code. Your code
says 'tdsysinfo_struct', yet this says 'TDSYSINFO_STRUCT'. Why the
difference?
having to pass them as function arguments when constructing the TDMR
array. And they are too big to be put to the stack anyway. Also, KVM
needs to use the TDSYSINFO_STRUCT to create TDX guests.
This is also a great place to mention that the tdsysinfo_struct contains
a *lot* of gunk which will not be used for a bit or that may never get
used.
@@ -40,6 +41,11 @@ static enum tdx_module_status_t tdx_module_status;/* Prevent concurrent attempts on TDX detection and initialization */staticDEFINE_MUTEX(tdx_module_lock);+/* Below two are used in TDH.SYS.INFO SEAMCALL ABI */+staticstructtdsysinfo_structtdx_sysinfo;+staticstructcmr_infotdx_cmr_array[MAX_CMRS]__aligned(CMR_INFO_ARRAY_ALIGNMENT);+staticinttdx_cmr_num;+/**DetectTDXprivateKeyIDstoseewhetherTDXhasbeenenabledbythe*BIOS.BothinitializingtheTDXmoduleandrunningTDXguestrequire
@@ -208,6 +214,121 @@ static int tdx_module_init_cpus(void)returnatomic_read(&sc.err);}+staticinlineboolis_cmr_empty(structcmr_info*cmr)+{+return!cmr->size;+}++staticinlineboolis_cmr_ok(structcmr_info*cmr)+{+/* CMR must be page aligned */+returnIS_ALIGNED(cmr->base,PAGE_SIZE)&&+IS_ALIGNED(cmr->size,PAGE_SIZE);+}++staticvoidprint_cmrs(structcmr_info*cmr_array,intcmr_num,+constchar*name)+{+inti;++for(i=0;i<cmr_num;i++){+structcmr_info*cmr=&cmr_array[i];++pr_info("%s : [0x%llx, 0x%llx)\n",name,+cmr->base,cmr->base+cmr->size);+}+}++/* Check CMRs reported by TDH.SYS.INFO, and trim tail empty CMRs. */+staticinttrim_empty_cmrs(structcmr_info*cmr_array,int*actual_cmr_num)+{+structcmr_info*cmr;+inti,cmr_num;++/*+*IntelTDXmodulespec,20.7.3CMR_INFO:+*+*TDH.SYS.INFOleaffunctionreturnsaMAX_CMRS(32)entry+*arrayofCMR_INFOentries.TheCMRsaresortedfromthe+*lowestbaseaddresstothehighestbaseaddress,andthey+*arenon-overlapping.+*+*ThisimpliesthatBIOSmaygenerateinvalidemptyentries+*iftotalCMRsarelessthan32.Needtoskipthemmanually.+*+*CMRalsomustbe4Kaligned.TDXdoesn'ttrustBIOS.TDX+*actuallyverifiesCMRsbeforeitgetsenabled,soanything+*doesn'tmeetabovemeanskernelbug(orTDXisbroken).+*/
I dislike comments like this that describe all the code below. Can't
you simply put the comment near the code that implements it?
quoted hunk
+ cmr = &cmr_array[0];+ /* There must be at least one valid CMR */+ if (WARN_ON_ONCE(is_cmr_empty(cmr) || !is_cmr_ok(cmr)))+ goto err;++ cmr_num = *actual_cmr_num;+ for (i = 1; i < cmr_num; i++) {+ struct cmr_info *cmr = &cmr_array[i];+ struct cmr_info *prev_cmr = NULL;++ /* Skip further empty CMRs */+ if (is_cmr_empty(cmr))+ break;++ /*+ * Do sanity check anyway to make sure CMRs:+ * - are 4K aligned+ * - don't overlap+ * - are in address ascending order.+ */+ if (WARN_ON_ONCE(!is_cmr_ok(cmr)))+ goto err;
Why does cmr_array[0] get a pass on the empty and sanity checks?
quoted hunk
+ prev_cmr = &cmr_array[i - 1];+ if (WARN_ON_ONCE((prev_cmr->base + prev_cmr->size) >+ cmr->base))+ goto err;+ }++ /* Update the actual number of CMRs */+ *actual_cmr_num = i;
That comment is not helpful. Yes, this is literally updating the number
of CMRs. Literally. That's the "what". But, the "why" is important.
Why is it doing this?
This is the point where I start to lose patience with these comments.
These are just a waste of space.
Also, I saw the loop above check 'cmr_num' CMRs for is_cmr_ok(). Now,
it'll print an 'actual_cmr_num=1' number of CMRs as being
"kernel-checked". Why? That makes zero sense.
quoted hunk
+ return 0;+err:+ pr_info("[TDX broken ?]: Invalid CMRs detected\n");+ print_cmrs(cmr_array, cmr_num, "BIOS-CMR");+ return -EINVAL;+}++static int tdx_get_sysinfo(void)+{+ struct tdx_module_output out;+ int ret;++ BUILD_BUG_ON(sizeof(struct tdsysinfo_struct) != TDSYSINFO_STRUCT_SIZE);++ ret = seamcall(TDH_SYS_INFO, __pa(&tdx_sysinfo), TDSYSINFO_STRUCT_SIZE,+ __pa(tdx_cmr_array), MAX_CMRS, NULL, &out);+ if (ret)+ return ret;++ /* R9 contains the actual entries written the CMR array. */+ tdx_cmr_num = out.r9;++ pr_info("TDX module: atributes 0x%x, vendor_id 0x%x, major_version %u, minor_version %u, build_date %u, build_num %u",+ tdx_sysinfo.attributes, tdx_sysinfo.vendor_id,+ tdx_sysinfo.major_version, tdx_sysinfo.minor_version,+ tdx_sysinfo.build_date, tdx_sysinfo.build_num);
This is a case where a little bit of vertical alignment will go a long way:
++ /*+ * trim_empty_cmrs() updates the actual number of CMRs by+ * dropping all tail empty CMRs.+ */+ return trim_empty_cmrs(tdx_cmr_array, &tdx_cmr_num);+}
Why does this both need to respect the "tdx_cmr_num = out.r9" value
*and* trim the empty ones? Couldn't it just ignore the "tdx_cmr_num =
out.r9" value and just trim the empty ones either way? It's not like
there is a billion of them. It would simplify the code for sure.
quoted hunk
/*
* Detect and initialize the TDX module.
*
@@ -232,6 +353,10 @@ static int init_tdx_module(void) if (ret) goto out;+ ret = tdx_get_sysinfo();+ if (ret)+ goto out;+ /* * Return -EINVAL until all steps of TDX module initialization * process are done.
@@ -15,10 +15,71 @@/**TDXmoduleSEAMCALLleaffunctions*/+#define TDH_SYS_INFO 32#define TDH_SYS_INIT 33#define TDH_SYS_LP_INIT 35#define TDH_SYS_LP_SHUTDOWN 44+structcmr_info{+u64base;+u64size;+}__packed;++#define MAX_CMRS 32+#define CMR_INFO_ARRAY_ALIGNMENT 512++structcpuid_config{+u32leaf;+u32sub_leaf;+u32eax;+u32ebx;+u32ecx;+u32edx;+}__packed;++#define TDSYSINFO_STRUCT_SIZE 1024+#define TDSYSINFO_STRUCT_ALIGNMENT 1024++structtdsysinfo_struct{+/* TDX-SEAM Module Info */+u32attributes;+u32vendor_id;+u32build_date;+u16build_num;+u16minor_version;+u16major_version;+u8reserved0[14];+/* Memory Info */+u16max_tdmrs;+u16max_reserved_per_tdmr;+u16pamt_entry_size;+u8reserved1[10];+/* Control Struct Info */+u16tdcs_base_size;+u8reserved2[2];+u16tdvps_base_size;+u8tdvps_xfam_dependent_size;+u8reserved3[9];+/* TD Capabilities */+u64attributes_fixed0;+u64attributes_fixed1;+u64xfam_fixed0;+u64xfam_fixed1;+u8reserved4[32];+u32num_cpuid_config;+/*+*TheactualnumberofCPUID_CONFIGdependsonabove+*'num_cpuid_config'.Thesizeof'structtdsysinfo_struct'+*is1024BdefinedbyTDXarchitecture.Useaunionwith+*specificpaddingtomake'sizeof(structtdsysinfo_struct)'+*equalto1024.+*/+union{+structcpuid_configcpuid_configs[0];+u8reserved5[892];+};
Can you double check what the "right" way to do variable arrays is these
days? I thought the [0] method was discouraged.
Also, it isn't *really* 892 bytes of reserved space, right? Anything
that's not cpuid_configs[] is reserved, I presume. Could you try to be
more precise there?
quoted hunk
+} __packed __aligned(TDSYSINFO_STRUCT_ALIGNMENT);+ /* * Do not put any hardware-defined TDX structure representations below * this comment!
From: Dave Hansen <hidden> Date: 2022-11-23 00:22:05
On 11/20/22 16:26, Kai Huang wrote:
TDX reports a list of "Convertible Memory Region" (CMR) to indicate all
memory regions that can possibly be used by the TDX module, but they are
not automatically usable to the TDX module. As a step of initializing
the TDX module, the kernel needs to choose a list of memory regions (out
from convertible memory regions) that the TDX module can use and pass
those regions to the TDX module. Once this is done, those "TDX-usable"
memory regions are fixed during module's lifetime. No more TDX-usable
memory can be added to the TDX module after that.
The initial support of TDX guests will only allocate TDX guest memory
from the global page allocator. To keep things simple, this initial
implementation simply guarantees all pages in the page allocator are TDX
memory. To achieve this, use all system memory in the core-mm at the
time of initializing the TDX module as TDX memory, and at the meantime,
refuse to add any non-TDX-memory in the memory hotplug.
Specifically, walk through all memory regions managed by memblock and
add them to a global list of "TDX-usable" memory regions, which is a
fixed list after the module initialization (or empty if initialization
fails). To reject non-TDX-memory in memory hotplug, add an additional
check in arch_add_memory() to check whether the new region is covered by
any region in the "TDX-usable" memory region list.
Note this requires all memory regions in memblock are TDX convertible
memory when initializing the TDX module. This is true in practice if no
new memory has been hot-added before initializing the TDX module, since
in practice all boot-time present DIMM is TDX convertible memory. If
any new memory has been hot-added, then initializing the TDX module will
fail due to that memory region is not covered by CMR.
This can be enhanced in the future, i.e. by allowing adding non-TDX
memory to a separate NUMA node. In this case, the "TDX-capable" nodes
and the "non-TDX-capable" nodes can co-exist, but the kernel/userspace
needs to guarantee memory pages for TDX guests are always allocated from
the "TDX-capable" nodes.
Note TDX assumes convertible memory is always physically present during
machine's runtime. A non-buggy BIOS should never support hot-removal of
any convertible memory. This implementation doesn't handle ACPI memory
removal but depends on the BIOS to behave correctly.
My eyes glazed over about halfway through that. Can you try to trim it
down a bit, or at least try to summarize it better up front?
+ * must be TDX memory, which is a fixed set of memory regions+ * that are passed to the TDX module. Reject the new region+ * if it is not TDX memory to guarantee above is true.+ */+ if (!tdx_cc_memory_compatible(start_pfn, start_pfn + nr_pages))+ return -EINVAL;
There's a real art to making a right-size comment. I don't think this
needs to be any more than:
/*
* Not all memory is compatible with TDX. Reject
* the addition of any incomatible memory.
*/
If you want to write a treatise, do it in Documentation or at the
tdx_cc_memory_compatible() definition.
@@ -46,6 +58,9 @@ static struct tdsysinfo_struct tdx_sysinfo; static struct cmr_info tdx_cmr_array[MAX_CMRS] __aligned(CMR_INFO_ARRAY_ALIGNMENT); static int tdx_cmr_num;+/* All TDX-usable memory regions */+static LIST_HEAD(tdx_memlist);+ /* * Detect TDX private KeyIDs to see whether TDX has been enabled by the * BIOS. Both initializing the TDX module and running TDX guest require
@@ -329,6 +344,107 @@ static int tdx_get_sysinfo(void) return trim_empty_cmrs(tdx_cmr_array, &tdx_cmr_num); }+/* Check whether the given pfn range is covered by any CMR or not. */+static bool pfn_range_covered_by_cmr(unsigned long start_pfn,+ unsigned long end_pfn)+{+ int i;++ for (i = 0; i < tdx_cmr_num; i++) {+ struct cmr_info *cmr = &tdx_cmr_array[i];+ unsigned long cmr_start_pfn;+ unsigned long cmr_end_pfn;++ cmr_start_pfn = cmr->base >> PAGE_SHIFT;+ cmr_end_pfn = (cmr->base + cmr->size) >> PAGE_SHIFT;++ if (start_pfn >= cmr_start_pfn && end_pfn <= cmr_end_pfn)+ return true;+ }
What if the pfn range overlaps two CMRs? It will never pass any
individual overlap test and will return false.
quoted hunk
+ return false;+}++/*+ * Add a memory region on a given node as a TDX memory block. The caller+ * to make sure all memory regions are added in address ascending order
s/to/must/
quoted hunk
+ * and don't overlap.+ */+static int add_tdx_memblock(unsigned long start_pfn, unsigned long end_pfn,+ int nid)+{+ struct tdx_memblock *tmb;++ tmb = kmalloc(sizeof(*tmb), GFP_KERNEL);+ if (!tmb)+ return -ENOMEM;++ INIT_LIST_HEAD(&tmb->list);+ tmb->start_pfn = start_pfn;+ tmb->end_pfn = end_pfn;+ tmb->nid = nid;++ list_add_tail(&tmb->list, &tdx_memlist);+ return 0;+}++static void free_tdx_memory(void)
This is named a bit too generically. How about free_tdx_memlist() or
something?
quoted hunk
+{+ while (!list_empty(&tdx_memlist)) {+ struct tdx_memblock *tmb = list_first_entry(&tdx_memlist,+ struct tdx_memblock, list);++ list_del(&tmb->list);+ kfree(tmb);+ }+}++/*+ * Add all memblock memory regions to the @tdx_memlist as TDX memory.+ * Must be called when get_online_mems() is called by the caller.+ */
Again, this explains the "what", but not the "why".
/*
* Ensure that all memblock memory regions are convertible to TDX
* memory. Once this has been established, stash the memblock
* ranges off in a secondary structure because $REASONS.
*/
Which makes me wonder: Why do you even need a secondary structure here?
What's wrong with the memblocks themselves?
quoted hunk
+static int build_tdx_memory(void)+{+ unsigned long start_pfn, end_pfn;+ int i, nid, ret;++ for_each_mem_pfn_range(i, MAX_NUMNODES, &start_pfn, &end_pfn, &nid) {+ /*+ * The first 1MB may not be reported as TDX convertible+ * memory. Manually exclude them as TDX memory.
I don't like the "may not" here very much.
quoted hunk
+ * This is fine as the first 1MB is already reserved in+ * reserve_real_mode() and won't end up to ZONE_DMA as+ * free page anyway.
^ free pages
+ */
This way too wishy washy. The TDX module may or may not... Then, it
doesn't matter since reserve_real_mode() does it anyway...
Then it goes and adds code to skip it!
Please just put a dang stake in the ground. If the other code deals
with this, then explain *why* more is needed here.
quoted hunk
+ /* Verify memory is truly TDX convertible memory */+ if (!pfn_range_covered_by_cmr(start_pfn, end_pfn)) {+ pr_info("Memory region [0x%lx, 0x%lx) is not TDX convertible memorry.\n",+ start_pfn << PAGE_SHIFT,+ end_pfn << PAGE_SHIFT);+ return -EINVAL;
... no 'goto err'? This leaks all the previous add_tdx_memblock()
structures, right?
quoted hunk
+ }++ /*+ * Add the memory regions as TDX memory. The regions in+ * memblock has already guaranteed they are in address+ * ascending order and don't overlap.+ */+ ret = add_tdx_memblock(start_pfn, end_pfn, nid);+ if (ret)+ goto err;+ }++ return 0;+err:+ free_tdx_memory();+ return ret;+}+ /* * Detect and initialize the TDX module. *
@@ -357,12 +473,56 @@ static int init_tdx_module(void) if (ret) goto out;+ /*+ * All memory regions that can be used by the TDX module must be+ * passed to the TDX module during the module initialization.+ * Once this is done, all "TDX-usable" memory regions are fixed+ * during module's runtime.+ *+ * The initial support of TDX guests only allocates memory from+ * the global page allocator. To keep things simple, for now+ * just make sure all pages in the page allocator are TDX memory.+ *+ * To achieve this, use all system memory in the core-mm at the+ * time of initializing the TDX module as TDX memory, and at the+ * meantime, reject any new memory in memory hot-add.+ *+ * This works as in practice, all boot-time present DIMM is TDX+ * convertible memory. However if any new memory is hot-added+ * before initializing the TDX module, the initialization will+ * fail due to that memory is not covered by CMR.+ *+ * This can be enhanced in the future, i.e. by allowing adding or+ * onlining non-TDX memory to a separate node, in which case the+ * "TDX-capable" nodes and the "non-TDX-capable" nodes can exist+ * together -- the userspace/kernel just needs to make sure pages+ * for TDX guests must come from those "TDX-capable" nodes.+ *+ * Build the list of TDX memory regions as mentioned above so+ * they can be passed to the TDX module later.+ */
This is msotly Documentation/, not a code comment. Please clean it up.
quoted hunk
+ get_online_mems();++ ret = build_tdx_memory();+ if (ret)+ goto out; /* * Return -EINVAL until all steps of TDX module initialization * process are done. */ ret = -EINVAL; out:+ /*+ * Memory hotplug checks the hot-added memory region against the+ * @tdx_memlist to see if the region is TDX memory.+ *+ * Do put_online_mems() here to make sure any modification to+ * @tdx_memlist is done while holding the memory hotplug read+ * lock, so that the memory hotplug path can just check the+ * @tdx_memlist w/o holding the @tdx_module_lock which may cause+ * deadlock.+ */
I'm honestly not following any of that.
quoted hunk
+ put_online_mems(); return ret; }
@@ -485,3 +645,26 @@ int tdx_enable(void) return ret; } EXPORT_SYMBOL_GPL(tdx_enable);++/*+ * Check whether the given range is TDX memory. Must be called between+ * mem_hotplug_begin()/mem_hotplug_done().+ */+bool tdx_cc_memory_compatible(unsigned long start_pfn, unsigned long end_pfn)+{+ struct tdx_memblock *tmb;++ /* Empty list means TDX isn't enabled successfully */+ if (list_empty(&tdx_memlist))+ return true;++ list_for_each_entry(tmb, &tdx_memlist, list) {+ /*+ * The new range is TDX memory if it is fully covered+ * by any TDX memory block.+ */+ if (start_pfn >= tmb->start_pfn && end_pfn <= tmb->end_pfn)+ return true;
Same bug. What if the start/end_pfn range is covered by more than one
tdx_memblock?
From: Huang, Kai <hidden> Date: 2022-11-23 00:31:11
On Tue, 2022-11-22 at 21:03 +0100, Thomas Gleixner wrote:
quoted
quoted
if (tdx_global_data_ptr->num_of_init_lps < tdx_global_data_ptr-
quoted
num_of_lps)
{
TDX_ERROR("Num of initialized lps %d is smaller than total num of
lps %d\n",
tdx_global_data_ptr->num_of_init_lps,
tdx_global_data_ptr->num_of_lps);
retval = TDX_SYS_CONFIG_NOT_PENDING;
goto EXIT;
}
tdx_global_data_ptr->num_of_init_lps is incremented at TDH.SYS.INIT
time. That if() is called at TDH.SYS.CONFIG time to help bring the
module up.
So, I think you're right. I don't see the docs that actually *explain*
this "you must seamcall all the things" requirement.
The code actually enforces this.
At TDH.SYS.INIT which is the first operation it gets the total number
of LPs from the sysinfo table:
src/vmm_dispatcher/api_calls/tdh_sys_init.c:
tdx_global_data_ptr->num_of_lps = sysinfo_table_ptr-
quoted
mcheck_fields.tot_num_lps;
Then TDH.SYS.LP.INIT increments the count of initialized LPs.
src/vmm_dispatcher/api_calls/tdh_sys_lp_init.c:
increment_num_of_lps(tdx_global_data_ptr)
_lock_xadd_32b(&tdx_global_data_ptr->num_of_init_lps, 1);
Finally TDH.SYS.CONFIG checks whether _ALL_ LPs have been initialized.
src/vmm_dispatcher/api_calls/tdh_sys_config.c:
if (tdx_global_data_ptr->num_of_init_lps < tdx_global_data_ptr-
quoted
num_of_lps)
Clearly that's nowhere spelled out in the documentation, but I don't
buy the 'architecturaly required' argument not at all. It's an
implementation detail of the TDX module.
Hi Thomas,
Thanks for review!
I agree on hardware level there shouldn't be such requirement (not 100% sure
though), but I guess from kernel's perspective, "the implementation detail of
the TDX module" is sort of "architectural requirement" -- at least Intel arch
guys think so I guess.
Technically there is IMO ZERO requirement to do so.
1) The TDX module is global
2) Seam-root and Seam-non-root operation are strictly a LP property.
The only architectural prerequisite for using Seam on a LP is that
obviously the encryption/decryption mechanics have been initialized
on the package to which the LP belongs.
I can see why it might be complicated to add/remove an LP after
initialization fact, but technically it should be possible.
"kernel soft offline" actually isn't an issue. We can bring down a logical cpu
after it gets initialized and then bring it up again.
Only add/removal of physical cpu will cause problem:
TDX MCHECK verifies all boot-time present cpus to make sure they are TDX-
compatible before it enables TDX in hardware. MCHECK cannot run on hot-added
CPU, so TDX cannot support physical CPU hotplug.
We tried to get it clarified in the specification, and below is what TDX/module
arch guys agreed to put to the TDX module spec (just checked it's not in latest
public spec yet, but they said it will be in next release):
"
4.1.3.2. CPU Configuration
During platform boot, MCHECK verifies all logical CPUs to ensure they meet TDX’s
security and certain functionality requirements, and MCHECK passes the following
CPU configuration information to the NP-SEAMLDR, P-SEAMLDR and the TDX Module:
· Total number of logical processors in the platform.
· Total number of installed packages in the platform.
· A table of per-package CPU family, model and stepping etc.
identification, as enumerated by CPUID(1).EAX.
The above information is static and does not change after platform boot and
MCHECK run.
Note: TDX doesn’t support adding or removing CPUs from TDX security
perimeter, as checked my MCHECK. BIOS should prevent CPUs from being hot-added
or hot-removed after platform boots.
The TDX module performs additional checks of the CPU’s configuration and
supported features, by reading MSRs and CPUID information as described in the
following sections.
"
TDX/Seam is not that special.
But what's absolutely annoying is that the documentation lacks any
information about the choice of enforcement which has been hardcoded
into the Seam module for whatever reasons.
Maybe I overlooked it, but then it's definitely well hidden.
It depends on the TDX Module implementation, which TDX arch guys think should be
"architectural" I think.
From: Huang, Kai <hidden> Date: 2022-11-23 01:14:49
On Wed, 2022-11-23 at 00:30 +0000, Huang, Kai wrote:
quoted
Clearly that's nowhere spelled out in the documentation, but I don't
buy the 'architecturaly required' argument not at all. It's an
implementation detail of the TDX module.
Hi Thomas,
Thanks for review!
I agree on hardware level there shouldn't be such requirement (not 100% sure
though), but I guess from kernel's perspective, "the implementation detail of
the TDX module" is sort of "architectural requirement" -- at least Intel arch
guys think so I guess.
Let me double check with the TDX module folks and figure out the root of the
requirement.
Thanks.
From: Huang, Kai <hidden> Date: 2022-11-23 01:15:46
On Tue, 2022-11-22 at 20:33 +0100, Peter Zijlstra wrote:
On Tue, Nov 22, 2022 at 11:24:48AM -0800, Dave Hansen wrote:
quoted
quoted
Not intialize TDX on busy NOHZ_FULL cpus and hard-limit the cpumask of
all TDX using tasks.
I don't think that works. As I mentioned to Thomas elsewhere, you don't
just need to initialize TDX on the CPUs where it is used. Before the
module will start working you need to initialize it on *all* the CPUs it
knows about. The module itself has a little counter where it tracks
this and will refuse to start being useful until it gets called
thoroughly enough.
That's bloody terrible, that is. How are we going to make that work with
the SMT mitigation crud that forces the SMT sibilng offline?
Then the counters don't match and TDX won't work.
Can we get this limitiation removed and simply let the module throw a
wobbly (error) when someone tries and use TDX without that logical CPU
having been properly initialized?
Dave kindly helped to raise this issue and I'll follow up with TDX module guys
to see whether we can remove/ease such limitation.
Thanks!