From: Dave Hansen <hidden> Date: 2016-02-23 01:11:10
As promised, here are the proposed new Memory Protection Keys
interfaces. These interfaces make it possible to do something
with pkeys other than execute-only support.
There are 5 syscalls here. I'm hoping for reviews of this set
which can help nail down what the final interfaces will be.
You can find a high-level overview of the feature and the new
syscalls here:
https://www.sr71.net/~dave/intel/pkeys.txt
===============================================================
To use memory protection keys (pkeys), an application absolutely
needs to be able to set the pkey field in the PTE (obviously has
to be done in-kernel) and make changes to the "rights" register
(using unprivileged instructions).
An application also needs to have an an allocator for the keys
themselves. If two different parts of an application both want
to protect their data with pkeys, they first need to know which
key to use for their individual purposes.
This set introduces 5 system calls, in 3 logical groups:
1. PTE pkey setting (sys_pkey_mprotect(), patches #1-3)
2. Key allocation (sys_pkey_alloc() / sys_pkey_free(), patch #4)
3. Rights register manipulation (sys_pkey_set/get(), patch #5)
These patches build on top of "core" support already in the tip tree,
specifically 62b5f7d013, which can currently be found at:
http://git.kernel.org/cgit/linux/kernel/git/tip/tip.git/log/?h=mm/pkeys
I have manpages written for some of these syscalls, and I will
submit a full set of manpages once we've reached some consensus
on what the interfaces should be.
This set is also available here:
git://git.kernel.org/pub/scm/linux/kernel/git/daveh/x86-pkeys.git pkeys-v024
I've written a set of unit tests for these interfaces, which is
available here:
https://www.sr71.net/~dave/intel/pkeys-test-2016-02-22/
I will submit that code for inclusion with the final version of
these patches.
=== diffstat ===
Dave Hansen (7):
x86, pkeys: Documentation
mm: implement new pkey_mprotect() system call
x86, pkeys: make mprotect_key() mask off additional vm_flags
x86: wire up mprotect_key() system call
x86, pkeys: allocation/free syscalls
x86, pkeys: add pkey set/get syscalls
pkeys: add details of system call use to Documentation/
Documentation/x86/protection-keys.txt | 91 +++++++++++++++++
arch/x86/entry/syscalls/syscall_32.tbl | 5 +
arch/x86/entry/syscalls/syscall_64.tbl | 5 +
arch/x86/include/asm/mmu.h | 8 ++
arch/x86/include/asm/mmu_context.h | 25 +++--
arch/x86/include/asm/pkeys.h | 83 ++++++++++++++-
arch/x86/kernel/fpu/xstate.c | 73 +++++++++++++-
arch/x86/mm/pkeys.c | 40 ++++++--
include/linux/pkeys.h | 39 ++++++--
include/uapi/asm-generic/mman-common.h | 5 +
mm/mprotect.c | 133 ++++++++++++++++++++++++-
11 files changed, 476 insertions(+), 31 deletions(-)
Cc: linux-api-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
Cc: linux-mm-Bw31MaZKKs3YtjvyW6yDsg@public.gmane.org
Cc: x86-DgEjT+Ai2ygdnm+yROfE0A@public.gmane.org
Cc: torvalds-de/tnXTf+JLsfHDXvbKv3WD2FQJk+8+b@public.gmane.org
Cc: akpm-de/tnXTf+JLsfHDXvbKv3WD2FQJk+8+b@public.gmane.org
From: Dave Hansen <hidden> Date: 2016-02-23 01:11:14
From: Dave Hansen <dave.hansen@linux.intel.com>
pkey_mprotect() is just like mprotect, except it also takes a
protection key as an argument. On systems that do not support
protection keys, it still works, but requires that key=0.
Otherwise it does exactly what mprotect does.
I expect it to get used like this, if you want to guarantee that
any mapping you create can *never* be accessed without the right
protection keys set up.
int real_prot = PROT_READ|PROT_WRITE;
pkey = pkey_alloc(0, PKEY_DENY_ACCESS);
ptr = mmap(NULL, PAGE_SIZE, PROT_NONE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0);
ret = pkey_mprotect(ptr, PAGE_SIZE, real_prot, pkey);
This way, there is *no* window where the mapping is accessible
since it was always either PROT_NONE or had a protection key set.
We settled on 'unsigned long' for the type of the key here. We
only need 4 bits on x86 today, but I figured that other
architectures might need some more space.
Semantically, we have a bit of a problem if we combine this
syscall with our previously-introduced execute-only support:
What do we do when we mix execute-only pkey use with
pkey_mprotect() use? For instance:
pkey_mprotect(ptr, PAGE_SIZE, PROT_WRITE, 6); // set pkey=6
mprotect(ptr, PAGE_SIZE, PROT_EXEC); // set pkey=X_ONLY_PKEY?
mprotect(ptr, PAGE_SIZE, PROT_WRITE); // is pkey=6 again?
To solve that, we make the plain-mprotect()-initiated execute-only
support only apply to VMAs that have the default protection key (0)
set on them.
Proposed semantics:
1. protection key 0 is special and represents the default,
unassigned protection key. It is always allocated.
2. mprotect() never affects a mapping's pkey_mprotect()-assigned
protection key. A protection key of 0 (even if set explicitly)
represents an unassigned protection key.
2a. mprotect(PROT_EXEC) on a mapping with an assigned protection
key may or may not result in a mapping with execute-only
properties. pkey_mprotect() plus pkey_set() on all threads
should be used to _guarantee_ execute-only semantics.
3. mprotect(PROT_EXEC) may result in an "execute-only" mapping. The
kernel will internally attempt to allocate and dedicate a
protection key for the purpose of execute-only mappings. This
may not be possible in cases where there are no free protection
keys available.
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: linux-api@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: x86@kernel.org
Cc: torvalds@linux-foundation.org
Cc: akpm@linux-foundation.org
---
b/arch/x86/include/asm/mmu_context.h | 15 ++++++++++-----
b/arch/x86/include/asm/pkeys.h | 11 +++++++++--
b/arch/x86/kernel/fpu/xstate.c | 15 ++++++++++++++-
b/arch/x86/mm/pkeys.c | 2 +-
b/mm/mprotect.c | 27 +++++++++++++++++++++++----
5 files changed, 57 insertions(+), 13 deletions(-)
diff -puN arch/x86/include/asm/mmu_context.h~pkeys-85-syscalls-mprotect_pkey arch/x86/include/asm/mmu_context.h
@@ -410,11 +413,12 @@ SYSCALL_DEFINE3(mprotect, unsigned long,for(nstart=start;;){unsignedlongnewflags;-intpkey=arch_override_mprotect_pkey(vma,prot,-1);+intvma_pkey;/* Here we know that vma->vm_start <= nstart < vma->vm_end. */-newflags=calc_vm_prot_bits(prot,pkey);+vma_pkey=arch_override_mprotect_pkey(vma,prot,pkey);+newflags=calc_vm_prot_bits(prot,vma_pkey);newflags|=(vma->vm_flags&~(VM_READ|VM_WRITE|VM_EXEC));/* newflags >> 4 shift VM_MAY% in place of VM_% */
_
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Dave Hansen <hidden> Date: 2016-02-23 01:11:20
From: Dave Hansen <dave.hansen-VuQAYsv1563Yd54FQh9/CA@public.gmane.org>
This patch adds two new system calls:
int pkey_alloc(unsigned long flags, unsigned long init_access_rights)
int pkey_free(int pkey);
These implement an "allocator" for the protection keys
themselves, which can be thought of as analogous to the allocator
that the kernel has for file descriptors. The kernel tracks
which numbers are in use, and only allows operations on keys that
are valid. A key which was not obtained by pkey_alloc() may not,
for instance, be passed to pkey_mprotect() (or the forthcoming
get/set syscalls).
These system calls are also very important given the kernel's use
of pkeys to implement execute-only support. These help ensure
that userspace can never assume that it has control of a key
unless it first asks the kernel.
The 'init_access_rights' argument to pkey_alloc() specifies the
rights that will be established for the returned pkey. For
instance:
pkey = pkey_alloc(flags, PKEY_DENY_WRITE);
will allocate 'pkey', but also sets the bits in PKRU[1] such that
writing to 'pkey' is already denied. This keeps userspace from
needing to have knowledge about manipulating PKRU with the
RDPKRU/WRPKRU instructions. Userspace is still free to use these
instructions as it wishes, but this facility ensures it is no
longer required.
The kernel does _not_ enforce that this interface must be used for
changes to PKRU, even for keys it does not control.
This allocation mechanism could be implemented in userspace.
Even if we did it in userspace, we would still need additional
user/kernel interfaces to tell userspace which keys are being
used by the kernel internally (such as for execute-only
mappings). Having the kernel provide this facility completely
removes the need for these additional interfaces, or having an
implementation of this in userspace at all.
1. PKRU is the Protection Key Rights User register. It is a
usermode-accessible register that controls whether writes
and/or access to each individual pkey is allowed or denied.
Signed-off-by: Dave Hansen <dave.hansen-VuQAYsv1563Yd54FQh9/CA@public.gmane.org>
Cc: linux-api-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
Cc: linux-mm-Bw31MaZKKs3YtjvyW6yDsg@public.gmane.org
Cc: x86-DgEjT+Ai2ygdnm+yROfE0A@public.gmane.org
Cc: torvalds-de/tnXTf+JLsfHDXvbKv3WD2FQJk+8+b@public.gmane.org
Cc: akpm-de/tnXTf+JLsfHDXvbKv3WD2FQJk+8+b@public.gmane.org
---
b/arch/x86/entry/syscalls/syscall_32.tbl | 2
b/arch/x86/entry/syscalls/syscall_64.tbl | 2
b/arch/x86/include/asm/mmu.h | 8 +++
b/arch/x86/include/asm/mmu_context.h | 10 +++
b/arch/x86/include/asm/pkeys.h | 78 +++++++++++++++++++++++++++++--
b/arch/x86/kernel/fpu/xstate.c | 3 +
b/arch/x86/mm/pkeys.c | 40 ++++++++++++---
b/include/linux/pkeys.h | 30 +++++++++--
b/include/uapi/asm-generic/mman-common.h | 5 +
b/mm/mprotect.c | 56 ++++++++++++++++++++++
10 files changed, 213 insertions(+), 21 deletions(-)
diff -puN arch/x86/entry/syscalls/syscall_32.tbl~pkeys-86-syscalls-allocation arch/x86/entry/syscalls/syscall_32.tbl
@@ -334,6 +334,8 @@ 325 common mlock2 sys_mlock2 326 common copy_file_range sys_copy_file_range 327 common pkey_mprotect sys_pkey_mprotect+328 common pkey_alloc sys_pkey_alloc+329 common pkey_free sys_pkey_free # # x32-specific system call numbers start at 512 to avoid cache impact
@@ -108,7 +108,16 @@ static inline void enter_lazy_tlb(structstaticinlineintinit_new_context(structtask_struct*tsk,structmm_struct*mm){+#ifdef CONFIG_X86_INTEL_MEMORY_PROTECTION_KEYS+if(boot_cpu_has(X86_FEATURE_OSPKE)){+/* pkey 0 is the default and always allocated */+mm->context.pkey_allocation_map=0x1;+/* -1 means unallocated or invalid */+mm->context.execute_only_pkey=-1;+}+#endifinit_new_context_ldt(tsk,mm);+return0;}staticinlinevoiddestroy_context(structmm_struct*mm)
@@ -23,6 +23,14 @@ typedef struct {conststructvdso_image*vdso_image;/* vdso image in use */atomic_tperf_rdpmc_allowed;/* nonzero if rdpmc is allowed */+#ifdef CONFIG_X86_INTEL_MEMORY_PROTECTION_KEYS+/*+*Onebitperprotectionkeysayswhetheruserspacecan+*useitornot.protectedbymmap_sem.+*/+u16pkey_allocation_map;+s16execute_only_pkey;+#endif}mm_context_t;#ifdef CONFIG_SMP
@@ -21,8 +21,19 @@int__execute_only_pkey(structmm_struct*mm){+boolneed_to_set_mm_pkey=false;+intexecute_only_pkey=mm->context.execute_only_pkey;intret;+/* Do we need to assign a pkey for mm's execute-only maps? */+if(execute_only_pkey==-1){+/* Go allocate one to use, which might fail */+execute_only_pkey=mm_pkey_alloc(mm);+if(!validate_pkey(execute_only_pkey))+return-1;+need_to_set_mm_pkey=true;+}+/**Wedonotwanttogothroughtherelativelycostly*dancetosetPKRUifwedonotneedto.Checkit
@@ -32,22 +43,33 @@ int __execute_only_pkey(struct mm_struct*canmakefpregsinactive.*/preempt_disable();-if(fpregs_active()&&-!__pkru_allows_read(read_pkru(),PKEY_DEDICATED_EXECUTE_ONLY)){+if(!need_to_set_mm_pkey&&+fpregs_active()&&+!__pkru_allows_read(read_pkru(),execute_only_pkey)){preempt_enable();-returnPKEY_DEDICATED_EXECUTE_ONLY;+returnexecute_only_pkey;}preempt_enable();-ret=__arch_set_user_pkey_access(current,PKEY_DEDICATED_EXECUTE_ONLY,-PKEY_DISABLE_ACCESS);+/*+*SetupPKRUsothatitdeniesaccessforeverything+*otherthanexecution.+*/+ret=__arch_set_user_pkey_access(current,execute_only_pkey,+PKEY_DISABLE_ACCESS);+/**IfthePKRU-setoperationfailedsomehow,justreturn*0andeffectivelydisableexecute-onlysupport.*/-if(ret)-return0;+if(ret){+mm_set_pkey_free(mm,execute_only_pkey);+return-1;+}-returnPKEY_DEDICATED_EXECUTE_ONLY;+/* We got one, store it and use it from here on out */+if(need_to_set_mm_pkey)+mm->context.execute_only_pkey=execute_only_pkey;+returnexecute_only_pkey;}staticinlineboolvma_is_pkey_exec_only(structvm_area_struct*vma)
@@ -55,7 +77,7 @@ static inline bool vma_is_pkey_exec_only/* Do this check first since the vm_flags should be hot */if((vma->vm_flags&(VM_READ|VM_WRITE|VM_EXEC))!=VM_EXEC)returnfalse;-if(vma_pkey(vma)!=PKEY_DEDICATED_EXECUTE_ONLY)+if(vma_pkey(vma)!=vma->vm_mm->context.execute_only_pkey)returnfalse;returntrue;
From: Dave Hansen <hidden> Date: 2016-02-23 01:11:45
From: Dave Hansen <dave.hansen@linux.intel.com>
This spells out all of the pkey-related system calls that we have
and provides some example code fragments to demonstrate how we
expect them to be used.
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: linux-api@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: x86@kernel.org
Cc: torvalds@linux-foundation.org
Cc: akpm@linux-foundation.org
---
b/Documentation/x86/protection-keys.txt | 63 ++++++++++++++++++++++++++++++++
1 file changed, 63 insertions(+)
diff -puN Documentation/x86/protection-keys.txt~pkeys-98-syscall-docs Documentation/x86/protection-keys.txt
@@ -19,6 +19,69 @@ even though there is theoretically space permissions are enforced on data access only and have no effect on instruction fetches.+=========================== Syscalls ===========================++There are 5 system calls which directly interact with pkeys:++ int pkey_alloc(unsigned long flags, unsigned long init_access_rights)+ int pkey_free(int pkey);+ int sys_pkey_mprotect(unsigned long start, size_t len,+ unsigned long prot, int pkey);+ unsigned long pkey_get(int pkey);+ int pkey_set(int pkey, unsigned long access_rights);++Before a pkey can be used, it must first be allocated with+pkey_alloc(). An application may either call pkey_set() or the+WRPKRU instruction directly in order to change access permissions+to memory covered with a key.++ int real_prot = PROT_READ|PROT_WRITE;+ pkey = pkey_alloc(0, PKEY_DENY_WRITE);+ ptr = mmap(NULL, PAGE_SIZE, PROT_NONE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0);+ ret = pkey_mprotect(ptr, PAGE_SIZE, real_prot, pkey);+ ... application runs here++Now, if the application needs to update the data at 'ptr', it can+gain access, do the update, then remove its write access:++ pkey_set(pkey, 0); // clear PKEY_DENY_WRITE+ *ptr = foo; // assign something+ pkey_set(pkey, PKEY_DENY_WRITE); // set PKEY_DENY_WRITE again++Now when it frees the memory, it will also free the pkey since it+is no longer in use:++ munmap(ptr, PAGE_SIZE);+ pkey_free(pkey);++=========================== Behavior ===========================++The kernel attempts to make protection keys consistent with the+behavior of a plain mprotect(). For instance if you do this:++ mprotect(ptr, size, PROT_NONE);+ something(ptr);++you can expect the same effects with protection keys when doing this:++ sys_pkey_alloc(0, PKEY_DISABLE_WRITE | PKEY_DISABLE_READ);+ sys_pkey_mprotect(ptr, size, PROT_READ|PROT_WRITE);+ something(ptr);++That should be true whether something() is a direct access to 'ptr'+like:++ *ptr = foo;++or when the kernel does the access on the application's behalf like+with a read():++ read(fd, ptr, 1);++The kernel will send a SIGSEGV in both cases, but si_code will be set+to SEGV_PKERR when violating protection keys versus SEGV_ACCERR when+the plain mprotect() permissions are violated.+ =========================== Config Option =========================== This config option adds approximately 1.5kb of text. and 50 bytes of
_
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Dave Hansen <hidden> Date: 2016-02-23 01:12:00
From: Dave Hansen <dave.hansen@linux.intel.com>
This establishes two more system calls for protection key management:
unsigned long pkey_get(int pkey);
int pkey_set(int pkey, unsigned long access_rights);
The return value from pkey_get() and the 'access_rights' passed
to pkey_set() are the same format: a bitmask containing
PKEY_DENY_WRITE and/or PKEY_DENY_ACCESS, or nothing set at all.
These can replace userspace's direct use of the new rdpkru/wrpkru
instructions.
With current hardware, the kernel can not enforce that it has
control over a given key. But, this at least allows the kernel
to indicate to userspace that userspace does not control a given
protection key. This makes it more likely that situations like
using a pkey after sys_pkey_free() can be detected.
The kernel does _not_ enforce that this interface must be used for
changes to PKRU, whether or not a key has been "allocated".
This syscall interface could also theoretically be replaced with a
pair of vsyscalls. The vsyscalls would just call WRPKRU/RDPKRU
directly in situations where they are drop-in equivalents for
what the kernel would be doing.
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: linux-api@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: x86@kernel.org
Cc: torvalds@linux-foundation.org
Cc: akpm@linux-foundation.org
---
b/arch/x86/entry/syscalls/syscall_32.tbl | 2 +
b/arch/x86/entry/syscalls/syscall_64.tbl | 2 +
b/arch/x86/include/asm/pkeys.h | 4 +-
b/arch/x86/kernel/fpu/xstate.c | 55 +++++++++++++++++++++++++++++--
b/include/linux/pkeys.h | 8 ++++
b/mm/mprotect.c | 41 +++++++++++++++++++++++
6 files changed, 109 insertions(+), 3 deletions(-)
diff -puN arch/x86/entry/syscalls/syscall_32.tbl~pkeys-87-syscalls-set-get arch/x86/entry/syscalls/syscall_32.tbl
@@ -336,6 +336,8 @@ 327 common pkey_mprotect sys_pkey_mprotect 328 common pkey_alloc sys_pkey_alloc 329 common pkey_free sys_pkey_free+330 common pkey_get sys_pkey_get+331 common pkey_set sys_pkey_set # # x32-specific system call numbers start at 512 to avoid cache impact
@@ -879,6 +880,9 @@ int __arch_set_user_pkey_access(struct tintpkey_shift=(pkey*PKRU_BITS_PER_PKEY);u32new_pkru_bits=0;+/* Only support manipulating current task for now */+if(tsk!=current)+return-EINVAL;/**ThischeckimpliesXSAVEsupport.OSPKEonlygets*setifweenableXSAVEandweenablePKUinXCR0.
@@ -904,7 +908,7 @@ int __arch_set_user_pkey_access(struct t*state.*/if(!old_pkru_state)-new_pkru_state.pkru=0;+new_pkru_state.pkru=PKRU_INIT_STATE;elsenew_pkru_state.pkru=old_pkru_state->pkru;
@@ -942,4 +946,51 @@ int arch_set_user_pkey_access(struct tasreturn-EINVAL;return__arch_set_user_pkey_access(tsk,pkey,init_val);}++/*+*Figuresoutwhattherightsarecurrentlyfor'pkey'.+*ConvertsfromPKRU'sformattotheuser-visiblePKEY_DISABLE_*+*format.+*/+unsignedlongarch_get_user_pkey_access(structtask_struct*tsk,intpkey)+{+structfpu*fpu=¤t->thread.fpu;+u32pkru_reg;+intret=0;++/* Only support manipulating current task for now */+if(tsk!=current)+return-1;+if(!boot_cpu_has(X86_FEATURE_OSPKE))+return-1;+/*+*ThecontentsofPKRUitselfareinvalid.Consultthe+*task'sXSAVEbufferforPKRUcontents.Thisismuch+*moreexpensivethanreadingPKRUdirectly,butshould+*berareorimpossiblewitheagerfpumode.+*/+if(!fpu->fpregs_active){+structxregs_state*xsave=&fpu->state.xsave;+structpkru_state*pkru_state=+get_xsave_addr(xsave,XFEATURE_MASK_PKRU);+/*+*PKRUisinitsinitstateandnotpresentin+*thebufferinasavedform.+*/+if(!pkru_state)+returnPKRU_INIT_STATE;++returnpkru_state->pkru;+}+/*+*Consulttheuserregisterdirectly.+*/+pkru_reg=read_pkru();+if(!__pkru_allows_read(pkru_reg,pkey))+ret|=PKEY_DISABLE_ACCESS;+if(!__pkru_allows_write(pkru_reg,pkey))+ret|=PKEY_DISABLE_WRITE;++returnret;+}#endif /* CONFIG_ARCH_HAS_PKEYS */
_
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Dave Hansen <hidden> Date: 2016-02-23 01:12:44
From: Dave Hansen <dave.hansen@linux.intel.com>
This is all that we need to get the new system call itself
working on x86.
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: linux-api@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: x86@kernel.org
Cc: torvalds@linux-foundation.org
Cc: akpm@linux-foundation.org
---
b/arch/x86/entry/syscalls/syscall_32.tbl | 1 +
b/arch/x86/entry/syscalls/syscall_64.tbl | 1 +
2 files changed, 2 insertions(+)
diff -puN arch/x86/entry/syscalls/syscall_32.tbl~pkeys-85b-x86-mprotect_key arch/x86/entry/syscalls/syscall_32.tbl
@@ -333,6 +333,7 @@ 324 common membarrier sys_membarrier 325 common mlock2 sys_mlock2 326 common copy_file_range sys_copy_file_range+327 common pkey_mprotect sys_pkey_mprotect # # x32-specific system call numbers start at 512 to avoid cache impact
_
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Dave Hansen <dave.hansen-VuQAYsv1563Yd54FQh9/CA@public.gmane.org>
This spells out all of the pkey-related system calls that we have
and provides some example code fragments to demonstrate how we
expect them to be used.
Signed-off-by: Dave Hansen <dave.hansen-VuQAYsv1563Yd54FQh9/CA@public.gmane.org>
Cc: linux-api-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
Cc: linux-mm-Bw31MaZKKs3YtjvyW6yDsg@public.gmane.org
Cc: x86-DgEjT+Ai2ygdnm+yROfE0A@public.gmane.org
Cc: torvalds-de/tnXTf+JLsfHDXvbKv3WD2FQJk+8+b@public.gmane.org
Cc: akpm-de/tnXTf+JLsfHDXvbKv3WD2FQJk+8+b@public.gmane.org
---
b/Documentation/x86/protection-keys.txt | 63 ++++++++++++++++++++++++++++++++
1 file changed, 63 insertions(+)
Please also add pkeys testcases to tools/tests/self-tests.
Thanks,
Ingo
From: Dave Hansen <dave.hansen@linux.intel.com>
This establishes two more system calls for protection key management:
unsigned long pkey_get(int pkey);
int pkey_set(int pkey, unsigned long access_rights);
The return value from pkey_get() and the 'access_rights' passed
to pkey_set() are the same format: a bitmask containing
PKEY_DENY_WRITE and/or PKEY_DENY_ACCESS, or nothing set at all.
These can replace userspace's direct use of the new rdpkru/wrpkru
instructions.
With current hardware, the kernel can not enforce that it has
control over a given key. But, this at least allows the kernel
to indicate to userspace that userspace does not control a given
protection key. This makes it more likely that situations like
using a pkey after sys_pkey_free() can be detected.
So it's analogous to file descriptor open()/close() syscalls: the kernel does not
enforce that different libraries of the same process do not interfere with each
other's file descriptors - but in practice it's not a problem because everyone
uses open()/close().
Resources that a process uses don't per se 'need' kernel level isolation to be
useful.
The kernel does _not_ enforce that this interface must be used for
changes to PKRU, whether or not a key has been "allocated".
Nor does the kernel enforce that open() must be used to get a file descriptor, so
code can do the following:
close(100);
and can interfere with a library that is holding a file open - but it's generally
not a problem and the above is considered poor code that will cause problems.
One thing that is different is that file descriptors are generally plentiful,
while of pkeys there are at most 16 - but I think it's still "large enough" to not
be an issue in practice.
We'll see ...
This syscall interface could also theoretically be replaced with a pair of
vsyscalls. The vsyscalls would just call WRPKRU/RDPKRU directly in situations
where they are drop-in equivalents for what the kernel would be doing.
Indeed.
Thanks,
Ingo
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Michael Kerrisk (man-pages) <hidden> Date: 2016-03-03 08:05:32
Hi Dave,
On 23 February 2016 at 02:11, Dave Hansen [off-list ref] wrote:
As promised, here are the proposed new Memory Protection Keys
interfaces. These interfaces make it possible to do something
with pkeys other than execute-only support.
There are 5 syscalls here. I'm hoping for reviews of this set
which can help nail down what the final interfaces will be.
You can find a high-level overview of the feature and the new
syscalls here:
https://www.sr71.net/~dave/intel/pkeys.txt
(That's pretty thin...)
===============================================================
To use memory protection keys (pkeys), an application absolutely
needs to be able to set the pkey field in the PTE (obviously has
to be done in-kernel) and make changes to the "rights" register
(using unprivileged instructions).
An application also needs to have an an allocator for the keys
themselves. If two different parts of an application both want
to protect their data with pkeys, they first need to know which
key to use for their individual purposes.
This set introduces 5 system calls, in 3 logical groups:
1. PTE pkey setting (sys_pkey_mprotect(), patches #1-3)
2. Key allocation (sys_pkey_alloc() / sys_pkey_free(), patch #4)
3. Rights register manipulation (sys_pkey_set/get(), patch #5)
These patches build on top of "core" support already in the tip tree,
specifically 62b5f7d013, which can currently be found at:
http://git.kernel.org/cgit/linux/kernel/git/tip/tip.git/log/?h=mm/pkeys
I have manpages written for some of these syscalls, and I will
submit a full set of manpages once we've reached some consensus
on what the interfaces should be.
Please don't do things in this order. Providing man pages up front
make it easier for people to understand, review, and critique the API.
Submitting man pages should be a foundational part of submitting a new
set of interfaces and discussing their design.
Thanks,
Michael
From: Dave Hansen <hidden> Date: 2016-03-03 23:49:15
On 03/03/2016 12:05 AM, Michael Kerrisk (man-pages) wrote:
quoted
quoted
I have manpages written for some of these syscalls, and I will
submit a full set of manpages once we've reached some consensus
on what the interfaces should be.
Please don't do things in this order. Providing man pages up front
make it easier for people to understand, review, and critique the API.
Submitting man pages should be a foundational part of submitting a new
set of interfaces and discussing their design.
Michael, thanks for taking a look, plus the very detailed previous
review you did of the first batch of man-pages that I posted.
I've posted a newer version including all of the new system calls, and
I've attempted to address all the earlier review comments you made.
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>