From: Dave Hansen <hidden> Date: 2015-09-28 19:25:30
I have addressed all known issues and review comments. I believe
they are ready to be pulled in to the x86 tree. Note that this
is also the first time anyone has seen the new 'selftests' code.
If there are issues limited to it, I'd prefer to fix those up
separately post-merge.
Changes from RFCv2 (Thanks Ingo and Thomas for most of these):
* few minor compile warnings
* changed 'nopku' interaction with cpuid bits. Now, we do not
clear the PKU cpuid bit, we just skip enabling it.
* changed __pkru_allows_write() to also check access disable bit
* removed the unused write_pkru()
* made si_pkey a u64 and added some patch description details.
Also made it share space in siginfo with MPX and clarified
comments.
* give some real text for the Processor Trace xsave state
* made vma_pkey() less ugly (and much more optimized actually)
* added SEGV_PKUERR to copy_siginfo_to_user()
* remove page table walk when filling in si_pkey, added some
big fat comments about it being inherently racy.
* added self test code
MM reviewers, if you are going to look at one thing, please look
at patch 14 which adds a bunch of additional vma/pte permission
checks.
This code contains a new system call: mprotect_key(), This needs
the usual amount of rigor around new interfaces. Review there
would be much appreciated.
This code is not runnable to anyone outside of Intel unless they
have some special hardware or a fancy simulator. If you are
interested in running this for real, please get in touch with me.
Hardware is available to a very small but nonzero number of
people.
This set is also available here (with the new syscall):
git://git.kernel.org/pub/scm/linux/kernel/git/daveh/x86-pkeys.git pkeys-v006
=== diffstat ===
(note that over half of this is kselftests)
Documentation/kernel-parameters.txt | 3
Documentation/x86/protection-keys.txt | 54 +
arch/powerpc/include/asm/mman.h | 5
arch/powerpc/include/asm/mmu_context.h | 11
arch/s390/include/asm/mmu_context.h | 11
arch/unicore32/include/asm/mmu_context.h | 11
arch/x86/Kconfig | 15
arch/x86/entry/syscalls/syscall_32.tbl | 1
arch/x86/entry/syscalls/syscall_64.tbl | 1
arch/x86/include/asm/cpufeature.h | 54 +
arch/x86/include/asm/disabled-features.h | 12
arch/x86/include/asm/fpu/types.h | 16
arch/x86/include/asm/fpu/xstate.h | 4
arch/x86/include/asm/mmu_context.h | 71 ++
arch/x86/include/asm/pgtable.h | 45 +
arch/x86/include/asm/pgtable_types.h | 34 -
arch/x86/include/asm/required-features.h | 4
arch/x86/include/asm/special_insns.h | 32 +
arch/x86/include/uapi/asm/mman.h | 23
arch/x86/include/uapi/asm/processor-flags.h | 2
arch/x86/kernel/cpu/common.c | 42 +
arch/x86/kernel/fpu/xstate.c | 7
arch/x86/kernel/process_64.c | 2
arch/x86/kernel/setup.c | 9
arch/x86/mm/fault.c | 143 +++-
arch/x86/mm/gup.c | 37 -
drivers/char/agp/frontend.c | 2
drivers/staging/android/ashmem.c | 9
fs/proc/task_mmu.c | 5
include/asm-generic/mm_hooks.h | 11
include/linux/mm.h | 13
include/linux/mman.h | 6
include/uapi/asm-generic/siginfo.h | 17
kernel/signal.c | 4
mm/Kconfig | 11
mm/gup.c | 28
mm/memory.c | 4
mm/mmap.c | 2
mm/mprotect.c | 20
mm/nommu.c | 2
tools/testing/selftests/x86/Makefile | 3
tools/testing/selftests/x86/pkey-helpers.h | 182 +++++
tools/testing/selftests/x86/protection_keys.c | 828 ++++++++++++++++++++++++++
43 files changed, 1705 insertions(+), 91 deletions(-)
Cc: linux-api@vger.kernel.org
Cc: linux-arch@vger.kernel.org
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Dave Hansen <hidden> Date: 2015-09-28 19:18:31
From: Dave Hansen <dave.hansen-VuQAYsv1563Yd54FQh9/CA@public.gmane.org>
This is all that we need to get the new system call itself
working on x86.
Signed-off-by: Dave Hansen <dave.hansen-VuQAYsv1563Yd54FQh9/CA@public.gmane.org>
Cc: linux-api-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
---
b/arch/x86/entry/syscalls/syscall_32.tbl | 1 +
b/arch/x86/entry/syscalls/syscall_64.tbl | 1 +
b/arch/x86/include/uapi/asm/mman.h | 7 +++++++
b/mm/Kconfig | 1 +
4 files changed, 10 insertions(+)
diff -puN arch/x86/entry/syscalls/syscall_32.tbl~pkeys-16-x86-mprotect_key arch/x86/entry/syscalls/syscall_32.tbl
@@ -689,4 +689,5 @@ config NR_PROTECTION_KEYS# Everything supports a _single_ key, so allow folks to# at least call APIs that take keys, but require that the# key be 0.+default16ifX86_INTEL_MEMORY_PROTECTION_KEYSdefault1
From: Dave Hansen <hidden> Date: 2015-09-28 19:20:30
From: Dave Hansen <dave.hansen@linux.intel.com>
mprotect_key() is just like mprotect, except it also takes a
protection key as an argument. On systems that do not support
protection keys, it still works, but requires that key=0.
Otherwise it does exactly what mprotect does.
I expect it to get used like this, if you want to guarantee that
any mapping you create can *never* be accessed without the right
protection keys set up.
pkey_deny_access(11); // random pkey
int real_prot = PROT_READ|PROT_WRITE;
ptr = mmap(NULL, PAGE_SIZE, PROT_NONE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0);
ret = mprotect_key(ptr, PAGE_SIZE, real_prot, 11);
This way, there is *no* window where the mapping is accessible
since it was always either PROT_NONE or had a protection key set.
We settled on 'unsigned long' for the type of the key here. We
only need 4 bits on x86 today, but I figured that other
architectures might need some more space.
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: linux-api@vger.kernel.org
---
b/mm/Kconfig | 7 +++++++
b/mm/mprotect.c | 20 +++++++++++++++++---
2 files changed, 24 insertions(+), 3 deletions(-)
diff -puN mm/Kconfig~pkeys-85-mprotect_pkey mm/Kconfig
@@ -683,3 +683,10 @@ config FRAME_VECTORconfigARCH_USES_HIGH_VMA_FLAGSbool++configNR_PROTECTION_KEYS+int+# Everything supports a _single_ key, so allow folks to+# at least call APIs that take keys, but require that the+# key be 0.+default1
From: Dave Hansen <hidden> Date: 2015-09-28 19:21:11
From: Dave Hansen <dave.hansen@linux.intel.com>
This plumbs a protection key through calc_vm_flag_bits().
We could of done this in calc_vm_prot_bits(), but I did not
feel super strongly which way to go. It was pretty arbitrary
which one to use.
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: linux-api@vger.kernel.org
Cc: linux-arch@vger.kernel.org
---
b/arch/powerpc/include/asm/mman.h | 5 +++--
b/drivers/char/agp/frontend.c | 2 +-
b/drivers/staging/android/ashmem.c | 9 +++++----
b/include/linux/mman.h | 6 +++---
b/mm/mmap.c | 2 +-
b/mm/mprotect.c | 2 +-
b/mm/nommu.c | 2 +-
7 files changed, 15 insertions(+), 13 deletions(-)
diff -puN arch/powerpc/include/asm/mman.h~pkeys-84-calc_vm_prot_bits arch/powerpc/include/asm/mman.h
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2015-09-29 06:39:51
On Mon, 2015-09-28 at 12:18 -0700, Dave Hansen wrote:
From: Dave Hansen <dave.hansen-VuQAYsv1563Yd54FQh9/CA@public.gmane.org>
mprotect_key() is just like mprotect, except it also takes a
protection key as an argument. On systems that do not support
protection keys, it still works, but requires that key=0.
I'm not sure how userspace is going to use the key=0 feature? ie. userspace
will still have to detect that keys are not supported and use key 0 everywhere.
At that point it could just as well skip the mprotect_key() syscalls entirely
couldn't it?
I expect it to get used like this, if you want to guarantee that
any mapping you create can *never* be accessed without the right
protection keys set up.
pkey_deny_access(11); // random pkey
int real_prot = PROT_READ|PROT_WRITE;
ptr = mmap(NULL, PAGE_SIZE, PROT_NONE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0);
ret = mprotect_key(ptr, PAGE_SIZE, real_prot, 11);
This way, there is *no* window where the mapping is accessible
since it was always either PROT_NONE or had a protection key set.
We settled on 'unsigned long' for the type of the key here. We
only need 4 bits on x86 today, but I figured that other
architectures might need some more space.
If the existing mprotect() syscall had a flags argument you could have just
used that. So is it worth just adding mprotect2() now and using it for this? ie:
int mprotect2(unsigned long start, size_t len, unsigned long prot, unsigned long flags) ..
And then you define bit zero of flags to say you're passing a pkey, and it's in
bits 1-63?
That way if other arches need to do something different you at least have the
flags available?
cheers
From: Dave Hansen <hidden> Date: 2015-09-29 14:16:48
On 09/28/2015 11:39 PM, Michael Ellerman wrote:
On Mon, 2015-09-28 at 12:18 -0700, Dave Hansen wrote:
quoted
From: Dave Hansen <dave.hansen@linux.intel.com>
mprotect_key() is just like mprotect, except it also takes a
protection key as an argument. On systems that do not support
protection keys, it still works, but requires that key=0.
I'm not sure how userspace is going to use the key=0 feature? ie. userspace
will still have to detect that keys are not supported and use key 0 everywhere.
At that point it could just as well skip the mprotect_key() syscalls entirely
couldn't it?
Yep.
Or, a new architecture could just skip mprotect() itself entirely and
only wire up mprotect_pkey(). I don't see this pkey=0 thing as an
important feature or anything. I just wanted to call out the behavior.
quoted
I expect it to get used like this, if you want to guarantee that
any mapping you create can *never* be accessed without the right
protection keys set up.
pkey_deny_access(11); // random pkey
int real_prot = PROT_READ|PROT_WRITE;
ptr = mmap(NULL, PAGE_SIZE, PROT_NONE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0);
ret = mprotect_key(ptr, PAGE_SIZE, real_prot, 11);
This way, there is *no* window where the mapping is accessible
since it was always either PROT_NONE or had a protection key set.
We settled on 'unsigned long' for the type of the key here. We
only need 4 bits on x86 today, but I figured that other
architectures might need some more space.
If the existing mprotect() syscall had a flags argument you could have just
used that. So is it worth just adding mprotect2() now and using it for this? ie:
int mprotect2(unsigned long start, size_t len, unsigned long prot, unsigned long flags) ..
And then you define bit zero of flags to say you're passing a pkey, and it's in
bits 1-63?
That way if other arches need to do something different you at least have the
flags available?
But what problem does that solve?
mprotect() itself has plenty of space in prot. Do any of the other
architectures need to pass in more than just an integer key to implement
storage/protection keys?
I'd much rather have a set of (relatively) arch-specific system calls
implementing protection keys rather than a single one with one
arch-specific argument.