Thread (71 messages) 71 messages, 6 authors, 21d ago

Re: [PATCH RFC v9 00/25] pkeys-based page table hardening

From: Kevin Brodsky <hidden>
Date: 2026-09-08 07:57:00
Also in: linux-hardening, linux-mm

On 07/09/2026 14:19, Linu Cherian wrote:
On Thu, Sep 03, 2026 at 06:47:50PM +0200, Kevin Brodsky wrote:
quoted
quoted
quoted
[...]

kpkeys
======

The use of pkeys involves two separate mechanisms: assigning a pkey to
pages, and defining the pkeys -> permissions mapping via the pkey
register. This is implemented through the following interface:

- Pages are assigned a pkey in the linear map using set_memory_pkey().
  This is sufficient for this series, but it is also plausible for
  higher-level allocators to support marking allocations with a given
  pkey.

- The pkey register is configured based on a *kpkeys context*. kpkeys
  contexts are represented as simple integers that correspond to a given
  configuration, for instance:

  KPKEYS_CTX_DEFAULT:
        RW access to KPKEYS_PKEY_DEFAULT
        RO access to any other KPKEYS_PKEY_*

  KPKEYS_CTX_<FEAT>:
        RW access to KPKEYS_PKEY_DEFAULT
        RW access to KPKEYS_PKEY_<FEAT>
        RO access to any other KPKEYS_PKEY_*

  Only pkeys that are managed by the kpkeys framework are impacted;
  permissions for other pkeys are left unchanged (this allows for other
  schemes using pkeys to be used in parallel, and arch-specific use of
  certain pkeys).
- Adding some basic details on what a scheme and context is quite helpful.

 - Giving some hints (may be an example) on how multiple schemes and multiple contexts
  play together would be quite helpful.
"scheme" doesn't mean anything precise, it's only the notion that pkeys
that aren't reserved for kpkeys (i.e. anything but 0 or 1 in this
series) may be used for other purposes. Happy to reword if you have a
suggestion.
Got it. IMHO, adding two definitions towards the start would make it easier to follow.

kpkeys: Set of pkeys reserved and managed by the kpkeys framework.
        Pkeys outside this set are left untouched.

kpkeys context: A permission state that defines the permissions for each pkey owned by
 	kpkeys

Or something better.
Got it, will add something along those lines, thanks!
quoted
quoted
[...]
quoted
Open questions
==============

A few aspects in this RFC that are debatable and/or worth discussing:

- There is currently no restriction on how kpkeys contexts map to pkeys
  permissions. A typical approach is to allocate one pkey per context and
  make it writable in that context only. As the number of contexts
Probably to avoid the assumption, may be we can we have something like
below 

For a pkey P, we could define
PKEY_P_PERM_CTXT_OTHERS	 //permission for pkey p in other contexts
PKEY_P_PERM_CTXT_SELF	 //permission for pkey p in self context

With the assumption of one pkey mapped for every context,
the permission for the default context would look something like,

PKEY_DEF_PERM_CTXT_SELF << PKEY_DEF_PKEY_SHIFT |
PKEY_CT0_PERM_CTXT_OTHERS << PKEY_CT0_PKEY_SHIFT | 
PKEY_CT1_PERM_CTXT_OTHERS << PKEY_CT1_PKEY_SHIFT |
...(for all valid contexts)

where,
Permission key, PKEY_DEF is associated with context DEFAULT,
Permission key, PKEY_CT0 is associated with context CT0,
Permission key, PKEY_CT1 is associated with context CT1
This adds assumptions rather than avoiding them. *Typically* when adding
a context you'd allocate a pkey that's only writable by this context,
but it doesn't have to be this way.
Okay agree. Then may be something like

Define permissions:

For default context,
KPKEYS_CTX_DEFAULT_PERM_PKEY_DEF
KPKEYS_CTX_DEFAULT_PERM_PKEY_CT0

For CT0 context,
KPKEYS_CTX_CT0_PERM_PKEY_DEF
KPKEYS_CTX_CT0_PERM_PKEY_CT0

Define POR_EL1:

For default context,
KPKEYS_POR_EL1_DEFAULT

For CT0 context,
KPKEYS_POR_EL1_CT0

Finally,
#define POR_EL1_INIT KPKEYS_POR_EL1_DEFAULT
We cannot do this because POR_EL1 is arm64-specific and its format is
not at all the same as x86's PKRS for instance.

I think what you're getting at is that the permissions for each pkeys in
a given context could be defined at the generic level. This could be
done, but I'm not sure this is essential, and we may not need all archs
to use exactly the same permissions. There's also the issue that x86
only encodes RW permissions directly, not X.
Probably using something similar would make the idea of kpkeys context 
more evident in the code as well ?
quoted
The configuration space is more easily understood by considering the
other use-cases we've investigated (struct cred protection and eBPF
isolation, linked further down). For instance, for cred protection, we
had KPKEYS_LVL_UNRESTRICTED with write access to all pkeys, and for eBPF
isolation, we need a level that is less privileged and therefore does
*not* have write access to pkey 0.
quoted
quoted
  increases, we may however run out of pkeys, especially on arm64 (just
  8 pkeys with POE). Depending on the use-cases, it may be acceptable to
  use the same pkey for the data associated to multiple contexts.
Lets say two contexts A and B, use the same pkey P as their permission matches.
But then, when we enter context A, permission for pkey P gets
relaxed, then that would relax permission for pages associated with
context B as well which is unintended ?
That may be exactly what is intended, it all depends on the use-case. C1
may have a private pkey P1, and C2 P2, and then P3 that is shared by C1
and C2 (writable by both)
Got it. With each context defining permissions for each pkey owned by
kpkeys makes sense. 

Also do we need to assume that nesting of different contexts is not valid ?
For example,
Default context:
	enter CTX 0
		enter CTX 1
		leave CTX 1
	leave CTX 0
Default context:
That is a good question. The enter/leave logic does support nesting,
since leave() restores the pkeys register as it was on enter(), but
whether the inner context has sufficient permissions will depend on the
situation. Certainly if nesting is expected then it has to be taken into
account when defining the permissions for each context.

- Kevin
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help