From: Maxime Bélair <hidden> Date: 2025-05-06 14:40:07
This patchset introduces a new syscall, lsm_manage_policy(), and the
associated Linux Security Module hook security_lsm_manage_policy(),
providing a unified interface for loading and managing LSM policies.
This syscall complements the existing per‑LSM pseudo‑filesystem mechanism
and works even when those filesystems are not mounted or available.
With this new syscall, administrators may lock down access to the
pseudo‑filesystem yet still manage LSM policies. A single, tightly scoped
entry point then replaces the many file operations exposed by those
filesystems, significantly reducing the attack surface. This is
particularly useful in containers or processes already confined by
Landlock, where these pseudo‑filesystems are typically unavailable.
Because it provides a logical and unified interface, lsm_manage_policy()
is simpler to use than several heterogeneous pseudo‑filesystems and
avoids edge cases such as partially loaded policies. It also eliminates
VFS overhead, yielding performance gains notably when many policies are
loaded, for instance at boot time.
This initial implementation is intentionally minimal to limit the scope
of changes. Currently, only policy loading is supported, and only
AppArmor registers this LSM hook. However, any LSM can adopt this
interface, and future patches could extend this syscall to support more
operations, such as replacing, removing, or querying loaded policies.
Landlock already provides three Landlock‑specific syscalls (e.g.
landlock_add_rule()) to restrict ambient rights for sets of processes
without touching any pseudo-filesystem. lsm_manage_policy() generalizes
that approach to the entire LSM layer, so any module can expose its
policy operations through one uniform interface and reap the advantages
outlined above.
This patchset is available at [1] and a minimal user space example
showing how to use this syscall with AppArmor is at [2].
[1] https://github.com/emixam16/linux/tree/lsm_syscall
[2] https://gitlab.com/emixam16/apparmor/tree/lsm_syscall
Maxime Bélair (3):
Wire up the lsm_manage_policy syscall
lsm: introduce security_lsm_manage_policy hook
AppArmor: add support for lsm_manage_policy
arch/alpha/kernel/syscalls/syscall.tbl | 1 +
arch/arm/tools/syscall.tbl | 1 +
arch/x86/entry/syscalls/syscall_32.tbl | 1 +
arch/x86/entry/syscalls/syscall_64.tbl | 1 +
include/linux/lsm_hook_defs.h | 2 ++
include/linux/security.h | 7 +++++++
include/linux/syscalls.h | 4 ++++
include/uapi/asm-generic/unistd.h | 4 +++-
include/uapi/linux/lsm.h | 8 +++++++
kernel/sys_ni.c | 1 +
security/apparmor/apparmorfs.c | 19 +++++++++++++++++
security/apparmor/include/apparmorfs.h | 3 +++
security/apparmor/lsm.c | 16 ++++++++++++++
security/lsm_syscalls.c | 11 ++++++++++
security/security.c | 21 +++++++++++++++++++
tools/include/uapi/asm-generic/unistd.h | 4 +++-
.../arch/x86/entry/syscalls/syscall_64.tbl | 1 +
17 files changed, 103 insertions(+), 2 deletions(-)
base-commit: 9c32cda43eb78f78c73aee4aa344b777714e259b
--
2.48.1
From: Maxime Bélair <hidden> Date: 2025-05-06 14:40:03
Enable users to manage AppArmor policies through the new hook
lsm_manage_policy. Currently, policies can be added but not replaced
using this new mechanism, ensuring that this interface can only further
confine the system.
Signed-off-by: Maxime Bélair <redacted>
---
security/apparmor/apparmorfs.c | 19 +++++++++++++++++++
security/apparmor/include/apparmorfs.h | 3 +++
security/apparmor/lsm.c | 16 ++++++++++++++++
3 files changed, 38 insertions(+)
@@ -1275,6 +1275,20 @@ static int apparmor_socket_shutdown(struct socket *sock, int how)returnaa_sock_perm(OP_SHUTDOWN,AA_MAY_SHUTDOWN,sock);}+staticintapparmor_lsm_manage_policy(u32lsm_id,u32op,void__user*buf,+size_tsize,u32flags)+{+loff_tpos=0;// Partial writing is not currently supported++if(lsm_id!=LSM_ID_APPARMOR)+return0;++if(op!=LSM_POLICY_LOAD||flags)+return-EOPNOTSUPP;++returnaa_profile_load_current_ns(buf,size,&pos);+}+#ifdef CONFIG_NETWORK_SECMARK/***apparmor_socket_sock_rcv_skb-checkpermsbeforeassociatingskbtosk
From: Maxime Bélair <hidden> Date: 2025-05-06 14:40:04
Add support for the new lsm_manage_policy syscall, providing a unified
API for loading and modifying LSM policies without requiring the LSM’s
pseudo-filesystem.
Benefits:
- Works even if the LSM pseudo-filesystem isn’t mounted or available
(e.g. in containers)
- Offers a logical and unified interface rather than multiple
heterogeneous pseudo-filesystems.
- Avoids overhead of other kernel interfaces for better efficiency
Signed-off-by: Maxime Bélair <redacted>
---
arch/alpha/kernel/syscalls/syscall.tbl | 1 +
arch/arm/tools/syscall.tbl | 1 +
arch/x86/entry/syscalls/syscall_32.tbl | 1 +
arch/x86/entry/syscalls/syscall_64.tbl | 1 +
include/linux/syscalls.h | 4 ++++
include/uapi/asm-generic/unistd.h | 4 +++-
kernel/sys_ni.c | 1 +
security/lsm_syscalls.c | 6 ++++++
tools/include/uapi/asm-generic/unistd.h | 4 +++-
tools/perf/arch/x86/entry/syscalls/syscall_64.tbl | 1 +
10 files changed, 22 insertions(+), 2 deletions(-)
@@ -507,3 +507,4 @@ 575 common listxattrat sys_listxattrat 576 common removexattrat sys_removexattrat 577 common open_tree_attr sys_open_tree_attr+578 common lsm_manage_policy sys_lsm_manage_policy
@@ -482,3 +482,4 @@ 465 common listxattrat sys_listxattrat 466 common removexattrat sys_removexattrat 467 common open_tree_attr sys_open_tree_attr+468 common lsm_manage_policy sys_lsm_manage_policy
@@ -391,6 +391,7 @@ 465 common listxattrat sys_listxattrat 466 common removexattrat sys_removexattrat 467 common open_tree_attr sys_open_tree_attr+468 common lsm_manage_policy sys_lsm_manage_policy # # Due to a historical design error, certain syscalls are numbered differently
@@ -391,6 +391,7 @@ 465 common listxattrat sys_listxattrat 466 common removexattrat sys_removexattrat 467 common open_tree_attr sys_open_tree_attr+468 common lsm_manage_policy sys_lsm_manage_policy # # Due to a historical design error, certain syscalls are numbered differently
From: Maxime Bélair <hidden> Date: 2025-05-06 14:40:06
Define a new LSM hook security_lsm_manage_policy and wire it into the
lsm_manage_policy() syscall so that LSMs can register a unified interface
for policy management. This initial, minimal implementation only supports
the LSM_POLICY_LOAD operation to limit changes.
Signed-off-by: Maxime Bélair <redacted>
---
include/linux/lsm_hook_defs.h | 2 ++
include/linux/security.h | 7 +++++++
include/uapi/linux/lsm.h | 8 ++++++++
security/lsm_syscalls.c | 7 ++++++-
security/security.c | 21 +++++++++++++++++++++
5 files changed, 44 insertions(+), 1 deletion(-)
From: Song Liu <song@kernel.org> Date: 2025-05-07 06:19:27
On Tue, May 6, 2025 at 7:40 AM Maxime Bélair
[off-list ref] wrote:
Define a new LSM hook security_lsm_manage_policy and wire it into the
lsm_manage_policy() syscall so that LSMs can register a unified interface
for policy management. This initial, minimal implementation only supports
the LSM_POLICY_LOAD operation to limit changes.
Signed-off-by: Maxime Bélair <redacted>
@@ -5883,6 +5883,27 @@ int security_bdev_setintegrity(struct block_device *bdev,}EXPORT_SYMBOL(security_bdev_setintegrity);+/**+*security_lsm_manage_policy()-ManagethepoliciesofLSMs+*@lsm_id:idofthelsmtotarget+*@op:Operationtoperform(oneoftheLSM_POLICY_XXXvalues)+*@buf:userspacepointertopolicydata+*@size:sizeof@buf+*@flags:lsmpolicymanagementflags+*+*ManagethepoliciesofaLSM.Thisnotablyallowstoupdatethemevenwhen+*thelsmfsisunavailableisrestricted.Currently,onlyLSM_POLICY_LOADis+*supported.+*+*Return:Returns0onsuccess,erroronfailure.+*/+intsecurity_lsm_manage_policy(u32lsm_id,u32op,void__user*buf,+size_tsize,u32flags)+{+returncall_int_hook(lsm_manage_policy,lsm_id,op,buf,size,flags);
If the LSM doesn't implement this hook, sys_lsm_manage_policy will return 0
for any inputs, right? This is gonna be so confusing for users.
Thanks,
Song
From: Song Liu <song@kernel.org> Date: 2025-05-07 06:26:33
On Tue, May 6, 2025 at 7:40 AM Maxime Bélair
[off-list ref] wrote:
Add support for the new lsm_manage_policy syscall, providing a unified
API for loading and modifying LSM policies without requiring the LSM’s
pseudo-filesystem.
Benefits:
- Works even if the LSM pseudo-filesystem isn’t mounted or available
(e.g. in containers)
- Offers a logical and unified interface rather than multiple
heterogeneous pseudo-filesystems.
These two do not feel like real benefits:
- Not working in containers is often not an issue, but a feature.
- One syscall cannot fit all use cases well...
- Avoids overhead of other kernel interfaces for better efficiency
.. and it is is probably less efficient, because everything need to
fit in the same API.
Overall, this set doesn't feel like a good change to me.
Thanks,
Song
syzbot will report user-controlled unbounded huge size memory allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this interface because
TOMOYO's policy is line-oriented ASCII text data where the destination is switched via
pseudo‑filesystem's filename; use of filename helps restricting which type of policy
can be manipulated by which process.
arch/riscv/include/asm/stacktrace.h:14:21: error: 'no_instrument_function' attribute applies only to functions
arch/riscv/include/asm/stacktrace.h:16:13: error: storage class specified for parameter 'dump_backtrace'
arch/riscv/include/asm/stacktrace.h:20:1: error: expected '=', ',', ';', 'asm' or '__attribute__' before '{' token
20 | {
| ^
In file included from arch/riscv/kernel/asm-offsets.c:17:
arch/riscv/include/asm/suspend.h:12:1: warning: empty declaration
12 | struct suspend_context {
| ^~~~~~
quoted
arch/riscv/include/asm/suspend.h:31:12: error: storage class specified for parameter 'in_suspend'
31 | extern int in_suspend;
| ^~~~~~~~~~
quoted
arch/riscv/kernel/asm-offsets.c:22:1: error: expected '=', ',', ';', 'asm' or '__attribute__' before '{' token
22 | {
| ^
include/linux/security.h:1607:12: error: old-style parameter declarations in prototyped function definition
1607 | static int security_lsm_manage_policy(u32 lsm_id, u32 op, void __user *buf,
| ^~~~~~~~~~~~~~~~~~~~~~~~~~
quoted
arch/riscv/kernel/asm-offsets.c:514: error: expected '{' at end of input
arch/riscv/kernel/asm-offsets.c:513:1: warning: no return statement in function returning non-void [-Wreturn-type]
From: Maxime Bélair <hidden> Date: 2025-05-07 15:37:27
On 5/7/25 08:26, Song Liu wrote:
On Tue, May 6, 2025 at 7:40 AM Maxime Bélair
[off-list ref] wrote:
quoted
Add support for the new lsm_manage_policy syscall, providing a unified
API for loading and modifying LSM policies without requiring the LSM’s
pseudo-filesystem.
Benefits:
- Works even if the LSM pseudo-filesystem isn’t mounted or available
(e.g. in containers)
- Offers a logical and unified interface rather than multiple
heterogeneous pseudo-filesystems.
These two do not feel like real benefits:
- One syscall cannot fit all use cases well...
This syscall is not intended to cover every case, nor to replace existing kernel
interfaces.
Each LSM can decide which operations it wants to support (if any). For example, when
loading policies, an LSM may choose to allow only policies that further restrict
privileges.
- Not working in containers is often not an issue, but a feature.
Indeed, using this syscall requires appropriate capabilities and will not permit
unprivileged containers to manage policies arbitrarily.
With this syscall, capability checks remain the responsibility of each LSM.
For instance, in the AppArmor patch, a profile can be loaded only if
aa_policy_admin_capable() succeeds (which requires CAP_MAC_ADMIN). Moreover, by design,
policies can be loaded only in the current namespace.
I see this syscall as a middle point between exposing the entire sysfs, creating a large
attack surface, and blocking everything.
Landlock’s existing syscalls already improve security by allowing processes to further
restrict their ambient rights while adding only a modest attack surface.
This syscall is a further step in that direction: it lets LSMs add restrictive policies
without requiring exposing every other interface.
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
quoted
- Avoids overhead of other kernel interfaces for better efficiency
.. and it is is probably less efficient, because everything need to
fit in the same API.
As shown below, the syscall can significantly improve the performance of policy management.
A more detailed benchmark is available in [1].
The following table presents the time required to load an AppArmor profile.
For every cell, the first value is the total time taken by aa-load, and the value in
parentheses is the time spent to load the policy in the kernel only (total - dry‑run).
Results are in microseconds and are averaged over 10 000 runs to reduce variance.
| t (µs) | syscall | pseudofs | Speedup |
|-----------|-------------|-------------|---------------|
| 1password | 4257 (1127) | 3333 (192) | x1.28 (x5.86) |
| Xorg | 6099 (2961) | 5167 (2020) | x1.18 (x1.47) |
If an LSM wants to allow several operations for a single LSM_POLICY_XXX it can multiplex a sub‑opcode in flags, and select the appropriate handler, this incurs negligible overhead.
Thanks,
Maxime
[1] https://gitlab.com/-/snippets/4840792
From: Maxime Bélair <hidden> Date: 2025-05-07 15:37:35
On 5/7/25 08:19, Song Liu wrote:
On Tue, May 6, 2025 at 7:40 AM Maxime Bélair
[off-list ref] wrote:
quoted
Define a new LSM hook security_lsm_manage_policy and wire it into the
lsm_manage_policy() syscall so that LSMs can register a unified interface
for policy management. This initial, minimal implementation only supports
the LSM_POLICY_LOAD operation to limit changes.
Signed-off-by: Maxime Bélair <redacted>
syzbot will report user-controlled unbounded huge size memory allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this interface because
TOMOYO's policy is line-oriented ASCII text data where the destination is switched via
pseudo‑filesystem's filename; use of filename helps restricting which type of policy
can be manipulated by which process.
First, like any LSM, TOMOYO is not obliged to implement every operation. It can simply
expose the one that makes sense for its use case. For instance, I don't think it needs an
equivalent of the manager interface.
If TOMOYO wants to support several sub‑operations, it can distinguish them with the
syscall’s flags parameter instead of filenames (as securityfs_if.c does today) and reuse
the code already employed by its pseudo‑fs, as in the AppArmor patch. Supporting this
syscall would therefore require only minimal changes.
Line‑oriented ASCII text is not a barrier, either. The syscall can pass that format just
fine. Because a typical TOMOYO line is very small, the performance gains from using the
syscall are actually greater. A brief benchmark is available in [1].
Thanks,
Maxime
[1] https://gitlab.com/-/snippets/4840792
syzbot will report user-controlled unbounded huge size memory allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this interface because
TOMOYO's policy is line-oriented ASCII text data where the destination is switched via
pseudo‑filesystem's filename ...
While Tetsuo's comment is limited to TOMOYO, I believe the argument
applies to a number of other LSMs as well. The reality is that there
is no one policy ideal shared across LSMs and that complicates things
like the lsm_manage_policy() proposal. I'm intentionally saying
"complicates" and not "prevents" because I don't want to flat out
reject something like this, but I think there needs to be a larger
discussion among the different LSM groups about what such an API
should look like. We may not need to get every LSM to support this
new API, but we need to get something that would work for a
significant majority and would be general/extensible enough that we
would expect it to work with the majority of future LSMs (as much as
we can predict the future anyway).
--
paul-moore.com
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
If the interpretation of "flags" argument varies across LSMs, it sounds like ioctl()'s
"cmd" argument. Also, there is prctl() which can already carry string-ish parameters
without involving open(). Why can't we use prctl() instead of lsm_manage_policy() ?
From: Song Liu <song@kernel.org> Date: 2025-05-08 06:07:07
On Wed, May 7, 2025 at 8:37 AM Maxime Bélair
[off-list ref] wrote:
[...]
quoted
These two do not feel like real benefits:
- One syscall cannot fit all use cases well...
This syscall is not intended to cover every case, nor to replace existing kernel
interfaces.
Each LSM can decide which operations it wants to support (if any). For example, when
loading policies, an LSM may choose to allow only policies that further restrict
privileges.
quoted
- Not working in containers is often not an issue, but a feature.
Indeed, using this syscall requires appropriate capabilities and will not permit
unprivileged containers to manage policies arbitrarily.
With this syscall, capability checks remain the responsibility of each LSM.
For instance, in the AppArmor patch, a profile can be loaded only if
aa_policy_admin_capable() succeeds (which requires CAP_MAC_ADMIN). Moreover, by design,
policies can be loaded only in the current namespace.
I see this syscall as a middle point between exposing the entire sysfs, creating a large
attack surface, and blocking everything.
Landlock’s existing syscalls already improve security by allowing processes to further
restrict their ambient rights while adding only a modest attack surface.
This syscall is a further step in that direction: it lets LSMs add restrictive policies
without requiring exposing every other interface.
I don't think a syscall makes the API more secure. If necessary, we can add
permission check to each pseudo file. The downside of the syscall, however,
is that all the permission checks are hard-coded in the kernel (except for
BPF LSM); while the sys admin can configure permissions of the pseudo
files in user space.
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
quoted
quoted
- Avoids overhead of other kernel interfaces for better efficiency
.. and it is is probably less efficient, because everything need to
fit in the same API.
As shown below, the syscall can significantly improve the performance of policy management.
A more detailed benchmark is available in [1].
The following table presents the time required to load an AppArmor profile.
For every cell, the first value is the total time taken by aa-load, and the value in
parentheses is the time spent to load the policy in the kernel only (total - dry‑run).
Results are in microseconds and are averaged over 10 000 runs to reduce variance.
| t (µs) | syscall | pseudofs | Speedup |
|-----------|-------------|-------------|---------------|
| 1password | 4257 (1127) | 3333 (192) | x1.28 (x5.86) |
| Xorg | 6099 (2961) | 5167 (2020) | x1.18 (x1.47) |
I am not sure the performance of loading security policies is on any
critical path.
The implementation calls the hook for each LSM, which is why I think the
syscall is not efficient.
Overall, I am still not convinced a syscall for all LSMs is needed. To
justify such
a syscall, I think we need to show that it is useful in multiple LSMs.
Also, if we
really want to have single set of APIs for all LSMs, we may also need
get_policy,
remove_policy, etc. This set as-is appears to be an incomplete design. The
implementation, with call_int_hook, is also problematic. It can easily
cause some
controversial behaviors.
Thanks,
Song
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-08 07:20:04
On 5/6/25 23:26, Song Liu wrote:
On Tue, May 6, 2025 at 7:40 AM Maxime Bélair
[off-list ref] wrote:
quoted
Add support for the new lsm_manage_policy syscall, providing a unified
API for loading and modifying LSM policies without requiring the LSM’s
pseudo-filesystem.
Benefits:
- Works even if the LSM pseudo-filesystem isn’t mounted or available
(e.g. in containers)
- Offers a logical and unified interface rather than multiple
heterogeneous pseudo-filesystems.
These two do not feel like real benefits:
- Not working in containers is often not an issue, but a feature.
and the LSM doesn't have to allow the syscall to function in a container
where appropriate. Its up to the LSM if the syscall is supported and
what kind of permissions are needed.
However having the ability to function in a container and not having to
mount securityfs, or procfs into a container. similar to what landlock
gets with its syscall can be beneficial.
- One syscall cannot fit all use cases well...
of course not, and for those other use cases new syscalls can be added.
quoted
- Avoids overhead of other kernel interfaces for better efficiency
.. and it is is probably less efficient, because everything need to
fit in the same API.
no not everything, just what fits into the syscall. Nor does an LSM
have to use the syscall it is still use what works for it.
This could be a little more efficient than the current fs interface
used by apparmor/selinux/smack but I don't think efficiency is going
to be a huge win for this.
Overall, this set doesn't feel like a good change to me.
Thanks,
Song
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-08 07:52:57
On 5/7/25 15:04, Tetsuo Handa wrote:
On 2025/05/08 0:37, Maxime Bélair wrote:
quoted
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
If the interpretation of "flags" argument varies across LSMs, it sounds like ioctl()'s
yes that does feel like ioctls(), on the other hand defining them at the LSM level won't
offer LSMs flexibility making it so the syscall covers fewer use cases. I am not opposed
to either, it just hashing out what people want, and what is acceptable.
"cmd" argument. Also, there is prctl() which can already carry string-ish parameters
without involving open(). Why can't we use prctl() instead of lsm_manage_policy() ?
prctl() can be used, I used it for the unprivileged policy demo. It has its own set of
problems. While LSM policy could be associated with the process doing the load/replacement
or what ever operation, it isn't necessarily tied to it. A lot of LSM policy is not
process specific making prctl() a poor fit.
prctl() requires allocating a global prctl()
prctl() are already being filtered/controlled by LSMs making them a poort fit for
use by an LSM in a stacking situation as it requires updating the policy of other
LSMs on the system. Yes seccomp can filter the syscall but that still is an easier
barrier to overcome than having to have instruction for how to allow your LSMs
prctl() in multiple LSMs.
Mickaël already argued the need for landlock to have syscalls. See
https://lore.kernel.org/lkml/20200511192156.1618284-7-mic@digikod.net/
and the numerous iterations before that.
Ideally those could have been LSM syscalls, with landlock leveraging them. AppArmor
is getting to where it has similar needs to landlock. Yes we can use ioctls, prctls,
netlink, the fs, etc. it doesn't mean that those are the best interfaces to do so,
and ideally any interface we use will be of benefit to some other LSMs in the future.
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-08 08:18:24
On 5/7/25 23:06, Song Liu wrote:
On Wed, May 7, 2025 at 8:37 AM Maxime Bélair
[off-list ref] wrote:
[...]
quoted
quoted
These two do not feel like real benefits:
- One syscall cannot fit all use cases well...
This syscall is not intended to cover every case, nor to replace existing kernel
interfaces.
Each LSM can decide which operations it wants to support (if any). For example, when
loading policies, an LSM may choose to allow only policies that further restrict
privileges.
quoted
- Not working in containers is often not an issue, but a feature.
Indeed, using this syscall requires appropriate capabilities and will not permit
unprivileged containers to manage policies arbitrarily.
With this syscall, capability checks remain the responsibility of each LSM.
For instance, in the AppArmor patch, a profile can be loaded only if
aa_policy_admin_capable() succeeds (which requires CAP_MAC_ADMIN). Moreover, by design,
policies can be loaded only in the current namespace.
I see this syscall as a middle point between exposing the entire sysfs, creating a large
attack surface, and blocking everything.
Landlock’s existing syscalls already improve security by allowing processes to further
restrict their ambient rights while adding only a modest attack surface.
This syscall is a further step in that direction: it lets LSMs add restrictive policies
without requiring exposing every other interface.
I don't think a syscall makes the API more secure. If necessary, we can add
It exposes a different attack surface. Requiring mounting of the fs to where it is visible
in the container, provides attack surface, and requires additional external configuration.
Then there is the whole issue of getting the various LSMs to allow another LSM in the
stack to be able manage its own policy.
permission check to each pseudo file. The downside of the syscall, however,
is that all the permission checks are hard-coded in the kernel (except for
The permission checks don't have to be hard coded. Each LSM can define how it handles
or manages the syscall. The default is that it isn't supported, but if an lsm decides
to support it, there is now reason that its policy can't determine the use of the
syscall.
BPF LSM); while the sys admin can configure permissions of the pseudo
files in user space.
Other LSMs also have policy that can control access to pseudo filesystems and
other resources. Again, the control doesn't have to be hard coded. And seccomp can
be used to block the syscall.
quoted
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
quoted
quoted
- Avoids overhead of other kernel interfaces for better efficiency
.. and it is is probably less efficient, because everything need to
fit in the same API.
As shown below, the syscall can significantly improve the performance of policy management.
A more detailed benchmark is available in [1].
The following table presents the time required to load an AppArmor profile.
For every cell, the first value is the total time taken by aa-load, and the value in
parentheses is the time spent to load the policy in the kernel only (total - dry‑run).
Results are in microseconds and are averaged over 10 000 runs to reduce variance.
| t (µs) | syscall | pseudofs | Speedup |
|-----------|-------------|-------------|---------------|
| 1password | 4257 (1127) | 3333 (192) | x1.28 (x5.86) |
| Xorg | 6099 (2961) | 5167 (2020) | x1.18 (x1.47) |
I am not sure the performance of loading security policies is on any
critical path.
generally speaking I agree, but I am also not going to turn down a
performance improvement either. Its a nice to have, but not a strong
argument for need.
The implementation calls the hook for each LSM, which is why I think the
syscall is not efficient.
it should only call the LSM identified by the lsmid in the call.
Overall, I am still not convinced a syscall for all LSMs is needed. To
justify such
its not needed by all LSMs, just a subset of them, and some nebulous
subset of potentially future LSMs that is entirely undefinable.
If we had had appropriate LSM syscalls landlock wouldn't have needed
to have landlock specific syscalls. Having another LSM go that route
feels wrong especially now that we have some LSM syscalls. If a
syscall is needed by an LSM its better to try hashing something out
that might have utility for multiple LSMs or at the very least,
potentially have utility in the future.
a syscall, I think we need to show that it is useful in multiple LSMs.
Also, if we
really want to have single set of APIs for all LSMs, we may also need
get_policy,
We are never going to get a single set of APIs for all LSMs. I will
settle for an api that has utility for a subset
remove_policy, etc. This set as-is appears to be an incomplete design. The
To have a complete design, there needs to be feedback and discussion
from multiple LSMs. This is a starting point.
implementation, with call_int_hook, is also problematic. It can easily
cause some> controversial behaviors.
agreed it shouldn't be doing a straight call_int_hook, it should only
call it against the lsm identified by the lsmid
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-08 08:20:44
On 5/7/25 08:37, Maxime Bélair wrote:
On 5/7/25 08:19, Song Liu wrote:
quoted
On Tue, May 6, 2025 at 7:40 AM Maxime Bélair
[off-list ref] wrote:
quoted
Define a new LSM hook security_lsm_manage_policy and wire it into the
lsm_manage_policy() syscall so that LSMs can register a unified interface
for policy management. This initial, minimal implementation only supports
the LSM_POLICY_LOAD operation to limit changes.
Signed-off-by: Maxime Bélair <redacted>
@@ -5883,6 +5883,27 @@ int security_bdev_setintegrity(struct block_device *bdev,}EXPORT_SYMBOL(security_bdev_setintegrity);+/**+*security_lsm_manage_policy()-ManagethepoliciesofLSMs+*@lsm_id:idofthelsmtotarget+*@op:Operationtoperform(oneoftheLSM_POLICY_XXXvalues)+*@buf:userspacepointertopolicydata+*@size:sizeof@buf+*@flags:lsmpolicymanagementflags+*+*ManagethepoliciesofaLSM.Thisnotablyallowstoupdatethemevenwhen+*thelsmfsisunavailableisrestricted.Currently,onlyLSM_POLICY_LOADis+*supported.+*+*Return:Returns0onsuccess,erroronfailure.+*/+intsecurity_lsm_manage_policy(u32lsm_id,u32op,void__user*buf,+size_tsize,u32flags)+{+returncall_int_hook(lsm_manage_policy,lsm_id,op,buf,size,flags);
If the LSM doesn't implement this hook, sys_lsm_manage_policy will return 0
for any inputs, right? This is gonna be so confusing for users.
Indeed, that was an oversight. It will return -EOPNOTSUPP in the next patch revision.
I think it needs to do more than that. I don't think this should call each LSM, the
infrastructure should filter it and only send it to the LSM identified by the lsm_id
syzbot will report user-controlled unbounded huge size memory allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this interface because
TOMOYO's policy is line-oriented ASCII text data where the destination is switched via
pseudo‑filesystem's filename; use of filename helps restricting which type of policy
can be manipulated by which process.
That is fine. But curious I am curious what the interface would look like to fit TOMOYO's
needs. I look at the current implementation as an opening discussion of what the syscall
should look like. I have no delusions that we are going to get something that will fit
all LSMs but without requirements, we won't be able to even attempt to hash something
better out.
syzbot will report user-controlled unbounded huge size memory allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this interface because
TOMOYO's policy is line-oriented ASCII text data where the destination is switched via
pseudo‑filesystem's filename ...
While Tetsuo's comment is limited to TOMOYO, I believe the argument
applies to a number of other LSMs as well. The reality is that there
is no one policy ideal shared across LSMs and that complicates things
like the lsm_manage_policy() proposal. I'm intentionally saying
"complicates" and not "prevents" because I don't want to flat out
reject something like this, but I think there needs to be a larger
discussion among the different LSM groups about what such an API
should look like. We may not need to get every LSM to support this
new API, but we need to get something that would work for a
significant majority and would be general/extensible enough that we
would expect it to work with the majority of future LSMs (as much as
we can predict the future anyway).
yep, I look at this is just a starting point for discussion. There
isn't going to be any discussion without some code, so here is a v1
that supports a single LSM let the bike shedding begin.
That is fine. But curious I am curious what the interface would look like to fit TOMOYO's
needs.
Stream (like "FILE *") with restart from the beginning (like rewind(fp)) support.
That is, the caller can read/write at least one byte at a time, and written data
is processed upon encountering '\n'.
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-08 14:44:47
On 5/8/25 05:55, Tetsuo Handa wrote:
On 2025/05/08 17:25, John Johansen wrote:
quoted
That is fine. But curious I am curious what the interface would look like to fit TOMOYO's
needs.
Stream (like "FILE *") with restart from the beginning (like rewind(fp)) support.
That is, the caller can read/write at least one byte at a time, and written data
is processed upon encountering '\n'.
that can be emulated within the current sycall, where the lsm maintains a buffer.
Are you asking to also read data back out as well, that could be added, but doing
a syscall per byte here or through the fs is going to have fairly high overhead.
Without understanding the requirement it would seem to me, that it would be
better to emulate that file buffer manipulation in userspace similar say C++
stringstreams, and then write the syscall when done.
That is fine. But curious I am curious what the interface would look like to fit TOMOYO's
needs.
Stream (like "FILE *") with restart from the beginning (like rewind(fp)) support.
That is, the caller can read/write at least one byte at a time, and written data
is processed upon encountering '\n'.
that can be emulated within the current sycall, where the lsm maintains a buffer.
That cannot be emulated, for there is no event that is automatically triggered when
the process terminates (i.e. implicit close() upon exit()) in order to release the
buffer the LSM maintains.
Are you asking to also read data back out as well, that could be added, but doing
a syscall per byte here or through the fs is going to have fairly high overhead.
At least one byte means arbitrary bytes; that is, the caller does not need to read
or write the whole policy at one syscall.
Without understanding the requirement it would seem to me, that it would be
better to emulate that file buffer manipulation in userspace similar say C++
stringstreams, and then write the syscall when done.
The size of the whole policy in byte varies a lot.
syzbot will report user-controlled unbounded huge size memory
allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this
interface because
TOMOYO's policy is line-oriented ASCII text data where the
destination is switched via
pseudo‑filesystem's filename ...
While Tetsuo's comment is limited to TOMOYO, I believe the argument
applies to a number of other LSMs as well. The reality is that there
is no one policy ideal shared across LSMs and that complicates things
like the lsm_manage_policy() proposal. I'm intentionally saying
"complicates" and not "prevents" because I don't want to flat out
reject something like this, but I think there needs to be a larger
discussion among the different LSM groups about what such an API
should look like. We may not need to get every LSM to support this
new API, but we need to get something that would work for a
significant majority and would be general/extensible enough that we
would expect it to work with the majority of future LSMs (as much as
we can predict the future anyway).
yep, I look at this is just a starting point for discussion. There
isn't going to be any discussion without some code, so here is a v1
that supports a single LSM let the bike shedding begin.
Aside from the issues with allocating a buffer for a big policy
I don't see a problem with this proposal. The system call looks
a lot like the other LSM interfaces, so any developer who likes
those ought to like this one. The infrastructure can easily check
the lsm_id and only call the appropriate LSM hook, so no one
is going to be interfering with other modules.
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-09 03:25:09
On 5/8/25 08:07, Tetsuo Handa wrote:
On 2025/05/08 23:44, John Johansen wrote:
quoted
On 5/8/25 05:55, Tetsuo Handa wrote:
quoted
On 2025/05/08 17:25, John Johansen wrote:
quoted
That is fine. But curious I am curious what the interface would look like to fit TOMOYO's
needs.
Stream (like "FILE *") with restart from the beginning (like rewind(fp)) support.
That is, the caller can read/write at least one byte at a time, and written data
is processed upon encountering '\n'.
that can be emulated within the current sycall, where the lsm maintains a buffer.
That cannot be emulated, for there is no event that is automatically triggered when
the process terminates (i.e. implicit close() upon exit()) in order to release the
buffer the LSM maintains.
security_task_free()
quoted
Are you asking to also read data back out as well, that could be added, but doing
a syscall per byte here or through the fs is going to have fairly high overhead.
At least one byte means arbitrary bytes; that is, the caller does not need to read
or write the whole policy at one syscall.
got it
quoted
Without understanding the requirement it would seem to me, that it would be
better to emulate that file buffer manipulation in userspace similar say C++
stringstreams, and then write the syscall when done.
The size of the whole policy in byte varies a lot.
sure, buffers can be variable length. AppArmor policy also varies a lot in size.
More than anything I am trying to understand TOMOYO's requirements. They do
align better with using an fs interface. Can they be met sure, but it would
be more work for TOMOYO.
One of the big motivations for the syscall from the apparmor side is getting
away from the need to have the vfs present or having to pass an fd into the
environment.
On Thu, May 08, 2025 at 12:52:55AM -0700, John Johansen wrote:
On 5/7/25 15:04, Tetsuo Handa wrote:
quoted
On 2025/05/08 0:37, Maxime Bélair wrote:
quoted
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
If the interpretation of "flags" argument varies across LSMs, it sounds like ioctl()'s
yes that does feel like ioctls(), on the other hand defining them at the LSM level won't
offer LSMs flexibility making it so the syscall covers fewer use cases. I am not opposed
to either, it just hashing out what people want, and what is acceptable.
quoted
"cmd" argument. Also, there is prctl() which can already carry string-ish parameters
without involving open(). Why can't we use prctl() instead of lsm_manage_policy() ?
prctl() can be used, I used it for the unprivileged policy demo. It has its own set of
problems. While LSM policy could be associated with the process doing the load/replacement
or what ever operation, it isn't necessarily tied to it. A lot of LSM policy is not
process specific making prctl() a poor fit.
prctl() requires allocating a global prctl()
prctl() are already being filtered/controlled by LSMs making them a poort fit for
use by an LSM in a stacking situation as it requires updating the policy of other
LSMs on the system. Yes seccomp can filter the syscall but that still is an easier
barrier to overcome than having to have instruction for how to allow your LSMs
prctl() in multiple LSMs.
Mickaël already argued the need for landlock to have syscalls. See
Landlock indeed requires syscalls mainly because of its unprivileged
nature.
This link might be misleading though, it points to an initial version of
the syscall proposal (v17) and it was then decided to create one syscall
per operation (v34), which is why we ended with 3 syscalls. See the
changelog:
https://lore.kernel.org/r/20210422154123.13086-9-mic@digikod.net
Ideally those could have been LSM syscalls, with landlock leveraging them.
I don't agree. The Landlock syscalls have a well-defined semantic, with
documented security requirements, and they deal with specific kernel
objects identified with file descriptors, including a dedicated one:
[landlock-ruleset]. For the features provided by these Landlock
syscalls, it would not have been a good idea to reuse existing syscalls,
nor to rely on the syscall proposed in this series because the interface
is too specific to some of the current privileged LSMs (i.e. ingest a
policy blob). Making this interface more generic would lead to even
less defined semantic though.
AppArmor
is getting to where it has similar needs to landlock. Yes we can use ioctls, prctls,
netlink, the fs, etc. it doesn't mean that those are the best interfaces to do so,
I think it would make sense to propose AppArmor-specific syscalls.
and ideally any interface we use will be of benefit to some other LSMs in the future.
The LSM syscalls may make sense to deal with LSM blobs managed by the
LSM framework (e.g. get/set properties) when the operations are
common/generic.
Security policies are specific to each LSM and they should implement
their own well-defined interface (e.g. filesystem, netlink, syscall).
The LSM framework doesn't provide nor manage any security policy, it
mainly provides a set of consistent and well-defined kernel hooks with
security blobs to enforce a security policy. I don't think it makes
sense to add LSM syscalls to manage things not managed by the LSM
framework.
On Thu, May 08, 2025 at 01:18:20AM -0700, John Johansen wrote:
On 5/7/25 23:06, Song Liu wrote:
quoted
On Wed, May 7, 2025 at 8:37 AM Maxime Bélair
[off-list ref] wrote:
[...]
quoted
quoted
These two do not feel like real benefits:
- One syscall cannot fit all use cases well...
This syscall is not intended to cover every case, nor to replace existing kernel
interfaces.
Each LSM can decide which operations it wants to support (if any). For example, when
loading policies, an LSM may choose to allow only policies that further restrict
privileges.
quoted
- Not working in containers is often not an issue, but a feature.
Indeed, using this syscall requires appropriate capabilities and will not permit
unprivileged containers to manage policies arbitrarily.
With this syscall, capability checks remain the responsibility of each LSM.
For instance, in the AppArmor patch, a profile can be loaded only if
aa_policy_admin_capable() succeeds (which requires CAP_MAC_ADMIN). Moreover, by design,
policies can be loaded only in the current namespace.
I see this syscall as a middle point between exposing the entire sysfs, creating a large
attack surface, and blocking everything.
Landlock’s existing syscalls already improve security by allowing processes to further
restrict their ambient rights while adding only a modest attack surface.
This syscall is a further step in that direction: it lets LSMs add restrictive policies
without requiring exposing every other interface.
I don't think a syscall makes the API more secure. If necessary, we can add
It exposes a different attack surface. Requiring mounting of the fs to where it is visible
in the container, provides attack surface, and requires additional external configuration.
We should also keep in mind that syscalls could be accessible from
everywhere, by everyone, which may increase the attack surface compared
to a privileged filesystem interface. Adding a second interface may
also introduce issues. Anyway, I'm definitely not against syscalls, but
I don't see why the filesystem interface would be "less secure" in this
context.
Then there is the whole issue of getting the various LSMs to allow another LSM in the
stack to be able manage its own policy.
Right, and it's a similar issue with seccomp policies wrt syscalls.
quoted
permission check to each pseudo file. The downside of the syscall, however,
is that all the permission checks are hard-coded in the kernel (except for
The permission checks don't have to be hard coded. Each LSM can define how it handles
or manages the syscall. The default is that it isn't supported, but if an lsm decides
to support it, there is now reason that its policy can't determine the use of the
syscall.
From an interface design point of view, it would be better to clearly
specify the scope of a command (e.g. which components could be impacted
by a command), and make sure the documentation reflect that as well.
Even better, have a syscalls per required privileges and impact (e.g.
privileged or unprivileged). Going this road, I'm not sure if a
privileged syscall would make sense given the existing filesystem
interface.
quoted
BPF LSM); while the sys admin can configure permissions of the pseudo
files in user space.
Other LSMs also have policy that can control access to pseudo filesystems and
other resources. Again, the control doesn't have to be hard coded. And seccomp can
be used to block the syscall.
quoted
quoted
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
quoted
quoted
- Avoids overhead of other kernel interfaces for better efficiency
.. and it is is probably less efficient, because everything need to
fit in the same API.
As shown below, the syscall can significantly improve the performance of policy management.
A more detailed benchmark is available in [1].
The following table presents the time required to load an AppArmor profile.
For every cell, the first value is the total time taken by aa-load, and the value in
parentheses is the time spent to load the policy in the kernel only (total - dry‑run).
Results are in microseconds and are averaged over 10 000 runs to reduce variance.
| t (µs) | syscall | pseudofs | Speedup |
|-----------|-------------|-------------|---------------|
| 1password | 4257 (1127) | 3333 (192) | x1.28 (x5.86) |
| Xorg | 6099 (2961) | 5167 (2020) | x1.18 (x1.47) |
I am not sure the performance of loading security policies is on any
critical path.
generally speaking I agree, but I am also not going to turn down a
performance improvement either. Its a nice to have, but not a strong
argument for need.
quoted
The implementation calls the hook for each LSM, which is why I think the
syscall is not efficient.
it should only call the LSM identified by the lsmid in the call.
quoted
Overall, I am still not convinced a syscall for all LSMs is needed. To
justify such
its not needed by all LSMs, just a subset of them, and some nebulous
subset of potentially future LSMs that is entirely undefinable.
If we had had appropriate LSM syscalls landlock wouldn't have needed
to have landlock specific syscalls. Having another LSM go that route
feels wrong especially now that we have some LSM syscalls.
I don't agree. Dedicated syscalls are a good thing. See my other
reply.
If a
syscall is needed by an LSM its better to try hashing something out
that might have utility for multiple LSMs or at the very least,
potentially have utility in the future.
quoted
a syscall, I think we need to show that it is useful in multiple LSMs.
Also, if we
really want to have single set of APIs for all LSMs, we may also need
get_policy,
We are never going to get a single set of APIs for all LSMs. I will
settle for an api that has utility for a subset
quoted
remove_policy, etc. This set as-is appears to be an incomplete design. The
To have a complete design, there needs to be feedback and discussion
from multiple LSMs. This is a starting point.
quoted
implementation, with call_int_hook, is also problematic. It can easily
cause some> controversial behaviors.
agreed it shouldn't be doing a straight call_int_hook, it should only
call it against the lsm identified by the lsmid
Yes, but then, I don't see the point of a "generic" LSM syscall.
syzbot will report user-controlled unbounded huge size memory
allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this
interface because
TOMOYO's policy is line-oriented ASCII text data where the
destination is switched via
pseudo‑filesystem's filename ...
While Tetsuo's comment is limited to TOMOYO, I believe the argument
applies to a number of other LSMs as well. The reality is that there
is no one policy ideal shared across LSMs and that complicates things
like the lsm_manage_policy() proposal. I'm intentionally saying
"complicates" and not "prevents" because I don't want to flat out
reject something like this, but I think there needs to be a larger
discussion among the different LSM groups about what such an API
should look like. We may not need to get every LSM to support this
new API, but we need to get something that would work for a
significant majority and would be general/extensible enough that we
would expect it to work with the majority of future LSMs (as much as
we can predict the future anyway).
yep, I look at this is just a starting point for discussion. There
isn't going to be any discussion without some code, so here is a v1
that supports a single LSM let the bike shedding begin.
Aside from the issues with allocating a buffer for a big policy
I don't see a problem with this proposal. The system call looks
a lot like the other LSM interfaces, so any developer who likes
those ought to like this one. The infrastructure can easily check
the lsm_id and only call the appropriate LSM hook, so no one
is going to be interfering with other modules.
We may not want to only be able to load buffers containing policies, but
also to leverage file descriptors like Landlock does. Getting a
property from a kernel object or updating it is mainly about dealing
with a buffer. And the current LSM syscalls do just that. Other kind
of operations may require more than that though.
I don't like multiplexer syscalls because they don't expose a clear
semantic and can be complex to manage and filter. This new syscall is
kind of a multiplexer that redirect commands to an arbitrary set of
kernel parts, which can then define their own semantic. I'd like to see
a clear set of well-defined operations and their required permission.
Even better, one syscall per operation should simplify their interface.
syzbot will report user-controlled unbounded huge size memory
allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this
interface because
TOMOYO's policy is line-oriented ASCII text data where the
destination is switched via
pseudo‑filesystem's filename ...
While Tetsuo's comment is limited to TOMOYO, I believe the argument
applies to a number of other LSMs as well. The reality is that there
is no one policy ideal shared across LSMs and that complicates things
like the lsm_manage_policy() proposal. I'm intentionally saying
"complicates" and not "prevents" because I don't want to flat out
reject something like this, but I think there needs to be a larger
discussion among the different LSM groups about what such an API
should look like. We may not need to get every LSM to support this
new API, but we need to get something that would work for a
significant majority and would be general/extensible enough that we
would expect it to work with the majority of future LSMs (as much as
we can predict the future anyway).
yep, I look at this is just a starting point for discussion. There
isn't going to be any discussion without some code, so here is a v1
that supports a single LSM let the bike shedding begin.
Aside from the issues with allocating a buffer for a big policy
I don't see a problem with this proposal. The system call looks
a lot like the other LSM interfaces, so any developer who likes
those ought to like this one. The infrastructure can easily check
the lsm_id and only call the appropriate LSM hook, so no one
is going to be interfering with other modules.
We may not want to only be able to load buffers containing policies, but
also to leverage file descriptors like Landlock does. Getting a
property from a kernel object or updating it is mainly about dealing
with a buffer. And the current LSM syscalls do just that. Other kind
of operations may require more than that though.
I don't like multiplexer syscalls because they don't expose a clear
semantic and can be complex to manage and filter. This new syscall is
kind of a multiplexer that redirect commands to an arbitrary set of
kernel parts, which can then define their own semantic. I'd like to see
a clear set of well-defined operations and their required permission.
Even better, one syscall per operation should simplify their interface.
The development and maintenance of system calls is expensive in both
time and effort. LSM specific system calls frighten me. When I was
young adding system calls was just not done. A system call would
never be allowed for a specific sub-system or optional feature. True,
there are issues with the LSM specific filesystem approach. But I
like it, as it allows the LSM more freedom in its interfaces and
won't clutter the API if the LSM goes away or quits using it.
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-11 10:47:33
On 5/9/25 03:26, Mickaël Salaün wrote:
On Thu, May 08, 2025 at 01:18:20AM -0700, John Johansen wrote:
quoted
On 5/7/25 23:06, Song Liu wrote:
quoted
On Wed, May 7, 2025 at 8:37 AM Maxime Bélair
[off-list ref] wrote:
[...]
quoted
quoted
These two do not feel like real benefits:
- One syscall cannot fit all use cases well...
This syscall is not intended to cover every case, nor to replace existing kernel
interfaces.
Each LSM can decide which operations it wants to support (if any). For example, when
loading policies, an LSM may choose to allow only policies that further restrict
privileges.
quoted
- Not working in containers is often not an issue, but a feature.
Indeed, using this syscall requires appropriate capabilities and will not permit
unprivileged containers to manage policies arbitrarily.
With this syscall, capability checks remain the responsibility of each LSM.
For instance, in the AppArmor patch, a profile can be loaded only if
aa_policy_admin_capable() succeeds (which requires CAP_MAC_ADMIN). Moreover, by design,
policies can be loaded only in the current namespace.
I see this syscall as a middle point between exposing the entire sysfs, creating a large
attack surface, and blocking everything.
Landlock’s existing syscalls already improve security by allowing processes to further
restrict their ambient rights while adding only a modest attack surface.
This syscall is a further step in that direction: it lets LSMs add restrictive policies
without requiring exposing every other interface.
I don't think a syscall makes the API more secure. If necessary, we can add
It exposes a different attack surface. Requiring mounting of the fs to where it is visible
in the container, provides attack surface, and requires additional external configuration.
We should also keep in mind that syscalls could be accessible from
everywhere, by everyone, which may increase the attack surface compared
to a privileged filesystem interface. Adding a second interface may
also introduce issues. Anyway, I'm definitely not against syscalls, but
I don't see why the filesystem interface would be "less secure" in this
context.
yes syscalls being accessible from everywhere is another form of attack
surface, that needs to be mediated.
the fs can be mediated, its expose is a multiple lsms with multiple
different interfaces on the files within it. What really is more
problematic is makng the fs available in the container. Yes a
container manager can do it but then you are dependent on the
container manager making your interface available.
Other wise you are looking at making mount available to your app
within the container.
quoted
Then there is the whole issue of getting the various LSMs to allow another LSM in the
stack to be able manage its own policy.
Right, and it's a similar issue with seccomp policies wrt syscalls.
yes, though seccomp I have found to be the easier one to deal with
quoted
quoted
permission check to each pseudo file. The downside of the syscall, however,
is that all the permission checks are hard-coded in the kernel (except for
The permission checks don't have to be hard coded. Each LSM can define how it handles
or manages the syscall. The default is that it isn't supported, but if an lsm decides
to support it, there is now reason that its policy can't determine the use of the
syscall.
From an interface design point of view, it would be better to clearly
specify the scope of a command (e.g. which components could be impacted
by a command), and make sure the documentation reflect that as well.
Even better, have a syscalls per required privileges and impact (e.g.
privileged or unprivileged). Going this road, I'm not sure if a
privileged syscall would make sense given the existing filesystem
interface.
uhhhmmm, not just privileged. As you well know we are looking to use
this for unprivileged policy. The LSM can limit to privileged if it
wants but it doesn't have to limit it to privileged policy.
quoted
quoted
BPF LSM); while the sys admin can configure permissions of the pseudo
files in user space.
Other LSMs also have policy that can control access to pseudo filesystems and
other resources. Again, the control doesn't have to be hard coded. And seccomp can
be used to block the syscall.
quoted
quoted
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
quoted
quoted
- Avoids overhead of other kernel interfaces for better efficiency
.. and it is is probably less efficient, because everything need to
fit in the same API.
As shown below, the syscall can significantly improve the performance of policy management.
A more detailed benchmark is available in [1].
The following table presents the time required to load an AppArmor profile.
For every cell, the first value is the total time taken by aa-load, and the value in
parentheses is the time spent to load the policy in the kernel only (total - dry‑run).
Results are in microseconds and are averaged over 10 000 runs to reduce variance.
| t (µs) | syscall | pseudofs | Speedup |
|-----------|-------------|-------------|---------------|
| 1password | 4257 (1127) | 3333 (192) | x1.28 (x5.86) |
| Xorg | 6099 (2961) | 5167 (2020) | x1.18 (x1.47) |
I am not sure the performance of loading security policies is on any
critical path.
generally speaking I agree, but I am also not going to turn down a
performance improvement either. Its a nice to have, but not a strong
argument for need.
quoted
The implementation calls the hook for each LSM, which is why I think the
syscall is not efficient.
it should only call the LSM identified by the lsmid in the call.
quoted
Overall, I am still not convinced a syscall for all LSMs is needed. To
justify such
its not needed by all LSMs, just a subset of them, and some nebulous
subset of potentially future LSMs that is entirely undefinable.
If we had had appropriate LSM syscalls landlock wouldn't have needed
to have landlock specific syscalls. Having another LSM go that route
feels wrong especially now that we have some LSM syscalls.
I don't agree. Dedicated syscalls are a good thing. See my other
reply.
I think we can just disagree on this point.
quoted
If a
syscall is needed by an LSM its better to try hashing something out
that might have utility for multiple LSMs or at the very least,
potentially have utility in the future.
quoted
a syscall, I think we need to show that it is useful in multiple LSMs.
Also, if we
really want to have single set of APIs for all LSMs, we may also need
get_policy,
We are never going to get a single set of APIs for all LSMs. I will
settle for an api that has utility for a subset
quoted
remove_policy, etc. This set as-is appears to be an incomplete design. The
To have a complete design, there needs to be feedback and discussion
from multiple LSMs. This is a starting point.
quoted
implementation, with call_int_hook, is also problematic. It can easily
cause some> controversial behaviors.
agreed it shouldn't be doing a straight call_int_hook, it should only
call it against the lsm identified by the lsmid
Yes, but then, I don't see the point of a "generic" LSM syscall.
its not a generic LSM syscall. Its a syscall or maybe a set of syscalls
for a specific scoped problem of loading/managing policy.
Can we come to something acceptable? I don't know but we are going to
look at it before trying for an apparmor specific syscall.
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-11 11:10:03
On 5/9/25 03:25, Mickaël Salaün wrote:
On Thu, May 08, 2025 at 12:52:55AM -0700, John Johansen wrote:
quoted
On 5/7/25 15:04, Tetsuo Handa wrote:
quoted
On 2025/05/08 0:37, Maxime Bélair wrote:
quoted
Again, each module decides which operations to expose through this syscall. In many cases
the operation will still require CAP_SYS_ADMIN or a similar capability, so environments
that choose this interface remain secure while gaining its advantages.
If the interpretation of "flags" argument varies across LSMs, it sounds like ioctl()'s
yes that does feel like ioctls(), on the other hand defining them at the LSM level won't
offer LSMs flexibility making it so the syscall covers fewer use cases. I am not opposed
to either, it just hashing out what people want, and what is acceptable.
quoted
"cmd" argument. Also, there is prctl() which can already carry string-ish parameters
without involving open(). Why can't we use prctl() instead of lsm_manage_policy() ?
prctl() can be used, I used it for the unprivileged policy demo. It has its own set of
problems. While LSM policy could be associated with the process doing the load/replacement
or what ever operation, it isn't necessarily tied to it. A lot of LSM policy is not
process specific making prctl() a poor fit.
prctl() requires allocating a global prctl()
prctl() are already being filtered/controlled by LSMs making them a poort fit for
use by an LSM in a stacking situation as it requires updating the policy of other
LSMs on the system. Yes seccomp can filter the syscall but that still is an easier
barrier to overcome than having to have instruction for how to allow your LSMs
prctl() in multiple LSMs.
Mickaël already argued the need for landlock to have syscalls. See
Landlock indeed requires syscalls mainly because of its unprivileged
nature.
This link might be misleading though, it points to an initial version of
the syscall proposal (v17) and it was then decided to create one syscall
per operation (v34), which is why we ended with 3 syscalls. See the
changelog:
https://lore.kernel.org/r/20210422154123.13086-9-mic@digikod.net
yes and no. I am well aware landlock's syscall got split into three syscalls.
All I was trying to do is reference to the start of the discussion on why
landlock needed a syscall(s). I thought the details of why you have three
etc, really didn't add to the discussion. But yeah not also pointing to
v34 could be considered misleading.
quoted
Ideally those could have been LSM syscalls, with landlock leveraging them.
I don't agree. The Landlock syscalls have a well-defined semantic, with
First I don't begrudge Landlock its syscalls, I think at the time it was
the only way forward.
documented security requirements, and they deal with specific kernel
objects identified with file descriptors, including a dedicated one:
[landlock-ruleset].
I am aware. Those semantics could have been kept and documented, within
a set of LSM syscalls. Yes landlock's syscalls shouldn't have been done
behind a single LSM syscall, I am not advocating for that but maybe
behind several LSM syscalls.
For the features provided by these Landlock
syscalls, it would not have been a good idea to reuse existing syscalls,
nor to rely on the syscall proposed in this series because the interface
is too specific to some of the current privileged LSMs (i.e. ingest a
policy blob). Making this interface more generic would lead to even
less defined semantic though.
Right, so again not a generic LSM syscall. But "generic" LSM syscalls
for certain purposes. Let me walk my statement back a little, what I
find unfortunate was that the landlock LSM syscalls didn't get discussed
as a set of generic LSM syscall's with landlock being the first to
implement them.
The question is hashing out where the generic semantics are vs. the
individual LSMs. Having an LSM syscall to deal with specific kernel
objects idenetified with file descriptors, and allowing each LSMs
to deal with that if it needs is possible.
Its a matter of figuring something out. It could be it turns out it is
not worth it. And some individual LSM syscalls like landlocks are the
way to go, its that it wasn't explored. I don't fault you, and think
it really wasn't even an option at the time.
quoted
AppArmor
is getting to where it has similar needs to landlock. Yes we can use ioctls, prctls,
netlink, the fs, etc. it doesn't mean that those are the best interfaces to do so,
I think it would make sense to propose AppArmor-specific syscalls.
that may be the case, but I think we should explore providing a more
LSM generic interface first.
quoted
and ideally any interface we use will be of benefit to some other LSMs in the future.
The LSM syscalls may make sense to deal with LSM blobs managed by the
LSM framework (e.g. get/set properties) when the operations are
common/generic.
Security policies are specific to each LSM and they should implement
their own well-defined interface (e.g. filesystem, netlink, syscall).
policies at some level are just blobs too. It is worth at least
exploring whether there can be a common interface.
The LSM framework doesn't provide nor manage any security policy, it
mainly provides a set of consistent and well-defined kernel hooks with
security blobs to enforce a security policy. I don't think it makes
sense to add LSM syscalls to manage things not managed by the LSM
framework.
we aren't talking about the LSM framework managing security policy,
just whether it makes sense for it to provide a common interface that
an LSM can choose to use to provide it a blob of policy that it
can then manage.
Its just a mechanism. This isn't all that different than using the
filesystem, netlink, or other mechanisms to shuttle the blob
between userspace to the kernel, and then the LSM manages its
policy and data.
The big difference is that using the syscall opens unprivileged
policy up to the LSM more broadly. If we are going to go the syscall
route for apparmor, we might as well see if we can't make that
mechanism more broadly available, and make it easier for other
LSMs in the future.
Again, it might turn out its a fools errand, and we have to do
an apparmor specific syscall, but it is worth exploring first.
syzbot will report user-controlled unbounded huge size memory
allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this
interface because
TOMOYO's policy is line-oriented ASCII text data where the
destination is switched via
pseudo‑filesystem's filename ...
While Tetsuo's comment is limited to TOMOYO, I believe the argument
applies to a number of other LSMs as well. The reality is that there
is no one policy ideal shared across LSMs and that complicates things
like the lsm_manage_policy() proposal. I'm intentionally saying
"complicates" and not "prevents" because I don't want to flat out
reject something like this, but I think there needs to be a larger
discussion among the different LSM groups about what such an API
should look like. We may not need to get every LSM to support this
new API, but we need to get something that would work for a
significant majority and would be general/extensible enough that we
would expect it to work with the majority of future LSMs (as much as
we can predict the future anyway).
yep, I look at this is just a starting point for discussion. There
isn't going to be any discussion without some code, so here is a v1
that supports a single LSM let the bike shedding begin.
Aside from the issues with allocating a buffer for a big policy
I don't see a problem with this proposal. The system call looks
a lot like the other LSM interfaces, so any developer who likes
those ought to like this one. The infrastructure can easily check
the lsm_id and only call the appropriate LSM hook, so no one
is going to be interfering with other modules.
We may not want to only be able to load buffers containing policies, but
also to leverage file descriptors like Landlock does. Getting a
I am not opposed to a syscall that leverages file desriptors like landlock
but that would be a different syscall with different semantics, and
something for an lsm that wants that semantic to introduce.
property from a kernel object or updating it is mainly about dealing
with a buffer. And the current LSM syscalls do just that. Other kind
of operations may require more than that though.
sure but they don't do it for the semantic of loading/managing policy.
I don't like multiplexer syscalls because they don't expose a clear
semantic and can be complex to manage and filter. This new syscall is
kind of a multiplexer that redirect commands to an arbitrary set of
kernel parts, which can then define their own semantic. I'd like to see
a clear set of well-defined operations and their required permission.
Even better, one syscall per operation should simplify their interface.
I am not opposed to that approach. This can be multiple syscalls. Its
a v1 to try and see if we can come to any agreement on a set of semantics
syzbot will report user-controlled unbounded huge size memory
allocation attempt. ;-)
This interface might be fine for AppArmor, but TOMOYO won't use this
interface because
TOMOYO's policy is line-oriented ASCII text data where the
destination is switched via
pseudo‑filesystem's filename ...
While Tetsuo's comment is limited to TOMOYO, I believe the argument
applies to a number of other LSMs as well. The reality is that there
is no one policy ideal shared across LSMs and that complicates things
like the lsm_manage_policy() proposal. I'm intentionally saying
"complicates" and not "prevents" because I don't want to flat out
reject something like this, but I think there needs to be a larger
discussion among the different LSM groups about what such an API
should look like. We may not need to get every LSM to support this
new API, but we need to get something that would work for a
significant majority and would be general/extensible enough that we
would expect it to work with the majority of future LSMs (as much as
we can predict the future anyway).
yep, I look at this is just a starting point for discussion. There
isn't going to be any discussion without some code, so here is a v1
that supports a single LSM let the bike shedding begin.
Aside from the issues with allocating a buffer for a big policy
I don't see a problem with this proposal. The system call looks
a lot like the other LSM interfaces, so any developer who likes
those ought to like this one. The infrastructure can easily check
the lsm_id and only call the appropriate LSM hook, so no one
is going to be interfering with other modules.
We may not want to only be able to load buffers containing policies, but
also to leverage file descriptors like Landlock does. Getting a
property from a kernel object or updating it is mainly about dealing
with a buffer. And the current LSM syscalls do just that. Other kind
of operations may require more than that though.
I don't like multiplexer syscalls because they don't expose a clear
semantic and can be complex to manage and filter. This new syscall is
kind of a multiplexer that redirect commands to an arbitrary set of
kernel parts, which can then define their own semantic. I'd like to see
a clear set of well-defined operations and their required permission.
Even better, one syscall per operation should simplify their interface.
The development and maintenance of system calls is expensive in both
time and effort. LSM specific system calls frighten me. When I was
young adding system calls was just not done. A system call would
never be allowed for a specific sub-system or optional feature. True,
there are issues with the LSM specific filesystem approach. But I
like it, as it allows the LSM more freedom in its interfaces and
won't clutter the API if the LSM goes away or quits using it.
I get the reticence on adding syscalls. Indeed its part of why I
want to explore LSM syscalls before going with an apparmor specific
syscall.
The current LSM specific fs approach has limitations that just can't
be reasonably worked around for some use cases, so that leaves going
with an alternate mechanism. For this use case, ioctls are problematic
like the fs. prctl could work for a subset and abused for the whole,
but a syscall feels cleaner.
I am open to other options.
On Sun, May 11, 2025 at 03:47:21AM -0700, John Johansen wrote:
On 5/9/25 03:26, Mickaël Salaün wrote:
quoted
On Thu, May 08, 2025 at 01:18:20AM -0700, John Johansen wrote:
quoted
On 5/7/25 23:06, Song Liu wrote:
quoted
On Wed, May 7, 2025 at 8:37 AM Maxime Bélair
[off-list ref] wrote:
[...]
quoted
quoted
quoted
permission check to each pseudo file. The downside of the syscall, however,
is that all the permission checks are hard-coded in the kernel (except for
The permission checks don't have to be hard coded. Each LSM can define how it handles
or manages the syscall. The default is that it isn't supported, but if an lsm decides
to support it, there is now reason that its policy can't determine the use of the
syscall.
From an interface design point of view, it would be better to clearly
specify the scope of a command (e.g. which components could be impacted
by a command), and make sure the documentation reflect that as well.
Even better, have a syscalls per required privileges and impact (e.g.
privileged or unprivileged). Going this road, I'm not sure if a
privileged syscall would make sense given the existing filesystem
interface.
uhhhmmm, not just privileged. As you well know we are looking to use
this for unprivileged policy. The LSM can limit to privileged if it
wants but it doesn't have to limit it to privileged policy.
Yes, I meant to say having a syscall for unprivileged actions, and maybe
another one for privileged ones, but this might be a hard sell. :)
To say it another way, for your use case, do you need this syscall(s)
for privileged operations? Do you plan to drop (or stop extending) the
filesystem interface or do you think it would be good for (AppArmor)
privileged operations too? I know syscalls might be attractive and
could be used for everything, but it's good to have a well-defined plan
and semantic to avoid using such syscall as another multiplexer with
unrelated operations and required privileges.
If this syscall should also be a way to do privileged operations, should
we also agree on a common set of permissions (e.g. global CAP_MAC_ADMIN
or user namespace one)?
[...]
quoted
quoted
quoted
Overall, I am still not convinced a syscall for all LSMs is needed. To
justify such
its not needed by all LSMs, just a subset of them, and some nebulous
subset of potentially future LSMs that is entirely undefinable.
If we had had appropriate LSM syscalls landlock wouldn't have needed
to have landlock specific syscalls. Having another LSM go that route
feels wrong especially now that we have some LSM syscalls.
I don't agree. Dedicated syscalls are a good thing. See my other
reply.
I think we can just disagree on this point.
quoted
quoted
If a
syscall is needed by an LSM its better to try hashing something out
that might have utility for multiple LSMs or at the very least,
potentially have utility in the future.
quoted
a syscall, I think we need to show that it is useful in multiple LSMs.
Also, if we
really want to have single set of APIs for all LSMs, we may also need
get_policy,
We are never going to get a single set of APIs for all LSMs. I will
settle for an api that has utility for a subset
quoted
remove_policy, etc. This set as-is appears to be an incomplete design. The
To have a complete design, there needs to be feedback and discussion
from multiple LSMs. This is a starting point.
quoted
implementation, with call_int_hook, is also problematic. It can easily
cause some> controversial behaviors.
agreed it shouldn't be doing a straight call_int_hook, it should only
call it against the lsm identified by the lsmid
Yes, but then, I don't see the point of a "generic" LSM syscall.
its not a generic LSM syscall. Its a syscall or maybe a set of syscalls
for a specific scoped problem of loading/managing policy.
Can we come to something acceptable? I don't know but we are going to
look at it before trying for an apparmor specific syscall.
I understand and it's good to have this discussion.
From: John Johansen <john.johansen@canonical.com> Date: 2025-05-17 07:59:39
On 5/12/25 03:20, Mickaël Salaün wrote:
On Sun, May 11, 2025 at 03:47:21AM -0700, John Johansen wrote:
quoted
On 5/9/25 03:26, Mickaël Salaün wrote:
quoted
On Thu, May 08, 2025 at 01:18:20AM -0700, John Johansen wrote:
quoted
On 5/7/25 23:06, Song Liu wrote:
quoted
On Wed, May 7, 2025 at 8:37 AM Maxime Bélair
[off-list ref] wrote:
[...]
quoted
quoted
quoted
quoted
permission check to each pseudo file. The downside of the syscall, however,
is that all the permission checks are hard-coded in the kernel (except for
The permission checks don't have to be hard coded. Each LSM can define how it handles
or manages the syscall. The default is that it isn't supported, but if an lsm decides
to support it, there is now reason that its policy can't determine the use of the
syscall.
From an interface design point of view, it would be better to clearly
specify the scope of a command (e.g. which components could be impacted
by a command), and make sure the documentation reflect that as well.
Even better, have a syscalls per required privileges and impact (e.g.
privileged or unprivileged). Going this road, I'm not sure if a
privileged syscall would make sense given the existing filesystem
interface.
uhhhmmm, not just privileged. As you well know we are looking to use
this for unprivileged policy. The LSM can limit to privileged if it
wants but it doesn't have to limit it to privileged policy.
Yes, I meant to say having a syscall for unprivileged actions, and maybe
another one for privileged ones, but this might be a hard sell. :)
indeed, in the apparmor case context would be important. Just exactly
what is privileged. It may be a privileged operation to load policy to one
namespace, but not to another that you are setting up for a child.
To say it another way, for your use case, do you need this syscall(s)
for privileged operations? Do you plan to drop (or stop extending) the
need, probably. That is to say, loading of policy have varying levels
of privilege. root within the container has privilege to load policy
to its namespace, but it might have authority to setup a child namespace
that does not require privilege for it to load policy into, and it
will determine if the child has privilege or unprivleged policy within
it.
Ideally we won't have to use the fs interface within the "privileged"
container, as there are cases where this is currently not done or
undesirable.
filesystem interface or do you think it would be good for (AppArmor)
privileged operations too? I know syscalls might be attractive and
could be used for everything, but it's good to have a well-defined plan
and semantic to avoid using such syscall as another multiplexer with
unrelated operations and required privileges.
sure. But the privilege level is use case dependent, to which policy
namespace is policy being loaded, replaced, ... The privilege level
very much will depend on what is in the stack/bounding of policy.
If this syscall should also be a way to do privileged operations, should
we also agree on a common set of permissions (e.g. global CAP_MAC_ADMIN
or user namespace one)?
I think requiring something like CAP_MAC_ADMIN would be a per LSM
decision.
[...]
quoted
quoted
quoted
quoted
Overall, I am still not convinced a syscall for all LSMs is needed. To
justify such
its not needed by all LSMs, just a subset of them, and some nebulous
subset of potentially future LSMs that is entirely undefinable.
If we had had appropriate LSM syscalls landlock wouldn't have needed
to have landlock specific syscalls. Having another LSM go that route
feels wrong especially now that we have some LSM syscalls.
I don't agree. Dedicated syscalls are a good thing. See my other
reply.
I think we can just disagree on this point.
quoted
quoted
If a
syscall is needed by an LSM its better to try hashing something out
that might have utility for multiple LSMs or at the very least,
potentially have utility in the future.
quoted
a syscall, I think we need to show that it is useful in multiple LSMs.
Also, if we
really want to have single set of APIs for all LSMs, we may also need
get_policy,
We are never going to get a single set of APIs for all LSMs. I will
settle for an api that has utility for a subset
quoted
remove_policy, etc. This set as-is appears to be an incomplete design. The
To have a complete design, there needs to be feedback and discussion
from multiple LSMs. This is a starting point.
quoted
implementation, with call_int_hook, is also problematic. It can easily
cause some> controversial behaviors.
agreed it shouldn't be doing a straight call_int_hook, it should only
call it against the lsm identified by the lsmid
Yes, but then, I don't see the point of a "generic" LSM syscall.
its not a generic LSM syscall. Its a syscall or maybe a set of syscalls
for a specific scoped problem of loading/managing policy.
Can we come to something acceptable? I don't know but we are going to
look at it before trying for an apparmor specific syscall.
I understand and it's good to have this discussion.