[PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
flat view
HOTtoday
From: Josef Bacik <josef@toxicpanda.com>
Date: 2026-10-06 15:44:42
Also in:
keyrings, linux-fsdevel, lkml
Hello,
Commit db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
stops a root task that has given up CAP_SETFCAP from creating a user
namespace that maps uid 0 and then writing file capabilities in it that
the initial namespace honours. The check only looks at the task that
creates the namespace, so if somebody who did have CAP_SETFCAP created
one, there are still several ways around it:
- setns() into their namespace and write uid_map from inside (patch 2)
- write "0 0 1" through a uid_map fd that they opened (patch 3)
- setns() into a namespace of theirs that already maps uid 0 and set
security.capability there, no map write needed (patch 4)
- ptrace one of their tasks and have it do any of the above (patch 5)
On an unmodified kernel we took a uid 0 task with CAP_SETFCAP dropped
from its permitted, effective and bounding sets and, through each of
these, ended up with a file that a uid 1000 user execs with
CAP_SYS_ADMIN in its effective set.
Patch 1 adds cred->setfcap_level, which records how far up the
namespace tree a task's CAP_SETFCAP reached when it entered its
namespace, and patches 2-5 check it. Patch 4 is the check that closes
the class, the map patches make the uid 0 map rule mean what
db2e718a4798 meant it to, and patch 5 keeps ptrace from borrowing what
the target is entitled to.
This does change behaviour. Everything new is -EPERM:
- a task that entered a namespace without CAP_SETFCAP outside can't map
uid 0 of the outside or write fscaps for that root user anymore
- a uid_map fd opened by a task with CAP_SETFCAP can't be used by a
task without it to map uid 0
- if a privileged task maps "0 0 1" from the parent for a namespace
created by a task without CAP_SETFCAP, that namespace can no longer
write fscaps honoured outside
- PTRACE_ATTACH and PTRACE_TRACEME fail when the tracer gave up
CAP_SETFCAP and CAP_SYS_PTRACE, the target didn't, and they share a
root user
Rootless containers, privileged runtimes writing the map from the
parent, nested unprivileged namespaces and containers that don't map
host uid 0 aren't affected. The ptrace check is one compare for
targets in the initial namespace and in namespaces entered without
the capability.
Testing: a set of flows run on the base and patched kernels, every
bypass route above gets -EPERM with the series and the 16 legitimate
flows behave the same. The capabilities, namespaces, ptrace, pidfd and
proc selftests give the same results before and after. Thanks,
Josef
---
Josef Bacik (5):
cred: record how far up CAP_SETFCAP reaches
userns: don't let setns() lend the right to map uid 0
userns: check the writer too before mapping uid 0
capabilities: limit fscaps to where CAP_SETFCAP reaches
capabilities: don't let ptrace borrow CAP_SETFCAP
include/linux/capability.h | 4 ++
include/linux/cred.h | 1 +
kernel/user_namespace.c | 40 +++++++++++----
security/commoncap.c | 119 +++++++++++++++++++++++++++++++++++++++++--
security/keys/process_keys.c | 1 +
5 files changed, 150 insertions(+), 15 deletions(-)
---
base-commit: 7909a3e30a05e40bbc8bfb7f5629ed642abeaab8
change-id: 20261006-b4-setfcap-userns-d63185a31ce3