Re: [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
flat view
From: Josef Bacik <josef@toxicpanda.com>
Date: 2026-10-08 15:40:56
Also in:
keyrings, linux-fsdevel, lkml
On Thu, Oct 08, 2026 at 08:57:56AM -0500, Serge E. Hallyn wrote:
Would you mind describing what other solutions you considered? I've been looking over this set since Tuesday, and finding it hard to reason about. (Part of that is certainly the nature of the problem, and it's possible that this is the best/simplest solution.)
Everything we looked at kept the state in the cred. We model checked the variants before writing the code, and these fell over: - a bool per cred for "had CAP_SETFCAP over the parent when it entered". It breaks on two hops: setns() into a namespace that maps 0, unshare again, and the bool says yes for the second namespace. Hence the level. - checking only at uid_map write time. That misses setxattr of security.capability in a namespace that already maps 0, hence patch 4. We didn't look at keeping the state on the namespace.
If we replaced the userns->parent_could_setfcap bool with a ref to the creator's cred, then at both setns and write we could check the actor's credentials, right? There are probably issues with that specific idea, but that's why it would be good to see what else you've considered.
Checking at write time alone doesn't work: once a task is inside the namespace its cred says nothing about what it could do outside, so a joiner and the creator look the same. It does work if setns() refuses to join a namespace that maps, or can still map, the parent's uid 0 unless the joiner has CAP_SETFCAP over the parent. Then everybody in a namespace has the same reach, and it can be a level stored on the namespace at create time instead of a cred ref, which would pin keyrings and the rest for the life of the namespace. The checks would be setns(), the map write (opener and writer are in the parent, so a plain capable check), setxattr of security.capability, and ptrace. The difference in behaviour is that the -EPERM moves to setns(): a root task without CAP_SETFCAP couldn't enter a root-owned container that maps host uid 0 at all, where with this series it can enter and is refused only for the map, fscaps and ptrace. I can prototype it if you prefer that.
Of course UID 0 will always continue to carry privileges even with an empty cap_eff. Here we're stopping it from writing filecaps to uid 0 owned files, but if it can open a 0 owned file on the host, like /bin/sh or a systemd init file, or ptrace a process (in a child ns that maps parent uid 0) doing so, it can still cause damage. My point being, we do need to keep in mind the tradeoff of keeping the code simple versus the realistic threat of the problem being addressed.
On the same kernels, the restricted root task can copy a binary and chmod 4755 it (it owns it, no capability needed), and a uid 1000 user runs it with a full CapEff. With SECBIT_NOROOT that setuid copy gives uid 1000 nothing, while the fscap file still gives it what's in the xattr on an unpatched kernel. So SECBIT_NOROOT is the case the series adds anything for, the same case db2e718a4798 covers. A smaller version is patches 1, 4 and 5, with cap_root_level() moved from 2 into 4. I built that and ran the same flows: every route that ends in a file capability still gets -EPERM and uid 1000 gets nothing. The uid 0 map writes refused by 2 and 3 go through again, but the fscap write after them is refused. Let me know which way you'd like to go and I'll rework it. Thanks, Josef