Thread (8 messages) 8 messages, 1 author, 3d ago
DORMANTno replies
Revisions (2)
  1. rfc [diff vs current]
  2. v2 current

[PATCH v2 7/7] Documentation: add failfs documentation

From: Christian Brauner <brauner@kernel.org>
Date: 2026-07-24 13:41:52
Also in: linux-doc, linux-fsdevel, lkml
Subsystem: documentation, the rest · Maintainers: Jonathan Corbet, Linus Torvalds

Document the failfs semantics, the FD_FAILFS_ROOT sentinel, the
fchroot() entry requirements, and the ways back out.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 Documentation/filesystems/failfs.rst | 73 ++++++++++++++++++++++++++++++++++++
 Documentation/filesystems/index.rst  |  1 +
 2 files changed, 74 insertions(+)
diff --git a/Documentation/filesystems/failfs.rst b/Documentation/filesystems/failfs.rst
new file mode 100644
index 000000000000..21ff2db7941d
--- /dev/null
+++ b/Documentation/filesystems/failfs.rst
@@ -0,0 +1,73 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+======
+failfs
+======
+
+failfs is a kernel-internal filesystem that fails every operation
+reaching it with ``EOPNOTSUPP``. It is the counterpart to nullfs. Where
+nullfs is permanently empty, failfs means "nothing is supported here".
+It cannot be mounted from userspace, nothing can be mounted on top of
+it. It cannot be cloned.
+
+The only way into it is the ``FD_FAILFS_ROOT`` file descriptor sentinel which
+is understood by ``fchdir(2)`` and ``fchroot(2)``.
+
+Semantics
+=========
+
+Every path walk of a component through failfs fails with
+``EOPNOTSUPP`` before that component is parsed, including ``.``.
+
+No path lookup can open the root, not even with ``O_PATH``.
+
+A process with its working directory in failfs fails every
+``AT_FDCWD``-relative lookup. As with any working directory that is
+unreachable from the process root, the ``getcwd(2)`` system call returns
+a path prefixed with ``(unreachable)``.
+
+A process with its root directory in failfs fails every absolute path
+lookup including absolute symlinks and the interpreter of dynamically
+linked binaries. In other words, this fails exec.
+
+Lookups anchored at explicit directory file descriptors keep working. It
+is the ``fs_struct`` equivalent of ``RESOLVE_BENEATH``. The process must
+anchor every lookup at a file descriptor it explicitly holds.
+
+Entering
+========
+
+``fchroot(FD_FAILFS_ROOT, 0)`` requires ``CAP_SYS_CHROOT`` in the
+caller's user namespace, mirroring ``chroot(2)``. Unprivileged callers
+may enter if all of the following hold:
+
+* ``no_new_privs`` is set: setuid binaries on regular mounts remain
+  reachable via inherited directory file descriptors and executing them
+  with an unusable root directory is the classic confused deputy.
+
+* The caller is not already chrooted: the root directory is what
+  confines ``..`` resolution and the failfs root can never be reached by
+  walking up a real mount tree, so moving the root of a chrooted task to
+  failfs would allow it to escape its chroot via ``openat(fd, "..")``.
+
+* The caller does not share its ``fs_struct``: ``no_new_privs`` is
+  checked on the calling thread, but the root lives in the ``fs_struct``.
+  A ``CLONE_FS`` sibling without ``no_new_privs`` could otherwise execute
+  a setuid binary with the failfs root, so entry requires ``fs->users ==
+  1``, the same restriction ``setns(2)`` applies for the mount and user
+  namespaces.
+
+Leaving
+=======
+
+Backing out is currently hard, but this is a property of the current
+implementation, not a guaranteed interface, and may be loosened later.
+For now a process that entered failfs counts as chrooted, so it cannot
+create user namespaces to regain ``CAP_SYS_CHROOT``, and ``chroot(2)``
+or ``fchroot(2)`` back out require ``CAP_SYS_CHROOT``. The remaining way
+out today is ``setns(2)`` with a mount namespace file descriptor, which
+requires ``CAP_SYS_ADMIN`` over the target mount namespace as well as
+``CAP_SYS_CHROOT`` and ``CAP_SYS_ADMIN`` in the caller's user namespace
+and resets both root and working directory. A process that holds no such
+file descriptor and restricts ``*chdir()``/``*chroot()``/``setns()`` via
+seccomp cannot currently get back out.
diff --git a/Documentation/filesystems/index.rst b/Documentation/filesystems/index.rst
index 1f71cf159547..734a45e51667 100644
--- a/Documentation/filesystems/index.rst
+++ b/Documentation/filesystems/index.rst
@@ -91,6 +91,7 @@ Documentation for filesystem implementations.
    ext3
    ext4/index
    f2fs
+   failfs
    gfs2/index
    hfs
    hfsplus
-- 
2.53.0
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help