From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:01:59
Here are a set of patches that adds system calls, that (a) allow
information about the VFS, mount topology, superblock and files to be
retrieved and (b) allow for notifications of mount topology rearrangement
events, mount and superblock attribute changes and other superblock events,
such as errors.
========================
FILESYSTEM NOTIFICATIONS
========================
The watch_mount() system call places a watch on a point in the mount
topology specified by the dirfd, path and at_flags parameters. All mount
topology change and mount attribute change notifications in the subtree
rooted at that point can be intercepted by the watch. Watches are ducted
through pipes:
int fd[2];
pipe2(fd, O_NOTIFICATION_PIPE);
ioctl(fd[0], IOC_WATCH_QUEUE_SET_SIZE, BUF_SIZE);
watch_mount(AT_FDCWD, "/", 0, fd[0], 0x02);
Events include:
- New mount made
- Mount unmounted
- Mount expired
- R/O state changed
- Other attribute changed
- Mount moved from
- Mount moved to
Using filtering, this may be limited in various ways (single mount watch vs
subtree watch, recursive vs non-recursive changes, to-R/O vs to-R/W, mount
vs submount).
Each mount now has a change counter. Whenever a mount is changed, this
gets incremented. It can be queried by fsinfo() using either
FSINFO_ATTR_MOUNT_INFO or FSINFO_ATTR_MOUNT_CHILDREN. The ID of the mount
on which the notification is generated is placed into the notification
message (triggered_on). If the event involves a second mount as well, such
as creation of a new mount, that gets returned too (changed_mount).
The watch_sb() system call places a watch on the superblock specified by
the dirfd, path and at_flags parameters. This allows various superblock
events to be monitored for, such as:
- Transition between R/W and R/O
- Filesystem errors
- Quota overrun
- Network status changes
Each superblock now gets a 64-bit unique superblock identifier and a
notification counter. The counter is incremented each time one of these
notifications would be generated. This attributes can be queried using
fsinfo() with FSINFO_ATTR_SB_NOTIFICATIONS. The identifier is placed into
notification messages.
============================
FILESYSTEM INFORMATION QUERY
============================
The fsinfo() system call allows information about the filesystem at a
particular path point to be queried as a set of attributes, some of which
may have more than one value.
Attribute values are of four basic types:
(1) Version dependent-length structure (size defined by type).
(2) Variable-length string (up to 4096, including NUL).
(3) List of structures (up to INT_MAX size).
(4) Opaque blob (up to INT_MAX size).
Attributes can have multiple values either as a sequence of values or a
sequence-of-sequences of values and all the values of a particular
attribute must be of the same type.
Note that the values of an attribute *are* allowed to vary between dentries
within a single superblock, depending on the specific dentry that you're
looking at, but all the values of an attribute have to be of the same type.
I've tried to make the interface as light as possible, so integer/enum
attribute selector rather than string and the core does all the allocation
and extensibility support work rather than leaving that to the filesystems.
That means that for the first two attribute types, the filesystem will
always see a sufficiently-sized buffer allocated. Further, this removes
the possibility of the filesystem gaining access to the userspace buffer.
fsinfo() allows a variety of information to be retrieved about a filesystem
and the mount topology:
(1) General superblock attributes:
- Filesystem identifiers (UUID, volume label, device numbers, ...)
- The limits on a filesystem's capabilities
- Information on supported statx fields and attributes and IOC flags.
- A variety single-bit flags indicating supported capabilities.
- Timestamp resolution and range.
- The amount of space/free space in a filesystem (as statfs()).
- Superblock notification counter.
(2) Filesystem-specific superblock attributes:
- Superblock-level timestamps.
- Cell name.
- Server names and addresses.
- Filesystem-specific information.
(3) VFS information:
- Mount topology information.
- Mount attributes.
- Mount notification counter.
(4) Information about what the fsinfo() syscall itself supports, including
the type and struct/element size of attributes.
The system is extensible:
(1) New attributes can be added. There is no requirement that a
filesystem implement every attribute. Note that the core VFS keeps a
table of types and sizes so it can handle future extensibility rather
than delegating this to the filesystems.
(2) Version length-dependent structure attributes can be made larger and
have additional information tacked on the end, provided it keeps the
layout of the existing fields. If an older process asks for a shorter
structure, it will only be given the bits it asks for. If a newer
process asks for a longer structure on an older kernel, the extra
space will be set to 0. In all cases, the size of the data actually
available is returned.
In essence, the size of a structure is that structure's version: a
smaller size is an earlier version and a later version includes
everything that the earlier version did.
(3) New single-bit capability flags can be added. This is a structure-typed
attribute and, as such, (2) applies. Any bits you wanted but the kernel
doesn't support are automatically set to 0.
fsinfo() may be called like the following, for example:
struct fsinfo_params params = {
.at_flags = AT_SYMLINK_NOFOLLOW,
.flags = FSINFO_FLAGS_QUERY_PATH,
.request = FSINFO_ATTR_AFS_SERVER_ADDRESSES,
.Nth = 2,
};
struct fsinfo_server_address address;
len = fsinfo(AT_FDCWD, "/afs/grand.central.org/doc", ¶ms,
&address, sizeof(address));
The above example would query an AFS filesystem to retrieve the address
list for the 3rd server, and:
struct fsinfo_params params = {
.at_flags = AT_SYMLINK_NOFOLLOW,
.flags = FSINFO_FLAGS_QUERY_PATH,
.request = FSINFO_ATTR_AFS_CELL_NAME;
};
char cell_name[256];
len = fsinfo(AT_FDCWD, "/afs/grand.central.org/doc", ¶ms,
&cell_name, sizeof(cell_name));
would retrieve the name of an AFS cell as a string.
In future, I want to make fsinfo() capable of querying a context created by
fsopen() or fspick(), e.g.:
fd = fsopen("ext4", 0);
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_PARAMETERS;
};
char buffer[65536];
fsinfo(fd, NULL, ¶ms, &buffer, sizeof(buffer));
even if that context doesn't currently have a superblock attached. I would
prefer this to contain length-prefixed strings so that there's no need to
insert escaping, especially as any character, including '\', can be used as
the separator in cifs and so that binary parameters can be returned (though
that is a lesser issue).
Two sample programs are provided, one to query filesystem attributes and
the other to display a mount subtree. Both of them can be given a path or
a mount ID to start at. Further, the watch_test sample program now watches
for mount events under "/" and for superblock events on whatever superblock
is backing "/mnt" when it the program is started.
The patches can be found here also:
https://git.kernel.org/pub/scm/linux/kernel/git/dhowells/linux-fs.git
on branch:
fsinfo-core
===================
SIGNIFICANT CHANGES
===================
ver #17:
(*) Applied comments from Jann Horn, Darrick Wong and Christian Brauner.
(*) Rearranged the order in which fsinfo() does things so that the
superblock operations table can have a function pointer rather than a
table pointer. The ->fsinfo() op is now called at least twice, once
to determine the size of buffer needed and then to retrieve the data.
If the retrieval step indicates yet more space is needed, the buffer
will be expanded and that step repeated.
(*) Merge the element size into the size in the fsinfo_attribute def and
don't set size for strings or opaques. Let a helper work that out.
This means that strings can actually get larger then 4K.
(*) A helper is provided to scan a list of attributes and call the
appropriate get function. This can be called from a filesystem's
->fsinfo() method multiple times. It also handles attribute
enumeration and info querying.
(*) Rearranged the patches to put all the notification patches first.
This allowed some of the bits to be squashed together. At some point,
I'll move the notification patches into a different branch.
ver #16:
(*) Split the features bits out of the fsinfo() core into their own patch
and got rid of the name encoding attributes.
(*) Renamed the 'array' type to 'list' and made AFS use it for returning
server address lists.
(*) Changed the ->fsinfo() method into an ->fsinfo_attributes[] table,
where each attribute has a ->get() method to deal with it. These
tables can then be returned with an fsinfo meta attribute.
(*) Dropped the fscontext query and parameter/description retrieval
attributes for now.
(*) Picked the mount topology attributes into this branch.
(*) Picked the mount notifications into this branch and rebased on top of
notifications-pipe-core.
(*) Picked the superblock notifications into this branch.
(*) Add sample code for Ext4 and NFS.
David
---
David Howells (17):
watch_queue: Add security hooks to rule on setting mount and sb watches
watch_queue: Implement mount topology and attribute change notifications
watch_queue: sample: Display mount tree change notifications
watch_queue: Introduce a non-repeating system-unique superblock ID
watch_queue: Add superblock notifications
watch_queue: sample: Display superblock notifications
fsinfo: Add fsinfo() syscall to query filesystem information
fsinfo: Provide a bitmap of supported features
fsinfo: Allow fsinfo() to look up a mount object by ID
fsinfo: Allow mount information to be queried
fsinfo: sample: Mount listing program
fsinfo: Allow the mount topology propogation flags to be retrieved
fsinfo: Query superblock unique ID and notification counter
fsinfo: Add API documentation
fsinfo: Add support for AFS
fsinfo: Add example support for Ext4
fsinfo: Add example support for NFS
Documentation/filesystems/fsinfo.rst | 491 ++++++++++++++++
arch/alpha/kernel/syscalls/syscall.tbl | 3
arch/arm/tools/syscall.tbl | 3
arch/arm64/include/asm/unistd.h | 2
arch/ia64/kernel/syscalls/syscall.tbl | 3
arch/m68k/kernel/syscalls/syscall.tbl | 3
arch/microblaze/kernel/syscalls/syscall.tbl | 3
arch/mips/kernel/syscalls/syscall_n32.tbl | 3
arch/mips/kernel/syscalls/syscall_n64.tbl | 3
arch/mips/kernel/syscalls/syscall_o32.tbl | 3
arch/parisc/kernel/syscalls/syscall.tbl | 3
arch/powerpc/kernel/syscalls/syscall.tbl | 3
arch/s390/kernel/syscalls/syscall.tbl | 3
arch/sh/kernel/syscalls/syscall.tbl | 3
arch/sparc/kernel/syscalls/syscall.tbl | 3
arch/x86/entry/syscalls/syscall_32.tbl | 3
arch/x86/entry/syscalls/syscall_64.tbl | 3
arch/xtensa/kernel/syscalls/syscall.tbl | 3
fs/Kconfig | 28 +
fs/Makefile | 2
fs/afs/internal.h | 1
fs/afs/super.c | 218 +++++++
fs/d_path.c | 2
fs/ext4/Makefile | 1
fs/ext4/ext4.h | 6
fs/ext4/fsinfo.c | 45 +
fs/ext4/super.c | 3
fs/fsinfo.c | 665 +++++++++++++++++++++
fs/internal.h | 12
fs/mount.h | 30 +
fs/mount_notify.c | 185 ++++++
fs/namespace.c | 323 ++++++++++
fs/nfs/Makefile | 1
fs/nfs/fsinfo.c | 230 +++++++
fs/nfs/internal.h | 6
fs/nfs/nfs4super.c | 3
fs/nfs/super.c | 3
fs/super.c | 156 +++++
include/linux/dcache.h | 1
include/linux/fs.h | 87 +++
include/linux/fsinfo.h | 110 ++++
include/linux/lsm_hooks.h | 24 +
include/linux/security.h | 16 +
include/linux/syscalls.h | 8
include/uapi/asm-generic/unistd.h | 8
include/uapi/linux/fsinfo.h | 361 ++++++++++++
include/uapi/linux/mount.h | 10
include/uapi/linux/watch_queue.h | 61 ++
include/uapi/linux/windows.h | 35 +
kernel/sys_ni.c | 7
samples/vfs/Makefile | 7
samples/vfs/test-fsinfo.c | 847 +++++++++++++++++++++++++++
samples/vfs/test-mntinfo.c | 243 ++++++++
samples/watch_queue/watch_test.c | 76 ++
security/security.c | 14
55 files changed, 4365 insertions(+), 11 deletions(-)
create mode 100644 Documentation/filesystems/fsinfo.rst
create mode 100644 fs/ext4/fsinfo.c
create mode 100644 fs/fsinfo.c
create mode 100644 fs/mount_notify.c
create mode 100644 fs/nfs/fsinfo.c
create mode 100644 include/linux/fsinfo.h
create mode 100644 include/uapi/linux/fsinfo.h
create mode 100644 include/uapi/linux/windows.h
create mode 100644 samples/vfs/test-fsinfo.c
create mode 100644 samples/vfs/test-mntinfo.c
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:02:05
Add security hooks that will allow an LSM to rule on whether or not a watch
may be set on a mount or on a superblock. More than one hook is required
as the watches watch different types of object.
Signed-off-by: David Howells <dhowells@redhat.com>
cc: Casey Schaufler <casey@schaufler-ca.com>
cc: Stephen Smalley <redacted>
cc: linux-security-module@vger.kernel.org
---
include/linux/lsm_hooks.h | 24 ++++++++++++++++++++++++
include/linux/security.h | 16 ++++++++++++++++
security/security.c | 14 ++++++++++++++
3 files changed, 54 insertions(+)
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:02:15
Add a mount notification facility whereby notifications about changes in
mount topology and configuration can be received. Note that this only
covers vfsmount topology changes and not superblock events. A separate
facility will be added for that.
Every mount is given a change counter than counts the number of topological
rearrangements in which it is involved and the number of attribute changes
it undergoes. This allows notification loss to be dealt with. Later
patches will provide a way to quickly retrieve this value, along with
information about topology and parameters for the superblock.
Firstly, an event queue needs to be created:
fd = open("/dev/event_queue", O_RDWR);
ioctl(fd, IOC_WATCH_QUEUE_SET_SIZE, page_size << n);
then a notification can be set up to report notifications via that queue:
struct watch_notification_filter filter = {
.nr_filters = 1,
.filters = {
[0] = {
.type = WATCH_TYPE_MOUNT_NOTIFY,
.subtype_filter[0] = UINT_MAX,
},
},
};
ioctl(fd, IOC_WATCH_QUEUE_SET_FILTER, &filter);
watch_mount(AT_FDCWD, "/", 0, fd, 0x02);
In this case, it would let me monitor the mount topology subtree rooted at
"/" for events. Mount notifications propagate up the tree towards the
root, so a watch will catch all of the events happening in the subtree
rooted at the watch.
After setting the watch, records will be placed into the queue when, for
example, as superblock switches between read-write and read-only. Records
are of the following format:
struct mount_notification {
struct watch_notification watch;
__u32 triggered_on;
__u32 changed_mount;
} *n;
Where:
n->watch.type will be WATCH_TYPE_MOUNT_NOTIFY.
n->watch.subtype will indicate the type of event, such as
NOTIFY_MOUNT_NEW_MOUNT.
n->watch.info & WATCH_INFO_LENGTH will indicate the length of the
record.
n->watch.info & WATCH_INFO_ID will be the fifth argument to
watch_mount(), shifted.
n->watch.info & NOTIFY_MOUNT_IN_SUBTREE if true indicates that the
notifcation was generated in the mount subtree rooted at the watch,
and not actually in the watch itself.
n->watch.info & NOTIFY_MOUNT_IS_RECURSIVE if true indicates that
the notifcation was generated by an event (eg. SETATTR) that was
applied recursively. The notification is only generated for the
object that initially triggered it.
n->watch.info & NOTIFY_MOUNT_IS_NOW_RO will be used for
NOTIFY_MOUNT_READONLY, being set if the superblock becomes R/O, and
being cleared otherwise, and for NOTIFY_MOUNT_NEW_MOUNT, being set
if the new mount is a submount (e.g. an automount).
n->watch.info & NOTIFY_MOUNT_IS_SUBMOUNT if true indicates that the
NOTIFY_MOUNT_NEW_MOUNT notification is in response to a mount
performed by the kernel (e.g. an automount).
n->triggered_on indicates the ID of the mount on which the watch
was installed.
n->changed_mount indicates the ID of the mount that was affected.
Note that it is permissible for event records to be of variable length -
or, at least, the length may be dependent on the subtype. Note also that
the queue can be shared between multiple notifications of various types.
Signed-off-by: David Howells <dhowells@redhat.com>
---
arch/alpha/kernel/syscalls/syscall.tbl | 1
arch/arm/tools/syscall.tbl | 1
arch/arm64/include/asm/unistd.h | 2
arch/ia64/kernel/syscalls/syscall.tbl | 1
arch/m68k/kernel/syscalls/syscall.tbl | 1
arch/microblaze/kernel/syscalls/syscall.tbl | 1
arch/mips/kernel/syscalls/syscall_n32.tbl | 1
arch/mips/kernel/syscalls/syscall_n64.tbl | 1
arch/mips/kernel/syscalls/syscall_o32.tbl | 1
arch/parisc/kernel/syscalls/syscall.tbl | 1
arch/powerpc/kernel/syscalls/syscall.tbl | 1
arch/s390/kernel/syscalls/syscall.tbl | 1
arch/sh/kernel/syscalls/syscall.tbl | 1
arch/sparc/kernel/syscalls/syscall.tbl | 1
arch/x86/entry/syscalls/syscall_32.tbl | 1
arch/x86/entry/syscalls/syscall_64.tbl | 1
arch/xtensa/kernel/syscalls/syscall.tbl | 1
fs/Kconfig | 9 +
fs/Makefile | 1
fs/mount.h | 30 ++++
fs/mount_notify.c | 185 +++++++++++++++++++++++++++
fs/namespace.c | 22 +++
include/linux/dcache.h | 1
include/linux/syscalls.h | 2
include/uapi/asm-generic/unistd.h | 4 -
include/uapi/linux/watch_queue.h | 32 +++++
kernel/sys_ni.c | 3
27 files changed, 304 insertions(+), 3 deletions(-)
create mode 100644 fs/mount_notify.c
@@ -477,3 +477,4 @@ # 545 reserved for clone3 547 common openat2 sys_openat2 548 common pidfd_getfd sys_pidfd_getfd+549 common watch_mount sys_watch_mount
@@ -451,3 +451,4 @@ 435 common clone3 sys_clone3 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd+439 common watch_mount sys_watch_mount
@@ -358,3 +358,4 @@ # 435 reserved for clone3 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd+439 common watch_mount sys_watch_mount
@@ -437,3 +437,4 @@ 435 common clone3 __sys_clone3 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd+439 common watch_mount sys_watch_mount
@@ -443,3 +443,4 @@ 435 common clone3 sys_clone3 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd+439 common watch_mount sys_watch_mount
@@ -435,3 +435,4 @@ 435 common clone3 sys_clone3_wrapper 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd+439 common watch_mount sys_watch_mount
@@ -440,3 +440,4 @@ # 435 reserved for clone3 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd+439 common watch_mount sys_watch_mount
@@ -483,3 +483,4 @@ # 435 reserved for clone3 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd+439 common watch_mount sys_watch_mount
@@ -359,6 +359,7 @@ 435 common clone3 __x64_sys_clone3/ptregs 437 common openat2 __x64_sys_openat2 438 common pidfd_getfd __x64_sys_pidfd_getfd+439 common watch_mount __x64_sys_watch_mount # # x32-specific system call numbers start at 512 to avoid cache impact
@@ -408,3 +408,4 @@ 435 common clone3 sys_clone3 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd+439 common watch_mount sys_watch_mount
@@ -72,6 +73,10 @@ struct mount {intmnt_expiry_mark;/* true if marked for expiry */structhlist_headmnt_pins;structhlist_headmnt_stuck_children;+atomic_tmnt_change_counter;/* Number of changed applied */+#ifdef CONFIG_MOUNT_NOTIFICATIONS+structwatch_list*mnt_watchers;/* Watches on dentries within this mount */+#endif}__randomize_layout;#define MNT_NS_INTERNAL ERR_PTR(-EINVAL) /* distinct from any mnt_namespace */
@@ -498,6 +498,9 @@ static int mnt_make_readonly(struct mount *mnt)smp_wmb();mnt->mnt.mnt_flags&=~MNT_WRITE_HOLD;unlock_mount_hash();+if(ret==0)+notify_mount(mnt,NULL,NOTIFY_MOUNT_READONLY,+NOTIFY_MOUNT_IS_NOW_RO);returnret;}
@@ -506,6 +509,7 @@ static int __mnt_unmake_readonly(struct mount *mnt)lock_mount_hash();mnt->mnt.mnt_flags&=~MNT_READONLY;unlock_mount_hash();+notify_mount(mnt,NULL,NOTIFY_MOUNT_READONLY,0);return0;}
@@ -819,6 +823,7 @@ static struct mountpoint *unhash_mnt(struct mount *mnt)*/staticvoidumount_mnt(structmount*mnt){+notify_mount(mnt->mnt_parent,mnt,NOTIFY_MOUNT_UNMOUNT,0);put_mountpoint(unhash_mnt(mnt));}
@@ -1159,6 +1164,11 @@ static void mntput_no_expire(struct mount *mnt)mnt->mnt.mnt_flags|=MNT_DOOMED;rcu_read_unlock();+#ifdef CONFIG_MOUNT_NOTIFICATIONS+if(mnt->mnt_watchers)+remove_watch_list(mnt->mnt_watchers,mnt->mnt_id);+#endif+list_del(&mnt->mnt_instance);if(unlikely(!list_empty(&mnt->mnt_mounts))){
@@ -2078,7 +2089,10 @@ static int attach_recursive_mnt(struct mount *source_mnt,lock_mount_hash();}if(moving){+notify_mount(source_mnt->mnt_parent,source_mnt,+NOTIFY_MOUNT_MOVE_FROM,0);unhash_mnt(source_mnt);+notify_mount(dest_mnt,source_mnt,NOTIFY_MOUNT_MOVE_TO,0);attach_mnt(source_mnt,dest_mnt,dest_mp);touch_mnt_namespace(source_mnt->mnt_ns);}else{
@@ -2087,6 +2101,11 @@ static int attach_recursive_mnt(struct mount *source_mnt,list_del_init(&source_mnt->mnt_ns->list);}mnt_set_mountpoint(dest_mnt,dest_mp,source_mnt);+notify_mount(dest_mnt,source_mnt,NOTIFY_MOUNT_NEW_MOUNT,+(source_mnt->mnt.mnt_sb->s_flags&SB_RDONLY?+NOTIFY_MOUNT_IS_NOW_RO:0)|+(source_mnt->mnt.mnt_sb->s_flags&SB_SUBMOUNT?+NOTIFY_MOUNT_IS_SUBMOUNT:0));commit_tree(source_mnt);}
@@ -2464,6 +2483,8 @@ static void set_mount_attributes(struct mount *mnt, unsigned int mnt_flags)mnt->mnt.mnt_flags=mnt_flags;touch_mnt_namespace(mnt->mnt_ns);unlock_mount_hash();+notify_mount(mnt,NULL,NOTIFY_MOUNT_SETATTR,+(mnt_flags&SB_RDONLY?NOTIFY_MOUNT_IS_NOW_RO:0));}staticvoidmnt_warn_timestamp_expiry(structpath*mountpoint,structvfsmount*mnt)
@@ -1003,6 +1003,8 @@ asmlinkage long sys_pidfd_send_signal(int pidfd, int sig,siginfo_t__user*info,unsignedintflags);asmlinkagelongsys_pidfd_getfd(intpidfd,intfd,unsignedintflags);+asmlinkagelongsys_watch_mount(intdfd,constchar__user*path,+unsignedintat_flags,intwatch_fd,intwatch_id);/**Architecture-specificsystemcalls
@@ -14,7 +14,8 @@enumwatch_notification_type{WATCH_TYPE_META=0,/* Special record */WATCH_TYPE_KEY_NOTIFY=1,/* Key change event notification */-WATCH_TYPE__NR=2+WATCH_TYPE_MOUNT_NOTIFY=2,/* Mount topology change notification */+WATCH_TYPE___NR=3};enumwatch_meta_notification_subtype{
@@ -101,4 +102,33 @@ struct key_notification {__u32aux;/* Per-type auxiliary data */};+/*+*Typeofmounttopologychangenotification.+*/+enummount_notification_subtype{+NOTIFY_MOUNT_NEW_MOUNT=0,/* New mount added */+NOTIFY_MOUNT_UNMOUNT=1,/* Mount removed manually */+NOTIFY_MOUNT_EXPIRY=2,/* Automount expired */+NOTIFY_MOUNT_READONLY=3,/* Mount R/O state changed */+NOTIFY_MOUNT_SETATTR=4,/* Mount attributes changed */+NOTIFY_MOUNT_MOVE_FROM=5,/* Mount moved from here */+NOTIFY_MOUNT_MOVE_TO=6,/* Mount moved to here (compare op_id) */+};++#define NOTIFY_MOUNT_IN_SUBTREE WATCH_INFO_FLAG_0 /* Event not actually at watched dentry */+#define NOTIFY_MOUNT_IS_RECURSIVE WATCH_INFO_FLAG_1 /* Change applied recursively */+#define NOTIFY_MOUNT_IS_NOW_RO WATCH_INFO_FLAG_2 /* Mount changed to R/O */+#define NOTIFY_MOUNT_IS_SUBMOUNT WATCH_INFO_FLAG_3 /* New mount is submount */++/*+*Mounttopology/configurationchangenotificationrecord.+*-watch.type=WATCH_TYPE_MOUNT_NOTIFY+*-watch.subtype=enummount_notification_subtype+*/+structmount_notification{+structwatch_notificationwatch;/* WATCH_TYPE_MOUNT_NOTIFY */+__u32triggered_on;/* The mount that the notify was on */+__u32changed_mount;/* The mount that got changed */+};+#endif /* _UAPI_LINUX_WATCH_QUEUE_H */
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:02:19
This is run like:
./watch_test
and watches "/" for changes to the mount topology and the attributes of
individual mount objects.
# mount -t tmpfs none /mnt
# mount -o remount,ro /mnt
# mount -o remount,rw /mnt
producing:
# ./watch_test
read() = 16
NOTIFY[000]: ty=000002 sy=00 i=02000010
MOUNT 00000060 change=0[new_mount] aux=416
read() = 16
NOTIFY[000]: ty=000002 sy=04 i=02010010
MOUNT 000001a0 change=4[setattr] aux=0
read() = 16
NOTIFY[000]: ty=000002 sy=04 i=02010010
MOUNT 000001a0 change=4[setattr] aux=0
Signed-off-by: David Howells <dhowells@redhat.com>
---
samples/watch_queue/watch_test.c | 39 +++++++++++++++++++++++++++++++++++++-
1 file changed, 38 insertions(+), 1 deletion(-)
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:02:29
Introduce an (effectively) non-repeating system-unique superblock ID that
can be used to determine that two object are in the same superblock without
risking reuse of the ID in the meantime (as is possible with device IDs).
The ID is time-based to make it harder to use it as a covert communications
channel.
In future patches, this ID will be used to tag superblock notification
messages. It will also be made queryable.
Signed-off-by: David Howells <dhowells@redhat.com>
---
fs/super.c | 24 ++++++++++++++++++++++++
include/linux/fs.h | 3 +++
2 files changed, 27 insertions(+)
@@ -1548,6 +1548,9 @@ struct super_block {spinlock_ts_inode_wblist_lock;structlist_heads_inodes_wb;/* writeback inodes */++/* Superblock event notifications */+u64s_unique_id;}__randomize_layout;/* Helper functions so that in most cases filesystems will
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:02:40
Add a superblock event notification facility whereby notifications about
superblock events, such as I/O errors (EIO), quota limits being hit
(EDQUOT) and running out of space (ENOSPC) can be reported to a monitoring
process asynchronously. Note that this does not cover vfsmount topology
changes. watch_mount() is used for that.
Firstly, an event queue needs to be created:
fd = open("/dev/event_queue", O_RDWR);
ioctl(fd, IOC_WATCH_QUEUE_SET_SIZE, page_size << n);
then a notification can be set up to report notifications via that queue:
struct watch_notification_filter filter = {
.nr_filters = 1,
.filters = {
[0] = {
.type = WATCH_TYPE_SB_NOTIFY,
.subtype_filter[0] = UINT_MAX,
},
},
};
ioctl(fd, IOC_WATCH_QUEUE_SET_FILTER, &filter);
watch_sb(AT_FDCWD, "/home/dhowells", 0, fd, 0x03);
In this case, it would let me monitor my own homedir for events. After
setting the watch, records will be placed into the queue when, for example,
as superblock switches between read-write and read-only. Records are of
the following format:
struct superblock_notification {
struct watch_notification watch;
__u64 sb_id;
} *n;
Where:
n->watch.type will be WATCH_TYPE_SB_NOTIFY.
n->watch.subtype will indicate the type of event, such as
NOTIFY_SUPERBLOCK_READONLY.
n->watch.info & WATCH_INFO_LENGTH will indicate the length of the
record.
n->watch.info & WATCH_INFO_ID will be the fifth argument to
watch_sb(), shifted.
n->watch.info & NOTIFY_SUPERBLOCK_IS_NOW_RO will be used for
NOTIFY_SUPERBLOCK_READONLY, being set if the superblock becomes
R/O, and being cleared otherwise.
n->sb_id will be the ID of the superblock, as can be retrieved with
the fsinfo() syscall, as part of the fsinfo_sb_notifications
attribute in the the watch_id field.
Note that it is permissible for event records to be of variable length -
or, at least, the length may be dependent on the subtype. Note also that
the queue can be shared between multiple notifications of various types.
Signed-off-by: David Howells <dhowells@redhat.com>
---
arch/alpha/kernel/syscalls/syscall.tbl | 1
arch/arm/tools/syscall.tbl | 1
arch/arm64/include/asm/unistd.h | 2
arch/ia64/kernel/syscalls/syscall.tbl | 1
arch/m68k/kernel/syscalls/syscall.tbl | 1
arch/microblaze/kernel/syscalls/syscall.tbl | 1
arch/mips/kernel/syscalls/syscall_n32.tbl | 1
arch/mips/kernel/syscalls/syscall_n64.tbl | 1
arch/mips/kernel/syscalls/syscall_o32.tbl | 1
arch/parisc/kernel/syscalls/syscall.tbl | 1
arch/powerpc/kernel/syscalls/syscall.tbl | 1
arch/s390/kernel/syscalls/syscall.tbl | 1
arch/sh/kernel/syscalls/syscall.tbl | 1
arch/sparc/kernel/syscalls/syscall.tbl | 1
arch/x86/entry/syscalls/syscall_32.tbl | 1
arch/x86/entry/syscalls/syscall_64.tbl | 1
arch/xtensa/kernel/syscalls/syscall.tbl | 1
fs/Kconfig | 12 ++
fs/super.c | 132 +++++++++++++++++++++++++++
include/linux/fs.h | 80 ++++++++++++++++
include/linux/syscalls.h | 2
include/uapi/asm-generic/unistd.h | 4 +
include/uapi/linux/watch_queue.h | 31 ++++++
kernel/sys_ni.c | 3 +
24 files changed, 279 insertions(+), 3 deletions(-)
@@ -478,3 +478,4 @@ 547 common openat2 sys_openat2 548 common pidfd_getfd sys_pidfd_getfd 549 common watch_mount sys_watch_mount+550 common watch_sb sys_watch_sb
@@ -452,3 +452,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -359,3 +359,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -438,3 +438,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -444,3 +444,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -436,3 +436,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -520,3 +520,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -441,3 +441,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -484,3 +484,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -360,6 +360,7 @@ 437 common openat2 __x64_sys_openat2 438 common pidfd_getfd __x64_sys_pidfd_getfd 439 common watch_mount __x64_sys_watch_mount+440 common watch_sb __x64_sys_watch_sb # # x32-specific system call numbers start at 512 to avoid cache impact
@@ -409,3 +409,4 @@ 437 common openat2 sys_openat2 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount+440 common watch_sb sys_watch_sb
@@ -993,6 +999,8 @@ int reconfigure_super(struct fs_context *fc)/* Needs to be ordered wrt mnt_is_readonly() */smp_wmb();sb->s_readonly_remount=0;+notify_sb(sb,NOTIFY_SUPERBLOCK_READONLY,+remount_ro?NOTIFY_SUPERBLOCK_IS_NOW_RO:0);/**Somefilesystemsmodifytheirmetadataviasomeotherpaththanthe
@@ -1891,3 +1899,127 @@ int thaw_super(struct super_block *sb)returnthaw_super_locked(sb);}EXPORT_SYMBOL(thaw_super);++#ifdef CONFIG_SB_NOTIFICATIONS+/*+*Postsuperblocknotifications.+*/+voidpost_sb_notification(structsuper_block*s,structsuperblock_notification*n)+{+post_watch_notification(s->s_watchers,&n->watch,current_cred(),+s->s_unique_id);+}++staticvoidsb_release_watch(structwatch*watch)+{+put_super(watch->private);+}++/**+*sys_watch_sb-Watchforsuperblockevents.+*@dfd:Basedirectorytopathwalkfromorfdreferringtosuperblock.+*@filename:Pathtosuperblocktoplacethewatchupon+*@at_flags:Pathwalkcontrolflags+*@watch_fd:Thewatchqueuetosendnotificationsto.+*@watch_id:ThewatchIDtobeplacedinthenotification(-1toremovewatch)+*/+SYSCALL_DEFINE5(watch_sb,+int,dfd,+constchar__user*,filename,+unsignedint,at_flags,+int,watch_fd,+int,watch_id)+{+structwatch_queue*wqueue;+structsuper_block*s;+structwatch_list*wlist=NULL;+structwatch*watch=NULL;+structpathpath;+unsignedintlookup_flags=+LOOKUP_DIRECTORY|LOOKUP_FOLLOW|LOOKUP_AUTOMOUNT;+booldrop_s_count=false;+intret;++if(watch_id<-1||watch_id>0xff)+return-EINVAL;+if((at_flags&~(AT_NO_AUTOMOUNT|AT_EMPTY_PATH))!=0)+return-EINVAL;+if(at_flags&AT_NO_AUTOMOUNT)+lookup_flags&=~LOOKUP_AUTOMOUNT;+if(at_flags&AT_EMPTY_PATH)+lookup_flags|=LOOKUP_EMPTY;++ret=user_path_at(dfd,filename,at_flags,&path);+if(ret)+returnret;++ret=inode_permission(path.dentry->d_inode,MAY_EXEC);+if(ret)+gotoerr_path;++wqueue=get_watch_queue(watch_fd);+if(IS_ERR(wqueue))+gotoerr_path;++s=path.dentry->d_sb;+if(watch_id>=0){+ret=-ENOMEM;+if(!READ_ONCE(s->s_watchers)){+wlist=kzalloc(sizeof(*wlist),GFP_KERNEL);+if(!wlist)+gotoerr_wqueue;+init_watch_list(wlist,sb_release_watch);+}++watch=kzalloc(sizeof(*watch),GFP_KERNEL);+if(!watch)+gotoerr_wlist;++init_watch(watch,wqueue);+watch->id=s->s_unique_id;+watch->private=s;+watch->info_id=(u32)watch_id<<24;++ret=security_watch_sb(watch,s);+if(ret<0)+gotoerr_watch;++down_write(&s->s_umount);+ret=-EIO;+if(atomic_read(&s->s_active)){+if(!s->s_watchers){+s->s_watchers=wlist;+wlist=NULL;+}++spin_lock(&sb_lock);+s->s_count++;+spin_unlock(&sb_lock);+ret=add_watch_to_object(watch,s->s_watchers);+if(ret==0)+watch=NULL;/* It worked */+else+drop_s_count=true;+}+up_write(&s->s_umount);+if(drop_s_count)+put_super(s);+}else{+ret=-EBADSLT;+down_write(&s->s_umount);+ret=remove_watch_from_object(s->s_watchers,wqueue,+s->s_unique_id,false);+up_write(&s->s_umount);+}++err_watch:+kfree(watch);+err_wlist:+kfree(wlist);+err_wqueue:+put_watch_queue(wqueue);+err_path:+path_put(&path);+returnret;+}+#endif
@@ -1551,6 +1552,11 @@ struct super_block {/* Superblock event notifications */u64s_unique_id;++#ifdef CONFIG_SB_NOTIFICATIONS+structwatch_list*s_watchers;+#endif+atomic_ts_notify_counter;}__randomize_layout;/* Helper functions so that in most cases filesystems will
@@ -1005,6 +1005,8 @@ asmlinkage long sys_pidfd_send_signal(int pidfd, int sig,asmlinkagelongsys_pidfd_getfd(intpidfd,intfd,unsignedintflags);asmlinkagelongsys_watch_mount(intdfd,constchar__user*path,unsignedintat_flags,intwatch_fd,intwatch_id);+asmlinkagelongsys_watch_sb(intdfd,constchar__user*path,+unsignedintat_flags,intwatch_fd,intwatch_id);/**Architecture-specificsystemcalls
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:02:55
Add a system call to allow filesystem information to be queried. A request
value can be given to indicate the desired attribute. Support is provided
for enumerating multi-value attributes.
===============
NEW SYSTEM CALL
===============
The new system call looks like:
int ret = fsinfo(int dfd,
const char *filename,
const struct fsinfo_params *params,
void *buffer,
size_t buf_size);
The params parameter optionally points to a block of parameters:
struct fsinfo_params {
__u32 at_flags;
__u32 flags;
__u32 request;
__u32 Nth;
__u32 Mth;
__u64 __reserved[3];
};
If params is NULL, it is assumed params->request should be
FSINFO_ATTR_STATFS, params->Nth should be 0, params->Mth should be 0,
params->at_flags should be 0 and params->flags should be 0.
If params is given, all of params->__reserved[] must be 0.
dfd, filename and params->at_flags indicate the file to query. There is no
equivalent of lstat() as that can be emulated with fsinfo() by setting
AT_SYMLINK_NOFOLLOW in params->at_flags. There is also no equivalent of
fstat() as that can be emulated by passing a NULL filename to fsinfo() with
the fd of interest in dfd. AT_NO_AUTOMOUNT can also be used to an allow
automount point to be queried without triggering it.
params->request indicates the attribute/attributes to be queried. This can
be one of:
FSINFO_ATTR_STATFS - statfs-style info
FSINFO_ATTR_IDS - Filesystem IDs
FSINFO_ATTR_LIMITS - Filesystem limits
FSINFO_ATTR_SUPPORTS - What's supported in statx(), IOC flags
FSINFO_ATTR_TIMESTAMP_INFO - Inode timestamp info
FSINFO_ATTR_VOLUME_ID - Volume ID (string)
FSINFO_ATTR_VOLUME_UUID - Volume UUID
FSINFO_ATTR_VOLUME_NAME - Volume name (string)
FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO - Information about attr Nth
FSINFO_ATTR_FSINFO_ATTRIBUTES - List of supported attrs
Some attributes (such as the servers backing a network filesystem) can have
multiple values. These can be enumerated by setting params->Nth and
params->Mth to 0, 1, ... until ENODATA is returned.
buffer and buf_size point to the reply buffer. The buffer is filled up to
the specified size, even if this means truncating the reply. The full size
of the reply is returned. In future versions, this will allow extra fields
to be tacked on to the end of the reply, but anyone not expecting them will
only get the subset they're expecting. If either buffer of buf_size are 0,
no copy will take place and the data size will be returned.
At the moment, this will only work on x86_64 and i386 as it requires the
system call to be wired up.
Signed-off-by: David Howells <dhowells@redhat.com>
cc: linux-api@vger.kernel.org
---
arch/alpha/kernel/syscalls/syscall.tbl | 1
arch/arm/tools/syscall.tbl | 1
arch/arm64/include/asm/unistd.h | 2
arch/ia64/kernel/syscalls/syscall.tbl | 1
arch/m68k/kernel/syscalls/syscall.tbl | 1
arch/microblaze/kernel/syscalls/syscall.tbl | 1
arch/mips/kernel/syscalls/syscall_n32.tbl | 1
arch/mips/kernel/syscalls/syscall_n64.tbl | 1
arch/mips/kernel/syscalls/syscall_o32.tbl | 1
arch/parisc/kernel/syscalls/syscall.tbl | 1
arch/powerpc/kernel/syscalls/syscall.tbl | 1
arch/s390/kernel/syscalls/syscall.tbl | 1
arch/sh/kernel/syscalls/syscall.tbl | 1
arch/sparc/kernel/syscalls/syscall.tbl | 1
arch/x86/entry/syscalls/syscall_32.tbl | 1
arch/x86/entry/syscalls/syscall_64.tbl | 1
arch/xtensa/kernel/syscalls/syscall.tbl | 1
fs/Kconfig | 7
fs/Makefile | 1
fs/fsinfo.c | 566 +++++++++++++++++++++++++
include/linux/fs.h | 4
include/linux/fsinfo.h | 72 +++
include/linux/syscalls.h | 4
include/uapi/asm-generic/unistd.h | 4
include/uapi/linux/fsinfo.h | 187 ++++++++
kernel/sys_ni.c | 1
samples/vfs/Makefile | 5
samples/vfs/test-fsinfo.c | 607 +++++++++++++++++++++++++++
28 files changed, 1474 insertions(+), 2 deletions(-)
create mode 100644 fs/fsinfo.c
create mode 100644 include/linux/fsinfo.h
create mode 100644 include/uapi/linux/fsinfo.h
create mode 100644 samples/vfs/test-fsinfo.c
@@ -479,3 +479,4 @@ 548 common pidfd_getfd sys_pidfd_getfd 549 common watch_mount sys_watch_mount 550 common watch_sb sys_watch_sb+551 common fsinfo sys_fsinfo
@@ -453,3 +453,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -360,3 +360,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -439,3 +439,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -445,3 +445,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -437,3 +437,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -521,3 +521,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -442,3 +442,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -485,3 +485,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -361,6 +361,7 @@ 438 common pidfd_getfd __x64_sys_pidfd_getfd 439 common watch_mount __x64_sys_watch_mount 440 common watch_sb __x64_sys_watch_sb+441 common fsinfo __x64_sys_fsinfo # # x32-specific system call numbers start at 512 to avoid cache impact
@@ -410,3 +410,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -15,6 +15,13 @@ config VALIDATE_FS_PARSEREnablethistoperformvalidationoftheparameterdescriptionforafilesystemwhenitisregistered.+configFSINFO+bool"Enable the fsinfo() system call"+help+Enablethefilesysteminformationqueryingsystemcalltoallow+comprehensiveinformationtoberetrievedaboutafilesystem,+superblockormountobject.+ifBLOCKconfigFS_IOMAP
@@ -0,0 +1,566 @@+// SPDX-License-Identifier: GPL-2.0+/* Filesystem information query.+*+*Copyright(C)2020RedHat,Inc.AllRightsReserved.+*WrittenbyDavidHowells(dhowells@redhat.com)+*/+#include<linux/syscalls.h>+#include<linux/fs.h>+#include<linux/file.h>+#include<linux/mount.h>+#include<linux/namei.h>+#include<linux/statfs.h>+#include<linux/security.h>+#include<linux/uaccess.h>+#include<linux/fsinfo.h>+#include<uapi/linux/mount.h>+#include"internal.h"++/**+*fsinfo_string-StoreaNUL-terminatedstringasanfsinfoattributevalue.+*@s:Thestringtostore(maybeNULL)+*@ctx:Theparametercontext+*/+intfsinfo_string(constchar*s,structfsinfo_context*ctx)+{+unsignedintlen;+char*p=ctx->buffer;+intret=0;++if(s){+len=min_t(size_t,strlen(s),ctx->buf_size-1);+if(!ctx->want_size_only){+memcpy(p,s,len);+p[len]=0;+}+ret=len;+}++returnret;+}+EXPORT_SYMBOL(fsinfo_string);++/*+*Getbasicfilesystemstatsfromstatfs.+*/+staticintfsinfo_generic_statfs(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_statfs*p=ctx->buffer;+structkstatfsbuf;+intret;++ret=vfs_statfs(path,&buf);+if(ret<0)+returnret;++p->f_blocks.lo=buf.f_blocks;+p->f_bfree.lo=buf.f_bfree;+p->f_bavail.lo=buf.f_bavail;+p->f_files.lo=buf.f_files;+p->f_ffree.lo=buf.f_ffree;+p->f_favail.lo=buf.f_ffree;+p->f_bsize=buf.f_bsize;+p->f_frsize=buf.f_frsize;+returnsizeof(*p);+}++staticintfsinfo_generic_ids(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_ids*p=ctx->buffer;+structsuper_block*sb;+structkstatfsbuf;+intret;++ret=vfs_statfs(path,&buf);+if(ret<0&&ret!=-ENOSYS)+returnret;+if(ret==0)+memcpy(&p->f_fsid,&buf.f_fsid,sizeof(p->f_fsid));++sb=path->dentry->d_sb;+p->f_fstype=sb->s_magic;+p->f_dev_major=MAJOR(sb->s_dev);+p->f_dev_minor=MINOR(sb->s_dev);+p->f_sb_id=sb->s_unique_id;+strlcpy(p->f_fs_name,sb->s_type->name,sizeof(p->f_fs_name));+returnsizeof(*p);+}++intfsinfo_generic_limits(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_limits*p=ctx->buffer;+structsuper_block*sb=path->dentry->d_sb;++p->max_file_size.hi=0;+p->max_file_size.lo=sb->s_maxbytes;+p->max_ino.hi=0;+p->max_ino.lo=UINT_MAX;+p->max_hard_links=sb->s_max_links;+p->max_uid=UINT_MAX;+p->max_gid=UINT_MAX;+p->max_projid=UINT_MAX;+p->max_filename_len=NAME_MAX;+p->max_symlink_len=PATH_MAX;+p->max_xattr_name_len=XATTR_NAME_MAX;+p->max_xattr_body_len=XATTR_SIZE_MAX;+p->max_dev_major=0xffffff;+p->max_dev_minor=0xff;+returnsizeof(*p);+}+EXPORT_SYMBOL(fsinfo_generic_limits);++intfsinfo_generic_supports(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_supports*p=ctx->buffer;+structsuper_block*sb=path->dentry->d_sb;++p->stx_mask=STATX_BASIC_STATS;+if(sb->s_d_op&&sb->s_d_op->d_automount)+p->stx_attributes|=STATX_ATTR_AUTOMOUNT;+returnsizeof(*p);+}+EXPORT_SYMBOL(fsinfo_generic_supports);++staticconststructfsinfo_timestamp_infofsinfo_default_timestamp_info={+.atime={+.minimum=S64_MIN,+.maximum=S64_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.mtime={+.minimum=S64_MIN,+.maximum=S64_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.ctime={+.minimum=S64_MIN,+.maximum=S64_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.btime={+.minimum=S64_MIN,+.maximum=S64_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+};++intfsinfo_generic_timestamp_info(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_timestamp_info*p=ctx->buffer;+structsuper_block*sb=path->dentry->d_sb;+s8exponent;++*p=fsinfo_default_timestamp_info;++if(sb->s_time_gran<1000000000){+if(sb->s_time_gran<1000)+exponent=-9;+elseif(sb->s_time_gran<1000000)+exponent=-6;+else+exponent=-3;++p->atime.gran_exponent=exponent;+p->mtime.gran_exponent=exponent;+p->ctime.gran_exponent=exponent;+p->btime.gran_exponent=exponent;+}++returnsizeof(*p);+}+EXPORT_SYMBOL(fsinfo_generic_timestamp_info);++staticintfsinfo_generic_volume_uuid(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_volume_uuid*p=ctx->buffer;+structsuper_block*sb=path->dentry->d_sb;++memcpy(p,&sb->s_uuid,sizeof(*p));+returnsizeof(*p);+}++staticintfsinfo_generic_volume_id(structpath*path,structfsinfo_context*ctx)+{+returnfsinfo_string(path->dentry->d_sb->s_id,ctx);+}++staticconststructfsinfo_attributefsinfo_common_attributes[]={+FSINFO_VSTRUCT(FSINFO_ATTR_STATFS,fsinfo_generic_statfs),+FSINFO_VSTRUCT(FSINFO_ATTR_IDS,fsinfo_generic_ids),+FSINFO_VSTRUCT(FSINFO_ATTR_LIMITS,fsinfo_generic_limits),+FSINFO_VSTRUCT(FSINFO_ATTR_SUPPORTS,fsinfo_generic_supports),+FSINFO_VSTRUCT(FSINFO_ATTR_TIMESTAMP_INFO,fsinfo_generic_timestamp_info),+FSINFO_STRING(FSINFO_ATTR_VOLUME_ID,fsinfo_generic_volume_id),+FSINFO_VSTRUCT(FSINFO_ATTR_VOLUME_UUID,fsinfo_generic_volume_uuid),++FSINFO_LIST(FSINFO_ATTR_FSINFO_ATTRIBUTES,(void*)123UL),+FSINFO_VSTRUCT_N(FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO,(void*)123UL),+{}+};++/*+*Determineanattribute'sminimumbuffersizeand,ifthebufferislarge+*enough,gettheattributevalue.+*/+staticintfsinfo_get_this_attribute(structpath*path,+structfsinfo_context*ctx,+conststructfsinfo_attribute*attr)+{+intbuf_size;++if(ctx->Nth!=0&&!(attr->flags&(FSINFO_FLAGS_N|FSINFO_FLAGS_NM)))+return-ENODATA;+if(ctx->Mth!=0&&!(attr->flags&FSINFO_FLAGS_NM))+return-ENODATA;++switch(attr->type){+caseFSINFO_TYPE_VSTRUCT:+ctx->clear_tail=true;+buf_size=attr->size;+break;+caseFSINFO_TYPE_STRING:+caseFSINFO_TYPE_OPAQUE:+caseFSINFO_TYPE_LIST:+buf_size=4096;+break;+default:+return-ENOPKG;+}++if(ctx->buf_size<buf_size)+returnbuf_size;++returnattr->get(path,ctx);+}++staticvoidfsinfo_attributes_insert(structfsinfo_context*ctx,+conststructfsinfo_attribute*attr)+{+__u32*p=ctx->buffer;+unsignedinti;++if(ctx->usage>=ctx->buf_size||+ctx->buf_size-ctx->usage<sizeof(__u32)){+ctx->usage+=sizeof(__u32);+return;+}++for(i=0;i<ctx->usage/sizeof(__u32);i++)+if(p[i]==attr->attr_id)+return;++p[i]=attr->attr_id;+ctx->usage+=sizeof(__u32);+}++staticintfsinfo_list_attributes(structpath*path,+structfsinfo_context*ctx,+conststructfsinfo_attribute*attributes)+{+conststructfsinfo_attribute*a;++for(a=attributes;a->get;a++)+fsinfo_attributes_insert(ctx,a);+return-EOPNOTSUPP;/* We want to go through all the lists */+}++staticintfsinfo_get_attribute_info(structpath*path,+structfsinfo_context*ctx,+conststructfsinfo_attribute*attributes)+{+conststructfsinfo_attribute*a;+structfsinfo_attribute_info*p=ctx->buffer;++if(!ctx->buf_size)+returnsizeof(*p);++for(a=attributes;a->get;a++){+if(a->attr_id==ctx->Nth){+p->attr_id=a->attr_id;+p->type=a->type;+p->flags=a->flags;+p->size=a->size;+p->size=a->size;+returnsizeof(*p);+}+}+return-EOPNOTSUPP;/* We want to go through all the lists */+}++/**+*fsinfo_get_attribute-Lookupandhandleanattribute+*@path:Theobjecttoquery+*@params:Parameterstodefinearequestandplacetostoreresult+*@attributes:Listofattributestosearch.+*+*Lookthroughalistofattributesforonethatmatchestherequested+*attributethencallthehandlerforit.+*/+intfsinfo_get_attribute(structpath*path,structfsinfo_context*ctx,+conststructfsinfo_attribute*attributes)+{+conststructfsinfo_attribute*a;++switch(ctx->requested_attr){+caseFSINFO_ATTR_FSINFO_ATTRIBUTE_INFO:+returnfsinfo_get_attribute_info(path,ctx,attributes);+caseFSINFO_ATTR_FSINFO_ATTRIBUTES:+returnfsinfo_list_attributes(path,ctx,attributes);+default:+for(a=attributes;a->get;a++)+if(a->attr_id==ctx->requested_attr)+returnfsinfo_get_this_attribute(path,ctx,a);+return-EOPNOTSUPP;+}+}+EXPORT_SYMBOL(fsinfo_get_attribute);++/**+*generic_fsinfo-Handleanfsinfoattributegenerically+*@path:Theobjecttoquery+*@params:Parameterstodefinearequestandplacetostoreresult+*/+staticintfsinfo_call(structpath*path,structfsinfo_context*ctx)+{+intret;++if(path->dentry->d_sb->s_op->fsinfo){+ret=path->dentry->d_sb->s_op->fsinfo(path,ctx);+if(ret!=-EOPNOTSUPP)+returnret;+}+ret=fsinfo_get_attribute(path,ctx,fsinfo_common_attributes);+if(ret!=-EOPNOTSUPP)+returnret;++switch(ctx->requested_attr){+caseFSINFO_ATTR_FSINFO_ATTRIBUTE_INFO:+return-ENODATA;+caseFSINFO_ATTR_FSINFO_ATTRIBUTES:+returnctx->usage;+default:+return-EOPNOTSUPP;+}+}++/**+*vfs_fsinfo-Retrievefilesysteminformation+*@path:Theobjecttoquery+*@params:Parameterstodefinearequestandplacetostoreresult+*+*Getanattributeonafilesystemoranobjectwithinafilesystem.The+*filesystemattributetobequeriedisindicatedby@ctx->requested_attr,and+*ifit'samulti-valuedattribute,theparticularvalueisselectedby+*@ctx->Nthandthen@ctx->Mth.+*+*Forcommonattributes,avaluemaybefabricatedifitisnotsupportedby+*thefilesystem.+*+*Onsuccess,thesizeoftheattribute'svalueisreturned(0isavalid+*size).Abufferwillhavebeenallocatedandwillbepointedtoby+*@ctx->buffer.Thecallermustfreethiswithkvfree().+*+*Errorscanalsobereturned:-ENOMEMifabuffercannotbeallocated,-EPERM+*or-EACCESifpermissionisdeniedbytheLSM,-EOPNOTSUPPifanattribute+*doesn'texistforthespecifiedobjector-ENODATAiftheattributeexists,+*buttheNth,Mthvaluedoesnotexist.-EMSGSIZEindicatesthatthevalueis+*unmanageableinternallyand-ENOPKGindicatesotherinternalfailure.+*+*Errorssuchas-EIOmayalsocomefromattemptstoaccessmediaorservers+*toobtaintherequestedinformationifit'snotimmediatelytohand.+*+*[*]Notethatthecallermayset@ctx->want_size_onlyifitonlywantsthe+*sizeofthevalueandnotthedata.Ifthisisset,abuffermaynotbe+*allocatedundersomecircumstances.Thisisintendedforsizequeryby+*userspace.+*+*[*]Notethat@ctx->clear_tailwillbereturnedsetifthedatashouldbe+*paddedoutwithzeroswhenwritingittouserspace.+*/+staticintvfs_fsinfo(structpath*path,structfsinfo_context*ctx)+{+structdentry*dentry=path->dentry;+intret;++ret=security_sb_statfs(dentry);+if(ret)+returnret;++/* Call the handler to find out the buffer size required. */+ctx->buf_size=0;+ret=fsinfo_call(path,ctx);+if(ret<0||ctx->want_size_only)+returnret;+ctx->buf_size=ret;++do{+/* Allocate a buffer of the requested size. */+if(ctx->buf_size>INT_MAX)+return-EMSGSIZE;+ctx->buffer=kvzalloc(ctx->buf_size,GFP_KERNEL);+if(!ctx->buffer)+return-ENOMEM;++ctx->usage=0;+ret=fsinfo_call(path,ctx);+if(IS_ERR_VALUE((long)ret))+returnret;+if((unsignedint)ret<=ctx->buf_size)+returnret;/* It fitted */++/* We need to resize the buffer */+ctx->buf_size=roundup(ret,PAGE_SIZE);+kvfree(ctx->buffer);+ctx->buffer=NULL;+}while(!signal_pending(current));++return-ERESTARTSYS;+}++staticintvfs_fsinfo_path(intdfd,constchar__user*pathname,+unsignedintat_flags,structfsinfo_context*ctx)+{+structpathpath;+unsignedlookup_flags=LOOKUP_FOLLOW|LOOKUP_AUTOMOUNT;+intret=-EINVAL;++if((at_flags&~(AT_SYMLINK_NOFOLLOW|AT_NO_AUTOMOUNT|+AT_EMPTY_PATH))!=0)+return-EINVAL;++if(at_flags&AT_SYMLINK_NOFOLLOW)+lookup_flags&=~LOOKUP_FOLLOW;+if(at_flags&AT_NO_AUTOMOUNT)+lookup_flags&=~LOOKUP_AUTOMOUNT;+if(at_flags&AT_EMPTY_PATH)+lookup_flags|=LOOKUP_EMPTY;++retry:+ret=user_path_at(dfd,pathname,lookup_flags,&path);+if(ret)+gotoout;++ret=vfs_fsinfo(&path,ctx);+path_put(&path);+if(retry_estale(ret,lookup_flags)){+lookup_flags|=LOOKUP_REVAL;+gotoretry;+}+out:+returnret;+}++staticintvfs_fsinfo_fd(unsignedintfd,structfsinfo_context*ctx)+{+structfdf=fdget_raw(fd);+intret=-EBADF;++if(f.file){+ret=vfs_fsinfo(&f.file->f_path,ctx);+fdput(f);+}+returnret;+}++/**+*sys_fsinfo-Systemcalltogetfilesysteminformation+*@dfd:Basedirectorytopathwalkfromorfdreferringtofilesystem.+*@pathname:FilesystemtoqueryorNULL.+*@_params:Parameterstodefinerequest(orNULLforenhancedstatfs).+*@user_buffer:Resultbuffer.+*@user_buf_size:Sizeofresultbuffer.+*+*Getinformationonafilesystem.Thefilesystemattributetobequeriedis+*indicatedby@_params->request,andsomeoftheattributescanhavemultiple+*values,indexedby@_params->Nthand@_params->Mth.If@_paramsisNULL,+*thenthe0thfsinfo_attr_statfsattributeisqueried.Ifanattributedoes+*notexist,EOPNOTSUPPisreturned;iftheNth,Mthvaluedoesnotexist,+*ENODATAisreturned.+*+*Onsuccess,thesizeoftheattribute'svalueisreturned.If+*@user_buf_sizeis0or@user_bufferisNULL,onlythesizeisreturned.If+*thesizeofthevalueislargerthan@user_buf_size,itwillbetruncatedby+*thecopy.Ifthesizeofthevalueissmallerthan@user_buf_sizethenthe+*excessbufferspacewillbecleared.Thefullsizeofthevaluewillbe+*returned,irrespectiveofhowmuchdataisactuallyplacedinthebuffer.+*/+SYSCALL_DEFINE5(fsinfo,+int,dfd,constchar__user*,pathname,+structfsinfo_params__user*,params,+void__user*,user_buffer,size_t,user_buf_size)+{+structfsinfo_contextctx;+structfsinfo_paramsuser_params;+unsignedintat_flags=0,result_size;+intret;++if(!user_buffer&&user_buf_size)+return-EINVAL;+if(user_buffer&&!user_buf_size)+return-EINVAL;+if(user_buf_size>UINT_MAX)+return-EOVERFLOW;++memset(&ctx,0,sizeof(ctx));+ctx.requested_attr=FSINFO_ATTR_STATFS;+if(user_buf_size==0)+ctx.want_size_only=true;++if(params){+if(copy_from_user(&user_params,params,sizeof(user_params)))+return-EFAULT;+if(user_params.__reserved32[0]||+user_params.__reserved[0]||+user_params.__reserved[1]||+user_params.__reserved[2]||+user_params.flags&~FSINFO_FLAGS_QUERY_MASK)+return-EINVAL;+at_flags=user_params.at_flags;+ctx.flags=user_params.flags;+ctx.requested_attr=user_params.request;+ctx.Nth=user_params.Nth;+ctx.Mth=user_params.Mth;+}++switch(ctx.flags&FSINFO_FLAGS_QUERY_MASK){+caseFSINFO_FLAGS_QUERY_PATH:+ret=vfs_fsinfo_path(dfd,pathname,at_flags,&ctx);+break;+caseFSINFO_FLAGS_QUERY_FD:+if(pathname)+return-EINVAL;+ret=vfs_fsinfo_fd(dfd,&ctx);+break;+default:+return-EINVAL;+}++if(ret<0)+gotoerror;++result_size=min_t(size_t,ret,user_buf_size);+if(result_size>0&&+copy_to_user(user_buffer,ctx.buffer,result_size)!=0){+ret=-EFAULT;+gotoerror;+}++/* Clear any part of the buffer that we won't fill if we're putting a+*structinthere.Strings,opaqueobjectsandarraysareexpectedto+*bevariablelength.+*/+if(ctx.clear_tail&&+user_buf_size>result_size&&+clear_user(user_buffer+result_size,user_buf_size-result_size)!=0){+ret=-EFAULT;+gotoerror;+}++error:+kvfree(ctx.buffer);+returnret;+}
@@ -0,0 +1,187 @@+/* SPDX-License-Identifier: GPL-2.0 WITH Linux-syscall-note */+/* fsinfo() definitions.+*+*Copyright(C)2020RedHat,Inc.AllRightsReserved.+*WrittenbyDavidHowells(dhowells@redhat.com)+*/+#ifndef _UAPI_LINUX_FSINFO_H+#define _UAPI_LINUX_FSINFO_H++#include<linux/types.h>+#include<linux/socket.h>++/*+*Thefilesystemattributesthatcanberequested.Notethatsomeattributes+*mayhavemultipleinstanceswhichcanbeswitchedintheparameterblock.+*/+#define FSINFO_ATTR_STATFS 0x00 /* statfs()-style state */+#define FSINFO_ATTR_IDS 0x01 /* Filesystem IDs */+#define FSINFO_ATTR_LIMITS 0x02 /* Filesystem limits */+#define FSINFO_ATTR_SUPPORTS 0x03 /* What's supported in statx, iocflags, ... */+#define FSINFO_ATTR_TIMESTAMP_INFO 0x04 /* Inode timestamp info */+#define FSINFO_ATTR_VOLUME_ID 0x05 /* Volume ID (string) */+#define FSINFO_ATTR_VOLUME_UUID 0x06 /* Volume UUID (LE uuid) */+#define FSINFO_ATTR_VOLUME_NAME 0x07 /* Volume name (string) */++#define FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO 0x100 /* Information about attr N (for path) */+#define FSINFO_ATTR_FSINFO_ATTRIBUTES 0x101 /* List of supported attrs (for path) */++/*+*Optionalfsinfo()parameterstructure.+*+*Ifthisisnotgiven,itisassumedthatfsinfo_attr_statfsinstance0,0is+*desired.+*/+structfsinfo_params{+__u32at_flags;/* AT_SYMLINK_NOFOLLOW and similar flags */+__u32flags;/* Flags controlling fsinfo() specifically */+#define FSINFO_FLAGS_QUERY_MASK 0x0007 /* What object should fsinfo() query? */+#define FSINFO_FLAGS_QUERY_PATH 0x0000 /* - path, specified by dirfd,pathname,AT_EMPTY_PATH */+#define FSINFO_FLAGS_QUERY_FD 0x0001 /* - fd specified by dirfd */+__u32request;/* ID of requested attribute */+__u32Nth;/* Instance of it (some may have multiple) */+__u32Mth;/* Subinstance of Nth instance */+__u32__reserved32[1];/* Reserved params; all must be 0 */+__u64__reserved[3];+};++enumfsinfo_value_type{+FSINFO_TYPE_VSTRUCT=0,/* Version-lengthed struct (up to 4096 bytes) */+FSINFO_TYPE_STRING=1,/* NUL-term var-length string (up to 4095 chars) */+FSINFO_TYPE_OPAQUE=2,/* Opaque blob (unlimited size) */+FSINFO_TYPE_LIST=3,/* List of ints/structs (unlimited size) */+};++/*+*Informationstructforfsinfo(FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO).+*+*Thisgivesinformationabouttheattributessupportedbyfsinfoforthe+*givenpath.+*/+structfsinfo_attribute_info{+unsignedintattr_id;/* The ID of the attribute */+enumfsinfo_value_typetype;/* The type of the attribute's value(s) */+unsignedintflags;+#define FSINFO_FLAGS_N 0x01 /* - Attr has a set of values */+#define FSINFO_FLAGS_NM 0x02 /* - Attr has a set of sets of values */+unsignedintsize;/* - Value size (FSINFO_STRUCT/FSINFO_LIST) */+};++#define FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO__STRUCT struct fsinfo_attribute_info+#define FSINFO_ATTR_FSINFO_ATTRIBUTES__STRUCT __u32++structfsinfo_u128{+#if defined(__BYTE_ORDER) ? __BYTE_ORDER == __BIG_ENDIAN : defined(__BIG_ENDIAN)+__u64hi;+__u64lo;+#elif defined(__BYTE_ORDER) ? __BYTE_ORDER == __LITTLE_ENDIAN : defined(__LITTLE_ENDIAN)+__u64lo;+__u64hi;+#endif+};++/*+*Informationstructforfsinfo(FSINFO_ATTR_STATFS).+*-Thisgivesextendedfilesysteminformation.+*/+structfsinfo_statfs{+structfsinfo_u128f_blocks;/* Total number of blocks in fs */+structfsinfo_u128f_bfree;/* Total number of free blocks */+structfsinfo_u128f_bavail;/* Number of free blocks available to ordinary user */+structfsinfo_u128f_files;/* Total number of file nodes in fs */+structfsinfo_u128f_ffree;/* Number of free file nodes */+structfsinfo_u128f_favail;/* Number of file nodes available to ordinary user */+__u64f_bsize;/* Optimal block size */+__u64f_frsize;/* Fragment size */+};++#define FSINFO_ATTR_STATFS__STRUCT struct fsinfo_statfs++/*+*Informationstructforfsinfo(FSINFO_ATTR_IDS).+*+*Listofbasicidentifiersasisnormallyfoundinstatfs().+*/+structfsinfo_ids{+charf_fs_name[15+1];/* Filesystem name */+__u64f_fsid;/* Short 64-bit Filesystem ID (as statfs) */+__u64f_sb_id;/* Internal superblock ID for sbnotify()/mntnotify() */+__u32f_fstype;/* Filesystem type from linux/magic.h [uncond] */+__u32f_dev_major;/* As st_dev_* from struct statx [uncond] */+__u32f_dev_minor;+__u32__padding[1];+};++#define FSINFO_ATTR_IDS__STRUCT struct fsinfo_ids++/*+*Informationstructforfsinfo(FSINFO_ATTR_LIMITS).+*+*Listofsupportedfilesystemlimits.+*/+structfsinfo_limits{+structfsinfo_u128max_file_size;/* Maximum file size */+structfsinfo_u128max_ino;/* Maximum inode number */+__u64max_uid;/* Maximum UID supported */+__u64max_gid;/* Maximum GID supported */+__u64max_projid;/* Maximum project ID supported */+__u64max_hard_links;/* Maximum number of hard links on a file */+__u64max_xattr_body_len;/* Maximum xattr content length */+__u32max_xattr_name_len;/* Maximum xattr name length */+__u32max_filename_len;/* Maximum filename length */+__u32max_symlink_len;/* Maximum symlink content length */+__u32max_dev_major;/* Maximum device major representable */+__u32max_dev_minor;/* Maximum device minor representable */+__u32__padding[1];+};++#define FSINFO_ATTR_LIMITS__STRUCT struct fsinfo_limits++/*+*Informationstructforfsinfo(FSINFO_ATTR_SUPPORTS).+*+*What'ssupportedinvariousmasks,suchasstatx()attributeandmaskbits+*andIOCflags.+*/+structfsinfo_supports{+__u64stx_attributes;/* What statx::stx_attributes are supported */+__u32stx_mask;/* What statx::stx_mask bits are supported */+__u32fs_ioc_getflags;/* What FS_IOC_GETFLAGS may return */+__u32fs_ioc_setflags_set;/* What FS_IOC_SETFLAGS may set */+__u32fs_ioc_setflags_clear;/* What FS_IOC_SETFLAGS may clear */+__u32win_file_attrs;/* What DOS/Windows FILE_* attributes are supported */+__u32__padding[1];+};++#define FSINFO_ATTR_SUPPORTS__STRUCT struct fsinfo_supports++structfsinfo_timestamp_one{+__s64minimum;/* Minimum timestamp value in seconds */+__s64maximum;/* Maximum timestamp value in seconds */+__u16gran_mantissa;/* Granularity(secs) = mant * 10^exp */+__s8gran_exponent;+__u8__padding[5];+};++/*+*Informationstructforfsinfo(FSINFO_ATTR_TIMESTAMP_INFO).+*/+structfsinfo_timestamp_info{+structfsinfo_timestamp_oneatime;/* Access time */+structfsinfo_timestamp_onemtime;/* Modification time */+structfsinfo_timestamp_onectime;/* Change time */+structfsinfo_timestamp_onebtime;/* Birth/creation time */+};++#define FSINFO_ATTR_TIMESTAMP_INFO__STRUCT struct fsinfo_timestamp_info++/*+*Informationstructforfsinfo(FSINFO_ATTR_VOLUME_UUID).+*/+structfsinfo_volume_uuid{+__u8uuid[16];+};++#define FSINFO_ATTR_VOLUME_UUID__STRUCT struct fsinfo_volume_uuid++#endif /* _UAPI_LINUX_FSINFO_H */
@@ -1,10 +1,15 @@# SPDX-License-Identifier: GPL-2.0-only# List of programs to build+hostprogs:=\+test-fsinfo\test-fsmount\test-statxalways-y:=$(hostprogs)+HOSTCFLAGS_test-fsinfo.o+=-I$(objtree)/usr/include+HOSTLDLIBS_test-fsinfo+=-static-lm+HOSTCFLAGS_test-fsmount.o+=-I$(objtree)/usr/includeHOSTCFLAGS_test-statx.o+=-I$(objtree)/usr/include
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:03:04
Provide a bitmap of features that a filesystem may provide for the path
being queried. Features include such things as:
(1) The general class of filesystem, such as kernel-interface,
block-based, flash-based, network-based.
(2) Supported inode features, such as which timestamps are supported,
whether simple numeric user, group or project IDs are supported and
whether user identification is actually more complex behind the
scenes.
(3) Supported volume features, such as it having a UUID, a name or a
filesystem ID.
(4) Supported filesystem features, such as what types of file are
supported, whether sparse files, extended attributes and quotas are
supported.
(5) Supported interface features, such as whether locking and leases are
supported, what open flags are honoured and how i_version is managed.
For some filesystems, this may be an immutable set and can just be memcpy'd
into the reply buffer.
Signed-off-by: David Howells <dhowells@redhat.com>
---
fs/fsinfo.c | 30 +++++++++++++++++++
include/linux/fsinfo.h | 38 ++++++++++++++++++++++++
include/uapi/linux/fsinfo.h | 67 ++++++++++++++++++++++++++++++++++++++++++
samples/vfs/test-fsinfo.c | 69 +++++++++++++++++++++++++++++++++++++++++++
4 files changed, 204 insertions(+)
@@ -22,6 +22,7 @@#define FSINFO_ATTR_VOLUME_ID 0x05 /* Volume ID (string) */#define FSINFO_ATTR_VOLUME_UUID 0x06 /* Volume UUID (LE uuid) */#define FSINFO_ATTR_VOLUME_NAME 0x07 /* Volume name (string) */+#define FSINFO_ATTR_FEATURES 0x08 /* Filesystem features (bits) */#define FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO 0x100 /* Information about attr N (for path) */#define FSINFO_ATTR_FSINFO_ATTRIBUTES 0x101 /* List of supported attrs (for path) */
@@ -155,6 +156,72 @@ struct fsinfo_supports {#define FSINFO_ATTR_SUPPORTS__STRUCT struct fsinfo_supports+/*+*Informationstructforfsinfo(FSINFO_ATTR_FEATURES).+*+*Bitmaskindicatingfilesystemfeatureswhererenderableassinglebits.+*/+enumfsinfo_feature{+FSINFO_FEAT_IS_KERNEL_FS=0,/* fs is kernel-special filesystem */+FSINFO_FEAT_IS_BLOCK_FS=1,/* fs is block-based filesystem */+FSINFO_FEAT_IS_FLASH_FS=2,/* fs is flash filesystem */+FSINFO_FEAT_IS_NETWORK_FS=3,/* fs is network filesystem */+FSINFO_FEAT_IS_AUTOMOUNTER_FS=4,/* fs is automounter special filesystem */+FSINFO_FEAT_IS_MEMORY_FS=5,/* fs is memory-based filesystem */+FSINFO_FEAT_AUTOMOUNTS=6,/* fs supports automounts */+FSINFO_FEAT_ADV_LOCKS=7,/* fs supports advisory file locking */+FSINFO_FEAT_MAND_LOCKS=8,/* fs supports mandatory file locking */+FSINFO_FEAT_LEASES=9,/* fs supports file leases */+FSINFO_FEAT_UIDS=10,/* fs supports numeric uids */+FSINFO_FEAT_GIDS=11,/* fs supports numeric gids */+FSINFO_FEAT_PROJIDS=12,/* fs supports numeric project ids */+FSINFO_FEAT_STRING_USER_IDS=13,/* fs supports string user identifiers */+FSINFO_FEAT_GUID_USER_IDS=14,/* fs supports GUID user identifiers */+FSINFO_FEAT_WINDOWS_ATTRS=15,/* fs has windows attributes */+FSINFO_FEAT_USER_QUOTAS=16,/* fs has per-user quotas */+FSINFO_FEAT_GROUP_QUOTAS=17,/* fs has per-group quotas */+FSINFO_FEAT_PROJECT_QUOTAS=18,/* fs has per-project quotas */+FSINFO_FEAT_XATTRS=19,/* fs has xattrs */+FSINFO_FEAT_JOURNAL=20,/* fs has a journal */+FSINFO_FEAT_DATA_IS_JOURNALLED=21,/* fs is using data journalling */+FSINFO_FEAT_O_SYNC=22,/* fs supports O_SYNC */+FSINFO_FEAT_O_DIRECT=23,/* fs supports O_DIRECT */+FSINFO_FEAT_VOLUME_ID=24,/* fs has a volume ID */+FSINFO_FEAT_VOLUME_UUID=25,/* fs has a volume UUID */+FSINFO_FEAT_VOLUME_NAME=26,/* fs has a volume name */+FSINFO_FEAT_VOLUME_FSID=27,/* fs has a volume FSID */+FSINFO_FEAT_IVER_ALL_CHANGE=28,/* i_version represents data + meta changes */+FSINFO_FEAT_IVER_DATA_CHANGE=29,/* i_version represents data changes only */+FSINFO_FEAT_IVER_MONO_INCR=30,/* i_version incremented monotonically */+FSINFO_FEAT_DIRECTORIES=31,/* fs supports (sub)directories */+FSINFO_FEAT_SYMLINKS=32,/* fs supports symlinks */+FSINFO_FEAT_HARD_LINKS=33,/* fs supports hard links */+FSINFO_FEAT_HARD_LINKS_1DIR=34,/* fs supports hard links in same dir only */+FSINFO_FEAT_DEVICE_FILES=35,/* fs supports bdev, cdev */+FSINFO_FEAT_UNIX_SPECIALS=36,/* fs supports pipe, fifo, socket */+FSINFO_FEAT_RESOURCE_FORKS=37,/* fs supports resource forks/streams */+FSINFO_FEAT_NAME_CASE_INDEP=38,/* Filename case independence is mandatory */+FSINFO_FEAT_NAME_NON_UTF8=39,/* fs has non-utf8 names */+FSINFO_FEAT_NAME_HAS_CODEPAGE=40,/* fs has a filename codepage */+FSINFO_FEAT_SPARSE=41,/* fs supports sparse files */+FSINFO_FEAT_NOT_PERSISTENT=42,/* fs is not persistent */+FSINFO_FEAT_NO_UNIX_MODE=43,/* fs does not support unix mode bits */+FSINFO_FEAT_HAS_ATIME=44,/* fs supports access time */+FSINFO_FEAT_HAS_BTIME=45,/* fs supports birth/creation time */+FSINFO_FEAT_HAS_CTIME=46,/* fs supports change time */+FSINFO_FEAT_HAS_MTIME=47,/* fs supports modification time */+FSINFO_FEAT_HAS_ACL=48,/* fs supports ACLs of some sort */+FSINFO_FEAT_HAS_INODE_NUMBERS=49,/* fs has inode numbers */+FSINFO_FEAT__NR+};++structfsinfo_features{+__u32nr_features;/* Number of supported features (FSINFO_FEAT__NR) */+__u8features[(FSINFO_FEAT__NR+7)/8];+};++#define FSINFO_ATTR_FEATURES__STRUCT struct fsinfo_features+structfsinfo_timestamp_one{__s64minimum;/* Minimum timestamp value in seconds */__s64maximum;/* Maximum timestamp value in seconds */
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:03:09
Allow the fsinfo() syscall to look up a mount object by ID rather than by
pathname. This is necessary as there can be multiple mounts stacked up at
the same pathname and there's no way to look through them otherwise.
This is done by passing FSINFO_FLAGS_QUERY_MOUNT to fsinfo() in the
parameters and then passing the mount ID as a string to fsinfo() in place
of the filename:
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_MOUNT,
.request = FSINFO_ATTR_IDS,
};
ret = fsinfo(AT_FDCWD, "21", ¶ms, buffer, sizeof(buffer));
The caller is only permitted to query a mount object if the root directory
of that mount connects directly to the current chroot if dfd == AT_FDCWD[*]
or the directory specified by dfd otherwise. Note that this is not
available to the pathwalk of any other syscall.
[*] This needs to be something other than AT_FDCWD, perhaps AT_FDROOT.
[!] This probably needs an LSM hook.
[!] This might want to check the permissions on all the intervening dirs -
but it would have to do that under RCU conditions.
[!] This might want to check a CAP_* flag.
Signed-off-by: David Howells <dhowells@redhat.com>
---
fs/fsinfo.c | 53 +++++++++++++++++++
fs/internal.h | 2 +
fs/namespace.c | 117 ++++++++++++++++++++++++++++++++++++++++++-
include/uapi/linux/fsinfo.h | 1
samples/vfs/test-fsinfo.c | 11 +++-
5 files changed, 179 insertions(+), 5 deletions(-)
@@ -63,7 +63,7 @@ static int __init set_mphash_entries(char *str)__setup("mphash_entries=",set_mphash_entries);staticu64event;-staticDEFINE_IDA(mnt_id_ida);+staticDEFINE_IDR(mnt_id_ida);staticDEFINE_IDA(mnt_group_ida);staticstructhlist_head*mount_hashtable__read_mostly;
@@ -104,17 +104,27 @@ static inline struct hlist_head *mp_hash(struct dentry *dentry)staticintmnt_alloc_id(structmount*mnt){-intres=ida_alloc(&mnt_id_ida,GFP_KERNEL);+intres;+/* Allocate an ID, but don't set the pointer back to the mount until+*later,asoncewedothat,wehavetofollowRCUprotocolstoget+*ridofthemountstruct.+*/+res=idr_alloc(&mnt_id_ida,NULL,0,INT_MAX,GFP_KERNEL);if(res<0)returnres;mnt->mnt_id=res;return0;}+staticvoidmnt_publish_id(structmount*mnt)+{+idr_replace(&mnt_id_ida,mnt,mnt->mnt_id);+}+staticvoidmnt_free_id(structmount*mnt){-ida_free(&mnt_id_ida,mnt->mnt_id);+idr_remove(&mnt_id_ida,mnt->mnt_id);}/*
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:03:18
Allow mount information, including information about the topology tree to
be queried with the fsinfo() system call. Setting AT_FSINFO_QUERY_MOUNT
allows overlapping mounts to be queried by indicating that the syscall
should interpet the pathname as a number indicating the mount ID.
To this end, four fsinfo() attributes are provided:
(1) FSINFO_ATTR_MOUNT_INFO.
This is a structure providing information about a mount, including:
- Mounted superblock ID.
- Mount ID (can be used with AT_FSINFO_QUERY_MOUNT).
- Parent mount ID.
- Mount attributes (eg. R/O, NOEXEC).
- A change counter.
Note that the parent mount ID is overridden to the ID of the queried
mount if the parent lies outside of the chroot or dfd tree.
(2) FSINFO_ATTR_MOUNT_DEVNAME.
This a string providing the device name associated with the mount.
Note that the device name may be a path that lies outside of the root.
(3) FSINFO_ATTR_MOUNT_POINT.
This is a string indicating the name of the mountpoint within the
parent mount, limited to the parent's mounted root and the chroot.
(4) FSINFO_ATTR_MOUNT_CHILDREN.
This produces an array of structures, one for each child and capped
with one for the argument mount (checked after listing all the
children). Each element contains the mount ID and the change counter
of the respective mount object.
Signed-off-by: David Howells <dhowells@redhat.com>
---
fs/d_path.c | 2
fs/fsinfo.c | 5 +
fs/internal.h | 10 ++
fs/namespace.c | 179 +++++++++++++++++++++++++++++++++++++++++++
include/uapi/linux/fsinfo.h | 34 ++++++++
samples/vfs/test-fsinfo.c | 27 ++++++
6 files changed, 256 insertions(+), 1 deletion(-)
@@ -229,7 +229,7 @@ static int prepend_unreachable(char **buffer, int *buflen)returnprepend(buffer,buflen,"(unreachable)",13);}-staticvoidget_fs_root_rcu(structfs_struct*fs,structpath*root)+voidget_fs_root_rcu(structfs_struct*fs,structpath*root){unsignedseq;
@@ -4108,3 +4109,181 @@ int lookup_mount_object(struct path *root, int mnt_id, struct path *_mntpt)unlock_mount_hash();gotoout_unlock;}++#ifdef CONFIG_FSINFO+/*+*Retrieveinformationaboutthenominatedmount.+*/+intfsinfo_generic_mount_info(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_mount_info*p=ctx->buffer;+structsuper_block*sb;+structmount*m;+structpathroot;+unsignedintflags;++if(!path->mnt)+return-ENODATA;++m=real_mount(path->mnt);+sb=m->mnt.mnt_sb;++p->f_sb_id=sb->s_unique_id;+p->mnt_id=m->mnt_id;+p->parent_id=m->mnt_parent->mnt_id;+p->change_counter=atomic_read(&m->mnt_change_counter);++get_fs_root(current->fs,&root);+if(path->mnt==root.mnt){+p->parent_id=p->mnt_id;+}else{+rcu_read_lock();+if(!are_paths_connected(&root,path))+p->parent_id=p->mnt_id;+rcu_read_unlock();+}+if(IS_MNT_SHARED(m))+p->group_id=m->mnt_group_id;+if(IS_MNT_SLAVE(m)){+intmaster=m->mnt_master->mnt_group_id;+intdom=get_dominating_id(m,&root);+p->master_id=master;+if(dom&&dom!=master)+p->from_id=dom;+}+path_put(&root);++flags=READ_ONCE(m->mnt.mnt_flags);+if(flags&MNT_READONLY)+p->attr|=MOUNT_ATTR_RDONLY;+if(flags&MNT_NOSUID)+p->attr|=MOUNT_ATTR_NOSUID;+if(flags&MNT_NODEV)+p->attr|=MOUNT_ATTR_NODEV;+if(flags&MNT_NOEXEC)+p->attr|=MOUNT_ATTR_NOEXEC;+if(flags&MNT_NODIRATIME)+p->attr|=MOUNT_ATTR_NODIRATIME;++if(flags&MNT_NOATIME)+p->attr|=MOUNT_ATTR_NOATIME;+elseif(flags&MNT_RELATIME)+p->attr|=MOUNT_ATTR_RELATIME;+else+p->attr|=MOUNT_ATTR_STRICTATIME;+returnsizeof(*p);+}++intfsinfo_generic_mount_devname(structpath*path,structfsinfo_context*ctx)+{+if(!path->mnt)+return-ENODATA;++returnfsinfo_string(real_mount(path->mnt)->mnt_devname,ctx);+}++/*+*Returnthepathofthismountrelativetoitsparentandclippedto+*thecurrentchroot.+*/+intfsinfo_generic_mount_point(structpath*path,structfsinfo_context*ctx)+{+structmountpoint*mp;+structmount*m,*parent;+structpathmountpoint,root;+size_tlen;+void*p;++if(!path->mnt)+return-ENODATA;++rcu_read_lock();++m=real_mount(path->mnt);+parent=m->mnt_parent;+if(parent==m)+gotoskip;+mp=READ_ONCE(m->mnt_mp);+if(mp)+gotofound;+skip:+rcu_read_unlock();+return-ENODATA;++found:+mountpoint.mnt=&parent->mnt;+mountpoint.dentry=READ_ONCE(mp->m_dentry);++get_fs_root_rcu(current->fs,&root);+if(path->mnt==root.mnt){+rcu_read_unlock();+len=snprintf(ctx->buffer,ctx->buf_size,"/");+}else{+if(root.mnt!=&parent->mnt){+root.mnt=&parent->mnt;+root.dentry=parent->mnt.mnt_root;+}++p=__d_path(&mountpoint,&root,ctx->buffer,ctx->buf_size);+rcu_read_unlock();++if(IS_ERR(p))+returnPTR_ERR(p);+if(!p)+return-EPERM;++len=(ctx->buffer+ctx->buf_size)-p;+memmove(ctx->buffer,p,len);+}+returnlen;+}++/*+*Storeamountrecordintothefsinfobuffer.+*/+staticvoidstore_mount_fsinfo(structfsinfo_context*ctx,+structfsinfo_mount_child*child)+{+unsignedintusage=ctx->usage;+unsignedinttotal=sizeof(*child);++if(ctx->usage>=INT_MAX)+return;+ctx->usage=usage+total;+if(ctx->buffer&&ctx->usage<=ctx->buf_size)+memcpy(ctx->buffer+usage,child,total);+}++/*+*Returninformationaboutthesubmountsrelativetopath.+*/+intfsinfo_generic_mount_children(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_mount_childrecord;+structmount*m,*child;++if(!path->mnt)+return-ENODATA;++m=real_mount(path->mnt);++rcu_read_lock();+list_for_each_entry_rcu(child,&m->mnt_mounts,mnt_child){+if(child->mnt_parent!=m)+continue;+record.mnt_id=child->mnt_id;+record.change_counter=atomic_read(&child->mnt_change_counter);+store_mount_fsinfo(ctx,&record);+}+rcu_read_unlock();++/* End the list with a copy of the parameter mount's details so that+*userspacecanquicklycheckforchanges.+*/+record.mnt_id=m->mnt_id;+record.change_counter=atomic_read(&m->mnt_change_counter);+store_mount_fsinfo(ctx,&record);+returnctx->usage;+}++#endif /* CONFIG_FSINFO */
@@ -27,6 +27,11 @@#define FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO 0x100 /* Information about attr N (for path) */#define FSINFO_ATTR_FSINFO_ATTRIBUTES 0x101 /* List of supported attrs (for path) */+#define FSINFO_ATTR_MOUNT_INFO 0x200 /* Mount object information */+#define FSINFO_ATTR_MOUNT_DEVNAME 0x201 /* Mount object device name (string) */+#define FSINFO_ATTR_MOUNT_POINT 0x202 /* Relative path of mount in parent (string) */+#define FSINFO_ATTR_MOUNT_CHILDREN 0x203 /* Children of this mount (list) */+/**Optionalfsinfo()parameterstructure.*
@@ -82,6 +88,34 @@ struct fsinfo_u128 {#endif};+/*+*Informationstructforfsinfo(FSINFO_ATTR_MOUNT_INFO).+*/+structfsinfo_mount_info{+__u64f_sb_id;/* Superblock ID */+__u32mnt_id;/* Mount identifier (use with AT_FSINFO_MOUNTID_PATH) */+__u32parent_id;/* Parent mount identifier */+__u32group_id;/* Mount group ID */+__u32master_id;/* Slave master group ID */+__u32from_id;/* Slave propagated from ID */+__u32attr;/* MOUNT_ATTR_* flags */+__u32change_counter;/* Number of changes applied. */+__u32__reserved[1];+};++#define FSINFO_ATTR_MOUNT_INFO__STRUCT struct fsinfo_mount_info++/*+*Informationstructelementforfsinfo(FSINFO_ATTR_MOUNT_CHILDREN).+*-Anextraelementisplacedontheendrepresentingtheparentmount.+*/+structfsinfo_mount_child{+__u32mnt_id;/* Mount identifier (use with AT_FSINFO_MOUNTID_PATH) */+__u32change_counter;/* Number of changes applied to mount. */+};++#define FSINFO_ATTR_MOUNT_CHILDREN__STRUCT struct fsinfo_mount_child+/**Informationstructforfsinfo(FSINFO_ATTR_STATFS).*-Thisgivesextendedfilesysteminformation.
@@ -117,4 +117,12 @@ enum fsconfig_command {#define MOUNT_ATTR_STRICTATIME 0x00000020 /* - Always perform atime updates */#define MOUNT_ATTR_NODIRATIME 0x00000080 /* Do not update directory access times */+/*+*Mountobjectpropogationattributes.+*/+#define MOUNT_PROPAGATION_UNBINDABLE 0x00000001 /* Mount is unbindable */+#define MOUNT_PROPAGATION_SLAVE 0x00000002 /* Mount is slave */+#define MOUNT_PROPAGATION_PRIVATE 0x00000000 /* Mount is private (ie. not shared) */+#define MOUNT_PROPAGATION_SHARED 0x00000004 /* Mount is shared */+#endif /* _UAPI_LINUX_MOUNT_H */
@@ -236,8 +236,8 @@ int main(int argc, char **argv)exit(2);}-printf("MOUNT MOUNT ID CHANGE# AT DEV TYPE\n");-printf("------------------------------------- ---------- -------- -- ----- --------\n");+printf("MOUNT MOUNT ID CHANGE# AT P DEV TYPE\n");+printf("------------------------------------- ---------- -------- -- - ----- --------\n");display_mount(mnt_id,0,path);return0;}
From: David Howells <dhowells@redhat.com> Date: 2020-02-21 18:03:41
Provide an fsinfo attribute to query the superblock unique ID and
notification counter. The unique ID is placed in notification events and
the counted it provided so that the changed superblock can be determined in
the event of a notification buffer overrun. This is accessed with:
struct fsinfo_params params = {
.request = FSINFO_ATTR_SB_NOTIFICATIONS,
};
and returns a structure that looks like:
struct fsinfo_sb_notifications {
__u64 watch_id;
__u32 notify_counter;
__u32 __reserved[1];
};
Where watch_id is a number uniquely identifying the superblock in
notification records and notify_counter is incremented for each
superblock notification posted.
Signed-off-by: David Howells <dhowells@redhat.com>
---
fs/fsinfo.c | 11 +++++++++++
include/uapi/linux/fsinfo.h | 12 ++++++++++++
include/uapi/linux/watch_queue.h | 2 +-
samples/vfs/test-fsinfo.c | 10 ++++++++++
4 files changed, 34 insertions(+), 1 deletion(-)
@@ -0,0 +1,491 @@+============================+Filesystem Information Query+============================++The fsinfo() system call allows the querying of filesystem and filesystem+security information beyond what stat(), statx() and statfs() can obtain. It+does not require a file to be opened as does ioctl().++fsinfo() may be called with a path, with open file descriptor or a with a mount+object identifier.++The fsinfo() system call needs to be configured on by enabling:++ "File systems"/"Enable the fsinfo() system call" (CONFIG_FSINFO)++This document has the following sections:++..contents:: :local:+++Overview+========++The fsinfo() system call retrieves one of a number of attributes, the IDs of+which can be found in include/uapi/linux/fsinfo.h::++ FSINFO_ATTR_STATFS - statfs()-style state+ FSINFO_ATTR_IDS - Filesystem IDs+ FSINFO_ATTR_LIMITS - Filesystem limits+ ...+ FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO - Information about an attribute+ FSINFO_ATTR_FSINFO_ATTRIBUTES - List of available attributes+ ...+ FSINFO_ATTR_MOUNT_INFO - Information about the mount topology+ ...++Each attribute can have zero or more values, which can be of one of the+following types:++*``VStruct``. This is a structure with a version-dependent length. New+ versions of the kernel may append more fields, though they are not+ permitted to remove or replace old ones.++ Older applications, expecting an older version of the field, can ask for a+ shorter struct and will only get the fields they requested; newer+ applications running on an older kernel will get the extra fields they+ requested filled with zeros. Either way, the system call returns the size+ of the internal struct, regardless of how much data it returned.++ This allows for struct-type fields to be extended in future.++*``String``. This is a variable-length string of up to 4096 characters (no+ NUL character is included). The returned string will be truncated if the+ output buffer is too small. The total size of the string is returned,+ regardless of any truncation.++*``Opaque``. This is a variable-length blob of indeterminate structure. It+ may be up to INT_MAX bytes in size.++*``List``. This is a variable-length list of fixed-size structures. The+ element size may not vary over time, so the element format must be designed+ with care. The maximum length is INT_MAX bytes, though this depends on the+ kernel being able to allocate an internal buffer large enough.++Value type is an inherent propery of an attribute and all the values of an+attribute must be of that type. Each attribute can have a single value, a+sequence of values or a sequence-of-sequences of values.+++Filesystem API+==============++If the filesystem wishes to provide a list of queryable attributes, it should+set the table pointer in the superblock::++ const struct fsinfo_attribute *fsinfo_attributes;++terminating it with a blank entry. Each entry is a ``struct fsinfo_attribute``+and these can be created with a set of helper macros::++ FSINFO_VSTRUCT(A,G)+ FSINFO_VSTRUCT_N(A,G)+ FSINFO_VSTRUCT_NM(A,G)+ FSINFO_STRING(A,G)+ FSINFO_STRING_N(A,G)+ FSINFO_STRING_NM(A,G)+ FSINFO_OPAQUE(A,G)+ FSINFO_LIST(A,G)+ FSINFO_LIST_N(A,G)++The names of the macro are a combination of type (vstruct, string, opaque and+list) and an optional qualifier, if the attribute has N values or N lots of M+values. ``A`` is the name of the attribute and ``G`` is a function to get a+value for that attribute.++For vstruct- and list-type attributes, it is expected that there is a macro+defined with the name ``A##__STRUCT`` that indicates the structure or element+type.++The get function needs to match the following type::++ int (*get)(struct path *path, struct fsinfo_context *ctx);++where "path" indicates the object to be queried and ctx is a context describing+the parameters and the output buffer. The function should return the total+size of the data it would like to produce or an error.++The parameter struct looks like::++ struct fsinfo_context {+ __u32 requested_attr;+ __u32 Nth;+ __u32 Mth;+ bool want_size_only;+ unsigned int buf_size;+ unsigned int usage;+ void *buffer;+ ...+ };++The fields relevant to the filesystem are as follows:++*``requested_attr``++ Which attribute is being requested. EOPNOTSUPP should be returned if the+ attribute is not supported by the filesystem or the LSM.++*``Nth`` and ``Mth``++ Which value of an attribute is being requested.++ For a single-value attribute Nth and Mth will both be 0.++ For a "1D" attribute, Nth will indicate which value and Mth will always+ be 0. Take, for example, FSINFO_ATTR_SERVER_NAME - for a network+ filesystem, the superblock will be backed by a number of servers. This will+ return the name of the Nth server. ENODATA will be returned if Nth goes+ beyond the end of the array.++ For a "2D" attribute, Mth will indicate the index in the Nth set of values.+ Take, for example, an attribute for a network filesystems that returns+ server addresses - each server may have one or more addresses. This could+ return the Mth address of the Nth server. ENODATA should be returned if the+ Nth set doesn't exist or the Mth element of the Nth set doesn't exist.++*``want_size_only``++ Is set to true if the caller only wants the size of the value so that the+ get function doesn't have to make expensive calculations or calls to+ retrieve the value.++*``buf_size``++ This indicates the current size of the buffer. For the list type and the+ opaque type this will be increased if the current buffer won't hold the+ value and the filesystem will be called again.++*``usage``++ This indicates how much of the buffer has been used so far for an list or+ opaque type attribute. This is updated by the fsinfo_note_param*()+ functions.++*``buffer``++ This points to the output buffer. For struct- and string-type attributes it+ will always be big enough; for list- and opaque-type, it will be buf_size in+ size and will be resized if the returned size is larger than this.++To simplify filesystem code, there will always be at least a minimal buffer+available if the ->fsinfo() method gets called - and the filesystem should+always write what it can into the buffer. It's possible that the fsinfo()+system call will then throw the contents away and just return the length.+++Helper Functions+================++The API includes a number of helper functions:++*``void fsinfo_set_feature(struct fsinfo_features *ft,+ enum fsinfo_feature feature);``++ This function sets a feature flag.++*``void fsinfo_clear_feature(struct fsinfo_features *ft,+ enum fsinfo_feature feature);``++ This function clears a feature flag.++*``void fsinfo_set_unix_features(struct fsinfo_features *ft);``++ Set feature flags appropriate to the features of a standard UNIX filesystem,+ such as having numeric UIDS and GIDS; allowing the creation of directories,+ symbolic links, hard links, device files, FIFO and socket files; permitting+ sparse files; and having access, change and modification times.+++Attribute Summary+=================++To summarise the attributes that are defined::++ Symbolic name Type+ ===================================== ===============+ FSINFO_ATTR_STATFS vstruct+ FSINFO_ATTR_IDS vstruct+ FSINFO_ATTR_LIMITS vstruct+ FSINFO_ATTR_SUPPORTS vstruct+ FSINFO_ATTR_FEATURES vstruct+ FSINFO_ATTR_TIMESTAMP_INFO vstruct+ FSINFO_ATTR_VOLUME_ID string+ FSINFO_ATTR_VOLUME_UUID vstruct+ FSINFO_ATTR_VOLUME_NAME string+ FSINFO_ATTR_NAME_ENCODING string+ FSINFO_ATTR_NAME_CODEPAGE string+ FSINFO_ATTR_FSINFO vstruct+ FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO vstruct+ FSINFO_ATTR_FSINFO_ATTRIBUTES list+ FSINFO_ATTR_MOUNT_INFO vstruct+ FSINFO_ATTR_MOUNT_DEVNAME string+ FSINFO_ATTR_MOUNT_POINT string+ FSINFO_ATTR_MOUNT_CHILDREN list+ FSINFO_ATTR_AFS_CELL_NAME string+ FSINFO_ATTR_AFS_SERVER_NAME N × string+ FSINFO_ATTR_AFS_SERVER_ADDRESS N × struct+++Attribute Catalogue+===================++A number of the attributes convey information about a filesystem superblock:++*``FSINFO_ATTR_STATFS``++ This struct-type attribute gives most of the equivalent data to statfs(),+ but with all the fields as unconditional 64-bit or 128-bit integers. Note+ that static data like IDs that don't change are retrieved with+ FSINFO_ATTR_IDS instead.++ Further, superblock flags (such as MS_RDONLY) are not exposed by this+ attribute; rather the parameters must be listed and the attributes picked+ out from that.++*``FSINFO_ATTR_IDS``++ This struct-type attribute conveys various identifiers used by the target+ filesystem. This includes the filesystem name, the NFS filesystem ID, the+ superblock ID used in notifications, the filesystem magic type number and+ the primary device ID.++*``FSINFO_ATTR_LIMITS``++ This struct-type attribute conveys the limits on various aspects of a+ filesystem, such as maximum file, symlink and xattr sizes, maxiumm filename+ and xattr name length, maximum number of symlinks, maximum device major and+ minor numbers and maximum UID, GID and project ID numbers.++*``FSINFO_ATTR_SUPPORTS``++ This struct-type attribute conveys information about the support the+ filesystem has for various UAPI features of a filesystem. This includes+ information about which bits are supported in various masks employed by the+ statx system call, what FS_IOC_* flags are supported by ioctls and what+ DOS/Windows file attribute flags are supported.++*``FSINFO_ATTR_TIMESTAMP_INFO``++ This struct-type attribute conveys information about the resolution and+ range of the timestamps available in a filesystem. The resolutions are+ given as a mantissa and exponent (resolution = mantissa * 10^exponent+ seconds), where the exponent can be negative to indicate a sub-second+ resolution (-9 being nanoseconds, for example).++*``FSINFO_ATTR_VOLUME_ID``++ This is a string-type attribute that conveys the superblock identifier for+ the volume. By default it will be filled in from the contents of s_id from+ the superblock. For a block-based filesystem, for example, this might be+ the name of the primary block device.++*``FSINFO_ATTR_VOLUME_UUID``++ This is a struct-type attribute that conveys the UUID identifier for the+ volume. By default it will be filled in from the contents of s_uuid from+ the superblock. If this doesn't exist, it will be an entirely zeros.++*``FSINFO_ATTR_VOLUME_NAME``++ This is a string-type attribute that conveys the name of the volume. By+ default it will return EOPNOTSUPP. For a disk-based filesystem, it might+ convey the partition label; for a network-based filesystem, it might convey+ the name of the remote volume.++*``FSINFO_ATTR_FEATURES``++ This is a special attribute, being a set of single-bit feature flags,+ formatted as struct-type attribute. The meanings of the feature bits are+ listed below - see the "Feature Bit Catalogue" section. The feature bits+ are grouped numerically into bytes, such that features 0-7 are in byte 0,+ 8-15 are in byte 1, 16-23 in byte 2 and so on.++ Any feature bit that's not supported by the kernel will be set to false if+ asked for. The highest supported feature can be obtained from attribute+ "FSINFO_ATTR_FSINFO".+++Some attributes give information about fsinfo itself:++*``FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO``++ This struct-type attribute gives metadata about the attribute with the ID+ specified by the Nth parameter, including its type, default size and+ element size.++*``FSINFO_ATTR_FSINFO_ATTRIBUTES``++ This list-type attribute gives a list of the attribute IDs available at the+ point of reference. FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO can then be used to+ query each attribute.++*``FSINFO_ATTR_FSINFO``++ This struct-type attribute gives information about the fsinfo() system call+ itself, including the maximum number of feature bits supported.+++Then there are filesystem-specific attributes, e.g.:++*``FSINFO_ATTR_AFS_CELL_NAME``++ This is a string-type attribute that retrieves the AFS cell name of the+ target object.++*``FSINFO_ATTR_AFS_SERVER_NAME``++ This is a string-type attribute that conveys the name of the Nth server+ backing a network-filesystem superblock.++*``FSINFO_ATTR_AFS_SERVER_ADDRESSES``++ This is a list-type attribute that conveys the Mth address of the Nth+ server, as returned by FSINFO_ATTR_SERVER_NAME.+++Feature Bit Catalogue+=====================++The feature bits convey single true/false assertions about a specific instance+of a filesystem (ie. a specific superblock). They are accessed using the+"FSINFO_ATTR_FEATURE" attribute:++*``FSINFO_FEAT_IS_KERNEL_FS``+*``FSINFO_FEAT_IS_BLOCK_FS``+*``FSINFO_FEAT_IS_FLASH_FS``+*``FSINFO_FEAT_IS_NETWORK_FS``+*``FSINFO_FEAT_IS_AUTOMOUNTER_FS``+*``FSINFO_FEAT_IS_MEMORY_FS``++ These indicate what kind of filesystem the target is: kernel API (proc),+ block-based (ext4), flash/nvm-based (jffs2), remote over the network (NFS),+ local quasi-filesystem that acts as a tray of mountpoints (autofs), plain+ in-memory filesystem (shmem).++*``FSINFO_FEAT_AUTOMOUNTS``++ This indicate if a filesystem may have objects that are automount points.++*``FSINFO_FEAT_ADV_LOCKS``+*``FSINFO_FEAT_MAND_LOCKS``+*``FSINFO_FEAT_LEASES``++ These indicate if a filesystem supports advisory locks, mandatory locks or+ leases.++*``FSINFO_FEAT_UIDS``+*``FSINFO_FEAT_GIDS``+*``FSINFO_FEAT_PROJIDS``++ These indicate if a filesystem supports/stores/transports numeric user IDs,+ group IDs or project IDs. The "FSINFO_ATTR_LIMITS" attribute can be used+ to find out the upper limits on the IDs values.++*``FSINFO_FEAT_STRING_USER_IDS``++ This indicates if a filesystem supports/stores/transports string user+ identifiers.++*``FSINFO_FEAT_GUID_USER_IDS``++ This indicates if a filesystem supports/stores/transports Windows GUIDs as+ user identifiers (eg. ntfs).++*``FSINFO_FEAT_WINDOWS_ATTRS``++ This indicates if a filesystem supports Windows FILE_* attribute bits+ (eg. cifs, jfs). The "FSINFO_ATTR_SUPPORTS" attribute can be used to find+ out which windows file attributes are supported by the filesystem.++*``FSINFO_FEAT_USER_QUOTAS``+*``FSINFO_FEAT_GROUP_QUOTAS``+*``FSINFO_FEAT_PROJECT_QUOTAS``++ These indicate if a filesystem supports quotas for users, groups or+ projects.++*``FSINFO_FEAT_XATTRS``++ These indicate if a filesystem supports extended attributes. The+ "FSINFO_ATTR_LIMITS" attribute can be used to find out the upper limits on+ the supported name and body lengths.++*``FSINFO_FEAT_JOURNAL``+*``FSINFO_FEAT_DATA_IS_JOURNALLED``++ These indicate whether the filesystem has a journal and whether data+ changes are logged to it.++*``FSINFO_FEAT_O_SYNC``+*``FSINFO_FEAT_O_DIRECT``++ These indicate whether the filesystem supports the O_SYNC and O_DIRECT+ flags.++*``FSINFO_FEAT_VOLUME_ID``+*``FSINFO_FEAT_VOLUME_UUID``+*``FSINFO_FEAT_VOLUME_NAME``+*``FSINFO_FEAT_VOLUME_FSID``++ These indicate whether ID, UUID, name and FSID identifiers actually exist+ in the filesystem and thus might be considered persistent.++*``FSINFO_FEAT_IVER_ALL_CHANGE``+*``FSINFO_FEAT_IVER_DATA_CHANGE``+*``FSINFO_FEAT_IVER_MONO_INCR``++ These indicate whether i_version in the inode is supported and, if so, what+ mode it operates in. The first two indicate if it's changed for any data+ or metadata change, or whether it's only changed for any data changes; the+ last indicates whether or not it's monotonically increasing for each such+ change.++*``FSINFO_FEAT_HARD_LINKS``+*``FSINFO_FEAT_HARD_LINKS_1DIR``++ These indicate whether the filesystem can have hard links made in it, and+ whether they can be made between directory or only within the same+ directory.++*``FSINFO_FEAT_DIRECTORIES``+*``FSINFO_FEAT_SYMLINKS``+*``FSINFO_FEAT_DEVICE_FILES``+*``FSINFO_FEAT_UNIX_SPECIALS``++ These indicate whether directories; symbolic links; device files; or pipes+ and sockets can be made within the filesystem.++*``FSINFO_FEAT_RESOURCE_FORKS``++ This indicates if the filesystem supports resource forks.++*``FSINFO_FEAT_NAME_CASE_INDEP``+*``FSINFO_FEAT_NAME_NON_UTF8``+*``FSINFO_FEAT_NAME_HAS_CODEPAGE``++ These indicate if the filesystem supports case-independent file names,+ whether the filenames are non-utf8 (see the "FSINFO_ATTR_NAME_ENCODING"+ attribute) and whether a codepage is in use to transliterate them (see+ the "FSINFO_ATTR_NAME_CODEPAGE" attribute).++*``FSINFO_FEAT_SPARSE``++ This indicates if a filesystem supports sparse files.++*``FSINFO_FEAT_NOT_PERSISTENT``++ This indicates if a filesystem is not persistent.++*``FSINFO_FEAT_NO_UNIX_MODE``++ This indicates if a filesystem doesn't support UNIX mode bits (though they+ may be manufactured from other bits, such as Windows file attribute flags).++*``FSINFO_FEAT_HAS_ATIME``+*``FSINFO_FEAT_HAS_BTIME``+*``FSINFO_FEAT_HAS_CTIME``+*``FSINFO_FEAT_HAS_MTIME``++ These indicate which timestamps a filesystem supports (access, birth,+ change, modify). The range and resolutions can be queried with the+ "FSINFO_ATTR_TIMESTAMPS" attribute).
@@ -760,3 +769,208 @@ static int afs_statfs(struct dentry *dentry, struct kstatfs *buf)returnret;}++#ifdef CONFIG_FSINFO+staticconststructfsinfo_timestamp_infoafs_timestamp_info={+.atime={+.minimum=0,+.maximum=UINT_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.mtime={+.minimum=0,+.maximum=UINT_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.ctime={+.minimum=0,+.maximum=UINT_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.btime={+.minimum=0,+.maximum=UINT_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+};++staticintafs_fsinfo_get_timestamp(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_timestamp_info*tsinfo=ctx->buffer;+*tsinfo=afs_timestamp_info;+returnsizeof(*tsinfo);+}++staticintafs_fsinfo_get_limits(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_limits*lim=ctx->buffer;++lim->max_file_size.hi=0;+lim->max_file_size.lo=MAX_LFS_FILESIZE;+/* Inode numbers can be 96-bit on YFS, but that's hard to determine. */+lim->max_ino.hi=0;+lim->max_ino.lo=UINT_MAX;+lim->max_hard_links=UINT_MAX;+lim->max_uid=UINT_MAX;+lim->max_gid=UINT_MAX;+lim->max_filename_len=AFSNAMEMAX-1;+lim->max_symlink_len=AFSPATHMAX-1;+returnsizeof(*lim);+}++staticintafs_fsinfo_get_supports(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_supports*p=ctx->buffer;++p->stx_mask=(STATX_TYPE|STATX_MODE|+STATX_NLINK|+STATX_UID|STATX_GID|+STATX_MTIME|STATX_INO|+STATX_SIZE);+p->stx_attributes=STATX_ATTR_AUTOMOUNT;+returnsizeof(*p);+}++staticintafs_fsinfo_get_features(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_features*p=ctx->buffer;++fsinfo_set_feature(p,FSINFO_FEAT_IS_NETWORK_FS);+fsinfo_set_feature(p,FSINFO_FEAT_AUTOMOUNTS);+fsinfo_set_feature(p,FSINFO_FEAT_ADV_LOCKS);+fsinfo_set_feature(p,FSINFO_FEAT_UIDS);+fsinfo_set_feature(p,FSINFO_FEAT_GIDS);+fsinfo_set_feature(p,FSINFO_FEAT_VOLUME_ID);+fsinfo_set_feature(p,FSINFO_FEAT_VOLUME_NAME);+fsinfo_set_feature(p,FSINFO_FEAT_IVER_MONO_INCR);+fsinfo_set_feature(p,FSINFO_FEAT_SYMLINKS);+fsinfo_set_feature(p,FSINFO_FEAT_HARD_LINKS_1DIR);+fsinfo_set_feature(p,FSINFO_FEAT_HAS_MTIME);+fsinfo_set_feature(p,FSINFO_FEAT_HAS_INODE_NUMBERS);+returnsizeof(*p);+}++staticintafs_dyn_fsinfo_get_features(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_features*p=ctx->buffer;++fsinfo_set_feature(p,FSINFO_FEAT_IS_AUTOMOUNTER_FS);+fsinfo_set_feature(p,FSINFO_FEAT_AUTOMOUNTS);+returnsizeof(*p);+}++staticintafs_fsinfo_get_volume_name(structpath*path,structfsinfo_context*ctx)+{+structafs_super_info*as=AFS_FS_S(path->dentry->d_sb);+structafs_volume*volume=as->volume;++memcpy(ctx->buffer,volume->name,volume->name_len);+returnvolume->name_len;+}++staticintafs_fsinfo_get_cell_name(structpath*path,structfsinfo_context*ctx)+{+structafs_super_info*as=AFS_FS_S(path->dentry->d_sb);+structafs_cell*cell=as->cell;++memcpy(ctx->buffer,cell->name,cell->name_len);+returncell->name_len;+}++staticintafs_fsinfo_get_server_name(structpath*path,structfsinfo_context*ctx)+{+structafs_server_list*slist;+structafs_super_info*as=AFS_FS_S(path->dentry->d_sb);+structafs_volume*volume=as->volume;+structafs_server*server;+intret=-ENODATA;++read_lock(&volume->servers_lock);+slist=volume->servers;+if(slist){+if(ctx->Nth<slist->nr_servers){+server=slist->servers[ctx->Nth].server;+ret=sprintf(ctx->buffer,"%pU",&server->uuid);+}+}++read_unlock(&volume->servers_lock);+returnret;+}++staticintafs_fsinfo_get_server_address(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_afs_server_address*p=ctx->buffer;+structafs_server_list*slist;+structafs_super_info*as=AFS_FS_S(path->dentry->d_sb);+structafs_addr_list*alist;+structafs_volume*volume=as->volume;+structafs_server*server;+structafs_net*net=afs_d2net(path->dentry);+unsignedinti;+intret=-ENODATA;++read_lock(&volume->servers_lock);+slist=afs_get_serverlist(volume->servers);+read_unlock(&volume->servers_lock);++if(ctx->Nth>=slist->nr_servers)+gotoput_slist;+server=slist->servers[ctx->Nth].server;++read_lock(&server->fs_lock);+alist=afs_get_addrlist(rcu_dereference_protected(+server->addresses,+lockdep_is_held(&server->fs_lock)));+read_unlock(&server->fs_lock);+if(!alist)+gotoput_slist;++ret=alist->nr_addrs*sizeof(*p);+if(ret<=ctx->buf_size){+for(i=0;i<alist->nr_addrs;i++)+memcpy(&p[i].address,&alist->addrs[i],+sizeof(structsockaddr_rxrpc));+}++afs_put_addrlist(alist);+put_slist:+afs_put_serverlist(net,slist);+returnret;+}++staticconststructfsinfo_attributeafs_fsinfo_attributes[]={+FSINFO_VSTRUCT(FSINFO_ATTR_TIMESTAMP_INFO,afs_fsinfo_get_timestamp),+FSINFO_VSTRUCT(FSINFO_ATTR_LIMITS,afs_fsinfo_get_limits),+FSINFO_VSTRUCT(FSINFO_ATTR_SUPPORTS,afs_fsinfo_get_supports),+FSINFO_VSTRUCT(FSINFO_ATTR_FEATURES,afs_fsinfo_get_features),+FSINFO_STRING(FSINFO_ATTR_VOLUME_NAME,afs_fsinfo_get_volume_name),+FSINFO_STRING(FSINFO_ATTR_AFS_CELL_NAME,afs_fsinfo_get_cell_name),+FSINFO_STRING_N(FSINFO_ATTR_AFS_SERVER_NAME,afs_fsinfo_get_server_name),+FSINFO_LIST_N(FSINFO_ATTR_AFS_SERVER_ADDRESSES,afs_fsinfo_get_server_address),+{}+};++staticconststructfsinfo_attributeafs_dyn_fsinfo_attributes[]={+FSINFO_VSTRUCT(FSINFO_ATTR_TIMESTAMP_INFO,afs_fsinfo_get_timestamp),+FSINFO_VSTRUCT(FSINFO_ATTR_FEATURES,afs_dyn_fsinfo_get_features),+{}+};++staticintafs_fsinfo(structpath*path,structfsinfo_context*ctx)+{+structafs_super_info*as=AFS_FS_S(path->dentry->d_sb);+intret;++if(as->dyn_root)+ret=fsinfo_get_attribute(path,ctx,afs_dyn_fsinfo_attributes);+else+ret=fsinfo_get_attribute(path,ctx,afs_fsinfo_attributes);+returnret;+}++#endif /* CONFIG_FSINFO */
@@ -33,6 +33,10 @@#define FSINFO_ATTR_MOUNT_POINT 0x202 /* Relative path of mount in parent (string) */#define FSINFO_ATTR_MOUNT_CHILDREN 0x203 /* Children of this mount (list) */+#define FSINFO_ATTR_AFS_CELL_NAME 0x300 /* AFS cell name (string) */+#define FSINFO_ATTR_AFS_SERVER_NAME 0x301 /* Name of the Nth server (string) */+#define FSINFO_ATTR_AFS_SERVER_ADDRESSES 0x302 /* List of addresses of the Nth server */+/**Optionalfsinfo()parameterstructure.*
@@ -0,0 +1,45 @@+// SPDX-License-Identifier: GPL-2.0+/* Filesystem information for ext4+*+*Copyright(C)2020RedHat,Inc.AllRightsReserved.+*WrittenbyDavidHowells(dhowells@redhat.com)+*/++#include<linux/mount.h>+#include"ext4.h"++staticintext4_fsinfo_get_volume_name(structpath*path,structfsinfo_context*ctx)+{+conststructext4_sb_info*sbi=EXT4_SB(path->mnt->mnt_sb);+conststructext4_super_block*es=sbi->s_es;++memcpy(ctx->buffer,es->s_volume_name,sizeof(es->s_volume_name));+returnstrlen(ctx->buffer);+}++staticintext4_fsinfo_get_timestamps(structpath*path,structfsinfo_context*ctx)+{+conststructext4_sb_info*sbi=EXT4_SB(path->mnt->mnt_sb);+conststructext4_super_block*es=sbi->s_es;+structfsinfo_ext4_timestamps*ts=ctx->buffer;++#define Z(R,S) R = S | (((u64)S##_hi) << 32)+Z(ts->mkfs_time,es->s_mkfs_time);+Z(ts->mount_time,es->s_mtime);+Z(ts->write_time,es->s_wtime);+Z(ts->last_check_time,es->s_lastcheck);+Z(ts->first_error_time,es->s_first_error_time);+Z(ts->last_error_time,es->s_last_error_time);+returnsizeof(*ts);+}++staticconststructfsinfo_attributeext4_fsinfo_attributes[]={+FSINFO_STRING(FSINFO_ATTR_VOLUME_NAME,ext4_fsinfo_get_volume_name),+FSINFO_VSTRUCT(FSINFO_ATTR_EXT4_TIMESTAMPS,ext4_fsinfo_get_timestamps),+{}+};++intext4_fsinfo(structpath*path,structfsinfo_context*ctx)+{+returnfsinfo_get_attribute(path,ctx,ext4_fsinfo_attributes);+}
@@ -37,6 +37,8 @@#define FSINFO_ATTR_AFS_SERVER_NAME 0x301 /* Name of the Nth server (string) */#define FSINFO_ATTR_AFS_SERVER_ADDRESSES 0x302 /* List of addresses of the Nth server */+#define FSINFO_ATTR_EXT4_TIMESTAMPS 0x400 /* Ext4 superblock timestamps */+/**Optionalfsinfo()parameterstructure.*
@@ -0,0 +1,230 @@+// SPDX-License-Identifier: GPL-2.0+/* Filesystem information for NFS+*+*Copyright(C)2020RedHat,Inc.AllRightsReserved.+*WrittenbyDavidHowells(dhowells@redhat.com)+*/++#include<linux/nfs_fs.h>+#include<linux/windows.h>+#include"internal.h"++staticconststructfsinfo_timestamp_infonfs_timestamp_info={+.atime={+.minimum=0,+.maximum=UINT_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.mtime={+.minimum=0,+.maximum=UINT_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.ctime={+.minimum=0,+.maximum=UINT_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.btime={+.minimum=0,+.maximum=UINT_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+};++staticintnfs_fsinfo_get_timestamp_info(structpath*path,structfsinfo_context*ctx)+{+conststructnfs_server*server=NFS_SB(path->dentry->d_sb);+structfsinfo_timestamp_info*r=ctx->buffer;+unsignedlonglongnsec;+unsignedintrem,mant;+intexp=-9;++*r=nfs_timestamp_info;++nsec=server->time_delta.tv_nsec;+nsec+=server->time_delta.tv_sec*1000000000ULL;+if(nsec==0)+gotoout;++do{+mant=nsec;+rem=do_div(nsec,10);+if(rem)+break;+exp++;+}while(nsec);++r->atime.gran_mantissa=mant;+r->atime.gran_exponent=exp;+r->btime.gran_mantissa=mant;+r->btime.gran_exponent=exp;+r->ctime.gran_mantissa=mant;+r->ctime.gran_exponent=exp;+r->mtime.gran_mantissa=mant;+r->mtime.gran_exponent=exp;++out:+returnsizeof(*r);+}++staticintnfs_fsinfo_get_info(structpath*path,structfsinfo_context*ctx)+{+conststructnfs_server*server=NFS_SB(path->dentry->d_sb);+conststructnfs_client*clp=server->nfs_client;+structfsinfo_nfs_info*r=ctx->buffer;++r->version=clp->rpc_ops->version;+r->minor_version=clp->cl_minorversion;+r->transport_proto=clp->cl_proto;+returnsizeof(*r);+}++staticintnfs_fsinfo_get_server_name(structpath*path,structfsinfo_context*ctx)+{+conststructnfs_server*server=NFS_SB(path->dentry->d_sb);+conststructnfs_client*clp=server->nfs_client;++returnfsinfo_string(clp->cl_hostname,ctx);+}++staticintnfs_fsinfo_get_server_addresses(structpath*path,structfsinfo_context*ctx)+{+conststructnfs_server*server=NFS_SB(path->dentry->d_sb);+conststructnfs_client*clp=server->nfs_client;+structfsinfo_nfs_server_address*addr=ctx->buffer;+intret;++ret=1*sizeof(*addr);+if(ret<=ctx->buf_size)+memcpy(&addr[0].address,&clp->cl_addr,clp->cl_addrlen);+returnret;++}++staticintnfs_fsinfo_get_gssapi_name(structpath*path,structfsinfo_context*ctx)+{+conststructnfs_server*server=NFS_SB(path->dentry->d_sb);+conststructnfs_client*clp=server->nfs_client;++returnfsinfo_string(clp->cl_acceptor,ctx);+}++staticintnfs_fsinfo_get_limits(structpath*path,structfsinfo_context*ctx)+{+conststructnfs_server*server=NFS_SB(path->dentry->d_sb);+structfsinfo_limits*lim=ctx->buffer;++lim->max_file_size.hi=0;+lim->max_file_size.lo=server->maxfilesize;+lim->max_ino.hi=0;+lim->max_ino.lo=U64_MAX;+lim->max_hard_links=UINT_MAX;+lim->max_uid=UINT_MAX;+lim->max_gid=UINT_MAX;+lim->max_filename_len=NAME_MAX-1;+lim->max_symlink_len=PATH_MAX-1;+returnsizeof(*lim);+}++staticintnfs_fsinfo_get_supports(structpath*path,structfsinfo_context*ctx)+{+conststructnfs_server*server=NFS_SB(path->dentry->d_sb);+structfsinfo_supports*sup=ctx->buffer;++/* Don't set STATX_INO as i_ino is fabricated and may not be unique. */++if(!(server->caps&NFS_CAP_MODE))+sup->stx_mask|=STATX_TYPE|STATX_MODE;+if(server->caps&NFS_CAP_OWNER)+sup->stx_mask|=STATX_UID;+if(server->caps&NFS_CAP_OWNER_GROUP)+sup->stx_mask|=STATX_GID;+if(server->caps&NFS_CAP_ATIME)+sup->stx_mask|=STATX_ATIME;+if(server->caps&NFS_CAP_CTIME)+sup->stx_mask|=STATX_CTIME;+if(server->caps&NFS_CAP_MTIME)+sup->stx_mask|=STATX_MTIME;+if(server->attr_bitmask[0]&FATTR4_WORD0_SIZE)+sup->stx_mask|=STATX_SIZE;+if(server->attr_bitmask[1]&FATTR4_WORD1_NUMLINKS)+sup->stx_mask|=STATX_NLINK;++if(server->attr_bitmask[0]&FATTR4_WORD0_ARCHIVE)+sup->win_file_attrs|=ATTR_ARCHIVE;+if(server->attr_bitmask[0]&FATTR4_WORD0_HIDDEN)+sup->win_file_attrs|=ATTR_HIDDEN;+if(server->attr_bitmask[1]&FATTR4_WORD1_SYSTEM)+sup->win_file_attrs|=ATTR_SYSTEM;++sup->stx_attributes=STATX_ATTR_AUTOMOUNT;+returnsizeof(*sup);+}++staticintnfs_fsinfo_get_features(structpath*path,structfsinfo_context*ctx)+{+conststructnfs_server*server=NFS_SB(path->dentry->d_sb);+structfsinfo_features*ft=ctx->buffer;++fsinfo_set_feature(ft,FSINFO_FEAT_IS_NETWORK_FS);+fsinfo_set_feature(ft,FSINFO_FEAT_AUTOMOUNTS);+fsinfo_set_feature(ft,FSINFO_FEAT_O_SYNC);+fsinfo_set_feature(ft,FSINFO_FEAT_O_DIRECT);+fsinfo_set_feature(ft,FSINFO_FEAT_ADV_LOCKS);+fsinfo_set_feature(ft,FSINFO_FEAT_DEVICE_FILES);+fsinfo_set_feature(ft,FSINFO_FEAT_UNIX_SPECIALS);+if(server->nfs_client->rpc_ops->version==4){+fsinfo_set_feature(ft,FSINFO_FEAT_LEASES);+fsinfo_set_feature(ft,FSINFO_FEAT_IVER_ALL_CHANGE);+}++if(server->caps&NFS_CAP_OWNER)+fsinfo_set_feature(ft,FSINFO_FEAT_UIDS);+if(server->caps&NFS_CAP_OWNER_GROUP)+fsinfo_set_feature(ft,FSINFO_FEAT_GIDS);+if(!(server->caps&NFS_CAP_MODE))+fsinfo_set_feature(ft,FSINFO_FEAT_NO_UNIX_MODE);+if(server->caps&NFS_CAP_ACLS)+fsinfo_set_feature(ft,FSINFO_FEAT_HAS_ACL);+if(server->caps&NFS_CAP_SYMLINKS)+fsinfo_set_feature(ft,FSINFO_FEAT_SYMLINKS);+if(server->caps&NFS_CAP_HARDLINKS)+fsinfo_set_feature(ft,FSINFO_FEAT_HARD_LINKS);+if(server->caps&NFS_CAP_ATIME)+fsinfo_set_feature(ft,FSINFO_FEAT_HAS_ATIME);+if(server->caps&NFS_CAP_CTIME)+fsinfo_set_feature(ft,FSINFO_FEAT_HAS_CTIME);+if(server->caps&NFS_CAP_MTIME)+fsinfo_set_feature(ft,FSINFO_FEAT_HAS_MTIME);++if(server->attr_bitmask[0]&FATTR4_WORD0_CASE_INSENSITIVE)+fsinfo_set_feature(ft,FSINFO_FEAT_NAME_CASE_INDEP);+if((server->attr_bitmask[0]&FATTR4_WORD0_ARCHIVE)||+(server->attr_bitmask[0]&FATTR4_WORD0_HIDDEN)||+(server->attr_bitmask[1]&FATTR4_WORD1_SYSTEM))+fsinfo_set_feature(ft,FSINFO_FEAT_WINDOWS_ATTRS);++returnsizeof(*ft);+}++staticconststructfsinfo_attributenfs_fsinfo_attributes[]={+FSINFO_VSTRUCT(FSINFO_ATTR_TIMESTAMP_INFO,nfs_fsinfo_get_timestamp_info),+FSINFO_VSTRUCT(FSINFO_ATTR_LIMITS,nfs_fsinfo_get_limits),+FSINFO_VSTRUCT(FSINFO_ATTR_SUPPORTS,nfs_fsinfo_get_supports),+FSINFO_VSTRUCT(FSINFO_ATTR_FEATURES,nfs_fsinfo_get_features),+FSINFO_VSTRUCT(FSINFO_ATTR_NFS_INFO,nfs_fsinfo_get_info),+FSINFO_STRING(FSINFO_ATTR_NFS_SERVER_NAME,nfs_fsinfo_get_server_name),+FSINFO_LIST(FSINFO_ATTR_NFS_SERVER_ADDRESSES,nfs_fsinfo_get_server_addresses),+FSINFO_STRING(FSINFO_ATTR_NFS_GSSAPI_NAME,nfs_fsinfo_get_gssapi_name),+{}+};++intnfs_fsinfo(structpath*path,structfsinfo_context*ctx)+{+returnfsinfo_get_attribute(path,ctx,nfs_fsinfo_attributes);+}
@@ -39,6 +39,11 @@#define FSINFO_ATTR_EXT4_TIMESTAMPS 0x400 /* Ext4 superblock timestamps */+#define FSINFO_ATTR_NFS_INFO 0x500 /* Information about an NFS mount */+#define FSINFO_ATTR_NFS_SERVER_NAME 0x501 /* Name of the server (string) */+#define FSINFO_ATTR_NFS_SERVER_ADDRESSES 0x502 /* List of addresses of the server */+#define FSINFO_ATTR_NFS_GSSAPI_NAME 0x503 /* GSSAPI acceptor name */+/**Optionalfsinfo()parameterstructure.*
From: James Bottomley <James.Bottomley@HansenPartnership.com> Date: 2020-02-21 20:21:42
On Fri, 2020-02-21 at 18:01 +0000, David Howells wrote:
[...]
============================
FILESYSTEM INFORMATION QUERY
============================
The fsinfo() system call allows information about the filesystem at a
particular path point to be queried as a set of attributes, some of
which may have more than one value.
Attribute values are of four basic types:
(1) Version dependent-length structure (size defined by type).
(2) Variable-length string (up to 4096, including NUL).
(3) List of structures (up to INT_MAX size).
(4) Opaque blob (up to INT_MAX size).
Attributes can have multiple values either as a sequence of values or
a sequence-of-sequences of values and all the values of a particular
attribute must be of the same type.
Note that the values of an attribute *are* allowed to vary between
dentries within a single superblock, depending on the specific dentry
that you're looking at, but all the values of an attribute have to be
of the same type.
I've tried to make the interface as light as possible, so
integer/enum attribute selector rather than string and the core does
all the allocation and extensibility support work rather than leaving
that to the filesystems. That means that for the first two attribute
types, the filesystem will always see a sufficiently-sized buffer
allocated. Further, this removes the possibility of the filesystem
gaining access to the userspace buffer.
fsinfo() allows a variety of information to be retrieved about a
filesystem and the mount topology:
(1) General superblock attributes:
- Filesystem identifiers (UUID, volume label, device numbers,
...)
- The limits on a filesystem's capabilities
- Information on supported statx fields and attributes and IOC
flags.
- A variety single-bit flags indicating supported capabilities.
- Timestamp resolution and range.
- The amount of space/free space in a filesystem (as statfs()).
- Superblock notification counter.
(2) Filesystem-specific superblock attributes:
- Superblock-level timestamps.
- Cell name.
- Server names and addresses.
- Filesystem-specific information.
(3) VFS information:
- Mount topology information.
- Mount attributes.
- Mount notification counter.
(4) Information about what the fsinfo() syscall itself supports,
including
the type and struct/element size of attributes.
The system is extensible:
(1) New attributes can be added. There is no requirement that a
filesystem implement every attribute. Note that the core VFS
keeps a
table of types and sizes so it can handle future extensibility
rather
than delegating this to the filesystems.
(2) Version length-dependent structure attributes can be made larger
and
have additional information tacked on the end, provided it keeps
the
layout of the existing fields. If an older process asks for a
shorter
structure, it will only be given the bits it asks for. If a
newer
process asks for a longer structure on an older kernel, the
extra
space will be set to 0. In all cases, the size of the data
actually
available is returned.
In essence, the size of a structure is that structure's version:
a
smaller size is an earlier version and a later version includes
everything that the earlier version did.
(3) New single-bit capability flags can be added. This is a
structure-typed
attribute and, as such, (2) applies. Any bits you wanted but
the kernel
doesn't support are automatically set to 0.
fsinfo() may be called like the following, for example:
struct fsinfo_params params = {
.at_flags = AT_SYMLINK_NOFOLLOW,
.flags = FSINFO_FLAGS_QUERY_PATH,
.request = FSINFO_ATTR_AFS_SERVER_ADDRESSES,
.Nth = 2,
};
struct fsinfo_server_address address;
len = fsinfo(AT_FDCWD, "/afs/grand.central.org/doc", ¶ms,
&address, sizeof(address));
The above example would query an AFS filesystem to retrieve the
address
list for the 3rd server, and:
struct fsinfo_params params = {
.at_flags = AT_SYMLINK_NOFOLLOW,
.flags = FSINFO_FLAGS_QUERY_PATH,
.request = FSINFO_ATTR_AFS_CELL_NAME;
};
char cell_name[256];
len = fsinfo(AT_FDCWD, "/afs/grand.central.org/doc", ¶ms,
&cell_name, sizeof(cell_name));
would retrieve the name of an AFS cell as a string.
In future, I want to make fsinfo() capable of querying a context
created by
fsopen() or fspick(), e.g.:
fd = fsopen("ext4", 0);
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_PARAMETERS;
};
char buffer[65536];
fsinfo(fd, NULL, ¶ms, &buffer, sizeof(buffer));
even if that context doesn't currently have a superblock attached. I
would prefer this to contain length-prefixed strings so that there's
no need to insert escaping, especially as any character, including
'\', can be used as the separator in cifs and so that binary
parameters can be returned (though that is a lesser issue).
Could I make a suggestion about how this should be done in a way that
doesn't actually require the fsinfo syscall at all: it could just be
done with fsconfig. The idea is based on something I've wanted to do
for configfd but couldn't because otherwise it wouldn't substitute for
fsconfig, but Christian made me think it was actually essential to the
ability of the seccomp and other verifier tools in the critique of
configfd and I belive the same critique applies here.
Instead of making fsconfig functionally configure ... as in you pass
the attribute name, type and parameters down into the fs specific
handler and the handler does a string match and then verifies the
parameters and then acts on them, make it table configured, so what
each fstype does is register a table of attributes which can be got and
optionally set (with each attribute having a get and optional set
function). We'd have multiple tables per fstype, so the generic VFS
can register a table of attributes it understands for every fstype
(things like name, uuid and the like) and then each fs type would
register a table of fs specific attributes following the same pattern.
The system would examine the fs specific table before the generic one,
allowing overrides. fsconfig would have the ability to both get and
set attributes, permitting retrieval as well as setting (which is how I
get rid of the fsinfo syscall), we'd have a global parameter, which
would retrieve the entire table by name and type so the whole thing is
introspectable because the upper layer knows a-priori all the
attributes which can be set for a given fs type and what type they are
(so we can make more of the parsing generic). Any attribute which
doesn't have a set routine would be read only and all attributes would
have to have a get routine meaning everything is queryable.
I think I know how to code this up in a way that would be fully
transparent to the existing syscalls.
James
On Fri, Feb 21, 2020 at 9:21 PM James Bottomley
[off-list ref] wrote:
On Fri, 2020-02-21 at 18:01 +0000, David Howells wrote:
[...]
quoted
============================
FILESYSTEM INFORMATION QUERY
============================
The fsinfo() system call allows information about the filesystem at a
particular path point to be queried as a set of attributes, some of
which may have more than one value.
Attribute values are of four basic types:
(1) Version dependent-length structure (size defined by type).
(2) Variable-length string (up to 4096, including NUL).
(3) List of structures (up to INT_MAX size).
(4) Opaque blob (up to INT_MAX size).
Attributes can have multiple values either as a sequence of values or
a sequence-of-sequences of values and all the values of a particular
attribute must be of the same type.
Note that the values of an attribute *are* allowed to vary between
dentries within a single superblock, depending on the specific dentry
that you're looking at, but all the values of an attribute have to be
of the same type.
I've tried to make the interface as light as possible, so
integer/enum attribute selector rather than string and the core does
all the allocation and extensibility support work rather than leaving
that to the filesystems. That means that for the first two attribute
types, the filesystem will always see a sufficiently-sized buffer
allocated. Further, this removes the possibility of the filesystem
gaining access to the userspace buffer.
fsinfo() allows a variety of information to be retrieved about a
filesystem and the mount topology:
(1) General superblock attributes:
- Filesystem identifiers (UUID, volume label, device numbers,
...)
- The limits on a filesystem's capabilities
- Information on supported statx fields and attributes and IOC
flags.
- A variety single-bit flags indicating supported capabilities.
- Timestamp resolution and range.
- The amount of space/free space in a filesystem (as statfs()).
- Superblock notification counter.
(2) Filesystem-specific superblock attributes:
- Superblock-level timestamps.
- Cell name.
- Server names and addresses.
- Filesystem-specific information.
(3) VFS information:
- Mount topology information.
- Mount attributes.
- Mount notification counter.
(4) Information about what the fsinfo() syscall itself supports,
including
the type and struct/element size of attributes.
The system is extensible:
(1) New attributes can be added. There is no requirement that a
filesystem implement every attribute. Note that the core VFS
keeps a
table of types and sizes so it can handle future extensibility
rather
than delegating this to the filesystems.
(2) Version length-dependent structure attributes can be made larger
and
have additional information tacked on the end, provided it keeps
the
layout of the existing fields. If an older process asks for a
shorter
structure, it will only be given the bits it asks for. If a
newer
process asks for a longer structure on an older kernel, the
extra
space will be set to 0. In all cases, the size of the data
actually
available is returned.
In essence, the size of a structure is that structure's version:
a
smaller size is an earlier version and a later version includes
everything that the earlier version did.
(3) New single-bit capability flags can be added. This is a
structure-typed
attribute and, as such, (2) applies. Any bits you wanted but
the kernel
doesn't support are automatically set to 0.
fsinfo() may be called like the following, for example:
struct fsinfo_params params = {
.at_flags = AT_SYMLINK_NOFOLLOW,
.flags = FSINFO_FLAGS_QUERY_PATH,
.request = FSINFO_ATTR_AFS_SERVER_ADDRESSES,
.Nth = 2,
};
struct fsinfo_server_address address;
len = fsinfo(AT_FDCWD, "/afs/grand.central.org/doc", ¶ms,
&address, sizeof(address));
The above example would query an AFS filesystem to retrieve the
address
list for the 3rd server, and:
struct fsinfo_params params = {
.at_flags = AT_SYMLINK_NOFOLLOW,
.flags = FSINFO_FLAGS_QUERY_PATH,
.request = FSINFO_ATTR_AFS_CELL_NAME;
};
char cell_name[256];
len = fsinfo(AT_FDCWD, "/afs/grand.central.org/doc", ¶ms,
&cell_name, sizeof(cell_name));
would retrieve the name of an AFS cell as a string.
In future, I want to make fsinfo() capable of querying a context
created by
fsopen() or fspick(), e.g.:
fd = fsopen("ext4", 0);
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_PARAMETERS;
};
char buffer[65536];
fsinfo(fd, NULL, ¶ms, &buffer, sizeof(buffer));
even if that context doesn't currently have a superblock attached. I
would prefer this to contain length-prefixed strings so that there's
no need to insert escaping, especially as any character, including
'\', can be used as the separator in cifs and so that binary
parameters can be returned (though that is a lesser issue).
Could I make a suggestion about how this should be done in a way that
doesn't actually require the fsinfo syscall at all: it could just be
done with fsconfig. The idea is based on something I've wanted to do
for configfd but couldn't because otherwise it wouldn't substitute for
fsconfig, but Christian made me think it was actually essential to the
ability of the seccomp and other verifier tools in the critique of
configfd and I belive the same critique applies here.
Instead of making fsconfig functionally configure ... as in you pass
the attribute name, type and parameters down into the fs specific
handler and the handler does a string match and then verifies the
parameters and then acts on them, make it table configured, so what
each fstype does is register a table of attributes which can be got and
optionally set (with each attribute having a get and optional set
function). We'd have multiple tables per fstype, so the generic VFS
can register a table of attributes it understands for every fstype
(things like name, uuid and the like) and then each fs type would
register a table of fs specific attributes following the same pattern.
The system would examine the fs specific table before the generic one,
allowing overrides. fsconfig would have the ability to both get and
set attributes, permitting retrieval as well as setting (which is how I
get rid of the fsinfo syscall), we'd have a global parameter, which
would retrieve the entire table by name and type so the whole thing is
introspectable because the upper layer knows a-priori all the
attributes which can be set for a given fs type and what type they are
(so we can make more of the parsing generic). Any attribute which
doesn't have a set routine would be read only and all attributes would
have to have a get routine meaning everything is queryable.
And that makes me wonder: would a
"/sys/class/fs/$ST_DEV/options/$OPTION" type interface be feasible for
this?
Thanks,
Miklos
From: James Bottomley <James.Bottomley@HansenPartnership.com> Date: 2020-02-24 14:55:42
On Mon, 2020-02-24 at 11:24 +0100, Miklos Szeredi wrote:
On Fri, Feb 21, 2020 at 9:21 PM James Bottomley
[off-list ref] wrote:
[...]
quoted
Could I make a suggestion about how this should be done in a way
that doesn't actually require the fsinfo syscall at all: it could
just be done with fsconfig. The idea is based on something I've
wanted to do for configfd but couldn't because otherwise it
wouldn't substitute for fsconfig, but Christian made me think it
was actually essential to the ability of the seccomp and other
verifier tools in the critique of configfd and I belive the same
critique applies here.
Instead of making fsconfig functionally configure ... as in you
pass the attribute name, type and parameters down into the fs
specific handler and the handler does a string match and then
verifies the parameters and then acts on them, make it table
configured, so what each fstype does is register a table of
attributes which can be got and optionally set (with each attribute
having a get and optional set function). We'd have multiple tables
per fstype, so the generic VFS can register a table of attributes
it understands for every fstype (things like name, uuid and the
like) and then each fs type would register a table of fs specific
attributes following the same pattern. The system would examine the
fs specific table before the generic one, allowing
overrides. fsconfig would have the ability to both get and
set attributes, permitting retrieval as well as setting (which is
how I get rid of the fsinfo syscall), we'd have a global parameter,
which would retrieve the entire table by name and type so the whole
thing is introspectable because the upper layer knows a-priori all
the attributes which can be set for a given fs type and what type
they are (so we can make more of the parsing generic). Any
attribute which doesn't have a set routine would be read only and
all attributes would have to have a get routine meaning everything
is queryable.
And that makes me wonder: would a
"/sys/class/fs/$ST_DEV/options/$OPTION" type interface be feasible
for this?
Once it's table driven, certainly a sysfs directory becomes possible.
The problem with ST_DEV is filesystems like btrfs and xfs that may have
multiple devices. The current fsinfo takes a fspick'd directory fd so
the input to the query is a path, which gets messy in sysfs, although I
could see something like /sys/class/fs/mount/<path>/$OPTION working.
James
On Mon, Feb 24, 2020 at 3:55 PM James Bottomley
[off-list ref] wrote:
Once it's table driven, certainly a sysfs directory becomes possible.
The problem with ST_DEV is filesystems like btrfs and xfs that may have
multiple devices.
For XFS there's always a single sb->s_dev though, that's what st_dev
will be set to on all files.
Btrfs subvolume is sort of a lightweight superblock, so basically all
such st_dev's are aliases of the same master superblock. So lookup of
all subvolume st_dev's could result in referencing the same underlying
struct super_block (just like /proc/$PID will reference the same
underlying task group regardless of which of the task group member's
PID is used).
Having this info in sysfs would spare us a number of issues that a set
of new syscalls would bring. The question is, would that be enough,
or is there a reason that sysfs can't be used to present the various
filesystem related information that fsinfo is supposed to present?
Thanks,
Miklos
From: Steven Whitehouse <hidden> Date: 2020-02-25 12:13:25
Hi,
On 24/02/2020 15:28, Miklos Szeredi wrote:
On Mon, Feb 24, 2020 at 3:55 PM James Bottomley
[off-list ref] wrote:
quoted
Once it's table driven, certainly a sysfs directory becomes possible.
The problem with ST_DEV is filesystems like btrfs and xfs that may have
multiple devices.
For XFS there's always a single sb->s_dev though, that's what st_dev
will be set to on all files.
Btrfs subvolume is sort of a lightweight superblock, so basically all
such st_dev's are aliases of the same master superblock. So lookup of
all subvolume st_dev's could result in referencing the same underlying
struct super_block (just like /proc/$PID will reference the same
underlying task group regardless of which of the task group member's
PID is used).
Having this info in sysfs would spare us a number of issues that a set
of new syscalls would bring. The question is, would that be enough,
or is there a reason that sysfs can't be used to present the various
filesystem related information that fsinfo is supposed to present?
Thanks,
Miklos
We need a unique id for superblocks anyway. I had wondered about using s_dev some time back, but for the reasons mentioned earlier in this thread I think it might just land up being confusing and difficult to manage. While fake s_devs are created for sbs that don't have a device, I can't help thinking that something closer to ifindex, but for superblocks, is needed here. That would avoid the issue of which device number to use.
In fact we need that anyway for the notifications, since without that there is a race that can lead to missing remounts of the same device, in case a umount/mount pair is missed due to an overrun, and then fsinfo returns the same device as before, with potentially the same mount options too. So I think a unique id for a superblock is a generically useful feature, which would also allow for sensible sysfs directory naming, if required,
Steve.
From: James Bottomley <James.Bottomley@HansenPartnership.com> Date: 2020-02-25 15:29:02
On Tue, 2020-02-25 at 12:13 +0000, Steven Whitehouse wrote:
Hi,
On 24/02/2020 15:28, Miklos Szeredi wrote:
quoted
On Mon, Feb 24, 2020 at 3:55 PM James Bottomley
[off-list ref] wrote:
quoted
Once it's table driven, certainly a sysfs directory becomes
possible. The problem with ST_DEV is filesystems like btrfs and
xfs that may have multiple devices.
For XFS there's always a single sb->s_dev though, that's what
st_dev will be set to on all files.
Btrfs subvolume is sort of a lightweight superblock, so basically
all such st_dev's are aliases of the same master superblock. So
lookup of all subvolume st_dev's could result in referencing the
same underlying struct super_block (just like /proc/$PID will
reference the same underlying task group regardless of which of the
task group member's PID is used).
Having this info in sysfs would spare us a number of issues that a
set of new syscalls would bring. The question is, would that be
enough, or is there a reason that sysfs can't be used to present
the various filesystem related information that fsinfo is supposed
to present?
Thanks,
Miklos
We need a unique id for superblocks anyway. I had wondered about
using s_dev some time back, but for the reasons mentioned earlier in
this thread I think it might just land up being confusing and
difficult to manage. While fake s_devs are created for sbs that don't
have a device, I can't help thinking that something closer to
ifindex, but for superblocks, is needed here. That would avoid the
issue of which device number to use.
In fact we need that anyway for the notifications, since without
that there is a race that can lead to missing remounts of the same
device, in case a umount/mount pair is missed due to an overrun, and
then fsinfo returns the same device as before, with potentially the
same mount options too. So I think a unique id for a superblock is a
generically useful feature, which would also allow for sensible sysfs
directory naming, if required,
But would this be informative and useful for the user? I'm sure we can
find a persistent id for a persistent superblock, but what about tmpfs
... that's going to have to change with every reboot. It's going to be
remarkably inconvenient if I want to get fsinfo on /run to have to keep
finding what the id is.
The other thing a file descriptor does that sysfs doesn't is that it
solves the information leak: if I'm in a mount namespace that has no
access to certain mounts, I can't fspick them and thus I can't see the
information. By default, with sysfs I can.
James
From: Steven Whitehouse <hidden> Date: 2020-02-25 15:47:52
Hi,
On 25/02/2020 15:28, James Bottomley wrote:
On Tue, 2020-02-25 at 12:13 +0000, Steven Whitehouse wrote:
quoted
Hi,
On 24/02/2020 15:28, Miklos Szeredi wrote:
quoted
On Mon, Feb 24, 2020 at 3:55 PM James Bottomley
[off-list ref] wrote:
quoted
Once it's table driven, certainly a sysfs directory becomes
possible. The problem with ST_DEV is filesystems like btrfs and
xfs that may have multiple devices.
For XFS there's always a single sb->s_dev though, that's what
st_dev will be set to on all files.
Btrfs subvolume is sort of a lightweight superblock, so basically
all such st_dev's are aliases of the same master superblock. So
lookup of all subvolume st_dev's could result in referencing the
same underlying struct super_block (just like /proc/$PID will
reference the same underlying task group regardless of which of the
task group member's PID is used).
Having this info in sysfs would spare us a number of issues that a
set of new syscalls would bring. The question is, would that be
enough, or is there a reason that sysfs can't be used to present
the various filesystem related information that fsinfo is supposed
to present?
Thanks,
Miklos
We need a unique id for superblocks anyway. I had wondered about
using s_dev some time back, but for the reasons mentioned earlier in
this thread I think it might just land up being confusing and
difficult to manage. While fake s_devs are created for sbs that don't
have a device, I can't help thinking that something closer to
ifindex, but for superblocks, is needed here. That would avoid the
issue of which device number to use.
In fact we need that anyway for the notifications, since without
that there is a race that can lead to missing remounts of the same
device, in case a umount/mount pair is missed due to an overrun, and
then fsinfo returns the same device as before, with potentially the
same mount options too. So I think a unique id for a superblock is a
generically useful feature, which would also allow for sensible sysfs
directory naming, if required,
But would this be informative and useful for the user? I'm sure we can
find a persistent id for a persistent superblock, but what about tmpfs
... that's going to have to change with every reboot. It's going to be
remarkably inconvenient if I want to get fsinfo on /run to have to keep
finding what the id is.
That is a different question though, or at least it might be... the idea of the superblock id is to uniquely identify a particular superblock. The mount notification should give you the association between that superblock and any devices (assuming those are applicable), or you can use fsinfo if you were not listening to the notifications at the time of the mount to get the same information.
If someone unmounts /run and remounts it, then the superblock id would change, but otherwise it would stay the same, so you know that it is the same mount that is being described in future notifications. One of the main aims here being to combine the fsinfo information with the notifications in a race free manner.
There are a number of ways one might want to specify a filesystem: by device, by uuid, by volume label and so forth but we can't use any of those very easily as a unique id. Someone might remove a drive and replace it with a different one (so same device, but different content) or they might have two filesystems with the same uuid if they've just done a dd copy to a new device. For the mount notifications we need something that doesn't suffer from these issues, but which can also be very easily associated with what in most cases are more convenient ways to specify a particular filesystem.
The other thing a file descriptor does that sysfs doesn't is that it
solves the information leak: if I'm in a mount namespace that has no
access to certain mounts, I can't fspick them and thus I can't see the
information. By default, with sysfs I can.
James
Yes, thats true, and I wasn't advocating for the sysfs method over fspick here, just pointing out that a unique superblock id would be a generically useful thing to have,
Steve.
On 2020-02-21, David Howells [off-list ref] wrote:
Add a system call to allow filesystem information to be queried. A request
value can be given to indicate the desired attribute. Support is provided
for enumerating multi-value attributes.
===============
NEW SYSTEM CALL
===============
The new system call looks like:
int ret = fsinfo(int dfd,
const char *filename,
const struct fsinfo_params *params,
void *buffer,
size_t buf_size);
The params parameter optionally points to a block of parameters:
struct fsinfo_params {
__u32 at_flags;
__u32 flags;
__u32 request;
__u32 Nth;
__u32 Mth;
__u64 __reserved[3];
};
If params is NULL, it is assumed params->request should be
FSINFO_ATTR_STATFS, params->Nth should be 0, params->Mth should be 0,
params->at_flags should be 0 and params->flags should be 0.
If params is given, all of params->__reserved[] must be 0.
I would suggest that rather than having a reserved field for future
extensions, you make use of copy_struct_from_user() and have extensible
structs:
int ret = fsinfo(int dfd,
const char *filename,
struct fsinfo_params *params,
size_t params_usize,
void *buffer,
size_t buf_usize);
struct fsinfo_params {
__u64 flags;
__u32 at_flags;
__u32 request;
__u32 Nth;
__u32 Mth;
};
I dropped the "const" on fsinfo_params because the planned CHECK_FiELDS
feature for extensible-struct syscalls requires writing to the struct. I
also switched the flags field to u64 because CHECK_FiELDS is intended to
use (1<<63) for all syscalls (this has the nice benefit of removing the
need of a padding field entirely).
dfd, filename and params->at_flags indicate the file to query. There is no
equivalent of lstat() as that can be emulated with fsinfo() by setting
AT_SYMLINK_NOFOLLOW in params->at_flags.
Minor gripe -- can we make the default be AT_SYMLINK_NOFOLLOW and you
need to explicitly pass AT_SYMLINK_FOLLOW? Accidentally following
symlinks is a constant source of security bugs.
There is also no equivalent of fstat() as that can be emulated by
passing a NULL filename to fsinfo() with the fd of interest in dfd.
Presumably you also need to pass AT_EMPTY_PATH?
params->request indicates the attribute/attributes to be queried. This can
be one of:
FSINFO_ATTR_STATFS - statfs-style info
FSINFO_ATTR_IDS - Filesystem IDs
FSINFO_ATTR_LIMITS - Filesystem limits
FSINFO_ATTR_SUPPORTS - What's supported in statx(), IOC flags
FSINFO_ATTR_TIMESTAMP_INFO - Inode timestamp info
FSINFO_ATTR_VOLUME_ID - Volume ID (string)
FSINFO_ATTR_VOLUME_UUID - Volume UUID
FSINFO_ATTR_VOLUME_NAME - Volume name (string)
FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO - Information about attr Nth
FSINFO_ATTR_FSINFO_ATTRIBUTES - List of supported attrs
Some attributes (such as the servers backing a network filesystem) can have
multiple values. These can be enumerated by setting params->Nth and
params->Mth to 0, 1, ... until ENODATA is returned.
buffer and buf_size point to the reply buffer. The buffer is filled up to
the specified size, even if this means truncating the reply. The full size
of the reply is returned. In future versions, this will allow extra fields
to be tacked on to the end of the reply, but anyone not expecting them will
only get the subset they're expecting. If either buffer of buf_size are 0,
no copy will take place and the data size will be returned.
Sounds good, though I think we should zero-fill the tail end of the
buffer (if the buffer is larger than the in-kernel one). This is
basically what a theoretical copy_struct_to_user() would do. It will
also ensure that CHECK_FiELDS will act consistently on a syscall that
has two extensible struct arguments.
quoted hunk
At the moment, this will only work on x86_64 and i386 as it requires the
system call to be wired up.
Signed-off-by: David Howells <dhowells@redhat.com>
cc: linux-api@vger.kernel.org
---
arch/alpha/kernel/syscalls/syscall.tbl | 1
arch/arm/tools/syscall.tbl | 1
arch/arm64/include/asm/unistd.h | 2
arch/ia64/kernel/syscalls/syscall.tbl | 1
arch/m68k/kernel/syscalls/syscall.tbl | 1
arch/microblaze/kernel/syscalls/syscall.tbl | 1
arch/mips/kernel/syscalls/syscall_n32.tbl | 1
arch/mips/kernel/syscalls/syscall_n64.tbl | 1
arch/mips/kernel/syscalls/syscall_o32.tbl | 1
arch/parisc/kernel/syscalls/syscall.tbl | 1
arch/powerpc/kernel/syscalls/syscall.tbl | 1
arch/s390/kernel/syscalls/syscall.tbl | 1
arch/sh/kernel/syscalls/syscall.tbl | 1
arch/sparc/kernel/syscalls/syscall.tbl | 1
arch/x86/entry/syscalls/syscall_32.tbl | 1
arch/x86/entry/syscalls/syscall_64.tbl | 1
arch/xtensa/kernel/syscalls/syscall.tbl | 1
fs/Kconfig | 7
fs/Makefile | 1
fs/fsinfo.c | 566 +++++++++++++++++++++++++
include/linux/fs.h | 4
include/linux/fsinfo.h | 72 +++
include/linux/syscalls.h | 4
include/uapi/asm-generic/unistd.h | 4
include/uapi/linux/fsinfo.h | 187 ++++++++
kernel/sys_ni.c | 1
samples/vfs/Makefile | 5
samples/vfs/test-fsinfo.c | 607 +++++++++++++++++++++++++++
28 files changed, 1474 insertions(+), 2 deletions(-)
create mode 100644 fs/fsinfo.c
create mode 100644 include/linux/fsinfo.h
create mode 100644 include/uapi/linux/fsinfo.h
create mode 100644 samples/vfs/test-fsinfo.c
@@ -479,3 +479,4 @@ 548 common pidfd_getfd sys_pidfd_getfd 549 common watch_mount sys_watch_mount 550 common watch_sb sys_watch_sb+551 common fsinfo sys_fsinfo
@@ -453,3 +453,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -360,3 +360,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -439,3 +439,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -445,3 +445,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -437,3 +437,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -521,3 +521,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -442,3 +442,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -485,3 +485,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -361,6 +361,7 @@ 438 common pidfd_getfd __x64_sys_pidfd_getfd 439 common watch_mount __x64_sys_watch_mount 440 common watch_sb __x64_sys_watch_sb+441 common fsinfo __x64_sys_fsinfo # # x32-specific system call numbers start at 512 to avoid cache impact
@@ -410,3 +410,4 @@ 438 common pidfd_getfd sys_pidfd_getfd 439 common watch_mount sys_watch_mount 440 common watch_sb sys_watch_sb+441 common fsinfo sys_fsinfo
@@ -15,6 +15,13 @@ config VALIDATE_FS_PARSEREnablethistoperformvalidationoftheparameterdescriptionforafilesystemwhenitisregistered.+configFSINFO+bool"Enable the fsinfo() system call"+help+Enablethefilesysteminformationqueryingsystemcalltoallow+comprehensiveinformationtoberetrievedaboutafilesystem,+superblockormountobject.+ifBLOCKconfigFS_IOMAP
@@ -0,0 +1,566 @@+// SPDX-License-Identifier: GPL-2.0+/* Filesystem information query.+*+*Copyright(C)2020RedHat,Inc.AllRightsReserved.+*WrittenbyDavidHowells(dhowells@redhat.com)+*/+#include<linux/syscalls.h>+#include<linux/fs.h>+#include<linux/file.h>+#include<linux/mount.h>+#include<linux/namei.h>+#include<linux/statfs.h>+#include<linux/security.h>+#include<linux/uaccess.h>+#include<linux/fsinfo.h>+#include<uapi/linux/mount.h>+#include"internal.h"++/**+*fsinfo_string-StoreaNUL-terminatedstringasanfsinfoattributevalue.+*@s:Thestringtostore(maybeNULL)+*@ctx:Theparametercontext+*/+intfsinfo_string(constchar*s,structfsinfo_context*ctx)+{+unsignedintlen;+char*p=ctx->buffer;+intret=0;++if(s){+len=min_t(size_t,strlen(s),ctx->buf_size-1);+if(!ctx->want_size_only){+memcpy(p,s,len);+p[len]=0;+}+ret=len;+}++returnret;+}+EXPORT_SYMBOL(fsinfo_string);++/*+*Getbasicfilesystemstatsfromstatfs.+*/+staticintfsinfo_generic_statfs(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_statfs*p=ctx->buffer;+structkstatfsbuf;+intret;++ret=vfs_statfs(path,&buf);+if(ret<0)+returnret;++p->f_blocks.lo=buf.f_blocks;+p->f_bfree.lo=buf.f_bfree;+p->f_bavail.lo=buf.f_bavail;+p->f_files.lo=buf.f_files;+p->f_ffree.lo=buf.f_ffree;+p->f_favail.lo=buf.f_ffree;+p->f_bsize=buf.f_bsize;+p->f_frsize=buf.f_frsize;+returnsizeof(*p);+}++staticintfsinfo_generic_ids(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_ids*p=ctx->buffer;+structsuper_block*sb;+structkstatfsbuf;+intret;++ret=vfs_statfs(path,&buf);+if(ret<0&&ret!=-ENOSYS)+returnret;+if(ret==0)+memcpy(&p->f_fsid,&buf.f_fsid,sizeof(p->f_fsid));++sb=path->dentry->d_sb;+p->f_fstype=sb->s_magic;+p->f_dev_major=MAJOR(sb->s_dev);+p->f_dev_minor=MINOR(sb->s_dev);+p->f_sb_id=sb->s_unique_id;+strlcpy(p->f_fs_name,sb->s_type->name,sizeof(p->f_fs_name));+returnsizeof(*p);+}++intfsinfo_generic_limits(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_limits*p=ctx->buffer;+structsuper_block*sb=path->dentry->d_sb;++p->max_file_size.hi=0;+p->max_file_size.lo=sb->s_maxbytes;+p->max_ino.hi=0;+p->max_ino.lo=UINT_MAX;+p->max_hard_links=sb->s_max_links;+p->max_uid=UINT_MAX;+p->max_gid=UINT_MAX;+p->max_projid=UINT_MAX;+p->max_filename_len=NAME_MAX;+p->max_symlink_len=PATH_MAX;+p->max_xattr_name_len=XATTR_NAME_MAX;+p->max_xattr_body_len=XATTR_SIZE_MAX;+p->max_dev_major=0xffffff;+p->max_dev_minor=0xff;+returnsizeof(*p);+}+EXPORT_SYMBOL(fsinfo_generic_limits);++intfsinfo_generic_supports(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_supports*p=ctx->buffer;+structsuper_block*sb=path->dentry->d_sb;++p->stx_mask=STATX_BASIC_STATS;+if(sb->s_d_op&&sb->s_d_op->d_automount)+p->stx_attributes|=STATX_ATTR_AUTOMOUNT;+returnsizeof(*p);+}+EXPORT_SYMBOL(fsinfo_generic_supports);++staticconststructfsinfo_timestamp_infofsinfo_default_timestamp_info={+.atime={+.minimum=S64_MIN,+.maximum=S64_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.mtime={+.minimum=S64_MIN,+.maximum=S64_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.ctime={+.minimum=S64_MIN,+.maximum=S64_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+.btime={+.minimum=S64_MIN,+.maximum=S64_MAX,+.gran_mantissa=1,+.gran_exponent=0,+},+};++intfsinfo_generic_timestamp_info(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_timestamp_info*p=ctx->buffer;+structsuper_block*sb=path->dentry->d_sb;+s8exponent;++*p=fsinfo_default_timestamp_info;++if(sb->s_time_gran<1000000000){+if(sb->s_time_gran<1000)+exponent=-9;+elseif(sb->s_time_gran<1000000)+exponent=-6;+else+exponent=-3;++p->atime.gran_exponent=exponent;+p->mtime.gran_exponent=exponent;+p->ctime.gran_exponent=exponent;+p->btime.gran_exponent=exponent;+}++returnsizeof(*p);+}+EXPORT_SYMBOL(fsinfo_generic_timestamp_info);++staticintfsinfo_generic_volume_uuid(structpath*path,structfsinfo_context*ctx)+{+structfsinfo_volume_uuid*p=ctx->buffer;+structsuper_block*sb=path->dentry->d_sb;++memcpy(p,&sb->s_uuid,sizeof(*p));+returnsizeof(*p);+}++staticintfsinfo_generic_volume_id(structpath*path,structfsinfo_context*ctx)+{+returnfsinfo_string(path->dentry->d_sb->s_id,ctx);+}++staticconststructfsinfo_attributefsinfo_common_attributes[]={+FSINFO_VSTRUCT(FSINFO_ATTR_STATFS,fsinfo_generic_statfs),+FSINFO_VSTRUCT(FSINFO_ATTR_IDS,fsinfo_generic_ids),+FSINFO_VSTRUCT(FSINFO_ATTR_LIMITS,fsinfo_generic_limits),+FSINFO_VSTRUCT(FSINFO_ATTR_SUPPORTS,fsinfo_generic_supports),+FSINFO_VSTRUCT(FSINFO_ATTR_TIMESTAMP_INFO,fsinfo_generic_timestamp_info),+FSINFO_STRING(FSINFO_ATTR_VOLUME_ID,fsinfo_generic_volume_id),+FSINFO_VSTRUCT(FSINFO_ATTR_VOLUME_UUID,fsinfo_generic_volume_uuid),++FSINFO_LIST(FSINFO_ATTR_FSINFO_ATTRIBUTES,(void*)123UL),+FSINFO_VSTRUCT_N(FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO,(void*)123UL),+{}+};++/*+*Determineanattribute'sminimumbuffersizeand,ifthebufferislarge+*enough,gettheattributevalue.+*/+staticintfsinfo_get_this_attribute(structpath*path,+structfsinfo_context*ctx,+conststructfsinfo_attribute*attr)+{+intbuf_size;++if(ctx->Nth!=0&&!(attr->flags&(FSINFO_FLAGS_N|FSINFO_FLAGS_NM)))+return-ENODATA;+if(ctx->Mth!=0&&!(attr->flags&FSINFO_FLAGS_NM))+return-ENODATA;++switch(attr->type){+caseFSINFO_TYPE_VSTRUCT:+ctx->clear_tail=true;+buf_size=attr->size;+break;+caseFSINFO_TYPE_STRING:+caseFSINFO_TYPE_OPAQUE:+caseFSINFO_TYPE_LIST:+buf_size=4096;+break;+default:+return-ENOPKG;+}++if(ctx->buf_size<buf_size)+returnbuf_size;++returnattr->get(path,ctx);+}++staticvoidfsinfo_attributes_insert(structfsinfo_context*ctx,+conststructfsinfo_attribute*attr)+{+__u32*p=ctx->buffer;+unsignedinti;++if(ctx->usage>=ctx->buf_size||+ctx->buf_size-ctx->usage<sizeof(__u32)){+ctx->usage+=sizeof(__u32);+return;+}++for(i=0;i<ctx->usage/sizeof(__u32);i++)+if(p[i]==attr->attr_id)+return;++p[i]=attr->attr_id;+ctx->usage+=sizeof(__u32);+}++staticintfsinfo_list_attributes(structpath*path,+structfsinfo_context*ctx,+conststructfsinfo_attribute*attributes)+{+conststructfsinfo_attribute*a;++for(a=attributes;a->get;a++)+fsinfo_attributes_insert(ctx,a);+return-EOPNOTSUPP;/* We want to go through all the lists */+}++staticintfsinfo_get_attribute_info(structpath*path,+structfsinfo_context*ctx,+conststructfsinfo_attribute*attributes)+{+conststructfsinfo_attribute*a;+structfsinfo_attribute_info*p=ctx->buffer;++if(!ctx->buf_size)+returnsizeof(*p);++for(a=attributes;a->get;a++){+if(a->attr_id==ctx->Nth){+p->attr_id=a->attr_id;+p->type=a->type;+p->flags=a->flags;+p->size=a->size;+p->size=a->size;+returnsizeof(*p);+}+}+return-EOPNOTSUPP;/* We want to go through all the lists */+}++/**+*fsinfo_get_attribute-Lookupandhandleanattribute+*@path:Theobjecttoquery+*@params:Parameterstodefinearequestandplacetostoreresult+*@attributes:Listofattributestosearch.+*+*Lookthroughalistofattributesforonethatmatchestherequested+*attributethencallthehandlerforit.+*/+intfsinfo_get_attribute(structpath*path,structfsinfo_context*ctx,+conststructfsinfo_attribute*attributes)+{+conststructfsinfo_attribute*a;++switch(ctx->requested_attr){+caseFSINFO_ATTR_FSINFO_ATTRIBUTE_INFO:+returnfsinfo_get_attribute_info(path,ctx,attributes);+caseFSINFO_ATTR_FSINFO_ATTRIBUTES:+returnfsinfo_list_attributes(path,ctx,attributes);+default:+for(a=attributes;a->get;a++)+if(a->attr_id==ctx->requested_attr)+returnfsinfo_get_this_attribute(path,ctx,a);+return-EOPNOTSUPP;+}+}+EXPORT_SYMBOL(fsinfo_get_attribute);++/**+*generic_fsinfo-Handleanfsinfoattributegenerically+*@path:Theobjecttoquery+*@params:Parameterstodefinearequestandplacetostoreresult+*/+staticintfsinfo_call(structpath*path,structfsinfo_context*ctx)+{+intret;++if(path->dentry->d_sb->s_op->fsinfo){+ret=path->dentry->d_sb->s_op->fsinfo(path,ctx);+if(ret!=-EOPNOTSUPP)+returnret;+}+ret=fsinfo_get_attribute(path,ctx,fsinfo_common_attributes);+if(ret!=-EOPNOTSUPP)+returnret;++switch(ctx->requested_attr){+caseFSINFO_ATTR_FSINFO_ATTRIBUTE_INFO:+return-ENODATA;+caseFSINFO_ATTR_FSINFO_ATTRIBUTES:+returnctx->usage;+default:+return-EOPNOTSUPP;+}+}++/**+*vfs_fsinfo-Retrievefilesysteminformation+*@path:Theobjecttoquery+*@params:Parameterstodefinearequestandplacetostoreresult+*+*Getanattributeonafilesystemoranobjectwithinafilesystem.The+*filesystemattributetobequeriedisindicatedby@ctx->requested_attr,and+*ifit'samulti-valuedattribute,theparticularvalueisselectedby+*@ctx->Nthandthen@ctx->Mth.+*+*Forcommonattributes,avaluemaybefabricatedifitisnotsupportedby+*thefilesystem.+*+*Onsuccess,thesizeoftheattribute'svalueisreturned(0isavalid+*size).Abufferwillhavebeenallocatedandwillbepointedtoby+*@ctx->buffer.Thecallermustfreethiswithkvfree().+*+*Errorscanalsobereturned:-ENOMEMifabuffercannotbeallocated,-EPERM+*or-EACCESifpermissionisdeniedbytheLSM,-EOPNOTSUPPifanattribute+*doesn'texistforthespecifiedobjector-ENODATAiftheattributeexists,+*buttheNth,Mthvaluedoesnotexist.-EMSGSIZEindicatesthatthevalueis+*unmanageableinternallyand-ENOPKGindicatesotherinternalfailure.+*+*Errorssuchas-EIOmayalsocomefromattemptstoaccessmediaorservers+*toobtaintherequestedinformationifit'snotimmediatelytohand.+*+*[*]Notethatthecallermayset@ctx->want_size_onlyifitonlywantsthe+*sizeofthevalueandnotthedata.Ifthisisset,abuffermaynotbe+*allocatedundersomecircumstances.Thisisintendedforsizequeryby+*userspace.+*+*[*]Notethat@ctx->clear_tailwillbereturnedsetifthedatashouldbe+*paddedoutwithzeroswhenwritingittouserspace.+*/+staticintvfs_fsinfo(structpath*path,structfsinfo_context*ctx)+{+structdentry*dentry=path->dentry;+intret;++ret=security_sb_statfs(dentry);+if(ret)+returnret;++/* Call the handler to find out the buffer size required. */+ctx->buf_size=0;+ret=fsinfo_call(path,ctx);+if(ret<0||ctx->want_size_only)+returnret;+ctx->buf_size=ret;++do{+/* Allocate a buffer of the requested size. */+if(ctx->buf_size>INT_MAX)+return-EMSGSIZE;+ctx->buffer=kvzalloc(ctx->buf_size,GFP_KERNEL);+if(!ctx->buffer)+return-ENOMEM;++ctx->usage=0;+ret=fsinfo_call(path,ctx);+if(IS_ERR_VALUE((long)ret))+returnret;+if((unsignedint)ret<=ctx->buf_size)+returnret;/* It fitted */++/* We need to resize the buffer */+ctx->buf_size=roundup(ret,PAGE_SIZE);+kvfree(ctx->buffer);+ctx->buffer=NULL;+}while(!signal_pending(current));++return-ERESTARTSYS;+}++staticintvfs_fsinfo_path(intdfd,constchar__user*pathname,+unsignedintat_flags,structfsinfo_context*ctx)+{+structpathpath;+unsignedlookup_flags=LOOKUP_FOLLOW|LOOKUP_AUTOMOUNT;+intret=-EINVAL;++if((at_flags&~(AT_SYMLINK_NOFOLLOW|AT_NO_AUTOMOUNT|+AT_EMPTY_PATH))!=0)+return-EINVAL;++if(at_flags&AT_SYMLINK_NOFOLLOW)+lookup_flags&=~LOOKUP_FOLLOW;+if(at_flags&AT_NO_AUTOMOUNT)+lookup_flags&=~LOOKUP_AUTOMOUNT;+if(at_flags&AT_EMPTY_PATH)+lookup_flags|=LOOKUP_EMPTY;++retry:+ret=user_path_at(dfd,pathname,lookup_flags,&path);+if(ret)+gotoout;++ret=vfs_fsinfo(&path,ctx);+path_put(&path);+if(retry_estale(ret,lookup_flags)){+lookup_flags|=LOOKUP_REVAL;+gotoretry;+}+out:+returnret;+}++staticintvfs_fsinfo_fd(unsignedintfd,structfsinfo_context*ctx)+{+structfdf=fdget_raw(fd);+intret=-EBADF;++if(f.file){+ret=vfs_fsinfo(&f.file->f_path,ctx);+fdput(f);+}+returnret;+}++/**+*sys_fsinfo-Systemcalltogetfilesysteminformation+*@dfd:Basedirectorytopathwalkfromorfdreferringtofilesystem.+*@pathname:FilesystemtoqueryorNULL.+*@_params:Parameterstodefinerequest(orNULLforenhancedstatfs).+*@user_buffer:Resultbuffer.+*@user_buf_size:Sizeofresultbuffer.+*+*Getinformationonafilesystem.Thefilesystemattributetobequeriedis+*indicatedby@_params->request,andsomeoftheattributescanhavemultiple+*values,indexedby@_params->Nthand@_params->Mth.If@_paramsisNULL,+*thenthe0thfsinfo_attr_statfsattributeisqueried.Ifanattributedoes+*notexist,EOPNOTSUPPisreturned;iftheNth,Mthvaluedoesnotexist,+*ENODATAisreturned.+*+*Onsuccess,thesizeoftheattribute'svalueisreturned.If+*@user_buf_sizeis0or@user_bufferisNULL,onlythesizeisreturned.If+*thesizeofthevalueislargerthan@user_buf_size,itwillbetruncatedby+*thecopy.Ifthesizeofthevalueissmallerthan@user_buf_sizethenthe+*excessbufferspacewillbecleared.Thefullsizeofthevaluewillbe+*returned,irrespectiveofhowmuchdataisactuallyplacedinthebuffer.+*/+SYSCALL_DEFINE5(fsinfo,+int,dfd,constchar__user*,pathname,+structfsinfo_params__user*,params,+void__user*,user_buffer,size_t,user_buf_size)+{+structfsinfo_contextctx;+structfsinfo_paramsuser_params;+unsignedintat_flags=0,result_size;+intret;++if(!user_buffer&&user_buf_size)+return-EINVAL;+if(user_buffer&&!user_buf_size)+return-EINVAL;+if(user_buf_size>UINT_MAX)+return-EOVERFLOW;++memset(&ctx,0,sizeof(ctx));+ctx.requested_attr=FSINFO_ATTR_STATFS;+if(user_buf_size==0)+ctx.want_size_only=true;++if(params){+if(copy_from_user(&user_params,params,sizeof(user_params)))+return-EFAULT;+if(user_params.__reserved32[0]||+user_params.__reserved[0]||+user_params.__reserved[1]||+user_params.__reserved[2]||+user_params.flags&~FSINFO_FLAGS_QUERY_MASK)+return-EINVAL;+at_flags=user_params.at_flags;+ctx.flags=user_params.flags;+ctx.requested_attr=user_params.request;+ctx.Nth=user_params.Nth;+ctx.Mth=user_params.Mth;+}++switch(ctx.flags&FSINFO_FLAGS_QUERY_MASK){+caseFSINFO_FLAGS_QUERY_PATH:+ret=vfs_fsinfo_path(dfd,pathname,at_flags,&ctx);+break;+caseFSINFO_FLAGS_QUERY_FD:+if(pathname)+return-EINVAL;+ret=vfs_fsinfo_fd(dfd,&ctx);+break;+default:+return-EINVAL;+}++if(ret<0)+gotoerror;++result_size=min_t(size_t,ret,user_buf_size);+if(result_size>0&&+copy_to_user(user_buffer,ctx.buffer,result_size)!=0){+ret=-EFAULT;+gotoerror;+}++/* Clear any part of the buffer that we won't fill if we're putting a+*structinthere.Strings,opaqueobjectsandarraysareexpectedto+*bevariablelength.+*/+if(ctx.clear_tail&&+user_buf_size>result_size&&+clear_user(user_buffer+result_size,user_buf_size-result_size)!=0){+ret=-EFAULT;+gotoerror;+}++error:+kvfree(ctx.buffer);+returnret;+}
@@ -0,0 +1,187 @@+/* SPDX-License-Identifier: GPL-2.0 WITH Linux-syscall-note */+/* fsinfo() definitions.+*+*Copyright(C)2020RedHat,Inc.AllRightsReserved.+*WrittenbyDavidHowells(dhowells@redhat.com)+*/+#ifndef _UAPI_LINUX_FSINFO_H+#define _UAPI_LINUX_FSINFO_H++#include<linux/types.h>+#include<linux/socket.h>++/*+*Thefilesystemattributesthatcanberequested.Notethatsomeattributes+*mayhavemultipleinstanceswhichcanbeswitchedintheparameterblock.+*/+#define FSINFO_ATTR_STATFS 0x00 /* statfs()-style state */+#define FSINFO_ATTR_IDS 0x01 /* Filesystem IDs */+#define FSINFO_ATTR_LIMITS 0x02 /* Filesystem limits */+#define FSINFO_ATTR_SUPPORTS 0x03 /* What's supported in statx, iocflags, ... */+#define FSINFO_ATTR_TIMESTAMP_INFO 0x04 /* Inode timestamp info */+#define FSINFO_ATTR_VOLUME_ID 0x05 /* Volume ID (string) */+#define FSINFO_ATTR_VOLUME_UUID 0x06 /* Volume UUID (LE uuid) */+#define FSINFO_ATTR_VOLUME_NAME 0x07 /* Volume name (string) */++#define FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO 0x100 /* Information about attr N (for path) */+#define FSINFO_ATTR_FSINFO_ATTRIBUTES 0x101 /* List of supported attrs (for path) */++/*+*Optionalfsinfo()parameterstructure.+*+*Ifthisisnotgiven,itisassumedthatfsinfo_attr_statfsinstance0,0is+*desired.+*/+structfsinfo_params{+__u32at_flags;/* AT_SYMLINK_NOFOLLOW and similar flags */+__u32flags;/* Flags controlling fsinfo() specifically */+#define FSINFO_FLAGS_QUERY_MASK 0x0007 /* What object should fsinfo() query? */+#define FSINFO_FLAGS_QUERY_PATH 0x0000 /* - path, specified by dirfd,pathname,AT_EMPTY_PATH */+#define FSINFO_FLAGS_QUERY_FD 0x0001 /* - fd specified by dirfd */+__u32request;/* ID of requested attribute */+__u32Nth;/* Instance of it (some may have multiple) */+__u32Mth;/* Subinstance of Nth instance */+__u32__reserved32[1];/* Reserved params; all must be 0 */+__u64__reserved[3];+};++enumfsinfo_value_type{+FSINFO_TYPE_VSTRUCT=0,/* Version-lengthed struct (up to 4096 bytes) */+FSINFO_TYPE_STRING=1,/* NUL-term var-length string (up to 4095 chars) */+FSINFO_TYPE_OPAQUE=2,/* Opaque blob (unlimited size) */+FSINFO_TYPE_LIST=3,/* List of ints/structs (unlimited size) */+};++/*+*Informationstructforfsinfo(FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO).+*+*Thisgivesinformationabouttheattributessupportedbyfsinfoforthe+*givenpath.+*/+structfsinfo_attribute_info{+unsignedintattr_id;/* The ID of the attribute */+enumfsinfo_value_typetype;/* The type of the attribute's value(s) */+unsignedintflags;+#define FSINFO_FLAGS_N 0x01 /* - Attr has a set of values */+#define FSINFO_FLAGS_NM 0x02 /* - Attr has a set of sets of values */+unsignedintsize;/* - Value size (FSINFO_STRUCT/FSINFO_LIST) */+};++#define FSINFO_ATTR_FSINFO_ATTRIBUTE_INFO__STRUCT struct fsinfo_attribute_info+#define FSINFO_ATTR_FSINFO_ATTRIBUTES__STRUCT __u32++structfsinfo_u128{+#if defined(__BYTE_ORDER) ? __BYTE_ORDER == __BIG_ENDIAN : defined(__BIG_ENDIAN)+__u64hi;+__u64lo;+#elif defined(__BYTE_ORDER) ? __BYTE_ORDER == __LITTLE_ENDIAN : defined(__LITTLE_ENDIAN)+__u64lo;+__u64hi;+#endif+};++/*+*Informationstructforfsinfo(FSINFO_ATTR_STATFS).+*-Thisgivesextendedfilesysteminformation.+*/+structfsinfo_statfs{+structfsinfo_u128f_blocks;/* Total number of blocks in fs */+structfsinfo_u128f_bfree;/* Total number of free blocks */+structfsinfo_u128f_bavail;/* Number of free blocks available to ordinary user */+structfsinfo_u128f_files;/* Total number of file nodes in fs */+structfsinfo_u128f_ffree;/* Number of free file nodes */+structfsinfo_u128f_favail;/* Number of file nodes available to ordinary user */+__u64f_bsize;/* Optimal block size */+__u64f_frsize;/* Fragment size */+};++#define FSINFO_ATTR_STATFS__STRUCT struct fsinfo_statfs++/*+*Informationstructforfsinfo(FSINFO_ATTR_IDS).+*+*Listofbasicidentifiersasisnormallyfoundinstatfs().+*/+structfsinfo_ids{+charf_fs_name[15+1];/* Filesystem name */+__u64f_fsid;/* Short 64-bit Filesystem ID (as statfs) */+__u64f_sb_id;/* Internal superblock ID for sbnotify()/mntnotify() */+__u32f_fstype;/* Filesystem type from linux/magic.h [uncond] */+__u32f_dev_major;/* As st_dev_* from struct statx [uncond] */+__u32f_dev_minor;+__u32__padding[1];+};++#define FSINFO_ATTR_IDS__STRUCT struct fsinfo_ids++/*+*Informationstructforfsinfo(FSINFO_ATTR_LIMITS).+*+*Listofsupportedfilesystemlimits.+*/+structfsinfo_limits{+structfsinfo_u128max_file_size;/* Maximum file size */+structfsinfo_u128max_ino;/* Maximum inode number */+__u64max_uid;/* Maximum UID supported */+__u64max_gid;/* Maximum GID supported */+__u64max_projid;/* Maximum project ID supported */+__u64max_hard_links;/* Maximum number of hard links on a file */+__u64max_xattr_body_len;/* Maximum xattr content length */+__u32max_xattr_name_len;/* Maximum xattr name length */+__u32max_filename_len;/* Maximum filename length */+__u32max_symlink_len;/* Maximum symlink content length */+__u32max_dev_major;/* Maximum device major representable */+__u32max_dev_minor;/* Maximum device minor representable */+__u32__padding[1];+};++#define FSINFO_ATTR_LIMITS__STRUCT struct fsinfo_limits++/*+*Informationstructforfsinfo(FSINFO_ATTR_SUPPORTS).+*+*What'ssupportedinvariousmasks,suchasstatx()attributeandmaskbits+*andIOCflags.+*/+structfsinfo_supports{+__u64stx_attributes;/* What statx::stx_attributes are supported */+__u32stx_mask;/* What statx::stx_mask bits are supported */+__u32fs_ioc_getflags;/* What FS_IOC_GETFLAGS may return */+__u32fs_ioc_setflags_set;/* What FS_IOC_SETFLAGS may set */+__u32fs_ioc_setflags_clear;/* What FS_IOC_SETFLAGS may clear */+__u32win_file_attrs;/* What DOS/Windows FILE_* attributes are supported */+__u32__padding[1];+};++#define FSINFO_ATTR_SUPPORTS__STRUCT struct fsinfo_supports++structfsinfo_timestamp_one{+__s64minimum;/* Minimum timestamp value in seconds */+__s64maximum;/* Maximum timestamp value in seconds */+__u16gran_mantissa;/* Granularity(secs) = mant * 10^exp */+__s8gran_exponent;+__u8__padding[5];+};++/*+*Informationstructforfsinfo(FSINFO_ATTR_TIMESTAMP_INFO).+*/+structfsinfo_timestamp_info{+structfsinfo_timestamp_oneatime;/* Access time */+structfsinfo_timestamp_onemtime;/* Modification time */+structfsinfo_timestamp_onectime;/* Change time */+structfsinfo_timestamp_onebtime;/* Birth/creation time */+};++#define FSINFO_ATTR_TIMESTAMP_INFO__STRUCT struct fsinfo_timestamp_info++/*+*Informationstructforfsinfo(FSINFO_ATTR_VOLUME_UUID).+*/+structfsinfo_volume_uuid{+__u8uuid[16];+};++#define FSINFO_ATTR_VOLUME_UUID__STRUCT struct fsinfo_volume_uuid++#endif /* _UAPI_LINUX_FSINFO_H */
@@ -1,10 +1,15 @@# SPDX-License-Identifier: GPL-2.0-only# List of programs to build+hostprogs:=\+test-fsinfo\test-fsmount\test-statxalways-y:=$(hostprogs)+HOSTCFLAGS_test-fsinfo.o+=-I$(objtree)/usr/include+HOSTLDLIBS_test-fsinfo+=-static-lm+HOSTCFLAGS_test-fsmount.o+=-I$(objtree)/usr/includeHOSTCFLAGS_test-statx.o+=-I$(objtree)/usr/include
On Tue, Feb 25, 2020 at 4:29 PM James Bottomley
[off-list ref] wrote:
The other thing a file descriptor does that sysfs doesn't is that it
solves the information leak: if I'm in a mount namespace that has no
access to certain mounts, I can't fspick them and thus I can't see the
information. By default, with sysfs I can.
That's true, but procfs/sysfs has to deal with various namespacing
issues anyway. If this is just about hiding a number of entries, then
I don't think that's going to be a big deal.
The syscall API is efficient: single syscall per query instead of
several, no parsing necessary.
However, it is difficult to extend, because the ABI must be updated,
possibly libc and util-linux also, so that scripts can also consume
the new parameter. With the sysfs approach only the kernel needs to
be updated, and possibly only the filesystem code, not even the VFS.
So I think the question comes down to: do we need a highly efficient
way to query the superblock parameters all at once, or not?
Thanks,
Miklos
From: Steven Whitehouse <hidden> Date: 2020-02-26 10:51:13
Hi,
On 26/02/2020 09:11, Miklos Szeredi wrote:
On Tue, Feb 25, 2020 at 4:29 PM James Bottomley
[off-list ref] wrote:
quoted
The other thing a file descriptor does that sysfs doesn't is that it
solves the information leak: if I'm in a mount namespace that has no
access to certain mounts, I can't fspick them and thus I can't see the
information. By default, with sysfs I can.
That's true, but procfs/sysfs has to deal with various namespacing
issues anyway. If this is just about hiding a number of entries, then
I don't think that's going to be a big deal.
The syscall API is efficient: single syscall per query instead of
several, no parsing necessary.
However, it is difficult to extend, because the ABI must be updated,
possibly libc and util-linux also, so that scripts can also consume
the new parameter. With the sysfs approach only the kernel needs to
be updated, and possibly only the filesystem code, not even the VFS.
So I think the question comes down to: do we need a highly efficient
way to query the superblock parameters all at once, or not?
Thanks,
Miklos
That is Ian's use case for autofs I think, and it will also be what is needed at start up of most applications using the fs notifications, as well as at resync time if there has been an overrun leading to lost fs notification messages. We do need a solution that can scale to large numbers of mounts efficiently. Being able to extend it is also an important consideration too, so hopefully David has a solution to that,
Steve.
From: Ian Kent <raven@themaw.net> Date: 2020-02-27 05:06:53
On Wed, 2020-02-26 at 10:11 +0100, Miklos Szeredi wrote:
On Tue, Feb 25, 2020 at 4:29 PM James Bottomley
[off-list ref] wrote:
quoted
The other thing a file descriptor does that sysfs doesn't is that
it
solves the information leak: if I'm in a mount namespace that has
no
access to certain mounts, I can't fspick them and thus I can't see
the
information. By default, with sysfs I can.
That's true, but procfs/sysfs has to deal with various namespacing
issues anyway. If this is just about hiding a number of entries,
then
I don't think that's going to be a big deal.
I didn't see name space considerations in sysfs when I was looking at
it recently. Obeying name space requirements is likely a lot of work
in sysfs.
The syscall API is efficient: single syscall per query instead of
several, no parsing necessary.
However, it is difficult to extend, because the ABI must be updated,
possibly libc and util-linux also, so that scripts can also consume
the new parameter. With the sysfs approach only the kernel needs to
be updated, and possibly only the filesystem code, not even the VFS.
So I think the question comes down to: do we need a highly efficient
way to query the superblock parameters all at once, or not?
Or a similar question could be, how could a sysfs interface work
to provide mount information.
Getting information about all mounts might not be too bad but the
sysfs directory structure that would be needed to represent all
system mounts (without considering name spaces) would likely
result in somewhat busy user space code.
For example, given a path, and the path is all I know, how do I
get mount information?
Ignoring possible multiple mounts on a mount point, call fsinfo()
with the path and get the id (the path walk is low overhead) to
use with fsinfo() to get the all the info I need ... done.
Again, ignoring possible multiple mounts on a mount point, and
assuming there is a sysfs tree enumerating all the system mounts.
I could open <sysfs base> + mount point path followed buy opening
and reading the individual attribute files ... a bit more busy
that one ... particularly if I need to do it for several thousand
mounts.
Then there's the code that would need to be added to maintain the
various views in the sysfs tree, which can't be restricted only to
the VFS because there's file system specific info needed too (the
maintain a table idea), and that's before considering name space
handling changes to sysfs.
At the least the question of "do we need a highly efficient way
to query the superblock parameters all at once" needs to be
extended to include mount table enumeration as well as getting
the info.
But this is just me thinking about mount table handling and the
quite significant problem we now have with user space scanning
the proc mount tables to get this information.
Ian
On Thu, Feb 27, 2020 at 6:06 AM Ian Kent [off-list ref] wrote:
At the least the question of "do we need a highly efficient way
to query the superblock parameters all at once" needs to be
extended to include mount table enumeration as well as getting
the info.
But this is just me thinking about mount table handling and the
quite significant problem we now have with user space scanning
the proc mount tables to get this information.
Right.
So the problem is that currently autofs needs to rescan the proc mount
table on every change. The solution to that is to
- add a notification mechanism
- and a way to selectively query mount/superblock information
right?
For the notification we have uevents in sysfs, which also supplies the
changed parameters. Taking aside namespace issues and addressing
mounts would this work for autofs?
Thanks,
Miklos
From: Ian Kent <raven@themaw.net> Date: 2020-02-27 11:34:54
On Thu, 2020-02-27 at 10:36 +0100, Miklos Szeredi wrote:
On Thu, Feb 27, 2020 at 6:06 AM Ian Kent [off-list ref] wrote:
quoted
At the least the question of "do we need a highly efficient way
to query the superblock parameters all at once" needs to be
extended to include mount table enumeration as well as getting
the info.
But this is just me thinking about mount table handling and the
quite significant problem we now have with user space scanning
the proc mount tables to get this information.
Right.
So the problem is that currently autofs needs to rescan the proc
mount
table on every change. The solution to that is to
Actually no, that's not quite the problem I see.
autofs handles large mount tables fairly well (necessarily) and
in time I plan to remove the need to read the proc tables at all
(that's proven very difficult but I'll get back to that).
This has to be done to resolve the age old problem of autofs not
being able to handle large direct mount maps. But, because of
the large number of mounts associated with large direct mount
maps, other system processes are badly affected too.
So the problem I want to see fixed is the effect of very large
mount tables on other user space applications, particularly the
effect when a large number of mounts or umounts are performed.
Clearly large mount tables not only result from autofs and the
problems caused by them are slightly different to the mount and
umount problem I describe. But they are a problem nevertheless
in the sense that frequent notifications that lead to reading
a large proc mount table has significant overhead that can't be
avoided because the table may have changed since the last time
it was read.
It's easy to cause several system processes to peg a fair number
of CPU's when a large number of mounts/umounts are being performed,
namely systemd, udisks2 and a some others. Also I've seen couple
of application processes badly affected purely by the presence of
a large number of mounts in the proc tables, that's not quite so
bad though.
- add a notification mechanism - lookup a mount based on path
- and a way to selectively query mount/superblock information
based on path ...
right?
For the notification we have uevents in sysfs, which also supplies
the
changed parameters. Taking aside namespace issues and addressing
mounts would this work for autofs?
The parameters supplied by the notification mechanism are important.
The place this is needed will be libmount since it catches a broad
number of user space applications, including those I mentioned above
(well at least systemd, I think also udisks2, very probably others).
So that means mount table info. needs to be maintained, whether that
can be achieved using sysfs I don't know. Creating and maintaining
the sysfs tree would be a big challenge I think.
But before trying to work out how to use a notification mechanism
just having a way to get the info provided by the proc tables using
a path alone should give initial immediate improvement in libmount.
Ian
On Thu, Feb 27, 2020 at 12:34 PM Ian Kent [off-list ref] wrote:
On Thu, 2020-02-27 at 10:36 +0100, Miklos Szeredi wrote:
quoted
On Thu, Feb 27, 2020 at 6:06 AM Ian Kent [off-list ref] wrote:
quoted
At the least the question of "do we need a highly efficient way
to query the superblock parameters all at once" needs to be
extended to include mount table enumeration as well as getting
the info.
But this is just me thinking about mount table handling and the
quite significant problem we now have with user space scanning
the proc mount tables to get this information.
Right.
So the problem is that currently autofs needs to rescan the proc
mount
table on every change. The solution to that is to
Actually no, that's not quite the problem I see.
autofs handles large mount tables fairly well (necessarily) and
in time I plan to remove the need to read the proc tables at all
(that's proven very difficult but I'll get back to that).
This has to be done to resolve the age old problem of autofs not
being able to handle large direct mount maps. But, because of
the large number of mounts associated with large direct mount
maps, other system processes are badly affected too.
So the problem I want to see fixed is the effect of very large
mount tables on other user space applications, particularly the
effect when a large number of mounts or umounts are performed.
Clearly large mount tables not only result from autofs and the
problems caused by them are slightly different to the mount and
umount problem I describe. But they are a problem nevertheless
in the sense that frequent notifications that lead to reading
a large proc mount table has significant overhead that can't be
avoided because the table may have changed since the last time
it was read.
It's easy to cause several system processes to peg a fair number
of CPU's when a large number of mounts/umounts are being performed,
namely systemd, udisks2 and a some others. Also I've seen couple
of application processes badly affected purely by the presence of
a large number of mounts in the proc tables, that's not quite so
bad though.
quoted
- add a notification mechanism - lookup a mount based on path
- and a way to selectively query mount/superblock information
based on path ...
quoted
right?
For the notification we have uevents in sysfs, which also supplies
the
changed parameters. Taking aside namespace issues and addressing
mounts would this work for autofs?
The parameters supplied by the notification mechanism are important.
The place this is needed will be libmount since it catches a broad
number of user space applications, including those I mentioned above
(well at least systemd, I think also udisks2, very probably others).
So that means mount table info. needs to be maintained, whether that
can be achieved using sysfs I don't know. Creating and maintaining
the sysfs tree would be a big challenge I think.
But before trying to work out how to use a notification mechanism
just having a way to get the info provided by the proc tables using
a path alone should give initial immediate improvement in libmount.
Adding Karel, Lennart, Zbigniew and util-linux@vger...
At a quick glance at libmount and systemd code, it appears that just
switching out the implementation in libmount will not be enough:
systemd is calling functions like mnt_table_parse_*() when it receives
a notification that the mount table changed.
What is the end purpose of parsing the mount tables? Can systemd guys
comment on that?
Thanks,
Miklos
From: Karel Zak <kzak@redhat.com> Date: 2020-02-27 15:14:37
On Thu, Feb 27, 2020 at 02:45:27PM +0100, Miklos Szeredi wrote:
quoted
So the problem I want to see fixed is the effect of very large
mount tables on other user space applications, particularly the
effect when a large number of mounts or umounts are performed.
Yes, now you have to generate (in kernel) and parse (in
userspace) all mount table to get information about just
one mount table entry. This is typical for umount or systemd.
quoted
quoted
- add a notification mechanism - lookup a mount based on path
- and a way to selectively query mount/superblock information
based on path ...
For umount-like use-cases we need mountpoint/ to mount entry
conversion; I guess something like open(mountpoint/) + fsinfo()
should be good enough.
For systemd we need the same, but triggered by notification. The ideal
solution is to get mount entry ID or FD from notification and later use this
ID or FD to ask for details about the mount entry (probably again fsinfo()).
The notification has to be usable with in epoll() set.
This solves 99% of our performance issues I guess.
quoted
So that means mount table info. needs to be maintained, whether that
can be achieved using sysfs I don't know. Creating and maintaining
the sysfs tree would be a big challenge I think.
It will be still necessary to get complete mount table sometimes, but
not in performance sensitive scenarios.
I'm not sure about sysfs/, you need somehow resolve namespaces, order
of the mount entries (which one is the last one), etc. IMHO translate
mountpoint path to sysfs/ path will be complicated.
quoted
But before trying to work out how to use a notification mechanism
just having a way to get the info provided by the proc tables using
a path alone should give initial immediate improvement in libmount.
Adding Karel, Lennart, Zbigniew and util-linux@vger...
At a quick glance at libmount and systemd code, it appears that just
switching out the implementation in libmount will not be enough:
systemd is calling functions like mnt_table_parse_*() when it receives
a notification that the mount table changed.
We're ready to change this stuff in systemd if there will be something
better (something per-mount-entry).
My plan is add new API to libmount to query information about one
mount entry (but I had no time to play with fsinfo yet).
What is the end purpose of parsing the mount tables? Can systemd guys
comment on that?
If mount/umount is triggered by systemd than it need verification
about success and final version of the mount options. It also reads
information from libmount to get userspace mount options (.e.g.
_netdev -- libmount uses mount source, target and fsroot to join
kernel and userpace stuff).
And don't forget that mount units are part of systemd dependencies, so
umount/mount is important event for systemd and it need details about
the changes (what, where, ... etc.)
Karel
--
Karel Zak [off-list ref]
http://karelzak.blogspot.com
From: Ian Kent <raven@themaw.net> Date: 2020-02-28 00:13:03
On Thu, 2020-02-27 at 14:45 +0100, Miklos Szeredi wrote:
On Thu, Feb 27, 2020 at 12:34 PM Ian Kent [off-list ref] wrote:
quoted
On Thu, 2020-02-27 at 10:36 +0100, Miklos Szeredi wrote:
quoted
On Thu, Feb 27, 2020 at 6:06 AM Ian Kent [off-list ref]
wrote:
quoted
At the least the question of "do we need a highly efficient way
to query the superblock parameters all at once" needs to be
extended to include mount table enumeration as well as getting
the info.
But this is just me thinking about mount table handling and the
quite significant problem we now have with user space scanning
the proc mount tables to get this information.
Right.
So the problem is that currently autofs needs to rescan the proc
mount
table on every change. The solution to that is to
Actually no, that's not quite the problem I see.
autofs handles large mount tables fairly well (necessarily) and
in time I plan to remove the need to read the proc tables at all
(that's proven very difficult but I'll get back to that).
This has to be done to resolve the age old problem of autofs not
being able to handle large direct mount maps. But, because of
the large number of mounts associated with large direct mount
maps, other system processes are badly affected too.
So the problem I want to see fixed is the effect of very large
mount tables on other user space applications, particularly the
effect when a large number of mounts or umounts are performed.
Clearly large mount tables not only result from autofs and the
problems caused by them are slightly different to the mount and
umount problem I describe. But they are a problem nevertheless
in the sense that frequent notifications that lead to reading
a large proc mount table has significant overhead that can't be
avoided because the table may have changed since the last time
it was read.
It's easy to cause several system processes to peg a fair number
of CPU's when a large number of mounts/umounts are being performed,
namely systemd, udisks2 and a some others. Also I've seen couple
of application processes badly affected purely by the presence of
a large number of mounts in the proc tables, that's not quite so
bad though.
quoted
- add a notification mechanism - lookup a mount based on path
- and a way to selectively query mount/superblock information
based on path ...
quoted
right?
For the notification we have uevents in sysfs, which also
supplies
the
changed parameters. Taking aside namespace issues and addressing
mounts would this work for autofs?
The parameters supplied by the notification mechanism are
important.
The place this is needed will be libmount since it catches a broad
number of user space applications, including those I mentioned
above
(well at least systemd, I think also udisks2, very probably
others).
So that means mount table info. needs to be maintained, whether
that
can be achieved using sysfs I don't know. Creating and maintaining
the sysfs tree would be a big challenge I think.
But before trying to work out how to use a notification mechanism
just having a way to get the info provided by the proc tables using
a path alone should give initial immediate improvement in libmount.
Adding Karel, Lennart, Zbigniew and util-linux@vger...
At a quick glance at libmount and systemd code, it appears that just
switching out the implementation in libmount will not be enough:
systemd is calling functions like mnt_table_parse_*() when it
receives
a notification that the mount table changed.
Maybe I wasn't clear, my bad, sorry about that.
There's no question that change notification handling is needed too.
I'm claiming that an initial change to use something that can get
the mount information without using the proc tables alone will give
an "initial immediate improvement".
The work needed to implement mount table change notification
handling will take much more time and exactly what changes that
will bring is not clear yet and I do plan to work on that too,
together with Karel.
Ian
From: Ian Kent <raven@themaw.net> Date: 2020-02-28 00:43:37
On Thu, 2020-02-27 at 16:14 +0100, Karel Zak wrote:
On Thu, Feb 27, 2020 at 02:45:27PM +0100, Miklos Szeredi wrote:
quoted
quoted
So the problem I want to see fixed is the effect of very large
mount tables on other user space applications, particularly the
effect when a large number of mounts or umounts are performed.
Yes, now you have to generate (in kernel) and parse (in
userspace) all mount table to get information about just
one mount table entry. This is typical for umount or systemd.
quoted
quoted
quoted
- add a notification mechanism - lookup a mount based on
path
- and a way to selectively query mount/superblock information
based on path ...
For umount-like use-cases we need mountpoint/ to mount entry
conversion; I guess something like open(mountpoint/) + fsinfo()
should be good enough.
For systemd we need the same, but triggered by notification. The
ideal
solution is to get mount entry ID or FD from notification and later
use this
ID or FD to ask for details about the mount entry (probably again
fsinfo()).
The notification has to be usable with in epoll() set.
This solves 99% of our performance issues I guess.
quoted
quoted
So that means mount table info. needs to be maintained, whether
that
can be achieved using sysfs I don't know. Creating and
maintaining
the sysfs tree would be a big challenge I think.
It will be still necessary to get complete mount table sometimes,
but
not in performance sensitive scenarios.
That was my understanding too.
Mount table enumeration is possible with fsinfo() but you still
have to handle each and every mount so improvement there is not
going to be as much as cases where the proc mount table needs to
be scanned independently for an individual mount. It will be
somewhat more straight forward without the need to dissect text
records though.
I'm not sure about sysfs/, you need somehow resolve namespaces, order
of the mount entries (which one is the last one), etc. IMHO translate
mountpoint path to sysfs/ path will be complicated.
I wonder about that too, after all sysfs contains a tree of nodes
from which the view is created unlike proc which translates kernel
information directly based on what the process should see.
We'll need to wait a bit and see what Miklos has in mind for mount
table enumeration and nothing has been said about name spaces yet.
While fsinfo() is not similar to proc it does handle name spaces
in a sensible way via. file handles, a bit similar to the proc fs,
and ordering is catered for in the fsinfo() enumeration in a natural
way. Not sure how that would be handled using sysfs ...
Ian
On Fri, Feb 28, 2020 at 1:43 AM Ian Kent [off-list ref] wrote:
quoted
I'm not sure about sysfs/, you need somehow resolve namespaces, order
of the mount entries (which one is the last one), etc. IMHO translate
mountpoint path to sysfs/ path will be complicated.
I wonder about that too, after all sysfs contains a tree of nodes
from which the view is created unlike proc which translates kernel
information directly based on what the process should see.
We'll need to wait a bit and see what Miklos has in mind for mount
table enumeration and nothing has been said about name spaces yet.
Adding Greg for sysfs knowledge.
As far as I understand the sysfs model is, basically:
- list of devices sorted by class and address
- with each class having a given set of attributes
Superblocks and mounts could get enumerated by a unique identifier.
mnt_id seems to be good for mounts, s_dev may or may not be good for
superblock, but s_id (as introduced in this patchset) could be used
instead.
As for namespaces, that's "just" an access control issue, AFAICS.
For example a task with a non-initial mount namespace should not have
access to attributes of mounts outside of its namespace. Checking
access to superblock attributes would be similar: scan the list of
mounts and only allow access if at least one mount would get access.
While fsinfo() is not similar to proc it does handle name spaces
in a sensible way via. file handles, a bit similar to the proc fs,
and ordering is catered for in the fsinfo() enumeration in a natural
way. Not sure how that would be handled using sysfs ...
I agree that the access control is much more straightforward with
fsinfo(2) and this may be the single biggest reason to introduce a new
syscall.
Let's see what others thing.
Thanks,
Miklos
On Fri, Feb 28, 2020 at 09:35:17AM +0100, Miklos Szeredi wrote:
On Fri, Feb 28, 2020 at 1:43 AM Ian Kent [off-list ref] wrote:
quoted
quoted
I'm not sure about sysfs/, you need somehow resolve namespaces, order
of the mount entries (which one is the last one), etc. IMHO translate
mountpoint path to sysfs/ path will be complicated.
I wonder about that too, after all sysfs contains a tree of nodes
from which the view is created unlike proc which translates kernel
information directly based on what the process should see.
We'll need to wait a bit and see what Miklos has in mind for mount
table enumeration and nothing has been said about name spaces yet.
Adding Greg for sysfs knowledge.
As far as I understand the sysfs model is, basically:
- list of devices sorted by class and address
- with each class having a given set of attributes
Close enough :)
Superblocks and mounts could get enumerated by a unique identifier.
mnt_id seems to be good for mounts, s_dev may or may not be good for
superblock, but s_id (as introduced in this patchset) could be used
instead.
So what would the sysfs tree look like with this?
As for namespaces, that's "just" an access control issue, AFAICS.
For example a task with a non-initial mount namespace should not have
access to attributes of mounts outside of its namespace. Checking
access to superblock attributes would be similar: scan the list of
mounts and only allow access if at least one mount would get access.
sysfs does handle namespaces, look at how networking does this. But,
it's not exactly the simplest thing to do so, so be careful with that as
this is going to be essential for this type of work.
thanks,
greg k-h
From: David Howells <dhowells@redhat.com> Date: 2020-02-28 14:44:53
Aleksa Sarai [off-list ref] wrote:
quoted
If params is given, all of params->__reserved[] must be 0.
I would suggest that rather than having a reserved field for future
extensions, you make use of copy_struct_from_user() and have extensible
structs:
Yeah. I seem to recall that special support was required for 6-arg syscalls
on some arches, though I could move the dfd argument into the parameter block
and make AT_FDCWD the default.
I dropped the "const" on fsinfo_params because the planned CHECK_FiELDS
feature for extensible-struct syscalls requires writing to the struct.
Ummm... Why? You shouldn't be trying to alter the parameters structure. It
could feasibly be stored static const in userspace (though I'm not sure how
likely it would be that someone would do that).
I also switched the flags field to u64 because CHECK_FiELDS is intended to
use (1<<63) for all syscalls (this has the nice benefit of removing the need
of a padding field entirely).
struct fsinfo_params {
__u32 flags;
__u32 at_flags;
__u32 request;
__u32 Nth;
__u32 Mth;
};
What padding? ;-)
Though possibly the struct does need forcing to 64-bit alignment for future
expansion.
quoted
dfd, filename and params->at_flags indicate the file to query. There is no
equivalent of lstat() as that can be emulated with fsinfo() by setting
AT_SYMLINK_NOFOLLOW in params->at_flags.
Minor gripe -- can we make the default be AT_SYMLINK_NOFOLLOW and you
need to explicitly pass AT_SYMLINK_FOLLOW? Accidentally following
symlinks is a constant source of security bugs.
Someone else has said that all new syscalls should be using RESOLVE_* flags in
preference to AT_* flags (even though RESOLVE_* flags are not a superset of
AT_* flags and appear to be in a header named specifically for the openat2()
syscall, not generic).
I'm not sure who authored openat2.h, but they went with a RESOLVE_NO_SYMLINKS
rather than a RESOLVE_SYMLINKS ;-)
quoted
There is also no equivalent of fstat() as that can be emulated by
passing a NULL filename to fsinfo() with the fd of interest in dfd.
Presumably you also need to pass AT_EMPTY_PATH?
Actually, you need to set FSINFO_FLAGS_QUERY_FD in fsinfo_params::flags. I
need to update the description for this.
Sounds good, though I think we should zero-fill the tail end of the
buffer (if the buffer is larger than the in-kernel one).
I do that. I should make it clearer in the patch description.
David
From: James Bottomley <James.Bottomley@HansenPartnership.com> Date: 2020-02-28 15:09:24
On Fri, 2020-02-28 at 09:35 +0100, Miklos Szeredi wrote:
On Fri, Feb 28, 2020 at 1:43 AM Ian Kent [off-list ref] wrote:
quoted
quoted
I'm not sure about sysfs/, you need somehow resolve namespaces,
order of the mount entries (which one is the last one), etc. IMHO
translate mountpoint path to sysfs/ path will be complicated.
I wonder about that too, after all sysfs contains a tree of nodes
from which the view is created unlike proc which translates kernel
information directly based on what the process should see.
We'll need to wait a bit and see what Miklos has in mind for mount
table enumeration and nothing has been said about name spaces yet.
Adding Greg for sysfs knowledge.
As far as I understand the sysfs model is, basically:
- list of devices sorted by class and address
- with each class having a given set of attributes
Superblocks and mounts could get enumerated by a unique identifier.
mnt_id seems to be good for mounts, s_dev may or may not be good for
superblock, but s_id (as introduced in this patchset) could be used
instead.
As for namespaces, that's "just" an access control issue, AFAICS.
That's an easy thing to say but not an easy thing to check: it can be
made so for label based namespaces like the network, but the mount
namespace is shared/cloned tree based. Assessing whether a given
superblock is within your current namespace root can become a large
search exercise. You can see how much of one in fs/proc_namespaces.c
which controls how /proc/self/mounts appears in your current namespace.
For example a task with a non-initial mount namespace should not have
access to attributes of mounts outside of its namespace. Checking
access to superblock attributes would be similar: scan the list of
mounts and only allow access if at least one mount would get access.
That scan can be expensive as I explained above. That's really why I
think this is a bad idea. Sysfs itself is nicely currently restricted
to system information that most containers don't need to know, so a lot
of the sysfs issues with containers can be solved by not mounting it.
If you suddenly make it required for filesystem information and
notifications, that security measure gets blown out of the water.
quoted
While fsinfo() is not similar to proc it does handle name spaces
in a sensible way via. file handles, a bit similar to the proc fs,
and ordering is catered for in the fsinfo() enumeration in a
natural way. Not sure how that would be handled using sysfs ...
I agree that the access control is much more straightforward with
fsinfo(2) and this may be the single biggest reason to introduce a
new syscall.
Let's see what others thing.
Containers are file based entities, so file descriptors are their most
natural thing and they have full ACL protection within the container
(can't open the file, can't then get the fd). The other reason
container people like file descriptors (all the Xat system calls that
have been introduced) is that if we do actually need to break the
boundaries or privileges of the container, we can do so by getting the
orchestration system to pass in a fd the interior of the container
wouldn't have access to.
James
On Fri, Feb 28, 2020 at 4:09 PM James Bottomley
[off-list ref] wrote:
Containers are file based entities, so file descriptors are their most
natural thing and they have full ACL protection within the container
(can't open the file, can't then get the fd). The other reason
container people like file descriptors (all the Xat system calls that
have been introduced) is that if we do actually need to break the
boundaries or privileges of the container, we can do so by getting the
orchestration system to pass in a fd the interior of the container
wouldn't have access to.
Yeah, agreed about the simplicity of fd based access. Then again a
filesystem access would allow immediate access to all scripts,
languages, etc. That, I think is a huge bonus compared to the
ioctl-like mess that the current proposal is, which would require
library, utility, language binding updates on all changes. Ugh.
One way to resolve that is to have the mount information
magic-symlinked from /proc/PID/fdmount/FD directly to the mountinfo
dir, which would then have a link into the sbinfo dir. With other
access denied to all except sysadmin.
Would that work?
Thanks,
Miklos
From: Christian Brauner <hidden> Date: 2020-02-28 15:53:02
On Tue, Feb 25, 2020 at 07:28:55AM -0800, James Bottomley wrote:
On Tue, 2020-02-25 at 12:13 +0000, Steven Whitehouse wrote:
quoted
Hi,
On 24/02/2020 15:28, Miklos Szeredi wrote:
quoted
On Mon, Feb 24, 2020 at 3:55 PM James Bottomley
[off-list ref] wrote:
quoted
Once it's table driven, certainly a sysfs directory becomes
possible. The problem with ST_DEV is filesystems like btrfs and
xfs that may have multiple devices.
For XFS there's always a single sb->s_dev though, that's what
st_dev will be set to on all files.
Btrfs subvolume is sort of a lightweight superblock, so basically
all such st_dev's are aliases of the same master superblock. So
lookup of all subvolume st_dev's could result in referencing the
same underlying struct super_block (just like /proc/$PID will
reference the same underlying task group regardless of which of the
task group member's PID is used).
Having this info in sysfs would spare us a number of issues that a
set of new syscalls would bring. The question is, would that be
enough, or is there a reason that sysfs can't be used to present
the various filesystem related information that fsinfo is supposed
to present?
Thanks,
Miklos
We need a unique id for superblocks anyway. I had wondered about
using s_dev some time back, but for the reasons mentioned earlier in
this thread I think it might just land up being confusing and
difficult to manage. While fake s_devs are created for sbs that don't
have a device, I can't help thinking that something closer to
ifindex, but for superblocks, is needed here. That would avoid the
issue of which device number to use.
In fact we need that anyway for the notifications, since without
that there is a race that can lead to missing remounts of the same
device, in case a umount/mount pair is missed due to an overrun, and
then fsinfo returns the same device as before, with potentially the
same mount options too. So I think a unique id for a superblock is a
generically useful feature, which would also allow for sensible sysfs
directory naming, if required,
But would this be informative and useful for the user? I'm sure we can
find a persistent id for a persistent superblock, but what about tmpfs
... that's going to have to change with every reboot. It's going to be
remarkably inconvenient if I want to get fsinfo on /run to have to keep
finding what the id is.
The other thing a file descriptor does that sysfs doesn't is that it
solves the information leak: if I'm in a mount namespace that has no
access to certain mounts, I can't fspick them and thus I can't see the
information. By default, with sysfs I can.
Difficult to figure out which part of the thread to reply too. :)
sysfs strikes me as fundamentally misguided for this task.
Init systems or any large-scale daemon will hate parsing things, there's
that and parts of the reason why mountinfo sucks is because of parsing a
possibly a potentially enormous file. Exposing information in sysfs will
require parsing again one way or the other. I've been discussing these
bottlenecks with Lennart quite a bit and reliable and performant mount
notifications without needing to parse stuff is very high on the issue
list. But even if that isn't an issue for some reason the namespace
aspect is definitely something I'd consider a no-go.
James has been poking at this a little already and I agree. More
specifically, sysfs and proc already are a security nightmare for
namespace-aware workloads and require special care. Not leaking
information in any way is a difficult task. I mean, over the last two
years I sent quite a lot of patches to the networking-namespace aware
part of sysfs alone either fixing information leaks, or making other
parts namespace aware that weren't and were causing issues (There's
another large-ish series sitting in Dave's tree right now.). And tbh,
network namespacing in sysfs is imho trivial compared to what we would
need to do to handle mount namespacing and especially mount propagation.
fsinfo() is way cleaner and ultimately simpler approach. We very much
want it file-descriptor based. The mount api opens up the road to secure
and _delegatable_ querying of filesystem information.
Christian
On Fri, Feb 28, 2020 at 1:27 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
Superblocks and mounts could get enumerated by a unique identifier.
mnt_id seems to be good for mounts, s_dev may or may not be good for
superblock, but s_id (as introduced in this patchset) could be used
instead.
So what would the sysfs tree look like with this?
For a start something like this:
mounts/$MOUNT_ID/
parent -> ../$PARENT_ID
super -> ../../supers/$SUPER_ID
root: path from mount root to fs root (could be optional as usually
they are the same)
mountpoint -> $MOUNTPOINT
flags: mount flags
propagation: mount propagation
children/$CHILD_ID -> ../../$CHILD_ID
supers/$SUPER_ID/
type: fstype
source: mount source (devname)
options: csv of mount options
Thanks,
Miklos
From: David Howells <dhowells@redhat.com> Date: 2020-02-28 16:36:21
sysfs also has some other disadvantages for this:
(1) There's a potential chicken-and-egg problem in that you have to create a
bunch of files and dirs in sysfs for every created mount and superblock
(possibly excluding special ones like the socket mount) - but this
includes sysfs itself. This might work - provided you create sysfs
first.
(2) sysfs is memory intensive. The directory structure has to be backed by
dentries and inodes that linger as long as the referenced object does
(procfs is more efficient in this regard for files that aren't being
accessed).
(3) It gives people extra, indirect ways to pin mount objects and
superblocks.
For the moment, fsinfo() gives you three ways of referring to a filesystem
object:
(a) Directly by path.
(b) By path associated with an fd.
(c) By mount ID (perm checked by working back up the tree).
but will need to add:
(d) By fscontext fd (which is hard to find in sysfs). Indeed, the superblock
may not even exist yet.
David
From: David Howells <dhowells@redhat.com> Date: 2020-02-28 16:42:39
Miklos Szeredi [off-list ref] wrote:
children/$CHILD_ID -> ../../$CHILD_ID
This would really suck. This bit would particularly affect rescanning time.
You also really want to read the entire child set atomically and, ideally,
include notification counters.
supers/$SUPER_ID/
type: fstype
source: mount source (devname)
options: csv of mount options
There's a lot more to fsinfo() than just this lot - and there's the
possibility that some of the values may change depending on exactly which file
you're looking at.
David
From: Al Viro <viro@zeniv.linux.org.uk> Date: 2020-02-28 17:15:22
On Fri, Feb 28, 2020 at 05:24:23PM +0100, Miklos Szeredi wrote:
On Fri, Feb 28, 2020 at 1:27 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
quoted
Superblocks and mounts could get enumerated by a unique identifier.
mnt_id seems to be good for mounts, s_dev may or may not be good for
superblock, but s_id (as introduced in this patchset) could be used
instead.
So what would the sysfs tree look like with this?
For a start something like this:
mounts/$MOUNT_ID/
parent -> ../$PARENT_ID
super -> ../../supers/$SUPER_ID
root: path from mount root to fs root (could be optional as usually
they are the same)
mountpoint -> $MOUNTPOINT
flags: mount flags
propagation: mount propagation
children/$CHILD_ID -> ../../$CHILD_ID
supers/$SUPER_ID/
type: fstype
source: mount source (devname)
options: csv of mount options
Oh, wonderful. So let me see if I got it right - any namespace operation
can create/destroy/move around an arbitrary amount of sysfs objects.
Better yet, we suddenly have to express the lifetime rules for struct mount
and struct superblock in terms of struct device garbage.
I'm less than thrilled by the entire fsinfo circus, but this really takes
the cake.
In case it needs to be spelled out: NAK.
On Fri, Feb 28, 2020 at 6:15 PM Al Viro [off-list ref] wrote:
On Fri, Feb 28, 2020 at 05:24:23PM +0100, Miklos Szeredi wrote:
quoted
On Fri, Feb 28, 2020 at 1:27 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
quoted
Superblocks and mounts could get enumerated by a unique identifier.
mnt_id seems to be good for mounts, s_dev may or may not be good for
superblock, but s_id (as introduced in this patchset) could be used
instead.
So what would the sysfs tree look like with this?
For a start something like this:
mounts/$MOUNT_ID/
parent -> ../$PARENT_ID
super -> ../../supers/$SUPER_ID
root: path from mount root to fs root (could be optional as usually
they are the same)
mountpoint -> $MOUNTPOINT
flags: mount flags
propagation: mount propagation
children/$CHILD_ID -> ../../$CHILD_ID
supers/$SUPER_ID/
type: fstype
source: mount source (devname)
options: csv of mount options
Oh, wonderful. So let me see if I got it right - any namespace operation
can create/destroy/move around an arbitrary amount of sysfs objects.
Parent/children symlinks may be excessive...
Better yet, we suddenly have to express the lifetime rules for struct mount
and struct superblock in terms of struct device garbage.
How so? struct mount and struct superblock would hold a ref on
struct device, not the other way round.
In any case, I'm not insistent on the use of sysfs device classes for
this; struct device (488B) does seem too heavy for struct mount
(328B).
What I'm pretty sure about is that a read(2) based interface would be
way more useful than the syscall multiplexer that the current proposal
is.
Thanks,
Miklos
On Fri, Feb 28, 2020 at 5:36 PM David Howells [off-list ref] wrote:
sysfs also has some other disadvantages for this:
(1) There's a potential chicken-and-egg problem in that you have to create a
bunch of files and dirs in sysfs for every created mount and superblock
(possibly excluding special ones like the socket mount) - but this
includes sysfs itself. This might work - provided you create sysfs
first.
Sysfs architecture looks something like this (I hope Greg will correct
me if I'm wrong):
device driver -> kobj tree <- sysfs tree
The kobj tree is created by the device driver, and the dentry tree is
created on demand from the kobj tree. Lifetime of kobjs is bound to
both the sysfs objects and the device but not the other way round.
I.e. device can go away while the sysfs object is still being
referenced, and sysfs can be freely mounted and unmounted
independently of device initialization.
So there's no ordering requirement between sysfs mounts and other
mounts. I might be wrong on the details, since mounts are created
very early in the boot process...
(2) sysfs is memory intensive. The directory structure has to be backed by
dentries and inodes that linger as long as the referenced object does
(procfs is more efficient in this regard for files that aren't being
accessed)
See above: I don't think dentries and inodes are pinned, only kobjs
and their associated cruft. Which may be too heavy, depending on the
details of the kobj tree.
(3) It gives people extra, indirect ways to pin mount objects and
superblocks.
See above.
For the moment, fsinfo() gives you three ways of referring to a filesystem
object:
(a) Directly by path.
A path is always representable by an O_PATH descriptor.
(b) By path associated with an fd.
See my proposal about linking from /proc/$PID/fdmount/$FD ->
/sys/devices/virtual/mounts/$MOUNT_ID.
(c) By mount ID (perm checked by working back up the tree).
Check that perm on lookup of /sys/devices/virtual/mounts/$MOUNT_ID.
The proc symlink would bypass the lookup check by directly jumping to
the mountinfo dir.
but will need to add:
(d) By fscontext fd (which is hard to find in sysfs). Indeed, the superblock
may not even exist yet.
Proc symlink would work for that too.
If sysfs is too heavy, this could be proc or a completely new
filesystem. The implementation is much less relevant at this stage of
the discussion than the interface.
Thanks,
Miklos
On Mon, Mar 02, 2020 at 10:09:51AM +0100, Miklos Szeredi wrote:
On Fri, Feb 28, 2020 at 5:36 PM David Howells [off-list ref] wrote:
quoted
sysfs also has some other disadvantages for this:
(1) There's a potential chicken-and-egg problem in that you have to create a
bunch of files and dirs in sysfs for every created mount and superblock
(possibly excluding special ones like the socket mount) - but this
includes sysfs itself. This might work - provided you create sysfs
first.
Sysfs architecture looks something like this (I hope Greg will correct
me if I'm wrong):
device driver -> kobj tree <- sysfs tree
The kobj tree is created by the device driver, and the dentry tree is
created on demand from the kobj tree. Lifetime of kobjs is bound to
both the sysfs objects and the device but not the other way round.
I.e. device can go away while the sysfs object is still being
referenced, and sysfs can be freely mounted and unmounted
independently of device initialization.
So there's no ordering requirement between sysfs mounts and other
mounts. I might be wrong on the details, since mounts are created
very early in the boot process...
quoted
(2) sysfs is memory intensive. The directory structure has to be backed by
dentries and inodes that linger as long as the referenced object does
(procfs is more efficient in this regard for files that aren't being
accessed)
See above: I don't think dentries and inodes are pinned, only kobjs
and their associated cruft. Which may be too heavy, depending on the
details of the kobj tree.
That is correct, they should not be pinned, that is what kernfs handles
and why we can handle 30k virtual block devices on a 31bit s390 instance
:)
So you shouldn't have to worry about memory for sysfs.
There are loads of other reasons probably not to use sysfs for this
instead :)
thanks,
greg k-h
From: Karel Zak <kzak@redhat.com> Date: 2020-03-02 10:34:18
On Fri, Feb 28, 2020 at 05:24:23PM +0100, Miklos Szeredi wrote:
ned-By: MIMEDefang 2.78 on 10.11.54.4
On Fri, Feb 28, 2020 at 1:27 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
quoted
Superblocks and mounts could get enumerated by a unique identifier.
mnt_id seems to be good for mounts, s_dev may or may not be good for
superblock, but s_id (as introduced in this patchset) could be used
instead.
So what would the sysfs tree look like with this?
For a start something like this:
mounts/$MOUNT_ID/
parent -> ../$PARENT_ID
super -> ../../supers/$SUPER_ID
root: path from mount root to fs root (could be optional as usually
they are the same)
mountpoint -> $MOUNTPOINT
flags: mount flags
propagation: mount propagation
children/$CHILD_ID -> ../../$CHILD_ID
supers/$SUPER_ID/
type: fstype
source: mount source (devname)
options:
What about use-cases where I have no ID, but I have mountpoint path
(e.g. "umount /foo")? In this case I have to go to open() + fsinfo()
and then sysfs does not make sense for me, right?
Karel
--
Karel Zak [off-list ref]
http://karelzak.blogspot.com
From: Ian Kent <raven@themaw.net> Date: 2020-03-03 05:36:03
On Mon, 2020-03-02 at 10:09 +0100, Miklos Szeredi wrote:
On Fri, Feb 28, 2020 at 5:36 PM David Howells [off-list ref]
wrote:
quoted
sysfs also has some other disadvantages for this:
(1) There's a potential chicken-and-egg problem in that you have
to create a
bunch of files and dirs in sysfs for every created mount and
superblock
(possibly excluding special ones like the socket mount) - but
this
includes sysfs itself. This might work - provided you create
sysfs
first.
Sysfs architecture looks something like this (I hope Greg will
correct
me if I'm wrong):
device driver -> kobj tree <- sysfs tree
The kobj tree is created by the device driver, and the dentry tree is
created on demand from the kobj tree. Lifetime of kobjs is bound to
both the sysfs objects and the device but not the other way round.
I.e. device can go away while the sysfs object is still being
referenced, and sysfs can be freely mounted and unmounted
independently of device initialization.
So there's no ordering requirement between sysfs mounts and other
mounts. I might be wrong on the details, since mounts are created
very early in the boot process...
quoted
(2) sysfs is memory intensive. The directory structure has to be
backed by
dentries and inodes that linger as long as the referenced
object does
(procfs is more efficient in this regard for files that aren't
being
accessed)
See above: I don't think dentries and inodes are pinned, only kobjs
and their associated cruft. Which may be too heavy, depending on the
details of the kobj tree.
quoted
(3) It gives people extra, indirect ways to pin mount objects and
superblocks.
See above.
quoted
For the moment, fsinfo() gives you three ways of referring to a
filesystem
object:
(a) Directly by path.
A path is always representable by an O_PATH descriptor.
quoted
(b) By path associated with an fd.
See my proposal about linking from /proc/$PID/fdmount/$FD ->
/sys/devices/virtual/mounts/$MOUNT_ID.
quoted
(c) By mount ID (perm checked by working back up the tree).
Check that perm on lookup of /sys/devices/virtual/mounts/$MOUNT_ID.
The proc symlink would bypass the lookup check by directly jumping to
the mountinfo dir.
quoted
but will need to add:
(d) By fscontext fd (which is hard to find in sysfs). Indeed, the
superblock
may not even exist yet.
Proc symlink would work for that too.
There's mounts enumeration too, ordering is required to identify the
top (or bottom depending on terminology) with more than one mount on
a mount point.
If sysfs is too heavy, this could be proc or a completely new
filesystem. The implementation is much less relevant at this stage
of
the discussion than the interface.
Ha, proc with the seq file interface, that's already proved to not
work properly and looks difficult to fix.
Ian
On Tue, Mar 3, 2020 at 6:28 AM Ian Kent [off-list ref] wrote:
On Mon, 2020-03-02 at 10:09 +0100, Miklos Szeredi wrote:
quoted
On Fri, Feb 28, 2020 at 5:36 PM David Howells [off-list ref]
wrote:
quoted
sysfs also has some other disadvantages for this:
(1) There's a potential chicken-and-egg problem in that you have
to create a
bunch of files and dirs in sysfs for every created mount and
superblock
(possibly excluding special ones like the socket mount) - but
this
includes sysfs itself. This might work - provided you create
sysfs
first.
Sysfs architecture looks something like this (I hope Greg will
correct
me if I'm wrong):
device driver -> kobj tree <- sysfs tree
The kobj tree is created by the device driver, and the dentry tree is
created on demand from the kobj tree. Lifetime of kobjs is bound to
both the sysfs objects and the device but not the other way round.
I.e. device can go away while the sysfs object is still being
referenced, and sysfs can be freely mounted and unmounted
independently of device initialization.
So there's no ordering requirement between sysfs mounts and other
mounts. I might be wrong on the details, since mounts are created
very early in the boot process...
quoted
(2) sysfs is memory intensive. The directory structure has to be
backed by
dentries and inodes that linger as long as the referenced
object does
(procfs is more efficient in this regard for files that aren't
being
accessed)
See above: I don't think dentries and inodes are pinned, only kobjs
and their associated cruft. Which may be too heavy, depending on the
details of the kobj tree.
quoted
(3) It gives people extra, indirect ways to pin mount objects and
superblocks.
See above.
quoted
For the moment, fsinfo() gives you three ways of referring to a
filesystem
object:
(a) Directly by path.
A path is always representable by an O_PATH descriptor.
quoted
(b) By path associated with an fd.
See my proposal about linking from /proc/$PID/fdmount/$FD ->
/sys/devices/virtual/mounts/$MOUNT_ID.
quoted
(c) By mount ID (perm checked by working back up the tree).
Check that perm on lookup of /sys/devices/virtual/mounts/$MOUNT_ID.
The proc symlink would bypass the lookup check by directly jumping to
the mountinfo dir.
quoted
but will need to add:
(d) By fscontext fd (which is hard to find in sysfs). Indeed, the
superblock
may not even exist yet.
Proc symlink would work for that too.
There's mounts enumeration too, ordering is required to identify the
top (or bottom depending on terminology) with more than one mount on
a mount point.
quoted
If sysfs is too heavy, this could be proc or a completely new
filesystem. The implementation is much less relevant at this stage
of
the discussion than the interface.
Ha, proc with the seq file interface, that's already proved to not
work properly and looks difficult to fix.
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Thanks,
Miklos
From: David Howells <dhowells@redhat.com> Date: 2020-03-03 09:13:04
Miklos Szeredi [off-list ref] wrote:
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Don't forget the efficiency criterion. One reason for going with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in the system is
slow (not to mention the global lock that is held during the read).
Now, going with sysfs files on top of procfs links might avoid the global
lock, and you can avoid rereading the options string if you export a change
notification, but you're going to end up injecting a whole lot of pathwalk
latency into the system.
On top of that, it isn't going to help with the case that I'm working towards
implementing where a container manager can monitor for mounts taking place
inside the container and supervise them. What I'm proposing is that during
the action phase (eg. FSCONFIG_CMD_CREATE), fsconfig() would hand an fd
referring to the context under construction to the manager, which would then
be able to call fsinfo() to query it and fsconfig() to adjust it, reject it or
permit it. Something like:
fd = receive_context_to_supervise();
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_OPTIONS,
};
fsinfo(fd, NULL, ¶ms, sizeof(params), buffer, sizeof(buffer));
supervise_parameters(buffer);
fsconfig(fd, FSCONFIG_SET_FLAG, "hard", NULL, 0);
fsconfig(fd, FSCONFIG_SET_STRING, "vers", "4.2", 0);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_CREATE, NULL, NULL, 0);
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_NOTIFICATIONS,
};
struct fsinfo_sb_notifications sbnotify;
fsinfo(fd, NULL, ¶ms, sizeof(params), &sbnotify, sizeof(sbnotify));
watch_super(fd, "", AT_EMPTY_PATH, watch_fd, 0x03);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_PERMIT, NULL, NULL, 0);
close(fd);
However, the supervised mount may be happening in a completely different set
of namespaces, in which case the supervisor presumably wouldn't be able to see
the links in procfs and the relevant portions of sysfs.
David
On Tue, Mar 3, 2020 at 10:13 AM David Howells [off-list ref] wrote:
Miklos Szeredi [off-list ref] wrote:
quoted
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Don't forget the efficiency criterion. One reason for going with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in the system is
slow (not to mention the global lock that is held during the read).
Now, going with sysfs files on top of procfs links might avoid the global
lock, and you can avoid rereading the options string if you export a change
notification, but you're going to end up injecting a whole lot of pathwalk
latency into the system.
Completely irrelevant. Cached lookup is so much optimized, that you
won't be able to see any of it.
No, I don't think this is going to be a performance issue at all, but
if anything we could introduce a syscall
ssize_t readfile(int dfd, const char *path, char *buf, size_t
bufsize, int flags);
that is basically the equivalent of open + read + close, or even a
vectored variant that reads multiple files. But that's off topic
again, since I don't think there's going to be any performance issue
even with plain I/O syscalls.
On top of that, it isn't going to help with the case that I'm working towards
implementing where a container manager can monitor for mounts taking place
inside the container and supervise them. What I'm proposing is that during
the action phase (eg. FSCONFIG_CMD_CREATE), fsconfig() would hand an fd
referring to the context under construction to the manager, which would then
be able to call fsinfo() to query it and fsconfig() to adjust it, reject it or
permit it. Something like:
fd = receive_context_to_supervise();
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_OPTIONS,
};
fsinfo(fd, NULL, ¶ms, sizeof(params), buffer, sizeof(buffer));
supervise_parameters(buffer);
fsconfig(fd, FSCONFIG_SET_FLAG, "hard", NULL, 0);
fsconfig(fd, FSCONFIG_SET_STRING, "vers", "4.2", 0);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_CREATE, NULL, NULL, 0);
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_NOTIFICATIONS,
};
struct fsinfo_sb_notifications sbnotify;
fsinfo(fd, NULL, ¶ms, sizeof(params), &sbnotify, sizeof(sbnotify));
watch_super(fd, "", AT_EMPTY_PATH, watch_fd, 0x03);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_PERMIT, NULL, NULL, 0);
close(fd);
However, the supervised mount may be happening in a completely different set
of namespaces, in which case the supervisor presumably wouldn't be able to see
the links in procfs and the relevant portions of sysfs.
It would be a "jump" link to the otherwise invisible directory.
Thanks,
Miklos
On Tue, Mar 3, 2020 at 10:26 AM Miklos Szeredi [off-list ref] wrote:
On Tue, Mar 3, 2020 at 10:13 AM David Howells [off-list ref] wrote:
quoted
Miklos Szeredi [off-list ref] wrote:
quoted
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Don't forget the efficiency criterion. One reason for going with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in the system is
slow (not to mention the global lock that is held during the read).
BTW, I do feel that there's room for improvement in userspace code as
well. Even quite big mount table could be scanned for *changes* very
efficiently. l.e. cache previous contents of /proc/self/mountinfo and
compare with new contents, line-by-line. Only need to parse the
changed/added/removed lines.
Also it would be pretty easy to throttle the number of updates so
systemd et al. wouldn't hog the system with unnecessary processing.
Thanks,
Miklos
From: Christian Brauner <hidden> Date: 2020-03-03 10:01:02
On Tue, Mar 03, 2020 at 10:26:21AM +0100, Miklos Szeredi wrote:
On Tue, Mar 3, 2020 at 10:13 AM David Howells [off-list ref] wrote:
quoted
Miklos Szeredi [off-list ref] wrote:
quoted
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Don't forget the efficiency criterion. One reason for going with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in the system is
slow (not to mention the global lock that is held during the read).
Now, going with sysfs files on top of procfs links might avoid the global
lock, and you can avoid rereading the options string if you export a change
notification, but you're going to end up injecting a whole lot of pathwalk
latency into the system.
Completely irrelevant. Cached lookup is so much optimized, that you
won't be able to see any of it.
No, I don't think this is going to be a performance issue at all, but
if anything we could introduce a syscall
ssize_t readfile(int dfd, const char *path, char *buf, size_t
bufsize, int flags);
that is basically the equivalent of open + read + close, or even a
vectored variant that reads multiple files. But that's off topic
again, since I don't think there's going to be any performance issue
even with plain I/O syscalls.
quoted
On top of that, it isn't going to help with the case that I'm working towards
implementing where a container manager can monitor for mounts taking place
inside the container and supervise them. What I'm proposing is that during
the action phase (eg. FSCONFIG_CMD_CREATE), fsconfig() would hand an fd
referring to the context under construction to the manager, which would then
be able to call fsinfo() to query it and fsconfig() to adjust it, reject it or
permit it. Something like:
fd = receive_context_to_supervise();
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_OPTIONS,
};
fsinfo(fd, NULL, ¶ms, sizeof(params), buffer, sizeof(buffer));
supervise_parameters(buffer);
fsconfig(fd, FSCONFIG_SET_FLAG, "hard", NULL, 0);
fsconfig(fd, FSCONFIG_SET_STRING, "vers", "4.2", 0);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_CREATE, NULL, NULL, 0);
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_NOTIFICATIONS,
};
struct fsinfo_sb_notifications sbnotify;
fsinfo(fd, NULL, ¶ms, sizeof(params), &sbnotify, sizeof(sbnotify));
watch_super(fd, "", AT_EMPTY_PATH, watch_fd, 0x03);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_PERMIT, NULL, NULL, 0);
close(fd);
However, the supervised mount may be happening in a completely different set
of namespaces, in which case the supervisor presumably wouldn't be able to see
the links in procfs and the relevant portions of sysfs.
It would be a "jump" link to the otherwise invisible directory.
More magic links to beam you around sounds like a bad idea. We had a
bunch of CVEs around them in containers and they were one of the major
reasons behind us pushing for openat2(). That's why it has a
RESOLVE_NO_MAGICLINKS flag.
Christian
On Tue, Mar 3, 2020 at 11:00 AM Christian Brauner
[off-list ref] wrote:
On Tue, Mar 03, 2020 at 10:26:21AM +0100, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 10:13 AM David Howells [off-list ref] wrote:
quoted
Miklos Szeredi [off-list ref] wrote:
quoted
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Don't forget the efficiency criterion. One reason for going with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in the system is
slow (not to mention the global lock that is held during the read).
Now, going with sysfs files on top of procfs links might avoid the global
lock, and you can avoid rereading the options string if you export a change
notification, but you're going to end up injecting a whole lot of pathwalk
latency into the system.
Completely irrelevant. Cached lookup is so much optimized, that you
won't be able to see any of it.
No, I don't think this is going to be a performance issue at all, but
if anything we could introduce a syscall
ssize_t readfile(int dfd, const char *path, char *buf, size_t
bufsize, int flags);
that is basically the equivalent of open + read + close, or even a
vectored variant that reads multiple files. But that's off topic
again, since I don't think there's going to be any performance issue
even with plain I/O syscalls.
quoted
On top of that, it isn't going to help with the case that I'm working towards
implementing where a container manager can monitor for mounts taking place
inside the container and supervise them. What I'm proposing is that during
the action phase (eg. FSCONFIG_CMD_CREATE), fsconfig() would hand an fd
referring to the context under construction to the manager, which would then
be able to call fsinfo() to query it and fsconfig() to adjust it, reject it or
permit it. Something like:
fd = receive_context_to_supervise();
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_OPTIONS,
};
fsinfo(fd, NULL, ¶ms, sizeof(params), buffer, sizeof(buffer));
supervise_parameters(buffer);
fsconfig(fd, FSCONFIG_SET_FLAG, "hard", NULL, 0);
fsconfig(fd, FSCONFIG_SET_STRING, "vers", "4.2", 0);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_CREATE, NULL, NULL, 0);
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_NOTIFICATIONS,
};
struct fsinfo_sb_notifications sbnotify;
fsinfo(fd, NULL, ¶ms, sizeof(params), &sbnotify, sizeof(sbnotify));
watch_super(fd, "", AT_EMPTY_PATH, watch_fd, 0x03);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_PERMIT, NULL, NULL, 0);
close(fd);
However, the supervised mount may be happening in a completely different set
of namespaces, in which case the supervisor presumably wouldn't be able to see
the links in procfs and the relevant portions of sysfs.
It would be a "jump" link to the otherwise invisible directory.
More magic links to beam you around sounds like a bad idea. We had a
bunch of CVEs around them in containers and they were one of the major
reasons behind us pushing for openat2(). That's why it has a
RESOLVE_NO_MAGICLINKS flag.
No, that link wouldn't beam you around at all, it would end up in an
internally mounted instance of a mountfs, a safe place where no
dangerous CVE's roam.
Thanks,
Miklos
From: Steven Whitehouse <hidden> Date: 2020-03-03 10:22:17
Hi,
On 03/03/2020 09:48, Miklos Szeredi wrote:
On Tue, Mar 3, 2020 at 10:26 AM Miklos Szeredi [off-list ref] wrote:
quoted
On Tue, Mar 3, 2020 at 10:13 AM David Howells [off-list ref] wrote:
quoted
Miklos Szeredi [off-list ref] wrote:
quoted
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Don't forget the efficiency criterion. One reason for going with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in the system is
slow (not to mention the global lock that is held during the read).
BTW, I do feel that there's room for improvement in userspace code as
well. Even quite big mount table could be scanned for *changes* very
efficiently. l.e. cache previous contents of /proc/self/mountinfo and
compare with new contents, line-by-line. Only need to parse the
changed/added/removed lines.
Also it would be pretty easy to throttle the number of updates so
systemd et al. wouldn't hog the system with unnecessary processing.
Thanks,
Miklos
At least having patches to compare would allow us to look at the performance here and gain some numbers, which would be helpful to frame the discussions. However I'm not seeing how it would be easy to throttle updates... they occur at whatever rate they are generated and this can be fairly high. Also I'm not sure that I follow how the notifications and the dumping of the whole table are synchronized in this case, either.
Al has pointed out before that a single mount operation on a subtree can generate a large number of changes on that subtree. That kind of scenario will need to be dealt with efficiently so that we don't miss things, and we also minimize the possibility of overruns, and additional overhead on the mount changes themselves, by keeping the notification messages small.
We should also look at what the likely worst case might be. I seem to remember from what Ian has said in the past that there can be tens of thousands of autofs mounts on some large systems. I assume that worst case might be something like that, but multiplied by however many containers might be on a system. Can anybody think of a situation which might require even more mounts?
The network subsystem had a similar problem... they use rtnetlink for the routing information, and just like the proposal here it contains a dump mechanism, and a way to listen to events (add/remove routes) which is synchronized with that dump. Ian did start looking at netlink some time ago, but it also has some issues (it is in the network namespace not the fs namespace, it also has various things accumulated over the years that we don't need for filesystems) but that was part of the original inspiration for the fs notifications.
There is also, of course, /proc/net/route which can be useful in many circumstances, but for efficiency and synchronization reasons if is not the interface of choice for routing protocols. David's proposal has a number of the important attributes of an rtnetlink-like (in a conceptual sense) solution, and I remain skeptical that a /sysfs or similar interface would be an efficient solution to the original problem, even if it might perhaps make a useful addition.
There is also the chicken-and-egg issue, in the sense that if the interface is via a filesystem (sysfs, proc or whatever), how does one receive a notification for that filesystem itself being mounted until after it has been mounted? Maybe that is not a particular problem, but I think a cleaner solution would not require a mount in order to watch for other mounts,
Steve.
From: Christian Brauner <hidden> Date: 2020-03-03 10:25:55
On Tue, Mar 03, 2020 at 11:13:50AM +0100, Miklos Szeredi wrote:
On Tue, Mar 3, 2020 at 11:00 AM Christian Brauner
[off-list ref] wrote:
quoted
On Tue, Mar 03, 2020 at 10:26:21AM +0100, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 10:13 AM David Howells [off-list ref] wrote:
quoted
Miklos Szeredi [off-list ref] wrote:
quoted
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Don't forget the efficiency criterion. One reason for going with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in the system is
slow (not to mention the global lock that is held during the read).
Now, going with sysfs files on top of procfs links might avoid the global
lock, and you can avoid rereading the options string if you export a change
notification, but you're going to end up injecting a whole lot of pathwalk
latency into the system.
Completely irrelevant. Cached lookup is so much optimized, that you
won't be able to see any of it.
No, I don't think this is going to be a performance issue at all, but
if anything we could introduce a syscall
ssize_t readfile(int dfd, const char *path, char *buf, size_t
bufsize, int flags);
that is basically the equivalent of open + read + close, or even a
vectored variant that reads multiple files. But that's off topic
again, since I don't think there's going to be any performance issue
even with plain I/O syscalls.
quoted
On top of that, it isn't going to help with the case that I'm working towards
implementing where a container manager can monitor for mounts taking place
inside the container and supervise them. What I'm proposing is that during
the action phase (eg. FSCONFIG_CMD_CREATE), fsconfig() would hand an fd
referring to the context under construction to the manager, which would then
be able to call fsinfo() to query it and fsconfig() to adjust it, reject it or
permit it. Something like:
fd = receive_context_to_supervise();
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_OPTIONS,
};
fsinfo(fd, NULL, ¶ms, sizeof(params), buffer, sizeof(buffer));
supervise_parameters(buffer);
fsconfig(fd, FSCONFIG_SET_FLAG, "hard", NULL, 0);
fsconfig(fd, FSCONFIG_SET_STRING, "vers", "4.2", 0);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_CREATE, NULL, NULL, 0);
struct fsinfo_params params = {
.flags = FSINFO_FLAGS_QUERY_FSCONTEXT,
.request = FSINFO_ATTR_SB_NOTIFICATIONS,
};
struct fsinfo_sb_notifications sbnotify;
fsinfo(fd, NULL, ¶ms, sizeof(params), &sbnotify, sizeof(sbnotify));
watch_super(fd, "", AT_EMPTY_PATH, watch_fd, 0x03);
fsconfig(fd, FSCONFIG_CMD_SUPERVISE_PERMIT, NULL, NULL, 0);
close(fd);
However, the supervised mount may be happening in a completely different set
of namespaces, in which case the supervisor presumably wouldn't be able to see
the links in procfs and the relevant portions of sysfs.
It would be a "jump" link to the otherwise invisible directory.
More magic links to beam you around sounds like a bad idea. We had a
bunch of CVEs around them in containers and they were one of the major
reasons behind us pushing for openat2(). That's why it has a
RESOLVE_NO_MAGICLINKS flag.
No, that link wouldn't beam you around at all, it would end up in an
internally mounted instance of a mountfs, a safe place where no
Even if it is a magic link to a safe place it's a magic link. They
aren't a great solution to this problem. fsinfo() is cleaner and
simpler as it creates a context for a supervised mount which gives the a
managing application fine-grained control and makes it easily
extendable.
Also, we're apparently at the point where it seems were suggesting
another (pseudo)filesystem to get information about filesystems.
Christian
On Tue, Mar 3, 2020 at 11:22 AM Steven Whitehouse [off-list ref] wrote:
Hi,
On 03/03/2020 09:48, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 10:26 AM Miklos Szeredi [off-list ref] wrote:
quoted
On Tue, Mar 3, 2020 at 10:13 AM David Howells [off-list ref] wrote:
quoted
Miklos Szeredi [off-list ref] wrote:
quoted
I'm doing a patch. Let's see how it fares in the face of all these
preconceptions.
Don't forget the efficiency criterion. One reason for going with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in the system is
slow (not to mention the global lock that is held during the read).
BTW, I do feel that there's room for improvement in userspace code as
well. Even quite big mount table could be scanned for *changes* very
efficiently. l.e. cache previous contents of /proc/self/mountinfo and
compare with new contents, line-by-line. Only need to parse the
changed/added/removed lines.
Also it would be pretty easy to throttle the number of updates so
systemd et al. wouldn't hog the system with unnecessary processing.
Thanks,
Miklos
At least having patches to compare would allow us to look at the
performance here and gain some numbers, which would be helpful to frame
the discussions. However I'm not seeing how it would be easy to throttle
updates... they occur at whatever rate they are generated and this can
be fairly high. Also I'm not sure that I follow how the notifications
and the dumping of the whole table are synchronized in this case, either.
What I meant is optimizing current userspace without additional kernel
infrastructure. Since currently there's only the monolithic
/proc/self/mountinfo, it's reasonable that if the rate of change is
very high, then we don't re-read this table on every change, only
within a reasonable time limit (e.g. 1s) to provide timely updates.
Re-reading the table on every change would (does?) slow down the
system so that the actual updates would even be slower, so throttling
in this case very much makes sense.
Once we have per-mount information from the kernel, throttling updates
probably does not make sense.
Thanks,
Miklos
From: Ian Kent <raven@themaw.net> Date: 2020-03-03 11:09:47
On Tue, 2020-03-03 at 11:32 +0100, Miklos Szeredi wrote:
On Tue, Mar 3, 2020 at 11:22 AM Steven Whitehouse <
swhiteho@redhat.com> wrote:
quoted
Hi,
On 03/03/2020 09:48, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 10:26 AM Miklos Szeredi <miklos@szeredi.hu
quoted
wrote:
On Tue, Mar 3, 2020 at 10:13 AM David Howells <
dhowells@redhat.com> wrote:
quoted
Miklos Szeredi [off-list ref] wrote:
quoted
I'm doing a patch. Let's see how it fares in the face of
all these
preconceptions.
Don't forget the efficiency criterion. One reason for going
with fsinfo(2) is
that scanning /proc/mounts when there are a lot of mounts in
the system is
slow (not to mention the global lock that is held during the
read).
BTW, I do feel that there's room for improvement in userspace
code as
well. Even quite big mount table could be scanned for *changes*
very
efficiently. l.e. cache previous contents of
/proc/self/mountinfo and
compare with new contents, line-by-line. Only need to parse the
changed/added/removed lines.
Also it would be pretty easy to throttle the number of updates so
systemd et al. wouldn't hog the system with unnecessary
processing.
Thanks,
Miklos
At least having patches to compare would allow us to look at the
performance here and gain some numbers, which would be helpful to
frame
the discussions. However I'm not seeing how it would be easy to
throttle
updates... they occur at whatever rate they are generated and this
can
be fairly high. Also I'm not sure that I follow how the
notifications
and the dumping of the whole table are synchronized in this case,
either.
What I meant is optimizing current userspace without additional
kernel
infrastructure. Since currently there's only the monolithic
/proc/self/mountinfo, it's reasonable that if the rate of change is
very high, then we don't re-read this table on every change, only
within a reasonable time limit (e.g. 1s) to provide timely updates.
Re-reading the table on every change would (does?) slow down the
system so that the actual updates would even be slower, so throttling
in this case very much makes sense.
Optimizing user space is a huge task.
For example, consider this (which is related to a recent upstream
discussion I had):
https://blog.janestreet.com/troubleshooting-systemd-with-systemtap/
Working on improving libmount is really useful but that can't help
with inherently inefficient approaches to keeping info. current
which is actually needed at times.
Once we have per-mount information from the kernel, throttling
updates
probably does not make sense.
And can easily lead to application problems. Throttling will
lead to an inability to have up to date information upon which
application decisions are made.
I don't think it's a viable solution to the separate problem
of a large number of notifications either.
Ian
On Tue, Mar 3, 2020 at 11:25 AM Christian Brauner
[off-list ref] wrote:
On Tue, Mar 03, 2020 at 11:13:50AM +0100, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 11:00 AM Christian Brauner
[off-list ref] wrote:
quoted
quoted
More magic links to beam you around sounds like a bad idea. We had a
bunch of CVEs around them in containers and they were one of the major
reasons behind us pushing for openat2(). That's why it has a
RESOLVE_NO_MAGICLINKS flag.
No, that link wouldn't beam you around at all, it would end up in an
internally mounted instance of a mountfs, a safe place where no
Even if it is a magic link to a safe place it's a magic link. They
aren't a great solution to this problem. fsinfo() is cleaner and
simpler as it creates a context for a supervised mount which gives the a
managing application fine-grained control and makes it easily
extendable.
Yeah, it's a nice and clean interface in the ioctl(2) sense. Sure,
fsinfo() is way better than ioctl(), but it at the core it's still the
same syscall multiplexer, do everything hack.
Also, we're apparently at the point where it seems were suggesting
another (pseudo)filesystem to get information about filesystems.
Implementation detail. Why would you care?
Thanks,
Miklos
From: Karel Zak <kzak@redhat.com> Date: 2020-03-03 11:38:27
On Tue, Mar 03, 2020 at 10:26:21AM +0100, Miklos Szeredi wrote:
No, I don't think this is going to be a performance issue at all, but
if anything we could introduce a syscall
ssize_t readfile(int dfd, const char *path, char *buf, size_t
bufsize, int flags);
off-topic, but I'll buy you many many beers if you implement it ;-),
because open + read + close is pretty common for /sys and /proc in
many userspace tools; for example ps, top, lsblk, lsmem, lsns, udevd
etc. is all about it.
Karel
--
Karel Zak [off-list ref]
http://karelzak.blogspot.com
From: Christian Brauner <hidden> Date: 2020-03-03 11:57:06
On Tue, Mar 03, 2020 at 12:33:48PM +0100, Miklos Szeredi wrote:
On Tue, Mar 3, 2020 at 11:25 AM Christian Brauner
[off-list ref] wrote:
quoted
On Tue, Mar 03, 2020 at 11:13:50AM +0100, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 11:00 AM Christian Brauner
[off-list ref] wrote:
quoted
quoted
quoted
More magic links to beam you around sounds like a bad idea. We had a
bunch of CVEs around them in containers and they were one of the major
reasons behind us pushing for openat2(). That's why it has a
RESOLVE_NO_MAGICLINKS flag.
No, that link wouldn't beam you around at all, it would end up in an
internally mounted instance of a mountfs, a safe place where no
Even if it is a magic link to a safe place it's a magic link. They
aren't a great solution to this problem. fsinfo() is cleaner and
simpler as it creates a context for a supervised mount which gives the a
managing application fine-grained control and makes it easily
extendable.
Yeah, it's a nice and clean interface in the ioctl(2) sense. Sure,
fsinfo() is way better than ioctl(), but it at the core it's still the
same syscall multiplexer, do everything hack.
In contrast to a generic ioctl() it's a domain-specific separate
syscall. You can't suddenly set kvm options through fsinfo() I would
hope. I find it at least debatable that a new filesystem is preferable.
And - feel free to simply dismiss the concerns I expressed - so far
there has not been a lot of excitement about this idea.
quoted
Also, we're apparently at the point where it seems were suggesting
another (pseudo)filesystem to get information about filesystems.
Implementation detail. Why would you care?
I wouldn't call this an implementation detail. That's quite a big
design choice; it's a separate fileystem. In addition, implementation
details need to be maintained.
Christian
On Tue, Mar 03, 2020 at 12:38:14PM +0100, Karel Zak wrote:
On Tue, Mar 03, 2020 at 10:26:21AM +0100, Miklos Szeredi wrote:
quoted
No, I don't think this is going to be a performance issue at all, but
if anything we could introduce a syscall
ssize_t readfile(int dfd, const char *path, char *buf, size_t
bufsize, int flags);
off-topic, but I'll buy you many many beers if you implement it ;-),
because open + read + close is pretty common for /sys and /proc in
many userspace tools; for example ps, top, lsblk, lsmem, lsns, udevd
etc. is all about it.
Unlimited beers for a 21-line kernel patch? Sign me up!
Totally untested, barely compiled patch below.
Actually, I like this idea (the syscall, not just the unlimited beers).
Maybe this could make a lot of sense, I'll write some actual tests for
it now that syscalls are getting "heavy" again due to CPU vendors
finally paying the price for their madness...
thanks,
greg k-h
-------------------
@@ -359,6 +359,7 @@ 435 common clone3 __x64_sys_clone3/ptregs 437 common openat2 __x64_sys_openat2 438 common pidfd_getfd __x64_sys_pidfd_getfd+439 common readfile __x86_sys_readfile # # x32-specific system call numbers start at 512 to avoid cache impact
On Tue, Mar 03, 2020 at 02:03:47PM +0100, Greg Kroah-Hartman wrote:
On Tue, Mar 03, 2020 at 12:38:14PM +0100, Karel Zak wrote:
quoted
On Tue, Mar 03, 2020 at 10:26:21AM +0100, Miklos Szeredi wrote:
quoted
No, I don't think this is going to be a performance issue at all, but
if anything we could introduce a syscall
ssize_t readfile(int dfd, const char *path, char *buf, size_t
bufsize, int flags);
off-topic, but I'll buy you many many beers if you implement it ;-),
because open + read + close is pretty common for /sys and /proc in
many userspace tools; for example ps, top, lsblk, lsmem, lsns, udevd
etc. is all about it.
Unlimited beers for a 21-line kernel patch? Sign me up!
Totally untested, barely compiled patch below.
Ok, that didn't even build, let me try this for real now...
On Tue, Mar 3, 2020 at 2:14 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
Unlimited beers for a 21-line kernel patch? Sign me up!
Totally untested, barely compiled patch below.
Ok, that didn't even build, let me try this for real now...
Some comments on the interface:
O_LARGEFILE can be unconditional, since offsets are not exposed to the caller.
Use the openat2 style arguments; limit the accepted flags to sane ones
(e.g. don't let this syscall create a file).
If buffer is too small to fit the whole file, return error.
Verify that the number of bytes read matches the file size, otherwise
return error (may need to loop?).
Thanks,
Miklos
On Tue, Mar 03, 2020 at 02:34:42PM +0100, Miklos Szeredi wrote:
On Tue, Mar 3, 2020 at 2:14 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
quoted
Unlimited beers for a 21-line kernel patch? Sign me up!
Totally untested, barely compiled patch below.
Ok, that didn't even build, let me try this for real now...
Some comments on the interface:
Ok, hey, let's do this proper :)
O_LARGEFILE can be unconditional, since offsets are not exposed to the caller.
Good point.
Use the openat2 style arguments; limit the accepted flags to sane ones
(e.g. don't let this syscall create a file).
Yeah, I just added that check to my local version:
/* Mask off all O_ flags as we only want to read from the file */
flags &= ~(VALID_OPEN_FLAGS);
flags |= O_RDONLY | O_LARGEFILE;
If buffer is too small to fit the whole file, return error.
Why? What's wrong with just returning the bytes asked for? If someone
only wants 5 bytes from the front of a file, it should be fine to give
that to them, right?
Verify that the number of bytes read matches the file size, otherwise
return error (may need to loop?).
No, we can't "match file size" as sysfs files do not really have a sane
"size". So I don't want to loop at all here, one-shot, that's all you
get :)
Let me actually do this and try it out for real.
/me has no idea what he is getting himself into...
thanks,
greg k-h
On Tue, Mar 03, 2020 at 02:43:16PM +0100, Greg Kroah-Hartman wrote:
On Tue, Mar 03, 2020 at 02:34:42PM +0100, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 2:14 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
quoted
Unlimited beers for a 21-line kernel patch? Sign me up!
Totally untested, barely compiled patch below.
Ok, that didn't even build, let me try this for real now...
Some comments on the interface:
Ok, hey, let's do this proper :)
Alright, how about this patch.
Actually tested with some simple sysfs files.
If people don't strongly object, I'll add "real" tests to it, hook it up
to all arches, write a manpage, and all the fun fluff a new syscall
deserves and submit it "for real".
It feels like I'm doing something wrong in that the actuall syscall
logic is just so small. Maybe I'll benchmark this thing to see if it
makes any real difference...
thanks,
greg k-h
From: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Subject: [PATCH] readfile: implement readfile syscall
It's a tiny syscall, meant to allow a user to do a single "open this
file, read into this buffer, and close the file" all in a single shot.
Should be good for reading "tiny" files like sysfs, procfs, and other
"small" files.
There is no restarting the syscall, am trying to keep it simple. At
least for now.
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
---
arch/x86/entry/syscalls/syscall_32.tbl | 1 +
arch/x86/entry/syscalls/syscall_64.tbl | 1 +
fs/open.c | 21 +++++++++++++++++++++
include/linux/syscalls.h | 2 ++
include/uapi/asm-generic/unistd.h | 4 +++-
5 files changed, 28 insertions(+), 1 deletion(-)
@@ -359,6 +359,7 @@ 435 common clone3 __x64_sys_clone3/ptregs 437 common openat2 __x64_sys_openat2 438 common pidfd_getfd __x64_sys_pidfd_getfd+439 common readfile __x64_sys_readfile # # x32-specific system call numbers start at 512 to avoid cache impact
@@ -1340,3 +1340,24 @@ int stream_open(struct inode *inode, struct file *filp)}EXPORT_SYMBOL(stream_open);++SYSCALL_DEFINE5(readfile,int,dfd,constchar__user*,filename,+char__user*,buffer,size_t,bufsize,int,flags)+{+intretval;+intfd;++/* Mask off all O_ flags as we only want to read from the file */+flags&=~(VALID_OPEN_FLAGS);+flags|=O_RDONLY|O_LARGEFILE;++fd=do_sys_open(dfd,filename,flags,0000);+if(fd<=0)+returnfd;++retval=ksys_read(fd,buffer,bufsize);++__close_fd(current->files,fd);++returnretval;+}
@@ -1003,6 +1003,8 @@ asmlinkage long sys_pidfd_send_signal(int pidfd, int sig,siginfo_t__user*info,unsignedintflags);asmlinkagelongsys_pidfd_getfd(intpidfd,intfd,unsignedintflags);+asmlinkagelongsys_readfile(intdfd,constchar__user*filename,+char__user*buffer,size_tbufsize,intflags);/**Architecture-specificsystemcalls
On Tue, Mar 3, 2020 at 2:43 PM Greg Kroah-Hartman
[off-list ref] wrote:
On Tue, Mar 03, 2020 at 02:34:42PM +0100, Miklos Szeredi wrote:
quoted
If buffer is too small to fit the whole file, return error.
Why? What's wrong with just returning the bytes asked for? If someone
only wants 5 bytes from the front of a file, it should be fine to give
that to them, right?
I think we need to signal in some way to the caller that the result
was truncated (see readlink(2), getxattr(2), getcwd(2)), otherwise the
caller might be surprised.
quoted
Verify that the number of bytes read matches the file size, otherwise
return error (may need to loop?).
No, we can't "match file size" as sysfs files do not really have a sane
"size". So I don't want to loop at all here, one-shot, that's all you
get :)
Hmm. I understand the no-size thing. But looping until EOF (i.e.
until read return zero) might be a good idea regardless, because short
reads are allowed.
Thanks,
Miklos
On Tue, Mar 3, 2020 at 3:10 PM Greg Kroah-Hartman
[off-list ref] wrote:
On Tue, Mar 03, 2020 at 02:43:16PM +0100, Greg Kroah-Hartman wrote:
quoted
On Tue, Mar 03, 2020 at 02:34:42PM +0100, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 2:14 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
quoted
Unlimited beers for a 21-line kernel patch? Sign me up!
Totally untested, barely compiled patch below.
Ok, that didn't even build, let me try this for real now...
Some comments on the interface:
Ok, hey, let's do this proper :)
Alright, how about this patch.
Actually tested with some simple sysfs files.
If people don't strongly object, I'll add "real" tests to it, hook it up
to all arches, write a manpage, and all the fun fluff a new syscall
deserves and submit it "for real".
Just FYI, io_uring is moving towards the same kind of thing... IIRC
you can already use it to batch a bunch of open() calls, then batch a
bunch of read() calls on all the new fds and close them at the same
time. And I think they're planning to add support for doing
open()+read()+close() all in one go, too, except that it's a bit
complicated because passing forward the file descriptor in a generic
way is a bit complicated.
It feels like I'm doing something wrong in that the actuall syscall
logic is just so small. Maybe I'll benchmark this thing to see if it
makes any real difference...
thanks,
greg k-h
From: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Subject: [PATCH] readfile: implement readfile syscall
It's a tiny syscall, meant to allow a user to do a single "open this
file, read into this buffer, and close the file" all in a single shot.
Should be good for reading "tiny" files like sysfs, procfs, and other
"small" files.
There is no restarting the syscall, am trying to keep it simple. At
least for now.
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
[...]
quoted hunk
+SYSCALL_DEFINE5(readfile, int, dfd, const char __user *, filename,+ char __user *, buffer, size_t, bufsize, int, flags)+{+ int retval;+ int fd;++ /* Mask off all O_ flags as we only want to read from the file */+ flags &= ~(VALID_OPEN_FLAGS);+ flags |= O_RDONLY | O_LARGEFILE;++ fd = do_sys_open(dfd, filename, flags, 0000);+ if (fd <= 0)+ return fd;++ retval = ksys_read(fd, buffer, bufsize);++ __close_fd(current->files, fd);++ return retval;+}
If you're gonna do something like that, wouldn't you want to also
elide the use of the file descriptor table completely? do_sys_open()
will have to do atomic operations in the fd table and stuff, which is
probably moderately bad in terms of cacheline bouncing if this is used
in a multithreaded context; and as a side effect, the fd would be
inherited by anyone who calls fork() concurrently. You'll probably
want to use APIs like do_filp_open() and filp_close(), or something
like that, instead.
If you can use dentry_open() and vfs_read(), you might be able to avoid
dealing with file descriptors entirely. That might make it worth a syscall.
You're going to be asked for writefile() you know ;-)
David
From: Christian Brauner <hidden> Date: 2020-03-03 14:24:10
On Tue, Mar 03, 2020 at 02:34:42PM +0100, Miklos Szeredi wrote:
On Tue, Mar 3, 2020 at 2:14 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
quoted
Unlimited beers for a 21-line kernel patch? Sign me up!
Totally untested, barely compiled patch below.
Ok, that didn't even build, let me try this for real now...
Some comments on the interface:
O_LARGEFILE can be unconditional, since offsets are not exposed to the caller.
Use the openat2 style arguments; limit the accepted flags to sane ones
(e.g. don't let this syscall create a file).
If we think this is worth it, might even good to either have it support
struct open_how or have it accept two flag arguments. We sure want
openat2()s RESOLVE_* flags in there.
Christian
On Tue, Mar 03, 2020 at 03:13:26PM +0100, Jann Horn wrote:
On Tue, Mar 3, 2020 at 3:10 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
On Tue, Mar 03, 2020 at 02:43:16PM +0100, Greg Kroah-Hartman wrote:
quoted
On Tue, Mar 03, 2020 at 02:34:42PM +0100, Miklos Szeredi wrote:
quoted
On Tue, Mar 3, 2020 at 2:14 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
quoted
Unlimited beers for a 21-line kernel patch? Sign me up!
Totally untested, barely compiled patch below.
Ok, that didn't even build, let me try this for real now...
Some comments on the interface:
Ok, hey, let's do this proper :)
Alright, how about this patch.
Actually tested with some simple sysfs files.
If people don't strongly object, I'll add "real" tests to it, hook it up
to all arches, write a manpage, and all the fun fluff a new syscall
deserves and submit it "for real".
Just FYI, io_uring is moving towards the same kind of thing... IIRC
you can already use it to batch a bunch of open() calls, then batch a
bunch of read() calls on all the new fds and close them at the same
time. And I think they're planning to add support for doing
open()+read()+close() all in one go, too, except that it's a bit
complicated because passing forward the file descriptor in a generic
way is a bit complicated.
It is complicated, I wouldn't recommend using io_ring for reading a
bunch of procfs or sysfs files, that feels like a ton of overkill with
too much setup/teardown to make it worth while.
But maybe not, will have to watch and see how it goes.
quoted
It feels like I'm doing something wrong in that the actuall syscall
logic is just so small. Maybe I'll benchmark this thing to see if it
makes any real difference...
thanks,
greg k-h
From: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Subject: [PATCH] readfile: implement readfile syscall
It's a tiny syscall, meant to allow a user to do a single "open this
file, read into this buffer, and close the file" all in a single shot.
Should be good for reading "tiny" files like sysfs, procfs, and other
"small" files.
There is no restarting the syscall, am trying to keep it simple. At
least for now.
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
[...]
quoted
+SYSCALL_DEFINE5(readfile, int, dfd, const char __user *, filename,+ char __user *, buffer, size_t, bufsize, int, flags)+{+ int retval;+ int fd;++ /* Mask off all O_ flags as we only want to read from the file */+ flags &= ~(VALID_OPEN_FLAGS);+ flags |= O_RDONLY | O_LARGEFILE;++ fd = do_sys_open(dfd, filename, flags, 0000);+ if (fd <= 0)+ return fd;++ retval = ksys_read(fd, buffer, bufsize);++ __close_fd(current->files, fd);++ return retval;+}
If you're gonna do something like that, wouldn't you want to also
elide the use of the file descriptor table completely? do_sys_open()
will have to do atomic operations in the fd table and stuff, which is
probably moderately bad in terms of cacheline bouncing if this is used
in a multithreaded context; and as a side effect, the fd would be
inherited by anyone who calls fork() concurrently. You'll probably
want to use APIs like do_filp_open() and filp_close(), or something
like that, instead.
Ah, nice, that does make more sense. I'll play around with that, and
benchmarking this thing later tonight. Have to go get some stable
kernels out first...
thanks for the quick review, much appreciated.
greg k-h
On Tue, Mar 03, 2020 at 03:10:50PM +0100, Miklos Szeredi wrote:
On Tue, Mar 3, 2020 at 2:43 PM Greg Kroah-Hartman
[off-list ref] wrote:
quoted
On Tue, Mar 03, 2020 at 02:34:42PM +0100, Miklos Szeredi wrote:
quoted
quoted
If buffer is too small to fit the whole file, return error.
Why? What's wrong with just returning the bytes asked for? If someone
only wants 5 bytes from the front of a file, it should be fine to give
that to them, right?
I think we need to signal in some way to the caller that the result
was truncated (see readlink(2), getxattr(2), getcwd(2)), otherwise the
caller might be surprised.
But that's not the way a "normal" read works. Short reads are fine, if
the file isn't big enough. That's how char device nodes work all the
time as well, and this kind of is like that, or some kind of "stream" to
read from.
If you think the file is bigger, then you, as the caller, can just pass
in a bigger buffer if you want to (i.e. you can stat the thing and
determine the size beforehand.)
Think of the "normal" use case here, a sysfs read with a PAGE_SIZE
buffer. That way userspace "knows" it will always read all of the data
it can from the file, we don't have to do any seeking or determining
real file size, or anything else like that.
We return the number of bytes read as well, so we "know" if we did a
short read, and also, you could imply, if the number of bytes read are
the exact same as the number of bytes of the buffer, maybe the file is
either that exact size, or bigger.
This should be "simple", let's not make it complex if we can help it :)
quoted
quoted
Verify that the number of bytes read matches the file size, otherwise
return error (may need to loop?).
No, we can't "match file size" as sysfs files do not really have a sane
"size". So I don't want to loop at all here, one-shot, that's all you
get :)
Hmm. I understand the no-size thing. But looping until EOF (i.e.
until read return zero) might be a good idea regardless, because short
reads are allowed.
If you want to loop, then do a userspace open/read-loop/close cycle.
That's not what this syscall should be for.
Should we call it: readfile-only-one-try-i-hope-my-buffer-is-big-enough()? :)
thanks,
greg k-h