From: Eric W. Biederman <hidden> Date: 2016-07-09 18:41:52
ebiederm-aS9lmoZGLiVWk0Htik3J/w@public.gmane.org (Eric W. Biederman) writes:
Andrew Vagin [off-list ref] writes:
quoted
All these thoughts about security make me thinking that kcmp is what we
should use here. It's maybe something like this:
kcmp(pid1, pid2, KCMP_NS_USERNS, fd1, fd2)
- to check if userns of the fd1 namepsace is equal to the fd2 userns
kcmp(pid1, pid2, KCMP_NS_PARENT, fd1, fd2)
- to check if a parent namespace of the fd1 pidns is equal to fd pidns.
fd1 and fd2 is file descriptors to namespace files.
So if we want to build a hierarchy, we need to collect all namespaces
and then enumerate them to check dependencies with help of kcmp.
That is certainly one way to go.
There is a funny case where we would want to compare a user namespace
file descriptor to a parent user namespace file descriptor.
Grumble, Grumble. I think this may actually a case for creating ioctls
for these two cases. Now that random nsfs file descriptors are bind
mountable the original reason for using proc files is not as pressing.
One ioctl for the user namespace that owns a file descriptor.
One ioctl for the parent namespace of a namespace file descriptor.
We also need some way to get a command file descriptor for a file system
super block. Al Viro has a pet project for cleaning up the mount API
and this might be the idea excuse to start looking at that.
(In principle we might be able to run commands through the namespace
file descriptor and using an ioctl feels dirty. But an ioctl that
only uses the fd and request argument does not suffer from the same
problems that ioctls that have to pass additional arguments suffer
from.)
Of course it should be an error perhaps -EINVAL to get a user
namespace owner or parent namespace that is outside of a processes
current user namespace or pid namespace. That way thing stay bounded
within the current namespaces the process is in. Which prevents any
leak possibilities, and keeps CRIU working.
Eric
From: Andrew Vagin <hidden> Date: 2016-07-13 03:43:56
On Sat, Jul 09, 2016 at 01:29:20PM -0500, Eric W. Biederman wrote:
ebiederm-aS9lmoZGLiVWk0Htik3J/w@public.gmane.org (Eric W. Biederman) writes:
quoted
Andrew Vagin [off-list ref] writes:
quoted
All these thoughts about security make me thinking that kcmp is what we
should use here. It's maybe something like this:
kcmp(pid1, pid2, KCMP_NS_USERNS, fd1, fd2)
- to check if userns of the fd1 namepsace is equal to the fd2 userns
kcmp(pid1, pid2, KCMP_NS_PARENT, fd1, fd2)
- to check if a parent namespace of the fd1 pidns is equal to fd pidns.
fd1 and fd2 is file descriptors to namespace files.
So if we want to build a hierarchy, we need to collect all namespaces
and then enumerate them to check dependencies with help of kcmp.
That is certainly one way to go.
There is a funny case where we would want to compare a user namespace
file descriptor to a parent user namespace file descriptor.
Grumble, Grumble. I think this may actually a case for creating ioctls
for these two cases. Now that random nsfs file descriptors are bind
mountable the original reason for using proc files is not as pressing.
One ioctl for the user namespace that owns a file descriptor.
One ioctl for the parent namespace of a namespace file descriptor.
We also need some way to get a command file descriptor for a file system
super block. Al Viro has a pet project for cleaning up the mount API
and this might be the idea excuse to start looking at that.
(In principle we might be able to run commands through the namespace
file descriptor and using an ioctl feels dirty. But an ioctl that
only uses the fd and request argument does not suffer from the same
problems that ioctls that have to pass additional arguments suffer
from.)
Of course it should be an error perhaps -EINVAL to get a user
namespace owner or parent namespace that is outside of a processes
current user namespace or pid namespace. That way thing stay bounded
within the current namespaces the process is in. Which prevents any
leak possibilities, and keeps CRIU working.
Overall this looks good to me (I left a handful of uninformed comments
inline ;).
It doesn't make it easy to walk leafward, but it doesn't look like the
kernel has a convenient way to list child namespaces either.
Something like /proc/<pid>/task/<tid>/children (with
CONFIG_PROC_CHILDREN) for namespaces would make it easier to get a
complete system overview (as far as your credentials and position in
the namespace hierarchies allow). But looking at the
CONFIG_PROC_CHILDREN implementation doesn't make me all that excited
about mimicking it for namespaces ;).
You can still brute-force it in userspace by walking the root-most
procfs's you can find and peeking at all the /proc/<pid>/ns/… entries
(but yuck ;). With mount and other namespaces not being hierarchical,
the “leafword” idea may not be all that useful anyway, but having a
more compact collection of mount namepaces (say) that you know about
would be nice. Where “know about” should probably means “know it
exists” but not necessarily “have permission to enter”. Still,
getting that figured out can happen independently to this parent/owner
work.
Cheers,
Trevor
--
This email may be signed or encrypted with GnuPG (http://www.gnupg.org).
For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy