Re: [CRIU] Introspecting userns relationships to other namespaces?
From: James Bottomley <hidden>
Date: 2016-07-09 10:33:20
Also in:
lkml
On July 9, 2016 4:26:28 PM GMT+09:00, Andrew Vagin [off-list ref] wrote:
On Fri, Jul 08, 2016 at 10:05:18PM -0500, Eric W. Biederman wrote:quoted
James Bottomley [off-list ref] writes:quoted
On Fri, 2016-07-08 at 18:52 -0500, Eric W. Biederman wrote:quoted
James Bottomley [off-list ref] writes:quoted
On July 8, 2016 1:38:19 PM PDT, Andrew Vagin[off-list ref]quoted
quoted
quoted
quoted
wrote:quoted
quoted
What do you think about the idea to mount nsfs and be able to look up any alive namespace by inum:I think I like it. It will give us a way to enter any extant namespace. It will work for Eric's fs namespaces as well.Perhapsquoted
quoted
quoted
quoted
a /process/ns/<inum> Directory?As you understood, I meant /proc/ns/<inum> (damn mobile phone completions).quoted
*Shivers* That makes it very easy to bypass any existing controls that existquoted
quoted
quoted
for getting at namespaces. It is true that everything of thatkindquoted
quoted
quoted
is directory based but still. Plus I think it would serve as information leak to information outside of the container. An operation to get a user namespace file descriptor from somekernelquoted
quoted
quoted
object sounds reasonably sane. A great big list of things sounds about as scary as it can get.Thisquoted
quoted
quoted
is not the time to be making it easier to escape from containers.To be honest, I think this argument is rubbish. If we're afraid of giving out a list of all the namespaces, it means we're afraidthere'squoted
quoted
some security bug and we're trying to obscure it by making the list hard to get. All we've done is allayed fears about the bug but the hackers still know the portals to get through. If such a bug exists, it will be possible to exploit it by simply reconstructing the information from the individual processdirectories,quoted
quoted
so obscurity doesn't protect us and all it does is give us a false sense of security. If such a bug doesn't exist, then all thesecurityquoted
quoted
mechanisms currently in place (like no re-entry to prior namespace) should protect us and we can give out the list. Let's deal with the world as we'd like it to be (no obscurenamespacequoted
quoted
bugs) and accept the consequences and the responsibility for fixing them if we turn out to be slightly incorrect. We'll end up in afarquoted
quoted
better place than security by obscurity would land us.No. That is not the fear. The permission checks on/proc/self/ns/xxxquoted
are different than if the namespace is bind mounted somewhere. That was done deliberately and with a reasonable amount offorethought.quoted
You are asking to throw those permission checks out. The answer isno.quoted
Furthermore there is a much clearer reason not to go with a list ofallquoted
namespaces. A list of all namespaces breaks CRIU. As you havedescribedquoted
it the list will change depending upon which machine you restore a checkpoint on. I honestly don't know what kind of havoc that willcausequoted
but it is certainly something we won't be able to checkpoint nomatterquoted
how hard we try.It's right. I hadn't thought about this.
Me neither. Sorry for the prior outburst. I think this means we're back to exposing owning userns in the /proc /<pid >/ns directory.
quoted
A global list of namespaces especially of the kind that you can open and get a handle to the namespace is just not appropriate. I know inode numbers comes darn close to names but they aren't really names and if it comes to it we can figure out how to preserve an applications view of it all across a checkpoint/restart. So far it hasn't proven necessary to preserve any inode numbers across checkpoint/restart but again it is theoretically possible if itbecomesquoted
necessary. Throwing away checkpoint/restart support for the sake of checkpoint/restart is a no-go. Containers fundamentally imply you don't have global visibility, and that is a good thing.All these thoughts about security make me thinking that kcmp is what we should use here. It's maybe something like this: kcmp(pid1, pid2, KCMP_NS_USERNS, fd1, fd2) - to check if userns of the fd1 namepsace is equal to the fd2 userns kcmp(pid1, pid2, KCMP_NS_PARENT, fd1, fd2) - to check if a parent namespace of the fd1 pidns is equal to fd pidns. fd1 and fd2 is file descriptors to namespace files. So if we want to build a hierarchy, we need to collect all namespaces and then enumerate them to check dependencies with help of kcmp.
Sure, but we need a method for opening the filehandles first .. . James
quoted
Eric
-- Sent from my Android device with K-9 Mail. Please excuse my brevity.