From: Richard Guy Briggs <hidden> Date: 2020-03-18 21:42:24
On 2020-03-18 17:01, Paul Moore wrote:
On Fri, Mar 13, 2020 at 3:23 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:42, Paul Moore wrote:
...
quoted
quoted
The thread has had a lot of starts/stops, so I may be repeating a
previous suggestion, but one idea would be to still emit a "death
record" when the final task in the audit container ID does die, but
block the particular audit container ID from reuse until it the
SIGNAL2 info has been reported. This gives us the timely ACID death
notification while still preventing confusion and ambiguity caused by
potentially reusing the ACID before the SIGNAL2 record has been sent;
there is a small nit about the ACID being present in the SIGNAL2
*after* its death, but I think that can be easily explained and
understood by admins.
Thinking quickly about possible technical solutions to this, maybe it
makes sense to have two counters on a contobj so that we know when the
last process in that container exits and can issue the death
certificate, but we still block reuse of it until all further references
to it have been resolved. This will likely also make it possible to
report the full contid chain in SIGNAL2 records. This will eliminate
some of the issues we are discussing with regards to passing a contobj
vs a contid to the audit_log_contid function, but won't eliminate them
all because there are still some contids that won't have an object
associated with them to make it impossible to look them up in the
contobj lists.
I'm not sure you need a full second counter, I imagine a simple flag
would be okay. I think you just something to indicate that this ACID
object is marked as "dead" but it still being held for sanity reasons
and should not be reused.
Ok, I see your point. This refcount can be changed to a flag easily
enough without change to the api if we can be sure that more than one
signal can't be delivered to the audit daemon *and* collected by sig2.
I'll have a more careful look at the audit daemon code to see if I can
determine this.
Steve, can you have a look and tell us if it is possible for the audit
daemon to make more than one signal_info (or signal_info2) record
request from the kernel after receiving a signal?
Another question occurs to me is that what if the audit daemon is sent a
signal and it cannot or will not collect the sig2 information from the
kernel (SIGKILL?)? Does that audit container identifier remain dead
until reboot, or do we institute some other form of reaping, possibly
time-based?
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-18 21:42:46
On Wed, Mar 18, 2020 at 5:27 PM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-18 16:56, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 2:59 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:29, Paul Moore wrote:
quoted
On Thu, Mar 12, 2020 at 3:30 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-02-13 16:44, Paul Moore wrote:
quoted
This is a bit of a thread-hijack, and for that I apologize, but
another thought crossed my mind while thinking about this issue
further ... Once we support multiple auditd instances, including the
necessary record routing and duplication/multiple-sends (the host
always sees *everything*), we will likely need to find a way to "trim"
the audit container ID (ACID) lists we send in the records. The
auditd instance running on the host/initns will always see everything,
so it will want the full container ACID list; however an auditd
instance running inside a container really should only see the ACIDs
of any child containers.
Agreed. This should be easy to check and limit, preventing an auditd
from seeing any contid that is a parent of its own contid.
quoted
For example, imagine a system where the host has containers 1 and 2,
each running an auditd instance. Inside container 1 there are
containers A and B. Inside container 2 there are containers Y and Z.
If an audit event is generated in container Z, I would expect the
host's auditd to see a ACID list of "1,Z" but container 1's auditd
should only see an ACID list of "Z". The auditd running in container
2 should not see the record at all (that will be relatively
straightforward). Does that make sense? Do we have the record
formats properly designed to handle this without too much problem (I'm
not entirely sure we do)?
I completely agree and I believe we have record formats that are able to
handle this already.
I'm not convinced we do. What about the cases where we have a field
with a list of audit container IDs? How do we handle that?
I don't understand the problem. (I think you crossed your 1/2 vs
A/B/Y/Z in your example.) ...
It looks like I did, sorry about that.
quoted
... Clarifying the example above, if as you
suggest an event happens in container Z, the hosts's auditd would report
Z,^2
and the auditd in container 2 would report
Z,^2
but if there were another auditd running in container Z it would report
Z
while the auditd in container 1 or A/B would see nothing.
Yes. My concern is how do we handle this to minimize duplicating and
rewriting the records? It isn't so much about the format, although
the format is a side effect.
Are you talking about caching, or about divulging more information than
necessary or even information leaks? Or even noticing that records that
need to be generated to two audit daemons share the same contid field
values and should be generated at the same time or information shared
between them? I'd see any of these as optimizations that don't affect
the api.
Imagine a record is generated in a container which has more than one
auditd in it's ancestry that should receive this record, how do we
handle that without completely killing performance? That's my
concern. If you've already thought up a plan for this - excellent,
please share :)
--
paul moore
www.paul-moore.com
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-18 21:48:04
On Wed, Mar 18, 2020 at 5:42 PM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-18 17:01, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 3:23 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:42, Paul Moore wrote:
...
quoted
quoted
The thread has had a lot of starts/stops, so I may be repeating a
previous suggestion, but one idea would be to still emit a "death
record" when the final task in the audit container ID does die, but
block the particular audit container ID from reuse until it the
SIGNAL2 info has been reported. This gives us the timely ACID death
notification while still preventing confusion and ambiguity caused by
potentially reusing the ACID before the SIGNAL2 record has been sent;
there is a small nit about the ACID being present in the SIGNAL2
*after* its death, but I think that can be easily explained and
understood by admins.
Thinking quickly about possible technical solutions to this, maybe it
makes sense to have two counters on a contobj so that we know when the
last process in that container exits and can issue the death
certificate, but we still block reuse of it until all further references
to it have been resolved. This will likely also make it possible to
report the full contid chain in SIGNAL2 records. This will eliminate
some of the issues we are discussing with regards to passing a contobj
vs a contid to the audit_log_contid function, but won't eliminate them
all because there are still some contids that won't have an object
associated with them to make it impossible to look them up in the
contobj lists.
I'm not sure you need a full second counter, I imagine a simple flag
would be okay. I think you just something to indicate that this ACID
object is marked as "dead" but it still being held for sanity reasons
and should not be reused.
Ok, I see your point. This refcount can be changed to a flag easily
enough without change to the api if we can be sure that more than one
signal can't be delivered to the audit daemon *and* collected by sig2.
I'll have a more careful look at the audit daemon code to see if I can
determine this.
Maybe I'm not understanding your concern, but this isn't really
different than any of the other things we track for the auditd signal
sender, right? If we are worried about multiple signals being sent
then it applies to everything, not just the audit container ID.
Another question occurs to me is that what if the audit daemon is sent a
signal and it cannot or will not collect the sig2 information from the
kernel (SIGKILL?)? Does that audit container identifier remain dead
until reboot, or do we institute some other form of reaping, possibly
time-based?
In order to preserve the integrity of the audit log that ACID value
would need to remain unavailable until the ACID which contains the
associated auditd is "dead" (no one can request the signal sender's
info if that container is dead).
--
paul moore
www.paul-moore.com
From: Richard Guy Briggs <hidden> Date: 2020-03-18 21:56:25
On 2020-03-18 17:42, Paul Moore wrote:
On Wed, Mar 18, 2020 at 5:27 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 16:56, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 2:59 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:29, Paul Moore wrote:
quoted
On Thu, Mar 12, 2020 at 3:30 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-02-13 16:44, Paul Moore wrote:
quoted
This is a bit of a thread-hijack, and for that I apologize, but
another thought crossed my mind while thinking about this issue
further ... Once we support multiple auditd instances, including the
necessary record routing and duplication/multiple-sends (the host
always sees *everything*), we will likely need to find a way to "trim"
the audit container ID (ACID) lists we send in the records. The
auditd instance running on the host/initns will always see everything,
so it will want the full container ACID list; however an auditd
instance running inside a container really should only see the ACIDs
of any child containers.
Agreed. This should be easy to check and limit, preventing an auditd
from seeing any contid that is a parent of its own contid.
quoted
For example, imagine a system where the host has containers 1 and 2,
each running an auditd instance. Inside container 1 there are
containers A and B. Inside container 2 there are containers Y and Z.
If an audit event is generated in container Z, I would expect the
host's auditd to see a ACID list of "1,Z" but container 1's auditd
should only see an ACID list of "Z". The auditd running in container
2 should not see the record at all (that will be relatively
straightforward). Does that make sense? Do we have the record
formats properly designed to handle this without too much problem (I'm
not entirely sure we do)?
I completely agree and I believe we have record formats that are able to
handle this already.
I'm not convinced we do. What about the cases where we have a field
with a list of audit container IDs? How do we handle that?
I don't understand the problem. (I think you crossed your 1/2 vs
A/B/Y/Z in your example.) ...
It looks like I did, sorry about that.
quoted
... Clarifying the example above, if as you
suggest an event happens in container Z, the hosts's auditd would report
Z,^2
and the auditd in container 2 would report
Z,^2
but if there were another auditd running in container Z it would report
Z
while the auditd in container 1 or A/B would see nothing.
Yes. My concern is how do we handle this to minimize duplicating and
rewriting the records? It isn't so much about the format, although
the format is a side effect.
Are you talking about caching, or about divulging more information than
necessary or even information leaks? Or even noticing that records that
need to be generated to two audit daemons share the same contid field
values and should be generated at the same time or information shared
between them? I'd see any of these as optimizations that don't affect
the api.
Imagine a record is generated in a container which has more than one
auditd in it's ancestry that should receive this record, how do we
handle that without completely killing performance? That's my
concern. If you've already thought up a plan for this - excellent,
please share :)
No, I haven't given that much thought other than the correctness and
security issues of making sure that each audit daemon is sufficiently
isolated to do its job but not jeopardize another audit domain. Audit
already kills performance, according to some...
We currently won't have that problem since there can only be one so far.
Fixing and optimizing this is part of the next phase of the challenge of
adding a second audit daemon.
Let's work on correctness and reasonable efficiency for this phase and
not focus on a problem we don't yet have. I wouldn't consider this
incurring technical debt at this point.
I could see cacheing a contid string from one starting point, but it may
be more work to search that cached string to truncate it or add to it
when another audit daemon requests a copy of a similar string. I
suppose every full contid string could be generated the first time it is
used and parts of it used (start/finish) as needed but that
search/indexing may not be worth it.
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-18 22:06:16
On Wed, Mar 18, 2020 at 5:56 PM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-18 17:42, Paul Moore wrote:
quoted
On Wed, Mar 18, 2020 at 5:27 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 16:56, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 2:59 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:29, Paul Moore wrote:
quoted
On Thu, Mar 12, 2020 at 3:30 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-02-13 16:44, Paul Moore wrote:
quoted
This is a bit of a thread-hijack, and for that I apologize, but
another thought crossed my mind while thinking about this issue
further ... Once we support multiple auditd instances, including the
necessary record routing and duplication/multiple-sends (the host
always sees *everything*), we will likely need to find a way to "trim"
the audit container ID (ACID) lists we send in the records. The
auditd instance running on the host/initns will always see everything,
so it will want the full container ACID list; however an auditd
instance running inside a container really should only see the ACIDs
of any child containers.
Agreed. This should be easy to check and limit, preventing an auditd
from seeing any contid that is a parent of its own contid.
quoted
For example, imagine a system where the host has containers 1 and 2,
each running an auditd instance. Inside container 1 there are
containers A and B. Inside container 2 there are containers Y and Z.
If an audit event is generated in container Z, I would expect the
host's auditd to see a ACID list of "1,Z" but container 1's auditd
should only see an ACID list of "Z". The auditd running in container
2 should not see the record at all (that will be relatively
straightforward). Does that make sense? Do we have the record
formats properly designed to handle this without too much problem (I'm
not entirely sure we do)?
I completely agree and I believe we have record formats that are able to
handle this already.
I'm not convinced we do. What about the cases where we have a field
with a list of audit container IDs? How do we handle that?
I don't understand the problem. (I think you crossed your 1/2 vs
A/B/Y/Z in your example.) ...
It looks like I did, sorry about that.
quoted
... Clarifying the example above, if as you
suggest an event happens in container Z, the hosts's auditd would report
Z,^2
and the auditd in container 2 would report
Z,^2
but if there were another auditd running in container Z it would report
Z
while the auditd in container 1 or A/B would see nothing.
Yes. My concern is how do we handle this to minimize duplicating and
rewriting the records? It isn't so much about the format, although
the format is a side effect.
Are you talking about caching, or about divulging more information than
necessary or even information leaks? Or even noticing that records that
need to be generated to two audit daemons share the same contid field
values and should be generated at the same time or information shared
between them? I'd see any of these as optimizations that don't affect
the api.
Imagine a record is generated in a container which has more than one
auditd in it's ancestry that should receive this record, how do we
handle that without completely killing performance? That's my
concern. If you've already thought up a plan for this - excellent,
please share :)
No, I haven't given that much thought other than the correctness and
security issues of making sure that each audit daemon is sufficiently
isolated to do its job but not jeopardize another audit domain. Audit
already kills performance, according to some...
We currently won't have that problem since there can only be one so far.
Fixing and optimizing this is part of the next phase of the challenge of
adding a second audit daemon.
Let's work on correctness and reasonable efficiency for this phase and
not focus on a problem we don't yet have. I wouldn't consider this
incurring technical debt at this point.
I agree, one stage at a time, but the choice we make here is going to
have a significant impact on what we can do later. We need to get
this as "right" as possible; this isn't something we should dismiss
with a hand-wave as a problem for the next stage. We don't need an
implementation, but I would like to see a rough design of how we would
address this problem.
I could see cacheing a contid string from one starting point, but it may
be more work to search that cached string to truncate it or add to it
when another audit daemon requests a copy of a similar string. I
suppose every full contid string could be generated the first time it is
used and parts of it used (start/finish) as needed but that
search/indexing may not be worth it.
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
--
paul moore
www.paul-moore.com
From: Richard Guy Briggs <hidden> Date: 2020-03-19 21:48:30
On 2020-03-18 17:47, Paul Moore wrote:
On Wed, Mar 18, 2020 at 5:42 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:01, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 3:23 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:42, Paul Moore wrote:
...
quoted
quoted
The thread has had a lot of starts/stops, so I may be repeating a
previous suggestion, but one idea would be to still emit a "death
record" when the final task in the audit container ID does die, but
block the particular audit container ID from reuse until it the
SIGNAL2 info has been reported. This gives us the timely ACID death
notification while still preventing confusion and ambiguity caused by
potentially reusing the ACID before the SIGNAL2 record has been sent;
there is a small nit about the ACID being present in the SIGNAL2
*after* its death, but I think that can be easily explained and
understood by admins.
Thinking quickly about possible technical solutions to this, maybe it
makes sense to have two counters on a contobj so that we know when the
last process in that container exits and can issue the death
certificate, but we still block reuse of it until all further references
to it have been resolved. This will likely also make it possible to
report the full contid chain in SIGNAL2 records. This will eliminate
some of the issues we are discussing with regards to passing a contobj
vs a contid to the audit_log_contid function, but won't eliminate them
all because there are still some contids that won't have an object
associated with them to make it impossible to look them up in the
contobj lists.
I'm not sure you need a full second counter, I imagine a simple flag
would be okay. I think you just something to indicate that this ACID
object is marked as "dead" but it still being held for sanity reasons
and should not be reused.
Ok, I see your point. This refcount can be changed to a flag easily
enough without change to the api if we can be sure that more than one
signal can't be delivered to the audit daemon *and* collected by sig2.
I'll have a more careful look at the audit daemon code to see if I can
determine this.
Maybe I'm not understanding your concern, but this isn't really
different than any of the other things we track for the auditd signal
sender, right? If we are worried about multiple signals being sent
then it applies to everything, not just the audit container ID.
Yes, you are right. In all other cases the information is simply
overwritten. In the case of the audit container identifier any
previous value is put before a new one is referenced, so only the last
signal is kept. So, we only need a flag. Does a flag implemented with
a rcu-protected refcount sound reasonable to you?
quoted
Another question occurs to me is that what if the audit daemon is sent a
signal and it cannot or will not collect the sig2 information from the
kernel (SIGKILL?)? Does that audit container identifier remain dead
until reboot, or do we institute some other form of reaping, possibly
time-based?
In order to preserve the integrity of the audit log that ACID value
would need to remain unavailable until the ACID which contains the
associated auditd is "dead" (no one can request the signal sender's
info if that container is dead).
I don't understand why it would be associated with the contid of the
audit daemon process rather than with the audit daemon process itself.
How does the signal collection somehow get transferred or delegated to
another member of that audit daemon's container?
Thinking aloud here, the audit daemon's exit when it calls audit_free()
needs to ..._put_sig and cancel that audit_sig_cid (which in the future
will be allocated per auditd rather than the global it is now since
there is only one audit daemon).
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Richard Guy Briggs <hidden> Date: 2020-03-19 22:03:21
On 2020-03-18 18:06, Paul Moore wrote:
On Wed, Mar 18, 2020 at 5:56 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:42, Paul Moore wrote:
quoted
On Wed, Mar 18, 2020 at 5:27 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 16:56, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 2:59 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:29, Paul Moore wrote:
quoted
On Thu, Mar 12, 2020 at 3:30 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-02-13 16:44, Paul Moore wrote:
quoted
This is a bit of a thread-hijack, and for that I apologize, but
another thought crossed my mind while thinking about this issue
further ... Once we support multiple auditd instances, including the
necessary record routing and duplication/multiple-sends (the host
always sees *everything*), we will likely need to find a way to "trim"
the audit container ID (ACID) lists we send in the records. The
auditd instance running on the host/initns will always see everything,
so it will want the full container ACID list; however an auditd
instance running inside a container really should only see the ACIDs
of any child containers.
Agreed. This should be easy to check and limit, preventing an auditd
from seeing any contid that is a parent of its own contid.
quoted
For example, imagine a system where the host has containers 1 and 2,
each running an auditd instance. Inside container 1 there are
containers A and B. Inside container 2 there are containers Y and Z.
If an audit event is generated in container Z, I would expect the
host's auditd to see a ACID list of "1,Z" but container 1's auditd
should only see an ACID list of "Z". The auditd running in container
2 should not see the record at all (that will be relatively
straightforward). Does that make sense? Do we have the record
formats properly designed to handle this without too much problem (I'm
not entirely sure we do)?
I completely agree and I believe we have record formats that are able to
handle this already.
I'm not convinced we do. What about the cases where we have a field
with a list of audit container IDs? How do we handle that?
I don't understand the problem. (I think you crossed your 1/2 vs
A/B/Y/Z in your example.) ...
It looks like I did, sorry about that.
quoted
... Clarifying the example above, if as you
suggest an event happens in container Z, the hosts's auditd would report
Z,^2
and the auditd in container 2 would report
Z,^2
but if there were another auditd running in container Z it would report
Z
while the auditd in container 1 or A/B would see nothing.
Yes. My concern is how do we handle this to minimize duplicating and
rewriting the records? It isn't so much about the format, although
the format is a side effect.
Are you talking about caching, or about divulging more information than
necessary or even information leaks? Or even noticing that records that
need to be generated to two audit daemons share the same contid field
values and should be generated at the same time or information shared
between them? I'd see any of these as optimizations that don't affect
the api.
Imagine a record is generated in a container which has more than one
auditd in it's ancestry that should receive this record, how do we
handle that without completely killing performance? That's my
concern. If you've already thought up a plan for this - excellent,
please share :)
No, I haven't given that much thought other than the correctness and
security issues of making sure that each audit daemon is sufficiently
isolated to do its job but not jeopardize another audit domain. Audit
already kills performance, according to some...
We currently won't have that problem since there can only be one so far.
Fixing and optimizing this is part of the next phase of the challenge of
adding a second audit daemon.
Let's work on correctness and reasonable efficiency for this phase and
not focus on a problem we don't yet have. I wouldn't consider this
incurring technical debt at this point.
I agree, one stage at a time, but the choice we make here is going to
have a significant impact on what we can do later. We need to get
this as "right" as possible; this isn't something we should dismiss
with a hand-wave as a problem for the next stage. We don't need an
implementation, but I would like to see a rough design of how we would
address this problem.
quoted
I could see cacheing a contid string from one starting point, but it may
be more work to search that cached string to truncate it or add to it
when another audit daemon requests a copy of a similar string. I
suppose every full contid string could be generated the first time it is
used and parts of it used (start/finish) as needed but that
search/indexing may not be worth it.
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format. The contid list
only exists in the kernel. When do you suggest generating the contid
list? It sounds like you are hinting at userspace generating that list
from multiple records over the span of audit logs since boot of the
machine.
Even if we had a binary format, the current design would require
generating that list at the time of record generation since it could be
any contiguous subset of a full nested contid list.
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-20 21:56:19
On Thu, Mar 19, 2020 at 5:48 PM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-18 17:47, Paul Moore wrote:
quoted
On Wed, Mar 18, 2020 at 5:42 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:01, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 3:23 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:42, Paul Moore wrote:
...
quoted
quoted
The thread has had a lot of starts/stops, so I may be repeating a
previous suggestion, but one idea would be to still emit a "death
record" when the final task in the audit container ID does die, but
block the particular audit container ID from reuse until it the
SIGNAL2 info has been reported. This gives us the timely ACID death
notification while still preventing confusion and ambiguity caused by
potentially reusing the ACID before the SIGNAL2 record has been sent;
there is a small nit about the ACID being present in the SIGNAL2
*after* its death, but I think that can be easily explained and
understood by admins.
Thinking quickly about possible technical solutions to this, maybe it
makes sense to have two counters on a contobj so that we know when the
last process in that container exits and can issue the death
certificate, but we still block reuse of it until all further references
to it have been resolved. This will likely also make it possible to
report the full contid chain in SIGNAL2 records. This will eliminate
some of the issues we are discussing with regards to passing a contobj
vs a contid to the audit_log_contid function, but won't eliminate them
all because there are still some contids that won't have an object
associated with them to make it impossible to look them up in the
contobj lists.
I'm not sure you need a full second counter, I imagine a simple flag
would be okay. I think you just something to indicate that this ACID
object is marked as "dead" but it still being held for sanity reasons
and should not be reused.
Ok, I see your point. This refcount can be changed to a flag easily
enough without change to the api if we can be sure that more than one
signal can't be delivered to the audit daemon *and* collected by sig2.
I'll have a more careful look at the audit daemon code to see if I can
determine this.
Maybe I'm not understanding your concern, but this isn't really
different than any of the other things we track for the auditd signal
sender, right? If we are worried about multiple signals being sent
then it applies to everything, not just the audit container ID.
Yes, you are right. In all other cases the information is simply
overwritten. In the case of the audit container identifier any
previous value is put before a new one is referenced, so only the last
signal is kept. So, we only need a flag. Does a flag implemented with
a rcu-protected refcount sound reasonable to you?
Well, if I recall correctly you still need to fix the locking in this
patchset so until we see what that looks like it is hard to say for
certain. Just make sure that the flag is somehow protected from
races; it is probably a lot like the "valid" flags you sometimes see
with RCU protected lists.
quoted
quoted
Another question occurs to me is that what if the audit daemon is sent a
signal and it cannot or will not collect the sig2 information from the
kernel (SIGKILL?)? Does that audit container identifier remain dead
until reboot, or do we institute some other form of reaping, possibly
time-based?
In order to preserve the integrity of the audit log that ACID value
would need to remain unavailable until the ACID which contains the
associated auditd is "dead" (no one can request the signal sender's
info if that container is dead).
I don't understand why it would be associated with the contid of the
audit daemon process rather than with the audit daemon process itself.
How does the signal collection somehow get transferred or delegated to
another member of that audit daemon's container?
Presumably once we support multiple audit daemons we will need a
struct to contain the associated connection state, with at most one
struct (and one auditd) allowed for a given ACID. I would expect that
the signal sender info would be part of that state included in that
struct. If a task sent a signal to it's associated auditd, and no one
ever queried the signal information stored in the per-ACID state
struct, I would expect that the refcount/flag/whatever would remain
held for the signal sender's ACID until the auditd state's ACID died
(the struct would be reaped as part of the ACID death). In cases
where the container orchestrator blocks sending signals across ACID
boundaries this really isn't an issue as it will all be the same ACID,
but since we don't want to impose any restrictions on what a container
*could* be it is important to make sure we handle the case where the
signal sender's ACID may be different from the associated auditd's
ACID.
Thinking aloud here, the audit daemon's exit when it calls audit_free()
needs to ..._put_sig and cancel that audit_sig_cid (which in the future
will be allocated per auditd rather than the global it is now since
there is only one audit daemon).
quoted
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-24 00:16:55
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format.
Yes, we are. That is another topic, and another set of changes I've
been deferring so as to not disrupt the audit container ID work.
I was thinking of what we do inside the kernel between when the record
triggering event happens and when we actually emit the record to
userspace. Perhaps we collect the ACID information while the event is
occurring, but we defer generating the record until later when we have
a better understanding of what should be included in the ACID list.
It is somewhat similar (but obviously different) to what we do for
PATH records (we collect the pathname info when the path is being
resolved).
--
paul moore
www.paul-moore.com
From: Richard Guy Briggs <hidden> Date: 2020-03-24 21:02:19
On 2020-03-23 20:16, Paul Moore wrote:
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
quoted
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format.
Yes, we are. That is another topic, and another set of changes I've
been deferring so as to not disrupt the audit container ID work.
I was thinking of what we do inside the kernel between when the record
triggering event happens and when we actually emit the record to
userspace. Perhaps we collect the ACID information while the event is
occurring, but we defer generating the record until later when we have
a better understanding of what should be included in the ACID list.
It is somewhat similar (but obviously different) to what we do for
PATH records (we collect the pathname info when the path is being
resolved).
Ok, now I understand your concern.
In the case of NETFILTER_PKT records, the CONTAINER_ID record is the
only other possible record and they are generated at the same time with
a local context.
In the case of any event involving a syscall, that CONTAINER_ID record
is generated at the time of the rest of the event record generation at
syscall exit.
The others are only generated when needed, such as the sig2 reply.
We generally just store the contobj pointer until we actually generate
the CONTAINER_ID (or CONTAINER_OP) record.
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Richard Guy Briggs <hidden> Date: 2020-03-25 12:29:24
On 2020-03-20 17:56, Paul Moore wrote:
On Thu, Mar 19, 2020 at 5:48 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:47, Paul Moore wrote:
quoted
On Wed, Mar 18, 2020 at 5:42 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:01, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 3:23 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:42, Paul Moore wrote:
...
quoted
quoted
The thread has had a lot of starts/stops, so I may be repeating a
previous suggestion, but one idea would be to still emit a "death
record" when the final task in the audit container ID does die, but
block the particular audit container ID from reuse until it the
SIGNAL2 info has been reported. This gives us the timely ACID death
notification while still preventing confusion and ambiguity caused by
potentially reusing the ACID before the SIGNAL2 record has been sent;
there is a small nit about the ACID being present in the SIGNAL2
*after* its death, but I think that can be easily explained and
understood by admins.
Thinking quickly about possible technical solutions to this, maybe it
makes sense to have two counters on a contobj so that we know when the
last process in that container exits and can issue the death
certificate, but we still block reuse of it until all further references
to it have been resolved. This will likely also make it possible to
report the full contid chain in SIGNAL2 records. This will eliminate
some of the issues we are discussing with regards to passing a contobj
vs a contid to the audit_log_contid function, but won't eliminate them
all because there are still some contids that won't have an object
associated with them to make it impossible to look them up in the
contobj lists.
I'm not sure you need a full second counter, I imagine a simple flag
would be okay. I think you just something to indicate that this ACID
object is marked as "dead" but it still being held for sanity reasons
and should not be reused.
Ok, I see your point. This refcount can be changed to a flag easily
enough without change to the api if we can be sure that more than one
signal can't be delivered to the audit daemon *and* collected by sig2.
I'll have a more careful look at the audit daemon code to see if I can
determine this.
Maybe I'm not understanding your concern, but this isn't really
different than any of the other things we track for the auditd signal
sender, right? If we are worried about multiple signals being sent
then it applies to everything, not just the audit container ID.
Yes, you are right. In all other cases the information is simply
overwritten. In the case of the audit container identifier any
previous value is put before a new one is referenced, so only the last
signal is kept. So, we only need a flag. Does a flag implemented with
a rcu-protected refcount sound reasonable to you?
Well, if I recall correctly you still need to fix the locking in this
patchset so until we see what that looks like it is hard to say for
certain. Just make sure that the flag is somehow protected from
races; it is probably a lot like the "valid" flags you sometimes see
with RCU protected lists.
This is like looking for a needle in a haystack. Can you point me to
some code that does "valid" flags with RCU protected lists.
quoted
quoted
quoted
Another question occurs to me is that what if the audit daemon is sent a
signal and it cannot or will not collect the sig2 information from the
kernel (SIGKILL?)? Does that audit container identifier remain dead
until reboot, or do we institute some other form of reaping, possibly
time-based?
In order to preserve the integrity of the audit log that ACID value
would need to remain unavailable until the ACID which contains the
associated auditd is "dead" (no one can request the signal sender's
info if that container is dead).
I don't understand why it would be associated with the contid of the
audit daemon process rather than with the audit daemon process itself.
How does the signal collection somehow get transferred or delegated to
another member of that audit daemon's container?
Presumably once we support multiple audit daemons we will need a
struct to contain the associated connection state, with at most one
struct (and one auditd) allowed for a given ACID. I would expect that
the signal sender info would be part of that state included in that
struct. If a task sent a signal to it's associated auditd, and no one
ever queried the signal information stored in the per-ACID state
struct, I would expect that the refcount/flag/whatever would remain
held for the signal sender's ACID until the auditd state's ACID died
(the struct would be reaped as part of the ACID death). In cases
where the container orchestrator blocks sending signals across ACID
boundaries this really isn't an issue as it will all be the same ACID,
but since we don't want to impose any restrictions on what a container
*could* be it is important to make sure we handle the case where the
signal sender's ACID may be different from the associated auditd's
ACID.
quoted
Thinking aloud here, the audit daemon's exit when it calls audit_free()
needs to ..._put_sig and cancel that audit_sig_cid (which in the future
will be allocated per auditd rather than the global it is now since
there is only one audit daemon).
quoted
paul moore
- RGB
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-29 03:11:58
On Tue, Mar 24, 2020 at 5:02 PM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-23 20:16, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
quoted
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format.
Yes, we are. That is another topic, and another set of changes I've
been deferring so as to not disrupt the audit container ID work.
I was thinking of what we do inside the kernel between when the record
triggering event happens and when we actually emit the record to
userspace. Perhaps we collect the ACID information while the event is
occurring, but we defer generating the record until later when we have
a better understanding of what should be included in the ACID list.
It is somewhat similar (but obviously different) to what we do for
PATH records (we collect the pathname info when the path is being
resolved).
Ok, now I understand your concern.
In the case of NETFILTER_PKT records, the CONTAINER_ID record is the
only other possible record and they are generated at the same time with
a local context.
In the case of any event involving a syscall, that CONTAINER_ID record
is generated at the time of the rest of the event record generation at
syscall exit.
The others are only generated when needed, such as the sig2 reply.
We generally just store the contobj pointer until we actually generate
the CONTAINER_ID (or CONTAINER_OP) record.
Perhaps I'm remembering your latest spin of these patches incorrectly,
but there is still a big gap between when the record is generated and
when it is sent up to the audit daemon. Most importantly in that gap
is the whole big queue/multicast/unicast mess.
You don't need to show me code, but I would like to see some sort of
plan for dealing with multiple nested audit daemons. Basically I just
want to make sure we aren't painting ourselves into a corner with this
approach; and if for some horrible reason we are, I at least want us
to be aware of what we are getting ourselves into.
--
paul moore
www.paul-moore.com
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-29 03:17:33
On Wed, Mar 25, 2020 at 8:29 AM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-20 17:56, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 5:48 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:47, Paul Moore wrote:
quoted
On Wed, Mar 18, 2020 at 5:42 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:01, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 3:23 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:42, Paul Moore wrote:
...
quoted
quoted
The thread has had a lot of starts/stops, so I may be repeating a
previous suggestion, but one idea would be to still emit a "death
record" when the final task in the audit container ID does die, but
block the particular audit container ID from reuse until it the
SIGNAL2 info has been reported. This gives us the timely ACID death
notification while still preventing confusion and ambiguity caused by
potentially reusing the ACID before the SIGNAL2 record has been sent;
there is a small nit about the ACID being present in the SIGNAL2
*after* its death, but I think that can be easily explained and
understood by admins.
Thinking quickly about possible technical solutions to this, maybe it
makes sense to have two counters on a contobj so that we know when the
last process in that container exits and can issue the death
certificate, but we still block reuse of it until all further references
to it have been resolved. This will likely also make it possible to
report the full contid chain in SIGNAL2 records. This will eliminate
some of the issues we are discussing with regards to passing a contobj
vs a contid to the audit_log_contid function, but won't eliminate them
all because there are still some contids that won't have an object
associated with them to make it impossible to look them up in the
contobj lists.
I'm not sure you need a full second counter, I imagine a simple flag
would be okay. I think you just something to indicate that this ACID
object is marked as "dead" but it still being held for sanity reasons
and should not be reused.
Ok, I see your point. This refcount can be changed to a flag easily
enough without change to the api if we can be sure that more than one
signal can't be delivered to the audit daemon *and* collected by sig2.
I'll have a more careful look at the audit daemon code to see if I can
determine this.
Maybe I'm not understanding your concern, but this isn't really
different than any of the other things we track for the auditd signal
sender, right? If we are worried about multiple signals being sent
then it applies to everything, not just the audit container ID.
Yes, you are right. In all other cases the information is simply
overwritten. In the case of the audit container identifier any
previous value is put before a new one is referenced, so only the last
signal is kept. So, we only need a flag. Does a flag implemented with
a rcu-protected refcount sound reasonable to you?
Well, if I recall correctly you still need to fix the locking in this
patchset so until we see what that looks like it is hard to say for
certain. Just make sure that the flag is somehow protected from
races; it is probably a lot like the "valid" flags you sometimes see
with RCU protected lists.
This is like looking for a needle in a haystack. Can you point me to
some code that does "valid" flags with RCU protected lists.
Sigh. Come on Richard, you've been playing in the kernel for some
time now. I can't think of one off the top of my head as I write
this, but there are several resources that deal with RCU protected
lists in the kernel, Google is your friend and Documentation/RCU is
your friend.
Spending time to learn how RCU works and how to use it properly is not
time wasted. It's a tricky thing to get right (I have to refresh my
memory on some of the more subtle details each time I write/review RCU
code), but it's very cool when done correctly.
--
paul moore
www.paul-moore.com
From: Richard Guy Briggs <hidden> Date: 2020-03-30 13:47:27
On 2020-03-28 23:11, Paul Moore wrote:
On Tue, Mar 24, 2020 at 5:02 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-23 20:16, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
quoted
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format.
Yes, we are. That is another topic, and another set of changes I've
been deferring so as to not disrupt the audit container ID work.
I was thinking of what we do inside the kernel between when the record
triggering event happens and when we actually emit the record to
userspace. Perhaps we collect the ACID information while the event is
occurring, but we defer generating the record until later when we have
a better understanding of what should be included in the ACID list.
It is somewhat similar (but obviously different) to what we do for
PATH records (we collect the pathname info when the path is being
resolved).
Ok, now I understand your concern.
In the case of NETFILTER_PKT records, the CONTAINER_ID record is the
only other possible record and they are generated at the same time with
a local context.
In the case of any event involving a syscall, that CONTAINER_ID record
is generated at the time of the rest of the event record generation at
syscall exit.
The others are only generated when needed, such as the sig2 reply.
We generally just store the contobj pointer until we actually generate
the CONTAINER_ID (or CONTAINER_OP) record.
Perhaps I'm remembering your latest spin of these patches incorrectly,
but there is still a big gap between when the record is generated and
when it is sent up to the audit daemon. Most importantly in that gap
is the whole big queue/multicast/unicast mess.
So you suggest generating that record on the fly once it reaches the end
of the audit_queue just before being sent? That sounds... disruptive.
Each audit daemon is going to have its own queues, so by the time it
ends up in a particular queue, we'll already know its scope and would
have the right list of contids to print in that record.
I don't see the point in deferring the generation of the contid list
beyond the point of submitting that record to the relevant audit_queue.
You don't need to show me code, but I would like to see some sort of
plan for dealing with multiple nested audit daemons. Basically I just
want to make sure we aren't painting ourselves into a corner with this
approach; and if for some horrible reason we are, I at least want us
to be aware of what we are getting ourselves into.
It wouldn't be significantly different from what we have, but as would
have to happen for *all* records generated to a particular auditd/queue
it would have to take the scope of that auditd into account, getting
references to PIDs right for that PID namespace, along with other
similar scope views including contid list range.
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-30 14:26:55
On Mon, Mar 30, 2020 at 9:47 AM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-28 23:11, Paul Moore wrote:
quoted
On Tue, Mar 24, 2020 at 5:02 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-23 20:16, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
quoted
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format.
Yes, we are. That is another topic, and another set of changes I've
been deferring so as to not disrupt the audit container ID work.
I was thinking of what we do inside the kernel between when the record
triggering event happens and when we actually emit the record to
userspace. Perhaps we collect the ACID information while the event is
occurring, but we defer generating the record until later when we have
a better understanding of what should be included in the ACID list.
It is somewhat similar (but obviously different) to what we do for
PATH records (we collect the pathname info when the path is being
resolved).
Ok, now I understand your concern.
In the case of NETFILTER_PKT records, the CONTAINER_ID record is the
only other possible record and they are generated at the same time with
a local context.
In the case of any event involving a syscall, that CONTAINER_ID record
is generated at the time of the rest of the event record generation at
syscall exit.
The others are only generated when needed, such as the sig2 reply.
We generally just store the contobj pointer until we actually generate
the CONTAINER_ID (or CONTAINER_OP) record.
Perhaps I'm remembering your latest spin of these patches incorrectly,
but there is still a big gap between when the record is generated and
when it is sent up to the audit daemon. Most importantly in that gap
is the whole big queue/multicast/unicast mess.
So you suggest generating that record on the fly once it reaches the end
of the audit_queue just before being sent? That sounds... disruptive.
Each audit daemon is going to have its own queues, so by the time it
ends up in a particular queue, we'll already know its scope and would
have the right list of contids to print in that record.
I'm not suggesting any particular solution, I'm just pointing out a
potential problem. It isn't clear to me that you've thought about how
we generate a multiple records, each with the correct ACID list
intended for a specific audit daemon, based on a single audit event.
Explain to me how you intend that to work and we are good. Be
specific because I'm not convinced we are talking on the same plane
here.
--
paul moore
www.paul-moore.com
From: Richard Guy Briggs <hidden> Date: 2020-03-30 15:24:07
On 2020-03-28 23:17, Paul Moore wrote:
On Wed, Mar 25, 2020 at 8:29 AM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-20 17:56, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 5:48 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:47, Paul Moore wrote:
quoted
On Wed, Mar 18, 2020 at 5:42 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 17:01, Paul Moore wrote:
quoted
On Fri, Mar 13, 2020 at 3:23 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-13 12:42, Paul Moore wrote:
...
quoted
quoted
The thread has had a lot of starts/stops, so I may be repeating a
previous suggestion, but one idea would be to still emit a "death
record" when the final task in the audit container ID does die, but
block the particular audit container ID from reuse until it the
SIGNAL2 info has been reported. This gives us the timely ACID death
notification while still preventing confusion and ambiguity caused by
potentially reusing the ACID before the SIGNAL2 record has been sent;
there is a small nit about the ACID being present in the SIGNAL2
*after* its death, but I think that can be easily explained and
understood by admins.
Thinking quickly about possible technical solutions to this, maybe it
makes sense to have two counters on a contobj so that we know when the
last process in that container exits and can issue the death
certificate, but we still block reuse of it until all further references
to it have been resolved. This will likely also make it possible to
report the full contid chain in SIGNAL2 records. This will eliminate
some of the issues we are discussing with regards to passing a contobj
vs a contid to the audit_log_contid function, but won't eliminate them
all because there are still some contids that won't have an object
associated with them to make it impossible to look them up in the
contobj lists.
I'm not sure you need a full second counter, I imagine a simple flag
would be okay. I think you just something to indicate that this ACID
object is marked as "dead" but it still being held for sanity reasons
and should not be reused.
Ok, I see your point. This refcount can be changed to a flag easily
enough without change to the api if we can be sure that more than one
signal can't be delivered to the audit daemon *and* collected by sig2.
I'll have a more careful look at the audit daemon code to see if I can
determine this.
Maybe I'm not understanding your concern, but this isn't really
different than any of the other things we track for the auditd signal
sender, right? If we are worried about multiple signals being sent
then it applies to everything, not just the audit container ID.
Yes, you are right. In all other cases the information is simply
overwritten. In the case of the audit container identifier any
previous value is put before a new one is referenced, so only the last
signal is kept. So, we only need a flag. Does a flag implemented with
a rcu-protected refcount sound reasonable to you?
Well, if I recall correctly you still need to fix the locking in this
patchset so until we see what that looks like it is hard to say for
certain. Just make sure that the flag is somehow protected from
races; it is probably a lot like the "valid" flags you sometimes see
with RCU protected lists.
This is like looking for a needle in a haystack. Can you point me to
some code that does "valid" flags with RCU protected lists.
Sigh. Come on Richard, you've been playing in the kernel for some
time now. I can't think of one off the top of my head as I write
this, but there are several resources that deal with RCU protected
lists in the kernel, Google is your friend and Documentation/RCU is
your friend.
Ok, I thought you were talking about a specific piece of code...
Spending time to learn how RCU works and how to use it properly is not
time wasted. It's a tricky thing to get right (I have to refresh my
memory on some of the more subtle details each time I write/review RCU
code), but it's very cool when done correctly.
I review Documentation/RCU almost every time I work on RCU...
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Richard Guy Briggs <hidden> Date: 2020-03-30 16:22:19
On 2020-03-30 10:26, Paul Moore wrote:
On Mon, Mar 30, 2020 at 9:47 AM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-28 23:11, Paul Moore wrote:
quoted
On Tue, Mar 24, 2020 at 5:02 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-23 20:16, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
quoted
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format.
Yes, we are. That is another topic, and another set of changes I've
been deferring so as to not disrupt the audit container ID work.
I was thinking of what we do inside the kernel between when the record
triggering event happens and when we actually emit the record to
userspace. Perhaps we collect the ACID information while the event is
occurring, but we defer generating the record until later when we have
a better understanding of what should be included in the ACID list.
It is somewhat similar (but obviously different) to what we do for
PATH records (we collect the pathname info when the path is being
resolved).
Ok, now I understand your concern.
In the case of NETFILTER_PKT records, the CONTAINER_ID record is the
only other possible record and they are generated at the same time with
a local context.
In the case of any event involving a syscall, that CONTAINER_ID record
is generated at the time of the rest of the event record generation at
syscall exit.
The others are only generated when needed, such as the sig2 reply.
We generally just store the contobj pointer until we actually generate
the CONTAINER_ID (or CONTAINER_OP) record.
Perhaps I'm remembering your latest spin of these patches incorrectly,
but there is still a big gap between when the record is generated and
when it is sent up to the audit daemon. Most importantly in that gap
is the whole big queue/multicast/unicast mess.
So you suggest generating that record on the fly once it reaches the end
of the audit_queue just before being sent? That sounds... disruptive.
Each audit daemon is going to have its own queues, so by the time it
ends up in a particular queue, we'll already know its scope and would
have the right list of contids to print in that record.
I'm not suggesting any particular solution, I'm just pointing out a
potential problem. It isn't clear to me that you've thought about how
we generate a multiple records, each with the correct ACID list
intended for a specific audit daemon, based on a single audit event.
Explain to me how you intend that to work and we are good. Be
specific because I'm not convinced we are talking on the same plane
here.
Well, every time a record gets generated, *any* record gets generated,
we'll need to check for which audit daemons this record is in scope and
generate a different one for each depending on the content and whether
or not the content is influenced by the scope. Some events will be
generated for some of the auditd/queues and not for others. Some fields
in some of the records will need to be tailored for that specific
auditd/queue for either contid scope or PID namespace base reference or
other scope differences.
Every auditd/queue will need its own serial number per event and maybe
even timestamp depending on whether that auditd is in a different time
namespace and beyond that PID and contid fields and maybe others will
need to be customized per auditd/queue. So, it may make sense to
generate the contents of each field for a generic record and then either
reuse content that is unchanged or generate new content for a field that
will be different in a different auditd/queue scope, then render the
final record per auditd/queue and enqueue it.
I see this as the primary work of ghak93 ("RFE: run multiple audit
daemons on one machine"). I don't see how our proposed contid field
value format changes with this path above.
This is getting closer and closer to a netlink binary format too...
This is also an argument for spreading fields out over more record types
rather than cramming as much information as we can into one record type
(subject attributes in particular).
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-30 17:34:46
On Mon, Mar 30, 2020 at 12:22 PM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-30 10:26, Paul Moore wrote:
quoted
On Mon, Mar 30, 2020 at 9:47 AM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-28 23:11, Paul Moore wrote:
quoted
On Tue, Mar 24, 2020 at 5:02 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-23 20:16, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
quoted
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format.
Yes, we are. That is another topic, and another set of changes I've
been deferring so as to not disrupt the audit container ID work.
I was thinking of what we do inside the kernel between when the record
triggering event happens and when we actually emit the record to
userspace. Perhaps we collect the ACID information while the event is
occurring, but we defer generating the record until later when we have
a better understanding of what should be included in the ACID list.
It is somewhat similar (but obviously different) to what we do for
PATH records (we collect the pathname info when the path is being
resolved).
Ok, now I understand your concern.
In the case of NETFILTER_PKT records, the CONTAINER_ID record is the
only other possible record and they are generated at the same time with
a local context.
In the case of any event involving a syscall, that CONTAINER_ID record
is generated at the time of the rest of the event record generation at
syscall exit.
The others are only generated when needed, such as the sig2 reply.
We generally just store the contobj pointer until we actually generate
the CONTAINER_ID (or CONTAINER_OP) record.
Perhaps I'm remembering your latest spin of these patches incorrectly,
but there is still a big gap between when the record is generated and
when it is sent up to the audit daemon. Most importantly in that gap
is the whole big queue/multicast/unicast mess.
So you suggest generating that record on the fly once it reaches the end
of the audit_queue just before being sent? That sounds... disruptive.
Each audit daemon is going to have its own queues, so by the time it
ends up in a particular queue, we'll already know its scope and would
have the right list of contids to print in that record.
I'm not suggesting any particular solution, I'm just pointing out a
potential problem. It isn't clear to me that you've thought about how
we generate a multiple records, each with the correct ACID list
intended for a specific audit daemon, based on a single audit event.
Explain to me how you intend that to work and we are good. Be
specific because I'm not convinced we are talking on the same plane
here.
Well, every time a record gets generated, *any* record gets generated,
we'll need to check for which audit daemons this record is in scope and
generate a different one for each depending on the content and whether
or not the content is influenced by the scope.
That's the problem right there - we don't want to have to generate a
unique record for *each* auditd on *every* record. That is a recipe
for disaster.
Solving this for all of the known audit records is not something we
need to worry about in depth at the moment (although giving it some
casual thought is not a bad thing), but solving this for the audit
container ID information *is* something we need to worry about right
now.
--
paul moore
www.paul-moore.com
From: Richard Guy Briggs <hidden> Date: 2020-03-30 17:50:01
On 2020-03-30 13:34, Paul Moore wrote:
On Mon, Mar 30, 2020 at 12:22 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-30 10:26, Paul Moore wrote:
quoted
On Mon, Mar 30, 2020 at 9:47 AM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-28 23:11, Paul Moore wrote:
quoted
On Tue, Mar 24, 2020 at 5:02 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-23 20:16, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
quoted
I hope we can do better than string manipulations in the kernel. I'd
much rather defer generating the ACID list (if possible), than
generating a list only to keep copying and editing it as the record is
sent.
At the moment we are stuck with a string-only format.
Yes, we are. That is another topic, and another set of changes I've
been deferring so as to not disrupt the audit container ID work.
I was thinking of what we do inside the kernel between when the record
triggering event happens and when we actually emit the record to
userspace. Perhaps we collect the ACID information while the event is
occurring, but we defer generating the record until later when we have
a better understanding of what should be included in the ACID list.
It is somewhat similar (but obviously different) to what we do for
PATH records (we collect the pathname info when the path is being
resolved).
Ok, now I understand your concern.
In the case of NETFILTER_PKT records, the CONTAINER_ID record is the
only other possible record and they are generated at the same time with
a local context.
In the case of any event involving a syscall, that CONTAINER_ID record
is generated at the time of the rest of the event record generation at
syscall exit.
The others are only generated when needed, such as the sig2 reply.
We generally just store the contobj pointer until we actually generate
the CONTAINER_ID (or CONTAINER_OP) record.
Perhaps I'm remembering your latest spin of these patches incorrectly,
but there is still a big gap between when the record is generated and
when it is sent up to the audit daemon. Most importantly in that gap
is the whole big queue/multicast/unicast mess.
So you suggest generating that record on the fly once it reaches the end
of the audit_queue just before being sent? That sounds... disruptive.
Each audit daemon is going to have its own queues, so by the time it
ends up in a particular queue, we'll already know its scope and would
have the right list of contids to print in that record.
I'm not suggesting any particular solution, I'm just pointing out a
potential problem. It isn't clear to me that you've thought about how
we generate a multiple records, each with the correct ACID list
intended for a specific audit daemon, based on a single audit event.
Explain to me how you intend that to work and we are good. Be
specific because I'm not convinced we are talking on the same plane
here.
Well, every time a record gets generated, *any* record gets generated,
we'll need to check for which audit daemons this record is in scope and
generate a different one for each depending on the content and whether
or not the content is influenced by the scope.
That's the problem right there - we don't want to have to generate a
unique record for *each* auditd on *every* record. That is a recipe
for disaster.
I don't see how we can get around this.
We will already have that problem for PIDs in different PID namespaces.
We already need to use a different serial number in each auditd/queue,
or else we serialize *all* audit events on the machine and either leak
information to the nested daemons that there are other events happenning
on the machine, or confuse the host daemon because it now thinks that we
are losing events due to serial numbers missing because some nested
daemon issued an event that was not relevant to the host daemon,
consuming a globally serial audit message sequence number.
Solving this for all of the known audit records is not something we
need to worry about in depth at the moment (although giving it some
casual thought is not a bad thing), but solving this for the audit
container ID information *is* something we need to worry about right
now.
If you think that a different nested contid value string per daemon is
not acceptable, then we are back to issuing a record that has only *one*
contid listed without any nesting information. This brings us back to
the original problem of keeping *all* audit log history since the boot
of the machine to be able to track the nesting of any particular contid.
What am I missing? What do you suggest?
paul moore
- RGB
--
Richard Guy Briggs [off-list ref]
Sr. S/W Engineer, Kernel Security, Base Operating Systems
Remote, Ottawa, Red Hat Canada
IRC: rgb, SunRaycer
Voice: +1.647.777.2635, Internal: (81) 32635
From: Paul Moore <paul@paul-moore.com> Date: 2020-03-30 19:55:51
On Mon, Mar 30, 2020 at 1:49 PM Richard Guy Briggs [off-list ref] wrote:
On 2020-03-30 13:34, Paul Moore wrote:
quoted
On Mon, Mar 30, 2020 at 12:22 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-30 10:26, Paul Moore wrote:
quoted
On Mon, Mar 30, 2020 at 9:47 AM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-28 23:11, Paul Moore wrote:
quoted
On Tue, Mar 24, 2020 at 5:02 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-23 20:16, Paul Moore wrote:
quoted
On Thu, Mar 19, 2020 at 6:03 PM Richard Guy Briggs [off-list ref] wrote:
quoted
On 2020-03-18 18:06, Paul Moore wrote:
...
quoted
quoted
Well, every time a record gets generated, *any* record gets generated,
we'll need to check for which audit daemons this record is in scope and
generate a different one for each depending on the content and whether
or not the content is influenced by the scope.
That's the problem right there - we don't want to have to generate a
unique record for *each* auditd on *every* record. That is a recipe
for disaster.
I don't see how we can get around this.
We will already have that problem for PIDs in different PID namespaces.
As I said below, let's not worry about this for all of the
known/current audit records, lets just think about how we solve this
for the ACID related information.
One of the bigger problems with translating namespace info (e.g. PIDs)
across ACIDs is that an ACID - by definition - has no understanding of
namespaces (both the concept as well as any given instance).
We already need to use a different serial number in each auditd/queue,
or else we serialize *all* audit events on the machine and either leak
information to the nested daemons that there are other events happenning
on the machine, or confuse the host daemon because it now thinks that we
are losing events due to serial numbers missing because some nested
daemon issued an event that was not relevant to the host daemon,
consuming a globally serial audit message sequence number.
This isn't really relevant to the ACID lists, but sure.
quoted
Solving this for all of the known audit records is not something we
need to worry about in depth at the moment (although giving it some
casual thought is not a bad thing), but solving this for the audit
container ID information *is* something we need to worry about right
now.
If you think that a different nested contid value string per daemon is
not acceptable, then we are back to issuing a record that has only *one*
contid listed without any nesting information. This brings us back to
the original problem of keeping *all* audit log history since the boot
of the machine to be able to track the nesting of any particular contid.
I'm not ruling anything out, except for the "let's just completely
regenerate every record for each auditd instance".
What am I missing? What do you suggest?
I'm missing a solution in this thread, since you are the person
driving this effort I'm asking you to get creative and present us with
some solutions. :)
--
paul moore
www.paul-moore.com