From: Alexander Potapenko <glider@google.com> Date: 2019-05-14 14:36:13
Add sock_alloc_send_pskb_noinit(), which is similar to
sock_alloc_send_pskb(), but allocates with __GFP_NO_AUTOINIT.
This helps reduce the slowdown on hackbench in the init_on_alloc mode
from 6.84% to 3.45%.
Slowdown for the initialization features compared to init_on_free=0,
init_on_alloc=0:
hackbench, init_on_free=1: +7.71% sys time (st.err 0.45%)
hackbench, init_on_alloc=1: +3.45% sys time (st.err 0.86%)
Linux build with -j12, init_on_free=1: +8.34% wall time (st.err 0.39%)
Linux build with -j12, init_on_free=1: +24.13% sys time (st.err 0.47%)
Linux build with -j12, init_on_alloc=1: -0.04% wall time (st.err 0.46%)
Linux build with -j12, init_on_alloc=1: +0.50% sys time (st.err 0.45%)
The slowdown for init_on_free=0, init_on_alloc=0 compared to the
baseline is within the standard error.
Signed-off-by: Alexander Potapenko <glider@google.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: Christoph Lameter <redacted>
To: Kees Cook <redacted>
Cc: Masahiro Yamada <redacted>
Cc: James Morris <jmorris@namei.org>
Cc: "Serge E. Hallyn" <serge@hallyn.com>
Cc: Nick Desaulniers <ndesaulniers@google.com>
Cc: Kostya Serebryany <redacted>
Cc: Dmitry Vyukov <dvyukov@google.com>
Cc: Sandeep Patil <redacted>
Cc: Laura Abbott <redacted>
Cc: Randy Dunlap <rdunlap@infradead.org>
Cc: Jann Horn <jannh@google.com>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: linux-mm@kvack.org
Cc: linux-security-module@vger.kernel.org
Cc: kernel-hardening@lists.openwall.com
---
v2:
- changed __GFP_NOINIT to __GFP_NO_AUTOINIT
---
include/net/sock.h | 5 +++++
net/core/sock.c | 29 +++++++++++++++++++++++++----
net/unix/af_unix.c | 13 +++++++------
3 files changed, 37 insertions(+), 10 deletions(-)
On Tue, May 14, 2019 at 04:35:37PM +0200, Alexander Potapenko wrote:
Add sock_alloc_send_pskb_noinit(), which is similar to
sock_alloc_send_pskb(), but allocates with __GFP_NO_AUTOINIT.
This helps reduce the slowdown on hackbench in the init_on_alloc mode
from 6.84% to 3.45%.
Out of curiosity, why the creation of the new function over adding a
gfp flag argument to sock_alloc_send_pskb() and updating callers? (There
are only 6 callers, and this change already updates 2 of those.)
Slowdown for the initialization features compared to init_on_free=0,
init_on_alloc=0:
hackbench, init_on_free=1: +7.71% sys time (st.err 0.45%)
hackbench, init_on_alloc=1: +3.45% sys time (st.err 0.86%)
In the commit log it might be worth mentioning that this is only
changing the init_on_alloc case (in case it's not already obvious to
folks). Perhaps there needs to be a split of __GFP_NO_AUTOINIT into
__GFP_NO_AUTO_ALLOC_INIT and __GFP_NO_AUTO_FREE_INIT? Right now __GFP_NO_AUTOINIT is only checked for init_on_alloc:
static inline bool want_init_on_alloc(gfp_t flags)
{
if (static_branch_unlikely(&init_on_alloc))
return !(flags & __GFP_NO_AUTOINIT);
return flags & __GFP_ZERO;
}
...
static inline bool want_init_on_free(void)
{
return static_branch_unlikely(&init_on_free);
}
On a related note, it might be nice to add an exclusion list to
the kmem_cache_create() cases, since it seems likely that further
tuning will be needed there. For example, with the init_on_free-similar
PAX_MEMORY_SANITIZE changes in the last public release of PaX/grsecurity,
the following were excluded from wipe-on-free:
buffer_head
names_cache
mm_struct
vm_area_struct
anon_vma
anon_vma_chain
skbuff_head_cache
skbuff_fclone_cache
Adding these and others (with details on why they were selected),
might improve init_on_free performance further without trading too
much coverage.
Having a kernel param with a comma-separated list of cache names and
the logic to add __GFP_NO_AUTOINIT at creation time would be a nice
(and cheap!) debug feature to let folks tune things for their specific
workloads, if they choose to. (And it could maybe also know what "none"
meant, to actually remove the built-in exclusions, similar to what
PaX's "pax_sanitize_slab=full" does.)
--
Kees Cook
On Thu, May 16, 2019 at 09:53:01AM -0700, Kees Cook wrote:
On Tue, May 14, 2019 at 04:35:37PM +0200, Alexander Potapenko wrote:
quoted
Add sock_alloc_send_pskb_noinit(), which is similar to
sock_alloc_send_pskb(), but allocates with __GFP_NO_AUTOINIT.
This helps reduce the slowdown on hackbench in the init_on_alloc mode
from 6.84% to 3.45%.
Out of curiosity, why the creation of the new function over adding a
gfp flag argument to sock_alloc_send_pskb() and updating callers? (There
are only 6 callers, and this change already updates 2 of those.)
quoted
Slowdown for the initialization features compared to init_on_free=0,
init_on_alloc=0:
hackbench, init_on_free=1: +7.71% sys time (st.err 0.45%)
hackbench, init_on_alloc=1: +3.45% sys time (st.err 0.86%)
So I've run some of my own wall-clock timings of kernel builds (which
should be an pretty big "worst case" situation, and I see much smaller
performance changes:
everything off
Run times: 289.18 288.61 289.66 287.71 287.67
Min: 287.67 Max: 289.66 Mean: 288.57 Std Dev: 0.79
baseline
init_on_alloc=1
Run times: 289.72 286.95 287.87 287.34 287.35
Min: 286.95 Max: 289.72 Mean: 287.85 Std Dev: 0.98
0.25% faster (within the std dev noise)
init_on_free=1
Run times: 303.26 301.44 301.19 301.55 301.39
Min: 301.19 Max: 303.26 Mean: 301.77 Std Dev: 0.75
4.57% slower
init_on_free=1 with the PAX_MEMORY_SANITIZE slabs excluded:
Run times: 299.19 299.85 298.95 298.23 298.64
Min: 298.23 Max: 299.85 Mean: 298.97 Std Dev: 0.55
3.60% slower
So the tuning certainly improved things by 1%. My perf numbers don't
show the 24% hit you were seeing at all, though.
In the commit log it might be worth mentioning that this is only
changing the init_on_alloc case (in case it's not already obvious to
folks). Perhaps there needs to be a split of __GFP_NO_AUTOINIT into
__GFP_NO_AUTO_ALLOC_INIT and __GFP_NO_AUTO_FREE_INIT? Right now
__GFP_NO_AUTOINIT is only checked for init_on_alloc:
I was obviously crazy here. :) GFP isn't present for free(), but a SLAB
flag works (as was done in PAX_MEMORY_SANITIZE). I'll send the patch I
used for the above timing test.
--
Kees Cook
In order to improve the init_on_free performance, some frequently
freed caches with less sensitive contents can be excluded from the
init_on_free behavior.
This patch is modified from Brad Spengler/PaX Team's code in the
last public patch of grsecurity/PaX based on my understanding of the
code. Changes or omissions from the original code are mine and don't
reflect the original grsecurity/PaX code.
Signed-off-by: Kees Cook <redacted>
---
fs/buffer.c | 2 +-
fs/dcache.c | 3 ++-
include/linux/slab.h | 3 +++
kernel/fork.c | 6 ++++--
mm/rmap.c | 5 +++--
mm/slab.h | 5 +++--
net/core/skbuff.c | 6 ++++--
7 files changed, 20 insertions(+), 10 deletions(-)
From: Alexander Potapenko <glider@google.com> Date: 2019-05-17 08:34:40
On Fri, May 17, 2019 at 2:50 AM Kees Cook [off-list ref] wrote:
In order to improve the init_on_free performance, some frequently
freed caches with less sensitive contents can be excluded from the
init_on_free behavior.
Did you see any notable performance improvement with this patch?
A similar one gave me only 1-2% on the parallel Linux build.
quoted hunk
This patch is modified from Brad Spengler/PaX Team's code in the
last public patch of grsecurity/PaX based on my understanding of the
code. Changes or omissions from the original code are mine and don't
reflect the original grsecurity/PaX code.
Signed-off-by: Kees Cook <redacted>
---
fs/buffer.c | 2 +-
fs/dcache.c | 3 ++-
include/linux/slab.h | 3 +++
kernel/fork.c | 6 ++++--
mm/rmap.c | 5 +++--
mm/slab.h | 5 +++--
net/core/skbuff.c | 6 ++++--
7 files changed, 20 insertions(+), 10 deletions(-)
--
Alexander Potapenko
Software Engineer
Google Germany GmbH
Erika-Mann-Straße, 33
80636 München
Geschäftsführer: Paul Manicle, Halimah DeLaine Prado
Registergericht und -nummer: Hamburg, HRB 86891
Sitz der Gesellschaft: Hamburg
From: Alexander Potapenko <glider@google.com> Date: 2019-05-17 08:49:17
On Fri, May 17, 2019 at 2:26 AM Kees Cook [off-list ref] wrote:
On Thu, May 16, 2019 at 09:53:01AM -0700, Kees Cook wrote:
quoted
On Tue, May 14, 2019 at 04:35:37PM +0200, Alexander Potapenko wrote:
quoted
Add sock_alloc_send_pskb_noinit(), which is similar to
sock_alloc_send_pskb(), but allocates with __GFP_NO_AUTOINIT.
This helps reduce the slowdown on hackbench in the init_on_alloc mode
from 6.84% to 3.45%.
Out of curiosity, why the creation of the new function over adding a
gfp flag argument to sock_alloc_send_pskb() and updating callers? (There
are only 6 callers, and this change already updates 2 of those.)
quoted
Slowdown for the initialization features compared to init_on_free=0,
init_on_alloc=0:
hackbench, init_on_free=1: +7.71% sys time (st.err 0.45%)
hackbench, init_on_alloc=1: +3.45% sys time (st.err 0.86%)
So I've run some of my own wall-clock timings of kernel builds (which
should be an pretty big "worst case" situation, and I see much smaller
performance changes:
How many cores were you using? I suspect the numbers may vary a bit
depending on that.
everything off
Run times: 289.18 288.61 289.66 287.71 287.67
Min: 287.67 Max: 289.66 Mean: 288.57 Std Dev: 0.79
baseline
init_on_alloc=1
Run times: 289.72 286.95 287.87 287.34 287.35
Min: 286.95 Max: 289.72 Mean: 287.85 Std Dev: 0.98
0.25% faster (within the std dev noise)
init_on_free=1
Run times: 303.26 301.44 301.19 301.55 301.39
Min: 301.19 Max: 303.26 Mean: 301.77 Std Dev: 0.75
4.57% slower
init_on_free=1 with the PAX_MEMORY_SANITIZE slabs excluded:
Run times: 299.19 299.85 298.95 298.23 298.64
Min: 298.23 Max: 299.85 Mean: 298.97 Std Dev: 0.55
3.60% slower
So the tuning certainly improved things by 1%. My perf numbers don't
show the 24% hit you were seeing at all, though.
Note that 24% is the _sys_ time slowdown. The wall time slowdown seen
in this case was 8.34%
quoted
In the commit log it might be worth mentioning that this is only
changing the init_on_alloc case (in case it's not already obvious to
folks). Perhaps there needs to be a split of __GFP_NO_AUTOINIT into
__GFP_NO_AUTO_ALLOC_INIT and __GFP_NO_AUTO_FREE_INIT? Right now
__GFP_NO_AUTOINIT is only checked for init_on_alloc:
I was obviously crazy here. :) GFP isn't present for free(), but a SLAB
flag works (as was done in PAX_MEMORY_SANITIZE). I'll send the patch I
used for the above timing test.
--
Kees Cook
--
Alexander Potapenko
Software Engineer
Google Germany GmbH
Erika-Mann-Straße, 33
80636 München
Geschäftsführer: Paul Manicle, Halimah DeLaine Prado
Registergericht und -nummer: Hamburg, HRB 86891
Sitz der Gesellschaft: Hamburg
From: Alexander Potapenko <glider@google.com> Date: 2019-05-17 13:51:00
On Fri, May 17, 2019 at 10:49 AM Alexander Potapenko [off-list ref] wrote:
On Fri, May 17, 2019 at 2:26 AM Kees Cook [off-list ref] wrote:
quoted
On Thu, May 16, 2019 at 09:53:01AM -0700, Kees Cook wrote:
quoted
On Tue, May 14, 2019 at 04:35:37PM +0200, Alexander Potapenko wrote:
quoted
Add sock_alloc_send_pskb_noinit(), which is similar to
sock_alloc_send_pskb(), but allocates with __GFP_NO_AUTOINIT.
This helps reduce the slowdown on hackbench in the init_on_alloc mode
from 6.84% to 3.45%.
Out of curiosity, why the creation of the new function over adding a
gfp flag argument to sock_alloc_send_pskb() and updating callers? (There
are only 6 callers, and this change already updates 2 of those.)
quoted
Slowdown for the initialization features compared to init_on_free=0,
init_on_alloc=0:
hackbench, init_on_free=1: +7.71% sys time (st.err 0.45%)
hackbench, init_on_alloc=1: +3.45% sys time (st.err 0.86%)
So I've run some of my own wall-clock timings of kernel builds (which
should be an pretty big "worst case" situation, and I see much smaller
performance changes:
How many cores were you using? I suspect the numbers may vary a bit
depending on that.
quoted
everything off
Run times: 289.18 288.61 289.66 287.71 287.67
Min: 287.67 Max: 289.66 Mean: 288.57 Std Dev: 0.79
baseline
init_on_alloc=1
Run times: 289.72 286.95 287.87 287.34 287.35
Min: 286.95 Max: 289.72 Mean: 287.85 Std Dev: 0.98
0.25% faster (within the std dev noise)
init_on_free=1
Run times: 303.26 301.44 301.19 301.55 301.39
Min: 301.19 Max: 303.26 Mean: 301.77 Std Dev: 0.75
4.57% slower
init_on_free=1 with the PAX_MEMORY_SANITIZE slabs excluded:
Run times: 299.19 299.85 298.95 298.23 298.64
Min: 298.23 Max: 299.85 Mean: 298.97 Std Dev: 0.55
3.60% slower
So the tuning certainly improved things by 1%. My perf numbers don't
show the 24% hit you were seeing at all, though.
Note that 24% is the _sys_ time slowdown. The wall time slowdown seen
in this case was 8.34%
I've collected more stats running QEMU with different numbers of cores.
The slowdown values of init_on_free compared to baseline are:
2 CPUs - 5.94% for wall time (20.08% for sys time)
6 CPUs - 7.43% for wall time (23.55% for sys time)
12 CPUs - 8.41% for wall time (24.25% for sys time)
24 CPUs - 9.49% for wall time (17.98% for sys time)
I'm building a defconfig of some fixed KMSAN tree with Clang, but that
shouldn't matter much.
quoted
quoted
In the commit log it might be worth mentioning that this is only
changing the init_on_alloc case (in case it's not already obvious to
folks). Perhaps there needs to be a split of __GFP_NO_AUTOINIT into
__GFP_NO_AUTO_ALLOC_INIT and __GFP_NO_AUTO_FREE_INIT? Right now
__GFP_NO_AUTOINIT is only checked for init_on_alloc:
I was obviously crazy here. :) GFP isn't present for free(), but a SLAB
flag works (as was done in PAX_MEMORY_SANITIZE). I'll send the patch I
used for the above timing test.
--
Kees Cook
--
Alexander Potapenko
Software Engineer
Google Germany GmbH
Erika-Mann-Straße, 33
80636 München
Geschäftsführer: Paul Manicle, Halimah DeLaine Prado
Registergericht und -nummer: Hamburg, HRB 86891
Sitz der Gesellschaft: Hamburg
--
Alexander Potapenko
Software Engineer
Google Germany GmbH
Erika-Mann-Straße, 33
80636 München
Geschäftsführer: Paul Manicle, Halimah DeLaine Prado
Registergericht und -nummer: Hamburg, HRB 86891
Sitz der Gesellschaft: Hamburg
On Fri, May 17, 2019 at 10:34:26AM +0200, Alexander Potapenko wrote:
On Fri, May 17, 2019 at 2:50 AM Kees Cook [off-list ref] wrote:
quoted
In order to improve the init_on_free performance, some frequently
freed caches with less sensitive contents can be excluded from the
init_on_free behavior.
Did you see any notable performance improvement with this patch?
A similar one gave me only 1-2% on the parallel Linux build.
Yup, that's in the other thread. I saw similar. But 1-2% on a 5% hit is
a lot. ;)
--
Kees Cook
On Fri, May 17, 2019 at 10:49:03AM +0200, Alexander Potapenko wrote:
On Fri, May 17, 2019 at 2:26 AM Kees Cook [off-list ref] wrote:
quoted
On Thu, May 16, 2019 at 09:53:01AM -0700, Kees Cook wrote:
quoted
On Tue, May 14, 2019 at 04:35:37PM +0200, Alexander Potapenko wrote:
quoted
Add sock_alloc_send_pskb_noinit(), which is similar to
sock_alloc_send_pskb(), but allocates with __GFP_NO_AUTOINIT.
This helps reduce the slowdown on hackbench in the init_on_alloc mode
from 6.84% to 3.45%.
Out of curiosity, why the creation of the new function over adding a
gfp flag argument to sock_alloc_send_pskb() and updating callers? (There
are only 6 callers, and this change already updates 2 of those.)
quoted
Slowdown for the initialization features compared to init_on_free=0,
init_on_alloc=0:
hackbench, init_on_free=1: +7.71% sys time (st.err 0.45%)
hackbench, init_on_alloc=1: +3.45% sys time (st.err 0.86%)
So I've run some of my own wall-clock timings of kernel builds (which
should be an pretty big "worst case" situation, and I see much smaller
performance changes:
How many cores were you using? I suspect the numbers may vary a bit
depending on that.
I was using 4.
quoted
init_on_alloc=1
Run times: 289.72 286.95 287.87 287.34 287.35
Min: 286.95 Max: 289.72 Mean: 287.85 Std Dev: 0.98
0.25% faster (within the std dev noise)
init_on_free=1
Run times: 303.26 301.44 301.19 301.55 301.39
Min: 301.19 Max: 303.26 Mean: 301.77 Std Dev: 0.75
4.57% slower
init_on_free=1 with the PAX_MEMORY_SANITIZE slabs excluded:
Run times: 299.19 299.85 298.95 298.23 298.64
Min: 298.23 Max: 299.85 Mean: 298.97 Std Dev: 0.55
3.60% slower
So the tuning certainly improved things by 1%. My perf numbers don't
show the 24% hit you were seeing at all, though.
Note that 24% is the _sys_ time slowdown. The wall time slowdown seen
in this case was 8.34%
Ah! Gotcha. Yeah, seems the impact for init_on_free is pretty
variable. The init_on_alloc appears close to free, though.
--
Kees Cook
Hi Kees,
On Fri, 17 May 2019 at 02:50, Kees Cook [off-list ref] wrote:
In order to improve the init_on_free performance, some frequently
freed caches with less sensitive contents can be excluded from the
init_on_free behavior.
This patch is modified from Brad Spengler/PaX Team's code in the
last public patch of grsecurity/PaX based on my understanding of the
code. Changes or omissions from the original code are mine and don't
reflect the original grsecurity/PaX code.
On Mon, May 20, 2019 at 08:10:19AM +0200, Mathias Krause wrote:
Hi Kees,
On Fri, 17 May 2019 at 02:50, Kees Cook [off-list ref] wrote:
quoted
In order to improve the init_on_free performance, some frequently
freed caches with less sensitive contents can be excluded from the
init_on_free behavior.
This patch is modified from Brad Spengler/PaX Team's code in the
last public patch of grsecurity/PaX based on my understanding of the
code. Changes or omissions from the original code are mine and don't
reflect the original grsecurity/PaX code.
Ah-ha! Thanks for finding the specific commit; I'll adjust
attribution. Are you able to describe how you chose the various excluded
kmem caches? (I assume it wasn't just "highest numbers in the stats
reporting".) And why run the ctor after wipe? Doesn't that means you're
just running the ctor again at the next allocation time?
Thanks!
-Kees