From: Mike Kazantsev <hidden> Date: 2012-10-19 14:51:02
Good day,
There seem to be a large slab memory leak in standard (kernel.org)
kernel with the specific configuration and workload I have here.
From what I can tell at the moment, it appears to be a leak in
IPSec-related xfrm code.
It was really noticeable on several different physical machines
with same kernel configuration but different worloads since I've
upgraded to kernel 3.5.0.
Graph of total slab usage (+ total available RAM) on these machines:
http://i.imgur.com/IyPqA.png
Presence of some leak can clearly be seen over time, and it caused
near-OOM condition several times now.
Sharp drops in memory usage indicates reboot, which, I'm afraid, with
such condition, has to be done at the regular intervals.
Initially I thought that it was triggered by heavy filesystem load, but
today finally got around to reboot one of the machines with
slub_debug=U and it doesn't seem to be the case.
slabtop showed "kmalloc-64" being the 99% offender in the past, but
with recent kernels (3.6.1), it has changed to "secpath_cache",
alloc_calls in /sys/kernel/slab/secpath_cache/ lists only the following:
2779138 secpath_dup+0x1b/0x5a age=400/169538/326767 pid=0-1543 cpus=0-3
And free_calls lists these two lines:
2543886 <not-available> age=4295223985 pid=0 cpus=0
235252 __secpath_destroy+0x3e/0x43 age=1651/174629/327902 pid=0-1519 cpus=0-3
Contents of all paths available in /sys/kernel/slab/secpath_cache/
and "slabtop -o" output should be attached to this mail.
These were taken after heavy network + fs i/o load (rsync from a
different machine over network) after ~10-20min.
"secpath_dup" seem to be ipsec-related call, and all machines in
question communicate over IPSec almost exclusively all the time
(openswan-2.6.37 userspace at the moment).
As noted, the problem is highly reproducible - all I have to do is to
run rsync or something similar between these nodes for a few minutes.
All machines in question have x86_64 kernel 3.6.1 now, but I'll
probably update it to 3.6.2 in a moment.
Keywords:
linux kernel networking mm slub slab secpath_dup secpath_cache xfrm
ipsec 3.5 3.6 memory leak oom slabtop x86 x86_64 amd64
/proc/version:
Linux version 3.6.1-fg.mf_master (root@anathema) (gcc version 4.6.3
(Exherbo gcc-4.6.3-r1) ) #1 SMP Sat Oct 13 04:21:08 YEKT 2012
Other information about the system (as per REPORTING-BUGS) is
attached, also including slabtop and slub_debug-related /sys paths
output/contents.
--
Mike Kazantsev // fraggod.net
From: Paul Moore <paul@paul-moore.com> Date: 2012-10-20 12:42:34
Thanks for the problem report. I'm not going to be in a position to start
looking into this until late Sunday, but hopefully it will be a quick fix.
Two quick questions (my apologies, I'm not able to dig through your logs
right now): do you see this leak on kernels < 3.5.0, and are you using any
labeled IPsec connections?
--
paul moore
www.paul-moore.com
From: Mike Kazantsev <hidden> Date: 2012-10-20 14:50:09
On Sat, 20 Oct 2012 08:42:33 -0400
Paul Moore [off-list ref] wrote:
Thanks for the problem report. I'm not going to be in a position to start
looking into this until late Sunday, but hopefully it will be a quick fix.
Two quick questions (my apologies, I'm not able to dig through your logs
right now): do you see this leak on kernels < 3.5.0, and are you using any
labeled IPsec connections?
As I understand, labelled connections are only used in SELinux
and SMACK LSM, which are not enabled (in Kconfig, i.e. not built) in any
of the kernels I use.
The only LSM I have enabled (and actually use on 2/4 of these machines)
is AppArmor, and though I think it doesn't attach any labels to network
connections yet (there's a "Wishlist" bug at
https://bugs.launchpad.net/ubuntu/+source/apparmor/+bug/796588, but I
can't seem to find an existing implementation).
I believe it has started with 3.5.0, according to all available logs I
have. I'm afraid laziness and other tasks have prevented me from
looking into and reporting the issue back then, but memory graph trends
start at the exact time of reboot into 3.5.0 kernels, and before that,
there're no such trends for slab memory usage.
I've been able to ignore and work around the problem for months now, so
I don't think there's any rush at all ;)
But that said, currently I've started git bisect process between v3.5
and v3.4 tags, so hopefully I'll get good-enough results of it before
you'll get to it (probably in a few hours to a few days).
Also, I've found that switching to "slab" allocator from "slub" doesn't
help the problem at all, so I guess something doesn't get freed in the
code indeed, though I hasn't been able to find anything relevant in the
logs for the sources where secpath_put and secpath_dup are used, and
decided to try bisect.
--
Mike Kazantsev // fraggod.net
From: Mike Kazantsev <hidden> Date: 2012-10-20 22:45:51
On Sat, 20 Oct 2012 20:49:58 +0600
Mike Kazantsev [off-list ref] wrote:
On Sat, 20 Oct 2012 08:42:33 -0400
Paul Moore [off-list ref] wrote:
quoted
Thanks for the problem report. I'm not going to be in a position to start
looking into this until late Sunday, but hopefully it will be a quick fix.
Two quick questions (my apologies, I'm not able to dig through your logs
right now): do you see this leak on kernels < 3.5.0, and are you using any
labeled IPsec connections?
As I understand, labelled connections are only used in SELinux
and SMACK LSM, which are not enabled (in Kconfig, i.e. not built) in any
of the kernels I use.
The only LSM I have enabled (and actually use on 2/4 of these machines)
is AppArmor, and though I think it doesn't attach any labels to network
connections yet (there's a "Wishlist" bug at
https://bugs.launchpad.net/ubuntu/+source/apparmor/+bug/796588, but I
can't seem to find an existing implementation).
I believe it has started with 3.5.0, according to all available logs I
have. I'm afraid laziness and other tasks have prevented me from
looking into and reporting the issue back then, but memory graph trends
start at the exact time of reboot into 3.5.0 kernels, and before that,
there're no such trends for slab memory usage.
I've been able to ignore and work around the problem for months now, so
I don't think there's any rush at all ;)
But that said, currently I've started git bisect process between v3.5
and v3.4 tags, so hopefully I'll get good-enough results of it before
you'll get to it (probably in a few hours to a few days).
Also, I've found that switching to "slab" allocator from "slub" doesn't
help the problem at all, so I guess something doesn't get freed in the
code indeed, though I hasn't been able to find anything relevant in the
logs for the sources where secpath_put and secpath_dup are used, and
decided to try bisect.
Sorry for yet another mail on the weekend, but I've finished the bisect
and here is the result:
a1c7fff7e18f59e684e07b0f9a770561cd39f395 is the first bad commit
commit a1c7fff7e18f59e684e07b0f9a770561cd39f395
Author: Eric Dumazet [off-list ref]
Date: Thu May 17 07:34:16 2012 +0000
net: netdev_alloc_skb() use build_skb()
netdev_alloc_skb() is used by networks driver in their RX path to
allocate an skb to receive an incoming frame.
With recent skb->head_frag infrastructure, it makes sense to change
netdev_alloc_skb() to use build_skb() and a frag allocator.
This permits a zero copy splice(socket->pipe), and better GRO or TCP
coalescing.
Signed-off-by: Eric Dumazet [off-list ref]
Signed-off-by: David S. Miller [off-list ref]
:040000 040000 17938b1b46bc38aa126cc23b7a7647259297657d 1e29cf65869391eb13552c51e0cf288fc7085fec M net
No skips, all "good" / "bad" decisions were very unambiguous and easy
to make - secpath_cache slabs either stayed at always-constant 20K
cumulative size (~5 of them) and were reported as 10-15% full in "good"
case, or were 99% full and eating memory at hudreds KiB/s (during same
rsync transfer) in "bad" case.
Reverting that commit in 3.6.2 kernel looks like a bad idea and doesn't
seem possible to do cleanly.
Being not a C coder and having only faint idea about how things should
be done with regards to socket buffers, I can't seem to find anything
to tweak based on that commit either.
kmemleak mechanism seem to provide stack traces and interesting calls
for debugging of whatever is allocating the non-freed objects, so guess
I'll see if I can get more definitive (to my ignorant eye) "look here"
hint from it, and might drop one more mail with data from there.
--
Mike Kazantsev // fraggod.net
From: Mike Kazantsev <hidden> Date: 2012-10-21 00:24:14
On Sun, 21 Oct 2012 04:45:40 +0600
Mike Kazantsev [off-list ref] wrote:
kmemleak mechanism seem to provide stack traces and interesting calls
for debugging of whatever is allocating the non-freed objects, so guess
I'll see if I can get more definitive (to my ignorant eye) "look here"
hint from it, and might drop one more mail with data from there.
kmemleak finds a lot (dozens megabytes of stack traces) of identical
paths leading to a leaks:
(for IPv6 packets)
unreferenced object 0xffff88002fa25b00 (size 56):
comm "softirq", pid 0, jiffies 4295009073 (age 295.620s)
hex dump (first 32 bytes):
01 00 00 00 01 00 00 00 00 fc 6e 30 00 88 ff ff ..........n0....
6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b kkkkkkkkkkkkkkkk
backtrace:
[<ffffffff814cfa2b>] kmemleak_alloc+0x21/0x3e
[<ffffffff810d9445>] kmem_cache_alloc+0xa5/0xb1
[<ffffffff8147dd35>] secpath_dup+0x1b/0x5a
[<ffffffff8147df39>] xfrm_input+0x64/0x484
[<ffffffff814b1d2c>] xfrm6_rcv_spi+0x19/0x1b
[<ffffffff814b1d4e>] xfrm6_rcv+0x20/0x22
[<ffffffff8148c19f>] ip6_input_finish+0x203/0x31b
[<ffffffff8148c622>] ip6_input+0x1e/0x50
[<ffffffff8148c31c>] ip6_rcv_finish+0x65/0x69
[<ffffffff8148c5a3>] ipv6_rcv+0x283/0x2e4
[<ffffffff813ff8ba>] __netif_receive_skb+0x599/0x64c
[<ffffffff813ffb08>] netif_receive_skb+0x47/0x78
[<ffffffff81400644>] napi_skb_finish+0x21/0x53
[<ffffffff81400778>] napi_gro_receive+0x102/0x10e
[<ffffffff8136978b>] rtl8169_poll+0x326/0x4f9
[<ffffffff813ffcda>] net_rx_action+0x9f/0x175
(for IPv4 packets)
unreferenced object 0xffff88003387e000 (size 56):
comm "softirq", pid 0, jiffies 4294915803 (age 563.583s)
hex dump (first 32 bytes):
01 00 00 00 01 00 00 00 00 48 be 30 00 88 ff ff .........H.0....
6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b kkkkkkkkkkkkkkkk
backtrace:
[<ffffffff814cfa2b>] kmemleak_alloc+0x21/0x3e
[<ffffffff810d9445>] kmem_cache_alloc+0xa5/0xb1
[<ffffffff8147dd35>] secpath_dup+0x1b/0x5a
[<ffffffff8147df39>] xfrm_input+0x64/0x484
[<ffffffff81474f7b>] xfrm4_rcv_encap+0x17/0x19
[<ffffffff81474f9c>] xfrm4_rcv+0x1f/0x21
[<ffffffff81430514>] ip_local_deliver_finish+0x170/0x22a
[<ffffffff81430706>] ip_local_deliver+0x46/0x78
[<ffffffff8143038d>] ip_rcv_finish+0x2bd/0x2d4
[<ffffffff81430969>] ip_rcv+0x231/0x28c
[<ffffffff813ff8ba>] __netif_receive_skb+0x599/0x64c
[<ffffffff813ffb08>] netif_receive_skb+0x47/0x78
[<ffffffff81400644>] napi_skb_finish+0x21/0x53
[<ffffffff81400778>] napi_gro_receive+0x102/0x10e
[<ffffffff8136978b>] rtl8169_poll+0x326/0x4f9
[<ffffffff813ffcda>] net_rx_action+0x9f/0x175
Object at the top and trace seem to be the same (between same
IP-family) everywhere, just ages and addresses are different.
IPv6 usage seem to be one important detail which I failed to mention.
IPv4 traces seem to be really rare (only several of them), but that
might be understandable because rsync was ran over IPv6.
Still wasn't able to figure out what might cause the get's/put's
disbalance with that commit, but was able to revert it, without
anything bad happening (so far), using the patch below (in case
issue might bite someone else before proper fix is found).
--
From: Eric Dumazet <hidden> Date: 2012-10-21 13:29:47
On Sun, 2012-10-21 at 06:24 +0600, Mike Kazantsev wrote:
quoted hunk
On Sun, 21 Oct 2012 04:45:40 +0600
Mike Kazantsev [off-list ref] wrote:
quoted
kmemleak mechanism seem to provide stack traces and interesting calls
for debugging of whatever is allocating the non-freed objects, so guess
I'll see if I can get more definitive (to my ignorant eye) "look here"
hint from it, and might drop one more mail with data from there.
kmemleak finds a lot (dozens megabytes of stack traces) of identical
paths leading to a leaks:
(for IPv6 packets)
unreferenced object 0xffff88002fa25b00 (size 56):
comm "softirq", pid 0, jiffies 4295009073 (age 295.620s)
hex dump (first 32 bytes):
01 00 00 00 01 00 00 00 00 fc 6e 30 00 88 ff ff ..........n0....
6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b kkkkkkkkkkkkkkkk
backtrace:
[<ffffffff814cfa2b>] kmemleak_alloc+0x21/0x3e
[<ffffffff810d9445>] kmem_cache_alloc+0xa5/0xb1
[<ffffffff8147dd35>] secpath_dup+0x1b/0x5a
[<ffffffff8147df39>] xfrm_input+0x64/0x484
[<ffffffff814b1d2c>] xfrm6_rcv_spi+0x19/0x1b
[<ffffffff814b1d4e>] xfrm6_rcv+0x20/0x22
[<ffffffff8148c19f>] ip6_input_finish+0x203/0x31b
[<ffffffff8148c622>] ip6_input+0x1e/0x50
[<ffffffff8148c31c>] ip6_rcv_finish+0x65/0x69
[<ffffffff8148c5a3>] ipv6_rcv+0x283/0x2e4
[<ffffffff813ff8ba>] __netif_receive_skb+0x599/0x64c
[<ffffffff813ffb08>] netif_receive_skb+0x47/0x78
[<ffffffff81400644>] napi_skb_finish+0x21/0x53
[<ffffffff81400778>] napi_gro_receive+0x102/0x10e
[<ffffffff8136978b>] rtl8169_poll+0x326/0x4f9
[<ffffffff813ffcda>] net_rx_action+0x9f/0x175
(for IPv4 packets)
unreferenced object 0xffff88003387e000 (size 56):
comm "softirq", pid 0, jiffies 4294915803 (age 563.583s)
hex dump (first 32 bytes):
01 00 00 00 01 00 00 00 00 48 be 30 00 88 ff ff .........H.0....
6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b 6b kkkkkkkkkkkkkkkk
backtrace:
[<ffffffff814cfa2b>] kmemleak_alloc+0x21/0x3e
[<ffffffff810d9445>] kmem_cache_alloc+0xa5/0xb1
[<ffffffff8147dd35>] secpath_dup+0x1b/0x5a
[<ffffffff8147df39>] xfrm_input+0x64/0x484
[<ffffffff81474f7b>] xfrm4_rcv_encap+0x17/0x19
[<ffffffff81474f9c>] xfrm4_rcv+0x1f/0x21
[<ffffffff81430514>] ip_local_deliver_finish+0x170/0x22a
[<ffffffff81430706>] ip_local_deliver+0x46/0x78
[<ffffffff8143038d>] ip_rcv_finish+0x2bd/0x2d4
[<ffffffff81430969>] ip_rcv+0x231/0x28c
[<ffffffff813ff8ba>] __netif_receive_skb+0x599/0x64c
[<ffffffff813ffb08>] netif_receive_skb+0x47/0x78
[<ffffffff81400644>] napi_skb_finish+0x21/0x53
[<ffffffff81400778>] napi_gro_receive+0x102/0x10e
[<ffffffff8136978b>] rtl8169_poll+0x326/0x4f9
[<ffffffff813ffcda>] net_rx_action+0x9f/0x175
Object at the top and trace seem to be the same (between same
IP-family) everywhere, just ages and addresses are different.
IPv6 usage seem to be one important detail which I failed to mention.
IPv4 traces seem to be really rare (only several of them), but that
might be understandable because rsync was ran over IPv6.
Still wasn't able to figure out what might cause the get's/put's
disbalance with that commit, but was able to revert it, without
anything bad happening (so far), using the patch below (in case
issue might bite someone else before proper fix is found).
--
Did you try linux-3.7-rc2 (or linux-3.7-rc1) ?
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Mike Kazantsev <hidden> Date: 2012-10-21 18:43:44
On Sun, 21 Oct 2012 19:57:01 +0600
Mike Kazantsev [off-list ref] wrote:
On Sun, 21 Oct 2012 15:29:43 +0200
Eric Dumazet [off-list ref] wrote:
quoted
Did you try linux-3.7-rc2 (or linux-3.7-rc1) ?
I did not, will do in a few hours, thanks for the pointer.
I just built "torvalds/linux-2.6" (v3.7-rc2) and rebooted into it,
started same rsync-over-net test and got kmalloc-64 leaking (it went up
to tens of MiB until I stopped rsync, normally these are fixed at ~500
KiB).
Unfortunately, I forgot to add slub_debug option and build kmemleak so
wasn't able to look at this case further, and when I rebooted with
these enabled/built, it was secpath_cache again.
So previously noted "slabtop showed 'kmalloc-64' being the 99% offender
in the past, but with recent kernels (3.6.1), it has changed to
'secpath_cache'" seem to be incorrect, as it seem to depend not on
kernel version, but some other factor.
Guess I'll try to reboot a few more times to see if I can catch
kmalloc-64 leaking (instead of secpath_cache) again.
--
Mike Kazantsev // fraggod.net
From: Mike Kazantsev <hidden> Date: 2012-10-21 19:51:45
On Mon, 22 Oct 2012 00:43:32 +0600
Mike Kazantsev [off-list ref] wrote:
quoted
On Sun, 21 Oct 2012 15:29:43 +0200
Eric Dumazet [off-list ref] wrote:
quoted
Did you try linux-3.7-rc2 (or linux-3.7-rc1) ?
I just built "torvalds/linux-2.6" (v3.7-rc2) and rebooted into it,
started same rsync-over-net test and got kmalloc-64 leaking (it went up
to tens of MiB until I stopped rsync, normally these are fixed at ~500
KiB).
Unfortunately, I forgot to add slub_debug option and build kmemleak so
wasn't able to look at this case further, and when I rebooted with
these enabled/built, it was secpath_cache again.
So previously noted "slabtop showed 'kmalloc-64' being the 99% offender
in the past, but with recent kernels (3.6.1), it has changed to
'secpath_cache'" seem to be incorrect, as it seem to depend not on
kernel version, but some other factor.
Guess I'll try to reboot a few more times to see if I can catch
kmalloc-64 leaking (instead of secpath_cache) again.
I haven't been able to catch the aforementioned condition, but noticed
that with v3.7-rc2, "hex dump" part seem to vary in kmemleak
traces, and contain all sorts of random stuff, for example:
unreferenced object 0xffff88002ae2de00 (size 56):
comm "softirq", pid 0, jiffies 4295006317 (age 213.066s)
hex dump (first 32 bytes):
01 00 00 00 01 00 00 00 20 9f f4 28 00 88 ff ff ........ ..(....
2f 6f 72 67 2f 66 72 65 65 64 65 73 6b 74 6f 70 /org/freedesktop
backtrace:
[<ffffffff814da4e3>] kmemleak_alloc+0x21/0x3e
[<ffffffff810dc1f7>] kmem_cache_alloc+0xa5/0xb1
[<ffffffff81487bf1>] secpath_dup+0x1b/0x5a
[<ffffffff81487df5>] xfrm_input+0x64/0x484
[<ffffffff814bbd70>] xfrm6_rcv_spi+0x19/0x1b
[<ffffffff814bbd92>] xfrm6_rcv+0x20/0x22
[<ffffffff814960c3>] ip6_input_finish+0x203/0x31b
[<ffffffff81496542>] ip6_input+0x1e/0x50
[<ffffffff81496240>] ip6_rcv_finish+0x65/0x69
[<ffffffff814964c3>] ipv6_rcv+0x27f/0x2e0
[<ffffffff8140a659>] __netif_receive_skb+0x5ba/0x65a
[<ffffffff8140a894>] netif_receive_skb+0x47/0x78
[<ffffffff8140b4bf>] napi_skb_finish+0x21/0x54
[<ffffffff8140b5ef>] napi_gro_receive+0xfd/0x10a
[<ffffffff81372b47>] rtl8169_poll+0x326/0x4fc
[<ffffffff8140ad44>] net_rx_action+0x9f/0x188
Not sure if it's relevant though.
--
Mike Kazantsev // fraggod.net
From: Eric Dumazet <hidden> Date: 2012-10-21 21:47:37
On Mon, 2012-10-22 at 01:51 +0600, Mike Kazantsev wrote:
On Mon, 22 Oct 2012 00:43:32 +0600
Mike Kazantsev [off-list ref] wrote:
quoted
quoted
On Sun, 21 Oct 2012 15:29:43 +0200
Eric Dumazet [off-list ref] wrote:
quoted
Did you try linux-3.7-rc2 (or linux-3.7-rc1) ?
I just built "torvalds/linux-2.6" (v3.7-rc2) and rebooted into it,
started same rsync-over-net test and got kmalloc-64 leaking (it went up
to tens of MiB until I stopped rsync, normally these are fixed at ~500
KiB).
Unfortunately, I forgot to add slub_debug option and build kmemleak so
wasn't able to look at this case further, and when I rebooted with
these enabled/built, it was secpath_cache again.
So previously noted "slabtop showed 'kmalloc-64' being the 99% offender
in the past, but with recent kernels (3.6.1), it has changed to
'secpath_cache'" seem to be incorrect, as it seem to depend not on
kernel version, but some other factor.
Guess I'll try to reboot a few more times to see if I can catch
kmalloc-64 leaking (instead of secpath_cache) again.
I haven't been able to catch the aforementioned condition, but noticed
that with v3.7-rc2, "hex dump" part seem to vary in kmemleak
traces, and contain all sorts of random stuff, for example:
unreferenced object 0xffff88002ae2de00 (size 56):
comm "softirq", pid 0, jiffies 4295006317 (age 213.066s)
hex dump (first 32 bytes):
01 00 00 00 01 00 00 00 20 9f f4 28 00 88 ff ff ........ ..(....
2f 6f 72 67 2f 66 72 65 65 64 65 73 6b 74 6f 70 /org/freedesktop
backtrace:
[<ffffffff814da4e3>] kmemleak_alloc+0x21/0x3e
[<ffffffff810dc1f7>] kmem_cache_alloc+0xa5/0xb1
[<ffffffff81487bf1>] secpath_dup+0x1b/0x5a
[<ffffffff81487df5>] xfrm_input+0x64/0x484
[<ffffffff814bbd70>] xfrm6_rcv_spi+0x19/0x1b
[<ffffffff814bbd92>] xfrm6_rcv+0x20/0x22
[<ffffffff814960c3>] ip6_input_finish+0x203/0x31b
[<ffffffff81496542>] ip6_input+0x1e/0x50
[<ffffffff81496240>] ip6_rcv_finish+0x65/0x69
[<ffffffff814964c3>] ipv6_rcv+0x27f/0x2e0
[<ffffffff8140a659>] __netif_receive_skb+0x5ba/0x65a
[<ffffffff8140a894>] netif_receive_skb+0x47/0x78
[<ffffffff8140b4bf>] napi_skb_finish+0x21/0x54
[<ffffffff8140b5ef>] napi_gro_receive+0xfd/0x10a
[<ffffffff81372b47>] rtl8169_poll+0x326/0x4fc
[<ffffffff8140ad44>] net_rx_action+0x9f/0x188
Not sure if it's relevant though.
OK, so some layer seems to have a bug if the skb->head is exactly
allocated, instead of having extra tailroom (because of kmalloc-powerof2
alignment)
Or some layer overwrites past skb->cb[] array
If you try to move sp field in sk_buff, does it change something ?
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Mike Kazantsev <hidden> Date: 2012-10-21 22:58:57
On Sun, 21 Oct 2012 23:47:33 +0200
Eric Dumazet [off-list ref] wrote:
OK, so some layer seems to have a bug if the skb->head is exactly
allocated, instead of having extra tailroom (because of kmalloc-powerof2
alignment)
Or some layer overwrites past skb->cb[] array
If you try to move sp field in sk_buff, does it change something ?
...
Also try to increase tailroom in __netdev_alloc_skb()
Applied both patches, but unfortunately, the problem seem to be still
there.
This time the leaking objects seem to show up as kmalloc-64.
OBJS ACTIVE USE OBJ SIZE SLABS OBJ/SLAB CACHE SIZE NAME
266760 265333 99% 0.30K 10260 26 82080K kmemleak_object
157440 157440 100% 0.06K 2460 64 9840K kmalloc-64
94458 94458 100% 0.10K 2422 39 9688K buffer_head
27573 27573 100% 0.19K 1313 21 5252K dentry
kmemleak traces:
unreferenced object 0xffff88002f38ec80 (size 64):
comm "softirq", pid 0, jiffies 4294900815 (age 142.346s)
hex dump (first 32 bytes):
01 00 00 00 01 00 00 00 00 08 03 2e 00 88 ff ff ................
2b 6f a0 ca 28 b2 4a f1 0a 74 33 74 5a 76 18 cb +o..(.J..t3tZv..
backtrace:
[<ffffffff814da4e3>] kmemleak_alloc+0x21/0x3e
[<ffffffff810dc1f7>] kmem_cache_alloc+0xa5/0xb1
[<ffffffff81487bf5>] secpath_dup+0x1b/0x5a
[<ffffffff81487df9>] xfrm_input+0x64/0x484
[<ffffffff8147eec3>] xfrm4_rcv_encap+0x17/0x19
[<ffffffff8147eee4>] xfrm4_rcv+0x1f/0x21
[<ffffffff8143b4e4>] ip_local_deliver_finish+0x170/0x22a
[<ffffffff8143b6d6>] ip_local_deliver+0x46/0x78
[<ffffffff8143b35d>] ip_rcv_finish+0x295/0x2ac
[<ffffffff8143b936>] ip_rcv+0x22e/0x288
[<ffffffff8140a65d>] __netif_receive_skb+0x5ba/0x65a
[<ffffffff8140a898>] netif_receive_skb+0x47/0x78
[<ffffffff8140b4c3>] napi_skb_finish+0x21/0x54
[<ffffffff8140b5f3>] napi_gro_receive+0xfd/0x10a
[<ffffffff81372b47>] rtl8169_poll+0x326/0x4fc
[<ffffffff8140ad48>] net_rx_action+0x9f/0x188
unreferenced object 0xffff880029b47580 (size 64):
comm "softirq", pid 0, jiffies 4294926900 (age 143.946s)
hex dump (first 32 bytes):
01 00 00 00 01 00 00 00 00 88 07 2e 00 88 ff ff ................
00 00 00 00 2f 6f 72 67 2f 66 72 65 65 64 65 73 ..../org/freedes
backtrace:
[<ffffffff814da4e3>] kmemleak_alloc+0x21/0x3e
[<ffffffff810dc1f7>] kmem_cache_alloc+0xa5/0xb1
[<ffffffff81487bf5>] secpath_dup+0x1b/0x5a
[<ffffffff81487df9>] xfrm_input+0x64/0x484
[<ffffffff814bbd74>] xfrm6_rcv_spi+0x19/0x1b
[<ffffffff814bbd96>] xfrm6_rcv+0x20/0x22
[<ffffffff814960c7>] ip6_input_finish+0x203/0x31b
[<ffffffff81496546>] ip6_input+0x1e/0x50
[<ffffffff81496244>] ip6_rcv_finish+0x65/0x69
[<ffffffff814964c7>] ipv6_rcv+0x27f/0x2e0
[<ffffffff8140a65d>] __netif_receive_skb+0x5ba/0x65a
[<ffffffff8140a898>] netif_receive_skb+0x47/0x78
[<ffffffff8140b4c3>] napi_skb_finish+0x21/0x54
[<ffffffff8140b5f3>] napi_gro_receive+0xfd/0x10a
[<ffffffff81372b47>] rtl8169_poll+0x326/0x4fc
[<ffffffff8140ad48>] net_rx_action+0x9f/0x188
I've grepped for "/org/free" specifically and sure enough, same scraps
of data seem to be in some of the (varied) dumps there.
--
Mike Kazantsev // fraggod.net
@@ -2977,6 +2977,9 @@ int netif_rx(struct sk_buff *skb){intret;+#ifdef CONFIG_XFRM+WARN_ON_ONCE(skb->sp);+#endif/* if netpoll wants it, pretend we never saw it */if(netpoll_rx(skb))returnNET_RX_DROP;
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Mike Kazantsev <hidden> Date: 2012-10-22 12:07:06
On Mon, 22 Oct 2012 10:15:43 +0200
Eric Dumazet [off-list ref] wrote:
On Mon, 2012-10-22 at 04:58 +0600, Mike Kazantsev wrote:
quoted
I've grepped for "/org/free" specifically and sure enough, same scraps
of data seem to be in some of the (varied) dumps there.
Content is not meaningful, as we dont initialize it.
So you see previous content.
Could you try the following :
...
With this patch on top of v3.7-rc2 (w/o patches from your previous
mail), leak seem to be still present.
If I understand correctly, WARN_ON_ONCE should've produced some output
in dmesg when the conditions passed to it were met.
They don't appear to be, as the only output in dmesg during
ipsec-related modules loading (I think openswan probes them manually)
is still "AVX instructions are not detected" (can be seen in tty on
boot) and the only post-boot dmesg output (incl. during leaks
happening) is from kmemleak ("kmemleak: ... new suspected memory
leaks").
Looks like kmem_cache_zalloc got rid of the content, though traces
still report it as "kmem_cache_alloc", but I guess it's because of its
"inline" nature.
--
Mike Kazantsev // fraggod.net
@@ -48,6 +48,7 @@#include<linux/inet.h>#include<linux/netfilter_ipv4.h>#include<net/inet_ecn.h>+#include<net/xfrm.h>/* NOTE. Logic of IP defragmentation is parallel to corresponding IPv6*codenow.Ifyouchangesomethinghere,_PLEASE_updateipv6/reassembly.c
From: Eric Dumazet <hidden> Date: 2012-10-22 15:22:22
On Mon, 2012-10-22 at 17:16 +0200, Eric Dumazet wrote:
OK, I believe I found the bug in IPv4 defrag / IPv6 reasm
Please test the following patch.
Thanks !
I'll send a more generic patch in a few minutes, changing
kfree_skb_partial() to call skb_release_head_state()
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Mike Kazantsev <hidden> Date: 2012-10-22 16:59:30
On Mon, 22 Oct 2012 17:28:02 +0200
Eric Dumazet [off-list ref] wrote:
On Mon, 2012-10-22 at 17:22 +0200, Eric Dumazet wrote:
quoted
On Mon, 2012-10-22 at 17:16 +0200, Eric Dumazet wrote:
quoted
OK, I believe I found the bug in IPv4 defrag / IPv6 reasm
Please test the following patch.
Thanks !
I'll send a more generic patch in a few minutes, changing
kfree_skb_partial() to call skb_release_head_state()
Here it is :
...
Problem is indeed gone in v3.7-rc2 with the proposed generic patch, I
haven't read the mail in time to test the first one, but I guess it's
not relevant now that the latter one works.
Thank you for taking your time to look into the problem and actually
fix it.
I'm unclear about policies in place on the matter, but I think this
patch might be a good candidate to backport into 3.5 and 3.6 kernels,
because they seem to suffer from the issue as well.
--
Mike Kazantsev // fraggod.net
From: Eric Dumazet <hidden> Date: 2012-10-22 17:24:13
On Mon, 2012-10-22 at 22:59 +0600, Mike Kazantsev wrote:
On Mon, 22 Oct 2012 17:28:02 +0200
Eric Dumazet [off-list ref] wrote:
quoted
On Mon, 2012-10-22 at 17:22 +0200, Eric Dumazet wrote:
quoted
On Mon, 2012-10-22 at 17:16 +0200, Eric Dumazet wrote:
quoted
OK, I believe I found the bug in IPv4 defrag / IPv6 reasm
Please test the following patch.
Thanks !
I'll send a more generic patch in a few minutes, changing
kfree_skb_partial() to call skb_release_head_state()
Here it is :
...
Problem is indeed gone in v3.7-rc2 with the proposed generic patch, I
haven't read the mail in time to test the first one, but I guess it's
not relevant now that the latter one works.
Thank you for taking your time to look into the problem and actually
fix it.
I'm unclear about policies in place on the matter, but I think this
patch might be a good candidate to backport into 3.5 and 3.6 kernels,
because they seem to suffer from the issue as well.
Thanks a lot Mike for your help.
Dont worry, I'll submit an official patch with details and all credits.
David Miller will forward it to stable teams.
Thanks !
From: Eric Dumazet <hidden> Date: 2012-10-22 19:03:47
From: Eric Dumazet <redacted>
Mike Kazantsev found 3.5 kernels and beyond were leaking memory,
and tracked the faulty commit to a1c7fff7e18f59e (net:
netdev_alloc_skb() use build_skb()
While this commit seems fine, it uncovered a bug introduced
in commit bad43ca8325 (net: introduce skb_try_coalesce()), in function
kfree_skb_partial() :
If head is stolen, we free the sk_buff,
without removing references on secpath (skb->sp).
So IPsec + IP defrag/reassembly (using skb coalescing), or
TCP coalescing could leak secpath objects.
Fix this bug by calling skb_release_head_state(skb) to properly
release all possible references to linked objects.
Reported-by: Mike Kazantsev <redacted>
Signed-off-by: Eric Dumazet <redacted>
Bisected-by: Mike Kazantsev [off-list ref]
Tested-by: Mike Kazantsev <redacted>
---
It seems TCP stack could immediately release secpath references instead
of waiting skb are eaten by consumer, thats will be a followup patch.
net/core/skbuff.c | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: David Miller <davem@davemloft.net> Date: 2012-10-22 19:18:03
From: Eric Dumazet <redacted>
Date: Mon, 22 Oct 2012 21:03:40 +0200
From: Eric Dumazet <redacted>
Mike Kazantsev found 3.5 kernels and beyond were leaking memory,
and tracked the faulty commit to a1c7fff7e18f59e (net:
netdev_alloc_skb() use build_skb()
While this commit seems fine, it uncovered a bug introduced
in commit bad43ca8325 (net: introduce skb_try_coalesce()), in function
kfree_skb_partial() :
If head is stolen, we free the sk_buff,
without removing references on secpath (skb->sp).
So IPsec + IP defrag/reassembly (using skb coalescing), or
TCP coalescing could leak secpath objects.
Fix this bug by calling skb_release_head_state(skb) to properly
release all possible references to linked objects.
Reported-by: Mike Kazantsev <redacted>
Signed-off-by: Eric Dumazet <redacted>
Bisected-by: Mike Kazantsev [off-list ref]
Tested-by: Mike Kazantsev <redacted>
Applied and queued up for -stable, thanks!
It seems TCP stack could immediately release secpath references instead
of waiting skb are eaten by consumer, thats will be a followup patch.
Indeed.
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>