sctp_close/sk_free: kernel BUG at slab.c:3074!

17 messages, 7 authors, 2012-09-06 · open the first message on its own page

sctp_close/sk_free: kernel BUG at slab.c:3074!

From: Fengguang Wu <hidden>
Date: 2012-09-04 13:59:37

Greetings,

The below oops happens somewhere in tree 

        git://gitorious.org/linux-can/linux-can-next master

Bisect has been time consuming and still in half way..

[  233.046014] kfree_debugcheck: out of range ptr ea6000000bb8h.
[  233.047399] ------------[ cut here ]------------
[  233.048393] kernel BUG at /c/kernel-tests/src/stable/mm/slab.c:3074!
[  233.048393] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC
[  233.048393] Modules linked in:
[  233.048393] CPU 0 
[  233.048393] Pid: 3929, comm: trinity-watchdo Not tainted 3.6.0-rc3+ #4192 Bochs Bochs
[  233.048393] RIP: 0010:[<ffffffff81169653>]  [<ffffffff81169653>] kfree_debugcheck+0x27/0x2d
[  233.048393] RSP: 0018:ffff88000facbca8  EFLAGS: 00010092
[  233.048393] RAX: 0000000000000031 RBX: 0000ea6000000bb8 RCX: 00000000a189a188
[  233.048393] RDX: 000000000000a189 RSI: ffffffff8108ad32 RDI: ffffffff810d30f9
[  233.048393] RBP: ffff88000facbcb8 R08: 0000000000000002 R09: ffffffff843846f0
[  233.048393] R10: ffffffff810ae37c R11: 0000000000000908 R12: 0000000000000202
[  233.048393] R13: ffffffff823dbd5a R14: ffff88000ec5bea8 R15: ffffffff8363c780
[  233.048393] FS:  00007faa6899c700(0000) GS:ffff88001f200000(0000) knlGS:0000000000000000
[  233.048393] CS:  0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[  233.048393] CR2: 00007faa6841019c CR3: 0000000012c82000 CR4: 00000000000006f0
[  233.048393] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[  233.048393] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[  233.048393] Process trinity-watchdo (pid: 3929, threadinfo ffff88000faca000, task ffff88000faec600)
[  233.048393] Stack:
[  233.048393]  0000000000000000 0000ea6000000bb8 ffff88000facbce8 ffffffff8116ad81
[  233.048393]  ffff88000ff588a0 ffff88000ff58850 ffff88000ff588a0 0000000000000000
[  233.048393]  ffff88000facbd08 ffffffff823dbd5a ffffffff823dbcb0 ffff88000ff58850
[  233.048393] Call Trace:
[  233.048393]  [<ffffffff8116ad81>] kfree+0x5f/0xca
[  233.048393]  [<ffffffff823dbd5a>] inet_sock_destruct+0xaa/0x13c
[  233.048393]  [<ffffffff823dbcb0>] ? inet_sk_rebuild_header+0x319/0x319
[  233.048393]  [<ffffffff8231c307>] __sk_free+0x21/0x14b
[  233.048393]  [<ffffffff8231c4bd>] sk_free+0x26/0x2a
[  233.048393]  [<ffffffff825372db>] sctp_close+0x215/0x224
[  233.048393]  [<ffffffff810d6835>] ? lock_release+0x16f/0x1b9
[  233.048393]  [<ffffffff823daf12>] inet_release+0x7e/0x85
[  233.048393]  [<ffffffff82317d15>] sock_release+0x1f/0x77
[  233.048393]  [<ffffffff82317d94>] sock_close+0x27/0x2b
[  233.048393]  [<ffffffff81173bbe>] __fput+0x101/0x20a
[  233.048393]  [<ffffffff81173cd5>] ____fput+0xe/0x10
[  233.048393]  [<ffffffff810a3794>] task_work_run+0x5d/0x75
[  233.048393]  [<ffffffff8108da70>] do_exit+0x290/0x7f5
[  233.048393]  [<ffffffff82707415>] ? retint_swapgs+0x13/0x1b
[  233.048393]  [<ffffffff8108e23f>] do_group_exit+0x7b/0xba
[  233.048393]  [<ffffffff8108e295>] sys_exit_group+0x17/0x17
[  233.048393]  [<ffffffff8270de10>] tracesys+0xdd/0xe2
[  233.048393] Code: 59 01 5d c3 55 48 89 e5 53 41 50 0f 1f 44 00 00 48 89 fb e8 d4 b0 f0 ff 84 c0 75 11 48 89 de 48 c7 c7 fc fa f7 82 e8 0d 0f 57 01 <0f> 0b 5f 5b 5d c3 55 48 89 e5 0f 1f 44 00 00 48 63 87 d8 00 00 
[  233.048393] RIP  [<ffffffff81169653>] kfree_debugcheck+0x27/0x2d
[  233.048393]  RSP <ffff88000facbca8>

Thanks,
Fengguang

sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Fengguang Wu <hidden>
Date: 2012-09-04 14:04:11

FYI, another kconfig triggering a slightly different oops on tree

        git://gitorious.org/linux-can/linux-can-next led-trigger

[   96.267311] ------------[ cut here ]------------
[   96.268294] kernel BUG at /c/kernel-tests/src/stable/arch/x86/mm/physaddr.c:18!
[   96.269988] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC
[   96.270636] Modules linked in:
[   96.270636] CPU 0 
[   96.270636] Pid: 2116, comm: trinity Not tainted 3.6.0-rc3+ #2679 Bochs Bochs
[   96.270636] RIP: 0010:[<ffffffff8102b22b>]  [<ffffffff8102b22b>] __phys_addr+0x46/0x6b
[   96.270636] RSP: 0018:ffff880019585c98  EFLAGS: 00010213
[   96.270636] RAX: ffff87ffffffffff RBX: 0000ea6000000bb8 RCX: 0000000000000000
[   96.270636] RDX: 0000000000000000 RSI: 0000000000000296 RDI: 0000ea6000000bb8
[   96.270636] RBP: ffff880019585c98 R08: 0000000000000058 R09: 0000000000000008
[   96.270636] R10: 000000000000000a R11: 0000000000000058 R12: ffff8800195f7718
[   96.270636] R13: ffffffff816521cf R14: ffffea0000000000 R15: 0000000000000000
[   96.270636] FS:  00007fa19b534700(0000) GS:ffff88001f200000(0000) knlGS:0000000000000000
[   96.270636] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[   96.270636] CR2: 00007fa19b03eba0 CR3: 000000001957b000 CR4: 00000000000006f0
[   96.270636] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[   96.270636] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[   96.270636] Process trinity (pid: 2116, threadinfo ffff880019584000, task ffff88001af2c680)
[   96.270636] Stack:
[   96.270636]  ffff880019585cd8 ffffffff811091d7 0000000000000000 ffff88001b1ef200
[   96.270636]  ffff88001b1ef4d0 0000000000000000 ffff88001b1617b0 0000000000000000
[   96.270636]  ffff880019585cf8 ffffffff816521cf ffff88001b1ef200 ffff88001b1ef248
[   96.270636] Call Trace:
[   96.270636]  [<ffffffff811091d7>] kfree+0x63/0x162
[   96.270636]  [<ffffffff816521cf>] inet_sock_destruct+0x112/0x1ca
[   96.270636]  [<ffffffff815f6fa4>] __sk_free+0x1d/0x114
[   96.270636]  [<ffffffff815f710b>] sk_free+0x1c/0x1e
[   96.270636]  [<ffffffff816d59d5>] sctp_close+0x21a/0x229
[   96.270636]  [<ffffffff810810f6>] ? lock_release_holdtime.part.6+0xb2/0xb7
[   96.270636]  [<ffffffff81651b3e>] ? inet_release+0x65/0xc3
[   96.270636]  [<ffffffff81651b93>] inet_release+0xba/0xc3
[   96.270636]  [<ffffffff81651af9>] ? inet_release+0x20/0xc3
[   96.270636]  [<ffffffff81674134>] inet6_release+0x30/0x3c
[   96.270636]  [<ffffffff815f2317>] sock_release+0x1f/0x77
[   96.270636]  [<ffffffff815f2396>] sock_close+0x27/0x2b
[   96.270636]  [<ffffffff8110ec22>] __fput+0xf0/0x24b
[   96.270636]  [<ffffffff8110ed8b>] ____fput+0xe/0x10
[   96.270636]  [<ffffffff8104f370>] task_work_run+0x5d/0x75
[   96.270636]  [<ffffffff81038a66>] do_exit+0x26b/0x7d7
[   96.270636]  [<ffffffff81725a95>] ? retint_swapgs+0x13/0x1b
[   96.270636]  [<ffffffff8103925b>] do_group_exit+0x7b/0xba
[   96.270636]  [<ffffffff810392b1>] sys_exit_group+0x17/0x17
[   96.270636]  [<ffffffff8172c78e>] tracesys+0xd0/0xd5
[   96.270636] Code: 00 80 48 01 c7 48 81 ff ff ff ff 1f 76 02 0f 0b 48 89 f8 48 03 05 f6 bd ae 00 eb 32 48 b8 ff ff ff ff ff 87 ff ff 48 39 c7 77 02 <0f> 0b 0f b6 0d 55 57 ba 00 48 b8 00 00 00 00 00 78 00 00 48 01 
[   96.270636] RIP  [<ffffffff8102b22b>] __phys_addr+0x46/0x6b
[   96.270636]  RSP <ffff880019585c98>

Thanks,
Fengguang

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Marc Kleine-Budde <mkl@pengutronix.de>
Date: 2012-09-04 17:11:00

On 09/04/2012 04:04 PM, Fengguang Wu wrote:
FYI, another kconfig triggering a slightly different oops on tree

        git://gitorious.org/linux-can/linux-can-next led-trigger
This in turn means the problem doesn't come from the CAN patches, as
both trees have different CAN patches. I'm adding Eric W. Biederman on
Cc as he contributed some sctp patches between v3.6 and net-next/master.

Marc
[   96.267311] ------------[ cut here ]------------
[   96.268294] kernel BUG at /c/kernel-tests/src/stable/arch/x86/mm/physaddr.c:18!
[   96.269988] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC
[   96.270636] Modules linked in:
[   96.270636] CPU 0 
[   96.270636] Pid: 2116, comm: trinity Not tainted 3.6.0-rc3+ #2679 Bochs Bochs
[   96.270636] RIP: 0010:[<ffffffff8102b22b>]  [<ffffffff8102b22b>] __phys_addr+0x46/0x6b
[   96.270636] RSP: 0018:ffff880019585c98  EFLAGS: 00010213
[   96.270636] RAX: ffff87ffffffffff RBX: 0000ea6000000bb8 RCX: 0000000000000000
[   96.270636] RDX: 0000000000000000 RSI: 0000000000000296 RDI: 0000ea6000000bb8
[   96.270636] RBP: ffff880019585c98 R08: 0000000000000058 R09: 0000000000000008
[   96.270636] R10: 000000000000000a R11: 0000000000000058 R12: ffff8800195f7718
[   96.270636] R13: ffffffff816521cf R14: ffffea0000000000 R15: 0000000000000000
[   96.270636] FS:  00007fa19b534700(0000) GS:ffff88001f200000(0000) knlGS:0000000000000000
[   96.270636] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[   96.270636] CR2: 00007fa19b03eba0 CR3: 000000001957b000 CR4: 00000000000006f0
[   96.270636] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[   96.270636] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[   96.270636] Process trinity (pid: 2116, threadinfo ffff880019584000, task ffff88001af2c680)
[   96.270636] Stack:
[   96.270636]  ffff880019585cd8 ffffffff811091d7 0000000000000000 ffff88001b1ef200
[   96.270636]  ffff88001b1ef4d0 0000000000000000 ffff88001b1617b0 0000000000000000
[   96.270636]  ffff880019585cf8 ffffffff816521cf ffff88001b1ef200 ffff88001b1ef248
[   96.270636] Call Trace:
[   96.270636]  [<ffffffff811091d7>] kfree+0x63/0x162
[   96.270636]  [<ffffffff816521cf>] inet_sock_destruct+0x112/0x1ca
[   96.270636]  [<ffffffff815f6fa4>] __sk_free+0x1d/0x114
[   96.270636]  [<ffffffff815f710b>] sk_free+0x1c/0x1e
[   96.270636]  [<ffffffff816d59d5>] sctp_close+0x21a/0x229
[   96.270636]  [<ffffffff810810f6>] ? lock_release_holdtime.part.6+0xb2/0xb7
[   96.270636]  [<ffffffff81651b3e>] ? inet_release+0x65/0xc3
[   96.270636]  [<ffffffff81651b93>] inet_release+0xba/0xc3
[   96.270636]  [<ffffffff81651af9>] ? inet_release+0x20/0xc3
[   96.270636]  [<ffffffff81674134>] inet6_release+0x30/0x3c
[   96.270636]  [<ffffffff815f2317>] sock_release+0x1f/0x77
[   96.270636]  [<ffffffff815f2396>] sock_close+0x27/0x2b
[   96.270636]  [<ffffffff8110ec22>] __fput+0xf0/0x24b
[   96.270636]  [<ffffffff8110ed8b>] ____fput+0xe/0x10
[   96.270636]  [<ffffffff8104f370>] task_work_run+0x5d/0x75
[   96.270636]  [<ffffffff81038a66>] do_exit+0x26b/0x7d7
[   96.270636]  [<ffffffff81725a95>] ? retint_swapgs+0x13/0x1b
[   96.270636]  [<ffffffff8103925b>] do_group_exit+0x7b/0xba
[   96.270636]  [<ffffffff810392b1>] sys_exit_group+0x17/0x17
[   96.270636]  [<ffffffff8172c78e>] tracesys+0xd0/0xd5
[   96.270636] Code: 00 80 48 01 c7 48 81 ff ff ff ff 1f 76 02 0f 0b 48 89 f8 48 03 05 f6 bd ae 00 eb 32 48 b8 ff ff ff ff ff 87 ff ff 48 39 c7 77 02 <0f> 0b 0f b6 0d 55 57 ba 00 48 b8 00 00 00 00 00 78 00 00 48 01 
[   96.270636] RIP  [<ffffffff8102b22b>] __phys_addr+0x46/0x6b
[   96.270636]  RSP <ffff880019585c98>

Thanks,
Fengguang

-- 
Pengutronix e.K.                  | Marc Kleine-Budde           |
Industrial Linux Solutions        | Phone: +49-231-2826-924     |
Vertretung West/Dortmund          | Fax:   +49-5121-206917-5555 |
Amtsgericht Hildesheim, HRA 2686  | http://www.pengutronix.de   |

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Eric W. Biederman <hidden>
Date: 2012-09-04 20:32:29

Marc Kleine-Budde [off-list ref] writes:
On 09/04/2012 04:04 PM, Fengguang Wu wrote:
quoted
FYI, another kconfig triggering a slightly different oops on tree

        git://gitorious.org/linux-can/linux-can-next led-trigger
This in turn means the problem doesn't come from the CAN patches, as
both trees have different CAN patches. I'm adding Eric W. Biederman on
Cc as he contributed some sctp patches between v3.6 and net-next/master.
Anything is possible, but this seems unlikely as I don't think I touched
anything close to that part of the code.

This most definitely looks like a memory stomp somewhere.

sk->inet_sk->inet_opt has a bad value.

I am puzzled though what are we doing with both ipv4 and ipv6 release
state doing on the same socket path?    Is this some crazy ipv6 socket
doing sctp with only ipv4 addresses?

Eric
Marc
quoted
[   96.267311] ------------[ cut here ]------------
[   96.268294] kernel BUG at /c/kernel-tests/src/stable/arch/x86/mm/physaddr.c:18!
[   96.269988] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC
[   96.270636] Modules linked in:
[   96.270636] CPU 0 
[   96.270636] Pid: 2116, comm: trinity Not tainted 3.6.0-rc3+ #2679 Bochs Bochs
[   96.270636] RIP: 0010:[<ffffffff8102b22b>]  [<ffffffff8102b22b>] __phys_addr+0x46/0x6b
[   96.270636] RSP: 0018:ffff880019585c98  EFLAGS: 00010213
[   96.270636] RAX: ffff87ffffffffff RBX: 0000ea6000000bb8 RCX: 0000000000000000
[   96.270636] RDX: 0000000000000000 RSI: 0000000000000296 RDI: 0000ea6000000bb8
[   96.270636] RBP: ffff880019585c98 R08: 0000000000000058 R09: 0000000000000008
[   96.270636] R10: 000000000000000a R11: 0000000000000058 R12: ffff8800195f7718
[   96.270636] R13: ffffffff816521cf R14: ffffea0000000000 R15: 0000000000000000
[   96.270636] FS:  00007fa19b534700(0000) GS:ffff88001f200000(0000) knlGS:0000000000000000
[   96.270636] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[   96.270636] CR2: 00007fa19b03eba0 CR3: 000000001957b000 CR4: 00000000000006f0
[   96.270636] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[   96.270636] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[   96.270636] Process trinity (pid: 2116, threadinfo ffff880019584000, task ffff88001af2c680)
[   96.270636] Stack:
[   96.270636]  ffff880019585cd8 ffffffff811091d7 0000000000000000 ffff88001b1ef200
[   96.270636]  ffff88001b1ef4d0 0000000000000000 ffff88001b1617b0 0000000000000000
[   96.270636]  ffff880019585cf8 ffffffff816521cf ffff88001b1ef200 ffff88001b1ef248
[   96.270636] Call Trace:
[   96.270636]  [<ffffffff811091d7>] kfree+0x63/0x162
[   96.270636]  [<ffffffff816521cf>] inet_sock_destruct+0x112/0x1ca
[   96.270636]  [<ffffffff815f6fa4>] __sk_free+0x1d/0x114
[   96.270636]  [<ffffffff815f710b>] sk_free+0x1c/0x1e
[   96.270636]  [<ffffffff816d59d5>] sctp_close+0x21a/0x229
[   96.270636]  [<ffffffff810810f6>] ? lock_release_holdtime.part.6+0xb2/0xb7
[   96.270636]  [<ffffffff81651b3e>] ? inet_release+0x65/0xc3
[   96.270636]  [<ffffffff81651b93>] inet_release+0xba/0xc3
[   96.270636]  [<ffffffff81651af9>] ? inet_release+0x20/0xc3
[   96.270636]  [<ffffffff81674134>] inet6_release+0x30/0x3c
[   96.270636]  [<ffffffff815f2317>] sock_release+0x1f/0x77
[   96.270636]  [<ffffffff815f2396>] sock_close+0x27/0x2b
[   96.270636]  [<ffffffff8110ec22>] __fput+0xf0/0x24b
[   96.270636]  [<ffffffff8110ed8b>] ____fput+0xe/0x10
[   96.270636]  [<ffffffff8104f370>] task_work_run+0x5d/0x75
[   96.270636]  [<ffffffff81038a66>] do_exit+0x26b/0x7d7
[   96.270636]  [<ffffffff81725a95>] ? retint_swapgs+0x13/0x1b
[   96.270636]  [<ffffffff8103925b>] do_group_exit+0x7b/0xba
[   96.270636]  [<ffffffff810392b1>] sys_exit_group+0x17/0x17
[   96.270636]  [<ffffffff8172c78e>] tracesys+0xd0/0xd5
[   96.270636] Code: 00 80 48 01 c7 48 81 ff ff ff ff 1f 76 02 0f 0b 48 89 f8 48 03 05 f6 bd ae 00 eb 32 48 b8 ff ff ff ff ff 87 ff ff 48 39 c7 77 02 <0f> 0b 0f b6 0d 55 57 ba 00 48 b8 00 00 00 00 00 78 00 00 48 01 
[   96.270636] RIP  [<ffffffff8102b22b>] __phys_addr+0x46/0x6b
[   96.270636]  RSP <ffff880019585c98>

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Marc Kleine-Budde <mkl@pengutronix.de>
Date: 2012-09-04 20:42:20

On 09/04/2012 10:32 PM, Eric W. Biederman wrote:
quoted
quoted
FYI, another kconfig triggering a slightly different oops on tree

        git://gitorious.org/linux-can/linux-can-next led-trigger
This in turn means the problem doesn't come from the CAN patches, as
both trees have different CAN patches. I'm adding Eric W. Biederman on
Cc as he contributed some sctp patches between v3.6 and net-next/master.
Anything is possible, but this seems unlikely as I don't think I touched
anything close to that part of the code.

This most definitely looks like a memory stomp somewhere.

sk->inet_sk->inet_opt has a bad value.

I am puzzled though what are we doing with both ipv4 and ipv6 release
state doing on the same socket path?    Is this some crazy ipv6 socket
doing sctp with only ipv4 addresses?
It's Wu's testcase, can you show us the code?
Eric, in case you haven't seen, this is another oops, from a slightly
different tree (a handfull of different CAN patches).
[  233.046014] kfree_debugcheck: out of range ptr ea6000000bb8h.
[  233.047399] ------------[ cut here ]------------
[  233.048393] kernel BUG at /c/kernel-tests/src/stable/mm/slab.c:3074!
[  233.048393] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC
[  233.048393] Modules linked in:
[  233.048393] CPU 0 
[  233.048393] Pid: 3929, comm: trinity-watchdo Not tainted 3.6.0-rc3+ #4192 Bochs Bochs
[  233.048393] RIP: 0010:[<ffffffff81169653>]  [<ffffffff81169653>] kfree_debugcheck+0x27/0x2d
[  233.048393] RSP: 0018:ffff88000facbca8  EFLAGS: 00010092
[  233.048393] RAX: 0000000000000031 RBX: 0000ea6000000bb8 RCX: 00000000a189a188
[  233.048393] RDX: 000000000000a189 RSI: ffffffff8108ad32 RDI: ffffffff810d30f9
[  233.048393] RBP: ffff88000facbcb8 R08: 0000000000000002 R09: ffffffff843846f0
[  233.048393] R10: ffffffff810ae37c R11: 0000000000000908 R12: 0000000000000202
[  233.048393] R13: ffffffff823dbd5a R14: ffff88000ec5bea8 R15: ffffffff8363c780
[  233.048393] FS:  00007faa6899c700(0000) GS:ffff88001f200000(0000) knlGS:0000000000000000
[  233.048393] CS:  0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[  233.048393] CR2: 00007faa6841019c CR3: 0000000012c82000 CR4: 00000000000006f0
[  233.048393] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[  233.048393] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[  233.048393] Process trinity-watchdo (pid: 3929, threadinfo ffff88000faca000, task ffff88000faec600)
[  233.048393] Stack:
[  233.048393]  0000000000000000 0000ea6000000bb8 ffff88000facbce8 ffffffff8116ad81
[  233.048393]  ffff88000ff588a0 ffff88000ff58850 ffff88000ff588a0 0000000000000000
[  233.048393]  ffff88000facbd08 ffffffff823dbd5a ffffffff823dbcb0 ffff88000ff58850
[  233.048393] Call Trace:
[  233.048393]  [<ffffffff8116ad81>] kfree+0x5f/0xca
[  233.048393]  [<ffffffff823dbd5a>] inet_sock_destruct+0xaa/0x13c
[  233.048393]  [<ffffffff823dbcb0>] ? inet_sk_rebuild_header+0x319/0x319
[  233.048393]  [<ffffffff8231c307>] __sk_free+0x21/0x14b
[  233.048393]  [<ffffffff8231c4bd>] sk_free+0x26/0x2a
[  233.048393]  [<ffffffff825372db>] sctp_close+0x215/0x224
[  233.048393]  [<ffffffff810d6835>] ? lock_release+0x16f/0x1b9
[  233.048393]  [<ffffffff823daf12>] inet_release+0x7e/0x85
[  233.048393]  [<ffffffff82317d15>] sock_release+0x1f/0x77
[  233.048393]  [<ffffffff82317d94>] sock_close+0x27/0x2b
[  233.048393]  [<ffffffff81173bbe>] __fput+0x101/0x20a
[  233.048393]  [<ffffffff81173cd5>] ____fput+0xe/0x10
[  233.048393]  [<ffffffff810a3794>] task_work_run+0x5d/0x75
[  233.048393]  [<ffffffff8108da70>] do_exit+0x290/0x7f5
[  233.048393]  [<ffffffff82707415>] ? retint_swapgs+0x13/0x1b
[  233.048393]  [<ffffffff8108e23f>] do_group_exit+0x7b/0xba
[  233.048393]  [<ffffffff8108e295>] sys_exit_group+0x17/0x17
[  233.048393]  [<ffffffff8270de10>] tracesys+0xdd/0xe2
[  233.048393] Code: 59 01 5d c3 55 48 89 e5 53 41 50 0f 1f 44 00 00 48 89 fb e8 d4 b0 f0 ff 84 c0 75 11 48 89 de 48 c7 c7 fc fa f7 82 e8 0d 0f 57 01 <0f> 0b 5f 5b 5d c3 55 48 89 e5 0f 1f 44 00 00 48 63 87 d8 00 00 
[  233.048393] RIP  [<ffffffff81169653>] kfree_debugcheck+0x27/0x2d
[  233.048393]  RSP <ffff88000facbca8>
Wu is running a bisect, let's hope that gives us a result.

Marc

-- 
Pengutronix e.K.                  | Marc Kleine-Budde           |
Industrial Linux Solutions        | Phone: +49-231-2826-924     |
Vertretung West/Dortmund          | Fax:   +49-5121-206917-5555 |
Amtsgericht Hildesheim, HRA 2686  | http://www.pengutronix.de   |

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Fengguang Wu <hidden>
Date: 2012-09-05 14:55:14

On Tue, Sep 04, 2012 at 01:32:21PM -0700, Eric W. Biederman wrote:
Marc Kleine-Budde [off-list ref] writes:
quoted
On 09/04/2012 04:04 PM, Fengguang Wu wrote:
quoted
FYI, another kconfig triggering a slightly different oops on tree

        git://gitorious.org/linux-can/linux-can-next led-trigger
This in turn means the problem doesn't come from the CAN patches, as
both trees have different CAN patches. I'm adding Eric W. Biederman on
Cc as he contributed some sctp patches between v3.6 and net-next/master.
Anything is possible, but this seems unlikely as I don't think I touched
anything close to that part of the code.
You are both right.  The bad commit turns out to be one of:

1bed966cc3bd4042110129f0fc51aeeb59c5b200 Merge branch 'tcp_fastopen_server'
168a8f58059a22feb9e9a2dcc1b8053dbbbc12ef tcp: TCP Fast Open Server - main code path
8336886f786fdacbc19b719c1f7ea91eb70706d4 tcp: TCP Fast Open Server - support TFO listeners

Thanks,
Fengguang
This most definitely looks like a memory stomp somewhere.

sk->inet_sk->inet_opt has a bad value.

I am puzzled though what are we doing with both ipv4 and ipv6 release
state doing on the same socket path?    Is this some crazy ipv6 socket
doing sctp with only ipv4 addresses?
quoted
quoted
[   96.267311] ------------[ cut here ]------------
[   96.268294] kernel BUG at /c/kernel-tests/src/stable/arch/x86/mm/physaddr.c:18!
[   96.269988] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC
[   96.270636] Modules linked in:
[   96.270636] CPU 0 
[   96.270636] Pid: 2116, comm: trinity Not tainted 3.6.0-rc3+ #2679 Bochs Bochs
[   96.270636] RIP: 0010:[<ffffffff8102b22b>]  [<ffffffff8102b22b>] __phys_addr+0x46/0x6b
[   96.270636] RSP: 0018:ffff880019585c98  EFLAGS: 00010213
[   96.270636] RAX: ffff87ffffffffff RBX: 0000ea6000000bb8 RCX: 0000000000000000
[   96.270636] RDX: 0000000000000000 RSI: 0000000000000296 RDI: 0000ea6000000bb8
[   96.270636] RBP: ffff880019585c98 R08: 0000000000000058 R09: 0000000000000008
[   96.270636] R10: 000000000000000a R11: 0000000000000058 R12: ffff8800195f7718
[   96.270636] R13: ffffffff816521cf R14: ffffea0000000000 R15: 0000000000000000
[   96.270636] FS:  00007fa19b534700(0000) GS:ffff88001f200000(0000) knlGS:0000000000000000
[   96.270636] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[   96.270636] CR2: 00007fa19b03eba0 CR3: 000000001957b000 CR4: 00000000000006f0
[   96.270636] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[   96.270636] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[   96.270636] Process trinity (pid: 2116, threadinfo ffff880019584000, task ffff88001af2c680)
[   96.270636] Stack:
[   96.270636]  ffff880019585cd8 ffffffff811091d7 0000000000000000 ffff88001b1ef200
[   96.270636]  ffff88001b1ef4d0 0000000000000000 ffff88001b1617b0 0000000000000000
[   96.270636]  ffff880019585cf8 ffffffff816521cf ffff88001b1ef200 ffff88001b1ef248
[   96.270636] Call Trace:
[   96.270636]  [<ffffffff811091d7>] kfree+0x63/0x162
[   96.270636]  [<ffffffff816521cf>] inet_sock_destruct+0x112/0x1ca
[   96.270636]  [<ffffffff815f6fa4>] __sk_free+0x1d/0x114
[   96.270636]  [<ffffffff815f710b>] sk_free+0x1c/0x1e
[   96.270636]  [<ffffffff816d59d5>] sctp_close+0x21a/0x229
[   96.270636]  [<ffffffff810810f6>] ? lock_release_holdtime.part.6+0xb2/0xb7
[   96.270636]  [<ffffffff81651b3e>] ? inet_release+0x65/0xc3
[   96.270636]  [<ffffffff81651b93>] inet_release+0xba/0xc3
[   96.270636]  [<ffffffff81651af9>] ? inet_release+0x20/0xc3
[   96.270636]  [<ffffffff81674134>] inet6_release+0x30/0x3c
[   96.270636]  [<ffffffff815f2317>] sock_release+0x1f/0x77
[   96.270636]  [<ffffffff815f2396>] sock_close+0x27/0x2b
[   96.270636]  [<ffffffff8110ec22>] __fput+0xf0/0x24b
[   96.270636]  [<ffffffff8110ed8b>] ____fput+0xe/0x10
[   96.270636]  [<ffffffff8104f370>] task_work_run+0x5d/0x75
[   96.270636]  [<ffffffff81038a66>] do_exit+0x26b/0x7d7
[   96.270636]  [<ffffffff81725a95>] ? retint_swapgs+0x13/0x1b
[   96.270636]  [<ffffffff8103925b>] do_group_exit+0x7b/0xba
[   96.270636]  [<ffffffff810392b1>] sys_exit_group+0x17/0x17
[   96.270636]  [<ffffffff8172c78e>] tracesys+0xd0/0xd5
[   96.270636] Code: 00 80 48 01 c7 48 81 ff ff ff ff 1f 76 02 0f 0b 48 89 f8 48 03 05 f6 bd ae 00 eb 32 48 b8 ff ff ff ff ff 87 ff ff 48 39 c7 77 02 <0f> 0b 0f b6 0d 55 57 ba 00 48 b8 00 00 00 00 00 78 00 00 48 01 
[   96.270636] RIP  [<ffffffff8102b22b>] __phys_addr+0x46/0x6b
[   96.270636]  RSP <ffff880019585c98>

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Marc Kleine-Budde <mkl@pengutronix.de>
Date: 2012-09-05 15:01:15

On 09/05/2012 04:55 PM, Fengguang Wu wrote:
quoted
quoted
This in turn means the problem doesn't come from the CAN patches, as
both trees have different CAN patches. I'm adding Eric W. Biederman on
Cc as he contributed some sctp patches between v3.6 and net-next/master.
Anything is possible, but this seems unlikely as I don't think I touched
anything close to that part of the code.
You are both right.  The bad commit turns out to be one of:

1bed966cc3bd4042110129f0fc51aeeb59c5b200 Merge branch 'tcp_fastopen_server'
168a8f58059a22feb9e9a2dcc1b8053dbbbc12ef tcp: TCP Fast Open Server - main code path
8336886f786fdacbc19b719c1f7ea91eb70706d4 tcp: TCP Fast Open Server - support TFO listeners

Thanks,
Fengguang
Thanks for your work Fengguang.

Marc

-- 
Pengutronix e.K.                  | Marc Kleine-Budde           |
Industrial Linux Solutions        | Phone: +49-231-2826-924     |
Vertretung West/Dortmund          | Fax:   +49-5121-206917-5555 |
Amtsgericht Hildesheim, HRA 2686  | http://www.pengutronix.de   |

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Eric Dumazet <hidden>
Date: 2012-09-05 15:30:52

On Wed, 2012-09-05 at 17:01 +0200, Marc Kleine-Budde wrote:
On 09/05/2012 04:55 PM, Fengguang Wu wrote:
quoted
quoted
quoted
This in turn means the problem doesn't come from the CAN patches, as
both trees have different CAN patches. I'm adding Eric W. Biederman on
Cc as he contributed some sctp patches between v3.6 and net-next/master.
Anything is possible, but this seems unlikely as I don't think I touched
anything close to that part of the code.
You are both right.  The bad commit turns out to be one of:

1bed966cc3bd4042110129f0fc51aeeb59c5b200 Merge branch 'tcp_fastopen_server'
168a8f58059a22feb9e9a2dcc1b8053dbbbc12ef tcp: TCP Fast Open Server - main code path
8336886f786fdacbc19b719c1f7ea91eb70706d4 tcp: TCP Fast Open Server - support TFO listeners

Thanks,
Fengguang
Thanks for your work Fengguang.

Marc
OK I have a good idea how to fix the bug, I will send a patch ASAP

Thanks

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Eric Dumazet <hidden>
Date: 2012-09-05 15:40:52

On Wed, 2012-09-05 at 17:30 +0200, Eric Dumazet wrote:
On Wed, 2012-09-05 at 17:01 +0200, Marc Kleine-Budde wrote:
quoted
On 09/05/2012 04:55 PM, Fengguang Wu wrote:
quoted
quoted
quoted
This in turn means the problem doesn't come from the CAN patches, as
both trees have different CAN patches. I'm adding Eric W. Biederman on
Cc as he contributed some sctp patches between v3.6 and net-next/master.
Anything is possible, but this seems unlikely as I don't think I touched
anything close to that part of the code.
You are both right.  The bad commit turns out to be one of:

1bed966cc3bd4042110129f0fc51aeeb59c5b200 Merge branch 'tcp_fastopen_server'
168a8f58059a22feb9e9a2dcc1b8053dbbbc12ef tcp: TCP Fast Open Server - main code path
8336886f786fdacbc19b719c1f7ea91eb70706d4 tcp: TCP Fast Open Server - support TFO listeners

Thanks,
Fengguang
Thanks for your work Fengguang.

Marc
OK I have a good idea how to fix the bug, I will send a patch ASAP
Could you test the following patch please ?

(Not sure why sctp doesnt memset/bzero its whole socket by the way...)

Thanks
diff --git a/net/ipv4/af_inet.c b/net/ipv4/af_inet.c
index 4f70ef0..845372b 100644
--- a/net/ipv4/af_inet.c
+++ b/net/ipv4/af_inet.c
@@ -149,11 +149,8 @@ void inet_sock_destruct(struct sock *sk)
 		pr_err("Attempt to release alive inet socket %p\n", sk);
 		return;
 	}
-	if (sk->sk_type == SOCK_STREAM) {
-		struct fastopen_queue *fastopenq =
-			inet_csk(sk)->icsk_accept_queue.fastopenq;
-		kfree(fastopenq);
-	}
+	if (sk->sk_protocol == IPPROTO_TCP)
+		kfree(inet_csk(sk)->icsk_accept_queue.fastopenq);
 
 	WARN_ON(atomic_read(&sk->sk_rmem_alloc));
 	WARN_ON(atomic_read(&sk->sk_wmem_alloc));


Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Eric Dumazet <hidden>
Date: 2012-09-05 16:57:10

On Wed, 2012-09-05 at 17:40 +0200, Eric Dumazet wrote:
Could you test the following patch please ?

(Not sure why sctp doesnt memset/bzero its whole socket by the way...)

Thanks
Here is a more complete patch, as there are three potential problems,
not only one :
diff --git a/net/ipv4/af_inet.c b/net/ipv4/af_inet.c
index 4f70ef0..845372b 100644
--- a/net/ipv4/af_inet.c
+++ b/net/ipv4/af_inet.c
@@ -149,11 +149,8 @@ void inet_sock_destruct(struct sock *sk)
 		pr_err("Attempt to release alive inet socket %p\n", sk);
 		return;
 	}
-	if (sk->sk_type == SOCK_STREAM) {
-		struct fastopen_queue *fastopenq =
-			inet_csk(sk)->icsk_accept_queue.fastopenq;
-		kfree(fastopenq);
-	}
+	if (sk->sk_protocol == IPPROTO_TCP)
+		kfree(inet_csk(sk)->icsk_accept_queue.fastopenq);
 
 	WARN_ON(atomic_read(&sk->sk_rmem_alloc));
 	WARN_ON(atomic_read(&sk->sk_wmem_alloc));
diff --git a/net/ipv4/inet_connection_sock.c b/net/ipv4/inet_connection_sock.c
index 8464b79..f0c5b9c 100644
--- a/net/ipv4/inet_connection_sock.c
+++ b/net/ipv4/inet_connection_sock.c
@@ -314,7 +314,7 @@ struct sock *inet_csk_accept(struct sock *sk, int flags, int *err)
 	newsk = req->sk;
 
 	sk_acceptq_removed(sk);
-	if (sk->sk_type == SOCK_STREAM && queue->fastopenq != NULL) {
+	if (sk->sk_protocol == IPPROTO_TCP && queue->fastopenq != NULL) {
 		spin_lock_bh(&queue->fastopenq->lock);
 		if (tcp_rsk(req)->listener) {
 			/* We are still waiting for the final ACK from 3WHS
@@ -775,7 +775,7 @@ void inet_csk_listen_stop(struct sock *sk)
 
 		percpu_counter_inc(sk->sk_prot->orphan_count);
 
-		if (sk->sk_type == SOCK_STREAM && tcp_rsk(req)->listener) {
+		if (sk->sk_protocol == IPPROTO_TCP && tcp_rsk(req)->listener) {
 			BUG_ON(tcp_sk(child)->fastopen_rsk != req);
 			BUG_ON(sk != tcp_rsk(req)->listener);
 

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Fengguang Wu <hidden>
Date: 2012-09-05 22:28:54

On Wed, Sep 05, 2012 at 06:57:00PM +0200, Eric Dumazet wrote:
On Wed, 2012-09-05 at 17:40 +0200, Eric Dumazet wrote:
quoted
Could you test the following patch please ?
It works - no single error for 1000 boots!

btw, the first bad commit has been bisected to 

        commit 8336886f786fdacbc19b719c1f7ea91eb70706d4
        Author: Jerry Chu [off-list ref]
        Date:   Fri Aug 31 12:29:12 2012 +0000

            tcp: TCP Fast Open Server - support TFO listeners
quoted
(Not sure why sctp doesnt memset/bzero its whole socket by the way...)

Thanks
Here is a more complete patch, as there are three potential problems,
not only one :
Great! I'll start tests for it.

Thanks,
Fengguang
quoted hunk
diff --git a/net/ipv4/af_inet.c b/net/ipv4/af_inet.c
index 4f70ef0..845372b 100644
--- a/net/ipv4/af_inet.c
+++ b/net/ipv4/af_inet.c
@@ -149,11 +149,8 @@ void inet_sock_destruct(struct sock *sk)
 		pr_err("Attempt to release alive inet socket %p\n", sk);
 		return;
 	}
-	if (sk->sk_type == SOCK_STREAM) {
-		struct fastopen_queue *fastopenq =
-			inet_csk(sk)->icsk_accept_queue.fastopenq;
-		kfree(fastopenq);
-	}
+	if (sk->sk_protocol == IPPROTO_TCP)
+		kfree(inet_csk(sk)->icsk_accept_queue.fastopenq);
 
 	WARN_ON(atomic_read(&sk->sk_rmem_alloc));
 	WARN_ON(atomic_read(&sk->sk_wmem_alloc));
diff --git a/net/ipv4/inet_connection_sock.c b/net/ipv4/inet_connection_sock.c
index 8464b79..f0c5b9c 100644
--- a/net/ipv4/inet_connection_sock.c
+++ b/net/ipv4/inet_connection_sock.c
@@ -314,7 +314,7 @@ struct sock *inet_csk_accept(struct sock *sk, int flags, int *err)
 	newsk = req->sk;
 
 	sk_acceptq_removed(sk);
-	if (sk->sk_type == SOCK_STREAM && queue->fastopenq != NULL) {
+	if (sk->sk_protocol == IPPROTO_TCP && queue->fastopenq != NULL) {
 		spin_lock_bh(&queue->fastopenq->lock);
 		if (tcp_rsk(req)->listener) {
 			/* We are still waiting for the final ACK from 3WHS
@@ -775,7 +775,7 @@ void inet_csk_listen_stop(struct sock *sk)
 
 		percpu_counter_inc(sk->sk_prot->orphan_count);
 
-		if (sk->sk_type == SOCK_STREAM && tcp_rsk(req)->listener) {
+		if (sk->sk_protocol == IPPROTO_TCP && tcp_rsk(req)->listener) {
 			BUG_ON(tcp_sk(child)->fastopen_rsk != req);
 			BUG_ON(sk != tcp_rsk(req)->listener);
 

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Jerry Chu <hidden>
Date: 2012-09-05 23:07:09

On Wed, Sep 5, 2012 at 3:28 PM, Fengguang Wu [off-list ref] wrote:
On Wed, Sep 05, 2012 at 06:57:00PM +0200, Eric Dumazet wrote:
quoted
On Wed, 2012-09-05 at 17:40 +0200, Eric Dumazet wrote:
quoted
Could you test the following patch please ?
It works - no single error for 1000 boots!
Sorry for introducing the bug, one of the casualties dealing with code
that is shared
outside of TCP. I did spend some effort adding special checks but I was wrong in
assuming inet_create() will zero all the field including fastopenq
inside icsk_accept_queue
inside struct inet_connection_sock - although this is true for TCP and
DCCP, SCTP doesn't
have inet_connection_sock hence inet_csk(sk) is bogus.

Kudo to Eric for fixing it quickly before I got to it.

Jerry
btw, the first bad commit has been bisected to

        commit 8336886f786fdacbc19b719c1f7ea91eb70706d4
        Author: Jerry Chu [off-list ref]
        Date:   Fri Aug 31 12:29:12 2012 +0000

            tcp: TCP Fast Open Server - support TFO listeners
quoted
quoted
(Not sure why sctp doesnt memset/bzero its whole socket by the way...)

Thanks
Here is a more complete patch, as there are three potential problems,
not only one :
Great! I'll start tests for it.

Thanks,
Fengguang
quoted
diff --git a/net/ipv4/af_inet.c b/net/ipv4/af_inet.c
index 4f70ef0..845372b 100644
--- a/net/ipv4/af_inet.c
+++ b/net/ipv4/af_inet.c
@@ -149,11 +149,8 @@ void inet_sock_destruct(struct sock *sk)
              pr_err("Attempt to release alive inet socket %p\n", sk);
              return;
      }
-     if (sk->sk_type == SOCK_STREAM) {
-             struct fastopen_queue *fastopenq =
-                     inet_csk(sk)->icsk_accept_queue.fastopenq;
-             kfree(fastopenq);
-     }
+     if (sk->sk_protocol == IPPROTO_TCP)
+             kfree(inet_csk(sk)->icsk_accept_queue.fastopenq);

      WARN_ON(atomic_read(&sk->sk_rmem_alloc));
      WARN_ON(atomic_read(&sk->sk_wmem_alloc));
diff --git a/net/ipv4/inet_connection_sock.c b/net/ipv4/inet_connection_sock.c
index 8464b79..f0c5b9c 100644
--- a/net/ipv4/inet_connection_sock.c
+++ b/net/ipv4/inet_connection_sock.c
@@ -314,7 +314,7 @@ struct sock *inet_csk_accept(struct sock *sk, int flags, int *err)
      newsk = req->sk;

      sk_acceptq_removed(sk);
-     if (sk->sk_type == SOCK_STREAM && queue->fastopenq != NULL) {
+     if (sk->sk_protocol == IPPROTO_TCP && queue->fastopenq != NULL) {
              spin_lock_bh(&queue->fastopenq->lock);
              if (tcp_rsk(req)->listener) {
                      /* We are still waiting for the final ACK from 3WHS
@@ -775,7 +775,7 @@ void inet_csk_listen_stop(struct sock *sk)

              percpu_counter_inc(sk->sk_prot->orphan_count);

-             if (sk->sk_type == SOCK_STREAM && tcp_rsk(req)->listener) {
+             if (sk->sk_protocol == IPPROTO_TCP && tcp_rsk(req)->listener) {
                      BUG_ON(tcp_sk(child)->fastopen_rsk != req);
                      BUG_ON(sk != tcp_rsk(req)->listener);
--
To unsubscribe from this list: send the line "unsubscribe netdev" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

Re: sctp_close/sk_free: kernel BUG at arch/x86/mm/physaddr.c:18!

From: Fengguang Wu <hidden>
Date: 2012-09-06 04:54:55

On Wed, Sep 05, 2012 at 06:57:00PM +0200, Eric Dumazet wrote:
Here is a more complete patch, as there are three potential problems,
not only one :
It's fine, too.

Tested-by: Fengguang Wu <redacted>

Thanks!

[PATCH net-next] tcp: fix TFO regression

From: Eric Dumazet <hidden>
Date: 2012-09-06 18:07:25

From: Eric Dumazet <edumazet@google.com>

Fengguang Wu reported various panics and bisected to commit
8336886f786fdac (tcp: TCP Fast Open Server - support TFO listeners)

Fix this by making sure socket is a TCP socket before accessing TFO data
structures.

[  233.046014] kfree_debugcheck: out of range ptr ea6000000bb8h.
[  233.047399] ------------[ cut here ]------------
[  233.048393] kernel BUG at /c/kernel-tests/src/stable/mm/slab.c:3074!
[  233.048393] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC
[  233.048393] Modules linked in:
[  233.048393] CPU 0 
[  233.048393] Pid: 3929, comm: trinity-watchdo Not tainted 3.6.0-rc3+
#4192 Bochs Bochs
[  233.048393] RIP: 0010:[<ffffffff81169653>]  [<ffffffff81169653>]
kfree_debugcheck+0x27/0x2d
[  233.048393] RSP: 0018:ffff88000facbca8  EFLAGS: 00010092
[  233.048393] RAX: 0000000000000031 RBX: 0000ea6000000bb8 RCX:
00000000a189a188
[  233.048393] RDX: 000000000000a189 RSI: ffffffff8108ad32 RDI:
ffffffff810d30f9
[  233.048393] RBP: ffff88000facbcb8 R08: 0000000000000002 R09:
ffffffff843846f0
[  233.048393] R10: ffffffff810ae37c R11: 0000000000000908 R12:
0000000000000202
[  233.048393] R13: ffffffff823dbd5a R14: ffff88000ec5bea8 R15:
ffffffff8363c780
[  233.048393] FS:  00007faa6899c700(0000) GS:ffff88001f200000(0000)
knlGS:0000000000000000
[  233.048393] CS:  0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[  233.048393] CR2: 00007faa6841019c CR3: 0000000012c82000 CR4:
00000000000006f0
[  233.048393] DR0: 0000000000000000 DR1: 0000000000000000 DR2:
0000000000000000
[  233.048393] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7:
0000000000000400
[  233.048393] Process trinity-watchdo (pid: 3929, threadinfo
ffff88000faca000, task ffff88000faec600)
[  233.048393] Stack:
[  233.048393]  0000000000000000 0000ea6000000bb8 ffff88000facbce8
ffffffff8116ad81
[  233.048393]  ffff88000ff588a0 ffff88000ff58850 ffff88000ff588a0
0000000000000000
[  233.048393]  ffff88000facbd08 ffffffff823dbd5a ffffffff823dbcb0
ffff88000ff58850
[  233.048393] Call Trace:
[  233.048393]  [<ffffffff8116ad81>] kfree+0x5f/0xca
[  233.048393]  [<ffffffff823dbd5a>] inet_sock_destruct+0xaa/0x13c
[  233.048393]  [<ffffffff823dbcb0>] ? inet_sk_rebuild_header
+0x319/0x319
[  233.048393]  [<ffffffff8231c307>] __sk_free+0x21/0x14b
[  233.048393]  [<ffffffff8231c4bd>] sk_free+0x26/0x2a
[  233.048393]  [<ffffffff825372db>] sctp_close+0x215/0x224
[  233.048393]  [<ffffffff810d6835>] ? lock_release+0x16f/0x1b9
[  233.048393]  [<ffffffff823daf12>] inet_release+0x7e/0x85
[  233.048393]  [<ffffffff82317d15>] sock_release+0x1f/0x77
[  233.048393]  [<ffffffff82317d94>] sock_close+0x27/0x2b
[  233.048393]  [<ffffffff81173bbe>] __fput+0x101/0x20a
[  233.048393]  [<ffffffff81173cd5>] ____fput+0xe/0x10
[  233.048393]  [<ffffffff810a3794>] task_work_run+0x5d/0x75
[  233.048393]  [<ffffffff8108da70>] do_exit+0x290/0x7f5
[  233.048393]  [<ffffffff82707415>] ? retint_swapgs+0x13/0x1b
[  233.048393]  [<ffffffff8108e23f>] do_group_exit+0x7b/0xba
[  233.048393]  [<ffffffff8108e295>] sys_exit_group+0x17/0x17
[  233.048393]  [<ffffffff8270de10>] tracesys+0xdd/0xe2
[  233.048393] Code: 59 01 5d c3 55 48 89 e5 53 41 50 0f 1f 44 00 00 48
89 fb e8 d4 b0 f0 ff 84 c0 75 11 48 89 de 48 c7 c7 fc fa f7 82 e8 0d 0f
57 01 <0f> 0b 5f 5b 5d c3 55 48 89 e5 0f 1f 44 00 00 48 63 87 d8 00 00 
[  233.048393] RIP  [<ffffffff81169653>] kfree_debugcheck+0x27/0x2d
[  233.048393]  RSP <ffff88000facbca8>

Reported-by: Fengguang Wu <redacted>
Tested-by: Fengguang Wu <redacted>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: "H.K. Jerry Chu" <redacted>
---
 net/ipv4/af_inet.c              |    7 ++-----
 net/ipv4/inet_connection_sock.c |    4 ++--
 2 files changed, 4 insertions(+), 7 deletions(-)
diff --git a/net/ipv4/af_inet.c b/net/ipv4/af_inet.c
index 4f70ef0..845372b 100644
--- a/net/ipv4/af_inet.c
+++ b/net/ipv4/af_inet.c
@@ -149,11 +149,8 @@ void inet_sock_destruct(struct sock *sk)
 		pr_err("Attempt to release alive inet socket %p\n", sk);
 		return;
 	}
-	if (sk->sk_type == SOCK_STREAM) {
-		struct fastopen_queue *fastopenq =
-			inet_csk(sk)->icsk_accept_queue.fastopenq;
-		kfree(fastopenq);
-	}
+	if (sk->sk_protocol == IPPROTO_TCP)
+		kfree(inet_csk(sk)->icsk_accept_queue.fastopenq);
 
 	WARN_ON(atomic_read(&sk->sk_rmem_alloc));
 	WARN_ON(atomic_read(&sk->sk_wmem_alloc));
diff --git a/net/ipv4/inet_connection_sock.c b/net/ipv4/inet_connection_sock.c
index 8464b79..f0c5b9c 100644
--- a/net/ipv4/inet_connection_sock.c
+++ b/net/ipv4/inet_connection_sock.c
@@ -314,7 +314,7 @@ struct sock *inet_csk_accept(struct sock *sk, int flags, int *err)
 	newsk = req->sk;
 
 	sk_acceptq_removed(sk);
-	if (sk->sk_type == SOCK_STREAM && queue->fastopenq != NULL) {
+	if (sk->sk_protocol == IPPROTO_TCP && queue->fastopenq != NULL) {
 		spin_lock_bh(&queue->fastopenq->lock);
 		if (tcp_rsk(req)->listener) {
 			/* We are still waiting for the final ACK from 3WHS
@@ -775,7 +775,7 @@ void inet_csk_listen_stop(struct sock *sk)
 
 		percpu_counter_inc(sk->sk_prot->orphan_count);
 
-		if (sk->sk_type == SOCK_STREAM && tcp_rsk(req)->listener) {
+		if (sk->sk_protocol == IPPROTO_TCP && tcp_rsk(req)->listener) {
 			BUG_ON(tcp_sk(child)->fastopen_rsk != req);
 			BUG_ON(sk != tcp_rsk(req)->listener);
 

Re: [PATCH net-next] tcp: fix TFO regression

From: Neal Cardwell <ncardwell@google.com>
Date: 2012-09-06 18:15:38

On Thu, Sep 6, 2012 at 2:07 PM, Eric Dumazet [off-list ref] wrote:
From: Eric Dumazet <edumazet@google.com>

Fengguang Wu reported various panics and bisected to commit
8336886f786fdac (tcp: TCP Fast Open Server - support TFO listeners)

Fix this by making sure socket is a TCP socket before accessing TFO data
structures.
...
Reported-by: Fengguang Wu <redacted>
Tested-by: Fengguang Wu <redacted>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: "H.K. Jerry Chu" <redacted>
Acked-by: Neal Cardwell <ncardwell@google.com>

neal

Re: [PATCH net-next] tcp: fix TFO regression

From: Jerry Chu <hidden>
Date: 2012-09-06 18:18:27

On Thu, Sep 6, 2012 at 11:07 AM, Eric Dumazet [off-list ref] wrote:
quoted hunk
From: Eric Dumazet <edumazet@google.com>

Fengguang Wu reported various panics and bisected to commit
8336886f786fdac (tcp: TCP Fast Open Server - support TFO listeners)

Fix this by making sure socket is a TCP socket before accessing TFO data
structures.

[  233.046014] kfree_debugcheck: out of range ptr ea6000000bb8h.
[  233.047399] ------------[ cut here ]------------
[  233.048393] kernel BUG at /c/kernel-tests/src/stable/mm/slab.c:3074!
[  233.048393] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC
[  233.048393] Modules linked in:
[  233.048393] CPU 0
[  233.048393] Pid: 3929, comm: trinity-watchdo Not tainted 3.6.0-rc3+
#4192 Bochs Bochs
[  233.048393] RIP: 0010:[<ffffffff81169653>]  [<ffffffff81169653>]
kfree_debugcheck+0x27/0x2d
[  233.048393] RSP: 0018:ffff88000facbca8  EFLAGS: 00010092
[  233.048393] RAX: 0000000000000031 RBX: 0000ea6000000bb8 RCX:
00000000a189a188
[  233.048393] RDX: 000000000000a189 RSI: ffffffff8108ad32 RDI:
ffffffff810d30f9
[  233.048393] RBP: ffff88000facbcb8 R08: 0000000000000002 R09:
ffffffff843846f0
[  233.048393] R10: ffffffff810ae37c R11: 0000000000000908 R12:
0000000000000202
[  233.048393] R13: ffffffff823dbd5a R14: ffff88000ec5bea8 R15:
ffffffff8363c780
[  233.048393] FS:  00007faa6899c700(0000) GS:ffff88001f200000(0000)
knlGS:0000000000000000
[  233.048393] CS:  0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[  233.048393] CR2: 00007faa6841019c CR3: 0000000012c82000 CR4:
00000000000006f0
[  233.048393] DR0: 0000000000000000 DR1: 0000000000000000 DR2:
0000000000000000
[  233.048393] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7:
0000000000000400
[  233.048393] Process trinity-watchdo (pid: 3929, threadinfo
ffff88000faca000, task ffff88000faec600)
[  233.048393] Stack:
[  233.048393]  0000000000000000 0000ea6000000bb8 ffff88000facbce8
ffffffff8116ad81
[  233.048393]  ffff88000ff588a0 ffff88000ff58850 ffff88000ff588a0
0000000000000000
[  233.048393]  ffff88000facbd08 ffffffff823dbd5a ffffffff823dbcb0
ffff88000ff58850
[  233.048393] Call Trace:
[  233.048393]  [<ffffffff8116ad81>] kfree+0x5f/0xca
[  233.048393]  [<ffffffff823dbd5a>] inet_sock_destruct+0xaa/0x13c
[  233.048393]  [<ffffffff823dbcb0>] ? inet_sk_rebuild_header
+0x319/0x319
[  233.048393]  [<ffffffff8231c307>] __sk_free+0x21/0x14b
[  233.048393]  [<ffffffff8231c4bd>] sk_free+0x26/0x2a
[  233.048393]  [<ffffffff825372db>] sctp_close+0x215/0x224
[  233.048393]  [<ffffffff810d6835>] ? lock_release+0x16f/0x1b9
[  233.048393]  [<ffffffff823daf12>] inet_release+0x7e/0x85
[  233.048393]  [<ffffffff82317d15>] sock_release+0x1f/0x77
[  233.048393]  [<ffffffff82317d94>] sock_close+0x27/0x2b
[  233.048393]  [<ffffffff81173bbe>] __fput+0x101/0x20a
[  233.048393]  [<ffffffff81173cd5>] ____fput+0xe/0x10
[  233.048393]  [<ffffffff810a3794>] task_work_run+0x5d/0x75
[  233.048393]  [<ffffffff8108da70>] do_exit+0x290/0x7f5
[  233.048393]  [<ffffffff82707415>] ? retint_swapgs+0x13/0x1b
[  233.048393]  [<ffffffff8108e23f>] do_group_exit+0x7b/0xba
[  233.048393]  [<ffffffff8108e295>] sys_exit_group+0x17/0x17
[  233.048393]  [<ffffffff8270de10>] tracesys+0xdd/0xe2
[  233.048393] Code: 59 01 5d c3 55 48 89 e5 53 41 50 0f 1f 44 00 00 48
89 fb e8 d4 b0 f0 ff 84 c0 75 11 48 89 de 48 c7 c7 fc fa f7 82 e8 0d 0f
57 01 <0f> 0b 5f 5b 5d c3 55 48 89 e5 0f 1f 44 00 00 48 63 87 d8 00 00
[  233.048393] RIP  [<ffffffff81169653>] kfree_debugcheck+0x27/0x2d
[  233.048393]  RSP <ffff88000facbca8>

Reported-by: Fengguang Wu <redacted>
Tested-by: Fengguang Wu <redacted>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: "H.K. Jerry Chu" <redacted>
---
 net/ipv4/af_inet.c              |    7 ++-----
 net/ipv4/inet_connection_sock.c |    4 ++--
 2 files changed, 4 insertions(+), 7 deletions(-)
diff --git a/net/ipv4/af_inet.c b/net/ipv4/af_inet.c
index 4f70ef0..845372b 100644
--- a/net/ipv4/af_inet.c
+++ b/net/ipv4/af_inet.c
@@ -149,11 +149,8 @@ void inet_sock_destruct(struct sock *sk)
                pr_err("Attempt to release alive inet socket %p\n", sk);
                return;
        }
-       if (sk->sk_type == SOCK_STREAM) {
-               struct fastopen_queue *fastopenq =
-                       inet_csk(sk)->icsk_accept_queue.fastopenq;
-               kfree(fastopenq);
-       }
+       if (sk->sk_protocol == IPPROTO_TCP)
+               kfree(inet_csk(sk)->icsk_accept_queue.fastopenq);

        WARN_ON(atomic_read(&sk->sk_rmem_alloc));
        WARN_ON(atomic_read(&sk->sk_wmem_alloc));
diff --git a/net/ipv4/inet_connection_sock.c b/net/ipv4/inet_connection_sock.c
index 8464b79..f0c5b9c 100644
--- a/net/ipv4/inet_connection_sock.c
+++ b/net/ipv4/inet_connection_sock.c
@@ -314,7 +314,7 @@ struct sock *inet_csk_accept(struct sock *sk, int flags, int *err)
        newsk = req->sk;

        sk_acceptq_removed(sk);
-       if (sk->sk_type == SOCK_STREAM && queue->fastopenq != NULL) {
+       if (sk->sk_protocol == IPPROTO_TCP && queue->fastopenq != NULL) {
                spin_lock_bh(&queue->fastopenq->lock);
                if (tcp_rsk(req)->listener) {
                        /* We are still waiting for the final ACK from 3WHS
@@ -775,7 +775,7 @@ void inet_csk_listen_stop(struct sock *sk)

                percpu_counter_inc(sk->sk_prot->orphan_count);

-               if (sk->sk_type == SOCK_STREAM && tcp_rsk(req)->listener) {
+               if (sk->sk_protocol == IPPROTO_TCP && tcp_rsk(req)->listener) {
                        BUG_ON(tcp_sk(child)->fastopen_rsk != req);
                        BUG_ON(sk != tcp_rsk(req)->listener);

Thanks, Eric.

Acked-by: H.K. Jerry Chu <redacted>

Re: [PATCH net-next] tcp: fix TFO regression

From: David Miller <davem@davemloft.net>
Date: 2012-09-06 18:23:34

From: Neal Cardwell <ncardwell@google.com>
Date: Thu, 6 Sep 2012 14:15:38 -0400
On Thu, Sep 6, 2012 at 2:07 PM, Eric Dumazet [off-list ref] wrote:
quoted
From: Eric Dumazet <edumazet@google.com>

Fengguang Wu reported various panics and bisected to commit
8336886f786fdac (tcp: TCP Fast Open Server - support TFO listeners)

Fix this by making sure socket is a TCP socket before accessing TFO data
structures.
...
quoted
Reported-by: Fengguang Wu <redacted>
Tested-by: Fengguang Wu <redacted>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: "H.K. Jerry Chu" <redacted>
Acked-by: Neal Cardwell <ncardwell@google.com>
Applied, thanks everyone.
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help