From: Jakub Kicinski <kuba@kernel.org> Date: 2021-06-21 23:13:11
Dave observed number of machines hitting OOM on the UDP send
path. The workload seems to be sending large UDP packets over
loopback. Since loopback has MTU of 64k kernel will try to
allocate an skb with up to 64k of head space. This has a good
chance of failing under memory pressure. What's worse if
the message length is <32k the allocation may trigger an
OOM killer.
This is entirely avoidable, we can use an skb with frags.
The scenario is unlikely and always using frags requires
an extra allocation so opt for using fallback, rather
then always using frag'ed/paged skb when payload is large.
Note that the size heuristic (header_len > PAGE_SIZE)
is not entirely accurate, __alloc_skb() will add ~400B
to size. Occasional order-1 allocation should be fine,
though, we are primarily concerned with order-3.
Reported-by: Dave Jones <redacted>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
---
include/net/sock.h | 11 +++++++++++
net/ipv4/ip_output.c | 19 +++++++++++++++++--
net/ipv6/ip6_output.c | 19 +++++++++++++++++--
3 files changed, 45 insertions(+), 4 deletions(-)
From: Paolo Abeni <pabeni@redhat.com> Date: 2021-06-22 10:07:40
On Mon, 2021-06-21 at 16:13 -0700, Jakub Kicinski wrote:
Dave observed number of machines hitting OOM on the UDP send
path. The workload seems to be sending large UDP packets over
loopback. Since loopback has MTU of 64k kernel will try to
allocate an skb with up to 64k of head space. This has a good
chance of failing under memory pressure. What's worse if
the message length is <32k the allocation may trigger an
OOM killer.
Out of sheer curiosity, are there a large number of UDP sockets in such
workload? did you increase rmem_default/rmem_max? If so, could tuning
udp_mem help?
From: Eric Dumazet <hidden> Date: 2021-06-22 14:12:17
On 6/22/21 1:13 AM, Jakub Kicinski wrote:
quoted hunk
Dave observed number of machines hitting OOM on the UDP send
path. The workload seems to be sending large UDP packets over
loopback. Since loopback has MTU of 64k kernel will try to
allocate an skb with up to 64k of head space. This has a good
chance of failing under memory pressure. What's worse if
the message length is <32k the allocation may trigger an
OOM killer.
This is entirely avoidable, we can use an skb with frags.
The scenario is unlikely and always using frags requires
an extra allocation so opt for using fallback, rather
then always using frag'ed/paged skb when payload is large.
Note that the size heuristic (header_len > PAGE_SIZE)
is not entirely accurate, __alloc_skb() will add ~400B
to size. Occasional order-1 allocation should be fine,
though, we are primarily concerned with order-3.
Reported-by: Dave Jones <redacted>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
---
include/net/sock.h | 11 +++++++++++
net/ipv4/ip_output.c | 19 +++++++++++++++++--
net/ipv6/ip6_output.c | 19 +++++++++++++++++--
3 files changed, 45 insertions(+), 4 deletions(-)
@@ -1095,9 +1095,24 @@ static int __ip_append_data(struct sock *sk,alloclen+=rt->dst.trailer_len;if(transhdrlen){-skb=sock_alloc_send_skb(sk,-alloclen+hh_len+15,+size_theader_len=alloclen+hh_len+15;+gfp_tsk_allocation;++if(header_len>PAGE_SIZE)+sk_allocation_push(sk,__GFP_NORETRY,+&sk_allocation);+skb=sock_alloc_send_skb(sk,header_len,(flags&MSG_DONTWAIT),&err);+if(header_len>PAGE_SIZE){+BUILD_BUG_ON(MAX_HEADER>=PAGE_SIZE);++sk_allocation_pop(sk,sk_allocation);+if(unlikely(!skb)&&!paged&&+rt->dst.dev->features&NETIF_F_SG){+paged=true;+gotoalloc_new_skb;+}+}
What about using sock_alloc_send_pskb(... PAGE_ALLOC_COSTLY_ORDER)
(as we did in unix_dgram_sendmsg() for large packets), for SG enabled interfaces ?
We do not _have_ to put all the payload in skb linear part,
we could instead use page frags (order-0 if high order pages are not available)
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-06-22 16:54:25
On Tue, 22 Jun 2021 16:12:11 +0200 Eric Dumazet wrote:
On 6/22/21 1:13 AM, Jakub Kicinski wrote:
quoted
Dave observed number of machines hitting OOM on the UDP send
path. The workload seems to be sending large UDP packets over
loopback. Since loopback has MTU of 64k kernel will try to
allocate an skb with up to 64k of head space. This has a good
chance of failing under memory pressure. What's worse if
the message length is <32k the allocation may trigger an
OOM killer.
This is entirely avoidable, we can use an skb with frags.
The scenario is unlikely and always using frags requires
an extra allocation so opt for using fallback, rather
then always using frag'ed/paged skb when payload is large.
Note that the size heuristic (header_len > PAGE_SIZE)
is not entirely accurate, __alloc_skb() will add ~400B
to size. Occasional order-1 allocation should be fine,
though, we are primarily concerned with order-3.
Reported-by: Dave Jones <redacted>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
@@ -1095,9 +1095,24 @@ static int __ip_append_data(struct sock *sk,alloclen+=rt->dst.trailer_len;if(transhdrlen){-skb=sock_alloc_send_skb(sk,-alloclen+hh_len+15,+size_theader_len=alloclen+hh_len+15;+gfp_tsk_allocation;++if(header_len>PAGE_SIZE)+sk_allocation_push(sk,__GFP_NORETRY,+&sk_allocation);+skb=sock_alloc_send_skb(sk,header_len,(flags&MSG_DONTWAIT),&err);+if(header_len>PAGE_SIZE){+BUILD_BUG_ON(MAX_HEADER>=PAGE_SIZE);++sk_allocation_pop(sk,sk_allocation);+if(unlikely(!skb)&&!paged&&+rt->dst.dev->features&NETIF_F_SG){+paged=true;+gotoalloc_new_skb;+}+}
What about using sock_alloc_send_pskb(... PAGE_ALLOC_COSTLY_ORDER)
(as we did in unix_dgram_sendmsg() for large packets), for SG enabled interfaces ?
PAGE_ALLOC_COSTLY_ORDER in itself is more of a problem than a solution.
AFAIU the app sends messages primarily above the ~60kB mark, which is
above COSTLY, and those do not trigger OOM kills. All OOM kills we see
have order=3. Checking with Rik and Johannes W that's expected, OOM
killer is only invoked for allocations <= COSTLY, larger ones will just
return NULL and let us deal with it (e.g. by falling back).
So adding GFP_NORETRY is key for 0 < order <= COSTLY,
skb_page_frag_refill()-style.
We do not _have_ to put all the payload in skb linear part,
we could instead use page frags (order-0 if high order pages are not available)
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-06-22 16:57:28
On Tue, 22 Jun 2021 12:07:27 +0200 Paolo Abeni wrote:
On Mon, 2021-06-21 at 16:13 -0700, Jakub Kicinski wrote:
quoted
Dave observed number of machines hitting OOM on the UDP send
path. The workload seems to be sending large UDP packets over
loopback. Since loopback has MTU of 64k kernel will try to
allocate an skb with up to 64k of head space. This has a good
chance of failing under memory pressure. What's worse if
the message length is <32k the allocation may trigger an
OOM killer.
Out of sheer curiosity, are there a large number of UDP sockets in such
workload? did you increase rmem_default/rmem_max? If so, could tuning
udp_mem help?
This is not thread safe.
Remember UDP sendmsg() does not lock the socket for non-corking sends.
Ugh, you're right :(
Hm, isn't it buggy to call sock_alloc_send_[p]skb() without holding the
lock in the first place, then? The knee jerk fix would be to add another
layer of specialization to the helpers:
From: Eric Dumazet <hidden> Date: 2021-06-22 17:48:49
On 6/22/21 6:54 PM, Jakub Kicinski wrote:
On Tue, 22 Jun 2021 16:12:11 +0200 Eric Dumazet wrote:
quoted
On 6/22/21 1:13 AM, Jakub Kicinski wrote:
quoted
Dave observed number of machines hitting OOM on the UDP send
path. The workload seems to be sending large UDP packets over
loopback. Since loopback has MTU of 64k kernel will try to
allocate an skb with up to 64k of head space. This has a good
chance of failing under memory pressure. What's worse if
the message length is <32k the allocation may trigger an
OOM killer.
This is entirely avoidable, we can use an skb with frags.
The scenario is unlikely and always using frags requires
an extra allocation so opt for using fallback, rather
then always using frag'ed/paged skb when payload is large.
Note that the size heuristic (header_len > PAGE_SIZE)
is not entirely accurate, __alloc_skb() will add ~400B
to size. Occasional order-1 allocation should be fine,
though, we are primarily concerned with order-3.
Reported-by: Dave Jones <redacted>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
@@ -1095,9 +1095,24 @@ static int __ip_append_data(struct sock *sk,alloclen+=rt->dst.trailer_len;if(transhdrlen){-skb=sock_alloc_send_skb(sk,-alloclen+hh_len+15,+size_theader_len=alloclen+hh_len+15;+gfp_tsk_allocation;++if(header_len>PAGE_SIZE)+sk_allocation_push(sk,__GFP_NORETRY,+&sk_allocation);+skb=sock_alloc_send_skb(sk,header_len,(flags&MSG_DONTWAIT),&err);+if(header_len>PAGE_SIZE){+BUILD_BUG_ON(MAX_HEADER>=PAGE_SIZE);++sk_allocation_pop(sk,sk_allocation);+if(unlikely(!skb)&&!paged&&+rt->dst.dev->features&NETIF_F_SG){+paged=true;+gotoalloc_new_skb;+}+}
What about using sock_alloc_send_pskb(... PAGE_ALLOC_COSTLY_ORDER)
(as we did in unix_dgram_sendmsg() for large packets), for SG enabled interfaces ?
PAGE_ALLOC_COSTLY_ORDER in itself is more of a problem than a solution.
AFAIU the app sends messages primarily above the ~60kB mark, which is
above COSTLY, and those do not trigger OOM kills. All OOM kills we see
have order=3. Checking with Rik and Johannes W that's expected, OOM
killer is only invoked for allocations <= COSTLY, larger ones will just
return NULL and let us deal with it (e.g. by falling back).
I really thought alloc_skb_with_frags() was already handling low-memory-conditions.
(alloc_skb_with_frags() is called from sock_alloc_send_pskb())
If it is not, lets fix it, because af_unix sockets will have the same issue ?
So adding GFP_NORETRY is key for 0 < order <= COSTLY,
skb_page_frag_refill()-style.
quoted
We do not _have_ to put all the payload in skb linear part,
we could instead use page frags (order-0 if high order pages are not available)
This is not thread safe.
Remember UDP sendmsg() does not lock the socket for non-corking sends.
Ugh, you're right :(
Hm, isn't it buggy to call sock_alloc_send_[p]skb() without holding the
lock in the first place, then? The knee jerk fix would be to add another
layer of specialization to the helpers:
It is not buggy. Please elaborate if you found it is.
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-06-22 18:09:51
On Tue, 22 Jun 2021 19:48:43 +0200 Eric Dumazet wrote:
quoted
quoted
What about using sock_alloc_send_pskb(... PAGE_ALLOC_COSTLY_ORDER)
(as we did in unix_dgram_sendmsg() for large packets), for SG enabled interfaces ?
PAGE_ALLOC_COSTLY_ORDER in itself is more of a problem than a solution.
AFAIU the app sends messages primarily above the ~60kB mark, which is
above COSTLY, and those do not trigger OOM kills. All OOM kills we see
have order=3. Checking with Rik and Johannes W that's expected, OOM
killer is only invoked for allocations <= COSTLY, larger ones will just
return NULL and let us deal with it (e.g. by falling back).
I really thought alloc_skb_with_frags() was already handling low-memory-conditions.
(alloc_skb_with_frags() is called from sock_alloc_send_pskb())
If it is not, lets fix it, because af_unix sockets will have the same issue ?
af_unix seems to cap at SKB_MAX_ALLOC which is order 2, AFAICT.
Perhaps that's a good enough fix in practice given we see OOMs with
order=3 only?
I'll review callers of alloc_skb_with_frags() and see if they depend
on the explicit geometry of the skb or we can safely fallback to pages.
From: Eric Dumazet <hidden> Date: 2021-06-22 18:48:03
On 6/22/21 8:09 PM, Jakub Kicinski wrote:
On Tue, 22 Jun 2021 19:48:43 +0200 Eric Dumazet wrote:
quoted
quoted
quoted
What about using sock_alloc_send_pskb(... PAGE_ALLOC_COSTLY_ORDER)
(as we did in unix_dgram_sendmsg() for large packets), for SG enabled interfaces ?
PAGE_ALLOC_COSTLY_ORDER in itself is more of a problem than a solution.
AFAIU the app sends messages primarily above the ~60kB mark, which is
above COSTLY, and those do not trigger OOM kills. All OOM kills we see
have order=3. Checking with Rik and Johannes W that's expected, OOM
killer is only invoked for allocations <= COSTLY, larger ones will just
return NULL and let us deal with it (e.g. by falling back).
I really thought alloc_skb_with_frags() was already handling low-memory-conditions.
(alloc_skb_with_frags() is called from sock_alloc_send_pskb())
If it is not, lets fix it, because af_unix sockets will have the same issue ?
af_unix seems to cap at SKB_MAX_ALLOC which is order 2, AFAICT.
It does not cap to SKB_MAX_ALLOC.
It definitely attempt big allocations if you send 64KB datagrams.
Please look at commit d14b56f508ad70eca3e659545aab3c45200f258c
net: cleanup gfp mask in alloc_skb_with_frags
This explains why we do not have __GFP_NORETRY there.
Perhaps that's a good enough fix in practice given we see OOMs with
order=3 only?
I'll review callers of alloc_skb_with_frags() and see if they depend
on the explicit geometry of the skb or we can safely fallback to pages.
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-06-22 19:04:35
On Tue, 22 Jun 2021 20:47:57 +0200 Eric Dumazet wrote:
On 6/22/21 8:09 PM, Jakub Kicinski wrote:
quoted
On Tue, 22 Jun 2021 19:48:43 +0200 Eric Dumazet wrote:
quoted
I really thought alloc_skb_with_frags() was already handling low-memory-conditions.
(alloc_skb_with_frags() is called from sock_alloc_send_pskb())
If it is not, lets fix it, because af_unix sockets will have the same issue ?
af_unix seems to cap at SKB_MAX_ALLOC which is order 2, AFAICT.
It does not cap to SKB_MAX_ALLOC.
It definitely attempt big allocations if you send 64KB datagrams.
Please look at commit d14b56f508ad70eca3e659545aab3c45200f258c
net: cleanup gfp mask in alloc_skb_with_frags
This explains why we do not have __GFP_NORETRY there.
Ah, right, slight misunderstanding.
Just to be 100% clear for UDP send we are allocating up to 64kB
in the _head_, AFAICT. Allocation of head does not clear GFP_WAIT.
Your memory was correct, alloc_skb_with_frags() does handle low-memory
when it comes to allocating frags. And what I was saying is af_unix
won't have the same problem as UDP as it caps head's size at
SKB_MAX_ALLOC, and frags are allocated with fallback.
For the UDP case we can either adapt the af_unix approach, and cap head
size to SKB_MAX_ALLOC or try to allocate the full skb and fall back.
Having alloc_skb_with_frags() itself re-balance head <> data
automatically does not feel right, no?
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-06-22 19:51:21
On Tue, 22 Jun 2021 12:04:26 -0700 Jakub Kicinski wrote:
For the UDP case we can either adapt the af_unix approach, and cap head
size to SKB_MAX_ALLOC or try to allocate the full skb and fall back.
Having alloc_skb_with_frags() itself re-balance head <> data
automatically does not feel right, no?
Actually looking closer at the UDP code it appears it only uses the
giant head it allocated if the underlying device doesn't have SG.
We can make the head smaller and probably only improve performance
for 99% of deployments. I'll send a v2.