From: Eric Dumazet <hidden> Date: 2017-01-24 22:57:37
From: Eric Dumazet <edumazet@google.com>
tcp_add_backlog() can use skb_condense() helper to get better
gains and less SKB_TRUESIZE() magic. This only happens when socket
backlog has to be used.
Some attacks involve specially crafted out of order tiny TCP packets,
clogging the ofo queue of (many) sockets.
Then later, expensive collapse happens, trying to copy all these skbs
into single ones.
This unfortunately does not work if each skb has no neighbor in TCP
sequence order.
By using skb_condense() if the skb could not be coalesced to a prior
one, we defeat these kind of threats, potentially saving 4K per skb
(or more, since this is one page fragment).
A typical NAPI driver allocates gro packets with GRO_MAX_HEAD bytes
in skb->head, meaning the copy done by skb_condense() is limited to
about 200 bytes.
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
net/ipv4/tcp_input.c | 1 +
net/ipv4/tcp_ipv4.c | 3 +--
2 files changed, 2 insertions(+), 2 deletions(-)
From: David Miller <davem@davemloft.net> Date: 2017-01-25 18:17:19
From: Eric Dumazet <redacted>
Date: Tue, 24 Jan 2017 14:57:36 -0800
From: Eric Dumazet <edumazet@google.com>
tcp_add_backlog() can use skb_condense() helper to get better
gains and less SKB_TRUESIZE() magic. This only happens when socket
backlog has to be used.
Some attacks involve specially crafted out of order tiny TCP packets,
clogging the ofo queue of (many) sockets.
Then later, expensive collapse happens, trying to copy all these skbs
into single ones.
This unfortunately does not work if each skb has no neighbor in TCP
sequence order.
By using skb_condense() if the skb could not be coalesced to a prior
one, we defeat these kind of threats, potentially saving 4K per skb
(or more, since this is one page fragment).
A typical NAPI driver allocates gro packets with GRO_MAX_HEAD bytes
in skb->head, meaning the copy done by skb_condense() is limited to
about 200 bytes.
Signed-off-by: Eric Dumazet <edumazet@google.com>
From: Eric Dumazet <hidden> Date: 2017-01-25 18:39:03
On Wed, 2017-01-25 at 13:17 -0500, David Miller wrote:
Applied, thanks Eric.
Thanks David.
It looks IPv6 potential big network headers are also a threat :
Various pskb_may_pull() to pull headers might reallocate skb->head,
but skb->truesize is not updated in __pskb_pull_tail()
We probably need to update skb->truesize, but it is tricky as the prior
skb->truesize value might have been used for memory accounting when skb
was stored in some queue.
Do you think we could change __pskb_pull_tail() right away and fix the
few places that would break, or should we add various helpers with extra
parameters to take a safe route ?
From: David Miller <davem@davemloft.net> Date: 2017-01-25 19:03:21
From: Eric Dumazet <redacted>
Date: Wed, 25 Jan 2017 10:38:52 -0800
Do you think we could change __pskb_pull_tail() right away and fix the
few places that would break, or should we add various helpers with extra
parameters to take a safe route ?
It should always be safe as long as we see no socket attached on RX,
right?
That's the only real case where truesize adjustments can cause trouble.
From: Eric Dumazet <hidden> Date: 2017-01-25 21:40:14
On Wed, 2017-01-25 at 14:03 -0500, David Miller wrote:
From: Eric Dumazet <redacted>
Date: Wed, 25 Jan 2017 10:38:52 -0800
quoted
Do you think we could change __pskb_pull_tail() right away and fix the
few places that would break, or should we add various helpers with extra
parameters to take a safe route ?
It should always be safe as long as we see no socket attached on RX,
right?
That's the only real case where truesize adjustments can cause trouble.
Queue can be virtual, as for xmit path, tracking skb->truesize in
sk->sk_wmem_alloc.
If a layer calls pskb_may_pull(), we can not change skb->truesize
without also changing skb->sk->sk_wmem_alloc, or sock_wfree() will
trigger bugs.