From: Paolo Abeni <pabeni@redhat.com> Date: 2021-03-30 14:26:20
Currently the mentioned helper can end-up freeing the socket wmem
without waking-up any processes waiting for more write memory.
If the partially orphaned skb is attached to an UDP (or raw) socket,
the lack of wake-up can hang the user-space.
Address the issue invoking the write_space callback after
releasing the memory, if the old skb destructor requires that.
Fixes: f6ba8d33cfbb ("netem: fix skb_orphan_partial()")
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
---
net/core/sock.c | 2 ++
1 file changed, 2 insertions(+)
From: Eric Dumazet <edumazet@google.com> Date: 2021-03-30 14:40:25
On Tue, Mar 30, 2021 at 4:25 PM Paolo Abeni [off-list ref] wrote:
quoted hunk
Currently the mentioned helper can end-up freeing the socket wmem
without waking-up any processes waiting for more write memory.
If the partially orphaned skb is attached to an UDP (or raw) socket,
the lack of wake-up can hang the user-space.
Address the issue invoking the write_space callback after
releasing the memory, if the old skb destructor requires that.
Fixes: f6ba8d33cfbb ("netem: fix skb_orphan_partial()")
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
---
net/core/sock.c | 2 ++
1 file changed, 2 insertions(+)
Interesting.
Why TCP is not a problem here ?
I would rather replace WARN_ON(refcount_sub_and_test(skb->truesize,
&sk->sk_wmem_alloc)) by :
skb_orphan(skb);
This will get rid of this suspect WARN_ON() at the same time ?
From: Eric Dumazet <edumazet@google.com> Date: 2021-03-30 14:41:59
On Tue, Mar 30, 2021 at 4:39 PM Eric Dumazet [off-list ref] wrote:
On Tue, Mar 30, 2021 at 4:25 PM Paolo Abeni [off-list ref] wrote:
quoted
Currently the mentioned helper can end-up freeing the socket wmem
without waking-up any processes waiting for more write memory.
If the partially orphaned skb is attached to an UDP (or raw) socket,
the lack of wake-up can hang the user-space.
Address the issue invoking the write_space callback after
releasing the memory, if the old skb destructor requires that.
Fixes: f6ba8d33cfbb ("netem: fix skb_orphan_partial()")
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
---
net/core/sock.c | 2 ++
1 file changed, 2 insertions(+)
Interesting.
Why TCP is not a problem here ?
I would rather replace WARN_ON(refcount_sub_and_test(skb->truesize,
&sk->sk_wmem_alloc)) by :
skb_orphan(skb);
And of course re-add
skb->sk = sk;
This will get rid of this suspect WARN_ON() at the same time ?
From: Paolo Abeni <pabeni@redhat.com> Date: 2021-03-30 15:19:21
On Tue, 2021-03-30 at 16:40 +0200, Eric Dumazet wrote:
On Tue, Mar 30, 2021 at 4:39 PM Eric Dumazet [off-list ref] wrote:
quoted
On Tue, Mar 30, 2021 at 4:25 PM Paolo Abeni [off-list ref] wrote:
quoted
Currently the mentioned helper can end-up freeing the socket wmem
without waking-up any processes waiting for more write memory.
If the partially orphaned skb is attached to an UDP (or raw) socket,
the lack of wake-up can hang the user-space.
Address the issue invoking the write_space callback after
releasing the memory, if the old skb destructor requires that.
Fixes: f6ba8d33cfbb ("netem: fix skb_orphan_partial()")
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
---
net/core/sock.c | 2 ++
1 file changed, 2 insertions(+)
AFAICS, tcp_wfree() does not call sk->sk_write_space(). Processes
waiting for wmem are woken by ack processing.
quoted
I would rather replace WARN_ON(refcount_sub_and_test(skb->truesize,
&sk->sk_wmem_alloc)) by :
skb_orphan(skb);
And of course re-add
skb->sk = sk;
Double checking to be sure. The patched slice of skb_orphan_partial()
will then look like:
if (can_skb_orphan_partial(skb)) {
struct sock *sk = skb->sk;
if (refcount_inc_not_zero(&sk->sk_refcnt)) {
skb_orphan(skb);
skb->sk = sk;
skb->destructor = sock_efree;
}
} // ...
Am I correct?
Thanks!
Paolo
From: Eric Dumazet <edumazet@google.com> Date: 2021-03-30 15:31:11
On Tue, Mar 30, 2021 at 5:18 PM Paolo Abeni [off-list ref] wrote:
On Tue, 2021-03-30 at 16:40 +0200, Eric Dumazet wrote:
quoted
On Tue, Mar 30, 2021 at 4:39 PM Eric Dumazet [off-list ref] wrote:
quoted
On Tue, Mar 30, 2021 at 4:25 PM Paolo Abeni [off-list ref] wrote:
quoted
Currently the mentioned helper can end-up freeing the socket wmem
without waking-up any processes waiting for more write memory.
If the partially orphaned skb is attached to an UDP (or raw) socket,
the lack of wake-up can hang the user-space.
Address the issue invoking the write_space callback after
releasing the memory, if the old skb destructor requires that.
Fixes: f6ba8d33cfbb ("netem: fix skb_orphan_partial()")
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
---
net/core/sock.c | 2 ++
1 file changed, 2 insertions(+)
AFAICS, tcp_wfree() does not call sk->sk_write_space(). Processes
waiting for wmem are woken by ack processing.
quoted
quoted
I would rather replace WARN_ON(refcount_sub_and_test(skb->truesize,
&sk->sk_wmem_alloc)) by :
skb_orphan(skb);
And of course re-add
skb->sk = sk;
Double checking to be sure. The patched slice of skb_orphan_partial()
will then look like:
if (can_skb_orphan_partial(skb)) {
struct sock *sk = skb->sk;
if (refcount_inc_not_zero(&sk->sk_refcnt)) {
skb_orphan(skb);
skb->sk = sk;
skb->destructor = sock_efree;
}
} // ...
Am I correct?
Yes.
We also could add a helper for the whole construct, since many other
paths do almost the same.
(They might use sock_hold(), but it seems safe to use the
refcount_inc_not_zero())
Or they omit the skb_orphan() (see can_skb_set_owner), which seems also risky.
static inline void skb_set_owner_sk_safe(sk, skb)
{
if (sk && refcount_inc_not_zero(&sk->sk_refcnt)) {
skb_orphan(skb);
skb->destructor = sock_efree;
skb->sk = sk;
}
}
From: Eric Dumazet <edumazet@google.com> Date: 2021-03-30 15:33:20
On Tue, Mar 30, 2021 at 5:18 PM Paolo Abeni [off-list ref] wrote:
On Tue, 2021-03-30 at 16:40 +0200, Eric Dumazet wrote:
quoted
quoted
Why TCP is not a problem here ?
AFAICS, tcp_wfree() does not call sk->sk_write_space(). Processes
waiting for wmem are woken by ack processing.
My concern was TCP Small Queue. If we do not call tcp_wfree() we might miss
an opportunity to queue more packets (and thus fallback to RTO or
incoming ACK is we are lucky)