From: peter enderborg <hidden> Date: 2016-06-17 13:58:31
From: Peter Enderborg <redacted>
When sending data the socket allocates memory for
payload on a cache or a page alloc. The page alloc
then might trigger compation that takes long time.
This can be avoided with smaller chunks. But
userspace can not know what is the right size for
the smaller sends. For this we add a SIZEHINT
getsocketopt where the userspace can get the size
for send that will fit into one page (order 0) or
the max for a slab cache allocation.
Signed-off-by: Peter Enderborg <redacted>
---
include/uapi/asm-generic/socket.h | 2 ++
include/uapi/linux/socket.h | 9 +++++++++
net/core/sock.c | 17 +++++++++++++++++
3 files changed, 28 insertions(+)
@@ -18,4 +20,11 @@ struct __kernel_sockaddr_storage {/* _SS_MAXSIZE value minus size of ss_family */}__attribute__((aligned(_K_SS_ALIGNSIZE)));/* force desired alignment */+structsock_sizehint{+__u32order_zero_size;+/* max payload size that can fit into one page in kernel */+__u32cache_size;+/* max payload size that can fit in socket slab cache */+};+#endif /* _UAPI_LINUX_SOCKET_H */
@@ -1254,6 +1254,23 @@ int sock_getsockopt(struct socket *sock, int level, int optname,v.val=sk->sk_incoming_cpu;break;+caseSO_SIZEHINT:+{+structsock_sizehinthint;++if(len>sizeof(hint))+len=sizeof(hint);++hint.order_zero_size=PAGE_SIZE-+SKB_DATA_ALIGN(sizeof(structskb_shared_info));+hint.cache_size=KMALLOC_MAX_CACHE_SIZE-+SKB_DATA_ALIGN(sizeof(structskb_shared_info));++if(copy_to_user(optval,&hint,len))+return-EFAULT;+gotolenout;+}+default:/* We implement the SO_SNDLOWAT etc to not be settable*(1003.1g7).
From: Eric Dumazet <hidden> Date: 2016-06-17 14:14:25
On Fri, 2016-06-17 at 15:58 +0200, peter enderborg wrote:
From: Peter Enderborg <redacted>
When sending data the socket allocates memory for
payload on a cache or a page alloc. The page alloc
then might trigger compation that takes long time.
This can be avoided with smaller chunks. But
userspace can not know what is the right size for
the smaller sends. For this we add a SIZEHINT
getsocketopt where the userspace can get the size
for send that will fit into one page (order 0) or
the max for a slab cache allocation.
For which kind of sockets exactly you hit a problem ?
Sorry, this patch is probably not helping in any way.
From: peter enderborg <hidden> Date: 2016-06-17 14:39:51
On 06/17/2016 04:14 PM, Eric Dumazet wrote:
On Fri, 2016-06-17 at 15:58 +0200, peter enderborg wrote:
quoted
From: Peter Enderborg <redacted>
When sending data the socket allocates memory for
payload on a cache or a page alloc. The page alloc
then might trigger compation that takes long time.
This can be avoided with smaller chunks. But
userspace can not know what is the right size for
the smaller sends. For this we add a SIZEHINT
getsocketopt where the userspace can get the size
for send that will fit into one page (order 0) or
the max for a slab cache allocation.
For which kind of sockets exactly you hit a problem ?
Sorry, this patch is probably not helping in any way.
It is mainly for af_unix sockets, and the effect is
quite significant when you hit a compaction, or with
this patch avoid get in to compaction, but it
can also be used for reducing the pressure on memory
for tcp. And the patches you suggested have been
applied (with the addition "af_unix: fix bug on large send()")
I see that there is a lot of other compaction fixes
recently but the problem are still there. And of course
to make any difference you need to change your
userland application too. But in our Qualcomm/Google
bastard to kernel. It makes a huge difference on the
behaviour of send(). But I also does not see this as
perfect solution. A wake-up function that has
the buffers reserved would be better.Or a pre allocated
send buffer would also be better. But I dont expect that
linux will have a real-time socket implementation in
near future.
From: Eric Dumazet <hidden> Date: 2016-06-17 16:04:00
On Fri, 2016-06-17 at 16:39 +0200, peter enderborg wrote:
On 06/17/2016 04:14 PM, Eric Dumazet wrote:
quoted
On Fri, 2016-06-17 at 15:58 +0200, peter enderborg wrote:
quoted
From: Peter Enderborg <redacted>
When sending data the socket allocates memory for
payload on a cache or a page alloc. The page alloc
then might trigger compation that takes long time.
This can be avoided with smaller chunks. But
userspace can not know what is the right size for
the smaller sends. For this we add a SIZEHINT
getsocketopt where the userspace can get the size
for send that will fit into one page (order 0) or
the max for a slab cache allocation.
For which kind of sockets exactly you hit a problem ?
Sorry, this patch is probably not helping in any way.
It is mainly for af_unix sockets, and the effect is
quite significant when you hit a compaction, or with
this patch avoid get in to compaction, but it
can also be used for reducing the pressure on memory
for tcp. And the patches you suggested have been
applied (with the addition "af_unix: fix bug on large send()")
I see that there is a lot of other compaction fixes
recently but the problem are still there. And of course
to make any difference you need to change your
userland application too. But in our Qualcomm/Google
bastard to kernel. It makes a huge difference on the
behaviour of send(). But I also does not see this as
perfect solution. A wake-up function that has
the buffers reserved would be better.Or a pre allocated
send buffer would also be better. But I dont expect that
linux will have a real-time socket implementation in
near future.
I have no evidence the problem you describe still exists in current
linux kernels.
Please patch your kernels, but do not send networking patches that seem
to work around a mm-layer problem, without notifying mm maintainers.
From: Eric Dumazet <hidden> Date: 2016-06-17 16:07:09
On Fri, 2016-06-17 at 16:39 +0200, peter enderborg wrote:
It is mainly for af_unix sockets, and the effect is
quite significant when you hit a compaction, or with
this patch avoid get in to compaction, but it
can also be used for reducing the pressure on memory
for tcp.
BTW, TCP always attempt order-3 allocations, even if you do a write(fd,
buffer, 4000)
So really your patch wont help.
We need to fix the mm layer (if needed), not add various works around.