From: Eric Dumazet <edumazet@google.com> Date: 2016-12-08 19:42:03
This patch series provides about 100 % performance increase under flood.
v2: added Paolo feedback on udp_rmem_release() for tiny sk_rcvbuf
added the last patch touching sk_rmem_alloc later
Eric Dumazet (4):
udp: add busylocks in RX path
udp: copy skb->truesize in the first cache line
udp: add batching to udp_rmem_release()
udp: udp_rmem_release() should touch sk_rmem_alloc later
include/linux/skbuff.h | 9 ++++++-
include/linux/udp.h | 3 +++
net/ipv4/udp.c | 71 ++++++++++++++++++++++++++++++++++++++++++++++----
3 files changed, 77 insertions(+), 6 deletions(-)
--
2.8.0.rc3.226.g39d4020
From: Eric Dumazet <edumazet@google.com> Date: 2016-12-08 19:42:05
Idea of busylocks is to let producers grab an extra spinlock
to relieve pressure on the receive_queue spinlock shared by consumer.
This behavior is requested only once socket receive queue is above
half occupancy.
Under flood, this means that only one producer can be in line
trying to acquire the receive_queue spinlock.
These busylock can be allocated on a per cpu manner, instead of a
per socket one (that would consume a cache line per socket)
This patch considerably improves UDP behavior under stress,
depending on number of NIC RX queues and/or RPS spread.
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
net/ipv4/udp.c | 43 ++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 42 insertions(+), 1 deletion(-)
@@ -1195,10 +1195,36 @@ void udp_skb_destructor(struct sock *sk, struct sk_buff *skb)}EXPORT_SYMBOL(udp_skb_destructor);+/* Idea of busylocks is to let producers grab an extra spinlock+*torelievepressureonthereceive_queuespinlocksharedbyconsumer.+*Underflood,thismeansthatonlyoneproducercanbeinline+*tryingtoacquirethereceive_queuespinlock.+*Thesebusylockcanbeallocatedonapercpumanner,insteadofa+*persocketone(thatwouldconsumeacachelinepersocket)+*/+staticintudp_busylocks_log__read_mostly;+staticspinlock_t*udp_busylocks__read_mostly;++staticspinlock_t*busylock_acquire(void*ptr)+{+spinlock_t*busy;++busy=udp_busylocks+hash_ptr(ptr,udp_busylocks_log);+spin_lock(busy);+returnbusy;+}++staticvoidbusylock_release(spinlock_t*busy)+{+if(busy)+spin_unlock(busy);+}+int__udp_enqueue_schedule_skb(structsock*sk,structsk_buff*skb){structsk_buff_head*list=&sk->sk_receive_queue;intrmem,delta,amt,err=-ENOMEM;+spinlock_t*busy=NULL;intsize;/* try to avoid the costly atomic add/sub pair when the receive
@@ -1214,8 +1240,11 @@ int __udp_enqueue_schedule_skb(struct sock *sk, struct sk_buff *skb)*-Lesscachelinemissesatcopyout()time*-Lessworkatconsume_skb()(lessalienpagefragfreeing)*/-if(rmem>(sk->sk_rcvbuf>>1))+if(rmem>(sk->sk_rcvbuf>>1)){skb_condense(skb);++busy=busylock_acquire(sk);+}size=skb->truesize;/* we drop only if the receive buf is full and the receive
From: Eric Dumazet <edumazet@google.com> Date: 2016-12-08 19:42:07
In UDP RX handler, we currently clear skb->dev before skb
is added to receive queue, because device pointer is no longer
available once we exit from RCU section.
Since this first cache line is always hot, lets reuse this space
to store skb->truesize and thus avoid a cache line miss at
udp_recvmsg()/udp_skb_destructor time while receive queue
spinlock is held.
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
include/linux/skbuff.h | 9 ++++++++-
net/ipv4/udp.c | 13 ++++++++++---
2 files changed, 18 insertions(+), 4 deletions(-)
@@ -645,8 +645,15 @@ struct sk_buff {structrb_noderbnode;/* used in netem & tcp stack */};structsock*sk;-structnet_device*dev;+union{+structnet_device*dev;+/* Some protocols might use this space to store information,+*whiledevicepointerwouldbeNULL.+*UDPreceivepathisoneuser.+*/+unsignedlongdev_scratch;+};/**Thisisthecontrolbuffer.Itisfreetouseforevery*layer.Pleaseputyourprivatevariablesthere.Ifyou
@@ -1188,10 +1188,14 @@ static void udp_rmem_release(struct sock *sk, int size, int partial)__sk_mem_reduce_allocated(sk,amt>>SK_MEM_QUANTUM_SHIFT);}-/* Note: called with sk_receive_queue.lock held */+/* Note: called with sk_receive_queue.lock held.+*Insteadofusingskb->truesizehere,findacopyofitinskb->dev_scratch+*Thisavoidsacachelinemisswhilereceive_queuelockisheld.+*Lookat__udp_enqueue_schedule_skb()tofindwherethiscopyisdone.+*/voidudp_skb_destructor(structsock*sk,structsk_buff*skb){-udp_rmem_release(sk,skb->truesize,1);+udp_rmem_release(sk,skb->dev_scratch,1);}EXPORT_SYMBOL(udp_skb_destructor);
@@ -1246,6 +1250,10 @@ int __udp_enqueue_schedule_skb(struct sock *sk, struct sk_buff *skb)busy=busylock_acquire(sk);}size=skb->truesize;+/* Copy skb->truesize into skb->dev_scratch to avoid a cache line miss+*inudp_skb_destructor()+*/+skb->dev_scratch=size;/* we drop only if the receive buf is full and the receive*queuecontainssomeotherskb
@@ -1272,7 +1280,6 @@ int __udp_enqueue_schedule_skb(struct sock *sk, struct sk_buff *skb)/* no need to setup a destructor, we will explicitly release the*forwardallocatedmemoryondequeue*/-skb->dev=NULL;sock_skb_set_dropcount(sk,skb);__skb_queue_tail(list,skb);
From: Eric Dumazet <edumazet@google.com> Date: 2016-12-08 19:42:09
If udp_recvmsg() constantly releases sk_rmem_alloc
for every read packet, it gives opportunity for
producers to immediately grab spinlocks and desperatly
try adding another packet, causing false sharing.
We can add a simple heuristic to give the signal
by batches of ~25 % of the queue capacity.
This patch considerably increases performance under
flood by about 50 %, since the thread draining the queue
is no longer slowed by false sharing.
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
include/linux/udp.h | 3 +++
net/ipv4/udp.c | 12 ++++++++++++
2 files changed, 15 insertions(+)
@@ -79,6 +79,9 @@ struct udp_sock {int(*gro_complete)(structsock*sk,structsk_buff*skb,intnhoff);++/* This field is dirtied by udp_recvmsg() */+intforward_deficit;};staticinlinestructudp_sock*udp_sk(conststructsock*sk)
From: Eric Dumazet <edumazet@google.com> Date: 2016-12-08 19:42:12
In flood situations, keeping sk_rmem_alloc at a high value
prevents producers from touching the socket.
It makes sense to lower sk_rmem_alloc only at the end
of udp_rmem_release() after the thread draining receive
queue in udp_recvmsg() finished the writes to sk_forward_alloc.
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
net/ipv4/udp.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
@@ -1191,13 +1191,14 @@ static void udp_rmem_release(struct sock *sk, int size, int partial)}up->forward_deficit=0;-atomic_sub(size,&sk->sk_rmem_alloc);sk->sk_forward_alloc+=size;amt=(sk->sk_forward_alloc-partial)&~(SK_MEM_QUANTUM-1);sk->sk_forward_alloc-=amt;if(amt)__sk_mem_reduce_allocated(sk,amt>>SK_MEM_QUANTUM_SHIFT);++atomic_sub(size,&sk->sk_rmem_alloc);}/* Note: called with sk_receive_queue.lock held.
From: Hannes Frederic Sowa <hidden> Date: 2016-12-08 19:45:51
Hi Eric,
On Thu, Dec 8, 2016, at 20:41, Eric Dumazet wrote:
Idea of busylocks is to let producers grab an extra spinlock
to relieve pressure on the receive_queue spinlock shared by consumer.
This behavior is requested only once socket receive queue is above
half occupancy.
Under flood, this means that only one producer can be in line
trying to acquire the receive_queue spinlock.
These busylock can be allocated on a per cpu manner, instead of a
per socket one (that would consume a cache line per socket)
This patch considerably improves UDP behavior under stress,
depending on number of NIC RX queues and/or RPS spread.
This patch mostly improves situation for non-connected sockets. Do you
think it makes sense to acquire the spinlock depending on the sockets
state? Connected UDP sockets flow in on one CPU anyway?
Otherwise the series looks really great, thanks!
From: Eric Dumazet <hidden> Date: 2016-12-08 19:58:31
On Thu, 2016-12-08 at 20:45 +0100, Hannes Frederic Sowa wrote:
Hi Eric,
This patch mostly improves situation for non-connected sockets. Do you
think it makes sense to acquire the spinlock depending on the sockets
state? Connected UDP sockets flow in on one CPU anyway?
We could do that, definitely.
However I could not measure any difference with a conditional test here
for connected socket.
Maybe it would show up if we are unlucky and multiple cpus compete on
the same cache line because of the hashed spinlock array.
From: David Miller <davem@davemloft.net> Date: 2016-12-10 03:13:16
From: Eric Dumazet <edumazet@google.com>
Date: Thu, 8 Dec 2016 11:41:53 -0800
This patch series provides about 100 % performance increase under flood.
v2: added Paolo feedback on udp_rmem_release() for tiny sk_rcvbuf
added the last patch touching sk_rmem_alloc later