From: "D. Wythe" <alibuda@linux.alibaba.com>
This patch set aims to optimizing performance of SMC in short-lived
links scenarios, which is quite unsatisfactory right now.
In our benchmark, we test it with follow scripts:
./wrk -c 10000 -t 4 -H 'Connection: Close' -d 20 http://smc-server
Current performance figures like that:
Running 20s test @ http://11.213.45.6
4 threads and 10000 connections
4956 requests in 20.06s, 3.24MB read
Socket errors: connect 0, read 0, write 672, timeout 0
Requests/sec: 247.07
Transfer/sec: 165.28KB
There are many reasons for this phenomenon, this patch set doesn't
solve it all though, but it can be well alleviated with it in.
Patch 1/5 (Make smc_tcp_listen_work() independent) :
Separate smc_tcp_listen_work() from smc_listen_work(), make them
independent of each other, the busy SMC handshake can not affect new TCP
connections visit any more. Avoid discarding a large number of TCP
connections after being overstock, which is undoubtedly raise the
connection establishment time.
Patch 2/5 (Limits SMC backlog connections):
Since patch 1 has separated smc_tcp_listen_work() from
smc_listen_work(), an unrestricted TCP accept have come into being. This
patch try to put a limit on SMC backlog connections refers to
implementation of TCP.
Patch 3/5 (Fallback when SMC handshake workqueue congested):
Considering the complexity of SMC handshake right now, in short-lived
links scenarios, this may not be the main scenario of SMC though, it's
performance is still quite poor. This Patch try to provide auto fallback
case when SMC handshake workqueue congested, which is the sign of SMC
handshake stacking in our opinion.
Patch 4/5 (Dynamic control SMC auto fallback by socket options)
This patch allow applications dynamically control the ability of SMC
auto fallback. Since SMC don't support set SMC socket option before,
this patch also have to support SMC's owns socket options.
Patch 5/5 (Add global configure for auto fallback by netlink)
This patch provides a way to get benefit of auto fallback without
modifying any code for applications, which is quite useful for most
existing applications.
After this patch set, performance figures like that:
Running 20s test @ http://11.213.45.6
4 threads and 10000 connections
693253 requests in 20.10s, 452.88MB read
Requests/sec: 34488.13
Transfer/sec: 22.53MB
That's a quite well performance improvement, about to 6 to 7 times in my
environment.
---
changelog:
v2 -> v1:
- fix compile warning
- fix invalid dependencies in kconfig
v3 -> v2:
- correct spelling mistakes
- fix useless variable declare
v4 -> v3
- make smc_tcp_ls_wq be static
v5 -> v4
- add dynamic control for SMC auto fallback by socket options
- add global configure for SMC auto fallback through netlink
v6 -> v5
- move auto fallback to net namespace scope
- remove auto fallback attribute in SMC_GEN_SYS_INFO
- add independent attributes for auto fallback
---
D. Wythe (5):
net/smc: Make smc_tcp_listen_work() independent
net/smc: Limit backlog connections
net/smc: Fallback when handshake workqueue congested
net/smc: Dynamic control auto fallback by socket options
net/smc: Add global configure for auto fallback by netlink
include/linux/socket.h | 1 +
include/linux/tcp.h | 1 +
include/net/netns/smc.h | 2 +
include/uapi/linux/smc.h | 15 ++++
net/ipv4/tcp_input.c | 3 +-
net/smc/af_smc.c | 182 ++++++++++++++++++++++++++++++++++++++++++++++-
net/smc/smc.h | 11 +++
net/smc/smc_netlink.c | 15 ++++
net/smc/smc_pnet.c | 3 +
9 files changed, 230 insertions(+), 3 deletions(-)
--
1.8.3.1
From: "D. Wythe" <alibuda@linux.alibaba.com>
In multithread and 10K connections benchmark, the backend TCP connection
established very slowly, and lots of TCP connections stay in SYN_SENT
state.
Client: smc_run wrk -c 10000 -t 4 http://server
the netstate of server host shows like:
145042 times the listen queue of a socket overflowed
145042 SYNs to LISTEN sockets dropped
One reason of this issue is that, since the smc_tcp_listen_work() shared
the same workqueue (smc_hs_wq) with smc_listen_work(), while the
smc_listen_work() do blocking wait for smc connection established. Once
the workqueue became congested, it's will block the accept() from TCP
listen.
This patch creates a independent workqueue(smc_tcp_ls_wq) for
smc_tcp_listen_work(), separate it from smc_listen_work(), which is
quite acceptable considering that smc_tcp_listen_work() runs very fast.
Signed-off-by: D. Wythe <alibuda@linux.alibaba.com>
---
net/smc/af_smc.c | 13 +++++++++++--
1 file changed, 11 insertions(+), 2 deletions(-)
@@ -59,6 +59,7 @@*creationonclient*/+staticstructworkqueue_struct*smc_tcp_ls_wq;/* wq for tcp listen work */structworkqueue_struct*smc_hs_wq;/* wq for handshake work */structworkqueue_struct*smc_close_wq;/* wq for close work */
From: "D. Wythe" <alibuda@linux.alibaba.com>
This patch intends to provide a mechanism to allow automatic fallback to
TCP according to the pressure of SMC handshake process. At present,
frequent visits will cause the incoming connections to be backlogged in
SMC handshake queue, raise the connections established time. Which is
quite unacceptable for those applications who base on short lived
connections.
There are two ways to implement this mechanism:
1. Fallback when TCP established.
2. Fallback before TCP established.
In the first way, we need to wait and receive CLC messages that the
client will potentially send, and then actively reply with a decline
message, in a sense, which is also a sort of SMC handshake, affect the
connections established time on its way.
In the second way, the only problem is that we need to inject SMC logic
into TCP when it is about to reply the incoming SYN, since we already do
that, it's seems not a problem anymore. And advantage is obvious, few
additional processes are required to complete the fallback.
This patch use the second way.
Link: https://lore.kernel.org/all/1641301961-59331-1-git-send-email-alibuda@linux.alibaba.com/
Signed-off-by: D. Wythe <alibuda@linux.alibaba.com>
---
include/linux/tcp.h | 1 +
net/ipv4/tcp_input.c | 3 ++-
net/smc/af_smc.c | 18 ++++++++++++++++++
3 files changed, 21 insertions(+), 1 deletion(-)
From: "D. Wythe" <alibuda@linux.alibaba.com>
Current implementation does not handling backlog semantics, one
potential risk is that server will be flooded by infinite amount
connections, even if client was SMC-incapable.
This patch works to put a limit on backlog connections, referring to the
TCP implementation, we divides SMC connections into two categories:
1. Half SMC connection, which includes all TCP established while SMC not
connections.
2. Full SMC connection, which includes all SMC established connections.
For half SMC connection, since all half SMC connections starts with TCP
established, we can achieve our goal by put a limit before TCP
established. Refer to the implementation of TCP, this limits will based
on not only the half SMC connections but also the full connections,
which is also a constraint on full SMC connections.
For full SMC connections, although we know exactly where it starts, it's
quite hard to put a limit before it. The easiest way is to block wait
before receive SMC confirm CLC message, while it's under protection by
smc_server_lgr_pending, a global lock, which leads this limit to the
entire host instead of a single listen socket. Another way is to drop
the full connections, but considering the cast of SMC connections, we
prefer to keep full SMC connections.
Even so, the limits of full SMC connections still exists, see commits
about half SMC connection below.
After this patch, the limits of backend connection shows like:
For SMC:
1. Client with SMC-capability can makes 2 * backlog full SMC connections
or 1 * backlog half SMC connections and 1 * backlog full SMC
connections at most.
2. Client without SMC-capability can only makes 1 * backlog half TCP
connections and 1 * backlog full TCP connections.
Signed-off-by: D. Wythe <alibuda@linux.alibaba.com>
---
net/smc/af_smc.c | 43 +++++++++++++++++++++++++++++++++++++++++++
net/smc/smc.h | 4 ++++
2 files changed, 47 insertions(+)
@@ -2266,6 +2300,15 @@ static int smc_listen(struct socket *sock, int backlog)smc->clcsock->sk->sk_data_ready=smc_clcsock_data_ready;smc->clcsock->sk->sk_user_data=(void*)((uintptr_t)smc|SK_USER_DATA_NOCOPY);++/* save origin ops */+smc->ori_af_ops=inet_csk(smc->clcsock->sk)->icsk_af_ops;++smc->af_ops=*smc->ori_af_ops;+smc->af_ops.syn_recv_sock=smc_tcp_syn_recv_sock;++inet_csk(smc->clcsock->sk)->icsk_af_ops=&smc->af_ops;+rc=kernel_listen(smc->clcsock,backlog);if(rc){smc->clcsock->sk->sk_data_ready=smc->clcsk_data_ready;
From: "D. Wythe" <alibuda@linux.alibaba.com>
Although we can control SMC auto fallback through socket options, which
means that applications who need it must modify their code. It's quite
troublesome for many existing applications. This patch modifies the
global default value of auto fallback through netlink, providing a way
to auto fallback without modifying any code for applications.
Suggested-by: Tony Lu <tonylu@linux.alibaba.com>
Signed-off-by: D. Wythe <alibuda@linux.alibaba.com>
---
include/net/netns/smc.h | 2 ++
include/uapi/linux/smc.h | 11 +++++++++++
net/smc/af_smc.c | 41 +++++++++++++++++++++++++++++++++++++++++
net/smc/smc.h | 6 ++++++
net/smc/smc_netlink.c | 15 +++++++++++++++
net/smc/smc_pnet.c | 3 +++
6 files changed, 78 insertions(+)
@@ -3006,6 +3044,9 @@ static int __smc_create(struct net *net, struct socket *sock, int protocol,smc->use_fallback=false;/* assume rdma capability first */smc->fallback_rsn=0;+/* default behavior from auto_fallback */+smc->auto_fallback=net->smc.auto_fallback;+rc=0;if(!clcsock){rc=sock_create_kern(net,family,SOCK_STREAM,IPPROTO_TCP,
@@ -111,6 +111,21 @@.flags=GENL_ADMIN_PERM,.doit=smc_nl_disable_seid,},+{+.cmd=SMC_NETLINK_DUMP_AUTO_FALLBACK,+/* can be retrieved by unprivileged users */+.dumpit=smc_nl_dump_auto_fallback,+},+{+.cmd=SMC_NETLINK_ENABLE_AUTO_FALLBACK,+.flags=GENL_ADMIN_PERM,+.doit=smc_nl_enable_auto_fallback,+},+{+.cmd=SMC_NETLINK_DISABLE_AUTO_FALLBACK,+.flags=GENL_ADMIN_PERM,+.doit=smc_nl_disable_auto_fallback,+},};staticconststructnla_policysmc_gen_nl_policy[2]={
@@ -868,6 +868,9 @@ int smc_pnet_net_init(struct net *net)smc_pnet_create_pnetids_list(net);+/* disable auto fallback by default */+net->smc.auto_fallback=0;+return0;}
From: "D. Wythe" <alibuda@linux.alibaba.com>
This patch aims to add dynamic control for SMC auto fallback, since we
don't have socket option level for SMC yet, which requires we need to
implement it at the same time.
This patch does the following:
- add new socket option level: SOL_SMC.
- add new SMC socket option: SMC_AUTO_FALLBACK.
- provide getter/setter for SMC socket options.
Signed-off-by: D. Wythe <alibuda@linux.alibaba.com>
---
include/linux/socket.h | 1 +
include/uapi/linux/smc.h | 4 +++
net/smc/af_smc.c | 69 +++++++++++++++++++++++++++++++++++++++++++++++-
net/smc/smc.h | 1 +
4 files changed, 74 insertions(+), 1 deletion(-)
@@ -2325,7 +2325,8 @@ static int smc_listen(struct socket *sock, int backlog)inet_csk(smc->clcsock->sk)->icsk_af_ops=&smc->af_ops;-tcp_sk(smc->clcsock->sk)->smc_in_limited=smc_is_in_limited;+if(smc->auto_fallback)+tcp_sk(smc->clcsock->sk)->smc_in_limited=smc_is_in_limited;rc=kernel_listen(smc->clcsock,backlog);if(rc){
@@ -2620,6 +2621,67 @@ static int smc_shutdown(struct socket *sock, int how)returnrc?rc:rc1;}+staticint__smc_getsockopt(structsocket*sock,intlevel,intoptname,+char__user*optval,int__user*optlen)+{+structsmc_sock*smc;+intval,len;++smc=smc_sk(sock->sk);++if(get_user(len,optlen))+return-EFAULT;++len=min_t(int,len,sizeof(int));++if(len<0)+return-EINVAL;++switch(optname){+caseSMC_AUTO_FALLBACK:+val=smc->auto_fallback;+break;+default:+return-EOPNOTSUPP;+}++if(put_user(len,optlen))+return-EFAULT;+if(copy_to_user(optval,&val,len))+return-EFAULT;++return0;+}++staticint__smc_setsockopt(structsocket*sock,intlevel,intoptname,+sockptr_toptval,unsignedintoptlen)+{+structsock*sk=sock->sk;+structsmc_sock*smc;+intval,rc;++smc=smc_sk(sk);++lock_sock(sk);+switch(optname){+caseSMC_AUTO_FALLBACK:+if(optlen<sizeof(int))+return-EINVAL;+if(copy_from_sockptr(&val,optval,sizeof(int)))+return-EFAULT;++smc->auto_fallback=!!val;+rc=0;+break;+default:+rc=-EOPNOTSUPP;+break;+}+release_sock(sk);++returnrc;+}+staticintsmc_setsockopt(structsocket*sock,intlevel,intoptname,sockptr_toptval,unsignedintoptlen){
@@ -2629,6 +2691,8 @@ static int smc_setsockopt(struct socket *sock, int level, int optname,if(level==SOL_TCP&&optname==TCP_ULP)return-EOPNOTSUPP;+elseif(level==SOL_SMC)+return__smc_setsockopt(sock,level,optname,optval,optlen);smc=smc_sk(sk);
@@ -2711,6 +2775,9 @@ static int smc_getsockopt(struct socket *sock, int level, int optname,structsmc_sock*smc;intrc;+if(level==SOL_SMC)+return__smc_getsockopt(sock,level,optname,optval,optlen);+smc=smc_sk(sock->sk);mutex_lock(&smc->clcsock_release_lock);if(!smc->clcsock){
From: "D. Wythe" <alibuda@linux.alibaba.com>
This patch intends to provide a mechanism to allow automatic fallback to
I would like to avoid the wording fallback all over here. The term SMC fallback
is used for SMC connections that are in our socket list, but use TCP because
something went wrong during handshake.
What you changes result in are TCP-only connections which are not handled by
the SMC module at all. So the comments should use a different naming for that.
What the patch actually does is to disable the SMC experimental TCP header option,
so the client receives no SMC indication and does not proceed with SMC.
Is this correct?
Please also see my comments below.
quoted hunk
TCP according to the pressure of SMC handshake process. At present,
frequent visits will cause the incoming connections to be backlogged in
SMC handshake queue, raise the connections established time. Which is
quite unacceptable for those applications who base on short lived
connections.
There are two ways to implement this mechanism:
1. Fallback when TCP established.
2. Fallback before TCP established.
In the first way, we need to wait and receive CLC messages that the
client will potentially send, and then actively reply with a decline
message, in a sense, which is also a sort of SMC handshake, affect the
connections established time on its way.
In the second way, the only problem is that we need to inject SMC logic
into TCP when it is about to reply the incoming SYN, since we already do
that, it's seems not a problem anymore. And advantage is obvious, few
additional processes are required to complete the fallback.
This patch use the second way.
Link: https://lore.kernel.org/all/1641301961-59331-1-git-send-email-alibuda@linux.alibaba.com/
Signed-off-by: D. Wythe <alibuda@linux.alibaba.com>
---
include/linux/tcp.h | 1 +
net/ipv4/tcp_input.c | 3 ++-
net/smc/af_smc.c | 18 ++++++++++++++++++
3 files changed, 21 insertions(+), 1 deletion(-)
From: "D. Wythe" <alibuda@linux.alibaba.com>
This patch aims to add dynamic control for SMC auto fallback, since we
Same here, we need a different wording. Maybe something like
"SMC handshake limitation".
quoted hunk
don't have socket option level for SMC yet, which requires we need to
implement it at the same time.
This patch does the following:
- add new socket option level: SOL_SMC.
- add new SMC socket option: SMC_AUTO_FALLBACK.
- provide getter/setter for SMC socket options.
Signed-off-by: D. Wythe <alibuda@linux.alibaba.com>
---
include/linux/socket.h | 1 +
include/uapi/linux/smc.h | 4 +++
net/smc/af_smc.c | 69 +++++++++++++++++++++++++++++++++++++++++++++++-
net/smc/smc.h | 1 +
4 files changed, 74 insertions(+), 1 deletion(-)
@@ -2325,7 +2325,8 @@ static int smc_listen(struct socket *sock, int backlog)inet_csk(smc->clcsock->sk)->icsk_af_ops=&smc->af_ops;-tcp_sk(smc->clcsock->sk)->smc_in_limited=smc_is_in_limited;+if(smc->auto_fallback)+tcp_sk(smc->clcsock->sk)->smc_in_limited=smc_is_in_limited;rc=kernel_listen(smc->clcsock,backlog);if(rc){
@@ -2620,6 +2621,67 @@ static int smc_shutdown(struct socket *sock, int how)returnrc?rc:rc1;}+staticint__smc_getsockopt(structsocket*sock,intlevel,intoptname,+char__user*optval,int__user*optlen)+{+structsmc_sock*smc;+intval,len;++smc=smc_sk(sock->sk);++if(get_user(len,optlen))+return-EFAULT;++len=min_t(int,len,sizeof(int));++if(len<0)+return-EINVAL;++switch(optname){+caseSMC_AUTO_FALLBACK:+val=smc->auto_fallback;+break;+default:+return-EOPNOTSUPP;+}++if(put_user(len,optlen))+return-EFAULT;+if(copy_to_user(optval,&val,len))+return-EFAULT;++return0;+}++staticint__smc_setsockopt(structsocket*sock,intlevel,intoptname,+sockptr_toptval,unsignedintoptlen)+{+structsock*sk=sock->sk;+structsmc_sock*smc;+intval,rc;++smc=smc_sk(sk);++lock_sock(sk);+switch(optname){+caseSMC_AUTO_FALLBACK:+if(optlen<sizeof(int))+return-EINVAL;+if(copy_from_sockptr(&val,optval,sizeof(int)))+return-EFAULT;++smc->auto_fallback=!!val;+rc=0;+break;+default:+rc=-EOPNOTSUPP;+break;+}+release_sock(sk);++returnrc;+}+staticintsmc_setsockopt(structsocket*sock,intlevel,intoptname,sockptr_toptval,unsignedintoptlen){
@@ -2629,6 +2691,8 @@ static int smc_setsockopt(struct socket *sock, int level, int optname,if(level==SOL_TCP&&optname==TCP_ULP)return-EOPNOTSUPP;+elseif(level==SOL_SMC)+return__smc_setsockopt(sock,level,optname,optval,optlen);smc=smc_sk(sk);
@@ -2711,6 +2775,9 @@ static int smc_getsockopt(struct socket *sock, int level, int optname,structsmc_sock*smc;intrc;+if(level==SOL_SMC)+return__smc_getsockopt(sock,level,optname,optval,optlen);+smc=smc_sk(sock->sk);mutex_lock(&smc->clcsock_release_lock);if(!smc->clcsock){
From: "D. Wythe" <alibuda@linux.alibaba.com>
Although we can control SMC auto fallback through socket options, which
means that applications who need it must modify their code. It's quite
troublesome for many existing applications. This patch modifies the
global default value of auto fallback through netlink, providing a way
to auto fallback without modifying any code for applications.
And of course also in this patch: no "auto fallback" in comments or as part
of variable names.
Do you plan to enhance the smc-tools user space part, too?
From: "D. Wythe" <alibuda@linux.alibaba.com>
This patch intends to provide a mechanism to allow automatic fallback to
I would like to avoid the wording fallback all over here. The term SMC fallback
is used for SMC connections that are in our socket list, but use TCP because
something went wrong during handshake.
What you changes result in are TCP-only connections which are not handled by
the SMC module at all. So the comments should use a different naming for that.
What the patch actually does is to disable the SMC experimental TCP header option,
so the client receives no SMC indication and does not proceed with SMC.
Is this correct?
Please also see my comments below.
I agree with you, the wording fallback doesn't fit here. I'll try
limitation.
quoted
TCP according to the pressure of SMC handshake process. At present,
frequent visits will cause the incoming connections to be backlogged in
SMC handshake queue, raise the connections established time. Which is
quite unacceptable for those applications who base on short lived
connections.
There are two ways to implement this mechanism:
1. Fallback when TCP established.
2. Fallback before TCP established.
In the first way, we need to wait and receive CLC messages that the
client will potentially send, and then actively reply with a decline
message, in a sense, which is also a sort of SMC handshake, affect the
connections established time on its way.
In the second way, the only problem is that we need to inject SMC logic
into TCP when it is about to reply the incoming SYN, since we already do
that, it's seems not a problem anymore. And advantage is obvious, few
additional processes are required to complete the fallback.
This patch use the second way.
Link: https://lore.kernel.org/all/1641301961-59331-1-git-send-email-alibuda@linux.alibaba.com/
Signed-off-by: D. Wythe <alibuda@linux.alibaba.com>
---
include/linux/tcp.h | 1 +
net/ipv4/tcp_input.c | 3 ++-
net/smc/af_smc.c | 18 ++++++++++++++++++
3 files changed, 21 insertions(+), 1 deletion(-)
From: "D. Wythe" <alibuda@linux.alibaba.com>
Although we can control SMC auto fallback through socket options, which
means that applications who need it must modify their code. It's quite
troublesome for many existing applications. This patch modifies the
global default value of auto fallback through netlink, providing a way
to auto fallback without modifying any code for applications.
And of course also in this patch: no "auto fallback" in comments or as part
of variable names.
I will fix all the wording and the naming issues in next series as soon
as possible.
Do you plan to enhance the smc-tools user space part, too?