Since commit dc7b3eb900aa ("ipvs: Fix reuse connection if real server is
dead"), new connections to dead servers are redistributed immediately to
new servers.
Then commit d752c3645717 ("ipvs: allow rescheduling of new connections when
port reuse is detected") disable expire_nodest_conn if conn_reuse_mode is
0. And new connection may be distributed to a real server with weight 0.
Signed-off-by: yangxingwu <redacted>
---
Documentation/networking/ipvs-sysctl.rst | 3 +--
net/netfilter/ipvs/ip_vs_core.c | 5 +++--
2 files changed, 4 insertions(+), 4 deletions(-)
@@ -37,8 +37,7 @@ conn_reuse_mode - INTEGER 0: disable any special handling on port reuse. The new connection will be delivered to the same real server that was- servicing the previous connection. This will effectively- disable expire_nodest_conn.+ servicing the previous connection. bit 1: enable rescheduling of new connections when it is safe. That is, whenever expire_nodest_conn and for TCP sockets, when
Since commit dc7b3eb900aa ("ipvs: Fix reuse connection if real server is
dead"), new connections to dead servers are redistributed immediately to
new servers.
Then commit d752c3645717 ("ipvs: allow rescheduling of new connections when
port reuse is detected") disable expire_nodest_conn if conn_reuse_mode is
0. And new connection may be distributed to a real server with weight 0.
Your change does not look correct to me. At the time
expire_nodest_conn was created, it was not checked when
weight is 0. At different places different terms are used
but in short, we have two independent states for real server:
- inhibited: weight=0 and no new connections should be served,
packets for existing connections can be routed to server
if it is still available and packets are not dropped
by expire_nodest_conn.
The new feature is that port reuse detection can
redirect the new TCP connection into a new IPVS conn and
to expire the existing cp/ct.
- unavailable (!IP_VS_DEST_F_AVAILABLE): server is removed,
can be temporary, drop traffic for existing connections
but on expire_nodest_conn we can select different server
The new conn_reuse_mode flag allows port reuse to
be detected. Only then expire_nodest_conn has the
opportunity with commit dc7b3eb900aa to check weight=0
and to consider the old traffic as finished. If a new
server is selected, any retrans from previous connection
would be considered as part from the new connection. It
is a rapid way to switch server without checking with
is_new_conn_expected() because we can not have many
conns/conntracks to different servers.
@@ -37,8 +37,7 @@ conn_reuse_mode - INTEGER 0: disable any special handling on port reuse. The new connection will be delivered to the same real server that was- servicing the previous connection. This will effectively- disable expire_nodest_conn.+ servicing the previous connection. bit 1: enable rescheduling of new connections when it is safe. That is, whenever expire_nodest_conn and for TCP sockets, when
thanks julian
What happens in this situation is that if we set the wait of the
realserver to 0 and do NOT remove the weight zero realserver with
sysctl settings (conn_reuse_mode == 0 && expire_nodest_conn == 1), and
the client reuses its source ports, the kernel will constantly
reuse connections and send the traffic to the weight 0 realserver.
you may check the details from
https://github.com/kubernetes/kubernetes/issues/81775
On Tue, Oct 26, 2021 at 2:12 AM Julian Anastasov [off-list ref] wrote:
Hello,
On Mon, 25 Oct 2021, yangxingwu wrote:
quoted
Since commit dc7b3eb900aa ("ipvs: Fix reuse connection if real server is
dead"), new connections to dead servers are redistributed immediately to
new servers.
Then commit d752c3645717 ("ipvs: allow rescheduling of new connections when
port reuse is detected") disable expire_nodest_conn if conn_reuse_mode is
0. And new connection may be distributed to a real server with weight 0.
Your change does not look correct to me. At the time
expire_nodest_conn was created, it was not checked when
weight is 0. At different places different terms are used
but in short, we have two independent states for real server:
- inhibited: weight=0 and no new connections should be served,
packets for existing connections can be routed to server
if it is still available and packets are not dropped
by expire_nodest_conn.
The new feature is that port reuse detection can
redirect the new TCP connection into a new IPVS conn and
to expire the existing cp/ct.
- unavailable (!IP_VS_DEST_F_AVAILABLE): server is removed,
can be temporary, drop traffic for existing connections
but on expire_nodest_conn we can select different server
The new conn_reuse_mode flag allows port reuse to
be detected. Only then expire_nodest_conn has the
opportunity with commit dc7b3eb900aa to check weight=0
and to consider the old traffic as finished. If a new
server is selected, any retrans from previous connection
would be considered as part from the new connection. It
is a rapid way to switch server without checking with
is_new_conn_expected() because we can not have many
conns/conntracks to different servers.
@@ -37,8 +37,7 @@ conn_reuse_mode - INTEGER 0: disable any special handling on port reuse. The new connection will be delivered to the same real server that was- servicing the previous connection. This will effectively- disable expire_nodest_conn.+ servicing the previous connection. bit 1: enable rescheduling of new connections when it is safe. That is, whenever expire_nodest_conn and for TCP sockets, when
thanks julian
What happens in this situation is that if we set the wait of the
realserver to 0 and do NOT remove the weight zero realserver with
sysctl settings (conn_reuse_mode == 0 && expire_nodest_conn == 1), and
the client reuses its source ports, the kernel will constantly
reuse connections and send the traffic to the weight 0 realserver.
What happens if you try conn_reuse_mode=1? The
one-second delay in previous kernels should be corrected with
commit f0a5e4d7a594e0fe237d3dfafb069bb82f80f42f
Date: Wed Jul 1 18:17:19 2020 +0300
ipvs: allow connection reuse for unconfirmed conntrack
On Tue, Oct 26, 2021 at 2:12 AM Julian Anastasov [off-list ref] wrote:
quoted
On Mon, 25 Oct 2021, yangxingwu wrote:
quoted
Since commit dc7b3eb900aa ("ipvs: Fix reuse connection if real server is
dead"), new connections to dead servers are redistributed immediately to
new servers.
Then commit d752c3645717 ("ipvs: allow rescheduling of new connections when
port reuse is detected") disable expire_nodest_conn if conn_reuse_mode is
0. And new connection may be distributed to a real server with weight 0.
Your change does not look correct to me. At the time
expire_nodest_conn was created, it was not checked when
weight is 0. At different places different terms are used
but in short, we have two independent states for real server:
- inhibited: weight=0 and no new connections should be served,
packets for existing connections can be routed to server
if it is still available and packets are not dropped
by expire_nodest_conn.
The new feature is that port reuse detection can
redirect the new TCP connection into a new IPVS conn and
to expire the existing cp/ct.
- unavailable (!IP_VS_DEST_F_AVAILABLE): server is removed,
can be temporary, drop traffic for existing connections
but on expire_nodest_conn we can select different server
The new conn_reuse_mode flag allows port reuse to
be detected. Only then expire_nodest_conn has the
opportunity with commit dc7b3eb900aa to check weight=0
and to consider the old traffic as finished. If a new
server is selected, any retrans from previous connection
would be considered as part from the new connection. It
is a rapid way to switch server without checking with
is_new_conn_expected() because we can not have many
conns/conntracks to different servers.
thanks Julian
yes, I know that the one-second delay issue has been fixed by commit
f0a5e4d7a594e0fe237d3dfafb069bb82f80f42f if we set conn_reuse_mode to
1
BUT it's still NOT what we expected with sysctl settings
(conn_reuse_mode == 0 && expire_nodest_conn == 1).
We run kubernetes in extremely diverse environments and this issue
happens a lot.
On Tue, Oct 26, 2021 at 1:44 PM Julian Anastasov [off-list ref] wrote:
Hello,
On Tue, 26 Oct 2021, yangxingwu wrote:
quoted
thanks julian
What happens in this situation is that if we set the wait of the
realserver to 0 and do NOT remove the weight zero realserver with
sysctl settings (conn_reuse_mode == 0 && expire_nodest_conn == 1), and
the client reuses its source ports, the kernel will constantly
reuse connections and send the traffic to the weight 0 realserver.
What happens if you try conn_reuse_mode=1? The
one-second delay in previous kernels should be corrected with
commit f0a5e4d7a594e0fe237d3dfafb069bb82f80f42f
Date: Wed Jul 1 18:17:19 2020 +0300
ipvs: allow connection reuse for unconfirmed conntrack
quoted
On Tue, Oct 26, 2021 at 2:12 AM Julian Anastasov [off-list ref] wrote:
quoted
On Mon, 25 Oct 2021, yangxingwu wrote:
quoted
Since commit dc7b3eb900aa ("ipvs: Fix reuse connection if real server is
dead"), new connections to dead servers are redistributed immediately to
new servers.
Then commit d752c3645717 ("ipvs: allow rescheduling of new connections when
port reuse is detected") disable expire_nodest_conn if conn_reuse_mode is
0. And new connection may be distributed to a real server with weight 0.
Your change does not look correct to me. At the time
expire_nodest_conn was created, it was not checked when
weight is 0. At different places different terms are used
but in short, we have two independent states for real server:
- inhibited: weight=0 and no new connections should be served,
packets for existing connections can be routed to server
if it is still available and packets are not dropped
by expire_nodest_conn.
The new feature is that port reuse detection can
redirect the new TCP connection into a new IPVS conn and
to expire the existing cp/ct.
- unavailable (!IP_VS_DEST_F_AVAILABLE): server is removed,
can be temporary, drop traffic for existing connections
but on expire_nodest_conn we can select different server
The new conn_reuse_mode flag allows port reuse to
be detected. Only then expire_nodest_conn has the
opportunity with commit dc7b3eb900aa to check weight=0
and to consider the old traffic as finished. If a new
server is selected, any retrans from previous connection
would be considered as part from the new connection. It
is a rapid way to switch server without checking with
is_new_conn_expected() because we can not have many
conns/conntracks to different servers.
Julian
what we want is if RS weight is 0, then no new connections should be
served even if conn_reuse_mode is 0, just as commit dc7b3eb900aa
("ipvs: Fix reuse connection if real server is
dead") trying to do
Pls let me know if there are any other issues of concern
On Tue, Oct 26, 2021 at 2:13 PM yangxingwu [off-list ref] wrote:
thanks Julian
yes, I know that the one-second delay issue has been fixed by commit
f0a5e4d7a594e0fe237d3dfafb069bb82f80f42f if we set conn_reuse_mode to
1
BUT it's still NOT what we expected with sysctl settings
(conn_reuse_mode == 0 && expire_nodest_conn == 1).
We run kubernetes in extremely diverse environments and this issue
happens a lot.
On Tue, Oct 26, 2021 at 1:44 PM Julian Anastasov [off-list ref] wrote:
quoted
Hello,
On Tue, 26 Oct 2021, yangxingwu wrote:
quoted
thanks julian
What happens in this situation is that if we set the wait of the
realserver to 0 and do NOT remove the weight zero realserver with
sysctl settings (conn_reuse_mode == 0 && expire_nodest_conn == 1), and
the client reuses its source ports, the kernel will constantly
reuse connections and send the traffic to the weight 0 realserver.
What happens if you try conn_reuse_mode=1? The
one-second delay in previous kernels should be corrected with
commit f0a5e4d7a594e0fe237d3dfafb069bb82f80f42f
Date: Wed Jul 1 18:17:19 2020 +0300
ipvs: allow connection reuse for unconfirmed conntrack
quoted
On Tue, Oct 26, 2021 at 2:12 AM Julian Anastasov [off-list ref] wrote:
quoted
On Mon, 25 Oct 2021, yangxingwu wrote:
quoted
Since commit dc7b3eb900aa ("ipvs: Fix reuse connection if real server is
dead"), new connections to dead servers are redistributed immediately to
new servers.
Then commit d752c3645717 ("ipvs: allow rescheduling of new connections when
port reuse is detected") disable expire_nodest_conn if conn_reuse_mode is
0. And new connection may be distributed to a real server with weight 0.
Your change does not look correct to me. At the time
expire_nodest_conn was created, it was not checked when
weight is 0. At different places different terms are used
but in short, we have two independent states for real server:
- inhibited: weight=0 and no new connections should be served,
packets for existing connections can be routed to server
if it is still available and packets are not dropped
by expire_nodest_conn.
The new feature is that port reuse detection can
redirect the new TCP connection into a new IPVS conn and
to expire the existing cp/ct.
- unavailable (!IP_VS_DEST_F_AVAILABLE): server is removed,
can be temporary, drop traffic for existing connections
but on expire_nodest_conn we can select different server
The new conn_reuse_mode flag allows port reuse to
be detected. Only then expire_nodest_conn has the
opportunity with commit dc7b3eb900aa to check weight=0
and to consider the old traffic as finished. If a new
server is selected, any retrans from previous connection
would be considered as part from the new connection. It
is a rapid way to switch server without checking with
is_new_conn_expected() because we can not have many
conns/conntracks to different servers.
what we want is if RS weight is 0, then no new connections should be
served even if conn_reuse_mode is 0, just as commit dc7b3eb900aa
("ipvs: Fix reuse connection if real server is
dead") trying to do
Pls let me know if there are any other issues of concern
My concern is with the behaviour people expect
from each sysctl var: conn_reuse_mode decides if port reuse
is considered for rescheduling and expire_nodest_conn
should have priority only for unavailable servers (nodest means
No Destination), not in this case.
We don't know how people use the conn_reuse_mode=0
mode, one may bind to a local port and try to send multiple
connections in a row with the hope they will go to same real
server, i.e. as part from same "session", even while weight=0.
If they do not want such behaviour (99% of the cases), they
will use the default conn_reuse_mode=1. OTOH, you have different
expectations for mode 0, not sure why but you do not want to use
the default mode=1 which is safer to use. May be the setups
forget to stay with conn_reuse_mode=1 on kernels 5.9+ and
set the var to 0 ?
The problem with mentioned commit dc7b3eb900aa is that
it breaks FTP and persistent connections while the goal of
weight=0 is graceful inhibition of the server. We made
the mistake to add priority for expire_nodest_conn when weight=0.
This can be fixed with a !cp->control check. We do not want
expire_nodest_conn to kill every connection during the
graceful period.
Regards
--
Julian Anastasov [off-list ref]
hello
On Thu, Oct 28, 2021 at 5:09 AM Julian Anastasov [off-list ref] wrote:
Hello,
On Wed, 27 Oct 2021, yangxingwu wrote:
quoted
what we want is if RS weight is 0, then no new connections should be
served even if conn_reuse_mode is 0, just as commit dc7b3eb900aa
("ipvs: Fix reuse connection if real server is
dead") trying to do
Pls let me know if there are any other issues of concern
My concern is with the behaviour people expect
from each sysctl var: conn_reuse_mode decides if port reuse
is considered for rescheduling and expire_nodest_conn
should have priority only for unavailable servers (nodest means
No Destination), not in this case.
We don't know how people use the conn_reuse_mode=0
mode, one may bind to a local port and try to send multiple
connections in a row with the hope they will go to same real
server, i.e. as part from same "session", even while weight=0.
If they do not want such behaviour (99% of the cases), they
will use the default conn_reuse_mode=1. OTOH, you have different
expectations for mode 0, not sure why but you do not want to use
the default mode=1 which is safer to use. May be the setups
forget to stay with conn_reuse_mode=1 on kernels 5.9+ and
set the var to 0 ?
The problem is we can NOT decide what the customers do, many of them
run kubernetes with old versions of kube-proxy. And most importantly,
upgrade to new version is a very long and painful process, that's why
we want to fix this at the kernel level
The problem with mentioned commit dc7b3eb900aa is that
it breaks FTP and persistent connections while the goal of
weight=0 is graceful inhibition of the server. We made
the mistake to add priority for expire_nodest_conn when weight=0.
This can be fixed with a !cp->control check. We do not want
expire_nodest_conn to kill every connection during the
graceful period.