[PATCH v2 net-next] tcp: fix ABC in tcp_slow_start()

Subsystems: networking [general], networking [tcp], the rest

STALE5137d

7 messages, 4 authors, 2012-07-20 · open the first message on its own page

[PATCH v2 net-next] tcp: fix ABC in tcp_slow_start()

From: Eric Dumazet <hidden>
Date: 2012-07-20 15:02:39

From: Eric Dumazet <edumazet@google.com>

When/if sysctl_tcp_abc > 1, we expect to increase cwnd by 2 if the
received ACK acknowledges more than 2*MSS bytes, in tcp_slow_start()

Problem is this RFC 3465 statement is not correctly coded, as
the while () loop increases snd_cwnd one by one.

Add a new variable to avoid this off-by one error.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: Tom Herbert <redacted>
Cc: Yuchung Cheng <redacted>
Cc: Neal Cardwell <ncardwell@google.com>
Cc: Nandita Dukkipati <redacted>
Cc: John Heffner <redacted>
Cc: Stephen Hemminger <redacted>
---
v2: added John suggestion

 net/ipv4/tcp_cong.c |    5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/net/ipv4/tcp_cong.c b/net/ipv4/tcp_cong.c
index 04dbd7a..4d4db16 100644
--- a/net/ipv4/tcp_cong.c
+++ b/net/ipv4/tcp_cong.c
@@ -307,6 +307,7 @@ EXPORT_SYMBOL_GPL(tcp_is_cwnd_limited);
 void tcp_slow_start(struct tcp_sock *tp)
 {
 	int cnt; /* increase in packets */
+	unsigned int delta = 0;
 
 	/* RFC3465: ABC Slow start
 	 * Increase only after a full MSS of bytes is acked
@@ -333,9 +334,9 @@ void tcp_slow_start(struct tcp_sock *tp)
 	tp->snd_cwnd_cnt += cnt;
 	while (tp->snd_cwnd_cnt >= tp->snd_cwnd) {
 		tp->snd_cwnd_cnt -= tp->snd_cwnd;
-		if (tp->snd_cwnd < tp->snd_cwnd_clamp)
-			tp->snd_cwnd++;
+		delta++;
 	}
+	tp->snd_cwnd = min(tp->snd_cwnd + delta, tp->snd_cwnd_clamp);
 }
 EXPORT_SYMBOL_GPL(tcp_slow_start);
 

Re: [PATCH v2 net-next] tcp: fix ABC in tcp_slow_start()

From: Yuchung Cheng <hidden>
Date: 2012-07-20 15:08:00

On Fri, Jul 20, 2012 at 8:02 AM, Eric Dumazet [off-list ref] wrote:
From: Eric Dumazet <edumazet@google.com>

When/if sysctl_tcp_abc > 1, we expect to increase cwnd by 2 if the
received ACK acknowledges more than 2*MSS bytes, in tcp_slow_start()

Problem is this RFC 3465 statement is not correctly coded, as
the while () loop increases snd_cwnd one by one.

Add a new variable to avoid this off-by one error.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Acked-by: Yuchung Cheng <redacted>
quoted hunk
Cc: Tom Herbert <redacted>
Cc: Yuchung Cheng <redacted>
Cc: Neal Cardwell <ncardwell@google.com>
Cc: Nandita Dukkipati <redacted>
Cc: John Heffner <redacted>
Cc: Stephen Hemminger <redacted>
---
v2: added John suggestion

 net/ipv4/tcp_cong.c |    5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/net/ipv4/tcp_cong.c b/net/ipv4/tcp_cong.c
index 04dbd7a..4d4db16 100644
--- a/net/ipv4/tcp_cong.c
+++ b/net/ipv4/tcp_cong.c
@@ -307,6 +307,7 @@ EXPORT_SYMBOL_GPL(tcp_is_cwnd_limited);
 void tcp_slow_start(struct tcp_sock *tp)
 {
        int cnt; /* increase in packets */
+       unsigned int delta = 0;

        /* RFC3465: ABC Slow start
         * Increase only after a full MSS of bytes is acked
@@ -333,9 +334,9 @@ void tcp_slow_start(struct tcp_sock *tp)
        tp->snd_cwnd_cnt += cnt;
        while (tp->snd_cwnd_cnt >= tp->snd_cwnd) {
                tp->snd_cwnd_cnt -= tp->snd_cwnd;
-               if (tp->snd_cwnd < tp->snd_cwnd_clamp)
-                       tp->snd_cwnd++;
+               delta++;
Nice! this also removes wasteful iteration when clamp << cwnd_cnt.
        }
+       tp->snd_cwnd = min(tp->snd_cwnd + delta, tp->snd_cwnd_clamp);
 }
 EXPORT_SYMBOL_GPL(tcp_slow_start);

Re: [PATCH v2 net-next] tcp: fix ABC in tcp_slow_start()

From: Neal Cardwell <ncardwell@google.com>
Date: 2012-07-20 16:03:46

On Fri, Jul 20, 2012 at 8:07 AM, Yuchung Cheng [off-list ref] wrote:
On Fri, Jul 20, 2012 at 8:02 AM, Eric Dumazet [off-list ref] wrote:
quoted
        tp->snd_cwnd_cnt += cnt;
        while (tp->snd_cwnd_cnt >= tp->snd_cwnd) {
Nice catch, Eric.

One thing that's always bothered me about the tp->snd_cwnd_cnt code is
that the slow start and congestion avoidance use different criteria
for incrementing snd_cwnd_cnt. tcp_slow_start() increments
snd_cwnd_cnt by snd_cwnd for each ACKed packet, and congestion
avoidance increases snd_cwnd_cnt by just 1 for each packet.

This means that if we exit slow start and enter congestion avoidance,
then we think we can have a "credit" for a bunch of ACKs that never
happened (up to snd_cwnd-1), so we can conceivably do our first
additive increase in congestion avoidance up to almost 1RTT too
early. Can we just get rid of the use of snd_cwnd_cnt in slow start,
and just use local variables in tcp_slow_start() rather than trying to
carry state between ACKs?

neal

Re: [PATCH v2 net-next] tcp: fix ABC in tcp_slow_start()

From: Eric Dumazet <hidden>
Date: 2012-07-20 16:09:01

On Fri, 2012-07-20 at 09:03 -0700, Neal Cardwell wrote:
On Fri, Jul 20, 2012 at 8:07 AM, Yuchung Cheng [off-list ref] wrote:
quoted
On Fri, Jul 20, 2012 at 8:02 AM, Eric Dumazet [off-list ref] wrote:
quoted
        tp->snd_cwnd_cnt += cnt;
        while (tp->snd_cwnd_cnt >= tp->snd_cwnd) {
Nice catch, Eric.

One thing that's always bothered me about the tp->snd_cwnd_cnt code is
that the slow start and congestion avoidance use different criteria
for incrementing snd_cwnd_cnt. tcp_slow_start() increments
snd_cwnd_cnt by snd_cwnd for each ACKed packet, and congestion
avoidance increases snd_cwnd_cnt by just 1 for each packet.

This means that if we exit slow start and enter congestion avoidance,
then we think we can have a "credit" for a bunch of ACKs that never
happened (up to snd_cwnd-1), so we can conceivably do our first
additive increase in congestion avoidance up to almost 1RTT too
early. Can we just get rid of the use of snd_cwnd_cnt in slow start,
and just use local variables in tcp_slow_start() rather than trying to
carry state between ACKs?
Apparently tcp_slow_start() needs the snd_cwnd_cnt in case 
"limited slow start"  is used :

cnt = sysctl_tcp_max_ssthresh >> 1;

So to address your point, maybe we should clear  snd_cwnd_cnt
when leaving slow start for congestion avoidance phase ?

Re: [PATCH v2 net-next] tcp: fix ABC in tcp_slow_start()

From: Neal Cardwell <ncardwell@google.com>
Date: 2012-07-20 16:50:37

On Fri, Jul 20, 2012 at 9:08 AM, Eric Dumazet [off-list ref] wrote:
So to address your point, maybe we should clear  snd_cwnd_cnt
when leaving slow start for congestion avoidance phase ?
Sounds good. That can be a separate commit to add the new logic to the
end of tcp_slow_start() to check to see if we've bumped into ssthresh
and reset snd_cwnd_cnt.

neal

Re: [PATCH v2 net-next] tcp: fix ABC in tcp_slow_start()

From: Neal Cardwell <ncardwell@google.com>
Date: 2012-07-20 17:58:28

On Fri, Jul 20, 2012 at 8:07 AM, Yuchung Cheng [off-list ref] wrote:
On Fri, Jul 20, 2012 at 8:02 AM, Eric Dumazet [off-list ref] wrote:
quoted
From: Eric Dumazet <edumazet@google.com>

When/if sysctl_tcp_abc > 1, we expect to increase cwnd by 2 if the
received ACK acknowledges more than 2*MSS bytes, in tcp_slow_start()

Problem is this RFC 3465 statement is not correctly coded, as
the while () loop increases snd_cwnd one by one.

Add a new variable to avoid this off-by one error.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Acked-by: Yuchung Cheng <redacted>
Acked-by: Neal Cardwell <ncardwell@google.com>

neal

Re: [PATCH v2 net-next] tcp: fix ABC in tcp_slow_start()

From: David Miller <davem@davemloft.net>
Date: 2012-07-20 18:01:53

From: Neal Cardwell <ncardwell@google.com>
Date: Fri, 20 Jul 2012 10:58:27 -0700
On Fri, Jul 20, 2012 at 8:07 AM, Yuchung Cheng [off-list ref] wrote:
quoted
On Fri, Jul 20, 2012 at 8:02 AM, Eric Dumazet [off-list ref] wrote:
quoted
From: Eric Dumazet <edumazet@google.com>

When/if sysctl_tcp_abc > 1, we expect to increase cwnd by 2 if the
received ACK acknowledges more than 2*MSS bytes, in tcp_slow_start()

Problem is this RFC 3465 statement is not correctly coded, as
the while () loop increases snd_cwnd one by one.

Add a new variable to avoid this off-by one error.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Acked-by: Yuchung Cheng <redacted>
Acked-by: Neal Cardwell <ncardwell@google.com>
Applied.
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help