From: Stephen Rothwell <hidden> Date: 2016-11-25 01:09:50
Hi all,
This is a typical user error report i.e. a net well specified one :-)
I am using a 6in4 tunnel from my Linux server at home (since my ISP
does not provide native IPv6) to another hosted Linus server (that has
native IPv6 connectivity). The throughput for IPv6 connections has
dropped from megabits per second to 10s of kilobits per second.
First, I am using Debian supplied kernels, so strike one, right?
Second, I don't actually remember when the problem started - it probably
started when I upgraded from a v4.4 based kernel to a v4.7 based one.
This server does not get rebooted very often as it runs hosted services
for quite a few people (its is ozlabs.org ...).
I tried creating the same tunnel to another hosted server I have access
to that is running a v3.16 based kernel and the performance is fine
(actually upward of 40MB/s).
I noticed from a tcpdump on the hosted server that (when I fetch a
large file over HTTP) the server is sending packets larger than the MTU
of the tunnel. These packets don't get acked and are later resent as
MTU sized packets. I will then send more larger packets and repeat ...
The mtu of the tunnel is set to 1280 (though leaving it unset and using
the default gave the same results). The tunnel is using sit and is
statically set up at both ends (though the hosted server end does not
specify a remote ipv4 end point).
Is there anything else I can tell you? Testing patches is a bit of a
pain, unfortunately, but I was hoping that someone may remember
something that may have caused this.
--
Cheers,
Stephen Rothwell
From: Eli Cooper <hidden> Date: 2016-11-25 02:18:23
Hi Stephen,
On 2016/11/25 9:09, Stephen Rothwell wrote:
Hi all,
This is a typical user error report i.e. a net well specified one :-)
I am using a 6in4 tunnel from my Linux server at home (since my ISP
does not provide native IPv6) to another hosted Linus server (that has
native IPv6 connectivity). The throughput for IPv6 connections has
dropped from megabits per second to 10s of kilobits per second.
First, I am using Debian supplied kernels, so strike one, right?
Second, I don't actually remember when the problem started - it probably
started when I upgraded from a v4.4 based kernel to a v4.7 based one.
This server does not get rebooted very often as it runs hosted services
for quite a few people (its is ozlabs.org ...).
I tried creating the same tunnel to another hosted server I have access
to that is running a v3.16 based kernel and the performance is fine
(actually upward of 40MB/s).
I noticed from a tcpdump on the hosted server that (when I fetch a
large file over HTTP) the server is sending packets larger than the MTU
of the tunnel. These packets don't get acked and are later resent as
MTU sized packets. I will then send more larger packets and repeat ...
Sounds like TSO/GSO packets are not properly segmented and therefore
dropped.
Could you first try turning off segmentation offloading for the tunnel
interface?
ethtool -K sit0 tso off gso off
The mtu of the tunnel is set to 1280 (though leaving it unset and using
the default gave the same results). The tunnel is using sit and is
statically set up at both ends (though the hosted server end does not
specify a remote ipv4 end point).
Is there anything else I can tell you? Testing patches is a bit of a
pain, unfortunately, but I was hoping that someone may remember
something that may have caused this.
From: Eric Dumazet <hidden> Date: 2016-11-25 02:30:17
On Fri, 2016-11-25 at 12:09 +1100, Stephen Rothwell wrote:
Hi all,
This is a typical user error report i.e. a net well specified one :-)
I am using a 6in4 tunnel from my Linux server at home (since my ISP
does not provide native IPv6) to another hosted Linus server (that has
native IPv6 connectivity). The throughput for IPv6 connections has
dropped from megabits per second to 10s of kilobits per second.
First, I am using Debian supplied kernels, so strike one, right?
Second, I don't actually remember when the problem started - it probably
started when I upgraded from a v4.4 based kernel to a v4.7 based one.
This server does not get rebooted very often as it runs hosted services
for quite a few people (its is ozlabs.org ...).
I tried creating the same tunnel to another hosted server I have access
to that is running a v3.16 based kernel and the performance is fine
(actually upward of 40MB/s).
I noticed from a tcpdump on the hosted server that (when I fetch a
large file over HTTP) the server is sending packets larger than the MTU
of the tunnel. These packets don't get acked and are later resent as
MTU sized packets. I will then send more larger packets and repeat ...
tcpdump shows big packets because SIT supports TSO (since linux-3.13)
lpaa23:~# ip -6 ro get 2002:af6:798::1
2002:af6:798::1 via fe80:: dev sixtofour0 src 2002:af6:797:: metric
1024 pref medium
lpaa23:~# ./netperf -H 2002:af6:798::1
MIGRATED TCP STREAM TEST from ::0 (::) port 0 AF_INET6 to
2002:af6:798::1 () port 0 AF_INET6
Recv Send Send
Socket Socket Message Elapsed
Size Size Size Time Throughput
bytes bytes bytes secs. 10^6bits/sec
87380 16384 16384 10.00 10374.64
lpaa23:~# ethtool -k sixtofour0|grep seg
tcp-segmentation-offload: on
tx-tcp-segmentation: on
tx-tcp-ecn-segmentation: on
tx-tcp6-segmentation: on
tx-tcp-psp-segmentation: off [fixed]
generic-segmentation-offload: on
tx-fcoe-segmentation: off [fixed]
tx-gre-segmentation: off [fixed]
tx-ipip-segmentation: off [fixed]
tx-sit-segmentation: off [fixed]
tx-udp_tnl-segmentation: off [fixed]
tx-mpls-segmentation: off [fixed]
tx-ggre-segmentation: off [fixed]
The mtu of the tunnel is set to 1280 (though leaving it unset and using
the default gave the same results). The tunnel is using sit and is
statically set up at both ends (though the hosted server end does not
specify a remote ipv4 end point).
Is there anything else I can tell you? Testing patches is a bit of a
pain, unfortunately, but I was hoping that someone may remember
something that may have caused this.
You could use "perf record -a -g -e skb:kfree_skb" to see where packets
are dropped on the sender.
You also could try to disable TSO and see if this makes a difference
ethtool -K sixtofour0 tso off
From: Stephen Rothwell <hidden> Date: 2016-11-25 02:45:22
Hi Eli,
On Fri, 25 Nov 2016 10:18:12 +0800 Eli Cooper [off-list ref] wrote:
Sounds like TSO/GSO packets are not properly segmented and therefore
dropped.
Could you first try turning off segmentation offloading for the tunnel
interface?
ethtool -K sit0 tso off gso off
On Thu, 24 Nov 2016 18:30:14 -0800 Eric Dumazet [off-list ref]
You also could try to disable TSO and see if this makes a difference
ethtool -K sixtofour0 tso off
So turning off tso brings performance up to IPv4 levels ...
Thanks for that, it solves my immediate problem.
--
Cheers,
Stephen Rothwell
On Fri, 25 Nov 2016 10:18:12 +0800 Eli Cooper [off-list ref] wrote:
quoted
Sounds like TSO/GSO packets are not properly segmented and therefore
dropped.
Could you first try turning off segmentation offloading for the tunnel
interface?
ethtool -K sit0 tso off gso off
On Thu, 24 Nov 2016 18:30:14 -0800 Eric Dumazet [off-list ref]
quoted
You also could try to disable TSO and see if this makes a difference
ethtool -K sixtofour0 tso off
So turning off tso brings performance up to IPv4 levels ...
Thanks for that, it solves my immediate problem.
Somehow this problem description really reminds me of a report on
netdev a bit ago, which the following patch fixed:
commit 9ee6c5dc816aa8256257f2cd4008a9291ec7e985
Author: Lance Richardson [off-list ref]
Date: Wed Nov 2 16:36:17 2016 -0400
ipv4: allow local fragmentation in ip_finish_output_gso()
Some configurations (e.g. geneve interface with default
MTU of 1500 over an ethernet interface with 1500 MTU) result
in the transmission of packets that exceed the configured MTU.
While this should be considered to be a "bad" configuration,
it is still allowed and should not result in the sending
of packets that exceed the configured MTU.
Could this be related?
I suppose it would be difficult to test this patch on this machine?
c'ya
sven-haegar
--
Three may keep a secret, if two of them are dead.
- Ben F.
From: Eli Cooper <hidden> Date: 2016-11-25 06:05:57
Hi Stephen,
On 2016/11/25 10:45, Stephen Rothwell wrote:
Hi Eli,
On Fri, 25 Nov 2016 10:18:12 +0800 Eli Cooper [off-list ref] wrote:
quoted
Sounds like TSO/GSO packets are not properly segmented and therefore
dropped.
Could you first try turning off segmentation offloading for the tunnel
interface?
ethtool -K sit0 tso off gso off
On Thu, 24 Nov 2016 18:30:14 -0800 Eric Dumazet [off-list ref]
quoted
You also could try to disable TSO and see if this makes a difference
ethtool -K sixtofour0 tso off
So turning off tso brings performance up to IPv4 levels ...
Thanks for that, it solves my immediate problem.
I think this is similar to the bug I fixed in commit ae148b085876
("ip6_tunnel: Update skb->protocol to ETH_P_IPV6 in ip6_tnl_xmit()").
I can reproduce a similar problem by applying xfrm to sit traffic.
TSO/GSO packets are dropped when IPSec is enabled, and IPv6 throughput
drops to 10s of Kbps. I am not sure if this is the same issue you
experienced, but I wrote a patch that fixed at least the issue I had.
Could you test the patch I sent to the mailing list just now?
Thanks,
Eli
From: Stephen Rothwell <hidden> Date: 2016-11-25 06:12:09
Hi Eric,
On Thu, 24 Nov 2016 19:54:04 -0800 Eric Dumazet [off-list ref] wrote:
Could you now report :
ethtool -k eth0
Features for eth0:
rx-checksumming: on
tx-checksumming: on
tx-checksum-ipv4: off [fixed]
tx-checksum-ip-generic: on
tx-checksum-ipv6: off [fixed]
tx-checksum-fcoe-crc: off [fixed]
tx-checksum-sctp: on
scatter-gather: on
tx-scatter-gather: on
tx-scatter-gather-fraglist: off [fixed]
tcp-segmentation-offload: on
tx-tcp-segmentation: on
tx-tcp-ecn-segmentation: off [fixed]
tx-tcp-mangleid-segmentation: off
tx-tcp6-segmentation: on
udp-fragmentation-offload: off [fixed]
generic-segmentation-offload: on
generic-receive-offload: on
large-receive-offload: off [fixed]
rx-vlan-offload: on
tx-vlan-offload: on
ntuple-filters: off
receive-hashing: on
highdma: on [fixed]
rx-vlan-filter: on [fixed]
vlan-challenged: off [fixed]
tx-lockless: off [fixed]
netns-local: off [fixed]
tx-gso-robust: off [fixed]
tx-fcoe-segmentation: off [fixed]
tx-gre-segmentation: on
tx-gre-csum-segmentation: on
tx-ipxip4-segmentation: on
tx-ipxip6-segmentation: on
tx-udp_tnl-segmentation: on
tx-udp_tnl-csum-segmentation: on
tx-gso-partial: on
fcoe-mtu: off [fixed]
tx-nocache-copy: off
loopback: off [fixed]
rx-fcs: off [fixed]
rx-all: off
tx-vlan-stag-hw-insert: off [fixed]
rx-vlan-stag-hw-parse: off [fixed]
rx-vlan-stag-filter: off [fixed]
l2-fwd-offload: off [fixed]
busy-poll: off [fixed]
hw-tc-offload: off [fixed]
--
Cheers,
Stephen Rothwell
From: Stephen Rothwell <hidden> Date: 2016-11-27 00:54:47
Hi Eli,
On Fri, 25 Nov 2016 14:05:04 +0800 Eli Cooper [off-list ref] wrote:
I think this is similar to the bug I fixed in commit ae148b085876
("ip6_tunnel: Update skb->protocol to ETH_P_IPV6 in ip6_tnl_xmit()").
I can reproduce a similar problem by applying xfrm to sit traffic.
TSO/GSO packets are dropped when IPSec is enabled, and IPv6 throughput
drops to 10s of Kbps. I am not sure if this is the same issue you
experienced, but I wrote a patch that fixed at least the issue I had.
Could you test the patch I sent to the mailing list just now?
Thanks for the patch!
Its a bit tricky to test since the problem only occurs in a production
machine (I tried reproducing in a VM, but the problem did not occur),
but I will try to just rebuild the sit module and see if I can insert
the modified one.
--
Cheers,
Stephen Rothwell
From: Stephen Rothwell <hidden> Date: 2016-11-27 02:02:35
Hi Eli,
On Sun, 27 Nov 2016 11:54:41 +1100 Stephen Rothwell [off-list ref] wrote:
On Fri, 25 Nov 2016 14:05:04 +0800 Eli Cooper [off-list ref] wrote:
quoted
I think this is similar to the bug I fixed in commit ae148b085876
("ip6_tunnel: Update skb->protocol to ETH_P_IPV6 in ip6_tnl_xmit()").
I can reproduce a similar problem by applying xfrm to sit traffic.
TSO/GSO packets are dropped when IPSec is enabled, and IPv6 throughput
drops to 10s of Kbps. I am not sure if this is the same issue you
experienced, but I wrote a patch that fixed at least the issue I had.
Could you test the patch I sent to the mailing list just now?
Thanks for the patch!
Its a bit tricky to test since the problem only occurs in a production
machine (I tried reproducing in a VM, but the problem did not occur),
but I will try to just rebuild the sit module and see if I can insert
the modified one.
OK, I tried your patch and unfortunately, it doesn't seem to have
worked ... I still get the large packets dropped and resent smaller.
--
Cheers,
Stephen Rothwell
From: Stephen Rothwell <hidden> Date: 2016-11-27 03:23:43
Hi Sven-Haegar,
On Fri, 25 Nov 2016 05:06:53 +0100 (CET) Sven-Haegar Koch [off-list ref] wrote:
Somehow this problem description really reminds me of a report on
netdev a bit ago, which the following patch fixed:
commit 9ee6c5dc816aa8256257f2cd4008a9291ec7e985
Author: Lance Richardson [off-list ref]
Date: Wed Nov 2 16:36:17 2016 -0400
ipv4: allow local fragmentation in ip_finish_output_gso()
Some configurations (e.g. geneve interface with default
MTU of 1500 over an ethernet interface with 1500 MTU) result
in the transmission of packets that exceed the configured MTU.
While this should be considered to be a "bad" configuration,
it is still allowed and should not result in the sending
of packets that exceed the configured MTU.
Could this be related?
I suppose it would be difficult to test this patch on this machine?
The kernel I am running on is based on 4.7.8, so the above patch
doesn't come close to applying. Most fo what it is reverting was
introduced in commit 359ebda25aa0 ("net/ipv4: Introduce IPSKB_FRAG_SEGS
bit to inet_skb_parm.flags") in v4.8-rc1.
--
Cheers,
Stephen Rothwell
From: Eli Cooper <hidden> Date: 2016-11-27 16:22:30
Hi Stephen,
On 2016/11/27 10:02, Stephen Rothwell wrote:
Hi Eli,
On Sun, 27 Nov 2016 11:54:41 +1100 Stephen Rothwell [off-list ref] wrote:
quoted
On Fri, 25 Nov 2016 14:05:04 +0800 Eli Cooper [off-list ref] wrote:
quoted
I think this is similar to the bug I fixed in commit ae148b085876
("ip6_tunnel: Update skb->protocol to ETH_P_IPV6 in ip6_tnl_xmit()").
I can reproduce a similar problem by applying xfrm to sit traffic.
TSO/GSO packets are dropped when IPSec is enabled, and IPv6 throughput
drops to 10s of Kbps. I am not sure if this is the same issue you
experienced, but I wrote a patch that fixed at least the issue I had.
Could you test the patch I sent to the mailing list just now?
Thanks for the patch!
Its a bit tricky to test since the problem only occurs in a production
machine (I tried reproducing in a VM, but the problem did not occur),
That's probably because the ethernet NIC in your VM does not support
segmentation offloading. You could, however, try reproducing it on
another (real) machine with the same driver.
quoted
but I will try to just rebuild the sit module and see if I can insert
the modified one.
OK, I tried your patch and unfortunately, it doesn't seem to have
worked ... I still get the large packets dropped and resent smaller.
It's a shame ... In my case, large packets are dropped only when xfrm is
in effect (therefore another output path is taken), and probably that's
not your case. Well, on the plus side, at least you reminded me that sit
device also needs to update skb's protocol.
Thanks,
Eli
From: "Stephen Rothwell" <redacted>
To: "Sven-Haegar Koch" <redacted>
Cc: "Eli Cooper" <redacted>, netdev@vger.kernel.org, "Eric Dumazet" <redacted>
Sent: Saturday, November 26, 2016 10:23:40 PM
Subject: Re: Large performance regression with 6in4 tunnel (sit)
Hi Sven-Haegar,
On Fri, 25 Nov 2016 05:06:53 +0100 (CET) Sven-Haegar Koch [off-list ref]
wrote:
quoted
Somehow this problem description really reminds me of a report on
netdev a bit ago, which the following patch fixed:
commit 9ee6c5dc816aa8256257f2cd4008a9291ec7e985
Author: Lance Richardson [off-list ref]
Date: Wed Nov 2 16:36:17 2016 -0400
ipv4: allow local fragmentation in ip_finish_output_gso()
Some configurations (e.g. geneve interface with default
MTU of 1500 over an ethernet interface with 1500 MTU) result
in the transmission of packets that exceed the configured MTU.
While this should be considered to be a "bad" configuration,
it is still allowed and should not result in the sending
of packets that exceed the configured MTU.
Could this be related?
I suppose it would be difficult to test this patch on this machine?
The kernel I am running on is based on 4.7.8, so the above patch
doesn't come close to applying. Most fo what it is reverting was
introduced in commit 359ebda25aa0 ("net/ipv4: Introduce IPSKB_FRAG_SEGS
bit to inet_skb_parm.flags") in v4.8-rc1.
--
Cheers,
Stephen Rothwell
@@ -224,8 +224,7 @@ static int ip_finish_output_gso(struct net *net, struct sock *sk,intret=0;/* common case: locally created skb or seglen is <= mtu */-if(((IPCB(skb)->flags&IPSKB_FORWARDED)==0)||-skb_gso_network_seglen(skb)<=mtu)+if(skb_gso_network_seglen(skb)<=mtu)returnip_finish_output2(net,sk,skb);/* Slowpath - GSO segment length is exceeding the dst MTU.
From: "Lance Richardson" <redacted>
To: "Stephen Rothwell" <redacted>
Cc: "Sven-Haegar Koch" <redacted>, "Eli Cooper" <redacted>, netdev@vger.kernel.org, "Eric Dumazet"
[off-list ref]
Sent: Monday, November 28, 2016 12:54:07 PM
Subject: Re: Large performance regression with 6in4 tunnel (sit)
quoted
From: "Stephen Rothwell" <redacted>
To: "Sven-Haegar Koch" <redacted>
Cc: "Eli Cooper" <redacted>, netdev@vger.kernel.org, "Eric
Dumazet" [off-list ref]
Sent: Saturday, November 26, 2016 10:23:40 PM
Subject: Re: Large performance regression with 6in4 tunnel (sit)
Hi Sven-Haegar,
On Fri, 25 Nov 2016 05:06:53 +0100 (CET) Sven-Haegar Koch
[off-list ref]
wrote:
quoted
Somehow this problem description really reminds me of a report on
netdev a bit ago, which the following patch fixed:
commit 9ee6c5dc816aa8256257f2cd4008a9291ec7e985
Author: Lance Richardson [off-list ref]
Date: Wed Nov 2 16:36:17 2016 -0400
ipv4: allow local fragmentation in ip_finish_output_gso()
Some configurations (e.g. geneve interface with default
MTU of 1500 over an ethernet interface with 1500 MTU) result
in the transmission of packets that exceed the configured MTU.
While this should be considered to be a "bad" configuration,
it is still allowed and should not result in the sending
of packets that exceed the configured MTU.
Could this be related?
I suppose it would be difficult to test this patch on this machine?
The kernel I am running on is based on 4.7.8, so the above patch
doesn't come close to applying. Most fo what it is reverting was
introduced in commit 359ebda25aa0 ("net/ipv4: Introduce IPSKB_FRAG_SEGS
bit to inet_skb_parm.flags") in v4.8-rc1.
--
Cheers,
Stephen Rothwell
@@ -224,8 +224,7 @@ static int ip_finish_output_gso(struct net *net, struct
sock *sk,
int ret = 0;
/* common case: locally created skb or seglen is <= mtu */
- if (((IPCB(skb)->flags & IPSKB_FORWARDED) == 0) ||
- skb_gso_network_seglen(skb) <= mtu)
+ if (skb_gso_network_seglen(skb) <= mtu)
return ip_finish_output2(net, sk, skb);
/* Slowpath - GSO segment length is exceeding the dst MTU.
BTW, I do think this would be worth trying. For the geneve case, I
measured on the order of a 10X-100X performance hit without this
patch, traces were similar to what you describe (too-large gso packets
were dropped, corresponding TCP segments were retransmitted later via
a non-gso code path).
Regards,
Lance
From: Alexander Duyck <hidden> Date: 2016-11-28 21:32:23
On Sat, Nov 26, 2016 at 7:23 PM, Stephen Rothwell [off-list ref] wrote:
Hi Sven-Haegar,
On Fri, 25 Nov 2016 05:06:53 +0100 (CET) Sven-Haegar Koch [off-list ref] wrote:
quoted
Somehow this problem description really reminds me of a report on
netdev a bit ago, which the following patch fixed:
commit 9ee6c5dc816aa8256257f2cd4008a9291ec7e985
Author: Lance Richardson [off-list ref]
Date: Wed Nov 2 16:36:17 2016 -0400
ipv4: allow local fragmentation in ip_finish_output_gso()
Some configurations (e.g. geneve interface with default
MTU of 1500 over an ethernet interface with 1500 MTU) result
in the transmission of packets that exceed the configured MTU.
While this should be considered to be a "bad" configuration,
it is still allowed and should not result in the sending
of packets that exceed the configured MTU.
Could this be related?
I suppose it would be difficult to test this patch on this machine?
The kernel I am running on is based on 4.7.8, so the above patch
doesn't come close to applying. Most fo what it is reverting was
introduced in commit 359ebda25aa0 ("net/ipv4: Introduce IPSKB_FRAG_SEGS
bit to inet_skb_parm.flags") in v4.8-rc1.
So I think I have this root caused. The problem seems to be the fact
that I chose to use lco_csum when trying to cancel out the inner IP
header from the checksum and it turns out that the transport offset is
never updated in the case of these tunnels.
For now a workaround is to just set tx-gso-partial to off on the
interface the tunnel is running over and you should be able to pass
traffic without any issues.
I have a patch for igb/igbvf that should be out in the next hour or so
which should address it.
Thanks.
- Alex
From: Stephen Rothwell <hidden> Date: 2016-11-28 22:38:22
Hi Alex,
On Mon, 28 Nov 2016 13:32:21 -0800 Alexander Duyck [off-list ref] wrote:
So I think I have this root caused. The problem seems to be the fact
that I chose to use lco_csum when trying to cancel out the inner IP
header from the checksum and it turns out that the transport offset is
never updated in the case of these tunnels.
For now a workaround is to just set tx-gso-partial to off on the
interface the tunnel is running over and you should be able to pass
traffic without any issues.
OK, so that works (even with gso and tso set to "on" on the sit
interface). Thanks.
I have a patch for igb/igbvf that should be out in the next hour or so
which should address it.
That will be a bit harder to test, but I will see what I can do.
--
Cheers,
Stephen Rothwell