Performance regression 3.11 with macvlan between 2 linux guests (bisected)

5 messages, 3 authors, 2013-08-22 · open the first message on its own page

Performance regression 3.11 with macvlan between 2 linux guests (bisected)

From: Christian Borntraeger <hidden>
Date: 2013-08-21 12:38:50

Vlad,

the patch 

commit 3e4f8b787370978733ca6cae452720a4f0c296b8
Author: Vlad Yasevich [off-list ref]
Date:   Tue Jun 25 16:04:22 2013 -0400

    macvtap: Perform GSO on forwarding path.

causes a severe performance regression for 2 Linux guests with virtio-net
connected via macvlan/vtap doing iperf workload. (2GBit vs. 20Gbit)

If I understand the patch correctly, we now check for gso depending on the
macvlan features. If the underlying hardware does not support the necessary
offloads then we will do segmentation, even if we keep the whole traffic internal
between two guests and even if both guests supports LRO,GSO etc.
---snip---
[...]
+       features = netif_skb_features(skb) & vlan->tap_features;
+       if (netif_needs_gso(skb, features)) {
+               struct sk_buff *segs = __skb_gso_segment(skb, features, false);
[...]

---snip---
Shouldnt we take the features of the target device or even better do the gsoing in the
target device driver?

Christian

FYI, the underlying HW has:

Features for eth0:
rx-checksumming: off [fixed]
tx-checksumming: off
	tx-checksum-ipv4: off [fixed]
	tx-checksum-ip-generic: off [fixed]
	tx-checksum-ipv6: off [fixed]
	tx-checksum-fcoe-crc: off [fixed]
	tx-checksum-sctp: off [fixed]
scatter-gather: off
	tx-scatter-gather: off [fixed]
	tx-scatter-gather-fraglist: off [fixed]
tcp-segmentation-offload: off
	tx-tcp-segmentation: off [fixed]
	tx-tcp-ecn-segmentation: off [fixed]
	tx-tcp6-segmentation: off [fixed]
udp-fragmentation-offload: off [fixed]
generic-segmentation-offload: off [requested on]
generic-receive-offload: on
large-receive-offload: off [fixed]
rx-vlan-offload: off [fixed]
tx-vlan-offload: off [fixed]
ntuple-filters: off [fixed]
receive-hashing: off [fixed]
highdma: off [fixed]
rx-vlan-filter: on [fixed]
vlan-challenged: off [fixed]
tx-lockless: off [fixed]
netns-local: off [fixed]
tx-gso-robust: off [fixed]
tx-fcoe-segmentation: off [fixed]
tx-gre-segmentation: off [fixed]
tx-udp_tnl-segmentation: off [fixed]
tx-mpls-segmentation: off [fixed]
fcoe-mtu: off [fixed]
tx-nocache-copy: off
loopback: off [fixed]
rx-fcs: off [fixed]
rx-all: off [fixed]
tx-vlan-stag-hw-insert: off [fixed]
rx-vlan-stag-hw-parse: off [fixed]
rx-vlan-stag-filter: off [fixed]

Re: Performance regression 3.11 with macvlan between 2 linux guests (bisected)

From: Vlad Yasevich <hidden>
Date: 2013-08-21 18:05:00

On 08/21/2013 08:38 AM, Christian Borntraeger wrote:
Vlad,

the patch

commit 3e4f8b787370978733ca6cae452720a4f0c296b8
Author: Vlad Yasevich [off-list ref]
Date:   Tue Jun 25 16:04:22 2013 -0400

     macvtap: Perform GSO on forwarding path.

causes a severe performance regression for 2 Linux guests with virtio-net
connected via macvlan/vtap doing iperf workload. (2GBit vs. 20Gbit)

If I understand the patch correctly, we now check for gso depending on the
macvlan features. If the underlying hardware does not support the necessary
offloads then we will do segmentation, even if we keep the whole traffic internal
between two guests and even if both guests supports LRO,GSO etc.
---snip---
[...]
+       features = netif_skb_features(skb) & vlan->tap_features;
+       if (netif_needs_gso(skb, features)) {
+               struct sk_buff *segs = __skb_gso_segment(skb, features, false);
[...]

---snip---
Shouldnt we take the features of the target device or even better do the gsoing in the
target device driver?
A corrected patch has been sent upstream.  We take into consideration 
the features on the target device that the user/vm has specified.
If the VM has enabled the TSO flags, then nothing will happen to the
GSO packet.  However, if the TSO flag is off, segmentation will be 
performed.

-vlad
Christian

FYI, the underlying HW has:

Features for eth0:
rx-checksumming: off [fixed]
tx-checksumming: off
	tx-checksum-ipv4: off [fixed]
	tx-checksum-ip-generic: off [fixed]
	tx-checksum-ipv6: off [fixed]
	tx-checksum-fcoe-crc: off [fixed]
	tx-checksum-sctp: off [fixed]
scatter-gather: off
	tx-scatter-gather: off [fixed]
	tx-scatter-gather-fraglist: off [fixed]
tcp-segmentation-offload: off
	tx-tcp-segmentation: off [fixed]
	tx-tcp-ecn-segmentation: off [fixed]
	tx-tcp6-segmentation: off [fixed]
udp-fragmentation-offload: off [fixed]
generic-segmentation-offload: off [requested on]
generic-receive-offload: on
large-receive-offload: off [fixed]
rx-vlan-offload: off [fixed]
tx-vlan-offload: off [fixed]
ntuple-filters: off [fixed]
receive-hashing: off [fixed]
highdma: off [fixed]
rx-vlan-filter: on [fixed]
vlan-challenged: off [fixed]
tx-lockless: off [fixed]
netns-local: off [fixed]
tx-gso-robust: off [fixed]
tx-fcoe-segmentation: off [fixed]
tx-gre-segmentation: off [fixed]
tx-udp_tnl-segmentation: off [fixed]
tx-mpls-segmentation: off [fixed]
fcoe-mtu: off [fixed]
tx-nocache-copy: off
loopback: off [fixed]
rx-fcs: off [fixed]
rx-all: off [fixed]
tx-vlan-stag-hw-insert: off [fixed]
rx-vlan-stag-hw-parse: off [fixed]
rx-vlan-stag-filter: off [fixed]

Re: Performance regression 3.11 with macvlan between 2 linux guests (bisected)

From: Vlad Yasevich <hidden>
Date: 2013-08-21 18:12:14

On 08/21/2013 02:04 PM, Vlad Yasevich wrote:
On 08/21/2013 08:38 AM, Christian Borntraeger wrote:
quoted
Vlad,

the patch

commit 3e4f8b787370978733ca6cae452720a4f0c296b8
Author: Vlad Yasevich [off-list ref]
Date:   Tue Jun 25 16:04:22 2013 -0400

     macvtap: Perform GSO on forwarding path.

causes a severe performance regression for 2 Linux guests with virtio-net
connected via macvlan/vtap doing iperf workload. (2GBit vs. 20Gbit)

If I understand the patch correctly, we now check for gso depending on
the
macvlan features. If the underlying hardware does not support the
necessary
offloads then we will do segmentation, even if we keep the whole
traffic internal
between two guests and even if both guests supports LRO,GSO etc.
---snip---
[...]
+       features = netif_skb_features(skb) & vlan->tap_features;
+       if (netif_needs_gso(skb, features)) {
+               struct sk_buff *segs = __skb_gso_segment(skb,
features, false);
[...]

---snip---
Shouldnt we take the features of the target device or even better do
the gsoing in the
target device driver?
A corrected patch has been sent upstream.  We take into consideration
the features on the target device that the user/vm has specified.
If the VM has enabled the TSO flags, then nothing will happen to the
GSO packet.  However, if the TSO flag is off, segmentation will be
performed.
Particularly.  This commit should fix the issue:
commit a567dd6252263c8147b7269df5d03d9e31463e11
  macvtap: simplify usage of tap_features

-vlad
-vlad
quoted
Christian

FYI, the underlying HW has:

Features for eth0:
rx-checksumming: off [fixed]
tx-checksumming: off
    tx-checksum-ipv4: off [fixed]
    tx-checksum-ip-generic: off [fixed]
    tx-checksum-ipv6: off [fixed]
    tx-checksum-fcoe-crc: off [fixed]
    tx-checksum-sctp: off [fixed]
scatter-gather: off
    tx-scatter-gather: off [fixed]
    tx-scatter-gather-fraglist: off [fixed]
tcp-segmentation-offload: off
    tx-tcp-segmentation: off [fixed]
    tx-tcp-ecn-segmentation: off [fixed]
    tx-tcp6-segmentation: off [fixed]
udp-fragmentation-offload: off [fixed]
generic-segmentation-offload: off [requested on]
generic-receive-offload: on
large-receive-offload: off [fixed]
rx-vlan-offload: off [fixed]
tx-vlan-offload: off [fixed]
ntuple-filters: off [fixed]
receive-hashing: off [fixed]
highdma: off [fixed]
rx-vlan-filter: on [fixed]
vlan-challenged: off [fixed]
tx-lockless: off [fixed]
netns-local: off [fixed]
tx-gso-robust: off [fixed]
tx-fcoe-segmentation: off [fixed]
tx-gre-segmentation: off [fixed]
tx-udp_tnl-segmentation: off [fixed]
tx-mpls-segmentation: off [fixed]
fcoe-mtu: off [fixed]
tx-nocache-copy: off
loopback: off [fixed]
rx-fcs: off [fixed]
rx-all: off [fixed]
tx-vlan-stag-hw-insert: off [fixed]
rx-vlan-stag-hw-parse: off [fixed]
rx-vlan-stag-filter: off [fixed]

Re: Performance regression 3.11 with macvlan between 2 linux guests (bisected)

From: Christian Borntraeger <hidden>
Date: 2013-08-22 08:23:39

On 21/08/13 20:12, Vlad Yasevich wrote:
quoted
A corrected patch has been sent upstream.  We take into consideration
the features on the target device that the user/vm has specified.
If the VM has enabled the TSO flags, then nothing will happen to the
GSO packet.  However, if the TSO flag is off, segmentation will be
performed.
Particularly.  This commit should fix the issue:
commit a567dd6252263c8147b7269df5d03d9e31463e11
 macvtap: simplify usage of tap_features
Tested-by: Christian Borntraeger <redacted>

The patch is still in net.git, but not in Linus git. Are we going to push this
for 3.11?

Re: Performance regression 3.11 with macvlan between 2 linux guests (bisected)

From: David Miller <davem@davemloft.net>
Date: 2013-08-22 08:29:57

From: Christian Borntraeger <redacted>
Date: Thu, 22 Aug 2013 10:23:32 +0200
On 21/08/13 20:12, Vlad Yasevich wrote:
quoted
quoted
A corrected patch has been sent upstream.  We take into consideration
the features on the target device that the user/vm has specified.
If the VM has enabled the TSO flags, then nothing will happen to the
GSO packet.  However, if the TSO flag is off, segmentation will be
performed.
Particularly.  This commit should fix the issue:
commit a567dd6252263c8147b7269df5d03d9e31463e11
 macvtap: simplify usage of tap_features
Tested-by: Christian Borntraeger <redacted>

The patch is still in net.git, but not in Linus git. Are we going to push this
for 3.11?
Of course.  Anything in 'net' is intended to make it into Linus's tree.

I generally push things to Linus every week or two.
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help