From: Igor Royzis <hidden> Date: 2014-05-20 11:24:37
Fix accessing GSO fragments memory (and a possible corruption therefore) after
reporting completion in a zero copy callback. The previous fix in the commit 1fd819ec
orphaned frags which eliminates zero copy advantages. The fix makes the completion
called after all the fragments were processed avoiding unnecessary orphaning/copying
from userspace.
The GSO fragments corruption issue was observed in a typical QEMU/KVM VM setup that
hosts a Windows guest (since QEMU virtio-net Windows driver doesn't support GRO).
The fix has been verified by running the HCK OffloadLSO test.
Signed-off-by: Igor Royzis <redacted>
Signed-off-by: Anton Nayshtut <redacted>
---
include/linux/skbuff.h | 1 +
net/core/skbuff.c | 18 +++++++++++++-----
2 files changed, 14 insertions(+), 5 deletions(-)
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2014-05-20 11:52:52
On Tue, May 20, 2014 at 02:24:21PM +0300, Igor Royzis wrote:
Fix accessing GSO fragments memory (and a possible corruption therefore) after
reporting completion in a zero copy callback. The previous fix in the commit 1fd819ec
orphaned frags which eliminates zero copy advantages. The fix makes the completion
called after all the fragments were processed avoiding unnecessary orphaning/copying
from userspace.
The GSO fragments corruption issue was observed in a typical QEMU/KVM VM setup that
hosts a Windows guest (since QEMU virtio-net Windows driver doesn't support GRO).
The fix has been verified by running the HCK OffloadLSO test.
Signed-off-by: Igor Royzis <redacted>
Signed-off-by: Anton Nayshtut <redacted>
OK but with 1fd819ec there's no corruption, correct?
So this patch is in fact an optimization?
If true, I'd like to see some performance numbers please.
Thanks!
From: Anton Nayshtut <hidden> Date: 2014-05-20 12:10:58
Hi Michael,
You're absolutely right.
We detected the actual corruption running MS HCK on earlier kernels,
before the 1fd819ec, so the patch was developed as a fix for this issue.
However, 1fd819ec fixes the corruption and now it's only an optimization
that re-enables the zero copy for this case.
We're collecting the numbers right now and will post them as soon as
possible.
Best Regards,
Anton
On 5/20/2014 2:50 PM, Michael S. Tsirkin wrote:
On Tue, May 20, 2014 at 02:24:21PM +0300, Igor Royzis wrote:
quoted
Fix accessing GSO fragments memory (and a possible corruption therefore) after
reporting completion in a zero copy callback. The previous fix in the commit 1fd819ec
orphaned frags which eliminates zero copy advantages. The fix makes the completion
called after all the fragments were processed avoiding unnecessary orphaning/copying
from userspace.
The GSO fragments corruption issue was observed in a typical QEMU/KVM VM setup that
hosts a Windows guest (since QEMU virtio-net Windows driver doesn't support GRO).
The fix has been verified by running the HCK OffloadLSO test.
Signed-off-by: Igor Royzis <redacted>
Signed-off-by: Anton Nayshtut <redacted>
OK but with 1fd819ec there's no corruption, correct?
So this patch is in fact an optimization?
If true, I'd like to see some performance numbers please.
Thanks!
From: Eric Dumazet <hidden> Date: 2014-05-20 14:29:03
On Tue, 2014-05-20 at 14:24 +0300, Igor Royzis wrote:
quoted hunk
Fix accessing GSO fragments memory (and a possible corruption therefore) after
reporting completion in a zero copy callback. The previous fix in the commit 1fd819ec
orphaned frags which eliminates zero copy advantages. The fix makes the completion
called after all the fragments were processed avoiding unnecessary orphaning/copying
from userspace.
The GSO fragments corruption issue was observed in a typical QEMU/KVM VM setup that
hosts a Windows guest (since QEMU virtio-net Windows driver doesn't support GRO).
The fix has been verified by running the HCK OffloadLSO test.
Signed-off-by: Igor Royzis <redacted>
Signed-off-by: Anton Nayshtut <redacted>
---
include/linux/skbuff.h | 1 +
net/core/skbuff.c | 18 +++++++++++++-----
2 files changed, 14 insertions(+), 5 deletions(-)
Before your patch :
sizeof(struct skb_shared_info)=0x140
offsetof(struct skb_shared_info, frags[1])=0x40
SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) -> 0x140
After your patch :
sizeof(struct skb_shared_info)=0x148
offsetof(struct skb_shared_info, frags[1])=0x48
SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) -> 0x180
Thats a serious bump, because it increases all skb truesizes, and
typical skb with one fragment will use 2 cache lines instead of one in
struct skb_shared_info, so this adds memory pressure in fast path.
So while this patch might increase performance for some workloads,
it generally decreases performance on many others.
From: Eric Dumazet <hidden> Date: 2014-05-20 16:05:45
On Tue, 2014-05-20 at 07:28 -0700, Eric Dumazet wrote:
On Tue, 2014-05-20 at 14:24 +0300, Igor Royzis wrote:
quoted
Fix accessing GSO fragments memory (and a possible corruption therefore) after
reporting completion in a zero copy callback. The previous fix in the commit 1fd819ec
orphaned frags which eliminates zero copy advantages. The fix makes the completion
called after all the fragments were processed avoiding unnecessary orphaning/copying
from userspace.
The GSO fragments corruption issue was observed in a typical QEMU/KVM VM setup that
hosts a Windows guest (since QEMU virtio-net Windows driver doesn't support GRO).
The fix has been verified by running the HCK OffloadLSO test.
It looks like all segments (generated by GSO segmentation) should share
original ubuf_info, and that it should be refcounted.
A nightmare I suppose...
(transferring the ubuf_info from original skb to last segment would be
racy, as the last segment could be freed _before_ previous ones, in case
a drop happens in qdisc layer, or packets are reordered by netem)
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2014-05-20 16:19:29
On Tue, May 20, 2014 at 09:05:38AM -0700, Eric Dumazet wrote:
On Tue, 2014-05-20 at 07:28 -0700, Eric Dumazet wrote:
quoted
On Tue, 2014-05-20 at 14:24 +0300, Igor Royzis wrote:
quoted
Fix accessing GSO fragments memory (and a possible corruption therefore) after
reporting completion in a zero copy callback. The previous fix in the commit 1fd819ec
orphaned frags which eliminates zero copy advantages. The fix makes the completion
called after all the fragments were processed avoiding unnecessary orphaning/copying
from userspace.
The GSO fragments corruption issue was observed in a typical QEMU/KVM VM setup that
hosts a Windows guest (since QEMU virtio-net Windows driver doesn't support GRO).
The fix has been verified by running the HCK OffloadLSO test.
It looks like all segments (generated by GSO segmentation) should share
original ubuf_info, and that it should be refcounted.
A nightmare I suppose...
That's what skb_frag_ref tried to do only for fragments, I guess.
(transferring the ubuf_info from original skb to last segment would be
racy, as the last segment could be freed _before_ previous ones, in case
a drop happens in qdisc layer, or packets are reordered by netem)
From: Igor Royzis <hidden> Date: 2014-05-25 10:54:58
If true, I'd like to see some performance numbers please.
The numbers have been obtained by running iperf between 2 QEMU Win2012
VMs, 4 vCPU/ 4GB RAM each.
iperf parameters: -w 256K -l 256K -t 300
Original kernel 3.15.0-rc5: 34.4 Gbytes transferred, 984
Mbits/sec bandwidth.
Kernel 3.15.0-rc5 with our patch: 42.5 Gbytes transferred, 1.22
Gbits/sec bandwidth.
Overall improvement is about 24%.
Below are raw iperf outputs.
kernel 3.15.0-rc5:
C:\iperf>iperf -c 192.168.11.2 -w 256K -l 256K -t 300
------------------------------------------------------------
Client connecting to 192.168.11.2, TCP port 5001
TCP window size: 256 KByte
------------------------------------------------------------
[ 3] local 192.168.11.1 port 49167 connected with 192.168.11.2 port 5001
[ ID] Interval Transfer Bandwidth
[ 3] 0.0-300.7 sec 34.4 GBytes 984 Mbits/sec
kernel 3.15.0-rc5-patched:
C:\iperf>iperf -c 192.168.11.2 -w 256K -l 256K -t 300
------------------------------------------------------------
Client connecting to 192.168.11.2, TCP port 5001
TCP window size: 256 KByte
------------------------------------------------------------
[ 3] local 192.168.11.1 port 49167 connected with 192.168.11.2 port 5001
[ ID] Interval Transfer Bandwidth
[ 3] 0.0-300.7 sec 42.5 GBytes 1.22 Gbits/sec
On Tue, May 20, 2014 at 2:50 PM, Michael S. Tsirkin [off-list ref] wrote:
On Tue, May 20, 2014 at 02:24:21PM +0300, Igor Royzis wrote:
quoted
Fix accessing GSO fragments memory (and a possible corruption therefore) after
reporting completion in a zero copy callback. The previous fix in the commit 1fd819ec
orphaned frags which eliminates zero copy advantages. The fix makes the completion
called after all the fragments were processed avoiding unnecessary orphaning/copying
from userspace.
The GSO fragments corruption issue was observed in a typical QEMU/KVM VM setup that
hosts a Windows guest (since QEMU virtio-net Windows driver doesn't support GRO).
The fix has been verified by running the HCK OffloadLSO test.
Signed-off-by: Igor Royzis <redacted>
Signed-off-by: Anton Nayshtut <redacted>
OK but with 1fd819ec there's no corruption, correct?
So this patch is in fact an optimization?
If true, I'd like to see some performance numbers please.
Thanks!
From: Igor Royzis <hidden> Date: 2014-05-25 11:09:07
On Tue, May 20, 2014 at 5:28 PM, Eric Dumazet [off-list ref] wrote:
On Tue, 2014-05-20 at 14:24 +0300, Igor Royzis wrote:
quoted
Fix accessing GSO fragments memory (and a possible corruption therefore) after
reporting completion in a zero copy callback. The previous fix in the commit 1fd819ec
orphaned frags which eliminates zero copy advantages. The fix makes the completion
called after all the fragments were processed avoiding unnecessary orphaning/copying
from userspace.
The GSO fragments corruption issue was observed in a typical QEMU/KVM VM setup that
hosts a Windows guest (since QEMU virtio-net Windows driver doesn't support GRO).
The fix has been verified by running the HCK OffloadLSO test.
Signed-off-by: Igor Royzis <redacted>
Signed-off-by: Anton Nayshtut <redacted>
---
include/linux/skbuff.h | 1 +
net/core/skbuff.c | 18 +++++++++++++-----
2 files changed, 14 insertions(+), 5 deletions(-)
Before your patch :
sizeof(struct skb_shared_info)=0x140
offsetof(struct skb_shared_info, frags[1])=0x40
SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) -> 0x140
After your patch :
sizeof(struct skb_shared_info)=0x148
offsetof(struct skb_shared_info, frags[1])=0x48
SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) -> 0x180
Thats a serious bump, because it increases all skb truesizes, and
typical skb with one fragment will use 2 cache lines instead of one in
struct skb_shared_info, so this adds memory pressure in fast path.
So while this patch might increase performance for some workloads,
it generally decreases performance on many others.
Would it "ease" the memory cache penalty if we moved the parent
fragment pointer from skb_shared_info to skbuff itself?