From: Alexander Lobakin <hidden> Date: 2021-02-18 20:51:31
This series introduces XSK generic zerocopy xmit by adding XSK umem
pages as skb frags instead of copying data to linear space.
The only requirement for this for drivers is to be able to xmit skbs
with skb_headlen(skb) == 0, i.e. all data including hard headers
starts from frag 0.
To indicate whether a particular driver supports this, a new netdev
priv flag, IFF_TX_SKB_NO_LINEAR, is added (and declared in virtio_net
as it's already capable of doing it). So consider implementing this
in your drivers to greatly speed-up generic XSK xmit.
The first bit adds missing IFF self-definition. It's a bit out, but
"while we are here".
The fourth patch adds headroom and tailroom reservations for the
allocated skbs on XSK generic xmit path. This ensures there won't
be any unwanted skb reallocations on fast-path due to headroom and/or
tailroom driver/device requirements (own headers/descriptors etc.).
The other three add a new private flag, declare it in virtio_net
driver and introduce generic XSK zerocopy xmit itself.
The main body of work is created and done by Xuan Zhuo. His original
cover letter:
v3:
Optimized code
v2:
1. add priv_flags IFF_TX_SKB_NO_LINEAR instead of netdev_feature
2. split the patch to three:
a. add priv_flags IFF_TX_SKB_NO_LINEAR
b. virtio net add priv_flags IFF_TX_SKB_NO_LINEAR
c. When there is support this flag, construct skb without linear
space
3. use ERR_PTR() and PTR_ERR() to handle the err
v1 message log:
---------------
This patch is used to construct skb based on page to save memory copy
overhead.
This has one problem:
We construct the skb by fill the data page as a frag into the skb. In
this way, the linear space is empty, and the header information is also
in the frag, not in the linear space, which is not allowed for some
network cards. For example, Mellanox Technologies MT27710 Family
[ConnectX-4 Lx] will get the following error message:
mlx5_core 0000:3b:00.1 eth1: Error cqe on cqn 0x817, ci 0x8,
qn 0x1dbb, opcode 0xd, syndrome 0x1, vendor syndrome 0x68
00000000: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
00000010: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
00000020: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
00000030: 00 00 00 00 60 10 68 01 0a 00 1d bb 00 0f 9f d2
WQE DUMP: WQ size 1024 WQ cur size 0, WQE index 0xf, len: 64
00000000: 00 00 0f 0a 00 1d bb 03 00 00 00 08 00 00 00 00
00000010: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
00000020: 00 00 00 2b 00 08 00 00 00 00 00 05 9e e3 08 00
00000030: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
mlx5_core 0000:3b:00.1 eth1: ERR CQE on SQ: 0x1dbb
I also tried to use build_skb to construct skb, but because of the
existence of skb_shinfo, it must be behind the linear space, so this
method is not working. We can't put skb_shinfo on desc->addr, it will be
exposed to users, this is not safe.
Finally, I added a feature NETIF_F_SKB_NO_LINEAR to identify whether the
network card supports the header information of the packet in the frag
and not in the linear space.
---------------- Performance Testing ------------
The test environment is Aliyun ECS server.
Test cmd:
xdpsock-ieth0-t-S-s<msgsize>
Test result data:
size 64 512 1024 1500
copy 1916747 1775988 1600203 1440054
page 1974058 1953655 1945463 1904478
percent 3.0% 10.0% 21.58% 32.3%
From v7 [4]:
- drop netdev priv flags rework (will be issued separately);
- pick up Acks from John.
From v6 [3]:
- rebase ontop of bpf-next after merge with net-next;
- address kdoc warnings.
From v5 [2]:
- fix a refcount leak in 0006 introduced in v4.
From v4 [1]:
- fix 0002 build error due to inverted static_assert() condition
(0day bot);
- collect two Acked-bys (Magnus).
From v3 [0]:
- refactor netdev_priv_flags to make it easier to add new ones and
prevent bitwidth overflow;
- add headroom (both standard and zerocopy) and tailroom (standard)
reservation in skb for drivers to avoid potential reallocations;
- fix skb->truesize accounting;
- misc comment rewords.
[0] https://lore.kernel.org/netdev/cover.1611236588.git.xuanzhuo@linux.alibaba.com
[1] https://lore.kernel.org/netdev/20210216113740.62041-1-alobakin@pm.me
[2] https://lore.kernel.org/netdev/20210216143333.5861-1-alobakin@pm.me
[3] https://lore.kernel.org/netdev/20210216172640.374487-1-alobakin@pm.me
[4] https://lore.kernel.org/netdev/20210217120003.7938-1-alobakin@pm.me
Alexander Lobakin (2):
netdevice: add missing IFF_PHONY_HEADROOM self-definition
xsk: respect device's headroom and tailroom on generic xmit path
Xuan Zhuo (3):
net: add priv_flags for allow tx skb without linear
virtio-net: support IFF_TX_SKB_NO_LINEAR
xsk: build skb by page (aka generic zerocopy xmit)
drivers/net/virtio_net.c | 3 +-
include/linux/netdevice.h | 5 ++
net/xdp/xsk.c | 114 ++++++++++++++++++++++++++++++++------
3 files changed, 103 insertions(+), 19 deletions(-)
--
2.30.1
From: Alexander Lobakin <hidden> Date: 2021-02-18 20:51:32
This is harmless for now, but can be fatal for future refactors.
Fixes: 871b642adebe3 ("netdev: introduce ndo_set_rx_headroom")
Signed-off-by: Alexander Lobakin <redacted>
Acked-by: John Fastabend <john.fastabend@gmail.com>
---
include/linux/netdevice.h | 1 +
1 file changed, 1 insertion(+)
From: Alexander Lobakin <hidden> Date: 2021-02-18 20:52:40
From: Xuan Zhuo <xuanzhuo@linux.alibaba.com>
Virtio net supports the case where the skb linear space is empty, so add
priv_flags.
Signed-off-by: Xuan Zhuo <xuanzhuo@linux.alibaba.com>
Acked-by: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Alexander Lobakin <redacted>
Acked-by: John Fastabend <john.fastabend@gmail.com>
---
drivers/net/virtio_net.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
@@ -2972,7 +2972,8 @@ static int virtnet_probe(struct virtio_device *vdev)return-ENOMEM;/* Set up network device as normal. */-dev->priv_flags|=IFF_UNICAST_FLT|IFF_LIVE_ADDR_CHANGE;+dev->priv_flags|=IFF_UNICAST_FLT|IFF_LIVE_ADDR_CHANGE|+IFF_TX_SKB_NO_LINEAR;dev->netdev_ops=&virtnet_netdev;dev->features=NETIF_F_HIGHDMA;--
From: Alexander Lobakin <hidden> Date: 2021-02-18 20:52:41
xsk_generic_xmit() allocates a new skb and then queues it for
xmitting. The size of new skb's headroom is desc->len, so it comes
to the driver/device with no reserved headroom and/or tailroom.
Lots of drivers need some headroom (and sometimes tailroom) to
prepend (and/or append) some headers or data, e.g. CPU tags,
device-specific headers/descriptors (LSO, TLS etc.), and if case
of no available space skb_cow_head() will reallocate the skb.
Reallocations are unwanted on fast-path, especially when it comes
to XDP, so generic XSK xmit should reserve the spaces declared in
dev->needed_headroom and dev->needed tailroom to avoid them.
Note on max(NET_SKB_PAD, L1_CACHE_ALIGN(dev->needed_headroom)):
Usually, output functions reserve LL_RESERVED_SPACE(dev), which
consists of dev->hard_header_len + dev->needed_headroom, aligned
by 16.
However, on XSK xmit hard header is already here in the chunk, so
hard_header_len is not needed. But it'd still be better to align
data up to cacheline, while reserving no less than driver requests
for headroom. NET_SKB_PAD here is to double-insure there will be
no reallocations even when the driver advertises no needed_headroom,
but in fact need it (not so rare case).
Fixes: 35fcde7f8deb ("xsk: support for Tx")
Signed-off-by: Alexander Lobakin <redacted>
Acked-by: Magnus Karlsson <magnus.karlsson@intel.com>
Acked-by: John Fastabend <john.fastabend@gmail.com>
---
net/xdp/xsk.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
From: Alexander Lobakin <hidden> Date: 2021-02-18 20:53:28
From: Xuan Zhuo <xuanzhuo@linux.alibaba.com>
This patch is used to construct skb based on page to save memory copy
overhead.
This function is implemented based on IFF_TX_SKB_NO_LINEAR. Only the
network card priv_flags supports IFF_TX_SKB_NO_LINEAR will use page to
directly construct skb. If this feature is not supported, it is still
necessary to copy data to construct skb.
---------------- Performance Testing ------------
The test environment is Aliyun ECS server.
Test cmd:
xdpsock-ieth0-t-S-s<msgsize>
Test result data:
size 64 512 1024 1500
copy 1916747 1775988 1600203 1440054
page 1974058 1953655 1945463 1904478
percent 3.0% 10.0% 21.58% 32.3%
Signed-off-by: Xuan Zhuo <xuanzhuo@linux.alibaba.com>
Reviewed-by: Dust Li <dust.li@linux.alibaba.com>
[ alobakin:
- expand subject to make it clearer;
- improve skb->truesize calculation;
- reserve some headroom in skb for drivers;
- tailroom is not needed as skb is non-linear ]
Signed-off-by: Alexander Lobakin <redacted>
Acked-by: Magnus Karlsson <magnus.karlsson@intel.com>
Acked-by: John Fastabend <john.fastabend@gmail.com>
---
net/xdp/xsk.c | 120 ++++++++++++++++++++++++++++++++++++++++----------
1 file changed, 96 insertions(+), 24 deletions(-)
@@ -454,56 +545,37 @@ static int xsk_generic_xmit(struct sock *sk)structsk_buff*skb;unsignedlongflags;interr=0;-u32hr,tr;mutex_lock(&xs->mutex);if(xs->queue_id>=xs->dev->real_num_tx_queues)gotoout;-hr=max(NET_SKB_PAD,L1_CACHE_ALIGN(xs->dev->needed_headroom));-tr=xs->dev->needed_tailroom;-while(xskq_cons_peek_desc(xs->tx,&desc,xs->pool)){-char*buffer;-u64addr;-u32len;-if(max_batch--==0){err=-EAGAIN;gotoout;}-len=desc.len;-skb=sock_alloc_send_skb(sk,hr+len+tr,1,&err);-if(unlikely(!skb))+skb=xsk_build_skb(xs,&desc);+if(IS_ERR(skb)){+err=PTR_ERR(skb);gotoout;+}-skb_reserve(skb,hr);-skb_put(skb,len);--addr=desc.addr;-buffer=xsk_buff_raw_get_data(xs->pool,addr);-err=skb_store_bits(skb,0,buffer,len);/* This is the backpressure mechanism for the Tx path.*Reservespaceinthecompletionqueueandonlyproceed*ifthereisspaceinit.Thisavoidshavingtoimplement*anybufferingintheTxpath.*/spin_lock_irqsave(&xs->pool->cq_lock,flags);-if(unlikely(err)||xskq_prod_reserve(xs->pool->cq)){+if(xskq_prod_reserve(xs->pool->cq)){spin_unlock_irqrestore(&xs->pool->cq_lock,flags);kfree_skb(skb);gotoout;}spin_unlock_irqrestore(&xs->pool->cq_lock,flags);-skb->dev=xs->dev;-skb->priority=sk->sk_priority;-skb->mark=sk->sk_mark;-skb_shinfo(skb)->destructor_arg=(void*)(long)desc.addr;-skb->destructor=xsk_destruct_skb;-err=__dev_direct_xmit(skb,xs->queue_id);if(err==NETDEV_TX_BUSY){/* Tell user-space to retry the send */--
From: Daniel Borkmann <daniel@iogearbox.net> Date: 2021-02-25 00:47:40
On 2/18/21 9:49 PM, Alexander Lobakin wrote:
This series introduces XSK generic zerocopy xmit by adding XSK umem
pages as skb frags instead of copying data to linear space.
The only requirement for this for drivers is to be able to xmit skbs
with skb_headlen(skb) == 0, i.e. all data including hard headers
starts from frag 0.
To indicate whether a particular driver supports this, a new netdev
priv flag, IFF_TX_SKB_NO_LINEAR, is added (and declared in virtio_net
as it's already capable of doing it). So consider implementing this
in your drivers to greatly speed-up generic XSK xmit.