The common case for AF_XDP sockets (xsks) is creating a single xsk on a queue for sending and
receiving frames as this is analogous to HW packet steering through RSS and other classification
methods in the NIC. AF_XDP uses the xdp redirect infrastructure to direct packets to the socket. It
was designed for the much more complicated case of DEVMAP xdp_redirects which directs traffic to
another netdev and thus potentially another driver. In the xsk redirect case, by skipping the
unnecessary parts of this common code we can significantly improve performance and pave the way
for batching in the driver. This RFC proposes one such way to simplify the infrastructure which
yields a 27% increase in throughput and a decrease in cycles per packet of 24 cycles [1]. The goal
of this RFC is to start a discussion on how best to simplify the single-socket datapath while
providing one method as an example.
Current approach:
1. XSK pointer: an xsk is created and a handle to the xsk is stored in the XSKMAP.
2. XDP program: bpf_redirect_map helper triggers the XSKMAP lookup which stores the result (handle
to the xsk) and the map type (XSKMAP) in the percpu bpf_redirect_info struct. The XDP_REDIRECT
action is returned.
3. XDP_REDIRECT handling called by the driver: the map type (XSKMAP) is read from the
bpf_redirect_info which selects the xsk_map_redirect path. The xsk pointer is retrieved from the
bpf_redirect_info and the XDP descriptor is pushed to the xsk's Rx ring. The socket is added to a
list for flushing later.
4. xdp_do_flush: iterate through the lists of all maps that can be used for redirect (CPUMAP,
DEVMAP and XSKMAP). When XSKMAP is flushed, go through all xsks that had any traffic redirected to
them and bump the Rx ring head pointer(s).
For the end goal of submitting the descriptor to the Rx ring and bumping the head pointer of that
ring, only some of these steps are needed. The rest is overhead. The bpf_redirect_map
infrastructure is needed for all other redirect operations, but is not necessary when redirecting
to a single AF_XDP socket. And similarly, flushing the list for every map type in step 4 is not
necessary when only one socket needs to be flushed.
Proposed approach:
1. XSK pointer: an xsk is created and a handle to the xsk is stored both in the XSKMAP and also the
netdev_rx_queue struct.
2. XDP program: new bpf_redirect_xsk helper returns XDP_REDIRECT_XSK.
3. XDP_REDIRECT_XSK handling called by the driver: the xsk pointer is retrieved from the
netdev_rx_queue struct and the XDP descriptor is pushed to the xsk's Rx ring.
4. xsk_flush: fetch the handle from the netdev_rx_queue and flush the xsk.
This fast path is triggered on XDP_REDIRECT_XSK if:
(i) AF_XDP socket SW Rx ring configured
(ii) Exactly one xsk attached to the queue
If any of these conditions are not met, fall back to the same behavior as the original approach:
xdp_redirect_map. This is handled under-the-hood in the new bpf_xdp_redirect_xsk helper so the user
does not need to be aware of these conditions.
Batching:
With this new approach it is possible to optimize the driver by submitting a batch of descriptors
to the Rx ring in Step 3 of the new approach by simply verifying that the action returned from
every program run of each packet in a batch equals XDP_REDIRECT_XSK. That's because with this
action we know the socket to redirect to will be the same for each packet in the batch. This is
not possible with XDP_REDIRECT because the xsk pointer is stored in the bpf_redirect_info and not
guaranteed to be the same for every packet in a batch.
[1] Performance:
The benchmarks were performed on VM running a 2.4GHz Ice Lake host with an i40e device passed
through. The xdpsock app was run on a single core with busy polling and configured in 'rxonly' mode.
./xdpsock -i <iface> -r -B
The improvement in throughput when using the new bpf helper and XDP action was measured at ~13% for
scalar processing, with reduction in cycles per packet of ~13. A further ~14% improvement in
throughput and reduction of ~11 cycles per packet was measured when the batched i40e driver path
was used, for a total improvement of ~27% in throughput and reduction of ~24 cycles per packet.
Other approaches considered:
Two other approaches were considered. The advantage of both being that neither involved introducing
a new XDP action. The first alternative approach considered was to create a new map type
BPF_MAP_TYPE_XSKMAP_DIRECT. When the XDP_REDIRECT action was returned, this map type could be
checked and used as an indicator to skip the map lookup and use the netdev_rx_queue xsk instead.
The second approach considered was similar and involved using a new bpf_redirect_info flag which
could be used in a similar fashion.
While both approaches yielded a performance improvement they were measured at about half of what
was measured for the approach outlined in this RFC. It seems using bpf_redirect_info is too
expensive.
Also, centralised processing of XDP actions was investigated. This would involve porting all drivers
to a common interface for handling XDP actions which would greatly simplify the work involved in
adding support for new XDP actions such as XDP_REDIRECT_XSK. However it was deemed at this point to
be more complex than adding support for the new action to every driver. Should this series be
considered worth pursuing for a proper patch set, the intention would be to update each driver
individually.
Thank you to Magnus Karlsson and Maciej Fijalkowski for several suggestions and insight provided.
TODO:
* Add selftest(s)
* Add support for all copy and zero copy drivers
* Libxdp support
The series applies on commit e5043894b21f ("bpftool: Use libbpf_get_error() to check error")
Thanks,
Ciara
Ciara Loftus (8):
xsk: add struct xdp_sock to netdev_rx_queue
bpf: add bpf_redirect_xsk helper and XDP_REDIRECT_XSK action
xsk: handle XDP_REDIRECT_XSK and expose xsk_rcv/flush
i40e: handle the XDP_REDIRECT_XSK action
xsk: implement a batched version of xsk_rcv
i40e: isolate descriptor processing in separate function
i40e: introduce batched XDP rx descriptor processing
libbpf: use bpf_redirect_xsk in the default program
drivers/net/ethernet/intel/i40e/i40e_txrx.c | 13 +-
.../ethernet/intel/i40e/i40e_txrx_common.h | 1 +
drivers/net/ethernet/intel/i40e/i40e_xsk.c | 285 +++++++++++++++---
include/linux/netdevice.h | 2 +
include/net/xdp_sock_drv.h | 49 +++
include/net/xsk_buff_pool.h | 22 ++
include/uapi/linux/bpf.h | 13 +
kernel/bpf/verifier.c | 7 +-
net/core/dev.c | 14 +
net/core/filter.c | 26 ++
net/xdp/xsk.c | 69 ++++-
net/xdp/xsk_queue.h | 31 ++
tools/include/uapi/linux/bpf.h | 13 +
tools/lib/bpf/xsk.c | 50 ++-
14 files changed, 551 insertions(+), 44 deletions(-)
--
2.17.1
Storing a reference to the XDP socket in the netdev_rx_queue structure
makes a single socket accessible without requiring a lookup in the XSKMAP.
A future commit will introduce the XDP_REDIRECT_XSK action which
indicates to use this reference instead of performing the lookup. Since
an rx ring is required for redirection, only store the reference if an
rx ring is configured.
When multiple sockets exist for a given context (netdev, qid), a
reference is not stored because in this case we fallback to the default
behavior of using the XSKMAP to redirect the packets.
Signed-off-by: Ciara Loftus <redacted>
---
include/linux/netdevice.h | 2 ++
net/xdp/xsk.c | 34 ++++++++++++++++++++++++++++++++++
2 files changed, 36 insertions(+)
@@ -728,6 +728,30 @@ static void xsk_unbind_dev(struct xdp_sock *xs)/* Wait for driver to stop using the xdp socket. */xp_del_xsk(xs->pool,xs);+if(xs->rx){+if(refcount_read(&dev->_rx[xs->queue_id].xsk_refcnt)==1){+refcount_set(&dev->_rx[xs->queue_id].xsk_refcnt,0);+WRITE_ONCE(xs->dev->_rx[xs->queue_id].xsk,NULL);+}else{+refcount_dec(&dev->_rx[xs->queue_id].xsk_refcnt);+/* If the refcnt returns to one again store the reference to the+*remainingsocketinthenetdev_rx_queue.+*/+if(refcount_read(&dev->_rx[xs->queue_id].xsk_refcnt)==1){+structnet*net=dev_net(dev);+structxdp_sock*xsk;+structsock*sk;++mutex_lock(&net->xdp.lock);+sk=sk_head(&net->xdp.list);+xsk=xdp_sk(sk);+mutex_lock(&xsk->mutex);+WRITE_ONCE(xs->dev->_rx[xs->queue_id].xsk,xsk);+mutex_unlock(&xsk->mutex);+mutex_unlock(&net->xdp.lock);+}+}+}xs->dev=NULL;synchronize_net();dev_put(dev);
@@ -972,6 +996,16 @@ static int xsk_bind(struct socket *sock, struct sockaddr *addr, int addr_len)xs->queue_id=qid;xp_add_xsk(xs->pool,xs);+if(xs->rx){+if(refcount_read(&dev->_rx[xs->queue_id].xsk_refcnt)==0){+WRITE_ONCE(dev->_rx[qid].xsk,xs);+refcount_set(&dev->_rx[qid].xsk_refcnt,1);+}else{+refcount_inc(&dev->_rx[qid].xsk_refcnt);+WRITE_ONCE(dev->_rx[qid].xsk,NULL);+}+}+out_unlock:if(err){dev_put(dev);
Add a new XDP redirect helper called bpf_redirect_xsk which simply
returns the new XDP_REDIRECT_XSK action if the xsk refcnt for the
netdev_rx_queue is equal to one. Checking this value verifies that the
AF_XDP socket Rx ring is configured and there is exactly one xsk attached
to the queue.
XDP_REDIRECT_XSK indicates to the driver that the XSKMAP lookup can be
skipped and the pointer to the socket to redirect to can instead be
retrieved from the netdev_rx_queue on which the packet was received.
If the aforementioned conditions are not met, fallback to the behavior of
xdp_redirect_map which returns XDP_REDIRECT for a successful XSKMAP lookup.
Signed-off-by: Ciara Loftus <redacted>
---
include/uapi/linux/bpf.h | 13 +++++++++++++
kernel/bpf/verifier.c | 7 ++++++-
net/core/filter.c | 22 ++++++++++++++++++++++
3 files changed, 41 insertions(+), 1 deletion(-)
@@ -4957,6 +4957,17 @@ union bpf_attr {***-ENOENT**if*task->mm*isNULL,ornovmacontains*addr*.***-EBUSY**iffailedtotrylockmmap_lock.***-EINVAL**forinvalid**flags**.+*+*longbpf_redirect_xsk(void*ctx,structbpf_map*map,u32key,u64flags)+*Description+*RedirectthepackettotheXDPsocketassociatedwiththenetdevqueueif+*thesockethasanrxringconfiguredandistheonlysocketattachedtothe+*queue.Fallbacktobpf_redirect_mapbehaviorifeitherconditionisnotmet.+*Return+***XDP_REDIRECT_XSK**ifsuccessful.+*+***XDP_REDIRECT**ifthefallbackwassuccessful,orthevalueofthe+*twolowerbitsofthe*flags*argumentonerror*/#define __BPF_FUNC_MAPPER(FN) \FN(unspec),\
@@ -5140,6 +5151,7 @@ union bpf_attr {FN(skc_to_unix_sock),\FN(kallsyms_lookup_name),\FN(find_vma),\+FN(redirect_xsk),\/* *//* integer value in 'imm' field of BPF_CALL instruction selects which helper
@@ -5520,6 +5532,7 @@ enum xdp_action {XDP_PASS,XDP_TX,XDP_REDIRECT,+XDP_REDIRECT_XSK,};/* user accessible metadata for XDP packet hook
Handle the XDP_REDIRECT_XSK action on the SKB path by retrieving the
handle to the socket from the netdev_rx_queue struct and immediately
calling the xsk_generic_rcv function. Also, prepare for supporting
this action in the drivers by exposing the xsk_rcv and xsk_flush functions
so they can be used directly in the drivers.
Signed-off-by: Ciara Loftus <redacted>
---
include/net/xdp_sock_drv.h | 21 +++++++++++++++++++++
net/core/dev.c | 14 ++++++++++++++
net/core/filter.c | 4 ++++
net/xdp/xsk.c | 6 ++++--
4 files changed, 43 insertions(+), 2 deletions(-)
If the BPF program returns XDP_REDIRECT_XSK, obtain the pointer to the
socket from the netdev_rx_queue struct and call the newly exposed xsk_rcv
function to push the XDP descriptor to the Rx ring. Then use xsk_flush to
flush the socket.
Signed-off-by: Ciara Loftus <redacted>
---
drivers/net/ethernet/intel/i40e/i40e_txrx.c | 13 +++++++++++-
.../ethernet/intel/i40e/i40e_txrx_common.h | 1 +
drivers/net/ethernet/intel/i40e/i40e_xsk.c | 21 +++++++++++++------
3 files changed, 28 insertions(+), 7 deletions(-)
@@ -144,13 +144,14 @@ int i40e_xsk_pool_setup(struct i40e_vsi *vsi, struct xsk_buff_pool *pool,*@rx_ring:Rxring*@xdp:xdp_buffusedasinputtotheXDPprogram*-*ReturnsanyofI40E_XDP_{PASS,CONSUMED,TX,REDIR}+*ReturnsanyofI40E_XDP_{PASS,CONSUMED,TX,REDIR,REDIR_XSK}**/staticinti40e_run_xdp_zc(structi40e_ring*rx_ring,structxdp_buff*xdp){interr,result=I40E_XDP_PASS;structi40e_ring*xdp_ring;structbpf_prog*xdp_prog;+structxdp_sock*xs;u32act;/* NB! xdp_prog will always be !NULL, due to the fact that
@@ -371,7 +380,7 @@ int i40e_clean_rx_irq_zc(struct i40e_ring *rx_ring, int budget)&rx_bytes,size,xdp_res);total_rx_packets+=rx_packets;total_rx_bytes+=rx_bytes;-xdp_xmit|=xdp_res&(I40E_XDP_TX|I40E_XDP_REDIR);+xdp_xmit|=xdp_res&(I40E_XDP_TX|I40E_XDP_REDIR|I40E_XDP_REDIR_XSK);next_to_clean=(next_to_clean+1)&count_mask;}
Introduce a batched version of xsk_rcv called xsk_rcv_batch which takes
an array of xdp_buffs and pushes them to the Rx ring. Also introduce a
batched version of xsk_buff_dma_sync_for_cpu.
Signed-off-by: Ciara Loftus <redacted>
---
include/net/xdp_sock_drv.h | 28 ++++++++++++++++++++++++++++
include/net/xsk_buff_pool.h | 22 ++++++++++++++++++++++
net/xdp/xsk.c | 29 +++++++++++++++++++++++++++++
net/xdp/xsk_queue.h | 31 +++++++++++++++++++++++++++++++
4 files changed, 110 insertions(+)
@@ -214,6 +214,28 @@ static inline void xp_release(struct xdp_buff_xsk *xskb)xskb->pool->free_heads[xskb->pool->free_heads_cnt++]=xskb;}+/* Release a batch of xdp_buffs back to an xdp_buff_pool.+*Thebatchofbuffsmustallcomefromthesamexdp_buff_pool.Thisway+*itissafetopushthebatchtothetopofthefree_headsstack,because+*atleastthesameamountwillhavebeenpoppedfromthestackearlierin+*thedatapath.+*/+staticinlinevoidxp_release_batch(structxdp_buff**bufs,intbatch_size)+{+structxdp_buff_xsk*xskb=container_of(*bufs,structxdp_buff_xsk,xdp);+structxsk_buff_pool*pool=xskb->pool;+u32tail=pool->free_heads_cnt;+u32i;++if(pool->unaligned){+for(i=0;i<batch_size;i++){+xskb=container_of(*(bufs+i),structxdp_buff_xsk,xdp);+pool->free_heads[tail+i]=xskb;+}+pool->free_heads_cnt+=batch_size;+}+}+staticinlineu64xp_get_handle(structxdp_buff_xsk*xskb){u64offset=xskb->xdp.data-xskb->xdp.data_hard_start;
@@ -399,6 +404,32 @@ static inline int xskq_prod_reserve_desc(struct xsk_queue *q,return0;}+staticinlineintxskq_prod_reserve_desc_batch(structxsk_queue*q,structxdp_buff**bufs,+intbatch_size)+{+structxdp_rxtx_ring*ring=(structxdp_rxtx_ring*)q->ring;+structxdp_buff_xsk*xskb;+u64addr;+u32len;+u32i;++if(xskq_prod_is_full_n(q,batch_size))+return-ENOSPC;++/* A, matches D */+for(i=0;i<batch_size;i++){+len=(*(bufs+i))->data_end-(*(bufs+i))->data;+xskb=container_of(*(bufs+i),structxdp_buff_xsk,xdp);+addr=xp_get_handle(xskb);+ring->desc[(q->cached_prod+i)&q->ring_mask].addr=addr;+ring->desc[(q->cached_prod+i)&q->ring_mask].len=len;+}++q->cached_prod+=batch_size;++return0;+}+staticinlinevoid__xskq_prod_submit(structxsk_queue*q,u32idx){smp_store_release(&q->ring->producer,idx);/* B, matches C */
To prepare for batched processing, first isolate descriptor processing
in a separate function to make it easier to introduce the batched
interfaces.
Signed-off-by: Cristian Dumitrescu <redacted>
Signed-off-by: Ciara Loftus <redacted>
---
drivers/net/ethernet/intel/i40e/i40e_xsk.c | 51 +++++++++++++++-------
1 file changed, 36 insertions(+), 15 deletions(-)
@@ -385,7 +380,33 @@ int i40e_clean_rx_irq_zc(struct i40e_ring *rx_ring, int budget)}rx_ring->next_to_clean=next_to_clean;-cleaned_count=(next_to_clean-rx_ring->next_to_use-1)&count_mask;+*stat_rx_packets=total_rx_packets;+*stat_rx_bytes=total_rx_bytes;+*xmit=xdp_xmit;+}++/**+*i40e_clean_rx_irq_zc-ConsumesRxpacketsfromthehardwarering+*@rx_ring:Rxring+*@budget:NAPIbudget+*+*Returnsamountofworkcompleted+**/+inti40e_clean_rx_irq_zc(structi40e_ring*rx_ring,intbudget)+{+unsignedinttotal_rx_bytes=0,total_rx_packets=0;+u16count_mask=rx_ring->count-1;+unsignedintxdp_xmit=0;+boolfailure=false;+u16cleaned_count;++i40e_clean_rx_desc_zc(rx_ring,+&total_rx_packets,+&total_rx_bytes,+&xdp_xmit,+budget);++cleaned_count=(rx_ring->next_to_clean-rx_ring->next_to_use-1)&count_mask;if(cleaned_count>=I40E_RX_BUFFER_WRITE)failure=!i40e_alloc_rx_buffers_zc(rx_ring,cleaned_count);
@@ -394,7 +415,7 @@ int i40e_clean_rx_irq_zc(struct i40e_ring *rx_ring, int budget)i40e_update_rx_stats(rx_ring,total_rx_bytes,total_rx_packets);if(xsk_uses_need_wakeup(rx_ring->xsk_pool)){-if(failure||next_to_clean==rx_ring->next_to_use)+if(failure||rx_ring->next_to_clean==rx_ring->next_to_use)xsk_set_rx_need_wakeup(rx_ring->xsk_pool);elsexsk_clear_rx_need_wakeup(rx_ring->xsk_pool);
Introduce batched processing of XDP frames in the i40e driver. The batch
size is fixed at 64.
First, the driver performs a lookahead in the rx ring to determine if
there are 64 contiguous descriptors available to be processed. If so,
and if the action returned from the bpf program run for each of the 64
descriptors is XDP_REDIRECT_XSK, the new xsk_rcv_batch API is used to
push the batch to the XDP socket Rx ring.
Logic to fallback to scalar processing is included for situations where
batch processing is not possible eg. not enough descriptors, ring wrap,
different actions returned from bpf program, etc.
Signed-off-by: Ciara Loftus <redacted>
Signed-off-by: Cristian Dumitrescu <redacted>
---
drivers/net/ethernet/intel/i40e/i40e_xsk.c | 219 +++++++++++++++++++--
1 file changed, 200 insertions(+), 19 deletions(-)
@@ -139,26 +142,12 @@ int i40e_xsk_pool_setup(struct i40e_vsi *vsi, struct xsk_buff_pool *pool,i40e_xsk_pool_disable(vsi,qid);}-/**-*i40e_run_xdp_zc-ExecutesanXDPprogramonanxdp_buff-*@rx_ring:Rxring-*@xdp:xdp_buffusedasinputtotheXDPprogram-*-*ReturnsanyofI40E_XDP_{PASS,CONSUMED,TX,REDIR,REDIR_XSK}-**/-staticinti40e_run_xdp_zc(structi40e_ring*rx_ring,structxdp_buff*xdp)+staticinti40e_handle_xdp_action(structi40e_ring*rx_ring,structxdp_buff*xdp,+structbpf_prog*xdp_prog,u32act){interr,result=I40E_XDP_PASS;structi40e_ring*xdp_ring;-structbpf_prog*xdp_prog;structxdp_sock*xs;-u32act;--/* NB! xdp_prog will always be !NULL, due to the fact that-*thispathisenabledbysettinganXDPprogram.-*/-xdp_prog=READ_ONCE(rx_ring->xdp_prog);-act=bpf_prog_run_xdp(xdp_prog,xdp);if(likely(act==XDP_REDIRECT_XSK)){xs=xsk_get_redirect_xsk(&rx_ring->netdev->_rx[xdp->rxq->queue_index]);
@@ -385,6 +391,172 @@ static inline void i40e_clean_rx_desc_zc(struct i40e_ring *rx_ring,*xmit=xdp_xmit;}+/**+*i40_rx_ring_lookahead-checkfornewdescriptorsintherxring+*@rx_ring:Rxring+*@budget:NAPIbudget+*+*Returnsthenumberofavailabledescriptorsincontiguousmemoryie.+*withoutaringwrap.+*+**/+staticinlineunsignedinti40_rx_ring_lookahead(structi40e_ring*rx_ring,+unsignedintbudget)+{+u32used=(rx_ring->next_to_clean-rx_ring->next_to_use-1)&(rx_ring->count-1);+unioni40e_rx_desc*rx_desc0=(unioni40e_rx_desc*)rx_ring->desc,*rx_desc;+u32next_to_clean=rx_ring->next_to_clean;+u32potential=rx_ring->count-used;+u16count_mask=rx_ring->count-1;+unsignedintsize;+u64qword;++budget&=I40E_XSK_BATCH_MASK;++while(budget){+if(budget>potential)+gotonext;+rx_desc=rx_desc0+((next_to_clean+budget-1)&count_mask);+qword=le64_to_cpu(rx_desc->wb.qword1.status_error_len);+dma_rmb();++size=(qword&I40E_RXD_QW1_LENGTH_PBUF_MASK)>>+I40E_RXD_QW1_LENGTH_PBUF_SHIFT;+if(size&&((next_to_clean+budget)<=count_mask))+returnbudget;++next:+budget>>=1;+budget&=I40E_XSK_BATCH_MASK;+}++return0;+}++/**+*i40e_run_xdp_zc_batch-ExecutesanXDPprogramonanarrayofxdp_buffs+*@rx_ring:Rxring+*@bufs:arrayofxdp_buffsusedasinputtotheXDPprogram+*@res:arrayofintswithresultforeachbufifanerroroccursorslowpathtaken.+*+*Returnszeroifallxdp_buffssuccessfullytookthefastpath(XDP_REDIRECT_XSK).+*Otherwisereturns-1andsetsindividualresultsforeachbufinthearray*res.+*IndividualresultsareoneofI40E_XDP_{PASS,CONSUMED,TX,REDIR,REDIR_XSK}+**/+staticinti40e_run_xdp_zc_batch(structi40e_ring*rx_ring,structxdp_buff**bufs,+structbpf_prog*xdp_prog,int*res)+{+u32last_act=XDP_REDIRECT_XSK;+intruns=0,ret=0,err,i;++while((runs<I40E_DESCS_PER_BATCH)&&(last_act==XDP_REDIRECT_XSK))+last_act=bpf_prog_run_xdp(xdp_prog,*(bufs+runs++));++if(likely(runs==I40E_DESCS_PER_BATCH)){+structxdp_sock*xs=+xsk_get_redirect_xsk(&rx_ring->netdev->_rx[(*bufs)->rxq->queue_index]);++err=xsk_rcv_batch(xs,bufs,I40E_DESCS_PER_BATCH);+if(unlikely(err)){+ret=-1;+for(i=0;i<I40E_DESCS_PER_BATCH;i++)+*(res+i)=I40E_XDP_PASS;+}+}else{+/* Handle the result of each program run individually */+u32act;++ret=-1;+for(i=0;i<I40E_DESCS_PER_BATCH;i++){+structxdp_buff*xdp=*(bufs+i);++/* The result of the first runs-2 programs was XDP_REDIRECT_XSK.+*Theresultofthesubsequentprogramrunwaslast_act.+*Anyremainingbufshavenotyethadtheprogramexecuted,so+*executeitnow.+*/++if(i<runs-2)+act=XDP_REDIRECT_XSK;+elseif(i==runs-1)+act=last_act;+else+act=bpf_prog_run_xdp(xdp_prog,xdp);++*(res+i)=i40e_handle_xdp_action(rx_ring,xdp,xdp_prog,act);+}+}++returnret;+}++staticinlinevoidi40e_clean_rx_desc_zc_batch(structi40e_ring*rx_ring,+structbpf_prog*xdp_prog,+unsignedint*total_rx_packets,+unsignedint*total_rx_bytes,+unsignedint*xdp_xmit)+{+u16next_to_clean=rx_ring->next_to_clean;+unsignedintxdp_res[I40E_DESCS_PER_BATCH];+unsignedintsize[I40E_DESCS_PER_BATCH];+unsignedintrx_packets,rx_bytes=0;+unioni40e_rx_desc*rx_desc;+structxdp_buff**bufs;+intj,ret;+u64qword;++rx_desc=I40E_RX_DESC(rx_ring,next_to_clean);++prefetch(rx_desc+I40E_DESCS_PER_BATCH);++for(j=0;j<I40E_DESCS_PER_BATCH;j++){+qword=le64_to_cpu((rx_desc+j)->wb.qword1.status_error_len);+size[j]=(qword&I40E_RXD_QW1_LENGTH_PBUF_MASK)>>+I40E_RXD_QW1_LENGTH_PBUF_SHIFT;+}++/* This memory barrier is needed to keep us from reading+*anyotherfieldsoutoftherx_descsuntilwehave+*verifiedthedescriptorshavebeenwrittenback.+*/+dma_rmb();++bufs=i40e_rx_bi(rx_ring,next_to_clean);++for(j=0;j<I40E_DESCS_PER_BATCH;j++)+xsk_buff_set_size(*(bufs+j),size[j]);++xsk_buff_dma_sync_for_cpu_batch(bufs,rx_ring->xsk_pool,I40E_DESCS_PER_BATCH);++ret=i40e_run_xdp_zc_batch(rx_ring,bufs,xdp_prog,xdp_res);++if(unlikely(ret)){+unsignedinterr_rx_packets=0,err_rx_bytes=0;++rx_packets=0;+rx_bytes=0;++for(j=0;j<I40E_DESCS_PER_BATCH;j++){+i40e_handle_xdp_result_zc(rx_ring,*(bufs+j),rx_desc+j,+&err_rx_packets,&err_rx_bytes,size[j],+xdp_res[j]);+*xdp_xmit|=(xdp_res[j]&(I40E_XDP_TX|I40E_XDP_REDIR|+I40E_XDP_REDIR_XSK));+rx_packets+=err_rx_packets;+rx_bytes+=err_rx_bytes;+}+}else{+rx_packets=I40E_DESCS_PER_BATCH;+for(j=0;j<I40E_DESCS_PER_BATCH;j++)+rx_bytes+=size[j];+*xdp_xmit|=I40E_XDP_REDIR_XSK;+}++rx_ring->next_to_clean+=I40E_DESCS_PER_BATCH;+*total_rx_packets+=rx_packets;+*total_rx_bytes+=rx_bytes;+}+/***i40e_clean_rx_irq_zc-ConsumesRxpacketsfromthehardwarering*@rx_ring:Rxring
NOTE: This will be committed to libxdp, not libbpf as the xsk support in
that library has been deprecated. It is only here to serve as an example
of what will be added into libxdp.
Use the new bpf_redirect_xsk helper in the default program if the kernel
supports it.
Signed-off-by: Ciara Loftus <redacted>
---
tools/include/uapi/linux/bpf.h | 13 +++++++++
tools/lib/bpf/xsk.c | 50 ++++++++++++++++++++++++++++++++--
2 files changed, 60 insertions(+), 3 deletions(-)
@@ -4957,6 +4957,17 @@ union bpf_attr {***-ENOENT**if*task->mm*isNULL,ornovmacontains*addr*.***-EBUSY**iffailedtotrylockmmap_lock.***-EINVAL**forinvalid**flags**.+*+*longbpf_redirect_xsk(void*ctx,structbpf_map*map,u32key,u64flags)+*Description+*RedirectthepackettotheXDPsocketassociatedwiththenetdevqueueif+*thesockethasanrxringconfiguredandistheonlysocketattachedtothe+*queue.Fallbacktobpf_redirect_mapbehaviorifeitherconditionisnotmet.+*Return+***XDP_REDIRECT_XSK**ifsuccessful.+*+***XDP_REDIRECT**ifthefallbackwassuccessful,orthevalueofthe+*twolowerbitsofthe*flags*argumentonerror*/#define __BPF_FUNC_MAPPER(FN) \FN(unspec),\
@@ -5140,6 +5151,7 @@ union bpf_attr {FN(skc_to_unix_sock),\FN(kallsyms_lookup_name),\FN(find_vma),\+FN(redirect_xsk),\/* *//* integer value in 'imm' field of BPF_CALL instruction selects which helper
@@ -5520,6 +5532,7 @@ enum xdp_action {XDP_PASS,XDP_TX,XDP_REDIRECT,+XDP_REDIRECT_XSK,};/* user accessible metadata for XDP packet hook
The common case for AF_XDP sockets (xsks) is creating a single xsk on a queue for sending and
receiving frames as this is analogous to HW packet steering through RSS and other classification
methods in the NIC. AF_XDP uses the xdp redirect infrastructure to direct packets to the socket. It
was designed for the much more complicated case of DEVMAP xdp_redirects which directs traffic to
another netdev and thus potentially another driver. In the xsk redirect case, by skipping the
unnecessary parts of this common code we can significantly improve performance and pave the way
for batching in the driver. This RFC proposes one such way to simplify the infrastructure which
yields a 27% increase in throughput and a decrease in cycles per packet of 24 cycles [1]. The goal
of this RFC is to start a discussion on how best to simplify the single-socket datapath while
providing one method as an example.
Current approach:
1. XSK pointer: an xsk is created and a handle to the xsk is stored in the XSKMAP.
2. XDP program: bpf_redirect_map helper triggers the XSKMAP lookup which stores the result (handle
to the xsk) and the map type (XSKMAP) in the percpu bpf_redirect_info struct. The XDP_REDIRECT
action is returned.
3. XDP_REDIRECT handling called by the driver: the map type (XSKMAP) is read from the
bpf_redirect_info which selects the xsk_map_redirect path. The xsk pointer is retrieved from the
bpf_redirect_info and the XDP descriptor is pushed to the xsk's Rx ring. The socket is added to a
list for flushing later.
4. xdp_do_flush: iterate through the lists of all maps that can be used for redirect (CPUMAP,
DEVMAP and XSKMAP). When XSKMAP is flushed, go through all xsks that had any traffic redirected to
them and bump the Rx ring head pointer(s).
For the end goal of submitting the descriptor to the Rx ring and bumping the head pointer of that
ring, only some of these steps are needed. The rest is overhead. The bpf_redirect_map
infrastructure is needed for all other redirect operations, but is not necessary when redirecting
to a single AF_XDP socket. And similarly, flushing the list for every map type in step 4 is not
necessary when only one socket needs to be flushed.
Proposed approach:
1. XSK pointer: an xsk is created and a handle to the xsk is stored both in the XSKMAP and also the
netdev_rx_queue struct.
2. XDP program: new bpf_redirect_xsk helper returns XDP_REDIRECT_XSK.
3. XDP_REDIRECT_XSK handling called by the driver: the xsk pointer is retrieved from the
netdev_rx_queue struct and the XDP descriptor is pushed to the xsk's Rx ring.
4. xsk_flush: fetch the handle from the netdev_rx_queue and flush the xsk.
This fast path is triggered on XDP_REDIRECT_XSK if:
(i) AF_XDP socket SW Rx ring configured
(ii) Exactly one xsk attached to the queue
If any of these conditions are not met, fall back to the same behavior as the original approach:
xdp_redirect_map. This is handled under-the-hood in the new bpf_xdp_redirect_xsk helper so the user
does not need to be aware of these conditions.
Batching:
With this new approach it is possible to optimize the driver by submitting a batch of descriptors
to the Rx ring in Step 3 of the new approach by simply verifying that the action returned from
every program run of each packet in a batch equals XDP_REDIRECT_XSK. That's because with this
action we know the socket to redirect to will be the same for each packet in the batch. This is
not possible with XDP_REDIRECT because the xsk pointer is stored in the bpf_redirect_info and not
guaranteed to be the same for every packet in a batch.
[1] Performance:
The benchmarks were performed on VM running a 2.4GHz Ice Lake host with an i40e device passed
through. The xdpsock app was run on a single core with busy polling and configured in 'rxonly' mode.
./xdpsock -i <iface> -r -B
The improvement in throughput when using the new bpf helper and XDP action was measured at ~13% for
scalar processing, with reduction in cycles per packet of ~13. A further ~14% improvement in
throughput and reduction of ~11 cycles per packet was measured when the batched i40e driver path
was used, for a total improvement of ~27% in throughput and reduction of ~24 cycles per packet.
Other approaches considered:
Two other approaches were considered. The advantage of both being that neither involved introducing
a new XDP action. The first alternative approach considered was to create a new map type
BPF_MAP_TYPE_XSKMAP_DIRECT. When the XDP_REDIRECT action was returned, this map type could be
checked and used as an indicator to skip the map lookup and use the netdev_rx_queue xsk instead.
The second approach considered was similar and involved using a new bpf_redirect_info flag which
could be used in a similar fashion.
While both approaches yielded a performance improvement they were measured at about half of what
was measured for the approach outlined in this RFC. It seems using bpf_redirect_info is too
expensive.
I think it was Bjørn that discovered that accessing the per CPU
bpf_redirect_info struct have an overhead of approx 2 ns (times 2.4GHz
~4.8 cycles). Your reduction in cycles per packet was ~13, where ~4.8
seem to be large.
The code access this_cpu_ptr(&bpf_redirect_info) two times.
One time in the BPF-helper redirect call and second in xdp_do_redirect.
(Hint xdp_redirect_map end-up calling __bpf_xdp_redirect_map)
Thus, it seems (as you say), the bpf_redirect_info approach is too
expensive. May be should look at storing bpf_redirect_info in a place
that doesn't requires the this_cpu_ptr() lookup... or cache the lookup
per NAPI cycle.
Have you tried this?
Also, centralised processing of XDP actions was investigated. This would involve porting all drivers
to a common interface for handling XDP actions which would greatly simplify the work involved in
adding support for new XDP actions such as XDP_REDIRECT_XSK. However it was deemed at this point to
be more complex than adding support for the new action to every driver. Should this series be
considered worth pursuing for a proper patch set, the intention would be to update each driver
individually.
I'm fine with adding a new helper, but I don't like introducing a new
XDP_REDIRECT_XSK action, which requires updating ALL the drivers.
With XDP_REDIRECT infra we beleived we didn't need to add more
XDP-action code to drivers, as we multiplex/add new features by
extending the bpf_redirect_info.
In this extreme performance case, it seems the this_cpu_ptr "lookup" of
bpf_redirect_info is the performance issue itself.
Could you experiement with different approaches that modify
xdp_do_redirect() to handle if new helper bpf_redirect_xsk was called,
prior to this_cpu_ptr() call.
(Thus, avoiding to introduce a new XDP-action).
Thank you to Magnus Karlsson and Maciej Fijalkowski for several suggestions and insight provided.
TODO:
* Add selftest(s)
* Add support for all copy and zero copy drivers
* Libxdp support
The series applies on commit e5043894b21f ("bpftool: Use libbpf_get_error() to check error")
Thanks,
Ciara
Ciara Loftus (8):
xsk: add struct xdp_sock to netdev_rx_queue
bpf: add bpf_redirect_xsk helper and XDP_REDIRECT_XSK action
xsk: handle XDP_REDIRECT_XSK and expose xsk_rcv/flush
i40e: handle the XDP_REDIRECT_XSK action
xsk: implement a batched version of xsk_rcv
i40e: isolate descriptor processing in separate function
i40e: introduce batched XDP rx descriptor processing
libbpf: use bpf_redirect_xsk in the default program
drivers/net/ethernet/intel/i40e/i40e_txrx.c | 13 +-
.../ethernet/intel/i40e/i40e_txrx_common.h | 1 +
drivers/net/ethernet/intel/i40e/i40e_xsk.c | 285 +++++++++++++++---
include/linux/netdevice.h | 2 +
include/net/xdp_sock_drv.h | 49 +++
include/net/xsk_buff_pool.h | 22 ++
include/uapi/linux/bpf.h | 13 +
kernel/bpf/verifier.c | 7 +-
net/core/dev.c | 14 +
net/core/filter.c | 26 ++
net/xdp/xsk.c | 69 ++++-
net/xdp/xsk_queue.h | 31 ++
tools/include/uapi/linux/bpf.h | 13 +
tools/lib/bpf/xsk.c | 50 ++-
14 files changed, 551 insertions(+), 44 deletions(-)
The common case for AF_XDP sockets (xsks) is creating a single xsk on a
queue for sending and
quoted
receiving frames as this is analogous to HW packet steering through RSS
and other classification
quoted
methods in the NIC. AF_XDP uses the xdp redirect infrastructure to direct
packets to the socket. It
quoted
was designed for the much more complicated case of DEVMAP
xdp_redirects which directs traffic to
quoted
another netdev and thus potentially another driver. In the xsk redirect
case, by skipping the
quoted
unnecessary parts of this common code we can significantly improve
performance and pave the way
quoted
for batching in the driver. This RFC proposes one such way to simplify the
infrastructure which
quoted
yields a 27% increase in throughput and a decrease in cycles per packet of
24 cycles [1]. The goal
quoted
of this RFC is to start a discussion on how best to simplify the single-socket
datapath while
quoted
providing one method as an example.
Current approach:
1. XSK pointer: an xsk is created and a handle to the xsk is stored in the
XSKMAP.
quoted
2. XDP program: bpf_redirect_map helper triggers the XSKMAP lookup
which stores the result (handle
quoted
to the xsk) and the map type (XSKMAP) in the percpu bpf_redirect_info
struct. The XDP_REDIRECT
quoted
action is returned.
3. XDP_REDIRECT handling called by the driver: the map type (XSKMAP) is
read from the
quoted
bpf_redirect_info which selects the xsk_map_redirect path. The xsk
pointer is retrieved from the
quoted
bpf_redirect_info and the XDP descriptor is pushed to the xsk's Rx ring. The
socket is added to a
quoted
list for flushing later.
4. xdp_do_flush: iterate through the lists of all maps that can be used for
redirect (CPUMAP,
quoted
DEVMAP and XSKMAP). When XSKMAP is flushed, go through all xsks that
had any traffic redirected to
quoted
them and bump the Rx ring head pointer(s).
For the end goal of submitting the descriptor to the Rx ring and bumping
the head pointer of that
quoted
ring, only some of these steps are needed. The rest is overhead. The
bpf_redirect_map
quoted
infrastructure is needed for all other redirect operations, but is not
necessary when redirecting
quoted
to a single AF_XDP socket. And similarly, flushing the list for every map type
in step 4 is not
quoted
necessary when only one socket needs to be flushed.
Proposed approach:
1. XSK pointer: an xsk is created and a handle to the xsk is stored both in
the XSKMAP and also the
quoted
netdev_rx_queue struct.
2. XDP program: new bpf_redirect_xsk helper returns XDP_REDIRECT_XSK.
3. XDP_REDIRECT_XSK handling called by the driver: the xsk pointer is
retrieved from the
quoted
netdev_rx_queue struct and the XDP descriptor is pushed to the xsk's Rx
ring.
quoted
4. xsk_flush: fetch the handle from the netdev_rx_queue and flush the
xsk.
quoted
This fast path is triggered on XDP_REDIRECT_XSK if:
(i) AF_XDP socket SW Rx ring configured
(ii) Exactly one xsk attached to the queue
If any of these conditions are not met, fall back to the same behavior as the
original approach:
quoted
xdp_redirect_map. This is handled under-the-hood in the new
bpf_xdp_redirect_xsk helper so the user
quoted
does not need to be aware of these conditions.
Batching:
With this new approach it is possible to optimize the driver by submitting a
batch of descriptors
quoted
to the Rx ring in Step 3 of the new approach by simply verifying that the
action returned from
quoted
every program run of each packet in a batch equals XDP_REDIRECT_XSK.
That's because with this
quoted
action we know the socket to redirect to will be the same for each packet in
the batch. This is
quoted
not possible with XDP_REDIRECT because the xsk pointer is stored in the
bpf_redirect_info and not
quoted
guaranteed to be the same for every packet in a batch.
[1] Performance:
The benchmarks were performed on VM running a 2.4GHz Ice Lake host
with an i40e device passed
quoted
through. The xdpsock app was run on a single core with busy polling and
configured in 'rxonly' mode.
quoted
./xdpsock -i <iface> -r -B
The improvement in throughput when using the new bpf helper and XDP
action was measured at ~13% for
quoted
scalar processing, with reduction in cycles per packet of ~13. A further ~14%
improvement in
quoted
throughput and reduction of ~11 cycles per packet was measured when
the batched i40e driver path
quoted
was used, for a total improvement of ~27% in throughput and reduction of
~24 cycles per packet.
quoted
Other approaches considered:
Two other approaches were considered. The advantage of both being that
neither involved introducing
quoted
a new XDP action. The first alternative approach considered was to create a
new map type
quoted
BPF_MAP_TYPE_XSKMAP_DIRECT. When the XDP_REDIRECT action was
returned, this map type could be
quoted
checked and used as an indicator to skip the map lookup and use the
netdev_rx_queue xsk instead.
quoted
The second approach considered was similar and involved using a new
bpf_redirect_info flag which
quoted
could be used in a similar fashion.
While both approaches yielded a performance improvement they were
measured at about half of what
quoted
was measured for the approach outlined in this RFC. It seems using
bpf_redirect_info is too
quoted
expensive.
I think it was Bjørn that discovered that accessing the per CPU
bpf_redirect_info struct have an overhead of approx 2 ns (times 2.4GHz
~4.8 cycles). Your reduction in cycles per packet was ~13, where ~4.8
seem to be large.
The code access this_cpu_ptr(&bpf_redirect_info) two times.
One time in the BPF-helper redirect call and second in xdp_do_redirect.
(Hint xdp_redirect_map end-up calling __bpf_xdp_redirect_map)
Thus, it seems (as you say), the bpf_redirect_info approach is too
expensive. May be should look at storing bpf_redirect_info in a place
that doesn't requires the this_cpu_ptr() lookup... or cache the lookup
per NAPI cycle.
Have you tried this?
quoted
Also, centralised processing of XDP actions was investigated. This would
involve porting all drivers
quoted
to a common interface for handling XDP actions which would greatly
simplify the work involved in
quoted
adding support for new XDP actions such as XDP_REDIRECT_XSK. However
it was deemed at this point to
quoted
be more complex than adding support for the new action to every driver.
Should this series be
quoted
considered worth pursuing for a proper patch set, the intention would be
to update each driver
quoted
individually.
I'm fine with adding a new helper, but I don't like introducing a new
XDP_REDIRECT_XSK action, which requires updating ALL the drivers.
With XDP_REDIRECT infra we beleived we didn't need to add more
XDP-action code to drivers, as we multiplex/add new features by
extending the bpf_redirect_info.
In this extreme performance case, it seems the this_cpu_ptr "lookup" of
bpf_redirect_info is the performance issue itself.
Could you experiement with different approaches that modify
xdp_do_redirect() to handle if new helper bpf_redirect_xsk was called,
prior to this_cpu_ptr() call.
(Thus, avoiding to introduce a new XDP-action).
Thanks for your feedback Jesper.
I understand the hesitation of adding a new action. If we can achieve the same improvement without
introducing a new action I would be very happy!
Without new the action we'll need a new way to indicate that the bpf_redirect_xsk helper was
called. Maybe another new field in the netdev alongside the xsk_refcnt. Or else extend
bpf_redirect_info - if we find a new home for it that it's too costly to access.
Thanks for your suggestions. I'll experiment as you suggested and report back.
Thanks,
Ciara
quoted
Thank you to Magnus Karlsson and Maciej Fijalkowski for several
suggestions and insight provided.
quoted
TODO:
* Add selftest(s)
* Add support for all copy and zero copy drivers
* Libxdp support
The series applies on commit e5043894b21f ("bpftool: Use
libbpf_get_error() to check error")
quoted
Thanks,
Ciara
Ciara Loftus (8):
xsk: add struct xdp_sock to netdev_rx_queue
bpf: add bpf_redirect_xsk helper and XDP_REDIRECT_XSK action
xsk: handle XDP_REDIRECT_XSK and expose xsk_rcv/flush
i40e: handle the XDP_REDIRECT_XSK action
xsk: implement a batched version of xsk_rcv
i40e: isolate descriptor processing in separate function
i40e: introduce batched XDP rx descriptor processing
libbpf: use bpf_redirect_xsk in the default program
drivers/net/ethernet/intel/i40e/i40e_txrx.c | 13 +-
.../ethernet/intel/i40e/i40e_txrx_common.h | 1 +
drivers/net/ethernet/intel/i40e/i40e_xsk.c | 285 +++++++++++++++---
include/linux/netdevice.h | 2 +
include/net/xdp_sock_drv.h | 49 +++
include/net/xsk_buff_pool.h | 22 ++
include/uapi/linux/bpf.h | 13 +
kernel/bpf/verifier.c | 7 +-
net/core/dev.c | 14 +
net/core/filter.c | 26 ++
net/xdp/xsk.c | 69 ++++-
net/xdp/xsk_queue.h | 31 ++
tools/include/uapi/linux/bpf.h | 13 +
tools/lib/bpf/xsk.c | 50 ++-
14 files changed, 551 insertions(+), 44 deletions(-)
I'm fine with adding a new helper, but I don't like introducing a new
XDP_REDIRECT_XSK action, which requires updating ALL the drivers.
With XDP_REDIRECT infra we beleived we didn't need to add more
XDP-action code to drivers, as we multiplex/add new features by
extending the bpf_redirect_info.
In this extreme performance case, it seems the this_cpu_ptr "lookup" of
bpf_redirect_info is the performance issue itself.
Could you experiement with different approaches that modify
xdp_do_redirect() to handle if new helper bpf_redirect_xsk was called,
prior to this_cpu_ptr() call.
(Thus, avoiding to introduce a new XDP-action).
Thanks for your feedback Jesper.
I understand the hesitation of adding a new action. If we can achieve the same improvement without
introducing a new action I would be very happy!
Without new the action we'll need a new way to indicate that the bpf_redirect_xsk helper was
called. Maybe another new field in the netdev alongside the xsk_refcnt. Or else extend
bpf_redirect_info - if we find a new home for it that it's too costly to access.
Thanks for your suggestions. I'll experiment as you suggested and
report back.
I'll add a +1 to the "let's try to solve this without a new return code" :)
Also, I don't think we need a new helper either; the bpf_redirect()
helper takes a flags argument, so we could just use ifindex=0,
flags=DEV_XSK or something like that.
Also, I think the batching in the driver idea can be generalised: we
just need to generalise the idea of "are all these packets going to the
same place" and have a batched version of xdp_do_redirect(), no? The
other map types do batching internally already, though, so I'm wondering
why batching in the driver helps XSK?
-Toke
On Tue, Nov 16, 2021 at 07:37:34AM +0000, Ciara Loftus wrote:
The common case for AF_XDP sockets (xsks) is creating a single xsk on a queue for sending and
receiving frames as this is analogous to HW packet steering through RSS and other classification
methods in the NIC. AF_XDP uses the xdp redirect infrastructure to direct packets to the socket. It
was designed for the much more complicated case of DEVMAP xdp_redirects which directs traffic to
another netdev and thus potentially another driver. In the xsk redirect case, by skipping the
unnecessary parts of this common code we can significantly improve performance and pave the way
for batching in the driver. This RFC proposes one such way to simplify the infrastructure which
yields a 27% increase in throughput and a decrease in cycles per packet of 24 cycles [1]. The goal
of this RFC is to start a discussion on how best to simplify the single-socket datapath while
providing one method as an example.
Current approach:
1. XSK pointer: an xsk is created and a handle to the xsk is stored in the XSKMAP.
2. XDP program: bpf_redirect_map helper triggers the XSKMAP lookup which stores the result (handle
to the xsk) and the map type (XSKMAP) in the percpu bpf_redirect_info struct. The XDP_REDIRECT
action is returned.
3. XDP_REDIRECT handling called by the driver: the map type (XSKMAP) is read from the
bpf_redirect_info which selects the xsk_map_redirect path. The xsk pointer is retrieved from the
bpf_redirect_info and the XDP descriptor is pushed to the xsk's Rx ring. The socket is added to a
list for flushing later.
4. xdp_do_flush: iterate through the lists of all maps that can be used for redirect (CPUMAP,
DEVMAP and XSKMAP). When XSKMAP is flushed, go through all xsks that had any traffic redirected to
them and bump the Rx ring head pointer(s).
For the end goal of submitting the descriptor to the Rx ring and bumping the head pointer of that
ring, only some of these steps are needed. The rest is overhead. The bpf_redirect_map
infrastructure is needed for all other redirect operations, but is not necessary when redirecting
to a single AF_XDP socket. And similarly, flushing the list for every map type in step 4 is not
necessary when only one socket needs to be flushed.
Proposed approach:
1. XSK pointer: an xsk is created and a handle to the xsk is stored both in the XSKMAP and also the
netdev_rx_queue struct.
2. XDP program: new bpf_redirect_xsk helper returns XDP_REDIRECT_XSK.
3. XDP_REDIRECT_XSK handling called by the driver: the xsk pointer is retrieved from the
netdev_rx_queue struct and the XDP descriptor is pushed to the xsk's Rx ring.
4. xsk_flush: fetch the handle from the netdev_rx_queue and flush the xsk.
This fast path is triggered on XDP_REDIRECT_XSK if:
(i) AF_XDP socket SW Rx ring configured
(ii) Exactly one xsk attached to the queue
If any of these conditions are not met, fall back to the same behavior as the original approach:
xdp_redirect_map. This is handled under-the-hood in the new bpf_xdp_redirect_xsk helper so the user
does not need to be aware of these conditions.
I don't think the micro optimization for specific use case warrants addition of new apis.
Please optimize it without adding new actions and new helpers.
I'm fine with adding a new helper, but I don't like introducing a new
XDP_REDIRECT_XSK action, which requires updating ALL the drivers.
With XDP_REDIRECT infra we beleived we didn't need to add more
XDP-action code to drivers, as we multiplex/add new features by
extending the bpf_redirect_info.
In this extreme performance case, it seems the this_cpu_ptr "lookup" of
bpf_redirect_info is the performance issue itself.
Could you experiement with different approaches that modify
xdp_do_redirect() to handle if new helper bpf_redirect_xsk was called,
prior to this_cpu_ptr() call.
(Thus, avoiding to introduce a new XDP-action).
Thanks for your feedback Jesper.
I understand the hesitation of adding a new action. If we can achieve the
same improvement without
quoted
introducing a new action I would be very happy!
Without new the action we'll need a new way to indicate that the
bpf_redirect_xsk helper was
quoted
called. Maybe another new field in the netdev alongside the xsk_refcnt. Or
else extend
quoted
bpf_redirect_info - if we find a new home for it that it's too costly to
access.
quoted
Thanks for your suggestions. I'll experiment as you suggested and
report back.
I'll add a +1 to the "let's try to solve this without a new return code" :)
Also, I don't think we need a new helper either; the bpf_redirect()
helper takes a flags argument, so we could just use ifindex=0,
flags=DEV_XSK or something like that.
The advantage of a new helper is that we can access the netdev
struct from it and check if there's a valid xsk stored in it, before
returning XDP_REDIRECT without the xskmap lookup. However,
I think your suggestion could work too. We would just
have to delay the check until xdp_do_redirect. At this point
though, if there isn't a valid xsk we might have to drop the packet
instead of falling back to the xskmap.
Also, I think the batching in the driver idea can be generalised: we
just need to generalise the idea of "are all these packets going to the
same place" and have a batched version of xdp_do_redirect(), no? The
other map types do batching internally already, though, so I'm wondering
why batching in the driver helps XSK?
With the current infrastructure figuring out if "all the packets are going
to the same place" looks like an expensive operation which could undo
the benefits of the batching that would come after it. We would need
to run the program N=batch_size times, store the actions and
bpf_redirect_info for each run and perform a series of compares. The new
action really helped here because it could easily indicate if all the
packets in a batch were going to the same place. But I understand it's
not an option. Maybe if we can mitigate the cost of accessing the
bpf_redirect_info as Jesper suggested, we can use a flag in it to signal
what the new action was signalling.
I'm not familiar with how the other map types and how they handle
batching so I will look into that.
Appreciate your feedback. I have a few avenues to explore.
Ciara
I'm fine with adding a new helper, but I don't like introducing a new
XDP_REDIRECT_XSK action, which requires updating ALL the drivers.
With XDP_REDIRECT infra we beleived we didn't need to add more
XDP-action code to drivers, as we multiplex/add new features by
extending the bpf_redirect_info.
In this extreme performance case, it seems the this_cpu_ptr "lookup" of
bpf_redirect_info is the performance issue itself.
Could you experiement with different approaches that modify
xdp_do_redirect() to handle if new helper bpf_redirect_xsk was called,
prior to this_cpu_ptr() call.
(Thus, avoiding to introduce a new XDP-action).
Thanks for your feedback Jesper.
I understand the hesitation of adding a new action. If we can achieve the
same improvement without
quoted
introducing a new action I would be very happy!
Without new the action we'll need a new way to indicate that the
bpf_redirect_xsk helper was
quoted
called. Maybe another new field in the netdev alongside the xsk_refcnt. Or
else extend
quoted
bpf_redirect_info - if we find a new home for it that it's too costly to
access.
quoted
Thanks for your suggestions. I'll experiment as you suggested and
report back.
I'll add a +1 to the "let's try to solve this without a new return code" :)
Also, I don't think we need a new helper either; the bpf_redirect()
helper takes a flags argument, so we could just use ifindex=0,
flags=DEV_XSK or something like that.
The advantage of a new helper is that we can access the netdev
struct from it and check if there's a valid xsk stored in it, before
returning XDP_REDIRECT without the xskmap lookup. However,
I think your suggestion could work too. We would just
have to delay the check until xdp_do_redirect. At this point
though, if there isn't a valid xsk we might have to drop the packet
instead of falling back to the xskmap.
I think it's OK to require the user to make sure that there is such a
socket loaded before using that flag...
quoted
Also, I think the batching in the driver idea can be generalised: we
just need to generalise the idea of "are all these packets going to the
same place" and have a batched version of xdp_do_redirect(), no? The
other map types do batching internally already, though, so I'm wondering
why batching in the driver helps XSK?
With the current infrastructure figuring out if "all the packets are going
to the same place" looks like an expensive operation which could undo
the benefits of the batching that would come after it. We would need
to run the program N=batch_size times, store the actions and
bpf_redirect_info for each run and perform a series of compares. The new
action really helped here because it could easily indicate if all the
packets in a batch were going to the same place. But I understand it's
not an option. Maybe if we can mitigate the cost of accessing the
bpf_redirect_info as Jesper suggested, we can use a flag in it to signal
what the new action was signalling.
Yes, it would probably require comparing the contents of the whole
bpf_redirect_info, at least as it is now; but assuming we can find a way
to mitigate the cost of accessing that structure, this may not be so
bad, and I believe there is some potential for compressing the state
further so we can get down to a single, or maybe two, 64-bit compares...
-Toke