[memnic PATCH 5/7] pmd: packet receiving optimization with prefetch
From: Hiroshi Shimamoto <hidden>
Date: 2014-09-11 07:50:52
Subsystem:
the rest · Maintainer:
Linus Torvalds
From: Hiroshi Shimamoto <h-shimamoto-ehU+Cx/zZe18UrSeD/g0lQ@public.gmane.org> Prefetch the next packet area to reduce memory stall cycles. Prefetching the next packet area could hide memory stall, because the next area will be accessed just after processing the current receive operations. We can see performance improvements with memnic-tester. Using Xeon E5-2697 v2 @ 2.70GHz, 4 vCPU. size | before | after 64 | 4.59Mpps | 5.54Mpps 128 | 4.87Mpps | 5.46Mpps 256 | 4.72Mpps | 5.21Mpps 512 | 4.41Mpps | 4.50Mpps 1024 | 3.64Mpps | 3.71Mpps 1280 | 3.15Mpps | 3.21Mpps 1518 | 2.87Mpps | 2.92Mpps Signed-off-by: Hiroshi Shimamoto <h-shimamoto-ehU+Cx/zZe18UrSeD/g0lQ@public.gmane.org> Reviewed-by: Hayato Momma <h-momma-JhyGz2TFV9J8UrSeD/g0lQ@public.gmane.org> --- pmd/pmd_memnic.c | 11 +++++++---- 1 file changed, 7 insertions(+), 4 deletions(-)
diff --git a/pmd/pmd_memnic.c b/pmd/pmd_memnic.c
index c22a14d..dbe5033 100644
--- a/pmd/pmd_memnic.c
+++ b/pmd/pmd_memnic.c@@ -286,7 +286,7 @@ static uint16_t memnic_recv_pkts(void *rx_queue, uint16_t nr; uint64_t pkts, bytes, errs; uint32_t framesz = adapter->framesz; - int idx; + int idx, next; struct rte_eth_stats *st = &adapter->stats[rte_lcore_id()]; if (!adapter->nic->hdr.valid)
@@ -298,6 +298,11 @@ static uint16_t memnic_recv_pkts(void *rx_queue, p = &data->packets[idx]; if (p->status != MEMNIC_PKT_ST_FILLED) break; + /* prefetch the next area */ + next = idx; + if (++next >= MEMNIC_NR_PACKET) + next = 0; + rte_prefetch0(&data->packets[next]); if (p->len > framesz) { errs++; goto drop;
@@ -318,9 +323,7 @@ static uint16_t memnic_recv_pkts(void *rx_queue, drop: rte_compiler_barrier(); p->status = MEMNIC_PKT_ST_FREE; - - if (++idx >= MEMNIC_NR_PACKET) - idx = 0; + idx = next; } adapter->up_idx = idx;
--
1.8.3.1