From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Ilias Apalodimas (2):
page_pool: DMA handling and allow to recycles frames via SKB
net: change users of __skb_frag_unref() and add an extra argument
Jesper Dangaard Brouer (1):
xdp: reduce size of struct xdp_mem_info
Matteo Croce (3):
mm: add a signature in struct page
mvpp2: recycle buffers
mvneta: recycle buffers
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/mvneta.c | 4 +-
.../net/ethernet/marvell/mvpp2/mvpp2_main.c | 17 +++----
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
include/linux/mm_types.h | 1 +
include/linux/skbuff.h | 33 +++++++++++--
include/net/page_pool.h | 15 ++++++
include/net/xdp.h | 5 +-
net/core/page_pool.c | 47 +++++++++++++++++++
net/core/skbuff.c | 20 +++++++-
net/core/xdp.c | 14 ++++--
net/tls/tls_device.c | 2 +-
13 files changed, 138 insertions(+), 26 deletions(-)
--
2.30.2
From: Jesper Dangaard Brouer <redacted>
It is possible to compress/reduce the size of struct xdp_mem_info.
This change reduce struct xdp_mem_info from 8 bytes to 4 bytes.
The member xdp_mem_info.id can be reduced to u16, as the mem_id_ht
rhashtable in net/core/xdp.c is already limited by MEM_ID_MAX=0xFFFE
which can safely fit in u16.
The member xdp_mem_info.type could be reduced more than u16, as it stores
the enum xdp_mem_type, but due to alignment it is only reduced to u16.
Signed-off-by: Jesper Dangaard Brouer <redacted>
---
include/net/xdp.h | 4 ++--
net/core/xdp.c | 8 ++++----
2 files changed, 6 insertions(+), 6 deletions(-)
@@ -48,8 +48,8 @@ enum xdp_mem_type {#define XDP_XMIT_FLAGS_MASK XDP_XMIT_FLUSHstructxdp_mem_info{-u32type;/* enum xdp_mem_type, but known size type */-u32id;+u16type;/* enum xdp_mem_type, but known size type */+u16id;};structpage_pool;
@@ -35,11 +35,11 @@ static struct rhashtable *mem_id_ht;staticu32xdp_mem_id_hashfn(constvoid*data,u32len,u32seed){-constu32*k=data;-constu32key=*k;+constu16*k=data;+constu16key=*k;BUILD_BUG_ON(sizeof_field(structxdp_mem_allocator,mem.id)-!=sizeof(u32));+!=sizeof(u16));/* Use cyclic increasing ID as direct hash key */returnkey;
@@ -49,7 +49,7 @@ static int xdp_mem_id_cmp(struct rhashtable_compare_arg *arg,constvoid*ptr){conststructxdp_mem_allocator*xa=ptr;-u32mem_id=*(u32*)arg->key;+u16mem_id=*(u16*)arg->key;returnxa->mem.id!=mem_id;}
@@ -232,6 +232,8 @@ static struct page *__page_pool_alloc_pages_slow(struct page_pool *pool,page_pool_dma_sync_for_device(pool,page,pool->p.max_len);skip_dma_map:+page->signature=PP_SIGNATURE;+/* Track how many pages are held 'in-flight' */pool->pages_state_hold_cnt++;
@@ -302,6 +304,8 @@ void page_pool_release_page(struct page_pool *pool, struct page *page)DMA_ATTR_SKIP_CPU_SYNC);page->dma_addr=0;skip_dma_unmap:+page->signature=0;+/* This may be the last page returned, releasing the pool, so*itisnotsafetoreferencepoolafterwards.*/
From: Ilias Apalodimas <ilias.apalodimas@linaro.org>
During skb_release_data() intercept the packet and if it's a buffer
coming from our page_pool API recycle it back to the pool for further
usage.
To achieve that we introduce a bit in struct sk_buff (pp_recycle:1) and
store the xdp_mem_info in page->private. The SKB bit is needed since
page->private is used by skb_copy_ubufs, so we can't rely solely on
page->private to trigger recycling.
The driver has to take care of the sync operations on it's own
during the buffer recycling since the buffer is never unmapped.
In order to enable recycling the driver must call skb_mark_for_recycle()
to store the information we need for recycling in page->private and
enabling the recycling bit
Storing the information in page->private allows us to recycle both SKBs
and their fragments
Signed-off-by: Ilias Apalodimas <ilias.apalodimas@linaro.org>
Signed-off-by: Jesper Dangaard Brouer <redacted>
Signed-off-by: Matteo Croce <redacted>
---
include/linux/skbuff.h | 33 +++++++++++++++++++++++++++----
include/net/page_pool.h | 13 +++++++++++++
include/net/xdp.h | 1 +
net/core/page_pool.c | 43 +++++++++++++++++++++++++++++++++++++++++
net/core/skbuff.c | 20 +++++++++++++++++--
net/core/xdp.c | 6 ++++++
6 files changed, 110 insertions(+), 6 deletions(-)
@@ -40,6 +40,9 @@#if IS_ENABLED(CONFIG_NF_CONNTRACK)#include<linux/netfilter/nf_conntrack_common.h>#endif+#if IS_BUILTIN(CONFIG_PAGE_POOL)+#include<net/page_pool.h>+#endif/* The interface for checksum offload between the stack and networking drivers*isasfollows...
@@ -243,4 +247,13 @@ static inline void page_pool_ring_unlock(struct page_pool *pool)spin_unlock_bh(&pool->ring.producer_lock);}+/* Store mem_info on struct page and use it while recycling skb frags */+staticinline+voidpage_pool_store_mem_info(structpage*page,structxdp_mem_info*mem)+{+u32*xmi=(u32*)mem;++set_page_private(page,*xmi);+}+#endif /* _NET_PAGE_POOL_H */
@@ -235,6 +235,7 @@ void xdp_return_buff(struct xdp_buff *xdp);voidxdp_flush_frame_bulk(structxdp_frame_bulk*bq);voidxdp_return_frame_bulk(structxdp_frame*xdpf,structxdp_frame_bulk*bq);+voidxdp_return_skb_frame(void*data,structxdp_mem_info*mem);/* When sending xdp_frame into the network stack, then there is no*returnpointcallback,whichisneededtoreleasee.g.DMA-mapping
@@ -17,12 +18,19 @@#include<linux/dma-mapping.h>#include<linux/page-flags.h>#include<linux/mm.h> /* for __put_page() */+#include<net/xdp.h>#include<trace/events/page_pool.h>#define DEFER_TIME (msecs_to_jiffies(1000))#define DEFER_WARN_INTERVAL (60 * HZ)+/* Used to store/retrieve hi/lo bytes from xdp_mem_info to page->private */+unionpage_pool_xmi{+u32raw;+structxdp_mem_infomem_info;+};+staticintpage_pool_init(structpage_pool*pool,conststructpage_pool_params*params){
@@ -587,3 +595,38 @@ void page_pool_update_nid(struct page_pool *pool, int new_nid)}}EXPORT_SYMBOL(page_pool_update_nid);++boolpage_pool_return_skb_page(void*data)+{+structxdp_mem_infomem_info;+unionpage_pool_xmiinfo;+structpage*page;++page=virt_to_head_page(data);+if(unlikely(page->signature!=PP_SIGNATURE))+returnfalse;++info.raw=page_private(page);+mem_info=info.mem_info;++/* If a buffer is marked for recycle and does not belong to+*MEM_TYPE_PAGE_POOL,thebufferswillbeeventuallyfreedfromthe+*networkstackandkfree_skb,buttheDMAregionwill*not*be+*correctlyunmapped.WARNherefortherecyclingmisusage+*/+if(unlikely(mem_info.type!=MEM_TYPE_PAGE_POOL)){+WARN_ONCE(true,"Tried to recycle non MEM_TYPE_PAGE_POOL");+returnfalse;+}++/* Driver set this to memory recycling info. Reset it on recycle+*Thiswill*not*workforNICusingasplit-pagememorymodel.+*Thepagewillbereturnedtothepoolhereregardlessofthe+*'flipped'fragmentbeinginuseornot+*/+set_page_private(page,0);+xdp_return_skb_frame(data,&mem_info);++returntrue;+}+EXPORT_SYMBOL(page_pool_return_skb_page);
@@ -3453,7 +3462,7 @@ int skb_shift(struct sk_buff *tgt, struct sk_buff *skb, int shiftlen)fragto=&skb_shinfo(tgt)->frags[merge];skb_frag_size_add(fragto,skb_frag_size(fragfrom));-__skb_frag_unref(fragfrom);+__skb_frag_unref(fragfrom,skb->pp_recycle);}/* Reposition in the original skb */
@@ -5234,6 +5243,13 @@ bool skb_try_coalesce(struct sk_buff *to, struct sk_buff *from,if(skb_cloned(to))returnfalse;+/* We can't coalesce skb that are allocated from slab and page_pool+*Therecyclemarkisontheskb,sothatmightenduptryingto+*recycleslaballocatedskb->head+*/+if(to->pp_recycle!=from->pp_recycle)+returnfalse;+if(len<=skb_tailroom(to)){if(len)BUG_ON(skb_copy_bits(from,0,skb_put(to,len),len));
From: Ilias Apalodimas <ilias.apalodimas@linaro.org>
On a previous patch we added an extra argument on __skb_frag_unref() to
handle recycling. Update the current users of the function with that.
Signed-off-by: Ilias Apalodimas <ilias.apalodimas@linaro.org>
Signed-off-by: Matteo Croce <redacted>
---
drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
net/tls/tls_device.c | 2 +-
4 files changed, 4 insertions(+), 4 deletions(-)
@@ -2125,7 +2125,7 @@ static int chcr_ktls_xmit(struct sk_buff *skb, struct net_device *dev)/* clear the frag ref count which increased locally before */for(i=0;i<record->num_frags;i++){/* clear the frag ref count */-__skb_frag_unref(&record->frags[i]);+__skb_frag_unref(&record->frags[i],false);}/* if any failure, come out from the loop. */if(ret){
From: Matteo Croce <redacted>
Use the new recycling API for page_pool.
In a drop rate test, the packet rate increased di 10%,
from 269 Kpps to 296 Kpps.
perf top on a stock system shows:
Overhead Shared Object Symbol
21.78% [kernel] [k] __pi___inval_dcache_area
21.66% [mvneta] [k] mvneta_rx_swbm
7.00% [kernel] [k] kmem_cache_alloc
6.05% [kernel] [k] eth_type_trans
4.44% [kernel] [k] kmem_cache_free.part.0
3.80% [kernel] [k] __netif_receive_skb_core
3.68% [kernel] [k] dev_gro_receive
3.65% [kernel] [k] get_page_from_freelist
3.43% [kernel] [k] page_pool_release_page
3.35% [kernel] [k] free_unref_page
And this is the same output with recycling enabled:
Overhead Shared Object Symbol
24.10% [kernel] [k] __pi___inval_dcache_area
23.02% [mvneta] [k] mvneta_rx_swbm
7.19% [kernel] [k] kmem_cache_alloc
6.50% [kernel] [k] eth_type_trans
4.93% [kernel] [k] __netif_receive_skb_core
4.77% [kernel] [k] kmem_cache_free.part.0
3.93% [kernel] [k] dev_gro_receive
3.03% [kernel] [k] build_skb
2.91% [kernel] [k] page_pool_put_page
2.85% [kernel] [k] __xdp_return
The test was done with mausezahn on the TX side with 64 byte raw
ethernet frames.
Signed-off-by: Matteo Croce <redacted>
---
drivers/net/ethernet/marvell/mvneta.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
On Mon, Mar 22, 2021 at 6:03 PM Matteo Croce [off-list ref] wrote:
From: Ilias Apalodimas <ilias.apalodimas@linaro.org>
During skb_release_data() intercept the packet and if it's a buffer
coming from our page_pool API recycle it back to the pool for further
usage.
To achieve that we introduce a bit in struct sk_buff (pp_recycle:1) and
store the xdp_mem_info in page->private. The SKB bit is needed since
page->private is used by skb_copy_ubufs, so we can't rely solely on
page->private to trigger recycling.
The driver has to take care of the sync operations on it's own
during the buffer recycling since the buffer is never unmapped.
In order to enable recycling the driver must call skb_mark_for_recycle()
to store the information we need for recycling in page->private and
enabling the recycling bit
Storing the information in page->private allows us to recycle both SKBs
and their fragments
Signed-off-by: Ilias Apalodimas <ilias.apalodimas@linaro.org>
Signed-off-by: Jesper Dangaard Brouer <redacted>
Signed-off-by: Matteo Croce <redacted>
---
Hi, the patch title really should be:
page_pool: DMA handling and frame recycling via SKBs
As in the previous RFC.
Sorry,
--
per aspera ad upstream
From: David Ahern <hidden> Date: 2021-03-23 14:58:39
On 3/22/21 11:02 AM, Matteo Croce wrote:
From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
patch 4 should be folded into 3; each patch should build without errors.
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Is the end goal of the page_pool changes to remove driver private caches?
Hi David,
On Tue, Mar 23, 2021 at 08:57:57AM -0600, David Ahern wrote:
On 3/22/21 11:02 AM, Matteo Croce wrote:
quoted
From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
patch 4 should be folded into 3; each patch should build without errors.
Yes
quoted
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Is the end goal of the page_pool changes to remove driver private caches?
Yes. The patchset doesn't currently support that , because all the >10gbit
interfaces split the page and we don't account for that. We should be able to
extend it though and account for that. I don't have any hardware
(Intel/mlx) available, but I'll be happy to talk to anyone that does and
figure out a way to support those cards properly.
Cheers
/Ilias
On Mon, 22 Mar 2021 18:03:01 +0100
Matteo Croce [off-list ref] wrote:
quoted hunk
From: Matteo Croce <redacted>
Use the new recycling API for page_pool.
In a drop rate test, the packet rate increased di 10%,
from 269 Kpps to 296 Kpps.
perf top on a stock system shows:
Overhead Shared Object Symbol
21.78% [kernel] [k] __pi___inval_dcache_area
21.66% [mvneta] [k] mvneta_rx_swbm
7.00% [kernel] [k] kmem_cache_alloc
6.05% [kernel] [k] eth_type_trans
4.44% [kernel] [k] kmem_cache_free.part.0
3.80% [kernel] [k] __netif_receive_skb_core
3.68% [kernel] [k] dev_gro_receive
3.65% [kernel] [k] get_page_from_freelist
3.43% [kernel] [k] page_pool_release_page
3.35% [kernel] [k] free_unref_page
And this is the same output with recycling enabled:
Overhead Shared Object Symbol
24.10% [kernel] [k] __pi___inval_dcache_area
23.02% [mvneta] [k] mvneta_rx_swbm
7.19% [kernel] [k] kmem_cache_alloc
6.50% [kernel] [k] eth_type_trans
4.93% [kernel] [k] __netif_receive_skb_core
4.77% [kernel] [k] kmem_cache_free.part.0
3.93% [kernel] [k] dev_gro_receive
3.03% [kernel] [k] build_skb
2.91% [kernel] [k] page_pool_put_page
2.85% [kernel] [k] __xdp_return
The test was done with mausezahn on the TX side with 64 byte raw
ethernet frames.
Signed-off-by: Matteo Croce <redacted>
---
drivers/net/ethernet/marvell/mvneta.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
This cause skb_mark_for_recycle() to set 'skb->pp_recycle=1' multiple
times, for the same SKB. (copy-pasted function below signature to help
reviewers).
This makes me question if we need an API for setting this per page
fragment?
Or if the API skb_mark_for_recycle() need to walk the page fragments in
the SKB and set the info stored in the page for each?
--
Best regards,
Jesper Dangaard Brouer
MSc.CS, Principal Kernel Engineer at Red Hat
LinkedIn: http://www.linkedin.com/in/brouer
From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Ilias Apalodimas (2):
page_pool: DMA handling and allow to recycles frames via SKB
net: change users of __skb_frag_unref() and add an extra argument
Jesper Dangaard Brouer (1):
xdp: reduce size of struct xdp_mem_info
Matteo Croce (3):
mm: add a signature in struct page
mvpp2: recycle buffers
mvneta: recycle buffers
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/mvneta.c | 4 +-
.../net/ethernet/marvell/mvpp2/mvpp2_main.c | 17 +++----
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
include/linux/mm_types.h | 1 +
include/linux/skbuff.h | 33 +++++++++++--
include/net/page_pool.h | 15 ++++++
include/net/xdp.h | 5 +-
net/core/page_pool.c | 47 +++++++++++++++++++
net/core/skbuff.c | 20 +++++++-
net/core/xdp.c | 14 ++++--
net/tls/tls_device.c | 2 +-
13 files changed, 138 insertions(+), 26 deletions(-)
Just for the reference, I've performed some tests on 1G SoC NIC with
this patchset on, here's direct link: [0]
From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Ilias Apalodimas (2):
page_pool: DMA handling and allow to recycles frames via SKB
net: change users of __skb_frag_unref() and add an extra argument
Jesper Dangaard Brouer (1):
xdp: reduce size of struct xdp_mem_info
Matteo Croce (3):
mm: add a signature in struct page
mvpp2: recycle buffers
mvneta: recycle buffers
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/mvneta.c | 4 +-
.../net/ethernet/marvell/mvpp2/mvpp2_main.c | 17 +++----
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
include/linux/mm_types.h | 1 +
include/linux/skbuff.h | 33 +++++++++++--
include/net/page_pool.h | 15 ++++++
include/net/xdp.h | 5 +-
net/core/page_pool.c | 47 +++++++++++++++++++
net/core/skbuff.c | 20 +++++++-
net/core/xdp.c | 14 ++++--
net/tls/tls_device.c | 2 +-
13 files changed, 138 insertions(+), 26 deletions(-)
Just for the reference, I've performed some tests on 1G SoC NIC with
this patchset on, here's direct link: [0]
Thanks for the testing!
Any chance you can get a perf measurement on this?
Is DMA syncing taking a substantial amount of your cpu usage?
Thanks
/Ilias
From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Ilias Apalodimas (2):
page_pool: DMA handling and allow to recycles frames via SKB
net: change users of __skb_frag_unref() and add an extra argument
Jesper Dangaard Brouer (1):
xdp: reduce size of struct xdp_mem_info
Matteo Croce (3):
mm: add a signature in struct page
mvpp2: recycle buffers
mvneta: recycle buffers
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/mvneta.c | 4 +-
.../net/ethernet/marvell/mvpp2/mvpp2_main.c | 17 +++----
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
include/linux/mm_types.h | 1 +
include/linux/skbuff.h | 33 +++++++++++--
include/net/page_pool.h | 15 ++++++
include/net/xdp.h | 5 +-
net/core/page_pool.c | 47 +++++++++++++++++++
net/core/skbuff.c | 20 +++++++-
net/core/xdp.c | 14 ++++--
net/tls/tls_device.c | 2 +-
13 files changed, 138 insertions(+), 26 deletions(-)
Just for the reference, I've performed some tests on 1G SoC NIC with
this patchset on, here's direct link: [0]
Thanks for the testing!
Any chance you can get a perf measurement on this?
I guess you mean perf-report (--stdio) output, right?
Is DMA syncing taking a substantial amount of your cpu usage?
From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Ilias Apalodimas (2):
page_pool: DMA handling and allow to recycles frames via SKB
net: change users of __skb_frag_unref() and add an extra argument
Jesper Dangaard Brouer (1):
xdp: reduce size of struct xdp_mem_info
Matteo Croce (3):
mm: add a signature in struct page
mvpp2: recycle buffers
mvneta: recycle buffers
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/mvneta.c | 4 +-
.../net/ethernet/marvell/mvpp2/mvpp2_main.c | 17 +++----
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
include/linux/mm_types.h | 1 +
include/linux/skbuff.h | 33 +++++++++++--
include/net/page_pool.h | 15 ++++++
include/net/xdp.h | 5 +-
net/core/page_pool.c | 47 +++++++++++++++++++
net/core/skbuff.c | 20 +++++++-
net/core/xdp.c | 14 ++++--
net/tls/tls_device.c | 2 +-
13 files changed, 138 insertions(+), 26 deletions(-)
Just for the reference, I've performed some tests on 1G SoC NIC with
this patchset on, here's direct link: [0]
Thanks for the testing!
Any chance you can get a perf measurement on this?
I guess you mean perf-report (--stdio) output, right?
Yea,
As hinted below, I am just trying to figure out if on Alexander's platform the
cost of syncing, is bigger that free-allocate. I remember one armv7 were that
was the case.
quoted
Is DMA syncing taking a substantial amount of your cpu usage?
From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Ilias Apalodimas (2):
page_pool: DMA handling and allow to recycles frames via SKB
net: change users of __skb_frag_unref() and add an extra argument
Jesper Dangaard Brouer (1):
xdp: reduce size of struct xdp_mem_info
Matteo Croce (3):
mm: add a signature in struct page
mvpp2: recycle buffers
mvneta: recycle buffers
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/mvneta.c | 4 +-
.../net/ethernet/marvell/mvpp2/mvpp2_main.c | 17 +++----
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
include/linux/mm_types.h | 1 +
include/linux/skbuff.h | 33 +++++++++++--
include/net/page_pool.h | 15 ++++++
include/net/xdp.h | 5 +-
net/core/page_pool.c | 47 +++++++++++++++++++
net/core/skbuff.c | 20 +++++++-
net/core/xdp.c | 14 ++++--
net/tls/tls_device.c | 2 +-
13 files changed, 138 insertions(+), 26 deletions(-)
Just for the reference, I've performed some tests on 1G SoC NIC with
this patchset on, here's direct link: [0]
Thanks for the testing!
Any chance you can get a perf measurement on this?
I guess you mean perf-report (--stdio) output, right?
Yea,
As hinted below, I am just trying to figure out if on Alexander's platform the
cost of syncing, is bigger that free-allocate. I remember one armv7 were that
was the case.
quoted
quoted
Is DMA syncing taking a substantial amount of your cpu usage?
That would be the same as for mvneta:
Overhead Shared Object Symbol
24.10% [kernel] [k] __pi___inval_dcache_area
23.02% [mvneta] [k] mvneta_rx_swbm
7.19% [kernel] [k] kmem_cache_alloc
Anyway, I tried to use the recycling *and* napi_build_skb on mvpp2,
and I get lower packet rate than recycling alone.
I don't know why, we should investigate it.
Regards,
--
per aspera ad upstream
From: Matteo Croce <redacted>
This series enables recycling of the buffers allocated with the page_pool API.
The first two patches are just prerequisite to save space in a struct and
avoid recycling pages allocated with other API.
Patch 2 was based on a previous idea from Jonathan Lemon.
The third one is the real recycling, 4 fixes the compilation of __skb_frag_unref
users, and 5,6 enable the recycling on two drivers.
In the last two patches I reported the improvement I have with the series.
The recycling as is can't be used with drivers like mlx5 which do page split,
but this is documented in a comment.
In the future, a refcount can be used so to support mlx5 with no changes.
Ilias Apalodimas (2):
page_pool: DMA handling and allow to recycles frames via SKB
net: change users of __skb_frag_unref() and add an extra argument
Jesper Dangaard Brouer (1):
xdp: reduce size of struct xdp_mem_info
Matteo Croce (3):
mm: add a signature in struct page
mvpp2: recycle buffers
mvneta: recycle buffers
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/mvneta.c | 4 +-
.../net/ethernet/marvell/mvpp2/mvpp2_main.c | 17 +++----
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
include/linux/mm_types.h | 1 +
include/linux/skbuff.h | 33 +++++++++++--
include/net/page_pool.h | 15 ++++++
include/net/xdp.h | 5 +-
net/core/page_pool.c | 47 +++++++++++++++++++
net/core/skbuff.c | 20 +++++++-
net/core/xdp.c | 14 ++++--
net/tls/tls_device.c | 2 +-
13 files changed, 138 insertions(+), 26 deletions(-)
Just for the reference, I've performed some tests on 1G SoC NIC with
this patchset on, here's direct link: [0]
Thanks for the testing!
Any chance you can get a perf measurement on this?
I guess you mean perf-report (--stdio) output, right?
Yea,
As hinted below, I am just trying to figure out if on Alexander's platform the
cost of syncing, is bigger that free-allocate. I remember one armv7 were that
was the case.
quoted
quoted
Is DMA syncing taking a substantial amount of your cpu usage?
(+1 this is an important question)
Sure, I'll drop perf tools to my test env and share the results,
maybe tomorrow or in a few days.
From what I know for sure about MIPS and my platform,
post-Rx synching (dma_sync_single_for_cpu()) is a no-op, and
pre-Rx (dma_sync_single_for_device() etc.) is a bit expensive.
I always have sane page_pool->pp.max_len value (smth about 1668
for MTU of 1500) to minimize the overhead.
By the word, IIRC, all machines shipped with mvpp2 have hardware
cache coherency units and don't suffer from sync routines at all.
That may be the reason why mvpp2 wins the most from this series.
That would be the same as for mvneta:
Overhead Shared Object Symbol
24.10% [kernel] [k] __pi___inval_dcache_area
23.02% [mvneta] [k] mvneta_rx_swbm
7.19% [kernel] [k] kmem_cache_alloc
Anyway, I tried to use the recycling *and* napi_build_skb on mvpp2,
and I get lower packet rate than recycling alone.
I don't know why, we should investigate it.
mvpp2 driver doesn't use napi_consume_skb() on its Tx completion path.
As a result, NAPI percpu caches get refilled only through
kmem_cache_alloc_bulk(), and most of skbuff_head recycling
doesn't work.
On Tue, Mar 23, 2021 at 04:55:31PM +0000, Alexander Lobakin wrote:
quoted
quoted
quoted
quoted
quoted
[...]
quoted
quoted
quoted
quoted
Thanks for the testing!
Any chance you can get a perf measurement on this?
I guess you mean perf-report (--stdio) output, right?
Yea,
As hinted below, I am just trying to figure out if on Alexander's platform the
cost of syncing, is bigger that free-allocate. I remember one armv7 were that
was the case.
quoted
quoted
Is DMA syncing taking a substantial amount of your cpu usage?
(+1 this is an important question)
Sure, I'll drop perf tools to my test env and share the results,
maybe tomorrow or in a few days.
From what I know for sure about MIPS and my platform,
post-Rx synching (dma_sync_single_for_cpu()) is a no-op, and
pre-Rx (dma_sync_single_for_device() etc.) is a bit expensive.
I always have sane page_pool->pp.max_len value (smth about 1668
for MTU of 1500) to minimize the overhead.
By the word, IIRC, all machines shipped with mvpp2 have hardware
cache coherency units and don't suffer from sync routines at all.
That may be the reason why mvpp2 wins the most from this series.
Yep exactly. It's also the reason why you explicitly have to opt-in using the
recycling (by marking the skb for it), instead of hiding the feature in the
page pool internals
Cheers
/Ilias
That would be the same as for mvneta:
Overhead Shared Object Symbol
24.10% [kernel] [k] __pi___inval_dcache_area
23.02% [mvneta] [k] mvneta_rx_swbm
7.19% [kernel] [k] kmem_cache_alloc
Anyway, I tried to use the recycling *and* napi_build_skb on mvpp2,
and I get lower packet rate than recycling alone.
I don't know why, we should investigate it.
mvpp2 driver doesn't use napi_consume_skb() on its Tx completion path.
As a result, NAPI percpu caches get refilled only through
kmem_cache_alloc_bulk(), and most of skbuff_head recycling
doesn't work.
On Tue, Mar 23, 2021 at 04:55:31PM +0000, Alexander Lobakin wrote:
quoted
quoted
quoted
quoted
quoted
quoted
[...]
quoted
quoted
quoted
quoted
quoted
Thanks for the testing!
Any chance you can get a perf measurement on this?
I guess you mean perf-report (--stdio) output, right?
Yea,
As hinted below, I am just trying to figure out if on Alexander's platform the
cost of syncing, is bigger that free-allocate. I remember one armv7 were that
was the case.
quoted
quoted
Is DMA syncing taking a substantial amount of your cpu usage?
(+1 this is an important question)
Sure, I'll drop perf tools to my test env and share the results,
maybe tomorrow or in a few days.
Oh we-e-e-ell...
Looks like I've been fooled by I-cache misses or smth like that.
That happens sometimes, not only on my machines, and not only on
MIPS if I'm not mistaken.
Sorry for confusing you guys.
I got drastically different numbers after I enabled CONFIG_KALLSYMS +
CONFIG_PERF_EVENTS for perf tools.
The only difference in code is that I rebased onto Mel's
mm-bulk-rebase-v6r4.
(lunar is my WIP NIC driver)
1. 5.12-rc3 baseline:
TCP: 566 Mbps
UDP: 615 Mbps
perf top:
4.44% [lunar] [k] lunar_rx_poll_page_pool
3.56% [kernel] [k] r4k_wait_irqoff
2.89% [kernel] [k] free_unref_page
2.57% [kernel] [k] dma_map_page_attrs
2.32% [kernel] [k] get_page_from_freelist
2.28% [lunar] [k] lunar_start_xmit
1.82% [kernel] [k] __copy_user
1.75% [kernel] [k] dev_gro_receive
1.52% [kernel] [k] cpuidle_enter_state_coupled
1.46% [kernel] [k] tcp_gro_receive
1.35% [kernel] [k] __rmemcpy
1.33% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.30% [kernel] [k] __dev_queue_xmit
1.22% [kernel] [k] pfifo_fast_dequeue
1.17% [kernel] [k] skb_release_data
1.17% [kernel] [k] skb_segment
free_unref_page() and get_page_from_freelist() consume a lot.
2. 5.12-rc3 + Page Pool recycling by Matteo:
TCP: 589 Mbps
UDP: 633 Mbps
perf top:
4.27% [lunar] [k] lunar_rx_poll_page_pool
2.68% [lunar] [k] lunar_start_xmit
2.41% [kernel] [k] dma_map_page_attrs
1.92% [kernel] [k] r4k_wait_irqoff
1.89% [kernel] [k] __copy_user
1.62% [kernel] [k] dev_gro_receive
1.51% [kernel] [k] cpuidle_enter_state_coupled
1.44% [kernel] [k] tcp_gro_receive
1.40% [kernel] [k] __rmemcpy
1.38% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.37% [kernel] [k] free_unref_page
1.35% [kernel] [k] __dev_queue_xmit
1.30% [kernel] [k] skb_segment
1.28% [kernel] [k] get_page_from_freelist
1.27% [kernel] [k] r4k_dma_cache_inv
+20 Mbps increase on both TCP and UDP. free_unref_page() and
get_page_from_freelist() dropped down the list significantly.
3. 5.12-rc3 + Page Pool recycling + PP bulk allocator (Mel & Jesper):
TCP: 596
UDP: 641
perf top:
4.38% [lunar] [k] lunar_rx_poll_page_pool
3.34% [kernel] [k] r4k_wait_irqoff
3.14% [kernel] [k] dma_map_page_attrs
2.49% [lunar] [k] lunar_start_xmit
1.85% [kernel] [k] dev_gro_receive
1.76% [kernel] [k] free_unref_page
1.76% [kernel] [k] __copy_user
1.65% [kernel] [k] inet_gro_receive
1.57% [kernel] [k] tcp_gro_receive
1.48% [kernel] [k] cpuidle_enter_state_coupled
1.43% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.42% [kernel] [k] __rmemcpy
1.25% [kernel] [k] skb_segment
1.21% [kernel] [k] r4k_dma_cache_inv
+10 Mbps on top of recycling.
get_page_from_freelist() is gone.
NAPI polling, CPU idle cycle (r4k_wait_irqoff) and DMA mapping
routine became the top consumers.
4-5. __always_inline for rmqueue_bulk() and __rmqueue_pcplist(),
removing 'noinline' from net/core/page_pool.c etc.
...makes absolutely no sense anymore.
I see Mel took Jesper's patch to make __rmqueue_pcplist() inline into
mm-bulk-rebase-v6r5, not sure if it's really needed now.
So I'm really glad we sorted out the things and I can see the real
performance improvements from both recycling and bulk allocations.
quoted
From what I know for sure about MIPS and my platform,
post-Rx synching (dma_sync_single_for_cpu()) is a no-op, and
pre-Rx (dma_sync_single_for_device() etc.) is a bit expensive.
I always have sane page_pool->pp.max_len value (smth about 1668
for MTU of 1500) to minimize the overhead.
By the word, IIRC, all machines shipped with mvpp2 have hardware
cache coherency units and don't suffer from sync routines at all.
That may be the reason why mvpp2 wins the most from this series.
Yep exactly. It's also the reason why you explicitly have to opt-in using the
recycling (by marking the skb for it), instead of hiding the feature in the
page pool internals
Cheers
/Ilias
That would be the same as for mvneta:
Overhead Shared Object Symbol
24.10% [kernel] [k] __pi___inval_dcache_area
23.02% [mvneta] [k] mvneta_rx_swbm
7.19% [kernel] [k] kmem_cache_alloc
Anyway, I tried to use the recycling *and* napi_build_skb on mvpp2,
and I get lower packet rate than recycling alone.
I don't know why, we should investigate it.
mvpp2 driver doesn't use napi_consume_skb() on its Tx completion path.
As a result, NAPI percpu caches get refilled only through
kmem_cache_alloc_bulk(), and most of skbuff_head recycling
doesn't work.
On Tue, Mar 23, 2021 at 04:55:31PM +0000, Alexander Lobakin wrote:
quoted
quoted
quoted
quoted
quoted
quoted
[...]
quoted
quoted
quoted
quoted
quoted
Thanks for the testing!
Any chance you can get a perf measurement on this?
I guess you mean perf-report (--stdio) output, right?
Yea,
As hinted below, I am just trying to figure out if on Alexander's platform the
cost of syncing, is bigger that free-allocate. I remember one armv7 were that
was the case.
quoted
quoted
Is DMA syncing taking a substantial amount of your cpu usage?
(+1 this is an important question)
Sure, I'll drop perf tools to my test env and share the results,
maybe tomorrow or in a few days.
Oh we-e-e-ell...
Looks like I've been fooled by I-cache misses or smth like that.
That happens sometimes, not only on my machines, and not only on
MIPS if I'm not mistaken.
Sorry for confusing you guys.
I got drastically different numbers after I enabled CONFIG_KALLSYMS +
CONFIG_PERF_EVENTS for perf tools.
The only difference in code is that I rebased onto Mel's
mm-bulk-rebase-v6r4.
(lunar is my WIP NIC driver)
1. 5.12-rc3 baseline:
TCP: 566 Mbps
UDP: 615 Mbps
perf top:
4.44% [lunar] [k] lunar_rx_poll_page_pool
3.56% [kernel] [k] r4k_wait_irqoff
2.89% [kernel] [k] free_unref_page
2.57% [kernel] [k] dma_map_page_attrs
2.32% [kernel] [k] get_page_from_freelist
2.28% [lunar] [k] lunar_start_xmit
1.82% [kernel] [k] __copy_user
1.75% [kernel] [k] dev_gro_receive
1.52% [kernel] [k] cpuidle_enter_state_coupled
1.46% [kernel] [k] tcp_gro_receive
1.35% [kernel] [k] __rmemcpy
1.33% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.30% [kernel] [k] __dev_queue_xmit
1.22% [kernel] [k] pfifo_fast_dequeue
1.17% [kernel] [k] skb_release_data
1.17% [kernel] [k] skb_segment
free_unref_page() and get_page_from_freelist() consume a lot.
2. 5.12-rc3 + Page Pool recycling by Matteo:
TCP: 589 Mbps
UDP: 633 Mbps
perf top:
4.27% [lunar] [k] lunar_rx_poll_page_pool
2.68% [lunar] [k] lunar_start_xmit
2.41% [kernel] [k] dma_map_page_attrs
1.92% [kernel] [k] r4k_wait_irqoff
1.89% [kernel] [k] __copy_user
1.62% [kernel] [k] dev_gro_receive
1.51% [kernel] [k] cpuidle_enter_state_coupled
1.44% [kernel] [k] tcp_gro_receive
1.40% [kernel] [k] __rmemcpy
1.38% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.37% [kernel] [k] free_unref_page
1.35% [kernel] [k] __dev_queue_xmit
1.30% [kernel] [k] skb_segment
1.28% [kernel] [k] get_page_from_freelist
1.27% [kernel] [k] r4k_dma_cache_inv
+20 Mbps increase on both TCP and UDP. free_unref_page() and
get_page_from_freelist() dropped down the list significantly.
3. 5.12-rc3 + Page Pool recycling + PP bulk allocator (Mel & Jesper):
TCP: 596
UDP: 641
perf top:
4.38% [lunar] [k] lunar_rx_poll_page_pool
3.34% [kernel] [k] r4k_wait_irqoff
3.14% [kernel] [k] dma_map_page_attrs
2.49% [lunar] [k] lunar_start_xmit
1.85% [kernel] [k] dev_gro_receive
1.76% [kernel] [k] free_unref_page
1.76% [kernel] [k] __copy_user
1.65% [kernel] [k] inet_gro_receive
1.57% [kernel] [k] tcp_gro_receive
1.48% [kernel] [k] cpuidle_enter_state_coupled
1.43% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.42% [kernel] [k] __rmemcpy
1.25% [kernel] [k] skb_segment
1.21% [kernel] [k] r4k_dma_cache_inv
+10 Mbps on top of recycling.
get_page_from_freelist() is gone.
NAPI polling, CPU idle cycle (r4k_wait_irqoff) and DMA mapping
routine became the top consumers.
Again, thanks for the extensive testing.
I assume you dont use page pool to map the buffers right?
Because if the ampping is preserved the only thing you have to do is sync it
after the packet reception
4-5. __always_inline for rmqueue_bulk() and __rmqueue_pcplist(),
removing 'noinline' from net/core/page_pool.c etc.
...makes absolutely no sense anymore.
I see Mel took Jesper's patch to make __rmqueue_pcplist() inline into
mm-bulk-rebase-v6r5, not sure if it's really needed now.
So I'm really glad we sorted out the things and I can see the real
performance improvements from both recycling and bulk allocations.
Those will probably be even bigger with and io(sm)/mu present
[...]
Cheers
/Ilias
This cause skb_mark_for_recycle() to set 'skb->pp_recycle=1' multiple
times, for the same SKB. (copy-pasted function below signature to help
reviewers).
This makes me question if we need an API for setting this per page
fragment?
Or if the API skb_mark_for_recycle() need to walk the page fragments in
the SKB and set the info stored in the page for each?
Considering just performances, I guess it is better open-code here since the
driver already performs a loop over fragments to build the skb, but I guess
this approach is quite risky and I would prefer to have a single utility
routine to take care of linear area + fragments. What do you think?
Regards,
Lorenzo
On Tue, Mar 23, 2021 at 04:55:31PM +0000, Alexander Lobakin wrote:
quoted
quoted
quoted
quoted
quoted
quoted
[...]
quoted
quoted
quoted
quoted
quoted
Thanks for the testing!
Any chance you can get a perf measurement on this?
I guess you mean perf-report (--stdio) output, right?
Yea,
As hinted below, I am just trying to figure out if on Alexander's platform the
cost of syncing, is bigger that free-allocate. I remember one armv7 were that
was the case.
quoted
quoted
Is DMA syncing taking a substantial amount of your cpu usage?
(+1 this is an important question)
Sure, I'll drop perf tools to my test env and share the results,
maybe tomorrow or in a few days.
Oh we-e-e-ell...
Looks like I've been fooled by I-cache misses or smth like that.
That happens sometimes, not only on my machines, and not only on
MIPS if I'm not mistaken.
Sorry for confusing you guys.
I got drastically different numbers after I enabled CONFIG_KALLSYMS +
CONFIG_PERF_EVENTS for perf tools.
The only difference in code is that I rebased onto Mel's
mm-bulk-rebase-v6r4.
(lunar is my WIP NIC driver)
1. 5.12-rc3 baseline:
TCP: 566 Mbps
UDP: 615 Mbps
perf top:
4.44% [lunar] [k] lunar_rx_poll_page_pool
3.56% [kernel] [k] r4k_wait_irqoff
2.89% [kernel] [k] free_unref_page
2.57% [kernel] [k] dma_map_page_attrs
2.32% [kernel] [k] get_page_from_freelist
2.28% [lunar] [k] lunar_start_xmit
1.82% [kernel] [k] __copy_user
1.75% [kernel] [k] dev_gro_receive
1.52% [kernel] [k] cpuidle_enter_state_coupled
1.46% [kernel] [k] tcp_gro_receive
1.35% [kernel] [k] __rmemcpy
1.33% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.30% [kernel] [k] __dev_queue_xmit
1.22% [kernel] [k] pfifo_fast_dequeue
1.17% [kernel] [k] skb_release_data
1.17% [kernel] [k] skb_segment
free_unref_page() and get_page_from_freelist() consume a lot.
2. 5.12-rc3 + Page Pool recycling by Matteo:
TCP: 589 Mbps
UDP: 633 Mbps
perf top:
4.27% [lunar] [k] lunar_rx_poll_page_pool
2.68% [lunar] [k] lunar_start_xmit
2.41% [kernel] [k] dma_map_page_attrs
1.92% [kernel] [k] r4k_wait_irqoff
1.89% [kernel] [k] __copy_user
1.62% [kernel] [k] dev_gro_receive
1.51% [kernel] [k] cpuidle_enter_state_coupled
1.44% [kernel] [k] tcp_gro_receive
1.40% [kernel] [k] __rmemcpy
1.38% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.37% [kernel] [k] free_unref_page
1.35% [kernel] [k] __dev_queue_xmit
1.30% [kernel] [k] skb_segment
1.28% [kernel] [k] get_page_from_freelist
1.27% [kernel] [k] r4k_dma_cache_inv
+20 Mbps increase on both TCP and UDP. free_unref_page() and
get_page_from_freelist() dropped down the list significantly.
3. 5.12-rc3 + Page Pool recycling + PP bulk allocator (Mel & Jesper):
TCP: 596
UDP: 641
perf top:
4.38% [lunar] [k] lunar_rx_poll_page_pool
3.34% [kernel] [k] r4k_wait_irqoff
3.14% [kernel] [k] dma_map_page_attrs
2.49% [lunar] [k] lunar_start_xmit
1.85% [kernel] [k] dev_gro_receive
1.76% [kernel] [k] free_unref_page
1.76% [kernel] [k] __copy_user
1.65% [kernel] [k] inet_gro_receive
1.57% [kernel] [k] tcp_gro_receive
1.48% [kernel] [k] cpuidle_enter_state_coupled
1.43% [nf_conntrack] [k] nf_conntrack_tcp_packet
1.42% [kernel] [k] __rmemcpy
1.25% [kernel] [k] skb_segment
1.21% [kernel] [k] r4k_dma_cache_inv
+10 Mbps on top of recycling.
get_page_from_freelist() is gone.
NAPI polling, CPU idle cycle (r4k_wait_irqoff) and DMA mapping
routine became the top consumers.
Again, thanks for the extensive testing.
I assume you dont use page pool to map the buffers right?
Because if the ampping is preserved the only thing you have to do is sync it
after the packet reception
No, I use Page Pool for both DMA mapping and synching for device.
The reason why DMA mapping takes a lot of CPU is that I test NATing,
so NIC firstly receives the frames and then xmits them with modified
headers -> this DMA map overhead is from lunar_start_xmit(), not
Rx path.
The actual Rx synching is r4k_dma_cache_inv() and it's not that
expensive.
quoted
4-5. __always_inline for rmqueue_bulk() and __rmqueue_pcplist(),
removing 'noinline' from net/core/page_pool.c etc.
...makes absolutely no sense anymore.
I see Mel took Jesper's patch to make __rmqueue_pcplist() inline into
mm-bulk-rebase-v6r5, not sure if it's really needed now.
So I'm really glad we sorted out the things and I can see the real
performance improvements from both recycling and bulk allocations.
Those will probably be even bigger with and io(sm)/mu present
Sure, DMA mapping is way more expensive through IOMMUs. I don't have
one on my boards, so can't collect any useful info.
This cause skb_mark_for_recycle() to set 'skb->pp_recycle=1' multiple
times, for the same SKB. (copy-pasted function below signature to help
reviewers).
This makes me question if we need an API for setting this per page
fragment?
Or if the API skb_mark_for_recycle() need to walk the page fragments in
the SKB and set the info stored in the page for each?
Considering just performances, I guess it is better open-code here since the
driver already performs a loop over fragments to build the skb, but I guess
this approach is quite risky and I would prefer to have a single utility
routine to take care of linear area + fragments. What do you think?
The mark_for_recycle does two things as you noticed,
set the pp_recyle bit on the skb head and update the struct page information we
need to trigger the recycling.
We could split those and be more explicit, but isn't the current approach a
bit simpler for the driver writer to get it right?
I don't think setting a single value to 1 will have any noticeable performance
impact, but we can always test it.