From: Yunsheng Lin <hidden> Date: 2021-08-20 06:58:05
Patch 1: Use relaxed atomic for release side accounting
Patch 2: Minor optimize for page_pool_dma_map() function
V2: Remove unnecessary unliky() mark as pointed out by
Heiner.
Yunsheng Lin (2):
page_pool: use relaxed atomic for release side accounting
page_pool: optimize the cpu sync operation when DMA mapping
net/core/page_pool.c | 11 ++++++-----
1 file changed, 6 insertions(+), 5 deletions(-)
--
2.7.4
From: Yunsheng Lin <hidden> Date: 2021-08-20 06:58:12
If the DMA_ATTR_SKIP_CPU_SYNC is not set, cpu syncing is
also done in dma_map_page_attrs(), so set the attrs according
to pool->p.flags to avoid calling cpu sync function again.
Signed-off-by: Yunsheng Lin <redacted>
---
net/core/page_pool.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
From: Yunsheng Lin <hidden> Date: 2021-08-20 06:58:14
There is no need to synchronize the account updating, so
use the relaxed atomic to avoid some memory barrier in the
data path.
Signed-off-by: Yunsheng Lin <redacted>
---
net/core/page_pool.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
@@ -370,7 +370,7 @@ void page_pool_release_page(struct page_pool *pool, struct page *page)/* This may be the last page returned, releasing the pool, so*itisnotsafetoreferencepoolafterwards.*/-count=atomic_inc_return(&pool->pages_state_release_cnt);+count=atomic_inc_return_relaxed(&pool->pages_state_release_cnt);trace_page_pool_state_release(pool,page,count);}EXPORT_SYMBOL(page_pool_release_page);
There is no need to synchronize the account updating, so
use the relaxed atomic to avoid some memory barrier in the
data path.
Signed-off-by: Yunsheng Lin <redacted>
@@ -370,7 +370,7 @@ void page_pool_release_page(struct page_pool *pool, struct page *page)/* This may be the last page returned, releasing the pool, so*itisnotsafetoreferencepoolafterwards.*/-count=atomic_inc_return(&pool->pages_state_release_cnt);+count=atomic_inc_return_relaxed(&pool->pages_state_release_cnt);trace_page_pool_state_release(pool,page,count);}EXPORT_SYMBOL(page_pool_release_page);
On Fri, Aug 20, 2021 at 02:56:51PM +0800, Yunsheng Lin wrote:
If the DMA_ATTR_SKIP_CPU_SYNC is not set, cpu syncing is
also done in dma_map_page_attrs(), so set the attrs according
to pool->p.flags to avoid calling cpu sync function again.
Isn't DMA_ATTR_SKIP_CPU_SYNC checked within dma_map_page_attrs() anyway?
Regards
/Ilias
From: Yunsheng Lin <hidden> Date: 2021-08-23 03:57:01
On 2021/8/20 17:39, Ilias Apalodimas wrote:
On Fri, Aug 20, 2021 at 02:56:51PM +0800, Yunsheng Lin wrote:
quoted
If the DMA_ATTR_SKIP_CPU_SYNC is not set, cpu syncing is
also done in dma_map_page_attrs(), so set the attrs according
to pool->p.flags to avoid calling cpu sync function again.
Isn't DMA_ATTR_SKIP_CPU_SYNC checked within dma_map_page_attrs() anyway?
Yes, the checking in dma_map_page_attrs() should save us from
calling dma_sync_single_for_device() again if we set the attrs
according to "pool->p.flags & PP_FLAG_DMA_SYNC_DEV".
As dma_sync_single_for_device() is EXPORT_SYMBOL()'ed, and
should be a no-op for dma coherent device, so there may be a
function calling overhead for dma coherent device, letting
dma_map_page_attrs() handling the sync seems to avoid the stack
pushing/poping overhead:
https://elixir.bootlin.com/linux/latest/source/kernel/dma/direct.h#L104
The one thing I am not sure about is that the pool->p.offset
and pool->p.max_len are used to decide the sync range before this
patch, while the sync range is the same as the map range when doing
the sync in dma_map_page_attrs().
I assumed the above is not a issue? only sync more than we need?
and it won't hurt the performance?
On Mon, Aug 23, 2021 at 11:56:48AM +0800, Yunsheng Lin wrote:
On 2021/8/20 17:39, Ilias Apalodimas wrote:
quoted
On Fri, Aug 20, 2021 at 02:56:51PM +0800, Yunsheng Lin wrote:
quoted
If the DMA_ATTR_SKIP_CPU_SYNC is not set, cpu syncing is
also done in dma_map_page_attrs(), so set the attrs according
to pool->p.flags to avoid calling cpu sync function again.
Isn't DMA_ATTR_SKIP_CPU_SYNC checked within dma_map_page_attrs() anyway?
Yes, the checking in dma_map_page_attrs() should save us from
calling dma_sync_single_for_device() again if we set the attrs
according to "pool->p.flags & PP_FLAG_DMA_SYNC_DEV".
But we aren't syncing anything right now when we allocate the pages since
this is called with DMA_ATTR_SKIP_CPU_SYNC. We are syncing the allocated
range on the end of the function, if the pool was created and was requested
to take care of the mappings for us.
As dma_sync_single_for_device() is EXPORT_SYMBOL()'ed, and
should be a no-op for dma coherent device, so there may be a
function calling overhead for dma coherent device, letting
dma_map_page_attrs() handling the sync seems to avoid the stack
pushing/poping overhead:
https://elixir.bootlin.com/linux/latest/source/kernel/dma/direct.h#L104
The one thing I am not sure about is that the pool->p.offset
and pool->p.max_len are used to decide the sync range before this
patch, while the sync range is the same as the map range when doing
the sync in dma_map_page_attrs().
I am not sure I am following here. We always sync the entire range as well
in the current code as the mapping function is called with max_len.
I assumed the above is not a issue? only sync more than we need?
and it won't hurt the performance?
We can sync more than we need, but if it's a non-coherent architecture,
there's a performance penalty.
Regards
/Ilias
From: Yunsheng Lin <hidden> Date: 2021-08-24 07:01:12
On 2021/8/23 20:42, Ilias Apalodimas wrote:
On Mon, Aug 23, 2021 at 11:56:48AM +0800, Yunsheng Lin wrote:
quoted
On 2021/8/20 17:39, Ilias Apalodimas wrote:
quoted
On Fri, Aug 20, 2021 at 02:56:51PM +0800, Yunsheng Lin wrote:
[..]
quoted
https://elixir.bootlin.com/linux/latest/source/kernel/dma/direct.h#L104
The one thing I am not sure about is that the pool->p.offset
and pool->p.max_len are used to decide the sync range before this
patch, while the sync range is the same as the map range when doing
the sync in dma_map_page_attrs().
I am not sure I am following here. We always sync the entire range as well
in the current code as the mapping function is called with max_len.
quoted
I assumed the above is not a issue? only sync more than we need?
and it won't hurt the performance?
We can sync more than we need, but if it's a non-coherent architecture,
there's a performance penalty.
Since I do not have any performance data to prove if there is a
performance penalty for non-coherent architecture, I will drop it:)
Hi Yunsheng,
+cc Lorenzo, which has done some tests on non-coherent platforms
On Tue, 24 Aug 2021 at 10:00, Yunsheng Lin [off-list ref] wrote:
On 2021/8/23 20:42, Ilias Apalodimas wrote:
quoted
On Mon, Aug 23, 2021 at 11:56:48AM +0800, Yunsheng Lin wrote:
quoted
On 2021/8/20 17:39, Ilias Apalodimas wrote:
quoted
On Fri, Aug 20, 2021 at 02:56:51PM +0800, Yunsheng Lin wrote:
[..]
quoted
quoted
https://elixir.bootlin.com/linux/latest/source/kernel/dma/direct.h#L104
The one thing I am not sure about is that the pool->p.offset
and pool->p.max_len are used to decide the sync range before this
patch, while the sync range is the same as the map range when doing
the sync in dma_map_page_attrs().
I am not sure I am following here. We always sync the entire range as well
in the current code as the mapping function is called with max_len.
quoted
I assumed the above is not a issue? only sync more than we need?
and it won't hurt the performance?
We can sync more than we need, but if it's a non-coherent architecture,
there's a performance penalty.
Since I do not have any performance data to prove if there is a
performance penalty for non-coherent architecture, I will drop it:)
I am pretty sure it does affect it. Unless I am missing something the
patch simply re-arranges calls to avoid calling dma_map_page_attrs()
right?
However since dma_map_page_attrs() won't do anything sync-related
since it's called with DMA_ATTR_SKIP_CPU_SYNC, I doubt calling it will
have any measurable difference. If there is, we should pick it up.
Regards
/Ilias