Thread (33 messages) flat view 33 messages, 6 authors, 2023-07-26

Re: [RFC 00/12] net: huge page backed page_pool

From: Yunsheng Lin <hidden>
Date: 2023-07-14 13:05:16

On 2023/7/13 1:01, Jakub Kicinski wrote:
On Wed, 12 Jul 2023 14:43:32 +0200 Jesper Dangaard Brouer wrote:
quoted
On 12/07/2023 13.47, Yunsheng Lin wrote:
quoted
On 2023/7/12 8:08, Jakub Kicinski wrote:  
quoted
Oh, I split the page into individual 4k pages after DMA mapping.
There's no need for the host memory to be a huge page. I mean,
the actual kernel identity mapping is a huge page AFAIU, and the
struct pages are allocated, anyway. We just need it to be a huge
page at DMA mapping time.

So the pages from the huge page provider only differ from normal
alloc_page() pages by the fact that they are a part of a 1G DMA
mapping.  
So, Jakub you are saying the PP refcnt's are still done "as usual" on 
individual pages.
Yes - other than coming from a specific 1G of physical memory 
the resulting pages are really pretty ordinary 4k pages.
quoted
quoted
If it is about DMA mapping, is it possible to use dma_map_sg()
to enable a big continuous dma map for a lot of discontinuous
4k pages to avoid allocating big huge page?

As the comment:
"The scatter gather list elements are merged together (if possible)
and tagged with the appropriate dma address and length."

https://elixir.free-electrons.com/linux/v4.16.18/source/arch/arm/mm/dma-mapping.c#L1805
  
This is interesting for two reasons.

(1) if this DMA merging helps IOTLB misses (?)
Maybe I misunderstand how IOMMU / virtual addressing works, but I don't
see how one can merge mappings from physically non-contiguous pages.
IOW we can't get 1G-worth of random 4k pages and hope that thru some
magic they get strung together and share an IOTLB entry (if that's
where Yunsheng's suggestion was going..)
From __arm_lpae_map(), it does seems that smmu in arm can install
pte in different level to point to page of different size.
quoted
(2) PP could use dma_map_sg() to amortize dma_map call cost.

For case (2) __page_pool_alloc_pages_slow() already does bulk allocation
of pages (alloc_pages_bulk_array_node()), and then loops over the pages
to DMA map them individually.  It seems like an obvious win to use
dma_map_sg() here?
For mapping, the above should work, the tricky problem is we need to ensure
all pages belonging to the same big dma mapping is released before we can do
the dma unmapping.
That could well be worth investigating!
.
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help