Re: [PATCH v3 net-next 08/14] mlx4: use order-0 pages for RX
From: Jesper Dangaard Brouer <hidden>
Date: 2017-02-14 12:12:14
Also in:
linux-mm
On Mon, 13 Feb 2017 15:16:35 -0800 Alexander Duyck [off-list ref] wrote: [...]
... As I'm sure Jesper will probably point out the atomic op for get_page/page_ref_inc can be pretty expensive if I recall correctly.
It is important to understand that there are two cases for the cost of an atomic op, which depend on the cache-coherency state of the cacheline. Measured on Skylake CPU i7-6700K CPU @ 4.00GHz (1) Local CPU atomic op : 27 cycles(tsc) 6.776 ns (2) Remote CPU atomic op: 260 cycles(tsc) 64.964 ns Notice the huge difference. And in case 2, it is enough that the remote CPU reads the cacheline and brings it into "Shared" (MESI) state, and the local CPU then does the atomic op. One key ideas behind the page_pool, is that remote CPUs read/detect refcnt==1 (Shared-state), and store the page in a small per-CPU array. When array is full, it gets bulk returned to the shared-ptr-ring pool. When "local" CPU need new pages, from the shared-ptr-ring it prefetchw during it's bulk refill, to latency-hide the MESI transitions needed. -- Best regards, Jesper Dangaard Brouer MSc.CS, Principal Kernel Engineer at Red Hat LinkedIn: http://www.linkedin.com/in/brouer -- To unsubscribe, send a message with 'unsubscribe linux-mm' in the body to majordomo@kvack.org. For more info on Linux MM, see: http://www.linux-mm.org/ . Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>