From: Eric Dumazet <hidden> Date: 2017-02-09 16:45:00
On Thu, Feb 9, 2017 at 8:41 AM, Tariq Toukan [off-list ref] wrote:
Hi Eric,
Thanks again for your series.
On 09/02/2017 3:58 PM, Eric Dumazet wrote:
As mentioned half a year ago, we better switch mlx4 driver to order-0
allocations and page recycling.
This reduces vulnerability surface thanks to better skb->truesize
tracking and provides better performance in most cases.
v2 provides an ethtool -S new counter (rx_alloc_pages) and
code factorization, plus Tariq fix.
I see that you made significant changes to the previous series, especially
patch 14 (RX CQE processing).
Please notice that our work week has just finished here in Israel.
I will review the series, especially the new patches (10 to 14), on Sunday.
We need to test this series again in our functional and performance
regression systems.
It will be running during the weekend, so we can analyze the results and
update you on Sunday.
Previous performance results showed a degradation, especially in:
- TCP single stream at 64KB length.
What RX ring size are you using ? I have not seen this at all.
- TCP 16 streams at 1KB length.
TCP does not really care, it coalesces all these into TSO skbs, full size...
This was probably because cache was too short, and many page allocations
were needed.
In CX4, we saw the same kind of degradation, much clearer and amplified as
it's 2.5 times faster (100G).
Regards,
Tariq Toukan
On Thu, Feb 9, 2017 at 8:41 AM, Tariq Toukan [off-list ref] wrote:
quoted
Hi Eric,
Thanks again for your series.
On 09/02/2017 3:58 PM, Eric Dumazet wrote:
As mentioned half a year ago, we better switch mlx4 driver to order-0
allocations and page recycling.
This reduces vulnerability surface thanks to better skb->truesize
tracking and provides better performance in most cases.
v2 provides an ethtool -S new counter (rx_alloc_pages) and
code factorization, plus Tariq fix.
I see that you made significant changes to the previous series, especially
patch 14 (RX CQE processing).
Please notice that our work week has just finished here in Israel.
I will review the series, especially the new patches (10 to 14), on Sunday.
We need to test this series again in our functional and performance
regression systems.
It will be running during the weekend, so we can analyze the results and
update you on Sunday.
Previous performance results showed a degradation, especially in:
- TCP single stream at 64KB length.
What RX ring size are you using ? I have not seen this at all.
Default, out of box.
quoted
- TCP 16 streams at 1KB length.
TCP does not really care, it coalesces all these into TSO skbs, full size...
But the kernel stack has to split it back accordingly in the receive side, no?
quoted
This was probably because cache was too short, and many page allocations
were needed.
In CX4, we saw the same kind of degradation, much clearer and amplified as
it's 2.5 times faster (100G).
Regards,
Tariq Toukan
From: Eric Dumazet <hidden> Date: 2017-02-09 17:26:32
On Thu, Feb 9, 2017 at 8:49 AM, Tariq Toukan [off-list ref] wrote:
On 09/02/2017 6:44 PM, Eric Dumazet wrote:
quoted
On Thu, Feb 9, 2017 at 8:41 AM, Tariq Toukan [off-list ref]
wrote:
quoted
Hi Eric,
Thanks again for your series.
On 09/02/2017 3:58 PM, Eric Dumazet wrote:
As mentioned half a year ago, we better switch mlx4 driver to order-0
allocations and page recycling.
This reduces vulnerability surface thanks to better skb->truesize
tracking and provides better performance in most cases.
v2 provides an ethtool -S new counter (rx_alloc_pages) and
code factorization, plus Tariq fix.
I see that you made significant changes to the previous series,
especially
patch 14 (RX CQE processing).
Please notice that our work week has just finished here in Israel.
I will review the series, especially the new patches (10 to 14), on
Sunday.
TCP does not really care, it coalesces all these into TSO skbs, full
size...
But the kernel stack has to split it back accordingly in the receive side,
no?
At 10Gbit or 40Gbit link speed, 16 TCP streams are sending 64KB TSO packets,
regardless of size of write() system calls.
Unless of course application uses write() with 1-byte, this might be
too expensive of course.
We are using 4096 slots per RX queue, this is why I could not reproduce
your results.
A single TCP flow easily can have more than 1024 MSS waiting in its
receive queue (typical receive window on linux is 6MB/2 )
I mentioned that having a slightly inflated skb->truesize might have an
impact in some workloads. (charging for 2048 bytes per MSS instead of
1536), but this is not related to mlx4 and should be tweaked in TCP
stack instead, since this 2048 bytes (half a page on x86) strategy is
now well spread.
We are using 4096 slots per RX queue, this is why I could not reproduce
your results.
Just so others understand this: The number of RX queue slots is
indirectly the size of the page-recycle "cache" in this scheme (that
depend on refcnt tricks to see if page can be reused).
A single TCP flow easily can have more than 1024 MSS waiting in its
receive queue (typical receive window on linux is 6MB/2 )
So, you do need to increase the page-"cache" size, and need this for
real-life cases, interesting.
I mentioned that having a slightly inflated skb->truesize might have an
impact in some workloads. (charging for 2048 bytes per MSS instead of
1536), but this is not related to mlx4 and should be tweaked in TCP
stack instead, since this 2048 bytes (half a page on x86) strategy is
now well spread.
From: Eric Dumazet <hidden> Date: 2017-02-13 00:33:12
On Sun, 2017-02-12 at 23:38 +0100, Jesper Dangaard Brouer wrote:
Just so others understand this: The number of RX queue slots is
indirectly the size of the page-recycle "cache" in this scheme (that
depend on refcnt tricks to see if page can be reused).
Note that the page recycle tricks only work on some occasions.
To provision correctly hosts dealing with TCP flows, one should not rely
on page recycling or any opportunistic (non guaranteed) behavior.
Page recycling, _if_ possible, will help to reduce system load
and thus lower latencies.
quoted
A single TCP flow easily can have more than 1024 MSS waiting in its
receive queue (typical receive window on linux is 6MB/2 )
So, you do need to increase the page-"cache" size, and need this for
real-life cases, interesting.
I believe this sizing was done mostly to cope with normal system
scheduling constraints [1], reducing packet losses under incast blasts.
Sizing happened before I did my patches to switch to order-0 pages
anyway.
The fact that it allowed page-recycling to happen more often was nice of
course.
[1]
- One can not really assume host will always have the ability to process
the RX ring in time, unless maybe CPU are fully dedicated to the napi
polling logic.
- Recent work to shift softirqs to ksoftirqd is potentially magnifying
the problem.