The new Mikrotik 10/25G NIC maintains compatibility with existing atl1c
driver. However it does have new features.
This patch set adds support for reporting cards higher link speed, max-mtu,
enables rx csum offload and improves tx performance.
Gatis Peisenieks (4):
atl1c: show correct link speed on Mikrotik 10/25G NIC
atl1c: improve performance by avoiding unnecessary pcie writes on xmit
atl1c: adjust max mtu according to Mikrotik 10/25G NIC ability
atl1c: enable rx csum offload on Mikrotik 10/25G NIC
drivers/net/ethernet/atheros/atl1c/atl1c.h | 3 ++
drivers/net/ethernet/atheros/atl1c/atl1c_hw.c | 9 +++++
drivers/net/ethernet/atheros/atl1c/atl1c_hw.h | 7 ++++
.../net/ethernet/atheros/atl1c/atl1c_main.c | 33 +++++++++++++++----
4 files changed, 46 insertions(+), 6 deletions(-)
base-commit: 3913ba732e972d88ebc391323999e780a9295852
--
2.31.1
Mikrotik 10/25G NIC supports hw checksum verification on rx for
IP/IPv6 + TCP/UDP packets. HW checksum offload helps reduce host
cpu load.
This enables the csum offload specifically for Mikrotik 10/25G NIC
as other HW supported by the driver is known to have problems with it.
TCP iperf3 to Threadripper 3960X with NIC improved 16.5 -> 20.0 Gbps
with mtu=1500.
Signed-off-by: Gatis Peisenieks <redacted>
---
drivers/net/ethernet/atheros/atl1c/atl1c.h | 2 ++
drivers/net/ethernet/atheros/atl1c/atl1c_main.c | 5 +++++
2 files changed, 7 insertions(+)
The new Mikrotik 10/25G NIC maintains compatibility with existing atl1c
driver. However it does have new features.
This defines some new register offsets, code for identifying the new type
of NIC and correct speed detection for the NIC.
Signed-off-by: Gatis Peisenieks <redacted>
---
drivers/net/ethernet/atheros/atl1c/atl1c.h | 1 +
drivers/net/ethernet/atheros/atl1c/atl1c_hw.c | 9 +++++++++
drivers/net/ethernet/atheros/atl1c/atl1c_hw.h | 7 +++++++
drivers/net/ethernet/atheros/atl1c/atl1c_main.c | 4 ++++
4 files changed, 21 insertions(+)
The kernel has xmit_more facility that hints the networking driver xmit
path about whether more packets are coming soon. This information can be
used to avoid unnecessary expensive PCIe transaction per tx packet at a
slight increase in latency.
Max TX pps on Mikrotik 10/25G NIC in a Threadripper 3960X system
improved from 1150Kpps to 1700Kpps.
Signed-off-by: Gatis Peisenieks <redacted>
---
drivers/net/ethernet/atheros/atl1c/atl1c_main.c | 13 +++++++++----
1 file changed, 9 insertions(+), 4 deletions(-)
@@ -2211,8 +2211,8 @@ static int atl1c_tx_map(struct atl1c_adapter *adapter,return-1;}-staticvoidatl1c_tx_queue(structatl1c_adapter*adapter,structsk_buff*skb,-structatl1c_tpd_desc*tpd,enumatl1c_trans_queuetype)+staticvoidatl1c_tx_queue(structatl1c_adapter*adapter,+enumatl1c_trans_queuetype){structatl1c_tpd_ring*tpd_ring=&adapter->tpd_ring[type];u16reg;
@@ -2238,6 +2238,7 @@ static netdev_tx_t atl1c_xmit_frame(struct sk_buff *skb,if(atl1c_tpd_avail(adapter,type)<tpd_req){/* no enough descriptor, just stop queue */+atl1c_tx_queue(adapter,type);netif_stop_queue(netdev);returnNETDEV_TX_BUSY;}
@@ -2246,6 +2247,7 @@ static netdev_tx_t atl1c_xmit_frame(struct sk_buff *skb,/* do TSO and check sum */if(atl1c_tso_csum(adapter,skb,&tpd,type)!=0){+atl1c_tx_queue(adapter,type);dev_kfree_skb_any(skb);returnNETDEV_TX_OK;}
The new Mikrotik 10/25G NIC supports jumbo frames. Jumbo frames are
supported for TSO as well.
This enables the support for mtu up to 9500 bytes.
Signed-off-by: Gatis Peisenieks <redacted>
---
drivers/net/ethernet/atheros/atl1c/atl1c_main.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
From: Eric Dumazet <hidden> Date: 2021-05-11 21:39:44
On 5/11/21 9:05 PM, Gatis Peisenieks wrote:
quoted hunk
The kernel has xmit_more facility that hints the networking driver xmit
path about whether more packets are coming soon. This information can be
used to avoid unnecessary expensive PCIe transaction per tx packet at a
slight increase in latency.
Max TX pps on Mikrotik 10/25G NIC in a Threadripper 3960X system
improved from 1150Kpps to 1700Kpps.
Signed-off-by: Gatis Peisenieks <redacted>
---
drivers/net/ethernet/atheros/atl1c/atl1c_main.c | 13 +++++++++----
1 file changed, 9 insertions(+), 4 deletions(-)
@@ -2211,8 +2211,8 @@ static int atl1c_tx_map(struct atl1c_adapter *adapter,return-1;}-staticvoidatl1c_tx_queue(structatl1c_adapter*adapter,structsk_buff*skb,-structatl1c_tpd_desc*tpd,enumatl1c_trans_queuetype)+staticvoidatl1c_tx_queue(structatl1c_adapter*adapter,+enumatl1c_trans_queuetype){structatl1c_tpd_ring*tpd_ring=&adapter->tpd_ring[type];u16reg;
@@ -2238,6 +2238,7 @@ static netdev_tx_t atl1c_xmit_frame(struct sk_buff *skb,if(atl1c_tpd_avail(adapter,type)<tpd_req){/* no enough descriptor, just stop queue */+atl1c_tx_queue(adapter,type);netif_stop_queue(netdev);returnNETDEV_TX_BUSY;}
@@ -2246,6 +2247,7 @@ static netdev_tx_t atl1c_xmit_frame(struct sk_buff *skb,/* do TSO and check sum */if(atl1c_tso_csum(adapter,skb,&tpd,type)!=0){+atl1c_tx_queue(adapter,type);dev_kfree_skb_any(skb);returnNETDEV_TX_OK;}
This is probably buggy.
You must check and use the return code of this function,
as in :
bool door_bell = __netdev_sent_queue(adapter->netdev, skb->len, netdev_xmit_more());
if (door_bell)
atl1c_tx_queue(adapter, type);
+ if (!more)
+ atl1c_tx_queue(adapter, type);
}
return NETDEV_TX_OK;
From: Chris Snook <chris.snook@gmail.com> Date: 2021-05-12 02:47:53
Increases in latency tend to hurt more on single-queue devices. Has
this been tested on the original gigabit atl1c?
- Chris
On Tue, May 11, 2021 at 12:05 PM Gatis Peisenieks [off-list ref] wrote:
quoted hunk
The kernel has xmit_more facility that hints the networking driver xmit
path about whether more packets are coming soon. This information can be
used to avoid unnecessary expensive PCIe transaction per tx packet at a
slight increase in latency.
Max TX pps on Mikrotik 10/25G NIC in a Threadripper 3960X system
improved from 1150Kpps to 1700Kpps.
Signed-off-by: Gatis Peisenieks <redacted>
---
drivers/net/ethernet/atheros/atl1c/atl1c_main.c | 13 +++++++++----
1 file changed, 9 insertions(+), 4 deletions(-)
@@ -2211,8 +2211,8 @@ static int atl1c_tx_map(struct atl1c_adapter *adapter,return-1;}-staticvoidatl1c_tx_queue(structatl1c_adapter*adapter,structsk_buff*skb,-structatl1c_tpd_desc*tpd,enumatl1c_trans_queuetype)+staticvoidatl1c_tx_queue(structatl1c_adapter*adapter,+enumatl1c_trans_queuetype){structatl1c_tpd_ring*tpd_ring=&adapter->tpd_ring[type];u16reg;
@@ -2238,6 +2238,7 @@ static netdev_tx_t atl1c_xmit_frame(struct sk_buff *skb,if(atl1c_tpd_avail(adapter,type)<tpd_req){/* no enough descriptor, just stop queue */+atl1c_tx_queue(adapter,type);netif_stop_queue(netdev);returnNETDEV_TX_BUSY;}
@@ -2246,6 +2247,7 @@ static netdev_tx_t atl1c_xmit_frame(struct sk_buff *skb,/* do TSO and check sum */if(atl1c_tso_csum(adapter,skb,&tpd,type)!=0){+atl1c_tx_queue(adapter,type);dev_kfree_skb_any(skb);returnNETDEV_TX_OK;}
This is probably buggy.
You must check and use the return code of this function,
as in :
bool door_bell = __netdev_sent_queue(adapter->netdev, skb->len,
netdev_xmit_more());
if (door_bell)
atl1c_tx_queue(adapter, type);
Eric, thank you for taking your time to look at this!
You are correct, tx queue can get stopped in __netdev_sent_queue
and if there were more packets coming the submit to HW would be
missed / unnecessarily delayed.
quoted
+ if (!more)
+ atl1c_tx_queue(adapter, type);
}
return NETDEV_TX_OK;
From: David Laight <hidden> Date: 2021-05-12 08:33:16
From: Chris Snook <chris.snook@gmail.com>
Sent: 12 May 2021 03:40
On Tue, May 11, 2021 at 12:05 PM Gatis Peisenieks [off-list ref] wrote:
quoted
The kernel has xmit_more facility that hints the networking driver xmit
path about whether more packets are coming soon. This information can be
used to avoid unnecessary expensive PCIe transaction per tx packet at a
slight increase in latency.
Increases in latency tend to hurt more on single-queue devices. Has
this been tested on the original gigabit atl1c?
It probably depends a lot on how expensive it is to 'kick' the mac unit.
A simple (posted) PCIe write when the PCIe host interface is idle (as is likely
when you've just been updating descriptors) is probably noise compared
to the rest of the cost of sending the packet.
(Eric will probably say they measured gains.)
OTOH if you have (as I have on one system) the e1000e driver and some
completely broken 'management interface' hardware which means it can
take a lot of microseconds to write to any MAC register you really
do need to look at netdev_xmit_more() [1].
Unfortunately it doesn't help that much.
netdev_xmit_more() reports the state of the tx queue when the current
skb transmit was passed to the mac driver.
It doesn't report the state of the queue at the time netdev_xmit_more()
is called - so any packets queued while the transmit setup is in
progress don't cause netdev_xmit_more() to return true.
I've traced this happening repeatedly...
[1] If the MI is active MAC writes are broken (may write to the
wrong register), so there is horrid code before each access that
(IIRC) effectively does:
while (mi_active())
mdelay(10);
This is just so broken (interrupts are even enabled).
David
-
Registered Address Lakeside, Bramley Road, Mount Farm, Milton Keynes, MK1 1PT, UK
Registration No: 1397386 (Wales)
Increases in latency tend to hurt more on single-queue devices. Has
this been tested on the original gigabit atl1c?
Thank you Chris, for checking this out!
I did test the atl1c driver with and without this change on actual
AR8151 hardware.
My test system was Intel(R) Core(TM) i7-4790K + RB44Ge.
That is a 4 port 1G AR8151 based card.
I measured latency with external traffic generator with test system
doing L2 forwarding. Receiving traffic on one atl1c interface and
transmitting over another atl1c interface. I had default 1000 packet
pfifo queue configured on the atl1c interfaces.
Max 64 byte packet L2 forward pps rate improved 860K -> 1070K.
Any latency difference at 800Kpps was lost in the noise - with the
particular traffic generator system (a linux based RouterOS
traffic-gen).
I measured average 285us for a 30 second run in both cases. Note that
this includes any traffic generator "internal" latency.
Note that I had to tweak atl1c tx interrupt moderation to get these
numbers. With default tx_imt = 1000 no matter what I get only 500
tx interrupts/sec. Since the tx clean is fast and do not get polled
repeatedly and ring size is 1024 I am limited to ~500Kpps.
tx_imt = 500 dobubles that, I used tx_imt = 200 for this test.
As a side note that still relates to latency discussion on AR8151
hardware what I did find out however is that rx interrupt moderation
timer value has a big effect on latency. Changing rx_imt
from 200 to 10 resulted in considerable improvement from 285us to 41us
average latency as measured by traffic generator. I do not have
enough knowledge of the quirks of all the hardware supported by
the driver to confidently put this in a patch though.
Mikrotik 10/25G NIC has its own interrupt moderation mechanism,
so this is not relevant to that if anyone is interested.
- Chris
On Tue, May 11, 2021 at 12:05 PM Gatis Peisenieks [off-list ref]
wrote:
quoted
The kernel has xmit_more facility that hints the networking driver
xmit
path about whether more packets are coming soon. This information can
be
used to avoid unnecessary expensive PCIe transaction per tx packet at
a
slight increase in latency.
Max TX pps on Mikrotik 10/25G NIC in a Threadripper 3960X system
improved from 1150Kpps to 1700Kpps.