@@ -2894,6 +2894,7 @@ static int ucc_geth_rx(struct ucc_geth_private *ugeth, u8 rxQ, int rx_work_limitu32bd_status;u8*bdBuffer;structnet_device*dev;+LIST_HEAD(rx_list);ugeth_vdbg("%s: IN",__func__);
@@ -2934,7 +2935,7 @@ static int ucc_geth_rx(struct ucc_geth_private *ugeth, u8 rxQ, int rx_work_limitdev->stats.rx_bytes+=length;/* Send the packet up the stack */-netif_receive_skb(skb);+list_add_tail(&skb->list,&rx_list);}skb=get_new_skb(ugeth,bd);
@@ -2960,6 +2961,8 @@ static int ucc_geth_rx(struct ucc_geth_private *ugeth, u8 rxQ, int rx_work_limitbd_status=in_be32((u32__iomem*)bd);}+netif_receive_skb_list(&rx_list);+ugeth->rxBd[rxQ]=bd;returnhowmany;}
From: Andrew Lunn <andrew@lunn.ch> Date: 2026-05-17 20:24:31
On Sun, May 17, 2026 at 12:28:56PM -0700, Rosen Penev wrote:
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
So my first through was, why is the core not doing this? The core NAPI
poll code can initialise the list. netif_receive_skb() withing the
driver poll would see there is a list and append to it. And when the
poll finished the NAPI core would pass the list up the stack? Maybe
this already exists and this driver is just using the wrong API?
Andrew
On Sun, May 17, 2026 at 1:24 PM Andrew Lunn [off-list ref] wrote:
On Sun, May 17, 2026 at 12:28:56PM -0700, Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
So my first through was, why is the core not doing this? The core NAPI
poll code can initialise the list. netif_receive_skb() withing the
driver poll would see there is a list and append to it. And when the
poll finished the NAPI core would pass the list up the stack? Maybe
this already exists and this driver is just using the wrong API?
I do not know. I know several drivers are already using
netif_receive_skb_list, some even which support hardware checksumming.
See 0a25d92c6f4facaf2852f1aac4cebfe01dd57a91
The core seems to use netif_receive_skb_list_internal. I do not know
the details.
Anyway, the performance difference is real.
From: Andrew Lunn <andrew@lunn.ch> Date: 2026-05-17 21:01:46
On Sun, May 17, 2026 at 01:44:40PM -0700, Rosen Penev wrote:
On Sun, May 17, 2026 at 1:24 PM Andrew Lunn [off-list ref] wrote:
quoted
On Sun, May 17, 2026 at 12:28:56PM -0700, Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
So my first through was, why is the core not doing this? The core NAPI
poll code can initialise the list. netif_receive_skb() withing the
driver poll would see there is a list and append to it. And when the
poll finished the NAPI core would pass the list up the stack? Maybe
this already exists and this driver is just using the wrong API?
I do not know. I know several drivers are already using
netif_receive_skb_list, some even which support hardware checksumming.
See 0a25d92c6f4facaf2852f1aac4cebfe01dd57a91
The core seems to use netif_receive_skb_list_internal. I do not know
the details.
Anyway, the performance difference is real.
I'm not disagreeing with that. But can a similar performance
difference be made for all drivers by doing this is the core?
That is the interesting question.
Andrew
From: Jakub Kicinski <kuba@kernel.org> Date: 2026-05-20 23:57:47
On Sun, 17 May 2026 12:28:56 -0700 Rosen Penev wrote:
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
GRO should be even better.
Speedup tested with bidirectional iperf3.
Please mention the platform / board as well.
--
pw-bot: cr
On Wed, May 20, 2026 at 4:57 PM Jakub Kicinski [off-list ref] wrote:
On Sun, 17 May 2026 12:28:56 -0700 Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
GRO should be even better.
GRO will result in slower routing performance because there is no
hardware checksum.
From: Jakub Kicinski <kuba@kernel.org> Date: 2026-05-21 00:45:47
On Wed, 20 May 2026 17:39:41 -0700 Rosen Penev wrote:
quoted
On Sun, 17 May 2026 12:28:56 -0700 Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
GRO should be even better.
GRO will result in slower routing performance because there is no
hardware checksum.
Mention this in the commit message too.
Network adapters without checksum offload are pretty rare these days.
Speaking of being old, do you know if this driver is used in practice?
Maybe we can delete it.
On Wed, May 20, 2026 at 5:45 PM Jakub Kicinski [off-list ref] wrote:
On Wed, 20 May 2026 17:39:41 -0700 Rosen Penev wrote:
quoted
quoted
On Sun, 17 May 2026 12:28:56 -0700 Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
GRO should be even better.
GRO will result in slower routing performance because there is no
hardware checksum.
Mention this in the commit message too.
Will do.
Network adapters without checksum offload are pretty rare these days.
Qualcomm continues to make adapters like these.
Speaking of being old, do you know if this driver is used in practice?
Yes. In OpenWrt with kernel 6.18.
Maybe we can delete it.
Way too early. I've tried previously to clean it up some but got rejected.
Hi Jakub,
Le 21/05/2026 à 02:45, Jakub Kicinski a écrit :
On Wed, 20 May 2026 17:39:41 -0700 Rosen Penev wrote:
quoted
quoted
On Sun, 17 May 2026 12:28:56 -0700 Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
GRO should be even better.
GRO will result in slower routing performance because there is no
hardware checksum.
Mention this in the commit message too.
Network adapters without checksum offload are pretty rare these days.
Speaking of being old, do you know if this driver is used in practice?
Maybe we can delete it.
That's way too early to remove that driver.
UCC is what provides Ethernet connectivity in the powerpc MPC83xx CPU
family. This family has just been declared End Of Life by January this
year with a Last Time Buy Date 30 Jul 2026 and Last Time Delivery Date
by 30 Apr 2027. We have that CPU on hundreds of boards spread all over
Europe and have to maintain those systems for the next 10 to 15 years if
not even more.
So we can maybe reconsider removing that driver by 2040 but unlikely before.
Christophe
From: Eric Dumazet <edumazet@google.com> Date: 2026-05-21 13:41:15
On Wed, May 20, 2026 at 5:39 PM Rosen Penev [off-list ref] wrote:
On Wed, May 20, 2026 at 4:57 PM Jakub Kicinski [off-list ref] wrote:
quoted
On Sun, 17 May 2026 12:28:56 -0700 Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
GRO should be even better.
GRO will result in slower routing performance because there is no
hardware checksum.
Then provide a knob or something, instead of trying to avoid GRO.
For end hosts (forwarding not enabled), checksum will need to be
computed anyway.
GRO should be faster for them.
Note that GRO also uses netif_receive_skb_list_internal()
On Thu, May 21, 2026 at 6:41 AM Eric Dumazet [off-list ref] wrote:
On Wed, May 20, 2026 at 5:39 PM Rosen Penev [off-list ref] wrote:
quoted
On Wed, May 20, 2026 at 4:57 PM Jakub Kicinski [off-list ref] wrote:
quoted
On Sun, 17 May 2026 12:28:56 -0700 Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
GRO should be even better.
GRO will result in slower routing performance because there is no
hardware checksum.
Then provide a knob or something, instead of trying to avoid GRO.
For end hosts (forwarding not enabled), checksum will need to be
computed anyway.
GRO should be faster for them.
Note that GRO also uses netif_receive_skb_list_internal()
so you recommend switching to napi_gro_receive even though there's no
RX hardware checksum?
From: Eric Dumazet <edumazet@google.com> Date: 2026-05-22 04:44:40
On Thu, May 21, 2026 at 4:29 PM Rosen Penev [off-list ref] wrote:
On Thu, May 21, 2026 at 6:41 AM Eric Dumazet [off-list ref] wrote:
quoted
On Wed, May 20, 2026 at 5:39 PM Rosen Penev [off-list ref] wrote:
quoted
On Wed, May 20, 2026 at 4:57 PM Jakub Kicinski [off-list ref] wrote:
quoted
On Sun, 17 May 2026 12:28:56 -0700 Rosen Penev wrote:
quoted
Collect received skbs on a local list during RX polling and pass the
completed batch to netif_receive_skb_list(). This lets the networking
stack process packets from a poll cycle in bulk instead of handing each
skb up individually.
GRO should be even better.
GRO will result in slower routing performance because there is no
hardware checksum.
Then provide a knob or something, instead of trying to avoid GRO.
For end hosts (forwarding not enabled), checksum will need to be
computed anyway.
GRO should be faster for them.
Note that GRO also uses netif_receive_skb_list_internal()
so you recommend switching to napi_gro_receive even though there's no
RX hardware checksum?
Certainly.
There is a reason we added support for sw checksum in GRO years ago.
Most linux hosts on this planet do not forward packets.
And if they do, there is a big chance the egress device supports TSO
or tx checksum offload.