The idea of moving the napi poll process out of softirq context to a
kernel thread based context is not new.
Paolo Abeni and Hannes Frederic Sowa have proposed patches to move napi
poll to kthread back in 2016. And Felix Fietkau has also proposed
patches of similar ideas to use workqueue to process napi poll just a
few weeks ago.
The main reason we'd like to push forward with this idea is that the
scheduler has poor visibility into cpu cycles spent in softirq context,
and is not able to make optimal scheduling decisions of the user threads.
For example, we see in one of the application benchmark where network
load is high, the CPUs handling network softirqs has ~80% cpu util. And
user threads are still scheduled on those CPUs, despite other more idle
cpus available in the system. And we see very high tail latencies. In this
case, we have to explicitly pin away user threads from the CPUs handling
network softirqs to ensure good performance.
With napi poll moved to kthread, scheduler is in charge of scheduling both
the kthreads handling network load, and the user threads, and is able to
make better decisions. In the previous benchmark, if we do this and we
pin the kthreads processing napi poll to specific CPUs, scheduler is
able to schedule user threads away from these CPUs automatically.
And the reason we prefer 1 kthread per napi, instead of 1 workqueue
entity per host, is that kthread is more configurable than workqueue,
and we could leverage existing tuning tools for threads, like taskset,
chrt, etc to tune scheduling class and cpu set, etc. Another reason is
if we eventually want to provide busy poll feature using kernel threads
for napi poll, kthread seems to be more suitable than workqueue.
Furthermore, for large platforms with 2 NICs attached to 2 sockets,
kthread is more flexible to be pinned to different sets of CPUs.
In this patch series, I revived Paolo and Hannes's patch in 2016 and
made modifications. Then there are changes proposed by Felix, Jakub,
Paolo and myself on top of those, with suggestions from Eric Dumazet.
In terms of performance, I ran tcp_rr tests with 1000 flows with
various request/response sizes, with RFS/RPS disabled, and compared
performance between softirq vs kthread vs workqueue (patchset proposed
by Felix Fietkau).
Host has 56 hyper threads and 100Gbps nic, 8 rx queues and only 1 numa
node. All threads are unpinned.
req/resp QPS 50%tile 90%tile 99%tile 99.9%tile
softirq 1B/1B 2.75M 337us 376us 1.04ms 3.69ms
kthread 1B/1B 2.67M 371us 408us 455us 550us
workq 1B/1B 2.56M 384us 435us 673us 822us
softirq 5KB/5KB 1.46M 678us 750us 969us 2.78ms
kthread 5KB/5KB 1.44M 695us 789us 891us 1.06ms
workq 5KB/5KB 1.34M 720us 905us 1.06ms 1.57ms
softirq 1MB/1MB 11.0K 79ms 166ms 306ms 630ms
kthread 1MB/1MB 11.0K 75ms 177ms 303ms 596ms
workq 1MB/1MB 11.0K 79ms 180ms 303ms 587ms
When running workqueue implementation, I found the number of threads
used is usually twice as much as kthread implementation. This probably
introduces higher scheduling cost, which results in higher tail
latencies in most cases.
I also ran an application benchmark, which performs fixed qps remote SSD
read/write operations, with various sizes. Again, both with RFS/RPS
disabled.
The result is as follows:
op_size QPS 50%tile 95%tile 99%tile 99.9%tile
softirq 4K 572.6K 385us 1.5ms 3.16ms 6.41ms
kthread 4K 572.6K 390us 803us 2.21ms 6.83ms
workq 4k 572.6K 384us 763us 3.12ms 6.87ms
softirq 64K 157.9K 736us 1.17ms 3.40ms 13.75ms
kthread 64K 157.9K 745us 1.23ms 2.76ms 9.87ms
workq 64K 157.9K 746us 1.23ms 2.76ms 9.96ms
softirq 1M 10.98K 2.03ms 3.10ms 3.7ms 11.56ms
kthread 1M 10.98K 2.13ms 3.21ms 4.02ms 13.3ms
workq 1M 10.98K 2.13ms 3.20ms 3.99ms 14.12ms
In this set of tests, the latency is predominant by the SSD operation.
Also, the user threads are much busier compared to tcp_rr tests. We have
to pin the kthreads/workqueue threads to limit to a few CPUs, to not
disturb user threads, and provide some isolation.
Changes since v8:
Added description for threaded param in struct net_device in patch 2.
Changes since v7:
Break napi_set_threaded() into 2 parts, one to create kthread called
from netif_napi_add(), the other to set threaded bit in napi_enable(),
to get rid of inconsistency through all napi in 1 dev.
Added documentation for /sys/class/net/<dev>/threaded.
Changes since v6:
Added memory barrier in napi_set_threaded().
Changed /sys/class/net/<dev>/thread to a ternary value.
Change dev->threaded to a bit instead of bool.
Changes since v5:
Removed ASSERT_RTNL() from napi_set_threaded() and removed rtnl_lock()
operation from napi_enable().
Changes since v4:
Recorded the threaded setting in dev and restore it in napi_enable().
Changes since v3:
Merged and rearranged patches in a logical order for easier review.
Changed sysfs control to be per device.
Changes since v2:
Corrected typo in patch 1, and updated the cover letter with more
detailed and updated test results.
Changes since v1:
Replaced kthread_create() with kthread_run() in patch 5 as suggested by
Felix Fietkau.
Changes since RFC:
Renamed the kthreads to be napi/<dev>-<napi_id> in patch 5 as suggested
by Hannes Frederic Sowa.
Felix Fietkau (1):
net: extract napi poll functionality to __napi_poll()
Wei Wang (2):
net: implement threaded-able napi poll loop support
net: add sysfs attribute to control napi threaded mode
Documentation/ABI/testing/sysfs-class-net | 15 ++
include/linux/netdevice.h | 23 +--
net/core/dev.c | 209 ++++++++++++++++++++--
net/core/net-sysfs.c | 50 ++++++
4 files changed, 273 insertions(+), 24 deletions(-)
--
2.30.0.365.g02bc693789-goog
From: Felix Fietkau <nbd@nbd.name>
This commit introduces a new function __napi_poll() which does the main
logic of the existing napi_poll() function, and will be called by other
functions in later commits.
This idea and implementation is done by Felix Fietkau [off-list ref] and
is proposed as part of the patch to move napi work to work_queue
context.
This commit by itself is a code restructure.
Signed-off-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Wei Wang <redacted>
---
net/core/dev.c | 35 +++++++++++++++++++++++++----------
1 file changed, 25 insertions(+), 10 deletions(-)
@@ -6768,15 +6768,10 @@ void __netif_napi_del(struct napi_struct *napi)}EXPORT_SYMBOL(__netif_napi_del);-staticintnapi_poll(structnapi_struct*n,structlist_head*repoll)+staticint__napi_poll(structnapi_struct*n,bool*repoll){-void*have;intwork,weight;-list_del_init(&n->poll_list);--have=netpoll_poll_lock(n);-weight=n->weight;/* This NAPI_STATE_SCHED test is for avoiding a race
@@ -6796,7 +6791,7 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll)n->poll,work,weight);if(likely(work<weight))-gotoout_unlock;+returnwork;/* Drivers must not modify the NAPI state if they*consumetheentireweight.Insuchcasesthiscode
@@ -6805,7 +6800,7 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll)*/if(unlikely(napi_disable_pending(n))){napi_complete(n);-gotoout_unlock;+returnwork;}/* The NAPI context has more processing work, but busy-polling
This patch adds a new sysfs attribute to the network device class.
Said attribute provides a per-device control to enable/disable the
threaded mode for all the napi instances of the given network device,
without the need for a device up/down.
User sets it to 1 or 0 to enable or disable threaded mode.
Co-developed-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Co-developed-by: Hannes Frederic Sowa <redacted>
Signed-off-by: Hannes Frederic Sowa <redacted>
Co-developed-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Wei Wang <redacted>
---
Documentation/ABI/testing/sysfs-class-net | 15 ++++++
include/linux/netdevice.h | 2 +
net/core/dev.c | 61 ++++++++++++++++++++++-
net/core/net-sysfs.c | 50 +++++++++++++++++++
4 files changed, 126 insertions(+), 2 deletions(-)
@@ -337,3 +337,18 @@ Contact: netdev@vger.kernel.org Description: 32-bit unsigned integer counting the number of times the link has been down++What: /sys/class/net/<iface>/threaded+Date: Jan 2021+KernelVersion: 5.12+Contact: netdev@vger.kernel.org+Description:+ Boolean value to control the threaded mode per device. User could+ set this value to enable/disable threaded mode for all napi+ belonging to this device, without the need to do device up/down.++ Possible values:+ == ==================================+ 0 threaded mode disabled for this dev+ 1 threaded mode enabled for this dev+ == ==================================
@@ -6740,6 +6741,62 @@ static void init_gro_hash(struct napi_struct *napi)napi->gro_bitmask=0;}+staticintnapi_set_threaded(structnapi_struct*n,boolthreaded)+{+interr=0;++if(threaded==!!test_bit(NAPI_STATE_THREADED,&n->state))+return0;++if(!threaded){+clear_bit(NAPI_STATE_THREADED,&n->state);+return0;+}++if(!n->thread){+err=napi_kthread_create(n);+if(err)+returnerr;+}++/* Make sure kthread is created before THREADED bit+*isset.+*/+smp_mb__before_atomic();+set_bit(NAPI_STATE_THREADED,&n->state);++return0;+}++staticvoiddev_disable_threaded_all(structnet_device*dev)+{+structnapi_struct*napi;++list_for_each_entry(napi,&dev->napi_list,dev_list)+napi_set_threaded(napi,false);+dev->threaded=0;+}++intdev_set_threaded(structnet_device*dev,boolthreaded)+{+structnapi_struct*napi;+intret;++dev->threaded=threaded;+list_for_each_entry(napi,&dev->napi_list,dev_list){+ret=napi_set_threaded(napi,threaded);+if(ret){+/* Error occurred on one of the napi,+*resetthreadedmodeonallnapi.+*/+dev_disable_threaded_all(dev);+break;+}+}++returnret;+}+voidnetif_napi_add(structnet_device*dev,structnapi_struct*napi,int(*poll)(structnapi_struct*,int),intweight){
This patch allows running each napi poll loop inside its own
kernel thread.
The kthread is created during netif_napi_add() if dev->threaded
is set. And threaded mode is enabled in napi_enable(). We will
provide a way to set dev->threaded and enable threaded mode
without a device up/down in the following patch.
Once that threaded mode is enabled and the kthread is
started, napi_schedule() will wake-up such thread instead
of scheduling the softirq.
The threaded poll loop behaves quite likely the net_rx_action,
but it does not have to manipulate local irqs and uses
an explicit scheduling point based on netdev_budget.
Co-developed-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Co-developed-by: Hannes Frederic Sowa <redacted>
Signed-off-by: Hannes Frederic Sowa <redacted>
Co-developed-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: Wei Wang <redacted>
---
include/linux/netdevice.h | 21 +++----
net/core/dev.c | 117 ++++++++++++++++++++++++++++++++++++++
2 files changed, 124 insertions(+), 14 deletions(-)
@@ -358,6 +359,7 @@ enum {NAPI_STATE_NO_BUSY_POLL,/* Do not add in napi_hash, no busy polling */NAPI_STATE_IN_BUSY_POLL,/* sk_busy_loop() owns this NAPI */NAPI_STATE_PREFER_BUSY_POLL,/* prefer busy-polling over softirq processing*/+NAPI_STATE_THREADED,/* The poll is performed inside its own thread*/};enum{
@@ -1493,6 +1494,37 @@ void netdev_notify_peers(struct net_device *dev)}EXPORT_SYMBOL(netdev_notify_peers);+staticintnapi_threaded_poll(void*data);++staticintnapi_kthread_create(structnapi_struct*n)+{+interr=0;++/* Create and wake up the kthread once to put it in+*TASK_INTERRUPTIBLEmodetoavoidtheblockedtask+*warningandworkwithloadavg.+*/+n->thread=kthread_run(napi_threaded_poll,n,"napi/%s-%d",+n->dev->name,n->napi_id);+if(IS_ERR(n->thread)){+err=PTR_ERR(n->thread);+pr_err("kthread_run failed with err %d\n",err);+n->thread=NULL;+}++returnerr;+}++staticvoidnapi_kthread_stop(structnapi_struct*n)+{+if(!n->thread)+return;++kthread_stop(n->thread);+clear_bit(NAPI_STATE_THREADED,&n->state);+n->thread=NULL;+}+staticint__dev_open(structnet_device*dev,structnetlink_ext_ack*extack){conststructnet_device_ops*ops=dev->netdev_ops;
@@ -4252,6 +4284,21 @@ int gro_normal_batch __read_mostly = 8;staticinlinevoid____napi_schedule(structsoftnet_data*sd,structnapi_struct*napi){+structtask_struct*thread;++if(test_bit(NAPI_STATE_THREADED,&napi->state)){+/* Paired with smp_mb__before_atomic() in+*napi_enable().UseREAD_ONCE()toguarantee+*acompletereadonnapi->thread.Onlycall+*wake_up_process()whenit'snotNULL.+*/+thread=READ_ONCE(napi->thread);+if(thread){+wake_up_process(thread);+return;+}+}+list_add_tail(&napi->poll_list,&sd->poll_list);__raise_softirq_irqoff(NET_RX_SOFTIRQ);}
@@ -6720,6 +6767,12 @@ void netif_napi_add(struct net_device *dev, struct napi_struct *napi,set_bit(NAPI_STATE_NPSVC,&napi->state);list_add_rcu(&napi->dev_list,&dev->napi_list);napi_hash_add(napi);+/* Create kthread for this napi if dev->threaded is set.+*Cleardev->threadedifkthreadcreationfailedsothat+*threadedmodewillnotbeenabledinnapi_enable().+*/+if(dev->threaded&&napi_kthread_create(napi))+dev->threaded=0;}EXPORT_SYMBOL(netif_napi_add);
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-02-03 00:29:36
On Fri, 29 Jan 2021 10:18:12 -0800 Wei Wang wrote:
This patch adds a new sysfs attribute to the network device class.
Said attribute provides a per-device control to enable/disable the
threaded mode for all the napi instances of the given network device,
without the need for a device up/down.
User sets it to 1 or 0 to enable or disable threaded mode.
Co-developed-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Co-developed-by: Hannes Frederic Sowa <redacted>
Signed-off-by: Hannes Frederic Sowa <redacted>
Co-developed-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Wei Wang <redacted>
+static int napi_set_threaded(struct napi_struct *n, bool threaded)
+{
+ int err = 0;
+
+ if (threaded == !!test_bit(NAPI_STATE_THREADED, &n->state))
+ return 0;
+
+ if (!threaded) {
+ clear_bit(NAPI_STATE_THREADED, &n->state);
Can we put a note in the commit message saying that stopping the
threads is slightly tricky but we'll do it if someone complains?
Or is there a stronger reason than having to wait for thread to finish
up with the NAPI not to stop them?
quoted hunk
+ return 0;
+ }
+
+ if (!n->thread) {
+ err = napi_kthread_create(n);
+ if (err)
+ return err;
+ }
+
+ /* Make sure kthread is created before THREADED bit
+ * is set.
+ */
+ smp_mb__before_atomic();
+ set_bit(NAPI_STATE_THREADED, &n->state);
+
+ return 0;
+}
+
+static void dev_disable_threaded_all(struct net_device *dev)
+{
+ struct napi_struct *napi;
+
+ list_for_each_entry(napi, &dev->napi_list, dev_list)
+ napi_set_threaded(napi, false);
+ dev->threaded = 0;
+}
+
+int dev_set_threaded(struct net_device *dev, bool threaded)
+{
+ struct napi_struct *napi;
+ int ret;
+
+ dev->threaded = threaded;
+ list_for_each_entry(napi, &dev->napi_list, dev_list) {
+ ret = napi_set_threaded(napi, threaded);
+ if (ret) {
+ /* Error occurred on one of the napi,
+ * reset threaded mode on all napi.
+ */
+ dev_disable_threaded_all(dev);
+ break;
+ }
+ }
+
+ return ret;
+}
+
void netif_napi_add(struct net_device *dev, struct napi_struct *napi,
int (*poll)(struct napi_struct *, int), int weight)
{
Maybe others disagree but I'd take this check out. What's wrong with
letting users see that threaded napi is disabled for devices without
NAPI?
This will also help a little devices which remove NAPIs when they are
down.
I've been caught off guard in the past by the fact that kernel returns
-ENOENT for XPS map when device has a single queue.
+ ret = sprintf(buf, fmt_dec, netdev->threaded);
+
+unlock:
+ rtnl_unlock();
+ return ret;
+}
+
+static int modify_napi_threaded(struct net_device *dev, unsigned long val)
+{
+ int ret;
+
+ if (list_empty(&dev->napi_list))
+ return -EOPNOTSUPP;
+
+ if (val != 0 && val != 1)
+ return -EOPNOTSUPP;
+
+ ret = dev_set_threaded(dev, val);
+
+ return ret;
On Tue, Feb 2, 2021 at 4:28 PM Jakub Kicinski [off-list ref] wrote:
On Fri, 29 Jan 2021 10:18:12 -0800 Wei Wang wrote:
quoted
This patch adds a new sysfs attribute to the network device class.
Said attribute provides a per-device control to enable/disable the
threaded mode for all the napi instances of the given network device,
without the need for a device up/down.
User sets it to 1 or 0 to enable or disable threaded mode.
Co-developed-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Co-developed-by: Hannes Frederic Sowa <redacted>
Signed-off-by: Hannes Frederic Sowa <redacted>
Co-developed-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Wei Wang <redacted>
quoted
+static int napi_set_threaded(struct napi_struct *n, bool threaded)
+{
+ int err = 0;
+
+ if (threaded == !!test_bit(NAPI_STATE_THREADED, &n->state))
+ return 0;
+
+ if (!threaded) {
+ clear_bit(NAPI_STATE_THREADED, &n->state);
Can we put a note in the commit message saying that stopping the
threads is slightly tricky but we'll do it if someone complains?
Or is there a stronger reason than having to wait for thread to finish
up with the NAPI not to stop them?
Yes. The main reason is the napi might be polled at the moment when
clearing this bit. We have to wait for the thread to finish this round
of polling.
Will add a comment on this.
quoted
+ return 0;
+ }
+
+ if (!n->thread) {
+ err = napi_kthread_create(n);
+ if (err)
+ return err;
+ }
+
+ /* Make sure kthread is created before THREADED bit
+ * is set.
+ */
+ smp_mb__before_atomic();
+ set_bit(NAPI_STATE_THREADED, &n->state);
+
+ return 0;
+}
+
+static void dev_disable_threaded_all(struct net_device *dev)
+{
+ struct napi_struct *napi;
+
+ list_for_each_entry(napi, &dev->napi_list, dev_list)
+ napi_set_threaded(napi, false);
+ dev->threaded = 0;
+}
+
+int dev_set_threaded(struct net_device *dev, bool threaded)
+{
+ struct napi_struct *napi;
+ int ret;
+
+ dev->threaded = threaded;
+ list_for_each_entry(napi, &dev->napi_list, dev_list) {
+ ret = napi_set_threaded(napi, threaded);
+ if (ret) {
+ /* Error occurred on one of the napi,
+ * reset threaded mode on all napi.
+ */
+ dev_disable_threaded_all(dev);
+ break;
+ }
+ }
+
+ return ret;
+}
+
void netif_napi_add(struct net_device *dev, struct napi_struct *napi,
int (*poll)(struct napi_struct *, int), int weight)
{
Maybe others disagree but I'd take this check out. What's wrong with
letting users see that threaded napi is disabled for devices without
NAPI?
This will also help a little devices which remove NAPIs when they are
down.
I've been caught off guard in the past by the fact that kernel returns
-ENOENT for XPS map when device has a single queue.
Ack.
quoted
+ ret = sprintf(buf, fmt_dec, netdev->threaded);
+
+unlock:
+ rtnl_unlock();
+ return ret;
+}
+
+static int modify_napi_threaded(struct net_device *dev, unsigned long val)
+{
+ int ret;
+
+ if (list_empty(&dev->napi_list))
+ return -EOPNOTSUPP;
+
+ if (val != 0 && val != 1)
+ return -EOPNOTSUPP;
+
+ ret = dev_set_threaded(dev, val);
+
+ return ret;
From: Alexander Duyck <hidden> Date: 2021-02-03 17:02:49
On Fri, Jan 29, 2021 at 10:20 AM Wei Wang [off-list ref] wrote:
quoted hunk
From: Felix Fietkau <nbd@nbd.name>
This commit introduces a new function __napi_poll() which does the main
logic of the existing napi_poll() function, and will be called by other
functions in later commits.
This idea and implementation is done by Felix Fietkau [off-list ref] and
is proposed as part of the patch to move napi work to work_queue
context.
This commit by itself is a code restructure.
Signed-off-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Wei Wang <redacted>
---
net/core/dev.c | 35 +++++++++++++++++++++++++----------
1 file changed, 25 insertions(+), 10 deletions(-)
@@ -6768,15 +6768,10 @@ void __netif_napi_del(struct napi_struct *napi)}EXPORT_SYMBOL(__netif_napi_del);-staticintnapi_poll(structnapi_struct*n,structlist_head*repoll)+staticint__napi_poll(structnapi_struct*n,bool*repoll){-void*have;intwork,weight;-list_del_init(&n->poll_list);--have=netpoll_poll_lock(n);-weight=n->weight;/* This NAPI_STATE_SCHED test is for avoiding a race
@@ -6796,7 +6791,7 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll)n->poll,work,weight);if(likely(work<weight))-gotoout_unlock;+returnwork;/* Drivers must not modify the NAPI state if they*consumetheentireweight.Insuchcasesthiscode
@@ -6805,7 +6800,7 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll)*/if(unlikely(napi_disable_pending(n))){napi_complete(n);-gotoout_unlock;+returnwork;}/* The NAPI context has more processing work, but busy-polling
@@ -6836,9 +6831,29 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll)if(unlikely(!list_empty(&n->poll_list))){pr_warn_once("%s: Budget exhausted after napi rescheduled\n",n->dev?n->dev->name:"backlog");-gotoout_unlock;+returnwork;}+*repoll=true;++returnwork;+}++staticintnapi_poll(structnapi_struct*n,structlist_head*repoll)+{+booldo_repoll=false;+void*have;+intwork;++list_del_init(&n->poll_list);++have=netpoll_poll_lock(n);++work=__napi_poll(n,&do_repoll);++if(!do_repoll)+gotoout_unlock;+list_add_tail(&n->poll_list,repoll);out_unlock:
Instead of using the out_unlock label why don't you only do the
list_add_tail if do_repoll is true? It will allow you to drop a few
lines of noise. Otherwise this looks good to me.
Reviewed-by: Alexander Duyck <alexanderduyck@fb.com>
From: Alexander Duyck <hidden> Date: 2021-02-03 17:21:19
On Fri, Jan 29, 2021 at 10:22 AM Wei Wang [off-list ref] wrote:
quoted hunk
This patch allows running each napi poll loop inside its own
kernel thread.
The kthread is created during netif_napi_add() if dev->threaded
is set. And threaded mode is enabled in napi_enable(). We will
provide a way to set dev->threaded and enable threaded mode
without a device up/down in the following patch.
Once that threaded mode is enabled and the kthread is
started, napi_schedule() will wake-up such thread instead
of scheduling the softirq.
The threaded poll loop behaves quite likely the net_rx_action,
but it does not have to manipulate local irqs and uses
an explicit scheduling point based on netdev_budget.
Co-developed-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Co-developed-by: Hannes Frederic Sowa <redacted>
Signed-off-by: Hannes Frederic Sowa <redacted>
Co-developed-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: Wei Wang <redacted>
---
include/linux/netdevice.h | 21 +++----
net/core/dev.c | 117 ++++++++++++++++++++++++++++++++++++++
2 files changed, 124 insertions(+), 14 deletions(-)
@@ -358,6 +359,7 @@ enum {NAPI_STATE_NO_BUSY_POLL,/* Do not add in napi_hash, no busy polling */NAPI_STATE_IN_BUSY_POLL,/* sk_busy_loop() owns this NAPI */NAPI_STATE_PREFER_BUSY_POLL,/* prefer busy-polling over softirq processing*/+NAPI_STATE_THREADED,/* The poll is performed inside its own thread*/};enum{
@@ -1493,6 +1494,37 @@ void netdev_notify_peers(struct net_device *dev)}EXPORT_SYMBOL(netdev_notify_peers);+staticintnapi_threaded_poll(void*data);++staticintnapi_kthread_create(structnapi_struct*n)+{+interr=0;++/* Create and wake up the kthread once to put it in+*TASK_INTERRUPTIBLEmodetoavoidtheblockedtask+*warningandworkwithloadavg.+*/+n->thread=kthread_run(napi_threaded_poll,n,"napi/%s-%d",+n->dev->name,n->napi_id);+if(IS_ERR(n->thread)){+err=PTR_ERR(n->thread);+pr_err("kthread_run failed with err %d\n",err);+n->thread=NULL;+}++returnerr;+}++staticvoidnapi_kthread_stop(structnapi_struct*n)+{+if(!n->thread)+return;++kthread_stop(n->thread);+clear_bit(NAPI_STATE_THREADED,&n->state);+n->thread=NULL;+}+
So I think the napi_kthread_stop should also be split into two parts
and distributed between the napi_disable and netif_napi_del functions.
We should probably be clearing the NAPI_STATE_THREADED bit in
napi_disable, and freeing the thread in netif_napi_del.
@@ -4252,6 +4284,21 @@ int gro_normal_batch __read_mostly = 8; static inline void ____napi_schedule(struct softnet_data *sd, struct napi_struct *napi) {+ struct task_struct *thread;++ if (test_bit(NAPI_STATE_THREADED, &napi->state)) {+ /* Paired with smp_mb__before_atomic() in+ * napi_enable(). Use READ_ONCE() to guarantee+ * a complete read on napi->thread. Only call+ * wake_up_process() when it's not NULL.+ */+ thread = READ_ONCE(napi->thread);+ if (thread) {+ wake_up_process(thread);+ return;+ }+ }+ list_add_tail(&napi->poll_list, &sd->poll_list); __raise_softirq_irqoff(NET_RX_SOFTIRQ); }
@@ -6720,6 +6767,12 @@ void netif_napi_add(struct net_device *dev, struct napi_struct *napi, set_bit(NAPI_STATE_NPSVC, &napi->state); list_add_rcu(&napi->dev_list, &dev->napi_list); napi_hash_add(napi);+ /* Create kthread for this napi if dev->threaded is set.+ * Clear dev->threaded if kthread creation failed so that+ * threaded mode will not be enabled in napi_enable().+ */+ if (dev->threaded && napi_kthread_create(napi))+ dev->threaded = 0; } EXPORT_SYMBOL(netif_napi_add);
So I think there may be an issue here since we had netif_napi_add
create the thread, but you are freeing it in napi_kthread_stop if I am
not mistaken. That is why I suggested making this only a clear_bit
call like the ones below to just clear the threaded flag from the
state.
quoted hunk
clear_bit(NAPI_STATE_PREFER_BUSY_POLL, &n->state);
clear_bit(NAPI_STATE_DISABLE, &n->state);
}
EXPORT_SYMBOL(napi_disable);
+/**
+ * napi_enable - enable NAPI scheduling
+ * @n: NAPI context
+ *
+ * Resume NAPI from being scheduled on this context.
+ * Must be paired with napi_disable.
+ */
+void napi_enable(struct napi_struct *n)
+{
+ BUG_ON(!test_bit(NAPI_STATE_SCHED, &n->state));
+ smp_mb__before_atomic();
+ clear_bit(NAPI_STATE_SCHED, &n->state);
+ clear_bit(NAPI_STATE_NPSVC, &n->state);
+ if (n->dev->threaded && n->thread)
+ set_bit(NAPI_STATE_THREADED, &n->state);
+}
+EXPORT_SYMBOL(napi_enable);
+
static void flush_gro_hash(struct napi_struct *napi)
{
int i;
@@ -6862,6 +6934,51 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll) return work; }+static int napi_thread_wait(struct napi_struct *napi)+{+ set_current_state(TASK_INTERRUPTIBLE);++ while (!kthread_should_stop() && !napi_disable_pending(napi)) {+ if (test_bit(NAPI_STATE_SCHED, &napi->state)) {+ WARN_ON(!list_empty(&napi->poll_list));+ __set_current_state(TASK_RUNNING);+ return 0;+ }++ schedule();+ set_current_state(TASK_INTERRUPTIBLE);+ }+ __set_current_state(TASK_RUNNING);+ return -1;+}++static int napi_threaded_poll(void *data)+{+ struct napi_struct *napi = data;+ void *have;++ while (!napi_thread_wait(napi)) {+ for (;;) {+ bool repoll = false;++ local_bh_disable();++ have = netpoll_poll_lock(napi);+ __napi_poll(napi, &repoll);+ netpoll_poll_unlock(have);++ __kfree_skb_flush();+ local_bh_enable();++ if (!repoll)+ break;++ cond_resched();+ }+ }+ return 0;+}+ static __latent_entropy void net_rx_action(struct softirq_action *h) { struct softnet_data *sd = this_cpu_ptr(&softnet_data);--
From: Alexander Duyck <hidden> Date: 2021-02-03 17:31:11
On Tue, Feb 2, 2021 at 5:01 PM Jakub Kicinski [off-list ref] wrote:
On Fri, 29 Jan 2021 10:18:12 -0800 Wei Wang wrote:
quoted
This patch adds a new sysfs attribute to the network device class.
Said attribute provides a per-device control to enable/disable the
threaded mode for all the napi instances of the given network device,
without the need for a device up/down.
User sets it to 1 or 0 to enable or disable threaded mode.
Co-developed-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Co-developed-by: Hannes Frederic Sowa <redacted>
Signed-off-by: Hannes Frederic Sowa <redacted>
Co-developed-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Wei Wang <redacted>
quoted
+static int napi_set_threaded(struct napi_struct *n, bool threaded)
+{
+ int err = 0;
+
+ if (threaded == !!test_bit(NAPI_STATE_THREADED, &n->state))
+ return 0;
+
+ if (!threaded) {
+ clear_bit(NAPI_STATE_THREADED, &n->state);
Can we put a note in the commit message saying that stopping the
threads is slightly tricky but we'll do it if someone complains?
Or is there a stronger reason than having to wait for thread to finish
up with the NAPI not to stop them?
Normally if we are wanting to shut down NAPI we would have to go
through a coordinated process with us setting the NAPI_STATE_DISABLE
bit and then having to sit on NAPI_STATE_SCHED. Doing that would
likely cause a traffic hiccup if somebody toggles this while the NIC
is active so probably best to not interfere.
I suspect this should be more than enough to have us switch in and out
of the threaded setup. I don't think leaving the threads allocated
after someone has enabled it once should be much of an issue. As far
as using just the bit to do the disable, I think the most it would
probably take is a second or so for the queues to switch over from
threaded to normal NAPI again.
quoted
+ return 0;
+ }
+
+ if (!n->thread) {
+ err = napi_kthread_create(n);
+ if (err)
+ return err;
+ }
+
+ /* Make sure kthread is created before THREADED bit
+ * is set.
+ */
+ smp_mb__before_atomic();
+ set_bit(NAPI_STATE_THREADED, &n->state);
+
+ return 0;
+}
+
+static void dev_disable_threaded_all(struct net_device *dev)
+{
+ struct napi_struct *napi;
+
+ list_for_each_entry(napi, &dev->napi_list, dev_list)
+ napi_set_threaded(napi, false);
+ dev->threaded = 0;
+}
+
+int dev_set_threaded(struct net_device *dev, bool threaded)
+{
+ struct napi_struct *napi;
+ int ret;
+
+ dev->threaded = threaded;
+ list_for_each_entry(napi, &dev->napi_list, dev_list) {
+ ret = napi_set_threaded(napi, threaded);
+ if (ret) {
+ /* Error occurred on one of the napi,
+ * reset threaded mode on all napi.
+ */
+ dev_disable_threaded_all(dev);
+ break;
+ }
+ }
+
+ return ret;
+}
+
void netif_napi_add(struct net_device *dev, struct napi_struct *napi,
int (*poll)(struct napi_struct *, int), int weight)
{
Maybe others disagree but I'd take this check out. What's wrong with
letting users see that threaded napi is disabled for devices without
NAPI?
This will also help a little devices which remove NAPIs when they are
down.
I've been caught off guard in the past by the fact that kernel returns
-ENOENT for XPS map when device has a single queue.
I agree there isn't any point to the check. I think this is a
hold-over from the original code that was querying each napi structure
assigned to the device.
quoted
+ ret = sprintf(buf, fmt_dec, netdev->threaded);
+
+unlock:
+ rtnl_unlock();
+ return ret;
+}
+
+static int modify_napi_threaded(struct net_device *dev, unsigned long val)
+{
+ int ret;
+
+ if (list_empty(&dev->napi_list))
+ return -EOPNOTSUPP;
+
+ if (val != 0 && val != 1)
+ return -EOPNOTSUPP;
+
+ ret = dev_set_threaded(dev, val);
+
+ return ret;
On Wed, Feb 3, 2021 at 9:20 AM Alexander Duyck
[off-list ref] wrote:
On Fri, Jan 29, 2021 at 10:22 AM Wei Wang [off-list ref] wrote:
quoted
This patch allows running each napi poll loop inside its own
kernel thread.
The kthread is created during netif_napi_add() if dev->threaded
is set. And threaded mode is enabled in napi_enable(). We will
provide a way to set dev->threaded and enable threaded mode
without a device up/down in the following patch.
Once that threaded mode is enabled and the kthread is
started, napi_schedule() will wake-up such thread instead
of scheduling the softirq.
The threaded poll loop behaves quite likely the net_rx_action,
but it does not have to manipulate local irqs and uses
an explicit scheduling point based on netdev_budget.
Co-developed-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Co-developed-by: Hannes Frederic Sowa <redacted>
Signed-off-by: Hannes Frederic Sowa <redacted>
Co-developed-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: Wei Wang <redacted>
---
include/linux/netdevice.h | 21 +++----
net/core/dev.c | 117 ++++++++++++++++++++++++++++++++++++++
2 files changed, 124 insertions(+), 14 deletions(-)
@@ -358,6 +359,7 @@ enum {NAPI_STATE_NO_BUSY_POLL,/* Do not add in napi_hash, no busy polling */NAPI_STATE_IN_BUSY_POLL,/* sk_busy_loop() owns this NAPI */NAPI_STATE_PREFER_BUSY_POLL,/* prefer busy-polling over softirq processing*/+NAPI_STATE_THREADED,/* The poll is performed inside its own thread*/};enum{
@@ -1493,6 +1494,37 @@ void netdev_notify_peers(struct net_device *dev)}EXPORT_SYMBOL(netdev_notify_peers);+staticintnapi_threaded_poll(void*data);++staticintnapi_kthread_create(structnapi_struct*n)+{+interr=0;++/* Create and wake up the kthread once to put it in+*TASK_INTERRUPTIBLEmodetoavoidtheblockedtask+*warningandworkwithloadavg.+*/+n->thread=kthread_run(napi_threaded_poll,n,"napi/%s-%d",+n->dev->name,n->napi_id);+if(IS_ERR(n->thread)){+err=PTR_ERR(n->thread);+pr_err("kthread_run failed with err %d\n",err);+n->thread=NULL;+}++returnerr;+}++staticvoidnapi_kthread_stop(structnapi_struct*n)+{+if(!n->thread)+return;++kthread_stop(n->thread);+clear_bit(NAPI_STATE_THREADED,&n->state);+n->thread=NULL;+}+
So I think the napi_kthread_stop should also be split into two parts
and distributed between the napi_disable and netif_napi_del functions.
We should probably be clearing the NAPI_STATE_THREADED bit in
napi_disable, and freeing the thread in netif_napi_del.
@@ -4252,6 +4284,21 @@ int gro_normal_batch __read_mostly = 8; static inline void ____napi_schedule(struct softnet_data *sd, struct napi_struct *napi) {+ struct task_struct *thread;++ if (test_bit(NAPI_STATE_THREADED, &napi->state)) {+ /* Paired with smp_mb__before_atomic() in+ * napi_enable(). Use READ_ONCE() to guarantee+ * a complete read on napi->thread. Only call+ * wake_up_process() when it's not NULL.+ */+ thread = READ_ONCE(napi->thread);+ if (thread) {+ wake_up_process(thread);+ return;+ }+ }+ list_add_tail(&napi->poll_list, &sd->poll_list); __raise_softirq_irqoff(NET_RX_SOFTIRQ); }
@@ -6720,6 +6767,12 @@ void netif_napi_add(struct net_device *dev, struct napi_struct *napi, set_bit(NAPI_STATE_NPSVC, &napi->state); list_add_rcu(&napi->dev_list, &dev->napi_list); napi_hash_add(napi);+ /* Create kthread for this napi if dev->threaded is set.+ * Clear dev->threaded if kthread creation failed so that+ * threaded mode will not be enabled in napi_enable().+ */+ if (dev->threaded && napi_kthread_create(napi))+ dev->threaded = 0; } EXPORT_SYMBOL(netif_napi_add);
So I think there may be an issue here since we had netif_napi_add
create the thread, but you are freeing it in napi_kthread_stop if I am
not mistaken. That is why I suggested making this only a clear_bit
call like the ones below to just clear the threaded flag from the
state.
Makes sense. I will split napi_kthread_stop().
quoted
clear_bit(NAPI_STATE_PREFER_BUSY_POLL, &n->state);
clear_bit(NAPI_STATE_DISABLE, &n->state);
}
EXPORT_SYMBOL(napi_disable);
+/**
+ * napi_enable - enable NAPI scheduling
+ * @n: NAPI context
+ *
+ * Resume NAPI from being scheduled on this context.
+ * Must be paired with napi_disable.
+ */
+void napi_enable(struct napi_struct *n)
+{
+ BUG_ON(!test_bit(NAPI_STATE_SCHED, &n->state));
+ smp_mb__before_atomic();
+ clear_bit(NAPI_STATE_SCHED, &n->state);
+ clear_bit(NAPI_STATE_NPSVC, &n->state);
+ if (n->dev->threaded && n->thread)
+ set_bit(NAPI_STATE_THREADED, &n->state);
+}
+EXPORT_SYMBOL(napi_enable);
+
static void flush_gro_hash(struct napi_struct *napi)
{
int i;
@@ -6862,6 +6934,51 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll) return work; }+static int napi_thread_wait(struct napi_struct *napi)+{+ set_current_state(TASK_INTERRUPTIBLE);++ while (!kthread_should_stop() && !napi_disable_pending(napi)) {+ if (test_bit(NAPI_STATE_SCHED, &napi->state)) {+ WARN_ON(!list_empty(&napi->poll_list));+ __set_current_state(TASK_RUNNING);+ return 0;+ }++ schedule();+ set_current_state(TASK_INTERRUPTIBLE);+ }+ __set_current_state(TASK_RUNNING);+ return -1;+}++static int napi_threaded_poll(void *data)+{+ struct napi_struct *napi = data;+ void *have;++ while (!napi_thread_wait(napi)) {+ for (;;) {+ bool repoll = false;++ local_bh_disable();++ have = netpoll_poll_lock(napi);+ __napi_poll(napi, &repoll);+ netpoll_poll_unlock(have);++ __kfree_skb_flush();+ local_bh_enable();++ if (!repoll)+ break;++ cond_resched();+ }+ }+ return 0;+}+ static __latent_entropy void net_rx_action(struct softirq_action *h) { struct softnet_data *sd = this_cpu_ptr(&softnet_data);--
On Wed, Feb 3, 2021 at 9:00 AM Alexander Duyck
[off-list ref] wrote:
On Fri, Jan 29, 2021 at 10:20 AM Wei Wang [off-list ref] wrote:
quoted
From: Felix Fietkau <nbd@nbd.name>
This commit introduces a new function __napi_poll() which does the main
logic of the existing napi_poll() function, and will be called by other
functions in later commits.
This idea and implementation is done by Felix Fietkau [off-list ref] and
is proposed as part of the patch to move napi work to work_queue
context.
This commit by itself is a code restructure.
Signed-off-by: Felix Fietkau <nbd@nbd.name>
Signed-off-by: Wei Wang <redacted>
---
net/core/dev.c | 35 +++++++++++++++++++++++++----------
1 file changed, 25 insertions(+), 10 deletions(-)
@@ -6768,15 +6768,10 @@ void __netif_napi_del(struct napi_struct *napi)}EXPORT_SYMBOL(__netif_napi_del);-staticintnapi_poll(structnapi_struct*n,structlist_head*repoll)+staticint__napi_poll(structnapi_struct*n,bool*repoll){-void*have;intwork,weight;-list_del_init(&n->poll_list);--have=netpoll_poll_lock(n);-weight=n->weight;/* This NAPI_STATE_SCHED test is for avoiding a race
@@ -6796,7 +6791,7 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll)n->poll,work,weight);if(likely(work<weight))-gotoout_unlock;+returnwork;/* Drivers must not modify the NAPI state if they*consumetheentireweight.Insuchcasesthiscode
@@ -6805,7 +6800,7 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll)*/if(unlikely(napi_disable_pending(n))){napi_complete(n);-gotoout_unlock;+returnwork;}/* The NAPI context has more processing work, but busy-polling
@@ -6836,9 +6831,29 @@ static int napi_poll(struct napi_struct *n, struct list_head *repoll)if(unlikely(!list_empty(&n->poll_list))){pr_warn_once("%s: Budget exhausted after napi rescheduled\n",n->dev?n->dev->name:"backlog");-gotoout_unlock;+returnwork;}+*repoll=true;++returnwork;+}++staticintnapi_poll(structnapi_struct*n,structlist_head*repoll)+{+booldo_repoll=false;+void*have;+intwork;++list_del_init(&n->poll_list);++have=netpoll_poll_lock(n);++work=__napi_poll(n,&do_repoll);++if(!do_repoll)+gotoout_unlock;+list_add_tail(&n->poll_list,repoll);out_unlock:
Instead of using the out_unlock label why don't you only do the
list_add_tail if do_repoll is true? It will allow you to drop a few
lines of noise. Otherwise this looks good to me.
Ack.
Reviewed-by: Alexander Duyck <alexanderduyck@fb.com>