From: Anthony Liguori <hidden> Date: 2012-02-17 23:02:16
Hi,
We've been studying small packet performance under KVM. Currently, this type of
workload performs fairly poorly in KVM compared to other hypervisors.
After a lot of study, we concluded that two factors currently are bottlenecking
performance.
1) vhost uses a single thread to read both the transmit and receive rings. The
code is clearly structured to support multiple threads but only does a single
thread today. As these patches show, this has a significant affect on
performance when running a two VCPU Linux guest under KVM.
2) When dealing with a workload like multiple TCP_RR instances, we process
packets off the transmit queue too quickly. It seems to be rare to actually
get significant batching. vhost does not use a timer for TX mitigation
instead relying on scheduling latency to encourage batching. But in an
unloaded system, the vhost thread simply gets scheduled too quickly and the
exit cost dominates the workload.
In the second patch, we introduce an mechanism to poll the transmit ring for
a short period of time. While this series doesn't show it, in our testing
with similar code, this can have a dramatic affect on throughput.
The second patch does manual tuning but we have a third patch that uses a simple
adaptive algorithm. I'll follow up soon with this patch and results for it.
We're looking to get some feedback on the approach here. I think the first
patch should be pretty non-controversial (other than the broadcast wake-up).
From: Anthony Liguori <hidden> Date: 2012-02-17 23:03:28
With workloads that are dominated by very high rates of small packets, we see
considerable overhead in virtio notifications.
The best strategy we've been able to come up with to deal with this is adaptive
polling. This patch simply adds the infrastructure needed to experiment with
polling strategies. It is not meant for inclusion.
Here are the results with various polling values. The spinning is not currently
a net win due to the high mutex contention caused by the broadcast wakeup. With
a patch attempting to signal wakeup, we see up to 170+ transactions per second
with TCP_RR 60 instance.
N Baseline Spin 0 Spin 1000 Spin 5000
TCP_RR
1 9,639.66 10,164.06 9,825.43 9,827.45 101.95%
10 62,819.55 54,059.78 63,114.30 60,767.23 96.73%
30 84,715.60 131,241.86 120,922.38 89,776.39 105.97%
60 124,614.71 148,720.66 158,678.08 141,400.05 113.47%
UDP_RR
1 9,652.50 10,343.72 9,493.95 9,569.54 99.14%
10 53,830.26 58,235.90 50,145.29 48,820.53 90.69%
30 89,471.01 97,634.53 95,108.34 91,263.65 102.00%
60 103,640.59 164,035.01 157,002.22 128,646.73 124.13%
TCP_STREAM
1 2,622.63 2,610.71 2,688.49 2,678.61 102.13%
4 4,928.02 4,812.05 4,971.00 5,104.57 103.58%
1 5,639.89 5,751.28 5,819.81 5,593.62 99.18%
4 5,874.72 6,575.55 6,324.87 6,502.33 110.68%
1 6,257.42 7,655.22 7,610.52 7,424.74 118.65%
4 5,370.78 6,044.83 5,784.23 6,209.93 115.62%
1 6,346.63 7,267.44 7,567.39 7,677.93 120.98%
4 5,198.02 5,657.12 5,528.94 5,792.42 111.44%
TCP_MAERTS
1 2,091.38 1,765.62 2,142.56 2,312.94 110.59%
4 5,319.52 5,619.49 5,544.50 5,645.81 106.13%
1 7,030.66 7,593.61 7,575.67 7,622.07 108.41%
4 9,040.53 7,275.84 7,322.07 6,681.34 73.90%
1 9,160.93 9,318.15 9,065.82 8,586.82 93.73%
4 9,372.49 8,875.63 8,959.03 9,056.07 96.62%
1 9,183.28 9,134.02 8,945.12 8,657.72 94.28%
4 9,377.17 8,877.52 8,959.54 9,071.53 96.74%
Cc: Tom Lendacky <redacted>
Cc: Cristian Viana <redacted>
Signed-off-by: Anthony Liguori <redacted>
---
drivers/vhost/net.c | 14 ++++++++++++++
1 files changed, 14 insertions(+), 0 deletions(-)
@@ -37,6 +37,10 @@ static int workers = 2;module_param(workers,int,0444);MODULE_PARM_DESC(workers,"Set the number of worker threads");+staticulongspin_threshold=0;+module_param(spin_threshold,ulong,0444);+MODULE_PARM_DESC(spin_threshold,"The polling threshold for the tx queue");+/* Max number of bytes transferred before requeueing the job.*Usingthislimitpreventsonevirtqueuefromstarvingothers.*/#define VHOST_NET_WEIGHT 0x80000
@@ -149,6 +154,7 @@ static void handle_tx(struct vhost_net *net)size_thdr_size;structsocket*sock;structvhost_ubuf_ref*uninitialized_var(ubufs);+size_tspin_count;boolzcopy;/* TODO: check that we are running from vhost_worker? */
From: Anthony Liguori <hidden> Date: 2012-02-17 23:04:38
This patch allows vhost to have multiple worker threads for devices such as
virtio-net which may have multiple virtqueues.
Since virtqueues are a lockless ring queue, in an ideal world data is being
produced by the producer as fast as data is being consumed by the consumer.
These loops will continue to consume data until none is left.
vhost currently multiplexes the consumer side of the queue on a single thread
by attempting to read from the queue until everything is read or it cannot
process anymore. This means that activity on one queue may stall another queue.
This is exacerbated when using any form of polling to read from the queues (as
we'll introduce in the next patch). By spawning a thread per-virtqueue, this
is addressed.
The only problem with this patch right now is how the wake up of the threads is
done. It's essentially a broadcast and we have seen lock contention as a
result. We've tried some approaches to signal a single thread but I'm not
confident that that code is correct yet so I'm only sending the broadcast
version.
Here are some performance results from this change. There's a modest
improvement with stream although a fair bit of variability too.
With RR, there's pretty significant improvements as the instance rate drives up.
Test, Size, Instance, Baseline, Patch, Relative
TCP_RR
Tx:256 Rx:256
1 9,639.66 10,164.06 105.44%
10 62,819.55 54,059.78 86.06%
30 84,715.60 131,241.86 154.92%
60 124,614.71 148,720.66 119.34%
UDP_RR
Tx:256 Rx:256
1 9,652.50 10,343.72 107.16%
10 53,830.26 58,235.90 108.18%
30 89,471.01 97,634.53 109.12%
60 103,640.59 164,035.01 158.27%
TCP_STREAM
Tx: 256
1 2,622.63 2,610.71 99.55%
4 4,928.02 4,812.05 97.65%
Tx: 1024
1 5,639.89 5,751.28 101.97%
4 5,874.72 6,575.55 111.93%
Tx: 4096
1 6,257.42 7,655.22 122.34%
4 5,370.78 6,044.83 112.55%
Tx: 16384
1 6,346.63 7,267.44 114.51%
4 5,198.02 5,657.12 108.83%
TCP_MAERTS
Rx: 256
1 2,091.38 1,765.62 84.42%
4 5,319.52 5,619.49 105.64%
Rx: 1024
1 7,030.66 7,593.61 108.01%
4 9,040.53 7,275.84 80.48%
Rx: 4096
1 9,160.93 9,318.15 101.72%
4 9,372.49 8,875.63 94.70%
Rx: 16384
1 9,183.28 9,134.02 99.46%
4 9,377.17 8,877.52 94.67%
106.46%
Cc: Tom Lendacky <redacted>
Cc: Cristian Viana <redacted>
Signed-off-by: Anthony Liguori <redacted>
---
drivers/vhost/net.c | 6 ++++-
drivers/vhost/vhost.c | 51 ++++++++++++++++++++++++++++++------------------
drivers/vhost/vhost.h | 8 +++++-
3 files changed, 43 insertions(+), 22 deletions(-)
@@ -33,6 +33,10 @@ static int experimental_zcopytx;module_param(experimental_zcopytx,int,0444);MODULE_PARM_DESC(experimental_zcopytx,"Enable Experimental Zero Copy TX");+staticintworkers=2;+module_param(workers,int,0444);+MODULE_PARM_DESC(workers,"Set the number of worker threads");+/* Max number of bytes transferred before requeueing the job.*Usingthislimitpreventsonevirtqueuefromstarvingothers.*/#define VHOST_NET_WEIGHT 0x80000
@@ -300,7 +303,9 @@ long vhost_dev_init(struct vhost_dev *dev,dev->mm=NULL;spin_lock_init(&dev->work_lock);INIT_LIST_HEAD(&dev->work_list);-dev->worker=NULL;+dev->nworkers=min(nworkers,VHOST_MAX_WORKERS);+for(i=0;i<dev->nworkers;i++)+dev->workers[i]=NULL;for(i=0;i<dev->nvqs;++i){dev->vqs[i].log=NULL;
@@ -354,7 +359,7 @@ static int vhost_attach_cgroups(struct vhost_dev *dev)staticlongvhost_dev_set_owner(structvhost_dev*dev){structtask_struct*worker;-interr;+interr,i;/* Is there an owner already? */if(dev->mm){
@@ -364,28 +369,34 @@ static long vhost_dev_set_owner(struct vhost_dev *dev)/* No owner, become one */dev->mm=get_task_mm(current);-worker=kthread_create(vhost_worker,dev,"vhost-%d",current->pid);-if(IS_ERR(worker)){-err=PTR_ERR(worker);-gotoerr_worker;-}+for(i=0;i<dev->nworkers;i++){+worker=kthread_create(vhost_worker,dev,"vhost-%d.%d",current->pid,i);+if(IS_ERR(worker)){+err=PTR_ERR(worker);+gotoerr_worker;+}-dev->worker=worker;-wake_up_process(worker);/* avoid contributing to loadavg */+dev->workers[i]=worker;+wake_up_process(worker);/* avoid contributing to loadavg */+}err=vhost_attach_cgroups(dev);if(err)-gotoerr_cgroup;+gotoerr_worker;err=vhost_dev_alloc_iovecs(dev);if(err)-gotoerr_cgroup;+gotoerr_worker;return0;-err_cgroup:-kthread_stop(worker);-dev->worker=NULL;+err_worker:+for(i=0;i<dev->nworkers;i++){+if(dev->workers[i]){+kthread_stop(dev->workers[i]);+dev->workers[i]=NULL;+}+}if(dev->mm)mmput(dev->mm);dev->mm=NULL;
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-02-19 14:42:02
On Fri, Feb 17, 2012 at 05:02:05PM -0600, Anthony Liguori wrote:
This patch allows vhost to have multiple worker threads for devices such as
virtio-net which may have multiple virtqueues.
Since virtqueues are a lockless ring queue, in an ideal world data is being
produced by the producer as fast as data is being consumed by the consumer.
These loops will continue to consume data until none is left.
vhost currently multiplexes the consumer side of the queue on a single thread
by attempting to read from the queue until everything is read or it cannot
process anymore. This means that activity on one queue may stall another queue.
There's actually an attempt to address this: look up
VHOST_NET_WEIGHT in the code. I take it, this isn't effective?
This is exacerbated when using any form of polling to read from the queues (as
we'll introduce in the next patch). By spawning a thread per-virtqueue, this
is addressed.
The only problem with this patch right now is how the wake up of the threads is
done. It's essentially a broadcast and we have seen lock contention as a
result.
On which lock?
We've tried some approaches to signal a single thread but I'm not
confident that that code is correct yet so I'm only sending the broadcast
version.
Yes, that looks like an obvious question.
Here are some performance results from this change. There's a modest
improvement with stream although a fair bit of variability too.
With RR, there's pretty significant improvements as the instance rate drives up.
Interesting. This was actually tested at one time and we saw
a significant performance improvement from using
a single thread especially with a single
stream in the guest. Profiling indicated that
with a single thread we get too many context
switches between TX and RX, since guest networking
tends to run TX and RX processing on the same
guest VCPU.
Maybe we were wrong or maybe this went away
for some reason. I'll see if this can be reproduced.
One point missing here is the ratio of
throughput divided by host CPU utilization.
The concern being to avoid hurting
heavily loaded systems.
These numbers also tend to be less variable as
they depend less on host scheduler decisions.
@@ -33,6 +33,10 @@ static int experimental_zcopytx;module_param(experimental_zcopytx,int,0444);MODULE_PARM_DESC(experimental_zcopytx,"Enable Experimental Zero Copy TX");+staticintworkers=2;+module_param(workers,int,0444);+MODULE_PARM_DESC(workers,"Set the number of worker threads");+/* Max number of bytes transferred before requeueing the job.*Usingthislimitpreventsonevirtqueuefromstarvingothers.*/#define VHOST_NET_WEIGHT 0x80000
@@ -300,7 +303,9 @@ long vhost_dev_init(struct vhost_dev *dev,dev->mm=NULL;spin_lock_init(&dev->work_lock);INIT_LIST_HEAD(&dev->work_list);-dev->worker=NULL;+dev->nworkers=min(nworkers,VHOST_MAX_WORKERS);+for(i=0;i<dev->nworkers;i++)+dev->workers[i]=NULL;for(i=0;i<dev->nvqs;++i){dev->vqs[i].log=NULL;
@@ -354,7 +359,7 @@ static int vhost_attach_cgroups(struct vhost_dev *dev)staticlongvhost_dev_set_owner(structvhost_dev*dev){structtask_struct*worker;-interr;+interr,i;/* Is there an owner already? */if(dev->mm){
@@ -364,28 +369,34 @@ static long vhost_dev_set_owner(struct vhost_dev *dev)/* No owner, become one */dev->mm=get_task_mm(current);-worker=kthread_create(vhost_worker,dev,"vhost-%d",current->pid);-if(IS_ERR(worker)){-err=PTR_ERR(worker);-gotoerr_worker;-}+for(i=0;i<dev->nworkers;i++){+worker=kthread_create(vhost_worker,dev,"vhost-%d.%d",current->pid,i);+if(IS_ERR(worker)){+err=PTR_ERR(worker);+gotoerr_worker;+}-dev->worker=worker;-wake_up_process(worker);/* avoid contributing to loadavg */+dev->workers[i]=worker;+wake_up_process(worker);/* avoid contributing to loadavg */+}err=vhost_attach_cgroups(dev);if(err)-gotoerr_cgroup;+gotoerr_worker;err=vhost_dev_alloc_iovecs(dev);if(err)-gotoerr_cgroup;+gotoerr_worker;return0;-err_cgroup:-kthread_stop(worker);-dev->worker=NULL;+err_worker:+for(i=0;i<dev->nworkers;i++){+if(dev->workers[i]){+kthread_stop(dev->workers[i]);+dev->workers[i]=NULL;+}+}if(dev->mm)mmput(dev->mm);dev->mm=NULL;
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-02-19 14:51:01
On Fri, Feb 17, 2012 at 05:02:06PM -0600, Anthony Liguori wrote:
With workloads that are dominated by very high rates of small packets, we see
considerable overhead in virtio notifications.
The best strategy we've been able to come up with to deal with this is adaptive
polling.
This patch simply adds the infrastructure needed to experiment with
polling strategies. It is not meant for inclusion.
Here are the results with various polling values. The spinning is not currently
a net win due to the high mutex contention caused by the broadcast wakeup. With
a patch attempting to signal wakeup, we see up to 170+ transactions per second
with TCP_RR 60 instance.
N Baseline Spin 0 Spin 1000 Spin 5000
TCP_RR
1 9,639.66 10,164.06 9,825.43 9,827.45 101.95%
10 62,819.55 54,059.78 63,114.30 60,767.23 96.73%
30 84,715.60 131,241.86 120,922.38 89,776.39 105.97%
60 124,614.71 148,720.66 158,678.08 141,400.05 113.47%
UDP_RR
1 9,652.50 10,343.72 9,493.95 9,569.54 99.14%
10 53,830.26 58,235.90 50,145.29 48,820.53 90.69%
30 89,471.01 97,634.53 95,108.34 91,263.65 102.00%
60 103,640.59 164,035.01 157,002.22 128,646.73 124.13%
TCP_STREAM
1 2,622.63 2,610.71 2,688.49 2,678.61 102.13%
4 4,928.02 4,812.05 4,971.00 5,104.57 103.58%
1 5,639.89 5,751.28 5,819.81 5,593.62 99.18%
4 5,874.72 6,575.55 6,324.87 6,502.33 110.68%
1 6,257.42 7,655.22 7,610.52 7,424.74 118.65%
4 5,370.78 6,044.83 5,784.23 6,209.93 115.62%
1 6,346.63 7,267.44 7,567.39 7,677.93 120.98%
4 5,198.02 5,657.12 5,528.94 5,792.42 111.44%
TCP_MAERTS
1 2,091.38 1,765.62 2,142.56 2,312.94 110.59%
4 5,319.52 5,619.49 5,544.50 5,645.81 106.13%
1 7,030.66 7,593.61 7,575.67 7,622.07 108.41%
4 9,040.53 7,275.84 7,322.07 6,681.34 73.90%
1 9,160.93 9,318.15 9,065.82 8,586.82 93.73%
4 9,372.49 8,875.63 8,959.03 9,056.07 96.62%
1 9,183.28 9,134.02 8,945.12 8,657.72 94.28%
4 9,377.17 8,877.52 8,959.54 9,071.53 96.74%
An obvious question would be how are BW divided by CPU
numbers affected.
@@ -37,6 +37,10 @@ static int workers = 2;module_param(workers,int,0444);MODULE_PARM_DESC(workers,"Set the number of worker threads");+staticulongspin_threshold=0;+module_param(spin_threshold,ulong,0444);+MODULE_PARM_DESC(spin_threshold,"The polling threshold for the tx queue");+/* Max number of bytes transferred before requeueing the job.*Usingthislimitpreventsonevirtqueuefromstarvingothers.*/#define VHOST_NET_WEIGHT 0x80000
@@ -149,6 +154,7 @@ static void handle_tx(struct vhost_net *net)size_thdr_size;structsocket*sock;structvhost_ubuf_ref*uninitialized_var(ubufs);+size_tspin_count;boolzcopy;/* TODO: check that we are running from vhost_worker? */
From: Tom Lendacky <hidden> Date: 2012-02-20 15:52:31
"Michael S. Tsirkin" [off-list ref] wrote on 02/19/2012 08:41:45 AM:
From: "Michael S. Tsirkin" <mst@redhat.com>
To: Anthony Liguori/Austin/IBM@IBMUS
Cc: netdev@vger.kernel.org, Tom Lendacky/Austin/IBM@IBMUS, Cristian
Viana [off-list ref]
Date: 02/19/2012 08:42 AM
Subject: Re: [PATCH 1/2] vhost: allow multiple workers threads
On Fri, Feb 17, 2012 at 05:02:05PM -0600, Anthony Liguori wrote:
quoted
This patch allows vhost to have multiple worker threads for devices
such as
quoted
virtio-net which may have multiple virtqueues.
Since virtqueues are a lockless ring queue, in an ideal world data is
being
quoted
produced by the producer as fast as data is being consumed by the
consumer.
quoted
These loops will continue to consume data until none is left.
vhost currently multiplexes the consumer side of the queue on a
single thread
quoted
by attempting to read from the queue until everything is read or it
cannot
quoted
process anymore. This means that activity on one queue may stall
another queue.
There's actually an attempt to address this: look up
VHOST_NET_WEIGHT in the code. I take it, this isn't effective?
quoted
This is exacerbated when using any form of polling to read from
the queues (as
quoted
we'll introduce in the next patch). By spawning a thread per-
virtqueue, this
quoted
is addressed.
The only problem with this patch right now is how the wake up of
the threads is
quoted
done. It's essentially a broadcast and we have seen lock contention as
a
quoted
result.
On which lock?
The mutex lock in the vhost_virtqueue struct. This really shows up when
running with patch 2/2 and increasing the spin_threshold. Both threads wake
up and try to acquire the mutex. As the spin_threshold increases you end
up
with one of the threads getting blocked for a longer and longer time and
unable to do any RX processing that might be needed.
Tom
quoted
We've tried some approaches to signal a single thread but I'm not
confident that that code is correct yet so I'm only sending the
broadcast
quoted
version.
Yes, that looks like an obvious question.
quoted
Here are some performance results from this change. There's a modest
improvement with stream although a fair bit of variability too.
With RR, there's pretty significant improvements as the instance
rate drives up.
Interesting. This was actually tested at one time and we saw
a significant performance improvement from using
a single thread especially with a single
stream in the guest. Profiling indicated that
with a single thread we get too many context
switches between TX and RX, since guest networking
tends to run TX and RX processing on the same
guest VCPU.
Maybe we were wrong or maybe this went away
for some reason. I'll see if this can be reproduced.
One point missing here is the ratio of
throughput divided by host CPU utilization.
The concern being to avoid hurting
heavily loaded systems.
These numbers also tend to be less variable as
they depend less on host scheduler decisions.
@@ -33,6 +33,10 @@ static int experimental_zcopytx;module_param(experimental_zcopytx,int,0444);MODULE_PARM_DESC(experimental_zcopytx,"Enable Experimental Zero Copy
TX");
quoted
+static int workers = 2;
+module_param(workers, int, 0444);
+MODULE_PARM_DESC(workers, "Set the number of worker threads");
+
/* Max number of bytes transferred before requeueing the job.
* Using this limit prevents one virtqueue from starving others. */
#define VHOST_NET_WEIGHT 0x80000
@@ -504,7 +508,7 @@ static int vhost_net_open(struct inode *inode,
struct file *f)
quoted
dev = &n->dev;
n->vqs[VHOST_NET_VQ_TX].handle_kick = handle_tx_kick;
n->vqs[VHOST_NET_VQ_RX].handle_kick = handle_rx_kick;
- r = vhost_dev_init(dev, n->vqs, VHOST_NET_VQ_MAX);
+ r = vhost_dev_init(dev, n->vqs, workers, VHOST_NET_VQ_MAX);
if (r < 0) {
kfree(n);
return r;
}
long vhost_dev_init(struct vhost_dev *dev,
- struct vhost_virtqueue *vqs, int nvqs)
+ struct vhost_virtqueue *vqs, int nworkers, int nvqs)
{
int i;
@@ -300,7 +303,9 @@ long vhost_dev_init(struct vhost_dev *dev, dev->mm = NULL; spin_lock_init(&dev->work_lock); INIT_LIST_HEAD(&dev->work_list);- dev->worker = NULL;+ dev->nworkers = min(nworkers, VHOST_MAX_WORKERS);+ for (i = 0; i < dev->nworkers; i++)+ dev->workers[i] = NULL; for (i = 0; i < dev->nvqs; ++i) { dev->vqs[i].log = NULL;
@@ -354,7 +359,7 @@ static int vhost_attach_cgroups(struct vhost_dev
*dev)
quoted
static long vhost_dev_set_owner(struct vhost_dev *dev)
{
struct task_struct *worker;
- int err;
+ int err, i;
/* Is there an owner already? */
if (dev->mm) {
@@ -364,28 +369,34 @@ static long vhost_dev_set_owner(struct vhost_dev
*dev)
quoted
/* No owner, become one */
dev->mm = get_task_mm(current);
- worker = kthread_create(vhost_worker, dev, "vhost-%d", current->
pid);
quoted
- if (IS_ERR(worker)) {
- err = PTR_ERR(worker);
- goto err_worker;
- }
+ for (i = 0; i < dev->nworkers; i++) {
+ worker = kthread_create(vhost_worker, dev, "vhost-%d.%d",
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-02-20 19:27:10
On Mon, Feb 20, 2012 at 09:50:37AM -0600, Tom Lendacky wrote:
"Michael S. Tsirkin" [off-list ref] wrote on 02/19/2012 08:41:45 AM:
quoted
From: "Michael S. Tsirkin" <mst@redhat.com>
To: Anthony Liguori/Austin/IBM@IBMUS
Cc: netdev@vger.kernel.org, Tom Lendacky/Austin/IBM@IBMUS, Cristian
Viana [off-list ref]
Date: 02/19/2012 08:42 AM
Subject: Re: [PATCH 1/2] vhost: allow multiple workers threads
On Fri, Feb 17, 2012 at 05:02:05PM -0600, Anthony Liguori wrote:
quoted
This patch allows vhost to have multiple worker threads for devices
such as
quoted
quoted
virtio-net which may have multiple virtqueues.
Since virtqueues are a lockless ring queue, in an ideal world data is
being
quoted
quoted
produced by the producer as fast as data is being consumed by the
consumer.
quoted
quoted
These loops will continue to consume data until none is left.
vhost currently multiplexes the consumer side of the queue on a
single thread
quoted
by attempting to read from the queue until everything is read or it
cannot
quoted
quoted
process anymore. This means that activity on one queue may stall
another queue.
There's actually an attempt to address this: look up
VHOST_NET_WEIGHT in the code. I take it, this isn't effective?
quoted
This is exacerbated when using any form of polling to read from
the queues (as
quoted
we'll introduce in the next patch). By spawning a thread per-
virtqueue, this
quoted
is addressed.
The only problem with this patch right now is how the wake up of
the threads is
quoted
done. It's essentially a broadcast and we have seen lock contention as
a
quoted
quoted
result.
On which lock?
The mutex lock in the vhost_virtqueue struct. This really shows up when
running with patch 2/2 and increasing the spin_threshold. Both threads wake
up and try to acquire the mutex. As the spin_threshold increases you end
up
with one of the threads getting blocked for a longer and longer time and
unable to do any RX processing that might be needed.
Tom
Weird, I had the impression each thread handles one vq.
Isn't this the design?
quoted
quoted
We've tried some approaches to signal a single thread but I'm not
confident that that code is correct yet so I'm only sending the
broadcast
quoted
quoted
version.
Yes, that looks like an obvious question.
quoted
Here are some performance results from this change. There's a modest
improvement with stream although a fair bit of variability too.
With RR, there's pretty significant improvements as the instance
rate drives up.
Interesting. This was actually tested at one time and we saw
a significant performance improvement from using
a single thread especially with a single
stream in the guest. Profiling indicated that
with a single thread we get too many context
switches between TX and RX, since guest networking
tends to run TX and RX processing on the same
guest VCPU.
Maybe we were wrong or maybe this went away
for some reason. I'll see if this can be reproduced.
One point missing here is the ratio of
throughput divided by host CPU utilization.
The concern being to avoid hurting
heavily loaded systems.
These numbers also tend to be less variable as
they depend less on host scheduler decisions.
@@ -33,6 +33,10 @@ static int experimental_zcopytx;module_param(experimental_zcopytx,int,0444);MODULE_PARM_DESC(experimental_zcopytx,"Enable Experimental Zero Copy
TX");
quoted
quoted
+static int workers = 2;
+module_param(workers, int, 0444);
+MODULE_PARM_DESC(workers, "Set the number of worker threads");
+
/* Max number of bytes transferred before requeueing the job.
* Using this limit prevents one virtqueue from starving others. */
#define VHOST_NET_WEIGHT 0x80000
@@ -504,7 +508,7 @@ static int vhost_net_open(struct inode *inode,
struct file *f)
quoted
dev = &n->dev;
n->vqs[VHOST_NET_VQ_TX].handle_kick = handle_tx_kick;
n->vqs[VHOST_NET_VQ_RX].handle_kick = handle_rx_kick;
- r = vhost_dev_init(dev, n->vqs, VHOST_NET_VQ_MAX);
+ r = vhost_dev_init(dev, n->vqs, workers, VHOST_NET_VQ_MAX);
if (r < 0) {
kfree(n);
return r;
}
long vhost_dev_init(struct vhost_dev *dev,
- struct vhost_virtqueue *vqs, int nvqs)
+ struct vhost_virtqueue *vqs, int nworkers, int nvqs)
{
int i;
@@ -300,7 +303,9 @@ long vhost_dev_init(struct vhost_dev *dev, dev->mm = NULL; spin_lock_init(&dev->work_lock); INIT_LIST_HEAD(&dev->work_list);- dev->worker = NULL;+ dev->nworkers = min(nworkers, VHOST_MAX_WORKERS);+ for (i = 0; i < dev->nworkers; i++)+ dev->workers[i] = NULL; for (i = 0; i < dev->nvqs; ++i) { dev->vqs[i].log = NULL;
@@ -354,7 +359,7 @@ static int vhost_attach_cgroups(struct vhost_dev
*dev)
quoted
quoted
static long vhost_dev_set_owner(struct vhost_dev *dev)
{
struct task_struct *worker;
- int err;
+ int err, i;
/* Is there an owner already? */
if (dev->mm) {
@@ -364,28 +369,34 @@ static long vhost_dev_set_owner(struct vhost_dev
*dev)
quoted
quoted
/* No owner, become one */
dev->mm = get_task_mm(current);
- worker = kthread_create(vhost_worker, dev, "vhost-%d", current->
pid);
quoted
quoted
- if (IS_ERR(worker)) {
- err = PTR_ERR(worker);
- goto err_worker;
- }
+ for (i = 0; i < dev->nworkers; i++) {
+ worker = kthread_create(vhost_worker, dev, "vhost-%d.%d",
From: Anthony Liguori <hidden> Date: 2012-02-20 19:46:28
On 02/20/2012 01:27 PM, Michael S. Tsirkin wrote:
On Mon, Feb 20, 2012 at 09:50:37AM -0600, Tom Lendacky wrote:
quoted
"Michael S. Tsirkin"[off-list ref] wrote on 02/19/2012 08:41:45 AM:
quoted
From: "Michael S. Tsirkin"<mst@redhat.com>
To: Anthony Liguori/Austin/IBM@IBMUS
Cc: netdev@vger.kernel.org, Tom Lendacky/Austin/IBM@IBMUS, Cristian
Viana[off-list ref]
Date: 02/19/2012 08:42 AM
Subject: Re: [PATCH 1/2] vhost: allow multiple workers threads
On Fri, Feb 17, 2012 at 05:02:05PM -0600, Anthony Liguori wrote:
quoted
This patch allows vhost to have multiple worker threads for devices
such as
quoted
quoted
virtio-net which may have multiple virtqueues.
Since virtqueues are a lockless ring queue, in an ideal world data is
being
quoted
quoted
produced by the producer as fast as data is being consumed by the
consumer.
quoted
quoted
These loops will continue to consume data until none is left.
vhost currently multiplexes the consumer side of the queue on a
single thread
quoted
by attempting to read from the queue until everything is read or it
cannot
quoted
quoted
process anymore. This means that activity on one queue may stall
another queue.
There's actually an attempt to address this: look up
VHOST_NET_WEIGHT in the code. I take it, this isn't effective?
quoted
This is exacerbated when using any form of polling to read from
the queues (as
quoted
we'll introduce in the next patch). By spawning a thread per-
virtqueue, this
quoted
is addressed.
The only problem with this patch right now is how the wake up of
the threads is
quoted
done. It's essentially a broadcast and we have seen lock contention as
a
quoted
quoted
result.
On which lock?
The mutex lock in the vhost_virtqueue struct. This really shows up when
running with patch 2/2 and increasing the spin_threshold. Both threads wake
up and try to acquire the mutex. As the spin_threshold increases you end
up
with one of the threads getting blocked for a longer and longer time and
unable to do any RX processing that might be needed.
Tom
Weird, I had the impression each thread handles one vq.
Isn't this the design?
Not the way the code is structured today. There is a single consumer/producer
work queue and either the vq notification or other actions may get placed on it.
It would be possible to do three threads, one for background tasks and then one
for each queue with a more invasive refactoring.
But I assumed that the reason the code was structured this was originally was
because you saw some value in having a single producer/consumer queue for
everything...
Regards,
Anthony Liguori
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-02-20 21:00:05
On Mon, Feb 20, 2012 at 01:46:03PM -0600, Anthony Liguori wrote:
On 02/20/2012 01:27 PM, Michael S. Tsirkin wrote:
quoted
On Mon, Feb 20, 2012 at 09:50:37AM -0600, Tom Lendacky wrote:
quoted
"Michael S. Tsirkin"[off-list ref] wrote on 02/19/2012 08:41:45 AM:
quoted
From: "Michael S. Tsirkin"<mst@redhat.com>
To: Anthony Liguori/Austin/IBM@IBMUS
Cc: netdev@vger.kernel.org, Tom Lendacky/Austin/IBM@IBMUS, Cristian
Viana[off-list ref]
Date: 02/19/2012 08:42 AM
Subject: Re: [PATCH 1/2] vhost: allow multiple workers threads
On Fri, Feb 17, 2012 at 05:02:05PM -0600, Anthony Liguori wrote:
quoted
This patch allows vhost to have multiple worker threads for devices
such as
quoted
quoted
virtio-net which may have multiple virtqueues.
Since virtqueues are a lockless ring queue, in an ideal world data is
being
quoted
quoted
produced by the producer as fast as data is being consumed by the
consumer.
quoted
quoted
These loops will continue to consume data until none is left.
vhost currently multiplexes the consumer side of the queue on a
single thread
quoted
by attempting to read from the queue until everything is read or it
cannot
quoted
quoted
process anymore. This means that activity on one queue may stall
another queue.
There's actually an attempt to address this: look up
VHOST_NET_WEIGHT in the code. I take it, this isn't effective?
quoted
This is exacerbated when using any form of polling to read from
the queues (as
quoted
we'll introduce in the next patch). By spawning a thread per-
virtqueue, this
quoted
is addressed.
The only problem with this patch right now is how the wake up of
the threads is
quoted
done. It's essentially a broadcast and we have seen lock contention as
a
quoted
quoted
result.
On which lock?
The mutex lock in the vhost_virtqueue struct. This really shows up when
running with patch 2/2 and increasing the spin_threshold. Both threads wake
up and try to acquire the mutex. As the spin_threshold increases you end
up
with one of the threads getting blocked for a longer and longer time and
unable to do any RX processing that might be needed.
Tom
Weird, I had the impression each thread handles one vq.
Isn't this the design?
Not the way the code is structured today. There is a single
consumer/producer work queue and either the vq notification or other
actions may get placed on it.
And then a random thread picks it up?
wont this cause packet reordering?
I'll go reread.
It would be possible to do three threads, one for background tasks
and then one for each queue with a more invasive refactoring.
But I assumed that the reason the code was structured this was
originally was because you saw some value in having a single
producer/consumer queue for everything...
Regards,
Anthony Liguori
The point was really to avoid scheduler overhead
as with tcp, tx and rx tend to run on the same cpu.
From: Shirley Ma <hidden> Date: 2012-02-21 01:04:18
On Mon, 2012-02-20 at 23:00 +0200, Michael S. Tsirkin wrote:
The point was really to avoid scheduler overhead
as with tcp, tx and rx tend to run on the same cpu.
We have tried different approaches in the past, like splitting vhost
thread to separate TX, RX threads; create per cpu vhost thread instead
of creating per VM per virtio_net vhost thread...
We think per cpu vhost thread is a better approach based on the data we
have collected. It will reduce both vhost resource and scheduler
overhead. It will not depend on host scheduler, has less various. The
patch is under testing, we hope we can post it soon.
Thanks
Shirley
From: Shirley Ma <hidden> Date: 2012-02-21 01:35:39
We tried similar approach before by using a minimum timer for handle_tx
to stay in the loop to accumulate more packets before enabling the guest
notification. It did have better TCP_RRs, UDP_RRs results. However, we
think this is just a debug patch. We really need to understand why
handle_tx can't see more packets to process for multiple instances
request/response type of workload first. Spinning in this loop is not a
good solution.
Thanks
Shirley
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-02-21 03:21:49
On Mon, Feb 20, 2012 at 05:04:10PM -0800, Shirley Ma wrote:
On Mon, 2012-02-20 at 23:00 +0200, Michael S. Tsirkin wrote:
quoted
The point was really to avoid scheduler overhead
as with tcp, tx and rx tend to run on the same cpu.
We have tried different approaches in the past, like splitting vhost
thread to separate TX, RX threads; create per cpu vhost thread instead
of creating per VM per virtio_net vhost thread...
We think per cpu vhost thread is a better approach based on the data we
have collected. It will reduce both vhost resource and scheduler
overhead. It will not depend on host scheduler, has less various. The
patch is under testing, we hope we can post it soon.
Thanks
Shirley
Yes, great, this is definitely interesting. I actually started with
a per-cpu one - it did not perform well but I did not
figure out why, switching to a single thread fixed it
and I did not dig into it.
--
MST
From: Jason Wang <hidden> Date: 2012-02-21 04:32:52
On 02/21/2012 03:46 AM, Anthony Liguori wrote:
On 02/20/2012 01:27 PM, Michael S. Tsirkin wrote:
quoted
On Mon, Feb 20, 2012 at 09:50:37AM -0600, Tom Lendacky wrote:
quoted
"Michael S. Tsirkin"[off-list ref] wrote on 02/19/2012 08:41:45 AM:
quoted
From: "Michael S. Tsirkin"<mst@redhat.com>
To: Anthony Liguori/Austin/IBM@IBMUS
Cc: netdev@vger.kernel.org, Tom Lendacky/Austin/IBM@IBMUS, Cristian
Viana[off-list ref]
Date: 02/19/2012 08:42 AM
Subject: Re: [PATCH 1/2] vhost: allow multiple workers threads
On Fri, Feb 17, 2012 at 05:02:05PM -0600, Anthony Liguori wrote:
quoted
This patch allows vhost to have multiple worker threads for devices
such as
quoted
quoted
virtio-net which may have multiple virtqueues.
Since virtqueues are a lockless ring queue, in an ideal world data is
being
quoted
quoted
produced by the producer as fast as data is being consumed by the
consumer.
quoted
quoted
These loops will continue to consume data until none is left.
vhost currently multiplexes the consumer side of the queue on a
single thread
quoted
by attempting to read from the queue until everything is read or it
cannot
quoted
quoted
process anymore. This means that activity on one queue may stall
another queue.
There's actually an attempt to address this: look up
VHOST_NET_WEIGHT in the code. I take it, this isn't effective?
quoted
This is exacerbated when using any form of polling to read from
the queues (as
quoted
we'll introduce in the next patch). By spawning a thread per-
virtqueue, this
quoted
is addressed.
The only problem with this patch right now is how the wake up of
the threads is
quoted
done. It's essentially a broadcast and we have seen lock
contention as
a
quoted
quoted
result.
On which lock?
The mutex lock in the vhost_virtqueue struct. This really shows up
when
running with patch 2/2 and increasing the spin_threshold. Both
threads wake
up and try to acquire the mutex. As the spin_threshold increases
you end
up
with one of the threads getting blocked for a longer and longer time
and
unable to do any RX processing that might be needed.
Tom
Weird, I had the impression each thread handles one vq.
Isn't this the design?
Not the way the code is structured today. There is a single
consumer/producer work queue and either the vq notification or other
actions may get placed on it.
It would be possible to do three threads, one for background tasks and
then one for each queue with a more invasive refactoring.
But I assumed that the reason the code was structured this was
originally was because you saw some value in having a single
producer/consumer queue for everything...
Regards,
Anthony Liguori
Not sure I'm reading the code correctly, looks like with this series,
two worker threads can try to handle the work of a same virtqueue? Looks
strange as the notification from other side (guest/net) should be
disabled even if vhost thread is spinning.
--
To unsubscribe from this list: send the line "unsubscribe netdev" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
From: Jason Wang <hidden> Date: 2012-02-21 04:51:53
On 02/19/2012 10:41 PM, Michael S. Tsirkin wrote:
On Fri, Feb 17, 2012 at 05:02:05PM -0600, Anthony Liguori wrote:
quoted
quoted
This patch allows vhost to have multiple worker threads for devices such as
virtio-net which may have multiple virtqueues.
Since virtqueues are a lockless ring queue, in an ideal world data is being
produced by the producer as fast as data is being consumed by the consumer.
These loops will continue to consume data until none is left.
vhost currently multiplexes the consumer side of the queue on a single thread
by attempting to read from the queue until everything is read or it cannot
process anymore. This means that activity on one queue may stall another queue.
There's actually an attempt to address this: look up
VHOST_NET_WEIGHT in the code. I take it, this isn't effective?
quoted
quoted
This is exacerbated when using any form of polling to read from the queues (as
we'll introduce in the next patch). By spawning a thread per-virtqueue, this
is addressed.
The only problem with this patch right now is how the wake up of the threads is
done. It's essentially a broadcast and we have seen lock contention as a
result.
On which lock?
quoted
quoted
We've tried some approaches to signal a single thread but I'm not
confident that that code is correct yet so I'm only sending the broadcast
version.
Yes, that looks like an obvious question.
quoted
quoted
Here are some performance results from this change. There's a modest
improvement with stream although a fair bit of variability too.
With RR, there's pretty significant improvements as the instance rate drives up.
Interesting. This was actually tested at one time and we saw
a significant performance improvement from using
a single thread especially with a single
stream in the guest. Profiling indicated that
with a single thread we get too many context
switches between TX and RX, since guest networking
tends to run TX and RX processing on the same
guest VCPU.
Maybe we were wrong or maybe this went away
for some reason. I'll see if this can be reproduced.
I've tried a similar test in Jan. The test uses one dedicated vhost
thread to handle tx requests and another one for rx. Test result shows
much degradation as the both of the #exits and #irq are increased. There
are some differences as I test between local host and guest, and the
guest does not have very recent virtio changes ( unlocked kick and
exposing index immediately ). I would try the recent kernel.
From: Jason Wang <hidden> Date: 2012-02-21 05:34:25
On 02/21/2012 09:35 AM, Shirley Ma wrote:
We tried similar approach before by using a minimum timer for handle_tx
to stay in the loop to accumulate more packets before enabling the guest
notification. It did have better TCP_RRs, UDP_RRs results. However, we
think this is just a debug patch. We really need to understand why
handle_tx can't see more packets to process for multiple instances
request/response type of workload first. Spinning in this loop is not a
good solution.
Spinning help for the latency, but looks like we need some adaptive
method to adjust the threshold dynamically such as monitor the minimum
time gap between two packets. For throughput, if we can improve the
batching of small packets we can improve it. I've tired to use event
index to delay the tx kick until a specified number of packets were
batched in the virtqueue. Test shows improvement of throughput in small
packets as the number of #exit were reduced greatly ( the packets/#exit
and cpu utilization were increased), but it damages the performance of
other. This is only for debug, but it confirms that there's something we
need to improve the batching.
Thanks
Shirley
--
To unsubscribe from this list: send the line "unsubscribe netdev" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
From: Shirley Ma <hidden> Date: 2012-02-21 05:42:53
On Tue, 2012-02-21 at 05:21 +0200, Michael S. Tsirkin wrote:
On Mon, Feb 20, 2012 at 05:04:10PM -0800, Shirley Ma wrote:
quoted
On Mon, 2012-02-20 at 23:00 +0200, Michael S. Tsirkin wrote:
quoted
The point was really to avoid scheduler overhead
as with tcp, tx and rx tend to run on the same cpu.
We have tried different approaches in the past, like splitting vhost
thread to separate TX, RX threads; create per cpu vhost thread
instead
quoted
of creating per VM per virtio_net vhost thread...
We think per cpu vhost thread is a better approach based on the data
we
quoted
have collected. It will reduce both vhost resource and scheduler
overhead. It will not depend on host scheduler, has less various.
The
quoted
patch is under testing, we hope we can post it soon.
Thanks
Shirley
Yes, great, this is definitely interesting. I actually started with
a per-cpu one - it did not perform well but I did not
figure out why, switching to a single thread fixed it
and I did not dig into it.
The patch includes per cpu vhost thread & vhost NUMA aware scheduling
It is very interesting. We are collecting performance data with
different workloads (streams, request/response) related to which VCPU
runs on which CPU, which vhost cpu thread is being scheduled, and which
NIC TX/RX queues is being used. The performance were different when
using different vhost scheduling approach for both TX/RX worker. The
results seems pretty good: like 60 UDP_RRs, the results event more than
doubled in our lab. However the TCP_RRs results couldn't catch up
UDP_RRs.
Thanks
Shirley
From: Shirley Ma <hidden> Date: 2012-02-21 06:29:08
On Tue, 2012-02-21 at 13:34 +0800, Jason Wang wrote:
On 02/21/2012 09:35 AM, Shirley Ma wrote:
quoted
We tried similar approach before by using a minimum timer for
handle_tx
quoted
to stay in the loop to accumulate more packets before enabling the
guest
quoted
notification. It did have better TCP_RRs, UDP_RRs results. However,
we
quoted
think this is just a debug patch. We really need to understand why
handle_tx can't see more packets to process for multiple instances
request/response type of workload first. Spinning in this loop is
not a
quoted
good solution.
Spinning help for the latency, but looks like we need some adaptive
method to adjust the threshold dynamically such as monitor the
minimum
time gap between two packets. For throughput, if we can improve the
batching of small packets we can improve it. I've tired to use event
index to delay the tx kick until a specified number of packets were
batched in the virtqueue. Test shows improvement of throughput in
small
packets as the number of #exit were reduced greatly ( the
packets/#exit
and cpu utilization were increased), but it damages the performance
of
other. This is only for debug, but it confirms that there's something
we
need to improve the batching.
Our test case was 60 instances 256/256 bytes tcp_rrs or udp_rrs. In
theory there should be multiple packets in the queue by the time vhost
gets notified, but from debugging output, there was only a few or even
one packet in the queue. So the questions here why the time gap between
two packets is that big?
Shirley
From: Jason Wang <hidden> Date: 2012-02-21 06:38:19
On 02/21/2012 02:28 PM, Shirley Ma wrote:
On Tue, 2012-02-21 at 13:34 +0800, Jason Wang wrote:
quoted
On 02/21/2012 09:35 AM, Shirley Ma wrote:
quoted
We tried similar approach before by using a minimum timer for
handle_tx
quoted
to stay in the loop to accumulate more packets before enabling the
guest
quoted
notification. It did have better TCP_RRs, UDP_RRs results. However,
we
quoted
think this is just a debug patch. We really need to understand why
handle_tx can't see more packets to process for multiple instances
request/response type of workload first. Spinning in this loop is
not a
quoted
good solution.
Spinning help for the latency, but looks like we need some adaptive
method to adjust the threshold dynamically such as monitor the
minimum
time gap between two packets. For throughput, if we can improve the
batching of small packets we can improve it. I've tired to use event
index to delay the tx kick until a specified number of packets were
batched in the virtqueue. Test shows improvement of throughput in
small
packets as the number of #exit were reduced greatly ( the
packets/#exit
and cpu utilization were increased), but it damages the performance
of
other. This is only for debug, but it confirms that there's something
we
need to improve the batching.
Our test case was 60 instances 256/256 bytes tcp_rrs or udp_rrs. In
theory there should be multiple packets in the queue by the time vhost
gets notified, but from debugging output, there was only a few or even
one packet in the queue. So the questions here why the time gap between
two packets is that big?
Shirley
Not sure whether it's related but did you try to disable the nagle
algorithm during the test?
On Tue, 2012-02-21 at 14:38 +0800, Jason Wang wrote:
On 02/21/2012 02:28 PM, Shirley Ma wrote:
quoted
On Tue, 2012-02-21 at 13:34 +0800, Jason Wang wrote:
quoted
On 02/21/2012 09:35 AM, Shirley Ma wrote:
quoted
We tried similar approach before by using a minimum timer for
handle_tx
quoted
to stay in the loop to accumulate more packets before enabling the
guest
quoted
notification. It did have better TCP_RRs, UDP_RRs results. However,
we
quoted
think this is just a debug patch. We really need to understand why
handle_tx can't see more packets to process for multiple instances
request/response type of workload first. Spinning in this loop is
not a
quoted
good solution.
Spinning help for the latency, but looks like we need some adaptive
method to adjust the threshold dynamically such as monitor the
minimum
time gap between two packets. For throughput, if we can improve the
batching of small packets we can improve it. I've tired to use event
index to delay the tx kick until a specified number of packets were
batched in the virtqueue. Test shows improvement of throughput in
small
packets as the number of #exit were reduced greatly ( the
packets/#exit
and cpu utilization were increased), but it damages the performance
of
other. This is only for debug, but it confirms that there's something
we
need to improve the batching.
Our test case was 60 instances 256/256 bytes tcp_rrs or udp_rrs. In
theory there should be multiple packets in the queue by the time vhost
gets notified, but from debugging output, there was only a few or even
one packet in the queue. So the questions here why the time gap between
two packets is that big?
Shirley
Not sure whether it's related but did you try to disable the nagle
algorithm during the test?
This is a 60-instance 256 byte request-response test. So each request
for each instance is independent and is sub-mtu size and the next
request is not sent until the response is received. So Nagle doesn't
delay any packets in this workload.
Thanks
Sridhar
From: Anthony Liguori <hidden> Date: 2012-03-05 13:21:50
On 02/20/2012 10:03 PM, Shirley Ma wrote:
On Tue, 2012-02-21 at 05:21 +0200, Michael S. Tsirkin wrote:
quoted
On Mon, Feb 20, 2012 at 05:04:10PM -0800, Shirley Ma wrote:
quoted
On Mon, 2012-02-20 at 23:00 +0200, Michael S. Tsirkin wrote:
quoted
The point was really to avoid scheduler overhead
as with tcp, tx and rx tend to run on the same cpu.
We have tried different approaches in the past, like splitting vhost
thread to separate TX, RX threads; create per cpu vhost thread
instead
quoted
of creating per VM per virtio_net vhost thread...
We think per cpu vhost thread is a better approach based on the data
we
quoted
have collected. It will reduce both vhost resource and scheduler
overhead. It will not depend on host scheduler, has less various.
The
quoted
patch is under testing, we hope we can post it soon.
Thanks
Shirley
Yes, great, this is definitely interesting. I actually started with
a per-cpu one - it did not perform well but I did not
figure out why, switching to a single thread fixed it
and I did not dig into it.
The patch includes per cpu vhost thread& vhost NUMA aware scheduling
Hi Shirley,
Are you planning on posting these patches soon?
Regards,
Anthony Liguori
From: Shirley Ma <hidden> Date: 2012-03-05 20:43:35
On Mon, 2012-03-05 at 07:21 -0600, Anthony Liguori wrote:
Hi Shirley,
Are you planning on posting these patches soon?
Regards,
Anthony Liguori
I thought to compare with different scheduling with different NICs
before submitting it. Maybe it's better to submit a RFC patch to get
review comments now. Let me clean up the patch and submit it here.
Thanks
Shirley
With workloads that are dominated by very high rates of small packets, we see
considerable overhead in virtio notifications.
The best strategy we've been able to come up with to deal with this is adaptive
polling. This patch simply adds the infrastructure needed to experiment with
polling strategies. It is not meant for inclusion.
Here are the results with various polling values. The spinning is not currently
a net win due to the high mutex contention caused by the broadcast wakeup. With
a patch attempting to signal wakeup, we see up to 170+ transactions per second
with TCP_RR 60 instance.
Do you really consider to add busy loop to the code? There would
probably be another way to solve it.
One of the options we talked about it to add a ethtool-controllable knob
in the guest to set the min/max time for wakeup.
@@ -37,6 +37,10 @@ static int workers = 2;module_param(workers,int,0444);MODULE_PARM_DESC(workers,"Set the number of worker threads");+staticulongspin_threshold=0;+module_param(spin_threshold,ulong,0444);+MODULE_PARM_DESC(spin_threshold,"The polling threshold for the tx queue");+/* Max number of bytes transferred before requeueing the job.*Usingthislimitpreventsonevirtqueuefromstarvingothers.*/#define VHOST_NET_WEIGHT 0x80000
@@ -149,6 +154,7 @@ static void handle_tx(struct vhost_net *net)size_thdr_size;structsocket*sock;structvhost_ubuf_ref*uninitialized_var(ubufs);+size_tspin_count;boolzcopy;/* TODO: check that we are running from vhost_worker? */