... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Jacek Konieczny <hidden> Date: 2014-05-15 13:04:36
On 05/15/14 13:59, Wei Liu wrote:
quoted hunk
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Unfortunately, this seems not enough to fix the problem I have reported
here:
http://lists.xenproject.org/archives/html/xen-devel/2014-05/msg01183.html
The dom0 network still stalls when using rate limiting on a VIF
interface after applying this patch to my 3.14.3 kernel (100% CPU#1
usage in the 'soft interrupts').
Greets,
Jacek
On Thu, May 15, 2014 at 03:04:36PM +0200, Jacek Konieczny wrote:
On 05/15/14 13:59, Wei Liu wrote:
quoted
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Unfortunately, this seems not enough to fix the problem I have reported
here:
http://lists.xenproject.org/archives/html/xen-devel/2014-05/msg01183.html
The dom0 network still stalls when using rate limiting on a VIF
interface after applying this patch to my 3.14.3 kernel (100% CPU#1
usage in the 'soft interrupts').
From: David Vrabel <hidden> Date: 2014-05-15 13:41:00
On 15/05/14 12:59, Wei Liu wrote:
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
No it does not. RING_FINAL_CHECK_FOR_REQUESTS() only advances the event
index if the ring is empty.
This will also result in xenvif_up() failing to properly enable the event.
I think Jacek's bug may be that netback fails to call napi_complete()
when credit is exceeded and there still outstanding requests on the
from-guest ring and thus napi repeatedly polls.
David
quoted hunk
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
On Thu, May 15, 2014 at 03:04:36PM +0200, Jacek Konieczny wrote:
On 05/15/14 13:59, Wei Liu wrote:
quoted
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Unfortunately, this seems not enough to fix the problem I have reported
here:
http://lists.xenproject.org/archives/html/xen-devel/2014-05/msg01183.html
The dom0 network still stalls when using rate limiting on a VIF
interface after applying this patch to my 3.14.3 kernel (100% CPU#1
usage in the 'soft interrupts').
Oh I was looking at the aggregated stats that's why the number looked
normal to me. :-/
On Thu, May 15, 2014 at 02:40:58PM +0100, David Vrabel wrote:
On 15/05/14 12:59, Wei Liu wrote:
quoted
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
No it does not. RING_FINAL_CHECK_FOR_REQUESTS() only advances the event
index if the ring is empty.
This will also result in xenvif_up() failing to properly enable the event.
I think Jacek's bug may be that netback fails to call napi_complete()
when credit is exceeded and there still outstanding requests on the
from-guest ring and thus napi repeatedly polls.
Correct. We should call napi_complete if this vif is rate limited.
Wei.
On Thu, May 15, 2014 at 03:04:36PM +0200, Jacek Konieczny wrote:
On 05/15/14 13:59, Wei Liu wrote:
quoted
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Unfortunately, this seems not enough to fix the problem I have reported
here:
http://lists.xenproject.org/archives/html/xen-devel/2014-05/msg01183.html
The dom0 network still stalls when using rate limiting on a VIF
interface after applying this patch to my 3.14.3 kernel (100% CPU#1
usage in the 'soft interrupts').
This is a patch for 3.14.4. I've tested it myself (and looking at the
right stats!) to confirm it works.
---8<---
From a4afed6c44027afff82d6fa7503faef83b01fffe Mon Sep 17 00:00:00 2001
From: Wei Liu <redacted>
Date: Thu, 15 May 2014 15:02:55 +0100
Subject: [PATCH] xen-netback: call napi_complete if vif is rate limited
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Cc: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
Cc: David Vrabel <redacted>
---
drivers/net/xen-netback/common.h | 2 +-
drivers/net/xen-netback/interface.c | 5 +++--
drivers/net/xen-netback/netback.c | 12 ++++++++----
3 files changed, 12 insertions(+), 7 deletions(-)
@@ -219,7 +219,7 @@ void xenvif_check_rx_xenvif(struct xenvif *vif);/* Prevent the device from generating any further traffic. */voidxenvif_carrier_off(structxenvif*vif);-intxenvif_tx_action(structxenvif*vif,intbudget);+intxenvif_tx_action(structxenvif*vif,intbudget,bool*rate_limited);intxenvif_kthread(void*data);voidxenvif_kick_thread(structxenvif*vif);
@@ -61,6 +61,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget){structxenvif*vif=container_of(napi,structxenvif,napi);intwork_done;+boolrate_limited;/* This vif is rogue, we pretend we've there is nothing to do*forthisviftodescheduleitfromNAPI.Butthisinterface
@@ -71,7 +72,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget)return0;}-work_done=xenvif_tx_action(vif,budget);+work_done=xenvif_tx_action(vif,budget,&rate_limited);if(work_done<budget){intmore_to_do=0;
@@ -96,7 +97,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget)local_irq_save(flags);RING_FINAL_CHECK_FOR_REQUESTS(&vif->tx,more_to_do);-if(!more_to_do)+if(!more_to_do||rate_limited)__napi_complete(napi);local_irq_restore(flags);
@@ -1382,7 +1386,7 @@ static int xenvif_tx_submit(struct xenvif *vif)}/* Called after netfront has transmitted */-intxenvif_tx_action(structxenvif*vif,intbudget)+intxenvif_tx_action(structxenvif*vif,intbudget,bool*rate_limited){unsignednr_gops;intwork_done;
@@ -1390,7 +1394,7 @@ int xenvif_tx_action(struct xenvif *vif, int budget)if(unlikely(!tx_work_todo(vif)))return0;-nr_gops=xenvif_tx_build_gops(vif,budget);+nr_gops=xenvif_tx_build_gops(vif,budget,rate_limited);if(nr_gops==0)return0;
From: Zoltan Kiss <hidden> Date: 2014-05-15 14:47:42
On 15/05/14 15:13, Wei Liu wrote:
quoted hunk
On Thu, May 15, 2014 at 03:04:36PM +0200, Jacek Konieczny wrote:
quoted
On 05/15/14 13:59, Wei Liu wrote:
quoted
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Unfortunately, this seems not enough to fix the problem I have reported
here:
http://lists.xenproject.org/archives/html/xen-devel/2014-05/msg01183.html
The dom0 network still stalls when using rate limiting on a VIF
interface after applying this patch to my 3.14.3 kernel (100% CPU#1
usage in the 'soft interrupts').
This is a patch for 3.14.4. I've tested it myself (and looking at the
right stats!) to confirm it works.
---8<---
From a4afed6c44027afff82d6fa7503faef83b01fffe Mon Sep 17 00:00:00 2001
From: Wei Liu <redacted>
Date: Thu, 15 May 2014 15:02:55 +0100
Subject: [PATCH] xen-netback: call napi_complete if vif is rate limited
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Cc: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
Cc: David Vrabel <redacted>
---
drivers/net/xen-netback/common.h | 2 +-
drivers/net/xen-netback/interface.c | 5 +++--
drivers/net/xen-netback/netback.c | 12 ++++++++----
3 files changed, 12 insertions(+), 7 deletions(-)
@@ -219,7 +219,7 @@ void xenvif_check_rx_xenvif(struct xenvif *vif);/* Prevent the device from generating any further traffic. */voidxenvif_carrier_off(structxenvif*vif);-intxenvif_tx_action(structxenvif*vif,intbudget);+intxenvif_tx_action(structxenvif*vif,intbudget,bool*rate_limited);intxenvif_kthread(void*data);voidxenvif_kick_thread(structxenvif*vif);
@@ -61,6 +61,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget){structxenvif*vif=container_of(napi,structxenvif,napi);intwork_done;+boolrate_limited;/* This vif is rogue, we pretend we've there is nothing to do*forthisviftodescheduleitfromNAPI.Butthisinterface
@@ -71,7 +72,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget)return0;}-work_done=xenvif_tx_action(vif,budget);+work_done=xenvif_tx_action(vif,budget,&rate_limited);if(work_done<budget){intmore_to_do=0;
@@ -96,7 +97,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget)local_irq_save(flags);RING_FINAL_CHECK_FOR_REQUESTS(&vif->tx,more_to_do);-if(!more_to_do)+if(!more_to_do||rate_limited)
How about calling timer_pending(&vif->credit_timeout) instead?
Also, can this __napi_complete and the callback's napi_schedule race
with each other? When napi_complete is between removing from the list
and clearing the bit, and napi_schedule is just test&set the bit, the
latter won't add the instance to the list again
@@ -1382,7 +1386,7 @@ static int xenvif_tx_submit(struct xenvif *vif)}/* Called after netfront has transmitted */-intxenvif_tx_action(structxenvif*vif,intbudget)+intxenvif_tx_action(structxenvif*vif,intbudget,bool*rate_limited){unsignednr_gops;intwork_done;
@@ -1390,7 +1394,7 @@ int xenvif_tx_action(struct xenvif *vif, int budget)if(unlikely(!tx_work_todo(vif)))return0;-nr_gops=xenvif_tx_build_gops(vif,budget);+nr_gops=xenvif_tx_build_gops(vif,budget,rate_limited);if(nr_gops==0)return0;
On Thu, May 15, 2014 at 03:47:38PM +0100, Zoltan Kiss wrote:
On 15/05/14 15:13, Wei Liu wrote:
quoted
On Thu, May 15, 2014 at 03:04:36PM +0200, Jacek Konieczny wrote:
quoted
On 05/15/14 13:59, Wei Liu wrote:
quoted
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Unfortunately, this seems not enough to fix the problem I have reported
here:
http://lists.xenproject.org/archives/html/xen-devel/2014-05/msg01183.html
The dom0 network still stalls when using rate limiting on a VIF
interface after applying this patch to my 3.14.3 kernel (100% CPU#1
usage in the 'soft interrupts').
This is a patch for 3.14.4. I've tested it myself (and looking at the
right stats!) to confirm it works.
---8<---
From a4afed6c44027afff82d6fa7503faef83b01fffe Mon Sep 17 00:00:00 2001
From: Wei Liu <redacted>
Date: Thu, 15 May 2014 15:02:55 +0100
Subject: [PATCH] xen-netback: call napi_complete if vif is rate limited
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Cc: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
Cc: David Vrabel <redacted>
---
drivers/net/xen-netback/common.h | 2 +-
drivers/net/xen-netback/interface.c | 5 +++--
drivers/net/xen-netback/netback.c | 12 ++++++++----
3 files changed, 12 insertions(+), 7 deletions(-)
@@ -219,7 +219,7 @@ void xenvif_check_rx_xenvif(struct xenvif *vif);/* Prevent the device from generating any further traffic. */voidxenvif_carrier_off(structxenvif*vif);-intxenvif_tx_action(structxenvif*vif,intbudget);+intxenvif_tx_action(structxenvif*vif,intbudget,bool*rate_limited);intxenvif_kthread(void*data);voidxenvif_kick_thread(structxenvif*vif);
@@ -61,6 +61,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget){structxenvif*vif=container_of(napi,structxenvif,napi);intwork_done;+boolrate_limited;/* This vif is rogue, we pretend we've there is nothing to do*forthisviftodescheduleitfromNAPI.Butthisinterface
@@ -71,7 +72,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget)return0;}-work_done=xenvif_tx_action(vif,budget);+work_done=xenvif_tx_action(vif,budget,&rate_limited);if(work_done<budget){intmore_to_do=0;
@@ -96,7 +97,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget)local_irq_save(flags);RING_FINAL_CHECK_FOR_REQUESTS(&vif->tx,more_to_do);-if(!more_to_do)+if(!more_to_do||rate_limited)
How about calling timer_pending(&vif->credit_timeout) instead?
timer_pending(&vif->credit_timeout) covers only one of two senarios of
"credit exceeded", see tx_credit_exceeded.
Also, can this __napi_complete and the callback's napi_schedule race with
each other? When napi_complete is between removing from the list and
clearing the bit, and napi_schedule is just test&set the bit, the latter
won't add the instance to the list again
I think it should be fine. How is it different from what we already have
now? Is this something similar to what David once posted?
[off-list ref]
Wei.
From: Zoltan Kiss <hidden> Date: 2014-05-15 16:34:13
On 15/05/14 16:30, Wei Liu wrote:
On Thu, May 15, 2014 at 03:47:38PM +0100, Zoltan Kiss wrote:
quoted
On 15/05/14 15:13, Wei Liu wrote:
quoted
On Thu, May 15, 2014 at 03:04:36PM +0200, Jacek Konieczny wrote:
quoted
On 05/15/14 13:59, Wei Liu wrote:
quoted
... otherwise the frontend will try to send TX event all the time, even
if no progress can be made. The pointer should only be advanced by the
routine that actually processes the ring (that is, xenvif_poll).
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Acked-by: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
---
drivers/net/xen-netback/netback.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Unfortunately, this seems not enough to fix the problem I have reported
here:
http://lists.xenproject.org/archives/html/xen-devel/2014-05/msg01183.html
The dom0 network still stalls when using rate limiting on a VIF
interface after applying this patch to my 3.14.3 kernel (100% CPU#1
usage in the 'soft interrupts').
This is a patch for 3.14.4. I've tested it myself (and looking at the
right stats!) to confirm it works.
---8<---
From a4afed6c44027afff82d6fa7503faef83b01fffe Mon Sep 17 00:00:00 2001
From: Wei Liu <redacted>
Date: Thu, 15 May 2014 15:02:55 +0100
Subject: [PATCH] xen-netback: call napi_complete if vif is rate limited
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Wei Liu <redacted>
Cc: Ian Campbell <redacted>
Cc: Paul Durrant <redacted>
Cc: David Vrabel <redacted>
---
drivers/net/xen-netback/common.h | 2 +-
drivers/net/xen-netback/interface.c | 5 +++--
drivers/net/xen-netback/netback.c | 12 ++++++++----
3 files changed, 12 insertions(+), 7 deletions(-)
@@ -219,7 +219,7 @@ void xenvif_check_rx_xenvif(struct xenvif *vif);/* Prevent the device from generating any further traffic. */voidxenvif_carrier_off(structxenvif*vif);-intxenvif_tx_action(structxenvif*vif,intbudget);+intxenvif_tx_action(structxenvif*vif,intbudget,bool*rate_limited);intxenvif_kthread(void*data);voidxenvif_kick_thread(structxenvif*vif);
@@ -61,6 +61,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget){structxenvif*vif=container_of(napi,structxenvif,napi);intwork_done;+boolrate_limited;/* This vif is rogue, we pretend we've there is nothing to do*forthisviftodescheduleitfromNAPI.Butthisinterface
@@ -71,7 +72,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget)return0;}-work_done=xenvif_tx_action(vif,budget);+work_done=xenvif_tx_action(vif,budget,&rate_limited);if(work_done<budget){intmore_to_do=0;
@@ -96,7 +97,7 @@ static int xenvif_poll(struct napi_struct *napi, int budget)local_irq_save(flags);RING_FINAL_CHECK_FOR_REQUESTS(&vif->tx,more_to_do);-if(!more_to_do)+if(!more_to_do||rate_limited)
How about calling timer_pending(&vif->credit_timeout) instead?
timer_pending(&vif->credit_timeout) covers only one of two senarios of
"credit exceeded", see tx_credit_exceeded.
The other scenario is when the packet size exceeds the credit. There is
no packet here actually, we just want to know if this vif ran out of
credit and waiting for the timer to fire.
quoted
Also, can this __napi_complete and the callback's napi_schedule race with
each other? When napi_complete is between removing from the list and
clearing the bit, and napi_schedule is just test&set the bit, the latter
won't add the instance to the list again
I think it should be fine. How is it different from what we already have
now? Is this something similar to what David once posted?
[off-list ref]
Unfortunately that discussion stalled, and my question were not
answered, so I bumped it again. But that's different a bit: it was about
racing between the NAPI instance (running in softirq context) and the
interrupt. Here the danger is that the NAPI instance and the softirq can
race. They both run in softirq context, and even if they were originally
on the same CPU, I'm sure if the instance move somewhere else, the timer
doesn't follow it.
Zoli
On Thu, May 15, 2014 at 05:34:09PM +0100, Zoltan Kiss wrote:
[...]
quoted
quoted
quoted
RING_FINAL_CHECK_FOR_REQUESTS(&vif->tx, more_to_do);
- if (!more_to_do)
+ if (!more_to_do || rate_limited)
How about calling timer_pending(&vif->credit_timeout) instead?
timer_pending(&vif->credit_timeout) covers only one of two senarios of
"credit exceeded", see tx_credit_exceeded.
The other scenario is when the packet size exceeds the credit. There is no
packet here actually, we just want to know if this vif ran out of credit and
waiting for the timer to fire.
quoted
Which place are you referring to? There's packet in the ring, right? So
you're saying in xenvif_poll "more_to_do" is true and "timer_pending" is
also true when we come to xenvif_poll again?
quoted
quoted
Also, can this __napi_complete and the callback's napi_schedule race with
each other? When napi_complete is between removing from the list and
clearing the bit, and napi_schedule is just test&set the bit, the latter
won't add the instance to the list again
I think it should be fine. How is it different from what we already have
now? Is this something similar to what David once posted?
[off-list ref]
Unfortunately that discussion stalled, and my question were not answered, so
I bumped it again. But that's different a bit: it was about racing between
the NAPI instance (running in softirq context) and the interrupt. Here the
danger is that the NAPI instance and the softirq can race. They both run in
softirq context, and even if they were originally on the same CPU, I'm sure
if the instance move somewhere else, the timer doesn't follow it.
This comes back to that original question, doesn't it? That's NAPI
running on CPU A and raised by CPU B.
Wei.
From: Zoltan Kiss <hidden> Date: 2014-05-15 17:03:18
On 15/05/14 17:53, Wei Liu wrote:
On Thu, May 15, 2014 at 05:34:09PM +0100, Zoltan Kiss wrote:
[...]
quoted
quoted
quoted
quoted
RING_FINAL_CHECK_FOR_REQUESTS(&vif->tx, more_to_do);
- if (!more_to_do)
+ if (!more_to_do || rate_limited)
How about calling timer_pending(&vif->credit_timeout) instead?
timer_pending(&vif->credit_timeout) covers only one of two senarios of
"credit exceeded", see tx_credit_exceeded.
The other scenario is when the packet size exceeds the credit. There is no
packet here actually, we just want to know if this vif ran out of credit and
waiting for the timer to fire.
quoted
Which place are you referring to? There's packet in the ring, right? So
you're saying in xenvif_poll "more_to_do" is true and "timer_pending" is
also true when we come to xenvif_poll again?
The goal of this patch to deschedule NAPI if the vif ran out of credit.
Either you can carry that information from build_gops via a bool, or you
can check whether the timer is pending. That's what tx_credit_exceeded
does as well, and then it checks if the actual packet fits in. But in
xenvif_poll you don't want to know whether an actual packet fits in, you
only need the information whether tx_credit_exceeded started the timer
or not.
If it is, you can be sure there is no more credit. If not, you can keep
the instance running.
quoted
quoted
quoted
Also, can this __napi_complete and the callback's napi_schedule race with
each other? When napi_complete is between removing from the list and
clearing the bit, and napi_schedule is just test&set the bit, the latter
won't add the instance to the list again
I think it should be fine. How is it different from what we already have
now? Is this something similar to what David once posted?
[off-list ref]
Unfortunately that discussion stalled, and my question were not answered, so
I bumped it again. But that's different a bit: it was about racing between
the NAPI instance (running in softirq context) and the interrupt. Here the
danger is that the NAPI instance and the softirq can race. They both run in
softirq context, and even if they were originally on the same CPU, I'm sure
if the instance move somewhere else, the timer doesn't follow it.
This comes back to that original question, doesn't it? That's NAPI
running on CPU A and raised by CPU B.
Wei.
On Thu, May 15, 2014 at 06:03:14PM +0100, Zoltan Kiss wrote:
On 15/05/14 17:53, Wei Liu wrote:
quoted
On Thu, May 15, 2014 at 05:34:09PM +0100, Zoltan Kiss wrote:
[...]
quoted
quoted
quoted
quoted
RING_FINAL_CHECK_FOR_REQUESTS(&vif->tx, more_to_do);
- if (!more_to_do)
+ if (!more_to_do || rate_limited)
How about calling timer_pending(&vif->credit_timeout) instead?
timer_pending(&vif->credit_timeout) covers only one of two senarios of
"credit exceeded", see tx_credit_exceeded.
The other scenario is when the packet size exceeds the credit. There is no
packet here actually, we just want to know if this vif ran out of credit and
waiting for the timer to fire.
quoted
Which place are you referring to? There's packet in the ring, right? So
you're saying in xenvif_poll "more_to_do" is true and "timer_pending" is
also true when we come to xenvif_poll again?
The goal of this patch to deschedule NAPI if the vif ran out of credit.
Either you can carry that information from build_gops via a bool, or you can
check whether the timer is pending. That's what tx_credit_exceeded does as
well, and then it checks if the actual packet fits in. But in xenvif_poll
you don't want to know whether an actual packet fits in, you only need the
information whether tx_credit_exceeded started the timer or not.
If it is, you can be sure there is no more credit. If not, you can keep the
instance running.
quoted
OK, you convinced me. I will try to rework this patch with your
approach.
Wei.