NAPI note (was Re: lockups with 2.4.20 (tg3? net/core/dev.c|deliver_to_old_ones))

6 messages, 4 authors, 2003-02-19 · open the first message on its own page

NAPI note (was Re: lockups with 2.4.20 (tg3? net/core/dev.c|deliver_to_old_ones))

From: Jeff Garzik <hidden>
Date: 2003-02-14 23:58:13

Manfred Spraul wrote:
It seems to be a generic NAPI restriction:
The caller of netif_receive_skb() must not own a spinlock that is 
acquired from an interrupt handler.

Thanks much for noticing this, Manfred.  tg3 is definitely buggy in this 
regard.  I've CC'd netdev as an FYI...  We should probably patch 
NAPI_HOWTO for this note.

I note that David pointed this out as an area for improvement, so he was 
already thinking in this direction anyway :)

	Jeff

Re: NAPI note

From: David S. Miller <hidden>
Date: 2003-02-18 02:57:19

   From: Jeff Garzik [off-list ref]
   Date: Fri, 14 Feb 2003 18:58:13 -0500

   Manfred Spraul wrote:
   > It seems to be a generic NAPI restriction:
   > The caller of netif_receive_skb() must not own a spinlock that is 
   > acquired from an interrupt handler.
   
   Thanks much for noticing this, Manfred.

I think this logic is buggy.

In the example I've seen posted, only a NAPI implementation bug
could cause the situation to occur.

If cpu1 is in ->poll() for the driver, then by definition the
device shall not cause interrupts.  The device's interrupts
are disabled before we enter the ->poll() handler, and as such
the "cpu2 take device interrupt and takes driver->lock" cannot
occur.

If anything we've found a bug in interrupt disabling in the tg3
driver.

Re: NAPI note

From: Manfred Spraul <hidden>
Date: 2003-02-18 16:31:20

David S. Miller wrote:
  From: Jeff Garzik [off-list ref]
  Date: Fri, 14 Feb 2003 18:58:13 -0500

  Manfred Spraul wrote:
  > It seems to be a generic NAPI restriction:
  > The caller of netif_receive_skb() must not own a spinlock that is 
  > acquired from an interrupt handler.
  
  Thanks much for noticing this, Manfred.

I think this logic is buggy.

In the example I've seen posted, only a NAPI implementation bug
could cause the situation to occur.

If cpu1 is in ->poll() for the driver, then by definition the
device shall not cause interrupts.  The device's interrupts
are disabled before we enter the ->poll() handler, and as such
the "cpu2 take device interrupt and takes driver->lock" cannot
occur.
 
No. I think the rule is that drivers that use the NAPI interface must 
not cause interrupts for packet receive and out-of-rx-buffers conditions.
But what about media error interrupts, or tx interrupts? Or MIB counter 
overflow, etc. What about shared pci interrupts?
All of them could occur, and if they take a spinlock that is held across 
netif_receive_skb(), then it can deadlock.

OTHO if it's guaranteed that no interrupt occurs, then the nic should 
not take a spinlock at all and rely on the synchronization provided by 
NAPI. (->poll is single-threaded).

--
    Manfred

Re: NAPI note

From: jamal <hidden>
Date: 2003-02-19 02:53:12


On Tue, 18 Feb 2003, Manfred Spraul wrote:
David S. Miller wrote:
quoted
  From: Jeff Garzik [off-list ref]
  Date: Fri, 14 Feb 2003 18:58:13 -0500

  Manfred Spraul wrote:
  > It seems to be a generic NAPI restriction:
  > The caller of netif_receive_skb() must not own a spinlock that is
  > acquired from an interrupt handler.

  Thanks much for noticing this, Manfred.

I think this logic is buggy.

In the example I've seen posted, only a NAPI implementation bug
could cause the situation to occur.

If cpu1 is in ->poll() for the driver, then by definition the
device shall not cause interrupts.  The device's interrupts
are disabled before we enter the ->poll() handler, and as such
the "cpu2 take device interrupt and takes driver->lock" cannot
occur.
No. I think the rule is that drivers that use the NAPI interface must
not cause interrupts for packet receive and out-of-rx-buffers conditions.
Ah, but that is only one of two rules.
Theres other drivers which dont follow this rule and just shutdown
all interupt sources. I know that the e1000 for example does this.
I am not sure about the tg3. I think the doc says this but may not
emphasize it as strongly.
So if tg3 uses method 2 then its as Dave says - a bug.
But what about media error interrupts, or tx interrupts? Or MIB counter
overflow, etc. What about shared pci interrupts?
Shared interupts should be interesting actually.
However if you are in poll mode and you receive an interupt you should be
able to quickly determine its not yours without much effect on shared
locks, no?
All of them could occur, and if they take a spinlock that is held across
netif_receive_skb(), then it can deadlock.
yes this could happen with method 1 of programming the driver; however,
tx, receive, link are essentially separate threads and would hardly share
locks.
OTHO if it's guaranteed that no interrupt occurs, then the nic should
not take a spinlock at all and rely on the synchronization provided by
NAPI. (->poll is single-threaded).
i havent studied the e1000 theres a lot of this happening already.
I dont think you need say to protect the tx ring for example from
tx completion interupts vs regular softirq path.

cheers,
jamal

Re: NAPI note

From: Jeff Garzik <hidden>
Date: 2003-02-19 03:14:55

jamal wrote:
On Tue, 18 Feb 2003, Manfred Spraul wrote:

quoted
David S. Miller wrote:

quoted
 From: Jeff Garzik [off-list ref]
 Date: Fri, 14 Feb 2003 18:58:13 -0500

 Manfred Spraul wrote:
 > It seems to be a generic NAPI restriction:
 > The caller of netif_receive_skb() must not own a spinlock that is
 > acquired from an interrupt handler.

 Thanks much for noticing this, Manfred.

I think this logic is buggy.

In the example I've seen posted, only a NAPI implementation bug
could cause the situation to occur.

If cpu1 is in ->poll() for the driver, then by definition the
device shall not cause interrupts.  The device's interrupts
are disabled before we enter the ->poll() handler, and as such
the "cpu2 take device interrupt and takes driver->lock" cannot
occur.
No. I think the rule is that drivers that use the NAPI interface must
not cause interrupts for packet receive and out-of-rx-buffers conditions.

Ah, but that is only one of two rules.
Theres other drivers which dont follow this rule and just shutdown
all interupt sources. I know that the e1000 for example does this.
I am not sure about the tg3. I think the doc says this but may not
emphasize it as strongly.
So if tg3 uses method 2 then its as Dave says - a bug.
tg3 shuts down all interrupt sources, and handles all interrupt events 
in dev->poll().

David and I hashed it out a bit on IRC.  The problem is that 
deliver_to_old_ones() waits, and thus the deadlock that Manfred 
described.  For 2.4.x, the solution is simply to avoid the deadlock in 
the driver.  For 2.5.x, David hinted that deliver_to_old_ones() may be 
going away.

quoted
But what about media error interrupts, or tx interrupts? Or MIB counter
overflow, etc. What about shared pci interrupts?

Shared interupts should be interesting actually.
However if you are in poll mode and you receive an interupt you should be
able to quickly determine its not yours without much effect on shared
locks, no?
Normally, yes.  However tg3 grabs a lock just about anytime it does 
anything.  ;-)  A long term project of mine is to slowly remove these 
locks, but that must wait until the driver stabilizes, and is overall a 
long process.  Most of the locks _are_ removeable, but we keep hit 
deadlock bugs like this, and hardware bugs which need workarounds, so 
those come first.

quoted
All of them could occur, and if they take a spinlock that is held across
netif_receive_skb(), then it can deadlock.

yes this could happen with method 1 of programming the driver; however,
tx, receive, link are essentially separate threads and would hardly share
locks.
They do in tg3's case.

The locks can be removed eventually, but such is the state of life right 
now.

	Jeff

Re: NAPI note

From: David S. Miller <hidden>
Date: 2003-02-19 03:18:41

   From: jamal [off-list ref]
   Date: Tue, 18 Feb 2003 21:53:12 -0500 (EST)
   
   Theres other drivers which dont follow this rule and just shutdown
   all interupt sources. I know that the e1000 for example does this.
   I am not sure about the tg3. I think the doc says this but may not
   emphasize it as strongly.
   So if tg3 uses method 2 then its as Dave says - a bug.
   
Right, but I forgot that it's not a bug in the shared interrupt
case where we need to grab a lock to access the hardware and
fetch the interrupt status.

   Shared interupts should be interesting actually.
   However if you are in poll mode and you receive an interupt you should be
   able to quickly determine its not yours without much effect on shared
   locks, no?

As Jeff has responded already, often you do need locks to do this
sanely.

In tg3, it's really complicated even though the chip writes the
interrupt status to a piece of memory shared with the cpu.

Any time you want to enable/disable tg3 chip interrupts you must
flip and/or check a status bit in this piece of memory.  So it
really needs a lock.

I personally think Jeff is overly optimistic about lock removal in
the driver.  :-)  The last time I attempted to be clever here, we
ended up with all sorts of deadlocks in tg3.  It requires real brain
time and heavy testing to make any kinds of changes in this area.
I also don't want to accomplish this by splitting up into seperate
lines of development of tg3, that's nuts as it would split up our
testing and make both lines get less testing than a unified mainline
driver would (which is what we happily do now).
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help