Thread (9 messages) flat view 9 messages, 4 authors, 2021-07-20

RE: [PATCH] qed: fix possible unpaired spin_{un}lock_bh in _qed_mcp_cmd_and_union()

From: Justin He <hidden>
Date: 2021-07-20 02:10:44
Also in: lkml

-----Original Message-----
From: Prabhakar Kushwaha <redacted>
Sent: Monday, July 19, 2021 10:51 PM
To: Justin He <redacted>
Cc: Ariel Elior <redacted>; GR-everest-linux-l2@marvell.com;
David S. Miller [off-list ref]; Jakub Kicinski [off-list ref];
netdev@vger.kernel.org; Linux Kernel Mailing List <linux-
kernel@vger.kernel.org>; nd [off-list ref]; Shai Malin [off-list ref];
Shai Malin [off-list ref]; Prabhakar Kushwaha [off-list ref]
Subject: Re: [PATCH] qed: fix possible unpaired spin_{un}lock_bh in
_qed_mcp_cmd_and_union()

Hi Justin,

On Mon, Jul 19, 2021 at 6:47 PM Justin He [off-list ref] wrote:
quoted
Hi Prabhakar
quoted
-----Original Message-----
From: Prabhakar Kushwaha <redacted>
Sent: Monday, July 19, 2021 6:36 PM
To: Justin He <redacted>
Cc: Ariel Elior <redacted>; GR-everest-linux-l2@marvell.com;
David S. Miller [off-list ref]; Jakub Kicinski [off-list ref];
netdev@vger.kernel.org; Linux Kernel Mailing List <linux-
kernel@vger.kernel.org>; nd [off-list ref]; Shai Malin
[off-list ref];
quoted
quoted
Shai Malin [off-list ref]; Prabhakar Kushwaha
[off-list ref]
quoted
quoted
Subject: Re: [PATCH] qed: fix possible unpaired spin_{un}lock_bh in
_qed_mcp_cmd_and_union()

Hi Jia,

On Thu, Jul 15, 2021 at 2:28 PM Jia He [off-list ref] wrote:
quoted
Liajian reported a bug_on hit on a ThunderX2 arm64 server with
FastLinQ
quoted
quoted
quoted
QL41000 ethernet controller:
 BUG: scheduling while atomic: kworker/0:4/531/0x00000200
  [qed_probe:488()]hw prepare failed
  kernel BUG at mm/vmalloc.c:2355!
  Internal error: Oops - BUG: 0 [#1] SMP
  CPU: 0 PID: 531 Comm: kworker/0:4 Tainted: G W 5.4.0-77-generic
#86-
quoted
quoted
Ubuntu
quoted
  pstate: 00400009 (nzcv daif +PAN -UAO)
 Call trace:
  vunmap+0x4c/0x50
  iounmap+0x48/0x58
  qed_free_pci+0x60/0x80 [qed]
  qed_probe+0x35c/0x688 [qed]
  __qede_probe+0x88/0x5c8 [qede]
  qede_probe+0x60/0xe0 [qede]
  local_pci_probe+0x48/0xa0
  work_for_cpu_fn+0x24/0x38
  process_one_work+0x1d0/0x468
  worker_thread+0x238/0x4e0
  kthread+0xf0/0x118
  ret_from_fork+0x10/0x18

In this case, qed_hw_prepare() returns error due to hw/fw error, but
in
quoted
quoted
quoted
theory work queue should be in process context instead of interrupt.

The root cause might be the unpaired spin_{un}lock_bh() in
_qed_mcp_cmd_and_union(), which causes botton half is disabled
incorrectly.
quoted
Reported-by: Lijian Zhang <redacted>
Signed-off-by: Jia He <redacted>
---
This patch is adding additional spin_{un}lock_bh().
Can you please enlighten about the exact flow causing this unpaired
spin_{un}lock_bh.
For instance:
_qed_mcp_cmd_and_union()
  In while loop
    spin_lock_bh()
    qed_mcp_has_pending_cmd() (assume false), will break the loop
I agree till here.
quoted
  if (cnt >= max_retries) {
...
    return -EAGAIN; <-- here returns -EAGAIN without invoking bh unlock
  }
Because of break, cnt has not been increased.
   - cnt is still less than max_retries.
  - if (cnt >= max_retries) will not be *true*, leading to spin_unlock_bh().
Hence pairing completed.
Sorry, indeed. Let me check other possibilities.
@David S. Miller Sorry for the inconvenience, could you please revert it
in netdev tree?

Apologies again.

--
Cheers,
Justin (Jia He)

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help