RE: [PATCH] qed: fix possible unpaired spin_{un}lock_bh in _qed_mcp_cmd_and_union()
From: Justin He <hidden>
Date: 2021-07-20 02:10:44
Also in:
lkml
-----Original Message----- From: Prabhakar Kushwaha <redacted> Sent: Monday, July 19, 2021 10:51 PM To: Justin He <redacted> Cc: Ariel Elior <redacted>; GR-everest-linux-l2@marvell.com; David S. Miller [off-list ref]; Jakub Kicinski [off-list ref]; netdev@vger.kernel.org; Linux Kernel Mailing List <linux- kernel@vger.kernel.org>; nd [off-list ref]; Shai Malin [off-list ref]; Shai Malin [off-list ref]; Prabhakar Kushwaha [off-list ref] Subject: Re: [PATCH] qed: fix possible unpaired spin_{un}lock_bh in _qed_mcp_cmd_and_union() Hi Justin, On Mon, Jul 19, 2021 at 6:47 PM Justin He [off-list ref] wrote:quoted
Hi Prabhakarquoted
-----Original Message----- From: Prabhakar Kushwaha <redacted> Sent: Monday, July 19, 2021 6:36 PM To: Justin He <redacted> Cc: Ariel Elior <redacted>; GR-everest-linux-l2@marvell.com; David S. Miller [off-list ref]; Jakub Kicinski [off-list ref]; netdev@vger.kernel.org; Linux Kernel Mailing List <linux- kernel@vger.kernel.org>; nd [off-list ref]; Shai Malin[off-list ref];quoted
quoted
Shai Malin [off-list ref]; Prabhakar Kushwaha[off-list ref]quoted
quoted
Subject: Re: [PATCH] qed: fix possible unpaired spin_{un}lock_bh in _qed_mcp_cmd_and_union() Hi Jia, On Thu, Jul 15, 2021 at 2:28 PM Jia He [off-list ref] wrote:quoted
Liajian reported a bug_on hit on a ThunderX2 arm64 server withFastLinQquoted
quoted
quoted
QL41000 ethernet controller: BUG: scheduling while atomic: kworker/0:4/531/0x00000200 [qed_probe:488()]hw prepare failed kernel BUG at mm/vmalloc.c:2355! Internal error: Oops - BUG: 0 [#1] SMP CPU: 0 PID: 531 Comm: kworker/0:4 Tainted: G W 5.4.0-77-generic#86-quoted
quoted
Ubuntuquoted
pstate: 00400009 (nzcv daif +PAN -UAO) Call trace: vunmap+0x4c/0x50 iounmap+0x48/0x58 qed_free_pci+0x60/0x80 [qed] qed_probe+0x35c/0x688 [qed] __qede_probe+0x88/0x5c8 [qede] qede_probe+0x60/0xe0 [qede] local_pci_probe+0x48/0xa0 work_for_cpu_fn+0x24/0x38 process_one_work+0x1d0/0x468 worker_thread+0x238/0x4e0 kthread+0xf0/0x118 ret_from_fork+0x10/0x18 In this case, qed_hw_prepare() returns error due to hw/fw error, butinquoted
quoted
quoted
theory work queue should be in process context instead of interrupt. The root cause might be the unpaired spin_{un}lock_bh() in _qed_mcp_cmd_and_union(), which causes botton half is disabledincorrectly.quoted
Reported-by: Lijian Zhang <redacted> Signed-off-by: Jia He <redacted> ---This patch is adding additional spin_{un}lock_bh(). Can you please enlighten about the exact flow causing this unpaired spin_{un}lock_bh.For instance: _qed_mcp_cmd_and_union() In while loop spin_lock_bh() qed_mcp_has_pending_cmd() (assume false), will break the loopI agree till here.quoted
if (cnt >= max_retries) { ... return -EAGAIN; <-- here returns -EAGAIN without invoking bh unlock }Because of break, cnt has not been increased. - cnt is still less than max_retries. - if (cnt >= max_retries) will not be *true*, leading to spin_unlock_bh(). Hence pairing completed.
Sorry, indeed. Let me check other possibilities. @David S. Miller Sorry for the inconvenience, could you please revert it in netdev tree? Apologies again. -- Cheers, Justin (Jia He)