From: Michael Chan <michael.chan@broadcom.com> Date: 2021-09-12 16:35:07
The first patch fixes an error recovery regression just introduced
about a week ago. The other two patches fix issues related to
freeing rings in the bnxt_close() path under error conditions.
Edwin Peer (1):
bnxt_en: make bnxt_free_skbs() safe to call after bnxt_free_mem()
Michael Chan (2):
bnxt_en: Fix error recovery regression
bnxt_en: Clean up completion ring page arrays completely
drivers/net/ethernet/broadcom/bnxt/bnxt.c | 33 ++++++++++++++++++++---
1 file changed, 29 insertions(+), 4 deletions(-)
--
2.18.1
From: Michael Chan <michael.chan@broadcom.com> Date: 2021-09-12 16:35:10
The recent patch has introduced a regression by not reading the reset
count in the ERROR_RECOVERY async event handler. We may have just
gone through a reset and the reset count has just incremented. If
we don't update the reset count in the ERROR_RECOVERY event handler,
the health check timer will see that the reset count has changed and
will initiate an unintended reset.
Restore the unconditional update of the reset count in
bnxt_async_event_process() if error recovery watchdog is enabled.
Also, update the reset count at the end of the reset sequence to
make it even more robust.
Fixes: 1b2b91831983 ("bnxt_en: Fix possible unintended driver initiated error recovery")
Reviewed-by: Edwin Peer <redacted>
Signed-off-by: Michael Chan <michael.chan@broadcom.com>
---
drivers/net/ethernet/broadcom/bnxt/bnxt.c | 12 ++++++++----
1 file changed, 8 insertions(+), 4 deletions(-)
@@ -2213,12 +2213,11 @@ static int bnxt_async_event_process(struct bnxt *bp,DIV_ROUND_UP(fw_health->polling_dsecs*HZ,bp->current_interval*10);fw_health->tmr_counter=fw_health->tmr_multiplier;-if(!fw_health->enabled){+if(!fw_health->enabled)fw_health->last_fw_heartbeat=bnxt_fw_health_readl(bp,BNXT_FW_HEARTBEAT_REG);-fw_health->last_fw_reset_cnt=-bnxt_fw_health_readl(bp,BNXT_FW_RESET_CNT_REG);-}+fw_health->last_fw_reset_cnt=+bnxt_fw_health_readl(bp,BNXT_FW_RESET_CNT_REG);netif_info(bp,drv,bp->dev,"Error recovery info: error recovery[1], master[%d], reset count[%u], health status: 0x%x\n",fw_health->master,fw_health->last_fw_reset_cnt,
@@ -12207,6 +12206,11 @@ static void bnxt_fw_reset_task(struct work_struct *work)return;}+if((bp->fw_cap&BNXT_FW_CAP_ERROR_RECOVERY)&&+bp->fw_health->enabled){+bp->fw_health->last_fw_reset_cnt=+bnxt_fw_health_readl(bp,BNXT_FW_RESET_CNT_REG);+}bp->fw_reset_state=0;/* Make sure fw_reset_state is 0 before clearing the flag */smp_mb__before_atomic();
From: Michael Chan <michael.chan@broadcom.com> Date: 2021-09-12 16:35:11
From: Edwin Peer <redacted>
The call to bnxt_free_mem(..., false) in the bnxt_half_open_nic() error
path will deallocate ring descriptor memory via bnxt_free_?x_rings(),
but because irq_re_init is false, the ring info itself is not freed.
To simplify error paths, deallocation functions have generally been
written to be safe when called on unallocated memory. It should always
be safe to call dev_close(), which calls bnxt_free_skbs() a second time,
even in this semi- allocated ring state.
Calling bnxt_free_skbs() a second time with the rings already freed will
cause NULL pointer dereference. Fix it by checking the rings are valid
before proceeding in bnxt_free_tx_skbs() and
bnxt_free_one_rx_ring_skbs().
Fixes: 975bc99a4a39 ("bnxt_en: Refactor bnxt_free_rx_skbs().")
Signed-off-by: Edwin Peer <redacted>
Signed-off-by: Michael Chan <michael.chan@broadcom.com>
---
drivers/net/ethernet/broadcom/bnxt/bnxt.c | 13 +++++++++++++
1 file changed, 13 insertions(+)
From: Michael Chan <michael.chan@broadcom.com> Date: 2021-09-12 16:35:11
We recently changed the completion ring page arrays to be dynamically
allocated to better support the expanded range of ring depths. The
cleanup path for this was not quite complete. It might cause the
shutdown path to crash if we need to abort before the completion ring
arrays have been allocated and initialized.
Fix it by initializing the ring_mem->pg_arr to NULL after freeing the
completion ring page array. Add a check in bnxt_free_ring() to skip
referencing the rmem->pg_arr if it is NULL.
Fixes: 03c7448790b8 ("bnxt_en: Don't use static arrays for completion ring pages")
Reviewed-by: Andy Gospodarek <redacted>
Reviewed-by: Edwin Peer <redacted>
Signed-off-by: Michael Chan <michael.chan@broadcom.com>
---
drivers/net/ethernet/broadcom/bnxt/bnxt.c | 8 ++++++++
1 file changed, 8 insertions(+)
Hello:
This series was applied to netdev/net.git (refs/heads/master):
On Sun, 12 Sep 2021 12:34:46 -0400 you wrote:
The first patch fixes an error recovery regression just introduced
about a week ago. The other two patches fix issues related to
freeing rings in the bnxt_close() path under error conditions.
Edwin Peer (1):
bnxt_en: make bnxt_free_skbs() safe to call after bnxt_free_mem()
[...]