Thread (6 messages) 6 messages, 2 authors, 16h ago

Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust

flat view

From: anmory <hidden>
Date: 2026-10-06 16:43:54
Also in: lkml

Hi Joe,

we appear to be hitting a very similar issue to the one described in
this RFC on two Broadcom BCM57504 systems.

One detail that may be particularly relevant is the ordering in our
reproductions: on both systems the AMD-Vi IO_PAGE_FAULT occurs before
the NETDEV WATCHDOG / TX timeout.

We originally encountered the problem after updating from Debian
6.12.107-1 to 6.12.111-1.

We subsequently reproduced it deliberately on two separate systems.

Test results:

affected-host-1:
BCM57504 [14e4:1751], bnxt_en
BIOS: 1.16.2
NIC FW: 36.11.55.00
FW mgmt: 236.1.153.0

6.12.107-1: stable
6.12.111-1: failure reproduced

affected-host-2:
BCM57504 [14e4:1751], bnxt_en
BIOS: 1.18.2
NIC FW: 36.11.73.00
FW mgmt: 236.1.173.0

6.12.107-1: stable
6.12.111-1: failure reproduced

The failure sequence on affected-host-1 was:

06:49:00 AMD-Vi IO_PAGE_FAULT
06:49:14 NETDEV WATCHDOG: transmit queue 0 timed out
06:49:14 TX timeout detected, starting reset task
06:49:18 HWRM/RING_FREE failures begin
06:49:25 HWRM_RING_ALLOC fails
06:49:25 bnxt_init_nic fails
06:49:25 nic open fails

On affected-host-2:

15:29:09 AMD-Vi IO_PAGE_FAULT
15:29:15 NETDEV WATCHDOG: transmit queue 4 timed out
15:29:15 TX timeout detected, starting reset task
15:29:19 HWRM/RING_FREE failures begin
15:29:37 HWRM_RING_ALLOC fails
15:29:37 bnxt_init_nic fails
15:29:37 nic open fails

So in our reproductions the IOMMU fault precedes the TX watchdog by
approximately 6-14 seconds.

We have also observed the timeout on different TX queues across
different occurrences, so it does not appear to be tied to a specific
queue.

Both interfaces were up and operating at 25 Gbit/s before the failure.
The systems use the IOMMU in translated mode.

After the reset attempt fails, the interface cannot be reopened and a
reboot is required.

The same problem therefore reproduces across:
- two physical systems
- two BIOS revisions
- two NIC firmware revisions
- different TX queues

while reverting to 6.12.107-1 has been stable with the same workload.

I've attached sanitized diagnostic output from both reproductions,
including PCI/device information, firmware versions, IOMMU
configuration and the complete failure event sequence.

We have not bisected the changes between 6.12.107 and 6.12.111 yet.

Given that the IO_PAGE_FAULT precedes the watchdog in our case, I
wonder whether this may provide some information about the initial
failure, in addition to the RING_FREE recovery problem addressed by
your RFC.

Please let me know if there are additional diagnostics that would be
useful, or if you would like us to test the RFC series on one of these
systems.

Thanks,
R.

https://proton.me/mail/home

Attachments

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help