Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
flat view
From: anmory <hidden>
Date: 2026-10-06 16:43:54
Also in:
lkml
Hi Joe, we appear to be hitting a very similar issue to the one described in this RFC on two Broadcom BCM57504 systems. One detail that may be particularly relevant is the ordering in our reproductions: on both systems the AMD-Vi IO_PAGE_FAULT occurs before the NETDEV WATCHDOG / TX timeout. We originally encountered the problem after updating from Debian 6.12.107-1 to 6.12.111-1. We subsequently reproduced it deliberately on two separate systems. Test results: affected-host-1: BCM57504 [14e4:1751], bnxt_en BIOS: 1.16.2 NIC FW: 36.11.55.00 FW mgmt: 236.1.153.0 6.12.107-1: stable 6.12.111-1: failure reproduced affected-host-2: BCM57504 [14e4:1751], bnxt_en BIOS: 1.18.2 NIC FW: 36.11.73.00 FW mgmt: 236.1.173.0 6.12.107-1: stable 6.12.111-1: failure reproduced The failure sequence on affected-host-1 was: 06:49:00 AMD-Vi IO_PAGE_FAULT 06:49:14 NETDEV WATCHDOG: transmit queue 0 timed out 06:49:14 TX timeout detected, starting reset task 06:49:18 HWRM/RING_FREE failures begin 06:49:25 HWRM_RING_ALLOC fails 06:49:25 bnxt_init_nic fails 06:49:25 nic open fails On affected-host-2: 15:29:09 AMD-Vi IO_PAGE_FAULT 15:29:15 NETDEV WATCHDOG: transmit queue 4 timed out 15:29:15 TX timeout detected, starting reset task 15:29:19 HWRM/RING_FREE failures begin 15:29:37 HWRM_RING_ALLOC fails 15:29:37 bnxt_init_nic fails 15:29:37 nic open fails So in our reproductions the IOMMU fault precedes the TX watchdog by approximately 6-14 seconds. We have also observed the timeout on different TX queues across different occurrences, so it does not appear to be tied to a specific queue. Both interfaces were up and operating at 25 Gbit/s before the failure. The systems use the IOMMU in translated mode. After the reset attempt fails, the interface cannot be reopened and a reboot is required. The same problem therefore reproduces across: - two physical systems - two BIOS revisions - two NIC firmware revisions - different TX queues while reverting to 6.12.107-1 has been stable with the same workload. I've attached sanitized diagnostic output from both reproductions, including PCI/device information, firmware versions, IOMMU configuration and the complete failure event sequence. We have not bisected the changes between 6.12.107 and 6.12.111 yet. Given that the IO_PAGE_FAULT precedes the watchdog in our case, I wonder whether this may provide some information about the initial failure, in addition to the RING_FREE recovery problem addressed by your RFC. Please let me know if there are additional diagnostics that would be useful, or if you would like us to test the RFC series on one of these systems. Thanks, R. https://proton.me/mail/home
Attachments
- affected-host-1-bnxt.txt [text/plain] 4027 bytes · preview
- affected-host-2-bnxt.txt [text/plain] 5471 bytes · preview