RE: [PATCH iwl-net 2/2] igc: Flush pending Tx descriptors before returning NETDEV_TX_BUSY
flat view
From: Loktionov, Aleksandr <hidden>
Date: 2026-10-05 10:24:56
Also in:
intel-wired-lan, lkml
quoted hunk ↗ jump to hunk
-----Original Message----- From: Benoit DE RANCOURT <redacted> Sent: Sunday, October 4, 2026 5:29 PM To: Nguyen, Anthony L <anthony.l.nguyen@intel.com>; Kitszel, Przemyslaw [off-list ref]; intel-wired- lan@lists.osuosl.org Cc: netdev@vger.kernel.org; Andrew Lunn <andrew+netdev@lunn.ch>; David S. Miller [off-list ref]; Eric Dumazet [off-list ref]; Jakub Kicinski [off-list ref]; Paolo Abeni [off-list ref]; Gomes, Vinicius [off-list ref]; Sasha Neftin [off-list ref]; linux-kernel@vger.kernel.org; Benoit DE RANCOURT [off-list ref] Subject: [PATCH iwl-net 2/2] igc: Flush pending Tx descriptors before returning NETDEV_TX_BUSY When the stack sends a burst with xmit_more set, igc_tx_map() defers the tail register write until the last skb of the burst or until the queue is stopped. If igc_xmit_frame_ring() then refuses an skb with NETDEV_TX_BUSY, it returns without flushing any pending tail write. Descriptors of packets already accepted in the burst can therefore remain unexposed to the hardware. The queue is stopped at that point and igc_clean_tx_irq() only wakes it once TX_WAKE_THRESHOLD descriptors are free. The hardware can only complete what the tail exposes, so if the unsignalled descriptors keep the free count below the threshold, the queue is never woken and the Tx watchdog resets the adapter. On an I226-V running 7.2.8, in all 25 Tx timeout episodes of the test described in the previous patch, the stalled queue had pending descriptors after a NETDEV_TX_BUSY return: the probes inferred at least 214 of the 256 ring descriptors beyond the last tail write. For that queue, the register dump showed TDH = TDT = the next_to_clean value recorded by the probe: the hardware had completed everything it had been given. For example, with queue 2 stopped, next_to_clean = 190 and next_to_use = 167: igc 0000:08:00.0 terra: NETDEV WATCHDOG: CPU: 1: transmit queue 2 timed out 5353 ms igc 0000:08:00.0 terra: TDH[0-3] 00000099 00000091 000000be 000000f5 igc 0000:08:00.0 terra: TDT[0-3] 00000099 00000091 000000be 000000f5 Write the tail before returning NETDEV_TX_BUSY. The refused skb is left untouched; only descriptors of already accepted packets are exposed to the hardware. Those packets have already been accounted to BQL, so flushing their descriptors lets completion processing progress; the refused skb is neither mapped nor accounted here. The previous patch makes this path rare again for common skbs, but it remains reachable, for instance when the linear part or a fragment of an skb exceeds IGC_MAX_DATA_PER_TXD. With this change alone (without the previous patch), three runs of the same test on 7.2.8 still produced 10751 NETDEV_TX_BUSY returns, but no Tx timeout, against 25 on the unpatched kernel. Fixes: 0507ef8a0372 ("igc: Add transmit and receive fastpath and interrupt handlers") Cc: stable@vger.kernel.org Assisted-by: LLM bpftrace Signed-off-by: Benoit DE RANCOURT <redacted> --- drivers/net/ethernet/intel/igc/igc_main.c | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-)diff --git a/drivers/net/ethernet/intel/igc/igc_main.cb/drivers/net/ethernet/intel/igc/igc_main.c index b94a08791102..16f2ca32f5f1 100644--- a/drivers/net/ethernet/intel/igc/igc_main.c +++ b/drivers/net/ethernet/intel/igc/igc_main.c@@ -1621,7 +1621,11 @@ static netdev_tx_t igc_xmit_frame_ring(structsk_buff *skb, &skb_shinfo(skb)->frags[f])); if (igc_maybe_stop_tx(tx_ring, count + 5)) { - /* this is a hard error */ + /* This is a hard error. Write the tail for packets deferred + * by xmit_more before returning busy, otherwise the stopped + * queue may never get enough completions to be woken. + */ + igc_flush_tx_descriptors(tx_ring); return NETDEV_TX_BUSY; } -- 2.55.0
Reviewed-by: Aleksandr Loktionov <redacted>