Thread (25 messages) 25 messages, 9 authors, 3d ago

Re: [PATCH net 3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells

From: Tony Nguyen <anthony.l.nguyen@intel.com>
Date: 2026-10-01 16:35:23
Also in: stable


On 9/29/2026 5:58 PM, netdev-bot+sashiko@kernel.org wrote:
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The patch fixes only the writer side of the ordering chain
   behind the dql_completed() BUG_ON.

--

ice: Restore Ordered MMIO Writes for Tx Doorbells

This commit changes the Tx ring and tstamp ring tail doorbell writes in
ice_tx_map() from writel_relaxed() back to writel(). The goal is to make
the dql->num_queued update from __netdev_tx_sent_queue() visible before
the NIC can start on the new descriptors. That avoids the BUG_ON in
dql_completed() seen on arm64.

Link: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2161572
quoted
diff --git a/drivers/net/ethernet/intel/ice/ice_txrx.c b/drivers/net/ethernet/intel/ice/ice_txrx.c
index 31303ab5be175..a2c7c4962882d 100644
--- a/drivers/net/ethernet/intel/ice/ice_txrx.c
+++ b/drivers/net/ethernet/intel/ice/ice_txrx.c
@@ -1561,10 +1561,10 @@ ice_tx_map(struct ice_tx_ring *tx_ring, struct ice_tx_buf *first,
  			}
  		}
  		tstamp_ring->next_to_use = j;
-		writel_relaxed(j, tstamp_ring->tail);
+		writel(j, tstamp_ring->tail);
  	} else {
  ring_kick:
-		writel_relaxed(i, tx_ring->tail);
+		writel(i, tx_ring->tail);
  	}
  	return;
[Severity: Medium]

This orders the xmit side. Is there a matching read barrier on the
completion side?

ice_clean_tx_irq() reads the DONE bit that the device writes, and then
completes the queue:

drivers/net/ethernet/intel/ice/ice_txrx.c:ice_clean_tx_irq() {
     ...
		if (!(eop_desc->cmd_type_offset_bsz &
		      cpu_to_le64(ICE_TX_DESC_DTYPE_DESC_DONE)))
			break;
     ...
	netdev_tx_completed_queue(txring_txq(tx_ring), total_pkts, total_bytes);
     ...
}

netdev_tx_completed_queue() then calls dql_completed(), which uses a
relaxed load:

lib/dynamic_queue_limits.c:dql_completed() {
	num_queued = READ_ONCE(dql->num_queued);
     ...
	BUG_ON(count > num_queued - dql->num_completed);
     ...
}

The smp_mb() in netdev_tx_completed_queue() runs after dql_completed(), so
it does not order this load. On arm64, a control dependency on the DONE
bit does not order a later load either.

The barriers that happen to sit on this path are not always there:

- The smp_rmb() at the top of the ice_clean_tx_irq() loop runs again only
   when the next tx_buf has a non-NULL next_to_watch. It is skipped when
   the loop stops on a NULL next_to_watch or when the budget runs out.

- skb_unref() issues smp_rmb() only when skb->users == 1 and
   CONFIG_DEBUG_NET is off. With users > 1 (for example pktgen with burst,
   which does refcount_add(burst, &skb->users)), napi_consume_skb() skips
   the defer path. The failing refcount_dec_and_test() then gives only
   release ordering.

- u64_stats_update_begin()/end() in ice_update_tx_ring_stats() are no-ops
   on 64-bit. dma_unmap_single() does nothing for coherent dma-direct
   without an IOMMU.

On that path, can the num_queued load complete before the DONE bit load?
If so, it would return a value from before the xmit CPU's dql_queued()
update and hit the same BUG_ON this commit is meant to fix.

The window is probably very small in practice. The num_queued store comes
before the doorbell, the descriptor fetch, the transmit and the DONE
write-back, and testing shows the observed crash is gone. This reader-side
pattern is also older than both this patch and commit ccde82e90946.
Seems this is pre-existing as the review states this is older than the patch and cited Fixes commit. Sashiko points out that this "window is probably very small in practice." This patch is essentially reverting the change from writel_relaxed() to writel() that was done in the cited Fixes: commit and existed prior to it. Can we take this as-is as it's resolving a real issue and this situation seems unlikely as stated by AI and we have not had any reports of this.

Thanks,
Tony
Would a dma_rmb() or smp_rmb() in ice_clean_tx_irq() help close this
fully? It could go after the DONE check or before the
netdev_tx_completed_queue() call.
  
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help