Thread (26 messages) 26 messages, 2 authors, 2d ago

Re: [PATCH net-next v4 14/15] net: macb: use context swapping in .set_ringparam()

From: Nicolai Buchwitz <hidden>
Date: 2026-07-19 10:53:34
Also in: lkml

Hi Théo

On 17.7.2026 21:48, Théo Lebrun wrote:
quoted hunk ↗ jump to hunk
ethtool_ops.set_ringparam() is implemented using the primitive close /
update ring size / reopen sequence. Under memory pressure this does not
fly: we free our buffers at close and cannot reallocate new ones at
open. Also, it triggers a slow PHY reinit.

Instead, exploit the new context mechanism and improve our sequence to:
 - allocate a new context (including buffers) first
 - if it fails, early return without any impact to the interface
 - stop interface
 - update global state (bp, netdev, etc)
 - pass buffer pointers to the hardware
 - start interface
 - free old context.

The HW disable sequence is inspired by macb_reset_hw() but avoids
(1) setting NCR bit CLRSTAT and (2) clearing register PBUFRXCUT.

The HW re-enable sequence is inspired by macb_mac_link_up(), skipping
over register writes which would be redundant (because values have not
changed).

The generic context swapping parts are isolated into helper functions
macb_context_swap_start|end(), reusable by other operations 
(change_mtu,
set_channels, etc).

Introduce a new locking primitive (mac_cfg_lock mutex) to serialise 
swap
with phylink MAC callbacks. Avoid stopping phylink to avoid a slow PHY
retrain. Those callbacks grab phydev->lock if it exists so we could
imagine grabbing that from the swap op, but phydev->lock doesn't exist
in the SFP case.

AT91 EMAC is handled differently as their buffer management is separate
and they don't do NAPI. We refuse them (-EBUSY) to avoid implementing
context swapping for them.

Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
---
 drivers/net/ethernet/cadence/macb.h      |   5 +
 drivers/net/ethernet/cadence/macb_main.c | 162 
+++++++++++++++++++++++++++++--
 2 files changed, 158 insertions(+), 9 deletions(-)
diff --git a/drivers/net/ethernet/cadence/macb.h 
b/drivers/net/ethernet/cadence/macb.h
index ac2f2d8065d7..93e513cf1fbb 100644
--- a/drivers/net/ethernet/cadence/macb.h
+++ b/drivers/net/ethernet/cadence/macb.h
@@ -1361,6 +1361,8 @@ struct macb {
 	struct macb_queue	queues[MACB_MAX_QUEUES];

 	spinlock_t		lock;
+	/* Serializes context swap against phylink MAC callbacks. */
+	struct mutex		mac_cfg_lock;
 	struct clk		*pclk;
 	struct clk		*hclk;
 	struct clk		*tx_clk;
@@ -1421,6 +1423,9 @@ struct macb {
 	struct delayed_work	tx_lpi_work;
 	u32			tx_lpi_timer;

+	/* ISR must not drive NAPI & BH mechanisms. Protected by bp->lock. */
+	bool			ctx_swap;
+
 	u32	rx_intr_mask;

 	struct macb_pm_data pm_data;
diff --git a/drivers/net/ethernet/cadence/macb_main.c 
b/drivers/net/ethernet/cadence/macb_main.c
index c832b6c1b98c..5792647eb0a6 100644
--- a/drivers/net/ethernet/cadence/macb_main.c
+++ b/drivers/net/ethernet/cadence/macb_main.c
[...]
+
+	for (q = 0, queue = bp->queues; q < bp->num_queues; ++q, ++queue) {
+		/* Must be done before NAPI is disabled. */
+		cancel_work_sync(&queue->tx_error_task);
+
+		napi_disable(&queue->napi_rx);
+		napi_disable(&queue->napi_tx);
+		netdev_tx_reset_queue(netdev_get_tx_queue(bp->netdev, q));
+	}
+
+	/* Must be done after napi_tx is disabled. */
+	cancel_delayed_work_sync(&bp->tx_lpi_work);
+
+	/* Can finally disable software Tx; need to wait until napi_tx and
+	 * tx_error_task cannot be scheduled as either might wakeup Tx.
+	 */
+	netif_tx_disable(bp->netdev);
Shouldn't netdev_tx_reset_queue() come after netif_tx_disable()? Tx is
still running here, so an xmit right after the reset adds bytes to the
DQL that never get completed (macb_free() frees the old skbs without
netdev_tx_completed_queue()). Not sure, but that could leave the queue
stopped by BQL forever?
[...]
Thanks
Nicolai
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help