[PATCH net V3 0/3] net/mlx5: SD LAG and devcom stability fixes
From: Tariq Toukan <tariqt@nvidia.com>
Date: 2026-09-15 11:35:57
Also in:
linux-rdma, lkml
Hi, This series by Shay fixes four bugs in the Socket Direct LAG and devcom subsystems, all related to initialization/teardown ordering and concurrent access to the LAG device. Regards, Tariq Internal sashiko comment:
quoted hunk
@@ -342,7 +342,14 @@ static void sd_lag_init(struct mlx5_core_dev *dev) return; } +recheck: mutex_lock(&ldev->lock); + if (ldev->mode_changes_in_progress) { + mutex_unlock(&ldev->lock); + msleep(100); + goto recheck; +
}
+
Does this open-coded retry loop reimplement a wait mechanism without immediate wakeups or fairness? It looks like we are polling the mode_changes_in_progress flag using a hard coded msleep(100). Could this unnecessarily delay the initialization path if the condition clears much sooner than 100ms? Would it be better to use a proper synchronization primitive like a waitqueue here instead of an ad-hoc flag loop? [SD] This is the same check as in mlx5_lag_remove_mdev(). I agree we need to change it, but this is net-next material V3: - Drop the "SD, serialize SD LAG init/cleanup against LAG mode changes" patch. - Elaborate the commit message in "net/mlx5: LAG, reload IB reps of LAG master before the rest". V2: https://lore.kernel.org/all/20260906071332.3759199-1-tariqt@nvidia.com/ (local) Shay Drory (3): net/mlx5: devcom, Base component size on linked devices net/mlx5: SD, unload reps on shared FDB create error path net/mlx5: LAG, reload IB reps of LAG master before the rest .../net/ethernet/mellanox/mlx5/core/lag/lag.c | 44 ++++++++++++++----- .../mellanox/mlx5/core/lag/shared_fdb.c | 1 + .../ethernet/mellanox/mlx5/core/lib/devcom.c | 5 ++- 3 files changed, 37 insertions(+), 13 deletions(-) base-commit: 23ca4ddc4fce2c233a49e9fd34d4b5b02bd7324e -- 2.44.0