From: Saeed Mahameed <saeedm@nvidia.com>
Hi Dave, Hi Jakub,
I know I missed the train for this week, this is V2 just in case there
will be another rc pull later next week.
Note: I cherry-picked the patch from net-next and it applies cleanly,
let me know if you face any issues with this pull.
v1->v2:
- Fixed missing space in commit message
- Cherry-picked an important fix from net-next patch #12
This series provides bug fixes to mlx5 driver.
Please pull and let me know if there is any problem.
Thanks,
Saeed.
The following changes since commit d1652b70d07cc3eed96210c876c4879e1655f20e:
asix: fix wrong return value in asix_check_host_enable() (2021-12-22 14:52:18 -0800)
are available in the Git repository at:
git://git.kernel.org/pub/scm/linux/kernel/git/saeed/linux.git tags/mlx5-fixes-2021-12-22
for you to fetch changes up to 4390c6edc0fb390e699d0f886f45575dfeafeb4b:
net/mlx5: Fix some error handling paths in 'mlx5e_tc_add_fdb_flow()' (2021-12-22 20:38:49 -0800)
----------------------------------------------------------------
mlx5-fixes-2021-12-22
----------------------------------------------------------------
Amir Tzin (1):
net/mlx5e: Wrap the tx reporter dump callback to extract the sq
Chris Mi (2):
net/mlx5: Fix tc max supported prio for nic mode
net/mlx5e: Delete forward rule for ct or sample action
Christophe JAILLET (1):
net/mlx5: Fix some error handling paths in 'mlx5e_tc_add_fdb_flow()'
Gal Pressman (1):
net/mlx5e: Fix skb memory leak when TC classifier action offloads are disabled
Maxim Mikityanskiy (2):
net/mlx5e: Fix interoperability between XSK and ICOSQ recovery flow
net/mlx5e: Fix ICOSQ recovery flow for XSK
Miaoqian Lin (1):
net/mlx5: DR, Fix NULL vs IS_ERR checking in dr_domain_init_resources
Moshe Shemesh (1):
net/mlx5: Fix SF health recovery flow
Shay Drory (2):
net/mlx5: Use first online CPU instead of hard coded CPU
net/mlx5: Fix error print in case of IRQ request failed
Yevgeny Kliteynik (1):
net/mlx5: DR, Fix querying eswitch manager vport for ECPF
drivers/net/ethernet/mellanox/mlx5/core/en.h | 5 ++-
.../net/ethernet/mellanox/mlx5/core/en/health.h | 2 ++
.../net/ethernet/mellanox/mlx5/core/en/rep/tc.h | 2 +-
.../ethernet/mellanox/mlx5/core/en/reporter_rx.c | 35 +++++++++++++++++++-
.../ethernet/mellanox/mlx5/core/en/reporter_tx.c | 10 +++++-
.../net/ethernet/mellanox/mlx5/core/en/xsk/setup.c | 16 +++++++++-
drivers/net/ethernet/mellanox/mlx5/core/en_main.c | 37 ++++++++++++++++------
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c | 31 +++++++++---------
.../ethernet/mellanox/mlx5/core/lib/fs_chains.c | 3 ++
drivers/net/ethernet/mellanox/mlx5/core/main.c | 11 ++++---
drivers/net/ethernet/mellanox/mlx5/core/pci_irq.c | 6 ++--
.../mellanox/mlx5/core/steering/dr_domain.c | 9 +++---
12 files changed, 121 insertions(+), 46 deletions(-)
From: Shay Drory <redacted>
In case IRQ layer failed to find or to request irq, the driver is
printing the first cpu of the provided affinity as part of the error
print. Empty affinity is a valid input for the IRQ layer, and it is
an error to call cpumask_first() on empty affinity.
Remove the first cpu print from the error message.
Fixes: c36326d38d93 ("net/mlx5: Round-Robin EQs over IRQs")
Signed-off-by: Shay Drory <redacted>
Reviewed-by: Moshe Shemesh <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/pci_irq.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
From: Moshe Shemesh <redacted>
SF do not directly control the PCI device. During recovery flow SF
should not be allowed to do pci disable or pci reset, its PF will do it.
It fixes the following kernel trace:
mlx5_core.sf mlx5_core.sf.25: mlx5_health_try_recover:387:(pid 40948): starting health recovery flow
mlx5_core 0000:03:00.0: mlx5_pci_slot_reset was called
mlx5_core 0000:03:00.0: wait vital counter value 0xab175 after 1 iterations
mlx5_core.sf mlx5_core.sf.25: firmware version: 24.32.532
mlx5_core.sf mlx5_core.sf.23: mlx5_health_try_recover:387:(pid 40946): starting health recovery flow
mlx5_core 0000:03:00.0: mlx5_pci_slot_reset was called
mlx5_core 0000:03:00.0: wait vital counter value 0xab193 after 1 iterations
mlx5_core.sf mlx5_core.sf.23: firmware version: 24.32.532
mlx5_core.sf mlx5_core.sf.25: mlx5_cmd_check:813:(pid 40948): ENABLE_HCA(0x104) op_mod(0x0) failed,
status bad resource state(0x9), syndrome (0x658908)
mlx5_core.sf mlx5_core.sf.25: mlx5_function_setup:1292:(pid 40948): enable hca failed
mlx5_core.sf mlx5_core.sf.25: mlx5_health_try_recover:389:(pid 40948): health recovery failed
Fixes: 1958fc2f0712 ("net/mlx5: SF, Add auxiliary device driver")
Signed-off-by: Moshe Shemesh <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/main.c | 11 ++++++-----
1 file changed, 6 insertions(+), 5 deletions(-)
From: Chris Mi <redacted>
Only prio 1 is supported if firmware doesn't support ignore flow
level for nic mode. The offending commit removed the check wrongly.
Add it back.
Fixes: 9a99c8f1253a ("net/mlx5e: E-Switch, Offload all chain 0 priorities when modify header and forward action is not supported")
Signed-off-by: Chris Mi <redacted>
Reviewed-by: Roi Dayan <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/lib/fs_chains.c | 3 +++
1 file changed, 3 insertions(+)
From: Chris Mi <redacted>
When there is ct or sample action, the ct or sample rule will be deleted
and return. But if there is an extra mirror action, the forward rule can't
be deleted because of the return.
Fix it by removing the return.
Fixes: 69e2916ebce4 ("net/mlx5: CT: Add support for mirroring")
Fixes: f94d6389f6a8 ("net/mlx5e: TC, Add support to offload sample action")
Signed-off-by: Chris Mi <redacted>
Reviewed-by: Roi Dayan <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c | 17 ++++++-----------
1 file changed, 6 insertions(+), 11 deletions(-)
From: Maxim Mikityanskiy <redacted>
Both regular RQ and XSKRQ use the same ICOSQ for UMRs. When doing
recovery for the ICOSQ, don't forget to deactivate XSKRQ.
XSK can be opened and closed while channels are active, so a new mutex
prevents the ICOSQ recovery from running at the same time. The ICOSQ
recovery deactivates and reactivates XSKRQ, so any parallel change in
XSK state would break consistency. As the regular RQ is running, it's
not enough to just flush the recovery work, because it can be
rescheduled.
Fixes: be5323c8379f ("net/mlx5e: Report and recover from CQE error on ICOSQ")
Signed-off-by: Maxim Mikityanskiy <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en.h | 2 ++
.../ethernet/mellanox/mlx5/core/en/health.h | 2 ++
.../mellanox/mlx5/core/en/reporter_rx.c | 35 ++++++++++++++++++-
.../mellanox/mlx5/core/en/xsk/setup.c | 16 ++++++++-
.../net/ethernet/mellanox/mlx5/core/en_main.c | 7 ++--
5 files changed, 58 insertions(+), 4 deletions(-)
@@ -70,7 +71,13 @@ static int mlx5e_rx_reporter_err_icosq_cqe_recover(void *ctx)interr;icosq=ctx;++mutex_lock(&icosq->channel->icosq_recovery_lock);++/* mlx5e_close_rq cancels this work before RQ and ICOSQ are killed. */rq=&icosq->channel->rq;+if(test_bit(MLX5E_RQ_STATE_ENABLED,&icosq->channel->xskrq.state))+xskrq=&icosq->channel->xskrq;mdev=icosq->channel->mdev;dev=icosq->channel->netdev;err=mlx5_core_query_sq_state(mdev,icosq->sqn,&state);
@@ -84,6 +91,9 @@ static int mlx5e_rx_reporter_err_icosq_cqe_recover(void *ctx)gotoout;mlx5e_deactivate_rq(rq);+if(xskrq)+mlx5e_deactivate_rq(xskrq);+err=mlx5e_wait_for_icosq_flush(icosq);if(err)gotoout;
@@ -97,15 +107,28 @@ static int mlx5e_rx_reporter_err_icosq_cqe_recover(void *ctx)gotoout;mlx5e_reset_icosq_cc_pc(icosq);+mlx5e_free_rx_in_progress_descs(rq);+if(xskrq)+mlx5e_free_rx_in_progress_descs(xskrq);+clear_bit(MLX5E_SQ_STATE_RECOVERING,&icosq->state);mlx5e_activate_icosq(icosq);-mlx5e_activate_rq(rq);+mlx5e_activate_rq(rq);rq->stats->recover++;++if(xskrq){+mlx5e_activate_rq(xskrq);+xskrq->stats->recover++;+}++mutex_unlock(&icosq->channel->icosq_recovery_lock);+return0;out:clear_bit(MLX5E_SQ_STATE_RECOVERING,&icosq->state);+mutex_unlock(&icosq->channel->icosq_recovery_lock);returnerr;}
@@ -4,6 +4,7 @@#include"setup.h"#include"en/params.h"#include"en/txrx.h"+#include"en/health.h"/* It matches XDP_UMEM_MIN_CHUNK_SIZE, but as this constant is private and may*changeunexpectedly,andmlx5ehasaminimumvalidstridesizeforstriding
@@ -170,7 +171,13 @@ void mlx5e_close_xsk(struct mlx5e_channel *c)voidmlx5e_activate_xsk(structmlx5e_channel*c){+/* ICOSQ recovery deactivates RQs. Suspend the recovery to avoid+*activatingXSKRQinthemiddleofrecovery.+*/+mlx5e_reporter_icosq_suspend_recovery(c);set_bit(MLX5E_RQ_STATE_ENABLED,&c->xskrq.state);+mlx5e_reporter_icosq_resume_recovery(c);+/* TX queue is created active. */spin_lock_bh(&c->async_icosq_lock);
@@ -180,6 +187,13 @@ void mlx5e_activate_xsk(struct mlx5e_channel *c)voidmlx5e_deactivate_xsk(structmlx5e_channel*c){-mlx5e_deactivate_rq(&c->xskrq);+/* ICOSQ recovery may reactivate XSKRQ if clear_bit is called in the+*middleofrecovery.Suspendtherecoverytoavoidit.+*/+mlx5e_reporter_icosq_suspend_recovery(c);+clear_bit(MLX5E_RQ_STATE_ENABLED,&c->xskrq.state);+mlx5e_reporter_icosq_resume_recovery(c);+synchronize_net();/* Sync with NAPI to prevent mlx5e_post_rx_wqes. */+/* TX queue is disabled on close. */}
@@ -2088,6 +2086,8 @@ static int mlx5e_open_queues(struct mlx5e_channel *c,if(err)gotoerr_close_xdpsq_cq;+mutex_init(&c->icosq_recovery_lock);+err=mlx5e_open_icosq(c,params,&cparam->icosq,&c->icosq);if(err)gotoerr_close_async_icosq;
@@ -2156,9 +2156,12 @@ static void mlx5e_close_queues(struct mlx5e_channel *c)mlx5e_close_xdpsq(&c->xdpsq);if(c->xdp)mlx5e_close_xdpsq(&c->rq_xdpsq);+/* The same ICOSQ is used for UMRs for both RQ and XSKRQ. */+cancel_work_sync(&c->icosq.recover_work);mlx5e_close_rq(&c->rq);mlx5e_close_sqs(c);mlx5e_close_icosq(&c->icosq);+mutex_destroy(&c->icosq_recovery_lock);mlx5e_close_icosq(&c->async_icosq);if(c->xdp)mlx5e_close_cq(&c->rq_xdpsq.cq);
From: Maxim Mikityanskiy <redacted>
There are two ICOSQs per channel: one is needed for RX, and the other
for async operations (XSK TX, kTLS offload). Currently, the recovery
flow for both is the same, and async ICOSQ is mistakenly treated like
the regular ICOSQ.
This patch prevents running the regular ICOSQ recovery on async ICOSQ.
The purpose of async ICOSQ is to handle XSK wakeup requests and post
kTLS offload RX parameters, it has nothing to do with RQ and XSKRQ UMRs,
so the regular recovery sequence is not applicable here.
Fixes: be5323c8379f ("net/mlx5e: Report and recover from CQE error on ICOSQ")
Signed-off-by: Maxim Mikityanskiy <redacted>
Reviewed-by: Aya Levin <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en.h | 3 --
.../net/ethernet/mellanox/mlx5/core/en_main.c | 30 ++++++++++++++-----
2 files changed, 22 insertions(+), 11 deletions(-)
From: Christophe JAILLET <redacted>
All the error handling paths of 'mlx5e_tc_add_fdb_flow()' end to 'err_out'
where 'flow_flag_set(flow, FAILED);' is called.
All but the new error handling paths added by the commits given in the
Fixes tag below.
Fix these error handling paths and branch to 'err_out'.
Fixes: 166f431ec6be ("net/mlx5e: Add indirect tc offload of ovs internal port")
Fixes: b16eb3c81fe2 ("net/mlx5: Support internal port as decap route device")
Signed-off-by: Christophe JAILLET <redacted>
Reviewed-by: Roi Dayan <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
(cherry picked from commit 31108d142f3632970f6f3e0224bd1c6781c9f87d)
---
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c | 14 +++++++++-----
1 file changed, 9 insertions(+), 5 deletions(-)
@@ -1456,13 +1456,15 @@ mlx5e_tc_add_fdb_flow(struct mlx5e_priv *priv,if(attr->chain){NL_SET_ERR_MSG_MOD(extack,"Internal port rule is only supported on chain 0");-return-EOPNOTSUPP;+err=-EOPNOTSUPP;+gotoerr_out;}if(attr->dest_chain){NL_SET_ERR_MSG_MOD(extack,"Internal port rule offload doesn't support goto action");-return-EOPNOTSUPP;+err=-EOPNOTSUPP;+gotoerr_out;}int_port=mlx5e_tc_int_port_get(mlx5e_get_int_port_priv(priv),
From: Gal Pressman <redacted>
When TC classifier action offloads are disabled (CONFIG_MLX5_CLS_ACT in
Kconfig), the mlx5e_rep_tc_receive() function which is responsible for
passing the skb to the stack (or freeing it) is defined as a nop, and
results in leaking the skb memory. Replace the nop with a call to
napi_gro_receive() to resolve the leak.
Fixes: 28e7606fa8f1 ("net/mlx5e: Refactor rx handler of represetor device")
Signed-off-by: Gal Pressman <redacted>
Reviewed-by: Ariel Levkovich <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en/rep/tc.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Hello:
This series was applied to netdev/net.git (master)
by Saeed Mahameed [off-list ref]:
On Thu, 23 Dec 2021 11:04:30 -0800 you wrote:
From: Miaoqian Lin <redacted>
The mlx5_get_uars_page() function returns error pointers.
Using IS_ERR() to check the return value to fix this.
Fixes: 4ec9e7b02697 ("net/mlx5: DR, Expose steering domain functionality")
Signed-off-by: Miaoqian Lin <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
[...]