From: Saeed Mahameed <saeedm@nvidia.com>
Hi Dave, Hi Jakub,
This series provides bug fixes to mlx5 driver.
Please pull and let me know if there is any problem.
Thanks,
Saeed.
The following changes since commit b0f38e15979fa8851e88e8aa371367f264e7b6e9:
natsemi: xtensa: fix section mismatch warnings (2021-11-30 18:13:37 -0800)
are available in the Git repository at:
git://git.kernel.org/pub/scm/linux/kernel/git/saeed/linux.git tags/mlx5-fixes-2021-11-30
for you to fetch changes up to 8c8cf0382257b28378eeff535150c087a653ca19:
net/mlx5e: SHAMPO, Fix constant expression result (2021-11-30 22:35:06 -0800)
----------------------------------------------------------------
mlx5-fixes-2021-11-30
----------------------------------------------------------------
Amir Tzin (1):
net/mlx5: Fix use after free in mlx5_health_wait_pci_up
Aya Levin (1):
net/mlx5: Fix access to a non-supported register
Ben Ben-Ishay (1):
net/mlx5e: SHAMPO, Fix constant expression result
Dmytro Linkin (2):
net/mlx5: E-switch, Respect BW share of the new group
net/mlx5: E-Switch, Check group pointer before reading bw_share value
Gal Pressman (1):
net/mlx5: Fix too early queueing of log timestamp work
Maor Dickman (1):
net/mlx5: E-Switch, Use indirect table only if all destinations support it
Maor Gottlieb (1):
net/mlx5: Lag, Fix recreation of VF LAG
Mark Bloch (1):
net/mlx5: E-Switch, fix single FDB creation on BlueField
Moshe Shemesh (1):
net/mlx5: Move MODIFY_RQT command to ignore list in internal error state
Raed Salem (2):
net/mlx5e: IPsec: Fix Software parser inner l3 type setting in case of encapsulation
net/mlx5e: Fix missing IPsec statistics on uplink representor
Tariq Toukan (1):
net/mlx5e: Sync TIR params updates against concurrent create/modify
drivers/net/ethernet/mellanox/mlx5/core/cmd.c | 2 +-
.../net/ethernet/mellanox/mlx5/core/en/rx_res.c | 41 +++++++++++++++++++++-
.../net/ethernet/mellanox/mlx5/core/en/rx_res.h | 6 ++--
.../mellanox/mlx5/core/en_accel/ipsec_rxtx.c | 2 +-
.../ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c | 24 +------------
drivers/net/ethernet/mellanox/mlx5/core/en_rep.c | 4 +++
drivers/net/ethernet/mellanox/mlx5/core/en_rx.c | 8 ++---
drivers/net/ethernet/mellanox/mlx5/core/esw/qos.c | 4 +--
.../ethernet/mellanox/mlx5/core/eswitch_offloads.c | 20 ++++++++---
drivers/net/ethernet/mellanox/mlx5/core/health.c | 5 +--
.../net/ethernet/mellanox/mlx5/core/lag/port_sel.c | 1 +
drivers/net/ethernet/mellanox/mlx5/core/lib/tout.c | 5 ++-
drivers/net/ethernet/mellanox/mlx5/core/lib/tout.h | 1 +
drivers/net/ethernet/mellanox/mlx5/core/main.c | 30 ++++++++--------
include/linux/mlx5/mlx5_ifc.h | 5 ++-
15 files changed, 97 insertions(+), 61 deletions(-)
From: Raed Salem <redacted>
Current code wrongly uses the skb->protocol field which reflects the
outer l3 protocol to set the inner l3 type in Software Parser (SWP)
fields settings in the ethernet segment (eseg) in flows where inner
l3 exists like in Vxlan over ESP flow, the above method wrongly use
the outer protocol type instead of the inner one. thus breaking cases
where inner and outer headers have different protocols.
Fix by setting the inner l3 type in SWP according to the inner l3 ip
header version.
Fixes: 2ac9cfe78223 ("net/mlx5e: IPSec, Add Innova IPSec offload TX data path")
Signed-off-by: Raed Salem <redacted>
Reviewed-by: Maor Dickman <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_rxtx.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Raed Salem <redacted>
The cited patch added the IPsec support to uplink representor, however
as uplink representors have his private statistics where IPsec stats
is not part of it, that effectively makes IPsec stats hidden when uplink
representor stats queried.
Resolve by adding IPsec stats to uplink representor private statistics.
Fixes: 5589b8f1a2c7 ("net/mlx5e: Add IPsec support to uplink representor")
Signed-off-by: Raed Salem <redacted>
Reviewed-by: Alaa Hleihel <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en_rep.c | 4 ++++
1 file changed, 4 insertions(+)
From: Tariq Toukan <tariqt@nvidia.com>
Transport Interface Receive (TIR) objects perform the packet processing and
reassembly and is also responsible for demultiplexing the packets into the
different RQs.
There are certain TIR context attributes that propagate to the pointed RQs
and applied to them (like packet_merge offloads (LRO/SHAMPO) and
tunneled_offload_en). When TIRs do not agree on attributes values, a "last
one wins" policy is applied. Hence, if not synced properly, a race between
TIR params update and a concurrent TIR create/modify operation might yield
to a mismatch between the shadow parameters in SW and the actual applied
state of the RQs in HW.
tunneled_offload_en is a fixed attribute per profile, while packet merge
offload state might be toggled and get out-of-sync. When this happens,
packet_merge offload might be working although not requested, or the
opposite.
All updates to packet_merge state and all create/modify operations of
regular redirection/steering TIRs are done under the same priv->state_lock,
so they do not run in parallel, and no race is possible.
However, there are other kind of TIRs (acceleration offloads TIRs, like TLS
TIRs) which are created on demand for each new connection without holding
the coarse priv->state_lock, hence might race.
Fix this by synchronizing all packet_merge state reads and writes against
all TIR create/modify operations. Include the modify operations of the
regular redirection steering TIRs under the new lock, for better code
layering and division of responsibilities.
Fixes: 1182f3659357 ("net/mlx5e: kTLS, Add kTLS RX HW offload support")
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Moshe Shemesh <redacted>
Reviewed-by: Maxim Mikityanskiy <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
.../ethernet/mellanox/mlx5/core/en/rx_res.c | 41 ++++++++++++++++++-
.../ethernet/mellanox/mlx5/core/en/rx_res.h | 6 +--
.../mellanox/mlx5/core/en_accel/ktls_rx.c | 24 +----------
3 files changed, 44 insertions(+), 27 deletions(-)
@@ -392,6 +395,7 @@ static int mlx5e_rx_res_ptp_init(struct mlx5e_rx_res *res)if(err)gotoout;+/* Separated from the channels RQs, does not share pkt_merge state with them */mlx5e_tir_builder_build_rqt(builder,res->mdev->mlx5e_res.hw_objs.td.tdn,mlx5e_rqt_get_rqtn(&res->ptp.rqt),inner_ft_support);
@@ -37,9 +37,6 @@ u32 mlx5e_rx_res_get_tirn_rss(struct mlx5e_rx_res *res, enum mlx5_traffic_typesu32mlx5e_rx_res_get_tirn_rss_inner(structmlx5e_rx_res*res,enummlx5_traffic_typestt);u32mlx5e_rx_res_get_tirn_ptp(structmlx5e_rx_res*res);-/* RQTN getters for modules that create their own TIRs */-u32mlx5e_rx_res_get_rqtn_direct(structmlx5e_rx_res*res,unsignedintix);-/* Activate/deactivate API */voidmlx5e_rx_res_channels_activate(structmlx5e_rx_res*res,structmlx5e_channels*chs);voidmlx5e_rx_res_channels_deactivate(structmlx5e_rx_res*res);
From: Moshe Shemesh <redacted>
When the device is in internal error state, command interface isn't
accessible and the driver decides which commands to fail and which
to ignore.
Move the MODIFY_RQT command to the ignore list in order to avoid
the following redundant warning messages in internal error state:
mlx5_core 0000:82:00.1: mlx5e_rss_disable:419:(pid 23754): Failed to redirect RQT 0x0 to drop RQ 0xc00848: err = -5
mlx5_core 0000:82:00.1: mlx5e_rx_res_channels_deactivate:598:(pid 23754): Failed to redirect direct RQT 0x1 to drop RQ 0xc00848 (channel 0): err = -5
mlx5_core 0000:82:00.1: mlx5e_rx_res_channels_deactivate:607:(pid 23754): Failed to redirect XSK RQT 0x19 to drop RQ 0xc00848 (channel 0): err = -5
Fixes: 43ec0f41fa73 ("net/mlx5e: Hide all implementation details of mlx5e_rx_res")
Signed-off-by: Moshe Shemesh <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/cmd.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Dmytro Linkin <redacted>
To enable transmit schduler on vport FW require non-zero configuration
for vport's TSAR. If vport added to the group which has configured BW
share value and TX rate values of the vport are zero, then scheduler
wouldn't be enabled on this vport.
Fix that by calling BW normalization if BW share of the new group is
configured.
Fixes: 0fe132eac38c ("net/mlx5: E-switch, Allow to add vports to rate groups")
Signed-off-by: Dmytro Linkin <redacted>
Reviewed-by: Roi Dayan <redacted>
Reviewed-by: Parav Pandit <redacted>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/esw/qos.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
@@ -423,7 +423,7 @@ static int esw_qos_vport_update_group(struct mlx5_eswitch *esw,returnerr;/* Recalculate bw share weights of old and new groups */-if(vport->qos.bw_share){+if(vport->qos.bw_share||new_group->bw_share){esw_qos_normalize_vports_min_rate(esw,curr_group,extack);esw_qos_normalize_vports_min_rate(esw,new_group,extack);}
From: Dmytro Linkin <redacted>
If log_esw_max_sched_depth is not supported group pointer of the vport
is NULL. Hence, check the pointer before reading bw_share value.
Fixes: 0fe132eac38c ("net/mlx5: E-switch, Allow to add vports to rate groups")
Signed-off-by: Dmytro Linkin <redacted>
Reviewed-by: Roi Dayan <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/esw/qos.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
@@ -130,7 +130,7 @@ static u32 esw_qos_calculate_min_rate_divider(struct mlx5_eswitch *esw,/* If vports min rate divider is 0 but their group has bw_share configured, then*needtosetbw_shareforvportstominimalvalue.*/-if(!group_level&&!max_guarantee&&group->bw_share)+if(!group_level&&!max_guarantee&&group&&group->bw_share)return1;return0;}
From: Mark Bloch <mbloch@nvidia.com>
Always use MLX5_FLOW_TABLE_OTHER_VPORT flag when creating egress ACL
table for single FDB. Not doing so on BlueField will make firmware fail
the command. On BlueField the E-Switch manager is the ECPF (vport 0xFFFE)
which is filled in the flow table creation command but as the
other_vport field wasn't set the firmware complains about a bad parameter.
This is different from a regular HCA where the E-Switch manager vport is
the PF (vport 0x0). Passing MLX5_FLOW_TABLE_OTHER_VPORT will make the
firmware happy both on BlueField and on regular HCAs without special
condition for each.
This fixes the bellow firmware syndrome:
mlx5_cmd_check:819:(pid 571): CREATE_FLOW_TABLE(0x930) op_mod(0x0) failed, status bad parameter(0x3), syndrome (0x754a4)
Fixes: db202995f503 ("net/mlx5: E-Switch, add logic to enable shared FDB")
Signed-off-by: Mark Bloch <mbloch@nvidia.com>
Reviewed-by: Maor Gottlieb <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/eswitch_offloads.c | 1 +
1 file changed, 1 insertion(+)
From: Maor Dickman <redacted>
When adding rule with multiple destinations, indirect table is used for all of
the destinations if at least one of the destinations support it, this can cause
creation of invalid indirect tables for the destinations that doesn't support it.
Fixed it by using indirect table only if all destinations support it.
Fixes: a508728a4c8b ("net/mlx5e: VF tunnel RX traffic offloading")
Signed-off-by: Maor Dickman <redacted>
Reviewed-by: Roi Dayan <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
.../mellanox/mlx5/core/eswitch_offloads.c | 19 +++++++++++++++----
1 file changed, 15 insertions(+), 4 deletions(-)
@@ -329,14 +329,25 @@ static boolesw_is_indir_table(structmlx5_eswitch*esw,structmlx5_flow_attr*attr){structmlx5_esw_flow_attr*esw_attr=attr->esw_attr;+boolresult=false;inti;-for(i=esw_attr->split_count;i<esw_attr->out_count;i++)+/* Indirect table is supported only for flows with in_port uplink+*andthedestinationisvportonthesameeswitchastheuplink,+*returnfalseincaseatleastoneofdestinationsdoesn'tmeet+*thiscriteria.+*/+for(i=esw_attr->split_count;i<esw_attr->out_count;i++){if(esw_attr->dests[i].rep&&mlx5_esw_indir_table_needed(esw,attr,esw_attr->dests[i].rep->vport,-esw_attr->dests[i].mdev))-returntrue;-returnfalse;+esw_attr->dests[i].mdev)){+result=true;+}else{+result=false;+break;+}+}+returnresult;}staticint
From: Ben Ben-Ishay <redacted>
mlx5e_build_shampo_hd_umr uses counters i and index incorrectly
as unsigned, thus the err state err_unmap could stuck in endless loop.
Change i to int to solve the first issue.
Reduce index check to solve the second issue, the caller function
validates that index could not rotate.
Fixes: 64509b052525 ("net/mlx5e: Add data path for SHAMPO feature")
Signed-off-by: Ben Ben-Ishay <redacted>
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en_rx.c | 8 +++-----
1 file changed, 3 insertions(+), 5 deletions(-)
From: Amir Tzin <redacted>
The device health recovery flow calls mlx5_health_wait_pci_up() which
queries the device for FW_RESET timeout after freeing the device
timeouts structure on mlx5_function_teardown(). Fix this bug by moving
timeouts structure init/cleanup to the device's init/uninit phases.
Since it is necessary to reset default software timeouts on function
reload, extract setting of defaults values from mlx5_tout_init() and
call mlx5_tout_set_def_val() directly from mlx5_function_setup().
Fixes: 5945e1adeab5 ("net/mlx5: Read timeout values from init segment")
Reported by: Niklas Schnelle [off-list ref]
Signed-off-by: Amir Tzin <redacted>
Signed-off-by: Moshe Shemesh <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
.../ethernet/mellanox/mlx5/core/lib/tout.c | 5 ++---
.../ethernet/mellanox/mlx5/core/lib/tout.h | 1 +
.../net/ethernet/mellanox/mlx5/core/main.c | 22 ++++++++++---------
3 files changed, 15 insertions(+), 13 deletions(-)
@@ -1476,6 +1469,12 @@ int mlx5_mdev_init(struct mlx5_core_dev *dev, int profile_idx)mlx5_debugfs_root);INIT_LIST_HEAD(&priv->traps);+err=mlx5_tout_init(dev);+if(err){+mlx5_core_err(dev,"Failed initializing timeouts, aborting\n");+gotoerr_timeout_init;+}+err=mlx5_health_init(dev);if(err)gotoerr_health_init;
@@ -1501,6 +1500,8 @@ int mlx5_mdev_init(struct mlx5_core_dev *dev, int profile_idx)err_pagealloc_init:mlx5_health_cleanup(dev);err_health_init:+mlx5_tout_cleanup(dev);+err_timeout_init:debugfs_remove(dev->priv.dbg_root);mutex_destroy(&priv->pgdir_mutex);mutex_destroy(&priv->alloc_mutex);
From: Gal Pressman <redacted>
The log timestamp work should not be queued before the command interface
is initialized, move it to a later stage in the init flow.
Fixes: 5a1023deeed0 ("net/mlx5: Add periodic update of host time to firmware")
Signed-off-by: Gal Pressman <redacted>
Reviewed-by: Moshe Shemesh <redacted>
Signed-off-by: Saeed Mahameed <saeedm@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/health.c | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
Hello:
This series was applied to netdev/net.git (master)
by Saeed Mahameed [off-list ref]:
On Tue, 30 Nov 2021 22:36:57 -0800 you wrote:
From: Raed Salem <redacted>
Current code wrongly uses the skb->protocol field which reflects the
outer l3 protocol to set the inner l3 type in Software Parser (SWP)
fields settings in the ethernet segment (eseg) in flows where inner
l3 exists like in Vxlan over ESP flow, the above method wrongly use
the outer protocol type instead of the inner one. thus breaking cases
where inner and outer headers have different protocols.
[...]