i.MX94 NETC (v4.3) integrates 802.1Q Ethernet switch functionality, the
switch provides advanced QoS with 8 traffic classes and a full range of
TSN standards capabilities. It has 3 user ports and 1 CPU port, and the
CPU port is connected to an internal ENETC through the pseduo link, so
instead of a back-to-back MAC, the lightweight "pseudo MAC" is used at
both ends of the pseudo link to transfer Ethernet frames. The pseudo
link provides a zero-copy interface (no serialization delay) and lower
power (less logic and memory).
Like most Ethernet switches, the NETC switch also supports a proprietary
switch tag, is used to carry in-band metadata information about frames.
This in-band metadata information can include the source port from which
the frame was received, what was the reason why this frame got forwarded
to the entity, and for the entity to indicate the precise destination
port of a frame. The NETC switch tag is added to frames after the source
MAC address. There are three types of switch tags, and each type has 1
to 4 subtypes, more details are as follows.
Forward switch tag (Type = 0): Represents forwarded frames.
- SubType = 0 - Normal frame processing.
To_Port switch tag (Type = 1): Represents frames that are to be sent to
a specific switch port.
- SubType = 0. No request to perform timestamping.
- SubType = 1. Request to perform one-step timestamping.
- SubType = 2. Request to perform two-step timestamping.
- SubType = 3. Request to perform both one-step timestamping and
two-step timestamping.
To_Host switch tag (Type = 2): Represents frames redirected or copied to
the switch management port.
- SubType = 0. Received frames redirected or copied to the switch
management port.
- SubType = 1. Received frames redirected or copied to the switch
management port with captured timestamp at the switch port where
the frame was received.
- SubType = 2. Transmit timestamp response (two-step timestamping).
Currently, this patch set supports Forward tag, SubType 0 of To_Port tag
and SubType 0 of To_Host tag. More tags will be supported in the future.
In addition, the switch supports NETC Table Management Protocol (NTMP),
some switch functionality is controlled using control messages sent to
the hardware using BD ring interface with 32B descriptors similar to the
packet Transmit BD ring used on ENETC. This interface is referred to as
the command BD ring. This is used to configure functionality where the
underlying resources may be shared between different entities or being
too large to configure using direct registers.
For this patch set, we have supported the following tables through the
command BD ring interface.
FDB Table: It contains forwarding and/or filtering information about MAC
addresses. The FDB table is used for MAC learning lookups and MAC
forwarding lookups.
VLAN Filter Table: It contains configuration and control information for
each VLAN configured on the switch.
Buffer Pool Table: It contains buffer pool configuration and operational
information. Each entry corresponds to a buffer pool. Currently, we use
this table to implement flow control feature on each port.
Ingress Port Filter Table: It contains a set of filters each capable of
classifying incoming traffic using a mix of L2, L3, and L4 parsed and
arbitrary field data. We use this table to implement host flood support
to the switch port.
The switch also supports other tables, and we will add more advanced
features through them in the future.
---
v2:
1. Use raw_smp_processor_id() in netc_select_cbdr() instead of
smp_processor_id().
2. Remove netc_port_free_mdio_bus() and netc_free_mdio_bus().
3. Correct the mask value in netc_port_set_mac_mode()
4. Rename net_port_set_rmii_mii_mac() to netc_port_set_rmii_mii_mac().
5. Check the return value of ntmp_bpt_update_entry() in
netc_switch_bpt_default_config().
6. Add some comments to avoid false positives from AI review.
v1 link: https://lore.kernel.org/imx/20260316094152.1558671-1-wei.fang@nxp.com/
---
Wei Fang (14):
dt-bindings: net: dsa: update the description of 'dsa,member' property
dt-bindings: net: dsa: add NETC switch
net: enetc: add pre-boot initialization for i.MX94 switch
net: enetc: add basic operations to the FDB table
net: enetc: add support for the "Add" operation to VLAN filter table
net: enetc: add support for the "Update" operation to buffer pool
table
net: enetc: add support for "Add" and "Delete" operations to IPFT
net: enetc: add multiple command BD rings support
net: dsa: add NETC switch tag support
net: dsa: netc: introduce NXP NETC switch driver for i.MX94
net: dsa: netc: add phylink MAC operations
net: dsa: netc: add more basic functions support
net: dsa: netc: initialize buffer bool table and implement
flow-control
net: dsa: netc: add support for the standardized counters
.../devicetree/bindings/net/dsa/dsa.yaml | 6 +-
.../bindings/net/dsa/nxp,netc-switch.yaml | 128 ++
MAINTAINERS | 11 +
drivers/net/dsa/Kconfig | 3 +
drivers/net/dsa/Makefile | 1 +
drivers/net/dsa/netc/Kconfig | 14 +
drivers/net/dsa/netc/Makefile | 3 +
drivers/net/dsa/netc/netc_ethtool.c | 192 ++
drivers/net/dsa/netc/netc_main.c | 1561 +++++++++++++++++
drivers/net/dsa/netc/netc_platform.c | 89 +
drivers/net/dsa/netc/netc_switch.h | 155 ++
drivers/net/dsa/netc/netc_switch_hw.h | 356 ++++
.../ethernet/freescale/enetc/netc_blk_ctrl.c | 188 +-
drivers/net/ethernet/freescale/enetc/ntmp.c | 391 ++++-
.../ethernet/freescale/enetc/ntmp_private.h | 120 ++
include/linux/dsa/tag_netc.h | 14 +
include/linux/fsl/netc_global.h | 6 +
include/linux/fsl/ntmp.h | 233 +++
include/net/dsa.h | 2 +
include/uapi/linux/if_ether.h | 1 +
net/dsa/Kconfig | 10 +
net/dsa/Makefile | 1 +
net/dsa/tag_netc.c | 180 ++
23 files changed, 3638 insertions(+), 27 deletions(-)
create mode 100644 Documentation/devicetree/bindings/net/dsa/nxp,netc-switch.yaml
create mode 100644 drivers/net/dsa/netc/Kconfig
create mode 100644 drivers/net/dsa/netc/Makefile
create mode 100644 drivers/net/dsa/netc/netc_ethtool.c
create mode 100644 drivers/net/dsa/netc/netc_main.c
create mode 100644 drivers/net/dsa/netc/netc_platform.c
create mode 100644 drivers/net/dsa/netc/netc_switch.h
create mode 100644 drivers/net/dsa/netc/netc_switch_hw.h
create mode 100644 include/linux/dsa/tag_netc.h
create mode 100644 net/dsa/tag_netc.c
--
2.34.1
The current description indicates that the 'dsa,member' property cannot
be set for a switch that is not part of any cluster. Vladimir thinks
that this is a case where the actual technical limitation was poorly
transposed into words when this restriction was first documented, in
commit 8c5ad1d6179d ("net: dsa: Document new binding").
The true technical limitation is that many DSA tagging protocols are
topology-unaware, and always call dsa_conduit_find_user() with a
switch_id of 0. Specifying a custom "dsa,member" property with a
non-zero switch_id would break them.
Therefore, for topology-aware switches, it is fine to specify this
property for them, even if they are not part of any cluster. Our NETC
switch is a good example which is topology-aware, the switch_id is
carried in the switch tag, but the switch_id 0 is reserved for VEPA
switch and cannot be used, so we need to use this property to assign
a non-zero switch_id for it.
Suggested-by: Vladimir Oltean <vladimir.oltean@nxp.com>
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
Documentation/devicetree/bindings/net/dsa/dsa.yaml | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
@@ -28,7 +28,11 @@ properties:A two element list indicates which DSA cluster, and position within thecluster a switch takes. <0 0> is cluster 0, switch 0. <0 1> is cluster 0,switch 1. <1 0> is cluster 1, switch 0. A switch not part of any cluster-(single device hanging off a CPU port) must not specify this property+(single device hanging off a CPU port) does not usually need to specify+this property, and then it becomes cluster 0, switch 0. For a topology+aware switch, its switch index can be specified through this property,+even if it is not part of any cluster. Also, topology-unaware switches+must always be defined as index 0 of their cluster.$ref:/schemas/types.yaml#/definitions/uint32-arrayadditionalProperties:true
Add bindings for NETC switch. This switch is a PCIe function of NETC IP,
it supports advanced QoS with 8 traffic classes and 4 drop resilience
levels, and a full range of TSN standards capabilities. The switch CPU
port connects to an internal ENETC port, which is also a PCIe function
of NETC IP. So these two ports use a light-weight "pseudo MAC" instead
of a back-to-back MAC, because the "pseudo MAC" provides the delineation
between switch and ENETC, this translates to lower power (less logic and
memory) and lower delay (as there is no serialization delay across this
link).
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
.../bindings/net/dsa/nxp,netc-switch.yaml | 128 ++++++++++++++++++
1 file changed, 128 insertions(+)
create mode 100644 Documentation/devicetree/bindings/net/dsa/nxp,netc-switch.yaml
@@ -0,0 +1,128 @@+# SPDX-License-Identifier: (GPL-2.0-only OR BSD-2-Clause)+%YAML1.2+---+$id:http://devicetree.org/schemas/net/dsa/nxp,netc-switch.yaml#+$schema:http://devicetree.org/meta-schemas/core.yaml#++title:NETC Switch family++description:+The NETC presents itself as a multi-function PCIe Root Complex Integrated+Endpoint (RCiEP) and provides full 802.1Q Ethernet switch functionality,+advanced QoS with 8 traffic classes and 4 drop resilience levels, and a+full range of TSN standards capabilities.++The CPU port of the switch connects to an internal ENETC. The switch and+the internal ENETC are fully integrated into the NETC IP, a back-to-back+MAC is not required. Instead, a light-weight "pseudo MAC" provides the+delineation between the switch and ENETC. This translates to lower power+(less logic and memory) and lower delay (as there is no serialization+delay across this link).++maintainers:+-Wei Fang <wei.fang@nxp.com>++properties:+compatible:+enum:+-pci1131,eef2++reg:+maxItems:1++dsa,member:+description:+The property indicates DSA cluster and switch index. For NETC switch,+The valid range of the switch index is 1 ~ 7, the value 0 is reserved+for VEPA switch.++$ref:dsa.yaml#++patternProperties:+"^(ethernet-)?ports$":+type:object+additionalProperties:true+patternProperties:+"^(ethernet-)?port@[0-9a-f]$":+type:object++$ref:dsa-port.yaml#++properties:+clocks:+items:+-description:MAC transmit/receive reference clock.++clock-names:+items:+-const:ref++mdio:+$ref:/schemas/net/mdio.yaml#+unevaluatedProperties:false+description:+Optional child node for switch port, otherwise use NETC EMDIO.++unevaluatedProperties:false++required:+-compatible+-reg+-dsa,member++allOf:+-$ref:/schemas/pci/pci-device.yaml++unevaluatedProperties:false++examples:+-|+pcie {+#address-cells = <3>;+#size-cells = <2>;++ethernet-switch@0,2 {+compatible = "pci1131,eef2";+reg = <0x200 0 0 0 0>;+dsa,member = <0 1>;+pinctrl-names = "default";+pinctrl-0 = <&pinctrl_switch>;++ports {+#address-cells = <1>;+#size-cells = <0>;++port@0 {+reg = <0>;+phy-handle = <ðphy0>;+phy-mode = "mii";+};++port@1 {+reg = <1>;+phy-handle = <ðphy1>;+phy-mode = "mii";+};++port@2 {+reg = <2>;+clocks = <&scmi_clk 103>;+clock-names = "ref";+phy-handle = <ðphy2>;+phy-mode = "rgmii-id";+};++port@3 {+reg = <3>;+ethernet = <&enetc3>;+phy-mode = "internal";++fixed-link {+speed = <2500>;+full-duplex;+pause;+};+};+};+};+};
Before probing the NETC switch driver, some pre-initialization needs to
be set in NETCMIX and IERB to ensure that the switch can work properly.
For example, i.MX94 NETC switch has three external ports and each port
is bound to a link. And each link needs to be configured so that it can
work properly, such as I/O variant and MII protocol.
In addition, the switch port 2 (MAC 2) and ENETC 0 (MAC 3) share the same
parallel interface, they cannot be used at the same time due to the SoC
constraint. And the MAC selection is controlled by the mac2_mac3_sel bit
of EXT_PIN_CONTROL register. Currently, the interface is set for ENETC 0
by default unless the switch port 2 is enabled in the DT node.
Like ENETC, each external port of the NETC switch can manage its external
PHY through its port MDIO registers. And the port can only access its own
external PHY by setting the PHY address to the LaBCR[MDIO_PHYAD_PRTAD].
If the accessed PHY address is not equal to LaBCR[MDIO_PHYAD_PRTAD], then
the MDIO access initiated by port MDIO will be invalid.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
.../ethernet/freescale/enetc/netc_blk_ctrl.c | 188 ++++++++++++++++--
1 file changed, 166 insertions(+), 22 deletions(-)
@@ -261,40 +261,112 @@ static int imx94_link_config(struct netc_blk_ctrl *priv,}staticintimx94_enetc_link_config(structnetc_blk_ctrl*priv,-structdevice_node*np)+structdevice_node*np,+bool*enetc0_en){intlink_id=imx94_enetc_get_link_id(np);if(link_id<0)returnlink_id;+if(link_id==IMX94_ENETC0_LINK&&of_device_is_available(np))+*enetc0_en=true;+returnimx94_link_config(priv,np,link_id);}+staticstructdevice_node*netc_get_switch_ports(structdevice_node*np)+{+structdevice_node*ports;++ports=of_get_child_by_name(np,"ports");+if(!ports)+ports=of_get_child_by_name(np,"ethernet-ports");++returnports;+}++staticintimx94_switch_link_config(structnetc_blk_ctrl*priv,+structdevice_node*np,+bool*swp2_en)+{+structdevice_node*ports;+intport_id,err=0;++ports=netc_get_switch_ports(np);+if(!ports)+return-ENODEV;++for_each_available_child_of_node_scoped(ports,child){+if(of_property_read_u32(child,"reg",&port_id)<0){+err=-ENODEV;+gotoend;+}++switch(port_id){+case0...2:/* External ports */+err=imx94_link_config(priv,child,port_id);+if(err)+gotoend;++if(port_id==2)+*swp2_en=true;++break;+case3:/* CPU port */+break;+default:+err=-EINVAL;+gotoend;+}+}++end:+of_node_put(ports);++returnerr;+}+staticintimx94_netcmix_init(structplatform_device*pdev){structnetc_blk_ctrl*priv=platform_get_drvdata(pdev);structdevice_node*np=pdev->dev.of_node;+boolenetc0_en=false,swp2_en=false;u32val;interr;for_each_child_of_node_scoped(np,child){for_each_child_of_node_scoped(child,gchild){-if(!of_device_is_compatible(gchild,"pci1131,e101"))-continue;--err=imx94_enetc_link_config(priv,gchild);-if(err)-returnerr;+if(of_device_is_compatible(gchild,"pci1131,e101")){+err=imx94_enetc_link_config(priv,gchild,+&enetc0_en);+if(err)+returnerr;+}elseif(of_device_is_compatible(gchild,+"pci1131,eef2")){+err=imx94_switch_link_config(priv,gchild,+&swp2_en);+if(err)+returnerr;+}}}-/* ENETC 0 and switch port 2 share the same parallel interface.-*Currently,theswitchisnotsupported,sothisinterfaceis-*usedbyENETC0bydefault.+if(enetc0_en&&swp2_en){+dev_err(&pdev->dev,+"Cannot enable swp2 and enetc0 at the same time\n");+return-EINVAL;+}++/* ENETC 0 and switch port 2 share the same parallel interface, they+*cannotbeenabledatthesametime.Theinterfaceissetforthe+*ENETC0bydefaultunlesstheswitchport2isenabledintheDTS.*/val=netc_reg_read(priv->netcmix,IMX94_EXT_PIN_CONTROL);-val|=MAC2_MAC3_SEL;+if(!swp2_en)+val|=MAC2_MAC3_SEL;+else+val&=~MAC2_MAC3_SEL;netc_reg_write(priv->netcmix,IMX94_EXT_PIN_CONTROL,val);return0;
@@ -610,6 +682,77 @@ static int imx94_enetc_mdio_phyaddr_config(struct netc_blk_ctrl *priv,return0;}+staticintimx94_ierb_enetc_init(structnetc_blk_ctrl*priv,+structdevice_node*np,+u32phy_mask)+{+interr;++err=imx94_enetc_update_tid(priv,np);+if(err)+returnerr;++returnimx94_enetc_mdio_phyaddr_config(priv,np,phy_mask);+}++staticintimx94_switch_mdio_phyaddr_config(structnetc_blk_ctrl*priv,+structdevice_node*np,+intport_id,u32phy_mask)+{+intaddr;++/* The switch has 3 external ports at most */+if(port_id>2)+return0;++addr=netc_get_phy_addr(np);+if(addr<0){+if(addr==-ENODEV)+return0;++returnaddr;+}++if(phy_mask&BIT(addr)){+dev_err(&priv->pdev->dev,+"Found same PHY address in EMDIO and switch node\n");+return-EINVAL;+}++netc_reg_write(priv->ierb,IERB_LBCR(port_id),+LBCR_MDIO_PHYAD_PRTAD(addr));++return0;+}++staticintimx94_ierb_switch_init(structnetc_blk_ctrl*priv,+structdevice_node*np,+u32phy_mask)+{+structdevice_node*ports;+intport_id,err=0;++ports=netc_get_switch_ports(np);+if(!ports)+return-ENODEV;++for_each_available_child_of_node_scoped(ports,child){+err=of_property_read_u32(child,"reg",&port_id);+if(err)+gotoend;++err=imx94_switch_mdio_phyaddr_config(priv,child,+port_id,phy_mask);+if(err)+gotoend;+}++end:+of_node_put(ports);++returnerr;+}+staticintimx94_ierb_init(structplatform_device*pdev){structnetc_blk_ctrl*priv=platform_get_drvdata(pdev);
@@ -625,17 +768,18 @@ static int imx94_ierb_init(struct platform_device *pdev)for_each_child_of_node_scoped(np,child){for_each_child_of_node_scoped(child,gchild){-if(!of_device_is_compatible(gchild,"pci1131,e101"))-continue;--err=imx94_enetc_update_tid(priv,gchild);-if(err)-returnerr;--err=imx94_enetc_mdio_phyaddr_config(priv,gchild,-phy_mask);-if(err)-returnerr;+if(of_device_is_compatible(gchild,"pci1131,e101")){+err=imx94_ierb_enetc_init(priv,gchild,+phy_mask);+if(err)+returnerr;+}elseif(of_device_is_compatible(gchild,+"pci1131,eef2")){+err=imx94_ierb_switch_init(priv,gchild,+phy_mask);+if(err)+returnerr;+}}}
The FDB table is used for MAC learning lookups and MAC forwarding lookups.
Each table entry includes information such as a FID and MAC address that
may be unicast or multicast and a forwarding destination field containing
a port bitmap identifying the associated port(s) with the MAC address.
FDB table entries can be static or dynamic. Static entries are added from
software whereby dynamic entries are added either by software or by the
hardware as MAC addresses are learned in the datapath.
The FDB table can only be managed by the command BD ring using table
management protocol version 2.0. Table management command operations Add,
Delete, Update and Query are supported. And the FDB table supports three
access methods: Entry ID, Exact Match Key Element and Search. This patch
adds the following basic supports to the FDB table.
ntmp_fdbt_update_entry() - update the configuration element data of a
specified FDB entry
ntmp_fdbt_delete_entry() - delete a specified FDB entry
ntmp_fdbt_add_entry() - add an entry into the FDB table
ntmp_fdbt_search_port_entry() - Search the FDB entry on the specified
port based on RESUME_ENTRY_ID.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/ethernet/freescale/enetc/ntmp.c | 199 ++++++++++++++++++
.../ethernet/freescale/enetc/ntmp_private.h | 59 ++++++
include/linux/fsl/ntmp.h | 67 ++++++
3 files changed, 325 insertions(+)
@@ -453,5 +459,198 @@ int ntmp_rsst_query_entry(struct ntmp_user *user, u32 *table, int count)}EXPORT_SYMBOL_GPL(ntmp_rsst_query_entry);+/**+*ntmp_fdbt_add_entry-addanentryintotheFDBtable+*@user:targetntmp_userstruct+*@entry_id:returnedvalue,theentryIDofthenewaddedentry+*@keye:keyelementdata+*@cfge:configurationelementdata+*+*Return:0onsuccess,otherwiseanegativeerrorcode+*/+intntmp_fdbt_add_entry(structntmp_user*user,u32*entry_id,+conststructfdbt_keye_data*keye,+conststructfdbt_cfge_data*cfge)+{+structntmp_dma_bufdata={+.dev=user->dev,+.size=sizeof(structfdbt_req_ua),+};+structfdbt_resp_query*resp;+structfdbt_req_ua*req;+unionnetc_cbdcbd;+u32len;+interr;++err=ntmp_alloc_data_mem(&data,(void**)&req);+if(err)+returnerr;++/* Request data */+ntmp_fill_crd(&req->crd,user->tbl.fdbt_ver,NTMP_QA_ENTRY_ID,+NTMP_GEN_UA_CFGEU);+req->ak.exact.keye=*keye;+req->cfge=*cfge;++len=NTMP_LEN(data.size,sizeof(*resp));+/* The entry ID is allotted by hardware, so we need to perform+*aqueryactionaftertheaddactiontogettheentryIDfrom+*hardware.+*/+ntmp_fill_request_hdr(&cbd,data.dma,len,NTMP_FDBT_ID,+NTMP_CMD_AQ,NTMP_AM_EXACT_KEY);+err=netc_xmit_ntmp_cmd(user,&cbd);+if(err){+dev_err(user->dev,"Failed to add %s entry, err: %pe\n",+ntmp_table_name(NTMP_FDBT_ID),ERR_PTR(err));+gotoend;+}++if(entry_id){+resp=(structfdbt_resp_query*)req;+*entry_id=le32_to_cpu(resp->entry_id);+}++end:+ntmp_free_data_mem(&data);++returnerr;+}+EXPORT_SYMBOL_GPL(ntmp_fdbt_add_entry);++/**+*ntmp_fdbt_update_entry-updatetheconfigurationelementdataofthe+*specifiedFDBentry+*@user:targetntmp_userstruct+*@entry_id:thespecifiedentryIDoftheFDBtable+*@cfge:configurationelementdata+*+*Return:0onsuccess,otherwiseanegativeerrorcode+*/+intntmp_fdbt_update_entry(structntmp_user*user,u32entry_id,+conststructfdbt_cfge_data*cfge)+{+structntmp_dma_bufdata={+.dev=user->dev,+.size=sizeof(structfdbt_req_ua),+};+structfdbt_req_ua*req;+unionnetc_cbdcbd;+u32len;+interr;++err=ntmp_alloc_data_mem(&data,(void**)&req);+if(err)+returnerr;++/* Request data */+ntmp_fill_crd(&req->crd,user->tbl.fdbt_ver,0,NTMP_GEN_UA_CFGEU);+req->ak.eid.entry_id=cpu_to_le32(entry_id);+req->cfge=*cfge;++/* Request header */+len=NTMP_LEN(data.size,NTMP_STATUS_RESP_LEN);+ntmp_fill_request_hdr(&cbd,data.dma,len,NTMP_FDBT_ID,+NTMP_CMD_UPDATE,NTMP_AM_ENTRY_ID);+err=netc_xmit_ntmp_cmd(user,&cbd);+if(err)+dev_err(user->dev,"Failed to update %s entry, err: %pe\n",+ntmp_table_name(NTMP_FDBT_ID),ERR_PTR(err));++ntmp_free_data_mem(&data);++returnerr;+}+EXPORT_SYMBOL_GPL(ntmp_fdbt_update_entry);++/**+*ntmp_fdbt_delete_entry-deletethespecifiedFDBentry+*@user:targetntmp_userstruct+*@entry_id:thespecifiedIDoftheFDBentry+*+*Return:0onsuccess,otherwiseanegativeerrorcode+*/+intntmp_fdbt_delete_entry(structntmp_user*user,u32entry_id)+{+u32req_len=sizeof(structfdbt_req_qd);++returnntmp_delete_entry_by_id(user,NTMP_FDBT_ID,+user->tbl.fdbt_ver,+entry_id,req_len,+NTMP_STATUS_RESP_LEN);+}+EXPORT_SYMBOL_GPL(ntmp_fdbt_delete_entry);++/**+*ntmp_fdbt_search_port_entry-SearchtheFDBentryonthespecified+*portbasedonRESUME_ENTRY_ID+*@user:targetntmp_userstruct+*@port:thespecifiedswitchportID+*@resume_entry_id:itisbothaninputandanoutput.Asaninput,it+*representstheFDBentryIDtobesearched.IfitisaNULLentryID,+*itindicatesthatthefirstFDBentryforthatportisbeingsearched.+*Asanoutput,itrepresentsthenextFDBentryIDtobesearched.+*@entry:returnedvalue,theresponsedataofthesearchedFDBentry+*+*Return:0onsuccess,otherwiseanegativeerrorcode+*/+intntmp_fdbt_search_port_entry(structntmp_user*user,intport,+u32*resume_entry_id,+structfdbt_entry_data*entry)+{+structntmp_dma_bufdata={+.dev=user->dev,+.size=sizeof(structfdbt_req_qd),+};+structfdbt_resp_query*resp;+structfdbt_req_qd*req;+unionnetc_cbdcbd;+u32len;+interr;++err=ntmp_alloc_data_mem(&data,(void**)&req);+if(err)+returnerr;++/* Request data */+ntmp_fill_crd(&req->crd,user->tbl.fdbt_ver,0,0);+req->ak.search.resume_eid=cpu_to_le32(*resume_entry_id);+req->ak.search.cfge.port_bitmap=cpu_to_le32(BIT(port));+/* Match CFGE_DATA[PORT_BITMAP] field */+req->ak.search.cfge_mc=FDBT_CFGE_MC_PORT_BITMAP;++/* Request header */+len=NTMP_LEN(data.size,sizeof(*resp));+ntmp_fill_request_hdr(&cbd,data.dma,len,NTMP_FDBT_ID,+NTMP_CMD_QUERY,NTMP_AM_SEARCH);++err=netc_xmit_ntmp_cmd(user,&cbd);+if(err){+dev_err(user->dev,+"Failed to search %s entry on port %d, err: %pe\n",+ntmp_table_name(NTMP_FDBT_ID),port,ERR_PTR(err));+gotoend;+}++if(!cbd.resp_hdr.num_matched){+entry->entry_id=NTMP_NULL_ENTRY_ID;+*resume_entry_id=NTMP_NULL_ENTRY_ID;+gotoend;+}++resp=(structfdbt_resp_query*)req;+*resume_entry_id=le32_to_cpu(resp->status);+entry->entry_id=le32_to_cpu(resp->entry_id);+entry->keye=resp->keye;+entry->cfge=resp->cfge;+entry->acte=resp->acte;++end:+ntmp_free_data_mem(&data);++returnerr;+}+EXPORT_SYMBOL_GPL(ntmp_fdbt_search_port_entry);+MODULE_DESCRIPTION("NXP NETC Library");MODULE_LICENSE("Dual BSD/GPL");
@@ -101,4 +103,61 @@ struct rsst_req_update {u8groups[];};+/* Access Key Format of FDB Table */+structfdbt_ak_eid{+__le32entry_id;+__le32resv[7];+};++structfdbt_ak_exact{+structfdbt_keye_datakeye;+__le32resv[5];+};++structfdbt_ak_search{+__le32resume_eid;+structfdbt_keye_datakeye;+structfdbt_cfge_datacfge;+u8acte;+u8keye_mc;+#define FDBT_KEYE_MAC GENMASK(1, 0)+u8cfge_mc;+#define FDBT_CFGE_MC GENMASK(2, 0)+#define FDBT_CFGE_MC_ANY 0+#define FDBT_CFGE_MC_DYNAMIC 1+#define FDBT_CFGE_MC_PORT_BITMAP 2+#define FDBT_CFGE_MC_DYNAMIC_AND_PORT_BITMAP 3+u8acte_mc;+#define FDBT_ACTE_MC BIT(0)+};++unionfdbt_access_key{+structfdbt_ak_eideid;+structfdbt_ak_exactexact;+structfdbt_ak_searchsearch;+};++/* FDB Table Request Data Buffer Format of Update and Add actions */+structfdbt_req_ua{+structntmp_cmn_req_datacrd;+unionfdbt_access_keyak;+structfdbt_cfge_datacfge;+};++/* FDB Table Request Data Buffer Format of Query and Delete actions */+structfdbt_req_qd{+structntmp_cmn_req_datacrd;+unionfdbt_access_keyak;+};++/* FDB Table Response Data Buffer Format of Query action */+structfdbt_resp_query{+__le32status;+__le32entry_id;+structfdbt_keye_datakeye;+structfdbt_cfge_datacfge;+u8acte;+u8resv[3];+};+#endif
The VLAN filter table contains configuration and control information for
each VLAN configured on the switch. Each VLAN entry includes the VLAN
port membership, which FID to use in the FDB lookup, which spanning tree
group to use, the egress frame modification actions to apply to a frame
exiting form this VLAN, and various configuration and control parameters
for this VLAN.
The VLAN filter table can only be managed by the command BD ring using
table management protocol version 2.0. The table supports Add, Delete,
Update and Query operations. And the table supports 3 access methods:
Entry ID, Exact Match Key Element and Search. But currently we only add
the ntmp_vft_add_entry() helper to support the upcoming switch driver to
add an entry to the VLAN filter table. Other interfaces will be added in
the future.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/ethernet/freescale/enetc/ntmp.c | 50 +++++++++++++++++++
.../ethernet/freescale/enetc/ntmp_private.h | 19 +++++++
include/linux/fsl/ntmp.h | 30 +++++++++++
3 files changed, 99 insertions(+)
@@ -160,4 +160,23 @@ struct fdbt_resp_query {u8resv[3];};+/* Access Key Format of VLAN Filter Table */+structvft_ak_exact{+__le16vid;/* bit0~11: VLAN ID, other bits are reserved */+__le16resv;+};++unionvft_access_key{+__le32entry_id;/* entry_id match */+structvft_ak_exactexact;+__le32resume_entry_id;/* search */+};++/* VLAN Filter Table Request Data Buffer Format of Update and Add actions */+structvft_req_ua{+structntmp_cmn_req_datacrd;+unionvft_access_keyak;+structvft_cfge_datacfge;+};+#endif
The buffer pool table contains buffer pool configuration and operational
information. Each entry corresponds to a buffer pool. The Entry ID value
represents the buffer pool ID to access.
The buffer pool table is a static bounded index table, buffer pools are
always present and enabled. It only supports Update and Query operations,
This patch only adds ntmp_bpt_update_entry() helper to support updating
the specified entry of the buffer pool table. Query action to the table
will be added in the future.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/ethernet/freescale/enetc/ntmp.c | 39 +++++++++++++++++++
.../ethernet/freescale/enetc/ntmp_private.h | 6 +++
include/linux/fsl/ntmp.h | 32 +++++++++++++++
3 files changed, 77 insertions(+)
@@ -179,4 +179,10 @@ struct vft_req_ua {structvft_cfge_datacfge;};+/* Buffer Pool Table Request Data Buffer Format of Update action */+structbpt_req_update{+structntmp_req_by_eidrbe;+structbpt_cfge_datacfge;+};+#endif
@@ -142,6 +166,8 @@ int ntmp_fdbt_search_port_entry(struct ntmp_user *user, int port,structfdbt_entry_data*entry);intntmp_vft_add_entry(structntmp_user*user,u16vid,conststructvft_cfge_data*cfge);+intntmp_bpt_update_entry(structntmp_user*user,u32entry_id,+conststructbpt_cfge_data*cfge);#elsestaticinlineintntmp_init_cbdr(structnetc_cbdr*cbdr,structdevice*dev,conststructnetc_cbdr_regs*regs)
The ingress port filter table (IPFT )contains a set of filters each
capable of classifying incoming traffic using a mix of L2, L3, and L4
parsed and arbitrary field data. As a result of a filter match, several
actions can be specified such as on whether to deny or allow a frame,
overriding internal QoS attributes associated with the frame and setting
parameters for the subsequent frame processing functions, such as stream
identification, policing, ingress mirroring. Each entry corresponds to a
filter. The ingress port filter entries are added using a precedence
value. If a frame matches multiple entries, the entry with the higher
precedence is used. Currently, this patch only adds "Add" and "Delete"
operations to the ingress port filter table. These two interfaces will
be used by both ENETC driver and NETC switch driver.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/ethernet/freescale/enetc/ntmp.c | 76 +++++++++++++
.../ethernet/freescale/enetc/ntmp_private.h | 36 ++++++
include/linux/fsl/ntmp.h | 104 ++++++++++++++++++
3 files changed, 216 insertions(+)
@@ -103,6 +103,42 @@ struct rsst_req_update {u8groups[];};+/* Ingress Port Filter Table Response Data Buffer Format of Query action */+structipft_resp_query{+__le32status;+__le32entry_id;+structipft_keye_datakeye;+__le64match_count;/* STSE_DATA */+structipft_cfge_datacfge;+}__packed;++structipft_ak_eid{+__le32entry_id;+__le32resv[52];+};++unionipft_access_key{+structipft_ak_eideid;+structipft_keye_datakeye;+};++/* Ingress Port Filter Table Request Data Buffer Format of Update and+*Addactions+*/+structipft_req_ua{+structntmp_cmn_req_datacrd;+unionipft_access_keyak;+structipft_cfge_datacfge;+};++/* Ingress Port Filter Table Request Data Buffer Format of Query and+*Deleteactions+*/+structipft_req_qd{+structntmp_req_by_eidrbe;+__le32resv[52];+};+/* Access Key Format of FDB Table */structfdbt_ak_eid{__le32entry_id;
All the tables of NETC switch are managed through the command BD ring,
but unlike ENETC, the switch has two command BD rings, if the current
ring is busy, the switch driver can switch to another ring to manage
the table. Currently, the NTMP driver does not support multiple rings.
Therefore, netc_select_cbdr() is added to select a appropriate ring to
execute the command for the switch.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/ethernet/freescale/enetc/ntmp.c | 27 ++++++++++++++++++---
1 file changed, 23 insertions(+), 4 deletions(-)
@@ -117,6 +117,25 @@ static void ntmp_clean_cbdr(struct netc_cbdr *cbdr)cbdr->next_to_clean=i;}+staticstructnetc_cbdr*netc_select_cbdr(structntmp_user*user)+{+intcpu,i;++for(i=0;i<user->cbdr_num;i++){+if(spin_is_locked(&user->ring[i].ring_lock))+continue;++return&user->ring[i];+}++/* If all the command BDRs are busy now, we select+*oneofthem,butneedtowaitforawhiletouse.+*/+cpu=raw_smp_processor_id();++return&user->ring[cpu%user->cbdr_num];+}+staticintnetc_xmit_ntmp_cmd(structntmp_user*user,unionnetc_cbd*cbd){unionnetc_cbd*cur_cbd;
@@ -125,10 +144,10 @@ static int netc_xmit_ntmp_cmd(struct ntmp_user *user, union netc_cbd *cbd)u16status;u32val;-/* Currently only i.MX95 ENETC is supported, and it only has one-*commandBDring-*/-cbdr=&user->ring[0];+if(user->cbdr_num==1)+cbdr=&user->ring[0];+else+cbdr=netc_select_cbdr(user);spin_lock_bh(&cbdr->ring_lock);
The NXP NETC switch tag is a proprietary header added to frames after the
source MAC address. The switch tag has 3 types, and each type has 1 ~ 4
subtypes, the details are as follows.
Forward NXP switch tag (Type=0): Represents forwarded frames.
- SubType = 0 - Normal frame processing.
To_Port NXP switch tag (Type=1): Represents frames that are to be sent
to a specific switch port.
- SubType = 0. No request to perform timestamping.
- SubType = 1. Request to perform one-step timestamping.
- SubType = 2. Request to perform two-step timestamping.
- SubType = 3. Request to perform both one-step timestamping and
two-step timestamping.
To_Host NXP switch tag (Type=2): Represents frames redirected or copied
to the switch management port.
- SubType = 0. Received frames redirected or copied to the switch
management port.
- SubType = 1. Received frames redirected or copied to the switch
management port with captured timestamp at the switch port where
the frame was received.
- SubType = 2. Transmit timestamp response (two-step timestamping).
In addition, the length of different type switch tag is different, the
minimum length is 6 bytes, the maximum length is 14 bytes. Currently,
Forward tag, SubType 0 of To_Port tag and Subtype 0 of To_Host tag are
supported. More tags will be supported in the future.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
include/linux/dsa/tag_netc.h | 14 +++
include/net/dsa.h | 2 +
include/uapi/linux/if_ether.h | 1 +
net/dsa/Kconfig | 10 ++
net/dsa/Makefile | 1 +
net/dsa/tag_netc.c | 180 ++++++++++++++++++++++++++++++++++
6 files changed, 208 insertions(+)
create mode 100644 include/linux/dsa/tag_netc.h
create mode 100644 net/dsa/tag_netc.c
@@ -123,6 +123,7 @@#define ETH_P_DSA_A5PSW 0xE001 /* A5PSW Tag Value [ NOT AN OFFICIALLY REGISTERED ID ] */#define ETH_P_IFE 0xED3E /* ForCES inter-FE LFB type */#define ETH_P_AF_IUCV 0xFBFB /* IBM af_iucv [ NOT AN OFFICIALLY REGISTERED ID ] */+#define ETH_P_NXP_NETC 0xFD3A /* NXP NETC DSA [ NOT AN OFFICIALLY REGISTERED ID ] */#define ETH_P_802_3_MIN 0x0600 /* If the value in the ethernet type is more than this value*thentheframeisEthernetII.Elseitis802.3*/
@@ -125,6 +125,16 @@ config NET_DSA_TAG_KSZSayYifyouwanttoenablesupportfortaggingframesfortheMicrochip8795/937x/9477/9893familiesofswitches.+configNET_DSA_TAG_NETC+tristate"Tag driver for NXP NETC switches"+help+SayYorMifyouwanttoenablesupportfortheNXPSwitchTag(NST),+asimplementedbyNXPNETCswitcheshavingversion4.3orlater.The+switchtagisaproprietaryheaderaddedtoframesafterthesource+MACaddress,ithas3typesandeachtypehasdifferentsubtypes,so+itslengthdepends onthetypeandsubtypeofthetag,themaximum+lengthis14bytes.+configNET_DSA_TAG_OCELOTtristate"Tag driver for Ocelot family of switches, using NPI port"selectPACKING
@@ -0,0 +1,180 @@+// SPDX-License-Identifier: GPL-2.0+/*+*Copyright2025-2026NXP+*/++#include<linux/dsa/tag_netc.h>++#include"tag.h"++#define NETC_NAME "nxp_netc"++/* Forward NXP switch tag */+#define NETC_TAG_FORWARD 0++/* To_Port NXP switch tag */+#define NETC_TAG_TO_PORT 1+/* SubType0: No request to perform timestamping */+#define NETC_TAG_TP_SUBTYPE0 0++/* To_Host NXP switch tag */+#define NETC_TAG_TO_HOST 2+/* SubType0: frames redirected or copied to CPU port */+#define NETC_TAG_TH_SUBTYPE0 0+/* SubType1: frames redirected or copied to CPU port with timestamp */+#define NETC_TAG_TH_SUBTYPE1 1+/* SubType2: Transmit timestamp response (two-step timestamping) */+#define NETC_TAG_TH_SUBTYPE2 2++/* NETC switch tag lengths */+#define NETC_TAG_FORWARD_LEN 6+#define NETC_TAG_TP_SUBTYPE0_LEN 6+#define NETC_TAG_TH_SUBTYPE0_LEN 6+#define NETC_TAG_TH_SUBTYPE1_LEN 14+#define NETC_TAG_TH_SUBTYPE2_LEN 14+#define NETC_TAG_CMN_LEN 5++#define NETC_TAG_SUBTYPE GENMASK(3, 0)+#define NETC_TAG_TYPE GENMASK(7, 4)+#define NETC_TAG_QV BIT(0)+#define NETC_TAG_IPV GENMASK(4, 2)+#define NETC_TAG_SWITCH GENMASK(2, 0)+#define NETC_TAG_PORT GENMASK(7, 3)++structnetc_tag_cmn{+__be16tpid;+u8type;+u8qos;+u8switch_port;+}__packed;++staticvoidnetc_fill_common_tag(structnetc_tag_cmn*tag,u8type,+u8subtype,u8sw_id,u8port,u8ipv)+{+tag->tpid=htons(ETH_P_NXP_NETC);+tag->type=FIELD_PREP(NETC_TAG_TYPE,type)|+FIELD_PREP(NETC_TAG_SUBTYPE,subtype);+tag->qos=NETC_TAG_QV|FIELD_PREP(NETC_TAG_IPV,ipv);+tag->switch_port=FIELD_PREP(NETC_TAG_SWITCH,sw_id)|+FIELD_PREP(NETC_TAG_PORT,port);+}++staticvoid*netc_fill_common_tp_tag(structsk_buff*skb,+structnet_device*ndev,+u8subtype,inttag_len)+{+structdsa_port*dp=dsa_user_to_port(ndev);+u16queue=skb_get_queue_mapping(skb);+u8ipv=netdev_txq_to_tc(ndev,queue);+void*tag;++skb_push(skb,tag_len);+dsa_alloc_etype_header(skb,tag_len);++tag=dsa_etype_header_pos_tx(skb);+memset(tag+NETC_TAG_CMN_LEN,0,tag_len-NETC_TAG_CMN_LEN);+netc_fill_common_tag(tag,NETC_TAG_TO_PORT,subtype,+dp->ds->index,dp->index,ipv);++returntag;+}++staticvoidnetc_fill_tp_tag_subtype0(structsk_buff*skb,+structnet_device*ndev)+{+netc_fill_common_tp_tag(skb,ndev,NETC_TAG_TP_SUBTYPE0,+NETC_TAG_TP_SUBTYPE0_LEN);+}++/* Currently only support To_Port tag, subtype 0 */+staticstructsk_buff*netc_xmit(structsk_buff*skb,+structnet_device*ndev)+{+netc_fill_tp_tag_subtype0(skb,ndev);++returnskb;+}++staticintnetc_get_rx_tag_len(intrx_type)+{+inttype=FIELD_GET(NETC_TAG_TYPE,rx_type);++if(type==NETC_TAG_TO_HOST){+u8subtype=rx_type&NETC_TAG_SUBTYPE;++if(subtype==NETC_TAG_TH_SUBTYPE1)+returnNETC_TAG_TH_SUBTYPE1_LEN;+elseif(subtype==NETC_TAG_TH_SUBTYPE2)+returnNETC_TAG_TH_SUBTYPE2_LEN;+else+returnNETC_TAG_TH_SUBTYPE0_LEN;+}++returnNETC_TAG_FORWARD_LEN;+}++staticstructsk_buff*netc_rcv(structsk_buff*skb,+structnet_device*ndev)+{+structnetc_tag_cmn*tag_cmn=dsa_etype_header_pos_rx(skb);+inttag_len=netc_get_rx_tag_len(tag_cmn->type);+intsw_id,port;++if(ntohs(tag_cmn->tpid)!=ETH_P_NXP_NETC){+dev_warn_ratelimited(&ndev->dev,"Unknown TPID 0x%04x\n",+ntohs(tag_cmn->tpid));++returnNULL;+}++if(tag_cmn->qos&NETC_TAG_QV)+skb->priority=FIELD_GET(NETC_TAG_IPV,tag_cmn->qos);++sw_id=NETC_TAG_SWITCH&tag_cmn->switch_port;+/* ENETC VEPA switch ID (0) is not supported yet */+if(!sw_id){+dev_warn_ratelimited(&ndev->dev,+"VEPA switch ID is not supported yet\n");++returnNULL;+}++port=FIELD_GET(NETC_TAG_PORT,tag_cmn->switch_port);+skb->dev=dsa_conduit_find_user(ndev,sw_id,port);+if(!skb->dev)+returnNULL;++if(tag_cmn->type==NETC_TAG_FORWARD)+dsa_default_offload_fwd_mark(skb);++/* Remove Switch tag from the frame */+skb_pull_rcsum(skb,tag_len);+dsa_strip_etype_header(skb,tag_len);++returnskb;+}++staticvoidnetc_flow_dissect(conststructsk_buff*skb,__be16*proto,+int*offset)+{+structnetc_tag_cmn*tag_cmn=(structnetc_tag_cmn*)(skb->data-2);+inttag_len=netc_get_rx_tag_len(tag_cmn->type);++*offset=tag_len;+*proto=((__be16*)skb->data)[(tag_len/2)-1];+}++staticconststructdsa_device_opsnetc_netdev_ops={+.name=NETC_NAME,+.proto=DSA_TAG_PROTO_NETC,+.xmit=netc_xmit,+.rcv=netc_rcv,+.needed_headroom=NETC_TAG_MAX_LEN,+.flow_dissect=netc_flow_dissect,+};++MODULE_DESCRIPTION("DSA tag driver for NXP NETC switch family");+MODULE_LICENSE("GPL");++MODULE_ALIAS_DSA_TAG_DRIVER(DSA_TAG_PROTO_NETC,NETC_NAME);+module_dsa_tag_driver(netc_netdev_ops);
For i.MX94 series, the NETC IP provides full 802.1Q Ethernet switch
functionality, advanced QoS with 8 traffic classes, and a full range of
TSN standards capabilities. The switch has 3 user ports and 1 CPU port,
the CPU port is connected to an internal ENETC. Since the switch and the
internal ENETC are fully integrated within the NETC IP, no back-to-back
MAC connection is required. Instead, a light-weight "pseudo MAC" is used
between the switch and the ENETC. This translates to lower power (less
logic and memory) and lower delay (as there is no serialization delay
across this link).
This patch introduces the initial NETC switch driver. At this stage,
only basic probe and remove functionality is supported. More features
will be supported in the subsequent patches.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
MAINTAINERS | 11 +
drivers/net/dsa/Kconfig | 3 +
drivers/net/dsa/Makefile | 1 +
drivers/net/dsa/netc/Kconfig | 14 +
drivers/net/dsa/netc/Makefile | 3 +
drivers/net/dsa/netc/netc_main.c | 672 ++++++++++++++++++++++++++
drivers/net/dsa/netc/netc_platform.c | 49 ++
drivers/net/dsa/netc/netc_switch.h | 92 ++++
drivers/net/dsa/netc/netc_switch_hw.h | 155 ++++++
9 files changed, 1000 insertions(+)
create mode 100644 drivers/net/dsa/netc/Kconfig
create mode 100644 drivers/net/dsa/netc/Makefile
create mode 100644 drivers/net/dsa/netc/netc_main.c
create mode 100644 drivers/net/dsa/netc/netc_platform.c
create mode 100644 drivers/net/dsa/netc/netc_switch.h
create mode 100644 drivers/net/dsa/netc/netc_switch_hw.h
@@ -0,0 +1,672 @@+// SPDX-License-Identifier: (GPL-2.0+ OR BSD-3-Clause)+/*+*NXPNETCswitchdriver+*Copyright2025-2026NXP+*/++#include<linux/clk.h>+#include<linux/etherdevice.h>+#include<linux/fsl/enetc_mdio.h>+#include<linux/if_vlan.h>+#include<linux/of_mdio.h>++#include"netc_switch.h"++staticenumdsa_tag_protocol+netc_get_tag_protocol(structdsa_switch*ds,intport,+enumdsa_tag_protocolmprot)+{+returnDSA_TAG_PROTO_NETC;+}++staticvoidnetc_port_rmw(structnetc_port*np,u32reg,+u32mask,u32val)+{+u32old,new;++WARN_ON((mask|val)!=mask);++old=netc_port_rd(np,reg);+new=(old&~mask)|val;+if(new==old)+return;++netc_port_wr(np,reg,new);+}++staticvoidnetc_mac_port_wr(structnetc_port*np,u32reg,u32val)+{+if(is_netc_pseudo_port(np))+return;++netc_port_wr(np,reg,val);+if(np->caps.pmac)+netc_port_wr(np,reg+NETC_PMAC_OFFSET,val);+}++staticvoidnetc_mac_port_rmw(structnetc_port*np,u32reg,+u32mask,u32val)+{+u32old,new;++if(is_netc_pseudo_port(np))+return;++WARN_ON((mask|val)!=mask);++old=netc_port_rd(np,reg);+new=(old&~mask)|val;+if(new==old)+return;++netc_port_wr(np,reg,new);+if(np->caps.pmac)+netc_port_wr(np,reg+NETC_PMAC_OFFSET,new);+}++staticvoidnetc_port_get_capability(structnetc_port*np)+{+u32val;++val=netc_port_rd(np,NETC_PMCAPR);+if(val&PMCAPR_HD)+np->caps.half_duplex=true;++if(FIELD_GET(PMCAPR_FP,val)==FP_SUPPORT)+np->caps.pmac=true;++val=netc_port_rd(np,NETC_PCAPR);+if(val&PCAPR_LINK_TYPE)+np->caps.pseudo_link=true;+}++staticintnetc_port_get_info_from_dt(structnetc_port*np,+structdevice_node*node,+structdevice*dev)+{+if(of_find_property(node,"clock-names",NULL)){+np->ref_clk=devm_get_clk_from_child(dev,node,"ref");+if(IS_ERR(np->ref_clk)){+dev_err(dev,"Port %d cannot get reference clock\n",+np->dp->index);+returnPTR_ERR(np->ref_clk);+}+}++return0;+}++staticintnetc_port_create_emdio_bus(structnetc_port*np,+structdevice_node*node)+{+structnetc_switch*priv=np->switch_priv;+structenetc_mdio_priv*mdio_priv;+structdevice*dev=priv->dev;+structenetc_hw*hw;+structmii_bus*bus;+interr;++hw=enetc_hw_alloc(dev,np->iobase);+if(IS_ERR(hw))+returndev_err_probe(dev,PTR_ERR(hw),+"Failed to allocate enetc_hw\n");++bus=devm_mdiobus_alloc_size(dev,sizeof(*mdio_priv));+if(!bus)+return-ENOMEM;++bus->name="NXP NETC switch external MDIO Bus";+bus->read=enetc_mdio_read_c22;+bus->write=enetc_mdio_write_c22;+bus->read_c45=enetc_mdio_read_c45;+bus->write_c45=enetc_mdio_write_c45;+bus->parent=dev;+mdio_priv=bus->priv;+mdio_priv->hw=hw;+mdio_priv->mdio_base=NETC_EMDIO_BASE;+snprintf(bus->id,MII_BUS_ID_SIZE,"%s-p%d-emdio",+dev_name(dev),np->dp->index);++err=devm_of_mdiobus_register(dev,bus,node);+if(err)+returndev_err_probe(dev,err,+"Cannot register EMDIO bus\n");++np->emdio=bus;++return0;+}++staticintnetc_port_create_mdio_bus(structnetc_port*np,+structdevice_node*node)+{+structdevice_node*mdio_node;+interr;++mdio_node=of_get_child_by_name(node,"mdio");+if(mdio_node){+err=netc_port_create_emdio_bus(np,mdio_node);+of_node_put(mdio_node);+if(err)+returnerr;+}++return0;+}++staticintnetc_init_switch_id(structnetc_switch*priv)+{+structnetc_switch_regs*regs=&priv->regs;+structdsa_switch*ds=priv->ds;++/* The value of 0 is reserved for the VEPA switch and cannot+*beused.+*/+if(ds->index>SWCR_SWID||!ds->index){+dev_err(priv->dev,"Switch index %d out of range\n",+ds->index);+return-ERANGE;+}++netc_base_wr(regs,NETC_SWCR,ds->index);++return0;+}++staticintnetc_init_all_ports(structnetc_switch*priv)+{+structdevice*dev=priv->dev;+structnetc_port*np;+structdsa_port*dp;+interr;++priv->ports=devm_kcalloc(dev,priv->info->num_ports,+sizeof(structnetc_port*),+GFP_KERNEL);+if(!priv->ports)+return-ENOMEM;++/* Some DSA interfaces may set the port even it is disabled, such+*as.port_disable(),.port_stp_state_set()andsoon.Toavoid+*crashcausedbyaccessingNULLportpointer,eachportis+*allocateditsownmemory.Otherwise,weneedtocheckwhether+*theportpointerisNULLintheseinterfaces.Thelatteris+*difficultforustocover.+*/+for(inti=0;i<priv->info->num_ports;i++){+np=devm_kzalloc(dev,sizeof(*np),GFP_KERNEL);+if(!np)+return-ENOMEM;++np->switch_priv=priv;+np->iobase=priv->regs.port+PORT_IOBASE(i);+netc_port_get_capability(np);+priv->ports[i]=np;+}++dsa_switch_for_each_available_port(dp,priv->ds){+np=priv->ports[dp->index];+np->dp=dp;+err=netc_port_get_info_from_dt(np,dp->dn,dev);+if(err)+returnerr;++if(dsa_port_is_user(dp)){+err=netc_port_create_mdio_bus(np,dp->dn);+if(err){+dev_err(dev,"Failed to create MDIO bus\n");+returnerr;+}+}+}++return0;+}++staticvoidnetc_init_ntmp_tbl_versions(structnetc_switch*priv)+{+structntmp_user*ntmp=&priv->ntmp;++/* All tables default to version 0 */+memset(&ntmp->tbl,0,sizeof(ntmp->tbl));+}++staticintnetc_init_all_cbdrs(structnetc_switch*priv)+{+structnetc_switch_regs*regs=&priv->regs;+structntmp_user*ntmp=&priv->ntmp;+inti,err;++ntmp->cbdr_num=NETC_CBDR_NUM;+ntmp->dev=priv->dev;+ntmp->ring=devm_kcalloc(ntmp->dev,ntmp->cbdr_num,+sizeof(structnetc_cbdr),+GFP_KERNEL);+if(!ntmp->ring)+return-ENOMEM;++for(i=0;i<ntmp->cbdr_num;i++){+structnetc_cbdr*cbdr=&ntmp->ring[i];+structnetc_cbdr_regscbdr_regs;++cbdr_regs.pir=regs->base+NETC_CBDRPIR(i);+cbdr_regs.cir=regs->base+NETC_CBDRCIR(i);+cbdr_regs.mr=regs->base+NETC_CBDRMR(i);+cbdr_regs.bar0=regs->base+NETC_CBDRBAR0(i);+cbdr_regs.bar1=regs->base+NETC_CBDRBAR1(i);+cbdr_regs.lenr=regs->base+NETC_CBDRLENR(i);++err=ntmp_init_cbdr(cbdr,ntmp->dev,&cbdr_regs);+if(err)+gotofree_cbdrs;+}++return0;++free_cbdrs:+for(i--;i>=0;i--)+ntmp_free_cbdr(&ntmp->ring[i]);++returnerr;+}++staticvoidnetc_remove_all_cbdrs(structnetc_switch*priv)+{+structntmp_user*ntmp=&priv->ntmp;++for(inti=0;i<NETC_CBDR_NUM;i++)+ntmp_free_cbdr(&ntmp->ring[i]);+}++staticintnetc_init_ntmp_user(structnetc_switch*priv)+{+netc_init_ntmp_tbl_versions(priv);++returnnetc_init_all_cbdrs(priv);+}++staticvoidnetc_free_ntmp_user(structnetc_switch*priv)+{+netc_remove_all_cbdrs(priv);+}++staticvoidnetc_switch_dos_default_config(structnetc_switch*priv)+{+structnetc_switch_regs*regs=&priv->regs;+u32val;++val=DOSL2CR_SAMEADDR|DOSL2CR_MSAMCC;+netc_base_wr(regs,NETC_DOSL2CR,val);++val=DOSL3CR_SAMEADDR|DOSL3CR_IPSAMCC;+netc_base_wr(regs,NETC_DOSL3CR,val);+}++staticvoidnetc_switch_vfht_default_config(structnetc_switch*priv)+{+structnetc_switch_regs*regs=&priv->regs;+u32val;++val=netc_base_rd(regs,NETC_VFHTDECR2);++/* if no match is found in the VLAN Filter table, then VFHTDECR2[MLO]+*willtakeeffect.VFHTDECR2[MLO]issetto"Software MAC learning+*secure" by default. Notice BPCR[MLO] will override VFHTDECR2[MLO]+*ifitsvalueisnotzero.+*/+val=u32_replace_bits(val,MLO_SW_SEC,VFHTDECR2_MLO);+val=u32_replace_bits(val,MFO_NO_MATCH_DISCARD,VFHTDECR2_MFO);+netc_base_wr(regs,NETC_VFHTDECR2,val);+}++staticvoidnetc_port_set_max_frame_size(structnetc_port*np,+u32max_frame_size)+{+netc_mac_port_wr(np,NETC_PM_MAXFRM(0),+PM_MAXFRAM&max_frame_size);+}++staticvoidnetc_switch_fixed_config(structnetc_switch*priv)+{+netc_switch_dos_default_config(priv);+netc_switch_vfht_default_config(priv);+}++staticvoidnetc_port_set_tc_max_sdu(structnetc_port*np,+inttc,u32max_sdu)+{+u32val=max_sdu&PTCTMSDUR_MAXSDU;++val|=FIELD_PREP(PTCTMSDUR_SDU_TYPE,SDU_TYPE_MPDU);+netc_port_wr(np,NETC_PTCTMSDUR(tc),val);+}++staticvoidnetc_port_set_all_tc_msdu(structnetc_port*np)+{+for(inttc=0;tc<NETC_TC_NUM;tc++)+netc_port_set_tc_max_sdu(np,tc,NETC_MAX_FRAME_LEN);+}++staticvoidnetc_port_set_mlo(structnetc_port*np,enumnetc_mlomlo)+{+netc_port_rmw(np,NETC_BPCR,BPCR_MLO,FIELD_PREP(BPCR_MLO,mlo));+}++staticvoidnetc_port_fixed_config(structnetc_port*np)+{+/* Default IPV and DR setting */+netc_port_rmw(np,NETC_PQOSMR,PQOSMR_VS|PQOSMR_VE,+PQOSMR_VS|PQOSMR_VE);++/* Enable L2 and L3 DOS */+netc_port_rmw(np,NETC_PCR,PCR_L2DOSE|PCR_L3DOSE,+PCR_L2DOSE|PCR_L3DOSE);+}++staticvoidnetc_port_default_config(structnetc_port*np)+{+netc_port_fixed_config(np);++/* Default VLAN unaware */+netc_port_rmw(np,NETC_BPDVR,BPDVR_RXVAM,BPDVR_RXVAM);++if(dsa_port_is_cpu(np->dp))+/* For CPU port, source port pruning is disabled and+*hardwareMAClearningisenabledbydefault.+*/+netc_port_rmw(np,NETC_BPCR,BPCR_SRCPRND|BPCR_MLO,+BPCR_SRCPRND|FIELD_PREP(BPCR_MLO,MLO_HW));+else+netc_port_set_mlo(np,MLO_DISABLE);++netc_port_set_max_frame_size(np,NETC_MAX_FRAME_LEN);+netc_port_set_all_tc_msdu(np);+netc_mac_port_rmw(np,NETC_PM_CMD_CFG(0),PM_CMD_CFG_TX_EN,+PM_CMD_CFG_TX_EN);+netc_port_rmw(np,NETC_POR,PCR_TXDIS,0);+}++staticintnetc_setup(structdsa_switch*ds)+{+structnetc_switch*priv=ds->priv;+structdsa_port*dp;+interr;++err=netc_init_switch_id(priv);+if(err)+returnerr;++err=netc_init_all_ports(priv);+if(err)+returnerr;++err=netc_init_ntmp_user(priv);+if(err)+returnerr;++netc_switch_fixed_config(priv);++/* default setting for ports */+dsa_switch_for_each_available_port(dp,ds)+netc_port_default_config(priv->ports[dp->index]);++return0;+}++staticvoidnetc_teardown(structdsa_switch*ds)+{+structnetc_switch*priv=ds->priv;++netc_free_ntmp_user(priv);+}++staticstructdevice_node*netc_get_switch_ports(structdevice_node*node)+{+structdevice_node*ports;++ports=of_get_child_by_name(node,"ports");+if(!ports)+ports=of_get_child_by_name(node,"ethernet-ports");++returnports;+}++staticboolnetc_port_is_emdio_consumer(structdevice_node*node)+{+structdevice_node*mdio_node;++/* If the port node has phy-handle property and it does+*notcontainamdiochildnode,thentheportisthe+*EMDIOconsumer.+*/+mdio_node=of_get_child_by_name(node,"mdio");+if(!mdio_node)+returntrue;++of_node_put(mdio_node);++returnfalse;+}++/* Currently, phylink_of_phy_connect() is called by dsa_user_create(),+*soiftheswitchusestheexternalMDIOcontroller(liketheEMDIO+*function)tomanagetheexternalPHYs.TheMDIObusmaynotbe+*createdwhenphylink_of_phy_connect()iscalled,soitwillreturn+*anerrorandcausetheswitchdrivertofailtoprobe.+*ThisworkaroundcanberemovedwhenDSAphylink_of_phy_connect()+*callsaremovedfromprobe()tondo_open().+*/+staticintnetc_switch_check_emdio_is_ready(structdevice*dev)+{+structdevice_node*ports,*phy_node;+structphy_device*phydev;+interr=0;++ports=netc_get_switch_ports(dev->of_node);+if(!ports)+return0;++for_each_available_child_of_node_scoped(ports,child){+/* If the node does not have phy-handle property, then+*theportdoesnotconnecttoaPHY,sotheportis+*nottheEMDIOconsumer.+*/+phy_node=of_parse_phandle(child,"phy-handle",0);+if(!phy_node)+continue;++if(!netc_port_is_emdio_consumer(child)){+of_node_put(phy_node);+continue;+}++phydev=of_phy_find_device(phy_node);+of_node_put(phy_node);+if(!phydev){+err=-EPROBE_DEFER;+gotoout;+}++put_device(&phydev->mdio.dev);+}++out:+of_node_put(ports);++returnerr;+}++staticintnetc_switch_pci_init(structpci_dev*pdev)+{+structdevice*dev=&pdev->dev;+structnetc_switch_regs*regs;+structnetc_switch*priv;+interr;++pcie_flr(pdev);+err=pci_enable_device_mem(pdev);+if(err)+returndev_err_probe(dev,err,"Failed to enable device\n");++/* The command BD rings and NTMP tables need DMA. No need to check+*thereturnvalue,becauseitneverreturnsfailwhenthemaskis+*DMA_BIT_MASK(64),seedma-api-howto.rst.+*/+dma_set_mask_and_coherent(dev,DMA_BIT_MASK(64));+err=pci_request_mem_regions(pdev,KBUILD_MODNAME);+if(err){+dev_err(dev,"Failed to request memory regions, err: %pe\n",+ERR_PTR(err));+gotodisable_pci_device;+}++pci_set_master(pdev);+priv=devm_kzalloc(dev,sizeof(*priv),GFP_KERNEL);+if(!priv){+err=-ENOMEM;+gotorelease_mem_regions;+}++priv->pdev=pdev;+priv->dev=dev;++regs=&priv->regs;+regs->base=pci_ioremap_bar(pdev,NETC_REGS_BAR);+if(!regs->base){+err=-ENXIO;+dev_err(dev,"pci_ioremap_bar() failed\n");+gotorelease_mem_regions;+}++regs->port=regs->base+NETC_REGS_PORT_BASE;+regs->global=regs->base+NETC_REGS_GLOBAL_BASE;+pci_set_drvdata(pdev,priv);++return0;++release_mem_regions:+pci_release_mem_regions(pdev);+disable_pci_device:+pci_disable_device(pdev);++returnerr;+}++staticvoidnetc_switch_pci_destroy(structpci_dev*pdev)+{+structnetc_switch*priv=pci_get_drvdata(pdev);++iounmap(priv->regs.base);+pci_release_mem_regions(pdev);+pci_disable_device(pdev);+}++staticvoidnetc_switch_get_ip_revision(structnetc_switch*priv)+{+structnetc_switch_regs*regs=&priv->regs;+u32val=netc_glb_rd(regs,NETC_IPBRR0);++priv->revision=val&IPBRR0_IP_REV;+}++staticconststructdsa_switch_opsnetc_switch_ops={+.get_tag_protocol=netc_get_tag_protocol,+.setup=netc_setup,+.teardown=netc_teardown,+};++staticintnetc_switch_probe(structpci_dev*pdev,+conststructpci_device_id*id)+{+structdevice_node*node=dev_of_node(&pdev->dev);+structdevice*dev=&pdev->dev;+structnetc_switch*priv;+structdsa_switch*ds;+interr;++if(!node)+returndev_err_probe(dev,-ENODEV,+"No DT bindings, skipping\n");++err=netc_switch_check_emdio_is_ready(dev);+if(err)+returnerr;++err=netc_switch_pci_init(pdev);+if(err)+returnerr;++priv=pci_get_drvdata(pdev);+netc_switch_get_ip_revision(priv);++err=netc_switch_platform_probe(priv);+if(err)+gotodestroy_netc_switch;++ds=devm_kzalloc(dev,sizeof(*ds),GFP_KERNEL);+if(!ds){+err=-ENOMEM;+gotodestroy_netc_switch;+}++ds->dev=dev;+ds->num_ports=priv->info->num_ports;+ds->num_tx_queues=NETC_TC_NUM;+ds->ops=&netc_switch_ops;+ds->priv=priv;++priv->ds=ds;++err=dsa_register_switch(ds);+if(err){+dev_err_probe(dev,err,"Failed to register DSA switch\n");+gotodestroy_netc_switch;+}++return0;++destroy_netc_switch:+netc_switch_pci_destroy(pdev);++returnerr;+}++staticvoidnetc_switch_remove(structpci_dev*pdev)+{+structnetc_switch*priv=pci_get_drvdata(pdev);++if(!priv)+return;++dsa_unregister_switch(priv->ds);+netc_switch_pci_destroy(pdev);+}++staticvoidnetc_switch_shutdown(structpci_dev*pdev)+{+structnetc_switch*priv=pci_get_drvdata(pdev);++if(!priv)+return;++dsa_switch_shutdown(priv->ds);+pci_set_drvdata(pdev,NULL);+}++staticconststructpci_device_idnetc_switch_ids[]={+{PCI_DEVICE(NETC_SWITCH_VENDOR_ID,NETC_SWITCH_DEVICE_ID)},+{}+};+MODULE_DEVICE_TABLE(pci,netc_switch_ids);++staticstructpci_drivernetc_switch_driver={+.name=KBUILD_MODNAME,+.id_table=netc_switch_ids,+.probe=netc_switch_probe,+.remove=netc_switch_remove,+.shutdown=netc_switch_shutdown,+};+module_pci_driver(netc_switch_driver);++MODULE_DESCRIPTION("NXP NETC Switch driver");+MODULE_LICENSE("Dual BSD/GPL");
@@ -0,0 +1,49 @@+// SPDX-License-Identifier: (GPL-2.0+ OR BSD-3-Clause)+/*+*NXPNETCswitchdriver+*Copyright2025-2026NXP+*/++#include"netc_switch.h"++structnetc_switch_platform{+u16revision;+conststructnetc_switch_info*info;+};++staticconststructnetc_switch_infoimx94_info={+.num_ports=4,+};++staticconststructnetc_switch_platformnetc_platforms[]={+{.revision=NETC_SWITCH_REV_4_3,.info=&imx94_info,},+{}+};++staticconststructnetc_switch_info*+netc_switch_get_info(structnetc_switch*priv)+{+inti;++/* Matching based on IP revision */+for(i=0;i<ARRAY_SIZE(netc_platforms);i++){+if(priv->revision==netc_platforms[i].revision)+returnnetc_platforms[i].info;+}++returnNULL;+}++intnetc_switch_platform_probe(structnetc_switch*priv)+{+conststructnetc_switch_info*info=netc_switch_get_info(priv);++if(!info){+dev_err(priv->dev,"Cannot find switch platform info\n");+return-EINVAL;+}++priv->info=info;++return0;+}
Different versions of NETC switches have different numbers of ports and
MAC capabilities, so add .phylink_get_caps() to struct netc_switch_info,
so that each version of the NETC switch can implement its own callback
to obtain MAC capabilities. In addition, related interfaces of struct
phylink_mac_ops are added, such as .mac_config(), .mac_link_up(), and
.mac_link_down().
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/dsa/netc/netc_main.c | 212 ++++++++++++++++++++++++++
drivers/net/dsa/netc/netc_platform.c | 40 +++++
drivers/net/dsa/netc/netc_switch.h | 4 +
drivers/net/dsa/netc/netc_switch_hw.h | 25 +++
4 files changed, 281 insertions(+)
@@ -569,10 +569,221 @@ static void netc_switch_get_ip_revision(struct netc_switch *priv)priv->revision=val&IPBRR0_IP_REV;}+staticvoidnetc_phylink_get_caps(structdsa_switch*ds,intport,+structphylink_config*config)+{+structnetc_switch*priv=ds->priv;++priv->info->phylink_get_caps(port,config);+}++staticvoidnetc_port_set_mac_mode(structnetc_port*np,+unsignedintmode,+phy_interface_tphy_mode)+{+u32mask=PM_IF_MODE_IFMODE|PM_IF_MODE_REVMII|PM_IF_MODE_ENA;+u32val=0;++switch(phy_mode){+casePHY_INTERFACE_MODE_RGMII:+casePHY_INTERFACE_MODE_RGMII_ID:+casePHY_INTERFACE_MODE_RGMII_RXID:+casePHY_INTERFACE_MODE_RGMII_TXID:+val|=IFMODE_RGMII;+/* Enable auto-negotiation for the MAC if its+*RGMIIinterfacesupportsIn-Bandstatus.+*/+if(phylink_autoneg_inband(mode))+val|=PM_IF_MODE_ENA;+break;+casePHY_INTERFACE_MODE_RMII:+val|=IFMODE_RMII;+break;+casePHY_INTERFACE_MODE_REVMII:+val|=PM_IF_MODE_REVMII;+fallthrough;+casePHY_INTERFACE_MODE_MII:+val|=IFMODE_MII;+break;+casePHY_INTERFACE_MODE_SGMII:+casePHY_INTERFACE_MODE_2500BASEX:+val|=IFMODE_SGMII;+break;+default:+break;+}++netc_mac_port_rmw(np,NETC_PM_IF_MODE(0),mask,val);+}++staticvoidnetc_mac_config(structphylink_config*config,unsignedintmode,+conststructphylink_link_state*state)+{+structdsa_port*dp=dsa_phylink_to_port(config);++netc_port_set_mac_mode(NETC_PORT(dp->ds,dp->index),mode,+state->interface);+}++staticvoidnetc_port_set_speed(structnetc_port*np,intspeed)+{+netc_port_rmw(np,NETC_PCR,PCR_PSPEED,PSPEED_SET_VAL(speed));+}++/* If the RGMII device does not support the In-Band Status (IBS), we need+*theMACdrivertogetthelinkspeedandduplexmodefromthePHYdriver.+*TheMACdriverthensetstheMACforthecorrectspeedandduplexmode+*tomatchthePHY.ThePHYdrivergetsthelinkstatusandspeedandduplex+*informationfromthePHYviatheMDIO/MDCinterface.+*/+staticvoidnetc_port_force_set_rgmii_mac(structnetc_port*np,+intspeed,intduplex)+{+u32mask,val;++mask=PM_IF_MODE_ENA|PM_IF_MODE_SSP|PM_IF_MODE_HD|+PM_IF_MODE_M10|PM_IF_MODE_REVMII;++switch(speed){+default:+caseSPEED_1000:+val=FIELD_PREP(PM_IF_MODE_SSP,SSP_1G);+break;+caseSPEED_100:+val=FIELD_PREP(PM_IF_MODE_SSP,SSP_100M);+break;+caseSPEED_10:+val=FIELD_PREP(PM_IF_MODE_SSP,SSP_10M);+break;+}++if(duplex!=DUPLEX_FULL)+val|=PM_IF_MODE_HD;++netc_mac_port_rmw(np,NETC_PM_IF_MODE(0),mask,val);+}++staticvoidnetc_port_set_rmii_mii_mac(structnetc_port*np,+intspeed,intduplex)+{+u32mask,val=0;++mask=PM_IF_MODE_ENA|PM_IF_MODE_SSP|PM_IF_MODE_HD|+PM_IF_MODE_M10;++if(speed==SPEED_10)+val|=PM_IF_MODE_M10;++if(duplex!=DUPLEX_FULL)+val|=PM_IF_MODE_HD;++netc_mac_port_rmw(np,NETC_PM_IF_MODE(0),mask,val);+}++staticvoidnetc_port_set_hd_flow_control(structnetc_port*np,boolen)+{+if(!np->caps.half_duplex)+return;++/* The HD_FCEN is used in conjunction with the PM_HD_FLOW_CTRL+*register,whichhasadefaultvalue,socurrentlywedonot+*setitinthedriver.Thehalfduplexflowcontrolworksby+*thebackpressure,andthebackpressureisessentiallyjust+*alongpreambletransmittedonthelinkintendedtocreate+*acollisionandgetthehalfduplexlinkpartnertodefer.+*/+netc_mac_port_rmw(np,NETC_PM_CMD_CFG(0),PM_CMD_CFG_HD_FCEN,+en?PM_CMD_CFG_HD_FCEN:0);+}++staticvoidnetc_port_mac_rx_enable(structnetc_port*np)+{+netc_port_rmw(np,NETC_POR,PCR_RXDIS,0);+netc_mac_port_rmw(np,NETC_PM_CMD_CFG(0),PM_CMD_CFG_RX_EN,+PM_CMD_CFG_RX_EN);+}++staticvoidnetc_port_wait_rx_empty(structnetc_port*np,intmac)+{+u32val;++if(read_poll_timeout(netc_port_rd,val,val&PM_IEVENT_RX_EMPTY,+100,10000,false,np,NETC_PM_IEVENT(mac)))+dev_warn(np->switch_priv->dev,+"MAC %d of swp%d RX is not empty\n",mac,+np->dp->index);+}++staticvoidnetc_port_mac_rx_graceful_stop(structnetc_port*np)+{+u32val;++if(is_netc_pseudo_port(np))+gotocheck_rx_busy;++if(np->caps.pmac){+netc_port_rmw(np,NETC_PM_CMD_CFG(1),PM_CMD_CFG_RX_EN,0);+netc_port_wait_rx_empty(np,1);+}++netc_port_rmw(np,NETC_PM_CMD_CFG(0),PM_CMD_CFG_RX_EN,0);+netc_port_wait_rx_empty(np,0);++check_rx_busy:+if(read_poll_timeout(netc_port_rd,val,!(val&PSR_RX_BUSY),+100,10000,false,np,NETC_PSR))+dev_warn(np->switch_priv->dev,"swp%d RX is busy\n",+np->dp->index);++netc_port_rmw(np,NETC_POR,PCR_RXDIS,PCR_RXDIS);+}++staticvoidnetc_mac_link_up(structphylink_config*config,+structphy_device*phy,unsignedintmode,+phy_interface_tinterface,intspeed,+intduplex,booltx_pause,boolrx_pause)+{+structdsa_port*dp=dsa_phylink_to_port(config);+structnetc_port*np;++np=NETC_PORT(dp->ds,dp->index);+netc_port_set_speed(np,speed);++if(phy_interface_mode_is_rgmii(interface)&&+!phylink_autoneg_inband(mode)){+netc_port_force_set_rgmii_mac(np,speed,duplex);+}++if(interface==PHY_INTERFACE_MODE_RMII||+interface==PHY_INTERFACE_MODE_REVMII||+interface==PHY_INTERFACE_MODE_MII){+netc_port_set_rmii_mii_mac(np,speed,duplex);+}++netc_port_set_hd_flow_control(np,duplex==DUPLEX_HALF);+netc_port_mac_rx_enable(np);+}++staticvoidnetc_mac_link_down(structphylink_config*config,+unsignedintmode,+phy_interface_tinterface)+{+structdsa_port*dp=dsa_phylink_to_port(config);++netc_port_mac_rx_graceful_stop(NETC_PORT(dp->ds,dp->index));+}++staticconststructphylink_mac_opsnetc_phylink_mac_ops={+.mac_config=netc_mac_config,+.mac_link_up=netc_mac_link_up,+.mac_link_down=netc_mac_link_down,+};+staticconststructdsa_switch_opsnetc_switch_ops={.get_tag_protocol=netc_get_tag_protocol,.setup=netc_setup,.teardown=netc_teardown,+.phylink_get_caps=netc_phylink_get_caps,};staticintnetc_switch_probe(structpci_dev*pdev,
@@ -613,6 +824,7 @@ static int netc_switch_probe(struct pci_dev *pdev,ds->num_ports=priv->info->num_ports;ds->num_tx_queues=NETC_TC_NUM;ds->ops=&netc_switch_ops;+ds->phylink_mac_ops=&netc_phylink_mac_ops;ds->priv=priv;priv->ds=ds;
@@ -11,8 +11,48 @@ struct netc_switch_platform {conststructnetc_switch_info*info;};+staticvoidimx94_switch_phylink_get_caps(intport,+structphylink_config*config)+{+config->mac_capabilities=MAC_ASYM_PAUSE|MAC_SYM_PAUSE|+MAC_1000FD;++switch(port){+case0...1:+__set_bit(PHY_INTERFACE_MODE_SGMII,+config->supported_interfaces);+__set_bit(PHY_INTERFACE_MODE_1000BASEX,+config->supported_interfaces);+__set_bit(PHY_INTERFACE_MODE_2500BASEX,+config->supported_interfaces);+config->mac_capabilities|=MAC_2500FD;+fallthrough;+case2:+config->mac_capabilities|=MAC_10|MAC_100;+__set_bit(PHY_INTERFACE_MODE_MII,+config->supported_interfaces);+__set_bit(PHY_INTERFACE_MODE_RMII,+config->supported_interfaces);+if(port==2)+__set_bit(PHY_INTERFACE_MODE_REVMII,+config->supported_interfaces);++phy_interface_set_rgmii(config->supported_interfaces);+break;+case3:/* CPU port */+__set_bit(PHY_INTERFACE_MODE_INTERNAL,+config->supported_interfaces);+config->mac_capabilities|=MAC_10FD|MAC_100FD|+MAC_2500FD;+break;+default:+break;+}+}+staticconststructnetc_switch_infoimx94_info={.num_ports=4,+.phylink_get_caps=imx94_switch_phylink_get_caps,};staticconststructnetc_switch_platformnetc_platforms[]={
This patch expands the NETC switch driver with several foundational
features, including FDB and MDB management, STP state handling, MTU
configuration, port setup/teardown, and host flooding support.
At this stage, the driver operates only in standalone port mode. Each
port uses VLAN 0 as its PVID, meaning ingress frames are internally
assigned VID 0 regardless of whether they arrive tagged or untagged.
Note that this does not inject a VLAN 0 header into the frame, the VID
is used purely for subsequent VLAN processing within the switch.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/dsa/netc/netc_main.c | 540 ++++++++++++++++++++++++++
drivers/net/dsa/netc/netc_switch.h | 33 ++
drivers/net/dsa/netc/netc_switch_hw.h | 11 +
3 files changed, 584 insertions(+)
@@ -386,6 +411,212 @@ static void netc_port_default_config(struct netc_port *np)netc_port_rmw(np,NETC_POR,PCR_TXDIS,0);}+staticu32netc_available_port_bitmap(structnetc_switch*priv)+{+structdsa_port*dp;+u32bitmap=0;++dsa_switch_for_each_available_port(dp,priv->ds)+bitmap|=BIT(dp->index);++returnbitmap;+}++staticintnetc_add_standalone_vlan_entry(structnetc_switch*priv)+{+u32bitmap_stg=VFT_STG_ID(0)|netc_available_port_bitmap(priv);+structvft_cfge_data*cfge;+u16cfg;+interr;++cfge=kzalloc_obj(*cfge);+if(!cfge)+return-ENOMEM;++cfge->bitmap_stg=cpu_to_le32(bitmap_stg);+cfge->et_eid=cpu_to_le32(NTMP_NULL_ENTRY_ID);+cfge->fid=cpu_to_le16(NETC_STANDALONE_PVID);++/* For standalone ports, MAC learning needs to be disabled, so frames+*fromotheruserportswillnotbeforwardedtothestandaloneports,+*becausetherearenoFDBentriesonthestandaloneports.Also,the+*framesreceivedbythestandaloneportscannotbefloodedtoother+*ports,soMACforwardingoptionneedstobesetto+*MFO_NO_MATCH_DISCARD,sotheframeswilldiscardedratherthan+*floodingtootherports.+*/+cfg=FIELD_PREP(VFT_MLO,MLO_DISABLE)|+FIELD_PREP(VFT_MFO,MFO_NO_MATCH_DISCARD);+cfge->cfg=cpu_to_le16(cfg);++err=ntmp_vft_add_entry(&priv->ntmp,NETC_STANDALONE_PVID,cfge);+if(err)+dev_err(priv->dev,+"Failed to add standalone VLAN entry\n");++kfree(cfge);++returnerr;+}++staticintnetc_port_add_fdb_entry(structnetc_port*np,+constunsignedchar*addr,u16vid)+{+structnetc_switch*priv=np->switch_priv;+structnetc_fdb_entry*entry;+structfdbt_keye_data*keye;+structfdbt_cfge_data*cfge;+intport=np->dp->index;+u32cfg=0;+interr;++entry=kzalloc_obj(*entry);+if(!entry)+return-ENOMEM;++keye=&entry->keye;+cfge=&entry->cfge;+ether_addr_copy(keye->mac_addr,addr);+keye->fid=cpu_to_le16(vid);++cfge->port_bitmap=cpu_to_le32(BIT(port));+cfge->cfg=cpu_to_le32(cfg);+cfge->et_eid=cpu_to_le32(NTMP_NULL_ENTRY_ID);++err=ntmp_fdbt_add_entry(&priv->ntmp,&entry->entry_id,keye,cfge);+if(err){+kfree(entry);++returnerr;+}++netc_add_fdb_entry(priv,entry);++return0;+}++staticintnetc_port_set_fdb_entry(structnetc_port*np,+constunsignedchar*addr,u16vid)+{+structnetc_switch*priv=np->switch_priv;+structnetc_fdb_entry*entry;+intport=np->dp->index;+u32port_bitmap;+interr=0;++mutex_lock(&priv->fdbt_lock);++entry=netc_lookup_fdb_entry(priv,addr,vid);+if(!entry){+err=netc_port_add_fdb_entry(np,addr,vid);+if(err)+dev_err(priv->dev,+"Failed to add FDB entry on port %d\n",+port);++gotounlock_fdbt;+}++port_bitmap=le32_to_cpu(entry->cfge.port_bitmap);+/* If the entry already exists on the port, return 0 directly */+if(unlikely(port_bitmap&BIT(port)))+gotounlock_fdbt;++/* If the entry already exists, but not on this port, we need to+*updatetheportbitmap.Ingeneral,itshouldonlybevalid+*formulticastorbroadcastaddress.+*/+port_bitmap^=BIT(port);+entry->cfge.port_bitmap=cpu_to_le32(port_bitmap);+err=ntmp_fdbt_update_entry(&priv->ntmp,entry->entry_id,+&entry->cfge);+if(err){+port_bitmap^=BIT(port);+entry->cfge.port_bitmap=cpu_to_le32(port_bitmap);+dev_err(priv->dev,"Failed to set FDB entry on port %d\n",+port);+}++unlock_fdbt:+mutex_unlock(&priv->fdbt_lock);++returnerr;+}++staticintnetc_port_del_fdb_entry(structnetc_port*np,+constunsignedchar*addr,u16vid)+{+structnetc_switch*priv=np->switch_priv;+structntmp_user*ntmp=&priv->ntmp;+structnetc_fdb_entry*entry;+intport=np->dp->index;+u32port_bitmap;+interr=0;++mutex_lock(&priv->fdbt_lock);++entry=netc_lookup_fdb_entry(priv,addr,vid);+if(unlikely(!entry))+gotounlock_fdbt;++port_bitmap=le32_to_cpu(entry->cfge.port_bitmap);+if(unlikely(!(port_bitmap&BIT(port))))+gotounlock_fdbt;++if(port_bitmap!=BIT(port)){+/* If the entry also exists on other ports, we need to+*updatetheentryintheFDBtable.+*/+port_bitmap^=BIT(port);+entry->cfge.port_bitmap=cpu_to_le32(port_bitmap);+err=ntmp_fdbt_update_entry(ntmp,entry->entry_id,+&entry->cfge);+if(err){+port_bitmap^=BIT(port);+entry->cfge.port_bitmap=cpu_to_le32(port_bitmap);+gotounlock_fdbt;+}+}else{+/* If the entry only exists on this port, just delete+*itfromtheFDBtable.+*/+err=ntmp_fdbt_delete_entry(ntmp,entry->entry_id);+if(err)+gotounlock_fdbt;++netc_del_fdb_entry(entry);+}++unlock_fdbt:+mutex_unlock(&priv->fdbt_lock);++returnerr;+}++staticintnetc_add_standalone_fdb_bcast_entry(structnetc_switch*priv)+{+constu8bcast[ETH_ALEN]={0xff,0xff,0xff,0xff,0xff,0xff};+structdsa_port*dp,*cpu_dp=NULL;++dsa_switch_for_each_cpu_port(dp,priv->ds){+cpu_dp=dp;+break;+}++if(!cpu_dp)+return-ENODEV;++/* If the user port acts as a standalone port, then its PVID is 0,+*MLOissetto"disable MAC learning"andMFOissetto"discard+*framesifnomatchingentryfoundinFDBtable". Therefore, we+*needtoaddabroadcastFDBentryontheCPUportsothatthe+*broadcastframesreceivedontheuserportcanbeforwardedto+*theCPUport.+*/+returnnetc_port_set_fdb_entry(NETC_PORT(priv->ds,cpu_dp->index),+bcast,NETC_STANDALONE_PVID);+}+staticintnetc_setup(structdsa_switch*ds){structnetc_switch*priv=ds->priv;
@@ -404,19 +635,61 @@ static int netc_setup(struct dsa_switch *ds)if(err)returnerr;+INIT_HLIST_HEAD(&priv->fdb_list);+mutex_init(&priv->fdbt_lock);+netc_switch_fixed_config(priv);/* default setting for ports */dsa_switch_for_each_available_port(dp,ds)netc_port_default_config(priv->ports[dp->index]);+err=netc_add_standalone_vlan_entry(priv);+if(err)+gotofree_lock_and_ntmp_user;++err=netc_add_standalone_fdb_bcast_entry(priv);+if(err)+gotofree_lock_and_ntmp_user;+return0;++free_lock_and_ntmp_user:+mutex_destroy(&priv->fdbt_lock);+netc_free_ntmp_user(priv);++returnerr;+}++staticvoidnetc_destroy_all_lists(structnetc_switch*priv)+{+netc_destroy_fdb_list(priv);+mutex_destroy(&priv->fdbt_lock);+}++staticvoidnetc_free_host_flood_rules(structnetc_switch*priv)+{+structdsa_port*dp;++dsa_switch_for_each_user_port(dp,priv->ds){+structnetc_port*np=priv->ports[dp->index];++/* No need to clear the hardware IPFT entry. Because PCIe+*FLRwillbeperformedwhentheswitchisre-registered,+*itwillresethardwarestate.Soonlyneedtofreethe+*memorytoavoidmemoryleak.+*/+kfree(np->host_flood);+np->host_flood=NULL;+}}staticvoidnetc_teardown(structdsa_switch*ds){structnetc_switch*priv=ds->priv;+netc_destroy_all_lists(priv);+netc_free_host_flood_rules(priv);netc_free_ntmp_user(priv);}
@@ -569,6 +842,261 @@ static void netc_switch_get_ip_revision(struct netc_switch *priv)priv->revision=val&IPBRR0_IP_REV;}+staticintnetc_port_enable(structdsa_switch*ds,intport,+structphy_device*phy)+{+structnetc_port*np=NETC_PORT(ds,port);+interr;++if(np->enable)+return0;++err=clk_prepare_enable(np->ref_clk);+if(err){+dev_err(ds->dev,+"Failed to enable enet_ref_clk of port %d\n",port);+returnerr;+}++np->enable=true;++return0;+}++staticvoidnetc_port_disable(structdsa_switch*ds,intport)+{+structnetc_port*np=NETC_PORT(ds,port);++/* When .port_disable() is called, .port_enable() may not have been+*called.Inthiscase,boththeprepare_countandenable_countof+*clockare0.Callingclk_disable_unprepare()atthistimewill+*causewarnings.+*/+if(!np->enable)+return;++clk_disable_unprepare(np->ref_clk);+np->enable=false;+}++staticvoidnetc_port_stp_state_set(structdsa_switch*ds,+intport,u8state)+{+structnetc_port*np=NETC_PORT(ds,port);+u32val;++switch(state){+caseBR_STATE_DISABLED:+caseBR_STATE_LISTENING:+caseBR_STATE_BLOCKING:+val=NETC_STG_STATE_DISABLED;+break;+caseBR_STATE_LEARNING:+val=NETC_STG_STATE_LEARNING;+break;+caseBR_STATE_FORWARDING:+val=NETC_STG_STATE_FORWARDING;+break;+default:+return;+}++netc_port_wr(np,NETC_BPSTGSR,val);+}++staticintnetc_port_change_mtu(structdsa_switch*ds,+intport,intmtu)+{+u32max_frame_size=mtu+VLAN_ETH_HLEN+ETH_FCS_LEN;+structnetc_port*np=NETC_PORT(ds,port);++if(dsa_is_cpu_port(ds,port))+max_frame_size+=NETC_TAG_MAX_LEN;++netc_port_set_max_frame_size(np,max_frame_size);++return0;+}++staticintnetc_port_max_mtu(structdsa_switch*ds,intport)+{+returnNETC_MAX_FRAME_LEN-VLAN_ETH_HLEN-ETH_FCS_LEN;+}++staticintnetc_port_fdb_add(structdsa_switch*ds,intport,+constunsignedchar*addr,u16vid,+structdsa_dbdb)+{+structnetc_port*np=NETC_PORT(ds,port);++/* Currently, we only support standalone port mode, so all VLANs+*shouldbeconvertedtoNETC_STANDALONE_PVID.+*/+returnnetc_port_set_fdb_entry(np,addr,NETC_STANDALONE_PVID);+}++staticintnetc_port_fdb_del(structdsa_switch*ds,intport,+constunsignedchar*addr,u16vid,+structdsa_dbdb)+{+structnetc_port*np=NETC_PORT(ds,port);++returnnetc_port_del_fdb_entry(np,addr,NETC_STANDALONE_PVID);+}++staticintnetc_port_fdb_dump(structdsa_switch*ds,intport,+dsa_fdb_dump_cb_t*cb,void*data)+{+structnetc_switch*priv=ds->priv;+u32resume_eid=NTMP_NULL_ENTRY_ID;+structfdbt_entry_data*entry;+structfdbt_keye_data*keye;+structfdbt_cfge_data*cfge;+boolis_static;+u32cfg;+interr;+u16vid;++entry=kmalloc_obj(*entry);+if(!entry)+return-ENOMEM;++keye=&entry->keye;+cfge=&entry->cfge;+mutex_lock(&priv->fdbt_lock);++do{+memset(entry,0,sizeof(*entry));+err=ntmp_fdbt_search_port_entry(&priv->ntmp,port,+&resume_eid,entry);+if(err||entry->entry_id==NTMP_NULL_ENTRY_ID)+break;++cfg=le32_to_cpu(cfge->cfg);+is_static=(cfg&FDBT_DYNAMIC)?false:true;+vid=le16_to_cpu(keye->fid);++err=cb(keye->mac_addr,vid,is_static,data);+if(err)+break;+}while(resume_eid!=NTMP_NULL_ENTRY_ID);++mutex_unlock(&priv->fdbt_lock);+kfree(entry);++returnerr;+}++staticintnetc_port_mdb_add(structdsa_switch*ds,intport,+conststructswitchdev_obj_port_mdb*mdb,+structdsa_dbdb)+{+returnnetc_port_fdb_add(ds,port,mdb->addr,mdb->vid,db);+}++staticintnetc_port_mdb_del(structdsa_switch*ds,intport,+conststructswitchdev_obj_port_mdb*mdb,+structdsa_dbdb)+{+returnnetc_port_fdb_del(ds,port,mdb->addr,mdb->vid,db);+}++staticintnetc_port_add_host_flood_rule(structnetc_port*np,+booluc,boolmc)+{+constu8dmac_mask[ETH_ALEN]={0x1,0,0,0,0,0};+structnetc_switch*priv=np->switch_priv;+structipft_entry_data*host_flood;+structipft_keye_data*keye;+structipft_cfge_data*cfge;+u16src_port;+u32cfg;+interr;++if(!uc&&!mc)+return0;++host_flood=kzalloc_obj(*host_flood);+if(!host_flood)+return-ENOMEM;++keye=&host_flood->keye;+cfge=&host_flood->cfge;++src_port=FIELD_PREP(IPFT_SRC_PORT,np->dp->index);+src_port|=IPFT_SRC_PORT_MASK;+keye->src_port=cpu_to_le16(src_port);++/* If either only unicast or only multicast need to be flooded+*tothehost,wealwayssetthemaskthatteststhefirstMAC+*DAoctet.Thevalueshouldbe0forthefirstbit(ifunicast+*hastobeflooded)or1(ifmulticast).Ifbothunicastand+*multicasthavetobeflooded,weleavethekeymaskempty,so+*itmatcheseverything.+*/+if(uc&&!mc)+ether_addr_copy(keye->dmac_mask,dmac_mask);++if(!uc&&mc){+ether_addr_copy(keye->dmac,dmac_mask);+ether_addr_copy(keye->dmac_mask,dmac_mask);+}++cfg=FIELD_PREP(IPFT_FLTFA,IPFT_FLTFA_REDIRECT);+cfg|=FIELD_PREP(IPFT_HR,NETC_HR_HOST_FLOOD);+cfge->cfg=cpu_to_le32(cfg);++err=ntmp_ipft_add_entry(&priv->ntmp,host_flood);+if(err){+kfree(host_flood);+returnerr;+}++np->uc=uc;+np->mc=mc;+np->host_flood=host_flood;+/* Enable ingress port filter table lookup */+netc_port_wr(np,NETC_PIPFCR,PIPFCR_EN);++return0;+}++staticvoidnetc_port_remove_host_flood(structnetc_port*np)+{+structnetc_switch*priv=np->switch_priv;++if(!np->host_flood)+return;++ntmp_ipft_delete_entry(&priv->ntmp,np->host_flood->entry_id);+kfree(np->host_flood);+np->host_flood=NULL;+np->uc=false;+np->mc=false;+/* Disable ingress port filter table lookup */+netc_port_wr(np,NETC_PIPFCR,0);+}++staticvoidnetc_port_set_host_flood(structdsa_switch*ds,intport,+booluc,boolmc)+{+structnetc_port*np=NETC_PORT(ds,port);++if(np->uc==uc&&np->mc==mc)+return;++/* IPFT does not support in-place updates to the KEYE element,+*soweneedtodeletetheoldIPFTentryandthenaddanew+*one.+*/+if(np->host_flood)+netc_port_remove_host_flood(np);++if(netc_port_add_host_flood_rule(np,uc,mc))+dev_err(ds->dev,"Failed to add host flood rule on port %d\n",+port);+}+staticvoidnetc_phylink_get_caps(structdsa_switch*ds,intport,structphylink_config*config){
The buffer pool is a quantity of memory available for buffering a group
of flows (e.g. frames having the same priority, frames received from the
same port), while waiting to be transmitted on a port. The buffer pool
tracks internal memory consumption with upper bound limits and optionally
a non-shared portion when associated with a shared buffer pool. Currently
the shared buffer pool is not supported, it will be added in the future.
For i.MX94, the switch has 4 ports and 8 buffer pools, so each port is
allocated two buffer pools. For frames with priorities of 0 to 3, they
will be mapped to the first buffer pool; For frames with priorities of
4 to 7, they will be mapped to the second buffer pool. Each buffer pool
has a flow control on threshold and a flow control off threshold. By
setting these threshold, add the flow control support to each port.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/dsa/netc/netc_main.c | 133 ++++++++++++++++++++++++++
drivers/net/dsa/netc/netc_switch.h | 9 ++
drivers/net/dsa/netc/netc_switch_hw.h | 12 +++
3 files changed, 154 insertions(+)
@@ -379,6 +379,8 @@ static void netc_port_set_mlo(struct netc_port *np, enum netc_mlo mlo)staticvoidnetc_port_fixed_config(structnetc_port*np){+u32pqnt=0xffff,qth=0xff00;+/* Default IPV and DR setting */netc_port_rmw(np,NETC_PQOSMR,PQOSMR_VS|PQOSMR_VE,PQOSMR_VS|PQOSMR_VE);
@@ -386,6 +388,15 @@ static void netc_port_fixed_config(struct netc_port *np)/* Enable L2 and L3 DOS */netc_port_rmw(np,NETC_PCR,PCR_L2DOSE|PCR_L3DOSE,PCR_L2DOSE|PCR_L3DOSE);++/* Set the quanta value of TX PAUSE frame */+netc_mac_port_wr(np,NETC_PM_PAUSE_QUANTA(0),pqnt);++/* When a quanta timer counts down and reaches this value,+*theMACsendsarefreshPAUSEframewiththeprogrammed+*fullquantavalueifapauseconditionstillexists.+*/+netc_mac_port_wr(np,NETC_PM_PAUSE_TRHESH(0),qth);}staticvoidnetc_port_default_config(structnetc_port*np)
@@ -617,6 +628,80 @@ static int netc_add_standalone_fdb_bcast_entry(struct netc_switch *priv)bcast,NETC_STANDALONE_PVID);}+staticu32netc_get_buffer_pool_num(structnetc_switch*priv)+{+returnnetc_base_rd(&priv->regs,NETC_BPCAPR)&BPCAPR_NUM_BP;+}++staticvoidnetc_port_set_pbpmcr(structnetc_port*np,u64mapping)+{+u32pbpmcr0=lower_32_bits(mapping);+u32pbpmcr1=upper_32_bits(mapping);++netc_port_wr(np,NETC_PBPMCR0,pbpmcr0);+netc_port_wr(np,NETC_PBPMCR1,pbpmcr1);+}++staticvoidnetc_ipv_to_buffer_pool_mapping(structnetc_switch*priv)+{+intnum_port_bp=priv->num_bp/priv->info->num_ports;+intq=NETC_IPV_NUM/num_port_bp;+intr=NETC_IPV_NUM%num_port_bp;+intnum=q+r;++/* IPV-to–buffer-pool mapping per port:+*Eachportisallocated'num_port_bp'bufferpoolsandsupports8+*IPVs,whereahigherIPVindicatesahigherframepriority.Each+*IPVcanbemappedtoonlyonebufferpool.+*+*Themappingruleisasfollows:+*-Thefirst'num'IPVssharetheport'sfirstbufferpool(index+*'base_id').+*-Afterthat,every'q'IPVsshareonebufferpool,withpool+*indicesincreasingsequentially.+*/+for(inti=0;i<priv->info->num_ports;i++){+u32base_id=i*num_port_bp;+u32bp_id=base_id;+u64mapping=0;++for(intipv=0;ipv<NETC_IPV_NUM;ipv++){+/* Update the buffer pool index */+if(ipv>=num)+bp_id=base_id+((ipv-num)/q)+1;++mapping|=(u64)bp_id<<(ipv*8);+}++netc_port_set_pbpmcr(priv->ports[i],mapping);+}+}++staticintnetc_switch_bpt_default_config(structnetc_switch*priv)+{+priv->num_bp=netc_get_buffer_pool_num(priv);+priv->bpt_list=devm_kcalloc(priv->dev,priv->num_bp,+sizeof(structbpt_cfge_data),+GFP_KERNEL);+if(!priv->bpt_list)+return-ENOMEM;++/* Initialize the maximum threshold of each buffer pool entry */+for(inti=0;i<priv->num_bp;i++){+structbpt_cfge_data*cfge=&priv->bpt_list[i];+interr;++cfge->max_thresh=cpu_to_le16(NETC_BP_THRESH);+err=ntmp_bpt_update_entry(&priv->ntmp,i,cfge);+if(err)+returnerr;+}++netc_ipv_to_buffer_pool_mapping(priv);++return0;+}+staticintnetc_setup(structdsa_switch*ds){structnetc_switch*priv=ds->priv;
@@ -644,6 +729,10 @@ static int netc_setup(struct dsa_switch *ds)dsa_switch_for_each_available_port(dp,ds)netc_port_default_config(priv->ports[dp->index]);+err=netc_switch_bpt_default_config(priv);+if(err)+gotofree_lock_and_ntmp_user;+err=netc_add_standalone_vlan_entry(priv);if(err)gotofree_lock_and_ntmp_user;
@@ -1288,7 +1411,17 @@ static void netc_mac_link_up(struct phylink_config *config,netc_port_set_rmii_mii_mac(np,speed,duplex);}+if(duplex==DUPLEX_HALF){+/* As per 802.3 annex 31B, PAUSE frames are only supported+*whenthelinkisconfiguredforfullduplexoperation.+*/+tx_pause=false;+rx_pause=false;+}+netc_port_set_hd_flow_control(np,duplex==DUPLEX_HALF);+netc_port_set_tx_pause(np,tx_pause);+netc_port_set_rx_pause(np,rx_pause);netc_port_mac_rx_enable(np);}
Each user port of the NETC switch supports 802.3 basic and mandatory
managed objects statistic counters and IETF Management Information
Database (MIB) package (RFC2665) and Remote Network Monitoring (RMON)
counters. And all of these counters are 64-bit registers. In addition,
some user ports support preemption, so these ports have two MACs, MAC
0 is the express MAC (eMAC), MAC 1 is the preemptible MAC (pMAC). So
for ports that support preemption, the statistics are the sum of the
pMAC and eMAC statistics.
Note that the current switch driver does not support preemption, all
frames are sent and received via the eMAC by default. The statistics
read from the pMAC should be zero.
Signed-off-by: Wei Fang <wei.fang@nxp.com>
---
drivers/net/dsa/netc/Makefile | 2 +-
drivers/net/dsa/netc/netc_ethtool.c | 192 ++++++++++++++++++++++++++
drivers/net/dsa/netc/netc_main.c | 4 +
drivers/net/dsa/netc/netc_switch.h | 17 +++
drivers/net/dsa/netc/netc_switch_hw.h | 153 ++++++++++++++++++++
include/linux/fsl/netc_global.h | 6 +
6 files changed, 373 insertions(+), 1 deletion(-)
create mode 100644 drivers/net/dsa/netc/netc_ethtool.c
From: "Russell King (Oracle)" <linux@armlinux.org.uk> Date: 2026-03-23 09:20:19
On Mon, Mar 23, 2026 at 02:07:51PM +0800, Wei Fang wrote:
quoted hunk
@@ -1288,7 +1411,17 @@ static void netc_mac_link_up(struct phylink_config *config, netc_port_set_rmii_mii_mac(np, speed, duplex); }+ if (duplex == DUPLEX_HALF) {+ /* As per 802.3 annex 31B, PAUSE frames are only supported+ * when the link is configured for full duplex operation.+ */+ tx_pause = false;+ rx_pause = false;+ }
I keep seeing this totally unnecessary code in reviews of mac_link_up()
methods. See phylink_resolve_an_pause().
--
RMK's Patch system: https://www.armlinux.org.uk/developer/patches/
FTTP is here! 80Mbps down 10Mbps up. Decent connectivity at last!
From: "Russell King (Oracle)" <linux@armlinux.org.uk> Date: 2026-03-23 09:30:56
On Mon, Mar 23, 2026 at 02:07:49PM +0800, Wei Fang wrote:
+static void netc_port_set_mac_mode(struct netc_port *np,
+ unsigned int mode,
+ phy_interface_t phy_mode)
+{
+ u32 mask = PM_IF_MODE_IFMODE | PM_IF_MODE_REVMII | PM_IF_MODE_ENA;
+ u32 val = 0;
+
+ switch (phy_mode) {
+ case PHY_INTERFACE_MODE_RGMII:
+ case PHY_INTERFACE_MODE_RGMII_ID:
+ case PHY_INTERFACE_MODE_RGMII_RXID:
+ case PHY_INTERFACE_MODE_RGMII_TXID:
+ val |= IFMODE_RGMII;
+ /* Enable auto-negotiation for the MAC if its
+ * RGMII interface supports In-Band status.
+ */
+ if (phylink_autoneg_inband(mode))
+ val |= PM_IF_MODE_ENA;
I would prefer newer drivers not to use phylink_autoneg_inband()
anymore. Note that there is no need to support RGMII inband in the
kernel (nor is there any proper support without a "phylink_pcs"
being present to provide the inband status.)
+static void netc_port_set_hd_flow_control(struct netc_port *np, bool en)
+{
+ if (!np->caps.half_duplex)
+ return;
+
+ /* The HD_FCEN is used in conjunction with the PM_HD_FLOW_CTRL
+ * register, which has a default value, so currently we do not
+ * set it in the driver. The half duplex flow control works by
+ * the backpressure, and the backpressure is essentially just
+ * a long preamble transmitted on the link intended to create
+ * a collision and get the half duplex link partner to defer.
+ */
+ netc_mac_port_rmw(np, NETC_PM_CMD_CFG(0), PM_CMD_CFG_HD_FCEN,
+ en ? PM_CMD_CFG_HD_FCEN : 0);
We don't support half duplex backpressure in the kernel. I notice
you always enable this whenever HD mode is negotiated, which means
there's no way for the user to disable it. Flow control can cause
problems. Ethernet relies on packet dropping for congestion
management.
The "case 2" above already ensures that port is 2 here.
--
RMK's Patch system: https://www.armlinux.org.uk/developer/patches/
FTTP is here! 80Mbps down 10Mbps up. Decent connectivity at last!
netc_port_set_rmii_mii_mac(np, speed, duplex);
}
+ if (duplex == DUPLEX_HALF) {
+ /* As per 802.3 annex 31B, PAUSE frames are only supported
+ * when the link is configured for full duplex operation.
+ */
+ tx_pause = false;
+ rx_pause = false;
+ }
I keep seeing this totally unnecessary code in reviews of mac_link_up()
methods. See phylink_resolve_an_pause().
+ u32 val = 0;
+
+ switch (phy_mode) {
+ case PHY_INTERFACE_MODE_RGMII:
+ case PHY_INTERFACE_MODE_RGMII_ID:
+ case PHY_INTERFACE_MODE_RGMII_RXID:
+ case PHY_INTERFACE_MODE_RGMII_TXID:
+ val |= IFMODE_RGMII;
+ /* Enable auto-negotiation for the MAC if its
+ * RGMII interface supports In-Band status.
+ */
+ if (phylink_autoneg_inband(mode))
+ val |= PM_IF_MODE_ENA;
I would prefer newer drivers not to use phylink_autoneg_inband()
anymore. Note that there is no need to support RGMII inband in the
kernel (nor is there any proper support without a "phylink_pcs"
being present to provide the inband status.)
Thanks for the info, I will remove it.
quoted
+static void netc_port_set_hd_flow_control(struct netc_port *np, bool en)
+{
+ if (!np->caps.half_duplex)
+ return;
+
+ /* The HD_FCEN is used in conjunction with the PM_HD_FLOW_CTRL
+ * register, which has a default value, so currently we do not
+ * set it in the driver. The half duplex flow control works by
+ * the backpressure, and the backpressure is essentially just
+ * a long preamble transmitted on the link intended to create
+ * a collision and get the half duplex link partner to defer.
+ */
+ netc_mac_port_rmw(np, NETC_PM_CMD_CFG(0),
PM_CMD_CFG_HD_FCEN,
quoted
+ en ? PM_CMD_CFG_HD_FCEN : 0);
We don't support half duplex backpressure in the kernel. I notice
you always enable this whenever HD mode is negotiated, which means
there's no way for the user to disable it. Flow control can cause
problems. Ethernet relies on packet dropping for congestion
management.
From: Paolo Abeni <pabeni@redhat.com> Date: 2026-03-24 08:13:40
On 3/23/26 10:30 AM, Russell King (Oracle) wrote:
On Mon, Mar 23, 2026 at 02:07:49PM +0800, Wei Fang wrote:
quoted
+static void netc_port_set_mac_mode(struct netc_port *np,
+ unsigned int mode,
+ phy_interface_t phy_mode)
+{
+ u32 mask = PM_IF_MODE_IFMODE | PM_IF_MODE_REVMII | PM_IF_MODE_ENA;
+ u32 val = 0;
+
+ switch (phy_mode) {
+ case PHY_INTERFACE_MODE_RGMII:
+ case PHY_INTERFACE_MODE_RGMII_ID:
+ case PHY_INTERFACE_MODE_RGMII_RXID:
+ case PHY_INTERFACE_MODE_RGMII_TXID:
+ val |= IFMODE_RGMII;
+ /* Enable auto-negotiation for the MAC if its
+ * RGMII interface supports In-Band status.
+ */
+ if (phylink_autoneg_inband(mode))
+ val |= PM_IF_MODE_ENA;
I would prefer newer drivers not to use phylink_autoneg_inband()
anymore. Note that there is no need to support RGMII inband in the
kernel (nor is there any proper support without a "phylink_pcs"
being present to provide the inband status.)
Would you mind sharing a patch adding a comment nearby
phylink_autoneg_inband(), explaining the above for future memory?
Thanks,
Paolo
From: Simon Horman <horms@kernel.org> Date: 2026-03-24 16:43:09
This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
For local reproduction steps see https://netdev-ai.bots.linux.dev/ai-local.html
---
net: dsa: netc: initialize buffer bool table and implement flow-control
The buffer pool is a quantity of memory available for buffering a group
of flows (e.g. frames having the same priority, frames received from the
same port), while waiting to be transmitted on a port. The buffer pool
tracks internal memory consumption with upper bound limits and optionally
a non-shared portion when associated with a shared buffer pool. Currently
the shared buffer pool is not supported, it will be added in the future.
For i.MX94, the switch has 4 ports and 8 buffer pools, so each port is
allocated two buffer pools. For frames with priorities of 0 to 3, they
will be mapped to the first buffer pool; For frames with priorities of
4 to 7, they will be mapped to the second buffer pool. Each buffer pool
has a flow control on threshold and a flow control off threshold. By
setting these threshold, add the flow control support to each port.
Can this cause a division by zero? If netc_get_buffer_pool_num() reads
the hardware register NETC_BPCAPR and returns 0, then num_port_bp would be
0 and this division would crash the kernel.
The call chain is:
netc_setup()
netc_switch_bpt_default_config()
priv->num_bp = netc_get_buffer_pool_num()
netc_ipv_to_buffer_pool_mapping()
num_port_bp = priv->num_bp / priv->info->num_ports;
q = NETC_IPV_NUM / num_port_bp; <- division by zero
Should netc_switch_bpt_default_config() validate that priv->num_bp is
non-zero before proceeding?
+ int r = NETC_IPV_NUM % num_port_bp;
+ int num = q + r;
+
+ /* IPV-to–buffer-pool mapping per port:
[ ... ]
+static int netc_switch_bpt_default_config(struct netc_switch *priv)
+{
+ priv->num_bp = netc_get_buffer_pool_num(priv);
+ priv->bpt_list = devm_kcalloc(priv->dev, priv->num_bp,
+ sizeof(struct bpt_cfge_data),
+ GFP_KERNEL);
+ if (!priv->bpt_list)
+ return -ENOMEM;
+
+ /* Initialize the maximum threshold of each buffer pool entry */
+ for (int i = 0; i < priv->num_bp; i++) {
+ struct bpt_cfge_data *cfge = &priv->bpt_list[i];
+ int err;
+
+ cfge->max_thresh = cpu_to_le16(NETC_BP_THRESH);
+ err = ntmp_bpt_update_entry(&priv->ntmp, i, cfge);
+ if (err)
+ return err;
+ }
+
+ netc_ipv_to_buffer_pool_mapping(priv);
+
+ return 0;
+}
+
[ ... ]
quoted hunk
@@ -1224,6 +1313,40 @@ static void netc_port_set_hd_flow_control(struct netc_port *np, bool en) en ? PM_CMD_CFG_HD_FCEN : 0); }+static void netc_port_set_tx_pause(struct netc_port *np, bool tx_pause)+{+ struct netc_switch *priv = np->switch_priv;+ int port = np->dp->index;+ int i, j, num_bp;++ num_bp = priv->num_bp / priv->info->num_ports;+ for (i = 0, j = port * num_bp; i < num_bp; i++, j++) {+ struct bpt_cfge_data *cfge;++ cfge = &priv->bpt_list[j];+ if (tx_pause) {+ cfge->fc_on_thresh = cpu_to_le16(NETC_FC_THRESH_ON);+ cfge->fc_off_thresh = cpu_to_le16(NETC_FC_THRESH_OFF);+ cfge->fccfg_sbpen = FIELD_PREP(BPT_FC_CFG,+ BPT_FC_CFG_EN_BPFC);+ cfge->fc_ports = cpu_to_le32(BIT(port));+ } else {+ cfge->fc_on_thresh = cpu_to_le16(0);+ cfge->fc_off_thresh = cpu_to_le16(0);+ cfge->fccfg_sbpen = 0;+ cfge->fc_ports = cpu_to_le32(0);+ }++ ntmp_bpt_update_entry(&priv->ntmp, j, cfge); ^^^^
Should the return value of ntmp_bpt_update_entry() be checked here?
Simon Horman noted in v1 review that other NTMP table operations in the
driver propagate errors from the update functions. While this was fixed
in netc_switch_bpt_default_config() in v2, the same issue remains here.
If the hardware update fails (for example, DMA allocation failure or
hardware timeout), the driver's bpt_list will have flow control values that
do not match the actual hardware state, potentially causing incorrect flow
control behavior.
https://lore.kernel.org/netdev/20260318145452.1898637-2-horms@kernel.org/
From: Simon Horman <horms@kernel.org> Date: 2026-03-24 16:54:28
This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
For local reproduction steps see https://netdev-ai.bots.linux.dev/ai-local.html
---
net: dsa: add NETC switch tag support
This commit introduces support for the NXP NETC switch tag protocol,
a proprietary header used by NETC switches. The tag supports three
types (Forward, To_Port, and To_Host) with variable length (6-14 bytes).
^^^^^^^^^^^^
Is the tag header guaranteed to be in the linear region of the skb?
Looking at other EtherType-based DSA tag drivers (tag_ar9331, tag_lan9303,
tag_dsa, tag_edsa, tag_a5psw, tag_brcm_legacy, tag_mxl-gsw1xx,
tag_gswip, and tag_mtk), they all call pskb_may_pull() before accessing
tag data to ensure the header is linearized.
For example, tag_ar9331.c:
static struct sk_buff *ar9331_tag_rcv(struct sk_buff *skb,
struct net_device *ndev)
{
if (unlikely(!pskb_may_pull(skb, AR9331_HDR_LEN)))
return NULL;
...
}
Similarly, tag_lan9303.c:
static struct sk_buff *lan9303_rcv(struct sk_buff *skb,
struct net_device *ndev)
{
if (unlikely(!pskb_may_pull(skb, LAN9303_TAG_LEN)))
return NULL;
...
}
Without this check, netc_rcv() accesses tag_cmn->type, tag_cmn->tpid,
tag_cmn->qos, and tag_cmn->switch_port before verifying these bytes are
linearized. The struct netc_tag_cmn is 5 bytes, with 3 bytes at
skb->data through skb->data+2 that may not be guaranteed to be in the
linear region.
Additionally, netc_get_rx_tag_len() can return up to 14 bytes
(NETC_TAG_TH_SUBTYPE1_LEN or NETC_TAG_TH_SUBTYPE2_LEN), and this value
is passed directly to skb_pull_rcsum():
+ int tag_len = netc_get_rx_tag_len(tag_cmn->type);
+ int sw_id, port;
+
+ if (ntohs(tag_cmn->tpid) != ETH_P_NXP_NETC) {
+ dev_warn_ratelimited(&ndev->dev, "Unknown TPID 0x%04x\n",
+ ntohs(tag_cmn->tpid));
+
+ return NULL;
+ }
+
+ if (tag_cmn->qos & NETC_TAG_QV)
+ skb->priority = FIELD_GET(NETC_TAG_IPV, tag_cmn->qos);
+
+ sw_id = NETC_TAG_SWITCH & tag_cmn->switch_port;
+ /* ENETC VEPA switch ID (0) is not supported yet */
+ if (!sw_id) {
+ dev_warn_ratelimited(&ndev->dev,
+ "VEPA switch ID is not supported yet\n");
+
+ return NULL;
+ }
+
+ port = FIELD_GET(NETC_TAG_PORT, tag_cmn->switch_port);
+ skb->dev = dsa_conduit_find_user(ndev, sw_id, port);
+ if (!skb->dev)
+ return NULL;
+
+ if (tag_cmn->type == NETC_TAG_FORWARD)
+ dsa_default_offload_fwd_mark(skb);
+
+ /* Remove Switch tag from the frame */
+ skb_pull_rcsum(skb, tag_len);
^^^^^^^^^^^^^^^^^^
The skb_pull_rcsum() function contains BUG_ON(len > skb->len). If a
truncated or malformed packet arrives with fewer than tag_len bytes
remaining, the kernel will panic rather than gracefully handling the
error.
Should a check like this be added at the beginning of netc_rcv():
if (unlikely(!pskb_may_pull(skb, NETC_TAG_MAX_LEN - 2)))
return NULL;
The -2 would account for the EtherType bytes already consumed by
eth_type_trans().
From: Simon Horman <horms@kernel.org> Date: 2026-03-24 16:58:39
This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
For local reproduction steps see https://netdev-ai.bots.linux.dev/ai-local.html
---
net: enetc: add multiple command BD rings support
This patch adds support for multiple command BD rings to the NTMP driver.
The switch has two command BD rings, and netc_select_cbdr() is added to
select an appropriate ring when the current one is busy.
Frank Li raised a concern about the use of spin_is_locked() here. Is this
lockless check reliable across all architectures? spin_is_locked() behavior
varies across configurations - it can return false negatives on UP builds and
may not provide the ordering guarantees needed for lockless checks.
Would spin_trylock() be more appropriate here? Or if this racy check is
intentional and acceptable, should it be documented why the potential race is
safe in this context?
quoted hunk
+
+ /* If all the command BDRs are busy now, we select
+ * one of them, but need to wait for a while to use.
+ */
+ cpu = raw_smp_processor_id();
+
+ return &user->ring[cpu % user->cbdr_num];
+}
+
static int netc_xmit_ntmp_cmd(struct ntmp_user *user, union netc_cbd *cbd)
{
union netc_cbd *cur_cbd;
@@ -125,10 +144,10 @@ static int netc_xmit_ntmp_cmd(struct ntmp_user *user, union netc_cbd *cbd) u16 status; u32 val;- /* Currently only i.MX95 ENETC is supported, and it only has one- * command BD ring- */- cbdr = &user->ring[0];+ if (user->cbdr_num == 1)+ cbdr = &user->ring[0];+ else+ cbdr = netc_select_cbdr(user);
This isn't a bug, but Frank Li suggested that the special case for
cbdr_num == 1 might be redundant. Since netc_select_cbdr() would handle the
single ring case correctly (the loop would find the only ring on the first
iteration), would it be simpler to just unconditionally call:
cbdr = netc_select_cbdr(user);
Can this cause a division by zero? If netc_get_buffer_pool_num() reads
From hardware perspective, num_bp is more than num_ports, so num_port_bp
will not be zero. Let me think how to improve it to avoid such false positive.
the hardware register NETC_BPCAPR and returns 0, then num_port_bp would
be
0 and this division would crash the kernel.
The call chain is:
netc_setup()
netc_switch_bpt_default_config()
priv->num_bp = netc_get_buffer_pool_num()
netc_ipv_to_buffer_pool_mapping()
num_port_bp = priv->num_bp / priv->info->num_ports;
q = NETC_IPV_NUM / num_port_bp; <- division by zero
Should netc_switch_bpt_default_config() validate that priv->num_bp is
non-zero before proceeding?
quoted
+ int r = NETC_IPV_NUM % num_port_bp;
+ int num = q + r;
+
+ /* IPV-to-buffer-pool mapping per port:
[ ... ]
quoted
+static int netc_switch_bpt_default_config(struct netc_switch *priv)
+{
+ priv->num_bp = netc_get_buffer_pool_num(priv);
+ priv->bpt_list = devm_kcalloc(priv->dev, priv->num_bp,
+ sizeof(struct bpt_cfge_data),
+ GFP_KERNEL);
+ if (!priv->bpt_list)
+ return -ENOMEM;
+
+ /* Initialize the maximum threshold of each buffer pool entry */
+ for (int i = 0; i < priv->num_bp; i++) {
+ struct bpt_cfge_data *cfge = &priv->bpt_list[i];
+ int err;
+
+ cfge->max_thresh = cpu_to_le16(NETC_BP_THRESH);
+ err = ntmp_bpt_update_entry(&priv->ntmp, i, cfge);
+ if (err)
+ return err;
+ }
+
+ netc_ipv_to_buffer_pool_mapping(priv);
+
+ return 0;
+}
+
Should the return value of ntmp_bpt_update_entry() be checked here?
Simon Horman noted in v1 review that other NTMP table operations in the
driver propagate errors from the update functions. While this was fixed
in netc_switch_bpt_default_config() in v2, the same issue remains here.
If the hardware update fails (for example, DMA allocation failure or
hardware timeout), the driver's bpt_list will have flow control values that
do not match the actual hardware state, potentially causing incorrect flow
control behavior.
^^^^^^^^^^^^
Is the tag header guaranteed to be in the linear region of the skb?
Looking at other EtherType-based DSA tag drivers (tag_ar9331, tag_lan9303,
tag_dsa, tag_edsa, tag_a5psw, tag_brcm_legacy, tag_mxl-gsw1xx,
tag_gswip, and tag_mtk), they all call pskb_may_pull() before accessing
tag data to ensure the header is linearized.
+static struct netc_cbdr *netc_select_cbdr(struct ntmp_user *user)
+{
+ int cpu, i;
+
+ for (i = 0; i < user->cbdr_num; i++) {
+ if (spin_is_locked(&user->ring[i].ring_lock))
+ continue;
+
+ return &user->ring[i];
+ }
Frank Li raised a concern about the use of spin_is_locked() here. Is this
lockless check reliable across all architectures? spin_is_locked() behavior
varies across configurations - it can return false negatives on UP builds and
may not provide the ordering guarantees needed for lockless checks.
Would spin_trylock() be more appropriate here? Or if this racy check is
intentional and acceptable, should it be documented why the potential race is
safe in this context?
I missed this comment in v1 for some reasons, I will improve this, thanks