From: Or Gerlitz <hidden> Date: 2012-08-01 17:10:20
Changes from V1:
- applied feedback from Ben H.
- add stats64
- use memcpy on ethtool
- remove double indexing on ethtool and unused variables
- remove extra tab ...
- remove HW_VLAN_FILTER and the add/kill_vid ndo entries
- applied feedback from Dave Miller to avoid using sysfs
- added rtnl_link_ops support in ipoib and use them to add/delete childs
- added usage to ndo_add/remove_slave in eipoib
- added support for eipoib VIFs in rtnetlink and new ndo op to configure
that to the eipoib device
- some more little cleanups of unused variables
changes from V0:
- applied feedback from Eric/Dave - RX flow uses only the last 20 bytes of skb->cb[]
- applied feedback from Ben H. on ethtool changes
- fix sparse error on function which should be made static
- made the netdev features related code of the driver more elegant/robust
- used _bh locking in some paths which used plain rw locking in V0
- some code rearrangements in flows that send ARPs
The eIPoIB driver provides a standard Ethernet netdevice over
the InfiniBand IPoIB interface.
Some services can run only on top of Ethernet L2 interfaces, and cannot be
bound to an IPoIB interface. With this new driver, these services can run
seamlessly.
Main use case of the driver is the Ethernet Virtual Switching used in
virtualized environments, where an eipoib netdevice can be used as a
Physical Interface (PIF) in the hypervisor domain, and allow other
guests Virtual Interfaces (VIF) connected to the same Virtual Switch
to run over the InfiniBand fabric.
This driver supports L2 Switching (Direct Bridging) as well as other L3
Switching modes (e.g. NAT).
Whenever an IPoIB interface is created, one eIPoIB PIF netdevice
will be created. The default naming scheme is as in other Ethernet
interfaces: ethX, for example, on a system with two IPoIB interfaces,
ib0 and ib1, two interfaces will be created ethX and ethX+1 When "X"
is the next free Ethernet number in the system.
Using "ethtool -i " over the new interface can tell on which IPoIB
PIF interface that interface is above. For example: driver: eth_ipoib:ib0
indicates that eth3 is the Ethernet interface over the ib0 IPoIB interface.
The driver can be used as independent interface or to serve in
virtualization environment as the physical layer for the virtual
interfaces on the virtual guest.
The driver interface (eipoib interface or which is also referred to as parent)
uses slave interfaces, IPoIB clones, which are the VIFs described above.
VIFs interfaces are enslaved/released from the eipoib driver on demand, according
to the management interface provided to user space.
The management interface for the driver uses rtnl interface. Via these interfaces
the driver gets details on new VIF's to manage. The driver can
enslave new VIF (IPoIB cloned interface) or detaches from it.
The driver can also configure the slave to support specific mac/vlan of virtual guest.
Here's an example script on how the managment interface looks
like with the new rtnetlink code we added, this is based
on a patch to iproute2 which will be posted once the interface
is agreed.
#!/bin/bash
# create IPoIB clone (ib0.1), enslave it to eIPoIB (eth4), add VIF to
# serve the master, the provided MAC is the master's one
ip link add link ib0 name ib0.1 type ipoib index 1
ip link set dev ib0.1 master eth4
ip link set dev eth4 vif ib0.1 mac 00:02:C9:43:3B:F1
# add a bridge whose uplink is the eIPOIB device
brctl addbr br2
brctl addif br2 eth4
ifconfig br2 12.134.41.1/16 up
# create IPoIB clone (ib0.2), enslave it to eIPoIB (eth4), add VIF to
# serve a guest, the provided MAC is the guest's one, vnet1 is a hypervisor
# NIC (e.g tap) that serves that guest
ip link add link ib0 name ib0.2 type ipoib index 2
ip link set dev ib0.2 master eth4
ip link set dev eth4 vif ib0.2 mac 52:54:00:67:E9:BD
brctl addif br2 vnet1
# repeat the above with vlan
ip link add link ib0 name ib0.8003 type ipoib pkey 0x8003
ip link set dev ib0.8003 master eth4
ip link set dev eth4 vif ib0.8003 mac 00:02:C9:43:3B:F1 vlan 3
vconfig add eth4 3
ifconfig eth4.3 up
brctl addbr br3
brctl addif br3 eth4.3
ifconfig br3 13.134.41.1/16 up
ip link add link ib0 name ib0.8003.1 type ipoib pkey 0x8003 index 1
ip link set dev ib0.8003.1 master eth4
ip link set dev eth4 vif ib0.8003.1 mac 52:54:00:DE:3A:23 vlan 3
cat /sys/class/net/eth4/eth/vifs
============= END OF SCRIPT =======
Note: Each ethX interface has at least one ibX.Y slave to serve the PIF
itself, in the VIFs list of ethX you'll notice that ibX.1 is always created
to serve applications running from the Hypervisor on top of ethX interface directly.
For IB applications that require native IPoIB interfaces (e.g. RDMA-CM), the
original ipoib interfaces ibX can still be used. For example, RDMA-CM and
eth_ipoib drivers can co-exist and make use of IPoIB
The last patch of this series was made such that the series works as is over net-next.
In parallel to effort a patch to modify IPoIB such that it doesn't assume dst/neighbour
on the skb was pushed upstream, its commit b63b70d87741 "IPoIB: Use a private hash table
for path lookup in xmit path", once present in net-next this patch can simply be dropped.
Erez Shitrit (10):
include/linux: Add private flags for IPoIB interfaces
IB/ipoib: Add support for acting as VIF
net: Add ndo_set_vif_param operation to serve eIPoIB VIFs
net/core: Add rtnetlink support to vif parameters
net/eipoib: Add private header file
net/eipoib: Add ethtool file support
net/eipoib: Add main driver functionality
net/eipoib: Add sysfs support
net/eipoib: Add Makefile, Kconfig and MAINTAINERS entries
IB/ipoib: Add support for transmission of skbs w.o dst/neighbour
Or Gerlitz (2):
IB/ipoib: Add rtnl_link_ops support
IB/ipoib: Add support for clones / multiple childs on the same partition
Documentation/infiniband/ipoib.txt | 15 +
MAINTAINERS | 6 +
drivers/infiniband/ulp/ipoib/Makefile | 3 +-
drivers/infiniband/ulp/ipoib/ipoib.h | 16 +
drivers/infiniband/ulp/ipoib/ipoib_cm.c | 9 +
drivers/infiniband/ulp/ipoib/ipoib_ib.c | 8 +-
drivers/infiniband/ulp/ipoib/ipoib_main.c | 38 +-
drivers/infiniband/ulp/ipoib/ipoib_netlink.c | 119 ++
drivers/infiniband/ulp/ipoib/ipoib_vlan.c | 89 +-
drivers/net/Kconfig | 15 +
drivers/net/Makefile | 1 +
drivers/net/eipoib/Makefile | 4 +
drivers/net/eipoib/eth_ipoib.h | 215 +++
drivers/net/eipoib/eth_ipoib_ethtool.c | 114 ++
drivers/net/eipoib/eth_ipoib_main.c | 1953 ++++++++++++++++++++++++++
drivers/net/eipoib/eth_ipoib_sysfs.c | 435 ++++++
include/linux/if.h | 2 +
include/linux/if_link.h | 16 +
include/linux/netdevice.h | 5 +-
include/rdma/e_ipoib.h | 54 +
net/core/rtnetlink.c | 42 +-
21 files changed, 3117 insertions(+), 42 deletions(-)
create mode 100644 drivers/infiniband/ulp/ipoib/ipoib_netlink.c
create mode 100644 drivers/net/eipoib/Makefile
create mode 100644 drivers/net/eipoib/eth_ipoib.h
create mode 100644 drivers/net/eipoib/eth_ipoib_ethtool.c
create mode 100644 drivers/net/eipoib/eth_ipoib_main.c
create mode 100644 drivers/net/eipoib/eth_ipoib_sysfs.c
create mode 100644 include/rdma/e_ipoib.h
Cc: Erez Shitrit <redacted>
From: Or Gerlitz <hidden> Date: 2012-08-01 17:09:59
From: Erez Shitrit <redacted>
The new 2 bits indicates whenever a device is considered PIF interface,
which means the "main" interfaces (ib0, ib1 etc), or cloned interfaces
(ib0.1, ib1.2 etc.) that is now in use by the eIPoIB driver.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
include/linux/if.h | 2 ++
1 files changed, 2 insertions(+), 0 deletions(-)
From: Or Gerlitz <hidden> Date: 2012-08-01 17:09:59
From: Erez Shitrit <redacted>
When IPoIB interface acts as a VIF for an eIPoIB interface, it uses
the skb cb storage area on the RX flow, to place information which
can be of use to the upper layer device.
One such usage example, is when an eIPoIB inteface needs to generate
a source mac for incoming Ethernet frames.
The IPoIB code checks the VIF private flag on the RX path, and according
to the value of the flag prepares the skb CB data, etc.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
drivers/infiniband/ulp/ipoib/ipoib.h | 5 +++
drivers/infiniband/ulp/ipoib/ipoib_cm.c | 9 +++++
drivers/infiniband/ulp/ipoib/ipoib_ib.c | 8 ++++-
drivers/infiniband/ulp/ipoib/ipoib_main.c | 21 +++++++++++
include/rdma/e_ipoib.h | 54 +++++++++++++++++++++++++++++
5 files changed, 96 insertions(+), 1 deletions(-)
create mode 100644 include/rdma/e_ipoib.h
@@ -452,6 +453,10 @@ static int ipoib_cm_req_handler(struct ib_cm_id *cm_id, struct ib_cm_event *evencm_id->context=p;p->state=IPOIB_CM_RX_LIVE;p->jiffies=jiffies;++/* used to keep track of base qpn in CM mode */+p->qpn=be32_to_cpu(data->qpn);+INIT_LIST_HEAD(&p->list);p->qp=ipoib_cm_create_rx_qp(dev,p);
@@ -669,6 +674,10 @@ copied:skb->dev=dev;/* XXX get correct PACKET_ type here */skb->pkt_type=PACKET_HOST;+/* if handler is registered on top of ipoib, set skb oob data. */+if(skb->dev->priv_flags&IFF_EIPOIB_VIF)+set_skb_oob_cb_data(skb,wc,NULL);+netif_receive_skb(skb);repost:
@@ -304,7 +304,13 @@ static void ipoib_ib_handle_rx_wc(struct net_device *dev, struct ib_wc *wc)likely(wc->wc_flags&IB_WC_IP_CSUM_OK))skb->ip_summed=CHECKSUM_UNNECESSARY;-napi_gro_receive(&priv->napi,skb);+/* if handler is registered on top of ipoib, set skb oob data */+if(dev->priv_flags&IFF_EIPOIB_VIF){+set_skb_oob_cb_data(skb,wc,&priv->napi);+/* the registered handler will take care of the skb */+netif_receive_skb(skb);+}else+napi_gro_receive(&priv->napi,skb);repost:if(unlikely(ipoib_ib_post_receive(dev,wr_id)))
@@ -91,6 +91,24 @@ static struct ib_client ipoib_client = {.remove=ipoib_remove_one};+voidset_skb_oob_cb_data(structsk_buff*skb,structib_wc*wc,+structnapi_struct*napi)+{+structipoib_cm_rx*p_cm_ctx=NULL;+structeipoib_cb_data*data=NULL;++p_cm_ctx=wc->qp->qp_context;+data=IPOIB_HANDLER_CB(skb);++data->rx.slid=wc->slid;+data->rx.sqpn=wc->src_qp;+data->rx.napi=napi;++/* in CM mode, use the "base" qpn as sqpn */+if(p_cm_ctx)+data->rx.sqpn=p_cm_ctx->qpn;+}+intipoib_open(structnet_device*dev){structipoib_dev_priv*priv=netdev_priv(dev);
From: Or Gerlitz <hidden> Date: 2012-08-01 17:10:00
Allow creating "clone" child interfaces which further partition an
IPoIB interface to sub interfaces who either use the same pkey as
their parent or use the same pkey as already created child interface.
Each child now has a child index, which together with the pkey is
used as the identifier of the created network device.
Clone interfaces can only be created using rtnl_link_ops, where the user
is allowed to provide a pkey and index. If no pkey is provided, the
parent pkey is used and if no index is provided an index of zero
is used.
A major use case for clone childs is for virtualization purposes, where
a per VM NIC is desired at the hypervisor level, such as the solution
provided by the newly introduced Ethernet IPoIB driver.
Signed-off-by: Or Gerlitz <redacted>
Signed-off-by: Erez Shitrit <redacted>
---
Documentation/infiniband/ipoib.txt | 12 ++++++++++++
drivers/infiniband/ulp/ipoib/ipoib.h | 7 +++++--
drivers/infiniband/ulp/ipoib/ipoib_netlink.c | 21 +++++++++++++++------
drivers/infiniband/ulp/ipoib/ipoib_vlan.c | 19 +++++++++++--------
4 files changed, 43 insertions(+), 16 deletions(-)
@@ -27,6 +27,18 @@ Partitions and P_Keys Child interface create/delete can also be done using IPoIB's rtnl_link_ops, where childs created using either way behave the same.+Clones+ Its possible to further partition an IPoIB interfaces, and create+ "clone" child interfaces which either use the same pkey as their+ parent, or as an already created child interface. Each child now has+ a child index, which together with the pkey is used as the identifier+ of the created network device.++ Clone interfaces can only be created using rtnl_link_ops, where the user+ is allowed to provide a pkey and index. If no pkey is provided, the+ parent pkey is used and if no index is provided an index of zero+ is used.+ Datagram vs Connected modes The IPoIB driver supports two modes of operation: datagram and
@@ -52,7 +54,7 @@ static int ipoib_new_child_link(struct net *src_net, struct net_device *dev,{structnet_device*pdev;structipoib_dev_priv*ppriv;-u16child_pkey;+u16child_pkey,child_index;interr;if(!tb[IFLA_LINK])
@@ -64,12 +66,18 @@ static int ipoib_new_child_link(struct net *src_net, struct net_device *dev,ppriv=netdev_priv(pdev);if(!data||!data[IFLA_IPOIB_CHILD_PKEY]){-ipoib_warn(ppriv,"no pkey specified, failing request\n");-return-EINVAL;+ipoib_warn(ppriv,"no pkey specified, using parent pkey\n");+child_pkey=ppriv->pkey;}elsechild_pkey=nla_get_u16(data[IFLA_IPOIB_CHILD_PKEY]);-err=__ipoib_vlan_add(ppriv,netdev_priv(dev),child_pkey);+if(!data||!data[IFLA_IPOIB_CHILD_INDEX]){+ipoib_warn(ppriv,"no index specified, using 0\n");+child_index=0;+}else+child_index=nla_get_u16(data[IFLA_IPOIB_CHILD_INDEX]);++err=__ipoib_vlan_add(ppriv,netdev_priv(dev),child_pkey,child_index);returnerr;}
From: Or Gerlitz <hidden> Date: 2012-08-01 17:10:03
From: Erez Shitrit <redacted>
The Ethernet IPoIB driver enslaves IPoIB devices and uses them as
VIFs (Virtual Interface) which serve an Ethernet NIC e.g present in a
guest OS. For each such slave that acts as a VIF, eIPoIB needs to know
the mac and optionally the vlan uses by that NIC, the new ndo opertaion
is used to associate the mac/vlan for that slave.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
include/linux/netdevice.h | 5 ++++-
1 files changed, 4 insertions(+), 1 deletions(-)
From: Or Gerlitz <hidden> Date: 2012-08-01 17:10:04
Add rtnl_link_ops to IPoIB, with the first usage being child device
create/delete through them. For that end, did little refactoring
of the ipoib_vlan_add/delete code which is now used by both the
sysfs and the rtnl_link_ops code.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
Documentation/infiniband/ipoib.txt | 3 +
drivers/infiniband/ulp/ipoib/Makefile | 3 +-
drivers/infiniband/ulp/ipoib/ipoib.h | 8 ++
drivers/infiniband/ulp/ipoib/ipoib_main.c | 10 ++-
drivers/infiniband/ulp/ipoib/ipoib_netlink.c | 110 ++++++++++++++++++++++++++
drivers/infiniband/ulp/ipoib/ipoib_vlan.c | 80 ++++++++++++-------
6 files changed, 181 insertions(+), 33 deletions(-)
create mode 100644 drivers/infiniband/ulp/ipoib/ipoib_netlink.c
@@ -24,6 +24,9 @@ Partitions and P_Keys The P_Key for any interface is given by the "pkey" file, and the main interface for a subinterface is in "parent."+ Child interface create/delete can also be done using IPoIB's+ rtnl_link_ops, where childs created using either way behave the same.+ Datagram vs Connected modes The IPoIB driver supports two modes of operation: datagram and
@@ -70,26 +64,16 @@ int ipoib_vlan_add(struct net_device *pdev, unsigned short pkey)*/if(ppriv->pkey==pkey){result=-ENOTUNIQ;-priv=NULL;gotoerr;}-list_for_each_entry(priv,&ppriv->child_intfs,list){-if(priv->pkey==pkey){+list_for_each_entry(tpriv,&ppriv->child_intfs,list){+if(tpriv->pkey==pkey){result=-ENOTUNIQ;-priv=NULL;gotoerr;}}-snprintf(intf_name,sizeofintf_name,"%s.%04x",-ppriv->dev->name,pkey);-priv=ipoib_intf_alloc(intf_name);-if(!priv){-result=-ENOMEM;-gotoerr;-}-priv->max_ib_mtu=ppriv->max_ib_mtu;/* MTU will be reset when mcast join happens */priv->dev->mtu=IPOIB_UD_MTU(priv->max_ib_mtu);
@@ -137,11 +121,10 @@ int ipoib_vlan_add(struct net_device *pdev, unsigned short pkey)list_add_tail(&priv->list,&ppriv->child_intfs);mutex_unlock(&ppriv->vlan_mutex);-rtnl_unlock();-return0;sysfs_failed:+result=-ENOMEM;ipoib_delete_debug_files(priv->dev);unregister_netdevice(priv->dev);
From: Or Gerlitz <hidden> Date: 2012-08-01 17:10:11
From: Erez Shitrit <redacted>
Via ethtool the driver describes its version, ABI version, on what PIF
interface it runs and various statistics.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
drivers/net/eipoib/eth_ipoib_ethtool.c | 114 ++++++++++++++++++++++++++++++++
1 files changed, 114 insertions(+), 0 deletions(-)
create mode 100644 drivers/net/eipoib/eth_ipoib_ethtool.c
From: Or Gerlitz <hidden> Date: 2012-08-01 17:10:11
From: Erez Shitrit <redacted>
The header file includes all structures, macros and non-static
functions which are of use by the driver.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
drivers/net/eipoib/eth_ipoib.h | 215 ++++++++++++++++++++++++++++++++++++++++
1 files changed, 215 insertions(+), 0 deletions(-)
create mode 100644 drivers/net/eipoib/eth_ipoib.h
@@ -0,0 +1,215 @@+/*+*Copyright(c)2012MellanoxTechnologies.Allrightsreserved+*+*Thissoftwareisavailabletoyouunderachoiceofoneoftwo+*licenses.YoumaychoosetobelicensedunderthetermsoftheGNU+*GeneralPublicLicense(GPL)Version2,availablefromthefile+*COPYINGinthemaindirectoryofthissourcetree,orthe+*openfabric.orgBSDlicensebelow:+*+*Redistributionanduseinsourceandbinaryforms,withor+*withoutmodification,arepermittedprovidedthatthefollowing+*conditionsaremet:+*+*-Redistributionsofsourcecodemustretaintheabove+*copyrightnotice,thislistofconditionsandthefollowing+*disclaimer.+*+*-Redistributionsinbinaryformmustreproducetheabove+*copyrightnotice,thislistofconditionsandthefollowing+*disclaimerinthedocumentationand/orothermaterials+*providedwiththedistribution.+*+*THESOFTWAREISPROVIDED"AS IS",WITHOUTWARRANTYOFANYKIND,+*EXPRESSORIMPLIED,INCLUDINGBUTNOTLIMITEDTOTHEWARRANTIESOF+*MERCHANTABILITY,FITNESSFORAPARTICULARPURPOSEAND+*NONINFRINGEMENT.INNOEVENTSHALLTHEAUTHORSORCOPYRIGHTHOLDERS+*BELIABLEFORANYCLAIM,DAMAGESOROTHERLIABILITY,WHETHERINAN+*ACTIONOFCONTRACT,TORTOROTHERWISE,ARISINGFROM,OUTOFORIN+*CONNECTIONWITHTHESOFTWAREORTHEUSEOROTHERDEALINGSINTHE+*SOFTWARE.+*/++#ifndef _LINUX_ETH_IPOIB_H+#define _LINUX_ETH_IPOIB_H++#include<linux/module.h>+#include<linux/errno.h>+#include<linux/netdevice.h>+#include<linux/skbuff.h>+#include<net/arp.h>+#include<linux/if_vlan.h>+#include<net/net_namespace.h>+#include<net/netns/generic.h>+#include<linux/if_infiniband.h>+#include<rdma/ib_verbs.h>++#include<rdma/e_ipoib.h>++/* macros and definitions */+#define DRV_VERSION "1.0.0"+#define DRV_RELDATE "June 1, 2012"+#define DRV_NAME "eth_ipoib"+#define SDRV_NAME "ipoib"+#define DRV_DESCRIPTION "IP-over-InfiniBand Para Virtualized Driver"+#define EIPOIB_ABI_VER 1++#undef pr_fmt+#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt++#define GID_LEN 16+#define GUID_LEN 8++#define PARENT_VLAN_FEATURES (NETIF_F_HW_VLAN_RX | NETIF_F_HW_VLAN_TX)++#define parent_for_each_slave(_parent, slave) \+list_for_each_entry(slave,&(_parent)->slave_list,list)\++#define PARENT_IS_OK(_parent) \+(((_parent)->dev->flags&IFF_UP)&&\+netif_running((_parent)->dev)&&\+((_parent)->slave_cnt>0))++#define IS_E_IPOIB_PROTO(_proto) \+(((_proto)==htons(ETH_P_ARP))||\+((_proto)==htons(ETH_P_RARP))||\+((_proto)==htons(ETH_P_IP)))++enumeipoib_emac_guest_info{+VALID,+MIGRATED_OUT,+INVALID,+};++/* structs */+structeth_arp_data{+u8arp_sha[ETH_ALEN];+__be32arp_sip;+u8arp_dha[ETH_ALEN];+__be32arp_dip;+}__packed;++structipoib_arp_data{+u8arp_sha[INFINIBAND_ALEN];+__be32arp_sip;+u8arp_dha[INFINIBAND_ALEN];+__be32arp_dip;+}__packed;++/* live migration support structures: */+structip_member{+__be32ip;+structlist_headlist;+};++/*+*foreachslave(emac)savesalltheipoverthatmac.+*theparentkeepsthatlistforlivemigration.+*/+structguest_emac_info{+u8emac[ETH_ALEN];+u16vlan;+structlist_headip_list;+structlist_headlist;+enumeipoib_emac_guest_inforec_state;+intnum_of_retries;+};++structneigh{+structlist_headlist;+u8emac[ETH_ALEN];+u8imac[INFINIBAND_ALEN];+/* this part is used for neigh_add_list */+charcmd[PAGE_SIZE];+};++structslave{+structnet_device*dev;+structlist_headlist;+u16pkey;+u16vlan;+u8emac[ETH_ALEN];+u8imac[INFINIBAND_ALEN];+structlist_headneigh_list;+};++structport_stats{+/* update PORT_STATS_LEN (number of stat fields)accordingly */+uint64_ttx_parent_dropped;+uint64_ttx_vif_miss;+uint64_ttx_neigh_miss;+uint64_ttx_vlan;+uint64_ttx_shared;+uint64_ttx_proto_errors;+uint64_ttx_skb_errors;+uint64_ttx_slave_err;++uint64_trx_parent_dropped;+uint64_trx_vif_miss;+uint64_trx_neigh_miss;+uint64_trx_vlan;+uint64_trx_shared;+uint64_trx_proto_errors;+uint64_trx_skb_errors;+uint64_trx_slave_err;+};++structparent{+structnet_device*dev;+intindex;+structneigh_parmsnparms;+structlist_headslave_list;+/* never change this value outside the attach/detach wrappers */+s32slave_cnt;+rwlock_tlock;+structport_statsport_stats;+structlist_headparent_list;+u16flags;+structworkqueue_struct*wq;+s8kill_timers;+structdelayed_workneigh_learn_work;+structdelayed_workvif_learn_work;+structlist_headneigh_add_list;+unionib_gidgid;+charipoib_main_interface[IFNAMSIZ];+/* live migration and bonding support */+structlist_heademac_ip_list;+structdelayed_workemac_ip_work;+structdelayed_workmigrate_out_work;+};++#define eipoib_slave_get_rcu(dev) \+((structslave*)rcu_dereference(dev->rx_handler_data))++/* name space support for sys/fs */+structeipoib_net{+structnet*net;/* Associated network namespace */+structclass_attributeclass_attr_eipoib_interfaces;+};++/* exported from main.c */+externinteipoib_net_id;+externstructlist_headparent_dev_list;++/* functions prototypes */+intmod_create_sysfs(structeipoib_net*eipoib_n);+voidmod_destroy_sysfs(structeipoib_net*eipoib_n);+voidparent_destroy_sysfs_entry(structparent*parent);+intparent_create_sysfs_entry(structparent*parent);+intcreate_slave_symlinks(structnet_device*master,+structnet_device*slave);+voiddestroy_slave_symlinks(structnet_device*master,+structnet_device*slave);+intparent_enslave(structnet_device*parent_dev,+structnet_device*slave_dev);+intparent_release_slave(structnet_device*parent_dev,+structnet_device*slave_dev);+structneigh*parent_get_neigh_cmd(charop,char*ifname,+u8*remac,u8*rimac);+structslave*parent_get_vif_cmd(charop,char*ifname,u8*lemac);+ssize_t__parent_store_neighs(structdevice*d,+structdevice_attribute*attr,+constchar*buffer,size_tcount);+voidparent_set_ethtool_ops(structnet_device*dev);++#endif /* _LINUX_ETH_IPOIB_H */
From: Or Gerlitz <hidden> Date: 2012-08-01 17:10:25
From: Erez Shitrit <redacted>
Guest IP packets sent by eIPoIB over an IPoIB VIF interface do not point
to dst or neighbour. This patch modifies IPoIB such that trasnmission
of such packets is possible. It does so by extending an already existing
path in the driver which was used so far only for unicast ARP probes.
This patch was made such that the series works as is over net-next, in
parallel to this driver a patch to modify IPoIB such that it doesn't
assume dst/neighbour on the skb was pushed upstream, its commit b63b70d87741
"IPoIB: Use a private hash table for path lookup in xmit path", once
present in net-next this patch can simply be dropped.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
drivers/infiniband/ulp/ipoib/ipoib_main.c | 7 ++++---
1 files changed, 4 insertions(+), 3 deletions(-)
@@ -800,7 +800,8 @@ static int ipoib_start_xmit(struct sk_buff *skb, struct net_device *dev)/* unicast GID -- should be ARP or RARP reply */if((be16_to_cpup((__be16*)skb->data)!=ETH_P_ARP)&&-(be16_to_cpup((__be16*)skb->data)!=ETH_P_RARP)){+(be16_to_cpup((__be16*)skb->data)!=ETH_P_RARP)&&+!(dev->priv_flags&IFF_EIPOIB_VIF)){ipoib_warn(priv,"Unicast, no %s: type %04x, QPN %06x %pI6\n",skb_dst(skb)?"neigh":"dst",be16_to_cpup((__be16*)skb->data),
@@ -850,7 +851,7 @@ static int ipoib_hard_header(struct sk_buff *skb,*destinationaddressintoskb->cbsowecanfigureoutwhere*tosendthepacketlater.*/-if(!skb_dst(skb)){+if(!skb_dst(skb)||dev->priv_flags&IFF_EIPOIB_VIF){structipoib_cb*cb=(structipoib_cb*)skb->cb;memcpy(cb->hwaddr,daddr,INFINIBAND_ALEN);}
From: Or Gerlitz <hidden> Date: 2012-08-01 17:10:26
From: Erez Shitrit <redacted>
The eipoib driver provides a standard Ethernet netdevice over
the InfiniBand IPoIB interface .
Some services can run only on top of Ethernet L2 interfaces, and cannot be
bound to an IPoIB interface. With this new driver, these services can run
seamlessly.
Main use case of the driver is the Ethernet Virtual Switching used in
virtualized environments, where an eipoib netdevice can be used as a
Physical Interface (PIF) in the hypervisor domain, and allow other
guests Virtual Interfaces (VIF) connected to the same Virtual Switch
to run over the InfiniBand fabric.
This driver supports L2 Switching (Direct Bridging) as well as other L3
Switching modes (e.g. NAT).
Whenever an IPoIB interface is created, one eIPoIB PIF netdevice
will be created. The default naming scheme is as in other Ethernet
interfaces: ethX, for example, on a system with two IPoIB interfaces,
ib0 and ib1, two interfaces will be created ethX and ethX+1 When "X"
is the next free Ethernet number in the system.
Using "ethtool -i " over the new interface can tell on which IPoIB
PIF interface that interface is above. For example: driver: eth_ipoib:ib0
indicates that eth3 is the Ethernet interface over the ib0 IPoIB interface.
The driver can be used as independent interface or to serve in
virtualization environment as the physical layer for the virtual
interfaces on the virtual guest.
The driver interface (eipoib interface or which is also referred to as parent)
uses slave interfaces, IPoIB clones, which are the VIFs described above.
VIFs interfaces are enslaved/released from the eipoib driver on demand, according
to the management interface provided to user space.
Note: Each ethX interface has at least one ibX.Y slave to serve the PIF
itself, in the VIFs list of ethX you'll notice that ibX.1 is always created
to serve applications running from the Hypervisor on top of ethX interface directly.
For IB applications that require native IPoIB interfaces (e.g. RDMA-CM), the
original ipoib interfaces ibX can still be used. For example, RDMA-CM and
eth_ipoib drivers can co-exist and make use of IPoIB
Support for Live migration:
The driver expose sysfs interface through which the manager can notify on
migrated out guests ("detached VIF"). The abandoned eIPoIB driver instance
sends ARP requests (defined number) to the network in order to triger the migrated
guest to publish its mac address of the new VIF that hosts it. When the guest
responds to that ARP, the eIPoIB driver on the host that owns that guest, sends
Gratuitous ARP in behalf of that guest such that all the peers on the network
which communicate with this VM are noticed on the new VIF address which is
used by the guest.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
drivers/net/eipoib/eth_ipoib_main.c | 1953 +++++++++++++++++++++++++++++++++++
1 files changed, 1953 insertions(+), 0 deletions(-)
create mode 100644 drivers/net/eipoib/eth_ipoib_main.c
@@ -0,0 +1,1953 @@+/*+*Copyright(c)2012MellanoxTechnologies.Allrightsreserved+*+*Thissoftwareisavailabletoyouunderachoiceofoneoftwo+*licenses.YoumaychoosetobelicensedunderthetermsoftheGNU+*GeneralPublicLicense(GPL)Version2,availablefromthefile+*COPYINGinthemaindirectoryofthissourcetree,orthe+*openfabric.orgBSDlicensebelow:+*+*Redistributionanduseinsourceandbinaryforms,withor+*withoutmodification,arepermittedprovidedthatthefollowing+*conditionsaremet:+*+*-Redistributionsofsourcecodemustretaintheabove+*copyrightnotice,thislistofconditionsandthefollowing+*disclaimer.+*+*-Redistributionsinbinaryformmustreproducetheabove+*copyrightnotice,thislistofconditionsandthefollowing+*disclaimerinthedocumentationand/orothermaterials+*providedwiththedistribution.+*+*THESOFTWAREISPROVIDED"AS IS",WITHOUTWARRANTYOFANYKIND,+*EXPRESSORIMPLIED,INCLUDINGBUTNOTLIMITEDTOTHEWARRANTIESOF+*MERCHANTABILITY,FITNESSFORAPARTICULARPURPOSEAND+*NONINFRINGEMENT.INNOEVENTSHALLTHEAUTHORSORCOPYRIGHTHOLDERS+*BELIABLEFORANYCLAIM,DAMAGESOROTHERLIABILITY,WHETHERINAN+*ACTIONOFCONTRACT,TORTOROTHERWISE,ARISINGFROM,OUTOFORIN+*CONNECTIONWITHTHESOFTWAREORTHEUSEOROTHERDEALINGSINTHE+*SOFTWARE.+*/++#include"eth_ipoib.h"+#include<net/ip.h>+#include<linux/if_link.h>++#define EMAC_IP_GC_TIME (10 * HZ)++#define MIG_OUT_ARP_REQ_ISSUE_TIME (0.5 * HZ)++#define MIG_OUT_MAX_ARP_RETRIES 5++#define LIVE_MIG_PACKET 1++#define PARENT_MAC_MASK 0xe7++/* forward declaration */+staticrx_handler_result_teipoib_handle_frame(structsk_buff**pskb);+staticinteipoib_device_event(structnotifier_block*unused,+unsignedlongevent,void*ptr);+staticvoidfree_ip_mem_in_rec(structguest_emac_info*emac_info);++staticconstchar*constversion=+DRV_DESCRIPTION": v"DRV_VERSION" ("DRV_RELDATE")\n";++LIST_HEAD(parent_dev_list);++/* name space sys/fs functions */+inteipoib_net_id__read_mostly;++staticint__net_initeipoib_net_init(structnet*net)+{+intrc;+structeipoib_net*eipoib_n=net_generic(net,eipoib_net_id);++eipoib_n->net=net;+rc=mod_create_sysfs(eipoib_n);++returnrc;+}++staticvoid__net_exiteipoib_net_exit(structnet*net)+{+structeipoib_net*eipoib_n=net_generic(net,eipoib_net_id);++mod_destroy_sysfs(eipoib_n);+}++staticstructpernet_operationseipoib_net_ops={+.init=eipoib_net_init,+.exit=eipoib_net_exit,+.id=&eipoib_net_id,+.size=sizeof(structeipoib_net),+};++/* set mac fields emac=<qpn><lid> */+staticinline+voidbuild_neigh_mac(u8*_mac,u32_qpn,u16_lid)+{+/* _qpn: 3B _lid: 2B */+*((__be32*)(_mac))=cpu_to_be32(_qpn);+*(u8*)(_mac)=0x2;/* set LG bit */+*(__be16*)(_mac+sizeof(_qpn))=cpu_to_be16(_lid);+}++staticinline+structslave*get_slave_by_dev(structparent*parent,+structnet_device*slave_dev)+{+structslave*slave,*slave_tmp;+intfound=0;++parent_for_each_slave(parent,slave_tmp){+if(slave_tmp->dev==slave_dev){+found=1;+slave=slave_tmp;+break;+}+}++returnfound?slave:NULL;+}++staticinline+structslave*get_slave_by_mac_and_vlan(structparent*parent,u8*mac,+u16vlan)+{+structslave*slave,*slave_tmp;+intfound=0;++parent_for_each_slave(parent,slave_tmp){+if((!memcmp(slave_tmp->emac,mac,ETH_ALEN))&&+(slave_tmp->vlan==vlan)){+found=1;+slave=slave_tmp;+break;+}+}++returnfound?slave:NULL;+}+++staticinline+structguest_emac_info*get_mac_ip_info_by_mac_and_vlan(structparent*parent,+u8*mac,u16vlan)+{+structguest_emac_info*emac_info,*emac_info_ret;+intfound=0;++list_for_each_entry(emac_info,&parent->emac_ip_list,list){+if((!memcmp(emac_info->emac,mac,ETH_ALEN))&&+vlan==emac_info->vlan){+found=1;+emac_info_ret=emac_info;+break;+}+}++returnfound?emac_info_ret:NULL;+}++/*+*searchesfortherelevantguest_emac_infointheparent.+*iffoundit,checkifitcontainstherequiredip+*ifnosuchguest_emac_infoobjectornoipreturn0,+*otherwisereturn1andifexistsettheguest_emac_infoobj.+*/+staticinline+intis_mac_info_contain_ip(structparent*parent,u8*mac,__be32ip,+structguest_emac_info*emac_info,u16vlan)+{+structip_member*ipm;+intfound=0;++emac_info=get_mac_ip_info_by_mac_and_vlan(parent,mac,vlan);+if(!emac_info)+return0;++list_for_each_entry(ipm,&emac_info->ip_list,list){+if(ipm->ip==ip){+found=1;+break;+}+}++returnfound;+}++staticinlineintnetdev_set_parent_master(structnet_device*slave,+structnet_device*master)+{+interr;++ASSERT_RTNL();++err=netdev_set_master(slave,master);+if(err)+returnerr;+if(master){+slave->priv_flags|=IFF_EIPOIB_VIF;+/* deny bonding from enslaving it. */;+slave->flags|=IFF_SLAVE;+}else{+slave->priv_flags&=~(IFF_EIPOIB_VIF);+slave->flags&=~(IFF_SLAVE);+}++return0;+}++staticinlineintis_driver_owner(structnet_device*dev,char*name)+{+structethtool_drvinfodrvinfo;++if(dev->ethtool_ops&&dev->ethtool_ops->get_drvinfo){+memset(&drvinfo,0,sizeof(drvinfo));+dev->ethtool_ops->get_drvinfo(dev,&drvinfo);+if(!strstr(drvinfo.driver,name))+return0;+}else+return0;++return1;+}++staticinlineintis_parent(structnet_device*dev)+{+return(dev->priv_flags&IFF_EIPOIB_PIF)&&+is_driver_owner(dev,DRV_NAME);+}++staticinlineintis_parent_mac(structnet_device*dev,u8*mac)+{+returnis_parent(dev)&&!memcmp(mac,dev->dev_addr,dev->addr_len);+}++staticinlineint__is_slave(structnet_device*dev)+{+returndev->master&&is_parent(dev->master);+}++staticinlineintis_slave(structnet_device*dev)+{+return(dev->priv_flags&IFF_EIPOIB_VIF)&&+is_driver_owner(dev,SDRV_NAME)&&__is_slave(dev);+}++/*+*-------------------------------Linkstatus------------------+*setparentcarrier:+*linkisupifatleastoneslavehaslinkup+*otherwise,bringlinkdown+*return1ifparentcarrierchanged,zerootherwise+*/+staticintparent_set_carrier(structparent*parent)+{+structslave*slave;++if(parent->slave_cnt==0)+gotodown;++/* bring parent link up if one slave (at least) is up */+parent_for_each_slave(parent,slave){+if(netif_carrier_ok(slave->dev)){+if(!netif_carrier_ok(parent->dev)){+netif_carrier_on(parent->dev);+return1;+}+return0;+}+}++down:+if(netif_carrier_ok(parent->dev)){+pr_info("bring down carrier\n");+netif_carrier_off(parent->dev);+return1;+}+return0;+}++staticintparent_set_mtu(structparent*parent)+{+structslave*slave,*f_slave;+unsignedintmtu;++if(parent->slave_cnt==0)+return0;++/* find min mtu */+f_slave=list_first_entry(&parent->slave_list,structslave,list);+mtu=f_slave->dev->mtu;++parent_for_each_slave(parent,slave)+mtu=min(slave->dev->mtu,mtu);++if(parent->dev->mtu!=mtu){+dev_set_mtu(parent->dev,mtu);+return1;+}++return0;+}++/*+*---------------------------slavelisthandling------+*+*Thisfunctionattachestheslavetotheendoflist.+*payattention,thecallershouldheldparen->lock+*/+staticvoidparent_attach_slave(structparent*parent,+structslave*new_slave)+{+list_add_tail(&new_slave->list,&parent->slave_list);+parent->slave_cnt++;+}++staticvoidparent_detach_slave(structparent*parent,structslave*slave)+{+list_del(&slave->list);+parent->slave_cnt--;+}++staticnetdev_features_tparent_fix_features(structnet_device*dev,+netdev_features_tfeatures)+{+structslave*slave;+structparent*parent=netdev_priv(dev);+netdev_features_tmask;++read_lock_bh(&parent->lock);++mask=features;+features&=~NETIF_F_ONE_FOR_ALL;+features|=NETIF_F_ALL_FOR_ALL;++parent_for_each_slave(parent,slave)+features=netdev_increment_features(features,+slave->dev->features,+mask);++features&=~NETIF_F_VLAN_CHALLENGED;+read_unlock_bh(&parent->lock);+returnfeatures;+}++staticintparent_compute_features(structparent*parent)+{+structnet_device*parent_dev=parent->dev;+u64hw_features,features;+structslave*slave;++if(list_empty(&parent->slave_list))+gotodone;++/* starts with the max set of features mask */+hw_features=features=~0LL;++/* gets the common features from all slaves */+parent_for_each_slave(parent,slave){+features&=slave->dev->features;+hw_features&=slave->dev->hw_features;+}++features=features|PARENT_VLAN_FEATURES;+hw_features=hw_features|PARENT_VLAN_FEATURES;++hw_features&=~NETIF_F_VLAN_CHALLENGED;+features&=hw_features;++parent_dev->hw_features=hw_features;+parent_dev->features=features;+parent_dev->vlan_features=parent_dev->features&~PARENT_VLAN_FEATURES;+done:+pr_info("%s: %s: Features: 0x%llx\n",+__func__,parent_dev->name,parent_dev->features);++return0;+}++staticinlineu16slave_get_pkey(structnet_device*dev)+{+u16pkey=(dev->broadcast[8]<<8)+dev->broadcast[9];++returnpkey;+}++staticvoidparent_setup_by_slave(structnet_device*parent_dev,+structnet_device*slave_dev)+{+structparent*parent=netdev_priv(parent_dev);+conststructnet_device_ops*slave_ops=slave_dev->netdev_ops;++parent_dev->mtu=slave_dev->mtu;+parent_dev->hard_header_len=slave_dev->hard_header_len;++slave_ops->ndo_neigh_setup(slave_dev,&parent->nparms);++}++/* enslave device <slave> to parent device <master> */+intparent_enslave(structnet_device*parent_dev,structnet_device*slave_dev)+{+structparent*parent=netdev_priv(parent_dev);+structslave*new_slave=NULL;+intold_features=parent_dev->features;+intres=0;+/* slave must be claimed by ipoib */+if(!is_driver_owner(slave_dev,SDRV_NAME))+return-EOPNOTSUPP;++/* parent must be initialized by parent_open() before enslaving */+if(!(parent_dev->flags&IFF_UP)){+pr_warn("%s parent is not up in "+"parent_enslave\n",+parent_dev->name);+return-EPERM;+}++/* already enslaved */+if((slave_dev->flags&IFF_SLAVE)||+(slave_dev->priv_flags&IFF_EIPOIB_VIF)){+pr_err("%s was already enslaved!!!\n",slave_dev->name);+return-EBUSY;+}++/* mark it as ipoib clone vif */+slave_dev->priv_flags|=IFF_EIPOIB_VIF;++/* set parent netdev attributes */+if(parent->slave_cnt==0)+parent_setup_by_slave(parent_dev,slave_dev);+else{+/* check netdev attr match */+if(slave_dev->hard_header_len!=parent_dev->hard_header_len){+pr_err("%s slave %s has different HDR len %d != %d\n",+parent_dev->name,slave_dev->name,+slave_dev->hard_header_len,+parent_dev->hard_header_len);+res=-EINVAL;+gotoerr_undo_flags;+}++if(slave_dev->type!=ARPHRD_INFINIBAND||+slave_dev->addr_len!=INFINIBAND_ALEN){+pr_err("%s slave type/addr_len is invalid (%d/%d)\n",+parent_dev->name,slave_dev->type,+slave_dev->addr_len);+res=-EINVAL;+gotoerr_undo_flags;+}+}+/*+*verfiythatthis(slave)devicebelongstotherelevantPIF+*abortifthenameoftheslaveisnotastheregularwayinipoib+*/+if(!strstr(slave_dev->name,parent->ipoib_main_interface)){+pr_err("%s slave name (%s) doesn't contain parent name (%s) ",+parent_dev->name,slave_dev->name,+parent->ipoib_main_interface);+res=-EINVAL;+gotoerr_undo_flags;+}++new_slave=kzalloc(sizeof(structslave),GFP_KERNEL);+if(!new_slave){+res=-ENOMEM;+gotoerr_undo_flags;+}++INIT_LIST_HEAD(&new_slave->neigh_list);++/* save slave's vlan */+new_slave->pkey=slave_get_pkey(slave_dev);++res=netdev_set_parent_master(slave_dev,parent_dev);+if(res){+pr_err("%s %d calling netdev_set_master\n",+slave_dev->name,res);+gotoerr_free;+}++res=dev_open(slave_dev);+if(res){+pr_info("open failed %s\n",+slave_dev->name);+gotoerr_unset_master;+}++new_slave->dev=slave_dev;++write_lock_bh(&parent->lock);++parent_attach_slave(parent,new_slave);++parent_compute_features(parent);++write_unlock_bh(&parent->lock);++read_lock_bh(&parent->lock);++parent_set_carrier(parent);++read_unlock_bh(&parent->lock);++res=create_slave_symlinks(parent_dev,slave_dev);+if(res)+gotoerr_close;++/* register handler */+res=netdev_rx_handler_register(slave_dev,eipoib_handle_frame,+new_slave);+if(res){+pr_warn("%s %d calling netdev_rx_handler_register\n",+parent_dev->name,res);+gotoerr_close;+}++pr_info("%s: enslaving %s\n",parent_dev->name,slave_dev->name);++/* enslave is successful */+return0;++/* Undo stages on error */+err_close:+dev_close(slave_dev);++err_unset_master:+netdev_set_parent_master(slave_dev,NULL);++err_free:+kfree(new_slave);++err_undo_flags:+parent_dev->features=old_features;++returnres;+}++staticvoidslave_free(structparent*parent,structslave*slave)+{+structneigh*neigh,*neigh_tmp;++list_for_each_entry_safe(neigh,neigh_tmp,&slave->neigh_list,list){+list_del(&neigh->list);+kfree(neigh);+}++netdev_rx_handler_unregister(slave->dev);++kfree(slave);+}++intparent_release_slave(structnet_device*parent_dev,+structnet_device*slave_dev)+{+structparent*parent=netdev_priv(parent_dev);+structslave*slave;+structguest_emac_info*emac_info;++/* slave is not a slave or master is not master of this slave */+if(!(slave_dev->flags&IFF_SLAVE)||+(slave_dev->master!=parent_dev)){+pr_err("%s cannot release %s.\n",+parent_dev->name,slave_dev->name);+return-EINVAL;+}++write_lock_bh(&parent->lock);++slave=get_slave_by_dev(parent,slave_dev);+if(!slave){+/* not a slave of this parent */+pr_warn("%s not enslaved %s\n",+parent_dev->name,slave_dev->name);+write_unlock_bh(&parent->lock);+return-EINVAL;+}++pr_info("%s: releasing interface %s\n",parent_dev->name,+slave_dev->name);++/* for live migration, mark its mac_ip record as invalid */+emac_info=get_mac_ip_info_by_mac_and_vlan(parent,slave->emac,slave->vlan);+if(!emac_info)+pr_warn("%s %s didn't find emac: %pM\n",+parent_dev->name,slave_dev->name,slave->emac);+else{+emac_info->rec_state=MIGRATED_OUT;+/* start GC work */+pr_info("%s: sending clean task for slave mac: %pM\n",+__func__,slave->emac);+queue_delayed_work(parent->wq,&parent->migrate_out_work,0);+queue_delayed_work(parent->wq,&parent->emac_ip_work,+EMAC_IP_GC_TIME);+}++/* release the slave from its parent */+parent_detach_slave(parent,slave);++parent_compute_features(parent);++if(parent->slave_cnt==0)+parent_set_carrier(parent);++write_unlock_bh(&parent->lock);++/* must do this from outside any spinlocks */+destroy_slave_symlinks(parent_dev,slave_dev);++netdev_set_parent_master(slave_dev,NULL);++dev_close(slave_dev);++slave_free(parent,slave);++return0;/* deletion OK */+}++staticintparent_release_all(structnet_device*parent_dev)+{+structparent*parent=netdev_priv(parent_dev);+structslave*slave,*slave_tmp;+structnet_device*slave_dev;+structneigh*neigh_cmd,*neigh_cmd_tmp;+structguest_emac_info*emac_info,*emac_info_tmp;+structslave;++write_lock_bh(&parent->lock);++netif_carrier_off(parent_dev);++if(parent->slave_cnt==0)+gotoout;++list_for_each_entry_safe(slave,slave_tmp,&parent->slave_list,list){+slave_dev=slave->dev;++/* remove slave from parent's slave-list */+parent_detach_slave(parent,slave);++parent_compute_features(parent);++write_unlock_bh(&parent->lock);++destroy_slave_symlinks(parent_dev,slave_dev);++netdev_set_parent_master(slave_dev,NULL);++dev_close(slave_dev);++slave_free(parent,slave);++write_lock_bh(&parent->lock);+}++list_for_each_entry_safe(neigh_cmd,neigh_cmd_tmp,+&parent->neigh_add_list,list){+list_del(&neigh_cmd->list);+kfree(neigh_cmd);+}++list_for_each_entry_safe(emac_info,emac_info_tmp,+&parent->emac_ip_list,list){+free_ip_mem_in_rec(emac_info);+list_del(&emac_info->list);+kfree(emac_info);+}++pr_info("%s: released all slaves\n",parent_dev->name);++out:+write_unlock_bh(&parent->lock);++return0;+}++/* -------------------------- Device entry points --------------------------- */+staticstructrtnl_link_stats64*parent_get_stats(structnet_device*parent_dev,+structrtnl_link_stats64*stats)+{+structparent*parent=netdev_priv(parent_dev);+structslave*slave;+structrtnl_link_stats64temp;++memset(stats,0,sizeof(*stats));++read_lock_bh(&parent->lock);++parent_for_each_slave(parent,slave){+conststructrtnl_link_stats64*sstats=+dev_get_stats(slave->dev,&temp);++stats->rx_packets+=sstats->rx_packets;+stats->rx_bytes+=sstats->rx_bytes;+stats->rx_errors+=sstats->rx_errors;+stats->rx_dropped+=sstats->rx_dropped;++stats->tx_packets+=sstats->tx_packets;+stats->tx_bytes+=sstats->tx_bytes;+stats->tx_errors+=sstats->tx_errors;+stats->tx_dropped+=sstats->tx_dropped;++stats->multicast+=sstats->multicast;+stats->collisions+=sstats->collisions;++stats->rx_length_errors+=sstats->rx_length_errors;+stats->rx_over_errors+=sstats->rx_over_errors;+stats->rx_crc_errors+=sstats->rx_crc_errors;+stats->rx_frame_errors+=sstats->rx_frame_errors;+stats->rx_fifo_errors+=sstats->rx_fifo_errors;+stats->rx_missed_errors+=sstats->rx_missed_errors;++stats->tx_aborted_errors+=sstats->tx_aborted_errors;+stats->tx_carrier_errors+=sstats->tx_carrier_errors;+stats->tx_fifo_errors+=sstats->tx_fifo_errors;+stats->tx_heartbeat_errors+=sstats->tx_heartbeat_errors;+stats->tx_window_errors+=sstats->tx_window_errors;+}++read_unlock_bh(&parent->lock);++returnstats;+}++/* ---------------------------- Main funcs ---------------------------------- */+staticstructneigh*neigh_cmd_find_by_mac(structslave*slave,u8*mac)+{+structnet_device*dev=slave->dev;+structnet_device*parent_dev=dev->master;+structparent*parent=netdev_priv(parent_dev);+structneigh*neigh;+intfound=0;++list_for_each_entry(neigh,&parent->neigh_add_list,list){+if(!memcmp(neigh->emac,mac,ETH_ALEN)){+found=1;+break;+}+}++returnfound?neigh:NULL;+}++staticstructneigh*neigh_find_by_mac(structslave*slave,u8*mac)+{+structneigh*neigh;+intfound=0;++list_for_each_entry(neigh,&slave->neigh_list,list){+if(!memcmp(neigh->emac,mac,ETH_ALEN)){+found=1;+break;+}+}++returnfound?neigh:NULL;+}++staticintneigh_learn(structslave*slave,structsk_buff*skb,u8*remac)+{+structnet_device*dev=slave->dev;+structnet_device*parent_dev=dev->master;+structparent*parent=netdev_priv(parent_dev);+structneigh*neigh_cmd;+u8*rimac;+intrc;++/* linearize to easy on reading the arp payload */+rc=skb_linearize(skb);+if(rc){+pr_err("%s: skb_linearize failed rc %d\n",dev->name,rc);+gotoout;+}else+rimac=skb->data+sizeof(structarphdr);++/* check if entry is being processed or already exists */+if(neigh_find_by_mac(slave,remac))+gotoout;++if(neigh_cmd_find_by_mac(slave,remac))+gotoout;++neigh_cmd=parent_get_neigh_cmd('+',slave->dev->name,remac,rimac);+if(!neigh_cmd){+pr_err("%s cannot build neigh cmd\n",slave->dev->name);+rc=-ENOMEM;+gotoout;+}++list_add_tail(&neigh_cmd->list,&parent->neigh_add_list);++/* calls neigh_learn_task() */+queue_delayed_work(parent->wq,&parent->neigh_learn_work,0);++out:+returnrc;+}++staticvoidneigh_learn_task(structwork_struct*work)+{+structparent*parent=container_of(work,structparent,+neigh_learn_work.work);+structneigh*neigh_cmd,*neigh_cmd_tmp;++write_lock_bh(&parent->lock);++if(parent->kill_timers)+gotoout;++list_for_each_entry_safe(neigh_cmd,neigh_cmd_tmp,+&parent->neigh_add_list,list){+__parent_store_neighs(&parent->dev->dev,NULL,+neigh_cmd->cmd,PAGE_SIZE);+list_del(&neigh_cmd->list);+kfree(neigh_cmd);+}++out:+write_unlock_bh(&parent->lock);+return;+}++staticvoidparent_work_cancel_all(structparent*parent)+{+write_lock_bh(&parent->lock);+parent->kill_timers=1;+write_unlock_bh(&parent->lock);++if(delayed_work_pending(&parent->neigh_learn_work))+cancel_delayed_work(&parent->neigh_learn_work);++if(delayed_work_pending(&parent->emac_ip_work))+cancel_delayed_work(&parent->emac_ip_work);++if(delayed_work_pending(&parent->migrate_out_work))+cancel_delayed_work(&parent->migrate_out_work);+}++staticstructparent*get_parent_by_pif_name(char*pif_name)+{+structparent*parent,*nxt;++list_for_each_entry_safe(parent,nxt,&parent_dev_list,parent_list){+if(!strcmp(parent->ipoib_main_interface,pif_name))+returnparent;+}+returnNULL;+}++staticvoidfree_ip_mem_in_rec(structguest_emac_info*emac_info)+{+structip_member*ipm,*tmp_ipm;+list_for_each_entry_safe(ipm,tmp_ipm,&emac_info->ip_list,list){+list_del(&ipm->list);+kfree(ipm);+}+}++staticinlinevoidfree_invalid_emac_ip_det(structparent*parent)+{+structguest_emac_info*emac_info,*emac_info_tmp;++list_for_each_entry_safe(emac_info,emac_info_tmp,+&parent->emac_ip_list,list){+if(emac_info->rec_state==INVALID){+free_ip_mem_in_rec(emac_info);+list_del(&emac_info->list);+kfree(emac_info);+}+}+}++staticvoidemac_info_clean_task(structwork_struct*work)+{+structparent*parent=container_of(work,structparent,+emac_ip_work.work);++write_lock_bh(&parent->lock);++if(parent->kill_timers)+gotoout;++free_invalid_emac_ip_det(parent);++out:+write_unlock_bh(&parent->lock);+return;+}++staticintmigrate_out_gen_arp_req(structparent*parent,u8*emac,+u16vlan)+{+structguest_emac_info*emac_info;+structip_member*ipm;+structslave*slave;+structsk_buff*nskb;+intret=0;++slave=get_slave_by_mac_and_vlan(parent,parent->dev->dev_addr,vlan);+if(unlikely(!slave)){+pr_info("%s: Failed to find parent slave !!! %pM\n",+__func__,parent->dev->dev_addr);+return-ENODEV;+}++emac_info=get_mac_ip_info_by_mac_and_vlan(parent,emac,vlan);++if(!emac_info)+return0;++/* go over all ip's attached to that mac */+list_for_each_entry(ipm,&emac_info->ip_list,list){+/* create and send arp request to that ip.*/+pr_info("%s: Sending arp For migrate_out event, to %pI4 "+"from 0.0.0.0\n",parent->dev->name,&(ipm->ip));++nskb=arp_create(ARPOP_REQUEST,+ETH_P_ARP,+ipm->ip,+slave->dev,+0,+slave->dev->broadcast,+slave->dev->broadcast,+slave->dev->broadcast);+if(nskb)+arp_xmit(nskb);+else{+pr_err("%s: %s failed creating skb\n",+__func__,slave->dev->name);+ret=-ENOMEM;+}+}+returnret;+}++staticvoidmigrate_out_work_task(structwork_struct*work)+{+structparent*parent=container_of(work,structparent,+migrate_out_work.work);+structguest_emac_info*emac_info;+intis_reschedule=0;+intret;++write_lock_bh(&parent->lock);++if(parent->kill_timers)+gotoout;++list_for_each_entry(emac_info,&parent->emac_ip_list,list){+if(emac_info->rec_state==MIGRATED_OUT){+if(emac_info->num_of_retries<+MIG_OUT_MAX_ARP_RETRIES){+ret=migrate_out_gen_arp_req(parent,emac_info->emac,+emac_info->vlan);+if(ret)+pr_err("%s: migrate_out_gen_arp failed: %d\n",+__func__,ret);++emac_info->num_of_retries=+emac_info->num_of_retries+1;+is_reschedule=1;+}else+emac_info->rec_state=INVALID;+}+}+/* issue arp request till the device removed that entry from list */+if(is_reschedule)+queue_delayed_work(parent->wq,&parent->migrate_out_work,+MIG_OUT_ARP_REQ_ISSUE_TIME);+out:+write_unlock_bh(&parent->lock);+return;+}++staticinlineintadd_emac_ip_info(structnet_device*slave_dev,__be32ip,+u8*mac,u16vlan)+{+structnet_device*parent_dev=slave_dev->master;+structparent*parent=netdev_priv(parent_dev);+structguest_emac_info*emac_info=NULL;+structip_member*ipm;+intret;+intis_just_alloc_emac_info=0;++ret=is_mac_info_contain_ip(parent,mac,ip,emac_info,vlan);+if(ret)+return0;++/* new ip add it to the emc_ip obj */+if(!emac_info){+emac_info=kzalloc(sizeof*emac_info,GFP_ATOMIC);+if(!emac_info){+pr_err("%s: Failed allocating emac_info\n",+parent_dev->name);+return-ENOMEM;+}+memcpy(emac_info->emac,mac,ETH_ALEN);+INIT_LIST_HEAD(&emac_info->ip_list);+emac_info->rec_state=VALID;+emac_info->vlan=vlan;+emac_info->num_of_retries=0;+list_add_tail(&emac_info->list,&parent->emac_ip_list);+is_just_alloc_emac_info=1;+}++ipm=kzalloc(sizeof*ipm,GFP_ATOMIC);+if(!ipm){+pr_err(" %s Failed allocating emac_info (ipm)\n",+parent_dev->name);+if(is_just_alloc_emac_info)+kfree(emac_info);+return-ENOMEM;+}++ipm->ip=ip;+list_add_tail(&ipm->list,&emac_info->ip_list);++return0;+}++/* build ipoib arp/rarp request/reply packet */+staticstructsk_buff*get_slave_skb_arp(structslave*slave,+structsk_buff*skb,+u8*rimac,int*ret)+{+structsk_buff*nskb;+structarphdr*arphdr=(structarphdr*)+(skb->data+sizeof(structethhdr));+structeth_arp_data*arp_data=(structeth_arp_data*)+(skb->data+sizeof(structethhdr)++sizeof(structarphdr));+u8t_addr[ETH_ALEN]={0};+interr=0;+/* mark regular packet handling */+*ret=0;++/*+*live-migrationsupport:keepsthenewmac/ipaddress:+*Inthatwayeachdriverknowswhichmac/vlan-IP'swhereonthe+*guestsabove,whenevermigrate_outeventcomesitwillsend+*arprequestforalltheseIP's.+*/+if(skb->protocol==htons(ETH_P_ARP))+err=add_emac_ip_info(slave->dev,arp_data->arp_sip,+arp_data->arp_sha,slave->vlan);+if(err)+pr_warn("%s: Failed creating: emac_ip_info for ip: %pI4",+__func__,&arp_data->arp_sip);+/*+*livemigrationsupport:+*1.checckifweareinlivemigrationprocess+*2.checkifthearpresponseisfortheparent+*3.ignorelocal-administratedbit,whichwassettomakesure+*thatthebridgewillnotdropit.+*/+arp_data->arp_dha[0]=arp_data->arp_dha[0]&0xFD;+if(htons(ARPOP_REPLY)==(arphdr->ar_op)&&+!memcmp(arp_data->arp_dha,slave->dev->master->dev_addr,ETH_ALEN)){+/*+*whenthesourceistheparentinterface,assumes+*thatweareinthemiddleoflivemigrationprocess,+*so,wewillsendgratuitousarp.+*/+pr_info("%s: Arp packet for parent: %s",+__func__,slave->dev->master->name);+/* create gratuitous ARP on behalf of the guest */+nskb=arp_create(ARPOP_REQUEST,+be16_to_cpu(skb->protocol),+arp_data->arp_sip,+slave->dev,+arp_data->arp_sip,+NULL,+slave->dev->dev_addr,+t_addr);+if(unlikely(!nskb))+pr_err("%s: %s live migration: failed creating skb\n",+__func__,slave->dev->name);+}else{+nskb=arp_create(be16_to_cpu(arphdr->ar_op),+be16_to_cpu(skb->protocol),+arp_data->arp_dip,+slave->dev,+arp_data->arp_sip,+rimac,+slave->dev->dev_addr,+NULL);+}++returnnskb;+}++/*+*buildipoibarprequestpacketaccordingtoipheader.+*usesforlive-migration,ormissingneighfornewvif.+*/+staticvoidget_slave_skb_arp_by_ip(structslave*slave,+structsk_buff*skb)+{+structsk_buff*nskb=NULL;+structiphdr*iph=ip_hdr(skb);++pr_info("Sending arp on behalf of slave %s, from %pI4"+" to %pI4",slave->dev->name,&(iph->saddr),+&(iph->daddr));++nskb=arp_create(ARPOP_REQUEST,+ETH_P_ARP,+iph->daddr,+slave->dev,+iph->saddr,+slave->dev->broadcast,+slave->dev->dev_addr,+NULL);+if(nskb)+arp_xmit(nskb);+else+pr_err("%s: %s failed creating skb\n",+__func__,slave->dev->name);+}++/* build ipoib ipv4/ipv6 packet */+staticstructsk_buff*get_slave_skb_ip(structslave*slave,+structsk_buff*skb)+{++skb_pull(skb,ETH_HLEN);+skb_reset_network_header(skb);++returnskb;+}++/*+*get_slave_skb--calledinTXflow+*getskbthatcanbesentthruslavexmitfunc,+*ifskbwasadjusted(cloned,pulled,etc..)successfully+*theoldskb(ifany)isfreedhere.+*/+staticstructsk_buff*get_slave_skb(structslave*slave,structsk_buff*skb)+{+structnet_device*dev=slave->dev;+structnet_device*parent_dev=dev->master;+structparent*parent=netdev_priv(parent_dev);+structsk_buff*nskb=NULL;+structethhdr*ethh=(structethhdr*)(skb->data);+structneigh*neigh=NULL;+u8rimac[INFINIBAND_ALEN];+intret=0;++/* set neigh mac */+if(is_multicast_ether_addr(ethh->h_dest)){+memcpy(rimac,dev->broadcast,INFINIBAND_ALEN);+}else{+neigh=neigh_find_by_mac(slave,ethh->h_dest);+if(neigh){+memcpy(rimac,neigh->imac,INFINIBAND_ALEN);+}else{+++parent->port_stats.tx_neigh_miss;+/*+*assumeVIFmigration,triestogettheneighby+*issuearprequestonbehalfofthevif.+*/+if(skb->protocol==htons(ETH_P_IP)){+pr_info("Missed neigh for slave: %s,"+"issue ARP request\n",+slave->dev->name);+get_slave_skb_arp_by_ip(slave,skb);+gotoout_arp_sent_instead;+}+}+}++if(skb->protocol==htons(ETH_P_ARP)||+skb->protocol==htons(ETH_P_RARP)){+nskb=get_slave_skb_arp(slave,skb,rimac,&ret);+if(!nskb&&LIVE_MIG_PACKET==ret){+pr_info("%s: live migration packets\n",__func__);+gotoerr;+}+}else{+if(!neigh)+gotoerr;+/* pull ethernet header here */+nskb=get_slave_skb_ip(slave,skb);+}++/* if new skb could not be adjusted/allocated, abort */+if(!nskb){+pr_err("%s get_slave_skb_ip/arp failed 0x%x\n",+dev->name,skb->protocol);+gotoerr;+}++if(neigh&&nskb==skb){/* ucast & non-arp/rarp */+/* dev_hard_header only for ucast, for arp done already.*/+if(dev_hard_header(nskb,dev,ntohs(skb->protocol),rimac,+dev->dev_addr,nskb->len)<0){+pr_warn("%s: dev_hard_header failed\n",+dev->name);+gotoerr;+}+}++/*+*newskbisreadytobesent,cleanoldskbifweholdaclone+*(oldskbisnotshared,alreadycheckedthat.)+*/+if((nskb!=skb))+dev_kfree_skb(skb);++nskb->dev=slave->dev;+returnnskb;++out_arp_sent_instead:/* whenever sent arp instead of ip packet */+err:+/* got error after nskb was adjusted/allocated */+if(nskb&&(nskb!=skb))+dev_kfree_skb(nskb);++returnNULL;+}++staticstructsk_buff*get_parent_skb_arp(structslave*slave,+structsk_buff*skb,+u8*remac)+{+structnet_device*dev=slave->dev->master;+structsk_buff*nskb;+structarphdr*arphdr=(structarphdr*)(skb->data);+structipoib_arp_data*arp_data=(structipoib_arp_data*)+(skb->data+sizeof(structarphdr));+u8*target_hw=slave->emac;+u8*dst_hw=slave->emac;+u8local_eth_addr[ETH_ALEN];++/* live migration: gets arp with broadcast src and dst */+if(!memcmp(arp_data->arp_sha,slave->dev->broadcast,INFINIBAND_ALEN)&&+!memcmp(arp_data->arp_dha,slave->dev->broadcast,INFINIBAND_ALEN)){+pr_info("%s: ARP with bcast src and dest send from src_hw: %pM\n",+__func__,slave->dev->master->dev_addr);+/* replace the src with the parent src: */+memcpy(local_eth_addr,slave->dev->master->dev_addr,ETH_ALEN);+/*+*setlocaladministratedbit,+*thatwaythebridgewillnotthrowsit+*/+local_eth_addr[0]=local_eth_addr[0]|0x2;+memcpy(remac,local_eth_addr,ETH_ALEN);+target_hw=NULL;+dst_hw=NULL;+}++nskb=arp_create(be16_to_cpu(arphdr->ar_op),+be16_to_cpu(skb->protocol),+arp_data->arp_dip,+dev,+arp_data->arp_sip,+dst_hw,+remac,+target_hw);++/* prepare place for the headers. */+if(nskb)+skb_reserve(nskb,ETH_HLEN);++returnnskb;+}++staticstructsk_buff*get_parent_skb_ip(structslave*slave,+structsk_buff*skb)+{+/* nop */+returnskb;+}++/* get_parent_skb -- called in RX flow */+staticstructsk_buff*get_parent_skb(structslave*slave,+structsk_buff*skb,u8*remac)+{+structnet_device*dev=slave->dev->master;+structsk_buff*nskb=NULL;+structethhdr*ethh;++if(skb->protocol==htons(ETH_P_ARP)||+skb->protocol==htons(ETH_P_RARP))+nskb=get_parent_skb_arp(slave,skb,remac);+else+nskb=get_parent_skb_ip(slave,skb);++/* if new skb could not be adjusted/allocated, abort */+if(!nskb)+gotoerr;++/* at this point, we can free old skb if it was cloned */+if(nskb&&(nskb!=skb))+dev_kfree_skb(skb);++skb=nskb;++/* build ethernet header */+ethh=(structethhdr*)skb_push(skb,ETH_HLEN);+ethh->h_proto=skb->protocol;+memcpy(ethh->h_source,remac,ETH_ALEN);+memcpy(ethh->h_dest,slave->emac,ETH_ALEN);++/* zero padding whenever is needed (arp for example).to ETH_ZLEN size */+if(unlikely((skb->len<ETH_ZLEN))){+if((skb->tail+(ETH_ZLEN-skb->len)>skb->end)||+skb_is_nonlinear(skb))+/* nothing */;+else+memset(skb_put(skb,ETH_ZLEN-skb->len),0,+ETH_ZLEN-skb->len);+}++/* set new skb fields */+skb->pkt_type=PACKET_HOST;+/*+*usemasterdev,toallownetpoll_receive_skb()+*innetif_receive_skb()+*/+skb->dev=dev;++/* pull the Ethernet header and update other fields */+skb->protocol=eth_type_trans(skb,skb->dev);++returnskb;++err:+/* got error after nskb was adjusted/allocated */+if(nskb&&(nskb!=skb))+dev_kfree_skb(nskb);++returnNULL;+}++staticintparent_rx(structsk_buff*skb,structslave*slave)+{+structnet_device*slave_dev=skb->dev;+structnet_device*parent_dev=slave_dev->master;+structparent*parent=netdev_priv(parent_dev);+structeipoib_cb_data*data=IPOIB_HANDLER_CB(skb);+structnapi_struct*napi=data->rx.napi;+structsk_buff*nskb;+intrc=0;+u8remac[ETH_ALEN];+intvlan_tag;++build_neigh_mac(remac,data->rx.sqpn,data->rx.slid);++read_lock_bh(&parent->lock);++if(unlikely(skb_headroom(skb)<ETH_HLEN)){+pr_warn("%s: small headroom %d < %d\n",+skb->dev->name,skb_headroom(skb),ETH_HLEN);+++parent->port_stats.rx_skb_errors;+gotodrop;+}++/* learn neighs based on ARP snooping */+if(unlikely(ntohs(skb->protocol)==ETH_P_ARP)){+read_unlock_bh(&parent->lock);+write_lock_bh(&parent->lock);+neigh_learn(slave,skb,remac);+write_unlock_bh(&parent->lock);+read_lock_bh(&parent->lock);+}++nskb=get_parent_skb(slave,skb,remac);+if(unlikely(!nskb)){+++parent->port_stats.rx_skb_errors;+pr_warn("%s: failed to create parent_skb\n",+skb->dev->name);+gotodrop;+}else+skb=nskb;++vlan_tag=slave->vlan&0xfff;+if(vlan_tag){+skb=__vlan_hwaccel_put_tag(skb,vlan_tag);+if(!skb){+pr_err("%s failed to insert VLAN tag\n",+skb->dev->name);+gotodrop;+}+++parent->port_stats.rx_vlan;+}++if(napi)+rc=napi_gro_receive(napi,skb);+else+rc=netif_receive_skb(skb);++read_unlock_bh(&parent->lock);++returnrc;++drop:+dev_kfree_skb_any(skb);+read_unlock_bh(&parent->lock);++returnNET_RX_DROP;+}++staticrx_handler_result_teipoib_handle_frame(structsk_buff**pskb)+{+structsk_buff*skb=*pskb;+structslave*slave;++slave=eipoib_slave_get_rcu(skb->dev);++parent_rx(skb,slave);++returnRX_HANDLER_CONSUMED;+}++staticnetdev_tx_tparent_tx(structsk_buff*skb,structnet_device*dev)+{+structparent*parent=netdev_priv(dev);+structslave*slave=NULL;+structethhdr*ethh=(structethhdr*)(skb->data);+structsk_buff*nskb;+intrc;+u16vlan;+u8mac_no_admin_bit[ETH_ALEN];++read_lock_bh(&parent->lock);++if(unlikely(!IS_E_IPOIB_PROTO(ethh->h_proto))){+++parent->port_stats.tx_proto_errors;+gotodrop;+}+/* assume: only orphan skb's */+if(unlikely(skb_shared(skb))){+++parent->port_stats.tx_shared;+gotodrop;+}++/* obtain VLAN information if present */+if(vlan_tx_tag_present(skb)){+vlan=vlan_tx_tag_get(skb)&0xfff;+++parent->port_stats.tx_vlan;+}else{+vlan=VLAN_N_VID;+}++/*+*forlivemigration:masktheadminbitifexists.+*onlyinARPpacketsthatcamefromparent'sVIFinterface.+*/+if(unlikely((htons(ETH_P_ARP)==ethh->h_proto)&&+!memcmp(parent->dev->dev_addr+1,ethh->h_source+1,ETH_ALEN-1))){+/* parent's VIF: */+memcpy(mac_no_admin_bit,ethh->h_source,ETH_ALEN);+mac_no_admin_bit[0]=mac_no_admin_bit[0]&0xFD;+/* get slave, and queue packet */+slave=get_slave_by_mac_and_vlan(parent,mac_no_admin_bit,vlan);+}+/* get slave, and queue packet */+if(!slave)+slave=get_slave_by_mac_and_vlan(parent,ethh->h_source,vlan);+if(unlikely(!slave)){+pr_info("vif: %pM with vlan: %d miss for parent: %s\n",+ethh->h_source,vlan,parent->ipoib_main_interface);+++parent->port_stats.tx_vif_miss;+gotodrop;+}++nskb=get_slave_skb(slave,skb);+if(unlikely(!nskb)){+++parent->port_stats.tx_skb_errors;+gotodrop;+}else+skb=nskb;++/*+*VSTmode:removesthevlantaginthetx(willadditintherx)+*theslaveisfromIPoIBanditisNETIF_F_VLAN_CHALLENGED,+*somustremovethevlantag.+*/+if(vlan!=VLAN_N_VID)+skb->vlan_tci=0;++/* arp packets: */+if(skb->protocol==htons(ETH_P_ARP)||+skb->protocol==htons(ETH_P_RARP)){+arp_xmit(skb);+gotoout;+}++/* ip packets */+skb_record_rx_queue(skb,skb_get_queue_mapping(skb));++rc=dev_queue_xmit(skb);+if(unlikely(rc)){+pr_err("slave tx method failed %d\n",rc);+++parent->port_stats.tx_slave_err;+dev_kfree_skb(skb);+}++gotoout;++drop:+++parent->port_stats.tx_parent_dropped;+dev_kfree_skb(skb);++out:+read_unlock_bh(&parent->lock);+returnNETDEV_TX_OK;+}++staticintparent_open(structnet_device*parent_dev)+{+structparent*parent=netdev_priv(parent_dev);++parent->kill_timers=0;+INIT_DELAYED_WORK(&parent->neigh_learn_work,neigh_learn_task);+INIT_DELAYED_WORK(&parent->emac_ip_work,emac_info_clean_task);+INIT_DELAYED_WORK(&parent->migrate_out_work,migrate_out_work_task);+return0;+}++staticintparent_close(structnet_device*parent_dev)+{+structparent*parent=netdev_priv(parent_dev);++write_lock_bh(&parent->lock);+parent->kill_timers=1;+write_unlock_bh(&parent->lock);++cancel_delayed_work(&parent->neigh_learn_work);+cancel_delayed_work(&parent->emac_ip_work);+cancel_delayed_work(&parent->migrate_out_work);++return0;+}+++staticvoidparent_deinit(structnet_device*parent_dev)+{+structparent*parent=netdev_priv(parent_dev);++list_del(&parent->parent_list);++parent_work_cancel_all(parent);+}++staticvoidparent_uninit(structnet_device*parent_dev)+{+structparent*parent=netdev_priv(parent_dev);++parent_deinit(parent_dev);+parent_destroy_sysfs_entry(parent);++if(parent->wq)+destroy_workqueue(parent->wq);+}++staticstructlock_class_keyparent_netdev_xmit_lock_key;+staticstructlock_class_keyparent_netdev_addr_lock_key;++staticvoidparent_set_lockdep_class_one(structnet_device*dev,+structnetdev_queue*txq,+void*_unused)+{+lockdep_set_class(&txq->_xmit_lock,+&parent_netdev_xmit_lock_key);+}++staticvoidparent_set_lockdep_class(structnet_device*dev)+{+lockdep_set_class(&dev->addr_list_lock,+&parent_netdev_addr_lock_key);+netdev_for_each_tx_queue(dev,parent_set_lockdep_class_one,NULL);+}++staticintparent_init(structnet_device*parent_dev)+{+structparent*parent=netdev_priv(parent_dev);++parent->wq=create_singlethread_workqueue(parent_dev->name);+if(!parent->wq)+return-ENOMEM;++parent_set_lockdep_class(parent_dev);++list_add_tail(&parent->parent_list,&parent_dev_list);++return0;+}++staticu16parent_select_q(structnet_device*dev,structsk_buff*skb)+{+returnskb_tx_hash(dev,skb);+}++staticintparent_add_vif_param(structnet_device*parent_dev,+structnet_device*new_vif_dev,+u16vlan,u8*mac)+{+structparent*parent=netdev_priv(parent_dev);+structslave*new_slave,*slave_tmp;+intret=0;++if(!is_valid_ether_addr(mac)){+pr_err("Invalid mac input for slave:%pM \n",mac);+return-EINVAL;+}++write_lock_bh(&parent->lock);++new_slave=get_slave_by_dev(parent,new_vif_dev);+if(!new_slave){+pr_err("%s: ERROR no slave:%s.!!!! \n",+__func__,new_vif_dev->name);+ret=-EINVAL;+gotoout;+}++if(!is_zero_ether_addr(new_slave->emac)){+pr_err("slave %s mac already set to %pM\n",+new_slave->dev->name,new_slave->emac);+ret=-EINVAL;+gotoout;+}++/* check another slave has this mac/vlan */+parent_for_each_slave(parent,slave_tmp){+if(!memcmp(slave_tmp->emac,mac,ETH_ALEN)&&+slave_tmp->vlan==new_slave->vlan){+pr_err("cannot update %s, slave %s already has"+" vlan 0x%x mac %pM\n",+parent->dev->name,new_slave->dev->name,+slave_tmp->vlan,+mac);+ret=-EINVAL;+gotoout;+}+}++/* ready to go */+pr_info("slave %s mac is set to %pM, vlan set to: %d\n",+new_slave->dev->name,mac,vlan);++memcpy(new_slave->emac,mac,ETH_ALEN);++new_slave->vlan=vlan;++out:+write_unlock_bh(&parent->lock);++returnret;+}++staticconststructnet_device_opsparent_netdev_ops={+.ndo_init=parent_init,+.ndo_uninit=parent_uninit,+.ndo_open=parent_open,+.ndo_stop=parent_close,+.ndo_start_xmit=parent_tx,+.ndo_select_queue=parent_select_q,+/* parnt mtu is min(slaves_mtus) */+.ndo_change_mtu=NULL,+.ndo_fix_features=parent_fix_features,+/*+*initialmacaddressisrandomized,canbechanged+*thruthisfunclater+*/+.ndo_set_mac_address=eth_mac_addr,+.ndo_get_stats64=parent_get_stats,+.ndo_add_slave=parent_enslave,+.ndo_del_slave=parent_release_slave,+.ndo_set_vif_param=parent_add_vif_param,+};++staticvoidparent_setup(structnet_device*parent_dev)+{+structparent*parent=netdev_priv(parent_dev);++/* initialize rwlocks */+rwlock_init(&parent->lock);++/* Initialize pointers */+parent->dev=parent_dev;+INIT_LIST_HEAD(&parent->neigh_add_list);+INIT_LIST_HEAD(&parent->slave_list);+INIT_LIST_HEAD(&parent->emac_ip_list);+/* Initialize the device entry points */+ether_setup(parent_dev);+/* parent_dev->hard_header_len is adjusted later */+parent_dev->netdev_ops=&parent_netdev_ops;+parent_set_ethtool_ops(parent_dev);++/* Initialize the device options */+parent_dev->tx_queue_len=0;+/* mark the parent intf as pif (master of other vifs.) */+parent_dev->priv_flags=IFF_EIPOIB_PIF;++parent_dev->hw_features=NETIF_F_SG|NETIF_F_IP_CSUM|+NETIF_F_RXCSUM|NETIF_F_GRO|NETIF_F_TSO;++parent_dev->features=parent_dev->hw_features;+parent_dev->vlan_features=parent_dev->hw_features;++parent_dev->features|=PARENT_VLAN_FEATURES;+}++/*+*Createanewparentbasedonthespecifiednameandparentparameters.+*CallermustNOTholdrtnl_lock;weneedtoreleaseitherebeforewe+*setupoursysfsentries.+*/+staticstructparent*parent_create(structnet_device*ibd)+{+structnet_device*parent_dev;+u32num_queues;+intrc;+unionib_gidgid;+structparent*parent=NULL;+inti,j;++memcpy(&gid,ibd->dev_addr+4,sizeof(unionib_gid));+num_queues=num_online_cpus();+num_queues=roundup_pow_of_two(num_queues);++parent_dev=alloc_netdev_mq(sizeof(structparent),"",+parent_setup,num_queues);+if(!parent_dev){+pr_err("%s failed to alloc netdev!\n",ibd->name);+rc=-ENOMEM;+gotoout_rtnl;+}++rc=dev_alloc_name(parent_dev,"eth%d");+if(rc<0)+gotoout_netdev;++/* eIPoIB interface mac format. */+for(i=0,j=0;i<8;i++){+if((PARENT_MAC_MASK>>i)&0x1){+if(j<6)/* only 6 bytes eth address */+parent_dev->dev_addr[j]=+gid.raw[GUID_LEN+i];+j++;+}+}++/* assuming that the ibd->dev.parent was alreadey been set. */+SET_NETDEV_DEV(parent_dev,ibd->dev.parent);++rc=register_netdevice(parent_dev);+if(rc<0)+gotoout_parent;++dev_net_set(parent_dev,&init_net);++rc=parent_create_sysfs_entry(netdev_priv(parent_dev));+if(rc<0)+gotoout_unreg;++parent=netdev_priv(parent_dev);+memcpy(parent->gid.raw,gid.raw,GID_LEN);+strncpy(parent->ipoib_main_interface,ibd->name,IFNAMSIZ);+parent_dev->dev_id=ibd->dev_id;++returnparent;++out_unreg:+unregister_netdevice(parent_dev);+out_parent:+parent_deinit(parent_dev);+out_netdev:+free_netdev(parent_dev);+out_rtnl:+returnERR_PTR(rc);+}+++staticvoidparent_free(structparent*parent)+{+structnet_device*parent_dev=parent->dev;++parent_work_cancel_all(parent);++parent_release_all(parent_dev);++unregister_netdevice(parent_dev);+}++staticvoidparent_free_all(void)+{+structparent*parent,*nxt;++list_for_each_entry_safe(parent,nxt,&parent_dev_list,parent_list)+parent_free(parent);+}++/* netdev events handlers */+staticinlineintis_ipoib_pif_intf(structnet_device*dev)+{+if(ARPHRD_INFINIBAND==dev->type&&dev->priv_flags&IFF_EIPOIB_PIF)+return1;+return0;+}++staticintparent_event_changename(structparent*parent)+{+parent_destroy_sysfs_entry(parent);++parent_create_sysfs_entry(parent);++returnNOTIFY_DONE;+}++staticintparent_master_netdev_event(unsignedlongevent,+structnet_device*parent_dev)+{+structparent*event_parent=netdev_priv(parent_dev);++switch(event){+caseNETDEV_CHANGENAME:+pr_info("%s: got NETDEV_CHANGENAME event",parent_dev->name);+returnparent_event_changename(event_parent);+default:+break;+}++returnNOTIFY_DONE;+}++staticintparent_slave_netdev_event(unsignedlongevent,+structnet_device*slave_dev)+{+structnet_device*parent_dev=slave_dev->master;+structparent*parent=netdev_priv(parent_dev);++if(!parent_dev){+pr_err("slave:%s has no parent.\n",slave_dev->name);+returnNOTIFY_DONE;+}++switch(event){+caseNETDEV_UNREGISTER:+parent_release_slave(parent_dev,slave_dev);+break;+caseNETDEV_CHANGE:+caseNETDEV_UP:+caseNETDEV_DOWN:+parent_set_carrier(parent);+break;+caseNETDEV_CHANGEMTU:+parent_set_mtu(parent);+break;+caseNETDEV_CHANGENAME:+break;+caseNETDEV_FEAT_CHANGE:+parent_compute_features(parent);+break;+default:+break;+}++returnNOTIFY_DONE;+}++staticinteipoib_netdev_event(structnotifier_block*this,+unsignedlongevent,void*ptr)+{+structnet_device*event_dev=(structnet_device*)ptr;++if(dev_net(event_dev)!=&init_net)+returnNOTIFY_DONE;++if(is_parent(event_dev))+returnparent_master_netdev_event(event,event_dev);++if(is_slave(event_dev))+returnparent_slave_netdev_event(event,event_dev);+/*+*generalnetworkdevicetriggersevent,checkifitisnew+*ibinterfacethatwewanttoenslave.+*/+returneipoib_device_event(this,event,ptr);+}++staticstructnotifier_blockparent_netdev_notifier={+.notifier_call=eipoib_netdev_event,+};++staticinteipoib_device_event(structnotifier_block*unused,+unsignedlongevent,void*ptr)+{+structnet_device*dev=ptr;+structparent*parent;++if(!is_ipoib_pif_intf(dev))+returnNOTIFY_DONE;++switch(event){+caseNETDEV_REGISTER:+parent=parent_create(dev);+if(IS_ERR(parent)){+pr_warn("failed to create parent for %s\n",+dev->name);+break;+}+break;+caseNETDEV_UNREGISTER:+parent=get_parent_by_pif_name(dev->name);+if(parent)+parent_free(parent);+break;+default:+break;+}++returnNOTIFY_DONE;+}++staticint__initmod_init(void)+{+intrc;++pr_info(DRV_NAME": %s",version);++rc=register_pernet_subsys(&eipoib_net_ops);+if(rc)+gotoout;++rc=register_netdevice_notifier(&parent_netdev_notifier);+if(rc){+pr_err("%s failed to register_netdevice_notifier, rc: 0x%x\n",+__func__,rc);+gotounreg_subsys;+}++gotoout;++unreg_subsys:+unregister_pernet_subsys(&eipoib_net_ops);+out:+returnrc;++}++staticvoid__exitmod_exit(void)+{+unregister_netdevice_notifier(&parent_netdev_notifier);++unregister_pernet_subsys(&eipoib_net_ops);++rtnl_lock();+parent_free_all();+rtnl_unlock();+}++module_init(mod_init);+module_exit(mod_exit);+MODULE_LICENSE("Dual BSD/GPL");+MODULE_VERSION(DRV_VERSION);+MODULE_DESCRIPTION(DRV_DESCRIPTION", v"DRV_VERSION);+MODULE_AUTHOR("Ali Ayoub && Erez Shitrit");
From: Or Gerlitz <hidden> Date: 2012-08-01 17:18:35
From: Erez Shitrit <redacted>
The management interface for the driver was using sysfs.
While most of this usage (all the "set" operations) was ported to use
rtnetlink, still some sysfs "show" entries were left in this posting,
till its clear if/how they can be implemented otherwise.
Here are few sysfs commands that are used in order to get
information from the driver,
1. get parent's slaves:
$ cat /sys/class/net/ethZ/eth/slaves
where ethZ is the driver interface
2. see the list of ipoib interfaces enslaved under eipoib interface,
$ cat /sys/class/net/ethX/eth/vifs
for example:
$ cat /sys/class/net/eth4/eth/vifs
SLAVE=ib0.1 MAC=9a:c2:1f:d7:3b:63 VLAN=N/A
SLAVE=ib0.2 MAC=52:54:00:60:55:88 VLAN=N/A
SLAVE=ib0.3 MAC=52:54:00:60:55:89 VLAN=N/A
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
drivers/net/eipoib/eth_ipoib_sysfs.c | 435 ++++++++++++++++++++++++++++++++++
1 files changed, 435 insertions(+), 0 deletions(-)
create mode 100644 drivers/net/eipoib/eth_ipoib_sysfs.c
@@ -0,0 +1,435 @@+/*+*Copyright(c)2012MellanoxTechnologies.Allrightsreserved+*+*Thissoftwareisavailabletoyouunderachoiceofoneoftwo+*licenses.YoumaychoosetobelicensedunderthetermsoftheGNU+*GeneralPublicLicense(GPL)Version2,availablefromthefile+*COPYINGinthemaindirectoryofthissourcetree,orthe+*openfabric.orgBSDlicensebelow:+*+*Redistributionanduseinsourceandbinaryforms,withor+*withoutmodification,arepermittedprovidedthatthefollowing+*conditionsaremet:+*+*-Redistributionsofsourcecodemustretaintheabove+*copyrightnotice,thislistofconditionsandthefollowing+*disclaimer.+*+*-Redistributionsinbinaryformmustreproducetheabove+*copyrightnotice,thislistofconditionsandthefollowing+*disclaimerinthedocumentationand/orothermaterials+*providedwiththedistribution.+*+*THESOFTWAREISPROVIDED"AS IS",WITHOUTWARRANTYOFANYKIND,+*EXPRESSORIMPLIED,INCLUDINGBUTNOTLIMITEDTOTHEWARRANTIESOF+*MERCHANTABILITY,FITNESSFORAPARTICULARPURPOSEAND+*NONINFRINGEMENT.INNOEVENTSHALLTHEAUTHORSORCOPYRIGHTHOLDERS+*BELIABLEFORANYCLAIM,DAMAGESOROTHERLIABILITY,WHETHERINAN+*ACTIONOFCONTRACT,TORTOROTHERWISE,ARISINGFROM,OUTOFORIN+*CONNECTIONWITHTHESOFTWAREORTHEUSEOROTHERDEALINGSINTHE+*SOFTWARE.+*/++#include<linux/kernel.h>+#include<linux/module.h>+#include<linux/device.h>+#include<linux/sched.h>+#include<linux/fs.h>+#include<linux/types.h>+#include<linux/string.h>+#include<linux/netdevice.h>+#include<linux/inetdevice.h>+#include<linux/in.h>+#include<linux/sysfs.h>+#include<linux/ctype.h>+#include<linux/inet.h>+#include<linux/rtnetlink.h>+#include<linux/etherdevice.h>+#include<net/net_namespace.h>++#include"eth_ipoib.h"++#define to_dev(obj) container_of(obj, struct device, kobj)+#define to_parent(cd) ((struct parent *)(netdev_priv(to_net_dev(cd))))+#define MOD_NA_STRING "N/A"++#define _sprintf(p, buf, format, arg...) \+((PAGE_SIZE-(int)(p-buf))<=0?0:\+scnprintf(p,PAGE_SIZE-(int)(p-buf),format,##arg))\++#define _end_of_line(_p, _buf) \+do{if(_p-_buf)/* eat the leftover space */\+buf[_p-_buf-1]='\n';\+}while(0)++/* helper functions */+staticintget_emac(u8*mac,char*s)+{+if(sscanf(s,"%hhx:%hhx:%hhx:%hhx:%hhx:%hhx",+mac+0,mac+1,mac+2,mac+3,mac+4,+mac+5)!=6)+return-1;++return0;+}++staticintget_imac(u8*mac,char*s)+{+if(sscanf(s,"%hhx:%hhx:%hhx:%hhx:%hhx:%hhx:%hhx:%hhx:"+"%hhx:%hhx:%hhx:%hhx:%hhx:%hhx:%hhx:%hhx:"+"%hhx:%hhx:%hhx:%hhx",+mac+0,mac+1,mac+2,mac+3,mac+4,+mac+5,mac+6,mac+7,mac+8,mac+9,+mac+10,mac+11,mac+12,mac+13,+mac+14,mac+15,mac+16,mac+17,+mac+18,mac+19)!=20)+return-1;++return0;+}++/* show/store functions per module (CLASS_ATTR) */+staticssize_tshow_parents(structclass*cls,structclass_attribute*attr,+char*buf)+{+char*p=buf;+structparent*parent;++rtnl_lock();/* because of parent_dev_list */++list_for_each_entry(parent,&parent_dev_list,parent_list){+p+=_sprintf(p,buf,"%s over IB port: %s\n",+parent->dev->name,+parent->ipoib_main_interface);+}+_end_of_line(p,buf);++rtnl_unlock();+return(ssize_t)(p-buf);+}++/* show/store functions per parent (DEVICE_ATTR) */+staticssize_tparent_show_neighs(structdevice*d,+structdevice_attribute*attr,char*buf)+{+structslave*slave;+structneigh*neigh;+structparent*parent=to_parent(d);+char*p=buf;++read_lock_bh(&parent->lock);+parent_for_each_slave(parent,slave){+list_for_each_entry(neigh,&slave->neigh_list,list){+p+=_sprintf(p,buf,"SLAVE=%-10s EMAC=%pM IMAC=%pM:%pM:%pM:%.2x:%.2x\n",+slave->dev->name,+neigh->emac,+neigh->imac,neigh->imac+6,neigh->imac+12,+neigh->imac[18],neigh->imac[19]);+}+}++read_unlock_bh(&parent->lock);++_end_of_line(p,buf);++return(ssize_t)(p-buf);+}++structneigh*parent_get_neigh_cmd(charop,+char*ifname,u8*remac,u8*rimac)+{+structneigh*neigh_cmd;++neigh_cmd=kzalloc(sizeof*neigh_cmd,GFP_ATOMIC);+if(!neigh_cmd){+pr_err("%s cannot allocate neigh struct\n",ifname);+gotoout;+}++/*+*populateemacfieldsoitcanbeusedeasily+*inneigh_cmd_find_by_mac()+*/+memcpy(neigh_cmd->emac,remac,ETH_ALEN);+memcpy(neigh_cmd->imac,rimac,INFINIBAND_ALEN);++/* prepare the command as a string */+sprintf(neigh_cmd->cmd,"%c%s %pM %pM:%pM:%pM:%.2x:%.2x",+op,ifname,remac,rimac,rimac+6,rimac+12,rimac[18],rimac[19]);+out:+returnneigh_cmd;+}++/* write_lock_bh(&parent->lock) must be held */+ssize_t__parent_store_neighs(structdevice*d,+structdevice_attribute*attr,+constchar*buffer,size_tcount)+{+charcommand[IFNAMSIZ+1]={0,};+charemac_str[ETH_ALEN*3]={0,};+u8emac[ETH_ALEN];+charimac_str[INFINIBAND_ALEN*3]={0,};+u8imac[INFINIBAND_ALEN];+char*ifname;+intfound=0,ret=count;+structslave*slave=NULL,*slave_tmp;+structneigh*neigh;+structparent*parent=to_parent(d);++sscanf(buffer,"%s %s %s",command,emac_str,imac_str);++/* check ifname */+ifname=command+1;+if((strlen(command)<=1)||!dev_valid_name(ifname)||+(command[0]!='+'&&command[0]!='-'))+gotoerr_no_cmd;++/* check if ifname exist */+parent_for_each_slave(parent,slave_tmp){+if(!strcmp(slave_tmp->dev->name,ifname)){+found=1;+slave=slave_tmp;+}+}++if(!found){+pr_err("%s could not find slave\n",ifname);+ret=-EINVAL;+gotoout;+}++if(get_emac(emac,emac_str)){+pr_err("%s bad emac %s\n",ifname,emac_str);+ret=-EINVAL;+gotoout;+}++if(get_imac(imac,imac_str)){+pr_err("%s bad imac %s\n",ifname,imac_str);+ret=-EINVAL;+gotoout;+}++/* process command */+if(command[0]=='+'){+found=0;+list_for_each_entry(neigh,&slave->neigh_list,list){+if(!memcmp(neigh->emac,emac,ETH_ALEN))+found=1;+}++if(found){+pr_err("%s: cannot update neigh, slave already has "+"this neigh mac %pM\n",+slave->dev->name,emac);+ret=-EINVAL;+gotoout;+}++neigh=kzalloc(sizeof*neigh,GFP_ATOMIC);+if(!neigh){+pr_err("%s cannot allocate neigh struct\n",+slave->dev->name);+ret=-ENOMEM;+gotoout;+}++/* ready to go */+pr_info("%s: slave %s neigh mac is set to %pM\n",+ifname,parent->dev->name,emac);+memcpy(neigh->emac,emac,ETH_ALEN);+memcpy(neigh->imac,imac,INFINIBAND_ALEN);++list_add_tail(&neigh->list,&slave->neigh_list);++gotoout;+}++if(command[0]=='-'){+found=0;+list_for_each_entry(neigh,&slave->neigh_list,list){+if(!memcmp(neigh->emac,emac,ETH_ALEN))+found=1;+}++if(!found){+pr_err("%s cannot delete neigh mac %pM\n",+ifname,emac);+ret=-EINVAL;+gotoout;+}++list_del(&neigh->list);+kfree(neigh);++gotoout;+}++err_no_cmd:+pr_err("%s USAGE: (-|+)ifname emac imac\n",DRV_NAME);+ret=-EPERM;++out:+returnret;+}++staticDEVICE_ATTR(neighs,S_IRUGO,parent_show_neighs,+NULL);++staticssize_tparent_show_vifs(structdevice*d,+structdevice_attribute*attr,char*buf)+{+structslave*slave;+structparent*parent=to_parent(d);+char*p=buf;++read_lock_bh(&parent->lock);+parent_for_each_slave(parent,slave){+if(is_zero_ether_addr(slave->emac)){+p+=_sprintf(p,buf,"SLAVE=%-10s MAC=%-17s "+"VLAN=%s\n",slave->dev->name,+MOD_NA_STRING,MOD_NA_STRING);+}elseif(slave->vlan==VLAN_N_VID){+p+=_sprintf(p,buf,"SLAVE=%-10s MAC=%pM VLAN=%s\n",+slave->dev->name,+slave->emac,+MOD_NA_STRING);+}else{+p+=_sprintf(p,buf,"SLAVE=%-10s MAC=%pM VLAN=%d\n",+slave->dev->name,+slave->emac,+slave->vlan);+}+}+read_unlock_bh(&parent->lock);++_end_of_line(p,buf);++return(ssize_t)(p-buf);+}++staticDEVICE_ATTR(vifs,S_IRUGO,parent_show_vifs,+NULL);++staticssize_tparent_show_slaves(structdevice*d,+structdevice_attribute*attr,char*buf)+{+structslave*slave;+structparent*parent=to_parent(d);+char*p=buf;++read_lock_bh(&parent->lock);+parent_for_each_slave(parent,slave)+p+=_sprintf(p,buf,"%s\n",slave->dev->name);+read_unlock_bh(&parent->lock);++_end_of_line(p,buf);++return(ssize_t)(p-buf);+}++staticDEVICE_ATTR(slaves,S_IRUGO,parent_show_slaves,+NULL);++/* sysfs create/destroy functions */+staticstructattribute*per_parent_attrs[]={+&dev_attr_slaves.attr,/* DEVICE_ATTR(slaves..) */+&dev_attr_vifs.attr,+&dev_attr_neighs.attr,+NULL,+};++/* name spcase support */+staticconstvoid*eipoib_namespace(structclass*cls,+conststructclass_attribute*attr)+{+conststructeipoib_net*eipoib_n=+container_of(attr,+structeipoib_net,class_attr_eipoib_interfaces);+returneipoib_n->net;+}++staticstructattribute_groupparent_group={+/* per parent sysfs files under: /sys/class/net/<IF>/eth/.. */+.name="eth",+.attrs=per_parent_attrs+};++intcreate_slave_symlinks(structnet_device*master,+structnet_device*slave)+{+charlinkname[IFNAMSIZ+7];+intret=0;++ret=sysfs_create_link(&(slave->dev.kobj),&(master->dev.kobj),+"eth_parent");+if(ret)+returnret;++sprintf(linkname,"slave_%s",slave->name);+ret=sysfs_create_link(&(master->dev.kobj),&(slave->dev.kobj),+linkname);+returnret;++}++voiddestroy_slave_symlinks(structnet_device*master,+structnet_device*slave)+{+charlinkname[IFNAMSIZ+7];++sysfs_remove_link(&(slave->dev.kobj),"eth_parent");+sprintf(linkname,"slave_%s",slave->name);+sysfs_remove_link(&(master->dev.kobj),linkname);+}++staticstructclass_attributeclass_attr_eth_ipoib_interfaces={+.attr={+.name="eth_ipoib_interfaces",+.mode=S_IWUSR|S_IRUGO,+},+.show=show_parents,+.namespace=eipoib_namespace,+};++/* per module sysfs file under: /sys/class/net/eth_ipoib_interfaces */+intmod_create_sysfs(structeipoib_net*eipoib_n)+{+intrc;+/* defined in CLASS_ATTR(eth_ipoib_interfaces..) */+eipoib_n->class_attr_eipoib_interfaces=+class_attr_eth_ipoib_interfaces;++sysfs_attr_init(&eipoib_n->class_attr_eipoib_interfaces.attr);++rc=netdev_class_create_file(&eipoib_n->class_attr_eipoib_interfaces);+if(rc)+pr_err("%s failed to create sysfs (rc %d)\n",+eipoib_n->class_attr_eipoib_interfaces.attr.name,rc);++returnrc;+}++voidmod_destroy_sysfs(structeipoib_net*eipoib_n)+{+netdev_class_remove_file(&eipoib_n->class_attr_eipoib_interfaces);+}++intparent_create_sysfs_entry(structparent*parent)+{+structnet_device*dev=parent->dev;+intrc;++rc=sysfs_create_group(&(dev->dev.kobj),&parent_group);+if(rc)+pr_info("failed to create sysfs group\n");++returnrc;+}++voidparent_destroy_sysfs_entry(structparent*parent)+{+structnet_device*dev=parent->dev;++sysfs_remove_group(&(dev->dev.kobj),&parent_group);+}
From: Ben Hutchings <hidden> Date: 2012-08-02 00:17:07
On Wed, 2012-08-01 at 20:09 +0300, Or Gerlitz wrote:
quoted hunk
From: Erez Shitrit <redacted>
The Ethernet IPoIB driver enslaves IPoIB devices and uses them as
VIFs (Virtual Interface) which serve an Ethernet NIC e.g present in a
guest OS. For each such slave that acts as a VIF, eIPoIB needs to know
the mac and optionally the vlan uses by that NIC, the new ndo opertaion
is used to associate the mac/vlan for that slave.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
include/linux/netdevice.h | 5 ++++-
1 files changed, 4 insertions(+), 1 deletions(-)
The semantics of this operation should be documented in the comment
above the structure definition. One detail worth covering is whether
'vlan' is just a VID or can also include priority+CFI bits.
If this is specific to eIPoIB, why not put that in the name of the
operation? If not, this *really* needs explaining because so far I have
no whether it is something I should consider implementing on a real
Ethernet device.
Ben.
int (*ndo_fdb_add)(struct ndmsg *ndm,
struct net_device *dev,
unsigned char *addr,
--
Ben Hutchings, Staff Engineer, Solarflare
Not speaking for my employer; that's the marketing department's job.
They asked us to note that Solarflare product names are trademarked.
[...]
if_nlmsg_size() returns the size of a message describing the interface.
But IFLA_VIF_INFO is write-only (why?) and therefore shouldn't be
included.
Ben.
--
Ben Hutchings, Staff Engineer, Solarflare
Not speaking for my employer; that's the marketing department's job.
They asked us to note that Solarflare product names are trademarked.
These must be null-terminated; therefore use strlcpy().
+ /* indicates ABI version */
+ snprintf(drvinfo->fw_version, 32, "%d", EIPOIB_ABI_VER);
[...]
This is an abuse of fw_version.
Ben.
--
Ben Hutchings, Staff Engineer, Solarflare
Not speaking for my employer; that's the marketing department's job.
They asked us to note that Solarflare product names are trademarked.
On Wed, 2012-08-01 at 20:09 +0300, Or Gerlitz wrote:
quoted
From: Erez Shitrit <redacted>
The Ethernet IPoIB driver enslaves IPoIB devices and uses them as
VIFs (Virtual Interface) which serve an Ethernet NIC e.g present in a
guest OS. For each such slave that acts as a VIF, eIPoIB needs to know
the mac and optionally the vlan uses by that NIC, the new ndo opertaion
is used to associate the mac/vlan for that slave.
Signed-off-by: Erez Shitrit <redacted>
Signed-off-by: Or Gerlitz <redacted>
---
include/linux/netdevice.h | 5 ++++-
1 files changed, 4 insertions(+), 1 deletions(-)
The semantics of this operation should be documented in the comment
above the structure definition. One detail worth covering is whether
'vlan' is just a VID or can also include priority+CFI bits.
We will add more documentation for that.
The idea was just for the VID, (in our driver at least)
If this is specific to eIPoIB, why not put that in the name of the
operation? If not, this *really* needs explaining because so far I have
no whether it is something I should consider implementing on a real
Ethernet device.
Ben.
Will add more documentation here, perhaps other drivers can use it as
well for visualization uses and more.
Thanks.
quoted
int (*ndo_fdb_add)(struct ndmsg *ndm,
struct net_device *dev,
unsigned char *addr,
These must be null-terminated; therefore use strlcpy().
ok, will fix.
quoted
+ /* indicates ABI version */
+ snprintf(drvinfo->fw_version, 32, "%d", EIPOIB_ABI_VER);
[...]
This is an abuse of fw_version.
Ben.
we took the idea from the bonding driver,
(snprintf(drvinfo->fw_version, 32, "%d", BOND_ABI_VERSION);)
Do you have any idea where can we keep the abi version?
Thanks, Erez
[...]
if_nlmsg_size() returns the size of a message describing the interface.
But IFLA_VIF_INFO is write-only (why?) and therefore shouldn't be
included.
Ben.
thanks, we missed that.
currently, we don't expose the vifs on its parent interface, we need to
think if we want to expose it, if yes we will add "get" option,
otherwise we will take it from if_nlmsg_size() function.
These must be null-terminated; therefore use strlcpy().
ok, will fix.
quoted
quoted
+ /* indicates ABI version */
+ snprintf(drvinfo->fw_version, 32, "%d", EIPOIB_ABI_VER);
[...]
This is an abuse of fw_version.
Ben.
we took the idea from the bonding driver,
(snprintf(drvinfo->fw_version, 32, "%d", BOND_ABI_VERSION);)
The bonding driver has lots of warts.
Do you have any idea where can we keep the abi version?
You don't need to, because David will insist that you will only change
the ABI in a backward-compatible way. :-)
(The bonding ABI version hasn't changed since Linux 2.6.3, and even then
it provided backward compatibility.)
Ben.
--
Ben Hutchings, Staff Engineer, Solarflare
Not speaking for my employer; that's the marketing department's job.
They asked us to note that Solarflare product names are trademarked.
From: Eric W. Biederman <hidden> Date: 2012-08-02 17:15:34
Or Gerlitz [off-list ref] writes:
From: Erez Shitrit <redacted>
The eipoib driver provides a standard Ethernet netdevice over
the InfiniBand IPoIB interface .
Some services can run only on top of Ethernet L2 interfaces, and cannot be
bound to an IPoIB interface. With this new driver, these services can run
seamlessly.
Do I read this code correctly that what you are doing is not tunneling
ethernet over IB but instead you are removing an ethernet header and
replacing it with an IB header?
Do I also read this code correctly if you can't find your destination
mac address in your ""neighbor table"" you do a normal IPoIB arp
for the infiniband GUID?
Do I read this right that if presented with a non-IPv4 or ARP packet
this code will do something undefined and unpredictable?
Maybe this makes some sense but just skimming it looks like you
are trying to force a square peg into a round hole resulting in
some weird code and some very weird maintainability issues.
I am honestly surprised at this approach. I would think it would be
faster and simpler to run an IB queue pair directly to the hypervisor or
possibly even the guest operating system bypassing the kernel and doing
all of this translation in userspace.
Eric
From: Ali Ayoub <hidden> Date: 2012-08-03 20:38:18
On 8/2/2012 10:15 AM, Eric W. Biederman wrote:
Or Gerlitz [off-list ref] writes:
quoted
From: Erez Shitrit <redacted>
The eipoib driver provides a standard Ethernet netdevice over
the InfiniBand IPoIB interface .
Some services can run only on top of Ethernet L2 interfaces, and cannot be
bound to an IPoIB interface. With this new driver, these services can run
seamlessly.
Do I read this code correctly that what you are doing is not tunneling
ethernet over IB but instead you are removing an ethernet header and
replacing it with an IB header?
Correct.
eIPoIB runs standard IPoIB on the wire, thus it doesn't encapsulate the
Ethernet frame on top of IPoIB, but rather translates it to an IPoIB
packet, this allows us to expose an Ethernet L2 network device, and
still keep interoperability with existing IPoIB endpoints. Running full
encapsulation (i.e. EoIPoIB) will break interoperability.
Do I also read this code correctly if you can't find your destination
mac address in your ""neighbor table"" you do a normal IPoIB arp
for the infiniband GUID?
Correct.
Wire protocol remains IPoIB.
Do I read this right that if presented with a non-IPv4 or ARP packet
this code will do something undefined and unpredictable?
The current code drops IPv6 packets (see IS_E_IPOIB_PROTO), IPv6 support
will be added later on.
Maybe this makes some sense but just skimming it looks like you
are trying to force a square peg into a round hole resulting in
some weird code and some very weird maintainability issues.
I am honestly surprised at this approach. I would think it would be
faster and simpler to run an IB queue pair directly to the hypervisor or
possibly even the guest operating system bypassing the kernel and doing
all of this translation in userspace.
With eIPoIB architecture, the VM sees standard Ethernet emulator,
allowing the administrator to enslave eIPoIB PIF to the vSwitch/vBridge
as if it was standard Ethernet. Other approaches that exposes IB QP to
the VM (with w/o bypassing the kernel) won't be possible with the
current emulators and management tools.
--
Ali;
From: David Miller <davem@davemloft.net> Date: 2012-08-03 21:33:16
From: Ali Ayoub <redacted>
Date: Fri, 03 Aug 2012 13:31:35 -0700
With eIPoIB architecture, the VM sees standard Ethernet emulator,
allowing the administrator to enslave eIPoIB PIF to the vSwitch/vBridge
as if it was standard Ethernet. Other approaches that exposes IB QP to
the VM (with w/o bypassing the kernel) won't be possible with the
current emulators and management tools.
So then fix the emulators and management tools to handle IB instead
of adding this bogus new protocol?
This new protocol seems to exist only because you don't want to have
to enhance the emulators and tools, and I'm sorry that isn't a valid
reason to do something like this.
From: Ali Ayoub <hidden> Date: 2012-08-03 22:39:38
On 8/3/2012 2:33 PM, David Miller wrote:
From: Ali Ayoub <redacted>
Date: Fri, 03 Aug 2012 13:31:35 -0700
quoted
With eIPoIB architecture, the VM sees standard Ethernet emulator,
allowing the administrator to enslave eIPoIB PIF to the vSwitch/vBridge
as if it was standard Ethernet. Other approaches that exposes IB QP to
the VM (with w/o bypassing the kernel) won't be possible with the
current emulators and management tools.
So then fix the emulators and management tools to handle IB instead
of adding this bogus new protocol?
This new protocol seems to exist only because you don't want to have
to enhance the emulators and tools, and I'm sorry that isn't a valid
reason to do something like this.
This driver exists to allow the user to have an Ethernet interface on
top of a high-speed InfiniBand (IB) interconnect.
Users would like to use sockets API from the VM without re-writing their
applications on top of IB verbs, this driver meant to allow such a user
to do so.
Exposing IB emulators and having IB support in the management tools for
the VM/Hypervisor won't address the usecases that this driver meant for.
With this driver, existing VMs, and their existing IP applications, can
run as-is on InfiniBand network.
From: David Miller <davem@davemloft.net> Date: 2012-08-03 23:36:36
From: Ali Ayoub <redacted>
Date: Fri, 03 Aug 2012 15:39:36 -0700
Users would like to use sockets API from the VM without re-writing their
applications on top of IB verbs, this driver meant to allow such a user
to do so.
That's what IPoIB was for, the application writers who don't want to have
to be knowledgable about IB verbs.
You're messing with the link layer here, and that's what is upsetting me.
It's a complete cop-out to changing the VM tools and emulators properly to
handle a new link layer.
The applications writers already have a way to use IB whilst using
something familiar, like IPv4, via IPoIB. You're doing something
completely different here, and it stinks.
From: Ali Ayoub <hidden> Date: 2012-08-04 00:02:37
On 8/3/2012 3:39 PM, Ali Ayoub wrote:
On 8/3/2012 2:33 PM, David Miller wrote:
quoted
From: Ali Ayoub <redacted>
Date: Fri, 03 Aug 2012 13:31:35 -0700
quoted
With eIPoIB architecture, the VM sees standard Ethernet emulator,
allowing the administrator to enslave eIPoIB PIF to the vSwitch/vBridge
as if it was standard Ethernet. Other approaches that exposes IB QP to
the VM (with w/o bypassing the kernel) won't be possible with the
current emulators and management tools.
So then fix the emulators and management tools to handle IB instead
of adding this bogus new protocol?
This new protocol seems to exist only because you don't want to have
to enhance the emulators and tools, and I'm sorry that isn't a valid
reason to do something like this.
This driver exists to allow the user to have an Ethernet interface on
top of a high-speed InfiniBand (IB) interconnect.
Users would like to use sockets API from the VM without re-writing their
applications on top of IB verbs, this driver meant to allow such a user
to do so.
Exposing IB emulators and having IB support in the management tools for
the VM/Hypervisor won't address the usecases that this driver meant for.
With this driver, existing VMs, and their existing IP applications, can
run as-is on InfiniBand network.
This driver exists to allow the user to have an Ethernet interface on
top of a high-speed InfiniBand (IB) interconnect.
Users would like to use sockets API from the VM without re-writing their
applications on top of IB verbs, this driver meant to allow such a user
to do so.
Exposing IB emulators and having IB support in the management tools for
the VM/Hypervisor won't address the usecases that this driver meant for.
With this driver, existing VMs, and their existing IP applications, can
run as-is on InfiniBand network.
From: David Miller <davem@davemloft.net> Date: 2012-08-04 00:05:59
From: Ali Ayoub <redacted>
Date: Fri, 03 Aug 2012 17:02:35 -0700
On 8/3/2012 3:39 PM, Ali Ayoub wrote:
quoted
On 8/3/2012 2:33 PM, David Miller wrote:
quoted
From: Ali Ayoub <redacted>
Date: Fri, 03 Aug 2012 13:31:35 -0700
quoted
With eIPoIB architecture, the VM sees standard Ethernet emulator,
allowing the administrator to enslave eIPoIB PIF to the vSwitch/vBridge
as if it was standard Ethernet. Other approaches that exposes IB QP to
the VM (with w/o bypassing the kernel) won't be possible with the
current emulators and management tools.
So then fix the emulators and management tools to handle IB instead
of adding this bogus new protocol?
This new protocol seems to exist only because you don't want to have
to enhance the emulators and tools, and I'm sorry that isn't a valid
reason to do something like this.
This driver exists to allow the user to have an Ethernet interface on
top of a high-speed InfiniBand (IB) interconnect.
Users would like to use sockets API from the VM without re-writing their
applications on top of IB verbs, this driver meant to allow such a user
to do so.
Exposing IB emulators and having IB support in the management tools for
the VM/Hypervisor won't address the usecases that this driver meant for.
With this driver, existing VMs, and their existing IP applications, can
run as-is on InfiniBand network.
This driver exists to allow the user to have an Ethernet interface on
top of a high-speed InfiniBand (IB) interconnect.
Users would like to use sockets API from the VM without re-writing their
applications on top of IB verbs, this driver meant to allow such a user
to do so.
Exposing IB emulators and having IB support in the management tools for
the VM/Hypervisor won't address the usecases that this driver meant for.
With this driver, existing VMs, and their existing IP applications, can
run as-is on InfiniBand network.
Just saying the same thing twice doesn't make your argument stronger.
From: Eric W. Biederman <hidden> Date: 2012-08-04 01:35:02
Ali Ayoub [off-list ref] wrote:
With this driver, existing VMs, and their existing IP applications, can
run as-is on InfiniBand network.
Actually it doesn't work like that.
If my application really needs ethernet it will not work with your driver. I can not run decnet or appletalk or ATAoE or PPPoE or LACP or use VLANs or any of a thousand other things that require real live ethernet to function.
Most VMs will talk to linux with a tap interface, and the output of a tap interface can be routed just fine, and that works with existing tools, and existing already deployed kernels.
Alternatively if you are silly you can implement a tap interface on top of IPoIB and have the same interface you do now, and it works with existing already deployed kernels.
So eIPoIB does not make sense from a time to market perspective.
Similarly eIPoIB does not make sense from a performance standpoint because the best performance requires the applications and hypervisors are infiniband and teaching the apps to cope.
eIPoIB also has considerable maintenance overhead as it is complex code doing some crazy things.
Now personally NAPT44 just about made the internet unusable for peer to peer appications. NAT66 aka network prefix translation introduces much less breakage but it still requires someone to run STUN on IPv6 to keep applications working, ick. NATEIB aka eIPoIB just looks complex and already breaks most of ethernet and history says that kind of breakage actually matters a lot. I think everyone who suggests any kind of NAT is a good idea should have to implement RFC 5245 Interacive Connectivity Establishment. It takes 2 100+ page rfcs to come up with a way that allows applications to with the address munging over the open internet. That way takes 3000+ lines of code in each application. That is the kind of complexity you are asking people up stream of you to deal with your NATTing o
f ethernet to infiniband. So I totally think eIPoIB stinks, and that it does not at all let existing applications work without modification.
Eric
From: Or Gerlitz <hidden> Date: 2012-08-04 21:23:21
On Sat, Aug 4, 2012 at 2:36 AM, David Miller [off-list ref] wrote:
From: Ali Ayoub <redacted>
Date: Fri, 03 Aug 2012 15:39:36 -0700
quoted
Users would like to use sockets API from the VM without re-writing their
applications on top of IB verbs, this driver meant to allow such a user
to do so.
That's what IPoIB was for, the application writers who don't want to have
to be knowledgable about IB verbs.
You're messing with the link layer here, and that's what is upsetting me.
It's a complete cop-out to changing the VM tools and emulators properly to
handle a new link layer.
The applications writers already have a way to use IB whilst using
something familiar, like IPv4, via IPoIB. You're doing something
completely different here, and it stinks.
Dave,
Just quick recap of things to make sure we're in sync on the facts
here and (me) trying to better follow your arguments:
Bringing IB to app usage is done through the IB verbs / RDMA CM, which
is supported
for quite a long while for native mode and BTW as for SRIOV, there's a
now running patch set @ linux-rdma submitted to Roland
http://marc.info/?l=linux-rdma&m=134398354428293&w=2
IPoIB is meant to allow IP apps to use IB, defined by IETF spec (RFC 4391/4392)
The idea in eIPoIB was to allow IP apps running on VMs under a
Para-Virtual set of mind, e.g when the Linux PV networking stack comes
into play, to use that stack w.o modifying it.
When looking on that, we thought so far so good, and went in the way
posted here. If reusing your last sentence... this driver provides a
way for apps to use the PV stack AND IB whilst using something
familiar, like IPv4.
I'd like to better understand your "messing with the link layer here,
and that's what is upsetting" claim --- since the service here is IP,
and this service has well defined mechanisms, eIPoIB is basically a
shim layer that allows to provide this service for VMs that interacts
with the emulators, drivers and tools used in Linux and the related
guest OS PV drivers (virtio and friends). For example, that messing
mostly goes down to translating ARP requests/replies from being
carried in Ethernet frame to be carried in IPoIB packet, is that
really too hacky? why and what?
Or.
From: Or Gerlitz <hidden> Date: 2012-08-04 21:33:09
On Sat, Aug 4, 2012 at 4:34 AM, Eric W. Biederman [off-list ref] wrote:
Ali Ayoub [off-list ref] wrote:
quoted
With this driver, existing VMs, and their existing IP applications, can
run as-is on InfiniBand network.
Actually it doesn't work like that.
If my application really needs ethernet it will not work with your driver.
Eric,
Starting to address your email
If your application really needs Ethernet, it will not work with
IPoIB, since the latter only supports IP apps. So this isn't the
problem we're trying to solve here, but rather allow IP app running in
PV mode to use IPoIB.
I can not run decnet or appletalk or ATAoE or PPPoE or LACP
[...]
or use VLANs
YES you can, vlan devices can be set over eipoib devices, and the
eipoib driver maps the vlan ID to IB Partition using a mapping defined
by the admin.
or any of a thousand other things that require real live ethernet to function.
[...]
Similarly eIPoIB does not make sense from a performance standpoint because
the best performance requires the applications and hypervisors are infiniband
and teaching the apps to cope.
agree 100% re the "best performance" part of this sentence, if your
app looks for best perf, yes you can use IB in either native or SRIOV
mode, as I mentioned in my response to Dave.
Or.
Or.
From: Or Gerlitz <hidden> Date: 2012-08-04 21:44:06
On Sun, Aug 5, 2012 at 12:23 AM, Or Gerlitz [off-list ref] wrote:
The idea in eIPoIB was to allow IP apps running on VMs under a Para-Virtual set of
mind, e.g when the Linux PV networking stack comes into play, to use that stack w.o > modifying it. When looking on that, we thought so far so good, and went in the way
posted here. If reusing your last sentence... this driver provides a way for apps to use
the PV stack AND IB whilst using something familiar, like IPv4.
OK, when I said the PV stack, I meant the portion of the PV stack
which assumes Ethernet link layer, ofcourse... If someone uses routing
they don't need this driver.
Again, since the app only uses IP, which is well defined, etc. the
work done by the eipoib driver, didn't seem as hackish messing, so I'm
again with that WW (Why/What) question from 5m ago.
Or.
From: Eric W. Biederman <hidden> Date: 2012-08-04 23:19:17
Or Gerlitz [off-list ref] wrote:
On Sun, Aug 5, 2012 at 12:23 AM, Or Gerlitz [off-list ref]
wrote:
quoted
The idea in eIPoIB was to allow IP apps running on VMs under a
Para-Virtual set of
quoted
mind, e.g when the Linux PV networking stack comes into play, to use
that stack w.o > modifying it. When looking on that, we thought so far
so good, and went in the way
quoted
posted here. If reusing your last sentence... this driver provides a
way for apps to use
quoted
the PV stack AND IB whilst using something familiar, like IPv4.
OK, when I said the PV stack, I meant the portion of the PV stack
which assumes Ethernet link layer, ofcourse... If someone uses routing
they don't need this driver.
Again, since the app only uses IP, which is well defined, etc. the
work done by the eipoib driver, didn't seem as hackish messing, so I'm
again with that WW (Why/What) question from 5m ago.
Something is missing from your sentence.
What is the alternative that you view as worse?
Juast as a point of information. In general when bridging is desired but not possible people deploy proxy arp.
I completely fail to see how having the VM output to a tun interface and then routing that, would not be supported by current solutions. I believe that is where VM solutions all started networking wise. Outputting to an interface is needed to support interfaces like 802.11 where bridging frequently does not work.
Eric
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-05 18:50:00
On Thu, Aug 02, 2012 at 10:15:23AM -0700, Eric W. Biederman wrote:
Or Gerlitz [off-list ref] writes:
quoted
From: Erez Shitrit <redacted>
The eipoib driver provides a standard Ethernet netdevice over
the InfiniBand IPoIB interface .
Some services can run only on top of Ethernet L2 interfaces, and cannot be
bound to an IPoIB interface. With this new driver, these services can run
seamlessly.
Do I read this code correctly that what you are doing is not tunneling
ethernet over IB but instead you are removing an ethernet header and
replacing it with an IB header?
Do I also read this code correctly if you can't find your destination
mac address in your ""neighbor table"" you do a normal IPoIB arp
for the infiniband GUID?
Do I read this right that if presented with a non-IPv4 or ARP packet
this code will do something undefined and unpredictable?
Maybe this makes some sense but just skimming it looks like you
are trying to force a square peg into a round hole resulting in
some weird code and some very weird maintainability issues.
I am honestly surprised at this approach. I would think it would be
faster and simpler to run an IB queue pair directly to the hypervisor or
possibly even the guest operating system bypassing the kernel and doing
all of this translation in userspace.
Eric
I'm on vacation and I have not looked at the patches, at Erez' request,
just reacting to the presentation and the discussion.
Bypassing the kernel has its own set of issues, not the
least of which is the need to lock all of guest memory which breaks
overcommit. Running an IB queue pair directly to the hypervisor
will also break live migration.
Another problem with exposing IB to guests has to do with the fact that
IB addresses such as combinations of LIDs, GIDs and QPNs to best of my
knowledge do not support soft hardware address setting, which interferes
with live migration.
So it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes
sense.
So far the things that make sense. Here are some that don't, to me:
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
Other things that I would expect to be addressed in such a spec is
interaction with other IPoIB features, such as connected
mode, checksum offloading etc, and IB features such as multipath etc.
- The way you encode LID/QPN in the MAC seems questionable. IIRC there's
more to IB addressing than just the LID. Since everyone on the subnet
need access to this translation, I think it makes sense to store it in
the SM. I think this would also obviate some IPv4 specific hacks
in kernel.
- IGMP/MAC snooping in a driver is just too hairy.
As you point out, bridge currently needs the uplink in promisc mode.
I don't think a driver should work around that limitation.
For some setups, it might be interesting to remove the
promisc mode requirement, failing that,
I think you could use macvtap passthrough.
- Currently migration works without host kernel help, would be
preferable to keep it that way.
Hope this helps,
MST
--
To unsubscribe from this list: send the line "unsubscribe netdev" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
From: Ali Ayoub <hidden> Date: 2012-08-07 00:14:27
On 8/3/2012 4:36 PM, David Miller wrote:
From: Ali Ayoub <redacted>
Date: Fri, 03 Aug 2012 15:39:36 -0700
quoted
Users would like to use sockets API from the VM without re-writing their
applications on top of IB verbs, this driver meant to allow such a user
to do so.
That's what IPoIB was for, the application writers who don't want to have
to be knowledgable about IB verbs.
You're messing with the link layer here, and that's what is upsetting me.
It's a complete cop-out to changing the VM tools and emulators properly to
handle a new link layer.
The applications writers already have a way to use IB whilst using
something familiar, like IPv4, via IPoIB. You're doing something
completely different here, and it stinks.
Indeed, IPoIB driver meant to allow IP applications to run over
InfiniBand, but IPoIB cannot serve IPoE traffic.
The goal of eIPoIB driver to show to the user an Ethernet L2 netdev; by
translating IPoE packets to IPoIB. It keeps the same IPoIB wire
protocol, and exposes to the host an ethX interface, all IPoIB packets
are then translated to IPoIB.
Among other things, the main benefit we're targeting is to allow IPoE
traffic within the VM to go through the (Ethernet) vBridge down to the
eIPoIB PIF, and eventually to IPoIB and to the IB network.
In Para virtualized environment, the VM emulator sends/receives packets
with Ethernet header, and the vBridge also performs L2 switching based
on the Ethernet header, in addition to other tools that expect an
Ethernet link layer. We'd like to support them on top of IPoIB.
I see your point to change the the tools and VM emulations to handle
IPoIB link layer, but this involves not only changing many
components/layers such netfront/vbridge/vconfig/etc.. but also -unlike
Ethernet- having IPoIB-aware network device emulation in the VM domain
requires giving access from the VM to the IB Hardware, and this
association will break upon VM migration, as Tsirkin indicated in the
other Email.
Another issue with IPoIB-aware emulator, is that the IPoIB frame doesn't
include the destination link layer per data packet (RFC 4391), therefor
the neighbor address needs to be passed from the VM domain down to the
ipoib-aware vBridge.
I don't see in other alternatives a solution for the problem we're
trying to solve. If there are changes/suggestions to improve eIPoIB
netdev driver to avoid "messing with the link layer" and make it
acceptable, we can discuss and apply them.
From: Eric W. Biederman <hidden> Date: 2012-08-07 00:44:18
Ali Ayoub [off-list ref] writes:
Among other things, the main benefit we're targeting is to allow IPoE
traffic within the VM to go through the (Ethernet) vBridge down to the
eIPoIB PIF, and eventually to IPoIB and to the IB network.
That works today without code changes. It is called routing.
In Para virtualized environment, the VM emulator sends/receives packets
with Ethernet header, and the vBridge also performs L2 switching based
on the Ethernet header, in addition to other tools that expect an
Ethernet link layer. We'd like to support them on top of IPoIB.
See routing. The code is already done.
I don't see in other alternatives a solution for the problem we're
trying to solve. If there are changes/suggestions to improve eIPoIB
netdev driver to avoid "messing with the link layer" and make it
acceptable, we can discuss and apply them.
Nothing needs to be applied the code is done. Routing from
IPoE to IPoIB works.
There is nothing in what anyone has posted as requirements that needs
work to implement.
I totally fail to see how getting packets of of the VM as ethernet
frames, and then IP layer routing those packets over IP is not an
option. What requirement am I missing.
All VMs should suport that mode of operation, and certainly the kernel
does.
Implementations involving bridges like macvlan and macvtap are
performance optimizations, and the optimizations don't even apply in
areas like 802.11, where only one mac address is supported per adapter.
Bridging can ocassionally also be an administrative simplification as
well, but you should be able to achieve the a similar simplification
with a dhcprelay and proxy arp.
Eric
Hi Eric, Ali
On Mon, 06 Aug 2012 17:44:06 -0700
ebiederm@xmission.com (Eric W. Biederman) wrote:
Implementations involving bridges like macvlan and macvtap are
performance optimizations, and the optimizations don't even apply in
areas like 802.11, where only one mac address is supported per adapter.
We had nearly equal situation perfomance result. share for you.
-----------------------------------------------------------------------
Ethernet over EoIB using gretap with "Raid STP" performance result.
http://twitpic.com/agd1pp
How to use EoIB with gretap.
http://twitpic.com/agd4io
ifconfig ib0 10.1.1.1/24 up
ip link add eoib0 type gretap remote 10.1.1.2 local 10.1.1.1
ip link add eoib1 type gretap remote 10.1.1.3 local 10.1.1.1
ip link add type veth
ifconfig eoib0 up up
ifconfig eoib1 up up
ifconfig veth0 up up
ifconfig veth1 up up
brctl addbr br0
brctl addif br0 eoib0
brctl addif br0 eoib1
bcrtl addif br0 veth1
rstpd
rstpctl rstp br0 on
rstpctl setbridgeprio br0 40960
rstpctl setportpathcost br0 vpn1 40000
rstpctl sethello br0 1
rstpctl setmaxage br0 6
rstpctl setfdelay br0 4
brctl stp br0 on
brctl sethello br0 1
brctl setfd br0 4
brctl setmaxage br0 6
-----------------------------------------------------------------------
regards,
--
SAKURA Internet Inc. / Senior Researcher
Naoto MATSUMOTO [off-list ref]
SAKURA Research Center <http://research.sakura.ad.jp/>
From: Eric W. Biederman <hidden> Date: 2012-08-07 03:33:19
ebiederm@xmission.com (Eric W. Biederman) writes:
Ali Ayoub [off-list ref] writes:
quoted
Among other things, the main benefit we're targeting is to allow IPoE
traffic within the VM to go through the (Ethernet) vBridge down to the
eIPoIB PIF, and eventually to IPoIB and to the IB network.
Oh yes. It just occurred to me there is huge problem with eIPoIB as
currently presented in these patches. It breaks DHCPv4 the same way
it breaks ARP, but DHCPv4 is not fixed up.
I am stunned to to realize there has been so much push for a solution
that doesn't even support dhcp.
How can anyone possibly argue that eIPoIB works or is useful? I just
can't imagine.
Totally stunned,
Eric
From: Joseph Glanville <hidden> Date: 2012-08-07 03:37:10
On 7 August 2012 10:44, Eric W. Biederman [off-list ref] wrote:
Ali Ayoub [off-list ref] writes:
quoted
Among other things, the main benefit we're targeting is to allow IPoE
traffic within the VM to go through the (Ethernet) vBridge down to the
eIPoIB PIF, and eventually to IPoIB and to the IB network.
That works today without code changes. It is called routing.
quoted
In Para virtualized environment, the VM emulator sends/receives packets
with Ethernet header, and the vBridge also performs L2 switching based
on the Ethernet header, in addition to other tools that expect an
Ethernet link layer. We'd like to support them on top of IPoIB.
See routing. The code is already done.
quoted
I don't see in other alternatives a solution for the problem we're
trying to solve. If there are changes/suggestions to improve eIPoIB
netdev driver to avoid "messing with the link layer" and make it
acceptable, we can discuss and apply them.
Nothing needs to be applied the code is done. Routing from
IPoE to IPoIB works.
There is nothing in what anyone has posted as requirements that needs
work to implement.
I totally fail to see how getting packets of of the VM as ethernet
frames, and then IP layer routing those packets over IP is not an
option. What requirement am I missing.
All VMs should suport that mode of operation, and certainly the kernel
does.
Implementations involving bridges like macvlan and macvtap are
performance optimizations, and the optimizations don't even apply in
areas like 802.11, where only one mac address is supported per adapter.
Bridging can ocassionally also be an administrative simplification as
well, but you should be able to achieve the a similar simplification
with a dhcprelay and proxy arp.
Eric
--
To unsubscribe from this list: send the line "unsubscribe netdev" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Hi,
While I agree this driver doesn't offer an adequate solution to the
problem I think the problem still very much exists.
Ultimately if you have an Infiniband fabric and you want to use it for
inter-guest communication across hosts you are left with 2 options,
tunneling and routing.
Routing can be ok sometimes but it's not a great solution for the vast
majority of people, it requires considerable edge intelligence in the
hypervisor and ISIS routing setups.
The biggest issue is that it's just not L2, guest applications on
guest operating systems expect to have access to fully fledged
Ethernet devices.
This is especially an issue for hosting providers etc where they have
little to no control over what applications their customers want to
use.
Tunnelling solves some of these problems, generally done over L3 and
using existing IP fabric to transport Ethernet L2 frames to the
destination hypervisors.
Better because now you have true L2 encapsulation but to be honest
it's slow and chews alot of cycles and generally has poor PPS
performance.
IMO this driver doesn't make sense compared to a true Ethernet over IB
encapsulation, which isn't actually all that much extra work. (If I am
right you already have a driver that does just this)
The point about retaining compatibility with existing IPoIB boggles my
mind a little.
It means little if no benefit is really added because only IP and ARP
will work.. no other L2 protocol will work correctly as others have
noted.
With full encapsulation you can make use of all the existing
infrastructure, linux bridge, OpenvSwitch, ebtables/netfilter etc.
Joseph.
--
CTO | Orion Virtualisation Solutions | www.orionvm.com.au
Phone: 1300 56 99 52 | Mobile: 0428 754 846
From: Or Gerlitz <hidden> Date: 2012-08-08 05:23:16
On Sun, Aug 5, 2012 at 9:50 PM, Michael S. Tsirkin [off-list ref] wrote:
[...]
So it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes sense.
[...]
Yep, that's among the points we're trying to make, the way you've put
it makes it clearer.
So far the things that make sense. Here are some that don't, to me:
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
To be precise, the solution uses 100% IPoIB wire-protocol, so we don't
see a need
for any spec change / standardization effort. This might go to the 1st
point you've
brought... improve the documentation, will do that. The pdf you looked
at was presented
in a conference.
Other things that I would expect to be addressed in such a spec is
interaction with other IPoIB features, such as connected
mode, checksum offloading etc, and IB features such as multipath etc.
For the eipoib interface, it doesn't really matters if the underlyind
ipoib clones used by it (we call them VIFs) use connected or datagram
mode, what does matter is the MTU and offload features supported by
these VIFs, for which the eipoib interface will have the min among all
these VIFs. Since for a given eipoib nic, all its VIFs must originated
from the same IPoIB PIF (e.g ib0) its easy admin job to make sure they
all have the same mtu / features which are needed for that eipoib nic,
e.g by using the same mode (connected/datagram for all of them), hope
this is clear.
- The way you encode LID/QPN in the MAC seems questionable. IIRC there's
more to IB addressing than just the LID. Since everyone on the subnet
need access to this translation, I think it makes sense to store it in
the SM. I think this would also obviate some IPv4 specific hacks in kernel.
The idead beyond the encoding was uniqueness, LID/QPN is unique per IB
HCA end-node. I wasn't sure to understand the comment re the IPv4 hacks.
- IGMP/MAC snooping in a driver is just too hairy.
mmm, any rough idea/direction how to do that otherwise?
As you point out, bridge currently needs the uplink in promisc mode.
I don't think a driver should work around that limitation.
For some setups, it might be interesting to remove the
promisc mode requirement, failing that,
I think you could use macvtap passthrough.
That's in the plans, the current code doesn't assume that the eipoib
has bridge on top, for VM networking it works with bridge + tap,
bridge + macvtap, but it would easily work with passthrough when we
allow to create multiple eipoib interfaces on the same ipoib PIF (e.g
today for the ib0 PIF we create eipoib eth0, and then two VIFs ib0.1
and ib0.2 that are enslaved by eth0, but next we will create eth1 and
eth2 which will use ib0.1 and ib0.2
respectively.
- Currently migration works without host kernel help, would be
preferable to keep it that way.
From: Or Gerlitz <hidden> Date: 2012-08-08 06:04:56
Eric W. Biederman [off-list ref] wrote:
quoted
Ali Ayoub [off-list ref] writes:
quoted
Among other things, the main benefit we're targeting is to allow IPoE
traffic within the VM to go through the (Ethernet) vBridge down to the
eIPoIB PIF, and eventually to IPoIB and to the IB network.
Oh yes. It just occurred to me there is huge problem with eIPoIB as
currently presented in these patches. It breaks DHCPv4 the same way
it breaks ARP, but DHCPv4 is not fixed up.
To put things in place, DHCPv4 is supported with eIPoIB, the DHCP
UDP/IP payload isn't touched, only need to set the BOOTP broadcast
flag in the dhcp server config file.
Or.
From: Or Gerlitz <hidden> Date: 2012-08-08 07:32:48
Eric W. Biederman [off-list ref] wrote:
Ali Ayoub [off-list ref] writes:
[...]
quoted
I don't see in other alternatives a solution for the problem we're
trying to solve. If there are changes/suggestions to improve eIPoIB
netdev driver to avoid "messing with the link layer" and make it
acceptable, we can discuss and apply them.
Nothing needs to be applied the code is done. Routing from
IPoE to IPoIB works. There is nothing in what anyone has posted as requirements
that needs work to implement.
I totally fail to see how getting packets of of the VM as ethernet
frames, and then IP layer routing those packets over IP is not an
option. What requirement am I missing.
As you've indicated routing w/w.o using proxy-arp is an option, however,
All VMs should suport that mode of operation, and certainly the kernel does.
Implementations involving bridges like macvlan and macvtap are
performance optimizations, and the optimizations don't even apply in
areas like 802.11, where only one mac address is supported per adapter.
Bridging can ocassionally also be an administrative simplification as
well, but you should be able to achieve the a similar simplification
with a dhcprelay and proxy arp.
as you wrote here, when performance and ease-of-use is under the spot,
VM deployments tend to not to use routing.
This is b/c it involves more over-head on the packet forwarding, and
more administration work, for example, for setting routing rules that
involve the VM IP address, something which AFAIK the hypervisor have
no clue on, also its unclear to me if/how live migration can work in
such setting.
From this exact reason, there's a bunch of use-cases by tools and
cloud stacks (such as open stack, ovirt, more) which do use bridged
mode and the rest of the Ethernet envelope, such as using virtual L2
vlan domains, ebtables based rules, etc etc. Where they and are not
application to ipoib, but are working file ith eipoib.
You mentioned that bridging mode doesn't apply to environment such as
802.11, and hence routing mode is used, we are trying to make a point
here that bridging mode applies to ipoib with the approach suggested
by eipoib.
Also, if we extend the discussion a bit, there are two more aspects to throw in:
The first is the performance thing we have already started to mention
-- specifically, the approach for RX zero copy (into the VM buffer),
use designs such as vhost + macvtap NIC in passthrough mode which is
likey to be set over a per VM hypervisor NIC, e.g such as the ones
provided by VMDQ patches John Fastabend started to post (see
http://marc.info/?l=linux-netdev&m=134264998405581&w=2) -- the ib0.N
clone child are IPoIB VMDQ NICs if you like, and setting an eipoib NIC
on top of each they can be plugged to that design.
The 2nd aspect, is NON VM environments where a NIC with Ethernet look
and feel is required for IP traffic, but this have to live within an
echo-system that fully uses IPoIB.
In other words, a use case where IPoIB has to be below the cover for
set of some specific apps, or nodes but do IP interaction with other
apps/nodes and gateways who use IPoIB, the eIPoIB driver provides that
functionality.
So, to sum up, routing / proxy-arp seems to be off where we are targeting.
Or.
From: Eric W. Biederman <hidden> Date: 2012-08-08 08:45:33
Or Gerlitz [off-list ref] writes:
Eric W. Biederman [off-list ref] wrote:
quoted
quoted
Ali Ayoub [off-list ref] writes:
quoted
quoted
Among other things, the main benefit we're targeting is to allow IPoE
traffic within the VM to go through the (Ethernet) vBridge down to the
eIPoIB PIF, and eventually to IPoIB and to the IB network.
quoted
Oh yes. It just occurred to me there is huge problem with eIPoIB as
currently presented in these patches. It breaks DHCPv4 the same way
it breaks ARP, but DHCPv4 is not fixed up.
To put things in place, DHCPv4 is supported with eIPoIB, the DHCP
UDP/IP payload isn't touched, only need to set the BOOTP broadcast
flag in the dhcp server config file.
Wrong. DHCPv4 is broken over eIPoIB.
Coming from ethernet
htype == 1 not 32 as required by RFC4390
hlen == 6 not 0 as required by RFC4390
The chaddr field is has 6 bytes of the ethernet mac address not the
required 16 bytes of 0.
The client-identifier field is optional over ethernet.
An ethernet DHCPv4 client simply does not generate a dhcp packet that
conforms to RFC4390.
Therefore DHCPv4 over eIPoIB is broken, and a dhcp server or relay
may reasonably look at the DHCP packet and drop it because it is
garbage.
You might find a forgiving dhcp server that doesn't drop insane packets
on the floor and tries to make things work.
I am sorry. eIPoIB is broken as designed. eIPoIB most assuredly is not
compatible with ethernet. eIPoIB most definitely does not work even for
the general case of transporing IP traffic. Claiming that eIPoIB is any
else is a lie.
Eric
From: Eric W. Biederman <hidden> Date: 2012-08-08 09:17:44
Or Gerlitz [off-list ref] writes:
Eric W. Biederman [off-list ref] wrote:
quoted
Ali Ayoub [off-list ref] writes:
[...]
quoted
quoted
I don't see in other alternatives a solution for the problem we're
trying to solve. If there are changes/suggestions to improve eIPoIB
netdev driver to avoid "messing with the link layer" and make it
acceptable, we can discuss and apply them.
quoted
Nothing needs to be applied the code is done. Routing from
IPoE to IPoIB works. There is nothing in what anyone has posted as requirements
that needs work to implement.
quoted
I totally fail to see how getting packets of of the VM as ethernet
frames, and then IP layer routing those packets over IP is not an
option. What requirement am I missing.
As you've indicated routing w/w.o using proxy-arp is an option, however,
quoted
All VMs should suport that mode of operation, and certainly the kernel does.
Implementations involving bridges like macvlan and macvtap are
performance optimizations, and the optimizations don't even apply in
areas like 802.11, where only one mac address is supported per adapter.
Bridging can ocassionally also be an administrative simplification as
well, but you should be able to achieve the a similar simplification
with a dhcprelay and proxy arp.
as you wrote here, when performance and ease-of-use is under the spot,
VM deployments tend to not to use routing.
This is b/c it involves more over-head on the packet forwarding, and
more administration work, for example, for setting routing rules that
involve the VM IP address, something which AFAIK the hypervisor have
no clue on, also its unclear to me if/how live migration can work in
such setting.
All you need to make proxy-arp essentially pain free is a smart dhcp
relay, that sets up the routes.
From this exact reason, there's a bunch of use-cases by tools and
cloud stacks (such as open stack, ovirt, more) which do use bridged
mode and the rest of the Ethernet envelope, such as using virtual L2
vlan domains, ebtables based rules, etc etc. Where they and are not
application to ipoib, but are working file ith eipoib.
Yes I am certain all of their IPv6 traffic works fine.
Regardless those are open source projects and can be modified to add
support to cleanly support inifinibnad.
You mentioned that bridging mode doesn't apply to environment such as
802.11, and hence routing mode is used, we are trying to make a pointn
here that bridging mode applies to ipoib with the approach suggested
by eipoib.
You are completely failing. Every time I look I see something about
eIPoIB that is even more broken. Given that eIPoIB is a NAT
implementation that isn't really a surprise but still.
eIPoIB imposes enough overhead that I expect that routing is cheaper,
so your performance advantges go right out the window.
eIPoIB is seriously incompatible with ethernet breaking almost
everything and barely allowing IPv4 to work.
Also, if we extend the discussion a bit, there are two more aspects to throw in:
The first is the performance thing we have already started to mention
-- specifically, the approach for RX zero copy (into the VM buffer),
use designs such as vhost + macvtap NIC in passthrough mode which is
likey to be set over a per VM hypervisor NIC, e.g such as the ones
provided by VMDQ patches John Fastabend started to post (see
http://marc.info/?l=linux-netdev&m=134264998405581&w=2) -- the ib0.N
clone child are IPoIB VMDQ NICs if you like, and setting an eipoib NIC
on top of each they can be plugged to that design.
If you care about performance link-layer NAT is not the way to go.
Teach the pieces you care about how to talk infiniband.
The 2nd aspect, is NON VM environments where a NIC with Ethernet look
and feel is required for IP traffic, but this have to live within an
echo-system that fully uses IPoIB.
In other words, a use case where IPoIB has to be below the cover for
set of some specific apps, or nodes but do IP interaction with other
apps/nodes and gateways who use IPoIB, the eIPoIB driver provides that
functionality.
ip link add type dummy.
There now you have an interface with ethernet look and feel, and
routing can happily avoid it.
So, to sum up, routing / proxy-arp seems to be off where we are
targeting.
My condolences.
The existence of router / proxy-arp means that solutions do exist
(unlike your previous claim) you just don't like the idea of deploying
them.
Infiniband is standard enough you could quite easily implement virtual
infiniband bridging as an alternative to ethernet bridging.
At this stage of the game eIPoIB is interesting the same way a zombie is
interesting. It is fascinating to see it still moving as chunks of
flesh fall to the floor removing any doubt that it is dead.
Eric
From: Or Gerlitz <hidden> Date: 2012-08-09 04:06:47
Eric W. Biederman [off-list ref] wrote:
Or Gerlitz [off-list ref] writes:
quoted
To put things in place, DHCPv4 is supported with eIPoIB, the DHCP
UDP/IP payload isn't touched, only need to set the BOOTP broadcast
flag in the dhcp server config file.
Wrong. DHCPv4 is broken over eIPoIB. Coming from ethernet
htype == 1 not 32 as required by RFC4390
hlen == 6 not 0 as required by RFC4390
The chaddr field is has 6 bytes of the ethernet mac address not the
required 16 bytes of 0. The client-identifier field is optional over ethernet.
An ethernet DHCPv4 client simply does not generate a dhcp packet that
conforms to RFC4390.
Therefore DHCPv4 over eIPoIB is broken, and a dhcp server or relay
may reasonably look at the DHCP packet and drop it because it is garbage.
You might find a forgiving dhcp server that doesn't drop insane packets
on the floor and tries to make things work.
Under the eIPoIB design, the VM DHCP interaction follows
Ethernet DHCP, and not the IPoIB DHCP (RFC 4390).
The DHCP server has no reason to drop such packets.
DHCP is a L7 (L5 to be precise) construct, I don't see
why that the fact IPoIB DHCP RFC exists, means/mandates
the DHCP server to care on the link layer type.
Or.
From: Or Gerlitz <hidden> Date: 2012-08-09 04:34:24
Eric W. Biederman [off-list ref] wrote:
Or Gerlitz [off-list ref] writes:
quoted
as you wrote here, when performance and ease-of-use is under the spot,
VM deployments tend to not to use routing.
This is b/c it involves more over-head on the packet forwarding, and
more administration work, for example, for setting routing rules that
involve the VM IP address, something which AFAIK the hypervisor have
no clue on, also its unclear to me if/how live migration can work in such setting.
All you need to make proxy-arp essentially pain free is a smart dhcp
relay, that sets up the routes.
So dhcp relay would help to maybe avoid some of the pain, however,
can you elaborate if/how live migration is supported under this scheme?
quoted
From this exact reason, there's a bunch of use-cases by tools and
cloud stacks (such as open stack, ovirt, more) which do use bridged
mode and the rest of the Ethernet envelope, such as using virtual L2
vlan domains, ebtables based rules, etc etc. Where they and are not
application to ipoib, but are working file ith eipoib.
Yes I am certain all of their IPv6 traffic works fine.
I'm not sure to follow this comment.
Regardless those are open source projects and can be modified to add
support to cleanly support inifiniband.
open source/code can be modified indeed, but since exposing IPoIB link
layer to tools/emulators and VMs doesn't really make sense (see below),
we brought that this approach as a way to go for allowing people to use
bridging mode when the fabric is IB.
quoted
You mentioned that bridging mode doesn't apply to environment such as
802.11, and hence routing mode is used, we are trying to make a pointn
here that bridging mode applies to ipoib with the approach suggested by eipoib.
You are completely failing. Every time I look I see something about
eIPoIB that is even more broken. Given that eIPoIB is a NAT
implementation that isn't really a surprise but still.
eIPoIB imposes enough overhead that I expect that routing is cheaper,
so your performance advantges go right out the window.
eIPoIB is seriously incompatible with ethernet breaking almost
everything and barely allowing IPv4 to work.
I don't agree on this incompatiblity statement, you had a claim
on DHCP and I addressed it, beyond that, you don't like the eIPoIB
basic idea/design but this can't base an incompatiblity argument.
quoted
Also, if we extend the discussion a bit, there are two more aspects to throw in:
The first is the performance thing we have already started to mention
-- specifically, the approach for RX zero copy (into the VM buffer),
use designs such as vhost + macvtap NIC in passthrough mode which is
likey to be set over a per VM hypervisor NIC, e.g such as the ones
provided by VMDQ patches John Fastabend started to post (see
http://marc.info/?l=linux-netdev&m=134264998405581&w=2) -- the ib0.N
clone child are IPoIB VMDQ NICs if you like, and setting an eipoib NIC
on top of each they can be plugged to that design.
If you care about performance link-layer NAT is not the way to go.
Teach the pieces you care about how to talk infiniband.
quoted
The 2nd aspect, is NON VM environments where a NIC with Ethernet look
and feel is required for IP traffic, but this have to live within an
echo-system that fully uses IPoIB.
In other words, a use case where IPoIB has to be below the cover for
set of some specific apps, or nodes but do IP interaction with other
apps/nodes and gateways who use IPoIB, the eIPoIB driver provides that
functionality.
ip link add type dummy.
There now you have an interface with ethernet look and feel, and
routing can happily avoid it.
again, not sure to follow, you mean "routing can happily use it", correct? that
is do routing between the dummy interface to IPoIB interface?
quoted
So, to sum up, routing / proxy-arp seems to be off where we are targeting.
My condolences.
The existence of router / proxy-arp means that solutions do exist
(unlike your previous claim) you just don't like the idea of deploying them.
Don't like them from set of arguments, which we are covering here, re
manageability it still needs to clarified if/how live migration work and what
does it mean to always mandate dhcp relay.
Infiniband is standard enough you could quite easily implement virtual
infiniband bridging as an alternative to ethernet bridging.
Not really, as Michael indicated in his response over this thread
http://marc.info/?l=linux-netdev&m=134419288218373&w=2
IPoIB link layer addresses use IB HW constructs for which soft
hardware address setting isn't supported, and this interferes
with live migration.
Or.
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-12 10:24:00
On Wed, Aug 08, 2012 at 08:23:15AM +0300, Or Gerlitz wrote:
On Sun, Aug 5, 2012 at 9:50 PM, Michael S. Tsirkin [off-list ref] wrote:
[...]
quoted
So it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes sense.
[...]
Yep, that's among the points we're trying to make, the way you've put
it makes it clearer.
quoted
So far the things that make sense. Here are some that don't, to me:
quoted
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
To be precise, the solution uses 100% IPoIB wire-protocol, so we don't
see a need
for any spec change / standardization effort.
Yes, I am guessing this is the real reason you pack LID/QPN
in the MAC - to make it all local. But it's a hack really,
and if you start storing it all in the SM you will need
to document the format so others can inter-operate.
This might go to the 1st
point you've
brought... improve the documentation, will do that. The pdf you looked
at was presented
in a conference.
quoted
Other things that I would expect to be addressed in such a spec is
interaction with other IPoIB features, such as connected
mode, checksum offloading etc, and IB features such as multipath etc.
For the eipoib interface, it doesn't really matters if the underlyind
ipoib clones used by it (we call them VIFs) use connected or datagram
mode, what does matter is the MTU and offload features supported by
these VIFs, for which the eipoib interface will have the min among all
these VIFs. Since for a given eipoib nic, all its VIFs must originated
from the same IPoIB PIF (e.g ib0) its easy admin job to make sure they
all have the same mtu / features which are needed for that eipoib nic,
e.g by using the same mode (connected/datagram for all of them), hope
this is clear.
Just pointing out all this needs to be documented.
quoted
- The way you encode LID/QPN in the MAC seems questionable. IIRC there's
more to IB addressing than just the LID. Since everyone on the subnet
need access to this translation, I think it makes sense to store it in
the SM. I think this would also obviate some IPv4 specific hacks in kernel.
The idead beyond the encoding was uniqueness, LID/QPN is unique per IB
HCA end-node.
But then it breaks with VM migration, IB failover, softmac setting in
guest, probably more?
I wasn't sure to understand the comment re the IPv4 hacks.
This refers to the ARP hack that you use to fix
VM migration.
quoted
- IGMP/MAC snooping in a driver is just too hairy.
mmm, any rough idea/direction how to do that otherwise?
Sure, even two ways, ideally you'd do both :)
A. fix macvtap
1. Use netdev_for_each_mc_addr etc to get multicast addresses
2. teach macvtap to fill that in (it currently floods multicasts
for guest to guest communication so we ned to fix it anyway)
B. fix bridge
teach bridge to work for VMs without using promisc mode
quoted
As you point out, bridge currently needs the uplink in promisc mode.
I don't think a driver should work around that limitation.
For some setups, it might be interesting to remove the
promisc mode requirement, failing that,
I think you could use macvtap passthrough.
That's in the plans, the current code doesn't assume that the eipoib
has bridge on top, for VM networking it works with bridge + tap,
bridge + macvtap, but it would easily work with passthrough when we
allow to create multiple eipoib interfaces on the same ipoib PIF (e.g
today for the ib0 PIF we create eipoib eth0, and then two VIFs ib0.1
and ib0.2 that are enslaved by eth0, but next we will create eth1 and
eth2 which will use ib0.1 and ib0.2
respectively.
The whole promisc mode emulation is there for the bridge, no?
Since you don't support promisc, ideally we'd check a hardware
capability and fail gracefully, though naturally this is not top
priority.
quoted
- Currently migration works without host kernel help, would be
preferable to keep it that way.
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-12 10:38:00
On Thu, Aug 09, 2012 at 07:34:23AM +0300, Or Gerlitz wrote:
quoted
Infiniband is standard enough you could quite easily implement virtual
infiniband bridging as an alternative to ethernet bridging.
Not really, as Michael indicated in his response over this thread
http://marc.info/?l=linux-netdev&m=134419288218373&w=2
IPoIB link layer addresses use IB HW constructs for which soft
hardware address setting isn't supported, and this interferes
with live migration.
Or.
But I'd just like to point out your code does not support softmac
either.
--
MST
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-12 10:55:37
On Wed, Aug 08, 2012 at 08:23:15AM +0300, Or Gerlitz wrote:
quoted
So far the things that make sense. Here are some that don't, to me:
quoted
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
To be precise, the solution uses 100% IPoIB wire-protocol, so we don't
see a need
for any spec change / standardization effort.
Thinking about it some more, the solution is not really 100% local:
- ARP packets used at migration are visible on the network
- QPN/LID in MAC are guest visible, and could be exported on the network
From: Or Gerlitz <hidden> Date: 2012-08-12 13:10:12
On 12/08/2012 13:22, Michael S. Tsirkin wrote:
On Wed, Aug 08, 2012 at 08:23:15AM +0300, Or Gerlitz wrote:
quoted
quoted
On Sun, Aug 5, 2012 at 9:50 PM, Michael S. Tsirkin[off-list ref] wrote:
[...]
quoted
quoted
So it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes sense.
[...]
Yep, that's among the points we're trying to make, the way you've put
it makes it clearer.
quoted
quoted
So far the things that make sense. Here are some that don't, to me:
quoted
quoted
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
To be precise, the solution uses 100% IPoIB wire-protocol, so we don't
see a need
for any spec change / standardization effort.
Yes, I am guessing this is the real reason you pack LID/QPN
in the MAC - to make it all local. But it's a hack really,
and if you start storing it all in the SM you will need
to document the format so others can inter-operate.
I'd like to review the way we generate these MAC addresses, maybe it
can be done differently.
quoted
quoted
This might go to the 1stpoint you've
brought... improve the documentation, will do that. The pdf you looked
at was presentedin a conference.
quoted
quoted
quoted
Other things that I would expect to be addressed in such a spec is
interaction with other IPoIB features, such as connected
mode, checksum offloading etc, and IB features such as multipath etc.
For the eipoib interface, it doesn't really matters if the underlyind
ipoib clones used by it (we call them VIFs) use connected or datagram
mode, what does matter is the MTU and offload features supported by
these VIFs, for which the eipoib interface will have the min among all
these VIFs. Since for a given eipoib nic, all its VIFs must originated
from the same IPoIB PIF (e.g ib0) its easy admin job to make sure they
all have the same mtu / features which are needed for that eipoib nic,
e.g by using the same mode (connected/datagram for all of them), hope
this is clear.
Just pointing out all this needs to be documented.
OK, will do
quoted
quoted
quoted
quoted
- The way you encode LID/QPN in the MAC seems questionable. IIRC there's
more to IB addressing than just the LID. Since everyone on the subnet
need access to this translation, I think it makes sense to store it in
the SM. I think this would also obviate some IPv4 specific hacks in kernel.
The idead beyond the encoding was uniqueness, LID/QPN is unique per IB
HCA end-node.
But then it breaks with VM migration, IB failover, softmac setting in guest, probably more?
With the current design/code the remote mac of a VM changes, when that
VM migrates or IB
LIDs are changed. As for softmac setting in the guest, we don't send
the guest MAC on the wire
anyway, since the Ethernet header is removed.
Or.
From: Or Gerlitz <hidden> Date: 2012-08-12 13:17:02
On 12/08/2012 13:22, Michael S. Tsirkin wrote:
quoted
quoted
quoted
quoted
- IGMP/MAC snooping in a driver is just too hairy.
quoted
mmm, any rough idea/direction how to do that otherwise?
Sure, even two ways, ideally you'd do both:)
A. fix macvtap
1. Use netdev_for_each_mc_addr etc to get multicast addresses
2. teach macvtap to fill that in (it currently floods multicasts
for guest to guest communication so we ned to fix it anyway)
B. fix bridge
teach bridge to work for VMs without using promisc mode
I wasn't sure to fully follow... need some more bits of info, the
macvtap fix
relates only to IGMP snooping, correct? as for the bridge fix, does this
somehow
relates to ARP snooping we do in the driver? how?
Or.
From: Or Gerlitz <hidden> Date: 2012-08-12 13:21:00
On 12/08/2012 13:54, Michael S. Tsirkin wrote:
On Wed, Aug 08, 2012 at 08:23:15AM +0300, Or Gerlitz wrote:
quoted
quoted
So far the things that make sense. Here are some that don't, to me:
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
To be precise, the solution uses 100% IPoIB wire-protocol, so we don't
see a need for any spec change / standardization effort.
Thinking about it some more, the solution is not really 100% local:
- ARP packets used at migration are visible on the network
The ARP packets on the wire are IPoIB ARP packets -- indeed, when the
guest moves, the IPoIB
link layer address which was associated with the VIF that used to serve
this guest on the host
we migrated from, changes to be of a VIF on the host that we migrated to.
- QPN/LID in MAC are guest visible, and could be exported on the network
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-12 13:43:05
On Sun, Aug 12, 2012 at 04:09:35PM +0300, Or Gerlitz wrote:
On 12/08/2012 13:22, Michael S. Tsirkin wrote:
quoted
On Wed, Aug 08, 2012 at 08:23:15AM +0300, Or Gerlitz wrote:
quoted
quoted
On Sun, Aug 5, 2012 at 9:50 PM, Michael S. Tsirkin[off-list ref] wrote:
[...]
quoted
quoted
So it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes sense.
[...]
Yep, that's among the points we're trying to make, the way you've put
it makes it clearer.
quoted
quoted
So far the things that make sense. Here are some that don't, to me:
quoted
quoted
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
To be precise, the solution uses 100% IPoIB wire-protocol, so we don't
see a need
for any spec change / standardization effort.
Yes, I am guessing this is the real reason you pack LID/QPN
in the MAC - to make it all local. But it's a hack really,
and if you start storing it all in the SM you will need
to document the format so others can inter-operate.
I'd like to review the way we generate these MAC addresses, maybe it
can be done differently.
quoted
quoted
quoted
This might go to the 1stpoint you've
brought... improve the documentation, will do that. The pdf you looked
at was presentedin a conference.
quoted
quoted
quoted
Other things that I would expect to be addressed in such a spec is
interaction with other IPoIB features, such as connected
mode, checksum offloading etc, and IB features such as multipath etc.
For the eipoib interface, it doesn't really matters if the underlyind
ipoib clones used by it (we call them VIFs) use connected or datagram
mode, what does matter is the MTU and offload features supported by
these VIFs, for which the eipoib interface will have the min among all
these VIFs. Since for a given eipoib nic, all its VIFs must originated
from the same IPoIB PIF (e.g ib0) its easy admin job to make sure they
all have the same mtu / features which are needed for that eipoib nic,
e.g by using the same mode (connected/datagram for all of them), hope
this is clear.
Just pointing out all this needs to be documented.
OK, will do
quoted
quoted
quoted
quoted
quoted
- The way you encode LID/QPN in the MAC seems questionable. IIRC there's
more to IB addressing than just the LID. Since everyone on the subnet
need access to this translation, I think it makes sense to store it in
the SM. I think this would also obviate some IPv4 specific hacks in kernel.
The idead beyond the encoding was uniqueness, LID/QPN is unique per IB
HCA end-node.
But then it breaks with VM migration, IB failover, softmac setting in guest, probably more?
With the current design/code the remote mac of a VM changes, when
that VM migrates or IB
LIDs are changed.
Which is exactly the problem with IB and VM migration.
As for softmac setting in the guest, we don't
send the guest MAC on the wire
anyway, since the Ethernet header is removed.
Or.
Yes but you generate remote addresses automatically so admin can not
change the local address without risking conflicts.
--
MST
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-12 13:56:59
On Sun, Aug 12, 2012 at 04:15:52PM +0300, Or Gerlitz wrote:
On 12/08/2012 13:22, Michael S. Tsirkin wrote:
quoted
quoted
quoted
quoted
quoted
- IGMP/MAC snooping in a driver is just too hairy.
quoted
mmm, any rough idea/direction how to do that otherwise?
Sure, even two ways, ideally you'd do both:)
A. fix macvtap
1. Use netdev_for_each_mc_addr etc to get multicast addresses
2. teach macvtap to fill that in (it currently floods multicasts
for guest to guest communication so we ned to fix it anyway)
B. fix bridge
teach bridge to work for VMs without using promisc mode
I wasn't sure to fully follow... need some more bits of info, the
macvtap fix
relates only to IGMP snooping, correct? as for the bridge fix, does
this somehow
relates to ARP snooping we do in the driver? how?
Or.
I didn't realize you do ARP snooping. Why?
I know you mangle outgoing ARP packets, this will go
away if you maintain a mapping in SM accessible to all guests.
--
MST
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-12 14:06:27
On Thu, Aug 09, 2012 at 07:06:46AM +0300, Or Gerlitz wrote:
Eric W. Biederman [off-list ref] wrote:
quoted
Or Gerlitz [off-list ref] writes:
quoted
quoted
To put things in place, DHCPv4 is supported with eIPoIB, the DHCP
UDP/IP payload isn't touched, only need to set the BOOTP broadcast
flag in the dhcp server config file.
quoted
Wrong. DHCPv4 is broken over eIPoIB. Coming from ethernet
htype == 1 not 32 as required by RFC4390
hlen == 6 not 0 as required by RFC4390
The chaddr field is has 6 bytes of the ethernet mac address not the
required 16 bytes of 0. The client-identifier field is optional over ethernet.
An ethernet DHCPv4 client simply does not generate a dhcp packet that
conforms to RFC4390.
Therefore DHCPv4 over eIPoIB is broken, and a dhcp server or relay
may reasonably look at the DHCP packet and drop it because it is garbage.
You might find a forgiving dhcp server that doesn't drop insane packets
on the floor and tries to make things work.
Under the eIPoIB design, the VM DHCP interaction follows
Ethernet DHCP, and not the IPoIB DHCP (RFC 4390).
The DHCP server has no reason to drop such packets.
DHCP is a L7 (L5 to be precise) construct, I don't see
why that the fact IPoIB DHCP RFC exists, means/mandates
the DHCP server to care on the link layer type.
Or.
For example DHCP server could be configured with
HW address/IP address table.
--
MST
From: Or Gerlitz <hidden> Date: 2012-08-12 14:14:14
On 12/08/2012 16:55, Michael S. Tsirkin wrote:
I didn't realize you do ARP snooping. Why? I know you mangle outgoing
ARP packets,
Maybe I wasn't accurate/clear, we do mangle outgoing/incoming ARP
packets, from/to Ethernet ARPs
to/from IPoIB ARPs.
this will go away if you maintain a mapping in SM accessible to all
guests.
guests don't interact with IB, I assume you referred to dom0 code, eIPoIB or
another driver in the host. But what mapping exactly? and remember that
this code (VM through eipoib) can talk to any IPoIB element on the
fabric, native,
virtualized, HW/SW gateways, etc etc.
Or.
From: Eric W. Biederman <hidden> Date: 2012-08-12 15:40:26
Or Gerlitz [off-list ref] writes:
On Sun, Aug 5, 2012 at 9:50 PM, Michael S. Tsirkin [off-list ref] wrote:
[...]
quoted
So it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes sense.
[...]
Yep, that's among the points we're trying to make, the way you've put
it makes it clearer.
quoted
- IGMP/MAC snooping in a driver is just too hairy.
mmm, any rough idea/direction how to do that otherwise?
Let me give you a non-hack recomendation.
- Give up on being wire compatible with IPoIB.
- Define and implement ethernet over inifiniband aka EoIB.
With EoIB:
- The SM would map ethernet address to inifiniband hardware addresses.
- You discover which multicast addresses are of interest from the
IP layer above so no snooping is necessary.
- You could run queue pairs directly to hosts.
Shrug. It is trivial and it will work. It will probably run into the
same problems that have historically been a problem for using IPoIB
(lack of stateless offloads) but shrug that is mostly a NIC firmware
problem. The switches will have no trouble and interoperability will
be assured.
If you want to map ethernet over infiniband please map ethernet over
infiniband. Don't poorly NAT ethernet into infiniband.
Eric
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-12 20:56:12
On Sun, Aug 12, 2012 at 05:13:43PM +0300, Or Gerlitz wrote:
On 12/08/2012 16:55, Michael S. Tsirkin wrote:
quoted
I didn't realize you do ARP snooping. Why? I know you mangle
outgoing ARP packets,
Maybe I wasn't accurate/clear, we do mangle outgoing/incoming ARP
packets, from/to Ethernet ARPs
to/from IPoIB ARPs.
quoted
this will go away if you maintain a mapping in SM accessible to
all guests.
guests don't interact with IB, I assume you referred to dom0 code, eIPoIB or
another driver in the host. But what mapping exactly?
Well we are getting into protocol design here. So here's a sketch
showing how you could build a protocol that does work. But note it is
not *exactly* IPoIB. It is I think close enough that you can
use existing NIC hardware/firmware, which is why it
differs slightly from what Eric described, and is
more complex. But it still shares the same property of no hacks, no
packet snooping in driver, etc.
And if you want to go that route, you really should talk to some IB
protocol people to figure out what works, write a spec and try to
standardize. lkml/netdev is not the right place.
But since you asked, if I had to, I would probably try to
do it like this:
- Each device registers with the SA specifying the
mac address (+ vlan?), SA stores the translation from that
to IPoIB address.
- alternatively, SA admin configures the translation statically
- you get a packet with 6 byte mac address,
query the SA for a mapping to IPoIB address, strip
ethernet frame and send
- multicast GID addresses can be similar, filled either when registering
for multicast or by SA admin
I think it's possible that you could also convert a mac address to
EUI-64 and prepend a prefix to get a legal GID. But maybe I'm missing
something. This could be handy for multicast.
In both cases:
- SA could return GID that you then resolve to
LID using another query, or it could return LID so you save a roundtrip
- results can be cached locally
- SA can send updates when translation changes to flush this cache
Now above means protocols such as ARP and DHCP use 48 bit addresses so
you can not mix this new protocol with IPoIB. Maybe IPoIB could simply
ignore irrelevant packets, but it's best not to try, get a
different all-broadcast group and CM ID instead to avoid confusion.
One other interesting thing you can do is forward multicast
registration data from the router, translate to mgid
by the SA and do appropriate IB mcast registrations.
My memory of how IB works is a bit rusty so I probably made some
mistakes above but should roughly work, I think.
and remember that
this code (VM through eipoib) can talk to any IPoIB element on the
fabric, native,
virtualized, HW/SW gateways, etc etc.
Or.
If you want this, then you really want a limited form of IPoIB bridging.
Alternatively, decide that what you do is not IPoIB, have a proper
protocol of your own.
--
MST
From: Or Gerlitz <hidden> Date: 2012-08-13 08:35:50
On 12/08/2012 18:40, Eric W. Biederman wrote:
Let me give you a non-hack recomendation.
- Give up on being wire compatible with IPoIB.
- Define and implement ethernet over inifiniband aka EoIB.
With EoIB:
- The SM would map ethernet address to inifiniband hardware addresses.
- You discover which multicast addresses are of interest from the
IP layer above so no snooping is necessary.
- You could run queue pairs directly to hosts.
Shrug. It is trivial and it will work. It will probably run into the
same problems that have historically been a problem for using IPoIB
(lack of stateless offloads) but shrug that is mostly a NIC firmware
problem. The switches will have no trouble and interoperability will
be assured.
If you want to map ethernet over infiniband please map ethernet over
infiniband. Don't poorly NAT ethernet into infiniband.
EoIB is a valid suggestion and we will look into it as well, BUT:
Providing EoIB is a separate discussion, obviously defining and
standardizing a new protocol takes what is takes (a lot of time, longish
term effort), and will also take time to develop/debug/mature e.g as you
mentioned, some of the features/offloads might require new NIC HW, etc
-- compared to IPoIB which is here for many years
In practice there is already a huge install base for IPoIB software and
hardware products, in different operating environments/OS. We can't just
through away everything and tell people to replace it all with a new
protocol, e.g. bridging devices, storage systems/appliances, VMware,
Windows, .. systems in production environments --- so
the interoperability concern you've mentioned gonna hit very hard.
The eIPoIB driver comes to provide a way to work with IPoIB in
virtualized environments, where still, the suggestions/concerns raised
in this thread should be addressed.
Or.
From: Eric W. Biederman <hidden> Date: 2012-08-13 16:08:56
Or Gerlitz [off-list ref] wrote:
On 12/08/2012 18:40, Eric W. Biederman wrote:
quoted
Let me give you a non-hack recomendation.
- Give up on being wire compatible with IPoIB.
- Define and implement ethernet over inifiniband aka EoIB.
With EoIB:
- The SM would map ethernet address to inifiniband hardware
addresses.
quoted
- You discover which multicast addresses are of interest from the
IP layer above so no snooping is necessary.
- You could run queue pairs directly to hosts.
Shrug. It is trivial and it will work. It will probably run into
the
quoted
same problems that have historically been a problem for using IPoIB
(lack of stateless offloads) but shrug that is mostly a NIC firmware
problem. The switches will have no trouble and interoperability will
be assured.
If you want to map ethernet over infiniband please map ethernet over
infiniband. Don't poorly NAT ethernet into infiniband.
EoIB is a valid suggestion and we will look into it as well, BUT:
Providing EoIB is a separate discussion, obviously defining and
standardizing a new protocol takes what is takes (a lot of time,
longish
term effort), and will also take time to develop/debug/mature e.g as
If you follow Michael Tirskins suggestion and use the same wire encoding as IPoIB and infer the mac address from the lids and queue pair numbers as you are already doing with for eIPoIB, except for defining exactly how to get the subnet manager to store the mac address to lid/qpn mapping you are done.
If you don't involve a comitte and simply define a defacto standard it will take less effort than this conversation and less effort than implementing your eIPoIB driver.
you
mentioned, some of the features/offloads might require new NIC HW, etc
-- compared to IPoIB which is here for many years
So deploy routing and proxy arp and you are done.
In practice there is already a huge install base for IPoIB software and
hardware products, in different operating environments/OS. We can't
just
through away everything and tell people to replace it all with a new
protocol, e.g. bridging devices, storage systems/appliances, VMware,
Windows, .. systems in production environments --- so
the interoperability concern you've mentioned gonna hit very hard.
There is no need to throw anything away. Just put them on different IP subnets.
Shrug.
The eIPoIB driver comes to provide a way to work with IPoIB in
virtualized environments, where still, the suggestions/concerns raised
in this thread should be addressed.
eIPoIB does not work.
I can't get an IP address with out a specially configured dhcp server, and special dhcp clients.
eIPoIB does not work with IPv6.
As David Miller already said this code has no chance of being merged.
Shrug. I have been polite and pointed out implementation choices that actually work. Solutions that are less effort and less code, and provide more interoperability.
If after patient explanation you can not appreciate why people consider eIPoIB to be totally unacceptable that is your problem.
Good luck in your future endeavours,
Eric
From: Or Gerlitz <hidden> Date: 2012-08-14 07:41:47
On 12/08/2012 13:22, Michael S. Tsirkin wrote:
quoted
quoted
quoted
quoted
quoted
quoted
- IGMP/MAC snooping in a driver is just too hairy.
quoted
mmm, any rough idea/direction how to do that otherwise?
Sure, even two ways, ideally you'd do both:)
A. fix macvtap
1. Use netdev_for_each_mc_addr etc to get multicast addresses
2. teach macvtap to fill that in (it currently floods multicasts
for guest to guest communication so we ned to fix it anyway)
B. fix bridge
teach bridge to work for VMs without using promisc mode
OK, I think I'm with you now... you suggest to avoid our direction of
implementing promiscuous multicast mode which is applied by today's
bridge, macvtap and friends by fixing these elements to support non
promisc multicast mode, yep, sure, sounds as win/win, which will
eliminate the need to do IGMP snooping in the driver.
Or.
From: Or Gerlitz <hidden> Date: 2012-08-14 08:46:07
On 12/08/2012 23:54, Michael S. Tsirkin wrote:
On Sun, Aug 12, 2012 at 05:13:43PM +0300, Or Gerlitz wrote:
quoted
On 12/08/2012 16:55, Michael S. Tsirkin wrote:
quoted
I didn't realize you do ARP snooping. Why? I know you mangle
outgoing ARP packets,
Maybe I wasn't accurate/clear, we do mangle outgoing/incoming ARP
packets, from/to Ethernet ARPs to/from IPoIB ARPs.
quoted
this will go away if you maintain a mapping in SM accessible to all guests.
guests don't interact with IB, I assume you referred to dom0 code, eIPoIB or
another driver in the host. But what mapping exactly?
Well we are getting into protocol design here.
wait... reading your responses again, I realized that we 1st and most
have to (try and) agree on the
problem statement before going/jumping to solutions.
AFAIU your email/s you maybe think that we mandate the admin to set a
specific MAC to the VM which is derived from the LID/QPN of the IPoIB
VIF serving it, well this is wrong, we don't, OTOH, indeed, the VM
source mac isn't sent on the wire, since the Ethernet header is dropped,
and on the receiving side is constructed from the LID/QPN
the IB packet arrived from, see next.
This reconstruction of what we call the REMAC (remote ethernet mac) is
based in the current submission on
the LID/QPN, and as I said earlier on this thread, we are revisiting
this approach -- where your idea below sounds
good: the eipoib driver can register with the SA an IB "service record"
entry mapping from LID/QPN to the VM mac, when ever a VM is to be served
by this eipoib instance, and remove the entry when the VM shouldn't be
served any more. This will allow to preserve 1:1 the Ethernet MAC header
sent by VMs on the receiving side.
So here's a sketch showing how you could build a protocol that does work. But note it is not *exactly* IPoIB.
HOWEVER, this doesn't touch the IPoIB wire protocol, and hence on the
wire it IS exactly IPoIB. We only make use of your lovely suggestion to
apply this SA assistance, so the change doesn't involve
hardware/firmware nor the wire protocol.
Or.
It is I think close enough that you can use existing NIC hardware/firmware, which is why it differs slightly from what Eric described, and is more complex. But it still shares the same property of no hacks, no packet snooping in driver, etc.
And if you want to go that route, you really should talk to some IB
protocol people to figure out what works, write a spec and try to
standardize. lkml/netdev is not the right place.
But since you asked, if I had to, I would probably try to
do it like this:
- Each device registers with the SA specifying the
mac address (+ vlan?), SA stores the translation from that
to IPoIB address.
- alternatively, SA admin configures the translation statically
- you get a packet with 6 byte mac address,
query the SA for a mapping to IPoIB address, strip
ethernet frame and send
- multicast GID addresses can be similar, filled either when registering
for multicast or by SA admin
I think it's possible that you could also convert a mac address to
EUI-64 and prepend a prefix to get a legal GID. But maybe I'm missing
something. This could be handy for multicast.
In both cases:
- SA could return GID that you then resolve to
LID using another query, or it could return LID so you save a roundtrip
- results can be cached locally
- SA can send updates when translation changes to flush this cache
Now above means protocols such as ARP and DHCP use 48 bit addresses so
you can not mix this new protocol with IPoIB. Maybe IPoIB could simply
ignore irrelevant packets, but it's best not to try, get a
different all-broadcast group and CM ID instead to avoid confusion.
One other interesting thing you can do is forward multicast
registration data from the router, translate to mgid
by the SA and do appropriate IB mcast registrations.
Hi All.
Our lab's additional infomation(result) for all.
Ethernet over L2TPv3 over IPoIB Pseudo-wire Multiplexing (RFC4719)
using "ip l2tp" command on linux. http://twitpic.com/ajn2nx
On Tue, 07 Aug 2012 10:21:33 +0900
Naoto MATSUMOTO [off-list ref] wrote:
-----------------------------------------------------------------------
Ethernet over EoIB using gretap with "Raid STP" performance result.
http://twitpic.com/agd1pp
How to use EoIB with gretap.
http://twitpic.com/agd4io
regards,
--
SAKURA Internet Inc. / Senior Researcher
Naoto MATSUMOTO [off-list ref]
SAKURA Research Center <http://research.sakura.ad.jp/>
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-08-20 18:56:44
On Sun, Aug 12, 2012 at 11:54:57PM +0300, Michael S. Tsirkin wrote:
quoted
and remember that
this code (VM through eipoib) can talk to any IPoIB element on the
fabric, native,
virtualized, HW/SW gateways, etc etc.
Or.
If you want this, then you really want a limited form of IPoIB bridging.
And to clarify that statement, here is how I would make such
IPoIB "bridging" work:
Guest side:
- Implement virtio-ipoib. This would be a device like virtio-net,
but instead of ethernet packets, it would pass packets
that consist of:
IPoIB destination address
IP packet
- this is passed to/from host without modifications, possibly with addition
of header such as virtio net header
- flags such as broadcast can also be added to header
- like virtio net get capabilities from host and expose
as netdev capabilities
Host side:
- create macvtap -passthrough like device that can sit on top of an
ipoib interface
- expose this device QPN and GID to guest as hardware address
- as we get packet forward it on UD QPN or CM as appropriate
depending on size,checksum and admin preference
- expose capabilities such as TSO
- can expose capability such as max MTU to guest too
Above means hardware address changes with migration.
So we need to notify guest when this happens.
This can be addressed from host by notifying all
neighbours.
Alternatively guest can notify all neighbours.
Notification can be done by broadcast.
This second option seems preferable.
this ipoib-vtap can support two modes
- bridge like mode:
guest to guest and guest to host packets
can be detected by macvtap and passed
to/from guest directly like macvlan bridge mode
- vepa like mode
guest to guest and guest to host packets
are sent out and looped back by IB switch
like macvlan vepa mode
As compared to the custom protocol I sent, it has -
Advantages: interoperates cleanly with ipoib
Disadvantages: no support for legacy (ethernet-only) guest
--
MST
From: Or Gerlitz <hidden> Date: 2012-08-23 06:45:50
On 20/08/2012 21:57, Michael S. Tsirkin wrote:
On Sun, Aug 12, 2012 at 11:54:57PM +0300, Michael S. Tsirkin wrote:
quoted
If you want this, then you really want a limited form of IPoIB bridging.
And to clarify that statement, here is how I would make such IPoIB "bridging" work:
Guest side:
- Implement virtio-ipoib. This would be a device like virtio-net,
but instead of ethernet packets, it would pass packets
that consist of:
IPoIB destination address
IP packet
- this is passed to/from host without modifications, possibly with addition
of header such as virtio net header
- flags such as broadcast can also be added to header
- like virtio net get capabilities from host and expose
as netdev capabilities
Host side:
- create macvtap -passthrough like device that can sit on top of an
ipoib interface
- expose this device QPN and GID to guest as hardware address
- as we get packet forward it on UD QPN or CM as appropriate
depending on size,checksum and admin preference
- expose capabilities such as TSO
- can expose capability such as max MTU to guest too
Above means hardware address changes with migration.
So we need to notify guest when this happens.
This can be addressed from host by notifying all neighbours.
Alternatively guest can notify all neighbours.
Notification can be done by broadcast.
This second option seems preferable.
this ipoib-vtap can support two modes
- bridge like mode:
guest to guest and guest to host packets
can be detected by macvtap and passed
to/from guest directly like macvlan bridge mode
- vepa like mode
guest to guest and guest to host packets
are sent out and looped back by IB switch
like macvlan vepa mode
As compared to the custom protocol I sent, it has -
Advantages: interoperates cleanly with ipoib
Disadvantages: no support for legacy (ethernet-only) guest
Hi Michael,
As you mentioned, the approach doesn't address legacy guests, who either
don't have the virtio
driver, or don't have a virtio driver patched to support virtio-ipoib
-- which doesn't go inline
with a strong requirement I got.
Other than this giant obstacle, I liked the suggestion and it seems
valid and viable -- BTW IB HW has
loopback capability, so the VM/VM packets wouldn't actually go to the IB
switch, but remain within the
HCA.
Or.
From: Or Gerlitz <hidden> Date: 2012-09-03 20:53:57
Michael S. Tsirkin [off-list ref] wrote:
[...] so it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes sense.
So far the things that make sense. Here are some that don't, to me:
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
Other things that I would expect to be addressed in such a spec is
interaction with other IPoIB features, such as connected
mode, checksum offloading etc, and IB features such as multipath etc.
- The way you encode LID/QPN in the MAC seems questionable. IIRC there's
more to IB addressing than just the LID. Since everyone on the subnet
need access to this translation, I think it makes sense to store it in
the SM. I think this would also obviate some IPv4 specific hacks in kernel.
- IGMP/MAC snooping in a driver is just too hairy.
As you point out, bridge currently needs the uplink in promisc mode.
I don't think a driver should work around that limitation.
For some setups, it might be interesting to remove the promisc
mode requirement, failing that, I think you could use macvtap passthrough.
- Currently migration works without host kernel help, would be
preferable to keep it that way.
Hi Michael,
If we rewind to this point, basically, you had few concerns
0. not enough documentation
1. the sender VM MAC isn't preserved when the packet is received
2. the IGMP snooping we planned to do within netdevice - isn't good practice
3. mangling of ARPs within netdevice - isn't good practice as well.
For 0,1,2 we have a way to address (see below)
So we are remained with #3 - the ARPs -- thinking on this a little
further, FWIW there --are-- components in the kernel which
mangle/generate ARPs and are exposing netdevice, such as openvswitch,
anyway:
does it make sense to forward ARPs received into / sent over the
eIPoIB netdevice (e.g using some sort of rule) to some outer entity
such as user-space
daemon for interception and later re-injection into eIPoIB?
Or.
Documentation we will fix,
Preserving remote VM mac at the receiver we have few directions for
solution, e.g either along your suggestion with SA records and/or with
using "alias GUIDs" (details TBD when the submission resumes).
Multicast we accept the direction you suggested - implement support
for multicast non promiscuous in the elements "above" eIPoIB (bridge,
macvtap, etc).
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-09-03 21:21:31
On Mon, Sep 03, 2012 at 11:53:56PM +0300, Or Gerlitz wrote:
Michael S. Tsirkin [off-list ref] wrote:
quoted
[...] so it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes sense.
quoted
So far the things that make sense. Here are some that don't, to me:
quoted
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
Other things that I would expect to be addressed in such a spec is
interaction with other IPoIB features, such as connected
mode, checksum offloading etc, and IB features such as multipath etc.
- The way you encode LID/QPN in the MAC seems questionable. IIRC there's
more to IB addressing than just the LID. Since everyone on the subnet
need access to this translation, I think it makes sense to store it in
the SM. I think this would also obviate some IPv4 specific hacks in kernel.
quoted
- IGMP/MAC snooping in a driver is just too hairy.
As you point out, bridge currently needs the uplink in promisc mode.
I don't think a driver should work around that limitation.
For some setups, it might be interesting to remove the promisc
mode requirement, failing that, I think you could use macvtap passthrough.
- Currently migration works without host kernel help, would be
preferable to keep it that way.
Hi Michael,
If we rewind to this point, basically, you had few concerns
I think some other people gave feedback too, you need to address it in
the patch (as opposed to by mail - even if it's in documentation or
comments) don't just focus on what I wrote.
0. not enough documentation
1. the sender VM MAC isn't preserved when the packet is received
2. the IGMP snooping we planned to do within netdevice - isn't good practice
3. mangling of ARPs within netdevice - isn't good practice as well.
For 0,1,2 we have a way to address (see below)
So we are remained with #3 - the ARPs -- thinking on this a little
further, FWIW there --are-- components in the kernel which
mangle/generate ARPs and are exposing netdevice, such as openvswitch,
anyway:
does it make sense to forward ARPs received into / sent over the
eIPoIB netdevice (e.g using some sort of rule) to some outer entity
such as user-space
daemon for interception and later re-injection into eIPoIB?
Or.
Well if this is all you want to do, you can bind a packet socket to the
interface, and drop them at the nic. It is harder to do for incoming
ARP requests though.
I would do something else: send ARPs out to some defined IB address.
This could be local host or queries from some SA property. Said remote
side could send you the responses in ethernet format so you do not need
to mangle responses at all. Similarly for incoming ARP requests.
The rule to do this can also just redirect non IP packets -
this is IPoIB after all.
Documentation we will fix,
And just to stress the point, document the limitations as well.
Preserving remote VM mac at the receiver we have few directions for
solution, e.g either along your suggestion with SA records and/or with
using "alias GUIDs" (details TBD when the submission resumes).
Multicast we accept the direction you suggested - implement support
for multicast non promiscuous in the elements "above" eIPoIB (bridge,
macvtap, etc).
From: Or Gerlitz <hidden> Date: 2012-09-04 18:50:11
On Tue, Sep 4, 2012 at 12:22 AM, Michael S. Tsirkin [off-list ref] wrote:
On Mon, Sep 03, 2012 at 11:53:56PM +0300, Or Gerlitz wrote:
quoted
So we are remained with #3 - the ARPs -- thinking on this a little
further, FWIW there --are-- components in the kernel which
mangle/generate ARPs and are exposing netdevice, such as openvswitch, anyway:
quoted
does it make sense to forward ARPs received into / sent over the
eIPoIB netdevice (e.g using some sort of rule) to some outer entity
such as user-space daemon for interception and later re-injection into eIPoIB?
Well if this is all you want to do, you can bind a packet socket to the
interface, and drop them at the nic. It is harder to do for incoming
ARP requests though.
I would do something else: send ARPs out to some defined IB address.
This could be local host or queries from some SA property. Said remote
side could send you the responses in ethernet format so you do not need
to mangle responses at all. Similarly for incoming ARP requests.
The rule to do this can also just redirect non IP packets - this is IPoIB after all.
Thanks for the heads up on the possible implementation route, will
look into that.
quoted
Documentation we will fix,
And just to stress the point, document the limitations as well.
sure, not that I see concrete limitations for the **user** at this point, but
if there are such, will put them clearly written.
Or.
From: Or Gerlitz <hidden> Date: 2012-09-04 18:57:02
On Tue, Sep 4, 2012 at 12:22 AM, Michael S. Tsirkin [off-list ref] wrote:
On Mon, Sep 03, 2012 at 11:53:56PM +0300, Or Gerlitz wrote:
quoted
Michael S. Tsirkin [off-list ref] wrote:
quoted
[...] so it seems that a sane solution would involve an extra level of
indirection, with guest addresses being translated to host IB addresses.
As long as you do this, maybe using an ethernet frame format makes sense.
quoted
So far the things that make sense. Here are some that don't, to me:
quoted
- Is a pdf presentation all you have in terms of documentation?
We are talking communication protocols here - I would expect a
proper spec, and some effort to standardize, otherwise where's the
guarantee it won't change in an incompatible way?
Other things that I would expect to be addressed in such a spec is
interaction with other IPoIB features, such as connected
mode, checksum offloading etc, and IB features such as multipath etc.
- The way you encode LID/QPN in the MAC seems questionable. IIRC there's
more to IB addressing than just the LID. Since everyone on the subnet
need access to this translation, I think it makes sense to store it in
the SM. I think this would also obviate some IPv4 specific hacks in kernel.
quoted
- IGMP/MAC snooping in a driver is just too hairy.
As you point out, bridge currently needs the uplink in promisc mode.
I don't think a driver should work around that limitation.
For some setups, it might be interesting to remove the promisc
mode requirement, failing that, I think you could use macvtap passthrough.
- Currently migration works without host kernel help, would be
preferable to keep it that way.
quoted
If we rewind to this point, basically, you had few concerns
I think some other people gave feedback too, you need to address it in
the patch (as opposed to by mail - even if it's in documentation or
comments) don't just focus on what I wrote.
The other feedback was:
1. suggesting to introduce new link layer for the para-virtualized
network stack, a direction pointed by Dave and you, for which I
responded that it doesn't address a hard requirement I got, which is
provide service to an arbitrary VM which have any OS and a virtual
Ethernet NIC on, emulated or para-virtualized.
2. Suggestions to use solutions which involve routing and/or
proxt-arp, for which I responded that in most/practical cases this
will require for the host to know the VM IP, something which isn't
valid assumption in many cases AND that bunch of cloud stacks,
specifically the leading ones, don't even support that option, which
also violates a hard requirement we got, to support these stacks which
are in use by customers.
3. Suggestions to invent EoIB -- I said, OK but this is long termish
process that we're looking on, might be around at some future point,
but still we are required to support whole echo systems which use
IPoIB, and there's no point to "route" between EoIB segment to IPoIB
segments, its back to #2 which we didn't accept
4. Other feedback saying the eIPoIB driver is messy and "we don't like
it" -- hard to exactly address in code changes
----> all in all, we were suggested few directions which don't allow
us to address the problem statement, so there's no way to change the
code so they are fullfilled, and some not concrete comments which are
also hard to address --> we took the route of design changes along
your **concrete** comments, doesn't it make sense?
Or.
From: Eric W. Biederman <hidden> Date: 2012-09-04 19:31:27
Or Gerlitz [off-list ref] writes:
On Tue, Sep 4, 2012 at 12:22 AM, Michael S. Tsirkin [off-list ref] wrote:
quoted
On Mon, Sep 03, 2012 at 11:53:56PM +0300, Or Gerlitz wrote:
quoted
quoted
So we are remained with #3 - the ARPs -- thinking on this a little
further, FWIW there --are-- components in the kernel which
mangle/generate ARPs and are exposing netdevice, such as openvswitch, anyway:
quoted
quoted
does it make sense to forward ARPs received into / sent over the
eIPoIB netdevice (e.g using some sort of rule) to some outer entity
such as user-space daemon for interception and later re-injection into eIPoIB?
quoted
Well if this is all you want to do, you can bind a packet socket to the
interface, and drop them at the nic. It is harder to do for incoming
ARP requests though.
quoted
I would do something else: send ARPs out to some defined IB address.
This could be local host or queries from some SA property. Said remote
side could send you the responses in ethernet format so you do not need
to mangle responses at all. Similarly for incoming ARP requests.
quoted
The rule to do this can also just redirect non IP packets - this is IPoIB after all.
Thanks for the heads up on the possible implementation route, will
look into that.
quoted
quoted
Documentation we will fix,
quoted
And just to stress the point, document the limitations as well.
sure, not that I see concrete limitations for the **user** at this point, but
if there are such, will put them clearly written.
So far you are still playing with a design that is strongly NOT
ethernet. So calling it eIPoIB will continue to be a LIE.
You are still playing with an implementation that doesn't even dream
of supporting IPv6 which makes it so far from ethernet I can't imagine
anyone taking your code seriously.
All ethernet protocols not working except IPv4 is a huge concrete
limitation.
Any implementation that breaks a naive ARP implementation also breaks
IPv6. Not to mention everything else that runs over ethernet.
If you are clever you can use the current IPoIB hardware accelleration
but you need to do something different so that you can either encode
or imply the MAC address so you won't have to munge ethernet protocols.
Just for fun you might want to consider what it takes to support 2 VMs
in the same VLAN that share the same IP address (but different MAC
addresses) for failover purposes.
Eric
From: Or Gerlitz <hidden> Date: 2012-09-04 19:47:42
On Tue, Sep 4, 2012 at 10:31 PM, Eric W. Biederman
[off-list ref] wrote:
Or Gerlitz [off-list ref] writes:
quoted
On Tue, Sep 4, 2012 at 12:22 AM, Michael S. Tsirkin [off-list ref] wrote:
quoted
quoted
Documentation we will fix,
And just to stress the point, document the limitations as well.
sure, not that I see concrete limitations for the **user** at this point, but
if there are such, will put them clearly written.
All ethernet protocols not working except IPv4 is a huge concrete limitation.
Oh, sure, that WILL be documented, currently eIPoIB can deliver only what
IPoIB can which is IPv4, ARP/RARP, IPv6+ND, for IPv6 see next
So far you are still playing with a design that is strongly NOT
ethernet. So calling it eIPoIB will continue to be a LIE.
You are still playing with an implementation that doesn't even dream
of supporting IPv6 which makes it so far from ethernet I can't imagine
anyone taking your code seriously.
This design can and will support IPv6, the IPv6 ND handling will follow the path
we are talking now for the IPv4 ARPs, e.g not within the driver, etc.
Could you be
more specific?
Any implementation that breaks a naive ARP implementation also breaks
IPv6. Not to mention everything else that runs over ethernet.
not sure to follow on the naive impl. comment, eIPoIB solution will
include kernel driver along with supporting user-space portion, this
is to follow a comment made by the community on mangling ARPs in
network driver.
If you are clever you can use the current IPoIB hardware accelleration
but you need to do something different so that you can either encode
or imply the MAC address so you won't have to munge ethernet protocols.
also here not sure to follow, we have a new design under which the
original VM MAC is preserved on the RX side, we will not generate MAC
for Ethernet frames sent by VMs which we reconstruct on the RX side
any more.
Just for fun you might want to consider what it takes to support 2 VMs
in the same VLAN that share the same IP address (but different MAC
addresses) for failover purposes.
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2012-09-04 21:20:48
On Tue, Sep 04, 2012 at 09:50:09PM +0300, Or Gerlitz wrote:
quoted
And just to stress the point, document the limitations as well.
sure, not that I see concrete limitations for the **user** at this point, but
if there are such, will put them clearly written.
Hmm, I'm afraid you mistook a short list of some major
bugs that jumped out at me for an exhaustive list.
This was not intended as such.
Here's how to find some of the limitations in your design
1. look through list archives. Some where pointed out to you
2. list everything ethernet does that you dont, or do differently
3. list everything ipoib does that you don't or do differently
4. list any extra setup work required on behalf of the user
5. check various overheads, compare with native ipoib and alternatives
such as routing
6. if you still have an empty list of limitations and disadvantages,
look again :)
The point is to have documentation that is useful both
for reviewers - so they can know which bugs are known
and which need to be reported; and for users -
who should not be expected to be familiar with
internals of your implementation.
--
MST