v1 version of this RFC patch was posted at http://www.spinics.net/lists/netdev/msg174245.html
Today macvtap used in virtualized environment does not have support to
propagate MAC, VLAN and interface flags from guest to lowerdev.
Which means to be able to register additional VLANs, unicast and multicast
addresses or change pkt filter flags in the guest, the lowerdev has to be
put in promisocous mode. Today the only macvlan mode that supports this is
the PASSTHRU mode and it puts the lower dev in promiscous mode.
PASSTHRU mode was added primarily for the SRIOV usecase. In PASSTHRU mode
there is a 1-1 mapping between macvtap and physical NIC or VF.
There are two problems with putting the lowerdev in promiscous mode (ie SRIOV
VF's):
- Some SRIOV cards dont support promiscous mode today (Thread on Intel
driver indicates that http://lists.openwall.net/netdev/2011/09/27/6)
- For the SRIOV NICs that support it, Putting the lowerdev in
promiscous mode leads to additional traffic being sent up to the
guest virtio-net to filter result in extra overheads.
Both the above problems can be solved by offloading filtering to the
lowerdev hw. ie lowerdev does not need to be in promiscous mode as
long as the guest filters are passed down to the lowerdev.
This patch basically adds the infrastructure to set and get MAC and VLAN
filters on an interface via rtnetlink. And adds support in macvlan and macvtap
to allow set and get filter operations.
Earlier version of this patch provided the TUNSETTXFILTER macvtap interface
for setting address filtering. In response to feedback, This version
introduces a netlink interface for the same.
Response to some of the questions raised during v1:
- Netlink interface:
This patch provides the following netlink interface to set mac and vlan
filters :
[IFLA_RX_FILTER] = {
[IFLA_ADDR_FILTER] = {
[IFLA_ADDR_FILTER_FLAGS]
[IFLA_ADDR_FILTER_UC_LIST] = {
[IFLA_ADDR_LIST_ENTRY]
}
[IFLA_ADDR_FILTER_MC_LIST] = {
[IFLA_ADDR_LIST_ENTRY]
}
}
[IFLA_VLAN_FILTER] = {
[IFLA_VLAN_BITMAP]
}
}
Note: The IFLA_VLAN_FILTER is a nested attribute and contains only
IFLA_VLAN_BITMAP today. The idea is that the IFLA_VLAN_FILTER can
be extended tomorrow to use a vlan list option if some implementations
prefer a list instead.
And it provides the following rtnl_link_ops to set/get MAC/VLAN filters:
int (*set_rx_addr_filter)(struct net_device *dev,
struct nlattr *tb[]);
int (*set_rx_vlan_filter)(struct net_device *dev,
struct nlattr *tb[]);
size_t (*get_rx_addr_filter_size)(const struct
net_device *dev);
size_t (*get_rx_vlan_filter_size)(const struct
net_device *dev);
int (*fill_rx_addr_filter)(struct sk_buff *skb,
const struct net_device *dev);
int (*fill_rx_vlan_filter)(struct sk_buff *skb,
const struct net_device *dev);
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from rtnl_link_ops
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
- Protection against address spoofing:
- This patch adds filtering support only for macvtap PASSTHRU
Mode. PASSTHRU mode is used mainly with SRIOV VF's. And SRIOV VF's
come with anti mac/vlan spoofing support. (Recently added
IFLA_VF_SPOOFCHK). In 802.1Qbh case the port profile has a knob to
enable/disable anti spoof check. Lowerdevice drivers also enforce limits
on the number of address registrations allowed.
- Support for multiqueue devices: Enable filtering on individual queues (?):
AFAIK, there is no netdev interface to install per queue hw
filters for a multi queue interface. And also I dont know of any hw
that provides an interface to set hw filters on a per queue basis.
A multi queue device appears as a single lowerdev (ie netdev) and
uses the same uc and mc lists to setup unicast and multicast hw filters.
So i dont see a huge problem with this patch coming in the way for
multi queue devices.
- Support for non-PASSTHRU mode:
I started implementing this. But there are a couple of problems.
- The lowerdev may not be a SRIOV VF and may not have
anti spoof capability
- Today, in non-PASSTHRU cases macvlan_handle_frame assumes that
every macvlan device on top of the lowerdev has a single unique mac.
And the macvlans are hashed on that single mac address.
To support filtering for non-PASSTHRU mode in addition to this
patch the following needs to be done:
- non-passthru mode with a single macvlan over a lower dev
can be treated as PASSTHRU case
- For non-PASSTHRU mode with multiple macvlans over a single
lower dev:
- Multiple unicast mac's now need to be hashed to the
same macvlan device. The macvlan hash needs to change
for lookup based on any one of the multiple unicast
addresses a macvlan is interested in
- We need to consider vlans during the lookup too
- So the macvlan device hash needs to hash on both mac
and vlan
- But the support for filtering in non-PASSTHRU mode can be
built on this patch
This patch series implements the following
01/8 rtnetlink: Netlink interface for setting MAC and VLAN filters
02/8 rtnetlink: Add rtnl link operations for MAC address and VLAN filtering
03/8 rtnetlink: Add support to set MAC/VLAN filters
04/8 rtnetlink: Add support to get MAC/VLAN filters
05/8 macvlan: Add support to set MAC/VLAN filter rtnl link operations
06/8 macvlan: Add support to get MAC/VLAN filter rtnl link operations
07/8 macvtap: Add support to set MAC/VLAN filter rtnl link operations
08/8 macvtap: Add support to get MAC/VLAN filter rtnl link operations
Please comment. Thanks.
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
From: Roopa Prabhu <redacted>
This patch adds the following rtnl_link_ops to set and get MAC and VLAN
filters
set_rx_addr_filter - to set address filter
set_rx_vlan_filter - To set vlan filter
get_rx_addr_filter_size - To get address filter size
get_rx_vlan_filter_size - To get vlan filter size
fill_rx_addr_filter - To fill addr filter
fill_rx_vlan_filter - To fill vlan filter
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
include/net/rtnetlink.h | 13 +++++++++++++
1 files changed, 13 insertions(+), 0 deletions(-)
From: Roopa Prabhu <redacted>
This patch adds support in rtnetlink for IFLA_RX_FILTER set.
It adds code in do_setlink to parse IFLA_RX_FILTER and call
the rtnl_link_ops->set_rx_addr_filter and
rtnl_link_ops->set_rx_vlan_filter to set MAC and VLAN filters.
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
net/core/rtnetlink.c | 67 ++++++++++++++++++++++++++++++++++++++++++++++++++
1 files changed, 67 insertions(+), 0 deletions(-)
From: Roopa Prabhu <redacted>
This patch adds support in rtnetlink for IFLA_RX_FILTER get.
It adds new function rtnl_rx_filter_get_size to get the size
of rx filters by calling rtnl_link_ops->get_rx_addr_filter_size
and rtnl_link_ops->get_rx_vlan_filter_size
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
net/core/rtnetlink.c | 90 +++++++++++++++++++++++++++++++++++++++++++++++++-
1 files changed, 89 insertions(+), 1 deletions(-)
From: Roopa Prabhu <redacted>
This patch adds support to set MAC and VLAN filter rtnl_link_ops
on a macvlan interface. It adds support for set_rx_addr_filter and
set_rx_vlan_filter rtnl link operations. It currently supports
only macvlan PASSTHRU mode.
For passthru mode,
- Address filters: macvlan netdev uc and mc lists are
updated to reflect the addresses in the filter.
- VLAN filter: Currently applied vlan bitmap is maintained in
struct macvlan_dev->vlan_filter. This vlan bitmap is updated to
reflect the new bitmap that came in the netlink msg.
lowerdev hw vlan filter is updated using macvlan netdev operations
ndo_vlan_rx_add_vid and ndo_vlan_rx_kill_vid (which inturn call
lowerdev vlan add/kill netdev ops)
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
drivers/net/macvlan.c | 296 ++++++++++++++++++++++++++++++++++++++++----
include/linux/if_macvlan.h | 8 +
2 files changed, 279 insertions(+), 25 deletions(-)
From: Roopa Prabhu <redacted>
This patch adds support to get MAC and VLAN filter rtnl_link_ops
on a macvlan interface. It adds support for get_rx_addr_filter_size,
get_rx_vlan_filter_size, fill_rx_addr_filter and fill_rx_vlan_filter
rtnl link operations.
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
drivers/net/macvlan.c | 126 ++++++++++++++++++++++++++++++++++++++++++++
include/linux/if_macvlan.h | 10 +++
2 files changed, 136 insertions(+), 0 deletions(-)
@@ -1014,6 +1014,128 @@ int macvlan_set_rx_addr_filter(struct net_device *dev,}EXPORT_SYMBOL(macvlan_set_rx_addr_filter);+staticsize_tmacvlan_get_rx_addr_filter_passthru_size(+conststructnet_device*dev)+{+size_tsize;++/* IFLA_ADDR_FILTER_FLAGS */+size=nla_total_size(sizeof(u32));++if(netdev_uc_count(dev))+/* IFLA_ADDR_FILTER_UC_LIST */+size+=nla_total_size(netdev_uc_count(dev)*+ETH_ALEN*sizeof(structnlattr));++if(netdev_mc_count(dev))+/* IFLA_ADDR_FILTER_MC_LIST */+size+=nla_total_size(netdev_mc_count(dev)*+ETH_ALEN*sizeof(structnlattr));++returnsize;+}++size_tmacvlan_get_rx_addr_filter_size(conststructnet_device*dev)+{+structmacvlan_dev*vlan=netdev_priv(dev);++switch(vlan->mode){+caseMACVLAN_MODE_PASSTHRU:+returnmacvlan_get_rx_addr_filter_passthru_size(dev);+default:+return0;+}+}+EXPORT_SYMBOL(macvlan_get_rx_addr_filter_size);++size_tmacvlan_get_rx_vlan_filter_size(conststructnet_device*dev)+{+structmacvlan_dev*vlan=netdev_priv(dev);++switch(vlan->mode){+caseMACVLAN_MODE_PASSTHRU:+/* IFLA_VLAN_BITMAP */+returnnla_total_size(VLAN_BITMAP_SIZE);+default:+return0;+}+}+EXPORT_SYMBOL(macvlan_get_rx_vlan_filter_size);++staticintmacvlan_fill_rx_addr_filter_passthru(structsk_buff*skb,+conststructnet_device*dev)+{+structnlattr*uninitialized_var(uc_list),*mc_list;+structnetdev_hw_addr*ha;++NLA_PUT_U32(skb,IFLA_ADDR_FILTER_FLAGS,dev->flags&RX_FILTER_FLAGS);++if(netdev_uc_count(dev)){+uc_list=nla_nest_start(skb,IFLA_ADDR_FILTER_UC_LIST);+if(uc_list==NULL)+gotonla_put_failure;++netdev_for_each_uc_addr(ha,dev){+NLA_PUT(skb,IFLA_ADDR_LIST_ENTRY,ETH_ALEN,ha->addr);+}+nla_nest_end(skb,uc_list);+}++if(netdev_mc_count(dev)){+mc_list=nla_nest_start(skb,IFLA_ADDR_FILTER_MC_LIST);+if(mc_list==NULL)+gotonla_uc_list_cancel;++netdev_for_each_mc_addr(ha,dev){+NLA_PUT(skb,IFLA_ADDR_LIST_ENTRY,ETH_ALEN,ha->addr);+}+nla_nest_end(skb,mc_list);+}++return0;++nla_uc_list_cancel:+if(netdev_uc_count(dev))+nla_nest_cancel(skb,uc_list);+nla_put_failure:+return-EMSGSIZE;+}++intmacvlan_fill_rx_addr_filter(structsk_buff*skb,+conststructnet_device*dev)+{+structmacvlan_dev*vlan=netdev_priv(dev);++switch(vlan->mode){+caseMACVLAN_MODE_PASSTHRU:+returnmacvlan_fill_rx_addr_filter_passthru(skb,dev);+default:+return-ENODATA;/* No data to Fill */+}+}+EXPORT_SYMBOL(macvlan_fill_rx_addr_filter);++intmacvlan_fill_rx_vlan_filter(structsk_buff*skb,+conststructnet_device*dev)+{+structmacvlan_dev*vlan=netdev_priv(dev);++switch(vlan->mode){+caseMACVLAN_MODE_PASSTHRU:+NLA_PUT(skb,IFLA_VLAN_BITMAP,VLAN_BITMAP_SIZE,+vlan->vlan_filter);+break;+default:+return-ENODATA;/* No data to Fill */+}++return0;++nla_put_failure:+return-EMSGSIZE;+}+EXPORT_SYMBOL(macvlan_fill_rx_vlan_filter);+staticconststructnla_policymacvlan_policy[IFLA_MACVLAN_MAX+1]={[IFLA_MACVLAN_MODE]={.type=NLA_U32},};
From: Roopa Prabhu <redacted>
This patch adds support to set MAC and VLAN filter rtnl_link_ops
on a macvtap interface. It adds support for set_rx_addr_filter and
set_rx_vlan_filter rtnl link operations. These operations inturn call the
equivalent operations defined in macvlan
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
drivers/net/macvtap.c | 22 ++++++++++++++++++----
1 files changed, 18 insertions(+), 4 deletions(-)
From: Roopa Prabhu <redacted>
This patch adds support to get MAC and VLAN filter rtnl_link_ops
on a macvtap interface. It adds support for get_rx_addr_filter_size,
get_rx_vlan_filter_size, fill_rx_addr_filter and fill_rx_vlan_filter
rtnl link operations. Calls equivalent macvlan operations.
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
drivers/net/macvtap.c | 27 +++++++++++++++++++++++++++
1 files changed, 27 insertions(+), 0 deletions(-)
From: Rose, Gregory V <hidden> Date: 2011-10-19 21:06:42
-----Original Message-----
From: netdev-owner@vger.kernel.org [mailto:netdev-owner@vger.kernel.org]
On Behalf Of Roopa Prabhu
Sent: Tuesday, October 18, 2011 11:26 PM
To: netdev@vger.kernel.org
Cc: sri@us.ibm.com; dragos.tatulea@gmail.com; arnd@arndb.de;
kvm@vger.kernel.org; mst@redhat.com; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com
Subject: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address filtering
support for passthru mode
[snip...]
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from rtnl_link_ops
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
I'm still digesting the rest of the RFC patches but I did want to quickly jump
in and push for adding this support in netdev_ops. I would like to see these
features available in more devices than just macvtap and macvlan. I can conceive
of use cases for multiple HW MAC and VLAN filters for a VF device that isn't
owned by a macvlan/macvtap interface and only has netdev_ops support. In this
case it would be necessary to program the filters directly to the VF device
interface or PF interface (or lowerdev as you refer to it) instead of going
through macvlan/macvtap.
This work dovetails nicely with some work I've been doing and I'd be very interested
in helping move this forward if we could work out the details that would allow support
of the features we (and the community) require.
- Greg Rose
LAN Access Division
Intel Corp.
On 10/19/11 2:06 PM, "Rose, Gregory V" [off-list ref] wrote:
quoted
-----Original Message-----
From: netdev-owner@vger.kernel.org [mailto:netdev-owner@vger.kernel.org]
On Behalf Of Roopa Prabhu
Sent: Tuesday, October 18, 2011 11:26 PM
To: netdev@vger.kernel.org
Cc: sri@us.ibm.com; dragos.tatulea@gmail.com; arnd@arndb.de;
kvm@vger.kernel.org; mst@redhat.com; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com
Subject: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address filtering
support for passthru mode
[snip...]
quoted
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from rtnl_link_ops
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
I'm still digesting the rest of the RFC patches but I did want to quickly jump
in and push for adding this support in netdev_ops. I would like to see these
features available in more devices than just macvtap and macvlan. I can
conceive
of use cases for multiple HW MAC and VLAN filters for a VF device that isn't
owned by a macvlan/macvtap interface and only has netdev_ops support. In this
case it would be necessary to program the filters directly to the VF device
interface or PF interface (or lowerdev as you refer to it) instead of going
through macvlan/macvtap.
This work dovetails nicely with some work I've been doing and I'd be very
interested
in helping move this forward if we could work out the details that would allow
support
of the features we (and the community) require.
Great. Thanks. I will definitely be interested to get this patch working for
any other use case you have.
Moving the ops to netdev should be trivial. You probably want the ops to
work on the VF via the PF, like the existing ndo_set_vf_mac etc.
Yes, lets work out the details and I can move this to netdev->ops. Let me
know.
Thanks,
Roopa
On Behalf Of Roopa Prabhu
Sent: Tuesday, October 18, 2011 11:26 PM
To: netdev@vger.kernel.org
Cc: sri@us.ibm.com; dragos.tatulea@gmail.com; arnd@arndb.de;
kvm@vger.kernel.org; mst@redhat.com; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com
Subject: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address filtering
support for passthru mode
[snip...]
quoted
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from rtnl_link_ops
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
I'm still digesting the rest of the RFC patches but I did want to
quickly jump
quoted
in and push for adding this support in netdev_ops. I would like to see
these
quoted
features available in more devices than just macvtap and macvlan. I can
conceive
of use cases for multiple HW MAC and VLAN filters for a VF device that
isn't
quoted
owned by a macvlan/macvtap interface and only has netdev_ops support.
In this
quoted
case it would be necessary to program the filters directly to the VF
device
quoted
interface or PF interface (or lowerdev as you refer to it) instead of
going
quoted
through macvlan/macvtap.
This work dovetails nicely with some work I've been doing and I'd be
very
quoted
interested
in helping move this forward if we could work out the details that would
allow
quoted
support
of the features we (and the community) require.
Great. Thanks. I will definitely be interested to get this patch working
for
any other use case you have.
Moving the ops to netdev should be trivial. You probably want the ops to
work on the VF via the PF, like the existing ndo_set_vf_mac etc.
That is correct, so we would need to add some way to pass the VF number to the op.
In addition, there are use cases for multiple MAC address filters for the Physical
Function (PF) so we would like to be able to identify to the netdev op that it is
supposed to perform the action on the PF filters instead of a VF.
An example of this would be when an administrator has created some number of VFs
for a given PF but is also running the PF in bridged (i.e. promiscuous) mode so that it
can support purely SW emulated network connections in some VMs that have low network
latency and bandwidth requirements while reserving the VFs for VMs that require the low latency, high throughput that directly assigned VFs can provide. In this case an
emulated SW interface in a VM is unable to properly communicate with VFs on the same
PF because the emulated SW interface's MAC address isn't programmed into the HW filters
on the PF. If we could use this op to program the MAC address and VLAN filters of
the emulated SW interfaces into the PF HW a VF could then properly communicate across
the NIC's internal VEB to the emulated SW interfaces.
Yes, lets work out the details and I can move this to netdev->ops. Let me
know.
I think essentially if you could add some parameter to the ops to specify whether it
is addressing a VF or the PF and then if it is a VF further specify the VF number we
would be very close to addressing the requirements of many valuable use cases in
addition to the ones you have identified in your RFC.
Does that sound reasonable?
Thanks,
- Greg
On Behalf Of Roopa Prabhu
Sent: Tuesday, October 18, 2011 11:26 PM
To: netdev@vger.kernel.org
Cc: sri@us.ibm.com; dragos.tatulea@gmail.com; arnd@arndb.de;
kvm@vger.kernel.org; mst@redhat.com; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com
Subject: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address
filtering
quoted
quoted
quoted
support for passthru mode
[snip...]
quoted
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from
rtnl_link_ops
quoted
quoted
quoted
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
I'm still digesting the rest of the RFC patches but I did want to
quickly jump
quoted
in and push for adding this support in netdev_ops. I would like to
see
quoted
these
quoted
features available in more devices than just macvtap and macvlan. I
can
quoted
quoted
conceive
of use cases for multiple HW MAC and VLAN filters for a VF device that
isn't
quoted
owned by a macvlan/macvtap interface and only has netdev_ops support.
In this
quoted
case it would be necessary to program the filters directly to the VF
device
quoted
interface or PF interface (or lowerdev as you refer to it) instead of
going
quoted
through macvlan/macvtap.
This work dovetails nicely with some work I've been doing and I'd be
very
quoted
interested
in helping move this forward if we could work out the details that
would
quoted
allow
quoted
support
of the features we (and the community) require.
Great. Thanks. I will definitely be interested to get this patch working
for
any other use case you have.
Moving the ops to netdev should be trivial. You probably want the ops to
work on the VF via the PF, like the existing ndo_set_vf_mac etc.
That is correct, so we would need to add some way to pass the VF number to
the op.
In addition, there are use cases for multiple MAC address filters for the
Physical
Function (PF) so we would like to be able to identify to the netdev op
that it is
supposed to perform the action on the PF filters instead of a VF.
An example of this would be when an administrator has created some number of VFs
for a given PF but is also running the PF in bridged (i.e. promiscuous)mode so
that it can support purely SW emulated network connections in some VMs that have
low network latency and bandwidth requirements while reserving the VFs for VMs that
On Behalf Of Roopa Prabhu
Sent: Tuesday, October 18, 2011 11:26 PM
To: netdev@vger.kernel.org
Cc: sri@us.ibm.com; dragos.tatulea@gmail.com; arnd@arndb.de;
kvm@vger.kernel.org; mst@redhat.com; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com
Subject: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address filtering
support for passthru mode
[snip...]
quoted
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from rtnl_link_ops
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
I'm still digesting the rest of the RFC patches but I did want to
quickly jump
quoted
in and push for adding this support in netdev_ops. I would like to see
these
quoted
features available in more devices than just macvtap and macvlan. I can
conceive
of use cases for multiple HW MAC and VLAN filters for a VF device that
isn't
quoted
owned by a macvlan/macvtap interface and only has netdev_ops support.
In this
quoted
case it would be necessary to program the filters directly to the VF
device
quoted
interface or PF interface (or lowerdev as you refer to it) instead of
going
quoted
through macvlan/macvtap.
This work dovetails nicely with some work I've been doing and I'd be
very
quoted
interested
in helping move this forward if we could work out the details that would
allow
quoted
support
of the features we (and the community) require.
Great. Thanks. I will definitely be interested to get this patch working
for
any other use case you have.
Moving the ops to netdev should be trivial. You probably want the ops to
work on the VF via the PF, like the existing ndo_set_vf_mac etc.
That is correct, so we would need to add some way to pass the VF number to the
op.
In addition, there are use cases for multiple MAC address filters for the
Physical
Function (PF) so we would like to be able to identify to the netdev op that it
is
supposed to perform the action on the PF filters instead of a VF.
An example of this would be when an administrator has created some number of
VFs
for a given PF but is also running the PF in bridged (i.e. promiscuous) mode
so that it
can support purely SW emulated network connections in some VMs that have low
network
latency and bandwidth requirements while reserving the VFs for VMs that
require the low latency, high throughput that directly assigned VFs can
provide. In this case an
emulated SW interface in a VM is unable to properly communicate with VFs on
the same
PF because the emulated SW interface's MAC address isn't programmed into the
HW filters
on the PF. If we could use this op to program the MAC address and VLAN
filters of
the emulated SW interfaces into the PF HW a VF could then properly communicate
across
the NIC's internal VEB to the emulated SW interfaces.
quoted
Yes, lets work out the details and I can move this to netdev->ops. Let me
know.
I think essentially if you could add some parameter to the ops to specify
whether it
is addressing a VF or the PF and then if it is a VF further specify the VF
number we
would be very close to addressing the requirements of many valuable use cases
in
addition to the ones you have identified in your RFC.
Does that sound reasonable?
Thanks for the details Greg. Sounds good. I will change it to provide netdev
ops with a vf argument and respin.
Thanks,
Roopa
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2011-10-24 05:46:20
On Tue, Oct 18, 2011 at 11:25:54PM -0700, Roopa Prabhu wrote:
v1 version of this RFC patch was posted at http://www.spinics.net/lists/netdev/msg174245.html
Today macvtap used in virtualized environment does not have support to
propagate MAC, VLAN and interface flags from guest to lowerdev.
Which means to be able to register additional VLANs, unicast and multicast
addresses or change pkt filter flags in the guest, the lowerdev has to be
put in promisocous mode. Today the only macvlan mode that supports this is
the PASSTHRU mode and it puts the lower dev in promiscous mode.
PASSTHRU mode was added primarily for the SRIOV usecase. In PASSTHRU mode
there is a 1-1 mapping between macvtap and physical NIC or VF.
There are two problems with putting the lowerdev in promiscous mode (ie SRIOV
VF's):
- Some SRIOV cards dont support promiscous mode today (Thread on Intel
driver indicates that http://lists.openwall.net/netdev/2011/09/27/6)
- For the SRIOV NICs that support it, Putting the lowerdev in
promiscous mode leads to additional traffic being sent up to the
guest virtio-net to filter result in extra overheads.
Both the above problems can be solved by offloading filtering to the
lowerdev hw. ie lowerdev does not need to be in promiscous mode as
long as the guest filters are passed down to the lowerdev.
This patch basically adds the infrastructure to set and get MAC and VLAN
filters on an interface via rtnetlink. And adds support in macvlan and macvtap
to allow set and get filter operations.
Looks sane to me. Some minor comments below.
Earlier version of this patch provided the TUNSETTXFILTER macvtap interface
for setting address filtering. In response to feedback, This version
introduces a netlink interface for the same.
Response to some of the questions raised during v1:
- Netlink interface:
This patch provides the following netlink interface to set mac and vlan
filters :
[IFLA_RX_FILTER] = {
[IFLA_ADDR_FILTER] = {
[IFLA_ADDR_FILTER_FLAGS]
[IFLA_ADDR_FILTER_UC_LIST] = {
[IFLA_ADDR_LIST_ENTRY]
}
[IFLA_ADDR_FILTER_MC_LIST] = {
[IFLA_ADDR_LIST_ENTRY]
}
}
[IFLA_VLAN_FILTER] = {
[IFLA_VLAN_BITMAP]
}
}
Note: The IFLA_VLAN_FILTER is a nested attribute and contains only
IFLA_VLAN_BITMAP today. The idea is that the IFLA_VLAN_FILTER can
be extended tomorrow to use a vlan list option if some implementations
prefer a list instead.
And it provides the following rtnl_link_ops to set/get MAC/VLAN filters:
int (*set_rx_addr_filter)(struct net_device *dev,
struct nlattr *tb[]);
int (*set_rx_vlan_filter)(struct net_device *dev,
struct nlattr *tb[]);
size_t (*get_rx_addr_filter_size)(const struct
net_device *dev);
size_t (*get_rx_vlan_filter_size)(const struct
net_device *dev);
int (*fill_rx_addr_filter)(struct sk_buff *skb,
const struct net_device *dev);
int (*fill_rx_vlan_filter)(struct sk_buff *skb,
const struct net_device *dev);
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from rtnl_link_ops
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
- Protection against address spoofing:
- This patch adds filtering support only for macvtap PASSTHRU
Mode. PASSTHRU mode is used mainly with SRIOV VF's. And SRIOV VF's
come with anti mac/vlan spoofing support. (Recently added
IFLA_VF_SPOOFCHK). In 802.1Qbh case the port profile has a knob to
enable/disable anti spoof check. Lowerdevice drivers also enforce limits
on the number of address registrations allowed.
- Support for multiqueue devices: Enable filtering on individual queues (?):
AFAIK, there is no netdev interface to install per queue hw
filters for a multi queue interface. And also I dont know of any hw
that provides an interface to set hw filters on a per queue basis.
VMDq hardware would support this, no?
A multi queue device appears as a single lowerdev (ie netdev) and
uses the same uc and mc lists to setup unicast and multicast hw filters.
So i dont see a huge problem with this patch coming in the way for
multi queue devices.
- Support for non-PASSTHRU mode:
I started implementing this. But there are a couple of problems.
- The lowerdev may not be a SRIOV VF and may not have
anti spoof capability
Anti-spoofing a really a separate feature, isn't it?
- Today, in non-PASSTHRU cases macvlan_handle_frame assumes that
every macvlan device on top of the lowerdev has a single unique mac.
And the macvlans are hashed on that single mac address.
To support filtering for non-PASSTHRU mode in addition to this
patch the following needs to be done:
- non-passthru mode with a single macvlan over a lower dev
can be treated as PASSTHRU case
- For non-PASSTHRU mode with multiple macvlans over a single
lower dev:
- Multiple unicast mac's now need to be hashed to the
same macvlan device. The macvlan hash needs to change
for lookup based on any one of the multiple unicast
addresses a macvlan is interested in
- We need to consider vlans during the lookup too
- So the macvlan device hash needs to hash on both mac
and vlan
It might be useful to expose the filters to the device.
- But the support for filtering in non-PASSTHRU mode can be
built on this patch
Agree, this can be added gradually.
This patch series implements the following
01/8 rtnetlink: Netlink interface for setting MAC and VLAN filters
02/8 rtnetlink: Add rtnl link operations for MAC address and VLAN filtering
03/8 rtnetlink: Add support to set MAC/VLAN filters
04/8 rtnetlink: Add support to get MAC/VLAN filters
05/8 macvlan: Add support to set MAC/VLAN filter rtnl link operations
06/8 macvlan: Add support to get MAC/VLAN filter rtnl link operations
07/8 macvtap: Add support to set MAC/VLAN filter rtnl link operations
08/8 macvtap: Add support to get MAC/VLAN filter rtnl link operations
Please comment. Thanks.
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2011-10-24 05:56:34
On Tue, Oct 18, 2011 at 11:26:36PM -0700, Roopa Prabhu wrote:
quoted hunk
From: Roopa Prabhu <redacted>
This patch adds support to get MAC and VLAN filter rtnl_link_ops
on a macvtap interface. It adds support for get_rx_addr_filter_size,
get_rx_vlan_filter_size, fill_rx_addr_filter and fill_rx_vlan_filter
rtnl link operations. Calls equivalent macvlan operations.
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
drivers/net/macvtap.c | 27 +++++++++++++++++++++++++++
1 files changed, 27 insertions(+), 0 deletions(-)
From: "Michael S. Tsirkin" <mst@redhat.com> Date: 2011-10-24 05:57:25
On Tue, Oct 18, 2011 at 11:26:30PM -0700, Roopa Prabhu wrote:
quoted hunk
From: Roopa Prabhu <redacted>
This patch adds support to set MAC and VLAN filter rtnl_link_ops
on a macvtap interface. It adds support for set_rx_addr_filter and
set_rx_vlan_filter rtnl link operations. These operations inturn call the
equivalent operations defined in macvlan
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
drivers/net/macvtap.c | 22 ++++++++++++++++++----
1 files changed, 18 insertions(+), 4 deletions(-)
On 10/23/11 10:47 PM, "Michael S. Tsirkin" [off-list ref] wrote:
On Tue, Oct 18, 2011 at 11:25:54PM -0700, Roopa Prabhu wrote:
quoted
v1 version of this RFC patch was posted at
http://www.spinics.net/lists/netdev/msg174245.html
Today macvtap used in virtualized environment does not have support to
propagate MAC, VLAN and interface flags from guest to lowerdev.
Which means to be able to register additional VLANs, unicast and multicast
addresses or change pkt filter flags in the guest, the lowerdev has to be
put in promisocous mode. Today the only macvlan mode that supports this is
the PASSTHRU mode and it puts the lower dev in promiscous mode.
PASSTHRU mode was added primarily for the SRIOV usecase. In PASSTHRU mode
there is a 1-1 mapping between macvtap and physical NIC or VF.
There are two problems with putting the lowerdev in promiscous mode (ie SRIOV
VF's):
- Some SRIOV cards dont support promiscous mode today (Thread on Intel
driver indicates that http://lists.openwall.net/netdev/2011/09/27/6)
- For the SRIOV NICs that support it, Putting the lowerdev in
promiscous mode leads to additional traffic being sent up to the
guest virtio-net to filter result in extra overheads.
Both the above problems can be solved by offloading filtering to the
lowerdev hw. ie lowerdev does not need to be in promiscous mode as
long as the guest filters are passed down to the lowerdev.
This patch basically adds the infrastructure to set and get MAC and VLAN
filters on an interface via rtnetlink. And adds support in macvlan and
macvtap
to allow set and get filter operations.
Looks sane to me. Some minor comments below.
quoted
Earlier version of this patch provided the TUNSETTXFILTER macvtap interface
for setting address filtering. In response to feedback, This version
introduces a netlink interface for the same.
Response to some of the questions raised during v1:
- Netlink interface:
This patch provides the following netlink interface to set mac and vlan
filters :
[IFLA_RX_FILTER] = {
[IFLA_ADDR_FILTER] = {
[IFLA_ADDR_FILTER_FLAGS]
[IFLA_ADDR_FILTER_UC_LIST] = {
[IFLA_ADDR_LIST_ENTRY]
}
[IFLA_ADDR_FILTER_MC_LIST] = {
[IFLA_ADDR_LIST_ENTRY]
}
}
[IFLA_VLAN_FILTER] = {
[IFLA_VLAN_BITMAP]
}
}
Note: The IFLA_VLAN_FILTER is a nested attribute and contains only
IFLA_VLAN_BITMAP today. The idea is that the IFLA_VLAN_FILTER can
be extended tomorrow to use a vlan list option if some implementations
prefer a list instead.
And it provides the following rtnl_link_ops to set/get MAC/VLAN filters:
int (*set_rx_addr_filter)(struct net_device *dev,
struct nlattr *tb[]);
int (*set_rx_vlan_filter)(struct net_device *dev,
struct nlattr *tb[]);
size_t (*get_rx_addr_filter_size)(const struct
net_device *dev);
size_t (*get_rx_vlan_filter_size)(const struct
net_device *dev);
int (*fill_rx_addr_filter)(struct sk_buff *skb,
const struct net_device
*dev);
int (*fill_rx_vlan_filter)(struct sk_buff *skb,
const struct net_device
*dev);
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from rtnl_link_ops
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
- Protection against address spoofing:
- This patch adds filtering support only for macvtap PASSTHRU
Mode. PASSTHRU mode is used mainly with SRIOV VF's. And SRIOV VF's
come with anti mac/vlan spoofing support. (Recently added
IFLA_VF_SPOOFCHK). In 802.1Qbh case the port profile has a knob to
enable/disable anti spoof check. Lowerdevice drivers also enforce limits
on the number of address registrations allowed.
- Support for multiqueue devices: Enable filtering on individual queues (?):
AFAIK, there is no netdev interface to install per queue hw
filters for a multi queue interface. And also I dont know of any hw
that provides an interface to set hw filters on a per queue basis.
VMDq hardware would support this, no?
Am not really sure. This patch uses netdev to pass filters to hw. And I
don't see any netdev infrastructure that would support per queue filters.
Maybe Greg (CC'ed) or anyone else from Intel can answer this.
Greg, michael had brought up this question during first version of these
patches as well. Will be nice to get the VMDq requirements for propagating
guest filters to hw clarified. Do you see any special VMDq nic requirement
we can cover in this patch. This is for VMDq queues directly connected to
guest nics. Thanks.
quoted
A multi queue device appears as a single lowerdev (ie netdev) and
uses the same uc and mc lists to setup unicast and multicast hw filters.
So i dont see a huge problem with this patch coming in the way for
multi queue devices.
- Support for non-PASSTHRU mode:
I started implementing this. But there are a couple of problems.
- The lowerdev may not be a SRIOV VF and may not have
anti spoof capability
Anti-spoofing a really a separate feature, isn't it?
Yes that is correct. It really should not be a concern with implementing
support for non-PASSTHRU mode. The only intent of adding the above line was
that eventually we should probably think of supporting anti-spoof feature on
Non-sriov devices if they are accepting filters from the guest.
I think I will move the above line to some place else more appropriate in
the comment log instead of covering it as part of the non-passthru macvlan
implementation.
quoted
- Today, in non-PASSTHRU cases macvlan_handle_frame assumes that
every macvlan device on top of the lowerdev has a single unique mac.
And the macvlans are hashed on that single mac address.
To support filtering for non-PASSTHRU mode in addition to this
patch the following needs to be done:
- non-passthru mode with a single macvlan over a lower dev
can be treated as PASSTHRU case
- For non-PASSTHRU mode with multiple macvlans over a single
lower dev:
- Multiple unicast mac's now need to be hashed to the
same macvlan device. The macvlan hash needs to change
for lookup based on any one of the multiple unicast
addresses a macvlan is interested in
- We need to consider vlans during the lookup too
- So the macvlan device hash needs to hash on both mac
and vlan
It might be useful to expose the filters to the device.
Yes
quoted
- But the support for filtering in non-PASSTHRU mode can be
built on this patch
Agree, this can be added gradually.
Ok thanks. Currently testing newer version of these patches, will post them
Sometime this week.
From: Rose, Gregory V <hidden> Date: 2011-10-24 21:51:29
-----Original Message-----
From: netdev-owner@vger.kernel.org [mailto:netdev-owner@vger.kernel.org]
On Behalf Of Roopa Prabhu
Sent: Monday, October 24, 2011 11:15 AM
To: Michael S. Tsirkin
Cc: netdev@vger.kernel.org; sri@us.ibm.com; dragos.tatulea@gmail.com;
arnd@arndb.de; kvm@vger.kernel.org; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com; Rose, Gregory V
Subject: Re: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address
filtering support for passthru mode
On 10/23/11 10:47 PM, "Michael S. Tsirkin" [off-list ref] wrote:
quoted
quoted
AFAIK, there is no netdev interface to install per queue hw
filters for a multi queue interface. And also I dont know of any hw
that provides an interface to set hw filters on a per queue basis.
VMDq hardware would support this, no?
Am not really sure. This patch uses netdev to pass filters to hw. And I
don't see any netdev infrastructure that would support per queue filters.
Maybe Greg (CC'ed) or anyone else from Intel can answer this.
Greg, michael had brought up this question during first version of these
patches as well. Will be nice to get the VMDq requirements for propagating
guest filters to hw clarified. Do you see any special VMDq nic requirement
we can cover in this patch. This is for VMDq queues directly connected to
guest nics. Thanks.
So far as I know there is no support for VMDq in the Linux kernel and while I know some folks have been working on it I can't really speak to that work or their plans. Much would depend on the implementation.
For now it makes sense to me to get support for multiple MAC and VLAN filters per virtual function (or virtual nic) and it seems to me you're going in the right direction for this. We'll have a look at your next set of patches and take it from there.
- Greg
From: Rose, Gregory V <hidden> Date: 2011-10-25 15:59:43
-----Original Message-----
From: Michael S. Tsirkin [mailto:mst@redhat.com]
Sent: Tuesday, October 25, 2011 8:46 AM
To: Roopa Prabhu
Cc: netdev@vger.kernel.org; sri@us.ibm.com; dragos.tatulea@gmail.com;
arnd@arndb.de; kvm@vger.kernel.org; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com; Rose, Gregory V
Subject: Re: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address
filtering support for passthru mode
On Mon, Oct 24, 2011 at 11:15:05AM -0700, Roopa Prabhu wrote:
quoted
quoted
quoted
... And also I dont know of any hw
that provides an interface to set hw filters on a per queue basis.
VMDq hardware would support this, no?
Am not really sure. This patch uses netdev to pass filters to hw. And I
don't see any netdev infrastructure that would support per queue
filters.
Sure. I was only saying that as far as I understand,
VMDq hardware does support this functionality.
Right, Greg?
In the case of Intel HW yes. I'll refrain from speaking for other HW vendors although I'm guessing it would be true in their cases also. YMMV, caveat emptor, etc. etc.
- Greg
On 10/23/11 10:56 PM, "Michael S. Tsirkin" [off-list ref] wrote:
On Tue, Oct 18, 2011 at 11:26:36PM -0700, Roopa Prabhu wrote:
quoted
From: Roopa Prabhu <redacted>
This patch adds support to get MAC and VLAN filter rtnl_link_ops
on a macvtap interface. It adds support for get_rx_addr_filter_size,
get_rx_vlan_filter_size, fill_rx_addr_filter and fill_rx_vlan_filter
rtnl link operations. Calls equivalent macvlan operations.
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
---
drivers/net/macvtap.c | 27 +++++++++++++++++++++++++++
1 files changed, 27 insertions(+), 0 deletions(-)
So why do we need the above wrappers? Can't use macvlanXXX directly?
I had followed the existing macvtap rtnl_link_ops convention here.
It seems cleaner this way. You can define the macvtap ops static and
Call equivalent macvlan functions from it if required. It also gives you
flexibility in adding any macvtap specific stuff before or after you call
the macvlan equivalent function (like some of the macvtap rtnl link ops
already do today)
In any case this part and the below empty line error goes away in the new
version.
Thanks,
Roopa
quoted
+
+
don't add double emoty lines pls.
quoted
static int macvtap_newlink(struct net *src_net,
struct net_device *dev,
struct nlattr *tb[],
From: Rose, Gregory V <hidden> Date: 2011-11-08 18:31:35
-----Original Message-----
From: netdev-owner@vger.kernel.org [mailto:netdev-owner@vger.kernel.org]
On Behalf Of Roopa Prabhu
Sent: Tuesday, October 18, 2011 11:26 PM
To: netdev@vger.kernel.org
Cc: sri@us.ibm.com; dragos.tatulea@gmail.com; arnd@arndb.de;
kvm@vger.kernel.org; mst@redhat.com; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com
Subject: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address filtering
support for passthru mode
v1 version of this RFC patch was posted at
http://www.spinics.net/lists/netdev/msg174245.html
Today macvtap used in virtualized environment does not have support to
propagate MAC, VLAN and interface flags from guest to lowerdev.
Which means to be able to register additional VLANs, unicast and multicast
addresses or change pkt filter flags in the guest, the lowerdev has to be
put in promisocous mode. Today the only macvlan mode that supports this is
the PASSTHRU mode and it puts the lower dev in promiscous mode.
PASSTHRU mode was added primarily for the SRIOV usecase. In PASSTHRU mode
there is a 1-1 mapping between macvtap and physical NIC or VF.
There are two problems with putting the lowerdev in promiscous mode (ie
SRIOV
VF's):
- Some SRIOV cards dont support promiscous mode today (Thread on
Intel
driver indicates that http://lists.openwall.net/netdev/2011/09/27/6)
- For the SRIOV NICs that support it, Putting the lowerdev in
promiscous mode leads to additional traffic being sent up to the
guest virtio-net to filter result in extra overheads.
Both the above problems can be solved by offloading filtering to the
lowerdev hw. ie lowerdev does not need to be in promiscous mode as
long as the guest filters are passed down to the lowerdev.
This patch basically adds the infrastructure to set and get MAC and VLAN
filters on an interface via rtnetlink. And adds support in macvlan and
macvtap
to allow set and get filter operations.
Earlier version of this patch provided the TUNSETTXFILTER macvtap
interface
for setting address filtering. In response to feedback, This version
introduces a netlink interface for the same.
Response to some of the questions raised during v1:
- Netlink interface:
This patch provides the following netlink interface to set mac and
vlan
filters :
[IFLA_RX_FILTER] = {
[IFLA_ADDR_FILTER] = {
[IFLA_ADDR_FILTER_FLAGS]
[IFLA_ADDR_FILTER_UC_LIST] = {
[IFLA_ADDR_LIST_ENTRY]
}
[IFLA_ADDR_FILTER_MC_LIST] = {
[IFLA_ADDR_LIST_ENTRY]
}
}
[IFLA_VLAN_FILTER] = {
[IFLA_VLAN_BITMAP]
}
}
Note: The IFLA_VLAN_FILTER is a nested attribute and contains only
IFLA_VLAN_BITMAP today. The idea is that the IFLA_VLAN_FILTER can
be extended tomorrow to use a vlan list option if some
implementations
prefer a list instead.
And it provides the following rtnl_link_ops to set/get MAC/VLAN
filters:
int (*set_rx_addr_filter)(struct net_device
*dev,
struct nlattr *tb[]);
int (*set_rx_vlan_filter)(struct net_device
*dev,
struct nlattr *tb[]);
size_t (*get_rx_addr_filter_size)(const struct
net_device *dev);
size_t (*get_rx_vlan_filter_size)(const struct
net_device *dev);
int (*fill_rx_addr_filter)(struct sk_buff *skb,
const struct net_device
*dev);
int (*fill_rx_vlan_filter)(struct sk_buff *skb,
const struct net_device
*dev);
Note: The choice of rtnl_link_ops was because I saw the use case for
this in virtual devices that need to do filtering in sw like
macvlan
and tun. Hw devices usually have filtering in hw with netdev->uc and
mc lists to indicate active filters. But I can move from
rtnl_link_ops
to netdev_ops if that is the preferred way to go and if there is a
need to support this interface on all kinds of interfaces.
Please suggest.
- Protection against address spoofing:
- This patch adds filtering support only for macvtap PASSTHRU
Mode. PASSTHRU mode is used mainly with SRIOV VF's. And SRIOV VF's
come with anti mac/vlan spoofing support. (Recently added
IFLA_VF_SPOOFCHK). In 802.1Qbh case the port profile has a knob to
enable/disable anti spoof check. Lowerdevice drivers also enforce
limits
on the number of address registrations allowed.
- Support for multiqueue devices: Enable filtering on individual queues
(?):
AFAIK, there is no netdev interface to install per queue hw
filters for a multi queue interface. And also I dont know of any hw
that provides an interface to set hw filters on a per queue basis.
A multi queue device appears as a single lowerdev (ie netdev) and
uses the same uc and mc lists to setup unicast and multicast hw
filters.
So i dont see a huge problem with this patch coming in the way for
multi queue devices.
- Support for non-PASSTHRU mode:
I started implementing this. But there are a couple of problems.
- The lowerdev may not be a SRIOV VF and may not have
anti spoof capability
- Today, in non-PASSTHRU cases macvlan_handle_frame assumes that
every macvlan device on top of the lowerdev has a single unique mac.
And the macvlans are hashed on that single mac address.
To support filtering for non-PASSTHRU mode in addition to this
patch the following needs to be done:
- non-passthru mode with a single macvlan over a lower dev
can be treated as PASSTHRU case
- For non-PASSTHRU mode with multiple macvlans over a single
lower dev:
- Multiple unicast mac's now need to be hashed to the
same macvlan device. The macvlan hash needs to change
for lookup based on any one of the multiple unicast
addresses a macvlan is interested in
- We need to consider vlans during the lookup too
- So the macvlan device hash needs to hash on both mac
and vlan
- But the support for filtering in non-PASSTHRU mode can be
built on this patch
This patch series implements the following
01/8 rtnetlink: Netlink interface for setting MAC and VLAN filters
02/8 rtnetlink: Add rtnl link operations for MAC address and VLAN
filtering
03/8 rtnetlink: Add support to set MAC/VLAN filters
04/8 rtnetlink: Add support to get MAC/VLAN filters
05/8 macvlan: Add support to set MAC/VLAN filter rtnl link operations
06/8 macvlan: Add support to get MAC/VLAN filter rtnl link operations
07/8 macvtap: Add support to set MAC/VLAN filter rtnl link operations
08/8 macvtap: Add support to get MAC/VLAN filter rtnl link operations
Please comment. Thanks.
I have finished my preliminary evaluation and testing of these RFC patches and find that the features that we would require are present and functional. I haven't fully fleshed out the 'get' operations in my own driver yet but the 'set' routines are all doing what we need them to do.
Thanks for your work on this Roopa!
- Greg
Signed-off-by: Roopa Prabhu <redacted>
Signed-off-by: Christian Benvenuti <redacted>
Signed-off-by: David Wang <redacted>
--
To unsubscribe from this list: send the line "unsubscribe netdev" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Moving the ops to netdev should be trivial. You probably want the ops to
work on the VF via the PF, like the existing ndo_set_vf_mac etc.
That is correct, so we would need to add some way to pass the VF number to the op.
In addition, there are use cases for multiple MAC address filters for the Physical
Function (PF) so we would like to be able to identify to the netdev op that it is
supposed to perform the action on the PF filters instead of a VF.
An example of this would be when an administrator has created some number of VFs
for a given PF but is also running the PF in bridged (i.e. promiscuous) mode so that it
can support purely SW emulated network connections in some VMs that have low network
latency and bandwidth requirements while reserving the VFs for VMs that require the low latency, high throughput that directly assigned VFs can provide. In this case an
emulated SW interface in a VM is unable to properly communicate with VFs on the same
PF because the emulated SW interface's MAC address isn't programmed into the HW filters
on the PF. If we could use this op to program the MAC address and VLAN filters of
the emulated SW interfaces into the PF HW a VF could then properly communicate across
the NIC's internal VEB to the emulated SW interfaces.
[...]
This would also be good for Solarflare's VF plugin architecture. The VF
driver works as a plugin for virtio or xen_netfront and can refuse
packets that need to be bridged to another (physically) local address.
The PF driver has to tell VFs what the local addresses are and currently
relies on some custom scripting to know about those extra addresses.
(No, none of that is upstream - I'm preparing for that now.)
Ben.
--
Ben Hutchings, Staff Engineer, Solarflare
Not speaking for my employer; that's the marketing department's job.
They asked us to note that Solarflare product names are trademarked.
From: Ben Hutchings <hidden> Date: 2011-11-17 23:43:41
On Tue, 2011-10-25 at 08:59 -0700, Rose, Gregory V wrote:
quoted
-----Original Message-----
From: Michael S. Tsirkin [mailto:mst@redhat.com]
Sent: Tuesday, October 25, 2011 8:46 AM
To: Roopa Prabhu
Cc: netdev@vger.kernel.org; sri@us.ibm.com; dragos.tatulea@gmail.com;
arnd@arndb.de; kvm@vger.kernel.org; davem@davemloft.net;
mchan@broadcom.com; dwang2@cisco.com; shemminger@vyatta.com;
eric.dumazet@gmail.com; kaber@trash.net; benve@cisco.com; Rose, Gregory V
Subject: Re: [net-next-2.6 PATCH 0/8 RFC v2] macvlan: MAC Address
filtering support for passthru mode
On Mon, Oct 24, 2011 at 11:15:05AM -0700, Roopa Prabhu wrote:
quoted
quoted
quoted
... And also I dont know of any hw
that provides an interface to set hw filters on a per queue basis.
VMDq hardware would support this, no?
Am not really sure. This patch uses netdev to pass filters to hw. And I
don't see any netdev infrastructure that would support per queue
filters.
Sure. I was only saying that as far as I understand,
VMDq hardware does support this functionality.
Right, Greg?
In the case of Intel HW yes. I'll refrain from speaking for other HW
vendors although I'm guessing it would be true in their cases also.
YMMV, caveat emptor, etc. etc.
Current Solarflare hardware supports:
- RX MAC address filters for queue selection (steering), which can be
combined with RSS (flow hashing)
- TX MAC address filters to prevent spoofing
Multiple filters can be associated with a single queue.
Ben.
--
Ben Hutchings, Staff Engineer, Solarflare
Not speaking for my employer; that's the marketing department's job.
They asked us to note that Solarflare product names are trademarked.