From: Jussi Maki <hidden> Date: 2021-06-09 13:56:07
This patchset introduces XDP support to the bonding driver.
Patch 1 contains the implementation, including support for
the recently introduced EXCLUDE_INGRESS. Patch 2 contains a
performance fix to the roundrobin mode which switches rr_tx_counter
to be per-cpu. Patch 3 contains the test suite for the implementation
using a pair of veth devices.
The vmtest.sh is modified to enable the bonding module and install
modules. The config change should probably be done in the libbpf
repository. Andrii: How would you like this done properly?
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
---
Jussi Maki (3):
net: bonding: Add XDP support to the bonding driver
net: bonding: Use per-cpu rr_tx_counter
selftests/bpf: Add tests for XDP bonding
drivers/net/bonding/bond_main.c | 459 +++++++++++++++---
include/linux/filter.h | 13 +-
include/linux/netdevice.h | 5 +
include/net/bonding.h | 3 +-
kernel/bpf/devmap.c | 34 +-
net/core/filter.c | 37 +-
.../selftests/bpf/prog_tests/xdp_bonding.c | 342 +++++++++++++
tools/testing/selftests/bpf/vmtest.sh | 30 +-
8 files changed, 843 insertions(+), 80 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/xdp_bonding.c
--
2.30.2
From: Jussi Maki <hidden> Date: 2021-06-09 13:56:01
The round-robin rr_tx_counter was shared across CPUs leading
to significant cache trashing at high packet rates. This patch
switches the round-robin mechanism to use a per-cpu counter to
decide the destination device.
On a 100Gbit 64 byte packet test this reduces the CPU load from
50% to 10% on the test system.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 18 +++++++++++++++---
include/net/bonding.h | 2 +-
2 files changed, 16 insertions(+), 4 deletions(-)
From: Jussi Maki <hidden> Date: 2021-06-09 13:56:01
XDP is implemented in the bonding driver by transparently delegating
the XDP program loading, removal and xmit operations to the bonding
slave devices. The overall goal of this work is that XDP programs
can be attached to a bond device *without* any further changes (or
awareness) necessary to the program itself, meaning the same XDP
program can be attached to a native device but also a bonding device.
Semantics of XDP_TX when attached to a bond are equivalent in such
setting to the case when a tc/BPF program would be attached to the
bond, meaning transmitting the packet out of the bond itself using one
of the bond's configured xmit methods to select a slave device (rather
than XDP_TX on the slave itself). Handling of XDP_TX to transmit
using the configured bonding mechanism is therefore implemented by
rewriting the BPF program return value in bpf_prog_run_xdp. To avoid
performance impact this check is guarded by a static key, which is
incremented when a XDP program is loaded onto a bond device. This
approach was chosen to avoid changes to drivers implementing XDP. If
the slave device does not match the receive device, then XDP_REDIRECT
is transparently used to perform the redirection in order to have
the network driver release the packet from its RX ring. The bonding
driver hashing functions have been refactored to allow reuse with
xdp_buff's to avoid code duplication.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 441 ++++++++++++++++++++++++++++----
include/linux/filter.h | 13 +-
include/linux/netdevice.h | 5 +
include/net/bonding.h | 1 +
kernel/bpf/devmap.c | 34 ++-
net/core/filter.c | 37 ++-
6 files changed, 467 insertions(+), 64 deletions(-)
@@ -317,6 +317,19 @@ bool bond_sk_check(struct bonding *bond)}}+staticboolbond_xdp_check(structbonding*bond)+{+switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_ACTIVEBACKUP:+caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+returntrue;+default:+returnfalse;+}+}+/*---------------------------------- VLAN -----------------------------------*//* In the following 2 functions, bond_vlan_rx_add_vid and bond_vlan_rx_kill_vid,
@@ -2001,6 +2014,28 @@ int bond_enslave(struct net_device *bond_dev, struct net_device *slave_dev,if(bond_mode_can_use_xmit_hash(bond))bond_update_slave_arr(bond,NULL);+if(bond->xdp_prog){+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=bond->xdp_prog,+.extack=extack,+};+if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+NL_SET_ERR_MSG(extack,"Slave does not support XDP");+slave_err(bond_dev,slave_dev,"Slave does not support XDP\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}+res=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(res<0){+/* ndo_bpf() sets extack error message */+slave_dbg(bond_dev,slave_dev,"Error %d calling ndo_bpf\n",res);+gotoerr_sysfs_del;+}+bpf_prog_inc(bond->xdp_prog);+}slave_info(bond_dev,slave_dev,"Enslaving as %s interface with %s link\n",bond_is_active_slave(new_slave)?"an active":"a backup",
@@ -2121,6 +2156,17 @@ static int __bond_release_one(struct net_device *bond_dev,/* recompute stats just before removing the slave */bond_get_stats(bond->dev,&bond->bond_stats);+if(bond->xdp_prog){+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=NULL,+.extack=NULL,+};+if(slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp))+slave_warn(bond_dev,slave_dev,"failed to unload XDP program\n");+}+bond_upper_dev_unlink(bond,slave);/* unregister rx_handler early so bond_handle_frame wouldn't be called*forthisslaveanymore.
@@ -3479,55 +3525,80 @@ static struct notifier_block bond_netdev_notifier = {/*---------------------------- Hashing Policies -----------------------------*/+/* Helper to access data in a packet, with or without a backing skb.+*Ifskbisgiventhedataislinearizedifnecessaryviapskb_may_pull.+*/+staticinlineconstvoid*bond_pull_data(structsk_buff*skb,+constvoid*data,inthlen,intn)+{+if(likely(n<=hlen))+returndata;+elseif(skb&&likely(pskb_may_pull(skb,n)))+returnskb->head;++returnNULL;+}+/* L2 hash helper */-staticinlineu32bond_eth_hash(structsk_buff*skb)+staticinlineu32bond_eth_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*ep,hdr_tmp;+structethhdr*ep;-ep=skb_header_pointer(skb,0,sizeof(hdr_tmp),&hdr_tmp);-if(ep)-returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;-return0;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;++ep=(structethhdr*)(data+mhoff);+returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;}-staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,-int*noff,int*proto,booll34)+staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,constvoid*data,+inthlen,intl2_proto,int*nhoff,int*ip_proto,booll34){conststructipv6hdr*iph6;conststructiphdr*iph;-if(skb->protocol==htons(ETH_P_IP)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph))))+if(l2_proto==htons(ETH_P_IP)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph));+if(!data)returnfalse;-iph=(conststructiphdr*)(skb->data+*noff);++iph=(conststructiphdr*)(data+*nhoff);iph_to_flow_copy_v4addrs(fk,iph);-*noff+=iph->ihl<<2;+*nhoff+=iph->ihl<<2;if(!ip_is_fragment(iph))-*proto=iph->protocol;-}elseif(skb->protocol==htons(ETH_P_IPV6)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph6))))+*ip_proto=iph->protocol;+}elseif(l2_proto==htons(ETH_P_IPV6)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph6));+if(!data)returnfalse;-iph6=(conststructipv6hdr*)(skb->data+*noff);++iph6=(conststructipv6hdr*)(data+*nhoff);iph_to_flow_copy_v6addrs(fk,iph6);-*noff+=sizeof(*iph6);-*proto=iph6->nexthdr;+*nhoff+=sizeof(*iph6);+*ip_proto=iph6->nexthdr;}else{returnfalse;}-if(l34&&*proto>=0)-fk->ports.ports=skb_flow_get_ports(skb,*noff,*proto);+if(l34&&*ip_proto>=0)+fk->ports.ports=__skb_flow_get_ports(skb,*nhoff,*ip_proto,data,hlen);returntrue;}-staticu32bond_vlan_srcmac_hash(structsk_buff*skb)+staticu32bond_vlan_srcmac_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*mac_hdr=(structethhdr*)skb_mac_header(skb);+structethhdr*mac_hdr;u32srcmac_vendor=0,srcmac_dev=0;u16vlan;inti;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;+mac_hdr=(structethhdr*)(data+mhoff);+for(i=0;i<3;i++)srcmac_vendor=(srcmac_vendor<<8)|mac_hdr->h_source[i];
@@ -3543,26 +3614,30 @@ static u32 bond_vlan_srcmac_hash(struct sk_buff *skb)}/* Extract the appropriate headers based on bond's xmit policy */-staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,+staticboolbond_flow_dissect(structbonding*bond,+structsk_buff*skb,+constvoid*data,+__be16l2_proto,+intnhoff,+inthlen,structflow_keys*fk){booll34=bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER34;-intnoff,proto=-1;+intip_proto=-1;switch(bond->params.xmit_policy){caseBOND_XMIT_POLICY_ENCAP23:caseBOND_XMIT_POLICY_ENCAP34:memset(fk,0,sizeof(*fk));return__skb_flow_dissect(NULL,skb,&flow_keys_bonding,-fk,NULL,0,0,0,0);+fk,data,l2_proto,nhoff,hlen,0);default:break;}fk->ports.ports=0;memset(&fk->icmp,0,sizeof(fk->icmp));-noff=skb_network_offset(skb);-if(!bond_flow_ip(skb,fk,&noff,&proto,l34))+if(!bond_flow_ip(skb,fk,data,hlen,l2_proto,&nhoff,&ip_proto,l34))returnfalse;/* ICMP error packets contains at least 8 bytes of the header
@@ -3601,33 +3674,30 @@ static u32 bond_ip_hash(u32 hash, struct flow_keys *flow)returnhash>>1;}-/**-*bond_xmit_hash-generateahashvaluebasedonthexmitpolicy-*@bond:bondingdevice-*@skb:buffertouseforheaders-*-*Thisfunctionwillextractthenecessaryheadersfromtheskbbufferanduse-*themtogenerateahashbasedonthexmit_policysetinthebondingdevice+/* Generate hash based on xmit policy. If @skb is given it is used to linearize+*thedataasrequired,butthisfunctioncanbeusedwithoutit.*/-u32bond_xmit_hash(structbonding*bond,structsk_buff*skb)+staticu32__bond_xmit_hash(structbonding*bond,+structsk_buff*skb,+constvoid*data,+__be16l2_proto,+intmhoff,+intnhoff,+inthlen){structflow_keysflow;u32hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP34&&-skb->l4_hash)-returnskb->hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_VLAN_SRCMAC)-returnbond_vlan_srcmac_hash(skb);+returnbond_vlan_srcmac_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER2||-!bond_flow_dissect(bond,skb,&flow))-returnbond_eth_hash(skb);+!bond_flow_dissect(bond,skb,data,l2_proto,nhoff,hlen,&flow))+returnbond_eth_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER23||bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP23){-hash=bond_eth_hash(skb);+hash=bond_eth_hash(skb,data,mhoff,hlen);}else{if(flow.icmp.id)memcpy(&hash,&flow.icmp,sizeof(hash));
@@ -4470,6 +4622,22 @@ static struct slave *bond_xmit_3ad_xor_slave_get(struct bonding *bond,returnslave;}+staticstructslave*bond_xdp_xmit_3ad_xor_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structbond_up_slave*slaves;+unsignedintcount;+u32hash;++hash=bond_xmit_hash_xdp(bond,xdp);+slaves=bond->usable_slaves;+count=slaves?READ_ONCE(slaves->count):0;+if(unlikely(!count))+returnNULL;++returnslaves->arr[hash%count];+}+/* Use this Xmit function for 3AD as well as XOR modes. The current*usableslavearrayisformedinthecontrolpath.Thexmitfunction*justcalculateshashandsendsthepacketout.
@@ -4754,6 +4922,164 @@ static netdev_tx_t bond_start_xmit(struct sk_buff *skb, struct net_device *dev)returnret;}+structnet_device*+bond_xdp_get_xmit_slave(structnet_device*bond_dev,structxdp_buff*xdp)+{+structbonding*bond=netdev_priv(bond_dev);+structslave*slave;++/* Caller needs to hold rcu_read_lock() */++switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+slave=bond_xdp_xmit_roundrobin_slave_get(bond,xdp);+break;++caseBOND_MODE_ACTIVEBACKUP:+slave=bond_xmit_activebackup_slave_get(bond);+break;++caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+slave=bond_xdp_xmit_3ad_xor_slave_get(bond,xdp);+break;++default:+/* Should never happen. Mode guarded by bond_xdp_check() */+netdev_err(bond_dev,"Unknown bonding mode %d for xdp xmit\n",BOND_MODE(bond));+WARN_ON_ONCE(1);+returnNULL;+}++if(slave)+returnslave->dev;++returnNULL;+}++staticintbond_xdp_xmit(structnet_device*bond_dev,+intn,structxdp_frame**frames,u32flags)+{+intnxmit,err=-ENXIO;++rcu_read_lock();++for(nxmit=0;nxmit<n;nxmit++){+structxdp_frame*frame=frames[nxmit];+structxdp_frame*frames1[]={frame};+structnet_device*slave_dev;+structxdp_buffxdp;++xdp_convert_frame_to_buff(frame,&xdp);++slave_dev=bond_xdp_get_xmit_slave(bond_dev,&xdp);+if(!slave_dev){+err=-ENXIO;+break;+}++err=slave_dev->netdev_ops->ndo_xdp_xmit(slave_dev,1,frames1,flags);+if(err<1)+break;+}++rcu_read_unlock();++/* If error happened on the first frame then we can pass the error up, otherwise+*reportthenumberofframesthatwerexmitted.+*/+if(err<0)+return(nxmit==0?err:nxmit);++returnnxmit;+}++staticintbond_xdp_set(structnet_device*dev,structbpf_prog*prog,+structnetlink_ext_ack*extack)+{+structbonding*bond=netdev_priv(dev);+structlist_head*iter;+structslave*slave,*rollback_slave;+structbpf_prog*old_prog;+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=prog,+.extack=extack,+};+interr;++ASSERT_RTNL();++if(!bond_xdp_check(bond))+return-EOPNOTSUPP;++old_prog=bond->xdp_prog;+bond->xdp_prog=prog;++bond_for_each_slave(bond,slave,iter){+structnet_device*slave_dev=slave->dev;++if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+NL_SET_ERR_MSG(extack,"Slave device does not support XDP");+slave_err(dev,slave_dev,"Slave does not support XDP\n");+err=-EOPNOTSUPP;+gotoerr;+}+err=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err<0){+/* ndo_bpf() sets extack error message */+slave_err(dev,slave_dev,"Error %d calling ndo_bpf\n",err);+gotoerr;+}+if(prog)+bpf_prog_inc(prog);+}++if(old_prog)+bpf_prog_put(old_prog);++if(prog)+static_branch_inc(&bpf_bond_redirect_enabled_key);+else+static_branch_dec(&bpf_bond_redirect_enabled_key);++return0;++err:+/* unwind the program changes */+bond->xdp_prog=old_prog;+xdp.prog=old_prog;+xdp.extack=NULL;/* do not overwrite original error */++bond_for_each_slave(bond,rollback_slave,iter){+structnet_device*slave_dev=rollback_slave->dev;+interr_unwind;++if(slave==rollback_slave)+break;++err_unwind=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err_unwind<0)+slave_err(dev,slave_dev,+"Error %d when unwinding XDP program change\n",err_unwind);+elseif(xdp.prog)+bpf_prog_inc(xdp.prog);+}+returnerr;+}++staticintbond_xdp(structnet_device*dev,structnetdev_bpf*xdp)+{+switch(xdp->command){+caseXDP_SETUP_PROG:+returnbond_xdp_set(dev,xdp->prog,xdp->extack);+default:+return-EINVAL;+}+}+staticu32bond_mode_bcast_speed(structslave*slave,u32speed){if(speed==0||speed==SPEED_UNKNOWN)
@@ -559,7 +568,7 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=READ_ONCE(dtab->netdev_map[i]);-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp,exclude_ifindex,exclude_ifindex_master))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -579,7 +588,9 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,head=dev_map_index_hash(dtab,i);hlist_for_each_entry_rcu(dst,head,index_hlist,lockdep_is_held(&dtab->index_lock)){-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp,+exclude_ifindex,+exclude_ifindex_master))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -646,16 +657,25 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,{structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);intexclude_ifindex=exclude_ingress?dev->ifindex:0;+intexclude_ifindex_master=0;structbpf_dtab_netdev*dst,*last_dst=NULL;structhlist_head*head;structhlist_node*next;unsignedinti;interr;+if(static_branch_unlikely(&bpf_bond_redirect_enabled_key)){+structnet_device*master=netdev_master_upper_dev_get_rcu(dev);++exclude_ifindex_master=(master&&exclude_ingress)?master->ifindex:0;+}+if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=READ_ONCE(dtab->netdev_map[i]);-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst||+dst->dev->ifindex==exclude_ifindex||+dst->dev->ifindex==exclude_ifindex_master)continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -674,7 +694,9 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,for(i=0;i<dtab->n_buckets;i++){head=dev_map_index_hash(dtab,i);hlist_for_each_entry_safe(dst,next,head,index_hlist){-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst||+dst->dev->ifindex==exclude_ifindex||+dst->dev->ifindex==exclude_ifindex_master)continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -2469,6 +2469,7 @@ int skb_do_redirect(struct sk_buff *skb)ri->flags=0;if(unlikely(!dev))gotoout_drop;+if(flags&BPF_F_PEER){conststructnet_device_ops*ops=dev->netdev_ops;
@@ -3947,6 +3948,40 @@ void bpf_clear_redirect_map(struct bpf_map *map)}}+DEFINE_STATIC_KEY_FALSE(bpf_bond_redirect_enabled_key);+EXPORT_SYMBOL_GPL(bpf_bond_redirect_enabled_key);+INDIRECT_CALLABLE_DECLARE(structnet_device*+bond_xdp_get_xmit_slave(structnet_device*bond_dev,structxdp_buff*xdp));++u32xdp_bond_redirect(structxdp_buff*xdp)+{+structnet_device*master,*slave;+structbpf_redirect_info*ri=this_cpu_ptr(&bpf_redirect_info);++master=netdev_master_upper_dev_get_rcu(xdp->rxq->dev);++#if IS_BUILTIN(CONFIG_BONDING)+slave=INDIRECT_CALL_1(master->netdev_ops->ndo_xdp_get_xmit_slave,+bond_xdp_get_xmit_slave,+master,xdp);+#else+slave=master->netdev_ops->ndo_xdp_get_xmit_slave(master,xdp);+#endif+if(slave&&slave!=xdp->rxq->dev){+/* The target device is different from the receiving device, so+*redirectittothenewdevice.+*UsingXDP_REDIRECTgetsthecorrectbehaviourfromXDPenabled+*driverstounmapthepacketfromtheirrxring.+*/+ri->tgt_index=slave->ifindex;+ri->map_id=INT_MAX;+ri->map_type=BPF_MAP_TYPE_UNSPEC;+returnXDP_REDIRECT;+}+returnXDP_TX;+}+EXPORT_SYMBOL_GPL(xdp_bond_redirect);+intxdp_do_redirect(structnet_device*dev,structxdp_buff*xdp,structbpf_prog*xdp_prog){
From: Jussi Maki <hidden> Date: 2021-06-09 13:56:20
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Signed-off-by: Jussi Maki <redacted>
---
.../selftests/bpf/prog_tests/xdp_bonding.c | 342 ++++++++++++++++++
tools/testing/selftests/bpf/vmtest.sh | 30 +-
2 files changed, 360 insertions(+), 12 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/xdp_bonding.c
@@ -0,0 +1,342 @@+// SPDX-License-Identifier: GPL-2.0++/**+*TestXDPbondingsupport+*+*Setsuptwobondedvethpairsbetweentwofreshnamespaces+*andverifiesthatXDP_TXprogramloadedonabonddevice+*arecorrectlyloadedontotheslavedevicesandXDP_TX'd+*packetsarebalancedusingbonding.+*/++#define _GNU_SOURCE+#include<sched.h>+#include<stdio.h>+#include<sys/types.h>+#include<sys/socket.h>+#include<fcntl.h>+#include<net/if.h>+#include<test_progs.h>+#include<network_helpers.h>+#include<linux/if_bonding.h>+#include<linux/limits.h>+#include<linux/if_ether.h>+#include<linux/udp.h>++#define BOND1_MAC {0x00, 0x11, 0x22, 0x33, 0x44, 0x55}+#define BOND1_MAC_STR "00:11:22:33:44:55"+#define BOND2_MAC {0x00, 0x22, 0x33, 0x44, 0x55, 0x66}+#define BOND2_MAC_STR "00:22:33:44:55:66"+#define NPACKETS 100++staticintroot_netns_fd=-1;++staticvoidrestore_root_netns(void)+{+ASSERT_OK(setns(root_netns_fd,CLONE_NEWNET),"restore_root_netns");+}++intsetns_by_name(char*name)+{+intnsfd,err;+charnspath[PATH_MAX];++snprintf(nspath,sizeof(nspath),"%s/%s","/var/run/netns",name);+nsfd=open(nspath,O_RDONLY|O_CLOEXEC);+if(nsfd<0)+return-1;++err=setns(nsfd,CLONE_NEWNET);+close(nsfd);+returnerr;+}++staticintget_rx_packets(constchar*iface)+{+FILE*f;+charline[512];+intiface_len=strlen(iface);++f=fopen("/proc/net/dev","r");+if(!f)+return-1;++while(fgets(line,sizeof(line),f)){+char*p=line;++while(*p==' ')+p++;/* skip whitespace */+if(!strncmp(p,iface,iface_len)){+p+=iface_len;+if(*p++!=':')+continue;+while(*p==' ')+p++;/* skip whitespace */+while(*p&&*p!=' ')+p++;/* skip rx bytes */+while(*p==' ')+p++;/* skip whitespace */+fclose(f);+returnatoi(p);+}+}+fclose(f);+return-1;+}++enum{+BOND_ONE_NO_ATTACH=0,+BOND_BOTH_AND_ATTACH,+};++staticintbonding_setup(intmode,intxmit_policy,intbond_both_attach)+{+#define SYS(fmt, ...) \+({\+charcmd[1024];\+snprintf(cmd,sizeof(cmd),fmt,##__VA_ARGS__);\+if(!ASSERT_OK(system(cmd),cmd))\+return-1;\+})++SYS("ip netns add ns_dst");+SYS("ip link add veth1_1 type veth peer name veth2_1 netns ns_dst");+SYS("ip link add veth1_2 type veth peer name veth2_2 netns ns_dst");++SYS("modprobe -r bonding &> /dev/null");+SYS("modprobe bonding mode=%d packets_per_slave=1 xmit_hash_policy=%d",mode,xmit_policy);++SYS("ip link add bond1 type bond");+SYS("ip link set bond1 address "BOND1_MAC_STR);+SYS("ip link set bond1 up");+SYS("ip -netns ns_dst link add bond2 type bond");+SYS("ip -netns ns_dst link set bond2 address "BOND2_MAC_STR);+SYS("ip -netns ns_dst link set bond2 up");++SYS("ip link set veth1_1 master bond1");+if(bond_both_attach==BOND_BOTH_AND_ATTACH){+SYS("ip link set veth1_2 master bond1");+}else{+SYS("ip link set veth1_2 up");+SYS("ip link set dev veth1_2 xdpdrv obj xdp_dummy.o sec xdp_dummy");+}++SYS("ip -netns ns_dst link set veth2_1 master bond2");++if(bond_both_attach==BOND_BOTH_AND_ATTACH)+SYS("ip -netns ns_dst link set veth2_2 master bond2");+else+SYS("ip -netns ns_dst link set veth2_2 up");++/* Load a dummy program on sending side as with veth peer needs to have a+*XDPprogramloadedaswell.+*/+SYS("ip link set dev bond1 xdpdrv obj xdp_dummy.o sec xdp_dummy");++if(bond_both_attach==BOND_BOTH_AND_ATTACH)+SYS("ip -netns ns_dst link set dev bond2 xdpdrv obj xdp_tx.o sec tx");++#undef SYS+return0;+}++staticvoidbonding_cleanup(void)+{+ASSERT_OK(system("ip link delete veth1_1"),"delete veth1_1");+ASSERT_OK(system("ip link delete veth1_2"),"delete veth1_2");+ASSERT_OK(system("ip netns delete ns_dst"),"delete ns_dst");+ASSERT_OK(system("modprobe -r bonding"),"unload bond");+}++staticintsend_udp_packets(intvary_dst_ip)+{+inti,s=-1;+intifindex;+uint8_tbuf[128]={};+structethhdreh={+.h_source=BOND1_MAC,+.h_dest=BOND2_MAC,+.h_proto=htons(ETH_P_IP),+};+structiphdr*iph=(structiphdr*)(buf+sizeof(eh));+structudphdr*uh=(structudphdr*)(buf+sizeof(eh)+sizeof(*iph));++s=socket(AF_PACKET,SOCK_RAW,IPPROTO_RAW);+if(!ASSERT_GE(s,0,"socket"))+gotoerr;++ifindex=if_nametoindex("bond1");+if(!ASSERT_GT(ifindex,0,"get bond1 ifindex"))+gotoerr;++memcpy(buf,&eh,sizeof(eh));+iph->ihl=5;+iph->version=4;+iph->tos=16;+iph->id=1;+iph->ttl=64;+iph->protocol=IPPROTO_UDP;+iph->saddr=1;+iph->daddr=2;+iph->tot_len=htons(sizeof(buf)-ETH_HLEN);+iph->check=0;++for(i=1;i<=NPACKETS;i++){+intn;+structsockaddr_llsaddr_ll={+.sll_ifindex=ifindex,+.sll_halen=ETH_ALEN,+.sll_addr=BOND2_MAC,+};++/* vary the UDP destination port for even distribution with roundrobin/xor modes */+uh->dest++;++if(vary_dst_ip)+iph->daddr++;++n=sendto(s,buf,sizeof(buf),0,(structsockaddr*)&saddr_ll,sizeof(saddr_ll));+if(!ASSERT_EQ(n,sizeof(buf),"sendto"))+gotoerr;+}++return0;++err:+if(s>=0)+close(s);+return-1;+}++voidtest_xdp_bonding_with_mode(char*name,intmode,intxmit_policy)+{+intbond1_rx;++if(!test__start_subtest(name))+return;++if(bonding_setup(mode,xmit_policy,BOND_BOTH_AND_ATTACH))+return;++if(send_udp_packets(xmit_policy!=BOND_XMIT_POLICY_LAYER34))+return;++bond1_rx=get_rx_packets("bond1");+ASSERT_TRUE(+bond1_rx>=NPACKETS,+"expected more received packets");++switch(mode){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_XOR:{+intveth1_rx=get_rx_packets("veth1_1");+intveth2_rx=get_rx_packets("veth1_2");+intdiff=abs(veth1_rx-veth2_rx);++ASSERT_GE(veth1_rx+veth2_rx,NPACKETS,"expected more packets");++switch(xmit_policy){+caseBOND_XMIT_POLICY_LAYER2:+ASSERT_GE(diff,NPACKETS/2,+"expected packets on only one of the interfaces");+break;+caseBOND_XMIT_POLICY_LAYER23:+caseBOND_XMIT_POLICY_LAYER34:+ASSERT_LT(diff,NPACKETS/2,+"expected even distribution of packets");+break;+default:+abort();+}+break;+}+default:+break;+}++bonding_cleanup();+}++voidtest_xdp_bonding_redirect_multi(void)+{+staticconstchar*constifaces[]={"bond2","veth2_1","veth2_2"};+intveth1_rx,veth2_rx;+interr;++if(!test__start_subtest("xdp_bonding_redirect_multi"))+return;++if(bonding_setup(BOND_MODE_ROUNDROBIN,BOND_XMIT_POLICY_LAYER23,BOND_ONE_NO_ATTACH))+gotoout;++err=system("ip -netns ns_dst link set dev bond2 xdpdrv "+"obj xdp_redirect_multi_kern.o sec xdp_redirect_map_multi");+if(!ASSERT_OK(err,"link set xdpdrv"))+gotoout;++/* populate the redirection devmap with the relevant interfaces */+if(!ASSERT_OK(setns_by_name("ns_dst"),"could not set netns to ns_dst"))+gotoout;++for(inti=0;i<ARRAY_SIZE(ifaces);i++){+charcmd[512];+intifindex=if_nametoindex(ifaces[i]);++if(!ASSERT_GT(ifindex,0,"could not get interface index"))+gotoout;++snprintf(cmd,sizeof(cmd),+"ip netns exec ns_dst bpftool map update name map_all key %d 0 0 0 value %d 0 0 0",+i,ifindex);++if(!ASSERT_OK(system(cmd),"bpftool map update"))+gotoout;+}+restore_root_netns();++send_udp_packets(BOND_MODE_ROUNDROBIN);++veth1_rx=get_rx_packets("veth1_1");+veth2_rx=get_rx_packets("veth1_2");++ASSERT_LT(veth1_rx,NPACKETS/2,"expected few packets on veth1");+ASSERT_GE(veth2_rx,NPACKETS,"expected more packets on veth2");+out:+restore_root_netns();+bonding_cleanup();+}++structbond_test_case{+char*name;+intmode;+intxmit_policy;+};++staticstructbond_test_casebond_test_cases[]={+{"xdp_bonding_roundrobin",BOND_MODE_ROUNDROBIN,BOND_XMIT_POLICY_LAYER23,},+{"xdp_bonding_activebackup",BOND_MODE_ACTIVEBACKUP,BOND_XMIT_POLICY_LAYER23},++{"xdp_bonding_xor_layer2",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER2,},+{"xdp_bonding_xor_layer23",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER23,},+{"xdp_bonding_xor_layer34",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER34,},+};++voidtest_xdp_bonding(void)+{+inti;++root_netns_fd=open("/proc/self/ns/net",O_RDONLY);+if(!ASSERT_GE(root_netns_fd,0,"open /proc/self/ns/net"))+return;++for(i=0;i<ARRAY_SIZE(bond_test_cases);i++){+structbond_test_case*test_case=&bond_test_cases[i];++test_xdp_bonding_with_mode(+test_case->name,+test_case->mode,+test_case->xmit_policy);+}++test_xdp_bonding_redirect_multi();+}
@@ -358,7 +364,7 @@ main()mkdir-p"${mount_dir}"update_kconfig"${kconfig_file}"-recompile_kernel"${kernel_checkout}""${make_command}"+recompile_kernel"${kernel_checkout}""${make_command}""${kconfig_file}"if[["${update_image}"=="no"&&!-f"${rootfs_img}"]];thenecho"rootfs image not found in ${rootfs_img}"
From: Maciej Fijalkowski <maciej.fijalkowski@intel.com> Date: 2021-06-09 22:19:49
On Wed, Jun 09, 2021 at 01:55:37PM +0000, Jussi Maki wrote:
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Cc: Magnus
Jussi,
AF_XDP selftests have very similar functionality just like you are trying
to introduce over here, e.g. we setup veth pair and generate traffic.
After a quick look seems that we could have a generic layer that would
be used by both AF_XDP and bonding selftests.
WDYT?
From: Maciej Fijalkowski <maciej.fijalkowski@intel.com> Date: 2021-06-09 22:42:34
On Wed, Jun 09, 2021 at 01:55:35PM +0000, Jussi Maki wrote:
XDP is implemented in the bonding driver by transparently delegating
the XDP program loading, removal and xmit operations to the bonding
slave devices. The overall goal of this work is that XDP programs
can be attached to a bond device *without* any further changes (or
awareness) necessary to the program itself, meaning the same XDP
program can be attached to a native device but also a bonding device.
Semantics of XDP_TX when attached to a bond are equivalent in such
setting to the case when a tc/BPF program would be attached to the
bond, meaning transmitting the packet out of the bond itself using one
of the bond's configured xmit methods to select a slave device (rather
than XDP_TX on the slave itself). Handling of XDP_TX to transmit
using the configured bonding mechanism is therefore implemented by
rewriting the BPF program return value in bpf_prog_run_xdp. To avoid
performance impact this check is guarded by a static key, which is
incremented when a XDP program is loaded onto a bond device. This
approach was chosen to avoid changes to drivers implementing XDP. If
the slave device does not match the receive device, then XDP_REDIRECT
is transparently used to perform the redirection in order to have
the network driver release the packet from its RX ring. The bonding
driver hashing functions have been refactored to allow reuse with
xdp_buff's to avoid code duplication.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 441 ++++++++++++++++++++++++++++----
include/linux/filter.h | 13 +-
include/linux/netdevice.h | 5 +
include/net/bonding.h | 1 +
kernel/bpf/devmap.c | 34 ++-
net/core/filter.c | 37 ++-
6 files changed, 467 insertions(+), 64 deletions(-)
Could this patch be broken down onto smaller chunks that would be easier
to review? Also please apply the Reverse Christmas Tree rule.
From: Jay Vosburgh <hidden> Date: 2021-06-09 23:29:29
Jussi Maki [off-list ref] wrote:
XDP is implemented in the bonding driver by transparently delegating
the XDP program loading, removal and xmit operations to the bonding
slave devices. The overall goal of this work is that XDP programs
can be attached to a bond device *without* any further changes (or
awareness) necessary to the program itself, meaning the same XDP
program can be attached to a native device but also a bonding device.
Semantics of XDP_TX when attached to a bond are equivalent in such
setting to the case when a tc/BPF program would be attached to the
bond, meaning transmitting the packet out of the bond itself using one
of the bond's configured xmit methods to select a slave device (rather
than XDP_TX on the slave itself). Handling of XDP_TX to transmit
using the configured bonding mechanism is therefore implemented by
rewriting the BPF program return value in bpf_prog_run_xdp. To avoid
performance impact this check is guarded by a static key, which is
incremented when a XDP program is loaded onto a bond device. This
approach was chosen to avoid changes to drivers implementing XDP. If
the slave device does not match the receive device, then XDP_REDIRECT
is transparently used to perform the redirection in order to have
the network driver release the packet from its RX ring. The bonding
driver hashing functions have been refactored to allow reuse with
xdp_buff's to avoid code duplication.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
The design adds logic around a bpf_bond_redirect_enabled_key
static key in the BPF core functions dev_map_enqueue_multi,
dev_map_redirect_multi and bpf_prog_run_xdp. Is this something that is
correctly implemented as a special case just for bonding (i.e., it will
never ever have to be extended), or is it possible that other
upper/lower type software devices will have similar XDP functionality
added in the future, e.g., bridge, VLAN, etc?
}
/* Extract the appropriate headers based on bond's xmit policy */
-static bool bond_flow_dissect(struct bonding *bond, struct sk_buff *skb,
+static bool bond_flow_dissect(struct bonding *bond,
+ struct sk_buff *skb,
+ const void *data,
+ __be16 l2_proto,
+ int nhoff,
+ int hlen,
struct flow_keys *fk)
Please compact the argument list down to fewer lines, in
conformance with usual coding practice in the kernel. The above style
of formatting occurs multiple times in this patch, both in function
declarations and function calls.
quoted hunk
{
bool l34 = bond->params.xmit_policy == BOND_XMIT_POLICY_LAYER34;
- int noff, proto = -1;
+ int ip_proto = -1;
switch (bond->params.xmit_policy) {
case BOND_XMIT_POLICY_ENCAP23:
case BOND_XMIT_POLICY_ENCAP34:
memset(fk, 0, sizeof(*fk));
return __skb_flow_dissect(NULL, skb, &flow_keys_bonding,
- fk, NULL, 0, 0, 0, 0);
+ fk, data, l2_proto, nhoff, hlen, 0);
default:
break;
}
fk->ports.ports = 0;
memset(&fk->icmp, 0, sizeof(fk->icmp));
- noff = skb_network_offset(skb);
- if (!bond_flow_ip(skb, fk, &noff, &proto, l34))
+ if (!bond_flow_ip(skb, fk, data, hlen, l2_proto, &nhoff, &ip_proto, l34))
return false;
/* ICMP error packets contains at least 8 bytes of the header
return hash >> 1;
}
-/**
- * bond_xmit_hash - generate a hash value based on the xmit policy
- * @bond: bonding device
- * @skb: buffer to use for headers
- *
- * This function will extract the necessary headers from the skb buffer and use
- * them to generate a hash based on the xmit_policy set in the bonding device
+/* Generate hash based on xmit policy. If @skb is given it is used to linearize
+ * the data as required, but this function can be used without it.
Please don't remove kernel-doc formatting; add your new
parameters to the documentation.
return slave;
}
+static struct slave *bond_xdp_xmit_3ad_xor_slave_get(struct bonding *bond,
+ struct xdp_buff *xdp)
+{
+ struct bond_up_slave *slaves;
+ unsigned int count;
+ u32 hash;
+
+ hash = bond_xmit_hash_xdp(bond, xdp);
+ slaves = bond->usable_slaves;
+ count = slaves ? READ_ONCE(slaves->count) : 0;
+ if (unlikely(!count))
+ return NULL;
+
+ return slaves->arr[hash % count];
+}
+
/* Use this Xmit function for 3AD as well as XOR modes. The current
* usable slave array is formed in the control path. The xmit function
* just calculates hash and sends the packet out.
if (map->map_type == BPF_MAP_TYPE_DEVMAP) {
for (i = 0; i < map->max_entries; i++) {
dst = READ_ONCE(dtab->netdev_map[i]);
- if (!is_valid_dst(dst, xdp, exclude_ifindex))
+ if (!is_valid_dst(dst, xdp, exclude_ifindex, exclude_ifindex_master))
continue;
/* we only need n-1 clones; last_dst enqueued below */
From: Jay Vosburgh <hidden> Date: 2021-06-10 00:04:36
Jussi Maki [off-list ref] wrote:
The round-robin rr_tx_counter was shared across CPUs leading
to significant cache trashing at high packet rates. This patch
"trashing" -> "thrashing" ?
quoted hunk
switches the round-robin mechanism to use a per-cpu counter to
decide the destination device.
On a 100Gbit 64 byte packet test this reduces the CPU load from
50% to 10% on the test system.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 18 +++++++++++++++---
include/net/bonding.h | 2 +-
2 files changed, 16 insertions(+), 4 deletions(-)
With the rr_tx_counter is per-cpu, each CPU is essentially doing
its own round-robin logic, independently of other CPUs, so the resulting
spread of transmitted packets may not be as evenly distributed (as
multiple CPUs could select the same interface to transmit on
approximately in lock-step). I'm not sure if this could cause actual
problems in practice, though, as particular flows shouldn't skip between
CPUs (and thus rr_tx_counters) very often, and round-robin already
shouldn't be the first choice if no packet reordering is a hard
requirement.
I think this patch could be submitted against net-next
independently of the rest of the series.
Acked-by: Jay Vosburgh <redacted>
-J
On Wed, Jun 9, 2021 at 6:55 AM Jussi Maki [off-list ref] wrote:
This patchset introduces XDP support to the bonding driver.
Patch 1 contains the implementation, including support for
the recently introduced EXCLUDE_INGRESS. Patch 2 contains a
performance fix to the roundrobin mode which switches rr_tx_counter
to be per-cpu. Patch 3 contains the test suite for the implementation
using a pair of veth devices.
The vmtest.sh is modified to enable the bonding module and install
modules. The config change should probably be done in the libbpf
repository. Andrii: How would you like this done properly?
I think vmtest.sh and CI setup doesn't support modules (not easily at
least). Can we just compile that driver in? Then you can submit a PR
against libbpf Github repo to adjust the config. We have also kernel
CI repo where we'll need to make this change.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
---
Jussi Maki (3):
net: bonding: Add XDP support to the bonding driver
net: bonding: Use per-cpu rr_tx_counter
selftests/bpf: Add tests for XDP bonding
drivers/net/bonding/bond_main.c | 459 +++++++++++++++---
include/linux/filter.h | 13 +-
include/linux/netdevice.h | 5 +
include/net/bonding.h | 3 +-
kernel/bpf/devmap.c | 34 +-
net/core/filter.c | 37 +-
.../selftests/bpf/prog_tests/xdp_bonding.c | 342 +++++++++++++
tools/testing/selftests/bpf/vmtest.sh | 30 +-
8 files changed, 843 insertions(+), 80 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/xdp_bonding.c
--
2.30.2
From: Jussi Maki <hidden> Date: 2021-06-14 07:55:54
On Thu, Jun 10, 2021 at 2:04 AM Jay Vosburgh [off-list ref] wrote:
Jussi Maki [off-list ref] wrote:
With the rr_tx_counter is per-cpu, each CPU is essentially doing
its own round-robin logic, independently of other CPUs, so the resulting
spread of transmitted packets may not be as evenly distributed (as
multiple CPUs could select the same interface to transmit on
approximately in lock-step). I'm not sure if this could cause actual
problems in practice, though, as particular flows shouldn't skip between
CPUs (and thus rr_tx_counters) very often, and round-robin already
shouldn't be the first choice if no packet reordering is a hard
requirement.
I think this patch could be submitted against net-next
independently of the rest of the series.
Yes this makes sense. I'll submit it separately against net-next today
and drop it off from this patchset.
From: Jussi Maki <hidden> Date: 2021-06-14 08:02:43
On Thu, Jun 10, 2021 at 1:29 AM Jay Vosburgh [off-list ref] wrote:
The design adds logic around a bpf_bond_redirect_enabled_key
static key in the BPF core functions dev_map_enqueue_multi,
dev_map_redirect_multi and bpf_prog_run_xdp. Is this something that is
correctly implemented as a special case just for bonding (i.e., it will
never ever have to be extended), or is it possible that other
upper/lower type software devices will have similar XDP functionality
added in the future, e.g., bridge, VLAN, etc?
Good point. For example the "team" driver would basically need pretty
much the same implementation. For that just using non-bond naming
would be enough. I don't think there's much of a cost for doing a more
generic mechanism, e.g. xdp "upper intercept" hook in netdev_ops, so
I'll try that out. At the very least I'll change the naming.
...
}
/* Extract the appropriate headers based on bond's xmit policy */
-static bool bond_flow_dissect(struct bonding *bond, struct sk_buff *skb,
+static bool bond_flow_dissect(struct bonding *bond,
+ struct sk_buff *skb,
+ const void *data,
+ __be16 l2_proto,
+ int nhoff,
+ int hlen,
struct flow_keys *fk)
Please compact the argument list down to fewer lines, in
conformance with usual coding practice in the kernel. The above style
of formatting occurs multiple times in this patch, both in function
declarations and function calls.
Thanks will do.
...
quoted
-/**
- * bond_xmit_hash - generate a hash value based on the xmit policy
- * @bond: bonding device
- * @skb: buffer to use for headers
- *
- * This function will extract the necessary headers from the skb buffer and use
- * them to generate a hash based on the xmit_policy set in the bonding device
+/* Generate hash based on xmit policy. If @skb is given it is used to linearize
+ * the data as required, but this function can be used without it.
Please don't remove kernel-doc formatting; add your new
parameters to the documentation.
The comment and the function declaration were untouched (see further
below in patch). I only introduced the common helper __bond_xmit_hash
used from bond_xmit_hash and bond_xmit_hash_xdp. Unfortunately the
generated diff was a bit confusing. I'll try and generate cleaner
diffs in the future.
}
+/**
+ * bond_xmit_hash_skb - generate a hash value based on the xmit policy
+ * @bond: bonding device
+ * @skb: buffer to use for headers
+ *
+ * This function will extract the necessary headers from the skb buffer and use
+ * them to generate a hash based on the xmit_policy set in the bonding device
+ */
+u32 bond_xmit_hash(struct bonding *bond, struct sk_buff *skb)
+{
+ if (bond->params.xmit_policy == BOND_XMIT_POLICY_ENCAP34 &&
+ skb->l4_hash)
+ return skb->hash;
+
+ return __bond_xmit_hash(bond, skb, skb->head, skb->protocol,
+ skb->mac_header,
+ skb->network_header,
+ skb_headlen(skb));
+}
From: Jussi Maki <hidden> Date: 2021-06-14 08:08:40
On Thu, Jun 10, 2021 at 12:19 AM Maciej Fijalkowski
[off-list ref] wrote:
On Wed, Jun 09, 2021 at 01:55:37PM +0000, Jussi Maki wrote:
quoted
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Cc: Magnus
Jussi,
AF_XDP selftests have very similar functionality just like you are trying
to introduce over here, e.g. we setup veth pair and generate traffic.
After a quick look seems that we could have a generic layer that would
be used by both AF_XDP and bonding selftests.
WDYT?
Sounds like a good idea to me to have more shared code in the
selftests and I don't see a reason not to use the AF_XDP datapath in
the bonding selftests. I'll look into it this week and get back to
you.
From: Magnus Karlsson <hidden> Date: 2021-06-14 08:49:22
On Mon, Jun 14, 2021 at 10:09 AM Jussi Maki [off-list ref] wrote:
On Thu, Jun 10, 2021 at 12:19 AM Maciej Fijalkowski
[off-list ref] wrote:
quoted
On Wed, Jun 09, 2021 at 01:55:37PM +0000, Jussi Maki wrote:
quoted
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Cc: Magnus
Jussi,
AF_XDP selftests have very similar functionality just like you are trying
to introduce over here, e.g. we setup veth pair and generate traffic.
After a quick look seems that we could have a generic layer that would
be used by both AF_XDP and bonding selftests.
WDYT?
Sounds like a good idea to me to have more shared code in the
selftests and I don't see a reason not to use the AF_XDP datapath in
the bonding selftests. I'll look into it this week and get back to
you.
Note, that I am currently rewriting a large part of the AF_XDP
selftests making it more amenable to adding various tests. A test is
in my patch set is described as a set of packets to send, a set of
packets that should be received in a certain order with specified
contents, and configuration/setup information for the sender and
receiver. The current code is riddled with test specific if-statements
that make it hard to extend and use generically. So please hold off
for a week or so and review my patch set when I send it to the list.
Better use of your time. Hopefully we can make it fit your bill too
with not too much work.
From: Jussi Maki <hidden> Date: 2021-06-14 12:20:43
On Mon, Jun 14, 2021 at 10:48 AM Magnus Karlsson
[off-list ref] wrote:
On Mon, Jun 14, 2021 at 10:09 AM Jussi Maki [off-list ref] wrote:
quoted
Sounds like a good idea to me to have more shared code in the
selftests and I don't see a reason not to use the AF_XDP datapath in
the bonding selftests. I'll look into it this week and get back to
you.
Note, that I am currently rewriting a large part of the AF_XDP
selftests making it more amenable to adding various tests. A test is
in my patch set is described as a set of packets to send, a set of
packets that should be received in a certain order with specified
contents, and configuration/setup information for the sender and
receiver. The current code is riddled with test specific if-statements
that make it hard to extend and use generically. So please hold off
for a week or so and review my patch set when I send it to the list.
Better use of your time. Hopefully we can make it fit your bill too
with not too much work.
Ok, thanks for the heads up! Looking forward to your patch set.
From: Jussi Maki <hidden> Date: 2021-06-14 12:26:55
On Thu, Jun 10, 2021 at 7:24 PM Andrii Nakryiko
[off-list ref] wrote:
On Wed, Jun 9, 2021 at 6:55 AM Jussi Maki [off-list ref] wrote:
quoted
This patchset introduces XDP support to the bonding driver.
Patch 1 contains the implementation, including support for
the recently introduced EXCLUDE_INGRESS. Patch 2 contains a
performance fix to the roundrobin mode which switches rr_tx_counter
to be per-cpu. Patch 3 contains the test suite for the implementation
using a pair of veth devices.
The vmtest.sh is modified to enable the bonding module and install
modules. The config change should probably be done in the libbpf
repository. Andrii: How would you like this done properly?
I think vmtest.sh and CI setup doesn't support modules (not easily at
least). Can we just compile that driver in? Then you can submit a PR
against libbpf Github repo to adjust the config. We have also kernel
CI repo where we'll need to make this change.
Unfortunately the mode and xmit_policy options of the bonding driver
are module params, so it'll need to be a module so the different modes
can be tested. I already modified vmtest.sh [1] to "make
module_install" into the rootfs and enable the bonding module via
scripts/config, but a cleaner approach would probably be to, as you
suggested, update latest.config in libbpf repo and probably get the
"modules_install" change into vmtest.sh separately (if you're happy
with this approach). What do you think?
[1] https://lore.kernel.org/netdev/20210609135537.1460244-1-joamaki@gmail.com/T/#maaf15ecd6b7c3af764558589118a3c6213e0af81
From: Jay Vosburgh <hidden> Date: 2021-06-14 15:37:43
Jussi Maki [off-list ref] wrote:
On Thu, Jun 10, 2021 at 7:24 PM Andrii Nakryiko
[off-list ref] wrote:
quoted
On Wed, Jun 9, 2021 at 6:55 AM Jussi Maki [off-list ref] wrote:
quoted
This patchset introduces XDP support to the bonding driver.
Patch 1 contains the implementation, including support for
the recently introduced EXCLUDE_INGRESS. Patch 2 contains a
performance fix to the roundrobin mode which switches rr_tx_counter
to be per-cpu. Patch 3 contains the test suite for the implementation
using a pair of veth devices.
The vmtest.sh is modified to enable the bonding module and install
modules. The config change should probably be done in the libbpf
repository. Andrii: How would you like this done properly?
I think vmtest.sh and CI setup doesn't support modules (not easily at
least). Can we just compile that driver in? Then you can submit a PR
against libbpf Github repo to adjust the config. We have also kernel
CI repo where we'll need to make this change.
Unfortunately the mode and xmit_policy options of the bonding driver
are module params, so it'll need to be a module so the different modes
can be tested. I already modified vmtest.sh [1] to "make
module_install" into the rootfs and enable the bonding module via
scripts/config, but a cleaner approach would probably be to, as you
suggested, update latest.config in libbpf repo and probably get the
"modules_install" change into vmtest.sh separately (if you're happy
with this approach). What do you think?
The bonding mode and xmit_hash_policy (and any other option) can
be changed via "ip link"; no module parameter needed, e.g.,
ip link set dev bond0 type bond xmit_hash_policy layer2
-J
On Mon, Jun 14, 2021 at 5:25 AM Jussi Maki [off-list ref] wrote:
On Thu, Jun 10, 2021 at 7:24 PM Andrii Nakryiko
[off-list ref] wrote:
quoted
On Wed, Jun 9, 2021 at 6:55 AM Jussi Maki [off-list ref] wrote:
quoted
This patchset introduces XDP support to the bonding driver.
Patch 1 contains the implementation, including support for
the recently introduced EXCLUDE_INGRESS. Patch 2 contains a
performance fix to the roundrobin mode which switches rr_tx_counter
to be per-cpu. Patch 3 contains the test suite for the implementation
using a pair of veth devices.
The vmtest.sh is modified to enable the bonding module and install
modules. The config change should probably be done in the libbpf
repository. Andrii: How would you like this done properly?
I think vmtest.sh and CI setup doesn't support modules (not easily at
least). Can we just compile that driver in? Then you can submit a PR
against libbpf Github repo to adjust the config. We have also kernel
CI repo where we'll need to make this change.
Unfortunately the mode and xmit_policy options of the bonding driver
are module params, so it'll need to be a module so the different modes
can be tested. I already modified vmtest.sh [1] to "make
module_install" into the rootfs and enable the bonding module via
scripts/config, but a cleaner approach would probably be to, as you
suggested, update latest.config in libbpf repo and probably get the
"modules_install" change into vmtest.sh separately (if you're happy
with this approach). What do you think?
If we can make modules work in vmtest.sh then it's great, regardless
if you need it still or not. It's not supported right now because no
one did work to support modules, not because we explicitly didn't want
modules in CI.
drivers/net/bonding/bond_main.c:4926:1: warning: no previous prototype for function 'bond_xdp_get_xmit_slave' [-Wmissing-prototypes]
bond_xdp_get_xmit_slave(struct net_device *bond_dev, struct xdp_buff *xdp)
^
drivers/net/bonding/bond_main.c:4925:1: note: declare 'static' if the function is not intended to be used outside of this translation unit
struct net_device *
^
static
1 warning generated.
--
quoted
drivers/net/bonding/bond_main.c:3720: warning: expecting prototype for bond_xmit_hash_skb(). Prototype was for bond_xmit_hash() instead
vim +/bond_xdp_get_xmit_slave +4926 drivers/net/bonding/bond_main.c
4924
4925 struct net_device *
From: kernel test robot <hidden> Date: 2021-06-22 07:25:08
Hi Jussi,
Thank you for the patch! Perhaps something to improve:
[auto build test WARNING on bpf-next/master]
url: https://github.com/0day-ci/linux/commits/Jussi-Maki/XDP-bonding-support/20210617-053146
base: https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git master
config: x86_64-randconfig-s031-20210622 (attached as .config)
compiler: gcc-9 (Debian 9.3.0-22) 9.3.0
reproduce:
# apt-get install sparse
# sparse version: v0.6.3-341-g8af24329-dirty
# https://github.com/0day-ci/linux/commit/61fabab38aec5b8e0cdc33867e35ea9740da84c8
git remote add linux-review https://github.com/0day-ci/linux
git fetch --no-tags linux-review Jussi-Maki/XDP-bonding-support/20210617-053146
git checkout 61fabab38aec5b8e0cdc33867e35ea9740da84c8
# save the attached .config to linux build tree
make W=1 C=1 CF='-fdiagnostic-prefix -D__CHECK_ENDIAN__' W=1 ARCH=x86_64
If you fix the issue, kindly add following tag as appropriate
Reported-by: kernel test robot <redacted>
sparse warnings: (new ones prefixed by >>)
drivers/net/bonding/bond_main.c:2660:26: sparse: sparse: restricted __be16 degrades to integer
drivers/net/bonding/bond_main.c:2666:20: sparse: sparse: restricted __be16 degrades to integer
drivers/net/bonding/bond_main.c:2713:40: sparse: sparse: incorrect type in assignment (different base types) @@ expected restricted __be16 [usertype] vlan_proto @@ got int @@
drivers/net/bonding/bond_main.c:2713:40: sparse: expected restricted __be16 [usertype] vlan_proto
drivers/net/bonding/bond_main.c:2713:40: sparse: got int
drivers/net/bonding/bond_main.c:3561:25: sparse: sparse: restricted __be16 degrades to integer
drivers/net/bonding/bond_main.c:3571:32: sparse: sparse: restricted __be16 degrades to integer
quoted
drivers/net/bonding/bond_main.c:3640:48: sparse: sparse: incorrect type in argument 5 (different base types) @@ expected int l2_proto @@ got restricted __be16 [usertype] l2_proto @@
drivers/net/bonding/bond_main.c:3640:48: sparse: expected int l2_proto
drivers/net/bonding/bond_main.c:3640:48: sparse: got restricted __be16 [usertype] l2_proto
drivers/net/bonding/bond_main.c:3661:58: sparse: sparse: incorrect type in argument 5 (different base types) @@ expected int l2_proto @@ got restricted __be16 [usertype] l2_proto @@
drivers/net/bonding/bond_main.c:3661:58: sparse: expected int l2_proto
drivers/net/bonding/bond_main.c:3661:58: sparse: got restricted __be16 [usertype] l2_proto
From: Jussi Maki <redacted>
This patchset introduces XDP support to the bonding driver.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter. Fix for this
has been already merged into net-next. The statistics were collected
using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Patch 1 prepares bond_xmit_hash for hashing xdp_buff's
Patch 2 adds hooks to implement redirection after bpf prog run
Patch 3 implements the hooks in the bonding driver.
Patch 4 modifies devmap to properly handle EXCLUDE_INGRESS with a slave device.
v1->v2:
- Split up into smaller easier to review patches and address cosmetic
review comments.
- Drop the INDIRECT_CALL optimization as it showed little improvement in tests.
- Drop the rr_tx_counter patch as that has already been merged into net-next.
- Separate the test suite into another patch set. This will follow later once a
patch set from Magnus Karlsson is merged and provides test utilities that can
be reused for XDP bonding tests. v2 contains no major functional changes and
was tested with the test suite included in v1.
(https://lore.kernel.org/bpf/202106221509.kwNvAAZg-lkp@intel.com/T/#m464146d47299125d5868a08affd6d6ce526dfad1)
---
Jussi Maki (4):
net: bonding: Refactor bond_xmit_hash for use with xdp_buff
net: core: Add support for XDP redirection to slave device
net: bonding: Add XDP support to the bonding driver
devmap: Exclude XDP broadcast to master device
drivers/net/bonding/bond_main.c | 431 +++++++++++++++++++++++++++-----
include/linux/filter.h | 13 +-
include/linux/netdevice.h | 5 +
include/net/bonding.h | 1 +
kernel/bpf/devmap.c | 34 ++-
net/core/filter.c | 25 ++
6 files changed, 445 insertions(+), 64 deletions(-)
--
2.27.0
From: Jussi Maki <redacted>
In preparation for adding XDP support to the bonding driver
refactor the packet hashing functions to be able to work with
any linear data buffer without an skb.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 147 +++++++++++++++++++-------------
1 file changed, 90 insertions(+), 57 deletions(-)
@@ -3479,55 +3479,80 @@ static struct notifier_block bond_netdev_notifier = {/*---------------------------- Hashing Policies -----------------------------*/+/* Helper to access data in a packet, with or without a backing skb.+*Ifskbisgiventhedataislinearizedifnecessaryviapskb_may_pull.+*/+staticinlineconstvoid*bond_pull_data(structsk_buff*skb,+constvoid*data,inthlen,intn)+{+if(likely(n<=hlen))+returndata;+elseif(skb&&likely(pskb_may_pull(skb,n)))+returnskb->head;++returnNULL;+}+/* L2 hash helper */-staticinlineu32bond_eth_hash(structsk_buff*skb)+staticinlineu32bond_eth_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*ep,hdr_tmp;+structethhdr*ep;-ep=skb_header_pointer(skb,0,sizeof(hdr_tmp),&hdr_tmp);-if(ep)-returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;-return0;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;++ep=(structethhdr*)(data+mhoff);+returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;}-staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,-int*noff,int*proto,booll34)+staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,constvoid*data,+inthlen,__be16l2_proto,int*nhoff,int*ip_proto,booll34){conststructipv6hdr*iph6;conststructiphdr*iph;-if(skb->protocol==htons(ETH_P_IP)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph))))+if(l2_proto==htons(ETH_P_IP)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph));+if(!data)returnfalse;-iph=(conststructiphdr*)(skb->data+*noff);++iph=(conststructiphdr*)(data+*nhoff);iph_to_flow_copy_v4addrs(fk,iph);-*noff+=iph->ihl<<2;+*nhoff+=iph->ihl<<2;if(!ip_is_fragment(iph))-*proto=iph->protocol;-}elseif(skb->protocol==htons(ETH_P_IPV6)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph6))))+*ip_proto=iph->protocol;+}elseif(l2_proto==htons(ETH_P_IPV6)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph6));+if(!data)returnfalse;-iph6=(conststructipv6hdr*)(skb->data+*noff);++iph6=(conststructipv6hdr*)(data+*nhoff);iph_to_flow_copy_v6addrs(fk,iph6);-*noff+=sizeof(*iph6);-*proto=iph6->nexthdr;+*nhoff+=sizeof(*iph6);+*ip_proto=iph6->nexthdr;}else{returnfalse;}-if(l34&&*proto>=0)-fk->ports.ports=skb_flow_get_ports(skb,*noff,*proto);+if(l34&&*ip_proto>=0)+fk->ports.ports=__skb_flow_get_ports(skb,*nhoff,*ip_proto,data,hlen);returntrue;}-staticu32bond_vlan_srcmac_hash(structsk_buff*skb)+staticu32bond_vlan_srcmac_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*mac_hdr=(structethhdr*)skb_mac_header(skb);+structethhdr*mac_hdr;u32srcmac_vendor=0,srcmac_dev=0;u16vlan;inti;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;+mac_hdr=(structethhdr*)(data+mhoff);+for(i=0;i<3;i++)srcmac_vendor=(srcmac_vendor<<8)|mac_hdr->h_source[i];
@@ -3543,26 +3568,25 @@ static u32 bond_vlan_srcmac_hash(struct sk_buff *skb)}/* Extract the appropriate headers based on bond's xmit policy */-staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,-structflow_keys*fk)+staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intnhoff,inthlen,structflow_keys*fk){booll34=bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER34;-intnoff,proto=-1;+intip_proto=-1;switch(bond->params.xmit_policy){caseBOND_XMIT_POLICY_ENCAP23:caseBOND_XMIT_POLICY_ENCAP34:memset(fk,0,sizeof(*fk));return__skb_flow_dissect(NULL,skb,&flow_keys_bonding,-fk,NULL,0,0,0,0);+fk,data,l2_proto,nhoff,hlen,0);default:break;}fk->ports.ports=0;memset(&fk->icmp,0,sizeof(fk->icmp));-noff=skb_network_offset(skb);-if(!bond_flow_ip(skb,fk,&noff,&proto,l34))+if(!bond_flow_ip(skb,fk,data,hlen,l2_proto,&nhoff,&ip_proto,l34))returnfalse;/* ICMP error packets contains at least 8 bytes of the header
@@ -3601,33 +3623,26 @@ static u32 bond_ip_hash(u32 hash, struct flow_keys *flow)returnhash>>1;}-/**-*bond_xmit_hash-generateahashvaluebasedonthexmitpolicy-*@bond:bondingdevice-*@skb:buffertouseforheaders-*-*Thisfunctionwillextractthenecessaryheadersfromtheskbbufferanduse-*themtogenerateahashbasedonthexmit_policysetinthebondingdevice+/* Generate hash based on xmit policy. If @skb is given it is used to linearize+*thedataasrequired,butthisfunctioncanbeusedwithoutitifthedatais+*knowntobelinear(e.g.withxdp_buff).*/-u32bond_xmit_hash(structbonding*bond,structsk_buff*skb)+staticu32__bond_xmit_hash(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intmhoff,intnhoff,inthlen){structflow_keysflow;u32hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP34&&-skb->l4_hash)-returnskb->hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_VLAN_SRCMAC)-returnbond_vlan_srcmac_hash(skb);+returnbond_vlan_srcmac_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER2||-!bond_flow_dissect(bond,skb,&flow))-returnbond_eth_hash(skb);+!bond_flow_dissect(bond,skb,data,l2_proto,nhoff,hlen,&flow))+returnbond_eth_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER23||bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP23){-hash=bond_eth_hash(skb);+hash=bond_eth_hash(skb,data,mhoff,hlen);}else{if(flow.icmp.id)memcpy(&hash,&flow.icmp,sizeof(hash));
From: Jussi Maki <redacted>
This adds the ndo_xdp_get_xmit_slave hook for transforming XDP_TX
into XDP_REDIRECT after BPF program run when the ingress device
is a bond slave.
Signed-off-by: Jussi Maki <redacted>
---
include/linux/filter.h | 13 ++++++++++++-
include/linux/netdevice.h | 5 +++++
net/core/filter.c | 25 +++++++++++++++++++++++++
3 files changed, 42 insertions(+), 1 deletion(-)
@@ -3947,6 +3947,31 @@ void bpf_clear_redirect_map(struct bpf_map *map)}}+DEFINE_STATIC_KEY_FALSE(bpf_master_redirect_enabled_key);+EXPORT_SYMBOL_GPL(bpf_master_redirect_enabled_key);++u32xdp_master_redirect(structxdp_buff*xdp)+{+structnet_device*master,*slave;+structbpf_redirect_info*ri=this_cpu_ptr(&bpf_redirect_info);++master=netdev_master_upper_dev_get_rcu(xdp->rxq->dev);+slave=master->netdev_ops->ndo_xdp_get_xmit_slave(master,xdp);+if(slave&&slave!=xdp->rxq->dev){+/* The target device is different from the receiving device, so+*redirectittothenewdevice.+*UsingXDP_REDIRECTgetsthecorrectbehaviourfromXDPenabled+*driverstounmapthepacketfromtheirrxring.+*/+ri->tgt_index=slave->ifindex;+ri->map_id=INT_MAX;+ri->map_type=BPF_MAP_TYPE_UNSPEC;+returnXDP_REDIRECT;+}+returnXDP_TX;+}+EXPORT_SYMBOL_GPL(xdp_master_redirect);+intxdp_do_redirect(structnet_device*dev,structxdp_buff*xdp,structbpf_prog*xdp_prog){
From: Jussi Maki <redacted>
XDP is implemented in the bonding driver by transparently delegating
the XDP program loading, removal and xmit operations to the bonding
slave devices. The overall goal of this work is that XDP programs
can be attached to a bond device *without* any further changes (or
awareness) necessary to the program itself, meaning the same XDP
program can be attached to a native device but also a bonding device.
Semantics of XDP_TX when attached to a bond are equivalent in such
setting to the case when a tc/BPF program would be attached to the
bond, meaning transmitting the packet out of the bond itself using one
of the bond's configured xmit methods to select a slave device (rather
than XDP_TX on the slave itself). Handling of XDP_TX to transmit
using the configured bonding mechanism is therefore implemented by
rewriting the BPF program return value in bpf_prog_run_xdp. To avoid
performance impact this check is guarded by a static key, which is
incremented when a XDP program is loaded onto a bond device. This
approach was chosen to avoid changes to drivers implementing XDP. If
the slave device does not match the receive device, then XDP_REDIRECT
is transparently used to perform the redirection in order to have
the network driver release the packet from its RX ring. The bonding
driver hashing functions have been refactored to allow reuse with
xdp_buff's to avoid code duplication.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter.
The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 284 ++++++++++++++++++++++++++++++++
include/net/bonding.h | 1 +
2 files changed, 285 insertions(+)
@@ -317,6 +317,19 @@ bool bond_sk_check(struct bonding *bond)}}+staticboolbond_xdp_check(structbonding*bond)+{+switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_ACTIVEBACKUP:+caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+returntrue;+default:+returnfalse;+}+}+/*---------------------------------- VLAN -----------------------------------*//* In the following 2 functions, bond_vlan_rx_add_vid and bond_vlan_rx_kill_vid,
@@ -2001,6 +2014,28 @@ int bond_enslave(struct net_device *bond_dev, struct net_device *slave_dev,if(bond_mode_can_use_xmit_hash(bond))bond_update_slave_arr(bond,NULL);+if(bond->xdp_prog){+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=bond->xdp_prog,+.extack=extack,+};+if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+NL_SET_ERR_MSG(extack,"Slave does not support XDP");+slave_err(bond_dev,slave_dev,"Slave does not support XDP\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}+res=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(res<0){+/* ndo_bpf() sets extack error message */+slave_dbg(bond_dev,slave_dev,"Error %d calling ndo_bpf\n",res);+gotoerr_sysfs_del;+}+bpf_prog_inc(bond->xdp_prog);+}slave_info(bond_dev,slave_dev,"Enslaving as %s interface with %s link\n",bond_is_active_slave(new_slave)?"an active":"a backup",
@@ -2121,6 +2156,17 @@ static int __bond_release_one(struct net_device *bond_dev,/* recompute stats just before removing the slave */bond_get_stats(bond->dev,&bond->bond_stats);+if(bond->xdp_prog){+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=NULL,+.extack=NULL,+};+if(slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp))+slave_warn(bond_dev,slave_dev,"failed to unload XDP program\n");+}+bond_upper_dev_unlink(bond,slave);/* unregister rx_handler early so bond_handle_frame wouldn't be called*forthisslaveanymore.
@@ -4288,6 +4354,47 @@ static struct slave *bond_xmit_roundrobin_slave_get(struct bonding *bond,returnNULL;}+staticstructslave*bond_xdp_xmit_roundrobin_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structslave*slave;+intslave_cnt;+u32slave_id;+conststructethhdr*eth;+void*data=xdp->data;++if(data+sizeof(structethhdr)>xdp->data_end)+gotonon_igmp;++eth=(structethhdr*)data;+data+=sizeof(structethhdr);++/* See comment on IGMP in bond_xmit_roundrobin_slave_get() */+if(eth->h_proto==htons(ETH_P_IP)){+conststructiphdr*iph;++if(data+sizeof(structiphdr)>xdp->data_end)+gotonon_igmp;++iph=(structiphdr*)data;++if(iph->protocol==IPPROTO_IGMP){+slave=rcu_dereference(bond->curr_active_slave);+if(slave)+returnslave;+returnbond_get_slave_by_id(bond,0);+}+}++non_igmp:+slave_cnt=READ_ONCE(bond->slave_cnt);+if(likely(slave_cnt)){+slave_id=bond_rr_gen_slave_id(bond)%slave_cnt;+returnbond_get_slave_by_id(bond,slave_id);+}+returnNULL;+}+staticnetdev_tx_tbond_xmit_roundrobin(structsk_buff*skb,structnet_device*bond_dev){
@@ -4503,6 +4610,22 @@ static struct slave *bond_xmit_3ad_xor_slave_get(struct bonding *bond,returnslave;}+staticstructslave*bond_xdp_xmit_3ad_xor_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structbond_up_slave*slaves;+unsignedintcount;+u32hash;++hash=bond_xmit_hash_xdp(bond,xdp);+slaves=bond->usable_slaves;+count=slaves?READ_ONCE(slaves->count):0;+if(unlikely(!count))+returnNULL;++returnslaves->arr[hash%count];+}+/* Use this Xmit function for 3AD as well as XOR modes. The current*usableslavearrayisformedinthecontrolpath.Thexmitfunction*justcalculateshashandsendsthepacketout.
@@ -4787,6 +4910,164 @@ static netdev_tx_t bond_start_xmit(struct sk_buff *skb, struct net_device *dev)returnret;}+staticstructnet_device*+bond_xdp_get_xmit_slave(structnet_device*bond_dev,structxdp_buff*xdp)+{+structbonding*bond=netdev_priv(bond_dev);+structslave*slave;++/* Caller needs to hold rcu_read_lock() */++switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+slave=bond_xdp_xmit_roundrobin_slave_get(bond,xdp);+break;++caseBOND_MODE_ACTIVEBACKUP:+slave=bond_xmit_activebackup_slave_get(bond);+break;++caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+slave=bond_xdp_xmit_3ad_xor_slave_get(bond,xdp);+break;++default:+/* Should never happen. Mode guarded by bond_xdp_check() */+netdev_err(bond_dev,"Unknown bonding mode %d for xdp xmit\n",BOND_MODE(bond));+WARN_ON_ONCE(1);+returnNULL;+}++if(slave)+returnslave->dev;++returnNULL;+}++staticintbond_xdp_xmit(structnet_device*bond_dev,+intn,structxdp_frame**frames,u32flags)+{+intnxmit,err=-ENXIO;++rcu_read_lock();++for(nxmit=0;nxmit<n;nxmit++){+structxdp_frame*frame=frames[nxmit];+structxdp_frame*frames1[]={frame};+structnet_device*slave_dev;+structxdp_buffxdp;++xdp_convert_frame_to_buff(frame,&xdp);++slave_dev=bond_xdp_get_xmit_slave(bond_dev,&xdp);+if(!slave_dev){+err=-ENXIO;+break;+}++err=slave_dev->netdev_ops->ndo_xdp_xmit(slave_dev,1,frames1,flags);+if(err<1)+break;+}++rcu_read_unlock();++/* If error happened on the first frame then we can pass the error up, otherwise+*reportthenumberofframesthatwerexmitted.+*/+if(err<0)+return(nxmit==0?err:nxmit);++returnnxmit;+}++staticintbond_xdp_set(structnet_device*dev,structbpf_prog*prog,+structnetlink_ext_ack*extack)+{+structbonding*bond=netdev_priv(dev);+structlist_head*iter;+structslave*slave,*rollback_slave;+structbpf_prog*old_prog;+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=prog,+.extack=extack,+};+interr;++ASSERT_RTNL();++if(!bond_xdp_check(bond))+return-EOPNOTSUPP;++old_prog=bond->xdp_prog;+bond->xdp_prog=prog;++bond_for_each_slave(bond,slave,iter){+structnet_device*slave_dev=slave->dev;++if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+NL_SET_ERR_MSG(extack,"Slave device does not support XDP");+slave_err(dev,slave_dev,"Slave does not support XDP\n");+err=-EOPNOTSUPP;+gotoerr;+}+err=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err<0){+/* ndo_bpf() sets extack error message */+slave_err(dev,slave_dev,"Error %d calling ndo_bpf\n",err);+gotoerr;+}+if(prog)+bpf_prog_inc(prog);+}++if(old_prog)+bpf_prog_put(old_prog);++if(prog)+static_branch_inc(&bpf_master_redirect_enabled_key);+else+static_branch_dec(&bpf_master_redirect_enabled_key);++return0;++err:+/* unwind the program changes */+bond->xdp_prog=old_prog;+xdp.prog=old_prog;+xdp.extack=NULL;/* do not overwrite original error */++bond_for_each_slave(bond,rollback_slave,iter){+structnet_device*slave_dev=rollback_slave->dev;+interr_unwind;++if(slave==rollback_slave)+break;++err_unwind=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err_unwind<0)+slave_err(dev,slave_dev,+"Error %d when unwinding XDP program change\n",err_unwind);+elseif(xdp.prog)+bpf_prog_inc(xdp.prog);+}+returnerr;+}++staticintbond_xdp(structnet_device*dev,structnetdev_bpf*xdp)+{+switch(xdp->command){+caseXDP_SETUP_PROG:+returnbond_xdp_set(dev,xdp->prog,xdp->extack);+default:+return-EINVAL;+}+}+staticu32bond_mode_bcast_speed(structslave*slave,u32speed){if(speed==0||speed==SPEED_UNKNOWN)
From: Jussi Maki <redacted>
If the ingress device is bond slave, do not broadcast back
through it or the bond master.
Signed-off-by: Jussi Maki <redacted>
---
kernel/bpf/devmap.c | 34 ++++++++++++++++++++++++++++------
1 file changed, 28 insertions(+), 6 deletions(-)
@@ -559,7 +568,7 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=READ_ONCE(dtab->netdev_map[i]);-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp,exclude_ifindex,exclude_ifindex_master))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -579,7 +588,9 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,head=dev_map_index_hash(dtab,i);hlist_for_each_entry_rcu(dst,head,index_hlist,lockdep_is_held(&dtab->index_lock)){-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp,+exclude_ifindex,+exclude_ifindex_master))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -646,16 +657,25 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,{structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);intexclude_ifindex=exclude_ingress?dev->ifindex:0;+intexclude_ifindex_master=0;structbpf_dtab_netdev*dst,*last_dst=NULL;structhlist_head*head;structhlist_node*next;unsignedinti;interr;+if(static_branch_unlikely(&bpf_master_redirect_enabled_key)){+structnet_device*master=netdev_master_upper_dev_get_rcu(dev);++exclude_ifindex_master=(master&&exclude_ingress)?master->ifindex:0;+}+if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=READ_ONCE(dtab->netdev_map[i]);-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst||+dst->dev->ifindex==exclude_ifindex||+dst->dev->ifindex==exclude_ifindex_master)continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -674,7 +694,9 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,for(i=0;i<dtab->n_buckets;i++){head=dev_map_index_hash(dtab,i);hlist_for_each_entry_safe(dst,next,head,index_hlist){-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst||+dst->dev->ifindex==exclude_ifindex||+dst->dev->ifindex==exclude_ifindex_master)continue;/* we only need n-1 clones; last_dst enqueued below */
From: Jay Vosburgh <hidden> Date: 2021-07-01 18:12:56
joamaki@gmail.com wrote:
quoted hunk
From: Jussi Maki <redacted>
If the ingress device is bond slave, do not broadcast back
through it or the bond master.
Signed-off-by: Jussi Maki <redacted>
---
kernel/bpf/devmap.c | 34 ++++++++++++++++++++++++++++------
1 file changed, 28 insertions(+), 6 deletions(-)
Will the above logic do what is intended if the device stacking
isn't a simple bond -> ethX arrangement? I.e., bond -> VLAN.?? -> ethX
or perhaps even bondA -> VLAN.?? -> bondB -> ethX ?
-J
quoted hunk
xdpf = xdp_convert_buff_to_frame(xdp);
if (unlikely(!xdpf))
return -EOVERFLOW;
if (map->map_type == BPF_MAP_TYPE_DEVMAP) {
for (i = 0; i < map->max_entries; i++) {
dst = READ_ONCE(dtab->netdev_map[i]);
- if (!is_valid_dst(dst, xdp, exclude_ifindex))
+ if (!is_valid_dst(dst, xdp, exclude_ifindex, exclude_ifindex_master))
continue;
/* we only need n-1 clones; last_dst enqueued below */
From: Jay Vosburgh <hidden> Date: 2021-07-01 18:20:14
joamaki@gmail.com wrote:
From: Jussi Maki <redacted>
This patchset introduces XDP support to the bonding driver.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter. Fix for this
has been already merged into net-next. The statistics were collected
using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
To be clear, the fact that the performance numbers for XDP_DROP
and XDP_TX are lower for "with patch, bond (1 dev)" than "without patch
(1 dev)" is expected, correct?
-J
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Patch 1 prepares bond_xmit_hash for hashing xdp_buff's
Patch 2 adds hooks to implement redirection after bpf prog run
Patch 3 implements the hooks in the bonding driver.
Patch 4 modifies devmap to properly handle EXCLUDE_INGRESS with a slave device.
v1->v2:
- Split up into smaller easier to review patches and address cosmetic
review comments.
- Drop the INDIRECT_CALL optimization as it showed little improvement in tests.
- Drop the rr_tx_counter patch as that has already been merged into net-next.
- Separate the test suite into another patch set. This will follow later once a
patch set from Magnus Karlsson is merged and provides test utilities that can
be reused for XDP bonding tests. v2 contains no major functional changes and
was tested with the test suite included in v1.
(https://lore.kernel.org/bpf/202106221509.kwNvAAZg-lkp@intel.com/T/#m464146d47299125d5868a08affd6d6ce526dfad1)
---
Jussi Maki (4):
net: bonding: Refactor bond_xmit_hash for use with xdp_buff
net: core: Add support for XDP redirection to slave device
net: bonding: Add XDP support to the bonding driver
devmap: Exclude XDP broadcast to master device
drivers/net/bonding/bond_main.c | 431 +++++++++++++++++++++++++++-----
include/linux/filter.h | 13 +-
include/linux/netdevice.h | 5 +
include/net/bonding.h | 1 +
kernel/bpf/devmap.c | 34 ++-
net/core/filter.c | 25 ++
6 files changed, 445 insertions(+), 64 deletions(-)
--
2.27.0
From: Jussi Maki <hidden> Date: 2021-07-05 10:33:04
On Thu, Jul 1, 2021 at 9:20 PM Jay Vosburgh [off-list ref] wrote:
joamaki@gmail.com wrote:
quoted
From: Jussi Maki <redacted>
This patchset introduces XDP support to the bonding driver.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter. Fix for this
has been already merged into net-next. The statistics were collected
using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
To be clear, the fact that the performance numbers for XDP_DROP
and XDP_TX are lower for "with patch, bond (1 dev)" than "without patch
(1 dev)" is expected, correct?
Yes that is correct. With the patch the ndo callback for choosing the
slave device is invoked which in this test (mode=xor) hashes L2&L3
headers (I seem to have failed to mention this in the original
message). In round-robin mode I recall it being about 16Mpps versus
the 18Mpps without the patch. I did also try "INDIRECT_CALL" to avoid
going via ndo_ops, but that had no discernible effect.
Will the above logic do what is intended if the device stacking
isn't a simple bond -> ethX arrangement? I.e., bond -> VLAN.?? -> ethX
or perhaps even bondA -> VLAN.?? -> bondB -> ethX ?
Good point. "bond -> VLAN -> eth" isn't an issue currently as vlan
devices do not support XDP. "bondA -> bondB -> ethX" however would be
supported, so I think it makes sense to change the code to collect all
upper devices and exclude them. I'll try to follow up with an updated
patch for this soon.
From: Jussi Maki <hidden> Date: 2021-07-07 13:13:56
This patchset introduces XDP support to the bonding driver.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter. Fix for this
has been already merged into net-next. The statistics were collected
using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Patch 1 prepares bond_xmit_hash for hashing xdp_buff's.
Patch 2 adds hooks to implement redirection after bpf prog run.
Patch 3 implements the hooks in the bonding driver.
Patch 4 modifies devmap to properly handle EXCLUDE_INGRESS with a slave device.
Patch 5 fixes an issue related to recent cleanup of rcu_read_lock in XDP context.
v2->v3:
- Address Jay's comment to properly exclude upper devices with EXCLUDE_INGRESS
when there are deeper nesting involved. Now all upper devices are excluded.
- Refuse to enslave devices that already have XDP programs loaded and refuse to
load XDP programs to slave devices. Earlier one could have a XDP program loaded
and after enslaving and loading another program onto the bond device the xdp_state
of the enslaved device would be pointing at an old program.
- Adapt netdev_lower_get_next_private_rcu so it can be called in the XDP context.
v1->v2:
- Split up into smaller easier to review patches and address cosmetic
review comments.
- Drop the INDIRECT_CALL optimization as it showed little improvement in tests.
- Drop the rr_tx_counter patch as that has already been merged into net-next.
- Separate the test suite into another patch set. This will follow later once a
patch set from Magnus Karlsson is merged and provides test utilities that can
be reused for XDP bonding tests. v2 contains no major functional changes and
was tested with the test suite included in v1.
(https://lore.kernel.org/bpf/202106221509.kwNvAAZg-lkp@intel.com/T/#m464146d47299125d5868a08affd6d6ce526dfad1)
---
Jussi Maki (5):
net: bonding: Refactor bond_xmit_hash for use with xdp_buff
net: core: Add support for XDP redirection to slave device
net: bonding: Add XDP support to the bonding driver
devmap: Exclude XDP broadcast to master device
net: core: Allow netdev_lower_get_next_private_rcu in bh context
drivers/net/bonding/bond_main.c | 450 ++++++++++++++++++++++++++++----
include/linux/filter.h | 13 +-
include/linux/netdevice.h | 6 +
include/net/bonding.h | 1 +
kernel/bpf/devmap.c | 67 ++++-
net/core/dev.c | 11 +-
net/core/filter.c | 25 ++
7 files changed, 504 insertions(+), 69 deletions(-)
--
2.27.0
From: Jussi Maki <hidden> Date: 2021-07-07 13:14:01
In preparation for adding XDP support to the bonding driver
refactor the packet hashing functions to be able to work with
any linear data buffer without an skb.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 147 +++++++++++++++++++-------------
1 file changed, 90 insertions(+), 57 deletions(-)
@@ -3487,55 +3487,80 @@ static struct notifier_block bond_netdev_notifier = {/*---------------------------- Hashing Policies -----------------------------*/+/* Helper to access data in a packet, with or without a backing skb.+*Ifskbisgiventhedataislinearizedifnecessaryviapskb_may_pull.+*/+staticinlineconstvoid*bond_pull_data(structsk_buff*skb,+constvoid*data,inthlen,intn)+{+if(likely(n<=hlen))+returndata;+elseif(skb&&likely(pskb_may_pull(skb,n)))+returnskb->head;++returnNULL;+}+/* L2 hash helper */-staticinlineu32bond_eth_hash(structsk_buff*skb)+staticinlineu32bond_eth_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*ep,hdr_tmp;+structethhdr*ep;-ep=skb_header_pointer(skb,0,sizeof(hdr_tmp),&hdr_tmp);-if(ep)-returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;-return0;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;++ep=(structethhdr*)(data+mhoff);+returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;}-staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,-int*noff,int*proto,booll34)+staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,constvoid*data,+inthlen,__be16l2_proto,int*nhoff,int*ip_proto,booll34){conststructipv6hdr*iph6;conststructiphdr*iph;-if(skb->protocol==htons(ETH_P_IP)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph))))+if(l2_proto==htons(ETH_P_IP)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph));+if(!data)returnfalse;-iph=(conststructiphdr*)(skb->data+*noff);++iph=(conststructiphdr*)(data+*nhoff);iph_to_flow_copy_v4addrs(fk,iph);-*noff+=iph->ihl<<2;+*nhoff+=iph->ihl<<2;if(!ip_is_fragment(iph))-*proto=iph->protocol;-}elseif(skb->protocol==htons(ETH_P_IPV6)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph6))))+*ip_proto=iph->protocol;+}elseif(l2_proto==htons(ETH_P_IPV6)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph6));+if(!data)returnfalse;-iph6=(conststructipv6hdr*)(skb->data+*noff);++iph6=(conststructipv6hdr*)(data+*nhoff);iph_to_flow_copy_v6addrs(fk,iph6);-*noff+=sizeof(*iph6);-*proto=iph6->nexthdr;+*nhoff+=sizeof(*iph6);+*ip_proto=iph6->nexthdr;}else{returnfalse;}-if(l34&&*proto>=0)-fk->ports.ports=skb_flow_get_ports(skb,*noff,*proto);+if(l34&&*ip_proto>=0)+fk->ports.ports=__skb_flow_get_ports(skb,*nhoff,*ip_proto,data,hlen);returntrue;}-staticu32bond_vlan_srcmac_hash(structsk_buff*skb)+staticu32bond_vlan_srcmac_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*mac_hdr=(structethhdr*)skb_mac_header(skb);+structethhdr*mac_hdr;u32srcmac_vendor=0,srcmac_dev=0;u16vlan;inti;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;+mac_hdr=(structethhdr*)(data+mhoff);+for(i=0;i<3;i++)srcmac_vendor=(srcmac_vendor<<8)|mac_hdr->h_source[i];
@@ -3551,26 +3576,25 @@ static u32 bond_vlan_srcmac_hash(struct sk_buff *skb)}/* Extract the appropriate headers based on bond's xmit policy */-staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,-structflow_keys*fk)+staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intnhoff,inthlen,structflow_keys*fk){booll34=bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER34;-intnoff,proto=-1;+intip_proto=-1;switch(bond->params.xmit_policy){caseBOND_XMIT_POLICY_ENCAP23:caseBOND_XMIT_POLICY_ENCAP34:memset(fk,0,sizeof(*fk));return__skb_flow_dissect(NULL,skb,&flow_keys_bonding,-fk,NULL,0,0,0,0);+fk,data,l2_proto,nhoff,hlen,0);default:break;}fk->ports.ports=0;memset(&fk->icmp,0,sizeof(fk->icmp));-noff=skb_network_offset(skb);-if(!bond_flow_ip(skb,fk,&noff,&proto,l34))+if(!bond_flow_ip(skb,fk,data,hlen,l2_proto,&nhoff,&ip_proto,l34))returnfalse;/* ICMP error packets contains at least 8 bytes of the header
@@ -3609,33 +3631,26 @@ static u32 bond_ip_hash(u32 hash, struct flow_keys *flow)returnhash>>1;}-/**-*bond_xmit_hash-generateahashvaluebasedonthexmitpolicy-*@bond:bondingdevice-*@skb:buffertouseforheaders-*-*Thisfunctionwillextractthenecessaryheadersfromtheskbbufferanduse-*themtogenerateahashbasedonthexmit_policysetinthebondingdevice+/* Generate hash based on xmit policy. If @skb is given it is used to linearize+*thedataasrequired,butthisfunctioncanbeusedwithoutitifthedatais+*knowntobelinear(e.g.withxdp_buff).*/-u32bond_xmit_hash(structbonding*bond,structsk_buff*skb)+staticu32__bond_xmit_hash(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intmhoff,intnhoff,inthlen){structflow_keysflow;u32hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP34&&-skb->l4_hash)-returnskb->hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_VLAN_SRCMAC)-returnbond_vlan_srcmac_hash(skb);+returnbond_vlan_srcmac_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER2||-!bond_flow_dissect(bond,skb,&flow))-returnbond_eth_hash(skb);+!bond_flow_dissect(bond,skb,data,l2_proto,nhoff,hlen,&flow))+returnbond_eth_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER23||bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP23){-hash=bond_eth_hash(skb);+hash=bond_eth_hash(skb,data,mhoff,hlen);}else{if(flow.icmp.id)memcpy(&hash,&flow.icmp,sizeof(hash));
From: Jussi Maki <hidden> Date: 2021-07-07 13:14:01
This adds the ndo_xdp_get_xmit_slave hook for transforming XDP_TX
into XDP_REDIRECT after BPF program run when the ingress device
is a bond slave.
The dev_xdp_prog_count is exposed so that slave devices can be checked
for loaded XDP programs in order to avoid the situation where both
bond master and slave have programs loaded according to xdp_state.
Signed-off-by: Jussi Maki <redacted>
---
include/linux/filter.h | 13 ++++++++++++-
include/linux/netdevice.h | 6 ++++++
net/core/dev.c | 9 ++++++++-
net/core/filter.c | 25 +++++++++++++++++++++++++
4 files changed, 51 insertions(+), 2 deletions(-)
@@ -9467,6 +9469,11 @@ static int dev_xdp_attach(struct net_device *dev, struct netlink_ext_ack *extackNL_SET_ERR_MSG(extack,"XDP_FLAGS_REPLACE is not specified");return-EINVAL;}+/* don't allow loading XDP programs to a bonded device */+if(netif_is_bond_slave(dev)){+NL_SET_ERR_MSG(extack,"XDP program can not be attached to a bond slave");+return-EINVAL;+}mode=dev_xdp_mode(dev,flags);/* can't replace attached link */
@@ -3950,6 +3950,31 @@ void bpf_clear_redirect_map(struct bpf_map *map)}}+DEFINE_STATIC_KEY_FALSE(bpf_master_redirect_enabled_key);+EXPORT_SYMBOL_GPL(bpf_master_redirect_enabled_key);++u32xdp_master_redirect(structxdp_buff*xdp)+{+structnet_device*master,*slave;+structbpf_redirect_info*ri=this_cpu_ptr(&bpf_redirect_info);++master=netdev_master_upper_dev_get_rcu(xdp->rxq->dev);+slave=master->netdev_ops->ndo_xdp_get_xmit_slave(master,xdp);+if(slave&&slave!=xdp->rxq->dev){+/* The target device is different from the receiving device, so+*redirectittothenewdevice.+*UsingXDP_REDIRECTgetsthecorrectbehaviourfromXDPenabled+*driverstounmapthepacketfromtheirrxring.+*/+ri->tgt_index=slave->ifindex;+ri->map_id=INT_MAX;+ri->map_type=BPF_MAP_TYPE_UNSPEC;+returnXDP_REDIRECT;+}+returnXDP_TX;+}+EXPORT_SYMBOL_GPL(xdp_master_redirect);+intxdp_do_redirect(structnet_device*dev,structxdp_buff*xdp,structbpf_prog*xdp_prog){
From: Jussi Maki <hidden> Date: 2021-07-07 13:14:01
XDP is implemented in the bonding driver by transparently delegating
the XDP program loading, removal and xmit operations to the bonding
slave devices. The overall goal of this work is that XDP programs
can be attached to a bond device *without* any further changes (or
awareness) necessary to the program itself, meaning the same XDP
program can be attached to a native device but also a bonding device.
Semantics of XDP_TX when attached to a bond are equivalent in such
setting to the case when a tc/BPF program would be attached to the
bond, meaning transmitting the packet out of the bond itself using one
of the bond's configured xmit methods to select a slave device (rather
than XDP_TX on the slave itself). Handling of XDP_TX to transmit
using the configured bonding mechanism is therefore implemented by
rewriting the BPF program return value in bpf_prog_run_xdp. To avoid
performance impact this check is guarded by a static key, which is
incremented when a XDP program is loaded onto a bond device. This
approach was chosen to avoid changes to drivers implementing XDP. If
the slave device does not match the receive device, then XDP_REDIRECT
is transparently used to perform the redirection in order to have
the network driver release the packet from its RX ring. The bonding
driver hashing functions have been refactored to allow reuse with
xdp_buff's to avoid code duplication.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 303 ++++++++++++++++++++++++++++++++
include/net/bonding.h | 1 +
2 files changed, 304 insertions(+)
@@ -317,6 +317,19 @@ bool bond_sk_check(struct bonding *bond)}}+staticboolbond_xdp_check(structbonding*bond)+{+switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_ACTIVEBACKUP:+caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+returntrue;+default:+returnfalse;+}+}+/*---------------------------------- VLAN -----------------------------------*//* In the following 2 functions, bond_vlan_rx_add_vid and bond_vlan_rx_kill_vid,
@@ -2010,6 +2023,39 @@ int bond_enslave(struct net_device *bond_dev, struct net_device *slave_dev,bond_update_slave_arr(bond,NULL);+if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+if(bond->xdp_prog){+NL_SET_ERR_MSG(extack,"Slave does not support XDP");+slave_err(bond_dev,slave_dev,"Slave does not support XDP\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}+}else{+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=bond->xdp_prog,+.extack=extack,+};++if(dev_xdp_prog_count(slave_dev)>0){+NL_SET_ERR_MSG(extack,"Slave has XDP program loaded, please unload before enslaving");+slave_err(bond_dev,slave_dev,"Slave has XDP program loaded, please unload before enslaving\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}++res=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(res<0){+/* ndo_bpf() sets extack error message */+slave_dbg(bond_dev,slave_dev,"Error %d calling ndo_bpf\n",res);+gotoerr_sysfs_del;+}+if(bond->xdp_prog)+bpf_prog_inc(bond->xdp_prog);+}+slave_info(bond_dev,slave_dev,"Enslaving as %s interface with %s link\n",bond_is_active_slave(new_slave)?"an active":"a backup",new_slave->link!=BOND_LINK_DOWN?"an up":"a down");
@@ -2129,6 +2175,17 @@ static int __bond_release_one(struct net_device *bond_dev,/* recompute stats just before removing the slave */bond_get_stats(bond->dev,&bond->bond_stats);+if(bond->xdp_prog){+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=NULL,+.extack=NULL,+};+if(slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp))+slave_warn(bond_dev,slave_dev,"failed to unload XDP program\n");+}+bond_upper_dev_unlink(bond,slave);/* unregister rx_handler early so bond_handle_frame wouldn't be called*forthisslaveanymore.
@@ -4296,6 +4373,47 @@ static struct slave *bond_xmit_roundrobin_slave_get(struct bonding *bond,returnNULL;}+staticstructslave*bond_xdp_xmit_roundrobin_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structslave*slave;+intslave_cnt;+u32slave_id;+conststructethhdr*eth;+void*data=xdp->data;++if(data+sizeof(structethhdr)>xdp->data_end)+gotonon_igmp;++eth=(structethhdr*)data;+data+=sizeof(structethhdr);++/* See comment on IGMP in bond_xmit_roundrobin_slave_get() */+if(eth->h_proto==htons(ETH_P_IP)){+conststructiphdr*iph;++if(data+sizeof(structiphdr)>xdp->data_end)+gotonon_igmp;++iph=(structiphdr*)data;++if(iph->protocol==IPPROTO_IGMP){+slave=rcu_dereference(bond->curr_active_slave);+if(slave)+returnslave;+returnbond_get_slave_by_id(bond,0);+}+}++non_igmp:+slave_cnt=READ_ONCE(bond->slave_cnt);+if(likely(slave_cnt)){+slave_id=bond_rr_gen_slave_id(bond)%slave_cnt;+returnbond_get_slave_by_id(bond,slave_id);+}+returnNULL;+}+staticnetdev_tx_tbond_xmit_roundrobin(structsk_buff*skb,structnet_device*bond_dev){
@@ -4511,6 +4629,22 @@ static struct slave *bond_xmit_3ad_xor_slave_get(struct bonding *bond,returnslave;}+staticstructslave*bond_xdp_xmit_3ad_xor_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structbond_up_slave*slaves;+unsignedintcount;+u32hash;++hash=bond_xmit_hash_xdp(bond,xdp);+slaves=bond->usable_slaves;+count=slaves?READ_ONCE(slaves->count):0;+if(unlikely(!count))+returnNULL;++returnslaves->arr[hash%count];+}+/* Use this Xmit function for 3AD as well as XOR modes. The current*usableslavearrayisformedinthecontrolpath.Thexmitfunction*justcalculateshashandsendsthepacketout.
@@ -4795,6 +4929,172 @@ static netdev_tx_t bond_start_xmit(struct sk_buff *skb, struct net_device *dev)returnret;}+staticstructnet_device*+bond_xdp_get_xmit_slave(structnet_device*bond_dev,structxdp_buff*xdp)+{+structbonding*bond=netdev_priv(bond_dev);+structslave*slave;++/* Caller needs to hold rcu_read_lock() */++switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+slave=bond_xdp_xmit_roundrobin_slave_get(bond,xdp);+break;++caseBOND_MODE_ACTIVEBACKUP:+slave=bond_xmit_activebackup_slave_get(bond);+break;++caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+slave=bond_xdp_xmit_3ad_xor_slave_get(bond,xdp);+break;++default:+/* Should never happen. Mode guarded by bond_xdp_check() */+netdev_err(bond_dev,"Unknown bonding mode %d for xdp xmit\n",BOND_MODE(bond));+WARN_ON_ONCE(1);+returnNULL;+}++if(slave)+returnslave->dev;++returnNULL;+}++staticintbond_xdp_xmit(structnet_device*bond_dev,+intn,structxdp_frame**frames,u32flags)+{+intnxmit,err=-ENXIO;++rcu_read_lock();++for(nxmit=0;nxmit<n;nxmit++){+structxdp_frame*frame=frames[nxmit];+structxdp_frame*frames1[]={frame};+structnet_device*slave_dev;+structxdp_buffxdp;++xdp_convert_frame_to_buff(frame,&xdp);++slave_dev=bond_xdp_get_xmit_slave(bond_dev,&xdp);+if(!slave_dev){+err=-ENXIO;+break;+}++err=slave_dev->netdev_ops->ndo_xdp_xmit(slave_dev,1,frames1,flags);+if(err<1)+break;+}++rcu_read_unlock();++/* If error happened on the first frame then we can pass the error up, otherwise+*reportthenumberofframesthatwerexmitted.+*/+if(err<0)+return(nxmit==0?err:nxmit);++returnnxmit;+}++staticintbond_xdp_set(structnet_device*dev,structbpf_prog*prog,+structnetlink_ext_ack*extack)+{+structbonding*bond=netdev_priv(dev);+structlist_head*iter;+structslave*slave,*rollback_slave;+structbpf_prog*old_prog;+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=prog,+.extack=extack,+};+interr;++ASSERT_RTNL();++if(!bond_xdp_check(bond))+return-EOPNOTSUPP;++old_prog=bond->xdp_prog;+bond->xdp_prog=prog;++bond_for_each_slave(bond,slave,iter){+structnet_device*slave_dev=slave->dev;++if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+NL_SET_ERR_MSG(extack,"Slave device does not support XDP");+slave_err(dev,slave_dev,"Slave does not support XDP\n");+err=-EOPNOTSUPP;+gotoerr;+}++if(dev_xdp_prog_count(slave_dev)>0){+NL_SET_ERR_MSG(extack,"Slave has XDP program loaded, please unload before enslaving");+slave_err(dev,slave_dev,"Slave has XDP program loaded, please unload before enslaving\n");+err=-EOPNOTSUPP;+gotoerr;+}++err=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err<0){+/* ndo_bpf() sets extack error message */+slave_err(dev,slave_dev,"Error %d calling ndo_bpf\n",err);+gotoerr;+}+if(prog)+bpf_prog_inc(prog);+}++if(old_prog)+bpf_prog_put(old_prog);++if(prog)+static_branch_inc(&bpf_master_redirect_enabled_key);+else+static_branch_dec(&bpf_master_redirect_enabled_key);++return0;++err:+/* unwind the program changes */+bond->xdp_prog=old_prog;+xdp.prog=old_prog;+xdp.extack=NULL;/* do not overwrite original error */++bond_for_each_slave(bond,rollback_slave,iter){+structnet_device*slave_dev=rollback_slave->dev;+interr_unwind;++if(slave==rollback_slave)+break;++err_unwind=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err_unwind<0)+slave_err(dev,slave_dev,+"Error %d when unwinding XDP program change\n",err_unwind);+elseif(xdp.prog)+bpf_prog_inc(xdp.prog);+}+returnerr;+}++staticintbond_xdp(structnet_device*dev,structnetdev_bpf*xdp)+{+switch(xdp->command){+caseXDP_SETUP_PROG:+returnbond_xdp_set(dev,xdp->prog,xdp->extack);+default:+return-EINVAL;+}+}+staticu32bond_mode_bcast_speed(structslave*slave,u32speed){if(speed==0||speed==SPEED_UNKNOWN)
From: Jussi Maki <hidden> Date: 2021-07-07 13:14:02
If the ingress device is bond slave, do not broadcast back
through it or the bond master.
Signed-off-by: Jussi Maki <redacted>
---
kernel/bpf/devmap.c | 67 +++++++++++++++++++++++++++++++++++++++------
1 file changed, 58 insertions(+), 9 deletions(-)
@@ -541,17 +540,48 @@ static int dev_map_enqueue_clone(struct bpf_dtab_netdev *obj,return0;}+staticinlineboolis_ifindex_excluded(int*excluded,intnum_excluded,intifindex)+{+while(num_excluded--){+if(ifindex==excluded[num_excluded])+returntrue;+}+returnfalse;+}++/* Get ifindex of each upper device. 'indexes' must be able to hold at+*leastMAX_NEST_DEVelements.+*Returnsthenumberofifindexesadded.+*/+staticintget_upper_ifindexes(structnet_device*dev,int*indexes)+{+structnet_device*upper;+structlist_head*iter;+intn=0;++netdev_for_each_upper_dev_rcu(dev,upper,iter){+indexes[n++]=upper->ifindex;+}+returnn;+}+intdev_map_enqueue_multi(structxdp_buff*xdp,structnet_device*dev_rx,structbpf_map*map,boolexclude_ingress){structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);-intexclude_ifindex=exclude_ingress?dev_rx->ifindex:0;structbpf_dtab_netdev*dst,*last_dst=NULL;+intexcluded_devices[1+MAX_NEST_DEV];structhlist_head*head;structxdp_frame*xdpf;+intnum_excluded=0;unsignedinti;interr;+if(exclude_ingress){+num_excluded=get_upper_ifindexes(dev_rx,excluded_devices);+excluded_devices[num_excluded++]=dev_rx->ifindex;+}+xdpf=xdp_convert_buff_to_frame(xdp);if(unlikely(!xdpf))return-EOVERFLOW;
@@ -559,7 +589,10 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=READ_ONCE(dtab->netdev_map[i]);-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp))+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -579,7 +612,10 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,head=dev_map_index_hash(dtab,i);hlist_for_each_entry_rcu(dst,head,index_hlist,lockdep_is_held(&dtab->index_lock)){-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp))+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -645,17 +681,26 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,boolexclude_ingress){structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);-intexclude_ifindex=exclude_ingress?dev->ifindex:0;structbpf_dtab_netdev*dst,*last_dst=NULL;+intexcluded_devices[1+MAX_NEST_DEV];structhlist_head*head;structhlist_node*next;+intnum_excluded=0;unsignedinti;interr;+if(exclude_ingress){+num_excluded=get_upper_ifindexes(dev,excluded_devices);+excluded_devices[num_excluded++]=dev->ifindex;+}+if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=READ_ONCE(dtab->netdev_map[i]);-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst)+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -669,12 +714,16 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,returnerr;last_dst=dst;+}}else{/* BPF_MAP_TYPE_DEVMAP_HASH */for(i=0;i<dtab->n_buckets;i++){head=dev_map_index_hash(dtab,i);hlist_for_each_entry_safe(dst,next,head,index_hlist){-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst)+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
This patchset introduces XDP support to the bonding driver.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter. Fix for this
has been already merged into net-next. The statistics were collected
using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Patch 1 prepares bond_xmit_hash for hashing xdp_buff's.
Patch 2 adds hooks to implement redirection after bpf prog run.
Patch 3 implements the hooks in the bonding driver.
Patch 4 modifies devmap to properly handle EXCLUDE_INGRESS with a slave device.
Patch 5 fixes an issue related to recent cleanup of rcu_read_lock in XDP context.
Patch 6 adds tests
v3->v4:
- Add back the test suite, while removing the vmtest.sh modifications to kernel
config new that CONFIG_BONDING=y is set. Discussed with Magnus Karlsson that
it makes sense right now to not reuse the code from xdpceiver.c for testing
XDP bonding.
v2->v3:
- Address Jay's comment to properly exclude upper devices with EXCLUDE_INGRESS
when there are deeper nesting involved. Now all upper devices are excluded.
- Refuse to enslave devices that already have XDP programs loaded and refuse to
load XDP programs to slave devices. Earlier one could have a XDP program loaded
and after enslaving and loading another program onto the bond device the xdp_state
of the enslaved device would be pointing at an old program.
- Adapt netdev_lower_get_next_private_rcu so it can be called in the XDP context.
v1->v2:
- Split up into smaller easier to review patches and address cosmetic
review comments.
- Drop the INDIRECT_CALL optimization as it showed little improvement in tests.
- Drop the rr_tx_counter patch as that has already been merged into net-next.
- Separate the test suite into another patch set. This will follow later once a
patch set from Magnus Karlsson is merged and provides test utilities that can
be reused for XDP bonding tests. v2 contains no major functional changes and
was tested with the test suite included in v1.
(https://lore.kernel.org/bpf/202106221509.kwNvAAZg-lkp@intel.com/T/#m464146d47299125d5868a08affd6d6ce526dfad1)
---
From: Jussi Maki <redacted>
In preparation for adding XDP support to the bonding driver
refactor the packet hashing functions to be able to work with
any linear data buffer without an skb.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 147 +++++++++++++++++++-------------
1 file changed, 90 insertions(+), 57 deletions(-)
@@ -3611,55 +3611,80 @@ static struct notifier_block bond_netdev_notifier = {/*---------------------------- Hashing Policies -----------------------------*/+/* Helper to access data in a packet, with or without a backing skb.+*Ifskbisgiventhedataislinearizedifnecessaryviapskb_may_pull.+*/+staticinlineconstvoid*bond_pull_data(structsk_buff*skb,+constvoid*data,inthlen,intn)+{+if(likely(n<=hlen))+returndata;+elseif(skb&&likely(pskb_may_pull(skb,n)))+returnskb->head;++returnNULL;+}+/* L2 hash helper */-staticinlineu32bond_eth_hash(structsk_buff*skb)+staticinlineu32bond_eth_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*ep,hdr_tmp;+structethhdr*ep;-ep=skb_header_pointer(skb,0,sizeof(hdr_tmp),&hdr_tmp);-if(ep)-returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;-return0;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;++ep=(structethhdr*)(data+mhoff);+returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;}-staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,-int*noff,int*proto,booll34)+staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,constvoid*data,+inthlen,__be16l2_proto,int*nhoff,int*ip_proto,booll34){conststructipv6hdr*iph6;conststructiphdr*iph;-if(skb->protocol==htons(ETH_P_IP)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph))))+if(l2_proto==htons(ETH_P_IP)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph));+if(!data)returnfalse;-iph=(conststructiphdr*)(skb->data+*noff);++iph=(conststructiphdr*)(data+*nhoff);iph_to_flow_copy_v4addrs(fk,iph);-*noff+=iph->ihl<<2;+*nhoff+=iph->ihl<<2;if(!ip_is_fragment(iph))-*proto=iph->protocol;-}elseif(skb->protocol==htons(ETH_P_IPV6)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph6))))+*ip_proto=iph->protocol;+}elseif(l2_proto==htons(ETH_P_IPV6)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph6));+if(!data)returnfalse;-iph6=(conststructipv6hdr*)(skb->data+*noff);++iph6=(conststructipv6hdr*)(data+*nhoff);iph_to_flow_copy_v6addrs(fk,iph6);-*noff+=sizeof(*iph6);-*proto=iph6->nexthdr;+*nhoff+=sizeof(*iph6);+*ip_proto=iph6->nexthdr;}else{returnfalse;}-if(l34&&*proto>=0)-fk->ports.ports=skb_flow_get_ports(skb,*noff,*proto);+if(l34&&*ip_proto>=0)+fk->ports.ports=__skb_flow_get_ports(skb,*nhoff,*ip_proto,data,hlen);returntrue;}-staticu32bond_vlan_srcmac_hash(structsk_buff*skb)+staticu32bond_vlan_srcmac_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*mac_hdr=(structethhdr*)skb_mac_header(skb);+structethhdr*mac_hdr;u32srcmac_vendor=0,srcmac_dev=0;u16vlan;inti;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;+mac_hdr=(structethhdr*)(data+mhoff);+for(i=0;i<3;i++)srcmac_vendor=(srcmac_vendor<<8)|mac_hdr->h_source[i];
@@ -3675,26 +3700,25 @@ static u32 bond_vlan_srcmac_hash(struct sk_buff *skb)}/* Extract the appropriate headers based on bond's xmit policy */-staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,-structflow_keys*fk)+staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intnhoff,inthlen,structflow_keys*fk){booll34=bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER34;-intnoff,proto=-1;+intip_proto=-1;switch(bond->params.xmit_policy){caseBOND_XMIT_POLICY_ENCAP23:caseBOND_XMIT_POLICY_ENCAP34:memset(fk,0,sizeof(*fk));return__skb_flow_dissect(NULL,skb,&flow_keys_bonding,-fk,NULL,0,0,0,0);+fk,data,l2_proto,nhoff,hlen,0);default:break;}fk->ports.ports=0;memset(&fk->icmp,0,sizeof(fk->icmp));-noff=skb_network_offset(skb);-if(!bond_flow_ip(skb,fk,&noff,&proto,l34))+if(!bond_flow_ip(skb,fk,data,hlen,l2_proto,&nhoff,&ip_proto,l34))returnfalse;/* ICMP error packets contains at least 8 bytes of the header
@@ -3733,33 +3755,26 @@ static u32 bond_ip_hash(u32 hash, struct flow_keys *flow)returnhash>>1;}-/**-*bond_xmit_hash-generateahashvaluebasedonthexmitpolicy-*@bond:bondingdevice-*@skb:buffertouseforheaders-*-*Thisfunctionwillextractthenecessaryheadersfromtheskbbufferanduse-*themtogenerateahashbasedonthexmit_policysetinthebondingdevice+/* Generate hash based on xmit policy. If @skb is given it is used to linearize+*thedataasrequired,butthisfunctioncanbeusedwithoutitifthedatais+*knowntobelinear(e.g.withxdp_buff).*/-u32bond_xmit_hash(structbonding*bond,structsk_buff*skb)+staticu32__bond_xmit_hash(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intmhoff,intnhoff,inthlen){structflow_keysflow;u32hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP34&&-skb->l4_hash)-returnskb->hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_VLAN_SRCMAC)-returnbond_vlan_srcmac_hash(skb);+returnbond_vlan_srcmac_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER2||-!bond_flow_dissect(bond,skb,&flow))-returnbond_eth_hash(skb);+!bond_flow_dissect(bond,skb,data,l2_proto,nhoff,hlen,&flow))+returnbond_eth_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER23||bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP23){-hash=bond_eth_hash(skb);+hash=bond_eth_hash(skb,data,mhoff,hlen);}else{if(flow.icmp.id)memcpy(&hash,&flow.icmp,sizeof(hash));
From: Jussi Maki <redacted>
This adds the ndo_xdp_get_xmit_slave hook for transforming XDP_TX
into XDP_REDIRECT after BPF program run when the ingress device
is a bond slave.
The dev_xdp_prog_count is exposed so that slave devices can be checked
for loaded XDP programs in order to avoid the situation where both
bond master and slave have programs loaded according to xdp_state.
Signed-off-by: Jussi Maki <redacted>
---
include/linux/filter.h | 13 ++++++++++++-
include/linux/netdevice.h | 6 ++++++
net/core/dev.c | 8 +++++++-
net/core/filter.c | 25 +++++++++++++++++++++++++
4 files changed, 50 insertions(+), 2 deletions(-)
@@ -9486,6 +9487,11 @@ static int dev_xdp_attach(struct net_device *dev, struct netlink_ext_ack *extackNL_SET_ERR_MSG(extack,"XDP_FLAGS_REPLACE is not specified");return-EINVAL;}+/* don't allow loading XDP programs to a bonded device */+if(netif_is_bond_slave(dev)){+NL_SET_ERR_MSG(extack,"XDP program can not be attached to a bond slave");+return-EINVAL;+}mode=dev_xdp_mode(dev,flags);/* can't replace attached link */
@@ -3950,6 +3950,31 @@ void bpf_clear_redirect_map(struct bpf_map *map)}}+DEFINE_STATIC_KEY_FALSE(bpf_master_redirect_enabled_key);+EXPORT_SYMBOL_GPL(bpf_master_redirect_enabled_key);++u32xdp_master_redirect(structxdp_buff*xdp)+{+structnet_device*master,*slave;+structbpf_redirect_info*ri=this_cpu_ptr(&bpf_redirect_info);++master=netdev_master_upper_dev_get_rcu(xdp->rxq->dev);+slave=master->netdev_ops->ndo_xdp_get_xmit_slave(master,xdp);+if(slave&&slave!=xdp->rxq->dev){+/* The target device is different from the receiving device, so+*redirectittothenewdevice.+*UsingXDP_REDIRECTgetsthecorrectbehaviourfromXDPenabled+*driverstounmapthepacketfromtheirrxring.+*/+ri->tgt_index=slave->ifindex;+ri->map_id=INT_MAX;+ri->map_type=BPF_MAP_TYPE_UNSPEC;+returnXDP_REDIRECT;+}+returnXDP_TX;+}+EXPORT_SYMBOL_GPL(xdp_master_redirect);+intxdp_do_redirect(structnet_device*dev,structxdp_buff*xdp,structbpf_prog*xdp_prog){
From: Jussi Maki <redacted>
If the ingress device is bond slave, do not broadcast back
through it or the bond master.
Signed-off-by: Jussi Maki <redacted>
---
kernel/bpf/devmap.c | 69 +++++++++++++++++++++++++++++++++++++++------
1 file changed, 60 insertions(+), 9 deletions(-)
@@ -562,17 +561,48 @@ static int dev_map_enqueue_clone(struct bpf_dtab_netdev *obj,return0;}+staticinlineboolis_ifindex_excluded(int*excluded,intnum_excluded,intifindex)+{+while(num_excluded--){+if(ifindex==excluded[num_excluded])+returntrue;+}+returnfalse;+}++/* Get ifindex of each upper device. 'indexes' must be able to hold at+*leastMAX_NEST_DEVelements.+*Returnsthenumberofifindexesadded.+*/+staticintget_upper_ifindexes(structnet_device*dev,int*indexes)+{+structnet_device*upper;+structlist_head*iter;+intn=0;++netdev_for_each_upper_dev_rcu(dev,upper,iter){+indexes[n++]=upper->ifindex;+}+returnn;+}+intdev_map_enqueue_multi(structxdp_buff*xdp,structnet_device*dev_rx,structbpf_map*map,boolexclude_ingress){structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);-intexclude_ifindex=exclude_ingress?dev_rx->ifindex:0;structbpf_dtab_netdev*dst,*last_dst=NULL;+intexcluded_devices[1+MAX_NEST_DEV];structhlist_head*head;structxdp_frame*xdpf;+intnum_excluded=0;unsignedinti;interr;+if(exclude_ingress){+num_excluded=get_upper_ifindexes(dev_rx,excluded_devices);+excluded_devices[num_excluded++]=dev_rx->ifindex;+}+xdpf=xdp_convert_buff_to_frame(xdp);if(unlikely(!xdpf))return-EOVERFLOW;
@@ -581,7 +611,10 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,for(i=0;i<map->max_entries;i++){dst=rcu_dereference_check(dtab->netdev_map[i],rcu_read_lock_bh_held());-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp))+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -601,7 +634,11 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,head=dev_map_index_hash(dtab,i);hlist_for_each_entry_rcu(dst,head,index_hlist,lockdep_is_held(&dtab->index_lock)){-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp))+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,+dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -675,18 +712,27 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,boolexclude_ingress){structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);-intexclude_ifindex=exclude_ingress?dev->ifindex:0;structbpf_dtab_netdev*dst,*last_dst=NULL;+intexcluded_devices[1+MAX_NEST_DEV];structhlist_head*head;structhlist_node*next;+intnum_excluded=0;unsignedinti;interr;+if(exclude_ingress){+num_excluded=get_upper_ifindexes(dev,excluded_devices);+excluded_devices[num_excluded++]=dev->ifindex;+}+if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=rcu_dereference_check(dtab->netdev_map[i],rcu_read_lock_bh_held());-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst)+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -700,12 +746,17 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,returnerr;last_dst=dst;+}}else{/* BPF_MAP_TYPE_DEVMAP_HASH */for(i=0;i<dtab->n_buckets;i++){head=dev_map_index_hash(dtab,i);hlist_for_each_entry_safe(dst,next,head,index_hlist){-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst)+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,+dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
From: Jussi Maki <redacted>
XDP is implemented in the bonding driver by transparently delegating
the XDP program loading, removal and xmit operations to the bonding
slave devices. The overall goal of this work is that XDP programs
can be attached to a bond device *without* any further changes (or
awareness) necessary to the program itself, meaning the same XDP
program can be attached to a native device but also a bonding device.
Semantics of XDP_TX when attached to a bond are equivalent in such
setting to the case when a tc/BPF program would be attached to the
bond, meaning transmitting the packet out of the bond itself using one
of the bond's configured xmit methods to select a slave device (rather
than XDP_TX on the slave itself). Handling of XDP_TX to transmit
using the configured bonding mechanism is therefore implemented by
rewriting the BPF program return value in bpf_prog_run_xdp. To avoid
performance impact this check is guarded by a static key, which is
incremented when a XDP program is loaded onto a bond device. This
approach was chosen to avoid changes to drivers implementing XDP. If
the slave device does not match the receive device, then XDP_REDIRECT
is transparently used to perform the redirection in order to have
the network driver release the packet from its RX ring. The bonding
driver hashing functions have been refactored to allow reuse with
xdp_buff's to avoid code duplication.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 309 +++++++++++++++++++++++++++++++-
include/net/bonding.h | 1 +
2 files changed, 309 insertions(+), 1 deletion(-)
@@ -317,6 +317,19 @@ bool bond_sk_check(struct bonding *bond)}}+staticboolbond_xdp_check(structbonding*bond)+{+switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_ACTIVEBACKUP:+caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+returntrue;+default:+returnfalse;+}+}+/*---------------------------------- VLAN -----------------------------------*//* In the following 2 functions, bond_vlan_rx_add_vid and bond_vlan_rx_kill_vid,
@@ -2133,6 +2146,41 @@ int bond_enslave(struct net_device *bond_dev, struct net_device *slave_dev,bond_update_slave_arr(bond,NULL);+if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+if(bond->xdp_prog){+NL_SET_ERR_MSG(extack,"Slave does not support XDP");+slave_err(bond_dev,slave_dev,"Slave does not support XDP\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}+}else{+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=bond->xdp_prog,+.extack=extack,+};++if(dev_xdp_prog_count(slave_dev)>0){+NL_SET_ERR_MSG(extack,+"Slave has XDP program loaded, please unload before enslaving");+slave_err(bond_dev,slave_dev,+"Slave has XDP program loaded, please unload before enslaving\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}++res=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(res<0){+/* ndo_bpf() sets extack error message */+slave_dbg(bond_dev,slave_dev,"Error %d calling ndo_bpf\n",res);+gotoerr_sysfs_del;+}+if(bond->xdp_prog)+bpf_prog_inc(bond->xdp_prog);+}+slave_info(bond_dev,slave_dev,"Enslaving as %s interface with %s link\n",bond_is_active_slave(new_slave)?"an active":"a backup",new_slave->link!=BOND_LINK_DOWN?"an up":"a down");
@@ -2252,6 +2300,17 @@ static int __bond_release_one(struct net_device *bond_dev,/* recompute stats just before removing the slave */bond_get_stats(bond->dev,&bond->bond_stats);+if(bond->xdp_prog){+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=NULL,+.extack=NULL,+};+if(slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp))+slave_warn(bond_dev,slave_dev,"failed to unload XDP program\n");+}+bond_upper_dev_unlink(bond,slave);/* unregister rx_handler early so bond_handle_frame wouldn't be called*forthisslaveanymore.
@@ -4420,6 +4499,47 @@ static struct slave *bond_xmit_roundrobin_slave_get(struct bonding *bond,returnNULL;}+staticstructslave*bond_xdp_xmit_roundrobin_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structslave*slave;+intslave_cnt;+u32slave_id;+conststructethhdr*eth;+void*data=xdp->data;++if(data+sizeof(structethhdr)>xdp->data_end)+gotonon_igmp;++eth=(structethhdr*)data;+data+=sizeof(structethhdr);++/* See comment on IGMP in bond_xmit_roundrobin_slave_get() */+if(eth->h_proto==htons(ETH_P_IP)){+conststructiphdr*iph;++if(data+sizeof(structiphdr)>xdp->data_end)+gotonon_igmp;++iph=(structiphdr*)data;++if(iph->protocol==IPPROTO_IGMP){+slave=rcu_dereference(bond->curr_active_slave);+if(slave)+returnslave;+returnbond_get_slave_by_id(bond,0);+}+}++non_igmp:+slave_cnt=READ_ONCE(bond->slave_cnt);+if(likely(slave_cnt)){+slave_id=bond_rr_gen_slave_id(bond)%slave_cnt;+returnbond_get_slave_by_id(bond,slave_id);+}+returnNULL;+}+staticnetdev_tx_tbond_xmit_roundrobin(structsk_buff*skb,structnet_device*bond_dev){
@@ -4635,6 +4755,22 @@ static struct slave *bond_xmit_3ad_xor_slave_get(struct bonding *bond,returnslave;}+staticstructslave*bond_xdp_xmit_3ad_xor_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structbond_up_slave*slaves;+unsignedintcount;+u32hash;++hash=bond_xmit_hash_xdp(bond,xdp);+slaves=rcu_dereference(bond->usable_slaves);+count=slaves?READ_ONCE(slaves->count):0;+if(unlikely(!count))+returnNULL;++returnslaves->arr[hash%count];+}+/* Use this Xmit function for 3AD as well as XOR modes. The current*usableslavearrayisformedinthecontrolpath.Thexmitfunction*justcalculateshashandsendsthepacketout.
@@ -4919,6 +5055,174 @@ static netdev_tx_t bond_start_xmit(struct sk_buff *skb, struct net_device *dev)returnret;}+staticstructnet_device*+bond_xdp_get_xmit_slave(structnet_device*bond_dev,structxdp_buff*xdp)+{+structbonding*bond=netdev_priv(bond_dev);+structslave*slave;++/* Caller needs to hold rcu_read_lock() */++switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+slave=bond_xdp_xmit_roundrobin_slave_get(bond,xdp);+break;++caseBOND_MODE_ACTIVEBACKUP:+slave=bond_xmit_activebackup_slave_get(bond);+break;++caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+slave=bond_xdp_xmit_3ad_xor_slave_get(bond,xdp);+break;++default:+/* Should never happen. Mode guarded by bond_xdp_check() */+netdev_err(bond_dev,"Unknown bonding mode %d for xdp xmit\n",BOND_MODE(bond));+WARN_ON_ONCE(1);+returnNULL;+}++if(slave)+returnslave->dev;++returnNULL;+}++staticintbond_xdp_xmit(structnet_device*bond_dev,+intn,structxdp_frame**frames,u32flags)+{+intnxmit,err=-ENXIO;++rcu_read_lock();++for(nxmit=0;nxmit<n;nxmit++){+structxdp_frame*frame=frames[nxmit];+structxdp_frame*frames1[]={frame};+structnet_device*slave_dev;+structxdp_buffxdp;++xdp_convert_frame_to_buff(frame,&xdp);++slave_dev=bond_xdp_get_xmit_slave(bond_dev,&xdp);+if(!slave_dev){+err=-ENXIO;+break;+}++err=slave_dev->netdev_ops->ndo_xdp_xmit(slave_dev,1,frames1,flags);+if(err<1)+break;+}++rcu_read_unlock();++/* If error happened on the first frame then we can pass the error up, otherwise+*reportthenumberofframesthatwerexmitted.+*/+if(err<0)+return(nxmit==0?err:nxmit);++returnnxmit;+}++staticintbond_xdp_set(structnet_device*dev,structbpf_prog*prog,+structnetlink_ext_ack*extack)+{+structbonding*bond=netdev_priv(dev);+structlist_head*iter;+structslave*slave,*rollback_slave;+structbpf_prog*old_prog;+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=prog,+.extack=extack,+};+interr;++ASSERT_RTNL();++if(!bond_xdp_check(bond))+return-EOPNOTSUPP;++old_prog=bond->xdp_prog;+bond->xdp_prog=prog;++bond_for_each_slave(bond,slave,iter){+structnet_device*slave_dev=slave->dev;++if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+NL_SET_ERR_MSG(extack,"Slave device does not support XDP");+slave_err(dev,slave_dev,"Slave does not support XDP\n");+err=-EOPNOTSUPP;+gotoerr;+}++if(dev_xdp_prog_count(slave_dev)>0){+NL_SET_ERR_MSG(extack,+"Slave has XDP program loaded, please unload before enslaving");+slave_err(dev,slave_dev,+"Slave has XDP program loaded, please unload before enslaving\n");+err=-EOPNOTSUPP;+gotoerr;+}++err=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err<0){+/* ndo_bpf() sets extack error message */+slave_err(dev,slave_dev,"Error %d calling ndo_bpf\n",err);+gotoerr;+}+if(prog)+bpf_prog_inc(prog);+}++if(old_prog)+bpf_prog_put(old_prog);++if(prog)+static_branch_inc(&bpf_master_redirect_enabled_key);+else+static_branch_dec(&bpf_master_redirect_enabled_key);++return0;++err:+/* unwind the program changes */+bond->xdp_prog=old_prog;+xdp.prog=old_prog;+xdp.extack=NULL;/* do not overwrite original error */++bond_for_each_slave(bond,rollback_slave,iter){+structnet_device*slave_dev=rollback_slave->dev;+interr_unwind;++if(slave==rollback_slave)+break;++err_unwind=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err_unwind<0)+slave_err(dev,slave_dev,+"Error %d when unwinding XDP program change\n",err_unwind);+elseif(xdp.prog)+bpf_prog_inc(xdp.prog);+}+returnerr;+}++staticintbond_xdp(structnet_device*dev,structnetdev_bpf*xdp)+{+switch(xdp->command){+caseXDP_SETUP_PROG:+returnbond_xdp_set(dev,xdp->prog,xdp->extack);+default:+return-EINVAL;+}+}+staticu32bond_mode_bcast_speed(structslave*slave,u32speed){if(speed==0||speed==SPEED_UNKNOWN)
From: Jussi Maki <redacted>
For the XDP bonding slave lookup to work in the NAPI poll context
in which the redudant rcu_read_lock() has been removed we have to
follow the same approach as in [1] and modify the WARN_ON to also
check rcu_read_lock_bh_held().
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=694cea395fded425008e93cd90cfdf7a451674af
Signed-off-by: Jussi Maki <redacted>
---
net/core/dev.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Jussi Maki <redacted>
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Signed-off-by: Jussi Maki <redacted>
---
.../selftests/bpf/prog_tests/xdp_bonding.c | 467 ++++++++++++++++++
1 file changed, 467 insertions(+)
@@ -0,0 +1,467 @@+// SPDX-License-Identifier: GPL-2.0++/**+*TestXDPbondingsupport+*+*Setsuptwobondedvethpairsbetweentwofreshnamespaces+*andverifiesthatXDP_TXprogramloadedonabonddevice+*arecorrectlyloadedontotheslavedevicesandXDP_TX'd+*packetsarebalancedusingbonding.+*/++#define _GNU_SOURCE+#include<sched.h>+#include<net/if.h>+#include<linux/if_link.h>+#include"test_progs.h"+#include"network_helpers.h"+#include<linux/if_bonding.h>+#include<linux/limits.h>+#include<linux/udp.h>++#define BOND1_MAC {0x00, 0x11, 0x22, 0x33, 0x44, 0x55}+#define BOND1_MAC_STR "00:11:22:33:44:55"+#define BOND2_MAC {0x00, 0x22, 0x33, 0x44, 0x55, 0x66}+#define BOND2_MAC_STR "00:22:33:44:55:66"+#define NPACKETS 100++staticintroot_netns_fd=-1;++staticvoidrestore_root_netns(void)+{+ASSERT_OK(setns(root_netns_fd,CLONE_NEWNET),"restore_root_netns");+}++intsetns_by_name(char*name)+{+intnsfd,err;+charnspath[PATH_MAX];++snprintf(nspath,sizeof(nspath),"%s/%s","/var/run/netns",name);+nsfd=open(nspath,O_RDONLY|O_CLOEXEC);+if(nsfd<0)+return-1;++err=setns(nsfd,CLONE_NEWNET);+close(nsfd);+returnerr;+}++staticintget_rx_packets(constchar*iface)+{+FILE*f;+charline[512];+intiface_len=strlen(iface);++f=fopen("/proc/net/dev","r");+if(!f)+return-1;++while(fgets(line,sizeof(line),f)){+char*p=line;++while(*p==' ')+p++;/* skip whitespace */+if(!strncmp(p,iface,iface_len)){+p+=iface_len;+if(*p++!=':')+continue;+while(*p==' ')+p++;/* skip whitespace */+while(*p&&*p!=' ')+p++;/* skip rx bytes */+while(*p==' ')+p++;/* skip whitespace */+fclose(f);+returnatoi(p);+}+}+fclose(f);+return-1;+}++enum{+BOND_ONE_NO_ATTACH=0,+BOND_BOTH_AND_ATTACH,+};++staticconstchar*constmode_names[]={+[BOND_MODE_ROUNDROBIN]="balance-rr",+[BOND_MODE_ACTIVEBACKUP]="active-backup",+[BOND_MODE_XOR]="balance-xor",+[BOND_MODE_BROADCAST]="broadcast",+[BOND_MODE_8023AD]="802.3ad",+[BOND_MODE_TLB]="balance-tlb",+[BOND_MODE_ALB]="balance-alb",+};++staticconstchar*constxmit_policy_names[]={+[BOND_XMIT_POLICY_LAYER2]="layer2",+[BOND_XMIT_POLICY_LAYER34]="layer3+4",+[BOND_XMIT_POLICY_LAYER23]="layer2+3",+[BOND_XMIT_POLICY_ENCAP23]="encap2+3",+[BOND_XMIT_POLICY_ENCAP34]="encap3+4",+};++#define MAX_LOADED 8+staticstructbpf_object*loaded_bpf_objects[MAX_LOADED]={};+staticintn_loaded_bpf_objects;++staticintload_xdp_program(constchar*filename,constchar*sec_name,constchar*iface)+{+structbpf_prog_load_attrprog_load_attr={+.prog_type=BPF_PROG_TYPE_XDP,+.file=filename,+};+structbpf_program*prog;+structbpf_object*obj;+intprog_fd=-1;+intifindex,err;++err=bpf_prog_load_xattr(&prog_load_attr,&obj,&prog_fd);+if(!ASSERT_OK(err,"prog load xattr"))+returnerr;++prog=bpf_object__find_program_by_title(obj,sec_name);+if(!ASSERT_OK_PTR(prog,"find program"))+gotoerr;++prog_fd=bpf_program__fd(prog);+if(!ASSERT_GE(prog_fd,0,"get program fd"))+gotoerr;++ifindex=if_nametoindex(iface);+if(!ASSERT_GT(ifindex,0,"get ifindex"))+gotoerr;++err=bpf_set_link_xdp_fd(ifindex,prog_fd,XDP_FLAGS_DRV_MODE|XDP_FLAGS_DRV_MODE);+if(!ASSERT_OK(err,"load xdp program"))+gotoerr;++loaded_bpf_objects[n_loaded_bpf_objects++]=obj;+if(n_loaded_bpf_objects==MAX_LOADED){+fprintf(stderr,"Too many loaded BPF objects\n");+gotoerr;+}++return0;++err:+bpf_object__close(obj);+return-1;+}++staticintbonding_setup(intmode,intxmit_policy,intbond_both_attach)+{+#define SYS(fmt, ...) \+({\+charcmd[1024];\+snprintf(cmd,sizeof(cmd),fmt,##__VA_ARGS__);\+if(!ASSERT_OK(system(cmd),cmd))\+return-1;\+})++SYS("ip netns add ns_dst");+SYS("ip link add veth1_1 type veth peer name veth2_1 netns ns_dst");+SYS("ip link add veth1_2 type veth peer name veth2_2 netns ns_dst");++SYS("ip link add bond1 type bond mode %s xmit_hash_policy %s",+mode_names[mode],xmit_policy_names[xmit_policy]);+SYS("ip link set bond1 up address "BOND1_MAC_STR" addrgenmode none");+SYS("ip -netns ns_dst link add bond2 type bond mode %s xmit_hash_policy %s",+mode_names[mode],xmit_policy_names[xmit_policy]);+SYS("ip -netns ns_dst link set bond2 up address "BOND2_MAC_STR" addrgenmode none");++SYS("ip link set veth1_1 master bond1");+if(bond_both_attach==BOND_BOTH_AND_ATTACH){+SYS("ip link set veth1_2 master bond1");+}else{+SYS("ip link set veth1_2 up addrgenmode none");++if(load_xdp_program("xdp_dummy.o","xdp_dummy","veth1_2"))+return-1;+}++SYS("ip -netns ns_dst link set veth2_1 master bond2");++if(bond_both_attach==BOND_BOTH_AND_ATTACH)+SYS("ip -netns ns_dst link set veth2_2 master bond2");+else+SYS("ip -netns ns_dst link set veth2_2 up addrgenmode none");++/* Load a dummy program on sending side as with veth peer needs to have a+*XDPprogramloadedaswell.+*/+if(load_xdp_program("xdp_dummy.o","xdp_dummy","bond1"))+return-1;++if(bond_both_attach==BOND_BOTH_AND_ATTACH){+if(!ASSERT_OK(setns_by_name("ns_dst"),"set netns to ns_dst"))+return-1;+if(load_xdp_program("xdp_tx.o","tx","bond2"))+return-1;+restore_root_netns();+}++#undef SYS+return0;+}++staticvoidbonding_cleanup(void)+{+restore_root_netns();+while(n_loaded_bpf_objects){+n_loaded_bpf_objects--;+bpf_object__close(loaded_bpf_objects[n_loaded_bpf_objects]);+}+ASSERT_OK(system("ip link delete bond1"),"delete bond1");+ASSERT_OK(system("ip link delete veth1_1"),"delete veth1_1");+ASSERT_OK(system("ip link delete veth1_2"),"delete veth1_2");+ASSERT_OK(system("ip netns delete ns_dst"),"delete ns_dst");+}++staticintsend_udp_packets(intvary_dst_ip)+{+structethhdreh={+.h_source=BOND1_MAC,+.h_dest=BOND2_MAC,+.h_proto=htons(ETH_P_IP),+};+uint8_tbuf[128]={};+structiphdr*iph=(structiphdr*)(buf+sizeof(eh));+structudphdr*uh=(structudphdr*)(buf+sizeof(eh)+sizeof(*iph));+inti,s=-1;+intifindex;++s=socket(AF_PACKET,SOCK_RAW,IPPROTO_RAW);+if(!ASSERT_GE(s,0,"socket"))+gotoerr;++ifindex=if_nametoindex("bond1");+if(!ASSERT_GT(ifindex,0,"get bond1 ifindex"))+gotoerr;++memcpy(buf,&eh,sizeof(eh));+iph->ihl=5;+iph->version=4;+iph->tos=16;+iph->id=1;+iph->ttl=64;+iph->protocol=IPPROTO_UDP;+iph->saddr=1;+iph->daddr=2;+iph->tot_len=htons(sizeof(buf)-ETH_HLEN);+iph->check=0;++for(i=1;i<=NPACKETS;i++){+intn;+structsockaddr_llsaddr_ll={+.sll_ifindex=ifindex,+.sll_halen=ETH_ALEN,+.sll_addr=BOND2_MAC,+};++/* vary the UDP destination port for even distribution with roundrobin/xor modes */+uh->dest++;++if(vary_dst_ip)+iph->daddr++;++n=sendto(s,buf,sizeof(buf),0,(structsockaddr*)&saddr_ll,sizeof(saddr_ll));+if(!ASSERT_EQ(n,sizeof(buf),"sendto"))+gotoerr;+}++return0;++err:+if(s>=0)+close(s);+return-1;+}++voidtest_xdp_bonding_with_mode(char*name,intmode,intxmit_policy)+{+intbond1_rx;++if(!test__start_subtest(name))+return;++if(bonding_setup(mode,xmit_policy,BOND_BOTH_AND_ATTACH))+gotoout;++if(send_udp_packets(xmit_policy!=BOND_XMIT_POLICY_LAYER34))+gotoout;++bond1_rx=get_rx_packets("bond1");+ASSERT_EQ(bond1_rx,NPACKETS,"expected more received packets");++switch(mode){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_XOR:{+intveth1_rx=get_rx_packets("veth1_1");+intveth2_rx=get_rx_packets("veth1_2");+intdiff=abs(veth1_rx-veth2_rx);++ASSERT_GE(veth1_rx+veth2_rx,NPACKETS,"expected more packets");++switch(xmit_policy){+caseBOND_XMIT_POLICY_LAYER2:+ASSERT_GE(diff,NPACKETS,+"expected packets on only one of the interfaces");+break;+caseBOND_XMIT_POLICY_LAYER23:+caseBOND_XMIT_POLICY_LAYER34:+ASSERT_LT(diff,NPACKETS/2,+"expected even distribution of packets");+break;+default:+PRINT_FAIL("Unimplemented xmit_policy=%d\n",xmit_policy);+break;+}+break;+}+caseBOND_MODE_ACTIVEBACKUP:{+intveth1_rx=get_rx_packets("veth1_1");+intveth2_rx=get_rx_packets("veth1_2");+intdiff=abs(veth1_rx-veth2_rx);++ASSERT_GE(diff,NPACKETS,+"expected packets on only one of the interfaces");+break;+}+default:+PRINT_FAIL("Unimplemented xmit_policy=%d\n",xmit_policy);+break;+}++out:+bonding_cleanup();+}+++/* Test the broadcast redirection using xdp_redirect_map_multi_prog and adding+*alltheinterfacestoitandcheckingthatbroadcastingwon'tsendthepacket+*toneithertheingressbonddevice(bond2)oritsslave(veth2_1).+*/+voidtest_xdp_bonding_redirect_multi(void)+{+staticconstchar*constifaces[]={"bond2","veth2_1","veth2_2"};+structbpf_prog_load_attrprog_load_attr={+.prog_type=BPF_PROG_TYPE_UNSPEC,+.file="xdp_redirect_multi_kern.o",+};+structbpf_program*redirect_prog;+intprog_fd,map_all_fd;+structbpf_object*obj;+intveth1_1_rx,veth1_2_rx;+interr;++if(!test__start_subtest("xdp_bonding_redirect_multi"))+return;++if(bonding_setup(BOND_MODE_ROUNDROBIN,BOND_XMIT_POLICY_LAYER23,BOND_ONE_NO_ATTACH))+gotoout;++err=bpf_prog_load_xattr(&prog_load_attr,&obj,&prog_fd);+if(!ASSERT_OK(err,"prog load xattr"))+gotoout;++map_all_fd=bpf_object__find_map_fd_by_name(obj,"map_all");+if(!ASSERT_GE(map_all_fd,0,"find map_all fd"))+gotoout;++redirect_prog=bpf_object__find_program_by_name(obj,"xdp_redirect_map_multi_prog");+if(!ASSERT_OK_PTR(redirect_prog,"find xdp_redirect_map_multi_prog"))+gotoout;++prog_fd=bpf_program__fd(redirect_prog);+if(!ASSERT_GE(prog_fd,0,"get prog fd"))+gotoout;++if(!ASSERT_OK(setns_by_name("ns_dst"),"could not set netns to ns_dst"))+gotoout;++/* populate the devmap with the relevant interfaces */+for(inti=0;i<ARRAY_SIZE(ifaces);i++){+intifindex=if_nametoindex(ifaces[i]);++if(!ASSERT_GT(ifindex,0,"could not get interface index"))+gotoout;++if(!ASSERT_OK(bpf_map_update_elem(map_all_fd,&ifindex,&ifindex,0),+"add interface to map_all"))+gotoout;+}++/* finally attach the program */+err=bpf_set_link_xdp_fd(if_nametoindex("bond2"),prog_fd,+XDP_FLAGS_DRV_MODE|XDP_FLAGS_UPDATE_IF_NOEXIST);+if(!ASSERT_OK(err,"set bond2 xdp"))+gotoout;++restore_root_netns();++if(send_udp_packets(BOND_MODE_ROUNDROBIN))+gotoout;++veth1_1_rx=get_rx_packets("veth1_1");+veth1_2_rx=get_rx_packets("veth1_2");++ASSERT_EQ(veth1_1_rx,0,"expected no packets on veth1_1");+ASSERT_GE(veth1_2_rx,NPACKETS,"expected packets on veth1_2");++out:+restore_root_netns();+bpf_object__close(obj);+bonding_cleanup();+}++staticintlibbpf_debug_print(enumlibbpf_print_levellevel,+constchar*format,va_listargs)+{+if(level!=LIBBPF_WARN)+vprintf(format,args);+return0;+}++structbond_test_case{+char*name;+intmode;+intxmit_policy;+};++staticstructbond_test_casebond_test_cases[]={+{"xdp_bonding_roundrobin",BOND_MODE_ROUNDROBIN,BOND_XMIT_POLICY_LAYER23,},+{"xdp_bonding_activebackup",BOND_MODE_ACTIVEBACKUP,BOND_XMIT_POLICY_LAYER23},++{"xdp_bonding_xor_layer2",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER2,},+{"xdp_bonding_xor_layer23",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER23,},+{"xdp_bonding_xor_layer34",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER34,},+};++voidtest_xdp_bonding(void)+{+libbpf_print_fn_told_print_fn;+inti;++old_print_fn=libbpf_set_print(libbpf_debug_print);++root_netns_fd=open("/proc/self/ns/net",O_RDONLY);+if(!ASSERT_GE(root_netns_fd,0,"open /proc/self/ns/net"))+return;++for(i=0;i<ARRAY_SIZE(bond_test_cases);i++){+structbond_test_case*test_case=&bond_test_cases[i];++test_xdp_bonding_with_mode(+test_case->name,+test_case->mode,+test_case->xmit_policy);+}++test_xdp_bonding_redirect_multi();++libbpf_set_print(old_print_fn);+close(root_netns_fd);+}
From: Jussi Maki <hidden> Date: 2021-08-03 09:41:28
On Tue, Aug 3, 2021 at 2:19 AM Andrii Nakryiko
[off-list ref] wrote:
On Mon, Aug 2, 2021 at 6:24 AM [off-list ref] wrote:
quoted
From: Jussi Maki <redacted>
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Signed-off-by: Jussi Maki <redacted>
---
Was there any reason not to use BPF skeleton in your new tests? And
also bpf_link-based XDP attachment instead of netlink-based?
Not really. I used the existing xdp_redirect_multi test as basis and
that used this approach. I'll give a go at changing this to use the
BPF skeletons.
From: Jussi Maki <hidden> Date: 2021-08-04 12:46:14
This patchset introduces XDP support to the bonding driver.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter. Fix for this
has been already merged into net-next. The statistics were collected
using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Patch 1 prepares bond_xmit_hash for hashing xdp_buff's.
Patch 2 adds hooks to implement redirection after bpf prog run.
Patch 3 implements the hooks in the bonding driver.
Patch 4 modifies devmap to properly handle EXCLUDE_INGRESS with a slave device.
Patch 5 fixes an issue related to recent cleanup of rcu_read_lock in XDP context.
Patch 6 fixes loading of xdp_tx.o by renaming section name.
Patch 7 adds tests.
v4->v5:
- As pointed by Andrii, use the generated BPF skeletons rather than libbpf
directly.
- Renamed section name in progs/xdp_tx.c as the BPF skeleton wouldn't load it
otherwise due to unknown program type.
- Daniel Borkmann noted that to retain backwards compatibility and allow some
use cases we should allow attaching XDP programs to a slave device when the
master does not have a program loaded. Modified the logic to allow this and
added tests for the different combinations of attaching a program.
v3->v4:
- Add back the test suite, while removing the vmtest.sh modifications to kernel
config new that CONFIG_BONDING=y is set. Discussed with Magnus Karlsson that
it makes sense right now to not reuse the code from xdpceiver.c for testing
XDP bonding.
v2->v3:
- Address Jay's comment to properly exclude upper devices with EXCLUDE_INGRESS
when there are deeper nesting involved. Now all upper devices are excluded.
- Refuse to enslave devices that already have XDP programs loaded and refuse to
load XDP programs to slave devices. Earlier one could have a XDP program loaded
and after enslaving and loading another program onto the bond device the xdp_state
of the enslaved device would be pointing at an old program.
- Adapt netdev_lower_get_next_private_rcu so it can be called in the XDP context.
v1->v2:
- Split up into smaller easier to review patches and address cosmetic
review comments.
- Drop the INDIRECT_CALL optimization as it showed little improvement in tests.
- Drop the rr_tx_counter patch as that has already been merged into net-next.
- Separate the test suite into another patch set. This will follow later once a
patch set from Magnus Karlsson is merged and provides test utilities that can
be reused for XDP bonding tests. v2 contains no major functional changes and
was tested with the test suite included in v1.
(https://lore.kernel.org/bpf/202106221509.kwNvAAZg-lkp@intel.com/T/#m464146d47299125d5868a08affd6d6ce526dfad1)
---
From: Jussi Maki <hidden> Date: 2021-08-04 12:46:21
If the ingress device is bond slave, do not broadcast back
through it or the bond master.
Signed-off-by: Jussi Maki <redacted>
---
kernel/bpf/devmap.c | 69 +++++++++++++++++++++++++++++++++++++++------
1 file changed, 60 insertions(+), 9 deletions(-)
@@ -562,17 +561,48 @@ static int dev_map_enqueue_clone(struct bpf_dtab_netdev *obj,return0;}+staticinlineboolis_ifindex_excluded(int*excluded,intnum_excluded,intifindex)+{+while(num_excluded--){+if(ifindex==excluded[num_excluded])+returntrue;+}+returnfalse;+}++/* Get ifindex of each upper device. 'indexes' must be able to hold at+*leastMAX_NEST_DEVelements.+*Returnsthenumberofifindexesadded.+*/+staticintget_upper_ifindexes(structnet_device*dev,int*indexes)+{+structnet_device*upper;+structlist_head*iter;+intn=0;++netdev_for_each_upper_dev_rcu(dev,upper,iter){+indexes[n++]=upper->ifindex;+}+returnn;+}+intdev_map_enqueue_multi(structxdp_buff*xdp,structnet_device*dev_rx,structbpf_map*map,boolexclude_ingress){structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);-intexclude_ifindex=exclude_ingress?dev_rx->ifindex:0;structbpf_dtab_netdev*dst,*last_dst=NULL;+intexcluded_devices[1+MAX_NEST_DEV];structhlist_head*head;structxdp_frame*xdpf;+intnum_excluded=0;unsignedinti;interr;+if(exclude_ingress){+num_excluded=get_upper_ifindexes(dev_rx,excluded_devices);+excluded_devices[num_excluded++]=dev_rx->ifindex;+}+xdpf=xdp_convert_buff_to_frame(xdp);if(unlikely(!xdpf))return-EOVERFLOW;
@@ -581,7 +611,10 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,for(i=0;i<map->max_entries;i++){dst=rcu_dereference_check(dtab->netdev_map[i],rcu_read_lock_bh_held());-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp))+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -601,7 +634,11 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,head=dev_map_index_hash(dtab,i);hlist_for_each_entry_rcu(dst,head,index_hlist,lockdep_is_held(&dtab->index_lock)){-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp))+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,+dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -675,18 +712,27 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,boolexclude_ingress){structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);-intexclude_ifindex=exclude_ingress?dev->ifindex:0;structbpf_dtab_netdev*dst,*last_dst=NULL;+intexcluded_devices[1+MAX_NEST_DEV];structhlist_head*head;structhlist_node*next;+intnum_excluded=0;unsignedinti;interr;+if(exclude_ingress){+num_excluded=get_upper_ifindexes(dev,excluded_devices);+excluded_devices[num_excluded++]=dev->ifindex;+}+if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=rcu_dereference_check(dtab->netdev_map[i],rcu_read_lock_bh_held());-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst)+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -700,12 +746,17 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,returnerr;last_dst=dst;+}}else{/* BPF_MAP_TYPE_DEVMAP_HASH */for(i=0;i<dtab->n_buckets;i++){head=dev_map_index_hash(dtab,i);hlist_for_each_entry_safe(dst,next,head,index_hlist){-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst)+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,+dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
From: Jussi Maki <hidden> Date: 2021-08-04 12:46:23
In preparation for adding XDP support to the bonding driver
refactor the packet hashing functions to be able to work with
any linear data buffer without an skb.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 147 +++++++++++++++++++-------------
1 file changed, 90 insertions(+), 57 deletions(-)
@@ -3611,55 +3611,80 @@ static struct notifier_block bond_netdev_notifier = {/*---------------------------- Hashing Policies -----------------------------*/+/* Helper to access data in a packet, with or without a backing skb.+*Ifskbisgiventhedataislinearizedifnecessaryviapskb_may_pull.+*/+staticinlineconstvoid*bond_pull_data(structsk_buff*skb,+constvoid*data,inthlen,intn)+{+if(likely(n<=hlen))+returndata;+elseif(skb&&likely(pskb_may_pull(skb,n)))+returnskb->head;++returnNULL;+}+/* L2 hash helper */-staticinlineu32bond_eth_hash(structsk_buff*skb)+staticinlineu32bond_eth_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*ep,hdr_tmp;+structethhdr*ep;-ep=skb_header_pointer(skb,0,sizeof(hdr_tmp),&hdr_tmp);-if(ep)-returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;-return0;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;++ep=(structethhdr*)(data+mhoff);+returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;}-staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,-int*noff,int*proto,booll34)+staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,constvoid*data,+inthlen,__be16l2_proto,int*nhoff,int*ip_proto,booll34){conststructipv6hdr*iph6;conststructiphdr*iph;-if(skb->protocol==htons(ETH_P_IP)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph))))+if(l2_proto==htons(ETH_P_IP)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph));+if(!data)returnfalse;-iph=(conststructiphdr*)(skb->data+*noff);++iph=(conststructiphdr*)(data+*nhoff);iph_to_flow_copy_v4addrs(fk,iph);-*noff+=iph->ihl<<2;+*nhoff+=iph->ihl<<2;if(!ip_is_fragment(iph))-*proto=iph->protocol;-}elseif(skb->protocol==htons(ETH_P_IPV6)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph6))))+*ip_proto=iph->protocol;+}elseif(l2_proto==htons(ETH_P_IPV6)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph6));+if(!data)returnfalse;-iph6=(conststructipv6hdr*)(skb->data+*noff);++iph6=(conststructipv6hdr*)(data+*nhoff);iph_to_flow_copy_v6addrs(fk,iph6);-*noff+=sizeof(*iph6);-*proto=iph6->nexthdr;+*nhoff+=sizeof(*iph6);+*ip_proto=iph6->nexthdr;}else{returnfalse;}-if(l34&&*proto>=0)-fk->ports.ports=skb_flow_get_ports(skb,*noff,*proto);+if(l34&&*ip_proto>=0)+fk->ports.ports=__skb_flow_get_ports(skb,*nhoff,*ip_proto,data,hlen);returntrue;}-staticu32bond_vlan_srcmac_hash(structsk_buff*skb)+staticu32bond_vlan_srcmac_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*mac_hdr=(structethhdr*)skb_mac_header(skb);+structethhdr*mac_hdr;u32srcmac_vendor=0,srcmac_dev=0;u16vlan;inti;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;+mac_hdr=(structethhdr*)(data+mhoff);+for(i=0;i<3;i++)srcmac_vendor=(srcmac_vendor<<8)|mac_hdr->h_source[i];
@@ -3675,26 +3700,25 @@ static u32 bond_vlan_srcmac_hash(struct sk_buff *skb)}/* Extract the appropriate headers based on bond's xmit policy */-staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,-structflow_keys*fk)+staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intnhoff,inthlen,structflow_keys*fk){booll34=bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER34;-intnoff,proto=-1;+intip_proto=-1;switch(bond->params.xmit_policy){caseBOND_XMIT_POLICY_ENCAP23:caseBOND_XMIT_POLICY_ENCAP34:memset(fk,0,sizeof(*fk));return__skb_flow_dissect(NULL,skb,&flow_keys_bonding,-fk,NULL,0,0,0,0);+fk,data,l2_proto,nhoff,hlen,0);default:break;}fk->ports.ports=0;memset(&fk->icmp,0,sizeof(fk->icmp));-noff=skb_network_offset(skb);-if(!bond_flow_ip(skb,fk,&noff,&proto,l34))+if(!bond_flow_ip(skb,fk,data,hlen,l2_proto,&nhoff,&ip_proto,l34))returnfalse;/* ICMP error packets contains at least 8 bytes of the header
@@ -3733,33 +3755,26 @@ static u32 bond_ip_hash(u32 hash, struct flow_keys *flow)returnhash>>1;}-/**-*bond_xmit_hash-generateahashvaluebasedonthexmitpolicy-*@bond:bondingdevice-*@skb:buffertouseforheaders-*-*Thisfunctionwillextractthenecessaryheadersfromtheskbbufferanduse-*themtogenerateahashbasedonthexmit_policysetinthebondingdevice+/* Generate hash based on xmit policy. If @skb is given it is used to linearize+*thedataasrequired,butthisfunctioncanbeusedwithoutitifthedatais+*knowntobelinear(e.g.withxdp_buff).*/-u32bond_xmit_hash(structbonding*bond,structsk_buff*skb)+staticu32__bond_xmit_hash(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intmhoff,intnhoff,inthlen){structflow_keysflow;u32hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP34&&-skb->l4_hash)-returnskb->hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_VLAN_SRCMAC)-returnbond_vlan_srcmac_hash(skb);+returnbond_vlan_srcmac_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER2||-!bond_flow_dissect(bond,skb,&flow))-returnbond_eth_hash(skb);+!bond_flow_dissect(bond,skb,data,l2_proto,nhoff,hlen,&flow))+returnbond_eth_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER23||bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP23){-hash=bond_eth_hash(skb);+hash=bond_eth_hash(skb,data,mhoff,hlen);}else{if(flow.icmp.id)memcpy(&hash,&flow.icmp,sizeof(hash));
From: Jussi Maki <hidden> Date: 2021-08-04 12:46:24
This adds the ndo_xdp_get_xmit_slave hook for transforming XDP_TX
into XDP_REDIRECT after BPF program run when the ingress device
is a bond slave.
The dev_xdp_prog_count is exposed so that slave devices can be checked
for loaded XDP programs in order to avoid the situation where both
bond master and slave have programs loaded according to xdp_state.
Signed-off-by: Jussi Maki <redacted>
---
include/linux/filter.h | 13 ++++++++++++-
include/linux/netdevice.h | 6 ++++++
net/core/dev.c | 13 ++++++++++++-
net/core/filter.c | 25 +++++++++++++++++++++++++
4 files changed, 55 insertions(+), 2 deletions(-)
@@ -9494,6 +9497,14 @@ static int dev_xdp_attach(struct net_device *dev, struct netlink_ext_ack *extackreturn-EBUSY;}+/* don't allow if an upper device already has a program */+netdev_for_each_upper_dev_rcu(dev,upper,iter){+if(dev_xdp_prog_count(upper)>0){+NL_SET_ERR_MSG(extack,"Cannot attach when an upper device already has a program");+return-EEXIST;+}+}+cur_prog=dev_xdp_prog(dev,mode);/* can't replace attached prog with link */if(link&&cur_prog){
@@ -3950,6 +3950,31 @@ void bpf_clear_redirect_map(struct bpf_map *map)}}+DEFINE_STATIC_KEY_FALSE(bpf_master_redirect_enabled_key);+EXPORT_SYMBOL_GPL(bpf_master_redirect_enabled_key);++u32xdp_master_redirect(structxdp_buff*xdp)+{+structnet_device*master,*slave;+structbpf_redirect_info*ri=this_cpu_ptr(&bpf_redirect_info);++master=netdev_master_upper_dev_get_rcu(xdp->rxq->dev);+slave=master->netdev_ops->ndo_xdp_get_xmit_slave(master,xdp);+if(slave&&slave!=xdp->rxq->dev){+/* The target device is different from the receiving device, so+*redirectittothenewdevice.+*UsingXDP_REDIRECTgetsthecorrectbehaviourfromXDPenabled+*driverstounmapthepacketfromtheirrxring.+*/+ri->tgt_index=slave->ifindex;+ri->map_id=INT_MAX;+ri->map_type=BPF_MAP_TYPE_UNSPEC;+returnXDP_REDIRECT;+}+returnXDP_TX;+}+EXPORT_SYMBOL_GPL(xdp_master_redirect);+intxdp_do_redirect(structnet_device*dev,structxdp_buff*xdp,structbpf_prog*xdp_prog){
From: Jussi Maki <hidden> Date: 2021-08-04 12:46:25
XDP is implemented in the bonding driver by transparently delegating
the XDP program loading, removal and xmit operations to the bonding
slave devices. The overall goal of this work is that XDP programs
can be attached to a bond device *without* any further changes (or
awareness) necessary to the program itself, meaning the same XDP
program can be attached to a native device but also a bonding device.
Semantics of XDP_TX when attached to a bond are equivalent in such
setting to the case when a tc/BPF program would be attached to the
bond, meaning transmitting the packet out of the bond itself using one
of the bond's configured xmit methods to select a slave device (rather
than XDP_TX on the slave itself). Handling of XDP_TX to transmit
using the configured bonding mechanism is therefore implemented by
rewriting the BPF program return value in bpf_prog_run_xdp. To avoid
performance impact this check is guarded by a static key, which is
incremented when a XDP program is loaded onto a bond device. This
approach was chosen to avoid changes to drivers implementing XDP. If
the slave device does not match the receive device, then XDP_REDIRECT
is transparently used to perform the redirection in order to have
the network driver release the packet from its RX ring. The bonding
driver hashing functions have been refactored to allow reuse with
xdp_buff's to avoid code duplication.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 309 +++++++++++++++++++++++++++++++-
include/net/bonding.h | 1 +
2 files changed, 309 insertions(+), 1 deletion(-)
@@ -317,6 +317,19 @@ bool bond_sk_check(struct bonding *bond)}}+staticboolbond_xdp_check(structbonding*bond)+{+switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_ACTIVEBACKUP:+caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+returntrue;+default:+returnfalse;+}+}+/*---------------------------------- VLAN -----------------------------------*//* In the following 2 functions, bond_vlan_rx_add_vid and bond_vlan_rx_kill_vid,
@@ -2133,6 +2146,41 @@ int bond_enslave(struct net_device *bond_dev, struct net_device *slave_dev,bond_update_slave_arr(bond,NULL);+if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+if(bond->xdp_prog){+NL_SET_ERR_MSG(extack,"Slave does not support XDP");+slave_err(bond_dev,slave_dev,"Slave does not support XDP\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}+}else{+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=bond->xdp_prog,+.extack=extack,+};++if(dev_xdp_prog_count(slave_dev)>0){+NL_SET_ERR_MSG(extack,+"Slave has XDP program loaded, please unload before enslaving");+slave_err(bond_dev,slave_dev,+"Slave has XDP program loaded, please unload before enslaving\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}++res=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(res<0){+/* ndo_bpf() sets extack error message */+slave_dbg(bond_dev,slave_dev,"Error %d calling ndo_bpf\n",res);+gotoerr_sysfs_del;+}+if(bond->xdp_prog)+bpf_prog_inc(bond->xdp_prog);+}+slave_info(bond_dev,slave_dev,"Enslaving as %s interface with %s link\n",bond_is_active_slave(new_slave)?"an active":"a backup",new_slave->link!=BOND_LINK_DOWN?"an up":"a down");
@@ -2252,6 +2300,17 @@ static int __bond_release_one(struct net_device *bond_dev,/* recompute stats just before removing the slave */bond_get_stats(bond->dev,&bond->bond_stats);+if(bond->xdp_prog){+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=NULL,+.extack=NULL,+};+if(slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp))+slave_warn(bond_dev,slave_dev,"failed to unload XDP program\n");+}+bond_upper_dev_unlink(bond,slave);/* unregister rx_handler early so bond_handle_frame wouldn't be called*forthisslaveanymore.
@@ -4420,6 +4499,47 @@ static struct slave *bond_xmit_roundrobin_slave_get(struct bonding *bond,returnNULL;}+staticstructslave*bond_xdp_xmit_roundrobin_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structslave*slave;+intslave_cnt;+u32slave_id;+conststructethhdr*eth;+void*data=xdp->data;++if(data+sizeof(structethhdr)>xdp->data_end)+gotonon_igmp;++eth=(structethhdr*)data;+data+=sizeof(structethhdr);++/* See comment on IGMP in bond_xmit_roundrobin_slave_get() */+if(eth->h_proto==htons(ETH_P_IP)){+conststructiphdr*iph;++if(data+sizeof(structiphdr)>xdp->data_end)+gotonon_igmp;++iph=(structiphdr*)data;++if(iph->protocol==IPPROTO_IGMP){+slave=rcu_dereference(bond->curr_active_slave);+if(slave)+returnslave;+returnbond_get_slave_by_id(bond,0);+}+}++non_igmp:+slave_cnt=READ_ONCE(bond->slave_cnt);+if(likely(slave_cnt)){+slave_id=bond_rr_gen_slave_id(bond)%slave_cnt;+returnbond_get_slave_by_id(bond,slave_id);+}+returnNULL;+}+staticnetdev_tx_tbond_xmit_roundrobin(structsk_buff*skb,structnet_device*bond_dev){
@@ -4635,6 +4755,22 @@ static struct slave *bond_xmit_3ad_xor_slave_get(struct bonding *bond,returnslave;}+staticstructslave*bond_xdp_xmit_3ad_xor_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structbond_up_slave*slaves;+unsignedintcount;+u32hash;++hash=bond_xmit_hash_xdp(bond,xdp);+slaves=rcu_dereference(bond->usable_slaves);+count=slaves?READ_ONCE(slaves->count):0;+if(unlikely(!count))+returnNULL;++returnslaves->arr[hash%count];+}+/* Use this Xmit function for 3AD as well as XOR modes. The current*usableslavearrayisformedinthecontrolpath.Thexmitfunction*justcalculateshashandsendsthepacketout.
@@ -4919,6 +5055,174 @@ static netdev_tx_t bond_start_xmit(struct sk_buff *skb, struct net_device *dev)returnret;}+staticstructnet_device*+bond_xdp_get_xmit_slave(structnet_device*bond_dev,structxdp_buff*xdp)+{+structbonding*bond=netdev_priv(bond_dev);+structslave*slave;++/* Caller needs to hold rcu_read_lock() */++switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+slave=bond_xdp_xmit_roundrobin_slave_get(bond,xdp);+break;++caseBOND_MODE_ACTIVEBACKUP:+slave=bond_xmit_activebackup_slave_get(bond);+break;++caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+slave=bond_xdp_xmit_3ad_xor_slave_get(bond,xdp);+break;++default:+/* Should never happen. Mode guarded by bond_xdp_check() */+netdev_err(bond_dev,"Unknown bonding mode %d for xdp xmit\n",BOND_MODE(bond));+WARN_ON_ONCE(1);+returnNULL;+}++if(slave)+returnslave->dev;++returnNULL;+}++staticintbond_xdp_xmit(structnet_device*bond_dev,+intn,structxdp_frame**frames,u32flags)+{+intnxmit,err=-ENXIO;++rcu_read_lock();++for(nxmit=0;nxmit<n;nxmit++){+structxdp_frame*frame=frames[nxmit];+structxdp_frame*frames1[]={frame};+structnet_device*slave_dev;+structxdp_buffxdp;++xdp_convert_frame_to_buff(frame,&xdp);++slave_dev=bond_xdp_get_xmit_slave(bond_dev,&xdp);+if(!slave_dev){+err=-ENXIO;+break;+}++err=slave_dev->netdev_ops->ndo_xdp_xmit(slave_dev,1,frames1,flags);+if(err<1)+break;+}++rcu_read_unlock();++/* If error happened on the first frame then we can pass the error up, otherwise+*reportthenumberofframesthatwerexmitted.+*/+if(err<0)+return(nxmit==0?err:nxmit);++returnnxmit;+}++staticintbond_xdp_set(structnet_device*dev,structbpf_prog*prog,+structnetlink_ext_ack*extack)+{+structbonding*bond=netdev_priv(dev);+structlist_head*iter;+structslave*slave,*rollback_slave;+structbpf_prog*old_prog;+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=prog,+.extack=extack,+};+interr;++ASSERT_RTNL();++if(!bond_xdp_check(bond))+return-EOPNOTSUPP;++old_prog=bond->xdp_prog;+bond->xdp_prog=prog;++bond_for_each_slave(bond,slave,iter){+structnet_device*slave_dev=slave->dev;++if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+NL_SET_ERR_MSG(extack,"Slave device does not support XDP");+slave_err(dev,slave_dev,"Slave does not support XDP\n");+err=-EOPNOTSUPP;+gotoerr;+}++if(dev_xdp_prog_count(slave_dev)>0){+NL_SET_ERR_MSG(extack,+"Slave has XDP program loaded, please unload before enslaving");+slave_err(dev,slave_dev,+"Slave has XDP program loaded, please unload before enslaving\n");+err=-EOPNOTSUPP;+gotoerr;+}++err=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err<0){+/* ndo_bpf() sets extack error message */+slave_err(dev,slave_dev,"Error %d calling ndo_bpf\n",err);+gotoerr;+}+if(prog)+bpf_prog_inc(prog);+}++if(old_prog)+bpf_prog_put(old_prog);++if(prog)+static_branch_inc(&bpf_master_redirect_enabled_key);+else+static_branch_dec(&bpf_master_redirect_enabled_key);++return0;++err:+/* unwind the program changes */+bond->xdp_prog=old_prog;+xdp.prog=old_prog;+xdp.extack=NULL;/* do not overwrite original error */++bond_for_each_slave(bond,rollback_slave,iter){+structnet_device*slave_dev=rollback_slave->dev;+interr_unwind;++if(slave==rollback_slave)+break;++err_unwind=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err_unwind<0)+slave_err(dev,slave_dev,+"Error %d when unwinding XDP program change\n",err_unwind);+elseif(xdp.prog)+bpf_prog_inc(xdp.prog);+}+returnerr;+}++staticintbond_xdp(structnet_device*dev,structnetdev_bpf*xdp)+{+switch(xdp->command){+caseXDP_SETUP_PROG:+returnbond_xdp_set(dev,xdp->prog,xdp->extack);+default:+return-EINVAL;+}+}+staticu32bond_mode_bcast_speed(structslave*slave,u32speed){if(speed==0||speed==SPEED_UNKNOWN)
From: Jussi Maki <hidden> Date: 2021-08-04 12:46:26
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Signed-off-by: Jussi Maki <redacted>
---
.../selftests/bpf/prog_tests/xdp_bonding.c | 533 ++++++++++++++++++
1 file changed, 533 insertions(+)
@@ -0,0 +1,533 @@+// SPDX-License-Identifier: GPL-2.0++/**+*TestXDPbondingsupport+*+*Setsuptwobondedvethpairsbetweentwofreshnamespaces+*andverifiesthatXDP_TXprogramloadedonabonddevice+*arecorrectlyloadedontotheslavedevicesandXDP_TX'd+*packetsarebalancedusingbonding.+*/++#define _GNU_SOURCE+#include<sched.h>+#include<net/if.h>+#include<linux/if_link.h>+#include"test_progs.h"+#include"network_helpers.h"+#include<linux/if_bonding.h>+#include<linux/limits.h>+#include<linux/udp.h>++#include"xdp_dummy.skel.h"+#include"xdp_redirect_multi_kern.skel.h"+#include"xdp_tx.skel.h"++#define BOND1_MAC {0x00, 0x11, 0x22, 0x33, 0x44, 0x55}+#define BOND1_MAC_STR "00:11:22:33:44:55"+#define BOND2_MAC {0x00, 0x22, 0x33, 0x44, 0x55, 0x66}+#define BOND2_MAC_STR "00:22:33:44:55:66"+#define NPACKETS 100++staticintroot_netns_fd=-1;++staticvoidrestore_root_netns(void)+{+ASSERT_OK(setns(root_netns_fd,CLONE_NEWNET),"restore_root_netns");+}++intsetns_by_name(char*name)+{+intnsfd,err;+charnspath[PATH_MAX];++snprintf(nspath,sizeof(nspath),"%s/%s","/var/run/netns",name);+nsfd=open(nspath,O_RDONLY|O_CLOEXEC);+if(nsfd<0)+return-1;++err=setns(nsfd,CLONE_NEWNET);+close(nsfd);+returnerr;+}++staticintget_rx_packets(constchar*iface)+{+FILE*f;+charline[512];+intiface_len=strlen(iface);++f=fopen("/proc/net/dev","r");+if(!f)+return-1;++while(fgets(line,sizeof(line),f)){+char*p=line;++while(*p==' ')+p++;/* skip whitespace */+if(!strncmp(p,iface,iface_len)){+p+=iface_len;+if(*p++!=':')+continue;+while(*p==' ')+p++;/* skip whitespace */+while(*p&&*p!=' ')+p++;/* skip rx bytes */+while(*p==' ')+p++;/* skip whitespace */+fclose(f);+returnatoi(p);+}+}+fclose(f);+return-1;+}++#define MAX_BPF_LINKS 8++structskeletons{+structxdp_dummy*xdp_dummy;+structxdp_tx*xdp_tx;+structxdp_redirect_multi_kern*xdp_redirect_multi_kern;++intnlinks;+structbpf_link*links[MAX_BPF_LINKS];+};++staticintxdp_attach(structskeletons*skeletons,structbpf_program*prog,char*iface)+{+structbpf_link*link;+intifindex;++ifindex=if_nametoindex(iface);+if(!ASSERT_GT(ifindex,0,"get ifindex"))+return-1;++if(!ASSERT_LE(skeletons->nlinks,MAX_BPF_LINKS,"too many XDP programs attached"))+return-1;++link=bpf_program__attach_xdp(prog,ifindex);+if(!ASSERT_OK_PTR(link,"attach xdp program"))+return-1;++skeletons->links[skeletons->nlinks++]=link;+return0;+}++enum{+BOND_ONE_NO_ATTACH=0,+BOND_BOTH_AND_ATTACH,+};++staticconstchar*constmode_names[]={+[BOND_MODE_ROUNDROBIN]="balance-rr",+[BOND_MODE_ACTIVEBACKUP]="active-backup",+[BOND_MODE_XOR]="balance-xor",+[BOND_MODE_BROADCAST]="broadcast",+[BOND_MODE_8023AD]="802.3ad",+[BOND_MODE_TLB]="balance-tlb",+[BOND_MODE_ALB]="balance-alb",+};++staticconstchar*constxmit_policy_names[]={+[BOND_XMIT_POLICY_LAYER2]="layer2",+[BOND_XMIT_POLICY_LAYER34]="layer3+4",+[BOND_XMIT_POLICY_LAYER23]="layer2+3",+[BOND_XMIT_POLICY_ENCAP23]="encap2+3",+[BOND_XMIT_POLICY_ENCAP34]="encap3+4",+};++staticintbonding_setup(structskeletons*skeletons,intmode,intxmit_policy,+intbond_both_attach)+{+#define SYS(fmt, ...) \+({\+charcmd[1024];\+snprintf(cmd,sizeof(cmd),fmt,##__VA_ARGS__);\+if(!ASSERT_OK(system(cmd),cmd))\+return-1;\+})++SYS("ip netns add ns_dst");+SYS("ip link add veth1_1 type veth peer name veth2_1 netns ns_dst");+SYS("ip link add veth1_2 type veth peer name veth2_2 netns ns_dst");++SYS("ip link add bond1 type bond mode %s xmit_hash_policy %s",+mode_names[mode],xmit_policy_names[xmit_policy]);+SYS("ip link set bond1 up address "BOND1_MAC_STR" addrgenmode none");+SYS("ip -netns ns_dst link add bond2 type bond mode %s xmit_hash_policy %s",+mode_names[mode],xmit_policy_names[xmit_policy]);+SYS("ip -netns ns_dst link set bond2 up address "BOND2_MAC_STR" addrgenmode none");++SYS("ip link set veth1_1 master bond1");+if(bond_both_attach==BOND_BOTH_AND_ATTACH){+SYS("ip link set veth1_2 master bond1");+}else{+SYS("ip link set veth1_2 up addrgenmode none");++if(xdp_attach(skeletons,skeletons->xdp_dummy->progs.xdp_dummy_prog,"veth1_2"))+return-1;+}++SYS("ip -netns ns_dst link set veth2_1 master bond2");++if(bond_both_attach==BOND_BOTH_AND_ATTACH)+SYS("ip -netns ns_dst link set veth2_2 master bond2");+else+SYS("ip -netns ns_dst link set veth2_2 up addrgenmode none");++/* Load a dummy program on sending side as with veth peer needs to have a+*XDPprogramloadedaswell.+*/+if(xdp_attach(skeletons,skeletons->xdp_dummy->progs.xdp_dummy_prog,"bond1"))+return-1;++if(bond_both_attach==BOND_BOTH_AND_ATTACH){+if(!ASSERT_OK(setns_by_name("ns_dst"),"set netns to ns_dst"))+return-1;++if(xdp_attach(skeletons,skeletons->xdp_tx->progs.xdp_tx,"bond2"))+return-1;++restore_root_netns();+}++return0;++#undef SYS+}++staticvoidbonding_cleanup(structskeletons*skeletons)+{+restore_root_netns();+while(skeletons->nlinks){+skeletons->nlinks--;+bpf_link__detach(skeletons->links[skeletons->nlinks]);+}+ASSERT_OK(system("ip link delete bond1"),"delete bond1");+ASSERT_OK(system("ip link delete veth1_1"),"delete veth1_1");+ASSERT_OK(system("ip link delete veth1_2"),"delete veth1_2");+ASSERT_OK(system("ip netns delete ns_dst"),"delete ns_dst");+}++staticintsend_udp_packets(intvary_dst_ip)+{+structethhdreh={+.h_source=BOND1_MAC,+.h_dest=BOND2_MAC,+.h_proto=htons(ETH_P_IP),+};+uint8_tbuf[128]={};+structiphdr*iph=(structiphdr*)(buf+sizeof(eh));+structudphdr*uh=(structudphdr*)(buf+sizeof(eh)+sizeof(*iph));+inti,s=-1;+intifindex;++s=socket(AF_PACKET,SOCK_RAW,IPPROTO_RAW);+if(!ASSERT_GE(s,0,"socket"))+gotoerr;++ifindex=if_nametoindex("bond1");+if(!ASSERT_GT(ifindex,0,"get bond1 ifindex"))+gotoerr;++memcpy(buf,&eh,sizeof(eh));+iph->ihl=5;+iph->version=4;+iph->tos=16;+iph->id=1;+iph->ttl=64;+iph->protocol=IPPROTO_UDP;+iph->saddr=1;+iph->daddr=2;+iph->tot_len=htons(sizeof(buf)-ETH_HLEN);+iph->check=0;++for(i=1;i<=NPACKETS;i++){+intn;+structsockaddr_llsaddr_ll={+.sll_ifindex=ifindex,+.sll_halen=ETH_ALEN,+.sll_addr=BOND2_MAC,+};++/* vary the UDP destination port for even distribution with roundrobin/xor modes */+uh->dest++;++if(vary_dst_ip)+iph->daddr++;++n=sendto(s,buf,sizeof(buf),0,(structsockaddr*)&saddr_ll,sizeof(saddr_ll));+if(!ASSERT_EQ(n,sizeof(buf),"sendto"))+gotoerr;+}++return0;++err:+if(s>=0)+close(s);+return-1;+}++voidtest_xdp_bonding_with_mode(structskeletons*skeletons,char*name,intmode,intxmit_policy)+{+intbond1_rx;++if(!test__start_subtest(name))+return;++if(bonding_setup(skeletons,mode,xmit_policy,BOND_BOTH_AND_ATTACH))+gotoout;++if(send_udp_packets(xmit_policy!=BOND_XMIT_POLICY_LAYER34))+gotoout;++bond1_rx=get_rx_packets("bond1");+ASSERT_EQ(bond1_rx,NPACKETS,"expected more received packets");++switch(mode){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_XOR:{+intveth1_rx=get_rx_packets("veth1_1");+intveth2_rx=get_rx_packets("veth1_2");+intdiff=abs(veth1_rx-veth2_rx);++ASSERT_GE(veth1_rx+veth2_rx,NPACKETS,"expected more packets");++switch(xmit_policy){+caseBOND_XMIT_POLICY_LAYER2:+ASSERT_GE(diff,NPACKETS,+"expected packets on only one of the interfaces");+break;+caseBOND_XMIT_POLICY_LAYER23:+caseBOND_XMIT_POLICY_LAYER34:+ASSERT_LT(diff,NPACKETS/2,+"expected even distribution of packets");+break;+default:+PRINT_FAIL("Unimplemented xmit_policy=%d\n",xmit_policy);+break;+}+break;+}+caseBOND_MODE_ACTIVEBACKUP:{+intveth1_rx=get_rx_packets("veth1_1");+intveth2_rx=get_rx_packets("veth1_2");+intdiff=abs(veth1_rx-veth2_rx);++ASSERT_GE(diff,NPACKETS,+"expected packets on only one of the interfaces");+break;+}+default:+PRINT_FAIL("Unimplemented xmit_policy=%d\n",xmit_policy);+break;+}++out:+bonding_cleanup(skeletons);+}+++/* Test the broadcast redirection using xdp_redirect_map_multi_prog and adding+*alltheinterfacestoitandcheckingthatbroadcastingwon'tsendthepacket+*toneithertheingressbonddevice(bond2)oritsslave(veth2_1).+*/+voidtest_xdp_bonding_redirect_multi(structskeletons*skeletons)+{+staticconstchar*constifaces[]={"bond2","veth2_1","veth2_2"};+intveth1_1_rx,veth1_2_rx;+interr;++if(!test__start_subtest("xdp_bonding_redirect_multi"))+return;++if(bonding_setup(skeletons,BOND_MODE_ROUNDROBIN,BOND_XMIT_POLICY_LAYER23,+BOND_ONE_NO_ATTACH))+gotoout;+++if(!ASSERT_OK(setns_by_name("ns_dst"),"could not set netns to ns_dst"))+gotoout;++/* populate the devmap with the relevant interfaces */+for(inti=0;i<ARRAY_SIZE(ifaces);i++){+intifindex=if_nametoindex(ifaces[i]);+intmap_fd=bpf_map__fd(skeletons->xdp_redirect_multi_kern->maps.map_all);++if(!ASSERT_GT(ifindex,0,"could not get interface index"))+gotoout;++err=bpf_map_update_elem(map_fd,&ifindex,&ifindex,0);+if(!ASSERT_OK(err,"add interface to map_all"))+gotoout;+}++if(xdp_attach(skeletons,+skeletons->xdp_redirect_multi_kern->progs.xdp_redirect_map_multi_prog,+"bond2"))+gotoout;++restore_root_netns();++if(send_udp_packets(BOND_MODE_ROUNDROBIN))+gotoout;++veth1_1_rx=get_rx_packets("veth1_1");+veth1_2_rx=get_rx_packets("veth1_2");++ASSERT_EQ(veth1_1_rx,0,"expected no packets on veth1_1");+ASSERT_GE(veth1_2_rx,NPACKETS,"expected packets on veth1_2");++out:+restore_root_netns();+bonding_cleanup(skeletons);+}++/* Test that XDP programs cannot be attached to both the bond master and slaves simultaneously */+voidtest_xdp_bonding_attach(structskeletons*skeletons)+{+structbpf_link*link=NULL;+structbpf_link*link2=NULL;+intveth,bond;+interr;++if(!test__start_subtest("xdp_bonding_attach"))+return;++if(!ASSERT_OK(system("ip link add veth type veth"),"add veth"))+gotoout;+if(!ASSERT_OK(system("ip link add bond type bond"),"add bond"))+gotoout;++veth=if_nametoindex("veth");+if(!ASSERT_GE(veth,0,"if_nametoindex veth"))+gotoout;+bond=if_nametoindex("bond");+if(!ASSERT_GE(bond,0,"if_nametoindex bond"))+gotoout;++/* enslaving with a XDP program loaded fails */+link=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,veth);+if(!ASSERT_OK_PTR(link,"attach program to veth"))+gotoout;++err=system("ip link set veth master bond");+if(!ASSERT_NEQ(err,0,"attaching slave with xdp program expected to fail"))+gotoout;++bpf_link__detach(link);+link=NULL;++err=system("ip link set veth master bond");+if(!ASSERT_OK(err,"set veth master"))+gotoout;++/* attaching to slave when master has no program is allowed */+link=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,veth);+if(!ASSERT_OK_PTR(link,"attach program to slave when enslaved"))+gotoout;++/* attaching to master not allowed when slave has program loaded */+link2=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,bond);+if(!ASSERT_ERR_PTR(link2,"attach program to master when slave has program"))+gotoout;++bpf_link__detach(link);+link=NULL;++/* attaching XDP program to master allowed when slave has no program */+link=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,bond);+if(!ASSERT_OK_PTR(link,"attach program to master"))+gotoout;++/* attaching to slave not allowed when master has program loaded */+link2=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,bond);+ASSERT_ERR_PTR(link2,"attach program to slave when master has program");++out:+if(link)+bpf_link__detach(link);+if(link2)+bpf_link__detach(link2);++system("ip link del veth");+system("ip link del bond");+}++staticintlibbpf_debug_print(enumlibbpf_print_levellevel,+constchar*format,va_listargs)+{+if(level!=LIBBPF_WARN)+vprintf(format,args);+return0;+}++structbond_test_case{+char*name;+intmode;+intxmit_policy;+};++staticstructbond_test_casebond_test_cases[]={+{"xdp_bonding_roundrobin",BOND_MODE_ROUNDROBIN,BOND_XMIT_POLICY_LAYER23,},+{"xdp_bonding_activebackup",BOND_MODE_ACTIVEBACKUP,BOND_XMIT_POLICY_LAYER23},++{"xdp_bonding_xor_layer2",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER2,},+{"xdp_bonding_xor_layer23",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER23,},+{"xdp_bonding_xor_layer34",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER34,},+};++voidtest_xdp_bonding(void)+{+libbpf_print_fn_told_print_fn;+structskeletonsskeletons={};+inti;++old_print_fn=libbpf_set_print(libbpf_debug_print);++root_netns_fd=open("/proc/self/ns/net",O_RDONLY);+if(!ASSERT_GE(root_netns_fd,0,"open /proc/self/ns/net"))+gotoout;++skeletons.xdp_dummy=xdp_dummy__open_and_load();+if(!ASSERT_OK_PTR(skeletons.xdp_dummy,"xdp_dummy__open_and_load"))+gotoout;++skeletons.xdp_tx=xdp_tx__open_and_load();+if(!ASSERT_OK_PTR(skeletons.xdp_tx,"xdp_tx__open_and_load"))+gotoout;++skeletons.xdp_redirect_multi_kern=xdp_redirect_multi_kern__open_and_load();+if(!ASSERT_OK_PTR(skeletons.xdp_redirect_multi_kern,+"xdp_redirect_multi_kern__open_and_load"))+gotoout;++test_xdp_bonding_attach(&skeletons);++for(i=0;i<ARRAY_SIZE(bond_test_cases);i++){+structbond_test_case*test_case=&bond_test_cases[i];++test_xdp_bonding_with_mode(+&skeletons,+test_case->name,+test_case->mode,+test_case->xmit_policy);+}++test_xdp_bonding_redirect_multi(&skeletons);++out:+if(skeletons.xdp_dummy)+xdp_dummy__destroy(skeletons.xdp_dummy);+if(skeletons.xdp_tx)+xdp_tx__destroy(skeletons.xdp_tx);+if(skeletons.xdp_redirect_multi_kern)+xdp_redirect_multi_kern__destroy(skeletons.xdp_redirect_multi_kern);++libbpf_set_print(old_print_fn);+if(root_netns_fd)+close(root_netns_fd);+}
From: Jussi Maki <hidden> Date: 2021-08-04 12:46:27
The program type cannot be deduced from 'tx' which causes an invalid
argument error when trying to load xdp_tx.o using the skeleton.
Rename the section name to "xdp/tx" so that libbpf can deduce the type.
Signed-off-by: Jussi Maki <redacted>
---
tools/testing/selftests/bpf/progs/xdp_tx.c | 2 +-
tools/testing/selftests/bpf/test_xdp_veth.sh | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
@@ -108,7 +108,7 @@ ip link set dev veth2 xdp pinned $BPF_DIR/progs/redirect_map_1 iplinksetdevveth3xdppinned$BPF_DIR/progs/redirect_map_2 ip-nns1linksetdevveth11xdpobjxdp_dummy.osecxdp_dummy-ip-nns2linksetdevveth22xdpobjxdp_tx.osectx+ip-nns2linksetdevveth22xdpobjxdp_tx.osecxdp/tx ip-nns3linksetdevveth33xdpobjxdp_dummy.osecxdp_dummytrapcleanupEXIT
On Wed, Aug 4, 2021 at 5:45 AM Jussi Maki [off-list ref] wrote:
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Signed-off-by: Jussi Maki <redacted>
---
.../selftests/bpf/prog_tests/xdp_bonding.c | 533 ++++++++++++++++++
1 file changed, 533 insertions(+)
[...]
+
+static int xdp_attach(struct skeletons *skeletons, struct bpf_program *prog, char *iface)
+{
+ struct bpf_link *link;
+ int ifindex;
+
+ ifindex = if_nametoindex(iface);
+ if (!ASSERT_GT(ifindex, 0, "get ifindex"))
+ return -1;
+
+ if (!ASSERT_LE(skeletons->nlinks, MAX_BPF_LINKS, "too many XDP programs attached"))
If it's already less or equal to MAX_BPF_LINKS, then you'll bump
nlinks below one more time and write beyond the array boundaries?
You want bpf_link__destroy, not bpf_link__detach (detach will leave
underlying BPF link FD open and ensure that bpf_link__destory() won't
do anything with it, just frees memory).
+/* Test the broadcast redirection using xdp_redirect_map_multi_prog and adding
+ * all the interfaces to it and checking that broadcasting won't send the packet
+ * to neither the ingress bond device (bond2) or its slave (veth2_1).
+ */
+void test_xdp_bonding_redirect_multi(struct skeletons *skeletons)
+{
+ static const char * const ifaces[] = {"bond2", "veth2_1", "veth2_2"};
+ int veth1_1_rx, veth1_2_rx;
+ int err;
+
+ if (!test__start_subtest("xdp_bonding_redirect_multi"))
+ return;
+
+ if (bonding_setup(skeletons, BOND_MODE_ROUNDROBIN, BOND_XMIT_POLICY_LAYER23,
+ BOND_ONE_NO_ATTACH))
+ goto out;
+
+
nit: another extra empty line, please check if there are more
+ if (!ASSERT_OK(setns_by_name("ns_dst"), "could not set netns to ns_dst"))
+ goto out;
+
[...]
+ /* enslaving with a XDP program loaded fails */
+ link = bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog, veth);
+ if (!ASSERT_OK_PTR(link, "attach program to veth"))
+ goto out;
+
+ err = system("ip link set veth master bond");
+ if (!ASSERT_NEQ(err, 0, "attaching slave with xdp program expected to fail"))
+ goto out;
+
+ bpf_link__detach(link);
same here and in few more places, you need destroy
+ link = NULL;
+
+ err = system("ip link set veth master bond");
+ if (!ASSERT_OK(err, "set veth master"))
+ goto out;
+
+ /* attaching to slave when master has no program is allowed */
+ link = bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog, veth);
+ if (!ASSERT_OK_PTR(link, "attach program to slave when enslaved"))
+ goto out;
+
+ /* attaching to master not allowed when slave has program loaded */
+ link2 = bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog, bond);
+ if (!ASSERT_ERR_PTR(link2, "attach program to master when slave has program"))
+ goto out;
+
+ bpf_link__detach(link);
+ link = NULL;
+
+ /* attaching XDP program to master allowed when slave has no program */
+ link = bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog, bond);
+ if (!ASSERT_OK_PTR(link, "attach program to master"))
+ goto out;
+
+ /* attaching to slave not allowed when master has program loaded */
+ link2 = bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog, bond);
+ ASSERT_ERR_PTR(link2, "attach program to slave when master has program");
+
+out:
+ if (link)
+ bpf_link__detach(link);
+ if (link2)
+ bpf_link__detach(link2);
bpf_link__destroy() handles NULLs just fine, you don't have to do extra checks
On Wed, Aug 4, 2021 at 5:45 AM Jussi Maki [off-list ref] wrote:
quoted hunk
The program type cannot be deduced from 'tx' which causes an invalid
argument error when trying to load xdp_tx.o using the skeleton.
Rename the section name to "xdp/tx" so that libbpf can deduce the type.
Signed-off-by: Jussi Maki <redacted>
---
tools/testing/selftests/bpf/progs/xdp_tx.c | 2 +-
tools/testing/selftests/bpf/test_xdp_veth.sh | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
@@ -108,7 +108,7 @@ ip link set dev veth2 xdp pinned $BPF_DIR/progs/redirect_map_1 iplinksetdevveth3xdppinned$BPF_DIR/progs/redirect_map_2 ip-nns1linksetdevveth11xdpobjxdp_dummy.osecxdp_dummy-ip-nns2linksetdevveth22xdpobjxdp_tx.osectx+ip-nns2linksetdevveth22xdpobjxdp_tx.osecxdp/tx ip-nns3linksetdevveth33xdpobjxdp_dummy.osecxdp_dummytrapcleanupEXIT--
From: Jussi Maki <hidden> Date: 2021-08-05 16:10:20
This patchset introduces XDP support to the bonding driver.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter. Fix for this
has been already merged into net-next. The statistics were collected
using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Patch 1 prepares bond_xmit_hash for hashing xdp_buff's.
Patch 2 adds hooks to implement redirection after bpf prog run.
Patch 3 implements the hooks in the bonding driver.
Patch 4 modifies devmap to properly handle EXCLUDE_INGRESS with a slave device.
Patch 5 fixes an issue related to recent cleanup of rcu_read_lock in XDP context.
Patch 6 fixes loading of xdp_tx.o by renaming section name.
Patch 7 adds tests.
v5->v6:
- Address Andrii's comments about the tests.
v4->v5:
- As pointed by Andrii, use the generated BPF skeletons rather than libbpf
directly.
- Renamed section name in progs/xdp_tx.c as the BPF skeleton wouldn't load it
otherwise due to unknown program type.
- Daniel Borkmann noted that to retain backwards compatibility and allow some
use cases we should allow attaching XDP programs to a slave device when the
master does not have a program loaded. Modified the logic to allow this and
added tests for the different combinations of attaching a program.
v3->v4:
- Add back the test suite, while removing the vmtest.sh modifications to kernel
config new that CONFIG_BONDING=y is set. Discussed with Magnus Karlsson that
it makes sense right now to not reuse the code from xdpceiver.c for testing
XDP bonding.
v2->v3:
- Address Jay's comment to properly exclude upper devices with EXCLUDE_INGRESS
when there are deeper nesting involved. Now all upper devices are excluded.
- Refuse to enslave devices that already have XDP programs loaded and refuse to
load XDP programs to slave devices. Earlier one could have a XDP program loaded
and after enslaving and loading another program onto the bond device the xdp_state
of the enslaved device would be pointing at an old program.
- Adapt netdev_lower_get_next_private_rcu so it can be called in the XDP context.
v1->v2:
- Split up into smaller easier to review patches and address cosmetic
review comments.
- Drop the INDIRECT_CALL optimization as it showed little improvement in tests.
- Drop the rr_tx_counter patch as that has already been merged into net-next.
- Separate the test suite into another patch set. This will follow later once a
patch set from Magnus Karlsson is merged and provides test utilities that can
be reused for XDP bonding tests. v2 contains no major functional changes and
was tested with the test suite included in v1.
(https://lore.kernel.org/bpf/202106221509.kwNvAAZg-lkp@intel.com/T/#m464146d47299125d5868a08affd6d6ce526dfad1)
---
From: Jussi Maki <hidden> Date: 2021-08-05 16:10:22
In preparation for adding XDP support to the bonding driver
refactor the packet hashing functions to be able to work with
any linear data buffer without an skb.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 147 +++++++++++++++++++-------------
1 file changed, 90 insertions(+), 57 deletions(-)
@@ -3611,55 +3611,80 @@ static struct notifier_block bond_netdev_notifier = {/*---------------------------- Hashing Policies -----------------------------*/+/* Helper to access data in a packet, with or without a backing skb.+*Ifskbisgiventhedataislinearizedifnecessaryviapskb_may_pull.+*/+staticinlineconstvoid*bond_pull_data(structsk_buff*skb,+constvoid*data,inthlen,intn)+{+if(likely(n<=hlen))+returndata;+elseif(skb&&likely(pskb_may_pull(skb,n)))+returnskb->head;++returnNULL;+}+/* L2 hash helper */-staticinlineu32bond_eth_hash(structsk_buff*skb)+staticinlineu32bond_eth_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*ep,hdr_tmp;+structethhdr*ep;-ep=skb_header_pointer(skb,0,sizeof(hdr_tmp),&hdr_tmp);-if(ep)-returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;-return0;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;++ep=(structethhdr*)(data+mhoff);+returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;}-staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,-int*noff,int*proto,booll34)+staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,constvoid*data,+inthlen,__be16l2_proto,int*nhoff,int*ip_proto,booll34){conststructipv6hdr*iph6;conststructiphdr*iph;-if(skb->protocol==htons(ETH_P_IP)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph))))+if(l2_proto==htons(ETH_P_IP)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph));+if(!data)returnfalse;-iph=(conststructiphdr*)(skb->data+*noff);++iph=(conststructiphdr*)(data+*nhoff);iph_to_flow_copy_v4addrs(fk,iph);-*noff+=iph->ihl<<2;+*nhoff+=iph->ihl<<2;if(!ip_is_fragment(iph))-*proto=iph->protocol;-}elseif(skb->protocol==htons(ETH_P_IPV6)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph6))))+*ip_proto=iph->protocol;+}elseif(l2_proto==htons(ETH_P_IPV6)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph6));+if(!data)returnfalse;-iph6=(conststructipv6hdr*)(skb->data+*noff);++iph6=(conststructipv6hdr*)(data+*nhoff);iph_to_flow_copy_v6addrs(fk,iph6);-*noff+=sizeof(*iph6);-*proto=iph6->nexthdr;+*nhoff+=sizeof(*iph6);+*ip_proto=iph6->nexthdr;}else{returnfalse;}-if(l34&&*proto>=0)-fk->ports.ports=skb_flow_get_ports(skb,*noff,*proto);+if(l34&&*ip_proto>=0)+fk->ports.ports=__skb_flow_get_ports(skb,*nhoff,*ip_proto,data,hlen);returntrue;}-staticu32bond_vlan_srcmac_hash(structsk_buff*skb)+staticu32bond_vlan_srcmac_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*mac_hdr=(structethhdr*)skb_mac_header(skb);+structethhdr*mac_hdr;u32srcmac_vendor=0,srcmac_dev=0;u16vlan;inti;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;+mac_hdr=(structethhdr*)(data+mhoff);+for(i=0;i<3;i++)srcmac_vendor=(srcmac_vendor<<8)|mac_hdr->h_source[i];
@@ -3675,26 +3700,25 @@ static u32 bond_vlan_srcmac_hash(struct sk_buff *skb)}/* Extract the appropriate headers based on bond's xmit policy */-staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,-structflow_keys*fk)+staticboolbond_flow_dissect(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intnhoff,inthlen,structflow_keys*fk){booll34=bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER34;-intnoff,proto=-1;+intip_proto=-1;switch(bond->params.xmit_policy){caseBOND_XMIT_POLICY_ENCAP23:caseBOND_XMIT_POLICY_ENCAP34:memset(fk,0,sizeof(*fk));return__skb_flow_dissect(NULL,skb,&flow_keys_bonding,-fk,NULL,0,0,0,0);+fk,data,l2_proto,nhoff,hlen,0);default:break;}fk->ports.ports=0;memset(&fk->icmp,0,sizeof(fk->icmp));-noff=skb_network_offset(skb);-if(!bond_flow_ip(skb,fk,&noff,&proto,l34))+if(!bond_flow_ip(skb,fk,data,hlen,l2_proto,&nhoff,&ip_proto,l34))returnfalse;/* ICMP error packets contains at least 8 bytes of the header
@@ -3733,33 +3755,26 @@ static u32 bond_ip_hash(u32 hash, struct flow_keys *flow)returnhash>>1;}-/**-*bond_xmit_hash-generateahashvaluebasedonthexmitpolicy-*@bond:bondingdevice-*@skb:buffertouseforheaders-*-*Thisfunctionwillextractthenecessaryheadersfromtheskbbufferanduse-*themtogenerateahashbasedonthexmit_policysetinthebondingdevice+/* Generate hash based on xmit policy. If @skb is given it is used to linearize+*thedataasrequired,butthisfunctioncanbeusedwithoutitifthedatais+*knowntobelinear(e.g.withxdp_buff).*/-u32bond_xmit_hash(structbonding*bond,structsk_buff*skb)+staticu32__bond_xmit_hash(structbonding*bond,structsk_buff*skb,constvoid*data,+__be16l2_proto,intmhoff,intnhoff,inthlen){structflow_keysflow;u32hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP34&&-skb->l4_hash)-returnskb->hash;-if(bond->params.xmit_policy==BOND_XMIT_POLICY_VLAN_SRCMAC)-returnbond_vlan_srcmac_hash(skb);+returnbond_vlan_srcmac_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER2||-!bond_flow_dissect(bond,skb,&flow))-returnbond_eth_hash(skb);+!bond_flow_dissect(bond,skb,data,l2_proto,nhoff,hlen,&flow))+returnbond_eth_hash(skb,data,mhoff,hlen);if(bond->params.xmit_policy==BOND_XMIT_POLICY_LAYER23||bond->params.xmit_policy==BOND_XMIT_POLICY_ENCAP23){-hash=bond_eth_hash(skb);+hash=bond_eth_hash(skb,data,mhoff,hlen);}else{if(flow.icmp.id)memcpy(&hash,&flow.icmp,sizeof(hash));
From: Jussi Maki <hidden> Date: 2021-08-05 16:10:25
This adds the ndo_xdp_get_xmit_slave hook for transforming XDP_TX
into XDP_REDIRECT after BPF program run when the ingress device
is a bond slave.
The dev_xdp_prog_count is exposed so that slave devices can be checked
for loaded XDP programs in order to avoid the situation where both
bond master and slave have programs loaded according to xdp_state.
Signed-off-by: Jussi Maki <redacted>
---
include/linux/filter.h | 13 ++++++++++++-
include/linux/netdevice.h | 6 ++++++
net/core/dev.c | 13 ++++++++++++-
net/core/filter.c | 25 +++++++++++++++++++++++++
4 files changed, 55 insertions(+), 2 deletions(-)
@@ -9494,6 +9497,14 @@ static int dev_xdp_attach(struct net_device *dev, struct netlink_ext_ack *extackreturn-EBUSY;}+/* don't allow if an upper device already has a program */+netdev_for_each_upper_dev_rcu(dev,upper,iter){+if(dev_xdp_prog_count(upper)>0){+NL_SET_ERR_MSG(extack,"Cannot attach when an upper device already has a program");+return-EEXIST;+}+}+cur_prog=dev_xdp_prog(dev,mode);/* can't replace attached prog with link */if(link&&cur_prog){
@@ -3950,6 +3950,31 @@ void bpf_clear_redirect_map(struct bpf_map *map)}}+DEFINE_STATIC_KEY_FALSE(bpf_master_redirect_enabled_key);+EXPORT_SYMBOL_GPL(bpf_master_redirect_enabled_key);++u32xdp_master_redirect(structxdp_buff*xdp)+{+structnet_device*master,*slave;+structbpf_redirect_info*ri=this_cpu_ptr(&bpf_redirect_info);++master=netdev_master_upper_dev_get_rcu(xdp->rxq->dev);+slave=master->netdev_ops->ndo_xdp_get_xmit_slave(master,xdp);+if(slave&&slave!=xdp->rxq->dev){+/* The target device is different from the receiving device, so+*redirectittothenewdevice.+*UsingXDP_REDIRECTgetsthecorrectbehaviourfromXDPenabled+*driverstounmapthepacketfromtheirrxring.+*/+ri->tgt_index=slave->ifindex;+ri->map_id=INT_MAX;+ri->map_type=BPF_MAP_TYPE_UNSPEC;+returnXDP_REDIRECT;+}+returnXDP_TX;+}+EXPORT_SYMBOL_GPL(xdp_master_redirect);+intxdp_do_redirect(structnet_device*dev,structxdp_buff*xdp,structbpf_prog*xdp_prog){
From: Jussi Maki <hidden> Date: 2021-08-05 16:10:26
XDP is implemented in the bonding driver by transparently delegating
the XDP program loading, removal and xmit operations to the bonding
slave devices. The overall goal of this work is that XDP programs
can be attached to a bond device *without* any further changes (or
awareness) necessary to the program itself, meaning the same XDP
program can be attached to a native device but also a bonding device.
Semantics of XDP_TX when attached to a bond are equivalent in such
setting to the case when a tc/BPF program would be attached to the
bond, meaning transmitting the packet out of the bond itself using one
of the bond's configured xmit methods to select a slave device (rather
than XDP_TX on the slave itself). Handling of XDP_TX to transmit
using the configured bonding mechanism is therefore implemented by
rewriting the BPF program return value in bpf_prog_run_xdp. To avoid
performance impact this check is guarded by a static key, which is
incremented when a XDP program is loaded onto a bond device. This
approach was chosen to avoid changes to drivers implementing XDP. If
the slave device does not match the receive device, then XDP_REDIRECT
is transparently used to perform the redirection in order to have
the network driver release the packet from its RX ring. The bonding
driver hashing functions have been refactored to allow reuse with
xdp_buff's to avoid code duplication.
The motivation for this change is to enable use of bonding (and
802.3ad) in hairpinning L4 load-balancers such as [1] implemented with
XDP and also to transparently support bond devices for projects that
use XDP given most modern NICs have dual port adapters. An alternative
to this approach would be to implement 802.3ad in user-space and
implement the bonding load-balancing in the XDP program itself, but
is rather a cumbersome endeavor in terms of slave device management
(e.g. by watching netlink) and requires separate programs for native
vs bond cases for the orchestrator. A native in-kernel implementation
overcomes these issues and provides more flexibility.
Below are benchmark results done on two machines with 100Gbit
Intel E810 (ice) NIC and with 32-core 3970X on sending machine, and
16-core 3950X on receiving machine. 64 byte packets were sent with
pktgen-dpdk at full rate. Two issues [2, 3] were identified with the
ice driver, so the tests were performed with iommu=off and patch [2]
applied. Additionally the bonding round robin algorithm was modified
to use per-cpu tx counters as high CPU load (50% vs 10%) and high rate
of cache misses were caused by the shared rr_tx_counter (see patch
2/3). The statistics were collected using "sar -n dev -u 1 10".
-----------------------| CPU |--| rxpck/s |--| txpck/s |----
without patch (1 dev):
XDP_DROP: 3.15% 48.6Mpps
XDP_TX: 3.12% 18.3Mpps 18.3Mpps
XDP_DROP (RSS): 9.47% 116.5Mpps
XDP_TX (RSS): 9.67% 25.3Mpps 24.2Mpps
-----------------------
with patch, bond (1 dev):
XDP_DROP: 3.14% 46.7Mpps
XDP_TX: 3.15% 13.9Mpps 13.9Mpps
XDP_DROP (RSS): 10.33% 117.2Mpps
XDP_TX (RSS): 10.64% 25.1Mpps 24.0Mpps
-----------------------
with patch, bond (2 devs):
XDP_DROP: 6.27% 92.7Mpps
XDP_TX: 6.26% 17.6Mpps 17.5Mpps
XDP_DROP (RSS): 11.38% 117.2Mpps
XDP_TX (RSS): 14.30% 28.7Mpps 27.4Mpps
--------------------------------------------------------------
RSS: Receive Side Scaling, e.g. the packets were sent to a range of
destination IPs.
[1]: https://cilium.io/blog/2021/05/20/cilium-110#standalonelb
[2]: https://lore.kernel.org/bpf/20210601113236.42651-1-maciej.fijalkowski@intel.com/T/#t
[3]: https://lore.kernel.org/bpf/CAHn8xckNXci+X_Eb2WMv4uVYjO2331UWB2JLtXr_58z0Av8+8A@mail.gmail.com/
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 309 +++++++++++++++++++++++++++++++-
include/net/bonding.h | 1 +
2 files changed, 309 insertions(+), 1 deletion(-)
@@ -317,6 +317,19 @@ bool bond_sk_check(struct bonding *bond)}}+staticboolbond_xdp_check(structbonding*bond)+{+switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_ACTIVEBACKUP:+caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+returntrue;+default:+returnfalse;+}+}+/*---------------------------------- VLAN -----------------------------------*//* In the following 2 functions, bond_vlan_rx_add_vid and bond_vlan_rx_kill_vid,
@@ -2133,6 +2146,41 @@ int bond_enslave(struct net_device *bond_dev, struct net_device *slave_dev,bond_update_slave_arr(bond,NULL);+if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+if(bond->xdp_prog){+NL_SET_ERR_MSG(extack,"Slave does not support XDP");+slave_err(bond_dev,slave_dev,"Slave does not support XDP\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}+}else{+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=bond->xdp_prog,+.extack=extack,+};++if(dev_xdp_prog_count(slave_dev)>0){+NL_SET_ERR_MSG(extack,+"Slave has XDP program loaded, please unload before enslaving");+slave_err(bond_dev,slave_dev,+"Slave has XDP program loaded, please unload before enslaving\n");+res=-EOPNOTSUPP;+gotoerr_sysfs_del;+}++res=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(res<0){+/* ndo_bpf() sets extack error message */+slave_dbg(bond_dev,slave_dev,"Error %d calling ndo_bpf\n",res);+gotoerr_sysfs_del;+}+if(bond->xdp_prog)+bpf_prog_inc(bond->xdp_prog);+}+slave_info(bond_dev,slave_dev,"Enslaving as %s interface with %s link\n",bond_is_active_slave(new_slave)?"an active":"a backup",new_slave->link!=BOND_LINK_DOWN?"an up":"a down");
@@ -2252,6 +2300,17 @@ static int __bond_release_one(struct net_device *bond_dev,/* recompute stats just before removing the slave */bond_get_stats(bond->dev,&bond->bond_stats);+if(bond->xdp_prog){+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=NULL,+.extack=NULL,+};+if(slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp))+slave_warn(bond_dev,slave_dev,"failed to unload XDP program\n");+}+bond_upper_dev_unlink(bond,slave);/* unregister rx_handler early so bond_handle_frame wouldn't be called*forthisslaveanymore.
@@ -4420,6 +4499,47 @@ static struct slave *bond_xmit_roundrobin_slave_get(struct bonding *bond,returnNULL;}+staticstructslave*bond_xdp_xmit_roundrobin_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structslave*slave;+intslave_cnt;+u32slave_id;+conststructethhdr*eth;+void*data=xdp->data;++if(data+sizeof(structethhdr)>xdp->data_end)+gotonon_igmp;++eth=(structethhdr*)data;+data+=sizeof(structethhdr);++/* See comment on IGMP in bond_xmit_roundrobin_slave_get() */+if(eth->h_proto==htons(ETH_P_IP)){+conststructiphdr*iph;++if(data+sizeof(structiphdr)>xdp->data_end)+gotonon_igmp;++iph=(structiphdr*)data;++if(iph->protocol==IPPROTO_IGMP){+slave=rcu_dereference(bond->curr_active_slave);+if(slave)+returnslave;+returnbond_get_slave_by_id(bond,0);+}+}++non_igmp:+slave_cnt=READ_ONCE(bond->slave_cnt);+if(likely(slave_cnt)){+slave_id=bond_rr_gen_slave_id(bond)%slave_cnt;+returnbond_get_slave_by_id(bond,slave_id);+}+returnNULL;+}+staticnetdev_tx_tbond_xmit_roundrobin(structsk_buff*skb,structnet_device*bond_dev){
@@ -4635,6 +4755,22 @@ static struct slave *bond_xmit_3ad_xor_slave_get(struct bonding *bond,returnslave;}+staticstructslave*bond_xdp_xmit_3ad_xor_slave_get(structbonding*bond,+structxdp_buff*xdp)+{+structbond_up_slave*slaves;+unsignedintcount;+u32hash;++hash=bond_xmit_hash_xdp(bond,xdp);+slaves=rcu_dereference(bond->usable_slaves);+count=slaves?READ_ONCE(slaves->count):0;+if(unlikely(!count))+returnNULL;++returnslaves->arr[hash%count];+}+/* Use this Xmit function for 3AD as well as XOR modes. The current*usableslavearrayisformedinthecontrolpath.Thexmitfunction*justcalculateshashandsendsthepacketout.
@@ -4919,6 +5055,174 @@ static netdev_tx_t bond_start_xmit(struct sk_buff *skb, struct net_device *dev)returnret;}+staticstructnet_device*+bond_xdp_get_xmit_slave(structnet_device*bond_dev,structxdp_buff*xdp)+{+structbonding*bond=netdev_priv(bond_dev);+structslave*slave;++/* Caller needs to hold rcu_read_lock() */++switch(BOND_MODE(bond)){+caseBOND_MODE_ROUNDROBIN:+slave=bond_xdp_xmit_roundrobin_slave_get(bond,xdp);+break;++caseBOND_MODE_ACTIVEBACKUP:+slave=bond_xmit_activebackup_slave_get(bond);+break;++caseBOND_MODE_8023AD:+caseBOND_MODE_XOR:+slave=bond_xdp_xmit_3ad_xor_slave_get(bond,xdp);+break;++default:+/* Should never happen. Mode guarded by bond_xdp_check() */+netdev_err(bond_dev,"Unknown bonding mode %d for xdp xmit\n",BOND_MODE(bond));+WARN_ON_ONCE(1);+returnNULL;+}++if(slave)+returnslave->dev;++returnNULL;+}++staticintbond_xdp_xmit(structnet_device*bond_dev,+intn,structxdp_frame**frames,u32flags)+{+intnxmit,err=-ENXIO;++rcu_read_lock();++for(nxmit=0;nxmit<n;nxmit++){+structxdp_frame*frame=frames[nxmit];+structxdp_frame*frames1[]={frame};+structnet_device*slave_dev;+structxdp_buffxdp;++xdp_convert_frame_to_buff(frame,&xdp);++slave_dev=bond_xdp_get_xmit_slave(bond_dev,&xdp);+if(!slave_dev){+err=-ENXIO;+break;+}++err=slave_dev->netdev_ops->ndo_xdp_xmit(slave_dev,1,frames1,flags);+if(err<1)+break;+}++rcu_read_unlock();++/* If error happened on the first frame then we can pass the error up, otherwise+*reportthenumberofframesthatwerexmitted.+*/+if(err<0)+return(nxmit==0?err:nxmit);++returnnxmit;+}++staticintbond_xdp_set(structnet_device*dev,structbpf_prog*prog,+structnetlink_ext_ack*extack)+{+structbonding*bond=netdev_priv(dev);+structlist_head*iter;+structslave*slave,*rollback_slave;+structbpf_prog*old_prog;+structnetdev_bpfxdp={+.command=XDP_SETUP_PROG,+.flags=0,+.prog=prog,+.extack=extack,+};+interr;++ASSERT_RTNL();++if(!bond_xdp_check(bond))+return-EOPNOTSUPP;++old_prog=bond->xdp_prog;+bond->xdp_prog=prog;++bond_for_each_slave(bond,slave,iter){+structnet_device*slave_dev=slave->dev;++if(!slave_dev->netdev_ops->ndo_bpf||+!slave_dev->netdev_ops->ndo_xdp_xmit){+NL_SET_ERR_MSG(extack,"Slave device does not support XDP");+slave_err(dev,slave_dev,"Slave does not support XDP\n");+err=-EOPNOTSUPP;+gotoerr;+}++if(dev_xdp_prog_count(slave_dev)>0){+NL_SET_ERR_MSG(extack,+"Slave has XDP program loaded, please unload before enslaving");+slave_err(dev,slave_dev,+"Slave has XDP program loaded, please unload before enslaving\n");+err=-EOPNOTSUPP;+gotoerr;+}++err=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err<0){+/* ndo_bpf() sets extack error message */+slave_err(dev,slave_dev,"Error %d calling ndo_bpf\n",err);+gotoerr;+}+if(prog)+bpf_prog_inc(prog);+}++if(old_prog)+bpf_prog_put(old_prog);++if(prog)+static_branch_inc(&bpf_master_redirect_enabled_key);+else+static_branch_dec(&bpf_master_redirect_enabled_key);++return0;++err:+/* unwind the program changes */+bond->xdp_prog=old_prog;+xdp.prog=old_prog;+xdp.extack=NULL;/* do not overwrite original error */++bond_for_each_slave(bond,rollback_slave,iter){+structnet_device*slave_dev=rollback_slave->dev;+interr_unwind;++if(slave==rollback_slave)+break;++err_unwind=slave_dev->netdev_ops->ndo_bpf(slave_dev,&xdp);+if(err_unwind<0)+slave_err(dev,slave_dev,+"Error %d when unwinding XDP program change\n",err_unwind);+elseif(xdp.prog)+bpf_prog_inc(xdp.prog);+}+returnerr;+}++staticintbond_xdp(structnet_device*dev,structnetdev_bpf*xdp)+{+switch(xdp->command){+caseXDP_SETUP_PROG:+returnbond_xdp_set(dev,xdp->prog,xdp->extack);+default:+return-EINVAL;+}+}+staticu32bond_mode_bcast_speed(structslave*slave,u32speed){if(speed==0||speed==SPEED_UNKNOWN)
From: Jussi Maki <hidden> Date: 2021-08-05 16:10:27
If the ingress device is bond slave, do not broadcast back
through it or the bond master.
Signed-off-by: Jussi Maki <redacted>
---
kernel/bpf/devmap.c | 69 +++++++++++++++++++++++++++++++++++++++------
1 file changed, 60 insertions(+), 9 deletions(-)
@@ -562,17 +561,48 @@ static int dev_map_enqueue_clone(struct bpf_dtab_netdev *obj,return0;}+staticinlineboolis_ifindex_excluded(int*excluded,intnum_excluded,intifindex)+{+while(num_excluded--){+if(ifindex==excluded[num_excluded])+returntrue;+}+returnfalse;+}++/* Get ifindex of each upper device. 'indexes' must be able to hold at+*leastMAX_NEST_DEVelements.+*Returnsthenumberofifindexesadded.+*/+staticintget_upper_ifindexes(structnet_device*dev,int*indexes)+{+structnet_device*upper;+structlist_head*iter;+intn=0;++netdev_for_each_upper_dev_rcu(dev,upper,iter){+indexes[n++]=upper->ifindex;+}+returnn;+}+intdev_map_enqueue_multi(structxdp_buff*xdp,structnet_device*dev_rx,structbpf_map*map,boolexclude_ingress){structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);-intexclude_ifindex=exclude_ingress?dev_rx->ifindex:0;structbpf_dtab_netdev*dst,*last_dst=NULL;+intexcluded_devices[1+MAX_NEST_DEV];structhlist_head*head;structxdp_frame*xdpf;+intnum_excluded=0;unsignedinti;interr;+if(exclude_ingress){+num_excluded=get_upper_ifindexes(dev_rx,excluded_devices);+excluded_devices[num_excluded++]=dev_rx->ifindex;+}+xdpf=xdp_convert_buff_to_frame(xdp);if(unlikely(!xdpf))return-EOVERFLOW;
@@ -581,7 +611,10 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,for(i=0;i<map->max_entries;i++){dst=rcu_dereference_check(dtab->netdev_map[i],rcu_read_lock_bh_held());-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp))+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -601,7 +634,11 @@ int dev_map_enqueue_multi(struct xdp_buff *xdp, struct net_device *dev_rx,head=dev_map_index_hash(dtab,i);hlist_for_each_entry_rcu(dst,head,index_hlist,lockdep_is_held(&dtab->index_lock)){-if(!is_valid_dst(dst,xdp,exclude_ifindex))+if(!is_valid_dst(dst,xdp))+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,+dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -675,18 +712,27 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,boolexclude_ingress){structbpf_dtab*dtab=container_of(map,structbpf_dtab,map);-intexclude_ifindex=exclude_ingress?dev->ifindex:0;structbpf_dtab_netdev*dst,*last_dst=NULL;+intexcluded_devices[1+MAX_NEST_DEV];structhlist_head*head;structhlist_node*next;+intnum_excluded=0;unsignedinti;interr;+if(exclude_ingress){+num_excluded=get_upper_ifindexes(dev,excluded_devices);+excluded_devices[num_excluded++]=dev->ifindex;+}+if(map->map_type==BPF_MAP_TYPE_DEVMAP){for(i=0;i<map->max_entries;i++){dst=rcu_dereference_check(dtab->netdev_map[i],rcu_read_lock_bh_held());-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst)+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
@@ -700,12 +746,17 @@ int dev_map_redirect_multi(struct net_device *dev, struct sk_buff *skb,returnerr;last_dst=dst;+}}else{/* BPF_MAP_TYPE_DEVMAP_HASH */for(i=0;i<dtab->n_buckets;i++){head=dev_map_index_hash(dtab,i);hlist_for_each_entry_safe(dst,next,head,index_hlist){-if(!dst||dst->dev->ifindex==exclude_ifindex)+if(!dst)+continue;++if(is_ifindex_excluded(excluded_devices,num_excluded,+dst->dev->ifindex))continue;/* we only need n-1 clones; last_dst enqueued below */
From: Jussi Maki <hidden> Date: 2021-08-05 16:10:36
The program type cannot be deduced from 'tx' which causes an invalid
argument error when trying to load xdp_tx.o using the skeleton.
Rename the section name to "xdp" so that libbpf can deduce the type.
Signed-off-by: Jussi Maki <redacted>
---
tools/testing/selftests/bpf/progs/xdp_tx.c | 2 +-
tools/testing/selftests/bpf/test_xdp_veth.sh | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
@@ -108,7 +108,7 @@ ip link set dev veth2 xdp pinned $BPF_DIR/progs/redirect_map_1 iplinksetdevveth3xdppinned$BPF_DIR/progs/redirect_map_2 ip-nns1linksetdevveth11xdpobjxdp_dummy.osecxdp_dummy-ip-nns2linksetdevveth22xdpobjxdp_tx.osectx+ip-nns2linksetdevveth22xdpobjxdp_tx.osecxdp ip-nns3linksetdevveth33xdpobjxdp_dummy.osecxdp_dummytrapcleanupEXIT
From: Jussi Maki <hidden> Date: 2021-08-05 16:10:43
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Signed-off-by: Jussi Maki <redacted>
---
.../selftests/bpf/prog_tests/xdp_bonding.c | 520 ++++++++++++++++++
1 file changed, 520 insertions(+)
@@ -0,0 +1,520 @@+// SPDX-License-Identifier: GPL-2.0++/**+*TestXDPbondingsupport+*+*Setsuptwobondedvethpairsbetweentwofreshnamespaces+*andverifiesthatXDP_TXprogramloadedonabonddevice+*arecorrectlyloadedontotheslavedevicesandXDP_TX'd+*packetsarebalancedusingbonding.+*/++#define _GNU_SOURCE+#include<sched.h>+#include<net/if.h>+#include<linux/if_link.h>+#include"test_progs.h"+#include"network_helpers.h"+#include<linux/if_bonding.h>+#include<linux/limits.h>+#include<linux/udp.h>++#include"xdp_dummy.skel.h"+#include"xdp_redirect_multi_kern.skel.h"+#include"xdp_tx.skel.h"++#define BOND1_MAC {0x00, 0x11, 0x22, 0x33, 0x44, 0x55}+#define BOND1_MAC_STR "00:11:22:33:44:55"+#define BOND2_MAC {0x00, 0x22, 0x33, 0x44, 0x55, 0x66}+#define BOND2_MAC_STR "00:22:33:44:55:66"+#define NPACKETS 100++staticintroot_netns_fd=-1;++staticvoidrestore_root_netns(void)+{+ASSERT_OK(setns(root_netns_fd,CLONE_NEWNET),"restore_root_netns");+}++staticintsetns_by_name(char*name)+{+intnsfd,err;+charnspath[PATH_MAX];++snprintf(nspath,sizeof(nspath),"%s/%s","/var/run/netns",name);+nsfd=open(nspath,O_RDONLY|O_CLOEXEC);+if(nsfd<0)+return-1;++err=setns(nsfd,CLONE_NEWNET);+close(nsfd);+returnerr;+}++staticintget_rx_packets(constchar*iface)+{+FILE*f;+charline[512];+intiface_len=strlen(iface);++f=fopen("/proc/net/dev","r");+if(!f)+return-1;++while(fgets(line,sizeof(line),f)){+char*p=line;++while(*p==' ')+p++;/* skip whitespace */+if(!strncmp(p,iface,iface_len)){+p+=iface_len;+if(*p++!=':')+continue;+while(*p==' ')+p++;/* skip whitespace */+while(*p&&*p!=' ')+p++;/* skip rx bytes */+while(*p==' ')+p++;/* skip whitespace */+fclose(f);+returnatoi(p);+}+}+fclose(f);+return-1;+}++#define MAX_BPF_LINKS 8++structskeletons{+structxdp_dummy*xdp_dummy;+structxdp_tx*xdp_tx;+structxdp_redirect_multi_kern*xdp_redirect_multi_kern;++intnlinks;+structbpf_link*links[MAX_BPF_LINKS];+};++staticintxdp_attach(structskeletons*skeletons,structbpf_program*prog,char*iface)+{+structbpf_link*link;+intifindex;++ifindex=if_nametoindex(iface);+if(!ASSERT_GT(ifindex,0,"get ifindex"))+return-1;++if(!ASSERT_LE(skeletons->nlinks+1,MAX_BPF_LINKS,"too many XDP programs attached"))+return-1;++link=bpf_program__attach_xdp(prog,ifindex);+if(!ASSERT_OK_PTR(link,"attach xdp program"))+return-1;++skeletons->links[skeletons->nlinks++]=link;+return0;+}++enum{+BOND_ONE_NO_ATTACH=0,+BOND_BOTH_AND_ATTACH,+};++staticconstchar*constmode_names[]={+[BOND_MODE_ROUNDROBIN]="balance-rr",+[BOND_MODE_ACTIVEBACKUP]="active-backup",+[BOND_MODE_XOR]="balance-xor",+[BOND_MODE_BROADCAST]="broadcast",+[BOND_MODE_8023AD]="802.3ad",+[BOND_MODE_TLB]="balance-tlb",+[BOND_MODE_ALB]="balance-alb",+};++staticconstchar*constxmit_policy_names[]={+[BOND_XMIT_POLICY_LAYER2]="layer2",+[BOND_XMIT_POLICY_LAYER34]="layer3+4",+[BOND_XMIT_POLICY_LAYER23]="layer2+3",+[BOND_XMIT_POLICY_ENCAP23]="encap2+3",+[BOND_XMIT_POLICY_ENCAP34]="encap3+4",+};++staticintbonding_setup(structskeletons*skeletons,intmode,intxmit_policy,+intbond_both_attach)+{+#define SYS(fmt, ...) \+({\+charcmd[1024];\+snprintf(cmd,sizeof(cmd),fmt,##__VA_ARGS__);\+if(!ASSERT_OK(system(cmd),cmd))\+return-1;\+})++SYS("ip netns add ns_dst");+SYS("ip link add veth1_1 type veth peer name veth2_1 netns ns_dst");+SYS("ip link add veth1_2 type veth peer name veth2_2 netns ns_dst");++SYS("ip link add bond1 type bond mode %s xmit_hash_policy %s",+mode_names[mode],xmit_policy_names[xmit_policy]);+SYS("ip link set bond1 up address "BOND1_MAC_STR" addrgenmode none");+SYS("ip -netns ns_dst link add bond2 type bond mode %s xmit_hash_policy %s",+mode_names[mode],xmit_policy_names[xmit_policy]);+SYS("ip -netns ns_dst link set bond2 up address "BOND2_MAC_STR" addrgenmode none");++SYS("ip link set veth1_1 master bond1");+if(bond_both_attach==BOND_BOTH_AND_ATTACH){+SYS("ip link set veth1_2 master bond1");+}else{+SYS("ip link set veth1_2 up addrgenmode none");++if(xdp_attach(skeletons,skeletons->xdp_dummy->progs.xdp_dummy_prog,"veth1_2"))+return-1;+}++SYS("ip -netns ns_dst link set veth2_1 master bond2");++if(bond_both_attach==BOND_BOTH_AND_ATTACH)+SYS("ip -netns ns_dst link set veth2_2 master bond2");+else+SYS("ip -netns ns_dst link set veth2_2 up addrgenmode none");++/* Load a dummy program on sending side as with veth peer needs to have a+*XDPprogramloadedaswell.+*/+if(xdp_attach(skeletons,skeletons->xdp_dummy->progs.xdp_dummy_prog,"bond1"))+return-1;++if(bond_both_attach==BOND_BOTH_AND_ATTACH){+if(!ASSERT_OK(setns_by_name("ns_dst"),"set netns to ns_dst"))+return-1;++if(xdp_attach(skeletons,skeletons->xdp_tx->progs.xdp_tx,"bond2"))+return-1;++restore_root_netns();+}++return0;++#undef SYS+}++staticvoidbonding_cleanup(structskeletons*skeletons)+{+restore_root_netns();+while(skeletons->nlinks){+skeletons->nlinks--;+bpf_link__destroy(skeletons->links[skeletons->nlinks]);+}+ASSERT_OK(system("ip link delete bond1"),"delete bond1");+ASSERT_OK(system("ip link delete veth1_1"),"delete veth1_1");+ASSERT_OK(system("ip link delete veth1_2"),"delete veth1_2");+ASSERT_OK(system("ip netns delete ns_dst"),"delete ns_dst");+}++staticintsend_udp_packets(intvary_dst_ip)+{+structethhdreh={+.h_source=BOND1_MAC,+.h_dest=BOND2_MAC,+.h_proto=htons(ETH_P_IP),+};+uint8_tbuf[128]={};+structiphdr*iph=(structiphdr*)(buf+sizeof(eh));+structudphdr*uh=(structudphdr*)(buf+sizeof(eh)+sizeof(*iph));+inti,s=-1;+intifindex;++s=socket(AF_PACKET,SOCK_RAW,IPPROTO_RAW);+if(!ASSERT_GE(s,0,"socket"))+gotoerr;++ifindex=if_nametoindex("bond1");+if(!ASSERT_GT(ifindex,0,"get bond1 ifindex"))+gotoerr;++memcpy(buf,&eh,sizeof(eh));+iph->ihl=5;+iph->version=4;+iph->tos=16;+iph->id=1;+iph->ttl=64;+iph->protocol=IPPROTO_UDP;+iph->saddr=1;+iph->daddr=2;+iph->tot_len=htons(sizeof(buf)-ETH_HLEN);+iph->check=0;++for(i=1;i<=NPACKETS;i++){+intn;+structsockaddr_llsaddr_ll={+.sll_ifindex=ifindex,+.sll_halen=ETH_ALEN,+.sll_addr=BOND2_MAC,+};++/* vary the UDP destination port for even distribution with roundrobin/xor modes */+uh->dest++;++if(vary_dst_ip)+iph->daddr++;++n=sendto(s,buf,sizeof(buf),0,(structsockaddr*)&saddr_ll,sizeof(saddr_ll));+if(!ASSERT_EQ(n,sizeof(buf),"sendto"))+gotoerr;+}++return0;++err:+if(s>=0)+close(s);+return-1;+}++staticvoidtest_xdp_bonding_with_mode(structskeletons*skeletons,intmode,intxmit_policy)+{+intbond1_rx;++if(bonding_setup(skeletons,mode,xmit_policy,BOND_BOTH_AND_ATTACH))+gotoout;++if(send_udp_packets(xmit_policy!=BOND_XMIT_POLICY_LAYER34))+gotoout;++bond1_rx=get_rx_packets("bond1");+ASSERT_EQ(bond1_rx,NPACKETS,"expected more received packets");++switch(mode){+caseBOND_MODE_ROUNDROBIN:+caseBOND_MODE_XOR:{+intveth1_rx=get_rx_packets("veth1_1");+intveth2_rx=get_rx_packets("veth1_2");+intdiff=abs(veth1_rx-veth2_rx);++ASSERT_GE(veth1_rx+veth2_rx,NPACKETS,"expected more packets");++switch(xmit_policy){+caseBOND_XMIT_POLICY_LAYER2:+ASSERT_GE(diff,NPACKETS,+"expected packets on only one of the interfaces");+break;+caseBOND_XMIT_POLICY_LAYER23:+caseBOND_XMIT_POLICY_LAYER34:+ASSERT_LT(diff,NPACKETS/2,+"expected even distribution of packets");+break;+default:+PRINT_FAIL("Unimplemented xmit_policy=%d\n",xmit_policy);+break;+}+break;+}+caseBOND_MODE_ACTIVEBACKUP:{+intveth1_rx=get_rx_packets("veth1_1");+intveth2_rx=get_rx_packets("veth1_2");+intdiff=abs(veth1_rx-veth2_rx);++ASSERT_GE(diff,NPACKETS,+"expected packets on only one of the interfaces");+break;+}+default:+PRINT_FAIL("Unimplemented xmit_policy=%d\n",xmit_policy);+break;+}++out:+bonding_cleanup(skeletons);+}++/* Test the broadcast redirection using xdp_redirect_map_multi_prog and adding+*alltheinterfacestoitandcheckingthatbroadcastingwon'tsendthepacket+*toneithertheingressbonddevice(bond2)oritsslave(veth2_1).+*/+staticvoidtest_xdp_bonding_redirect_multi(structskeletons*skeletons)+{+staticconstchar*constifaces[]={"bond2","veth2_1","veth2_2"};+intveth1_1_rx,veth1_2_rx;+interr;++if(bonding_setup(skeletons,BOND_MODE_ROUNDROBIN,BOND_XMIT_POLICY_LAYER23,+BOND_ONE_NO_ATTACH))+gotoout;+++if(!ASSERT_OK(setns_by_name("ns_dst"),"could not set netns to ns_dst"))+gotoout;++/* populate the devmap with the relevant interfaces */+for(inti=0;i<ARRAY_SIZE(ifaces);i++){+intifindex=if_nametoindex(ifaces[i]);+intmap_fd=bpf_map__fd(skeletons->xdp_redirect_multi_kern->maps.map_all);++if(!ASSERT_GT(ifindex,0,"could not get interface index"))+gotoout;++err=bpf_map_update_elem(map_fd,&ifindex,&ifindex,0);+if(!ASSERT_OK(err,"add interface to map_all"))+gotoout;+}++if(xdp_attach(skeletons,+skeletons->xdp_redirect_multi_kern->progs.xdp_redirect_map_multi_prog,+"bond2"))+gotoout;++restore_root_netns();++if(send_udp_packets(BOND_MODE_ROUNDROBIN))+gotoout;++veth1_1_rx=get_rx_packets("veth1_1");+veth1_2_rx=get_rx_packets("veth1_2");++ASSERT_EQ(veth1_1_rx,0,"expected no packets on veth1_1");+ASSERT_GE(veth1_2_rx,NPACKETS,"expected packets on veth1_2");++out:+restore_root_netns();+bonding_cleanup(skeletons);+}++/* Test that XDP programs cannot be attached to both the bond master and slaves simultaneously */+staticvoidtest_xdp_bonding_attach(structskeletons*skeletons)+{+structbpf_link*link=NULL;+structbpf_link*link2=NULL;+intveth,bond;+interr;++if(!ASSERT_OK(system("ip link add veth type veth"),"add veth"))+gotoout;+if(!ASSERT_OK(system("ip link add bond type bond"),"add bond"))+gotoout;++veth=if_nametoindex("veth");+if(!ASSERT_GE(veth,0,"if_nametoindex veth"))+gotoout;+bond=if_nametoindex("bond");+if(!ASSERT_GE(bond,0,"if_nametoindex bond"))+gotoout;++/* enslaving with a XDP program loaded fails */+link=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,veth);+if(!ASSERT_OK_PTR(link,"attach program to veth"))+gotoout;++err=system("ip link set veth master bond");+if(!ASSERT_NEQ(err,0,"attaching slave with xdp program expected to fail"))+gotoout;++bpf_link__destroy(link);+link=NULL;++err=system("ip link set veth master bond");+if(!ASSERT_OK(err,"set veth master"))+gotoout;++/* attaching to slave when master has no program is allowed */+link=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,veth);+if(!ASSERT_OK_PTR(link,"attach program to slave when enslaved"))+gotoout;++/* attaching to master not allowed when slave has program loaded */+link2=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,bond);+if(!ASSERT_ERR_PTR(link2,"attach program to master when slave has program"))+gotoout;++bpf_link__destroy(link);+link=NULL;++/* attaching XDP program to master allowed when slave has no program */+link=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,bond);+if(!ASSERT_OK_PTR(link,"attach program to master"))+gotoout;++/* attaching to slave not allowed when master has program loaded */+link2=bpf_program__attach_xdp(skeletons->xdp_dummy->progs.xdp_dummy_prog,bond);+ASSERT_ERR_PTR(link2,"attach program to slave when master has program");++out:+bpf_link__destroy(link);+bpf_link__destroy(link2);++system("ip link del veth");+system("ip link del bond");+}++staticintlibbpf_debug_print(enumlibbpf_print_levellevel,+constchar*format,va_listargs)+{+if(level!=LIBBPF_WARN)+vprintf(format,args);+return0;+}++structbond_test_case{+char*name;+intmode;+intxmit_policy;+};++staticstructbond_test_casebond_test_cases[]={+{"xdp_bonding_roundrobin",BOND_MODE_ROUNDROBIN,BOND_XMIT_POLICY_LAYER23,},+{"xdp_bonding_activebackup",BOND_MODE_ACTIVEBACKUP,BOND_XMIT_POLICY_LAYER23},++{"xdp_bonding_xor_layer2",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER2,},+{"xdp_bonding_xor_layer23",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER23,},+{"xdp_bonding_xor_layer34",BOND_MODE_XOR,BOND_XMIT_POLICY_LAYER34,},+};++voidtest_xdp_bonding(void)+{+libbpf_print_fn_told_print_fn;+structskeletonsskeletons={};+inti;++old_print_fn=libbpf_set_print(libbpf_debug_print);++root_netns_fd=open("/proc/self/ns/net",O_RDONLY);+if(!ASSERT_GE(root_netns_fd,0,"open /proc/self/ns/net"))+gotoout;++skeletons.xdp_dummy=xdp_dummy__open_and_load();+if(!ASSERT_OK_PTR(skeletons.xdp_dummy,"xdp_dummy__open_and_load"))+gotoout;++skeletons.xdp_tx=xdp_tx__open_and_load();+if(!ASSERT_OK_PTR(skeletons.xdp_tx,"xdp_tx__open_and_load"))+gotoout;++skeletons.xdp_redirect_multi_kern=xdp_redirect_multi_kern__open_and_load();+if(!ASSERT_OK_PTR(skeletons.xdp_redirect_multi_kern,+"xdp_redirect_multi_kern__open_and_load"))+gotoout;++if(!test__start_subtest("xdp_bonding_attach"))+test_xdp_bonding_attach(&skeletons);++for(i=0;i<ARRAY_SIZE(bond_test_cases);i++){+structbond_test_case*test_case=&bond_test_cases[i];++if(!test__start_subtest(test_case->name))+test_xdp_bonding_with_mode(+&skeletons,+test_case->mode,+test_case->xmit_policy);+}++if(!test__start_subtest("xdp_bonding_redirect_multi"))+test_xdp_bonding_redirect_multi(&skeletons);++out:+xdp_dummy__destroy(skeletons.xdp_dummy);+xdp_tx__destroy(skeletons.xdp_tx);+xdp_redirect_multi_kern__destroy(skeletons.xdp_redirect_multi_kern);++libbpf_set_print(old_print_fn);+if(root_netns_fd)+close(root_netns_fd);+}
On Thu, Aug 5, 2021 at 9:10 AM Jussi Maki [off-list ref] wrote:
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Signed-off-by: Jussi Maki <redacted>
---
.../selftests/bpf/prog_tests/xdp_bonding.c | 520 ++++++++++++++++++
1 file changed, 520 insertions(+)
I don't pretend to understand what's going on in this selftests, but
it looks good from the generic selftest standpoint. One and half small
issues below, please double-check (and probably fix the fd close
issue).
Acked-by: Andrii Nakryiko <andrii@kernel.org>
[...]
+
+/* Test the broadcast redirection using xdp_redirect_map_multi_prog and adding
+ * all the interfaces to it and checking that broadcasting won't send the packet
+ * to neither the ingress bond device (bond2) or its slave (veth2_1).
+ */
+static void test_xdp_bonding_redirect_multi(struct skeletons *skeletons)
+{
+ static const char * const ifaces[] = {"bond2", "veth2_1", "veth2_2"};
+ int veth1_1_rx, veth1_2_rx;
+ int err;
+
+ if (bonding_setup(skeletons, BOND_MODE_ROUNDROBIN, BOND_XMIT_POLICY_LAYER23,
+ BOND_ONE_NO_ATTACH))
+ goto out;
+
+
+ if (!ASSERT_OK(setns_by_name("ns_dst"), "could not set netns to ns_dst"))
+ goto out;
+
+ /* populate the devmap with the relevant interfaces */
+ for (int i = 0; i < ARRAY_SIZE(ifaces); i++) {
+ int ifindex = if_nametoindex(ifaces[i]);
+ int map_fd = bpf_map__fd(skeletons->xdp_redirect_multi_kern->maps.map_all);
+
+ if (!ASSERT_GT(ifindex, 0, "could not get interface index"))
+ goto out;
+
+ err = bpf_map_update_elem(map_fd, &ifindex, &ifindex, 0);
+ if (!ASSERT_OK(err, "add interface to map_all"))
+ goto out;
+ }
+
+ if (xdp_attach(skeletons,
+ skeletons->xdp_redirect_multi_kern->progs.xdp_redirect_map_multi_prog,
+ "bond2"))
+ goto out;
+
+ restore_root_netns();
the "goto out" below might call restore_root_netns() again, is that ok?
+
+ if (send_udp_packets(BOND_MODE_ROUNDROBIN))
+ goto out;
+
+ veth1_1_rx = get_rx_packets("veth1_1");
+ veth1_2_rx = get_rx_packets("veth1_2");
+
+ ASSERT_EQ(veth1_1_rx, 0, "expected no packets on veth1_1");
+ ASSERT_GE(veth1_2_rx, NPACKETS, "expected packets on veth1_2");
+
+out:
+ restore_root_netns();
+ bonding_cleanup(skeletons);
+}
+
technically, fd could be 0, so for fds we have if (fd >= 0)
everywhere. Also, if open() above fails, root_netns_fd will be -1 and
you'll still attempt to close it.
On Thu, Aug 5, 2021 at 9:10 AM Jussi Maki [off-list ref] wrote:
The program type cannot be deduced from 'tx' which causes an invalid
argument error when trying to load xdp_tx.o using the skeleton.
Rename the section name to "xdp" so that libbpf can deduce the type.
Signed-off-by: Jussi Maki <redacted>
---
@@ -108,7 +108,7 @@ ip link set dev veth2 xdp pinned $BPF_DIR/progs/redirect_map_1 iplinksetdevveth3xdppinned$BPF_DIR/progs/redirect_map_2 ip-nns1linksetdevveth11xdpobjxdp_dummy.osecxdp_dummy-ip-nns2linksetdevveth22xdpobjxdp_tx.osectx+ip-nns2linksetdevveth22xdpobjxdp_tx.osecxdp ip-nns3linksetdevveth33xdpobjxdp_dummy.osecxdp_dummytrapcleanupEXIT--
From: Jussi Maki <hidden> Date: 2021-08-09 14:25:11
On Sat, Aug 7, 2021 at 12:50 AM Andrii Nakryiko
[off-list ref] wrote:
On Thu, Aug 5, 2021 at 9:10 AM Jussi Maki [off-list ref] wrote:
quoted
Add a test suite to test XDP bonding implementation
over a pair of veth devices.
Signed-off-by: Jussi Maki <redacted>
---
.../selftests/bpf/prog_tests/xdp_bonding.c | 520 ++++++++++++++++++
1 file changed, 520 insertions(+)
I don't pretend to understand what's going on in this selftests, but
it looks good from the generic selftest standpoint. One and half small
issues below, please double-check (and probably fix the fd close
issue).
the "goto out" below might call restore_root_netns() again, is that ok?
Yep that's fine.
quoted
+ if (!test__start_subtest("xdp_bonding_redirect_multi"))
+ test_xdp_bonding_redirect_multi(&skeletons);
+
+out:
+ xdp_dummy__destroy(skeletons.xdp_dummy);
+ xdp_tx__destroy(skeletons.xdp_tx);
+ xdp_redirect_multi_kern__destroy(skeletons.xdp_redirect_multi_kern);
+
+ libbpf_set_print(old_print_fn);
+ if (root_netns_fd)
technically, fd could be 0, so for fds we have if (fd >= 0)
everywhere. Also, if open() above fails, root_netns_fd will be -1 and
you'll still attempt to close it.
Good catch. Daniel, could you fix this when applying to be "if
(root_netns_fd >= 0)"?
From: Daniel Borkmann <daniel@iogearbox.net> Date: 2021-08-09 21:41:37
On 8/9/21 4:24 PM, Jussi Maki wrote:
[...]
quoted
quoted
+ if (!test__start_subtest("xdp_bonding_redirect_multi"))
+ test_xdp_bonding_redirect_multi(&skeletons);
+
+out:
+ xdp_dummy__destroy(skeletons.xdp_dummy);
+ xdp_tx__destroy(skeletons.xdp_tx);
+ xdp_redirect_multi_kern__destroy(skeletons.xdp_redirect_multi_kern);
+
+ libbpf_set_print(old_print_fn);
+ if (root_netns_fd)
technically, fd could be 0, so for fds we have if (fd >= 0)
everywhere. Also, if open() above fails, root_netns_fd will be -1 and
you'll still attempt to close it.
Good catch. Daniel, could you fix this when applying to be "if
(root_netns_fd >= 0)"?
Yep, done now, I had to rebase due to 220ade77452c ("bonding: 3ad: fix the concurrency
between __bond_release_one() and bond_3ad_state_machine_handler()") which this series
here didn't take into account. Please double check.
Thanks everyone,
Daniel
From: Jonathan Toppins <hidden> Date: 2021-08-11 01:52:59
On 7/31/21 1:57 AM, Jussi Maki wrote:
quoted hunk
In preparation for adding XDP support to the bonding driver
refactor the packet hashing functions to be able to work with
any linear data buffer without an skb.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 147 +++++++++++++++++++-------------
1 file changed, 90 insertions(+), 57 deletions(-)
@@ -3611,55 +3611,80 @@ static struct notifier_block bond_netdev_notifier = {/*---------------------------- Hashing Policies -----------------------------*/+/* Helper to access data in a packet, with or without a backing skb.+*Ifskbisgiventhedataislinearizedifnecessaryviapskb_may_pull.+*/+staticinlineconstvoid*bond_pull_data(structsk_buff*skb,+constvoid*data,inthlen,intn)+{+if(likely(n<=hlen))+returndata;+elseif(skb&&likely(pskb_may_pull(skb,n)))+returnskb->head;++returnNULL;+}+/* L2 hash helper */-staticinlineu32bond_eth_hash(structsk_buff*skb)+staticinlineu32bond_eth_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*ep,hdr_tmp;+structethhdr*ep;-ep=skb_header_pointer(skb,0,sizeof(hdr_tmp),&hdr_tmp);-if(ep)-returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;-return0;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;++ep=(structethhdr*)(data+mhoff);+returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;}-staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,-int*noff,int*proto,booll34)+staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,constvoid*data,+inthlen,__be16l2_proto,int*nhoff,int*ip_proto,booll34){conststructipv6hdr*iph6;conststructiphdr*iph;-if(skb->protocol==htons(ETH_P_IP)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph))))+if(l2_proto==htons(ETH_P_IP)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph));+if(!data)returnfalse;-iph=(conststructiphdr*)(skb->data+*noff);++iph=(conststructiphdr*)(data+*nhoff);iph_to_flow_copy_v4addrs(fk,iph);-*noff+=iph->ihl<<2;+*nhoff+=iph->ihl<<2;if(!ip_is_fragment(iph))-*proto=iph->protocol;-}elseif(skb->protocol==htons(ETH_P_IPV6)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph6))))+*ip_proto=iph->protocol;+}elseif(l2_proto==htons(ETH_P_IPV6)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph6));+if(!data)returnfalse;-iph6=(conststructipv6hdr*)(skb->data+*noff);++iph6=(conststructipv6hdr*)(data+*nhoff);iph_to_flow_copy_v6addrs(fk,iph6);-*noff+=sizeof(*iph6);-*proto=iph6->nexthdr;+*nhoff+=sizeof(*iph6);+*ip_proto=iph6->nexthdr;}else{returnfalse;}-if(l34&&*proto>=0)-fk->ports.ports=skb_flow_get_ports(skb,*noff,*proto);+if(l34&&*ip_proto>=0)+fk->ports.ports=__skb_flow_get_ports(skb,*nhoff,*ip_proto,data,hlen);returntrue;}-staticu32bond_vlan_srcmac_hash(structsk_buff*skb)+staticu32bond_vlan_srcmac_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*mac_hdr=(structethhdr*)skb_mac_header(skb);+structethhdr*mac_hdr;u32srcmac_vendor=0,srcmac_dev=0;u16vlan;inti;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;+mac_hdr=(structethhdr*)(data+mhoff);
The XDP changes are not introduced in this patch but this section looks
consistent in later patches in the series. So assuming the XDP buff
passed gets to this point how will a NULL dereference be avoided given
skb == NULL, in the XDP call path, as skb is dereferenced later in the
function?
By this section:
...
if (!skb_vlan_tag_present(skb))
return srcmac_vendor ^ srcmac_dev;
vlan = skb_vlan_tag_get(skb);
...
referencing net-next/master id: d1a4e0a9576fd2b29a0d13b306a9f52440908ab4
quoted hunk
+
for (i = 0; i < 3; i++)
srcmac_vendor = (srcmac_vendor << 8) | mac_hdr->h_source[i];
@@ -3675,26 +3700,25 @@ static u32 bond_vlan_srcmac_hash(struct sk_buff *skb) } /* Extract the appropriate headers based on bond's xmit policy */-static bool bond_flow_dissect(struct bonding *bond, struct sk_buff *skb,- struct flow_keys *fk)+static bool bond_flow_dissect(struct bonding *bond, struct sk_buff *skb, const void *data,+ __be16 l2_proto, int nhoff, int hlen, struct flow_keys *fk) { bool l34 = bond->params.xmit_policy == BOND_XMIT_POLICY_LAYER34;- int noff, proto = -1;+ int ip_proto = -1; switch (bond->params.xmit_policy) { case BOND_XMIT_POLICY_ENCAP23: case BOND_XMIT_POLICY_ENCAP34: memset(fk, 0, sizeof(*fk)); return __skb_flow_dissect(NULL, skb, &flow_keys_bonding,- fk, NULL, 0, 0, 0, 0);+ fk, data, l2_proto, nhoff, hlen, 0); default: break; } fk->ports.ports = 0; memset(&fk->icmp, 0, sizeof(fk->icmp));- noff = skb_network_offset(skb);- if (!bond_flow_ip(skb, fk, &noff, &proto, l34))+ if (!bond_flow_ip(skb, fk, data, hlen, l2_proto, &nhoff, &ip_proto, l34)) return false; /* ICMP error packets contains at least 8 bytes of the header
@@ -3702,22 +3726,20 @@ static bool bond_flow_dissect(struct bonding *bond, struct sk_buff *skb, * to correlate ICMP error packets within the same flow which * generated the error. */- if (proto == IPPROTO_ICMP || proto == IPPROTO_ICMPV6) {- skb_flow_get_icmp_tci(skb, &fk->icmp, skb->data,- skb_transport_offset(skb),- skb_headlen(skb));- if (proto == IPPROTO_ICMP) {+ if (ip_proto == IPPROTO_ICMP || ip_proto == IPPROTO_ICMPV6) {+ skb_flow_get_icmp_tci(skb, &fk->icmp, data, nhoff, hlen);+ if (ip_proto == IPPROTO_ICMP) { if (!icmp_is_err(fk->icmp.type)) return true;- noff += sizeof(struct icmphdr);- } else if (proto == IPPROTO_ICMPV6) {+ nhoff += sizeof(struct icmphdr);+ } else if (ip_proto == IPPROTO_ICMPV6) { if (!icmpv6_is_err(fk->icmp.type)) return true;- noff += sizeof(struct icmp6hdr);+ nhoff += sizeof(struct icmp6hdr); }- return bond_flow_ip(skb, fk, &noff, &proto, l34);+ return bond_flow_ip(skb, fk, data, hlen, l2_proto, &nhoff, &ip_proto, l34); } return true;
@@ -3733,33 +3755,26 @@ static u32 bond_ip_hash(u32 hash, struct flow_keys *flow) return hash >> 1; }-/**- * bond_xmit_hash - generate a hash value based on the xmit policy- * @bond: bonding device- * @skb: buffer to use for headers- *- * This function will extract the necessary headers from the skb buffer and use- * them to generate a hash based on the xmit_policy set in the bonding device+/* Generate hash based on xmit policy. If @skb is given it is used to linearize+ * the data as required, but this function can be used without it if the data is+ * known to be linear (e.g. with xdp_buff). */-u32 bond_xmit_hash(struct bonding *bond, struct sk_buff *skb)+static u32 __bond_xmit_hash(struct bonding *bond, struct sk_buff *skb, const void *data,+ __be16 l2_proto, int mhoff, int nhoff, int hlen) { struct flow_keys flow; u32 hash;- if (bond->params.xmit_policy == BOND_XMIT_POLICY_ENCAP34 &&- skb->l4_hash)- return skb->hash;- if (bond->params.xmit_policy == BOND_XMIT_POLICY_VLAN_SRCMAC)- return bond_vlan_srcmac_hash(skb);+ return bond_vlan_srcmac_hash(skb, data, mhoff, hlen); if (bond->params.xmit_policy == BOND_XMIT_POLICY_LAYER2 ||- !bond_flow_dissect(bond, skb, &flow))- return bond_eth_hash(skb);+ !bond_flow_dissect(bond, skb, data, l2_proto, nhoff, hlen, &flow))+ return bond_eth_hash(skb, data, mhoff, hlen); if (bond->params.xmit_policy == BOND_XMIT_POLICY_LAYER23 || bond->params.xmit_policy == BOND_XMIT_POLICY_ENCAP23) {- hash = bond_eth_hash(skb);+ hash = bond_eth_hash(skb, data, mhoff, hlen); } else { if (flow.icmp.id) memcpy(&hash, &flow.icmp, sizeof(hash));
@@ -3770,6 +3785,25 @@ u32 bond_xmit_hash(struct bonding *bond, struct sk_buff *skb) return bond_ip_hash(hash, &flow); }+/**+ * bond_xmit_hash - generate a hash value based on the xmit policy+ * @bond: bonding device+ * @skb: buffer to use for headers+ *+ * This function will extract the necessary headers from the skb buffer and use+ * them to generate a hash based on the xmit_policy set in the bonding device+ */+u32 bond_xmit_hash(struct bonding *bond, struct sk_buff *skb)+{+ if (bond->params.xmit_policy == BOND_XMIT_POLICY_ENCAP34 &&+ skb->l4_hash)+ return skb->hash;++ return __bond_xmit_hash(bond, skb, skb->head, skb->protocol,+ skb->mac_header, skb->network_header,+ skb_headlen(skb));+}+ /*-------------------------- Device entry points ----------------------------*/ void bond_work_init_all(struct bonding *bond)
From: Jussi Maki <hidden> Date: 2021-08-11 08:22:46
Hi Jonathan,
Thanks for catching this. You're right, this will NULL deref if XDP
bonding is used with the VLAN_SRCMAC xmit policy. I think what
happened was that a very early version restricted the xmit policies
that were applicable, but it got dropped when this was refactored.
I'll look into this today and will add in support (or refuse) the
VLAN_SRCMAC xmit policy and extend the tests to cover this.
On Wed, Aug 11, 2021 at 3:52 AM Jonathan Toppins [off-list ref] wrote:
On 7/31/21 1:57 AM, Jussi Maki wrote:
quoted
In preparation for adding XDP support to the bonding driver
refactor the packet hashing functions to be able to work with
any linear data buffer without an skb.
Signed-off-by: Jussi Maki <redacted>
---
drivers/net/bonding/bond_main.c | 147 +++++++++++++++++++-------------
1 file changed, 90 insertions(+), 57 deletions(-)
@@ -3611,55 +3611,80 @@ static struct notifier_block bond_netdev_notifier = {/*---------------------------- Hashing Policies -----------------------------*/+/* Helper to access data in a packet, with or without a backing skb.+*Ifskbisgiventhedataislinearizedifnecessaryviapskb_may_pull.+*/+staticinlineconstvoid*bond_pull_data(structsk_buff*skb,+constvoid*data,inthlen,intn)+{+if(likely(n<=hlen))+returndata;+elseif(skb&&likely(pskb_may_pull(skb,n)))+returnskb->head;++returnNULL;+}+/* L2 hash helper */-staticinlineu32bond_eth_hash(structsk_buff*skb)+staticinlineu32bond_eth_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*ep,hdr_tmp;+structethhdr*ep;-ep=skb_header_pointer(skb,0,sizeof(hdr_tmp),&hdr_tmp);-if(ep)-returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;-return0;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;++ep=(structethhdr*)(data+mhoff);+returnep->h_dest[5]^ep->h_source[5]^ep->h_proto;}-staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,-int*noff,int*proto,booll34)+staticboolbond_flow_ip(structsk_buff*skb,structflow_keys*fk,constvoid*data,+inthlen,__be16l2_proto,int*nhoff,int*ip_proto,booll34){conststructipv6hdr*iph6;conststructiphdr*iph;-if(skb->protocol==htons(ETH_P_IP)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph))))+if(l2_proto==htons(ETH_P_IP)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph));+if(!data)returnfalse;-iph=(conststructiphdr*)(skb->data+*noff);++iph=(conststructiphdr*)(data+*nhoff);iph_to_flow_copy_v4addrs(fk,iph);-*noff+=iph->ihl<<2;+*nhoff+=iph->ihl<<2;if(!ip_is_fragment(iph))-*proto=iph->protocol;-}elseif(skb->protocol==htons(ETH_P_IPV6)){-if(unlikely(!pskb_may_pull(skb,*noff+sizeof(*iph6))))+*ip_proto=iph->protocol;+}elseif(l2_proto==htons(ETH_P_IPV6)){+data=bond_pull_data(skb,data,hlen,*nhoff+sizeof(*iph6));+if(!data)returnfalse;-iph6=(conststructipv6hdr*)(skb->data+*noff);++iph6=(conststructipv6hdr*)(data+*nhoff);iph_to_flow_copy_v6addrs(fk,iph6);-*noff+=sizeof(*iph6);-*proto=iph6->nexthdr;+*nhoff+=sizeof(*iph6);+*ip_proto=iph6->nexthdr;}else{returnfalse;}-if(l34&&*proto>=0)-fk->ports.ports=skb_flow_get_ports(skb,*noff,*proto);+if(l34&&*ip_proto>=0)+fk->ports.ports=__skb_flow_get_ports(skb,*nhoff,*ip_proto,data,hlen);returntrue;}-staticu32bond_vlan_srcmac_hash(structsk_buff*skb)+staticu32bond_vlan_srcmac_hash(structsk_buff*skb,constvoid*data,intmhoff,inthlen){-structethhdr*mac_hdr=(structethhdr*)skb_mac_header(skb);+structethhdr*mac_hdr;u32srcmac_vendor=0,srcmac_dev=0;u16vlan;inti;+data=bond_pull_data(skb,data,hlen,mhoff+sizeof(structethhdr));+if(!data)+return0;+mac_hdr=(structethhdr*)(data+mhoff);
The XDP changes are not introduced in this patch but this section looks
consistent in later patches in the series. So assuming the XDP buff
passed gets to this point how will a NULL dereference be avoided given
skb == NULL, in the XDP call path, as skb is dereferenced later in the
function?
By this section:
...
if (!skb_vlan_tag_present(skb))
return srcmac_vendor ^ srcmac_dev;
vlan = skb_vlan_tag_get(skb);
...
referencing net-next/master id: d1a4e0a9576fd2b29a0d13b306a9f52440908ab4
quoted
+
for (i = 0; i < 3; i++)
srcmac_vendor = (srcmac_vendor << 8) | mac_hdr->h_source[i];
@@ -3675,26 +3700,25 @@ static u32 bond_vlan_srcmac_hash(struct sk_buff *skb) } /* Extract the appropriate headers based on bond's xmit policy */-static bool bond_flow_dissect(struct bonding *bond, struct sk_buff *skb,- struct flow_keys *fk)+static bool bond_flow_dissect(struct bonding *bond, struct sk_buff *skb, const void *data,+ __be16 l2_proto, int nhoff, int hlen, struct flow_keys *fk) { bool l34 = bond->params.xmit_policy == BOND_XMIT_POLICY_LAYER34;- int noff, proto = -1;+ int ip_proto = -1; switch (bond->params.xmit_policy) { case BOND_XMIT_POLICY_ENCAP23: case BOND_XMIT_POLICY_ENCAP34: memset(fk, 0, sizeof(*fk)); return __skb_flow_dissect(NULL, skb, &flow_keys_bonding,- fk, NULL, 0, 0, 0, 0);+ fk, data, l2_proto, nhoff, hlen, 0); default: break; } fk->ports.ports = 0; memset(&fk->icmp, 0, sizeof(fk->icmp));- noff = skb_network_offset(skb);- if (!bond_flow_ip(skb, fk, &noff, &proto, l34))+ if (!bond_flow_ip(skb, fk, data, hlen, l2_proto, &nhoff, &ip_proto, l34)) return false; /* ICMP error packets contains at least 8 bytes of the header
@@ -3702,22 +3726,20 @@ static bool bond_flow_dissect(struct bonding *bond, struct sk_buff *skb, * to correlate ICMP error packets within the same flow which * generated the error. */- if (proto == IPPROTO_ICMP || proto == IPPROTO_ICMPV6) {- skb_flow_get_icmp_tci(skb, &fk->icmp, skb->data,- skb_transport_offset(skb),- skb_headlen(skb));- if (proto == IPPROTO_ICMP) {+ if (ip_proto == IPPROTO_ICMP || ip_proto == IPPROTO_ICMPV6) {+ skb_flow_get_icmp_tci(skb, &fk->icmp, data, nhoff, hlen);+ if (ip_proto == IPPROTO_ICMP) { if (!icmp_is_err(fk->icmp.type)) return true;- noff += sizeof(struct icmphdr);- } else if (proto == IPPROTO_ICMPV6) {+ nhoff += sizeof(struct icmphdr);+ } else if (ip_proto == IPPROTO_ICMPV6) { if (!icmpv6_is_err(fk->icmp.type)) return true;- noff += sizeof(struct icmp6hdr);+ nhoff += sizeof(struct icmp6hdr); }- return bond_flow_ip(skb, fk, &noff, &proto, l34);+ return bond_flow_ip(skb, fk, data, hlen, l2_proto, &nhoff, &ip_proto, l34); } return true;
@@ -3733,33 +3755,26 @@ static u32 bond_ip_hash(u32 hash, struct flow_keys *flow) return hash >> 1; }-/**- * bond_xmit_hash - generate a hash value based on the xmit policy- * @bond: bonding device- * @skb: buffer to use for headers- *- * This function will extract the necessary headers from the skb buffer and use- * them to generate a hash based on the xmit_policy set in the bonding device+/* Generate hash based on xmit policy. If @skb is given it is used to linearize+ * the data as required, but this function can be used without it if the data is+ * known to be linear (e.g. with xdp_buff). */-u32 bond_xmit_hash(struct bonding *bond, struct sk_buff *skb)+static u32 __bond_xmit_hash(struct bonding *bond, struct sk_buff *skb, const void *data,+ __be16 l2_proto, int mhoff, int nhoff, int hlen) { struct flow_keys flow; u32 hash;- if (bond->params.xmit_policy == BOND_XMIT_POLICY_ENCAP34 &&- skb->l4_hash)- return skb->hash;- if (bond->params.xmit_policy == BOND_XMIT_POLICY_VLAN_SRCMAC)- return bond_vlan_srcmac_hash(skb);+ return bond_vlan_srcmac_hash(skb, data, mhoff, hlen); if (bond->params.xmit_policy == BOND_XMIT_POLICY_LAYER2 ||- !bond_flow_dissect(bond, skb, &flow))- return bond_eth_hash(skb);+ !bond_flow_dissect(bond, skb, data, l2_proto, nhoff, hlen, &flow))+ return bond_eth_hash(skb, data, mhoff, hlen); if (bond->params.xmit_policy == BOND_XMIT_POLICY_LAYER23 || bond->params.xmit_policy == BOND_XMIT_POLICY_ENCAP23) {- hash = bond_eth_hash(skb);+ hash = bond_eth_hash(skb, data, mhoff, hlen); } else { if (flow.icmp.id) memcpy(&hash, &flow.icmp, sizeof(hash));
@@ -3770,6 +3785,25 @@ u32 bond_xmit_hash(struct bonding *bond, struct sk_buff *skb) return bond_ip_hash(hash, &flow); }+/**+ * bond_xmit_hash - generate a hash value based on the xmit policy+ * @bond: bonding device+ * @skb: buffer to use for headers+ *+ * This function will extract the necessary headers from the skb buffer and use+ * them to generate a hash based on the xmit_policy set in the bonding device+ */+u32 bond_xmit_hash(struct bonding *bond, struct sk_buff *skb)+{+ if (bond->params.xmit_policy == BOND_XMIT_POLICY_ENCAP34 &&+ skb->l4_hash)+ return skb->hash;++ return __bond_xmit_hash(bond, skb, skb->head, skb->protocol,+ skb->mac_header, skb->network_header,+ skb_headlen(skb));+}+ /*-------------------------- Device entry points ----------------------------*/ void bond_work_init_all(struct bonding *bond)
From: Jonathan Toppins <hidden> Date: 2021-08-11 14:05:19
On 8/11/21 4:22 AM, Jussi Maki wrote:
Hi Jonathan,
Thanks for catching this. You're right, this will NULL deref if XDP
bonding is used with the VLAN_SRCMAC xmit policy. I think what
happened was that a very early version restricted the xmit policies
that were applicable, but it got dropped when this was refactored.
I'll look into this today and will add in support (or refuse) the
VLAN_SRCMAC xmit policy and extend the tests to cover this.
In support of some customer requests and to stop adding more and more
hashing policies I was looking at adding a custom policy that exposes a
bitfield so userspace can select which header items should be included
in the hash. I was looking at a flow dissector implementation to parse
the packet and then generate the hash from the flow data pulled. It
looks like the outer hashing functions as they exist now,
bond_xmit_hash() and bond_xmit_hash_xdp(), could make the correctly
formatted call to __skb_flow_dissect(). We would then pass around the
resultant struct flow_keys, or bonding specific one to add MAC header
parsing support, and it appears we could avoid making the actual hashing
functions know if they need to hash an sk_buff vs xdp_buff. What do you
think?
-Jon
From: Jussi Maki <hidden> Date: 2021-08-16 09:06:27
On Wed, Aug 11, 2021 at 4:05 PM Jonathan Toppins [off-list ref] wrote:
On 8/11/21 4:22 AM, Jussi Maki wrote:
quoted
Hi Jonathan,
Thanks for catching this. You're right, this will NULL deref if XDP
bonding is used with the VLAN_SRCMAC xmit policy. I think what
happened was that a very early version restricted the xmit policies
that were applicable, but it got dropped when this was refactored.
I'll look into this today and will add in support (or refuse) the
VLAN_SRCMAC xmit policy and extend the tests to cover this.
In support of some customer requests and to stop adding more and more
hashing policies I was looking at adding a custom policy that exposes a
bitfield so userspace can select which header items should be included
in the hash. I was looking at a flow dissector implementation to parse
the packet and then generate the hash from the flow data pulled. It
looks like the outer hashing functions as they exist now,
bond_xmit_hash() and bond_xmit_hash_xdp(), could make the correctly
formatted call to __skb_flow_dissect(). We would then pass around the
resultant struct flow_keys, or bonding specific one to add MAC header
parsing support, and it appears we could avoid making the actual hashing
functions know if they need to hash an sk_buff vs xdp_buff. What do you
think?
That sounds great! I wasn't particularly happy about how it works with
skb being optional as that was just waiting to break (as it did). The
team driver does the hashing using a user-space provided bpf program
and I'm looking to figure out how to support XDP with it. I wonder if
we could have a single approach that would work for both bonding and
team (e.g. use bpf to hash). CC'ing Jiri as he wrote the team driver.