From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:05
Hi,
This patchset proposes a new fast forwarding path infrastructure that
combines the GRO/GSO and the flowtable infrastructures. The idea is to
add a hook at the GRO layer that is invoked before the standard GRO
protocol offloads. This allows us to build custom packet chains that we
can quickly pass in one go to the neighbour layer to define fast
forwarding path for flows.
For each packet that gets into the GRO layer, we first check if there is
an entry in the flowtable, if so, the packet is placed in a list until
the GRO infrastructure decides to send the batch from gro_complete to
the neighbour layer. The first packet in the list takes the route from
the flowtable entry, so we avoid reiterative routing lookups.
In case no entry is found in the flowtable, the packet is passed up to
the classic GRO offload handlers. Thus, this packet follows the standard
forwarding path. Note that the initial packets of the flow always go
through the standard IPv4/IPv6 netfilter forward hook, that is used to
configure what flows are placed in the flowtable. Therefore, only a few
(initial) packets follow the standard forwarding path while most of the
follow up packets take this new fast forwarding path.
The fast forwarding path is enabled through explicit user policy, so the
user needs to request this behaviour from control plane, the following
example shows how to place flows in the new fast forwarding path from
the netfilter forward chain:
table x {
flowtable f {
hook early_ingress priority 0; devices = { eth0, eth1 }
}
chain y {
type filter hook forward priority 0;
ip protocol tcp flow offload @f
}
}
The example above defines a fastpath for TCP flows that are placed in
the flowtable 'f', this flowtable is hooked at the new early_ingress
hook. The initial TCP packets that match this rule from the standard
fowarding path create an entry in the flowtable, thus, GRO creates chain
of packets for those that find an entry in the flowtable and send
them through the neighbour layer.
This new hook is happening before the ingress taps, therefore, packets
that follow this new fast forwarding path are not shown by tcpdump.
This patchset supports both layer 3 IPv4 and IPv6, and layer 4 TCP and
UDP protocols. This fastpath also integrates with the IPSec
infrastructure and the ESP protocol.
We have collected performance numbers:
TCP TSO TCP Fast Forward
32.5 Gbps 35.6 Gbps
UDP UDP Fast Forward
17.6 Gbps 35.6 Gbps
ESP ESP Fast Forward
6 Gbps 7.5 Gbps
For UDP, this is doubling performance, and we almost achieve line rate
with one single CPU using the Intel i40e NIC. We got similar numbers
with the Mellanox ConnectX-4. For TCP, this is slightly improving things
even if TSO is being defeated given that we need to segment the packet
chain in software. We would like to explore HW GRO support with hardware
vendors with this new mode, we think that should improve the TCP numbers
we are showing above even more. For ESP traffic, performance improvement
is ~25%, in this case, perf shows the bottleneck becomes the crypto layer.
This patchset is co-authored work with Steffen Klassert.
Comments are welcome, thanks.
Pablo Neira Ayuso (6):
netfilter: nft_chain_filter: add support for early ingress
netfilter: nf_flow_table: add hooknum to flowtable type
netfilter: nf_flow_table: add flowtable for early ingress hook
netfilter: nft_flow_offload: enable offload after second packet is seen
netfilter: nft_flow_offload: remove secpath check
netfilter: nft_flow_offload: make sure route is not stale
Steffen Klassert (7):
net: Add a helper to get the packet offload callbacks by priority.
net: Change priority of ipv4 and ipv6 packet offloads.
net: Add a GSO feature bit for the netfilter forward fastpath.
net: Use one bit of NAPI_GRO_CB for the netfilter fastpath.
netfilter: add early ingress hook for IPv4
netfilter: add early ingress support for IPv6
netfilter: add ESP support for early ingress
include/linux/netdev_features.h | 4 +-
include/linux/netdevice.h | 6 +-
include/linux/netfilter.h | 6 +
include/linux/netfilter_ingress.h | 1 +
include/linux/skbuff.h | 2 +
include/net/netfilter/early_ingress.h | 24 +++
include/net/netfilter/nf_flow_table.h | 4 +
include/uapi/linux/netfilter.h | 1 +
net/core/dev.c | 50 ++++-
net/ipv4/af_inet.c | 1 +
net/ipv4/netfilter/Makefile | 1 +
net/ipv4/netfilter/early_ingress.c | 327 +++++++++++++++++++++++++++++
net/ipv4/netfilter/nf_flow_table_ipv4.c | 12 ++
net/ipv6/ip6_offload.c | 1 +
net/ipv6/netfilter/Makefile | 1 +
net/ipv6/netfilter/early_ingress.c | 315 ++++++++++++++++++++++++++++
net/ipv6/netfilter/nf_flow_table_ipv6.c | 1 +
net/netfilter/Kconfig | 8 +
net/netfilter/Makefile | 1 +
net/netfilter/core.c | 35 +++-
net/netfilter/early_ingress.c | 361 ++++++++++++++++++++++++++++++++
net/netfilter/nf_flow_table_inet.c | 1 +
net/netfilter/nf_flow_table_ip.c | 72 +++++++
net/netfilter/nf_tables_api.c | 120 ++++++-----
net/netfilter/nft_chain_filter.c | 6 +-
net/netfilter/nft_flow_offload.c | 13 +-
net/xfrm/xfrm_output.c | 4 +
27 files changed, 1297 insertions(+), 81 deletions(-)
create mode 100644 include/net/netfilter/early_ingress.h
create mode 100644 net/ipv4/netfilter/early_ingress.c
create mode 100644 net/ipv6/netfilter/early_ingress.c
create mode 100644 net/netfilter/early_ingress.c
--
2.11.0
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:07
From: Steffen Klassert <steffen.klassert@secunet.com>
With this helper it is possible to request callbacks with
a certain priority. This will be used in the upcoming forward
fastpath to pass packets to the standard GRO path.
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/linux/netdevice.h | 1 +
net/core/dev.c | 14 ++++++++++++++
2 files changed, 15 insertions(+)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:08
From: Steffen Klassert <steffen.klassert@secunet.com>
The forward fastpath needs to insert callbacks with
higher priority than the standard callbacks. So change
the priority of ipv4 and ipv6 packet offloads from zero
to one. With this we are able to insert callbacks with
priotity zero if needed.
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/ipv4/af_inet.c | 1 +
net/ipv6/ip6_offload.c | 1 +
2 files changed, 2 insertions(+)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:09
From: Steffen Klassert <steffen.klassert@secunet.com>
The netfilter forward fastpath has its own logic to create
GSO packets. So add a feature bit that we can catch GSO
packets that are generated by the fastpath GRO handler.
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/linux/netdev_features.h | 4 +++-
include/linux/netdevice.h | 1 +
include/linux/skbuff.h | 2 ++
3 files changed, 6 insertions(+), 1 deletion(-)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:11
From: Steffen Klassert <steffen.klassert@secunet.com>
This patch adds a is_ffwd bit to the NAPI_GRO_CB to indicate
fastpath packtes in the GRO layer. It also implements the
logic we need for this in the generic codepath. The rest
of the needed logic is implemented within netfilter and
introduced with a followup patch.
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/linux/netdevice.h | 2 +-
net/core/dev.c | 36 +++++++++++++++++++++++++++---------
2 files changed, 28 insertions(+), 10 deletions(-)
@@ -2238,7 +2238,7 @@ struct napi_gro_cb {/* Number of gro_receive callbacks this packet already went through */u8recursion_counter:4;-/* 1 bit hole */+u8is_ffwd:1;/* used to support CHECKSUM_COMPLETE for tunneling protocols */__wsumcsum;
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:17
From: Steffen Klassert <steffen.klassert@secunet.com>
Add the new early ingress hook for the netdev family, this new hook is
called from the GRO layer before the standard ipv4 GRO layers.
This hook allows us to perform early packet filtering and to define fast
forwarding path through packet chaining and flowtables using the new GSO
netfilter type. Packet that don't follow the fast path are passed up to
the standard GRO path for aggregation as usual.
This patch adds the GRO and GSO logic for this custom packet chaining.
The chaining uses the frag_list pointer so this means we do not need to
mangle the packets, therefore the aggregation strategy we follow does
not modify the packet as in the standard GRO path - we have no need to
recalculate checksum. This chain of packets is sent from the
.gro_complete callback directly to the neighbour layer. The first packet
in the chain holds a reference to the destination route.
Supported layer 4 protocols for this custom GRO packet chaining include
TCP and UDP.
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/linux/netdevice.h | 2 +
include/linux/netfilter.h | 6 +
include/linux/netfilter_ingress.h | 1 +
include/net/netfilter/early_ingress.h | 20 +++
include/uapi/linux/netfilter.h | 1 +
net/ipv4/netfilter/Makefile | 1 +
net/ipv4/netfilter/early_ingress.c | 319 +++++++++++++++++++++++++++++++++
net/netfilter/Kconfig | 8 +
net/netfilter/Makefile | 1 +
net/netfilter/core.c | 35 +++-
net/netfilter/early_ingress.c | 323 ++++++++++++++++++++++++++++++++++
11 files changed, 716 insertions(+), 1 deletion(-)
create mode 100644 include/net/netfilter/early_ingress.h
create mode 100644 net/ipv4/netfilter/early_ingress.c
create mode 100644 net/netfilter/early_ingress.c
@@ -2,6 +2,7 @@## Makefile for the netfilter modules on top of IPv4.#+obj-$(CONFIG_NETFILTER_EARLY_INGRESS)+=early_ingress.o# objects for l3 independent conntracknf_conntrack_ipv4-y:=nf_conntrack_l3proto_ipv4.onf_conntrack_proto_icmp.o
@@ -0,0 +1,319 @@+#include<linux/kernel.h>+#include<linux/netfilter.h>+#include<linux/types.h>+#include<net/xfrm.h>+#include<net/arp.h>+#include<net/udp.h>+#include<net/tcp.h>+#include<net/protocol.h>+#include<net/netfilter/early_ingress.h>++staticconststructnet_offload__rcu*nft_ip_offloads[MAX_INET_PROTOS]__read_mostly;++staticstructsk_buff*nft_udp4_gso_segment(structsk_buff*skb,+netdev_features_tfeatures)+{+skb_push(skb,sizeof(structiphdr));+returnnft_skb_segment(skb);+}++staticstructsk_buff*nft_tcp4_gso_segment(structsk_buff*skb,+netdev_features_tfeatures)+{+skb_push(skb,sizeof(structiphdr));+returnnft_skb_segment(skb);+}++staticstructsk_buff*nft_ipv4_gso_segment(structsk_buff*skb,+netdev_features_tfeatures)+{+structsk_buff*segs=ERR_PTR(-EINVAL);+conststructnet_offload*ops;+structpacket_offload*ptype;+structiphdr*iph;+intproto;+intihl;++if(!(skb_shinfo(skb)->gso_type&SKB_GSO_NFT)){+ptype=dev_get_packet_offload(skb->protocol,1);+if(ptype)+returnptype->callbacks.gso_segment(skb,features);++returnERR_PTR(-EPROTONOSUPPORT);+}++if(SKB_GSO_CB(skb)->encap_level==0){+iph=ip_hdr(skb);+skb_reset_network_header(skb);+}else{+iph=(structiphdr*)skb->data;+}++if(unlikely(!pskb_may_pull(skb,sizeof(*iph))))+gotoout;++ihl=iph->ihl*4;+if(ihl<sizeof(*iph))+gotoout;++SKB_GSO_CB(skb)->encap_level+=ihl;++if(unlikely(!pskb_may_pull(skb,ihl)))+gotoout;++__skb_pull(skb,ihl);++proto=iph->protocol;++segs=ERR_PTR(-EPROTONOSUPPORT);++ops=rcu_dereference(nft_ip_offloads[proto]);+if(likely(ops&&ops->callbacks.gso_segment))+segs=ops->callbacks.gso_segment(skb,features);++out:+returnsegs;+}++staticintnft_ipv4_gro_complete(structsk_buff*skb,intnhoff)+{+structiphdr*iph=(structiphdr*)(skb->data+nhoff);+structdst_entry*dst=skb_dst(skb);+structrtable*rt=(structrtable*)dst;+conststructnet_offload*ops;+structpacket_offload*ptype;+structnet_device*dev;+structneighbour*neigh;+unsignedinthh_len;+interr=0;+u32nexthop;+u16count;++count=NAPI_GRO_CB(skb)->count;++if(!NAPI_GRO_CB(skb)->is_ffwd){+ptype=dev_get_packet_offload(skb->protocol,1);+if(ptype)+returnptype->callbacks.gro_complete(skb,nhoff);++return0;+}++rcu_read_lock();+ops=rcu_dereference(nft_ip_offloads[iph->protocol]);+if(!ops||!ops->callbacks.gro_complete)+gotoout_unlock;++/* Only need to add sizeof(*iph) to get to the next hdr below+*becauseanyhdrwithoptionwillhavebeenflushedin+*inet_gro_receive().+*/+err=ops->callbacks.gro_complete(skb,nhoff+sizeof(*iph));++out_unlock:+rcu_read_unlock();++if(err)+returnerr;++skb_shinfo(skb)->gso_type|=SKB_GSO_NFT;+skb_shinfo(skb)->gso_segs=count;++dev=dst->dev;+dev_hold(dev);+skb->dev=dev;++if(skb_dst(skb)->xfrm){+err=dst_output(dev_net(dev),NULL,skb);+if(err!=-EREMOTE)+return-EINPROGRESS;+}++if(count<=1)+skb_gso_reset(skb);++hh_len=LL_RESERVED_SPACE(dev);++if(unlikely(skb_headroom(skb)<hh_len&&dev->header_ops)){+structsk_buff*skb2;++skb2=skb_realloc_headroom(skb,LL_RESERVED_SPACE(dev));+if(!skb2){+kfree_skb(skb);+return-ENOMEM;+}+consume_skb(skb);+skb=skb2;+}+rcu_read_lock();+nexthop=(__forceu32)rt_nexthop(rt,iph->daddr);+neigh=__ipv4_neigh_lookup_noref(dev,nexthop);+if(unlikely(!neigh))+neigh=__neigh_create(&arp_tbl,&nexthop,dev,false);+if(!IS_ERR(neigh))+neigh_output(neigh,skb);+rcu_read_unlock();++return-EINPROGRESS;+}++staticstructsk_buff**nft_ipv4_gro_receive(structsk_buff**head,+structsk_buff*skb)+{+conststructnet_offload*ops;+structpacket_offload*ptype;+structsk_buff**pp=NULL;+structsk_buff*p;+structiphdr*iph;+unsignedinthlen;+unsignedintoff;+intproto,ret;++off=skb_gro_offset(skb);+hlen=off+sizeof(*iph);++iph=skb_gro_header_slow(skb,hlen,off);+if(unlikely(!iph)){+pp=ERR_PTR(-EPERM);+gotoout;+}++proto=iph->protocol;++rcu_read_lock();++if(*(u8*)iph!=0x45){+kfree_skb(skb);+pp=ERR_PTR(-EPERM);+gotoout_unlock;+}++if(unlikely(ip_fast_csum((u8*)iph,5))){+kfree_skb(skb);+pp=ERR_PTR(-EPERM);+gotoout_unlock;+}++if(ip_is_fragment(iph))+gotoout_unlock;++ret=nf_hook_early_ingress(skb);+switch(ret){+caseNF_STOLEN:+break;+caseNF_ACCEPT:+ptype=dev_get_packet_offload(skb->protocol,1);+if(ptype)+pp=ptype->callbacks.gro_receive(head,skb);++gotoout_unlock;+caseNF_DROP:+pp=ERR_PTR(-EPERM);+gotoout_unlock;+}++ops=rcu_dereference(nft_ip_offloads[proto]);+if(!ops||!ops->callbacks.gro_receive)+gotoout_unlock;++if(iph->ttl<=1){+kfree_skb(skb);+pp=ERR_PTR(-EPERM);+gotoout_unlock;+}++skb->ip_summed=CHECKSUM_UNNECESSARY;++for(p=*head;p;p=p->next){+structiphdr*iph2;++if(!NAPI_GRO_CB(p)->same_flow)+continue;++iph2=ip_hdr(p);+/* The above works because, with the exception of the top+*(innermost)layer,weonlyaggregatepktswiththesame+*hdrlengthsoallthehdrswe'llneedtoverifywillstart+*atthesameoffset.+*/+if((iph->protocol^iph2->protocol)|+((__forceu32)iph->saddr^(__forceu32)iph2->saddr)|+((__forceu32)iph->daddr^(__forceu32)iph2->daddr)){+NAPI_GRO_CB(p)->same_flow=0;+continue;+}++if(!NAPI_GRO_CB(p)->is_ffwd)+continue;++if(!skb_dst(p))+continue;++/* All fields must match except length and checksum. */+NAPI_GRO_CB(p)->flush|=+((iph->ttl-1)^iph2->ttl)|+(iph->tos^iph2->tos)|+((iph->frag_off^iph2->frag_off)&htons(IP_DF));++pp=&p;++break;+}++NAPI_GRO_CB(skb)->is_atomic=!!(iph->frag_off&htons(IP_DF));++ip_decrease_ttl(iph);+skb->priority=rt_tos2priority(iph->tos);++skb_pull(skb,off);+NAPI_GRO_CB(skb)->data_offset=sizeof(*iph);+skb_reset_network_header(skb);+skb_set_transport_header(skb,sizeof(*iph));++pp=call_gro_receive(ops->callbacks.gro_receive,head,skb);+out_unlock:+rcu_read_unlock();++out:+NAPI_GRO_CB(skb)->data_offset=0;+returnpp;+}++staticstructpacket_offloadnft_ipv4_packet_offload__read_mostly={+.type=cpu_to_be16(ETH_P_IP),+.priority=0,+.callbacks={+.gro_receive=nft_ipv4_gro_receive,+.gro_complete=nft_ipv4_gro_complete,+.gso_segment=nft_ipv4_gso_segment,+},+};++staticconststructnet_offloadnft_udp4_offload={+.callbacks={+.gso_segment=nft_udp4_gso_segment,+.gro_receive=nft_udp_gro_receive,+},+};++staticconststructnet_offloadnft_tcp4_offload={+.callbacks={+.gso_segment=nft_tcp4_gso_segment,+.gro_receive=nft_tcp_gro_receive,+},+};++staticconststructnet_offload__rcu*nft_ip_offloads[MAX_INET_PROTOS]__read_mostly={+[IPPROTO_UDP]=&nft_udp4_offload,+[IPPROTO_TCP]=&nft_tcp4_offload,+};++voidnf_early_ingress_ip_enable(void)+{+dev_add_offload(&nft_ipv4_packet_offload);+}++voidnf_early_ingress_ip_disable(void)+{+dev_remove_offload(&nft_ipv4_packet_offload);+}
@@ -306,6 +306,11 @@ nf_hook_entry_head(struct net *net, int pf, unsigned int hooknum,return&dev->nf_hooks_ingress;}#endif+if(hooknum==NF_NETDEV_EARLY_INGRESS){+if(dev&&dev_net(dev)==net)+return&dev->nf_hooks_early_ingress;+}+WARN_ON_ONCE(1);returnNULL;}
@@ -321,7 +326,8 @@ static int __nf_register_net_hook(struct net *net, int pf,if(reg->hooknum==NF_NETDEV_INGRESS)return-EOPNOTSUPP;#endif-if(reg->hooknum!=NF_NETDEV_INGRESS||+if((reg->hooknum!=NF_NETDEV_INGRESS&&+reg->hooknum!=NF_NETDEV_EARLY_INGRESS)||!reg->dev||dev_net(reg->dev)!=net)return-EINVAL;}
@@ -347,6 +353,9 @@ static int __nf_register_net_hook(struct net *net, int pf,if(pf==NFPROTO_NETDEV&®->hooknum==NF_NETDEV_INGRESS)net_inc_ingress_queue();#endif+if(pf==NFPROTO_NETDEV&®->hooknum==NF_NETDEV_EARLY_INGRESS)+nf_early_ingress_enable();+#ifdef HAVE_JUMP_LABELstatic_key_slow_inc(&nf_hooks_needed[pf][reg->hooknum]);#endif
@@ -404,6 +413,9 @@ static void __nf_unregister_net_hook(struct net *net, int pf,#ifdef CONFIG_NETFILTER_INGRESSif(pf==NFPROTO_NETDEV&®->hooknum==NF_NETDEV_INGRESS)net_dec_ingress_queue();++if(pf==NFPROTO_NETDEV&®->hooknum==NF_NETDEV_EARLY_INGRESS)+nf_early_ingress_disable();#endif#ifdef HAVE_JUMP_LABELstatic_key_slow_dec(&nf_hooks_needed[pf][reg->hooknum]);
@@ -535,6 +547,27 @@ int nf_hook_slow(struct sk_buff *skb, struct nf_hook_state *state,}EXPORT_SYMBOL(nf_hook_slow);+intnf_hook_netdev(structsk_buff*skb,structnf_hook_state*state,+conststructnf_hook_entries*e)+{+unsignedintverdict,s,v=NF_ACCEPT;++for(s=0;s<e->num_hook_entries;s++){+verdict=nf_hook_entry_hookfn(&e->hooks[s],skb,state);+v=verdict&NF_VERDICT_MASK;+switch(v){+caseNF_ACCEPT:+break;+caseNF_DROP:+kfree_skb(skb);+/* Fall through */+default:+returnv;+}+}++returnv;+}intskb_make_writable(structsk_buff*skb,unsignedintwritable_len){
@@ -0,0 +1,323 @@+#include<linux/kernel.h>+#include<linux/netfilter.h>+#include<linux/types.h>+#include<net/xfrm.h>+#include<net/arp.h>+#include<net/udp.h>+#include<net/tcp.h>+#include<net/protocol.h>+#include<crypto/aead.h>+#include<net/netfilter/early_ingress.h>++/* XXX: Maybe export this from net/core/skbuff.c+*insteadofholdingalocalcopy*/+staticvoidskb_headers_offset_update(structsk_buff*skb,intoff)+{+/* Only adjust this if it actually is csum_start rather than csum */+if(skb->ip_summed==CHECKSUM_PARTIAL)+skb->csum_start+=off;+/* {transport,network,mac}_header and tail are relative to skb->head */+skb->transport_header+=off;+skb->network_header+=off;+if(skb_mac_header_was_set(skb))+skb->mac_header+=off;+skb->inner_transport_header+=off;+skb->inner_network_header+=off;+skb->inner_mac_header+=off;+}++structsk_buff*nft_skb_segment(structsk_buff*head_skb)+{+unsignedintheadroom;+structsk_buff*nskb;+structsk_buff*segs=NULL;+structsk_buff*tail=NULL;+unsignedintdoffset=head_skb->data-skb_mac_header(head_skb);+structsk_buff*list_skb=skb_shinfo(head_skb)->frag_list;+unsignedinttnl_hlen=skb_tnl_header_len(head_skb);+unsignedintdelta_segs,delta_len,delta_truesize;++__skb_push(head_skb,doffset);++headroom=skb_headroom(head_skb);++delta_segs=delta_len=delta_truesize=0;++skb_shinfo(head_skb)->frag_list=NULL;++segs=skb_clone(head_skb,GFP_ATOMIC);+if(unlikely(!segs))+returnERR_PTR(-ENOMEM);++do{+nskb=list_skb;++list_skb=list_skb->next;++if(!tail)+segs->next=nskb;+else+tail->next=nskb;++tail=nskb;++delta_len+=nskb->len;+delta_truesize+=nskb->truesize;++skb_push(nskb,doffset);++nskb->dev=head_skb->dev;+nskb->queue_mapping=head_skb->queue_mapping;+nskb->network_header=head_skb->network_header;+nskb->mac_len=head_skb->mac_len;+nskb->mac_header=head_skb->mac_header;+nskb->transport_header=head_skb->transport_header;++if(!secpath_exists(nskb))+nskb->sp=secpath_get(head_skb->sp);++skb_headers_offset_update(nskb,skb_headroom(nskb)-headroom);++skb_copy_from_linear_data_offset(head_skb,-tnl_hlen,+nskb->data-tnl_hlen,+doffset+tnl_hlen);++}while(list_skb);++segs->len=head_skb->len-delta_len;+segs->data_len=head_skb->data_len-delta_len;+segs->truesize+=head_skb->data_len-delta_truesize;++head_skb->len=segs->len;+head_skb->data_len=segs->data_len;+head_skb->truesize+=segs->truesize;++skb_shinfo(segs)->gso_size=0;+skb_shinfo(segs)->gso_segs=0;+skb_shinfo(segs)->gso_type=0;++segs->prev=tail;++returnsegs;+}++staticintnft_skb_gro_receive(structsk_buff**head,structsk_buff*skb)+{+structsk_buff*p=*head;++if(unlikely((!NAPI_GRO_CB(p)->is_ffwd)||!skb_dst(p)))+return-EINVAL;++if(NAPI_GRO_CB(p)->last==p)+skb_shinfo(p)->frag_list=skb;+else+NAPI_GRO_CB(p)->last->next=skb;+NAPI_GRO_CB(p)->last=skb;++NAPI_GRO_CB(p)->count++;+p->data_len+=skb->len;+p->truesize+=skb->truesize;+p->len+=skb->len;++NAPI_GRO_CB(skb)->same_flow=1;+return0;+}++staticstructsk_buff**udp_gro_ffwd_receive(structsk_buff**head,+structsk_buff*skb,+structudphdr*uh)+{+structsk_buff*p=NULL;+structsk_buff**pp=NULL;+structudphdr*uh2;+intflush=0;++for(;(p=*head);head=&p->next){++if(!NAPI_GRO_CB(p)->same_flow)+continue;++uh2=udp_hdr(p);++/* Match ports and either checksums are either both zero+*ornonzero.+*/+if((*(u32*)&uh->source!=*(u32*)&uh2->source)||+(!uh->check^!uh2->check)){+NAPI_GRO_CB(p)->same_flow=0;+continue;+}++gotofound;+}++gotoout;++found:+p=*head;++if(nft_skb_gro_receive(head,skb))+flush=1;++out:+if(p&&(!NAPI_GRO_CB(skb)->same_flow||flush))+pp=head;++NAPI_GRO_CB(skb)->flush|=flush;+returnpp;+}++structsk_buff**nft_udp_gro_receive(structsk_buff**head,structsk_buff*skb)+{+structudphdr*uh;++uh=skb_gro_header_slow(skb,skb_transport_offset(skb)+sizeof(structudphdr),+skb_transport_offset(skb));++if(unlikely(!uh))+gotoflush;++if(NAPI_GRO_CB(skb)->flush)+gotoflush;++if(NAPI_GRO_CB(skb)->is_ffwd)+returnudp_gro_ffwd_receive(head,skb,uh);++flush:+NAPI_GRO_CB(skb)->flush=1;+returnNULL;+}++structsk_buff**nft_tcp_gro_receive(structsk_buff**head,structsk_buff*skb)+{+structsk_buff**pp=NULL;+structsk_buff*p;+structtcphdr*th;+structtcphdr*th2;+unsignedintlen;+unsignedintthlen;+__be32flags;+unsignedintmss=1;+unsignedinthlen;+intflush=1;+inti;++th=skb_gro_header_slow(skb,skb_transport_offset(skb)+sizeof(structtcphdr),+skb_transport_offset(skb));+if(unlikely(!th))+gotoout;++thlen=th->doff*4;+if(thlen<sizeof(*th))+gotoout;++hlen=skb_transport_offset(skb)+thlen;++th=skb_gro_header_slow(skb,hlen,skb_transport_offset(skb));+if(unlikely(!th))+gotoout;++skb_gro_pull(skb,thlen);+len=skb_gro_len(skb);+flags=tcp_flag_word(th);++for(;(p=*head);head=&p->next){+if(!NAPI_GRO_CB(p)->same_flow)+continue;++th2=tcp_hdr(p);++if(*(u32*)&th->source^*(u32*)&th2->source){+NAPI_GRO_CB(p)->same_flow=0;+continue;+}++gotofound;+}++gotoout_check_final;++found:+flush=NAPI_GRO_CB(p)->flush;+flush|=(__forceint)(flags&TCP_FLAG_CWR);+flush|=(__forceint)((flags^tcp_flag_word(th2))&+~(TCP_FLAG_CWR|TCP_FLAG_FIN|TCP_FLAG_PSH));+flush|=(__forceint)(th->ack_seq^th2->ack_seq);+for(i=sizeof(*th);i<thlen;i+=4)+flush|=*(u32*)((u8*)th+i)^+*(u32*)((u8*)th2+i);++mss=skb_shinfo(p)->gso_size;++flush|=(len-1)>=mss;+flush|=(ntohl(th2->seq)+(skb_gro_len(p)-(hlen*(NAPI_GRO_CB(p)->count-1))))^ntohl(th->seq);++if(flush||nft_skb_gro_receive(head,skb)){+mss=1;+gotoout_check_final;+}++p=*head;++out_check_final:+flush=len<mss;+flush|=(__forceint)(flags&(TCP_FLAG_URG|TCP_FLAG_PSH|+TCP_FLAG_RST|TCP_FLAG_SYN|+TCP_FLAG_FIN));++if(p&&(!NAPI_GRO_CB(skb)->same_flow||flush))+pp=head;++out:+NAPI_GRO_CB(skb)->flush|=(flush!=0);++returnpp;+}++staticinlineboolnf_hook_early_ingress_active(conststructsk_buff*skb)+{+#ifdef HAVE_JUMP_LABEL+if(!static_key_false(&nf_hooks_needed[NFPROTO_NETDEV][NF_NETDEV_EARLY_INGRESS]))+returnfalse;+#endif+returnrcu_access_pointer(skb->dev->nf_hooks_early_ingress);+}++intnf_hook_early_ingress(structsk_buff*skb)+{+structnf_hook_entries*e=+rcu_dereference(skb->dev->nf_hooks_early_ingress);+structnf_hook_statestate;+intret=NF_ACCEPT;++if(nf_hook_early_ingress_active(skb)){+if(unlikely(!e))+return0;++nf_hook_state_init(&state,NF_NETDEV_EARLY_INGRESS,+NFPROTO_NETDEV,skb->dev,NULL,NULL,+dev_net(skb->dev),NULL);++ret=nf_hook_netdev(skb,&state,e);+}++returnret;+}++/* protected by nf_hook_mutex. */+staticintnf_early_ingress_use;++voidnf_early_ingress_enable(void)+{+if(nf_early_ingress_use++==0){+nf_early_ingress_use++;+nf_early_ingress_ip_enable();+}+}++voidnf_early_ingress_disable(void)+{+if(--nf_early_ingress_use==0){+nf_early_ingress_ip_disable();+}+}
@@ -2,6 +2,7 @@## Makefile for the netfilter modules on top of IPv6.#+obj-$(CONFIG_NETFILTER_EARLY_INGRESS)+=early_ingress.o# Link order matters here.obj-$(CONFIG_IP6_NF_IPTABLES)+=ip6_tables.o
@@ -0,0 +1,307 @@+#include<linux/kernel.h>+#include<linux/netfilter.h>+#include<linux/types.h>+#include<net/xfrm.h>+#include<net/arp.h>+#include<net/udp.h>+#include<net/tcp.h>+#include<net/protocol.h>+#include<net/netfilter/early_ingress.h>+#include<net/ip6_route.h>++staticconststructnet_offload__rcu*nft_ip6_offloads[MAX_INET_PROTOS]__read_mostly;++staticstructsk_buff*nft_udp6_gso_segment(structsk_buff*skb,+netdev_features_tfeatures)+{+skb_push(skb,sizeof(structipv6hdr));+returnnft_skb_segment(skb);+}++staticstructsk_buff*nft_tcp6_gso_segment(structsk_buff*skb,+netdev_features_tfeatures)+{+skb_push(skb,sizeof(structipv6hdr));+returnnft_skb_segment(skb);+}++staticstructsk_buff*nft_ipv6_gso_segment(structsk_buff*skb,+netdev_features_tfeatures)+{+structsk_buff*segs=ERR_PTR(-EINVAL);+conststructnet_offload*ops;+structpacket_offload*ptype;+structipv6hdr*iph;+intproto;++if(!(skb_shinfo(skb)->gso_type&SKB_GSO_NFT)){+ptype=dev_get_packet_offload(skb->protocol,1);+if(ptype)+returnptype->callbacks.gso_segment(skb,features);++returnERR_PTR(-EPROTONOSUPPORT);+}++if(SKB_GSO_CB(skb)->encap_level==0){+iph=ipv6_hdr(skb);+skb_reset_network_header(skb);+}else{+iph=(structipv6hdr*)skb->data;+}++if(unlikely(!pskb_may_pull(skb,sizeof(*iph))))+gotoout;++SKB_GSO_CB(skb)->encap_level+=sizeof(*iph);++if(unlikely(!pskb_may_pull(skb,sizeof(*iph))))+gotoout;++__skb_pull(skb,sizeof(*iph));++proto=iph->nexthdr;++segs=ERR_PTR(-EPROTONOSUPPORT);++ops=rcu_dereference(nft_ip6_offloads[proto]);+if(likely(ops&&ops->callbacks.gso_segment))+segs=ops->callbacks.gso_segment(skb,features);++out:+returnsegs;+}++staticintnft_ipv6_gro_complete(structsk_buff*skb,intnhoff)+{+structipv6hdr*iph=(structipv6hdr*)(skb->data+nhoff);+structdst_entry*dst=skb_dst(skb);+structrt6_info*rt=(structrt6_info*)dst;+conststructnet_offload*ops;+structpacket_offload*ptype;+intproto=iph->nexthdr;+structin6_addr*nexthop;+structneighbour*neigh;+structnet_device*dev;+unsignedinthh_len;+interr=0;+u16count;++count=NAPI_GRO_CB(skb)->count;++if(!NAPI_GRO_CB(skb)->is_ffwd){+ptype=dev_get_packet_offload(skb->protocol,1);+if(ptype)+returnptype->callbacks.gro_complete(skb,nhoff);++return0;+}++rcu_read_lock();+ops=rcu_dereference(nft_ip6_offloads[proto]);+if(!ops||!ops->callbacks.gro_complete)+gotoout_unlock;++/* Only need to add sizeof(*iph) to get to the next hdr below+*becauseanyhdrwithoptionwillhavebeenflushedin+*inet_gro_receive().+*/+err=ops->callbacks.gro_complete(skb,nhoff+sizeof(*iph));++out_unlock:+rcu_read_unlock();++if(err)+returnerr;++skb_shinfo(skb)->gso_type|=SKB_GSO_NFT;+skb_shinfo(skb)->gso_segs=count;++dev=dst->dev;+dev_hold(dev);+skb->dev=dev;++if(skb_dst(skb)->xfrm){+err=dst_output(dev_net(dev),NULL,skb);+if(err!=-EREMOTE)+return-EINPROGRESS;+}++if(count<=1)+skb_gso_reset(skb);++hh_len=LL_RESERVED_SPACE(dev);++if(unlikely(skb_headroom(skb)<hh_len&&dev->header_ops)){+structsk_buff*skb2;++skb2=skb_realloc_headroom(skb,LL_RESERVED_SPACE(dev));+if(!skb2){+kfree_skb(skb);+return-ENOMEM;+}+consume_skb(skb);+skb=skb2;+}+rcu_read_lock();+nexthop=rt6_nexthop(rt,&iph->daddr);+neigh=__ipv6_neigh_lookup_noref(dev,nexthop);+if(unlikely(!neigh))+neigh=__neigh_create(&arp_tbl,&nexthop,dev,false);+if(!IS_ERR(neigh))+neigh_output(neigh,skb);+rcu_read_unlock();++return-EINPROGRESS;+}++staticstructsk_buff**nft_ipv6_gro_receive(structsk_buff**head,+structsk_buff*skb)+{+conststructnet_offload*ops;+structpacket_offload*ptype;+structsk_buff**pp=NULL;+structsk_buff*p;+structipv6hdr*iph;+unsignedintnlen;+unsignedinthlen;+unsignedintoff;+intproto,ret;++off=skb_gro_offset(skb);+hlen=off+sizeof(*iph);++iph=skb_gro_header_slow(skb,hlen,off);+if(unlikely(!iph))+gotoout;++proto=iph->nexthdr;++rcu_read_lock();++if(iph->version!=6)+gotoout_unlock;++nlen=skb_network_header_len(skb);++ret=nf_hook_early_ingress(skb);+switch(ret){+caseNF_STOLEN:+break;+caseNF_ACCEPT:+ptype=dev_get_packet_offload(skb->protocol,1);+if(ptype)+pp=ptype->callbacks.gro_receive(head,skb);++gotoout_unlock;+caseNF_DROP:+pp=ERR_PTR(-EPERM);+gotoout_unlock;+}++ops=rcu_dereference(nft_ip6_offloads[proto]);+if(!ops||!ops->callbacks.gro_receive)+gotoout_unlock;++if(iph->hop_limit<=1)+gotoout_unlock;++skb->ip_summed=CHECKSUM_UNNECESSARY;++for(p=*head;p;p=p->next){+structipv6hdr*iph2;+__be32first_word;/* <Version:4><Traffic_Class:8><Flow_Label:20> */++if(!NAPI_GRO_CB(p)->same_flow)+continue;++if(!NAPI_GRO_CB(p)->is_ffwd){+NAPI_GRO_CB(p)->same_flow=0;+continue;+}++if(!skb_dst(p)){+NAPI_GRO_CB(p)->same_flow=0;+continue;+}++iph2=ipv6_hdr(p);+first_word=*(__be32*)iph^*(__be32*)iph2;++/* All fields must match except length and Traffic Class.+*XXXskbsonthegro_listhaveallbeenparsedandpulled+*alreadysowedon'tneedtocomparenlen+*(nlen!=(sizeof(*iph2)+ipv6_exthdrs_len(iph2,&ops)))+*memcmp()alonebelowissuffcient,right?+*/+if((first_word&htonl(0xF00FFFFF))||+memcmp(&iph->nexthdr,&iph2->nexthdr,+nlen-offsetof(structipv6hdr,nexthdr))){+NAPI_GRO_CB(p)->same_flow=0;+continue;+}+/* flush if Traffic Class fields are different */+NAPI_GRO_CB(p)->flush|=!!(first_word&htonl(0x0FF00000));++NAPI_GRO_CB(skb)->is_ffwd=1;+skb_dst_set_noref(skb,skb_dst(p));+pp=&p;++break;+}++NAPI_GRO_CB(skb)->is_atomic=true;++iph->hop_limit--;++skb_pull(skb,off);+NAPI_GRO_CB(skb)->data_offset=sizeof(*iph);+skb_reset_network_header(skb);+skb_set_transport_header(skb,sizeof(*iph));++pp=call_gro_receive(ops->callbacks.gro_receive,head,skb);+out_unlock:+rcu_read_unlock();++out:+NAPI_GRO_CB(skb)->data_offset=0;+returnpp;+}++staticstructpacket_offloadnft_ip6_packet_offload__read_mostly={+.type=cpu_to_be16(ETH_P_IPV6),+.priority=0,+.callbacks={+.gro_receive=nft_ipv6_gro_receive,+.gro_complete=nft_ipv6_gro_complete,+.gso_segment=nft_ipv6_gso_segment,+},+};++staticconststructnet_offloadnft_udp6_offload={+.callbacks={+.gso_segment=nft_udp6_gso_segment,+.gro_receive=nft_udp_gro_receive,+},+};++staticconststructnet_offloadnft_tcp6_offload={+.callbacks={+.gso_segment=nft_tcp6_gso_segment,+.gro_receive=nft_tcp_gro_receive,+},+};++staticconststructnet_offload__rcu*nft_ip6_offloads[MAX_INET_PROTOS]__read_mostly={+[IPPROTO_UDP]=&nft_udp6_offload,+[IPPROTO_TCP]=&nft_tcp6_offload,+};++voidnf_early_ingress_ip6_enable(void)+{+dev_add_offload(&nft_ip6_packet_offload);+}++voidnf_early_ingress_ip6_disable(void)+{+dev_remove_offload(&nft_ip6_packet_offload);+}
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:21
From: Steffen Klassert <steffen.klassert@secunet.com>
This patch adds the GSO logic for ESP and the codepath that allows
the xfrm infrastructure to signal the GRO layer that the packet is
following the fast forwarding path.
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/net/netfilter/early_ingress.h | 2 ++
net/ipv4/netfilter/early_ingress.c | 8 ++++++++
net/ipv6/netfilter/early_ingress.c | 8 ++++++++
net/netfilter/early_ingress.c | 36 +++++++++++++++++++++++++++++++++++
net/xfrm/xfrm_output.c | 4 ++++
5 files changed, 58 insertions(+)
@@ -146,6 +146,10 @@ int xfrm_output_resume(struct sk_buff *skb, int err)while(likely((err=xfrm_output_one(skb,err))==0)){nf_reset(skb);+if(!skb_dst(skb)->xfrm&&skb->sp&&+(skb_shinfo(skb)->gso_type&SKB_GSO_NFT))+return-EREMOTE;+err=skb_dst(skb)->ops->local_out(net,skb->sk,skb);if(unlikely(err!=1))gotoout;
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:22
This patch adds the new filter chain at the early ingress hook.
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
---
net/netfilter/nft_chain_filter.c | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:28
Add the new flowtable type for the early ingress hook, this allows
us to combine the custom GRO chaining with the flowtable abstraction
to define fastpaths.
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
---
include/net/netfilter/nf_flow_table.h | 3 ++
net/ipv4/netfilter/nf_flow_table_ipv4.c | 11 ++++++
net/netfilter/nf_flow_table_ip.c | 62 +++++++++++++++++++++++++++++++++
3 files changed, 76 insertions(+)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:32
Once we have a confirmed conntrack, ie. a packet went through the stack
and a conntrack was added, then allow second packet to configure the
flowtable offload.
This allows UDP media traffic going in only one direction to enable offloads.
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
---
net/netfilter/nft_flow_offload.c | 11 +++--------
1 file changed, 3 insertions(+), 8 deletions(-)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:34
It is safe to place a flow that is coming from IPSec into the flowtable.
So decapsulated can benefit from the flowtable fastpath.
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
---
net/netfilter/nft_flow_offload.c | 2 --
1 file changed, 2 deletions(-)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2018-06-14 14:20:53
Use dst_check() to validate that route is still valid, otherwise,
tear down the flow entry and pass up packet to the standard forwarding
path so we have a chance to cache the fresh route again.
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com>
---
net/netfilter/nf_flow_table_ip.c | 10 ++++++++++
1 file changed, 10 insertions(+)
From: Willem de Bruijn <willemdebruijn.kernel@gmail.com> Date: 2018-06-14 15:51:26
This patchset supports both layer 3 IPv4 and IPv6, and layer 4 TCP and
UDP protocols. This fastpath also integrates with the IPSec
infrastructure and the ESP protocol.
We have collected performance numbers:
TCP TSO TCP Fast Forward
32.5 Gbps 35.6 Gbps
UDP UDP Fast Forward
17.6 Gbps 35.6 Gbps
ESP ESP Fast Forward
6 Gbps 7.5 Gbps
For UDP, this is doubling performance, and we almost achieve line rate
with one single CPU using the Intel i40e NIC. We got similar numbers
with the Mellanox ConnectX-4. For TCP, this is slightly improving things
even if TSO is being defeated given that we need to segment the packet
chain in software.
The difference between TCP and UDP stems from lack of GRO for UDP. We
recently added UDP GSO to allow for batch traversal of the UDP stack on
transmission. Adding a UDP GRO handler can probably extend batching to
the forwarding path in a similar way without the need for a new infrastructure.
From: Eric Dumazet <hidden> Date: 2018-06-14 15:57:56
On 06/14/2018 07:19 AM, Pablo Neira Ayuso wrote:
Hi,
We have collected performance numbers:
TCP TSO TCP Fast Forward
32.5 Gbps 35.6 Gbps
UDP UDP Fast Forward
17.6 Gbps 35.6 Gbps
ESP ESP Fast Forward
6 Gbps 7.5 Gbps
For UDP, this is doubling performance, and we almost achieve line rate
with one single CPU using the Intel i40e NIC. We got similar numbers
with the Mellanox ConnectX-4. For TCP, this is slightly improving things
even if TSO is being defeated given that we need to segment the packet
chain in software. We would like to explore HW GRO support with hardware
vendors with this new mode, we think that should improve the TCP numbers
we are showing above even more.
Hi Pablo
Not very convincing numbers, because it is unclear what traffic patterns were used.
We normally use packets per second to measure a forwarding workload,
and it is not clear if you tried a DDOS, or/and a mix of packets being locally
delivered and packets being forwarded.
Presumably adding cache line misses (to probe for flows) will slow down the things.
I suspect the NIC you use has some kind of bottleneck on sending TSO packets,
or that you hit the issue that GRO might cook suboptimal packets for forwarding workloads
(eg setting frag_list)
This path series add yet more code to GRO engine which is already very fat
to the point many people advocate to turn it off.
Saving cpu cycles on moderate load is not okay if added complexity
slows down the DDOS (or stress) by 10 % :/
To me, GRO is specialized to optimize the non-forwarding case,
so it is counter-intuitive to base a fast forwarding path on top of it.
From: David Miller <davem@davemloft.net> Date: 2018-06-14 17:18:34
From: Pablo Neira Ayuso <pablo@netfilter.org>
Date: Thu, 14 Jun 2018 16:19:34 +0200
This patchset proposes a new fast forwarding path infrastructure
that combines the GRO/GSO and the flowtable infrastructures. The
idea is to add a hook at the GRO layer that is invoked before the
standard GRO protocol offloads. This allows us to build custom
packet chains that we can quickly pass in one go to the neighbour
layer to define fast forwarding path for flows.
We have full, complete, customizability of the packet path via XDP
and eBPF.
XDP and eBPF supports everything necessary to accomplish that,
there are implementations of forwarding implementations in
the tree and elsewhere.
And most importantly, XDP and eBPF are optimized in drivers and
offloaded to hardware.
There really is no need for something like what you are proposing.
From: Pablo Neira Ayuso <pablo@netfilter.org>
Date: Thu, 14 Jun 2018 16:19:34 +0200
quoted
This patchset proposes a new fast forwarding path infrastructure
that combines the GRO/GSO and the flowtable infrastructures. The
idea is to add a hook at the GRO layer that is invoked before the
standard GRO protocol offloads. This allows us to build custom
packet chains that we can quickly pass in one go to the neighbour
layer to define fast forwarding path for flows.
We have full, complete, customizability of the packet path via XDP
and eBPF.
XDP and eBPF supports everything necessary to accomplish that,
there are implementations of forwarding implementations in
the tree and elsewhere.
And most importantly, XDP and eBPF are optimized in drivers and
offloaded to hardware.
There really is no need for something like what you are proposing.
I see one possible upside to that approach here which is the low end
MIPS/ARM/PowerPC 32-bit based routers that do not have an eBPF JIT
available (that's only MIPS32 and PowerPC AFAICT), it would be great to
see what happens on those systems and if we do get any performance
improvements for a traditional forwarding/routing workload. On those
platforms there are a number of things that just literally kill the
routing performance: small I and D caches, small or not L2, limited
bandwidth DRAM, huge call depths, big struct sk_buff layout, you name it.
--
Florian
From: Tom Herbert <hidden> Date: 2018-06-14 20:52:04
On Thu, Jun 14, 2018 at 7:19 AM, Pablo Neira Ayuso [off-list ref] wrote:
Hi,
This patchset proposes a new fast forwarding path infrastructure that
combines the GRO/GSO and the flowtable infrastructures. The idea is to
add a hook at the GRO layer that is invoked before the standard GRO
protocol offloads. This allows us to build custom packet chains that we
can quickly pass in one go to the neighbour layer to define fast
forwarding path for flows.
For each packet that gets into the GRO layer, we first check if there is
an entry in the flowtable, if so, the packet is placed in a list until
the GRO infrastructure decides to send the batch from gro_complete to
the neighbour layer. The first packet in the list takes the route from
the flowtable entry, so we avoid reiterative routing lookups.
In case no entry is found in the flowtable, the packet is passed up to
the classic GRO offload handlers. Thus, this packet follows the standard
forwarding path. Note that the initial packets of the flow always go
through the standard IPv4/IPv6 netfilter forward hook, that is used to
configure what flows are placed in the flowtable. Therefore, only a few
(initial) packets follow the standard forwarding path while most of the
follow up packets take this new fast forwarding path.
IIRC, there was a similar proposal a while back that want to bundle
packets of the same flow together (without doing GRO) so that they
could be processed by various functions by looking at just one
representative packet in the group. The concept had some promise, but
in the end it created quite a bit of complexity since at some point
the packet bundle needed to be undone to go back to processing the
individual packets.
Tom
The fast forwarding path is enabled through explicit user policy, so the
user needs to request this behaviour from control plane, the following
example shows how to place flows in the new fast forwarding path from
the netfilter forward chain:
table x {
flowtable f {
hook early_ingress priority 0; devices = { eth0, eth1 }
}
chain y {
type filter hook forward priority 0;
ip protocol tcp flow offload @f
}
}
The example above defines a fastpath for TCP flows that are placed in
the flowtable 'f', this flowtable is hooked at the new early_ingress
hook. The initial TCP packets that match this rule from the standard
fowarding path create an entry in the flowtable, thus, GRO creates chain
of packets for those that find an entry in the flowtable and send
them through the neighbour layer.
This new hook is happening before the ingress taps, therefore, packets
that follow this new fast forwarding path are not shown by tcpdump.
This patchset supports both layer 3 IPv4 and IPv6, and layer 4 TCP and
UDP protocols. This fastpath also integrates with the IPSec
infrastructure and the ESP protocol.
We have collected performance numbers:
TCP TSO TCP Fast Forward
32.5 Gbps 35.6 Gbps
UDP UDP Fast Forward
17.6 Gbps 35.6 Gbps
ESP ESP Fast Forward
6 Gbps 7.5 Gbps
For UDP, this is doubling performance, and we almost achieve line rate
with one single CPU using the Intel i40e NIC. We got similar numbers
with the Mellanox ConnectX-4. For TCP, this is slightly improving things
even if TSO is being defeated given that we need to segment the packet
chain in software. We would like to explore HW GRO support with hardware
vendors with this new mode, we think that should improve the TCP numbers
we are showing above even more. For ESP traffic, performance improvement
is ~25%, in this case, perf shows the bottleneck becomes the crypto layer.
This patchset is co-authored work with Steffen Klassert.
Comments are welcome, thanks.
Pablo Neira Ayuso (6):
netfilter: nft_chain_filter: add support for early ingress
netfilter: nf_flow_table: add hooknum to flowtable type
netfilter: nf_flow_table: add flowtable for early ingress hook
netfilter: nft_flow_offload: enable offload after second packet is seen
netfilter: nft_flow_offload: remove secpath check
netfilter: nft_flow_offload: make sure route is not stale
Steffen Klassert (7):
net: Add a helper to get the packet offload callbacks by priority.
net: Change priority of ipv4 and ipv6 packet offloads.
net: Add a GSO feature bit for the netfilter forward fastpath.
net: Use one bit of NAPI_GRO_CB for the netfilter fastpath.
netfilter: add early ingress hook for IPv4
netfilter: add early ingress support for IPv6
netfilter: add ESP support for early ingress
include/linux/netdev_features.h | 4 +-
include/linux/netdevice.h | 6 +-
include/linux/netfilter.h | 6 +
include/linux/netfilter_ingress.h | 1 +
include/linux/skbuff.h | 2 +
include/net/netfilter/early_ingress.h | 24 +++
include/net/netfilter/nf_flow_table.h | 4 +
include/uapi/linux/netfilter.h | 1 +
net/core/dev.c | 50 ++++-
net/ipv4/af_inet.c | 1 +
net/ipv4/netfilter/Makefile | 1 +
net/ipv4/netfilter/early_ingress.c | 327 +++++++++++++++++++++++++++++
net/ipv4/netfilter/nf_flow_table_ipv4.c | 12 ++
net/ipv6/ip6_offload.c | 1 +
net/ipv6/netfilter/Makefile | 1 +
net/ipv6/netfilter/early_ingress.c | 315 ++++++++++++++++++++++++++++
net/ipv6/netfilter/nf_flow_table_ipv6.c | 1 +
net/netfilter/Kconfig | 8 +
net/netfilter/Makefile | 1 +
net/netfilter/core.c | 35 +++-
net/netfilter/early_ingress.c | 361 ++++++++++++++++++++++++++++++++
net/netfilter/nf_flow_table_inet.c | 1 +
net/netfilter/nf_flow_table_ip.c | 72 +++++++
net/netfilter/nf_tables_api.c | 120 ++++++-----
net/netfilter/nft_chain_filter.c | 6 +-
net/netfilter/nft_flow_offload.c | 13 +-
net/xfrm/xfrm_output.c | 4 +
27 files changed, 1297 insertions(+), 81 deletions(-)
create mode 100644 include/net/netfilter/early_ingress.h
create mode 100644 net/ipv4/netfilter/early_ingress.c
create mode 100644 net/ipv6/netfilter/early_ingress.c
create mode 100644 net/netfilter/early_ingress.c
--
2.11.0
On those platforms there are a number of things that just literally
kill the routing performance: small I and D caches, small or not L2,
limited bandwidth DRAM, huge call depths, big struct sk_buff layout,
you name it.
Another reason to work on a 64-bit MIPS eBPF JIT.
We have a model, and game plan for this kind of application. And it's
XDP and eBPF with JITs.
We are fully commited to this approach, and I see anything else that
tries to slip in and approach some sub-part of the problem as a
complete distraction and a step backwards.
All of the effort on this work could have been spent filling in the
missing pieces you mention.
And guess what? Then millions of possibilities would have been
openned up, rather than just this one special case.
So, I ask, please see the larger picture.
Thank you.
From: David Miller <davem@davemloft.net> Date: 2018-06-14 23:58:35
From: Tom Herbert <redacted>
Date: Thu, 14 Jun 2018 13:52:03 -0700
IIRC, there was a similar proposal a while back that want to bundle
packets of the same flow together (without doing GRO) so that they
could be processed by various functions by looking at just one
representative packet in the group. The concept had some promise, but
in the end it created quite a bit of complexity since at some point
the packet bundle needed to be undone to go back to processing the
individual packets.
You're probably talking about Edward Cree's SKB list stuff, and as
per his presenation at netconf 2 weeks ago he plans to revitalize
it given how Spectre et al. gives cause to reevaluate all bulking
techniques.
On Thu, Jun 14, 2018 at 11:50:49AM -0400, Willem de Bruijn wrote:
quoted
This patchset supports both layer 3 IPv4 and IPv6, and layer 4 TCP and
UDP protocols. This fastpath also integrates with the IPSec
infrastructure and the ESP protocol.
We have collected performance numbers:
TCP TSO TCP Fast Forward
32.5 Gbps 35.6 Gbps
UDP UDP Fast Forward
17.6 Gbps 35.6 Gbps
ESP ESP Fast Forward
6 Gbps 7.5 Gbps
For UDP, this is doubling performance, and we almost achieve line rate
with one single CPU using the Intel i40e NIC. We got similar numbers
with the Mellanox ConnectX-4. For TCP, this is slightly improving things
even if TSO is being defeated given that we need to segment the packet
chain in software.
The difference between TCP and UDP stems from lack of GRO for UDP.
Right.
We
recently added UDP GSO to allow for batch traversal of the UDP stack on
transmission. Adding a UDP GRO handler can probably extend batching to
the forwarding path in a similar way without the need for a new infrastructure.
That's more or less what we did. The batching method ist just
optimized for the forwarding path. We are generating skb chains
by chaning at the frag_list pointer of the first skb. With that,
we don't need to mange packet. We keep the packets in the native
form, so the 'segmentation' is rather easy.
The rest is just to be able to configure this and to make
sure that we handle only flows that are going to be (fast)
forwarded, as the upper stack can not (yet) handle such
skb chains.
On Thu, Jun 14, 2018 at 08:57:20AM -0700, Eric Dumazet wrote:
On 06/14/2018 07:19 AM, Pablo Neira Ayuso wrote:
quoted
Hi,
quoted
We have collected performance numbers:
TCP TSO TCP Fast Forward
32.5 Gbps 35.6 Gbps
UDP UDP Fast Forward
17.6 Gbps 35.6 Gbps
ESP ESP Fast Forward
6 Gbps 7.5 Gbps
For UDP, this is doubling performance, and we almost achieve line rate
with one single CPU using the Intel i40e NIC. We got similar numbers
with the Mellanox ConnectX-4. For TCP, this is slightly improving things
even if TSO is being defeated given that we need to segment the packet
chain in software. We would like to explore HW GRO support with hardware
vendors with this new mode, we think that should improve the TCP numbers
we are showing above even more.
Hi Pablo
Not very convincing numbers, because it is unclear what traffic patterns were used.
We normally use packets per second to measure a forwarding workload,
and it is not clear if you tried a DDOS, or/and a mix of packets being locally
delivered and packets being forwarded.
Yes, these number need some more explaination. We used my IPsec
forwarding test setup for this. It looks like this:
------------ ------------
-->| router 1 |-------->| router 2 |--
| ------------ ------------ |
| |
| -------------------- |
--------|Spirent Testcenter|<----------
--------------------
The numbers are from single stream forwarding tests, no local delivery.
Packet size in the UDP case was 1460 byte. I used this packet size
because such packets still fit into the mtu when encapsulated by IPsec.
Presumably adding cache line misses (to probe for flows) will slow down the things.
I suspect the NIC you use has some kind of bottleneck on sending TSO packets,
or that you hit the issue that GRO might cook suboptimal packets for forwarding workloads
(eg setting frag_list)
That might be, I was a bit surprised about the TCP numbers myself.
I was more focused on UDP and IPsec because these don't have
hardware segmentation support. I've just added a TCP handler to
see what happens, the numbers looked ok, so I kept it.
All this is based on the approach I pesented last year at the nefilter
workshop.
This path series add yet more code to GRO engine which is already very fat
to the point many people advocate to turn it off.
We tried to stay away from the generic codepath as much as possible.
Currently we need five 'if' statements, two of them are in error
paths (Patch 4).
Saving cpu cycles on moderate load is not okay if added complexity
slows down the DDOS (or stress) by 10 % :/
Why 10%?
To me, GRO is specialized to optimize the non-forwarding case,
so it is counter-intuitive to base a fast forwarding path on top of it.
It is optimized for the non-forwarding case, but it seems that forwarding
can benefit from that too with very little cost for the non-forwarding case.
On Thu, Jun 14, 2018 at 10:18:31AM -0700, David Miller wrote:
From: Pablo Neira Ayuso <pablo@netfilter.org>
Date: Thu, 14 Jun 2018 16:19:34 +0200
quoted
This patchset proposes a new fast forwarding path infrastructure
that combines the GRO/GSO and the flowtable infrastructures. The
idea is to add a hook at the GRO layer that is invoked before the
standard GRO protocol offloads. This allows us to build custom
packet chains that we can quickly pass in one go to the neighbour
layer to define fast forwarding path for flows.
We have full, complete, customizability of the packet path via XDP
and eBPF.
XDP and eBPF supports everything necessary to accomplish that,
there are implementations of forwarding implementations in
the tree and elsewhere.
And most importantly, XDP and eBPF are optimized in drivers and
offloaded to hardware.
There really is no need for something like what you are proposing.
I started with this last year because I wanted to improve
the IPsec (and UDP) forwarding path. Batching packets
at layer2 and send them directly to the output path
seemed to be a good method to improve this.
In particular, we need to do only one IPsec lookup
for the whole packet chain. So it relaxes the pain
from reomoving the IPsec flowcache a bit. It can be
only a first step, but we need some improvements here
as people start to complain about that.
On Thu, Jun 14, 2018 at 01:52:03PM -0700, Tom Herbert wrote:
On Thu, Jun 14, 2018 at 7:19 AM, Pablo Neira Ayuso [off-list ref] wrote:
quoted
Hi,
This patchset proposes a new fast forwarding path infrastructure that
combines the GRO/GSO and the flowtable infrastructures. The idea is to
add a hook at the GRO layer that is invoked before the standard GRO
protocol offloads. This allows us to build custom packet chains that we
can quickly pass in one go to the neighbour layer to define fast
forwarding path for flows.
For each packet that gets into the GRO layer, we first check if there is
an entry in the flowtable, if so, the packet is placed in a list until
the GRO infrastructure decides to send the batch from gro_complete to
the neighbour layer. The first packet in the list takes the route from
the flowtable entry, so we avoid reiterative routing lookups.
In case no entry is found in the flowtable, the packet is passed up to
the classic GRO offload handlers. Thus, this packet follows the standard
forwarding path. Note that the initial packets of the flow always go
through the standard IPv4/IPv6 netfilter forward hook, that is used to
configure what flows are placed in the flowtable. Therefore, only a few
(initial) packets follow the standard forwarding path while most of the
follow up packets take this new fast forwarding path.
IIRC, there was a similar proposal a while back that want to bundle
packets of the same flow together (without doing GRO) so that they
could be processed by various functions by looking at just one
representative packet in the group. The concept had some promise, but
in the end it created quite a bit of complexity since at some point
the packet bundle needed to be undone to go back to processing the
individual packets.
With the way we chain the packets it is not too complicated to
undo this chaining (nft_skb_segment in patch 5 implements this).
After that, this looks like a chain of usual segments, so we
trigger xmit_more with every packet chain.
On Thu, Jun 14, 2018 at 04:58:34PM -0700, David Miller wrote:
From: Tom Herbert <redacted>
Date: Thu, 14 Jun 2018 13:52:03 -0700
quoted
IIRC, there was a similar proposal a while back that want to bundle
packets of the same flow together (without doing GRO) so that they
could be processed by various functions by looking at just one
representative packet in the group. The concept had some promise, but
in the end it created quite a bit of complexity since at some point
the packet bundle needed to be undone to go back to processing the
individual packets.
You're probably talking about Edward Cree's SKB list stuff, and as
per his presenation at netconf 2 weeks ago he plans to revitalize
it given how Spectre et al. gives cause to reevaluate all bulking
techniques.
Are there patches for the proposal Edward did a while ago,
or was it just a concept?
Maybe we can somehow put things together, I just need some
batching method that works for IPsec and UDP. It does not
need to be exactly the one we proposing here.
From: Edward Cree <hidden> Date: 2018-06-15 12:19:04
On 15/06/18 07:34, Steffen Klassert wrote:
On Thu, Jun 14, 2018 at 04:58:34PM -0700, David Miller wrote:
quoted
You're probably talking about Edward Cree's SKB list stuff, and as
per his presenation at netconf 2 weeks ago he plans to revitalize
it given how Spectre et al. gives cause to reevaluate all bulking
techniques.
Are there patches for the proposal Edward did a while ago,
or was it just a concept?
Old patches are at http://lists.openwall.net/netdev/2016/04/19/89
(note that the absolute numbers I give in the cover letter are wrong;
I quoted them as Mpps but they're actually Mbps which is 8x higher).
I hope to have a new series ready shortly after net-next reopens.
-Ed
From: Eric Dumazet <hidden> Date: 2018-06-15 13:01:29
On 06/14/2018 11:03 PM, Steffen Klassert wrote:
On Thu, Jun 14, 2018 at 08:57:20AM -0700, Eric Dumazet wrote:
quoted
quoted
Saving cpu cycles on moderate load is not okay if added complexity
slows down the DDOS (or stress) by 10 % :/
Why 10%?
GRO adds a ~6 % cost on UDP receive path at this moment, depending on the state
of GRO engine (number of packets in the napi->gro_list)
Adding yet another conditions and icache pressure might raise the cost to 10%,
but we do not know because the numbers presented in this RFC do not include that.
(Early demux is also adding extra costs for UDP on 'non connected sockets' BTW)
Most linux hosts are not routers, but end hosts, lets not forget this...
From: Daniel Borkmann <daniel@iogearbox.net> Date: 2018-06-15 13:22:29
Hi Steffen,
On 06/15/2018 08:17 AM, Steffen Klassert wrote:
On Thu, Jun 14, 2018 at 10:18:31AM -0700, David Miller wrote:
quoted
From: Pablo Neira Ayuso <pablo@netfilter.org>
Date: Thu, 14 Jun 2018 16:19:34 +0200
quoted
This patchset proposes a new fast forwarding path infrastructure
that combines the GRO/GSO and the flowtable infrastructures. The
idea is to add a hook at the GRO layer that is invoked before the
standard GRO protocol offloads. This allows us to build custom
packet chains that we can quickly pass in one go to the neighbour
layer to define fast forwarding path for flows.
We have full, complete, customizability of the packet path via XDP
and eBPF.
XDP and eBPF supports everything necessary to accomplish that,
there are implementations of forwarding implementations in
the tree and elsewhere.
And most importantly, XDP and eBPF are optimized in drivers and
offloaded to hardware.
There really is no need for something like what you are proposing.
I started with this last year because I wanted to improve
the IPsec (and UDP) forwarding path. Batching packets
at layer2 and send them directly to the output path
seemed to be a good method to improve this.
In particular, we need to do only one IPsec lookup
for the whole packet chain. So it relaxes the pain
from reomoving the IPsec flowcache a bit. It can be
only a first step, but we need some improvements here
as people start to complain about that.
But did you also experiment with XDP on this? Would be curious about
the numbers. You'd get implicit batching for the forwarding via devmap
as well if you're required to flush it out via different device with
XDP_REDIRECT; otherwise XDP_TX of course. Given we have recently
integrated helpers for XDP to do a FIB and neighbor lookup from the
kernel tables, where it's thus shared and integrated with the rest of
the stack and tooling, it would be awesome to get to the same point
with xfrm as well. Eyal recently did a start on that for xfrm for tc
progs; would be nice to have integration on XDP as well, potentially
it might also result in a bigger plus on the forwarding numbers.
Thanks,
Daniel
From: Tom Herbert <hidden> Date: 2018-06-15 20:12:30
On Thu, Jun 14, 2018 at 4:58 PM, David Miller [off-list ref] wrote:
From: Tom Herbert <redacted>
Date: Thu, 14 Jun 2018 13:52:03 -0700
quoted
IIRC, there was a similar proposal a while back that want to bundle
packets of the same flow together (without doing GRO) so that they
could be processed by various functions by looking at just one
representative packet in the group. The concept had some promise, but
in the end it created quite a bit of complexity since at some point
the packet bundle needed to be undone to go back to processing the
individual packets.
You're probably talking about Edward Cree's SKB list stuff, and as
per his presenation at netconf 2 weeks ago he plans to revitalize
it given how Spectre et al. gives cause to reevaluate all bulking
techniques.nearly
The use case for that will be an interesting question. GSO/GRO solves
the problem for TCP and this extends to nearly all cases where TCP is
in an encapsulated packet. Super efficient forwarding can be done in
XDP/BPF (without needing overhead of GSO/GRO). That pretty much leaves
UDP as non-encapsulation end protocol, which I guess these days pretty
much means QUIC :-) I am still interested to see if we can implement
GSO/GRO for QUIC (via a generic GSO/GRO BPF function so we don't
hardcode any QUIC protocol or other application protocols in kernel).
Tom
Hi Daniel,
On Fri, Jun 15, 2018 at 03:22:24PM +0200, Daniel Borkmann wrote:
Hi Steffen,
On 06/15/2018 08:17 AM, Steffen Klassert wrote:
quoted
I started with this last year because I wanted to improve
the IPsec (and UDP) forwarding path. Batching packets
at layer2 and send them directly to the output path
seemed to be a good method to improve this.
In particular, we need to do only one IPsec lookup
for the whole packet chain. So it relaxes the pain
from reomoving the IPsec flowcache a bit. It can be
only a first step, but we need some improvements here
as people start to complain about that.
But did you also experiment with XDP on this?
I've already tried to figure out what I have to to
do to get XDP with forwarding, but still don't realy
know how to set this up.
Maybe it is time to have a deeper look into BPF/XDP,
but for now I feel a bit lost with this.
Would be curious about
the numbers. You'd get implicit batching for the forwarding via devmap
as well if you're required to flush it out via different device with
XDP_REDIRECT; otherwise XDP_TX of course. Given we have recently
integrated helpers for XDP to do a FIB and neighbor lookup from the
kernel tables, where it's thus shared and integrated with the rest of
the stack and tooling, it would be awesome to get to the same point
with xfrm as well. Eyal recently did a start on that for xfrm for tc
progs; would be nice to have integration on XDP as well, potentially
it might also result in a bigger plus on the forwarding numbers.
It might make sense to intrgrate XDP with xfrm to
be able to compare numbers etc. But I need a working
XDP setup and some understanding about it first, what
could take some time.
From: Daniel Borkmann <daniel@iogearbox.net> Date: 2018-06-19 22:22:09
On 06/17/2018 11:23 AM, Steffen Klassert wrote:
[...]
quoted
Would be curious about
the numbers. You'd get implicit batching for the forwarding via devmap
as well if you're required to flush it out via different device with
XDP_REDIRECT; otherwise XDP_TX of course. Given we have recently
integrated helpers for XDP to do a FIB and neighbor lookup from the
kernel tables, where it's thus shared and integrated with the rest of
the stack and tooling, it would be awesome to get to the same point
with xfrm as well. Eyal recently did a start on that for xfrm for tc
progs; would be nice to have integration on XDP as well, potentially
it might also result in a bigger plus on the forwarding numbers.
It might make sense to intrgrate XDP with xfrm to
be able to compare numbers etc. But I need a working
XDP setup and some understanding about it first, what
could take some time.
Okay, no prob. If you have any questions feel free to shoot an email.
Thanks,
Daniel
From: Andrew Collins <hidden> Date: 2018-06-20 00:56:46
On Thu, Jun 14, 2018 at 5:55 PM, David Miller [off-list ref] wrote:
And guess what? Then millions of possibilities would have been
openned up, rather than just this one special case.
So, I ask, please see the larger picture.
+cc netdev/etc
This is perhaps unrelated to the topic at hand, but as someone who's shipped
a bunch of devices over the years using the linux kernel forwarding path and
needs performance but wants to avoid moving to out of tree userspace offload
for all the reasons that you and many others have stated, is the long
term vision that the existing kernel forwarding path will transparently take
advantage of eBPF (ala bpfilter), or that users will write custom/individualized
eBPF forwarding paths for their usecases as necessary?
I (and I suspect many others) will start on the latter anyways, I'm just curious
whether it's desired/expected that such custom fastpath users will eventually
be rolled back into/replaced by a transparent upstream in-kernel equivalent.