From: Sun Shouxin <hidden> Date: 2021-12-10 13:09:10
Since ipv6 neighbor solicitation and advertisement messages
isn't handled gracefully in bonding6 driver, we can see packet
drop due to inconsistency bewteen mac address in the option
message and source MAC .
Another examples is ipv6 neighbor solicitation and advertisement
messages from VM via tap attached to host brighe, the src mac
mighe be changed through balance-alb mode, but it is not synced
with Link-layer address in the option message.
The patch implements bond6's tx handle for ipv6 neighbor
solicitation and advertisement messages.
Border-Leaf
/ \
/ \
Tunnel1 Tunnel2
/ \
/ \
Leaf-1--Tunnel3--Leaf-2
\ /
\ /
\ /
\ /
NIC1 NIC2
\ /
server
We can see in our lab the Border-Leaf receives occasionally
a NA packet which is assigned to NIC1 mac in ND/NS option
message, but actaully send out via NIC2 mac due to tx-alb,
as a result, it will cause inconsistency between MAC table
and ND Table in Border-Leaf, i.e, NIC1 = Tunnel2 in ND table
and NIC1 = Tunnel1 in mac table.
And then, Border-Leaf starts to forward packet destinated
to the Server, it will only check the ND table entry in some
switch to encapsulate the destination MAC of the message as
NIC1 MAC, and then send it out from Tunnel2 by ND table.
Then, Leaf-2 receives the packet, it notices the destination
MAC of message is NIC1 MAC and should forword it to Tunne1
by Tunnel3.
However, this traffic forward will be failure due to split
horizon of VxLAN tunnels.
Suggested-by: Hu Yadi <redacted>
Signed-off-by: Sun Shouxin <redacted>
---
drivers/net/bonding/bond_alb.c | 131 +++++++++++++++++++++++++++++++++++++++++
1 file changed, 131 insertions(+)
@@ -1269,6 +1270,119 @@ static int alb_set_mac_address(struct bonding *bond, void *addr)returnres;}+/*determine if the packet is NA or NS*/+staticboolalb_determine_nd(structicmp6hdr*hdr)+{+if(hdr->icmp6_type==NDISC_NEIGHBOUR_ADVERTISEMENT||+hdr->icmp6_type==NDISC_NEIGHBOUR_SOLICITATION){+returntrue;+}++returnfalse;+}++staticvoidalb_change_nd_option(structsk_buff*skb,void*data)+{+structnd_msg*msg=(structnd_msg*)skb_transport_header(skb);+structnd_opt_hdr*nd_opt=(structnd_opt_hdr*)msg->opt;+structnet_device*dev=skb->dev;+structicmp6hdr*icmp6h=icmp6_hdr(skb);+structipv6hdr*ip6hdr=ipv6_hdr(skb);+u8*lladdr=NULL;+u32ndoptlen=skb_tail_pointer(skb)-(skb_transport_header(skb)++offsetof(structnd_msg,opt));++while(ndoptlen){+intl;++switch(nd_opt->nd_opt_type){+caseND_OPT_SOURCE_LL_ADDR:+caseND_OPT_TARGET_LL_ADDR:+lladdr=ndisc_opt_addr_data(nd_opt,dev);+break;++default:+lladdr=NULL;+break;+}++l=nd_opt->nd_opt_len<<3;++if(ndoptlen<l||l==0)+return;++if(lladdr){+memcpy(lladdr,data,dev->addr_len);+icmp6h->icmp6_cksum=0;++icmp6h->icmp6_cksum=csum_ipv6_magic(&ip6hdr->saddr,+&ip6hdr->daddr,+ntohs(ip6hdr->payload_len),+IPPROTO_ICMPV6,+csum_partial(icmp6h,+ntohs(ip6hdr->payload_len),0));+}+ndoptlen-=l;+nd_opt=((void*)nd_opt)+l;+}+}++staticu8*alb_get_lladdr(structsk_buff*skb)+{+structnd_msg*msg=(structnd_msg*)skb_transport_header(skb);+structnd_opt_hdr*nd_opt=(structnd_opt_hdr*)msg->opt;+structnet_device*dev=skb->dev;+u8*lladdr=NULL;+u32ndoptlen=skb_tail_pointer(skb)-(skb_transport_header(skb)++offsetof(structnd_msg,opt));++while(ndoptlen){+intl;++switch(nd_opt->nd_opt_type){+caseND_OPT_SOURCE_LL_ADDR:+caseND_OPT_TARGET_LL_ADDR:+lladdr=ndisc_opt_addr_data(nd_opt,dev);+break;++default:+break;+}++l=nd_opt->nd_opt_len<<3;++if(ndoptlen<l||l==0)+returnlladdr;++if(lladdr)+returnlladdr;++ndoptlen-=l;+nd_opt=((void*)nd_opt)+l;+}++returnlladdr;+}++staticvoidalb_set_nd_option(structsk_buff*skb,structbonding*bond,+structslave*tx_slave)+{+structipv6hdr*ip6hdr;+structicmp6hdr*hdr=NULL;++if(skb->protocol==htons(ETH_P_IPV6)){+if(tx_slave&&tx_slave!=+rcu_access_pointer(bond->curr_active_slave)){+ip6hdr=ipv6_hdr(skb);+if(ip6hdr->nexthdr==IPPROTO_ICMPV6){+hdr=icmp6_hdr(skb);+if(alb_determine_nd(hdr))+alb_change_nd_option(skb,tx_slave->dev->dev_addr);+}+}+}+}+/************************ exported alb functions ************************/intbond_alb_initialize(structbonding*bond,intrlb_enabled)
@@ -1415,6 +1529,7 @@ struct slave *bond_xmit_alb_slave_get(struct bonding *bond,}caseETH_P_IPV6:{conststructipv6hdr*ip6hdr;+structicmp6hdr*hdr=NULL;/* IPv6 doesn't really use broadcast mac address, but leave*thatherejustincase.
drivers/net/bonding/bond_alb.c:1318:47: error: implicit declaration of function 'csum_ipv6_magic'; did you mean 'csum_tcpudp_magic'? [-Werror=implicit-function-declaration]
Hi,all
Any comments will be appreciated.
thanks a lot.
Yadi
主题: [PATCH V2] net: bonding: Add support for IPV6 ns/na
Since ipv6 neighbor solicitation and advertisement messages isn't handled
gracefully in bonding6 driver, we can see packet drop due to inconsistency
bewteen mac address in the option message and source MAC .
Another examples is ipv6 neighbor solicitation and advertisement messages
from VM via tap attached to host brighe, the src mac mighe be changed
through balance-alb mode, but it is not synced with Link-layer address in
the option message.
The patch implements bond6's tx handle for ipv6 neighbor solicitation and
advertisement messages.
Border-Leaf
/ \
/ \
Tunnel1 Tunnel2
/ \
/ \
Leaf-1--Tunnel3--Leaf-2
\ /
\ /
\ /
\ /
NIC1 NIC2
\ /
server
We can see in our lab the Border-Leaf receives occasionally a NA packet
which is assigned to NIC1 mac in ND/NS option message, but actaully send out
via NIC2 mac due to tx-alb, as a result, it will cause inconsistency between
MAC table and ND Table in Border-Leaf, i.e, NIC1 = Tunnel2 in ND table and
NIC1 = Tunnel1 in mac table.
And then, Border-Leaf starts to forward packet destinated to the Server, it
will only check the ND table entry in some switch to encapsulate the
destination MAC of the message as
NIC1 MAC, and then send it out from Tunnel2 by ND table.
Then, Leaf-2 receives the packet, it notices the destination MAC of message
is NIC1 MAC and should forword it to Tunne1 by Tunnel3.
However, this traffic forward will be failure due to split horizon of VxLAN
tunnels.
Suggested-by: Hu Yadi <redacted>
Signed-off-by: Sun Shouxin <redacted>
---
drivers/net/bonding/bond_alb.c | 131
+++++++++++++++++++++++++++++++++++++++++
1 file changed, 131 insertions(+)
From: Eric Dumazet <hidden> Date: 2021-12-14 08:05:26
On 12/10/21 5:08 AM, Sun Shouxin wrote:
quoted hunk
Since ipv6 neighbor solicitation and advertisement messages
isn't handled gracefully in bonding6 driver, we can see packet
drop due to inconsistency bewteen mac address in the option
message and source MAC .
Another examples is ipv6 neighbor solicitation and advertisement
messages from VM via tap attached to host brighe, the src mac
mighe be changed through balance-alb mode, but it is not synced
with Link-layer address in the option message.
The patch implements bond6's tx handle for ipv6 neighbor
solicitation and advertisement messages.
Border-Leaf
/ \
/ \
Tunnel1 Tunnel2
/ \
/ \
Leaf-1--Tunnel3--Leaf-2
\ /
\ /
\ /
\ /
NIC1 NIC2
\ /
server
We can see in our lab the Border-Leaf receives occasionally
a NA packet which is assigned to NIC1 mac in ND/NS option
message, but actaully send out via NIC2 mac due to tx-alb,
as a result, it will cause inconsistency between MAC table
and ND Table in Border-Leaf, i.e, NIC1 = Tunnel2 in ND table
and NIC1 = Tunnel1 in mac table.
And then, Border-Leaf starts to forward packet destinated
to the Server, it will only check the ND table entry in some
switch to encapsulate the destination MAC of the message as
NIC1 MAC, and then send it out from Tunnel2 by ND table.
Then, Leaf-2 receives the packet, it notices the destination
MAC of message is NIC1 MAC and should forword it to Tunne1
by Tunnel3.
However, this traffic forward will be failure due to split
horizon of VxLAN tunnels.
Suggested-by: Hu Yadi <redacted>
Signed-off-by: Sun Shouxin <redacted>
---
drivers/net/bonding/bond_alb.c | 131 +++++++++++++++++++++++++++++++++++++++++
1 file changed, 131 insertions(+)
@@ -1269,6 +1270,119 @@ static int alb_set_mac_address(struct bonding *bond, void *addr)returnres;}+/*determine if the packet is NA or NS*/+staticboolalb_determine_nd(structicmp6hdr*hdr)+{+if(hdr->icmp6_type==NDISC_NEIGHBOUR_ADVERTISEMENT||+hdr->icmp6_type==NDISC_NEIGHBOUR_SOLICITATION){+returntrue;+}++returnfalse;+}++staticvoidalb_change_nd_option(structsk_buff*skb,void*data)+{+structnd_msg*msg=(structnd_msg*)skb_transport_header(skb);+structnd_opt_hdr*nd_opt=(structnd_opt_hdr*)msg->opt;+structnet_device*dev=skb->dev;+structicmp6hdr*icmp6h=icmp6_hdr(skb);+structipv6hdr*ip6hdr=ipv6_hdr(skb);+u8*lladdr=NULL;+u32ndoptlen=skb_tail_pointer(skb)-(skb_transport_header(skb)++offsetof(structnd_msg,opt));++while(ndoptlen){+intl;++switch(nd_opt->nd_opt_type){+caseND_OPT_SOURCE_LL_ADDR:+caseND_OPT_TARGET_LL_ADDR:+lladdr=ndisc_opt_addr_data(nd_opt,dev);+break;++default:+lladdr=NULL;+break;+}++l=nd_opt->nd_opt_len<<3;++if(ndoptlen<l||l==0)+return;++if(lladdr){+memcpy(lladdr,data,dev->addr_len);
I am not sure it is allowed to change skb content without
making sure skb ->head is private.
(Think of tcpdump -i slaveX : we want to see the packet content before your change)
I would think skb_cow_head() or something similar is needed.
This is tricky of course, since all cached pointers (icmp6h, ip6hdr, msg, nd_opt)
would need to be fetched again, since skb->head/data might be changed
by skb_cow_head().
@@ -1415,6 +1529,7 @@ struct slave *bond_xmit_alb_slave_get(struct bonding *bond, } case ETH_P_IPV6: { const struct ipv6hdr *ip6hdr;+ struct icmp6hdr *hdr = NULL; /* IPv6 doesn't really use broadcast mac address, but leave * that here just in case.
Since ipv6 neighbor solicitation and advertisement messages
isn't handled gracefully in bonding6 driver, we can see packet
drop due to inconsistency bewteen mac address in the option
message and source MAC .
Another examples is ipv6 neighbor solicitation and advertisement
messages from VM via tap attached to host brighe, the src mac
mighe be changed through balance-alb mode, but it is not synced
with Link-layer address in the option message.
The patch implements bond6's tx handle for ipv6 neighbor
solicitation and advertisement messages.
Border-Leaf
/ \
/ \
Tunnel1 Tunnel2
/ \
/ \
Leaf-1--Tunnel3--Leaf-2
\ /
\ /
\ /
\ /
NIC1 NIC2
\ /
server
We can see in our lab the Border-Leaf receives occasionally
a NA packet which is assigned to NIC1 mac in ND/NS option
message, but actaully send out via NIC2 mac due to tx-alb,
as a result, it will cause inconsistency between MAC table
and ND Table in Border-Leaf, i.e, NIC1 = Tunnel2 in ND table
and NIC1 = Tunnel1 in mac table.
And then, Border-Leaf starts to forward packet destinated
to the Server, it will only check the ND table entry in some
switch to encapsulate the destination MAC of the message as
NIC1 MAC, and then send it out from Tunnel2 by ND table.
Then, Leaf-2 receives the packet, it notices the destination
MAC of message is NIC1 MAC and should forword it to Tunne1
by Tunnel3.
However, this traffic forward will be failure due to split
horizon of VxLAN tunnels.
Suggested-by: Hu Yadi <redacted>
Signed-off-by: Sun Shouxin <redacted>
---
drivers/net/bonding/bond_alb.c | 131 +++++++++++++++++++++++++++++++++++++++++
1 file changed, 131 insertions(+)
I am not sure it is allowed to change skb content without
making sure skb ->head is private.
(Think of tcpdump -i slaveX : we want to see the packet content before your change)
I would think skb_cow_head() or something similar is needed.
This is tricky of course, since all cached pointers (icmp6h, ip6hdr, msg, nd_opt)
would need to be fetched again, since skb->head/data might be changed
by skb_cow_head().
The tcpdump should show the last packet which sent off from NIC in the end.
could you light me up specific conditions?
}
case ETH_P_IPV6: {
const struct ipv6hdr *ip6hdr;
+ struct icmp6hdr *hdr = NULL;
/* IPv6 doesn't really use broadcast mac address, but leave
* that here just in case.
From: Eric Dumazet <hidden> Date: 2021-12-20 15:25:23
On 12/19/21 5:57 PM, 孙守鑫 wrote:
在 2021/12/14 16:05, Eric Dumazet 写道:
quoted
On 12/10/21 5:08 AM, Sun Shouxin wrote:
quoted
Since ipv6 neighbor solicitation and advertisement messages
isn't handled gracefully in bonding6 driver, we can see packet
drop due to inconsistency bewteen mac address in the option
message and source MAC .
Another examples is ipv6 neighbor solicitation and advertisement
messages from VM via tap attached to host brighe, the src mac
mighe be changed through balance-alb mode, but it is not synced
with Link-layer address in the option message.
The patch implements bond6's tx handle for ipv6 neighbor
solicitation and advertisement messages.
Border-Leaf
/ \
/ \
Tunnel1 Tunnel2
/ \
/ \
Leaf-1--Tunnel3--Leaf-2
\ /
\ /
\ /
\ /
NIC1 NIC2
\ /
server
We can see in our lab the Border-Leaf receives occasionally
a NA packet which is assigned to NIC1 mac in ND/NS option
message, but actaully send out via NIC2 mac due to tx-alb,
as a result, it will cause inconsistency between MAC table
and ND Table in Border-Leaf, i.e, NIC1 = Tunnel2 in ND table
and NIC1 = Tunnel1 in mac table.
And then, Border-Leaf starts to forward packet destinated
to the Server, it will only check the ND table entry in some
switch to encapsulate the destination MAC of the message as
NIC1 MAC, and then send it out from Tunnel2 by ND table.
Then, Leaf-2 receives the packet, it notices the destination
MAC of message is NIC1 MAC and should forword it to Tunne1
by Tunnel3.
However, this traffic forward will be failure due to split
horizon of VxLAN tunnels.
Suggested-by: Hu Yadi <redacted>
Signed-off-by: Sun Shouxin <redacted>
---
drivers/net/bonding/bond_alb.c | 131 +++++++++++++++++++++++++++++++++++++++++
1 file changed, 131 insertions(+)
I am not sure it is allowed to change skb content without
making sure skb ->head is private.
(Think of tcpdump -i slaveX : we want to see the packet content before your change)
I would think skb_cow_head() or something similar is needed.
This is tricky of course, since all cached pointers (icmp6h, ip6hdr, msg, nd_opt)
would need to be fetched again, since skb->head/data might be changed
by skb_cow_head().
The tcpdump should show the last packet which sent off from NIC in the end.
could you light me up specific conditions?
I think I have been clear.
You can not modify skb->head unless it is allowed to.
tcpdump on the slave must show the exact packet being received,
before your modifications in bonding driver.
}
case ETH_P_IPV6: {
const struct ipv6hdr *ip6hdr;
+ struct icmp6hdr *hdr = NULL;
/* IPv6 doesn't really use broadcast mac address, but leave
* that here just in case.