From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2021-12-09 00:08:55
Hi,
The following patchset contains Netfilter fixes for net:
1) Fix bogus compilter warning in nfnetlink_queue, from Florian Westphal.
2) Don't run conntrack on vrf with !dflt qdisc, from Nicolas Dichtel.
3) Fix nft_pipapo bucket load in AVX2 lookup routine for six 8-bit
groups, from Stefano Brivio.
4) Break rule evaluation on malformed TCP options.
5) Use socat instead of nc in selftests/netfilter/nft_zones_many.sh,
also from Florian
6) Fix KCSAN data-race in conntrack timeout updates, from Eric Dumazet.
Please, pull these changes from:
git://git.kernel.org/pub/scm/linux/kernel/git/pablo/nf.git
Thanks.
----------------------------------------------------------------
The following changes since commit 34d8778a943761121f391b7921f79a7adbe1feaf:
MAINTAINERS: s390/net: add Alexandra and Wenjia as maintainer (2021-11-30 12:20:07 +0000)
are available in the Git repository at:
git://git.kernel.org/pub/scm/linux/kernel/git/pablo/nf.git HEAD
for you to fetch changes up to 802a7dc5cf1bef06f7b290ce76d478138408d6b1:
netfilter: conntrack: annotate data-races around ct->timeout (2021-12-08 01:29:15 +0100)
----------------------------------------------------------------
Eric Dumazet (1):
netfilter: conntrack: annotate data-races around ct->timeout
Florian Westphal (2):
netfilter: nfnetlink_queue: silence bogus compiler warning
selftests: netfilter: switch zone stress to socat
Nicolas Dichtel (1):
vrf: don't run conntrack on vrf with !dflt qdisc
Pablo Neira Ayuso (1):
netfilter: nft_exthdr: break evaluation if setting TCP option fails
Stefano Brivio (2):
nft_set_pipapo: Fix bucket load in AVX2 lookup routine for six 8-bit groups
selftests: netfilter: Add correctness test for mac,net set type
drivers/net/vrf.c | 8 +++---
include/net/netfilter/nf_conntrack.h | 6 ++---
net/netfilter/nf_conntrack_core.c | 6 ++---
net/netfilter/nf_conntrack_netlink.c | 2 +-
net/netfilter/nf_flow_table_core.c | 4 +--
net/netfilter/nfnetlink_queue.c | 2 +-
net/netfilter/nft_exthdr.c | 11 +++++---
net/netfilter/nft_set_pipapo_avx2.c | 2 +-
tools/testing/selftests/netfilter/conntrack_vrf.sh | 30 +++++++++++++++++++---
.../selftests/netfilter/nft_concat_range.sh | 24 ++++++++++++++---
.../testing/selftests/netfilter/nft_zones_many.sh | 19 +++++++++-----
11 files changed, 82 insertions(+), 32 deletions(-)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2021-12-09 00:08:59
From: Florian Westphal <fw@strlen.de>
net/netfilter/nfnetlink_queue.c:601:36: warning: variable 'ctinfo' is
uninitialized when used here [-Wuninitialized]
if (ct && nfnl_ct->build(skb, ct, ctinfo, NFQA_CT, NFQA_CT_INFO) < 0)
ctinfo is only uninitialized if ct == NULL. Init it to 0 to silence this.
Reported-by: kernel test robot <redacted>
Signed-off-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/nfnetlink_queue.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2021-12-09 00:09:00
From: Stefano Brivio <redacted>
The existing net,mac test didn't cover the issue recently reported
by Nikita Yushchenko, where MAC addresses wouldn't match if given
as first field of a concatenated set with AVX2 and 8-bit groups,
because there's a different code path covering the lookup of six
8-bit groups (MAC addresses) if that's the first field.
Add a similar mac,net test, with MAC address and IPv4 address
swapped in the set specification.
Signed-off-by: Stefano Brivio <redacted>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
.../selftests/netfilter/nft_concat_range.sh | 24 ++++++++++++++++---
1 file changed, 21 insertions(+), 3 deletions(-)
@@ -23,8 +23,8 @@ TESTS="reported_issues correctness concurrency timeout"# Set types, defined by TYPE_ variables belowTYPES="net_port port_net net6_port port_proto net6_port_mac net6_port_mac_proto-net_port_netnet_macnet_mac_icmpnet6_mac_icmpnet6_port_net6_port-net_port_mac_proto_net"+net_port_netnet_macmac_netnet_mac_icmpnet6_mac_icmp+net6_port_net6_portnet_port_mac_proto_net"# Reported bugs, also described by TYPE_ variables belowBUGS="flush_remove_add"
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2021-12-09 00:09:01
From: Nicolas Dichtel <redacted>
After the below patch, the conntrack attached to skb is set to "notrack" in
the context of vrf device, for locally generated packets.
But this is true only when the default qdisc is set to the vrf device. When
changing the qdisc, notrack is not set anymore.
In fact, there is a shortcut in the vrf driver, when the default qdisc is
set, see commit dcdd43c41e60 ("net: vrf: performance improvements for
IPv4") for more details.
This patch ensures that the behavior is always the same, whatever the qdisc
is.
To demonstrate the difference, a new test is added in conntrack_vrf.sh.
Fixes: 8c9c296adfae ("vrf: run conntrack only in context of lower/physdev for locally generated packets")
Signed-off-by: Nicolas Dichtel <redacted>
Acked-by: Florian Westphal <fw@strlen.de>
Reviewed-by: David Ahern <dsahern@kernel.org>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
drivers/net/vrf.c | 8 ++---
.../selftests/netfilter/conntrack_vrf.sh | 30 ++++++++++++++++---
2 files changed, 30 insertions(+), 8 deletions(-)
@@ -150,11 +150,27 @@ EOF# oifname is the vrf device. test_masquerade_vrf(){+localqdisc=$1++if["$qdisc"!="default"];then+tc-net$ns0qdiscadddevtvrfroot$qdisc+fi+ipnetnsexec$ns0conntrack-F2>/dev/null ipnetnsexec$ns0nft-f-<<EOF flushruleset tableipnat{+chainrawout{+typefilterhookoutputpriorityraw;++oiftvrfctstateuntrackedcounter+}+chainpostrouting2{+typefilterhookpostroutingprioritymangle;++oiftvrfctstateuntrackedcounter+}chainpostrouting{typenathookpostroutingpriority0;# NB: masquerade should always be combined with 'oif(name) bla',
@@ -171,13 +187,18 @@ EOFfi# must also check that nat table was evaluated on second (lower device) iteration.-ipnetnsexec$ns0nftlisttableipnat|grep-q'counter packets 2'+ipnetnsexec$ns0nftlisttableipnat|grep-q'counter packets 2'&&+ipnetnsexec$ns0nftlisttableipnat|grep-q'untracked counter packets [1-9]'if[$?-eq0];then-echo"PASS: iperf3 connect with masquerade + sport rewrite on vrf device"+echo"PASS: iperf3 connect with masquerade + sport rewrite on vrf device ($qdisc qdisc)"else-echo"FAIL: vrf masq rule has unexpected counter value"+echo"FAIL: vrf rules have unexpected counter value"ret=1fi++if["$qdisc"!="default"];then+tc-net$ns0qdiscdeldevtvrfroot+fi}# add masq rule that gets evaluated w. outif set to veth device.
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2021-12-09 00:09:03
From: Stefano Brivio <redacted>
The sixth byte of packet data has to be looked up in the sixth group,
not in the seventh one, even if we load the bucket data into ymm6
(and not ymm5, for convenience of tracking stalls).
Without this fix, matching on a MAC address as first field of a set,
if 8-bit groups are selected (due to a small set size) would fail,
that is, the given MAC address would never match.
Reported-by: Nikita Yushchenko <redacted>
Cc: <redacted> # 5.6.x
Fixes: 7400b063969b ("nft_set_pipapo: Introduce AVX2-based lookup implementation")
Signed-off-by: Stefano Brivio <redacted>
Tested-By: Nikita Yushchenko <redacted>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/nft_set_pipapo_avx2.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
@@ -886,7 +886,7 @@ static int nft_pipapo_avx2_lookup_8b_6(unsigned long *map, unsigned long *fill,NFT_PIPAPO_AVX2_BUCKET_LOAD8(4,lt,4,pkt[4],bsize);NFT_PIPAPO_AVX2_AND(5,0,1);-NFT_PIPAPO_AVX2_BUCKET_LOAD8(6,lt,6,pkt[5],bsize);+NFT_PIPAPO_AVX2_BUCKET_LOAD8(6,lt,5,pkt[5],bsize);NFT_PIPAPO_AVX2_AND(7,2,3);/* Stall */
@@ -1036,7 +1036,7 @@ static int nf_ct_resolve_clash_harder(struct sk_buff *skb, u32 repl_idx)}/* We want the clashing entry to go away real soon: 1 second timeout. */-loser_ct->timeout=nfct_time_stamp+HZ;+WRITE_ONCE(loser_ct->timeout,nfct_time_stamp+HZ);/* IPS_NAT_CLASH removes the entry automatically on the first*reply.AlsopreventsUDPtrackerfrommovingtheentryto
@@ -1560,7 +1560,7 @@ __nf_conntrack_alloc(struct net *net,/* save hash for reusing when confirming */*(unsignedlong*)(&ct->tuplehash[IP_CT_DIR_REPLY].hnnode.pprev)=hash;ct->status=0;-ct->timeout=0;+WRITE_ONCE(ct->timeout,0);write_pnet(&ct->ct_net,net);memset(&ct->__nfct_init_offset,0,offsetof(structnf_conn,proto)-
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2021-12-09 00:09:06
From: Florian Westphal <fw@strlen.de>
centos9 has nmap-ncat which doesn't like the '-q' option, use socat.
While at it, mark test skipped if needed tools are missing.
Signed-off-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
.../selftests/netfilter/nft_zones_many.sh | 19 +++++++++++++------
1 file changed, 13 insertions(+), 6 deletions(-)
@@ -18,11 +18,17 @@ cleanup()ipnetnsdel$ns}-ipnetnsadd$ns-if[$?-ne0];then-echo"SKIP: Could not create net namespace $gw"-exit$ksft_skip-fi+checktool(){+if!$1>/dev/null2>&1;then+echo"SKIP: Could not $2"+exit$ksft_skip+fi+}++checktool"nft --version""run test without nft tool"+checktool"ip -Version""run test without ip tool"+checktool"socat -V""run test without socat tool"+checktool"ip netns add $ns""create net namespace"trapcleanupEXIT
@@ -71,7 +77,8 @@ EOFlocalstart=$(date+%s%3N)i=$((i+10000))j=$((j+1))-ddif=/dev/zeroof=/dev/stdoutbs=8kcount=100002>/dev/null|ipnetnsexec"$ns"nc-w1-q1-u-p12345127.0.0.112345>/dev/null+# nft rule in output places each packet in a different zone.+ddif=/dev/zeroof=/dev/stdoutbs=8kcount=100002>/dev/null|ipnetnsexec"$ns"socatSTDINUDP:127.0.0.1:12345,sourceport=12345if[$?-ne0];thenret=1break
Hello:
This series was applied to netdev/net.git (master)
by Pablo Neira Ayuso [off-list ref]:
On Thu, 9 Dec 2021 01:08:41 +0100 you wrote:
From: Florian Westphal <fw@strlen.de>
net/netfilter/nfnetlink_queue.c:601:36: warning: variable 'ctinfo' is
uninitialized when used here [-Wuninitialized]
if (ct && nfnl_ct->build(skb, ct, ctinfo, NFQA_CT, NFQA_CT_INFO) < 0)
ctinfo is only uninitialized if ct == NULL. Init it to 0 to silence this.
[...]