From: Julius Bairaktaris <hidden> Date: 2026-09-10 09:00:57
A host that forwards one direction of a TCP connection only sees the
client's SYN and then an ACK continuing from it; the answer took another
path. That ACK has no entry in the transition table, so it and every
packet after it are invalid, the conntrack entry stays in SYN_SENT
[UNREPLIED] with a single packet, and the connection reaches neither the
stateful part of a ruleset nor a flowtable.
Patch 1 takes such a connection over as the mid-stream pickup it is.
Patch 2 withholds the reply direction of a flow offloaded in one
direction until conntrack has seen a reply, and offloads it once
conntrack has. Patch 3 offers a connection whose reply was never seen to
the flowtable in the original direction. Patch 4 adds a selftest arm for
the path; without patches 1 to 3 it fails, with the router forwarding the
SYN of the connection and nothing else.
Verified on an IPQ8074 router with hardware flow offload, iperf3 -P2
between two hosts on different subnets of one bridge with the reply
direction bypassing the router:
CPU port switch port throughput
asymmetric, without the series 81k pps 81k pps 944-947 Mbit/s
asymmetric, with the series 2-11 pps 81k pps 948-949 Mbit/s
symmetric, with the series 11-38 pps 81k pps 926-948 Mbit/s
I wrote this series with the help of an AI coding assistant, as the
Assisted-by tags record. I have reviewed and tested it myself.
Gary Dotzler (2):
netfilter: conntrack: pick up a TCP flow whose SYN was never answered
netfilter: nft_flow_offload: offload a TCP flow that has no reply
Julius Bairaktaris (2):
netfilter: flowtable: promote a flow offloaded in one direction only
selftests: netfilter: cover a TCP flow whose reply is never seen
include/net/netfilter/nf_conntrack_l4proto.h | 7 ++
net/netfilter/nf_conntrack_proto_tcp.c | 19 +++++
net/netfilter/nf_flow_table_ip.c | 26 +++++++
net/netfilter/nft_flow_offload.c | 6 +-
.../selftests/net/netfilter/nft_flowtable.sh | 73 +++++++++++++++++++
5 files changed, 129 insertions(+), 2 deletions(-)
--
2.53.0
From: Julius Bairaktaris <hidden> Date: 2026-09-10 09:00:56
From: Gary Dotzler <redacted>
When only one direction of a connection passes the host, the SYN is
seen, the answer to it is not, and the next packet is an ACK continuing
from where the SYN left off. The transition table has no entry for that,
so that ACK and every packet after it are invalid and the entry sits in
SYN_SENT [UNREPLIED] with one packet.
Treat the connection as a mid-stream pickup. Delete the entry and look
the packet up again, so it creates a new entry through the loose path.
That path fills in the unseen direction from the packet and stops
window checking in both directions.
Signed-off-by: Gary Dotzler <redacted>
Tested-by: Julius Bairaktaris <redacted>
Signed-off-by: Julius Bairaktaris <redacted>
Assisted-by: Claude:claude-opus-5
---
net/netfilter/nf_conntrack_proto_tcp.c | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
@@ -1173,6 +1173,25 @@ int nf_conntrack_tcp_packet(struct nf_conn *ct,returnNF_ACCEPT;}+/* The answer to the SYN never passed the host, as happens+*whenthereplydirectiontakesanotherpath,andtheclient+*continuesfromwhereitsSYNleftoff.Taketheconnection+*overasamid-streampickup:deletetheentrysothepacket+*createsanewonethatseedstheunseendirectionfromthe+*packetitself.+*/+if(tn->tcp_loose&&!nfct_synproxy(ct)&&+old_state==TCP_CONNTRACK_SYN_SENT&&+index==TCP_ACK_SET&&dir==IP_CT_DIR_ORIGINAL&&+!test_bit(IPS_SEEN_REPLY_BIT,&ct->status)&&+ntohl(th->seq)==ct->proto.tcp.seen[dir].td_end){+spin_unlock_bh(&ct->lock);++if(nf_ct_kill(ct))+return-NF_REPEAT;+returnNF_DROP;+}+/* Invalid packet */spin_unlock_bh(&ct->lock);nf_ct_l4proto_log_invalid(skb,ct,state,
From: Julius Bairaktaris <hidden> Date: 2026-09-10 09:00:57
A flow offloaded in the original direction alone carries a reply tuple
that conntrack has never seen a packet for. Leave that direction on the
classic path, so conntrack tracks it, and offload it as well once the
connection is assured.
act_ct promotes its unidirectional UDP flows on the same condition.
Signed-off-by: Julius Bairaktaris <redacted>
Assisted-by: Claude:claude-opus-5
---
net/netfilter/nf_flow_table_ip.c | 26 ++++++++++++++++++++++++++
1 file changed, 26 insertions(+)
@@ -465,6 +465,26 @@ nf_flow_offload_lookup(struct nf_flowtable_ctx *ctx,returnflow_offload_lookup(flow_table,&tuple);}+/* The reply direction of a flow offloaded in one direction only stays on the+*classicpathsothatconntrackseesit.Onceconntrackhas,theflowis+*offloadedinbothdirections.+*/+staticboolnf_flow_reply_unoffloaded(structnf_flowtable*flow_table,+structflow_offload*flow,+enumflow_offload_tuple_dirdir)+{+if(dir!=FLOW_OFFLOAD_DIR_REPLY||+test_bit(NF_FLOW_HW_BIDIRECTIONAL,&flow->flags))+returnfalse;++if(test_bit(IPS_ASSURED_BIT,&flow->ct->status)){+set_bit(NF_FLOW_HW_BIDIRECTIONAL,&flow->flags);+flow_offload_refresh(flow_table,flow,true);+}++returntrue;+}+staticintnf_flow_offload_forward(structnf_flowtable_ctx*ctx,structnf_flowtable*flow_table,structflow_offload_tuple_rhash*tuplehash,
@@ -478,6 +498,9 @@ static int nf_flow_offload_forward(struct nf_flowtable_ctx *ctx,dir=tuplehash->tuple.dir;flow=container_of(tuplehash,structflow_offload,tuplehash[dir]);+if(nf_flow_reply_unoffloaded(flow_table,flow,dir))+return0;+mtu=flow->tuplehash[dir].tuple.mtu+ctx->offset;if(flow->tuplehash[!dir].tuple.tun_num)mtu-=sizeof(*iph);
@@ -1074,6 +1097,9 @@ static int nf_flow_offload_ipv6_forward(struct nf_flowtable_ctx *ctx,dir=tuplehash->tuple.dir;flow=container_of(tuplehash,structflow_offload,tuplehash[dir]);+if(nf_flow_reply_unoffloaded(flow_table,flow,dir))+return0;+mtu=flow->tuplehash[dir].tuple.mtu+ctx->offset;if(flow->tuplehash[!dir].tuple.tun_num)mtu-=sizeof(*ip6h);
From: Julius Bairaktaris <hidden> Date: 2026-09-10 09:00:58
From: Gary Dotzler <redacted>
A connection whose reply never reaches the host cannot become assured,
and only assured TCP connections are offered to the flowtable, so every
packet of it stays on the classic path.
Offer such a connection to the flowtable in the original direction. The
reply direction follows once conntrack has seen a reply, so a flow that
starts one-directional picks it up if the path turns symmetric.
Signed-off-by: Gary Dotzler <redacted>
Co-developed-by: Julius Bairaktaris <redacted>
Signed-off-by: Julius Bairaktaris <redacted>
Assisted-by: Claude:claude-opus-5
---
include/net/netfilter/nf_conntrack_l4proto.h | 7 +++++++
net/netfilter/nft_flow_offload.c | 6 ++++--
2 files changed, 11 insertions(+), 2 deletions(-)
@@ -208,6 +208,13 @@ static inline bool nf_conntrack_tcp_established(const struct nf_conn *ct)returnct->proto.tcp.state==TCP_CONNTRACK_ESTABLISHED&&test_bit(IPS_ASSURED_BIT,&ct->status);}++/* A flow picked up without its reply cannot become assured. */+staticinlineboolnf_conntrack_tcp_unreplied(conststructnf_conn*ct)+{+returnct->proto.tcp.state==TCP_CONNTRACK_ESTABLISHED&&+!test_bit(IPS_SEEN_REPLY_BIT,&ct->status);+}#endif#ifdef CONFIG_NF_CT_PROTO_SCTP
From: Julius Bairaktaris <hidden> Date: 2026-09-10 09:01:04
Add an arm to nft_flowtable.sh in which ns2 answers over a direct link,
so that nsr1 sees the original direction only. The forward hook counter
then has to stay far below the size of the transferred file, which only
happens if the flowtable takes the connection over.
Signed-off-by: Julius Bairaktaris <redacted>
Assisted-by: Claude:claude-opus-5
---
.../selftests/net/netfilter/nft_flowtable.sh | 73 +++++++++++++++++++
1 file changed, 73 insertions(+)
@@ -516,6 +516,79 @@ elseret=1fi+# Asymmetric path test:+# ns2 answers over a direct link, so nsr1 sees the original direction only.+# Such a connection never becomes assured, but the flowtable is expected to+# take over the direction that nsr1 does see.+check_orig_offloaded()+{+localwhat=$1++localorig+orig=$(ipnetnsexec"$nsr1"nftresetcounterinetfilterrouted_orig|greppackets)+localorig_cnt=${orig#*bytes}++localfs+fs=$(du-sb"$nsin")+localmax_orig=$((${fs%%/*}/2))++# the flowtable takes over after the first few packets, so the forward+# hook must see a small fraction of the transferred file.+if["$orig_cnt"-gt"$max_orig"];then+echo"FAIL: $what: original counter $orig_cnt exceeds expected value $max_orig"1>&2+ret=1+return1+fi++echo"PASS: $what"+}++test_asymmetric_path()+{+iplinkaddnameeth1netns"$ns1"typevethpeernameeth1netns"$ns2"+ip-net"$ns1"addradd10.0.9.99/24deveth1+ip-net"$ns2"addradd10.0.9.98/24deveth1+ip-net"$ns1"addradddead:9::99/64deveth1nodad+ip-net"$ns2"addradddead:9::98/64deveth1nodad+ip-net"$ns1"linkseteth1up+ip-net"$ns2"linkseteth1up++# ns1 keeps sending through nsr1, ns2 answers on the direct link.+ip-net"$ns2"routeadd10.0.1.99via10.0.9.99deveth1+ip-6-net"$ns2"routeadddead:1::99viadead:9::99deveth1++ipnetnsexec"$ns1"sysctl-qnet.ipv4.ip_no_pmtu_disc=0+ipnetnsexec"$ns2"sysctl-qnet.ipv4.ip_no_pmtu_disc=0++ipnetnsexec"$nsr1"nftresetcounterstableinetfilter>/dev/null++iftest_tcp_forwarding"$ns1""$ns2"1410.0.2.9912345;then+check_orig_offloaded"flow offloaded for ns1/ns2 without reply"+else+echo"FAIL: flow offload for ns1/ns2 without reply"1>&2+ipnetnsexec"$nsr1"nftlistruleset1>&2+ret=1+fi++ipnetnsexec"$nsr1"nftresetcounterstableinetfilter>/dev/null++iftest_tcp_forwarding"$ns1""$ns2"16"[dead:2::99]"12345;then+check_orig_offloaded"IPv6 flow offloaded for ns1/ns2 without reply"+else+echo"FAIL: IPv6 flow offload for ns1/ns2 without reply"1>&2+ipnetnsexec"$nsr1"nftlistruleset1>&2+ret=1+fi++ipnetnsexec"$ns1"sysctl-qnet.ipv4.ip_no_pmtu_disc=1+ipnetnsexec"$ns2"sysctl-qnet.ipv4.ip_no_pmtu_disc=1++ip-net"$ns1"linkdeleth1+ipnetnsexec"$nsr1"nftresetcounterstableinetfilter>/dev/null+}++test_asymmetric_path+# delete default route, i.e. ns2 won't be able to reach ns1 and# will depend on ns1 being masqueraded in nsr1.# expect ns1 has nsr1 address.
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2026-09-11 11:42:01
On Thu, Sep 10, 2026 at 11:00:48AM +0200, Julius Bairaktaris wrote:
A host that forwards one direction of a TCP connection only sees the
client's SYN and then an ACK continuing from it; the answer took another
path. That ACK has no entry in the transition table, so it and every
packet after it are invalid, the conntrack entry stays in SYN_SENT
[UNREPLIED] with a single packet, and the connection reaches neither the
stateful part of a ruleset nor a flowtable.
conntrack needs to see packets in both directions, are you assuming a
packet-based load balancer in front of it?
Patch 1 takes such a connection over as the mid-stream pickup it is.
Patch 2 withholds the reply direction of a flow offloaded in one
direction until conntrack has seen a reply, and offloads it once
conntrack has. Patch 3 offers a connection whose reply was never seen to
the flowtable in the original direction. Patch 4 adds a selftest arm for
the path; without patches 1 to 3 it fails, with the router forwarding the
SYN of the connection and nothing else.
Verified on an IPQ8074 router with hardware flow offload, iperf3 -P2
between two hosts on different subnets of one bridge with the reply
direction bypassing the router:
CPU port switch port throughput
asymmetric, without the series 81k pps 81k pps 944-947 Mbit/s
asymmetric, with the series 2-11 pps 81k pps 948-949 Mbit/s
symmetric, with the series 11-38 pps 81k pps 926-948 Mbit/s
I wrote this series with the help of an AI coding assistant, as the
Assisted-by tags record. I have reviewed and tested it myself.
Gary Dotzler (2):
netfilter: conntrack: pick up a TCP flow whose SYN was never answered
netfilter: nft_flow_offload: offload a TCP flow that has no reply
Julius Bairaktaris (2):
netfilter: flowtable: promote a flow offloaded in one direction only
selftests: netfilter: cover a TCP flow whose reply is never seen
include/net/netfilter/nf_conntrack_l4proto.h | 7 ++
net/netfilter/nf_conntrack_proto_tcp.c | 19 +++++
net/netfilter/nf_flow_table_ip.c | 26 +++++++
net/netfilter/nft_flow_offload.c | 6 +-
.../selftests/net/netfilter/nft_flowtable.sh | 73 +++++++++++++++++++
5 files changed, 129 insertions(+), 2 deletions(-)
--
2.53.0
From: Julius Bairaktaris <hidden> Date: 2026-09-11 12:10:40
Hi Pablo,
thanks for taking your time to review.
conntrack needs to see packets in both directions, are you assuming a
packet-based load balancer in front of it?
No, an asymmetric route is in front of it, with two subnets sharing one
L2 segment and the server answering over that link, so the router only
ever sees one direction.
Conntrack already picks such a connection up when it misses the
handshake entirely, and patch 1 routes the SYN-seen case into that same
path, under the same nf_conntrack_tcp_loose gate.
net/netfilter/nf_conntrack_proto_tcp.c, tcp_new():
} else if (tn->tcp_loose == 0) {
/* Don't try to pick up connections. */
return false;
} else {
...
ct->proto.tcp.seen[0].flags =
ct->proto.tcp.seen[1].flags = IP_CT_TCP_FLAG_SACK_PERM |
IP_CT_TCP_FLAG_BE_LIBERAL;
Julius
Am Fr., 11. Sept. 2026 um 11:42 Uhr schrieb Pablo Neira Ayuso
[off-list ref]:
On Thu, Sep 10, 2026 at 11:00:48AM +0200, Julius Bairaktaris wrote:
quoted
A host that forwards one direction of a TCP connection only sees the
client's SYN and then an ACK continuing from it; the answer took another
path. That ACK has no entry in the transition table, so it and every
packet after it are invalid, the conntrack entry stays in SYN_SENT
[UNREPLIED] with a single packet, and the connection reaches neither the
stateful part of a ruleset nor a flowtable.
conntrack needs to see packets in both directions, are you assuming a
packet-based load balancer in front of it?
quoted
Patch 1 takes such a connection over as the mid-stream pickup it is.
Patch 2 withholds the reply direction of a flow offloaded in one
direction until conntrack has seen a reply, and offloads it once
conntrack has. Patch 3 offers a connection whose reply was never seen to
the flowtable in the original direction. Patch 4 adds a selftest arm for
the path; without patches 1 to 3 it fails, with the router forwarding the
SYN of the connection and nothing else.
Verified on an IPQ8074 router with hardware flow offload, iperf3 -P2
between two hosts on different subnets of one bridge with the reply
direction bypassing the router:
CPU port switch port throughput
asymmetric, without the series 81k pps 81k pps 944-947 Mbit/s
asymmetric, with the series 2-11 pps 81k pps 948-949 Mbit/s
symmetric, with the series 11-38 pps 81k pps 926-948 Mbit/s
I wrote this series with the help of an AI coding assistant, as the
Assisted-by tags record. I have reviewed and tested it myself.
Gary Dotzler (2):
netfilter: conntrack: pick up a TCP flow whose SYN was never answered
netfilter: nft_flow_offload: offload a TCP flow that has no reply
Julius Bairaktaris (2):
netfilter: flowtable: promote a flow offloaded in one direction only
selftests: netfilter: cover a TCP flow whose reply is never seen
include/net/netfilter/nf_conntrack_l4proto.h | 7 ++
net/netfilter/nf_conntrack_proto_tcp.c | 19 +++++
net/netfilter/nf_flow_table_ip.c | 26 +++++++
net/netfilter/nft_flow_offload.c | 6 +-
.../selftests/net/netfilter/nft_flowtable.sh | 73 +++++++++++++++++++
5 files changed, 129 insertions(+), 2 deletions(-)
--
2.53.0
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2026-09-11 14:09:53
Hi Julius,
On Fri, Sep 11, 2026 at 12:10:04PM +0000, Julius Bairaktaris wrote:
Hi Pablo,
thanks for taking your time to review.
quoted
conntrack needs to see packets in both directions, are you assuming a
packet-based load balancer in front of it?
No, an asymmetric route is in front of it, with two subnets sharing one
L2 segment and the server answering over that link, so the router only
ever sees one direction.
I see this requirement to support asymmetric path keeps coming, but
how hard is really to maintain this TCP state machine to deal with all
possible scenarios? ie. invalid transitions, retransmissions, etc.
this all without having access to full TCP connection. Is it that you
need NAT and the stateless NAT in nftables does not fulfill your
requirements?
Conntrack already picks such a connection up when it misses the
handshake entirely, and patch 1 routes the SYN-seen case into that same
path, under the same nf_conntrack_tcp_loose gate.
net/netfilter/nf_conntrack_proto_tcp.c, tcp_new():
} else if (tn->tcp_loose == 0) {
/* Don't try to pick up connections. */
return false;
} else {
...
ct->proto.tcp.seen[0].flags =
ct->proto.tcp.seen[1].flags = IP_CT_TCP_FLAG_SACK_PERM |
IP_CT_TCP_FLAG_BE_LIBERAL;
Julius
Am Fr., 11. Sept. 2026 um 11:42 Uhr schrieb Pablo Neira Ayuso
[off-list ref]:
quoted
On Thu, Sep 10, 2026 at 11:00:48AM +0200, Julius Bairaktaris wrote:
quoted
A host that forwards one direction of a TCP connection only sees the
client's SYN and then an ACK continuing from it; the answer took another
path. That ACK has no entry in the transition table, so it and every
packet after it are invalid, the conntrack entry stays in SYN_SENT
[UNREPLIED] with a single packet, and the connection reaches neither the
stateful part of a ruleset nor a flowtable.
conntrack needs to see packets in both directions, are you assuming a
packet-based load balancer in front of it?
quoted
Patch 1 takes such a connection over as the mid-stream pickup it is.
Patch 2 withholds the reply direction of a flow offloaded in one
direction until conntrack has seen a reply, and offloads it once
conntrack has. Patch 3 offers a connection whose reply was never seen to
the flowtable in the original direction. Patch 4 adds a selftest arm for
the path; without patches 1 to 3 it fails, with the router forwarding the
SYN of the connection and nothing else.
Verified on an IPQ8074 router with hardware flow offload, iperf3 -P2
between two hosts on different subnets of one bridge with the reply
direction bypassing the router:
CPU port switch port throughput
asymmetric, without the series 81k pps 81k pps 944-947 Mbit/s
asymmetric, with the series 2-11 pps 81k pps 948-949 Mbit/s
symmetric, with the series 11-38 pps 81k pps 926-948 Mbit/s
I wrote this series with the help of an AI coding assistant, as the
Assisted-by tags record. I have reviewed and tested it myself.
Gary Dotzler (2):
netfilter: conntrack: pick up a TCP flow whose SYN was never answered
netfilter: nft_flow_offload: offload a TCP flow that has no reply
Julius Bairaktaris (2):
netfilter: flowtable: promote a flow offloaded in one direction only
selftests: netfilter: cover a TCP flow whose reply is never seen
include/net/netfilter/nf_conntrack_l4proto.h | 7 ++
net/netfilter/nf_conntrack_proto_tcp.c | 19 +++++
net/netfilter/nf_flow_table_ip.c | 26 +++++++
net/netfilter/nft_flow_offload.c | 6 +-
.../selftests/net/netfilter/nft_flowtable.sh | 73 +++++++++++++++++++
5 files changed, 129 insertions(+), 2 deletions(-)
--
2.53.0
From: Pablo Neira Ayuso <pablo@netfilter.org> Date: 2026-09-11 14:16:19
On Fri, Sep 11, 2026 at 04:09:51PM +0200, Pablo Neira Ayuso wrote:
Hi Julius,
On Fri, Sep 11, 2026 at 12:10:04PM +0000, Julius Bairaktaris wrote:
quoted
Hi Pablo,
thanks for taking your time to review.
quoted
conntrack needs to see packets in both directions, are you assuming a
packet-based load balancer in front of it?
No, an asymmetric route is in front of it, with two subnets sharing one
L2 segment and the server answering over that link, so the router only
ever sees one direction.
I see this requirement to support asymmetric path keeps coming, but
how hard is really to maintain this TCP state machine to deal with all
possible scenarios? ie. invalid transitions, retransmissions, etc.
this all without having access to full TCP connection. Is it that you
need NAT and the stateless NAT in nftables does not fulfill your
requirements?
I can reply myself, you're targetting at offloading this flow via the
flowtable.
Let me take a look what can be done here for this assymetric case.