[PATCH] bpf, sockmap: Fix self-redirect copied_seq double-counting
From: Geliang Tang <geliang@kernel.org>
Date: 2026-08-29 02:00:49
Also in:
bpf
Subsystem:
bpf [l7 framework] (sockmap), networking [general], the rest · Maintainers:
John Fastabend, Jakub Sitnicki, Jiayuan Chen, "David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, Linus Torvalds
From: Geliang Tang <redacted>
When a BPF stream_verdict program redirects an skb back to the same
socket (self-redirect with BPF_F_INGRESS), sk_psock_verdict_apply()
calls tcp_eat_skb() which advances tcp_sk->copied_seq. However, the
skb is then delivered to the socket's psock ingress queue and later
read by tcp_bpf_recvmsg_parser(), which also advances copied_seq via
the copied_from_self accounting path. This double-counting causes
copied_seq to advance by 2x the actual data length, triggering:
TCP recvmsg seq # bug 2: copied BF2E806, seq BF2E7FD, \
rcvnxt BF2E806, fl 0
WARNING: net/ipv4/tcp.c:2745 at tcp_recvmsg_locked+0x72b/0x2640
Call Trace:
tcp_recvmsg+0x10a/0x500
sock_recvmsg+0x168/0x1d0
__sys_recvfrom+0x19a/0x2a0
__x64_sys_recvfrom+0xe4/0x1f0
do_syscall_64+0xf7/0x530
entry_SYSCALL_64_after_hwframe+0x77/0x7f
cleanup rbuf bug: copied BF2E806 seq BF2E806 rcvnxt BF2E806
WARNING: net/ipv4/tcp.c:1609 at tcp_cleanup_rbuf+0xf2/0x1c0
Call Trace:
tcp_recvmsg_locked+0x8d1/0x2640
tcp_recvmsg+0x10a/0x500
sock_recvmsg+0x168/0x1d0
__sys_recvfrom+0x19a/0x2a0
__x64_sys_recvfrom+0xe4/0x1f0
do_syscall_64+0xf7/0x530
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Fix this by checking if the redirect destination is the same socket.
For self-redirect (dst == psock->sk), skip tcp_eat_skb() since the
copied_seq will be advanced when the data is actually read from the
ingress queue. For cross-socket redirects, tcp_eat_skb() is still
needed to account for data leaving the source socket.
Fixes: e5c6de5fa025 ("bpf, sockmap: Incorrectly handling copied_seq")
Signed-off-by: Geliang Tang <redacted>
---
Hi,
I encountered this while adding MPTCP BPF sockmap support. The existing
TCP sockmap selftests don't cover self-redirect, but the MPTCP tests do,
exposing this latent issue.
With this fix, both TCP and MPTCP tests pass, validating self-redirect
functionality.
---
net/core/skmsg.c | 8 ++++++--
1 file changed, 6 insertions(+), 2 deletions(-)
diff --git a/net/core/skmsg.c b/net/core/skmsg.c
index 2521b643fa05..5fa7b9639eef 100644
--- a/net/core/skmsg.c
+++ b/net/core/skmsg.c@@ -1039,10 +1039,14 @@ static int sk_psock_verdict_apply(struct sk_psock *psock, struct sk_buff *skb, goto out_free; } break; - case __SK_REDIRECT: - tcp_eat_skb(psock->sk, skb); + case __SK_REDIRECT: { + struct sock *dst = skb_bpf_redirect_fetch(skb); + + if (dst != psock->sk) + tcp_eat_skb(psock->sk, skb); err = sk_psock_skb_redirect(psock, skb); break; + } case __SK_DROP: default: out_free:
--
2.53.0