From: Wang Yufen <hidden> Date: 2022-03-14 12:34:48
A tcp socket in a sockmap. If user invokes bpf_map_delete_elem to delete
the sockmap element, the tcp socket will switch to use the TCP protocol
stack to send and receive packets. The switching process may cause some
issues, such as if some msgs exist in the ingress queue and are cleared
by sk_psock_drop(), the packets are lost, and the tcp data is abnormal.
Signed-off-by: Wang Yufen <redacted>
---
include/uapi/linux/bpf.h | 3 +++
kernel/bpf/syscall.c | 2 ++
net/core/sock_map.c | 3 +++
3 files changed, 8 insertions(+)
@@ -1218,6 +1218,9 @@ enum {/* Create a map that is suitable to be an inner map with dynamic max entries */BPF_F_INNER_MAP=(1U<<12),++/* This should only be used for bpf_map_delete_elem called by user. */+BPF_F_TCP_SOCKMAP=(1U<<13),};/* Flags for BPF_PROG_QUERY. */
From: Jakub Sitnicki <jakub@cloudflare.com> Date: 2022-03-14 15:35:24
On Mon, Mar 14, 2022 at 08:44 PM +08, Wang Yufen wrote:
A tcp socket in a sockmap. If user invokes bpf_map_delete_elem to delete
the sockmap element, the tcp socket will switch to use the TCP protocol
stack to send and receive packets. The switching process may cause some
issues, such as if some msgs exist in the ingress queue and are cleared
by sk_psock_drop(), the packets are lost, and the tcp data is abnormal.
Signed-off-by: Wang Yufen <redacted>
---
Can you please tell us a bit more about the life-cycle of the socket in
your workload? Questions that come to mind:
1) What triggers the removal of the socket from sockmap in your case?
2) Would it still be a problem if removal from sockmap did not cause any
packets to get dropped?
[...]
On Mon, Mar 14, 2022 at 08:44 PM +08, Wang Yufen wrote:
quoted
A tcp socket in a sockmap. If user invokes bpf_map_delete_elem to delete
the sockmap element, the tcp socket will switch to use the TCP protocol
stack to send and receive packets. The switching process may cause some
issues, such as if some msgs exist in the ingress queue and are cleared
by sk_psock_drop(), the packets are lost, and the tcp data is abnormal.
Signed-off-by: Wang Yufen <redacted>
---
Can you please tell us a bit more about the life-cycle of the socket in
your workload? Questions that come to mind:
1) What triggers the removal of the socket from sockmap in your case?
We use sk_msg to redirect with sock hash, like this:
skA redirect skB
Tx <-----------> skB,Rx
And construct a scenario where the packet sending speed is high, the
packet receiving speed is slow, so the packets are stacked in the ingress
queue on the receiving side. In this case, if run bpf_map_delete_elem() to
delete the sockmap entry, will trigger the following procedure:
sock_hash_delete_elem()
sock_map_unref()
sk_psock_put()
sk_psock_drop()
sk_psock_stop()
__sk_psock_zap_ingress()
__sk_psock_purge_ingress_msg()
2) Would it still be a problem if removal from sockmap did not cause any
packets to get dropped?
Yes, it still be a problem. If removal from sockmap did not cause any
packets to get dropped, packet receiving process switches to use TCP
protocol stack. The packets in the psock ingress queue cannot be received
by the user.
Thanks.
From: Jakub Sitnicki <jakub@cloudflare.com> Date: 2022-03-15 12:15:21
On Tue, Mar 15, 2022 at 03:24 PM +08, wangyufen wrote:
在 2022/3/14 23:30, Jakub Sitnicki 写道:
quoted
On Mon, Mar 14, 2022 at 08:44 PM +08, Wang Yufen wrote:
quoted
A tcp socket in a sockmap. If user invokes bpf_map_delete_elem to delete
the sockmap element, the tcp socket will switch to use the TCP protocol
stack to send and receive packets. The switching process may cause some
issues, such as if some msgs exist in the ingress queue and are cleared
by sk_psock_drop(), the packets are lost, and the tcp data is abnormal.
Signed-off-by: Wang Yufen <redacted>
---
Can you please tell us a bit more about the life-cycle of the socket in
your workload? Questions that come to mind:
1) What triggers the removal of the socket from sockmap in your case?
We use sk_msg to redirect with sock hash, like this:
skA redirect skB
Tx <-----------> skB,Rx
And construct a scenario where the packet sending speed is high, the
packet receiving speed is slow, so the packets are stacked in the ingress
queue on the receiving side. In this case, if run bpf_map_delete_elem() to
delete the sockmap entry, will trigger the following procedure:
sock_hash_delete_elem()
sock_map_unref()
sk_psock_put()
sk_psock_drop()
sk_psock_stop()
__sk_psock_zap_ingress()
__sk_psock_purge_ingress_msg()
quoted
2) Would it still be a problem if removal from sockmap did not cause any
packets to get dropped?
Yes, it still be a problem. If removal from sockmap did not cause any
packets to get dropped, packet receiving process switches to use TCP
protocol stack. The packets in the psock ingress queue cannot be received
by the user.
Thanks for the context. So, if I understand correctly, you want to avoid
breaking the network pipe by updating the sockmap from user-space.
This sounds awfully similar to BPF_MAP_FREEZE. Have you considered that?
From: Daniel Borkmann <daniel@iogearbox.net> Date: 2022-03-15 16:25:34
On 3/15/22 1:12 PM, Jakub Sitnicki wrote:
On Tue, Mar 15, 2022 at 03:24 PM +08, wangyufen wrote:
quoted
在 2022/3/14 23:30, Jakub Sitnicki 写道:
quoted
On Mon, Mar 14, 2022 at 08:44 PM +08, Wang Yufen wrote:
quoted
A tcp socket in a sockmap. If user invokes bpf_map_delete_elem to delete
the sockmap element, the tcp socket will switch to use the TCP protocol
stack to send and receive packets. The switching process may cause some
issues, such as if some msgs exist in the ingress queue and are cleared
by sk_psock_drop(), the packets are lost, and the tcp data is abnormal.
Signed-off-by: Wang Yufen <redacted>
---
Can you please tell us a bit more about the life-cycle of the socket in
your workload? Questions that come to mind:
1) What triggers the removal of the socket from sockmap in your case?
We use sk_msg to redirect with sock hash, like this:
skA redirect skB
Tx <-----------> skB,Rx
And construct a scenario where the packet sending speed is high, the
packet receiving speed is slow, so the packets are stacked in the ingress
queue on the receiving side. In this case, if run bpf_map_delete_elem() to
delete the sockmap entry, will trigger the following procedure:
sock_hash_delete_elem()
sock_map_unref()
sk_psock_put()
sk_psock_drop()
sk_psock_stop()
__sk_psock_zap_ingress()
__sk_psock_purge_ingress_msg()
quoted
2) Would it still be a problem if removal from sockmap did not cause any
packets to get dropped?
Yes, it still be a problem. If removal from sockmap did not cause any
packets to get dropped, packet receiving process switches to use TCP
protocol stack. The packets in the psock ingress queue cannot be received
by the user.
Thanks for the context. So, if I understand correctly, you want to avoid
breaking the network pipe by updating the sockmap from user-space.
This sounds awfully similar to BPF_MAP_FREEZE. Have you considered that?
+1
Aside from that, the patch as-is also fails BPF CI in a lot of places, please
make sure to check selftests:
https://github.com/kernel-patches/bpf/runs/5537367301?check_suite_focus=true
[...]
#145/73 sockmap_listen/sockmap IPv6 test_udp_redir:OK
#145/74 sockmap_listen/sockmap IPv6 test_udp_unix_redir:OK
#145/75 sockmap_listen/sockmap Unix test_unix_redir:OK
#145/76 sockmap_listen/sockmap Unix test_unix_redir:OK
./test_progs:test_ops_cleanup:1424: map_delete: expected EINVAL/ENOENT: Operation not supported
test_ops_cleanup:FAIL:1424
./test_progs:test_ops_cleanup:1424: map_delete: expected EINVAL/ENOENT: Operation not supported
test_ops_cleanup:FAIL:1424
#145/77 sockmap_listen/sockhash IPv4 TCP test_insert_invalid:FAIL
./test_progs:test_ops_cleanup:1424: map_delete: expected EINVAL/ENOENT: Operation not supported
test_ops_cleanup:FAIL:1424
./test_progs:test_ops_cleanup:1424: map_delete: expected EINVAL/ENOENT: Operation not supported
test_ops_cleanup:FAIL:1424
#145/78 sockmap_listen/sockhash IPv4 TCP test_insert_opened:FAIL
./test_progs:test_ops_cleanup:1424: map_delete: expected EINVAL/ENOENT: Operation not supported
test_ops_cleanup:FAIL:1424
./test_progs:test_ops_cleanup:1424: map_delete: expected EINVAL/ENOENT: Operation not supported
test_ops_cleanup:FAIL:1424
#145/79 sockmap_listen/sockhash IPv4 TCP test_insert_bound:FAIL
./test_progs:test_ops_cleanup:1424: map_delete: expected EINVAL/ENOENT: Operation not supported
test_ops_cleanup:FAIL:1424
./test_progs:test_ops_cleanup:1424: map_delete: expected EINVAL/ENOENT: Operation not supported
test_ops_cleanup:FAIL:1424
[...]
Thanks,
Daniel
From: Cong Wang <hidden> Date: 2022-03-16 00:36:57
On Tue, Mar 15, 2022 at 01:12:08PM +0100, Jakub Sitnicki wrote:
On Tue, Mar 15, 2022 at 03:24 PM +08, wangyufen wrote:
quoted
在 2022/3/14 23:30, Jakub Sitnicki 写道:
quoted
On Mon, Mar 14, 2022 at 08:44 PM +08, Wang Yufen wrote:
quoted
A tcp socket in a sockmap. If user invokes bpf_map_delete_elem to delete
the sockmap element, the tcp socket will switch to use the TCP protocol
stack to send and receive packets. The switching process may cause some
issues, such as if some msgs exist in the ingress queue and are cleared
by sk_psock_drop(), the packets are lost, and the tcp data is abnormal.
Signed-off-by: Wang Yufen <redacted>
---
Can you please tell us a bit more about the life-cycle of the socket in
your workload? Questions that come to mind:
1) What triggers the removal of the socket from sockmap in your case?
We use sk_msg to redirect with sock hash, like this:
skA redirect skB
Tx <-----------> skB,Rx
And construct a scenario where the packet sending speed is high, the
packet receiving speed is slow, so the packets are stacked in the ingress
queue on the receiving side. In this case, if run bpf_map_delete_elem() to
delete the sockmap entry, will trigger the following procedure:
sock_hash_delete_elem()
sock_map_unref()
sk_psock_put()
sk_psock_drop()
sk_psock_stop()
__sk_psock_zap_ingress()
__sk_psock_purge_ingress_msg()
quoted
2) Would it still be a problem if removal from sockmap did not cause any
packets to get dropped?
Yes, it still be a problem. If removal from sockmap did not cause any
packets to get dropped, packet receiving process switches to use TCP
protocol stack. The packets in the psock ingress queue cannot be received
by the user.
Thanks for the context. So, if I understand correctly, you want to avoid
breaking the network pipe by updating the sockmap from user-space.
This sounds awfully similar to BPF_MAP_FREEZE. Have you considered that?
Doesn't BPF_MAP_FREEZE only freeze write operations from syscalls?
For sockmap, receiving packets is not a part of map write operation.
The problem here is that skmsg can only be consumed when the socket is
still in the map, as it uses a separate queue and a separate type of
message (skmsg vs. skb). So, esstentially this behavior is by design.
Thanks.
On Tue, Mar 15, 2022 at 03:24 PM +08, wangyufen wrote:
quoted
在 2022/3/14 23:30, Jakub Sitnicki 写道:
quoted
On Mon, Mar 14, 2022 at 08:44 PM +08, Wang Yufen wrote:
quoted
A tcp socket in a sockmap. If user invokes bpf_map_delete_elem to delete
the sockmap element, the tcp socket will switch to use the TCP protocol
stack to send and receive packets. The switching process may cause some
issues, such as if some msgs exist in the ingress queue and are cleared
by sk_psock_drop(), the packets are lost, and the tcp data is abnormal.
Signed-off-by: Wang Yufen <redacted>
---
Can you please tell us a bit more about the life-cycle of the socket in
your workload? Questions that come to mind:
1) What triggers the removal of the socket from sockmap in your case?
We use sk_msg to redirect with sock hash, like this:
skA redirect skB
Tx <-----------> skB,Rx
And construct a scenario where the packet sending speed is high, the
packet receiving speed is slow, so the packets are stacked in the ingress
queue on the receiving side. In this case, if run bpf_map_delete_elem() to
delete the sockmap entry, will trigger the following procedure:
sock_hash_delete_elem()
sock_map_unref()
sk_psock_put()
sk_psock_drop()
sk_psock_stop()
__sk_psock_zap_ingress()
__sk_psock_purge_ingress_msg()
quoted
2) Would it still be a problem if removal from sockmap did not cause any
packets to get dropped?
Yes, it still be a problem. If removal from sockmap did not cause any
packets to get dropped, packet receiving process switches to use TCP
protocol stack. The packets in the psock ingress queue cannot be received
by the user.
Thanks for the context. So, if I understand correctly, you want to avoid
breaking the network pipe by updating the sockmap from user-space.
This sounds awfully similar to BPF_MAP_FREEZE. Have you considered that?
.
Sorry, I didn't notice this. I used BPF_MAP_FREEZE to verify, can solve
my problem, thanks.