[RFC 2/2] dlm: fix socket hang on send buffer limitation with socket error
flat view
From: Alexander Aring <aahringo@redhat.com>
Date: 2026-10-02 23:21:22
Also in:
gfs2
Subsystem:
distributed lock manager (dlm), filesystems (vfs and infrastructure), the rest · Maintainers:
Alexander Aring, David Teigland, Alexander Viro, Christian Brauner, Linus Torvalds
As Sashiko AI bot mentioned [0], a socket connection can get stuck when
the send buffer limitation is reached and a socket error simultaneously
occurs.
This issue can be reproduced with the following setup:
1. Set sk_sndbuf to SOCK_MIN_SNDBUF.
2. Slow down DLM connections using netem.
3. Use tcpkill to randomly send TCP resets to DLM connections.
I instrumented debug printouts to confirm that the sk_write_space()
notifier callbacks were being executed, using step 3 to trigger random
socket errors.
With these changes, I can no longer reproduce the hang.
Changes included:
- Removed sk_write_pending counting, as this should not be modified
at the socket application layer (or is at least unnecessary).
- Moved clearing the CF_SEND_PENDING bit—which allows re-queuing
swork (send worker for sendmsg())—to the sk_write_space() callback,
since this callback notifies us that the underlying socket is no longer
constrained by its send buffer.
- Handled the race condition between sendmsg() and evaluating
SOCK_NOSPACE after sendmsg(), where sk_write_space() could be called
in between, using CF_APP_LIMITED:
- In sk_write_space(), queue swork again if CF_APP_LIMITED is set.
- If CF_APP_LIMITED is not set, do nothing as send_to_sock() will handle
it, confirming the race occurred.
- Introduced new handling in lowcomms_error_report() when a socket error
occurs during send buffer limitation. If CF_APP_LIMITED is set, swork
will be re-queued, which will fail and trigger a reconnect.
- Added various comments explaining the interaction with CF_APP_LIMITED.
[0] https://lore.kernel.org/netdev/179090395863.434549.3668493667259120759@kernel.org/ (local)
Signed-off-by: Alexander Aring <aahringo@redhat.com>
---
fs/dlm/lowcomms.c | 46 ++++++++++++++++++++++++++++++++++++----------
1 file changed, 36 insertions(+), 10 deletions(-)
diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
index c3a414d2e32b..29ee1b52d0df 100644
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c@@ -521,12 +521,20 @@ static void lowcomms_write_space(struct sock *sk) sk_clear_nospace(sk); - spin_lock_bh(&con->writequeue_lock); - if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) - con->sock->sk->sk_write_pending--; - - lowcomms_queue_swork(con); - spin_unlock_bh(&con->writequeue_lock); + if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) { + /* signal to send again by clearing + * CF_SEND_PENDING and queue swork. + */ + spin_lock_bh(&con->writequeue_lock); + clear_bit(CF_SEND_PENDING, &con->flags); + lowcomms_queue_swork(con); + spin_unlock_bh(&con->writequeue_lock); + } else { + /* CF_APP_LIMITED is cleared, so send_to_sock() will + * simply reschedule work without hitting the + * CF_APP_LIMITED path. + */ + } } static void lowcomms_state_change(struct sock *sk)
@@ -623,6 +631,23 @@ static void lowcomms_error_report(struct sock *sk) break; } + /* if waiting on sk_write_space() and an sk_err occurs, the callback + * won't fire. Clear CF_SEND_PENDING and if CF_APP_LIMITED was set + * so resend tasks can re-queue swork, triggering a sendmsg() failure + * to initiate reconnection. + */ + if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) { + spin_lock_bh(&con->writequeue_lock); + clear_bit(CF_SEND_PENDING, &con->flags); + /* dlm_midcomms_unack_msg_resend() does not always + * trigger lowcomms_queue_swork() as it tries to + * avoid to put pending messages into the lowcomms + * sending buffer. Force it here again. + */ + lowcomms_queue_swork(con); + spin_unlock_bh(&con->writequeue_lock); + } + dlm_midcomms_unack_msg_resend(con->nodeid); listen_sock.sk_error_report(sk);
@@ -1391,18 +1416,19 @@ static int send_to_sock(struct connection *con) spin_lock_bh(&con->writequeue_lock); if (test_bit(SOCK_NOSPACE, &con->sock->flags) && !test_and_set_bit(CF_APP_LIMITED, &con->flags)) { - con->sock->sk->sk_write_pending++; - - clear_bit(CF_SEND_PENDING, &con->flags); spin_unlock_bh(&con->writequeue_lock); release_sock(con->sock->sk); - /* wait for write_space() event */ + /* wait for sk_write_space() event */ return DLM_IO_END; } spin_unlock_bh(&con->writequeue_lock); release_sock(con->sock->sk); + /* the sk_write_space() came in between sock_sendmsg() + * and check on SOCK_NOSPACE and the socket became + * writeable again so just resched swork. + */ return DLM_IO_RESCHED; } else if (ret < 0) { return ret;
--
2.43.0