Thread (4 messages) 4 messages, 1 author, 5d ago

[RFC 2/2] dlm: fix socket hang on send buffer limitation with socket error

flat view

From: Alexander Aring <aahringo@redhat.com>
Date: 2026-10-02 23:21:22
Also in: gfs2
Subsystem: distributed lock manager (dlm), filesystems (vfs and infrastructure), the rest · Maintainers: Alexander Aring, David Teigland, Alexander Viro, Christian Brauner, Linus Torvalds

As Sashiko AI bot mentioned [0], a socket connection can get stuck when
the send buffer limitation is reached and a socket error simultaneously
occurs.

This issue can be reproduced with the following setup:
1. Set sk_sndbuf to SOCK_MIN_SNDBUF.
2. Slow down DLM connections using netem.
3. Use tcpkill to randomly send TCP resets to DLM connections.

I instrumented debug printouts to confirm that the sk_write_space()
notifier callbacks were being executed, using step 3 to trigger random
socket errors.

With these changes, I can no longer reproduce the hang.

Changes included:
- Removed sk_write_pending counting, as this should not be modified
  at the socket application layer (or is at least unnecessary).
- Moved clearing the CF_SEND_PENDING bit—which allows re-queuing
  swork (send worker for sendmsg())—to the sk_write_space() callback,
  since this callback notifies us that the underlying socket is no longer
  constrained by its send buffer.
- Handled the race condition between sendmsg() and evaluating
  SOCK_NOSPACE after sendmsg(), where sk_write_space() could be called
  in between, using CF_APP_LIMITED:
  - In sk_write_space(), queue swork again if CF_APP_LIMITED is set.
  - If CF_APP_LIMITED is not set, do nothing as send_to_sock() will handle
    it, confirming the race occurred.
- Introduced new handling in lowcomms_error_report() when a socket error
  occurs during send buffer limitation. If CF_APP_LIMITED is set, swork
  will be re-queued, which will fail and trigger a reconnect.
- Added various comments explaining the interaction with CF_APP_LIMITED.

[0] https://lore.kernel.org/netdev/179090395863.434549.3668493667259120759@kernel.org/ (local)

Signed-off-by: Alexander Aring <aahringo@redhat.com>
---
 fs/dlm/lowcomms.c | 46 ++++++++++++++++++++++++++++++++++++----------
 1 file changed, 36 insertions(+), 10 deletions(-)
diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
index c3a414d2e32b..29ee1b52d0df 100644
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c
@@ -521,12 +521,20 @@ static void lowcomms_write_space(struct sock *sk)
 
 	sk_clear_nospace(sk);
 
-	spin_lock_bh(&con->writequeue_lock);
-	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags))
-		con->sock->sk->sk_write_pending--;
-
-	lowcomms_queue_swork(con);
-	spin_unlock_bh(&con->writequeue_lock);
+	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) {
+		/* signal to send again by clearing
+		 * CF_SEND_PENDING and queue swork.
+		 */
+		spin_lock_bh(&con->writequeue_lock);
+		clear_bit(CF_SEND_PENDING, &con->flags);
+		lowcomms_queue_swork(con);
+		spin_unlock_bh(&con->writequeue_lock);
+	} else {
+		/* CF_APP_LIMITED is cleared, so send_to_sock() will
+		 * simply reschedule work without hitting the
+		 * CF_APP_LIMITED path.
+		 */
+	}
 }
 
 static void lowcomms_state_change(struct sock *sk)
@@ -623,6 +631,23 @@ static void lowcomms_error_report(struct sock *sk)
 		break;
 	}
 
+	/* if waiting on sk_write_space() and an sk_err occurs, the callback
+	 * won't fire. Clear CF_SEND_PENDING and if CF_APP_LIMITED was set
+	 * so resend tasks can re-queue swork, triggering a sendmsg() failure
+	 * to initiate reconnection.
+	 */
+	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) {
+		spin_lock_bh(&con->writequeue_lock);
+		clear_bit(CF_SEND_PENDING, &con->flags);
+		/* dlm_midcomms_unack_msg_resend() does not always
+		 * trigger lowcomms_queue_swork() as it tries to
+		 * avoid to put pending messages into the lowcomms
+		 * sending buffer. Force it here again.
+		 */
+		lowcomms_queue_swork(con);
+		spin_unlock_bh(&con->writequeue_lock);
+	}
+
 	dlm_midcomms_unack_msg_resend(con->nodeid);
 
 	listen_sock.sk_error_report(sk);
@@ -1391,18 +1416,19 @@ static int send_to_sock(struct connection *con)
 		spin_lock_bh(&con->writequeue_lock);
 		if (test_bit(SOCK_NOSPACE, &con->sock->flags) &&
 		    !test_and_set_bit(CF_APP_LIMITED, &con->flags)) {
-			con->sock->sk->sk_write_pending++;
-
-			clear_bit(CF_SEND_PENDING, &con->flags);
 			spin_unlock_bh(&con->writequeue_lock);
 			release_sock(con->sock->sk);
 
-			/* wait for write_space() event */
+			/* wait for sk_write_space() event */
 			return DLM_IO_END;
 		}
 		spin_unlock_bh(&con->writequeue_lock);
 		release_sock(con->sock->sk);
 
+		/* the sk_write_space() came in between sock_sendmsg()
+		 * and check on SOCK_NOSPACE and the socket became
+		 * writeable again so just resched swork.
+		 */
 		return DLM_IO_RESCHED;
 	} else if (ret < 0) {
 		return ret;
-- 
2.43.0
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help