Thread (8 messages) flat view 8 messages, 3 authors, 7d ago
COOLING7d

[PATCH net 1/2] tls: don't over-fill the plaintext sk_msg ring in tls_sw_sendmsg_splice()

From: chanyoung <hidden>
Date: 2026-07-26 10:56:49
Also in: linux-kselftest, stable
Subsystem: networking [general], networking [tls], the rest · Maintainers: "David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, John Fastabend, Sabrina Dubroca, Linus Torvalds

tls_sw_sendmsg_splice() appends pages to the open record's plaintext
sk_msg ring with sk_msg_page_add(), which performs no fullness check of
its own, and the loop only tests sk_msg_full() at the bottom of its
do-while.

If the ring is already full when the function is entered, the first
sk_msg_page_add() writes the reserved slot and sk_msg_iter_next() wraps
sg.end around to sg.start.  sk_msg_iter_dist() then returns 0, so
sk_msg_full() reports the ring as empty, the loop keeps running, and each
further add overwrites a live entry without putting its page reference
while sg.size keeps growing.  sg.size is then larger than the data
reachable by walking the logical [sg.start, sg.end) ring.

tls_push_record() marks the end of the scatterlist at the logical last
entry but passes the inflated msg_pl->sg.size to tls_do_encryption() as
cryptlen, so the AEAD scatterwalk runs past the end-marked entry and
dereferences the NULL returned by sg_next():

  BUG: kernel NULL pointer dereference, address: 0000000000000008
  CPU: 1 UID: 1000 PID: 204 Comm: exploit Not tainted 7.2.0-rc4+ #1 PREEMPTLAZY
  RIP: 0010:memcpy_from_scatterwalk+0x32/0xc0
  Call Trace:
   <TASK>
   skcipher_walk_next+0x1d1/0x2c0
   gcm_encrypt_aesni_avx+0x1e9/0x220
   bpf_exec_tx_verdict+0x3bb/0x860
   tls_sw_sendmsg+0xa1a/0xca0
   __sys_sendto+0x1da/0x1f0
   do_syscall_64+0xdc/0x520
   entry_SYSCALL_64_after_hwframe+0x76/0x7e
   </TASK>

An unprivileged user can reach this on a plain loopback TCP socket with
the "tls" ULP attached.  A full but unpushed plaintext ring survives
across a syscall through the copy path: sk_msg_clone() returns 0 rather
than -ENOSPC for the frag that makes the ring exactly full, because its
guard is "if (i == src->sg.end && len)" and len reaches 0 as that frag is
added, so full_record is never set and MSG_MORE keeps eor clear.  Since
record_room is a byte count, a frag-exhausted ring that holds only a few
hundred bytes still admits the next splice(), which then re-enters
tls_sw_sendmsg_splice() on a full ring.

The caller already handles a ring that becomes full during the splice by
testing sk_msg_full() afterwards and setting full_record to push the
record, so the loop condition only needs to be evaluated before the first
sk_msg_page_add() rather than after it.  Turn the do-while into a while
loop: when the ring is full on entry the function returns without adding
anything, the caller pushes the record, and the next iteration of the
caller's loop starts from a fresh, empty ring.

Cc: stable@vger.kernel.org
Fixes: fe1e81d4f73b ("tls/sw: Support MSG_SPLICE_PAGES")
Signed-off-by: Chanyoung Park <redacted>
---
 net/tls/tls_sw.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index d4afc90fd79..0c413d05bb1 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -738,7 +738,7 @@ static int tls_sw_sendmsg_splice(struct sock *sk, struct msghdr *msg,
 {
 	struct page *page = NULL, **pages = &page;
 
-	do {
+	while (try_to_copy && !sk_msg_full(msg_pl)) {
 		ssize_t part;
 		size_t off;
 
@@ -758,7 +758,7 @@ static int tls_sw_sendmsg_splice(struct sock *sk, struct msghdr *msg,
 		sk_mem_charge(sk, part);
 		*copied += part;
 		try_to_copy -= part;
-	} while (try_to_copy && !sk_msg_full(msg_pl));
+	}
 
 	return 0;
 }
-- 
2.43.0
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help