Thread (26 messages) 26 messages, 2 authors, 12d ago

Re: [PATCH net-next v3 13/13] net/rds: hold a connection reference from struct rds_incoming

From: netdev-bot+sashiko@kernel.org
Date: 2026-09-17 09:39:06
Also in: linux-rdma

Thank you for your contribution! Sashiko AI review found 3 potential
issue(s) to consider:

Critical: 0 · High: 1 · Medium: 0 · Low: 2

- [High] rds_send_probe() (net/rds/send.c) now takes a connection
  reference for rm->m_inc.i_conn (line 1583) while linking the probe…
- [Low] The cross-netns comment in rds_recv_incoming()
  (net/rds/recv.c:424-431) is now factually wrong: it justifies dropping…
- [Low] The commit fixes a KASAN-reported slab-use-after-free
  (rds6_inc_info_copy() via rds_info_getsockopt(), plus a second panic…

--

net/rds: hold a connection reference from struct rds_incoming

This makes every assignment of struct rds_incoming::i_conn take a
connection reference, dropping it from rds_inc_put() for received
messages and from a new rds_message_free() helper for m_inc.  The goal is
to keep a connection allocated while incs that point at it sit on socket
receive or send queues.

Closes: https://lore.kernel.org/netdev/20260720184955.3008978-1-nicoyip.dev@gmail.com/ (local)

A question on the trailers before the diff:
Reported-by: Chengfeng Ye <redacted>
Closes: https://lore.kernel.org/netdev/20260720184955.3008978-1-nicoyip.dev@gmail.com/ (local)
[Severity: Low]
The changelog fixes a KASAN-reported slab-use-after-free (rds6_inc_info_copy()
via rds_info_getsockopt(), plus the rds_ib_inc_free()/rds_clear_recv_queue()
panic) and carries Reported-by:/Closes:, but there is no Fixes: tag.

The unreferenced i_conn assignment goes back to bdbe6fbc6a2f ("RDS:
recv.c").  Would adding

  Fixes: bdbe6fbc6a2f ("RDS: recv.c")

be worthwhile so backporters can see which releases are affected, even
with the "This is not a stable candidate" note in the changelog that
explains this particular fix depends on the connection reference counting
added earlier in the series?
quoted hunk ↗ jump to hunk
diff --git a/net/rds/recv.c b/net/rds/recv.c
index 6204e577a90ae..b031c0b43af83 100644
--- a/net/rds/recv.c
+++ b/net/rds/recv.c
[ ... ]
quoted hunk ↗ jump to hunk
@@ -325,6 +337,13 @@ void rds_recv_incoming(struct rds_connection *conn, struct in6_addr *saddr,
 	unsigned long flags;
 	struct rds_conn_path *cp;
 
+	/* every caller initialized @inc with rds_inc_init() or
+	 * rds_inc_path_init() first, so i_conn already holds a reference.
+	 * Take the new one before dropping the old, so that re-pointing an
+	 * inc at the connection it already refers to cannot free it.
+	 */
+	rds_conn_get(conn);
+	rds_conn_put(inc->i_conn);
 	inc->i_conn = conn;
 	inc->i_rx_jiffies = jiffies;
 	if (conn->c_trans->t_mp_capable)
[Severity: Low]
Does the cross-netns comment further down in rds_recv_incoming() still
describe the code correctly after this change?  It currently reads:

	/*
	 * rds_find_bound() uses a global (netns-agnostic) hash table.
	 * An RDS connection created in netns A can match a socket bound
	 * in the init netns, delivering inc cross-netns with inc->i_conn
	 * pointing into netns A.  When cleanup_net() then frees that conn,
	 * any subsequent dereference of inc->i_conn is a use-after-free.
	 * Drop the inc if the receiving socket lives in a different netns.
	 */

With i_conn now owning a reference from rds_inc_init()/rds_inc_path_init(),
transferred here and released only after c_trans->inc_free() in
rds_inc_put(), can cleanup_net() -> rds_conn_destroy() still free the
connection while the inc is queued?  The test itself may still be wanted
for netns isolation and because c_net would outlive the freed struct net,
but should the stated rationale be updated?
quoted hunk ↗ jump to hunk
diff --git a/net/rds/send.c b/net/rds/send.c
index 94d6ac174dde8..f59a487b079d5 100644
--- a/net/rds/send.c
+++ b/net/rds/send.c
[ ... ]
quoted hunk ↗ jump to hunk
@@ -960,6 +962,7 @@ static int rds_send_queue_rm(struct rds_sock *rs, struct rds_connection *conn,
 		/* The code ordering is a little weird, but we're
 		   trying to minimize the time we hold c_lock */
 		rds_message_populate_header(&rm->m_inc.i_hdr, sport, dport, 0);
+		rds_conn_get(conn);	/* put in rds_message_put() */
 		rm->m_inc.i_conn = conn;
 		rm->m_inc.i_conn_path = cp;
 		rds_message_addref(rm);
@@ -1577,6 +1580,7 @@ rds_send_probe(struct rds_conn_path *cp, __be16 sport,
 	list_add_tail(&rm->m_conn_item, &cp->cp_send_queue);
 	set_bit(RDS_MSG_ON_CONN, &rm->m_flags);
 	rds_message_addref(rm);
+	rds_conn_get(cp->cp_conn);	/* put in rds_message_put() */
 	rm->m_inc.i_conn = cp->cp_conn;
 	rm->m_inc.i_conn_path = cp;
[Severity: High]
Can this new reference in rds_send_probe() be stranded when the
connection is being destroyed?

rds_send_queue_rm() takes the same kind of reference for the same kind of
enqueue, but tests the destroy predicate under cp_lock first:

	spin_lock(&cp->cp_lock);
	if (rds_destroy_pending(conn)) {
		spin_unlock(&cp->cp_lock);
		*queued = -EAGAIN;
		goto unlock;
	}
	...
	rds_conn_get(conn);	/* put in rds_message_put() */

rds_send_probe() has no such test inside its cp_lock section, and
rds_conn_path_quiesce() drains cp_send_queue exactly once under that same
lock:

net/rds/connection.c:rds_conn_path_quiesce() {
	spin_lock_irqsave(&cp->cp_lock, flags);
	list_splice_init(&cp->cp_send_queue, &purge);
	spin_unlock_irqrestore(&cp->cp_lock, flags);
	list_for_each_entry_safe(rm, rtmp, &purge, m_conn_item) {
		...
		rds_message_put(rm);
	}
}

So if rds_send_probe() acquires cp_lock after that splice, the probe rm is
added to a queue nobody drains again, its RDS_MSG_ON_CONN reference is
never dropped, and the trailing

	rcu_read_lock();
	if (!rds_destroy_pending(cp->cp_conn))
		queue_delayed_work(cp->cp_wq, &cp->cp_send_w, 1);
	rcu_read_unlock();

skips the work, so rds_message_put() never reaches the new
rds_message_free().  Does that leak the struct rds_connection reference
taken above?

The window looks reachable from user context and from remote input:

rds_sendmsg() -> rds_send_ping() -> rds_send_probe()
rds_recv_incoming() -> rds_send_pong() -> rds_send_probe()
rds_tcp_accept_one() -> rds_send_probe()

with the rds_destroy_pending() check in rds_sendmsg() done before the
user data copy, well ahead of the rds_send_ping() call.

If the reference is stranded, does t_conn_count then stay non-zero and
leave rds_conn_wait_conns_freed() looping forever on module unload?

net/rds/connection.c:rds_conn_wait_conns_freed() {
	while (!wait_event_timeout(rds_conn_freed_waitq,
				   !atomic_read(&trans->t_conn_count),
				   msecs_to_jiffies(RDS_CONN_FREE_POLL_MS))) {
	...
}

Would adding the same rds_destroy_pending(cp->cp_conn) test inside
rds_send_probe()'s cp_lock section, before the list_add_tail() and
rds_conn_get(), keep the six reference sites consistent?

The commit message says of the six assignment sites that "each of them now
takes a reference"; should it also say what happens when a probe races
teardown?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260914033719.138057-1-achender%40kernel.org
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help