[BUG] net/unix: blocking AF_UNIX sendmsg can wedge permanently via send-buffer exhaustion while GC triggers stay unreachable
From: Waseem Abbas Khan <hidden>
Date: 2026-09-20 19:46:48
Subject: [BUG] af_unix: blocking sendmsg() can wedge permanently - GC
trigger starvation with pinned fd cycles
Hi,
I'd like to report a reproducible issue in net/unix/garbage.c that
affects all current
kernels including latest mainline (verified on 6.12.93, 6.12.104,
7.2-rc5, 7.2.6 and a
6.6-based kernel): a userspace process can permanently wedge its OWN blocking
sendmsg() by creating only "pinned" fd cycles, because no GC trigger
is ever reachable
before the sender's socket send-buffer fills with skbs that can never be freed.
Reproducer (userspace, unprivileged; no threads, no races, ~140 rounds):
X = socket(AF_UNIX, SOCK_DGRAM, 0);
loop {
A, B = fresh AF_UNIX/SOCK_DGRAM sockets bound to abstract names;
sendmsg(X -> A, SCM_RIGHTS carrying B); // blocking send
sendmsg(X -> B, SCM_RIGHTS carrying A); // blocking send
close(A); close(B);
}
After ~139 rounds the third sendmsg blocks forever:
task stack: unix_dgram_sendmsg -> sock_alloc_send_pskb+0x161
(schedule_timeout loop)
Observed state at the wedge (via QEMU monitor + symbols):
- sender SO_MEMINFO: wmem_alloc=213504 >= sk_sndbuf=212992, wmem_queued=212992
- /proc/net/unix: ~271 leftover sockets; process has only 9 fds
- unix_tot_inflight = 278; gc_in_progress = 0; no unix_gc work queued
- unix_graph_maybe_cyclic = 1
Analysis (why it can never recover on its own):
- close(A)/close(B) never fully release the sockets: the fpl in the
*other* socket's
queue holds the only remaining file reference, so
unix_release_sock() - and with it
the only reachable GC trigger (unix_gc() called on release when
unix_tot_inflight
is non-zero) - is never reached. The cycles are fully "pinned".
- The send-path trigger requires unix_tot_inflight >
UNIX_INFLIGHT_TRIGGER_GC (16000).
- The send-buffer limit (sk_sndbuf, default 212992) is exhausted after
~278 unconsumed
fd-skbs - i.e. roughly 60x earlier than the 16000 trigger - so a
blocking sender
parks in sock_alloc_send_pskb() forever before any GC can be scheduled.
- The per-user Unix inflight sane threshold (SCM_MAX_FD*8 = 2024, used for the
flush/wait penalty) is likewise unreachable: the wmem stall happens
~7x earlier.
Falsification / healing experiment:
A companion process that closes an *unpinned* socketpair every 500ms forces
release-path GCs; with it, the same loop completes 2,000,000 rounds
without wedging,
and the sender's wmem_alloc shows a clean leak/collect sawtooth. This
confirms the
issue is purely GC trigger starvation, not memory corruption (2M
rounds KASAN-clean).
The wedge persists across all four kernels above; nothing in mainline history
(garbage.c commits up to 4a4263dfeaba, 2026-09-12) addresses these triggers.
Impact:
- Any app doing blocking fd-passing cycles in a quiet system
(container, embedded,
single-app) permanently hangs its send path; the pinned
file/skbs/sockets (bounded
by RLIMIT_NOFILE per user) leak until some unrelated socket release
triggers a GC.
- Related in spirit to the historical 2008/2010 sendmsg-GC stall
reports; the current
trigger set leaves an unreachable-trigger window that this
reproducer lands in.
Possible fix directions:
- Have the wmem wait in unix sendmsg paths schedule/flush the GC
before/while parking
(e.g. unix_schedule_gc(user) when sock_alloc would block), or
- trigger on pinned-cycle conditions (e.g. consider graph state / cyclic_sccs at
lower inflight thresholds), or
- align thresholds so the wmem cap cannot be reached while GC triggers
stay unreachable.
I can provide the exact reproducer source (gcracer9 mode 0) and the
healer companion
used for the experiments, plus monitor-derived state dumps.
Thanks,
[reporter]