[BUG] net/unix: blocking AF_UNIX sendmsg can wedge permanently via send-buffer exhaustion while GC triggers stay unreachable

From: Waseem Abbas Khan <hidden>
Date: 2026-09-20 19:46:48

Subject: [BUG] af_unix: blocking sendmsg() can wedge permanently - GC
trigger starvation with pinned fd cycles

Hi,

I'd like to report a reproducible issue in net/unix/garbage.c that
affects all current
kernels including latest mainline (verified on 6.12.93, 6.12.104,
7.2-rc5, 7.2.6 and a
6.6-based kernel): a userspace process can permanently wedge its OWN blocking
sendmsg() by creating only "pinned" fd cycles, because no GC trigger
is ever reachable
before the sender's socket send-buffer fills with skbs that can never be freed.

Reproducer (userspace, unprivileged; no threads, no races, ~140 rounds):

  X = socket(AF_UNIX, SOCK_DGRAM, 0);
  loop {
      A, B = fresh AF_UNIX/SOCK_DGRAM sockets bound to abstract names;
      sendmsg(X -> A, SCM_RIGHTS carrying B); // blocking send
      sendmsg(X -> B, SCM_RIGHTS carrying A); // blocking send
      close(A); close(B);
  }

After ~139 rounds the third sendmsg blocks forever:

  task stack: unix_dgram_sendmsg -> sock_alloc_send_pskb+0x161
(schedule_timeout loop)

Observed state at the wedge (via QEMU monitor + symbols):
  - sender SO_MEMINFO: wmem_alloc=213504 >= sk_sndbuf=212992, wmem_queued=212992
  - /proc/net/unix: ~271 leftover sockets; process has only 9 fds
  - unix_tot_inflight = 278; gc_in_progress = 0; no unix_gc work queued
  - unix_graph_maybe_cyclic = 1

Analysis (why it can never recover on its own):
- close(A)/close(B) never fully release the sockets: the fpl in the
*other* socket's
  queue holds the only remaining file reference, so
unix_release_sock() - and with it
  the only reachable GC trigger (unix_gc() called on release when
unix_tot_inflight
  is non-zero) - is never reached. The cycles are fully "pinned".
- The send-path trigger requires unix_tot_inflight >
UNIX_INFLIGHT_TRIGGER_GC (16000).
- The send-buffer limit (sk_sndbuf, default 212992) is exhausted after
~278 unconsumed
  fd-skbs - i.e. roughly 60x earlier than the 16000 trigger - so a
blocking sender
  parks in sock_alloc_send_pskb() forever before any GC can be scheduled.
- The per-user Unix inflight sane threshold (SCM_MAX_FD*8 = 2024, used for the
  flush/wait penalty) is likewise unreachable: the wmem stall happens
~7x earlier.

Falsification / healing experiment:
A companion process that closes an *unpinned* socketpair every 500ms forces
release-path GCs; with it, the same loop completes 2,000,000 rounds
without wedging,
and the sender's wmem_alloc shows a clean leak/collect sawtooth. This
confirms the
issue is purely GC trigger starvation, not memory corruption (2M
rounds KASAN-clean).
The wedge persists across all four kernels above; nothing in mainline history
(garbage.c commits up to 4a4263dfeaba, 2026-09-12) addresses these triggers.

Impact:
- Any app doing blocking fd-passing cycles in a quiet system
(container, embedded,
  single-app) permanently hangs its send path; the pinned
file/skbs/sockets (bounded
  by RLIMIT_NOFILE per user) leak until some unrelated socket release
triggers a GC.
- Related in spirit to the historical 2008/2010 sendmsg-GC stall
reports; the current
  trigger set leaves an unreachable-trigger window that this
reproducer lands in.

Possible fix directions:
- Have the wmem wait in unix sendmsg paths schedule/flush the GC
before/while parking
  (e.g. unix_schedule_gc(user) when sock_alloc would block), or
- trigger on pinned-cycle conditions (e.g. consider graph state / cyclic_sccs at
  lower inflight thresholds), or
- align thresholds so the wmem cap cannot be reached while GC triggers
stay unreachable.

I can provide the exact reproducer source (gcracer9 mode 0) and the
healer companion
used for the experiments, plus monitor-derived state dumps.

Thanks,
[reporter]
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help