Re: [RFC PATCH 0/1] close(): stop exposing non-retryable EINTR
From: Rich Felker <dalias@libc.org>
Date: 2026-09-17 17:48:41
Also in:
linux-fsdevel, lkml
On Thu, Sep 17, 2026 at 10:03:36AM -0700, Andy Lutomirski wrote:
On Thu, Sep 17, 2026 at 3:37 AM Mikko Rantalainen [off-list ref] wrote:quoted
On 9/17/26 02:16, Andy Lutomirski wrote:quoted
On Mon, Sep 14, 2026 at 2:44 AM Mikko Rantalainen [off-list ref] wrote:quoted
NFS and devices make the question more interesting, though. Some implementations put meaningful state transitions into their close path. Tape devices are an obvious historical example: depending on the device/mode, close may write filemarks, flush tape buffers, rewind, etc. Those operations are not necessarily replaceable by fsync() before close().I realize this would be complex and maybe a can of worms, but maybe we should have a new improved syscall here. close() is indeed deeply problematic. On the one hand, we have the actual task of closing an fd in the sense of removing it from the table. This should really be doable without blocking or without side effects (except possibly for zapping old-style POSIX locks -- it doesn't really make sense to be able to close all fds to a file while still keeping it locked, especially since we report the pid of the lock-holding process). On the other hand, close has actual *meaningful* effects, many of which you've mentioned in your email. IMO it would be really nice to be able to explicitly *do* those effects separately from closing the fd, maybe even asynchronously via io_uring. Would it be so bad to have a new operation to do (possibly with moderately fine control) the close work and another one to just release fd- and process-associated locks and drop the fd?I think introducing yet another syscall would be bad replacement for the *inability to decide the exact semantics* we want to use for the current syscalls, especially close(). It's pretty clear that close() is poorly defined right now. The only thing that we know for sure is that it releases the file descriptor and in case of Linux, this happens for sure unless EBADF is returned.I think I mostly agree with your analysis except that I really don't like the ENOSPC. The kernel actually has the ability to commit to having space for pending writes on an fd *without syncing*, then I think that capability should be expose to userspace. fsync is a pretty poor alternative for code that just wants to make sure that, in the absence of a bug, crash or physical failure, the data won't be lost.
Yes, if close() can give ENOSPC, that's indicative of a bug in write(). I think the problem is everyone likes their bad NFS implementations cutting corners to be fast... Rich