Thread (2 messages) flat view 2 messages, 2 authors, 2011-11-13

Re: [PATCH 4/4] sunrpc: use SKB fragment destructors to delay completion until page is released by network stack.

From: "Michael S. Tsirkin" <mst@redhat.com>
Date: 2011-11-13 10:16:09
Also in: linux-nfs

On Fri, Nov 11, 2011 at 01:20:27PM +0000, Ian Campbell wrote:
On Fri, 2011-11-11 at 12:38 +0000, Michael S. Tsirkin wrote:
quoted
On Wed, Nov 09, 2011 at 03:02:07PM +0000, Ian Campbell wrote:
quoted
This prevents an issue where an ACK is delayed, a retransmit is queued (either
at the RPC or TCP level) and the ACK arrives before the retransmission hits the
wire. If this happens to an NFS WRITE RPC then the write() system call
completes and the userspace process can continue, potentially modifying data
referenced by the retransmission before the retransmission occurs.

Signed-off-by: Ian Campbell <redacted>
Acked-by: Trond Myklebust <redacted>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: Neil Brown <redacted>
Cc: "J. Bruce Fields" <redacted>
Cc: linux-nfs@vger.kernel.org
Cc: netdev@vger.kernel.org
So this blocks the system call until all page references
are gone, right?
Right. The alternative is to return to userspace while the network stack
still has a reference to the buffer which was passed in -- that's the
exact class of problem this patch is supposed to fix.
BTW, the increased latency and the overhead extra wakeups might for some
workloads be greater than the cost of the data copy.
quoted
But, there's no upper limit on how long the
page is referenced, correct?
Correct.
quoted
 consider a bridged setup
with an skb queued at a tap device - this cause one process
to block another one by virtue of not consuming a cloned skb?
Hmm, yes.

One approach might be to introduce the concept of an skb timeout to the
stack as a whole and cancel (or deep copy) after that timeout occurs.
That's going to be tricky though I suspect...
Further, an application might use signals such as SIGALARM,
delaying them significantly will cause trouble.
A simpler option would be to have an end points such as a tap device
Which end points would that be? Doesn't this affect a packet socket
with matching filters? A userspace TCP socket that happens to
reside on the same box?  It also seems that at least with a tap device
an skb can get queued in a qdisk, also indefinitely, right?
which can swallow skbs for arbitrary times implement a policy in this
regard, either to deep copy or drop after a timeout?

Ian.
Or deep copy immediately?

-- 
MST
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help