Thread (53 messages) 53 messages, 6 authors, 2013-01-11

Re: ppoll() stuck on POLLIN while TCP peer is sending

From: Mel Gorman <mgorman@suse.de>
Date: 2013-01-09 13:42:54
Also in: linux-mm, lkml

On Tue, Jan 08, 2013 at 06:32:29PM -0800, Eric Dumazet wrote:
On Tue, 2013-01-08 at 18:14 -0800, Eric Dumazet wrote:
quoted
On Tue, 2013-01-08 at 23:23 +0000, Eric Wong wrote:
quoted
Mel Gorman [off-list ref] wrote:
quoted
Please try the following patch. However, even if it works the benefit of
capture may be so marginal that partially reverting it and simplifying
compaction.c is the better decision.
I already got my VM stuck on this one.  I had two twosleepy instances,
2774 was the one that got stuck (also confirmed by watching top).

Btw, have you been able to reproduce this on your end?

I think the easiest reproduction on my 2-core VM is by running 2
twosleepy processes and doing the following to dirty a lot of pages:
Given the persistent sk_stream_wait_memory() traces I suspect a plain
TCP bug, triggered by some extra wait somewhere.

Please mm guys don't spend too much time right now, I'll try to
reproduce the problem.

Don't be confused by sk_stream_wait_memory() name.
A thread is stuck here because TCP stack is failing to wake it.
Hmm, it seems sk_filter() can return -ENOMEM because skb has the
pfmemalloc() set.
The skb should not have pfmemalloc set in most cases, particularly after
cfd19c5a (mm: only set page->pfmemalloc when ALLOC_NO_WATERMARKS was used)
but the capture patch also failed to clear pfmemalloc properly so it could
be set in error.

-- 
Mel Gorman
SUSE Labs
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help