Re: [PATCH] btrfs: Avoid BUG_ON()s because of ENOMEM caused by kmalloc() failure
From: Satoru Takeuchi <hidden>
Date: 2016-02-17 06:06:04
On 2016/02/16 2:53, David Sterba wrote:
On Mon, Feb 15, 2016 at 02:38:09PM +0900, Satoru Takeuchi wrote:quoted
There are some BUG_ON()'s after kmalloc() as follows. ===== foo = kmalloc(); BUG_ON(!foo); /* -ENOMEM case */ ===== A Docker + memory cgroup user hit these BUG_ON()s. https://bugzilla.kernel.org/show_bug.cgi?id=112101 Since it's very hard to handle these ENOMEMs properly, preventing these kmalloc() failures to avoid these BUG_ON()s for now, are a bit better than the current implementation anyway.Beware that the NOFAIL semantics is can cause deadlocks if it's on the critical writeback path or and can be reentered from itself through the reclaim. Unless you're sure that this is not the case, please do not add them just because it would seemingly fix the allocation failures.
About the all cases I changed, kmalloc()s can block since gfp_flags_allow_blocking() are true. Then no locks are acquired here and deadlocks don't happen. Am I missing something?
In the docker example, the memory is limited by cgroups so the NOFAIL mode can exhaust all reserves and just loop endlessly waiting for the OOM killer to get some memory or just waiting without any chance to progress.
I consider triggering OOM killer and killing processes in a cgroup are better than killing whole system. About the possibility of endless loop, there are many such problems in the whole kernel. Of course it can be said to Btrfs. ========================================== $ grep -rnH __GFP_NOFAIL fs/btrfs/ fs/btrfs/extent-tree.c:5970: GFP_NOFS | __GFP_NOFAIL); fs/btrfs/extent-tree.c:6043: bytenr + num_bytes - 1, GFP_NOFS | __GFP_NOFAIL); fs/btrfs/extent_io.c:4643: eb = kmem_cache_zalloc(extent_buffer_cache, GFP_NOFS|__GFP_NOFAIL); fs/btrfs/extent_io.c:4909: p = find_or_create_page(mapping, index, GFP_NOFS|__GFP_NOFAIL); ========================================== I understand fixing these problems cooperate with memory cgroup guys is the best in the long run. However, I consider bypassing this problem for now is better than the current implementation. Thanks, Satoru
-- To unsubscribe from this list: send the line "unsubscribe linux-btrfs" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html