Thread (4 messages) 4 messages, 2 authors, 2025-09-26

Re: [REGRESSION] workqueue/writeback: Severe CPU hang due to kworker proliferation during I/O flush and cgroup cleanup

From: Chenglong Tang <hidden>
Date: 2025-09-26 20:08:10
Also in: regressions, stable

cc'ing GKE folks

On Fri, Sep 26, 2025 at 12:59 PM Tejun Heo [off-list ref] wrote:
cc'ing Jan.

On Fri, Sep 26, 2025 at 12:54:29PM -0700, Chenglong Tang wrote:
quoted
Just did more testing here. Confirmed that the system hang's still
there but less frequently(6/40) with the patches
http://lkml.kernel.org/r/20250912103522.2935-1-jack@suse.cz appied to
v6.17-rc7. In the bad instances, the kworker count climbed to over
600+ and caused the hang over 80+ seconds.

So I think the patches didn't fully solve the issue.
I wonder how the number of workers still exploded to 600+. Are there that
many cgroups being shut down? Does clamping down @max_active resolve the
problem? There's no reason to have really high concurrency for this.

Thanks.

--
tejun
  
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help