Re: [PATCH] mm/memcontrol: avoid stuck FLUSHING_CACHED_CHARGE bit on isolated cpus
From: Shakeel Butt <shakeel.butt@linux.dev>
Date: 2026-08-28 15:41:07
Also in:
linux-mm, lkml, stable
On Fri, Aug 28, 2026 at 04:25:41PM +0200, Michal Hocko wrote:
quoted hunk ↗ jump to hunk
On Fri 28-08-26 09:46:21, Rik van Riel wrote:quoted
drain_all_stock() can leave FLUSHING_CACHED_CHARGE set after the work is dropped. It sets the bit before checking isolation and schedule_drain_work() checks isolation and queues in a separate RCU critical section, so housekeeping_update()'s synchronize_rcu() can race the second check. drain_local_stock() only clears the bit for work that ran, so the bit remains set and the stock is never drained again. Reorganize the drain_all_stock() loop, reducing nesting, splitting out local vs remote cpu handling, and skipping everything on isolated cpus, which solves the stuck FLUSHING_CACHED_CHARGE flag.Is there any reason why we cannot simply clear the flag if the work is not scheduled? ---diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 69b37f63a307..907ac63e5067 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c@@ -2261,8 +2261,10 @@ static bool is_memcg_drain_needed(struct memcg_stock_pcp *stock, return flush; } -static void schedule_drain_work(int cpu, struct work_struct *work) +static bool schedule_drain_work(int cpu, struct work_struct *work) { + int ret = false; + /* * Protect housekeeping cpumask read and work enqueue together * in the same RCU critical section so that later cpuset isolated@@ -2270,8 +2272,12 @@ static void schedule_drain_work(int cpu, struct work_struct *work) * pending work on newly isolated CPUs. */ guard(rcu)(); - if (!cpu_is_isolated(cpu)) - queue_work_on(cpu, memcg_wq, work); + if (!cpu_is_isolated(cpu)) { + queue_work_on(cpu, memcg_wq, &memcg_st->work); + ret = true; + } + + return ret;
Let's go with this patch. We can simplify above by inversing the check: if (cpu_is_isolated(cpu)) return false; queue_work_on(cpu, memcg_wq, &memcg_st->work); return true;
quoted hunk ↗ jump to hunk
} /*@@ -2303,8 +2309,8 @@ void drain_all_stock(struct mem_cgroup *root_memcg) &memcg_st->flags)) { if (cpu == curcpu) drain_local_memcg_stock(&memcg_st->work); - else - schedule_drain_work(cpu, &memcg_st->work); + else if (!schedule_drain_work(cpu, &memcg_st->work)) + clear_bit(FLUSHING_CACHED_CHARGE, &memcg_st->flags) } if (!test_bit(FLUSHING_CACHED_CHARGE, &obj_st->flags) &&@@ -2313,8 +2319,8 @@ void drain_all_stock(struct mem_cgroup *root_memcg) &obj_st->flags)) { if (cpu == curcpu) drain_local_obj_stock(&obj_st->work); - else - schedule_drain_work(cpu, &obj_st->work); + else if (!schedule_drain_work(cpu, &obj_st->work)) + clear_bit(FLUSHING_CACHED_CHARGE, &obj_st->flags); } } migrate_enable();-- Michal Hocko SUSE Labs