Thread (31 messages) flat view 31 messages, 4 authors, 2d ago

Re: [PATCH 01/12] mm/huge_memory: zap deposited page tables after an RCU grace period

From: Kiryl Shutsemau <kas@kernel.org>
Date: 2026-09-01 15:29:12
Also in: linux-alpha, linux-arch, linux-m68k, linux-mips, linux-mm, linux-riscv, linux-s390, linux-sh, linux-um, lkml, loongarch, sparclinux

On Tue, Sep 01, 2026 at 03:41:17PM +0100, Lorenzo Stoakes (ARM) wrote:
On Tue, Sep 01, 2026 at 11:24:08AM -0300, Jason Gunthorpe wrote:
quoted
On Tue, Sep 01, 2026 at 03:12:45PM +0100, Lorenzo Stoakes (ARM) wrote:
quoted
It won't be costly at the time of the calls obviously as its deferred. Maybe
increase some time spent in softirq but again is 512x that big of a deal?

I'm not sure how you'd both defer the free and somehow utilise mmu_gather here
either really, certainly not without it becoming extremely messy.
The less costly version is to thread the page to be freed onto the
mmu_gather through a linked list in the struct page memory. This is
super cheap since it is just a singly linked list operation.

Then when the mmu_gather is flushed it does a single call_rcu using
the rcu head of the struct page of the head of the list. The callback
clears the entire linked list of pages.

Since you have to tlb flush anyhow, it makes sense to always use the
mmu_gather. For example the design I ended up with for iommupt
accumulates all the invalidations and all the free-able memory into a
gather then invalidates and frees.

This allows maximizing the tlbi efficiency too. You can't do call_srcu
until you flush the tlb and if you call once per table then you are
also tlb flushing once per table too.

So if the kernel really does want to clear out 512 leaf tables the
optimal implementation is one range tlbi for 512 entries followed by
one call_rcu to free the memory. Hence the gather..
I think there's some confusion here.

This isn't the path in which a page table is being freed, the _deposited_
table is zapped, in zap_deposited_table().

That is, the page table kept in reserve for THP split, that is not
currently mapped.

It amounts to a __free_pages() call.

The TLB operations are in e.g. zap_huge_pmd() etc. and nobody has
complained about inefficiencies there.

So, unless I'm missing something here, TLB flushes play no role in this
whatsoever.

The issue Kiryl raised was that instead of immediately freeing page tables,
they are now batched up individually by call_rcu().

I personally find it difficult to imagine the numbers here would be
problematic or certainly cause anything observable beyond what is
observable now.

So I'm going to have to say, unless it can be clearly demonstrated this is
problematic, I don't think there's any reason to add additional complexity
here.
It would be nice to measure munmap() overhead here.

I am worried about hitting DEFAULT_MAX_RCU_BLIMIT and trigger
rcu_force_quiescent_state() which can be disruptive to the system.

DEFAULT_MAX_RCU_BLIMIT is 10K, so it is ~20G of THP unmapped on x86.

munmap() of 64G worth of THP should be enough to demonstrate the
problem.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help