Re: [PATCH v3 0/5] mm: Unconditional per-VMA locks and cleanups
From: Suren Baghdasaryan <surenb@google.com>
Date: 2026-08-03 17:51:41
Also in:
linux-mm, lkml
On Sun, Aug 2, 2026 at 7:11 PM Barry Song [off-list ref] wrote:
On Mon, Aug 3, 2026 at 5:58 AM Suren Baghdasaryan [off-list ref] wrote:quoted
v2 version of this patchset [1] was written by Dave Hansen and per his request, I'm taking over this series. tl;dr: Make per-VMA locks available in all configs. Simplify some of the per-VMA lock users now that they can rely on them being always available. Binder and networking folks: Your code is the target of the cleanups. I'm cc'ing you now on v2 because there's emerging consensus on the mm side that the approach here is sane. I'm not quite sure how this pile would get merged, but ack/review tags would be appreciated if this looks good to you. Longer version: When working on some x86 shadow stack code, it was a real pain to avoid causing recursive locking problems with mmap_lock. One way to avoid those was to avoid mmap_lock and use per-VMA locks instead. They are great, but they are not available in all configs which makes them unusable in generic code, or if you want to completely avoid mmap_lock. Make per-VMA locks available in all configs. Right now, they are only available on select architectures when SMP and MMU are enabled. But all of the primitives that per-VMA locks are built on (RCU, maple trees, refcounts) work just fine without SMP or MMU. The only real downside is that making VMAs a wee bit bigger on !MMU and !SMP builds. The upside is much cleaner code, lower complexity and less #ifdeffery. Clean up a binder VMA locking site now that it can rely on per-VMA locks. Building on top of universally-available per-VMA locks, introduce a new helper. Since the new API does not require callers to have a fallback to mmap_lock, it's much easier to use. Callers can potentially replace this very common kernel idiom: mmap_read_lock(mm); vma = vma_lookup() // fiddle with vma mmap_read_unlock(mm); with: vma = vma_start_read_unlocked(mm, address); // fiddle with vma vma_end_read(vma); Which avoids mmap_lock entirely in the fast path. Use that new API for another binder site and one in the TCP code.Nice, Suren and Dave. I wonder if we could use the same approach in the page fault path. Instead of falling back to mmap_lock when lock_vma_under_rcu() fails the first time, could we wait for the writer to finish and then retry acquiring the VMA lock?
Yeah, we might be able to do that. Matthew is working on moving common page-fault handling code into a single arch-independent place. Your suggested change would be simpler if done after Matthew's refactoring.
quoted hunk ↗ jump to hunk
For example:diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c index 85e23388f9bb..684f38cc4e74 100644 --- a/arch/arm64/mm/fault.c +++ b/arch/arm64/mm/fault.c@@ -677,7 +677,7 @@ static int __kprobes do_page_fault(unsigned longfar, unsigned long esr, if (!(mm_flags & FAULT_FLAG_USER)) goto lock_mmap; - vma = lock_vma_under_rcu(mm, addr); + vma = vma_start_read_unlocked(mm, addr); if (!vma) goto lock_mmap;diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c index 45b99c3b1442..a3a4c4741e30 100644 --- a/arch/x86/mm/fault.c +++ b/arch/x86/mm/fault.c@@ -1331,7 +1331,7 @@ void do_user_addr_fault(struct pt_regs *regs, if (!(flags & FAULT_FLAG_USER)) goto lock_mmap; - vma = lock_vma_under_rcu(mm, address); + vma = vma_start_read_unlocked(mm, address); if (!vma) goto lock_mmap;Best Regards Barry