[PATCH 00/11] Convert barrier pairs to acquire/release for better performance
From: Jinjie Ruan <hidden>
Date: 2026-08-25 09:53:50
Also in:
linux-can, linux-ext4, linux-fsdevel, lkml
Hi, This series converts some existing smp_wmb()/smp_rmb() barrier pairs to smp_store_release()/smp_load_acquire() across various subsystems. Background ========== Many architectures support load acquire and store release instructions which can replace explicit memory barriers and save cycles. As noted in the ARM architecture reference [1]: "Weaker ordering requirements that are imposed by Load-Acquire and Store-Release instructions allow for micro-architectural optimizations, which could reduce some of the performance impacts that are otherwise imposed by an explicit memory barrier. If the ordering requirement is satisfied using either a Load-Acquire or Store-Release, then it would be preferable to use these instructions instead of a DMB." On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB barriers. Replacing the read barrier with smp_load_acquire() reduces this to 8 cycles on an Ampere Altra. We also observed significant barrier overhead while profiling Unxibench syscall test on arm64: a single getuid() call is ~8ns slower than on a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(), which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows the use of LDAR, eliminating the measurable overhead. This motivated a broader search for existing barrier pairs that can be converted to the lighter acquire/release semantics. Changes ======= Each patch in this series targets a specific barrier pair where the publish/subscribe pattern is already present: - Writers populate data, then publish a flag/count/pointer via smp_store_release() - Readers load the flag/count/pointer via smp_load_acquire(), then consume the data This preserves the existing memory ordering guarantees while allowing architectures with native acquire/release instructions (e.g. arm64's STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD). On architectures without native support, the generated code is generally no worse than the explicit barrier pair. The conversions are mechanical and no functional change is intended. [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba Jinjie Ruan (11): user_namespace: Use acquire/release for nr_extents synchronization lib/vsprintf: Use acquire/release for ptr_key publication fs: aio: Use acquire/release for ring->tail publication fs: Use acquire/release for fdtable resize synchronization pidfs: Use test_bit_acquire() for attr flag tests super: Use acquire for SB_BORN check in super_cache_count() ext4: Convert group-count barrier protocol to acquire/release soreuseport: publish num_socks with acquire/release net: sched: act_gact: use acquire/release for tcfg_ptype 8021q: publish vlan_devices_arrays entries with acquire/release can: isotp: publish tx.state with smp_store_release() fs/aio.c | 12 +++++------- fs/ext4/ext4.h | 10 +++------- fs/ext4/mballoc.c | 6 ++---- fs/ext4/resize.c | 19 +++++++++++-------- fs/file.c | 10 ++++------ fs/pidfs.c | 6 ++---- fs/super.c | 7 +++---- kernel/user_namespace.c | 24 +++++++++++++----------- lib/vsprintf.c | 11 ++++------- net/8021q/vlan.c | 6 ++---- net/8021q/vlan.h | 8 +++----- net/can/isotp.c | 4 ++-- net/core/sock_reuseport.c | 20 ++++++++------------ net/sched/act_gact.c | 12 ++++-------- 14 files changed, 66 insertions(+), 89 deletions(-) -- 2.34.1