Re: [PATCH v2 13/20] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops
From: Mark Rutland <mark.rutland@arm.com>
Date: 2026-08-06 12:03:02
Also in:
stable
On Thu, Aug 06, 2026 at 01:32:52PM +0200, David Hildenbrand (Arm) wrote:
On 8/6/26 13:21, Mark Rutland wrote:quoted
On Wed, Aug 05, 2026 at 08:47:08AM +0200, David Hildenbrand (Arm) wrote:quoted
On 8/5/26 08:45, David Hildenbrand (Arm) wrote:quoted
FWIW, in a recent discussion on some prototype hacking [1] we saw some overhead in micro-benchmarks that would really hammer on a path that would now do a preempt_disable()+preempt_enable(). Switching from preempt_disable() to preempt_enable_no_resched() made it turn to noise. Of course, that has other undesirable impacts, and I am not sure if we are in the territory of code layout changes affecting the numbers. Just mentioning it as some data point.[1] https://lore.kernel.org/linux-mm/20260630174852-mutt-send-email-mst@kernel.org/ (local)Thanks for the pointer. IIUC in those cases you're using preempt_disable() .. preempt_enable() directly, not this_cpu_*(), right?It was purely preempt_disable/preempt_enable experiments without any percpu stuff.quoted
If so, patches 5 and 6 of this series [2,3] might have an impact, but I wouldn't expect a significant change unless you're calling preempt_enable a lot. Please beware that it's not safe to use preempt_enable_no_resched() UNLESS it is immediately followed by a call to schedule(). That's not documented today (and I couldn't find a good reference), so more folk are likely to be tempted to use it...Yes, that's also why we abandoned that (including for various other reasons :) ).
:)
preempt_enable_no_resched() helped to identify that the preempt_enable() was really causing the noticeable overhead, not the other minor stuff we added on some hot paths.
Understood! If we seeeing particularly noticeable overhead from preempt_enable() in some workloads, there are some options we could investigate to reduce that impact (e.g. using __preserve_most or a trampoline like x86's preempt_schedule_thunk to reduce necessary spills and register pressure). Please let me know if you see anything that stands out, as any examples would be useful for investigation. We'd want to figure out how much of the overhead comes from register pressure, and how much of it comes from the conditional call itself. If you're testing with PREEMPT_DYNAMIC=y, on some architectures (including arm64) you might see overhead reduced by: https://lore.kernel.org/lkml/20260803191731.3244294-1-mark.rutland@arm.com/ (local) ... but IIUC on x86 that won't change the cost of preempt_enable[_notrace](). Today that makes a static call to preempt_schedule[_notrace]_thunk, and a plain call will be the same cost. Mark.