Thread (39 messages) flat view 39 messages, 8 authors, 1d ago

Re: [PATCH v5 0/8] Generic IRQ entry/exit support for powerpc

From: Jirka Hladky <hidden>
Date: 2026-09-10 19:59:21
Also in: lkml

Oh, I forgot to add the actual results from the LPAR:

Model name:                  POWER10 (architected), altivec supported
Model:                     2.0 (pvr 0080 0200)
Type                                         : Dedicated
Mode                                         : Capped
Entitled Capacity                            : 1.00
Memory Mode                                  : Dedicated

Results (median bogo-ops/sec, N=15 per cell):
  Mode          7.1        7.2-rc4      Delta
  Enforcing     68,275     70,154       +2.7%
  Permissive    68,082     67,166       -1.3%
  Disabled      99,654     100,510      +0.9%

On Thu, Sep 10, 2026 at 9:54 PM Jirka Hladky [off-list ref] wrote:
Hi Mukesh,

Following up on the ppc64le SELinux/kill() regression - it doesn't reproduce
with proper multi-iteration sampling. stress-ng --kill has ~30%
run-to-run variance
on our LPARs. Part of the variance comes from LPARs being much less
isolated than I'd assumed - we saw contention bleed across LPARs on the same
physical server, even between "Dedicated" partitions. With 15 iterations per
config, 7.1 and 7.2-rc4 are statistically even (within ~3%) across enforcing/
permissive/disabled, confirmed on a second, physically separate LPAR (7.1 vs
7.2-rc6). This matches what you said - you couldn't reproduce it either.

I also built 7.2-rc4 from source with your proposed
rcu_ct_kernel_exit_qs() patch and re-ran the same sweep - it shows no
consistent effect vs. stock 7.2-rc4 (differences within the same noise
band), which is consistent with there being no real regression here for it
to address.

Sorry for the false alarm, and thank you for the patch and for taking the
time to look into this - I don't think it's needed for this issue, though
it may still be a reasonable efficiency improvement on its own merits (it
does avoid the unconditional rcu_preempt_deferred_qs() call on the common
fast path). I'll update the internal RHEL tracking ticket (RHEL-216740) to
reflect that this doesn't reproduce.

Thank you!
Jirka


On Tue, Sep 8, 2026 at 6:01 PM Jirka Hladky [off-list ref] wrote:
quoted
Hi Mukesh,

Re: policy details -  I'm running an ELN kernel on RHEL-10.2, with SELinux
in enforcing mode using the default (targeted) policy, nothing custom
or hardened on top of it. Let me know if you need anything more specific
(policy version, booleans, etc.) and I'll pull it.

Re: the rcu_ct_kernel_exit_qs() diff — thanks, that guard against the
unconditional rcu_preempt_deferred_qs() call lines up well with what I
was seeing on the RCU-barrier side. I'll apply it on top of 7.2 on
POWER10 box and re-run the stress-ng --kill numbers
(enforcing/permissive/disabled) to see how much of the SELinux tax it
recovers. No need to prepare anything separately.

I'm OOO tomorrow, but I'll get to this by the end of the week and report
back with results.

Thanks
Jirka


On Tue, Sep 8, 2026 at 3:26 PM Mukesh Kumar Chaurasiya
[off-list ref] wrote:
quoted
On Thu, Jul 30, 2026 at 01:24:46PM +0200, Jirka Hladky wrote:
quoted
On Thu, Jul 30, 2026 at 11:36 AM Mukesh Kumar Chaurasiya
[off-list ref] wrote:
quoted
On Mon, Jul 27, 2026 at 11:41:07AM +0200, Jirka Hladky wrote:
quoted
Hi Mukesh, Shrikanth, Madhavan,

I'm seeing a ~19% throughput regression on ppc64le (POWER10) in
syscall-heavy workloads between kernel 7.1 and 7.2-rc4, which I've
narrowed down to the GENERIC_ENTRY enablement for powerpc (commit
bee25f97ad24 "powerpc: Enable GENERIC_ENTRY feature").

The regression manifests specifically when SELinux is enabled. Regression
disappears with selinux=0. With SELinux disabled, 7.2 is actually
slightly *faster* than 7.1.

Test: stress-ng --kill 1 -t 23 (single-thread kill() syscall loop)
Machine: IBM,9080-HEX POWER10, 8 vCPUs (1 core SMT-8), 64 GiB

Results (bogo-ops/sec, higher is better):

Kernel   Enforcing   Permissive   Disabled    SELinux tax
------   ---------   ----------   --------    -----------
7.1      69,107      70,248       93,566      -26.1%
7.2-rc4  56,247      59,075       99,240      -43.3%
Delta    -18.6%      -15.9%       +6.1%

Key observations:
- SELinux permissive ~ enforcing on both kernels (overhead is in the
  code path, not policy evaluation)
- SELinux disabled: 7.2 is 6% faster than 7.1 -- the generic entry
  path itself is not slower for the base syscall
- The SELinux "tax" nearly doubles from 26% to 43% between 7.1->7.2
- No changes to security/selinux/avc.c between v7.1 and v7.2-rc4

The kill() hot path (sys_kill -> check_kill_permission -> security_task_kill
-> selinux_task_kill -> avc_has_perm -> avc_lookup) wraps every AVC lookup
in rcu_read_lock()/rcu_read_unlock(). On ppc64le with PREEMPT_RCU, each
pair requires lwsync/isync barriers. I suspect the generic entry path
changes something in how context tracking, tracing, or RCU interacts
with the syscall fast path that amplifies the per-call RCU cost.

Note: this is on top of a separate ~36% regression from 6.12->7.1 caused
by HAVE_PREEMPT_DYNAMIC_KEY enabling CONFIG_PREEMPT_RCU on ppc64le
(commit 6ad7751537e8), which I reported separately on the v4 thread.

Thank you
Jirka
[...]
Hey Jirka,

Thanks for reporting this.
I'll take a look.

Thanks,
Mukesh
Thank you, Mukesh! I'm happy to run additional tests on the POWER10 machine
if you have patches to try or specific experiments you'd like me to do.

Jirka
can you try this with the below diff?
diff --git a/include/linux/rcutiny.h b/include/linux/rcutiny.h
index e56ded733b1b..dcad641eb2c2 100644
--- a/include/linux/rcutiny.h
+++ b/include/linux/rcutiny.h
@@ -120,6 +120,7 @@ static inline bool rcu_preempt_need_deferred_qs(struct task_struct *t)
        return false;
 }
 static inline void rcu_preempt_deferred_qs(struct task_struct *t) { }
+static inline void rcu_ct_kernel_exit_qs(void) { }
 void rcu_scheduler_starting(void);
 static inline void rcu_end_inkernel_boot(void) { }
 static inline bool rcu_inkernel_boot_has_ended(void) { return true; }
diff --git a/include/linux/rcutree.h b/include/linux/rcutree.h
index 16a04202888b..d623f2a7d3fc 100644
--- a/include/linux/rcutree.h
+++ b/include/linux/rcutree.h
@@ -87,6 +87,7 @@ static inline void rcu_irq_exit_check_preempt(void) { }

 struct task_struct;
 void rcu_preempt_deferred_qs(struct task_struct *t);
+void rcu_ct_kernel_exit_qs(void);

 void exit_rcu(void);
diff --git a/kernel/context_tracking.c b/kernel/context_tracking.c
index a743e7ffa6c0..011018214c6d 100644
--- a/kernel/context_tracking.c
+++ b/kernel/context_tracking.c
@@ -118,7 +118,7 @@ static void noinstr ct_kernel_exit(bool user, int offset)
        lockdep_assert_irqs_disabled();
        trace_rcu_watching(TPS("End"), ct_nesting(), 0, ct_rcu_watching());
        WARN_ON_ONCE(IS_ENABLED(CONFIG_RCU_EQS_DEBUG) && !user && !is_idle_task(current));
-       rcu_preempt_deferred_qs(current);
+       rcu_ct_kernel_exit_qs();

        // instrumentation for the noinstr ct_kernel_exit_state()
        instrument_atomic_write(&ct->state, sizeof(ct->state));
diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
index 96848fc1f02b..c23478f70c17 100644
--- a/kernel/rcu/tree.c
+++ b/kernel/rcu/tree.c
@@ -368,6 +368,25 @@ notrace void rcu_momentary_eqs(void)
 }
 EXPORT_SYMBOL_GPL(rcu_momentary_eqs);

+/**
+ * rcu_ct_kernel_exit_qs - report deferred QS on syscall/exception exit if needed
+ *
+ * Called from ct_kernel_exit() on every return to userspace.  Guards the
+ * rcu_preempt_deferred_qs() call with rcu_preempt_need_deferred_qs() so that
+ * on the common fast path -- where nothing is deferred -- we avoid the
+ * cache-line traffic on current->rcu_read_unlock_special that the unconditional
+ * call causes.  This is particularly significant on weakly-ordered architectures
+ * (e.g. ppc64le) where rcu_read_lock/unlock issue lwsync/isync barriers and
+ * already touch that cache line in the syscall body.
+ *
+ * Follows the same pattern used by rcu_flavor_sched_clock_irq().
+ */
+notrace void rcu_ct_kernel_exit_qs(void)
+{
+       if (rcu_preempt_need_deferred_qs(current))
+               rcu_preempt_deferred_qs(current);
+}
+
 /**
  * rcu_is_cpu_rrupt_from_idle - see if 'interrupted' from idle
  *

--
-Jirka


--
-Jirka


-- 
-Jirka

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help