From: Anton Blanchard <hidden> Date: 2016-10-03 06:41:05
From: Anton Blanchard <redacted>
During context switch, switch_mm() sets our current CPU in mm_cpumask.
We can avoid this atomic sequence in most cases by checking before
setting the bit.
Testing on a POWER8 using our context switch microbenchmark:
tools/testing/selftests/powerpc/benchmarks/context_switch \
--process --no-fp --no-altivec --no-vector
Performance improves 2%.
Signed-off-by: Anton Blanchard <redacted>
---
arch/powerpc/include/asm/mmu_context.h | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
@@ -72,7 +72,8 @@ static inline void switch_mm(struct mm_struct *prev, struct mm_struct *next,structtask_struct*tsk){/* Mark this context has been used on the new CPU */-cpumask_set_cpu(smp_processor_id(),mm_cpumask(next));+if(!cpumask_test_cpu(smp_processor_id(),mm_cpumask(next)))+cpumask_set_cpu(smp_processor_id(),mm_cpumask(next));/* 32-bit keeps track of the current PGDIR in the thread struct */#ifdef CONFIG_PPC32
From: Anton Blanchard <redacted>
During context switch, switch_mm() sets our current CPU in mm_cpumask.
We can avoid this atomic sequence in most cases by checking before
setting the bit.
Testing on a POWER8 using our context switch microbenchmark:
tools/testing/selftests/powerpc/benchmarks/context_switch \
--process --no-fp --no-altivec --no-vector
Performance improves 2%.
Signed-off-by: Anton Blanchard <redacted>
---
arch/powerpc/include/asm/mmu_context.h | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
@@ -72,7 +72,8 @@ static inline void switch_mm(struct mm_struct *prev, struct mm_struct *next,structtask_struct*tsk){/* Mark this context has been used on the new CPU */-cpumask_set_cpu(smp_processor_id(),mm_cpumask(next));+if(!cpumask_test_cpu(smp_processor_id(),mm_cpumask(next)))+cpumask_set_cpu(smp_processor_id(),mm_cpumask(next));
I think this makes sense, in fact I think in the longer term we can
even use __set_bit() reorder-able version since we have a sync
coming out of schedule(). The read side for TLB flush can use a RMB
Acked-by: Balbir Singh <bsingharora@gmail.com>
Balbir Singh.
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2016-10-03 23:58:28
On Tue, 2016-10-04 at 10:25 +1100, Balbir Singh wrote:
I think this makes sense, in fact I think in the longer term we can
even use __set_bit() reorder-able version since we have a sync
coming out of schedule(). The read side for TLB flush can use a RMB
No, that wouldn't be atomic vs. other threads accessing the same
bitmap.
Ben.
On Tue, 2016-10-04 at 10:25 +1100, Balbir Singh wrote:
quoted
I think this makes sense, in fact I think in the longer term we can
even use __set_bit() reorder-able version since we have a sync
coming out of schedule(). The read side for TLB flush can use a RMB
No, that wouldn't be atomic vs. other threads accessing the same
bitmap.
Somethings distorted my thought process. Thanks!
Balbir Singh.
From: Michael Ellerman <hidden> Date: 2016-10-05 02:36:29
On Mon, 2016-03-10 at 06:40:29 UTC, Anton Blanchard wrote:
From: Anton Blanchard <redacted>
During context switch, switch_mm() sets our current CPU in mm_cpumask.
We can avoid this atomic sequence in most cases by checking before
setting the bit.
Testing on a POWER8 using our context switch microbenchmark:
tools/testing/selftests/powerpc/benchmarks/context_switch \
--process --no-fp --no-altivec --no-vector
Performance improves 2%.
Signed-off-by: Anton Blanchard <redacted>
Acked-by: Balbir Singh <bsingharora@gmail.com>