Commit 582f95835a8f ("arm64: entry: convert el0_sync to C") converted lots
of functions from assembly to C, this greatly improves readability. But
el0_svc()/el0_svc_compat() is in response to system call requests from
user mode and may be in the hot path.
Although the SVC is in the first case of the switch statement in C, the
compiler optimizes the switch statement as a whole, and does not give SVC
a small boost.
Use "likely()" to help SVC directly invoke its handler after a simple
judgment to avoid entering the switch table lookup process.
After:
0000000000000ff0 <el0t_64_sync_handler>:
ff0: d503245f bti c
ff4: d503233f paciasp
ff8: a9bf7bfd stp x29, x30, [sp, #-16]!
ffc: 910003fd mov x29, sp
1000: d5385201 mrs x1, esr_el1
1004: 531a7c22 lsr w2, w1, #26
1008: f100545f cmp x2, #0x15
100c: 540000a1 b.ne 1020 <el0t_64_sync_handler+0x30>
1010: 97fffe14 bl 860 <el0_svc>
1014: a8c17bfd ldp x29, x30, [sp], #16
1018: d50323bf autiasp
101c: d65f03c0 ret
1020: f100705f cmp x2, #0x1c
Execute "./lat_syscall null" on my board (BogoMIPS : 200.00), it can save
about 10ns.
Before:
Simple syscall: 0.2365 microseconds
Simple syscall: 0.2354 microseconds
Simple syscall: 0.2339 microseconds
After:
Simple syscall: 0.2255 microseconds
Simple syscall: 0.2254 microseconds
Simple syscall: 0.2256 microseconds
Signed-off-by: Zhen Lei <redacted>
---
arch/arm64/kernel/entry-common.c | 18 ++++++++++++------
1 file changed, 12 insertions(+), 6 deletions(-)
From: Mark Rutland <mark.rutland@arm.com> Date: 2021-09-14 09:55:28
Hi,
On Fri, Sep 03, 2021 at 08:19:50PM +0800, Zhen Lei wrote:
Commit 582f95835a8f ("arm64: entry: convert el0_sync to C") converted lots
of functions from assembly to C, this greatly improves readability. But
el0_svc()/el0_svc_compat() is in response to system call requests from
user mode and may be in the hot path.
Although the SVC is in the first case of the switch statement in C, the
compiler optimizes the switch statement as a whole, and does not give SVC
a small boost.
Use "likely()" to help SVC directly invoke its handler after a simple
judgment to avoid entering the switch table lookup process.
After:
0000000000000ff0 <el0t_64_sync_handler>:
ff0: d503245f bti c
ff4: d503233f paciasp
ff8: a9bf7bfd stp x29, x30, [sp, #-16]!
ffc: 910003fd mov x29, sp
1000: d5385201 mrs x1, esr_el1
1004: 531a7c22 lsr w2, w1, #26
1008: f100545f cmp x2, #0x15
100c: 540000a1 b.ne 1020 <el0t_64_sync_handler+0x30>
1010: 97fffe14 bl 860 <el0_svc>
1014: a8c17bfd ldp x29, x30, [sp], #16
1018: d50323bf autiasp
101c: d65f03c0 ret
1020: f100705f cmp x2, #0x1c
It would be helpful if you could state which toolchain and config was
used to generate the above.
For comparison, what was the code generation like before? I assume
el0_svc wasn't the target of the first test and branch? Assuming so, how
many tests and branches were there before the call to el0_svc()?
At a high-level, I'm not too keen on special-casing things unless
necessary.
I wonder if we could get similar results without special-casing by using
a static const array of handlers indexed by the EC, since (with GCC
11.1.0 from the kernel.org crosstool page) that can result in code like:
0000000000001010 <el0t_64_sync_handler>:
1010: d503245f bti c
1014: d503233f paciasp
1018: a9bf7bfd stp x29, x30, [sp, #-16]!
101c: 910003fd mov x29, sp
1020: d5385201 mrs x1, esr_el1
1024: 90000002 adrp x2, 0 <el0t_64_sync_handlers>
1028: 531a7c23 lsr w3, w1, #26
102c: 91000042 add x2, x2, #:lo12:<el0t_64_sync_handlers>
1030: f8637842 ldr x2, [x2, x3, lsl #3]
1034: d63f0040 blr x2
1038: a8c17bfd ldp x29, x30, [sp], #16
103c: d50323bf autiasp
1040: d65f03c0 ret
... which might do better by virtue of reducing a chain of potential
mispredicts down to a single potential mispredict, and dynamic branch
prediction hopefully does a good job of predicting the common case at
runtime. That said, the resulting tables will be pretty big...
Execute "./lat_syscall null" on my board (BogoMIPS : 200.00), it can save
about 10ns.
Before:
Simple syscall: 0.2365 microseconds
Simple syscall: 0.2354 microseconds
Simple syscall: 0.2339 microseconds
After:
Simple syscall: 0.2255 microseconds
Simple syscall: 0.2254 microseconds
Simple syscall: 0.2256 microseconds
I appreciate this can be seen by a microbenchmark, but does this have an
impact on a real workload? I'd imagine that real syscall usage will
dominate this in practice, and this would fall into the noise.
Thanks,
Mark.
Hi,
On Fri, Sep 03, 2021 at 08:19:50PM +0800, Zhen Lei wrote:
quoted
Commit 582f95835a8f ("arm64: entry: convert el0_sync to C") converted lots
of functions from assembly to C, this greatly improves readability. But
el0_svc()/el0_svc_compat() is in response to system call requests from
user mode and may be in the hot path.
Although the SVC is in the first case of the switch statement in C, the
compiler optimizes the switch statement as a whole, and does not give SVC
a small boost.
Use "likely()" to help SVC directly invoke its handler after a simple
judgment to avoid entering the switch table lookup process.
After:
0000000000000ff0 <el0t_64_sync_handler>:
ff0: d503245f bti c
ff4: d503233f paciasp
ff8: a9bf7bfd stp x29, x30, [sp, #-16]!
ffc: 910003fd mov x29, sp
1000: d5385201 mrs x1, esr_el1
1004: 531a7c22 lsr w2, w1, #26
1008: f100545f cmp x2, #0x15
100c: 540000a1 b.ne 1020 <el0t_64_sync_handler+0x30>
1010: 97fffe14 bl 860 <el0_svc>
1014: a8c17bfd ldp x29, x30, [sp], #16
1018: d50323bf autiasp
101c: d65f03c0 ret
1020: f100705f cmp x2, #0x1c
It would be helpful if you could state which toolchain and config was
used to generate the above.
gcc version 7.3.0 (GCC), make defconfig
For comparison, what was the code generation like before? I assume
el0_svc wasn't the target of the first test and branch? Assuming so, how
many tests and branches were there before the call to el0_svc()?
At a high-level, I'm not too keen on special-casing things unless
necessary.
I wonder if we could get similar results without special-casing by using
a static const array of handlers indexed by the EC, since (with GCC
11.1.0 from the kernel.org crosstool page) that can result in code like:
0000000000001010 <el0t_64_sync_handler>:
1010: d503245f bti c
1014: d503233f paciasp
1018: a9bf7bfd stp x29, x30, [sp, #-16]!
101c: 910003fd mov x29, sp
1020: d5385201 mrs x1, esr_el1
1024: 90000002 adrp x2, 0 <el0t_64_sync_handlers>
1028: 531a7c23 lsr w3, w1, #26
102c: 91000042 add x2, x2, #:lo12:<el0t_64_sync_handlers>
1030: f8637842 ldr x2, [x2, x3, lsl #3]
1034: d63f0040 blr x2
1038: a8c17bfd ldp x29, x30, [sp], #16
103c: d50323bf autiasp
1040: d65f03c0 ret
... which might do better by virtue of reducing a chain of potential
mispredicts down to a single potential mispredict, and dynamic branch
prediction hopefully does a good job of predicting the common case at
runtime. That said, the resulting tables will be pretty big...
a48: 38624862 ldrb w2, [x3, w2, uxtw]
a4c: 10000063 adr x3, a58 <el0_sync_handler+0x48>
a50: 8b228862 add x2, x3, w2, sxtb #2
a54: d61f0040 br x2
The original implementation also generated a query table, but yours is
more concise. I will try to test it. Looks like a better solution.
quoted
Execute "./lat_syscall null" on my board (BogoMIPS : 200.00), it can save
about 10ns.
Before:
Simple syscall: 0.2365 microseconds
Simple syscall: 0.2354 microseconds
Simple syscall: 0.2339 microseconds
After:
Simple syscall: 0.2255 microseconds
Simple syscall: 0.2254 microseconds
Simple syscall: 0.2256 microseconds
I appreciate this can be seen by a microbenchmark, but does this have an
impact on a real workload? I'd imagine that real syscall usage will
dominate this in practice, and this would fall into the noise.
The product side has a test plan, but the progress will be slow.
Hi,
On Fri, Sep 03, 2021 at 08:19:50PM +0800, Zhen Lei wrote:
quoted
Commit 582f95835a8f ("arm64: entry: convert el0_sync to C") converted lots
of functions from assembly to C, this greatly improves readability. But
el0_svc()/el0_svc_compat() is in response to system call requests from
user mode and may be in the hot path.
Although the SVC is in the first case of the switch statement in C, the
compiler optimizes the switch statement as a whole, and does not give SVC
a small boost.
Use "likely()" to help SVC directly invoke its handler after a simple
judgment to avoid entering the switch table lookup process.
After:
0000000000000ff0 <el0t_64_sync_handler>:
ff0: d503245f bti c
ff4: d503233f paciasp
ff8: a9bf7bfd stp x29, x30, [sp, #-16]!
ffc: 910003fd mov x29, sp
1000: d5385201 mrs x1, esr_el1
1004: 531a7c22 lsr w2, w1, #26
1008: f100545f cmp x2, #0x15
100c: 540000a1 b.ne 1020 <el0t_64_sync_handler+0x30>
1010: 97fffe14 bl 860 <el0_svc>
1014: a8c17bfd ldp x29, x30, [sp], #16
1018: d50323bf autiasp
101c: d65f03c0 ret
1020: f100705f cmp x2, #0x1c
It would be helpful if you could state which toolchain and config was
used to generate the above.
gcc version 7.3.0 (GCC), make defconfig
quoted
For comparison, what was the code generation like before? I assume
el0_svc wasn't the target of the first test and branch? Assuming so, how
many tests and branches were there before the call to el0_svc()?
At a high-level, I'm not too keen on special-casing things unless
necessary.
I wonder if we could get similar results without special-casing by using
a static const array of handlers indexed by the EC, since (with GCC
11.1.0 from the kernel.org crosstool page) that can result in code like:
0000000000001010 <el0t_64_sync_handler>:
1010: d503245f bti c
1014: d503233f paciasp
1018: a9bf7bfd stp x29, x30, [sp, #-16]!
101c: 910003fd mov x29, sp
1020: d5385201 mrs x1, esr_el1
1024: 90000002 adrp x2, 0 <el0t_64_sync_handlers>
1028: 531a7c23 lsr w3, w1, #26
102c: 91000042 add x2, x2, #:lo12:<el0t_64_sync_handlers>
1030: f8637842 ldr x2, [x2, x3, lsl #3]
1034: d63f0040 blr x2
1038: a8c17bfd ldp x29, x30, [sp], #16
103c: d50323bf autiasp
1040: d65f03c0 ret
... which might do better by virtue of reducing a chain of potential
mispredicts down to a single potential mispredict, and dynamic branch
prediction hopefully does a good job of predicting the common case at
runtime. That said, the resulting tables will be pretty big...
a48: 38624862 ldrb w2, [x3, w2, uxtw]
a4c: 10000063 adr x3, a58 <el0_sync_handler+0x48>
a50: 8b228862 add x2, x3, w2, sxtb #2
a54: d61f0040 br x2
The original implementation also generated a query table, but yours is
more concise. I will try to test it. Looks like a better solution.
quoted
quoted
Execute "./lat_syscall null" on my board (BogoMIPS : 200.00), it can save
about 10ns.
Before:
Simple syscall: 0.2365 microseconds
Simple syscall: 0.2354 microseconds
Simple syscall: 0.2339 microseconds
After:
Simple syscall: 0.2255 microseconds
Simple syscall: 0.2254 microseconds
Simple syscall: 0.2256 microseconds
I appreciate this can be seen by a microbenchmark, but does this have an
impact on a real workload? I'd imagine that real syscall usage will
dominate this in practice, and this would fall into the noise.
The product side has a test plan, but the progress will be slow.
Hi,
On Tue, Sep 14, 2021 at 10:55:16AM +0100, Mark Rutland wrote:
Hi,
At a high-level, I'm not too keen on special-casing things unless
necessary.
I wonder if we could get similar results without special-casing by using
a static const array of handlers indexed by the EC, since (with GCC
11.1.0 from the kernel.org crosstool page) that can result in code like:
0000000000001010 <el0t_64_sync_handler>:
1010: d503245f bti c
1014: d503233f paciasp
1018: a9bf7bfd stp x29, x30, [sp, #-16]!
101c: 910003fd mov x29, sp
1020: d5385201 mrs x1, esr_el1
1024: 90000002 adrp x2, 0 <el0t_64_sync_handlers>
1028: 531a7c23 lsr w3, w1, #26
102c: 91000042 add x2, x2, #:lo12:<el0t_64_sync_handlers>
1030: f8637842 ldr x2, [x2, x3, lsl #3]
1034: d63f0040 blr x2
1038: a8c17bfd ldp x29, x30, [sp], #16
103c: d50323bf autiasp
1040: d65f03c0 ret
... which might do better by virtue of reducing a chain of potential
mispredicts down to a single potential mispredict, and dynamic branch
prediction hopefully does a good job of predicting the common case at
runtime. That said, the resulting tables will be pretty big...
I tested Mark's branch which implements this (found at
https://git.kernel.org/pub/scm/linux/kernel/git/mark/linux.git/log/?h=arm64/entry/switch-table)
I also took lmbench from https://github.com/intel/lmbench.git and built
`lat_syscall` with:
gcc lat_syscall.c lib_*.c -l m -o lat_syscall -static
These are the results I got from benchmarking on my MacBook Air M1, with
the following command:
./lat_syscall null &> /dev/null ; uname -a ; for i in 0 1 2 3 4 ; do ./lat_syscall null ; done
The kernel was based on arm64_defconfig that was then stripped of as much as possible.
GCC 11.1.0 from kernel.org crosstool page.
Clang build fom git b041b613e6fff713fc9ad6dbc73024286fb2fc93.
gcc:
master: 0.14300
switch-table: 0.14350
likely: 0.13962
clang:
master: 0.14354
switch-table: 0.14642
likely: 0.14256
The generated code looks similar to what Leizhen has posted, so I didn't
post it again.
So it seems the table approach actually performs worse in my testing,
and Leizhen's approach is slightly better than master (d0ee23f9d78be5531c4b055ea424ed0b489dfe9b).
Thanks,
Joey
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel