From: Mark Brown <broonie@kernel.org> Date: 2021-05-11 16:09:29
This series is a combination of factoring out some duplicated code and a
very minor optimisation to the performance of handling converting FPSIMD
state to SVE in the live registers for 128 bit SVE vectors.
v2:
- Combine P and FFR flushing into a single macro.
Mark Brown (3):
arm64/sve: Split _sve_flush macro into separate Z and predicate
flushes
arm64/sve: Use the sve_flush macros in sve_load_from_fpsimd_state()
arm64/sve: Skip flushing Z registers with 128 bit vectors
arch/arm64/include/asm/fpsimd.h | 2 +-
arch/arm64/include/asm/fpsimdmacros.h | 4 +++-
arch/arm64/kernel/entry-fpsimd.S | 19 ++++++++++++-------
arch/arm64/kernel/fpsimd.c | 6 ++++--
4 files changed, 20 insertions(+), 11 deletions(-)
base-commit: 6efb943b8616ec53a5e444193dccf1af9ad627b5
--
2.20.1
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel
From: Mark Brown <broonie@kernel.org> Date: 2021-05-11 16:09:28
Trivial refactoring to support further work, no change to generated code.
Signed-off-by: Mark Brown <broonie@kernel.org>
---
arch/arm64/include/asm/fpsimdmacros.h | 4 +++-
arch/arm64/kernel/entry-fpsimd.S | 3 ++-
2 files changed, 5 insertions(+), 2 deletions(-)
From: Mark Brown <broonie@kernel.org> Date: 2021-05-11 16:10:15
This makes the code a bit clearer and as a result we can also make the
indentation more normal, there is no change to the generated code.
Signed-off-by: Mark Brown <broonie@kernel.org>
---
arch/arm64/kernel/entry-fpsimd.S | 9 ++++-----
1 file changed, 4 insertions(+), 5 deletions(-)
From: Mark Brown <broonie@kernel.org> Date: 2021-05-11 16:10:54
When the SVE vector length is 128 bits then there are no bits in the Z
registers which are not shared with the V registers so we can skip them
when zeroing state not shared with FPSIMD, this results in a minor
performance improvement.
Signed-off-by: Mark Brown <broonie@kernel.org>
---
arch/arm64/include/asm/fpsimd.h | 2 +-
arch/arm64/kernel/entry-fpsimd.S | 9 +++++++--
arch/arm64/kernel/fpsimd.c | 6 ++++--
3 files changed, 12 insertions(+), 5 deletions(-)
From: Dave Martin <Dave.Martin@arm.com> Date: 2021-05-12 13:43:26
On Tue, May 11, 2021 at 05:04:45PM +0100, Mark Brown wrote:
This makes the code a bit clearer and as a result we can also make the
indentation more normal, there is no change to the generated code.
Signed-off-by: Mark Brown <broonie@kernel.org>
From: Dave Martin <Dave.Martin@arm.com> Date: 2021-05-12 13:51:42
On Tue, May 11, 2021 at 05:04:46PM +0100, Mark Brown wrote:
quoted hunk
When the SVE vector length is 128 bits then there are no bits in the Z
registers which are not shared with the V registers so we can skip them
when zeroing state not shared with FPSIMD, this results in a minor
performance improvement.
Signed-off-by: Mark Brown <broonie@kernel.org>
---
arch/arm64/include/asm/fpsimd.h | 2 +-
arch/arm64/kernel/entry-fpsimd.S | 9 +++++++--
arch/arm64/kernel/fpsimd.c | 6 ++++--
3 files changed, 12 insertions(+), 5 deletions(-)
This does require that ZCR_EL1.LEN has already been set to match x0, and
is not changed again before entering userspace.
It would be a good idea to at least describe this in a comment so that
this doesn't get forgotten later on, but there's a limit to how
foolproof this low-level backend code needs to be...
quoted hunk
+ */
SYM_FUNC_START(sve_flush_live)
+ cbz x0, 1f // A VQ-1 of 0 is 128 bits so no extra Z state
sve_flush_z
- sve_flush_p_ffr
+1: sve_flush_p_ffr
ret
SYM_FUNC_END(sve_flush_live)
With a comment added as outlined above,
Reviewed-by: Dave Martin <Dave.Martin@arm.com>
Cheers
---Dave
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel
From: Mark Brown <broonie@kernel.org> Date: 2021-05-12 14:24:57
On Wed, May 12, 2021 at 02:49:09PM +0100, Dave Martin wrote:
On Tue, May 11, 2021 at 05:04:46PM +0100, Mark Brown wrote:
quoted
+/*
+ * Zero all SVE registers but the first 128-bits of each vector
+ *
+ * x0 = VQ - 1
This does require that ZCR_EL1.LEN has already been set to match x0, and
is not changed again before entering userspace.
It would be a good idea to at least describe this in a comment so that
this doesn't get forgotten later on, but there's a limit to how
foolproof this low-level backend code needs to be...
Similar concerns exist for huge chunks of the existing SVE code (eg, the
no further changes on vector length constraint is pretty much universal
and is I'd say largely more of a "make sure you handle the register
contents" thing on anything that changes the vector length).
On Tue, May 11, 2021 at 05:04:43PM +0100, Mark Brown wrote:
This series is a combination of factoring out some duplicated code and a
very minor optimisation to the performance of handling converting FPSIMD
state to SVE in the live registers for 128 bit SVE vectors.
v2:
- Combine P and FFR flushing into a single macro.
Mark Brown (3):
arm64/sve: Split _sve_flush macro into separate Z and predicate
flushes
arm64/sve: Use the sve_flush macros in sve_load_from_fpsimd_state()
arm64/sve: Skip flushing Z registers with 128 bit vectors
The series makes sense to me:
Acked-by: Catalin Marinas <catalin.marinas@arm.com>
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel