Re: [PATCH v3 8/8] xor: Add AVX-512 optimized xor_gen()
From: sashiko-bot@kernel.org
Date: 2026-08-13 02:39:50
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: - [Medium] Missing `vzeroupper` instruction after using AVX-512 (ZMM) registers, leading to severe AVX-to-SSE transition penalties in kernel threads. --
--- Patch [8]: [PATCH v3 8/8] xor: Add AVX-512 optimized xor_gen() ---commit 0074b2adffd00d86b2a03e4d30495b21794ecb20 Author: Eric Biggers [off-list ref] xor: Add AVX-512 optimized xor_gen() This commit introduces an AVX-512 optimized implementation of xor_gen() using 512-bit ZMM registers and the vpternlogq instruction. It is enabled on x86_64 CPUs with AVX512F that do not have the PREFER_YMM feature set.
quoted hunk ↗ jump to hunk
diff --git a/lib/raid/xor/x86/xor-avx512.c b/lib/raid/xor/x86/xor-avx512.c new file mode 100644 index 0000000000000..17f57900d8274 --- /dev/null +++ b/lib/raid/xor/x86/xor-avx512.c
[ ... ]
+static void xor_gen_avx512(void *dest, void **srcs, unsigned int src_cnt,
+ unsigned int bytes)
+{
+ kernel_fpu_begin();
+ xor_gen_avx512_inner(dest, srcs, src_cnt, bytes);[Severity: Medium] Does this code need a vzeroupper instruction before releasing the FPU to avoid AVX-to-SSE transition penalties? If a kernel thread modifies the ZMM registers here and subsequently executes legacy SSE instructions (for example, in kcryptd) without an intervening user-space context switch to clear the state, the CPU could incur a severe transition penalty on affected architectures. While this problem already existed in the xor-avx.c implementation using YMM registers, the newly introduced xor-avx512.c appears to repeat the defect.
+ kernel_fpu_end(); +}
-- Sashiko AI review · https://sashiko.dev/#/patchset/20260813021506.55129-1-ebiggers@kernel.org?part=8