Re: [PATCH RESEND] xor: add missing vzeroupper to AVX code
From: David Laight <hidden>
Date: 2026-09-02 16:07:08
Also in:
lkml, stable
On Wed, 2 Sep 2026 15:37:06 +0200 Christoph Hellwig [off-list ref] wrote:
On Mon, Aug 31, 2026 at 02:22:48PM -0700, Eric Biggers wrote:quoted
Since the AVX optimized XOR code uses YMM registers, execute vzeroupper before returning from it. This is needed to avoid degrading the performance of any later SSE code that may happen to be executed. Fixes: ea4d26ae24e5 ("raid5: add AVX optimized RAID5 checksumming") Cc: stable@vger.kernel.org Signed-off-by: Eric Biggers <ebiggers@kernel.org> --- This didn't get taken through the x86 tree. Andrew, it seems you're taking patches to lib/raid/. Can you apply this one?Can we do kernel_avx_{begin,end} instead of having to open code and document this everywhere, please?
In which case I think you want the vzeroupper in kernel_avx_begin(). Actually, for some cpu at least, you need vzeroupper in kernel_fpu_begin() even if the code only uses the SSE registers. See: https://stackoverflow.com/questions/41303780/why-is-this-sse-code-6-times-slower-without-vzeroupper-on-skylake Basically, on Skylake, all SSE hit a penalty if any ymm high bits might be non-zero. Don't know what has changed since... I don't have the Intel optimisation manual downloaded (ok, I might have it but have NFI where), and the link on that page is broken. David