Thread (9 messages) flat view 9 messages, 4 authors, 3d ago

Re: [PATCH RESEND] xor: add missing vzeroupper to AVX code

From: David Laight <hidden>
Date: 2026-09-02 16:07:08
Also in: lkml, stable

On Wed, 2 Sep 2026 15:37:06 +0200
Christoph Hellwig [off-list ref] wrote:
On Mon, Aug 31, 2026 at 02:22:48PM -0700, Eric Biggers wrote:
quoted
Since the AVX optimized XOR code uses YMM registers, execute vzeroupper
before returning from it.  This is needed to avoid degrading the
performance of any later SSE code that may happen to be executed.

Fixes: ea4d26ae24e5 ("raid5: add AVX optimized RAID5 checksumming")
Cc: stable@vger.kernel.org
Signed-off-by: Eric Biggers <ebiggers@kernel.org>
---

This didn't get taken through the x86 tree.  Andrew, it seems you're
taking patches to lib/raid/.  Can you apply this one?  
Can we do kernel_avx_{begin,end} instead of having to open code
and document this everywhere, please?
In which case I think you want the vzeroupper in kernel_avx_begin().

Actually, for some cpu at least, you need vzeroupper in kernel_fpu_begin()
even if the code only uses the SSE registers.

See: https://stackoverflow.com/questions/41303780/why-is-this-sse-code-6-times-slower-without-vzeroupper-on-skylake
Basically, on Skylake, all SSE hit a penalty if any ymm high bits might be non-zero.
Don't know what has changed since...

I don't have the Intel optimisation manual downloaded (ok, I might have
it but have NFI where), and the link on that page is broken.

David
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help