Thread (9 messages) flat view 9 messages, 4 authors, 9d ago

Re: [PATCH 0/2] crypto: zstd - avoid initializing the workspace twice

From: Nick Terrell <hidden>
Date: 2026-09-10 22:31:15
Also in: lkml

On Tue, Sep 8, 2026 at 11:39 AM Usama Arif [off-list ref] wrote:
On Tue, 25 Aug 2026 15:06:00 -0700 Usama Arif [off-list ref] wrote:
quoted
Both zstd_compress() and zstd_decompress() set up the shared per-CPU
workspace as a C/DStream before walking the request, and then, when the
first source and destination fragments each span the whole request, hand
off to zstd_compress_one()/zstd_decompress_one(), which immediately
overwrite that same ctx->wksp with a CCtx/DCtx. The stream setup is
discarded without a byte having been processed.

That one-shot path is not a corner case: zswap always takes it when
storing, and takes it for a load whenever the stored object lies within a
single zsmalloc page.

These two patches defer the stream initialization to the first walk
iteration that actually streams, guarded by a flag because that iteration
can be reached more than once.

A 4 KiB crypto_acomp benchmark [1], twelve runs of nine 30,000-operation
rounds. Bare metal is an Intel Xeon Platinum 8321HC, turbo off,
performance governor, pinned to one core; the VM is a one-vCPU KVM guest
on a faster host.

                  baseline    patched     delta
  bare metal
    compress      52,283 ns   51,038 ns   1,245 ns   2.4%
    decompress     2,317 ns    1,998 ns     319 ns  13.8%
  one-vCPU KVM
    compress      16,675 ns   15,050 ns   1,625 ns   9.8%
    decompress     3,516 ns    2,265 ns   1,251 ns  35.6%

The guest numbers are larger because the two CPUID instructions in
ZSTD_cpuid() become unconditional VM exits there.

[1] https://gist.github.com/uarif1/5cf02f0e22c23f0d1b3d84348f12914c
Hi,

Just wanted to check if there was any feedback or review of the series.

I think its a nice optimization and even with the CPUID instructions
getting cached [1], this is still needed. The improvement in baremetal
is not coming (just) from CPUID instructions.
Agreed, this optimization makes sense to me as well.
[1] https://lore.kernel.org/all/20260901110850.1805747-1-usama.arif@linux.dev/ (local)

Thanks,
Usama
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help