Thread (25 messages) 25 messages, 11 authors, 2018-08-16

Re: [PATCH] Performance Improvement in CRC16 Calculations.

flat view

From: Nicolas Pitre <hidden>
Date: 2018-08-10 20:02:28
Also in: linux-crypto, linux-scsi, lkml

On Fri, 10 Aug 2018, Joe Perches wrote:
On Fri, 2018-08-10 at 14:12 -0500, Jeff Lien wrote:
quoted
This patch provides a performance improvement for the CRC16 calculations done in read/write
workloads using the T10 Type 1/2/3 guard field.  For example, today with sequential write
workloads (one thread/CPU of IO) we consume 100% of the CPU because of the CRC16 computation
bottleneck.  Today's block devices are considerably faster, but the CRC16 calculation prevents
folks from utilizing the throughput of such devices.  To speed up this calculation and expose
the block device throughput, we slice the old single byte for loop into a 16 byte for loop,
with a larger CRC table to match.  The result has shown 5x performance improvements on various
big endian and little endian systems running the 4.18.0 kernel version.
Thanks.

This seems a sensible tradeoff for the 4k text size increase.
More like 7.5KB.  Would be best if this was configurable so the small 
version remained available.


Nicolas
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help