Dmitry Potapov [off-list ref] writes:
file size = 1Kb; Hashing 100000 files
Before:
0.63user 0.86system 0:01.49elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k
0inputs+0outputs (0major+100428minor)pagefaults 0swaps
After:
0.54user 0.53system 0:01.07elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k
0inputs+0outputs (0major+421minor)pagefaults 0swaps
As you can see, in all tests the read() version performed better than
mmap() though the difference decreases with increase of the file size.
While for 1Kb files, the speed up is 39% (based on the elapsed time),
it is mere 1% for 1Mb file size.
Sounds good. Summarizing your numbers,
1Kb 39.25%
2Kb 27.27%
4Kb 17.06%
8Kb 11.21%
16Kb 7.00%
32Kb 4.81%
64Kb 3.31%
128Kb 2.29%
256Kb 3.51%
512Kb 2.92%
1024Kb 1.14%
32*1024 sounds like a better cut-off to me. After that, doubling the size
does not get comparable gain, and numbers get unstable (notice the glitch
around 256kB).
On Sun, Feb 21, 2010 at 11:32:10AM -0800, Junio C Hamano wrote:
32*1024 sounds like a better cut-off to me. After that, doubling the size
does not get comparable gain, and numbers get unstable (notice the glitch
around 256kB).
The reduction of speed-up after 32Kb is most likely due to L1 cache
size, which is 32Kb data per core on Core 2, and L2 cache is shared
among cores and is considerably slow. I have run my test a few more
times, and here are results:
1 - 39.25%
2 - 30.00%
4 - 17.79%
8 - 11.76%
16 - 7.58%
32 - 5.38%
64 - 3.89%
128 - 2.87%
256 - 2.31%
512 - 2.92%
1024 - 1.14%
and here is one more re-run starting with 32Kb:
32 - 5.38%
64 - 3.89%
128 - 2.29%
256 - 2.91%
512 - 2.92%
1024 - 1.14%
If you look at speed-up numbers, you can think that the numbers are
unstable, but in fact, the best time in 5 runs does not differ more
than 0.01s between those trials. But because difference for >=128Kb
is 0.05s or less, the accuracy of the above numbers is less than 25%.
But overall the outcome is clear -- read() is always a winner.
It would be interesting to see what difference Nehalem, which has a
smaller but much faster L2 cache than Core 2. It may perform better
at larger sizes up to 256Kb.
Anyway, based on above data, I believe that the proper cut-off should be
at least 64Kb, because additional 32Kb (from 32Kb to 64Kb) is about of
2.5% of total memory that git consumes anyway, and it gives you speed-up
around 3.5%...
Dmitry