Thread (8 messages) 8 messages, 2 authors, 2016-06-15

Re: Cloning speed comparison, round II

From: Linus Torvalds <torvalds@osdl.org>
Date: 2016-06-15 22:42:11


On Sat, 12 Nov 2005, Petr Baudis wrote:
Dear diary, on Sat, Nov 12, 2005 at 08:40:11PM CET, I got a letter
where Linus Torvalds [off-list ref] said that...
quoted

On Sat, 12 Nov 2005, Petr Baudis wrote:
quoted
             rsync   git+ssh(*)   git(**)   http

git.git      0m45s   0m34s        5m30s     4m01s (++)

cogito.git   2m09s   1m54s (+)    4m30s     15m11s (only single run)
Well, at the time of fetching, master.kernel.org with git+ssh had load
about ~3.5 and some wild gzip was eating most of the CPU there. So if
the git protocol still manages to be TEN times slower while rsync goes
full speed from that machine, I would say that this means the git server
requires way too much CPU.
Look again.

master.kernel.org was _faster_ than rsync using the native git protocol, 
despite being under a load of 3.5.

Look at the numbers: 45 secs for rsync, 34 secs for git protocol to 
master.

Now, I don't know which rsync machine you used (you can rsync both from 
master and from rync.kernel.org), since you don't say. 

Now, it's unquestionably true that rsync can be faster under many 
circumstances. Most notably when disk IO is really slow, since the native 
git protocol will do a lot more synchronous operations, since it actually 
tests what it is doing.

But I _guarantee_ you that rsync is at least ten times slower than the 
native git protocol in many circumstances. It can't handle repacking 
(which is critical for good server performance).

And in fact it can't handle totally unpacked directories and small updates 
well either (the reason I totally stopped doing rsync was because it took 
minutes to go through the whole list of unpacked objects for a small 
update, while the native protocol would just fetch the needed objects and 
be done with it.

So sometimes rsync is faster, sometimes the git protocol is faster. But 
the git protocol is _always_ better from a sanity standpoint.

The things you get with the native git protocol:

 - you don't have to trust the other end. If the other end lies about the 
   SHA1's of its objects, rsync will never know. It will just download the 
   thing, and you may have a corrupt database.

   With the git protocol, we just get the objects, and recompute their 
   names. The other end can't lie about what their SHA is.

 - the rsync protocol totally breaks down with multiple branches. It 
   fetches stuff it shouldn't because it doesn't know better.

 - the rsync protocol scales with project size, not with change size. This 
   works well for small projects, where the changes are usually not all 
   that hugely different from the total size of the project, but it really 
   sucks for big projects.

 - the rsync protocol fundamentally cannot handle two differently packed 
   trees well. That doesn't matter if you only track one tree, but it 
   matters _hugely_ for people (like me) who pull from tens of different 
   trees.

So the fact is: rsync is often slower, and _always_ less capable. 

			Linus
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help