Thread (20 messages) flat view 20 messages, 8 authors, 2016-08-11

Re: git-fetching from a big repository is slow

From: Andreas Ericsson <hidden>
Date: 2016-08-11 19:16:46

Johannes Schindelin wrote:
Hi,

On Thu, 14 Dec 2006, Andreas Ericsson wrote:
quoted
Andy Parkins wrote:
quoted
Hello,

I've got a big repository.  I've got two computers.  One has the repository
up-to-date (164M after repack); one is behind (30M ish).

I used git-fetch to try and update; and the sync took HOURS.  I zipped the
.git directory and transferred that and it took about 15 minutes to
transfer.

Am I doing something wrong?  The git-fetch was done with a git+ssh:// URL.
The zip transfer with scp (so ssh shouldn't be a factor).
This seems to happen if your repository consists of many large binary files,
especially many large binary files of several versions that do not deltify
well against each other. Perhaps it's worth adding gzip compression detecion
to git? I imagine more people than me are tracking gzipped/bzip2'ed content
that pretty much never deltifies well against anything else.
Or we add something like the heuristics we discovered in another thread, 
where rename detection (which is related to delta candidate searching) is 
not started if the sizes differ drastically.
It wouldn't work for this particular case though. In our distribution 
repository we have ~300 bzip2 compressed tarballs with an average size 
of 3MiB. 240 of those are between 2.5 and 4 MiB, so they don't 
drastically differ, but neither do they delta well.

One option would be to add some sort of config option to skip attempting 
deltas of files with a certain suffix. That way we could just tell it to 
ignore *.gz,*.tgz,*.bz2 and everything would work just as it does today, 
but a lot faster.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help