git-fetching from a big repository is slow

20 messages, 8 authors, 2016-08-11 · open the first message on its own page

git-fetching from a big repository is slow

From: Andy Parkins <hidden>
Date: 2016-08-11 20:06:24

Hello,

I've got a big repository.  I've got two computers.  One has the repository 
up-to-date (164M after repack); one is behind (30M ish).

I used git-fetch to try and update; and the sync took HOURS.  I zipped 
the .git directory and transferred that and it took about 15 minutes to 
transfer.

Am I doing something wrong?  The git-fetch was done with a git+ssh:// URL.  
The zip transfer with scp (so ssh shouldn't be a factor).



Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE

Re: git-fetching from a big repository is slow

From: Andreas Ericsson <hidden>
Date: 2016-08-11 19:16:46

Johannes Schindelin wrote:
Hi,

On Thu, 14 Dec 2006, Andreas Ericsson wrote:
quoted
Andy Parkins wrote:
quoted
Hello,

I've got a big repository.  I've got two computers.  One has the repository
up-to-date (164M after repack); one is behind (30M ish).

I used git-fetch to try and update; and the sync took HOURS.  I zipped the
.git directory and transferred that and it took about 15 minutes to
transfer.

Am I doing something wrong?  The git-fetch was done with a git+ssh:// URL.
The zip transfer with scp (so ssh shouldn't be a factor).
This seems to happen if your repository consists of many large binary files,
especially many large binary files of several versions that do not deltify
well against each other. Perhaps it's worth adding gzip compression detecion
to git? I imagine more people than me are tracking gzipped/bzip2'ed content
that pretty much never deltifies well against anything else.
Or we add something like the heuristics we discovered in another thread, 
where rename detection (which is related to delta candidate searching) is 
not started if the sizes differ drastically.
It wouldn't work for this particular case though. In our distribution 
repository we have ~300 bzip2 compressed tarballs with an average size 
of 3MiB. 240 of those are between 2.5 and 4 MiB, so they don't 
drastically differ, but neither do they delta well.

One option would be to add some sort of config option to skip attempting 
deltas of files with a certain suffix. That way we could just tell it to 
ignore *.gz,*.tgz,*.bz2 and everything would work just as it does today, 
but a lot faster.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se

Re: git-fetching from a big repository is slow

From: Geert Bosch <hidden>
Date: 2016-08-11 19:22:12

On Dec 14, 2006, at 10:06, Andreas Ericsson wrote:
It wouldn't work for this particular case though. In our  
distribution repository we have ~300 bzip2 compressed tarballs with  
an average size of 3MiB. 240 of those are between 2.5 and 4 MiB, so  
they don't drastically differ, but neither do they delta well.

One option would be to add some sort of config option to skip  
attempting deltas of files with a certain suffix. That way we could  
just tell it to ignore *.gz,*.tgz,*.bz2 and everything would work  
just as it does today, but a lot faster.
Such special magic based on filenames is always a bad idea. Tomorrow  
somebody
comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
compressed content. In the end git will be doing lots of magic and  
still perform
badly on unknown compressed content.

There is a very simple way of detecting compressed files: just look  
at the
size of the compressed blob and compare against the size of the  
expanded blob.
If the compressed blob has a non-trivial size which is close to the  
expanded
size, assume the file is not interesting as source or target for deltas.

Example:
    if (compressed_size > expanded_size / 4 * 3 + 1024) {
      /* don't try to deltify if blob doesn't compress well */
      return ...;
    }

Re: git-fetching from a big repository is slow

From: Andy Parkins <hidden>
Date: 2016-08-11 19:25:29

On Thursday 2006 December 14 13:53, Andreas Ericsson wrote:
This seems to happen if your repository consists of many large binary
files, especially many large binary files of several versions that do
not deltify well against each other. Perhaps it's worth adding gzip
It's actually just every released patch to the linux kernel ever issued.  
Almost entirely ASCII and every revision (save the first) created by patching 
the previous.


Andy

-- 
Dr Andy Parkins, M Eng (hons), MIEE

Re: git-fetching from a big repository is slow

From: Junio C Hamano <hidden>
Date: 2016-08-11 19:30:43

Johannes Schindelin [off-list ref] writes:
git-show-ref traverses every single _local_ tag when called. This is to 
overcome the problem that tags can be packed now, so a simple file 
existence check is not sufficient.
Is "traverses every single _local_ tag" a fact?  It might go
through every single _local_ (possibly stale) packed tag in
memory but it should not traverse $GIT_DIR/refs/tags.

If I recall correctly, show-ref (1) first checks the filesystem
"$GIT_DIR/$named_ref" and says Ok if found and valid; otherwise
(2) checks packed refs (reads $GIT_DIR/packed-refs if not
already).  So that would be at most one open (which may fail in
(1)) and one open+read (in (2)).  Unless we are talking about
fork+exec overhead, that "traverse" should be reasonably fast.

Where is the bottleneck?

Re: git-fetching from a big repository is slow

From: Nicolas Pitre <hidden>
Date: 2016-08-11 19:40:34

On Fri, 15 Dec 2006, Johannes Schindelin wrote:
Hi,

On Thu, 14 Dec 2006, Shawn Pearce wrote:
quoted
Geert Bosch [off-list ref] wrote:
quoted
Such special magic based on filenames is always a bad idea. Tomorrow  
somebody
comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
compressed content. In the end git will be doing lots of magic and  
still perform
badly on unknown compressed content.

There is a very simple way of detecting compressed files: just look  
at the
size of the compressed blob and compare against the size of the  
expanded blob.
If the compressed blob has a non-trivial size which is close to the  
expanded
size, assume the file is not interesting as source or target for deltas.

Example:
   if (compressed_size > expanded_size / 4 * 3 + 1024) {
     /* don't try to deltify if blob doesn't compress well */
     return ...;
   }
And yet I get good delta compression on a number of ZIP formatted files 
which don't get good additional zlib compression (<3%). Doing the above 
would cause those packfiles to explode to about 10x their current size.
A pity. Geert's proposition sounded good to me.

However, there's got to be a way to cut short the search for a delta 
base/deltification when a certain (maybe even configurable) amount of time 
has been spent on it.
Yes! Run git-repack -a -d on the remote repository.

Re: git-fetching from a big repository is slow

From: Johannes Schindelin <hidden>
Date: 2016-08-11 19:40:40

Hi,

On Thu, 14 Dec 2006, Andy Parkins wrote:
On Thursday 2006 December 14 15:45, Han-Wen Nienhuys wrote:
quoted
I just noticed that git-fetch now runs git-show-ref --verify on every
tag it encounters. This seems to slow down fetch over here.
There aren't any tags in this repository :-)
git-show-ref traverses every single _local_ tag when called. This is to 
overcome the problem that tags can be packed now, so a simple file 
existence check is not sufficient.

It would be much faster, probably, if you pack the local refs. IIRC I once 
argued for automatically packing refs (and all refs), but this has not 
been picked up, and I do not really care about it either.

Ciao,
Dscho

Re: git-fetching from a big repository is slow

From: Geert Bosch <hidden>
Date: 2016-08-11 19:48:52

On Dec 14, 2006, at 14:46, Shawn Pearce wrote:
And yet I get good delta compression on a number of ZIP formatted
files which don't get good additional zlib compression (<3%).
Doing the above would cause those packfiles to explode to about
10x their current size.
Yes, that's because for zip files each file in the archive is
compressed independently. Similar things might happen when
checking in uncompressed tar files with JPG's. The question
is whether you prefer bad time usage or bad space usage when
handling large binary blobs. Maybe we should use a faster,
less precise algorithm instead of giving up.

Still, I think doing anything based on filename is a mistake.
If we want to have a heuristic to prevent spending too much time
on deltifying large compressed files, the heuristic should be
based on content, not filename.

Maybe we could some "magic" as used by the file(1) command
that allows git to say a bit more about the content of blobs.
This could be used both for ordering files during deltification
and to determine wether to try deltification at all.

   -Geert

Re: git-fetching from a big repository is slow

From: Johannes Schindelin <hidden>
Date: 2016-08-11 19:55:29

Hi,

On Thu, 14 Dec 2006, Andreas Ericsson wrote:
Andy Parkins wrote:
quoted
Hello,

I've got a big repository.  I've got two computers.  One has the repository
up-to-date (164M after repack); one is behind (30M ish).

I used git-fetch to try and update; and the sync took HOURS.  I zipped the
.git directory and transferred that and it took about 15 minutes to
transfer.

Am I doing something wrong?  The git-fetch was done with a git+ssh:// URL.
The zip transfer with scp (so ssh shouldn't be a factor).
This seems to happen if your repository consists of many large binary files,
especially many large binary files of several versions that do not deltify
well against each other. Perhaps it's worth adding gzip compression detecion
to git? I imagine more people than me are tracking gzipped/bzip2'ed content
that pretty much never deltifies well against anything else.
Or we add something like the heuristics we discovered in another thread, 
where rename detection (which is related to delta candidate searching) is 
not started if the sizes differ drastically.

Ciao,
Dscho

Re: git-fetching from a big repository is slow

From: Johannes Schindelin <hidden>
Date: 2016-08-11 20:00:09

Hi,

On Thu, 14 Dec 2006, Shawn Pearce wrote:
Geert Bosch [off-list ref] wrote:
quoted
Such special magic based on filenames is always a bad idea. Tomorrow  
somebody
comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
compressed content. In the end git will be doing lots of magic and  
still perform
badly on unknown compressed content.

There is a very simple way of detecting compressed files: just look  
at the
size of the compressed blob and compare against the size of the  
expanded blob.
If the compressed blob has a non-trivial size which is close to the  
expanded
size, assume the file is not interesting as source or target for deltas.

Example:
   if (compressed_size > expanded_size / 4 * 3 + 1024) {
     /* don't try to deltify if blob doesn't compress well */
     return ...;
   }
And yet I get good delta compression on a number of ZIP formatted files 
which don't get good additional zlib compression (<3%). Doing the above 
would cause those packfiles to explode to about 10x their current size.
A pity. Geert's proposition sounded good to me.

However, there's got to be a way to cut short the search for a delta 
base/deltification when a certain (maybe even configurable) amount of time 
has been spent on it.

Ciao,
Dscho

Re: git-fetching from a big repository is slow

From: Andreas Ericsson <hidden>
Date: 2016-08-11 20:00:27

Andy Parkins wrote:
Hello,

I've got a big repository.  I've got two computers.  One has the repository 
up-to-date (164M after repack); one is behind (30M ish).

I used git-fetch to try and update; and the sync took HOURS.  I zipped 
the .git directory and transferred that and it took about 15 minutes to 
transfer.

Am I doing something wrong?  The git-fetch was done with a git+ssh:// URL.  
The zip transfer with scp (so ssh shouldn't be a factor).
This seems to happen if your repository consists of many large binary 
files, especially many large binary files of several versions that do 
not deltify well against each other. Perhaps it's worth adding gzip 
compression detecion to git? I imagine more people than me are tracking 
gzipped/bzip2'ed content that pretty much never deltifies well against 
anything else.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se

Re: git-fetching from a big repository is slow

From: Shawn Pearce <hidden>
Date: 2016-08-11 20:07:17

Johannes Schindelin [off-list ref] wrote:
On Thu, 14 Dec 2006, Shawn Pearce wrote:
quoted
I'm OK with a small increase in packfile size as a result of slightly 
less optimal delta base selection on the really large binary files due 
to something like the above, but 10x is insane.
Not if it is a server having to do all the work. Along with all the work 
for all other clients. When you do a fetch, you really should be nice to 
the serving side.
Yes, that's true.

But I fail to see what that has to do with the part you quoted above.
A 1% increase in transfer bandwidth may be better for a server if
it halves the CPU usage or disk IO usage if the server has more
bandwidth than those available; likewise a 1% decrease in transfer
bandwidth may be better for a server if it has lots of CPU to spare
but very little network bandwidth available.

Since every server is different its not like we can tune for just
one of those cases and cross our fingers.

-- 

Re: git-fetching from a big repository is slow

From: Shawn Pearce <hidden>
Date: 2016-08-11 20:09:33

Johannes Schindelin [off-list ref] wrote:
On Thu, 14 Dec 2006, Shawn Pearce wrote:
quoted
Geert Bosch [off-list ref] wrote:
quoted
   if (compressed_size > expanded_size / 4 * 3 + 1024) {
     /* don't try to deltify if blob doesn't compress well */
     return ...;
   }
And yet I get good delta compression on a number of ZIP formatted files 
which don't get good additional zlib compression (<3%). Doing the above 
would cause those packfiles to explode to about 10x their current size.
A pity. Geert's proposition sounded good to me.

However, there's got to be a way to cut short the search for a delta 
base/deltification when a certain (maybe even configurable) amount of time 
has been spent on it.
I'm not sure time is the best rule there.

Maybe if the object is large (e.g. over 512 KiB or some configured
limit) and did not compress well when we last deflated it
(e.g. Geert's rule above) then only try to delta it against another
object whose hinted filename is very close/exactly matches and
whose size is very close, and don't make nearly as many attempts
on the matching hunks within any two files if the file appears to
be binary and not text.

I'm OK with a small increase in packfile size as a result of slightly
less optimal delta base selection on the really large binary files
due to something like the above, but 10x is insane.

-- 

Re: git-fetching from a big repository is slow

From: Andy Parkins <hidden>
Date: 2016-08-11 20:22:53

On Thursday 2006 December 14 15:45, Han-Wen Nienhuys wrote:
I just noticed that git-fetch now runs git-show-ref --verify on every
tag it encounters. This seems to slow down fetch over here.
There aren't any tags in this repository :-)

Andy

-- 
Dr Andy Parkins, M Eng (hons), MIEE

Re: git-fetching from a big repository is slow

From: Johannes Schindelin <hidden>
Date: 2016-08-11 20:24:40

Hi,

On Thu, 14 Dec 2006, Shawn Pearce wrote:
I'm OK with a small increase in packfile size as a result of slightly 
less optimal delta base selection on the really large binary files due 
to something like the above, but 10x is insane.
Not if it is a server having to do all the work. Along with all the work 
for all other clients. When you do a fetch, you really should be nice to 
the serving side.

Ciao,
Dscho

Re: git-fetching from a big repository is slow

From: Johannes Schindelin <hidden>
Date: 2016-08-11 20:27:54

Hi,

On Thu, 14 Dec 2006, Junio C Hamano wrote:
Johannes Schindelin [off-list ref] writes:
quoted
git-show-ref traverses every single _local_ tag when called. This is to 
overcome the problem that tags can be packed now, so a simple file 
existence check is not sufficient.
Is "traverses every single _local_ tag" a fact?  It might go
through every single _local_ (possibly stale) packed tag in
memory but it should not traverse $GIT_DIR/refs/tags.

If I recall correctly, show-ref (1) first checks the filesystem
"$GIT_DIR/$named_ref" and says Ok if found and valid; otherwise
(2) checks packed refs (reads $GIT_DIR/packed-refs if not
already).
If I read builtin-show-ref.c correctly, it _always_ calls 
for_each_ref(show_ref, NULL);

The only reason that the loop in for_each_ref can stop early is if 
show_ref returns something different than 0. But it does not! Every single 
return in show_ref() returns 0. It does not matter, though (see below).
So that would be at most one open (which may fail in (1)) and one 
open+read (in (2)).  Unless we are talking about fork+exec overhead, 
that "traverse" should be reasonably fast.

Where is the bottleneck?
The problem is that so many stat()s _do_ take time. Again, if I read the 
code correctly, it not only stat()s every loose ref, but also resolves the 
refs in get_ref_dir(), which is called from get_loose_refs(), which is 
unconditionally called in for_each_ref().

Even if the refs are packed, it takes quite _long_ (I confirmed this). And 
it is not at all necessary! Instead of a O(n^2) we can easily reduce this 
to O(n*log(n)), and we can reduce the n fork()&exec()s of git-show-ref by 
a single one.

Ciao,
Dscho

Re: git-fetching from a big repository is slow

From: Shawn Pearce <hidden>
Date: 2016-08-11 20:28:34

Geert Bosch [off-list ref] wrote:
Such special magic based on filenames is always a bad idea. Tomorrow  
somebody
comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
compressed content. In the end git will be doing lots of magic and  
still perform
badly on unknown compressed content.

There is a very simple way of detecting compressed files: just look  
at the
size of the compressed blob and compare against the size of the  
expanded blob.
If the compressed blob has a non-trivial size which is close to the  
expanded
size, assume the file is not interesting as source or target for deltas.

Example:
   if (compressed_size > expanded_size / 4 * 3 + 1024) {
     /* don't try to deltify if blob doesn't compress well */
     return ...;
   }
And yet I get good delta compression on a number of ZIP formatted
files which don't get good additional zlib compression (<3%).
Doing the above would cause those packfiles to explode to about
10x their current size.

-- 

Re: git-fetching from a big repository is slow

From: Andreas Ericsson <hidden>
Date: 2016-08-11 20:35:16

Geert Bosch wrote:
On Dec 14, 2006, at 10:06, Andreas Ericsson wrote:
quoted
It wouldn't work for this particular case though. In our distribution 
repository we have ~300 bzip2 compressed tarballs with an average size 
of 3MiB. 240 of those are between 2.5 and 4 MiB, so they don't 
drastically differ, but neither do they delta well.

One option would be to add some sort of config option to skip 
attempting deltas of files with a certain suffix. That way we could 
just tell it to ignore *.gz,*.tgz,*.bz2 and everything would work just 
as it does today, but a lot faster.
Such special magic based on filenames is always a bad idea. Tomorrow 
somebody
comes with .zip files (oh, and of course .ZIP), then it's .jpg's other
compressed content. In the end git will be doing lots of magic and still 
perform
badly on unknown compressed content.
Hence config option. People can tell git to skip trying to delta 
whatever they want. For this particular mothership repo, we only ever 
work against it when we're at the office, meaning resulting datasize is 
not an issue, but data computation can be a real bottle-neck.
There is a very simple way of detecting compressed files: just look at the
size of the compressed blob and compare against the size of the expanded 
blob.
If the compressed blob has a non-trivial size which is close to the 
expanded
size, assume the file is not interesting as source or target for deltas.

Example:
   if (compressed_size > expanded_size / 4 * 3 + 1024) {
     /* don't try to deltify if blob doesn't compress well */
     return ...;
   }
Many compression algorithms generate similar output for similar input. 
Most source-code projects change relatively little between releases, so 
they *could* delta well, it's just that in our repo they don't.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se

Re: git-fetching from a big repository is slow

From: Nicolas Pitre <hidden>
Date: 2016-08-11 20:38:55

On Thu, 14 Dec 2006, Andreas Ericsson wrote:
Andy Parkins wrote:
quoted
Hello,

I've got a big repository.  I've got two computers.  One has the repository
up-to-date (164M after repack); one is behind (30M ish).

I used git-fetch to try and update; and the sync took HOURS.  I zipped the
.git directory and transferred that and it took about 15 minutes to
transfer.

Am I doing something wrong?  The git-fetch was done with a git+ssh:// URL.
The zip transfer with scp (so ssh shouldn't be a factor).
This seems to happen if your repository consists of many large binary files,
especially many large binary files of several versions that do not deltify
well against each other. Perhaps it's worth adding gzip compression detecion
to git? I imagine more people than me are tracking gzipped/bzip2'ed content
that pretty much never deltifies well against anything else.
If your remote repository is fully packed in a single pack that should 
not have any impact on the transfer latency since no attempt to 
redeltify objects against each other is attempted by default when those 
objects are in the same pack.

Re: git-fetching from a big repository is slow

From: Han-Wen Nienhuys <hidden>
Date: 2016-08-11 20:41:31

Andy Parkins escreveu:
On Thursday 2006 December 14 13:53, Andreas Ericsson wrote:
quoted
This seems to happen if your repository consists of many large binary
files, especially many large binary files of several versions that do
not deltify well against each other. Perhaps it's worth adding gzip
It's actually just every released patch to the linux kernel ever issued.  
Almost entirely ASCII and every revision (save the first) created by patching 
the previous.
I just noticed that git-fetch now runs git-show-ref --verify on every
tag it encounters. This seems to slow down fetch over here.

-- 
 Han-Wen Nienhuys - hanwen@xs4all.nl - http://www.xs4all.nl/~hanwen
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help