Thread (10 messages) flat view 10 messages, 6 authors, 2016-06-15

Re: dumb transports not being welcomed..

From: Johannes Schindelin <hidden>
Date: 2016-06-15 22:42:06

Possibly related (same subject, not in this thread)

Hi,

On Tue, 13 Sep 2005, Linus Torvalds wrote:
On Wed, 14 Sep 2005, Johannes Schindelin wrote:
quoted
IMHO the culprit is git-rev-list, which takes ages and ages for big 
repositories (beware: this could be my Darwin client which might be 
incapable to stop the rev enumeration in time; but if that can be done 
unintentionally, this can be intentionally, too!).
Packed too?
Yes. Almost all of it.
git-rev-list will take a long time if the tree is unpacked and not in the 
cache. It's all disk seeks. That's _especially_ true of a full clone 
(which will walk the whole way down).
That could be the case, but my test case is a CVS project I track on one 
side, and I fetch on the other side. Therefore, your diagram from your 
other mail does not really apply. My history looks more or less like this:

a b c d origin
          |
          |
          |
          |
\ \ \ \   |
 ---------|

So, the origin is a linear CVS project. With many, many, many commits. At 
one stage I broke off new git branches. I kept tracking the CVS project, 
though.

What I see when fetching all heads (thanks to Junio, this is one call to 
git-fetch now), where all but origin are up to date, is that it takes a 
very long time. Swapping kicks in, and top tells me that 26.6% of the 
memory is occupied by git-rev-list (The server has 128M, with 1G swap, and 
I am unfortunately not the only user of this machine).

I fail to see why it should need those amounts of memory. (I tested this 
over the ssh protocol, which should essentially do the same as git-daemon, 
right?) After all, the merge point between the branches should be marked 
uninteresting after one single step from each of my private branches.
But I have tons of memory in my machines, and I haven't looked at how 
badly it does if you don't have that. I know that master.kernel.org is 
certainly not having any trouble at all with me pulling from lots of 
trees.. Maybe git-rev-list uses up lots of your memory.
That certainly is the case.

As for master.kernel.org: Unfortunately, you will not be the only puller. 
And if your process needs just 5% of the RAM, then 21 pullers will be too 
many.
That said, I do think that --objects handling is _very_ CPU-hungry.
In my experience, before the swapping started, the process did not get 
more than 20% CPU.

Nevertheless, I still think that it would be a good idea to reuse the 
files created for the dumb transport for the intelligent transport. 
Especially for a project which is more often fetched than uploaded.

I also see other strange things like packing 0 objects, and packing >0 
objects after just having fetched from that repository. Hopefully I will 
have time to look into that (and understand the code to begin with).

Ciao,
Dscho
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help