On Wed, 14 Sep 2005, Johannes Schindelin wrote:
IMHO the culprit is git-rev-list, which takes ages and ages for big
repositories (beware: this could be my Darwin client which might be
incapable to stop the rev enumeration in time; but if that can be done
unintentionally, this can be intentionally, too!).
Packed too?
git-rev-list will take a long time if the tree is unpacked and not in the
cache. It's all disk seeks. That's _especially_ true of a full clone
(which will walk the whole way down).
But I have tons of memory in my machines, and I haven't looked at how
badly it does if you don't have that. I know that master.kernel.org is
certainly not having any trouble at all with me pulling from lots of
trees.. Maybe git-rev-list uses up lots of your memory.
I'm seeing 14 seconds of CPU-time for a _full_ kernel history, with
"--objects". Yes, it's not exactly cheap, and maybe I should optimize it
(it's all in the "--objects" handling and probably a large portion of it
is because trees actually pack very well indeed, so it's actually
unpacking a lot of trees), but considering that that is preparing the
metadata for pulling down a hundred megs of stuff..
That said, I do think that --objects handling is _very_ CPU-hungry. The
offender is this old commit of mine:
4311d328fee11fbd80862e3c5de06a26a0e80046
Author: Linus Torvalds [off-list ref]
Date: Sat Jul 23 10:01:49 2005 -0700
Be more aggressive about marking trees uninteresting
...
which is much better about avoiding objects in old trees, but it does so
at the expense of being _horribly_ CPU-inefficient. It will walk through
every tree of every commit that we decided was uninteresting.
You can try to just undo that one commit - it will make pack-files have a
few extraneous objects, but I think it will make a huge difference in the
CPU cost of "small pulls" (it won't matter at all for the "git clone"
case: for that case we just always have to walk the whole object tree).
Linus