Junio,
I just ran git clone against the mainline git repository using both http
and rsync. http was still quite slow compared to rsync. I expected that
the http time would be much faster than in the past due to the pack
file.
Is there something simple I'm missing?
--
Darrin
From: Junio C Hamano <hidden> Date: 2016-06-15 22:42:02
Darrin Thompson [off-list ref] writes:
I just ran git clone against the mainline git repository using both http
and rsync. http was still quite slow compared to rsync. I expected that
the http time would be much faster than in the past due to the pack
file.
Is there something simple I'm missing?
No, the only thing you missed was that I did not write it to
make it fast, but just to make it work ;-). The commit walker
simply does not work against a dumb http server repository that
is packed and prune-packed, which is already the case for both
kernel and git repositories.
The thing is, the base pack for the git repository is 1.8MB
currently containing 4500+ objects, while we accumulated 600+
unpacked objects since then which is about ~5MB. The commit
walker needs to fetched the latter one by one in the old way.
When packed incrementally on top of the base pack, these 600+
unpacked objects compress down to something like 400KB, and I
was hoping we could wait until we accumulate enough to produce
an incremental about a meg or so ...
On Thu, 2005-07-28 at 19:24 -0700, Junio C Hamano wrote:
The thing is, the base pack for the git repository is 1.8MB
currently containing 4500+ objects, while we accumulated 600+
unpacked objects since then which is about ~5MB. The commit
walker needs to fetched the latter one by one in the old way.
When packed incrementally on top of the base pack, these 600+
unpacked objects compress down to something like 400KB, and I
was hoping we could wait until we accumulate enough to produce
an incremental about a meg or so ...
Ok... so lets check my assumptions:
1. Pack files should reduce the number of http round trips.
2. What I'm seeing when I check out mainline git is the acquisition of a
single large pack, then 600+ more recent objects. Better than before,
but still hundreds of round trips.
3. If I wanted to further speed up the initial checkout on my own
repositories I could frequently repack my most recent few hundred
objects.
4. If curl had pipelining then less pack management would be needed.
Where is the code for gitweb? (i.e. http://kernel.org/git ) Seems like
it could benefit from some git-send-pack superpowers.
--
Darrin
From: Ryan Anderson <hidden> Date: 2016-06-15 22:42:03
On Fri, Jul 29, 2005 at 09:03:41AM -0500, Darrin Thompson wrote:
Where is the code for gitweb? (i.e. http://kernel.org/git ) Seems like
it could benefit from some git-send-pack superpowers.
http://www.kernel.org/pub/software/scm/gitweb/
It occurs to me that pulling this into the main git repository might not
be a bad idea, since it is currently living outside any revision
tracking at the moment.
--
Ryan Anderson
sometimes Pug Majere
On Fri, 2005-07-29 at 10:48 -0400, Ryan Anderson wrote:
On Fri, Jul 29, 2005 at 09:03:41AM -0500, Darrin Thompson wrote:
quoted
Where is the code for gitweb? (i.e. http://kernel.org/git ) Seems like
it could benefit from some git-send-pack superpowers.
http://www.kernel.org/pub/software/scm/gitweb/
It occurs to me that pulling this into the main git repository might not
be a bad idea, since it is currently living outside any revision
tracking at the moment.
From: Junio C Hamano <hidden> Date: 2016-06-15 22:42:03
Darrin Thompson [off-list ref] writes:
Ok... so lets check my assumptions:
1. Pack files should reduce the number of http round trips.
2. What I'm seeing when I check out mainline git is the acquisition of a
single large pack, then 600+ more recent objects. Better than before,
but still hundreds of round trips.
3. If I wanted to further speed up the initial checkout on my own
repositories I could frequently repack my most recent few hundred
objects.
4. If curl had pipelining then less pack management would be needed.
All true. Another possibility is to make multiple requests in
parallel; if curl does not do pipelining, either switch to
something that does, or have more then one process using curl.
The dumb server preparation creates three files, two of which is
currently used by clone (one is list of packs, the other is list
of branches and tags). The third one is commit ancestry
information. The commit walker could be taught to read it to
figure out what commits it still needs to fetch without waiting
for the commit being retrieved to be parsed.
Sorry, I am not planning to write that part myself.
One potential low hanging fruit is that even for cloning via
git:// URL we _might_ be better off starting with the dumb
server protocol; get the list of statically prepared packs and
obtain them upfront before starting the clone-pack/upload-pack
protocol pair.