q: faster way to integrate/merge lots of topic branches?

22 messages, 10 authors, 2016-06-15 · open the first message on its own page

q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

I've got the following, possibly stupid question: is there a way to 
merge a healthy number of topic branches into the master branch in a 
quicker way, when most of the branches are already merged up?

Right now i've got something like this scripted up:

  for B in $(git-branch | cut -c3- ); do git-merge $B; done 

It takes a lot of time to run on even a 3.45GHz box:

  real    0m53.228s
  user    0m41.134s
  sys     0m11.405s

I just had a workflow incident where i forgot that this script was 
running in one window (53 seconds are a _long_ time to start doing some 
other stuff :-), i switched branches and the script merrily chugged away 
merging branches into a topic branch i did not intend.

It iterates over 140 branches - but all of them are already merged up.

Anyone can simulate it by switching to the linus/master branch of the 
current Linux kernel tree, and doing:

   time for ((i=0; i<140; i++)); do git-merge v2.6.26; done

   real    1m26.397s
   user    1m10.048s
   sys     0m13.944s

One could argue that determining whether it's all merged up already is a 
complex task, but but even this seemingly trivial merge of HEAD into 
HEAD is quite slow:

   time for ((i=0; i<140; i++)); do git-merge HEAD; done

   real    0m17.871s
   user    0m8.977s
   sys     0m8.396s

I'm wondering whether there are tricks to speed this up. The real script 
i'm using is much longer and obscured with boring details like errors, 
conflicts, etc. - but the above is the gist of it. (and that is what 
makes it slow primarily)

Using a speculative Octopus might be one approach, but that runs into 
the octopus merge limitation at 24 branches, and it also is quite slow 
as well. (and is not equivalent to the serial merge of 140 branches)

I have thought of using the last CommitDate of the topic branch and 
compare it with the last CommitDate of the master branch [and i can 
trust those values] - that would be a lot faster - but maybe i'm missing 
something trivial that makes that approach unworkable. It would also be 
nice to have a builtin shortcut for that instead of having to go via 
"git-log --pretty=fuller" to dump the CommitDate field.

builtin-integrate.c perhaps? ;-)

	Ingo

Re: q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

* Ingo Molnar [off-list ref] wrote:
I have thought of using the last CommitDate of the topic branch and 
compare it with the last CommitDate of the master branch [and i can 
trust those values] - that would be a lot faster - but maybe i'm 
missing something trivial that makes that approach unworkable. It 
would also be nice to have a builtin shortcut for that instead of 
having to go via "git-log --pretty=fuller" to dump the CommitDate 
field.
hm, this method would be fragile if done purely within my integration 
script, as the timestamp of the head would have to be updated 
atomically, while always merging all the topic branches in one such 
transaction. (so that the timestamps do not get out of sync and a topic 
branch is not skipped by accident)

So i guess it's better to just create a separate .git/refs/merge-cache/ 
hierarchy with timestamps of last merged branches and their head sha1 
... but maybe i'm banging on open doors?

	Ingo

Re: q: faster way to integrate/merge lots of topic branches?

From: Andreas Ericsson <hidden>
Date: 2016-06-15 22:45:00

Ingo Molnar wrote:
I've got the following, possibly stupid question: is there a way to 
merge a healthy number of topic branches into the master branch in a 
quicker way, when most of the branches are already merged up?

Right now i've got something like this scripted up:

  for B in $(git-branch | cut -c3- ); do git-merge $B; done 

It takes a lot of time to run on even a 3.45GHz box:

  real    0m53.228s
  user    0m41.134s
  sys     0m11.405s

I just had a workflow incident where i forgot that this script was 
running in one window (53 seconds are a _long_ time to start doing some 
other stuff :-), i switched branches and the script merrily chugged away 
merging branches into a topic branch i did not intend.

It iterates over 140 branches - but all of them are already merged up.
With the builtin merge (which is in next), this should be doable with
an octopus merge, which will eliminate the branches that are already
fully merged, resulting in a less-than-140-way merge (thank gods...).
It also doesn't have the 24-way cap that the scripted version suffers
from.

If it does a good job at your rather extreme use-case, I'd say it's
good enough for 'master' pretty soon :-)

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Tel: +46 8-230225                  Fax: +46 8-230231

Re: q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

* Ingo Molnar [off-list ref] wrote:
So i guess it's better to just create a separate 
.git/refs/merge-cache/ hierarchy with timestamps of last merged 
branches and their head sha1 ... but maybe i'm banging on open doors?
here's the git-fastmerge script i've whipped up in 10 minutes. It does 
the trick nicely for me:

first run:

  real    0m53.228s
  user    0m41.134s
  sys     0m11.405s

second run:

  real    0m2.751s
  user    0m1.280s
  sys     0m1.491s

or a 20x speedup. Yummie! :-)

It properly notices when i commit to a topic branch, and it maintains a 
proper matrix of <A> <- <B> merge timestamps. It even embedds the sha1's 
in the timestamp path so it should be quite complete. It should work 
fine across resets, re-merges, etc. too i think. It should work well 
with renamed branches as well i think. (although i dont do that all that 
often)

In fact even if i delete the whole .git/mergecache/ hierarchy and run a 
'cold' merge, it's much faster:

  real    0m32.129s
  user    0m24.456s
  sys     0m7.603s

Because many of the branches have the same sha1 so it's already 
half-optimized even on the first run.

Much of the remaining 2.7 seconds overhead comes from the git-log runs 
to retrieve the sha1s, so i guess it could all be made even faster.

Now this scheme assumes that there's a sane underlying filesystem that 
can take these long pathnames and which has good timestamps (which i 
have, so it's not a worry for me).

Hm?

	Ingo

-----------------{ git-fastmerge }--------------------->
#!/bin/bash

usage () {
  echo 'usage: git-fastmerge <refspec>..'
  exit -1
}

[ $# = 0 ] && usage

BRANCH=$1

MERGECACHE=.git/mergecache

[ ! -d $MERGECACHE ] && { mkdir $MERGECACHE || usage; }

HEAD_SHA1=$(git-log -1 --pretty=format:"%H")
BRANCH_SHA1=$(git-log -1 --pretty=format:"%H" $BRANCH)

CACHE=$MERGECACHE/$HEAD_SHA1/$BRANCH_SHA1

[ -f "$CACHE" -a "$CACHE" -nt .git/refs/heads/$BRANCH_SHA1 ] && {
  echo "merge-cache hit on HEAD <= $1"
  exit 0
}

git-merge $1 && {
  mkdir -p $(dirname $CACHE)
  touch $CACHE
}

Re: q: faster way to integrate/merge lots of topic branches?

From: SZEDER Gábor <hidden>
Date: 2016-06-15 22:45:00

Hi,

On Wed, Jul 23, 2008 at 03:05:18PM +0200, Ingo Molnar wrote:
I've got the following, possibly stupid question: is there a way to 
merge a healthy number of topic branches into the master branch in a 
quicker way, when most of the branches are already merged up?

Right now i've got something like this scripted up:

  for B in $(git-branch | cut -c3- ); do git-merge $B; done 
you cound use 'git branch --no-merged' to list only those branches
that have not been merged into your current HEAD.


Üdv,
Gábor

Re: q: faster way to integrate/merge lots of topic branches?

From: Sergey Vlasov <hidden>
Date: 2016-06-15 22:45:00

On Wed, 23 Jul 2008 15:05:18 +0200 Ingo Molnar wrote:
Anyone can simulate it by switching to the linus/master branch of the
current Linux kernel tree, and doing:

   time for ((i=0; i<140; i++)); do git-merge v2.6.26; done

   real    1m26.397s
   user    1m10.048s
   sys     0m13.944s
Timing results here (E6750 @ 2.66GHz):
41.61s user 3.71s system 99% cpu 45.530 total

However, testing whether there is something new to merge could be
performed significantly faster:

$ time sh -c 'for ((i=0; i<140; i++)); do [ -n "$(git rev-list --max-count=1 v2.6.26 ^HEAD)" ]; done'
sh -c   5.49s user 0.26s system 99% cpu 5.786 total

The same loop with "git merge-base v2.6.26 HEAD" takes about 40
seconds here - apparently finding the merge base is the expensive
part, and it makes sense to avoid it if you expect that most of your
branches do not contain anything new to merge.

Re: q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

* Andreas Ericsson [off-list ref] wrote:
Ingo Molnar wrote:
quoted
I've got the following, possibly stupid question: is there a way to  
merge a healthy number of topic branches into the master branch in a  
quicker way, when most of the branches are already merged up?

Right now i've got something like this scripted up:

  for B in $(git-branch | cut -c3- ); do git-merge $B; done 

It takes a lot of time to run on even a 3.45GHz box:

  real    0m53.228s
  user    0m41.134s
  sys     0m11.405s

I just had a workflow incident where i forgot that this script was  
running in one window (53 seconds are a _long_ time to start doing some 
other stuff :-), i switched branches and the script merrily chugged 
away merging branches into a topic branch i did not intend.

It iterates over 140 branches - but all of them are already merged up.
With the builtin merge (which is in next), this should be doable with 
an octopus merge, which will eliminate the branches that are already 
fully merged, resulting in a less-than-140-way merge (thank gods...). 
It also doesn't have the 24-way cap that the scripted version suffers 
from.

If it does a good job at your rather extreme use-case, I'd say it's 
good enough for 'master' pretty soon :-)
hm, while i do love octopus merges [*] for release and bisection-quality 
purposes, for throw-away (delta-)integration runs it's more manageable 
to do a predictable series of one-on-one merges.

It results in better git-rerere behavior, has easier (to the human) 
conflict resolutions and the octopus merge also falls apart quite easily 
when it runs into conflicts. Furthermore, i've often seen octopus merges 
fail while a series of 1:1 merges succeeded.

What i could try is to do a speculative octopus merge, in the hope of it 
just going fine - and then fall back to the serial merge if it fails?

The git-fastmerge approach is probably still faster though - and 
certainly simpler from a workflow POV.

	Ingo

[*] take a look at these in the Linux kernel -git repo:

      gitk 3c1ca43fafea41e38cb2d0c1684119af4c1de547
      gitk 6924d1ab8b7bbe5ab416713f5701b3316b2df85b

Re: q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

* SZEDER Gábor [off-list ref] wrote:
Hi,

On Wed, Jul 23, 2008 at 03:05:18PM +0200, Ingo Molnar wrote:
quoted
I've got the following, possibly stupid question: is there a way to 
merge a healthy number of topic branches into the master branch in a 
quicker way, when most of the branches are already merged up?

Right now i've got something like this scripted up:

  for B in $(git-branch | cut -c3- ); do git-merge $B; done 
you cound use 'git branch --no-merged' to list only those branches
that have not been merged into your current HEAD.
hm, it's very slow:

  $ time git branch --no-merged
  [...]

  real    0m9.177s
  user    0m9.027s
  sys     0m0.129s

when running it on tip/master:

  http://people.redhat.com/mingo/tip.git/README

	Ingo

Re: q: faster way to integrate/merge lots of topic branches?

From: Björn Steinbrink <hidden>
Date: 2016-06-15 22:45:00

On 2008.07.23 15:05:18 +0200, Ingo Molnar wrote:
I've got the following, possibly stupid question: is there a way to 
merge a healthy number of topic branches into the master branch in a 
quicker way, when most of the branches are already merged up?

Right now i've got something like this scripted up:

  for B in $(git-branch | cut -c3- ); do git-merge $B; done 
Not yet in any release (AFAICT), but with git.git master, you could use:

for B in $(git branch --no-merged); do git-merge $B; done


Or with earlier versions, this should work, but it's a lot slower:

for B in $(git branch | cut -c3- ); do
	[[ -n "$(git rev-list -1 HEAD..$B)" ]] && git merge $B;
done

Björn

Re: q: faster way to integrate/merge lots of topic branches?

From: Santi Béjar <hidden>
Date: 2016-06-15 22:45:00

On Wed, Jul 23, 2008 at 15:05, Ingo Molnar [off-list ref] wrote:
I've got the following, possibly stupid question: is there a way to
merge a healthy number of topic branches into the master branch in a
quicker way, when most of the branches are already merged up?
You could filter upfront the branches that are already merged up with:

git show-branch --independent <commits>

but it has a limit of 25 refs.

Santi

Re: q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

* Sergey Vlasov [off-list ref] wrote:
On Wed, 23 Jul 2008 15:05:18 +0200 Ingo Molnar wrote:
quoted
Anyone can simulate it by switching to the linus/master branch of the
current Linux kernel tree, and doing:

   time for ((i=0; i<140; i++)); do git-merge v2.6.26; done

   real    1m26.397s
   user    1m10.048s
   sys     0m13.944s
Timing results here (E6750 @ 2.66GHz):
41.61s user 3.71s system 99% cpu 45.530 total

However, testing whether there is something new to merge could be
performed significantly faster:

$ time sh -c 'for ((i=0; i<140; i++)); do [ -n "$(git rev-list --max-count=1 v2.6.26 ^HEAD)" ]; done'
sh -c   5.49s user 0.26s system 99% cpu 5.786 total

The same loop with "git merge-base v2.6.26 HEAD" takes about 40 
seconds here - apparently finding the merge base is the expensive 
part, and it makes sense to avoid it if you expect that most of your 
branches do not contain anything new to merge.
using git-fastmerge i get 2.4 seconds:

  $ time for ((i=0; i<140; i++)); do git-fastmerge v2.6.26; done
  [...]
  real    0m2.388s
  user    0m1.211s
  sys     0m1.131s

for something that 'progresses' in a forward manner (which merges do 
fundamentally) nothing beats the performance of a timestamped cache i 
think.

at least for my usecase.

Even assuming that the filesystem is sane, is my merge-cache 
implementation semantically equivalent to a git-merge? One detail is 
that i suspect it is not equivalent in the git-merge --no-ff case. (but 
that is a not too interesting non-default case anyway)

	Ingo

Re: q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

* Ingo Molnar [off-list ref] wrote:
Even assuming that the filesystem is sane, is my merge-cache 
implementation semantically equivalent to a git-merge? One detail is 
that i suspect it is not equivalent in the git-merge --no-ff case. 
(but that is a not too interesting non-default case anyway)
actually, since --no-ff creates a merge commit and thus propagates the 
head sha1, this should work fine as well.

(besides the small detail that my script has $1 hardcoded so parameters 
are not properly passed onto.)

	Ingo

Re: q: faster way to integrate/merge lots of topic branches?

From: Jay Soffian <hidden>
Date: 2016-06-15 22:45:00

On Wed, Jul 23, 2008 at 9:49 AM, Ingo Molnar [off-list ref] wrote:
#!/bin/bash

usage () {
 echo 'usage: git-fastmerge <refspec>..'
 exit -1
}

[ $# = 0 ] && usage

BRANCH=$1

MERGECACHE=.git/mergecache

[ ! -d $MERGECACHE ] && { mkdir $MERGECACHE || usage; }

HEAD_SHA1=$(git-log -1 --pretty=format:"%H")
BRANCH_SHA1=$(git-log -1 --pretty=format:"%H" $BRANCH)

CACHE=$MERGECACHE/$HEAD_SHA1/$BRANCH_SHA1

[ -f "$CACHE" -a "$CACHE" -nt .git/refs/heads/$BRANCH_SHA1 ] && {
Shouldn't this be:

[ -f "$CACHE" -a "$CACHE" -nt .git/refs/heads/$BRANCH ] && {

?

j.

Re: q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

* Jay Soffian [off-list ref] wrote:
quoted
CACHE=$MERGECACHE/$HEAD_SHA1/$BRANCH_SHA1

[ -f "$CACHE" -a "$CACHE" -nt .git/refs/heads/$BRANCH_SHA1 ] && {
Shouldn't this be:

[ -f "$CACHE" -a "$CACHE" -nt .git/refs/heads/$BRANCH ] && {

?
yeah, i just figured it out too ... the hard way :)

Updated script below. This works fine across resets in the master 
branch.

While it's fast in the empty-merge case, it's not as fast as i'd like it 
to be in the almost-empty-merge case.

	Ingo

--------------{ git-fastmerge }-------------------->
#!/bin/bash

usage () {
  echo 'usage: tip-fastmerge <refspec>..'
  exit -1
}

[ $# = 0 ] && usage

BRANCH=$1

MERGECACHE=.git/mergecache

[ ! -d $MERGECACHE ] && { mkdir $MERGECACHE || usage; }

HEADREF=.git/$(cut -d' ' -f2 .git/HEAD)

HEAD_SHA1=$(git-log -1 --pretty=format:"%H")
BRANCH_SHA1=$(git-log -1 --pretty=format:"%H" $BRANCH)

CACHE=$MERGECACHE/$HEAD_SHA1/$BRANCH_SHA1

[ -f "$CACHE" -a "$CACHE" -nt "$HEADREF" ] && {
# echo "merge-cache hit on HEAD <= $1"
  exit 0
}

git-merge $1 && {
  mkdir -p $(dirname $CACHE)
  touch $CACHE
}

Re: q: faster way to integrate/merge lots of topic branches?

From: Miklos Vajna <hidden>
Date: 2016-06-15 22:45:00

On Wed, Jul 23, 2008 at 03:40:41PM +0200, Andreas Ericsson [off-list ref] wrote:
With the builtin merge (which is in next)
Just a small correction: it's already in master.

Re: q: faster way to integrate/merge lots of topic branches?

From: Ingo Molnar <hidden>
Date: 2016-06-15 22:45:00

* Ingo Molnar [off-list ref] wrote:
quoted
Shouldn't this be:

[ -f "$CACHE" -a "$CACHE" -nt .git/refs/heads/$BRANCH ] && {

?
yeah, i just figured it out too ... the hard way :)

Updated script below. This works fine across resets in the master 
branch.

While it's fast in the empty-merge case, it's not as fast as i'd like 
it to be in the almost-empty-merge case.
When i update a topic branch, i first get a relatively fast run:

  earth4:~/tip> time todo-merge-all
  merging all branches ...
  Auto-merged arch/x86/kernel/genx2apic_uv_x.c
  Merge made by recursive.
   arch/x86/kernel/genx2apic_uv_x.c |    1 -
   1 files changed, 0 insertions(+), 1 deletions(-)
  ... merge done.

  real    0m6.625s
  user    0m3.740s
  sys     0m2.563s

Then on the next run it's slower:

  earth4:~/tip> time todo-merge-all
  merging all branches ...
  ... merge done.

  real    0m30.823s
  user    0m23.403s
  sys     0m7.545s

that's unfortunate. The freshly updated topic branch was at the end of 
the run, now all other topic branches will have to run slow at least 
once until they become cached again.

Perhaps the cache should update all other current topics to the new 
sha1, to establish the fact that they were not merged this time. (and 
that they are still not to be merged)

(It's still much faster than completely uncached though, because of the 
overlap in sha1's.)

Third (empty) run is fast again, because it's fully cached:

  earth4:~/tip> time todo-merge-all
  merging all branches ...
  ... merge done.

  real    0m3.036s
  user    0m1.360s
  sys     0m1.782s

But it would be nice if the cache worked more intelligently in the 
one-topic-updated-only case as well.

	Ingo

Re: q: faster way to integrate/merge lots of topic branches?

From: Linus Torvalds <torvalds@linux-foundation.org>
Date: 2016-06-15 22:45:00


On Wed, 23 Jul 2008, Ingo Molnar wrote:
I've got the following, possibly stupid question: is there a way to 
merge a healthy number of topic branches into the master branch in a 
quicker way, when most of the branches are already merged up?

Right now i've got something like this scripted up:

  for B in $(git-branch | cut -c3- ); do git-merge $B; done 

It takes a lot of time to run on even a 3.45GHz box:

  real    0m53.228s
  user    0m41.134s
  sys     0m11.405s
This is almost certainly because a lot of your branches are a long way 
back in the history, and just parsing the commit history is old.

For example, doing a no-op merge of something old like v2.6.24 (which is 
obviously already merged) takes half a second for me:

	[torvalds@woody linux]$ time git merge v2.6.24
	Already up-to-date.

	real	0m0.546s
	user	0m0.488s
	sys	0m0.008s

and it gets worse the further back in history you go (going back to 2.6.14 
takes a second and a half - plus any IO needed, of course).

And just about _all_ of it is literally just unpacking the commits as you 
start going backwards from the current point, eg:

	[torvalds@woody linux]$ time ~/git/git merge v2.6.14
	Already up-to-date.
	real	0m1.540s

vs

	[torvalds@woody linux]$ time git rev-list ..v2.6.14
	real	0m1.407s

(The merge loop isn't quite as optimized as the regular revision 
traversal, so you see it being slower, but you can still see that it's 
roughly in the same class).

The merge gets a bit more expensive still if you have enabled merge 
summaries (because now it traverses the lists twice - once for merge 
bases, once for logs), but that's still a secondary effect (ie it adds 
another 10% or so to the cost, but the base cost is still very much about 
the parsing of the commits).

In fact, the two top entries in a profile look roughly like:

	102161   70.2727  libz.so.1.2.3            libz.so.1.2.3            (no symbols)
	7685      5.2862  git                      git                      find_pack_entry_one
	...

ie 70% of the time is just purely unpacking the data, and another 5% is 
just finding it. We could perhaps improve on it, but not a whole lot.

Now, quite frankly, I don't think that times on the order of one second 
are worth worrying about for _regular_ merges, and the whole (and only) 
reason you see this as a performance problem is that you're basically 
automating it over a ton of branches, with most of them being old and 
already merged.

But that also points to a solution: instead of trying to merge them one at 
a time, and doing the costly revision traversal over and over and over 
again, do the costly thing _once_, and then you can just filter out the 
branches that aren't interesting.

So instead of doing

	for B in $(git-branch | cut -c3- ); do git-merge $B; done

the obvious optimization is to add "--no-merged" to the "git branch" call. 
That itself is expensive (ie doing "git branch --no-merged" will have to 
traverse at least as far back as the oldest branch), so that phase will be 
AT LEAST as expensive as one of the merges (and probably quite a bit more: 
I suspect "--no-merged" isn't very heavily optimized), but if a lot of 
your branches are already fully merged, it will do all that work _once_, 
and then avoid it for the merges themselves.

So the _trivial_ solution is to just change it to

	for B in $(git branch --no-merged | cut -c3- ); do git-merge $B; done

and that may already fix it in practice for you, bringing the cost down by 
a factor of two or more, depending on the exact pattern (of course, it 
could also make the cost go _up_ - if it turns out that none of the 
branches are merged).

Other solutions exist, but they get much uglier. Octopus merges are more 
efficient, for example, for all the same reasons - it keeps the commit 
traversal in a single process, and thus avoids having to re-parse the 
whole history down to the common base. But they have other problems, of 
course.

			Linus

Re: q: faster way to integrate/merge lots of topic branches?

From: Linus Torvalds <torvalds@linux-foundation.org>
Date: 2016-06-15 22:45:00


Ahh, missed that somebody already suggested it.

On Wed, 23 Jul 2008, Ingo Molnar wrote:
hm, it's very slow:

  $ time git branch --no-merged
  [...]

  real    0m9.177s
  user    0m9.027s
  sys     0m0.129s

when running it on tip/master:

  http://people.redhat.com/mingo/tip.git/README
We can probably speed it up, but more importantly, even if we don't, it's 
slow _once_.

It's worth taking a 9s hit, if that means that you can then skip half of 
the merges entirely, and thus win half of the 53s cost. 

But I'll look if there's a way to cut it down from 9s. I suspect it has to 
traverse the whole history to make 100% sure that something isn't merged, 
but even that should be faster than 9s.

		Linus

Re: q: faster way to integrate/merge lots of topic branches?

From: Linus Torvalds <torvalds@linux-foundation.org>
Date: 2016-06-15 22:45:00


On Wed, 23 Jul 2008, Linus Torvalds wrote:
But I'll look if there's a way to cut it down from 9s. I suspect it has to 
traverse the whole history to make 100% sure that something isn't merged, 
but even that should be faster than 9s.
Heh. It should be trivially doable _much_ faster, but the has_commit() 
logic really relies on re-doing the "in_merge_base()" thing over and over 
again (clearing the bits), instead of just populating the object list with 
a "already seen" bit and lettign that expand over time.

So using "git branch --no-merged" does avoid re-parsing the commits over 
and over again (which is a pretty big win), but the way the code is 
written it does end up traversing the commit list fully for every single 
branch. That's quite horrible.

Lars added to Cc list in the hope that he'll be embarrassed enough about 
the performance to try to fix it ;)

		Linus

Re: q: faster way to integrate/merge lots of topic branches?

From: Pierre Habouzit <hidden>
Date: 2016-06-15 22:45:00

On Wed, Jul 23, 2008 at 05:59:01PM +0000, Linus Torvalds wrote:
In fact, the two top entries in a profile look roughly like:

	102161   70.2727  libz.so.1.2.3            libz.so.1.2.3            (no symbols)
	7685      5.2862  git                      git                      find_pack_entry_one
	...

ie 70% of the time is just purely unpacking the data, and another 5% is 
just finding it. We could perhaps improve on it, but not a whole lot.
  Well there is an easy way though, that could reduce that: using
adaptative compression. I proposed a patch once upon a time, that set
the compression strengh to 0 for "small" objects with a configurable
cut-off. If you do that, most trees, commits messages and so on aren't
compressed, and it will reduce (with IIRC a 5-liner) this time quite
dramatically.

  I could maybe resurect it to see if for people that do the kind of
things Ingo does it helps. By setting the cut-off at 1k, I had packs
being less than 1% bigger IIRC. I'll try to find it again and run your
tests with it to see how much it helps.

  [ Of course, it doesn't invalidate the rest of your mail about being
    more clever with git-merge, but still, we could reduce this 70% of
    zlib time quite a lot with that ]

-- 
·O·  Pierre Habouzit
··O                                                madcoder@debian.org
OOO                                                http://www.madism.org

Re: q: faster way to integrate/merge lots of topic branches?

From: Pierre Habouzit <hidden>
Date: 2016-06-15 22:45:00

On Wed, Jul 23, 2008 at 07:09:20PM +0000, Pierre Habouzit wrote:
On Wed, Jul 23, 2008 at 05:59:01PM +0000, Linus Torvalds wrote:
quoted
In fact, the two top entries in a profile look roughly like:

	102161   70.2727  libz.so.1.2.3            libz.so.1.2.3            (no symbols)
	7685      5.2862  git                      git                      find_pack_entry_one
	...

ie 70% of the time is just purely unpacking the data, and another 5% is 
just finding it. We could perhaps improve on it, but not a whole lot.
  Well there is an easy way though, that could reduce that: using
adaptative compression. I proposed a patch once upon a time, that set
the compression strengh to 0 for "small" objects with a configurable
cut-off. If you do that, most trees, commits messages and so on aren't
compressed, and it will reduce (with IIRC a 5-liner) this time quite
dramatically.

  I could maybe resurect it to see if for people that do the kind of
things Ingo does it helps. By setting the cut-off at 1k, I had packs
being less than 1% bigger IIRC. I'll try to find it again and run your
tests with it to see how much it helps.
  Unsurprisingly with a 1024o cutoff, the numbers are (first run is
forced cold-cache with /proc/.../drop_caches, second is the best run of 5):

default git:

    3.10user 0.16system 0:08.10elapsed 40%CPU (0avgtext+0avgdata 0maxresident)k
    116152inputs+0outputs (671major+35286minor)pagefaults 0swaps

    2.01user 0.11system 0:02.12elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k
    0inputs+0outputs (0major+35958minor)pagefaults 0swaps

With a 1024k cutoff:

    1.16user 0.13system 0:08.29elapsed 15%CPU (0avgtext+0avgdata 0maxresident)k
    154208inputs+0outputs (947major+39777minor)pagefaults 0swaps

    0.76user 0.06system 0:00.82elapsed 100%CPU (0avgtext+0avgdata 0maxresident)k
    0inputs+0outputs (0major+40724minor)pagefaults 0swaps


According to [0], a 1k cutoff meant something like a 10% larger pack. 512o
meant an almost identical pack in size, but with reduced performance
improvements.


  [0] http://thread.gmane.org/gmane.comp.version-control.git/70019/focus=70250
-- 
·O·  Pierre Habouzit
··O                                                madcoder@debian.org
OOO                                                http://www.madism.org

Re: q: faster way to integrate/merge lots of topic branches?

From: Pierre Habouzit <hidden>
Date: 2016-06-15 22:45:00

On mer, jui 23, 2008 at 08:27:22 +0000, Pierre Habouzit wrote:
On Wed, Jul 23, 2008 at 07:09:20PM +0000, Pierre Habouzit wrote:
quoted
On Wed, Jul 23, 2008 at 05:59:01PM +0000, Linus Torvalds wrote:
quoted
In fact, the two top entries in a profile look roughly like:

	102161   70.2727  libz.so.1.2.3            libz.so.1.2.3            (no symbols)
	7685      5.2862  git                      git                      find_pack_entry_one
	...

ie 70% of the time is just purely unpacking the data, and another 5% is 
just finding it. We could perhaps improve on it, but not a whole lot.
  Well there is an easy way though, that could reduce that: using
adaptative compression. I proposed a patch once upon a time, that set
the compression strengh to 0 for "small" objects with a configurable
cut-off. If you do that, most trees, commits messages and so on aren't
compressed, and it will reduce (with IIRC a 5-liner) this time quite
dramatically.

  I could maybe resurect it to see if for people that do the kind of
things Ingo does it helps. By setting the cut-off at 1k, I had packs
being less than 1% bigger IIRC. I'll try to find it again and run your
tests with it to see how much it helps.
  Unsurprisingly with a 1024o cutoff, the numbers are (first run is
forced cold-cache with /proc/.../drop_caches, second is the best run of 5):

default git:

    3.10user 0.16system 0:08.10elapsed 40%CPU (0avgtext+0avgdata 0maxresident)k
    116152inputs+0outputs (671major+35286minor)pagefaults 0swaps

    2.01user 0.11system 0:02.12elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k
    0inputs+0outputs (0major+35958minor)pagefaults 0swaps

With a 1024k cutoff:

    1.16user 0.13system 0:08.29elapsed 15%CPU (0avgtext+0avgdata 0maxresident)k
    154208inputs+0outputs (947major+39777minor)pagefaults 0swaps

    0.76user 0.06system 0:00.82elapsed 100%CPU (0avgtext+0avgdata 0maxresident)k
    0inputs+0outputs (0major+40724minor)pagefaults 0swaps
With a 512o cutoff:

    1.49user 0.17system 0:07.50elapsed 22%CPU (0avgtext+0avgdata 0maxresident)k
    127648inputs+0outputs (780major+36687minor)pagefaults 0swaps

    1.54user 0.07system 0:01.61elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k
    0inputs+0outputs (0major+37467minor)pagefaults 0swaps


What I bench, I see I forgot to mention, is: git merge v2.6.14. And
the respective pack sizes:

    214M  .git-0/objects/pack/pack-bfeec11abed1ec6d046bc954b94d70ba81716356.pack
    225M  .git-512/objects/pack/pack-bfeec11abed1ec6d046bc954b94d70ba81716356.pack
    243M  .git-1024/objects/pack/pack-bfeec11abed1ec6d046bc954b94d70ba81716356.pack

.git-0 is cheating because it was generated with a way deeper window and
memory window that the other ones, but it allow to give rough
impressions.


-- 
·O·  Pierre Habouzit
··O                                                madcoder@debian.org
OOO                                                http://www.madism.org
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help