From: Jeff King <hidden> Date: 2018-08-17 20:54:31
This series more aggressively reuses on-disk deltas to serve fetches
when reachability bitmaps tell us a more complete picture of what the
client has. That saves server CPU and results in smaller packs. See the
final patch for numbers and more discussion.
It's a resurrection of this very old series:
https://public-inbox.org/git/20140326072215.GA31739@sigill.intra.peff.net/
The core idea is good, but it got left as "I should dig into this more
to see if we can do even better". In fact, I _did_ do some of that
digging, as you can see in the thread, so I'm mildly embarrassed not to
have reposted it before now.
We've been using that original at GitHub since 2014, which I think helps
demonstrate the correctness of the approach (and the numbers here and in
that thread show that performance is generally a net win over the status
quo).
I's rebased on top of the current master, since the original made some
assumptions about struct object_entry that are no longer true post-v2.18
(due to the struct-shrinking exercise). So I fixed that and a few other
rough edges. But that also means you're not getting code with 4-years of
production testing behind it. :)
The other really ugly thing in the original was the way it leaked
object_entry structs (though in practice that didn't really matter since
we needed them until the end of the program anyway). This version fixes
that.
[1/6]: t/perf: factor boilerplate out of test_perf
[2/6]: t/perf: factor out percent calculations
[3/6]: t/perf: add infrastructure for measuring sizes
[4/6]: t/perf: add perf tests for fetches from a bitmapped server
[5/6]: pack-bitmap: save "have" bitmap from walk
[6/6]: pack-objects: reuse on-disk deltas for thin "have" objects
builtin/pack-objects.c | 28 +++++++----
pack-bitmap.c | 23 +++++++++-
pack-bitmap.h | 7 +++
pack-objects.c | 19 ++++++++
pack-objects.h | 20 +++++++-
t/perf/README | 25 ++++++++++
t/perf/aggregate.perl | 69 ++++++++++++++++++++++------
t/perf/p5311-pack-bitmaps-fetch.sh | 45 ++++++++++++++++++
t/perf/perf-lib.sh | 74 +++++++++++++++++++-----------
9 files changed, 258 insertions(+), 52 deletions(-)
create mode 100755 t/perf/p5311-pack-bitmaps-fetch.sh
-Peff
From: Jeff King <hidden> Date: 2018-08-17 20:55:09
About half of test_perf() is boilerplate preparing to run
_any_ test, and the other half is specifically running a
timing test. Let's split it into two functions, so that we
can reuse the boilerplate in future commits.
Signed-off-by: Jeff King <redacted>
---
Best viewed with "-w".
t/perf/perf-lib.sh | 61 ++++++++++++++++++++++++++--------------------
1 file changed, 35 insertions(+), 26 deletions(-)
From: Jeff King <hidden> Date: 2018-08-17 20:55:27
This will let us reuse the code when we add new values to
aggregate besides times.
Signed-off-by: Jeff King <redacted>
---
t/perf/aggregate.perl | 21 ++++++++++++---------
1 file changed, 12 insertions(+), 9 deletions(-)
From: Jeff King <hidden> Date: 2018-08-17 20:56:40
The main objective of scripts in the perf framework is to
run "test_perf", which measures the time it takes to run
some operation. However, it can also be interesting to see
the change in the output size of certain operations.
This patch introduces test_size, which records a single
numeric output from the test and shows it in the aggregated
output (with pretty printing and relative size comparison).
Signed-off-by: Jeff King <redacted>
---
t/perf/README | 25 ++++++++++++++++++++++
t/perf/aggregate.perl | 48 ++++++++++++++++++++++++++++++++++++++-----
t/perf/perf-lib.sh | 13 ++++++++++++
3 files changed, 81 insertions(+), 5 deletions(-)
@@ -168,3 +168,28 @@ that While we have tried to make sure that it can cope with embedded whitespace and other special characters, it will not work with multi-line data.++Rather than tracking the performance by run-time as `test_perf` does, you+may also track output size by using `test_size`. The stdout of the+function should be a single numeric value, which will be captured and+shown in the aggregated output. For example:++ test_perf 'time foo' '+ ./foo >foo.out+ '++ test_size 'output size'+ wc -c <foo.out+ '++might produce output like:++ Test origin HEAD+ -------------------------------------------------------------+ 1234.1 time foo 0.37(0.79+0.02) 0.26(0.51+0.02) -29.7%+ 1234.2 output size 4.3M 3.6M -14.7%++The item being measured (and its units) is up to the test; the context+and the test title should make it clear to the user whether bigger or+smaller numbers are better. Unlike test_perf, the test code will only be+run once, since output sizes tend to be more deterministic than timings.
@@ -32,9 +38,15 @@ sub relative_change {subformat_times{my($r,$u,$s,$firstr)=@_;+# no value means we did not finish the testif(!defined$r){return"<missing>";}+# a single value means we have a size, not times+if(!defined$u){+returnformat_size($r,$firstr);+}+# otherwise, we have real/user/system timesmy$out=sprintf"%.2f(%.2f+%.2f)",$r,$u,$s;$out.=' '.relative_change($r,$firstr)ifdefined$firstr;return$out;
@@ -54,6 +66,25 @@ sub usage {exit(1);}+subhuman_size{+my$n=shift;+my@units=('',qw(K M G));+while($n>900&&@units>1){+$n/=1000;+shift@units;+}+return$nunlesslength$units[0];+returnsprintf'%.1f%s',$n,$units[0];+}++subformat_size{+my($size,$first)=@_;+# match the width of a time: 0.00(0.00+0.00)+my$out=sprintf'%15s',human_size($size);+$out.=' '.relative_change($size,$first)ifdefined$first;+return$out;+}+my(@dirs,%dirnames,%dirabbrevs,%prefixes,@tests,$codespeed,$sortby,$subsection,$reponame);
@@ -184,7 +215,14 @@ sub print_default_results {my$firstr;formy$i(0..$#dirs){my$d=$dirs[$i];-$times{$prefixes{$d}.$t}=[get_times("$resultsdir/$prefixes{$d}$t.times")];+my$base="$resultsdir/$prefixes{$d}$t";+$times{$prefixes{$d}.$t}=[];+foreachmy$type(qw(times size)){+if(-e"$base.$type"){+$times{$prefixes{$d}.$t}=[get_times("$base.$type")];+last;+}+}my($r,$u,$s)=@{$times{$prefixes{$d}.$t}};my$w=lengthformat_times($r,$u,$s,$firstr);$colwidth[$i]=$wif$w>$colwidth[$i];
@@ -231,6 +231,19 @@ test_perf () {test_wrapper_test_perf_"$@"}+test_size_(){+say>&3"running: $2"+iftest_eval_"$2"3>"$base".size;then+test_ok_"$1"+else+test_failure_"$@"+fi+}++test_size(){+test_wrapper_test_size_"$@"+}+# We extend test_done to print timings at the end (./run disables this# and does it after running everything) test_at_end_hook_(){
From: Jeff King <hidden> Date: 2018-08-17 20:57:44
A server with bitmapped packs can serve a clone very
quickly. However, fetches are not necessarily made any
faster, because we spend a lot less time in object traversal
(which is what bitmaps help with) and more time finding
deltas (because we may have to throw out on-disk deltas if
the client does not have the base).
As a first step to making this faster, this patch introduces
a new perf script to measure fetches into a repo of various
ages from a fully-bitmapped server.
We separately measure the work done by the server (in
pack-objects) and that done by the client (in index-pack).
Furthermore, we measure the size of the resulting pack.
Breaking it down like this (instead of just doing a regular
"git fetch") lets us see how much each side benefits from
any changes. And since we know the pack size, if we estimate
the network speed, then one could calculate a complete
wall-clock time for the operation (though the script does
not do this automatically).
Signed-off-by: Jeff King <redacted>
---
t/perf/p5311-pack-bitmaps-fetch.sh | 45 ++++++++++++++++++++++++++++++
1 file changed, 45 insertions(+)
create mode 100755 t/perf/p5311-pack-bitmaps-fetch.sh
@@ -0,0 +1,45 @@+#!/bin/sh++test_description='performance of fetches from bitmapped packs'+../perf-lib.sh++test_perf_default_repo++test_expect_success'create bitmapped server repo''+gitconfigpack.writebitmapstrue&&+gitconfigpack.writebitmaphashcachetrue&&+gitrepack-ad+'++# simulate a fetch from a repository that last fetched N days ago, for+# various values of N. We do so by following the first-parent chain,+# and assume the first entry in the chain that is N days older than the current+# HEAD is where the HEAD would have been then.+fordaysin1248163264128;do+title=$(printf'%10s'"($days days)")+test_expect_success"setup revs from $days days ago"'+now=$(gitlog-1--format=%ctHEAD)&&+then=$(($now-($days*86400)))&&+tip=$(gitrev-list-1--first-parent--until=$thenHEAD)&&+{+echoHEAD&&+echo^$tip+}>revs+'++test_perf"server $title"'+gitpack-objects--stdout--revs\+--thin--delta-base-offset\+<revs>tmp.pack+'++test_size"size $title"'+wc-c<tmp.pack+'++test_perf"client $title"'+gitindex-pack--stdin--fix-thin<tmp.pack+'+done++test_done
From: Jeff King <hidden> Date: 2018-08-17 20:59:24
When we do a bitmap walk, we save the result, which
represents (WANTs & ~HAVEs); i.e., every object we care
about visiting in our walk. However, we throw away the
haves bitmap, which can sometimes be useful, too. Save it
and provide an access function so code which has performed a
walk can query it.
A few notes on the accessor interface:
- the bitmap code calls these "haves" because it grew out
of the want/have negotiation for fetches. But really,
these are simply the objects that would be flagged
UNINTERESTING in a regular traversal. Let's use that
more universal nomenclature for the external module
interface. We may want to change the internal naming
inside the bitmap code, but that's outside the scope of
this patch.
- it still uses a bare "sha1" rather than "oid". That's
true of all of the bitmap code. And in this particular
instance, our caller in pack-objects is dealing with the
bare sha1 that comes from a packed REF_DELTA (we're
pointing directly to the mmap'd pack on disk). That's
something we'll have to deal with as we transition to a
new hash, but we can wait and see how the caller ends up
being fixed and adjust this interface accordingly.
Signed-off-by: Jeff King <redacted>
---
Funny story: the earlier version of this series called it bitmap_have().
That caused a bug later when somebody tried to build on it, thinking it
was "does the bitmap have this object in the result". Oops. Hence the
more descriptive name.
pack-bitmap.c | 23 ++++++++++++++++++++++-
pack-bitmap.h | 7 +++++++
2 files changed, 29 insertions(+), 1 deletion(-)
@@ -86,6 +86,9 @@ struct bitmap_index {/* Bitmap result of the last performed walk */structbitmap*result;+/* "have" bitmap from the last performed walk */+structbitmap*haves;+/* Version of the bitmap index */unsignedintversion;
@@ -1114,5 +1117,23 @@ void free_bitmap_index(struct bitmap_index *b)free(b->ext_index.objects);free(b->ext_index.hashes);bitmap_free(b->result);+bitmap_free(b->haves);free(b);}++intbitmap_has_sha1_in_uninteresting(structbitmap_index*bitmap_git,+constunsignedchar*sha1)+{+intpos;++if(!bitmap_git)+return0;/* no bitmap loaded */+if(!bitmap_git->haves)+return0;/* walk had no "haves" */++pos=bitmap_position_packfile(bitmap_git,sha1);+if(pos<0)+return0;++returnbitmap_get(bitmap_git->haves,pos);+}
From: Jeff King <hidden> Date: 2018-08-17 21:06:09
When we serve a fetch, we pass the "wants" and "haves" from
the fetch negotiation to pack-objects. That tells us not
only which objects we need to send, but we also use the
boundary commits as "preferred bases": their trees and blobs
are candidates for delta bases, both for reusing on-disk
deltas and for finding new ones.
However, this misses some opportunities. Modulo some special
cases like shallow or partial clones, we know that every
object reachable from the "haves" could be a preferred base.
We don't use them all for two reasons:
1. It's expensive to traverse the whole history and
enumerate all of the objects the other side has.
2. The delta search is expensive, so we want to keep the
number of candidate bases sane. The boundary commits
are the most likely to work.
When we have reachability bitmaps, though, reason 1 no
longer applies. We can efficiently compute the set of
reachable objects on the other side (and in fact already did
so as part of the bitmap set-difference to get the list of
interesting objects). And using this set conveniently
covers the shallow and partial cases, since we have to
disable the use of bitmaps for those anyway.
The second reason argues against using these bases in the
search for new deltas. But there's one case where we can use
this information for free: when we have an existing on-disk
delta that we're considering reusing, we can do so if we
know the other side has the base object. This in fact saves
time during the delta search, because it's one less delta we
have to compute.
And that's exactly what this patch does: when we're
considering whether to reuse an on-disk delta, if bitmaps
tell us the other side has the object (and we're making a
thin-pack), then we reuse it.
Here are the results on p5311 using linux.git, which
simulates a client fetching after `N` days since their last
fetch:
Test origin HEAD
--------------------------------------------------------------------------
5311.3: server (1 days) 0.27(0.27+0.04) 0.12(0.09+0.03) -55.6%
5311.4: size (1 days) 0.9M 237.0K -73.7%
5311.5: client (1 days) 0.04(0.05+0.00) 0.10(0.10+0.00) +150.0%
5311.7: server (2 days) 0.34(0.42+0.04) 0.13(0.10+0.03) -61.8%
5311.8: size (2 days) 1.5M 347.7K -76.5%
5311.9: client (2 days) 0.07(0.08+0.00) 0.16(0.15+0.01) +128.6%
5311.11: server (4 days) 0.56(0.77+0.08) 0.13(0.10+0.02) -76.8%
5311.12: size (4 days) 2.8M 566.6K -79.8%
5311.13: client (4 days) 0.13(0.15+0.00) 0.34(0.31+0.02) +161.5%
5311.15: server (8 days) 0.97(1.39+0.11) 0.30(0.25+0.05) -69.1%
5311.16: size (8 days) 4.3M 1.0M -76.0%
5311.17: client (8 days) 0.20(0.22+0.01) 0.53(0.52+0.01) +165.0%
5311.19: server (16 days) 1.52(2.51+0.12) 0.30(0.26+0.03) -80.3%
5311.20: size (16 days) 8.0M 2.0M -74.5%
5311.21: client (16 days) 0.40(0.47+0.03) 1.01(0.98+0.04) +152.5%
5311.23: server (32 days) 2.40(4.44+0.20) 0.31(0.26+0.04) -87.1%
5311.24: size (32 days) 14.1M 4.1M -70.9%
5311.25: client (32 days) 0.70(0.90+0.03) 1.81(1.75+0.06) +158.6%
5311.27: server (64 days) 11.76(26.57+0.29) 0.55(0.50+0.08) -95.3%
5311.28: size (64 days) 89.4M 47.4M -47.0%
5311.29: client (64 days) 5.71(9.31+0.27) 15.20(15.20+0.32) +166.2%
5311.31: server (128 days) 16.15(36.87+0.40) 0.91(0.82+0.14) -94.4%
5311.32: size (128 days) 134.8M 100.4M -25.5%
5311.33: client (128 days) 9.42(16.86+0.49) 25.34(25.80+0.46) +169.0%
In all cases we save CPU time on the server (sometimes
significant) and the resulting pack is smaller. We do spend
more CPU time on the client side, because it has to
reconstruct more deltas. But that's the right tradeoff to
make, since clients tend to outnumber servers. It just means
the thin pack mechanism is doing its job.
From the user's perspective, the end-to-end time of the
operation will generally be faster. E.g., in the 128-day
case, we saved 15s on the server at a cost of 16s on the
client. Since the resulting pack is 34MB smaller, this is a
net win if the network speed is less than 270Mbit/s. And
that's actually the worst case. The 64-day case saves just
over 11s at a cost of just under 11s. So it's a slight win
at any network speed, and the 40MB saved is pure bonus. That
trend continues for the smaller fetches.
The implementation itself is mostly straightforward, with
the new logic going into check_object(). But there are two
tricky bits.
The first is that check_object() needs access to the
relevant information (the thin flag and bitmap result). We
can do this by pushing these into program-lifetime globals.
The second is that the rest of the code assumes that any
reused delta will point to another "struct object_entry" as
its base. But by definition, we don't have such an entry!
I looked at a number of options that didn't quite work:
- we could use a different flag for reused deltas. But it's
not a single bit for "I'm being reused". We have to
actually store the oid of the base, which is normally
done by pointing to the existing object_entry. And we'd
have to modify all the code which looks at deltas.
- we could add the reused bases to the end of the existing
object_entry array. While this does create some extra
work as later stages consider the extra entries, it's
actually not too bad (we're not sending them, so they
don't cost much in the delta search, and at most we'd
have 2*N of them).
But there's a more subtle problem. Adding to the existing
array means we might need to grow it with realloc, which
could move the earlier entries around. While many of the
references to other entries are done by integer index,
some (including ones on the stack) use pointers, which
would become invalidated.
This isn't insurmountable, but it would require quite a
bit of refactoring (and it's hard to know that you've got
it all, since it may work _most_ of the time and then
fail subtly based on memory allocation patterns).
- we could allocate a new one-off entry for the base. In
fact, this is what an earlier version of this patch did.
However, since the refactoring brought in by ad635e82d6
(Merge branch 'nd/pack-objects-pack-struct', 2018-05-23),
the delta_idx code requires that both entries be in the
main packing list.
So taking all of those options into account, what I ended up
with is a separate list of "external bases" that are not
part of the main packing list. Each delta entry that points
to an external base has a single-bit flag to do so; we have a
little breathing room in the bitfield section of
object_entry.
This lets us limit the change primarily to the oe_delta()
and oe_set_delta_ext() functions. And as a bonus, most of
the rest of the code does not consider these dummy entries
at all, saving both runtime CPU and code complexity.
Signed-off-by: Jeff King <redacted>
---
builtin/pack-objects.c | 28 +++++++++++++++++++---------
pack-objects.c | 19 +++++++++++++++++++
pack-objects.h | 20 ++++++++++++++++++--
3 files changed, 56 insertions(+), 11 deletions(-)
@@ -177,3 +177,22 @@ struct object_entry *packlist_alloc(struct packing_data *pdata,returnnew_entry;}++voidoe_set_delta_ext(structpacking_data*pdata,+structobject_entry*delta,+constunsignedchar*sha1)+{+structobject_entry*base;++ALLOC_GROW(pdata->ext_bases,pdata->nr_ext+1,pdata->alloc_ext);+base=&pdata->ext_bases[pdata->nr_ext++];+memset(base,0,sizeof(*base));+hashcpy(base->idx.oid.hash,sha1);++/* These flags mark that we are not part of the actual pack output. */+base->preferred_base=1;+base->filled=1;++delta->ext_base=1;+delta->delta_idx=base-pdata->ext_bases+1;+}
If the traversal has not been performed, we pretend the
object was not reachable?
Is this a good API design, as it can be used when you do not
have done all preparations? similarly to prepare_bitmap_walk
we could have
if (!bitmap_git->result)
BUG("failed to perform bitmap walk before querying");
You seem to have rebased it to master resolving conflicts only. ;-)
Do we want to talk about object ids here instead?
(This is what I get to think about when reviewing this series
"bottom up". I use "git log -w -p master..HEAD" after applying
the patches, probably I should also use --reverse, such that I
get to see the commit message before the code for each commit
and yet only need to scroll in one direction.)
If the traversal has not been performed, we pretend the
object was not reachable?
If the traversal hasn't been performed, the results are not defined
(though I suspect yeah, it happens to say "no").
Is this a good API design, as it can be used when you do not
have done all preparations? similarly to prepare_bitmap_walk
we could have
if (!bitmap_git->result)
BUG("failed to perform bitmap walk before querying");
From: Stefan Beller <hidden> Date: 2018-08-17 22:59:54
On Fri, Aug 17, 2018 at 2:06 PM Jeff King [off-list ref] wrote:
When we serve a fetch, we pass the "wants" and "haves" from
the fetch negotiation to pack-objects. That tells us not
only which objects we need to send, but we also use the
boundary commits as "preferred bases": their trees and blobs
are candidates for delta bases, both for reusing on-disk
deltas and for finding new ones.
However, this misses some opportunities. Modulo some special
cases like shallow or partial clones, we know that every
object reachable from the "haves" could be a preferred base.
We don't use them all for two reasons:
s/all/at all/ ?
The first is that check_object() needs access to the
relevant information (the thin flag and bitmap result). We
can do this by pushing these into program-lifetime globals.
I discussed internally if extending the fetch protocol to include
submodule packs would be a good idea, as then you can get all
the superproject+submodule updates via one connection. This
gives some benefits, such as a more consistent view from the
superproject as well as already knowing the have/wants for
the submodule.
With this background story, moving things into globals
makes me sad, but I guess we can flip this decision once
we actually move towards "submodule packs in the
main connection".
The second is that the rest of the code assumes that any
reused delta will point to another "struct object_entry" as
its base. But by definition, we don't have such an entry!
I got lost here by the definition (which def?).
The delta that we look up from the bitmap, doesn't may
not be in the pack, but it could be based off of an object
the client already has in its object store and for that
there is no struct object_entry in memory.
Is that correct?
So taking all of those options into account, what I ended up
with is a separate list of "external bases" that are not
part of the main packing list. Each delta entry that points
to an external base has a single-bit flag to do so; we have a
little breathing room in the bitfield section of
object_entry.
This lets us limit the change primarily to the oe_delta()
and oe_set_delta_ext() functions. And as a bonus, most of
the rest of the code does not consider these dummy entries
at all, saving both runtime CPU and code complexity.
From: Jeff King <hidden> Date: 2018-08-17 23:32:37
On Fri, Aug 17, 2018 at 03:57:18PM -0700, Stefan Beller wrote:
On Fri, Aug 17, 2018 at 2:06 PM Jeff King [off-list ref] wrote:
quoted
When we serve a fetch, we pass the "wants" and "haves" from
the fetch negotiation to pack-objects. That tells us not
only which objects we need to send, but we also use the
boundary commits as "preferred bases": their trees and blobs
are candidates for delta bases, both for reusing on-disk
deltas and for finding new ones.
However, this misses some opportunities. Modulo some special
cases like shallow or partial clones, we know that every
object reachable from the "haves" could be a preferred base.
We don't use them all for two reasons:
s/all/at all/ ?
No, I meant "we don't use all of them". As in the Pokemon "gotta catch
'em all" slogan. ;)
Probably writing out "all of them" is better.
quoted
The first is that check_object() needs access to the
relevant information (the thin flag and bitmap result). We
can do this by pushing these into program-lifetime globals.
I discussed internally if extending the fetch protocol to include
submodule packs would be a good idea, as then you can get all
the superproject+submodule updates via one connection. This
gives some benefits, such as a more consistent view from the
superproject as well as already knowing the have/wants for
the submodule.
With this background story, moving things into globals
makes me sad, but I guess we can flip this decision once
we actually move towards "submodule packs in the
main connection".
I don't think it significantly changes the existing code, which is
already relying on a ton of globals (most notably to_pack). The first
step in doing multiple packs in the same process is going to be to shove
all of that into a "struct pack_objects_context" or similar, and these
can just follow the rest.
quoted
The second is that the rest of the code assumes that any
reused delta will point to another "struct object_entry" as
its base. But by definition, we don't have such an entry!
I got lost here by the definition (which def?).
The delta that we look up from the bitmap, doesn't may
not be in the pack, but it could be based off of an object
the client already has in its object store and for that
there is no struct object_entry in memory.
Is that correct?
Right, we are interested in objects that we _couldn't_ find a struct
for. I agree this could be more clear.
-Peff
From: Jeff King <hidden> Date: 2018-08-21 19:06:26
On Fri, Aug 17, 2018 at 04:54:27PM -0400, Jeff King wrote:
This series more aggressively reuses on-disk deltas to serve fetches
when reachability bitmaps tell us a more complete picture of what the
client has. That saves server CPU and results in smaller packs. See the
final patch for numbers and more discussion.
Here's a v2, with just a few cosmetic fixes to address the comments on
v1 (range-diff below).
[1/6]: t/perf: factor boilerplate out of test_perf
[2/6]: t/perf: factor out percent calculations
[3/6]: t/perf: add infrastructure for measuring sizes
[4/6]: t/perf: add perf tests for fetches from a bitmapped server
[5/6]: pack-bitmap: save "have" bitmap from walk
[6/6]: pack-objects: reuse on-disk deltas for thin "have" objects
builtin/pack-objects.c | 28 +++++++----
pack-bitmap.c | 25 +++++++++-
pack-bitmap.h | 7 +++
pack-objects.c | 19 ++++++++
pack-objects.h | 20 +++++++-
t/perf/README | 25 ++++++++++
t/perf/aggregate.perl | 69 ++++++++++++++++++++++------
t/perf/p5311-pack-bitmaps-fetch.sh | 45 ++++++++++++++++++
t/perf/perf-lib.sh | 74 +++++++++++++++++++-----------
9 files changed, 260 insertions(+), 52 deletions(-)
create mode 100755 t/perf/p5311-pack-bitmaps-fetch.sh
1: 89fa0ec8d8 ! 1: 3e1b94d7d6 pack-bitmap: save "have" bitmap from walk
@@ -69,6 +69,8 @@
+
+ if (!bitmap_git)
+ return 0; /* no bitmap loaded */
++ if (!bitmap_git->result)
++ BUG("failed to perform bitmap walk before querying");
+ if (!bitmap_git->haves)
+ return 0; /* walk had no "haves" */
+
2: f7ca0d59e3 ! 2: b8b2416aac pack-objects: reuse on-disk deltas for thin "have" objects
@@ -12,7 +12,7 @@
However, this misses some opportunities. Modulo some special
cases like shallow or partial clones, we know that every
object reachable from the "haves" could be a preferred base.
- We don't use them all for two reasons:
+ We don't use all of them for two reasons:
1. It's expensive to traverse the whole history and
enumerate all of the objects the other side has.
@@ -100,15 +100,16 @@
The second is that the rest of the code assumes that any
reused delta will point to another "struct object_entry" as
- its base. But by definition, we don't have such an entry!
+ its base. But of course the case we are interested in here
+ is the one where don't have such an entry!
I looked at a number of options that didn't quite work:
- - we could use a different flag for reused deltas. But it's
- not a single bit for "I'm being reused". We have to
- actually store the oid of the base, which is normally
- done by pointing to the existing object_entry. And we'd
- have to modify all the code which looks at deltas.
+ - we could use a flag to signal a reused delta, but it's
+ not a single bit. We have to actually store the oid of
+ the base, which is normally done by pointing to the
+ existing object_entry. And we'd have to modify all the
+ code which looks at deltas.
- we could add the reused bases to the end of the existing
object_entry array. While this does create some extra
@@ -173,7 +174,7 @@
static int depth = 50;
static int delta_search_threads;
static int pack_to_stdout;
-+static int thin = 0;
++static int thin;
static int num_preferred_base;
static struct progress *progress_state;
From: Jeff King <hidden> Date: 2018-08-21 19:08:35
On Tue, Aug 21, 2018 at 03:06:22PM -0400, Jeff King wrote:
On Fri, Aug 17, 2018 at 04:54:27PM -0400, Jeff King wrote:
quoted
This series more aggressively reuses on-disk deltas to serve fetches
when reachability bitmaps tell us a more complete picture of what the
client has. That saves server CPU and results in smaller packs. See the
final patch for numbers and more discussion.
Here's a v2, with just a few cosmetic fixes to address the comments on
v1 (range-diff below).