From: Junio C Hamano <hidden> Date: 2016-06-15 23:08:36
Jeff King [off-list ref] writes:
On Tue, Mar 01, 2016 at 03:36:26PM -0800, Junio C Hamano wrote:
quoted
This will be necessary when we start reading from a split bundle
where the header and the thin-pack data live in different files.
The in-core bundle header will read from a file that has the header,
and will record the path to that file. We would find the name of
the file that hosts the thin-pack data from the header, and we would
take that name as relative to the file we read the header from.
Neat. I'm hoping this means you're working on split bundles. :)
Let's just say that during the -rc freeze period, because I can stop
looking at or queuing completely new topics to encourage people who
are responsible for topics in the upcoming release to focus more on
responding to regressions and follow-up fixes necessary, I have a
better chance of having some leftover time to look into things
myself, at least enough to figure out what needs to be done, ;-)
What are the memory ownership rules for header.bundle_file?
Here you assign from either an argv parameter or a stack buffer, and
here...
quoted
@@ -112,6 +111,8 @@ void release_bundle_header(struct bundle_header *header) for (i = 0; i < header->references.nr; i++) free(header->references.list[i].name); free(header->references.list);++ free((void *)header->bundle_file); }
You free it.
The call in get_refs_from_bundle does do an xstrdup().
Good eyes.
Should we have:
void init_bundle_header(struct bundle_header *header, const char *file)
{
memset(header, 0, sizeof(*header));
header.bundle_file = xstrdup(file);
}
to abstract the whole procedure?
Maybe, maybe not. I'll decide after adding the bundle_version field
to the structure (which will be read from an existing bundle, but
which will have to be set for a bundle being created).
Thanks.
From: Junio C Hamano <hidden> Date: 2016-06-15 23:08:36
Even though the command does read the bundle header and checks to
see if it looks reasonable, the thin-pack data stream that follows
the header in the bundle file is not checked.
The documentation gives an incorrect impression that the data
contained in the bundle is validated, but the command is to validate
that the receiving repository is ready to accept the bundle, not to
check the validity of a bundle file itself.
Rephrase the paragraph to clarify this.
Signed-off-by: Junio C Hamano <redacted>
---
Documentation/git-bundle.txt | 9 ++++-----
1 file changed, 4 insertions(+), 5 deletions(-)
@@ -38,11 +38,10 @@ create <file>:: 'git-rev-list-args' arguments to define the bundle contents. verify <file>::- Used to check that a bundle file is valid and will apply- cleanly to the current repository. This includes checks on the- bundle format itself as well as checking that the prerequisite- commits exist and are fully linked in the current repository.- 'git bundle' prints a list of missing commits, if any, and exits+ Verifies that the given 'file' has a valid-looking bundle+ header, and that your repository has all prerequisite+ objects necessary to unpack the file as a bundle. The+ command prints a list of missing commits, if any, and exits with a non-zero status. list-heads <file>::
From: Junio C Hamano <hidden> Date: 2016-06-15 23:08:36
Here is a preview of the "split bundle" stuff that may serve as one
of the enabling technology to offload "git clone" traffic off of the
server core network to CDN.
Changes:
- The "checksum" bit about the in-bundle packdata, which was
incorrect, was dropped from the proposed log message of 1/4.
- "init_bundle_header()" helper has been added to 3/4; the name of
the new field in bundle_header structure is now "filename", not
"bundle_file". It is silly to name a field with "bundle" in it
when the structure is about a bundle already.
- 4/4 is new, and implements the unbundling part, i.e. running
either "git clone" or "git bundle unbundle" on the two files
after a split bundle is made available locally.
Junio C Hamano (4):
bundle doc: 'verify' is not about verifying the bundle
bundle: plug resource leak
bundle: keep a copy of bundle file name in the in-core bundle header
bundle v3: the beginning
Documentation/git-bundle.txt | 9 ++-
builtin/bundle.c | 6 +-
bundle.c | 142 +++++++++++++++++++++++++++++++++++++------
bundle.h | 8 ++-
t/t5704-bundle.sh | 64 +++++++++++++++++++
transport.c | 4 +-
6 files changed, 204 insertions(+), 29 deletions(-)
--
2.8.0-rc0-114-g0b3e5e5
From: Junio C Hamano <hidden> Date: 2016-06-15 23:08:36
The bundle v3 format introduces an ability to have the bundle header
(which describes what references in the bundled history can be
fetched, and what objects the receiving repository must have in
order to unbundle it successfully) in one file, and the bundled pack
stream data in a separate file.
A v3 bundle file begins with a line with "# v3 git bundle", followed
by zero or more "extended header" lines, and an empty line, finally
followed by the list of prerequisites and references in the same
format as v2 bundle. If it uses the "split bundle" feature, there
is a "data: $URL" extended header line, and nothing follows the list
of prerequisites and references. Also, "sha1: " extended header
line may exist to help validating that the pack stream data matches
the bundle header.
A typical expected use of a split bundle is to help initial clone
that involves a huge data transfer, and would go like this:
- Any repository people would clone and fetch from would regularly
be repacked, and it is expected that there would be a packfile
without prerequisites that holds all (or at least most) of the
history of it (call it pack-$name.pack).
- After arranging that packfile to be downloadable over popular
transfer methods used for serving static files (such as HTTP or
HTTPS) that are easily resumable as $URL/pack-$name.pack, a v3
bundle file (call it $name.bndl) can be prepared with an extended
header "data: $URL/pack-$name.pack" to point at the download
location for the packfile, and be served at "$URL/$name.bndl".
- An updated Git client, when trying to "git clone" from such a
repository, may be redirected to $URL/$name.bndl", which would be
a tiny text file (when split bundle feature is used).
- The client would then inspect the downloaded $name.bndl, learn
that the corresponding packfile exists at $URL/pack-$name.pack,
and downloads it as pack-$name.pack, until the download succeeds.
This can easily be done with "wget --continue" equivalent over an
unreliable link. The checksum recorded on the "sha1: " header
line is expected to be used by this downloader (not written yet).
- After fully downloading $name.bndl and pack-$name.pack and
storing them next to each other, the client would clone from the
$name.bndl; this would populate the newly created repository with
reasonably recent history.
- Then the client can issue "git fetch" against the original
repository to obtain the most recent part of the history created
since the bundle was made.
Signed-off-by: Junio C Hamano <redacted>
---
bundle.c | 103 +++++++++++++++++++++++++++++++++++++++++++++++++-----
bundle.h | 3 ++
t/t5704-bundle.sh | 64 +++++++++++++++++++++++++++++++++
3 files changed, 161 insertions(+), 9 deletions(-)
@@ -33,16 +34,55 @@ static int parse_bundle_header(int fd, struct bundle_header *header, int quiet)intstatus=0;/* The bundle header begins with the signature */-if(strbuf_getwholeline_fd(&buf,fd,'\n')||-strcmp(buf.buf,bundle_signature)){+if(strbuf_getwholeline_fd(&buf,fd,'\n')){+bad_bundle:if(!quiet)-error(_("'%s' does not look like a v2 bundle file"),+error(_("'%s' does not look like a supported bundle file"),header->filename);status=-1;gotoabort;}-/* The bundle header ends with an empty line */+if(!strcmp(buf.buf,bundle_signature_v2))+header->bundle_version=2;+elseif(!strcmp(buf.buf,bundle_signature_v3))+header->bundle_version=3;+else+gotobad_bundle;++if(header->bundle_version==3){+/*+*bundleversionv3hasextendedheadersbeforethe+*listofprerequisitesandreferences.Theextended+*headersendwithanemptyline.+*/+while(!strbuf_getwholeline_fd(&buf,fd,'\n')){+constchar*cp;+if(buf.len&&buf.buf[buf.len-1]=='\n')+buf.buf[--buf.len]='\0';+if(!buf.len)+break;+if(skip_prefix(buf.buf,"data: ",&cp)){+header->datafile=xstrdup(cp);+continue;+}+if(skip_prefix(buf.buf,"sha1: ",&cp)){+unsignedcharsha1[GIT_SHA1_RAWSZ];+if(get_sha1_hex(cp,sha1)||+cp[GIT_SHA1_HEXSZ])+gotobad_bundle;+hashcpy(header->csum,sha1);+continue;+}++gotobad_bundle;+}+}++/*+*Thebundleheaderlistsprerequisitesand+*references,andthelistendswithanemptyline.+*/while(!strbuf_getwholeline_fd(&buf,fd,'\n')&&buf.len&&buf.buf[0]!='\n'){unsignedcharsha1[20];
@@ -77,7 +117,8 @@ static int parse_bundle_header(int fd, struct bundle_header *header, int quiet)abort:if(status){-close(fd);+if(0<=fd)+close(fd);fd=-1;}strbuf_release(&buf);
@@ -71,4 +71,68 @@ test_expect_success 'prerequisites with an empty commit message' 'gitbundleverifybundle'+# bundle v3 (experimental)+test_expect_success'clone from v3''++# as "bundle create" does not exist yet for v3+# prepare it by hand here+head=$(gitrev-parseHEAD)&&+name=$(echo$head|gitpack-objects--revsv3)&&+test_when_finished"rm v3-$name.pack v3-$name.idx"&&+cat>v3.bndl<<-EOF&&+# v3 git bundle+data:v3-$name.pack++$headHEAD+$headrefs/heads/master+EOF++gitbundleverifyv3.bndl&&+gitbundlelist-headsv3.bndl>actual&&+cat>expect<<-EOF&&+$headHEAD+$headrefs/heads/master+EOF+test_cmpexpectactual&&++gitclonev3.bndlv3dst&&+git-Cv3dstfor-each-ref--format="%(objectname) %(refname)">actual&&+cat>expect<<-EOF&&+$headrefs/heads/master+$headrefs/remotes/origin/HEAD+$headrefs/remotes/origin/master+EOF+test_cmpexpectactual&&+git-Cv3dstfsck&&++# an "inline" v3 is still possible.+cat>v3i.bndl<<-EOF&&+# v3 git bundle++$headHEAD+$headrefs/heads/master++EOF+catv3-$name.pack>>v3i.bndl&&+test_when_finished"rm v3i.bndl"&&++gitbundleverifyv3i.bndl&&+gitbundlelist-headsv3i.bndl>actual&&+cat>expect<<-EOF&&+$headHEAD+$headrefs/heads/master+EOF+test_cmpexpectactual&&++gitclonev3i.bndlv3idst&&+git-Cv3idstfor-each-ref--format="%(objectname) %(refname)">actual&&+cat>expect<<-EOF&&+$headrefs/heads/master+$headrefs/remotes/origin/HEAD+$headrefs/remotes/origin/master+EOF+test_cmpexpectactual&&+git-Cv3idstfsck+'+ test_done
From: Junio C Hamano <hidden> Date: 2016-06-15 23:08:36
The bundle header structure holds two lists of refs and object
names, which should be released when the user is done with it.
Signed-off-by: Junio C Hamano <redacted>
---
bundle.c | 12 ++++++++++++
bundle.h | 1 +
transport.c | 1 +
3 files changed, 14 insertions(+)
@@ -102,6 +102,18 @@ int is_bundle(const char *path, int quiet)return(fd>=0);}+voidrelease_bundle_header(structbundle_header*header)+{+inti;++for(i=0;i<header->prerequisites.nr;i++)+free(header->prerequisites.list[i].name);+free(header->prerequisites.list);+for(i=0;i<header->references.nr;i++)+free(header->references.list[i].name);+free(header->references.list);+}+staticintlist_refs(structref_list*r,intargc,constchar**argv){inti;
@@ -23,5 +23,6 @@ int verify_bundle(struct bundle_header *header, int verbose);intunbundle(structbundle_header*header,intbundle_fd,intflags);intlist_bundle_refs(structbundle_header*header,intargc,constchar**argv);+voidrelease_bundle_header(structbundle_header*);#endif
@@ -107,6 +107,7 @@ static int close_bundle(struct transport *transport)structbundle_transport_data*data=transport->data;if(data->fd>0)close(data->fd);+release_bundle_header(&data->header);free(data);return0;}
From: Junio C Hamano <hidden> Date: 2016-06-15 23:08:36
This will be necessary when we start reading from a split bundle
where the header and the thin-pack data live in different files.
The in-core bundle header will read from a file that has the header,
and will record the path to that file. We would find the name of
the file that hosts the thin-pack data from the header, and we would
take that name as relative to the file we read the header from.
Helped-by: Jeff King [off-list ref]
Signed-off-by: Junio C Hamano <redacted>
---
builtin/bundle.c | 6 +++---
bundle.c | 27 +++++++++++++++++----------
bundle.h | 4 +++-
transport.c | 3 ++-
4 files changed, 25 insertions(+), 15 deletions(-)
@@ -30,9 +35,9 @@ static int parse_bundle_header(int fd, struct bundle_header *header,/* The bundle header begins with the signature */if(strbuf_getwholeline_fd(&buf,fd,'\n')||strcmp(buf.buf,bundle_signature)){-if(report_path)+if(!quiet)error(_("'%s' does not look like a v2 bundle file"),-report_path);+header->filename);status=-1;gotoabort;}
@@ -79,13 +84,13 @@ static int parse_bundle_header(int fd, struct bundle_header *header,returnfd;}-intread_bundle_header(constchar*path,structbundle_header*header)+intread_bundle_header(structbundle_header*header){-intfd=open(path,O_RDONLY);+intfd=open(header->filename,O_RDONLY);if(fd<0)-returnerror(_("could not open '%s'"),path);-returnparse_bundle_header(fd,header,path);+returnerror(_("could not open '%s'"),header->filename);+returnparse_bundle_header(fd,header,0);}intis_bundle(constchar*path,intquiet)
@@ -96,7 +101,7 @@ int is_bundle(const char *path, int quiet)if(fd<0)return0;memset(&header,0,sizeof(header));-fd=parse_bundle_header(fd,&header,quiet?NULL:path);+fd=parse_bundle_header(fd,&header,quiet);if(fd>=0)close(fd);return(fd>=0);
...this is const, even though we know it is allocated on the heap.
I am OK if we want to keep it "conceptually const" so that anybody
looking at the struct knows they should not touch it. But I am also OK
with just making it "char *".
-Peff
On Thu, Mar 3, 2016 at 3:32 AM, Junio C Hamano [off-list ref] wrote:
- After arranging that packfile to be downloadable over popular
transfer methods used for serving static files (such as HTTP or
HTTPS) that are easily resumable as $URL/pack-$name.pack, a v3
bundle file (call it $name.bndl) can be prepared with an extended
header "data: $URL/pack-$name.pack" to point at the download
location for the packfile, and be served at "$URL/$name.bndl".
Extra setup to offload things to CDN is great and all. But would it be
ok if we introduced a minimal resumable download service via
git-daemon to enable this feature with very little setup? Like
git-shell, you can only download certain packfiles for this use case
and nothing else with this service.
--
Duy
From: Christian Couder <hidden> Date: 2016-06-16 02:19:32
I am responding to this 2+ month old email because I am investigating
adding an alternate object store at the same level as loose and packed
objects. This alternate object store could be used for large files. I
am working on this for GitLab. (Yeah, I am working, as a freelance,
for both Booking.com and GitLab these days.)
On Wed, Mar 2, 2016 at 9:32 PM, Junio C Hamano [off-list ref] wrote:
The bundle v3 format introduces an ability to have the bundle header
(which describes what references in the bundled history can be
fetched, and what objects the receiving repository must have in
order to unbundle it successfully) in one file, and the bundled pack
stream data in a separate file.
A v3 bundle file begins with a line with "# v3 git bundle", followed
by zero or more "extended header" lines, and an empty line, finally
followed by the list of prerequisites and references in the same
format as v2 bundle. If it uses the "split bundle" feature, there
is a "data: $URL" extended header line, and nothing follows the list
of prerequisites and references. Also, "sha1: " extended header
line may exist to help validating that the pack stream data matches
the bundle header.
A typical expected use of a split bundle is to help initial clone
that involves a huge data transfer, and would go like this:
- Any repository people would clone and fetch from would regularly
be repacked, and it is expected that there would be a packfile
without prerequisites that holds all (or at least most) of the
history of it (call it pack-$name.pack).
- After arranging that packfile to be downloadable over popular
transfer methods used for serving static files (such as HTTP or
HTTPS) that are easily resumable as $URL/pack-$name.pack, a v3
bundle file (call it $name.bndl) can be prepared with an extended
header "data: $URL/pack-$name.pack" to point at the download
location for the packfile, and be served at "$URL/$name.bndl".
- An updated Git client, when trying to "git clone" from such a
repository, may be redirected to $URL/$name.bndl", which would be
a tiny text file (when split bundle feature is used).
- The client would then inspect the downloaded $name.bndl, learn
that the corresponding packfile exists at $URL/pack-$name.pack,
and downloads it as pack-$name.pack, until the download succeeds.
This can easily be done with "wget --continue" equivalent over an
unreliable link. The checksum recorded on the "sha1: " header
line is expected to be used by this downloader (not written yet).
I wonder if this mechanism could also be used or extended to clone and
fetch an alternate object database.
In [1], [2] and [3], and this was also discussed during the
Contributor Summit last month, Peff says that he started working on
alternate object database support a long time ago, and that the hard
part is a protocol extension to tell remotes that you can access some
objects in a different way.
If a Git client would download a "$name.bndl" v3 bundle file that
would have a "data: $URL/alt-odb-$name.odb" extended header, the Git
client would just need to download "$URL/alt-odb-$name.odb" and use
the alternate object database support on this file.
This way it would know all it has to know to access the objects in the
alternate database. The alternate object database may not contain the
real objects, if they are too big for example, but just files that
describe how to get the real objects.
- After fully downloading $name.bndl and pack-$name.pack and
storing them next to each other, the client would clone from the
$name.bndl; this would populate the newly created repository with
reasonably recent history.
- Then the client can issue "git fetch" against the original
repository to obtain the most recent part of the history created
since the bundle was made.
On Fri, May 20, 2016 at 7:39 PM, Christian Couder
[off-list ref] wrote:
I am responding to this 2+ month old email because I am investigating
adding an alternate object store at the same level as loose and packed
objects. This alternate object store could be used for large files. I
am working on this for GitLab. (Yeah, I am working, as a freelance,
for both Booking.com and GitLab these days.)
I'm also interested in this from a different angle, narrow clone that
potentially allows to skip download some large blobs (likely old ones
from the past that nobody will bother).
On Wed, Mar 2, 2016 at 9:32 PM, Junio C Hamano [off-list ref] wrote:
quoted
The bundle v3 format introduces an ability to have the bundle header
(which describes what references in the bundled history can be
fetched, and what objects the receiving repository must have in
order to unbundle it successfully) in one file, and the bundled pack
stream data in a separate file.
A v3 bundle file begins with a line with "# v3 git bundle", followed
by zero or more "extended header" lines, and an empty line, finally
followed by the list of prerequisites and references in the same
format as v2 bundle. If it uses the "split bundle" feature, there
is a "data: $URL" extended header line, and nothing follows the list
of prerequisites and references. Also, "sha1: " extended header
line may exist to help validating that the pack stream data matches
the bundle header.
A typical expected use of a split bundle is to help initial clone
that involves a huge data transfer, and would go like this:
- Any repository people would clone and fetch from would regularly
be repacked, and it is expected that there would be a packfile
without prerequisites that holds all (or at least most) of the
history of it (call it pack-$name.pack).
- After arranging that packfile to be downloadable over popular
transfer methods used for serving static files (such as HTTP or
HTTPS) that are easily resumable as $URL/pack-$name.pack, a v3
bundle file (call it $name.bndl) can be prepared with an extended
header "data: $URL/pack-$name.pack" to point at the download
location for the packfile, and be served at "$URL/$name.bndl".
- An updated Git client, when trying to "git clone" from such a
repository, may be redirected to $URL/$name.bndl", which would be
a tiny text file (when split bundle feature is used).
- The client would then inspect the downloaded $name.bndl, learn
that the corresponding packfile exists at $URL/pack-$name.pack,
and downloads it as pack-$name.pack, until the download succeeds.
This can easily be done with "wget --continue" equivalent over an
unreliable link. The checksum recorded on the "sha1: " header
line is expected to be used by this downloader (not written yet).
I wonder if this mechanism could also be used or extended to clone and
fetch an alternate object database.
In [1], [2] and [3], and this was also discussed during the
Contributor Summit last month, Peff says that he started working on
alternate object database support a long time ago, and that the hard
part is a protocol extension to tell remotes that you can access some
objects in a different way.
If a Git client would download a "$name.bndl" v3 bundle file that
would have a "data: $URL/alt-odb-$name.odb" extended header, the Git
client would just need to download "$URL/alt-odb-$name.odb" and use
the alternate object database support on this file.
What does this file contain exactly? A list of SHA-1 that can be
retrieved from this remote/alternate odb? I wonder if we could just
git-replace for this marking. The replaced content could contain the
uri pointing to the alt odb. We could optionally contact alt odb to
retrieve real content, or just show the replaced/fake data when alt
odb is out of reach. Transferring git-replace is basically ref
exchange, which may be fine if you don't have a lot of objects in this
alt odb. If you do, well, we need to deal with lots of refs anyway.
This may benefit from it too.
From: Christian Couder <hidden> Date: 2016-06-16 02:19:40
On Tue, May 31, 2016 at 2:43 PM, Duy Nguyen [off-list ref] wrote:
On Fri, May 20, 2016 at 7:39 PM, Christian Couder
[off-list ref] wrote:
quoted
I am responding to this 2+ month old email because I am investigating
adding an alternate object store at the same level as loose and packed
objects. This alternate object store could be used for large files. I
am working on this for GitLab. (Yeah, I am working, as a freelance,
for both Booking.com and GitLab these days.)
I'm also interested in this from a different angle, narrow clone that
potentially allows to skip download some large blobs (likely old ones
from the past that nobody will bother).
Interesting!
[...]
quoted
I wonder if this mechanism could also be used or extended to clone and
fetch an alternate object database.
In [1], [2] and [3], and this was also discussed during the
Contributor Summit last month, Peff says that he started working on
alternate object database support a long time ago, and that the hard
part is a protocol extension to tell remotes that you can access some
objects in a different way.
If a Git client would download a "$name.bndl" v3 bundle file that
would have a "data: $URL/alt-odb-$name.odb" extended header, the Git
client would just need to download "$URL/alt-odb-$name.odb" and use
the alternate object database support on this file.
What does this file contain exactly? A list of SHA-1 that can be
retrieved from this remote/alternate odb?
It would depend on the external odb. Git could support different
external odb that have different trade-offs.
I wonder if we could just
git-replace for this marking. The replaced content could contain the
uri pointing to the alt odb.
Yeah, interesting!
That's indeed another possibility that might not need the transfer of
any external odb.
But in this case it might be cleaner to just have a separate ref hierarchy like:
refs/external-odbs/my-ext-odb/<sha1>
instead of using the replace one.
Or maybe:
refs/replace/external-odbs/my-ext-odb/<sha1>
if we really want to use the replace hierarchy.
We could optionally contact alt odb to
retrieve real content, or just show the replaced/fake data when alt
odb is out of reach.
Yeah, I wonder if that really needs the replace mechanism.
Transferring git-replace is basically ref
exchange, which may be fine if you don't have a lot of objects in this
alt odb.
Yeah sure, great idea!
By the way this makes me wonder if we could implement resumable clone
using some kind of replace ref.
The client while cloning nearly as usual would download one or more
special replace refs that would points to objects with links to
download bundles using standard protocols.
Just after the clone, the client would read these objects and download
the bundles from these objects.
And then it would clone from these bundles.
If you do, well, we need to deal with lots of refs anyway.
This may benefit from it too.
I have rebased, fixed and improved it a bit. I added write support for
blobs. But the result is not very clean right now.
I was going to send a RFC patch series after cleaning the result, but
as you ask, here are some links to some branches:
- https://github.com/chriscool/git/commits/gl-external-odb3 (the
updated patches from Peff, plus 2 small patches from me)
- https://github.com/chriscool/git/commits/gl-external-odb7 (the same
as above, plus a number of WIP patches to add blob write support)
Thanks,
Christian.
It's now "jk/external-odb-wip" at the same repo. I wouldn't be surprised
if it doesn't even compile, though. I basically rebase my topics daily
against Junio's "master", so it may be carried forward, but things
marked "-wip" aren't part of my daily git build, and generally don't
even get compile-tested (usually if the rebase looks too hairy or awful,
I'll drop it completely, though, and I haven't done that here).
You're probably better off looking whatever Christian produces. :)
-Peff
From: Jeff King <hidden> Date: 2016-06-16 02:19:40
On Fri, May 20, 2016 at 02:39:06PM +0200, Christian Couder wrote:
I wonder if this mechanism could also be used or extended to clone and
fetch an alternate object database.
In [1], [2] and [3], and this was also discussed during the
Contributor Summit last month, Peff says that he started working on
alternate object database support a long time ago, and that the hard
part is a protocol extension to tell remotes that you can access some
objects in a different way.
If a Git client would download a "$name.bndl" v3 bundle file that
would have a "data: $URL/alt-odb-$name.odb" extended header, the Git
client would just need to download "$URL/alt-odb-$name.odb" and use
the alternate object database support on this file.
This way it would know all it has to know to access the objects in the
alternate database. The alternate object database may not contain the
real objects, if they are too big for example, but just files that
describe how to get the real objects.
I'm not sure about this strategy. I see two complications:
1. I don't think bundles need to be a part of this "external odb"
strategy at all. If I understand correctly, I think you want to use
it as a place to stuff metadata that the server tells the client,
like "by the way, go here if you want another way to access some
objects".
But there are lots of cases where the server might want to tell
the client that don't involve bundles at all.
2. A server pointing the client to another object store is actually
the least interesting bit of the protocol.
The more interesting cases (to me) are:
a. The receiving side of a connection (e.g., a fetch client)
somehow has out-of-band access to some objects. How does it
tell the other side "do not bother sending me these objects; I
can get them in another way"?
b. The receiving side of a connection has out-of-band access to
some objects. Some of these will be expensive to get (e.g.,
requiring a large download), and some may be fast (e.g.,
they've already been fetched to a local cache). How do we tell
the sending side not to assume we have cheap access to these
objects (e.g., for use as a delta base)?
So I don't think you want to tie this into bundles due to (1), and I
think that bundles would be insufficient anyway because of (2).
Or maybe I'm misunderstanding what you propose.
-Peff
On Tue, May 31, 2016 at 8:18 PM, Christian Couder
[off-list ref] wrote:
quoted
quoted
I wonder if this mechanism could also be used or extended to clone and
fetch an alternate object database.
In [1], [2] and [3], and this was also discussed during the
Contributor Summit last month, Peff says that he started working on
alternate object database support a long time ago, and that the hard
part is a protocol extension to tell remotes that you can access some
objects in a different way.
If a Git client would download a "$name.bndl" v3 bundle file that
would have a "data: $URL/alt-odb-$name.odb" extended header, the Git
client would just need to download "$URL/alt-odb-$name.odb" and use
the alternate object database support on this file.
What does this file contain exactly? A list of SHA-1 that can be
retrieved from this remote/alternate odb?
It would depend on the external odb. Git could support different
external odb that have different trade-offs.
quoted
I wonder if we could just
git-replace for this marking. The replaced content could contain the
uri pointing to the alt odb.
Yeah, interesting!
That's indeed another possibility that might not need the transfer of
any external odb.
But in this case it might be cleaner to just have a separate ref hierarchy like:
refs/external-odbs/my-ext-odb/<sha1>
instead of using the replace one.
Or maybe:
refs/replace/external-odbs/my-ext-odb/<sha1>
if we really want to use the replace hierarchy.
Yep. replace hierarchy crossed my mind. But then I thought about
performance degradation when there are more than one pack (we have to
search through them all for every SHA-1) and discarded it because we
would need to do the same linear search here. I guess we will most
likely have one or two name spaces so it probably won't matter.
quoted
We could optionally contact alt odb to
retrieve real content, or just show the replaced/fake data when alt
odb is out of reach.
Yeah, I wonder if that really needs the replace mechanism.
Replace mechanism provides good hook point. But it really depends how
invasive this remote odb is. If a fake content is enough to avoid
breakages up high, git-replace is enough. If you really need to pass
remote odb info up so higher levels can do something more fancy, then
it's insufficient.
By the way this makes me wonder if we could implement resumable clone
using some kind of replace ref.
The client while cloning nearly as usual would download one or more
special replace refs that would points to objects with links to
download bundles using standard protocols.
Just after the clone, the client would read these objects and download
the bundles from these objects.
And then it would clone from these bundles.
I thought we have settled on resumable clone, just waiting for an
implementation :) Doing it your way, you would need to download these
special objects too (in a pack?) and come back download some more
bundles. It would be more efficient to show the bundle uri early and
go download the bundle on the side while you go on to get the
addition/smaller pack that contains the rest.
--
Duy
I have rebased, fixed and improved it a bit. I added write support for
blobs. But the result is not very clean right now.
I was going to send a RFC patch series after cleaning the result, but
as you ask, here are some links to some branches:
- https://github.com/chriscool/git/commits/gl-external-odb3 (the
updated patches from Peff, plus 2 small patches from me)
- https://github.com/chriscool/git/commits/gl-external-odb7 (the same
as above, plus a number of WIP patches to add blob write support)
Thanks. I had a super quick look. It would be nice if you could give a
high level overview on this (if you're going to spend a lot more time on it).
One random thought, maybe it's better to have a daemon for external
odb right from the start (one for all odbs, or one per odb, I don't
know). It could do fancy stuff like object caching if necessary, and
it can avoid high cost handshake (e.g. via tls) every time a git
process runs and gets one object. Reducing process spawn would
definitely receive a big cheer from Windows crowd.
Any thought on object streaming support? It could be a big deal (might
affect some design decisions). I would also think about how pack v4
fits in this (e.g. how a tree walker can still walk fast, a big
promise of pack v4; I suppose if you still maintain "pack" concept
over external odb then it might work). Not that it really matters.
Pack v4 is the future, but the future can never be "today" :)
--
Duy
I have rebased, fixed and improved it a bit. I added write support for
blobs. But the result is not very clean right now.
I was going to send a RFC patch series after cleaning the result, but
as you ask, here are some links to some branches:
- https://github.com/chriscool/git/commits/gl-external-odb3 (the
updated patches from Peff, plus 2 small patches from me)
- https://github.com/chriscool/git/commits/gl-external-odb7 (the same
as above, plus a number of WIP patches to add blob write support)
Thanks. I had a super quick look. It would be nice if you could give a
high level overview on this (if you're going to spend a lot more time on it).
Sorry about the late answer.
Here is a new series after some cleanup:
https://github.com/chriscool/git/commits/gl-external-odb12
The high level overview of the patch series I would like to send
really soon now could go like this:
---
Git can store its objects only in the form of loose objects in
separate files or packed objects in a pack file.
To be able to better handle some kind of objects, for example big
blobs, it would be nice if Git could store its objects in other object
databases (ODB).
To do that, this patch series makes it possible to register commands,
using "odb.<odbname>.command" config variables, to access external
ODBs. Each specified command will then be called the following ways:
- "<command> have": the command should output the sha1, size and
type of all the objects the external ODB contains, one object per
line.
- "<command> get <sha1>": the command should then read from the
external ODB the content of the object corresponding to <sha1> and
output it on stdout.
- "<command> put <sha1> <size> <type>": the command should then read
from stdin an object and store it in the external ODB.
This RFC patch series does not address the following important parts
of a complete solution:
- There is no way to transfer external ODB content using Git.
- No real external ODB has been interfaced with Git. The tests use
another git repo in a separate directory for this purpose which is
probably useless in the real world.
---
One random thought, maybe it's better to have a daemon for external
odb right from the start (one for all odbs, or one per odb, I don't
know). It could do fancy stuff like object caching if necessary, and
it can avoid high cost handshake (e.g. via tls) every time a git
process runs and gets one object. Reducing process spawn would
definitely receive a big cheer from Windows crowd.
The caching could be done inside Git and I am not sure it's worth
optimizing this now.
It could also make it more difficult to write support for an external
ODB if we required a daemon.
Maybe later we can add support for "odb.<odbname>.daemon" if we think
that this is worth it.
Any thought on object streaming support?
No I didn't think about this. In fact I am not sure what this means.
It could be a big deal (might
affect some design decisions).
Could you elaborate on this?
I would also think about how pack v4
fits in this (e.g. how a tree walker can still walk fast, a big
promise of pack v4; I suppose if you still maintain "pack" concept
over external odb then it might work). Not that it really matters.
Pack v4 is the future, but the future can never be "today" :)
Sorry I haven't really followed pack v4 and I forgot what it is about.
Thanks,
Christian.
From: Mike Hommey <hidden> Date: 2016-06-16 02:19:46
On Tue, Jun 07, 2016 at 10:46:07AM +0200, Christian Couder wrote:
The high level overview of the patch series I would like to send
really soon now could go like this:
---
Git can store its objects only in the form of loose objects in
separate files or packed objects in a pack file.
To be able to better handle some kind of objects, for example big
blobs, it would be nice if Git could store its objects in other object
databases (ODB).
To do that, this patch series makes it possible to register commands,
using "odb.<odbname>.command" config variables, to access external
ODBs. Each specified command will then be called the following ways:
- "<command> have": the command should output the sha1, size and
type of all the objects the external ODB contains, one object per
line.
- "<command> get <sha1>": the command should then read from the
external ODB the content of the object corresponding to <sha1> and
output it on stdout.
- "<command> put <sha1> <size> <type>": the command should then read
from stdin an object and store it in the external ODB.
(disclaimer: I didn't look at the patch series)
Does this mean you're going to fork/exec() a new <command> for each of
these? It would probably be better if it was "batched", where the
executable is invoked once and the commands are passed to its stdin.
Mike
On Tue, Jun 7, 2016 at 3:46 PM, Christian Couder
[off-list ref] wrote:
quoted
Any thought on object streaming support?
No I didn't think about this. In fact I am not sure what this means.
quoted
It could be a big deal (might
affect some design decisions).
Could you elaborate on this?
Object streaming api is in streaming.h. Normally objects are small and
we can inflate the whole thing in memory before doing anything with
them. For really large objects (which I guess is one of the reasons
for remote odb) we don't want to do that. It takes lots of memory and
you could have objects larger than your physical memory. In some cases
when can ignore those objects (e.g. mark them binary and choose not to
diff). In some other cases (e.g. checkout), we use streaming interface
to process an object while we're inflating it to keep memory usage
down. It's easy to add a new streaming backend, once you settle on how
remote odb streams stuff.
quoted
I would also think about how pack v4
fits in this (e.g. how a tree walker can still walk fast, a big
promise of pack v4; I suppose if you still maintain "pack" concept
over external odb then it might work). Not that it really matters.
Pack v4 is the future, but the future can never be "today" :)
Sorry I haven't really followed pack v4 and I forgot what it is about.
It's a new pack format (and practically vaporware at this point) that
promises much faster access when you need to walk through trees and
commits (think rev-list --objects --all, or git-blame). Because we are
(or I am) still not sure if pack v4 will ever get to the state where
it can be merged to git.git, I think it's ok for you to ignore it too
if you want. You can read more about the format here [1] and go even
further back to [2] when Nicolas teased us with the pack size
(smaller, which is a nice side effect). The potential issue with pack
v4 is, the tree walker (struct tree_desc and related funcs in
walk-tree.h) needs to know about pack v4 in order to walk fast.
Current tree walker does not care if an object is packed (using what
format) at all. Remote odb for pack v4 must have some way that allows
to read pack data directly, something close to "mmap", it's not just
about an api to "get me the canonical content of this object".
[1] http://article.gmane.org/gmane.comp.version-control.git/234012
[2] http://article.gmane.org/gmane.comp.version-control.git/233038
--
Duy
From: Christian Couder <hidden> Date: 2016-06-16 02:19:47
On Wed, Jun 1, 2016 at 12:31 AM, Jeff King [off-list ref] wrote:
On Fri, May 20, 2016 at 02:39:06PM +0200, Christian Couder wrote:
quoted
I wonder if this mechanism could also be used or extended to clone and
fetch an alternate object database.
In [1], [2] and [3], and this was also discussed during the
Contributor Summit last month, Peff says that he started working on
alternate object database support a long time ago, and that the hard
part is a protocol extension to tell remotes that you can access some
objects in a different way.
If a Git client would download a "$name.bndl" v3 bundle file that
would have a "data: $URL/alt-odb-$name.odb" extended header, the Git
client would just need to download "$URL/alt-odb-$name.odb" and use
the alternate object database support on this file.
This way it would know all it has to know to access the objects in the
alternate database. The alternate object database may not contain the
real objects, if they are too big for example, but just files that
describe how to get the real objects.
I'm not sure about this strategy.
I am also not sure that this is the best strategy, but I think it's
worth discussing.
I see two complications:
1. I don't think bundles need to be a part of this "external odb"
strategy at all. If I understand correctly, I think you want to use
it as a place to stuff metadata that the server tells the client,
like "by the way, go here if you want another way to access some
objects".
Yeah, basically I think it might be possible to use the bundle
mechanism to transfer what an external ODB on the client would need to
be initialized or updated.
But there are lots of cases where the server might want to tell
the client that don't involve bundles at all.
The idea is also that anytime the server needs to send external ODB
data to the client, it would ask its own external ODB to prepare a
kind of bundle with that data and use the bundle v3 mechanism to send
it.
That may need the bundle v3 mechanism to be extended, but I don't see
in which cases it would not work.
2. A server pointing the client to another object store is actually
the least interesting bit of the protocol.
The more interesting cases (to me) are:
a. The receiving side of a connection (e.g., a fetch client)
somehow has out-of-band access to some objects. How does it
tell the other side "do not bother sending me these objects; I
can get them in another way"?
I don't see a difference with regular objects that the fetch client
already has. If it already has some regular objects, a way to tell the
server "don't bother sending me these objects" is useful already and
it should be possible to use it to tell the server that there is no
need to send some objects stored in the external ODB too.
Also something like this is needed for shallow clones and narrow clones anyway.
b. The receiving side of a connection has out-of-band access to
some objects. Some of these will be expensive to get (e.g.,
requiring a large download), and some may be fast (e.g.,
they've already been fetched to a local cache). How do we tell
the sending side not to assume we have cheap access to these
objects (e.g., for use as a delta base)?
I don't think we need to tell the sending side we have cheap access or
not to some objects.
If the objects are managed by the external ODB, it's the external ODB
on the server and on the client that will manage these objects. They
should not be used as delta bases.
Perhaps there is no mechanism to say that some objects (basically all
external ODB managed objects) should not be used as delta bases, but
that could be added.
Thanks,
Christian.
From: Christian Couder <hidden> Date: 2016-06-16 02:19:47
On Wed, Jun 1, 2016 at 3:37 PM, Duy Nguyen [off-list ref] wrote:
On Tue, May 31, 2016 at 8:18 PM, Christian Couder
[off-list ref] wrote:
quoted
quoted
quoted
I wonder if this mechanism could also be used or extended to clone and
fetch an alternate object database.
In [1], [2] and [3], and this was also discussed during the
Contributor Summit last month, Peff says that he started working on
alternate object database support a long time ago, and that the hard
part is a protocol extension to tell remotes that you can access some
objects in a different way.
If a Git client would download a "$name.bndl" v3 bundle file that
would have a "data: $URL/alt-odb-$name.odb" extended header, the Git
client would just need to download "$URL/alt-odb-$name.odb" and use
the alternate object database support on this file.
What does this file contain exactly? A list of SHA-1 that can be
retrieved from this remote/alternate odb?
It would depend on the external odb. Git could support different
external odb that have different trade-offs.
quoted
I wonder if we could just
git-replace for this marking. The replaced content could contain the
uri pointing to the alt odb.
Yeah, interesting!
That's indeed another possibility that might not need the transfer of
any external odb.
But in this case it might be cleaner to just have a separate ref hierarchy like:
refs/external-odbs/my-ext-odb/<sha1>
instead of using the replace one.
Or maybe:
refs/replace/external-odbs/my-ext-odb/<sha1>
if we really want to use the replace hierarchy.
Yep. replace hierarchy crossed my mind. But then I thought about
performance degradation when there are more than one pack (we have to
search through them all for every SHA-1) and discarded it because we
would need to do the same linear search here. I guess we will most
likely have one or two name spaces so it probably won't matter.
Yeah.
quoted
quoted
We could optionally contact alt odb to
retrieve real content, or just show the replaced/fake data when alt
odb is out of reach.
Yeah, I wonder if that really needs the replace mechanism.
Replace mechanism provides good hook point. But it really depends how
invasive this remote odb is. If a fake content is enough to avoid
breakages up high, git-replace is enough. If you really need to pass
remote odb info up so higher levels can do something more fancy, then
it's insufficient.
quoted
By the way this makes me wonder if we could implement resumable clone
using some kind of replace ref.
The client while cloning nearly as usual would download one or more
special replace refs that would points to objects with links to
download bundles using standard protocols.
Just after the clone, the client would read these objects and download
the bundles from these objects.
And then it would clone from these bundles.
I thought we have settled on resumable clone, just waiting for an
implementation :) Doing it your way, you would need to download these
special objects too (in a pack?) and come back download some more
bundles. It would be more efficient to show the bundle uri early and
go download the bundle on the side while you go on to get the
addition/smaller pack that contains the rest.
Yeah, something like the bundle v3 mechanism is probably more efficient.
Thanks,
Christian.
From: Jeff King <hidden> Date: 2016-06-16 02:19:47
On Tue, Jun 07, 2016 at 03:19:46PM +0200, Christian Couder wrote:
quoted
But there are lots of cases where the server might want to tell
the client that don't involve bundles at all.
The idea is also that anytime the server needs to send external ODB
data to the client, it would ask its own external ODB to prepare a
kind of bundle with that data and use the bundle v3 mechanism to send
it.
That may need the bundle v3 mechanism to be extended, but I don't see
in which cases it would not work.
Ah, I see we do not have the same underlying mental model.
I think the external odb is purely the _client's_ business. The server
does not have to have an external odb at all, and does not need to know
about the client's. The client is responsible for telling the server
during the git protocol anything it would need to know (like "do not
bother sending objects over 50MB; I can get them elsewhere").
This makes the problem much more complicated, but it is more flexible
and decentralized.
quoted
a. The receiving side of a connection (e.g., a fetch client)
somehow has out-of-band access to some objects. How does it
tell the other side "do not bother sending me these objects; I
can get them in another way"?
I don't see a difference with regular objects that the fetch client
already has. If it already has some regular objects, a way to tell the
server "don't bother sending me these objects" is useful already and
it should be possible to use it to tell the server that there is no
need to send some objects stored in the external ODB too.
The way to do that with normal objects is by finding shared commit tips,
and assuming the normal git repository property of "if you have X, you
have all of the objects reachable from X".
This whole idea is essentially creating "holes" in that property. You
can enumerate all of the holes, but I am not sure that scales well. We
get a lot of efficiency by communicating only ref tips during the
negotiation, and not individual object names.
Also something like this is needed for shallow clones and narrow
clones anyway.
Yes, and I don't think it scales well there, either. A single shallow
cutoff works OK. But if you repeatedly shallow-fetch into a repository,
you end up with a patchwork of disconnected "islands" of history. The
CPU required on the server side to serve those fetch requests is much
greater than what would normally be needed. You can't use things like
reachability bitmaps, and you have to open up the trees for each island
to see which objects the other side actually has.
quoted
b. The receiving side of a connection has out-of-band access to
some objects. Some of these will be expensive to get (e.g.,
requiring a large download), and some may be fast (e.g.,
they've already been fetched to a local cache). How do we tell
the sending side not to assume we have cheap access to these
objects (e.g., for use as a delta base)?
I don't think we need to tell the sending side we have cheap access or
not to some objects.
If the objects are managed by the external ODB, it's the external ODB
on the server and on the client that will manage these objects. They
should not be used as delta bases.
Perhaps there is no mechanism to say that some objects (basically all
external ODB managed objects) should not be used as delta bases, but
that could be added.
Yes, I agree that _if_ the server can access the list of objects
available in the external odb, this becomes much easier. I'm just not
convinced that level of coupling is a good idea.
Note that the server would also want to take this into account during
repacking, as otherwise you end up with fetches that are very expensive
to serve (you want to send X which is a delta based on Y, but you know
that Y is available via the external odb, and therefore should not be
used as a base. So you have to throw out the delta for X and either send
it whole or compute a new one. That's much more expensive than blitting
the delta from disk, which is what a normal clone would do).
-Peff