read-only working copies using links

6 messages, 3 authors, 2016-06-15 · open the first message on its own page

read-only working copies using links

From: Chad Dombrova <hidden>
Date: 2016-06-15 22:46:01

hi all,

there's a major feature for working with large binaries that has not  
yet been addressed by git:  the ability to check out a file as a  
symbolic/hard link to a blob in the repository, instead of duplicating  
the file into the working copy.

imagine a scenario where one user is putting large binary files into a  
git repo on a networked server.  100 other users on the server need  
read-only access to this repo.  they clone the repo using --shared or  
--local, which saves disk space for the object files, but each of  
these 100 working copies also creates copies of all the binary files  
at the HEAD revision. it would be 100x as efficient in both disk space  
and checkout speeds if, in place of these files, symbolic or hard  
links were made to the blob files in .git/objects.

the crux of the issue is that the blob objects would have to be stored  
as exact copies of the original files.  it would seem there are two  
things that currently prevent this from happening.  1) blobs are  
stored with compression and 2) they include a small header.   
compression can be disabled by setting core.loosecompression to 0, so  
that seems like less of an issue.  as for the header, wouldn't it be  
possible to store it separately?  in other words, store two files per  
blob directory, a small stub file with the header info and the  
unaltered file data.

what are the caveats to a system like this?  has anyone looked into  
this before?

-chad

p.s.
i tried submitting a post through nabble a few days and it said that  
it was still pending, so i thought i'd try submitting directly to the  
mailing list.  sorry, if i end up double-posting

Re: read-only working copies using links

From: Sverre Rabbelier <hidden>
Date: 2016-06-15 22:46:01

Heya,

On Sat, Jan 24, 2009 at 10:17, Chad Dombrova [off-list ref] wrote:
the crux of the issue is that the blob objects would have to be stored as
exact copies of the original files.  it would seem there are two things that
currently prevent this from happening.  1) blobs are stored with compression
and 2) they include a small header.  compression can be disabled by setting
core.loosecompression to 0, so that seems like less of an issue.  as for the
header, wouldn't it be possible to store it separately?  in other words,
store two files per blob directory, a small stub file with the header info
and the unaltered file data.
I think Tim Ansell (cced) was talking about this at the gittogether
(storing the metadata seperately), as it would benefit sparse/narrow
checkout, another advantage supporting his case?

-- 
Cheers,

Sverre Rabbelier

Re: read-only working copies using links

From: Chad Dombrova <hidden>
Date: 2016-06-15 22:46:01

I think Tim Ansell (cced) was talking about this at the gittogether
(storing the metadata seperately), as it would benefit sparse/narrow
checkout, another advantage supporting his case?
what's the case against it, other than the obvious, that it will take  
more work?


-chad

Re: read-only working copies using links

From: Sverre Rabbelier <hidden>
Date: 2016-06-15 22:46:01

On Sat, Jan 24, 2009 at 19:39, Chad Dombrova [off-list ref] wrote:
what's the case against it, other than the obvious, that it will take more
work?
Good question, I think it was mostly that, someone has to implement it
(possibly as part of packv4). Backwards compatibility is of course
always an concern, but I'm not too familiar with the subject, perhaps
other people on the list (or even those were at the gittogether) can
comment?

-- 
Cheers,

Sverre Rabbelier

Re: read-only working copies using links

From: Jeff King <hidden>
Date: 2016-06-15 22:46:01

On Sat, Jan 24, 2009 at 10:39:46AM -0800, Chad Dombrova wrote:
quoted
I think Tim Ansell (cced) was talking about this at the gittogether
(storing the metadata seperately), as it would benefit sparse/narrow
checkout, another advantage supporting his case?
what's the case against it, other than the obvious, that it will take  
more work?
I'm not sure this is actually the same as Tim's proposal. Tim wanted to
store the commit and tree information separately from the blob
information (since his use case was that blobs are enormous, but the
rest is reasonable).

AIUI, Chad's proposal is about storing the actual blob data itself
separate from the blob object's metadata (i.e., its object type and
length headers). Which means that the normal loose object format is not
acceptable, and you would end up with something like (for example):

  .git/objects/pack/pack-full-of-your-regular-stuff.{pack,idx}
  .git/objects/[0-9a-f]{2}/[0-9a-f]{38}/header
  .git/objects/[0-9a-f]{2}/[0-9a-f]{38}/data

or something similar. Then you could hardlink directly to the 'data'
portion. So you would need:

  - to teach everything that ever looks for loose objects how to read
    this new format. In theory, it's all nicely encapsulated in
    sha1_file.c

  - to teach checkout routines to hardlink such a case instead of
    copying the file

The obvious downsides that I can think of are:

  - it has the potential to make object reading, which is a core part of
    git (read: very performance- and correctness- sensitive) a lot more
    complex. But maybe the implementation would not be that painful;
    somebody would have to look very closely to see.

  - it interacts badly with smudge/clean filters and crlf conversion.
    In those cases you can't hardlink. If you treat this like an
    optimization, though, it's not so bad: we only do the optimization
    when we _can_, and fall back to regular checkout if those other
    options are in effect.

  - it's somewhat dangerous to your repository's health. Git's model is
    that object files are immutable (since they are, after all, named
    after their contents). But now you are linking them into your
    working tree, which makes them susceptible to some third party tool
    munging them. So yes, most tools will probably behave, but any tool
    that misbehaves will actually corrupt your repository.

-Peff

Re: read-only working copies using links

From: Jeff King <hidden>
Date: 2016-06-15 22:46:01

On Sat, Jan 24, 2009 at 07:43:20PM +0100, Sverre Rabbelier wrote:
On Sat, Jan 24, 2009 at 19:39, Chad Dombrova [off-list ref] wrote:
quoted
what's the case against it, other than the obvious, that it will take more
work?
Good question, I think it was mostly that, someone has to implement it
(possibly as part of packv4). Backwards compatibility is of course
always an concern, but I'm not too familiar with the subject, perhaps
other people on the list (or even those were at the gittogether) can
comment?
If I understand his proposal correctly, such objects must _not_ be part
of a pack. The whole idea is splitting them _more_, not less.

-Peff
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help