Re: [PATCH v2 00/14] Sparse checkout

8 messages, 4 authors, 2016-06-15 · open the first message on its own page

Re: [PATCH v2 00/14] Sparse checkout

From: Junio C Hamano <hidden>
Date: 2016-06-15 22:45:23

"Nguyen Thai Ngoc Duy" [off-list ref] writes:
On 9/21/08, Jakub Narebski [off-list ref] wrote:
...
quoted
 >>  BTW I think that the same rules are used in gitattributes, aren't
 >>  they?
 >
 > They have different implementations. Though the rules may be the same.

Were you able to reuse either one?
No. .gitignore is tied to read_directory() while .gitattributes has
attributes attached. So I rolled out another one for index.
I am sorry, but that sounds like a rather lame excuse.  It certainly is
possible to introduce an "ignored" attribute and have .gitattributes file
specify that, instead of having an entry in .gitignore file, if you teach
read_directory() to pay attention to the attributes mechanism.  If we had
from day one that a more generic gitattributes mechanism, I would imagine
we wouldn't even had a separate .gitignore codepath but used the attribute
mechanism throughout the system.

Now I do not think we are ever going to deprecate gitignore and move
everybody to "ignored" attributes, because such a transition would not buy
the end users anything, but it technically is possible and would have been
the right thing to do, if we were building the system from scratch.  We
still could add it as an optional feature (i.e. if a path has the
attribute that says "ignored" or "not ignored", then that determines the
fate of the path, otherwise we look at gitignore).

I wouldn't be surprised if an alternative implementation of your code to
assign "sparseness" to each path internally used "to-be-checked-out"
attribute, and used that attribute to control how ls-files filters its
output.

A better excuse might have been that "I am not reading these patterns from
anywhere but command line", but that got me thinking further.

How would that --narrow-match that is not stored anywhere on the
filesystem but used only for filtering the output be any more useful than
a grep that filters ls-files output in practice?

I would imagine it would be much more useful if .git/info/attributes can
specify "checkout" attribute that is defined like this:

        `checkout`
        ^^^^^^^^^^

        This attribute controls if the path can be left not checked-out to the
        working tree.

        Unset::
                Unsetting the `checkout` marks the path not to be checked out.

        Unspecified::
                A path which does not have any `checkout` attribute specified is
                handled in no special way.

        Any value set to `checkout` is ignored, and git acts as if the
        attribute is left unspecified.

Then whenever a new path enters the index, you _could_ check with the
attribute mechanism to set the CE_NOCHECKOUT flag.  Just like an already
tracked path is not ignored even if it matches .gitignore pattern, a path
without CE_NOCHECKOUT that is in the index is checked out even if it has
checkout attribute Unset.

Hmm?

Re: [PATCH v2 00/14] Sparse checkout

From: Nguyen Thai Ngoc Duy <hidden>
Date: 2016-06-15 22:45:23

On 9/21/08, Junio C Hamano [off-list ref] wrote:
"Nguyen Thai Ngoc Duy" [off-list ref] writes:

 > On 9/21/08, Jakub Narebski [off-list ref] wrote:
quoted
...
quoted
quoted
 >>  BTW I think that the same rules are used in gitattributes, aren't
 >>  >>  they?
 >>  >
 >>  > They have different implementations. Though the rules may be the same.
 >>
 >> Were you able to reuse either one?
 >
 > No. .gitignore is tied to read_directory() while .gitattributes has
 > attributes attached. So I rolled out another one for index.


I am sorry, but that sounds like a rather lame excuse.  It certainly is
 possible to introduce an "ignored" attribute and have .gitattributes file
 specify that, instead of having an entry in .gitignore file, if you teach
 read_directory() to pay attention to the attributes mechanism.  If we had
 from day one that a more generic gitattributes mechanism, I would imagine
 we wouldn't even had a separate .gitignore codepath but used the attribute
 mechanism throughout the system.

 Now I do not think we are ever going to deprecate gitignore and move
 everybody to "ignored" attributes, because such a transition would not buy
 the end users anything, but it technically is possible and would have been
 the right thing to do, if we were building the system from scratch.  We
 still could add it as an optional feature (i.e. if a path has the
 attribute that says "ignored" or "not ignored", then that determines the
 fate of the path, otherwise we look at gitignore).

 I wouldn't be surprised if an alternative implementation of your code to
 assign "sparseness" to each path internally used "to-be-checked-out"
 attribute, and used that attribute to control how ls-files filters its
 output.

 A better excuse might have been that "I am not reading these patterns from
 anywhere but command line", but that got me thinking further.
That "from command line" piece makes a bit of difference. For example
patterns separated by colons and backslash escape, but that does not
stop it from reusing attr.c.
 How would that --narrow-match that is not stored anywhere on the
 filesystem but used only for filtering the output be any more useful than
 a grep that filters ls-files output in practice?
Well, it works exactly like 'grep' internally.
 I would imagine it would be much more useful if .git/info/attributes can
 specify "checkout" attribute that is defined like this:

        `checkout`
        ^^^^^^^^^^

        This attribute controls if the path can be left not checked-out to the
        working tree.

        Unset::
                Unsetting the `checkout` marks the path not to be checked out.

        Unspecified::
                A path which does not have any `checkout` attribute specified is
                handled in no special way.

        Any value set to `checkout` is ignored, and git acts as if the
        attribute is left unspecified.

 Then whenever a new path enters the index, you _could_ check with the
 attribute mechanism to set the CE_NOCHECKOUT flag.  Just like an already
 tracked path is not ignored even if it matches .gitignore pattern, a path
 without CE_NOCHECKOUT that is in the index is checked out even if it has
 checkout attribute Unset.

 Hmm?
Well I think people would want to save no-checkout rules eventually.
But I don't know how they want to use it. Will the saved rules be hard
restriction, that no files can be checked out outside defined areas?
Will it be to save a couple of keystrokes,   that is, instead of
typing "--reset-sparse=blah" all the time, now just "--reset-sparse"
and default rules will be applied? Your suggestion would be the third,
applying on new files only.

Anyway I will try to extend attr.c a bit to take input from command
line, then move "sparse patterns" over to use attr.c.
-- 
Duy

Re: [PATCH v2 00/14] Sparse checkout

From: Jakub Narebski <hidden>
Date: 2016-06-15 22:45:23

On Sun, 21 Sep 2008, Nguyen Thai Ngoc Duy wrote:
On 9/21/08, Junio C Hamano [off-list ref] wrote:
quoted
 How would that --narrow-match that is not stored anywhere on the
 filesystem but used only for filtering the output be any more useful than
 a grep that filters ls-files output in practice?
Well, it works exactly like 'grep' internally.
quoted
 I would imagine it would be much more useful if .git/info/attributes can
 specify "checkout" attribute that is defined like this:

        `checkout`
        ^^^^^^^^^^
[...]
quoted
 Then whenever a new path enters the index, you _could_ check with the
 attribute mechanism to set the CE_NOCHECKOUT flag.  Just like an already
 tracked path is not ignored even if it matches .gitignore pattern, a path
 without CE_NOCHECKOUT that is in the index is checked out even if it has
 checkout attribute Unset.

 Hmm?
Well I think people would want to save no-checkout rules eventually.
But I don't know how they want to use it. Will the saved rules be hard
restriction, that no files can be checked out outside defined areas?
Will it be to save a couple of keystrokes,   that is, instead of
typing "--reset-sparse=blah" all the time, now just "--reset-sparse"
and default rules will be applied? Your suggestion would be the third,
applying on new files only.

Anyway I will try to extend attr.c a bit to take input from command
line, then move "sparse patterns" over to use attr.c.
First, I think that this was Junio asking for discussion more than
for changing the design.

Second, while unifying the "check the match" part of gitignore,
gitattribute and sparse checkout would be IMVHO a good idea, I'm
not sure if trying to use/reuse attr.c literally would be a good
idea, at least not without larger surgery.  AFAIK, IIUC gitattributes
have some limitations, one of which that they are read from working
area (and there is no API for reading from tree); although this could
be enough for `checkout' attribute, which is not that different in
work from `smudge' attribute, or `crlf` attribute.

-- 
Jakub Narebski
Poland

Re: [PATCH v2 00/14] Sparse checkout

From: Nguyen Thai Ngoc Duy <hidden>
Date: 2016-06-15 22:45:23

On 9/21/08, Jakub Narebski [off-list ref] wrote:
On Sun, 21 Sep 2008, Nguyen Thai Ngoc Duy wrote:
 > On 9/21/08, Junio C Hamano [off-list ref] wrote:

quoted
quoted
 How would that --narrow-match that is not stored anywhere on the
 > >  filesystem but used only for filtering the output be any more useful than
 > >  a grep that filters ls-files output in practice?
 >
 > Well, it works exactly like 'grep' internally.
 >
 > >  I would imagine it would be much more useful if .git/info/attributes can
 > >  specify "checkout" attribute that is defined like this:
 > >
 > >         `checkout`
 > >         ^^^^^^^^^^

[...]


 > >  Then whenever a new path enters the index, you _could_ check with the
 > >  attribute mechanism to set the CE_NOCHECKOUT flag.  Just like an already
 > >  tracked path is not ignored even if it matches .gitignore pattern, a path
 > >  without CE_NOCHECKOUT that is in the index is checked out even if it has
 > >  checkout attribute Unset.
 > >
 > >  Hmm?
 >
 > Well I think people would want to save no-checkout rules eventually.
 > But I don't know how they want to use it. Will the saved rules be hard
 > restriction, that no files can be checked out outside defined areas?
 > Will it be to save a couple of keystrokes,   that is, instead of
 > typing "--reset-sparse=blah" all the time, now just "--reset-sparse"
 > and default rules will be applied? Your suggestion would be the third,
 > applying on new files only.
 >
 > Anyway I will try to extend attr.c a bit to take input from command
 > line, then move "sparse patterns" over to use attr.c.


First, I think that this was Junio asking for discussion more than
 for changing the design.
I just tried to see if it was feasible. Checking the source again, I
misunderstood  gitattributes/gitingore's leading '/' notion (in a good
way). Leading '/' means './' and that would be fine for
.git{attributes,ignore}. In sparse patterns, leading '/' means
toplevel directory because you may want to checkout some more from a
subdirectory without moving up to toplevel directory. Now
.git{ignore,attributes} and sparse patterns are incompatible, gaah...
 Second, while unifying the "check the match" part of gitignore,
 gitattribute and sparse checkout would be IMVHO a good idea, I'm
It is surely good. Optimization like 68492fc (Speedup scanning for
excluded files.) could be applied to .gitattributes too. Now I know
why I was confused when reading the matching part of
.git{attributes,ignore}.
-- 
Duy

Re: [PATCH v2 00/14] Sparse checkout

From: Jakub Narebski <hidden>
Date: 2016-06-15 22:45:23

On Sun, 21 Sep 2008, Nguyen Thai Ngoc Duy wrote:
On 9/21/08, Jakub Narebski [off-list ref] wrote:
quoted
On Sun, 21 Sep 2008, Nguyen Thai Ngoc Duy wrote:
quoted
On 9/21/08, Junio C Hamano [off-list ref] wrote:
[...] Checking the source again, I
misunderstood  gitattributes/gitingore's leading '/' notion (in a good
way). Leading '/' means './' and that would be fine for
.git{attributes,ignore}. 
By the way it would be nice if gitignore accepted './' as equivalent
to current '/', as this is something I think (from questions here
and on #git) that people expect to work (not reading documentation
carefully enough).  This is something that for example `ls' would use,
or something that `find' returns.
In sparse patterns, leading '/' means toplevel directory because you
may want to checkout some more from a subdirectory without moving up
to toplevel directory. Now .git{ignore,attributes} and sparse patterns
are incompatible, gaah... 
Well, this doesn't make sense in a _file_, but makes perfect sense when
invoked from _command line_, as option argument.

But I was thinking more about centralizing pattern matching wrt either
full pathname (with prefix stripped, or not), or basename of a file.
If match check is centralized, then if you enhance pattern language (for
selecting which files to mark no-checkout in sparse checkout for example
by allowing '**' which matches also '/' (if you don't go route of 'tar'
with '--wildcards-match-slash' option)), then it would enhance gitignore
patterns and gitattributes patterns too (well, excluding the fact that
they are delimited differently).
 
quoted
 Second, while unifying the "check the match" part of gitignore,
 gitattribute and sparse checkout would be IMVHO a good idea, [...]
It is surely good. Optimization like 68492fc (Speedup scanning for
excluded files.) could be applied to .gitattributes too. Now I know
why I was confused when reading the matching part of
.git{attributes,ignore}.
And all speedups (well, perhaps not all) would apply to all classes
of matching against patterns as well.
-- 
Jakub Narebski
Poland

Re: [PATCH v2 00/14] Sparse checkout

From: Santi Béjar <hidden>
Date: 2016-06-15 22:45:23

On Sun, Sep 21, 2008 at 12:11 PM, Nguyen Thai Ngoc Duy
[off-list ref] wrote:
On 9/21/08, Junio C Hamano [off-list ref] wrote:
quoted
"Nguyen Thai Ngoc Duy" [off-list ref] writes:

 > On 9/21/08, Jakub Narebski [off-list ref] wrote:
quoted
...
quoted
quoted
 >>  BTW I think that the same rules are used in gitattributes, aren't
 >>  >>  they?
 >>  >
 >>  > They have different implementations. Though the rules may be the same.
 >>
 >> Were you able to reuse either one?
 >
 > No. .gitignore is tied to read_directory() while .gitattributes has
 > attributes attached. So I rolled out another one for index.


I am sorry, but that sounds like a rather lame excuse.  It certainly is
 possible to introduce an "ignored" attribute and have .gitattributes file
 specify that, instead of having an entry in .gitignore file, if you teach
 read_directory() to pay attention to the attributes mechanism.  If we had
 from day one that a more generic gitattributes mechanism, I would imagine
 we wouldn't even had a separate .gitignore codepath but used the attribute
 mechanism throughout the system.

 Now I do not think we are ever going to deprecate gitignore and move
 everybody to "ignored" attributes, because such a transition would not buy
 the end users anything, but it technically is possible and would have been
 the right thing to do, if we were building the system from scratch.  We
 still could add it as an optional feature (i.e. if a path has the
 attribute that says "ignored" or "not ignored", then that determines the
 fate of the path, otherwise we look at gitignore).

 I wouldn't be surprised if an alternative implementation of your code to
 assign "sparseness" to each path internally used "to-be-checked-out"
 attribute, and used that attribute to control how ls-files filters its
 output.

 A better excuse might have been that "I am not reading these patterns from
 anywhere but command line", but that got me thinking further.
That "from command line" piece makes a bit of difference. For example
patterns separated by colons and backslash escape, but that does not
stop it from reusing attr.c.
quoted
 How would that --narrow-match that is not stored anywhere on the
 filesystem but used only for filtering the output be any more useful than
 a grep that filters ls-files output in practice?
Well, it works exactly like 'grep' internally.
quoted
 I would imagine it would be much more useful if .git/info/attributes can
 specify "checkout" attribute that is defined like this:

        `checkout`
        ^^^^^^^^^^

        This attribute controls if the path can be left not checked-out to the
        working tree.

        Unset::
                Unsetting the `checkout` marks the path not to be checked out.

        Unspecified::
                A path which does not have any `checkout` attribute specified is
                handled in no special way.

        Any value set to `checkout` is ignored, and git acts as if the
        attribute is left unspecified.

 Then whenever a new path enters the index, you _could_ check with the
 attribute mechanism to set the CE_NOCHECKOUT flag.  Just like an already
 tracked path is not ignored even if it matches .gitignore pattern, a path
 without CE_NOCHECKOUT that is in the index is checked out even if it has
 checkout attribute Unset.

 Hmm?
Well I think people would want to save no-checkout rules eventually.
But I don't know how they want to use it. Will the saved rules be hard
restriction, that no files can be checked out outside defined areas?
Will it be to save a couple of keystrokes,   that is, instead of
typing "--reset-sparse=blah" all the time, now just "--reset-sparse"
and default rules will be applied? Your suggestion would be the third,
applying on new files only.

Anyway I will try to extend attr.c a bit to take input from command
line, then move "sparse patterns" over to use attr.c.

While I agree that the checkout attr looks like an attribute (so
reusing attr.c is a good idea) and $GIT_DIR/info/gitattributes seems a
good place to specify them, I think it will be better in the config
$GIT_DIR/config. There it is clear that it is a local thing and you
have "git config" to read and write them. Additionally you could have
different patterns in the config (sparse.default, sparse.doc,
sparse.src,...), although maybe it is not very useful.

I think the main UI to sparse checkout should be a default sparse
pattern that is used for "all" commands, like merge, reset, and
checkout. Now it is too easy to escape from the sparse checkout, when
you merge or checkout a branch with new files, when doing a "git reset
--hard" (when you abort a failed merge), or when doing a diff
(specially when you pull).

Santi

Re: [PATCH v2 00/14] Sparse checkout

From: Nguyen Thai Ngoc Duy <hidden>
Date: 2016-06-15 22:45:23

On Tue, Sep 23, 2008 at 01:06:47PM +0200, =?ISO-8859-1?Q?Santi_B=E9jar_ wrote:
While I agree that the checkout attr looks like an attribute (so
reusing attr.c is a good idea) and $GIT_DIR/info/gitattributes seems a
good place to specify them, I think it will be better in the config
$GIT_DIR/config. There it is clear that it is a local thing and you
have "git config" to read and write them. Additionally you could have
different patterns in the config (sparse.default, sparse.doc,
sparse.src,...), although maybe it is not very useful.

I think the main UI to sparse checkout should be a default sparse
pattern that is used for "all" commands, like merge, reset, and
checkout. Now it is too easy to escape from the sparse checkout, when
you merge or checkout a branch with new files, when doing a "git reset
--hard" (when you abort a failed merge), or when doing a diff
(specially when you pull).
It should not escape that easy (except newly added files). There is a
bug in my apply_narrow_spec() that effectively disables sparse
checkkout for other unpack_trees() calls except checkout/clone. Try
the below patch.
diff --git a/unpack-trees.c b/unpack-trees.c
index 10f377c..5bbe016 100644
--- a/unpack-trees.c
+++ b/unpack-trees.c
@@ -146,8 +146,10 @@ static int apply_narrow_spec(struct unpack_trees_options *o)
 	struct index_state *index = &o->result;
 	int i;
 
+	/*
 	if (!(o->new_narrow_path | o->add_narrow_path | o->remove_narrow_path))
 		return 0;
+	*/
 
 	for (i = 0; i < index->cache_nr; i++) {
 		struct cache_entry *ce = index->cache[i];
-- 
Duy

Re: [PATCH v2 00/14] Sparse checkout

From: Nguyen Thai Ngoc Duy <hidden>
Date: 2016-06-15 22:45:24

On 9/23/08, Santi Béjar [off-list ref] wrote:
While I agree that the checkout attr looks like an attribute (so
 reusing attr.c is a good idea) and $GIT_DIR/info/gitattributes seems a
 good place to specify them, I think it will be better in the config
 $GIT_DIR/config. There it is clear that it is a local thing and you
 have "git config" to read and write them. Additionally you could have
 different patterns in the config (sparse.default, sparse.doc,
 sparse.src,...), although maybe it is not very useful.

 I think the main UI to sparse checkout should be a default sparse
 pattern that is used for "all" commands, like merge, reset, and
 checkout. Now it is too easy to escape from the sparse checkout, when
 you merge or checkout a branch with new files, when doing a "git reset
 --hard" (when you abort a failed merge), or when doing a diff
 (specially when you pull).
I have made a patch to save default sparse patterns, something to play
with so we can have better idea how to do it properly.

There is another option --default-sparse in "git clone" and "git
checkout". The option can be used to save default sparse patterns
(specified by --sparse-checkout in "git clone" or --reset-sparse in
"git checkout"). Something like this:

git clone --default-sparse --sparse-checkout=Documentation/ git.git
git checkout --default-sparse --reset-sparse=t/

Default sparse patterns will be used for other unpack_trees()-related
commands like reset, read-tree, merge, pull... For "git checkout" it
will only be used when neither --full, --reset-sparse,
--include-sparse nor --exclude-sparse is present. And it only applies
to newly-added files.

Patch series is in http://repo.or.cz/w/git/pclouds.git (branch
sparse-checkout). Note that it also incorporates fixes and some option
renames.
-- 
Duy
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help