From: James Bowes <hidden> Date: 2016-06-15 22:42:59
The following two patches make git-gc a builtin command.
The first patch modifies run-command.*, making two public functions that take
va_lists (one of these existed already, of course), so that less code has to be duplicated in builtin-gc.c for error handling. The second patch contains the
builtin-gc.c code itself.
-James
@@ -0,0 +1,81 @@+/*+*gitgcbuiltincommand+*+*Cleanupunreachablefilesandoptimizetherepository.+*+*Copyright(c)2007JamesBowes+*+*Basedongit-gc.sh,whichis+*+*Copyright(c)2006ShawnO.Pearce+*/++#include"cache.h"+#include"run-command.h"++staticconstcharbuiltin_gc_usage[]="git-gc [--prune]";++staticintpack_refs;++staticintgc_config(constchar*var,constchar*value)+{+if(!strcmp(var,"gc.packrefs"))+if(strlen(value)==0||!strcmp(value,"notbare"))+pack_refs=!is_bare_repository();+else+pack_refs=git_config_bool(var,value);+else+returngit_default_config(var,value);+return0;+}++staticvoidrun_command_or_die(constchar*cmd,...)+{+interr;+va_listparams;++va_start(params,cmd);+err=run_command_va(cmd,params);+va_end(params);++switch(err){+case0:+return;+case-ERR_RUN_COMMAND_FORK:+die("unable to fork for %s",cmd);+case-ERR_RUN_COMMAND_EXEC:+die("unable to exec %s",cmd);+default:+die("%s died with strange error",cmd);+}+}++intcmd_gc(intargc,constchar**argv,constchar*prefix)+{+inti;+intprune=0;++git_config(gc_config);++for(i=1;i<argc;i++){+constchar*arg=argv[i];+if(!strcmp(arg,"--prune")){+prune=1;+continue;+}+/* perhaps other parameters later... */+break;+}+if(i!=argc)+usage(builtin_gc_usage);++if(pack_refs)+run_command_or_die("git-pack-refs","--prune",NULL);+run_command_or_die("git-reflog","expire","--all",NULL);+run_command_or_die("git-repack","-a","-d","-l",NULL);+if(prune)+run_command_or_die("git-prune",NULL);+run_command_or_die("git-rerere","gc",NULL);++return0;+}
Shawn recently sent a series which discourages the va_list versions of
run_command. I think that makes sense. So, using
run_command_v_opt(argv_pack_refs, RUN_GIT_CMD) would be better IMHO.
And instead of die()ing, I'd rather do something like
return (pack_refs || run_command_v_opt(argv_pack_refs, RUN_GIT_CMD) &&
run_command_v_opt(argv_reflog_expire, RUN_GIT_CMD) &&
run_command_v_opt(argv_repack, RUN_GIT_CMD) &&
(prune || run_command_v_opt(argv_prune, RUN_GIT_CMD) &&
run_command_v_opt(argv_rerere, RUN_GIT_CMD);
Hmm?
Ciao,
Dscho
On Sun, Mar 11, 2007 at 06:06:56PM -0400, James Bowes wrote:
The following two patches make git-gc a builtin command.
What's the advantage in making git-gc a builtin command? It's not
like it's going to help performance a whole lot (especially since
you're just forking separate processes to run git-prune,
git-pack-refs, et.al.), and as a shell script it's a lot easier to
explain to people what git-gc is actually doing, so there is
pedagogical value to keeping it as a shell script.
Regards,
- Ted
On Mon, Mar 12, 2007 at 12:23:41PM +0100, Johannes Schindelin wrote:
Hi,
On Sun, 11 Mar 2007, Theodore Tso wrote:
quoted
On Sun, Mar 11, 2007 at 06:06:56PM -0400, James Bowes wrote:
quoted
The following two patches make git-gc a builtin command.
What's the advantage in making git-gc a builtin command?
Portability. Plus, James wanted to get involved in Git development, and
building in gc really was the shortest path into that.
I'm not sure I understand the portability argument? All of the
platforms that git currently supports will handle shell scripts,
right?
Heck, git-commit is still a shell script, and that's a rather, ah,
fundamental command, isn't it?
- Ted
From: Shawn O. Pearce <hidden> Date: 2016-06-15 22:42:59
Johannes Schindelin [off-list ref] wrote:
On Sun, 11 Mar 2007, Theodore Tso wrote:
quoted
On Sun, Mar 11, 2007 at 06:06:56PM -0400, James Bowes wrote:
quoted
The following two patches make git-gc a builtin command.
What's the advantage in making git-gc a builtin command?
Portability. Plus, James wanted to get involved in Git development, and
building in gc really was the shortest path into that.
Actually, git-gc.sh is pretty portable. To POSIX systems.
Windows ain't POSIX. Getting rid of some of those shell scripts
just makes us more portable, even to Windows. (Yes, people really
do still get forced to use that non-operating system.)
Ted talked about git-commit.sh being more important, but Dsco
clipped it. ;-)
I think git-commit.sh and git-merge.sh should both get ported to
builtins too, as both are somewhat hairy in shell, are quite core
to the system, and would be faster on Windows if written in C
(less forking == more speed there).
But they are so core that any rewrite must be undertaken carefully.
--
Shawn.
I'm not sure I understand the portability argument? All of the
platforms that git currently supports will handle shell scripts,
right?
Git "supports" MinGW, or at least wants to. And yes, you can put bash in
there, but we'd be *so* much better off if we had no shell scripting at
all.
Another thing I find annoying (even as a UNIX user) is that whenever I do
any tracing for performance data, shell is absolutely horrid. It's *so*
much nicer to do 'strace' on built-in programs that it's not even funny.
It's also sad how many performance issues we've had with shell, just
because even something really simple (like a few hundred refs) is just too
slow for shell scripting.
Heck, git-commit is still a shell script, and that's a rather, ah,
fundamental command, isn't it?
Yeah, and that's probably my pet peeve. I'd love to see a built-in "git
commit" and "git fetch". The "fetch--tool" thing in next gets rid of some
of the latter (and apparently the worst performance problems), but it's
sad how we have a really nice builtin "push", but our "fetch" is still
mostly really hairy shell-code (not just "git-fetch.sh" itself, but
"git-parse-remote.sh".
A gold star for whoever gets rid of any of of commit/clone/fetch or
ls-remote
(ls-remote isn't that big or hairy, but I mention it because it's a user
of "parse-remote", so making even just ls-remote built-in is probably
going to help with fetch/clone eventually).
Linus
From: Jakub Narebski <hidden> Date: 2016-06-15 22:42:59
Linus Torvalds wrote:
Another thing I find annoying (even as a UNIX user) is that whenever I do
any tracing for performance data, shell is absolutely horrid. It's *so*
much nicer to do 'strace' on built-in programs that it's not even funny.
Isn't that what GIT_TRACE was made for?
--
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Another thing I find annoying (even as a UNIX user) is that whenever I do
any tracing for performance data, shell is absolutely horrid. It's *so*
much nicer to do 'strace' on built-in programs that it's not even funny.
Isn't that what GIT_TRACE was made for?
That just shows the high-level git commands.
If you look for performance issues or correctness issues (like when I
tried to figure out if O_LARGEFILE was set for "git clone"), GIT_TRACE
does nothing. You want to do "strace -f -o trace-file".
And shell scripts look horrible there, and make it much harder to follow
things. In fact, it doesn't even need to be shell per se, but fork/exec
already makes things harder to see, shell just tends to (a) make it even
more so (try stracing though a shell startup, ugh) and (b) cause tons of
fork/exec cases.
For example, when we made patch generation a built-in, it suddenly became
*hugely* easier to follow what was going on in the traces, because it got
much more streamlined. In general I find that "high performance" == "easy
to trace".
Linus