Thomas Gummerer [off-list ref] writes:
== Work done in the previous 12 weeks ==
- Definition of a tentative index file v5 format [1]. This differs
from the proposal in making it possible to bisect the directory
entries and file entries, to do a binary search. The exact bits
for each section were also defined. To further compress the index,
along with prefix compression, the stat data is hashed, since
it's only used for comparison, but the plain data is never used.
s/comparison/equality comparison/ perhaps?
Thanks to Michael Haggerty, Nguyen Thai Ngoc Duy, Thomas Rast
and Robin Rosenberg for feedback.
- Read the index format format and translate it to the current in
s/format format/on-disk file format/ or something?
memory format. This doesn't include reading any of the current
extensions, which are now part of the main index. The code again
is on github. [4] Thanks for reviewing the first steps to Thomas
Rast.
- Started implementing the writer, which extracts the directories from
the in-memory format, and writes the header and the directories to
disk.
- I found a few bugs in the algorithm for extracting the directories
and decided to completely rewrite it, using a hash table instead of
simple lists, since the old one would have to many corner cases to
handle.
What does "the algorithm" refer to? Is it the one described in the
previous bullet point, or is it the code in production? If latter,
it would help to separate out the task to fix the breakage, as
people with the current or previous versions of Git will be
negatively affected until that bug is fixed. If former, I am not
sure if this task needs to be described in two bullet points ("I did
X, X had bug so I redid X in a different way" is still a single task
to do X).
== Work done int the last week ==
- Polished the patch for the ce_namelen field. The thread for the
patch can be found at [5].
Thanks for this one; I think it is ready for 'next', but if you are
still not satisfied I do not mind waiting for further perfection.
On 07/16, Junio C Hamano wrote:
Thomas Gummerer [off-list ref] writes:
quoted
== Work done in the previous 12 weeks ==
- Definition of a tentative index file v5 format [1]. This differs
from the proposal in making it possible to bisect the directory
entries and file entries, to do a binary search. The exact bits
for each section were also defined. To further compress the index,
along with prefix compression, the stat data is hashed, since
it's only used for comparison, but the plain data is never used.
s/comparison/equality comparison/ perhaps?
Exactly, thanks.
quoted
Thanks to Michael Haggerty, Nguyen Thai Ngoc Duy, Thomas Rast
and Robin Rosenberg for feedback.
quoted
- Read the index format format and translate it to the current in
s/format format/on-disk file format/ or something?
Yes, thanks.
quoted
memory format. This doesn't include reading any of the current
extensions, which are now part of the main index. The code again
is on github. [4] Thanks for reviewing the first steps to Thomas
Rast.
quoted
- Started implementing the writer, which extracts the directories from
the in-memory format, and writes the header and the directories to
disk.
- I found a few bugs in the algorithm for extracting the directories
and decided to completely rewrite it, using a hash table instead of
simple lists, since the old one would have to many corner cases to
handle.
What does "the algorithm" refer to? Is it the one described in the
previous bullet point, or is it the code in production? If latter,
it would help to separate out the task to fix the breakage, as
people with the current or previous versions of Git will be
negatively affected until that bug is fixed. If former, I am not
sure if this task needs to be described in two bullet points ("I did
X, X had bug so I redid X in a different way" is still a single task
to do X).
It refers to the algorithm in the previous bullet point, which
extracts the directories, and can be included in the above bullet
point. Sorry for the confusion.
quoted
== Work done int the last week ==
- Polished the patch for the ce_namelen field. The thread for the
patch can be found at [5].
Thanks for this one; I think it is ready for 'next', but if you are
still not satisfied I do not mind waiting for further perfection.
Thanks, I'm satisfied with it, for me it can be merged to 'next'.
Thanks Junio for reading the progress report, this is just
corrected version without the errors that he pointed out.
== Work done in the previous 12 weeks ==
- Definition of a tentative index file v5 format [1]. This differs
from the proposal in making it possible to bisect the directory
entries and file entries, to do a binary search. The exact bits
for each section were also defined. To further compress the index,
along with prefix compression, the stat data is hashed, since
it's only used for equality comparison, but the plain data is
never used.
Thanks to Michael Haggerty, Nguyen Thai Ngoc Duy, Thomas Rast
and Robin Rosenberg for feedback.
- Prototype of a converter from the index format v2/v3 to the index
format v5. [2] The converter reads the index from a git repository,
can output parts of the index (header, index entries as in
git ls-files --debug, cache tree as in test-dump-cache-tree, or
the reuc data). Then it writes the v5 index file format to
.git/index-v5. Thanks to Michael Haggerty for the code review.
- Prototype of a reader for the new index file format. [3] The
reader has mainly the purpose to show the algorithm used to read
the index lexicographically sorted after the full name which is
required by the current internal memory format. Big thanks for
reviewing this code and giving me advice on refactoring goes
to Michael Haggerty.
- Read the on-disk index file format and translate it to the current
in memory format. This doesn't include reading any of the current
extensions, which are now part of the main index. The code again
is on github. [4] Thanks for reviewing the first steps to Thomas
Rast.
- Read the cache-tree data (formerly an extension, now it's integrated
with the rest of the directory data) from the new ondisk format.
There are still a few optimizations to do in this algorithm.
- Started implementing the API (suggested by Duy), but it's still
in the very early stages. There is one commit for this on GitHub [1],
but it's a very early work in progress.
- Started implementing the writer, which extracts the directories from
the in-memory format, and writes the header and the directories to
disk. The algorithm uses a hash-table instead of a simple list,
to avoid many corner cases.
- Implemented writing the file block to disk, and basic tests from the
test suite are running fine, not including tests that require
conflicted data or the cache-tree to work, which both are not
implemented yet.
- Started implementing a patch to introduce a ce_namelen field in
struct cache_entry and drop the name length from the flags. [5]
== Work done int the last week ==
- Polished the patch for the ce_namelen field. The thread for the
patch can be found at [5]. Again thanks to Junio, Duy and Thomas
for reviewing it and giving me suggestions for improving it.
- Implemented the cache-tree and conflict data writing to the
index-v5 file.
== Outlook for the next week ==
- There are still a few bugs in the conflict writing, which will
be fixed, to make the test suite pass with index-v5.
- Once the test suite passes, the code still needs to be refactored
and optimized.
- If the two points above go well, I'll continue working on the api
that Duy suggested.
[1] https://github.com/tgummerer/git/wiki/Index-file-format-v5
[2] https://github.com/tgummerer/git/blob/pythonprototype/git-convert-index.py
[3] https://github.com/tgummerer/git/blob/pythonprototype/git-read-index-v5.py
[4] https://github.com/tgummerer/git/tree/index-v5
[5] http://thread.gmane.org/gmane.comp.version-control.git/200997