From: Junio C Hamano <hidden> Date: 2016-06-15 22:48:18
Nicolas Pitre [off-list ref] writes:
quoted
Honesty is very good. An alternative implementation that does not hurt
performance as much as the "paranoia" would, and checks "the input well
enough" would be very welcome.
Can't we rely on the mtime of the source file? Sample it before
starting hashing it, then make sure it didn't change when done.
I suspect that opening to mmap(2), hashing once to compute the object
name, and deflating it to write it out, will all happen within the same
second, unless you are talking about a really huge file, or you started at
very near a second boundary.
I am perfectly Ok that it will have false negatives that way, but if the
probability of it catching problems is too low because git is too fast
relative to the file timestamp granularity, then it doesn't sound very
useful in practice, unless of course you are on better filesystems.
It won't have false positives and that is a very good thing, though.
From: Nicolas Pitre <nico@fluxnic.net> Date: 2016-06-15 22:48:18
On Thu, 18 Feb 2010, Junio C Hamano wrote:
Nicolas Pitre [off-list ref] writes:
quoted
quoted
Honesty is very good. An alternative implementation that does not hurt
performance as much as the "paranoia" would, and checks "the input well
enough" would be very welcome.
Can't we rely on the mtime of the source file? Sample it before
starting hashing it, then make sure it didn't change when done.
I suspect that opening to mmap(2), hashing once to compute the object
name, and deflating it to write it out, will all happen within the same
second, unless you are talking about a really huge file, or you started at
very near a second boundary.
How is the index dealing with this? Surely if a file is added to the
index and modified within the same second then 'git status' will fail to
notice the changes. I'm not familiar enough with that part of Git.
Alternatively, you could use the initial mtime sample to determine the
filesystem's time granularity by noticing how many LSBs are zero.
Let's say FAT should have a granularity of one second. Then if the
mtime of the file is less than one second away before starting to hash
then just wait for one second. If one second later the mtime has
changed and still less than a second away then abort. If after the hash
the mtime has changed then abort.
On a recent filesystem, it is likely that the mtime granularity is a
nanosecond. Nevertheless the above algorithm should just work all the
same, although it is unlikely that the mtime will be within the current
nanosecond, hence the probability for having to do an initial wait is
almost zero. On kernels without hires timers the granularity will be
like 10 ms.
Of course you might be unlucky and the initial mtime sample happens to
be right on a whole second even on a high resolution mtime filesystem,
in which case the delay test will consider one second instead of 10 ms
or whatever. but the probability is rather small that you'll end up
with all sub-second bits to be all zeroes causing a longer delay than
actually necessary, and this would matter only for files that would have
been modified within that second. I don't think there is a reliable way
to enquire a filesystem+OS time stamping granularity.
Nicolas
From: Jonathan Nieder <hidden> Date: 2016-06-15 22:48:18
Nicolas Pitre wrote:
On Thu, 18 Feb 2010, Junio C Hamano wrote:
quoted
I suspect that opening to mmap(2), hashing once to compute the object
name, and deflating it to write it out, will all happen within the same
second, unless you are talking about a really huge file, or you started at
very near a second boundary.
How is the index dealing with this? Surely if a file is added to the
index and modified within the same second then 'git status' will fail to
notice the changes. I'm not familiar enough with that part of Git.
See Documentation/technical/racy-git.txt and t/t0010-racy-git.sh.
Short version: in the awful case, the timestamp of the index is the
same as (or before) the timestamp of the file. Git will notice this
and re-hash the tracked file.
Alternatively, you could use the initial mtime sample to determine the
filesystem's time granularity by noticing how many LSBs are zero.
Yuck.
If such detection is going to happen, I would prefer to see it used
once to determine the initial value of a per-repository configuration
variable asking to speed up ‘git add’ and friends.
Note that we are currently not using the nsec timestamps to make any
important decisions, probably because in some filesystems they are
unreliable when inode cache entries are evicted (not sure about the
current status; does this work in NFS, for example?). Within the
short runtime of ‘git add’, I guess this would not be as much of a
problem.
Jonathan
On Thu, Feb 18, 2010 at 07:04:56PM -0600, Jonathan Nieder wrote:
Nicolas Pitre wrote:
quoted
On Thu, 18 Feb 2010, Junio C Hamano wrote:
quoted
I suspect that opening to mmap(2), hashing once to compute the object
name, and deflating it to write it out, will all happen within the same
second, unless you are talking about a really huge file, or you started at
very near a second boundary.
How is the index dealing with this? Surely if a file is added to the
index and modified within the same second then 'git status' will fail to
notice the changes. I'm not familiar enough with that part of Git.
See Documentation/technical/racy-git.txt and t/t0010-racy-git.sh.
Short version: in the awful case, the timestamp of the index is the
same as (or before) the timestamp of the file. Git will notice this
and re-hash the tracked file.
As far as I can tell, the index doesn't handle this case at all.
Suppose the file is modified during git add near the beginning of the
file, after git add has read that part of the file, but the modifications
finish before git add does. Now the mtime of the file is earlier
than the index timestamp, but the file contents don't match the index.
This holds even if the objects git adds to the index aren't corrupted.
Actually right now you can have all four combinations: index up to date
or not, and object matching its sha1 hash or not, depending on where and
when you modify data during an index update.
racy-git.txt doesn't discuss concurrent modification of files with the
index. It only discusses low-resolution file timestamps and modifications
at times that are close to, but not concurrent with, index modifications.
Git probably also doesn't handle things like NTP time corrections
(especially those where time moves backward by sub-second intervals) and
mismatched server/client clocks on remote filesystems either (mind you,
I know of no SCM that currently handles that case, and CVS in particular
is unusually bad at it).
Personally, I find the combination of nanosecond-precision timestamps
and network file systems amusing. At nanosecond precision, relativistic
effects start to matter across a volume of space the size of my laptop.
I'm not sure how timestamps at any resolution could be a reliable metric
for detecting changes to file contents in the general case. A valuable
hint in many cases, but not authoritative (unless they all come from a
single monotonic high-resolution clock guaranteed to increment faster than
git--but they don't).
rsync solves this sort of problem with a 'modification window' parameter,
which is a time interval that is "close enough" to consider two timestamps
to be equal. Some of rsync's use cases set that window to six months.
Git would use a modification window for the opposite reason rsync
does--rsync uses the window to avoid unnecessarily examining files that
have different timestamps, while git would use it to re-examine files
even when it appears to be unnecessary.
Git probably wants the modification window to be the maximum clock
offset between a network filesystem client and server plus the minimum
representable interval in the filesystem's timestamp data type--which
is a value git couldn't possibly know for some cases, so it needs input
from the user.
From: Junio C Hamano <hidden> Date: 2016-06-15 22:48:18
Zygo Blaxell [off-list ref] writes:
As far as I can tell, the index doesn't handle this case at all.
...
racy-git.txt doesn't discuss concurrent modification of files with the
index. It only discusses low-resolution file timestamps and modifications
at times that are close to, but not concurrent with, index modifications.
Correct. As I said a few times in this thread, a use case with concurrent
modifications is outside of the original design scope of git.
As you may have realized, racy-git solution actually _relies_ on lack of
concurrent modifications. The document does not even _talk_ about this
assumption, exactly because at least back then it was a common knowledge
shared by everybody that users are not supposed to muck with files in the
work tree until they get control back from git and they can keep both
halves if they get a new broken loose object if they did so ;-).
On Fri, Feb 19, 2010 at 09:52:10AM -0800, Junio C Hamano wrote:
As you may have realized, racy-git solution actually _relies_ on lack of
concurrent modifications.
It relies on 1) no concurrent modifications, 2) strictly increasing
timestamps, and 3) consistent timestamps between the working directory
and index. Believe it or not, out of those three I actually think the
first assumption is the most reasonable, because it's something a
user could prevent without using administrative privileges or changing
filesystems.
For performance reasons I frequently work with GIT_DIR on ext3 and working
directory on tmpfs (though a mix of ext3 and ext4 has the same issues).
One has second resolution, the other nanosecond. If two event timestamps
with different precisions are compared as values at maximum precision,
the later event might appear to occur before the earlier one.
I try to avoid using anything that relies on mtime on network filesystems
because a lot more than just Git breaks in those cases.
NTP breakage is rarer, mostly because NTP step events are rare, and
negative ones ever rarer. I've seen negative steps on laptops when they
lose time accuracy during suspend or when Internet access is unavailable
(or asymmetrically laggy), then get it back later. On the other hand,
I don't actually test for effects of this event anywhere, so I can't say
it hasn't caused breakage I don't know about, although I can say it hasn't
cause breakage I do know about either.