From: Junio C Hamano <hidden> Date: 2016-06-15 23:07:43
Jeff King [off-list ref] writes:
I also notice that if we are deleting, we _do_ set
RESOLVE_REF_NO_RECURSE from the very beginning, which means we would
generally not get a valid lock->old_oid.hash for a symref. But I'm not
sure what it would mean to delete a symref while asking for its current
value (it cannot have one!). So I don't think it is a bug.
I started scratching my head after noticing that the NO_RECURSE bit
set in the DELETING codepath before reading the above, and I am
still doing so.
A transaction that attempts to delete an existing symref presumably
wants to make sure that the "old" value it read hasn't changed, but
ensuring the object name (obtained by reading the ref that is
pointed by the symref by dereferencing) are the same is not the
right way to ensure nobody raced with us in the meantime anyway (we
should rather be making sure that the symref is still pointing at
the same ref), so in that sense, in the context of acquiring the
lock, old oid value is meaningless for symrefs.
This patch is a strict improvement as the behaviour for REF_DELETING
case is unchanged by it (an idempotent resolve-ref-unsafe may be
called one more time in some cases), and other cases are better, I
think.
From: Jeff King <hidden> Date: 2016-06-15 23:07:43
On Tue, Jan 12, 2016 at 11:41:17AM -0800, Junio C Hamano wrote:
Jeff King [off-list ref] writes:
quoted
I also notice that if we are deleting, we _do_ set
RESOLVE_REF_NO_RECURSE from the very beginning, which means we would
generally not get a valid lock->old_oid.hash for a symref. But I'm not
sure what it would mean to delete a symref while asking for its current
value (it cannot have one!). So I don't think it is a bug.
I started scratching my head after noticing that the NO_RECURSE bit
set in the DELETING codepath before reading the above, and I am
still doing so.
A transaction that attempts to delete an existing symref presumably
wants to make sure that the "old" value it read hasn't changed, but
ensuring the object name (obtained by reading the ref that is
pointed by the symref by dereferencing) are the same is not the
right way to ensure nobody raced with us in the meantime anyway (we
should rather be making sure that the symref is still pointing at
the same ref), so in that sense, in the context of acquiring the
lock, old oid value is meaningless for symrefs.
Right, that's the point I was trying to make. Though I think there is
something even more subtle going in (see the end of this message).
In theory you might want the old sha1 for logging purposes, but since we
delete the reflog along with the symref, I don't think it matters there,
either.
I'm not sure we actually delete symrefs very often, though. Grepping for
delete_ref and REF_NODEREF shows the callers expecting symrefs to mostly
be "git remote", which never passes in an old_sha1.
The only caller which does so is "branch -d", but I think it doesn't
affect symrefs. It gets the "old" sha1 by calling resolve_ref_unsafe()
itself with RESOLVE_REF_NO_RECURSE, so it will unconditionally remove a
symref you ask it to, even if somebody else raced and put something in
it.
You can also call "update-ref --no-ref -d" with an "old" sha1, but I
doubt anyone ever does so.
This patch is a strict improvement as the behaviour for REF_DELETING
case is unchanged by it (an idempotent resolve-ref-unsafe may be
called one more time in some cases), and other cases are better, I
think.
Yeah. My gut feeling is that the REF_DELETING special-handling of
REF_NODEREF could just be folded into what I've added in this series.
But absent a case that is demonstrably broken, I'm inclined not to muck
with it too much.
I had thought that this:
git init
git commit --allow-empty -m foo
git symbolic-ref refs/foo refs/heads/master
old=$(git rev-parse foo)
git update-ref --no-deref -d refs/foo $old
might trigger a problem (because reading refs/foo with NODEREF will give
us a blank sha1 to compare against). But of course that is nonsense. The
actual lock verification is not done by this initial resolve_ref. It
happens _after_ we take the lock (as it must to avoid races), when
verify_lock() calls read_ref_full().
But now I'm doubly confused. When we call read_ref_full(), it is _also_
into lock->old_oid.hash. Which should be overwriting what the earlier
resolve_ref_unsafe() wrote into it. Which would mean my whole commit is
wrong; we can just unconditionally do a non-recursive resolution in the
first place. But when I did so, I ended up with funny reflog values
(which is why I wrote the patch as I did).
Let me try to dig a little further into that case and see what is going
on.
-Peff
From: Jeff King <hidden> Date: 2016-06-15 23:07:43
On Tue, Jan 12, 2016 at 03:22:51PM -0500, Jeff King wrote:
I had thought that this:
git init
git commit --allow-empty -m foo
git symbolic-ref refs/foo refs/heads/master
old=$(git rev-parse foo)
git update-ref --no-deref -d refs/foo $old
might trigger a problem (because reading refs/foo with NODEREF will give
us a blank sha1 to compare against). But of course that is nonsense. The
actual lock verification is not done by this initial resolve_ref. It
happens _after_ we take the lock (as it must to avoid races), when
verify_lock() calls read_ref_full().
But now I'm doubly confused. When we call read_ref_full(), it is _also_
into lock->old_oid.hash. Which should be overwriting what the earlier
resolve_ref_unsafe() wrote into it. Which would mean my whole commit is
wrong; we can just unconditionally do a non-recursive resolution in the
first place. But when I did so, I ended up with funny reflog values
(which is why I wrote the patch as I did).
Let me try to dig a little further into that case and see what is going
on.
Ah, I see. When calling git-symbolic-ref, we don't provide an old_sha1,
and therefore never call verify_lock(). And we get whatever value in
lock->old_oid we happened to read earlier in resolve_ref_unsafe(). Which
happened outside of a lock. Yikes. It seems like we could racily write
the wrong reflog entry in such a case.
So I think we'd want something like this:
@@ -1845,7 +1845,7 @@ static int verify_lock(struct ref_lock *lock,errno=save_errno;return-1;}-if(hashcmp(lock->old_oid.hash,old_sha1)){+if(old_sha1&&hashcmp(lock->old_oid.hash,old_sha1)){strbuf_addf(err,"ref %s is at %s but expected %s",lock->ref_name,sha1_to_hex(lock->old_oid.hash),
to make sure that the value in lock->old_oid always comes from what we
read under the lock. And then the resolve_ref() calls in
lock_ref_sha1_basic() really don't matter. They are just about making
sure there is space to create the lockfile.
The patch above is not quite right; I'll work up a series that takes
this approach.
-Peff
From: Jeff King <hidden> Date: 2016-06-15 23:07:43
On Tue, Jan 12, 2016 at 03:42:29PM -0500, Jeff King wrote:
Ah, I see. When calling git-symbolic-ref, we don't provide an old_sha1,
and therefore never call verify_lock(). And we get whatever value in
lock->old_oid we happened to read earlier in resolve_ref_unsafe(). Which
happened outside of a lock. Yikes. It seems like we could racily write
the wrong reflog entry in such a case.
[...]
The patch above is not quite right; I'll work up a series that takes
this approach.
OK, here it is. This replaces the top two patches of jk/symbolic-ref
(i.e., everything in this thread I've sent in the last day or two).
Besides fixing the race (which is detailed in patch 2/3), I think the
resulting 3/3 is much cleaner. Sorry for all the false starts. The more
I looked at this, the more complex it seemed to get. But I _think_ this
is the right solution. :)
[1/3]: checkout,clone: check return value of create_symref
[2/3]: lock_ref_sha1_basic: always fill old_oid while holding lock
[3/3]: lock_ref_sha1_basic: handle REF_NODEREF with invalid refs
-Peff
From: Jeff King <hidden> Date: 2016-06-15 23:07:43
It's unlikely that we would fail to create or update a
symbolic ref (especially HEAD), but if we do, we should
notice and complain. Note that there's no need to give more
details in our error message; create_symref will already
have done so.
While we're here, let's also fix a minor memory leak in
clone.
Signed-off-by: Jeff King <redacted>
---
Same as before (but with the extra test added in v2, in case you didn't
that up yet).
builtin/checkout.c | 3 ++-
builtin/clone.c | 11 +++++++----
t/t2011-checkout-invalid-head.sh | 6 ++++++
3 files changed, 15 insertions(+), 5 deletions(-)
From: Jeff King <hidden> Date: 2016-06-15 23:07:43
Our basic strategy for taking a ref lock is:
1. Create $ref.lock to take the lock
2. Read the ref again while holding the lock (during which
time we know that nobody else can be updating it).
3. Compare the value we read to the expected "old_sha1"
The value we read in step (2) is returned to the caller via
the lock->old_oid field, who may use it for other purposes
(such as writing a reflog).
If we have no "old_sha1" (i.e., we are unconditionally
taking the lock), then we obviously must omit step 3. But we
_also_ omit step 2. This seems like a nice optimization, but
it means that the caller sees only whatever was left in
lock->old_oid from previous calls to resolve_ref_unsafe(),
which happened outside of the lock.
We can demonstrate this race pretty easily. Imagine you have
three commits, $one, $two, and $three. One script just flips
between $one and $two, without providing an old-sha1:
while true; do
git update-ref -m one refs/heads/foo $one
git update-ref -m two refs/heads/foo $two
done
Meanwhile, another script tries to set the value to $three,
also not using an old-sha1:
while true; do
git update-ref -m three refs/heads/foo $three
done
If these run simultaneously, we'll see a lot of lock
contention, but each of the writes will succeed some of the
time. The reflog may record movements between any of the
three refs, but we would expect it to provide a consistent
log: the "from" field of each log entry should be the same
as the "two" field of the previous one.
But if we check this:
perl -alne '
print "mismatch on line $."
if defined $last && $F[0] ne $last;
$last = $F[1];
' .git/logs/refs/heads/foo
we'll see many mismatches. Why?
Because sometimes, in the time between lock_ref_sha1_basic
filling lock->old_oid via resolve_ref_unsafe() and it taking
the lock, there may be a complete write by another process.
And the "from" field in our reflog entry will be wrong, and
will refer to an older value.
This is probably quite rare in practice. It requires writers
which do not provide an old-sha1 value, and it is a very
quick race. However, it is easy to fix: we simply perform
step (2), the read-under-lock, whether we have an old-sha1
or not. Then the value we hand back to the caller is always
atomic.
Signed-off-by: Jeff King <redacted>
---
refs/files-backend.c | 17 +++++++++++------
1 file changed, 11 insertions(+), 6 deletions(-)
From: Jeff King <hidden> Date: 2016-06-15 23:07:43
We sometimes call lock_ref_sha1_basic with REF_NODEREF
to operate directly on a symbolic ref. This is used, for
example, to move to a detached HEAD, or when updating
the contents of HEAD via checkout or symbolic-ref.
However, the first step of the function is to resolve the
refname to get the "old" sha1, and we do so without telling
resolve_ref_unsafe() that we are only interested in the
symref. As a result, we may detect a problem there not with
the symref itself, but with something it points to.
The real-world example I found (and what is used in the test
suite) is a HEAD pointing to a ref that cannot exist,
because it would cause a directory/file conflict with other
existing refs. This situation is somewhat broken, of
course, as trying to _commit_ on that HEAD would fail. But
it's not explicitly forbidden, and we should be able to move
away from it. However, neither "git checkout" nor "git
symbolic-ref" can do so. We try to take the lock on HEAD,
which is pointing to a non-existent ref. We bail from
resolve_ref_unsafe() with errno set to EISDIR, and the lock
code thinks we are attempting to create a d/f conflict.
Of course we're not. The problem is that the lock code has
no idea what level we were at when we got EISDIR, so trying
to diagnose or remove empty directories for HEAD is not
useful.
To make things even more complicated, we only get EISDIR in
the loose-ref case. If the refs are packed, the resolution
may "succeed", giving us the pointed-to ref in "refname",
but a null oid. Later, we say "ah, the null oid means we are
creating; let's make sure there is room for it", but
mistakenly check against the _resolved_ refname, not the
original.
We can fix this by making two tweaks:
1. Call resolve_ref_unsafe() with RESOLVE_REF_NO_RECURSE
when REF_NODEREF is set. This means any errors
we get will be from the orig_refname, and we can act
accordingly.
We already do this in the REF_DELETING case, but we
should do it for update, too.
2. If we do get a "refname" return from
resolve_ref_unsafe(), even with RESOLVE_REF_NO_RECURSE
it may be the name of the ref pointed-to by a symref.
We already normalize this back to orig_refname before
taking the lockfile, but we need to do so before the
null_oid check.
While we're rearranging the REF_NODEREF handling, we can
also bump the initialization of lflags to the top of the
function, where we are setting up other flags. This saves us
from having yet another conditional block on REF_NODEREF
just to set it later.
Signed-off-by: Jeff King <redacted>
---
refs/files-backend.c | 19 ++++++++++---------
t/t1401-symbolic-ref.sh | 7 +++++++
t/t2011-checkout-invalid-head.sh | 33 +++++++++++++++++++++++++++++++++
3 files changed, 50 insertions(+), 9 deletions(-)
@@ -25,4 +25,37 @@ test_expect_success 'checkout notices failure to lock HEAD' 'test_must_failgitcheckout-bother'+test_expect_success'create ref directory/file conflict scenario''+gitupdate-refrefs/heads/outer/innermaster&&++# do not rely on symbolic-ref to get a known state,+# as it may use the same code we are testing+reset_to_df(){+echo"ref: refs/heads/outer">.git/HEAD+}+'++test_expect_success'checkout away from d/f HEAD (unpacked, to branch)''+reset_to_df&&+gitcheckoutmaster+'++test_expect_success'checkout away from d/f HEAD (unpacked, to detached)''+reset_to_df&&+gitcheckout--detachmaster+'++test_expect_success'pack refs''+gitpack-refs--all--prune+'++test_expect_success'checkout away from d/f HEAD (packed, to branch)''+reset_to_df&&+gitcheckoutmaster+'++test_expect_success'checkout away from d/f HEAD (packed, to detached)''+reset_to_df&&+gitcheckout--detachmaster+' test_done
From: Eric Sunshine <hidden> Date: 2016-06-15 23:07:43
On Tue, Jan 12, 2016 at 4:44 PM, Jeff King [off-list ref] wrote:
Our basic strategy for taking a ref lock is:
[...]
If these run simultaneously, we'll see a lot of lock
contention, but each of the writes will succeed some of the
time. The reflog may record movements between any of the
three refs, but we would expect it to provide a consistent
log: the "from" field of each log entry should be the same
as the "two" field of the previous one.
From: Jeff King <hidden> Date: 2016-06-15 23:07:43
On Tue, Jan 12, 2016 at 08:25:45PM -0500, Eric Sunshine wrote:
On Tue, Jan 12, 2016 at 4:44 PM, Jeff King [off-list ref] wrote:
quoted
Our basic strategy for taking a ref lock is:
[...]
If these run simultaneously, we'll see a lot of lock
contention, but each of the writes will succeed some of the
time. The reflog may record movements between any of the
three refs, but we would expect it to provide a consistent
log: the "from" field of each log entry should be the same
as the "two" field of the previous one.
s/two/to/
Whoops. Was tweaking my scripts "one" and "two" to test the race while I
wrote up the commit message. :)
-Peff