From: Jeff King <hidden> Date: 2016-06-15 22:57:44
When we try to load an object from disk and fail, our
general strategy is to see if we can get it from somewhere
else (e.g., a loose object). That lets users fix corruption
problems by copying known-good versions of objects into the
object database.
We already handle the case where we were not able to read
the delta from disk. However, when we find that the delta we
read does not apply, we simply die. This case is harder to
trigger, as corruption in the delta data itself would
trigger a crc error from zlib. However, a corruption that
pointed us at the wrong delta base might cause it.
We can do the same "fail and try to find the object
elsewhere" trick instead of dying. This not only gives us a
chance to recover, but also puts us on code paths that will
alert the user to the problem (with the current message,
they do not even know which sha1 caused the problem).
Signed-off-by: Jeff King <redacted>
---
I needed this earlier today to recover from a corrupted packfile (I
fortunately had an older version of the repo in backups). Still tracking
down the exact nature of the corruption.
sha1_file.c | 11 ++++++++++-
1 file changed, 10 insertions(+), 1 deletion(-)
From: Nicolas Pitre <nico@fluxnic.net> Date: 2016-06-15 22:57:44
On Thu, 13 Jun 2013, Jeff King wrote:
When we try to load an object from disk and fail, our
general strategy is to see if we can get it from somewhere
else (e.g., a loose object). That lets users fix corruption
problems by copying known-good versions of objects into the
object database.
We already handle the case where we were not able to read
the delta from disk. However, when we find that the delta we
read does not apply, we simply die. This case is harder to
trigger, as corruption in the delta data itself would
trigger a crc error from zlib. However, a corruption that
pointed us at the wrong delta base might cause it.
We can do the same "fail and try to find the object
elsewhere" trick instead of dying. This not only gives us a
chance to recover, but also puts us on code paths that will
alert the user to the problem (with the current message,
they do not even know which sha1 caused the problem).
Signed-off-by: Jeff King <redacted>
That makes sense.
Could you produce a test case to go along with this change?
quoted hunk
---
I needed this earlier today to recover from a corrupted packfile (I
fortunately had an older version of the repo in backups). Still tracking
down the exact nature of the corruption.
sha1_file.c | 11 ++++++++++-
1 file changed, 10 insertions(+), 1 deletion(-)
From: Jeff King <hidden> Date: 2016-06-15 22:57:45
On Thu, Jun 13, 2013 at 08:05:21PM -0400, Nicolas Pitre wrote:
quoted
We already handle the case where we were not able to read
the delta from disk. However, when we find that the delta we
read does not apply, we simply die. This case is harder to
trigger, as corruption in the delta data itself would
trigger a crc error from zlib. However, a corruption that
pointed us at the wrong delta base might cause it.
That makes sense.
Could you produce a test case to go along with this change?
Yes. I was a little worried I would have trouble doing it without
relying on a lot of pack internals, but the infrastructure you set up in
t5303 makes it relatively easy (and we do not have to make any
assumptions that t5303 does not already make).
Here is a re-roll; the first patch is a small cleanup in t5303 that is
required for the new tests to work.
[1/2]: t5303: drop "count=1" from corruption dd
[2/2]: unpack_entry: do not die when we fail to apply a delta
-Peff
From: Jeff King <hidden> Date: 2016-06-15 22:57:45
This test corrupts pack objects by using "dd" with a seek
command. It passes "count=1 bs=1" to munge just a single
byte. However, the test added in commit b3118bdc wants to
munge two bytes, and the second byte of corruption is
silently ignored.
This turned out not to impact the test, however. The idea
was to reduce the "size of this entry" part of the header so
that zlib runs out of input bytes while inflating the entry.
That header is two bytes long, and the test reduced the
value of both bytes; since we experience the problem if we
are off by even 1 byte, it is sufficient to munge only the
first one.
Even though the test would have worked with only a single
byte munged, and we could simply tweak the test to use a
single byte, it makes sense to lift this 1-byte restriction
from do_corrupt_object. It will allow future tests that do
need to change multiple bytes to do so.
Signed-off-by: Jeff King <redacted>
---
t/t5303-pack-corruption-resilience.sh | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Jeff King <hidden> Date: 2016-06-15 22:57:45
When we try to load an object from disk and fail, our
general strategy is to see if we can get it from somewhere
else (e.g., a loose object). That lets users fix corruption
problems by copying known-good versions of objects into the
object database.
We already handle the case where we were not able to read
the delta from disk. However, when we find that the delta we
read does not apply, we simply die. This case is harder to
trigger, as corruption in the delta data itself would
trigger a crc error from zlib. However, a corruption that
pointed us at the wrong delta base might cause it.
We can do the same "fail and try to find the object
elsewhere" trick instead of dying. This not only gives us a
chance to recover, but also puts us on code paths that will
alert the user to the problem (with the current message,
they do not even know which sha1 caused the problem).
Note that unlike some other pack corruptions, we do not
recover automatically from this case when doing a repack.
There is nothing apparently wrong with the delta, as it
points to a valid, accessible object, and we realize the
error only when the resulting size does not match up. And in
theory, one could even have a case where the corrupted size
is the same, and the problem would only be noticed by
recomputing the sha1.
We can get around this by recomputing the deltas with
--no-reuse-delta, which our test does (and this is probably
good advice for anyone recovering from pack corruption).
Signed-off-by: Jeff King <redacted>
---
I don't know if it is worth showing that pack-objects without
"--no-reuse-delta" does not help this case with a test_expect_failure. I
don't have plans to work on it, and I'm not sure that it is a big deal.
sha1_file.c | 11 ++++++++++-
t/t5303-pack-corruption-resilience.sh | 27 +++++++++++++++++++++++++++
2 files changed, 37 insertions(+), 1 deletion(-)
@@ -276,6 +276,33 @@ test_expect_success \gitcat-fileblob$blob_3>/dev/null' test_expect_success\+'corruption of delta base reference pointing to wrong object'\+'create_new_pack--delta-base-offset&&+gitprune-packed&&+printf"\220\033"|do_corrupt_object$blob_32&&+gitcat-fileblob$blob_1>/dev/null&&+gitcat-fileblob$blob_2>/dev/null&&+test_must_failgitcat-fileblob$blob_3>/dev/null'++test_expect_success\+'... but having a loose copy allows for full recovery'\+'mv${pack}.idxtmp&&+githash-object-tblob-wfile_3&&+mvtmp${pack}.idx&&+gitcat-fileblob$blob_1>/dev/null&&+gitcat-fileblob$blob_2>/dev/null&&+gitcat-fileblob$blob_3>/dev/null'++test_expect_success\+'... and then a repack "clears" the corruption'\+'do_repack--delta-base-offset--no-reuse-delta&&+gitprune-packed&&+gitverify-pack${pack}.pack&&+gitcat-fileblob$blob_1>/dev/null&&+gitcat-fileblob$blob_2>/dev/null&&+gitcat-fileblob$blob_3>/dev/null'++test_expect_success\'corrupting header to have too small output buffer fails unpack'\'create_new_pack&&gitprune-packed&&