Thread (49 messages) flat view 49 messages, 6 authors, 5d ago

Re: [PATCH 2/2] packfile: recover when a multi-pack-index names a removed pack

From: Jeff King <hidden>
Date: 2026-08-24 04:48:30

On Thu, Aug 20, 2026 at 06:36:09PM -0700, Elijah Newren wrote:
quoted
quoted
The false negative is not limited to one caller.  Any reader
(cat-file, rev-list, pack-objects, ...) can spuriously fail with
"unable to read object", and callers that only ask whether an object
exists get a wrong answer too, since the OBJECT_INFO_QUICK path never
retries.  Writers that merge in-core, such as "git replay", are hit
hardest: merge-ort treats the unreadable tree as a premature abort, sets
result.clean < 0, and returns without a result tree.
Hm. Isn't there a slight variant of the race though for any caller that
does not use OBJECT_INFO_QUICK?

Namely, the packfile containing our object disappears and is being
written to a new packfile, and that file is the only one containing it.
Without OBJECT_INFO_QUICK we would be fine: we notice the object could
not be found, and then we perform a second read that makes the "packed"
backend reload its packfiles. It would find the new packfile, and
because it's not covered by its MIDX it would use it to surface the
object. But without OBJECT_INFO_QUICK that's not the case, as we would
skip reloading packfiles altogether, and hence we would not be able to
find that object at all.

As far as I can see though, we don't seem to pass OBJECT_INFO_QUICK in
any of the mentioned readers. I could very well be missing something
here, but I would have thought that those readers are fine in this
scenario?
Nicely caught -- and you're right that the readers named above are
fine: they're all non-QUICK, so the second read reloads the packfiles
and finds the object in its new, non-MIDX-covered home, exactly as you
describe.
OK, so do I understand correctly that you _can't_ get the "unable to
read object" result that the commit message claims? I.e., the reprepare
/ packfile reload is helps us (just like it does for the non-midx case
when an idx has been mapped but the pack disappears before we open it).

So there is no bug there for non-QUICK callers. But then...
But the variant you describe is a real bug for QUICK callers that
don't get that second read -- e.g. upload-pack's object-existence
checks and mktree --batch.  I have three more race-condition patches
to clean up and submit, and this is one of them: it forces the reload
even under OBJECT_INFO_QUICK once we notice a pack has vanished out
from under us.
This seems wrong. The whole point of the QUICK flag is that the caller
is OK producing a false negative for an object lookup, and it would
prefer that outcome to spending the time to reload. If there are callers
passing QUICK that aren't OK with false negatives, they are broken and
the fix should be there. But repreparing the packs for a QUICK miss is
going to reintroduce the performance problems that QUICK was introduced
to help.

So between the two cases, it sounds like things (or at least the
low-level lookups) are working as designed, and there is no bug. Or am I
misunderstanding something?
Your wording also makes me realize that my fix in this unsubmitted
patch still has a hole: it triggers when opening the pack .idx fails,
but if the timing is such that the .idx is already mmapped and only
the .pack has gone missing, it won't fire.  I'll look into that before
submitting...and then clean up/submit my two other race fixes as well.
I think it would be fine, for the same reason that regular idx lookups
are fine. In packfile_fill_entry() we call is_pack_valid(), checking
that the pack is still there (and relying on its side effect of leaving
the fd/mmap open so that it remains accessible even if the file is
deleted).

-Peff
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help