Thread (18 messages) 18 messages, 4 authors, 4d ago

Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)

From: Pablo Sabater <hidden>
Date: 2026-09-30 19:06:25

On Wed Sep 30, 2026 at 7:14 PM WEST, Derrick Stolee wrote:
On 9/29/2026 8:21 PM, Pablo Sabater wrote:
quoted
[Cc'd Derrick Stolee for his work in the backfill(1) command]

This series adds a --dry-run option to git-backfill(1) that reports how
many missing blobs would be fetched and, when the remote server
supports the object-info capability, their total size:

        $ git backfill --dry-run
        After backfill, 48 blobs would be fetched (1.20 KiB).

If the server does not advertise object-info, only the count is shown.
This is a helpful capability, but I'm not sure the size check counts
as a "dry run" because it involves a network call (and possibly many
depending on --min-batch-size).

Perhaps a different argument would be better, such as --info=(count|size)
to make it clear what level of information you want to know in advance
and thus how much effort are you willing to put in to discover this. 
Makes sense to have it as an --info option.
quoted
I am not a git-backfill(1) user myself, but it seemed useful for users
to know how much data a backfill would bring in before running it.
I'm not sure that we want to add a feature based on speculation. Git
is a collection of "itches" that the contributors needed scratched.
The work is motivated by real needs.

While I can see some benefit to curiosity, I'm not sure how much this
would prevent users from making their decision as to whether they
should run backfill or not.
quoted
The number of missing blobs is the sum of the number of blobs to be
fetched in each batch. The object-info capability lets us ask the server
for the size of each blob without downloading it, so summing them gives
an estimate of the total.

Note that this is an upper bound rather than the exact disk usage:
object-info reports the uncompressed size of each object, while the
objects end up stored compressed and possibly deltified in a packfile,
so the space actually used on disk will usually be smaller.
I don't think the uncompressed size is a useful metric here, as it is
likely astronomically larger than what will be downloaded. How will
this help a user make a decision?
Yes, that's one of the itches I have with the object-info protocol: it
cannot give you a reliable compressed size, and that's why only the
total size is supported.

The object-info protocol could be extended to support the
objectsize:disk attribute, either by having the server know what we
already have and the oids that we want, or by directly having the
server send us the compressed size of its local copy (to avoid too much
work).  Even with the first option, because it goes in batches, it
would still be an estimate, just a closer one.

Given that I don't use backfill, I Cc'd you because I wasn't sure if it
was really useful, and the main motivation was the "I'm going to check
--dry-run before backfilling" case, so it helps to decide.  If it's not
that helpful and seems to end up as a decoration option, it might be
better to drop it.
Thanks,
-Stolee
Thanks for taking a look,
Pablo
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help