Re: [PATCH] packed-refs: use `fwrite()` when passing refs verbatim
From: Toon Claes <hidden>
Date: 2026-10-01 20:30:31
Karthik Nayak [off-list ref] writes:
The `write_with_updates()` function uses a `struct ref_iterator` to iterate over all refs to write to the temporary packfile. It receives
I would say in general "packfile" is only used for objects, but it seems there is one mention of "packfile" in this file, so I'm not sure.
the iterator from `packed_ref_iterator_begin()` which takes a snapshot of the 'packed-refs' file. While writing to the new packfile, writes are routed via `write_packed_entry()` which uses `fprintf()`. Even for references which haven't changed, we use the same mechanism. Instead, let's track the position of unchanged references in the snapshot iterator and directly use `fwrite()`.
There are a few sanitizations that get lost by using fwrite():
* oid_to_hex() is no longer called, so if an OID would be written in the
old packed-ref as uppercase, it would be copied as such
* next_record() uses isspace(3) to split the oid from the refname. If
the record would be separated by a tab instead of a space, that would
be copied over.
* next_record() calls check_refname_format(), if it ain't good, oidclr()
is called. Before this patch, that means the zero OID will be written
for such ref. That changes with this patch, and the original OID is
written if refname_is_safe() passes.
This won't make a difference in Git though, because upon reading back
from the packed-refs, the zero OID is read into memory anyway, so
there is no observable difference.
* similar to the item above, when the ref is peelable, and REF_ISBROKEN
(around line 985), the `peeled_oid` is not filled in. This means for a
refname not matching the format will not contain a peeled OID for that
ref.
For example, in the source file:
1937b04ead6e53d949a4e2eb97ed743b42692fde refs/tags/a-bad~tag
^67c1263aae5c1750e9d9211995e6d1bb38314fca
is converted to:
0000000000000000000000000000000000000000 refs/tags/a-bad~tag
After this patch, the ref and it's peeled commit OID is copied
verbatim.
I'm not saying any of this is bad. These are just some side-effects of
your fwrite() approach. And some might be worht mentioning in the commit
message.
This removes the unnecessary formatting operation involved. We can see a
consistent ~20% performance improvement when deleting from packed
references.
Benchmark 1: update-ref: delete ref (refcount = 100000, revision = master)
Time (mean ± σ): 28.7 ms ± 1.7 ms [User: 22.5 ms, System: 5.9 ms]
Range (min … max): 26.7 ms … 33.3 ms 46 runs
Benchmark 2: update-ref: delete ref (refcount = 100000, revision = b4/kn-speedup-packed-refs)
Time (mean ± σ): 23.8 ms ± 1.2 ms [User: 17.5 ms, System: 6.0 ms]
Range (min … max): 22.1 ms … 27.7 ms 56 runs
Summary
update-ref: delete ref (refcount = 100000, revision = b4/kn-speedup-packed-refs) ran
1.21 ± 0.09 times faster than update-ref: delete ref (refformat = files, refcount = 100000, revision = master)Not bad.
quoted hunk ↗ jump to hunk
Signed-off-by: Karthik Nayak <redacted> --- refs/packed-backend.c | 43 ++++++++++++++++++++++++++++++++----------- 1 file changed, 32 insertions(+), 11 deletions(-)diff --git a/refs/packed-backend.c b/refs/packed-backend.c index a73fc6aca7..ef952cdba6 100644 --- a/refs/packed-backend.c +++ b/refs/packed-backend.c@@ -879,6 +879,12 @@ struct packed_ref_iterator { /* The current position in the snapshot's buffer: */ const char *pos; + /* + * Start of the current record, set when advancing `pos`. Used to + * pass records verbatim to `fwrite()`.
I assume there was a version you've been working on that was calling fwrite() directly from ref_transaction_error write_with_updates()? With that being wrapped by a helper function, I think it's better to name that function here.
quoted hunk ↗ jump to hunk
+ */ + const char *record_start; + /* The end of the part of the buffer that will be iterated over: */ const char *eof;@@ -933,6 +939,7 @@ static int next_record(struct packed_ref_iterator *iter) if (iter->pos == iter->eof) return ITER_DONE; + iter->record_start = iter->pos; iter->base.ref.flags = REF_ISPACKED; p = iter->pos;@@ -1218,17 +1225,27 @@ static struct ref_iterator *packed_ref_iterator_begin( /* * Write an entry to the packed-refs file for the specified refname. - * If peeled is non-NULL, write it as the entry's peeled value. On - * error, return a nonzero value and leave errno set at the value left - * by the failing call to `fprintf()`. + * + * If the raw data is available, skip the formatting and directly write to + * the file using `fwrite()`. e.g. when deleting references and remaining + * refs need to be written verbatim. Otherwise, use `fprintf()`. + * + * If peeled is non-NULL, write it as the entry's peeled value. + * + * On error, return a nonzero value and leave errno set at the value left + * by the failing call to `fwrite()` or `fprintf()`. */ -static int write_packed_entry(FILE *fh, const char *refname, - const struct object_id *oid, +static int write_packed_entry(FILE *fh, const char *raw, size_t raw_len, + const char *refname, const struct object_id *oid, const struct object_id *peeled)
I'm not convinced it's worth to have both ways of writing in a single
function.
I rather keep this function as-is and add a function:
static int write_packed_entry_preformatted(FILE *fh,
const char *line,
size_t line_len)
{
if (fwrite(raw, raw_len, 1, fh) != 1)
return -1;
return 0;
}
Or maybe even:
static int write_packed_entry_from_iter(FILE *fh,
struct packed_ref_iterator *packed_iter)
{
size_t ret = fwrite(packed_iter->record_start,
packed_iter->pos - packed_iter->record_start,
1, fh);
if (ret != 1)
return -1;
return 0;
}
quoted hunk ↗ jump to hunk
{ - if (fprintf(fh, "%s %s\n", oid_to_hex(oid), refname) < 0 || - (peeled && fprintf(fh, "^%s\n", oid_to_hex(peeled)) < 0)) + if (raw) { + if (fwrite(raw, raw_len, 1, fh) != 1)
I see in some places, for example in fwrite_or_die(), `len` and `1` are swapped and the return value is compared against `len`. This is useful when the length can be 0. But that cannot the case here, so no need to change that.
quoted hunk ↗ jump to hunk
+ return -1; + } else if (fprintf(fh, "%s %s\n", oid_to_hex(oid), refname) < 0 || + (peeled && fprintf(fh, "^%s\n", oid_to_hex(peeled)) < 0)) {
For what's it's worth, I think it's really ugly to do this in a single if() statement, but that just could be me.
quoted hunk ↗ jump to hunk
return -1; + } return 0; }@@ -1530,9 +1547,13 @@ static enum ref_transaction_error write_with_updates(struct packed_ref_store *re } if (cmp < 0) { - /* Pass the old reference through. */ - if (write_packed_entry(out, iter->ref.name, - iter->ref.oid, iter->ref.peeled_oid)) + const struct packed_ref_iterator *packed_iter = + (const struct packed_ref_iterator *)iter; + size_t len = packed_iter->pos - packed_iter->record_start;
This is nice. Because packed_iter->pos points at the next record, the `len` will include the peeled OID line as well. Which is fwrite(3)'n in one go. That's a nice win.
quoted hunk ↗ jump to hunk
+ + if (write_packed_entry(out, packed_iter->record_start, + len, iter->ref.name, iter->ref.oid, + iter->ref.peeled_oid)) goto write_error; if ((ok = ref_iterator_advance(iter)) != ITER_OK) {@@ -1551,7 +1572,7 @@ static enum ref_transaction_error write_with_updates(struct packed_ref_store *re } else { bool peeled = update->flags & REF_HAVE_PEELED; - if (write_packed_entry(out, update->refname, + if (write_packed_entry(out, NULL, 0, update->refname, &update->new_oid, peeled ? &update->peeled : NULL)) goto write_error;--- base-commit: a018953688f1b10bddf91bff8747068f5f4746a4 change-id: 20260930-kn-speedup-packed-refs-9868f5d0abe9 Thanks - Karthik
-- Laters, Toon