Re: Calculating tree nodes
From: Andreas Ericsson <hidden>
Date: 2016-06-15 22:43:33
Jon Smirl wrote:
On 9/4/07, David Tweed [off-list ref] wrote:quoted
On 9/4/07, Jon Smirl [off-list ref] wrote:quoted
Git has picked up the hierarchical storage scheme since it was built on a hierarchical file system.FWIW my memory is that initial git used path-to-blob lists (as you're describing but without delta-ing) and tree nodes were added after a couple of weeks, the motivation _at the time_ being they were a natural way to dramatically reduce the size of repos. One of the nice things about tree nodes is that for doing a diff between versions you can, to overwhelming probability, decide equality/inequality of two arbitrarily deep and complicated subtrees by comparing 40 characters, regardless of how remote and convoluted their common ancestry. With delta chains don't you end up having to trace back to a common "entry" in the history? (Of course, I don't know how packs affect this - presumably there's some delta chasing to get to the bare objects as well.)While it is a 40 character compare, how many disk accesses were needed to get those two SHAs into memory?
One more than there would have been to read only the commit, and one more per level of recursion, assuming you never ever pack your repository. If you *do* pack it, the tree(s) needed to compare are likely already inside the sliding packfile window. In that case, there are no extra disk accesses. -- Andreas Ericsson andreas.ericsson@op5.se OP5 AB www.op5.se Tel: +46 8-230225 Fax: +46 8-230231