Re: [PATCH v2 1/4] odb: decouple source path comparisons from `the_repository`
Jeff King <[email protected]>
| Newsgroups | org.kernel.vger.git |
|---|---|
| Message-ID | <[email protected]> |
On Fri, Aug 14, 2026 at 12:03:43PM -0700, Junio C Hamano wrote: > Jeff King <[email protected]> writes: > > > How bad is a duplicate alternate? It's a minor performance issue, I'd > > think. We would add its packs to the list (though hardly ever look > > through them, as the "first" copy would satisfy most requests, and the > > unused second copies end up at the back of the MRU list). You'd only pay > > the extra lookup cost for an object which we fail to find entirely, > > which is rare-ish (mostly speculative lookups for fetches). > > There may be a future application to be written to go through list > of alternates---enumerate all objects that exist in the first one, > and then remove them as duplicates to other alternates. Oops, there > was a duplicated entry and we ended up removing the objects from the > first one registered under a different spelling. Yeah, that would be dangerous. You _might_ even be able to trigger that now with an object directory that points to itself as an alternate, and then doing "git repack -adl" or similar. I don't recall offhand whether we normalize the names or if we'd be fooled by symlinks. Or for that matter if we are even careful about comparing alternates to the main odb directory. I hate to be cavalier about conditions that could cause data loss, but at the same time...it kind of feels like you'd have to be _trying_ to shoot yourself in the foot to create such a situation. > > Alternatively, I think we could probably make the check more thorough in > > a similar way. Always consider a pair of case-insensitive matches as > > possible duplicates, and then for each possible duplicate use stat() to > > check their st_dev and st_ino values. That keeps things cheap for normal > > cases, and we pay only the stat() before de-duping. It's correct and > > doesn't rely on the repo, though it is a bit more somewhat complicated > > code. > > Hmph, I prefer not to trust st_dev and st_ino on platforms where > case insensitivity can possibly become an issue, though. Yeah, I would prefer not to go down that road, either. There are a lot of complexity and portability headaches. I offered it mostly as a "you probably _could_ do this super-carefully" option, but my take is that we don't need to be super-careful. > > [1] Even on a single filesystem I think case-sensitivity check is not > > completely sufficient either. We know that filesystems do more > > complicated one-way transformations than just case folding, like > > unicode normalization or even removing some funky code points. > > We'd miss those "equivalent" spellings. > > macOS? Naturally. :) -Peff