Re: [PATCH v2 1/4] odb: decouple source path comparisons from `the_repository`

Jeff King <[email protected]>
Newsgroups org.kernel.vger.git
Message-ID <[email protected]>
On Fri, Aug 14, 2026 at 12:03:43PM -0700, Junio C Hamano wrote:

> Jeff King <[email protected]> writes:
> 
> > How bad is a duplicate alternate? It's a minor performance issue, I'd
> > think. We would add its packs to the list (though hardly ever look
> > through them, as the "first" copy would satisfy most requests, and the
> > unused second copies end up at the back of the MRU list). You'd only pay
> > the extra lookup cost for an object which we fail to find entirely,
> > which is rare-ish (mostly speculative lookups for fetches).
> 
> There may be a future application to be written to go through list
> of alternates---enumerate all objects that exist in the first one,
> and then remove them as duplicates to other alternates.  Oops, there
> was a duplicated entry and we ended up removing the objects from the
> first one registered under a different spelling.

Yeah, that would be dangerous. You _might_ even be able to trigger that
now with an object directory that points to itself as an alternate, and
then doing "git repack -adl" or similar. I don't recall offhand whether
we normalize the names or if we'd be fooled by symlinks. Or for that
matter if we are even careful about comparing alternates to the main odb
directory.

I hate to be cavalier about conditions that could cause data loss, but
at the same time...it kind of feels like you'd have to be _trying_ to
shoot yourself in the foot to create such a situation.

> > Alternatively, I think we could probably make the check more thorough in
> > a similar way. Always consider a pair of case-insensitive matches as
> > possible duplicates, and then for each possible duplicate use stat() to
> > check their st_dev and st_ino values. That keeps things cheap for normal
> > cases, and we pay only the stat() before de-duping. It's correct and
> > doesn't rely on the repo, though it is a bit more somewhat complicated
> > code.
> 
> Hmph, I prefer not to trust st_dev and st_ino on platforms where
> case insensitivity can possibly become an issue, though.

Yeah, I would prefer not to go down that road, either. There are a lot
of complexity and portability headaches. I offered it mostly as a "you
probably _could_ do this super-carefully" option, but my take is that we
don't need to be super-careful.

> > [1] Even on a single filesystem I think case-sensitivity check is not
> >     completely sufficient either. We know that filesystems do more
> >     complicated one-way transformations than just case folding, like
> >     unicode normalization or even removing some funky code points.
> >     We'd miss those "equivalent" spellings.
> 
> macOS?

Naturally. :)

-Peff
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.