Re: [PATCH 3/7] odb/streaming: support streaming arbitrary object types

Justin Tobler <[email protected]> Tue, 4 Aug 2026 13:03:41 -0500
Newsgroups org.kernel.vger.git
Message-ID <anInniMjCtU9Qae7@denethor>
On 26/08/04 09:25AM, Patrick Steinhardt wrote:
> The object database supports the ability to write object streams into
> it. This functionality is used when we encounter a blob that is larger
> than "core.bigFileThreshold" so that we don't have to soak large files
> into memory.
> 
> As we only ever write large files, the infrastructure doesn't support
> specifying any other object type than "blob". This limitation is quite
> artificial though: there is no reason why we shouldn't support writing
> arbitrary large objects with a stream. While it's very unlikely that we
> encounter a huge object other than a blob, users are known to be
> creative and sometimes like to inflict pain on themselves by creating
> commits or trees that are huge.
> 
> Extend the infrastructure to support streaming arbitrary object types.
> For now we don't use this functionality anywhere, but it brings us a bit
> closer to unify `struct odb_read_stream` and `struct odb_write_stream`.

Very happy to see this change. :)

> Signed-off-by: Patrick Steinhardt <[email protected]>
> ---
>  builtin/unpack-objects.c      |  1 +
>  object-file.c                 | 31 +++++++++++++++----------------
>  odb/source-inmemory.c         |  2 +-
>  odb/source-loose.c            |  2 +-
>  odb/streaming.c               |  3 ++-
>  odb/streaming.h               |  3 ++-
>  t/unit-tests/u-odb-inmemory.c |  7 +++++--
>  7 files changed, 27 insertions(+), 22 deletions(-)

Just FYI, there is also a comment in "odb/transaction.h" for the
`write_object_stream` callback that is also now outdated due to this
change. We may want to update that too.

[snip]
> @@ -953,7 +953,7 @@ int index_fd(struct index_state *istate, struct object_id *oid,
>  				 type, path, flags);
>  	} else {
>  		struct odb_write_stream stream;
> -		odb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size));
> +		odb_write_stream_from_fd(&stream, fd, xsize_t(st->st_size), OBJ_BLOB);

We still only target large blobs for streaming here, but the underlying
infrastructure is now generic which is nice.

[snip]
> diff --git a/odb/streaming.h b/odb/streaming.h
> index 5e8e6e532e..3c8ed55129 100644
> --- a/odb/streaming.h
> +++ b/odb/streaming.h
> @@ -56,6 +56,7 @@ struct odb_write_stream {
>  	ssize_t (*read)(struct odb_write_stream *, unsigned char *, size_t);
>  	void *data;
>  	size_t size;
> +	enum object_type type;

We now store the object type in the stream itself. Similar to size
information, the type information is always known in advance when
creating the object stream.

The rest of this patch is just updating call sites accordingly. Looks
good.

-Justin