[RFC] block: add an overlaybd image format driver
Huiba Li <[email protected]>
| Newsgroups | org.nongnu.qemu-devel |
|---|---|
| Message-ID | <CA+AFTT-iZ8wbUOkEaU_SCM0rnqRJsefkOr=CTgC6T22nz_ZFoA@mail.gmail.com> |
Hi,
I would like to propose adding a QEMU block driver for the
overlaybd image format, and I am looking for early feedback on
scope and design before sending patches.
What overlaybd is
-----------------
Overlaybd is an open-source image format (
https://containerd.github.io/overlaybd/)
hosted as a sub-project of containerd (CNCF graduated). A disk
image is a stack of read-only snapshot blobs plus a writable top
file, in a layout compatible with OCI image registries: a base
image is stored once and shared by every derivative image, and
data is fetched on demand rather than pulled in full before boot.
Two properties distinguish it from qcow2/VHDX/VMDK-style chains:
* Lookup cost is O(1) in chain depth. At open time the per-file
indices are merged into a single in-memory index (an
extent-based log-structured merge tree), so a read never walks
backing files, and per-file index caches do not multiply with
snapshots.
* The merged index is small (about 300 KB average for 50+ GB
images in our production data), and because extents are
variable-length, writes land at 512-byte granularity with no
cluster copy-on-write.
The format is openly specified:
https://github.com/containerd/overlaybd/tree/main/docs/specs/lsmt.md
https://github.com/containerd/overlaybd/tree/main/docs/specs/zfile.md
Reference implementation:
https://github.com/containerd/overlaybd
Overlaybd has been in production at Alibaba for years (Taobao,
Tmall, Alibaba Cloud, Function Compute) and is used by Azure AKS
Artifact Streaming, Databricks serverless compute, and several
microVM-based sandbox platforms. The design is documented in two
USENIX ATC papers (DADI, ATC'20; FaaSNet, ATC'21). I am one of the
overlaybd/DADI authors.
Why a QEMU driver
-----------------
Today overlaybd images are served to VMs through a kernel blk dev
and a userspace daemon. That fits container workloads well, but
a native QEMU driver would let any QEMU-based stack (libvirt,
KubeVirt, OpenStack, plain qemu-system-*) boot overlaybd images
directly, with QEMU's usual amenities (backing files, block jobs,
throttling, migration) and no extra daemon on the host. This is
beneficial to agent sandboxes.
Proposed implementation
-----------------------
A new block format driver, block/overlaybd.c, written in C from the
open spec (like the existing vmdk/vhdx drivers) rather than binding
the C++ reference implementation:
* v1: read-only support. Local layer files, index merge at open,
raw and ZFile blobs (lz4/zstd are already optional QEMU
dependencies). Writes can be served by a qcow2/raw file on top,
exactly like booting a read-only backing image today.
* v2: native writable layer, and lazy fetching of remote blobs —
either composed over the existing curl driver or with a small
built-in HTTP fetcher; guidance welcome.
Questions
---------
1. Is a new in-tree format driver acceptable in principle, given
the format is stable, openly specified, and deployed at scale?
2. Is read-only-first a reasonable merge scope for the initial
series (with iotests and a docs/interop/ spec page)?
3. For remote blobs: compose over block/curl, or fetch in-driver
with a local cache?
4. CLI surface: point -drive at the OCI-style JSON image config
(which lists the layer blobs), or expose per-layer options?
I will write the iotests and the docs/interop format page, and add
myself to MAINTAINERS in the first series.
Thanks for any comments,
Huiba Li