Skip to content

zarr-metadata

Basic tools for modelling Zarr metadata, with minimal dependencies.

zarr-metadata is developed in the zarr-python repository and released independently of zarr itself. Install it with:

pip install zarr-metadata

Who needs this

This library might be useful to you if your software interacts with Zarr metadata documents.

What this is

This library is not a full Zarr implementation. Instead, it's a collection of data structures and routines that closely model the content of the Zarr specifications, such as:

  • Typed JSON shapes (zarr_metadata.v2 and zarr_metadata.v3): TypedDict definitions and Literal aliases for the JSON documents specified by the Zarr v2 and Zarr v3 specifications, plus types for zarr-extensions and a few widely-used-but-unspecified entities (e.g. consolidated metadata).
  • Document models (zarr_metadata.model): canonical frozen-dataclass models of whole metadata documents, with validators, loc-aware parsers, and store-key (de)serialization. A document produced by to_json shares no mutable state with the model that produced it.
  • Optional Pydantic integration (zarr_metadata.pydantic, requires Pydantic 2.13 or newer): each model as a Pydantic field type that validates raw documents through the same strict parser.

What this is for

The public TypedDict definitions describe the static JSON shape of Zarr metadata. For strict, loc-aware validation of JSON loaded from disk, use the model parser:

import json
from zarr_metadata.model import ZarrV3ArrayMetadata

with open("zarr.json", "rb") as f:
    raw = json.load(f)

metadata = ZarrV3ArrayMetadata.from_json(raw)

The optional Pydantic integration delegates raw input to the same strict parser and returns the same normalized model class:

from pydantic import TypeAdapter
import zarr_metadata.pydantic as zmp

metadata = TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python(raw)
encoded = metadata.to_key_value()["zarr.json"]

A bare TypeAdapter over a public document TypedDict is a coercive shape adapter, not a Zarr conformance validator; it may coerce values or discard members that the strict model parser rejects.

Validation boundary

The model validators enforce the declared document structure and a small set of context-free consistency rules, including fixed format literals, finite JSON numbers, non-negative dimensions, and one dimension_names entry per array dimension. In a v3 document they also read each extension point -- the data type, chunk grid, chunk key encoding, each codec and each storage transformer -- through the definition that claims its name in a scope, CORE_AND_EXTENSIONS unless a context is passed: a configuration its definition refuses is refused, and a key it does not declare is reported as unknown_key. A name nothing in the scope claims is left unjudged, and whether to support it is the consumer's decision. A v3 fill value is judged against the data type it names, by that data type's definition, the chunk grid against the shape, by the grid's definition, and the codecs as a pipeline: in order, each judged by its definition against the chunk it is handed. The validators do not read a shard's inner codecs as a pipeline, and do none of a codec's arithmetic: whether a fill value survives a cast_value round trip is not judged.

Scope

At minimum, this library supports what Zarr-Python needs: the complete Zarr v2 and v3 specs, consolidated metadata, and a subset of the metadata defined in zarr-extensions. We are generally open to contributions that add types, models, or validation for Zarr metadata with a published spec.

Runtime array behavior is out of scope: nothing here encodes or decodes chunks, resolves codec or data type names to implementations, or performs store I/O. The models begin and end at the metadata documents themselves — from_key_value / to_key_value map documents to store keys and bytes, and everything past that belongs to consumer libraries.

Reference