Skip to content

Shared design for DataTree tables with duckdb-zarr: one schema per node #262

Description

@alxmrs

duckdb-zarr is adding support for Zarr stores with nested groups, such as those written by xarray.DataTree.to_zarr(). We'd like both projects to turn the same store into the same tables, with the same names, so this issue proposes a shared design. It builds on #82.

The duckdb-zarr version is decision 8 in xqlsystems/duckdb-zarr#52 (design only, no code yet).

Proposal

  1. A table is one dimension group in one node. It's identified by (node path, dims). Arrays in different nodes never share a table, even when their dimension names match. This is the DataTree model: simulation/coarse and simulation/fine can both have foo(x, y) with different lengths of x.
  2. Tables inside a node are named with the existing rules. default_table_name ("_".join(dims) in dimension order, scalar for none) and resolve_table_names (overrides keyed by the dims tuple, case-insensitive collision check), applied per node.
  3. A node maps to a schema. Tree → catalog, node path → schema, dimension group → table:
    • The root node is the default schema: main in DuckDB, public in DataFusion.
    • A nested path is one quoted schema name: sim."simulation/fine".x_y.
    • In DataFusion this is a CatalogProvider whose SchemaProviders are the nodes, as suggested in Support for xarray-datatrees? #82.
  4. Coordinates are inherited like xarray. A node's tables include coordinates from its ancestors for dimensions the node shares with them, the same view dt["simulation/fine"].to_dataset() gives.
  5. Opening one node uses xarray's word. duckdb-zarr adds read_zarr(store, group := 'simulation/fine'), matching xr.open_zarr(store, group=...). With no group, only the root node is read.

Questions for xarray-sql

  • Does a catalog per tree and a schema per node fit the planned from_datatree API? Or would you rather have flat table names in one schema (for example simulation_fine__x_y)?
  • How should the root node look? Is it the default schema, or a schema named after the tree?
  • Should table_names overrides for a tree be keyed by (node_path, dims), with the per-node dims keys kept as they are?
  • Is anything in xarray-sql's pivot tied to one flat Dataset that would make inherited coordinates hard?

Why now

The immediate user is AnnData (xqlsystems/duckdb-zarr#40). An AnnData store is a tree: obs and var are nodes, and X, layers and obsm share their axes. The AnnData-specific parts come after this: naming its axes and decoding its categorical, nullable and sparse groups. Agreeing on the tree mapping first means both projects expose AnnData the same way.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions