Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
858704d
Add an ADBC engine adapter
alxmrs Sep 26, 2026
793d5d4
Don't fall back to flat tables through an aborted transaction
alxmrs Sep 26, 2026
41e94e0
Support ClickHouse in the ADBC adapter
alxmrs Sep 26, 2026
8d57591
Document ClickHouse in the ADBC section
alxmrs Sep 26, 2026
cac9fc6
Keep missing values in ClickHouse float columns (#257)
kentstephen Sep 27, 2026
82a2c6d
Compare zone-aware result times with zone-labeled window literals
alxmrs Sep 27, 2026
af18619
Quote identifiers in each database's own SQL; register on MySQL
alxmrs Sep 27, 2026
9ab110a
Round-trip timedelta coordinates through every ADBC database
alxmrs Sep 27, 2026
2a68d57
Document ClickHouse details from review; run ClickHouse tests on chDB…
alxmrs Sep 27, 2026
3d32144
Narrow widened types back to the template; parse text times on spill
alxmrs Sep 27, 2026
bbc8e32
Keep each database's ADBC differences in one dialect table
alxmrs Sep 27, 2026
6738827
Run one ADBC contract suite against every available database
alxmrs Sep 27, 2026
af90d89
Document the tested databases, temporary tables, and ingest options
alxmrs Sep 27, 2026
9242344
Return a str from the test's driver-path helper (mypy)
alxmrs Sep 27, 2026
2067d87
Run the ADBC contract against database servers in CI (first pass)
alxmrs Sep 27, 2026
763f346
Support SQL Server in the ADBC adapter
alxmrs Sep 27, 2026
869add6
Automate the per-database corner cases from review in the contract
alxmrs Sep 27, 2026
5d3b314
Never lose data silently on ingest: a type pass per database
alxmrs Sep 27, 2026
adc96d3
Run the ADBC contract against GizmoSQL (DuckDB over Flight SQL) in CI
alxmrs Sep 27, 2026
e9051cb
Add realistic ARCO-ERA5 queries across every ADBC database
alxmrs Sep 27, 2026
dbd9dac
Infer result dims from a mixed-dimension template
alxmrs Sep 27, 2026
d8d3207
Run the ARCO-ERA5 queries on every database in the adbc databases wor…
alxmrs Sep 27, 2026
8cb75a9
Write SQLite's whole-second times exactly as datetime() does
alxmrs Sep 27, 2026
229f38a
Parse spilled text times as ISO 8601, not an inferred format
alxmrs Sep 27, 2026
324442f
ANALYZE tables after ingest on PostgreSQL
alxmrs Sep 27, 2026
673432c
Document MariaDB's nested-loop joins, Trino ingest, and PostgreSQL AN…
alxmrs Sep 27, 2026
8aee30a
Run the database CI as one job per database; cancel superseded runs
alxmrs Sep 27, 2026
8a6108c
Restore a widened dtype only for plain selects, decided by dims
alxmrs Sep 28, 2026
95cab62
Catch only driver errors in the ADBC adapter; replace atomically on C…
alxmrs Sep 28, 2026
ce3b79a
Pin CI's database images and ADBC drivers; require adbc-driver-manage…
alxmrs Sep 28, 2026
a48896f
Type driver_error's lookup (mypy)
alxmrs Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
162 changes: 162 additions & 0 deletions .github/workflows/adbc-databases.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,162 @@
# Runs the ADBC adapter's contract tests and the ARCO-ERA5 queries against
# real database servers in service containers, one job per database so
# each pulls only its own image and a failure names its database. The
# main CI job runs the contract on the in-process backends (SQLite,
# DuckDB, chDB, DataFusion); the "in-process" job here adds their
# ARCO-ERA5 queries.
name: adbc databases

on:
push:
branches: [ main ]
pull_request:
paths:
- "xarray_sql/backends/adbc.py"
- "xarray_sql/backends/_adbc_dialects.py"
- "xarray_sql/ds.py"
- "xarray_sql/lazyscan.py"
- "xarray_sql/roundtrip.py"
- "tests/_adbc.py"
- "tests/test_adbc_backend.py"
- "tests/test_adbc_era5_integration.py"
- ".github/workflows/adbc-databases.yml"
workflow_dispatch:

# A newer push to the same pull request supersedes a running check.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}

jobs:
databases:
name: ${{ matrix.backend }}
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
include:
- backend: sqlite,duckdb,chdb,datafusion
- backend: postgresql
image: postgres:18
port: 5432
uri: postgresql://postgres:xql@localhost:5432/postgres
- backend: mysql
image: mysql:8.4
port: 3306
uri: mysql://root:xql@127.0.0.1:3306/xql
- backend: mariadb
image: mariadb:11.4
port: 3306
uri: mysql://root:xql@127.0.0.1:3306/xql
- backend: clickhouse
image: clickhouse/clickhouse-server:26.8
port: 8123
uri: http://localhost:8123/?user=default&password=xql
- backend: trino
image: trinodb/trino:483
port: 8080
uri: http://ci@localhost:8080?catalog=memory&schema=default
- backend: mssql
# Floating tags pinned to the digests of the last green run.
image: mcr.microsoft.com/mssql/server:2022-latest@sha256:4402d880dd4c34bfa7d8705e56a86cd6c88da80a1f6bbbe741f999e76264a090
port: 1433
uri: sqlserver://sa:XqlPassw0rd@localhost:1433?database=master
- backend: flightsql
image: gizmodata/gizmosql:latest@sha256:7d42a760fe9ba0bf6afb3580a7d962125eb68a3082d19768e13508a4db7c8a96
port: 31337
uri: grpc://localhost:31337
services:
# An empty image (the in-process databases) starts no container.
# Each image reads only its own variables below.
db:
image: ${{ matrix.image }}
ports:
- ${{ matrix.port || 1 }}:${{ matrix.port || 1 }}
env:
POSTGRES_PASSWORD: xql
MYSQL_ROOT_PASSWORD: xql
MYSQL_DATABASE: xql
MARIADB_ROOT_PASSWORD: xql
MARIADB_DATABASE: xql
CLICKHOUSE_PASSWORD: xql
ACCEPT_EULA: "Y"
MSSQL_SA_PASSWORD: XqlPassw0rd
TLS_ENABLED: "0"
GIZMOSQL_USERNAME: xql
GIZMOSQL_PASSWORD: xql
env:
VIRTUAL_ENV: ${{ github.workspace }}/.venv
XARRAY_SQL_TEST_ONLY: ${{ matrix.backend }}
XARRAY_SQL_TEST_FLIGHTSQL_USERNAME: xql
XARRAY_SQL_TEST_FLIGHTSQL_PASSWORD: xql
steps:
- uses: actions/checkout@v4

- uses: dtolnay/rust-toolchain@stable

- name: Setup sccache
uses: mozilla-actions/sccache-action@v0.0.9

- name: Configure sccache
run: |
echo "SCCACHE_GHA_ENABLED=true" >> $GITHUB_ENV
echo "RUSTC_WRAPPER=sccache" >> $GITHUB_ENV

- uses: astral-sh/setup-uv@v5
with:
python-version: "3.12"
enable-cache: true

- name: Install xarray_sql
run: uv sync --dev --no-install-package xarray-sql
- name: build rust
run: uv run --no-project maturin develop --uv

- name: Install ADBC drivers
run: |
# Pinned, so an upstream release cannot turn this red unannounced.
uv pip install dbc adbc-driver-postgresql==1.12.0 adbc-driver-flightsql==1.12.0
for driver in chdb=26.7.0 datafusion=0.27.0 mysql=0.6.1 \
clickhouse=0.1.1 trino=0.5.3 mssql=1.6.2; do
uv run --no-project dbc install "$driver"
done

- name: Point the tests' backend table at the server
if: matrix.uri
run: |
echo "XARRAY_SQL_TEST_${BACKEND^^}_URI=${{ matrix.uri }}" >> "$GITHUB_ENV"
env:
BACKEND: ${{ matrix.backend }}

- name: Wait for the server
if: matrix.uri
run: |
uv run --no-project python - <<'PY'
import os, sys, time
sys.path.insert(0, os.getcwd())
from tests._adbc import BACKENDS
backend = next(b for b in BACKENDS if b.name == "${{ matrix.backend }}")
for attempt in range(100):
try:
con = backend.connect()
cur = con.cursor()
cur.execute("SELECT 1")
cur.fetchall()
print(f"{backend.name} is up")
break
except Exception as exc:
print(f"waiting for {backend.name}: {str(exc)[:120]}")
time.sleep(3)
else:
sys.exit(f"{backend.name} did not start")
PY

- name: Run the ADBC contract tests
if: matrix.uri
run: uv run --no-project pytest -v -rs tests/test_adbc_backend.py

- name: Run realistic ARCO-ERA5 queries
# Reads a regional subset anonymously from the public bucket.
run: >-
uv run --no-project pytest -v -rs -m integration
tests/test_adbc_era5_integration.py
12 changes: 12 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -68,5 +68,17 @@ jobs:
run: uv sync --dev --no-install-package xarray-sql
- name: build rust
run: uv run --no-project maturin develop --uv
- name: Install in-process ADBC drivers
# chDB (embedded ClickHouse) and DataFusion run the ADBC adapter's
# contract tests without a server. dbc installs into the active
# virtualenv.
env:
VIRTUAL_ENV: ${{ github.workspace }}/.venv
run: |
uv pip install dbc
uv run --no-project dbc install "chdb=26.7.0"
uv run --no-project dbc install "datafusion=0.27.0"
- name: Run unit tests
env:
VIRTUAL_ENV: ${{ github.workspace }}/.venv
run: uv run --no-project pytest -v . -m "not integration"
7 changes: 6 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,10 +77,15 @@ rel = con.sql('SELECT time, AVG("air") AS air FROM air GROUP BY time ORDER BY ti
xql.to_dataset(rel, template=ds) # any engine's Arrow result round-trips
```

Any database with an [ADBC](https://arrow.apache.org/adbc/) driver
(PostgreSQL, SQLite, Snowflake, BigQuery, ...) works too: `xql.register`
ingests the Dataset into a table there, and `xql.to_dataset(cursor, ...)`
brings results back.

`table_names` (below) works the same way on every engine, so a query written
against `era5.surface` is not tied to the engine it was written for.

See [Engines](https://xqlsystems.github.io/xarray-sql/latest/engines/) for the support matrix, DuckDB/Polars details,
See [Engines](https://xqlsystems.github.io/xarray-sql/latest/engines/) for the support matrix, DuckDB/Polars/ADBC details,
and the lazy chunked round-trip.

## A bigger example: ARCO-ERA5
Expand Down
176 changes: 165 additions & 11 deletions docs/engines.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,22 +190,176 @@ chunked round-trip is fully supported: windows re-execute on Polars'
streaming engine.


## ADBC (adapter; any database with a driver)

```sh
pip install 'xarray-sql[adbc]' adbc-driver-postgresql # or -sqlite, -snowflake, ...
```

[ADBC](https://arrow.apache.org/adbc/) is a database-neutral API whose
drivers speak Arrow natively. `xql.register` accepts any ADBC DBAPI
connection, so PostgreSQL, SQLite, Snowflake, BigQuery, Flight SQL,
DuckDB, and every other database with an ADBC driver share one code
path:

```python
import adbc_driver_postgresql.dbapi
import xarray_sql as xql

con = adbc_driver_postgresql.dbapi.connect("postgresql://localhost/weather")
xql.register(con, "era5", ds) # seam 1: ingest
con.commit()

cur = con.cursor()
cur.execute("""
SELECT time, lat, lon, AVG(t2m) AS t2m
FROM era5
WHERE lat BETWEEN 40 AND 41
GROUP BY time, lat, lon
ORDER BY time, lat, lon
""")
out = xql.to_dataset(cur, template=ds) # seam 2
```

**Registration copies the data.** An ADBC database usually runs in
another process or on another machine, so it cannot call back into
Python to scan a lazy Dataset while a query runs. The adapter instead
streams the Dataset into a new table with ADBC's bulk ingest: chunks
are read on the same prefetching scan the DuckDB adapter uses
(`batch_size`, `prefetch`, `prefetch_bytes`, `coalesce_rows` tune it),
so memory stays bounded while the driver writes, and queries afterwards
run entirely in the database. Ingest what you intend to query —
`ds.sel(...)` a region or `ds[[...]]` a few variables first — rather
than a whole archive.

Options specific to this adapter:

- `mode="create"` (default) raises if the table exists; `"replace"`
drops and recreates it; `"append"` and `"create_append"` add rows,
which is how to load a long time series in slices.
- `temporary=True` creates temporary tables that the database drops
when the connection closes — the closest match to the other
engines' register-for-this-session behavior. Where the driver cannot
create them (see the table below), registration raises rather than
risk a permanent table: Trino's driver, for one, silently ignores
the request.
- `ingest_options={...}` sets driver-specific options on each ingest
statement — e.g. Spark's staging area,
`{"spark.ingest.staging_area_uri": "s3://bucket/path"}`.
- **Commit after registering.** Ingest runs inside the connection's
current transaction, as DB-API prescribes. The tables are visible to
this connection at once, but other connections — a BI tool, a
separate reader — see nothing until you call `con.commit()` (SQL
Server even blocks them on the lock), and `con.rollback()` discards
the tables. DuckDB's driver autocommits; MySQL commits DDL itself.

Mixed-dimension Datasets are ingested into a database schema named
after the Dataset, so `era5.surface` is the same SQL here as on
DataFusion and DuckDB. An existing schema is used as is, so a role
granted only that schema can register into it. On databases without
schemas (SQLite), and for temporary tables, the groups are created as
flat `era5_surface` tables instead (with a warning in the first case).
Where creating the schema fails *and* the failure aborts the
transaction (PostgreSQL without the `CREATE` privilege), registration
raises instead: call `con.rollback()`, then create the schema
beforehand or pass `temporary=True`.

**ClickHouse.** ClickHouse's
[ADBC driver](https://adbc-drivers.org/drivers/clickhouse/) (a preview
at the time of writing) can only append, so on ClickHouse the adapter
creates each table itself and then appends to it:

```python
from adbc_driver_manager import dbapi

# dbc install clickhouse
con = dbapi.connect(
driver="clickhouse",
db_kwargs={"uri": "http://localhost:8123/?user=default&password=..."},
)
xql.register(con, "era5", ds)
```

Pass credentials as URI query parameters, as above; credentials in the
URI's user-info part are not used. Driver 0.1.1 works with ClickHouse
26.8 but fails every query against 26.9 (`decompression error: incorrect
magic number`). [chDB](https://clickhouse.com/docs/chdb), ClickHouse
embedded in-process, takes the same path with no server:
`dbc install chdb`, then `driver="chdb"` and `uri="chdb://"`.

Tables are `MergeTree` sorted by their dimensions
(`ORDER BY (time, latitude, longitude)`), so ClickHouse's primary index
skips data on dimension filters much as chunk pruning does elsewhere.
Timestamps are declared `DateTime64(p, 'UTC')`, with the precision `p`
following the coordinate's resolution (9 for `datetime64[ns]`, 6 for
`datetime64[us]`), so a literal like `time >= '2020-01-01'` means UTC
rather than the server's local zone.
Mixed-dimension Datasets go into a ClickHouse *database* named after
the Dataset (`era5.surface`), and `temporary=True` creates `Memory`
tables. To choose the engine or sort key yourself, create the table
first and register with `mode="append"`. ClickHouse has no
transactions, so `mode="replace"` ingests into a staging table and
swaps it in with `EXCHANGE TABLES` only once the ingest succeeds; a
failed replace leaves the old table as it was.

**Tested databases.** Databases differ in how they quote identifiers,
whether they have schemas, which types they store, and what their
drivers support; the adapter keeps those facts in one table of
dialects and has a single code path. The test suite runs the same
contract — round-trips, every mode, temporary tables, mixed-dimension
naming, missing values, timedelta coordinates, the chunked round-trip —
against each database it can reach:

| Database | `name.group` as | Temporary tables | Notes |
|---|---|---|---|
| SQLite | flat `name_group` (no schemas) | yes | times stored as text (`2021-01-01 04:00:00…`), so plain literals compare correctly; timedeltas and unsigned integers as integers |
| DuckDB | schema | yes | |
| PostgreSQL | schema | yes | tables are `ANALYZE`d after ingest; a failed statement aborts the transaction; times to the microsecond; mixed-case names need quotes |
| MySQL, MariaDB | database | yes | backtick identifiers; the driver ignores the target schema, so the adapter switches the default database for the ingest; times to the microsecond; MariaDB joins without hash joins by default (see limitations) |
| ClickHouse, chDB | database | yes (`Memory`) | tables created by the adapter (above) |
| DataFusion | schema | no | mixed-case names need quotes |
| Trino | schema | no | ingest is slow, about 10k rows/s (see limitations) |
| SQL Server | schema | yes (queried as `#name`) | timedeltas and unsigned integers as integers; times to the microsecond |

Spark, BigQuery, Databricks, and Snowflake follow their drivers'
published feature tables (no temporary tables; backtick identifiers in
Spark, BigQuery, and Databricks; no target schema in Spark, whose
groups are flat) but are not exercised by the test suite; other
databases get standard SQL.

Where a database lacks a type, the adapter converts on the way in
rather than let the driver lose data silently. Unsigned integers widen
to the next signed width where there are none (PostgreSQL would
otherwise wrap a `uint64` above the int64 range around to a negative
number); a `uint64` too large for int64 raises `ValueError` instead.
Registering a time coordinate with sub-microsecond values on a database
that stores microseconds warns, and a name the database folds
(`Weather` on PostgreSQL) warns that it must be quoted in queries.
Where a database widens a type — SQLite
stores `float32` as `float64` and `bool` as an integer, MySQL `bool` as
`int8`, interval or text durations — `to_dataset` narrows a plain
`SELECT` back to the template's type; derived values such as an `AVG`
keep the result's type.

The cursor is a one-shot Arrow stream: `xql.to_dataset(cur, ...)`
round-trips eagerly, and `chunks=` needs `spill=True`.

## Engine support matrix

What each integration provides. Known issues and constraints live on
[Known issues & limitations](limitations.md).

| | DataFusion | DuckDB | Polars |
|---|---|---|---|
| Register | `XarrayContext` / any `SessionContext` | `xql.register(con, name, ds)` | `pl.scan_pyarrow_dataset(xql.arrow_dataset(ds))` |
| Projection pushdown | yes | yes | yes |
| Chunk pruning on dim predicates | yes | yes | yes |
| Eager round-trip (`xql.to_dataset`) | yes | yes | yes |
| Chunked round-trip (`chunks=`) | re-execution | `spill=True` [^spill-only] | re-execution (streaming engine) |
| `geometry` column ([geospatial](geospatial.md#geoarrow-point-geometry-columns)) | annotated WKB passes through | native `GEOMETRY` (`"wkb"` encoding) | plain binary/struct |
| Mixed-dimension datasets | one schema, `name.group` tables | `name.group` views over `name_group` tables | `xql.arrow_datasets(ds, name)`, one per group |
| Naming those tables (`table_names=`) | yes | yes | yes |
| Version floor | bundled (core dependency) | `duckdb >= 1.4` (tested on 1.5) | tested on `polars 1.42` |
| | DataFusion | DuckDB | Polars | ADBC |
|---|---|---|---|---|
| Register | `XarrayContext` / any `SessionContext` | `xql.register(con, name, ds)` | `pl.scan_pyarrow_dataset(xql.arrow_dataset(ds))` | `xql.register(con, name, ds)` (copies into the database) |
| Projection pushdown | yes | yes | yes | n/a (the database's own tables) |
| Chunk pruning on dim predicates | yes | yes | yes | n/a (the database's own indexes) |
| Eager round-trip (`xql.to_dataset`) | yes | yes | yes | yes (pass the cursor) |
| Chunked round-trip (`chunks=`) | re-execution | `spill=True` [^spill-only] | re-execution (streaming engine) | `spill=True` |
| `geometry` column ([geospatial](geospatial.md#geoarrow-point-geometry-columns)) | annotated WKB passes through | native `GEOMETRY` (`"wkb"` encoding) | plain binary/struct | driver-dependent |
| Mixed-dimension datasets | one schema, `name.group` tables | `name.group` views over `name_group` tables | `xql.arrow_datasets(ds, name)`, one per group | `name.group` tables in a schema; `name_group` without schemas |
| Naming those tables (`table_names=`) | yes | yes | yes | yes |
| Version floor | bundled (core dependency) | `duckdb >= 1.4` (tested on 1.5) | tested on `polars 1.42` | `adbc-driver-manager >= 1.12` (see [tested databases](#adbc-adapter-any-database-with-a-driver)) |

[^spill-only]: Why DuckDB relations do not re-execute — and two other
engine-specific issues worth knowing — is explained on
Expand Down
Loading
Loading