Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 11 additions & 1 deletion src/about/whats-new-23.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

DataJoint 2.3 adds a first-class **upstream read surface** — `Diagram.trace` and `self.upstream` — which make "a computed row derives only from its declared upstream inputs" easy to follow inside `make()` and easy to query afterward. It also ships the **SparkAdapter Codec Protocol** for typed rendering to Spark-native types, **`dj.deploy.set_replica_identity`** for PostgreSQL change-data-capture, and a **cascade fix** for Part-of-Part and renamed-foreign-key chains.

Later releases on the line added **redesigned diagram rendering** on the DataJoint brand palette, with a light/dark [`display.diagram_theme`](../reference/configuration.md#display-settings) setting, and **S3 stores that resolve an ambient AWS identity** instead of requiring static keys.
Later releases on the line added **redesigned diagram rendering** on the DataJoint brand palette, with a light/dark [`display.diagram_theme`](../reference/configuration.md#display-settings) setting, **S3 stores that resolve an ambient AWS identity** instead of requiring static keys, and **extrinsic provenance at the pipeline's entry points** — a hidden `_prov` attribute the framework records on every row arriving from outside.

> **Upgrading from 2.0, 2.1, or 2.2?** No API breaks — every feature on the 2.3 line is additive.
> Two fixes in 2.3.3 do, however, reject input that previously passed silently: a misspelled native
Expand All @@ -13,6 +13,16 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi

> **Citation:** Yatsenko D, Nguyen TT. *DataJoint 2.0: A Computational Substrate for Agentic Scientific Workflows.* arXiv:2602.16585. 2026. [doi:10.48550/arXiv.2602.16585](https://doi.org/10.48550/arXiv.2602.16585)

## Changes in 2.3.4

2.3.4 is a patch release on the 2.3 line. If you are upgrading from **2.3.3**:

- **Rows that enter the pipeline record where they came from.** Every `Manual` table now carries a hidden `_prov` attribute, declared by default and filled on insert. Inside a pipeline provenance is structural — a computed row cannot exist without its declared upstream, so the foreign-key graph is the lineage — but at the boundary that runs out, and until now each pipeline invented its own record of external origin. `_prov` gives the boundary one shape, so that "which rows have no recorded origin" is a query rather than an audit. **No author writes it:** `insert()` takes no provenance argument, and the content comes from deployment configuration (`provenance.source`), from the connection (user, host, database, time, code version), and — for an insert running inside a `make()` — from the ingesting table and key. That last source is what makes the [fan-out ingestion pattern](../explanation/fan-out-ingestion.md) traceable without a foreign key. A field the operator can set is weaker evidence than one the system sets, which is the point for an audit; anything an author wants to record deliberately belongs in the data model as a visible attribute, where queries can reach it. Tables declared before 2.3.4 have no column and record nothing until `dj.deploy.add_prov_column` adds it. See [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md), [Record Data Origin](../how-to/record-data-origin.md) and [#1547](https://github.com/datajoint/datajoint-python/issues/1547).
- **New settings: `provenance.capture` and `provenance.source`.** Capture defaults to **on** — a slot absent from most tables is one nothing can rely on. `provenance.source` names the external system a process draws from, set per deployment through `DJ_PROVENANCE_SOURCE`, the config file, or the secrets directory rather than in pipeline code. See [Configuration](../reference/configuration.md#provenance-settings).
- **`dj.Entry`, `dj.Ingest` and `dj.Compute` name the tiers on one axis.** The three tiers that differ by *how a row arrives* now read that way: `dj.Entry` for a row entered from outside, `dj.Ingest` for one loaded by `make()` from an external source, `dj.Compute` for one derived from upstream rows. `dj.Manual`, `dj.Imported` and `dj.Computed` are permanent aliases, not deprecations — the old names declare the same tables, `describe()` and the diagram are unchanged, and no pipeline needs editing. `dj.Lookup` and `dj.Part` keep their names, which already say what they are. See [Relational Workflow Model](../explanation/relational-workflow-model.md) and [#1546](https://github.com/datajoint/datajoint-python/issues/1546).
- **Every platform-managed column goes through the DataJoint type system.** `_job_start_time`, `_job_duration`, `_job_version`, `_prov` and `_singleton` are declared in DataJoint notation and compiled by the same path as a user attribute, so each backend's mapping comes from one place. This fixes two PostgreSQL defects: `_job_start_time` was declared as a bare `timestamp` and so took microsecond precision where MySQL got `datetime(3)` ([#1566](https://github.com/datajoint/datajoint-python/issues/1566)), and `migrate.add_job_metadata_columns` built its `ALTER` with backtick-quoted identifiers, which is a syntax error there — the retrofit had never run on PostgreSQL. Existing tables are untouched; only new declarations and retrofits change. See [#1567](https://github.com/datajoint/datajoint-python/issues/1567).
- **Codecs receive connection context explicitly.** `encode()` and `decode()` now take a `context` argument carrying `schema`, `table`, `field` and `config`, so `key` means what it means everywhere else in DataJoint — the primary key. Passing the calling connection's `config` matters: `_build_path` and `_get_backend` fall back to the global `dj.config` without it, which in a process holding connections for several users belongs to none of them and resolves a different store silently rather than raising. Two third-party codecs hit exactly that, both following the `SchemaCodec` docstring, which is updated here. **Nothing breaks:** DataJoint passes `context` only to codecs whose signature declares it, the pre-2.3.4 underscore keys are still populated in `key`, and `_extract_context(key)` still accepts one argument — it warns only when it falls back to those keys. Codecs written before 2.3.4 keep working unchanged; `key` reverting to primary-key-only, and `config` becoming required, wait for 2.4. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb) and [#1550](https://github.com/datajoint/datajoint-python/issues/1550).

## Changes in 2.3.3

2.3.3 is a patch release on the 2.3 line. If you are upgrading from **2.3.2**:
Expand Down
Loading