From c098091ce52304b63b2129c44d829b9a23f6eaf4 Mon Sep 17 00:00:00 2001 From: Dimitri Yatsenko Date: Wed, 30 Sep 2026 15:50:18 -0500 Subject: [PATCH 1/5] docs: what's new in 2.3.4 Adds the 2.3.4 section to the 2.3 release notes, covering the hidden `_prov` attribute and the two provenance settings, and names the release's shape in the page intro. Also corrects a forward reference the 2.3.3 notes carried: the codec `context=` parameter that replaces the `key["_config"]` threading was said to be scheduled for 2.3.4, but #1550 now sits on the 2.3.5 milestone. --- src/about/whats-new-23.md | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/src/about/whats-new-23.md b/src/about/whats-new-23.md index 1cd3915a..f4d744be 100644 --- a/src/about/whats-new-23.md +++ b/src/about/whats-new-23.md @@ -2,7 +2,7 @@ DataJoint 2.3 adds a first-class **upstream read surface** — `Diagram.trace` and `self.upstream` — which make "a computed row derives only from its declared upstream inputs" easy to follow inside `make()` and easy to query afterward. It also ships the **SparkAdapter Codec Protocol** for typed rendering to Spark-native types, **`dj.deploy.set_replica_identity`** for PostgreSQL change-data-capture, and a **cascade fix** for Part-of-Part and renamed-foreign-key chains. -Later releases on the line added **redesigned diagram rendering** on the DataJoint brand palette, with a light/dark [`display.diagram_theme`](../reference/configuration.md#display-settings) setting, and **S3 stores that resolve an ambient AWS identity** instead of requiring static keys. +Later releases on the line added **redesigned diagram rendering** on the DataJoint brand palette, with a light/dark [`display.diagram_theme`](../reference/configuration.md#display-settings) setting, **S3 stores that resolve an ambient AWS identity** instead of requiring static keys, and **extrinsic provenance at the pipeline's entry points** — a hidden `_prov` attribute the framework records on every row arriving from outside. > **Upgrading from 2.0, 2.1, or 2.2?** No API breaks — every feature on the 2.3 line is additive. > Two fixes in 2.3.3 do, however, reject input that previously passed silently: a misspelled native @@ -13,6 +13,13 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi > **Citation:** Yatsenko D, Nguyen TT. *DataJoint 2.0: A Computational Substrate for Agentic Scientific Workflows.* arXiv:2602.16585. 2026. [doi:10.48550/arXiv.2602.16585](https://doi.org/10.48550/arXiv.2602.16585) +## Changes in 2.3.4 + +2.3.4 is a patch release on the 2.3 line. If you are upgrading from **2.3.3**: + +- **Rows that enter the pipeline record where they came from.** Every `Manual` table now carries a hidden `_prov` attribute, declared by default and filled on insert. Inside a pipeline provenance is structural — a computed row cannot exist without its declared upstream, so the foreign-key graph is the lineage — but at the boundary that runs out, and until now each pipeline invented its own record of external origin. `_prov` gives the boundary one shape, so that "which rows have no recorded origin" is a query rather than an audit. **No author writes it:** `insert()` takes no provenance argument, and the content comes from deployment configuration (`provenance.source`), from the connection (user, host, database, time, code version), and — for an insert running inside a `make()` — from the ingesting table and key. That last source is what makes the [fan-out ingestion pattern](../explanation/fan-out-ingestion.md) traceable without a foreign key. A field the operator can set is weaker evidence than one the system sets, which is the point for an audit; anything an author wants to record deliberately belongs in the data model as a visible attribute, where queries can reach it. Tables declared before 2.3.4 have no column and record nothing until `dj.deploy.add_prov_column` adds it. See [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md), [Record Data Origin](../how-to/record-data-origin.md) and [#1547](https://github.com/datajoint/datajoint-python/issues/1547). +- **New settings: `provenance.capture` and `provenance.source`.** Capture defaults to **on** — a slot absent from most tables is one nothing can rely on. `provenance.source` names the external system a process draws from, set per deployment through `DJ_PROVENANCE_SOURCE`, the config file, or the secrets directory rather than in pipeline code. See [Configuration](../reference/configuration.md#provenance-settings). + ## Changes in 2.3.3 2.3.3 is a patch release on the 2.3 line. If you are upgrading from **2.3.2**: @@ -25,7 +32,7 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi - **Attribute types are validated at declaration and insert.** Four related gaps closed. A misspelled native type (`int24`, `intbanana`) is now rejected with `Unsupported attribute type` instead of being passed to the server as invalid DDL. `decimal(M,D) unsigned` is accepted again — it was rejected in 2.x while the equivalent `numeric(M,D) unsigned` passed. A bare `blob` is accepted alongside `tinyblob`/`longblob`. And inserting a NumPy array into a **native** `blob` attribute now raises instead of silently storing the array's text representation; use a `` codec attribute to store arrays. See [#1527](https://github.com/datajoint/datajoint-python/issues/1527), [#1528](https://github.com/datajoint/datajoint-python/issues/1528), [#1529](https://github.com/datajoint/datajoint-python/issues/1529) and [#1530](https://github.com/datajoint/datajoint-python/issues/1530). - **File-protocol stores are safe on Windows.** Paths in `file://` store URLs are now built with POSIX separators on every platform. Previously a Windows backslash in a stored path did not match the forward-slash form used during reference discovery, so `gc.collect()` could classify live files as orphans and delete them. Windows is now covered by CI. See [Clean Up Object Storage](../how-to/garbage-collection.md) and [#1520](https://github.com/datajoint/datajoint-python/issues/1520). - **Foreign-key columns are indexed on PostgreSQL.** Table declaration emitted an index for unique foreign keys only, leaving ordinary ones unindexed and making joins and cascading deletes scan. Non-unique foreign-key columns are now indexed, skipping any already covered by an existing index. Existing tables are unaffected until redeclared. See [#1512](https://github.com/datajoint/datajoint-python/issues/1512). -- **Custom codecs must resolve stores against the calling connection.** The `SchemaCodec` example previously omitted `config=` when calling `_build_path` and `_get_backend`, so a codec written from it resolved its store against the module-level `dj.config` rather than the connection actually in use — reading the wrong store, or failing validation, in any process holding more than one connection. The example now threads `key["_config"]` through both, as the built-in `object` and `npy` codecs already did. If you maintain a codec, check both call sites. This threading is scheduled to be replaced by an explicit `context=` parameter in 2.3.4 — see [#1550](https://github.com/datajoint/datajoint-python/issues/1550) — so the underscore key will keep working with a deprecation warning rather than changing under you. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb). +- **Custom codecs must resolve stores against the calling connection.** The `SchemaCodec` example previously omitted `config=` when calling `_build_path` and `_get_backend`, so a codec written from it resolved its store against the module-level `dj.config` rather than the connection actually in use — reading the wrong store, or failing validation, in any process holding more than one connection. The example now threads `key["_config"]` through both, as the built-in `object` and `npy` codecs already did. If you maintain a codec, check both call sites. This threading is scheduled to be replaced by an explicit `context=` parameter in 2.3.5 — see [#1550](https://github.com/datajoint/datajoint-python/issues/1550) — so the underscore key will keep working with a deprecation warning rather than changing under you. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb). ## Changes in 2.3.2 From 31db3f2295f2b5c7c8fa5a3f6e30f459470a57ea Mon Sep 17 00:00:00 2001 From: Dimitri Yatsenko Date: Wed, 30 Sep 2026 16:45:36 -0500 Subject: [PATCH 2/5] docs: 2.3.4 release notes cover the codec context argument #1550 moves onto the 2.3.4 milestone, so the 2.3.3 forward reference is correct again and the release gains a bullet of its own. The bullet leads on what the argument is for and states plainly that codecs written before 2.3.4 keep working, since that is the first question a codec author will have. --- src/about/whats-new-23.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/src/about/whats-new-23.md b/src/about/whats-new-23.md index f4d744be..e40564bd 100644 --- a/src/about/whats-new-23.md +++ b/src/about/whats-new-23.md @@ -19,6 +19,7 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi - **Rows that enter the pipeline record where they came from.** Every `Manual` table now carries a hidden `_prov` attribute, declared by default and filled on insert. Inside a pipeline provenance is structural — a computed row cannot exist without its declared upstream, so the foreign-key graph is the lineage — but at the boundary that runs out, and until now each pipeline invented its own record of external origin. `_prov` gives the boundary one shape, so that "which rows have no recorded origin" is a query rather than an audit. **No author writes it:** `insert()` takes no provenance argument, and the content comes from deployment configuration (`provenance.source`), from the connection (user, host, database, time, code version), and — for an insert running inside a `make()` — from the ingesting table and key. That last source is what makes the [fan-out ingestion pattern](../explanation/fan-out-ingestion.md) traceable without a foreign key. A field the operator can set is weaker evidence than one the system sets, which is the point for an audit; anything an author wants to record deliberately belongs in the data model as a visible attribute, where queries can reach it. Tables declared before 2.3.4 have no column and record nothing until `dj.deploy.add_prov_column` adds it. See [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md), [Record Data Origin](../how-to/record-data-origin.md) and [#1547](https://github.com/datajoint/datajoint-python/issues/1547). - **New settings: `provenance.capture` and `provenance.source`.** Capture defaults to **on** — a slot absent from most tables is one nothing can rely on. `provenance.source` names the external system a process draws from, set per deployment through `DJ_PROVENANCE_SOURCE`, the config file, or the secrets directory rather than in pipeline code. See [Configuration](../reference/configuration.md#provenance-settings). +- **Codecs receive connection context explicitly.** `encode()` and `decode()` now take a `context` argument carrying `schema`, `table`, `field` and `config`, so `key` means what it means everywhere else in DataJoint — the primary key. Passing the calling connection's `config` matters: `_build_path` and `_get_backend` fall back to the global `dj.config` without it, which in a process holding connections for several users belongs to none of them and resolves a different store silently rather than raising. Two third-party codecs hit exactly that, both following the `SchemaCodec` docstring, which is updated here. **Nothing breaks:** DataJoint passes `context` only to codecs whose signature declares it, the pre-2.3.4 underscore keys are still populated in `key`, and `_extract_context(key)` still accepts one argument — it warns only when it falls back to those keys. Codecs written before 2.3.4 keep working unchanged; `key` reverting to primary-key-only, and `config` becoming required, wait for 2.4. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb) and [#1550](https://github.com/datajoint/datajoint-python/issues/1550). ## Changes in 2.3.3 @@ -32,7 +33,7 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi - **Attribute types are validated at declaration and insert.** Four related gaps closed. A misspelled native type (`int24`, `intbanana`) is now rejected with `Unsupported attribute type` instead of being passed to the server as invalid DDL. `decimal(M,D) unsigned` is accepted again — it was rejected in 2.x while the equivalent `numeric(M,D) unsigned` passed. A bare `blob` is accepted alongside `tinyblob`/`longblob`. And inserting a NumPy array into a **native** `blob` attribute now raises instead of silently storing the array's text representation; use a `` codec attribute to store arrays. See [#1527](https://github.com/datajoint/datajoint-python/issues/1527), [#1528](https://github.com/datajoint/datajoint-python/issues/1528), [#1529](https://github.com/datajoint/datajoint-python/issues/1529) and [#1530](https://github.com/datajoint/datajoint-python/issues/1530). - **File-protocol stores are safe on Windows.** Paths in `file://` store URLs are now built with POSIX separators on every platform. Previously a Windows backslash in a stored path did not match the forward-slash form used during reference discovery, so `gc.collect()` could classify live files as orphans and delete them. Windows is now covered by CI. See [Clean Up Object Storage](../how-to/garbage-collection.md) and [#1520](https://github.com/datajoint/datajoint-python/issues/1520). - **Foreign-key columns are indexed on PostgreSQL.** Table declaration emitted an index for unique foreign keys only, leaving ordinary ones unindexed and making joins and cascading deletes scan. Non-unique foreign-key columns are now indexed, skipping any already covered by an existing index. Existing tables are unaffected until redeclared. See [#1512](https://github.com/datajoint/datajoint-python/issues/1512). -- **Custom codecs must resolve stores against the calling connection.** The `SchemaCodec` example previously omitted `config=` when calling `_build_path` and `_get_backend`, so a codec written from it resolved its store against the module-level `dj.config` rather than the connection actually in use — reading the wrong store, or failing validation, in any process holding more than one connection. The example now threads `key["_config"]` through both, as the built-in `object` and `npy` codecs already did. If you maintain a codec, check both call sites. This threading is scheduled to be replaced by an explicit `context=` parameter in 2.3.5 — see [#1550](https://github.com/datajoint/datajoint-python/issues/1550) — so the underscore key will keep working with a deprecation warning rather than changing under you. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb). +- **Custom codecs must resolve stores against the calling connection.** The `SchemaCodec` example previously omitted `config=` when calling `_build_path` and `_get_backend`, so a codec written from it resolved its store against the module-level `dj.config` rather than the connection actually in use — reading the wrong store, or failing validation, in any process holding more than one connection. The example now threads `key["_config"]` through both, as the built-in `object` and `npy` codecs already did. If you maintain a codec, check both call sites. This threading is scheduled to be replaced by an explicit `context=` parameter in 2.3.4 — see [#1550](https://github.com/datajoint/datajoint-python/issues/1550) — so the underscore key will keep working with a deprecation warning rather than changing under you. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb). ## Changes in 2.3.2 From 3dbcdefd4b04822d8b0c656e2ab3e3a4ad75b7cb Mon Sep 17 00:00:00 2001 From: Dimitri Yatsenko Date: Thu, 1 Oct 2026 17:11:49 -0500 Subject: [PATCH 3/5] docs: 2.3.4 release notes cover the tier names and the platform columns Two user-visible changes in the milestone were missing from the notes. `dj.Entry`, `dj.Ingest` and `dj.Compute` are permanent aliases, verified as the same classes rather than subclasses, so nothing about `describe()` or the diagram changes. The platform-column refactor is a fix, not only tidying: on PostgreSQL `_job_start_time` had microsecond precision where MySQL had millisecond, and the job-metadata retrofit had never run there at all. --- src/about/whats-new-23.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/src/about/whats-new-23.md b/src/about/whats-new-23.md index e40564bd..ae8be4dd 100644 --- a/src/about/whats-new-23.md +++ b/src/about/whats-new-23.md @@ -19,6 +19,8 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi - **Rows that enter the pipeline record where they came from.** Every `Manual` table now carries a hidden `_prov` attribute, declared by default and filled on insert. Inside a pipeline provenance is structural — a computed row cannot exist without its declared upstream, so the foreign-key graph is the lineage — but at the boundary that runs out, and until now each pipeline invented its own record of external origin. `_prov` gives the boundary one shape, so that "which rows have no recorded origin" is a query rather than an audit. **No author writes it:** `insert()` takes no provenance argument, and the content comes from deployment configuration (`provenance.source`), from the connection (user, host, database, time, code version), and — for an insert running inside a `make()` — from the ingesting table and key. That last source is what makes the [fan-out ingestion pattern](../explanation/fan-out-ingestion.md) traceable without a foreign key. A field the operator can set is weaker evidence than one the system sets, which is the point for an audit; anything an author wants to record deliberately belongs in the data model as a visible attribute, where queries can reach it. Tables declared before 2.3.4 have no column and record nothing until `dj.deploy.add_prov_column` adds it. See [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md), [Record Data Origin](../how-to/record-data-origin.md) and [#1547](https://github.com/datajoint/datajoint-python/issues/1547). - **New settings: `provenance.capture` and `provenance.source`.** Capture defaults to **on** — a slot absent from most tables is one nothing can rely on. `provenance.source` names the external system a process draws from, set per deployment through `DJ_PROVENANCE_SOURCE`, the config file, or the secrets directory rather than in pipeline code. See [Configuration](../reference/configuration.md#provenance-settings). +- **`dj.Entry`, `dj.Ingest` and `dj.Compute` name the tiers on one axis.** The three tiers that differ by *how a row arrives* now read that way: `dj.Entry` for a row entered from outside, `dj.Ingest` for one loaded by `make()` from an external source, `dj.Compute` for one derived from upstream rows. `dj.Manual`, `dj.Imported` and `dj.Computed` are permanent aliases, not deprecations — the old names declare the same tables, `describe()` and the diagram are unchanged, and no pipeline needs editing. `dj.Lookup` and `dj.Part` keep their names, which already say what they are. See [Relational Workflow Model](../explanation/relational-workflow-model.md) and [#1546](https://github.com/datajoint/datajoint-python/issues/1546). +- **Every platform-managed column goes through the DataJoint type system.** `_job_start_time`, `_job_duration`, `_job_version`, `_prov` and `_singleton` are declared in DataJoint notation and compiled by the same path as a user attribute, so each backend's mapping comes from one place. This fixes two PostgreSQL defects: `_job_start_time` was declared as a bare `timestamp` and so took microsecond precision where MySQL got `datetime(3)` ([#1566](https://github.com/datajoint/datajoint-python/issues/1566)), and `migrate.add_job_metadata_columns` built its `ALTER` with backtick-quoted identifiers, which is a syntax error there — the retrofit had never run on PostgreSQL. Existing tables are untouched; only new declarations and retrofits change. See [#1567](https://github.com/datajoint/datajoint-python/issues/1567). - **Codecs receive connection context explicitly.** `encode()` and `decode()` now take a `context` argument carrying `schema`, `table`, `field` and `config`, so `key` means what it means everywhere else in DataJoint — the primary key. Passing the calling connection's `config` matters: `_build_path` and `_get_backend` fall back to the global `dj.config` without it, which in a process holding connections for several users belongs to none of them and resolves a different store silently rather than raising. Two third-party codecs hit exactly that, both following the `SchemaCodec` docstring, which is updated here. **Nothing breaks:** DataJoint passes `context` only to codecs whose signature declares it, the pre-2.3.4 underscore keys are still populated in `key`, and `_extract_context(key)` still accepts one argument — it warns only when it falls back to those keys. Codecs written before 2.3.4 keep working unchanged; `key` reverting to primary-key-only, and `config` becoming required, wait for 2.4. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb) and [#1550](https://github.com/datajoint/datajoint-python/issues/1550). ## Changes in 2.3.3 From ee13c106d33b0c10ba34bc83d912bb464e83d81c Mon Sep 17 00:00:00 2001 From: Dimitri Yatsenko Date: Fri, 2 Oct 2026 10:20:30 -0500 Subject: [PATCH 4/5] docs: 2.3.4 adds the names; adopting them is 2.4 The bullet read as though the rename had happened. What ships in 2.3.4 is the second set of names (#1558); making them primary is #1546, slated for 2.4, and the docs keep the original names until then. Also dropped "which already say what they are" of Lookup and Part -- a comment on the other names by implication, and the release notes have no reason to make it. --- src/about/whats-new-23.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/src/about/whats-new-23.md b/src/about/whats-new-23.md index ae8be4dd..2f33d5f9 100644 --- a/src/about/whats-new-23.md +++ b/src/about/whats-new-23.md @@ -19,7 +19,7 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi - **Rows that enter the pipeline record where they came from.** Every `Manual` table now carries a hidden `_prov` attribute, declared by default and filled on insert. Inside a pipeline provenance is structural — a computed row cannot exist without its declared upstream, so the foreign-key graph is the lineage — but at the boundary that runs out, and until now each pipeline invented its own record of external origin. `_prov` gives the boundary one shape, so that "which rows have no recorded origin" is a query rather than an audit. **No author writes it:** `insert()` takes no provenance argument, and the content comes from deployment configuration (`provenance.source`), from the connection (user, host, database, time, code version), and — for an insert running inside a `make()` — from the ingesting table and key. That last source is what makes the [fan-out ingestion pattern](../explanation/fan-out-ingestion.md) traceable without a foreign key. A field the operator can set is weaker evidence than one the system sets, which is the point for an audit; anything an author wants to record deliberately belongs in the data model as a visible attribute, where queries can reach it. Tables declared before 2.3.4 have no column and record nothing until `dj.deploy.add_prov_column` adds it. See [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md), [Record Data Origin](../how-to/record-data-origin.md) and [#1547](https://github.com/datajoint/datajoint-python/issues/1547). - **New settings: `provenance.capture` and `provenance.source`.** Capture defaults to **on** — a slot absent from most tables is one nothing can rely on. `provenance.source` names the external system a process draws from, set per deployment through `DJ_PROVENANCE_SOURCE`, the config file, or the secrets directory rather than in pipeline code. See [Configuration](../reference/configuration.md#provenance-settings). -- **`dj.Entry`, `dj.Ingest` and `dj.Compute` name the tiers on one axis.** The three tiers that differ by *how a row arrives* now read that way: `dj.Entry` for a row entered from outside, `dj.Ingest` for one loaded by `make()` from an external source, `dj.Compute` for one derived from upstream rows. `dj.Manual`, `dj.Imported` and `dj.Computed` are permanent aliases, not deprecations — the old names declare the same tables, `describe()` and the diagram are unchanged, and no pipeline needs editing. `dj.Lookup` and `dj.Part` keep their names, which already say what they are. See [Relational Workflow Model](../explanation/relational-workflow-model.md) and [#1546](https://github.com/datajoint/datajoint-python/issues/1546). +- **Three tiers gain a second name.** `dj.Entry`, `dj.Ingest` and `dj.Compute` join `dj.Manual`, `dj.Imported` and `dj.Computed`. Each pair is one class rather than a subclass, so a table declared either way is identical — same SQL prefix, same tier detection, same `describe()` and diagram — and both names are permanent, with no deprecation and nothing to migrate. `dj.Lookup` and `dj.Part` are unchanged. The documentation keeps the original names for now; adopting the new ones as primary is slated for 2.4 ([#1546](https://github.com/datajoint/datajoint-python/issues/1546)). See [Relational Workflow Model](../explanation/relational-workflow-model.md) and [#1546](https://github.com/datajoint/datajoint-python/issues/1546). - **Every platform-managed column goes through the DataJoint type system.** `_job_start_time`, `_job_duration`, `_job_version`, `_prov` and `_singleton` are declared in DataJoint notation and compiled by the same path as a user attribute, so each backend's mapping comes from one place. This fixes two PostgreSQL defects: `_job_start_time` was declared as a bare `timestamp` and so took microsecond precision where MySQL got `datetime(3)` ([#1566](https://github.com/datajoint/datajoint-python/issues/1566)), and `migrate.add_job_metadata_columns` built its `ALTER` with backtick-quoted identifiers, which is a syntax error there — the retrofit had never run on PostgreSQL. Existing tables are untouched; only new declarations and retrofits change. See [#1567](https://github.com/datajoint/datajoint-python/issues/1567). - **Codecs receive connection context explicitly.** `encode()` and `decode()` now take a `context` argument carrying `schema`, `table`, `field` and `config`, so `key` means what it means everywhere else in DataJoint — the primary key. Passing the calling connection's `config` matters: `_build_path` and `_get_backend` fall back to the global `dj.config` without it, which in a process holding connections for several users belongs to none of them and resolves a different store silently rather than raising. Two third-party codecs hit exactly that, both following the `SchemaCodec` docstring, which is updated here. **Nothing breaks:** DataJoint passes `context` only to codecs whose signature declares it, the pre-2.3.4 underscore keys are still populated in `key`, and `_extract_context(key)` still accepts one argument — it warns only when it falls back to those keys. Codecs written before 2.3.4 keep working unchanged; `key` reverting to primary-key-only, and `config` becoming required, wait for 2.4. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb) and [#1550](https://github.com/datajoint/datajoint-python/issues/1550). From a0917922ab5ee2879893ca9ea239d9006b124f71 Mon Sep 17 00:00:00 2001 From: Dimitri Yatsenko Date: Fri, 2 Oct 2026 17:31:31 -0500 Subject: [PATCH 5/5] docs: the 2.3.4 notes said capture defaults on Follows datajoint/datajoint-python#1568. The notes promised a column on every Manual table and argued the case for defaulting on, both of which are now wrong. Capture is a deployment's decision, and the upgrade note is the stronger line: an unchanged schema declares under 2.3.4 exactly what it declared under 2.3.3. --- src/about/whats-new-23.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/src/about/whats-new-23.md b/src/about/whats-new-23.md index 2f33d5f9..b174bd18 100644 --- a/src/about/whats-new-23.md +++ b/src/about/whats-new-23.md @@ -17,8 +17,8 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi 2.3.4 is a patch release on the 2.3 line. If you are upgrading from **2.3.3**: -- **Rows that enter the pipeline record where they came from.** Every `Manual` table now carries a hidden `_prov` attribute, declared by default and filled on insert. Inside a pipeline provenance is structural — a computed row cannot exist without its declared upstream, so the foreign-key graph is the lineage — but at the boundary that runs out, and until now each pipeline invented its own record of external origin. `_prov` gives the boundary one shape, so that "which rows have no recorded origin" is a query rather than an audit. **No author writes it:** `insert()` takes no provenance argument, and the content comes from deployment configuration (`provenance.source`), from the connection (user, host, database, time, code version), and — for an insert running inside a `make()` — from the ingesting table and key. That last source is what makes the [fan-out ingestion pattern](../explanation/fan-out-ingestion.md) traceable without a foreign key. A field the operator can set is weaker evidence than one the system sets, which is the point for an audit; anything an author wants to record deliberately belongs in the data model as a visible attribute, where queries can reach it. Tables declared before 2.3.4 have no column and record nothing until `dj.deploy.add_prov_column` adds it. See [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md), [Record Data Origin](../how-to/record-data-origin.md) and [#1547](https://github.com/datajoint/datajoint-python/issues/1547). -- **New settings: `provenance.capture` and `provenance.source`.** Capture defaults to **on** — a slot absent from most tables is one nothing can rely on. `provenance.source` names the external system a process draws from, set per deployment through `DJ_PROVENANCE_SOURCE`, the config file, or the secrets directory rather than in pipeline code. See [Configuration](../reference/configuration.md#provenance-settings). +- **Rows that enter the pipeline record where they came from.** A `Manual` table can carry a hidden `_prov` attribute, declared when a deployment enables `provenance.capture` and filled on insert. Inside a pipeline provenance is structural — a computed row cannot exist without its declared upstream, so the foreign-key graph is the lineage — but at the boundary that runs out, and until now each pipeline invented its own record of external origin. `_prov` gives the boundary one shape, so that "which rows have no recorded origin" is a query rather than an audit. **No author writes it:** `insert()` takes no provenance argument, and the content comes from deployment configuration (`provenance.source`), from the connection (user, host, database, time, code version), and — for an insert running inside a `make()` — from the ingesting table and key. That last source is what makes the [fan-out ingestion pattern](../explanation/fan-out-ingestion.md) traceable without a foreign key. A field the operator can set is weaker evidence than one the system sets, which is the point for an audit; anything an author wants to record deliberately belongs in the data model as a visible attribute, where queries can reach it. A table declared before capture was enabled has no column and records nothing until `dj.deploy.add_prov_column` adds it. See [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md), [Record Data Origin](../how-to/record-data-origin.md) and [#1547](https://github.com/datajoint/datajoint-python/issues/1547). +- **New settings: `provenance.capture` and `provenance.source`.** Capture defaults to **off**, like `jobs.add_job_metadata` and for the same reason: it changes the DDL of every `Manual` table declared afterwards, so upgrading to 2.3.4 leaves an unchanged schema declaring exactly what it declared under 2.3.3. `provenance.source` names the external system a process draws from, set per deployment through `DJ_PROVENANCE_SOURCE`, the config file, or the secrets directory rather than in pipeline code. See [Configuration](../reference/configuration.md#provenance-settings). - **Three tiers gain a second name.** `dj.Entry`, `dj.Ingest` and `dj.Compute` join `dj.Manual`, `dj.Imported` and `dj.Computed`. Each pair is one class rather than a subclass, so a table declared either way is identical — same SQL prefix, same tier detection, same `describe()` and diagram — and both names are permanent, with no deprecation and nothing to migrate. `dj.Lookup` and `dj.Part` are unchanged. The documentation keeps the original names for now; adopting the new ones as primary is slated for 2.4 ([#1546](https://github.com/datajoint/datajoint-python/issues/1546)). See [Relational Workflow Model](../explanation/relational-workflow-model.md) and [#1546](https://github.com/datajoint/datajoint-python/issues/1546). - **Every platform-managed column goes through the DataJoint type system.** `_job_start_time`, `_job_duration`, `_job_version`, `_prov` and `_singleton` are declared in DataJoint notation and compiled by the same path as a user attribute, so each backend's mapping comes from one place. This fixes two PostgreSQL defects: `_job_start_time` was declared as a bare `timestamp` and so took microsecond precision where MySQL got `datetime(3)` ([#1566](https://github.com/datajoint/datajoint-python/issues/1566)), and `migrate.add_job_metadata_columns` built its `ALTER` with backtick-quoted identifiers, which is a syntax error there — the retrofit had never run on PostgreSQL. Existing tables are untouched; only new declarations and retrofits change. See [#1567](https://github.com/datajoint/datajoint-python/issues/1567). - **Codecs receive connection context explicitly.** `encode()` and `decode()` now take a `context` argument carrying `schema`, `table`, `field` and `config`, so `key` means what it means everywhere else in DataJoint — the primary key. Passing the calling connection's `config` matters: `_build_path` and `_get_backend` fall back to the global `dj.config` without it, which in a process holding connections for several users belongs to none of them and resolves a different store silently rather than raising. Two third-party codecs hit exactly that, both following the `SchemaCodec` docstring, which is updated here. **Nothing breaks:** DataJoint passes `context` only to codecs whose signature declares it, the pre-2.3.4 underscore keys are still populated in `key`, and `_extract_context(key)` still accepts one argument — it warns only when it falls back to those keys. Codecs written before 2.3.4 keep working unchanged; `key` reverting to primary-key-only, and `config` becoming required, wait for 2.4. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb) and [#1550](https://github.com/datajoint/datajoint-python/issues/1550).