diff --git a/src/about/whats-new-23.md b/src/about/whats-new-23.md index 1cd3915a..b174bd18 100644 --- a/src/about/whats-new-23.md +++ b/src/about/whats-new-23.md @@ -2,7 +2,7 @@ DataJoint 2.3 adds a first-class **upstream read surface** — `Diagram.trace` and `self.upstream` — which make "a computed row derives only from its declared upstream inputs" easy to follow inside `make()` and easy to query afterward. It also ships the **SparkAdapter Codec Protocol** for typed rendering to Spark-native types, **`dj.deploy.set_replica_identity`** for PostgreSQL change-data-capture, and a **cascade fix** for Part-of-Part and renamed-foreign-key chains. -Later releases on the line added **redesigned diagram rendering** on the DataJoint brand palette, with a light/dark [`display.diagram_theme`](../reference/configuration.md#display-settings) setting, and **S3 stores that resolve an ambient AWS identity** instead of requiring static keys. +Later releases on the line added **redesigned diagram rendering** on the DataJoint brand palette, with a light/dark [`display.diagram_theme`](../reference/configuration.md#display-settings) setting, **S3 stores that resolve an ambient AWS identity** instead of requiring static keys, and **extrinsic provenance at the pipeline's entry points** — a hidden `_prov` attribute the framework records on every row arriving from outside. > **Upgrading from 2.0, 2.1, or 2.2?** No API breaks — every feature on the 2.3 line is additive. > Two fixes in 2.3.3 do, however, reject input that previously passed silently: a misspelled native @@ -13,6 +13,16 @@ Later releases on the line added **redesigned diagram rendering** on the DataJoi > **Citation:** Yatsenko D, Nguyen TT. *DataJoint 2.0: A Computational Substrate for Agentic Scientific Workflows.* arXiv:2602.16585. 2026. [doi:10.48550/arXiv.2602.16585](https://doi.org/10.48550/arXiv.2602.16585) +## Changes in 2.3.4 + +2.3.4 is a patch release on the 2.3 line. If you are upgrading from **2.3.3**: + +- **Rows that enter the pipeline record where they came from.** A `Manual` table can carry a hidden `_prov` attribute, declared when a deployment enables `provenance.capture` and filled on insert. Inside a pipeline provenance is structural — a computed row cannot exist without its declared upstream, so the foreign-key graph is the lineage — but at the boundary that runs out, and until now each pipeline invented its own record of external origin. `_prov` gives the boundary one shape, so that "which rows have no recorded origin" is a query rather than an audit. **No author writes it:** `insert()` takes no provenance argument, and the content comes from deployment configuration (`provenance.source`), from the connection (user, host, database, time, code version), and — for an insert running inside a `make()` — from the ingesting table and key. That last source is what makes the [fan-out ingestion pattern](../explanation/fan-out-ingestion.md) traceable without a foreign key. A field the operator can set is weaker evidence than one the system sets, which is the point for an audit; anything an author wants to record deliberately belongs in the data model as a visible attribute, where queries can reach it. A table declared before capture was enabled has no column and records nothing until `dj.deploy.add_prov_column` adds it. See [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md), [Record Data Origin](../how-to/record-data-origin.md) and [#1547](https://github.com/datajoint/datajoint-python/issues/1547). +- **New settings: `provenance.capture` and `provenance.source`.** Capture defaults to **off**, like `jobs.add_job_metadata` and for the same reason: it changes the DDL of every `Manual` table declared afterwards, so upgrading to 2.3.4 leaves an unchanged schema declaring exactly what it declared under 2.3.3. `provenance.source` names the external system a process draws from, set per deployment through `DJ_PROVENANCE_SOURCE`, the config file, or the secrets directory rather than in pipeline code. See [Configuration](../reference/configuration.md#provenance-settings). +- **Three tiers gain a second name.** `dj.Entry`, `dj.Ingest` and `dj.Compute` join `dj.Manual`, `dj.Imported` and `dj.Computed`. Each pair is one class rather than a subclass, so a table declared either way is identical — same SQL prefix, same tier detection, same `describe()` and diagram — and both names are permanent, with no deprecation and nothing to migrate. `dj.Lookup` and `dj.Part` are unchanged. The documentation keeps the original names for now; adopting the new ones as primary is slated for 2.4 ([#1546](https://github.com/datajoint/datajoint-python/issues/1546)). See [Relational Workflow Model](../explanation/relational-workflow-model.md) and [#1546](https://github.com/datajoint/datajoint-python/issues/1546). +- **Every platform-managed column goes through the DataJoint type system.** `_job_start_time`, `_job_duration`, `_job_version`, `_prov` and `_singleton` are declared in DataJoint notation and compiled by the same path as a user attribute, so each backend's mapping comes from one place. This fixes two PostgreSQL defects: `_job_start_time` was declared as a bare `timestamp` and so took microsecond precision where MySQL got `datetime(3)` ([#1566](https://github.com/datajoint/datajoint-python/issues/1566)), and `migrate.add_job_metadata_columns` built its `ALTER` with backtick-quoted identifiers, which is a syntax error there — the retrofit had never run on PostgreSQL. Existing tables are untouched; only new declarations and retrofits change. See [#1567](https://github.com/datajoint/datajoint-python/issues/1567). +- **Codecs receive connection context explicitly.** `encode()` and `decode()` now take a `context` argument carrying `schema`, `table`, `field` and `config`, so `key` means what it means everywhere else in DataJoint — the primary key. Passing the calling connection's `config` matters: `_build_path` and `_get_backend` fall back to the global `dj.config` without it, which in a process holding connections for several users belongs to none of them and resolves a different store silently rather than raising. Two third-party codecs hit exactly that, both following the `SchemaCodec` docstring, which is updated here. **Nothing breaks:** DataJoint passes `context` only to codecs whose signature declares it, the pre-2.3.4 underscore keys are still populated in `key`, and `_extract_context(key)` still accepts one argument — it warns only when it falls back to those keys. Codecs written before 2.3.4 keep working unchanged; `key` reverting to primary-key-only, and `config` becoming required, wait for 2.4. See [Custom Codecs](../tutorials/advanced/custom-codecs.ipynb) and [#1550](https://github.com/datajoint/datajoint-python/issues/1550). + ## Changes in 2.3.3 2.3.3 is a patch release on the 2.3 line. If you are upgrading from **2.3.2**: