Skip to content

Hidden _prov attribute for extrinsic provenance at Entry tables - #1555

Merged
dimitri-yatsenko merged 4 commits into
masterfrom
feat/entry-provenance
Oct 1, 2026
Merged

dimitri-yatsenko merged 4 commits into
masterfrom
feat/entry-provenance

Conversation

@dimitri-yatsenko

Copy link
Copy Markdown
Member

Closes #1547. Implements the settled design. Documentation is tracked in datajoint/datajoint-docs#284, which this unblocks — the reference spec lands there as reference/specs/boundary-provenance.md.

What this adds

A hidden _prov JSON attribute on Entry (dj.Manual) tables, declared by default and filled on insert.

No author ever writes it. insert gains no argument, and the content comes from three places, none of them the call site:

Source Supplies
config.provenance.source the deployment constant — which external system this process draws from
Ambient connection state conn_info user/host/database, insert time, _get_job_version
Ambient execution state the ingesting table and key, when the insert runs inside a make()

The third is what makes the fan-out pattern traceable: a row written into an Entry table from inside an ingesting make() records what wrote it, with no foreign key back.

That ownership is the point rather than a limitation. A field an operator can set is weaker evidence than one the system sets, which is what ALCOA+ attributability wants — and it retires the adoption risk an author-supplied field carries, since there is nothing left for a pipeline to neglect. Anything an author wants to record deliberately belongs in the model as a visible attribute, where queries can reach it; hidden attributes are excluded from query composition by design.

Entry tables only

Compute and Ingest get nothing. Their provenance is entailed by the foreign-key graph, and config.jobs.add_job_metadata already records agent, time and version for them — a _prov there would record the same facts twice while missing the half that matters (which external resource a make() actually read, which is per-row and belongs with consumed-input capture). A part inherits its master's.

Capture defaults on

add_job_metadata defaults to False, which is why migrate.add_job_metadata_columns exists. Repeating that would leave no consumer able to assume the column exists, and would make "which Entry rows have no recorded origin" conditional on each table's declaration-time config rather than a query. Stated plainly: an Entry table declared under this change differs in DDL from one declared before it. The difference is additive and hidden.

Files

File Change
provenance.py (new) payload assembly, tier test, the ingesting context var
settings.py ProvenanceSettings, env_prefix="DJ_PROVENANCE_", as config.provenance
declare.py the column on Entry tables at declaration
adapters/{base,mysql,postgres}.py provenance_columns() — json / jsonb
table.py _attach_provenance on the insert path
autopopulate.py scope the ingesting context to make(), released in the existing finally
deploy.py add_prov_column retrofit

On the retrofit's home. It sits in deploy rather than migrate because migrate is deprecated for removal in 2.4/2.5, while every existing deployment needs this retrofit and will still need it afterwards whenever capture has been off. It is idempotent by construction, which is deploy's stated contract.

Verification

  • Full integration suite: 679 passed, 14 skipped, 0 failed, across MySQL and PostgreSQL — the regression that mattered, since _prov now lands on every Entry table by default.
  • 403 unit tests pass, including 16 new ones.
  • 6 new integration tests: tier selection, insert without author involvement, exclusion from heading/fetch/join, rejection of an author-supplied _prov, the fan-out write recording the ingesting key, and the capture-off → add_prov_column → retrofit path including that pre-existing rows keep NULL.
  • ruff, ruff-format, codespell and mypy pass via pre-commit; mypy caught two real type errors that are fixed here.
  • Verified DJ_PROVENANCE_SOURCE (JSON) and DJ_PROVENANCE_CAPTURE inject through the environment, which is the Platform's per-project path.

The diff is purely additive apart from one docstring line in deploy.py.

Left open deliberately

  • Changing configuration mid-process silently shifts what later rows record. A note for the spec, not a mechanism.
  • Range queries on capture time go through JSON extraction; a deployment adds a generated column and indexes it if that becomes hot. Field filtering itself is already portable — condition.py translates JSON paths to json_value() on MySQL and jsonb_extract_path_text() on PostgreSQL.
  • Whether allow_direct_insert should require _prov stays out of scope: that is enforcement, which is a Platform concern.

@dimitri-yatsenko

Copy link
Copy Markdown
Member Author

CI note: the red integration jobs are infrastructure, not this branch.

All eight matrix jobs report 1052 passed, 33 errors, 0 failed. Every one of the 33 errors is a collection error in an object-storage test — test_gc, test_npy_codec, test_codecs, test_attach — and each traces to the same cause:

docker.errors.ImageNotFound: 404 Client Error ... No such image: minio/minio:latest

The MinIO testcontainer image could not be pulled, so those suites never ran. Nothing in this branch touches object storage, and no test failed.

The provenance tests passed in CI as part of that 1052, and the full suite ran clean locally: 679 passed, 14 skipped, 0 failed across MySQL and PostgreSQL.

Re-running the jobs once the image is pullable should clear it.

@ttngu207

Copy link
Copy Markdown
Contributor

Thanks @dimitri-yatsenko, reviewing, should we address #1553 as well in this PR?

@dimitri-yatsenko

Copy link
Copy Markdown
Member Author

Correction to my CI note above. I said re-running the jobs once the image is pullable should clear it. That was wrong — it will not clear on its own.

MinIO deleted minio/minio from Docker Hub on 2026-09-11 (the whole repository, every tag, so pinning does not help), and quay.io/minio/minio began requiring authentication on 2026-09-24. Free community builds stopped in October 2025 and the open-source repo was archived in February 2026. It is a withdrawal, not an outage.

Fix is #1559 — switches the fixture to Chainguard's public build, verified with a cold pull: 118 passed across the five object-storage suites. Merge that first and these go green.

@ttngu207

Copy link
Copy Markdown
Contributor

The shape of this is right — Entry-only, framework-owned, and capture on by default. That last one especially: defaulting off would have made this another setting nobody turns on, and the add_job_metadata precedent shows how that goes.

Two things will fire on merge, though.

Jobs tables get _prov

is_entry_table is a prefix test, and ~ passes both clauses:

>>> is_entry_table("~~analysis")
True

The column really lands, and every job row gets a payload:

tables: ['__analysis', 'src', '~lineage', '~~analysis']
  '~~analysis'  has _prov: True
  '~~analysis'  rows total / with _prov: (2, 2)

(verified on fb4e15a against MySQL 8.0 — it takes a jobs.refresh() to materialise the table, so a plain populate() won't show it. ~lineage is fine; it's created by raw DDL rather than declare().)

On a 100k-key queue that's a lot of JSON nobody asked for. And since reserve/complete run inside _populate1's try, I think a nested populate would stamp the inner job rows with the outer table's context — haven't verified that part.

Gating on Manual.tier_regexp rather than the prefix would fix it. Cheapest possible test: add ("~~analysis", False) to the is_entry_table parametrization — no database needed.

An unserializable source breaks every insert

source is dict[str, Any], pydantic accepts anything, and serialize is a bare json.dumps:

>>> dj.config.provenance.source = {"system": "PyRat", "pulled_at": datetime.date(2026, 1, 1)}   # accepted
>>> Src.insert1({'id': 99})
TypeError: Object of type date is not JSON serializable

Every insert into every Entry table in the process, with an error that never mentions provenance. _jsonable covers the ingesting key but not this. A default=str in serialize plus a validator on the field would close it — and wrapping the payload work in _attach_provenance would make "must never break an insert" actually total, rather than just the version call.

Smaller things

  • insert(QueryExpression) records nothing. That branch builds INSERT … SELECT and returns before _insert_rows, so those rows land with _prov NULL (verified). Copying rows in from another schema is about as boundary-crossing as an insert gets, so it's worth either closing or calling out in the docstring.
  • git rev-parse per insert and per _populate1, unmemoized — roughly 1.9 ms vs 0.003 ms per build_payload. 100k inserts is ~190 s of fork/exec. autopopulate.py:693 also makes it unconditional, where the old path only paid it on success.
  • add_prov_column leaves the cached heading stale. I think inserts keep writing NULL until the process reconnects — the manual _init_from_database() in the test is what hides it.
  • provenance_columns as a new @abstractmethod makes any out-of-tree adapter uninstantiable. A default return [] on the base would avoid it.
  • settings.py:329 and table.py:969 both say datajoint.migrate.add_prov_column — it's in deploy, which the deploy docstring argues for at length. The settings one is a pydantic description, so it shows up in config help.
  • Dry-run counters increment outside the if not dry_run guard, unlike set_replica_identity right next door.
  • Threads. A ThreadPoolExecutor fan-out inside a make() loses context silently. Multiprocessing is fine — fork, and each worker sets and resets its own.

One that isn't yours to fix but is now twice as sharp: _prov is hidden, so there's no public way to read it back, same as _job_* (#1553). Every test here reaches for raw SQL. Whatever lands for that should probably cover both.

dimitri-yatsenko added a commit that referenced this pull request Oct 1, 2026
…insert

Both from review of #1555.

Job tables were getting `_prov`. `is_entry_table` excluded the other tiers'
prefixes one at a time -- `_`, `#`, and `__` for parts -- and `~` was not among
them, so `~~analysis`, `~jobs` and `~lineage` all passed. The column really
landed and every job row carried a payload; on a large queue that is a lot of
JSON nobody asked for. It only surfaced after `jobs.refresh()` materialises the
table, which is why a plain populate() in the first round of tests missed it.

Now matched against `Manual.tier_regexp`. Enumerating what to exclude makes
every tier added later an Entry table until someone remembers this function;
matching the tier definition inverts that, so a name is an Entry table only if
the library says it is.

A `source` that json could not render broke every insert.
`config.provenance.source` is deployment-supplied and typed `dict[str, Any]`,
so a `date` in it raised `TypeError: Object of type date is not JSON
serializable` from inside every insert into every Entry table, naming neither
provenance nor the setting responsible. Three layers close it:

- `serialize` passes `default=str`, so ordinary values a deployment would
  actually set -- dates, paths -- record correctly rather than failing;
- a validator on the field rejects what remains (non-string keys, cycles) at
  assignment, where the error belongs;
- `_attach_provenance` catches and logs, so recording where a row came from can
  never stop the row being written. That guarantee previously covered only the
  version call.

Also fixes the `datajoint.migrate.add_prov_column` reference in the settings
description -- it lives in `deploy`, and that string shows up in config help.

Regression tests fail against the previous code: 4 of them, including the
integration one that drives a real job queue.
@dimitri-yatsenko

Copy link
Copy Markdown
Member Author

@ttngu207 Both blockers fixed in fa51011 — thank you, these were real and I'd missed both.

Job tables

Confirmed exactly as you described: ~~analysis, ~jobs and ~lineage all passed is_entry_table, and the column really landed.

Taken your suggestion and gated on Manual.tier_regexp. The underlying mistake was enumerating what to exclude — that makes every tier added later an Entry table until someone remembers this function, which is precisely how ~ got in. Matching the tier definition inverts it: a name is an Entry table only if the library says so.

Your parametrization case is in, plus ~jobs and ~lineage, and an integration test that drives a real job queue — jobs.refresh() first, as you noted, since populate() alone never materialises the table. Verified the tests fail against the previous code: 4 failures including the integration one, so they are not vacuous.

Unserializable source

Also confirmed. Closed in three layers rather than one, because default=str alone silently stringifies and a validator alone does not cover everything:

  • serialize passes default=str, so values a deployment would realistically set — dates, paths — record correctly ("2026-01-01") rather than failing;
  • a validator on the field rejects what remains, non-string keys and cycles, at assignment, where the error names the setting;
  • _attach_provenance catches and logs, so the "must never break an insert" guarantee is now total rather than covering just the version call — your point.

Also fixed the stale datajoint.migrate.add_prov_column in the settings description. Good catch that it is a pydantic description and surfaces in config help.

Full suite: 680 passed, 14 skipped, 0 failed.

On the rest

All verified, none disputed. Holding them out of this PR deliberately rather than dismissing them:

  • insert(QueryExpression) returns at table.py:101 before _insert_rows — confirmed. Agreed it is boundary-crossing and should be closed properly rather than at speed.
  • Unmemoized git rev-parse, and I made it unconditional in autopopulate.py where the old path only paid it on success. Nearly free to fix, but it touches the populate hot path.
  • Stale heading after add_prov_column — you are right that the manual _init_from_database() in my test is what hides it.
  • provenance_columns as @abstractmethod breaking out-of-tree adapters — a default return [] is the fix.
  • Dry-run counters outside the guard, inconsistent with set_replica_identity next door.
  • Threads — correct, contextvars do not propagate into ThreadPoolExecutor workers. Documentation, not a fix.

I will file these as a single follow-up unless you would rather see any of them in this PR.

#1553

Agreed it is not this PR's to fix, and agreed the pressure doubles — every test here reaches for raw SQL for exactly that reason. A public read path should cover _job_* and _prov together; folding a new public API into a release-eve PR is the wrong trade.

dimitri-yatsenko added a commit that referenced this pull request Oct 1, 2026
…T ... SELECT

Three follow-ups from review of #1555.

The Entry-table test is now inlined at its two call sites rather than wrapped in
`provenance.is_entry_table`. The import stays deferred inside the function:
user_tables imports table, which imports declare, so a module-scope import is a
cycle -- which is what the wrapper had been hiding.

`insert(QueryExpression)` builds INSERT ... SELECT and returned before the
provenance was attached, so copied rows landed with NULL. It now carries the
source's `_prov` across. A copied row did not originate in the destination, so
the source's record is the true one; re-stamping it here would claim an origin
that is not where the data came from.

`heading.as_sql` resolves an explicitly named hidden attribute against the full
attribute set. Default field lists are built from `attributes` and still never
contain hidden names, so nothing else changes.

681 passed, 14 skipped.
Inside the pipeline provenance is structural: a Compute row cannot exist
unless its declared upstream does, so the foreign-key graph is the lineage.
At the boundary that runs out. Rows arrive in Entry tables from a person, an
instrument, or a feed, and until now each pipeline invented its own record of
where they came from -- a source_file column here, a notes varchar there, or
nothing at all.

Adds a hidden `_prov` JSON attribute to Entry (dj.Manual) tables, declared by
default and filled on insert.

The attribute is framework-owned: no author ever writes it. `insert` gains no
argument, and the content comes from three places, none of them the call site:

- configuration -- `config.provenance.source`, set per deployment
- ambient connection state -- user, host, database, insert time, code version
- ambient execution state -- the ingesting table and key, inside a `make()`

That third source is what makes a fan-out write traceable: a row written into
an Entry table from inside an ingesting `make()` records what wrote it without
any foreign key. Ownership is the point -- a field an operator can set is
weaker evidence than one the system sets, and nothing is left for a pipeline
to neglect. Anything an author wants to record deliberately belongs in the
model as a visible attribute, where queries can reach it.

Entry tables only. Compute and Ingest have no use for the slot -- their
provenance is entailed by the graph, and job metadata already records agent,
time and version for them -- while a part inherits its master's.

Capture defaults on. Off by default would leave no consumer able to assume the
column exists and would turn "which Entry rows have no recorded origin" into a
question conditional on each table's declaration-time config.

`deploy.add_prov_column` adds the slot to tables declared earlier. It sits in
deploy rather than migrate because it is idempotent and stays useful for as
long as capture can be turned off, which outlives the migration module's
scheduled removal.
…insert

Both from review of #1555.

Job tables were getting `_prov`. `is_entry_table` excluded the other tiers'
prefixes one at a time -- `_`, `#`, and `__` for parts -- and `~` was not among
them, so `~~analysis`, `~jobs` and `~lineage` all passed. The column really
landed and every job row carried a payload; on a large queue that is a lot of
JSON nobody asked for. It only surfaced after `jobs.refresh()` materialises the
table, which is why a plain populate() in the first round of tests missed it.

Now matched against `Manual.tier_regexp`. Enumerating what to exclude makes
every tier added later an Entry table until someone remembers this function;
matching the tier definition inverts that, so a name is an Entry table only if
the library says it is.

A `source` that json could not render broke every insert.
`config.provenance.source` is deployment-supplied and typed `dict[str, Any]`,
so a `date` in it raised `TypeError: Object of type date is not JSON
serializable` from inside every insert into every Entry table, naming neither
provenance nor the setting responsible. Three layers close it:

- `serialize` passes `default=str`, so ordinary values a deployment would
  actually set -- dates, paths -- record correctly rather than failing;
- a validator on the field rejects what remains (non-string keys, cycles) at
  assignment, where the error belongs;
- `_attach_provenance` catches and logs, so recording where a row came from can
  never stop the row being written. That guarantee previously covered only the
  version call.

Also fixes the `datajoint.migrate.add_prov_column` reference in the settings
description -- it lives in `deploy`, and that string shows up in config help.

Regression tests fail against the previous code: 4 of them, including the
integration one that drives a real job queue.
…T ... SELECT

Three follow-ups from review of #1555.

The Entry-table test is now inlined at its two call sites rather than wrapped in
`provenance.is_entry_table`. The import stays deferred inside the function:
user_tables imports table, which imports declare, so a module-scope import is a
cycle -- which is what the wrapper had been hiding.

`insert(QueryExpression)` builds INSERT ... SELECT and returned before the
provenance was attached, so copied rows landed with NULL. It now carries the
source's `_prov` across. A copied row did not originate in the destination, so
the source's record is the true one; re-stamping it here would claim an origin
that is not where the data came from.

`heading.as_sql` resolves an explicitly named hidden attribute against the full
attribute set. Default field lists are built from `attributes` and still never
contain hidden names, so nothing else changes.

681 passed, 14 skipped.
The provenance serialization test asserted the rendered Path equalled
'/mnt/raw'. On Windows str(Path('/mnt/raw')) is '\\mnt\\raw', so the
assertion failed on both Windows jobs while passing everywhere else.

The behavior under test is unchanged and was always correct: default=str
stringifies a Path rather than raising. Only the expectation was platform-
specific.
@ttngu207

ttngu207 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Re-checked on 4702d24 — the ~ names all fail Manual.tier_regexp now and job tables come out clean, a date in source records as a string instead of taking the insert down, and INSERT … SELECT carries _prov across. Matching the tier definition instead of enumerating exclusions is the right call.

Two worth carrying forward, neither holding this up:

  • add_prov_column leaves the cached heading stale, so inserts in the same process keep writing NULL until it reconnects — and that window never gets backfilled. A line on the docstring would cover it.
  • git rev-parse per insert and per _populate1 is still unmemoized; worth a follow-up issue so it doesn't evaporate.

LGTM.

@dimitri-yatsenko
dimitri-yatsenko merged commit 6a3adc9 into master Oct 1, 2026
17 checks passed
@dimitri-yatsenko
dimitri-yatsenko deleted the feat/entry-provenance branch October 1, 2026 18:21
@dimitri-yatsenko dimitri-yatsenko added the enhancement Indicates new improvements label Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement Indicates new improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a hidden _prov attribute for extrinsic provenance at pipeline boundaries

2 participants