Skip to content

Document extrinsic provenance at entry tables - #288

Open
dimitri-yatsenko wants to merge 6 commits into
mainfrom
docs/entry-provenance
Open

dimitri-yatsenko wants to merge 6 commits into
mainfrom
docs/entry-provenance

Conversation

@dimitri-yatsenko

Copy link
Copy Markdown
Member

Closes #284. Documents the _prov attribute implemented in datajoint/datajoint-python#1555.

One property shapes all of this: the attribute is framework-owned and no author ever writes it. insert takes no provenance argument, and the content comes from deployment configuration, from the connection, and — inside a make() — from the ingesting table and key. So the feature's entire author-facing surface is configuration, and the documentation has to carry what the API does not.

New

reference/specs/boundary-provenance.md — modelled on specs/job-metadata.md. Covers the attribute and why it is one JSON column rather than typed ones; which tiers carry it, with the reasoning for Imported specifically (job metadata already records agent/time/version there, and the half that is missing is per-row and cannot come from configuration); the three content sources; why an author cannot write it and what to model instead; configuration including the environment and secrets forms; retrofitting via deploy.add_prov_column; and querying a hidden JSON attribute portably.

how-to/record-data-origin.md — the task-oriented half: set the source, confirm rows are carrying an origin, find the ones that are not, turn capture off, retrofit existing tables, and model a per-row link when _prov is the wrong tool for it.

examples/fan_out_provenance.py — the fan-out shape with the origin recorded, showing both the framework's record and the modelled source_file column doing their separate jobs.

Revised

  • reference/configuration.md — the config.provenance settings table.
  • explanation/fan-out-ingestion.md — the substantive change. "The responsibility it carries: record where the data came from" told authors to hand-roll the origin in every insert1. That responsibility has moved to configuration, so the section now separates the two jobs. The worked example keeps its source_file column: hidden attributes are excluded from query composition by design, so _prov is the audit record and a modelled column is the domain link — the page says why both exist rather than dropping either.
  • explanation/comparison-to-provenance-systems.md — the page already said entry-point tables record source identity; it now has a standard shape behind the claim. Its "opt-in and configured explicitly" framing for inbound/outbound integration is unchanged and still accurate — that integration is not what shipped.
  • tutorials/basics/03-data-entry.ipynb — entry is the boundary, so this is where a tutorial reader should learn what gets recorded and how to read it back.
  • mkdocs.yaml — Entry Provenance under Data Operations specs, Record Data Origin after Insert Data.

Naming

Per the decision recorded in #267 and datajoint/datajoint-python#1546, this material teaches dj.Manual and notes in a version-added admonition that dj.Entry is available as a permanent alias from 2.3.4. The 2.4 terminology sweep renames these pages with every other.

Two departures from the plan in #284

  • The retrofit helper is datajoint.deploy.add_prov_column, not a migrate function. datajoint.migrate is deprecated for removal in 2.4/2.5, while this retrofit stays useful for as long as capture can be turned off.
  • One _prov JSON attribute, not typed _prov_* columns — the key set is deliberately unsettled, and field filtering is portable across both backends anyway.

Verification

mkdocs build --strict clean; check_links.py passes across 144 built pages (up from 142), including the new tutorial → how-to cross-reference and the spec's links into job-metadata.md, type-system.md, and both explanation pages. check_notebook_versions.py clean. All six admonitions on the spec render.

The tutorial addition is a markdown cell only — verified programmatically that every code cell and output is byte-identical to HEAD, so no re-execution is needed. That also means nothing here executes against _prov: the notebooks install DataJoint from master, and datajoint/datajoint-python#1555 is not merged yet. Worth holding this PR until it is, so the docs and the library land together.

Documents the hidden `_prov` attribute added in datajoint-python 2.3.4. The
feature's entire author-facing surface is configuration — insert takes no
argument and no author can write the attribute — so the documentation carries
what the API does not.

New:
- reference/specs/boundary-provenance.md — the specification: the attribute,
  which tiers carry it and why Imported does not, the three content sources,
  the framework-ownership invariant, configuration, retrofitting through
  deploy.add_prov_column, and querying a hidden JSON attribute
- how-to/record-data-origin.md — configure the source, confirm rows are
  carrying an origin, find rows that are not, and what to model instead when
  a row needs something only it can say
- examples/fan_out_provenance.py — the fan-out shape with the origin recorded

Revised:
- reference/configuration.md — the config.provenance settings
- explanation/fan-out-ingestion.md — the section that told authors to record
  each row's origin by hand now separates the two jobs: the framework records
  the audit trail, and the author models the link the pipeline itself queries.
  The worked example keeps its source_file column, since hidden attributes are
  excluded from query composition and cannot replace a modelled one.
- explanation/comparison-to-provenance-systems.md — the claim that entry-point
  tables record source identity now has a standard shape behind it
- tutorials/basics/03-data-entry.ipynb — entry is the boundary, so the data
  entry tutorial says what is recorded there and how to read it back
- mkdocs.yaml — two nav entries

The tutorial addition is a markdown cell only; no code cells or outputs change,
so the notebook needs no re-execution.
The slot is granted by matching the Manual tier rather than by excluding the
other tiers' prefixes, so job queues and lineage tables are excluded by
construction. Worth stating: the first implementation enumerated prefixes and
job tables slipped through, which is the kind of thing a reader should be able
to check against the spec.
Three pages advertised a mapping restriction on `_prov`
(`& {"_prov.system": "PyRat"}`). That form is silently ignored on a hidden
attribute: no WHERE clause is emitted and every row comes back. Documenting it
was worse than documenting nothing, since the natural uses are audit questions
and the wrong answer looks like a right one. Replaced with the string form,
which works, plus a warning pointing at datajoint-python#1561.

Reading a hidden attribute back needs SQL until 2.4. job-metadata.md had
advertised `to_arrays('_job_start_time', ...)` as the access path since the
feature shipped; that call has never worked, because the heading excludes
hidden names and every accessor derives from it. The spec now says so and shows
the SQL, pointing at datajoint-python#1562 for the supported accessor.

Covers reference/specs/boundary-provenance.md, how-to/record-data-origin.md,
the data-entry tutorial, and reference/specs/job-metadata.md — the last being
the original claim behind datajoint-docs#1553.

Tutorial edit is a markdown cell; no code cells or outputs change.
The previous commit framed that ignoring as a defect. It is not: a mapping
restriction drops attributes it cannot match so that `Session & key` works when
`key` carries attributes from a more detailed downstream table. Passing a full
key dict down the graph is the normal idiom and `make()` depends on it.

The narrow point stands — a hidden attribute is invisible to that matching, so
naming one in a mapping drops the predicate and returns every row — but the
reason is the design working as intended on an attribute it cannot see, not a
bug in the ignoring. Reworded across all four pages so a reader does not come
away thinking unmatched attributes should raise.

The guidance is unchanged: write the condition as a string.
The recommended workaround used MySQL's JSON_VALUE, which fails on PostgreSQL
with UndefinedFunction — verified on postgres:15. Half the supported backends
were given a condition that does not run.

Both spellings are now shown, with the reason stated: DataJoint's own JSON-path
translation is what the mapping form would have provided, and that is exactly
the path a hidden attribute cannot take. IS NULL / IS NOT NULL remain portable.
The previous wording read as a general claim that filtering inside a JSON
attribute has no portable spelling. It does: {"data.system": "PyRat"} on an
ordinary json attribute returns the right rows on MySQL 8.0 and postgres:15
alike, because translate_attribute hands the path to adapter.json_path_expr.

The backend-specific SQL is needed here only because _prov is hidden and the
mapping form cannot reach a hidden attribute — which is datajoint-python#1561,
not a JSON defect. Reworded so a reader does not conclude JSON paths are
generally unportable.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Document extrinsic provenance capture at Entry tables (datajoint-python 2.3.4)

1 participant