Skip to content

Declare all platform-managed columns through one mechanism, in DataJoint notation #1567

Description

@dimitri-yatsenko

Platform-managed columns were injected three different ways. Unifying them on the DataJoint type system — and letting them be written in DataJoint notation — removes two adapter methods, fixes a live cross-backend defect (#1566), and makes a new backend inherit every platform column for free.

Implemented in #1565, for v2.3.4. The sections below are what that PR does; the original proposal is unchanged in substance.

Where things stood

column how it was declared type system? :type: comment?
_singleton core_type_to_sql + format_column_definition yes yes
_prov same yes yes
_job_start_time, _job_duration, _job_version hand-written SQL per adapter (adapter.job_metadata_columns()) no no

_singleton was the right precedent all along; _job_* was the remaining holdout.

What the holdout cost

Verified against real servers, not inferred:

  • A cross-backend defect. _job_start_time was datetime(3) on MySQL and timestamp without time zone on PostgreSQL — _job_start_time precision differs by backend: datetime(3) on MySQL, microsecond timestamp on PostgreSQL #1566. The type system already produced the right thing (timestamp(3)); the hand-written string is what lost it.
  • original_type was None for all three on both backends. Cosmetic only: nothing reads it for a hidden attribute, and decoding derives json/uuid/numeric/dtype from the SQL type (heading.py:499-503). It comes free with the type-system path rather than being a reason on its own.
  • Two adapter methods to keep in sync, and a third backend would have needed a third copy of each.

Grammar separated from policy

The definition language could not express a hidden attribute, for two independent reasons:

  1. attribute_parser's name rule required a leading lowercase letter, so _prov = null : json raised ParseException before any check ran.
  2. compile_attribute rejected a leading underscore outright.

Both were policy expressed at the wrong level. The grammar can now spell a hidden attribute; what stays forbidden is a user declaring one.

Grammar — one character class:

attribute_name = pp.Word(pp.srange("[a-z]"),  ...)     # was
attribute_name = pp.Word(pp.srange("[a-z_]"), ...)     # now

Policy — the underscore check moved out of compile_attribute, which the framework itself calls, into _reject_user_hidden_attributes, called from declare() before anything is appended. That is the single point a user's definition enters the system. No allow_reserved flag to thread through and forget, and a user declaring _foo : int still gets the same error.

One registry, in the notation the rest of the schema uses

declare.py now holds all five, and _append_platform_attributes splices them into the definition string before it is parsed:

_singleton      = 1    : bool          # singleton primary key
_job_start_time = null : datetime(3)   # when computation began
_job_duration   = null : float32       # computation duration in seconds
_job_version    = ''   : varchar(64)   # code version
_prov           = null : json          # extrinsic provenance for a row that entered from outside

Each is parsed by the normal machinery, so every platform column gets backend type mapping, a :type: comment, and original_type — with no Python construction and no adapter method. job_metadata_columns and provenance_columns are both gone. A new adapter inherits all five by implementing core_type_to_sql, which it must implement anyway.

The two retrofit paths go through the same compile_attribute call, so a column added to an existing table is indistinguishable in the catalog from a declared one:

  • deploy.add_prov_column for _prov
  • migrate.add_job_metadata_columns for _job_*, which also emitted backtick-quoted identifiers and so had never run on PostgreSQL at all

Constraints, and how each is met

  • The user-facing ban keeps working, with the same error. tests/unit/test_declare_hidden_attribute.py pins the message and now also pins that the guard fires from a definition string rather than only from compile_attribute directly, since the check moved.
  • describe() keeps omitting hidden attributes, so the describe → prepare_declare → alter round trip holds.
  • Existing tables are unaffected; only new declarations change.
  • tests/integration/test_migrate_job_metadata.py declares the same Computed table twice — once with the columns, once without them and then migrated — and asserts the catalog cannot tell them apart on either backend.

Sequencing

Merged into v2.3.4 rather than deferred: the parser change is one character class, and the guard it relocates is pinned by tests that moved with it. #1566 closes with it.

#1562 — a public way to read hidden attributes — remains open for 2.4. The two work the same ground from opposite sides: this one makes hidden attributes uniformly declarable, that one makes them readable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementIndicates new improvements

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions