Repository navigation
Admit verified local SmolVLA datasets through normal preflight - #347
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Local corrected-dataset artifacts could not reach Tether's normal training preflight: the CLI accepted only an HF dataset ID/revision, and metadata checks always used Hub reads. This adds explicit --dataset-root and --dataset-manifest-sha256 inputs for one bounded lane: pretrained SmolVLA LoRA using a completed verified-moments-v1 export.
A stdlib verifier binds the canonical completion receipt, complete file inventory and actual hashes/sizes, v3 feature/statistics shapes, selected development split/provenance, exact declared producer pins and explicit SmolVLA normalization evidence. Missing local files and malformed receipts fail before metadata lookup; they never silently substitute Hub dataset metadata during Tether admission. Existing schema/padding and episode-count checks read local info.json, verify real selected-base metadata, and remain blocking for local input. Command construction revalidates before adding dataset.root.
Local requests reject skip-preflight, resume, revision ambiguity, unsupported model/training lanes and extra-argument escapes that could change dataset selection, output ownership, normalization, features or LoRA method. Existing remote-only flows and positional FinetuneConfig callers remain compatible. ImageNet default behavior is preserved and reported separately from explicit SmolVLA VISUAL=IDENTITY.
SmolVLA is the first lane because it already has meaningful base/action compatibility preflight. ACT's from-scratch path lacks that contract and remains outside this change. Studio's companion draft resolves workspace-owned export IDs, verifies source/split/runtime identity and supplies a separate job-owned snapshot; this public verifier alone is not an authenticity signature or filesystem race-proofing mechanism. Downstream LeRobot retains its own fallback behavior if callers later mutate/remove files, so the input must remain immutable through use.
Validation: 172 focused tests passed, including 111 new admission tests and 61 existing finetune/preflight/resume checks. Ruff and compilation passed. Independent cross-repository review has no unresolved findings for this offline admission scope. New public fixtures deliberately use opaque synthetic Parquet slots and synthetic base metadata; they do not establish loader or training qualification. Companion Studio verification in FastCrest/tether-studio#134 uses a real exported SmolVLA artifact and passed 23 offline checks plus one targeted recovery regression, stopping before trainer dispatch.
The original top-level tether.cli registration test also passed separately after its missing exact-candidate source/import closure was materialized and all 19 required Git blobs were verified. That check made zero network/subprocess attempts or model-runtime imports; no code or dependency edits were needed. The 172 tests were reused. The edited finetune Typer command also passed a dry-run with both new options. Full package/doctor and training acceptance are not claimed. GitHub Actions were not inspected or used.
Base main: 18157ec. Reviewed composed candidate: 7606360 (tree 2d4cfa4fc01538a0858c0d97935658b8ef640cdc). It includes the exact CUDA evaluation fix from Studio pin 731779f as a second parent: local_runner passes verification_device="cuda" and the regression asserts it. Both files match that pin byte-for-byte. Seven affected offline evaluation tests passed with zero network/subprocess attempts or model/provider imports; the unchanged admission/CLI evidence above is reused. Independent review confirms the full tree preserves both source lines and the artifact-identity primitive remains byte-identical. Studio's runtime-pin update is a separate dependency step.
No model/checkpoint download, training, GPU/provider call, HF upload, credentials or production data. This source integration does not qualify actual training, platform packaging or customer acceptance.