You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Implement the September 21, 2026 PET performance, architecture, refactoring, and test-coverage audit in an evidence-driven order: fix demonstrated reliability failures first, make measurements reflect client behavior next, then simplify and bound concurrency. Do not start with a broad rewrite or an async-runtime migration.
This plan links 12 newly filed implementation issues, reuses #525 for the UTF-8 panic, and tracks #522 as an existing quality-workflow prerequisite.
Evidence and audit scope
The audit examined 4e523bad8bb9be8a84c01800a3e400a6a771cf8e. At filing, main is d586c60ec6b7191256fa3f47a2c866cc5739914e; the intervening change only removes redundant Conda vector drains and does not address these findings.
Windows default-feature validation: cargo test --workspace --offline --locked passed 647 tests, with 0 failures and 2 ignored documentation examples. Quality-tooling Python tests: 47 passed.
Bounded reproductions confirmed interpreter pipe backpressure, non-UTF-8 worker panic, EOF busy-looping, request-ID/framing defects, and refresh glob expansion blocking the dispatcher. Local timings are Windows debug diagnostics, not release budgets.
Exact-revision performance artifacts and coverage artifacts support the measurement/coverage findings. Source-level architectural risks and optimization opportunities are explicitly distinguished from reproduced failures in each ticket.
Existing work: reuse, do not duplicate
#525 / PR #526: existing non-UTF-8 interpreter-output fix. Review/validate and integrate it before overlapping subprocess-runner changes.
#522 / PR #524: existing fork-safe quality-workflow fix. Complete it so fork contributions execute the substantive quality gates. Runtime fixes do not need to wait for comment-publishing changes; never weaken the gates or grant untrusted PRs write credentials as a workaround.
Ordered implementation plan
Order below is the recommended landing sequence, not a requirement to serialize independent investigation or test preparation. P1/P2/P3 are relative priorities within this audit, not estimates of effort. The dependency column distinguishes required foundations from coordination.
String/large IDs round-trip; fragmented/extra-header/invalid/oversized frames have explicit outcomes
Phase exit: reproduced process/transport failures are covered by deterministic regressions; normal request behavior, streaming, and locator priority remain intact.
Phase 2 - Trustworthy measurements, responsiveness, and coverage
Order
Priority
Individual issue
Prerequisite / coordination
Exit evidence
5
P1
#531 - Gate client-observed latency and define TTFE boundaries
Can start alongside Phase 1; coordinate quality jobs with #522
Pre-discovery/queueing costs appear in gated metrics; schema transition and comparator tests preserve exact-base checks
6
P1
#535 - Move glob expansion off dispatch and bound traversal
Stable identity-based inventory plus client latency/resource measurements across increasing sizes and long-lived churn
Phase exit: the roughly 400 ms client / 2 ms reported-duration reproduction is no longer invisible to the performance gate; black-box server testing contributes usable coverage; repeatable workloads establish a baseline for architectural changes.
Phase 3 - Coherent ownership, bounded work, and focused refactoring
Order
Priority
Individual issue
Prerequisite / coordination
Exit evidence
9
P2
#536 - Give find/resolve coherent configuration snapshots
#533 regression harness and #531 metrics; the small find-lock fix may land earlier if isolated
Barrier-driven tests prove old-or-new, never mixed, configuration and no global lock over discovery I/O
10
P2
#539 - Bound scheduling and share failed in-flight probes
Behavior-preserving components replace duplicate module instances and tangled responsibilities without coverage/latency regression
Phase exit: ownership/resource limits are explicit and tested. Optimizations are supported by measurements. The locator framework and public behavior are preserved rather than replaced wholesale.
Parallel work and review boundaries
Transport shutdown/framing and subprocess-runner work can use separate PRs; coordinate the latter with #526. Metric/coverage test preparation can run alongside those fixes. Snapshot, scheduler, and writer changes should not be merged as one large refactor: establish ownership first, then independently validate scheduling and output behavior. Poetry profiling can run in parallel once the workload harness is ready.
Purely mechanical library-module reuse or the find read-guard fix may be isolated earlier; do not smuggle behavioral changes into those cleanups. Do not delay demonstrated reliability fixes for the optional Poetry optimization or final module cleanup.
Shared implementation and verification contract
Each ticket includes its own scope and acceptance tests. Implementations must preserve locator ordering, complete environment/manager information, platform path/symlink behavior, refresh coalescing, generation filtering, and scoped state synchronization. Use explicit errors, not success-shaped fallback data. Measure intentional metric/protocol changes and update directly related documentation.
Run targeted tests first, then relevant workspace/platform/feature jobs. Before committing Rust changes, run cargo fmt --all and cargo clippy --all -- -D warnings; CI's all-targets/all-features lint and the existing exact-base quality gates must also remain green. Concurrency tests should use deterministic barriers/channels rather than timing guesses. Test fixtures must bound and clean up their own subprocesses and files.
Close this tracking issue only when each linked item is complete or explicitly deferred with a recorded reason. Check issue/PR state when starting work; links above describe filing-time status, not a substitute for current GitHub state.
Goal
Implement the September 21, 2026 PET performance, architecture, refactoring, and test-coverage audit in an evidence-driven order: fix demonstrated reliability failures first, make measurements reflect client behavior next, then simplify and bound concurrency. Do not start with a broad rewrite or an async-runtime migration.
This plan links 12 newly filed implementation issues, reuses #525 for the UTF-8 panic, and tracks #522 as an existing quality-workflow prerequisite.
Evidence and audit scope
The audit examined
4e523bad8bb9be8a84c01800a3e400a6a771cf8e. At filing, main isd586c60ec6b7191256fa3f47a2c866cc5739914e; the intervening change only removes redundant Conda vector drains and does not address these findings.cargo test --workspace --offline --lockedpassed 647 tests, with 0 failures and 2 ignored documentation examples. Quality-tooling Python tests: 47 passed.Existing work: reuse, do not duplicate
#525 / PR #526: existing non-UTF-8 interpreter-output fix. Review/validate and integrate it before overlapping subprocess-runner changes.
#522 / PR #524: existing fork-safe quality-workflow fix. Complete it so fork contributions execute the substantive quality gates. Runtime fixes do not need to wait for comment-publishing changes; never weaken the gates or grant untrusted PRs write credentials as a workaround.
Ordered implementation plan
Order below is the recommended landing sequence, not a requirement to serialize independent investigation or test preparation. P1/P2/P3 are relative priorities within this audit, not estimates of effort. The dependency column distinguishes required foundations from coordination.
Phase 1 - Reliability boundaries
Phase exit: reproduced process/transport failures are covered by deterministic regressions; normal request behavior, streaming, and locator priority remain intact.
Phase 2 - Trustworthy measurements, responsiveness, and coverage
Phase exit: the roughly 400 ms client / 2 ms reported-duration reproduction is no longer invisible to the performance gate; black-box server testing contributes usable coverage; repeatable workloads establish a baseline for architectural changes.
Phase 3 - Coherent ownership, bounded work, and focused refactoring
Phase exit: ownership/resource limits are explicit and tested. Optimizations are supported by measurements. The locator framework and public behavior are preserved rather than replaced wholesale.
Parallel work and review boundaries
Transport shutdown/framing and subprocess-runner work can use separate PRs; coordinate the latter with #526. Metric/coverage test preparation can run alongside those fixes. Snapshot, scheduler, and writer changes should not be merged as one large refactor: establish ownership first, then independently validate scheduling and output behavior. Poetry profiling can run in parallel once the workload harness is ready.
Purely mechanical library-module reuse or the find read-guard fix may be isolated earlier; do not smuggle behavioral changes into those cleanups. Do not delay demonstrated reliability fixes for the optional Poetry optimization or final module cleanup.
Shared implementation and verification contract
Each ticket includes its own scope and acceptance tests. Implementations must preserve locator ordering, complete environment/manager information, platform path/symlink behavior, refresh coalescing, generation filtering, and scoped state synchronization. Use explicit errors, not success-shaped fallback data. Measure intentional metric/protocol changes and update directly related documentation.
Run targeted tests first, then relevant workspace/platform/feature jobs. Before committing Rust changes, run
cargo fmt --allandcargo clippy --all -- -D warnings; CI's all-targets/all-features lint and the existing exact-base quality gates must also remain green. Concurrency tests should use deterministic barriers/channels rather than timing guesses. Test fixtures must bound and clean up their own subprocesses and files.Milestones
Close this tracking issue only when each linked item is complete or explicitly deferred with a recorded reason. Check issue/PR state when starting work; links above describe filing-time status, not a substitute for current GitHub state.