Skip to content

Give find and resolve coherent configuration snapshots without lock-held discovery #536

Description

Tracking plan: #528
Priority: P2. Evidence: source-confirmed ownership inconsistency; add deterministic race tests before claiming a reproduced user failure.

Problem

Refresh uses a transient locator graph and a configuration-generation snapshot. Find/resolve still use the long-lived shared graph while configure mutates its locators one by one before publishing the new generation. Those operations can therefore observe partially applied configuration even though refresh has a coherent boundary.

Additionally, find passes a chained configuration read/clone expression directly into discovery; the temporary read guard can live through that call, retaining the configuration lock during filesystem work.

Sources: configure prepare/mutate/publish, resolve, find, refresh-state contract.

Scope

  • First make find's configuration snapshot an explicit local value and release its guard before discovery.
  • Define one coherent request-snapshot boundary for configure/refresh/find/resolve. Prefer constructing the replacement configuration/locator graph off-lock and atomically publishing it, rather than extending broad locks across I/O.
  • Separate deliberately shared performance caches from configured input and correctness-critical discovery state. Preserve current generation-gated notifications, locator ordering, scoped refresh-state sync, and coalescing semantics.
  • Replace unnecessary mutable-graph rollback complexity only where the new ownership model demonstrably removes the failure mode; do not bundle unrelated rewrites.

Acceptance criteria

  • Barrier/channel-driven tests prove find and resolve see either the old or new configuration, never a mix across locators.
  • Slow discovery/configuration fixtures do not hold a global configuration lock over filesystem or subprocess I/O.
  • Concurrent configures publish in a defined order; failures leave the previous usable snapshot intact.
  • Full/workspace/kind-filtered refresh sync, stale notification suppression, joined refresh replies, and manager fidelity retain existing behavior.
  • In-flight cache invalidation/clear and configuration changes have explicit tested semantics.
  • No regression in client-observed latency or long-lived cache behavior, and locator-state documentation includes the actual lifetime of every locator's mutable state.

Dependencies and prior work

Use #531 metrics and the #533 long-lived/concurrency workload before broad ownership changes. This establishes the ownership foundation for #539 and #540; the complete order is in #528. #385 fixed configure versus refresh isolation; #461 removed locator I/O from the configuration write lock. This issue extends consistency to find/resolve and must not regress either fix.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    debtCode quality issues

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions