Skip to content

v0.9.4: db contention fixes, databricks genie, snowflake cortex, additional search connectors - #8356

Merged
waleedlatif1 merged 32 commits into
mainfrom
staging
Sep 28, 2026
Merged

waleedlatif1 merged 32 commits into
mainfrom
staging

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

waleedlatif1 and others added 21 commits September 26, 2026 13:11
…une finished outbox events (#8330)

- Processing recovery discovered candidates by walking the global `(uploaded_at, id)` recovery index, joining each document to its knowledge base and connector, and rejecting it only afterwards. Retained inputs of paused or dormant sources sit at the front of that order, so every call read through all of them first. Follow-up batches (`NOT IN` attempted connectors) walked the same range again
- Discovery now starts from the eligible connectors and reads each one's oldest candidates through a new partial index `doc_connector_processing_recovery_idx (connector_id, uploaded_at, id)` with a `CROSS JOIN LATERAL`, then takes the oldest overall. Paused sources cost nothing. Batch size, the follow-up batches for refused connectors, blocked knowledge bases, the lock-time recheck, and liveness are unchanged, so recovery picks the same candidates as before
- The old `doc_processing_recovery_idx` stays for now and is dropped in a later contract migration
- Outbox retention: `outbox_event` never deleted completed rows. The outbox processor now prunes `completed` rows older than 7 days in bounded, `SKIP LOCKED` batches, only for `knowledge.document.processing.recover` and `knowledge.document.storage.cleanup`. Both have random ids, and nothing reads their completed rows. Every other event type is kept, including idempotency-keyed ones and the checkpoint expiry events, whose completed rows are read back by id, as are pending, processing, and dead-letter rows
…nd event (#8328)

- Each chat turn loaded the chat's full transcript through `resolveOrCreateChat` just to read the MCP server ids earlier user messages were tagged with. `resolveOrCreateChat` now never loads the transcript, and a new `loadChatMcpServerIds` reads only the MCP context ids with one jsonb query, in the same first-tagged order
- Split the chat loaders: `getAccessibleCopilotChatDetail` (row plus authorization, no transcript) backs both `resolveOrCreateChat` and `getAccessibleCopilotChatWithMessages`. Dropped the `includeTranscript` option and the `conversationHistory` field, which nothing read
- Client: a `completed` event for the viewer's own live stream no longer refetches the chat detail, because the client's own stream finalization already refetches it. `renamed` marks the detail stale without refetching; the lists that show titles still refetch
…indexes (#8333)

- Since #8314, workspace knowledge search ranks on `embedding`/`document`. The only live reader of `embedding_keyword_search` and `embedding_keyword_tin` is dormant indexed search, and only for bases that are search indexes. Yet every chunk write still upserted a keyword row, ran a Tin `DELETE` on insert, and every document insert fanned an ACL/connector sync out to the projections
- Script migration `0025_scope_keyword_projections`, all `CREATE OR REPLACE` in one transaction with the 0024 lock-timeout retry:
  - `sync_embedding_keyword_search` takes the shared membership lock and writes keyword rows only for `is_search_index` bases. On an update that moves a chunk out of a search index, it deletes the chunk's row
  - A new flip trigger backfills a base's keyword rows when it becomes a search index and deletes them when it stops being one, under the exclusive membership lock. It is separate from the Tin flip trigger, which only exists where Tin is installed. Its upsert skips unchanged rows
  - The Tin chunk trigger skips its `DELETE` on insert, since a new chunk can't have a Tin row
  - `projection_source_acl_sync` becomes `AFTER UPDATE OF connector_id, acl ... WHEN` the value actually changed. A new document has no chunks yet (FK), so the insert arm never matched anything
- The projector's keyword page writes only for search-index bases, keyed on the base flag
- Fresh installs, `db:push`, and late Tin adoption end with the same trigger definitions: 0019 installs the final guarded, key-share-locked Tin triggers, and 0025 re-runs after 0016. A `db:push` re-run of the legacy 0016 backfill may refill keyword rows for non-search bases; they're unread and left in place like the existing ones
- Both adoption backfills read their chunks `FOR KEY SHARE`, so a chunk delete racing an adoption is waited out instead of failing it
- Rollback floor: v0.9.1 and earlier ranked workspace keyword search through `embedding_keyword_search`. Once 0025 has run, don't roll back below v0.9.2, and self-hosters should upgrade through v0.9.2 or later
- Existing keyword rows for non-search bases are left in place (unread) and removed separately
- Every transaction-scoped advisory lock in `apps/sim` ran the same statement, `SELECT pg_advisory_xact_lock(hashtextextended($1, 0))`, so query insights showed all lock waits as one fingerprint with no way to tell which caller was waiting
- New `acquireAdvisoryXactLock(tx, tag, key)` and `tryAcquireAdvisoryXactLock(tx, tag, key)` in `lib/db/advisory-locks.ts` issue the same call with a trailing SQLCommenter tag (`/*lock='<family>'*/`), which PlanetScale query tags can filter on. The tag must be a static snake_case identifier, so it can't break out of the comment
- About 45 call sites now go through the helper, each tagged with its lock family (`user_table_rows_pos`, `organization_membership`, `usage_log_event`, …). Keys, lock types, timeouts, and transaction boundaries are unchanged
- Left alone: statements whose SQL text is already unique (`packages/db` locks, the file-search dispatch `generate_series` claim) and one untyped repair script
…s processes (#8325)

- Enterprise reporting-window usage sums (a year-long ledger scan) are now shared across processes through Redis with a 30s TTL. Trigger.dev runs every task in a fresh process, so the existing in-process LRU was always cold there, and the scan ran on nearly every document-processing and execution admission check
- The in-process LRU stays in front and still coalesces concurrent misses. The Redis GET has a 250ms deadline, and any Redis error, timeout, or unreadable value falls through to the exact sum. The write is a fire-and-forget `SET … NX` with a jittered TTL, so a slower, older sum never overwrites a fresher one or extends its life. There is no lock or lease
- The usage threshold email is now level-triggered, instead of being edge-triggered off an exact before/after org sum on every workflow completion. A claim keyed on (billing period, limit) (`claimCreditsThreshold`) sends each threshold at most once. A new period or a changed limit re-arms it with no reset write, and concurrent completions can't both send
- Recipients are resolved before claiming, so a period's email isn't used up when nobody can receive it. Zero-cost completions don't claim. The personal usage baseline and the claimed period come from the same billing context
- Invoicing, cycle close, overage, and threshold billing still read the ledger exactly
- Rollout note: accounts already at or above 80% this billing period get one threshold email after deploy; there is no backfill
…e unique index (#8327)

- The File block's URL fetch no longer saves every fetched URL into workspace Files. Nothing ever read the saved copy: the parser keeps its own execution file, and the reuse the save once served was removed earlier. Repeated fetches were piling up `name (N).html` copies in the Files root
- Renamed `fetchExternalUrlToWorkspace` to `fetchExternalUrl` and removed its save, permission, and upload branch
- Name existence checks (`fileNameExistsInWorkspaceFolder`, `getWorkspaceFileByName`) now spell the folder predicate as `coalesce(folder_id, '') = $folder`, matching the unique `(workspace_id, coalesce(folder_id, ''), original_name)` index, so each check is a point lookup. The old `folder_id IS NULL` form made root lookups scan the folder-id index
- `allocateUniqueWorkspaceFileName` probes the base name plus `(1)`…`(20)`, then falls back to a short-id suffix instead of probing up to 1,000 candidates and then failing. The unique index and the existing conflict retry remain the authority
…where counts are shown (#8329)

- v1 knowledge base list, detail, and update returned `docCount`/`tokenCount` over every document in the base, including connector documents the caller cannot read. They now count through the caller's access (`resolveV1KnowledgeReadAccess`), the same way the v1 documents routes already do. v1 detail also now returns its real `connectorTypes`
- The internal KB list joined `document` through the full access predicate on every call, but only the Knowledge page shows the totals. The list contract gains `includeCounts` (default false), and without it the query reads only `knowledge_base` columns. The service option is `countsFor: access`, so a count without an access filter can't be written
- React Query: counted lists get their own key beside `list()`, used only by the Knowledge page and its server prefetch. Document, upload, and connector mutations refresh only the counted lists; KB create, rename, delete, restore, and move refresh both
- Removed `getKnowledgeBaseById`, whose unfiltered count join ran on every context resolution. Callers use `getActiveKnowledgeBaseReference`, and single-base totals come only from `attachKnowledgeBaseConnectors(kb, access)`
- The v2 list keeps returning totals, since its public contract requires them
* fix(copilot): bound agent CLI grep matching

* fix(copilot): bound grep context expansion
* fix(api): constrain workflow response headers

* fix(api): reserve additional browser policy headers
…#8332)

- Drops 8 indexes on hot write tables in one `DROP INDEX CONCURRENTLY` migration (0386), following the 0239 pattern: `COMMIT` breakpoint, `lock_timeout 0`, idempotent replay
- `{emb,doc}_date1_idx` / `date2_idx`: tag filters compare `col::date`, which a plain timestamp btree can never match, so these are written on every chunk insert and document update and never serve a query
- Four left-prefix duplicates of an existing non-partial superset, which keeps serving equality lookups and FK cascades on the leading column:
  - `emb_kb_id_idx`: covered by `emb_kb_enabled_idx` / `emb_kb_model_idx`
  - `emb_doc_id_idx`: covered by `emb_doc_chunk_idx` / `emb_doc_enabled_idx`
  - `usage_log_workspace_id_idx`: covered by `usage_log_workspace_created_at_idx`
  - `copilot_runs_execution_id_idx`: covered by `copilot_runs_execution_started_at_idx`
- Kept on purpose:
  - Every `number` and `boolean` tag-slot index. The vector leg's short-result rescue probe checks tag filters with an `EXISTS` over `embedding` that relies on them, and without them a selective number/boolean filter turns that probe into a table-wide scan
  - The small `copilot_runs` chat/workspace indexes
…8343)

* feat(databricks): add Genie agent operations to the Databricks block

* fix(databricks): fail Genie asks on nested errors and document per-tool outputs
* fix(execution): bound reference and context processing

* fix(execution): preserve compatibility and tighten input budgets
…8346)

* fix(desktop): wait for native update staging and verify installation

* fix(desktop): distinguish refresh failures from native update errors
…#8347)

* feat(snowflake): add Cortex Analyst operations to the Snowflake block

* fix(snowflake): keep Cortex Analyst SQL context for quoted and multi-source views and replay suggestions

* fix(snowflake): run Cortex Analyst SQL with the resolved token, the chosen source's schema, and document access requirements
* feat(search): expand live providers and discussion reads

* fix(search): preserve transcript details and account search

* fix(search): require complete review and date evidence

* fix(search): align mixed-query acceptance limits

* fix(search): bind acceptance evidence to matched comments
…tamps stay off the index (#8334)

- Connector sync stamps `document.source_seen_at` on every listed document. `source_seen_at` is a key column of `doc_connector_reconciliation_idx`, so every stamp was a non-HOT update that wrote every index on `document`. This is release A of two: move every reader off that key so release B can drop it and the stamps become HOT
- New `doc_connector_reconciliation_v2_idx (connector_id, id)` with the same partial predicate, built concurrently. The old index stays until release B
- The absence walks (ACL revoke, soft delete, hard delete) and the member resurrection walk page by id instead of `(seen, id)`, with absence as a plain filter. Each window is built by a recursive keyset walk that fetches one row per step (`id > previous ORDER BY id LIMIT 1`) up to 5,000 ids, so every statement reads at most one window whatever plan the database picks. A single `ORDER BY id LIMIT` could be planned as a bitmap read of the whole connector plus a sort. Matches are filtered within the window, so cost is bounded by ids scanned, not matches found. A walk whose absence count is zero is skipped
- Both seen stamps skip rows already stamped at or after the run start (`staleSeen`), so current rows aren't rewritten and a later stamp is never overwritten by an earlier one
* fix(chat): use hostname brand icons for inline sources

* fix(chat): recognize www aliases in source branding
* feat(chat): add shared find menu

* fix(chat): preserve find focus and scroll ownership

* fix(chat): resume pending send scroll after find closes
@vercel

vercel Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 28, 2026 4:00am UTC

Request Review

@github-actions github-actions Bot added the requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep label Sep 28, 2026
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

⚠️ Cross-repo companion check

One or more companion PRs aren't merged into main yet (aggregated across the feature PRs in this release). Merging this without them will leave copilot and sim out of sync — merge them in lockstep.

  • ⚠️ simstudioai/mothership#539 — merged into staging (this PR targets main) — feat(search): expand live provider contracts and query guidance
  • ⚠️ simstudioai/mothership#544 — merged into staging (this PR targets main) — fix(search): reject empty Notion queries before dispatch

Comment thread apps/sim/scripts/test-search-discussions-live.ts Dismissed
Comment thread apps/sim/scripts/test-search-discussions-live.ts Dismissed

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 278 files

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

Comment thread apps/sim/tools/snowflake/cortex_analyst_ask.ts
Comment thread apps/sim/lib/api/contracts/v2/knowledge.ts
Comment thread apps/sim/tools/databricks/utils.ts Outdated
Comment thread apps/sim/lib/logs/execution/logger.ts
Comment thread apps/sim/lib/db/advisory-locks.ts Outdated
Comment thread packages/db/knowledge-projection.ts
Comment thread apps/sim/scripts/test-search-discussions-live.ts
Comment thread apps/sim/lib/sim-search/live/meeting-content.ts Outdated
Comment thread apps/desktop/e2e/fixtures/updater.ts
@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[High risk] Database contention fixes and new search connectors with API integration changes.

The PR should not merge until chat find counts and highlights the same rendered matches; the updater checkpoint and legacy-index concerns are follow-up improvements.

Findings

  1. P1 Citation spacing breaks chat find ▶
  2. P2 Native errors leave stale checkpoints ▶
  3. P2 Legacy index preserves write cost ▶

Summary

This release combines database and knowledge-index contention work with new live-search and warehouse integrations, chat improvements, and desktop updater staging verification.

  • Knowledge recovery, reconciliation, projections, counts, and outbox retention receive new bounded or scoped paths.
  • Chat gains a find menu and avoids repeated transcript loads; Databricks Genie, Snowflake Cortex Analyst, and additional search providers gain operations.
  • Desktop updates now wait for native staging and record installation outcomes.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Connector documents] --> B[Recovery and id-window reconciliation]
  B --> C[Knowledge projections]
  C --> D[Indexed and live search]
  E[Chat transcript] --> F[Chat find index]
  F --> G[Rendered-match highlighting]
  H[Desktop download] --> I[Native staging]
  I --> J[Install checkpoint]
  J --> K[Next-launch result]
Loading

Reviews (1) · Last reviewed commit: "feat(chat): add shared find menu (#8351)"

Comment thread apps/desktop/src/main/updater.ts
Comment thread packages/db/schema.ts
…al coverage (#8359)

* fix(search): let cancelled discussion reads stop instead of reporting partial coverage

* fix(search): judge live-read cancellation by the signal, not the error shape

* fix(search): stop a live read cancelled during current-scope verification
… anchor shared sums to their start (#8360)

* fix(billing): key the usage email claim on the exact period start and anchor shared sums to their start

* fix(billing): anchor shared sum expiry with a portable PX and a floor for slow sums

* docs(billing): note the reconnect delay in the shared sum staleness bound
…atch mention chips next to punctuation (#8361)

* fix(chat): keep find stable across footnotes and pending chats, and match mention chips next to punctuation

* fix(chat): keep the longest overlapping mention chip and centralize the integration matcher mock

* fix(chat): avoid Array.at in the mention range merge

* fix(chat): only treat a period as a mention boundary when no name continues after it
…osts (#8362)

* fix(databricks): send workspace tokens only to Databricks workspace hosts

* fix(databricks): accept DoD, custom-URL, and trailing-dot workspace hosts

* fix(security): strip a trailing FQDN dot from vendor-hosted URLs

* test(databricks): assert every tool refuses a foreign host with the allowlist error
…k and bound discussion case text (#8363)

* test(knowledge): accept either ordered index in the scale window check and bound discussion case text

* chore: document the GitHub text-limit helper and brace the scale plan assertion
* refactor(db): require a transaction for transaction-scoped advisory locks

* chore(db): document why the advisory lock helpers take a transaction
…cursor (#8365)

* fix(knowledge): resume the member resurrection walk from a persisted cursor

* chore(knowledge): brace the resurrection cursor branches and drop a vacuous assertion
…action (#8366)

* fix(search): validate a new provider server before the approval transaction

* fix(search): recheck for the provider server after a failed pre-approval lookup
…#8367)

* fix(knowledge): serve list totals when a request omits the count flag

* test(knowledge): cover the internal list request a pre-flag page sends
@waleedlatif1
waleedlatif1 merged commit 8725250 into main Sep 28, 2026
66 of 67 checks passed

This branch was previously deployed

1 inactive deployment
Preview — a657493f Deployed Sep 28, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants