feat(codex): run Codex on the official SDK with direct Responses, MCP tools and resumable threads - #1153
Open
tiantt wants to merge 17 commits into
Open
feat(codex): run Codex on the official SDK with direct Responses, MCP tools and resumable threads#1153tiantt wants to merge 17 commits into
tiantt wants to merge 17 commits into
Conversation
The SDK is stable now and pins its matching Codex CLI binary exactly,
so only `openai-codex` is pinned (in the extra and in the AgentKit
example requirements).
- Pass `personality` / `effort` as plain strings instead of importing
the SDK's private `openai_codex.generated` enums.
- Pin three settings in the generated config.toml that CLI 0.159
changed underneath the runtime:
- `unbounded_connection_retries = false`: an unreachable shim was
retried forever, hanging the invocation (no turn timeout exists).
- `goals = false`: goal tools were advertised to the backend model
although every thread here is ephemeral.
- `model_reasoning_summary = "none"`: Ark's Responses API rejected
every request over `reasoning.summary` ("unknown field").
- Explicitly ignore the four notification types the new SDK adds; none
of them can leave a turn waiting on the runtime.
- Add baseline tests for cancellation (while waiting on the model,
during a tool, with a failing interrupt), the pinned config settings,
the SDK/binary version pairing, and a real-binary smoke test that an
unreachable shim fails the turn in bounded time.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two ways the history the model sees drifted from what actually happened, both older than the SDK upgrade (reproduced on 0.1.0b3 as well). On a multi-step investigation they made the model repeat itself until the call budget ran out. - The shim appended its ADK tool pairs to the tail of every request, so the ADK results always read as the model's most recent step. It kept announcing it had just fetched its data and re-ran the same analysis. Each pair is now anchored to the Codex-visible function call of the reply it preceded and spliced back in before that reply; pairs whose anchor is gone (e.g. after compaction) still fall back to the tail. - Commands were recorded with Codex's login-shell wrapper (`/bin/zsh -lc '...'`). The next invocation's model saw that in its replayed history, copied it into its own commands, and Codex wrapped them again. The recorded command is now the one the model sent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pins the observable contract every agent runtime must keep, so the Codex runtime redesign (thread resume, MCP tool bridge, steer) can change its internals against a fixed bar. Scenarios run offline through the real Agent and Runner for adk, codex and piagent via per-runtime adapters; a new runtime only adds an adapter and its capability set. - Required: single turn, streaming, multi-turn, session isolation, cancellation (waiting on the model and mid-tool), tool call, tool failure, concurrent sessions, usage, lifecycle cleanup. - Capability-gated: approvals, MCP tools, skills; and resume across a restart, steer, turn timeout and compaction, written against adapter hooks no runtime implements yet, so skipped until a runtime declares them. - Two real gaps are strict xfails with their cause: runtime="adk" lets a raising tool crash the invocation, and runtime="piagent" never finishes a turn cancelled while a bridged tool runs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A fixed two-turn scenario (one ADK tool, three shell steps, then a follow-up on the same workspace) that a Codex runtime PR can run once against a real model for ~75K tokens, instead of full examples that cost millions. It fails on the regressions seen so far: backend-rejected request fields, looping on out-of-order tool history, the model copying Codex's shell wrapper, tools re-run, lost multi-turn context, and a token total over CODEX_PROBE_MAX_TOKENS. Skipped unless CODEX_RUN_PROBE=1 and MODEL_AGENT_* are set; never in CI. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex speaks the Responses API, and Ark / BytePlus ModelArk / OpenAI serve it compatibly, so for those hosts Codex can call the model directly instead of going through the in-process Responses-to-chat shim. This adds the routing layer the runtime integration will use: resolve_transport picks "direct" or "shim" (explicit CodexRuntimeConfig.model_transport, env VEADK_CODEX_MODEL_TRANSPORT, or host-based auto), and direct_route / shim_route build the provider as a thread_start(config=...) override. The key only travels in route.env (excluded from repr), never in the provider config Codex is handed. lean_codex_config pins the settings the shim path needed (no reasoning summary, bounded connection retries) and trims tools a VeADK turn cannot use (multi_agent, web_search, view_image, goals, request_user_input), roughly halving input tokens per request. Every key was checked against CLI 0.159.2 at thread level by capturing the request the real binary sends; an opt-in codex_smoke test keeps that proof. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ADK tools currently run inside the Responses shim, invisible to Codex, so the shim has to own the tool loop and its history. Serving them to Codex as an MCP server injected per thread lets Codex own the loop instead. This adds that server as a standalone module; wiring it into the runtime is a separate change. One in-process streamable-HTTP server per event loop (get_bridge) serves every turn. Each turn registers its specs and executors and gets a random bearer token; tools/list and tools/call only ever see that token's tools, and unknown tokens get 401. tools/call runs the executor under the turn's OTel context with the model's call_id from `_meta.callId`, returns the executor's JSON as structuredContent, records interrupt statuses (pending / authentication_required / confirmation_required / transferred) for the runtime to stop the turn, and turns executor exceptions into isError results. A dropped HTTP request or `notifications/cancelled` cancels the executor. The transport answers with plain JSON, not SSE: sse-starlette latches a process-global exit flag when it sees a uvicorn server shut down with a stream open, after which every later SSE response ends immediately, so one bridge stop mid-call would break every later bridge. The server also no longer captures the host's SIGINT/SIGTERM (uvicorn's capture_signals). codex_server_config() emits the Codex keys verified against 0.159.2: default_tools_approval_mode="approve" (deny_all blocks MCP calls otherwise), supports_parallel_tool_calls, tool_timeout_sec (above the ADK timeout), startup_timeout_sec and required=true. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The codex runtime is moving to a direct mode: Codex calls the model
provider itself (thread config model_providers.<id>) and reaches ADK tools
through VeADK's local streamable-HTTP MCP bridge (mcp_servers.veadk).
ShimDrivingCodex can only exercise the Responses shim, so none of that
path was testable offline.
DirectDrivingCodex mimics what codex 0.159.2 was observed to do in the
MCP spikes: it reads the provider and MCP servers from
thread_start(config=...) and credentials from CodexConfig.env, calls the
model through the patched litellm.aresponses with no shim, connects to each
server with the real mcp streamable-HTTP client (bearer from the env var),
advertises the tools as a {"type": "namespace", "name": "mcp__<server>"}
tool, executes namespaced function calls via tools/call with _meta.callId
(concurrently when supports_parallel_tool_calls), feeds results back with
codex's "Wall time ... Output:" framing, and emits mcpToolCall item
notifications, per-model-call token usage and turn/completed (interrupted
when interrupt() cancels in-flight model/MCP calls). Notification payloads
carry every field the real SDK models require, so real models are built
when openai-codex is installed.
ScriptedBackend gains, additively, the ability to answer with a
namespaced call when a plan names a tool only advertised inside a
namespace tool, and records namespaced tool names as plain names so the
two arms stay comparable.
The self-tests run the fake against a real FastMCP server on loopback and
reset sse_starlette's process-wide AppStatus.should_exit, which otherwise
latches after the first uvicorn shutdown and hangs every later bridge.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…r MCP Until now every Codex turn went through the in-process Responses->chat shim, which also ran the agent's ADK tools in a loop Codex could not see. For backends that speak the Responses API (Ark, BytePlus, OpenAI) neither is needed: a spike showed Codex calls Ark directly with none of the shim's compatibility patches, and reaches ADK tools through a local MCP server. That makes Codex the single owner of the tool loop, which the planned thread resume depends on (shim-run tools never reach the Codex rollout). - `model_transport` (auto/direct/shim, env VEADK_CODEX_MODEL_TRANSPORT) picks the path; `auto` is direct for Ark, BytePlus and OpenAI hosts and keeps the shim for everything else. - Direct: the provider and a trimmed tool set go in the thread config, the key only in the subprocess env; ADK tools are registered on the MCP bridge under a per-turn bearer token. - Codex's mirror of a bridged MCP call is dropped, so each ADK tool call appears once, with callbacks and state deltas applied. - `max_llm_calls` is charged per usage update (Codex announces nothing before a request), and the turn is interrupted once it is exceeded. - A bridged tool that needs a credential or confirmation, or transfers control, interrupts the turn instead of letting the model continue. - The conformance suite gains a codex-direct adapter; the example and runtime docs describe both transports. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
runtime="codex" is moving to one persistent Codex thread per VeADK session (thread_start(ephemeral=False), then thread_resume). In production (AgentKit) turns land on different stateless instances with ephemeral disks, and a thread's whole context lives in one rollout JSONL file under CODEX_HOME; copying that file into a fresh CODEX_HOME is enough to resume. So VeADK has to carry rollouts between instances itself. - rollout_io: find/export/import a thread's rollout. Import is atomic (tmp + fsync + rename), 0600, and refuses relpaths outside CODEX_HOME/sessions or not named for the thread, so stored data cannot overwrite config.toml/auth.json or escape via symlinked dirs. - thread_store: CodexThreadStore keyed by (app, user, session, agent) with compare-and-set versions so racing instances cannot silently drop a turn, plus the instruction hash the thread started with. Backends: in-memory, local dir (flock), and a SQL table (gzip blob, LONGBLOB on MySQL) on the short-term memory's own DatabaseSessionService.db_engine, so sqlite, mysql and postgresql (including its search_path schema) are covered with no extra configuration. select_thread_store picks it from ShortTermMemory. Nothing is wired into runtime.py yet. Rollout contents are never logged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Keeping one Codex thread per session makes the app-server's turn semantics load-bearing, and against Codex 0.159.2 they are hostile to naive use: turn() on a thread with an active turn joins it (new input lost when it races an interrupt), interrupt() right after turn() is rejected with -32600 "no active turn" and the turn then runs to completion, the interrupted turn stays active until its turn/completed arrives 10-70 ms later, and compact() only starts a compaction turn whose events reach no handle and which rejects turn() with -32603 ActiveTurnNotSteerable. turn_control.py turns that into small primitives the runtime can wire in without further protocol knowledge: a per-session asyncio lock that drops idle entries, an active-turn registry whose steer() never starts a turn, TurnCompletion fed by the stream pump, interrupt_turn() that retries the early rejection and waits for turn/completed, start_fresh_turn() that refuses a joined turn and waits out a compaction, compact_and_wait() that polls thread.read(include_turns=True), and run_with_turn_timeout() raising CodexTurnTimeout. Errors are classified through the SDK's public InvalidRequestError / InternalRpcError. Not wired into runtime.py yet. Unit tests cover every branch on fakes; codex_smoke tests (CODEX_RUN_SMOKE=1) check the five behaviours against the real binary with a stub backend. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every invocation used to start an ephemeral Codex thread and replay the session as a JSON transcript in the prompt. Codex then never saw its own earlier commands, outputs and plans, and the replayed commands taught the model to copy Codex's shell wrapper. With ADK tools now reaching Codex over MCP (so they are in its history), the thread itself can be kept. - `thread_mode` (resume/ephemeral, env VEADK_CODEX_THREAD_MODE), resume by default on the direct transport; the shim stays ephemeral. - The thread's rollout is saved to the thread store after every turn (failures and cancellation included) and imported into the next invocation's fresh CODEX_HOME, so any instance can resume it. - A resumed turn sends the current message only, plus what the user or other agents said since this agent last replied, and the results of tool calls that ran once the user confirmed them. - Settings are passed again on every resume: Codex falls back to its defaults for model and sandbox otherwise. - A changed instruction starts a new thread from the transcript (Codex ignores new developer instructions on resume); a failed resume or a corrupt record does the same instead of failing the turn. - A per-session lock keeps two invocations off one thread; a lost optimistic save is logged, not raised. - Usage is reported per turn, not as the resumed thread's running total. - The offline fake persists and resumes threads; the conformance suite's codex-direct adapter now declares resume_across_restart. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With persistent threads a turn has to end cleanly and on time, can be redirected while it runs, and needs its history kept in check. - `turn_timeout_seconds`: past the deadline the turn is interrupted, given a grace period, and the invocation fails with a TimeoutError. The watchdog owns the pump, so a pump it has to cancel surfaces as the timeout rather than as a cancellation. - Cancellation waits until the Codex turn has really stopped (retrying an interrupt that lands before the model request starts), so the next invocation never joins a dying turn. - `Runner.steer(session_id, text)` adds an instruction to the turn in flight through a generic `BaseRuntime.steer`; the Codex runtime registers its running turn per session (direct transport only: a steered message through the shim would carry no turn marker). - `auto_compact_token_limit` hands compaction to Codex itself (`model_auto_compact_token_limit`); no summarizer of our own. - The turn's span records `veadk.codex.*` metadata: transport, thread and turn ids, resume, model, sandbox, approval mode, tool count, status and duration -- no prompts, arguments or keys. - The fake app-server steers and compacts like Codex 0.159; the conformance suite's codex adapters now declare turn_timeout, and codex-direct also steer and compaction. Docs gain an architecture section with the local/managed backend roadmap. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lled - examples/codex_session_lifecycle: one session through MCP and function tools, a skill, sandboxed writes, streaming, a steer, cancelling a running turn, and resuming the session's Codex thread. Verified against Ark. - When the caller's task is cancelled, ADK's Runner closes the runtime generator (GeneratorExit) instead of raising CancelledError into it. That path logged the turn as failed and never interrupted Codex; it is now a cancellation with the same safe interrupt. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From an anr code review of the Codex runtime work: - Credentials leaked into the sandboxed shell: `env` run by Codex printed the direct transport's model key (and would show the shim turn token and the MCP bridge token). The generated config now pins a shell environment policy excluding VEADK_CODEX_*, *API_KEY*, *SECRET* and *TOKEN*; a real-binary smoke test runs `env` and checks the key is gone. - `max_tool_iterations` only bounded the shim: on the direct transport the runtime now counts the MCP bridge's calls and interrupts the turn with CodexToolIterationLimitError past the limit. - `extra_body` was silently dropped on the direct transport. Under `auto` an agent with a body to forward stays on the shim; an explicit `direct` logs a warning. VeADK's default Ark body (caching) is dropped on Responses anyway and does not force the shim. - `turn_timeout_seconds` defaults to 1800: an unbounded turn would also block its session's queued invocations. - Saving the rollout no longer swallows cancellation, and cleanup always runs; rollout import/export move off the event loop. - The backfill handed to a resumed thread is capped (message count, message size, resumed tool result size). - Failure logs carry thread/turn/session ids and a bounded error message; a codex_turn_started line correlates turn ids with invocations. - Docs and the AgentKit example (veadk-python>=1.1.15, moved together with the openai-codex pin) describe the above. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… save races A review of the runtime="codex" work found guarded behaviour with no test pinning it, so a regression in any of these would have gone unnoticed: - Thread stores: LocalDir and Database now must raise ThreadStoreCorrupt for an undecodable blob or a sha mismatch, and _load_thread must delete the corrupt record so the next expected_version=None save succeeds (otherwise the session is pinned to throwaway threads for good). - MCP bridge: oversized bodies get 413 before any executor runs, paths other than /mcp get 404, and non-loopback Host/Origin are refused (DNS-rebinding protection), each paired with a loopback positive. - Config: VEADK_CODEX_THREAD_MODE accepted/rejected values, invalid thread_mode, and zero/negative turn_timeout_seconds and auto_compact_token_limit. - Runner.steer: routing to a nested codex sub-agent with the runner's app_name and default user_id, False when nothing delivers, first delivery wins, and a pure adk tree never consults a runtime. - Cross-instance saves: a ThreadStoreConflict on save is logged as codex_thread_save_conflict and does not fail the invocation, and a turn after another writer's save resumes that writer's record and version. - Conformance compaction: the compaction request advertises no agent tools while the agent's own turns do, and the next turn no longer carries turn 1's answer. The module docstring now lists which adapters declare steer/turn_timeout/compaction/resume_across_restart. Every new assertion was mutation-checked against a temporarily broken veadk/ (or, for the compaction tool list, the fake Codex) and restored byte-identically. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From the semantic axes of an anr review: - A turn whose rollout save was lost (cross-instance conflict, store error, cancellation) vanished from the Codex thread: the backfill anchored on this agent's last reply, and that reply belonged to the lost turn. Records now store the last invocation their rollout covers (`covered_invocation_id`), and a resumed turn is handed everything after it, including this agent's own lost replies. - A codex agent transferred to within the same invocation lost what the parent agent said in it: the backfill dropped the whole current invocation. It now skips only the current user message (rendered as the prompt) and this agent's own events. - The settings pinned for Codex lived twice (hand-written config.toml and the direct thread override) and had already drifted. `pinned_codex_settings()` is now the single source for both. - Docs: the per-session lock applies to direct + resume only; `max_tool_iterations` counts individual calls on the direct transport. turn_control's module doc now describes the real wiring and names the helpers that are not wired in. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ntime From a second round of review on the Codex runtime: - Sensitive model headers (Authorization, Cookie, or names with key / token / secret / password / signature) no longer land in Codex's config file on the direct transport: they go through `env_http_headers`, with values only in VEADK_CODEX_* env vars the sandboxed shell cannot see. A real-binary test checks the header reaches the backend and appears in neither the file nor the shell. - Resume no longer abandons a thread's native context on any error: transient overload is retried (bounded, with backoff); unknown threads and invalid parameters still fall back to a new thread at once. - A resumed turn on an empty workspace (new instance, restart) is told its earlier files may be gone. - Rollouts over `VEADK_CODEX_MAX_ROLLOUT_BYTES` (32 MiB) are refused; the binding is dropped so the next turn starts fresh. The in-process store keeps rollouts compressed and evicts least-recently-used ones. - OpenTelemetry metrics with low-cardinality attributes: resume and save outcomes, turns by status/transport with duration, startup time, and tokens. - Docs: limitations (workspace not persisted, no cross-instance lease, best-effort max_llm_calls on direct, rollout cap), sensitive headers, resume retries, metrics and the new env vars. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Reworks
runtime="codex"so Codex owns the agent loop, the tool loop, the sandbox and short-term context, while VeADK owns session binding, tool access, event translation, policy, tracing and deployment. The public API (Agent(runtime="codex")) is unchanged.What changes
openai-codex0.1.0b3 → 0.159.2 (stable; it pins its matching CLI). No imports fromopenai_codex.generated. Pins three CLI 0.159 defaults that broke the runtime (Ark rejectedreasoning.summary; unbounded connection retries could hang a turn; goal tools on ephemeral threads).model_transport, defaultauto): for Ark / BytePlus / OpenAI, Codex calls the Responses API itself, and the agent's ADK tools reach Codex through a per-turn local MCP bridge, so Codex drives the tool loop. Chat-only backends keep the shim.thread_mode, defaultresume, direct transport): one Codex thread per session; its rollout is stored with the session (same database for sqlite/mysql/postgresql short-term memory) and resumed on any instance. No transcript replay; the resumed turn gets only the current message plus what others said since, and results of tools approved in between.turn_timeout_seconds(default 1800), safe interrupt on cancel,Runner.steer(session_id, text)via a genericBaseRuntime.steer, Codex-native auto compaction (auto_compact_token_limit).max_tool_iterationson the direct transport,extra_bodykeeps the shim underauto.veadk.codex.*span attributes (transport, thread/turn ids, resume, model, sandbox, approval mode, tool count, status, duration).Behavior changes for existing Ark users (and how to roll back)
max_llm_callsis charged after each call (one in-flight request may be aborted)VEADK_CODEX_MODEL_TRANSPORT=shimormodel_transport="shim"veadk_codex_threadstable is created in the session database (needs DDL permission)VEADK_CODEX_THREAD_MODE=ephemeralorthread_mode="ephemeral"turn_timeout_seconds=NoneReview follow-ups (latest commits)
env_http_headersand never land in Codex's config file.covered_invocation_id); a transferred-in agent sees the parent's words.VEADK_CODEX_MAX_ROLLOUT_BYTES, 32 MiB) and LRU eviction for the in-process store;max_tool_iterationsenforced on direct;extra_bodykeeps the shim underauto.veadk.codex.thread.resume,veadk.codex.thread.save,veadk.codex.turn(+.duration),veadk.codex.turn.startup,veadk.codex.turn.tokens(low-cardinality attributes only).max_llm_callson direct is best effort.Tests
tests/runtime+tests/runnerwithCODEX_RUN_SMOKE=1: 529 passed. New behavior is mutation-checked (the test fails when the guarded code is broken).codex_with_skill_and_mcp,codex_data_analysis, newcodex_session_lifecycleexample (tools, skill, sandbox, streaming, steer, cancel, resume), andcodex_ops_assistantondeepseek-v4-pro: both turns completed, ticket filed, turn 2 resumed the thread.Follow-ups (not in this PR)
veadk_codex_threads; rollout size cap; eviction for the in-memory thread store (localbackend).🤖 Generated with Claude Code