Skip to content

fix: recover the local stack after a port clash - #4

Merged
sijie merged 1 commit into
mainfrom
fix/local-port-clash-recovery
Oct 3, 2026
Merged

sijie merged 1 commit into
mainfrom
fix/local-port-clash-recovery

Conversation

@sijie

@sijie sijie commented Oct 3, 2026

Copy link
Copy Markdown
Member

Summary

  • local/engine.sh removes the Agent Engine's containers that are not running before ork local start, so a start never reuses one that Docker left without a network.
  • Troubleshooting for the Local course gives a port-clash recovery that works, for both stacks.
  • Lab 0 steps 1 and 2 point to Troubleshooting when a port is taken.

Why

Found while taking the Local course on the CLI path with port 8080 held by another container.

On Docker Engine 29.2.1, a container whose port could not be bound is left without its network. Every later start of that same container has loopback only, even after the port is free. Only a new container recovers.

  • Lab 0 step 2. With 8080 taken, local/engine.sh stops with Bind for 0.0.0.0:8080 failed: port is already allocated. local/engine.sh --check then prints fix: local/engine.sh, and that rerun fails with registry-1 exited (1) (getaddrinfo EAI_AGAIN postgres in its log). Stopping the other program does not help. Changing the port worked only because Compose builds a new container.
  • Lab 0 step 1. With a streaming port taken (8000 in my run), stopping the program and running up again exits 0 and the step's check lists all six services, but the MCP container has no published port and cannot be reached.

Changes

  • local/engine.sh, local/lib.sh: before ork local start, remove the engine's containers that are created, exited or dead. Running containers and other projects' containers are left alone. The engine's data is in volumes.
  • local/tests/run.sh: two tests for that, a fake ork, and a fix for the // that a TMPDIR ending in / leaves in the test path on macOS.
  • labs/local/troubleshooting.md:
    • Streaming port clash: stop the program, local/down.sh, then up again.
    • Engine port clash: stop the program and run local/engine.sh again, or move the registry with ORCA_LOCAL_REGISTRY_PORT. Names the second port the engine publishes, 18082.
    • Doctor cannot reach a service that is up: a container lost its network, so local/down.sh and start again.
    • New row for .venv/bin/python: No such file or directory on the CLI path.
  • labs/local/00-set-up.md: steps 1 and 2 point to Troubleshooting for a port clash.

Test plan

  • local/tests/run.sh: 43 passed. The two new tests fail against the previous engine.sh and lib.sh (41 passed, 2 failed).
  • shellcheck for local/, cli/ and scripts/, scripts/check-labs.sh, scripts/tests/run.sh (25), cli/tests/run.sh (151), TypeScript tests (226) and typecheck.
  • A throwaway engine with its own data directory on port 38080 (Docker Engine 29.2.1, Compose 5.1.1, ork 0.6.0):
    • Previous script: bind error, then registry-1 exited (1), and still exited (1) after the port was free.
    • This script: the same bind error while the port is taken, four PASS lines once it is free.
    • Run again on a healthy engine: the running containers keep their ids; only the two one-shot containers are replaced.
    • After every engine container was stopped and the port taken: recovers once the port is free, and the workspace key is still accepted.
  • Lab 0 step 2 as written, on the default port: its check passes, and step 3's check prints true.
  • Streaming stack with port 8000 taken: local/down.sh, then up again, publishes the MCP port again.
  • Python tests were not run locally (no python/.venv in the checkout). No Python file changed, and CI runs them.

Not verified: a program that is not a container holding the port. The troubleshooting row says "if it is a container" for that reason.

Docker Engine 29.2 leaves a container whose port could not be bound
without its network. Every later start of that container has loopback
only, even once the port is free. Only a new container recovers.

For the Local course that meant:

- Lab 0 step 2: after `Bind for 0.0.0.0:8080 failed`, running
  local/engine.sh again (what --check says to do) failed with
  `registry-1 exited (1)`, and stopping the other program did not help.
- Lab 0 step 1: after stopping the other program, `up` exited 0 and the
  step's check listed six services, but the container had no published
  port.

local/engine.sh now removes the engine's containers that are not
running before `ork local start`. The engine's data is in volumes. A
rerun gives the same bind error while the port is taken, and a working
engine once it is free.

Troubleshooting gives a recovery that works for both stacks, covers a
doctor that cannot reach a container that lost its network, and has a
row for a missing python/.venv on the CLI path. Lab 0 steps 1 and 2
point to it.
@sijie
sijie merged commit 8f472c1 into main Oct 3, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant