CB-306: spawn-readiness gate — don't hand a session to bridge_send until the peer's bridge MCP has handshaked #4

Closed
opened 2026-07-18 15:00:59 +02:00 by ltms · 2 comments
Owner

Problem

bridge_spawn returns a session as soon as the worker pane/agent is created over herdr, but a
freshly spawned Claude Code worker is not actually reachable over the bus until its bridge MCP
client has connected and handshaked
with the daemon. Between those two moments the worker exists
but cannot receive an injected turn.

The worst-case manifestation is the pre-REPL folder-trust stall: a worker launched in a
not-yet-trusted directory sits at the CLI trust prompt and never loads its MCP config, so its bridge
client never connects. A bridge_send to that session blocks against a peer that will never answer,
and the caller only learns via a timeout (~caller MCP client cap, ~60s) — presenting as a silent
"communication break" rather than a spawn failure.

Root asymmetry: worker delivery is a status-gated push (Injector.enqueue, gated on
agent_status ∈ {idle, blocked}), but "the pane is up" is inferred from herdr, not from a positive
signal that the bridge protocol is live end-to-end.

Proposal — a readiness gate

Introduce an explicit peer-readiness state distinct from "pane exists":

  • Track a presence per session: SPAWNING → READY → … where READY is set only when the worker's
    bridge MCP client has completed its first handshake with the daemon (a positive protocol signal,
    not a herdr-output inference).
  • bridge_send / injector delivery must not target a session until it is READY. Before READY,
    either block on a bounded readiness wait or reject fast.
  • Add a spawn-readiness timeout (config knob, e.g. spawn_ready_timeout, default ~15–20s). If a
    peer doesn't reach READY within it, bridge_spawn fails explicitly ("peer did not become
    reachable — likely stuck pre-REPL / untrusted folder") instead of returning a half-live session
    that later times out on first send.
  • Surface presence in bridge_list / GET /workers roster so a stuck-spawning worker is visible.

Why now

This turns the pre-REPL stall (and any other "pane up but MCP not connected" case) from a delayed,
misleading bridge_send timeout into a fast, explicit spawn failure at the right layer.

Notes / scope

  • The handshake signal likely rides on the worker's first bridge MCP call reaching the daemon
    (server side sees the client connect); confirm the daemon can attribute that connection to a
    specific spawned session (name/nonce correlation, as with WORKER_NAME).
  • Complements but is distinct from CB-307 (reliable primary delivery): CB-306 makes spawn
    honest about reachability; CB-307 makes worker→primary delivery reliable once both are live.
  • Related existing guards: async+poll (CB-107), completion fallback (CB-106), status-gated injector +
    WorkerPresence, idle_ttl reaper (CB-303), orphan reap (CB-117).

Layer: session (SessionManager presence FSM) + inject/presence + spawn verbs.
Type: resilience / guard. Stage: hardening (post Stage-3/4).

## Problem `bridge_spawn` returns a session as soon as the worker **pane/agent** is created over herdr, but a freshly spawned Claude Code worker is not actually reachable over the bus until its **bridge MCP client has connected and handshaked** with the daemon. Between those two moments the worker exists but cannot receive an injected turn. The worst-case manifestation is the **pre-REPL folder-trust stall**: a worker launched in a not-yet-trusted directory sits at the CLI trust prompt and never loads its MCP config, so its bridge client never connects. A `bridge_send` to that session blocks against a peer that will never answer, and the caller only learns via a timeout (~caller MCP client cap, ~60s) — presenting as a silent "communication break" rather than a spawn failure. Root asymmetry: worker delivery is a **status-gated push** (`Injector.enqueue`, gated on `agent_status ∈ {idle, blocked}`), but "the pane is up" is inferred from herdr, not from a positive signal that the *bridge protocol* is live end-to-end. ## Proposal — a readiness gate Introduce an explicit **peer-readiness** state distinct from "pane exists": - Track a `presence` per session: `SPAWNING → READY → …` where `READY` is set only when the worker's bridge MCP client has completed its first handshake with the daemon (a positive protocol signal, not a herdr-output inference). - `bridge_send` / injector delivery must not target a session until it is `READY`. Before `READY`, either block on a bounded readiness wait or reject fast. - Add a **spawn-readiness timeout** (config knob, e.g. `spawn_ready_timeout`, default ~15–20s). If a peer doesn't reach `READY` within it, `bridge_spawn` fails **explicitly** ("peer did not become reachable — likely stuck pre-REPL / untrusted folder") instead of returning a half-live session that later times out on first send. - Surface `presence` in `bridge_list` / `GET /workers` roster so a stuck-spawning worker is visible. ## Why now This turns the pre-REPL stall (and any other "pane up but MCP not connected" case) from a delayed, misleading `bridge_send` timeout into a **fast, explicit spawn failure** at the right layer. ## Notes / scope - The handshake signal likely rides on the worker's first bridge MCP call reaching the daemon (server side sees the client connect); confirm the daemon can attribute that connection to a specific spawned session (name/nonce correlation, as with `WORKER_NAME`). - Complements but is distinct from **CB-307** (reliable primary delivery): CB-306 makes *spawn* honest about reachability; CB-307 makes *worker→primary* delivery reliable once both are live. - Related existing guards: async+poll (CB-107), completion fallback (CB-106), status-gated injector + `WorkerPresence`, idle_ttl reaper (CB-303), orphan reap (CB-117). **Layer:** `session` (SessionManager presence FSM) + `inject`/presence + spawn verbs. **Type:** resilience / guard. **Stage:** hardening (post Stage-3/4).
Author
Owner

Placement decision (2026-07-18): readiness belongs in the PeerLauncher adapter (ClaudeCodeLauncher), not core

"Is the worker up and usable in its terminal?" is peer-specific knowledge — a Claude Code worker
is ready once it clears the folder-trust prompt and its REPL/bridge-MCP is live; a Codex peer would
signal readiness differently. Per the CB-401 thesis (the launcher owns how to bring a peer of kind X
to life
), confirming the peer is alive-and-usable is part of birthing it. So the readiness gate is
delegated to dev.ltms.bridged.worker.ClaudeCodeLauncher (the PeerLauncher adapter);
SessionManager/core stays peer-neutral and never hardcodes "MCP handshake" or "past trust prompt."

SPI contract

  • PeerLauncher.spawn(SpawnRequest) blocks until the peer is READY and returns a ready
    PeerHandle, or throws a spawn-failure (PeerUnreachableException) on spawn_ready_timeout
    (~15–20s). Core gets a usable handle or a clean failure — it never hands a half-live session to
    bridge_send.
  • Core still owns the roster presence (SPAWNING/READY/FAILED in bridge_list /
    GET /workers), fed by the launcher, so a stuck spawn is visible.

Readiness signal (adapter-private, in ClaudeCodeLauncher)

  • Primary = terminal-usability probe via herdr — the launcher already owns AgentControl; it
    observes the pane clear the trust prompt and reach an interactive/idle Claude REPL. Directly
    catches the pre-REPL folder-trust stall ("usable in the worker terminal").
  • Optional strengthening = MCP handshake (worker's bridge client actually reached the daemon) for
    a proven round-trip — observed at the MCP-server layer and reported to the launcher/roster. Can
    land later; terminal-usability is the v1 bar.
  • On timeout → fail fast + auto-reap the dead pane (tie into CB-117) so a stuck spawn never
    becomes an orphan.

Bonus: resolves CB-401 deferral #2

This justifies PeerHandle.terminalId() — the launcher needs terminal/herdr coordinates to run
the readiness probe, so terminalId is a legitimate adapter coordinate, not a herdr leak. The
"revisit at CB-402" note can close in favour of keeping it.

## Placement decision (2026-07-18): readiness belongs in the PeerLauncher adapter (`ClaudeCodeLauncher`), not core "Is the worker up and usable in its terminal?" is **peer-specific** knowledge — a Claude Code worker is ready once it clears the folder-trust prompt and its REPL/bridge-MCP is live; a Codex peer would signal readiness differently. Per the CB-401 thesis (the launcher owns *how to bring a peer of kind X to life*), confirming the peer is alive-and-usable is part of birthing it. So the readiness gate is **delegated to `dev.ltms.bridged.worker.ClaudeCodeLauncher`** (the `PeerLauncher` adapter); `SessionManager`/core stays peer-neutral and never hardcodes "MCP handshake" or "past trust prompt." ### SPI contract - `PeerLauncher.spawn(SpawnRequest)` **blocks until the peer is READY** and returns a ready `PeerHandle`, or throws a spawn-failure (`PeerUnreachableException`) on `spawn_ready_timeout` (~15–20s). Core gets a usable handle or a clean failure — it never hands a half-live session to `bridge_send`. - Core still owns the **roster presence** (`SPAWNING`/`READY`/`FAILED` in `bridge_list` / `GET /workers`), fed by the launcher, so a stuck spawn is visible. ### Readiness signal (adapter-private, in `ClaudeCodeLauncher`) - **Primary = terminal-usability probe via herdr** — the launcher already owns `AgentControl`; it observes the pane clear the trust prompt and reach an interactive/idle Claude REPL. Directly catches the **pre-REPL folder-trust stall** ("usable in the worker terminal"). - **Optional strengthening = MCP handshake** (worker's bridge client actually reached the daemon) for a proven round-trip — observed at the MCP-server layer and reported to the launcher/roster. Can land later; terminal-usability is the v1 bar. - On timeout → fail fast **+ auto-reap the dead pane** (tie into CB-117) so a stuck spawn never becomes an orphan. ### Bonus: resolves CB-401 deferral #2 This justifies `PeerHandle.terminalId()` — the launcher **needs** terminal/herdr coordinates to run the readiness probe, so `terminalId` is a legitimate adapter coordinate, not a herdr leak. The "revisit at CB-402" note can close in favour of keeping it.
Author
Owner

CB-306 shipped — merged to main @ 7dd6c46 (design note 3a5cdc5), pushed.

What landed (launcher-owned terminal-readiness gate, per the design in docs/CB-306-Spawn-Readiness-Gate.md):

  • ClaudeCodeLauncher.spawn(SpawnRequest) now blocks until the worker pane reports an injectable herdr state (IDLE/BLOCKED/DONE via AgentStatus.injectable()) or the timeout elapses. On timeout it self-reaps the pane (stop(paneId)) and throws the new PeerUnreachableException (dev.ltms.bridged.peer) — no orphan left behind.
  • Gate is opt-out: spawnReadyTimeoutMs == 0 disables it (legacy non-blocking spawn). Config knobs spawn_ready_timeout_ms (default 20000) + spawn_ready_poll_ms (default 300) in BridgedConfig + bridged.example.yaml.
  • Testability seam: injectable monotonic clock (LongSupplier) + Runnable sleeper so unit tests never real-sleep.
  • Clean failure propagation: BridgeMcp.spawn → tool error, BridgedApp → HTTP 502 spawn_timeout (not an uncaught 500). SessionManager.acquire registers no session on a spawn throw (new test proves the roster stays empty).

Design note — what this ticket did not rebuild: the SPAWNING→READY transition already existed in core (SessionManager/PresenceBridge.markPresent() flips on first MCP contact from the worker — the delivery lifecycle). CB-306's gate is the complementary terminal-side signal owned by the launcher and is deliberately layered on top; the MCP-contact transition is untouched.

Verification (primary gate — the worker can't run these):

  • mvn clean install on main: BUILD SUCCESS, 188 tests, 0 failures / 0 errors / 0 skipped (was 183 pre-CB-306; +5 new).
  • IDE diagnostics: 7/8 changed files clean. The 2 "PeerUnreachableException never used" warnings are stale-index false positives — the class is compiled and referenced (throw + two catches); the IDE's sync/reverse-usage index couldn't refresh in this session, while the referencing files themselves analyze clean.

Delegation: implemented by an off-subscription gx10 worker over the bridge in a pre-trusted worktree, then primary-verified + integrated. Resolves CB-401 deferral #2 (the launcher legitimately needs PeerHandle.terminalId() for the probe). Closing.

**CB-306 shipped** — merged to `main` @ `7dd6c46` (design note `3a5cdc5`), pushed. **What landed** (launcher-owned terminal-readiness gate, per the design in `docs/CB-306-Spawn-Readiness-Gate.md`): - `ClaudeCodeLauncher.spawn(SpawnRequest)` now **blocks until the worker pane reports an injectable herdr state** (`IDLE`/`BLOCKED`/`DONE` via `AgentStatus.injectable()`) or the timeout elapses. On timeout it **self-reaps the pane** (`stop(paneId)`) and throws the new `PeerUnreachableException` (`dev.ltms.bridged.peer`) — no orphan left behind. - Gate is **opt-out**: `spawnReadyTimeoutMs == 0` disables it (legacy non-blocking spawn). Config knobs `spawn_ready_timeout_ms` (default 20000) + `spawn_ready_poll_ms` (default 300) in `BridgedConfig` + `bridged.example.yaml`. - **Testability seam**: injectable monotonic clock (`LongSupplier`) + `Runnable` sleeper so unit tests never real-sleep. - Clean failure propagation: `BridgeMcp.spawn` → tool error, `BridgedApp` → HTTP 502 `spawn_timeout` (not an uncaught 500). `SessionManager.acquire` registers **no** session on a spawn throw (new test proves the roster stays empty). **Design note — what this ticket did *not* rebuild:** the `SPAWNING→READY` transition already existed in core (`SessionManager`/`PresenceBridge.markPresent()` flips on first MCP contact from the worker — the *delivery* lifecycle). CB-306's gate is the complementary **terminal-side** signal owned by the launcher and is deliberately layered on top; the MCP-contact transition is untouched. **Verification (primary gate — the worker can't run these):** - `mvn clean install` on `main`: **BUILD SUCCESS**, **188 tests, 0 failures / 0 errors / 0 skipped** (was 183 pre-CB-306; +5 new). - IDE diagnostics: 7/8 changed files clean. The 2 "`PeerUnreachableException` never used" warnings are **stale-index false positives** — the class is compiled and referenced (throw + two catches); the IDE's sync/reverse-usage index couldn't refresh in this session, while the referencing files themselves analyze clean. **Delegation:** implemented by an off-subscription gx10 worker over the bridge in a pre-trusted worktree, then primary-verified + integrated. Resolves CB-401 deferral #2 (the launcher legitimately needs `PeerHandle.terminalId()` for the probe). Closing.
ltms closed this issue 2026-07-18 16:31:27 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#4