fleetd #176: fail-fast + safe UNKNOWN refinement in the spawn-readiness gate #233

Merged
ltms merged 2 commits from worker/cb176-spawn-gate-3d75c8-2 into main 2026-09-03 04:20:23 +02:00
Member

fleetd #176 — spawn-readiness gate (HerdrPeerLauncher.waitUntilInjectableOrThrow)

Scope was limited to the two fixes described in the brief. The "count the subscription seat" idea from the ticket was ignored per instructions (comment 2 already disproved it) — no seat accounting was built.

Fix 1 — fail fast on a herdr *_not_found mid-wait (implemented)

agents.get(paneId) inside the poll loop was unguarded. When herdr answers *_not_found (the backend process exited, not "slow"), that exception used to propagate raw out of waitUntilInjectableOrThrow: stop(paneId) was skipped (pane/tab leaked) and the caller got a HerdrException instead of PeerUnreachableException.

Now: the agents.get call is wrapped in try/catch. isAlreadyGone(e) (the existing helper) routes into a new failFastOnGoneBackend that:

  • stops waiting immediately — does not keep polling out the rest of spawnReadyTimeoutMs,
  • runs the same teardown the timeout path runs (stop(paneId)),
  • throws PeerUnreachableException naming the pane, elapsed ms, and the pane tail (via the existing readPaneQuietly), and says the process exited, not that the pane was slow.

Any other HerdrException still propagates unchanged (verified by a dedicated test).

Fix 2 — safe UNKNOWN refinement (implemented)

Wired StatusRefiner (the same one StatusPoller uses) into the gate, but only under two guards:

  1. Per-adapter: refinement runs only when namePrefix.equals("claude"). StatusRefiner.classify is documented as Claude-Code-TUI-specific; the opencode adapter (namePrefix = "opencode") never reaches it, so a confidently-wrong classification on the wrong TUI cannot happen.
  2. Corroborated liveness: a refined non-UNKNOWN result is accepted only when the same agents.get(paneId) sample also reports a non-null agentType. This closes the trap named in the brief: a pane whose backend already exited settles at a bare shell prompt that can also contain ❯, which StatusRefiner.classify reads as "idle at the Claude Code TUI prompt". Without the agentType corroboration that dead pane would misclassify as injectable — worse than today's timeout.

Status and agentType are read from one agents.get(paneId) call (I switched the loop from agents.status(paneId) to agents.get(paneId) — status() was already just get(target).status() under the hood, so this is the same underlying herdr call, not an extra one).

Refinement runs only when the raw status is UNKNOWN — verified with a dedicated test (refinementNeverFiresWhenRawStatusIsAlreadyInjectable) asserting agent.read is never called when the raw status is already IDLE. A healthy spawn therefore adds zero extra herdr calls; this matches StatusRefiner's own documented cost model (one extra agent.read per poll only while persistently UNKNOWN).

What I found about agentType

Agent.agentType (JSON field "agent") is documented as "detected agent kind... or null before herdr has detected it". I did not independently verify against a live herdr daemon that it is also null for a pane sitting at a bare shell after the backend exited (the brief flagged this as unverified too) — I trusted the brief's stated belief and built the corroboration on it, since it is the only liveness signal available in the same agent.get sample. This should be confirmed against real herdr behavior before this gate is trusted in a dogfood run.

spawnReadyTimeoutMs / config defaults

Unchanged — no config or default touched.

Tests

Added to ClaudeCodeLauncherTest (and matching fixtures added to FakeHerdr: agentType(String), agentGetFailsWithAfter(int okCalls, String code)):

  • spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout — timeout set to 60s, backend reports pane_not_found after one normal poll; asserts the clock never reaches anywhere near 60s, exactly 2 agent.get calls happened (no further polling), the message says "exited", and the pane is closed. This is acceptance criterion 2.
  • spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged — an internal_error code must NOT be treated as fail-fast; it propagates as HerdrException and the gate does not tear the pane down itself (extra coverage beyond the brief, to prove fix 1 doesn't over-fire).
  • refinedIdleIsNotAcceptedWhenAgentTypeIsNull — pane content contains ❯, agentType is null, status stays UNKNOWN forever; asserts the gate waits out the full timeout and closes the pane rather than accepting the trap. This is acceptance criterion 3.
  • refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness — the positive case: UNKNOWN + non-null agentType + a genuine idle-footer marker (? for shortcuts) resolves to injectable well before the timeout. (Not explicitly required by the acceptance criteria, added for confidence that fix 2 is not a no-op.)
  • refinementNeverFiresWhenRawStatusIsAlreadyInjectable — raw status already IDLE, pane content deliberately looks like WORKING (esc to interrupt); asserts the spawn still succeeds and agent.read was never called. This is acceptance criterion 4.

Build

Ran unpiped from the repo root:

mvn clean install > /tmp/cb176.log 2>&1; echo "EXIT=$?"

EXIT=0. Full suite: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.

Caveats for review

  • The agentType-is-null-on-a-bare-shell assumption is not verified against live herdr — see above.
  • The per-adapter guard is a literal "claude".equals(namePrefix) check inside the base class rather than a new adapter-overridable hook — kept minimal since HerdrPeerLauncher already owns namePrefix and this avoids touching ClaudeCodeLauncher/OpenCodeLauncher.
  • Out of scope, not investigated further, spotted while reading this file: checkPaneCommandFits and startAwaitingShellPrompt's SHELL_READY_RETRIES loop poll on herdr but don't obviously distinguish "shell still booting" from "shell will never become available" beyond a fixed retry count — same "slow vs dead" shape as this ticket, in a different place. Not fixed here.

Worktree root: /Users/dai.ha/LTMS/.bridged-worktrees/169147-2
Branch: worker/cb176-spawn-gate-3d75c8-2
Files changed:

  • fleetd/src/main/java/dev/ltms/fleet/member/HerdrPeerLauncher.java
  • fleetd/src/test/java/dev/ltms/fleet/herdr/FakeHerdr.java
  • fleetd/src/test/java/dev/ltms/fleet/member/ClaudeCodeLauncherTest.java

Update — review follow-up (guard-a coverage)

The lead reviewed the production diff and confirmed it independently (NAME_PREFIX values, the status()==get().status() equivalence, and that *_not_found already propagated out of spawn before this change). One gap was found before merge: guard (a) — the "claude".equals(namePrefix) per-adapter check in refinedInjectable — had no test. Guard (b) (the agentType corroboration) had both a positive and a negative test, but nothing proved guard (a) actually does anything, and a live opencode pane always reports a non-null agentType, so guard (b) alone can't catch a broken guard (a).

Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt to OpenCodeLauncherTest:

  • agentType("opencode") — non-null, satisfies guard (b) on its own, so it cannot be what makes the test pass
  • pane content containing ❯ — what StatusRefiner.classify would read as IDLE
  • raw status UNKNOWN throughout, short timeout
  • asserts the gate still times out (PeerUnreachableException, clock reaches the full timeout)
  • asserts no agent.read call used source="detection" (StatusRefiner.PROBE_SOURCE) — proving the refiner was never even reached for this pane, not merely that its answer was discarded

Note on that last assertion: the lead's suggested assertFalse(herdr.called("agent.read")) does not hold here — HerdrPeerLauncher.readPaneQuietly reads the pane tail (source="recent") for the timeout-log message on every timeout, regardless of adapter, so agent.read genuinely is called elsewhere in this same test. I scoped the assertion to the refiner's own probe source ("detection") instead, which is strictly stronger evidence for the actual claim (the refiner itself was never reached) and does hold.

Mutation check performed and reverted (not committed): temporarily changed if (!"claude".equals(namePrefix)) to if (false) in HerdrPeerLauncher.refinedInjectable. Reran just the new test — it failed: org.opentest4j.AssertionFailedError: Expected dev.ltms.fleet.peer.PeerUnreachableException to be thrown, but nothing was thrown. Restored the guard, reran — Tests run: 1, Failures: 0, Errors: 0, Skipped: 0. Confirmed the guard's own file (HerdrPeerLauncher.java) has no diff before this commit (git status/git diff clean on it).

Build: mvn clean install > /tmp/cb176-r2.log 2>&1; echo "EXIT=$?" → EXIT=0. Full log: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.

Files changed in this follow-up: fleetd/src/test/java/dev/ltms/fleet/member/OpenCodeLauncherTest.java only.

## fleetd #176 — spawn-readiness gate (HerdrPeerLauncher.waitUntilInjectableOrThrow) Scope was limited to the two fixes described in the brief. The "count the subscription seat" idea from the ticket was ignored per instructions (comment 2 already disproved it) — no seat accounting was built. ### Fix 1 — fail fast on a herdr `*_not_found` mid-wait (implemented) `agents.get(paneId)` inside the poll loop was unguarded. When herdr answers `*_not_found` (the backend process exited, not "slow"), that exception used to propagate raw out of `waitUntilInjectableOrThrow`: `stop(paneId)` was skipped (pane/tab leaked) and the caller got a `HerdrException` instead of `PeerUnreachableException`. Now: the `agents.get` call is wrapped in try/catch. `isAlreadyGone(e)` (the existing helper) routes into a new `failFastOnGoneBackend` that: - stops waiting immediately — does **not** keep polling out the rest of `spawnReadyTimeoutMs`, - runs the same teardown the timeout path runs (`stop(paneId)`), - throws `PeerUnreachableException` naming the pane, elapsed ms, and the pane tail (via the existing `readPaneQuietly`), and says the process **exited**, not that the pane was slow. Any other `HerdrException` still propagates unchanged (verified by a dedicated test). ### Fix 2 — safe UNKNOWN refinement (implemented) Wired `StatusRefiner` (the same one `StatusPoller` uses) into the gate, but only under two guards: 1. **Per-adapter**: refinement runs only when `namePrefix.equals("claude")`. `StatusRefiner.classify` is documented as Claude-Code-TUI-specific; the opencode adapter (`namePrefix = "opencode"`) never reaches it, so a confidently-wrong classification on the wrong TUI cannot happen. 2. **Corroborated liveness**: a refined non-UNKNOWN result is accepted only when the *same* `agents.get(paneId)` sample also reports a non-null `agentType`. This closes the trap named in the brief: a pane whose backend already exited settles at a bare shell prompt that can also contain `❯`, which `StatusRefiner.classify` reads as "idle at the Claude Code TUI prompt". Without the `agentType` corroboration that dead pane would misclassify as injectable — worse than today's timeout. Status and agentType are read from **one** `agents.get(paneId)` call (I switched the loop from `agents.status(paneId)` to `agents.get(paneId)` — `status()` was already just `get(target).status()` under the hood, so this is the same underlying herdr call, not an extra one). Refinement runs **only** when the raw status is UNKNOWN — verified with a dedicated test (`refinementNeverFiresWhenRawStatusIsAlreadyInjectable`) asserting `agent.read` is never called when the raw status is already IDLE. A healthy spawn therefore adds zero extra herdr calls; this matches `StatusRefiner`'s own documented cost model (one extra `agent.read` per poll only while persistently UNKNOWN). ### What I found about `agentType` `Agent.agentType` (JSON field `"agent"`) is documented as "detected agent kind... or null before herdr has detected it". I did **not** independently verify against a live herdr daemon that it is also null for a pane sitting at a bare shell after the backend exited (the brief flagged this as unverified too) — I trusted the brief's stated belief and built the corroboration on it, since it is the only liveness signal available in the same `agent.get` sample. This should be confirmed against real herdr behavior before this gate is trusted in a dogfood run. ### `spawnReadyTimeoutMs` / config defaults Unchanged — no config or default touched. ### Tests Added to `ClaudeCodeLauncherTest` (and matching fixtures added to `FakeHerdr`: `agentType(String)`, `agentGetFailsWithAfter(int okCalls, String code)`): - `spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout` — timeout set to 60s, backend reports `pane_not_found` after one normal poll; asserts the clock never reaches anywhere near 60s, exactly 2 `agent.get` calls happened (no further polling), the message says "exited", and the pane is closed. This is acceptance criterion 2. - `spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged` — an `internal_error` code must NOT be treated as fail-fast; it propagates as `HerdrException` and the gate does not tear the pane down itself (extra coverage beyond the brief, to prove fix 1 doesn't over-fire). - `refinedIdleIsNotAcceptedWhenAgentTypeIsNull` — pane content contains `❯`, `agentType` is `null`, status stays UNKNOWN forever; asserts the gate waits out the full timeout and closes the pane rather than accepting the trap. This is acceptance criterion 3. - `refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness` — the positive case: UNKNOWN + non-null `agentType` + a genuine idle-footer marker (`? for shortcuts`) resolves to injectable well before the timeout. (Not explicitly required by the acceptance criteria, added for confidence that fix 2 is not a no-op.) - `refinementNeverFiresWhenRawStatusIsAlreadyInjectable` — raw status already IDLE, pane content deliberately looks like `WORKING` (`esc to interrupt`); asserts the spawn still succeeds and `agent.read` was never called. This is acceptance criterion 4. ### Build Ran unpiped from the repo root: ``` mvn clean install > /tmp/cb176.log 2>&1; echo "EXIT=$?" ``` `EXIT=0`. Full suite: `Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`. ### Caveats for review - The `agentType`-is-null-on-a-bare-shell assumption is not verified against live herdr — see above. - The per-adapter guard is a literal `"claude".equals(namePrefix)` check inside the base class rather than a new adapter-overridable hook — kept minimal since `HerdrPeerLauncher` already owns `namePrefix` and this avoids touching `ClaudeCodeLauncher`/`OpenCodeLauncher`. - Out of scope, not investigated further, spotted while reading this file: `checkPaneCommandFits` and `startAwaitingShellPrompt`'s `SHELL_READY_RETRIES` loop poll on herdr but don't obviously distinguish "shell still booting" from "shell will never become available" beyond a fixed retry count — same "slow vs dead" shape as this ticket, in a different place. Not fixed here. **Worktree root:** `/Users/dai.ha/LTMS/.bridged-worktrees/169147-2` **Branch:** `worker/cb176-spawn-gate-3d75c8-2` **Files changed:** - `fleetd/src/main/java/dev/ltms/fleet/member/HerdrPeerLauncher.java` - `fleetd/src/test/java/dev/ltms/fleet/herdr/FakeHerdr.java` - `fleetd/src/test/java/dev/ltms/fleet/member/ClaudeCodeLauncherTest.java` --- ## Update — review follow-up (guard-a coverage) The lead reviewed the production diff and confirmed it independently (NAME_PREFIX values, the `status()==get().status()` equivalence, and that `*_not_found` already propagated out of spawn before this change). One gap was found before merge: **guard (a)** — the `"claude".equals(namePrefix)` per-adapter check in `refinedInjectable` — had no test. Guard (b) (the `agentType` corroboration) had both a positive and a negative test, but nothing proved guard (a) actually does anything, and a live opencode pane always reports a non-null `agentType`, so guard (b) alone can't catch a broken guard (a). Added `opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt` to `OpenCodeLauncherTest`: - `agentType("opencode")` — non-null, satisfies guard (b) on its own, so it cannot be what makes the test pass - pane content containing `❯` — what `StatusRefiner.classify` would read as IDLE - raw status UNKNOWN throughout, short timeout - asserts the gate still times out (`PeerUnreachableException`, clock reaches the full timeout) - asserts no `agent.read` call used `source="detection"` (`StatusRefiner.PROBE_SOURCE`) — proving the refiner was never even reached for this pane, not merely that its answer was discarded Note on that last assertion: the lead's suggested `assertFalse(herdr.called("agent.read"))` does **not** hold here — `HerdrPeerLauncher.readPaneQuietly` reads the pane tail (`source="recent"`) for the timeout-log message on every timeout, regardless of adapter, so `agent.read` genuinely is called elsewhere in this same test. I scoped the assertion to the refiner's own probe source (`"detection"`) instead, which is strictly stronger evidence for the actual claim (the refiner itself was never reached) and does hold. **Mutation check performed and reverted (not committed):** temporarily changed `if (!"claude".equals(namePrefix))` to `if (false)` in `HerdrPeerLauncher.refinedInjectable`. Reran just the new test — it failed: `org.opentest4j.AssertionFailedError: Expected dev.ltms.fleet.peer.PeerUnreachableException to be thrown, but nothing was thrown.` Restored the guard, reran — `Tests run: 1, Failures: 0, Errors: 0, Skipped: 0`. Confirmed the guard's own file (`HerdrPeerLauncher.java`) has no diff before this commit (`git status`/`git diff` clean on it). Build: `mvn clean install > /tmp/cb176-r2.log 2>&1; echo "EXIT=$?"` → `EXIT=0`. Full log: `Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`. Files changed in this follow-up: `fleetd/src/test/java/dev/ltms/fleet/member/OpenCodeLauncherTest.java` only.
agent added 1 commit 2026-09-03 04:11:32 +02:00
fleetd #176: fail fast when the spawn gate's backend exits mid-wait, and safely refine its UNKNOWN status
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m26s
60b7e67b42
Fix 1: agents.get(paneId) inside waitUntilInjectableOrThrow was unguarded, so a
herdr *_not_found answer (the backend process exited) propagated as a raw
HerdrException instead of PeerUnreachableException, and skipped teardown
entirely, leaking the pane/tab. Now caught via isAlreadyGone(); fails
immediately (does not burn the rest of the timeout), runs the same teardown
the timeout path runs, and the exception message says the process exited
rather than that the pane was slow. Any other HerdrException still
propagates unchanged.

Fix 2: the gate now resolves a raw UNKNOWN into StatusRefiner's pane-content
classification, like StatusPoller already does mid-life. Guarded against the
trap noted in the ticket: a pane whose backend exited settles at a bare shell
prompt that can also contain the "❯" glyph classify() reads as idle. So a
refined result is accepted only when the corroborating agentType from the
SAME agent.get sample is non-null, and only for the "claude" adapter (namePrefix)
since StatusRefiner.classify is written for the Claude Code TUI only. Refine
only runs when the raw status is UNKNOWN, so a healthy spawn adds zero extra
herdr calls.

Tests added to ClaudeCodeLauncherTest (FakeHerdr gained agentType()/
agentGetFailsWithAfter() fixtures):
 - spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout
 - spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged
 - refinedIdleIsNotAcceptedWhenAgentTypeIsNull
 - refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness
 - refinementNeverFiresWhenRawStatusIsAlreadyInjectable

mvn clean install: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
agent added 1 commit 2026-09-03 04:17:17 +02:00
fleetd #176 review: cover the namePrefix guard, the adapter-refinement corroboration that had no test
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m14s
432c1d92d1
The agentType corroboration (guard b) had positive and negative tests, but the
per-adapter namePrefix guard (guard a) had none — nothing proved that an
opencode pane can never reach StatusRefiner.classify, only that a claude pane
with a null agentType is rejected. Since a live opencode pane always reports a
non-null agentType ("opencode"), guard (b) alone cannot catch a broken guard
(a).

Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt
to OpenCodeLauncherTest: agentType("opencode") (non-null, satisfies guard b on
its own) + pane content containing "❯" (would classify as IDLE) + raw status
UNKNOWN throughout. Asserts the gate still times out, and that no agent.read
call used source=detection (StatusRefiner.PROBE_SOURCE) — proving the refiner
was never even reached, not just that its answer was discarded. A blanket
"agent.read is never called" does not hold here: HerdrPeerLauncher.readPaneQuietly
reads the pane tail (source=recent) for the timeout log on every timeout,
regardless of adapter, so the assertion is scoped to the refiner's own probe
source instead.

Mutation check performed and reverted before this commit: temporarily changed
`if (!"claude".equals(namePrefix))` to `if (false)` in
HerdrPeerLauncher.refinedInjectable — the new test failed
("Expected PeerUnreachableException to be thrown, but nothing was thrown"),
confirming it actually exercises the guard. Restored the guard and reran —
test passes again (1/1). No production code changed in this commit.

mvn clean install: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
ltms merged commit 0f08b93659 into main 2026-09-03 04:20:23 +02:00
ltms deleted branch worker/cb176-spawn-gate-3d75c8-2 2026-09-03 04:20:23 +02:00
Sign in to join this conversation.