Compare commits

...

5 Commits

Author SHA1 Message Date
Dai Ha 7aa45b234b fleetd #201/#227: refine backend outage work 2026-09-03 10:05:45 +07:00
ltms 0f08b93659 fleetd #176: spawn gate fails fast on a dead backend, and refines UNKNOWN safely
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m47s
Two fixes to HerdrPeerLauncher.waitUntilInjectableOrThrow.

Fix 1: the status call is now guarded. A herdr *_not_found answer — what happens
when the backend process exited rather than being slow — used to escape as a raw
HerdrException, skipping stop() and leaking the pane and tab, with a message that
blamed a slow pane. It now fails immediately, runs the same teardown, and says the
process exited.

Fix 2: the gate can now resolve UNKNOWN with the same StatusRefiner the status
poller uses, behind two guards. It only runs for the claude adapter, because
StatusRefiner.classify reads a Claude Code TUI. And a refined result is accepted
only when the same agents.get sample still reports a non-null agentType. That
second guard closes a trap: a dead pane sits at a shell prompt containing the same
❯ glyph the classifier reads as idle, so refining without corroboration would turn
"the backend died" into "ready to inject".

Verified by the lead before merge: NAME_PREFIX really is "claude"/"opencode" so the
adapter split holds; AgentControl.status(t) was already get(t).status(), so moving
the loop to get() adds no herdr call; and a *_not_found already failed the spawn
before this change, so fix 1 improves an existing failure rather than creating one.
The adapter guard was mutation-tested independently (set it to `if (false)`, the new
opencode test fails with "Expected PeerUnreachableException to be thrown, but nothing
was thrown"; restored, it passes). Independent build: 1123 tests, 0 failures,
0 errors, 0 skipped.

Not fixed, and not claimed: that herdr reports a null agentType for a bare shell
pane is unverified against a live daemon. To be proved by a live spawn on both
backends after deploy. The seat-accounting suggestion in #176 was deliberately not
built — its cause was tested in that issue and not reproduced.
2026-09-03 04:20:23 +02:00
Dai Ha 432c1d92d1 fleetd #176 review: cover the namePrefix guard, the adapter-refinement corroboration that had no test
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m14s
The agentType corroboration (guard b) had positive and negative tests, but the
per-adapter namePrefix guard (guard a) had none — nothing proved that an
opencode pane can never reach StatusRefiner.classify, only that a claude pane
with a null agentType is rejected. Since a live opencode pane always reports a
non-null agentType ("opencode"), guard (b) alone cannot catch a broken guard
(a).

Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt
to OpenCodeLauncherTest: agentType("opencode") (non-null, satisfies guard b on
its own) + pane content containing "❯" (would classify as IDLE) + raw status
UNKNOWN throughout. Asserts the gate still times out, and that no agent.read
call used source=detection (StatusRefiner.PROBE_SOURCE) — proving the refiner
was never even reached, not just that its answer was discarded. A blanket
"agent.read is never called" does not hold here: HerdrPeerLauncher.readPaneQuietly
reads the pane tail (source=recent) for the timeout log on every timeout,
regardless of adapter, so the assertion is scoped to the refiner's own probe
source instead.

Mutation check performed and reverted before this commit: temporarily changed
`if (!"claude".equals(namePrefix))` to `if (false)` in
HerdrPeerLauncher.refinedInjectable — the new test failed
("Expected PeerUnreachableException to be thrown, but nothing was thrown"),
confirming it actually exercises the guard. Restored the guard and reran —
test passes again (1/1). No production code changed in this commit.

mvn clean install: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:17:10 +07:00
Dai Ha 60b7e67b42 fleetd #176: fail fast when the spawn gate's backend exits mid-wait, and safely refine its UNKNOWN status
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m26s
Fix 1: agents.get(paneId) inside waitUntilInjectableOrThrow was unguarded, so a
herdr *_not_found answer (the backend process exited) propagated as a raw
HerdrException instead of PeerUnreachableException, and skipped teardown
entirely, leaking the pane/tab. Now caught via isAlreadyGone(); fails
immediately (does not burn the rest of the timeout), runs the same teardown
the timeout path runs, and the exception message says the process exited
rather than that the pane was slow. Any other HerdrException still
propagates unchanged.

Fix 2: the gate now resolves a raw UNKNOWN into StatusRefiner's pane-content
classification, like StatusPoller already does mid-life. Guarded against the
trap noted in the ticket: a pane whose backend exited settles at a bare shell
prompt that can also contain the "❯" glyph classify() reads as idle. So a
refined result is accepted only when the corroborating agentType from the
SAME agent.get sample is non-null, and only for the "claude" adapter (namePrefix)
since StatusRefiner.classify is written for the Claude Code TUI only. Refine
only runs when the raw status is UNKNOWN, so a healthy spawn adds zero extra
herdr calls.

Tests added to ClaudeCodeLauncherTest (FakeHerdr gained agentType()/
agentGetFailsWithAfter() fixtures):
 - spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout
 - spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged
 - refinedIdleIsNotAcceptedWhenAgentTypeIsNull
 - refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness
 - refinementNeverFiresWhenRawStatusIsAlreadyInjectable

mvn clean install: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:10:49 +07:00
ltms 3fbd43fe3f fleetd #175: read back the model opencode actually resolved, and quarantine a silent substitution
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m29s
Verified by the lead before merge. Read the full production diff, checkModelMatch, parseModel and actualModelForDirectory. Confirmed the check runs on the real #209 late-resolve path (not at spawn, which is why #203 was closed), that unknown/incomplete evidence never quarantines, and that claude-code is structurally excluded because SessionAwareHandle is only built by OpenCodeLauncher.spawn(). Round 2 closed the one gap I found: model JSON with an id but no providerID used to read as a mismatch for a provider-prefixed profile. Independent build: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped.
2026-09-02 13:01:33 +02:00
5 changed files with 851 additions and 9 deletions
+531
View File
@@ -0,0 +1,531 @@
# CB-201 and CB-227 refinement
Date: 2026-09-03
## Decision
#201 and #227 are one delivery program, but they are not one implementation unit.
#201 has a real seam: `CompletionResolver` can publish a typed backend-error event only after its
waiter resolution wins. #227 can consume that event without knowing any pane text. The classifier
must land before the final #227 wiring. However, the policy engine, roster state, and lead nudge can
be built in parallel with the classifier.
I propose five units. Units 1 to 4 own separate files and can run in parallel. Unit 5 owns all
composition files and lands after them. It also depends on the #234 defect 2 fix named in the task.
```mermaid
flowchart LR
U1["Unit 1: typed backend-error classification"] --> U5["Unit 5: wire policy, spawn gate, and fleet views"]
U2["Unit 2: credential outage policy"] --> U5
U3["Unit 3: lead outage nudge"] --> U5
U4["Unit 4: durable member outcome"] --> U5
D234["#234 defect 2: fail-loud target resolution"] --> U5
```
*Figure 1. Four file-disjoint foundations feed one composition unit.*
This split keeps `Fleetd.java` under one owner. It also keeps every other production file under one
unit in this plan.
## Evidence checked in the current branch
I read both issue pages in full. Each page reports zero comments.
| Evidence | What the code says now |
|---|---|
| `inject/CompletionResolver.java:229-237` | A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today. |
| `inject/CompletionResolver.java:260-276` and `:332-367` | #211 already added raw-screen classification when `lastAssistantBlock` is empty. The “dead-code question” in #201 is stale on this branch. |
| `inject/CompletionResolver.java:288-317` | Exhaustion wins before the hard-coded `API Error:` match. A backend error then goes through generic `fail(...)`. |
| `inject/CompletionResolver.java:311-316` | The code admits that the pattern is a heuristic. A member report which quotes an API error may match it. |
| `inject/CompletionResolver.java:449-467` | Startup coverage exists only for `exhaustedPattern`. |
| `inject/ExhaustedPatternLookup.java:13-25` | The current lookup and explicit `none()` value are a good shape for the new classifier seam. |
| `Fleetd.java:196-207` | One `BackendQuarantine` is shared by placement and the exhaustion sink. Its cooldown comes from `quarantineCooldownSeconds`. |
| `Fleetd.java:322-363` | Pattern compilation, target-to-profile lookup, and the live `ExhaustionSink` are composed in `Fleetd.main`. The sink on this branch still ends in `.ifPresent(...)`. This plan assumes #234 replaces that silent path. |
| `placement/BackendQuarantine.java:60-87` | A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock. |
| `member/CompositePeerLauncher.java:260-317` | Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off. |
| `member/CompositePeerLauncher.java:347-379` | Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”. |
| `placement/PlacementContext.java:10-22` and `PlacementPolicyUtil.java:14-83` | Automatic placement has only one transient exclusion set named `quarantined`. Reusing it would make outage errors say “backend exhausted”. |
| `mcp/FleetMcp.java:913-1025` | `fleet_list` sets `free: 0` and adds `credentialId` plus `quarantinedForSeconds` when quarantine is active. |
| `session/MemberSession.java:51-59` | The roster has `DONE` and generic `FAILED`, but no backend-error state or stored reason. |
| `session/SessionManager.java:695-773` | A normal boundary moves `BUSY` to `DONE`. A failure moves any non-released session to `FAILED`. The async completion resolver can race the `DONE` update. |
| `session/SessionManager.java:648-687` | `rosterView` reports the session state, but it reports no terminal reason. |
| `msg/MessageService.java:922-940` | CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage. |
| `msg/ReplyPushLoop.java:20-48` | Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns. |
| `msg/ReplyPushLoop.java:305-395` | Each push entry point resolves the owning lead through `PrimaryRegistry`. Missing ownership is logged and the durable or pending item remains the backstop. |
| `msg/ReplyPushLoop.java:496-547` | One tick builds one combined nudge. Pending items have separate reminder counts. |
| `health/FleetHealthMonitor.java:91-143` | Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications. |
| `health/FleetHealthMonitor.java:206-208` | `healthCoverage` means health enabled plus webhook configured. It does not describe lead-pane alerts. |
| `Fleetd.java:465-486` | Health stays `detection-only` without the webhook notification setting. |
I also read the related unit tests for `CompletionResolver`, `ReplyPushLoop`, `BackendQuarantine`,
`CompositePeerLauncher`, `PlacementPolicyUtil`, `SessionManager`, `MessageService`, and `FleetMcp`.
I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No
peer architect was named, so I did not exchange a design with one.
## Required behaviour
The policy should use these first values:
- Threshold: **2** classified backend errors.
- Window: **60 seconds**, measured from the first error to the second.
- Cool-off: **60 seconds**, starting when the threshold is reached.
- Correlation key: `credentialId`, never profile name and never error text.
- Incident rule: one active incident per credential. Errors during its cool-off do not extend it and
do not create more lead notices.
- Rearm rule: after cool-off ends, two fresh errors are needed for another incident.
Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window
fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without
turning a short backend fault into the default 1,800-second exhaustion quarantine.
A single classified error still fails its send and marks its member `backend_error`. It does not
cool a credential and does not send an outage notice. This is what “a single error changes nothing”
must mean at the credential level. It cannot mean that the failed member still looks successful.
```mermaid
sequenceDiagram
participant R1 as Resolver for member A
participant R2 as Resolver for member B
participant P as Outage policy
participant S as Spawn gate
participant N as Lead push loop
participant L as Lead pane
R1->>P: backend error for credential C
Note over P: Count 1, no cool-off
R2->>P: backend error for credential C within 60s
P->>P: Start one 60s incident
P->>S: Credential C is cooling off
P->>N: Queue one incident notice
N->>L: Inject when lead is idle, blocked, or done
L->>S: Request another spawn on credential C
S-->>L: Refuse and report remaining cool-off
```
*Figure 2. The second independent classification creates the fleet-level event.*
Against the 2026-09-01 case, the second failed member would start cool-off. `fleet_list` would show
zero free capacity and both members as `backend_error`. The push loop would inject one outage notice
even if the lead had not polled either ticket yet. The design reports the outage. It does not recover
uncommitted work from the members.
## Unit 1 — Typed backend-error classification
### Scope
Replace the direct hard-coded check inside `CompletionResolver` with a lookup and a sink. Keep the
public send result as a failed send. The typed internal event is the seam #227 consumes.
The lookup returns the pattern for a target. The sink receives the target, matched line, and full
failure reason. It fires only after `Rendezvous.resolveFailure(...)` wins for that exact captured
waiter. This copies the race rule already used by `ExhaustionSink`.
The classifier must run in all three current paths:
1. a normal non-empty assistant block;
2. the #211 raw scrape fallback;
3. a turn inside `MIN_TURN_NANOS`, before it becomes a generic too-fast failure.
In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic
failure or completion. A fast turn still fails when no configured pattern matches.
Keep `(?i)\bAPI Error\s*:` as a compatibility pattern for profiles without `errorPattern` until the
operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name
profiles using this weaker legacy default.
### Files owned
- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorPatternLookup.java`.
- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorSink.java`.
- Change `fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java`.
- Change `fleetd/src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java`.
No other unit may edit these files.
### Acceptance criteria
1. A target-specific error pattern matches a normal assistant block and resolves the send as failed.
2. The same match calls `BackendErrorSink` exactly once after the waiter resolution wins.
3. A late classification which loses to `fleet_reply` does not call the sink.
4. An exhausted line that also matches the generic error pattern stays `BACKEND_EXHAUSTED`. It calls
only `ExhaustionSink`.
5. A raw pane with leading Terminal User Interface (TUI) chrome and no assistant marker still uses
the #211 fallback and calls the backend-error sink.
6. A matching error inside the two-second floor is typed and sent to the sink. A non-matching fast
turn stays a generic failure.
7. An unchanged delivery baseline which contains old backend-error text is suppressed. It never
increments outage evidence.
8. A non-match keeps the existing completion result and text.
9. Constructors used by current callers keep compiling. They use the legacy default lookup and an
explicit inert sink until Unit 5 supplies the production objects.
10. Unit tests pass. The developer runs the focused test first, then `mvn clean install` from
`fleetd/`.
### Dependencies
None. Unit 1 can run with Units 2, 3, and 4.
Unit 5 depends on its new lookup, sink, and constructor.
### What to report back
- The exact classifier order in all three paths.
- The focused test command and result.
- The test which proves a losing waiter race does not publish an event.
- The test which proves a fast matching failure is typed.
- The final `mvn clean install` result.
- Any constructor kept only for transition and where Unit 5 replaces it.
## Unit 2 — Credential outage policy
### Scope
Build a small credential-keyed state machine. It accepts already-classified backend-error events.
It does not read pane text, profiles, sessions, or lead state.
Use an injected monotonic clock. A call records `credentialId`, target, and reason. It returns a new
incident only on the threshold crossing. The incident contains a stable event id, credential id,
the distinct affected targets, evidence count, window, and remaining cool-off.
This class owns both correlation and short cool-off. Keeping them together makes threshold crossing
and the cool-off deadline one atomic state change.
### Files owned
- Add `fleetd/src/main/java/dev/ltms/fleet/placement/BackendOutagePolicy.java`.
- Add `fleetd/src/test/java/dev/ltms/fleet/placement/BackendOutagePolicyTest.java`.
No other unit may edit these files.
### Acceptance criteria
1. One error creates no incident and no cool-off.
2. Two errors for one credential within 60 seconds create exactly one incident and a 60-second
cool-off.
3. Two errors more than 60 seconds apart do not create an incident.
4. The exact 60-second boundary has a pinned result. Use inclusive `<= 60s` so scheduler delay does
not discard evidence at the boundary.
5. Different credentials never share evidence.
6. Different profiles which supply the same credential id do share evidence. The policy itself only
sees the credential id.
7. More errors during active cool-off do not extend its deadline and do not return another incident.
8. After expiry, old evidence is cleared. Two fresh errors are needed to create the next incident.
9. Remaining seconds round up, matching `BackendQuarantine` reporting.
10. Concurrent second and third errors cannot return two incidents.
11. The focused tests and `mvn clean install` pass.
### Dependencies
None. Unit 2 can run with Units 1, 3, and 4.
Unit 5 depends on the policy API.
### What to report back
- The state transition table and locking method.
- The exact threshold, window, cool-off, and boundary rule.
- The test which proves one incident under concurrent calls.
- The focused test and `mvn clean install` results.
## Unit 3 — Lead outage nudge
### Scope
Add backend incidents as a fourth pending source in `ReplyPushLoop`. Do not create another scheduler
or call `AgentControl.send` from `Fleetd`. The existing combined per-lead schedule is the control
which prevents competing injected turns.
The entry point takes an incident id, affected worker targets, credential id, affected profile
names, and remaining cool-off. It resolves distinct owning leads through `PrimaryRegistry`.
Each `(incidentId, lead)` item is one-shot. It waits while the lead is not injectable. After one
successful `agents.send`, remove it. A send exception keeps it pending for a bounded retry. It never
uses the repeated reminder behaviour of an uncollected ticket.
Also add a fail-loud entry point for a classified target that Unit 5 cannot map to a credential. It
uses `PrimaryRegistry.nudgeTargetFor(target)` and says that correlation could not run. If no lead is
known, log at `WARN`, not `DEBUG`.
### Files owned
- Change `fleetd/src/main/java/dev/ltms/fleet/msg/ReplyPushLoop.java`.
- Change `fleetd/src/test/java/dev/ltms/fleet/msg/ReplyPushLoopTest.java`.
No other unit may edit these files.
### Acceptance criteria
1. One incident affecting two workers owned by one lead causes one successful pane injection.
2. Two affected workers owned by two leads cause one successful injection per affected lead. This
is one notice per event per lead, not one notice per member.
3. Repeating the same incident id is idempotent.
4. A busy or unknown lead is not injected. The item stays pending until the lead becomes injectable
or its attempt cap is reached.
5. After one successful injection, later ticks do not mention that incident again.
6. A failed `agents.send` is retried within the existing bound. A successful retry still gives only
one successful send.
7. A pending ticket and an outage incident for one lead appear in one combined nudge, not two
competing turns.
8. The text names the credential, profiles, affected workers, and remaining cool-off. It tells the
lead to run `fleet_list`.
9. An unmapped target produces a direct warning notice when a lead is known. If no lead is known,
the code logs a `WARN` naming the target and reason.
10. `stop()` clears incident state as it clears other push state.
11. Existing reply, ticket, and question tests stay green. The focused tests and
`mvn clean install` pass.
### Dependencies
None. The API uses plain values, not the Unit 2 incident class. This lets Unit 3 run in parallel.
Unit 5 adapts the Unit 2 incident into this entry point.
### What to report back
- The exact one-shot and retry rules.
- The test showing one combined nudge with a failed ticket.
- The test showing one successful send for two affected workers.
- The focused test and `mvn clean install` results.
## Unit 4 — Durable member backend-error outcome
### Scope
Make a classified backend failure remain visible after its ticket is collected or expires.
Add `BACKEND_ERROR` to `MemberSession.State`. Add a nullable failure detail to `MemberSession` and
render it as `failureReason` in `SessionManager.rosterView`. Add
`SessionManager.onBackendError(target, reason)`.
The transition must handle both completion orderings:
- `BUSY -> BACKEND_ERROR` when classification wins before the normal completion state update;
- `DONE -> BACKEND_ERROR` when the async resolver runs after `SessionManager.onTurnComplete`.
It must use a compare-and-set retry or another atomic update. `RELEASED` must never return to the
roster. A backend-error member is terminal and cannot accept another delivery.
### Files owned
- Change `fleetd/src/main/java/dev/ltms/fleet/session/MemberSession.java`.
- Change `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java`.
- Change `fleetd/src/test/java/dev/ltms/fleet/session/SessionManagerTest.java`.
No other unit may edit these files.
### Acceptance criteria
1. `onBackendError` moves a `BUSY` member to `BACKEND_ERROR` and stores the reason.
2. It also moves `DONE` to `BACKEND_ERROR`, covering the resolver race.
3. A later normal `onTurnComplete` cannot change `BACKEND_ERROR` back to `DONE`.
4. A released or unknown member is not recreated. The unknown case logs at `WARN` and returns an
explicit false result to its caller.
5. `onDelivered` refuses a `BACKEND_ERROR` member, just as it refuses generic `FAILED`.
6. `rosterView` reports `state: backend_error` and `failureReason` after the send ticket is gone.
7. Ordinary members do not gain a blank or invented `failureReason` field.
8. Existing constructors keep source compatibility for tests and adapters.
9. The focused tests and `mvn clean install` pass.
### Dependencies
None. Unit 4 can run with Units 1, 2, and 3.
Unit 5 calls the new session method from the production sink.
### What to report back
- The two race orderings and the tests for both.
- The exact roster JSON shape.
- The unknown-target result and log level.
- The focused test and `mvn clean install` results.
## Unit 5 — Production wiring, spawn gate, and fleet views
### Scope
Compose Units 1 to 4 in production. This is the only unit which edits `Fleetd.java`.
Add per-profile `errorPattern` config beside `exhaustedPattern`. Compile both once at startup. A
configured pattern wins over the legacy default. Report configured profiles and legacy-default
profiles separately at startup. A bad regex must stop startup with the profile and key in the
message.
Wire one production `BackendErrorSink` with this order:
1. mark the member `backend_error` with its reason;
2. resolve the profile and its current `effectiveCredentialId()` through the fail-loud #234 seam;
3. record the error in `BackendOutagePolicy`;
4. on a new incident, submit one event to `ReplyPushLoop`.
If target metadata cannot be resolved, do not end in `Optional.ifPresent`. Log an error and call the
Unit 3 unmapped-target notice. The failed send still reaches its ticket through CB-588.
Teach both spawn paths about a separate cool-off source. Exhaustion quarantine has priority when
both states are active. Automatic placement needs a distinct `coolingOff` set so its refusal does
not say “exhausted”.
Extend the MCP (Model Context Protocol) views:
- A cooling profile has `free: 0`, `credentialId`, and `coolingOffForSeconds` in `fleet_list`.
- It does not have `quarantinedForSeconds` unless exhaustion quarantine is also active.
- `fleet_profiles` has a separate `coolingOff` map, not an entry in `quarantined`.
- A direct spawn refusal says the credential is cooling off after repeated backend errors and gives
the remaining seconds.
Do not change `FleetHealthMonitor.coverage`. It still describes the periodic health webhook path.
Lead-pane outage delivery is a separate capability.
### Files owned
- `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java`
- `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java`
- `fleetd/src/main/java/dev/ltms/fleet/config/ConfigRef.java`
- `fleetd/src/main/java/dev/ltms/fleet/member/CompositePeerLauncher.java`
- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementContext.java`
- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementPolicyUtil.java`
- `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java`
- `fleetd/src/test/java/dev/ltms/fleet/config/FleetConfigTest.java`
- `fleetd/src/test/java/dev/ltms/fleet/config/ConfigRefTest.java`
- `fleetd/src/test/java/dev/ltms/fleet/member/CompositePeerLauncherTest.java`
- `fleetd/src/test/java/dev/ltms/fleet/placement/PlacementPolicyTest.java`
- `fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java`
- Add `fleetd/src/test/java/dev/ltms/fleet/BackendOutageFlowTest.java`.
- `fleetd/fleetd.example.yaml`
- `CLAUDE.md`
No earlier unit edits these files.
The lead, not a worker, must update `wiki/7-Use-Cases.md`, `wiki/9-Implementation.md`, and
`wiki/11-Features.md`. Project rules forbid workers from committing `wiki/`. The portable block in
`CLAUDE.md` and `wiki/7-Use-Cases.md` must remain byte-identical.
### Acceptance criteria
1. `errorPattern` binds per profile. Blank uses the legacy default and is reported as degraded
coverage. A config reload which changes it is reported as deferred because patterns are compiled
at startup.
2. A malformed `errorPattern` stops startup and names `profiles.<name>.errorPattern`.
3. The real production sink never silently drops an unknown target. A test captures its error log
and the fallback notice call.
4. One real backend-error classification through `CompletionResolver` marks only its member. It does
not cool the credential and does not send an outage notice.
5. Two real classifications for one credential within 60 seconds start one incident.
6. The integration test then calls the real explicit-profile spawn gate. It is refused before any
adapter spawn call, with “cooling off” and remaining seconds in the message.
7. The same test calls an automatic placement path. A cooling candidate is skipped. If every
candidate is cooling, the error names cool-off rather than exhaustion.
8. Profiles sharing the credential are all blocked. A profile on another credential stays usable.
9. `fleet_list` from the same fixture shows both members as `backend_error`, preserves each failure
reason, and reports `free: 0`, the credential, and `coolingOffForSeconds`.
10. `fleet_profiles` reports cool-off separately from quarantine.
11. The real `ReplyPushLoop` receives one incident and makes one successful lead-pane send. Existing
failed-ticket notice content may share that same combined send.
12. Exhaustion still wins when a line matches both patterns. A simultaneous exhaustion quarantine
also wins in spawn errors and fleet views.
13. After the 60-second cool-off, spawn is allowed again. A new incident needs two fresh errors.
14. `healthCoverage` has the same value before and after this change for the same health config.
15. `fleetd.example.yaml` explains `errorPattern`, the legacy fallback, 2/60/60 policy, and the
difference between cool-off and exhaustion quarantine.
16. `CLAUDE.md` tells leads how `fleet_profiles` and `fleet_list` report cool-off. The lead later
applies the matching wiki updates and runs the documented byte-sync check.
17. The developer records the new end-to-end test failing before implementation, then passing. The
focused suites and `mvn clean install` pass.
### Dependencies
Unit 5 starts only after Units 1 to 4 are merged or rebased into its branch. It also starts after the
#234 defect 2 fix lands, because both areas touch the same target-resolution control path.
### What to report back
- The exact commits used for Units 1 to 4 and #234.
- The startup coverage line with one configured and one legacy-default profile.
- The red test output before the implementation and its green result after.
- The explicit and automatic spawn refusal text.
- Sample `fleet_list` and `fleet_profiles` JSON for cool-off and exhaustion.
- The number and text of lead-pane sends in the real-path test.
- The focused test commands and final `mvn clean install` result.
- The exact `CLAUDE.md` change and the wiki edits the lead must apply.
## File ownership summary
| Area | Unit | Shared edit risk |
|---|---:|---|
| Completion classification | 1 | Only Unit 1 edits `CompletionResolver` and its test. |
| Correlation and cool-off state | 2 | New files only. |
| Lead push scheduling | 3 | Only Unit 3 edits `ReplyPushLoop` and its test. |
| Member terminal state | 4 | Only Unit 4 edits `MemberSession`, `SessionManager`, and their test. |
| Main composition, config, placement, MCP views, shipped prompt | 5 | Only Unit 5 edits `Fleetd`, `FleetConfig`, `CompositePeerLauncher`, placement context, `FleetMcp`, and `CLAUDE.md`. |
| Wiki propagation | Lead after Unit 5 | Workers do not commit the wiki submodule. |
## What I would not build
1. **Do not reuse `BackendQuarantine` for outages.** Its repeat call restarts a long credential
quarantine. Its fields and errors say “exhausted”. That is wrong for a short outage.
2. **Do not merge the exhaustion and generic error patterns.** Exhaustion must win because it has a
different policy and duration.
3. **Do not group by error string.** One outage can produce different text. The shared operational
limit is the credential.
4. **Do not mark a profile unusable until config changes.** The current classifier cannot safely
tell a permanent malformed request from a transient service fault. A permanent state would need
a stronger error taxonomy first.
5. **Do not quarantine on the first generic backend error.** That would turn one bad request or one
false pattern match into a fleet-wide capacity loss.
6. **Do not add this to `FleetHealthMonitor`.** The monitor samples slow member health. The exact
backend event already exists at completion resolution, and moving it to polling would lose type
and time.
7. **Do not add another direct lead injector.** `ReplyPushLoop` already owns status gating,
per-lead coalescing, retry bounds, and heartbeat stand-down.
8. **Do not change `healthCoverage` to `full`.** That field still means a webhook notification sink
exists for periodic health. A backend outage nudge does not make every health event visible.
9. **Do not persist incident history across daemon restart in this work.** Existing exhaustion
quarantine is also in memory. A 60-second state does not justify a new durable store.
10. **Do not build work recovery.** The PR-body survival story proves why checkpoint-first work is
useful, but these tickets are about detection, capacity, and signalling.
11. **Do not remove the legacy `API Error:` fallback in the first release.** Doing so would turn an
unedited config back into a false successful completion. Report it as degraded coverage instead.
12. **Do not reorder or add the old `visibleTurn` fallback from #201.** #211 already implemented the
narrow raw-scrape fallback at `CompletionResolver.classifyRawScrapeFallback`.
## Riskiest assumption and cheapest experiment
The riskiest assumption is that a configured error regex means “the backend failed this turn”. The
current code and test already show the counterexample: a worker may quote `API Error:` while writing
a valid report. Two such false matches on one credential would now remove capacity for 60 seconds.
The cheapest experiment is a replay corpus before Unit 5 ships:
1. Save the full pane text from the measured 2026-09-01 outage.
2. Produce one safe failure per backend with a disposable invalid endpoint or request.
3. Save one valid member report which quotes each error line.
4. Replay all samples through the real `CompletionResolver` test fixture.
5. Require outage samples to match and quoted-report samples not to match after assistant-block
extraction and baseline checks.
This costs no outage deployment and no real sleep. If quoted reports still match, narrow the profile
patterns before enabling correlation. Do not raise the threshold to hide a bad classifier.
## Sequencing with three developers
First wave:
1. Developer A: Unit 1, typed classification.
2. Developer B: Unit 2, credential outage policy.
3. Developer C: Unit 3, lead outage nudge.
As soon as one slot is free, start Unit 4. It is file-disjoint from every first-wave unit. Merge and
review Units 1 to 4 independently.
Start Unit 5 only after all four foundations and #234 are available. Unit 5 is the only high-conflict
integration branch, so no other active unit should touch its file list.
## Checks performed for this refinement
- Read issue #201 and issue #227 through their Gitea pages. Both showed zero comments.
- Read the source and tests named in the evidence section.
- Ran `git status --short --branch`; the branch was clean before this document was added.
- Ran `git log --oneline -12` to identify the branch base.
- I did not run Maven because this change adds only a design document.
- Rendered both Mermaid blocks with `npx @mermaid-js/mermaid-cli`; both commands succeeded.
@@ -3,11 +3,13 @@ package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.StatusRefiner;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.MemberRole;
@@ -151,6 +153,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
/**
* fleetd #176 fix 2: the same UNKNOWN-refinement {@code dev.ltms.fleet.inject.StatusPoller}
* uses, reused here for the spawn-readiness gate. Constructed once from {@link #agents} — see
* {@link #refinedInjectable(String, Agent)} for the corroboration that keeps it from firing on
* a dead pane's bare shell prompt.
*/
private final StatusRefiner statusRefiner;
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
@@ -272,6 +282,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
this.memberCredentials = memberCredentials;
this.hostEnvNames = hostEnvNames != null ? hostEnvNames : () -> System.getenv().keySet();
this.config = config;
this.statusRefiner = new StatusRefiner(agents);
}
// --- adapter seams -------------------------------------------------------------------------
@@ -909,16 +920,35 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* Poll {@link AgentControl#get} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*
* <p>fleetd #176 fix 1: {@code agents.get} was previously unguarded here, so a herdr
* {@code *_not_found} answer — which is what happens when the backend process EXITED rather
* than being slow — propagated as a raw {@link HerdrException} instead of the
* {@link PeerUnreachableException} every other failure path of this gate produces, and skipped
* teardown ({@link #stop}) entirely, leaking the pane/tab. {@link #failFastOnGoneBackend} closes
* that gap: it stops waiting immediately (never burns the rest of the timeout), runs the same
* teardown the timeout path below runs, and throws with a message that says the backend exited
* rather than that the pane was slow. Any other {@link HerdrException} still propagates
* unchanged — this gate does not know how to recover from it.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
Object lastStatus = null;
long start = nowMillis.getAsLong();
long deadline = start + spawnReadyTimeoutMs;
AgentStatus lastStatus = null;
while (nowMillis.getAsLong() < deadline) {
var status = agents.status(paneId);
lastStatus = status;
if (status.injectable()) {
Agent sample;
try {
sample = agents.get(paneId);
} catch (HerdrException e) {
if (isAlreadyGone(e)) {
failFastOnGoneBackend(paneId, e, nowMillis.getAsLong() - start);
}
throw e; // any other herdr failure is not ours to interpret — let it propagate
}
lastStatus = sample.status();
if (lastStatus.injectable() || refinedInjectable(paneId, sample)) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
@@ -937,6 +967,66 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
+ spawnReadyTimeoutMs + "ms");
}
/**
* fleetd #176 fix 1: the backend process exited while the gate was still waiting — herdr
* answered {@code *_not_found} instead of ever reporting an injectable status. Fails
* immediately (never burns the rest of {@link #spawnReadyTimeoutMs}), runs the exact same
* teardown {@link #waitUntilInjectableOrThrow}'s timeout path runs, and throws a
* {@link PeerUnreachableException} whose message says the process exited rather than that the
* pane was slow.
*
* @throws PeerUnreachableException always — this method never returns normally
*/
private void failFastOnGoneBackend(String paneId, HerdrException cause, long elapsedMs) {
String tail = readPaneQuietly(paneId); // read before stop() closes the pane
log.warn("peer pane={} backend process exited after {}ms while waiting for injectable "
+ "state (herdr: {}) — closing. Pane tail:\n{}",
paneId, elapsedMs, cause.getMessage(), tail);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " backend process exited after " + elapsedMs
+ "ms while waiting to become injectable (herdr reported: " + cause.getMessage()
+ "). Pane tail:\n" + tail);
}
/**
* fleetd #176 fix 2: resolve a raw {@link AgentStatus#UNKNOWN} sample into a trustworthy
* injectable state via pane content, the same refinement {@code StatusPoller} applies during a
* peer's working life — but corroborated, because the trap this gate is exposed to that the
* poller is not: a pane whose backend has already exited settles at a plain shell prompt, and
* that prompt commonly contains the same {@code ❯} glyph {@link StatusRefiner#classify} treats
* as "idle at the Claude Code TUI prompt". Naively wiring the refiner in would turn "the backend
* died" into "ready to inject" — strictly worse than today's timeout.
*
* <p>Two guards, both required:
* <ul>
* <li>{@link StatusRefiner#classify} is written for the Claude Code TUI only (its own javadoc
* says so), so refinement only ever runs for the {@code claude} adapter — never for
* {@code opencode} or any future non-Claude backend, whatever its pane looks like.</li>
* <li>The refined status is accepted only when the <em>same</em> {@code agents.get} sample
* ({@code sample}, taken once by the caller) still reports a non-null
* {@link Agent#agentType()} — herdr's own "a supported backend is still detected here"
* signal, read from the very sample the raw status came from so the two can never
* disagree. A bare shell prompt left by an exited backend reports no {@code agentType}.</li>
* </ul>
*
* <p>Called only when {@code sample.status() == UNKNOWN}, so a healthy spawn — which never sees
* {@code UNKNOWN} — triggers zero extra herdr calls; only a persistently-{@code unknown} pane
* pays the one extra {@code agent.read} {@link StatusRefiner#refine} performs.
*/
private boolean refinedInjectable(String paneId, Agent sample) {
if (sample.status() != AgentStatus.UNKNOWN) {
return false;
}
if (!"claude".equals(namePrefix)) {
return false; // StatusRefiner.classify reads a Claude Code TUI prompt specifically
}
if (sample.agentType() == null) {
return false; // no corroborating liveness signal — could be a dead pane's bare shell
}
return statusRefiner.refine(paneId, AgentStatus.UNKNOWN, agents).injectable();
}
/**
* fleetd #220: the pane's recent output, clipped, for the readiness-gate timeout log — or a
* short note when it cannot be read. Best-effort by construction: this runs on a path that is
@@ -44,10 +44,13 @@ public final class FakeHerdr implements HerdrClient {
private String agentSendErrorCode = null;
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -103,6 +106,27 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Set the detected agent kind ({@code "agent"} field) that {@code agent.get} reports —
* {@code null} models herdr not (or no longer) detecting a supported backend in the pane, e.g.
* a bare shell prompt (fleetd #176 fix 2 corroboration test).
*/
public FakeHerdr agentType(String type) {
this.agentType = type;
return this;
}
/**
* Make {@code agent.get} succeed normally for its first {@code okCalls} invocations, then fail
* every call after that with {@code code} — fleetd #176 fix 1's "backend exited mid-wait"
* fixture. {@code okCalls == 0} fails from the very first call.
*/
public FakeHerdr agentGetFailsWithAfter(int okCalls, String code) {
this.agentGetOkCalls = okCalls;
this.agentGetFailCode = code;
return this;
}
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
public FakeHerdr readText(String text) {
this.readText = text;
@@ -208,10 +232,21 @@ public final class FakeHerdr implements HerdrClient {
}
yield mapper.readTree("{\"type\":\"ok\"}");
}
case "agent.get" -> mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":"claude",
case "agent.get" -> {
if (agentGetFailCode != null) {
long getCalls = calls.stream().filter(c -> c.method().equals("agent.get")).count();
if (getCalls > agentGetOkCalls) {
throw new HerdrException(
"herdr error [" + agentGetFailCode + "]: agent target not found",
agentGetFailCode, null);
}
}
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
yield mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
.formatted(agentField, agentStatus));
}
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
@@ -963,6 +963,152 @@ class ClaudeCodeLauncherTest {
assertDoesNotThrow(() -> UUID.fromString(handle.id()));
}
// --- fleetd #176 fix 1: fail fast when the backend process exits mid-wait -------------------
@Test
void spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout() {
// First agent.get sees UNKNOWN (one normal tick); the second reports the pane gone, exactly
// what herdr answers when the backend process has already exited. The timeout is generous
// (60s) so a test that wrongly falls through to the old unguarded call — and therefore
// waits out the whole window — is unambiguously distinguishable from one that fails fast.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
herdr.agentGetFailsWithAfter(1, "pane_not_found");
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pRoot_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().toLowerCase().contains("exited"),
"exception message says the backend exited, not that the pane was slow: "
+ ex.getMessage());
assertTrue(clock[0] < 60_000,
"the gate must not burn the rest of the 60s timeout: clock only reached " + clock[0]);
long getCalls = herdr.calls.stream().filter(c -> c.method().equals("agent.get")).count();
assertEquals(2, getCalls,
"exactly one normal poll then the not_found answer — no further polling after that: "
+ getCalls);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"the pane is torn down (no orphan) even on the fail-fast path");
}
@Test
void spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged() {
// Fix 1 must only special-case a "*_not_found" answer. Any other herdr failure keeps
// propagating as-is — this gate does not know how to recover from it.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
herdr.agentGetFailsWithAfter(0, "internal_error");
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
dev.ltms.fleet.herdr.HerdrException ex = assertThrows(
dev.ltms.fleet.herdr.HerdrException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertEquals("internal_error", ex.code());
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"an error this gate does not recognize is not this gate's teardown to run");
}
// --- fleetd #176 fix 2: corroborated UNKNOWN refinement --------------------------------------
@Test
void refinedIdleIsNotAcceptedWhenAgentTypeIsNull() {
// The trap fix 2 must close: a pane sitting at a bare shell prompt after its backend exited
// still contains the same "❯" glyph StatusRefiner.classify treats as "idle at the Claude
// Code TUI prompt". Without the agentType corroboration this would be misread as injectable
// and the gate would hand back a peer that never started. herdr reports no agentType for
// that bare shell.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
herdr.agentType(null); // no supported backend detected — could be a bare shell
herdr.readText("some-host:~ user$ ❯ "); // looks exactly like an idle Claude Code prompt
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
500, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 500,
"a null agentType must not let the '❯' prompt refine to injectable — the gate has "
+ "to wait out the full timeout: clock only reached " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"the pane is torn down on timeout, same as any other never-injectable spawn");
}
@Test
void refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness() {
// The positive case: a genuinely live claude pane that herdr misreports as UNKNOWN (CB-115)
// still resolves to injectable once agentType corroborates that a supported backend is
// detected in the same sample.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN at the raw level
herdr.agentType("claude"); // herdr still detects a live claude backend
herdr.readText("? for shortcuts"); // StatusRefiner.classify's idle footer marker
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "a corroborated refined-IDLE sample lets the spawn succeed");
assertTrue(clock[0] < 60_000,
"refinement must resolve well before the timeout: clock reached " + clock[0]);
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no pane close — the peer is genuinely injectable");
}
@Test
void refinementNeverFiresWhenRawStatusIsAlreadyInjectable() {
// Acceptance criterion 4: refinement must trigger only on a raw UNKNOWN sample. Seed the
// pane content with an active-generation marker that StatusRefiner.classify would read as
// WORKING (never injectable) if refine() were wrongly invoked here — proving that a raw
// IDLE status short-circuits before refine() (and its extra agent.read call) ever runs.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("idle"); // already injectable at the raw level
herdr.agentType("claude");
herdr.readText("esc to interrupt"); // would classify as WORKING if refine() ran anyway
long[] clock = {0};
ClaudeCodeLauncher gated = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
5000, () -> clock[0], () -> clock[0] += 300);
PeerHandle handle = gated.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "an already-injectable raw status succeeds without ever refining");
assertFalse(herdr.called("agent.read"),
"refine() must never run (and so never call agent.read) when the raw status is "
+ "already injectable");
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@Test
@@ -396,6 +396,46 @@ class OpenCodeLauncherTest {
assertNotNull(ex.getMessage());
}
/**
* fleetd #176 fix 2's per-adapter guard: {@code StatusRefiner.classify} is written for the
* Claude Code TUI only, and the spawn-readiness gate must never run it against an opencode
* pane. {@code agentType("opencode")} deliberately satisfies the OTHER guard (the corroborating
* liveness check) so it cannot be what makes this test pass — only the {@code namePrefix}
* check can be. If that check were ever removed, this pane's {@code ❯} content would refine
* straight to IDLE and the gate would report ready before the backend actually was.
*/
@Test
void opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
herdr.agentType("opencode"); // non-null — satisfies the liveness guard on its own
herdr.readText("some-host:~ user$ ❯ "); // content StatusRefiner.classify reads as IDLE
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root, root);
assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000,
"an opencode pane must never refine to injectable, however its content reads — "
+ "the gate has to wait out the full timeout: clock only reached " + clock[0]);
// A blanket "agent.read is never called" does not hold here: the timeout path itself reads
// the pane tail (source=recent) for its own log message, on every timeout, regardless of
// adapter — see HerdrPeerLauncher.readPaneQuietly. So assert on the refiner's OWN probe
// source (StatusRefiner.PROBE_SOURCE = "detection") instead — that call happens only inside
// StatusRefiner.refine, so its absence proves the refiner itself was never reached for this
// opencode pane, not merely that its answer was discarded.
long detectionReads = herdr.calls.stream()
.filter(c -> c.method().equals("agent.read"))
.filter(c -> "detection".equals(((Map<?, ?>) c.params()).get("source")))
.count();
assertEquals(0, detectionReads,
"the refiner's own pane probe (source=detection) must never run against an "
+ "opencode pane — the namePrefix guard has to stop it before that call");
}
@Test
void spawnReturnsHandleWhenGateDisabled(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();