Compare commits

..

5 Commits

Author SHA1 Message Date
Dai Ha 776743cbe2 fleetd#201 Unit 1: typed backend-error classification in CompletionResolver
CI / build (pull_request) Successful in 1m19s
CI / contract (pull_request) Successful in 1m45s
Replace the hardcoded API-Error check with a target-keyed BackendErrorPatternLookup
plus a BackendErrorSink, mirroring the existing ExhaustedPatternLookup/ExhaustionSink
pair. Classifies in all three paths (normal block, #211 raw-scrape fallback, and the
fleetd#164 MIN_TURN_NANOS floor). The sink fires only after Rendezvous.resolveFailure
wins for the exact waiter. A target with no configured pattern still falls back to the
narrow (?i)\bAPI Error\s*: compatibility pattern. Existing constructors keep compiling
via BackendErrorPatternLookup.legacy() / BackendErrorSink.none() defaults.

Public send result is unchanged (still a failed send) — the typed sink event is the
internal seam Unit 5 will consume.
2026-09-03 10:33:59 +07:00
ltms 0f08b93659 fleetd #176: spawn gate fails fast on a dead backend, and refines UNKNOWN safely
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m47s
Two fixes to HerdrPeerLauncher.waitUntilInjectableOrThrow.

Fix 1: the status call is now guarded. A herdr *_not_found answer — what happens
when the backend process exited rather than being slow — used to escape as a raw
HerdrException, skipping stop() and leaking the pane and tab, with a message that
blamed a slow pane. It now fails immediately, runs the same teardown, and says the
process exited.

Fix 2: the gate can now resolve UNKNOWN with the same StatusRefiner the status
poller uses, behind two guards. It only runs for the claude adapter, because
StatusRefiner.classify reads a Claude Code TUI. And a refined result is accepted
only when the same agents.get sample still reports a non-null agentType. That
second guard closes a trap: a dead pane sits at a shell prompt containing the same
❯ glyph the classifier reads as idle, so refining without corroboration would turn
"the backend died" into "ready to inject".

Verified by the lead before merge: NAME_PREFIX really is "claude"/"opencode" so the
adapter split holds; AgentControl.status(t) was already get(t).status(), so moving
the loop to get() adds no herdr call; and a *_not_found already failed the spawn
before this change, so fix 1 improves an existing failure rather than creating one.
The adapter guard was mutation-tested independently (set it to `if (false)`, the new
opencode test fails with "Expected PeerUnreachableException to be thrown, but nothing
was thrown"; restored, it passes). Independent build: 1123 tests, 0 failures,
0 errors, 0 skipped.

Not fixed, and not claimed: that herdr reports a null agentType for a bare shell
pane is unverified against a live daemon. To be proved by a live spawn on both
backends after deploy. The seat-accounting suggestion in #176 was deliberately not
built — its cause was tested in that issue and not reproduced.
2026-09-03 04:20:23 +02:00
Dai Ha 432c1d92d1 fleetd #176 review: cover the namePrefix guard, the adapter-refinement corroboration that had no test
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m14s
The agentType corroboration (guard b) had positive and negative tests, but the
per-adapter namePrefix guard (guard a) had none — nothing proved that an
opencode pane can never reach StatusRefiner.classify, only that a claude pane
with a null agentType is rejected. Since a live opencode pane always reports a
non-null agentType ("opencode"), guard (b) alone cannot catch a broken guard
(a).

Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt
to OpenCodeLauncherTest: agentType("opencode") (non-null, satisfies guard b on
its own) + pane content containing "❯" (would classify as IDLE) + raw status
UNKNOWN throughout. Asserts the gate still times out, and that no agent.read
call used source=detection (StatusRefiner.PROBE_SOURCE) — proving the refiner
was never even reached, not just that its answer was discarded. A blanket
"agent.read is never called" does not hold here: HerdrPeerLauncher.readPaneQuietly
reads the pane tail (source=recent) for the timeout log on every timeout,
regardless of adapter, so the assertion is scoped to the refiner's own probe
source instead.

Mutation check performed and reverted before this commit: temporarily changed
`if (!"claude".equals(namePrefix))` to `if (false)` in
HerdrPeerLauncher.refinedInjectable — the new test failed
("Expected PeerUnreachableException to be thrown, but nothing was thrown"),
confirming it actually exercises the guard. Restored the guard and reran —
test passes again (1/1). No production code changed in this commit.

mvn clean install: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:17:10 +07:00
Dai Ha 60b7e67b42 fleetd #176: fail fast when the spawn gate's backend exits mid-wait, and safely refine its UNKNOWN status
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m26s
Fix 1: agents.get(paneId) inside waitUntilInjectableOrThrow was unguarded, so a
herdr *_not_found answer (the backend process exited) propagated as a raw
HerdrException instead of PeerUnreachableException, and skipped teardown
entirely, leaking the pane/tab. Now caught via isAlreadyGone(); fails
immediately (does not burn the rest of the timeout), runs the same teardown
the timeout path runs, and the exception message says the process exited
rather than that the pane was slow. Any other HerdrException still
propagates unchanged.

Fix 2: the gate now resolves a raw UNKNOWN into StatusRefiner's pane-content
classification, like StatusPoller already does mid-life. Guarded against the
trap noted in the ticket: a pane whose backend exited settles at a bare shell
prompt that can also contain the "❯" glyph classify() reads as idle. So a
refined result is accepted only when the corroborating agentType from the
SAME agent.get sample is non-null, and only for the "claude" adapter (namePrefix)
since StatusRefiner.classify is written for the Claude Code TUI only. Refine
only runs when the raw status is UNKNOWN, so a healthy spawn adds zero extra
herdr calls.

Tests added to ClaudeCodeLauncherTest (FakeHerdr gained agentType()/
agentGetFailsWithAfter() fixtures):
 - spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout
 - spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged
 - refinedIdleIsNotAcceptedWhenAgentTypeIsNull
 - refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness
 - refinementNeverFiresWhenRawStatusIsAlreadyInjectable

mvn clean install: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:10:49 +07:00
ltms 3fbd43fe3f fleetd #175: read back the model opencode actually resolved, and quarantine a silent substitution
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m29s
Verified by the lead before merge. Read the full production diff, checkModelMatch, parseModel and actualModelForDirectory. Confirmed the check runs on the real #209 late-resolve path (not at spawn, which is why #203 was closed), that unknown/incomplete evidence never quarantines, and that claude-code is structurally excluded because SessionAwareHandle is only built by OpenCodeLauncher.spawn(). Round 2 closed the one gap I found: model JSON with an id but no providerID used to read as a mismatch for a provider-prefixed profile. Independent build: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped.
2026-09-02 13:01:33 +02:00
8 changed files with 700 additions and 29 deletions
@@ -0,0 +1,35 @@
package dev.ltms.fleet.inject;
import java.util.regex.Pattern;
/**
* Per-target lookup for a profile's configured backend-error pattern (fleetd#201 / #227): how
* {@link CompletionResolver} recognizes a backend that failed a turn outright (a credential
* outage, a provider 5xx) from whatever ends up on the member's pane, apart from a genuine
* completion.
*
* <p>The pattern is always profile config, never a vendor string in Java source: every backend
* words its failure differently, so a hardcoded sentence would only ever match one of them. A
* target with no configured pattern is not "off" the way {@link ExhaustedPatternLookup#none()}
* is — {@link CompletionResolver} falls back to its narrow {@code (?i)\bAPI Error\s*:}
* compatibility pattern instead, so classification still happens, just without a profile-specific
* match. {@link #legacy()} is the explicit stand-in existing (pre-fleetd#201) callers pass to get
* exactly that fallback-only behavior until a caller wires a real, profile-driven lookup
* (fleetd#201 Unit 5).
*/
@FunctionalInterface
public interface BackendErrorPatternLookup {
/** The compiled pattern configured for {@code target}'s profile, or {@code null} if none. */
Pattern patternFor(String target);
/**
* Legacy lookup — no profile has a configured pattern, so {@link CompletionResolver} classifies
* every target using only its built-in {@code (?i)\bAPI Error\s*:} compatibility pattern. The
* explicit stand-in existing constructors pass so their behavior is unchanged until a caller
* wires a real, profile-driven lookup.
*/
static BackendErrorPatternLookup legacy() {
return target -> null;
}
}
@@ -0,0 +1,35 @@
package dev.ltms.fleet.inject;
/**
* Notified when {@link CompletionResolver} actually delivers a typed backend-error classification
* to a waiting send (fleetd#201 / #227) — never on a race that lost. {@link CompletionResolver}
* calls this only after {@code Rendezvous.resolveFailure} returns {@code true} for that exact
* waiter, mirroring the win-only race rule {@link ExhaustionSink} already uses.
*
* <p>The public send result is unchanged by this classification — it is still a failed send
* ({@code Rendezvous.Kind#FAILED}); this sink is the internal seam a later stage (fleetd#201 Unit
* 5, the policy that cools a credential off after two such failures in 60 seconds and tells the
* lead) consumes. {@link CompletionResolver} knows only {@code target} (a herdr terminal id); it
* has no notion of profiles or credentials, so mapping {@code target} to whatever should be
* quarantined is entirely the sink's job — exactly like {@link ExhaustionSink}.
*/
@FunctionalInterface
public interface BackendErrorSink {
/**
* @param target the herdr terminal id whose turn was classified as a backend error
* @param matchedLine the single pane line that matched the backend-error pattern
* @param reason the full failure reason carried by the classification (the matched line
* plus whatever pane context the classifying path attaches)
*/
void onBackendError(String target, String matchedLine, String reason);
/**
* Inert sink — nothing happens on a backend-error classification. The explicit stand-in a
* caller (or a test not exercising this feature) passes instead of a defaulting overload,
* exactly like {@link ExhaustionSink#none()}.
*/
static BackendErrorSink none() {
return (target, matchedLine, reason) -> { };
}
}
@@ -89,6 +89,8 @@ public final class CompletionResolver implements TurnListener {
private final Rendezvous rendezvous;
private final ExhaustedPatternLookup exhaustedPatterns;
private final ExhaustionSink exhaustionSink;
private final BackendErrorPatternLookup backendErrorPatterns;
private final BackendErrorSink backendErrorSink;
private final LongSupplier nowNanos;
/**
@@ -129,10 +131,48 @@ public final class CompletionResolver implements TurnListener {
* classification actually resolves a waiter. Required for the same
* reason as {@code exhaustedPatterns} — pass {@link ExhaustionSink#none()}
* to opt out.
*
* <p>Transition constructor (fleetd#201 Unit 1): delegates to the full constructor below with
* {@link BackendErrorPatternLookup#legacy()} and {@link BackendErrorSink#none()}, so every
* existing caller keeps today's behavior — the narrow built-in {@code API Error:} match, no
* sink notified — until a caller wires a real, profile-driven backend-error lookup and sink
* (fleetd#201 Unit 5).
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink) {
this(agents, rendezvous, exhaustedPatterns, exhaustionSink, System::nanoTime);
this(agents, rendezvous, exhaustedPatterns, exhaustionSink,
BackendErrorPatternLookup.legacy(), BackendErrorSink.none(), System::nanoTime);
}
/**
* Transition constructor (fleetd#201 Unit 1): same legacy backend-error defaults as the 4-arg
* constructor above, but with the injectable clock. Kept so existing fleetd#164 timing tests
* (and any other caller that wants a controllable clock but not the backend-error feature)
* keep compiling unchanged.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink, LongSupplier nowNanos) {
this(agents, rendezvous, exhaustedPatterns, exhaustionSink,
BackendErrorPatternLookup.legacy(), BackendErrorSink.none(), nowNanos);
}
/**
* Production constructor (fleetd#201 Unit 1): adds the target-keyed backend-error pattern
* lookup and sink alongside the existing exhaustion pair. Required, like {@code exhaustedPatterns}
* and {@code exhaustionSink} — pass {@link BackendErrorPatternLookup#legacy()} and
* {@link BackendErrorSink#none()} to opt out.
*
* @param backendErrorPatterns fleetd#201: per-target lookup for a profile's configured
* backend-error pattern; a target with none configured is matched
* against the built-in compatibility pattern instead (never "off").
* @param backendErrorSink fleetd#201: notified when a backend-error classification actually
* resolves a waiter — never on a race that lost.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink, BackendErrorPatternLookup backendErrorPatterns,
BackendErrorSink backendErrorSink) {
this(agents, rendezvous, exhaustedPatterns, exhaustionSink, backendErrorPatterns, backendErrorSink,
System::nanoTime);
}
/**
@@ -145,11 +185,14 @@ public final class CompletionResolver implements TurnListener {
* package.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink, LongSupplier nowNanos) {
ExhaustionSink exhaustionSink, BackendErrorPatternLookup backendErrorPatterns,
BackendErrorSink backendErrorSink, LongSupplier nowNanos) {
this.agents = agents;
this.rendezvous = rendezvous;
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
this.backendErrorPatterns = Objects.requireNonNull(backendErrorPatterns, "backendErrorPatterns");
this.backendErrorSink = Objects.requireNonNull(backendErrorSink, "backendErrorSink");
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
}
@@ -232,7 +275,7 @@ public final class CompletionResolver implements TurnListener {
// screen, since that is usually the backend's own error.
long elapsedNanos = nowNanos.getAsLong() - turn.deliveredAtNanos();
if (elapsedNanos < MIN_TURN_NANOS) {
fail(target, turn, tooFastReason(target, elapsedNanos));
failTooFast(target, turn, waiter, elapsedNanos);
return;
}
String tail;
@@ -302,18 +345,27 @@ public final class CompletionResolver implements TurnListener {
}
return;
}
// fleetd#164 (part 2): a scrape that read cleanly and produced content still isn't a real
// reply when that content is the backend's own rejection (e.g. an HTTP 400 before the worker
// did any work). Classify it as a failure naming the member, rather than handing the caller a
// scrape that reads like a completed answer.
String backendError = firstMatchingLine(assistantBlock, BACKEND_ERROR);
// fleetd#164 (part 2) / fleetd#201: a scrape that read cleanly and produced content still
// isn't a real reply when that content is the backend's own rejection (e.g. an HTTP 400
// before the worker did any work). Classify it as a failure naming the member, rather than
// handing the caller a scrape that reads like a completed answer, and — only on the
// resolution that actually wins the race, mirroring the exhaustion sink above — notify the
// typed backend-error sink so a later stage can act on repeated failures.
String backendError = firstMatchingLine(assistantBlock, backendErrorPatternOrFallback(target));
if (backendError != null) {
// Carry the whole scrape, not just the matched line. The pattern is a heuristic: a member
// that forgot fleet_reply while reporting *about* a backend error matches it too. Failing
// is still right — the caller must not read a scrape as an answer — but dropping the rest
// of the pane would destroy the report, which is the same defect fleetd#164 is about.
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + tail);
String reason = "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + tail;
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
// fleetd#201 Unit 1: only on the resolution that actually won the race — a late
// duplicate must never double-count one backend failure.
backendErrorSink.onBackendError(target, backendError, reason);
}
return;
}
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
@@ -353,14 +405,20 @@ public final class CompletionResolver implements TurnListener {
}
return true;
}
String backendError = firstMatchingLine(raw, BACKEND_ERROR);
String backendError = firstMatchingLine(raw, backendErrorPatternOrFallback(target));
if (backendError != null) {
// Carry the pane, not just the matched line — the same fleetd#164 rule the normal path
// above applies. Here it matters more, not less: the trimmed block was empty, so the raw
// scrape is the ONLY copy of whatever the member managed to say. Clipped to the same cap
// the normal path uses, since a raw screen has no boundary trimming to bound it.
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + clip(raw));
String reason = "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + clip(raw);
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback from the raw scrape: {}", target, reason);
// fleetd#201 Unit 1: only on the resolution that actually won the race.
backendErrorSink.onBackendError(target, backendError, reason);
}
return true;
}
return false;
@@ -402,22 +460,53 @@ public final class CompletionResolver implements TurnListener {
}
/**
* fleetd#164: the failure reason for a turn that settled inside {@link #MIN_TURN_NANOS} — names
* the member and both timings, and appends whatever the pane shows (usually the backend's own
* error) so the caller sees the cause, not just "it failed".
* fleetd#164 (floor) / fleetd#201 (classification): fail a turn that settled inside
* {@link #MIN_TURN_NANOS} — a crash signature (e.g. a backend HTTP 400 before the worker did
* anything) that a bare {@code BUSY -> DONE} transition cannot be told apart from a genuinely
* fast completion. Runs the same backend-error classification the normal and raw-scrape paths
* apply, against whatever is on screen right now: a match is a typed failure that notifies
* {@link #backendErrorSink} (only on the resolution that wins the race); a non-match stays the
* original generic too-fast failure, naming the member and both timings, with whatever the pane
* shows appended so the caller sees the cause, not just "it failed".
*/
private String tooFastReason(String target, long elapsedNanos) {
private void failTooFast(String target, InFlight turn, CompletableFuture<Rendezvous.Resolution> waiter,
long elapsedNanos) {
String scrape;
try {
scrape = clip(agents.read(target, SCRAPE_SOURCE));
scrape = agents.read(target, SCRAPE_SOURCE);
} catch (RuntimeException e) {
scrape = "";
}
String reason = String.format(
String clippedScrape = clip(scrape);
String baseReason = String.format(
"member %s went BUSY -> DONE in %dms (floor %dms) — too fast to be real work, most "
+ "likely a backend error before any work started",
target, elapsedNanos / 1_000_000, MIN_TURN_NANOS / 1_000_000);
return scrape.isBlank() ? reason : reason + ": " + scrape;
String backendError = firstMatchingLine(scrape, backendErrorPatternOrFallback(target));
if (backendError != null) {
String reason = baseReason + ": " + clippedScrape;
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
// fleetd#201 Unit 1: only on the resolution that actually won the race.
backendErrorSink.onBackendError(target, backendError, reason);
}
return;
}
fail(target, turn, clippedScrape.isBlank() ? baseReason : baseReason + ": " + clippedScrape);
}
/**
* fleetd#201: the pattern to classify a backend error against for {@code target} — its
* profile's configured {@link BackendErrorPatternLookup} entry when there is one, else the
* built-in {@link #BACKEND_ERROR} compatibility pattern. A target with no configured pattern is
* never "off": it always falls back to this narrow default, exactly like the classification
* behaved before fleetd#201 (Unit 5 reports a target relying on this fallback separately, as
* legacy-default coverage rather than full per-profile coverage).
*/
private Pattern backendErrorPatternOrFallback(String target) {
Pattern configured = backendErrorPatterns.patternFor(target);
return configured != null ? configured : BACKEND_ERROR;
}
/**
@@ -3,11 +3,13 @@ package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.StatusRefiner;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.MemberRole;
@@ -151,6 +153,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
/**
* fleetd #176 fix 2: the same UNKNOWN-refinement {@code dev.ltms.fleet.inject.StatusPoller}
* uses, reused here for the spawn-readiness gate. Constructed once from {@link #agents} — see
* {@link #refinedInjectable(String, Agent)} for the corroboration that keeps it from firing on
* a dead pane's bare shell prompt.
*/
private final StatusRefiner statusRefiner;
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
@@ -272,6 +282,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
this.memberCredentials = memberCredentials;
this.hostEnvNames = hostEnvNames != null ? hostEnvNames : () -> System.getenv().keySet();
this.config = config;
this.statusRefiner = new StatusRefiner(agents);
}
// --- adapter seams -------------------------------------------------------------------------
@@ -909,16 +920,35 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* Poll {@link AgentControl#get} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*
* <p>fleetd #176 fix 1: {@code agents.get} was previously unguarded here, so a herdr
* {@code *_not_found} answer — which is what happens when the backend process EXITED rather
* than being slow — propagated as a raw {@link HerdrException} instead of the
* {@link PeerUnreachableException} every other failure path of this gate produces, and skipped
* teardown ({@link #stop}) entirely, leaking the pane/tab. {@link #failFastOnGoneBackend} closes
* that gap: it stops waiting immediately (never burns the rest of the timeout), runs the same
* teardown the timeout path below runs, and throws with a message that says the backend exited
* rather than that the pane was slow. Any other {@link HerdrException} still propagates
* unchanged — this gate does not know how to recover from it.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
Object lastStatus = null;
long start = nowMillis.getAsLong();
long deadline = start + spawnReadyTimeoutMs;
AgentStatus lastStatus = null;
while (nowMillis.getAsLong() < deadline) {
var status = agents.status(paneId);
lastStatus = status;
if (status.injectable()) {
Agent sample;
try {
sample = agents.get(paneId);
} catch (HerdrException e) {
if (isAlreadyGone(e)) {
failFastOnGoneBackend(paneId, e, nowMillis.getAsLong() - start);
}
throw e; // any other herdr failure is not ours to interpret — let it propagate
}
lastStatus = sample.status();
if (lastStatus.injectable() || refinedInjectable(paneId, sample)) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
@@ -937,6 +967,66 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
+ spawnReadyTimeoutMs + "ms");
}
/**
* fleetd #176 fix 1: the backend process exited while the gate was still waiting — herdr
* answered {@code *_not_found} instead of ever reporting an injectable status. Fails
* immediately (never burns the rest of {@link #spawnReadyTimeoutMs}), runs the exact same
* teardown {@link #waitUntilInjectableOrThrow}'s timeout path runs, and throws a
* {@link PeerUnreachableException} whose message says the process exited rather than that the
* pane was slow.
*
* @throws PeerUnreachableException always — this method never returns normally
*/
private void failFastOnGoneBackend(String paneId, HerdrException cause, long elapsedMs) {
String tail = readPaneQuietly(paneId); // read before stop() closes the pane
log.warn("peer pane={} backend process exited after {}ms while waiting for injectable "
+ "state (herdr: {}) — closing. Pane tail:\n{}",
paneId, elapsedMs, cause.getMessage(), tail);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " backend process exited after " + elapsedMs
+ "ms while waiting to become injectable (herdr reported: " + cause.getMessage()
+ "). Pane tail:\n" + tail);
}
/**
* fleetd #176 fix 2: resolve a raw {@link AgentStatus#UNKNOWN} sample into a trustworthy
* injectable state via pane content, the same refinement {@code StatusPoller} applies during a
* peer's working life — but corroborated, because the trap this gate is exposed to that the
* poller is not: a pane whose backend has already exited settles at a plain shell prompt, and
* that prompt commonly contains the same {@code ❯} glyph {@link StatusRefiner#classify} treats
* as "idle at the Claude Code TUI prompt". Naively wiring the refiner in would turn "the backend
* died" into "ready to inject" — strictly worse than today's timeout.
*
* <p>Two guards, both required:
* <ul>
* <li>{@link StatusRefiner#classify} is written for the Claude Code TUI only (its own javadoc
* says so), so refinement only ever runs for the {@code claude} adapter — never for
* {@code opencode} or any future non-Claude backend, whatever its pane looks like.</li>
* <li>The refined status is accepted only when the <em>same</em> {@code agents.get} sample
* ({@code sample}, taken once by the caller) still reports a non-null
* {@link Agent#agentType()} — herdr's own "a supported backend is still detected here"
* signal, read from the very sample the raw status came from so the two can never
* disagree. A bare shell prompt left by an exited backend reports no {@code agentType}.</li>
* </ul>
*
* <p>Called only when {@code sample.status() == UNKNOWN}, so a healthy spawn — which never sees
* {@code UNKNOWN} — triggers zero extra herdr calls; only a persistently-{@code unknown} pane
* pays the one extra {@code agent.read} {@link StatusRefiner#refine} performs.
*/
private boolean refinedInjectable(String paneId, Agent sample) {
if (sample.status() != AgentStatus.UNKNOWN) {
return false;
}
if (!"claude".equals(namePrefix)) {
return false; // StatusRefiner.classify reads a Claude Code TUI prompt specifically
}
if (sample.agentType() == null) {
return false; // no corroborating liveness signal — could be a dead pane's bare shell
}
return statusRefiner.refine(paneId, AgentStatus.UNKNOWN, agents).injectable();
}
/**
* fleetd #220: the pane's recent output, clipped, for the readiness-gate timeout log — or a
* short note when it cannot be read. Best-effort by construction: this runs on a path that is
@@ -44,10 +44,13 @@ public final class FakeHerdr implements HerdrClient {
private String agentSendErrorCode = null;
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -103,6 +106,27 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Set the detected agent kind ({@code "agent"} field) that {@code agent.get} reports —
* {@code null} models herdr not (or no longer) detecting a supported backend in the pane, e.g.
* a bare shell prompt (fleetd #176 fix 2 corroboration test).
*/
public FakeHerdr agentType(String type) {
this.agentType = type;
return this;
}
/**
* Make {@code agent.get} succeed normally for its first {@code okCalls} invocations, then fail
* every call after that with {@code code} — fleetd #176 fix 1's "backend exited mid-wait"
* fixture. {@code okCalls == 0} fails from the very first call.
*/
public FakeHerdr agentGetFailsWithAfter(int okCalls, String code) {
this.agentGetOkCalls = okCalls;
this.agentGetFailCode = code;
return this;
}
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
public FakeHerdr readText(String text) {
this.readText = text;
@@ -208,10 +232,21 @@ public final class FakeHerdr implements HerdrClient {
}
yield mapper.readTree("{\"type\":\"ok\"}");
}
case "agent.get" -> mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":"claude",
case "agent.get" -> {
if (agentGetFailCode != null) {
long getCalls = calls.stream().filter(c -> c.method().equals("agent.get")).count();
if (getCalls > agentGetOkCalls) {
throw new HerdrException(
"herdr error [" + agentGetFailCode + "]: agent target not found",
agentGetFailCode, null);
}
}
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
yield mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
.formatted(agentField, agentStatus));
}
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
@@ -4,8 +4,11 @@ import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TestTurnTokens;
import dev.ltms.fleet.msg.TurnToken;
@@ -759,4 +762,202 @@ class CompletionResolverTest {
assertEquals("partial (configured: [terra]; not configured: [gx10])",
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra")));
}
// --- fleetd#201 Unit 1: target-keyed backend-error pattern + typed sink ----------------------
@Test
void aConfiguredBackendErrorPatternClassifiesAMatchAsAFailureAndNotifiesTheSinkOnce() {
String block = "⏺ 503 Service Unavailable: upstream credential rejected\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
java.util.List<String> notified = new java.util.ArrayList<>();
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a configured-pattern match still resolves the blocked send");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals(1, notified.size(), "the sink fires exactly once for the winning classification");
assertEquals("term_a: 503 Service Unavailable: upstream credential rejected", notified.get(0));
}
@Test
void aLosingBackendErrorClassificationNeverNotifiesTheSink() {
// The waiter was already resolved (e.g. by the worker's own fleet_reply) before this scrape
// landed — resolveFailure loses the race and must return false, so the sink must not fire.
String block = "⏺ 503 Service Unavailable: upstream credential rejected\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
java.util.List<String> notified = new java.util.ArrayList<>();
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
assertTrue(rendezvous.resolveCompletion(waiter, "already replied")); // fleet_reply won first
resolver.resolve("term_a", turn);
assertTrue(notified.isEmpty(), "a classification that loses the race must not fire the sink");
assertEquals("already replied", waiter.getNow(null).text(), "the earlier resolution stands untouched");
}
@Test
void aBackendErrorClassificationThatLosesARealRaceDuringTheScrapeNeverNotifiesTheSink() {
// The test above pre-resolves the waiter BEFORE calling resolve(), so it is caught by
// resolve()'s own top-of-method isDone() guard before ever reaching the win-gated
// resolveFailure call this unit adds — it proves the outcome, not the specific gate. This
// test forces the actual race window fleetd#201 must be safe in: a genuine fleet_reply lands
// WHILE the herdr scrape for this turn is in flight, i.e. AFTER resolve() has already passed
// the top guard (waiter not done yet) and committed to reading the pane, but BEFORE it
// reaches the "if (rendezvous.resolveFailure(...))" line below. The side effect is attached
// to the one real I/O step resolve() performs between those two points: the herdr
// "agent.read" call itself.
Rendezvous rendezvous = new Rendezvous();
var waiter = rendezvous.open("term_a");
ObjectMapper mapper = new ObjectMapper();
HerdrClient racingDuringScrape = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
if ("agent.read".equals(method)) {
// The real fleet_reply that wins the race, landing mid-scrape.
assertTrue(rendezvous.resolve("term_a", "a genuine fleet_reply landed first"),
"the racing reply must land while the waiter is still open");
}
try {
return mapper.readTree(mapper.writeValueAsString(java.util.Map.of(
"type", "agent_read",
"read", java.util.Map.of("text",
"⏺ 503 Service Unavailable: upstream credential rejected\n❯ "))));
} catch (Exception e) {
throw new RuntimeException(e);
}
}
@Override
public void close() {
}
};
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
java.util.List<String> notified = new java.util.ArrayList<>();
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
CompletionResolver resolver = new CompletionResolver(new AgentControl(racingDuringScrape), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink);
var turn = new CompletionResolver.InFlight(waiter, null); // not done yet — passes the top guard
resolver.resolve("term_a", turn);
assertTrue(notified.isEmpty(),
"a classification that loses a real race during the scrape must not fire the sink");
assertEquals(Rendezvous.Kind.REPLY, waiter.getNow(null).kind(), "the genuine fleet_reply stands");
assertEquals("a genuine fleet_reply landed first", waiter.getNow(null).text());
}
@Test
void exhaustionKeepsWinningOverBackendErrorEvenWhenBothPatternsMatchTheSameLine() {
// A line matching both a configured exhausted pattern AND a configured backend-error pattern
// must classify BACKEND_EXHAUSTED and call only the ExhaustionSink — the ordering fleetd#578
// already relies on must not change.
String block = "⏺ The usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup exhausted = target -> Pattern.compile("usage limit has been reached");
BackendErrorPatternLookup backendErrors = target -> Pattern.compile("(?i)usage limit");
java.util.List<String> exhaustedNotified = new java.util.ArrayList<>();
java.util.List<String> backendErrorNotified = new java.util.ArrayList<>();
ExhaustionSink exhaustionSink = (target, reason) -> exhaustedNotified.add(target);
BackendErrorSink backendErrorSink = (target, matchedLine, reason) -> backendErrorNotified.add(target);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
exhausted, exhaustionSink, backendErrors, backendErrorSink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"exhaustion must keep winning over backend error");
assertEquals(1, exhaustedNotified.size(), "the exhaustion sink fires");
assertTrue(backendErrorNotified.isEmpty(), "the backend-error sink must never fire for this turn");
}
@Test
void aConfiguredPatternAlsoClassifiesTheRawScrapeFallbackAndNotifiesTheSink() {
// No ⏺ marker and leading TUI chrome ⇒ lastAssistantBlock() yields "", so classification must
// fall back to the raw scrape (fleetd#211) — and it must use the configured pattern too.
String block = """
╭──────────────────────────────────────╮
503 Service Unavailable: upstream credential rejected
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
java.util.List<String> notified = new java.util.ArrayList<>();
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a raw-scrape configured-pattern match still resolves the blocked send");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"classified from the raw scrape even though the trimmed block was empty");
assertEquals(1, notified.size(), "the sink fires exactly once from the raw-scrape path");
assertTrue(notified.get(0).contains("503 Service Unavailable"), notified.get(0));
}
// --- fleetd#201 Unit 1: classification inside the fleetd#164 MIN_TURN_NANOS floor -------------
@Test
void aMatchingErrorInsideTheFloorIsTypedAndNotifiesTheSink() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ 503 Service Unavailable: upstream credential rejected\n❯ ");
Rendezvous rendezvous = new Rendezvous();
long[] clock = {10_000_000_000L};
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
java.util.List<String> notified = new java.util.ArrayList<>();
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink, () -> clock[0]);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null, clock[0]); // delivered "now"
clock[0] += CompletionResolver.MIN_TURN_NANOS - 1; // inside the floor
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone());
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals(1, notified.size(), "a matching error inside the floor notifies the sink");
assertTrue(notified.get(0).contains("503 Service Unavailable"), notified.get(0));
}
@Test
void aNonMatchInsideTheFloorStaysGenericAndNeverNotifiesTheSink() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ still starting up\n❯ ");
Rendezvous rendezvous = new Rendezvous();
long[] clock = {10_000_000_000L};
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
java.util.List<String> notified = new java.util.ArrayList<>();
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink, () -> clock[0]);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null, clock[0]); // delivered "now"
clock[0] += CompletionResolver.MIN_TURN_NANOS - 1; // inside the floor
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone());
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"still a failure — the floor itself, not the pattern, is why");
assertTrue(waiter.getNow(null).text().contains("too fast to be real work"),
"a non-match inside the floor stays the generic too-fast reason: " + waiter.getNow(null).text());
assertTrue(notified.isEmpty(), "a non-match must never notify the typed sink");
}
}
@@ -963,6 +963,152 @@ class ClaudeCodeLauncherTest {
assertDoesNotThrow(() -> UUID.fromString(handle.id()));
}
// --- fleetd #176 fix 1: fail fast when the backend process exits mid-wait -------------------
@Test
void spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout() {
// First agent.get sees UNKNOWN (one normal tick); the second reports the pane gone, exactly
// what herdr answers when the backend process has already exited. The timeout is generous
// (60s) so a test that wrongly falls through to the old unguarded call — and therefore
// waits out the whole window — is unambiguously distinguishable from one that fails fast.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
herdr.agentGetFailsWithAfter(1, "pane_not_found");
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pRoot_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().toLowerCase().contains("exited"),
"exception message says the backend exited, not that the pane was slow: "
+ ex.getMessage());
assertTrue(clock[0] < 60_000,
"the gate must not burn the rest of the 60s timeout: clock only reached " + clock[0]);
long getCalls = herdr.calls.stream().filter(c -> c.method().equals("agent.get")).count();
assertEquals(2, getCalls,
"exactly one normal poll then the not_found answer — no further polling after that: "
+ getCalls);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"the pane is torn down (no orphan) even on the fail-fast path");
}
@Test
void spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged() {
// Fix 1 must only special-case a "*_not_found" answer. Any other herdr failure keeps
// propagating as-is — this gate does not know how to recover from it.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
herdr.agentGetFailsWithAfter(0, "internal_error");
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
dev.ltms.fleet.herdr.HerdrException ex = assertThrows(
dev.ltms.fleet.herdr.HerdrException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertEquals("internal_error", ex.code());
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"an error this gate does not recognize is not this gate's teardown to run");
}
// --- fleetd #176 fix 2: corroborated UNKNOWN refinement --------------------------------------
@Test
void refinedIdleIsNotAcceptedWhenAgentTypeIsNull() {
// The trap fix 2 must close: a pane sitting at a bare shell prompt after its backend exited
// still contains the same "❯" glyph StatusRefiner.classify treats as "idle at the Claude
// Code TUI prompt". Without the agentType corroboration this would be misread as injectable
// and the gate would hand back a peer that never started. herdr reports no agentType for
// that bare shell.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
herdr.agentType(null); // no supported backend detected — could be a bare shell
herdr.readText("some-host:~ user$ ❯ "); // looks exactly like an idle Claude Code prompt
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
500, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 500,
"a null agentType must not let the '❯' prompt refine to injectable — the gate has "
+ "to wait out the full timeout: clock only reached " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"the pane is torn down on timeout, same as any other never-injectable spawn");
}
@Test
void refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness() {
// The positive case: a genuinely live claude pane that herdr misreports as UNKNOWN (CB-115)
// still resolves to injectable once agentType corroborates that a supported backend is
// detected in the same sample.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN at the raw level
herdr.agentType("claude"); // herdr still detects a live claude backend
herdr.readText("? for shortcuts"); // StatusRefiner.classify's idle footer marker
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "a corroborated refined-IDLE sample lets the spawn succeed");
assertTrue(clock[0] < 60_000,
"refinement must resolve well before the timeout: clock reached " + clock[0]);
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no pane close — the peer is genuinely injectable");
}
@Test
void refinementNeverFiresWhenRawStatusIsAlreadyInjectable() {
// Acceptance criterion 4: refinement must trigger only on a raw UNKNOWN sample. Seed the
// pane content with an active-generation marker that StatusRefiner.classify would read as
// WORKING (never injectable) if refine() were wrongly invoked here — proving that a raw
// IDLE status short-circuits before refine() (and its extra agent.read call) ever runs.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("idle"); // already injectable at the raw level
herdr.agentType("claude");
herdr.readText("esc to interrupt"); // would classify as WORKING if refine() ran anyway
long[] clock = {0};
ClaudeCodeLauncher gated = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
5000, () -> clock[0], () -> clock[0] += 300);
PeerHandle handle = gated.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "an already-injectable raw status succeeds without ever refining");
assertFalse(herdr.called("agent.read"),
"refine() must never run (and so never call agent.read) when the raw status is "
+ "already injectable");
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@Test
@@ -396,6 +396,46 @@ class OpenCodeLauncherTest {
assertNotNull(ex.getMessage());
}
/**
* fleetd #176 fix 2's per-adapter guard: {@code StatusRefiner.classify} is written for the
* Claude Code TUI only, and the spawn-readiness gate must never run it against an opencode
* pane. {@code agentType("opencode")} deliberately satisfies the OTHER guard (the corroborating
* liveness check) so it cannot be what makes this test pass — only the {@code namePrefix}
* check can be. If that check were ever removed, this pane's {@code ❯} content would refine
* straight to IDLE and the gate would report ready before the backend actually was.
*/
@Test
void opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
herdr.agentType("opencode"); // non-null — satisfies the liveness guard on its own
herdr.readText("some-host:~ user$ ❯ "); // content StatusRefiner.classify reads as IDLE
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root, root);
assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000,
"an opencode pane must never refine to injectable, however its content reads — "
+ "the gate has to wait out the full timeout: clock only reached " + clock[0]);
// A blanket "agent.read is never called" does not hold here: the timeout path itself reads
// the pane tail (source=recent) for its own log message, on every timeout, regardless of
// adapter — see HerdrPeerLauncher.readPaneQuietly. So assert on the refiner's OWN probe
// source (StatusRefiner.PROBE_SOURCE = "detection") instead — that call happens only inside
// StatusRefiner.refine, so its absence proves the refiner itself was never reached for this
// opencode pane, not merely that its answer was discarded.
long detectionReads = herdr.calls.stream()
.filter(c -> c.method().equals("agent.read"))
.filter(c -> "detection".equals(((Map<?, ?>) c.params()).get("source")))
.count();
assertEquals(0, detectionReads,
"the refiner's own pane probe (source=detection) must never run against an "
+ "opencode pane — the namePrefix guard has to stop it before that call");
}
@Test
void spawnReturnsHandleWhenGateDisabled(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();