maxLoad counted panes, never subscription seats: a subscription:true
profile's lead is itself a live claude session on that same account,
so free overstated capacity by the lead's own seat (measured free:1
with a real ceiling of 0, and free:3 on an idle fleet with a real
ceiling of 2).
Add FleetMcp.LeadSeatSource (same shape as QuarantineSource/
OutageSource) and Fleetd.leadSeatLookup, which derives the seat count
from fleet.leaders.<name>.profile matched against the target profile
by effectiveCredentialId() - no hardcoded "-1", and no new config key:
profile: already exists for this exact "which account does this lead
share" question. maxLoad itself is left untouched; only free (and a
new, additive-only leadSeats field) changes.
Exhaustion quarantine (cause 2 in the ticket) already forced free to 0
via the same BackendQuarantine capacityView already reads - confirmed
by reading the exhaustionSink wiring, no code change needed there.
Found by reading a real boot log after the redeploy, not by a test.
coverage() is shared by two call sites — CB-578's exhaustedPattern line
and Unit 5's errorPattern line — but its 'off' branch hard-coded the
word exhaustedPattern. So this daemon printed:
backend-exhausted classification (CB-578 stage A): partial
(configured: [sol, terra]; not configured: [...])
backend-error classification (fleetd #201 Unit 5): off
(no profile has an exhaustedPattern configured; profiles: [...])
Two lines, one directly under the other, disagreeing about whether any
profile has an exhaustedPattern. Both were individually defensible and
together they were nonsense. Worse, the message sends an operator to
set the wrong key: the thing that is missing is errorPattern.
coverage now takes the key name. I changed the signature rather than
adding an overload, so the compiler found all three existing callers
instead of leaving them silently on the old path.
Every earlier coverage test passed the exhaustion case only, which is
why none of them could see this. The new test pins the errorPattern
case. Reverting the fix turns it red with 0 compile errors.
1229 tests, 0 failures, BUILD SUCCESS.
This is the second defect in two hours found only by reading the live
startup log — see #115, where the noise of a false warning had been
hiding a correct line saying a whole feature was off.
Before this, dropping either #241's worktree lookup or Unit 5's
backend-error pair at Fleetd.main's new CompletionResolver(...) call
left all 1216 tests green with 0 compile errors. Every existing test
built its own CompletionResolver, so they proved the class and never
the wiring. BackendOutageFlowTest was the sharpest case: it copies
main's sink lambda line-for-line, so it proves the copy and cannot
notice the original being deleted.
The three inline arguments are now package-private static factories on
Fleetd, following the deliverableTo pattern the file already had, each
with its own behaviour test. backendErrorSink is public so a
cross-package test can drive the real production object rather than a
hand-mirrored copy.
The test that was actually missing is a source-text assertion. That is
the honest fallback for a composition root with no seam, and it is
labelled [SOURCE TEXT] in every test name and message so it cannot be
misread as a behaviour check. It is not vacuous: two tests pin that the
variables are assigned from the factories, and two pin that those
variables reach the call site, so renaming a variable while assigning
an inert value does not slip through.
Known cost, accepted: the assertions match exact source substrings, so
reformatting that statement will break them. That is the price of
covering a main method, and a spurious failure here is loud and
obvious, which is the right direction to fail.
Verified by the lead, both mutations re-run against the merged code —
see the merge check.
PR #251
Fleetd.main built three of CompletionResolver's 8 constructor arguments inline
(a worktree/branch lookup lambda, and the backend-error pattern lookup + sink
locals). Dropping any of them at the call site compiled clean and left every
existing test green, because every existing test constructs its own
CompletionResolver and only ever proves the class, never main's wiring.
Extract each into a static factory on Fleetd (worktreeBranchLookup,
backendErrorPatternLookup, backendErrorSink — the same static-factory pattern
Fleetd.deliverableTo already uses), test each factory's own behaviour, and add
a source-text assertion (FleetdCompletionResolverWiringTest) proving main's
CompletionResolver call still passes all three. backendErrorSink is public so
BackendOutageFlowTest can exercise the real production sink directly instead
of the hand-mirrored copy its own class doc used to describe.
No production behaviour changes — mechanical extraction only.
The default is now [.env], not [.env, .envrc]. .env is data, so copying
it into a worker worktree can only move values. .envrc is executable
shell that direnv runs on every cd, so copying it moves behaviour. Those
are different risks and should not share a default.
The knob is unchanged. An operator who wants .envrc copied writes
parityOverlay: ['.env', '.envrc'] and owns that choice; a new test pins
that escape hatch, because without it this would be a removal rather
than a re-default.
Decision recorded on the ticket, with the evidence it asked for first:
this checkout has no .env and no .envrc, and direnv is not on PATH, so
there was no live exposure. Point 1 (extend the credential scrub to
direnv) is declined and the reason is on the ticket — the scrub is a
one-shot .zlogin and a direnv hook runs on every cd, so no amount of
work on the scrub can cover it. Not copying the executable file is the
smaller change and removes the need.
The worker also fixed WorktreeSessionManagerTest, which hardcoded the
same default at another layer and broke the build. Outside its named
scope, correctly flagged rather than done silently.
Verified by the lead: 1216 tests, 0 failures, 0 compile errors.
PR #250
.env is data; .envrc is executable shell that direnv runs on every cd, so
copying it into a worker moves behaviour, not just values. The default
parityOverlay is now [.env] only. The knob is unchanged: an operator who
wants .envrc copied can still write parityOverlay: [.env, .envrc]
explicitly.
Updates FleetConfig's default and javadoc, fleetd.example.yaml's two
mentions of the default, and the FleetConfigTest coverage: renamed the
default test, added parityOverlayExplicitEnvrcOptInStillWorks to prove
the .envrc opt-in escape hatch still works, and fixed
WorktreeSessionManagerTest#worktreeAcquireRunsParityOverlayWithProfileDefaults
which also hardcoded the old default.
The completion fallback scrapes a member's pane when a turn ends with no
fleet_reply. If the pane still shows the brief the lead injected, the
scrape returned that brief, and the lead read its own words as the
member's answer. A silent member looked like a member that had reported.
echoesInjectedBrief now recognises that case and refuses it. Round 1 used
plain containment in both directions, which destroyed real reports: a
genuine report that quotes the brief contains it. Round 2 keeps the safe
direction unbounded (the brief contains the scrape) and bounds the other
one at MAX_ECHO_EXCESS_CHARS, so a scrape only counts as an echo when it
adds almost nothing to the brief.
Merge note — the Fleetd.java conflict:
This call site was changed by both #201/#227 Unit 5 (backendErrorPatterns
+ backendErrorSink) and by this ticket (the worktree/branch lookup). I
resolved it onto the full 8-argument constructor so neither feature is
dropped; nowNanos has to be passed explicitly to reach that overload.
Verified by the lead: 1215 tests, 0 failures, 0 compile errors.
I also measured whether the resolution itself is protected, and it is
NOT. Both mutations at this call site stay green:
- drop the worktree lookup (pass _ -> null): 1215 tests, 0 failures
- drop Unit 5's patterns/sink (legacy()/none()): 1215 tests, 0 failures
Nothing in the suite covers Fleetd's composition root, so either feature
could be silently unwired here and the build would still be clean. The
tests prove the seams, not the caller. Filed separately rather than
fixed in a merge commit.
PR #245, branch worker/cb241-fallback-echo-1175e9-11
A profile's credential that throws two distinct backend errors inside 60
seconds now cools off for 60 seconds. Automatic placement skips it,
an explicit fleet_spawn naming it is refused before the adapter is
called, and fleet_list/fleet_profiles report it as coolingOffForSeconds
next to the separate CB-578 quarantinedForSeconds.
Verified by the lead: see the merge check below. The worker ran 7
mutations, all killed with 0 compile errors; M7 was NOT killed on the
first pass (the assertion only checked .contains("quarantined"), which
is true of both the correct message and the mutated fallback), and the
worker strengthened it to assertEquals on the exact literal and kept
that change. That is the right call and it is reported honestly.
Two deviations, both justified in the PR:
- FixedPlacementPolicy needed the same coolingOff filter because it
filters candidates inline instead of using PlacementPolicyUtil.
- BackendOutageFlowTest sits in dev.ltms.fleet.inject because
CompletionResolver.InFlight is package-private there.
The startup coverage log line is a code-reading claim, not a captured
line from a live daemon. The worker said so rather than overclaiming.
PR #246, branch worker/cb201-unit5-wiring-6c12e6-8
Wires the already-merged units into production:
- Per-profile errorPattern config (beside exhaustedPattern), compiled once at
startup; falls back to the legacy (?i)\bAPI Error\s*: pattern when unset.
Startup logs configured-vs-legacy coverage, same as exhaustedPattern.
- One production BackendErrorSink in Fleetd.java: mark backend_error on the
session, resolve the profile's credential fail-loud (never
Optional.ifPresent), record it in BackendOutagePolicy, and push a lead
nudge on a new incident.
- CompositePeerLauncher's explicit and automatic spawn paths both refuse a
cooling-off credential; exhaustion quarantine wins when both are active.
PlacementContext gets a separate coolingOff set so refusal text says
"cooling off", never "exhausted".
- fleet_list/fleet_profiles report coolingOffForSeconds as an independent
fact from quarantinedForSeconds; both can appear together.
- fleetd.example.yaml documents errorPattern and the 2/60/60 cool-off policy;
CLAUDE.md tells leads how to read the two independent outage states.
Also: FixedPlacementPolicy.java, not in the original file list, needed the
same coolingOff filtering as PlacementPolicyUtil (it does its own inline
candidate filtering rather than delegating).
1190 tests, 0 failures (up from the 1163 baseline); BUILD SUCCESS.
Claude Code asks 'is this a project you trust?' the first time it starts in a directory it
has not seen. It is interactive with no timeout, and every member spawned with worktree:true
lands in a brand-new directory. The member never reaches its first turn and never replies,
while herdr reports blocked/interactive_ready — which reads as healthy.
ClaudeCodeLauncher now seeds projects.<cwd>.hasTrustDialogAccepted in the profile's
.claude.json before the process starts. This is not a new grant: the operator already
trusted the repo by configuring the profile against it, and a worktree is a checkout of it.
The write is gated on isProvisionedWorktree(cwd) — a .git that is a regular gitdir-pointer
file, never a real checkout. That gate exists because an earlier revision of this change,
run under mutation testing, wrote to the operator's real ~/.claude.json and truncated it
from 72KB to 919 bytes. Tests using a null configDir fall back to the real user.home, so
an ungated seed reaches real files.
The write is atomic (sibling temp file + ATOMIC_MOVE, never truncate-in-place) and the
whole read-modify-write is under a lock, because .claude.json is large, live, and rewritten
by Claude Code itself while fleetd runs. Two parallel spawns are normal here.
Verified by the lead: 1173 tests, 0 failures. A truncating write turns the torn-read test
red; removing the lock turns concurrentSeedsForDifferentCwdsBothSurvive red. Both with 0
compile errors. copyPosixPermissionsIfPresent is NOT covered by a test — its mutation stays
green — but createTempFile is 0600 on POSIX by default, so the not-world-readable property
holds without it; the line only preserves a non-default mode.
Round 1 used plain bidirectional containment. The direction that catches the real bug --
the pane holds the brief plus a status bar, so the scrape contains the brief -- also fires
when a member restates the whole brief and then writes a genuine report under it. That
threw the report away and told the lead nothing was produced, which is worse than the bug
being fixed: it destroys a delivery instead of merely obscuring one.
The safe direction (the scrape is a fragment of the brief) stays unbounded, because a
fragment of the brief is by definition not a report. The dangerous direction now requires
the scrape to add at most MAX_ECHO_EXCESS_CHARS beyond the brief, which is the amount of
TUI chrome a real echo carries.
Work by the cb241 worker, committed by the lead: its backend stopped answering after the
fix was written, so two turns ended with no commit and no reply. Verified by the lead:
1169 tests, 0 failures; removing the bound turns pinsTheMaximumTuiChromeExcess and
completionFallbackKeepsARealReportThatRestatesTheWholeBrief red with 0 compile errors.
Files.writeString truncates the target in place before writing, so there
was a window where .claude.json could be observed empty or half-written
- exactly the shape of the incident this ticket already hit once, but
reachable in production too: a crash/kill mid-write, or two concurrent
claude-code spawns (normal here - several run in parallel routinely)
racing a naive read-modify-write and silently discarding one spawn's
entry.
Two independent fixes, each with its own dedicated test proving it (not
the other):
- ClaudeCodeLauncher.writeAtomically: serialise to a sibling temp file in
the same directory, then Files.move with ATOMIC_MOVE + REPLACE_EXISTING,
preserving the target's existing POSIX permissions (.claude.json ships
0600). A reader now only ever observes the fully-old or fully-new file,
never a torn one. Package-visible so a test can drive it directly.
- TRUST_JSON_LOCK: a process-wide lock around seedTrustDialog's whole
read-modify-write, so two concurrent spawns for different cwds both
keep their entry instead of the second write discarding the first.
Sufficient because every spawn on this daemon runs in one JVM; it does
NOT protect against a second daemon process or the operator's own live
Claude Code writing at the same instant - writeAtomically covers that
case instead.
Both fail soft, same as before: any I/O failure here must never block a
spawn.
Four new tests: a large (30-project) existing file survives without
collapsing (asserted on the restored key set, not just that the result
parses); two concurrent spawns for different cwds both keep their entry
(CountDownLatch-synchronised, not a sleep); existing 0600 permissions
survive the write; and a direct test of writeAtomically with a busy-poll
reader thread proving a concurrent reader never observes a torn file.
See PR body for the full mutation-testing table, including an
honest note on which of these tests the atomicity mutation actually
caught (not the one implied by the numbering in review) and why.
The daemon log and the worktree git config are both new in 205ad82, but neither helps a
lead who is writing a brief. This is the line that does: never brief a worker to edit
.mcp.json, opencode.json or .autoenv in its worktree, because the edit cannot be
committed and nothing will say so.
Goes in the project addendum, not the canonical block, so the byte-sync with
wiki/7-Use-Cases.md is unaffected — re-checked and still true.
isolateToolSurface replaces .mcp.json, opencode.json and .autoenv with stubs in every
provisioned worktree and marks them --skip-worktree. That neutralisation is correct and
is unchanged here — the committed files would mount the primary's credentials.
The problem was that it was invisible. A worker told to edit opencode.json read a 3-byte
stub and reported, truthfully and wrongly, that the mount key did not exist. A missing
file would have prompted a question; a plausible stub did not.
Two changes, both visibility only. The daemon now logs one info summary per provisioning
with the denominator, the files neutralised, and the consequence. And the list is recorded
in worktree-scoped git config (fleet.neutralizedConfig / fleet.neutralizedConfigNote) so a
worker can discover it from inside its own worktree with
'git config --worktree --get-all fleet.neutralizedConfig'.
Worktree-scoped config was chosen over a file in the working tree because it lives in
.git/worktrees/<nonce>/config.worktree and so can never appear in git status, and because
configureEnvironmentCredentialHelper already uses the same mechanism in the same add() call.
Verified by the lead: baseline 1172 tests, 0 failures. Reverting the summary to log.debug
goes red (2 tests), and recording into --local rather than --worktree — which would leak
the record into the shared repo config — goes red too. Both with 0 compile errors.
isolateToolSurface replaced .mcp.json/opencode.json/.autoenv with neutral stubs and
marked them --skip-worktree, but said nothing anywhere. A real worker read a 3-byte
{} stub for opencode.json, where the repo's real file is 30+ lines, and truthfully
(but wrongly) reported a mount key did not exist.
Two readers, two fixes:
- the daemon operator gets one info log per provisioning, naming the denominator,
what was neutralized, and why anything was not (same shape as overlayParity's
fix in #148 point 3).
- the worker gets the same fact recorded in worktree-scoped git config
(fleet.neutralizedConfig / fleet.neutralizedConfigNote), discoverable with
`git config --worktree --get-all fleet.neutralizedConfig` from inside its own
worktree, without asking the lead. Not a working-tree file: this repo already
uses worktree-scoped config for the credential helper and the SSH->HTTPS
rewrite, and it lives under .git/worktrees/<nonce>/ so it can never appear in
`git status` for the worker to trip on or commit.
The neutralization itself (stub content, --skip-worktree marking) is unchanged.
A claude-code member spawned into a fresh worktree hits an interactive,
un-timed workspace-trust prompt on its first start in a directory it has
never seen. It never reaches its first turn and never mounts the bridge.
Fix: ClaudeCodeLauncher.seedTrustDialog writes
projects.<cwd>.hasTrustDialogAccepted / hasCompletedProjectOnboarding into
the profile's configDir/.claude.json (or ~/.claude.json when configDir is
unset) BEFORE the herdr spawn call, additively (existing keys/projects are
preserved). Gated to isProvisionedWorktree(cwd) - a .git that is a regular
gitdir-pointer file, never a real checkout's .git directory - the same
signal writeIdeOverlay already used, now shared between both.
That gate is a fix for a real incident hit while building this: an
earlier ungated version ran against this file's own pre-existing tests
(configDir=null, no cwd -> falls back to the real user.dir and
~/.claude.json) and corrupted the operator's actual ~/.claude.json down
to a single entry during a mutation-testing run. See PR body for the
full incident report.
FakeHerdr gained onAgentStart(Runnable) so a test can assert the seed
is on disk at the exact instant herdr's agent.start call is reached -
i.e. strictly before the peer process itself would start.
Both defects lived in one method. overlayParity logged every step at debug, so at the
default level the copy was silent and nobody could tell which overlay files a member
actually got. It also marked a copied tracked file --skip-worktree and said nothing, so a
worker editing that file later found git ignoring the change with no error anywhere.
The summary now reports the denominator, not a bare count: 'copied 1 of 2 candidates:
.env (.envrc absent)'. A bare 'copied 1' is the same under-reporting shape as #113.
Neutralised files are named with the consequence in the message itself.
No marker file is written into the worktree: acceptance criterion 1 requires the worktree
to hold exactly the configured overlay set, so a marker would violate the fix it documents.
The copy and mark logic is unchanged — only logging is new.
Verified by the lead: baseline 1168 tests, 0 failures. Reverting either log.info to
log.debug goes red (2 reds and 1 red, 0 compile errors each), which is the regression that
matters since the whole fix is the log level.
overlayParity logged everything at debug, so at the default level nobody could
tell which overlay files a spawn actually received (#148 pt 3), and a tracked
file marked --skip-worktree gave no warning that it can no longer be edited
from that worktree (#134).
Report the outcome at info: a per-spawn summary naming the denominator (every
configured candidate), what was copied, and why anything was not — plus a
separate line naming every file marked --skip-worktree, stating plainly that
it cannot be committed from this worktree. No worktree-local marker file: the
worktree must hold exactly the configured overlay set and nothing else, so an
extra file would violate that invariant.
CompletionResolver already reads the pane on every turn, so the classifier lives there
rather than in a new watcher. A matched pattern resolves the waiter as a failure and
fires BackendErrorSink, always inside the resolveFailure win-gate so exactly one thread
reports one incident.
The too-fast path now takes a fresh scrape instead of reusing the pre-turn text, so a
backend that dies immediately is still classified. BackendErrorSink is a single-method
functional interface by design — see #234 for what a default overload does to a lambda.
An exhausted opencode member could not always be mapped back to a credential, because
the launcher knew the profile and the sink did not. The sink now takes a profile hint.
The interface is inverted on purpose: the three-argument method is the single abstract
method and the two-argument one is the default. A lambda can only implement the abstract
method, so every lambda is now forced to carry the profile. The first round of this fix
added the third argument as a default overload, and the production forwarder in Fleetd
was a two-argument lambda — so the fix compiled, passed its tests, and never ran.
Verified by the lead: deleting the forwardingTo factory's override, and rewriting the
Fleetd call site as a plain lambda, both fail to compile now rather than passing silently.
Backend incidents become a fourth source inside ReplyPushLoop, not a new scheduler — a
second injector would race the one control that already owns lead-pane delivery. One
notice per incident per affected lead (not per member), one-shot, waiting while the lead
pane is not injectable, and combined into the same nudge as any pending failed ticket.
A classified target that cannot be mapped to a credential gets its own truthful notice
and its own pending/delivered records. It previously reused the incident message, which
told the lead a credential was 'cooling for 0 remaining seconds' when nothing was cooling,
and smuggled the free-text reason into the profiles field.
Verified by the lead: 1129 tests green; routing the unmapped path back through
onBackendIncident turns unmappedBackendTargetUsesTheKnownLeadSchedule red with 0 compile
errors, and that test now asserts the whole rendered message rather than two substrings.
Adds BackendOutagePolicy — a credential-keyed state machine on an injected monotonic
clock. Two classified backend errors from two DISTINCT targets on one credential inside
60 seconds mint one incident and start a 60-second cool-off. Errors during cool-off
neither extend it nor mint another; expiry clears evidence, so two fresh errors rearm.
Correlated on credentialId, never on profile name or error text. Deliberately not
BackendQuarantine: that restarts a 1800-second cooldown per exhaustion, and its name
would make every refusal say 'backend exhausted', which is a different condition.
Threshold counts distinct targets rather than raw events (lead decision): the classifier
is a heuristic and a valid member report can quote an 'API Error:' line, so one member
repeating that line must not remove a healthy credential's capacity. A real outage hits
every member on the credential, so true detection is unaffected.
Verified by the lead: 1135 tests green; reverting evidenceCount() to reasons.size()
turns two BackendOutagePolicyTest cases red with 0 compile errors.
Round 3's factory fixed the two known call sites but the underlying shape
was still there: a lambda written against ExhaustionSink binds to whichever
overload is abstract, and the 2-arg form held that position, so ANY lambda
-- a call-site forwarder, a hand-built test double, a future caller who has
never heard of fleetd #234 -- could still silently take the hint-dropping
default. Two rounds shipped exactly that mistake in two different places.
Fix: made the 3-arg onExhausted(target, reason, profile) the interface's
single abstract method; the 2-arg form is now a default that delegates with
a null profile. A lambda declared against ExhaustionSink today is forced by
the compiler to take three parameters -- there is no overload left for it to
bind to that can drop the hint. This is enforced by the type system, not by
a test that has to remember to check for it.
Knock-on changes:
- ExhaustionSink.none() -- a 3-arg lambda, still a genuine no-op, now safe
by construction rather than by care.
- ExhaustionSink.forwardingTo(...) -- collapses to a one-line 3-arg lambda;
kept as a named factory (round 3's lesson: a test must call the real
object, not rebuild its shape).
- Fleetd.java's real sink and the two OpenCodeLauncherTest sinks that used
to be anonymous classes overriding both overloads are now plain lambdas
too -- the 2-arg override each carried was pure boilerplate once the
interface provides it as a default.
- CompletionResolver.java itself: UNCHANGED, zero diff (confirmed via
`git diff --stat` before staging) -- its two call sites still call the
2-arg onExhausted(target, reason), which is now the default and behaves
identically. CompletionResolverTest (41 tests, 0 failures) proves this;
its five ExhaustionSink lambdas needed a mechanical third parameter added
to keep compiling against the new abstract method, no assertion changed.
Mutation proof, re-run against the new shape: forwardingTo's body edited to
call the 2-arg default instead of passing the hint through (the equivalent
of round 3's "delete the 3-arg override" now that there is only one method
to break) -- both new tests go red with the same assertions as round 3:
ExhaustionSinkForwardingHazardTest...: expected: <gx> but was: <null>
OpenCodeLauncherTest...ForwardingHop: expected: <true> but was: <false>
Tests run: 68, Failures: 2
Restored, re-ran: green (Tests run: 109, Failures: 0, including
CompletionResolverTest).
Compiler proof (not committed -- a scratch file outside the worktree,
compiled with the real ExhaustionSink.java on the classpath, then deleted):
ExhaustionSink forwarder = (target, reason) -> System.out.println(target + reason);
error: incompatible types: incompatible parameter types in lambda expression
A 2-arg lambda against this interface no longer compiles at all.
mvn clean install: Tests run: 1129, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
Adds MemberSession.State.BACKEND_ERROR with a nullable failureReason, surfaced in
rosterView, and SessionManager.onBackendError(target, reason). The CAS loop accepts
both sides of the completion race (BUSY and DONE); BACKEND_ERROR is terminal.
completeTurn now returns early when its CAS loses, so a stale DONE copy can no longer
release the pane or reset its context behind a member that just went BACKEND_ERROR.
Verified by the lead: 1130 tests green; mutating the completeTurn early return back to
the old fall-through turns losingCompletionDoesNotReleaseOrClearABackendErrorMember red
with 0 compile errors.
Two errors from the same target inside the window must never trip the
outage threshold on their own (a valid member report can legitimately
quote an "API Error:" line twice) — only two DIFFERENT targets on the
same credential do. Change evidenceCount() to targets.size() instead
of reasons.size(); reasons() still keeps every event, including
same-target repeats, so it can be longer than evidenceCount(). A real
outage still hits every target on the credential, so this loses no
true-positive coverage while cutting a real false-positive path.
Replace the hardcoded API-Error check with a target-keyed BackendErrorPatternLookup
plus a BackendErrorSink, mirroring the existing ExhaustedPatternLookup/ExhaustionSink
pair. Classifies in all three paths (normal block, #211 raw-scrape fallback, and the
fleetd#164 MIN_TURN_NANOS floor). The sink fires only after Rendezvous.resolveFailure
wins for the exact waiter. A target with no configured pattern still falls back to the
narrow (?i)\bAPI Error\s*: compatibility pattern. Existing constructors keep compiling
via BackendErrorPatternLookup.legacy() / BackendErrorSink.none() defaults.
Public send result is unchanged (still a failed send) — the typed sink event is the
internal seam Unit 5 will consume.
Add BackendOutagePolicy: two classified backend errors on the same
credentialId within a 60s window mint one Incident and start a 60s
cool-off for that credential, one atomic ConcurrentHashMap.compute()
per credentialId so a concurrent second and third event can never
both cross the threshold. Errors during cool-off are ignored outright
(no extension, no incident); once cool-off elapses the next error
clears old evidence, requiring two fresh errors to rearm. This is a
new class, deliberately not BackendQuarantine (wrong store, wrong
1800s duration, misleading "exhausted" semantics for a 60s transient
fault). Knows nothing about panes, profiles, sessions, launchers, or
leads — takes events in, returns incidents out.
Round 2's tests never reached Fleetd.java at all: both new tests declared
their OWN local copy of the forwarding shape instead of calling production's.
Mutating Fleetd.java's real forwarder back into the broken lambda left those
copies untouched, so the whole suite stayed green while production had
regressed to exactly the bug being fixed -- proven live by the reviewer.
Fix: extracted the forwarding shape into one named factory,
ExhaustionSink.forwardingTo(Supplier<ExhaustionSink> target), with the
"why a lambda here is wrong" explanation moved onto it (the one place the
shape is now written). Fleetd.java's forwarder collapses to one line:
ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);
Both new tests now call this same factory instead of rebuilding an anonymous
class inline, so they exercise the identical object production builds:
- ExhaustionSinkForwardingHazardTest: calls ExhaustionSink.forwardingTo
directly and asserts the hint reaches the real sink through it.
- OpenCodeLauncherTest#theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop:
same factory call, inside the full Fleetd-shaped construction order
(forwarder built first, real sink pointed at via the AtomicReference
afterward), driven through the real SessionManager.acquire() path.
Mutation proof, this time on production code only: deleted the factory's
3-arg override (falls back to the interface default, dropping the hint) --
both new tests go red with no test file touched:
ExhaustionSinkForwardingHazardTest...: expected: <gx> but was: <null>
OpenCodeLauncherTest...ForwardingHop: expected: <true> but was: <false>
Tests run: 68, Failures: 2
Restored, re-ran: green (Tests run: 68, Failures: 0). Confirmed Fleetd.java
carries no lambda ExhaustionSink anywhere (grep). ExhaustionSink.none() stays
a lambda on purpose -- both its overloads are true no-ops regardless of
arity, so there is no hint to drop.
mvn clean install: Tests run: 1129, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
The round-1 fix was dead on the real production path. Fleetd.java:177 builds
a forwarding sink (needed because the adapters are constructed before
`sessions` exists, breaking a genuine cycle) as a LAMBDA:
ExhaustionSink forwardingExhaustionSink =
(target, reason) -> exhaustionSinkRef.get().onExhausted(target, reason);
A lambda can only implement the interface's one abstract method (the 2-arg
overload), so it silently inherited the 3-arg overload's default body, which
drops the profile hint and calls back into the 2-arg method. OpenCodeLauncher
is constructed with this forwarder, so the hint it supplies (its own
already-known profile name) was thrown away before it ever reached the real
sink built later in Fleetd.main -- reproducing the exact silent no-op round 1
was sent to fix. The 1127 tests from round 1 all injected a sink directly
into OpenCodeLauncher and never went through this forwarding hop, so none of
them could see it.
Fix: forwardingExhaustionSink is now an anonymous class overriding both
overloads, each delegating to whatever exhaustionSinkRef currently holds.
Audited every other ExhaustionSink value in main/: the only other one is
ExhaustionSink.none() (a lambda), which is safe regardless of arity since
both its 2-arg body and the inherited 3-arg default are true no-ops.
New tests:
- ExhaustionSinkForwardingHazardTest: isolates the hazard at the interface
level (a lambda forwarder drops the hint; an anonymous-class forwarder
does not), independent of Fleetd.java's specific wiring.
- OpenCodeLauncherTest#theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop:
replicates Fleetd.java's actual construction order (forwarder built and
handed to the launcher first, real sink built and pointed at via the
AtomicReference afterward) and drives the quarantine through it via the
real SessionManager.acquire() path.
Both proven by mutation: temporarily rewriting each fixed forwarder back
into the pre-fix lambda makes its test fail with a real assertion message
(both matched exactly: "expected: <gx> but was: <null>" for the interface
proof, "expected: <true> but was: <false>" for the composed-wiring test);
restoring makes it pass again. No reverts were committed.
mvn clean install: Tests run: 1130, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
Defect 1: OpenCodeSessionDiscovery.actualModelForDirectory queried
WHERE directory = ?, the same heuristic sessionIdForDirectory uses. Since a
default fleet_spawn (no worktree:) shares the lead's cwd with every other
worker and every past session ever run there, the model read-back could
silently compare against a DIFFERENT session's row. Renamed to
actualModelForSessionId(sessionId), keyed on the primary key id instead, and
made OpenCodeLauncher's SessionAwareHandle cache the resolved id once
non-null (AtomicReference) so a later sibling row in the same directory can
never flip which session's evidence is read. sessionIdForDirectory (#209) is
left directory-based on purpose, with a comment explaining why the heuristic
is unavoidable at that layer.
Defect 2: the ERROR log claimed "quarantining this profile's credential" but
Fleetd's ExhaustionSink lambda resolved target -> roster -> profile ->
credential, while OpenCodeLauncher's model-mismatch check fires from
agentSessionId() during SessionManager.acquire(), before the session is
registered in the roster -- the lookup found nothing and silently no-opped.
Added a default 3-arg ExhaustionSink.onExhausted(target, reason, profile)
overload (defaults to the 2-arg method, so CompletionResolver's two call
sites are unchanged); OpenCodeLauncher now passes its own already-known
profile name; Fleetd's sink became an anonymous class that tries the roster
first, falls back to the hint, and logs loudly at ERROR naming target/reason
when neither resolves, instead of silently no-oping.
Both fixes proven by mutation: reverting each independently makes its new
test fail with a real assertion message, restoring makes it pass again.
mvn clean install: Tests run: 1127, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
tokenEnv no longer defaults to FLEETD_WORKER_TOKEN. A profile that names no
tokenEnv is stating it needs none, which is different from one naming a variable
that turns out to be unset. requiredSecretEnvVars now warns only for an explicitly
declared tokenEnv, so the permanent false alarm about a token no profile needs is
gone and the real warnings beside it stay trustworthy.
Decision recorded in a code comment: an explicit tokenEnv is checked for every
kind, opencode included. opencode can use its own provider credentials, but an
explicit tokenEnv declares a required host secret for its configured provider.
Second commit fixes a crash the first one introduced. Making tokenEnv nullable
changed what every reader of that value can receive, and ClaudeCodeLauncher:261
passed it straight to env.apply — System::getenv in production, which throws on a
null name. The live profile local-direct (kind: claude-code, baseUrl set, no
tokenEnv) would have crashed on spawn. It is weight: 0 today, so the failure would
have surfaced whenever someone re-enabled it. Now uses the superclass helper
resolveEnv, which already tolerates a null name and which OpenCodeLauncher was
already using.
Reader audit, all four: ClaudeCodeLauncher fixed; OpenCodeLauncher already safe via
resolveEnv; ConfigRef uses Objects.equals; MemberEnvAllowList drops null and blank
names in addIfPresent.
Verified by the lead before merge: reverting the resolveEnv fix makes the new test
fail with a NullPointerException from the null variable name, and restoring it
passes. Independent build: 1125 tests, 0 failures, 0 errors, 0 skipped.
Not verified: the ticket's acceptance criterion 4, a real boot on this host showing
no FLEETD_WORKER_TOKEN line while the WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN lines
are unchanged. A worker cannot restart the daemon it talks through. The lead checks
that at the next redeploy.
Two fixes to HerdrPeerLauncher.waitUntilInjectableOrThrow.
Fix 1: the status call is now guarded. A herdr *_not_found answer — what happens
when the backend process exited rather than being slow — used to escape as a raw
HerdrException, skipping stop() and leaking the pane and tab, with a message that
blamed a slow pane. It now fails immediately, runs the same teardown, and says the
process exited.
Fix 2: the gate can now resolve UNKNOWN with the same StatusRefiner the status
poller uses, behind two guards. It only runs for the claude adapter, because
StatusRefiner.classify reads a Claude Code TUI. And a refined result is accepted
only when the same agents.get sample still reports a non-null agentType. That
second guard closes a trap: a dead pane sits at a shell prompt containing the same
❯ glyph the classifier reads as idle, so refining without corroboration would turn
"the backend died" into "ready to inject".
Verified by the lead before merge: NAME_PREFIX really is "claude"/"opencode" so the
adapter split holds; AgentControl.status(t) was already get(t).status(), so moving
the loop to get() adds no herdr call; and a *_not_found already failed the spawn
before this change, so fix 1 improves an existing failure rather than creating one.
The adapter guard was mutation-tested independently (set it to `if (false)`, the new
opencode test fails with "Expected PeerUnreachableException to be thrown, but nothing
was thrown"; restored, it passes). Independent build: 1123 tests, 0 failures,
0 errors, 0 skipped.
Not fixed, and not claimed: that herdr reports a null agentType for a bare shell
pane is unverified against a live daemon. To be proved by a live spawn on both
backends after deploy. The seat-accounting suggestion in #176 was deliberately not
built — its cause was tested in that issue and not reproduced.
The agentType corroboration (guard b) had positive and negative tests, but the
per-adapter namePrefix guard (guard a) had none — nothing proved that an
opencode pane can never reach StatusRefiner.classify, only that a claude pane
with a null agentType is rejected. Since a live opencode pane always reports a
non-null agentType ("opencode"), guard (b) alone cannot catch a broken guard
(a).
Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt
to OpenCodeLauncherTest: agentType("opencode") (non-null, satisfies guard b on
its own) + pane content containing "❯" (would classify as IDLE) + raw status
UNKNOWN throughout. Asserts the gate still times out, and that no agent.read
call used source=detection (StatusRefiner.PROBE_SOURCE) — proving the refiner
was never even reached, not just that its answer was discarded. A blanket
"agent.read is never called" does not hold here: HerdrPeerLauncher.readPaneQuietly
reads the pane tail (source=recent) for the timeout log on every timeout,
regardless of adapter, so the assertion is scoped to the refiner's own probe
source instead.
Mutation check performed and reverted before this commit: temporarily changed
`if (!"claude".equals(namePrefix))` to `if (false)` in
HerdrPeerLauncher.refinedInjectable — the new test failed
("Expected PeerUnreachableException to be thrown, but nothing was thrown"),
confirming it actually exercises the guard. Restored the guard and reran —
test passes again (1/1). No production code changed in this commit.
mvn clean install: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
Fix 1: agents.get(paneId) inside waitUntilInjectableOrThrow was unguarded, so a
herdr *_not_found answer (the backend process exited) propagated as a raw
HerdrException instead of PeerUnreachableException, and skipped teardown
entirely, leaking the pane/tab. Now caught via isAlreadyGone(); fails
immediately (does not burn the rest of the timeout), runs the same teardown
the timeout path runs, and the exception message says the process exited
rather than that the pane was slow. Any other HerdrException still
propagates unchanged.
Fix 2: the gate now resolves a raw UNKNOWN into StatusRefiner's pane-content
classification, like StatusPoller already does mid-life. Guarded against the
trap noted in the ticket: a pane whose backend exited settles at a bare shell
prompt that can also contain the "❯" glyph classify() reads as idle. So a
refined result is accepted only when the corroborating agentType from the
SAME agent.get sample is non-null, and only for the "claude" adapter (namePrefix)
since StatusRefiner.classify is written for the Claude Code TUI only. Refine
only runs when the raw status is UNKNOWN, so a healthy spawn adds zero extra
herdr calls.
Tests added to ClaudeCodeLauncherTest (FakeHerdr gained agentType()/
agentGetFailsWithAfter() fixtures):
- spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout
- spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged
- refinedIdleIsNotAcceptedWhenAgentTypeIsNull
- refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness
- refinementNeverFiresWhenRawStatusIsAlreadyInjectable
mvn clean install: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
Verified by the lead before merge. Read the full production diff, checkModelMatch, parseModel and actualModelForDirectory. Confirmed the check runs on the real #209 late-resolve path (not at spawn, which is why #203 was closed), that unknown/incomplete evidence never quarantines, and that claude-code is structurally excluded because SessionAwareHandle is only built by OpenCodeLauncher.spawn(). Round 2 closed the one gap I found: model JSON with an id but no providerID used to read as a mismatch for a provider-prefixed profile. Independent build: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped.
parseModel already tolerates a model JSON with an id but no providerID (a real shape
opencode can write). checkModelMatch's providerMatches check did not: a provider-
prefixed profile whose id matched but whose evidence had no providerID was reported
as a mismatch and quarantined on incomplete data, which acceptance rule 4 forbids.
Compare the provider only when BOTH the profile requested one AND the evidence has
one. A genuine id mismatch is still caught either way — narrows the check, does not
disable it.
opencode does not fail on an unknown -m <model> flag — it silently falls back to a
default model, which can be a paid credential. Extends the existing late-resolve
path (#209's SessionManager -> handle.agentSessionId() re-poll) so that once the
opencode session row exists, OpenCodeSessionDiscovery also reads its `model` JSON
column and OpenCodeLauncher's SessionAwareHandle compares it against the profile's
configured model.
Comparison rule: split the profile's model on the first '/' into provider+id. Compare
id always; compare provider only when the profile specified one. A bare model name
with no '/' matches on id alone. Absent/unparseable evidence is UNKNOWN, never a
mismatch, so a working profile is never quarantined on missing data. A real mismatch
logs an ERROR naming both models and the profile, then quarantines through the
existing ExhaustionSink path (wired via an AtomicReference forwarding sink in
Fleetd.java to break the sessions/workers/adapters construction cycle).
Verified by the lead before merge: read the full production diff, confirmed the memberHerdrSocket-absent branch is the literal unmodified Files.createTempFile call in its own branch, and that the refusal names the missing key. Measured behaviour recorded: claude 2.1.258 exits 1 immediately on an unreadable --append-system-prompt-file, so the pre-fix bug was the loud readiness-gate failure, not a silent charter-less member. The /tmp full-suite failure the worker reported was the two other #225 copies, fixed by #230 which is already on main; main is built and checked after this merge.
Verified by the lead before merge. Read the full production diff: locking is consistent (every MemberRegistry method uses synchronized(terminalToSlot)), and the refusal happens before the launcher starts a process. Proved the restored fallback test is real by removing the DEV fallback from SessionManager and re-running it: it failed with "expected: <DEV> but was: <ARCHITECT>", then reverted. Independent build: BUILD SUCCESS, 1092 tests, 0 failures, 0 skipped.
The helper I copied from OpenCodeLauncherTest read the CWD's owning
group instead of the process's real primary group, so it silently
picked up whatever group owns the directory Maven was started from
(staff in a home checkout, wheel under /private/tmp on macOS) rather
than a group the operator is actually in. Replaced with the id -gn
based resolution that landed on #230 for the other two copies of this
helper, same shape and skip wording.
Verified by the lead before merge: read the full production diff, confirmed `group` is normalised to null at GitWorktrees:148 so the `group == null` guard is complete, and confirmed the refusal runs before `git worktree add` so a failure leaves no half-made worktree. Independent build in the worker's worktree: BUILD SUCCESS, 1093 tests, 0 failures, 0 skipped.