The line at Injector.java:383-387 printed two numbers that read as measurements
but were both compile-time constants: READINESS_GRACE_POLLS for the poll count
(the loop's own Target.notReadySincePoll counter was in scope at the same call
site), and READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000 for the elapsed
time — arithmetic on two constants, never a measurement, and wrong in the
direction that says everything ran on schedule.
Fix, copying the LongSupplier-clock shape LeadRollover already uses:
- print t.notReadySincePoll instead of the constant for the poll count (the two
agree by construction on this branch, so no test can tell them apart — the
comment says so honestly).
- add Target.notReadySinceMillis, stamped at the first non-ready sample and
reset at all three sites notReadySincePoll already resets (:304, :355, :389
pre-fix line numbers), to compute a real elapsed time at expiry.
- inject a LongSupplier nowMillis (defaulting to System::currentTimeMillis)
through new package-private constructor overloads so a test can supply a
clock whose advance does not track POLL_INTERVAL_MILLIS.
Tests use a ListAppender to assert on the log message contents, per
LeadRolloverTest's pattern. The elapsed-time test drives the loop with a stub
clock returning two literal, non-derived values so it can fail if the fix
regresses to the constant-arithmetic line — proved by mutation: reverting the
elapsed calculation to READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS turns that
one test red (1 failure); reverting the poll-count print to the constant is an
equivalent mutant (0 failures), because the counter and the constant are
identical at that exact call site by construction.
fleetd clean install: Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0.
Three commits. All verified by the lead in a detached worktree, not taken from the worker's report.
3fb3311 — the /clear-settle and success paths return measured elapsed time and a real nudge count
c87cc25 — print the measured nudge count; fix the same defect on the sibling turn-settle timeout line
e966cba — the poll count on the same line was still a constant; the test could not tell the difference
Lead verification at e966cba:
mvn clean install -> Tests run: 1681, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
LeadRolloverTest -> Tests run: 33, Failures: 0, Errors: 0, Skipped: 0
Gitea CI on e966cba: success
Mutation M2 (PICKUP_GRACE_POLLS 8 -> 5), run by the lead: 2 failures,
clearGraceReleaseLogIsWarnWithMeasuredElapsed and successLogPrintsMeasuredElapsedForTheWholeRoll.
The emitted line under the mutant read "after 5 consecutive IDLE/DONE polls (4 of those were
nudged)", so both numbers now follow the loop. Restored, shasum byte-identical, control 33/0 green.
Known and deliberate limit, stated in the code comment: on the grace-release branch both counters
equal the constants by construction (one write site each), so no test can prove the difference on
that line. The change buys one source of truth, not provable coverage. The place nudges genuinely
varies with the run is the /clear-timeout warn, which is covered by a discriminating test.
Two more fixes on the same line, LeadRollover.java:600-624:
1. The poll-count argument (second, was PICKUP_GRACE_POLLS) now prints the
loop's own idlePollsAwaitingPickup counter instead of the constant. Same
defect shape as the nudges fix from c87cc25, one argument over.
2. LeadRolloverTest's clearGraceReleaseLogIsWarnWithMeasuredElapsed asserted
its nudge-count expectation as `LeadRollover.PICKUP_GRACE_POLLS - 1` — the
same expression the production code used to build the log line from, so it
could not discriminate a reverted fix. Rewritten to plain literals
("after 8 consecutive IDLE/DONE polls (7 of those were nudged)"), proven to
trip when PICKUP_GRACE_POLLS's value changes (2 failures under a
PICKUP_GRACE_POLLS=5 mutation: this test and successLogPrintsMeasured...).
Also corrected the comment above the log line: idlePollsAwaitingPickup and
nudges each have exactly one write site on this loop's release branch, so they
cannot differ from PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 at this call
site — confirmed by re-deriving the loop's control flow, and by reverting just
the nudges argument (M1) and observing 33/0 stayed green even after the test
rewrite. That is an equivalent mutant on this line, not a gap the test rewrite
could close; the honest value of printing the counters is one source of truth
for the loop, not a provable-by-test difference here. The place nudges truly
varies with the run — and is covered by a test that can tell it apart from a
constant — is the /clear-timeout warn's clearResult.nudges() in runRollover.
mvn -Dtest=LeadRolloverTest test: 33/0. mvn clean install: 1681/0, BUILD SUCCESS.
Same-shape sweep (found, not fixed, per instructions):
Injector.java:383-387 — the readiness-grace warn prints the constant
READINESS_GRACE_POLLS where the measured per-target counter
t.notReadySincePoll is in scope (single increment site, just reached the
threshold at this call site — same equivalent-mutant situation as this fix).
- waitForClearPickupAndSettle's grace-release warn now prints the measured 'nudges' counter
instead of the constant PICKUP_GRACE_POLLS - 1. The two happen to agree today, but the
constant expression was wrong once before (printed PICKUP_GRACE_POLLS itself, claiming 8
nudges where 7 went out) and a code read did not catch it — only a mutation test did.
Printing the counter cannot drift from the loop's real behaviour.
- waitUntilAtTurnBoundary (the FIRST wait, ~line 386-391) had the identical 'configured value
printed as if measured' defect as the three lines fixed in the original #494 commit, but was
out of scope because the brief named specific lines instead of the shape. Fixed the same way:
it now returns a TurnSettleResult(settled, elapsedMillis) instead of a bare boolean, and the
timeout warn prints 'configured={}s elapsed={}ms' instead of presenting cfg.turnSettleSeconds()
as the measured wait.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Test additions: clearGraceReleaseLogIsWarnWithMeasuredElapsed now also asserts the measured
nudge count; successLogPrintsMeasuredElapsedForTheWholeRoll's expected elapsed value is
updated (7000ms, not 6500ms) to account for waitUntilAtTurnBoundary's own new clock read; a
new turnTimeoutLogPrintsMeasuredElapsedNotJustConfigured test pins the sibling line.
- Proved the nudges fix with a temporary mutation: set PICKUP_GRACE_POLLS to 5, confirmed via
two greps that the mutant applied and the original constant was gone, ran the grace-release
test and read the actual log line — nudge count followed to 4 (= 5 - 1), then restored to 8
and reran the full LeadRolloverTest suite as a control (33/33 green).
- runRollover's /clear-timeout warn now prints configured/elapsed/nudges, each labelled,
instead of presenting cfg.clearSettleSeconds() as if it were the measured wait.
- waitForClearPickupAndSettle's pickup-grace release is now log.warn (was log.info) and
prints the measured elapsed time next to the target pane — this is the exact path that
reported a false-success roll in the real incident (438ms of a 20s budget).
- The success line ('lead-rollover: rolled') now prints the measured elapsed time for the
whole roll.
- waitForClearPickupAndSettle now returns a ClearSettleResult(settled, elapsedMillis, nudges)
instead of a bare boolean, so callers can log the measured values instead of the config.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Adds 3 tests to LeadRolloverTest pinning the content of each changed log line, using a
self-advancing fake clock so the measured elapsed/nudge values are deterministic and
provably distinct from the configured budget.
On this host the lead's cwd IS the fleetd checkout, so fleetd/.gitignore
protects the handover file and the distinction is invisible. On fleet01 the
lead works in /home/ltms/LTMS/kb while fleetd sits in a different directory,
and that repo has no .handover rule — measured 2026-09-12.
Tell the lead to check its own workspace's .gitignore before writing, and to
report it rather than committing the file or editing someone's .gitignore.
Refs fleetd #491, #487, #480.
The skill said the roll 'does nothing until #489 is merged and redeployed'.
Both have happened, so that sentence now reads as a green light. It is not
one: no roll has bootstrapped a fresh session end to end yet. Say that
plainly, and tell the lead to warn the operator before confirming.
Refs fleetd #489, #480.
Fixes the live defect measured on 2026-09-12: the roll joined /clear and
bootstrapText into one line and Claude Code refused it as
"Unknown command: /clearFresh".
Verified by the lead before merge:
- mvn clean install: Tests run: 1677, Failures: 0, BUILD SUCCESS
- LeadRolloverTest baseline: 29 tests, 0 failures
- mutation PICKUP_GRACE_POLLS 8 -> 1: 1 failure (the regression test)
- mutation deleting the !clearSettled guard: 3 failures, 2 pre-existing tests
- mutation replacing waitForClearPickupAndSettle with "return true": 5 failures,
including the strengthened pickup test
- positive control after each restore: 29 tests, 0 failures
- CI run 1738 on f687046: success
A separate mutation of the FIRST gate hung the suite instead of failing it.
That is fleetd #486, not a regression here; the reproduction is recorded there.
Three review corrections on top of the previous commit:
1. The class javadoc's four-step continuation list (lines 47-58) was stale.
Step 1 said "report an injectable state", but waitUntilAtTurnBoundary's
own javadoc excludes BLOCKED - fixed to say IDLE or DONE. Step 3 still
described the old plain re-check ("the original, pre-correction wait...
still here") - fixed to describe what waitForClearPickupAndSettle
actually does: nudge while unpicked-up, then wait for a real WORKING ->
IDLE/DONE boundary, releasing rather than wedging if WORKING never shows.
2. PICKUP_GRACE_POLLS=8 bounds the number of consecutive not-yet-picked-up
polls, not the number of nudges - the 8th poll releases instead of
nudging again, so 8 polls produce 7 nudges. The log.info in the release
branch and two javadoc spots said "8 nudges"; fixed all three to state
the poll count and the nudge count separately and correctly. Behavior
and the constant are unchanged.
3. pickupSeenStopsNudgingAndBootstrapTextIsSent asserted only promptCallCount
and sendKeysCallCount, both of which a return-true stub also satisfies.
Added an assertion on the already-tracked postClearGetCalls counter
(>= 2), which only a real post-/clear poll loop can produce - this is
what makes the test fail against a return-true mutant.
LeadRollover.runRollover's second wait (after /clear) was a no-op: it polled
for IDLE/DONE, which /clear itself never leaves since it starts no real turn,
so it always returned true on the first poll. Combined with a direct
agents.send bypassing Injector (deliberate, to avoid wedging the pane), the
submit Enter that accompanies /clear could race the paste and leave it
unsubmitted — bootstrapText then landed concatenated onto the same input
line, exactly as measured live on 2026-09-12.
Replace that second wait with waitForClearPickupAndSettle, which copies the
pickup-nudge pattern Injector already ships for its own post-turn /clear
housekeeping (fleetd #306): nudge agents.submit while the pane hasn't
reported WORKING yet, release after PICKUP_GRACE_POLLS=8 nudges rather than
wedge, and require a real WORKING -> IDLE/DONE boundary once a pickup is
observed. BLOCKED stays excluded from both the nudge and the boundary check,
same as the (unchanged) first wait — a paused live turn is not settled, and
nudging Enter into an open prompt could wrongly answer it.
Adds four tests to LeadRolloverTest covering the paste-race regression
(nudge ordered between /clear and bootstrapText), a confirmed pickup, a
deadline expiry with no boundary ever reached, and a throwing submit().
Section 11 said the bootstrap prompt landing was 'not yet proven
end-to-end'. It is now measured failing: the first real roll joined
/clear and bootstrapText into one line. Nothing was cleared, so the
failure is safe, but a lead that reads the old wording would reach for
fleet_handover expecting it to work.
Refs fleetd #489, #480.
The lead rollover handover file now lives inside the workspace, at the relative
path fleetd.yaml's leadRollover.handoverPath names. It is a snapshot of one
moment's live state, so it must never enter git history.
The handover skill also now says the handoverPath fleetd hands back is always
absolute, even when the configured value is relative — a lead that resolves it
itself can pick a different file from the one the daemon checks.
LeadRollover.resolveHandoverPath's relative branch resolved the configured
handoverPath against leadWorkspace.apply(...) (fleet.leaders.<name>.cwd) but
never forced the result absolute. If an operator writes a RELATIVE cwd, the
returned path stays relative, silently breaking the "always absolute"
contract documented on PendingRollover.
Fix: call toAbsolutePath() unconditionally on both branches (the
already-absolute input branch, where it is a no-op, and the relative
branch), so neither branch trusts isAbsolute() alone to already imply what
toAbsolutePath() enforces. Method javadoc now states the absolute result is
guaranteed, not merely usual.
Added a test: a lead with a RELATIVE cwd and a relative handoverPath still
yields an absolute PendingRollover.handoverPath. Asserts both isAbsolute()
and the exact resolved value, since isAbsolute() alone would also pass for a
path resolved against the wrong base.
Proved the test discriminates: reverting the toAbsolutePath() calls (keeping
the test) made it fail with an AssertionFailedError ("expected: <true> but
was: <false>"); restoring the fix made it pass again.
Note: FleetConfig has no validation on fleet.leaders.<name>.cwd at config
load (grep across every validate* method: 0 matches for .cwd()) — a relative
cwd is silently accepted. Not adding validation here per instruction; that
is a separate ticket.
Add FleetdLeadRolloverWorkspaceLookupTest, calling the package-private Fleetd.leadRollover(...)
factory directly (with a real ConfigRef built from a temp fleetd.yaml, never the gitignored live
one) to prove the terminal -> lead-name -> Leader.cwd() lookup it builds actually works: a relative
handoverPath resolves against the calling lead's configured cwd; a terminal absent from the live
lead-terminal map falls back to user.dir; and the lookup is read live, not snapshotted at
construction time (a lead discovered by the tab scan after leadRollover(...) was built still
resolves correctly).
Proved this closes the gap: mutating the factory's lambda body (String leadName = null;, always
"no lead found", which forces the daemon-cwd fallback this ticket exists to fix) left the full
1669-test suite green before this commit. With the new test added, the same one-line mutation now
fails 2 of its 3 cases; reverting it goes green again (3/3). Mutation applied/reverted only during
verification and is not part of this commit (git diff on Fleetd.java is empty).
FleetdLeadRolloverWiringTest's class javadoc corrected: it previously claimed no behavioural test
could catch this wiring dropping out, which was true only before this commit and only covered the
factory's own body, not its call site. Restated what each test class actually covers: the source-
text pin covers the call site's argument list; the new behavioural test covers the lambda's body.
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the
calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that
lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores
only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed
back to the lead in the fleet_handover open response, and the default bootstrapText sentence
all see the same absolute location instead of a value resolved against whatever directory the
daemon process happened to start in.
FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it
would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor
(resolvedHandoverPath) method builds the default sentence from the resolved path instead.
Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to
lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on
every call, never off a startup snapshot.
fleetd #480 merged in #483, #484 and #485, so the instruction surface has to
catch up. Per CLAUDE.md's own rule, a code change that silently invalidates the
canonical block is an incomplete change.
CLAUDE.md + wiki/7-Use-Cases.md: one new intent-to-tool row for fleet_handover.
Both copies edited identically; the byte-identical sync check prints True.
The row carries the trap rather than just the call: open FIRST, then write the
file, then confirm — because confirm refuses the file as stale unless its
modified time is later than the open request, so the obvious order fails.
.claude/skills/handover/SKILL.md: the skill said "do not call a fleet_handover
tool: it does not exist." That was true when it was written this morning and is
now false, which is the worst state for an instruction file to be in. Replaced
with the two real paths (by hand, or with the tool when leadRollover: is
configured), plus a new section 11 giving the three-step order and the five
things that surprise a caller — chiefly that `accepted` does not mean the pane
has been cleared, and that operatorConfirmed is a report of what a human said,
not a confidence level.
Section 11 also states plainly that the bootstrap prompt landing in a freshly
cleared pane is not yet proven end-to-end, and that the recovery is the manual
path — which is why the file is written before confirm, never after.
fleet_handover{action: "open"|"confirm"|"cancel"} drives LeadRollover, which #483 and
#484 landed with nothing calling it. Primary-only via a new Authz.Action.HANDOVER, on
the same case line as SPAWN/STOP/DRAIN.
The tool has NO terminal, session or leadTerminal parameter of any kind — the pane is
always callerTerminal(exchange), resolved from the connection. A lead can therefore only
ever roll itself, never another lead. That is charter invariant 3, and it is the second
of the two corrections recorded in LeadRollover's class javadoc.
Registered unconditionally, so the tool surface does not vary with config: with
leadRollover: absent, every action returns a clean NOT_CONFIGURED refusal instead of
failing, and open()'s IllegalStateException (config removed by a hot reload after
construction) is caught and turned into the same refusal. A config-dependent tool set
would have collided with #474's charter tool-surface gate and McpContractDocTest.
Correction round applied before merge, and it is the reason this took two passes.
The unit first shipped with a defaulted 15-argument FleetMcp constructor delegating to
the new 16-argument one with leadRollover = null. I mutated the wiring rather than
reasoning about it: deleting just the leadRollover argument from Fleetd.main's FleetMcp
call compiled with 0 errors and passed all 1659 tests, BUILD SUCCESS — while the live
daemon would have answered NOT_CONFIGURED to every fleet_handover call for ever.
Neither FleetMcpHandoverTest (it builds its own FleetMcp) nor FleetdLeadRolloverWiringTest
(it pins that LeadRollover is constructed, not that it is passed on) could see it.
The worker then found the defect was wider than I had named: all five shorter
constructors (11/12/13/14/15-arg) formed one defaulting chain into the 16-arg one, each
silently supplying another feature's "off" value — leadChannel, outage, leadSeats, peers,
and finally leadRollover. All five are deleted. FleetMcp now has exactly one public
constructor, so every one of those features is compile-enforced at its call site, not
just this one.
Verified by the lead before merge, on PR head merged with current main (0176378):
- CI run 1728 green on eb05576.
- The identical mutation re-run by me: control against the original -> 1; mutant present
(MUT-DROP-ARG2) -> 1; original gone -> 0; `mvn -o -q compile` now FAILS —
"constructor FleetMcp ... cannot be applied to given types; reason: actual and formal
argument lists differ in length" at Fleetd.java:[698,24]. The antidote holds: required
parameter = compile error, defaulted overload = silent survivor a green suite vouches for.
- Restored clean (0 changed files), then mvn clean install on the merged tree:
BUILD SUCCESS 1, BUILD FAILURE 0, Tests run: 1662, Failures: 0, Errors: 0, Skipped: 0.
- Three call sites of `new FleetMcp(` measured, all accounted for: Fleetd.java,
FleetMcpAuthzTest.java, FleetMcpHandoverTest.java.
FleetMcp had a defaulted 15-argument constructor that delegated to the new
16-argument one with an implicit null for leadRollover. Dropping the
leadRollover argument from Fleetd.main's FleetMcp(...) call fell back to that
shorter overload, compiled fine, and left all 1659 tests green — the live
daemon would then answer NOT_CONFIGURED to fleet_handover forever with
nothing going red.
Delete every overload that could reach the 16-arg constructor with a
silently-defaulted leadRollover (11/12/13/14/15-arg forms all chained to it),
leaving the 16-arg constructor as FleetMcp's sole public constructor. Update
FleetMcpAuthzTest's call site to pass every parameter explicitly (leadChannel
null, OutageSource.none(), LeadSeatSource.none(), List.of(), leadRollover
null) — Fleetd.java and FleetMcpHandoverTest already called the full form.
Proved with mvn -o -q compile: removing the leadRollover argument from
Fleetd.main now fails to compile instead of silently defaulting.
No behaviour changes — NOT_CONFIGURED refusals are unchanged.
LeadRollover's two settle waits used AgentStatus#injectable(), which is
IDLE || BLOCKED || DONE. That is the right rule for Injector ("may I deliver a message
without stepping on a live turn") and the wrong one here ("has the turn actually
ended"), because BLOCKED is a live turn that is merely paused — a pane sitting on an
approval prompt.
Path in: the lead calls confirm(); its turn carries on and hits anything needing
approval; herdr reports blocked; within 250ms the deferred continuation reads that as
settled; /clear is typed into an open prompt. That destroys the lead's live context
mid-turn, which is exactly what the turnSettleSeconds gate added in #483 exists to
prevent. The window is turnSettleSeconds, default 20s.
Both waits now require a real turn boundary — IDLE or DONE. AgentStatus#injectable() is
untouched: it is correct for Injector, LeadHeartbeatLoop, LeadCoordLoop, ReplyPushLoop
and HerdrPeerLauncher's two readiness checks, all of which are delivery gates.
Verified by the lead before merge:
- CI run 1726 green on head e2a91e8.
- Mutation on the half the worker did NOT mutate: dropped the DONE arm, leaving
`if (status == AgentStatus.IDLE)`. Mutant proven applied with two unrelated proofs
using different search strings (MUT-DROP-DONE present -> 1; 'IDLE || status' gone
-> 0). Same single test, same 100s cap: unmutated passes in 0.172s; mutated is still
running at 100s and does not pass. So doneStatusStillCompletesTheFullRoll really does
pin the DONE arm, and the fix did not over-tighten to IDLE-only.
Why the earlier #483 battery could not have caught this: it proved the guard FIRES, not
that the guard's predicate was tight enough. LeadRolloverTest drove the fake status with
only "idle" and "working"; neither value discriminates this defect, only "blocked" does.
Correction round applied before merge: both warning lines said "never went idle" and
"did not become injectable". Retiring that word is the whole point of the change, so a
future reader debugging from those lines would have been handed back the wrong mental
model. Both now say "did not reach a turn boundary (IDLE or DONE)".
Filed separately as #486, deliberately not fixed here: with the DONE arm removed the
test hangs rather than fails, because waitUntilAtTurnBoundary is bounded only by an
injected clock that the tests hold fixed. Not a production bug — nowMillis is
System::currentTimeMillis there — but it turns a future regression into a CI hang
instead of a red test.
Both waitUntilAtTurnBoundary guard messages still said "never went idle" /
"did not become injectable" — the old mental model the rename was meant to
retire. Made both say what the code now actually waits for: a turn boundary
(IDLE or DONE).
Adds the fleet_handover tool (open/confirm/cancel) as a thin adapter over
LeadRollover, registered unconditionally so the charter tool-surface gate
sees a stable set regardless of whether leadRollover: is configured. With a
null LeadRollover every action degrades to a clean NOT_CONFIGURED refusal
instead of throwing. Gated on a new Authz.Action.HANDOVER (primary-only,
same as SPAWN/STOP/DRAIN). The caller's own connection-resolved terminal is
the only lead identity ever used — the tool's input schema carries no
terminal/session/leadTerminal parameter, so a lead can only ever roll
itself. Fleetd.main now passes its existing leadRollover local into FleetMcp
via a new trailing constructor parameter.
LeadRollover's waitUntilInjectable used AgentStatus#injectable(), which
accepts BLOCKED. A BLOCKED pane is paused mid-turn on a prompt, not
settled — reusing injectable() let /clear (or the bootstrap text after
it) fire into an open approval prompt within the 20s settle window,
destroying the lead's live context.
Renamed the helper to waitUntilAtTurnBoundary and restricted both waits
to IDLE or DONE only, with a comment explaining why this class does not
reuse injectable() (it answers "may I deliver", not "has the turn
ended"). Added tests for BLOCKED-forever on both waits (zero sends /
exactly one send) and for DONE still completing the full roll.
Two corrections were applied to the original unit before this merge, and the PR
description above still describes the pre-correction shape:
1. confirm() no longer rolls inline. It validates every gate, then hands a one-shot
continuation to continuationRunner and returns. confirm() is called BY the lead FROM
its own turn, so the pane is WORKING and cannot report injectable until confirm()
returns; the old code sent /clear first and then timed out waiting, which destroyed
the lead's context and started no fresh session. The continuation waits for the
calling turn to settle FIRST (turnSettleSeconds, a new knob), and if that wait times
out it sends no /clear at all.
2. The pane to roll comes from the caller's terminal, not PrimaryRegistry. A single-slot
lookup let lead X's confirm() clear lead Y's pane. open() records the caller's
terminal; confirm() refuses with NOT_YOUR_ROLLOVER on a mismatch.
Verified by the lead before merge:
- CI run 1722 green on head a942622.
- Mutation battery on the two safety-critical branches, in a scratch worktree at a942622.
Controls measured against the ORIGINAL first; each mutant proven applied with two
unrelated proofs using different search strings.
M1 `if (!turnSettled)` -> `if (false)`: turnThatNeverSettlesSendsNoClearAtAll FAILS
(LeadRolloverTest:247, expected 0 sends but was 1).
M2 ownership check -> `if (false)`: aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover
FAILS (LeadRolloverTest:266). Unmutated harness-proof run: exit 0 with both guard
WARN lines logged, so the branches really are exercised.
Nothing calls LeadRollover yet. The fleet_handover MCP tool is Unit C.
Two defects found after the fact, both from the original brief, both fixed here.
1. confirm() is called FROM the calling lead's own turn, so its pane is still
WORKING and can never report injectable inside that same call. The old
confirm() sent /clear before polling for that — the poll always timed out,
but only after /clear had already fired and queued, destroying the lead's
context with no fresh session ever started and a refusal return that lied
about what had happened.
Fix: confirm() now only validates and, if every gate passes, hands a
one-shot continuation to a new continuationRunner (a real virtual thread in
production, Runnable::run in tests) and returns RollDecision.approved()
immediately - "scheduled", not "rolled". The continuation itself does the
actual work, once the calling turn has ended: wait for the SAME pane to
report injectable again (new turnSettleSeconds config key, default 20) -
if this never happens, /clear is NEVER sent, at all - then /clear, then
wait again (clearSettleSeconds, as before), then bootstrapText. The "no
timer/scheduler, only confirm() can roll" invariant is restated precisely
in LeadRollover's class javadoc: it is about initiative, not synchronicity
- a single-shot continuation of an already-approved confirm() call still
satisfies it; a recurring background loop would not.
2. confirm() resolved the pane to clear via PrimaryRegistry.primaryTerminal(),
a single-slot lookup that is correct for a background loop with no caller
but wrong here: on a daemon with more than one labelled lead tab, lead X's
confirm() could clear lead Y's pane, violating the charter's "identity
comes from the connection, never an argument" invariant.
Fix: open() and confirm() now take the caller's terminal id as a parameter
(resolved by the MCP layer from the connection - the later MCP-tool unit
must pass it in, never accept it as a request field). confirm() refuses
with a new NOT_YOUR_ROLLOVER reason unless it matches the terminal open()
recorded. LeadRollover no longer depends on PrimaryRegistry at all.
Also: renamed RollResult to RollDecision (rolled -> accepted) to reflect the
new meaning - approved and scheduled, not necessarily cleared yet. Added
turnSettleSeconds to the leadRollover: config block (documented in
fleetd.example.yaml alongside the existing keys) and updated Fleetd.java's
leadRollover(...) factory to drop the primaryRegistry parameter, with
FleetdLeadRolloverWiringTest's source-text pin updated to match.
New tests: turnThatNeverSettlesSendsNoClearAtAll (the branch that matters
most - a turn that never ends means /clear is never sent) and
aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover (NOT_YOUR_ROLLOVER),
plus a settle-after-clear timeout test and an open() input-validation test.
LeadRolloverTest: 11 -> 14 tests.
Adds the opt-in leadRollover: config block and LeadRollover, the executor a
later unit's MCP tool will call. A lead writes a handover file, then open()
records a token and confirm() verifies it (exists, non-empty, fresh) and an
operator confirmation before clearing the lead's own pane via /clear (sent
directly through AgentControl, bypassing Injector, same as
ClaudeCodeLauncher#clearContext) and bootstrapping a fresh session. Nothing
but an explicit confirm() call can ever roll a pane - no timer, no heartbeat,
no background thread anywhere in this class.
Wired into Fleetd.java exactly like LeadHeartbeatLoop: constructed only when
leadRollover: is present at startup, and nothing calls it yet - the MCP tool
is a separate, later unit.
Classified leadRollover: as HOT in ConfigRef (joins placement/
memberCredentials/memberLoginShell/models): the executor holds
Supplier<FleetConfig.LeadRollover> and reads every field fresh per call,
unlike LeadHeartbeatLoop's frozen final fields. The one caveat: the object's
construction is still gated on presence in the startup config snapshot, so a
freshly-added block needs a restart before anything exists to call.
Tests: LeadRolloverTest (14 cases covering the 6 hard requirements - no
object without the config block, only confirm() can roll, missing/empty/
stale handover file each refuse by name, requireOperatorConfirm gating, and
the injected wall-clock supplier) and FleetdLeadRolloverWiringTest (source-
text pin on Fleetd.main's construction call, mirroring
FleetdCompletionResolverWiringTest). Also updated the existing
FleetConfigValidateAllTest, FleetConfigWithDefaultsPreservesEveryComponentTest,
ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest to
account for the new record component.
The skill shipped in #481 described fleetd #480's automation in the present
tense: "fleetd ticket #480 lets a lead session hand off to a fresh one". The
tool is not built. A session loading the skill would look for a fleet_handover
tool it cannot call, and this repo already knows what a confidently wrong
instruction costs — a wrong comment stops the reader investigating with a false
conclusion, which is worse than no comment.
Now says plainly: today the handoff is manual, #480 will automate the same
cycle, and do not call fleet_handover because it does not exist. The nine
procedure rules are unchanged — only who performs the swap changes when #480
ships, not what the file must contain.
My wording to fix, not the worker's: the brief handed them the present tense.
Adds .claude/skills/handover/SKILL.md — the outgoing lead's procedure for writing
the file a fresh lead session inherits.
Reviewed by the lead against the eight requirements in the brief: every number
carries its command, measured-vs-reported is marked, open decisions name their
owner, live hazards, what is not owed, a first section of three re-measurement
commands, a time-and-commit stamp, and a plain statement that the file goes
stale. All eight are present, plus a leave-out section and a template.
Verified by the lead, not taken from the worker's report:
- CLAUDE.md diff is the addendum's skill list only, in the primary-side group.
- The canonical block is byte-identical with the wiki template (sync check True).
- CI run 1718 green on 325d0771, the same commit reviewed.
Follow-up owed by the lead: the skill describes #480's automation in the present
tense, but the tool does not exist yet. The lead will soften that wording so the
skill does not promise a fleet_handover tool a session cannot call.
Adds .claude/skills/handover/SKILL.md, the procedure an outgoing lead
follows to write the handover file a fresh lead session inherits when
fleetd clears the pane. Registers the new skill in CLAUDE.md's
primary-side skills list; no other change to CLAUDE.md.
CLAUDE.md loads into every session. The daemon-redeploy procedure is
needed only when someone redeploys, so it paid for context it did not
use: 58 lines, about 918 est. tokens, every session.
The procedure now lives in .claude/skills/redeploy-fleetd/SKILL.md,
which loads only when invoked. The moved text is byte-identical to
what was removed, plus one new paragraph documenting --no-build
(scripts/redeploy-fleetd.sh:34,62,266) — the script and the wiki
already had that flag, CLAUDE.md never did.
CLAUDE.md keeps an 11-line pointer, because two rules must stay
resident: a merge is not a deployment, and workers must never
redeploy. A session learns it needs the procedure before it needs
the skill, then the next line names the skill.
Also updates the addendum's primary-side skill list, as this file's
own rule for .claude/skills/** changes requires.
CLAUDE.md: 34,442 -> 31,591 chars.
Canonical block untouched — the wiki sync check still prints True.
My battery on the #474 merge found one survivor: reverting Fleetd.java:154
from the three-argument ConfigRef constructor to the plain two-argument
one turns the live reload gate off and leaves all 1633 tests green. Both
new #474 tests build their own ConfigRef with the method reference, so
neither reads what main chose.
This adds FleetdConfigRefWiringTest, following the three source-text
precedents already in the tree (FleetdBackendQuarantineWiringTest,
FleetdLeadSeatWiringTest, FleetdCompletionResolverWiringTest) rather
than the weaker sibling pattern that builds the wiring itself. No
production change.
The worker branched fresh off 4466ee0 rather than continuing its old
branch off the stranded 435e022 base. That was its own call and it was
the right one: one merge base, a clean 80-line diff, and no cherry-pick
needed this time.
Its vacuity guard is worth keeping in mind for the next source-text
test: it asserts the file it read contains 'public final class Fleetd'
before asserting anything about the mutation, with a message saying the
assertFalse below would pass vacuously on a broken read. An unrelated
anchor is the right choice there, because a guard that shares the
mutation's text cannot tell a bad read from a real change.
Fleetd.main's own choice of the three-argument ConfigRef constructor
(with Fleetd::assertChartersNameOnlyRegisteredTools as extraValidation)
was unpinned. Reverting Fleetd.java:154 to the plain two-argument
constructor compiled with 0 errors and left the whole suite green,
because ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest
each build their own ConfigRef directly rather than through main.
Adds FleetdConfigRefWiringTest, a source-text check on Fleetd.java
following the FleetdBackendQuarantineWiringTest/FleetdLeadSeatWiringTest/
FleetdCompletionResolverWiringTest precedent: asserts the exact
three-argument construction is present, asserts the plain two-argument
form is absent, and guards against a vacuous pass on a broken/empty
source read by first asserting an unrelated anchor is present.
A charter naming an MCP tool the server does not register refused Fleetd.main
at startup but slipped through ConfigRef.reload(), because reload() only ran
FleetConfig.validateAll(), which never looks at what a charter's text names.
CharterToolSurface stays in the mcp package (config must not depend on it), so
ConfigRef now accepts the check as a Consumer<FleetConfig> extraValidation,
run inside reload()'s same try/catch as validateAll(). Fleetd.main wires a new
package-private adapter, Fleetd.assertChartersNameOnlyRegisteredTools, into
both the startup call site and ConfigRef's constructor, so the two call sites
can never check different things.
Tests: ConfigRefTest (reload refuses/accepts, via a locally-built equivalent
consumer since Fleetd's method is package-private to dev.ltms.fleet) and the
new FleetdConfigRefCharterToolSurfaceWiringTest (same proof through the exact
Fleetd::assertChartersNameOnlyRegisteredTools reference production uses).
Verified deleting the new extraValidation.accept(fresh) call site fails both
new "refuses" tests by name.
(cherry picked from commit 97e4c1d658)
Lead review. Cherry-picked, not merged: the worker branched from 435e022,
which is not an ancestor of main (I had reset and rewritten that commit's
message while the worker was already on it). Merging its branch would have
added a second merge commit for work already on main. Same content, no
duplicated history. Rule learned: once a worker is spawned, its base
commit is published.
My own mutation battery, five cells, each a full `mvn -B clean test` on
this commit's own tree:
CONTROL 1, unmutated 1633 tests, 0 failures, BUILD SUCCESS
M1 delete accept(fresh) KILLED 2 by name
M3 3-arg ctor stores no-op KILLED 2 by name
M4 set(fresh) before valid KILLED 6
M5 accept outside try KILLED 2 errors
M2 main uses the 2-arg ctor SURVIVED
M2 is a real gap and it is a test gap, not a defect: the production code
here is correct. Reverting Fleetd.java:154 to `new ConfigRef(configPath,
cfg)` turns the live reload gate off and leaves all 1633 tests green,
because both new tests build their own ConfigRef with the method
reference rather than reading what main wires.
That is the fourth Fleetd.main call site to survive a battery (#446 M5
and M7, #466 M1, now this). Extracting a check into a well-tested helper
moves the untested surface UP, into the line that chooses to call it. The
repo already has the answer in three source-text wiring tests
(FleetdBackendQuarantineWiringTest, FleetdLeadSeatWiringTest,
FleetdCompletionResolverWiringTest); the worker followed the other
sibling pattern, which builds the wiring itself and so cannot pin main.
Delegated as a follow-up, with the surviving mutation as its acceptance
criterion.
Also mine to correct: my M3 cell's annotation said the no-op store "must
be 2" and measured 1. The 2-arg constructor delegates rather than
assigning, so my pattern only ever matched the mutant. The cell still
stands on its other proof (requireNonNull left = 0). Third battery in a
row with a wrong CONTROL annotation of my own.
The escalating cooldown landed in 789b6a8, but the only thing an operator
could see was quarantinedForSeconds. A long number does not say whether
this is the first exhaustion or the fifth, and after escalation those look
the same from outside: 3600 seconds could be a big base cooldown or a
credential on its fourth strike.
Both reports now carry quarantineAttempt beside quarantinedForSeconds.
The part that matters is HOW, not that the field exists.
BackendQuarantine#status(credentialId) does ONE quarantines.get() and
returns both values from the same QuarantineState. It is not two accessors
that a caller pairs up. That is the pattern
CompositePeerLauncher.modelGateState() already documents: a report that
reads a different source from the behaviour it describes will eventually
disagree with it, and the disagreement is invisible because both halves
look right on their own.
This is reporting only. Escalation, the ceiling and the quiet-gap reset are
untouched. The flat two-argument constructor still reports a real growing
attempt count even though its cooldown stays flat - which is the honest
answer, since the streak is real whether or not the cooldown uses it.
The worker's report was lost: its fleet_reply never arrived and its inbox
was empty. Nothing was lost with it, because the brief required pushing
first and putting the report in the PR body. That is the second time that
practice has saved a turn.
Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
The comment above the charter writer said flipping that call back to
putArray was "proven load-bearing" and pointed at two named tests. I
mutated exactly that, on the merge commit, and it SURVIVED at 1618 green -
so neither named test covers it, and neither can.
The reason is structural, not a missing test: that writer runs first
against an empty array, so putArray has nothing to replace and the two
idioms are equivalent there. No test can distinguish them. The worker
measured the same thing independently and said so in their report; the
comment was left over from the pre-fix state, where the mutation being
described was a DIFFERENT one (a later writer destroying the charter).
A comment that names tests which do not cover the line is worse than no
comment. The next person mutates the line, sees green, and concludes the
tests are broken.
The corrected version states what each cell actually measures: this line
is unpinnable and why, the later writers ARE pinned and by which test
names, and deleting this line entirely fails the charter-reaches-the-member
assertion - the hole that predates #393.
memberSkills: copied skill folders into every provisioned worktree's
.claude/skills/ and stopped there. Claude Code reads that directory
natively; opencode never does. So the feature was INERT for opencode
members rather than broken: the copy succeeded, the files were correct,
and nothing ever read them. No test failed because there was nothing to
fail - the feature worked at the only layer it implemented.
An opencode member's only channel for static guidance is the instructions[]
array in its generated config. OpenCodeLauncher now adds each seeded
skill's SKILL.md there. A folder with no SKILL.md is never delivered and
the log names it.
The second half is an ordering hazard fleet01 found by reading, and that I
then measured. Three writers append to instructions[]: the role charter,
the seeded skills, and the IDE rules. The charter used putArray
(CREATE-OR-REPLACE) while the other two used withArray (get-or-create).
That was safe only because the charter ran first against an empty array -
an undeclared constraint that nothing tested. Measured on the earlier
merge: making the skills writer use putArray left 1603 tests green while
silently deleting the charter entry, so an opencode member would launch
with no role contract at all. Worse than the bug being fixed, and
invisible.
All three writers now use withArray, and three tests pin the array's
CONTENTS (never its size - a size assertion passes when putArray swaps two
entries for two others) across the combinations that matter:
charter-only, charter+ide, charter+skills+ide.
WHY NOTHING CAUGHT IT, MEASURED RATHER THAN ASSUMED. fleet01 first said
the missing axis was the COMBINATION of writers, then revised that to a
stronger claim: that nothing asserted the charter reaches an opencode
member at all. I checked the second claim on this merge and it is FALSE.
Deleting the charter writer outright fails 6 tests, and 3 of those existed
before #393: writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet,
roleCharterWithoutMcpOrCustomProviderStillWritesAConfig, and
aPinnedEndpointAndTheFleetMcpCoexistInOneConfig. The charter reaching a
member was already pinned.
So their FIRST diagnosis was the right one. The surviving mutation did not
delete the charter write; it made a LATER writer replace the whole array.
Every pre-existing test had exactly one writer active, and with one writer
putArray and withArray are indistinguishable. The gap was never "is the
charter delivered" - it was "are two writers ever active at once", which is
the combination axis. Recording this because the stronger claim is the more
quotable one and it would have sent the next reader looking for a hole that
is not there.
One honest residue: the charter writer's own idiom cannot be pinned.
Flipping it back to putArray leaves the suite green, and always will,
because it runs first against an empty array where the two idioms are
equivalent. The edit removes an undeclared constraint for the next person
to add a writer; it is not a change any test can detect. The comment in
the source claimed two named tests cover it - that was wrong, and the
commit after this one corrects it rather than leaving a false claim beside
the code.
What this does NOT do: opencode has no equivalent of Claude Code's Skill
tool, so the content is static system-prompt text present from spawn, not
something a member can invoke by name. This closes the DELIVERY gap and
cannot close the ACTIVATION gap. That residue is opencode's design.
Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
Add BackendQuarantine#status(credentialId) -> Optional<Status>, a single
QuarantineState read that answers both remainingSeconds and repeatCount
together -- the same "one accessor" pattern CompositePeerLauncher.
modelGateState() already uses, so the two facts can never disagree.
fleet_profiles/REST GET /profiles and fleet_list's capacity rows
(FleetMcp.profilesView/capacityView) now call status() instead of
remainingSeconds() and add a "quarantineAttempt" field beside
"quarantinedForSeconds": 1 for a first occurrence, 2 for the second
in a row, and so on.
No change to the escalation, ceiling, or reset logic itself -- this
unit is reporting only.
A role charter is free text in config that tells a member which tools to
call, and nothing checked that those tools exist. A charter naming
bridge_send - a name CB-634 removed - started the daemon cleanly, and the
member found out at run time by calling something that was not there.
#464 shipped a test for this, but it wrote its own charter into a @TempDir,
so nothing anyone put in the real config could fail it. That was a defect
in my acceptance criteria, not in that work.
Now: FleetTool is one enum of the 11 registered wire names, and every
reader goes through it.
- FleetMcp's schema builders pass FleetTool.X.wireName() instead of a
literal.
- FleetMcp's constructor asserts at startup that what it registers with the
SDK equals FleetTool.wireNames() exactly, in both directions.
- toolAction(String, Map) resolves arbitrary wire input against
FleetTool.byWireName() and keeps its run-time throw, which is necessary -
network input has no closed compile-time form. It then hands off to
authzAction(FleetTool, Map), a switch over the enum with NO default, so a
new tool is a compile error at that layer.
- CharterToolSurface lives in mcp, not config, and Fleetd.main calls it
right after validateAll(). Config must not depend on the MCP server:
config loads before the server exists.
My brief undercounted the problem and the worker corrected it. I said there
were two tool-name inventories plus a test fixture. There were FIVE: the
registrations, the authz switch, and three separate source-text scrapes of
FleetMcp.java in CharterToolSurfaceTest, FleetMcpAuthzTest and
McpContractDocTest - none of which the ticket mentioned. Fixing those three
was required, not scope creep: once the literals moved into FleetTool their
regexes matched zero names, so one would have failed on its vacuity guard
and the other two would have gone quietly vacuous. Inventories after: one.
Verified independently on origin/main before accepting the wider diff:
three test files did read FleetMcp.java as source text, with a control file
at zero to prove the search discriminated.
Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
OpenCodeLauncher.writeConfig has three writers into the instructions[]
array (charter, seeded skills, IDE rules). The charter writer used
putArray (create-or-REPLACE) instead of withArray (get-or-create), which
"worked" only because it happened to run first against a still-empty
array — an undeclared ordering dependency nothing tested. Found by the
fleet01 lead and verified on this branch's merge: flipping the skills
writer to putArray left the full 1603-test suite green while silently
deleting the charter entry, which would launch an opencode member with
no role contract at all.
Fix: charter's putArray -> withArray (one-word change, behavior-identical
today). Add three tests asserting instructions[] CONTENT as an exact
ordered list (not size) across writer combinations: charter only,
charter + IDE rules, and charter + IDE rules + seeded skills. Mutation
testing (see PR body) shows the skills and IDE-rules writers are each
independently detectable by name; the charter writer's own mutation is
not detectable by any test, because it structurally always runs first
against an empty array, so putArray and withArray are equivalent there.
FleetConfig.validateCharters() only checked that a charter key is a role
wire name and its text is non-blank; #464's CharterToolSurfaceTest compared
charter text against the registered tool surface, but wrote its own charter
into a @TempDir fixture, so nothing anyone wrote into the live fleetd.yaml
could ever fail it.
Add FleetTool, an enum in dev.ltms.fleet.mcp holding the one canonical set
of registered tool wire names. FleetMcp's tool schemas now derive their
names from it, its constructor asserts at startup that what it actually
registers with the SDK equals FleetTool.wireNames() exactly, and its authz
dispatch (toolAction/authzAction) resolves the wire string against FleetTool
before switching on the enum itself with no default -- adding a tool without
pinning its Authz.Action is now a compile error, not just a test gap.
Add CharterToolSurface (mcp package, not config -- config loads before the
MCP server exists) and call it from Fleetd.main right after
cfg.validateAll(), so a charter naming a tool the server does not register
refuses the daemon's startup, naming both the charter key and the unknown
tool. FleetdStartupValidationTest proves this through Fleetd.main itself
against a live-shaped config fixture (bridge_send, CB-634's own removed
name).
CharterToolSurfaceTest, FleetMcpAuthzTest and McpContractDocTest each kept
an independent regex scrape of FleetMcp.java's source for the registered
side of their own comparison -- three more copies of the same list nothing
tied together. All three now read FleetTool.wireNames() instead.
Proved canonical by removal: deleting FleetTool.ACK while ackTool() still
referenced it broke mvn compile in two places (FleetMcp.java:940,:1898);
registering a schema under a literal not backed by FleetTool
("fleet_ack_v2") failed FleetMcp's new startup assertion in every test that
constructs it (12 errors, IllegalStateException at FleetMcp.<init>). Both
reverted before this commit.
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.
BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.
Three things on the record because they are judgement calls, not facts:
- The reset is a TIME PROXY, not a success signal. Nothing in this codebase
reports a spawn success back to this class, so "it started working again"
cannot be observed here. A base cooldown of quiet is the best available
evidence, and the class doc says that plainly instead of implying the
stronger thing.
- No automatic probing. That was the operator's design constraint and the
implementation respects it: the daemon warns and waits, it never pokes a
limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
so no new ConfigRef hot/cold/deferred question arises.
quarantineCooldownSeconds stays Deferred and is now the BASE of the
backoff; fleetd.example.yaml and FleetConfig's javadoc say so.
Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.
The old two-argument constructor is behaviourally unchanged - the same
formula with multiplier 1.0 and a ceiling equal to the base, which collapses
to the original flat "now + cooldown". Every existing call site keeps its
shape.
The second commit exists because my own mutation on the first merge found
the wiring unpinned: putting Fleetd.main back on the flat constructor left
all 1608 tests green, so the factory was pinned and the decision to use it
was not. FleetdBackendQuarantineWiringTest closes that, following the five
existing *WiringTest files rather than inventing an idiom. It checks source
TEXT, and its class doc says so: it does not prove the call executes, and it
cannot tell "wrong factory" apart from "renamed the anchor" - both fail the
same assertion. That limit is real and recorded rather than papered over.
Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
BackendQuarantineTest proves BackendQuarantine.withEscalation itself
escalates, is ceilinged, and resets. Nothing proved Fleetd.main still
calls it: reverting main to the flat two-argument constructor compiled
with 0 errors and left the full 1608-test suite green, because every
other test builds its own BackendQuarantine directly.
Adds FleetdBackendQuarantineWiringTest, a source-text assertion on
Fleetd.java (same idiom as FleetdLeadSeatWiringTest and
FleetdCompletionResolverWiringTest) pinning that main's declaration is
built from BackendQuarantine.withEscalation(...), not `new
BackendQuarantine(...)`. Measured: reverting to the flat constructor
fails it (expected true, was false); renaming the anchored local
variable while keeping behaviour identical also fails it loudly, not
silently. The class doc states plainly this checks source text only,
not that the call executes or that escalation reaches a real backend.
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.
BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.
Three things I want on the record because they are judgement calls, not
facts:
- The reset is a TIME PROXY, not a success signal. Nothing in this
codebase reports a spawn success back to this class, so "it started
working again" cannot be observed here. A base cooldown of quiet is the
best available evidence. The class doc says this plainly rather than
implying the stronger thing.
- No automatic probing. That was the operator's design constraint and the
implementation respects it: the daemon warns and waits, it never pokes a
limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
so no new ConfigRef hot/cold/deferred classification question arises.
quarantineCooldownSeconds stays Deferred and is now the BASE of the
backoff; fleetd.example.yaml and FleetConfig's javadoc say so.
Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and is deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.
The old two-argument constructor is behaviourally unchanged - it is the
same formula with multiplier 1.0 and a ceiling equal to the base, which
collapses to the original flat "now + cooldown". Every existing call site
keeps its shape.
Build number and my own mutation results are reported in the PR and on the
ticket, measured on this merge commit rather than on the branch.
GitWorktrees.seedSkills copies memberSkills:-seeded skill folders into every
provisioned worktree's .claude/skills/ and logged "skill seeding: N of M" as
if that were success — but .claude/skills/ is a Claude Code CLI convention.
opencode has no such discovery, so a kind: opencode member never actually
read a seeded skill even though the log said N of M succeeded.
Two changes, both required:
1. Deliver it. OpenCodeLauncher.skillInstructionFiles scans
<cwd>/.claude/skills/*/SKILL.md at spawn time (the one point the launcher
knows both the kind and the cwd) and appends each to the generated
opencode.json's instructions[] array, the same channel already used for
the member charter and IDE rules. A skill folder with no SKILL.md is
named and skipped rather than silently dropped.
2. Stop claiming it where the claim can't be verified. GitWorktrees.seedSkills'
log now says explicitly that consumption depends on the member's kind and
points at the launcher's own log; OpenCodeLauncher logs its own kind-aware
"skill delivery: M of N ..." line once the kind is actually known, naming
any folder it could not turn into an instructions[] entry.
fleetd.example.yaml's memberSkills: doc previously claimed "Claude Code
members only; an opencode member reads a different path (.opencode/agent)
this key does not touch" — false as of this fix, corrected to name both
kinds and how each consumes it.
Tests: OpenCodeLauncherTest gains two cases driving the real
GitWorktrees#add seeding path (not a hand-built fixture) into an
opencode-kind spawn — one asserting a seeded skill's SKILL.md lands in
instructions[] plus the honest log line, one covering a skill folder
without SKILL.md (delivered skills still flow, the malformed one is named
in the log and excluded from instructions[]). ClaudeCodeLauncher is
untouched — its native .claude/skills/ discovery already worked and is out
of scope.
mvn -B clean test: Tests run: 1603, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
A flat 30-minute quarantine retries a weekly subscription limit about
336 times before the window resets. BackendQuarantine now doubles the
cooldown on each consecutive exhaustion of the same credential (no
more than one base cooldown after the previous quarantine's deadline),
capped at 12x the base cooldown (~6h at the 1800s default), and resets
back to the base cooldown once a base-cooldown's worth of quiet has
passed with no further exhaustion.
The flat two-argument constructor is unchanged (equivalent to
multiplier 1.0 / ceiling == base), so all ~20 existing call sites keep
their current shape and behaviour. Production wiring (Fleetd.main)
switches to the new BackendQuarantine.withEscalation factory.
This only touches the exhaustion path (BackendQuarantine's one
production caller is Fleetd.exhaustionSink, fired on BACKEND_EXHAUSTED
alone) and never the separate, unescalated cooling-off mechanism
(BackendOutagePolicy, fixed 60s) that guards against a transient
backend-error storm.
Three rounds. Round 1 made exhaustedPattern a hot config key, made the
quarantine warning name the fix instead of only the fact, and added model/reason
to fleet_profiles' quarantined rows. Round 2 extracted the warning text into
usageLimitFixWarning/usageLimitFixWarningNoModel and pinned both. Round 3
extracted the sink itself into a static exhaustionSink(...) factory and pinned
what it actually logs, using a ListAppender on this class's own logger.
Where the mutation numbers below come from, stated exactly. The battery ran on
merge commit 3d2d521, tree 1dcca23: origin/main at 235644c plus this branch.
This merge commit's tree is 953ce11, which is that tree plus one file --
CharterToolSurfaceTest, from the #464 merge (49df792) that landed on main while
the battery was running. So the battery did not run on this exact tree. That one
added file was verified green on its own merge. The build on THIS tree, tree
953ce11, is: Tests run: 1601, Failures: 0, Errors: 0, Skipped: 0, BUILD
SUCCESS, 0 compile-error blocks. That is the battery tree's 1600 plus that one
test, which is the arithmetic the two trees predict. Saying this rather than
implying one tree.
On the battery tree: 1600 tests green, 0 compile-error blocks. FleetMcp.java
auto-merged there against #463's change to the same file, so that build was also
the gate on the auto-merge -- a clean auto-merge is not a compiling merge.
- M5, the round-2 survivor reproduced verbatim: the call site stops using either
pinned method and logs a literal instead. Round 2 left this green across 1592
tests. Now KILLED by
FleetdExhaustionSinkWarningTest.profileWithModelGetsTheActionableFix and
.profileWithNoModelGetsTheFallbackNotAFix.
- M6, the ternary's two branches swapped, so a profile WITH a model gets the
no-model text and vice versa. Both extracted methods and both their unit tests
untouched. KILLED by the same two tests, independently.
- M7, THE HALF THIS ROUND DID NOT PIN, and it SURVIVED. main()'s call to the
factory replaced with an inert lambda: the factory and its test stay perfect,
the daemon quarantines nothing and logs nothing. 1600 green.
M7 is the same defect shape as round 2, moved one level up, and it is worth
naming plainly. Extracting a thing in order to pin it CREATES the seam the test
then lands on. Round 2 extracted the warning TEXT, pinned the two methods, and
left the call site that selects between them unpinned -- M5. Round 3 extracted
the SINK, pinned the factory's behaviour, and left main()'s wiring of it
unpinned -- M7. The question to ask on any such fix is which of three things a
test now reaches: the value, the call site, or the selection between values.
Extraction only ever answers the first.
M7 is not a reason to hold this merge. Proving main() wires this factory means
driving daemon startup, which nothing here does -- that is fleetd #460. A
cheaper option exists and is recorded there: a source-reading assertion of the
kind #439 used, which would kill M7 without starting the daemon.
Round 3's own caveat, disclosed by the worker and worth keeping: its first Cell A
run was corrupted by a second concurrent mvn against the same module directory.
It killed that run, checked for stray ForkedBooter processes, and re-ran. The
numbers above are mine, from this battery, not that run.
CharterToolSurfaceTest extracts every fleet_* / bridge_* token from configured
launch charters and every tool("fleet_...") FleetMcp registers, then asserts the
first set is a subset of the second.
Verified on the merge commit. Its three acceptance criteria are met:
- Catches the ticket's own example. Fixture charter naming bridge_send, the tool
CB-634 renamed away: KILLED.
- Fails loudly with no charter text. Fixture stripped of every tool name: KILLED
by its named.isEmpty() guard, not a silent pass.
- Fails loudly with no registered tools. The tool("...") scrape broken so it
matches nothing: KILLED by its registered.isEmpty() guard.
- And one cell of my own: the server stops registering fleet_reply, which the
fixture names. KILLED. This is what proves the 'registered' half reads real
production source and is not a second fixture.
The scrape finds 11 registered tools: ack, ask, list, poll, profiles, reply,
send, spawn, status, stop, whoami. An independent count of every "fleet_x"
literal in FleetMcp.java is also 11.
WHAT THIS DOES NOT CLOSE, and it is the ticket's actual gap. The charter half is
a @TempDir fixture the test writes itself, so no charter text anyone writes can
make this test fail. Measured: the test mentions fleetd.yaml 0 times, and the
commit changes 0 production files -- FleetConfig.validateCharters() still never
reads charter text (0 lines of its body mention a tool name). So this pins the
comparison logic and acts as a rename tripwire for the two tools the fixture
names. It does not check the live config. That needs a production-side check and
is filed as a follow-up.
That residue is my ticket's fault, not the worker's: the three criteria I wrote
are exactly the three it met.
A note on my own battery, because it nearly published four false kills. The
first run showed rc=1 on all four mutation cells and I would have read that as
four kills. It was zsh: unquoted parameters are NOT word-split, so
'mvn -B $scope test' passed '-Dtest=X -DfailIfNoTests=false' as ONE argument and
surefire ran zero tests. The tell was a missing 'Tests run:' line. The rerun
proves the harness first -- selector alone must report 'Tests run: 1' -- and
every cell now prints its surefire summary count so a void cell cannot pass for
a kill.
A wrapper overload that is called with no callerIsPrimary argument used to
default it to true, so a forgotten argument silently handed out the lead's
coordination state. It now defaults to false: a missing identity fails closed.
Verified on the merge commit, not the branch:
- The funnel is real. FleetMcp has 7 listFleet declarations and 7 real calls
(an 8th 'listFleet(' match is a javadoc {@link}). Exactly one call writes a
literal for the new boolean, and it writes false; exactly one writes the real
predicate, coordinatorVisibleTo(principal(exchange)) in the MCP handler. No
call writes true.
- Control battery on the merge: 1584 tests green unmutated.
- M1, the fix reverted at the one line that writes the default (false -> true):
KILLED by listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey.
- M2, the half this round weakened. The worker rewrote
listIsByteForByteUnchangedForThePrimaryCaller and dropped its byte-for-byte
equality assertion, which was correct because that comparison ran against the
implicit-default overload -- now the path #463 closes. So: does anything still
notice if the primary's coordinator row silently loses a field? Dropped
heldCount: KILLED by two tests, that same test and
listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero.
- M3, a regression check on #439's own gate, 'if (callerIsPrimary)' -> 'if (true)':
KILLED by three tests.
What is no longer pinned, stated plainly: the primary's output is now checked
field by field, not as a whole string. A field that no test names could
disappear without failing anything. Every field the tests do name is pinned,
proven by M2. The whole-output answer belongs to fleetd #460.
One control in my own battery was wrong and is worth recording: I labelled
', true);' in FleetMcp.java 'must be 0' and it is 2 -- row.put("configured",
true) and m.put("self", true), neither a listFleet delegation. The pattern was
too wide. The count above comes from a walk over each declaration and call
instead.
Round 2 pinned usageLimitFixWarning/usageLimitFixWarningNoModel's TEXT via
FleetdUsageLimitFixWarningTest, but a mutation battery against the merged PR
proved two gaps in the caller that builds main()'s real quarantine
ExhaustionSink: nothing proved the sink's log.warn actually invokes either
method (Cell A), and nothing proved it picks the right one for a profile
with vs without a configured model: (Cell B).
Extract the inline lambda into a new static Fleetd.exhaustionSink(...)
factory (same refactor-for-testability class the lead approved in round 2
for the two static warning methods), behaviourally unchanged from the
lambda it replaces. FleetdExhaustionSinkWarningTest drives this factory's
return value directly and asserts on the real text a ListAppender attached
to Fleetd's own logger captures, covering both cells from one mechanism:
- Cell A (ternary result replaced by a literal string): confirmed red,
1594 run / 1 failure, restored, confirmed green (1594/0).
- Cell B (ternary's two branches swapped): confirmed red, 1594 run /
2 failures (both test methods independently caught it), restored,
confirmed green (1594/0).
Verified by me on a local merge of 29e7a06 onto f5e02fe:
- One file, one hunk. 5 changed lines inside the canonical block, all of them
invariant 5; a line-by-line diff of the block against origin/main shows
nothing else moved.
- The addendum is untouched: both 'Herdr socket tests (measured 2026-09-10)'
and 'Only the lead can run that check' appear once on each side.
- Block size: origin/main 17467 chars / 17626 bytes, the merge 17557 chars /
17718 bytes. The worker's numbers were bytes and were right; I am naming both
units because len(str) in python is characters and this block is full of em
dashes, which is a mistake I have made before.
- mvn -B clean test: Tests run: 1583, Failures: 0, Errors: 0, Skipped: 0.
BUILD SUCCESS, rc=0.
The three things I reserved from the worker, because a provisioned worktree
cannot do them:
- wiki/7-Use-Cases.md updated to match, in the wiki repo's own commit.
- The sync script run in this clone: in sync: True.
- Other projects carrying the block: NONE. 29 CLAUDE.md files under the
operator's project roots, 18 of them carry the block and 17 of those are this
repo's own worktrees, which inherit it from git. Only
LTMS/claude-bridge/CLAUDE.md is a real second copy, and it is this file.
On the worker's criterion-4 question: the addendum note stays. The restated
invariant removes the unsatisfiability, which was the sharp problem, but the
note still names the exact four files the carve-out covers and records why
(fleetd #449, where a fake and the real herdr disagreed for weeks under a green
suite). A rule that is merely satisfiable is not the same as a rule a worker can
apply without re-deriving it.
The compat overload at FleetMcp.java:1373 defaulted callerIsPrimary to a
literal true, so a caller that forgot the argument silently got the
coordinator row (this daemon's coord-id, mailbox state, held-mail previews,
peer reachability) -- lead-to-lead state fleetd #439 just gated. Flip the
default to false: a forgotten argument now yields a missing row instead of
a leaked one.
Six FleetMcpTest methods relied on the implicit true to see the coordinator
row at all; they now pass true explicitly through the canonical overload.
listIsByteForByteUnchangedForThePrimaryCaller's own premise (comparing the
implicit-default path against an explicit-true path) was the shape of the
bug, so it now only exercises the explicit-true path.
Added listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey
to pin the new default: a compat overload called with no callerIsPrimary
argument, against a fully-configured lead channel, must produce a result
with the coordinator key absent -- not empty, not redacted, absent.
This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.
Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:
- Phases 1 to 3 are done. Both systemd user units exist, are active and
enabled, linger is on, one java process (so the old restart.sh
double-daemon problem is gone), the live fleetd.yaml is on the host, and
the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at 4887731, 77 commits behind
origin/main, 0 ahead. So fleet01's daemon runs code from before this
week's merges.
- Phases 6 and 7 are not done. The Mac still runs the daemon this fleet
uses, and fleet01 has 2 weeks of uptime, so no reboot proof exists.
- Section 9 is out of date: claude is installed on fleet01 and a fleet01
lead is live on the coordination channel, so the headless-login blocker
it names is solved.
Also records fleet01's four profiles with their weights, because
placement: weighted plus gx at 100 and xf at 80 means an unqualified spawn
there almost never lands on a claude-code profile.
And one search that found the opposite of what I expected, written down so
the next session does not repeat it: opencode IS installed on fleet01, at
~/.opencode/bin/opencode, but only on the INTERACTIVE PATH. zsh -ic finds
it, zsh -lc does not. A herdr pane is a non-login interactive zsh and
fleetd types the launch command into the pane, so gx and xf are not broken
by this. The daemon's own login-shell ExecStart cannot see it, which is the
mirror of the trap that ExecStart exists to fix: credentials live on the
login side, ~/.opencode/bin on the interactive side.
Invariant 5 banned 'driving the terminal multiplexer directly', naming
herdr CLI and socket as the banned tool. That bans a mechanism. What it
protects is the control plane: nobody may move a fleet session, pane or
peer by a route that skips the bridge's policy checks.
In this repo herdr is itself the subject under test, so four contract
tests must open its socket on purpose (see the addendum's 'Herdr socket
tests' note). Under the old wording, a worker assigned to that code
reads invariant 5 and finds its only path to finish the task banned.
Restate the invariant by purpose: never move fleet state except through
the bridge. herdr's CLI and socket stay as the named example of the
banned route, not the definition of it.
Verified by me on a local merge of c1ca627 onto 1348287:
- mvn -B clean test: see the totals below. Ran in fleetd/, redirected to a file,
exit code captured on its own line.
- Mutation battery on merge 8402923 (same two parents, earlier base 5f1b260),
4 cells, each with a proof gate on the occurrence count:
* CONTROL, unmutated: 1583 tests, 0 failures, rc=0.
* M2r - the handler stops asking who called: coordinatorVisibleTo(principal(exchange))
replaced by a literal true. KILLED by
FleetMcpAuthzTest.theFleetListHandlerActuallyConsultsCoordinatorVisibleTo.
This is the mutant that survived round 1, so the gap the follow-up was for is closed.
* M3 - is the new source-reading detector vacuous? Renamed its anchor
(listHandler -> listHandlerX, 2 sites, behaviour identical). The detector FAILED,
rc=1, as a source-reading test must: it cannot silently pass on an empty scrape.
* M4 - the predicate itself always says yes: return caller.isPrimary() replaced by
return true. KILLED by FleetMcpAuthzTest.onlyThePrimaryMaySeeTheCoordinatorRow.
Tree verified clean before the battery and restored after each cell.
- ANON is pinned too, not only worker and architect: FleetMcpAuthzTest asserts
coordinatorVisibleTo(ANON) is false.
- listFleet( appears in exactly one main file, mcp/FleetMcp.java, and no REST class builds
the coordinator row, so the detector's single-file scope covers every live call site
today. Control for that sweep: 109 main .java files matched a string they all contain.
Not in this PR, and my call, not the worker's: the six compat overloads of listFleet still
default callerIsPrimary = true, which fails open. Safe today because the one production
call site passes the computed value. Filed separately.
The javadoc threw away the whole loaded run as "not clean evidence". The
fixed cell order makes that too strong in one direction. A cell that runs
last on a climbing load has a free explanation for FAILING. It has no free
explanation for PASSING: surviving a worse condition than a fair order
would have given it is evidence in the safe direction. So the old version's
0 of 3 is still discarded, and the three passes are kept with the load each
one ran at.
Also names the cell that tests the swallow explanation head-on, which the
old text left as "nobody has managed that yet". SHELL_READY_TIMEOUT_MS at 0
types input at once — the worst case for "typed before the prompt" — and it
passed 3 of 3 at load 18.42 to 23.65. The wider read window cannot explain
that away, because a swallowed keystroke is lost, not late: the command
never runs, so no amount of polling makes its output appear.
And a warning not to carry the raw load average to another host. Load
average counts differently per core and per operating system, so only load
per core compares. I broke that rule myself when comparing this run with
another host's numbers.
Both points came from the fleet01 lead reviewing 2af13ab.
Review of PR #462 found M2: the coordinatorVisibleTo gate (then an inline
principal(exchange).isPrimary() check) could survive a mutation that
replaced the argument with a literal true at the one production call
site, because every existing test drove listFleet directly and supplied
the boolean itself -- nothing exercised the handler's own call.
- Name the decision: FleetMcp.coordinatorVisibleTo(Principal), a small
package-private predicate next to denyFor/recordPrimarySingleton. The
fleet_list handler now calls coordinatorVisibleTo(principal(exchange))
instead of inlining .isPrimary().
- Pin the predicate's role table in FleetMcpAuthzTest
(onlyThePrimaryMaySeeTheCoordinatorRow), covering primary/worker/
architect and, newly, anonymous.
- Add a source-reading detector at the boundary
(theFleetListHandlerActuallyConsultsCoordinatorVisibleTo), same idiom as
toolsTheServerRegisters/everyRegisteredToolHasItsHandlerActionPinned: it
reads FleetMcp.java, isolates the listHandler block, asserts (as a
control) that the block actually contains a listFleet( call, then
asserts the call's trailing boolean argument is exactly
coordinatorVisibleTo(principal(exchange)) -- not a literal true/false.
Both mutations from the review were reproduced and killed by these tests,
then reverted; see the PR body for the full break-and-restore transcript.
I put "the sync script must print in sync: True" in a worker's acceptance
criteria for #455. The worker could not run it and said so, honestly, instead
of inventing a pass. My brief was the defect.
A member's provisioned worktree has wiki/ uninitialized, so the script dies
with FileNotFoundError. Measured in three worker worktrees: git submodule
status printed a leading '-' and wiki/ held 0 entries. The primary's own clone
printed a leading '+' and the file was there.
The note is dated, gives the re-measure command, says what each outcome means,
and says to delete it once it stops reproducing - as this file requires of any
measurement in an addendum.
Canonical block untouched: the sync script itself reports in sync: True.
A mutation battery run against merged PR #457 (d02dd1b) proved criteria 2 and 3
shipped without a test that could catch them breaking: deleting either
row.put("model", model) or row.put("reason", reason) in FleetMcp.profilesView
left all 1586 tests green (rc=0), and renaming the "usage-limit fix:" log tag
to something meaningless did too. Criterion 1's own mutation (LiveExhaustedPatterns
snapshotting instead of reading live) was correctly killed by the existing
FleetdExhaustionDetectionArmedWiringTest/LiveExhaustedPatternsTest — only 2 and 3
were unguarded.
- Extracted the exhaustionSink WARNING text out of two inline SLF4J {}-placeholder
log.warn calls into two static methods, Fleetd.usageLimitFixWarning(profile, model)
and Fleetd.usageLimitFixWarningNoModel(profile, quarantineCooldownSeconds), the
same extracted-static-method + dedicated-test idiom as
exhaustedPatternCoverageLine/errorPatternCoverageLine. log.warn is now called with
each method's return value as a single already-formatted argument, so the string a
test asserts on is byte-identical to what fleetd.out receives. No behaviour change:
same text, same two branches, same call site.
- New FleetdUsageLimitFixWarningTest pins the leading "usage-limit fix:" grep tag and
three specific facts (profile name, model name, "enabled: false" under
models.allow) rather than the whole sentence, plus the no-model fallback's profile
name and cooldown-seconds substitution and its explicit absence of "enabled: false".
- New FleetProfilesQuarantineModelReasonFieldsTest exercises FleetMcp.QuarantineSource
with a modelFor/reasonFor that actually return values (every existing test used
QuarantineSource.none() or a 3-arg form defaulting both to null), asserting the
quarantined row's model/reason are present when supplied and absent (not null, not
blank) when modelFor/reasonFor return null or a blank string.
- Each of the three: implemented, broken by hand (row.put deleted / tag renamed),
confirmed the new test goes red, restored, confirmed green again. See PR body and
this ticket's fleet_reply for the verbatim failure output of all three.
Doc-only, one file, 13 added lines, inside §Project addendum. Verified by me:
- Canonical block byte-identical: 17468 chars on main and on the branch.
- The sync script in CLAUDE.md prints "in sync: True" both before and after.
- The note's own re-measure command works. Run verbatim, it returns exactly the 4 files
the note names, and no others:
fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java
fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java
fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java
fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java
Control that the search reaches the tree: 120 test java files, 63 of them mention herdr.
That control matters here — a zero-match re-measure command would tell a future session to
delete a live restriction.
- All four required elements are present: the carve-out, what stays banned (the control plane),
who it applies to, and the perishable half (dated 2026-09-10, the command, what each outcome
means, and delete-when-stale).
- Cross-references #458 for the canonical restatement, which is deliberately not in this change.
Pre-send check applied to the note itself: a worker assigned to AgentControlContractTest can now
name one legal action that finishes its task — let the test open the herdr socket in a throwaway
workspace it tears down.
The coordinator row is lead-to-lead coordination state (coord-ids, mailbox
facts, held-message previews). fleet_list returned it to every caller,
including a worker or an architect, because coordinatorView() had no way to
know who was asking.
Gate at the call site inside listFleet: a new overload takes
callerIsPrimary and only assembles/attaches the coordinator row when it is
true, so the key is absent (not empty) for a worker or an architect. The
MCP handler now passes principal(exchange).isPrimary(); every other
listFleet overload keeps passing true, so callers with no caller identity
(existing unit tests, the no-op wrappers) are unaffected -- confirmed by a
byte-for-byte comparison test against the pre-fix overload.
Authz's READ case is untouched: it stays shared by fleet_status,
fleet_profiles and fleet_whoami, and the gate here is purely inside
fleet_list's own result assembly.
#456's first paragraph said "HerdrPeerLauncher's own {@link
#spawn(SpawnRequest, PlacementDecision)} javadoc". That link resolves to this
interface's own abstract declaration, not to HerdrPeerLauncher's override, so a
reader who follows it lands on the wrong text. The second paragraph already
used the plain {@code HerdrPeerLauncher.spawn(...)} form; both now match.
Also says who is actually forced to read the paragraph, because #456's
reasoning rests on it and the two cases differ. A class that implements this
interface directly must write a body for spawn(SpawnRequest,
PlacementDecision) - it is abstract here - so it reads this javadoc. A
subclass of HerdrPeerLauncher does not: HerdrPeerLauncher already implements
that method (member/HerdrPeerLauncher.java:611) and the subclass inherits the
body. For a subclass the paragraph is advice, not a gate.
javadoc -Ddoclint=reference: 5 "reference not found", the same 5 in the same 5
untouched files as origin/main at 2af13ab, and none in PeerLauncher.java. Those
5 are ticket #459. mvn compile rc=0.
Doc-only, one file, 17 added lines. Verified by me on a local merge of 0788d84
onto 2af13ab (merge commit 5b46538):
- mvn -B clean test: Tests run: 1578, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS, rc=0.
- javadoc reference lint (mvn javadoc:javadoc -Ddoclint=reference): the merge has 5
"reference not found" errors; origin/main at 2af13ab has the same 5, in the same 5
untouched files. So this PR adds no broken link. Both runs rc=1 for that pre-existing
reason, which is now ticket #459.
- The quoted phrase is real: "are unoverridden here and just wrap {@link #defaultProfile()}"
is at member/HerdrPeerLauncher.java:604. place() is not overloaded, so {@link #place}
is unambiguous.
One inaccuracy I am fixing in a follow-up commit rather than sending the PR back: the first
new paragraph writes "HerdrPeerLauncher's own {@link #spawn(SpawnRequest, PlacementDecision)}
javadoc", but that link resolves to PeerLauncher's own abstract declaration, not to
HerdrPeerLauncher's override. The second paragraph already gets this right with plain
{@code HerdrPeerLauncher.spawn(...)}.
The model gate can be turned off at runtime (models.allow[].enabled: false, hot
since fleetd #422), but the usage-limit detector it's meant to react to was
compiled once at Fleetd.main startup into a frozen Map<String,Pattern> — arming
or disarming exhaustedPattern needed a daemon restart. Backwards for a feature
meant to react live.
- New LiveExhaustedPatterns: reads exhaustedPattern off the live config supplier
per lookup (matching CompositePeerLauncher#models0's live-supplier pattern),
caching compiled Pattern objects by PROFILE NAME (not pattern text — pattern
text would grow unboundedly as an operator tunes a regex across reloads;
profile names are bounded by the small, restart-gated set of configured
profiles). patternFor() backs CompletionResolver's classification; armed()
backs fleet_profiles' exhaustionDetectionArmed — both read the same object,
the fleetd #404 single-accessor rule CompositePeerLauncher.modelGateState()
established for the model gate.
- Fleetd.java: on BACKEND_EXHAUSTED, log a WARNING naming the profile's model
and the exact fix (enabled: false under models.allow, hot, no restart; remove
it again once the window resets) — or, when the profile has no model:
configured, say quarantine is the only thing keeping spawns off it.
- FleetMcp.QuarantineSource gains modelFor/reasonFor; fleet_profiles'
quarantined rows gain model/reason fields so a lead can see why without
reading the daemon log. capacityView (fleet_list) intentionally untouched —
scoped to fleet_profiles only.
- ConfigRef/FleetConfig docs + fleetd.example.yaml updated: exhaustedPattern
moves from Deferred to Hot. errorPattern stays deferred on purpose (out of
scope for this ticket).
- Tests: LiveExhaustedPatternsTest (new, unit-level hotness/caching proof),
FleetdExhaustionDetectionArmedWiringTest (rewritten — same QuarantineSource
object, read before and after a reload, asserts the answer flips with no
restart), ConfigRefTest/ConfigRefProfileCoverageTest updated for the new
Hot/Deferred classification.
The previous comment named a cause with no cell behind it. fleet01's review
made the point: a confident wrong mechanism gets copied, and a confident
under-determined one gets copied the same way.
So the comment now names the cell that settles it. Hold the old 1000ms write
sleep and change only the read - 800ms fixed sleep becomes a 5s poll - and the
test goes 0 of 3 to 3 of 3. The read deadline was the whole story.
It also says where: a 12-core macOS host near idle (load 2.6 to 5.9). The
loaded run agreed, but its load climbed from 7 to 50 while the cells ran and
the old version ran last, so it is not clean evidence and the comment says so.
Comment only. No test or production code changed.
Decision: leave both as default methods (option 1), not abstract. Neither
default is a live defect today — HerdrPeerLauncher is the sole single-profile
implementer and the degenerate answer (ignore role, always defaultProfile())
is correct for it. Making them abstract would force ~10 boilerplate one-line
overrides across 5 unrelated PeerLauncher test doubles (NeverSpawnsLauncher,
RaceLauncher, NoResumeLauncher, ClearContextSpyLauncher, LazyIdLauncher) that
never call either method, for a risk that is speculative (no multi-profile
HerdrPeerLauncher subclass exists or is planned).
Strengthens both javadocs with an explicit MUST-override warning and cross-
references HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)'s existing
#450 javadoc, which already names both methods as "unoverridden here" and
ties that to being a single-profile adapter -- the concrete place a future
multi-profile launcher author would read, since #450 made that method
abstract and any subclass must write its body.
The polling fix that landed in #452 is right, but its javadoc named a
mechanism nobody measured: that input typed before the shell's prompt was
swallowed by the shell's own startup.
I mutated the settle poll away — SHELL_READY_TIMEOUT_MS = 0, so input is
typed at once with no wait — and the test passed 3 of 3. So waitForText is
the load-bearing half, and the proven cause is the old 800ms READ deadline,
not the 1000ms write delay.
The direction is the point: typing at 0ms works where typing at 1000ms
failed. If early input were swallowed, 0ms would be worse than 1000ms. It is
better, so the swallow explanation is unsupported.
waitUntilSettled stays as cheap insurance, now labelled as insurance rather
than as the fix. Comment-only; AgentControlContractTest still green.
Verified on a locally built merge onto c11ad71 (the PR was branched from 822327e, before #448):
1578 unit tests green, then `-Dgroups=contract` gives 30 tests / 0 failures / 0 skipped here.
CI's own run of the tag on d4f93a7: 29 tests, 0 failures, 6 skipped — all 6 herdr tests skip on
the runner, and LeadMailboxTest's 15 tests run there for the first time.
Two mutations killed: restoring `assertEquals(14` fails the test, and pointing the expected env
value at a wrong URL fails it with the real pane text in the message.
Verified by the lead on head 5289eb5, and then again on the MERGE.
Battery on the branch tip:
FULL BUILD Tests run: 1577, Failures: 0, Errors: 0, Skipped: 0 BUILD SUCCESS
compile errors: 0
CONTROL contract test, real broker: Tests run: 9, Failures: 0, Skipped: 0
M3a AmqpReplyInbox.ack's `h == null` branch reports true
-> KILLED AmqpReplyInboxContractTest.ackReportsHitVsMissAgainstARealBroker
M3b AmqpReplyInbox.ack's `perTarget == null` branch reports true
-> KILLED same method
M3a is the gap I found in round 1: reporting true for something never held
survived both the default suite and -Pcontract against a real broker. It is now
closed in the adapter this daemon actually runs.
Both branches die, and they die through DIFFERENT assertions, which the worker
worked out and I confirmed by reading the code. held.get(target) is populated by
the deliver callback, not by own(), so after own() with nothing ever delivered
the map entry is still null — the "never held" assertion therefore exercises
`perTarget == null`, and only the double-ack assertion reaches `h == null`. Both
assertions are load-bearing; neither is redundant.
This PR was pushed before #451 landed, so the battery above tested the branch,
not the result. Merging main (bdcf285) in gave 0 conflicts and #448 adds no
PeerLauncher implementer, but a clean auto-merge is not a compiling merge, so I
built the merged tree 0829542:
FULL BUILD Tests run: 1578, Failures: 0 BUILD SUCCESS, 0 compile errors
contract, real broker Tests run: 9, Failures: 0
CompositePeerLauncherTest (#447's guarantee) Tests run: 76, Failures: 0
The worker also declined to point the contract test at the shared local LavinMQ
broker, because this adapter never deletes queues and a run would leave orphaned
durable queues on the instance backing the live fleet. It started a disposable
rabbitmq:3.13-management container on a throwaway port instead, then removed it.
That was its own judgment and it was correct.
Held peer mail is deliberately NOT ackable: fleet_ack against a coord-id errors
and names fleet_poll{coordId}. See #437 for why refusing is the right answer
today, and why that is a policy choice rather than a structural one.
Verified by the lead on head cfebc57 (base 822327e is current main).
FULL BUILD Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0 BUILD SUCCESS
compile errors: 0
M1 delete HerdrPeerLauncher's new override
-> BUILD FAILURE, 1 COMPILATION ERROR block, naming exactly
ClaudeCodeLauncher.java:[49,14] and OpenCodeLauncher.java:[59,14]
M2 CONTROL for M1: put the interface method back to a `default` AND delete
the override
-> BUILD SUCCESS, 0 compile errors
This is the row that makes M1 mean something. Without it, M1 only shows
that the build broke; with it, the break is attributable to the method
being abstract rather than to anything else the edit disturbed.
M3 CompositePeerLauncher's override reverts to the re-entering form
-> KILLED, Errors: 1
CompositePeerLauncherTest
.spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace
So #447's guarantee survives this refactor of the interface it rests on.
Tree restored clean after each mutation (git status --porcelain empty).
The worker corrected my ticket, and it was right. My #450 body listed five
src/main implementers of PeerLauncher and quoted `grep -rln 'implements
PeerLauncher'` as the source; that command returns two files. I had run a wider
pattern that also matched a comment in ConfigRef and the `extends
HerdrPeerLauncher` line in two subclasses, then quoted the narrow command beside
the wide command's output. Ground truth: ConfigRef implements
Supplier<FleetConfig>; ClaudeCodeLauncher and OpenCodeLauncher extend
HerdrPeerLauncher. So one override in that parent serves both, which is what the
worker built. The ticket body is corrected.
Out-of-scope note carried forward from the worker: defaultProfileFor(MemberRole)
and place(MemberRole) are two more default methods with the same shape. Noted,
not fixed here.
- HerdrClient.java, HerdrCodec.java, HerdrContractTest.java: the herdr port to
protocol 19 (CB-521) left the client javadoc and the contract test's own
assertion still saying protocol 14 / herdr 0.7.0. Updated to 19 / 0.8.0 and
renamed pingReturnsProtocol14 -> pingReturnsProtocol19. Verified the
assertion is real by temporarily changing the expected value to 20 (fails),
then restoring 19 (passes).
- AgentControlContractTest.java: tabCreateInjectsEnvIntoTheSeedShell was
failing, not skipping, on a host with a live herdr socket. Diagnosed with a
temporary instrumented run (not committed) that polled the pane every
200ms before and after sending input: the seed shell reliably takes ~2.5s
to reach its prompt (measured 3x), while the test's fixed 1000ms sleep
raced that startup. Input typed too early was swallowed by the shell's own
startup, leaving the typed line followed by the "Restored session" banner
and no command output — indistinguishable at a glance from the env map
never reaching the shell. Once the shell was actually ready, the injected
env value showed up in ~200ms, ruling out an env-seam defect. Replaced both
fixed sleeps with bounded polling on the actual conditions (pane text
settling, then the expected output appearing). Ran the fixed test 3x
standalone, all green.
- .gitea/workflows/ci.yml: the "Contract tests" step ran exactly one class by
name (-Dtest=AmqpReplyInboxContractTest), silently excluding every other
@Tag("contract") test from CI including the herdr ones above -- which is
how the stale protocol 14 assertion went unnoticed. Changed to
-Dgroups=contract, which selects the whole tagged group and picks up
future contract tests automatically.
The default re-entered the single-argument spawn(SpawnRequest), which re-runs
checks that can refuse the profile place() just chose (#444's window). Only
CompositePeerLauncher overrode it; a future placement-doing launcher could
have inherited the wrong body silently.
Give every current implementer an explicit override, chosen by what it does:
- HerdrPeerLauncher (base of ClaudeCodeLauncher/OpenCodeLauncher, neither of
which overrides spawn(req) or place()) does no placement filtering of its
own, so it gets the re-entering form.
- CompositePeerLauncher's routing-form override is untouched.
- 5 test-fake PeerLauncher implementers (SessionManagerTest, FleetdBackendErrorSinkTest)
get overrides matching their existing spawn(SpawnRequest) shape: delegating
wrappers delegate, unreachable stubs throw, the single-profile fake re-enters.
ConfigRef does not implement PeerLauncher at all (confirmed in this tree at
822327e) despite the ticket listing it as a src/main implementer.
Verified by the lead on the exact tree that lands (head e4c703a, base 82fae94 is
an ancestor, so this is the tree I measured):
FULL BUILD Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0 BUILD SUCCESS
compile errors: 0
CONTROL CompositePeerLauncherTest Tests run: 76, Failures: 0 -> GREEN
M1 the 2-arg spawn re-enters the 1-arg spawn (the inherited default this
ticket forbids for a multi-profile launcher)
-> KILLED Errors: 1
CompositePeerLauncherTest
.spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace
M2 drop the stamping: route to the decided profile but do not carry it
(SpawnRequest routedReq = req)
-> KILLED Failures: 1
same test method
Both mutations proved applied by printing the mutated method, and the tree was
restored clean after each (git status --porcelain empty).
Round 1 of this PR had a fixture weakness I found by mutation: StubLauncher's own
fallback default was "sol", the same profile place() decides, so an unstamped
request landed on spawnCount("sol") by coincidence and M2 survived. e4c703a gives
the adapter "b" as its fallback instead. One fixture now kills both mutations.
src/main is javadoc-only in this PR: 12 added lines, 0 added code lines, measured
by filtering the main-side diff.
AmqpReplyInboxContractTest is the one contract-group class CI actually
runs, and it never asserted on ack()'s return value at all — so the
exact defect this ticket fixes (reporting success for an ack that
removed nothing) was unpinned in the adapter fleetd runs live.
Add ackReportsHitVsMissAgainstARealBroker: a msgId never held for an
owned target returns false without throwing, a real held reply returns
true and is removed, and acking the same msgId again returns false.
Ran against both broker modes the class supports: Testcontainers
(AMQP_URI unset) and an external broker via AMQP_URI (the CI shape,
using a disposable container — not the shared local LavinMQ instance).
Review found the fixture's StubLauncher fell back to 'sol' too — the
same profile place() decides — so an UNSTAMPED request could land on
spawnCount('sol') by coincidence, and the assertion's claim that the
request 'actually carried sol' was unproven. Dropping the stamping
(SpawnRequest routedReq = req) while keeping the routing survived the
test unchanged.
Fix: give the adapter 'b' as its own fallback default instead, so an
unstamped request counts against 'b', not 'sol'. Verified both
mutations against the single test in isolation:
- drop-stamping (routedReq = req): RED, expected <sol> but was <b>
- re-entering (return spawn(req.withProfile(decision.profile())))
i.e. M1 from the first round: still RED, PlacementException
naming the now-quarantined 'sol'
Restored both; full suite green at 1575 tests.
No changes to src/main — PeerLauncher's javadoc from the first round
is unchanged.
ReplyInbox.ack now returns boolean (true = removed, false = nothing to
remove) instead of void, so FleetMcp.ack can finally tell a hit from a
miss. FleetMcp.ack returns an error when the boolean is false, naming
fleet_poll{coordId} for held peer mail, which has no route through this
call. MessageService.ackReply propagates the boolean; drainReplies keeps
ignoring it (its own javadoc already documents that loss window as
deliberate). Updated the tool schema's target description to match.
Rewrote FleetMcpTest's ack tests to publish a real message before
asserting success, and added tests for a never-queued id and a coord-id
target, both now erroring. Added boolean assertions to
InMemoryReplyInboxTest's existing ack cases.
Add a test that resolves place(role) while nothing is quarantined,
then quarantines the resolved profile's credential BEFORE spawning
against the held PlacementDecision. CompositePeerLauncher.spawn(req,
decision) must still honor the decision and land on the quarantined
profile, since it never re-runs the explicit-profile enforce* checks.
Verified the test kills the regression: with the override's body
replaced by the re-entering spawn(req.withProfile(...)) form, this
exact test goes RED with a PlacementException naming the now-
quarantined profile; restored, the full suite is green (1575 tests).
Also documents on PeerLauncher's default spawn(req, decision) that a
launcher routing across more than one profile MUST override it,
naming the four enforce* checks the default's re-entry re-applies.
Test written by a worker whose backend died before it could report; evidence
re-run by the lead against the merged tree.
Verified: merge of current main clean (0 conflicts); full build 1574 tests,
0 failures, BUILD SUCCESS, 0 compile errors; control green; deleting each of
reportGitHostShape, reportMemberTrustModel, reportMemberCredentialsGap and
reportExhaustedPatternGap from main() is KILLED by
mainReportsEveryStartupGapBeforeValidationAborts.
The new test never names List — only ListAppender, which has its own
import. An unused import is an IDE warning, and this repo treats
warnings as gates. No behaviour change: FleetdStartupReportTest still
runs 1 test, 0 failures, BUILD SUCCESS, 0 compile errors.
Found by the fleet01 lead reviewing #438 after I had merged it. Verified
independently before merging.
The implementation choice is the load-bearing part: LeadMailbox.own() now
assigns queueDeclare's durable flag and basicConsume's autoAck flag to named
locals, passes those SAME locals into the two real AMQP calls (:203/:204), and
derives heldDurable from them (:207). So the reported fact cannot drift from a
duplicate constant - a mutation to either call's argument moves the behaviour
and the report together. LeadChannel.heldDurable() is abstract, so a future
implementer gets a compile error rather than a silent default.
My own battery, merged tree, control green, tree restored clean:
- M1 revert to the literal true -> KILLED by
FleetMcpTest.listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable
- M3 heldCount forced to 0 (a half the worker did not touch) -> KILLED by
FleetMcpTest.listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero
- M2 break the derivation itself -> SURVIVED under plain clean install, exactly
as the worker reported. LeadMailbox needs a real broker, so the only test that
reaches the real queueDeclare/basicConsume is @Tag("contract"), excluded from
the default build. The worker ran that arm with -Pcontract and got 3 reds
including its own new test. Pre-existing structural limit of this class, not
introduced here, and the worker flagged it rather than hiding it.
Full build, my own run: Tests run: 1573, Failures: 0, Errors: 0, Skipped: 0 -
BUILD SUCCESS, 0 compile errors.
Not blocking, noted for a possible follow-up: one boolean over two independent
facts cannot say WHICH fact was lost. The merged field is still strictly better
than the literal it replaces, because it can now go false at all.
Round 4 pins the fix. Verified independently before merging.
My own battery in the worker's tree (control green, tree restored clean):
- M1 revert SessionManager:677 to launcher.spawn(spawnReq) -> KILLED by
SessionManagerTest.acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedThe
WorktreeForUnderARotatingPolicy. This is the ticket's own deliverable and it
survived 186 tests in round 3.
- M3 drop the withProfile stamping in the 2-arg spawn -> KILLED by 3 tests.
- M2 make the 2-arg spawn re-enter the refusing branch -> SURVIVED, but it is a
near-equivalent mutant, not a gap in this work. Post-#435 all four enforce*
conditions are ones the routing branch already filtered on, so the two paths
differ only if placement state moves between place() and spawn(). Filed
separately.
Full build, my own run this turn, whole log redirected and grepped:
Tests run: 1572, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, 0 compile
errors. Branch already contains current main.
Read the src/main diff. The new test uses roundRobin() (stateful) and asserts
agreement between the overlay profile and the spawned profile, rather than a
hardcoded name, which is the right shape - select() is stateful, so two calls
disagree by design.
FleetMcp.coordinatorView wrote heldDurable as a literal true, so a change
that broke either the durable queue declare or the manual-ack consume in
LeadMailbox would leave the field, and the full suite, green.
- LeadChannel gets a new heldDurable() method: the conclusion of a durable
queue declare AND a manual-ack consumer, derived by the implementation
from what it actually did, never asserted.
- LeadMailbox.own() captures the exact booleans it passes to
queueDeclare/basicConsume and stores their conjunction; heldDurable()
returns it.
- FleetMcp.coordinatorView now reads channel.heldDurable() instead of a
literal; updated the javadoc to say where the fact comes from.
- FakeLeadChannel gets a heldDurable field (default true) + withHeldDurable
setter so FleetMcpTest can prove the field goes false.
- FleetMcpTest: new test asserts heldDurable:false when the channel says so.
- LeadMailboxTest (contract, real broker): new test asserts heldDurable()
true against a real LeadMailbox. Verified by hand that flipping own()'s
autoAck local to true turns this test (and two pre-existing redelivery
tests) red, and restoring it turns them green again.
SessionManager.acquireWithWorktree's unqualified branch must carry the
PlacementDecision it already resolved via launcher.place() into
launcher.spawn(spawnReq, decision) rather than re-deriving it through a
blank-profile launcher.spawn(spawnReq). Every existing test in this file uses
PlacementPolicies.fixed(), which answers select() the same way on every call,
so dropping the decision (handle = launcher.spawn(spawnReq);) was invisible:
186 tests stayed green under that mutation.
acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy
uses PlacementPolicies.roundRobin() instead — deterministic AND stateful, so
two select() calls on the same policy instance disagree (index 0 then index 1
across a two-profile pool). It asserts AGREEMENT between the profile the
worktree's parity overlay was provisioned for and the profile the member
actually spawned on, never a hardcoded expected profile name.
Verified as a real mutation, not a no-op: applying the exact mutation
(handle = launcher.spawn(spawnReq);) turns it red — expected [b.mcp.json]
but was [a.mcp.json] — and reverting turns it green again. Full build:
1572 tests, 0 failures, 0 errors, BUILD SUCCESS.
fleet_poll{coordId} peeks this daemon's own held lead-to-lead mail and
returns full bodies without acking. New Authz.Action COORD_READ, primary
only — not the architect, which holds READ today. pollAction is now
argument-derived over both target and coordId.
heldView and HELD_PREVIEW_MAX_CHARS are untouched: fleet_list stays a
cheap always-safe scan, and the full read is a separately authorized call.
Verified by me, not taken from the report: merged tree builds 1562 green
(main 1555 + 7 new methods), 0 compile errors, no merge conflicts.
Mutation battery on lines the worker did NOT mutate, control 111 green:
- COORD_READ widened to the architect KILLED (AuthzTest + FleetMcpAuthzTest)
- self-coord-id guard removed KILLED (FleetMcpTest)
- full body swapped for the preview KILLED (FleetMcpTest)
The third mutation targets the ticket's own deliverable, and it is pinned.
Authz.permits has no default, so a new action is a compile error rather
than a silently unhandled case.
Follow-up filed as #439: fleet_list's coordinator row is still READ-gated,
so a worker sees peer coord-ids and 80-char previews of lead-to-lead
bodies. Pre-existing; the implementer flagged it and left it alone.
fleetd #435 (merged to main) made FixedPlacementPolicy evaluate maxLoad during automatic
selection, the same way weighted/round-robin already did. Six comments across
CompositePeerLauncher.java, PeerLauncher.java, PlacementDecision.java, and SessionManager.java
justified round 2's place()/spawn(req, decision) mechanism by saying fixed "deliberately never
evaluates maxLoad" — that claim is now false, and needed restating, not just deleting.
The honest case after #435: the two-path shape (routing branch falls through an excluded
candidate; explicit-profile branch refuses on it) is still real and still deliberate — an
operator who names a profile should get a refusal, not a silent substitution. What round 1 got
wrong, and what round 2 still needs to prevent, is turning a fall-through into a refusal by
accident: resolving a name via place() and then feeding it back to spawn(SpawnRequest) as an
explicit profile. Before #435 that accident was reachable through maxLoad specifically, because
fixed never evaluated it; #435 closed that specific gap, so a PlacementDecision can no longer be
at-cap in the first place. What survives as the justification for spawn(req, decision): it never
re-evaluates a condition place() already decided, and it closes the window between that decision
and the spawn in which the underlying state could otherwise move — not a failure #435 already
prevents.
Re-measured the sibling paths this round exists to keep in agreement (one profile at maxLoad: 1,
liveCount pinned at 1, PlacementPolicies.fixed(), unqualified spawn): both the with-worktree and
without-worktree paths now throw the identical PlacementException — "worker profile 'a' is at
maxLoad (1 live >= 1 cap), and no available candidate remains" — closed upstream by #435, at
place()/select(), before either path ever reaches a spawn call. The observable asymmetry this PR
was filed to fix is gone; what remains is the structural argument above.
No behavior change: place()/PlacementDecision/spawn(req, decision) are untouched, and
FixedPlacementPolicy/PlacementPolicyUtil are taken wholesale from main's merge.
fleet_list truncated held lead-to-lead messages to an 80-char preview with
no way to read the full body, and fleet_poll{target} drained the wrong
inbox (a worker's reply queue, not the coordinator mailbox) -- it silently
returned []. fleet_ack would have destroyed the message unread.
Add a non-destructive read: fleet_poll{coordId} peeks (never acks) this
daemon's own held mail via LeadChannel.peek(). The coordId must equal the
caller's own selfCoordId -- passing a peer's id is refused with a reason,
instead of repeating the original silent-[] confusion.
This is authorization-sensitive: mapping it to the existing READ action
would let any worker read every peer lead's mail in full. READ's openness
rests on "the roster carries no secrets" (Authz.java), which does not hold
for lead-to-lead coordination bodies. Added Authz.Action.COORD_READ,
primary-only (not even the architect, which holds READ today), and made
pollAction's signature depend on both target and coordId so every call
site states explicitly what it passes.
Also fixes fleet_list's "pending: 0" trap: mailbox.pending only counts
broker-ready messages, so a healthy held mailbox reads as empty. Added
heldCount/heldDurable beside held[] so the durability fact isn't implied
only by reading the code.
Mutation-tested: pollAction's COORD_READ->READ mapping, the peek->ack
substitution, the 80-char preview cap widened to 81, and the Authz case
widened to include caller.isWorker() -- each breaks exactly its matching
test and nothing else. The first attempt at the preview-cap test used a
homogeneous "x"*200 body, which a widened cap slipped through unnoticed
(contains() found a shifted match); replaced with a sentinel character at
index 80 to actually pin the boundary.
Updates CLAUDE.md's intent->tool table for fleet_poll's new coordId
semantics, per this repo's own "prompt is part of the product" rule.
wiki/ is a submodule and not committable from a worker's worktree --
wiki-bound content is in the PR body instead.
fixed was the one automatic policy that ignored maxLoad, and it is the
default for an absent placement: key. It now gates on the same shared
atCap predicate weighted and round-robin use, and falls through to the
next candidate rather than refusing.
Verified here: 1555 tests green (main was 1548, +7 new methods), 0
compile errors. Mutation battery, control 113 green:
- atCap boundary >= -> > KILLED (15 tests, across all 3 policies)
- drop cap check in fixed walk KILLED (3 tests)
- weightExcluded -> false KILLED (2 pre-existing tests)
Neither live host is affected today: both run placement: weighted (Mac
fleetd.yaml:210, fleet01 fleetd.yaml:172), and weighted already skipped
at-cap candidates via PlacementPolicyUtil.available. This aligns fixed
with the other policies and with maxLoad's documented contract.
FixedPlacementPolicy (the default placement policy) never consulted
maxLoad, so an at-cap default was chosen anyway on every unqualified
spawn -- the cap was advisory, not enforced, for the one policy every
config uses by default. weighted/round-robin already gated on it via
PlacementPolicyUtil.available().
Extract the "at cap" predicate into PlacementPolicyUtil.atCap(ctx, c)
so all three policies share one definition, and consult it at both of
FixedPlacementPolicy's filter sites (the default fast path and the
candidate walk), mirroring the existing weightExcluded pattern. An
at-cap default now falls through to the next candidate instead of
refusing the spawn -- only when every candidate is unusable does the
policy still throw, naming the cap in the message. Update the class
javadoc (five exceptions -> six) and the reason-priority comments to
match CompositePeerLauncher's explicit-spawn order (quarantine,
cooling off, max load, model-off).
Round 1 closed quarantine/cool-off/model-off routing for acquireWithWorktree by resolving the
profile through routedProfileFor(role) and handing that name back to launcher.spawn(SpawnRequest)
as an EXPLICIT profile. That re-resolution has a cost the lead measured directly: naming a
profile explicitly makes CompositePeerLauncher.spawn take its THROWING branch (enforceMaxLoad
included), while the routing branch a blank spawn takes never calls enforceMaxLoad at all, and
FixedPlacementPolicy (the default) deliberately never evaluates maxLoad during automatic
selection. So an at-cap pool-first profile that placement itself would have picked for a plain
unqualified spawn could die at enforceMaxLoad one call later, purely because the worktree path's
route to the spawn passed through an explicit profile name — a new failure a worktree-less
unqualified spawn never hits.
This closes the two-path shape instead of moving it: PeerLauncher gains place(role), returning an
opaque PlacementDecision, and spawn(req, decision), which honors that decision through the SAME
routing branch a blank spawn uses — no enforce* check is newly applied. SessionManager.
acquireWithWorktree now keeps the PlacementDecision from place() and hands it to
spawn(req, decision) for an unqualified request, instead of re-resolving through an explicit
profile name. An explicitly-named profile is unaffected: it still goes through spawn(req) and its
throwing branch, exactly as before.
Also corrects the acquireWithWorktree comment's false claim that round 1 "loses nothing else" —
maxLoad was lost too, as a new hard failure, not a retry. The comment now names it explicitly.
Kept the four round-1 tests (still pass — routedProfileFor now just delegates to place()). Added
one class asserting the invariant itself: an unqualified spawn on a maxLoad-capped profile must
land the same outcome with and without a worktree, asserting on the pair rather than a hardcoded
direction, so it stays correct however fleetd #435 (not this ticket) resolves whether maxLoad
should gate an unqualified spawn at all.
Verified in the worker's tree at 7fd914d: 1544 tests green (base 1535 + 9),
0 compile errors, unpiped mvn clean install. merge-tree against f8b0d42 reports
no conflicts and the two sides share no files.
Read all four production diffs. The design is right: one ModelGateState record
carrying both "armed" and "off" from a single models0() read, so the startup
log line, fleet_profiles' modelGateArmed and the spawn gate cannot disagree —
the fleetd #404 lesson applied properly. disabledModels() now delegates to it
rather than being a second independent read.
Mutated the two subtlest lines, with a control in the same script and the
changed line echoed back with its number:
- modelGateState(): sentinel identity check replaced by the naive
"m.offIds().isEmpty()" -> 2 failures. FleetProfilesModelGateStateTest
.modelsBlockWithNothingOffReportsGateArmedAndZeroOff:86 and
CompositePeerLauncherTest.modelGateStateIsHotReloadedThroughARealConfigRef
:1519, both "expected: <true> but was: <false>". So the one state this
ticket exists to expose — a block present with nothing off — is pinned.
- notConfigured() returning a non-empty off set, breaking the invariant the
record's javadoc states but does not enforce -> 3 failures, including
disabledModelsIsEmptyWithNoModelsConfigured:1581. So the invariant is
observable even though the constructor does not check it.
Control run unmutated: 1544 green.
Follow-up on me, not a merge blocker: fleet_profiles gains an operator-visible
field, so this needs a wiki/11-Features.md entry. Workers cannot commit the
wiki submodule, so I am adding it.
The javadoc said "what the spawn lifecycle reads". Nothing in src/main calls
profileForSlot at all, in either the ".profileForSlot(" or the
"::profileForSlot" form. My own #431 ticket text repeated that sentence as a
fact and ranked the three accessors by it, and the #432 worker copied it into
the test file's comment and one assertion message. Corrected in all three
places; the ticket correction is posted on #431.
What the spawn lifecycle actually reads for an architect's profile is the
SlotReservation that reserve() returns — SessionManager.java:225,
"reservation == null ? profile : reservation.profile()".
The corrected ranking, measured rather than read off the javadoc:
- nameForSlot is wired, at CallerResolver.java:137 (method reference, which is
why a ".nameForSlot(" grep missed it)
- isSlot is reached through bind, called at MemberRegistry.java:235 and :376
- profileForSlot has no caller at all
Prose only. No behaviour change.
Verified in my own tree at a196d34: 1539 tests green, 0 compile errors, 0 files
changed under src/main (test-only, as reported).
Mutated two halves the worker's own proof did not cover, with a control in the
same script and the changed line echoed back each time:
- roleForSlot returning ARCHITECT for any configured slot (flatten is
role-blind) -> 2 failures, incl. CallerResolverTest
.aBoundNonArchitectSlotStillResolvesAsAWorker:454. So the role-blind flatten
cannot grant ARCHITECT through a non-architect pool; that was already pinned.
- nameForSlot parsing the key suffix instead of reading config -> 1 failure,
MemberRegistryLiveTest.nameForSlotReflectsANameChangedByReload:255. The new
test has teeth beyond the freeze the worker ran.
Control run unmutated: green.
The ticket's severity ranking was wrong and I corrected it on #431. profileForSlot
has no caller in src/main in either the "." or "::" form, so its javadoc ("what
the spawn lifecycle reads") names a caller that does not exist; nameForSlot is
wired at CallerResolver.java:137; isSlot is reached through bind, at :235 and
:376. A follow-up commit fixes the test prose that repeated my claim.
PeerLauncher.disabledModels() reported an empty set both when no
models: block exists and when a block exists with nothing off, so
fleet_profiles/GET /profiles and the startup log could not tell an
inert gate from an armed one reporting zero. Add
PeerLauncher.ModelGateState (configured + off), a modelGateState()
default method disabledModels() now delegates to, and a
CompositePeerLauncher override that reads models0() once and
distinguishes the null-supplier case (no models: block) from a real,
config-supplied block via identity against the NO_MODELS_CONFIGURED
sentinel — reusing the exact accessor the spawn gate itself reads, per
the fleetd #404 lesson.
Wires the state into a new startup log line (Fleetd.modelGateCoverageLine)
and a new modelGateArmed field in FleetMcp.profilesView, reported
unconditionally alongside the existing modelsOff set.
e1d7dde (PR #430) kept two good fixes and one regressed one. Kept: (1)
CompositePeerLauncher.defaultProfile() delegating to
defaultProfileFor(MemberRole.DEV) so fleet_profiles' "default" tracks a live
reload, and (2) PeerLauncher.defaultProfileFor(MemberRole). Redone:
acquireWithWorktree's profile pre-resolution.
The regression: acquireWithWorktree pre-resolved via
launcher.defaultProfileFor(memberRole), which just returns the role pool's
FIRST entry, blind to quarantine/cool-off/model-off. That name was then
passed to launcher.spawn as an EXPLICIT profile, which takes
CompositePeerLauncher.spawn's THROWING branch (enforceNotQuarantined /
enforceMaxLoad / enforceModelEnabled) instead of the ROUTING branch a blank
profile gets. So a quarantined or model-off pool-first profile turned a
routine unqualified spawn into a hard PlacementException -- undermining
fleetd #429's "the fleet keeps working when a model is turned off"
guarantee for every worktree spawn.
Fix: add PeerLauncher.routedProfileFor(MemberRole), the profile an
unqualified spawn of that role would actually be routed to right now --
same candidate list, same quarantined/coolingOff/modelOff filtering, same
PlacementPolicy spawn() itself consults. CompositePeerLauncher implements it
by extracting spawn()'s context-building into a shared private
placementContextFor(role, unreachable), so spawn() and routedProfileFor()
can never disagree about which conditions apply to which candidate.
acquireWithWorktree now calls routedProfileFor once and reuses that name for
repoRoot, parityOverlay, and the spawn -- the fleetd #425 defect (the three
disagreeing) stays fixed, now on the routed path instead of the blind one.
An explicit profile named by the caller is untouched -- it still hits the
throwing branch, which is correct for an operator override.
Trade-off carried over from e1d7dde, now precisely scoped: an unqualified
worktree spawn still loses CompositePeerLauncher's cross-candidate retry on
a live PeerUnreachableException (a transport failure at spawn time, which
placement cannot see in advance) -- but NOT the quarantine/cool-off/
model-off routing, which routedProfileFor already resolved before spawn
ever runs. Accepted: a worktree provisioned for the wrong backend is worse
than a spawn that fails cleanly and can be retried by the caller.
Tests: CompositePeerLauncherTest gains
routedProfileForSkipsAQuarantinedPoolFirstProfileUnderFixedPolicy and
...ModelOff..., both under PlacementPolicies.fixed() (the default policy,
not weighted() -- the previous round's tests all used weighted() and never
exercised FixedPlacementPolicy's own inline filter, which is exactly what
regressed). SessionManagerTest gains
acquireWithWorktreeRoutesAroundAQuarantinedPoolFirstProfile, proving
repoRoot/parityOverlay/spawn agree on the ROUTED profile, not just the
pool-reordered-by-reload one the existing #425 tests already covered.
#424 made MemberRegistry.slots() re-read fleet: on every call, but only
roleForSlot was tested against a real reload. profileForSlot, isSlot and
nameForSlot all have the same live-read line and none was pinned — proved by
freezing each to a construction-time snapshot and watching the full suite
stay green.
Adds 4 tests to MemberRegistryLiveTest, each driving a real ConfigRef.reload()
against a @TempDir config file (never two frozen registries compared in
memory, which would test the constructor instead of the reload):
- profileForSlotReflectsAProfileChangedByReload
- isSlotStopsReportingASlotRemovedByReload / isSlotStartsReportingASlotAddedByReload
- nameForSlotReflectsANameChangedByReload
No production change. profileForSlot has no call site anywhere in src/main
yet, so there is no spawn-lifecycle seam to drive the test through beyond the
accessor itself.
fleet_profiles' "default" was CompositePeerLauncher.defaultProfile, a value
frozen at construction from cfg.effectiveDefaultProfile(). An unqualified
fleet_spawn instead resolves the dev pool live via defaultProfileFor(DEV) on
every call, so reordering fleet.developers and reloading changed where a
spawn landed without ever changing what fleet_profiles reported.
- CompositePeerLauncher.defaultProfile() now delegates to
defaultProfileFor(MemberRole.DEV) -- the same live, reload-aware pool read
placement already uses -- falling back to the frozen field only when no
profiles are configured at all.
- PeerLauncher gains a default defaultProfileFor(MemberRole) method so a
generic PeerLauncher reference can ask for a role's live default; the
default implementation delegates to defaultProfile() for launchers with no
pool concept of their own.
- SessionManager.acquireWithWorktree resolved a profile via
launcher.defaultProfile() (DEV-only) to provision repoRoot/parityOverlay,
then spawned with the original (possibly blank) profile, which re-resolves
independently through placement -- for any non-DEV role, or across a config
reload between the two reads, the two resolutions could disagree and
provision a worktree for a profile the member never runs on. Fixed by
resolving once, through defaultProfileFor(the caller's actual role), and
reusing that same resolved name for repoRoot, parityOverlay, and the spawn
itself. Trade-off: this path now spawns with an explicit profile rather
than a blank one, so it loses CompositePeerLauncher's cross-candidate retry
on PeerUnreachableException -- accepted because a worktree provisioned for
the wrong backend is worse than a spawn that fails cleanly and can be
retried.
Tests: CompositePeerLauncherTest (live dev-pool reorder + empty-pool
fallback), FleetProfilesLiveDefaultTest (drives FleetMcp.profilesView
directly), SessionManagerTest (worktree overlay follows a reorder, and a
non-DEV role's worktree spawn uses that role's pool, not DEV's).
fleetd #424. MemberRegistry.slots() now re-reads fleet: through a supplier,
so removing an architect slot demotes the bound pane on its very next request.
The boundEntries cache and entryFor() fallback from the first round are gone:
the slot OCCUPANCY (terminalToSlot) survives a reload, the ARCHITECT role does
not. That split was my own ticket wording's fault -- I asked for a test that a
bound architect "survives the rebuild", which conflated the binding with the
privilege.
Conflict resolved by hand in ConfigRef.java: #422 (models:) and #424
(architects) both rewrote the same Hot bullet. Kept both.
Also corrected two claims #424's own second commit left stale -- 7f672f0
reversed the behaviour but never touched ConfigRef, whose whole job is to tell
the operator what a reload does:
- the Hot bullet said MemberRegistry's rule "keeps a live session's identity
even after its slot is removed from config"
- the reload-report comment said "only a NEW bind is refused"
Both now say what the code does: removal revokes ARCHITECT on the next
request, and only the slot occupancy survives.
Verified by the lead: 1535 tests, 0 failures, 0 compile errors, BUILD SUCCESS
on the merged tree.
Mutation of three halves the worker's own proof did not cover -- profileForSlot,
nameForSlot and isSlot each pointed at a frozen snapshot taken at construction
(live readers 5 -> 4, each mutation naming its method and line). All three
PASSED at 1535. The ticket's own fix is well pinned; these three sibling live
reads are not. Follow-up filed.
fleetd #422 + follow-up. Verified by the lead: 1524 tests green, 0 compile errors.
Mutation proof of three halves the worker's own proof did not cover:
- modelOffProfiles() -> Set.of() (the feed into ctx.modelOff): kills 3, incl. fixedPlacementSkipsAnOffModelProfileToo
- enforceModelEnabled() -> no-op (the explicit-spawn gate): kills 3, incl. the real-ConfigRef hot-reload test
- disabledModels() -> Set.of() (the reporting accessor): kills 1, alone
Correction to the #424 fix in PR #428: the ticket asked to revoke a removed
architect slot, but the previous change (boundEntries) kept BOTH the binding
and the ARCHITECT privilege alive for an already-bound session after its slot
left config. That left the ticket's actual headline defect half-open.
The corrected rule: config governs both what may be bound next AND what a
bound slot still grants. Removing a slot now demotes its bound session to
worker on the very next request (roleForSlot/nameForSlot read slots() with no
cache, so CallerResolver.resolve falls through to Principal.worker(...)). The
terminalToSlot binding itself is untouched by a reload, on purpose: dropping
it would double-book the slot key and break unbind's compare-safe contract.
- Delete boundEntries and entryFor; profileForSlot/roleForSlot/nameForSlot/
isSlot all read slots() directly, live, with no cache.
- Rewrite the class doc's binding rule for the corrected semantic.
- Replace the old "survives removal" test with anArchitectAlreadyBoundToASlotIsDemotedByReload,
asserted through a real CallerResolver.resolve (not the roleForSlot seam),
plus two tests for what must NOT change: the binding still occupies the
slot after removal (a second terminal cannot claim it, even once the slot
returns to config), and unbind still succeeds for the original terminal.
Verified snapshot()/CallerResolver.members() need no change: snapshot() only
ever reported raw terminalToSlot occupancy, and CallerResolver.members() has
no production caller.
FixedPlacementPolicy is the DEFAULT placement policy (PlacementPolicies.fromName
returns it for an absent/blank name) and it built its own inline candidate
filter instead of calling PlacementPolicyUtil.available(). That filter checked
quarantined/coolingOff/unreachable/excluded() but never modelOff(), so an
unqualified fleet_spawn on any fleet without an explicit placement: policy
could still land on a profile whose model the operator turned off.
- Add ctx.modelOff() to both filter sites: the default-profile fast path and
the fallback walk over ctx.candidates().
- Add a modelOff refusal reason to the default-profile reasons list, worded as
an operator decision ("turned off in models.allow"), matching
enforceModelEnabled. Quarantine and cooling off still take priority when a
profile is also model-off, matching CompositePeerLauncher's explicit-spawn
check order.
- Update the two stale "excluded from automatic selection" messages to name
model-off, consistent with PlacementPolicyUtil.emptyException.
- Update the class javadoc: four exceptions -> five, with a new bullet for
model-off (fleetd #422).
Tests: PlacementPolicyTest gains fixedSkipsModelOffDefault (fast-path),
fixedFallbackWalkSkipsModelOffCandidate (fallback walk),
fixedThrowsWhenDefaultAndEveryCandidateModelOff (all-off refusal wording), and
fixedReportsQuarantineNotModelOffWhenBothApply (priority). CompositePeerLauncherTest
gains fixedPlacementSkipsAnOffModelProfileToo, an integration-level mirror of
the existing placementSkipsAnOffModelProfileAndRoutesToAnotherOne but under
PlacementPolicies.fixed(). The two existing weighted()-based tests are
untouched.
Ships the two halves left out of the earlier allow-list ticket in one PR,
since apart they are inert: a gate with no flag always allows, and a flag
nothing reads does nothing.
- FleetConfig.Models.ModelEntry gains `enabled` (default on; absent/true =
on, false = off). Turning a model off never removes it from `allow:` —
validateModels() checks membership only, so an off model stays valid
config and a still-configured profile naming it does not refuse reload.
Models.offIds() is the one live accessor both the gate and the status
report read.
- CompositePeerLauncher.enforceModelEnabled is a FOURTH, independent
spawn-refusal reason (operator intent) — never layered onto
BackendQuarantine/BackendOutagePolicy, which are backend-reported outage.
Wired into the explicit-profile branch. modelOffProfiles() feeds the same
off-model exclusion into PlacementContext for unqualified spawns via
PlacementPolicyUtil (a new modelOff set, counted into its own bucket in
emptyException so "all off" is named as the cause, not generic).
Both read models0(), a live Supplier<FleetConfig.Models>, so a reload
reaches the very next spawn — no restart.
- PeerLauncher.disabledModels() (default empty) lets fleet_profiles/
GET /profiles report off models by reading the exact same accessor the
gate reads (the fleetd #404 lesson: a status field must read the source
the behaviour reads).
- ConfigRef: `models` reclassified from deferred to hot-excluded — nothing
about it is baked into a startup-built object anymore; membership is
re-validated in full on every reload via validateAll(), and the on/off
half is read live everywhere. Tally: 5 cold, 13 deferred, 3 split, 4
hot-excluded (25 total). ConfigRefTopLevelCoverageTest and
ConfigRefTopLevelReportingCoverageTest updated with no new exclusion
added just to force green.
Tests: FleetConfigTest (old-style fixture stays on; an off model is still
valid config; one model disables every profile naming it),
CompositePeerLauncherTest (explicit refusal wording distinct from
quarantine/cool-off; unqualified spawn skips an off candidate and names
model-off when every candidate is off; a real ConfigRef.reload() proves
the hot path; disabledModels() matches the gate).
MemberRegistry used to flatten fleet.architects into an unmodifiable map at
construction, so removing (revoking) an architect slot from config never
took effect: reserve()/requireSlotFor() kept granting spawns against the
frozen snapshot forever, while ConfigRef told the operator "already
applied" for the wrong consumer.
- MemberRegistry gains a live constructor (MemberRegistry.live(Supplier))
that re-flattens fleet.architects/developers/reviewers on every
slots()/slotsFor() call, so reserve() and requireSlotFor() (which both
read through slotsFor) govern the NEXT spawn with no restart. The frozen
single-arg constructor is kept for tests and fixed/code-built configs.
- Binding rule: config governs what may be bound next; it never
retroactively unbinds a live session. A slot removed from config while a
terminal is bound to it keeps that binding. To keep the bound terminal's
IDENTITY too (CallerResolver.resolve reads roleForSlot/nameForSlot on
every request), every successful bind now caches the slot's Entry into a
new boundEntries map; roleForSlot/nameForSlot/profileForSlot/isSlot fall
back to it when the slot is no longer live, and unbind clears it in the
same critical section it clears the binding.
- Fleetd.java now wires MemberRegistry.live(() -> config.get().fleet())
instead of the frozen constructor.
- ConfigRef: corrected the fleet.leaders split-key message and the class
doc's Hot bullet — architects is now hot for two independent consumers
(CompositePeerLauncher for placement, MemberRegistry for identity), not
only the one the message used to name. An architects-only edit still
reports nothing beyond "config reloaded", which is now honest since the
key really is fully hot for both consumers.
Added MemberRegistryLiveTest: real ConfigRef.reload() against a @TempDir
file, both directions (slot removed / slot added) for requireSlotFor and
reserve tested separately, plus a bound-architect-survives-removal test
that checks the binding AND the identity (roleForSlot/nameForSlot).
Mutation-tested: reverting requireSlotFor to a frozen snapshot fails
requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload and its mirror;
reverting reserve the same way fails the two reserve tests; dropping the
boundEntries fallback fails the survives-removal test's roleForSlot
assertion. All three restored before commit.
mvn clean install: Tests run: 1510, Failures: 0, Errors: 0, Skipped: 0 —
BUILD SUCCESS.
Review found a gap: the earlier tests all called CompletionResolver.coverage()
directly, supplying the UnsetMeaning themselves — proving the enum's wording,
never that Fleetd's two call sites pair the right meaning with the right key.
Swapping the two UnsetMeaning arguments at those call sites (recreating #415's
defect with exhaustedPattern and errorPattern exchanged) compiled with 0 errors
and left all 1506 tests green.
Extract the two coverage-line call sites out of main() into package-private
static factories (Fleetd.exhaustedPatternCoverageLine /
errorPatternCoverageLine), the same pattern already used for capacitySource
and worktreeBranchLookup. Add FleetdPatternCoverageLineTest, which calls both
factories directly and asserts the actual wording each produces for the same
empty-coverage input, including that the two differ.
Also recorded the swap-mutation measurement (0 errors, 1506 green) in
UnsetMeaning's javadoc so a future reader does not delete the new test as
redundant with CompletionResolverTest.
CompletionResolver.coverage() measured pattern coverage (how many profiles set
a key) but its 'off' wording read as feature state. That is false for
errorPattern: an unset errorPattern still runs the classification against the
built-in BACKEND_ERROR pattern (CompletionResolver.java:84), so the empty case
is not off.
Add CompletionResolver.UnsetMeaning (OFF / BUILT_IN_DEFAULT), a required
parameter every coverage() call must supply — no defaulted overload, so a
future third pattern key cannot compile without stating what unset means for
it. Fleetd.java now passes UnsetMeaning.OFF for exhaustedPattern (no fallback
exists) and UnsetMeaning.BUILT_IN_DEFAULT for errorPattern.
Tests: updated the three existing empty/full/partial cases to pass the new
parameter, corrected the one test that pinned the old (wrong) errorPattern
wording, and added a test that asserts the same empty-coverage input produces
different wording for the two keys.
`profiles` is a DEFERRED key: HerdrPeerLauncher takes Map.copyOf(profiles) once at
construction, so a profile added only to the hot-reloaded map can never be spawned.
CapacitySource was built with `() -> config.get().profiles().keySet()` — the live map —
so fleet_list reported a hot-added profile as free while fleet_spawn on that same
profile failed with "unknown worker profile". fleet_list's contract for `free` says it
runs "the same check the spawn gate itself runs"; it did not.
The set now comes from the startup snapshot, the same shape as the coordinator.peers
wiring three lines below, whose comment already stated the rule. maxLoad stays live on
purpose — ConfigRef documents it as hot, like credentialId and weight — so a maxLoad
edit still takes effect without a restart.
Found by the fleet01 lead, who proved it with a live ghost profile rather than an
argument. Implemented by a sonnet member; its reply was lost to an empty scrape, so the
work was salvaged uncommitted from its worktree and both proof steps were run by the
lead instead:
full build Tests run: 1505, Failures: 0, 0 compile errors
mutation A: live keySet restored liveOnlyProfileIsNotListed FAILS (alone)
mutation B: permanently empty set startupProfileIsListed FAILS (alone)
Mutation B is the point of the second direction: per #404, a test that only ever checks
the absent case cannot tell a correct lookup from one that returns nothing at all.
anAskThatLeavesByThrowingStillClosesItsQuestion barriered on Phase.ASKING,
which markAsyncQuestion sets in ask()'s FIRST step. The assertion right
after it depends on ask()'s THIRD step (pushLoop.onQuestionOpened), which
is what actually populates ReplyPushLoop's pendingQuestions map. Under
load the asker thread can be descheduled between those two steps, so the
barrier released before decide() had anything to see, and it correctly
returned STOP instead of the expected INJECT.
Add ReplyPushLoop#pendingQuestionTurnIdsForTest, a package-private test
seam (modeled on MessageService#isCompletionStampedForTest) exposing the
private pendingQuestionTurnIdsFor. The test now waits for its own turnId
to appear there before asserting on decide() — not for decide() itself to
return INJECT, which would make the barrier assert nothing.
Checked every other awaitTicketPhaseOn(..., Phase.ASKING) in the file
(two, in the CB-582 nudge tests): both are followed by a real awaitNudge()
that waits for an actual agent.prompt push-loop call before any assertion
depends on push-loop state, so they are not exposed to this race.
No production behaviour changed.
Widens the completed-hook's real race window (normally instructions-wide, needing
~2x-core host load to hit by chance per #399) by injecting a bounded sleep into the
test clock's completion-stamp read. This makes the ordering invariant — a test must
wait for isCompletionStampedForTest, not just DONE, before advancing the clock past
the TTL — fail deterministically on the first run when the barrier is removed, and
pass deterministically with it present. No production code changed.
The armed lookup now uses the same compiled map detection uses. Second test added during review after a mutation proved the true direction was unpinned.
Mutation-tested during review: replacing the armed lambda with
`profile -> false` left the whole suite green at 1475 tests, because the
only existing test passes an EMPTY startup map. That mutation would make
#395's visibility feature silently dead.
With this test the same mutation fails, and it is the only test that
fails, so nothing else covers this direction.
The blanking loop classified a name as blanked from eval's exit status.
zsh coerces a bare NAME= assignment on an integer special parameter
(SECONDS, RANDOM, SHLVL, HISTSIZE, COLUMNS, LINES, USERNAME) to a number
instead of failing, so eval returned 0 with the value untouched. Measured
7 false receipts in 10 names. This is a security receipt, so a count that
overstates the scrub is worse than no count.
The loop now runs the eval unconditionally and decides from the observed
value, read back with the (P) indirection flag. One check covers all
three shapes a name can take: a real blank, a fatal error eval merely
contained, and this silent no-op. Exit status plays no part.
Lead review: mutation removing the '!' unblankable report line is CAUGHT
(2 failures in EnvAllowListScrubTest, both asserting the name is reported
rather than silently dropped). The eval-site identifier guard is
untouched.
Unplanted evidence the fix works: the #394 test
unblankableNameInTheMiddleDoesNotAbortNamesAfterIt began failing under
the fix, because zsh auto-exports SHLVL and the old exit-status bug had
been miscounting it as blanked all along. Its exact-count assertion had
only ever passed because of the bug beside it; it is now a presence
check, since which names a zsh version auto-exports is not this test's
to pin.
Still not pinned, tracked in #394's follow-up: the eval-site identifier
guard has no test behind it.
models: allow: is a single place that names every model the fleet may
use. Absent or empty keeps today's behaviour, so this ships inert until
configured. Once set, a profile naming a model outside the list refuses
to start, and refuses a reload, rather than reaching a backend adapter
as a free-form string.
The allow-list cannot be checked against a provider catalogue: for an
opencode profile fleetd SYNTHESIZES the provider from provider/model
plus baseUrl (OpenCodeLauncher:562-583), so a valid fleetd model id
appears in no published catalogue. An operator-owned list is therefore
the only workable gate.
Includes the #398 follow-up: FleetConfig.validateAll() reflectively
sweeps this class's validateXxx() methods, and both real call sites
(Fleetd.main and ConfigRef.reload) call that one method. Before this,
deleting a validateXxx() call from either caller left the whole suite
green. FleetdStartupValidationTest now drives the real Fleetd.main.
Lead review: mutation on the reload call site is caught (ConfigRefTest,
2 failures). Mutation replacing the reflective sweep with a hardcoded
list is NOT caught (1491 green) — so the sweep is a convenience and the
denominator test is the real guarantee; two false statements in the test
javadoc were corrected to say so (af4c88d).
Recovered work: the worker's agent died mid-turn with the follow-up
uncommitted and the startup call left disabled as
'// MUTATION-TEST-TEMP: cfg.validateAll();'. I restored it before
committing.
NOT covered: the five log-only reportXxx(cfg) calls in main are still
unpinned — filed as #407.
Measured at review: reverting validateAll() to a hardcoded list of
today's six calls leaves the suite green (1491 tests, 0 failures). The
class javadoc claimed that mutation fails a test. It does not — claim 1
pins the generic helper on an unrelated class, claim 2 pins today's six,
and a hardcoded list satisfies both.
The interaction was the real hazard. The denominator assertion IS a
tripwire (declaring a seventh validator fails it), but its failure
message said the sweep reaches new validators 'by construction' and told
the author to just update the expected set. If the sweep were ever
replaced by a name list, the one assertion that fires would hand back a
false all-clear at the moment it fired.
Javadoc now states the measurement, and names the denominator test as
the actual guarantee. The assertion message now says to confirm
validateAll() still delegates to invokeAllValidators(this) BEFORE
updating the expected set.
Mutation testing found that deleting a cfg.validateXxx() call from
Fleetd.main left the full suite green: every test called a validator
directly and none exercised main as the caller.
FleetConfig.validateAll() sweeps this class's own public no-arg void
validateXxx() methods by reflection and invokes each in alphabetical
order, so a newly written validator is wired into both callers
(Fleetd.main and ConfigRef.reload) with no second step to forget.
FleetdStartupValidationTest calls the real Fleetd.main with six configs,
each failing exactly one validator.
Recovered by the lead: the worker's agent died mid-turn with this work
uncommitted, and had left the startup call commented out as
'// MUTATION-TEST-TEMP: cfg.validateAll();' from its own mutation run.
I restored the call before committing. Build after restoring:
Tests run: 1491, Failures: 0, BUILD SUCCESS.
NOT covered, and not claimed to be: the five log-only reporters in
main (reportRequiredSecrets, reportGitHostShape, reportMemberTrustModel,
reportMemberCredentialsGap, and reportExhaustedPatternGap on current
main) are not validateXxx() methods, so the sweep does not reach them
and their call sites stay unpinned.
poll() can report Phase.DONE for a ticket before the whenComplete hook that
stamps Task.completedNanos has run — CompletableFuture.complete() publishes
its result and only then runs dependents. The TTL tests advanced an injected
clock right after observing DONE, so on a host where the hook runs late it
stamps the ADVANCED time and the eviction never happens (fails on Linux,
passes on macOS).
Add a package-private test seam, MessageService.isCompletionStampedForTest,
that reports whether completedNanos is stamped. Both TTL tests now wait on
that (a real volatile read/write happens-before edge) before advancing the
clock, instead of on Phase.DONE. The prune condition in pruneTerminalTickets
is untouched.
eval "export NAME=" can return success even when zsh coerces the bare
assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
HISTSIZE, COLUMNS, LINES, USERNAME) instead of failing, leaving the
value unchanged. The old exit-status check then reported the name as
blanked when it was not -- a false receipt.
Classify on the observed effect instead: attempt the export, then read
the name's value back with the (P) indirection flag and decide from
whether it is now empty. One check now covers all three shapes a name
can take here -- a genuine blank, a fatal read-only error eval merely
contains, and this silent no-op -- with the exit status playing no
part in the decision.
Adds a test driving all three shapes through the real scrubScript in
one run (a normal name, LINENO for the fatal case, SECONDS for the
silent no-op), with the parent environment explicitly carrying those
names since a cleared ProcessBuilder parent does not expose them on
its own. Also corrects the previously-merged
unblankableNameInTheMiddleDoesNotAbortNamesAfterIt test, whose "exactly
one failed name" assertion turned out to only pass by accident: zsh
itself auto-exports SHLVL on every shell start, and the old exit-status
bug was silently miscounting it as blanked. The fixed classification
now correctly reports it unblankable too, so the test asserts presence
rather than an exact count.
A profile with no exhaustedPattern has usage-limit detection silently
disabled. Startup now reports every unarmed profile (a louder warning for
subscription profiles, which are the ones a limit actually stops), and
fleet_profiles / GET /profiles carry exhaustionDetectionArmed per profile.
Reviewed by the lead: all three requested mutations fail a test, and the
back-compat QuarantineSource ctor defaults to 'not armed' when the source
is unknown. KNOWN GAP, not fixed here: deleting the
reportExhaustedPatternGap(cfg) call at Fleetd.java:140 leaves the suite
green (Tests run: 1472, Failures: 0). That is the same unpinned-startup-
call shape as PR #398's six validators, and fleetd #398's ticket owns it.
exhaustedPattern is opt-in per profile: unset means a usage-limit
refusal on that profile is never classified BACKEND_EXHAUSTED and
never quarantines its credential, with nothing telling the operator.
Add a startup WARN naming every unarmed profile (a louder, separate
WARN for a subscription: true profile, since that is the operator's
own metered plan). Surface the same fact per profile in fleet_profiles
as exhaustionDetectionArmed, so an operator can tell "healthy" from
"can never be caught" without reading fleetd.yaml.
The CB-633 allow-list scrub has been dying mid-loop on every fleet01 member
pane and saying nothing. `export UID=` in zsh is not a failed command — it is
a fatal parameter error that terminates the whole sourced file. The blanking
loop is wrapped in `{ ... } 2>/dev/null`, so the message was swallowed and the
report block after the loop never ran.
Root cause found by the fleet01 lead, with xtrace on a live pane's own ZDOTDIR:
+scrub.zsh:28> _cb633_n=UID
+scrub.zsh:28> export 'UID='
+zsh:1> rc=126 <- file aborted
The severity is the INVERSION, and this is their finding, quoted:
"env lists inherited names first and the names a startup file exports last.
So the loop blanks the harmless inherited half and dies immediately before
the operator's own exports — exactly the credentials the policy exists to
remove. The selection is inverted, not merely partial."
Measured there: UID is name 42 of 57, and a ~/.zshrc decoy at 58 survived on
8 of 8 spawns. "Partial scrub" reads as "we got most of it"; it got precisely
the wrong half.
Fixed with `eval "export ${n}=" 2>/dev/null` rather than a skip-list of the
known-fatal names (UID EUID GID EGID PPID LINENO). A skip-list has to be
complete forever and this is a security control; eval needs no list. Measured:
plain export dies at UID and every later name keeps its value, while the eval
form completes the loop and blanks all of them. PR #396 proposed the skip-list
and is closed in favour of this; its claim that the abort happens "however the
assignment is wrapped" holds for a direct `if ! export` but not for eval,
which reparses in a nested context.
The report now carries `failed N` and `!`-prefixed unblankable names, and
HerdrPeerLauncher WARNs when any name could not be blanked. The old "no report"
WARN no longer claims the daemon knows what the member saw.
Why the suite stayed green: EnvAllowListScrubTest starts zsh from
pb.environment().clear(), and under a cleared parent UID is not an exported
name at all, so the abort could not reproduce in that harness.
Reviewed by mutation, which found a second gap now also closed: the eval is
only safe because names are filtered to ^[A-Za-z_][A-Za-z0-9_]*$. Replacing
that pattern with .* left the class green, so the line the security property
rests on was unpinned. The guard is now re-asserted at the eval site and pinned
by a test. The reachable vector is a VALUE with an embedded newline, not a
hostile name — measured: zsh strips non-identifier env names outright, while
MULTI=$'keep\njunk.fragment' forges 'junk.fragment' as a candidate name out of
its own value.
Closes#394. Refs #396, #388.
EnvAllowListScrub's blanking loop splices each name into a string
handed to eval ("export ${n}="). That is only safe because every name
reaching _cb633_blank already passed an identifier check in the
enumeration loop -- 20 lines away, in a different loop. Before eval
was introduced a non-conforming name reaching plain `export "$n="`
was inert either way (the quoting neutralized it); eval removed that
safety net, so the enumeration loop's guard became the ONLY thing
standing between a non-identifier string and code execution in the
member's pane, with nothing at the eval site itself defending that
property.
Re-assert the same [A-Za-z_][A-Za-z0-9_]* check immediately before
the eval call, independent of the enumeration loop's own guard (left
untouched, not moved). A name that fails it is counted unblankable
rather than silently dropped, so a bypass of the upstream guard would
leave real evidence in the report.
New test exploits the "junk from multi-line values" gap the
enumeration loop's own comment already documents: a value with an
embedded newline makes `command env`'s text output split into a
spurious extra "name" line that was never a real variable. Runs the
real generated scrubScript() end-to-end under zsh and asserts the
non-conforming fragment is neither blanked nor counted unblankable.
The fragment used is merely non-conforming (contains a dot) --
never command-shaped.
Mutation-verified both guards. Weakening the enumeration guard alone
DOES break the new test (the fragment then reaches the new eval-site
guard and gets counted unblankable, failing the "not unblankable"
assertion). Removing the new eval-site guard alone, with the
enumeration guard intact, does NOT break it: _cb633_blank has exactly
one producer (the enumeration loop), so nothing can reach the eval
site without already having passed the identical check there. That is
expected given the single-source architecture, and it is exactly why
the eval-site guard is defense-in-depth against a future change that
adds a second path into _cb633_blank or decouples the two loops --
not a currently independently-observable divergence.
`fleet_reply` has no route to a peer lead. `AmqpReplyInbox` publishes to
`agent.<target>.inbox`, mandatory, and a lead's own terminal has no such
queue, so the publish is refused. `MessageService.reply()` has no peer
branch at all — `grep -c 'coord\|LeadMailbox'` on it returns 0. The charter
told every lead to use a tool that cannot work, and both leads here hit it.
Three edits to the canonical block, byte-identical with the wiki template
(pushed as 803726a; the in-sync check in this file reports True):
- the intent->tool row now says `fleet_send{coordId}`, or `{sessionId}` for
a peer on the same host, and says plainly that `fleet_reply` is refused
- the prose says WHY: `fleet_reply` resolves a member's blocked `fleet_send`,
while a peer's coord-id message is durable and non-blocking, so there is
nothing for it to resolve
- lead<->lead item 3 gains the data-point rule: N observations are N data
points only if they differ in the axis you are trusting
Wording for all three drafted by the fleet01 lead, who verified the missing
queue namespace independently in its own tree. The data-point rule has now
caught three separate errors in a day, in both directions: one cause blamed
for N failures, and N agreeing measurements that shared a single instrument.
The refusal message itself is still wrong — it says "queue not declared or
owned", which sends the reader to the broker instead of to this file. That
half stays open on #391.
Tracked as fleetd #391.
Add an optional top-level `models:` block (Models{allow: List<ModelEntry>})
naming the models any profiles: entry may use. Absent/empty allow: keeps
today's behaviour exactly (no check, no warning). When configured,
FleetConfig.validateModels() fails config load (and reload, via ConfigRef)
naming both the model and the profile, if any profile's model: is outside
the list. The check is one-way: editing profiles: alone can never widen
what is permitted, only models.allow: can.
Wired into Fleetd.main() alongside the other validateXxx() calls, and into
ConfigRef.reload()/DEFERRED_KEYS so a bad edit can't slip in through a
reload either. Each ModelEntry is its own record (not a bare string) so a
later unit can add per-model on/off or load-limit state without changing
the YAML shape. One flat string namespace covers both a bare Claude id and
an opencode provider-prefixed id.
EnvAllowListScrub's blanking loop used a plain `export "$n="` on every
name not on the allow-list. For a zsh read-only/special parameter (e.g.
UID) that is a FATAL parameter error, and since the loop runs inside the
sourced startup file, the error aborts the whole file: every name still
to come is never blanked, and scrub-report.txt is never written at all
-- silently, because 2>/dev/null on the group swallows it.
Route each blanking attempt through `eval` instead, which contains the
error to that one iteration. The loop always finishes; a name it could
not blank is now counted separately ("failed" on the report's first
line) and listed !-prefixed rather than disappearing. No skip-list of
known-bad names is added -- every enumerated name is still attempted,
so a name nobody has thought of is still tried and, if it fails, still
counted.
HerdrPeerLauncher: log a WARN when a pane's report carries a nonzero
failed count, and reword the "no report at all" WARN so it no longer
claims the daemon knows the member "saw the full host environment" --
a partial vs. a missing scrub are different situations and only the
first is now distinguishable from the report alone.
EnvAllowListScrub generated four zsh startup files but only .zshrc and
.zlogin sourced the scrub — .zshenv (the one file zsh always reads) did
not. A pane shell that is neither login nor interactive reads only
.zshenv and stops, so it was never scrubbed at all (measured on fleet01,
issue #388).
Adding an unguarded scrub to .zshenv (the ticket's own suggested fix) is
wrong: .zshenv is read by every zsh, including a short-lived `zsh -c`
a member's own tooling forks for a single command. Those children are
also neither login nor interactive, so they would scrub the environment
their parent deliberately set for them (GIT_DIR, VIRTUAL_ENV, ...), and
the rewritten scrub-report.txt would describe the last child to exit
instead of the pane.
Fix (per comment 15387, measured): keep .zshrc/.zlogin unconditional,
and add to .zshenv a pass guarded on the exact condition that defines
the gap (neither login nor interactive), plus a per-pane sentinel
(_CB633_SCRUBBED) so it runs once per pane, not once per process. The
sentinel is exported only after the scrub runs, and is folded into the
scrub's own allow-list so a later pass in the same pane cannot blank it
back to empty.
Also corrects the class javadoc's wrong premise (a bare argv[0] proves
NOT login, not "therefore interactive") and its now-stale two-file
walkthrough.
Tests: two new real-zsh tests in EnvAllowListScrubTest run actual
non-login/non-interactive zsh processes (never string-match the
generated files) to prove: a neither-shell pane is scrubbed; a child
that pane forks keeps variables the pane deliberately set for it; the
child does not re-scrub; and scrub-report.txt still describes the pane
after the child exits. Both fail without the production fix (verified
by reverting it and re-running: AssertionFailedError on the sentinel
and on the decoy secret surviving).
The two tests merged with #386 both start with the member already BUSY, so a
single global drift baseline passes them. This one sleeps the host while nothing
is busy and only then starts a turn, which fails without the per-member map.
FleetHealthMonitor.tick's stalled check compared two monotonic-clock
readings (System.nanoTime(), which macOS freezes across a host sleep),
so a member BUSY for 101 real minutes was never flagged.
The monitor now also takes a wall-clock LongSupplier (realtimeClock),
used only inside the stall check. Each tick measures how far the two
clocks moved apart since the previous tick and folds any positive
divergence into a running total; when a single tick's divergence
exceeds one tick interval (the signature of a sleep, since a tick
cannot run while the process itself is suspended) it logs one WARN
naming how long the detector could not see. The correction is applied
per member, keyed to when that member's current lastActivityAtNanos
was first observed BUSY - not since monitor start - so a sleep that
happened before a member went busy is never charged to it.
Every other use of the monitor's clock (readiness grace, snapshot
timestamp) is unchanged. Backend quarantine/cool-off, the lead tab
scan, the completion resolver, the session reaper and the message
service TTLs are untouched, per the ticket's decision.
Existing FleetHealthMonitor/FleetHealth tests pass unmodified (none of
them ticks a BUSY session more than once, so the drift path never
engages for them). Two new tests: a frozen monotonic clock past the
real-time threshold produces STALL_SUSPECTED, and a single sleep gap
logs the divergence exactly once, not once per tick.
CompositePeerLauncher:372 rebuilt a routed SpawnRequest from a literal
new SpawnRequest(...) call listing six of the original request's own
accessors. That call is only correct because it happens to match the
canonical 6-arg constructor today; add a 7th component plus the
established back-compat constructor at the old (now-shorter) arity and
this call would silently rebind to it, dropping the new field on every
profile-routed spawn with no compile error — the same defect shape
already guarded on FleetConfig.withDefaults() (#357) and MemberSession
(#358).
Add SpawnRequest.withProfile(String), modeled on
MemberSession.withState/withActivity, and use it at the call site
instead. Add a guard test that resolves the true canonical constructor
by exact component types (never by argument count), gives every
component a distinctive value, and asserts every component but
profileName survives withProfile() unchanged.
Proved the guard against the real mechanism: temporarily dropped the
last (role) argument from withProfile()'s constructor call so it bound
to the 5-arg back-compat constructor — it still compiled, and the new
test failed, catching the silently-defaulted role. Restored the fix
and reconfirmed green.
The old line said profiles differ in model and cost. That is true and it is
not the reason the default hurts. Measured on two hosts: a default sitting on
an exhausted or withdrawn credential either fails the spawn loudly or, worse,
produces a member that starts fine and then returns nothing.
Wording proposed by the fleet01 lead; merged with the existing 'not in tier'
clause, which is still right. Applied byte-identically to the wiki template.
Verification found the fix's dedup check could be moved behind the
injectable-status gate and every test still passed. That placement matters:
this lead is mid-turn most of the time, so gating the ack on an idle pane
leaves the redelivered message held, and the next recovery delivers it again.
The new test fails on that mutant and passes on the fix.
An AMQP recovery clears the held delivery tags, the broker redelivers with
fresh ones, and the coordination loop wrote the same peer message into the
lead pane again. Measured on the live daemon: one msgId reached the pane 12
times in 9 hours, across 19 recovery events.
LeadCoordLoop now remembers the msgIds it has written to a pane (bounded at
1024) and acks a redelivery without a second write. LeadMailbox.ack no longer
returns quietly for an unknown msgId: it throws, so an ack that never reached
the broker is reported instead of hidden. A repeat ack that this connection
already completed stays quiet, tracked in a bounded set.
The 9 tests the member wrote failed 7 of 9 on first build. The production
code was correct; the test harness was not. logback-test.xml sets
dev.ltms.fleet to WARN, so the INFO shape lines were dropped by the level
check before any appender saw them.
The member copied attach()/detach() from MemberTrustModelReportTest but not
the setLevel(INFO) those siblings do at each call site. Doing it inside
attach()/detach() covers all nine at once, and restores the original level
(null, meaning inherit) rather than a concrete one.
Proven by mutation: making the code log the value fails 4 tests, including
theValueNeverAppearsInLogOutput — 'the full GITEA_HOST value must never
reach the log'. Full build 1459 tests green.
A member receives the git host as GITEA_HOST and often gets a full URL
(scheme, trailing slash) where it expects a bare host, so it builds
https://https://... and the request never leaves the machine.
Logs set/unset, length, startsWithScheme and trailingSlash next to the
existing secret report. The value itself is never logged, and the value
is still passed to members unchanged — rewriting it here would change
what works on one host and breaks on another.
Written by the gx member on branch worker/t377-7f587b-7; it could not
build, commit or push because the host command classifier refused every
shell command (see #381). Build and verification are mine.
The label was renamed in the #365 merge. This table still named the old one.
'sent' counts the herdr paste-and-submit call returning, never a confirmation
that the pane read it.
A turn that settles inside MIN_TURN_NANOS is still reported FAILED. Only the claim
about WHY is withdrawn, and the pane is still carried.
Measured on fleet01 on 2026-09-08 UTC: an opencode member on mimo-v2.5-free answered
a real question in 1575ms, below the 2000ms floor. The daemon reported the turn as
failed with 'most likely a backend error before any work started'. The answer was
right there in the scrape. With no errorPattern configured — the live state on both
hosts, which both log as 'backend-error classification: off' — that sentence is a
guess, and a reader who believes it stops looking at the pane.
WHAT I REJECTED, because the next person will try it. A worker implemented the
ticket's first suggested direction: inspect the pane inside the floor and resolve a
COMPLETION when the text looks like a real reply. Its test for 'looks like a real
reply' was non-blank plus a '.', '!' or '?' anywhere in the text. That is unsafe
twice over. lastAssistantBlock falls back to the WHOLE pane when it finds no U+23FA
marker, so on a crash the candidate reply is the entire screen; and a crash pane
almost always contains a full stop, in a file path, a version or a hostname. I ran
that implementation against the new guard test and it resolved
Error: connection reset while loading src/main/java/Foo.java v1.2.3
as a COMPLETION — expected: <FAILED> but was: <COMPLETION>. A loud wrong answer
became a silent one, which is the trade the ticket brief forbade.
The obvious repair does not work either. Requiring the U+23FA marker as positive
evidence would be safe, but that marker is Claude Code chrome and an opencode pane
never carries it — and an opencode member is what raised this ticket. There is no
reliable cross-backend marker for 'this is a real reply', so this path must not try
to judge one. That reasoning is now in the failTooFast javadoc.
Two guard tests. aPlausibleLookingReplyInsideTheFloorStillFails pins the safety
property against exactly the rejected approach; it passes today and fails against
that implementation, which is how it was verified rather than assumed.
theTooFastFailureDoesNotAssertACauseItCannotKnow pins the wording.
One existing test changed. aNonMatchInsideTheFloorStaysGenericAndNeverNotifiesTheSink
asserted the phrase 'too fast to be real work', which carried the withdrawn claim. It
now asserts what it was really guarding: the floor alone fails the turn, the reason
stays generic, the pane is carried, and the typed sink is never notified.
MIN_TURN_NANOS is unchanged at 2000ms. mvn clean install: 1441 tests, 0 failures.
fleetd #362 review finding 2 protects GitWorktrees#previouslyEffectiveExcludesFileContent's
Java-side XDG_CONFIG_HOME/HOME read (it never goes through a git subprocess, so no
GIT_CONFIG_GLOBAL/GIT_CONFIG_SYSTEM isolation reaches it) with a gitEnv constructor seam. A
mutation run during the #372/#369 merge found that seam unpinned: stripping hermeticGitEnv(tmp)
from seedingGitWorktrees left every test green, poisoned XDG_CONFIG_HOME or not.
Adds seedingGitWorktreesResolvesTheExcludesFileFallbackInsideItsThrowawayDirectory, which asserts
the property directly (a GitWorktrees built for seeding resolves the fallback inside its own
throwaway directory) using a self-contained marker instead of relying on an externally poisoned
env var. Refactors hermeticGitEnv/seedingGitWorktrees into two-argument overloads (one taking an
explicit XDG_CONFIG_HOME / gitEnv) so the new test can pre-populate the marker before construction
while still going through the same production construction every other seeding test uses; no
behavior change for the 4 existing call sites.
fleetd #365. fleet_reply always returned the literal "delivered" and
POST /sessions/{id}/reply always returned {"delivered": true}, whether
the reply resolved a live waiting send/ticket or was merely queued in
the inbox for a later drain (CB-307) — both are successes, but not the
same fact.
MessageService.reply() now returns a ReplyOutcome (RESOLVED_SEND,
RESOLVED_ASYNC_TICKET, or QUEUED) instead of an always-true boolean.
FleetMcp.reply and FleetApp.replyMessage both read it: the MCP tool
result names which happened, and the REST body's "delivered" field is
now accurate, with an added "outcome" field.
Also renames the heartbeat/push-loop nudge metric's "delivered" outcome
to "sent" (LeadHeartbeatLoop, ReplyPushLoop, FleetMetrics): it only
records that the herdr agent.prompt paste-and-submit call succeeded,
never that the lead's pane actually read it — there is no read-receipt
concept at that layer, so "delivered" overclaimed there too.
Tests: MessageServiceTest/FleetMcpTest/FleetAppTest strengthened to
assert the specific outcome per case (a resolved send, a resolved async
ticket, and a queued reply); ReplyPushLoopTest updated for the outcome
rename.
Follow-up to #357 (FleetConfig.withDefaults()). Reflectively enumerate each
record's own components, resolve the canonical constructor by exact
component types, build a real non-null value per component, run each
rebuild site, and assert every component survives (except the one it is
documented to change). Exclusion lists pinned at 0 for both.
Re-counted the Profile back-compat ladder directly against the source:
8 constructors (arities 25, 24, 22, 20, 18, 15, 14, 12) against a
canonical arity of 26 — the ticket's own number was explicitly untrusted.
The operator chose this and its size; the reasoning below is the fleet01 lead's.
Placed in the canonical block's boundary paragraph rather than the orchestration
body. That paragraph already talks about the addendum layer instead of protocol,
every project that mounts the bridge inherits it, and it sits about 3800 characters
before the primary's step list, so it does not dilute the steps a lead reads while
working. Perishability is structurally an addendum problem: the block is
byte-identical across projects by construction, so a dated local measurement in the
block body would already be a layering violation.
What happened. The fleet01 lead's kb addendum held a dated merge-refusal section
that carried an instruction to delete itself once it stopped reproducing. On
2026-09-08 UTC the operator granted merge rights on akb/kb, the lead re-ran the
probe, got 409 'head out of date' where the identical request had returned 405
'User not allowed to merge PR' on 2026-09-06, and deleted the section as instructed.
Why four parts and not one. The lead's finding is that the banner did not work
because it was emphatic. It worked because the falsification condition was
executable: it carried the exact probe, the reason for the all-zeroes
head_commit_id, and what each response code meant. The lead did not have to
reconstruct the experiment or decide what would count as refutation, and just ran
it. A banner saying 'this may be out of date, verify before relying on it' costs
the same space and does nothing, because deciding what would falsify a claim is the
expensive step and a reader in the middle of another task will not pay it. So: the
date, the command, what each outcome means, and the instruction to delete. The
fourth without the second is decoration.
The closing clause is the justification for the machinery. Most stale notes are
merely wrong. This one went stale in the dangerous direction: it would have told a
future lead it could not merge at the exact moment merging became its job, silently
and with confidence. A note that goes harmlessly stale does not need this.
Note what is NOT centralized here. The banner text itself cannot be. What fired for
the lead was a specific instruction sitting on top of the specific stale fact, which
it could not read past on its way to acting. A rule elsewhere saying 'date your
measurements' would not have fired, because nobody reads that rule at the moment
they re-measure. This sentence sets the convention; the trigger still has to live
next to the fact it governs.
Propagated to the wiki template in the same turn, wiki 8c4f152 on main; the sync
check in this file's addendum reports 'in sync: True'.
Step 8 gained a refusal paragraph in 2f71a30, which said what a lead does once the
forge refuses a merge. It did not say how a lead establishes that it was refused.
Both halves of this amendment come from the fleet01 lead, measured on akb/kb on
2026-09-08 UTC, and both are ways to be wrong about a permission you never tested.
Do not read a refusal off a permissions field. After the operator granted merge
rights, the lead re-ran its probe: POST .../pulls/53/merge with an all-zeroes
head_commit_id, chosen so the request cannot succeed on its merits and a rejection
can only mean the refusal. It returned 409 'head out of date' where the identical
request returned 405 'User not allowed to merge PR' on 2026-09-06. A 409 is payload
validation and sits after the permission gate, so the grant took. The lead reports
the repository permissions object did not change across that flip -- still
admin:false, push:true, pull:true. I did not read that object myself; my forge token
is a different identity and would return a different one, so this stays the lead's
measurement and not mine. Merge rights on a protected branch live in branch
protection, so a permissions field can be wrong in both directions.
Do not count a transport failure as a refusal. The lead's first attempt returned
HTTP 000, because GITEA_HOST already carries a scheme and a trailing slash and the
URL came out as https://https://git.ltms.dev//api/... Under a 'not 200' test that is
indistinguishable from being refused. A probe exists to separate a refusal from
everything else, so an error that never reached the gate has to be a third answer
that concludes nothing.
Propagated to the wiki template in the same turn, wiki 8c2ef96 on main; the sync
check in this file's addendum reports 'in sync: True'.
Lands PR #355 (fleetd #354's sibling), rebased onto current main by a worker
after 39 commits of drift left it unmergeable.
The problem, measured on the original branch: a fleetd host idle-slept after as
little as one minute (pmset -g custom reported 'sleep 1' on battery). Overnight
the daemon's AMQP link dropped 13 times, and every drop minute had a sleep or
wake event in pmset -g log in the same minute or the one before. The AMQP churn
is the visible symptom; the real cost is a member mid-turn freezing with the
host, and a long turn with nobody typing is exactly the case that goes idle.
IdleSleepGuard holds an OS-level assertion for as long as at least one member is
live. It is driven by SessionManager's existing onAcquire/onRelease hooks rather
than a second member count kept in parallel, so it reads the same registry
fleet_list's numbers come from, and only a real 0->1 or 1->0 crossing touches the
OS. It fails safe: a mechanism that cannot acquire means nothing is ever held,
and it never throws, never blocks a spawn, a release, or shutdown.
Conflict resolution was the whole job, and all three were in config plumbing:
ConfigRef, FleetConfig and ConfigRefTopLevelReportingCoverageTest. The power
package is byte-identical to the original branch commit.
Verified on this merge, not taken from the worker's report:
mvn clean install -> Tests run: 1439, Failures: 0, Errors: 0, BUILD SUCCESS
(1425 on main + 14 new: 4 caffeinate, 5 guard, 1 wiring, 4 config)
The denominator recount, which the worker flagged as its own weakest number
because this file's count has drifted three times before (#330/#333/#337). I
counted it mechanically rather than reading it: FleetConfig has 24 canonical
record components; COLD_KEYS 5, DEFERRED_KEYS 13, SPLIT_KEYS 3, plus the 3 the
javadoc names as hot-excluded (placement, memberCredentials, memberLoginShell).
5+13+3+3 = 24. The javadoc's '24 components: 5 cold, 13 deferred, 3 split, 3
hot-excluded' is correct. The worker's prose called idleSleepGuard the 25th
constructor argument; it is the 24th. The code is right, the report was off by
one.
Mutation run on merge, on the half the worker verified by READING rather than by
proving -- it said it had checked that withDefaults()'s final call binds the true
canonical constructor. I dropped the trailing idleSleepGuard argument so the call
silently binds the 23-arg back-compat overload. It compiles, which is the whole
hazard. Caught: 1 failure, 3 errors, BUILD FAILURE, and
FleetConfigWithDefaultsPreservesEveryComponentTest names the dropped component
and prints its own denominator -- '24 components, 24 checked, 0 excluded, 23
survived'. That test was added on main after this exact defect happened live when
idleSleepGuard was added on a sibling branch; the worker had to add the missing
entry to it, and doing so is what makes the guard cover this component at all.
FleetConfigWithDefaultsPreservesEveryComponentTest was added on main after
the sleep-guard commit was cherry-picked, specifically to catch this exact
rebase hazard (a withDefaults() call silently rebinding to a stale-arity
back-compat constructor). Its baseValues() map didn't know about the new
idleSleepGuard component yet, so the coverage test itself failed the
name-drift check. Add a real, non-null value for it, consistent with how
the sibling ConfigRefTopLevelReportingCoverageTest already covers it.
Adds a small IdleSleepGuard (dev.ltms.fleet.power) that holds a macOS
caffeinate -i child while at least one fleet member is live, and
releases it once none are. It hangs off SessionManager's existing
onAcquire/onRelease hooks and SessionManager#size() rather than
tracking members a second way. New idleSleepGuard: config block,
on by default, following the FleetConfig.Health/ConfigReload pattern.
Fourth item under Lead <-> lead. The addendum is instruction surface -- every
future session on that host obeys it, and a wrong one is obeyed as faithfully as
a right one -- but nothing in the block said to have anyone check it.
Evidence is one addendum, written on fleet01 this week, and it carried two
defects that its author did not see and a non-author did. First: it described a
forge permission wall as if it were policy, so a token regrade would have made
it tell a session NOT to merge at the moment merging became its job. Second,
found only because the first was raised: it DID carry the 'this is a dated
measurement' caveat, at the bottom, after the prohibition -- so a session that
read the instruction and stopped had already taken it as policy. The caveat sat
downstream of the thing it qualified.
Neither is a writing slip. Both are the author being unable to see their own
qualifier placement, which is what a second reader is for.
Deliberately conditional: not every operator runs two leads, so the rule ends
with what a lone lead can still do -- ask which sentence goes false first, and
whether a reader reaches the caveat before acting.
Weaker evidence than the step 8 amendment (n=1 addendum, 2 defects, versus a
measured 405). Recorded as such so it can be dropped if it does not earn the
context it costs.
Step 8 said 'then merge' and nothing about a refusal. fleet01 is the first host
to hit that: on akb/kb its lead gets 405 'User not allowed to merge PR' from the
API, and main is protected, so merging locally and pushing is refused too. Both
routes shut. The charter was telling a lead to do something the forge would not
let it do, and that gap was latent in every copy of the block.
The amendment keeps the rule that the merge decision is never delegated, while
admitting the mechanical merge may not be the lead's to make.
The second sentence is the one that earns its keep, and it came from the fleet01
lead rather than from me: never call a PR 'ready to merge' without having read
the diff. A refusal is exactly when that shortcut is tempting, because no action
is left that forces the lead to look. Without it, a refused merge quietly turns
step 8 from 'read it yourself, then merge' into 'forward the reviewer's verdict'
— the proxy-delegation the same step forbids two lines earlier.
Propagated to the wiki template in the same turn; the sync check in this file's
addendum reports 'in sync: True'.
The comment claimed no test in this class can reach the real machine's home
directory. Measured on the merge: 53 of the 58 new GitWorktrees(...)
constructions here pass no env override, and stripping the override from
seedingGitWorktrees leaves the class green under the poison command that the
same comment cites as proof. Say what it covers and what it does not.
fleetd #369. Every git subprocess the test class starts now goes through one
gitProcessBuilder factory that applies the hermetic environment. Before this,
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus set nothing at all -- so the tests inherited the JVM's
whole real environment, including the operator's default excludes file. That
file applies with no core.excludesFile configured at all, and /dev/null for the
global config does not stop it.
Round 1 pinned this with a call-site count: exactly 2 literal
new ProcessBuilder( occurrences. That catches a NEW helper built the old way,
but it is a proxy, not the property. I measured the gap -- deleting
pb.environment().putAll(hermeticEnv()) from inside the factory left every call
site unchanged, the count stayed 2, and 1413 tests stayed green. Round 2 added
gitProcessBuilderCarriesTheFullHermeticEnvironment, which asserts on what the
factory actually hands to ProcessBuilder#start(). Both checks are kept: they
catch different regressions.
Verified on this merge, not taken from the worker's report:
mvn clean install -> Tests run: 1414, Failures: 0, Errors: 0, BUILD SUCCESS
poison control on the PRE-FIX file:
XDG_CONFIG_HOME=<dir with a '*' git/ignore> mvn test -Dtest=GitWorktreesTest
-> tests=59 failures=56, so the poison genuinely reaches these tests
Mutation run on merge, on a half neither the worker nor the reviewer touched --
made seedingGitWorktrees pass null instead of hermeticGitEnv(tmp), which strips
the hermetic environment from the PRODUCTION GitWorktrees instances rather than
from the test's own subprocesses: tests=61 failures=0 unpoisoned AND poisoned.
That half is unpinned. It is fleetd #362's protection, not this ticket's, and it
guards a different path -- the Java-side XDG read in
previouslyEffectiveExcludesFileContent, which no assertion observes. Out of
scope here; filed as a follow-up rather than held against this PR.
One javadoc sentence corrected in the merge: hermeticGitEnv claimed "no test in
this class can reach the real machine's home directory". 53 of the 58
new GitWorktrees(...) constructions in this file pass no env override at all, so
the claim is true of the 5 seeding sites and of every test-started subprocess,
not of the class.
fleetd #368. PrimaryRegistry.forgetDelegation fires only when a WORKER is
released, never when the delegating LEAD goes away. The map is keyed by the
worker, so a closed, crashed or relaunched lead left its bindings behind. A
stale non-null entry then beat nudgeTargetFor's single-primary fallback every
time -- and that fallback's own javadoc argues it is correct precisely in the
case the stale entry was hiding.
ReplyPushLoop now probes the recorded lead with agents.status before trusting
it, at all five entry points, and a lead that is really gone is forgotten so
resolution reaches the fallback.
Round 1 caught any RuntimeException and treated it as death, and death here
calls forgetDelegation -- destructive and permanent on ONE reading. That is the
#359 mistake repeated two days later: a socket blip or a decode error on a live
lead would silently unbind it forever. I measured the breadth was unpinned
(narrowing it left 1414 green), and round 2 narrowed it to the one affirmative
signal AgentControl.agentCall itself uses, agent_not_found. Every other failure
is treated as live, because guessing wrong costs one extra retry next tick while
guessing wrong the other way is unrecoverable.
The fallback it lands on is live, not stale: PrimaryRegistry.record overwrites
the single slot on every orchestration-side MCP call, and neither host pins
primary.terminal, so a relaunched lead re-registers on its first tool call.
Verified on this merge, not taken from the worker's report:
mvn clean install -> Tests run: 1415, Failures: 0, Errors: 0, BUILD SUCCESS
Mutation run on merge, on a half the worker did not touch -- dropped the
forgetDelegation call while still returning the fallback, so behaviour on the
first nudge is identical and only the self-healing is lost: 2 failures,
BUILD FAILURE. The cleanup is pinned, not just the fallback.
everyGitSubprocessGoesThroughTheHermeticFactory counts ProcessBuilder("git",
...) call sites, so it catches a new helper built the old way, but a
reviewer proved it does not catch gitProcessBuilder itself being gutted:
removing pb.environment().putAll(hermeticEnv()) from inside the factory
leaves every call site unchanged, the count stays 2, and the whole
unpoisoned suite stays green.
Add gitProcessBuilderCarriesTheFullHermeticEnvironment, which inspects what
the factory actually hands to ProcessBuilder#start(): every hermetic key
present with the isolating value, and XDG_CONFIG_HOME pointed inside the
class's own throwaway directory rather than left unset or pointing at the
operator's real one. This fails the moment the hermetic environment stops
being applied, on any machine, with no poison needed. Keep the call-site
count check too — the two catch different regressions.
isLive treated any RuntimeException from the liveness probe as "the lead is
gone", which forgetDelegation then acted on destructively and permanently.
That made a transient herdr hiccup (socket blip, decode error) on a perfectly
live lead indistinguishable from the lead actually being dead — the same
one-bad-reading mistake #359 shipped a guard against for lead-tab liveness.
Narrow isLive to match AgentControl.agentCall's own rule: only an affirmative
HerdrException("agent_not_found") counts as gone. Every other failure is
treated as still live and the binding is left alone.
Adds aTransientLivenessFailureMustNotForgetABindingToAStillLiveLead, which
fails with the bare RuntimeException catch and passes with the narrowed one.
PR #370 shipped both units but its install block stopped at
`systemctl --user enable --now`. Without lingering a user manager starts at
your first login and stops at your last logout, so the units do not come back
after a reboot -- which is the whole reason this ticket moved fleet01 off the
setsid scripts.
It is easy to miss because leaving it out looks like success: `systemctl --user
enable` reports "enabled" and both units run while you stay logged in. The
issue named this and the PR did not carry it over.
fleet01 itself is fine -- measured `Linger=yes`, both units `enabled`. This is
about the next host that follows these instructions.
Comment only; SystemdUnitSafetyTest still 8 green (a commented line is not an
active directive).
fleetd #360. deploy/fleetd.service shipped four mount-namespacing directives
(ProtectSystem, ProtectHome, ProtectKernelTunables, ProtectControlGroups),
PrivateTmp=true, and an ExecStart that ran java directly. Each one starts green
and breaks the daemon in a way nothing logs: lsof goes blind so every caller is
resolved ANONYMOUS and refused; the member ZDOTDIR scrub becomes a no-op; every
credential is empty. deploy/herdr.service did not exist at all, though
fleetd.service's After=/Wants= already named it.
Both units are now the ones running on fleet01, comments included -- the
bisected lsof counts and the reasons live in the files, because the next person
to 'harden' this needs the reason, not the rule.
A unit file has no compile step, so SystemdUnitSafetyTest reads both units plus
deploy/herdr-inner.sh and fails on an active forbidden directive, on
PrivateTmp=true, on an ExecStart that skips the login shell, on a herdr-inner.sh
that does not exec a login shell, and on one that does not set a non-zero pty
size. Each message names the consequence. A vacuity guard pins that all three
files exist and that the DO-NOT-add comment block still mentions every forbidden
directive, so 'comment survives, directive does not' is actually exercised.
Verified on this merge, not taken from the worker's report:
mvn clean install -> Tests run: 1420, Failures: 0, Errors: 0, BUILD SUCCESS
Round 1 shipped herdr-inner.sh untested; I mutated it (dropped both the login
shell and the stty sizing) and got 6 green. Round 2 added those two checks.
Mutation run on merge, on a half the worker never touched -- ProtectHome=read-only
added to herdr.service, the unit it only ever mutated fleetd.service for:
Tests run: 8, Failures: 1, BUILD FAILURE. The parameterisation really covers
both files.
herdr.service's ExecStart only names deploy/herdr-inner.sh, so that script was the only
place herdr's login-shell and pty-size properties lived, and nothing was reading it -- the
same silent-at-startup shape the ticket was about, one file further down the chain. A
mutation dropping both the login shell and the stty sizing left SystemdUnitSafetyTest green.
Add herdrInnerScriptUsesALoginShell and herdrInnerScriptSetsANonZeroPtySize, extend the
vacuity guard to cover herdr-inner.sh too, and tighten execStartUsesALoginShell's check from
a bare contains("-lc") substring match to the same login-shell-invocation regex the new
checks use.
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus (plus every other raw git subprocess in this class) set no
isolation at all — inheriting the JVM's real environment, including the
operator's real ~/.gitconfig and default excludes file. Measured: with a
poisoned XDG_CONFIG_HOME, 56 of 59 tests failed.
Centralize every git subprocess this test starts through one factory,
gitProcessBuilder, which always applies the existing hermeticGitEnv isolation
(extended with XDG_CONFIG_HOME, the same fix#366 already applied to the
production-instance seam). Add a self-check test that counts direct
ProcessBuilder("git", ...) constructions in this file's own source and fails
if a future helper bypasses the factory, so the omission that caused this
ticket is caught by name instead of rediscovered on a poisoned machine.
PrimaryRegistry.forgetDelegation only fires when a WORKER is released, never when
the delegating LEAD terminal itself disappears (closed, crashed, or relaunched).
A stale, non-null leadByTarget entry always beat nudgeTargetFor's single-primary
fallback, so a dead lead silently swallowed every reply nudge for its workers.
Fix: ReplyPushLoop now verifies (via the same agents.status check decide() already
uses every tick) that a recorded delegating lead is actually live before trusting
it. A dead lead is treated as if never recorded — self-healing the binding
(mirroring AgentControl.paneByTerminal's self-heal on agent_not_found) and falling
through to PrimaryRegistry's existing fallback.
deploy/fleetd.service started clean on fleet01 but broke the daemon in three ways nothing
logs: ProtectSystem/ProtectHome/ProtectKernelTunables/ProtectControlGroups each put the unit
in its own mount namespace, which blinds fleetd's lsof-based caller lookup and falls every
caller back to ANONYMOUS; PrivateTmp=true silently no-ops the credential scrub the member
pane depends on; and running java directly from ExecStart skips the login shell that sources
the daemon's secrets, so it boots with empty credentials.
Replace the unit with the version verified working on fleet01 for a day, and add the
deploy/herdr.service companion unit it was already depending on via After=/Wants= but which
did not exist in the repo. Add deploy/herdr-inner.sh as the login-shell template
herdr.service's ExecStart wraps in a pty.
Add SystemdUnitSafetyTest (fleetd/src/test/java/dev/ltms/fleet/deploy) to read both unit
files from disk and fail if a forbidden mount-namespacing directive is active, PrivateTmp is
true, or fleetd.service's ExecStart does not go through a login shell -- the only guard
possible for a unit file with no compile step.
fleetd #359. LeadTabScanner used to join labelled tabs straight to terminals
with no liveness check, and its javadoc excused that ("a stale name costs
nothing here"). It cost plenty: LeadCoordLoop reads that map to pick which
pane a peer message goes into, so a dead tab was a candidate it could pick.
LeadLauncher, meanwhile, had no cleanup path at all -- every reconcile that
found 0 live created another tab and left the old one.
Both now cross-check agent.list, and neither trusts a single reading of it.
That matters because the daemon's own evidence on fleet01 was agent.list
reporting 0 live while ps showed one real claude. A first cut of this fix
closed tabs on that single reading, which would have closed the operator's
live lead instead of leaving a spare tab. So: a dead tab is flagged, not
closed, and only closed when a later reconcile still finds it dead; and the
scanner grants one grace scan to a terminal it already knew was live.
Verified on this merge, not taken from the worker's report:
mvn clean install -> Tests run: 1412, Failures: 0, Errors: 0, BUILD SUCCESS
Mutation run on merge (PendingCloseMarker.strip made identity, so a flagged
tab stops matching its configured label): 4 failures, BUILD FAILURE.
Known and accepted: ensureLeads() runs at startup, so the second reading
arrives at the next restart. A tab that dies mid-session stays flagged and
open until then. Deliberate -- the bug is about repeated restarts, and one
leftover tab is cheaper than closing a live session on unverified evidence.
Finding 1 (LeadLauncher): closing a labelled tab on a single agent.list
miss could destroy a live lead's session — the ticket's own evidence
showed that exact signal missing a genuinely running agent. A dead
reading now only flags the tab (PendingCloseMarker); it is closed only
if a later, independently-connected reconcile still finds it dead
while flagged. A tab found live again has its flag cleared instead.
Finding 2 (LeadTabScanner): the new agent.list cross-check in scan()
was not covered by get()'s "keep the cache on a failed scan" contract,
which only fires on a thrown HerdrException. A successful-but-short
agent.list could silently drop a lead CallerResolver had already
resolved, demoting it to Role.WORKER. A terminal already reported live
now gets one grace scan before being dropped; a terminal never
reported live gets none, so the original #359 exclusion is unaffected.
Both mechanisms were mutation-tested: reverting either change turns
exactly its own new tests red and nothing else.
fleetd #362 item 3. A member spawned against a repo that does not ship its
own .claude/skills/ could not load implementer, reviewer or hunter at all.
Every brief starts with "Load the <name> skill", and outside this repo that
line was silently a no-op. memberSkills: <dir> now copies those folders into
each provisioned worktree, skipping any name the target repo already ships.
Two review rounds, both about the same hazard: core.excludesFile is
single-valued, so pointing it at fleetd's own file would SHADOW the
operator's. It now composes instead of replacing, and the XDG default
excludes file is carried forward too.
Verified on this merge, not taken from the worker's report:
mvn clean install -> Tests run: 1402, Failures: 0, Errors: 0, BUILD SUCCESS
Two mutations run on merge:
drop the XDG fallback -> 1 failure, BUILD FAILURE
remove the composition itself -> 2 failures, BUILD FAILURE
Both directions are pinned.
Finding 1 (lead): the XDG fallback branch of previouslyEffectiveExcludesFileContent was
unpinned — deleting it left the suite green (Tests run: 1386, Failures: 0). Added
seedSkillsComposesWithTheXdgDefaultExcludesFileWhenNoneIsConfigured to pin it: isolates
XDG_CONFIG_HOME via the gitEnv seam at a temp dir carrying a synthetic git/ignore, points
GIT_CONFIG_GLOBAL at an empty file so core.excludesFile is genuinely unset (forcing the
fallback branch), seeds a skill, and asserts a file matching the XDG-default pattern still
reads as clean. Reverting the fix (mutating the fallback to resolve to "") turns this test
red with a real pasted failure (see PR body): "expected: <> but was: <?? xdg-fallback-marker>".
Finding 2 (lead, the one that actually needed a code fix): the fallback read XDG_CONFIG_HOME
and HOME straight from the JVM's own environment, not through the gitEnv seam every git
subprocess in this class already honours — so no test could isolate it, and on a machine
carrying a real ~/.config/git/ignore (this dev machine does), every seeding test silently
composed with that real file. Added resolveEnv/resolveHome, which check gitEnv first and
fall back to the JVM's real environment only when the seam doesn't supply a value (production
behaviour, where gitEnv is always Map.of(), is unchanged). Added a hermeticGitEnv() test
helper and routed every seeding test in GitWorktreesTest through it, so no test in the class
can reach the real machine's home directory for this fallback.
Also documents two non-defects the lead asked for one javadoc line each on: the composed
excludesFile is a snapshot taken at seed time, not a live reference to the operator's file;
and excludeSeededSkillsFromGitStatus assumes a fresh worktree (not idempotent, but the
double-seed path does not exist today, so no guard was added for it).
LeadTabScanner.scan() joined labelled tabs to panes with no liveness check at
all, so a tab left behind by a crashed/relaunched lead read as a live lead
forever -- exactly the hazard its own javadoc predicted but excused. It now
cross-checks agent.list, the same signal LeadLauncher already trusted, and
drops any labelled tab with no agent running in it. That alone removes the
duplicate candidates LeadCoordLoop.resolveLocalLead() could pick from,
including the dead one its own WARN's advice (name a lead after
coordinator.selfId) could land a message in.
LeadLauncher never actually stopped the accumulation: relaunching on "0 live"
always created a brand new tab and left the old dead-labelled one right where
it was, so any restart that found 0 live for any reason (a real crash, or a
herdr read that missed a still-running agent) added one more dead tab,
forever. ensureLeads() now closes every dead-labelled tab for a lead as part
of the same reconcile that decides to relaunch, so at most one tab survives
per configured lead once a restart's reconcile has run.
fleetd #361. fleet_list's coordination block now carries this daemon's own
mailbox state and one row per operator-declared peer, each as a tri-state
status (exists / absent / unknown) rather than a boolean. pending and
consumers appear only when status is "exists", so an unresolved probe can
never render as a measured zero.
Verified on this merge, not taken from the worker's report:
mvn clean install -> Tests run: 1394, Failures: 0, Errors: 0, BUILD SUCCESS
Mutation run on merge (isMissingQueue always returns true, which restores
the exact defect the ticket fixes): 4 failures, BUILD FAILURE. The
discriminator's false branch is pinned.
The reviewer's mutation (isMissingQueue always returns true) restored
the exact overstatement fleetd #361 exists to fix -- every declare
failure reading as a confirmed absence -- and still left mvn clean
install green (1389/1389), because no test drove a non-404 shape
through inspect(). The false branch was the whole discriminator
between MailboxState.absent() and MailboxState.unknown(), unpinned.
Widened LeadMailbox.isMissingQueue from private to package-private and
added LeadMailboxIsMissingQueueTest: five hermetic tests (no broker)
covering the true case and all three false shapes isMissingQueue's own
javadoc lists -- a different reply code, a ShutdownSignalException
whose reason isn't a Channel.Close, and an IOException with no such
cause at all (plus an IOException wrapping an unrelated exception
type). Re-ran the reviewer's exact mutation locally: 4 of 5 new tests
went red with the expected assertion messages; reverted, and mvn clean
install is green again at 1394/1394 (1389 + 5 new).
Also added a one-line javadoc note on LeadMailbox.inspect being honest
about which of its two RuntimeException catches is proven by a test
(the createChannel() one, end-to-end against a real broker) and which
stays purely defensive (the declare-site one, for a connection-drops-
mid-call race no test drives on purpose).
core.excludesFile is single-valued, so pointing it at fleetd's own seeded-skill exclude file
with --replace-all at worktree scope was SHADOWING whatever excludesFile the worktree already
resolved (an operator's global config, most commonly) instead of adding to it. This repo's own
.gitignore does not ignore target/ — only an operator's global excludesFile does — so every
worker's `mvn clean install` would make target/ show up as untracked, and CB-576's deliberately
untracked-inclusive hasUncommitted would then read every such worktree as dirty forever, so it
is never cleaned up.
excludeSeededSkillsFromGitStatus now reads whatever core.excludesFile resolves to BEFORE writing
anything (falling back to git's own $XDG_CONFIG_HOME/git/ignore default when the key is unset
entirely, per gitignore(5)), and writes that content into fleetd's own exclude file ahead of the
seeded skill patterns, so every operator-configured pattern keeps applying inside the seeded
worktree. Proven with a new test, seedSkillsComposesWithAnAlreadyEffectiveGlobalExcludesFile,
which isolates a synthetic "operator's global config" via a new gitEnv test seam on GitWorktrees
(GIT_CONFIG_GLOBAL pointed at a throwaway temp file, never the real machine's config) and drives
the real add() path end to end.
Also documents (FleetConfig javadoc + fleetd.example.yaml) that memberSkills copies every
non-hidden subdirectory of its source wholesale, with no per-file allowlist.
Three findings from review of #364, fixed on the same branch:
1. LeadChannel.MailboxState.absent() was returned both for a genuinely
absent mailbox AND for "the probe could not determine anything"
(timeout, unreachable broker, other declare failure) -- exactly the
overstatement #361 exists to fix, one level down. MailboxState now
carries a Presence enum (EXISTS/ABSENT/UNKNOWN) with exists()/known()
accessors; LeadMailbox.inspect classifies a real AMQP 404 (measured
against a live broker, not assumed: an IOException wrapping a
ShutdownSignalException whose Channel.Close reply code is 404) as
ABSENT and everything else as UNKNOWN. fleet_list's mailbox/peer rows
now render a "status" of exists/absent/unknown and only include
pending/consumers when status is "exists", so an unresolved self- or
peer-probe can never render as a measured zero.
2. FleetMcp.probe's get(timeoutMs) left a timed-out inspect() task
running forever on its own virtual thread, holding the AMQP channel
it had already opened -- against a hung (not down) broker this would
orphan one channel per fleet_list call until the connection's
channel-max was exhausted, breaking publish() too. probe() now holds
the Future and calls cancel(true) on timeout/failure so the orphaned
task is interrupted instead of abandoned, and now returns
MailboxState.unknown() (never absent()) on timeout/exception.
3. LeadMailbox.inspect only caught IOException, but createChannel() on
an already-closed connection throws AlreadyClosedException, an
unchecked RuntimeException (measured against a live broker) -- so it
could escape the "never throws" contract. Both places in inspect now
also catch RuntimeException and report unknown().
Tests: MailboxState.exists()/absent()/unknown() call sites updated
across FleetMcpTest/FleetMcpLeadCoordTest; new hermetic tests cover the
tri-state fleet_list rendering (self-probe unknown, a peer that's
absent vs. one that's unknown) and probe cancellation (a LeadChannel
fake that blocks until interrupted, proving probe() doesn't just give
up on it); new @Tag("contract") LeadMailboxTest cases pin the real
exception shapes for both the 404 and the already-closed-connection
paths and prove inspect() reports unknown (never throws) when the
connection is already closed.
Add memberSkills: <dir> to FleetConfig. GitWorktrees#add copies each
skill folder from that directory into <worktree>/.claude/skills/ so a
member spawned against ANY repo — not only one that already ships its
own skills — can load a bridge skill (e.g. implementer). A skill the
target repo already carries is never overwritten.
Every seeded path is hidden from `git status` in that worktree ONLY,
via a --worktree-scoped core.excludesFile pointing at a file under the
worktree's own private git dir (outside the working tree, so it can
never be committed) — not the shared .git/info/exclude, which a linked
worktree resolves to the repo's common git dir and would otherwise leak
visibility changes into the primary checkout and every sibling
worktree. Proven with a real `git status --porcelain` in
GitWorktreesTest, not by reasoning.
Seeding is best-effort like the existing overlayParity/isolateToolSurface
steps: a missing/unreadable source or a copy/exclude failure is logged
and skipped, never fails the spawn. memberSkills is triaged as a
DEFERRED config key in ConfigRef (baked once into GitWorktrees at
startup, like worktreeGroup), with its own changedDeferredKeys branch
and coverage-test entries.
Lead-to-lead AMQP coordination had a send half with tools and a receive
half without. This closes three blind spots:
- LeadChannel gains inspect(coordId) -> MailboxState(exists, pending,
consumers), implemented in LeadMailbox with a throwaway probe channel
(never the long-lived publish/consume channels) so a passive-declare
404 on a missing queue can never take down publish() on the same
instance.
- FleetConfig.Coordinator gains peers: List<String> (defaults to empty,
blank entries dropped) so a daemon can declare which peer coord-ids
it expects to reach.
- fleet_list reports coordination state via a new CoordinationSource
(own coord-id, own mailbox state, held messages as msgId/from/preview
only, and one row per configured peer with reachability/pending/
consumers), following the existing OutageSource/QuarantineSource
"Source record with none()" idiom instead of growing listFleet's
overload chain by another positional parameter. Every peer probe is
bounded by a 1.5s timeout on a virtual-thread pool and degrades to
absent rather than ever slowing or failing fleet_list.
- fleet_send{coordId}'s success text now says "durably confirmed by the
broker" instead of "delivered", and warns (while still reporting
success) when the target mailbox has zero consumers attached.
Tests: hermetic unit tests for exists/absent/zero-consumer/old-config-
no-peers-key/new fleet_list shape using FakeLeadChannel, plus a
@Tag("contract") LeadMailboxTest.inspectingAMissingMailboxNeverBreaks
PublishOnTheSameInstance proving the invariant against a real broker.
CB-527 shipped a Claude Code plugin and a marketplace in this repo. Nothing in
CLAUDE.md or docs/ ever named it, so a later session planned the same feature
from scratch. The wiki Features entry existed and was correct, but wiki/ is a
submodule whose pointer is never advanced, so no session reads it.
Visibility:
- CLAUDE.md addendum now names plugin/ and both structural limits, so every
session sees it. This is the change that stops the rebuild happening again.
- wiki/11-Features.md records the rename and why the entry alone was not enough.
Drift (each measured against the code, not assumed):
- mount name fleetd -> fleet, matching PeerLauncher.MCP_MOUNT_NAME. The old name
gave a lead with both a project .mcp.json and the plugin two mounts of one
daemon and a duplicated fleet_* tool set.
- url is now ${FLEETD_MCP_URL} instead of a hardcoded address, so one plugin can
serve hosts running the daemon on different ports. Plain ${VAR}, the form
kb-alms proves works here; ${VAR:-default} is untested and not used.
- plugin claude-bridge -> fleet, marketplace claude-bridge -> fleetd, version
0.2.0. Breaking for a 0.1.0 install: mcp__fleetd__* becomes mcp__fleet__*.
- README install path ltms/claude-bridge -> the fleet/fleetd remote.
- the setup skill's §5 told operators to pin primary.terminal:. CB-579 replaced
that with fleet.leaders.*.tab. Replaced, with the duplicate-tab warning (#359).
Scope: the plugin is lead-side only, and cannot be otherwise. The launcher adds
--agent only when <worktree>/.claude/agents/<role>.md exists in the member's own
tree (ClaudeCodeLauncher.java:371,391), and a member's CLAUDE_CONFIG_DIR points
at its profile's config dir (ClaudeCodeLauncher.java:285), so a member never
reads the operator's plugin store. On this Mac all four Claude profiles set
configDir, and the four ccs instances hold four separate copies of the plugin
store -- same md5, different inodes. Seeding member skills through the worktree
is #362 scope item 3, implemented separately.
Note for anyone verifying a plugin: `claude plugin validate` does NOT read
.mcp.json. Replacing it with `{ this is not json at all` still passes, exit 0.
Refs #362, #359
The gate counted ERROR lines since RESTART_MARK. On a laptop that idle-sleeps after one
minute on battery that meant 6 ERROR lines for an AMQP link that recovered every time,
and a gate that cries wolf is a gate nobody reads.
It now reports three states: no errors; only errors proven to have recovered (quiet, and
the gate passes); anything else (the old warning, unchanged). Attribution is per
connection, using the names #356 put into the log -- a lead-mailbox recovery can no
longer clear an unrecovered reply-inbox reset. A candidate carrying neither name is
unattributable and stays LOUD.
Two earlier rounds were rejected. Round 1 was inert: it matched nothing in the real log,
because the layout abbreviates the logger and 'Connection reset' sits in the stack trace,
not on the ERROR line -- my brief had pointed the worker at fleetd.out, which is untracked
and so absent from its worktree. Round 2 was correct and honest but could not attribute
anything, which is what motivated #356.
Verified on merge beyond the worker's own mutations:
- ran the classifier against the REAL log, which is still in the pre-#356 format: 6 total,
0 recovered, 6 unexplained. Old-format lines carry no connection name, so they stay loud
-- the safe direction, on genuine data rather than a fixture.
- adversarial fixture the worker did not write: a lead recovery BEFORE any failure banks
no credit; 2 inbox resets with 1 recovery leaves 1 unexplained; a non-AMQP ERROR stays
loud. total=3 recovered=1 unexplained=2, as intended.
- RESTART_MARK still anchors the scanned region.
Caveat carried from the PR: the patterns are source-derived. The daemon has not been
redeployed, so they are not yet confirmed against a live log.
Both connections were already named at newConnection() -- 'fleetd-reply-inbox' and
'fleetd-lead-mailbox' -- and neither name ever reached the log: 0 occurrences in
fleetd.out, and both connections logged under the same thread name
'[AMQP Connection 10.10.20.13:5672]'. So when one of the two died and never came
back, the log could not say which.
AmqpConnectionFailureLogger extends DefaultExceptionHandler and overrides only the
protected log(String, Throwable) sink that every handle* method calls virtually, so
the identity is added without changing any handler action.
My brief caused a defect here and the correction is the interesting part. I told the
worker the client 'currently uses ForgivingExceptionHandler', read off the log line
c.r.c.i.ForgivingExceptionHandler -- which names where the LOGGER FIELD is declared,
not the instance's class. javap on the jar shows ConnectionFactory's constructor does
'new DefaultExceptionHandler', and DefaultExceptionHandler extends StrictExceptionHandler
extends ForgivingExceptionHandler. The first version therefore extended the base and
silently dropped strict channel-closing on four listener/consumer paths. Now pinned by
a type assertion on both factories plus a behavioural test that handleConsumerException
still closes the channel once.
Verified on merge with a mutation the worker did not run: it mutated the parent class,
so I mutated the copied private-static isSocketClosedOrConnectionReset in the DANGEROUS
direction (always true => every failure logs at WARN and vanishes from the redeploy
gate's ERROR count). Caught: 'inbox failure line ==> expected: <ERROR> but was: <WARN>'.
Merged main in first; the auto-merge compiled. 1379 green, unpiped.
Test-only. FleetConfig.java itself is unchanged.
The hazard is the back-compat constructor ladder (21/20/18/17/16/15/14 alongside the
22-arg canonical). Add a component and leave withDefaults()'s call at the old arity and
it binds to a back-compat constructor: it compiles, the suite passes, and the new key is
silently defaulted away on every load().
Verified on merge with a mutation the worker did not run: I made withDefaults() issue a
21-arg call, reproducing the real binding rather than an explicit null. It compiled, and
the guard failed by name -- 'memberLoginShell: ... a component silently dropped by
withDefaults(), the shape of the defect this test exists to catch'.
Exclusion list is empty and its size is pinned, so a future exemption must touch an
assertion rather than grow quietly.
Adding a component to FleetConfig follows an established pattern: the
record grows by one arg, and a back-compat constructor is added at the
OLD arity so existing callers keep compiling. That back-compat
constructor also silently captures withDefaults()'s own literal-arity
'return new FleetConfig(...)' call the next time this happens, since
that call is now a legal overload match too. It compiles, every other
test passes, and the new component is defaulted away on every load().
This is not hypothetical - it happened live while building the (now
parked) idle-sleep-guard PR, caught only because that branch's own new
tests asserted on the new field.
Add a reflective test that builds a FleetConfig through the true
canonical constructor (resolved by record-component types, not arg
count - the same pattern ConfigRefTopLevelReportingCoverageTest already
uses in this file) with a real, non-null value in every component, runs
the real withDefaults(), and asserts every value survives unchanged.
Never hardcodes the arity - it enumerates
FleetConfig.class.getRecordComponents() - so it keeps working as the
record grows. No back-compat constructor is touched or removed.
ask()'s TimeoutException catch used to run clearAsyncQuestion(turnId, true) -- forgetting the
Task's asyncTasksByTurn mapping -- before rendezvous.closeAsk(turnId) ran in the shared finally.
Between those two calls the ask was still "answerable" (askSession(turnId) non-null) but the Task
mapping was already gone, so a racing answer() call found task == null, skipped
finishAsyncTask, and stranded the async ticket at PENDING even though answer() itself reported a
result. #329 fixed one step of this same race; this closes the remaining one.
The fix reorders the fresh owner's teardown: closeAsk runs first, then markAskTimedOut and
clearAsyncQuestion. A racing answer() call now either sees the ask still open (and the Task
mapping guaranteed intact) or sees it already closed (STALE_TURN, before it ever reaches
asyncTasksByTurn). It also gates the whole block by ticket.fresh(), matching the invariant the
finally block already states ("only the fresh owner tears down the shared turn") -- a duplicate
coalesced ask() timing out no longer forgets bookkeeping the fresh owner still needs.
Adds aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket, which pins the exact window
with a new test-only hook (askTimeoutRaceHookForTest) and proves both invariants: a late answer()
racing the timeout sees STALE_TURN, and the async ticket still resolves DONE from the worker's
real reply. Reverting the reorder (verified locally, not committed) makes this test fail with
"expected STALE_TURN but was TIMED_OUT_WORKING".
Site 1 (abandon()'s matching loop, reachable): the recovery/put-back branch calls
inbox.publish, which AmqpReplyInbox implements as a real broker round trip that
throws IllegalStateException on an unroutable/unconfirmed/interrupted publish.
An uncaught throw there aborted the loop, stranding every task after it in
`matching` PENDING forever. Fixed by recording each task's own future.complete()
result before any cleanup runs, then wrapping the cleanup in try/catch so one
task's failure cannot stop its siblings from getting their outcome. Reaching the
throwing branch by real timing needs a race the file's own #137 follow-up already
found unreachable through the public API, so the reproducing test uses a
test-only hook (same technique as the existing fleetd #324/#329 hooks) to inject
the throw at that exact point.
Site 2 (sendAsync's task.future.whenComplete, reachable): the returned stage is
discarded, so an uncaught throw from pushLoop.onTicketTerminal vanished with no
log line. Reproduced for real: Fleetd's shutdown hook runs messages.close()
(stops the async executor from taking new work, but does not cancel a send
already in flight) before pushLoop.close() (shuts its scheduler down
immediately) — a ticket completing in that window makes onTicketTerminal's own
scheduler.schedule(...) throw a genuine RejectedExecutionException. Fixed with a
try/catch(Throwable) plus log.error inside the whenComplete action.
Site 3 (the two `finally { asyncTasksByWaiter.remove(reply); rendezvous.close(...);
}` blocks in send() and answer()): read Rendezvous.close/closeAsk and the
ConcurrentHashMap operations behind them — both are plain map ops on a non-null
key with no user-overridable code, so neither can throw. Left unchanged; not a
defect.
Mutation-proven: reverting either fix reproduces the failure it exists to catch
— removing site 1's try/catch aborts abandon() with the injected exception
(MessageServiceTest#aPerTaskCleanupFailureDoesNotStrandTheRemainingMatchingTasks
errors); removing site 2's try/catch leaves the RejectedExecutionException
unlogged (MessageServiceTest#aTicketTerminalPushFailureDoesNotVanishSilently
fails its log assertion). Full suite: mvn clean install, Tests run: 1365,
Failures: 0, Errors: 0, BUILD SUCCESS.
HerdrPeerLauncher.stop() used to gate spaces.locatePane() on usesTabPlacement(),
which reads the delegate's OWN configured profiles. When CompositePeerLauncher's
single-daemon stop() shortcut hands a pane to a delegate that never spawned it
(spawnedBy empty after a daemon restart, herdrDaemonCount()==1), that delegate's
placement config says nothing true about how the pane was actually placed, and a
dedicated tab could be skipped and leaked.
Resolve the tab unconditionally instead — WorkspaceControl#locatePane already
tolerates a missing pane by returning null — and let the existing single-occupant
check (tabPaneCount()==1) be the only thing that decides whether to close it, same
as it already protects a shared tab regardless of declared placement.
Adds a mixed-placement CompositePeerLauncherTest (every existing stop-fallback test
configured both adapters as tab placement, so the mis-routing never showed) and
updates FleetAppTest#stopWorkerInPanePlacementClosesOnlyThePane, whose old
assertion (no pane.get on pane placement) documented exactly the skip this fix
removes.
#339 stopped a member's own prose about an error from recording a credential
outage, by requiring the pattern at the start of its matched line. A bare
lookingAt also rejected a genuine error line rendered as
| 503 Service Unavailable: upstream credential rejected
The send still failed, but the outage was never recorded. That is the false
negative #339's own invariant 3 named as worse than the false positive it set
out to fix: an unrecorded outage leaves the fleet spawning into a dead
credential.
Measured with a throwaway probe on the raw-scrape path, whose own comment says
to expect leading chrome there: kind=FAILED, sinkNotified=0.
startsWithBackendError now skips a leading run of non-letter, non-digit
characters before the check. That keeps #339's intent: prose still does not
match, because there the pattern sits after words rather than after chrome.
The worker's own prose test still passes.
Mutation: restoring the bare lookingAt fails the new test.
The fix reported every distinct unprotected name, but nothing held it there.
Mutation: replacing .filter(unprotectedGapNamesWarned::add) with a filter that
adds and always returns true - so every name is logged on every spawn - left
all 1358 tests green. The Set behaved; nothing proved this class used it as a
guard rather than as a record.
Two tests added:
- theSameUnprotectedNameIsWarnedAboutOnlyOnceAcrossSpawns pins invariant 1, the
noise control. It now fails on that mutation, showing both duplicate WARNs.
- anAllowListWarnDoesNotSuppressALaterDenyByDefaultWarnForADifferentName covers
the reverse policy order. The defect was found going deny-by-default then
allow-list; a guard fixed in one direction is not fixed in the other.
ConfigRefTopLevelReportingCoverageTest (added by #333) proved every COLD_KEYS
and SPLIT_KEYS member has a real comparison behind it, but left
DEFERRED_TOP_LEVEL_KEYS unexercised. Re-measured by mutation (drop each
key's branch from changedDeferredKeys, run the suite, restore): 6 of the 11
deferred keys had no behavioural test naming them — guard, leadHeartbeat,
worktreeRoot, spawnReadyTimeoutMs, spawnReadyPollMs, quarantineCooldownSeconds
— which corrects the issue's own guessed list in two ways: lifecycle is
actually covered (ConfigRefTest.aDeferredChangeIsAppliedAndReported), and
worktreeRoot was missing from the issue's list entirely.
Promoted the test-side DEFERRED_TOP_LEVEL_KEYS copy into ConfigRef.DEFERRED_KEYS
(package-private, alongside COLD_KEYS/SPLIT_KEYS) so the reflective test reads
the same set changedDeferredKeys is compared against, and made
changedDeferredKeys package-private so the test can call it directly. Every
DEFERRED_KEYS component turned out to be a scalar or a simple record, so no
exclusion set was needed.
Mutation proof: dropping guard's branch from changedDeferredKeys leaves the
whole suite green except the new
everyDeferredKeyIsActuallyReportedByChangedDeferredKeys test, which fails
naming guard exactly.
unprotectedGapLogged was one AtomicBoolean guarding two WARN branches in
logCredentialGap that name different env var names (the allow-list
keptByDerivedList branch, and warnGapUnprotected's deny-by-default /
non-zsh-fallback branch). memberCredentials is a live, re-read-per-spawn
supplier, so between two spawns a policy reload can change which names are
in the gap: spawn 1 warns about name A and trips the shared flag, and
spawn 2's gap containing a different name B never gets its WARN.
Replace the AtomicBoolean with unprotectedGapNamesWarned, a
ConcurrentHashMap-backed Set<String> guard keyed per name (same shape as
OpenCodeLauncher.modelCheckSkippedWarned), so each distinct credential-shaped
name is warned about exactly once, ever, regardless of which branch or
which spawn first reports it. allowListGapLogged (the separate INFO guard,
#192) is untouched. Neither WARN's wording changed.
memberHerdrSocket was missing from the prose. The #333 worker spotted it and
correctly left it alone as outside its scope.
Fixed by pointing the bullet at COLD_KEYS instead of re-listing its contents,
so the prose and the set cannot drift apart a second time.
F1: fleet: was sitting in ConfigRefTopLevelCoverageTest's HOT_EXCLUDED_TOP_LEVEL_KEYS
escape hatch, even though fleet.leaders is read only at startup (LeadTabScanner's
identity map, LeadLauncher.ensureLeads) while the rest of fleet: (role pools,
charters, tabLabel) is live. A reload changing only fleet.leaders reported a bare
"config reloaded" -- the operator edits a lead's tab: label, sees the reload
succeed, and the pane keeps resolving as a worker. Moved fleet into
ConfigRef.SPLIT_KEYS; changedSplitKeys now compares fleet.leaders specifically
(not the whole Fleet record, which would over-claim "restart" for a tabLabel-only
change) and names both halves in the message.
F2: membership in SPLIT_KEYS/COLD_KEYS never proved a matching branch existed in
changedSplitKeys/changedColdKeys -- measured by dropping the coordinator branch
while leaving "coordinator" in SPLIT_KEYS: both ConfigRefTopLevelCoverageTest and
the in-method "kept in step" assert stayed green. Added
ConfigRefTopLevelReportingCoverageTest, the ConfigRefProfileCoverageTest mechanism
one level up: reflection-built FleetConfig pairs that differ in exactly one
top-level component, calling the real (now package-private) changedColdKeys/
changedSplitKeys to prove each COLD_KEYS/SPLIT_KEYS member is actually reported.
Scoped to split+cold, not deferred -- see the new test's javadoc for why and what
that leaves open.
Both findings carry a behavioural test in ConfigRefTest plus a mutation proof
(revert -> real failure -> restore) recorded in the PR description.
The comment that landed with #329 said a null task means the turn was never
an async ticket. That is wrong, and it makes the guard read as complete.
A genuine async ticket also reaches answer() with task == null. ask() runs
clearAsyncQuestion(turnId, true) in its catch block, which drops the
asyncTasksByTurn entry, while rendezvous.closeAsk(turnId) runs later, in its
finally. Between the two the ask is still answerable and the map entry is
already gone, so answer()'s lookup returns null and the ticket is stranded.
Measured with a throwaway probe firing only that first half: answer() reported
REPLIED while the ticket stayed PENDING with a null reply. The probe used
forgetTurnForTest, so it omits markAskTimedOut; that cannot change the outcome,
because askTimedOut is read only by askAnsweredAsyncTasks, which reply() never
reaches while answer()'s own waiter is live.
Open as fleetd #334. The comment now says so.
F2 (sendAsync executor catch): log when completeExceptionally returns
false, so an exception thrown after finishAsyncTask already completed
the ticket's future is no longer silently lost.
F1 (answer()'s stranded async ticket): reuse the Task reference answer()
already looked up before rendezvous.answerAsk(), instead of a second
asyncTasksByTurn lookup by turnId in finishAsyncTask. The second lookup
raced ask()'s unlocked timeout cleanup, which could forget turnId first
and leave the ticket stuck PENDING even though answer() itself returned
REPLIED. The #282 chained-ask guard is unaffected: it is still keyed on
result.outcome() == QUESTION, not on this lookup. Removed the now-unused
finishAsyncTask(String, Reply) overload.
F3 (reply()'s orphan recovery path): read orphan.turnId once instead of
twice, closing the same double-read shape fleetd #324 fixed in
finishAsyncTask.
Each fix has its own test plus a test-only race hook (mirroring #324's
finishAsyncTaskRaceHook) to force the exact interleaving deterministically.
Mutation-tested each fix by reverting it, confirming the real failure
(swallowed exception / PENDING ticket / NullPointerException), then
restoring it.
mvn clean install: Tests run: 1345, Failures: 0, Errors: 0, Skipped: 0,
BUILD SUCCESS.
Unit 1: ConfigRef gets a fourth reload class, `split`, for keys read both
off the startup snapshot and live off config.get() at different sites
(health:, coordinator:). A split change is accepted (Outcome.applied()
stays true) and reported by name, naming which half is live and which
needs a restart, via a new Outcome.split() field kept separate from
deferred() since the two carry different guarantees for any caller that
branches on them, not just prose in summary(). Class doc updated: four
classes now, denominator note no longer calls health/coordinator
undecided.
Unit 2: ConfigRefTopLevelCoverageTest enumerates FleetConfig's 22
top-level record components and requires each to sit in exactly one of
COLD_KEYS, a pinned "compared in changedDeferredKeys" set, SPLIT_KEYS, or
a pinned hot-exclusion escape hatch — printing its own denominator and
pinning the escape hatch's exact contents the way #323 asked for.
Deviates from the issue's starting values by one key: `profiles` moves
from the suggested Hot bucket into the deferred bucket, because
changedDeferredKeys demonstrably compares it (add/remove and launch
settings), and citing "read live off the config supplier" for the whole
key would be false — most Profile fields are not read live, only
weight/maxLoad/credentialId are (and those are already covered by
ConfigRefProfileCoverageTest). Cold=5, split=2, deferred=11, hot=4,
total=22 — verified against the record and against ConfigRef's code, not
copied from the issue.
The merged javadoc said a changed primary.terminal leaves a lead 'unresolved as
primary until a restart'. That over-claims. CB-532 made the pin deprecated:
identity comes from leaders:/leadScan:, and Fleetd.java:511 warns about the pin
at startup. A changed pin still needs a restart, but for the fallback nudge
destination, the deprecated identity path, and pushReminders/pushBackoffMs -
not for a lead that uses leaders:.
Also record what I measured. FleetConfig has 22 top-level components; four are
named nowhere in ConfigRef. memberCredentials and memberLoginShell are hot and
correctly absent (both read live off config.get() at spawn). health and
coordinator are undecided, not hot. 'Absent' looks the same for both kinds, and
twice now the forgotten kind hid among the correct kind.
ConfigRef.changedDeferredKeys only classified seven top-level FleetConfig
keys (#323 fixed the profile side). Two more keys are read only off the
startup snapshot and were missing:
- primary: Fleetd.java:506/519/520 feed PrimaryRegistry and ReplyPushLoop
at construction; neither is rebuilt on reload.
- configReload: Fleetd.java:679-680 decide once at startup whether to
build a ConfigWatcher at all, and with what interval; the watcher that
would apply a later change is itself built once, so it is deferred
(not cold — no already-open resource goes inconsistent, a running
watcher just keeps its original settings).
health and coordinator are deliberately left unclassified: both are read
both off the startup snapshot AND live off the config supplier at a
second call site, so no single bucket is correct for either — see the
PR body for the options writeup and the coordinator.uriEnv exposure
question the issue asked to be answered.
Each fix is proven with a failing-first test in ConfigRefTest and a
revert-quote-restore mutation check (see PR body for the transcripts).
answer() holds sessionLocks while finishAsyncTask reads the volatile Task.turnId twice — once to
check it is non-null, once as the ConcurrentHashMap.remove key. ask()'s own timeout path mutates
the same field with no lock, via clearAsyncQuestion(turnId, true). volatile makes each read fresh
but not the pair atomic, so the field can go null between the two reads and remove(null, task)
throws NullPointerException on the lead's own answer() call, even though the reply already
completed on the line above.
Capture task.turnId into a local once and use that for both the check and the removal.
Added a package-private test seam (finishAsyncTaskRaceHook + forgetTurnForTest) so a test can force
the exact interleaving deterministically, by running the identical clearAsyncQuestion(turnId, true)
cleanup ask() uses, at the point between finishAsyncTask's former two reads. Both are inert (null)
in production.
ConfigRef.sameLaunchSettings' javadoc claimed it compares every component
the launcher reads at spawn. It missed ideProjectDir, ideOpenCommand and
autoCompactWindow, and changedDeferredKeys separately missed worktreeGroup
(baked into the same GitWorktrees as worktreeRoot, Fleetd.java:251). A
reload that changed only one of those keys reported "config reloaded" with
nothing deferred, and the running daemon kept the old value.
Fix the four instances, and add ConfigRefProfileCoverageTest: it enumerates
every FleetConfig.Profile record component by reflection, mutates each one
not in the new ConfigRef.LAUNCH_SETTINGS_EXCLUDED set on a base profile,
and asserts sameLaunchSettings actually notices — so a fifth missed field
fails the build by name instead of drifting silently. It also prints its
own denominator (26 components, 23 compared, 3 excluded) per the ticket's
requirement that a checker must be able to state what it checked.
Also add the `profile` field itself to the comparison (it was neither
compared nor excluded before this fix — the coverage test surfaced it).
Rewrote the sameLaunchSettings javadoc to describe what the coverage test
actually guarantees instead of repeating the unchecked claim.
SessionManager.releaseRemoved read hasUncommitted() once, while the worker
could still write, then used that stale boolean after launcher.stop() to
authorise `git worktree remove --force`. The same stale read also gated
trySnapshot, so a worker that wrote between the read and the stop lost its
work with neither a preserve nor a snapshot.
Add a second, best-effort hasUncommitted read immediately before the
removal, taken only on the path that is actually about to delete something
(never on a release that already decided to preserve, and never for
SHUTDOWN, which preserves unconditionally). If the tree is now dirty,
preserve it and attempt a fresh snapshot, since the original snapshot never
ran when the pre-stop read said clean. A failing re-check also preserves,
matching the existing CB-581 fail-safe rule.
AmqpReplyInbox.release() used held.remove(target) then iterated the old
map. A delivery landing on the consumer work-pool thread after the
remove (basicCancel does not flush one already handed to that pool) hit
deliverCallback's computeIfAbsent, found the key gone, and created a
brand-new map release() never looks at again — delivered-but-unacked
forever, never requeued, never redelivered (#298 only closed the
"already in held when release runs" case).
Fix: release() swaps in a RELEASED tombstone via held.compute(...)
instead of held.remove(...). ConcurrentHashMap serializes compute/
computeIfAbsent calls for the same key against each other, so whichever
of release() and a concurrent deliverCallback runs first is fully
visible to the other — no gap. deliverCallback checks for the
tombstone and nacks-with-requeue instead of recreating a map; peek/ack
treat it as empty; own() clears a stale tombstone so a target is never
poisoned if its id is ever reused (the issue's own text says id reuse
doesn't happen, but the tombstone would otherwise sit in `held` forever
either way).
New test AmqpReplyInboxReleaseRaceTest forces the actual interleaving
with a latch (blocks release() inside its nack loop, which is only
reachable after the tombstone swap, then fires a concurrent delivery)
rather than a sequential call — a sequential test would not have caught
this, since #298's own contract test forces settlement before release()
runs. Mutation-tested: reverting the fix makes this test fail with
"expected: <2> but was: <1>" (m1 never nacked); restored after
confirming that failure.
ConnectionIdentity.resolve() called pids.pidForLocalPort(remotePort),
which returns -1 both on a real failure and (silently, no log line)
when lsof just finds no matching process. terminalForPid(-1) then
matches no pane, so CallerResolver's loopback-trust fallback could not
tell that caller apart from a genuine primary and handed it
Principal.primary(...) — granting SPAWN, STOP, SEND and DRAIN to a
worker whose PID lookup failed. This is the escalation PaneLocator's
own javadoc already names; CB-161's ancestry walk only helps once a
candidate pid exists, and a failed lookup has none.
Fix: ConnectionIdentity.Caller gets a resolved() predicate (pid > 0),
centralised next to the -1 sentinel it tests for the same reason
isLoopback() is centralised (fleetd #305: two independent copies of
one rule already drifted once). CallerResolver's loopback-trust
fallback now requires c.resolved() before granting PRIMARY; an
unresolved caller gets Principal.anonymous() — the same already-tested
"authenticated as nothing" outcome used everywhere else in that
method, so the refusal is a clean, named, unsurprising result rather
than something that looks like a bug.
Also logs the previously-silent "lsof ran clean, found no match" case
in LsofPeerPidLookup at DEBUG, since that (not a slow lsof — the
waitFor result was already discarded) is the likelier real trigger.
loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary is untouched
and still green: a real pid that owns no pane (the actual primary) is
still resolved() and still PRIMARY. Token mode is unaffected — it
never consults c.pid() at all.
Mutation-tested: reverting only the CallerResolver.java guard
reproduces the escalation exactly (aFailedPeerPidLookupIsRefusedNotPromotedToPrimary
fails with "expected: <ANONYMOUS> but was: <PRIMARY>").
FixedPlacementPolicy's class javadoc still opened with "This ignores caps
and reachability" after the previous commit added reachability as the
fourth carve-out that is explicitly NOT ignored — caught by a shape-check
survey run against this same file as part of #315's own request ("look in
placement/ ... for the same shape: a caller/comment that documents an
expectation ... where an implementation does not meet it"). Reworded the
opening sentence: fixed still ignores caps (maxLoad) by design, but
reachability is now a narrower, per-call retry exclusion, not an ignored
concern.
CompositePeerLauncher.spawn retries a failed candidate on the next one and
rebuilds PlacementContext "so the policy excludes this profile" (its own
comment), but FixedPlacementPolicy.select never read ctx.unreachable(). Under
the default `fixed` placement policy (used when `placement` is unset or set
to `fixed`), every retry re-picked the same dead default and a second,
healthy, configured profile was never tried. This also covers the wiring-bug
branch (a candidate profile with no owning adapter), which hit the exact same
symptom for the same reason.
Not live on this fleet: fleetd.yaml sets placement: weighted, which already
consults ctx.unreachable() via PlacementPolicyUtil.available(). This is live
only for a deployment that leaves placement unset or sets it to fixed.
Fix is in FixedPlacementPolicy: consult ctx.unreachable() in the same two
places it already consults quarantined/coolingOff (the default check and the
fallback walk over candidates()), and add a fourth reason to the "no
candidate remains" exception. Considered fixing this in
CompositePeerLauncher's retry loop instead (break when select() returns an
already-unreachable profile), but that only fails faster on the same dead
profile — it cannot make the loop advance to a different candidate, because
only the policy decides which candidate is next. The defect is that one
policy implementation does not honor the loop's stated contract, so the fix
belongs in that policy, matching how weighted/round-robin already behave.
Also fixed: the "no reachable worker profile" exception message said
"trying N candidate(s)" where N was unreachable.size(), a count of DISTINCT
profiles (a HashSet dedupes a profile added twice), under wording that reads
as a count of attempts. Reworded to "N distinct candidate(s)" so the count
matches what is measured and the profile list that follows it.
Tests: two new failover tests next to the three existing ones in
CompositePeerLauncherTest (which all use PlacementPolicies.weighted(), which
is why this had no coverage) — one pinned to PlacementPolicies.fixed() for
the unreachable-default case, one for the wiring-bug (no adapter) case.
Mutation-proofed: reverted FixedPlacementPolicy.java, both new tests failed
with the exact bug ("no reachable worker profile available after trying 1
distinct candidate(s): a" / "...c"), then restored the fix.
MessageService.reply()'s async-recovery path (askAnsweredAsyncTasks)
required a live Task.turnId, but ask()'s own TimeoutException handler
calls clearAsyncQuestion(turnId, true) — deliberately forgetting turnId
so hasAsyncQuestion() stops reporting the target BUSY. That made a
worker's eventual real fleet_reply, after an unanswered fleet_ask, fall
through to the inbox: fleet_poll{ticket} stayed PENDING forever and was
later force-failed with the false reason "session released before it
replied".
Fix: a new Task.askTimedOut marker is set (markAskTimedOut) right
before the turnId is forgotten, and askAnsweredAsyncTasks accepts it in
place of a live turnId. The marker never touches asyncTasksByTurn, so
the BUSY-release behaviour (invariant 1) is untouched. The existing
ambiguity guard (candidates.size() > 1 -> inbox, never guess) still
applies unchanged, but is now genuinely reachable rather than pure
defence in depth, since an ask timeout frees its target for a fresh,
independent delegation — the affected javadocs are updated to say so.
Tests: MessageServiceTest.aReplyAfterAnAskTimeoutStillCompletesTheAsyncTicket
(positive, mutation-proven) and
.twoAskTimedOutTicketsOnOneTargetFallBackToTheInboxRatherThanGuess
(negative/ambiguity). FleetMcpTest's
unansweredAsyncAskReturnsTheTicketToPending was renamed and its final
assertion updated — it had pinned the old (buggy) inbox-stranding
behaviour as expected.
The compare-and-release declines silently. This race is unobservable by
construction, so a reaper that quietly stops reaping is the hardest kind of
behaviour to diagnose later. One debug line names the pane and the likely
cause.
drainAll iterated a one-shot registry snapshot with nothing to refuse a new
fleet_spawn while the drain was still running (mcp.close() only runs 8 calls
after sessions.close() in the shutdown hook). A session registered in that
window was never visited by the drain loop: its pane kept running and its
worktree was never preserved, with the in-memory registry gone at exit.
Fix, both mechanisms as the issue asked for (neither alone is complete):
- SessionManager.acquire now checks a `draining` flag, flipped true at the
very start of drainAll before the registry snapshot is even taken, and
throws the new ShuttingDownException (invariant 3: fail loudly, say why).
FleetMcp.spawn and FleetApp.spawnMember surface it as a clean error/503
rather than an uncaught RuntimeException.
- The flag alone cannot close the whole race: a caller already past the
check can still be mid-launcher.spawn() (a real herdr round trip) when
drainAll snapshots the registry. drainAll now re-reads the registry once
its main pass finishes and drains whatever straggler landed there too,
bounded by the SAME whole-drain deadline (invariant 1: timeoutNanos stays
a budget for the whole drain, never extended for a straggler).
- ReleaseCause.SHUTDOWN still preserves worktrees for both the initial pass
and the sweep (invariant 2, unchanged release() path).
Tests (SessionManagerTest): a guard test proving acquire() throws once
drainAll has started, and a race test using a launcher double that blocks
the second spawn() and the first stop() call to force, deterministically,
the exact interleaving where a spawn passes the guard before drainAll flips
it and only registers after the initial snapshot — proving the post-loop
sweep catches it.
Shape check (SessionManager.java only, not fixed): reapIdle has the same
shape — a decision made from a roster() snapshot, then acted on via
release(s.paneId()) with no re-check of the session's current state.
Four latches gate delivery in Injector, and only awaitingCompletion had a way
out of a sustained unknown streak. CB-109 added that escape because a worker
stuck in a state herdr cannot classify never produces a working->idle
boundary. The same is true during post-turn housekeeping, but the escape was
never extended there.
awaitingPostTurnPickup and postTurnObserved are both released only on an
injectable sample, so a worker that goes unknown and stays there wedges: the
target is polled forever, every later message to it is blocked by the delivery
gate, and no onTurnFailed fires, so the session sits at DONE and looks healthy.
The counter did not even increment, since ++unknownSinceTurn sits inside the
awaitingCompletion short-circuit.
postTurnPending needs no escape; it is cleared on the line after the listener
call that sets it.
The escape does not set turnFailed. The delegated turn already completed and
its waiter already resolved — what is outstanding is the /clear. Failing the
turn would drive SessionManager.onFailed on a session that genuinely finished.
Not reachable in the live configuration: the path needs lifecycle.clearAfterTurn,
which fleetd.yaml does not set. It becomes reachable as soon as anyone turns
that supported knob on.
Fixes#306
ConnectionIdentity and CallerResolver each kept their own isLoopback. They
drifted: the identity resolver accepted only 127.0.0.1, the authorization
check accepted all of 127.0.0.0/8.
A caller from 127.0.0.2 therefore had its identity resolution skipped, so it
carried no terminal, and CallerResolver reads a missing terminal as "not a
worker" — which under loopback-trust, the default mode, is the primary. A
worker got spawn, stop, send and drain. The skip also happens before the PID
ancestry walk, so that defence is bypassed too.
Being strict in ConnectionIdentity was not the safe direction. That predicate
decides whether identity is resolved at all, and resolution is what demotes a
worker, so every address it excluded was one where a worker became the lead.
Measured, not assumed: on Linux the whole 127.0.0.0/8 is bound to lo, and
binding a source of 127.0.0.2 on the fleet host succeeds (curl rc=7, the
connect refused rather than the bind). On macOS the source bind fails (rc=45),
so this workstation was never exposed.
The shared predicate also accepts the IPv4-mapped IPv6 form, which neither
copy handled. That one failed in the safe direction: a primary on
::ffff:127.0.0.1 was refused as anonymous.
No transport-level test binds a real 127.0.0.2 source — it cannot run on
macOS. The reasoning is recorded on the issue.
Fixes#305
POST /members and DELETE /members/{paneId} were the two routes in FleetApp
with no catch (HerdrException). FleetApp has no Javalin exception mapper, so
the exception escaped as the default 500 with the body "Server Error" — no
herdr code, no herdr message. fleet_spawn and fleet_stop catch the same
exception and report a named error, so this was the same one-door-guarded
shape as #297.
Both now go through the existing herdrError mapper: 404 when herdr says the
target is gone, 502 otherwise. That is an answer the caller can act on.
stopMember matters more than spawnMember. SessionManager.release deregisters
the session, notifies the release listener and preserves a dirty worktree
before it calls launcher.stop, so a throw from that stop arrives after the
teardown the caller asked for has already happened. A bare 500 told the caller
to retry and carried nothing to explain what went wrong.
The two existing tests that asserted 500 now assert 502 and check the error
body. Neither was about the status code: one guards that a failed teardown is
not reported as a successful 204, the other that a failed spawn still closes
its tab. Both properties are unchanged.
Fixes#304
MessageService.reply now throws on blank content. FleetMcp.reply guarded only
against null, and its handler is a bare BiFunction with no try/catch, so a
whitespace-only fleet_reply left the handler as an uncaught
IllegalArgumentException instead of the clean tool error null already got.
fleet_send has always used isBlank here; reply now matches it.
The worker found that null/isBlank difference and reported it as a correction
to my ticket, which had quoted the guard wrongly. It was right: I grepped the
error string and assumed the condition matched its sibling.
REST read content with .path("content").asText(""), so a body missing the key
became an empty string that resolved the lead's waiter. The turn completed and
the lead saw a member that finished and reported nothing, indistinguishable
from one that genuinely said nothing. The guard went into MessageService.reply,
which both doors call, rather than being written a second time in FleetApp.
release() cancelled the consumer and dropped its local record of deliveries the
broker still held as outstanding. Cancelling a consumer does not requeue them:
they stay unacked on the still-open channel until it or the connection closes.
So a held-but-undrained reply became permanently unreachable — a worker's
report lost with no error and no log line. It now nacks with requeue, after
cancelling, so a later own() can still receive it.
The merged change shared the QuarantineSource and OutageSource instances
between fleet_profiles and GET /profiles, so the two doors read identical
facts. It then rendered those facts through a character-for-character copy of
the loop, in a different file. Shared inputs do not make duplicated computation
safe: a later edit to the row shape lands on one door and not the other, and
the two disagree about a live outage. That is what #284 was.
The ticket caused this. It said 'read from the same shared instances' and 'do
not change FleetMcp', and together those made copying the loop the only legal
move. Extracting FleetMcp.profilesView and calling it from both is what the
ticket should have asked for.
GET /agents and GET /members now map a herdr transport failure into the same
{error, detail} envelope every other handler in FleetApp uses, instead of
letting it escape to Javalin's default handling. GET /profiles now reports the
quarantined and coolingOff states, read from the same shared sources FleetMcp
reads. REST is the door a lead falls back to when its MCP mount drops, so it
was weakest exactly when it was load-bearing.
release(target) used to cancel the target's consumer and clear held's local
record for it. Cancelling a consumer does not requeue the broker's in-flight
deliveries — they stay unacked on the still-open channel until a real
connection drop. So a held-but-undrained reply became permanently
unreachable: never acked, never nacked, never requeued, invisible to peek.
Fix: cancel the consumer first (so it can no longer receive redeliveries),
then nack-with-requeue every held delivery for that target before dropping
the local record. Nacking before the cancel was tried first but a real
broker demonstrated a race: the still-active consumer immediately received
the requeued message back, racing held.remove and leaving peek non-empty.
Cancelling first avoids that. A failed requeue is logged at WARN and does
not abort release(), matching the best-effort teardown style #293 settled
for HerdrPeerLauncher.stop().
Extends AmqpReplyInboxContractTest.releaseCancelsConsumer... to assert the
held delivery is recoverable via a later own(), not just absent from peek.
Two REST-only visibility gaps, both against the same shared instances FleetMcp
reads (BackendQuarantine/BackendOutagePolicy), never recomputed:
- GET /agents and GET /members let a HerdrException escape uncaught, outside
the {error, detail} envelope every other failure path in FleetApp uses.
Both now route through the existing herdrError() helper, matching healthz/
sessionStatus. GET /members is the endpoint's own comment names as the
out-of-band path a lead falls back to when its MCP mount drops.
- GET /profiles omitted the two outage states fleet_profiles already reports:
quarantined (CB-578 stage B) and coolingOff (fleetd #201 Unit 5). FleetApp
now takes the SAME FleetMcp.QuarantineSource/OutageSource instances Fleetd
wires into FleetMcp (extracted to local vars in Fleetd.java so both doors
share one object, not two independently-built copies of the same rule).
FleetMcp itself is unchanged. Item 3 of the ticket (a capacity block on
GET /members) is explicitly out of scope and was not added.
The javadoc said 'for each session that is BUSY, poll up to timeoutNanos',
which reads as a per-session grace period. The deadline is taken once, before
the loop, so the first BUSY session can spend all of it. That is deliberate and
is the safer of the two designs: the drain is one phase of a shutdown sequence
that must finish inside launchd's exit window, and a per-session grace would
overrun it and get the daemon SIGKILLed part-way through, leaving the sessions
not yet reached with no clean release, no preserved-worktree log and no
snapshot. Found by a read-only hunt that read the code correctly and drew the
opposite conclusion from the wording.
Two exits created a pane and left it running. The readiness gate propagated an
unrelated herdr error without teardown, and spawnAsPane never closed the pane it
split when the peer failed to start. Neither could be cleaned up by the caller:
SessionManager.acquire never learns the pane id, because spawn throws before it
returns one. It removed the worktree anyway, so the leak was a live backend with
a deleted cwd, invisible to fleet_list and holding a seat nothing decremented.
No behaviour change today. HerdrCodec wraps every encode/decode failure
and UnixSocketHerdrClient wraps every IOException, so HerdrException is
all closeTab can currently throw.
But releaseZdotdir five lines below catches RuntimeException, and the
whole point of this fix is that nothing here may mask the cleanups
below. Guarding against the expected exception type and staying bare
against any other is the same asymmetry the ticket exists to remove,
one level down. This stops a later change inside
WorkspaceControl.closeTab reopening it.
The pane is already closed by the time spaces.closeTab runs, so a failing
tab.close is cosmetic workspace tidying, not a real teardown failure. Left
bare, it propagated out of stop() and masked releaseZdotdir (ZDOTDIR leak)
and, worse, SessionManager.release()'s worktree removal (no self-heal,
no retry — the registry entry is already gone by then).
Wrap it in a try/catch that logs a WARN naming the tab id, matching the
"must not mask a real teardown failure above" comment already on
releaseZdotdir. isAlreadyGone is untouched — this continues past *any*
tab.close failure, not just *_not_found, since the failure is cosmetic
regardless of its cause.
Adds FakeHerdr#tabCloseFailsWith/tabCloseFailsForTab (the tab.close
counterpart to #290's paneCloseFailsForPane) plus two tests: one proving
releaseZdotdir still runs (the generated ZDOTDIR is deleted) and one
proving SessionManager.release() still removes the worktree, both with a
non-not_found tab.close failure.
The delayed re-check reads `states` from its own scheduled task, while
`tick` writes and prunes it. Both run on the single-threaded scheduler
Fleetd passes in today, so they are serialised — but nothing in the
class enforces that, and an unsynchronised HashMap read racing a resize
can spin a CPU forever rather than fail visibly.
`priors` and `orphanStreaks` stay plain maps: `tick` is still their only
toucher. The comment says which is which, so the next person does not
have to re-derive it.
FleetHealthMonitor's fire-once-per-transition rule (CB-580) means a target
that is genuinely mid-fleet_ask when health first classifies it GONE is
correctly skipped (sweepAsking=false). But nothing re-fires abandon() once
that ask lapses on its own 55-115s later: FleetHealth.decide keeps reporting
GONE every tick, and reportTransition's previous==next guard never lets the
sweep run again. The ticket then sat PENDING forever, the same destination
#275 fixed for an explicit teardown, reached here by a health guess instead.
SessionManager.reapIdle only reaps READY/DONE sessions (SessionManager.java:844),
and a session mid-turn (including mid-ask) stays BUSY the whole time
(onDelivered sets BUSY, nothing clears it until the turn completes) — so
SessionReaper never releases such a session and onRelease's sweepAsking=true
path is never reached.
Fix: schedule one bounded, delayed re-check per terminal transition (not a
per-tick retry — that shape was rejected by CB-580). It fires failTerminalTarget
again after a delay that exceeds the worst-case ask-lapse window, and only if
the target is still classified in the same terminal state at that time, so a
recovered or since-released target is never reached into. sweepAsking stays
false throughout, so a ticket whose ask has not yet lapsed is still never
touched — same invariant abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer pins.
Proven with a mutation: neutering recheckTerminalTarget's body made
delayedRecheckSweepsATicketWhoseAskLapsedAfterGoneWasFirstObserved fail with
"expected: <FAILED> but was: <PENDING>", all 20 other FleetHealthMonitorTest
cases still green; restored and reran clean (1300 tests, 0 failures).
#283 fixed release() to catch and log a worktree-removal failure, which closed off
reapIdleCountsAllThreeSessionsWhenOnlyItsWorktreeRemovalFails as a trigger for
reapIdle's own per-session try/catch (CB-581) — that test now proves a different,
still-real thing (a swallowed removal failure doesn't shrink the reaped count),
but the try/catch itself lost its test.
Add FakeHerdr.paneCloseFailsForPane(paneId, code) so a test can make exactly one
session's launcher.stop() fail while its siblings still tear down normally
(paneCloseFailsWith already existed but fails every pane, which cannot isolate
one session in a three-session reap). Add
reapIdleSurvivesOneSessionWhoseLauncherStopFails beside the #283 test, using
launcher.stop() as the trigger the ticket names, and prove it catches removal of
reapIdle's try/catch: deleting the guard makes the test fail with the
HerdrException propagating out of reapIdle uncaught (quoted in the PR body).
Two corrections on top of the merged worker branches.
#284: I told the worker to report a BACKEND_ERROR/FAILED session as
reclaimable. That half of my own ticket was wrong. Once the live count
stops counting a terminal session, its seat is already in `free`;
counting it in `reclaimable` too reports the same seat twice, and
`free + reclaimable` reads as more capacity than maxLoad allows. Worse,
only the profile-level count was widened, so the same fleet_list
response said `reclaimable: 2` while every member row said
`reclaimable: false`.
Both views now call one shared predicate, FleetMcp.reclaimable, so they
cannot drift apart. A test runs it over every MemberSession.State value,
so a state added later cannot slip through unconsidered.
#285: the new per-file chgrp+chmod helper moved from ClaudeCodeLauncher
into EnvAllowListScrub as shareFileWithGroup, next to the directory-wide
shareWithGroup it was copied from. It now reuses that class's own
setGroupAndPermissions and also catches UnsupportedOperationException,
which the copy missed — on a filesystem without POSIX group ownership
the copy threw a raw runtime exception instead of the sibling's
UncheckedIOException.
Measured after merging: removing the guard alone leaves the new test green,
because ask() calls markAsyncQuestion before resolveQuestion, so the task has
already moved to the new turnId. The PR claimed each half was necessary; only
the pair is. Keeping the guard, with the ordering written down so nobody
deletes it as dead code or trusts it as the only protection.
seedTrustDialog gated only on isProvisionedWorktree(cwd) and, being static, could not
see memberHerdrSocketConfigured() — unlike its sibling writeCharterFile, which already
refuses the spawn when it cannot place a file where a different-uid member can read it.
Under memberHerdrSocket + configDir unset, seedTrustDialog wrote fleetd's OWN
~/.claude.json while believing it was seeding the member's, reintroducing the fleetd
#149 failure (interactive trust dialog, no fleet_reply, silent readiness timeout) for
this one config combination.
Makes seedTrustDialog an instance method so it can see memberHerdrSocketConfigured()
and memberGroup(), and applies writeCharterFile's "refuse, don't degrade" rule: under
memberHerdrSocket it now requires both configDir and worktreeGroup before touching any
file, naming exactly which is missing, and shares the written .claude.json group-
readable (rw-r-----) via a new shareTrustJsonWithGroup so the member's OS user can
actually open it. The memberHerdrSocket-absent path (today's only live mode) is
unchanged.
answer() opened a fresh forward waiter but, unlike send(), never
registered it in asyncTasksByWaiter. So when a worker chained a
second fleet_ask inside the same resumed turn (before calling
fleet_reply), markAsyncQuestion had no Task to re-associate, and
answer() then completed the async ticket's future with the second
QUESTION as if it were a terminal reply — fleet_poll reported FAILED
while the worker was still alive and mid-conversation.
Fix: register answer()'s waiter in asyncTasksByWaiter (mirroring
send()) so a chained ask can re-arm the ticket under its new turnId,
and guard answer()'s finishAsyncTask call the same way sendAsync's
own lambda already does (skip on Outcome.QUESTION). Also drop the
stale asyncTasksByTurn entry left behind when markAsyncQuestion
re-arms a task under a new turnId, a leak the fix makes reachable
for the first time.
Reachability confirmed by driving the exact sequence through the
public API (sendAsync -> ask -> answer -> ask again) in a new test;
reverting the production change makes it fail with
"expected: <ASKING> but was: <FAILED>", confirming it catches the
regression.
Two teardown-cleanup leaks in SessionManager, same shape as #274.
Defect 1: release()'s last step (removing a released session's worktree)
was the one cleanup step in the method left unguarded, even though every
sibling step is wrapped because exec() can throw on a non-zero exit or its
own 30s timeout. By the time it ran, the registry entry, retained handle,
and pane were already gone, so a throw here escaped release() with no
retry path and made a fully-torn-down session look like a failed stop.
Now wrapped in try/catch with a WARN, matching the pattern already used
by every other step in this method.
Defect 2: acquireWithWorktree's catch (covering failures after add()
returns — overlayParity, shareWithGroup, launcher.spawn) removed the
worktree but left the branch it provisioned orphaned. #274 already fixed
the sibling failure inside add() itself (GitWorktrees.cleanupAfterAddFailure
deletes both). Extracted that branch-delete into a new Worktrees.deleteBranch
method, reused by both cleanupAfterAddFailure and this catch, so a routine
spawn failure (quarantined credential, backend refusal) no longer leaks a
worker/<slug>-<nonce> branch.
A normal release() still never deletes a branch — only the failed-provision
path does. releaseRemovesWorktreeButDoesNotDeleteBranch pins this, and
spawnFailureAfterAddDeletesTheOrphanedBranch / unchangedRegression* prove
the two paths stay apart.
A member torn down while parked in fleet_ask left its async ticket pending
for good. resolveQuestion had already closed the forward waiter, so
abandon()'s waiter branch found nothing; the 'question == null' guard then
excluded the task from the matching loop. By the time the worker's own ask
lapsed (~55-115s), the released session was gone from the roster, so
nothing was left to call abandon() on that target again. fleet_poll{ticket}
reported PENDING forever.
The ticket told the worker to drop the 'question == null' guard. That was
wrong, and the worker said so with evidence: an existing test
(abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer) deliberately pins that
an ASKING ticket must SURVIVE abandon(), because the primary may be mid
answer() for that same turn. Widening the shared method would have traded
this bug for a worse one — a health guess killing a live conversation.
So the fix splits the two callers by what they actually know:
- sessions.onRelease (fleet_stop / idle reaper) knows the pane is being
stopped right now, so it sweeps: sweepAsking=true.
- FleetHealthMonitor keeps sweepAsking=false. GONE/NEVER_READY is a
classification from the live agent list, not a teardown it performed.
I verified the reachability chain myself rather than taking it on trust.
FleetHealth.decide returns GONE before it can ever return
DELEGATION_ORPHANED; FleetHealthMonitor.reportTransition returns early when
previous == next; and terminal() is GONE/NEVER_READY only. So after the one
GONE transition fires and no-ops, nothing re-fires. Every link holds.
Verified: the real merge into current main builds green (1283 tests), the
protective test still passes untouched, and my own mutation — reverting the
sweepAsking widening — fails the new test with 'a released target's open ask
can never resume, so it must fail right here'.
GitWorktrees.add() created the worktree and branch, then ran more steps that
can throw — requireCredentialFreeHttpsOrigin among them, which is an
intended security refusal, not an IO accident. Any throw meant add() never
returned, so SessionManager.acquireWithWorktree never learned the path, its
'if (path != null)' cleanup could not fire, and the worktree and branch
leaked with nothing tracking them. Every OTHER exit from that method was
cleaned up correctly; only the exits inside add() were uncounted.
add() now cleans up what it created before rethrowing, reusing remove() and
additionally deleting the branch — a branch that never finished provisioning
has no session and no PR behind it. Worktree first, since a checked-out
branch cannot be deleted. Cleanup failure is logged and never masks the
original exception.
Verified by me: the real merge into current main builds green (1281 tests),
and I reran the mutation myself without git stash — dropping the cleanup
call fails the new test with 'the worktree directory leaked after a
post-creation step threw'.
The test drives add() itself through the existing afterWorktreeAdded seam,
so the failure happens after the worktree exists rather than downstream in
another caller.
Confirmed reachable: a target torn down for good (fleet_stop / the idle
reaper) while its async ticket sits in fleet_ask (Phase.ASKING) got
permanently stuck. resolveQuestion already closes the forward waiter, the
question == null guard excluded the task from abandon()'s sweep, and by the
time the worker's own fleet_ask lapses (~55-115s) the released session no
longer appears in FleetHealthMonitor's roster, so nothing ever calls
abandon() again. fleet_poll{ticket} then reports PENDING forever.
Add abandon(target, reason, sweepAsking) — sessions.onRelease (a definite
teardown: the pane is being stopped right now) passes true and now fails the
ASKING ticket and closes its reverse-rendezvous ask. FleetHealthMonitor's
health-classification call keeps the 2-arg overload (sweepAsking=false):
a GONE/NEVER_READY reading is a guess from the live agent list, not a
teardown it performed, and abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer
already covers why an active ask must survive that guess (the primary may
be mid-answer for the same turn). hasOrphanedDelegation is left unchanged
for the same reason — it must not flag a live, active ask as orphaned.
Proven with a test driving the real public sequence (sendAsync -> ask ->
abandon(..., true)), not a hand-built task map; reverted the widening to
confirm it goes red, then restored it.
A worker's worktree is isolated; refs/stash is not. It is one stack shared
by the primary's checkout and every worker worktree of this repo.
This bit a real worker today. Two ran in parallel; one called git stash
while the other was mid-edit, and the second worker's in-progress change
was silently overwritten by the first's stashed content. It recovered by
retyping the edit and diffing to confirm, and pushed the other worker's
change back onto the stack untouched — but nothing warned either of them,
and nothing would have.
Measured before writing this: 'git stash list' from a worker worktree and
from the primary's checkout return byte-identical output, and refs/stash
is a single common ref, not a per-worktree one.
The branch already IS the isolation, so the skill now points at committing
a wip commit or writing a patch file instead.
A malformed exhaustedPattern passed FleetConfig.load and then crashed the
daemon at startup, in Fleetd.main's unguarded Pattern.compile, with a
message naming neither the profile nor the key. Its sibling errorPattern
had a load-time validator whose own javadoc explains exactly why that is
bad. The validator was correct; its coverage was not.
rejectMalformedErrorPattern becomes rejectMalformedProfilePatterns and now
compiles both keys, reporting failures from either in one message.
Verified by me, not taken on the worker's word: the actual merge of this
branch into main builds green (1280 tests), and I reran the mutation myself
— narrowing the loop back to errorPattern turns exactly the two new tests
red, one of them with 'Expected IllegalStateException to be thrown, but
nothing was thrown', which is the defect stated out loud.
GitWorktrees.add() created the worktree and branch, then ran several more
steps that can throw (requireCredentialFreeHttpsOrigin — an intended
security refusal, not only an IO accident — plus the credential-helper and
tool-surface isolation steps). Any exception there meant add() never
returned, so its caller (SessionManager#acquireWithWorktree) never learned
the path: its local `path` stayed null, the `if (path != null)` cleanup
guard never ran, and the worktree directory and branch leaked on disk
forever with nothing tracking them.
Wrap those steps in try/catch; on failure, clean up via the same
`git worktree remove --force` path remove() already uses, additionally
force-delete the new branch (remove() alone deliberately leaves a
released session's branch behind, but a branch that never finished
provisioning has nothing else pointing at it), log the cleanup outcome,
and rethrow the original exception so it is never masked.
Test drives add() itself via the existing afterWorktreeAdded seam with a
mutation that trips requireCredentialFreeHttpsOrigin after the worktree
exists, then asserts both the worktree directory and the branch are gone.
Reverting the fix (git stash on GitWorktrees.java, test unchanged) turns
it red: "the worktree directory leaked after a post-creation step threw
==> expected: <false> but was: <true>". Restored afterward.
mvn clean install: BUILD SUCCESS, Tests run: 1275, Failures: 0, Errors: 0
FleetConfig.rejectMalformedErrorPattern only compiled errorPattern eagerly
at config load. exhaustedPattern was compiled unguarded in Fleetd.main,
so profiles.<name>.exhaustedPattern: "[" passed load() and then crashed
the whole daemon at boot with a raw PatternSyntaxException naming neither
the profile nor the key.
Rename the validator to rejectMalformedProfilePatterns and extend it to
also compile every non-blank exhaustedPattern, reporting
profiles.<name>.exhaustedPattern ("<value>"): <message> in the same style
as errorPattern. Both keys are collected and reported together from a
single load. Fleetd.java's compile site is left as-is per scope — it is
now safe because load already rejects a bad value.
Added tests covering: a bad exhaustedPattern is refused; a bad pattern in
each key is reported together in one message; valid patterns still load;
a blank/absent exhaustedPattern is ignored.
A gap in my own #269 fix. That ticket stopped four sites claiming things
about the member's environment that fleetd cannot see when memberHerdrSocket
is configured, and gave the WARN in logCredentialGap a guard. The INFO line
called two lines earlier never got one:
logAllowListCoverage(allowed); // no guard
logCredentialGap(creds, allowed); // guarded since #269
Read plainly, "member credentials: allowed 7 of 39" is a statement about the
member's credentials. Under memberHerdrSocket the pane is routed to a second
herdr whose environment fleetd has no channel to inspect, so those counts
come from fleetd's own process instead. Same overclaim #269 existed to
remove, in the line next door.
The method's javadoc does carry the caveat, by cross-reference to another
field's javadoc. That does not help the operator reading fleetd.out.
The counts stay useful, so this is not a WARN and not a refusal — only the
claim is narrowed. The unguarded path keeps its exact original wording, so
the existing assertion on "member credentials: allowed 1 of 3" still holds.
The new test pins the pair together so a later edit cannot fix one line and
leave the other. Mutation-proved: with the guard removed it fails printing
the old line verbatim.
Also worth recording: no test covered #269's own guard — that WARN wording
shipped unverified, and still has no coverage.
fleet_poll is two operations behind one tool name. With `ticket` it observes
an async delegation and changes nothing. With `target` it calls
MessageService.drainReplies, which REMOVES the replies — a second call
returns nothing.
The handler gated both branches with a constant Authz.Action.READ, and did
not pass the target at all. READ is open to every authenticated role, so any
worker could read a peer's sessionId out of fleet_list and destroy the
replies that peer had queued for the primary. The gate failed open, and a
drained reply is not recoverable.
Three things already said the tight gate was intended:
- fleet_ack, four lines below, gates the same drain as DRAIN, with a
comment giving the exact reasoning missed here ("Acking removes a reply
from the inbox, so it is a drain, not a read").
- the REST path checks DRAIN in FleetApp.drainReplies.
- wiki/2-Message-Server.md lists fleet_poll as lead-only, and the tool
schema says "drain that worker's inbox".
Nothing that works today breaks: the documented flow is fleet_poll{target}
then fleet_ack{target,msgId}, and fleet_ack is already primary-only. A
worker could never complete that flow — only destroy its first half.
The required action is a function of the arguments, but the handler chose it
before looking at them. pollAction(target) makes that choice explicit. The
ticket branch stays READ on purpose: an architect may fleet_send, so it owns
tickets and must be able to poll them.
Why the suite missed it: FleetMcpAuthzTest checks every Action against every
Role, including "a worker may not DRAIN", and passed the whole time. The
policy table was right; the action fed to it was wrong, and nothing tested
that mapping. The new tests assert against pollAction itself, so the handler
keeps no private copy of the rule.
Mutation-proved: reverting pollAction to a constant READ turns exactly the
two new defect tests red and leaves the ticket-branch test green.
Introduced in 9daf1ec, where Authz.READ's own javadoc ("...task polling")
describes only the ticket half.
OpenCodeLauncher.SessionAwareHandle.agentSessionId() is the only caller of
checkModelMatch (fleetd #175), and it sits behind the fleetd #249 worktree
gate. A spawn with no worktree:true — the ordinary shape of most opencode
spawns — never reached the check at all, and the gap was totally silent.
The check cannot be decoupled from agentSessionId()'s resolved id: doing so
would re-derive 'whatever is newest in the shared directory' and reintroduce
the false-positive risk fleetd #234 fixed (a sibling's differently-configured
model looking like a mismatch for a profile that never actually ran it). The
#249 gate is correct and stays as-is.
Instead, log once per profile at WARN, naming the profile, the same
treatment discoveryUnavailable already gets a few lines above — a logged
UNKNOWN beats a check that silently never runs.
Adds PackageCyclesTest, which fails the build on any new cycle between
the top-level dev.ltms.fleet.* packages. Today's five real cycles are
recorded as narrow, explicit exceptions (ignoreDependency per named
pair, both directions), each commented with the ticket step (or a note
that it needs its own) that removes it. No package moves in this PR.
archunit-junit5 1.5.0 (current stable, newer than an earlier 1.4.1
draft). Main code only (DO_NOT_INCLUDE_TESTS) and importPackages(...)
instead of a working-directory-relative target/classes path.
HerdrPeerLauncher asserted, as established fact, that member panes run under
a different OS user whenever memberHerdrSocket is configured. fleetd has no
channel to see the uid at the other end of a herdr unix socket — an operator
may point memberHerdrSocket at a second herdr under the SAME user for pane
isolation, in which case members do inherit fleetd's environment and the
count this WARN told them to disregard is the real gap.
Reworded the class javadoc on hostEnvNames, the WARN in
warnUnknownMemberEnvironment, the javadoc on memberHerdrSocketConfigured(),
and warnCannotShareScrubDirectory's "unreadable by another uid" claim to say
what is actually true: fleetd cannot confirm what OS user the second herdr
runs as, so the member credential gap is UNKNOWN, not known-clean or
known-dirty. No behaviour change — the fallback paths and the honest
UNKNOWN conclusion stay the same, only the stated reason changes.
Matches the framing already used by Fleetd.reportMemberTrustModel on main.
Added unknownEnvironmentWarnStatesUncertaintyNotAnAssertedDifferentUser to
HerdrPeerLauncherAllowListWiringTest asserting the new WARN wording and that
it no longer claims a different OS user as fact.
The correction removed a false claim (blocking the socket breaks git over
SSH) but took a true one with it: the socket is a live handle to the agent,
so a member holding it can sign with every key the agent holds. Without that,
the entry reads as if the setting does not matter, and an operator has no
reason left not to set it to allow. Fixing an overclaim must not leave an
underclaim.
Also record the measurement and the mistake behind the old claim, so the next
person does not re-argue it from scratch.
The merged fix told the operator to look in journalctl for
"Fleetd.reportRequiredSecrets". That string never appears there — it is a
method name. The logger is d.l.f.Fleetd and the lines read
"startup secret NAME: set|MISSING", so the advice sent the operator looking
for text that does not exist.
Give the grep instead, and say what the report does not cover: it lists only
names a configured profile references, so a secret nothing references is never
reported.
fleet_list's free row subtracted leadSeatCount, but the real spawn gate
(CompositePeerLauncher#enforceMaxLoad) only ever compares live against
maxLoad and never reads leadSeatCount. So free could report 0 while a
fleet_spawn on that exact profile still succeeded, and a lead trusting
free:0 gave up on capacity the gate would still grant.
free now always equals max(0, maxLoad - live); leadSeats stays in the
row as an informational fact, never subtracted. Documented in the
fleet_list tool description and fleetd.example.yaml.
TRUST_JSON_LOCK only serialises seedTrustDialog calls this launcher makes
inside its own JVM. It cannot reach the one writer that actually shares the
target file on a real host: the operator's own live Claude Code, whose
CLAUDE_CONFIG_DIR is routinely the very configDir a profile is given, so the
file fleetd writes on every claude-code spawn is that session's own config.
A plain read-modify-write there is a routine lost update: fleetd reads v1,
the operator's session writes v2, fleetd's ATOMIC_MOVE lands v3 built from
v1 and v2 is gone, atomically.
Add a bounded compare-and-swap: before the move, re-read the target's exact
bytes and compare with what the update was built from; on a mismatch,
rebuild from the fresh bytes and retry (up to 5 attempts). Exhausting the
retries writes nothing and logs a WARN naming the file — the member shows
the trust dialog and fails to reach an injectable state instead, which is
visible and recoverable, unlike silently overwriting the operator's live
config. Also warn every time the seed is about to target the default
~/.claude.json (configDir unset), since that is the unsafe default.
This narrows the lost-update window, it does not close it: a write landing
between the final re-read and the ATOMIC_MOVE itself is still lost.
Ports the intent of the lost CB-576 commit c393600 (worker/cb576-01a04b-17,
never merged, package dev.ltms.bridged.*) onto main's dev.ltms.fleet.*
tree. Adds FakeWorktrees.markGone (tracks which add()'d worktree paths
still "exist", mirroring GitWorktrees.hasUncommitted's Files.exists guard
for the already-gone case) and a regression test,
releaseStillStopsPaneAndNotifiesWhenWorktreeIsGone, asserting that
SessionManager.release still fires notifyReleased, stops the pane, and
falls through to remove() when the worktree is already gone.
Follow-up to #259, found by verifying the guard rather than trusting it.
The scrape read FleetApp's raw source, so a registration disabled with `//`
still matched. Measured: commenting out `app.get("/tasks/{ticket}", ...)` left
the test GREEN, while deleting the same line was caught. Only the commented-out
shape was blind, and it is the silent direction — the inventory would keep
claiming a route the server no longer serves.
Drop whole-line comments before scraping. Only lines starting with //, * or /*
are dropped, deliberately not every // on a line: that would also cut a string
literal containing // (a URL) and could silently delete a real registration
sharing the line. The remaining gap is a trailing comment beside real code; no
registration here has that shape, and the vacuity test catches a scrape that
loses registrations wholesale.
Proof: with the fix, the comment-out mutation fails naming
"Removed ... [GET /tasks/{ticket}]"; reverted, FleetApp.java confirmed clean.
mvn clean install: 0 compile errors, 1257 tests, BUILD SUCCESS.
FleetApp's route list has never been checked against anything and has
already drifted once (GET /member-credentials shipped hours before the
#252 ticket and was missing from its list). Add
RestRouteInventoryTest, modelled on McpContractDocTest, which scrapes
FleetApp.java's app.<verb>("path") calls with a regex and compares
them against an explicit expected inventory, failing loudly with the
added/removed routes when they diverge.
seedTrustDialog targets ~/.claude.json when a profile sets no configDir.
Two IDE-overlay fixtures built a worktree-shaped @TempDir, which opens the
#149 isProvisionedWorktree gate, and left configDir null — so every test run
added two project entries to the operator's real file. 116 had accumulated,
none of them still existing on disk, and 32 of those came from current code.
The #149 gate only closed the opposite case: a fixture with cwd unset falling
back to user.dir. A fixture that builds a worktree on purpose walks straight
through it.
- ideProfile/ideProfileModule now take configDir first and mandatory, so each
fixture states where the trust seed goes.
- noFixtureSeededTheDefaultClaudeJson snapshots the temp-dir project keys in
@BeforeAll and fails in @AfterAll on any key this class added. Differential,
not absolute: an absolute check would fail on every host still carrying the
historical entries, and such a check gets deleted rather than fixed.
Not done: a blanket -Duser.home redirect in surefire. EnvAllowListScrubTest
tests the credential scrub against the operator's real login chain and guards
with assumeTrue($HOME/.zshrc exists), so the redirect would silently skip two
security tests.
Proof: with the bug put back on one fixture the guard fails and names the path;
reverted and confirmed identical with diff -q. A full suite run under a fake
home now creates no .claude.json at all. mvn clean install: 0 compile errors,
1255 tests, BUILD SUCCESS.
scripts/probe-member-credentials.sh carried its own NAMES array of 31
names. The live policy has 34. The probe reported 26 blocked against a
policy that blocks 29, exited 0, and printed a table that looked
complete. A verification tool that under-reports is worse than none,
because its clean output stops anyone looking.
Same defect as #114, fixed the same way: DELETE the second copy rather
than correct it. The NAMES array is gone, not updated.
The daemon now serves GET /member-credentials — names and counts, never
a value; MemberCredentialPolicyView reads no environment at all, so
there is nothing to redact by construction. The probe fetches it and
refuses with a non-zero exit when the daemon is unreachable, the policy
is absent or empty, or knownCount disagrees with the length of known[].
No local fallback: a verification tool must not quietly degrade into a
weaker check.
The startup log line and the endpoint now share that one class, so the
counting exists once. That also protects a subtlety I measured before
briefing this: blocked is NOT known - allowed. Live, known=34 and
allow=7, but only 5 of those 7 appear in known, so blocked=29 and the
naive subtraction gives 27. The view reuses creds.blockedSet(), the
existing derivation, so it keeps 29.
Worker's mutation: MemberCredentialPolicyView.of(...) forced to return
ABSENT turned 4 tests red with 0 compile errors — including
MemberCredentialsGapReportTest, which proves the startup log really
does run through this path. Reverted and confirmed with diff -q.
It also caught a bug in its own first draft: jq's // operator treats
false and 0 as missing, so `.present // empty` turned a genuine
"present": false into "unknown". Fixed by reading the fields directly.
NOT yet verified: acceptance criterion 5, the live 34/29/5 run. The
route does not exist until the daemon is redeployed onto this jar, so
that check comes next and is mine, not the worker's.
Merged clean, then built on the merged tree: 1255 tests, 0 failures,
0 compile errors.
policy=allow-list is enforced by a ZDOTDIR scrub, and a non-zsh login
shell ignores ZDOTDIR entirely, so no scrub runs. The launcher already
DETECTED this and logged a WARN — then degraded to the weaker overlay
and spawned anyway. The operator asked for the blocking control and
silently got the weaker one, which is the defect the ticket is about.
Detection existed; refusal did not. Under policy=allow-list a non-zsh
shell now throws IllegalArgumentException before any ZDOTDIR or env
work, naming the actual shell and giving three ways out. Under
policy=deny-by-default nothing changes: that overlay is applied to the
pane before any shell runs, so it does not depend on the shell.
Checked against the live config myself, because this refuses spawns and
no worker can see fleetd.yaml:
memberHerdrSocket : NOT set -> the shell comes from fleetd's own
$SHELL, not the unset memberLoginShell
policy : allow-list
fleetd's $SHELL : zsh, proven by behaviour rather than by reading
the process env — the daemon log shows the ZDOTDIR
scrub generating a directory 147 times, most
recently minutes ago, and that only happens when
isZshShell() returned true
So the new refusal cannot fire on this host. Had memberHerdrSocket been
set, the unset memberLoginShell would have read as "<unset>", non-zsh,
and refused every spawn — worth knowing before anyone sets that key.
Worker's mutation evidence, re-stated: `if (!zsh)` -> `if (false)` turned
the refusal test RED with 0 compile errors, then reverted clean.
NOT verified: a live non-zsh member spawn. Forcing it means changing the
daemon's own environment, and the value of the test does not justify
that. The unit tests drive the real launcher.spawn entry point.
Merged clean, then built on the merged tree: 1250 tests, 0 failures,
0 compile errors.
Stage 1 shipped INERT on this host and every test was green. The matcher
compared effectiveCredentialId(), which fell back to the profile's own
NAME when credentialId was unset. This host runs the lead on `opus` and
members on `sonnet`; both are subscription:true with no credentialId, so
it compared "opus" against "sonnet", never matched, and charged 0 seats.
Every stage-1 test put the lead on the SAME profile name as the target,
so the fixture encoded the one shape the live config does not have.
Stage 2 returns a "<subscription>" sentinel when credentialId is unset
and subscription is true. An explicit credentialId still wins, so an
operator with two genuinely separate Claude logins can keep them apart.
Verified by me on the live config shape, not by reasoning:
opus.effectiveCredentialId() = <subscription>
sonnet.effectiveCredentialId() = <subscription>
seats charged to sonnet = 1 (was 0 before this change)
Only opus and sonnet join the sentinel group on this host; local,
local-direct, gx, xf, sol and terra are unaffected. free is clamped
with Math.max(0, ...), so the subtraction cannot report a negative.
Second, wider consequence, flagged by the worker and confirmed here:
CompositePeerLauncher.credentialIdFor feeds enforceNotQuarantined and
enforceNotCoolingOff, so quarantining one subscription profile now also
refuses spawns on the other. That is correct — one Claude subscription
hitting a usage limit really does take out every profile on it — but it
is a behavioural change beyond fleet_list's numbers.
Checked all 5 logical callers of effectiveCredentialId(); every one
wants "this account", none wants "this exact profile".
Merged clean, then built: 1250 tests, 0 failures, 0 compile errors.
An auto-merge with no conflicts is not a compiling merge, so the build
was run on the merged tree before this landed.
scripts/probe-member-credentials.sh carried its own hand-maintained NAMES array
(31 names, recorded 2026-08-16), so a name added later to fleetd.yaml's
memberCredentials.known was never checked and the probe still exited 0 with a
clean-looking table. Same drift shape as #114's tool catalogue.
- New dev.ltms.fleet.member.MemberCredentialPolicyView: the single place that
turns a MemberCredentials policy into names + counts (never a value). Reused
by Fleetd.reportMemberCredentialsGap (startup log line) and by the new
GET /member-credentials REST endpoint (FleetApp), so the two can no longer
drift apart the way the probe and the policy did.
- FleetApp gains one route + handler + a Supplier<MemberCredentialPolicyView>
constructor param (legacy constructors default to ::absent, so existing call
sites are unaffected).
- probe-member-credentials.sh now fetches its name list from
GET /member-credentials instead of carrying one. No local fallback: an
unreachable daemon, an empty/absent policy, or a knownCount/known[] length
mismatch all refuse with a non-zero exit rather than silently checking zero
names. Prints "policy contains N; this run checked N" so the two numbers are
visibly equal.
The existing guard walks KNOWN_TOP_LEVEL_KEYS and anchors its regex at
column 0, so it sees only top-level keys. Every nested key was outside
its scope and nothing said so, which is the shape #113 collects: a
checker narrower than it looks, whose green run stops anyone looking.
Two derived guards replace the assumption:
everyNestedConfigKeyIsDocumentedInTheExample
walks FleetConfig's record components (17 records, 83 distinct
key names) and requires each to be documented in the example.
everyLiveKeyInTheExampleBindsToARecordComponent
resolves every live key path in the example against the record
tree, so a documented key that binds to nothing fails here
instead of being silently ignored in production.
Neither carries a list, so a key added to any nested record is covered
the moment it compiles (criterion 2).
Both mutations run through the real caller, not the helper (criterion 1,
which asks for exactly that):
removed every mention of paneProbeIntervalSeconds from the example
-> FAILS, naming health.paneProbeIntervalSeconds
added a live bind.totallyMadeUpKnob to the example
-> FAILS, naming bind.totallyMadeUpKnob
0 compile errors in both; both reverted and confirmed with diff -q.
The first attempt at mutation 1 removed only the `key:` line and the
run stayed green — correctly, because the key was still documented in
prose. An incomplete mutation proves nothing, so it was redone.
Denominators (criterion 3): both guards print how many keys they
checked, and the floor for "did the walk descend?" is derived from
KNOWN_TOP_LEVEL_KEYS.size() rather than being a literal.
Scope is stated in the javadoc rather than implied: the guards do not
check a key sits at the right path, do not parse commented prose for
the reverse direction, and do not prove a parsed key is read by
anything. paneProbeIntervalSeconds is parsed and read by nothing, and
these guards pass it -- the example already says so in its own text.
everyOptionalKnobDocumentedInTheExampleBinds keeps its hand-written
list but is re-documented as a value-binding spot check, explicitly
not a coverage guard; coverage now comes from the two derived tests.
broker.uri is documented only in the example's prose convention
(`# uri -> ...`), never as a copy-pasteable `uri:` key, because
writing it out invites pasting a password into a file -- the thing
uriEnv exists to avoid. The matcher accepts that convention rather
than pushing the file toward doing it.
Full build: 1236 tests, 0 failures, 0 compile errors.
The ZDOTDIR scrub that enforces policy=allow-list only runs on zsh. The daemon
already detected a non-zsh login shell (isZshShell/warnNonZsh, from #213), but
degraded to the weaker CB-596 overlay and spawned anyway — the exact "control
silently does nothing" defect this ticket is about. Now a non-zsh shell under
policy=allow-list refuses the spawn (IllegalArgumentException, naming the
shell), surfaced by FleetMcp.spawn's existing catch(IllegalArgumentException).
policy=deny-by-default is unaffected in substance (its overlay never depended
on the shell) but now also logs a one-time WARN naming the shell, since the
stronger allow-list control is unavailable there.
The worktreeRoot/worktreeGroup-missing degrade path under memberHerdrSocket
is untouched — that gap is fleetd #213's scope, not this one.
Stage 1's lead-seat matcher (leadSeatLookup) was correct but inert on
the live host: the lead runs on profile 'opus', members on 'sonnet',
both subscription:true with no explicit credentialId. Because
effectiveCredentialId() fell back to the profile's own name, opus and
sonnet never matched even though they share one Claude login, so the
matcher charged zero seats.
FleetConfig.Profile.effectiveCredentialId() now falls back to a shared
sentinel (SUBSCRIPTION_CREDENTIAL_ID = "<subscription>") instead of the
profile name when subscription:true and credentialId is unset. An
explicit credentialId still wins, so two separate Claude logins on one
host can still be kept apart.
This is also BackendQuarantine's and BackendOutagePolicy's grouping
key and CompositePeerLauncher's spawn-time enforcement key, so the fix
also links quarantine/cool-off across subscription profiles sharing an
account -- intentional: one usage limit really does take out every
profile on that login, mirroring credentialId: openai-shared already
doing this for off-subscription profiles. Every caller was reviewed;
none wants "this exact profile" over "this account".
Tests added:
- FleetdLeadSeatLookupTest: the live shape itself (lead on a
DIFFERENT subscription profile than the target, same account,
neither sets credentialId) -- the case stage 1's suite never covered
- FleetMcpTest: quarantining one subscription profile's shared
account zeroes free on another sharing it, via the same
effectiveCredentialId()-driven wiring Fleetd.main uses
Mutation-tested: reverting the subscription branch to the old
fall-back-to-profile-name behavior sends both new tests RED with 0
compile errors; reverting the mutation restores byte-identical
(diff -q) source and green tests.
fleetd.example.yaml's fleetd #176 notes are rewritten for the sentinel
semantics and when to override it with an explicit credentialId.
BackendOutageFlowTest held a ~30-line hand-copy of the lambda in
Fleetd.main, under a comment promising it mirrored production "EXACTLY".
That promise was the defect. The test proved the copy, so any change to
the real sink left the flow test green.
#248 made Fleetd.backendErrorSink(...) public for exactly this reason.
The test now calls it.
Measured, same mutation in the real sink (an early return after
sessions.onBackendError, dropping the cool-off and the lead nudge):
old test (hand-copy): Tests run: 5, Failures: 0 -- blind
new test (real sink): Tests run: 5, Failures: 4 -- catches it
0 compile errors in both runs, so both are real results. Production
reverted and confirmed with diff -q.
Full build: 1234 tests, 0 failures, 0 compile errors.
A member spawned without a worktree inherits the lead's cwd, which holds
many old opencode session rows. sessionIdForDirectory picks the most
recently updated row for that directory, so a brand-new member - which
has not written its own row yet - resolves to somebody else's session.
Measured: a row three days old, from a different profile.
The damage was at the tool surface. fleet_list told the lead that
agentSessionId is the id to pass as resumeSessionId, so acting on it
would resume a stranger's conversation, with foreign context, and
nothing to distinguish that from a correct resume.
Fixed by refusing to answer rather than by making the heuristic smarter.
#234 already established the heuristic cannot be made reliable at that
layer, and its javadoc records why, so the SQL is untouched.
agentSessionId() now returns null for a non-provisioned cwd, and
spawn() refuses a resumeSessionId request for one outright, before
anything starts. isProvisionedWorktree moved to HerdrPeerLauncher so
both adapters share it. fleet_list and fleet_spawn descriptions no
longer describe the id as always safe to resume.
Verified rather than taken on trust:
- the refusal reaches the lead as a readable message, not a stack trace
- FleetMcp.spawn already catches IllegalArgumentException and returns
error(e.getMessage()).
- the message tells the lead to pass fleet_spawn{worktree:<slug>}, which
is valid: worktree is typed string, 'true' or a ticket slug.
- 13 existing tests moved off a placeholder "/work/dir" onto a real
provisioned-worktree fixture. They cover #175/#234 model-mismatch
machinery and would otherwise have tripped the new gate incidentally.
Worker's mutation evidence, both reverted and diff-confirmed:
- gate at OpenCodeLauncher:812 -> if(false): RED at
OpenCodeLauncherTest:448, expected <null> but was <ses_someone_elses>.
- resume refusal at OpenCodeLauncher:683 -> 'false &&': RED at
OpenCodeLauncherTest:379, expected IllegalArgumentException.
Both 0 compile errors.
Closes#249. PR #253.
maxLoad counted panes, never subscription seats: a subscription:true
profile's lead is itself a live claude session on that same account,
so free overstated capacity by the lead's own seat (measured free:1
with a real ceiling of 0, and free:3 on an idle fleet with a real
ceiling of 2).
Add FleetMcp.LeadSeatSource (same shape as QuarantineSource/
OutageSource) and Fleetd.leadSeatLookup, which derives the seat count
from fleet.leaders.<name>.profile matched against the target profile
by effectiveCredentialId() - no hardcoded "-1", and no new config key:
profile: already exists for this exact "which account does this lead
share" question. maxLoad itself is left untouched; only free (and a
new, additive-only leadSeats field) changes.
Exhaustion quarantine (cause 2 in the ticket) already forced free to 0
via the same BackendQuarantine capacityView already reads - confirmed
by reading the exhaustionSink wiring, no code change needed there.
docs/MCP-Contract.md was written 2026-07-14, before any MCP code existed,
and never caught up. CLAUDE.md points every session in the fleet at it.
Audited against the code today. The drift was not confined to the tool
table the ticket reported:
section 3 still described the OLD identity rule - "any connection that
does not map to a known worker is treated as a primary".
That was a real privilege bug, fixed since by the ancestry
walk in #161. The page still taught it.
section 4 names port 8080 (the mount is 8765) and says the pom does
not yet carry an MCP dependency.
section 5 named fleet_read and fleet_cancel, which do not exist, and
omitted fleet_poll, fleet_ack, fleet_profiles, fleet_whoami
and fleet_list, which do.
section 8 says turn_id where the code says turnId, and has no row for
the exhausted outcome CB-578 added.
sections
9, 10, 11 pre-build planning: "new work" columns, open decisions long
since decided, CB-1xx placeholders.
Every one of those is the same defect: a hand-maintained second copy of
something the code already states. So the copy is deleted rather than
corrected - correcting it just restarts the clock.
What survives is the flows and the status gating, because a flow is a
shape rather than a name, and shapes are what this page was ever good
for. They are rewritten with the names checked against the code, and
extended with what has been learned since: the ~60s cap on a blocking
send, the ~55s ask window, and the three ways the turn-done fallback
loses a report (clipped, echoed brief, slow member).
389 lines -> 188.
The names that remain are guarded. McpContractDocTest fails if the page
names a fleet_* tool FleetMcp does not register, and - because an empty
set is a subset of everything - a second test pins that both sides
actually found names, so the check cannot pass by checking nothing. A
third pins the "this is not the tool reference" sentence, which is the
fix itself: without it someone helpfully re-adds a tool table.
Mutation-tested both ways, 0 compile errors each: adding `fleet_read` to
the doc fails theDocNamesNoToolThatDoesNotExist ("names [fleet_read] ...
Checked 6 name(s)"); removing the disclaimer fails
theDocStillDisclaimsBeingTheToolReference.
All 5 mermaid diagrams render under mermaid-cli.
CLAUDE.md's pointer said "section 6 only" and now names the guard
instead. It is in the project addendum, so the canonical block is
untouched - verified still byte-identical with the wiki template.
REST is split out to #252: 14 routes, documented nowhere, and it IS a
supported operator surface - one of them drains on read.
1232 tests, 0 failures.
OpenCodeSessionDiscovery.sessionIdForDirectory keys on the worker's cwd, which
is reliable only when fleetd provisioned a unique git worktree for that
member. Without one (the default no-worktree spawn), the cwd is shared with
other sessions, and "most recently updated row for this directory" can pick a
stranger's session — fleet_list would then hand a lead an agentSessionId that
resumes someone else's conversation.
Move isProvisionedWorktree from ClaudeCodeLauncher to the shared
HerdrPeerLauncher base (both adapters need it now). OpenCodeLauncher.spawn now
refuses a resumeSessionId spawn outright when the target cwd is not a
provisioned worktree (fleetd can never verify or re-report that identity), and
SessionAwareHandle.agentSessionId() withholds the id — returns null rather
than guessing — for any member spawned without one, resumed or not. Corrected
fleet_list/fleet_spawn's tool descriptions, which previously implied
agentSessionId is always a safe resume handle.
seedTrustDialog wrote two keys into the shared .claude.json:
hasTrustDialogAccepted and hasCompletedProjectOnboarding. Only the first
one survives.
Measured live on 2026-09-03, minutes after a spawn seeded the file:
hasTrustDialogAccepted: 28 of 28 project entries
hasCompletedProjectOnboarding: 0 of 28 project entries
Our entry was written by the running jar and the key was already gone,
so it was written and then removed. It is absent from the 27 entries
Claude Code wrote for itself too, which says Claude Code normalises the
whole file when it saves and drops that key every time.
That reframes #247. I filed it as a race - a save landing between our
read and our ATOMIC_MOVE. It is not a race. The other writer removes
this key as its steady-state behaviour, with no window involved. So the
compare-and-swap retry proposed there would not have helped: it would
re-add a key that gets stripped again on the next save.
The seed's whole job is to stop the workspace-trust dialog blocking a
member (#149). The live probe reached idle with hasTrustDialogAccepted
alone, so the second key was never doing that job. Writing it only added
a contested key to a file two processes share, and made the next reader
think it mattered.
The atomic write and the lock stay. Both are still correct, both are
cheap, and hasTrustDialogAccepted is genuinely shared state.
The new assertion is assertFalse, not a deletion. Removing the old
assertion would leave nothing to stop someone re-adding the key later as
a plausible-looking completeness fix. Mutation-tested: restoring the
production line fails
seedTrustDialogWritesOnlyTheTrustFlagAndNotTheOnboardingKey:2198 with 0
compile errors.
1229 tests, 0 failures.
Found by reading a real boot log after the redeploy, not by a test.
coverage() is shared by two call sites — CB-578's exhaustedPattern line
and Unit 5's errorPattern line — but its 'off' branch hard-coded the
word exhaustedPattern. So this daemon printed:
backend-exhausted classification (CB-578 stage A): partial
(configured: [sol, terra]; not configured: [...])
backend-error classification (fleetd #201 Unit 5): off
(no profile has an exhaustedPattern configured; profiles: [...])
Two lines, one directly under the other, disagreeing about whether any
profile has an exhaustedPattern. Both were individually defensible and
together they were nonsense. Worse, the message sends an operator to
set the wrong key: the thing that is missing is errorPattern.
coverage now takes the key name. I changed the signature rather than
adding an overload, so the compiler found all three existing callers
instead of leaving them silently on the old path.
Every earlier coverage test passed the exhaustion case only, which is
why none of them could see this. The new test pins the errorPattern
case. Reverting the fix turns it red with 0 compile errors.
1229 tests, 0 failures, BUILD SUCCESS.
This is the second defect in two hours found only by reading the live
startup log — see #115, where the noise of a false warning had been
hiding a correct line saying a whole feature was off.
Before this, dropping either #241's worktree lookup or Unit 5's
backend-error pair at Fleetd.main's new CompletionResolver(...) call
left all 1216 tests green with 0 compile errors. Every existing test
built its own CompletionResolver, so they proved the class and never
the wiring. BackendOutageFlowTest was the sharpest case: it copies
main's sink lambda line-for-line, so it proves the copy and cannot
notice the original being deleted.
The three inline arguments are now package-private static factories on
Fleetd, following the deliverableTo pattern the file already had, each
with its own behaviour test. backendErrorSink is public so a
cross-package test can drive the real production object rather than a
hand-mirrored copy.
The test that was actually missing is a source-text assertion. That is
the honest fallback for a composition root with no seam, and it is
labelled [SOURCE TEXT] in every test name and message so it cannot be
misread as a behaviour check. It is not vacuous: two tests pin that the
variables are assigned from the factories, and two pin that those
variables reach the call site, so renaming a variable while assigning
an inert value does not slip through.
Known cost, accepted: the assertions match exact source substrings, so
reformatting that statement will break them. That is the price of
covering a main method, and a spurious failure here is loud and
obvious, which is the right direction to fail.
Verified by the lead, both mutations re-run against the merged code —
see the merge check.
PR #251
Fleetd.main built three of CompletionResolver's 8 constructor arguments inline
(a worktree/branch lookup lambda, and the backend-error pattern lookup + sink
locals). Dropping any of them at the call site compiled clean and left every
existing test green, because every existing test constructs its own
CompletionResolver and only ever proves the class, never main's wiring.
Extract each into a static factory on Fleetd (worktreeBranchLookup,
backendErrorPatternLookup, backendErrorSink — the same static-factory pattern
Fleetd.deliverableTo already uses), test each factory's own behaviour, and add
a source-text assertion (FleetdCompletionResolverWiringTest) proving main's
CompletionResolver call still passes all three. backendErrorSink is public so
BackendOutageFlowTest can exercise the real production sink directly instead
of the hand-mirrored copy its own class doc used to describe.
No production behaviour changes — mechanical extraction only.
The default is now [.env], not [.env, .envrc]. .env is data, so copying
it into a worker worktree can only move values. .envrc is executable
shell that direnv runs on every cd, so copying it moves behaviour. Those
are different risks and should not share a default.
The knob is unchanged. An operator who wants .envrc copied writes
parityOverlay: ['.env', '.envrc'] and owns that choice; a new test pins
that escape hatch, because without it this would be a removal rather
than a re-default.
Decision recorded on the ticket, with the evidence it asked for first:
this checkout has no .env and no .envrc, and direnv is not on PATH, so
there was no live exposure. Point 1 (extend the credential scrub to
direnv) is declined and the reason is on the ticket — the scrub is a
one-shot .zlogin and a direnv hook runs on every cd, so no amount of
work on the scrub can cover it. Not copying the executable file is the
smaller change and removes the need.
The worker also fixed WorktreeSessionManagerTest, which hardcoded the
same default at another layer and broke the build. Outside its named
scope, correctly flagged rather than done silently.
Verified by the lead: 1216 tests, 0 failures, 0 compile errors.
PR #250
.env is data; .envrc is executable shell that direnv runs on every cd, so
copying it into a worker moves behaviour, not just values. The default
parityOverlay is now [.env] only. The knob is unchanged: an operator who
wants .envrc copied can still write parityOverlay: [.env, .envrc]
explicitly.
Updates FleetConfig's default and javadoc, fleetd.example.yaml's two
mentions of the default, and the FleetConfigTest coverage: renamed the
default test, added parityOverlayExplicitEnvrcOptInStillWorks to prove
the .envrc opt-in escape hatch still works, and fixed
WorktreeSessionManagerTest#worktreeAcquireRunsParityOverlayWithProfileDefaults
which also hardcoded the old default.
The completion fallback scrapes a member's pane when a turn ends with no
fleet_reply. If the pane still shows the brief the lead injected, the
scrape returned that brief, and the lead read its own words as the
member's answer. A silent member looked like a member that had reported.
echoesInjectedBrief now recognises that case and refuses it. Round 1 used
plain containment in both directions, which destroyed real reports: a
genuine report that quotes the brief contains it. Round 2 keeps the safe
direction unbounded (the brief contains the scrape) and bounds the other
one at MAX_ECHO_EXCESS_CHARS, so a scrape only counts as an echo when it
adds almost nothing to the brief.
Merge note — the Fleetd.java conflict:
This call site was changed by both #201/#227 Unit 5 (backendErrorPatterns
+ backendErrorSink) and by this ticket (the worktree/branch lookup). I
resolved it onto the full 8-argument constructor so neither feature is
dropped; nowNanos has to be passed explicitly to reach that overload.
Verified by the lead: 1215 tests, 0 failures, 0 compile errors.
I also measured whether the resolution itself is protected, and it is
NOT. Both mutations at this call site stay green:
- drop the worktree lookup (pass _ -> null): 1215 tests, 0 failures
- drop Unit 5's patterns/sink (legacy()/none()): 1215 tests, 0 failures
Nothing in the suite covers Fleetd's composition root, so either feature
could be silently unwired here and the build would still be clean. The
tests prove the seams, not the caller. Filed separately rather than
fixed in a merge commit.
PR #245, branch worker/cb241-fallback-echo-1175e9-11
A profile's credential that throws two distinct backend errors inside 60
seconds now cools off for 60 seconds. Automatic placement skips it,
an explicit fleet_spawn naming it is refused before the adapter is
called, and fleet_list/fleet_profiles report it as coolingOffForSeconds
next to the separate CB-578 quarantinedForSeconds.
Verified by the lead: see the merge check below. The worker ran 7
mutations, all killed with 0 compile errors; M7 was NOT killed on the
first pass (the assertion only checked .contains("quarantined"), which
is true of both the correct message and the mutated fallback), and the
worker strengthened it to assertEquals on the exact literal and kept
that change. That is the right call and it is reported honestly.
Two deviations, both justified in the PR:
- FixedPlacementPolicy needed the same coolingOff filter because it
filters candidates inline instead of using PlacementPolicyUtil.
- BackendOutageFlowTest sits in dev.ltms.fleet.inject because
CompletionResolver.InFlight is package-private there.
The startup coverage log line is a code-reading claim, not a captured
line from a live daemon. The worker said so rather than overclaiming.
PR #246, branch worker/cb201-unit5-wiring-6c12e6-8
Wires the already-merged units into production:
- Per-profile errorPattern config (beside exhaustedPattern), compiled once at
startup; falls back to the legacy (?i)\bAPI Error\s*: pattern when unset.
Startup logs configured-vs-legacy coverage, same as exhaustedPattern.
- One production BackendErrorSink in Fleetd.java: mark backend_error on the
session, resolve the profile's credential fail-loud (never
Optional.ifPresent), record it in BackendOutagePolicy, and push a lead
nudge on a new incident.
- CompositePeerLauncher's explicit and automatic spawn paths both refuse a
cooling-off credential; exhaustion quarantine wins when both are active.
PlacementContext gets a separate coolingOff set so refusal text says
"cooling off", never "exhausted".
- fleet_list/fleet_profiles report coolingOffForSeconds as an independent
fact from quarantinedForSeconds; both can appear together.
- fleetd.example.yaml documents errorPattern and the 2/60/60 cool-off policy;
CLAUDE.md tells leads how to read the two independent outage states.
Also: FixedPlacementPolicy.java, not in the original file list, needed the
same coolingOff filtering as PlacementPolicyUtil (it does its own inline
candidate filtering rather than delegating).
1190 tests, 0 failures (up from the 1163 baseline); BUILD SUCCESS.
Claude Code asks 'is this a project you trust?' the first time it starts in a directory it
has not seen. It is interactive with no timeout, and every member spawned with worktree:true
lands in a brand-new directory. The member never reaches its first turn and never replies,
while herdr reports blocked/interactive_ready — which reads as healthy.
ClaudeCodeLauncher now seeds projects.<cwd>.hasTrustDialogAccepted in the profile's
.claude.json before the process starts. This is not a new grant: the operator already
trusted the repo by configuring the profile against it, and a worktree is a checkout of it.
The write is gated on isProvisionedWorktree(cwd) — a .git that is a regular gitdir-pointer
file, never a real checkout. That gate exists because an earlier revision of this change,
run under mutation testing, wrote to the operator's real ~/.claude.json and truncated it
from 72KB to 919 bytes. Tests using a null configDir fall back to the real user.home, so
an ungated seed reaches real files.
The write is atomic (sibling temp file + ATOMIC_MOVE, never truncate-in-place) and the
whole read-modify-write is under a lock, because .claude.json is large, live, and rewritten
by Claude Code itself while fleetd runs. Two parallel spawns are normal here.
Verified by the lead: 1173 tests, 0 failures. A truncating write turns the torn-read test
red; removing the lock turns concurrentSeedsForDifferentCwdsBothSurvive red. Both with 0
compile errors. copyPosixPermissionsIfPresent is NOT covered by a test — its mutation stays
green — but createTempFile is 0600 on POSIX by default, so the not-world-readable property
holds without it; the line only preserves a non-default mode.
Round 1 used plain bidirectional containment. The direction that catches the real bug --
the pane holds the brief plus a status bar, so the scrape contains the brief -- also fires
when a member restates the whole brief and then writes a genuine report under it. That
threw the report away and told the lead nothing was produced, which is worse than the bug
being fixed: it destroys a delivery instead of merely obscuring one.
The safe direction (the scrape is a fragment of the brief) stays unbounded, because a
fragment of the brief is by definition not a report. The dangerous direction now requires
the scrape to add at most MAX_ECHO_EXCESS_CHARS beyond the brief, which is the amount of
TUI chrome a real echo carries.
Work by the cb241 worker, committed by the lead: its backend stopped answering after the
fix was written, so two turns ended with no commit and no reply. Verified by the lead:
1169 tests, 0 failures; removing the bound turns pinsTheMaximumTuiChromeExcess and
completionFallbackKeepsARealReportThatRestatesTheWholeBrief red with 0 compile errors.
Files.writeString truncates the target in place before writing, so there
was a window where .claude.json could be observed empty or half-written
- exactly the shape of the incident this ticket already hit once, but
reachable in production too: a crash/kill mid-write, or two concurrent
claude-code spawns (normal here - several run in parallel routinely)
racing a naive read-modify-write and silently discarding one spawn's
entry.
Two independent fixes, each with its own dedicated test proving it (not
the other):
- ClaudeCodeLauncher.writeAtomically: serialise to a sibling temp file in
the same directory, then Files.move with ATOMIC_MOVE + REPLACE_EXISTING,
preserving the target's existing POSIX permissions (.claude.json ships
0600). A reader now only ever observes the fully-old or fully-new file,
never a torn one. Package-visible so a test can drive it directly.
- TRUST_JSON_LOCK: a process-wide lock around seedTrustDialog's whole
read-modify-write, so two concurrent spawns for different cwds both
keep their entry instead of the second write discarding the first.
Sufficient because every spawn on this daemon runs in one JVM; it does
NOT protect against a second daemon process or the operator's own live
Claude Code writing at the same instant - writeAtomically covers that
case instead.
Both fail soft, same as before: any I/O failure here must never block a
spawn.
Four new tests: a large (30-project) existing file survives without
collapsing (asserted on the restored key set, not just that the result
parses); two concurrent spawns for different cwds both keep their entry
(CountDownLatch-synchronised, not a sleep); existing 0600 permissions
survive the write; and a direct test of writeAtomically with a busy-poll
reader thread proving a concurrent reader never observes a torn file.
See PR body for the full mutation-testing table, including an
honest note on which of these tests the atomicity mutation actually
caught (not the one implied by the numbering in review) and why.
The daemon log and the worktree git config are both new in 205ad82, but neither helps a
lead who is writing a brief. This is the line that does: never brief a worker to edit
.mcp.json, opencode.json or .autoenv in its worktree, because the edit cannot be
committed and nothing will say so.
Goes in the project addendum, not the canonical block, so the byte-sync with
wiki/7-Use-Cases.md is unaffected — re-checked and still true.
isolateToolSurface replaces .mcp.json, opencode.json and .autoenv with stubs in every
provisioned worktree and marks them --skip-worktree. That neutralisation is correct and
is unchanged here — the committed files would mount the primary's credentials.
The problem was that it was invisible. A worker told to edit opencode.json read a 3-byte
stub and reported, truthfully and wrongly, that the mount key did not exist. A missing
file would have prompted a question; a plausible stub did not.
Two changes, both visibility only. The daemon now logs one info summary per provisioning
with the denominator, the files neutralised, and the consequence. And the list is recorded
in worktree-scoped git config (fleet.neutralizedConfig / fleet.neutralizedConfigNote) so a
worker can discover it from inside its own worktree with
'git config --worktree --get-all fleet.neutralizedConfig'.
Worktree-scoped config was chosen over a file in the working tree because it lives in
.git/worktrees/<nonce>/config.worktree and so can never appear in git status, and because
configureEnvironmentCredentialHelper already uses the same mechanism in the same add() call.
Verified by the lead: baseline 1172 tests, 0 failures. Reverting the summary to log.debug
goes red (2 tests), and recording into --local rather than --worktree — which would leak
the record into the shared repo config — goes red too. Both with 0 compile errors.
isolateToolSurface replaced .mcp.json/opencode.json/.autoenv with neutral stubs and
marked them --skip-worktree, but said nothing anywhere. A real worker read a 3-byte
{} stub for opencode.json, where the repo's real file is 30+ lines, and truthfully
(but wrongly) reported a mount key did not exist.
Two readers, two fixes:
- the daemon operator gets one info log per provisioning, naming the denominator,
what was neutralized, and why anything was not (same shape as overlayParity's
fix in #148 point 3).
- the worker gets the same fact recorded in worktree-scoped git config
(fleet.neutralizedConfig / fleet.neutralizedConfigNote), discoverable with
`git config --worktree --get-all fleet.neutralizedConfig` from inside its own
worktree, without asking the lead. Not a working-tree file: this repo already
uses worktree-scoped config for the credential helper and the SSH->HTTPS
rewrite, and it lives under .git/worktrees/<nonce>/ so it can never appear in
`git status` for the worker to trip on or commit.
The neutralization itself (stub content, --skip-worktree marking) is unchanged.
A claude-code member spawned into a fresh worktree hits an interactive,
un-timed workspace-trust prompt on its first start in a directory it has
never seen. It never reaches its first turn and never mounts the bridge.
Fix: ClaudeCodeLauncher.seedTrustDialog writes
projects.<cwd>.hasTrustDialogAccepted / hasCompletedProjectOnboarding into
the profile's configDir/.claude.json (or ~/.claude.json when configDir is
unset) BEFORE the herdr spawn call, additively (existing keys/projects are
preserved). Gated to isProvisionedWorktree(cwd) - a .git that is a regular
gitdir-pointer file, never a real checkout's .git directory - the same
signal writeIdeOverlay already used, now shared between both.
That gate is a fix for a real incident hit while building this: an
earlier ungated version ran against this file's own pre-existing tests
(configDir=null, no cwd -> falls back to the real user.dir and
~/.claude.json) and corrupted the operator's actual ~/.claude.json down
to a single entry during a mutation-testing run. See PR body for the
full incident report.
FakeHerdr gained onAgentStart(Runnable) so a test can assert the seed
is on disk at the exact instant herdr's agent.start call is reached -
i.e. strictly before the peer process itself would start.
Both defects lived in one method. overlayParity logged every step at debug, so at the
default level the copy was silent and nobody could tell which overlay files a member
actually got. It also marked a copied tracked file --skip-worktree and said nothing, so a
worker editing that file later found git ignoring the change with no error anywhere.
The summary now reports the denominator, not a bare count: 'copied 1 of 2 candidates:
.env (.envrc absent)'. A bare 'copied 1' is the same under-reporting shape as #113.
Neutralised files are named with the consequence in the message itself.
No marker file is written into the worktree: acceptance criterion 1 requires the worktree
to hold exactly the configured overlay set, so a marker would violate the fix it documents.
The copy and mark logic is unchanged — only logging is new.
Verified by the lead: baseline 1168 tests, 0 failures. Reverting either log.info to
log.debug goes red (2 reds and 1 red, 0 compile errors each), which is the regression that
matters since the whole fix is the log level.
overlayParity logged everything at debug, so at the default level nobody could
tell which overlay files a spawn actually received (#148 pt 3), and a tracked
file marked --skip-worktree gave no warning that it can no longer be edited
from that worktree (#134).
Report the outcome at info: a per-spawn summary naming the denominator (every
configured candidate), what was copied, and why anything was not — plus a
separate line naming every file marked --skip-worktree, stating plainly that
it cannot be committed from this worktree. No worktree-local marker file: the
worktree must hold exactly the configured overlay set and nothing else, so an
extra file would violate that invariant.
CompletionResolver already reads the pane on every turn, so the classifier lives there
rather than in a new watcher. A matched pattern resolves the waiter as a failure and
fires BackendErrorSink, always inside the resolveFailure win-gate so exactly one thread
reports one incident.
The too-fast path now takes a fresh scrape instead of reusing the pre-turn text, so a
backend that dies immediately is still classified. BackendErrorSink is a single-method
functional interface by design — see #234 for what a default overload does to a lambda.
An exhausted opencode member could not always be mapped back to a credential, because
the launcher knew the profile and the sink did not. The sink now takes a profile hint.
The interface is inverted on purpose: the three-argument method is the single abstract
method and the two-argument one is the default. A lambda can only implement the abstract
method, so every lambda is now forced to carry the profile. The first round of this fix
added the third argument as a default overload, and the production forwarder in Fleetd
was a two-argument lambda — so the fix compiled, passed its tests, and never ran.
Verified by the lead: deleting the forwardingTo factory's override, and rewriting the
Fleetd call site as a plain lambda, both fail to compile now rather than passing silently.
Backend incidents become a fourth source inside ReplyPushLoop, not a new scheduler — a
second injector would race the one control that already owns lead-pane delivery. One
notice per incident per affected lead (not per member), one-shot, waiting while the lead
pane is not injectable, and combined into the same nudge as any pending failed ticket.
A classified target that cannot be mapped to a credential gets its own truthful notice
and its own pending/delivered records. It previously reused the incident message, which
told the lead a credential was 'cooling for 0 remaining seconds' when nothing was cooling,
and smuggled the free-text reason into the profiles field.
Verified by the lead: 1129 tests green; routing the unmapped path back through
onBackendIncident turns unmappedBackendTargetUsesTheKnownLeadSchedule red with 0 compile
errors, and that test now asserts the whole rendered message rather than two substrings.
Adds BackendOutagePolicy — a credential-keyed state machine on an injected monotonic
clock. Two classified backend errors from two DISTINCT targets on one credential inside
60 seconds mint one incident and start a 60-second cool-off. Errors during cool-off
neither extend it nor mint another; expiry clears evidence, so two fresh errors rearm.
Correlated on credentialId, never on profile name or error text. Deliberately not
BackendQuarantine: that restarts a 1800-second cooldown per exhaustion, and its name
would make every refusal say 'backend exhausted', which is a different condition.
Threshold counts distinct targets rather than raw events (lead decision): the classifier
is a heuristic and a valid member report can quote an 'API Error:' line, so one member
repeating that line must not remove a healthy credential's capacity. A real outage hits
every member on the credential, so true detection is unaffected.
Verified by the lead: 1135 tests green; reverting evidenceCount() to reasons.size()
turns two BackendOutagePolicyTest cases red with 0 compile errors.
Round 3's factory fixed the two known call sites but the underlying shape
was still there: a lambda written against ExhaustionSink binds to whichever
overload is abstract, and the 2-arg form held that position, so ANY lambda
-- a call-site forwarder, a hand-built test double, a future caller who has
never heard of fleetd #234 -- could still silently take the hint-dropping
default. Two rounds shipped exactly that mistake in two different places.
Fix: made the 3-arg onExhausted(target, reason, profile) the interface's
single abstract method; the 2-arg form is now a default that delegates with
a null profile. A lambda declared against ExhaustionSink today is forced by
the compiler to take three parameters -- there is no overload left for it to
bind to that can drop the hint. This is enforced by the type system, not by
a test that has to remember to check for it.
Knock-on changes:
- ExhaustionSink.none() -- a 3-arg lambda, still a genuine no-op, now safe
by construction rather than by care.
- ExhaustionSink.forwardingTo(...) -- collapses to a one-line 3-arg lambda;
kept as a named factory (round 3's lesson: a test must call the real
object, not rebuild its shape).
- Fleetd.java's real sink and the two OpenCodeLauncherTest sinks that used
to be anonymous classes overriding both overloads are now plain lambdas
too -- the 2-arg override each carried was pure boilerplate once the
interface provides it as a default.
- CompletionResolver.java itself: UNCHANGED, zero diff (confirmed via
`git diff --stat` before staging) -- its two call sites still call the
2-arg onExhausted(target, reason), which is now the default and behaves
identically. CompletionResolverTest (41 tests, 0 failures) proves this;
its five ExhaustionSink lambdas needed a mechanical third parameter added
to keep compiling against the new abstract method, no assertion changed.
Mutation proof, re-run against the new shape: forwardingTo's body edited to
call the 2-arg default instead of passing the hint through (the equivalent
of round 3's "delete the 3-arg override" now that there is only one method
to break) -- both new tests go red with the same assertions as round 3:
ExhaustionSinkForwardingHazardTest...: expected: <gx> but was: <null>
OpenCodeLauncherTest...ForwardingHop: expected: <true> but was: <false>
Tests run: 68, Failures: 2
Restored, re-ran: green (Tests run: 109, Failures: 0, including
CompletionResolverTest).
Compiler proof (not committed -- a scratch file outside the worktree,
compiled with the real ExhaustionSink.java on the classpath, then deleted):
ExhaustionSink forwarder = (target, reason) -> System.out.println(target + reason);
error: incompatible types: incompatible parameter types in lambda expression
A 2-arg lambda against this interface no longer compiles at all.
mvn clean install: Tests run: 1129, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
Adds MemberSession.State.BACKEND_ERROR with a nullable failureReason, surfaced in
rosterView, and SessionManager.onBackendError(target, reason). The CAS loop accepts
both sides of the completion race (BUSY and DONE); BACKEND_ERROR is terminal.
completeTurn now returns early when its CAS loses, so a stale DONE copy can no longer
release the pane or reset its context behind a member that just went BACKEND_ERROR.
Verified by the lead: 1130 tests green; mutating the completeTurn early return back to
the old fall-through turns losingCompletionDoesNotReleaseOrClearABackendErrorMember red
with 0 compile errors.
Two errors from the same target inside the window must never trip the
outage threshold on their own (a valid member report can legitimately
quote an "API Error:" line twice) — only two DIFFERENT targets on the
same credential do. Change evidenceCount() to targets.size() instead
of reasons.size(); reasons() still keeps every event, including
same-target repeats, so it can be longer than evidenceCount(). A real
outage still hits every target on the credential, so this loses no
true-positive coverage while cutting a real false-positive path.
Replace the hardcoded API-Error check with a target-keyed BackendErrorPatternLookup
plus a BackendErrorSink, mirroring the existing ExhaustedPatternLookup/ExhaustionSink
pair. Classifies in all three paths (normal block, #211 raw-scrape fallback, and the
fleetd#164 MIN_TURN_NANOS floor). The sink fires only after Rendezvous.resolveFailure
wins for the exact waiter. A target with no configured pattern still falls back to the
narrow (?i)\bAPI Error\s*: compatibility pattern. Existing constructors keep compiling
via BackendErrorPatternLookup.legacy() / BackendErrorSink.none() defaults.
Public send result is unchanged (still a failed send) — the typed sink event is the
internal seam Unit 5 will consume.
Add BackendOutagePolicy: two classified backend errors on the same
credentialId within a 60s window mint one Incident and start a 60s
cool-off for that credential, one atomic ConcurrentHashMap.compute()
per credentialId so a concurrent second and third event can never
both cross the threshold. Errors during cool-off are ignored outright
(no extension, no incident); once cool-off elapses the next error
clears old evidence, requiring two fresh errors to rearm. This is a
new class, deliberately not BackendQuarantine (wrong store, wrong
1800s duration, misleading "exhausted" semantics for a 60s transient
fault). Knows nothing about panes, profiles, sessions, launchers, or
leads — takes events in, returns incidents out.
Round 2's tests never reached Fleetd.java at all: both new tests declared
their OWN local copy of the forwarding shape instead of calling production's.
Mutating Fleetd.java's real forwarder back into the broken lambda left those
copies untouched, so the whole suite stayed green while production had
regressed to exactly the bug being fixed -- proven live by the reviewer.
Fix: extracted the forwarding shape into one named factory,
ExhaustionSink.forwardingTo(Supplier<ExhaustionSink> target), with the
"why a lambda here is wrong" explanation moved onto it (the one place the
shape is now written). Fleetd.java's forwarder collapses to one line:
ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);
Both new tests now call this same factory instead of rebuilding an anonymous
class inline, so they exercise the identical object production builds:
- ExhaustionSinkForwardingHazardTest: calls ExhaustionSink.forwardingTo
directly and asserts the hint reaches the real sink through it.
- OpenCodeLauncherTest#theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop:
same factory call, inside the full Fleetd-shaped construction order
(forwarder built first, real sink pointed at via the AtomicReference
afterward), driven through the real SessionManager.acquire() path.
Mutation proof, this time on production code only: deleted the factory's
3-arg override (falls back to the interface default, dropping the hint) --
both new tests go red with no test file touched:
ExhaustionSinkForwardingHazardTest...: expected: <gx> but was: <null>
OpenCodeLauncherTest...ForwardingHop: expected: <true> but was: <false>
Tests run: 68, Failures: 2
Restored, re-ran: green (Tests run: 68, Failures: 0). Confirmed Fleetd.java
carries no lambda ExhaustionSink anywhere (grep). ExhaustionSink.none() stays
a lambda on purpose -- both its overloads are true no-ops regardless of
arity, so there is no hint to drop.
mvn clean install: Tests run: 1129, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
The round-1 fix was dead on the real production path. Fleetd.java:177 builds
a forwarding sink (needed because the adapters are constructed before
`sessions` exists, breaking a genuine cycle) as a LAMBDA:
ExhaustionSink forwardingExhaustionSink =
(target, reason) -> exhaustionSinkRef.get().onExhausted(target, reason);
A lambda can only implement the interface's one abstract method (the 2-arg
overload), so it silently inherited the 3-arg overload's default body, which
drops the profile hint and calls back into the 2-arg method. OpenCodeLauncher
is constructed with this forwarder, so the hint it supplies (its own
already-known profile name) was thrown away before it ever reached the real
sink built later in Fleetd.main -- reproducing the exact silent no-op round 1
was sent to fix. The 1127 tests from round 1 all injected a sink directly
into OpenCodeLauncher and never went through this forwarding hop, so none of
them could see it.
Fix: forwardingExhaustionSink is now an anonymous class overriding both
overloads, each delegating to whatever exhaustionSinkRef currently holds.
Audited every other ExhaustionSink value in main/: the only other one is
ExhaustionSink.none() (a lambda), which is safe regardless of arity since
both its 2-arg body and the inherited 3-arg default are true no-ops.
New tests:
- ExhaustionSinkForwardingHazardTest: isolates the hazard at the interface
level (a lambda forwarder drops the hint; an anonymous-class forwarder
does not), independent of Fleetd.java's specific wiring.
- OpenCodeLauncherTest#theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop:
replicates Fleetd.java's actual construction order (forwarder built and
handed to the launcher first, real sink built and pointed at via the
AtomicReference afterward) and drives the quarantine through it via the
real SessionManager.acquire() path.
Both proven by mutation: temporarily rewriting each fixed forwarder back
into the pre-fix lambda makes its test fail with a real assertion message
(both matched exactly: "expected: <gx> but was: <null>" for the interface
proof, "expected: <true> but was: <false>" for the composed-wiring test);
restoring makes it pass again. No reverts were committed.
mvn clean install: Tests run: 1130, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
Defect 1: OpenCodeSessionDiscovery.actualModelForDirectory queried
WHERE directory = ?, the same heuristic sessionIdForDirectory uses. Since a
default fleet_spawn (no worktree:) shares the lead's cwd with every other
worker and every past session ever run there, the model read-back could
silently compare against a DIFFERENT session's row. Renamed to
actualModelForSessionId(sessionId), keyed on the primary key id instead, and
made OpenCodeLauncher's SessionAwareHandle cache the resolved id once
non-null (AtomicReference) so a later sibling row in the same directory can
never flip which session's evidence is read. sessionIdForDirectory (#209) is
left directory-based on purpose, with a comment explaining why the heuristic
is unavoidable at that layer.
Defect 2: the ERROR log claimed "quarantining this profile's credential" but
Fleetd's ExhaustionSink lambda resolved target -> roster -> profile ->
credential, while OpenCodeLauncher's model-mismatch check fires from
agentSessionId() during SessionManager.acquire(), before the session is
registered in the roster -- the lookup found nothing and silently no-opped.
Added a default 3-arg ExhaustionSink.onExhausted(target, reason, profile)
overload (defaults to the 2-arg method, so CompletionResolver's two call
sites are unchanged); OpenCodeLauncher now passes its own already-known
profile name; Fleetd's sink became an anonymous class that tries the roster
first, falls back to the hint, and logs loudly at ERROR naming target/reason
when neither resolves, instead of silently no-oping.
Both fixes proven by mutation: reverting each independently makes its new
test fail with a real assertion message, restoring makes it pass again.
mvn clean install: Tests run: 1127, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
tokenEnv no longer defaults to FLEETD_WORKER_TOKEN. A profile that names no
tokenEnv is stating it needs none, which is different from one naming a variable
that turns out to be unset. requiredSecretEnvVars now warns only for an explicitly
declared tokenEnv, so the permanent false alarm about a token no profile needs is
gone and the real warnings beside it stay trustworthy.
Decision recorded in a code comment: an explicit tokenEnv is checked for every
kind, opencode included. opencode can use its own provider credentials, but an
explicit tokenEnv declares a required host secret for its configured provider.
Second commit fixes a crash the first one introduced. Making tokenEnv nullable
changed what every reader of that value can receive, and ClaudeCodeLauncher:261
passed it straight to env.apply — System::getenv in production, which throws on a
null name. The live profile local-direct (kind: claude-code, baseUrl set, no
tokenEnv) would have crashed on spawn. It is weight: 0 today, so the failure would
have surfaced whenever someone re-enabled it. Now uses the superclass helper
resolveEnv, which already tolerates a null name and which OpenCodeLauncher was
already using.
Reader audit, all four: ClaudeCodeLauncher fixed; OpenCodeLauncher already safe via
resolveEnv; ConfigRef uses Objects.equals; MemberEnvAllowList drops null and blank
names in addIfPresent.
Verified by the lead before merge: reverting the resolveEnv fix makes the new test
fail with a NullPointerException from the null variable name, and restoring it
passes. Independent build: 1125 tests, 0 failures, 0 errors, 0 skipped.
Not verified: the ticket's acceptance criterion 4, a real boot on this host showing
no FLEETD_WORKER_TOKEN line while the WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN lines
are unchanged. A worker cannot restart the daemon it talks through. The lead checks
that at the next redeploy.
Two fixes to HerdrPeerLauncher.waitUntilInjectableOrThrow.
Fix 1: the status call is now guarded. A herdr *_not_found answer — what happens
when the backend process exited rather than being slow — used to escape as a raw
HerdrException, skipping stop() and leaking the pane and tab, with a message that
blamed a slow pane. It now fails immediately, runs the same teardown, and says the
process exited.
Fix 2: the gate can now resolve UNKNOWN with the same StatusRefiner the status
poller uses, behind two guards. It only runs for the claude adapter, because
StatusRefiner.classify reads a Claude Code TUI. And a refined result is accepted
only when the same agents.get sample still reports a non-null agentType. That
second guard closes a trap: a dead pane sits at a shell prompt containing the same
❯ glyph the classifier reads as idle, so refining without corroboration would turn
"the backend died" into "ready to inject".
Verified by the lead before merge: NAME_PREFIX really is "claude"/"opencode" so the
adapter split holds; AgentControl.status(t) was already get(t).status(), so moving
the loop to get() adds no herdr call; and a *_not_found already failed the spawn
before this change, so fix 1 improves an existing failure rather than creating one.
The adapter guard was mutation-tested independently (set it to `if (false)`, the new
opencode test fails with "Expected PeerUnreachableException to be thrown, but nothing
was thrown"; restored, it passes). Independent build: 1123 tests, 0 failures,
0 errors, 0 skipped.
Not fixed, and not claimed: that herdr reports a null agentType for a bare shell
pane is unverified against a live daemon. To be proved by a live spawn on both
backends after deploy. The seat-accounting suggestion in #176 was deliberately not
built — its cause was tested in that issue and not reproduced.
The agentType corroboration (guard b) had positive and negative tests, but the
per-adapter namePrefix guard (guard a) had none — nothing proved that an
opencode pane can never reach StatusRefiner.classify, only that a claude pane
with a null agentType is rejected. Since a live opencode pane always reports a
non-null agentType ("opencode"), guard (b) alone cannot catch a broken guard
(a).
Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt
to OpenCodeLauncherTest: agentType("opencode") (non-null, satisfies guard b on
its own) + pane content containing "❯" (would classify as IDLE) + raw status
UNKNOWN throughout. Asserts the gate still times out, and that no agent.read
call used source=detection (StatusRefiner.PROBE_SOURCE) — proving the refiner
was never even reached, not just that its answer was discarded. A blanket
"agent.read is never called" does not hold here: HerdrPeerLauncher.readPaneQuietly
reads the pane tail (source=recent) for the timeout log on every timeout,
regardless of adapter, so the assertion is scoped to the refiner's own probe
source instead.
Mutation check performed and reverted before this commit: temporarily changed
`if (!"claude".equals(namePrefix))` to `if (false)` in
HerdrPeerLauncher.refinedInjectable — the new test failed
("Expected PeerUnreachableException to be thrown, but nothing was thrown"),
confirming it actually exercises the guard. Restored the guard and reran —
test passes again (1/1). No production code changed in this commit.
mvn clean install: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
Fix 1: agents.get(paneId) inside waitUntilInjectableOrThrow was unguarded, so a
herdr *_not_found answer (the backend process exited) propagated as a raw
HerdrException instead of PeerUnreachableException, and skipped teardown
entirely, leaking the pane/tab. Now caught via isAlreadyGone(); fails
immediately (does not burn the rest of the timeout), runs the same teardown
the timeout path runs, and the exception message says the process exited
rather than that the pane was slow. Any other HerdrException still
propagates unchanged.
Fix 2: the gate now resolves a raw UNKNOWN into StatusRefiner's pane-content
classification, like StatusPoller already does mid-life. Guarded against the
trap noted in the ticket: a pane whose backend exited settles at a bare shell
prompt that can also contain the "❯" glyph classify() reads as idle. So a
refined result is accepted only when the corroborating agentType from the
SAME agent.get sample is non-null, and only for the "claude" adapter (namePrefix)
since StatusRefiner.classify is written for the Claude Code TUI only. Refine
only runs when the raw status is UNKNOWN, so a healthy spawn adds zero extra
herdr calls.
Tests added to ClaudeCodeLauncherTest (FakeHerdr gained agentType()/
agentGetFailsWithAfter() fixtures):
- spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout
- spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged
- refinedIdleIsNotAcceptedWhenAgentTypeIsNull
- refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness
- refinementNeverFiresWhenRawStatusIsAlreadyInjectable
mvn clean install: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
Verified by the lead before merge. Read the full production diff, checkModelMatch, parseModel and actualModelForDirectory. Confirmed the check runs on the real #209 late-resolve path (not at spawn, which is why #203 was closed), that unknown/incomplete evidence never quarantines, and that claude-code is structurally excluded because SessionAwareHandle is only built by OpenCodeLauncher.spawn(). Round 2 closed the one gap I found: model JSON with an id but no providerID used to read as a mismatch for a provider-prefixed profile. Independent build: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped.
parseModel already tolerates a model JSON with an id but no providerID (a real shape
opencode can write). checkModelMatch's providerMatches check did not: a provider-
prefixed profile whose id matched but whose evidence had no providerID was reported
as a mismatch and quarantined on incomplete data, which acceptance rule 4 forbids.
Compare the provider only when BOTH the profile requested one AND the evidence has
one. A genuine id mismatch is still caught either way — narrows the check, does not
disable it.
opencode does not fail on an unknown -m <model> flag — it silently falls back to a
default model, which can be a paid credential. Extends the existing late-resolve
path (#209's SessionManager -> handle.agentSessionId() re-poll) so that once the
opencode session row exists, OpenCodeSessionDiscovery also reads its `model` JSON
column and OpenCodeLauncher's SessionAwareHandle compares it against the profile's
configured model.
Comparison rule: split the profile's model on the first '/' into provider+id. Compare
id always; compare provider only when the profile specified one. A bare model name
with no '/' matches on id alone. Absent/unparseable evidence is UNKNOWN, never a
mismatch, so a working profile is never quarantined on missing data. A real mismatch
logs an ERROR naming both models and the profile, then quarantines through the
existing ExhaustionSink path (wired via an AtomicReference forwarding sink in
Fleetd.java to break the sessions/workers/adapters construction cycle).
Verified by the lead before merge: read the full production diff, confirmed the memberHerdrSocket-absent branch is the literal unmodified Files.createTempFile call in its own branch, and that the refusal names the missing key. Measured behaviour recorded: claude 2.1.258 exits 1 immediately on an unreadable --append-system-prompt-file, so the pre-fix bug was the loud readiness-gate failure, not a silent charter-less member. The /tmp full-suite failure the worker reported was the two other #225 copies, fixed by #230 which is already on main; main is built and checked after this merge.
Verified by the lead before merge. Read the full production diff: locking is consistent (every MemberRegistry method uses synchronized(terminalToSlot)), and the refusal happens before the launcher starts a process. Proved the restored fallback test is real by removing the DEV fallback from SessionManager and re-running it: it failed with "expected: <DEV> but was: <ARCHITECT>", then reverted. Independent build: BUILD SUCCESS, 1092 tests, 0 failures, 0 skipped.
The helper I copied from OpenCodeLauncherTest read the CWD's owning
group instead of the process's real primary group, so it silently
picked up whatever group owns the directory Maven was started from
(staff in a home checkout, wheel under /private/tmp on macOS) rather
than a group the operator is actually in. Replaced with the id -gn
based resolution that landed on #230 for the other two copies of this
helper, same shape and skip wording.
Verified by the lead before merge: read the full production diff, confirmed `group` is normalised to null at GitWorktrees:148 so the `group == null` guard is complete, and confirmed the refusal runs before `git worktree add` so a failure leaves no half-made worktree. Independent build in the worker's worktree: BUILD SUCCESS, 1093 tests, 0 failures, 0 skipped.
#224: GitWorktrees#add created worktreeRoot with the daemon's umask and never shared it with
worktreeGroup, even though shareWithGroup shares every child underneath it (each worktree, and
the repo's common git dir). Under memberHerdrSocket: the member pane runs as a different OS
user, which needs execute on every ancestor directory to reach anything underneath, no matter
how carefully each child is shared — so a member could not read the opencode.json #219 places
under this root, could not reach its own worktree, and could not read #213's ZDOTDIR scrub when
placed here either.
Fix: add() now calls a new shareRootWithGroup(root) right after creating the root, chgrp+chmod
g+x on the root itself (non-recursive — each child is still shared individually by its own call
site). No-op when worktreeGroup is unset, so behaviour is byte-identical in today's only live
mode. On failure (group missing, or operator not a member of it) the spawn is refused with a
WorktreeException naming the root, its current mode, and the group — mirroring shareWithGroup's
existing refusal shape — before `git worktree add` ever runs, so no partial worktree is left
behind.
Also adds the assertion the #221 reviewer flagged as missing: a test driving
EnvAllowListScrub#shareWithGroup directly against a directory holding several flat files
(opencode.json, member-charter.md, ide-rules.md, plus an unrelated one) and asserting every one
of them gets group-readable/never-group-writable permissions, not just the two files someone
happened to think of.
#225: OpenCodeLauncherTest/HerdrPeerLauncherAllowListWiringTest's currentUserGroup() read the
group that owns the current working directory, not the process's own primary group, despite its
comment claiming the latter. Those coincide only by accident: a home checkout is typically owned
by a group the operator belongs to (staff), while a checkout under /private/tmp on macOS is
group wheel, which the operator is usually not a member of — so the same test fails for real
depending on where the repo happens to be checked out, and the existing assumeTrue only guarded
against "no POSIX groups at all", never "a resolvable but wrong group". Fixed by resolving the
process's REAL primary group via `id -gn` instead, with assumeTrue (skip, not fail) only when
that itself cannot be resolved on the host. The permission assertions these tests exist for are
unchanged.
Verified `mvn clean install` green from both a home checkout and a /private/tmp copy (mirroring
the exact repro in #225): 1093 tests, 0 failures, 0 errors in both locations.
ClaudeCodeLauncher#writeCharterFile used Files.createTempFile with no
directory argument, which resolves against fleetd's own java.io.tmpdir
(macOS: the per-user $TMPDIR, mode 0700). Under memberHerdrSocket: the
member pane runs as a different OS user and cannot read that directory,
and since #220 the charter file is the ONLY delivery path for
--append-system-prompt-file. Following #219's refusal decision (a
charter is the member's turn contract, not a degradable control): with
memberHerdrSocket configured, the charter now goes into a fresh
per-spawn directory under worktreeRoot, shared read-only via
EnvAllowListScrub.shareWithGroup (reusing #213/#219's mechanism); a
missing worktreeRoot/worktreeGroup refuses the spawn by name instead of
writing an unreadable file. With memberHerdrSocket absent the path is
unchanged.
An explicit-profile spawn bypasses role-pool placement (CompositePeerLauncher
only constrains an UNQUALIFIED spawn to fleet.<role>), so it was the one path
that could ask for role=architect on a profile no architect slot carries.
MemberRegistry silently held the session as a plain worker while GET /members
still reported the requested "architect" and only fleet_whoami (which reads
live bindings, not the request) told the truth.
- MemberLifecycle.requireSlotFor(role, profile): refuses the acquire before
anything spawns when no configured architect slot carries the profile,
naming the role, the profile, and the pools that do carry it. No-op for
dev/reviewer, which are placement candidates only, never a live identity
binding — refusing a profile mismatch there would break the documented
fleet_spawn{profile:"opus"} (role defaults to dev) flow.
- MemberLifecycle.acquired(...) now returns the role the session actually
holds, so a residual race (a slot exists but every instance is already
bound to a different terminal) still falls back to dev honestly instead of
lying — this case logs at WARN (was INFO), naming profile and terminal.
- SessionManager now records the role acquired() returns on MemberSession,
never the requested role, so GET /members and fleet_list can no longer
report a role the member does not hold; no changes needed to memberView/
rosterView, which just read session.role().
An architect's identity IS the slot it is bound to — binding a role with no
slot to bind means inventing an identity out of nothing, which is the quiet
failure the whole role system exists to prevent.
Tests: SessionManagerTest and FleetMcpTest each drive a real spawn through
FleetMcp.spawn -> SessionManager.acquire -> the real ClaudeCodeLauncher (via
FakeHerdr), then assert on GET /members and fleet_whoami for that same
session — not on MemberRegistry.bind directly (fleetd issue #113's mistake).
Site 1 (config root): under memberHerdrSocket, writeConfig() now places the
ephemeral opencode.json directory under worktreeRoot and shares it read-only
with worktreeGroup, reusing EnvAllowListScrub#shareWithGroup (widened to
package-private and generalized) — the same mechanism #213 built for the
ZDOTDIR scrub, rather than a second copy. Unlike the ZDOTDIR scrub's
degrade-to-overlay fallback, a missing worktreeRoot/worktreeGroup here
REFUSES the spawn (IllegalStateException from buildLaunch): this file is the
member's only way to learn where the bridge MCP is, so writing it somewhere
unreadable would just produce an undeliverable member with no signal
pointing at the cause. memberHerdrSocket absent stays byte-identical.
Site 2 (discovery root): under memberHerdrSocket, agentSessionId() now
declares session discovery unavailable and logs one WARN per launcher
instance instead of silently scanning fleetd's own $HOME (opencode.db lives
under the MEMBER's home under this config key). Decision + reasoning for why
this is a declare-unavailable rather than a new config key is in
defaultDiscoveryRoot()'s javadoc.
Widened HerdrPeerLauncher#memberHerdrSocketConfigured/memberScrubParentDir/
memberGroup to package-private so OpenCodeLauncher reuses the exact same
config resolution rather than re-deriving it.
Same-shape finding (not fixed, out of scope): ClaudeCodeLauncher#writeCharterFile
(line ~465) writes the role-charter temp file via Files.createTempFile with no
directory argument, i.e. under java.io.tmpdir — the same site-1 shape, unfixed
for the Claude Code adapter.
herdr does not exec a member's launch command — it TYPES it into the pane,
and a pty line buffer holds 1024 bytes (BSD/macOS MAX_CANON). Past that the
tail is dropped and NOTHING reports it: herdr answers "agent started", the
backend exits on the mangled argument it was handed, the pane closes, and the
only symptom is the readiness gate timing out 20 seconds later with no reason.
That is what broke every claude-code spawn after #214. The reply charter rode
inline on --append-system-prompt, so the command was already 978 bytes; adding
--session-id <uuid> made it 1028, and the 4 bytes cut off the end turned
--autocompact 250000 into --autocompact 25, which claude rejects. Measured on
the live pane, the cut is at byte 1024 exactly.
- ClaudeCodeLauncher: the charter ALWAYS travels as --append-system-prompt-file.
The file path already existed for the two-charter case; the inline form only
ever saved a temp file, and it cost ~800 bytes of the line budget. This takes
the prose off the command line for good.
- HerdrPeerLauncher.checkPaneCommandFits: refuse a command that cannot fit,
naming the byte count and the longest argument, instead of spawning something
that cannot work. The estimate is deliberately conservative — fleetd cannot
see herdr's quoting, and an under-estimate would let the silent truncation
back in.
- HerdrPeerLauncher.waitUntilInjectableOrThrow: log the pane tail and the last
herdr status BEFORE stop() closes the pane. Without it the gate reports only
that it timed out, which is true of every cause. This is what found the bug,
and it stays.
The guard also catches a case that was already over the limit: a profile with
ideMcpUrl set assembles 1084 bytes. It is now impossible to ship that silently.
3 tests, all watched failing first: with the inline charter restored the guard
fires in the new fit test, in the pre-existing autocompact test and in the IDE
mount test. Full suite 1081 tests green. Proven live: sonnet spawns again, the
member obeys the file-delivered charter and ends its turn with fleet_reply, and
fleet_list reports the #214 agentSessionId.
The normal backend-error path appends the pane tail to the failure reason on
purpose (fleetd#164): the BACKEND_ERROR pattern is a heuristic, and a member
that reported *about* an error while forgetting fleet_reply matches it too, so
dropping the rest of the pane destroys the report.
The new raw-scrape fallback did not do that. It matters more there, not less:
the fallback only runs when the trimmed assistant block was empty, so the raw
scrape is the ONLY copy of whatever the member managed to say. A lead read the
matched line and nothing else.
Clipped to the same cap the normal path uses, since a raw screen has no
boundary trimming to bound its size.
The assertion was watched failing without the fix:
AssertionFailedError: fleetd#164: the failure must carry the pane, not only
the matched line ... expected: <true> but was: <false>
mvn clean install: Tests run: 1079, Failures: 0, Errors: 0, Skipped: 0
The memberCredentials.policy: allow-list ZDOTDIR scrub decided zsh-vs-not
using fleetd's own process $SHELL and wrote the generated scrub dir into
fleetd's own java.io.tmpdir. Under memberHerdrSocket: (member panes run as
a different OS user than fleetd's own process) this silently protects
nothing: the wrong shell decides the gate, and the directory can be
unreachable to the member.
- New FleetConfig.memberLoginShell: the member OS user's login shell,
only ever read when memberHerdrSocket: is configured; fleetd's own
$SHELL keeps deciding everything when memberHerdrSocket: is absent
(byte-identical to before).
- HerdrPeerLauncher.applyEnvironmentAllowListPolicy: memberHerdrSocket +
memberLoginShell not configured/non-zsh falls back to the CB-596
sentinel overlay (warn loudly, never refuse to spawn). memberHerdrSocket
+ zsh memberLoginShell generates the ZDOTDIR under the configured
worktreeRoot instead of java.io.tmpdir, shared with the existing
worktreeGroup (reused, not a new key).
- EnvAllowListScrub: new generate(parentDir, allowedNames, group) overload
shares the generated directory via pure-Java POSIX group ownership
(rwxr-x--- dir, rw-r----- files) — no external process spawn.
- fleetd.example.yaml documents memberLoginShell: and worktreeGroup:'s
reuse for the scrub directory (the live fleetd.yaml is gitignored).
4 new tests in HerdrPeerLauncherAllowListWiringTest cover the acceptance
criteria; 3 of the 4 were watched failing against the pre-fix code.
CompletionResolver.resolve() returned an empty-scrape failure before the
BACKEND_EXHAUSTED / BACKEND_ERROR classification ever ran, whenever
lastAssistantBlock() found no usable text — most commonly a pane with no ⏺
marker at all, whose boundary scan starts at the top of the raw screen and
breaks immediately on the first line of TUI chrome. Since BACKEND_EXHAUSTED
is the only caller of exhaustionSink, this meant an exhausted backend was
recorded as "produced nothing" instead of being quarantined.
Fix: run the same two classifications against the raw (untrimmed) scrape as
a fallback, only inside the empty-scrape failure branch. A pane that already
yields a usable assistant block never reaches this branch, so the existing
narrow match is unchanged. lastAssistantBlock stays the sole source of the
reply text; only classification ever consults the raw scrape.
abandon() drained the target's stranded reply once and then reused that same
Reply for every open task it walked past. Two open tickets on one target
therefore both came back REPLIED with the same text — one of them a reply the
worker never gave for that delegation.
One worker answer can settle at most one delegation. It now goes to the oldest
open task (lowest createdNanos) and every other open task keeps the ordinary
WORKER_FAILED path. If the chosen task turns out to be already resolved by
another path, the drained reply is published back to the inbox instead of being
dropped.
reply()'s matching side gets the same rule: more than one candidate means the
reply goes to the inbox rather than to a guess.
Reachability, checked rather than assumed:
- matching.size() >= 2 alone is reachable and was a real bug before this change.
- More than one candidate in reply() is not reachable today — hasAsyncQuestion
matches any task with a stamped turnId, and answer()'s
clearAsyncQuestion(turnId, false) leaves that stamp until the resumed turn
resolves. Kept as defensive code, documented, no test seam added.
- matching.size() >= 2 together with a live strand is not reachable either:
send() and answer() are the only two lock holders and both open a Rendezvous
waiter inside the lock, so "lock held" and "waiter open" are one fact, and an
acceptance always clears the strand first.
I checked the last point by building it in a scratch worktree: it can be forced
by closing an accepted send's waiter directly through Rendezvous, and it does go
red against the pre-fix loop — but that breaks the lock-and-waiter invariant
from outside the class, so no such test is added. A note in abandon()'s javadoc
says so, to save the next reader the same round trip.
Worker's REPORT-cb137.md left out of main.
mvn clean install: Tests run: 1071, Failures: 0, Errors: 0, Skipped: 0
Defect 1 (reply()'s askAnsweredAsyncTasks returning >1 candidate): confirmed
unreachable today. Documented why in three places — hasAsyncQuestion matches
any task with a stamped turnId (not just an open question), and answer()'s
clearAsyncQuestion(turnId, false) leaves that stamp in place until the
resumed turn's own future resolves — so a second task can never reach the
same eligible state while a first one holds it. Kept the defensive
inbox-fallback branch as defence in depth against that guarantee weakening,
per review instruction; no test seam added.
Defect 2 (abandon()'s broader `matching` filter applying a stranded reply to
more than one task): matching.size() >= 2 alone IS reachable (already
covered by abandonFailsEveryPendingAsyncTicketForTheReleasedTarget) and was
a real pre-fix bug (97f6c33's parent reused one drained reply for every
matching task). But hadStrandedReply == true together with matching.size()
>= 2, at the instant abandon() runs, is not constructible through the public
API: send() and answer() are the only two sites that ever hold a target's
session lock, and both open a Rendezvous waiter for that target as the first
thing they do while holding it — so "lock held" and "waiter open" are the
same fact throughout this class, and reply()'s fast path always resolves an
open waiter directly instead of stranding. A strand can only be created
while no task is accepted, and the moment the lock is next taken, that
acceptance clears the strand again before abandon() can observe both facts
together. Documented this in abandon()'s javadoc and removed the earlier
attempt at a deterministic test for the conjunction, whose apparent failure
was an invalid premise (the "accepted" task's own acceptance silently
cleared the strand it was meant to race against), not the fix being absent.
The CB-205-recovery fix in #205 assumed a target has at most one open
async task, with no guard. Fix three consequences:
- reply(): askAnsweredAsyncTask -> askAnsweredAsyncTasks (List). Exactly
one candidate completes it (unchanged). Zero falls to the inbox
(unchanged). More than one now ALSO falls to the inbox instead of
picking an arbitrary ConcurrentHashMap iteration order, and logs a
WARN naming the target and every candidate ticket.
- abandon(): a stranded reply now settles at most one matching task —
the oldest by Task#createdNanos (a new field, the tiebreaker). Every
other matching task keeps WORKER_FAILED, same as today.
- abandon(): if the chosen recovery task's complete() loses a race
(another path resolved it first), the drained reply is republished
to the inbox instead of being silently dropped.
Single-task behavior is unchanged; only the ambiguous case changes.
SessionManager called handle.agentSessionId() once at spawn and froze it in the
immutable MemberSession. For opencode that value is always null - the session
row does not exist yet when the pane is created - so fleet_list never reported
an agentSessionId and fleet_spawn{resumeSessionId} was unusable for that
backend. The handle's javadoc said 'the caller re-calls later'; no caller did,
and SessionManager did not even retain the handle.
Retain the PeerHandle per pane and re-resolve while the stored id is still
null, sticky once found, CAS-swapped into the registry.
roster() stays non-resolving - it is the roster supplier for LeadHeartbeatLoop
and FleetHealthMonitor, and resolving there would open opencode's 841MB SQLite
database on every tick, for every unresolved member, forever. rosterResolved()
carries the resolve and is used only by fleet_list and the REST roster, the two
surfaces that report the id. get(paneId) and release() resolve too, both
caller-driven.
A test pins the split: the plain roster() must never call agentSessionId()
again.
Under memberHerdrSocket: member panes run as a different OS user, so
hostEnvNames (fleetd's own environment) no longer describes what a member
inherits. logCredentialGap now reports "unknown, not clean" in that mode
instead of its usual UNBLOCKED/blanked conclusions, once per launcher, naming
the config key and scoping its count to fleetd's own process.
Byte-identical when memberHerdrSocket: is absent, pinned by a test.
The endpoint became /members in the CB-634 rename but the body key stayed
"workers", so a caller that read "members" saw an empty fleet and reported
no members at all.
Emit the canonical "members" key. Keep "workers" as a deprecated alias so an
existing REST consumer keeps working - the out-of-band path a lead falls back
to when its MCP mount drops reads this endpoint.
The test now pins both keys and asserts they carry the same rows, so the alias
cannot silently drift. Watched failing without the fix:
"GET /members must return its rows under \"members\" ==> expected: not <null>"
roster() is the supplier for LeadHeartbeatLoop and FleetHealthMonitor
(both timer-driven) and for placement/exhaustion checks and the
metrics scrape — none of which read agentSessionId. Resolving there
meant every tick could open a lazy-resolving adapter's (opencode's)
on-disk session database once per member whose id was still unknown,
with no bound: a member whose id never appears would pay that cost for
the life of the process.
roster() goes back to its pre-#209 behavior (no resolve, no I/O). A
new rosterResolved() carries the resolve logic, and is used only by
the two surfaces that actually report agentSessionId to a caller:
fleet_list (FleetMcp.listFleet) and the REST roster
(FleetApp.listMembers). fleet_whoami's roster().stream() at
FleetMcp.java:768 does not surface the field, so it stays on the plain
roster(). get(paneId) (fleet_status) and the release() resolve are
caller-driven, not timers, and are unchanged.
Retargeted the roster-facing tests from #209 at rosterResolved(), and
added plainRosterDoesNotResolveAgentSessionId, which pins the split by
asserting the handle's agentSessionId() is not called again by
roster().
When memberHerdrSocket: is configured, member panes run under a different OS
user than fleetd's own process, so hostEnvNames (fleetd's own environment)
no longer describes what a member pane inherits. logCredentialGap now checks
for that config key and, when set, logs a single WARN saying the gap is
UNKNOWN (not clean) and names the key, instead of printing the "inherits
them UNBLOCKED" / "the scrub blanks them" conclusions as fact. Behaviour is
byte-identical when memberHerdrSocket is absent (the default and only mode
this host runs).
SessionManager used to call PeerHandle.agentSessionId() exactly once at
spawn and freeze the answer into the immutable MemberSession. For
opencode that call always came back null, because opencode has not
written its on-disk session row yet when the pane is created, and no
caller ever re-asked the handle — it went out of scope at the end of
the spawn method. fleet_list therefore never reported agentSessionId
for an opencode member, and resumeSessionId was unusable for it.
Retain each spawn's PeerHandle in SessionManager, keyed by paneId, and
re-resolve a still-null agentSessionId against it from roster(), get(),
and release() (so a released member's detail also carries a
late-resolved id). Resolution is bounded: only sessions with a still-
null id do any work, a resolved id is never looked up again, and a
throwing handle degrades to "unresolved" rather than breaking the
caller. MemberSession gains a withAgentSessionId wither in the same
style as withState/withActivity.
Stage 3 of #185. A provisioned worktree and the repo's git store are made
group-writable when worktreeGroup names an OS group; absent, nothing changes.
The share pass runs after overlayParity, not inside add(), because
overlayParity copies more files in after add() returns.
This isolates credentials, not the repository: a member in the group can
still write the operator's git objects and refs.
opencode migrated its session store from a JSON file tree to SQLite in
January. OpenCodeSessionDiscovery still scanned the frozen tree, so it
returned null for every member: agentSessionId was never known and
resumeSessionId silently did nothing for every opencode profile, through
57 member spawns, with nothing logging that the search found nothing.
setReadOnly(true) is the whole thing keeping fleetd out of the operator's
live 841MB opencode.db, and no test failed when it was removed.
The obvious test does not work. Making the database file unwritable and
checking the read still succeeds passes either way, because SQLite silently
downgrades a read-write open of an unwritable file to read-only. I wrote that
test, watched it pass with the flag removed, and threw it away.
What works: extract a package-private openReadOnly(), then ask that connection
to INSERT and require the refusal. Watched red with the flag removed, green
with it restored.
Also switches the test's INSERT helper to a PreparedStatement -- hand-escaped
SQL in a test is a pattern that gets copied into main code.
Two defects in the stage-3 share pass, both of which would have failed EVERY
provisioning spawn once worktreeGroup was set, not only the two-user case.
.git/logs was handed to chgrp unguarded while packed-refs was guarded. It does
not exist with core.logAllRefUpdates=false, or before the first ref update, and
chgrp on a missing path exits non-zero -- surfacing as a WorktreeException that
blames a group which is in fact fine. Every path is now skipped when absent.
repoRoot + "/.git" was hardcoded. That is a FILE, not a directory, when the
checkout is itself a linked worktree -- the very thing this class creates for
every member. It now asks git: rev-parse --git-common-dir, resolved against
repoRoot because git answers relatively for an ordinary checkout.
Both new tests were watched failing with the fix removed before being kept.
Adds worktreeGroup (top-level FleetConfig key), Worktrees.shareWithGroup
(GitWorktrees impl: git config core.sharedRepository group + one-time
chgrp/chmod g+rwX/setgid fix-up over the worktree, .git/objects, refs,
logs, worktrees, and packed-refs when present), and wires SessionManager
to call it AFTER overlayParity so overlay files are covered too. Off by
default (byte-identical behaviour when unset). Documents the
credentials-not-repository caveat in the javadoc and example config.
opencode migrated its session store to SQLite in January 2026; the JSON tree under
storage/session/<projectID>/ses_*.json stopped being written, so OpenCodeSessionDiscovery
returned null for every member forever, and fleet_spawn{resumeSessionId} was unreachable.
- Add org.xerial:sqlite-jdbc 3.53.4.0, opened read-only (SQLiteConfig.setReadOnly), so it
never disturbs a live opencode process writing the WAL-mode database.
- Rewrite sessionIdForDirectory to run a parameterized SELECT ... WHERE directory = ?
ORDER BY time_updated DESC LIMIT 1 against the session table. Still never throws: a
missing database, a locked/corrupt one, or no matching row all return null.
- Log the silence that let this go unnoticed: WARN once per instance when opencode.db
itself is missing (the layout moved again), DEBUG when it exists but no row matches
yet (the normal interim answer right after a spawn).
- Replace the JSON-fixture tests with a synthetic-SQLite-db fixture; delete the tests
that only proved the old JSON scan worked.
A wait:false delegation whose worker used fleet_ask ended with fleet_poll{ticket}
reporting 'the worker session was released before it replied' -- naming a
worktree, a branch and a snapshot commit, so it read as lost work. The worker had
in fact replied in full.
The ticket guessed the ask rendezvous detached the turn. The real cause is
narrower: answer() (behind fleet_send{turnId}) waits only for the lead's own
bounded MCP call window. A resumed turn doing real work -- edits, a build, a
push, a PR -- routinely outlives it. On timeout answer()'s finally closed the
waiter, so the worker's later fleet_reply found none and fell to the session
inbox, leaving the ticket's future unresolved until fleet_stop forced it FAILED.
reply() now looks for the async task parked on this exact answered turn and
completes it with the real reply. That is safe against completing the wrong
ticket: answer() calls clearAsyncQuestion(turnId, false), so the task keeps its
turnId and stays in asyncTasksByTurn, and hasAsyncQuestion therefore still makes
send() return BUSY for a second async send to that target. At most one candidate
task can exist per target.
abandon() keeps an independent check: if a reply was stranded, a released session
reports REPLIED with that text rather than a failure -- so the recovery hint that
implies lost work never prints once a reply exists.
Verified before merging: build green unpiped, and both new tests drive the full
delegation path (send async, ask, answer, reply, poll the ticket) rather than
handing a Reply to a sink, which is the trap this ticket called out.
Co-authored-by: fleetd worker <worker@ltms.dev>
broker.uriEnv names an environment variable holding amqp://user:password@host,
and it was reaching every member. Its name is not credential-shaped -- no TOKEN,
KEY or SECRET in it -- so every name-pattern heuristic missed it, and it sat on
neither credential list.
fleetd already knows the name: the operator wrote it in broker.uriEnv. So derive
the exclusion from the config rather than hoping an operator also remembers to
deny it. coordinator.uriEnv has the same shape and is excluded too; on this host
both resolve to the same variable.
Excluded even when the operator lists the name under memberCredentials.allow:,
following the SSH_AUTH_SOCK precedent. There is no override, because a member has
no legitimate use for the broker password.
Reviewer finding, recorded rather than overstated: this is only a hard guarantee
under policy: allow-list, where the ZDOTDIR scrub runs after the pane's shell has
sourced the operator's chain. Under deny-list the name is removed from the
pre-shell env only, and a login shell re-exports it. That is deny-list's existing
weakness rather than a regression here, but the javadoc now says so plainly
instead of implying a guarantee that path cannot give.
Co-authored-by: fleetd worker <worker@ltms.dev>
PaneLocator's javadoc always claimed it found 'the agent pane whose process
tree contains' a pid. It did not: paneOwnsPid matched only the pane's shell_pid
and its foreground_processes. A process a member spawned -- python3, curl, any
helper opening its own connection to 127.0.0.1:8765 -- matched no pane, so
CallerResolver fell through to loopback-trust and resolved it as the PRIMARY.
A member escalated to lead by shelling out.
terminalForPid now builds the caller's ancestor set once (bounded at 32
generations, with a cycle guard) and matches any ancestor against a pane's pids.
The set is reused across both herdr clients on the CB-185 two-daemon path.
This only ever ADDS matches, which is the safe direction: the failure mode of
the fix is a member correctly restricted, while the failure mode of the bug is a
member acting as the lead. The no-match case still returns null, so the lead --
which maps to a pane named by leaders: -- still resolves as primary.
Ancestry is walked through a new ParentResolver seam so the tests drive it from
a fake pid->parent map rather than spawning real processes.
Co-authored-by: fleetd worker <worker@ltms.dev>
fleet_send{turnId} (MessageService.answer) blocks the primary only for its
own bounded MCP-call window (25s default, 120s max) — far shorter than a
worker's resumed turn can genuinely take. When that window expires, answer()
closes its rendezvous waiter, so the worker's eventual fleet_reply has no
live waiter to resolve and falls back to the session inbox. The async
ticket's future was never completed by that path, so fleet_poll{ticket}
stayed PENDING until fleet_stop's abandon() forced it FAILED with a
misleading "the worker session was released before it replied" reason,
even though the reply had genuinely arrived.
- MessageService.reply(): before falling to the inbox, look for the async
task this exact turn belongs to (already answered — question cleared,
turnId still stamped — but not yet resolved) and complete it directly
with the real reply, so fleet_poll{ticket} returns it.
- MessageService.abandon(): defense in depth, independent of the above —
never write a false WORKER_FAILED once a reply reached the inbox for
this target; recover and use its real content instead.
- Two new tests drive the full delegation path (async send -> ask ->
answer with a short timeout -> reply -> poll/abandon), not a reply sink
directly; both fail with the fix disabled and pass with it restored.
main already shipped the core of #164 in 3bfa828: the MIN_TURN_NANOS floor and
the hard fail on an empty or unreadable scrape. This adds the one case that was
still resolving as a success -- a scrape that reads cleanly but whose content is
the backend's own rejection (e.g. "API Error: 400 invalid request body").
The BACKEND_ERROR pattern is deliberately narrow. A growing list of ad-hoc error
strings rots as backends change their wording, and broader backend-error
surfacing is #164 point 3.
Because the pattern is a heuristic, it also matches a member that forgot
fleet_reply while reporting *about* a backend error. So the failure reason
carries the whole pane tail, not just the matched line: a genuine backend error
reads as before, and a false positive keeps its report instead of losing it.
Checked before merging: 3bfa828 is an ancestor of main; the branch was current
with main; mvn clean install green unpiped (1037 tests, 0 [ERROR] lines); and
each of the 5 new tests fails with the fix commented out.
Co-authored-by: fleetd worker <worker@ltms.dev>
Lead-authored and lead-verified: full mvn clean install green at 1032 tests, and the new test proved by restoring the old comparison and watching it fail.
pruneTerminalTickets compared the cutoff against createdNanos, so the real
window to collect a reply was "TTL minus however long the task ran". A
delegation that ran longer than the 10-minute TTL was already past the cutoff
the moment it finished, so the next prune destroyed its reply.
That is the normal case here, not an edge case. Real delegated work runs well
past ten minutes. Three workers in one session did, and two of their complete
reports were lost. The reply lives only in Task.future, so pruning it discards
the worker's whole report, and fleet_poll{target} returns [] rather than
holding it — there is no fallback.
Task now stamps completedNanos from a whenComplete hook registered in its
constructor, so every completion path stamps it (a reply, the completion
fallback, a timeout, a failure, an abandon on teardown) without each one having
to remember to. The stamp is a boxed Long, not a long with a sentinel:
System.nanoTime may return any value, so no number can mean "not stamped yet".
A task that is done but not yet stamped is left for the next sweep.
createdNanos had no other reader and is removed.
The TTL still bounds tasks — an uncollected finished ticket is still evicted
once the TTL passes since it finished. Both halves are pinned by a test, and
the first one was proved by restoring the old comparison and watching it fail.
Tests run: 1032, Failures: 0, Errors: 0, Skipped: 0
Lead-verified: merged onto current main (which already carries #194 and #196), full `mvn clean install` green at 1030 tests. Confirmed no plain `exec` call that reads a remote URL remains.
Review round 2 closed the half-fix: execRedacted was added but applied only to the new calls, leaving the three pre-existing URL readers (lines 143, 171, 334) still copying stdout into exception messages — the exact hole CB-189 named.
Kept the check before the origin strip, on the worker's reasoning: the credential really is in the config at that moment, so WARN-found followed by INFO-fixed is the full audit trail, whereas moving it after would silence the origin case entirely.
Lead-verified: merged with #194 onto an integration branch off main, full `mvn clean install` green at 1025 tests.
Review round 2 fixed the stale ambiguous-pane message and, more importantly, a real abort: probeOwner called list() unguarded, so one unreachable daemon made panes on a different healthy daemon un-stoppable too — the same bug blocker 1 exists to fix, through a new door. Worker proved it by removing the guard and quoting the failure.
Lead-verified: merged with #196 onto an integration branch off main, full `mvn clean install` green at 1025 tests. Deny-by-default WARN confirmed byte-identical to main.
Review round 2 fixed a defect I found in round 1: the new INFO asserted that gap names were not on the derived allow-list without ever checking, so a name derived from a profile's tokenEnv/gitTokenEnv/env: would be reported as safe while the member actually inherited it. Now split with MemberEnvAllowList.keeps — the same predicate the generated scrub evaluates.
Review found that execRedacted was applied only to the new remote-enumeration code
and left three pre-existing calls reading remote.origin.url through the plain,
unredacted exec: removeUserInfoFromHttpsOrigin, requireCredentialFreeHttpsOrigin, and
configureHttpsUrlRewriteForSshOrigin. A non-zero exit or timeout on any of those could
still have copied the credentialed URL into a WorktreeException message. Switches all
three to execRedacted; the set-url write in removeUserInfoFromHttpsOrigin is left on
plain exec with a comment explaining why (it writes the already-stripped URL, not a
read).
Widens the shared exec(Map, boolean, String...) overload to package-private, the same
test-seam pattern already used by the afterWorktreeAdded constructor parameter, and
adds a test that drives it directly with a synthetic failing command whose stdout
carries a marker (passed via env, not argv, so the always-printed command line can't
carry it) and asserts the marker never reaches the exception message.
Lead review of PR #196 found two issues in CompositePeerLauncher.probeOwner:
1. The "more than one daemon claims this pane" throw kept the old
pre-fix message ("no owning herdr daemon was recorded"), which was
only true of the code it replaced. Reworded to say what actually
happened: N configured herdr daemons report this pane, so it is
genuinely ambiguous. Updated the one test pinning the old string.
2. probeOwner let list() propagate straight out of the probe loop, so
one unreachable daemon aborted the whole probe and made a pane on a
DIFFERENT, healthy daemon un-stoppable too — resurrecting the exact
bug blocker 1 fixes. Now catches HerdrException per daemon, logs the
exception class only, and treats that daemon as not knowing the pane
so probing continues. New test proves this: verified it fails with
the try/catch removed (HerdrException propagates and the stop that
should succeed via the healthy daemon throws instead), then restored.
Full mvn clean install: 1019 tests, 0 failures, 0 errors.
Lead review on PR #194 found that the allow-list INFO wording claimed the
whole gap ("credential-shaped names on neither known: nor allow:") is blanked
by the scrub, without checking that against effectiveAllowed. effectiveAllowed
is a SUPERSET of known+allow — MemberEnvAllowList.derive also unions in every
profile's gitTokenEnv/gitHostEnv/tokenEnv/env: keys, and derivedAllowedNames
further unions in the spawn's own env keys — so a gap name can still be kept
by the derived list (e.g. a profile's tokenEnv names it) and reach the member
unblocked while the INFO said "no member pane keeps them". That inversion is
exactly what #192 exists to remove.
logCredentialGap now splits the gap with MemberEnvAllowList.keeps (the same
predicate the generated scrub itself evaluates, so this cannot drift from
what the scrub does): names it keeps get a WARN, guarded by the same
unprotectedGapLogged flag as the deny-by-default case (same severity — a name
reaching a member unprotected is equally serious either way); names it
blanks keep the existing INFO, guarded by allowListGapLogged. The
deny-by-default WARN text and the non-zsh fallback are untouched.
Added allowListWarnsWhenTheDerivedAllowListKeepsAnUncoveredName and
allowListSplitsAMixedGapBetweenTheWarnAndTheInfo to ClaudeCodeLauncherTest.
1. CompositePeerLauncher.stop() was permanently un-stoppable for any
member that survived a daemon restart, because spawnedBy is in-memory
only. On a cache miss with more than one configured herdr daemon, probe
each distinct daemon's agent.list() for the pane instead of refusing
outright: exactly one owner routes and caches; zero owners is treated
as already-stopped (a no-op, matching the tolerance HerdrPeerLauncher
already gives an already-gone pane); more than one owner is the
genuine per-daemon-pane-id ambiguity and still throws.
2. FleetApp#healthz always reported the LEAD daemon's herdr version/
protocol even when a second (member) daemon was configured, so a
member-daemon protocol mismatch was invisible behind a green
/healthz while every spawn silently failed. Added a separate "member"
key alongside the unchanged "herdr" key, and a "protocolMismatch"
flag when the two differ. Verified scripts/redeploy-fleetd.sh and
scripts/rename-checkout.sh only check the HTTP status code and print
the body verbatim — neither parses a specific field — so adding a key
is safe.
Both fixes are covered by tests written to fail without the fix
(verified by reverting each fix and watching the new tests fail, then
restoring). Full `mvn clean install`: 1018 tests, 0 failures, 0 errors.
GitWorktrees only ever inspected origin's HTTPS fetch URL for embedded credentials. A
credential on any other remote, on a pushurl, or on a plain http:// URL passed through
unreported. Adds an additive, reporting-only check that enumerates every remote and both
its fetch and push URLs, flagging non-empty user-info on any non-SSH-family scheme.
The existing origin/https strip-and-refuse behaviour is untouched. The new check is
wrapped so it can never abort a provision, and on failure logs only the exception's
class, never its message, since the enumerating `git remote` call is not redacted.
Also adds execRedacted, an exec variant that never copies captured stdout into a
WorktreeException message, for commands whose stdout may itself be a credentialed URL.
The memberHerdrSocket section exists at 11-Features.md:2174; what was absent was
its row in the index table. My omission when I added the section. Wiki fixed at
b24965c.
logCredentialGap(creds) always emitted the WARN wording ("every member pane
inherits them UNBLOCKED"), even under memberCredentials.policy: allow-list on
a zsh login shell, where the generated ZDOTDIR scrub genuinely blanks the
name. The line reported the control working as though it were a hole.
Pass an effectiveAllowed set instead: null keeps the WARN (deny-by-default,
and the allow-list non-zsh fallback, where nothing is ever scrubbed); the
derived allow-list set (only reachable after applyEnvironmentAllowListPolicy's
own zsh gate) selects a new INFO wording that says the scrub will blank the
name instead of claiming it is inherited unblocked.
Also split the single credentialGapLogged AtomicBoolean into two guards
(unprotectedGapLogged / allowListGapLogged) — one per report kind. Since
memberCredentials is a live, re-read-per-spawn supplier, a shared flag let a
harmless allow-list INFO on one spawn permanently suppress a later spawn's
real deny-by-default WARN after a policy reload.
Fixes gitea #192.
memberHerdrSocket splits lead operations from member operations onto two herdr
daemons. Three seams still assumed one shared daemon and broke silently when the
two clients differ (all three collapse to today's behaviour when they are the
same object):
1. ConnectionIdentity's PaneLocator was pinned to the member daemon only, so a
lead's own MCP connection (which lives on the LEAD daemon) resolved to
terminal == null, breaking fleet_reply/fleet_ask/fleet_whoami for a lead.
PaneLocator now searches the lead client first, then the member client.
2. StatusPoller's StatusRefiner was pinned to the member daemon, so refining an
UNKNOWN status for a lead target read the wrong daemon's pane content and
never left UNKNOWN, wedging status-gated delivery to that lead forever.
StatusRefiner gained a refine(target, raw, control) overload and the poller
now refines through the same AgentControl the raw status was sampled from.
3. FleetApp was constructed with the raw lead-only herdr client, so /healthz
stayed green while the member daemon was down (every spawn then fails
invisibly) and GET /sessions silently dropped every member workspace.
FleetApp now takes both clients: healthz requires both to answer, sessions
merges workspaces from both.
Each fix has a test proven to fail without it (verified by reverting the
production change and re-running): FleetdConnectionIdentityConstructionTest /
FleetdFleetAppConstructionTest assert the actual Fleetd.java wiring (the same
technique as FleetdHerdrControlConstructionTest); StatusPollerRoutingTest and
the new PaneLocatorTest/FleetAppTwoDaemonTest cases exercise the real
production classes end to end rather than a hand-built object graph.
Three fixes on top of the pane-id PR, from my own read and the reviewer's:
- stop() removed the spawnedBy record BEFORE the delegate accepted the stop. A
delegate that threw left the pane alive with its owner forgotten, so the retry
fell into the ambiguous branch and refused the id for good. Remove after.
- list() deduplicated on the raw pane id. Pane ids are per-daemon counters, so
two daemons can each hold w1:p1 on different panes, and one of the two real
agents was silently dropped from fleet_list and every view built on it. The
key is now (owning daemon, pane id). Delegates sharing one daemon still
collapse, which is what the dedupe was for.
- The class javadoc still stated the single-herdr-connection premise as fact,
next to the bullet this PR had just corrected for stop(). Fixed there too.
Also drops a redundantly qualified java.util.Collections.
Tests: 993 run, 0 failures, BUILD SUCCESS.
The skill said SSH to fleet01 is denied, so every report wrote 'not
reachable' for that fleet's daemon facts. That is true only for the user
dai.ha. The host alias fleet01 maps to user ltms and key auth works.
Checked 2026-08-28 while measuring #185: ssh fleet01 connects, and ltms
has passwordless sudo there. So fleet01's PID, uptime, jar and /healthz
can be reported over SSH even though its REST port is unreachable.
The javadoc said a member cannot authenticate at all once memberCredentials
blocks SSH_AUTH_SOCK, "there is no private key file on this host, only an
ssh-agent socket". That is wrong, and it was written after looking only in
~/.ssh, which holds nothing but Include lines.
Measured: ssh -G git.ltms.dev resolves an IdentityFile under the shared-env
directory. That file exists, is readable by this user, and has no passphrase.
A live member with SSH_AUTH_SOCK blanked pushed to the forge over SSH.
The rewrite itself is unchanged and still worth having. Only its stated reason
was wrong: it routes a member through its own scoped token instead of the
operator's ssh identity, which is what makes a member's pushes attributable
and revocable. It is not what stands between a member and the forge.
The four log lines added with the worktree HTTPS rewrite echoed the origin
URL verbatim, and one of them echoed the ssh:// authority, which carries
user-info. An ssh authority is normally just git@, so in practice this
changes nothing -- but a remote URL is not obviously a credential channel,
and that is precisely why one has leaked here three times (#157, #182).
Redact at the log call, not after it surprises someone.
Git never consults a credential.helper for an SSH transport, so #177's helper was inert on this repo — whose origin is ssh://. Once allow-list policy blocks SSH_AUTH_SOCK, a member on an SSH origin cannot authenticate at all: there is no private key file on this host, only an agent socket.
A worktree-scoped `url.<https>.insteadOf <ssh>` gives the member HTTPS for fetch and push while the primary checkout keeps SSH untouched. Host and port are parsed from the origin, never hardcoded — a test with a synthetic host proves it. The scp-like shorthand is left alone deliberately, since its host:path split is defined by ssh_config aliases rather than URI syntax.
Verified by the lead in an independent worktree: Tests run: 986, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
The member credential scrub removes environment variables. It cannot remove a token written into git config inside the repo the member works in, so `git remote -v` handed a member a credential it was deliberately not given.
Provisioning now strips HTTPS user info from the origin before `git worktree add`, refuses the worktree if user info survives, and configures a per-worktree credential helper that reads WORKER_GITEA_TOKEN at call time. Nothing is persisted.
The helper emits BOTH username and password, and resets the inherited helper list first. An earlier revision emitted only `username=`, which made git fall through to the next helper — on a Mac that is osxkeychain, so a member would have authenticated with the operator's stored credential while every test passed and `git remote -v` looked clean. See #182.
`worktreeCredentialHelperCompletesWithoutUsingAnInheritedHelper` plants a synthetic operator helper in an isolated global config and proves the worktree helper wins. The worker confirmed it fails when the reset is removed. All credential tests pin GIT_CONFIG_GLOBAL and GIT_CONFIG_SYSTEM so they can neither read nor write real credentials.
Verified by the lead in an independent worktree: Tests run: 973, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
A turn that dies on a backend error produces the same working -> idle transition as a real one, just faster and with nothing on screen. The resolver accepted that as a completed turn and handed the caller HTTP 200 with an empty reply, so a lost turn and a successful empty answer were indistinguishable.
Now: an empty or unreadable scrape fails, naming the member; and a BUSY -> DONE inside MIN_TURN_NANOS (2s) fails as a crash signature.
One existing test encoded the bug — it asserted a failed scrape resolved as a success carrying "" — and has been inverted rather than worked around.
Verified by the lead in an independent worktree: Tests run: 973, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
`policy: allow-list` silently ignored every name an operator wrote under `allow:` unless a profile happened to carry it too, so turning the policy on would have blanked credentials working members depend on. Derivation now unions the operator's list.
`SSH_AUTH_SOCK` stays governed only by `sshAuthSock`, even when listed under `allow:` — it is a live handle to the operator's ssh-agent, not a value.
Adds one INFO line per allow-list spawn, `member credentials: allowed N of M`, emitted only after the shell gate so it can never report coverage on a path where the scrub does not run.
Verified by the lead in an independent worktree: Tests run: 976, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
No behaviour change. #154 supposed that ownership drains a whole queue into the in-memory `held` map, making `x-max-length` and per-message TTL decorative. Measurement says otherwise: `basicConsume` is manual-ack, `deliverCallback` acks only duplicates, and `basicQos` is set on the one shared channel before any consumer starts — so total `held` is bounded by the prefetch window across all targets.
Adds a fake-broker test that drives the real `own()` path and fails if receipt ever starts acking, plus a javadoc line naming the prefetch window at the point of first mention.
Verified by the lead in an independent worktree: Tests run: 971, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
CompletionResolver.resolve() used to hand the caller a successful "" reply whenever a
turn's scrape came back empty (whether the read failed, or genuinely produced nothing),
making a lost turn indistinguishable from a real empty answer. It also had no way to
tell a crashed backend's near-instant BUSY -> DONE transition apart from a genuine
completion.
Add MIN_TURN_NANOS (2s), a named floor below which a completed turn is treated as a
crash signature and failed rather than resolved as a reply. Fail on any empty scrape
(read failure or a clean-but-empty read) instead of resolving with "". Both failures
name the member and carry whatever is on the pane for context.
Thread an injectable LongSupplier clock through CompletionResolver (matching the
SessionManager/MessageService nowNanos pattern) so the floor is testable without a
real sleep.
The coverage line was logged before the zsh gate, so a non-zsh
spawn (where nothing is scrubbed — overlayBlockedCredentials is the
fallback instead) printed 'allowed N of M' as if the derived
allow-list scrub had run. Move the log after the gate so it only
fires on the path that actually generates the ZDOTDIR scrub; the
non-zsh fallback keeps logCredentialGap's WARN as its only signal.
Added a test proving no 'allowed N of M' line is emitted on the
non-zsh fallback, through the real HerdrPeerLauncher#spawn path.
Share one deliverability predicate between Injector and FleetApp, so the status endpoint reports the same answer the injector acts on instead of re-deriving it from one of that predicate's two inputs.
Verified by the lead in an independent worktree: Tests run: 971, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
MemberEnvAllowList.derive only ever looked at profile fields, so
memberCredentials.allow: was silently ignored under
policy: allow-list — turning the policy on would have blanked
credentials working members already depended on.
- derive(profiles, configuredAllow) unions memberCredentials.allow
into the derived set, with SSH_AUTH_SOCK explicitly excluded from
that union (it stays governed only by sshAuthSock: allow).
- HerdrPeerLauncher threads MemberCredentials.allowSet() into the
derivation instead of calling the profiles-only overload.
- Added a per-spawn INFO log 'member credentials: allowed N of M'
(N/M from the daemon's own env, the existing hostEnvNames proxy),
never logging a blocked name or a value.
CB-640 published the three message-layer facts and CB-641 wired the herdr
and time ones. This joins them, so every HealthSnapshot field now carries
real evidence and the NOT_YET_OBSERVED placeholder is gone. That constant
is what made 8 of the 9 fault states unreachable, GONE and NEVER_READY
included, which is why CB-580's failTarget never fired.
hasOrphanedDelegation is a true snapshot, but it can read true for one
tick during an ordinary race: an async ticket exists before its virtual
thread reaches rendezvous.open, so for that instant nothing is accepted or
queued behind it. decide maps the field straight to DELEGATION_ORPHANED
with no smoothing, so one racy read would log a fault that clears on the
next tick. The monitor now requires two consecutive observations. That
costs one interval on a real orphan and removes the false positive.
Two tests drive real ticks against a genuinely orphaned ticket (an
unanswered fleet_ask that lapsed back to PENDING), not the seam: one tick
reports nothing, two report once, and a single clean tick in between
resets the streak.
Also correct two config comments. paneProbeIntervalSeconds is parsed and
read by nothing, so its "minimum 60" note promised a floor that does not
exist.
970 tests green.
hasQueuedDelivery/hasStrandedReply/hasOrphanedDelegation surface three of the
message-layer facts FleetHealthMonitor needs but currently hardcodes to
NOT_YET_OBSERVED. Additive only — no existing public method's signature or
behavior changes.
scrubWritesAnAllowedNofMReport spawned /bin/zsh with no guard, so it
errored on the Gitea CI runner (Linux ARM64 container, no zsh) while
the two sibling zsh tests already skipped there via assumeTrue. Add the
same guard so CI skips instead of failing; the Mac build still runs it.
The alias removal left bridge_* tool names in prose. Fix them:
- README no longer claims the old bridge_* names still answer (they were removed).
- pom + LeadTabScanner comments name fleet_* tools.
- FleetMcp comment no longer mentions the removed deprecated twin.
- docs/MCP-Contract.md and e2e swept bridge_* -> fleet_*; e2e ask files renamed.
The historical mcp__bridge__* mount-name note in CLAUDE.md is kept on purpose.
949 tests pass.
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
- CB-634: deliver IDE guidance as an on-disk CLAUDE.local.md / opencode overlay,
pinned to the module dir, best-effort auto-open (opt-in per profile via ideMcpUrl).
- CB-636: per-profile autoCompactWindow -> --autocompact for claude-code,
provider.<p>.models.<m>.limit.context for opencode. Range-validated [100000,1000000].
- CB-637: cross-host lead-to-lead over a shared AMQP coordination vhost
(coordinator: block, fleet_send{coordId}, LeadMailbox + LeadCoordLoop).
952 unit tests + 5 LeadMailbox contract tests green. Live E2E verified on the Mac daemon.
The lead-to-lead wiring shipped with CB-635 in its comments, but CB-635 is
already the broker.uriEnv / unreachable-broker work. Relabel the mailbox +
fleet_send{coordId} + receive loop to CB-637 so a ticket number names one
feature. Add the cross-host peer-lead row to the primary intent->tool table
(kept byte-identical with the wiki template).
The sibling ticket landed the mechanism (LeadMailbox, LeadMessage, the
coordinator: config block) but nothing opened it, nothing sent through it, and
nothing read it. This is the wiring.
- LeadChannel: a small interface LeadMailbox now implements (publish/peek/ack
plus a selfCoordId() accessor). It exists so FleetMcp and the receive loop can
be tested with a fake instead of a live broker. LeadMailbox's AMQP logic is
untouched — the diff is the implements clause, four @Override marks and the
accessor.
- Fleetd.openLeadMailbox: opens this daemon's mailbox after the reply inbox is
selected, with the same env-injected seam selectReplyInbox uses. Every "off"
path returns null and the daemon still starts: no coordinator block (silent),
a uriEnv that does not resolve (INFO), a configured broker with no selfId
(WARN — a mailbox is named after the coord-id that owns it), or a broker that
refuses at boot (WARN, credentials stripped). Closed in the ordered shutdown
hook, after the loop that reads it has stopped.
- fleet_send{coordId}: publishes a LeadMessage(from=selfCoordId, to=coordId) to
the peer's mailbox and returns the broker-confirmed receipt. coordId is
mutually exclusive with sessionId/turnId and is rejected by name rather than
resolved by precedence. An unroutable/nacked/timed-out publish comes back as a
tool error naming the coordId, never a crash. The worker send/reply path is
not touched.
- LeadCoordLoop: the receive half. Each tick peeks the mailbox, resolves the
local lead pane, and — only at a turn boundary — injects "[lead <from>] <text>"
and acks. Anything not delivered stays unacked and is retried, so a message is
never dropped; one message per tick, so every delivery is gated on a status
read that already saw the previous one.
- fleet_list reports {selfId, configured} when coordination is on, so an
operator can find the coord-id a peer must use to reach them. Omitted
entirely when it is off.
Tests: 20 new hermetic tests (no broker) across routing, delivery and startup
selection. mvn clean install: Tests run: 944, Failures: 0, Errors: 0, Skipped: 0
— BUILD SUCCESS.
Adds the broker-side mechanism for lead-to-lead messages across daemons/hosts
(unit 1 of 2): a LeadMessage envelope carrying from/to coord-ids, an
AMQP-backed LeadMailbox modeled closely on AmqpReplyInbox (consume-and-hold,
deferred manual ack, confirm-mode publish, recovery handling), and a new
optional coordinator: config block (separate vhost from broker:, leader
traffic only). Config parsing + accessors only — FleetMcp/Fleetd/Injector/
MessageService and the send path are untouched; wiring is a separate ticket.
Add opt-in Integer autoCompactWindow to FleetConfig.Profile (last field,
null/unset = today's behaviour). Validated at config load to [100000,
1000000] — the band Claude Code's own --autocompact flag accepts.
Claude Code: appends --autocompact <window> to argv (mirrors --model),
so it survives the ccs <profile> wrapper.
opencode: has no absolute compact-at-N knob (only compaction.auto/prune/
reserved/tail_turns/preserve_recent_tokens), so the window is applied as
the resolved model's own limit.context (+ a required limit.output:16384
default) in the generated opencode.json, merged via get-or-create nodes
so it does not clobber a custom-provider block. Only applies when model:
resolves to "provider/model"; otherwise logs a WARN naming the profile
rather than silently doing nothing.
Docs added to fleetd.example.yaml explaining the cross-backend semantics
difference (compacts AT the window vs. WITHIN it).
The overlay pinned project_path to the worktree root. For a repo whose Maven
module is a subdir (this repo's pom is in `bridged/`, not at the root), opening
the root imports no module and every ide_* call resolves nothing. Pin and open
the module dir instead.
Two new opt-in per-Profile keys, both read only when ideMcpUrl is set:
- ideProjectDir: repo-relative module dir the IDE opens and the overlay pins;
blank keeps the old worktree-root behaviour.
- ideOpenCommand: host command that opens that dir in the IDE at spawn, with
{dir} substituted and run through /bin/sh -c so env (e.g. DISPLAY) can be set
inline. Best-effort and non-fatal — a failure never fails the spawn. Blank
keeps the manual-open behaviour. No close half yet (deferred).
Shared helpers PeerLauncher.ideProjectPath / openInIde back both launchers.
The two Profile fields ride a back-compat constructor, so every existing call
site and YAML compiles and behaves unchanged.
Tests: overlay content pins the module dir when ideProjectDir is set;
ideProjectPath resolution; openInIde no-op on a blank command. 918 tests green.
git reads info/exclude from the common dir for a linked worktree (only
info/sparse-checkout is per-worktree), so the entry written into
<common>/worktrees/<name>/info/exclude was never honoured and CLAUDE.local.md
showed as untracked -- at risk of being swept into a worker's PR. Derive the
common dir (<common>/worktrees/<name> -> <common>) and write there. Found by
dogfooding a real spawn on fleet01; the test now uses the real worktree layout
and asserts the entry lands in the common dir, not the per-worktree gitdir.
Move the IDE guidance text to PeerLauncher.ideOverlayText (shared by both
launchers). ClaudeCodeLauncher drops it from the reply-charter file and writes
CLAUDE.local.md into a provisioned worktree instead, gated on a .git FILE
(safety: never writes into the primary's real .git-DIRECTORY checkout) and
registers it in info/exclude. OpenCodeLauncher mounts the intellij server and
adds the rules file to the instructions array.
Adds `ideMcpUrl` to FleetConfig.Profile (default off). When set, the
Claude Code launcher mounts the IDE Index MCP as a second inline
--mcp-config server named `intellij`, and appends an IDE charter that
pins every ide_* call to the member's own worktree (spec.cwd()). The
charter order is role -> ide -> reply, one --append-system-prompt-file,
reply last (CB-618). The mount gate now fires on ideMcpUrl alone, not
only mcpUrl. ConfigRef treats an ideMcpUrl change as deferred, like the
other launch flags.
Never touches .mcp.json or CLAUDE.md — the mount and the rule arrive as
launch flags, so a project's own config is untouched.
Not yet done (see fleetd #162): the bridged-owned IDE lifecycle
(open on provision, close before worktree removal), and the opencode
adapter (separate ticket). fleetd.example.yaml documents ideMcpUrl and
fixes the stale parityOverlay default.
911 tests green.
An empty uriEnv no longer stops the daemon (#152), so the failure is quiet: bridged
starts, falls back to the in-memory reply inbox, and held reports stop surviving a
restart. --check is the only thing that says so before the fact. The var name is read
out of bridged.yaml so a renamed key cannot make the check lie.
#151: Broker gains uriEnv beside uri, taking the AMQP URI from an env var so
the password stays out of fleetd.yaml (same pattern as auth.tokenEnv). uriEnv
wins when set; isConfigured() treats a uriEnv naming an unset/blank variable
as unconfigured. A configured uriEnv is added to the startup required-secrets
report. Never logs the resolved URI (it carries the password).
#152: AmqpReplyInbox.open throwing at boot no longer stops the daemon. The
selection at the call site catches the failure and falls back to the in-memory
inbox for the process lifetime, warning loudly that durable cross-restart
delivery is off and logging the failed URI with credentials stripped.
Round 2 fixes the defect that mattered: the scrub lived in .zlogin only, and a
herdr pane is a login shell on macOS but a plain interactive one on Linux. It
would have protected nothing on the vhost it was built for, in silence.
Verified by the lead: 901 tests, 0 failures, mvn clean install green. Both new
tests mutation-checked -- unwiring the control fails one, reverting the scrub to
.zlogin alone fails the other while the login-shell test still passes.
NOT yet deployed: the running daemon still holds the old jar.
Six review findings, from the peer lead `vms` and a reviewer worker. The first
one is a real defect that would have shipped as a dead control.
1. The scrub only ran in a login shell. It lived in the generated `.zlogin`,
and zsh reads `.zlogin` only for a login shell. herdr does not open the same
kind of shell everywhere: measured on herdr 0.8.0, a macOS pane runs `-zsh`
(login) while a Linux pane runs a plain `/usr/bin/zsh`. So on the vhost this
was being built for, `.zlogin` never ran and every member kept the whole
secret store, in silence.
The scrub body now lives in a generated `scrub.zsh` that BOTH `.zshrc` and
`.zlogin` source, each after sourcing its own `$HOME` counterpart. Linux
runs the first, macOS runs both, and the second pass is not merely harmless
-- it re-scrubs anything the operator's `~/.zlogin` exported after `~/.zshrc`
had finished. Re-running is idempotent.
2. `INFRASTRUCTURE_PASSTHROUGH` listed names that are not infrastructure:
ANTHROPIC_AUTH_TOKEN, GITEA_TOKEN, GITEA_HOST, ANTHROPIC_BASE_URL,
ANTHROPIC_MODEL, CLAUDE_CONFIG_DIR, OPENCODE_CONFIG, BRIDGED_MEMBER. I read
each injection point and confirmed every one of them reaches `launch.env()`
only when actually injected, so `allowed.addAll(launch.env().keySet())`
already covers the legitimate case. As static entries they were pure leak
surface: a host that happened to export ANTHROPIC_AUTH_TOKEN would have had
it passed straight through.
3. A missing `scrub-report.txt` at teardown was logged at debug. The report is
the only evidence the scrub ran at all. Its absence has an innocent reading
and a serious one, and we cannot tell them apart from the daemon -- so it is
now a WARN that says exactly that. Logging it at debug is how a control that
quietly stopped working stays unnoticed.
4. Nothing tested that the control was wired in. Deleting the single
`applyEnvironmentAllowListPolicy(cfg, launch)` line left all 896 tests green
while turning the feature completely off -- the CB-586/CB-611 shape again.
`HerdrPeerLauncherAllowListWiringTest` starts a real spawn and asserts on the
env that reached herdr. Mutation-checked: unwiring that line fails it.
5. `EnvAllowListScrubTest` now also runs `zsh -i` with no `-l`, which is the
Linux pane shape, so finding 1 is tested from a Mac. Mutation-checked:
putting the scrub back in `.zlogin` alone fails that test alone, while the
login-shell test still passes -- which is exactly the blind spot that let
the bug through.
6. Two ZDOTDIR leaks closed. A failed spawn has no pane id, so its directory
was never keyed for teardown; it is now removed on the way out. And
`deleteOnExit` covers a clean shutdown and nothing else, so `generate` now
reaps sibling directories older than 24h left by a killed daemon.
Also: `policy:` is lowercased with Locale.ROOT, and the `.zlogin`-only claim is
corrected in fleetd.example.yaml, FleetConfig and HerdrPeerLauncher.
901 tests, 0 failures, `mvn clean install` green.
Move member environment control out of the pane-creation env overlay
(defeated by any file the login shell sources) into a per-spawn ZDOTDIR
directory whose .zlogin runs LAST, after the operator's whole chain, and
blanks every exported variable not on an allow-list DERIVED from what the
launcher itself injects (profiles' tokenEnv/gitTokenEnv/gitHostEnv/env
keys + an infrastructure set) — never hand-typed.
- memberCredentials.policy: allow-list (deny-by-default/deny-list stay
default and unchanged); known:/allow: become reporting only under it.
- memberCredentials.sshAuthSock knob, blocked by default; allowing it is
an explicit decision (operator ssh-agent handle).
- Non-zsh login shell: loud WARN, protection off, fallback to the old
enumerated-name overlay.
- Scrub writes an 'allowed N of M' denominator report, read at teardown;
credential-shaped blanked names go to WARN (names only, never values).
- Equality test against a real login zsh from a clean parent: surviving
non-empty exports EQUAL baseline ∩ derived allow-list.
Every launcher writes the bridge's MCP server into the config it hands its peer,
and it named that server "bridge". So a member addressed its tools as
mcp__bridge__fleet_send while the tools themselves are already fleet_*. The
mount is named "fleet" now, and a member's tools are mcp__fleet__*.
The name was a bare literal in three files: ClaudeCodeLauncher and LeadLauncher
build a --mcp-config JSON string, OpenCodeLauncher writes an opencode.json node.
Three hand-written copies of one name is how a rename lands in two of them, so
the name is now one constant, PeerLauncher.MCP_MOUNT_NAME.
The mount name is local to the peer — it is the label its own client puts on the
server, and nothing in the daemon reads it back. Renaming it changes no wire
call.
Tests. Each launcher's test now asserts the mount is named fleet AND that
nothing writes "bridge"; the second half is the part that would have caught a
half-done rename. LeadLauncherTest never checked the name at all, only the URL,
so it gained the assertion rather than had one updated.
CLAUDE.md's role-detection ladder quoted mcp__bridge__* as the marker of a
spawned member. It names mcp__fleet__* now, and says that a member spawned
before this change still reports the old prefix. The portable block stays
byte-identical with the wiki template (wiki 569a917).
Build: cd bridged && mvn clean install, then read target/surefire-reports/*.xml
directly — 884 tests, 0 failures, 0 errors.
Unit 3 renamed bridged.example.yaml to fleetd.example.yaml and taught Fleetd to
read fleetd.yaml first. Two things it left behind.
bridged/.gitignore still ignored only bridged.yaml. An operator who follows the
new comment and copies the example to fleetd.yaml gets an untracked live config
holding tokens, and git offers to commit it. Both names are ignored now, because
both names work until the cutover.
Three design docs still pointed readers at bridged.example.yaml, a file that no
longer exists under that name.
Part of #145 (CB-632). Doing this NOW, ahead of the rest of the path
renames, for one reason: the vms lead is about to wire a monitoring
dashboard to these series. Renaming a metric after a dashboard points at
it breaks continuity and silently leaves a dead panel. Renaming it before
costs nothing, so it goes first rather than at the cutover.
Nine series renamed, all declared in FleetMetrics.
Two real defects found while doing it:
- FleetApp had "bridged_auth_failures_total" written as a LITERAL
instead of using FleetMetrics.AUTH_FAILURES -- a second hand-written
copy of a name, which is how these drift. It now uses the constant,
so there is one source for that name again.
- Nothing guarded the prefix. One test does assert a wire name
(FleetAppAuthTest checks the real /metrics body for fleet_sessions),
which is good, but it covers one series out of nine. The literal
above was covered by nothing at all.
So this adds MetricNamesTest, which reads the constants reflectively
rather than listing them -- a test that lists the nine names is itself a
second hand-written copy, and would pass while a tenth went unchecked.
It asserts its own denominator too: "no name starts with bridged_" is
true of an empty set, so a sweep that found nothing would pass loudly.
Asserting the count of 9 makes a broken sweep fail instead.
Also renamed three herdr contract-test workspace labels, __bridged_* ->
__fleet_*. Those are throwaway workspaces created by ensureWorkspace, not
metrics, but they are the same word.
Verified: mvn clean install green, 52 classes, 881 tests, 0 failures.
Test count is up by 3 -- the new prefix, denominator and uniqueness
checks. No bridged_ string remains anywhere outside wiki/.
Part of #145 (CB-632). Documentation only, plus one internal literal.
Unit 1 renamed the package and classes, which left every doc describing
classes that no longer exist. This fixes the prose across README.md,
docs/ and bridged/docs/ -- 18 files.
Renamed: dev.ltms.bridged -> dev.ltms.fleet, the five class names, and
"bridged" where it names the daemon as a product rather than a path.
Also renamed two literals, because a doc that disagrees with the code is
worse than one that is out of date:
- bridged-local-noauth -> fleetd-local-noauth. A placeholder apiKey
OpenCodeLauncher sends when a profile resolves no token, to a local
endpoint that does not check it. No test asserts the old string.
- the vnd.ltms.bridged.* media type in the M4 design doc. It appears
in no Java file, so nothing implements it yet.
Deliberately NOT renamed, because each is still literally true today and
changes only at the cutover:
- paths: bridged/, bridged.yaml, bridged.example.yaml, bridged.jar,
.bridged-worktrees, deploy/dev.ltms.bridged.plist,
scripts/redeploy-bridged.sh, bridged-launchd-wrapper.sh
- bridged_* metric names -- renaming these after the monitoring is
wired would break dashboard continuity, so they move before it is
- bridge_* MCP tool names, which answer alongside fleet_* on purpose
- BRIDGED_* environment variables, read by a file outside this repo
Method note: perl, not sed. BSD sed has no \b and no lookaround, and a
word-boundary expression there fails silently. The prose replace uses
(?<![\w./-])bridged(?![\w./-]) so it cannot touch a path or an
identifier, then every remaining hit was read by hand.
Verified: mvn clean install green, 51 classes, 878 tests, 0 failures.
mcp/BridgeMcp is now mcp/FleetMcp. This line is in the project addendum,
not the canonical block, so the wiki template is untouched -- the sync
check still returns True.
Part of #145.
Part of #145 (CB-632), under epic #125.
The product is called fleet and the daemon is called fleetd, but the code
still said bridge everywhere. This renames the Java half:
package dev.ltms.bridged -> dev.ltms.fleet
Bridged -> Fleetd (the main class)
BridgedConfig -> FleetConfig
BridgeMcp -> FleetMcp
BridgedApp -> FleetApp
BridgedMetrics -> FleetMetrics
The package root is dev.ltms.fleet, not dev.ltms.fleetd. The trailing d
means daemon, which names a process, not a namespace.
What this commit deliberately does NOT change:
- The module directory stays bridged/, and <finalName> stays bridged.
The installed launchd plist names bridged/target/bridged.jar and its
KeepAlive is armed, so renaming the jar on its own strands a restart.
Both change at the cutover, together with the plist, in one step.
- The bridge_* MCP tool aliases. CB-622 shipped both names on purpose.
One test names a local variable viaBridge because it holds the result
of the deprecated call; the rename collided with it and the compiler
caught it. That variable is back.
- BRIDGED_* env var names, and bridged.yaml. Both are operator
contracts and need a read-both shim, which is a later unit.
Two things a plain search-and-replace would have missed:
- logback.xml and logback-test.xml name the package twice, once as a
turboFilter class= attribute. The compiler never checks those.
- BSD sed does not support \b. The word-boundary expression matched
nothing and said nothing, while the other ten in the same command
worked. Checked the leftovers instead of trusting the exit code.
Verified: mvn clean install green, 51 test classes, 878 tests, 0 failures
-- the same count as before the rename.
Adds the rename script for CB-624 (#128), with a read-only --check mode.
Two workers ran this ticket in parallel and were kept apart on purpose: one wrote the script, one inventoried the same rename read-only without seeing it. The inventory found three surfaces the script had missed, all verified on the live machine before being acted on: ~/.claude.json's projects key (trust state, enabledMcpjsonServers, allowedTools, lastSessionId), ~/.config/herdr/session.json, and the JetBrains recentProjects.xml / trusted-paths.xml pair in 2026.1 and 2026.2.
All three are report-only by decision, not by oversight, and the reason is written into the script header: ~/.claude.json is global config written by live sessions, herdr/session.json is live process state, and IntelliJ rewrites its own files when the project is reopened.
While adding them the author found a trap worth recording: JetBrains stores $USER_HOME$/LTMS/claude-bridge, not the absolute path, so a plain absolute-path grep reports 0 hits on files full of them. The script now counts both forms and says which form matched.
Every apply-mode branch is unexecuted. Running it moves the directory this system runs from and stops the daemon the workers talk through, so the script is well-reasoned, not proven. Its first real run is its test. Verified before merge: CI run 215 green on both jobs, bash -n passes. shellcheck is not installed on this machine, so it never ran.
--check gains section 6 counting ~/.claude.json (the projects entry keyed
by the old absolute path), ~/.config/herdr/session.json, and JetBrains
recentProjects.xml / trusted-paths.xml across every IntelliJIdea* version.
Apply mode's closing summary lists the same three with manual follow-ups.
None of the three is rewritten automatically, on purpose: ~/.claude.json
is global live Claude Code config, herdr session.json is live process
state, and JetBrains rewrites its own files when the project is reopened
at the new path. Header comment states the reason for each so it does not
read as an oversight.
JetBrains stores these paths as its $USER_HOME$ macro rather than a
literal absolute path, so the counter matches both forms and says which
form the hits used.
Renames ~/LTMS/claude-bridge -> ~/LTMS/fleetd as one auditable command,
modeled on redeploy-bridged.sh. --check reports every surface holding the
old absolute path (checkout files, launchd plist, Claude Code project
state, worktree .git pointers, running daemon). Apply mode stops the
daemon first (launchctl-aware), moves the checkout and the Claude Code
project-state slug dir derived from both paths, repairs worktree gitdir
pointers, rewrites bridged.yaml and the installed plist if they exist,
restarts, and verifies /healthz plus a fresh 'bridged listening' line
anchored to a pre-stop marker.
The contract job reaches the broker by network alias (AMQP_URI=amqp://guest:guest@rabbitmq:5672), so the host mapping 5672:5672 was never used. It only bound a port on the runner host, which made two concurrent runs collide and the service container fail to start.
Seen on 2026-08-22: run 209 (PR) contract=success and run 210 (the merge of that same code) contract=failure with every step, including checkout, marked cancelled. Both started at 22:08. Re-running 210 alone on the identical commit passed.
PR #139's own run 212 is green on both jobs with the mapping gone, which proves the alias path still works.
Also the worker-PR proof for CB-623 (#127): opened by the agent account against fleet/fleetd after the org transfer.
The repo moved lms/claude-bridge -> fleet/claude-bridge -> fleet/fleetd.
Gitea redirects hold, so most of this is not urgent, but one line was a
real break: the implementer skill posts a worker's PR to a hardcoded
repo path, so every worker PR would have gone to the old address.
.claude/skills/implementer/SKILL.md the worker PR endpoint (functional)
.gitmodules wiki submodule URL
CLAUDE.md + wiki/7-Use-Cases.md the canonical block, kept byte-identical
README.md clone command and wiki link
deploy/bridged.service Documentation=
plugin/.claude-plugin/plugin.json homepage + repository
docs/*.md issue and wiki links
The wiki is not a separate repo. /repos/lms/claude-bridge.wiki returns 404
and lms owned no .wiki entity, so the wiki moved with the repo; both the old
and the new wiki SSH URLs resolve to the same sha. Ticket step 4 assumed a
second transfer that does not exist.
The eleven tool names become fleet_* throughout the portable block. The
fallback ladder keeps mcp__bridge__* unchanged and that is deliberate: both
launchers still mount the server as "bridge", so a spawned member really
does see that prefix. Renaming the mount is a separate surface CB-622 did
not touch.
Separately, a pre-existing defect: the ladder quoted the reply charter as
"You are an off-subscription worker in the claude-bridge fleet", while
REPLY_CHARTER says "You are a spawned member in the claude-bridge fleet".
A member matching that quote found nothing, so the ladder's first rung
could never fire. Now quoted verbatim.
wiki/7-Use-Cases.md is advanced to the matching template commit; the sync
check prints in sync: True.
1. opencode.json — mount key bridged -> fleetd. Unit C could not do this:
the file is neutralized by the worktree overlay, so every worker sees a
stub and correctly reported it held no mount key. Tracked as CB-628.
2. e2e/bridge_ask_transcript.md keeps its name. It is a dated record of a
run on 2026-07-16 that really did call bridge_ask, and its first line
says so. The harness now writes fleet_ask_transcript.md for new runs;
its header already says 'Live fleet_ask', so the two now agree.
3. docs/MCP-Contract.md line 20 — the historical banner describes section
6, and section 6 now says fleet_ask.
Lead-verified: 27 .md files, 165 occurrences, no wiki/, no CLAUDE.md. The docs/MCP-Contract.md boundary held exactly — every hunk falls in 223-286, inside section 6 (210-293). The transcript rename is reverted by the lead in a follow-up: that file is a dated record of a run that really did call bridge_ask.
Lead-verified: 9 files, +68/-68, nothing under wiki/, root .mcp.json untouched. The worker's opencode.json finding was correct about its worktree and led to CB-628 (#134); the tracked file is fixed separately by the lead.
Lead-verified: both names reach the same handler instance; warn-once uses a per-name Set, not a numeric sentinel; Authz keys on an action enum so the deprecated names keep their authorization. 878 tests green in a local integration with #132, #133 and the lead's opencode.json fix. Live MCP verification is still outstanding and is the lead's step after redeploy.
fleet_* is now the documented tool name for all eleven MCP tools; each
bridge_* twin is registered against the exact same handler (no logic
duplication) and its description leads with a DEPRECATED notice. A
bridge_* call logs one WARN naming the old and new name, once per name
for the life of the process (a Set, not a numeric sentinel).
REPLY_CHARTER in HerdrPeerLauncher now tells a spawned member to call
fleet_reply — the one rule that must survive with no repo checkout.
All other bridge_* string literals across mcp/, Javadoc, and tests were
renamed to fleet_* for consistency with the new documented name.
Rename the eleven MCP tool names (bridge_ack/ask/list/poll/profiles/reply/
send/spawn/status/stop/whoami) to their fleet_* names across the Markdown
documentation. fleet_* is written as the normal name; one deprecation note
in README.md says bridge_* still works for one release.
docs/MCP-Contract.md is renamed only inside section 6 (lines 210-293):
sections 1-5 and 7-11 are stale pre-build design text (CB-609) and are
deliberately left with old names so dead text does not look maintained.
Also renames e2e/bridge_ask_transcript.md to e2e/fleet_ask_transcript.md
to match its content. CLAUDE.md, wiki/, plugin/skills/setup/SKILL.md and
.claude/skills/port-to-opencode/SKILL.md are owned by other units and are
untouched.
Two defects in CB-617 that only a live spawn could find. Both were shipped
green: every test passed because every test read the argv we built, and none
ran the binary that has to accept it.
1. Claude Code refuses to start when both --append-system-prompt and
--append-system-prompt-file are on the command line:
Error: Cannot use both --append-system-prompt and
--append-system-prompt-file. Please use only one.
CB-617 put the role charter on the file flag and left the reply charter on
the inline flag, so every Claude-profile spawn with a role charter died at
launch. The pane exited on its own and bridged reported it as
spawn_timeout ("did not reach injectable state within 20000ms"), which
hides the real cause.
Both charters now go in the one file, role charter first and reply charter
last — last is where the reply rule must sit, because it is the rule that
must survive. A member with only a reply charter keeps the proven inline
flag, which is also the only form that reaches a member with no repo
checkout.
2. A Claude Code agent definition needs `name:` in its frontmatter. Ours had
only `description:`, so the files were skipped and --agent architect failed
with "not found. Available agents: claude, Explore, ...". Added to all
three. The OpenCode files take their name from the filename and are
unchanged.
Checked on this host, in this repo, with the real binary:
claude --model claude-sonnet-5 --agent architect \
--append-system-prompt-file /tmp/combined.md -p '...'
-> ROLEOK, REPLYOK, yes
so the agent definition, the role charter and the reply charter all compose.
The updated test now asserts the constraint that actually binds: with a role
charter present, --append-system-prompt must be absent, and the file must open
with the role charter and end with the reply charter.
874 tests pass.
herdr refused to shell-encode a multi-line inline --append-system-prompt
argument (invalid_agent_argument), which broke any Claude Code profile with
a configured multi-line fleet.charters.<role>. Split delivery: the role
charter now writes to a temp file and mounts via
--append-system-prompt-file; the one-line REPLY_CHARTER keeps its inline
--append-system-prompt delivery, since it must reach a member with no repo
checkout. Also pass --agent <role> when a role's agent-definition file
exists under the worker's cwd (.claude/agents/<role>.md for claude-code,
.opencode/agent/<role>.md for opencode); absent either file, nothing extra
is added and the member still spawns.
The base's LaunchSpec now carries roleCharter/replyCharter/cwd alongside
the existing composed charter field, so CharterReceipt keeps fingerprinting
the same composed text it always did -- unchanged, since the digest covers
the full logical charter content regardless of how it is delivered.
A page audit found the dead host still named in four places outside the wiki.
ollama.ltms.dev no longer exists; the one front door is llm.ltms.dev, and the
path differs by backend - /anthropic for claude-code, /v1 for opencode.
Anyone following README.md today points a member at nothing.
README.md also now links the new wiki chapter 13, the operator guide, since
the README's own next step used to be the never-written Setup page.
Not touched: SubscriptionGuard.java:14 still names ollama.ltms.dev, but only
in a javadoc line describing the Stage-1 example allowlist. It is a comment
about history, not a live default, so it stays until that class is next
edited for its own reasons.
profiles.<name>.subscription is read in five places (BridgedConfig record +
isSubscription + validateSubscriptionProfiles, ClaudeCodeLauncher, and the
startup secret check that deliberately skips it) and was in bridged.example.yaml
nowhere. bridged.yaml is gitignored, so the example is the only place an operator
can learn a key exists - which meant a fresh host had no way to discover the one
switch that moves cost onto the operator's own subscription. Two profiles here
have it set.
Documents what changes when it is true, that it is mutually exclusive with
baseUrl and refused at load, that the startup secret check skips such profiles
so a clean secrets report says nothing about them, and that maxLoad is the only
throttle against the operator's plan - there is no metering or budget refusal.
94 config tests pass; the example still loads.
The instruction surface named a parameter the tool refuses. bridge_ack requires
target and msgId (BridgeMcp.java:1001) and errors with 'target and msgId are
required' otherwise, but step 5 of the primary procedure said bridge_ack{ticket,
msgId}. The table two sections below it was already correct, so the file
contradicted itself and the wrong half was in the numbered steps a lead follows.
The same line is fixed in the byte-identical wiki template.
docs/MCP-Contract.md still opened with 'Greenfield - no MCP code exists yet' and
CLAUDE.md pointed every session at it. Audited against BridgeMcp.java: it names
two tools that do not exist, omits four that do, gets nearly every parameter name
wrong, and uses /workers paths the daemon does not serve. Its status banner now
says so and names the live schema as the authority; CLAUDE.md's pointer is
narrowed to section 6, which is the part that did survive. Rewrite is CB-609.
Replaces CB-592's single hardcoded GITEA_ACCESS_TOKEN shadow with a
memberCredentials: block: policy + allow + known, validated at config load
in the shape CB-606 established.
What this does and does not do, because the distinction matters:
IT DOES enumerate, report and validate. No credential name is hardcoded in
Java any more. An unrecognised policy: refuses at load. A credential-shaped
env var on neither known: nor allow: is named in a WARN. An absent block
warns loudly at startup, because absence now blocks nothing at all and there
is no hardcoded fallback left.
IT DOES NOT, on its own, block anything the login shell re-exports. The env
overlay is applied at pane creation and the pane's login shell runs after it.
An exec-time fix was attempted and has no seam: herdr protocol 19's
agent.start takes a fixed kind plus trailing CLI args, with no env map and no
argv[0] control. The enforcement therefore lives in the operator's secrets.sh,
which sets BRIDGED_MEMBER-guarded sentinels for the 29 blocked names.
Verified by me, not taken on report: my own unpiped mvn clean install gives
870 tests, BUILD SUCCESS; loading the live gitignored bridged.yaml with the
new code yields 34 known / 5 allow / 29 blocked with an empty intersection;
and that blocked set diffs clean against the secrets.sh block, so the two
lists agree today.
herdr protocol 19's agent.start takes a fixed `kind` (herdr resolves the
executable) plus trailing CLI args for that binary; only tab.create/pane.split
accept an env map, and that IS the round-1 pane-creation overlay already
shipped. There is no argv/env control point that runs after the pane's login
shell and before the agent process starts, so the proposed `env NAME=value ...`
argv prefix cannot be implemented against this API. Documented the finding and
corrected bridged.example.yaml's round-1 comments, which had overclaimed that
the overlay survives the login shell.
Kept everything else: added a startup WARN (Bridged.reportMemberCredentialsGap)
when memberCredentials: is absent or its known: list is empty, so CB-592's
protection loss is never silent, mirroring CB-594's reportRequiredSecrets.
Replaces CB-592's single hardcoded GITEA_ACCESS_TOKEN shadow in HerdrPeerLauncher
with BridgedConfig.MemberCredentials (memberCredentials: policy/allow/known).
Every known name not also allowed is overlaid with a sentinel; an unrecognized
policy value refuses at load; a credential-shaped host env var on neither list
is logged as a gap (name only, never a value). bridged.yaml is gitignored, so
the 31 measured names + 4-name allowlist ship as a commented block in
bridged.example.yaml for the operator to apply live.
SessionReaper.lastWipSweepNanos started at Long.MIN_VALUE as a "never swept
yet" sentinel. That sentinel cannot be compared by subtraction. nanoTime() is
positive on this platform, so `now - Long.MIN_VALUE` wraps to a large negative
number, the gate `delta < WIP_SWEEP_INTERVAL_NANOS` reads it as "swept moments
ago", and the method returns before the assignment that would have fixed the
field. The sweep never ran once, for the life of the process, and nothing in
the log said so.
Measured:
System.nanoTime() = 31305820625625 (positive)
now - Long.MIN_VALUE = -9223340731034150183
interval (6h in nanos) = 21600000000000
gate 'delta < interval' -> true => returns early, every iteration, forever
Fix: a separate `sweptOnce` boolean holds "never yet", so the subtraction only
runs once both operands come from nanoTime. The first pass always sweeps — a
restart is a fine moment for it, the 24h age floor keeps it safe, and the
feature becomes observable right after a redeploy instead of six hours later.
The existing tests all passed because they call SessionManager.sweepWipRefs
directly, which walks around the gate. The new test asserts through the reaper
loop instead: it spawns a worktree session, starts the reaper, and requires a
real prune call at the seam with the 24h floor intact. Removing the fix makes
it fail with "it never reached the seam".
860 tests, mvn clean install, BUILD SUCCESS.
A refs/wip snapshot ref is deleted only when both hold: its commit's tree is
already reachable from main, and it is older than 24h. Reachability is the safety
floor — a snapshot exists because the work was committed nowhere else, so an
unreachable one is the last copy and is never swept. Every deletion logs the ref
and the sha.
/members gains wipRefs{count,costBytes} so the growth is visible.
Verified against real git, not only the fakes: a recoverable+old ref is deleted,
a recoverable+young one survives the age floor, and the last copy survives. A repo
with no main deletes nothing, and a repo with no snapshots is a clean no-op.
Closes#67
The prompt is part of the product: CB-582 changed what bridge_status returns, so
the primary's step 5 was no longer the whole truth.
The second half matters more than the first. A nudge makes a lead more likely to
notice an ask; it does not widen the ~55s window, which is bounded by the
WORKER's own MCP client timeout, not by anything the daemon chooses. Without that
sentence a lead reads 'the ask now nudges me' as 'asking works now' and briefs a
worker to ask — which is the failure CB-582 was filed about.
Add the CB-586 retention rule to GitWorktrees and drive it from the reaper:
a snapshot is deleted only when its tree content is already reachable from
main AND the ref is older than 24h. Reachability keeps the last copy of a
worker's work; the age floor stops a fresh snapshot being swept while a
lead is still looking at it. Every deletion logs the ref name and commit
sha so it is recoverable from the reflog. The /members response gains a
wipRefs{count,costBytes} census the operator can read without shelling
into the repo.
Issue #82 step 1 is a measurement, and the classifier refuses an ad-hoc pipeline
that enumerates credential names inside a member — correctly. This is the seam:
one file the operator reads once and then runs, instead of approving a shell
pipeline they have to take on trust.
It never prints a credential value or any part of one. #82's criterion 1 asked
for a 6-character prefix; this prints a truncated SHA-256 instead. A prefix of a
short secret is most of the secret and would end up pasted into a ticket, while
the hash answers every question the prefix was for — is it set, is it the same
value as over there, is it the CB-592 sentinel.
Refuses to run unless BRIDGED_MEMBER=1, since the finding is what a MEMBER holds;
--allow-outside-member takes the comparison reading and labels it as such.
I have not run the reading path. That is the operator's call, which is the whole
point of the ticket.
Three fields had CB-604's shape — lower-cased, compared against one string,
never checked against the valid set. The auth.mode one was the worst: a typo of
'token' silently behaved as loopback-trust, and validateAuthExposure() only fires
on a non-loopback bind, so a loopback bind hid it end to end. The daemon started
clean and authenticated nobody while the operator believed token mode was on.
All three now refuse at config load, naming the value and the accepted set. The
top-level placement policy was already validated but only lazily at first spawn;
it now calls PlacementPolicies.fromName eagerly at load, so a bad name cannot
start a daemon that merely looks healthy.
Verified by probing the real BridgedConfig.load with eleven values, including the
critical auth.mode typo on a loopback bind, and by loading the live gitignored
bridged.yaml — which the worker cannot see and so could not check.
Closes#106
An auth.mode typo (e.g. "toekn") used to silently fall back to loopback-trust with no signal
anywhere — validateAuthExposure() only checks the pairing on a non-loopback bind, so on the
common loopback bind the daemon started cleanly and authenticated nobody. Per-profile placement
had the same shape, falling back to legacy pane placement. The top-level placement policy name
was already validated by PlacementPolicies.fromName, but only lazily at first spawn through
CompositePeerLauncher's Supplier; it is now checked eagerly at load, calling fromName itself as
the single source of truth.
Found reviewing CB-582 before the merge, not by the implementer.
ask() clears its question on three paths — no-waiter, timed out, and (from
answer()) answered. It can also leave by throwing: an interrupt while blocked on
the answer, or an ExecutionException from the answer future. Those run only the
finally block, which tore down the rendezvous turn but not the push loop's copy.
The result was a question that stayed pending for good: named in every nudge
until it hit its own cap, then left in pendingQuestions with no remover at all.
Teardown now happens where the rendezvous teardown already happens, so the two
cannot drift apart again. Closing a turnId that was never pending is a no-op, so
the normal paths are unaffected.
The new test fails on the pre-fix code with expected: <STOP> but was: <INJECT>.
An async (wait:false) delegation opens a ~55s reverse-rendezvous window when its
worker calls bridge_ask. A lead polling on its normal minutes-long cadence never
sees that window, so the worker times out and proceeds without an answer.
The question is now a third source in the CB-588 per-lead push schedule, with its
own per-item nudge count (CB-598's shape), and it is surfaced by bridge_status and
by REST /sessions/{id}/status and /tasks/{ticket} (which previously dropped turnId
on an ASKING phase, so a REST caller could see the question but not answer it).
Closes#61
kind: was lower-cased and compared against one string, so a typo like
'opencod' was accepted and routed to the claude-code adapter. With argv:
unset the launch command became the misspelled string itself, and the
daemon tried to run a program named after the typo. Nothing said a word
until the spawn failed.
It now throws at load, naming the profile, the bad value and the
accepted set - matching rejectNegativeMaxLoad and the duplicate-adapter
check, which already treat routing mistakes as fatal.
Verified here by probing the real BridgedConfig.load with five values:
opencod refused with the full message; opencode, OpenCode, claude-code
and an absent kind all accepted with the right adapter. 832 tests,
BUILD SUCCESS, unpiped.
Closes#102
bridge_ask blocks the worker's turn for ~55s by default (BridgedApp.java,
BridgeMcp.java) — a value deliberately kept just under the worker's own MCP
client's ~60s call cap so the daemon can return a clean timeout before the
client severs the call, not a value that can usefully be widened. A lead
following the charter's wait:false + poll cadence is minutes away, so the
window closes long before a poll would ever see the question — and until now
bridge_poll on such a ticket just read as ordinary "pending" progress.
bridge_poll(ticket) already surfaced Phase.ASKING with the question and
turnId (CB-205); this ships the two pieces that were still missing:
- The lead's own pane is now nudged the instant a question opens, reusing
the CB-588 ReplyPushLoop push mechanism (a third source alongside queued
replies and terminal tickets) rather than a new path. The nudge is capped
by the loop's existing maxReminders budget, and stops the moment the
question is answered or lapses.
- bridge_status(sessionId) and REST GET /sessions/{id}/status now also show
an open question and how to answer it, via a new
MessageService.pendingAsk() lookup — covering the case where a lead checks
status directly rather than the ticket.
- The REST /tasks/{ticket} endpoint was silently missing turnId on an ASKING
phase (only the MCP layer's formatted text carried it) — fixed as part of
making the state genuinely visible over both surfaces.
An unanswered question still behaves as today: the worker proceeds and its
reply says the ask went unanswered — not a hard failure.
Most of issue #65 was already shipped in 5d5b3bd - MemberSession
records agentSessionId, PeerHandle.agentSessionId() has no default, the
roster exposes it, and bridge_spawn accepts sessionName and
resumeSessionId. That commit left one item for follow-up: carrying the
id on ReleaseDetail.
This does that item. CB-578 stage C already lets a lead re-dispatch onto
the same worktree after a failure; without the session id that is a cold
start. With it, the work and the thread both survive.
Also updates the bridge_spawn and bridge_list rows in the CLAUDE.md
intent table, which the earlier commit missed.
Verified here: trial merge onto main builds 830 tests BUILD SUCCESS,
unpiped. Confirmed against the code that items 1-4 really were already
on main, so the scope-down is correct rather than work skipped.
Closes#65
The weight ratio does not express a preference order. weighted spreads
spawns across every profile with a free slot, so paid spawns happen
while the free box is idle - and a profile at maxLoad freezes its score,
so it can lose the next pick after a slot frees.
The real fix is a cost-first policy (CB-589). This documents the
workaround and its trap next to the key, because bridged.yaml is
gitignored: a fresh host starts without the workaround and quietly pays,
with nothing to tell the operator why.
Comments only. BridgedConfigTest: 85 tests, BUILD SUCCESS.
Closes the one piece issue #65 deliberately left out of 5d5b3bd: a
released member's agentSessionId now rides alongside worktree, branch
and snapshotRef on ReleaseDetail, and Bridged's onRelease handler
names it in the abandon reason, so a lead can resume the member's
conversation instead of only re-dispatching a fresh one onto the same
files.
Also updates the CLAUDE.md bridge_spawn/bridge_list table row, which
5d5b3bd shipped the sessionName/resumeSessionId/agentSessionId surface
for but never updated.
Three gaps that only bite once the agent is loaded, plus one wrong
comment.
The script computed its log path from its own location while the plist
hard-codes one. Run from a different checkout, every post-restart check
would read the wrong file and report a clean restart while the daemon
crash-looped. It now compares the two and fails, not warns.
A failed 'launchctl load' after a successful 'unload -w' left the agent
stopped AND persistently disabled - worse than before the redeploy. It
now retries once, then dies naming the exact recovery command.
The plist now says plainly that ThrottleInterval paces restarts but does
not bound them, and what actually stops the loop.
Verified here: ran the script with --check from the merged tree and it
behaves exactly as before, so the unsupervised path - my only restart
route - is intact. Exercised the log-path check against match, mismatch
and missing-plist fixtures using a truncated copy with no mutating code
in it: ok/1/1. 829 tests BUILD SUCCESS.
Closes#91
bridged.yaml is gitignored, so bridged.example.yaml is the only
committed description of the config schema. Two tests already covered
example -> code; nothing covered code -> example, so a brand-new key
could ship undocumented and no test would notice.
A new test compares BridgedConfig.KNOWN_TOP_LEVEL_KEYS against the
example scanned as TEXT, so a key documented only as a comment counts as
documented. That is what makes the guard correct rather than annoying:
most of the example is commented on purpose.
Verified here: added an undocumented key and watched the test fail with
an actionable message naming it; then documented that key as a comment
only and watched it pass. Probe reverted, tree clean.
Closes#96
The reminder count was one counter per lead per source, carried forward
across ticks. A counter carried forward has no memory of which item it
counted, so work arriving during the backoff window inherited an
already-capped count and was never named in a nudge.
tick() now recomputes each source's count fresh from the minimum count
among the items actually pending, tracked per item. A fresh item keeps
its source eligible; an older capped item still rides along in the text
without spending more budget. decide() is unchanged.
Verified here: read the diff; the bumped set is exactly the set named in
the nudge, and the empty early-return skips the bump. Trial merge onto
main builds 826 tests BUILD SUCCESS, unpiped.
Closes#87
Background loops call the fake from their own scheduler threads while a
test polls called() from the test thread. The list was a plain
ArrayList, so a nudge landing mid-stream threw
ConcurrentModificationException out of called().
It surfaced while I was verifying CB-598, which nudges more often, but
the race is on main today and is unrelated to that change.
824 tests, BUILD SUCCESS.
- redeploy-bridged.sh now refuses (not warns) a supervised restart when
its computed log path disagrees with the loaded plist's StandardOutPath
— otherwise every post-restart check reads the wrong file and can
report a clean restart while the daemon crash-loops. The check is a
pure, testable function; the script gained a source-for-test guard so
it can be exercised without installing the agent or touching launchd.
- a failed 'launchctl load' after a successful 'unload' now retries once
and, on ultimate failure, tells the operator the agent is stopped AND
disabled plus the exact recovery command, instead of leaving that
silently worse than the pre-redeploy state.
- the plist documents honestly that the crash loop launchd retries is
unbounded (ThrottleInterval only paces it), and what actually stops it.
- fixed the requiredSecretEnvVars javadoc: the auth.tokenEnv startup
throw is ~370 lines below its call site, not a few lines above it, and
only fires in auth.mode: token.
The test asserted one of two interleavings that are both correct, and
steered toward it with a 5 ms Thread.sleep. Under load the other
interleaving happened and main went red on a correct implementation.
The head start is now a latch counted down from inside the sweep's
guarded loop, so the ordering is guaranteed, not likely. Test file only;
the production guard is unchanged.
Verified here: 822 tests BUILD SUCCESS; 10/10 passes while a full clean
install ran in parallel; 5/5 failures with the guard removed, so the test
still catches the bug it exists for.
Closes#95
BridgedConfig.KNOWN_TOP_LEVEL_KEYS is now package-private so a test can assert
every key the parser accepts appears in bridged.example.yaml — live or
commented-out, since the file is gitignored and the example is the only
committed description of the config schema. The existing tests only checked
the example->code direction; this adds code->example.
My premise for this ticket was wrong and the worker corrected it. I reported five whole sections missing from the example; nothing was missing. My comparison script only counted uncommented lines, so every section documented as a commented-out example looked absent. Earlier tickets had each updated the example alongside their feature.
What it found instead is more useful than what I asked for — real inaccuracies, found by tracing each field through the parser and its consumers:
- `health.workingSuspectAfterSeconds` and `paneProbeIntervalSeconds` documented enforced minimums that do not exist. I checked: both names appear **only** in the `Health` record declaration and are read by nothing. Only `intervalSeconds` is clamped, and it is silently raised to 15 rather than rejected.
- `notifications.mode: webhook` only flips what `bridge_list` reports as `healthCoverage`. It sends no webhook — "webhook" appears in one `configured()` boolean and there is no delivery code in the repo.
- `lifecycle.clearAfterTurn` was undocumented, and is a no-op for any peer kind other than claude-code.
- The reload doc claimed the whole `fleet:` block is hot; `fleet.leaders` is built once at startup and is not rebuilt, so a change is silently accepted and does nothing until a restart.
- The `fleet.leaders` demotion consequence is now stated next to the block itself: an unmatched pane is silently an ordinary worker and every orchestration call it makes is refused, with no startup error.
Documenting a knob as dead is worth more than documenting it as working. Someone tuning `workingSuspectAfterSeconds` would otherwise have concluded their monitor was broken.
Comments only — no parsing or production code touched. Verified by the lead: parses cleanly under the project's own snakeyaml 1.30, top-level live keys `[bind, herdrSocket, profiles, placement, fleet, guard]`, the rest correctly commented examples. Both dead-knob claims verified by grep against `src/main` rather than taken on the worker's word.
`PlacementException extends IllegalStateException`, and neither spawn path caught that type, so it escaped to Javalin's default handler as a bare `500 Server Error` with a text/plain body — while every other failure on the same endpoint returned structured JSON. The reason existed and was good, but only in the daemon log.
I hit this live while orchestrating: asked for a member on a full profile, got a blank 500, guessed another profile, got a blank 500 again, and spent two round trips learning things the daemon already knew.
Both surfaces now catch it. REST returns 503 with `{"error":"no_capacity","detail":...}`; MCP returns the same reason in the `isError` shape it already uses for every other spawn failure. 503 is right because the request was valid and will likely succeed later — the caller did nothing wrong, so 400 would have been a lie.
The other throw sites all funnel through the same type, so quarantine cooldowns, weight-0 exclusion, and the all-at-cap / all-quarantined / all-unreachable messages now reach callers too. That last group matters most: those three distinguish "wait a moment" from "your backends are gone", and all three used to arrive as the identical blank 500.
Tests assert the caller can read the *reason*, not merely that the status changed — one per surface.
Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 824, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.
No exception message was reworded. This change delivers messages that were already written.
Work that arrived during the ~15s push_backoff_ms window between two
ticks landed in the pending map before the next tick's start-of-tick
snapshot, so a shared per-lead-per-source counter (carried forward via
scheduleNext(lead, count+1, ...)) already treated it as exhausted
backlog even though no nudge had ever named it. ReplyPushLoop.tick now
recomputes each source's reminder count fresh every tick as the
minimum nudge count among that source's currently pending items, so a
freshly-arrived item (count 0) keeps its source eligible regardless of
how depleted an older, still-undrained sibling's count is. decide()
itself is unchanged.
PlacementException extends IllegalStateException, which neither BridgedApp
nor BridgeMcp's spawn catch blocks handled, so a maxLoad/quarantine/
all-exhausted refusal fell through to a blank 500 on REST and lost its
message on MCP. Catch it on both surfaces, ahead of the unrelated
IllegalArgumentException(unknown_profile) mapping, and return its message
structured: REST as {"error":"no_capacity","detail":...} with status 503,
MCP as an isError result prefixed "no capacity: ...".
Audited bridged.example.yaml against BridgedConfig's KNOWN_TOP_LEVEL_KEYS and
found every top-level key already documented (broker, health, lifecycle,
configReload, quarantineCooldownSeconds, fleet.leaders/architects/reviewers
included) — CB-573/CB-566/CB-559/CB-579/CB-527/528 each updated the example
alongside their feature. What was actually wrong:
- health.workingSuspectAfterSeconds/paneProbeIntervalSeconds claimed enforced
minimums (300/60) that don't exist in code — only intervalSeconds is
clamped (floor 15); the other two are parsed but never read anywhere.
- notifications.mode: webhook was undocumented as only flipping the
healthCoverage label bridge_list reports — no webhook is ever sent.
- lifecycle.clearAfterTurn was missing entirely.
- the HOT bullet under configReload claimed the whole fleet: block reloads
live, but ConfigRef's own javadoc carves out fleet.leaders as needing a
restart with no deferred-list warning — added that exception.
- fleet.leaders' demotion consequence (unmatched tab -> silent WORKER
demotion, no startup error) is now stated inline next to the block, not
just implied by the multi-lead rationale higher up.
The launchd unit was a CHANGEME template that had never been installed, and it could not have worked if it were: launchd does not source a login shell, so the daemon would have started with no forge or gateway token, and the failure would only appear much later as workers unable to open a PR.
Four parts: a wrapper that execs one login shell in place so the job inherits the secret store; a startup report naming which required secret env vars resolved and which are MISSING, by name only, never a value; the plist filled in with this host's real verified paths; and `redeploy-bridged.sh` detecting the agent and switching stop/start to `launchctl unload -w` / `load -w`, falling back to the existing kill + nohup when it is not installed.
Verified by the lead. My own unpiped build of the branch: 813 tests, BUILD SUCCESS. Trial-merged onto current main (which had moved twice) and rebuilt: Tests run: 822, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. All three host paths in the plist exist; no CHANGEME remains; the secret report leaks no values.
An independent reviewer, briefed only from the diff, verified the parts that matter today and said merge. It confirmed live that the unsupervised path is unchanged, that `launchctl list` exits 113 for an absent agent so the detection reads correctly, that the wrapper round-trips arguments containing spaces and quotes and stays a single exec chain, and that neither branch of the secret report can print a value.
Its two open findings only bite once the agent is actually loaded, which has not happened and is the operator's call. Filed as #91 — that must land before the agent is ever installed. The important one is that the script computes its log path from its own location while the plist hard-codes an absolute one; if they ever disagree, the post-restart ERROR check reads the wrong file and reports "ok" while the daemon crash-loops.
The agent is deliberately NOT installed by this merge. Nothing here changes how the daemon runs today.
Carries PR #84 (its head `78ca24d` is an ancestor of this one), so this single merge delivers both rounds.
CB-590 collapses the CB-307 reply schedule and the CB-588 ticket schedule into one per lead, which closes the double-injection race. The review round found that this also collapsed the two reminder caps into one shared budget — so a busy reply stream could exhaust the cap and a ticket arriving afterwards would never be nudged at all. That was a regression, not a pre-existing wart: before this PR the two sources had independent counters.
The follow-up keeps the single schedule and gives each source its own budget. `decide()` returns INJECT while either source has pending work under its own cap, and STOP only when neither does. `stopOrRestart` and its snapshot-diff race logic are untouched.
Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 814, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. `ReplyPushLoopTest`: 41 tests.
Known and deliberately out of scope: an item arriving during the 15s backoff is already inside the tick's "before" snapshot, so at-cap work can be abandoned rather than nudged once. That shape predates this PR in the ticket-only loop and is filed as #87 (CB-598).
Carries PR #83 (its head is an ancestor of this one), so this single merge delivers both rounds.
CB-527 bounds prefetch. CB-528 makes a publish wait for its broker confirm, so a failed publish is never reported as success, and correlates a `mandatory` Return back to the right publish.
The review round on #83 found one real race: `failPendingPublishesOnRecovery` swept the pending maps without holding `publishChannelLock`, so a publish issued on the already-recovered channel could be failed by the sweep — the exact inversion of what CB-528 exists to prevent, and undetectable by dedup because `MessageService.reply` mints a fresh msgId per call. Fixed by taking the same lock. `close()` now fails in-flight publishes promptly instead of letting them time out after 10s.
Verified by the lead, not taken on the worker's word:
- `mvn -f bridged/pom.xml clean install` unpiped: Tests run: 809, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
- `mvn test -Pcontract -Dtest=AmqpReplyInboxContractTest` against a real broker: Tests run: 8 — BUILD SUCCESS
That second run mattered: #83's own new tests are contract-tagged, so the default suite and CI prove nothing about them.
decide(lead, reminderCount) shared one counter across the reply and
ticket sources after PR #84 collapsed both onto a single per-lead
schedule. A reply stream that used up the whole budget could then
make decide() STOP even for a ticket that had never been nudged and
had coalesced onto the same still-active schedule — stranding it with
no live schedule left, since stopOrRestart's racedIn check does not
save work that was already present in the "before" snapshot.
decide() now tracks a per-source count (replyReminderCount,
ticketReminderCount) and returns INJECT while either source is still
under its own cap, STOP only when both are exhausted. Still exactly
one schedule per lead; stopOrRestart's snapshot-diff logic is
untouched.
failPendingPublishesOnRecovery walked and cleared pendingBySeq/pendingByMsgId
without holding publishChannelLock, so a publish() that registered while the
sweep was still iterating could be failed even though it published on the
already-recovered channel — a successful publish reported as failed, and
since MessageService.reply mints a fresh msgId per retry, dedup can't catch
the resulting duplicate. Guard the sweep with publishChannelLock: publish()
only holds it for the seq/map-put/basicPublish, so the sweep can only ever
wait for an in-flight basicPublish to return, never a broker round trip.
Also make close() fail in-flight publishes immediately with a clear message
instead of leaving them to idle out the 10s confirm timeout, and record the
(currently unreachable) msgId-uniqueness assumption pendingByMsgId relies on.
AmqpReplyInboxRecoveryRaceTest drives the sweep and a real publish() against
each other directly (no live broker reconnect) using Proxy-backed fake AMQP
channels and a large in-flight backlog to make the race window observable;
confirmed it fails without the guard (reply "fresh" wrongly failed as
"connection recovered mid-publish") and passes with it.
Adds scripts/bridged-launchd-wrapper.sh so the launchd-run daemon still gets
WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN by execing through a login shell (launchd
never sources secrets.sh itself). bridged now logs at startup which required
token env vars (derived from each profile's tokenEnv/gitTokenEnv, not a
hand-written list) resolved or are MISSING, by name only. Fills in the real
paths in deploy/dev.ltms.bridged.plist for this host and points it at the
wrapper. scripts/redeploy-bridged.sh now detects a loaded launchd agent and
uses launchctl unload/load instead of a raw kill+nohup, because a bare
SIGTERM exits this JVM at 143 (measured) which KeepAlive.SuccessfulExit=false
reads as a crash and would race the script's own restart; --check reports
installed/loaded state and stays read-only.
Both reply-queued and ticket-terminal nudges could independently decide
to inject into the same lead pane in the same window, since they ran as
two separate schedules keyed differently (worker target vs. lead) that
never checked each other. Replace both with a single per-lead schedule
(activeLeads) that drains pending reply targets and pending tickets
together, sends at most one combined nudge per tick, and shares one
reminder cap across both sources — so two injections into the same pane
can no longer overlap, and work queued while the lead is busy is never
lost, only deferred.
CB-527: basicQos(prefetch) on the consume channel before basicConsume, configurable
via broker.prefetch (default 32), so an undrained inbox backlog stays on the broker
instead of growing the JVM heap without limit.
CB-528: publish moves to its own confirm-mode channel with mandatory=true and a
return listener, so an unroutable or unconfirmed reply now throws instead of
vanishing silently. The confirm callback checks the per-message returned flag
(set by the return listener, which the broker always fires before the matching
confirm) so an acked-but-returned publish is still reported as a failure. The ack
path stays on its own channel/lock and never waits on a publish confirm.
CLAUDE.md told every member 'You mount only the bridge MCP' and told the lead
'a worker mounts only the bridge MCP and cannot run your other tooling'. Both
were false for Claude Code members.
Measured by spawning one member per backend and asking each what it actually has:
opencode (gx) 11 bridge tools only claim TRUE
claude-code (local) 11 bridge + 45 gitea + 2 context7 claim FALSE
Source is ~/.claude.json user-scope mcpServers; --mcp-config adds to that scope
rather than replacing it, so the worktree parity overlay (which correctly
neutralises .mcp.json and opencode.json) cannot see or stop it.
The forge tools are mounted but not usable: CB-592's blocked sentinel means
get_me and list_issues both fail with 'invalid username, password or token'.
That is defence in depth working in a path it was not designed for, so the text
now says a mounted tool is not a working tool rather than pretending the tools
are absent.
Block propagated byte-identically to wiki 7-Use-Cases.md (wiki 1d95e3f).
systems/vms moved the LLM route timeout again, 1800s -> 86400s (24h), after
the silent-truncation risk was discussed. They tried request: 0s first: it
removes the total-duration timer, but on an AIGatewayRoute the idle timeout
is derived from the request timeout, so 0s also removed any bound on a
stalled connection.
At 86400s our own MessageService.ASYNC_TIMEOUT_MS (30 min) binds first, so a
runaway request now ends as a clean FAILED ticket we raised instead of a
silently truncated 200. While the gateway sat at 1800s the two numbers were
equal and did not nest.
`local` runs on /anthropic and `gx` on /v1, both weight 100; `local-direct`
stays weight 0 as the escape hatch.
systems/vms fixed both blockers, and each was re-checked from this side rather
than taken on trust:
listener buffer 32 KiB -> 32 Mi ours: 1.2 MB body -> 200 (was 413)
LLM route timeout 60s -> 1800s ours: 101s stream -> 200,
message_stop present, 4000/4000
Neither was deliberate: 32 KiB was Envoy Gateway's default
per_connection_buffer_limit_bytes, and 60s was Envoy AI Gateway's own default.
The 60s bounded GENERATION as well as prompt size — a tiny prompt with a long
answer returned 504 at 60.05s.
Verified with real workloads, not liveness probes. A `local` member read this
document and CLAUDE.md in full — 48,344 bytes of file content, comfortably past
the old 32,768 ceiling — and answered four questions correctly, including the
document's length (said ~456, actual 455). A `gx` member did the same. The
trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it
no longer counts as proof here.
§7.2 is new and is the part that matters later. One risk is ACCEPTED, not
solved: on a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked
encoding cleanly instead of resetting, so a truncated answer arrives as HTTP 200
with no error and no terminator (envoyproxy/envoy#17186, acknowledged 2021,
never fixed; the Dec 2025 fix#42269 is HTTP/2 only and SSE here is HTTP/1.1).
Measured at the old 60s: 200, 61.07s, 2473 of 4000 emitted, message_stop 0,
error events 0, ending on a well-formed frame.
The recommended defence — reject a stream with no terminator — does NOT
transfer to us: Claude Code and opencode are third-party clients and we do not
own their SSE parsing. So this is acceptable because a request would have to run
1800s to trip it, not because we could detect it. If a member ever returns a
confident but truncated answer, suspect this before anything in our own code.
Also recorded, from the upstream bisection: ClientTrafficPolicy is honoured in
standalone `aigw run` but BackendTrafficPolicy is silently ignored, and nothing
external distinguishes them (envoyproxy/gateway#9513). Same silent-default shape
this repo keeps hitting.
bridged.yaml carries the same notes inline (gitignored, so not in this commit).
Refs: gitea #76
I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.
Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:
listener default/llm/http per_connection_buffer_limit_bytes: 32768
Nobody chose 32 KiB; it was inherited.
Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.
Consequences recorded in the doc:
* DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
profiles to it would be designing around a bug.
* Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
deployed, pending their operator's approval. We do not re-test until they
confirm — a half-changed system gives a number neither side can trust.
* In standalone `aigw run` a SecurityPolicy is accepted and then silently
ignored, so "the config was accepted" proves nothing there. They will verify
by re-reading the live config_dump and sending a large request. Same
silent-default shape this repo keeps hitting, one layer down.
bridged.yaml carries the same correction (gitignored, so not in this commit).
Refs: gitea #76
Deployed U1-U2c, restarted, spawned both new profiles for real, then reverted.
llm.ltms.dev answers HTTP 413 above 32 KiB (32768 bytes), on BOTH surfaces:
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
That is far below one agent turn. It is an edge limit (Caddy request_body
max_size, and/or Envoy), so the fix is in systems/vms, not here.
The part worth recording is how it nearly passed. Two members, same message,
same moment: `local` finished in 66s, `gx` never finished at all. `local`
passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB — the profile looked healthy and was a
landmine set to fire on the first turn that reads a file. So §7's checklist
was not wrong, it was too easy; it now says to use a file-reading task.
opencode's failure mode is worse than a crash: it catches the 413, compacts
its context, retries, and loops. Observed 10+ minutes BUSY with no reply. From
the lead's side that is indistinguishable from a slow worker. Reproduced
outside the bridge with the launcher's own generated config, which is how it
became a one-line error instead of a hang; §7.1 records that procedure.
Everything else about the migration checked out and is recorded so it is not
re-tested: token accepted on both surfaces, unauthenticated 401 (the Caddy
proxy does gate, whatever the gateway's own fail-open policy does),
/v1/models exactly ["deepseek-v4-flash"], the guard allowlist accepted
llm.ltms.dev, and the generated opencode provider block is correct with a real
llmk- key.
Also answers §3b's open question: reasoning survives BOTH surfaces —
/anthropic returns a real "type":"thinking" block and /v1 returns a populated
reasoning_content. The feared /v1 translation loss did not happen.
Config state (bridged.yaml is gitignored, so it is described rather than
committed): `local` back on http://gx00.gw:8000, `gx` kept at weight 0,
`local-direct` kept, llm.ltms.dev left in the guard allowlist. The file
carries these numbers and the exact two-key edit to switch back.
Verified after the revert with a task that reads two large files: correct on
all three questions. Daemon pid 66745, jar f1fd659423e6.
Refs: gitea #76
Live check on a member pane showed the CB-592 shadow did NOT take effect:
GITEA_ACCESS_TOKEN inside the pane was still the real admin token.
Measured cause. The overlay itself works — GITEA_TOKEN is injected the same
way, is exported by no shell file, and does reach the pane. The sentinel loses
one step later. A herdr pane runs a LOGIN shell, ~/.zprofile line 41 sources
${SHARED_ENV}/tools/secrets.sh, and that file does a plain unconditional
`export GITEA_ACCESS_TOKEN=...`. A login shell overwrites a value already in
the environment, so the real token is put back before the member starts.
Confirmed directly:
GITEA_ACCESS_TOKEN=cb592-sentinel zsh -lc ...
-> RESULT: sentinel was OVERWRITTEN by the login shell
This defeats any launcher-side overlay for any name secrets.sh exports. No
change in this repo can win it alone.
So this adds the half that does survive: BRIDGED_MEMBER=1, a name secrets.sh
never exports. It is a no-op until the operator guards the export:
[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...
Setting it now costs nothing and makes that one line the whole remaining fix.
The sentinel stays: it is correct for any peer kind whose pane does not start
a login shell, and it keeps the intent explicit where every adapter passes.
Also corrects the javadoc and the test javadoc, which both claimed a
protection that was measured not to hold.
The other reported failure was my own bad test, not a regression. The probe
called /api/v1/user, which a minimal write:repository token cannot read. Same
token on the repo endpoint answers 200, so CB-302 is intact:
GITEA_ACCESS_TOKEN: /user=200 /repos/lms/claude-bridge=200
WORKER_GITEA_TOKEN: /user=403 /repos/lms/claude-bridge=200
Tests 805 -> 807. Both new tests proved to discriminate by reverting the
marker: everySpawnMarksThePaneAsAMember and
aProfileEnvEntryCannotClearTheMemberMarker both fail without it.
Refs: gitea #77
A live probe showed every spawned member carried the admin GITEA_ACCESS_TOKEN: 108
environment variables in a member's pane against 99 in the primary's. The operator's
rule is that only the leader and architects may use it; everyone else uses
WORKER_GITEA_TOKEN. We were not enforcing that at all.
The cause is invisible from inside the launcher. baseEnv builds a fresh map holding only
PATH and the profile's env:, so a member looks like it gets a small explicit environment.
That map is an OVERLAY: WorkspaceControl.createTab/splitPane send only the keys it
contains, and herdr spawns the pane from its own login-shell environment, so every key we
never mention passes straight through — admin token included.
The fix puts a non-blank sentinel over the key in baseEnv, applied AFTER the profile's
env: so no profile, present or future, can restore the real token by naming it in config.
One place, every adapter, including peer kinds not yet written — deliberately not a
per-profile bridged.yaml entry, which is the silent-default shape this repo has shipped
nine times.
A non-blank sentinel rather than the empty string, on purpose: whether an empty overlay
value overrides an inherited variable or is skipped as blank cannot be settled from this
repo, because herdr's merge happens in an external process. baseEnv's own PATH seeding
(CB-511) already relies on a non-blank value replacing an inherited one, so this reuses
the shape that is demonstrated to work rather than the one that is merely plausible.
CB-302's repo-scoped GITEA_TOKEN grant is untouched — a worker can still open its own PR.
The subscription boundary was checked and is unaffected: the primary's pane carries no
ANTHROPIC_* at all, so nothing is inherited there.
Closes gitea #77. Live verification follows separately: the daemon must be redeployed
before this reaches any pane.
herdr spawns a pane from its own login-shell process env and layers our map on
top, so any key baseEnv never mentions passes straight through — including the
admin forge token. baseEnv now puts a non-blank sentinel for
GITEA_ACCESS_TOKEN, applied after the profile's own env: so no profile can
restore it. One place, every adapter, every profile including future ones.
CB-302's GITEA_TOKEN grant (applyGitToken) is untouched.
Operator's rule, 2026-08-15: only the leader and architects may use GITEA_ACCESS_TOKEN;
everyone else uses WORKER_GITEA_TOKEN.
opencode.json is TRACKED, so it ships in every worker worktree, and it mounted the gitea
MCP with {env:GITEA_ACCESS_TOKEN}. A live probe confirmed that variable actually resolves
inside a member: herdr spawns each pane from its own login-shell environment and layers
the launcher's map on top, so a member sees 108 variables rather than the small explicit
set baseEnv appears to build. That gave an opencode member admin forge TOOLS — enough to
merge its own PR, which both CLAUDE.md and the member contract forbid.
This is the narrow half of the fix: it removes the tooling. The admin token is still
present as a string in every member's environment, which is the real defect and is
tracked as CB-592 (gitea #77) — that fix belongs in the launcher, in one place, not
per-profile in bridged.yaml where a sixth profile would silently reopen it.
.mcp.json keeps GITEA_ACCESS_TOKEN and is correct to: it is skip-worktree, the primary's
own local copy, and the primary is the lead. That is the pattern this change follows —
the shared tracked file grants least privilege, and anything needing more overrides
locally.
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic
surface and an OpenAI surface, so both member kinds can point at it. The plan is in
docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work.
The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible
endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That
matters more than it looks — every opencode member today is sol or terra, and both sit
on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks
out both. A gateway-backed opencode profile is free and off that credential, which
retires a single point of failure rather than only adding capacity.
Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is
resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token
added to secrets.sh after the daemon started is simply absent: the launcher injects an
empty token and the gateway answers 401, long after the restart and with nothing tying
the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same
login-shell check that never prints the value.
An async delegation ticket (bridge_send wait:false) resolves on MessageService.reply's
rendezvous fast path, which returns before onReplyQueued. So CB-307's push loop only ever
heard about the durable-inbox case, and the mode CLAUDE.md tells leads to prefer never
nudged anyone. Closes gitea #72.
Adds a second, independent reminder schedule keyed by the lead terminal, so several
tickets finishing together coalesce into one nudge. The CB-307 path is untouched.
Three defects were found in review and fixed before merge:
* a pendingTickets entry outlived the ticket it named. poll() returns null once
pruneTerminalTickets drops a ticket, so ticketCollected was never reached and the
entry leaked for the daemon's life, riding along on every later nudge and sending
the lead after a ticket bridge_poll can no longer find.
* a lost nudge: a ticket landing between decideTickets returning STOP and
activeLeads.remove coalesced onto a schedule that was about to die. That is the
exact failure this ticket exists to remove, reintroduced in a narrow window.
* the success direction was unpinned in tests, and the comment listing the paths that
complete the future was short by several.
The obvious fix for the second one was wrong: restarting on any pending ticket defeats
the reminder cap, because a never-collected ticket at cap is expected to still be there.
The fix diffs against a snapshot taken before the decision, so only a ticket that truly
arrived during the window restarts the schedule.
Verified on my own unpiped build: 802 tests, 0 failures, BUILD SUCCESS.
Two reviewers on the diff; the loop-gating finding they raised is split out as CB-590.
- ReplyPushLoop.stopOrRestartTicketLoop: after releasing a lead's active-schedule
slot on STOP, restart only if a ticket landed that the pre-decision snapshot
did not already account for. A naive "restart on any pending ticket" version
was tried first and reverted: it defeated the reminder cap by restarting
forever on a stale, never-collected ticket (broke
successfulTicketNudgeIncrementsDelivered and ticketNudgesSendUpToCapThenStop).
Diffing against a pendingBefore snapshot distinguishes a genuine race arrival
from stale cap-exhausted backlog.
- Two new ReplyPushLoopTest cases exercise stopOrRestartTicketLoop directly
(now package-private) rather than forcing the underlying thread race:
aTicketStillPendingWhenTheLoopStopsIsNotStranded (the race must restart) and
aStaleUncollectedTicketAtCapDoesNotRestartTheLoop (the cap must still hold).
Both were verified to fail against deliberately-reverted versions of the fix
before being restored to green.
- MessageServiceTest: pin the success-path nudge test's negative direction too
(must not contain "FAILED"), not just the failure-path test.
- MessageService: complete the whenComplete comment's list of completion paths
(TIMED_OUT/BUSY/BACKEND_EXHAUSTED via finishAsyncTask, completeExceptionally
on throw, and answer() -> finishAsyncTask(turnId, result)).
tasks is the sole authority on whether a ticket exists, but
pruneTerminalTickets dropped entries from it without telling
ReplyPushLoop. ticketCollected(ticket) was only ever called from
MessageService.poll's terminal branch, which pruneTerminalTickets
short-circuits past once a ticket is gone (poll returns null at the
top). A ticket the lead never polled — or one the reminder cap already
gave up on — was pruned from tasks but never collected in
ReplyPushLoop.pendingTickets, so it rode along on every later nudge to
the same lead forever, naming a ticket bridge_poll could no longer
find, and the map itself never shrank.
pruneTerminalTickets now calls pushLoop.ticketCollected for every
ticket it actually removes (guarded on pushLoop != null), so a pending
nudge entry lives exactly as long as its ticket is pollable. Reused
tasks as the only removal trigger rather than adding a second live
query back into MessageService — no new source of truth.
Added an injectable clock (LongSupplier nowNanos, defaulting to
System::nanoTime) to MessageService, mirroring the SessionManager/
SessionReaper nowNanos seam, so a test can cross the 10-minute
TICKET_TTL_NANOS deterministically instead of sleeping for real.
Confirmed the new regression test fails against the prior
pruneTerminalTickets (a stale ticket rides along on a later coalesced
nudge) before restoring the fix.
An async bridge_send(wait:false) registers a rendezvous waiter, so its
reply always takes MessageService.reply's fast path and returns before
ReplyPushLoop.onReplyQueued is ever called — the exact mode the charter
tells leads to prefer never nudged.
Add ReplyPushLoop.onTicketTerminal(ticket, target, failed), a second
entry point reusing the loop's status gating, bounded/backoff reminders
and metrics, keyed by the nudge-receiving lead so several tickets
finishing together coalesce into one nudge naming the count. MessageService
wires it via task.future.whenComplete in sendAsync (covers reply,
completion fallback, wedge, and abandon() alike) and calls the new
ticketCollected(ticket) from poll() once a terminal view is handed back,
so an already-collected ticket is never nudged again. The CB-307 inbox
path (onReplyQueued/decide/NUDGE_FORMAT) is untouched.
The comment said non-positive means unlimited at load. That stopped being
true when CB-585 made an explicit maxLoad: 0 survive as a real cap of zero
and made a negative value refuse config load. The code below it was already
right — only the comment described the old normalisation. Flagged by the
CB-585 worker, which correctly stayed out of a file not on its list.
Wires up the two links that made session resume unreachable: MemberSession
now records agentSessionId from PeerHandle at acquire (both the plain and
worktree paths), and it survives onto the roster (rosterView, so both
bridge_list and GET /members show it). bridge_spawn accepts sessionName and
resumeSessionId, threading them into the SpawnRequest fields that already
existed but were never reachable from the MCP surface.
A resumeSessionId now requires an explicit profile (a resumed conversation
is tied to the specific backend that started it, so an unqualified spawn
routed by placement has no safe candidate to check) and is refused, naming
Capability.SESSION_RESUME, when that profile's adapter does not declare it
— PeerLauncher gains capabilitiesFor(profileName) so a mixed fleet is
checked per-adapter rather than against the fleet-wide capability union.
PeerHandle.agentSessionId() loses its default, the same fix CB-571 (c00a86b)
made for charterReceipt() one method above it — both existing adapters
already overrode it, so this only closes the landmine for a future one.
Out of scope, left for follow-up: carrying agentSessionId on a failed
ticket's ReleaseDetail (issue #65 criterion 5) — CB-578 stage C just
landed in that file and this ticket deliberately stayed out of it.
--skip-worktree is an index flag. GitWorktrees.snapshot staged into a fresh
empty temp index, which carried none of the real index's skip-worktree bits,
so add -A staged local on-disk content for files git status correctly hides
(e.g. .mcp.json). Copy the real index (resolved via git rev-parse --git-path
index, correct for linked worktrees) into the temp index before staging, so
add -A skips exactly what git status skips.
The doc claimed an in-flight failure re-surfaces the messages on a later
drain, because the ack is local. That was true when InMemoryReplyInbox was
the only inbox, and became false without anyone noticing when the AMQP
adapter landed: there the ack is a broker-side basicAck, so a crash while
writing the response loses the reply outright. Re-polling cannot recover it,
since the broker has already forgotten it.
Behaviour is unchanged and the window stays accepted — acking on the next
poll instead would double-deliver on every normal drain. The point is that
the comment sat exactly where the next implementer would read it and said
the opposite of what happens.
Closes#12.
Found by running it. The script launched the daemon with a relative jar path
(cwd is bridged/) but detected it with an absolute one, so pgrep never matched.
The daemon restarted correctly and booted clean, and the script still failed
with 'no process appeared' — the worst shape of bug for a deploy tool, because
it invites a second restart on a daemon that is already healthy.
Detection now matches both path forms, and the launch uses the absolute path so
ps names which checkout is running.
The lead already owned the redeploy, but the command classifier refuses a bare
kill on the daemon, so in practice every deploy still needed the operator to
approve a stop and a start by hand. A single script is the seam that fixes
that: the operator allow-lists one auditable command instead of two ad-hoc
ones.
It also stops the procedure from living only in a checklist people read after
things go wrong. It builds before it stops anything, so a failed build never
leaves the fleet down; waits for the old process to exit instead of assuming;
polls /healthz; and anchors its log checks to a line marker taken before the
restart, so old errors cannot be misread as new ones.
The check with no log line anywhere in the daemon is the reason --check exists:
bridged inherits WORKER_GITEA_TOKEN from the shell that starts it, and starting
from a non-login shell empties it. The daemon boots fine, healthz is green, and
the failure only appears later as workers that cannot open a PR. --check tests
whether the name resolves and never prints the value.
A CLAUDE.md rule grants intent, not tool permission. Tried the redeploy right
after writing the section and the classifier refused `kill`, so the section
would have been misleading as written. States the honest split until a Bash
permission rule exists: the lead builds and verifies, the operator runs the
stop and start with `!`.
A merge is not a deployment: the running bridged holds the jar it started
with, so merged code does nothing until the daemon is rebuilt and restarted.
Calling that work shipped is a false report. This makes the redeploy the
lead's job rather than something handed back to the operator, and records the
five things that have gone wrong doing it here — chiefly that starting the
daemon from a non-login shell empties WORKER_GITEA_TOKEN, which nothing logs
and which only surfaces later as workers that cannot open a PR.
Goes in the project addendum, not the canonical block; the block is unchanged
and still byte-identical with the wiki template.
CB-578 stage C commits a preserved dirty worktree to refs/wip/<branch> with
git add -A. The parity overlay copies the primary's environment files into
every worktree, so a non-gitignored overlay path now reaches a durable git
object instead of only sitting on disk. add -A respects .gitignore, which is
what stops it — so the rule documented for CB-581 is now what keeps a secret
out of a commit, not just out of a directory.
Adds Worktrees.snapshot(worktreePath, branch, message): stages into a
temporary GIT_INDEX_FILE (never the worker's real index/HEAD), writes
the tree, commit-trees it onto the worktree's current HEAD, and points
refs/wip/<branch> at the result. add -A (never -f) respects .gitignore.
SessionManager.release() now snapshots any dirty worktree before the
preserve-or-remove decision, regardless of release cause (COMPLETED or
SHUTDOWN) — preserving on disk alone is one `worktree remove --force`
away from gone. A failing snapshot never escalates: the worktree is
still preserved, the pane still stops, and the release listener is
still notified.
onRelease's listener now receives a ReleaseDetail (terminal, worktree
path, branch, snapshot ref) instead of a bare terminal id, so Bridged's
WORKER_FAILED wiring can put the same three facts into a failed
ticket's detail — a lead can re-dispatch onto the same tree instead of
starting from the base commit.
A quarantined profile's capacity row now forces free:0 and names the
quarantine (credentialId, quarantinedForSeconds), reusing the same
QuarantineSource bridge_profiles already reads instead of a second
lookup. An ordinary fleet's capacity rows are unchanged (no new keys).
Closes issue #50. Verified by the lead: own unpiped build of the branch merged
onto main — 740 tests, BUILD SUCCESS, exit 0.
Stage A only classified a usage-limit refusal. Nothing acted on it, so the
fleet would spawn another member onto the same exhausted account and fail the
same way. Now a BACKEND_EXHAUSTED classification quarantines the CREDENTIAL
for a cooldown: an explicit spawn onto it is refused with the profile, the
credential and the seconds remaining, placement skips it under every policy,
and it lifts itself on an injected clock.
Keyed by credential, not profile name, because sol and terra are two models on
one OpenAI account. Quarantining only the profile that reported the refusal
would leave its sibling live, and the next spawn walks onto the same dead
account. A profile that sets no credentialId quarantines alone under its own
name, so an existing config behaves exactly as before.
The implementer also found a real bug outside the brief: ConfigRef's
sameLaunchSettings never compared exhaustedPattern, so a reload changing only
that key reported 'config reloaded' while the value is actually deferred —
the precise failure that file's own doc calls the worst outcome a reload can
produce. Fixed with a regression test.
Two things I changed at review. The example.yaml conflict with e09cac6 was
resolved toward the branch, which was a strict superset. And
BackendQuarantine.none() held a clock frozen at 0, so a quarantine() call on
it would have locked a credential out for the life of the daemon — two
production CompositePeerLauncher constructors default to none(). It is now a
real no-op with a test.
Left open deliberately: the implementer's bridge_ask about the credentialId
name timed out after 55s with no answer, because I was polling minutes apart.
It proceeded on its own judgment and chose well. The 55s ask window against an
async lead is a bridge problem, not a worker problem.
Added at merge review. none() held a clock frozen at 0 with a 1ns cooldown, so
a quarantine() call on it recorded a deadline that could never pass — the
credential would be locked out for the life of the daemon. Two production
CompositePeerLauncher constructors default to none(), so that failure would
have been silent and permanent.
The implementer documented the limitation honestly rather than hiding it, but
a stand-in named none() should not need the caveat. quarantine() is now a
no-op on that instance, with a test asserting it. An inert value must omit the
fact, never invent one.
A BACKEND_EXHAUSTED classification (stage A) now puts that profile's
credential into a BackendQuarantine for a configurable cooldown. A spawn
onto a quarantined profile is refused naming the credential and roughly
when it lifts; weighted/round-robin/fixed placement skip a quarantined
candidate; the quarantine lifts itself on the injected clock; and it is
visible on bridge_profiles.
Keyed by credential, not by profile name, via the new Profile.credentialId
(profiles sharing one credential quarantine together — e.g. two models on
one account) and effectiveCredentialId() (unset ⇒ quarantines alone,
today's behaviour unchanged). Fixed a related gap along the way: a reload
changing exhaustedPattern was silently reported "applied" even though it's
deferred — sameLaunchSettings() now catches it too.
Issue #57 criterion 5 asked for an explicit decision on this rather than
silence. It is closer to live than the issue assumed.
The DEFAULT parityOverlay is List.of(".env", ".envrc"), so it applies to every
profile, and neither path was gitignored. Since CB-576 a release preserves any
worktree that git status --porcelain calls dirty, and that deliberately counts
untracked files — the work lost in CB-576 was a file nobody had added. So one
.env at the repo root would make every COMPLETED release preserve its
worktree, and worktrees would accumulate with no error to notice.
Inert today only because neither file exists here. Both are now gitignored,
which they deserve on their own as environment files. The javadoc carries the
rule for the next overlay path: it must be gitignored, or tracked and
skip-worktree'd.
Verified by the lead: own build of the branch merged onto main — 717 tests,
BUILD SUCCESS, exit 0.
release() ran an unprotected sequence. hasUncommitted shells out to git status
and throws on a non-zero exit, which skipped both notifyReleased and
launcher.stop. The session was already out of the registry, so nothing retried
it: a live pane kept burning a fleet slot while absent from the roster, and any
send blocked on it was never resolved. reapIdle called release bare inside a
loop, so one such session aborted the whole pass and skipped every session
after it.
Now the dirty-check block is guarded and fails toward preserving the worktree —
'we could not tell' must not be treated as 'it is clean', because deleting on a
guess destroys work with no other copy. notifyReleased runs in a finally and
launcher.stop runs unconditionally, so the pane always stops. reapIdle catches
per session, matching the shape drainAll already used.
Checked while reviewing: notifyReleased cannot throw out of the finally — it
already guards each listener and only logs. The pane stop is genuinely
unconditional.
hasUncommitted shells out to git and can throw; release() now catches that
inside a try/finally so notifyReleased and launcher.stop always run, and
defaults to preserving the worktree on a throw (can't tell dirty vs clean,
so don't risk deleting unrecoverable work). reapIdle wraps each per-session
release in try/catch, matching drainAll, so one bad session no longer
skips the rest of the reaping pass.
Unit 2 was written as one block, but CB-571/576/580 and unit 2a have since
built parts of it. A symbol survey of main source plus the merge history puts
it at 1 of 21 criteria done, with three partials.
Criterion 15 says a normal COMPLETED release removes the worktree. Since
CB-576 that is false on purpose — a dirty worktree is preserved, because
deleting it destroys work that cannot be recovered. The criterion is wrong,
not the code.
The patterns are compiled once in Bridged.main from the startup config
snapshot, so adding one to a profile does nothing until a restart. Neither
the per-key docs nor the reload-class table said so.
Verified by the lead: own build of this branch merged onto main — 710 tests,
BUILD SUCCESS, exit 0.
Supersedes PR #54. A reviewer found that OpenCodeLauncher.SessionAwareHandle
wrapped the base WorkerHandle but never overrode charterReceipt(), so it
inherited the interface default of null while the real receipt sat on its
delegate — sol and terra would have shown no charterSource/charterSha256 on
the roster while Claude Code members showed both.
The fix is the root one, not the one-line override: PeerHandle.charterReceipt()
is no longer a default, so the compiler forces every implementation to answer.
This repo had shipped that same class of defect — a defaulted dependency that
compiles, passes tests, and quietly turns a feature off — eight times before
this one.
Verified by the lead: own build of the branch merged onto main — 704 tests,
BUILD SUCCESS, exit 0. No vendor wording in any Java source (grep clean); a
profile with no exhaustedPattern keeps today's completion-fallback path exactly.
Accepted the implementer's deviation from the brief. The brief asked for a
terminal HealthState BACKEND_EXHAUSTED. FleetHealth.decide() is a pure
classifier over HealthSnapshot, which carries only booleans, AgentStatus and
MemberSession.State — never pane text. FleetHealth's own javadoc already says
ERROR_ON_SCREEN 'is not decided yet because it needs a bounded pane detection
read'. A second declared-but-unproduced value would repeat a gap the class
already documents as a problem, so the signal was put where the evidence
actually lives: the completion scrape.
SessionAwareHandle wrapped the base's WorkerHandle but never overrode
charterReceipt(), so it silently inherited the interface default (null)
while the real receipt sat on its delegate. sol/terra never got a
charterSource/charterSha256 roster row.
Deletes the default so every PeerHandle must answer explicitly; the
compiler now catches this class of gap instead of a roster field
quietly going missing.
CB-575 already names the merged MCP-cancellation-filter change, so the
charter-receipt comments used the wrong number. Retag to CB-571, the number
this work was authored against.
Record a CharterReceipt (role, source, sha-256 digest, byte count) for every
launch, store it on the MemberSession, expose it in bridge_list and GET
/members, and log it at spawn as digest+role only. Redact the charter argv
argument in the legacy pane-placement spawn log so the charter text never
reaches the daemon log. The charter prose itself is never recorded.
2026-08-15 08:10:56 +02:00
365 changed files with 73209 additions and 19249 deletions
"description":"Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"name":"fleetd",
"description":"Tooling for orchestrating a fleet of delegated coding agents through the fleetd MCP gateway.",
"owner":{
"name":"LTMS"
},
"plugins":[
{
"name":"claude-bridge",
"name":"fleet",
"source":"./plugin",
"description":"Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version":"0.1.0",
"description":"Mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
description: Report the status of every fleet that shares one LavinMQ instance. Use for local daemon health, broker-wide fleet presence, and cross-host lead coordination checks.
---
# Status of every fleet on the shared LavinMQ instance
**The headline: always report what is missing.** This skill starts with the local fleet, then adds
broker-wide facts when its read-only credential exists. A missing fleet must appear as `unknown` or
`not reachable`, with the reason and the fix. Never leave it out.
The known topology has one LavinMQ instance on `10.10.20.13` (`fleet01`). AMQP uses port `5672`,
and the management API uses port `15672`. The Mac fleet owns vhost `/mac`. The fleet01 fleet owns
vhost `/fleet01`.
## 1. Protect credentials before any probe
**Hard rule — never print `LAVINMQ_URI`.** It is an AMQP URI with its password inline. It only
resolves in a login shell because `${SHARED_ENV}/tools/secrets.sh` supplies it. A non-login shell
can make every broker probe look empty.
- Never run `echo "$LAVINMQ_URI"`.
- Never put `${LAVINMQ_URI:-something}` in output. That form expands to the secret value when set.
- Parse the user, host, and password into shell or Python variables. Use them without printing them.
- Prefer `resolves` or `does not resolve` over any part of the value.
- Every command that can read `LAVINMQ_URI` must send all output through this redaction before it
reaches the report:
```bash
sed -E 's#://[^@]*@#://<redacted>@#g'
```
**The `g` flag is not optional.** Without it `sed` replaces only the first match on each line, so a
line carrying two URIs leaks the second one. `scripts/redeploy-fleetd.sh --check` prints lines like
that. Checked on 2026-08-27: without `g`, `amqp://u1:p1@h1/mac and http://u2:p2@h2:15672/api`
redacts the first pair and prints `u2:p2` in the clear.
Keep `pipefail` on when applying that filter. Otherwise the filter can hide a failed probe. Apply
the same no-print rule to the management password below, even though it is not in an AMQP URI.
## 2. Tier 1 — this fleet (always run)
Start here even when the broker tier is blocked. Work from the local fleetd checkout.
First run the read-only deployment check. It already checks the daemon process, deployed jar versus
the checkout `HEAD`, launchd state, and whether each configured token resolves in a login shell.
Do not copy those checks into new shell code. The script reads `LAVINMQ_URI`, so redact all output:
```bash
set -o pipefail
scripts/redeploy-fleetd.sh --check 2>&1\
| sed -E 's#://[^@]*@#://<redacted>@#g'
git rev-parse HEAD
```
Treat jar drift as a top-level warning. A merge is not a deployment. State the running jar result
as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into a match.
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'target/fleetd.jar'||true)"
if[ -z "$PIDS"];then
printf'%s\n''fleetd: not running'
else
for PID in $PIDS;do
ps -p "$PID" -o pid=,etime=,lstart=,command=
done
fi
```
Read the full health response. Keep the HTTP status because `503` means fleetd is running but herdr
is not reachable. Report both `herdr.version` and `herdr.protocol` when present:
| Mac (`/mac`) | PID + uptime | jar vs `HEAD` | health + version + protocol | `fleet_list` + `healthCoverage` | facts or blocked reason | inbox facts or open question | exact missing facts |
| fleet01 (`/fleet01`) | reachable/down/unknown | value or `not reachable` | value or `not reachable` | value or `not reachable` | facts or blocked reason | inbox facts plus coordinator-vhost question | exact missing facts and fix |
Add rows for unknown vhosts. End with three short sections: `Current warnings`, `Checks that were
blocked`, and `Operator action`. Until the management user exists, `Operator action` must say:
> Create a read-only LavinMQ management user with the `monitoring` tag and access to `/mac` and
> `/fleet01`. Put its user and password in `${SHARED_ENV}/tools/secrets.sh` as
> `LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD`.
description: Procedure for an outgoing lead to write the handover file that a fresh lead session inherits. Load this when your context is filling up and you are about to be replaced, whether you hand off by hand or fleetd does it for you. The file is the new lead's only inheritance — follow it exactly.
---
# Handover — write the file the next lead depends on
A lead session fills up its context and has to be replaced by a fresh one. The outgoing lead
writes a handover file, and the new session reads that file and carries on.
**There are two ways to hand off, and the file is the same either way.**
- **By hand.** You write the file, then tell the operator where it is. The operator starts the new
session and points it at the file. This always works.
- **With `fleet_handover`** (fleetd #480, merged 2026-09-11). You ask fleetd to do the swap: it
checks the file, clears your pane, and tells the fresh session to read it. This needs
`leadRollover:` in `fleetd.yaml`; without it every action answers a clean refusal naming
`NOT_CONFIGURED`, and you fall back to the manual path. Section 11 below is the procedure.
Nothing else in this skill changes between the two. Only who performs the swap changes.
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
through at the end of a session.
This skill is the procedure for writing it. Every rule below earned its place because a past
handover got it wrong.
## 1. Confirm you are the right session to write this
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
## 2. Every number needs a command, run in this turn
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
you write one, run the command that produces it — now, in this turn, against the live state.
Never take a number from:
- earlier in your own conversation — the state has moved since then,
- a peer lead's report — that is their measurement, not yours,
- your own memory of an earlier session.
Put the command, or its real output, next to the number. That lets the next lead re-run it and
check it still matches. If you cannot measure something yourself, say so instead of guessing:
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
## 3. Say what you measured and what you did not
Mark every claim as one of two things:
- **"I checked this myself, in the code or on this host, at `<time>`."**
- **"I did not check this myself; `<who>` reported it."**
Never present someone else's measurement as your own. This matters most for cross-host claims —
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
re-verify from here.
## 4. Record open decisions, and who owns them
List three things:
- what the operator actually asked for, in their own words where you have them,
- what is still unanswered,
- any question you decided yourself instead of asking, with your reason.
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
reopens a closed question or assumes the operator chose something they never did.
## 5. Record live hazards
List anything that will break if the next lead does the obvious thing next. This includes:
- unpushed commits or unmerged branches,
- a build, a spawn, or a redeploy still running,
- code merged to `main` but not yet redeployed to the live daemon,
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
"tricky."
## 6. Record what is explicitly not owed
List work that is finished, and work that another party has said they do not want touched. Name
who said so and when. Without this line, the next lead re-does closed work or reopens a question
a peer already declined to revisit.
## 7. Open the file with three re-measurement commands
The file's own first section must give the next lead three concrete commands to run before
acting on anything else in the file:
1. confirm role — for example `fleet_whoami`,
2. confirm the state of the working tree — for example `git status` and
`git rev-list origin/main..HEAD`,
3. read the live fleet — for example `fleet_list`.
Record what each command answered when you wrote the file, and tell the reader to run it again
rather than trust your answer. The point of this section is that the reader checks live state
before acting on any claim in the rest of the file, including yours.
## 8. Stamp the file with time and commit
Near the top of the file, write:
- the date and time you wrote it,
- the commit the tree was on (`git rev-parse HEAD`),
- whether the tree was clean (`git status`).
Without this, nobody can tell how old the file is, or which code it describes.
## 9. State plainly that the file goes stale fast
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
of one moment, not a live fact.
## 10. What to leave out
Do not include:
- narration of how the session felt, or how hard something was,
- anything the repo already records — code structure, git history, or a rule already written in
`CLAUDE.md`. Point at it instead of repeating it,
- advice that is only true for the session that is ending — a half-open terminal, a local
variable, a train of thought with no state behind it.
A handover file is a record of state and decisions. It is not a diary.
## 11. Using `fleet_handover` (only if `leadRollover:` is configured)
**Run the three steps in this order. The order is not a style choice — the wrong order is
refused.**
1.**`fleet_handover{action: "open", reason: "<why now>"}`.** It returns a `token` and the
`handoverPath` you must write to. Nothing has happened to your pane yet.
**Write to exactly that path, and do not resolve it yourself.** It is always absolute, even when
the operator configured a relative `handoverPath`: fleetd resolves a relative one against your
own workspace before it hands it to you. The daemon and your pane can run in different
directories, so a path you resolve yourself can point at a different file from the one the daemon
will check.
**Check that the path is ignored by git before you write to it (#491).** A relative
`handoverPath` resolves inside YOUR workspace, which is usually a repository — and usually not
the `fleetd` one, so an ignore rule added to `fleetd` does not protect it. Run
`grep -n handover <your workspace>/.gitignore`. No output means the file you are about to write
will show up as untracked content in that repo. The file is a snapshot of live state and must
never be committed, so tell the operator rather than committing it or silently editing their
`.gitignore`.
2.**Write the handover file at that path**, following sections 1–10 above.
3.**Ask the operator, then `fleet_handover{action: "confirm", token, operatorConfirmed: true}`.**
Why that order: `confirm` refuses with `HANDOVER_STALE` unless the file was modified **after** the
`open` request. That check stops a leftover file from an earlier session being accepted as this
one's handover. So writing the file first and then calling `open` — the obvious order — always
fails.
`{action: "cancel", token}` drops a pending request without rolling.
**Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
get another one.
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
from your connection, so you can only ever roll yourself.
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
defaults to `true` and this is the only thing standing between a judgement call and a wiped
session.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
the operator so before you confirm.** The recovery is the same either way: the file is already
written, so the operator starts a session and points it at the file. That is why you write the
file before you confirm, and never the other way round.
## Writing style
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
class, method, file, flag, and config key exactly as it appears in the code — replacing a
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
first time you use it.
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
brackets, colons, or slashes.
## Template
```markdown
# Handover — <fleet name> lead session, <date and time>
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
spawns, or restarts.
## 0. Do these three things first
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
## 1. What the operator asked for
<the live instructions, in their words where you have them; what is still open; any decision
you made yourself, and why>
## 2. Open decisions, and who owns them
<one line per decision: who owns it, what is unanswered>
## 3. Live hazards
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
## 4. What is not owed
<finished work, and work another party has declined; name who said so and when>
description: Defect-hunt procedure for a fleetd worker — sweep an assigned package for real bugs and report several ranked findings without fixing anything. Load this when the lead asks you to hunt or audit a scope rather than review one diff. Do NOT load `reviewer` for this; the two want different output.
---
# Hunter worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies.
**This skill is not `reviewer`.**`reviewer` judges one diff and reports the *single* most
important issue in about 90 words. A hunt sweeps a whole package and reports *several* findings
in a long structured form. Loading both gives you two contradictory output contracts, and the
usual result is a worker that writes a good report into its terminal and ends the turn without
sending it. Load exactly one.
## 0. Read this before you read code: how the report gets home
Your terminal reaches nobody. The lead sees **only** the text inside your `fleet_reply` call.
A long report is exactly the case where this goes wrong, so plan for it:
- **Write the report into the `fleet_reply` argument itself.** Do not compose it in your terminal
and then summarise it into the call.
- If the report is long, **send it anyway** — one `fleet_reply` with everything.
- If you end the turn without replying, the bridge scrapes your pane instead. That scrape carries
at most the last 4000 characters, and on a hunt it usually captures the tail of the lead's own
brief rather than your findings. The lead then has nothing and has to ask you again.
## 1. Change nothing
A hunt is read-only. Do not edit a production file, do not "quickly fix" what you find, and do
not run a formatter. You may run the build and tests to *check* a claim, and you should say so
when you did.
## 2. Read the whole scope first
Read every file in the assigned package before you judge any of it. A defect that a caller
elsewhere in the same package makes unreachable is not a defect, and you cannot know that from
one file.
Stay inside the scope. If a defect there depends on a class outside it, read that class to
confirm — but the defect itself must live in the scope you were given.
## 3. The bar — this matters more than the count
**Name the path into the bad state.** Say which caller, in which state, reaches it. A defect on
paper is not a reachable defect. If you cannot name that path, keep the finding but mark it
`unproven` and say exactly what you could not check. Do not drop it, and do not dress it up.
**Say which direction the harm goes.** Data loss, privilege escalation and silent wrong answers
are worth reporting even when the window is narrow. A finding whose worst outcome is a worse log
line is not worth a block.
Two workers once ran the same scope: the one that applied the direction-of-harm filter found ten
real defects, the one that did not found none. Fewer findings the lead can act on beat many the
lead has to triage.
## 4. Shapes that have produced real merged fixes here
Read for these first:
1.**A one-way gate.** A guard added after an incident closes only the direction that incident
came from. Do not only ask what closes the gate — ask **which states still open it**.
2.**A value read once, then used later to authorise something destructive**, after something
else has had a chance to change it.
3.**A failure downgraded to a value that looks like a legitimate result** — `-1`, `null`, an
empty list, `false` — which a caller then trusts.
4.**A lock held for one half of a read-modify-write and not the other**, or two collections
updated under different locks.
5.**A comment or javadoc stating an invariant the code no longer keeps.** Comments are
load-bearing in this repo; a stale one has already caused a bug.
## 5. What you cannot check, and must not claim you did
-`fleetd/fleetd.yaml` is gitignored and **absent from your worktree**. You cannot read it. If a
finding depends on live configuration, name the key and say you could not check it.
-`.mcp.json`, `opencode.json` and `.autoenv` in your worktree are neutralised stubs, not the
repo's real files.
- The `wiki/` submodule pointer is months old. Do not cite it.
Reporting a fact you took from the lead's brief as something you measured yourself is a false
report, even when the fact is correct. Say where each fact came from.
## 6. The report — what goes in `fleet_reply`
One block per finding, most severe first:
```
FINDING N — <one line>
file:line
Path in: <which caller, in which state, reaches this>
description: Implementer-role procedure for a bridged worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over bridged.
description: Implementer-role procedure for a fleetd worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over fleetd.
---
# Implementer worker — procedure
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
@@ -39,6 +39,19 @@ a worker made all 59 of its edits in the primary's tree and never noticed.
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
---
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
description: Reviewer-role procedure for a bridged worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over bridged.
description: Reviewer-role procedure for a fleetd worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over fleetd.
---
# Reviewer worker — procedure
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
**Wrong skill for a sweep.** This one reviews *one* diff or scope and reports the *single* most
important issue. If the lead asked you to hunt or audit a whole package for several defects, load
`hunter` instead and ignore this file — the two want different output, and following both is how a
worker ends its turn with a good report that never gets sent.
## 1. Read the whole scope before you judge
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
@@ -24,13 +29,13 @@ wrong.
covers what it was given.
- Do **not** edit files or run the build. You review; the owner acts.
## 3. Reach for `bridge_ask` only for a genuine fork
## 3. Reach for `fleet_ask` only for a genuine fork
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
fixes with different consequences — those are the lead's call, and guessing produces a
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
## 4. The finding — what goes in `bridge_reply`
## 4. The finding — what goes in `fleet_reply`
Report the **single most important** real issue in the scope, in these four lines, under
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `bridge_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
only `fleet_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `bridge_reply`,
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
@@ -41,7 +49,7 @@ and the sender silently receives nothing. Fail toward the recoverable error.
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2.**The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
side cannot see your screen. An answer that isn't in a `fleet_*` call is silently discarded.
3.**Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
@@ -50,17 +58,18 @@ and the sender silently receives nothing. Fail toward the recoverable error.
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
5.**Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
5.**Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
### Primary (lead) — run this on every task, in order
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
does this?" is **a worker**, not you. Reach for `fleet_send` before you reach for `Edit`. The steps
below are the procedure — run them in order, every task, not only the big ones.
0.**Know your role** — `bridge_whoami`, once per session, before anything else.
0.**Know your role** — `fleet_whoami`, once per session, before anything else.
1.**Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
ready to delegate — refine it or keep it.
@@ -70,20 +79,25 @@ below are the procedure — run them in order, every task, not only the big ones
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3.**Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
4.**Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
3.**Spawn every delegated unit first** — `fleet_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model, cost and
LIVENESS, not in tier, so the default is rarely what you want. The default is whatever the
daemon reports, and on a host where it sits on an exhausted or withdrawn credential every
unqualified spawn fails — sometimes loudly, sometimes as a member that spawns fine and then
produces nothing. `fleet_profiles` reports the default; check it once per session.
4.**Then send them all** — `fleet_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
of your context, your plan, or your screen.
5.**Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not**`sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6.**Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
5.**Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
with `fleet_send{turnId, content}` — **not**`sessionId`. A worker gone quiet is diagnosed with
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
6.**Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge tools it appears to have hold a blocked credential and fail, and a piped command
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
7.**Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -91,11 +105,23 @@ below are the procedure — run them in order, every task, not only the big ones
read it yourself.
8.**Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
your attention; they never substitute for it. Then merge, then `fleet_stop{paneId}`.
**If the forge refuses you the merge** — a protected branch, a token without the grant — the
adjudication is still yours. Read the diff, decide, and hand the operator a merge-ready queue
with the refusal quoted. Never report a PR as merged, and never call one "ready to merge"
without having read the diff yourself. A refusal is exactly when that shortcut is tempting,
because no action is left that forces you to look, and taking it turns this step into
forwarding a reviewer's verdict — which is delegating the merge by proxy, two lines above.
**Test a refusal; do not read it off a permissions field.** A protected branch holds its merge
rights separately from the repository permissions, so that field can say yes while the merge is
refused, and still say no after a grant makes it work. Probe instead, with a request that cannot
succeed on its merits, so a rejection can only mean the refusal. Treat a transport failure as a
third answer that proves nothing: a timeout, a DNS error or a bad URL is not a refusal, and
counting it as one makes you sure of something you never measured.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
prefer `wait:false` + `fleet_poll` for anything non-trivial: a blocking `fleet_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
@@ -103,21 +129,24 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Intent | Tool |
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `members` · one peer's state: `bridge_status{sessionId}` |
| Message a **peer lead**on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}`if they are on this host. **Not**`fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after*`open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
### Lead ↔ lead — coordinate, never delegate
`bridge_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
`fleet_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own members, and its own judgment. An empty
`members` array means no members are spawned; it says nothing about peers.
@@ -137,27 +166,42 @@ The traffic between leads is coordination and nothing else:
3.**Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
**N observations are N data points only if they differ in the axis you are trusting.** This cuts
both ways. N *failures* blamed on one cause are one data point when the cases share what you are
not varying. N *agreeing measurements* are also one data point when they share an instrument —
two hosts, two operators and the same formula is one formula, not two confirmations.
4.**Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
author is the worst reader of their own qualifier placement — measured here, one addendum carried
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
sentence goes false first, and would a reader reach the caveat before acting?"
Being messaged by a peer does not make you its worker: answer with `bridge_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
Being messaged by a peer does not make you its worker: answer the way you would open —
`fleet_send{coordId}` for another daemon, `fleet_send{sessionId}` on this host — and push back on
the substance if it is wrong. `fleet_reply` resolves a member's blocked `fleet_send`; a peer's
coord-id message is durable and non-blocking, so there is nothing for it to resolve. A peer that
simply complies has thrown away the reason there are two of you.
### Member (worker or architect) — the turn contract
1.**Load the playbook skill the lead named** before doing anything else.
2.**Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
3.**`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
3.**`fleet_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4.**End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
4.**End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `fleet_reply` ⇒ the sender gets nothing and the exchange stalls.
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
5.**Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
5.**Report honestly.** State only what you actually ran and its real output, including failures,
and never claim the result of a check you had no way to run. **Measure your own tools; do not
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
find there holds a deliberately blocked credential and fails every call, by design.
6.**Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
@@ -165,29 +209,101 @@ you.
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every spawned member, at launch, every peer kind — never a lead |
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `fleet_reply`* | every spawned member, at launch, every peer kind — never a lead |
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
| role agent definition files | role contract and per-job procedure | a member whose launcher binds its role to the matching file in its worktree |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
| the bridge's own docs | design detail, flows, error model | on demand |
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
`CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they* must obey belongs in the
charter, not here.
A rule belongs in **exactly one** layer — the outermost one that must obey it. A member without a
repo checkout still gets the launcher's reply charter, which is why that one rule stays there.
Peers that don't read `CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they*
must obey belongs in the charter, not here.
## Project addendum — claude-bridge (not part of the canonical block)
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Herdr socket tests (measured 2026-09-10).** In this repo, herdr is a subject under test. A
worker assigned to herdr code, and the lead, may let a test open the herdr socket directly in a
throwaway workspace that the test tears down. This only covers
`fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java`, and
`fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java`. It is not a
general licence. Using the herdr CLI or socket to move a real fleet session, pane, or peer stays
banned. That is the control plane that invariant 5 protects. Re-measure with
`grep -rl 'UnixSocketHerdrClient.connect()' fleetd/src/test/java --include='*.java'`. A non-empty
result means tests still open the socket and this note still applies. An empty result means nobody
does this any more; delete this section. Canonical invariant 5 restatement is tracked in #458 and
is not part of this change.
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
distinct backend errors (a non-exhaustion failure such as an HTTP 5xx) within 60 seconds — a
short, fixed 60s cooldown, not configurable per profile. Each check runs independently, so a
profile can show both at once. In the JSON: a cooling profile carries `credentialId` and
`coolingOffForSeconds`; a quarantined profile carries `quarantinedForSeconds`; a profile hit by
both carries all three fields, and either state alone already sets that profile's `free` to `0`.
A `fleet_spawn` naming a cooling-off profile is refused before it ever reaches the backend
adapter, with a message naming the credential and the remaining seconds ("cooling off after
repeated backend errors") — distinct wording from a quarantine refusal, so don't conflate the
two when reading a spawn failure.
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR),
`reviewer` (one diff → one structured finding) and `hunter` (sweep a package → several ranked
findings, change nothing). Name exactly one in every delegation. **`reviewer` and `hunter` are
not interchangeable** — `reviewer` caps the answer at one finding in about 90 words, so naming
it for a multi-finding sweep hands the worker two contradictory output contracts. That has
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace).
`port-to-opencode` (make an OpenCode session a participant in this workspace),
`fleets-status` (report every fleet that shares one LavinMQ instance),
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
fleetd #480).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
unmentioned by every instruction file, so it drifted and a later session planned it from scratch
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
worktree; and a plugin cannot deliver anything to members at all, because
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it,
so a member never reads the operator's plugin store. **The plugin is the lead-side surface;
member-facing assets travel in the worktree.**
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
committed copies would otherwise mount the primary's IDE and forge servers (fleetd #134). The
worktree's copy of each is a stub, **not** the repo's real file, so a worker that reads one and
reports what it found is reporting on the stub. The daemon logs a per-spawn summary, but the
worker cannot see that log. From inside its own worktree a worker — or a lead debugging one —
reads the list with `git config --worktree --get-all fleet.neutralizedConfig`, and the
consequence with `git config --worktree --get fleet.neutralizedConfigNote`. Never brief a worker
to edit one of these files: the edit cannot be committed, and it will not tell you so.
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, the turn-done
fallback and status gating — are diagrammed in `docs/MCP-Contract.md`. That page is now flows
only: its pre-build tool catalogue, parameter tables and REST paths were deleted rather than
corrected, because a hand-maintained second copy of the tool surface is what drifted for a month
while this line pointed every session at it (CB-609 / #114). **The live MCP schema is the tool
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
worker has no business restarting the daemon it is talking through, and stopping it kills the
worker's own channel mid-turn.
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
and its flags, the drain step, the operator's permission grant, and the five checks that have each
gone wrong here before. Do not hand-roll the steps from memory.
### The prompt is part of the product — update it with the code (mandatory)
@@ -200,15 +316,15 @@ Before you call any work done, check the row that matches what you touched:
| You changed… | Re-read and update… |
|---|---|
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| a `fleet_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
| `ConnectionIdentity` / how a caller is resolved | the `fleet_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__fleet__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `bridged.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
| **anything an operator can use, configure, or observe** — an MCP tool, a `fleetd.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
That last row is not bookkeeping. Chapters 1–10 answer *how is this built* and *why this way*;
none of them has a home for *what can it do and how do I turn it on*, so for twenty tickets a
@@ -219,7 +335,7 @@ is a Roadmap line. A change that touches none of the three earns no entry, and t
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
*Figure 1. Four file-disjoint foundations feed one composition unit.*
This split keeps `Fleetd.java` under one owner. It also keeps every other production file under one
unit in this plan.
## Evidence checked in the current branch
I read both issue pages in full. Each page reports zero comments.
| Evidence | What the code says now |
|---|---|
| `inject/CompletionResolver.java:229-237` | A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today. |
| `inject/CompletionResolver.java:260-276` and `:332-367` | #211 already added raw-screen classification when `lastAssistantBlock` is empty. The “dead-code question” in #201 is stale on this branch. |
| `inject/CompletionResolver.java:288-317` | Exhaustion wins before the hard-coded `API Error:` match. A backend error then goes through generic `fail(...)`. |
| `inject/CompletionResolver.java:311-316` | The code admits that the pattern is a heuristic. A member report which quotes an API error may match it. |
| `inject/CompletionResolver.java:449-467` | Startup coverage exists only for `exhaustedPattern`. |
| `inject/ExhaustedPatternLookup.java:13-25` | The current lookup and explicit `none()` value are a good shape for the new classifier seam. |
| `Fleetd.java:196-207` | One `BackendQuarantine` is shared by placement and the exhaustion sink. Its cooldown comes from `quarantineCooldownSeconds`. |
| `Fleetd.java:322-363` | Pattern compilation, target-to-profile lookup, and the live `ExhaustionSink` are composed in `Fleetd.main`. The sink on this branch still ends in `.ifPresent(...)`. This plan assumes #234 replaces that silent path. |
| `placement/BackendQuarantine.java:60-87` | A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock. |
| `member/CompositePeerLauncher.java:260-317` | Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off. |
| `member/CompositePeerLauncher.java:347-379` | Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”. |
| `placement/PlacementContext.java:10-22` and `PlacementPolicyUtil.java:14-83` | Automatic placement has only one transient exclusion set named `quarantined`. Reusing it would make outage errors say “backend exhausted”. |
| `mcp/FleetMcp.java:913-1025` | `fleet_list` sets `free: 0` and adds `credentialId` plus `quarantinedForSeconds` when quarantine is active. |
| `session/MemberSession.java:51-59` | The roster has `DONE` and generic `FAILED`, but no backend-error state or stored reason. |
| `session/SessionManager.java:695-773` | A normal boundary moves `BUSY` to `DONE`. A failure moves any non-released session to `FAILED`. The async completion resolver can race the `DONE` update. |
| `session/SessionManager.java:648-687` | `rosterView` reports the session state, but it reports no terminal reason. |
| `msg/MessageService.java:922-940` | CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage. |
| `msg/ReplyPushLoop.java:20-48` | Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns. |
| `msg/ReplyPushLoop.java:305-395` | Each push entry point resolves the owning lead through `PrimaryRegistry`. Missing ownership is logged and the durable or pending item remains the backstop. |
| `msg/ReplyPushLoop.java:496-547` | One tick builds one combined nudge. Pending items have separate reminder counts. |
| `health/FleetHealthMonitor.java:91-143` | Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications. |
| `health/FleetHealthMonitor.java:206-208` | `healthCoverage` means health enabled plus webhook configured. It does not describe lead-pane alerts. |
| `Fleetd.java:465-486` | Health stays `detection-only` without the webhook notification setting. |
I also read the related unit tests for `CompletionResolver`, `ReplyPushLoop`, `BackendQuarantine`,
`CompositePeerLauncher`, `PlacementPolicyUtil`, `SessionManager`, `MessageService`, and `FleetMcp`.
I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No
peer architect was named, so I did not exchange a design with one.
## Required behaviour
The policy should use these first values:
- Threshold: **2** classified backend errors.
- Window: **60 seconds**, measured from the first error to the second.
- Cool-off: **60 seconds**, starting when the threshold is reached.
- Correlation key: `credentialId`, never profile name and never error text.
- Incident rule: one active incident per credential. Errors during its cool-off do not extend it and
do not create more lead notices.
- Rearm rule: after cool-off ends, two fresh errors are needed for another incident.
Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window
fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without
turning a short backend fault into the default 1,800-second exhaustion quarantine.
A single classified error still fails its send and marks its member `backend_error`. It does not
cool a credential and does not send an outage notice. This is what “a single error changes nothing”
must mean at the credential level. It cannot mean that the failed member still looks successful.
```mermaid
sequenceDiagram
participant R1 as Resolver for member A
participant R2 as Resolver for member B
participant P as Outage policy
participant S as Spawn gate
participant N as Lead push loop
participant L as Lead pane
R1->>P: backend error for credential C
Note over P: Count 1, no cool-off
R2->>P: backend error for credential C within 60s
P->>P: Start one 60s incident
P->>S: Credential C is cooling off
P->>N: Queue one incident notice
N->>L: Inject when lead is idle, blocked, or done
L->>S: Request another spawn on credential C
S-->>L: Refuse and report remaining cool-off
```
*Figure 2. The second independent classification creates the fleet-level event.*
Against the 2026-09-01 case, the second failed member would start cool-off. `fleet_list` would show
zero free capacity and both members as `backend_error`. The push loop would inject one outage notice
even if the lead had not polled either ticket yet. The design reports the outage. It does not recover
uncommitted work from the members.
## Unit 1 — Typed backend-error classification
### Scope
Replace the direct hard-coded check inside `CompletionResolver` with a lookup and a sink. Keep the
public send result as a failed send. The typed internal event is the seam #227 consumes.
The lookup returns the pattern for a target. The sink receives the target, matched line, and full
failure reason. It fires only after `Rendezvous.resolveFailure(...)` wins for that exact captured
waiter. This copies the race rule already used by `ExhaustionSink`.
The classifier must run in all three current paths:
1. a normal non-empty assistant block;
2. the #211 raw scrape fallback;
3. a turn inside `MIN_TURN_NANOS`, before it becomes a generic too-fast failure.
In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic
failure or completion. A fast turn still fails when no configured pattern matches.
Keep `(?i)\bAPI Error\s*:` as a compatibility pattern for profiles without `errorPattern` until the
operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name
@@ -18,7 +18,7 @@ The design rests on three pieces (the shape this ticket proposes):
1. **Dedicated per-agent channels** — every agent has its own addressable inbox on the broker.
2. **A federated agent directory** — a global "who/where/status" lookup, assembled from per-host
presence, not a central database.
3. **A per-host gateway** — each host runs a `bridged` that owns its local herdr, registers/manages
3. **A per-host gateway** — each host runs a `fleetd` that owns its local herdr, registers/manages
its own sessions, and proxies messages to/from other hosts over the broker.
## 2. What is single-host today (the assumptions to break)
@@ -27,7 +27,7 @@ The design rests on three pieces (the shape this ticket proposes):
flowchart TB
subgraph host["Single host (today)"]
primary["primary<br/>(MCP client)"]
daemon["bridged daemon<br/>127.0.0.1:8765"]
daemon["fleetd daemon<br/>127.0.0.1:8765"]
reg["in-process registry<br/>keyed by PeerHandle.id() == paneId"]
herdr["herdr<br/>(local unix-socket PTY mux)"]
w1["worker pane wQ:p1"]
@@ -48,21 +48,21 @@ Three concrete bake-ins assume one host:
|---|---|---|
| **herdr is local** | `herdr/` unix socket `~/.config/herdr/herdr.sock` | You cannot drive another host's PTYs → each host **must** own its herdr. This is why a per-host gateway is mandatory. |
| **registry is in-process, keyed by `paneId`** | `session/SessionManager` | `paneId` (e.g. `wQ:p2B`) is a herdr-local coordinate — meaningless off-host. Routing needs a host-unique id. |
| **loopback, no authn** | `rest/BridgedApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
| **loopback, no authn** | `rest/FleetApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
## 3. Target architecture
```mermaid
flowchart TB
subgraph hostA["HOST A"]
gA["gateway = bridged A"]
gA["gateway = fleetd A"]
regA["local registry + herdr"]
primary["primary (MCP client)"]
gA --- regA
primary --- gA
end
subgraph hostB["HOST B"]
gB["gateway = bridged B"]
gB["gateway = fleetd B"]
regB["local registry + herdr"]
wb["worker panes"]
gB --- regB
@@ -96,13 +96,13 @@ host's terminals.*
delayed-message exchange (the remind/backoff loop for free) — the same reasons CB-307 picked it.
- **Federated agent directory** = a **soft-state, bridge-owned** roster, *not* a broker-stored
database. Per the persistence-boundary decision (bridged is soft-state; the broker owns *message*
database. Per the persistence-boundary decision (fleetd is soft-state; the broker owns *message*
durability, not *who/where/status*), each gateway announces its local agents `(globalId, host,
status, capabilities)` on a `roster.*` presence topic with periodic heartbeats. Every gateway
builds an eventually-consistent **union view** — literally CB-304's `rosterView`, federated. A
stale entry expires by missed heartbeat (reuses CB-303's idle/TTL thinking).
U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
U1 --> U2 --> U3
U4 --> U5
U3 --> U5
```
| # | Scope | Who | Why |
|---|---|---|---|
| U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own |
| U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `fleetd.yaml` is gitignored, so a worker cannot see or edit it |
| U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same |
| U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop |
| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon |
| U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained |
| U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria |
U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.
**Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles
need the same restart there is no reason to do two, but there is still a reason to *verify* in order:
`gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream.
`local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them
in that order separates the two causes instead of confusing them.
> **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN`
> (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login
> shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed.
---
## 6. Traps carried over from the wiki
Each of these cost someone real debugging time upstream. They apply to us.
1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us
that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's
report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first**
(`fleet_list` → `fleet_poll` anything wanted → `fleet_stop`), then rotate.
2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently
ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and
nothing else. Never reason as if the gateway authenticates.
3. **Exact model name.** A regex match routes fine but returns an **empty**`/v1/models` list while
completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway.
4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the
gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done.
5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty
means the model match is a regex. Do not conflate them when diagnosing.
---
## 7. Verification — what would prove this works
Merging config is not proving it. The checks, in order:
1. `fleet_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in
`fleet_reply`. This is the first proof of the token, the URL and the model name, and it risks
nothing the fleet depends on.
2. `fleet_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a
loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws,
for the same reason.
3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy
and the gateway.
4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is
the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does —
nothing else distinguishes a working passthrough from a translator quietly dropping thinking
deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it.
5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not
`legacy`. That is the whole point of taking our own token.
6. `fleet_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than
theoretical.
7. `fleet_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot
quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.
---
## 7.1 What the live run actually found — 2026-08-15
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The
migration was then **reverted**. This section is the result, so none of it has to be re-derived.
### The blocker
`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both
`bridged` today has **exactly one security control: the loopback bind**. Every other guarantee
`fleetd` today has **exactly one security control: the loopback bind**. Every other guarantee
rests on it.
The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone —
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `bridged` host); the
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `fleetd` host); the
token path is the split-host fallback."* The token path does not exist yet.
That leaves a seam that is **latent today and load-bearing the moment the bind moves**:
@@ -67,7 +67,7 @@ it will be disabled and the stage is wasted. So:
auth:
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
# mode: token # every non-worker caller must present a valid bearer token
# tokenEnv: BRIDGED_API_TOKEN # host env var holding the token; never the literal value
# tokenEnv: FLEETD_API_TOKEN # host env var holding the token; never the literal value
```
`mode: loopback-trust` is the current behaviour, named honestly and now *chosen* rather than
@@ -112,9 +112,9 @@ The roadmap says "systemd unit". **This host is macOS — there is no systemd on
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
`java -jar`. Ship **both**:
- `deploy/dev.ltms.bridged.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
- `deploy/dev.ltms.fleet.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
ordered start after herdr.
- `deploy/bridged.service` — systemd unit for the Linux gateways CB-308 introduces.
- `deploy/fleetd.service` — systemd unit for the Linux gateways CB-308 introduces.
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon
must **retry the socket rather than exit** — supervision ordering is a nicety, socket-retry is the
@@ -158,10 +158,10 @@ configured. It already leaks nothing but herdr's version and up/down.
### 3.1 There are TWO entry paths, and only one of them has identity today
The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not
literally true, and the difference is security-relevant.** `BridgeMcp` calls `MessageService` /
literally true, and the difference is security-relevant.** `FleetMcp` calls `MessageService` /
`SessionManager`*directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is
mounted as a raw servlet on Jetty's `ServletContextHandler`
(`BridgedApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
(`FleetApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
Javalin's `before` filters at all.
The current split is the mirror image of what you'd expect:
@@ -172,7 +172,7 @@ The current split is the mirror image of what you'd expect:
| REST routes | ❌ **none at all** — the session id is taken from the URL path and trusted | ❌ none |
So REST is the *more* exposed surface: `POST /sessions/{id}/reply` accepts any `{id}` from the
path, whereas the MCP `bridge_reply` derives the worker from the connection and refuses to read it
path, whereas the MCP `fleet_reply` derives the worker from the connection and refuses to read it
from an argument. Loopback-only bind is what makes this safe today.
**Therefore CB-505 must enforce on both paths against one shared resolver** — not at a single
@@ -188,15 +188,15 @@ Deliberately small; every one maps to a failure mode we have actually hit.
| Metric | Type | Why it exists |
|---|---|---|
| `bridged_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
| `bridged_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
| `bridged_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
| `bridged_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers |
| `fleet_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
| `fleet_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
| `fleet_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ sent\|exhausted — a rising `exhausted` means the primary is not draining. `sent` was called `delivered` until fleetd #365; it counts the herdr paste-and-submit call returning, never a confirmation the pane read it |
@@ -70,7 +72,7 @@ post-teardown invariant so a later regression becomes `DELEGATION_ORPHANED`.
### 2.2 Corrections made during design
The first state table missed `BUSY` in bridged plus `IDLE` or `DONE` in herdr. It would have found
The first state table missed `BUSY` in fleetd plus `IDLE` or `DONE` in herdr. It would have found
the real trace only through a late, weak stall timer. The final model adds
`TURN_BOUNDARY_LOST` as a strong disagreement state.
@@ -143,7 +145,7 @@ Idle may drive configured resource cleanup. It never opens an incident and never
| `TURN_BOUNDARY_LOST` | Same session turn stays `BUSY`; same accepted task stays open; two raw snapshots show `IDLE` or `DONE` | Strong disagreement. Strict reconciliation may repair it. |
| `ERROR_ON_SCREEN` | Suspicious non-working state survives grace; `detection` matches a tested adapter-specific fatal signature | Certain only for the matched signature. A bare word such as `Exception` is not enough. |
| `STALL_SUSPECTED` | Open turn is older than the configured threshold; two normalised `recent_unwrapped` digests are unchanged; no boundary or reply occurs | Not certain. A long valid API call can look the same. Lead decides. |
| `MUTE` | Turn resolves through completion fallback instead of `bridge_reply` | Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident. |
| `MUTE` | Turn resolves through completion fallback instead of `fleet_reply` | Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident. |
| `REPLY_STRANDED` | Typed reply or health message remains after owning-lead push reaches its cap | Collection failed. This does not explain whether the lead is busy, dead, or ignoring the nudge. |
| `DELEGATION_ORPHANED` | Target is gone, failed, or released, but one or more tasks remain `PENDING` after reconciliation grace | Certain bridge invariant failure. This is not an inbox-drain fault. |
| `WORK_PRODUCT_AT_RISK` | Provisioned branch has commits after its recorded base; member is `DONE`, `FAILED`, or preserved after release; no turn or inbox item remains; long-idle threshold passed | A warning, not proof of loss. Work may already have an open pull request or a squash merge. |
@@ -159,7 +161,7 @@ request or merge state. If committed work appears with `REPLY_STRANDED` or
| State | Exact evidence | Certainty and action |
|---|---|---|
| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the bridged-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. |
| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the fleetd-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. |
A failed fleet list alone is not a dead-member claim. A single `_not_found` with a healthy global
link is a target fault, not a control-link fault.
@@ -228,8 +230,8 @@ Only the lead may:
- choose how to use partial work in a worktree;
- restart herdr or change network, model, credentials, backend, or configuration.
Reports include literal safe tool calls such as `bridge_status(sessionId="...")`,
`bridge_poll(ticket="...")`, `bridge_list()`, and optional `bridge_stop(paneId="...")`. A judgement
Reports include literal safe tool calls such as `fleet_status(sessionId="...")`,
`fleet_poll(ticket="...")`, `fleet_list()`, and optional `fleet_stop(paneId="...")`. A judgement
state never presents stop as the only action.
### 5.3 Release causes and worktree safety
@@ -249,7 +251,7 @@ preserving cause.
Before abnormal release removes the live session, M4 writes an atomic manifest under the worktree
root. It records session identity, owner, role, profile, repository, path, branch, base commit,
release cause, release time, state, and pending task ids. `bridge_list.preservedWorktrees` loads these
release cause, release time, state, and pending task ids. `fleet_list.preservedWorktrees` loads these
manifests after restart. Stop output and WARN logs also name the path and cause. M4 never
auto-deletes a preserved worktree.
@@ -296,7 +298,7 @@ Compiled brakes apply even if config asks for more:
- at most two pane reads occur in one fleet tick;
- targets rotate fairly;
- only a normalised digest and optional clipped local excerpt are stored;
- no pane excerpt leaves bridged in a human webhook.
- no pane excerpt leaves fleetd in a human webhook.
Use `detection` for tested screen signatures. Use normalised `recent_unwrapped` only for progress
comparison.
@@ -328,7 +330,7 @@ The reader uses AMQP `content_type`, never body sniffing:
| 12 | CB-568 failure operation | via CB-580 | **Partial.** CB-580 (`0af902e`) routes `GONE` and `NEVER_READY` into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call. |
| 14 | `DELEGATION_ORPHANED` | 3 files | **Partial.** The health state exists. The teardown-invariant check that creates it, and the retry rule, do not. |
@@ -115,9 +115,9 @@ earlier `STATE.md` idea — a PR is reviewable, mergeable, and self-describing.*
## Infra facts (verified this session)
- **Remote:**`ssh://git@git.ltms.dev:2224/lms/claude-bridge.git` (gitea). Push is over **SSH** —
- **Remote:**`ssh://git@git.ltms.dev:2224/fleet/fleetd.git` (gitea). Push is over **SSH** —
a worker running as the same user with the same keys can `git push`**with no extra credential**.
- **gitea is NOT in the project `.mcp.json`** (only `jetbrains`, `intellij-index`, `bridged`). The
- **gitea is NOT in the project `.mcp.json`** (only `jetbrains`, `intellij-index`, `fleetd`). The
primary's gitea MCP comes from a global/user config, so **workers do not inherit it**. A worker
gets only the `bridge` MCP mounted (via `--mcp-config` launch flag).
- **No gitea CLI** (`tea`) installed; `glab` is present but is the GitLab CLI (wrong backend).
@@ -128,7 +128,7 @@ earlier `STATE.md` idea — a PR is reviewable, mergeable, and self-describing.*
| Option | Mechanism | Trade-off |
|---|---|---|
| **A. gitea REST + token** | Worker `curl`s `POST /api/v1/repos/lms/claude-bridge/pulls` with a scoped token injected by the daemon into the worker env | Minimal, no new server; token lives in the off-subscription worker's env (scope it tightly) |
| **A. gitea REST + token** | Worker `curl`s `POST /api/v1/repos/fleet/fleetd/pulls` with a scoped token injected by the daemon into the worker env | Minimal, no new server; token lives in the off-subscription worker's env (scope it tightly) |
| **B. mount gitea MCP into workers** | Add the gitea MCP to the worker's `--mcp-config` alongside `bridge` | Clean tool call, but the gitea MCP's own auth/token must be provisioned per worker; more moving parts |
| **C. install `tea` CLI** | Worker runs `tea pr create` with a token | Another dependency to install + configure; same token question as A |
@@ -140,7 +140,7 @@ and the token is a single scoped secret the daemon injects like it already injec
- Off-subscription workers already *could* push (SSH, same user). The **incremental grant is
PR-create**, i.e. a gitea API token.
- Scope the token **minimally**: the `lms/claude-bridge` repo, `write:repository` (create branch +
- Scope the token **minimally**: the `fleet/fleetd` repo, `write:repository` (create branch +
PR), **not** merge/admin/org. A leaked token can open PRs, not merge them — the primary/human is
still the merge gate.
- Inject via the daemon (env var, e.g. `GITEA_TOKEN`), never written to the worker's config dir —
@@ -153,11 +153,11 @@ and the token is a single scoped secret the daemon injects like it already injec
| **Config-parity overlay** | **CB-301 ext** — `SessionManager.acquire`, after `git worktree add` | symlink/copy the `parityOverlay` set into the worktree so the worker is a full peer; **this is what makes worktrees viable, not a dead-end** |
# Live bridge_ask — reverse rendezvous — 2026-07-16 16:30
# Live fleet_ask — reverse rendezvous — 2026-07-16 16:30
One worker paused its delegated turn to ask the primary, then resumed with the answer (profile `default`). Result: **`OK`**.
## Round-trip
1. **primary → worker** (delegation): the ask-forcing task.
2. **worker → primary** (`bridge_ask`, 6.6s): 'PICK A COLOR: red or blue?' — surfaced on the primary's blocked send as a `question` with `turnId=term_656bb47d2c42a9e#1`.
2. **worker → primary** (`fleet_ask`, 6.6s): 'PICK A COLOR: red or blue?' — surfaced on the primary's blocked send as a `question` with `turnId=term_656bb47d2c42a9e#1`.
3. **primary → worker** (answer on that turn): `blue`.
> **OK:** asked, resumed the same turn, and the reply reflected the primary's answer
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.