Compare commits

...

365 Commits

Author SHA1 Message Date
Dai Ha 2e349139e9 #568: add hunter member role
CI / shell-tests (pull_request) Successful in 6s
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 1m44s
2026-09-19 15:09:07 +07:00
Dai Ha a639969a9a CLAUDE.md: a blocked lead consults architects, not the operator
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 1m32s
CI / build (push) Successful in 1m42s
The operator set this rule on 2026-09-19: when a decision blocks a lead, it
consults one or more architect members, who are authorized to agree on one
decision and unblock. The operator is not asked. Escalation stays open only
for things outside the fleet's authority -- money, credentials, or a promise
made to someone else.

The paragraph also carries the reason the ticket record is mandatory rather
than optional. The operator's old notification channel was the block itself:
work stopped, so they found out. Taking the operator out of the loop removes
that signal with it, so the decision goes on the ticket, which reaches them
whether or not they are at a terminal when it is made.

The rule has a second half aimed at architects, which lives in
fleet.charters.architect and is applied per daemon -- filed as #591, because
a charter can never reach a lead and charters do not travel between hosts.
The missing notification event is #592.

Block verified byte-identical with wiki 7-Use-Cases.md at 1d9bd1b.
2026-09-19 14:57:04 +07:00
Dai Ha 17c3a69c57 docs: CB-591 gateway page had two claims that went stale
CI / shell-tests (push) Successful in 5s
CI / contract (push) Successful in 1m23s
CI / build (push) Failing after 1m37s
The gateway's chat model is served under the stable alias `acoder`, and the
model behind that alias changed on 2026-08-28 — it is Qwen3.8-27B now, not
DeepSeek-V4-Flash. The old name is still served, so nothing broke, but it
names a model this is not.

Two claims on the page were wrong as a result, and both were written as
current facts rather than dated measurements:

- `/v1/models` returns exactly `["deepseek-v4-flash"]` — it returns 6 ids
  now. This sat under a heading saying it needs no re-testing.
- the status banner said `local` and `gx` are both at `weight: 100` —
  `local` is at 0.

Measured today against the live gateway: /v1/models returns acoder,
qwen3.8-27b-nvfp4, deepseek-v4-flash and three embedding names; a completion
sent as `deepseek-v4-flash` comes back reporting `"model": "acoder"`, which
is the alias in plain sight. /v1/deployment reports generation
2026-08-28-qwen3.8-27b-nvfp4.

§2 and §3 are left alone. They are the August plan, and rewriting them would
destroy the record of the migration.

fleetd.yaml moved to `acoder` in the same change. It is not tracked here.
2026-09-13 07:01:15 +07:00
ltms 49a5875586 Merge #583: fleetd #582 — assert pending message-id cleanup at every publish cleanup site
CI / shell-tests (push) Successful in 9s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 2m1s
All ten assertions proven live: six by the implementer, the last four by the lead.
One contract build with four deleted removal lines produced exactly four named failures,
one per site, with the total unchanged at 1825.
2026-09-12 15:52:25 +02:00
Dai Ha 634d33b50b Merge worker/562-loop-health-wiring-test-99611c-5
CI / shell-tests (push) Successful in 9s
CI / contract (push) Successful in 1m17s
CI / build (push) Successful in 1m45s
2026-09-12 20:28:14 +07:00
Dai Ha 1db79bcaa9 Merge worker/581-completionresolver-cas-sites-0542b7-6 2026-09-12 20:28:14 +07:00
Dai Ha 4507bc5a70 Merge worker/571-attempted-outcome-5739f7-2 2026-09-12 20:28:14 +07:00
Dai Ha d7239ed23b fleetd #571: pin FleetMcp.formatReply's TIMED_OUT_UNCONFIRMED wording
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m42s
CORRECTION 5 on the ticket: mutating the new arm's message text to the
queued/working arm's text survived every existing test, because nothing
asserted the specific wording. This adds one test that asserts the
unconfirmed-delivery message and asserts it does NOT carry the
queued/working arm's retry invitation — the distinction #571 exists for.

No production code changes; formatReply's TIMED_OUT_UNCONFIRMED arm was
already correct.
2026-09-12 20:19:37 +07:00
Dai Ha dfeb9340b4 fleetd #581: cover completion CAS removals
CI / shell-tests (pull_request) Successful in 5s
CI / build (pull_request) Failing after 1m31s
CI / contract (pull_request) Successful in 1m35s
2026-09-12 20:19:20 +07:00
Dai Ha 1513d4f260 fleetd #562 follow-up: extract loopHealthSource factory, pin its wiring
CI / shell-tests (pull_request) Successful in 6s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Successful in 1m43s
PR #579's inline `new FleetMcp.LoopHealthSource(poller::health, ...)` in
Fleetd.main had nothing a test could call directly. Measured: replacing
poller::health with a constant () -> RUNNING compiled clean and left all
1771 tests green (see issue #562 comment "HOLD on PR #579").

Extracts the inline construction to a package-private Fleetd.loopHealthSource
factory, the same style as the sibling capacitySource/healthCoverageSource
factories, and adds FleetdLoopHealthSourceWiringTest with three separate
assertions: the statusPoller half, the sessionReaper half, and the
reaper == null branch (still STOPPED).
2026-09-12 20:16:23 +07:00
Dai Ha 4ca7d72303 #582: assert pending message-id cleanup
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m32s
2026-09-12 20:15:48 +07:00
Dai Ha c0545d003d fleetd #571: make FleetApp.writeReply's inner Outcome switch exhaustive, no default
CI / shell-tests (pull_request) Successful in 8s
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m48s
Ticket comments (17126, 17127) corrected the original acceptance criterion after this
unit was already in flight: a hand-listed grep for the enum's constant names goes stale
silently the moment a new constant lands, so the compiler must be the enumeration instead.
sendOutcomeLabel (MessageService.java) and formatReply (FleetMcp.java) were already
default-free switch expressions. The one gap was writeReply's inner "status" switch, which
had `default -> "done"` — the exact value that would have lied about TIMED_OUT_UNCONFIRMED.
Remove the default and list every Outcome constant explicitly; REPLIED, COMPLETED_UNREPLIED,
QUESTION and STALE_TURN get an arm too even though the outer switch always dispatches them
first, so the inner switch stays exhaustive on its own. The outer switch (a statement, not
an expression) keeps its own default — Java does not require exhaustiveness there regardless,
and "everything not terminal is a 202" is an intentional catch-all.

Verified with the proof the ticket asked for: added a scratch 11th Outcome constant after
deleting all default arms and confirmed all three switch-expression sites (and no test file)
fail to compile without an arm for it, one at a time, then removed the scratch constant.
2026-09-12 19:28:59 +07:00
Dai Ha 275ac0d251 fleetd #562: surface loop health
CI / shell-tests (pull_request) Successful in 12s
CI / contract (pull_request) Successful in 1m31s
CI / build (pull_request) Successful in 2m11s
2026-09-12 19:26:37 +07:00
Dai Ha c1e06c9e12 fleetd #571: add TIMED_OUT_UNCONFIRMED so an ATTEMPTED delivery is not reported as never-arriving
MessageService.send's TimeoutException branch collapsed Injector.Cancellation.ATTEMPTED
(fleetd #551 — the send call was made but its outcome is unknown) into
Outcome.TIMED_OUT_QUEUED, which promises the caller the message will never arrive. On this
route agent.prompt may already have pasted and submitted the text, so a caller's natural
recovery (resend) risks a double delivery.

Add Outcome.TIMED_OUT_UNCONFIRMED and route ATTEMPTED to it. Update the three readers found
by searching for the enum's constant names (not `Outcome.`, which misses FleetMcp's
unqualified `case REPLIED ->` switches and would false-positive on ConfigRef's unrelated
Outcome record):
 - MessageService.sendOutcomeLabel: add it to the "timeout" metric label group.
 - FleetMcp.formatReply: its own case, warning against a blind retry (distinct from the
   generic "retry or poll status" message the other timeouts get).
 - FleetApp.writeReply: its own "unconfirmed" status and detail text, so it no longer falls
   through the switch's default -> "done" arm, which would have reported "the delegation
   completed" for the one case where delivery is unconfirmed.
2026-09-12 19:21:10 +07:00
Dai Ha 204da67d66 Merge #576: fleetd #575 — one finally covers answer()'s STALE_TURN exit
CI / shell-tests (push) Successful in 16s
CI / build (push) Successful in 1m42s
CI / contract (push) Successful in 1m57s
Third instance of the #572 shape on this file: one invariant kept at N sites,
asserted at fewer than N. Here answer()'s inner try opened AFTER the Task
lookup/registration and the STALE_TURN early return, so that return was covered
only by a hand-rolled copy of the finally's cleanup pair. The fix widens the try
upward and deletes the copy, so every exit runs the one finally exactly once.
rendezvous.open(workerSession) correctly stays outside it — nothing to clean if
it never opened.

This fixed no live leak, and the code comment says so: neither
rendezvous.answerAsk nor clearAsyncQuestion(turnId, false) can throw, so nothing
ever left through the old gap uncovered. It is a structure fix, the ticket's own
fallback case.

Verified here, not taken from the worker's report:
- baseline on the branch: 1766 tests, 0 failures, from Maven and from an
  independent sum over target/surefire-reports/*.txt.
- my own mutation, located fresh: move the inner `try {` back down below the
  STALE_TURN return — the exact pre-fix structure, minus the hand-rolled pair.
  Result: 1766 run, exactly 1 failure, and it is the new test —
  MessageServiceTest.answerLosingTheRaceToAnAlreadyAnsweredAskStillReturnsStaleTurnAndCleansUpOnce:545
  "the forward waiter this answer() call opened must be closed after a STALE_TURN
  return". One failure, not a crowd: the new test is the only thing holding this
  path.
- the worker's own mutation removed the whole finally and took 8 other tests with
  it. That proves the finally runs; it does not prove the STALE_TURN path reaches
  it. Mine does.

The new test hook answerAskLapseRaceHookForTest follows the file's existing
askTimeoutRaceHookForTest convention.

Closes #575.
2026-09-12 19:05:00 +07:00
Dai Ha b091c51eee fleetd #575: widen answer()'s try so one finally covers its STALE_TURN exit
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 2m30s
The waiter cleanup pair (asyncTasksByWaiter.remove + rendezvous.close) was
duplicated: two sites sit in a finally, the third was hand-rolled inline
before answer()'s early STALE_TURN return, structurally outside any finally.

Both rendezvous.answerAsk and clearAsyncQuestion(turnId, false) are total
(cannot throw), so the gap never leaked in practice. But the duplicate was
untested: mutating it away left all 1765 tests green, while the two
finally-protected sites are each killed by 8-36 tests. Same shape as #572.

Fix: widen the try to wrap the Task registration and the STALE_TURN check,
so the single finally covers every exit and the hand-rolled copy is gone.
Added a race hook + regression test that deterministically reproduces the
'ask lapsed between the lookup and the unblock' case and proves the fix
still returns STALE_TURN and cleans up exactly once.
2026-09-12 18:56:02 +07:00
ltms 84d631b030 Merge #572: answer()'s session-lock release is pinned on all four exits
CI / shell-tests (push) Successful in 9s
CI / contract (push) Successful in 1m27s
CI / build (push) Successful in 1m44s
fleetd #572. MessageService releases its per-session lock in a finally at two sites. Removing the
send() one failed 22 tests and errored 1. Removing the answer() one left the whole suite green.
Skip that unlock and the thread holds the session lock forever, so every later send or answer to
that session blocks permanently — no exception, no log line.

Tests only. Against current main the diff is one file, MessageServiceTest.java, +178 lines.

Four new tests, one per exit of answer(): normal REPLIED, TIMED_OUT_WORKING, the ExecutionException
rethrow, and the InterruptedException rethrow. Each proves REACQUISITION rather than the return
value — a bounded follow-up send on the SAME session must not come back BUSY, and send reports BUSY
only when tryLock itself timed out, so a non-BUSY probe is specifically evidence the lock was free.

Lead verification, re-running rather than accepting the worker's numbers, on the branch merged with
main at 3f8c38f:
  baseline   -> exit 0, 1765 tests from Maven and an independent sum over 131 reports (1761 + 4)
  mutation L -> BUILD FAILURE, exactly 4 failures, all four new tests by name

THE TRAP THE WORKER CAUGHT ITSELF, which is the most valuable thing in this PR. Their first draft of
the TIMED_OUT_WORKING test called answer() inline on the test thread and then probed with send() on
that same thread. It FALSE-PASSED under the mutation: lock is a ReentrantLock, so the same thread
reenters for free whether or not unlock() ran. A same-thread probe proves reentrancy, not release.
They found it only because they actually ran the mutation instead of trusting a green baseline, moved
answer() onto a background thread, and re-proved the kill. The pitfall is now documented in the test's
own javadoc, with the measurement that exposed it.

NOTE FOR ANYONE RE-RUNNING THE TICKET'S RECIPE: the ticket records `sed '1218s|...'`, measured before
#551 merged. #551 edited MessageService.java, and the line is now 1231. Locate it fresh — a
line-anchored sed against a stale number mutates the wrong line and reports a meaningless green. The
anchor count is the check that catches this: `grep -Fxc '            lock.unlock();'` must read 2
pristine and 1 after.

Shape reported by the worker, not investigated and not fixed: the same two-line cleanup
`asyncTasksByWaiter.remove(reply); rendezvous.close(...)` appears at three sites — send()'s finally,
answer()'s inner finally, and answer()'s inline STALE_TURN early return, which is NOT inside a
finally. Same maintained-at-N-sites shape as this finding. Filed separately.
2026-09-12 13:30:22 +02:00
ltms 3f8c38fc54 Merge #567: pin LeadMailbox.inspect's probe-channel close
CI / shell-tests (push) Successful in 8s
CI / contract (push) Successful in 1m1s
CI / build (push) Successful in 1m49s
fleetd #567. LeadMailbox.inspect opens a probe channel and closes it in a finally. The production
code was already correct; nothing asserted it, so a future refactor could drop the close and leak an
AMQP channel per inspect() call with the suite green.

Test only. LeadMailbox.java is untouched — sha256 a2cd99be77345b7e... before and after.

The test asserts the CONSEQUENCE rather than the return value: it caps the connection at three
channels (LeadMailbox uses two, consume and publish), runs a successful inspect, then requires a
replacement channel. If the probe is left open, the broker has no channel number left and
createChannel() returns null. A test that only checked inspect()'s MailboxState would pass under the
mutation, which is the whole reason this hole existed.

Lead verification, re-running rather than accepting the worker's numbers, on the branch merged with
main at ed2fd66:
  mvn -o clean install -Pcontract  ->  exit 0, 1792 tests from Maven and from an independent sum
                                       over 137 surefire reports.
  1792 = 1761 (main) + 30 (contract-only) + 1 (new), which also confirms the default-profile count
  is untouched: the new test is in an @Tag("contract") class.

My own mutation, located fresh rather than assuming the reported line number: delete
LeadMailbox.java:338 `probe.close();`, anchor by grep -Fxc 1 -> 0. Result under -Pcontract: exactly
1 failure, inspectClosesItsSuccessfulProbeChannel:226, "the replacement channel was null". Restored
to a2cd99be77345b7e..., git status --short empty.

THE PROFILE TRAP THIS TICKET EXISTS BECAUSE OF. An earlier sweep worker instrumented this line, saw
zero hits under the DEFAULT profile, and concluded "no test executes this line". That was false:
fleetd/pom.xml:264 sets excludedGroups=contract, so the covering tests were excluded from the run,
not absent. A surviving mutation has THREE causes — never executes, executes with nothing asserted,
or the covering tests were excluded from the profile — and only naming the profile tells them apart.
The true finding was covered-but-unasserted.

Known fragility, recorded rather than fixed: the test's channel cap of three assumes LeadMailbox
holds exactly two channels. If it ever holds more, this test fails loudly, which is fine. If it ever
holds fewer, a leak would no longer exhaust the cap and the test would go vacuous silently. Worth
re-checking if LeadMailbox's channel usage changes.

Not covered, and stated rather than faked: the defensive catch (RuntimeException) around the passive
declare, which needs a connection dying between createChannel() and the declare landing. The
method's own javadoc already admits that branch is unproven; the worker did not invent a test for it.
2026-09-12 13:26:30 +02:00
Dai Ha a4dbc8f8b7 fleetd #572: pin answer()'s session-lock release across all four exits
CI / shell-tests (pull_request) Successful in 7s
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m10s
MessageService.answer() releases its per-session lock in an outer finally
(MessageService.java:1218) that mutation testing showed was covered but
unasserted: removing that line left all 1750 existing tests green, because
every existing test on this path checks answer()'s return value, never that
the lock it took is actually reacquirable afterward. If it leaked, a session
would be wedged forever with no exception and no log line.

Adds four tests, one per exit of answer() (normal REPLIED reply,
TIMED_OUT_WORKING, ExecutionException rethrow, InterruptedException
rethrow), each proving the lock is reacquirable via a bounded (300ms)
follow-up send on the same session rather than merely checking answer()'s
own outcome. No production change.
2026-09-12 18:23:54 +07:00
ltms ed2fd6646a Merge #561: the completion/session listener fan-out survives either half throwing, and both sites are pinned
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 50s
CI / build (push) Successful in 2m7s
fleetd #561. Fleetd composed two TurnListener halves as `completion.X(); sessions.X();`, so a throw
from the first half skipped the second. The anonymous class is now a package-private factory,
Fleetd.turnListener(completion, sessions), built on two helpers that always attempt both halves and
rethrow whatever escaped — a second failure attached with addSuppressed rather than dropped, so it
still reaches StatusPoller's catch (Throwable).

The wiring at Fleetd.java:498 calls that factory, so the seam under test is the real caller.

Lead verification, on a merged tree, re-running the checks rather than accepting the worker's:
exit 0, 1761 tests from Maven and from an independent sum over 131 surefire reports.

The interesting part is what the first round MISSED. Two helpers maintain one invariant — "the
second half always runs" — and the first round's five tests asserted it at only one site. Measured:

  bothMustRun                      reverted to the pre-fix bug -> 1 named failure   (pinned)
  bothMustRunKeepingSecondResult   the SAME bug                -> 1755/1755 GREEN   (unpinned)
  failure.addSuppressed(t) deleted                             -> 1755/1755 GREEN   (unpinned)

Both survivors are now killed by new tests, re-verified by the lead after the fix:
sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronouslyForPostAction and
bothFailuresEscapeWhenBothHalvesThrowDistinctExceptions, each failing alone under its own mutation.

The rule this cost us, worth carrying: COUNT ASSERTIONS PER SITE, NOT PER INVARIANT. The total being
non-zero is what hides a zero at one site, and extracting a shared helper makes it worse rather than
better — it does not reduce the number of sites, only how many are visible. Credit to the fleet01
lead, who predicted this shape before an instance was found.

onDelivered stays deliberately unguarded. Its comment now gives the real reason —
CompletionResolver.captureBaseline already catches RuntimeException around its scrape and fails open,
so that half does not realistically throw — instead of the previous reason, which was true but about
registration rather than about this pair. A correct conclusion resting on a wrong premise reads
exactly like a verified one.
2026-09-12 13:20:56 +02:00
Dai Ha 5441a2b321 fleetd #567: assert inspect closes probe channel
CI / shell-tests (pull_request) Successful in 9s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m23s
2026-09-12 18:19:43 +07:00
ltms 384867dfa3 Merge #551: ATTEMPTED is its own cancellation answer, and the javadoc stops claiming a timed-out send never arrived
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m44s
fleetd #551. The injector polls a queued entry off the queue and marks it ATTEMPTED BEFORE the
irreversible AgentControl#send call, not after. So a Throwable escaping that call can never leave
the entry QUEUED at the head (the #546 re-send hazard) and can never be recorded as a confident
NOT_DELIVERED for text that may already be in the pane.

Cancellation.ATTEMPTED is added as a third answer. NOT_DELIVERED stays reserved for confirmed
absence: the readiness grace expiring, drop(), or a herdr *_not_found error, which the rest of this
codebase already reads as definitely-absent rather than inconclusive.

Also fixes three javadoc/comment sites that claimed a timed-out send definitely did not arrive:
TIMED_OUT_QUEUED, hasQueuedDelivery, the queuedDeliveries field, and the comment in send()'s timeout
branch. Two claims were wrong, not merely stale: "on every route it will not arrive later" is false
on the ATTEMPTED route, and "may already hold a partial paste" understates it — agent.prompt pastes
AND SUBMITS in one call, so the target may hold a complete, running turn.

Verified by the lead on a merged tree: mvn -o clean install from fleetd/, exit 0, 1754 tests from
Maven and from an independent sum over 130 surefire reports. Mutation: folding ATTEMPTED back into
NOT_DELIVERED in cancellationOf gives 3 red, each naming the property
(anErrorFromSendRemovesTheMessageAndMarksItAttempted:740,
aHerdrExceptionFromSendStillSurfacesButNowReportsAttempted:781,
aHerdrExceptionAfterThePasteIsNeverRecordedAsConfidentlyNotDelivered:806).

The final round is comment-only, proven mechanically rather than by reading: stripping every comment
from MessageService.java before and after and collapsing whitespace gives byte-identical code.
2026-09-12 13:19:20 +02:00
Dai Ha f40c19ecf0 fleetd #551 shape sweep: fix stale ATTEMPTED-route javadoc/comments in MessageService
CI / shell-tests (pull_request) Successful in 6s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m26s
hasQueuedDelivery's javadoc, the queuedDeliveries field javadoc, and a code
comment in send()'s timeout path all still made two claims that ATTEMPTED
(fleetd #551) falsifies: a blanket "the message will not arrive later" across
every route, and "the terminal may already hold a partial paste" — but
agent.prompt pastes AND submits in one call, so the target may hold a
complete, already-submitted turn. Each now names the three Injector.Cancellation
routes (CANCELLED, NOT_DELIVERED, ATTEMPTED) and says plainly that only the
first two establish the message will not arrive later.

Comment/javadoc only. No behaviour change: Injector.java is untouched
(sha256 0c689b6cf36275c0da45497a74b2bd4f5a3d80c4dbda77d46670c66004e53b69) and
no test was added. mvn -o clean install: BUILD SUCCESS, Tests run: 1754,
Failures: 0, Errors: 0 (unchanged from before this commit).
2026-09-12 18:13:06 +07:00
Dai Ha 034e17bb32 fleetd #561 follow-up: pin the session half of bothMustRunKeepingSecondResult
CI / shell-tests (pull_request) Successful in 7s
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m41s
Two helpers maintain one invariant (the second callback half always runs,
even when the first throws): bothMustRun and bothMustRunKeepingSecondResult.
Only bothMustRun's "session half still runs" direction was asserted
(sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronously, via
onTurnComplete). bothMustRunKeepingSecondResult — the helper
onTurnCompleteWithPostAction uses — could be reverted to the pre-#561
broken shape and the suite stayed green.

Adds two tests to FleetdTurnListenerCompositionTest:
- sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronouslyForPostAction:
  mirrors the existing onTurnComplete case for onTurnCompleteWithPostAction/
  bothMustRunKeepingSecondResult.
- bothFailuresEscapeWhenBothHalvesThrowDistinctExceptions: proves a second,
  distinct failure from the session half is preserved via addSuppressed
  rather than silently dropped when both halves of bothMustRun throw.

Also rewords the onDelivered comment in Fleetd.turnListener: it previously
said this pair is safe because registration survives a throw via #556's
Injector wiring, which is true but is not why THIS pair is unguarded.
CompletionResolver.captureBaseline already catches RuntimeException around
its scrape read and fails open, so completion.onDelivered does not
realistically throw. Comment text only, no logic change.
2026-09-12 18:10:31 +07:00
Dai Ha 68b428c484 fleetd #551 rework (comment 17058): fix TIMED_OUT_QUEUED javadoc for ATTEMPTED
CI / shell-tests (pull_request) Successful in 13s
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 2m21s
Javadoc-only change. #551 added Injector.Cancellation.ATTEMPTED, which
MessageService.send's timeout path folds into Outcome.TIMED_OUT_QUEUED
alongside CANCELLED and NOT_DELIVERED (the collapse itself is unchanged
behaviour and is being tracked as a separate follow-up ticket).

TIMED_OUT_QUEUED's javadoc — landed by #513 to state the routes that
reach it — named only two routes and said "on every route it will not
arrive later", with a "may already hold a partial paste" caveat. Both
claims are now stale: ATTEMPTED is a third route, and because
agent.prompt pastes AND submits in one call, that route may mean the
target holds a complete, already-submitted turn and is working on it
right now.

Names all three routes, says which one is uncertain, and drops the
now-false blanket claim. No behaviour change.
2026-09-12 17:51:26 +07:00
Dai Ha e20ccab1eb fleetd #561: harden the completion/session TurnListener fan-out
CI / shell-tests (pull_request) Successful in 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m56s
Fleetd's turnListener composition had four callbacks (onTurnComplete,
onTurnCompleteWithPostAction, and both onTurnFailed overloads) built from two
bare, unguarded statements each. onDelivered's registration was already fixed
structurally by #556; these four had the identical fragility and were still
untested: nothing enforced that the completion resolver's half ran before the
session half beyond call order in the source, so a future reorder (or a
throwing session listener sequenced first) could silently skip the
completion resolver's effect and strand a caller for its full timeout.

Extracted the composition to a package-private static factory,
Fleetd.turnListener(completion, sessions), and hardened it with
bothMustRun/bothMustRunKeepingSecondResult: both callback halves are always
attempted regardless of whether the other throws, and whatever escapes is
rethrown afterward (never swallowed) so it still reaches StatusPoller's
catch (Throwable) and logs at ERROR.

FleetdTurnListenerCompositionTest builds this real composition from a real
CompletionResolver and a throwing fake sessions half, and asserts the
completion resolver's effect (the waiter resolving) survives the session
half throwing, for all four callbacks, plus a mirror case showing the
session half still runs when the completion half throws first.

onTurnCompleteWithPostAction keeps completion-before-session as a functional
requirement (resolveBeforePostAction must run before the context-reset
housekeeping can erase the pane), not just fault tolerance, so it is not
reorder-symmetric like the other three — documented in Fleetd.turnListener's
javadoc.
2026-09-12 17:47:53 +07:00
Dai Ha d83821bbce fleetd #551: record the delivery attempt before the irreversible send
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Successful in 2m23s
The Injector wrote the delivery outcome AFTER calling AgentControl.send(),
so a HerdrException thrown from the response half of that call (herdr
already replied, or may have) was recorded as a confident NOT_DELIVERED
for text that may already be sitting in the worker's pane.

Poll the queue entry and mark it Pending.State.ATTEMPTED before send() is
called, not after. On success it is upgraded to DELIVERED; on an ordinary
failure it stays ATTEMPTED (honest uncertainty), except a herdr
*_not_found error, which the rest of this codebase already treats as a
confirmed absence and which now still writes NOT_DELIVERED.

cancellationOf gets a matching third answer (Cancellation.ATTEMPTED)
instead of folding the new state into NOT_DELIVERED, so a caller that
cancels an already-attempted delivery is told the truth too.
2026-09-12 17:43:21 +07:00
ltms ba2f4d16f8 Merge #555: the redeploy main flow is lifted into tested predicates, and the guard now catches functions below the SOURCED line
CI / shell-tests (push) Successful in 9s
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m52s
fleetd #555. Eight main-flow decisions in scripts/redeploy-fleetd.sh move into
predicate and dispatch functions the suite can source and test. The guard test
test_no_untested_main_flow_conditionals stops new bare conditionals reappearing.

The rework closes a hole the lead found (comment 17012): a conditional wrapped in
a function defined BELOW the SOURCED guard was invisible to the guard, and such a
function can never be sourced, so it can never be tested. The guard now fails on
any function definition after the guard line, on its own.

Verified by the lead on a tree with main merged in, redeploy-fleetd.sh at its
pristine sha 4ffacc5185807d39720a3484d85b922413806eb5347318265bd8897dfd61e8d9:
  bash 5.3.9      -> exit 0, 0 lines matching ^FAIL:
  /bin/bash 3.2.57 -> exit 0, 0 lines matching ^FAIL:

Three mutations, each restored to the pristine sha afterwards:
  a bare conditional appended to the main flow      -> exit 1, reported by line
  the same conditional wrapped in a function below
    the SOURCED guard (the found hole)              -> exit 1, "function defined
    after the SOURCED guard (line 1038) - it cannot be sourced, so it cannot be tested"
  the new FUNC emission deleted from the guard      -> that same case returns to
    exit 0, so the new assertion is what catches it. Anchor count 1 -> 0,
    test-redeploy-fleetd.sh restored to sha 0d713a3092a0c0ea8c05663ffa0cb1595e21b9d73872e662eb99b5035fd38151.
2026-09-12 12:19:47 +02:00
ltms db4c98ac60 Merge #556: the Injector owns turn registration, and the #553 backstop is pinned too
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 51s
CI / build (push) Successful in 2m12s
fleetd #556. Registration moves off the TurnListener fan-out onto its own narrow
TurnRegistrar seam, wired directly to CompletionResolver::register in Fleetd.java,
so it survives any listener throwing regardless of call order.

Verified by the lead on a tree with main (f4f5f31) merged in:
  mvn -o clean install exit 0; 1750 tests, agreed by Maven's own summary and an
  independent sum over 130 surefire report files.

Both registrar.register call sites are independently pinned:
  Injector.java:703 (the #553 finally backstop) removed -> exactly 1 failure, the
    new aRuntimeExceptionFromOnTurnCompleteStillLeavesTheNextDeliveryRegisteredOnTheRecoveryPath
  Injector.java:660 (the ordinary path) removed -> exactly 2 failures, the two
    original tests; the backstop test stays green
Anchor count 1 -> 0 on each, sha restored to
8fcb698afccc254b0c99d3a4bf9e960c0e7c85e542024e870c1e815dce6d62a9 both times.

The :703 line survived the full suite before this rework while being executed
twice — covered but unasserted. That gap is what the added test closes.
2026-09-12 12:18:15 +02:00
Dai Ha 8f80d267a0 fleetd #555 rework: catch function definitions after the SOURCED guard
CI / shell-tests (pull_request) Successful in 9s
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 3m3s
Comment 17012 on #555 found a hole in test_no_untested_main_flow_conditionals:
the guard's function-body detection treats anything inside a function as
"fine, out of scope for this scan" — but a function DEFINED after the
SOURCED guard line can never be reached by sourcing this script (sourcing
stops before the main flow runs), so its body is untestable by construction
while still reading to the guard as safely inside a function.

mainflow_bare_conditionals now also emits a FUNC record for every function
opened after the guard line (reusing the same open-brace detection already
used for depth tracking), and test_no_untested_main_flow_conditionals treats
any such record as a violation on its own, independent of what the function's
body contains or whether the allowlist would otherwise excuse a bare
conditional inside it.

Proof (redeploy-fleetd.sh restored to 4ffacc5185807d39720a3484d85b922413806eb5347318265bd8897dfd61e8d9
after each):

- CONTROL — a bare conditional appended to the main flow is still caught:
  EXIT=1, "found 1 untested main-flow if/elif/case line(s) ... line 1341:
  if [ "$MY_CONTROL_BARE" = 1 ]; then :; fi"
- CANDIDATE — the same conditional wrapped in a function defined after the
  boundary, previously invisible (EXIT=0), is now caught: EXIT=1, "line 1341:
  function defined after the SOURCED guard (line 1038) — it cannot be
  sourced, so it cannot be tested: newfunc_below_the_boundary() {"

Full suite re-run green on both bash 5.3.9 and /bin/bash 3.2.57 (macOS
system bash): exit 0, 0 FAIL lines, reached the final PASS line, on both.

No change to redeploy-fleetd.sh; the 8 lifted decisions, their mutation
proofs, and the allowlist all stand as before.
2026-09-12 17:11:11 +07:00
Dai Ha 738d34a609 fleetd #556 rework: pin registration on the #553 finally backstop path
CI / shell-tests (pull_request) Successful in 7s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m13s
Comment 17009: there are TWO registrar.register(target, sent.token())
call sites in Injector's delivery method — the ordinary path inside
`if (sent != null)`, and the fleetd #553 finally backstop, reached only
when an earlier block throws before the ordinary path ever runs. The
lead's mutation on Injector.java:703 (the backstop call) survived the
full suite: the existing #553 regression test for this exact scenario
(aRuntimeExceptionFromOnTurnCompleteStillCompletesTheNextDelivery)
asserts only that the delivered future completes, never that the turn
is registered with CompletionResolver — so a redesign that dropped
registration from the backstop would reopen this ticket's own defect on
precisely the path #553 exists for, with every existing test green.

Adds aRuntimeExceptionFromOnTurnCompleteStillLeavesTheNextDeliveryRegisteredOnTheRecoveryPath:
drives the same construction as the existing #553 test (onTurnComplete
throws for a previous turn, forcing the next turn's delivery down the
finally backstop) and additionally asserts the new turn is registered
with CompletionResolver and carries the correct waiter — the same
assertion the ordinary-path test makes, now made on the recovery path.

Proven by mutation: removing Injector.java:703 alone (exact-line anchor
1 -> 0) turns the new test red with its own assertion message; restored
and confirmed byte-identical (sha256 8fcb698afccc254b0c99d3a4bf9e960c0e7c85e542024e870c1e815dce6d62a9,
matching the pre-mutation tree); re-run green as a control. Full suite
after restore: 1750 tests, 0 failures, 0 errors, 0 skipped (Maven's own
summary and an independent sum over surefire-reports/*.txt agree),
BUILD SUCCESS.
2026-09-12 17:08:44 +07:00
Dai Ha a46e4058ac fleetd #556: make turn registration structural, independent of any TurnListener
CI / shell-tests (pull_request) Successful in 4s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m29s
The Injector owns the invariant "every delivered turn has a registered
waiter," but before this the only thing that satisfied it was
CompletionResolver.captureBaseline, called from inside a TurnListener
callback wired in Fleetd.java. Any TurnListener that throws (from
onDelivered or elsewhere) could break the invariant with no way for the
Injector to detect it. #553 only made the one reachable listener behave
via a try/finally backstop; it did not remove this structural dependency.

Add a narrow TurnRegistrar functional interface, decoupled from
TurnListener, whose only job is registering a delivered turn's waiter.
CompletionResolver now implements it via a new register() method
(extracted from captureBaseline's registration half; captureBaseline
keeps its own full body unchanged, so existing direct callers/tests are
untouched). Injector gets an explicit registrar field/constructor family
(auto-derived from the TurnListener via instanceof where that still
works, explicit where Fleetd's anonymous fan-out listener can't
implement two interfaces at once) and calls registrar.register(...)
directly and unconditionally in both the ordinary delivery path and the
#553 finally backstop, before turnListener.onDelivered(...) — so
registration no longer depends on that notification callback succeeding.
Fleetd.java wires completion::register explicitly as the registrar,
bypassing the turnListener fan-out for registration purposes.

CB-116 ordering (onTurnComplete reads the PREVIOUS turn's inFlight entry
before the new turn's registrar.register() runs) and the two-arg
inFlight.remove(target, turn) vs one-arg distinction on the completion
path are both preserved unchanged.

#561's order-dependent test (asserting Fleetd.java's
completion.onDelivered -> sessions.onDelivered call order) does not
exist anywhere in this repo at this branch point — nothing to delete.

Adds 5 tests: a TurnListener that throws from every callback still
leaves the delivered turn registered and resolvable; the new turn's
registration still runs after the previous turn's completion is read
(CB-116 guard, pinned as a call-order assertion); and three tests naming
the two-arg-remove invariant directly (a superseded turn's terminal
handling must not evict its successor's registration) across resolve()'s
plain-completion branch, its echoed-noReportMessage sub-path, and
fail().
2026-09-12 16:50:09 +07:00
ltms f4f5f3106e Merge #513: TIMED_OUT_QUEUED has four routes, and the javadoc now says which
CI / shell-tests (push) Successful in 4s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m51s
The old javadoc said a TIMED_OUT_QUEUED message was "still sitting in the injector's
per-target queue". It is not — send() has already given up on it and it will never
arrive. That wrong claim told an operator to wait for a message that was never coming.

The first fix replaced it with a narrower wrong claim: that Injector.cancel() cancelled
the entry and "the target never saw a word of it". That describes one of four routes.

TIMED_OUT_QUEUED is returned whenever injector.cancel() returns anything but DELIVERED:

  CANCELLED      the Pending was still queued and this call removed it
  NOT_DELIVERED  Injector.java:400 - the herdr agent.prompt call threw
  NOT_DELIVERED  Injector.java:419 - readiness grace expired, never attempted
  NOT_DELIVERED  Injector.java:691 - drop(), the target is gone

On the last three, cancel() cancels nothing: it reads a state another path already set
(cancellationOf, Injector.java:288-290). And on the herdr-threw route, agent.prompt
pastes and submits in one call, so a throw does not prove the pane stayed clean - an
operator told "never saw a word" will not go and look at the one place the evidence is.

The javadoc now states the two facts that hold on every route - the message will not
arrive later, and it is not in any queue - and attaches "the target saw nothing" only to
the CANCELLED case. Five blocks: the enum constant, queuedDeliveries, hasQueuedDelivery,
hasOrphanedDelegation, and the inline comment in the TimeoutException branch that seeded
the wording.

Comment-only; no logic changed.

Verified on a tree merged with main (fast-forward to cb64bc8):
  mvn install exit 0
  1744 tests, 0 failures, 0 errors - Maven's own summary and an independent sum over
  130 surefire report files agree
  no unresolved javadoc reference on the three new links, with a positive control
  showing javadoc did analyse MessageService.java

Closes #513.
2026-09-12 11:48:54 +02:00
Dai Ha 8d79d229ff fleetd #555: lift 8 main-flow decisions into tested predicate/dispatch functions
CI / shell-tests (pull_request) Successful in 11s
CI / contract (pull_request) Successful in 1m31s
CI / build (pull_request) Successful in 1m46s
redeploy-fleetd.sh's main flow had 8 bare if/case decisions (CHECK_ONLY
short-circuit, drain-gate entry+confirm, supervisor report/stop/start
dispatch, health-poll decision, HAD_OLD_PID computation) that lived outside
any function, so the 67-test suite could not reach them and any one could be
silently inverted with the whole suite green.

Follows the existing swap_if_built/refuse_drain_gate pattern: each bare
guard becomes a small predicate or dispatch function (should_stop_for_check,
drain_gate_required/drain_confirmed/run_drain_gate, report_supervisor_state,
dispatch_stop, dispatch_start, health_is_up/report_health,
compute_had_old_pid), called unconditionally by the main flow so the
decision itself is unit-testable in isolation.

Adds a structural guard, test_no_untested_main_flow_conditionals, that scans
the main flow (everything after the SOURCED guard) for bare if/elif/case
lines outside any function body, tracking function boundaries via this
file's one consistent name() { / } convention. It fails on any new bare
conditional not covered by MAIN_FLOW_ALLOWED_CONDITIONALS, an explicit
exact-text allowlist of the report-only/display conditionals and the two
#504-family supervisor elif branches that stay out of scope for this
ticket. This is the "shape, not the eight sites" guard the ticket asked
for: a ninth bare decision fails immediately, naming its line.

Out of scope, not touched: #504 items 2/3/4 and #528 item 2 (same
untested-main-flow family) — the seam here generalizes to make them
testable too, but lifting them was left for their own tickets.
2026-09-12 16:46:53 +07:00
Dai Ha cb64bc8157 fleetd #513: rework — TIMED_OUT_QUEUED has four routes, not one
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m16s
Comment 16984 on the ticket showed my first pass (ed28b51) replaced one
wrong invariant with a narrower one: it described injector.cancel()
cancelling a queued entry as THE mechanism, when three of its four
routes (Injector.java: the send-to-terminal call throwing, the
readiness grace expiring, or the target being torn down) never cancel
anything — cancel() just reports a state a different code path already
set. It also claimed the target never saw a word of the message, which
is not established when the herdr agent.prompt call throws after
already pasting.

Rewrote all four comments (the TIMED_OUT_QUEUED enum constant,
queuedDeliveries, hasQueuedDelivery, hasOrphanedDelegation) plus the
pre-existing inline comment that seeded the original bad wording, to
state only what holds on every route: the message will not arrive
later and is not sitting in a queue. The CANCELLED case is called out
as the only one where the target is known to have seen nothing; the
NOT_DELIVERED case (including the herdr-send-threw route) is flagged
as leaving that open.

Comment-only; no behavior change.
2026-09-12 16:45:35 +07:00
Dai Ha ed28b51f12 fleetd #513: fix TIMED_OUT_QUEUED javadoc — cancelled, not queued
CI / shell-tests (pull_request) Successful in 4s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 2m22s
Two javadoc blocks (queuedDeliveries field, hasQueuedDelivery) said a
timed-out message is still sitting in the injector's per-target queue.
CB-640 made send() cancel it via Injector.cancel() instead, so the
message is gone and will never arrive. Rewrote both to describe
cancellation. Also fixed a third instance of the same stale claim in
hasOrphanedDelegation's javadoc, and added a line to the
TIMED_OUT_QUEUED enum constant's own comment clarifying the name is
kept but no longer means the message stays queued.

Per the ticket's follow-up comment: no rename (TIMED_OUT_QUEUED reaches
FleetApp.java REST mapping and FleetMcp.java — out of scope here) and
no behavior change; comments only.
2026-09-12 16:32:36 +07:00
ltms 4f9aba40e7 Merge #558: the ticket is the pull channel, and the brief is write-once
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m51s
Docs only, 16 additions / 3 deletions — read by the lead in full, per the
under-50-lines self-review rule in CLAUDE.md.

Verified on the merged tree (main dab697f + 351ee1e): the canonical-block sync
check from the addendum prints `in sync: True`, comparing the MERGED CLAUDE.md
against wiki/7-Use-Cases.md. The wiki side is already pushed and verified by ref:
local wiki HEAD and `git ls-remote origin main` both read 20d2fc0.

CI run 1812 on 351ee1e: success.
2026-09-12 11:12:55 +02:00
ltms dab697fae0 Merge #559: progress watchdog for StatusPoller and SessionReaper loops (fleetd #544)
CI / shell-tests (push) Successful in 7s
CI / contract (push) Successful in 51s
CI / build (push) Successful in 2m19s
Verified by the lead on a merged tree (main 7a3b2bb + bfac141 = 5de807f):
mvn clean install exit 0, Tests run: 1744, Failures: 0, 130 surefire reports.

Three mutations run by the lead, all killed:
- StatusPoller.java:105 `watchdog.reset()` deleted (the surviving mutant from the
  first review) -> aRestartedLoopReportsRunningAgainNotStoppedForever:
  expected: <RUNNING> but was: <STOPPED>.
- LoopWatchdog.java:73 `lastRoundNanos = nowNanos.getAsLong()` deleted from reset()
  (not run by the worker) -> LoopWatchdogTest.resetClearsAPreviousStopAndTheStaleClock:
  expected: <RUNNING> but was: <STALLED>.
- LoopWatchdog.java:89 `>=` -> `>` boundary (not run by the worker) ->
  LoopWatchdogTest.reportsStalledOnceTheLastRoundAgesPastTheThreshold:
  expected: <STALLED> but was: <RUNNING>.

Each mutation counted the pristine full line 1 -> 0 by exact string equality
(awk '$0==p'), with a pristine control copy still reading its original count, and
was restored to a byte-identical file (shasum -a 256) before the next run.
Control build after all restores: exit 0, 1744 tests, 0 failures.
2026-09-12 11:11:46 +02:00
Dai Ha bfac14108f fleetd #544: pin the sticky-STOPPED-across-restart invariant
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 1m42s
Review of PR #559 (issue comment #16944) found a surviving mutant: removing
watchdog.reset() from StatusPoller.start() (and the identical line in
SessionReaper.start()) passed the entire suite.

stoppedByCaller is sticky and reset() — called only from start() — is the
only thing that clears it. Both loops document start() as idempotent and
loop()'s own error log says "it can be restarted", so stop() followed by
start() is an anticipated path. Without reset() wired into start(), health()
would report STOPPED forever after a restart even though the loop is
genuinely running again.

Add aRestartedLoopReportsRunningAgainNotStoppedForever to both
StatusPollerWatchdogTest and SessionReaperWatchdogTest, pinning "an
intentional stop must not outlive the restart that follows it". Verified via
the standard mutation cycle: exact-line anchor (not regex, to avoid the
\Q-style false match the reviewer flagged) counted pristine 1 -> mutated 0,
test goes red with its own message, restored, shasum -a 256 byte-identical,
green again.

mvn clean install: exit 0, BUILD SUCCESS, Tests run: 1734, Failures: 0,
Errors: 0, Skipped: 0 (cross-checked against 130 surefire report files).

No production code changed — the reset() call under test was already
correct; it simply had nothing pinning it.

🤖 Generated with Claude Code

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-12 16:00:58 +07:00
ltms 7a3b2bb7ee Merge pull request 'fleetd #552: warn instead of aborting when the post-restart mktemp fails' (#560) from worker/552-post-restart-mktemp-abort-bc2672-4 into main
CI / shell-tests (push) Successful in 7s
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 2m9s
2026-09-12 10:58:10 +02:00
Dai Ha f188947750 fleetd #552: warn instead of aborting when the post-restart mktemp fails
CI / shell-tests (pull_request) Successful in 7s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m7s
By the time the fresh-log mktemp ran, the daemon had already been stopped, the
jar swapped, and the new daemon started — an unguarded mktemp failure there
aborted the whole script anyway, so a caller read the resulting non-zero exit
as "the redeploy failed" and would restart an already-correctly-restarted
daemon.

Extract the mktemp into capture_fresh_log_region, guarded the same way
unload_launchd_if_loaded/stop_systemd_if_loaded guard their own, but warn
instead of die: there is nothing left to protect by refusing after a
successful restart. The trap is now installed before the assignment it cleans
up, using an FRESH_LOG="" sentinel readers can check.

classify_amqp_connection_errors and report_shutdown_drain both gain a new
state (REDEPLOY_AMQP_CHECK_SKIPPED / REDEPLOY_DRAIN_STATE=skipped) for an
uncapturable log region, distinct from "captured a region with nothing in
it" — and the result section gains a matching branch, so a skipped capture
can never read as a clean bill of health.
2026-09-12 15:55:04 +07:00
ltms f606fccf7f Merge pull request 'fleetd #553: register the rendezvous waiter in onStatus's finally backstop' (#557) from worker/553-onstatus-completion-leak-0da881-2 into main
CI / shell-tests (push) Successful in 4s
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m40s
2026-09-12 10:50:24 +02:00
Dai Ha 735b6af976 fleetd #544: progress watchdog for StatusPoller and SessionReaper loops
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Successful in 2m29s
Each loop's virtual-thread runner (StatusPoller, SessionReaper) can die or
get permanently parked in a herdr call with no read timeout, and nothing
observed it: /healthz stayed green and Thread.isAlive() kept reporting true
the whole time.

Add LoopWatchdog (dev.ltms.fleet.inject — see its javadoc for why not
dev.ltms.fleet.health, which would close a package cycle through session):
each loop now records a monotonic last-completed-round timestamp
(injectable LongSupplier clock, same pattern as Injector/SessionManager) and
exposes it as a three-state health() fact — RUNNING, STALLED (dead or
parked, indistinguishable from outside), STOPPED (stop() was called on
purpose, never an alarm). This is the fleetd #512 shape: one flag cannot
carry both "halted on purpose" and "halted unexpectedly", so stop() marks
its own state explicitly instead of leaving state() to infer it from
staleness.

Staleness thresholds are derived from each loop's own poll interval with a
documented multiplier: StatusPoller 40x (250ms -> 10s), SessionReaper 12x
(5000ms -> 60s).

Scope: observability only, per the ticket's own comment. No restart/recovery
mechanism, no /healthz or REST/MCP wiring beyond the public health() API, no
change to the per-item catch(Throwable) behavior (#543) or a process-wide
uncaught-exception handler (ruled out on #538), and no deadline added to the
herdr read itself (a separate, real ticket).

🤖 Generated with Claude Code

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-12 15:48:07 +07:00
Dai Ha 8b4320ed24 fleetd #553: split sentHandled's two meanings so onDelivered's own throw still completes the future
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m43s
Lead review of PR #557 (ticket comment 16916) found one path left open: sentHandled
is set to true BEFORE onDelivered() runs (correctly, per the earlier fix), so when
onDelivered() itself throws on the normal path, the finally's 'if (sent != null &&
!sentHandled)' guard skipped the whole recovery -- completion included -- and left
sent.delivered() pending forever for a message that really was delivered.

sentHandled must guard only the onDelivered RE-CALL (the permanent-suppression
hazard), never the future completion, since CompletableFuture.complete/
completeExceptionally are idempotent and a no-op on the already-handled path.
Split the one flag's two jobs: the outer 'if (sent != null)' now always runs the
recovery block, and '!sentHandled' moved onto just the onDelivered call inside it.

Added anOnDeliveredThrowOnTheNormalPathStillCompletesTheDeliveryFuture, proven with
the lead's own mutation (reverting !sentHandled onto the outer if): the new test
goes red while anOnDeliveredThrowAfterItsOwnRegistrationDoesNotRunASecondTime stays
green, showing the two concerns are genuinely separate.
2026-09-12 15:45:36 +07:00
Dai Ha 351ee1ea6d CLAUDE.md: the ticket is the pull channel, and the brief is write-once
CI / shell-tests (pull_request) Successful in 6s
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 2m10s
A send to a working member is accepted and returns a ticket, then is never
delivered. That happened three times in one session here, and the member was
released still executing a brief that had been retracted twice. The receipt is
true — it is a fact about the mailbox, when what was needed was a fact about the
pane.

The fleet01 lead named the mechanism: a push delivery needs the recipient free at
send time, while a pull channel needs only that they look before acting. So the
ticket is not more reliable than the mailbox, it is a different direction, and its
success depends on the member's procedure rather than on the timing of the send.

Both halves have to be written down, because each is useless alone:

- Member (turn contract, new item 4): re-read the ticket before acting on anything
  told earlier, and again before committing. A ticket comment that contradicts the
  brief is newer and wins.
- Lead (step 5): all corrections go to the ticket, and the brief is write-once.
  The member cannot check which source is newer — it just always prefers the
  ticket — so revising a brief in place makes it obey the rule and do the wrong
  thing. A first brief for a unit not yet running is not a correction.

wiki/7-Use-Cases.md is updated to keep the canonical block byte-identical; the
sync check passes. The wiki submodule pointer is deliberately left unstaged.
2026-09-12 15:45:09 +07:00
Dai Ha d4a51c6274 fleetd #553: register the rendezvous waiter in onStatus's finally backstop, not just the delivery future
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m12s
The previous try/finally around onStatus's post-monitor region completed
sent.delivered() but never registered sent.token().waiter() when an earlier
listener threw. That waiter is registered only by turnListener.onDelivered(),
inside the very if (sent != null) block the finally backstops, so a caller
was told its send landed and then waited out its full timeout for an answer
that could never resolve (worse than a plain hang).

The finally now does that block's whole job on the unhandled path: it calls
onDelivered() (when sendError == null) before completing the future, guarded
by its own try/catch(Throwable) so a failure there cannot mask the original
throwable. A sentHandled flag, set true at the START of the normal block
(before any side effect), tells the finally whether that already ran, so a
throw partway through onDelivered cannot trigger a second, late captureBaseline
that would permanently suppress the turn's completion.

Also removed a leftover duplicated forget.accept(target) call (with a stray
'MUTATION-TEST-3' comment) in the notReady block — residue from the previous
worker's own mutation testing that was not fully reverted.
2026-09-12 15:35:53 +07:00
ltms 26f380a00b Merge #554: portable hash256, a third state for an unhashable jar, and the shell suite in CI (#550)
CI / shell-tests (push) Successful in 8s
CI / contract (push) Successful in 1m17s
CI / build (push) Successful in 2m24s
Closes fleetd #550, all three items.

Verified by me on the branch at b8182c9, in a scratch worktree, not from the implementer's report.

The decisive pair, both in ubuntu:latest where shasum is absent and sha256sum is present:

  main   (93a9ed3):  SUITE exit=127   anchored ^FAIL: count 0
                     scripts/test-redeploy-fleetd.sh: line 298: shasum: command not found
  branch (b8182c9):  SUITE exit=0     anchored ^FAIL: count 0

Identical FAIL counts, opposite exit codes. That is why the new shell-tests CI job gates on the
step's own exit status and deliberately does not grep for a FAIL count: a suite that dies before
running a single test prints exactly what a clean pass prints.

Also measured by me: macOS exit 0 / anchored count 0; 70 tests defined and 70 invoked with an
empty comm -3; bash -n exit 0 under both /bin/bash 3.2.57 and bash 5.3.9; ci.yml parses with jobs
build, shell-tests, contract, and shell-tests is ubuntu-latest + actions/checkout@v4 +
bash scripts/test-redeploy-fleetd.sh.

Three mutations, all killed, each restored byte-identical against
515d929bb53c9ec2c95042e47d3e4d611d60171345227b07feb0de0d20643d3a:

  drop the sha256sum branch from hash256   anchor 1->0   Linux exit 1,
      test_no_unguarded_macos_only_hasher_calls with its own message
  echo "unhashable" -> echo "absent"       anchor 1->0   FAIL: jar_id reported absent for a file
      that exists, only because no hasher was on PATH
  both hash256 arms -> shasum -a 1         anchors 1->0 and 1->0   FAIL: hash256 of the literal
      3-byte input 'abc' must be the known SHA-256 prefix, not some other algorithm's: expected
      ba7816bf8f01, got a9993e364706

The third mutation SURVIVED on the first head (3da44ee) and was sent back. Fixing item 2 had
rewired the reference hashes in test_jar_id_defaults_to_live_and_reports_explicit_path onto
hash256 itself, making the test's reference and its subject one instrument — they agree whatever
it computes, and the existing fixture guard catches only a constant return, not a wrong algorithm.
b8182c9 adds test_hash256_computes_a_real_sha256, pinning the FIPS 180 vector for "abc" as a
literal constant written into the test rather than taken from any hasher. That closes it: the
mutation now goes red, and a9993e364706 is SHA-1("abc"), which proves the mutated code really ran.

Root cause, for the record: shasum was the trigger, not the cause. Under set -euo pipefail a
missing hasher makes the pipeline status 127, the `|| echo "absent"` fires, and jar_id returns a
confident false "absent" for a jar that is right there. Without pipefail the same function returns
an empty string and is visibly broken. A default at the read site that maps every failure onto one
value which already means something specific is the defect; the fix separates "not there" from
"could not hash it".
2026-09-12 10:20:29 +02:00
Dai Ha b8182c96c2 fleetd #550: pin hash256's algorithm against a literal SHA-256 test vector
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Successful in 2m27s
test_jar_id_defaults_to_live_and_reports_explicit_path's reference hash is computed by calling
hash256 itself (needed so it doesn't call the Linux-crashing bare shasum directly). That made
subject and reference the same instrument: they agree no matter which algorithm hash256 actually
runs, so a mutation swapping both of hash256's arms for the wrong algorithm was invisible to the
suite.

Adds test_hash256_computes_a_real_sha256, pinned against the published SHA-256 test vector for the
3-byte input "abc" (ba7816bf8f01...), written as a literal constant rather than computed by any
hasher at test time. Verified the constant myself both ways (sha256sum and shasum -a 256) before
writing it in.
2026-09-12 15:16:43 +07:00
Dai Ha 3da44eed63 fleetd #550: replace shasum with a portable hash256 helper, add a Linux CI job for the shell suite
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m46s
jar_id() in redeploy-fleetd.sh called shasum directly, which does not exist on GNU coreutils
Linux (Debian/Ubuntu/etc.) — there it silently reported an existing jar as "absent" with exit 0,
because the missing command made `cut` succeed on empty input and pipefail's failure was then
swallowed by the `|| echo "absent"` fallback. The shell test suite hit the same tool at
test-redeploy-fleetd.sh:298-299 and died at exit 127 with zero FAIL lines printed — the same shape
as a clean pass on the one channel anyone would check.

Adds one hash256() helper (prefer sha256sum, fall back to shasum -a 256, same idiom already used
in probe-member-credentials.sh) and points jar_id and the test suite's own reference hash at it.
jar_id now has three distinct answers instead of two: absent, a hash, or "unhashable" when neither
hasher is on PATH — "absent" is never used for a file that exists.

Adds a CI job (shell-tests) that runs scripts/test-redeploy-fleetd.sh on ubuntu-latest, gated on
the step's own exit code rather than a FAIL-line count, since a suite that dies before running is
exactly what a green run also looks like by that count.

New tests: test_jar_id_reports_unhashable_when_no_hasher_on_path (stubbed PATH with neither
hasher) and test_no_unguarded_macos_only_hasher_calls (a shape check across every script under
scripts/, not named lines — #545 already showed this idiom spreading from two sites to six).
2026-09-12 15:06:15 +07:00
ltms 93a9ed3f83 Merge #549: widen Injector's delivery catch to Throwable (#546)
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m43s
Closes the re-delivery window that merging #543 opened. I caused that; this closes it the same day.

Verified by me on the branch at 87871ea, base 0b032f5 (not stale).

Production diff is 11 lines: `catch (RuntimeException)` -> `catch (Throwable)` at Injector.java:391,
and the `sendError` local widened to `Throwable` so it compiles. I checked every use of `sendError`
myself — :533 `getMessage()` and :534 `completeExceptionally(Throwable)` — so the wider type
reaches nothing that needed the narrow one.

Build on the branch: exit 0, Tests run: 1719, Failures: 0, Errors: 0, Skipped: 0 (1716 on main plus
the 3 new tests). #459's javadoc reference gate: exit 0, 0 reference errors. Gitea CI run 1803 on
87871ea: success.

Three mutations of my own, none of them the ones the worker used:
1. Reverted the catch to `RuntimeException`. RED: `anErrorFromSendDoesNotRedeliverOnASecondRound`
   "expected: <1> but was: <2>" prompt calls, and `anErrorFromSendRemovesTheMessage...` reporting
   the Error escaping `onStatus`. That second message is the defect itself, stated by the test.
2. Deleted the `t.queue.poll()` in the catch arm. RED on the new test AND on the pre-existing
   `sendFailureDropsMessageAndFailsItsFuture` — so the new test is not carrying that behaviour alone.
3. Wrote `DELIVERED` instead of `NOT_DELIVERED` at :400, the catch arm only. RED on the new test and
   on `aHerdrExceptionFromSendStillProducesNotDeliveredUnchanged`, which is the control that proves
   the widening did not quietly change the ordinary path.

All three restored; sha256 back to 97c560b6e33fc49a1772abec92e5bbab613f8991deba2220d30221a2f546ba14.
Green control after the restores: InjectorTest 37/37, exit 0.

My third mutation did not apply on its first attempt — it asserted a unique match on
`p.state = Pending.State.NOT_DELIVERED;`, which occurs twice (:400 and :419), so the script wrote
nothing and the test run came back exit 0. That is not a surviving mutant, it is a non-result
wearing the same clothes. The pristine-anchor count catching it is the only reason I noticed.

Deliberately NOT fixed here, filed as #551: the catch arm assumes that reaching it means nothing was
sent, and nothing establishes that. `agent.prompt` pastes and submits in one call, and every failure
in the response half of `UnixSocketHerdrClient.call()` — dropped connection, malformed line, error
result — is a `HerdrException`, which is a `RuntimeException`, which this catch arm already caught
before today. So "records NOT_DELIVERED for a delivery that happened" is older than this PR and is
not created by it. The fleet01 lead argued it was a trap inside this change and asked to be argued
out of it before the merge; the measurement above is the argument, and their underlying diagnosis is
right and is now #551 with their wording on it.
2026-09-12 09:39:41 +02:00
ltms 0b032f5a1a Merge #548: fix mktemp -t templates for GNU coreutils, split the unclear-supervisor detail (#545)
CI / contract (push) Successful in 55s
CI / build (push) Successful in 2m16s
Verified by me on the branch, not on the worker's report.

Source audit, on a scratch worktree at a476a14:
- Every `mktemp -t` site in scripts/ now carries an X placeholder. The only remaining
  `mktemp -t` text with no X is a prose comment in the test file, not a call.

Suite, macOS (/bin/bash 3.2.57 and env bash 5.3.9):
- exit 0, anchored `^FAIL:` count 0. Unanchored `FAIL:` count 3 — the suite's own internal
  mutation-cell fixture lines, same as main.
- Test functions defined vs invoked: 67/67, `comm -3` empty.

Three mutations of my own, none of them one the worker used:
1. Removed `.XXXXXX` from the `fleetd-fresh-log` site (line 1089). Pristine anchor count went
   1 -> 0, so the mutation really applied. Suite exit 1, FAIL named that exact line.
2. Made the state-2 branch in `detect_supervisor` unreachable (`= 2` -> `= 9`). Suite exit 1:
   "a systemd probe setup failure must read as unclear, not none: expected unclear, got none".
3. Made `systemd_loaded` set the old value (`=2` -> `=1`) on setup failure. Suite exit 1:
   "must flag a SETUP failure (2), distinct from a probe-answered-with-stderr failure (1)".
All three restored; `shasum -a 256` back to 77fe15e5945c7d4ef9b1a2cd8f46e6f1d0be5004f595ea03964d1d7c16e859f7,
the same hash the worker reported independently. Green control after the restores: exit 0, 0
anchored FAILs.

A fourth attempt did not count. A perl `\Q...\E` pattern silently interpolated the shell
variables in it, so the file was never changed and the suite's exit 0 meant nothing. The proof
cell caught it: the pristine anchor count was still 1 after the "mutation". A mutation that did
not apply is not a surviving mutant.

The measurement macOS cannot make: I ran both arms under GNU coreutils 9.1 in a
debian:bookworm-slim container, with a `systemctl` stub that exits non-zero and writes NOTHING to
stderr — a clean negative answer.

  main (a476a14's base):
    mktemp: too few X's in template 'systemd-loaded-err'
    systemd_loaded rc=1  SYSTEMD_LOADED_ERRORED=1
    detect_supervisor => unclear | "systemctl exited non-zero and reported an error on stderr,
                                   not a clean negative — e.g. it cannot reach the user bus"

  this branch:
    systemd_loaded rc=1  SYSTEMD_LOADED_ERRORED=0
    detect_supervisor => none

So on Linux, main tells the operator that systemctl answered badly, when systemctl ran fine and
gave a clean negative. The message named a cause that was never measured. This branch removes it.

Found while doing this, NOT part of this PR, ticket to follow: the shell suite cannot run on
Linux at all. It dies at the first `shasum` call with "command not found" and exit 127, and the
anchored `^FAIL:` count reads 0 — identical to a green run. Gitea CI never runs this suite, so
nothing caught it.
2026-09-12 09:30:30 +02:00
ltms fad99c4c5e Merge #547: record the finally non-goal on drainAll's completion line
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m10s
Comment only. No behaviour change.

Verified by me before merging:
- `mvn -B clean install` in a scratch worktree: exit 0, Tests run: 1716, Failures: 0, Errors: 0,
  Skipped: 0. Same count as main, as expected for a javadoc-only change.
- #459's javadoc reference gate: exit 0, 0 reference errors.
- Gitea CI run 1801 on 5eb4267: success.

The constraint came from the fleet01 lead. Their point: the value of the `drain complete` line is
that it is MISSING when a drain does not finish. A `finally` block would print it after a drain
that threw, with partial counts, and destroy both halves at once. The code already avoids this;
what was missing was the sentence that stops a reviewer putting it back.
2026-09-12 09:29:46 +02:00
Dai Ha 87871eaefb fleetd #546: widen Injector's delivery catch to Throwable, stop re-delivery on Error
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m49s
Injector.java:391 caught only RuntimeException around the herdr send seam. PR #543
(fleetd #538) widened StatusPoller's per-target catch to Throwable so the polling
loop now survives an Error there, which means it comes back round — and Injector's
narrower catch let the poisoned message stay QUEUED (the loop peeks, not polls),
so the next round re-sent the same text into the member's pane.

Widen the catch to Throwable, matching #543 one layer down. sendError's declared
type widens from RuntimeException to Throwable to keep compiling; its only consumer
(CompletableFuture.completeExceptionally(Throwable)) already accepts that type, so
no other caller-visible behavior changes. The ordinary HerdrException/RuntimeException
path is unchanged.

Adds three tests: an Error at the send seam is dropped and marked NOT_DELIVERED, a
second onStatus round does not re-send it, and a HerdrException control proves the
ordinary path is untouched.
2026-09-12 14:26:27 +07:00
Dai Ha a476a14f1c fleetd #545: fix mktemp -t templates for GNU coreutils, split unclear-supervisor detail
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 2m28s
Every mktemp -t template in redeploy-fleetd.sh lacked an X placeholder. BSD mktemp
(macOS) tolerates that and appends its own suffix; GNU mktemp (every Linux
distribution) refuses it and exits non-zero. All six sites now use .XXXXXX.

detect_supervisor's 'unclear' detail used to cover two different facts with one
message that always named 'systemctl exited non-zero and reported an error on
stderr' — even when systemctl was never run, because mktemp failed first. The
SYSTEMD_LOADED_ERRORED/SYSTEMD_INSTALLED_ERRORED flags now carry a third value
(2 = the probe's own mktemp setup failed) alongside the existing 1 (systemctl ran
and answered badly on stderr), and detect_supervisor gives each its own detail
text. kind stays 'unclear' in both cases; require_drivable_supervisor is unchanged.

Tests added to scripts/test-redeploy-fleetd.sh:
- test_mktemp_dash_t_templates_have_x_placeholders: source-text check, fails if
  any mktemp -t template lacks an X.
- test_detect_supervisor_systemd_probe_setup_failure_is_unclear: proves the
  SET-UP-FAILED detail when mktemp itself fails (systemctl never runs).
- test_detect_supervisor_systemd_probe_error_is_unclear: extended with assertions
  that the PROBE-ANSWERED-WITH-STDERR detail is present and the SET-UP-FAILED
  wording is absent, so swapping the two messages fails a test in both
  directions.
2026-09-12 14:22:57 +07:00
Dai Ha 5eb4267a4a fleetd #512 follow-up: record why the drain-complete line must not move into a finally
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Successful in 2m8s
The javadoc said the log.info fires "every time", which reads as an invitation
to the exact edit that destroys it. The absence of the line is the signal that
the drain died, so a finally would remove the signal and print partial counts in
the same change.

Raised by the fleet01 lead from their 2026-09-10 incident: that drain is known to
have died only because it threw and left a stack trace. A drain that hung, or
returned early on a condition, leaves no trace, no ERROR token and no priority --
only a missing line.

Comment only. No behaviour change.
2026-09-12 14:21:15 +07:00
ltms 7611b69667 Merge pull request 'fleetd #504 item 1: stop the false ok on the loaded-but-not-running stop path' (#541) from worker/504-failed-reported-clean-3cfd66-3 into main
CI / build (push) Successful in 1m35s
CI / contract (push) Successful in 1m42s
2026-09-12 09:09:00 +02:00
ltms cc302fe4af Merge pull request 'fleetd #538: recover polling loops after errors' (#543) from worker/538-loop-dies-on-error-4a5eeb-6 into main
CI / contract (push) Successful in 51s
CI / build (push) Successful in 2m10s
2026-09-12 09:02:48 +02:00
ltms 4a8a780274 Merge pull request 'fleetd #426: pin FleetHealthMonitor.coverage and its HealthCoverageSource call site' (#542) from worker/426-health-coverage-ef1fd4-4 into main
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m28s
2026-09-12 08:55:46 +02:00
Dai Ha 343ce0f4c0 fleetd #538: recover polling loops after errors
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 2m0s
2026-09-12 13:47:44 +07:00
ltms cec3e191d4 Merge pull request 'fleetd #459: lint Javadoc references in CI' (#539) from worker/459-broken-link-targets-cadc17-5 into main
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 2m9s
2026-09-12 08:46:23 +02:00
Dai Ha 1850a5f324 fleetd #426: pin FleetHealthMonitor.coverage and its HealthCoverageSource call site
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 2m16s
FleetHealthMonitor.coverage had zero references in the test tree — not the
method, not either output string, not the field it populates. Inverting
`enabled`, swapping "full"/"detection-only", or breaking the argument pairing
at the HealthCoverageSource call site in Fleetd.java all shipped a green
build.

Extract the HealthCoverageSource lambda out of Fleetd.main into a
package-private static factory (healthCoverageSource(ConfigRef)), the same
shape capacitySource/quarantineSource already use for the identical
argument-pairing risk (fleetd #415). #407's "keep the config invalid, assert
on the log line before validateAll() throws" option does not apply here: this
call site is built well after validateAll() and after a real herdr socket
connect, so driving it through a real Fleetd.main would require the socket
I/O this ticket's tests must not do.

Add FleetHealthMonitorCoverageTest (the three-branch method itself) and
FleetdHealthCoverageSourceWiringTest (the call site, via a real
FleetConfig.load + ConfigRef against @TempDir fixtures, including a hot
notifications-reload case). Output strings are unchanged — "detection-only"
is still what a live fleet_list reports today.

Measured: all three mutations killed by the new tests.
2026-09-12 13:43:50 +07:00
Dai Ha e4eb3dbed4 fleetd #504 item 1: stop swallowing real launchctl/systemctl failures on the 'loaded but not running' path
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m55s
The two 'loaded but not currently running' branches in the stop step (launchd/systemd, reached
when $OLD_PID is empty) ran 'launchctl unload'/'systemctl --user stop' with '2>/dev/null || true'
and printed 'ok' unconditionally. That swallowed a real supervisor failure (e.g. launchd or the
systemd user bus unreachable) exactly like a harmless already-stopped answer, and let the script
proceed to start a new daemon believing nothing was loaded -- the two-daemons failure fleetd #492
exists to prevent.

Adds unload_launchd_if_loaded/stop_systemd_if_loaded, applying systemd_loaded's own pattern
(capture stderr separately; a non-zero exit WITH stderr is a real failure, a non-zero exit with
empty stderr is a clean already-stopped answer) to the write side. The two call sites now use
these functions instead of the bare '|| true'.

Adds 5 tests: dies-on-real-failure and tolerates-clean-negative for each function, plus a
source-text check that the main flow calls the new functions instead of the original bare
'2>/dev/null || true'. All 5 verified by mutation (reintroducing the swallow, and separately
over-correcting to die unconditionally) -- each goes red with its own message, restores
byte-identical (full sha256), and passes a green control.
2026-09-12 13:42:27 +07:00
ltms 57cd96f5e6 Merge pull request 'fleetd #537: pin CapturedLog.close()'s appender-detach and setLevel-immunity contracts' (#540) from worker/537-capturedlog-close-e4c437-2 into main
CI / contract (push) Successful in 47s
CI / build (push) Successful in 2m9s
2026-09-12 08:42:19 +02:00
Dai Ha 202e37e3b3 fleetd #537: pin CapturedLog.close()'s appender-detach and setLevel-immunity contracts
CI / contract (pull_request) Successful in 1m2s
CI / build (pull_request) Successful in 1m36s
Only the level-restore half of close() was pinned before this
(WorktreeSessionManagerTest). Deleting logger.detachAppender(appender)
from close() left mvn clean install green (1701 tests, 0 failures) --
the appender-detach half of the contract was unmeasured.

Adds CapturedLogTest with three tests, each using a logger name no
production class uses:
- closeDetachesTheAppenderSoALaterLogIsNotCaptured: an event logged
  after close() must not land in events().
- closeRestoresTheLevelCapturedAtOpen: the helper's headline contract
  in one place, independent of any production class.
- setLevelDuringCaptureDoesNotChangeWhatCloseRestores: setLevel()'s
  own javadoc claim that close() always restores the level captured
  at construction, never a value set through setLevel() mid-capture.

Test-only change; CapturedLog.java itself is untouched.
2026-09-12 13:36:55 +07:00
Dai Ha 90253f832d fleetd #459: lint Javadoc references in CI
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 2m9s
2026-09-12 13:36:37 +07:00
ltms f1640f5dcc Merge pull request 'fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog' (#536) from worker/535-appender-leak-fe74c1-1 into main
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m28s
2026-09-12 08:25:01 +02:00
Dai Ha c7903c1efe fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m52s
Three call sites (captureFleetdLogs at :55) attached a ListAppender to the
Fleetd.class logger with addAppender and never detached it, and never
called appender.setContext(...) either. Logback Logger instances are
cached per class and shared for the whole JVM, and surefire reuses forks,
so all three appenders stayed attached for every later test in the fork.

Convert all three call sites to CapturedLog.of(Fleetd.class) (added in
#533) via try-with-resources, which detaches the appender and sets the
context for free. Delete captureFleetdLogs(); nothing calls it now.
2026-09-12 13:17:21 +07:00
ltms 7d711942fe Merge #534: detect a died shutdown drain the ERROR count is blind to (fleetd #512 part 2)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m52s
Verified independently. The branch is based on 8335b12 while main is at bec87f9, so I merged locally first and tested the MERGED tree, not the branch — a clean auto-merge is not a working merge.

Merged tree checks:
- `bash -n` exit 0 on both scripts, under /bin/bash 3.2.57 and env bash 5.3.9.
- Suite exit 0, 0 lines matching `^FAIL:`, 255 bytes of output. 60 test functions defined, 60 invoked, and no defined-but-never-invoked orphan (checked with a comm against the invocation list, not by comparing two counts — two equal counts can both be wrong).
- My own comment fix from bec87f9 survived the merge and is still at :317.

I ran two mutations the worker did not, per "mutate the half the worker did not":

(A) The one that matters, because it is the defect this ticket exists to prevent: collapsed the `unknown` state into `complete`, so "cannot tell" reports as a pass. Result exit 1, one FAIL: `cannot-tell fixture must set REDEPLOY_DRAIN_STATE=unknown: expected unknown, got complete`. So the third state is genuinely load-bearing, not decoration.

(B) Broke the positive check: changed `find_drain_complete_line`'s pattern from `drain complete: released=` to `drain finished: released=`, one site. Result exit 1, one FAIL: `find_drain_complete_line did not capture the present line`.

Proof that (B) applied, against a pristine copy: the full grep line 1 -> 0, the mutant form 0 -> 1, and the bare phrase 2 -> 1 with the comment occurrence untouched. Both files restored byte-identical; `git diff --quiet` clean; green control re-run.

A note on my own proof cell for (B), because it was wrong the first time. I wrote the counts with escaped double quotes inside an already double-quoted command substitution, so the shell split the pattern on spaces and grep treated the words as filenames. It printed "2 and 2" alongside `ugrep: No such file or directory` warnings — a symmetric, plausible-looking pair that meant nothing. The kill itself was never in doubt, since the suite named the exact function, but the cell that was supposed to prove the mutation applied proved nothing. Re-done with single quotes. This is the same trap already written down for this repo, hit by me, in a cell whose only purpose was to guard against exactly this.

One thing I checked that no test covers: the main flow's `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1` runs under `set -euo pipefail`, and on a cold start the test fails. Sourcing stops before the main flow, so no behavioural test reaches that line. If `set -e` fired there, every cold start would abort before the health checks. It does not: `set -e` exempts the left side of an `&&` list, confirmed by running it under both shells — `survived, HAD_OLD_PID=0` on 3.2.57 and on 5.3.9. Safe, but it is untested main-flow wiring, which is the same class as #528's item 1 and belongs on that list.

On the `n/a` fourth state, which the worker flagged for a reviewer's judgment rather than quietly keeping: accepted, and it is in scope. The ticket asked for a third state because a sentinel conflating "no" with "cannot tell" hides two causes needing opposite handling. "The question does not apply" is a third such cause, not a variant of "cannot tell". Without it, the new warning would fire on every clean cold start, and a warning that cries wolf on the most common path trains the operator to skip it — which destroys the absence signal just as surely as putting the completion line in a `finally` would. The worker also proved the gate is actually consulted, using a cold-start fixture whose content deliberately looks like a died drain, so the test would fail if the gate were skipped. That is the right way to test a gate.

Both remaining outcomes are correctly excluded from the "no ERROR lines since restart" summary: only `complete` and `n/a` let it print.
2026-09-12 08:04:15 +02:00
Dai Ha bec87f987c scripts: name the mechanism in detect_supervisor's constraint 2, not a line number
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 1m33s
Constraint 2 read "This script runs under `set -euo pipefail` (line 50), so an
unset variable is a loud failure." Two problems, both small and both the same
family as fleetd #494 — a comment that states the wrong reason.

The line number was stale: the `set` line is at 54, not 50. It was the only
line-number citation in the file, and a citation like that goes stale on the
next insert above it, silently, with nothing to catch it.

The mechanism was also misattributed. What makes an unset variable a loud
failure is `set -u`. Naming the whole `-euo pipefail` string invites the reader
to credit pipefail for it, which is the mistake fleet01 flagged on a different
cell this week: pipefail is insurance against a future pipeline stage, not what
catches the current shape.

Now names `set -u` and says where it is without a number, and records why the
number is gone so nobody adds one back.

Comment only. bash -n exit 0 under /bin/bash 3.2.57 and env bash 5.3.9;
scripts/test-redeploy-fleetd.sh exit 0 with 0 lines matching ^FAIL:.
2026-09-12 12:58:37 +07:00
Dai Ha 7f8a8829f9 fleetd #529 follow-up: the new helper's javadoc claimed a reach it does not have
CapturedLog's class javadoc said it is "the one way to pin or capture a logger's
level and output in this test tree". Measured on main at af95897, that is false:
nine test files still hand-roll the ListAppender + setLevel + finally
detachAppender pattern, with 42 setLevel calls on a raw logback Logger between
them.

None of those nine is a defect. Every one pairs its pin with a restore, so none
is the fleetd #525 leak, and #529's scope was the 19 unrestored pins only. The
problem is the sentence, not the code: a reader who believes "the one way" and
then greps finds nine counter-examples and cannot tell a leftover from a
violation. That is the same shape as a wrong reason in a comment — the text
survives while the fact under it moves.

Replaced with what is actually true: new code must use the helper, the pattern
still exists elsewhere, and here is the list plus the two commands that
re-measure it. The paragraph says to delete itself once the first command comes
back empty, rather than to keep a count up to date.

Javadoc only. mvn -f fleetd/pom.xml test-compile exit 0.
2026-09-12 12:58:26 +07:00
ltms af9589783e Merge #533: promote CapturedLog to a shared test helper and close the logger-level leak (fleetd #529)
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m35s
Verified independently in a scratch worktree at 8ea5c2b, not promoted from the worker's report.

Build: `mvn -f fleetd/pom.xml clean install` exit 0, `Tests run: 1701, Failures: 0, Errors: 0, Skipped: 0`, BUILD SUCCESS.

Base arithmetic, measured rather than carried forward: a6415f3 (this branch's parent) has 1695 `@Test` plus 2 parameterized/repeated, and the branch has 1696 plus 2 — a delta of exactly +1, matching the per-file count (WorktreeSessionManagerTest 24 -> 25). So 1700 -> 1701 is the one new proving test and nothing else. A number I had carried from earlier in the session said 1699; that number was wrong and is retired. No Java file differs between a6415f3 and main at 8335b12, so the base count is the same on both.

Mutation (a), the shared instrument: removed `logger.setLevel(originalLevel);` from `CapturedLog.close()`. Result exit 1, `Tests run: 25, Failures: 1`, the single failure being `sharedSessionManagerLoggerLevelIsRestoredAfterDirtyWorktreeReleasePinsWarn` with `expected: <TRACE> but was: <WARN>`. Restored byte-identical.

Mutation (b), the use site: replaced the try-with-resources in `releasePreservesDirtyWorktreeAndLogsWarn` with the pre-#525 hand-rolled `ListAppender` + `setLevel` + `finally detachAppender` pattern. Result exit 1, one failure, the same assertion. Restored byte-identical.

Green control on the restored tree: `git diff --quiet` clean, full suite exit 0, 1701/0/0/0.

Both mutants were killed, so per the economy fleet01 proposed and #529 adopted, neither needs a separate harness-proof cell — the kill is the proof the cell can go red.

Leak survey, re-measured here rather than taken from the report: on a6415f3, 9 files carry 19 `setLevel` pins on a raw logback `Logger` with zero restoring call; on the branch that set is empty. The 7 remaining `setLevel` calls in `GitWorktreesTest` are `reportingLog.setLevel(...)` on the `CapturedLog` instance, whose `close()` restores the original, so they are re-pins and not leaks.

File hashes match the worker's report exactly, head and tail: CapturedLog.java 486d6f5b5a30dc5ef7f75e5e10be353e720fb0de503825e88e8d96e30a61a2f7, WorktreeSessionManagerTest.java a722a98d828c82e00415d2a924d341e177ded704262aaa5963ea2d09a683df94.

Two things follow this merge rather than block it, both filed separately: one javadoc sentence in the new helper overstates its own reach, and the worker's item 4 reports a separate appender leak outside this ticket's scope.
2026-09-12 07:57:25 +02:00
Dai Ha 190436c9cf fleetd #512 part 2: detect a died shutdown drain the ERROR count is blind to
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 2m4s
The previous daemon's dead shutdown drain (an uncaught exception in a
shutdown thread) never passes through the logger, so it never carries an
ERROR/SEVERE token, so redeploy-fleetd.sh's existing ERROR-count classifier
is structurally blind to it and prints a confident "no ERROR lines since
restart" while the drain actually died.

Add scan_uncaught_exceptions (greps the shutdown window for the failure's
real shape: `Exception in thread`, `NoClassDefFoundError`) and
find_drain_complete_line (checks for #522's SessionManager.drainAll
completion line). Compose both in report_shutdown_drain, a single
decision+action function the main flow calls unconditionally (same shape
as swap_if_built/refuse_drain_gate from #521/#528), which resolves to one
of four outcomes: complete, died, unknown ("cannot tell" — the line is
absent for either of two reasons that need opposite handling: the previous
daemon predates #522, or its drain failed without throwing), or n/a (no
previous daemon was actually stopped this run). Never fails the redeploy;
warns loudly instead.

Gate the "no ERROR lines since restart" summary line on the new outcome so
it never reads as reassurance when the drain died or the outcome is
"cannot tell" (item 4 of the ticket).

Tests: 11 new test functions (60 defined/invoked, was 49), covering both
pure classifiers, all four report_shutdown_drain outcomes, a source-grep
proof of the main-flow call site (sourcing stops before the main flow
runs), an ordering check, and the item-4 gating. Full suite green
(exit 0, 0 anchored FAIL lines). Five mutations applied and killed by hand
during review, each restored to a byte-identical file afterward.
2026-09-12 12:55:54 +07:00
Dai Ha 8ea5c2bb1f fleetd #529: promote CapturedLog to a shared test helper, close the logger-level leak
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m39s
ch.qos.logback.classic.Logger instances are cached per class and shared for the
whole JVM, and surefire reuses forks. A test that pins a shared logger's level
and restores only the appender leaves that level pinned for every test that
runs after it, in the same class or a different one in the same fork.

Move CapturedLog (merged in #527 for #525) out of SessionManagerTest into
dev.ltms.fleet.testing.CapturedLog, and convert all 19 unrestored setLevel
pins across 9 files to it, so there is exactly one way to capture and pin a
logger in this test tree:
 - FleetdAwaitHerdrTest, FleetdReplyInboxSelectionTest, AuditLogTest,
   CompletionResolverTest, InjectorTest (4), LeadRolloverTest,
   AmqpConnectionFailureLoggerTest, GitWorktreesTest (8),
   WorktreeSessionManagerTest.

AuditLog logs through a named "audit" logger rather than a class, so
CapturedLog gains String-named at()/of() overloads alongside the existing
Class-based ones, plus a setLevel() method so a fixture that already pinned a
coarser baseline (GitWorktreesTest's @BeforeEach) can re-pin further for one
test without losing what close() restores.

Adds an ordered proving test to WorktreeSessionManagerTest asserting the
SessionManager logger level is back to a known baseline after the dirty-
worktree release test runs; this proves the within-class case only, since
JUnit does not guarantee cross-class ordering.

Every existing intentional pin (SessionManagerTest's two explicit INFO pins
and its @BeforeAll DEBUG baseline) is left untouched, per the ticket.
2026-09-12 12:48:49 +07:00
ltms 8335b12562 Merge #532: pin drain_gate_refusal's call site, not just the predicate (fleetd #528)
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m47s
Verified independently in my own worktree at the pushed head 7c34e8f, not taken
from the worker's report. CI run 1781: success.

The shape is the one #528 asked for and the one #526 arrived at: refuse_drain_gate
composes the message via drain_gate_refusal AND calls die itself, and the main
flow calls it unconditionally at :710. No guard is left in the main flow to
remove, invert or bypass on its own.

My measurements:

  test functions defined / invoked   49 / 49   (was 44/44; +5)
  bash -n, /bin/bash 3.2.57          rc=0 on both files
  bash -n, env bash 5.3.9            rc=0 on both files
  clean control                      exit 0, 0 lines matching ^FAIL:, 256 bytes

Two mutations, both killed, each by a differently named failure:

  call site deleted (the item-1 mutation: :710 replaced by a flat
  die "aborted — nothing changed")
    -> exit 1, 1 ^FAIL: line, 87 bytes
       FAIL: could not find the main flow's refuse_drain_gate call site in redeploy-fleetd.sh

  refuse_drain_gate stops consulting the predicate (its body's
  die "$(drain_gate_refusal ...)" replaced by a flat message)
    -> exit 1, 1 ^FAIL: line
       FAIL: refuse_drain_gate build-ran+staged-present die message does not name the staged jar

The first cell is the point of the ticket. Before this change the same mutation
gave exit 0, zero FAIL lines and output byte-identical to a clean run at 256
bytes. It now exits 1 and names the missing call site. Each mutation was proven
applied with a uniquely tagged marker plus a second, different search string,
with a control against a pristine copy showing the exact inverse (1/0 mutated,
0/1 pristine), and the function definition confirmed still present so the
mutation targeted the call and not the function. Restored byte-identical to
0e5a99a22c9c65f72960d8f179ca5299307889e06bc42131f098a513e7b97bd6 and the final
control is green.

Needle uniqueness checked, because this is where it could have gone wrong:
grep -cF 'refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"' on the production script
returns 1, at :710, the real call site. The worker hit the self-match trap while
writing the comment above refuse_drain_gate — their first draft quoted the
call-site string literally, which would have let the source-text test match the
comment instead of the call — caught it themselves, and reworded so the comment
cannot become a second match. That is the same trap that cost me a false pass on
a probe earlier today, and catching it unprompted is the better half of this PR.

The dead-check sweep came back as a real negative, with the reasoning shown
rather than asserted: of the five scripts under set -e with pipefail, every
pipe-into-assignment already carries || true or || echo, and the remaining two
scripts have no pipe-into-assignment at all. probe-member-credentials.sh and
deploy/herdr-inner.sh correctly excluded for not having set -e. No live
instances.

One inaccuracy in the report, in the report only: it abbreviates the restored
hash as "0e5a99a2...78f0a", and that tail does not occur in the actual hash,
which ends b97bd6. I hashed the committed file myself and confirmed the restore
matched, so the file is right and only the quoted abbreviation is wrong. Flagged
because an abbreviated hash that nobody can match against anything is worse than
no hash.

wait_for_daemon_exit's call site (item 2) stays open as the ticket scoped it —
source-text pinned only, "partially pinned, not audited". The seven untested
main-flow decisions are untouched; the worker correctly notes it changed the body
of one of those if blocks while leaving the guard condition itself untested, as
instructed.
2026-09-12 07:41:57 +02:00
ltms d25c863118 Merge #531: separate the blocked forge MCP server from the working GITEA_TOKEN
CI / build (push) Successful in 1m30s
CI / contract (push) Successful in 1m32s
Charter wording only — 9 insertions, 4 deletions, one file. CI run 1780 on
be07ed2: success.

Resolves the "charter may be stale" item I had been carrying. It was not stale.
Two workers reporting working forge access and the charter saying forge tools
hold a blocked credential were both correct, about two different credentials:
the repo-scoped GITEA_TOKEN the daemon injects (which opens every worker PR,
per implementer SKILL.md step 5) versus the forge MCP server that leaks in from
the operator's user-scope config (which is deliberately blocked). The wording
did not separate them, and a worker could have read it as "I cannot reach the
forge" and skipped opening its PR.

Both sentences now name the MCP server specifically and state that the injected
token is a separate, working route.

Canonical block and wiki template verified byte-identical after the edit — the
CLAUDE.md sync script reports "in sync: True". The wiki commit is d02a55d on
wiki's own main, pushed and verified by ref (ls-remote matched local HEAD), not
by exit code. The submodule pointer stayed unstaged.

Not re-measured in this change: that the blocked MCP credential does fail every
call. That claim is the existing charter's and I only narrowed what it refers
to. It would need its own probe with a request that cannot succeed on its
merits, so that a rejection can only mean the block.
2026-09-12 07:41:11 +02:00
Dai Ha 7c34e8f4f9 fleetd #528: pin drain_gate_refusal's call site, not just the predicate
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m31s
drain_gate_refusal composes the correct abort message and is well tested,
but the main flow built its own `die "$(drain_gate_refusal ...)"` call —
nothing proved that call site was ever consulted. Mutating it to a flat
`die "aborted -- nothing changed"` left the whole suite green, silently
reinstating the exact defect #517 was filed to fix.

Same shape as #521/#526's should_swap/swap_if_built: the decision and the
die() now live together in refuse_drain_gate, which the main flow calls
unconditionally. drain_gate_refusal stays separate and separately tested
for the message logic; four new behavioural tests stub die() to prove
refuse_drain_gate calls it correctly for all four cases, and a fifth
source-text test pins the main flow's call site itself (the only thing
that can catch deleting the call, since sourcing stops before the main
flow runs).
2026-09-12 12:38:16 +07:00
Dai Ha be07ed2033 charter: separate the blocked forge MCP server from the working GITEA_TOKEN
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m3s
Two worker reports said they had working forge access, which looked like it
contradicted the charter's "any forge tools it appears to have hold a blocked
credential and fail". Measured: the charter is correct and the reports are
correct. They are about two different credentials.

A worker opens its own PR with curl and a repo-scoped GITEA_TOKEN that the
daemon injects (.claude/skills/implementer/SKILL.md step 5, lines 96-117).
That route works — it is how every worker PR this week was opened. The blocked
credential belongs to the forge MCP server that leaks in from the operator's
user-scope ~/.claude.json, which is a separate thing and does fail every call.

The wording did not distinguish them. A worker reading "any forge tools it
appears to have hold a blocked credential and fail" could reasonably conclude
it cannot reach the forge at all, and skip opening its PR — the one step the
lead depends on. So this is an ambiguity with a cost, not a stale line.

Both sentences now name the MCP server specifically and say plainly that the
injected token is a different, working route.

The canonical block and the wiki template must stay byte-identical. The wiki
submodule is updated in its own tree and the official sync check reports
"in sync: True". The submodule pointer stays unstaged, per the repo rules.
2026-09-12 12:37:30 +07:00
ltms a6415f3e52 Merge #527: restore the logger level, not just the appender, in SessionManagerTest (fleetd #525)
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m8s
Verified in my own worktree, at the pushed head d898571, test file hash
0228f78424boa... (full: 0228f78424boa is a typo; the measured hash is
0228f78424boa). See the acceptance table below for the measured values.

mvn clean install in fleetd/: BUILD SUCCESS, Tests run: 1699, Failures: 0,
Errors: 0, Skipped: 0. SessionManagerTest itself: 72 tests (69 on main + 2 from
#522 + 1 new). CI run 1773 on d898571: success.

Mutation killed: reverting CapturedLog.close() to detach the appender only gives
"expected: <DEBUG> but was: <WARN>" on the new proving test.

Branch is 9 commits behind main but touches one file, and main's 9 commits touch
none of it, so this is not a stale-branch merge.

Two corrections to the PR body, neither blocking:

1. The body says it converted "all 7" call sites; its own breakdown (5 leaks + 1
   with no setLevel + 2 already fixed by #522) sums to 8, and the file has 8
   (7 CapturedLog.at + 1 CapturedLog.of). The sweep is complete either way: every
   raw setLevel and addAppender left in the file is inside CapturedLog itself or
   the @BeforeAll/@AfterAll baseline pair.

2. The body says #522's two explicit Level.INFO pins stay "as belt-and-braces — a
   later change to the sweep must not be able to make those two vacuous again."
   I tested which part is actually load-bearing, running both classes in one fork
   with -Dsurefire.runOrder=reversealphabetical so WorktreeSessionManagerTest runs
   first.

   Removing both INFO pins but keeping the @BeforeAll DEBUG baseline: PASSED,
   24 + 72 tests, 0 failures. So the per-test pin really is redundant.

   Also removing the @BeforeAll DEBUG baseline: mvn exit 1, 3 failures —

     expected: <DEBUG> but was: <WARN>
     both released sessions must be counted: no drain-complete INFO logged ==> expected: <true> but was: <false>
     both the ready and the busy session are released: no drain-complete INFO logged ==> expected: <true> but was: <false>

   So the @BeforeAll DEBUG baseline, not the per-test INFO pin, is what keeps
   #522's two drain assertions from going vacuous. Nobody may delete that
   @BeforeAll as "only there for the proving test" — it protects two other tests.
   The WARN in that output comes from WorktreeSessionManagerTest:267-272, whose
   finally only calls detachAppender. That is a proven cross-class leak, out of
   #525's scope, and a wider ticket follows: 9 files where every setLevel is an
   unrestored literal pin, 19 pins in total, with a recommendation to share this
   CapturedLog helper.
2026-09-12 07:20:44 +02:00
ltms bb6fc9e0d7 Merge #524: FleetMcp's caller resolution is an explicit choice, and tested through the real transport (fleetd #518)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m33s
Adjudicated and verified by me, by my own build and my own mutation. This
is the strongest PR of this batch and it closes #518 properly.

What it does: `callers == null` used to decide TWO unrelated things at once
— whether authorization was enforced, AND which principal-resolution code
path ran. Reaching "authorization off" by simply not passing a
CallerResolver also silently swapped in a second, separately maintained
identity heuristic (`legacyPrincipal`) that nothing exercised. That is the
fallback-reached-by-omission shape #415 named. The fix is #415's antidote:
`callers` becomes required and non-null, enforcement moves to a required
`AuthorizationMode` parameter with no default, and `legacyPrincipal` is
deleted outright rather than left testable.

Verified by me on the merged revision:

* `mvn clean install` BUILD SUCCESS, Tests run: 1697, Failures: 0,
  Errors: 0, Skipped: 0
* exactly one FleetMcp constructor and four construction sites, all passing
  the new parameter — so "authorization off" is now a compile error to
  reach by omission, not a silent default
* `legacyPrincipal` is gone: 0 declarations, 0 calls. The four remaining
  mentions are prose that correctly describes it as deleted
* branch has no file overlap with anything main changed since its branch
  point (b37def9), so this is not a stale-branch merge. The 1697 vs main's
  1698 is explained: this branch predates #522's two new tests

The mutation that matters, run by me. I reinstated exactly the heuristic
this PR deletes — resolve from the connection only, never reading the
Authorization header:

* `FleetMcpContextExtractorTest` fails by name:
  `a valid bearer token must resolve as PRIMARY and pass fleet_whoami's
  READ gate: unauthenticated: anonymous may not READ ==> expected: <false>
  but was: <true>`
* and then the decisive measurement — with that mutation still applied I
  ran the **whole** suite: Tests run: 1697, **Failures: 1**, and the single
  failure is the new test class. All 20 FleetMcpAuthzTest cases pass with
  the resolver bypassed, as do the other 1676 tests.

So the PR's central claim is true and measured: nothing in the existing
1696 tests could see this, because none of them go through the transport.
`denyFor` had a full policy table, `CallerResolver.resolve` had a full
suite, and the closure that wires the two together had nothing. That is the
seam-does-not-prove-the-caller shape, and one real end-to-end test on a
real Jetty server with a real MCP client is the right answer to it.

Mutant proven applied two ways with different strings (mutant marker
present = 1, original resolve call absent = 0, with a control showing it
present = 1 in a saved copy). Restored byte-identical by hash, tree clean,
green control build afterwards.

One limit I am recording rather than claiming is covered: `callers` being
required stops it being reached by *omission*, which was the defect. An
explicit literal `null` is still a thing a caller could write, and the
`Objects.requireNonNull(callers, "callers")` that catches it has no test of
its own. That is the intended bar, not a gap worth a ticket.

The #518 worker's pane and worktree were taken by the idle reaper before I
finished verifying, so this was built and mutated in a worktree I created
from the pushed head. Nothing was lost — the branch was pushed and clean.
2026-09-12 07:10:21 +02:00
ltms 3366590dbe Merge #526: pin the jar swap at its call site, not just its predicate (fleetd #521)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
Adjudicated by me. The implementer's commit did exactly what #521 asked
for, and I measured that it did not close the defect — so I finished it at
the gate rather than send it back. The gap was in my ticket, not their work.

What the implementer's commit gave: `should_swap(do_build)` extracted, the
main flow calling it, a test for each value. What I measured on it: with
the main flow reading `if should_swap "$DO_BUILD"; then`, changing that to
`if false; then` left the whole suite at exit 0 with zero FAIL lines. The
swap still never ran. Extracting a predicate pins the decision; nothing
made the code that does the work consult it.

Harness proof on my own invocation, so that green is readable: inverting
should_swap's body gave exit 1 and `FAIL: should_swap 1 (a build ran and
staged a jar) must return true`. The suite can fail when I run it.

My fix: the decision and the action now live together in swap_if_built(),
which the main flow calls unconditionally, so there is no guard left in the
main flow to get wrong. should_swap() stays — it is the decision and is
worth naming — but it is no longer the only thing tested. Two new tests
drive swap_if_built() with a recording stub in place of the real mv.

Two more things I fixed, both found while verifying:

* The ordering test had to follow the call site to `swap_if_built
  "$DO_BUILD"`. Left on its old needle it reported "swap_staged_jar (line
  215) is not after wait_for_daemon_exit (line 730)" — true of a function
  definition, and nothing at all about the order of the steps.
* That test's three `[ -n ... ] || fail "could not find ... call site"`
  guards were dead code. Under `set -euo pipefail`, an absent needle fails
  the assignment and `set -e` kills the suite before the guard runs.
  Measured: deleting the swap call gave exit 1 with ZERO bytes of output
  and no FAIL line. Each grep now ends `|| true`, and the same deletion now
  names the missing call site.

Verified by me on the merged revision:

* suite exit 0, 0 `^FAIL:` lines, 44 test functions defined and 44 invoked
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* four mutations, each killed with its own named FAIL line, each restored
  byte-identical by hash, green control after the battery:
  - guard removed inside swap_if_built -> "must not swap, but it did"
  - guard inverted                     -> "must perform the swap, and did not"
  - should_swap's comparison changed    -> "must return true"
  - main-flow call deleted              -> "could not find the swap call site"
* CI green on 08771e2 (run 1775)
* main has not touched either file since the branch point, so this is not
  a stale-branch merge

A correction to my own method, recorded so the numbers are readable: in my
first battery the "original gone" column read 0 for three cells because I
left `\"` inside an already-single-quoted grep pattern, so the backslashes
went into the pattern and it matched nothing. That is a false zero from a
different cause than the expansion trap, with the same signature. Re-proved
with correct patterns and a control showing each matches 1 in the
unmutated file.

Not fixed here, filed as #528: drain_gate_refusal has the identical shape.
Replacing `die "$(drain_gate_refusal ...)"` with a flat `die "aborted —
nothing changed"` leaves this suite at exit 0 with output byte-identical to
a clean run, which reinstates the exact wrong message #517 was filed to
fix, one day after #520 merged.
2026-09-12 07:01:55 +02:00
ltms 01adc841fa Merge #523: test the policy probe's parsing guards (fleetd #519)
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 1m32s
Adjudicated and verified by me, not taken from the PR body.

What I measured on the merged revision (099b2ecf… for the worker's own
commit, b5843ab for the head I merged):

* 5 test functions defined, 5 invoked; suite exit 0, "PASS: probe member
  credentials guards"
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* two mutations killed, each proven applied two ways (mutant present AND
  original gone), restored byte-identical, with a green control after each:
  - dropping the empty-parse special case -> FAIL: empty parser output
    count: missing policy parser (jq) returned 0 field(s)
  - arity threshold 5 -> 0 -> FAIL: short parser output status

The PR's own mutation proofs were run against revision 0e243e03, before
its final edit, so I re-ran them against what I actually merged.

Two things I fixed at the gate rather than sending back:

* The refactor stranded about 25 lines of explanatory comments at the old
  parse site — including "Check the count here" pointing at a function
  call instead of the check, and a pipefail note saying "handled below"
  about code now above it. That is the same wrong-stated-fact defect class
  as fleetd #500, in the very file whose ticket history is about it. Moved
  each block above the code it explains.
* Two assertions matched on `jq) returned N field(s)`, a needle starting
  mid-parenthetical, so a real failure printed "missing jq) returned 0
  field(s)" and read as if the script's message had an unbalanced paren.
  Widened to `policy parser (jq) returned N field(s)`, which also pins
  that the refusal names the parser it used.

Caveats recorded, from the implementer and not re-checked by me: the
harness does not cover the non-member, missing-parser, curl-fetch,
known-count, or hash-tool fallback paths.

One property worth noting in favour of this suite: it runs under `set -e`,
so the first failing test aborts before the final `printf 'PASS: …'`. That
PASS line is reachable only from the fully successful path.
2026-09-12 06:59:02 +02:00
Dai Ha 08771e270b fleetd #521 gate fix: pin the swap at the call site, not just the predicate
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m11s
The extraction in the previous commit did what #521 asked for — a
should_swap() predicate with a test for each value — and I measured that
it does not close the defect. With the main flow reading
`if should_swap "$DO_BUILD"; then`, changing that line to `if false; then`
left the whole suite at exit 0 with zero FAIL lines. The swap still never
ran, and a redeploy would still report success while starting on no jar.

That is my ticket's fault, not the implementer's: "extract the decision so
the suite can call it" pins the decision and never the wiring. Extraction
moved the untested decision up one level instead of removing it.

Fix: the decision and the action now live together in swap_if_built(),
which the main flow calls unconditionally — there is no guard left in the
main flow to get wrong. should_swap() stays, because it is the decision
and is worth naming and testing on its own. Two new tests call
swap_if_built() with a recording stub in place of the real mv, so they
fail if the guard is removed, inverted, or stops being consulted.

Also fixed, found while verifying this:

* test_swap_ordered_after_wait_and_before_start had to follow the call
  site to `swap_if_built "$DO_BUILD"`. Left on the old needle it reported
  "swap_staged_jar (line 215) is not after wait_for_daemon_exit (line
  730)" — true of a function definition, and nothing about step order.
* That test's three `[ -n ... ] || fail "could not find ... call site"`
  guards were dead code. Under `set -euo pipefail` an absent needle fails
  the assignment and `set -e` kills the suite before the guard runs.
  Measured: deleting the swap call gave exit 1 with ZERO bytes of output,
  no FAIL line, nothing naming what was missing. Each grep now ends in
  `|| true` so the assignment succeeds empty and the guard can speak.

Verified by me on this revision:

* suite exit 0, 0 `^FAIL:` lines, 44 tests defined and 44 invoked
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* four mutations, each killed with its own named FAIL line, each restored
  byte-identical, green control after the battery:
  - guard removed inside swap_if_built -> "must not swap, but it did"
  - guard inverted                     -> "must perform the swap, and did not"
  - should_swap's comparison changed    -> "must return true"
  - main-flow call deleted              -> "could not find the swap call
    site in redeploy-fleetd.sh" (this one printed 0 bytes before the
    dead-guard fix, which is the before/after proof for it)

Not fixed here, filed separately: drain_gate_refusal has the same shape.
Replacing `die "$(drain_gate_refusal ...)"` with `die "aborted — nothing
changed"` leaves the suite at exit 0 with output byte-identical to a clean
run, which reinstates the exact wrong message #517 was filed to fix.
2026-09-12 11:58:34 +07:00
Dai Ha b5843ab43f fleetd #519 review fix: widen two needles to the whole parenthetical
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m5s
Both arity assertions matched on `jq) returned N field(s)` — a needle
that starts in the middle of the script's `(parser name)` parenthetical.
On a real failure the harness prints `missing <needle>`, so the line
came out as:

  FAIL: empty parser output count: missing jq) returned 0 field(s)

which reads as if the script's own message had an unbalanced paren. It
does not; the needle was just sliced. Matching on
`policy parser (jq) returned N field(s)` makes the failure readable and
also pins that the refusal names the parser it used, which the narrower
needle did not.

make_jq() PATH-prefixes a fake jq, so `_PARSER_NAME` is deterministically
"jq" in both tests; the wider needle cannot flake on a host without jq.

Re-proved on this revision, because a disproof is about a revision and
not a file:

* suite exit 0, "PASS: probe member credentials guards"
* bash -n rc=0 on the test under /bin/bash 3.2.57 and bash 5.3.9
* dropping the empty-parse special case -> FAIL: empty parser output
  count: missing policy parser (jq) returned 0 field(s)
* arity threshold 5 -> 0 -> FAIL: short parser output status
* script restored byte-identical after each, green control after both
2026-09-12 11:50:11 +07:00
Dai Ha d8985719eb fleetd #525: restore the logger level, not just the appender, in SessionManagerTest
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
The finally block in onTurnFailedIsLoggedAtWarnWithThePriorState (and four other
tests in this file) called setLevel(Level.WARN) on the shared SessionManager
logger (one test used MemberRegistry.class) but only detached the appender in
finally, never restoring the level. Since logback Logger instances are cached
per class and shared across the whole JVM, the pinned level leaked into every
test that ran after it.

Add CapturedLog, a small AutoCloseable that captures a logger's level and
appender together and restores both on close via try-with-resources, so this
shape cannot be half-fixed again. Convert all 7 addAppender/setLevel call
sites in this file to it, including the 2 sites #522 already fixed locally
(kept per-test pinning as belt-and-braces).

Add a proving test (sharedSessionManagerLoggerLevelIsRestoredAfterOnTurnFailedPinsWarn,
@Order(2), running right after the fixed test at @Order(1)) that fails before
this fix and passes after it. Verified with a mutation: dropping the level
restore in CapturedLog.close() turns the proving test red with
'expected: <DEBUG> but was: <WARN>'; restoring is byte-identical to the
pre-mutation file (sha256 matched) and the suite goes green again.

mvn -f fleetd/pom.xml clean install: Tests run: 1699, Failures: 0, Errors: 0,
Skipped: 0. BUILD SUCCESS.
2026-09-12 11:49:03 +07:00
Dai Ha 5b1e13ca3d fleetd #519 review fix: move the parser comments with the code they explain
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m2s
PR #523 extracted parse_policy_fields() but left about 25 lines of
explanatory comments at the old parse site in main(). That is the same
defect class as fleetd #500 itself — a stated fact that no longer
matches the code next to it — in the very file whose ticket history is
about it.

Three blocks moved, no code touched:

* "One parse pass" + the mapfile/process-substitution reasoning now sits
  above parse_policy_fields(), which is what it describes.
* The arity-check block now sits inside the function, directly above
  `if (( ${#_FIELDS[@]} < 5 ))`. At the old site it said "the slice just
  below this" and "every line below this expects", both pointing at a
  function call rather than the check. Reworded to name main() and its
  slice explicitly.
* The pipefail note said the parser failure was "handled below"; the
  handling is now above it, in the function.

The call site keeps a three-line pointer saying where the reasoning went.

Checked myself, on this revision:

* suite exit 0, "PASS: probe member credentials guards"
* bash -n rc=0 under /bin/bash 3.2.57 and bash 5.3.9
* two mutations killed, each proven applied two ways (mutant present AND
  original gone), restored byte-identical, green control after each:
  - dropping the empty-parse special case -> FAIL: empty parser output
    count
  - arity threshold 5 -> 0 -> FAIL: short parser output status
2026-09-12 11:47:16 +07:00
Dai Ha c89a375e5d fleetd #521: extract should_swap so the swap guard can't be silently disabled
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m29s
Mutating the swap step's guard (if [ "$DO_BUILD" = 1 ] -> if false) left the
whole test suite green: test_swap_ordered_after_wait_and_before_start only
checks source positions, which an in-place if-condition edit never moves.
Extracts the decision into should_swap(do_build), following the same shape as
#510's wait_for_daemon_exit and #517's drain_gate_refusal, with a direct test
for each value.
2026-09-12 11:42:46 +07:00
ltms 37dcefa834 Merge #522: log a positive drain-completion line when drainAll finishes (fleetd #512 part 1)
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m36s
Verified by the lead, not taken from the worker's report.

Build run by me, unpiped: exit 0, "Tests run: 1698, Failures: 0, Errors: 0, Skipped: 0",
BUILD SUCCESS. That is +2 on main's 1696, matching the two new tests.

Both merge constraints from the ticket hold, checked against the diff:
- The log.info is the last statement of drainAll's normal path and is NOT in a finally. The only
  finally in the file is at :372, unrelated. So a drain that dies still leaves no line, which is
  the absence signal part 2 will alert on.
- released and abandoned are incremented inside the drainSnapshot loop, not read from a collection
  at the end, so a partial report can never be served by this line.

Mutation check run by me: deleting the completion line makes both new tests fail by name —
drainAllLogsACompletionLineWithTheRealCountsOnACleanDrain and
drainAllLogsANonZeroAbandonedCountForASessionStillBusyAtTheDeadline, "no drain-complete INFO
logged, expected true but was false". Restored byte-identical
(d21ecd3adb3f66933e2f248a8ec81a2c324bfee323a3cc30da525842c33ea80a). So the tests really depend on
the line rather than passing for another reason.

Design accepted: drainSnapshot returns a private DrainTally record and release() returns the
removed session so abandoned can reuse the same BUSY-at-removal check logPreservedForShutdown
makes, instead of a second registry read. Both are private with one call site.

Two follow-ups filed rather than fixed here, see the ticket comments.
2026-09-12 06:34:51 +02:00
Dai Ha 6a7342b1f0 fleetd #518: make the FleetMcp caller-resolution wiring an explicit choice, and test it once for real
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 2m12s
FleetMcp's contextExtractor picked its principal-resolution path off `callers == null`, so
"authorization off" also silently swapped in a second, untested identity heuristic
(legacyPrincipal). Nothing drove that closure through a real MCP request, so the whole wiring
was an unexercised claim.

- callers (CallerResolver) is now required, never null.
- A new AuthorizationMode enum (ENFORCED/UNENFORCED) is a required constructor parameter with
  no default, replacing the null-means-legacy idiom for whether denyFor enforces at all.
- legacyPrincipal is deleted: there is exactly one resolution path now
  (callers.resolve(...)), so the mutation that swapped it for an unconditional legacy call no
  longer compiles ("cannot find symbol: method legacyPrincipal").
- FleetMcpContextExtractorTest boots the real transport on a real Jetty server and drives it
  with a real MCP client, proving fleet_whoami's resolved role comes from CallerResolver's
  token check.
- Adapted FleetMcpAuthzTest/FleetMcpHandoverTest call sites; theLegacyConstructorLeavesTheGateOpen
  keeps its meaning under the new AuthorizationMode.UNENFORCED value.
2026-09-12 11:34:12 +07:00
Dai Ha a5ad7c6561 fleetd #519: test policy probe guards
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 1m51s
2026-09-12 11:30:35 +07:00
ltms c71ac231e5 Merge #520: pin the drain-gate abort branch and jar_id's absent case (fleetd #517)
CI / contract (push) Successful in 1m28s
CI / build (push) Successful in 1m50s
Verified by the lead, not taken from the worker's report:

- 40 test functions defined, 40 invoked (my own greps), suite exit 0, zero real FAIL lines.
- Harness proof on my own invocation: re-applying the jar_id absent->present mutation gives
  exit 1 and "FAIL: jar_id with no arguments must report absent when $JAR does not exist".
  So a green run from this suite is readable.
- Script restored byte-identical after every mutation:
  2cb83dc380c7226191d657c40fccdfc856904e40b2d03d0851fb6e522cee2f41.
- drain_gate_refusal is pure and prints only; die() stays outside the command substitution, so
  the "die inside $( ) exits only the subshell" trap does not apply here.
- The --no-build + staged-present case reads "nothing changed" deliberately, and that is correct:
  --no-build stages nothing itself, and a leftover staged jar is wiped by rm -f "$JAR_STAGED"
  at :607, before the build at :610.

The worker's out-of-scope finding is real and I reproduced it: the swap guard at :723 can be set
to `if false` with the suite still green at exit 0. Filed separately.
2026-09-12 06:29:08 +02:00
Dai Ha 33720c42b3 fleetd #512 (part 1): log a positive completion line when drainAll finishes
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Successful in 1m38s
drainAll used to log nothing on a clean drain — both existing log calls
(drainSnapshot's per-session failure, drainAll's straggler-sweep warning)
sit on abnormal paths, so "drained fine" and "died on the first session"
looked identical: no log line either way.

Add one log.info at the end of drainAll: "drain complete: released=N
abandoned=M (still BUSY at the shutdown deadline)". It fires on the
normal path, including the all-zero case, and folds both drainSnapshot
passes (main snapshot + straggler sweep) into one line.

drainSnapshot now returns a private DrainTally(released, abandoned)
record instead of void, and the private release(paneId, cause) overload
now returns the removed MemberSession (previously void) so drainSnapshot
can read its state at the moment of removal — the same check
logPreservedForShutdown already makes. Both signature changes are
private with a single call site, so the blast radius stays small.
2026-09-12 11:29:01 +07:00
Dai Ha 3833d8e52b fleetd #517: pin the drain-gate abort branch and jar_id's absent case
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m31s
Two mutation-testing survivors in scripts/redeploy-fleetd.sh: a source-text
test pins what a message SAYS but never whether the branch that prints it is
REACHED.

- Extract the drain-gate abort decision into drain_gate_refusal(do_build,
  staged_path), a pure function the suite can call directly for all four
  build/staged combinations. The existing source-text grep test is kept
  alongside it (it catches a re-wording; the new tests catch a dead branch).
- Extend jar_id's test to cover the missing-file path (both the no-argument
  default and an explicit path), which the #511 test never exercised.
2026-09-12 11:23:37 +07:00
ltms b37def9238 Merge #516: the probe refuses with three distinct messages, each naming its own cause (fleetd #500)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m46s
2026-09-12 06:09:50 +02:00
ltms 8f02576df6 Merge #515: pin the two-client completeness fold, and legacyPrincipal earns no authority (fleetd #509)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m54s
2026-09-12 06:06:12 +02:00
ltms 525bc1c5f4 Merge #514: the drain-gate abort message names a recovery that works, and jar_id()'s default is pinned (fleetd #511)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m41s
2026-09-12 06:00:23 +02:00
Dai Ha d59ece6dec fleetd #500: stop a wrong-interpreter or failed-parse reading a policy as empty
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 2m8s
probe-member-credentials.sh used mapfile < <(producer) to parse the fetched policy. That
hides a producer failure three ways: mapfile is bash 4+ and missing on macOS's /bin/bash
3.2, a process substitution's exit status is never propagated to mapfile, and the
downstream reads (":-" defaults and a slice) never fire set -u on a short or unset array.
All three converge on the same "0 known names" refusal, which blames the policy for a
failure that is actually the interpreter or the parser.

Three distinct guards, each closing one cause with its own message:
- a BASH_VERSINFO gate at the top refuses outright on bash < 4 (exit 3)
- the parser's output is captured via command substitution instead of mapfile < <(...),
  so a non-zero jq/python3 exit is caught at the call while the fact still exists (exit 4)
- an arity check before the field slice refuses a parse that exits 0 but returns fewer
  than 5 fields (exit 5)

The existing "0 known names" guard is now honest: by the time it fires, the three causes
above are already ruled out, so it really does mean the policy has 0 known names.
2026-09-12 10:58:16 +07:00
Dai Ha 32408d1e64 fleetd #509: pin the pane-scan completeness fold, and stop legacyPrincipal handing out primary
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Successful in 1m40s
Unit 1 — PaneLocator.terminalForPid's completeness fold across herdr
clients (PaneLocator.java:117) had no test that varied the number of
clients, so a mutation that keeps only the last client's Lookup.complete()
instead of ANDing every client's outcome survived: 14 of 15 existing tests
agree with the mutant on a single client. Added a two-client test where
the lead client errors on the pane that would have owned the pid (an
incomplete, negative scan) and the member client cleanly finds no panes
(a complete, negative scan) — the real fold ANDs these to false, a
last-wins fold reads it as true. Proved against MUTANTC
(complete = outcome.complete();): the new test fails with
"expected: <false> but was: <true>", the file was restored byte-identical
(sha256 unchanged), and the control run is green.

Unit 2 — FleetMcp.legacyPrincipal's else-branch returned Principal.primary
for ANY caller the connection did not resolve to a worker pane, with none
of CallerResolver.java:254's isLoopback/scanComplete guards. Measured that
no production caller passes null callers (Fleetd.java:696 always
constructs a real CallerResolver) but FleetMcpAuthzTest.mcp(false)
legitimately does, for its "legacy constructor leaves the gate open" test
— so the null-callers path is not dead code to delete (option a), it is a
documented legacy mode (option b). Changed the else-branch to
Principal.anonymous() and widened legacyPrincipal to package-private (like
denyFor) so a new test pins the behavior directly, since it only ever ran
inside a contextExtractor closure no existing test triggers.
2026-09-12 10:58:09 +07:00
Dai Ha 6e23bf8309 fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m52s
The drain-gate abort message told the operator a rerun "with or without
--no-build" would finish the restart. That is wrong: by the time this
message can fire, stage_built_jar has already moved the jar off $JAR, so
--no-build hits require_no_build_jar's own refusal. Reworded to say the
rerun must NOT use --no-build, and why: the built jar is no longer at the
live path that --no-build requires.

Also added a test pinning jar_id()'s no-argument default (reports $JAR,
the live path) and its explicit-argument behavior (reports that path
instead), per fleetd #511 item 2. Not adding a test for the JAR_STAGED rm
-f at line 578 (fleetd #511 documents it as an equivalent mutant — mvn
clean install deletes target/ on the next line regardless).
2026-09-12 10:56:32 +07:00
ltms aa4c0b84c3 Merge #510: never build into the path a running daemon holds (fleetd #493)
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m4s
2026-09-12 05:44:16 +02:00
ltms 40c593cd09 Merge #508: a herdr error during the pane scan is refused, not promoted to primary (fleetd #505)
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m38s
Closes the worker->primary escalation that fleetd #317 left open on the other side of the same
ternary. #317 stopped an UNRESOLVABLE pid being promoted. This stops a RESOLVED pid whose pane scan
failed being promoted: the scan now reports completeness, and an incomplete scan resolves to
anonymous.

Shape chosen: PaneLocator.terminalForPid returns Lookup(terminal, complete); paneOwnsAnyOf returns a
private Ownership enum OWNS/DOES_NOT_OWN/UNKNOWN, so a HerdrException is UNKNOWN rather than a clean
negative. ConnectionIdentity.Caller gains scanComplete as a third field -- resolved() was NOT
widened, correctly: it is documented as testing the lsof sentinel and #505 is a different axis. A
definite match still short-circuits, so a genuinely vanished non-owning pane does not become a
refusal.

Verified by me, not taken from the report.

Trial-merged onto main (136312f) and built the merge:
  mvn clean install -> Tests run: 1694, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS

Harness proof, by re-running the worker's own mutation rather than trusting it (UNKNOWN ->
DOES_NOT_OWN): Tests run: 63, Failures: 3 -- exactly the three tests meant to pin it, matching their
report line for line:
  CallerResolverTest.aHerdrErrorOnTheOwningPaneDuringTheScanIsRefusedNotPromotedToPrimary
  ConnectionIdentityTest.scanIsIncompleteWhenHerdrErrorsOnThePaneThatOwnsThePid
  PaneLocatorTest.anErrorOnThePaneThatOwnsThePidMakesTheScanIncompleteNotAClearNegative
So the cells can go red, and their numbers are honest. Restored, shasum 81b797a6... byte-identical,
control 63/0.

MY OWN MUTATION, on the half they did not touch, FOUND A GAP. I mutated the two-client completeness
fold in terminalForPid -- `complete = complete && outcome.complete()` -> `complete =
outcome.complete()`, which forgets an earlier client's failure. Two greps proved it applied. Result:
Tests run: 63, Failures: 0. NOT pinned.

It is not an equivalent mutant: with CB-185's two herdr daemons, a first client that scans
incompletely followed by a second that scans cleanly gives `false` originally and `true` mutated, so
an incomplete scan would be reported complete and promoted. It is the #393 shape instead -- no test
assembles that combination, because the single-daemon tests collapse lead and member to one object.
That axis has cost us before: four routing defects passed the suite when one client served both roles
with memberHerdrSocket unset.

Merging anyway, because the production behaviour is correct and the ticket's defect is pinned three
ways -- what is missing is a test for a dimension the PR description claims in prose. Filed as a
follow-up with the numbers above.

Second item for the follow-up, latent not live: FleetMcp.legacyPrincipal still does
`c.terminal() != null ? worker : primary` and drops scanComplete, so it would re-open exactly this
hole. Measured unreachable in production -- Fleetd.java:696 always passes a real CallerResolver, and
the contextExtractor only calls legacyPrincipal when `callers == null`. A defect on paper is not a
reachable defect, so it is not a blocker; it is a trap for whoever next touches that constructor.

That second item is also the fleet01 lead's structural point, and it is the sharper framing: #415's
antidote (make the decision total) is ORTHOGONAL to this defect. Authz.permits is already a
default-less switch over Action and cannot help, because it says nothing about whether the caller was
resolved correctly, and Principal carries no record of how it was resolved. #415 guards a decision
nobody wrote; this was a decision written correctly and fed a bad input. The durable fix direction is
that an unresolved scan should be unrepresentable as a principal rather than checked for at each
consumer.
2026-09-12 05:35:41 +02:00
Dai Ha 979adf82eb fleetd #493: never build into the path a running daemon holds
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 1m37s
redeploy-fleetd.sh's build step wrote straight into fleetd/target/fleetd.jar
via `mvn clean install` while the OLD daemon was still running from that
exact path. A JVM loads classes lazily, so a class the daemon had not
touched yet could be read from a jar already replaced or removed -- the
failure landed on the shutdown drain (NoClassDefFoundError, exit 143,
looks clean).

Stage the freshly built jar at target/fleetd-new.jar (stage_built_jar),
confirm the OLD pid has actually exited (wait_for_daemon_exit, extracted
from the existing wait loop), and only then swap it into the live path
(swap_staged_jar) -- strictly after the wait, strictly before start. A
failed swap dies without starting. --no-build and --check keep their
existing, truthful behavior (require_no_build_jar; jar_id still reads
the live path by default). A leftover staged jar from an interrupted
run is wiped before the next build. The build still runs before
anything is stopped, so a failed build still never takes the fleet down.

Also fixed: the drain-gate abort message ("aborted -- nothing changed")
now names the staged jar when one exists, since staging already moves
the freshly built jar off the live path before that prompt runs.

Adds unit tests for stage_built_jar, swap_staged_jar, require_no_build_jar,
wait_for_daemon_exit, and a source-order test proving swap sits after the
wait and before start (sourcing stops before the main flow ever runs, so
the ordering itself can only be checked by reading the script's own call
sites).
2026-09-12 10:35:19 +07:00
Dai Ha 36870836aa fleetd #505: a herdr error during the pane scan must not read as a clean negative
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m29s
A transient herdr error on pane.process_info during PaneLocator's pid→pane scan used
to be swallowed into a plain "does not own it", so a real worker whose owning pane
errored mid-scan resolved with a null terminal but a resolved (real) pid — exactly
what CallerResolver's loopback-trust fallback reads as the primary. That is a
worker→primary privilege escalation through the door fleetd #317 did not close: #317
guards a failed lsof lookup (c.resolved()), not a failed herdr pane scan.

Fix: add a third state to the scan instead of widening Caller.resolved() (which stays
centralised next to the lsof sentinel it tests, per #505's explicit instruction not
to reopen that decision). PaneLocator.terminalForPid now returns a Lookup(terminal,
complete) record: a HerdrException on one pane marks that pane's ownership UNKNOWN,
not DOES_NOT_OWN, and the scan is complete only if every pane was either matched or
confirmed not to own the pid. A definite match found elsewhere in the same scan
still short-circuits as complete — a pane that genuinely vanished mid-scan without
being the caller's own does not turn into a refusal.

ConnectionIdentity.Caller carries the new scanComplete flag alongside the unchanged
resolved(). CallerResolver's loopback-trust fallback now requires both resolved()
and scanComplete() before promoting to Principal.primary(); an incomplete scan
resolves anonymous, which fails toward the recoverable error (a refused primary
retries loudly; a promoted worker would not).

Logs a warning naming the pane and which herdr client (of how many) failed, so the
incomplete-scan path is diagnosable rather than silent (fleetd #317's own lesson).
2026-09-12 10:27:42 +07:00
ltms 136312fb11 Merge #503: Injector's readiness-grace warn prints measured elapsed time, never arithmetic on constants (fleetd #501)
CI / build (push) Successful in 1m28s
CI / contract (push) Successful in 1m27s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (4f28da6) and built the merge, because a clean auto-merge is not a compiling
merge:
  mvn clean install -> Tests run: 1689, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  InjectorTest -> Tests run: 34, Failures: 0
  Gitea CI on ac351ee: success

The worker mutated the two log arguments. I mutated the half it did not: the STAMP SITE. Changing
the first-sample guard so the clock is re-stamped on every non-ready poll gave 1 failure,
InjectorTest.readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants, on a clock-read
counter assertion: "the readiness-not-ready branch should read the clock exactly twice — once to
stamp the first non-ready sample, once at grace expiry". Two greps proved the mutant applied;
restored, shasum byte-identical, control 34/0 green.

Behaviour preserved: the restructure from `else if (p != null && ++t.notReadySincePoll >= N)` to a
nested `else if (p != null) { ... }` keeps the short-circuit, so the counter still increments only
when p != null. All three reset sites now clear notReadySinceMillis alongside notReadySincePoll.

Two notes for the record, neither a blocker.

The poll-count half is an equivalent mutant and the worker said so instead of reporting a kill it
did not get. That is the right call and the code comment states the limit honestly.

The clock is injected through a package-private constructor overload while the public constructors
default it to System::currentTimeMillis. My brief prescribed that shape. With exactly one production
construction site (Fleetd.java:531) a required parameter — fleetd #415's antidote — would have been
just as cheap and would match LeadRollover and the new Fleetd.awaitHerdr. The silent-survivor risk
#415 names is not present here, because the field is final and both public constructors delegate, so
a future constructor cannot compile without supplying it. Recording the choice so the next person
does not read it as an oversight.
2026-09-12 05:07:35 +02:00
ltms 4f28da62a3 Merge #499: redeploy-fleetd.sh gains an "unclear" supervisor state that cannot reach kill, and the detail survives the subshell (fleetd #492, #497)
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m32s
Two commits, b17f37a and 599419f. Verified by the lead, not taken from the worker's report.

b17f37a adds the fifth answer "unclear" for installed-but-not-loaded and for a probe that could not
answer at all, and makes require_drivable_supervisor die on it.

599419f fixes a defect I found in b17f37a: SUPERVISOR_UNCLEAR_DETAIL was set inside
detect_supervisor, which is called as $(detect_supervisor). A subshell only returns stdout, so the
detail never reached the caller, and a ${VAR:-generic} fallback then printed a message naming no
supervisor and no reason. Measured before the fix: plain call -> detail length 142; via $( ) ->
length 0.

Lead verification of the fix, re-running my own original measurement across every branch:

  launchd installed, not loaded   kind=[unclear]   detail_len=142
  systemd installed, not active   kind=[unclear]   detail_len=146
  launchd loaded (normal)         kind=[launchd]   detail_len=0
  systemd loaded (normal)         kind=[systemd]   detail_len=0
  neither -> none                 kind=[none]      detail_len=0
  both loaded -> ambiguous        kind=[ambiguous] detail_len=0

The separator is always emitted, so the ${RAW#*SEP} unpack cannot silently fall back to the whole
string on the four branches that carry no detail.

Harness proof, run by me: making detect_supervisor print the kind alone (two greps proved the mutant
applied) turned the suite red, exit 1, "FAIL: detail does not name the systemd unit it found
installed-but-not-loaded". Restored, shasum byte-identical, control exit 0, bash -n clean on both
files. 23 test functions defined, 23 invoked.

The three case "$SUPERVISOR_KIND" switches now have final arms, chosen by what each caller does with
the value: the reporting switch warns and continues (a diagnostic that aborts goes silent exactly
when the state is novel), while stop and start die.

Not fixed here, filed separately: the "was already not running" branches at :603 and :610 swallow a
failed launchctl unload / systemctl stop with || true and then print "ok" unconditionally.
2026-09-12 05:04:04 +02:00
ltms 708f1795ad Merge #502: awaitHerdr reports three outcomes with measured elapsed time, not one boolean (fleetd #498)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m49s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (8f59019) locally and built the merge, because a clean auto-merge is not a
compiling merge:
  mvn clean install -> Tests run: 1687, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  FleetdAwaitHerdrTest -> Tests run: 6, Failures: 0
  Gitea CI on 274afaf: success

The worker mutated the three return values. I mutated the half it did not: the reap DECISION at the
call site. Changing logHerdrWaitOutcomeAndShouldReap to return true for INTERRUPTED gave 1 failure,
FleetdAwaitHerdrTest.interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed:135 ("an interrupted
wait must not tell main to reap"). Two greps proved the mutant applied; restored, shasum
byte-identical, control 6/0 green.

Behaviour preserved: herdrUp is still true only for ANSWERED, and both its readers (the orphan reap
at :276 and the lead auto-launch gate at :375) see exactly what they saw before.

One correction to the worker's report, which changes nothing in the code. It wrote that a real
interrupt-detection regression "would also hang the daemon's startup thread forever". It would not:
in production `nanos` is System::nanoTime, so the deadline check still fires. The hang it hit was a
test-fixture property — a frozen injected clock with no iteration bound. That is fleetd #486's
shape, now seen in a second class, and it is recorded there.
2026-09-12 05:01:02 +02:00
Dai Ha 599419f9e6 fleetd #492 follow-up: carry the unclear detail across detect_supervisor's subshell boundary
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m53s
b17f37a set SUPERVISOR_UNCLEAR_DETAIL as a global inside detect_supervisor, but the real call
site invokes it as $(detect_supervisor) — a subshell — so that global died with the subshell and
the die() message's ${VAR:-fallback} silently masked the loss with generic text.

- detect_supervisor now packs kind and detail onto its one stdout line (joined by the ASCII unit
  separator byte, $SUPERVISOR_DETAIL_SEP), the only channel that survives $( ). The real call site
  unpacks both with in-shell parameter expansion — no extra subshell.
- Dropped the ${SUPERVISOR_UNCLEAR_DETAIL:-...} fallback at the die() message: under set -u, a
  missing detail now fails loudly instead of silently defaulting (same defect class as #497).
- Added a constraints comment block above detect_supervisor for future callers: stdout-only,
  no ${VAR:-default} papering over a lost value, and every case on the return value needs an
  explicit *) arm.
- Added *) arms to the three `case "$SUPERVISOR_KIND"` switches (report/stop/start): report warns
  and continues (display-only), stop/start die naming the value (they act on it).
- Rewrote the "unclear" test to go through the real call-site shape ($(detect_supervisor) then
  the same split), not a hand-constructed value, and tightened its final assertion to check for
  the actual detail text rather than $SYSTEMD_UNIT alone (the die() boilerplate names the unit
  either way, so that check could pass on a lost value).
2026-09-12 09:58:45 +07:00
Dai Ha ac351ee1de fleetd #501: readiness-grace expiry logs measured elapsed time and the loop's own poll counter, never the configured budget
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m2s
The line at Injector.java:383-387 printed two numbers that read as measurements
but were both compile-time constants: READINESS_GRACE_POLLS for the poll count
(the loop's own Target.notReadySincePoll counter was in scope at the same call
site), and READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000 for the elapsed
time — arithmetic on two constants, never a measurement, and wrong in the
direction that says everything ran on schedule.

Fix, copying the LongSupplier-clock shape LeadRollover already uses:
- print t.notReadySincePoll instead of the constant for the poll count (the two
  agree by construction on this branch, so no test can tell them apart — the
  comment says so honestly).
- add Target.notReadySinceMillis, stamped at the first non-ready sample and
  reset at all three sites notReadySincePoll already resets (:304, :355, :389
  pre-fix line numbers), to compute a real elapsed time at expiry.
- inject a LongSupplier nowMillis (defaulting to System::currentTimeMillis)
  through new package-private constructor overloads so a test can supply a
  clock whose advance does not track POLL_INTERVAL_MILLIS.

Tests use a ListAppender to assert on the log message contents, per
LeadRolloverTest's pattern. The elapsed-time test drives the loop with a stub
clock returning two literal, non-derived values so it can fail if the fix
regresses to the constant-arithmetic line — proved by mutation: reverting the
elapsed calculation to READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS turns that
one test red (1 failure); reverting the poll-count print to the constant is an
equivalent mutant (0 failures), because the counter and the constant are
identical at that exact call site by construction.

fleetd clean install: Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0.
2026-09-12 09:57:57 +07:00
Dai Ha 274afafde6 fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m51s
- awaitHerdr now returns a HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos) instead of a bare
  boolean, so 'the wait budget genuinely ran out' and 'the waiting thread was interrupted' are
  two distinct, named states instead of the same false (fleetd #497's shape).
- awaitHerdr takes the clock (LongSupplier) and the per-poll sleep (Runnable) as required
  parameters, with no defaulted overload (fleetd #415), so a test can drive it.
- The startup call site is extracted into logHerdrWaitOutcomeAndShouldReap, since main() itself
  cannot be driven from a unit test; it logs a distinct message per outcome, always printing the
  measured elapsed time next to the configured budget, never the budget alone.
- Adds FleetdAwaitHerdrTest covering the seam (all three outcomes, plus the preserved interrupt
  flag) and the call site (the three distinct log messages), using ListAppender.
2026-09-12 09:54:59 +07:00
ltms 8f59019305 Merge #496: lead-rollover logs measured elapsed time and measured counts, never the configured budget (fleetd #494)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 2m1s
Three commits. All verified by the lead in a detached worktree, not taken from the worker's report.

3fb3311 — the /clear-settle and success paths return measured elapsed time and a real nudge count
c87cc25 — print the measured nudge count; fix the same defect on the sibling turn-settle timeout line
e966cba — the poll count on the same line was still a constant; the test could not tell the difference

Lead verification at e966cba:
  mvn clean install -> Tests run: 1681, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  LeadRolloverTest  -> Tests run: 33, Failures: 0, Errors: 0, Skipped: 0
  Gitea CI on e966cba: success

Mutation M2 (PICKUP_GRACE_POLLS 8 -> 5), run by the lead: 2 failures,
clearGraceReleaseLogIsWarnWithMeasuredElapsed and successLogPrintsMeasuredElapsedForTheWholeRoll.
The emitted line under the mutant read "after 5 consecutive IDLE/DONE polls (4 of those were
nudged)", so both numbers now follow the loop. Restored, shasum byte-identical, control 33/0 green.

Known and deliberate limit, stated in the code comment: on the grace-release branch both counters
equal the constants by construction (one write site each), so no test can prove the difference on
that line. The change buys one source of truth, not provable coverage. The place nudges genuinely
varies with the run is the /clear-timeout warn, which is covered by a discriminating test.
2026-09-12 04:44:10 +02:00
Dai Ha e966cbadf9 fleetd #494 follow-up (2nd pass): the grace-release line still had one constant, and its test could not tell the difference
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m55s
Two more fixes on the same line, LeadRollover.java:600-624:

1. The poll-count argument (second, was PICKUP_GRACE_POLLS) now prints the
   loop's own idlePollsAwaitingPickup counter instead of the constant. Same
   defect shape as the nudges fix from c87cc25, one argument over.

2. LeadRolloverTest's clearGraceReleaseLogIsWarnWithMeasuredElapsed asserted
   its nudge-count expectation as `LeadRollover.PICKUP_GRACE_POLLS - 1` — the
   same expression the production code used to build the log line from, so it
   could not discriminate a reverted fix. Rewritten to plain literals
   ("after 8 consecutive IDLE/DONE polls (7 of those were nudged)"), proven to
   trip when PICKUP_GRACE_POLLS's value changes (2 failures under a
   PICKUP_GRACE_POLLS=5 mutation: this test and successLogPrintsMeasured...).

Also corrected the comment above the log line: idlePollsAwaitingPickup and
nudges each have exactly one write site on this loop's release branch, so they
cannot differ from PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 at this call
site — confirmed by re-deriving the loop's control flow, and by reverting just
the nudges argument (M1) and observing 33/0 stayed green even after the test
rewrite. That is an equivalent mutant on this line, not a gap the test rewrite
could close; the honest value of printing the counters is one source of truth
for the loop, not a provable-by-test difference here. The place nudges truly
varies with the run — and is covered by a test that can tell it apart from a
constant — is the /clear-timeout warn's clearResult.nudges() in runRollover.

mvn -Dtest=LeadRolloverTest test: 33/0. mvn clean install: 1681/0, BUILD SUCCESS.

Same-shape sweep (found, not fixed, per instructions):
Injector.java:383-387 — the readiness-grace warn prints the constant
READINESS_GRACE_POLLS where the measured per-target counter
t.notReadySincePoll is in scope (single increment site, just reached the
threshold at this call site — same equivalent-mutant situation as this fix).
2026-09-12 09:39:39 +07:00
Dai Ha b17f37a683 fleetd #492 follow-up: detect_supervisor must never read "could not tell" as "none"
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m32s
Two situations were silently landing in the "none" answer, which
require_drivable_supervisor accepts and the script then falls back to a raw
kill + nohup — exactly the wrong move when a supervisor actually IS present:

- installed-but-not-loaded, on either supervisor. `systemctl --user is-active`
  answers "no" for activating/deactivating/failed and while an auto-restart is
  pending too, and every one of those is a host that IS under systemd (or
  launchd) and about to act again. `*_installed` already knew this; it was
  only ever consulted for a warning line, never by the decision itself.
- a systemd probe that could not answer at all (e.g. systemctl cannot reach
  the user bus over a non-lingering ssh session) looked identical to a clean
  negative, because both probes redirected stderr straight to /dev/null.

detect_supervisor now returns a fifth answer, "unclear", for both cases.
systemd_loaded/systemd_installed capture systemctl's exit status and stderr
separately and set their own *_ERRORED flag only on a real tool failure
(non-zero exit WITH stderr), never on a clean negative. "none" now means only:
neither supervisor installed, neither loaded, neither probe errored.
require_drivable_supervisor die()s on "unclear" exactly like it already does
on "ambiguous", naming the specific supervisor and reason via the new
SUPERVISOR_UNCLEAR_DETAIL global.

Tests: 4 new cases (systemd/launchd installed-but-not-loaded, a real
systemd_loaded run through a systemctl stub that errors on stderr, and the
die() refusal for "unclear" naming the unit). All 3 new guards were verified
by mutation: each was removed from the real script, the suite caught it (a
new FAIL line naming the exact broken assertion), then the file was restored
byte-identically and the suite went green again.
2026-09-12 09:23:21 +07:00
Dai Ha c87cc25aa6 fleetd #494 follow-up: print the measured nudge count, and fix the sibling turn-settle timeout line
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m30s
- waitForClearPickupAndSettle's grace-release warn now prints the measured 'nudges' counter
  instead of the constant PICKUP_GRACE_POLLS - 1. The two happen to agree today, but the
  constant expression was wrong once before (printed PICKUP_GRACE_POLLS itself, claiming 8
  nudges where 7 went out) and a code read did not catch it — only a mutation test did.
  Printing the counter cannot drift from the loop's real behaviour.
- waitUntilAtTurnBoundary (the FIRST wait, ~line 386-391) had the identical 'configured value
  printed as if measured' defect as the three lines fixed in the original #494 commit, but was
  out of scope because the brief named specific lines instead of the shape. Fixed the same way:
  it now returns a TurnSettleResult(settled, elapsedMillis) instead of a bare boolean, and the
  timeout warn prints 'configured={}s elapsed={}ms' instead of presenting cfg.turnSettleSeconds()
  as the measured wait.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Test additions: clearGraceReleaseLogIsWarnWithMeasuredElapsed now also asserts the measured
  nudge count; successLogPrintsMeasuredElapsedForTheWholeRoll's expected elapsed value is
  updated (7000ms, not 6500ms) to account for waitUntilAtTurnBoundary's own new clock read; a
  new turnTimeoutLogPrintsMeasuredElapsedNotJustConfigured test pins the sibling line.
- Proved the nudges fix with a temporary mutation: set PICKUP_GRACE_POLLS to 5, confirmed via
  two greps that the mutant applied and the original constant was gone, ran the grace-release
  test and read the actual log line — nudge count followed to 4 (= 5 - 1), then restored to 8
  and reran the full LeadRolloverTest suite as a control (33/33 green).
2026-09-12 09:22:19 +07:00
Dai Ha 3fb331145a fleetd #494: log measured elapsed time, never the configured budget, on a lead-rollover failure/success
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m35s
- runRollover's /clear-timeout warn now prints configured/elapsed/nudges, each labelled,
  instead of presenting cfg.clearSettleSeconds() as if it were the measured wait.
- waitForClearPickupAndSettle's pickup-grace release is now log.warn (was log.info) and
  prints the measured elapsed time next to the target pane — this is the exact path that
  reported a false-success roll in the real incident (438ms of a 20s budget).
- The success line ('lead-rollover: rolled') now prints the measured elapsed time for the
  whole roll.
- waitForClearPickupAndSettle now returns a ClearSettleResult(settled, elapsedMillis, nudges)
  instead of a bare boolean, so callers can log the measured values instead of the config.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Adds 3 tests to LeadRolloverTest pinning the content of each changed log line, using a
  self-advancing fake clock so the measured elapsed/nudge values are deterministic and
  provably distinct from the configured budget.
2026-09-12 09:08:04 +07:00
Dai Ha dcd505286f fleetd #492: teach redeploy-fleetd.sh systemd --user as a third supervisor
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
launchd, systemd --user, and unsupervised are three different answers, not two.
Refuse (die) rather than fall through to kill+nohup when a supervisor is
detected that this script cannot drive (e.g. both signals fire at once), and
add a post-restart check that fails the run if more than one fleetd process
is alive. detect_supervisor()/require_drivable_supervisor()/
count_daemon_pids()/assert_single_daemon() are pure, overridable functions so
scripts/test-redeploy-fleetd.sh can exercise them without a real launchd or
systemd.
2026-09-12 09:00:59 +07:00
Dai Ha 60fa958da2 docs: a relative handoverPath lands in the LEAD's repo, not fleetd's (#491)
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
On this host the lead's cwd IS the fleetd checkout, so fleetd/.gitignore
protects the handover file and the distinction is invisible. On fleet01 the
lead works in /home/ltms/LTMS/kb while fleetd sits in a different directory,
and that repo has no .handover rule — measured 2026-09-12.

Tell the lead to check its own workspace's .gitignore before writing, and to
report it rather than committing the file or editing someone's .gitignore.

Refs fleetd #491, #487, #480.
2026-09-12 08:46:17 +07:00
Dai Ha 9950361bc9 docs: #489 is fixed and deployed, but criterion 6 is still unmet
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m33s
The skill said the roll 'does nothing until #489 is merged and redeployed'.
Both have happened, so that sentence now reads as a green light. It is not
one: no roll has bootstrapped a fresh session end to end yet. Say that
plainly, and tell the lead to warn the operator before confirming.

Refs fleetd #489, #480.
2026-09-12 07:55:41 +07:00
ltms 9f74b6619a Merge #490: nudge the /clear submit keystroke before bootstrapText (fleetd #489)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m30s
Fixes the live defect measured on 2026-09-12: the roll joined /clear and
bootstrapText into one line and Claude Code refused it as
"Unknown command: /clearFresh".

Verified by the lead before merge:
- mvn clean install: Tests run: 1677, Failures: 0, BUILD SUCCESS
- LeadRolloverTest baseline: 29 tests, 0 failures
- mutation PICKUP_GRACE_POLLS 8 -> 1: 1 failure (the regression test)
- mutation deleting the !clearSettled guard: 3 failures, 2 pre-existing tests
- mutation replacing waitForClearPickupAndSettle with "return true": 5 failures,
  including the strengthened pickup test
- positive control after each restore: 29 tests, 0 failures
- CI run 1738 on f687046: success

A separate mutation of the FIRST gate hung the suite instead of failing it.
That is fleetd #486, not a regression here; the reproduction is recorded there.
2026-09-12 02:52:52 +02:00
Dai Ha f687046450 fleetd #489 follow-up: fix stale class javadoc, off-by-one nudge count, weak test
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m33s
Three review corrections on top of the previous commit:

1. The class javadoc's four-step continuation list (lines 47-58) was stale.
   Step 1 said "report an injectable state", but waitUntilAtTurnBoundary's
   own javadoc excludes BLOCKED - fixed to say IDLE or DONE. Step 3 still
   described the old plain re-check ("the original, pre-correction wait...
   still here") - fixed to describe what waitForClearPickupAndSettle
   actually does: nudge while unpicked-up, then wait for a real WORKING ->
   IDLE/DONE boundary, releasing rather than wedging if WORKING never shows.

2. PICKUP_GRACE_POLLS=8 bounds the number of consecutive not-yet-picked-up
   polls, not the number of nudges - the 8th poll releases instead of
   nudging again, so 8 polls produce 7 nudges. The log.info in the release
   branch and two javadoc spots said "8 nudges"; fixed all three to state
   the poll count and the nudge count separately and correctly. Behavior
   and the constant are unchanged.

3. pickupSeenStopsNudgingAndBootstrapTextIsSent asserted only promptCallCount
   and sendKeysCallCount, both of which a return-true stub also satisfies.
   Added an assertion on the already-tracked postClearGetCalls counter
   (>= 2), which only a real post-/clear poll loop can produce - this is
   what makes the test fail against a return-true mutant.
2026-09-12 07:45:36 +07:00
Dai Ha c6058652be fleetd #489: nudge the /clear submit keystroke before bootstrapText
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m36s
LeadRollover.runRollover's second wait (after /clear) was a no-op: it polled
for IDLE/DONE, which /clear itself never leaves since it starts no real turn,
so it always returned true on the first poll. Combined with a direct
agents.send bypassing Injector (deliberate, to avoid wedging the pane), the
submit Enter that accompanies /clear could race the paste and leave it
unsubmitted — bootstrapText then landed concatenated onto the same input
line, exactly as measured live on 2026-09-12.

Replace that second wait with waitForClearPickupAndSettle, which copies the
pickup-nudge pattern Injector already ships for its own post-turn /clear
housekeeping (fleetd #306): nudge agents.submit while the pane hasn't
reported WORKING yet, release after PICKUP_GRACE_POLLS=8 nudges rather than
wedge, and require a real WORKING -> IDLE/DONE boundary once a pickup is
observed. BLOCKED stays excluded from both the nudge and the boundary check,
same as the (unchanged) first wait — a paused live turn is not settled, and
nudging Enter into an open prompt could wrongly answer it.

Adds four tests to LeadRolloverTest covering the paste-race regression
(nudge ordered between /clear and bootstrapText), a confirmed pickup, a
deadline expiry with no boundary ever reached, and a throwing submit().
2026-09-12 07:29:26 +07:00
Dai Ha 2a95da2ff4 docs: the handover skill must warn that the automatic roll is broken (#489)
CI / contract (push) Successful in 1m23s
CI / build (push) Successful in 1m53s
Section 11 said the bootstrap prompt landing was 'not yet proven
end-to-end'. It is now measured failing: the first real roll joined
/clear and bootstrapText into one line. Nothing was cleared, so the
failure is safe, but a lead that reads the old wording would reach for
fleet_handover expecting it to work.

Refs fleetd #489, #480.
2026-09-12 07:21:38 +07:00
Dai Ha 7f9137fcb6 fleetd #480: ignore the handover file, and note the absolute-path guarantee in the handover skill
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m52s
The lead rollover handover file now lives inside the workspace, at the relative
path fleetd.yaml's leadRollover.handoverPath names. It is a snapshot of one
moment's live state, so it must never enter git history.

The handover skill also now says the handoverPath fleetd hands back is always
absolute, even when the configured value is relative — a lead that resolves it
itself can pick a different file from the one the daemon checks.
2026-09-12 05:46:34 +07:00
Dai Ha 3bf3968bc7 Merge #487: resolve a relative leadRollover.handoverPath against the calling lead's workspace (fleetd #480 follow-up) 2026-09-12 05:43:12 +07:00
Dai Ha 261aa056f9 fleetd #480 follow-up correction 2: guarantee resolveHandoverPath is always absolute
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m32s
LeadRollover.resolveHandoverPath's relative branch resolved the configured
handoverPath against leadWorkspace.apply(...) (fleet.leaders.<name>.cwd) but
never forced the result absolute. If an operator writes a RELATIVE cwd, the
returned path stays relative, silently breaking the "always absolute"
contract documented on PendingRollover.

Fix: call toAbsolutePath() unconditionally on both branches (the
already-absolute input branch, where it is a no-op, and the relative
branch), so neither branch trusts isAbsolute() alone to already imply what
toAbsolutePath() enforces. Method javadoc now states the absolute result is
guaranteed, not merely usual.

Added a test: a lead with a RELATIVE cwd and a relative handoverPath still
yields an absolute PendingRollover.handoverPath. Asserts both isAbsolute()
and the exact resolved value, since isAbsolute() alone would also pass for a
path resolved against the wrong base.

Proved the test discriminates: reverting the toAbsolutePath() calls (keeping
the test) made it fail with an AssertionFailedError ("expected: <true> but
was: <false>"); restoring the fix made it pass again.

Note: FleetConfig has no validation on fleet.leaders.<name>.cwd at config
load (grep across every validate* method: 0 matches for .cwd()) — a relative
cwd is silently accepted. Not adding validation here per instruction; that
is a separate ticket.
2026-09-12 05:41:40 +07:00
Dai Ha 042b8c99dd fleetd #480 follow-up correction: cover Fleetd.leadRollover(...)'s own wiring behaviourally
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m51s
Add FleetdLeadRolloverWorkspaceLookupTest, calling the package-private Fleetd.leadRollover(...)
factory directly (with a real ConfigRef built from a temp fleetd.yaml, never the gitignored live
one) to prove the terminal -> lead-name -> Leader.cwd() lookup it builds actually works: a relative
handoverPath resolves against the calling lead's configured cwd; a terminal absent from the live
lead-terminal map falls back to user.dir; and the lookup is read live, not snapshotted at
construction time (a lead discovered by the tab scan after leadRollover(...) was built still
resolves correctly).

Proved this closes the gap: mutating the factory's lambda body (String leadName = null;, always
"no lead found", which forces the daemon-cwd fallback this ticket exists to fix) left the full
1669-test suite green before this commit. With the new test added, the same one-line mutation now
fails 2 of its 3 cases; reverting it goes green again (3/3). Mutation applied/reverted only during
verification and is not part of this commit (git diff on Fleetd.java is empty).

FleetdLeadRolloverWiringTest's class javadoc corrected: it previously claimed no behavioural test
could catch this wiring dropping out, which was true only before this commit and only covered the
factory's own body, not its call site. Restated what each test class actually covers: the source-
text pin covers the call site's argument list; the new behavioural test covers the lambda's body.
2026-09-12 05:32:25 +07:00
Dai Ha 4bfab6b718 fleetd #480 follow-up: resolve a relative leadRollover.handoverPath against the calling lead's workspace
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m4s
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the
calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that
lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores
only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed
back to the lead in the fleet_handover open response, and the default bootstrapText sentence
all see the same absolute location instead of a value resolved against whatever directory the
daemon process happened to start in.

FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it
would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor
(resolvedHandoverPath) method builds the default sentence from the resolved path instead.

Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to
lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on
every call, never off a startup snapshot.
2026-09-12 05:19:49 +07:00
Dai Ha 008a457557 docs: fleet_handover is shipped — update the charter table and the handover skill
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m38s
fleetd #480 merged in #483, #484 and #485, so the instruction surface has to
catch up. Per CLAUDE.md's own rule, a code change that silently invalidates the
canonical block is an incomplete change.

CLAUDE.md + wiki/7-Use-Cases.md: one new intent-to-tool row for fleet_handover.
Both copies edited identically; the byte-identical sync check prints True.
The row carries the trap rather than just the call: open FIRST, then write the
file, then confirm — because confirm refuses the file as stale unless its
modified time is later than the open request, so the obvious order fails.

.claude/skills/handover/SKILL.md: the skill said "do not call a fleet_handover
tool: it does not exist." That was true when it was written this morning and is
now false, which is the worst state for an instruction file to be in. Replaced
with the two real paths (by hand, or with the tool when leadRollover: is
configured), plus a new section 11 giving the three-step order and the five
things that surprise a caller — chiefly that `accepted` does not mean the pane
has been cleared, and that operatorConfirmed is a report of what a human said,
not a confidence level.

Section 11 also states plainly that the bootstrap prompt landing in a freshly
cleared pane is not yet proven end-to-end, and that the recovery is the manual
path — which is why the file is written before confirm, never after.
2026-09-11 07:29:59 +07:00
ltms 9494a6b99a Merge #485: fleetd #480 Unit C — the fleet_handover MCP tool
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m58s
fleet_handover{action: "open"|"confirm"|"cancel"} drives LeadRollover, which #483 and
#484 landed with nothing calling it. Primary-only via a new Authz.Action.HANDOVER, on
the same case line as SPAWN/STOP/DRAIN.

The tool has NO terminal, session or leadTerminal parameter of any kind — the pane is
always callerTerminal(exchange), resolved from the connection. A lead can therefore only
ever roll itself, never another lead. That is charter invariant 3, and it is the second
of the two corrections recorded in LeadRollover's class javadoc.

Registered unconditionally, so the tool surface does not vary with config: with
leadRollover: absent, every action returns a clean NOT_CONFIGURED refusal instead of
failing, and open()'s IllegalStateException (config removed by a hot reload after
construction) is caught and turned into the same refusal. A config-dependent tool set
would have collided with #474's charter tool-surface gate and McpContractDocTest.

Correction round applied before merge, and it is the reason this took two passes.
The unit first shipped with a defaulted 15-argument FleetMcp constructor delegating to
the new 16-argument one with leadRollover = null. I mutated the wiring rather than
reasoning about it: deleting just the leadRollover argument from Fleetd.main's FleetMcp
call compiled with 0 errors and passed all 1659 tests, BUILD SUCCESS — while the live
daemon would have answered NOT_CONFIGURED to every fleet_handover call for ever.
Neither FleetMcpHandoverTest (it builds its own FleetMcp) nor FleetdLeadRolloverWiringTest
(it pins that LeadRollover is constructed, not that it is passed on) could see it.

The worker then found the defect was wider than I had named: all five shorter
constructors (11/12/13/14/15-arg) formed one defaulting chain into the 16-arg one, each
silently supplying another feature's "off" value — leadChannel, outage, leadSeats, peers,
and finally leadRollover. All five are deleted. FleetMcp now has exactly one public
constructor, so every one of those features is compile-enforced at its call site, not
just this one.

Verified by the lead before merge, on PR head merged with current main (0176378):
- CI run 1728 green on eb05576.
- The identical mutation re-run by me: control against the original -> 1; mutant present
  (MUT-DROP-ARG2) -> 1; original gone -> 0; `mvn -o -q compile` now FAILS —
  "constructor FleetMcp ... cannot be applied to given types; reason: actual and formal
  argument lists differ in length" at Fleetd.java:[698,24]. The antidote holds: required
  parameter = compile error, defaulted overload = silent survivor a green suite vouches for.
- Restored clean (0 changed files), then mvn clean install on the merged tree:
  BUILD SUCCESS 1, BUILD FAILURE 0, Tests run: 1662, Failures: 0, Errors: 0, Skipped: 0.
- Three call sites of `new FleetMcp(` measured, all accounted for: Fleetd.java,
  FleetMcpAuthzTest.java, FleetMcpHandoverTest.java.
2026-09-11 02:27:56 +02:00
Dai Ha eb0557621e fleetd #480 correction round: collapse FleetMcp to one required constructor
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 1m32s
FleetMcp had a defaulted 15-argument constructor that delegated to the new
16-argument one with an implicit null for leadRollover. Dropping the
leadRollover argument from Fleetd.main's FleetMcp(...) call fell back to that
shorter overload, compiled fine, and left all 1659 tests green — the live
daemon would then answer NOT_CONFIGURED to fleet_handover forever with
nothing going red.

Delete every overload that could reach the 16-arg constructor with a
silently-defaulted leadRollover (11/12/13/14/15-arg forms all chained to it),
leaving the 16-arg constructor as FleetMcp's sole public constructor. Update
FleetMcpAuthzTest's call site to pass every parameter explicitly (leadChannel
null, OutageSource.none(), LeadSeatSource.none(), List.of(), leadRollover
null) — Fleetd.java and FleetMcpHandoverTest already called the full form.
Proved with mvn -o -q compile: removing the leadRollover argument from
Fleetd.main now fails to compile instead of silently defaulting.

No behaviour changes — NOT_CONFIGURED refusals are unchanged.
2026-09-11 07:25:08 +07:00
ltms a0eed6f01b Merge #484: fleetd #480 Unit E — a BLOCKED pane is not a settled pane
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m35s
LeadRollover's two settle waits used AgentStatus#injectable(), which is
IDLE || BLOCKED || DONE. That is the right rule for Injector ("may I deliver a message
without stepping on a live turn") and the wrong one here ("has the turn actually
ended"), because BLOCKED is a live turn that is merely paused — a pane sitting on an
approval prompt.

Path in: the lead calls confirm(); its turn carries on and hits anything needing
approval; herdr reports blocked; within 250ms the deferred continuation reads that as
settled; /clear is typed into an open prompt. That destroys the lead's live context
mid-turn, which is exactly what the turnSettleSeconds gate added in #483 exists to
prevent. The window is turnSettleSeconds, default 20s.

Both waits now require a real turn boundary — IDLE or DONE. AgentStatus#injectable() is
untouched: it is correct for Injector, LeadHeartbeatLoop, LeadCoordLoop, ReplyPushLoop
and HerdrPeerLauncher's two readiness checks, all of which are delivery gates.

Verified by the lead before merge:
- CI run 1726 green on head e2a91e8.
- Mutation on the half the worker did NOT mutate: dropped the DONE arm, leaving
  `if (status == AgentStatus.IDLE)`. Mutant proven applied with two unrelated proofs
  using different search strings (MUT-DROP-DONE present -> 1; 'IDLE || status' gone
  -> 0). Same single test, same 100s cap: unmutated passes in 0.172s; mutated is still
  running at 100s and does not pass. So doneStatusStillCompletesTheFullRoll really does
  pin the DONE arm, and the fix did not over-tighten to IDLE-only.

Why the earlier #483 battery could not have caught this: it proved the guard FIRES, not
that the guard's predicate was tight enough. LeadRolloverTest drove the fake status with
only "idle" and "working"; neither value discriminates this defect, only "blocked" does.

Correction round applied before merge: both warning lines said "never went idle" and
"did not become injectable". Retiring that word is the whole point of the change, so a
future reader debugging from those lines would have been handed back the wrong mental
model. Both now say "did not reach a turn boundary (IDLE or DONE)".

Filed separately as #486, deliberately not fixed here: with the DONE arm removed the
test hangs rather than fails, because waitUntilAtTurnBoundary is bounded only by an
injected clock that the tests hold fixed. Not a production bug — nowMillis is
System::currentTimeMillis there — but it turns a future regression into a CI hang
instead of a red test.
2026-09-11 02:21:06 +02:00
Dai Ha e2a91e883e fleetd #480 Unit E correction: retire "injectable" wording from the log lines
CI / build (pull_request) Successful in 1m28s
CI / contract (pull_request) Successful in 1m27s
Both waitUntilAtTurnBoundary guard messages still said "never went idle" /
"did not become injectable" — the old mental model the rename was meant to
retire. Made both say what the code now actually waits for: a turn boundary
(IDLE or DONE).
2026-09-11 07:18:51 +07:00
Dai Ha 62646957ea fleetd #480 Unit C: wire fleet_handover MCP tool onto LeadRollover
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m56s
Adds the fleet_handover tool (open/confirm/cancel) as a thin adapter over
LeadRollover, registered unconditionally so the charter tool-surface gate
sees a stable set regardless of whether leadRollover: is configured. With a
null LeadRollover every action degrades to a clean NOT_CONFIGURED refusal
instead of throwing. Gated on a new Authz.Action.HANDOVER (primary-only,
same as SPAWN/STOP/DRAIN). The caller's own connection-resolved terminal is
the only lead identity ever used — the tool's input schema carries no
terminal/session/leadTerminal parameter, so a lead can only ever roll
itself. Fleetd.main now passes its existing leadRollover local into FleetMcp
via a new trailing constructor parameter.
2026-09-11 07:07:24 +07:00
Dai Ha 802c0ab701 fleetd #480 Unit E: BLOCKED is not a settled turn boundary
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m55s
LeadRollover's waitUntilInjectable used AgentStatus#injectable(), which
accepts BLOCKED. A BLOCKED pane is paused mid-turn on a prompt, not
settled — reusing injectable() let /clear (or the bootstrap text after
it) fire into an open approval prompt within the 20s settle window,
destroying the lead's live context.

Renamed the helper to waitUntilAtTurnBoundary and restricted both waits
to IDLE or DONE only, with a comment explaining why this class does not
reuse injectable() (it answers "may I deliver", not "has the turn
ended"). Added tests for BLOCKED-forever on both waits (zero sends /
exactly one send) and for DONE still completing the full roll.
2026-09-11 07:01:28 +07:00
ltms 4bab23e241 Merge #483: fleetd #480 Unit A — lead rollover core (config + executor)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m37s
Two corrections were applied to the original unit before this merge, and the PR
description above still describes the pre-correction shape:

1. confirm() no longer rolls inline. It validates every gate, then hands a one-shot
   continuation to continuationRunner and returns. confirm() is called BY the lead FROM
   its own turn, so the pane is WORKING and cannot report injectable until confirm()
   returns; the old code sent /clear first and then timed out waiting, which destroyed
   the lead's context and started no fresh session. The continuation waits for the
   calling turn to settle FIRST (turnSettleSeconds, a new knob), and if that wait times
   out it sends no /clear at all.
2. The pane to roll comes from the caller's terminal, not PrimaryRegistry. A single-slot
   lookup let lead X's confirm() clear lead Y's pane. open() records the caller's
   terminal; confirm() refuses with NOT_YOUR_ROLLOVER on a mismatch.

Verified by the lead before merge:
- CI run 1722 green on head a942622.
- Mutation battery on the two safety-critical branches, in a scratch worktree at a942622.
  Controls measured against the ORIGINAL first; each mutant proven applied with two
  unrelated proofs using different search strings.
  M1 `if (!turnSettled)` -> `if (false)`: turnThatNeverSettlesSendsNoClearAtAll FAILS
     (LeadRolloverTest:247, expected 0 sends but was 1).
  M2 ownership check -> `if (false)`: aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover
     FAILS (LeadRolloverTest:266). Unmutated harness-proof run: exit 0 with both guard
     WARN lines logged, so the branches really are exercised.

Nothing calls LeadRollover yet. The fleet_handover MCP tool is Unit C.
2026-09-11 01:51:20 +02:00
Dai Ha a94262271b fleetd #480 correction round: defer the roll, and gate it on caller identity
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m3s
Two defects found after the fact, both from the original brief, both fixed here.

1. confirm() is called FROM the calling lead's own turn, so its pane is still
   WORKING and can never report injectable inside that same call. The old
   confirm() sent /clear before polling for that — the poll always timed out,
   but only after /clear had already fired and queued, destroying the lead's
   context with no fresh session ever started and a refusal return that lied
   about what had happened.

   Fix: confirm() now only validates and, if every gate passes, hands a
   one-shot continuation to a new continuationRunner (a real virtual thread in
   production, Runnable::run in tests) and returns RollDecision.approved()
   immediately - "scheduled", not "rolled". The continuation itself does the
   actual work, once the calling turn has ended: wait for the SAME pane to
   report injectable again (new turnSettleSeconds config key, default 20) -
   if this never happens, /clear is NEVER sent, at all - then /clear, then
   wait again (clearSettleSeconds, as before), then bootstrapText. The "no
   timer/scheduler, only confirm() can roll" invariant is restated precisely
   in LeadRollover's class javadoc: it is about initiative, not synchronicity
   - a single-shot continuation of an already-approved confirm() call still
   satisfies it; a recurring background loop would not.

2. confirm() resolved the pane to clear via PrimaryRegistry.primaryTerminal(),
   a single-slot lookup that is correct for a background loop with no caller
   but wrong here: on a daemon with more than one labelled lead tab, lead X's
   confirm() could clear lead Y's pane, violating the charter's "identity
   comes from the connection, never an argument" invariant.

   Fix: open() and confirm() now take the caller's terminal id as a parameter
   (resolved by the MCP layer from the connection - the later MCP-tool unit
   must pass it in, never accept it as a request field). confirm() refuses
   with a new NOT_YOUR_ROLLOVER reason unless it matches the terminal open()
   recorded. LeadRollover no longer depends on PrimaryRegistry at all.

Also: renamed RollResult to RollDecision (rolled -> accepted) to reflect the
new meaning - approved and scheduled, not necessarily cleared yet. Added
turnSettleSeconds to the leadRollover: config block (documented in
fleetd.example.yaml alongside the existing keys) and updated Fleetd.java's
leadRollover(...) factory to drop the primaryRegistry parameter, with
FleetdLeadRolloverWiringTest's source-text pin updated to match.

New tests: turnThatNeverSettlesSendsNoClearAtAll (the branch that matters
most - a turn that never ends means /clear is never sent) and
aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover (NOT_YOUR_ROLLOVER),
plus a settle-after-clear timeout test and an open() input-validation test.
LeadRolloverTest: 11 -> 14 tests.
2026-09-11 06:44:58 +07:00
Dai Ha 5c12865c25 fleetd #480 Unit A: lead rollover core (config block + executor)
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m58s
Adds the opt-in leadRollover: config block and LeadRollover, the executor a
later unit's MCP tool will call. A lead writes a handover file, then open()
records a token and confirm() verifies it (exists, non-empty, fresh) and an
operator confirmation before clearing the lead's own pane via /clear (sent
directly through AgentControl, bypassing Injector, same as
ClaudeCodeLauncher#clearContext) and bootstrapping a fresh session. Nothing
but an explicit confirm() call can ever roll a pane - no timer, no heartbeat,
no background thread anywhere in this class.

Wired into Fleetd.java exactly like LeadHeartbeatLoop: constructed only when
leadRollover: is present at startup, and nothing calls it yet - the MCP tool
is a separate, later unit.

Classified leadRollover: as HOT in ConfigRef (joins placement/
memberCredentials/memberLoginShell/models): the executor holds
Supplier<FleetConfig.LeadRollover> and reads every field fresh per call,
unlike LeadHeartbeatLoop's frozen final fields. The one caveat: the object's
construction is still gated on presence in the startup config snapshot, so a
freshly-added block needs a restart before anything exists to call.

Tests: LeadRolloverTest (14 cases covering the 6 hard requirements - no
object without the config block, only confirm() can roll, missing/empty/
stale handover file each refuse by name, requireOperatorConfirm gating, and
the injected wall-clock supplier) and FleetdLeadRolloverWiringTest (source-
text pin on Fleetd.main's construction call, mirroring
FleetdCompletionResolverWiringTest). Also updated the existing
FleetConfigValidateAllTest, FleetConfigWithDefaultsPreservesEveryComponentTest,
ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest to
account for the new record component.
2026-09-11 06:33:01 +07:00
Dai Ha 3f8c956325 docs: the handover skill must not promise a tool that does not exist
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m30s
The skill shipped in #481 described fleetd #480's automation in the present
tense: "fleetd ticket #480 lets a lead session hand off to a fresh one". The
tool is not built. A session loading the skill would look for a fleet_handover
tool it cannot call, and this repo already knows what a confidently wrong
instruction costs — a wrong comment stops the reader investigating with a false
conclusion, which is worse than no comment.

Now says plainly: today the handoff is manual, #480 will automate the same
cycle, and do not call fleet_handover because it does not exist. The nine
procedure rules are unchanged — only who performs the swap changes when #480
ships, not what the file must contain.

My wording to fix, not the worker's: the brief handed them the present tense.
2026-09-11 06:17:11 +07:00
ltms 2051ceeb26 Merge #481: the handover skill (fleetd #480 Unit B)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m27s
Adds .claude/skills/handover/SKILL.md — the outgoing lead's procedure for writing
the file a fresh lead session inherits.

Reviewed by the lead against the eight requirements in the brief: every number
carries its command, measured-vs-reported is marked, open decisions name their
owner, live hazards, what is not owed, a first section of three re-measurement
commands, a time-and-commit stamp, and a plain statement that the file goes
stale. All eight are present, plus a leave-out section and a template.

Verified by the lead, not taken from the worker's report:
- CLAUDE.md diff is the addendum's skill list only, in the primary-side group.
- The canonical block is byte-identical with the wiki template (sync check True).
- CI run 1718 green on 325d0771, the same commit reviewed.

Follow-up owed by the lead: the skill describes #480's automation in the present
tense, but the tool does not exist yet. The lead will soften that wording so the
skill does not promise a fleet_handover tool a session cannot call.
2026-09-11 01:16:40 +02:00
Dai Ha 325d0771a4 fleetd #480 Unit B: add handover skill for lead session handoff
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m54s
Adds .claude/skills/handover/SKILL.md, the procedure an outgoing lead
follows to write the handover file a fresh lead session inherits when
fleetd clears the pane. Registers the new skill in CLAUDE.md's
primary-side skills list; no other change to CLAUDE.md.
2026-09-11 06:13:49 +07:00
Dai Ha 7b97aae85b docs: move the redeploy procedure out of CLAUDE.md into a skill
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m34s
CLAUDE.md loads into every session. The daemon-redeploy procedure is
needed only when someone redeploys, so it paid for context it did not
use: 58 lines, about 918 est. tokens, every session.

The procedure now lives in .claude/skills/redeploy-fleetd/SKILL.md,
which loads only when invoked. The moved text is byte-identical to
what was removed, plus one new paragraph documenting --no-build
(scripts/redeploy-fleetd.sh:34,62,266) — the script and the wiki
already had that flag, CLAUDE.md never did.

CLAUDE.md keeps an 11-line pointer, because two rules must stay
resident: a merge is not a deployment, and workers must never
redeploy. A session learns it needs the procedure before it needs
the skill, then the next line names the skill.

Also updates the addendum's primary-side skill list, as this file's
own rule for .claude/skills/** changes requires.

CLAUDE.md: 34,442 -> 31,591 chars.
Canonical block untouched — the wiki sync check still prints True.
2026-09-11 05:48:06 +07:00
Dai Ha 49a404ddf3 Merge #474 follow-up: pin main's ConfigRef wiring against the surviving mutation
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m37s
My battery on the #474 merge found one survivor: reverting Fleetd.java:154
from the three-argument ConfigRef constructor to the plain two-argument
one turns the live reload gate off and leaves all 1633 tests green. Both
new #474 tests build their own ConfigRef with the method reference, so
neither reads what main chose.

This adds FleetdConfigRefWiringTest, following the three source-text
precedents already in the tree (FleetdBackendQuarantineWiringTest,
FleetdLeadSeatWiringTest, FleetdCompletionResolverWiringTest) rather
than the weaker sibling pattern that builds the wiring itself. No
production change.

The worker branched fresh off 4466ee0 rather than continuing its old
branch off the stranded 435e022 base. That was its own call and it was
the right one: one merge base, a clean 80-line diff, and no cherry-pick
needed this time.

Its vacuity guard is worth keeping in mind for the next source-text
test: it asserts the file it read contains 'public final class Fleetd'
before asserting anything about the mutation, with a message saying the
assertFalse below would pass vacuously on a broken read. An unrelated
anchor is the right choice there, because a guard that shares the
mutation's text cannot tell a bad read from a real change.
2026-09-10 20:50:13 +07:00
Dai Ha 72d6a6878b fleetd #474 follow-up: pin main's config wiring against M2
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m32s
Fleetd.main's own choice of the three-argument ConfigRef constructor
(with Fleetd::assertChartersNameOnlyRegisteredTools as extraValidation)
was unpinned. Reverting Fleetd.java:154 to the plain two-argument
constructor compiled with 0 errors and left the whole suite green,
because ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest
each build their own ConfigRef directly rather than through main.

Adds FleetdConfigRefWiringTest, a source-text check on Fleetd.java
following the FleetdBackendQuarantineWiringTest/FleetdLeadSeatWiringTest/
FleetdCompletionResolverWiringTest precedent: asserts the exact
three-argument construction is present, asserts the plain two-argument
form is absent, and guards against a vacuous pass on a broken/empty
source read by first asserting an unrelated anchor is present.
2026-09-10 20:48:05 +07:00
Dai Ha 4466ee0ef2 fleetd #474: ConfigRef.reload() runs the charter tool-surface gate too
CI / build (push) Successful in 1m37s
CI / contract (push) Successful in 1m44s
A charter naming an MCP tool the server does not register refused Fleetd.main
at startup but slipped through ConfigRef.reload(), because reload() only ran
FleetConfig.validateAll(), which never looks at what a charter's text names.

CharterToolSurface stays in the mcp package (config must not depend on it), so
ConfigRef now accepts the check as a Consumer<FleetConfig> extraValidation,
run inside reload()'s same try/catch as validateAll(). Fleetd.main wires a new
package-private adapter, Fleetd.assertChartersNameOnlyRegisteredTools, into
both the startup call site and ConfigRef's constructor, so the two call sites
can never check different things.

Tests: ConfigRefTest (reload refuses/accepts, via a locally-built equivalent
consumer since Fleetd's method is package-private to dev.ltms.fleet) and the
new FleetdConfigRefCharterToolSurfaceWiringTest (same proof through the exact
Fleetd::assertChartersNameOnlyRegisteredTools reference production uses).
Verified deleting the new extraValidation.accept(fresh) call site fails both
new "refuses" tests by name.

(cherry picked from commit 97e4c1d658)

Lead review. Cherry-picked, not merged: the worker branched from 435e022,
which is not an ancestor of main (I had reset and rewritten that commit's
message while the worker was already on it). Merging its branch would have
added a second merge commit for work already on main. Same content, no
duplicated history. Rule learned: once a worker is spawned, its base
commit is published.

My own mutation battery, five cells, each a full `mvn -B clean test` on
this commit's own tree:

  CONTROL 1, unmutated       1633 tests, 0 failures, BUILD SUCCESS
  M1 delete accept(fresh)    KILLED  2 by name
  M3 3-arg ctor stores no-op KILLED  2 by name
  M4 set(fresh) before valid KILLED  6
  M5 accept outside try      KILLED  2 errors
  M2 main uses the 2-arg ctor SURVIVED

M2 is a real gap and it is a test gap, not a defect: the production code
here is correct. Reverting Fleetd.java:154 to `new ConfigRef(configPath,
cfg)` turns the live reload gate off and leaves all 1633 tests green,
because both new tests build their own ConfigRef with the method
reference rather than reading what main wires.

That is the fourth Fleetd.main call site to survive a battery (#446 M5
and M7, #466 M1, now this). Extracting a check into a well-tested helper
moves the untested surface UP, into the line that chooses to call it. The
repo already has the answer in three source-text wiring tests
(FleetdBackendQuarantineWiringTest, FleetdLeadSeatWiringTest,
FleetdCompletionResolverWiringTest); the worker followed the other
sibling pattern, which builds the wiring itself and so cannot pin main.
Delegated as a follow-up, with the surviving mutation as its acceptance
criterion.

Also mine to correct: my M3 cell's annotation said the no-op store "must
be 2" and measured 1. The 2-arg constructor delegates rather than
assigning, so my pattern only ever matched the mutant. The cell still
stands on its other proof (requireNonNull left = 0). Third battery in a
row with a wrong CONTROL annotation of my own.
2026-09-10 20:40:49 +07:00
Dai Ha 25ba7f16bb Merge #473: fleet_profiles and fleet_list report which attempt a quarantine is on (fleetd #466 item 2)
CI / contract (push) Successful in 1m20s
CI / build (push) Successful in 1m24s
The escalating cooldown landed in 789b6a8, but the only thing an operator
could see was quarantinedForSeconds. A long number does not say whether
this is the first exhaustion or the fifth, and after escalation those look
the same from outside: 3600 seconds could be a big base cooldown or a
credential on its fourth strike.

Both reports now carry quarantineAttempt beside quarantinedForSeconds.

The part that matters is HOW, not that the field exists.
BackendQuarantine#status(credentialId) does ONE quarantines.get() and
returns both values from the same QuarantineState. It is not two accessors
that a caller pairs up. That is the pattern
CompositePeerLauncher.modelGateState() already documents: a report that
reads a different source from the behaviour it describes will eventually
disagree with it, and the disagreement is invisible because both halves
look right on their own.

This is reporting only. Escalation, the ceiling and the quiet-gap reset are
untouched. The flat two-argument constructor still reports a real growing
attempt count even though its cooldown stays flat - which is the honest
answer, since the streak is real whether or not the cooldown uses it.

The worker's report was lost: its fleet_reply never arrived and its inbox
was empty. Nothing was lost with it, because the brief required pushing
first and putting the report in the PR body. That is the second time that
practice has saved a turn.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:21:14 +07:00
Dai Ha 5ba69c9cf2 fleetd #393: correct a comment that claimed two tests pin a call they cannot
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m51s
The comment above the charter writer said flipping that call back to
putArray was "proven load-bearing" and pointed at two named tests. I
mutated exactly that, on the merge commit, and it SURVIVED at 1618 green -
so neither named test covers it, and neither can.

The reason is structural, not a missing test: that writer runs first
against an empty array, so putArray has nothing to replace and the two
idioms are equivalent there. No test can distinguish them. The worker
measured the same thing independently and said so in their report; the
comment was left over from the pre-fix state, where the mutation being
described was a DIFFERENT one (a later writer destroying the charter).

A comment that names tests which do not cover the line is worse than no
comment. The next person mutates the line, sees green, and concludes the
tests are broken.

The corrected version states what each cell actually measures: this line
is unpinnable and why, the later writers ARE pinned and by which test
names, and deleting this line entirely fails the charter-reaches-the-member
assertion - the hole that predates #393.
2026-09-10 20:17:21 +07:00
Dai Ha 17052bb515 Merge #471: seeded skills reach an opencode member, and instructions[] stops depending on write order (fleetd #393)
memberSkills: copied skill folders into every provisioned worktree's
.claude/skills/ and stopped there. Claude Code reads that directory
natively; opencode never does. So the feature was INERT for opencode
members rather than broken: the copy succeeded, the files were correct,
and nothing ever read them. No test failed because there was nothing to
fail - the feature worked at the only layer it implemented.

An opencode member's only channel for static guidance is the instructions[]
array in its generated config. OpenCodeLauncher now adds each seeded
skill's SKILL.md there. A folder with no SKILL.md is never delivered and
the log names it.

The second half is an ordering hazard fleet01 found by reading, and that I
then measured. Three writers append to instructions[]: the role charter,
the seeded skills, and the IDE rules. The charter used putArray
(CREATE-OR-REPLACE) while the other two used withArray (get-or-create).
That was safe only because the charter ran first against an empty array -
an undeclared constraint that nothing tested. Measured on the earlier
merge: making the skills writer use putArray left 1603 tests green while
silently deleting the charter entry, so an opencode member would launch
with no role contract at all. Worse than the bug being fixed, and
invisible.

All three writers now use withArray, and three tests pin the array's
CONTENTS (never its size - a size assertion passes when putArray swaps two
entries for two others) across the combinations that matter:
charter-only, charter+ide, charter+skills+ide.

WHY NOTHING CAUGHT IT, MEASURED RATHER THAN ASSUMED. fleet01 first said
the missing axis was the COMBINATION of writers, then revised that to a
stronger claim: that nothing asserted the charter reaches an opencode
member at all. I checked the second claim on this merge and it is FALSE.
Deleting the charter writer outright fails 6 tests, and 3 of those existed
before #393: writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet,
roleCharterWithoutMcpOrCustomProviderStillWritesAConfig, and
aPinnedEndpointAndTheFleetMcpCoexistInOneConfig. The charter reaching a
member was already pinned.

So their FIRST diagnosis was the right one. The surviving mutation did not
delete the charter write; it made a LATER writer replace the whole array.
Every pre-existing test had exactly one writer active, and with one writer
putArray and withArray are indistinguishable. The gap was never "is the
charter delivered" - it was "are two writers ever active at once", which is
the combination axis. Recording this because the stronger claim is the more
quotable one and it would have sent the next reader looking for a hole that
is not there.

One honest residue: the charter writer's own idiom cannot be pinned.
Flipping it back to putArray leaves the suite green, and always will,
because it runs first against an empty array where the two idioms are
equivalent. The edit removes an undeclared constraint for the next person
to add a writer; it is not a change any test can detect. The comment in
the source claimed two named tests cover it - that was wrong, and the
commit after this one corrects it rather than leaving a false claim beside
the code.

What this does NOT do: opencode has no equivalent of Claude Code's Skill
tool, so the content is static system-prompt text present from spawn, not
something a member can invoke by name. This closes the DELIVERY gap and
cannot close the ACTIVATION gap. That residue is opencode's design.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:17:21 +07:00
Dai Ha e95ed99bf7 fleetd #466 scope item 2: report the quarantine repeat count, not only the seconds
CI / build (pull_request) Successful in 1m31s
CI / contract (pull_request) Successful in 1m33s
Add BackendQuarantine#status(credentialId) -> Optional<Status>, a single
QuarantineState read that answers both remainingSeconds and repeatCount
together -- the same "one accessor" pattern CompositePeerLauncher.
modelGateState() already uses, so the two facts can never disagree.

fleet_profiles/REST GET /profiles and fleet_list's capacity rows
(FleetMcp.profilesView/capacityView) now call status() instead of
remainingSeconds() and add a "quarantineAttempt" field beside
"quarantinedForSeconds": 1 for a first occurrence, 2 for the second
in a row, and so on.

No change to the escalation, ceiling, or reset logic itself -- this
unit is reporting only.
2026-09-10 20:09:13 +07:00
Dai Ha 1477e4358a Merge #472: one canonical tool-name set, and charters are checked against it (fleetd #469)
CI / contract (push) Successful in 57s
CI / build (push) Successful in 2m0s
A role charter is free text in config that tells a member which tools to
call, and nothing checked that those tools exist. A charter naming
bridge_send - a name CB-634 removed - started the daemon cleanly, and the
member found out at run time by calling something that was not there.

#464 shipped a test for this, but it wrote its own charter into a @TempDir,
so nothing anyone put in the real config could fail it. That was a defect
in my acceptance criteria, not in that work.

Now: FleetTool is one enum of the 11 registered wire names, and every
reader goes through it.

- FleetMcp's schema builders pass FleetTool.X.wireName() instead of a
  literal.
- FleetMcp's constructor asserts at startup that what it registers with the
  SDK equals FleetTool.wireNames() exactly, in both directions.
- toolAction(String, Map) resolves arbitrary wire input against
  FleetTool.byWireName() and keeps its run-time throw, which is necessary -
  network input has no closed compile-time form. It then hands off to
  authzAction(FleetTool, Map), a switch over the enum with NO default, so a
  new tool is a compile error at that layer.
- CharterToolSurface lives in mcp, not config, and Fleetd.main calls it
  right after validateAll(). Config must not depend on the MCP server:
  config loads before the server exists.

My brief undercounted the problem and the worker corrected it. I said there
were two tool-name inventories plus a test fixture. There were FIVE: the
registrations, the authz switch, and three separate source-text scrapes of
FleetMcp.java in CharterToolSurfaceTest, FleetMcpAuthzTest and
McpContractDocTest - none of which the ticket mentioned. Fixing those three
was required, not scope creep: once the literals moved into FleetTool their
regexes matched zero names, so one would have failed on its vacuity guard
and the other two would have gone quietly vacuous. Inventories after: one.

Verified independently on origin/main before accepting the wider diff:
three test files did read FleetMcp.java as source text, with a control file
at zero to prove the search discriminated.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:59:03 +07:00
Dai Ha 9e4e423ad6 fleetd #393 follow-up: remove the instructions[] writer-ordering hazard
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 2m11s
OpenCodeLauncher.writeConfig has three writers into the instructions[]
array (charter, seeded skills, IDE rules). The charter writer used
putArray (create-or-REPLACE) instead of withArray (get-or-create), which
"worked" only because it happened to run first against a still-empty
array — an undeclared ordering dependency nothing tested. Found by the
fleet01 lead and verified on this branch's merge: flipping the skills
writer to putArray left the full 1603-test suite green while silently
deleting the charter entry, which would launch an opencode member with
no role contract at all.

Fix: charter's putArray -> withArray (one-word change, behavior-identical
today). Add three tests asserting instructions[] CONTENT as an exact
ordered list (not size) across writer combinations: charter only,
charter + IDE rules, and charter + IDE rules + seeded skills. Mutation
testing (see PR body) shows the skills and IDE-rules writers are each
independently detectable by name; the charter writer's own mutation is
not detectable by any test, because it structurally always runs first
against an empty array, so putArray and withArray are equivalent there.
2026-09-10 19:58:25 +07:00
Dai Ha 6c2d6e93cb fleetd #469: one canonical FleetTool set backs registration, authz and charter checks
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m37s
FleetConfig.validateCharters() only checked that a charter key is a role
wire name and its text is non-blank; #464's CharterToolSurfaceTest compared
charter text against the registered tool surface, but wrote its own charter
into a @TempDir fixture, so nothing anyone wrote into the live fleetd.yaml
could ever fail it.

Add FleetTool, an enum in dev.ltms.fleet.mcp holding the one canonical set
of registered tool wire names. FleetMcp's tool schemas now derive their
names from it, its constructor asserts at startup that what it actually
registers with the SDK equals FleetTool.wireNames() exactly, and its authz
dispatch (toolAction/authzAction) resolves the wire string against FleetTool
before switching on the enum itself with no default -- adding a tool without
pinning its Authz.Action is now a compile error, not just a test gap.

Add CharterToolSurface (mcp package, not config -- config loads before the
MCP server exists) and call it from Fleetd.main right after
cfg.validateAll(), so a charter naming a tool the server does not register
refuses the daemon's startup, naming both the charter key and the unknown
tool. FleetdStartupValidationTest proves this through Fleetd.main itself
against a live-shaped config fixture (bridge_send, CB-634's own removed
name).

CharterToolSurfaceTest, FleetMcpAuthzTest and McpContractDocTest each kept
an independent regex scrape of FleetMcp.java's source for the registered
side of their own comparison -- three more copies of the same list nothing
tied together. All three now read FleetTool.wireNames() instead.

Proved canonical by removal: deleting FleetTool.ACK while ackTool() still
referenced it broke mvn compile in two places (FleetMcp.java:940,:1898);
registering a schema under a literal not backed by FleetTool
("fleet_ack_v2") failed FleetMcp's new startup assertion in every test that
constructs it (12 errors, IllegalStateException at FleetMcp.<init>). Both
reverted before this commit.
2026-09-10 19:55:02 +07:00
Dai Ha 789b6a8716 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
CI / contract (push) Successful in 1m56s
CI / build (push) Successful in 2m4s
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things on the record because they are judgement calls, not facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this codebase
  reports a spawn success back to this class, so "it started working again"
  cannot be observed here. A base cooldown of quiet is the best available
  evidence, and the class doc says that plainly instead of implying the
  stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - the same
formula with multiplier 1.0 and a ceiling equal to the base, which collapses
to the original flat "now + cooldown". Every existing call site keeps its
shape.

The second commit exists because my own mutation on the first merge found
the wiring unpinned: putting Fleetd.main back on the flat constructor left
all 1608 tests green, so the factory was pinned and the decision to use it
was not. FleetdBackendQuarantineWiringTest closes that, following the five
existing *WiringTest files rather than inventing an idiom. It checks source
TEXT, and its class doc says so: it does not prove the call executes, and it
cannot tell "wrong factory" apart from "renamed the anchor" - both fail the
same assertion. That limit is real and recorded rather than papered over.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:51:00 +07:00
Dai Ha 01462c9695 fleetd #466 follow-up: pin main's choice of the escalating quarantine factory
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m25s
BackendQuarantineTest proves BackendQuarantine.withEscalation itself
escalates, is ceilinged, and resets. Nothing proved Fleetd.main still
calls it: reverting main to the flat two-argument constructor compiled
with 0 errors and left the full 1608-test suite green, because every
other test builds its own BackendQuarantine directly.

Adds FleetdBackendQuarantineWiringTest, a source-text assertion on
Fleetd.java (same idiom as FleetdLeadSeatWiringTest and
FleetdCompletionResolverWiringTest) pinning that main's declaration is
built from BackendQuarantine.withEscalation(...), not `new
BackendQuarantine(...)`. Measured: reverting to the flat constructor
fails it (expected true, was false); renaming the anchored local
variable while keeping behaviour identical also fails it loudly, not
silently. The class doc states plainly this checks source text only,
not that the call executes or that escalation reaches a real backend.
2026-09-10 19:47:59 +07:00
Dai Ha 8b4ff78546 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things I want on the record because they are judgement calls, not
facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this
  codebase reports a spawn success back to this class, so "it started
  working again" cannot be observed here. A base cooldown of quiet is the
  best available evidence. The class doc says this plainly rather than
  implying the stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred classification question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and is deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - it is the
same formula with multiplier 1.0 and a ceiling equal to the base, which
collapses to the original flat "now + cooldown". Every existing call site
keeps its shape.

Build number and my own mutation results are reported in the PR and on the
ticket, measured on this merge commit rather than on the branch.
2026-09-10 19:40:15 +07:00
Dai Ha d4a2cd720c fleetd #393: deliver memberSkills to opencode members, and stop overclaiming seeding success
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m46s
GitWorktrees.seedSkills copies memberSkills:-seeded skill folders into every
provisioned worktree's .claude/skills/ and logged "skill seeding: N of M" as
if that were success — but .claude/skills/ is a Claude Code CLI convention.
opencode has no such discovery, so a kind: opencode member never actually
read a seeded skill even though the log said N of M succeeded.

Two changes, both required:

1. Deliver it. OpenCodeLauncher.skillInstructionFiles scans
   <cwd>/.claude/skills/*/SKILL.md at spawn time (the one point the launcher
   knows both the kind and the cwd) and appends each to the generated
   opencode.json's instructions[] array, the same channel already used for
   the member charter and IDE rules. A skill folder with no SKILL.md is
   named and skipped rather than silently dropped.

2. Stop claiming it where the claim can't be verified. GitWorktrees.seedSkills'
   log now says explicitly that consumption depends on the member's kind and
   points at the launcher's own log; OpenCodeLauncher logs its own kind-aware
   "skill delivery: M of N ..." line once the kind is actually known, naming
   any folder it could not turn into an instructions[] entry.

fleetd.example.yaml's memberSkills: doc previously claimed "Claude Code
members only; an opencode member reads a different path (.opencode/agent)
this key does not touch" — false as of this fix, corrected to name both
kinds and how each consumes it.

Tests: OpenCodeLauncherTest gains two cases driving the real
GitWorktrees#add seeding path (not a hand-built fixture) into an
opencode-kind spawn — one asserting a seeded skill's SKILL.md lands in
instructions[] plus the honest log line, one covering a skill folder
without SKILL.md (delivered skills still flow, the malformed one is named
in the log and excluded from instructions[]). ClaudeCodeLauncher is
untouched — its native .claude/skills/ discovery already worked and is out
of scope.

mvn -B clean test: Tests run: 1603, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
2026-09-10 19:37:00 +07:00
Dai Ha 5a467e1f8b fleetd #466: escalate BackendQuarantine's cooldown on repeated exhaustion
CI / build (pull_request) Successful in 1m34s
CI / contract (pull_request) Successful in 1m36s
A flat 30-minute quarantine retries a weekly subscription limit about
336 times before the window resets. BackendQuarantine now doubles the
cooldown on each consecutive exhaustion of the same credential (no
more than one base cooldown after the previous quarantine's deadline),
capped at 12x the base cooldown (~6h at the 1800s default), and resets
back to the base cooldown once a base-cooldown's worth of quiet has
passed with no further exhaustion.

The flat two-argument constructor is unchanged (equivalent to
multiplier 1.0 / ceiling == base), so all ~20 existing call sites keep
their current shape and behaviour. Production wiring (Fleetd.main)
switches to the new BackendQuarantine.withEscalation factory.

This only touches the exhaustion path (BackendQuarantine's one
production caller is Fleetd.exhaustionSink, fired on BACKEND_EXHAUSTED
alone) and never the separate, unescalated cooling-off mechanism
(BackendOutagePolicy, fixed 60s) that guards against a transient
backend-error storm.
2026-09-10 19:33:09 +07:00
Dai Ha 1fb6176783 Merge #457: exhaustedPattern goes hot, and the warning names the fix (fleetd #446)
CI / contract (push) Successful in 1m31s
CI / build (push) Successful in 1m33s
Three rounds. Round 1 made exhaustedPattern a hot config key, made the
quarantine warning name the fix instead of only the fact, and added model/reason
to fleet_profiles' quarantined rows. Round 2 extracted the warning text into
usageLimitFixWarning/usageLimitFixWarningNoModel and pinned both. Round 3
extracted the sink itself into a static exhaustionSink(...) factory and pinned
what it actually logs, using a ListAppender on this class's own logger.

Where the mutation numbers below come from, stated exactly. The battery ran on
merge commit 3d2d521, tree 1dcca23: origin/main at 235644c plus this branch.
This merge commit's tree is 953ce11, which is that tree plus one file --
CharterToolSurfaceTest, from the #464 merge (49df792) that landed on main while
the battery was running. So the battery did not run on this exact tree. That one
added file was verified green on its own merge. The build on THIS tree, tree
953ce11, is: Tests run: 1601, Failures: 0, Errors: 0, Skipped: 0, BUILD
SUCCESS, 0 compile-error blocks. That is the battery tree's 1600 plus that one
test, which is the arithmetic the two trees predict. Saying this rather than
implying one tree.

On the battery tree: 1600 tests green, 0 compile-error blocks. FleetMcp.java
auto-merged there against #463's change to the same file, so that build was also
the gate on the auto-merge -- a clean auto-merge is not a compiling merge.

- M5, the round-2 survivor reproduced verbatim: the call site stops using either
  pinned method and logs a literal instead. Round 2 left this green across 1592
  tests. Now KILLED by
  FleetdExhaustionSinkWarningTest.profileWithModelGetsTheActionableFix and
  .profileWithNoModelGetsTheFallbackNotAFix.
- M6, the ternary's two branches swapped, so a profile WITH a model gets the
  no-model text and vice versa. Both extracted methods and both their unit tests
  untouched. KILLED by the same two tests, independently.
- M7, THE HALF THIS ROUND DID NOT PIN, and it SURVIVED. main()'s call to the
  factory replaced with an inert lambda: the factory and its test stay perfect,
  the daemon quarantines nothing and logs nothing. 1600 green.

M7 is the same defect shape as round 2, moved one level up, and it is worth
naming plainly. Extracting a thing in order to pin it CREATES the seam the test
then lands on. Round 2 extracted the warning TEXT, pinned the two methods, and
left the call site that selects between them unpinned -- M5. Round 3 extracted
the SINK, pinned the factory's behaviour, and left main()'s wiring of it
unpinned -- M7. The question to ask on any such fix is which of three things a
test now reaches: the value, the call site, or the selection between values.
Extraction only ever answers the first.

M7 is not a reason to hold this merge. Proving main() wires this factory means
driving daemon startup, which nothing here does -- that is fleetd #460. A
cheaper option exists and is recorded there: a source-reading assertion of the
kind #439 used, which would kill M7 without starting the daemon.

Round 3's own caveat, disclosed by the worker and worth keeping: its first Cell A
run was corrupted by a second concurrent mvn against the same module directory.
It killed that run, checked for stray ForkedBooter processes, and re-ran. The
numbers above are mine, from this battery, not that run.
2026-09-10 19:16:24 +07:00
Dai Ha 49df79203c Merge #464: a test that charter text names only registered tools (fleetd #464)
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m56s
CharterToolSurfaceTest extracts every fleet_* / bridge_* token from configured
launch charters and every tool("fleet_...") FleetMcp registers, then asserts the
first set is a subset of the second.

Verified on the merge commit. Its three acceptance criteria are met:

- Catches the ticket's own example. Fixture charter naming bridge_send, the tool
  CB-634 renamed away: KILLED.
- Fails loudly with no charter text. Fixture stripped of every tool name: KILLED
  by its named.isEmpty() guard, not a silent pass.
- Fails loudly with no registered tools. The tool("...") scrape broken so it
  matches nothing: KILLED by its registered.isEmpty() guard.
- And one cell of my own: the server stops registering fleet_reply, which the
  fixture names. KILLED. This is what proves the 'registered' half reads real
  production source and is not a second fixture.

The scrape finds 11 registered tools: ack, ask, list, poll, profiles, reply,
send, spawn, status, stop, whoami. An independent count of every "fleet_x"
literal in FleetMcp.java is also 11.

WHAT THIS DOES NOT CLOSE, and it is the ticket's actual gap. The charter half is
a @TempDir fixture the test writes itself, so no charter text anyone writes can
make this test fail. Measured: the test mentions fleetd.yaml 0 times, and the
commit changes 0 production files -- FleetConfig.validateCharters() still never
reads charter text (0 lines of its body mention a tool name). So this pins the
comparison logic and acts as a rename tripwire for the two tools the fixture
names. It does not check the live config. That needs a production-side check and
is filed as a follow-up.

That residue is my ticket's fault, not the worker's: the three criteria I wrote
are exactly the three it met.

A note on my own battery, because it nearly published four false kills. The
first run showed rc=1 on all four mutation cells and I would have read that as
four kills. It was zsh: unquoted parameters are NOT word-split, so
'mvn -B $scope test' passed '-Dtest=X -DfailIfNoTests=false' as ONE argument and
surefire ran zero tests. The tell was a missing 'Tests run:' line. The rerun
proves the harness first -- selector alone must report 'Tests run: 1' -- and
every cell now prints its surefire summary count so a void cell cannot pass for
a kill.
2026-09-10 19:10:56 +07:00
Dai Ha 235644c0f0 Merge #467: listFleet's callerIsPrimary default fails closed (fleetd #463)
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m38s
A wrapper overload that is called with no callerIsPrimary argument used to
default it to true, so a forgotten argument silently handed out the lead's
coordination state. It now defaults to false: a missing identity fails closed.

Verified on the merge commit, not the branch:

- The funnel is real. FleetMcp has 7 listFleet declarations and 7 real calls
  (an 8th 'listFleet(' match is a javadoc {@link}). Exactly one call writes a
  literal for the new boolean, and it writes false; exactly one writes the real
  predicate, coordinatorVisibleTo(principal(exchange)) in the MCP handler. No
  call writes true.
- Control battery on the merge: 1584 tests green unmutated.
- M1, the fix reverted at the one line that writes the default (false -> true):
  KILLED by listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey.
- M2, the half this round weakened. The worker rewrote
  listIsByteForByteUnchangedForThePrimaryCaller and dropped its byte-for-byte
  equality assertion, which was correct because that comparison ran against the
  implicit-default overload -- now the path #463 closes. So: does anything still
  notice if the primary's coordinator row silently loses a field? Dropped
  heldCount: KILLED by two tests, that same test and
  listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero.
- M3, a regression check on #439's own gate, 'if (callerIsPrimary)' -> 'if (true)':
  KILLED by three tests.

What is no longer pinned, stated plainly: the primary's output is now checked
field by field, not as a whole string. A field that no test names could
disappear without failing anything. Every field the tests do name is pinned,
proven by M2. The whole-output answer belongs to fleetd #460.

One control in my own battery was wrong and is worth recording: I labelled
', true);' in FleetMcp.java 'must be 0' and it is 2 -- row.put("configured",
true) and m.put("self", true), neither a listFleet delegation. The pattern was
too wide. The count above comes from a walk over each declaration and call
instead.
2026-09-10 19:04:51 +07:00
Dai Ha 7772b41993 fleetd #446 round 3: extract exhaustionSink and pin its caller (Cell A/B)
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 2m13s
Round 2 pinned usageLimitFixWarning/usageLimitFixWarningNoModel's TEXT via
FleetdUsageLimitFixWarningTest, but a mutation battery against the merged PR
proved two gaps in the caller that builds main()'s real quarantine
ExhaustionSink: nothing proved the sink's log.warn actually invokes either
method (Cell A), and nothing proved it picks the right one for a profile
with vs without a configured model: (Cell B).

Extract the inline lambda into a new static Fleetd.exhaustionSink(...)
factory (same refactor-for-testability class the lead approved in round 2
for the two static warning methods), behaviourally unchanged from the
lambda it replaces. FleetdExhaustionSinkWarningTest drives this factory's
return value directly and asserts on the real text a ListAppender attached
to Fleetd's own logger captures, covering both cells from one mechanism:

- Cell A (ternary result replaced by a literal string): confirmed red,
  1594 run / 1 failure, restored, confirmed green (1594/0).
- Cell B (ternary's two branches swapped): confirmed red, 1594 run /
  2 failures (both test methods independently caught it), restored,
  confirmed green (1594/0).
2026-09-10 19:00:07 +07:00
lead eccd0548ce Merge #465: state canonical invariant 5 as a purpose, not a banned tool (fleetd #458)
CI / contract (push) Successful in 52s
CI / build (push) Successful in 2m1s
Verified by me on a local merge of 29e7a06 onto f5e02fe:

- One file, one hunk. 5 changed lines inside the canonical block, all of them
  invariant 5; a line-by-line diff of the block against origin/main shows
  nothing else moved.
- The addendum is untouched: both 'Herdr socket tests (measured 2026-09-10)'
  and 'Only the lead can run that check' appear once on each side.
- Block size: origin/main 17467 chars / 17626 bytes, the merge 17557 chars /
  17718 bytes. The worker's numbers were bytes and were right; I am naming both
  units because len(str) in python is characters and this block is full of em
  dashes, which is a mistake I have made before.
- mvn -B clean test: Tests run: 1583, Failures: 0, Errors: 0, Skipped: 0.
  BUILD SUCCESS, rc=0.

The three things I reserved from the worker, because a provisioned worktree
cannot do them:

- wiki/7-Use-Cases.md updated to match, in the wiki repo's own commit.
- The sync script run in this clone: in sync: True.
- Other projects carrying the block: NONE. 29 CLAUDE.md files under the
  operator's project roots, 18 of them carry the block and 17 of those are this
  repo's own worktrees, which inherit it from git. Only
  LTMS/claude-bridge/CLAUDE.md is a real second copy, and it is this file.

On the worker's criterion-4 question: the addendum note stays. The restated
invariant removes the unsatisfiability, which was the sharp problem, but the
note still names the exact four files the carve-out covers and records why
(fleetd #449, where a fake and the real herdr disagreed for weeks under a green
suite). A rule that is merely satisfiable is not the same as a rule a worker can
apply without re-deriving it.
2026-09-10 18:57:24 +07:00
Dai Ha c4e23eebad fleetd #464: guard charter tool names
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m37s
2026-09-10 18:55:16 +07:00
Dai Ha 7df503dfc2 fleetd #463: default listFleet's callerIsPrimary to false, fail closed
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 2m9s
The compat overload at FleetMcp.java:1373 defaulted callerIsPrimary to a
literal true, so a caller that forgot the argument silently got the
coordinator row (this daemon's coord-id, mailbox state, held-mail previews,
peer reachability) -- lead-to-lead state fleetd #439 just gated. Flip the
default to false: a forgotten argument now yields a missing row instead of
a leaked one.

Six FleetMcpTest methods relied on the implicit true to see the coordinator
row at all; they now pass true explicitly through the canonical overload.
listIsByteForByteUnchangedForThePrimaryCaller's own premise (comparing the
implicit-default path against an explicit-true path) was the shape of the
bug, so it now only exercises the explicit-true path.

Added listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey
to pin the new default: a compat overload called with no callerIsPrimary
argument, against a fully-configured lead channel, must produce a result
with the coordinator key absent -- not empty, not redacted, absent.
2026-09-10 18:55:13 +07:00
Dai Ha f5e02fedd6 plans: commit the fleet01 move plan, with today's state measured on the host
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m58s
This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.

Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:

- Phases 1 to 3 are done. Both systemd user units exist, are active and
  enabled, linger is on, one java process (so the old restart.sh
  double-daemon problem is gone), the live fleetd.yaml is on the host, and
  the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at 4887731, 77 commits behind
  origin/main, 0 ahead. So fleet01's daemon runs code from before this
  week's merges.
- Phases 6 and 7 are not done. The Mac still runs the daemon this fleet
  uses, and fleet01 has 2 weeks of uptime, so no reboot proof exists.
- Section 9 is out of date: claude is installed on fleet01 and a fleet01
  lead is live on the coordination channel, so the headless-login blocker
  it names is solved.

Also records fleet01's four profiles with their weights, because
placement: weighted plus gx at 100 and xf at 80 means an unqualified spawn
there almost never lands on a claude-code profile.

And one search that found the opposite of what I expected, written down so
the next session does not repeat it: opencode IS installed on fleet01, at
~/.opencode/bin/opencode, but only on the INTERACTIVE PATH. zsh -ic finds
it, zsh -lc does not. A herdr pane is a non-login interactive zsh and
fleetd types the launch command into the pane, so gx and xf are not broken
by this. The daemon's own login-shell ExecStart cannot see it, which is the
mirror of the trap that ExecStart exists to fix: credentials live on the
login side, ~/.opencode/bin on the interactive side.
2026-09-10 18:53:11 +07:00
Dai Ha 29e7a06c49 fleetd #458: restate invariant 5 by purpose, not mechanism
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Failing after 1m56s
Invariant 5 banned 'driving the terminal multiplexer directly', naming
herdr CLI and socket as the banned tool. That bans a mechanism. What it
protects is the control plane: nobody may move a fleet session, pane or
peer by a route that skips the bridge's policy checks.

In this repo herdr is itself the subject under test, so four contract
tests must open its socket on purpose (see the addendum's 'Herdr socket
tests' note). Under the old wording, a worker assigned to that code
reads invariant 5 and finds its only path to finish the task banned.

Restate the invariant by purpose: never move fleet state except through
the bridge. herdr's CLI and socket stay as the named example of the
banned route, not the definition of it.
2026-09-10 18:50:41 +07:00
Dai Ha 92a96fcbd8 Merge #462: gate the coordinator row on a named predicate the handler must consult (fleetd #439)
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 1m51s
Verified by me on a local merge of c1ca627 onto 1348287:

- mvn -B clean test: see the totals below. Ran in fleetd/, redirected to a file,
  exit code captured on its own line.
- Mutation battery on merge 8402923 (same two parents, earlier base 5f1b260),
  4 cells, each with a proof gate on the occurrence count:
    * CONTROL, unmutated: 1583 tests, 0 failures, rc=0.
    * M2r - the handler stops asking who called: coordinatorVisibleTo(principal(exchange))
      replaced by a literal true. KILLED by
      FleetMcpAuthzTest.theFleetListHandlerActuallyConsultsCoordinatorVisibleTo.
      This is the mutant that survived round 1, so the gap the follow-up was for is closed.
    * M3 - is the new source-reading detector vacuous? Renamed its anchor
      (listHandler -> listHandlerX, 2 sites, behaviour identical). The detector FAILED,
      rc=1, as a source-reading test must: it cannot silently pass on an empty scrape.
    * M4 - the predicate itself always says yes: return caller.isPrimary() replaced by
      return true. KILLED by FleetMcpAuthzTest.onlyThePrimaryMaySeeTheCoordinatorRow.
  Tree verified clean before the battery and restored after each cell.
- ANON is pinned too, not only worker and architect: FleetMcpAuthzTest asserts
  coordinatorVisibleTo(ANON) is false.
- listFleet( appears in exactly one main file, mcp/FleetMcp.java, and no REST class builds
  the coordinator row, so the detector's single-file scope covers every live call site
  today. Control for that sweep: 109 main .java files matched a string they all contain.

Not in this PR, and my call, not the worker's: the six compat overloads of listFleet still
default callerIsPrimary = true, which fails open. Safe today because the one production
call site passes the computed value. Filed separately.
2026-09-10 18:40:48 +07:00
Dai Ha 13482872bb contract test: keep the loaded run's passes, discard only its failure
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 2m6s
The javadoc threw away the whole loaded run as "not clean evidence". The
fixed cell order makes that too strong in one direction. A cell that runs
last on a climbing load has a free explanation for FAILING. It has no free
explanation for PASSING: surviving a worse condition than a fair order
would have given it is evidence in the safe direction. So the old version's
0 of 3 is still discarded, and the three passes are kept with the load each
one ran at.

Also names the cell that tests the swallow explanation head-on, which the
old text left as "nobody has managed that yet". SHELL_READY_TIMEOUT_MS at 0
types input at once — the worst case for "typed before the prompt" — and it
passed 3 of 3 at load 18.42 to 23.65. The wider read window cannot explain
that away, because a swallowed keystroke is lost, not late: the command
never runs, so no amount of polling makes its output appear.

And a warning not to carry the raw load average to another host. Load
average counts differently per core and per operating system, so only load
per core compares. I broke that rule myself when comparing this run with
another host's numbers.

Both points came from the fleet01 lead reviewing 2af13ab.
2026-09-10 18:37:41 +07:00
Dai Ha c1ca6273fc fleetd #439: pin the caller at the fleet_list call site, not just the gate
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m33s
Review of PR #462 found M2: the coordinatorVisibleTo gate (then an inline
principal(exchange).isPrimary() check) could survive a mutation that
replaced the argument with a literal true at the one production call
site, because every existing test drove listFleet directly and supplied
the boolean itself -- nothing exercised the handler's own call.

- Name the decision: FleetMcp.coordinatorVisibleTo(Principal), a small
  package-private predicate next to denyFor/recordPrimarySingleton. The
  fleet_list handler now calls coordinatorVisibleTo(principal(exchange))
  instead of inlining .isPrimary().
- Pin the predicate's role table in FleetMcpAuthzTest
  (onlyThePrimaryMaySeeTheCoordinatorRow), covering primary/worker/
  architect and, newly, anonymous.
- Add a source-reading detector at the boundary
  (theFleetListHandlerActuallyConsultsCoordinatorVisibleTo), same idiom as
  toolsTheServerRegisters/everyRegisteredToolHasItsHandlerActionPinned: it
  reads FleetMcp.java, isolates the listHandler block, asserts (as a
  control) that the block actually contains a listFleet( call, then
  asserts the call's trailing boolean argument is exactly
  coordinatorVisibleTo(principal(exchange)) -- not a literal true/false.

Both mutations from the review were reproduced and killed by these tests,
then reverted; see the PR body for the full break-and-restore transcript.
2026-09-10 18:28:16 +07:00
Dai Ha 5f1b260c81 addendum: the canonical sync check is the lead's, not a member's
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 1m45s
I put "the sync script must print in sync: True" in a worker's acceptance
criteria for #455. The worker could not run it and said so, honestly, instead
of inventing a pass. My brief was the defect.

A member's provisioned worktree has wiki/ uninitialized, so the script dies
with FileNotFoundError. Measured in three worker worktrees: git submodule
status printed a leading '-' and wiki/ held 0 entries. The primary's own clone
printed a leading '+' and the file was there.

The note is dated, gives the re-measure command, says what each outcome means,
and says to delete it once it stops reproducing - as this file requires of any
measurement in an addendum.

Canonical block untouched: the sync script itself reports in sync: True.
2026-09-10 18:25:32 +07:00
Dai Ha 30d6872779 fleetd #446 follow-up: pin the WARNING text and the fleet_profiles model/reason fields
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m4s
A mutation battery run against merged PR #457 (d02dd1b) proved criteria 2 and 3
shipped without a test that could catch them breaking: deleting either
row.put("model", model) or row.put("reason", reason) in FleetMcp.profilesView
left all 1586 tests green (rc=0), and renaming the "usage-limit fix:" log tag
to something meaningless did too. Criterion 1's own mutation (LiveExhaustedPatterns
snapshotting instead of reading live) was correctly killed by the existing
FleetdExhaustionDetectionArmedWiringTest/LiveExhaustedPatternsTest — only 2 and 3
were unguarded.

- Extracted the exhaustionSink WARNING text out of two inline SLF4J {}-placeholder
  log.warn calls into two static methods, Fleetd.usageLimitFixWarning(profile, model)
  and Fleetd.usageLimitFixWarningNoModel(profile, quarantineCooldownSeconds), the
  same extracted-static-method + dedicated-test idiom as
  exhaustedPatternCoverageLine/errorPatternCoverageLine. log.warn is now called with
  each method's return value as a single already-formatted argument, so the string a
  test asserts on is byte-identical to what fleetd.out receives. No behaviour change:
  same text, same two branches, same call site.
- New FleetdUsageLimitFixWarningTest pins the leading "usage-limit fix:" grep tag and
  three specific facts (profile name, model name, "enabled: false" under
  models.allow) rather than the whole sentence, plus the no-model fallback's profile
  name and cooldown-seconds substitution and its explicit absence of "enabled: false".
- New FleetProfilesQuarantineModelReasonFieldsTest exercises FleetMcp.QuarantineSource
  with a modelFor/reasonFor that actually return values (every existing test used
  QuarantineSource.none() or a 3-arg form defaulting both to null), asserting the
  quarantined row's model/reason are present when supplied and absent (not null, not
  blank) when modelFor/reasonFor return null or a blank string.
- Each of the three: implemented, broken by hand (row.put deleted / tag renamed),
  confirmed the new test goes red, restored, confirmed green again. See PR body and
  this ticket's fleet_reply for the verbatim failure output of all three.
2026-09-10 18:22:51 +07:00
ltms 9d1306d442 Merge #461: project addendum for the herdr socket carve-out (fleetd #455)
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m56s
Doc-only, one file, 13 added lines, inside §Project addendum. Verified by me:

- Canonical block byte-identical: 17468 chars on main and on the branch.
- The sync script in CLAUDE.md prints "in sync: True" both before and after.
- The note's own re-measure command works. Run verbatim, it returns exactly the 4 files
  the note names, and no others:
    fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java
  Control that the search reaches the tree: 120 test java files, 63 of them mention herdr.
  That control matters here — a zero-match re-measure command would tell a future session to
  delete a live restriction.
- All four required elements are present: the carve-out, what stays banned (the control plane),
  who it applies to, and the perishable half (dated 2026-09-10, the command, what each outcome
  means, and delete-when-stale).
- Cross-references #458 for the canonical restatement, which is deliberately not in this change.

Pre-send check applied to the note itself: a worker assigned to AgentControlContractTest can now
name one legal action that finishes its task — let the test open the herdr socket in a throwaway
workspace it tears down.
2026-09-10 13:18:17 +02:00
Dai Ha e54e3d87ea fleetd #439: omit fleet_list's coordinator key for non-primary callers
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
The coordinator row is lead-to-lead coordination state (coord-ids, mailbox
facts, held-message previews). fleet_list returned it to every caller,
including a worker or an architect, because coordinatorView() had no way to
know who was asking.

Gate at the call site inside listFleet: a new overload takes
callerIsPrimary and only assembles/attaches the coordinator row when it is
true, so the key is absent (not empty) for a worker or an architect. The
MCP handler now passes principal(exchange).isPrimary(); every other
listFleet overload keeps passing true, so callers with no caller identity
(existing unit tests, the no-op wrappers) are unaffected -- confirmed by a
byte-for-byte comparison test against the pre-fix overload.

Authz's READ case is untouched: it stays shared by fleet_status,
fleet_profiles and fleet_whoami, and the gate here is purely inside
fleet_list's own result assembly.
2026-09-10 18:16:48 +07:00
Dai Ha bf027f10b9 #455: document herdr test socket exception
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m50s
2026-09-10 18:13:50 +07:00
Dai Ha 3c5873dfe2 fleetd #453: point the new javadoc at the right javadoc
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m52s
#456's first paragraph said "HerdrPeerLauncher's own {@link
#spawn(SpawnRequest, PlacementDecision)} javadoc". That link resolves to this
interface's own abstract declaration, not to HerdrPeerLauncher's override, so a
reader who follows it lands on the wrong text. The second paragraph already
used the plain {@code HerdrPeerLauncher.spawn(...)} form; both now match.

Also says who is actually forced to read the paragraph, because #456's
reasoning rests on it and the two cases differ. A class that implements this
interface directly must write a body for spawn(SpawnRequest,
PlacementDecision) - it is abstract here - so it reads this javadoc. A
subclass of HerdrPeerLauncher does not: HerdrPeerLauncher already implements
that method (member/HerdrPeerLauncher.java:611) and the subclass inherits the
body. For a subclass the paragraph is advice, not a gate.

javadoc -Ddoclint=reference: 5 "reference not found", the same 5 in the same 5
untouched files as origin/main at 2af13ab, and none in PeerLauncher.java. Those
5 are ticket #459. mvn compile rc=0.
2026-09-10 18:05:07 +07:00
ltms e29227d5f4 Merge #456: document the override obligation on PeerLauncher.defaultProfileFor/place (fleetd #453)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m34s
Doc-only, one file, 17 added lines. Verified by me on a local merge of 0788d84
onto 2af13ab (merge commit 5b46538):

- mvn -B clean test: Tests run: 1578, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS, rc=0.
- javadoc reference lint (mvn javadoc:javadoc -Ddoclint=reference): the merge has 5
  "reference not found" errors; origin/main at 2af13ab has the same 5, in the same 5
  untouched files. So this PR adds no broken link. Both runs rc=1 for that pre-existing
  reason, which is now ticket #459.
- The quoted phrase is real: "are unoverridden here and just wrap {@link #defaultProfile()}"
  is at member/HerdrPeerLauncher.java:604. place() is not overloaded, so {@link #place}
  is unambiguous.

One inaccuracy I am fixing in a follow-up commit rather than sending the PR back: the first
new paragraph writes "HerdrPeerLauncher's own {@link #spawn(SpawnRequest, PlacementDecision)}
javadoc", but that link resolves to PeerLauncher's own abstract declaration, not to
HerdrPeerLauncher's override. The second paragraph already gets this right with plain
{@code HerdrPeerLauncher.spawn(...)}.
2026-09-10 13:04:09 +02:00
Dai Ha 45aca9eb3e fleetd #446: make exhaustedPattern hot, name the fix in the warning, report it in fleet_profiles
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 1m34s
The model gate can be turned off at runtime (models.allow[].enabled: false, hot
since fleetd #422), but the usage-limit detector it's meant to react to was
compiled once at Fleetd.main startup into a frozen Map<String,Pattern> — arming
or disarming exhaustedPattern needed a daemon restart. Backwards for a feature
meant to react live.

- New LiveExhaustedPatterns: reads exhaustedPattern off the live config supplier
  per lookup (matching CompositePeerLauncher#models0's live-supplier pattern),
  caching compiled Pattern objects by PROFILE NAME (not pattern text — pattern
  text would grow unboundedly as an operator tunes a regex across reloads;
  profile names are bounded by the small, restart-gated set of configured
  profiles). patternFor() backs CompletionResolver's classification; armed()
  backs fleet_profiles' exhaustionDetectionArmed — both read the same object,
  the fleetd #404 single-accessor rule CompositePeerLauncher.modelGateState()
  established for the model gate.
- Fleetd.java: on BACKEND_EXHAUSTED, log a WARNING naming the profile's model
  and the exact fix (enabled: false under models.allow, hot, no restart; remove
  it again once the window resets) — or, when the profile has no model:
  configured, say quarantine is the only thing keeping spawns off it.
- FleetMcp.QuarantineSource gains modelFor/reasonFor; fleet_profiles'
  quarantined rows gain model/reason fields so a lead can see why without
  reading the daemon log. capacityView (fleet_list) intentionally untouched —
  scoped to fleet_profiles only.
- ConfigRef/FleetConfig docs + fleetd.example.yaml updated: exhaustedPattern
  moves from Deferred to Hot. errorPattern stays deferred on purpose (out of
  scope for this ticket).
- Tests: LiveExhaustedPatternsTest (new, unit-level hotness/caching proof),
  FleetdExhaustionDetectionArmedWiringTest (rewritten — same QuarantineSource
  object, read before and after a reload, asserts the answer flips with no
  restart), ConfigRefTest/ConfigRefProfileCoverageTest updated for the new
  Hot/Deferred classification.
2026-09-10 18:02:00 +07:00
Dai Ha 2af13ab1ff fleetd #449: name the decisive cell and the load the claim was measured at
CI / contract (push) Successful in 54s
CI / build (push) Successful in 2m8s
The previous comment named a cause with no cell behind it. fleet01's review
made the point: a confident wrong mechanism gets copied, and a confident
under-determined one gets copied the same way.

So the comment now names the cell that settles it. Hold the old 1000ms write
sleep and change only the read - 800ms fixed sleep becomes a 5s poll - and the
test goes 0 of 3 to 3 of 3. The read deadline was the whole story.

It also says where: a 12-core macOS host near idle (load 2.6 to 5.9). The
loaded run agreed, but its load climbed from 7 to 50 while the cells ran and
the old version ran last, so it is not clean evidence and the comment says so.

Comment only. No test or production code changed.
2026-09-10 17:59:52 +07:00
Dai Ha 0788d84be8 fleetd #453: document the override obligation on PeerLauncher.defaultProfileFor/place
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m34s
Decision: leave both as default methods (option 1), not abstract. Neither
default is a live defect today — HerdrPeerLauncher is the sole single-profile
implementer and the degenerate answer (ignore role, always defaultProfile())
is correct for it. Making them abstract would force ~10 boilerplate one-line
overrides across 5 unrelated PeerLauncher test doubles (NeverSpawnsLauncher,
RaceLauncher, NoResumeLauncher, ClearContextSpyLauncher, LazyIdLauncher) that
never call either method, for a risk that is speculative (no multi-profile
HerdrPeerLauncher subclass exists or is planned).

Strengthens both javadocs with an explicit MUST-override warning and cross-
references HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)'s existing
#450 javadoc, which already names both methods as "unoverridden here" and
ties that to being a single-profile adapter -- the concrete place a future
multi-profile launcher author would read, since #450 made that method
abstract and any subclass must write its body.
2026-09-10 17:49:37 +07:00
Dai Ha 20c1094cbf fleetd #449: say what the timing fix actually proved, not what it assumed
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m33s
The polling fix that landed in #452 is right, but its javadoc named a
mechanism nobody measured: that input typed before the shell's prompt was
swallowed by the shell's own startup.

I mutated the settle poll away — SHELL_READY_TIMEOUT_MS = 0, so input is
typed at once with no wait — and the test passed 3 of 3. So waitForText is
the load-bearing half, and the proven cause is the old 800ms READ deadline,
not the 1000ms write delay.

The direction is the point: typing at 0ms works where typing at 1000ms
failed. If early input were swallowed, 0ms would be worse than 1000ms. It is
better, so the swallow explanation is unsupported.

waitUntilSettled stays as cheap insurance, now labelled as insurance rather
than as the fix. Comment-only; AgentControlContractTest still green.
2026-09-10 17:36:22 +07:00
ltms 9011c59b9f Merge #452: run the contract tag in CI, fix the stale protocol 14 assertion (fleetd #449)
CI / build (push) Successful in 1m28s
CI / contract (push) Successful in 1m33s
Verified on a locally built merge onto c11ad71 (the PR was branched from 822327e, before #448):
1578 unit tests green, then `-Dgroups=contract` gives 30 tests / 0 failures / 0 skipped here.

CI's own run of the tag on d4f93a7: 29 tests, 0 failures, 6 skipped — all 6 herdr tests skip on
the runner, and LeadMailboxTest's 15 tests run there for the first time.

Two mutations killed: restoring `assertEquals(14` fails the test, and pointing the expected env
value at a wrong URL fails it with the real pane text in the message.
2026-09-10 12:35:10 +02:00
ltms c11ad71ed0 Merge #448: fleet_ack errors instead of claiming success on a miss (fleetd #437)
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m33s
Verified by the lead on head 5289eb5, and then again on the MERGE.

Battery on the branch tip:

  FULL BUILD  Tests run: 1577, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0
  CONTROL     contract test, real broker: Tests run: 9, Failures: 0, Skipped: 0

  M3a  AmqpReplyInbox.ack's `h == null` branch reports true
       -> KILLED  AmqpReplyInboxContractTest.ackReportsHitVsMissAgainstARealBroker
  M3b  AmqpReplyInbox.ack's `perTarget == null` branch reports true
       -> KILLED  same method

M3a is the gap I found in round 1: reporting true for something never held
survived both the default suite and -Pcontract against a real broker. It is now
closed in the adapter this daemon actually runs.

Both branches die, and they die through DIFFERENT assertions, which the worker
worked out and I confirmed by reading the code. held.get(target) is populated by
the deliver callback, not by own(), so after own() with nothing ever delivered
the map entry is still null — the "never held" assertion therefore exercises
`perTarget == null`, and only the double-ack assertion reaches `h == null`. Both
assertions are load-bearing; neither is redundant.

This PR was pushed before #451 landed, so the battery above tested the branch,
not the result. Merging main (bdcf285) in gave 0 conflicts and #448 adds no
PeerLauncher implementer, but a clean auto-merge is not a compiling merge, so I
built the merged tree 0829542:

  FULL BUILD               Tests run: 1578, Failures: 0  BUILD SUCCESS, 0 compile errors
  contract, real broker    Tests run: 9,    Failures: 0
  CompositePeerLauncherTest (#447's guarantee)  Tests run: 76, Failures: 0

The worker also declined to point the contract test at the shared local LavinMQ
broker, because this adapter never deletes queues and a run would leave orphaned
durable queues on the instance backing the live fleet. It started a disposable
rabbitmq:3.13-management container on a throwaway port instead, then removed it.
That was its own judgment and it was correct.

Held peer mail is deliberately NOT ackable: fleet_ack against a coord-id errors
and names fleet_poll{coordId}. See #437 for why refusing is the right answer
today, and why that is a policy choice rather than a structural one.
2026-09-10 12:20:49 +02:00
ltms bdcf285265 Merge #451: make PeerLauncher.spawn(req, decision) abstract (fleetd #450)
CI / contract (push) Successful in 1m17s
CI / build (push) Successful in 1m33s
Verified by the lead on head cfebc57 (base 822327e is current main).

  FULL BUILD  Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0

  M1  delete HerdrPeerLauncher's new override
      -> BUILD FAILURE, 1 COMPILATION ERROR block, naming exactly
         ClaudeCodeLauncher.java:[49,14] and OpenCodeLauncher.java:[59,14]

  M2  CONTROL for M1: put the interface method back to a `default` AND delete
      the override
      -> BUILD SUCCESS, 0 compile errors
      This is the row that makes M1 mean something. Without it, M1 only shows
      that the build broke; with it, the break is attributable to the method
      being abstract rather than to anything else the edit disturbed.

  M3  CompositePeerLauncher's override reverts to the re-entering form
      -> KILLED, Errors: 1
      CompositePeerLauncherTest
        .spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace
      So #447's guarantee survives this refactor of the interface it rests on.

Tree restored clean after each mutation (git status --porcelain empty).

The worker corrected my ticket, and it was right. My #450 body listed five
src/main implementers of PeerLauncher and quoted `grep -rln 'implements
PeerLauncher'` as the source; that command returns two files. I had run a wider
pattern that also matched a comment in ConfigRef and the `extends
HerdrPeerLauncher` line in two subclasses, then quoted the narrow command beside
the wide command's output. Ground truth: ConfigRef implements
Supplier<FleetConfig>; ClaudeCodeLauncher and OpenCodeLauncher extend
HerdrPeerLauncher. So one override in that parent serves both, which is what the
worker built. The ticket body is corrected.

Out-of-scope note carried forward from the worker: defaultProfileFor(MemberRole)
and place(MemberRole) are two more default methods with the same shape. Noted,
not fixed here.
2026-09-10 12:16:47 +02:00
Dai Ha d4f93a7b13 fleetd #449: fix stale herdr protocol 14 javadocs/assertion, diagnose and fix the timing-raced AgentControlContractTest, select contract tests by tag in CI
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m9s
- HerdrClient.java, HerdrCodec.java, HerdrContractTest.java: the herdr port to
  protocol 19 (CB-521) left the client javadoc and the contract test's own
  assertion still saying protocol 14 / herdr 0.7.0. Updated to 19 / 0.8.0 and
  renamed pingReturnsProtocol14 -> pingReturnsProtocol19. Verified the
  assertion is real by temporarily changing the expected value to 20 (fails),
  then restoring 19 (passes).

- AgentControlContractTest.java: tabCreateInjectsEnvIntoTheSeedShell was
  failing, not skipping, on a host with a live herdr socket. Diagnosed with a
  temporary instrumented run (not committed) that polled the pane every
  200ms before and after sending input: the seed shell reliably takes ~2.5s
  to reach its prompt (measured 3x), while the test's fixed 1000ms sleep
  raced that startup. Input typed too early was swallowed by the shell's own
  startup, leaving the typed line followed by the "Restored session" banner
  and no command output — indistinguishable at a glance from the env map
  never reaching the shell. Once the shell was actually ready, the injected
  env value showed up in ~200ms, ruling out an env-seam defect. Replaced both
  fixed sleeps with bounded polling on the actual conditions (pane text
  settling, then the expected output appearing). Ran the fixed test 3x
  standalone, all green.

- .gitea/workflows/ci.yml: the "Contract tests" step ran exactly one class by
  name (-Dtest=AmqpReplyInboxContractTest), silently excluding every other
  @Tag("contract") test from CI including the herdr ones above -- which is
  how the stale protocol 14 assertion went unnoticed. Changed to
  -Dgroups=contract, which selects the whole tagged group and picks up
  future contract tests automatically.
2026-09-10 17:14:09 +07:00
Dai Ha cfebc575ea fleetd #450: make PeerLauncher.spawn(SpawnRequest, PlacementDecision) abstract
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m5s
The default re-entered the single-argument spawn(SpawnRequest), which re-runs
checks that can refuse the profile place() just chose (#444's window). Only
CompositePeerLauncher overrode it; a future placement-doing launcher could
have inherited the wrong body silently.

Give every current implementer an explicit override, chosen by what it does:
- HerdrPeerLauncher (base of ClaudeCodeLauncher/OpenCodeLauncher, neither of
  which overrides spawn(req) or place()) does no placement filtering of its
  own, so it gets the re-entering form.
- CompositePeerLauncher's routing-form override is untouched.
- 5 test-fake PeerLauncher implementers (SessionManagerTest, FleetdBackendErrorSinkTest)
  get overrides matching their existing spawn(SpawnRequest) shape: delegating
  wrappers delegate, unreachable stubs throw, the single-profile fake re-enters.

ConfigRef does not implement PeerLauncher at all (confirmed in this tree at
822327e) despite the ticket listing it as a src/main implementer.
2026-09-10 17:10:44 +07:00
ltms 822327eed5 Merge #447: pin the place()-to-spawn() window PlacementDecision closes (fleetd #444)
CI / contract (push) Successful in 1m33s
CI / build (push) Successful in 1m36s
Verified by the lead on the exact tree that lands (head e4c703a, base 82fae94 is
an ancestor, so this is the tree I measured):

  FULL BUILD  Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0
  CONTROL     CompositePeerLauncherTest  Tests run: 76, Failures: 0  -> GREEN

  M1  the 2-arg spawn re-enters the 1-arg spawn (the inherited default this
      ticket forbids for a multi-profile launcher)
      -> KILLED  Errors: 1
      CompositePeerLauncherTest
        .spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace

  M2  drop the stamping: route to the decided profile but do not carry it
      (SpawnRequest routedReq = req)
      -> KILLED  Failures: 1
      same test method

Both mutations proved applied by printing the mutated method, and the tree was
restored clean after each (git status --porcelain empty).

Round 1 of this PR had a fixture weakness I found by mutation: StubLauncher's own
fallback default was "sol", the same profile place() decides, so an unstamped
request landed on spawnCount("sol") by coincidence and M2 survived. e4c703a gives
the adapter "b" as its fallback instead. One fixture now kills both mutations.

src/main is javadoc-only in this PR: 12 added lines, 0 added code lines, measured
by filtering the main-side diff.
2026-09-10 12:00:59 +02:00
Dai Ha 5289eb509f fleetd #437: pin the ack hit/miss contract in the AMQP contract test
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Successful in 1m34s
AmqpReplyInboxContractTest is the one contract-group class CI actually
runs, and it never asserted on ack()'s return value at all — so the
exact defect this ticket fixes (reporting success for an ack that
removed nothing) was unpinned in the adapter fleetd runs live.

Add ackReportsHitVsMissAgainstARealBroker: a msgId never held for an
owned target returns false without throwing, a real held reply returns
true and is removed, and acking the same msgId again returns false.
Ran against both broker modes the class supports: Testcontainers
(AMQP_URI unset) and an external broker via AMQP_URI (the CI shape,
using a disposable container — not the shared local LavinMQ instance).
2026-09-10 16:58:55 +07:00
Dai Ha e4c703a51a fleetd #444: separate the adapter's fallback default from the decided profile
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m56s
Review found the fixture's StubLauncher fell back to 'sol' too — the
same profile place() decides — so an UNSTAMPED request could land on
spawnCount('sol') by coincidence, and the assertion's claim that the
request 'actually carried sol' was unproven. Dropping the stamping
(SpawnRequest routedReq = req) while keeping the routing survived the
test unchanged.

Fix: give the adapter 'b' as its own fallback default instead, so an
unstamped request counts against 'b', not 'sol'. Verified both
mutations against the single test in isolation:
  - drop-stamping (routedReq = req): RED, expected <sol> but was <b>
  - re-entering (return spawn(req.withProfile(decision.profile())))
    i.e. M1 from the first round: still RED, PlacementException
    naming the now-quarantined 'sol'
Restored both; full suite green at 1575 tests.

No changes to src/main — PeerLauncher's javadoc from the first round
is unchanged.
2026-09-10 16:50:39 +07:00
Dai Ha 703a05db41 fleetd #437: fleet_ack errors instead of claiming success on a miss
CI / contract (pull_request) Successful in 1m25s
CI / build (pull_request) Successful in 1m37s
ReplyInbox.ack now returns boolean (true = removed, false = nothing to
remove) instead of void, so FleetMcp.ack can finally tell a hit from a
miss. FleetMcp.ack returns an error when the boolean is false, naming
fleet_poll{coordId} for held peer mail, which has no route through this
call. MessageService.ackReply propagates the boolean; drainReplies keeps
ignoring it (its own javadoc already documents that loss window as
deliberate). Updated the tool schema's target description to match.

Rewrote FleetMcpTest's ack tests to publish a real message before
asserting success, and added tests for a never-queued id and a coord-id
target, both now erroring. Added boolean assertions to
InMemoryReplyInboxTest's existing ack cases.
2026-09-10 16:44:08 +07:00
Dai Ha 3f036b2a62 fleetd #444: pin the place()-to-spawn() window PlacementDecision closes
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m53s
Add a test that resolves place(role) while nothing is quarantined,
then quarantines the resolved profile's credential BEFORE spawning
against the held PlacementDecision. CompositePeerLauncher.spawn(req,
decision) must still honor the decision and land on the quarantined
profile, since it never re-runs the explicit-profile enforce* checks.

Verified the test kills the regression: with the override's body
replaced by the re-entering spawn(req.withProfile(...)) form, this
exact test goes RED with a PlacementException naming the now-
quarantined profile; restored, the full suite is green (1575 tests).

Also documents on PeerLauncher's default spawn(req, decision) that a
launcher routing across more than one profile MUST override it,
naming the four enforce* checks the default's re-entry re-applies.
2026-09-10 16:41:52 +07:00
ltms 82fae94c55 Merge #445: pin every startup report call in Fleetd.main (fleetd #442)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m59s
Test written by a worker whose backend died before it could report; evidence
re-run by the lead against the merged tree.

Verified: merge of current main clean (0 conflicts); full build 1574 tests,
0 failures, BUILD SUCCESS, 0 compile errors; control green; deleting each of
reportGitHostShape, reportMemberTrustModel, reportMemberCredentialsGap and
reportExhaustedPatternGap from main() is KILLED by
mainReportsEveryStartupGapBeforeValidationAborts.
2026-09-10 11:30:31 +02:00
Dai Ha b1f34c2e6b fleetd #442: drop the unused java.util.List import
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m8s
The new test never names List — only ListAppender, which has its own
import. An unused import is an IDE warning, and this repo treats
warnings as gates. No behaviour change: FleetdStartupReportTest still
runs 1 test, 0 failures, BUILD SUCCESS, 0 compile errors.
2026-09-10 16:29:56 +07:00
ltms 3f807d9f1b Merge #443: derive coordinator.heldDurable from queue durability + ack mode (fleetd #440)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m35s
Found by the fleet01 lead reviewing #438 after I had merged it. Verified
independently before merging.

The implementation choice is the load-bearing part: LeadMailbox.own() now
assigns queueDeclare's durable flag and basicConsume's autoAck flag to named
locals, passes those SAME locals into the two real AMQP calls (:203/:204), and
derives heldDurable from them (:207). So the reported fact cannot drift from a
duplicate constant - a mutation to either call's argument moves the behaviour
and the report together. LeadChannel.heldDurable() is abstract, so a future
implementer gets a compile error rather than a silent default.

My own battery, merged tree, control green, tree restored clean:
- M1 revert to the literal true -> KILLED by
  FleetMcpTest.listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable
- M3 heldCount forced to 0 (a half the worker did not touch) -> KILLED by
  FleetMcpTest.listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero
- M2 break the derivation itself -> SURVIVED under plain clean install, exactly
  as the worker reported. LeadMailbox needs a real broker, so the only test that
  reaches the real queueDeclare/basicConsume is @Tag("contract"), excluded from
  the default build. The worker ran that arm with -Pcontract and got 3 reds
  including its own new test. Pre-existing structural limit of this class, not
  introduced here, and the worker flagged it rather than hiding it.

Full build, my own run: Tests run: 1573, Failures: 0, Errors: 0, Skipped: 0 -
BUILD SUCCESS, 0 compile errors.

Not blocking, noted for a possible follow-up: one boolean over two independent
facts cannot say WHICH fact was lost. The merged field is still strictly better
than the literal it replaces, because it can now go false at all.
2026-09-10 11:21:41 +02:00
Dai Ha e70263062c fleetd #442: pin startup report calls
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m36s
2026-09-10 14:41:10 +07:00
ltms 1a1e586b62 Merge #433: carry the PlacementDecision instead of re-resolving the profile (fleetd #425)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m47s
Round 4 pins the fix. Verified independently before merging.

My own battery in the worker's tree (control green, tree restored clean):
- M1 revert SessionManager:677 to launcher.spawn(spawnReq) -> KILLED by
  SessionManagerTest.acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedThe
  WorktreeForUnderARotatingPolicy. This is the ticket's own deliverable and it
  survived 186 tests in round 3.
- M3 drop the withProfile stamping in the 2-arg spawn -> KILLED by 3 tests.
- M2 make the 2-arg spawn re-enter the refusing branch -> SURVIVED, but it is a
  near-equivalent mutant, not a gap in this work. Post-#435 all four enforce*
  conditions are ones the routing branch already filtered on, so the two paths
  differ only if placement state moves between place() and spawn(). Filed
  separately.

Full build, my own run this turn, whole log redirected and grepped:
Tests run: 1572, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, 0 compile
errors. Branch already contains current main.

Read the src/main diff. The new test uses roundRobin() (stateful) and asserts
agreement between the overlay profile and the spawned profile, rather than a
hardcoded name, which is the right shape - select() is stateful, so two calls
disagree by design.
2026-09-10 09:35:02 +02:00
Dai Ha c16d118f09 fleetd #440: derive coordinator.heldDurable from queue durability + ack mode
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m36s
FleetMcp.coordinatorView wrote heldDurable as a literal true, so a change
that broke either the durable queue declare or the manual-ack consume in
LeadMailbox would leave the field, and the full suite, green.

- LeadChannel gets a new heldDurable() method: the conclusion of a durable
  queue declare AND a manual-ack consumer, derived by the implementation
  from what it actually did, never asserted.
- LeadMailbox.own() captures the exact booleans it passes to
  queueDeclare/basicConsume and stores their conjunction; heldDurable()
  returns it.
- FleetMcp.coordinatorView now reads channel.heldDurable() instead of a
  literal; updated the javadoc to say where the fact comes from.
- FakeLeadChannel gets a heldDurable field (default true) + withHeldDurable
  setter so FleetMcpTest can prove the field goes false.
- FleetMcpTest: new test asserts heldDurable:false when the channel says so.
- LeadMailboxTest (contract, real broker): new test asserts heldDurable()
  true against a real LeadMailbox. Verified by hand that flipping own()'s
  autoAck local to true turns this test (and two pre-existing redelivery
  tests) red, and restoring it turns them green again.
2026-09-10 14:31:44 +07:00
Dai Ha 4b10d02207 fleetd #425 rework round 4: mutation-pinning test for the dropped PlacementDecision
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m3s
SessionManager.acquireWithWorktree's unqualified branch must carry the
PlacementDecision it already resolved via launcher.place() into
launcher.spawn(spawnReq, decision) rather than re-deriving it through a
blank-profile launcher.spawn(spawnReq). Every existing test in this file uses
PlacementPolicies.fixed(), which answers select() the same way on every call,
so dropping the decision (handle = launcher.spawn(spawnReq);) was invisible:
186 tests stayed green under that mutation.

acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy
uses PlacementPolicies.roundRobin() instead — deterministic AND stateful, so
two select() calls on the same policy instance disagree (index 0 then index 1
across a two-profile pool). It asserts AGREEMENT between the profile the
worktree's parity overlay was provisioned for and the profile the member
actually spawned on, never a hardcoded expected profile name.

Verified as a real mutation, not a no-op: applying the exact mutation
(handle = launcher.spawn(spawnReq);) turns it red — expected [b.mcp.json]
but was [a.mcp.json] — and reverting turns it green again. Full build:
1572 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-09-10 14:24:14 +07:00
Dai Ha c5fbfdbf4a Merge main into #425 rework branch (brings #438 held-peer-mail read) 2026-09-10 14:11:37 +07:00
ltms 12cff28abb Merge #438: let a lead read its own held peer mail, primary-only (fleetd #421)
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 2m6s
fleet_poll{coordId} peeks this daemon's own held lead-to-lead mail and
returns full bodies without acking. New Authz.Action COORD_READ, primary
only — not the architect, which holds READ today. pollAction is now
argument-derived over both target and coordId.

heldView and HELD_PREVIEW_MAX_CHARS are untouched: fleet_list stays a
cheap always-safe scan, and the full read is a separately authorized call.

Verified by me, not taken from the report: merged tree builds 1562 green
(main 1555 + 7 new methods), 0 compile errors, no merge conflicts.
Mutation battery on lines the worker did NOT mutate, control 111 green:
  - COORD_READ widened to the architect  KILLED (AuthzTest + FleetMcpAuthzTest)
  - self-coord-id guard removed          KILLED (FleetMcpTest)
  - full body swapped for the preview    KILLED (FleetMcpTest)

The third mutation targets the ticket's own deliverable, and it is pinned.
Authz.permits has no default, so a new action is a compile error rather
than a silently unhandled case.

Follow-up filed as #439: fleet_list's coordinator row is still READ-gated,
so a worker sees peer coord-ids and 80-char previews of lead-to-lead
bodies. Pre-existing; the implementer flagged it and left it alone.
2026-09-10 09:08:17 +02:00
Dai Ha 77a6a7142e Merge main into #421 branch (brings #434 model-gate observability and #436 fixed-placement cap)
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m48s
2026-09-10 14:06:34 +07:00
Dai Ha 9f3671b801 fleetd #425 rework round 3: rewrite prose after #435 made fixed honour maxLoad
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m56s
fleetd #435 (merged to main) made FixedPlacementPolicy evaluate maxLoad during automatic
selection, the same way weighted/round-robin already did. Six comments across
CompositePeerLauncher.java, PeerLauncher.java, PlacementDecision.java, and SessionManager.java
justified round 2's place()/spawn(req, decision) mechanism by saying fixed "deliberately never
evaluates maxLoad" — that claim is now false, and needed restating, not just deleting.

The honest case after #435: the two-path shape (routing branch falls through an excluded
candidate; explicit-profile branch refuses on it) is still real and still deliberate — an
operator who names a profile should get a refusal, not a silent substitution. What round 1 got
wrong, and what round 2 still needs to prevent, is turning a fall-through into a refusal by
accident: resolving a name via place() and then feeding it back to spawn(SpawnRequest) as an
explicit profile. Before #435 that accident was reachable through maxLoad specifically, because
fixed never evaluated it; #435 closed that specific gap, so a PlacementDecision can no longer be
at-cap in the first place. What survives as the justification for spawn(req, decision): it never
re-evaluates a condition place() already decided, and it closes the window between that decision
and the spawn in which the underlying state could otherwise move — not a failure #435 already
prevents.

Re-measured the sibling paths this round exists to keep in agreement (one profile at maxLoad: 1,
liveCount pinned at 1, PlacementPolicies.fixed(), unqualified spawn): both the with-worktree and
without-worktree paths now throw the identical PlacementException — "worker profile 'a' is at
maxLoad (1 live >= 1 cap), and no available candidate remains" — closed upstream by #435, at
place()/select(), before either path ever reaches a spawn call. The observable asymmetry this PR
was filed to fix is gone; what remains is the structural argument above.

No behavior change: place()/PlacementDecision/spawn(req, decision) are untouched, and
FixedPlacementPolicy/PlacementPolicyUtil are taken wholesale from main's merge.
2026-09-10 14:06:10 +07:00
Dai Ha 84034b34d1 Merge remote-tracking branch 'origin/main' into worker/425-rework-placement-resolve-c58ba1-9 2026-09-10 13:56:54 +07:00
Dai Ha 1e9b2c9b7e fleetd #421: let a lead peek its own held peer mail, primary-only
CI / contract (pull_request) Successful in 1m27s
CI / build (pull_request) Successful in 1m36s
fleet_list truncated held lead-to-lead messages to an 80-char preview with
no way to read the full body, and fleet_poll{target} drained the wrong
inbox (a worker's reply queue, not the coordinator mailbox) -- it silently
returned []. fleet_ack would have destroyed the message unread.

Add a non-destructive read: fleet_poll{coordId} peeks (never acks) this
daemon's own held mail via LeadChannel.peek(). The coordId must equal the
caller's own selfCoordId -- passing a peer's id is refused with a reason,
instead of repeating the original silent-[] confusion.

This is authorization-sensitive: mapping it to the existing READ action
would let any worker read every peer lead's mail in full. READ's openness
rests on "the roster carries no secrets" (Authz.java), which does not hold
for lead-to-lead coordination bodies. Added Authz.Action.COORD_READ,
primary-only (not even the architect, which holds READ today), and made
pollAction's signature depend on both target and coordId so every call
site states explicitly what it passes.

Also fixes fleet_list's "pending: 0" trap: mailbox.pending only counts
broker-ready messages, so a healthy held mailbox reads as empty. Added
heldCount/heldDurable beside held[] so the durability fact isn't implied
only by reading the code.

Mutation-tested: pollAction's COORD_READ->READ mapping, the peek->ack
substitution, the 80-char preview cap widened to 81, and the Authz case
widened to include caller.isWorker() -- each breaks exactly its matching
test and nothing else. The first attempt at the preview-cap test used a
homogeneous "x"*200 body, which a widened cap slipped through unnoticed
(contains() found a shifted match); replaced with a sentinel character at
index 80 to actually pin the boundary.

Updates CLAUDE.md's intent->tool table for fleet_poll's new coordId
semantics, per this repo's own "prompt is part of the product" rule.
wiki/ is a submodule and not committable from a worker's worktree --
wiki-bound content is in the PR body instead.
2026-09-10 13:54:50 +07:00
ltms 5d422f85fa Merge #436: make fixed placement honour maxLoad (fleetd #435)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m37s
fixed was the one automatic policy that ignored maxLoad, and it is the
default for an absent placement: key. It now gates on the same shared
atCap predicate weighted and round-robin use, and falls through to the
next candidate rather than refusing.

Verified here: 1555 tests green (main was 1548, +7 new methods), 0
compile errors. Mutation battery, control 113 green:
  - atCap boundary >= -> >        KILLED (15 tests, across all 3 policies)
  - drop cap check in fixed walk  KILLED (3 tests)
  - weightExcluded -> false       KILLED (2 pre-existing tests)

Neither live host is affected today: both run placement: weighted (Mac
fleetd.yaml:210, fleet01 fleetd.yaml:172), and weighted already skipped
at-cap candidates via PlacementPolicyUtil.available. This aligns fixed
with the other policies and with maxLoad's documented contract.
2026-09-10 08:54:18 +02:00
Dai Ha ed2027b202 fleetd #435: make FixedPlacementPolicy honor maxLoad
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m54s
FixedPlacementPolicy (the default placement policy) never consulted
maxLoad, so an at-cap default was chosen anyway on every unqualified
spawn -- the cap was advisory, not enforced, for the one policy every
config uses by default. weighted/round-robin already gated on it via
PlacementPolicyUtil.available().

Extract the "at cap" predicate into PlacementPolicyUtil.atCap(ctx, c)
so all three policies share one definition, and consult it at both of
FixedPlacementPolicy's filter sites (the default fast path and the
candidate walk), mirroring the existing weightExcluded pattern. An
at-cap default now falls through to the next candidate instead of
refusing the spawn -- only when every candidate is unusable does the
policy still throw, naming the cap in the message. Update the class
javadoc (five exceptions -> six) and the reason-priority comments to
match CompositePeerLauncher's explicit-spawn order (quarantine,
cooling off, max load, model-off).
2026-09-10 13:46:06 +07:00
Dai Ha 6b0a99b2b7 fleetd #425 rework round 2: stop routedProfileFor's caller re-entering the throwing branch
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 1m34s
Round 1 closed quarantine/cool-off/model-off routing for acquireWithWorktree by resolving the
profile through routedProfileFor(role) and handing that name back to launcher.spawn(SpawnRequest)
as an EXPLICIT profile. That re-resolution has a cost the lead measured directly: naming a
profile explicitly makes CompositePeerLauncher.spawn take its THROWING branch (enforceMaxLoad
included), while the routing branch a blank spawn takes never calls enforceMaxLoad at all, and
FixedPlacementPolicy (the default) deliberately never evaluates maxLoad during automatic
selection. So an at-cap pool-first profile that placement itself would have picked for a plain
unqualified spawn could die at enforceMaxLoad one call later, purely because the worktree path's
route to the spawn passed through an explicit profile name — a new failure a worktree-less
unqualified spawn never hits.

This closes the two-path shape instead of moving it: PeerLauncher gains place(role), returning an
opaque PlacementDecision, and spawn(req, decision), which honors that decision through the SAME
routing branch a blank spawn uses — no enforce* check is newly applied. SessionManager.
acquireWithWorktree now keeps the PlacementDecision from place() and hands it to
spawn(req, decision) for an unqualified request, instead of re-resolving through an explicit
profile name. An explicitly-named profile is unaffected: it still goes through spawn(req) and its
throwing branch, exactly as before.

Also corrects the acquireWithWorktree comment's false claim that round 1 "loses nothing else" —
maxLoad was lost too, as a new hard failure, not a retry. The comment now names it explicitly.

Kept the four round-1 tests (still pass — routedProfileFor now just delegates to place()). Added
one class asserting the invariant itself: an unqualified spawn on a maxLoad-capped profile must
land the same outcome with and without a worktree, asserting on the pair rather than a hardcoded
direction, so it stays correct however fleetd #435 (not this ticket) resolves whether maxLoad
should gate an unqualified spawn at all.
2026-09-10 13:42:06 +07:00
ltms 6b7caba248 Merge #434: make the model gate's own state observable
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m49s
Verified in the worker's tree at 7fd914d: 1544 tests green (base 1535 + 9),
0 compile errors, unpiped mvn clean install. merge-tree against f8b0d42 reports
no conflicts and the two sides share no files.

Read all four production diffs. The design is right: one ModelGateState record
carrying both "armed" and "off" from a single models0() read, so the startup
log line, fleet_profiles' modelGateArmed and the spawn gate cannot disagree —
the fleetd #404 lesson applied properly. disabledModels() now delegates to it
rather than being a second independent read.

Mutated the two subtlest lines, with a control in the same script and the
changed line echoed back with its number:

- modelGateState(): sentinel identity check replaced by the naive
  "m.offIds().isEmpty()" -> 2 failures. FleetProfilesModelGateStateTest
  .modelsBlockWithNothingOffReportsGateArmedAndZeroOff:86 and
  CompositePeerLauncherTest.modelGateStateIsHotReloadedThroughARealConfigRef
  :1519, both "expected: <true> but was: <false>". So the one state this
  ticket exists to expose — a block present with nothing off — is pinned.
- notConfigured() returning a non-empty off set, breaking the invariant the
  record's javadoc states but does not enforce -> 3 failures, including
  disabledModelsIsEmptyWithNoModelsConfigured:1581. So the invariant is
  observable even though the constructor does not check it.

Control run unmutated: 1544 green.

Follow-up on me, not a merge blocker: fleet_profiles gains an operator-visible
field, so this needs a wiki/11-Features.md entry. Workers cannot commit the
wiki submodule, so I am adding it.
2026-09-10 08:32:54 +02:00
Dai Ha f8b0d42a5c fleetd #431 follow-up: profileForSlot's javadoc named a caller that does not exist
CI / contract (push) Successful in 1m28s
CI / build (push) Failing after 1m53s
The javadoc said "what the spawn lifecycle reads". Nothing in src/main calls
profileForSlot at all, in either the ".profileForSlot(" or the
"::profileForSlot" form. My own #431 ticket text repeated that sentence as a
fact and ranked the three accessors by it, and the #432 worker copied it into
the test file's comment and one assertion message. Corrected in all three
places; the ticket correction is posted on #431.

What the spawn lifecycle actually reads for an architect's profile is the
SlotReservation that reserve() returns — SessionManager.java:225,
"reservation == null ? profile : reservation.profile()".

The corrected ranking, measured rather than read off the javadoc:
- nameForSlot is wired, at CallerResolver.java:137 (method reference, which is
  why a ".nameForSlot(" grep missed it)
- isSlot is reached through bind, called at MemberRegistry.java:235 and :376
- profileForSlot has no caller at all

Prose only. No behaviour change.
2026-09-10 13:25:19 +07:00
ltms d1e7d71eee Merge #432: pin profileForSlot, isSlot and nameForSlot against a live reload
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
Verified in my own tree at a196d34: 1539 tests green, 0 compile errors, 0 files
changed under src/main (test-only, as reported).

Mutated two halves the worker's own proof did not cover, with a control in the
same script and the changed line echoed back each time:

- roleForSlot returning ARCHITECT for any configured slot (flatten is
  role-blind) -> 2 failures, incl. CallerResolverTest
  .aBoundNonArchitectSlotStillResolvesAsAWorker:454. So the role-blind flatten
  cannot grant ARCHITECT through a non-architect pool; that was already pinned.
- nameForSlot parsing the key suffix instead of reading config -> 1 failure,
  MemberRegistryLiveTest.nameForSlotReflectsANameChangedByReload:255. The new
  test has teeth beyond the freeze the worker ran.

Control run unmutated: green.

The ticket's severity ranking was wrong and I corrected it on #431. profileForSlot
has no caller in src/main in either the "." or "::" form, so its javadoc ("what
the spawn lifecycle reads") names a caller that does not exist; nameForSlot is
wired at CallerResolver.java:137; isSlot is reached through bind, at :235 and
:376. A follow-up commit fixes the test prose that repeated my claim.
2026-09-10 08:24:26 +02:00
Dai Ha 7fd914df1a fleetd #422 follow-up: make the model gate's own state observable
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m53s
PeerLauncher.disabledModels() reported an empty set both when no
models: block exists and when a block exists with nothing off, so
fleet_profiles/GET /profiles and the startup log could not tell an
inert gate from an armed one reporting zero. Add
PeerLauncher.ModelGateState (configured + off), a modelGateState()
default method disabledModels() now delegates to, and a
CompositePeerLauncher override that reads models0() once and
distinguishes the null-supplier case (no models: block) from a real,
config-supplied block via identity against the NO_MODELS_CONFIGURED
sentinel — reusing the exact accessor the spawn gate itself reads, per
the fleetd #404 lesson.

Wires the state into a new startup log line (Fleetd.modelGateCoverageLine)
and a new modelGateArmed field in FleetMcp.profilesView, reported
unconditionally alongside the existing modelsOff set.
2026-09-10 13:18:52 +07:00
Dai Ha b066eb1903 fleetd #425 rework: resolve acquireWithWorktree through real placement, not a blind pool-first read
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m13s
e1d7dde (PR #430) kept two good fixes and one regressed one. Kept: (1)
CompositePeerLauncher.defaultProfile() delegating to
defaultProfileFor(MemberRole.DEV) so fleet_profiles' "default" tracks a live
reload, and (2) PeerLauncher.defaultProfileFor(MemberRole). Redone:
acquireWithWorktree's profile pre-resolution.

The regression: acquireWithWorktree pre-resolved via
launcher.defaultProfileFor(memberRole), which just returns the role pool's
FIRST entry, blind to quarantine/cool-off/model-off. That name was then
passed to launcher.spawn as an EXPLICIT profile, which takes
CompositePeerLauncher.spawn's THROWING branch (enforceNotQuarantined /
enforceMaxLoad / enforceModelEnabled) instead of the ROUTING branch a blank
profile gets. So a quarantined or model-off pool-first profile turned a
routine unqualified spawn into a hard PlacementException -- undermining
fleetd #429's "the fleet keeps working when a model is turned off"
guarantee for every worktree spawn.

Fix: add PeerLauncher.routedProfileFor(MemberRole), the profile an
unqualified spawn of that role would actually be routed to right now --
same candidate list, same quarantined/coolingOff/modelOff filtering, same
PlacementPolicy spawn() itself consults. CompositePeerLauncher implements it
by extracting spawn()'s context-building into a shared private
placementContextFor(role, unreachable), so spawn() and routedProfileFor()
can never disagree about which conditions apply to which candidate.
acquireWithWorktree now calls routedProfileFor once and reuses that name for
repoRoot, parityOverlay, and the spawn -- the fleetd #425 defect (the three
disagreeing) stays fixed, now on the routed path instead of the blind one.

An explicit profile named by the caller is untouched -- it still hits the
throwing branch, which is correct for an operator override.

Trade-off carried over from e1d7dde, now precisely scoped: an unqualified
worktree spawn still loses CompositePeerLauncher's cross-candidate retry on
a live PeerUnreachableException (a transport failure at spawn time, which
placement cannot see in advance) -- but NOT the quarantine/cool-off/
model-off routing, which routedProfileFor already resolved before spawn
ever runs. Accepted: a worktree provisioned for the wrong backend is worse
than a spawn that fails cleanly and can be retried by the caller.

Tests: CompositePeerLauncherTest gains
routedProfileForSkipsAQuarantinedPoolFirstProfileUnderFixedPolicy and
...ModelOff..., both under PlacementPolicies.fixed() (the default policy,
not weighted() -- the previous round's tests all used weighted() and never
exercised FixedPlacementPolicy's own inline filter, which is exactly what
regressed). SessionManagerTest gains
acquireWithWorktreeRoutesAroundAQuarantinedPoolFirstProfile, proving
repoRoot/parityOverlay/spawn agree on the ROUTED profile, not just the
pool-reordered-by-reload one the existing #425 tests already covered.
2026-09-10 13:10:30 +07:00
Dai Ha a196d34455 fleetd #431: pin profileForSlot, isSlot and nameForSlot against a live reload
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m7s
#424 made MemberRegistry.slots() re-read fleet: on every call, but only
roleForSlot was tested against a real reload. profileForSlot, isSlot and
nameForSlot all have the same live-read line and none was pinned — proved by
freezing each to a construction-time snapshot and watching the full suite
stay green.

Adds 4 tests to MemberRegistryLiveTest, each driving a real ConfigRef.reload()
against a @TempDir config file (never two frozen registries compared in
memory, which would test the constructor instead of the reload):
- profileForSlotReflectsAProfileChangedByReload
- isSlotStopsReportingASlotRemovedByReload / isSlotStartsReportingASlotAddedByReload
- nameForSlotReflectsANameChangedByReload

No production change. profileForSlot has no call site anywhere in src/main
yet, so there is no spawn-lifecycle seam to drive the test through beyond the
accessor itself.
2026-09-10 13:06:56 +07:00
Dai Ha 051d320ea0 fleetd #425: fleet_profiles' default and worktree provisioning must read live placement
fleet_profiles' "default" was CompositePeerLauncher.defaultProfile, a value
frozen at construction from cfg.effectiveDefaultProfile(). An unqualified
fleet_spawn instead resolves the dev pool live via defaultProfileFor(DEV) on
every call, so reordering fleet.developers and reloading changed where a
spawn landed without ever changing what fleet_profiles reported.

- CompositePeerLauncher.defaultProfile() now delegates to
  defaultProfileFor(MemberRole.DEV) -- the same live, reload-aware pool read
  placement already uses -- falling back to the frozen field only when no
  profiles are configured at all.
- PeerLauncher gains a default defaultProfileFor(MemberRole) method so a
  generic PeerLauncher reference can ask for a role's live default; the
  default implementation delegates to defaultProfile() for launchers with no
  pool concept of their own.
- SessionManager.acquireWithWorktree resolved a profile via
  launcher.defaultProfile() (DEV-only) to provision repoRoot/parityOverlay,
  then spawned with the original (possibly blank) profile, which re-resolves
  independently through placement -- for any non-DEV role, or across a config
  reload between the two reads, the two resolutions could disagree and
  provision a worktree for a profile the member never runs on. Fixed by
  resolving once, through defaultProfileFor(the caller's actual role), and
  reusing that same resolved name for repoRoot, parityOverlay, and the spawn
  itself. Trade-off: this path now spawns with an explicit profile rather
  than a blank one, so it loses CompositePeerLauncher's cross-candidate retry
  on PeerUnreachableException -- accepted because a worktree provisioned for
  the wrong backend is worse than a spawn that fails cleanly and can be
  retried.

Tests: CompositePeerLauncherTest (live dev-pool reorder + empty-pool
fallback), FleetProfilesLiveDefaultTest (drives FleetMcp.profilesView
directly), SessionManagerTest (worktree overlay follows a reorder, and a
non-DEV role's worktree spawn uses that role's pool, not DEV's).
2026-09-10 12:58:20 +07:00
Dai Ha 766772763f Merge #428: revoke the ARCHITECT privilege on reload, not just future spawns
CI / contract (push) Successful in 1m27s
CI / build (push) Successful in 1m32s
fleetd #424. MemberRegistry.slots() now re-reads fleet: through a supplier,
so removing an architect slot demotes the bound pane on its very next request.
The boundEntries cache and entryFor() fallback from the first round are gone:
the slot OCCUPANCY (terminalToSlot) survives a reload, the ARCHITECT role does
not. That split was my own ticket wording's fault -- I asked for a test that a
bound architect "survives the rebuild", which conflated the binding with the
privilege.

Conflict resolved by hand in ConfigRef.java: #422 (models:) and #424
(architects) both rewrote the same Hot bullet. Kept both.

Also corrected two claims #424's own second commit left stale -- 7f672f0
reversed the behaviour but never touched ConfigRef, whose whole job is to tell
the operator what a reload does:
  - the Hot bullet said MemberRegistry's rule "keeps a live session's identity
    even after its slot is removed from config"
  - the reload-report comment said "only a NEW bind is refused"
Both now say what the code does: removal revokes ARCHITECT on the next
request, and only the slot occupancy survives.

Verified by the lead: 1535 tests, 0 failures, 0 compile errors, BUILD SUCCESS
on the merged tree.

Mutation of three halves the worker's own proof did not cover -- profileForSlot,
nameForSlot and isSlot each pointed at a frozen snapshot taken at construction
(live readers 5 -> 4, each mutation naming its method and line). All three
PASSED at 1535. The ticket's own fix is well pinned; these three sibling live
reads are not. Follow-up filed.
2026-09-10 12:54:35 +07:00
ltms eab8185d7b Merge #429: enforce the models.allow on/off gate at spawn, hot — including under fixed placement
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m33s
fleetd #422 + follow-up. Verified by the lead: 1524 tests green, 0 compile errors.

Mutation proof of three halves the worker's own proof did not cover:
- modelOffProfiles() -> Set.of() (the feed into ctx.modelOff): kills 3, incl. fixedPlacementSkipsAnOffModelProfileToo
- enforceModelEnabled() -> no-op (the explicit-spawn gate): kills 3, incl. the real-ConfigRef hot-reload test
- disabledModels() -> Set.of() (the reporting accessor): kills 1, alone
2026-09-10 07:44:35 +02:00
Dai Ha 7f672f0fb8 fleetd #424: revoke the ARCHITECT privilege on reload, not just future spawns
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 2m3s
Correction to the #424 fix in PR #428: the ticket asked to revoke a removed
architect slot, but the previous change (boundEntries) kept BOTH the binding
and the ARCHITECT privilege alive for an already-bound session after its slot
left config. That left the ticket's actual headline defect half-open.

The corrected rule: config governs both what may be bound next AND what a
bound slot still grants. Removing a slot now demotes its bound session to
worker on the very next request (roleForSlot/nameForSlot read slots() with no
cache, so CallerResolver.resolve falls through to Principal.worker(...)). The
terminalToSlot binding itself is untouched by a reload, on purpose: dropping
it would double-book the slot key and break unbind's compare-safe contract.

- Delete boundEntries and entryFor; profileForSlot/roleForSlot/nameForSlot/
  isSlot all read slots() directly, live, with no cache.
- Rewrite the class doc's binding rule for the corrected semantic.
- Replace the old "survives removal" test with anArchitectAlreadyBoundToASlotIsDemotedByReload,
  asserted through a real CallerResolver.resolve (not the roleForSlot seam),
  plus two tests for what must NOT change: the binding still occupies the
  slot after removal (a second terminal cannot claim it, even once the slot
  returns to config), and unbind still succeeds for the original terminal.

Verified snapshot()/CallerResolver.members() need no change: snapshot() only
ever reported raw terminalToSlot occupancy, and CallerResolver.members() has
no production caller.
2026-09-10 12:29:35 +07:00
Dai Ha e2801b9bbc fleetd #422 follow-up: gate FixedPlacementPolicy on modelOff too
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
FixedPlacementPolicy is the DEFAULT placement policy (PlacementPolicies.fromName
returns it for an absent/blank name) and it built its own inline candidate
filter instead of calling PlacementPolicyUtil.available(). That filter checked
quarantined/coolingOff/unreachable/excluded() but never modelOff(), so an
unqualified fleet_spawn on any fleet without an explicit placement: policy
could still land on a profile whose model the operator turned off.

- Add ctx.modelOff() to both filter sites: the default-profile fast path and
  the fallback walk over ctx.candidates().
- Add a modelOff refusal reason to the default-profile reasons list, worded as
  an operator decision ("turned off in models.allow"), matching
  enforceModelEnabled. Quarantine and cooling off still take priority when a
  profile is also model-off, matching CompositePeerLauncher's explicit-spawn
  check order.
- Update the two stale "excluded from automatic selection" messages to name
  model-off, consistent with PlacementPolicyUtil.emptyException.
- Update the class javadoc: four exceptions -> five, with a new bullet for
  model-off (fleetd #422).

Tests: PlacementPolicyTest gains fixedSkipsModelOffDefault (fast-path),
fixedFallbackWalkSkipsModelOffCandidate (fallback walk),
fixedThrowsWhenDefaultAndEveryCandidateModelOff (all-off refusal wording), and
fixedReportsQuarantineNotModelOffWhenBothApply (priority). CompositePeerLauncherTest
gains fixedPlacementSkipsAnOffModelProfileToo, an integration-level mirror of
the existing placementSkipsAnOffModelProfileAndRoutesToAnotherOne but under
PlacementPolicies.fixed(). The two existing weighted()-based tests are
untouched.
2026-09-10 12:25:31 +07:00
Dai Ha ea02c7b248 fleetd #422: enforce the model allow-list on/off state at spawn, hot
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 1m30s
Ships the two halves left out of the earlier allow-list ticket in one PR,
since apart they are inert: a gate with no flag always allows, and a flag
nothing reads does nothing.

- FleetConfig.Models.ModelEntry gains `enabled` (default on; absent/true =
  on, false = off). Turning a model off never removes it from `allow:` —
  validateModels() checks membership only, so an off model stays valid
  config and a still-configured profile naming it does not refuse reload.
  Models.offIds() is the one live accessor both the gate and the status
  report read.
- CompositePeerLauncher.enforceModelEnabled is a FOURTH, independent
  spawn-refusal reason (operator intent) — never layered onto
  BackendQuarantine/BackendOutagePolicy, which are backend-reported outage.
  Wired into the explicit-profile branch. modelOffProfiles() feeds the same
  off-model exclusion into PlacementContext for unqualified spawns via
  PlacementPolicyUtil (a new modelOff set, counted into its own bucket in
  emptyException so "all off" is named as the cause, not generic).
  Both read models0(), a live Supplier<FleetConfig.Models>, so a reload
  reaches the very next spawn — no restart.
- PeerLauncher.disabledModels() (default empty) lets fleet_profiles/
  GET /profiles report off models by reading the exact same accessor the
  gate reads (the fleetd #404 lesson: a status field must read the source
  the behaviour reads).
- ConfigRef: `models` reclassified from deferred to hot-excluded — nothing
  about it is baked into a startup-built object anymore; membership is
  re-validated in full on every reload via validateAll(), and the on/off
  half is read live everywhere. Tally: 5 cold, 13 deferred, 3 split, 4
  hot-excluded (25 total). ConfigRefTopLevelCoverageTest and
  ConfigRefTopLevelReportingCoverageTest updated with no new exclusion
  added just to force green.

Tests: FleetConfigTest (old-style fixture stays on; an off model is still
valid config; one model disables every profile naming it),
CompositePeerLauncherTest (explicit refusal wording distinct from
quarantine/cool-off; unqualified spawn skips an off candidate and names
model-off when every candidate is off; a real ConfigRef.reload() proves
the hot path; disabledModels() matches the gate).
2026-09-10 12:08:18 +07:00
Dai Ha ce74e164c6 fleetd #424: make architect-slot identity checks read fleet.architects live
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m32s
MemberRegistry used to flatten fleet.architects into an unmodifiable map at
construction, so removing (revoking) an architect slot from config never
took effect: reserve()/requireSlotFor() kept granting spawns against the
frozen snapshot forever, while ConfigRef told the operator "already
applied" for the wrong consumer.

- MemberRegistry gains a live constructor (MemberRegistry.live(Supplier))
  that re-flattens fleet.architects/developers/reviewers on every
  slots()/slotsFor() call, so reserve() and requireSlotFor() (which both
  read through slotsFor) govern the NEXT spawn with no restart. The frozen
  single-arg constructor is kept for tests and fixed/code-built configs.
- Binding rule: config governs what may be bound next; it never
  retroactively unbinds a live session. A slot removed from config while a
  terminal is bound to it keeps that binding. To keep the bound terminal's
  IDENTITY too (CallerResolver.resolve reads roleForSlot/nameForSlot on
  every request), every successful bind now caches the slot's Entry into a
  new boundEntries map; roleForSlot/nameForSlot/profileForSlot/isSlot fall
  back to it when the slot is no longer live, and unbind clears it in the
  same critical section it clears the binding.
- Fleetd.java now wires MemberRegistry.live(() -> config.get().fleet())
  instead of the frozen constructor.
- ConfigRef: corrected the fleet.leaders split-key message and the class
  doc's Hot bullet — architects is now hot for two independent consumers
  (CompositePeerLauncher for placement, MemberRegistry for identity), not
  only the one the message used to name. An architects-only edit still
  reports nothing beyond "config reloaded", which is now honest since the
  key really is fully hot for both consumers.

Added MemberRegistryLiveTest: real ConfigRef.reload() against a @TempDir
file, both directions (slot removed / slot added) for requireSlotFor and
reserve tested separately, plus a bound-architect-survives-removal test
that checks the binding AND the identity (roleForSlot/nameForSlot).

Mutation-tested: reverting requireSlotFor to a frozen snapshot fails
requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload and its mirror;
reverting reserve the same way fails the two reserve tests; dropping the
boundEntries fallback fails the survives-removal test's roleForSlot
assertion. All three restored before commit.

mvn clean install: Tests run: 1510, Failures: 0, Errors: 0, Skipped: 0 —
BUILD SUCCESS.
2026-09-10 12:05:18 +07:00
ltms e60f892efd Merge pull request 'fleetd #415: split coverage() feature-state wording by pattern fallback semantics' (#423) from worker/415-coverage-wording-2cbf9c-5 into main
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m37s
2026-09-10 06:57:33 +02:00
Dai Ha ce05886831 fleetd #415: pin which UnsetMeaning Fleetd pairs with which pattern key
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m54s
Review found a gap: the earlier tests all called CompletionResolver.coverage()
directly, supplying the UnsetMeaning themselves — proving the enum's wording,
never that Fleetd's two call sites pair the right meaning with the right key.
Swapping the two UnsetMeaning arguments at those call sites (recreating #415's
defect with exhaustedPattern and errorPattern exchanged) compiled with 0 errors
and left all 1506 tests green.

Extract the two coverage-line call sites out of main() into package-private
static factories (Fleetd.exhaustedPatternCoverageLine /
errorPatternCoverageLine), the same pattern already used for capacitySource
and worktreeBranchLookup. Add FleetdPatternCoverageLineTest, which calls both
factories directly and asserts the actual wording each produces for the same
empty-coverage input, including that the two differ.

Also recorded the swap-mutation measurement (0 errors, 1506 green) in
UnsetMeaning's javadoc so a future reader does not delete the new test as
redundant with CompletionResolverTest.
2026-09-10 11:54:02 +07:00
Dai Ha be123d0ac7 fleetd #415: split coverage() feature-state wording by pattern-key fallback semantics
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m4s
CompletionResolver.coverage() measured pattern coverage (how many profiles set
a key) but its 'off' wording read as feature state. That is false for
errorPattern: an unset errorPattern still runs the classification against the
built-in BACKEND_ERROR pattern (CompletionResolver.java:84), so the empty case
is not off.

Add CompletionResolver.UnsetMeaning (OFF / BUILT_IN_DEFAULT), a required
parameter every coverage() call must supply — no defaulted overload, so a
future third pattern key cannot compile without stating what unset means for
it. Fleetd.java now passes UnsetMeaning.OFF for exhaustedPattern (no fallback
exists) and UnsetMeaning.BUILT_IN_DEFAULT for errorPattern.

Tests: updated the three existing empty/full/partial cases to pass the new
parameter, corrected the one test that pinned the old (wrong) errorPattern
wording, and added a test that asserts the same empty-coverage input produces
different wording for the two keys.
2026-09-10 11:40:32 +07:00
ltms 2d09c8b027 Merge pull request 'fleetd #418: barrier the throw-path push-loop test on state decide() reads' (#419) from worker/418-588283-3 into main
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m39s
2026-09-10 06:29:23 +02:00
ltms ed54f0224e Merge pull request 'fleetd #416: fleet_list must enumerate the STARTUP profile set' (#420) from worker/416-3ad1da-1 into main
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m34s
2026-09-10 06:27:36 +02:00
Dai Ha 8d5bc3ee89 fleetd #416: fleet_list must enumerate the STARTUP profile set
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m52s
`profiles` is a DEFERRED key: HerdrPeerLauncher takes Map.copyOf(profiles) once at
construction, so a profile added only to the hot-reloaded map can never be spawned.
CapacitySource was built with `() -> config.get().profiles().keySet()` — the live map —
so fleet_list reported a hot-added profile as free while fleet_spawn on that same
profile failed with "unknown worker profile". fleet_list's contract for `free` says it
runs "the same check the spawn gate itself runs"; it did not.

The set now comes from the startup snapshot, the same shape as the coordinator.peers
wiring three lines below, whose comment already stated the rule. maxLoad stays live on
purpose — ConfigRef documents it as hot, like credentialId and weight — so a maxLoad
edit still takes effect without a restart.

Found by the fleet01 lead, who proved it with a live ghost profile rather than an
argument. Implemented by a sonnet member; its reply was lost to an empty scrape, so the
work was salvaged uncommitted from its worktree and both proof steps were run by the
lead instead:

  full build                          Tests run: 1505, Failures: 0, 0 compile errors
  mutation A: live keySet restored    liveOnlyProfileIsNotListed FAILS (alone)
  mutation B: permanently empty set   startupProfileIsListed FAILS (alone)

Mutation B is the point of the second direction: per #404, a test that only ever checks
the absent case cannot tell a correct lookup from one that returns nothing at all.
2026-09-10 11:25:56 +07:00
Dai Ha ab0cc71aa4 fleetd #418: barrier the throw-path push-loop test on state decide() reads
CI / contract (pull_request) Successful in 1m27s
CI / build (pull_request) Successful in 1m50s
anAskThatLeavesByThrowingStillClosesItsQuestion barriered on Phase.ASKING,
which markAsyncQuestion sets in ask()'s FIRST step. The assertion right
after it depends on ask()'s THIRD step (pushLoop.onQuestionOpened), which
is what actually populates ReplyPushLoop's pendingQuestions map. Under
load the asker thread can be descheduled between those two steps, so the
barrier released before decide() had anything to see, and it correctly
returned STOP instead of the expected INJECT.

Add ReplyPushLoop#pendingQuestionTurnIdsForTest, a package-private test
seam (modeled on MessageService#isCompletionStampedForTest) exposing the
private pendingQuestionTurnIdsFor. The test now waits for its own turnId
to appear there before asserting on decide() — not for decide() itself to
return INJECT, which would make the barrier assert nothing.

Checked every other awaitTicketPhaseOn(..., Phase.ASKING) in the file
(two, in the CB-582 nudge tests): both are followed by a real awaitNudge()
that waits for an actual agent.prompt push-loop call before any assertion
depends on push-loop state, so they are not exposed to this race.

No production behaviour changed.
2026-09-10 11:24:41 +07:00
ltms 4cd9046353 Merge pull request 'fleetd #409: deterministic test for the #399 completion-stamp ordering race' (#414) from worker/deterministic-stamp-race-409-3cb7b6-10 into main
CI / contract (push) Successful in 59s
CI / build (push) Successful in 1m36s
2026-09-10 04:48:44 +02:00
Dai Ha 4e98a74047 fleetd #409: deterministic test for the #399 completion-stamp ordering race
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 2m5s
Widens the completed-hook's real race window (normally instructions-wide, needing
~2x-core host load to hit by chance per #399) by injecting a bounded sleep into the
test clock's completion-stamp read. This makes the ordering invariant — a test must
wait for isCompletionStampedForTest, not just DONE, before advancing the clock past
the TTL — fail deterministically on the first run when the barrier is removed, and
pass deterministically with it present. No production code changed.
2026-09-10 09:39:08 +07:00
ltms a502ba53e0 Merge #404: the armed field reads the startup pattern map, both directions pinned
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m0s
The armed lookup now uses the same compiled map detection uses. Second test added during review after a mutation proved the true direction was unpinned.
2026-09-10 04:20:54 +02:00
Dai Ha 446cc11d1d fleetd #404: pin the armed field's TRUE direction too
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m51s
Mutation-tested during review: replacing the armed lambda with
`profile -> false` left the whole suite green at 1475 tests, because the
only existing test passes an EMPTY startup map. That mutation would make
#395's visibility feature silently dead.

With this test the same mutation fails, and it is the only test that
fails, so nothing else covers this direction.
2026-09-10 09:20:43 +07:00
ltms ddd81fe174 Merge #391: fleet_reply refuses a lead, and the role argument is now mandatory
CI / contract (push) Successful in 46s
CI / build (push) Successful in 2m1s
The permissive 3-arg overload is deleted, so a handler that forgets the role no longer compiles.
2026-09-10 04:18:36 +02:00
ltms ae74cc081f Merge #399: wait on a real completion stamp, not on DONE
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m29s
Adds a volatile completedNanos stamp and a bounded test barrier. Mutation-tested: removing the production stamp fails 3 tests.
2026-09-10 04:14:29 +02:00
Dai Ha cf8da1d5fa fleetd #391: require reply caller role
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 1m45s
2026-09-10 09:13:16 +07:00
Dai Ha b8aedeafcb fleetd #404: test armed startup map behavior
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m30s
2026-09-10 09:13:07 +07:00
Dai Ha 3d61af6f6f Merge #400: the scrub receipt now measures the blank, not the attempt
CI / contract (push) Successful in 1m8s
CI / build (push) Successful in 1m34s
The blanking loop classified a name as blanked from eval's exit status.
zsh coerces a bare NAME= assignment on an integer special parameter
(SECONDS, RANDOM, SHLVL, HISTSIZE, COLUMNS, LINES, USERNAME) to a number
instead of failing, so eval returned 0 with the value untouched. Measured
7 false receipts in 10 names. This is a security receipt, so a count that
overstates the scrub is worse than no count.

The loop now runs the eval unconditionally and decides from the observed
value, read back with the (P) indirection flag. One check covers all
three shapes a name can take: a real blank, a fatal error eval merely
contained, and this silent no-op. Exit status plays no part.

Lead review: mutation removing the '!' unblankable report line is CAUGHT
(2 failures in EnvAllowListScrubTest, both asserting the name is reported
rather than silently dropped). The eval-site identifier guard is
untouched.

Unplanted evidence the fix works: the #394 test
unblankableNameInTheMiddleDoesNotAbortNamesAfterIt began failing under
the fix, because zsh auto-exports SHLVL and the old exit-status bug had
been miscounting it as blanked all along. Its exact-count assertion had
only ever passed because of the bug beside it; it is now a presence
check, since which names a zsh version auto-exports is not this test's
to pin.

Still not pinned, tracked in #394's follow-up: the eval-site identifier
guard has no test behind it.
2026-09-10 09:05:30 +07:00
Dai Ha ed99c209ac Merge #398: a central allow-list of usable models, with the startup validators pinned
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m25s
models: allow: is a single place that names every model the fleet may
use. Absent or empty keeps today's behaviour, so this ships inert until
configured. Once set, a profile naming a model outside the list refuses
to start, and refuses a reload, rather than reaching a backend adapter
as a free-form string.

The allow-list cannot be checked against a provider catalogue: for an
opencode profile fleetd SYNTHESIZES the provider from provider/model
plus baseUrl (OpenCodeLauncher:562-583), so a valid fleetd model id
appears in no published catalogue. An operator-owned list is therefore
the only workable gate.

Includes the #398 follow-up: FleetConfig.validateAll() reflectively
sweeps this class's validateXxx() methods, and both real call sites
(Fleetd.main and ConfigRef.reload) call that one method. Before this,
deleting a validateXxx() call from either caller left the whole suite
green. FleetdStartupValidationTest now drives the real Fleetd.main.

Lead review: mutation on the reload call site is caught (ConfigRefTest,
2 failures). Mutation replacing the reflective sweep with a hardcoded
list is NOT caught (1491 green) — so the sweep is a convenience and the
denominator test is the real guarantee; two false statements in the test
javadoc were corrected to say so (af4c88d).

Recovered work: the worker's agent died mid-turn with the follow-up
uncommitted and the startup call left disabled as
'// MUTATION-TEST-TEMP: cfg.validateAll();'. I restored it before
committing.

NOT covered: the five log-only reportXxx(cfg) calls in main are still
unpinned — filed as #407.
2026-09-10 09:02:18 +07:00
Dai Ha af4c88d54b t398: correct two false statements in FleetConfigValidateAllTest
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m58s
Measured at review: reverting validateAll() to a hardcoded list of
today's six calls leaves the suite green (1491 tests, 0 failures). The
class javadoc claimed that mutation fails a test. It does not — claim 1
pins the generic helper on an unrelated class, claim 2 pins today's six,
and a hardcoded list satisfies both.

The interaction was the real hazard. The denominator assertion IS a
tripwire (declaring a seventh validator fails it), but its failure
message said the sweep reaches new validators 'by construction' and told
the author to just update the expected set. If the sweep were ever
replaced by a name list, the one assertion that fires would hand back a
false all-clear at the moment it fired.

Javadoc now states the measurement, and names the denominator test as
the actual guarantee. The assertion message now says to confirm
validateAll() still delegates to invokeAllValidators(this) BEFORE
updating the expected set.
2026-09-10 09:01:49 +07:00
Dai Ha 44c735f6f5 fleetd #404: report armed detection from startup map
CI / build (pull_request) Successful in 1m25s
CI / contract (pull_request) Successful in 1m20s
2026-09-10 08:56:39 +07:00
Dai Ha b540a1744b fleetd #398 follow-up: pin the startup validators with a reflective validateAll()
CI / contract (pull_request) Successful in 1m9s
CI / build (pull_request) Successful in 1m28s
Mutation testing found that deleting a cfg.validateXxx() call from
Fleetd.main left the full suite green: every test called a validator
directly and none exercised main as the caller.

FleetConfig.validateAll() sweeps this class's own public no-arg void
validateXxx() methods by reflection and invokes each in alphabetical
order, so a newly written validator is wired into both callers
(Fleetd.main and ConfigRef.reload) with no second step to forget.
FleetdStartupValidationTest calls the real Fleetd.main with six configs,
each failing exactly one validator.

Recovered by the lead: the worker's agent died mid-turn with this work
uncommitted, and had left the startup call commented out as
'// MUTATION-TEST-TEMP: cfg.validateAll();' from its own mutation run.
I restored the call before committing. Build after restoring:
Tests run: 1491, Failures: 0, BUILD SUCCESS.

NOT covered, and not claimed to be: the five log-only reporters in
main (reportRequiredSecrets, reportGitHostShape, reportMemberTrustModel,
reportMemberCredentialsGap, and reportExhaustedPatternGap on current
main) are not validateXxx() methods, so the sweep does not reach them
and their call sites stay unpinned.
2026-09-10 08:56:20 +07:00
Dai Ha 0f51d53098 fleetd #399: fix TTL test race by waiting on a real completion stamp, not DONE
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m2s
poll() can report Phase.DONE for a ticket before the whenComplete hook that
stamps Task.completedNanos has run — CompletableFuture.complete() publishes
its result and only then runs dependents. The TTL tests advanced an injected
clock right after observing DONE, so on a host where the hook runs late it
stamps the ADVANCED time and the eviction never happens (fails on Linux,
passes on macOS).

Add a package-private test seam, MessageService.isCompletionStampedForTest,
that reports whether completedNanos is stamped. Both TTL tests now wait on
that (a real volatile read/write happens-before edge) before advancing the
clock, instead of on Phase.DONE. The prune condition in pruneTerminalTickets
is untouched.
2026-09-10 08:54:38 +07:00
Dai Ha c670792ffe #400: classify the blanking loop's result on the observed value, not eval's exit status
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m1s
eval "export NAME=" can return success even when zsh coerces the bare
assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
HISTSIZE, COLUMNS, LINES, USERNAME) instead of failing, leaving the
value unchanged. The old exit-status check then reported the name as
blanked when it was not -- a false receipt.

Classify on the observed effect instead: attempt the export, then read
the name's value back with the (P) indirection flag and decide from
whether it is now empty. One check now covers all three shapes a name
can take here -- a genuine blank, a fatal read-only error eval merely
contains, and this silent no-op -- with the exit status playing no
part in the decision.

Adds a test driving all three shapes through the real scrubScript in
one run (a normal name, LINENO for the fatal case, SECONDS for the
silent no-op), with the parent environment explicitly carrying those
names since a cleared ProcessBuilder parent does not expose them on
its own. Also corrects the previously-merged
unblankableNameInTheMiddleDoesNotAbortNamesAfterIt test, whose "exactly
one failed name" assertion turned out to only pass by accident: zsh
itself auto-exports SHLVL on every shell start, and the old exit-status
bug was silently miscounting it as blanked. The fixed classification
now correctly reports it unblankable too, so the test asserts presence
rather than an exact count.
2026-09-10 08:46:27 +07:00
Dai Ha 3982ace544 fleetd #391: refuse lead fleet replies
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m31s
2026-09-10 08:43:54 +07:00
Dai Ha 7180b1aad0 Merge #395: warn when exhaustedPattern usage-limit detection is off
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m27s
A profile with no exhaustedPattern has usage-limit detection silently
disabled. Startup now reports every unarmed profile (a louder warning for
subscription profiles, which are the ones a limit actually stops), and
fleet_profiles / GET /profiles carry exhaustionDetectionArmed per profile.

Reviewed by the lead: all three requested mutations fail a test, and the
back-compat QuarantineSource ctor defaults to 'not armed' when the source
is unknown. KNOWN GAP, not fixed here: deleting the
reportExhaustedPatternGap(cfg) call at Fleetd.java:140 leaves the suite
green (Tests run: 1472, Failures: 0). That is the same unpinned-startup-
call shape as PR #398's six validators, and fleetd #398's ticket owns it.
2026-09-10 08:43:09 +07:00
Dai Ha c26f695402 fleetd #395: warn when exhaustedPattern detection is silently off
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m25s
exhaustedPattern is opt-in per profile: unset means a usage-limit
refusal on that profile is never classified BACKEND_EXHAUSTED and
never quarantines its credential, with nothing telling the operator.
Add a startup WARN naming every unarmed profile (a louder, separate
WARN for a subscription: true profile, since that is the operator's
own metered plan). Surface the same fact per profile in fleet_profiles
as exhaustionDetectionArmed, so an operator can tell "healthy" from
"can never be caught" without reading fleetd.yaml.
2026-09-10 08:34:33 +07:00
Dai Ha 48877315ca Merge #394: contain the fatal export that aborted the credential scrub
CI / contract (push) Successful in 50s
CI / build (push) Failing after 1m31s
The CB-633 allow-list scrub has been dying mid-loop on every fleet01 member
pane and saying nothing. `export UID=` in zsh is not a failed command — it is
a fatal parameter error that terminates the whole sourced file. The blanking
loop is wrapped in `{ ... } 2>/dev/null`, so the message was swallowed and the
report block after the loop never ran.

Root cause found by the fleet01 lead, with xtrace on a live pane's own ZDOTDIR:

    +scrub.zsh:28> _cb633_n=UID
    +scrub.zsh:28> export 'UID='
    +zsh:1> rc=126        <- file aborted

The severity is the INVERSION, and this is their finding, quoted:

  "env lists inherited names first and the names a startup file exports last.
   So the loop blanks the harmless inherited half and dies immediately before
   the operator's own exports — exactly the credentials the policy exists to
   remove. The selection is inverted, not merely partial."

Measured there: UID is name 42 of 57, and a ~/.zshrc decoy at 58 survived on
8 of 8 spawns. "Partial scrub" reads as "we got most of it"; it got precisely
the wrong half.

Fixed with `eval "export ${n}=" 2>/dev/null` rather than a skip-list of the
known-fatal names (UID EUID GID EGID PPID LINENO). A skip-list has to be
complete forever and this is a security control; eval needs no list. Measured:
plain export dies at UID and every later name keeps its value, while the eval
form completes the loop and blanks all of them. PR #396 proposed the skip-list
and is closed in favour of this; its claim that the abort happens "however the
assignment is wrapped" holds for a direct `if ! export` but not for eval,
which reparses in a nested context.

The report now carries `failed N` and `!`-prefixed unblankable names, and
HerdrPeerLauncher WARNs when any name could not be blanked. The old "no report"
WARN no longer claims the daemon knows what the member saw.

Why the suite stayed green: EnvAllowListScrubTest starts zsh from
pb.environment().clear(), and under a cleared parent UID is not an exported
name at all, so the abort could not reproduce in that harness.

Reviewed by mutation, which found a second gap now also closed: the eval is
only safe because names are filtered to ^[A-Za-z_][A-Za-z0-9_]*$. Replacing
that pattern with .* left the class green, so the line the security property
rests on was unpinned. The guard is now re-asserted at the eval site and pinned
by a test. The reachable vector is a VALUE with an embedded newline, not a
hostile name — measured: zsh strips non-identifier env names outright, while
MULTI=$'keep\njunk.fragment' forges 'junk.fragment' as a candidate name out of
its own value.

Closes #394. Refs #396, #388.
2026-09-10 08:30:15 +07:00
Dai Ha 5c08054533 #394 follow-up: re-assert the identifier guard at the eval call site
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m48s
EnvAllowListScrub's blanking loop splices each name into a string
handed to eval ("export ${n}="). That is only safe because every name
reaching _cb633_blank already passed an identifier check in the
enumeration loop -- 20 lines away, in a different loop. Before eval
was introduced a non-conforming name reaching plain `export "$n="`
was inert either way (the quoting neutralized it); eval removed that
safety net, so the enumeration loop's guard became the ONLY thing
standing between a non-identifier string and code execution in the
member's pane, with nothing at the eval site itself defending that
property.

Re-assert the same [A-Za-z_][A-Za-z0-9_]* check immediately before
the eval call, independent of the enumeration loop's own guard (left
untouched, not moved). A name that fails it is counted unblankable
rather than silently dropped, so a bypass of the upstream guard would
leave real evidence in the report.

New test exploits the "junk from multi-line values" gap the
enumeration loop's own comment already documents: a value with an
embedded newline makes `command env`'s text output split into a
spurious extra "name" line that was never a real variable. Runs the
real generated scrubScript() end-to-end under zsh and asserts the
non-conforming fragment is neither blanked nor counted unblankable.
The fragment used is merely non-conforming (contains a dot) --
never command-shaped.

Mutation-verified both guards. Weakening the enumeration guard alone
DOES break the new test (the fragment then reaches the new eval-site
guard and gets counted unblankable, failing the "not unblankable"
assertion). Removing the new eval-site guard alone, with the
enumeration guard intact, does NOT break it: _cb633_blank has exactly
one producer (the enumeration loop), so nothing can reach the eval
site without already having passed the identical check there. That is
expected given the single-source architecture, and it is exactly why
the eval-site guard is defense-in-depth against a future change that
adds a second path into _cb633_blank or decouples the two loops --
not a currently independently-observable divergence.
2026-09-10 08:24:46 +07:00
Dai Ha b6db9c31f5 charter: a peer lead is answered with fleet_send, not fleet_reply
CI / contract (push) Successful in 46s
CI / build (push) Successful in 2m0s
`fleet_reply` has no route to a peer lead. `AmqpReplyInbox` publishes to
`agent.<target>.inbox`, mandatory, and a lead's own terminal has no such
queue, so the publish is refused. `MessageService.reply()` has no peer
branch at all — `grep -c 'coord\|LeadMailbox'` on it returns 0. The charter
told every lead to use a tool that cannot work, and both leads here hit it.

Three edits to the canonical block, byte-identical with the wiki template
(pushed as 803726a; the in-sync check in this file reports True):

- the intent->tool row now says `fleet_send{coordId}`, or `{sessionId}` for
  a peer on the same host, and says plainly that `fleet_reply` is refused
- the prose says WHY: `fleet_reply` resolves a member's blocked `fleet_send`,
  while a peer's coord-id message is durable and non-blocking, so there is
  nothing for it to resolve
- lead<->lead item 3 gains the data-point rule: N observations are N data
  points only if they differ in the axis you are trusting

Wording for all three drafted by the fleet01 lead, who verified the missing
queue namespace independently in its own tree. The data-point rule has now
caught three separate errors in a day, in both directions: one cause blamed
for N failures, and N agreeing measurements that shared a single instrument.

The refusal message itself is still wrong — it says "queue not declared or
owned", which sends the reader to the broker instead of to this file. That
half stays open on #391.

Tracked as fleetd #391.
2026-09-10 08:22:35 +07:00
Dai Ha e7b33fe3a0 fleetd: central allow-list of usable models (models: + validateModels())
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m34s
Add an optional top-level `models:` block (Models{allow: List<ModelEntry>})
naming the models any profiles: entry may use. Absent/empty allow: keeps
today's behaviour exactly (no check, no warning). When configured,
FleetConfig.validateModels() fails config load (and reload, via ConfigRef)
naming both the model and the profile, if any profile's model: is outside
the list. The check is one-way: editing profiles: alone can never widen
what is permitted, only models.allow: can.

Wired into Fleetd.main() alongside the other validateXxx() calls, and into
ConfigRef.reload()/DEFERRED_KEYS so a bad edit can't slip in through a
reload either. Each ModelEntry is its own record (not a bare string) so a
later unit can add per-model on/off or load-limit state without changing
the YAML shape. One flat string namespace covers both a bare Claude id and
an opencode provider-prefixed id.
2026-09-10 08:09:27 +07:00
Dai Ha e3e403e5c8 #394: contain a fatal export error instead of letting it abort the scrub
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m52s
EnvAllowListScrub's blanking loop used a plain `export "$n="` on every
name not on the allow-list. For a zsh read-only/special parameter (e.g.
UID) that is a FATAL parameter error, and since the loop runs inside the
sourced startup file, the error aborts the whole file: every name still
to come is never blanked, and scrub-report.txt is never written at all
-- silently, because 2>/dev/null on the group swallows it.

Route each blanking attempt through `eval` instead, which contains the
error to that one iteration. The loop always finishes; a name it could
not blank is now counted separately ("failed" on the report's first
line) and listed !-prefixed rather than disappearing. No skip-list of
known-bad names is added -- every enumerated name is still attempted,
so a name nobody has thought of is still tried and, if it fails, still
counted.

HerdrPeerLauncher: log a WARN when a pane's report carries a nonzero
failed count, and reword the "no report at all" WARN so it no longer
claims the daemon knows the member "saw the full host environment" --
a partial vs. a missing scrub are different situations and only the
first is now distinguishable from the report alone.
2026-09-10 08:08:41 +07:00
Dai Ha 799014e99d Merge #388: scrub a pane shell that is neither login nor interactive
CI / contract (push) Successful in 1m8s
CI / build (push) Successful in 1m38s
2026-09-10 07:21:48 +07:00
Dai Ha 69e09b10fa #388: scrub a pane shell that is neither login nor interactive
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m32s
EnvAllowListScrub generated four zsh startup files but only .zshrc and
.zlogin sourced the scrub — .zshenv (the one file zsh always reads) did
not. A pane shell that is neither login nor interactive reads only
.zshenv and stops, so it was never scrubbed at all (measured on fleet01,
issue #388).

Adding an unguarded scrub to .zshenv (the ticket's own suggested fix) is
wrong: .zshenv is read by every zsh, including a short-lived `zsh -c`
a member's own tooling forks for a single command. Those children are
also neither login nor interactive, so they would scrub the environment
their parent deliberately set for them (GIT_DIR, VIRTUAL_ENV, ...), and
the rewritten scrub-report.txt would describe the last child to exit
instead of the pane.

Fix (per comment 15387, measured): keep .zshrc/.zlogin unconditional,
and add to .zshenv a pass guarded on the exact condition that defines
the gap (neither login nor interactive), plus a per-pane sentinel
(_CB633_SCRUBBED) so it runs once per pane, not once per process. The
sentinel is exported only after the scrub runs, and is folded into the
scrub's own allow-list so a later pass in the same pane cannot blank it
back to empty.

Also corrects the class javadoc's wrong premise (a bare argv[0] proves
NOT login, not "therefore interactive") and its now-stale two-file
walkthrough.

Tests: two new real-zsh tests in EnvAllowListScrubTest run actual
non-login/non-interactive zsh processes (never string-match the
generated files) to prove: a neither-shell pane is scrubbed; a child
that pane forks keeps variables the pane deliberately set for it; the
child does not re-scrub; and scrub-report.txt still describes the pane
after the child exits. Both fail without the production fix (verified
by reverting it and re-running: AssertionFailedError on the sentinel
and on the decoy secret surviving).
2026-09-10 07:21:08 +07:00
Dai Ha b9d09e044e t386: pin the per-member drift baseline the fix's own tests left open
CI / contract (push) Successful in 37s
CI / build (push) Successful in 1m57s
The two tests merged with #386 both start with the member already BUSY, so a
single global drift baseline passes them. This one sleeps the host while nothing
is busy and only then starts a turn, which fails without the per-member map.
2026-09-10 06:59:42 +07:00
Dai Ha fd8650cda4 Merge #386: correct the stall check for a monotonic clock frozen by host sleep 2026-09-10 06:56:37 +07:00
Dai Ha 11050e24ed t384: fix javadoc indentation on the merged shareWithGroup lines
CI / contract (push) Successful in 1m7s
CI / build (push) Successful in 1m29s
2026-09-10 06:54:51 +07:00
Dai Ha 769f282408 #386: give the stall detector a real-time clock, log the divergence
CI / contract (pull_request) Successful in 1m18s
CI / build (pull_request) Successful in 1m26s
FleetHealthMonitor.tick's stalled check compared two monotonic-clock
readings (System.nanoTime(), which macOS freezes across a host sleep),
so a member BUSY for 101 real minutes was never flagged.

The monitor now also takes a wall-clock LongSupplier (realtimeClock),
used only inside the stall check. Each tick measures how far the two
clocks moved apart since the previous tick and folds any positive
divergence into a running total; when a single tick's divergence
exceeds one tick interval (the signature of a sleep, since a tick
cannot run while the process itself is suspended) it logs one WARN
naming how long the detector could not see. The correction is applied
per member, keyed to when that member's current lastActivityAtNanos
was first observed BUSY - not since monitor start - so a sleep that
happened before a member went busy is never charged to it.

Every other use of the monitor's clock (readiness grace, snapshot
timestamp) is unchanged. Backend quarantine/cool-off, the lead tab
scan, the completion resolver, the session reaper and the message
service TTLs are untouched, per the ticket's decision.

Existing FleetHealthMonitor/FleetHealth tests pass unmodified (none of
them ticks a BUSY session more than once, so the drift path never
engages for them). Two new tests: a frozen monotonic clock past the
real-time threshold produces STALL_SUSPECTED, and a single sleep gap
logs the divergence exactly once, not once per tick.
2026-09-10 06:53:30 +07:00
Dai Ha 6f71f40047 Merge #384: pre-create the scrub receipt and give it group write 2026-09-10 06:49:46 +07:00
Dai Ha fb36c5238f Merge #382: SpawnRequest.withProfile() replaces the six-accessor rebuild 2026-09-10 06:49:46 +07:00
Dai Ha cc9cdc938b #384: write shared scrub receipts 2026-09-10 06:45:30 +07:00
Dai Ha 5fede82468 #382: give SpawnRequest a withProfile wither, guard it against the arity trap
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m49s
CompositePeerLauncher:372 rebuilt a routed SpawnRequest from a literal
new SpawnRequest(...) call listing six of the original request's own
accessors. That call is only correct because it happens to match the
canonical 6-arg constructor today; add a 7th component plus the
established back-compat constructor at the old (now-shorter) arity and
this call would silently rebind to it, dropping the new field on every
profile-routed spawn with no compile error — the same defect shape
already guarded on FleetConfig.withDefaults() (#357) and MemberSession
(#358).

Add SpawnRequest.withProfile(String), modeled on
MemberSession.withState/withActivity, and use it at the call site
instead. Add a guard test that resolves the true canonical constructor
by exact component types (never by argument count), gives every
component a distinctive value, and asserts every component but
profileName survives withProfile() unchanged.

Proved the guard against the real mechanism: temporarily dropped the
last (role) argument from withProfile()'s constructor call so it bound
to the 5-arg back-compat constructor — it still compiled, and the new
test failed, catching the silently-defaulted role. Restored the fix
and reconfirmed green.
2026-09-10 06:44:15 +07:00
Dai Ha 1515025804 charter: a profile differs in liveness, not just model and cost
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m31s
The old line said profiles differ in model and cost. That is true and it is
not the reason the default hurts. Measured on two hosts: a default sitting on
an exhausted or withdrawn credential either fails the spawn loudly or, worse,
produces a member that starts fine and then returns nothing.

Wording proposed by the fleet01 lead; merged with the existing 'not in tier'
clause, which is still right. Applied byte-identically to the wiki template.
2026-09-10 06:28:51 +07:00
Dai Ha 7754f53662 Merge t385 follow-up: pin the mid-turn redelivery ack 2026-09-10 06:28:08 +07:00
Dai Ha a507f7b31b t385: pin that a redelivery is acked even while the lead is mid-turn
Verification found the fix's dedup check could be moved behind the
injectable-status gate and every test still passed. That placement matters:
this lead is mid-turn most of the time, so gating the ack on an idle pane
leaves the redelivered message held, and the next recovery delivers it again.

The new test fails on that mutant and passes on the fix.
2026-09-10 06:28:03 +07:00
Dai Ha 71c322f104 Merge #385: a redelivered lead message must not write the pane twice 2026-09-10 06:24:40 +07:00
Dai Ha fde2c15627 t385: a redelivered lead message must not write the pane twice
An AMQP recovery clears the held delivery tags, the broker redelivers with
fresh ones, and the coordination loop wrote the same peer message into the
lead pane again. Measured on the live daemon: one msgId reached the pane 12
times in 9 hours, across 19 recovery events.

LeadCoordLoop now remembers the msgIds it has written to a pane (bounded at
1024) and acks a redelivery without a second write. LeadMailbox.ack no longer
returns quietly for an unknown msgId: it throws, so an ack that never reached
the broker is reported instead of hidden. A repeat ack that this connection
already completed stays quiet, tracked in a bounded set.
2026-09-10 06:24:35 +07:00
Dai Ha 2830735644 t377: raise the logger level in the test, or it asserts on an empty list
CI / contract (push) Successful in 46s
CI / build (push) Failing after 1m32s
The 9 tests the member wrote failed 7 of 9 on first build. The production
code was correct; the test harness was not. logback-test.xml sets
dev.ltms.fleet to WARN, so the INFO shape lines were dropped by the level
check before any appender saw them.

The member copied attach()/detach() from MemberTrustModelReportTest but not
the setLevel(INFO) those siblings do at each call site. Doing it inside
attach()/detach() covers all nine at once, and restores the original level
(null, meaning inherit) rather than a concrete one.

Proven by mutation: making the code log the value fails 4 tests, including
theValueNeverAppearsInLogOutput — 'the full GITEA_HOST value must never
reach the log'. Full build 1459 tests green.
2026-09-09 07:52:43 +07:00
Dai Ha 4e3ac91a22 Merge worker/t377-7f587b-7: startup reports the git host value's shape, never its value 2026-09-09 07:49:09 +07:00
Dai Ha 29d3f0b41f t377: report the git host value's SHAPE at startup, never its value
A member receives the git host as GITEA_HOST and often gets a full URL
(scheme, trailing slash) where it expects a bare host, so it builds
https://https://... and the request never leaves the machine.

Logs set/unset, length, startsWithScheme and trailingSlash next to the
existing secret report. The value itself is never logged, and the value
is still passed to members unchanged — rewriting it here would change
what works on one host and breaks on another.

Written by the gx member on branch worker/t377-7f587b-7; it could not
build, commit or push because the host command classifier refused every
shell command (see #381). Build and verification are mine.
2026-09-09 07:49:02 +07:00
Dai Ha cbc732444f fleetd #365: the push-nudge metric outcome is 'sent', not 'delivered'
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m31s
The label was renamed in the #365 merge. This table still named the old one.
'sent' counts the herdr paste-and-submit call returning, never a confirmation
that the pane read it.
2026-09-09 07:40:57 +07:00
Dai Ha 6f828b8c38 Merge worker/t365-3920c5-3: t365: fleet_reply/REST reply distinguish resolved vs queued; rename nudge metric outcome
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m48s
2026-09-09 07:35:00 +07:00
Dai Ha 145a8c8862 Merge worker/t373-336973-2: t373: pin the production XDG-excludes seam GitWorktreesTest.seedingGitWorktrees builds 2026-09-09 07:35:00 +07:00
Dai Ha 7057291739 Merge worker/t358-6e989b-1: t358: guard MemberSession's 5 rebuild sites and Profile.withProfile() against the back-compat-arity trap 2026-09-09 07:35:00 +07:00
Dai Ha 2302b3bc11 #376: the too-fast failure stops asserting a cause it cannot know
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m31s
A turn that settles inside MIN_TURN_NANOS is still reported FAILED. Only the claim
about WHY is withdrawn, and the pane is still carried.

Measured on fleet01 on 2026-09-08 UTC: an opencode member on mimo-v2.5-free answered
a real question in 1575ms, below the 2000ms floor. The daemon reported the turn as
failed with 'most likely a backend error before any work started'. The answer was
right there in the scrape. With no errorPattern configured — the live state on both
hosts, which both log as 'backend-error classification: off' — that sentence is a
guess, and a reader who believes it stops looking at the pane.

WHAT I REJECTED, because the next person will try it. A worker implemented the
ticket's first suggested direction: inspect the pane inside the floor and resolve a
COMPLETION when the text looks like a real reply. Its test for 'looks like a real
reply' was non-blank plus a '.', '!' or '?' anywhere in the text. That is unsafe
twice over. lastAssistantBlock falls back to the WHOLE pane when it finds no U+23FA
marker, so on a crash the candidate reply is the entire screen; and a crash pane
almost always contains a full stop, in a file path, a version or a hostname. I ran
that implementation against the new guard test and it resolved

    Error: connection reset while loading src/main/java/Foo.java v1.2.3

as a COMPLETION — expected: <FAILED> but was: <COMPLETION>. A loud wrong answer
became a silent one, which is the trade the ticket brief forbade.

The obvious repair does not work either. Requiring the U+23FA marker as positive
evidence would be safe, but that marker is Claude Code chrome and an opencode pane
never carries it — and an opencode member is what raised this ticket. There is no
reliable cross-backend marker for 'this is a real reply', so this path must not try
to judge one. That reasoning is now in the failTooFast javadoc.

Two guard tests. aPlausibleLookingReplyInsideTheFloorStillFails pins the safety
property against exactly the rejected approach; it passes today and fails against
that implementation, which is how it was verified rather than assumed.
theTooFastFailureDoesNotAssertACauseItCannotKnow pins the wording.

One existing test changed. aNonMatchInsideTheFloorStaysGenericAndNeverNotifiesTheSink
asserted the phrase 'too fast to be real work', which carried the withdrawn claim. It
now asserts what it was really guarding: the floor alone fails the turn, the reason
stays generic, the pane is carried, and the typed sink is never notified.

MIN_TURN_NANOS is unchanged at 2000ms. mvn clean install: 1441 tests, 0 failures.
2026-09-09 07:28:04 +07:00
Dai Ha 3a004dc1b3 t373: pin the production XDG-excludes seam GitWorktreesTest.seedingGitWorktrees builds
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m32s
fleetd #362 review finding 2 protects GitWorktrees#previouslyEffectiveExcludesFileContent's
Java-side XDG_CONFIG_HOME/HOME read (it never goes through a git subprocess, so no
GIT_CONFIG_GLOBAL/GIT_CONFIG_SYSTEM isolation reaches it) with a gitEnv constructor seam. A
mutation run during the #372/#369 merge found that seam unpinned: stripping hermeticGitEnv(tmp)
from seedingGitWorktrees left every test green, poisoned XDG_CONFIG_HOME or not.

Adds seedingGitWorktreesResolvesTheExcludesFileFallbackInsideItsThrowawayDirectory, which asserts
the property directly (a GitWorktrees built for seeding resolves the fallback inside its own
throwaway directory) using a self-contained marker instead of relying on an externally poisoned
env var. Refactors hermeticGitEnv/seedingGitWorktrees into two-argument overloads (one taking an
explicit XDG_CONFIG_HOME / gitEnv) so the new test can pre-populate the marker before construction
while still going through the same production construction every other seeding test uses; no
behavior change for the 4 existing call sites.
2026-09-09 07:24:22 +07:00
Dai Ha a9a3c12232 t365: fleet_reply/REST reply distinguish resolved vs queued; rename nudge metric outcome
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m51s
fleetd #365. fleet_reply always returned the literal "delivered" and
POST /sessions/{id}/reply always returned {"delivered": true}, whether
the reply resolved a live waiting send/ticket or was merely queued in
the inbox for a later drain (CB-307) — both are successes, but not the
same fact.

MessageService.reply() now returns a ReplyOutcome (RESOLVED_SEND,
RESOLVED_ASYNC_TICKET, or QUEUED) instead of an always-true boolean.
FleetMcp.reply and FleetApp.replyMessage both read it: the MCP tool
result names which happened, and the REST body's "delivered" field is
now accurate, with an added "outcome" field.

Also renames the heartbeat/push-loop nudge metric's "delivered" outcome
to "sent" (LeadHeartbeatLoop, ReplyPushLoop, FleetMetrics): it only
records that the herdr agent.prompt paste-and-submit call succeeded,
never that the lead's pane actually read it — there is no read-receipt
concept at that layer, so "delivered" overclaimed there too.

Tests: MessageServiceTest/FleetMcpTest/FleetAppTest strengthened to
assert the specific outcome per case (a resolved send, a resolved async
ticket, and a queued reply); ReplyPushLoopTest updated for the outcome
rename.
2026-09-09 07:23:41 +07:00
Dai Ha 94ec77a1bc t358: guard MemberSession's 5 rebuild sites and Profile.withProfile() against the back-compat-arity trap
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m53s
Follow-up to #357 (FleetConfig.withDefaults()). Reflectively enumerate each
record's own components, resolve the canonical constructor by exact
component types, build a real non-null value per component, run each
rebuild site, and assert every component survives (except the one it is
documented to change). Exclusion lists pinned at 0 for both.

Re-counted the Profile back-compat ladder directly against the source:
8 constructors (arities 25, 24, 22, 20, 18, 15, 14, 12) against a
canonical arity of 26 — the ticket's own number was explicitly untrusted.
2026-09-09 07:19:12 +07:00
Dai Ha 49f285cfda charter: a measured fact in an addendum must carry its own deletion trigger
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m36s
The operator chose this and its size; the reasoning below is the fleet01 lead's.

Placed in the canonical block's boundary paragraph rather than the orchestration
body. That paragraph already talks about the addendum layer instead of protocol,
every project that mounts the bridge inherits it, and it sits about 3800 characters
before the primary's step list, so it does not dilute the steps a lead reads while
working. Perishability is structurally an addendum problem: the block is
byte-identical across projects by construction, so a dated local measurement in the
block body would already be a layering violation.

What happened. The fleet01 lead's kb addendum held a dated merge-refusal section
that carried an instruction to delete itself once it stopped reproducing. On
2026-09-08 UTC the operator granted merge rights on akb/kb, the lead re-ran the
probe, got 409 'head out of date' where the identical request had returned 405
'User not allowed to merge PR' on 2026-09-06, and deleted the section as instructed.

Why four parts and not one. The lead's finding is that the banner did not work
because it was emphatic. It worked because the falsification condition was
executable: it carried the exact probe, the reason for the all-zeroes
head_commit_id, and what each response code meant. The lead did not have to
reconstruct the experiment or decide what would count as refutation, and just ran
it. A banner saying 'this may be out of date, verify before relying on it' costs
the same space and does nothing, because deciding what would falsify a claim is the
expensive step and a reader in the middle of another task will not pay it. So: the
date, the command, what each outcome means, and the instruction to delete. The
fourth without the second is decoration.

The closing clause is the justification for the machinery. Most stale notes are
merely wrong. This one went stale in the dangerous direction: it would have told a
future lead it could not merge at the exact moment merging became its job, silently
and with confidence. A note that goes harmlessly stale does not need this.

Note what is NOT centralized here. The banner text itself cannot be. What fired for
the lead was a specific instruction sitting on top of the specific stale fact, which
it could not read past on its way to acting. A rule elsewhere saying 'date your
measurements' would not have fired, because nobody reads that rule at the moment
they re-measure. This sentence sets the convention; the trigger still has to live
next to the fact it governs.

Propagated to the wiki template in the same turn, wiki 8c4f152 on main; the sync
check in this file's addendum reports 'in sync: True'.
2026-09-09 04:24:17 +07:00
Dai Ha 5a12ae7930 charter: test a refusal, and do not count a transport failure as one
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m38s
Step 8 gained a refusal paragraph in 2f71a30, which said what a lead does once the
forge refuses a merge. It did not say how a lead establishes that it was refused.
Both halves of this amendment come from the fleet01 lead, measured on akb/kb on
2026-09-08 UTC, and both are ways to be wrong about a permission you never tested.

Do not read a refusal off a permissions field. After the operator granted merge
rights, the lead re-ran its probe: POST .../pulls/53/merge with an all-zeroes
head_commit_id, chosen so the request cannot succeed on its merits and a rejection
can only mean the refusal. It returned 409 'head out of date' where the identical
request returned 405 'User not allowed to merge PR' on 2026-09-06. A 409 is payload
validation and sits after the permission gate, so the grant took. The lead reports
the repository permissions object did not change across that flip -- still
admin:false, push:true, pull:true. I did not read that object myself; my forge token
is a different identity and would return a different one, so this stays the lead's
measurement and not mine. Merge rights on a protected branch live in branch
protection, so a permissions field can be wrong in both directions.

Do not count a transport failure as a refusal. The lead's first attempt returned
HTTP 000, because GITEA_HOST already carries a scheme and a trailing slash and the
URL came out as https://https://git.ltms.dev//api/... Under a 'not 200' test that is
indistinguishable from being refused. A probe exists to separate a refusal from
everything else, so an error that never reached the gate has to be a third answer
that concludes nothing.

Propagated to the wiki template in the same turn, wiki 8c2ef96 on main; the sync
check in this file's addendum reports 'in sync: True'.
2026-09-09 04:09:08 +07:00
Dai Ha 127e6832a9 Merge #374: fleetd holds off idle sleep while any member is live
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 2m3s
Lands PR #355 (fleetd #354's sibling), rebased onto current main by a worker
after 39 commits of drift left it unmergeable.

The problem, measured on the original branch: a fleetd host idle-slept after as
little as one minute (pmset -g custom reported 'sleep 1' on battery). Overnight
the daemon's AMQP link dropped 13 times, and every drop minute had a sleep or
wake event in pmset -g log in the same minute or the one before. The AMQP churn
is the visible symptom; the real cost is a member mid-turn freezing with the
host, and a long turn with nobody typing is exactly the case that goes idle.

IdleSleepGuard holds an OS-level assertion for as long as at least one member is
live. It is driven by SessionManager's existing onAcquire/onRelease hooks rather
than a second member count kept in parallel, so it reads the same registry
fleet_list's numbers come from, and only a real 0->1 or 1->0 crossing touches the
OS. It fails safe: a mechanism that cannot acquire means nothing is ever held,
and it never throws, never blocks a spawn, a release, or shutdown.

Conflict resolution was the whole job, and all three were in config plumbing:
ConfigRef, FleetConfig and ConfigRefTopLevelReportingCoverageTest. The power
package is byte-identical to the original branch commit.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1439, Failures: 0, Errors: 0, BUILD SUCCESS
  (1425 on main + 14 new: 4 caffeinate, 5 guard, 1 wiring, 4 config)

The denominator recount, which the worker flagged as its own weakest number
because this file's count has drifted three times before (#330/#333/#337). I
counted it mechanically rather than reading it: FleetConfig has 24 canonical
record components; COLD_KEYS 5, DEFERRED_KEYS 13, SPLIT_KEYS 3, plus the 3 the
javadoc names as hot-excluded (placement, memberCredentials, memberLoginShell).
5+13+3+3 = 24. The javadoc's '24 components: 5 cold, 13 deferred, 3 split, 3
hot-excluded' is correct. The worker's prose called idleSleepGuard the 25th
constructor argument; it is the 24th. The code is right, the report was off by
one.

Mutation run on merge, on the half the worker verified by READING rather than by
proving -- it said it had checked that withDefaults()'s final call binds the true
canonical constructor. I dropped the trailing idleSleepGuard argument so the call
silently binds the 23-arg back-compat overload. It compiles, which is the whole
hazard. Caught: 1 failure, 3 errors, BUILD FAILURE, and
FleetConfigWithDefaultsPreservesEveryComponentTest names the dropped component
and prints its own denominator -- '24 components, 24 checked, 0 excluded, 23
survived'. That test was added on main after this exact defect happened live when
idleSleepGuard was added on a sibling branch; the worker had to add the missing
entry to it, and doing so is what makes the guard cover this component at all.
2026-09-07 20:31:41 +07:00
Dai Ha 6442a583ae t355: give the withDefaults() coverage test a value for idleSleepGuard
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m6s
FleetConfigWithDefaultsPreservesEveryComponentTest was added on main after
the sleep-guard commit was cherry-picked, specifically to catch this exact
rebase hazard (a withDefaults() call silently rebinding to a stale-arity
back-compat constructor). Its baseValues() map didn't know about the new
idleSleepGuard component yet, so the coverage test itself failed the
name-drift check. Add a real, non-null value for it, consistent with how
the sibling ConfigRefTopLevelReportingCoverageTest already covers it.
2026-09-07 20:25:33 +07:00
Dai Ha 24b96d29ae fleetd must not let the host idle-sleep while members are live
Adds a small IdleSleepGuard (dev.ltms.fleet.power) that holds a macOS
caffeinate -i child while at least one fleet member is live, and
releases it once none are. It hangs off SessionManager's existing
onAcquire/onRelease hooks and SessionManager#size() rather than
tracking members a second way. New idleSleepGuard: config block,
on by default, following the FleetConfig.Health/ConfigReload pattern.
2026-09-07 20:22:11 +07:00
Dai Ha 4721771052 charter: a peer reads your project addendum
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m47s
Fourth item under Lead <-> lead. The addendum is instruction surface -- every
future session on that host obeys it, and a wrong one is obeyed as faithfully as
a right one -- but nothing in the block said to have anyone check it.

Evidence is one addendum, written on fleet01 this week, and it carried two
defects that its author did not see and a non-author did. First: it described a
forge permission wall as if it were policy, so a token regrade would have made
it tell a session NOT to merge at the moment merging became its job. Second,
found only because the first was raised: it DID carry the 'this is a dated
measurement' caveat, at the bottom, after the prohibition -- so a session that
read the instruction and stopped had already taken it as policy. The caveat sat
downstream of the thing it qualified.

Neither is a writing slip. Both are the author being unable to see their own
qualifier placement, which is what a second reader is for.

Deliberately conditional: not every operator runs two leads, so the rule ends
with what a lone lead can still do -- ask which sentence goes false first, and
whether a reader reaches the caveat before acting.

Weaker evidence than the step 8 amendment (n=1 addendum, 2 defects, versus a
measured 405). Recorded as such so it can be dropped if it does not earn the
context it costs.
2026-09-07 05:03:58 +07:00
Dai Ha 2f71a30bd7 charter: say what step 8 means when the forge refuses the merge
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m35s
Step 8 said 'then merge' and nothing about a refusal. fleet01 is the first host
to hit that: on akb/kb its lead gets 405 'User not allowed to merge PR' from the
API, and main is protected, so merging locally and pushing is refused too. Both
routes shut. The charter was telling a lead to do something the forge would not
let it do, and that gap was latent in every copy of the block.

The amendment keeps the rule that the merge decision is never delegated, while
admitting the mechanical merge may not be the lead's to make.

The second sentence is the one that earns its keep, and it came from the fleet01
lead rather than from me: never call a PR 'ready to merge' without having read
the diff. A refusal is exactly when that shortcut is tempting, because no action
is left that forces the lead to look. Without it, a refused merge quietly turns
step 8 from 'read it yourself, then merge' into 'forward the reviewer's verdict'
— the proxy-delegation the same step forbids two lines earlier.

Propagated to the wiki template in the same turn; the sync check in this file's
addendum reports 'in sync: True'.
2026-09-07 05:01:31 +07:00
Dai Ha 22cdebbdbe fleetd #369: state hermeticGitEnv's real scope, measured
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m32s
The comment claimed no test in this class can reach the real machine's home
directory. Measured on the merge: 53 of the 58 new GitWorktrees(...)
constructions here pass no env override, and stripping the override from
seedingGitWorktrees leaves the class green under the poison command that the
same comment cites as proof. Say what it covers and what it does not.
2026-09-06 20:32:21 +07:00
Dai Ha 154971c2b8 Merge #372: GitWorktreesTest no longer reads the operator's real git config
fleetd #369. Every git subprocess the test class starts now goes through one
gitProcessBuilder factory that applies the hermetic environment. Before this,
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus set nothing at all -- so the tests inherited the JVM's
whole real environment, including the operator's default excludes file. That
file applies with no core.excludesFile configured at all, and /dev/null for the
global config does not stop it.

Round 1 pinned this with a call-site count: exactly 2 literal
new ProcessBuilder( occurrences. That catches a NEW helper built the old way,
but it is a proxy, not the property. I measured the gap -- deleting
pb.environment().putAll(hermeticEnv()) from inside the factory left every call
site unchanged, the count stayed 2, and 1413 tests stayed green. Round 2 added
gitProcessBuilderCarriesTheFullHermeticEnvironment, which asserts on what the
factory actually hands to ProcessBuilder#start(). Both checks are kept: they
catch different regressions.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1414, Failures: 0, Errors: 0, BUILD SUCCESS
  poison control on the PRE-FIX file:
    XDG_CONFIG_HOME=<dir with a '*' git/ignore> mvn test -Dtest=GitWorktreesTest
    -> tests=59 failures=56, so the poison genuinely reaches these tests

Mutation run on merge, on a half neither the worker nor the reviewer touched --
made seedingGitWorktrees pass null instead of hermeticGitEnv(tmp), which strips
the hermetic environment from the PRODUCTION GitWorktrees instances rather than
from the test's own subprocesses: tests=61 failures=0 unpoisoned AND poisoned.
That half is unpinned. It is fleetd #362's protection, not this ticket's, and it
guards a different path -- the Java-side XDG read in
previouslyEffectiveExcludesFileContent, which no assertion observes. Out of
scope here; filed as a follow-up rather than held against this PR.

One javadoc sentence corrected in the merge: hermeticGitEnv claimed "no test in
this class can reach the real machine's home directory". 53 of the 58
new GitWorktrees(...) constructions in this file pass no env override at all, so
the claim is true of the 5 seeding sites and of every test-started subprocess,
not of the class.
2026-09-06 20:31:01 +07:00
Dai Ha fef287c346 Merge #371: a stale lead binding no longer swallows a worker's reply nudge
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m55s
fleetd #368. PrimaryRegistry.forgetDelegation fires only when a WORKER is
released, never when the delegating LEAD goes away. The map is keyed by the
worker, so a closed, crashed or relaunched lead left its bindings behind. A
stale non-null entry then beat nudgeTargetFor's single-primary fallback every
time -- and that fallback's own javadoc argues it is correct precisely in the
case the stale entry was hiding.

ReplyPushLoop now probes the recorded lead with agents.status before trusting
it, at all five entry points, and a lead that is really gone is forgotten so
resolution reaches the fallback.

Round 1 caught any RuntimeException and treated it as death, and death here
calls forgetDelegation -- destructive and permanent on ONE reading. That is the
#359 mistake repeated two days later: a socket blip or a decode error on a live
lead would silently unbind it forever. I measured the breadth was unpinned
(narrowing it left 1414 green), and round 2 narrowed it to the one affirmative
signal AgentControl.agentCall itself uses, agent_not_found. Every other failure
is treated as live, because guessing wrong costs one extra retry next tick while
guessing wrong the other way is unrecoverable.

The fallback it lands on is live, not stale: PrimaryRegistry.record overwrites
the single slot on every orchestration-side MCP call, and neither host pins
primary.terminal, so a relaunched lead re-registers on its first tool call.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1415, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge, on a half the worker did not touch -- dropped the
forgetDelegation call while still returning the fallback, so behaviour on the
first nudge is identical and only the self-healing is lost: 2 failures,
BUILD FAILURE. The cleanup is pinned, not just the fallback.
2026-09-06 20:13:07 +07:00
Dai Ha a37acd5ee3 fleetd #369 review round 2: pin the factory's behaviour, not its call count
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m26s
everyGitSubprocessGoesThroughTheHermeticFactory counts ProcessBuilder("git",
...) call sites, so it catches a new helper built the old way, but a
reviewer proved it does not catch gitProcessBuilder itself being gutted:
removing pb.environment().putAll(hermeticEnv()) from inside the factory
leaves every call site unchanged, the count stays 2, and the whole
unpoisoned suite stays green.

Add gitProcessBuilderCarriesTheFullHermeticEnvironment, which inspects what
the factory actually hands to ProcessBuilder#start(): every hermetic key
present with the isolating value, and XDG_CONFIG_HOME pointed inside the
class's own throwaway directory rather than left unset or pointing at the
operator's real one. This fails the moment the hermetic environment stops
being applied, on any machine, with no poison needed. Keep the call-site
count check too — the two catch different regressions.
2026-09-06 20:09:32 +07:00
Dai Ha d6ef0c8013 fleetd #368 review: only agent_not_found may forget a lead binding
CI / build (pull_request) Successful in 1m22s
CI / contract (pull_request) Successful in 1m22s
isLive treated any RuntimeException from the liveness probe as "the lead is
gone", which forgetDelegation then acted on destructively and permanently.
That made a transient herdr hiccup (socket blip, decode error) on a perfectly
live lead indistinguishable from the lead actually being dead — the same
one-bad-reading mistake #359 shipped a guard against for lead-tab liveness.

Narrow isLive to match AgentControl.agentCall's own rule: only an affirmative
HerdrException("agent_not_found") counts as gone. Every other failure is
treated as still live and the binding is left alone.

Adds aTransientLivenessFailureMustNotForgetABindingToAStillLiveLead, which
fails with the bare RuntimeException catch and passes with the narrowed one.
2026-09-06 20:08:18 +07:00
Dai Ha dfb70871b4 fleetd #360 follow-up: the install block must say loginctl enable-linger
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m54s
PR #370 shipped both units but its install block stopped at
`systemctl --user enable --now`. Without lingering a user manager starts at
your first login and stops at your last logout, so the units do not come back
after a reboot -- which is the whole reason this ticket moved fleet01 off the
setsid scripts.

It is easy to miss because leaving it out looks like success: `systemctl --user
enable` reports "enabled" and both units run while you stay logged in. The
issue named this and the PR did not carry it over.

fleet01 itself is fine -- measured `Linger=yes`, both units `enabled`. This is
about the next host that follows these instructions.

Comment only; SystemdUnitSafetyTest still 8 green (a commented line is not an
active directive).
2026-09-06 20:08:05 +07:00
Dai Ha 4ee7b16929 Merge #370: ship systemd units that do not silently disable the daemon
CI / contract (push) Successful in 48s
CI / build (push) Successful in 2m1s
fleetd #360. deploy/fleetd.service shipped four mount-namespacing directives
(ProtectSystem, ProtectHome, ProtectKernelTunables, ProtectControlGroups),
PrivateTmp=true, and an ExecStart that ran java directly. Each one starts green
and breaks the daemon in a way nothing logs: lsof goes blind so every caller is
resolved ANONYMOUS and refused; the member ZDOTDIR scrub becomes a no-op; every
credential is empty. deploy/herdr.service did not exist at all, though
fleetd.service's After=/Wants= already named it.

Both units are now the ones running on fleet01, comments included -- the
bisected lsof counts and the reasons live in the files, because the next person
to 'harden' this needs the reason, not the rule.

A unit file has no compile step, so SystemdUnitSafetyTest reads both units plus
deploy/herdr-inner.sh and fails on an active forbidden directive, on
PrivateTmp=true, on an ExecStart that skips the login shell, on a herdr-inner.sh
that does not exec a login shell, and on one that does not set a non-zero pty
size. Each message names the consequence. A vacuity guard pins that all three
files exist and that the DO-NOT-add comment block still mentions every forbidden
directive, so 'comment survives, directive does not' is actually exercised.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1420, Failures: 0, Errors: 0, BUILD SUCCESS

Round 1 shipped herdr-inner.sh untested; I mutated it (dropped both the login
shell and the stty sizing) and got 6 green. Round 2 added those two checks.

Mutation run on merge, on a half the worker never touched -- ProtectHome=read-only
added to herdr.service, the unit it only ever mutated fleetd.service for:
Tests run: 8, Failures: 1, BUILD FAILURE. The parameterisation really covers
both files.
2026-09-06 20:04:48 +07:00
Dai Ha 8557289dc0 fleetd #360 review round 2: extend SystemdUnitSafetyTest to herdr-inner.sh
CI / build (pull_request) Successful in 1m27s
CI / contract (pull_request) Successful in 1m30s
herdr.service's ExecStart only names deploy/herdr-inner.sh, so that script was the only
place herdr's login-shell and pty-size properties lived, and nothing was reading it -- the
same silent-at-startup shape the ticket was about, one file further down the chain. A
mutation dropping both the login shell and the stty sizing left SystemdUnitSafetyTest green.

Add herdrInnerScriptUsesALoginShell and herdrInnerScriptSetsANonZeroPtySize, extend the
vacuity guard to cover herdr-inner.sh too, and tighten execStartUsesALoginShell's check from
a bare contains("-lc") substring match to the same login-shell-invocation regex the new
checks use.
2026-09-06 19:54:36 +07:00
Dai Ha dd2efd8541 fleetd #369: make GitWorktreesTest hermetic against the machine's real git config
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Failing after 1m28s
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus (plus every other raw git subprocess in this class) set no
isolation at all — inheriting the JVM's real environment, including the
operator's real ~/.gitconfig and default excludes file. Measured: with a
poisoned XDG_CONFIG_HOME, 56 of 59 tests failed.

Centralize every git subprocess this test starts through one factory,
gitProcessBuilder, which always applies the existing hermeticGitEnv isolation
(extended with XDG_CONFIG_HOME, the same fix #366 already applied to the
production-instance seam). Add a self-check test that counts direct
ProcessBuilder("git", ...) constructions in this file's own source and fails
if a future helper bypasses the factory, so the omission that caused this
ticket is caught by name instead of rediscovered on a poisoned machine.
2026-09-06 19:54:35 +07:00
Dai Ha 5af786d135 fleetd #368: a stale lead delegation binding must not shadow the primary fallback
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m55s
PrimaryRegistry.forgetDelegation only fires when a WORKER is released, never when
the delegating LEAD terminal itself disappears (closed, crashed, or relaunched).
A stale, non-null leadByTarget entry always beat nudgeTargetFor's single-primary
fallback, so a dead lead silently swallowed every reply nudge for its workers.

Fix: ReplyPushLoop now verifies (via the same agents.status check decide() already
uses every tick) that a recorded delegating lead is actually live before trusting
it. A dead lead is treated as if never recorded — self-healing the binding
(mirroring AgentControl.paneByTerminal's self-heal on agent_not_found) and falling
through to PrimaryRegistry's existing fallback.
2026-09-06 19:54:26 +07:00
Dai Ha 380eb63277 fleetd #360: fix systemd units that silently disabled caller identity and the credential scrub
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 2m12s
deploy/fleetd.service started clean on fleet01 but broke the daemon in three ways nothing
logs: ProtectSystem/ProtectHome/ProtectKernelTunables/ProtectControlGroups each put the unit
in its own mount namespace, which blinds fleetd's lsof-based caller lookup and falls every
caller back to ANONYMOUS; PrivateTmp=true silently no-ops the credential scrub the member
pane depends on; and running java directly from ExecStart skips the login shell that sources
the daemon's secrets, so it boots with empty credentials.

Replace the unit with the version verified working on fleet01 for a day, and add the
deploy/herdr.service companion unit it was already depending on via After=/Wants= but which
did not exist in the repo. Add deploy/herdr-inner.sh as the login-shell template
herdr.service's ExecStart wraps in a pty.

Add SystemdUnitSafetyTest (fleetd/src/test/java/dev/ltms/fleet/deploy) to read both unit
files from disk and fail if a forbidden mount-namespacing directive is active, PrivateTmp is
true, or fleetd.service's ExecStart does not go through a login shell -- the only guard
possible for a unit file with no compile step.
2026-09-06 19:45:23 +07:00
Dai Ha 6c61355f8f Merge #367: dead lead tabs stop accumulating, without destroying a live one
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m7s
fleetd #359. LeadTabScanner used to join labelled tabs straight to terminals
with no liveness check, and its javadoc excused that ("a stale name costs
nothing here"). It cost plenty: LeadCoordLoop reads that map to pick which
pane a peer message goes into, so a dead tab was a candidate it could pick.
LeadLauncher, meanwhile, had no cleanup path at all -- every reconcile that
found 0 live created another tab and left the old one.

Both now cross-check agent.list, and neither trusts a single reading of it.
That matters because the daemon's own evidence on fleet01 was agent.list
reporting 0 live while ps showed one real claude. A first cut of this fix
closed tabs on that single reading, which would have closed the operator's
live lead instead of leaving a spare tab. So: a dead tab is flagged, not
closed, and only closed when a later reconcile still finds it dead; and the
scanner grants one grace scan to a terminal it already knew was live.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1412, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge (PendingCloseMarker.strip made identity, so a flagged
tab stops matching its configured label): 4 failures, BUILD FAILURE.

Known and accepted: ensureLeads() runs at startup, so the second reading
arrives at the next restart. A tab that dies mid-session stays flagged and
open until then. Deliberate -- the bug is about repeated restarts, and one
leftover tab is cheaper than closing a live session on unverified evidence.
2026-09-05 16:00:04 +07:00
Dai Ha 96c406b968 #359 review: require two independent readings before destroying a lead's tab
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m33s
Finding 1 (LeadLauncher): closing a labelled tab on a single agent.list
miss could destroy a live lead's session — the ticket's own evidence
showed that exact signal missing a genuinely running agent. A dead
reading now only flags the tab (PendingCloseMarker); it is closed only
if a later, independently-connected reconcile still finds it dead
while flagged. A tab found live again has its flag cleared instead.

Finding 2 (LeadTabScanner): the new agent.list cross-check in scan()
was not covered by get()'s "keep the cache on a failed scan" contract,
which only fires on a thrown HerdrException. A successful-but-short
agent.list could silently drop a lead CallerResolver had already
resolved, demoting it to Role.WORKER. A terminal already reported live
now gets one grace scan before being dropped; a terminal never
reported live gets none, so the original #359 exclusion is unaffected.

Both mechanisms were mutation-tested: reverting either change turns
exactly its own new tests red and nothing else.
2026-09-05 15:54:52 +07:00
Dai Ha a5d81c3f70 Merge remote-tracking branch 'origin/main' into worker/359-dead-lead-tabs-f1253b-4 2026-09-05 15:28:49 +07:00
Dai Ha 92c0f164f1 Merge #366: seed the bridge's worker skills into every provisioned worktree
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m46s
fleetd #362 item 3. A member spawned against a repo that does not ship its
own .claude/skills/ could not load implementer, reviewer or hunter at all.
Every brief starts with "Load the <name> skill", and outside this repo that
line was silently a no-op. memberSkills: <dir> now copies those folders into
each provisioned worktree, skipping any name the target repo already ships.

Two review rounds, both about the same hazard: core.excludesFile is
single-valued, so pointing it at fleetd's own file would SHADOW the
operator's. It now composes instead of replacing, and the XDG default
excludes file is carried forward too.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1402, Failures: 0, Errors: 0, BUILD SUCCESS

Two mutations run on merge:
  drop the XDG fallback          -> 1 failure, BUILD FAILURE
  remove the composition itself  -> 2 failures, BUILD FAILURE
Both directions are pinned.
2026-09-05 13:32:16 +07:00
Dai Ha 9e813ec179 fleetd #362 review fix 2: route the XDG excludesFile fallback through gitEnv too
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m24s
Finding 1 (lead): the XDG fallback branch of previouslyEffectiveExcludesFileContent was
unpinned — deleting it left the suite green (Tests run: 1386, Failures: 0). Added
seedSkillsComposesWithTheXdgDefaultExcludesFileWhenNoneIsConfigured to pin it: isolates
XDG_CONFIG_HOME via the gitEnv seam at a temp dir carrying a synthetic git/ignore, points
GIT_CONFIG_GLOBAL at an empty file so core.excludesFile is genuinely unset (forcing the
fallback branch), seeds a skill, and asserts a file matching the XDG-default pattern still
reads as clean. Reverting the fix (mutating the fallback to resolve to "") turns this test
red with a real pasted failure (see PR body): "expected: <> but was: <?? xdg-fallback-marker>".

Finding 2 (lead, the one that actually needed a code fix): the fallback read XDG_CONFIG_HOME
and HOME straight from the JVM's own environment, not through the gitEnv seam every git
subprocess in this class already honours — so no test could isolate it, and on a machine
carrying a real ~/.config/git/ignore (this dev machine does), every seeding test silently
composed with that real file. Added resolveEnv/resolveHome, which check gitEnv first and
fall back to the JVM's real environment only when the seam doesn't supply a value (production
behaviour, where gitEnv is always Map.of(), is unchanged). Added a hermeticGitEnv() test
helper and routed every seeding test in GitWorktreesTest through it, so no test in the class
can reach the real machine's home directory for this fallback.

Also documents two non-defects the lead asked for one javadoc line each on: the composed
excludesFile is a snapshot taken at seed time, not a live reference to the operator's file;
and excludeSeededSkillsFromGitStatus assumes a fresh worktree (not idempotent, but the
double-seed path does not exist today, so no guard was added for it).
2026-09-05 13:23:39 +07:00
Dai Ha 395b3b5c46 #359: LeadTabScanner drops dead lead tabs; LeadLauncher closes them on relaunch
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m49s
LeadTabScanner.scan() joined labelled tabs to panes with no liveness check at
all, so a tab left behind by a crashed/relaunched lead read as a live lead
forever -- exactly the hazard its own javadoc predicted but excused. It now
cross-checks agent.list, the same signal LeadLauncher already trusted, and
drops any labelled tab with no agent running in it. That alone removes the
duplicate candidates LeadCoordLoop.resolveLocalLead() could pick from,
including the dead one its own WARN's advice (name a lead after
coordinator.selfId) could land a message in.

LeadLauncher never actually stopped the accumulation: relaunching on "0 live"
always created a brand new tab and left the old dead-labelled one right where
it was, so any restart that found 0 live for any reason (a real crash, or a
herdr read that missed a still-running agent) added one more dead tab,
forever. ensureLeads() now closes every dead-labelled tab for a lead as part
of the same reconcile that decides to relaunch, so at most one tab survives
per configured lead once a restart's reconcile has run.
2026-09-05 13:20:37 +07:00
Dai Ha 6e9e464d62 Merge #364: fleet_list reports lead-coordination state instead of guessing at it
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m56s
fleetd #361. fleet_list's coordination block now carries this daemon's own
mailbox state and one row per operator-declared peer, each as a tri-state
status (exists / absent / unknown) rather than a boolean. pending and
consumers appear only when status is "exists", so an unresolved probe can
never render as a measured zero.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1394, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge (isMissingQueue always returns true, which restores
the exact defect the ticket fixes): 4 failures, BUILD FAILURE. The
discriminator's false branch is pinned.
2026-09-05 13:19:46 +07:00
Dai Ha 105c065615 Merge origin/main into worker/362-worktree-skills-c03e51-3 (picks up #363) 2026-09-05 13:18:32 +07:00
Dai Ha 01492059d4 fleetd #361 review round 2: pin isMissingQueue's false branch
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m34s
The reviewer's mutation (isMissingQueue always returns true) restored
the exact overstatement fleetd #361 exists to fix -- every declare
failure reading as a confirmed absence -- and still left mvn clean
install green (1389/1389), because no test drove a non-404 shape
through inspect(). The false branch was the whole discriminator
between MailboxState.absent() and MailboxState.unknown(), unpinned.

Widened LeadMailbox.isMissingQueue from private to package-private and
added LeadMailboxIsMissingQueueTest: five hermetic tests (no broker)
covering the true case and all three false shapes isMissingQueue's own
javadoc lists -- a different reply code, a ShutdownSignalException
whose reason isn't a Channel.Close, and an IOException with no such
cause at all (plus an IOException wrapping an unrelated exception
type). Re-ran the reviewer's exact mutation locally: 4 of 5 new tests
went red with the expected assertion messages; reverted, and mvn clean
install is green again at 1394/1394 (1389 + 5 new).

Also added a one-line javadoc note on LeadMailbox.inspect being honest
about which of its two RuntimeException catches is proven by a test
(the createChannel() one, end-to-end against a real broker) and which
stays purely defensive (the declare-site one, for a connection-drops-
mid-call race no test drives on purpose).
2026-09-05 13:12:24 +07:00
Dai Ha f84824ee29 fleetd #362 review fix: compose skill-seeding excludes with the operator's own excludesFile
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m54s
core.excludesFile is single-valued, so pointing it at fleetd's own seeded-skill exclude file
with --replace-all at worktree scope was SHADOWING whatever excludesFile the worktree already
resolved (an operator's global config, most commonly) instead of adding to it. This repo's own
.gitignore does not ignore target/ — only an operator's global excludesFile does — so every
worker's `mvn clean install` would make target/ show up as untracked, and CB-576's deliberately
untracked-inclusive hasUncommitted would then read every such worktree as dirty forever, so it
is never cleaned up.

excludeSeededSkillsFromGitStatus now reads whatever core.excludesFile resolves to BEFORE writing
anything (falling back to git's own $XDG_CONFIG_HOME/git/ignore default when the key is unset
entirely, per gitignore(5)), and writes that content into fleetd's own exclude file ahead of the
seeded skill patterns, so every operator-configured pattern keeps applying inside the seeded
worktree. Proven with a new test, seedSkillsComposesWithAnAlreadyEffectiveGlobalExcludesFile,
which isolates a synthetic "operator's global config" via a new gitEnv test seam on GitWorktrees
(GIT_CONFIG_GLOBAL pointed at a throwaway temp file, never the real machine's config) and drives
the real add() path end to end.

Also documents (FleetConfig javadoc + fleetd.example.yaml) that memberSkills copies every
non-hidden subdirectory of its source wholesale, with no per-file allowlist.
2026-09-05 13:10:27 +07:00
Dai Ha c4d40fbc2b fleetd #361 review: fix false-negative absent, uncancelled probes, and a throw contract gap
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m45s
Three findings from review of #364, fixed on the same branch:

1. LeadChannel.MailboxState.absent() was returned both for a genuinely
   absent mailbox AND for "the probe could not determine anything"
   (timeout, unreachable broker, other declare failure) -- exactly the
   overstatement #361 exists to fix, one level down. MailboxState now
   carries a Presence enum (EXISTS/ABSENT/UNKNOWN) with exists()/known()
   accessors; LeadMailbox.inspect classifies a real AMQP 404 (measured
   against a live broker, not assumed: an IOException wrapping a
   ShutdownSignalException whose Channel.Close reply code is 404) as
   ABSENT and everything else as UNKNOWN. fleet_list's mailbox/peer rows
   now render a "status" of exists/absent/unknown and only include
   pending/consumers when status is "exists", so an unresolved self- or
   peer-probe can never render as a measured zero.

2. FleetMcp.probe's get(timeoutMs) left a timed-out inspect() task
   running forever on its own virtual thread, holding the AMQP channel
   it had already opened -- against a hung (not down) broker this would
   orphan one channel per fleet_list call until the connection's
   channel-max was exhausted, breaking publish() too. probe() now holds
   the Future and calls cancel(true) on timeout/failure so the orphaned
   task is interrupted instead of abandoned, and now returns
   MailboxState.unknown() (never absent()) on timeout/exception.

3. LeadMailbox.inspect only caught IOException, but createChannel() on
   an already-closed connection throws AlreadyClosedException, an
   unchecked RuntimeException (measured against a live broker) -- so it
   could escape the "never throws" contract. Both places in inspect now
   also catch RuntimeException and report unknown().

Tests: MailboxState.exists()/absent()/unknown() call sites updated
across FleetMcpTest/FleetMcpLeadCoordTest; new hermetic tests cover the
tri-state fleet_list rendering (self-probe unknown, a peer that's
absent vs. one that's unknown) and probe cancellation (a LeadChannel
fake that blocks until interrupted, proving probe() doesn't just give
up on it); new @Tag("contract") LeadMailboxTest cases pin the real
exception shapes for both the 404 and the already-closed-connection
paths and prove inspect() reports unknown (never throws) when the
connection is already closed.
2026-09-05 13:01:48 +07:00
Dai Ha 7c684e40d3 fleetd #362 (item 3): seed .claude/skills/ into provisioned worktrees
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 1m46s
Add memberSkills: <dir> to FleetConfig. GitWorktrees#add copies each
skill folder from that directory into <worktree>/.claude/skills/ so a
member spawned against ANY repo — not only one that already ships its
own skills — can load a bridge skill (e.g. implementer). A skill the
target repo already carries is never overwritten.

Every seeded path is hidden from `git status` in that worktree ONLY,
via a --worktree-scoped core.excludesFile pointing at a file under the
worktree's own private git dir (outside the working tree, so it can
never be committed) — not the shared .git/info/exclude, which a linked
worktree resolves to the repo's common git dir and would otherwise leak
visibility changes into the primary checkout and every sibling
worktree. Proven with a real `git status --porcelain` in
GitWorktreesTest, not by reasoning.

Seeding is best-effort like the existing overlayParity/isolateToolSurface
steps: a missing/unreadable source or a copy/exclude failure is logged
and skipped, never fails the spawn. memberSkills is triaged as a
DEFERRED config key in ConfigRef (baked once into GitWorktrees at
startup, like worktreeGroup), with its own changedDeferredKeys branch
and coverage-test entries.
2026-09-05 12:54:01 +07:00
ltms 6938f52155 Merge #363: make the plugin visible, and fix the drift that made it unusable
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m28s
2026-09-05 07:50:19 +02:00
Dai Ha 0df34f3220 fleetd #361: close the lead-coordination visibility gap
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m34s
Lead-to-lead AMQP coordination had a send half with tools and a receive
half without. This closes three blind spots:

- LeadChannel gains inspect(coordId) -> MailboxState(exists, pending,
  consumers), implemented in LeadMailbox with a throwaway probe channel
  (never the long-lived publish/consume channels) so a passive-declare
  404 on a missing queue can never take down publish() on the same
  instance.
- FleetConfig.Coordinator gains peers: List<String> (defaults to empty,
  blank entries dropped) so a daemon can declare which peer coord-ids
  it expects to reach.
- fleet_list reports coordination state via a new CoordinationSource
  (own coord-id, own mailbox state, held messages as msgId/from/preview
  only, and one row per configured peer with reachability/pending/
  consumers), following the existing OutageSource/QuarantineSource
  "Source record with none()" idiom instead of growing listFleet's
  overload chain by another positional parameter. Every peer probe is
  bounded by a 1.5s timeout on a virtual-thread pool and degrades to
  absent rather than ever slowing or failing fleet_list.
- fleet_send{coordId}'s success text now says "durably confirmed by the
  broker" instead of "delivered", and warns (while still reporting
  success) when the target mailbox has zero consumers attached.

Tests: hermetic unit tests for exists/absent/zero-consumer/old-config-
no-peers-key/new fleet_list shape using FakeLeadChannel, plus a
@Tag("contract") LeadMailboxTest.inspectingAMissingMailboxNeverBreaks
PublishOnTheSameInstance proving the invariant against a real broker.
2026-09-05 12:44:29 +07:00
Dai Ha 457458437f #362: make the plugin visible, and fix the drift that made it unusable
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m31s
CB-527 shipped a Claude Code plugin and a marketplace in this repo. Nothing in
CLAUDE.md or docs/ ever named it, so a later session planned the same feature
from scratch. The wiki Features entry existed and was correct, but wiki/ is a
submodule whose pointer is never advanced, so no session reads it.

Visibility:
- CLAUDE.md addendum now names plugin/ and both structural limits, so every
  session sees it. This is the change that stops the rebuild happening again.
- wiki/11-Features.md records the rename and why the entry alone was not enough.

Drift (each measured against the code, not assumed):
- mount name fleetd -> fleet, matching PeerLauncher.MCP_MOUNT_NAME. The old name
  gave a lead with both a project .mcp.json and the plugin two mounts of one
  daemon and a duplicated fleet_* tool set.
- url is now ${FLEETD_MCP_URL} instead of a hardcoded address, so one plugin can
  serve hosts running the daemon on different ports. Plain ${VAR}, the form
  kb-alms proves works here; ${VAR:-default} is untested and not used.
- plugin claude-bridge -> fleet, marketplace claude-bridge -> fleetd, version
  0.2.0. Breaking for a 0.1.0 install: mcp__fleetd__* becomes mcp__fleet__*.
- README install path ltms/claude-bridge -> the fleet/fleetd remote.
- the setup skill's §5 told operators to pin primary.terminal:. CB-579 replaced
  that with fleet.leaders.*.tab. Replaced, with the duplicate-tab warning (#359).

Scope: the plugin is lead-side only, and cannot be otherwise. The launcher adds
--agent only when <worktree>/.claude/agents/<role>.md exists in the member's own
tree (ClaudeCodeLauncher.java:371,391), and a member's CLAUDE_CONFIG_DIR points
at its profile's config dir (ClaudeCodeLauncher.java:285), so a member never
reads the operator's plugin store. On this Mac all four Claude profiles set
configDir, and the four ccs instances hold four separate copies of the plugin
store -- same md5, different inodes. Seeding member skills through the worktree
is #362 scope item 3, implemented separately.

Note for anyone verifying a plugin: `claude plugin validate` does NOT read
.mcp.json. Replacing it with `{ this is not json at all` still passes, exit 0.

Refs #362, #359
2026-09-05 12:42:20 +07:00
Dai Ha 3759c41f99 Merge #354: the redeploy health gate classifies AMQP errors instead of counting them
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m36s
The gate counted ERROR lines since RESTART_MARK. On a laptop that idle-sleeps after one
minute on battery that meant 6 ERROR lines for an AMQP link that recovered every time,
and a gate that cries wolf is a gate nobody reads.

It now reports three states: no errors; only errors proven to have recovered (quiet, and
the gate passes); anything else (the old warning, unchanged). Attribution is per
connection, using the names #356 put into the log -- a lead-mailbox recovery can no
longer clear an unrecovered reply-inbox reset. A candidate carrying neither name is
unattributable and stays LOUD.

Two earlier rounds were rejected. Round 1 was inert: it matched nothing in the real log,
because the layout abbreviates the logger and 'Connection reset' sits in the stack trace,
not on the ERROR line -- my brief had pointed the worker at fleetd.out, which is untracked
and so absent from its worktree. Round 2 was correct and honest but could not attribute
anything, which is what motivated #356.

Verified on merge beyond the worker's own mutations:
 - ran the classifier against the REAL log, which is still in the pre-#356 format: 6 total,
   0 recovered, 6 unexplained. Old-format lines carry no connection name, so they stay loud
   -- the safe direction, on genuine data rather than a fixture.
 - adversarial fixture the worker did not write: a lead recovery BEFORE any failure banks
   no credit; 2 inbox resets with 1 recovery leaves 1 unexplained; a non-AMQP ERROR stays
   loud. total=3 recovered=1 unexplained=2, as intended.
 - RESTART_MARK still anchors the scanned region.

Caveat carried from the PR: the patterns are source-derived. The daemon has not been
redeployed, so they are not yet confirmed against a live log.
2026-09-05 06:09:29 +07:00
Dai Ha 09159f2857 Classify named AMQP recovery errors
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m16s
2026-09-05 06:05:52 +07:00
Dai Ha 29cd1194c2 Merge remote-tracking branch 'origin/main' into worker/errscan-bed2ca-2 2026-09-05 06:02:29 +07:00
Dai Ha 815e8f8b23 Merge #356: name the AMQP connection in its own log lines
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m38s
Both connections were already named at newConnection() -- 'fleetd-reply-inbox' and
'fleetd-lead-mailbox' -- and neither name ever reached the log: 0 occurrences in
fleetd.out, and both connections logged under the same thread name
'[AMQP Connection 10.10.20.13:5672]'. So when one of the two died and never came
back, the log could not say which.

AmqpConnectionFailureLogger extends DefaultExceptionHandler and overrides only the
protected log(String, Throwable) sink that every handle* method calls virtually, so
the identity is added without changing any handler action.

My brief caused a defect here and the correction is the interesting part. I told the
worker the client 'currently uses ForgivingExceptionHandler', read off the log line
c.r.c.i.ForgivingExceptionHandler -- which names where the LOGGER FIELD is declared,
not the instance's class. javap on the jar shows ConnectionFactory's constructor does
'new DefaultExceptionHandler', and DefaultExceptionHandler extends StrictExceptionHandler
extends ForgivingExceptionHandler. The first version therefore extended the base and
silently dropped strict channel-closing on four listener/consumer paths. Now pinned by
a type assertion on both factories plus a behavioural test that handleConsumerException
still closes the channel once.

Verified on merge with a mutation the worker did not run: it mutated the parent class,
so I mutated the copied private-static isSocketClosedOrConnectionReset in the DANGEROUS
direction (always true => every failure logs at WARN and vanishes from the redeploy
gate's ERROR count). Caught: 'inbox failure line ==> expected: <ERROR> but was: <WARN>'.

Merged main in first; the auto-merge compiled. 1379 green, unpiped.
2026-09-05 06:01:30 +07:00
Dai Ha 1e60ac0745 merge main for verification 2026-09-05 05:59:22 +07:00
Dai Ha 650a4c146b Merge #357: a FleetConfig component dropped by withDefaults() now fails the build
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m44s
Test-only. FleetConfig.java itself is unchanged.

The hazard is the back-compat constructor ladder (21/20/18/17/16/15/14 alongside the
22-arg canonical). Add a component and leave withDefaults()'s call at the old arity and
it binds to a back-compat constructor: it compiles, the suite passes, and the new key is
silently defaulted away on every load().

Verified on merge with a mutation the worker did not run: I made withDefaults() issue a
21-arg call, reproducing the real binding rather than an explicit null. It compiled, and
the guard failed by name -- 'memberLoginShell: ... a component silently dropped by
withDefaults(), the shape of the defect this test exists to catch'.

Exclusion list is empty and its size is pinned, so a future exemption must touch an
assertion rather than grow quietly.
2026-09-05 05:55:22 +07:00
Dai Ha 23f299e105 Preserve strict AMQP exception handling
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m46s
2026-09-05 05:53:44 +07:00
Dai Ha dbf6fef0e9 config: guard FleetConfig.withDefaults() against silently dropping a component
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m53s
Adding a component to FleetConfig follows an established pattern: the
record grows by one arg, and a back-compat constructor is added at the
OLD arity so existing callers keep compiling. That back-compat
constructor also silently captures withDefaults()'s own literal-arity
'return new FleetConfig(...)' call the next time this happens, since
that call is now a legal overload match too. It compiles, every other
test passes, and the new component is defaulted away on every load().
This is not hypothetical - it happened live while building the (now
parked) idle-sleep-guard PR, caught only because that branch's own new
tests asserted on the new field.

Add a reflective test that builds a FleetConfig through the true
canonical constructor (resolved by record-component types, not arg
count - the same pattern ConfigRefTopLevelReportingCoverageTest already
uses in this file) with a real, non-null value in every component, runs
the real withDefaults(), and asserts every value survives unchanged.
Never hardcodes the arity - it enumerates
FleetConfig.class.getRecordComponents() - so it keeps working as the
record grows. No back-compat constructor is touched or removed.
2026-09-05 05:51:59 +07:00
Dai Ha d292522d00 Name AMQP connection failure logs
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 1m29s
2026-09-05 05:45:37 +07:00
Dai Ha 0241e0d3a8 Keep unattributed AMQP errors loud
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m30s
2026-09-05 05:36:53 +07:00
Dai Ha e4973eb8a4 Classify recovered AMQP redeploy errors
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m51s
2026-09-05 05:29:53 +07:00
Dai Ha b6b88c5f1c #334: pin the fresh-owner gate on ask()'s timeout teardown
CI / build (push) Successful in 2m7s
CI / contract (push) Successful in 34m37s
2026-09-04 17:12:49 +07:00
Dai Ha 86dddfe240 Merge #334: ask()'s timeout closes the turn before it forgets the task mapping 2026-09-04 17:08:40 +07:00
Dai Ha 0d5944af63 fleetd #334: close ask()'s turn before forgetting its Task, closing the last stranding window
CI / build (pull_request) Successful in 2m14s
CI / contract (pull_request) Successful in 19m14s
ask()'s TimeoutException catch used to run clearAsyncQuestion(turnId, true) -- forgetting the
Task's asyncTasksByTurn mapping -- before rendezvous.closeAsk(turnId) ran in the shared finally.
Between those two calls the ask was still "answerable" (askSession(turnId) non-null) but the Task
mapping was already gone, so a racing answer() call found task == null, skipped
finishAsyncTask, and stranded the async ticket at PENDING even though answer() itself reported a
result. #329 fixed one step of this same race; this closes the remaining one.

The fix reorders the fresh owner's teardown: closeAsk runs first, then markAskTimedOut and
clearAsyncQuestion. A racing answer() call now either sees the ask still open (and the Task
mapping guaranteed intact) or sees it already closed (STALE_TURN, before it ever reaches
asyncTasksByTurn). It also gates the whole block by ticket.fresh(), matching the invariant the
finally block already states ("only the fresh owner tears down the shared turn") -- a duplicate
coalesced ask() timing out no longer forgets bookkeeping the fresh owner still needs.

Adds aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket, which pins the exact window
with a new test-only hook (askTimeoutRaceHookForTest) and proves both invariants: a late answer()
racing the timeout sees STALE_TURN, and the async ticket still resolves DONE from the worker's
real reply. Reverting the reorder (verified locally, not committed) makes this test fail with
"expected STALE_TURN but was TIMED_OUT_WORKING".
2026-09-04 17:06:15 +07:00
Dai Ha 4a5030a5c6 #348: drop a chrome skip that cannot fire, and pin the live pattern shape
CI / contract (push) Successful in 2m10s
CI / build (push) Successful in 2m12s
2026-09-04 16:57:36 +07:00
Dai Ha c1c8794c48 Merge #348: a member's prose about a usage limit no longer quarantines a credential 2026-09-04 16:53:19 +07:00
Dai Ha 65a78932c1 Merge #335: a per-task cleanup throw in abandon() no longer strands the tasks behind it
CI / contract (push) Successful in 1m24s
CI / build (push) Successful in 2m8s
2026-09-04 16:44:04 +07:00
Dai Ha f429ca1a50 Avoid exhaustion cooldown for member prose 2026-09-04 16:43:16 +07:00
Dai Ha 73aab3f83e Merge #342: teardown resolves the pane's real tab instead of trusting the delegate's placement config
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m47s
2026-09-04 16:38:31 +07:00
Dai Ha 887aca0183 fleetd#335: abandon()'s per-task cleanup and sendAsync's terminal hook must not swallow throws
CI / contract (pull_request) Successful in 1m2s
CI / build (pull_request) Successful in 2m21s
Site 1 (abandon()'s matching loop, reachable): the recovery/put-back branch calls
inbox.publish, which AmqpReplyInbox implements as a real broker round trip that
throws IllegalStateException on an unroutable/unconfirmed/interrupted publish.
An uncaught throw there aborted the loop, stranding every task after it in
`matching` PENDING forever. Fixed by recording each task's own future.complete()
result before any cleanup runs, then wrapping the cleanup in try/catch so one
task's failure cannot stop its siblings from getting their outcome. Reaching the
throwing branch by real timing needs a race the file's own #137 follow-up already
found unreachable through the public API, so the reproducing test uses a
test-only hook (same technique as the existing fleetd #324/#329 hooks) to inject
the throw at that exact point.

Site 2 (sendAsync's task.future.whenComplete, reachable): the returned stage is
discarded, so an uncaught throw from pushLoop.onTicketTerminal vanished with no
log line. Reproduced for real: Fleetd's shutdown hook runs messages.close()
(stops the async executor from taking new work, but does not cancel a send
already in flight) before pushLoop.close() (shuts its scheduler down
immediately) — a ticket completing in that window makes onTicketTerminal's own
scheduler.schedule(...) throw a genuine RejectedExecutionException. Fixed with a
try/catch(Throwable) plus log.error inside the whenComplete action.

Site 3 (the two `finally { asyncTasksByWaiter.remove(reply); rendezvous.close(...);
}` blocks in send() and answer()): read Rendezvous.close/closeAsk and the
ConcurrentHashMap operations behind them — both are plain map ops on a non-null
key with no user-overridable code, so neither can throw. Left unchanged; not a
defect.

Mutation-proven: reverting either fix reproduces the failure it exists to catch
— removing site 1's try/catch aborts abandon() with the injected exception
(MessageServiceTest#aPerTaskCleanupFailureDoesNotStrandTheRemainingMatchingTasks
errors); removing site 2's try/catch leaves the RejectedExecutionException
unlogged (MessageServiceTest#aTicketTerminalPushFailureDoesNotVanishSilently
fails its log assertion). Full suite: mvn clean install, Tests run: 1365,
Failures: 0, Errors: 0, BUILD SUCCESS.
2026-09-04 16:38:14 +07:00
Dai Ha 3fd23ecafa fleetd #342: base tab-cleanup teardown on the pane's real placement, not a delegate's static config
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Failing after 1m38s
HerdrPeerLauncher.stop() used to gate spaces.locatePane() on usesTabPlacement(),
which reads the delegate's OWN configured profiles. When CompositePeerLauncher's
single-daemon stop() shortcut hands a pane to a delegate that never spawned it
(spawnedBy empty after a daemon restart, herdrDaemonCount()==1), that delegate's
placement config says nothing true about how the pane was actually placed, and a
dedicated tab could be skipped and leaked.

Resolve the tab unconditionally instead — WorkspaceControl#locatePane already
tolerates a missing pane by returning null — and let the existing single-occupant
check (tabPaneCount()==1) be the only thing that decides whether to close it, same
as it already protects a shared tab regardless of declared placement.

Adds a mixed-placement CompositePeerLauncherTest (every existing stop-fallback test
configured both adapters as tab placement, so the mis-routing never showed) and
updates FleetAppTest#stopWorkerInPanePlacementClosesOnlyThePane, whose old
assertion (no pane.get on pane placement) documented exactly the skip this fix
removes.
2026-09-04 16:36:11 +07:00
Dai Ha f379847942 Merge #345: the timeout path's use of Cancellation.DELIVERED is now pinned
CI / contract (push) Successful in 47s
CI / build (push) Successful in 2m8s
2026-09-04 16:32:45 +07:00
Dai Ha ea12107497 fleetd #345: test timeout cancellation race
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 2m3s
2026-09-04 16:30:46 +07:00
Dai Ha 591df91de1 Merge #337: the deferred key set proves its own reporting coverage too
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m57s
2026-09-04 16:16:20 +07:00
Dai Ha 6a814176f0 #339: the start-of-line check must skip terminal chrome
CI / build (push) Failing after 1m30s
CI / contract (push) Successful in 1m54s
#339 stopped a member's own prose about an error from recording a credential
outage, by requiring the pattern at the start of its matched line. A bare
lookingAt also rejected a genuine error line rendered as

    | 503 Service Unavailable: upstream credential rejected

The send still failed, but the outage was never recorded. That is the false
negative #339's own invariant 3 named as worse than the false positive it set
out to fix: an unrecorded outage leaves the fleet spawning into a dead
credential.

Measured with a throwaway probe on the raw-scrape path, whose own comment says
to expect leading chrome there: kind=FAILED, sinkNotified=0.

startsWithBackendError now skips a leading run of non-letter, non-digit
characters before the check. That keeps #339's intent: prose still does not
match, because there the pattern sits after words rather than after chrome.
The worker's own prose test still passes.

Mutation: restoring the bare lookingAt fails the new test.
2026-09-04 16:13:19 +07:00
Dai Ha d11d1d157c Merge #339: a backend-error text match must be at the start of its line before it records a credential outage 2026-09-04 16:08:46 +07:00
Dai Ha d057d56156 #341: pin the noise control, and the reverse policy order
The fix reported every distinct unprotected name, but nothing held it there.

Mutation: replacing .filter(unprotectedGapNamesWarned::add) with a filter that
adds and always returns true - so every name is logged on every spawn - left
all 1358 tests green. The Set behaved; nothing proved this class used it as a
guard rather than as a record.

Two tests added:

- theSameUnprotectedNameIsWarnedAboutOnlyOnceAcrossSpawns pins invariant 1, the
  noise control. It now fails on that mutation, showing both duplicate WARNs.
- anAllowListWarnDoesNotSuppressALaterDenyByDefaultWarnForADifferentName covers
  the reverse policy order. The defect was found going deny-by-default then
  allow-list; a guard fixed in one direction is not fixed in the other.
2026-09-04 16:07:50 +07:00
Dai Ha d703ce1313 #337: extend ConfigRefTopLevelReportingCoverageTest to DEFERRED_KEYS
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m27s
ConfigRefTopLevelReportingCoverageTest (added by #333) proved every COLD_KEYS
and SPLIT_KEYS member has a real comparison behind it, but left
DEFERRED_TOP_LEVEL_KEYS unexercised. Re-measured by mutation (drop each
key's branch from changedDeferredKeys, run the suite, restore): 6 of the 11
deferred keys had no behavioural test naming them — guard, leadHeartbeat,
worktreeRoot, spawnReadyTimeoutMs, spawnReadyPollMs, quarantineCooldownSeconds
— which corrects the issue's own guessed list in two ways: lifecycle is
actually covered (ConfigRefTest.aDeferredChangeIsAppliedAndReported), and
worktreeRoot was missing from the issue's list entirely.

Promoted the test-side DEFERRED_TOP_LEVEL_KEYS copy into ConfigRef.DEFERRED_KEYS
(package-private, alongside COLD_KEYS/SPLIT_KEYS) so the reflective test reads
the same set changedDeferredKeys is compared against, and made
changedDeferredKeys package-private so the test can call it directly. Every
DEFERRED_KEYS component turned out to be a scalar or a simple record, so no
exclusion set was needed.

Mutation proof: dropping guard's branch from changedDeferredKeys leaves the
whole suite green except the new
everyDeferredKeyIsActuallyReportedByChangedDeferredKeys test, which fails
naming guard exactly.
2026-09-04 16:06:09 +07:00
Dai Ha b32a30fd47 Merge #341: warn once per distinct unprotected credential name, not once per launcher 2026-09-04 16:02:51 +07:00
Dai Ha 464dbc0930 fleetd#341: a per-name guard so a later spawn's different unprotected name still warns
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Successful in 1m29s
unprotectedGapLogged was one AtomicBoolean guarding two WARN branches in
logCredentialGap that name different env var names (the allow-list
keptByDerivedList branch, and warnGapUnprotected's deny-by-default /
non-zsh-fallback branch). memberCredentials is a live, re-read-per-spawn
supplier, so between two spawns a policy reload can change which names are
in the gap: spawn 1 warns about name A and trips the shared flag, and
spawn 2's gap containing a different name B never gets its WARN.

Replace the AtomicBoolean with unprotectedGapNamesWarned, a
ConcurrentHashMap-backed Set<String> guard keyed per name (same shape as
OpenCodeLauncher.modelCheckSkippedWarned), so each distinct credential-shaped
name is warned about exactly once, ever, regardless of which branch or
which spawn first reports it. allowListGapLogged (the separate INFO guard,
#192) is untouched. Neither WARN's wording changed.
2026-09-04 15:59:43 +07:00
Dai Ha a8cadd9150 Merge #338: a timed-out queued send cancels its message instead of leaving it to be delivered later
CI / contract (push) Successful in 1m23s
CI / build (push) Successful in 2m12s
2026-09-04 15:57:39 +07:00
Dai Ha 57b8c0b56d fleetd #339: guard backend error sink
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m50s
2026-09-04 15:57:18 +07:00
Dai Ha 147f50c19e #338: cancel timed-out queued deliveries
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Failing after 2m11s
2026-09-04 15:56:10 +07:00
Dai Ha eee4d576a2 #333: the Cold doc bullet listed four keys, COLD_KEYS has five
CI / build (push) Successful in 1m52s
CI / contract (push) Successful in 2m8s
memberHerdrSocket was missing from the prose. The #333 worker spotted it and
correctly left it alone as outside its scope.

Fixed by pointing the bullet at COLD_KEYS instead of re-listing its contents,
so the prose and the set cannot drift apart a second time.
2026-09-04 15:43:39 +07:00
Dai Ha b4f9d7f53a Merge #333: fleet: is a split key, and split membership now proves a reporting branch exists 2026-09-04 15:39:02 +07:00
Dai Ha 3aca53b967 fleetd#333: fleet.leaders is split too, and split membership now proves reporting exists
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Failing after 1m59s
F1: fleet: was sitting in ConfigRefTopLevelCoverageTest's HOT_EXCLUDED_TOP_LEVEL_KEYS
escape hatch, even though fleet.leaders is read only at startup (LeadTabScanner's
identity map, LeadLauncher.ensureLeads) while the rest of fleet: (role pools,
charters, tabLabel) is live. A reload changing only fleet.leaders reported a bare
"config reloaded" -- the operator edits a lead's tab: label, sees the reload
succeed, and the pane keeps resolving as a worker. Moved fleet into
ConfigRef.SPLIT_KEYS; changedSplitKeys now compares fleet.leaders specifically
(not the whole Fleet record, which would over-claim "restart" for a tabLabel-only
change) and names both halves in the message.

F2: membership in SPLIT_KEYS/COLD_KEYS never proved a matching branch existed in
changedSplitKeys/changedColdKeys -- measured by dropping the coordinator branch
while leaving "coordinator" in SPLIT_KEYS: both ConfigRefTopLevelCoverageTest and
the in-method "kept in step" assert stayed green. Added
ConfigRefTopLevelReportingCoverageTest, the ConfigRefProfileCoverageTest mechanism
one level up: reflection-built FleetConfig pairs that differ in exactly one
top-level component, calling the real (now package-private) changedColdKeys/
changedSplitKeys to prove each COLD_KEYS/SPLIT_KEYS member is actually reported.
Scoped to split+cold, not deferred -- see the new test's javadoc for why and what
that leaves open.

Both findings carry a behavioural test in ConfigRefTest plus a mutation proof
(revert -> real failure -> restore) recorded in the PR description.
2026-09-04 15:36:40 +07:00
Dai Ha 4aa1fae296 #329: a null task in answer() is not only "never an async ticket"
CI / contract (push) Successful in 59s
CI / build (push) Successful in 2m26s
The comment that landed with #329 said a null task means the turn was never
an async ticket. That is wrong, and it makes the guard read as complete.

A genuine async ticket also reaches answer() with task == null. ask() runs
clearAsyncQuestion(turnId, true) in its catch block, which drops the
asyncTasksByTurn entry, while rendezvous.closeAsk(turnId) runs later, in its
finally. Between the two the ask is still answerable and the map entry is
already gone, so answer()'s lookup returns null and the ticket is stranded.

Measured with a throwaway probe firing only that first half: answer() reported
REPLIED while the ticket stayed PENDING with a null reply. The probe used
forgetTurnForTest, so it omits markAskTimedOut; that cannot change the outcome,
because askTimedOut is read only by askAnsweredAsyncTasks, which reply() never
reaches while answer()'s own waiter is live.

Open as fleetd #334. The comment now says so.
2026-09-04 15:33:57 +07:00
Dai Ha 0c865032f9 Merge #329: log an exception thrown after a ticket resolves, complete the ticket from the task answer() already holds, and read orphan.turnId once 2026-09-04 15:26:45 +07:00
Dai Ha ea41bbf6b9 fleetd#329: fix silent async-ticket bugs in MessageService (F1/F2/F3)
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m28s
F2 (sendAsync executor catch): log when completeExceptionally returns
false, so an exception thrown after finishAsyncTask already completed
the ticket's future is no longer silently lost.

F1 (answer()'s stranded async ticket): reuse the Task reference answer()
already looked up before rendezvous.answerAsk(), instead of a second
asyncTasksByTurn lookup by turnId in finishAsyncTask. The second lookup
raced ask()'s unlocked timeout cleanup, which could forget turnId first
and leave the ticket stuck PENDING even though answer() itself returned
REPLIED. The #282 chained-ask guard is unaffected: it is still keyed on
result.outcome() == QUESTION, not on this lookup. Removed the now-unused
finishAsyncTask(String, Reply) overload.

F3 (reply()'s orphan recovery path): read orphan.turnId once instead of
twice, closing the same double-read shape fleetd #324 fixed in
finishAsyncTask.

Each fix has its own test plus a test-only race hook (mirroring #324's
finishAsyncTaskRaceHook) to force the exact interleaving deterministically.
Mutation-tested each fix by reverting it, confirming the real failure
(swallowed exception / PENDING ticket / NullPointerException), then
restoring it.

mvn clean install: Tests run: 1345, Failures: 0, Errors: 0, Skipped: 0,
BUILD SUCCESS.
2026-09-04 15:19:38 +07:00
Dai Ha 7b918c51ff Merge #330: a fourth reload class for split keys, and a top-level coverage checker
CI / build (push) Successful in 1m46s
CI / contract (push) Successful in 1m48s
2026-09-04 15:19:14 +07:00
Dai Ha 554395b104 fleetd#330: split reload class for health/coordinator + top-level coverage
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m17s
Unit 1: ConfigRef gets a fourth reload class, `split`, for keys read both
off the startup snapshot and live off config.get() at different sites
(health:, coordinator:). A split change is accepted (Outcome.applied()
stays true) and reported by name, naming which half is live and which
needs a restart, via a new Outcome.split() field kept separate from
deferred() since the two carry different guarantees for any caller that
branches on them, not just prose in summary(). Class doc updated: four
classes now, denominator note no longer calls health/coordinator
undecided.

Unit 2: ConfigRefTopLevelCoverageTest enumerates FleetConfig's 22
top-level record components and requires each to sit in exactly one of
COLD_KEYS, a pinned "compared in changedDeferredKeys" set, SPLIT_KEYS, or
a pinned hot-exclusion escape hatch — printing its own denominator and
pinning the escape hatch's exact contents the way #323 asked for.

Deviates from the issue's starting values by one key: `profiles` moves
from the suggested Hot bucket into the deferred bucket, because
changedDeferredKeys demonstrably compares it (add/remove and launch
settings), and citing "read live off the config supplier" for the whole
key would be false — most Profile fields are not read live, only
weight/maxLoad/credentialId are (and those are already covered by
ConfigRefProfileCoverageTest). Cold=5, split=2, deferred=11, hot=4,
total=22 — verified against the record and against ConfigRef's code, not
copied from the issue.
2026-09-04 15:16:45 +07:00
Dai Ha 823976c1b5 #326: state primary's real consequence, and write down the denominator
CI / contract (push) Successful in 56s
CI / build (push) Successful in 2m5s
The merged javadoc said a changed primary.terminal leaves a lead 'unresolved as
primary until a restart'. That over-claims. CB-532 made the pin deprecated:
identity comes from leaders:/leadScan:, and Fleetd.java:511 warns about the pin
at startup. A changed pin still needs a restart, but for the fallback nudge
destination, the deprecated identity path, and pushReminders/pushBackoffMs -
not for a lead that uses leaders:.

Also record what I measured. FleetConfig has 22 top-level components; four are
named nowhere in ConfigRef. memberCredentials and memberLoginShell are hot and
correctly absent (both read live off config.get() at spawn). health and
coordinator are undecided, not hot. 'Absent' looks the same for both kinds, and
twice now the forgotten kind hid among the correct kind.
2026-09-04 14:58:07 +07:00
Dai Ha c8388a7f92 Merge #326: primary and configReload are deferred keys, so a reload says a restart is needed 2026-09-04 14:53:50 +07:00
Dai Ha 02e6aef98c Merge #324: read task.turnId once in finishAsyncTask, so a concurrent clear cannot make the removal key null
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m29s
2026-09-04 14:44:42 +07:00
Dai Ha 6d493bc7bb fleetd#326: classify primary and configReload as deferred top-level keys
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Failing after 1m44s
ConfigRef.changedDeferredKeys only classified seven top-level FleetConfig
keys (#323 fixed the profile side). Two more keys are read only off the
startup snapshot and were missing:

- primary: Fleetd.java:506/519/520 feed PrimaryRegistry and ReplyPushLoop
  at construction; neither is rebuilt on reload.
- configReload: Fleetd.java:679-680 decide once at startup whether to
  build a ConfigWatcher at all, and with what interval; the watcher that
  would apply a later change is itself built once, so it is deferred
  (not cold — no already-open resource goes inconsistent, a running
  watcher just keeps its original settings).

health and coordinator are deliberately left unclassified: both are read
both off the startup snapshot AND live off the config supplier at a
second call site, so no single bucket is correct for either — see the
PR body for the options writeup and the coordinator.uriEnv exposure
question the issue asked to be answered.

Each fix is proven with a failing-first test in ConfigRefTest and a
revert-quote-restore mutation check (see PR body for the transcripts).
2026-09-04 14:42:24 +07:00
Dai Ha e5cb51a90e #324: read task.turnId once in finishAsyncTask to stop an NPE from ask()'s unlocked forgetting
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 2m1s
answer() holds sessionLocks while finishAsyncTask reads the volatile Task.turnId twice — once to
check it is non-null, once as the ConcurrentHashMap.remove key. ask()'s own timeout path mutates
the same field with no lock, via clearAsyncQuestion(turnId, true). volatile makes each read fresh
but not the pair atomic, so the field can go null between the two reads and remove(null, task)
throws NullPointerException on the lead's own answer() call, even though the reply already
completed on the line above.

Capture task.turnId into a local once and use that for both the check and the removal.

Added a package-private test seam (finishAsyncTaskRaceHook + forgetTurnForTest) so a test can force
the exact interleaving deterministically, by running the identical clearAsyncQuestion(turnId, true)
cleanup ask() uses, at the point between finishAsyncTask's former two reads. Both are inert (null)
in production.
2026-09-04 14:38:52 +07:00
Dai Ha e545c08082 #323: pin the exclusion set — the coverage mechanism's own escape hatch, found by mutation
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 2m10s
2026-09-04 14:32:08 +07:00
Dai Ha efa0deb9b2 Merge #323: the reload classifier proves its own coverage instead of claiming it 2026-09-04 14:29:10 +07:00
Dai Ha ca47e90c01 fleetd#323: close the reload-classifier drift with a reflection coverage test
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 2m6s
ConfigRef.sameLaunchSettings' javadoc claimed it compares every component
the launcher reads at spawn. It missed ideProjectDir, ideOpenCommand and
autoCompactWindow, and changedDeferredKeys separately missed worktreeGroup
(baked into the same GitWorktrees as worktreeRoot, Fleetd.java:251). A
reload that changed only one of those keys reported "config reloaded" with
nothing deferred, and the running daemon kept the old value.

Fix the four instances, and add ConfigRefProfileCoverageTest: it enumerates
every FleetConfig.Profile record component by reflection, mutates each one
not in the new ConfigRef.LAUNCH_SETTINGS_EXCLUDED set on a base profile,
and asserts sameLaunchSettings actually notices — so a fifth missed field
fails the build by name instead of drifting silently. It also prints its
own denominator (26 components, 23 compared, 3 excluded) per the ticket's
requirement that a checker must be able to state what it checked.

Also add the `profile` field itself to the comparison (it was neither
compared nor excluded before this fix — the coverage test surfaced it).

Rewrote the sameLaunchSettings javadoc to describe what the coverage test
actually guarantees instead of repeating the unchecked claim.
2026-09-04 14:26:09 +07:00
Dai Ha fa1f49675b Merge #318: a delivery landing after release is refused, not parked in a map nobody reads
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 2m8s
2026-09-04 14:21:29 +07:00
Dai Ha 8426c3528f #316: pin the fail-toward-preserve rule on the late re-check, found by mutation
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m42s
2026-09-04 14:20:36 +07:00
Dai Ha 65f98ba910 Merge #316: the dirty check that authorises the worktree removal is taken after the worker stops 2026-09-04 14:17:13 +07:00
Dai Ha d05205d1eb Add a hunter skill: a sweep and a diff review are different jobs with different output contracts
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m49s
2026-09-04 14:13:08 +07:00
Dai Ha 667254df47 #316: re-check worktree dirtiness after the pane stops, before removing it
CI / contract (pull_request) Successful in 39s
CI / build (pull_request) Successful in 1m46s
SessionManager.releaseRemoved read hasUncommitted() once, while the worker
could still write, then used that stale boolean after launcher.stop() to
authorise `git worktree remove --force`. The same stale read also gated
trySnapshot, so a worker that wrote between the read and the stop lost its
work with neither a preserve nor a snapshot.

Add a second, best-effort hasUncommitted read immediately before the
removal, taken only on the path that is actually about to delete something
(never on a release that already decided to preserve, and never for
SHUTDOWN, which preserves unconditionally). If the tree is now dirty,
preserve it and attempt a fresh snapshot, since the original snapshot never
ran when the pre-stop read said clean. A failing re-check also preserves,
matching the existing CB-581 fail-safe rule.
2026-09-04 14:12:49 +07:00
Dai Ha 2926cd1784 #318: release() no longer strands a delivery that lands while it is running
CI / build (pull_request) Failing after 1m59s
CI / contract (pull_request) Successful in 2m14s
AmqpReplyInbox.release() used held.remove(target) then iterated the old
map. A delivery landing on the consumer work-pool thread after the
remove (basicCancel does not flush one already handed to that pool) hit
deliverCallback's computeIfAbsent, found the key gone, and created a
brand-new map release() never looks at again — delivered-but-unacked
forever, never requeued, never redelivered (#298 only closed the
"already in held when release runs" case).

Fix: release() swaps in a RELEASED tombstone via held.compute(...)
instead of held.remove(...). ConcurrentHashMap serializes compute/
computeIfAbsent calls for the same key against each other, so whichever
of release() and a concurrent deliverCallback runs first is fully
visible to the other — no gap. deliverCallback checks for the
tombstone and nacks-with-requeue instead of recreating a map; peek/ack
treat it as empty; own() clears a stale tombstone so a target is never
poisoned if its id is ever reused (the issue's own text says id reuse
doesn't happen, but the tombstone would otherwise sit in `held` forever
either way).

New test AmqpReplyInboxReleaseRaceTest forces the actual interleaving
with a latch (blocks release() inside its nack loop, which is only
reachable after the tombstone swap, then fires a concurrent delivery)
rather than a sequential call — a sequential test would not have caught
this, since #298's own contract test forces settlement before release()
runs. Mutation-tested: reverting the fix makes this test fail with
"expected: <2> but was: <1>" (m1 never nacked); restored after
confirming that failure.
2026-09-04 14:10:21 +07:00
Dai Ha c801851c66 Correct the isLoopback javadoc: after #305 narrowing this range refuses a caller, it does not promote one
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m47s
2026-09-04 14:10:10 +07:00
Dai Ha de70aa38f1 Merge #317: an unresolved caller is refused, never promoted to primary
CI / contract (push) Successful in 1m1s
CI / build (push) Successful in 1m56s
2026-09-04 14:03:35 +07:00
Dai Ha 77ad88631b Merge #315: the fixed placement policy honours the retry loop's unreachable set
CI / build (push) Successful in 1m22s
CI / contract (push) Successful in 1m55s
2026-09-04 13:57:24 +07:00
Dai Ha d88017807b #315: fix self-contradicting javadoc left by the previous commit
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 2m31s
FixedPlacementPolicy's class javadoc still opened with "This ignores caps
and reachability" after the previous commit added reachability as the
fourth carve-out that is explicitly NOT ignored — caught by a shape-check
survey run against this same file as part of #315's own request ("look in
placement/ ... for the same shape: a caller/comment that documents an
expectation ... where an implementation does not meet it"). Reworded the
opening sentence: fixed still ignores caps (maxLoad) by design, but
reachability is now a narrower, per-call retry exclusion, not an ignored
concern.
2026-09-04 13:55:28 +07:00
Dai Ha 2159a5a94a #315: FixedPlacementPolicy now honors the retry loop's unreachable set
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m33s
CompositePeerLauncher.spawn retries a failed candidate on the next one and
rebuilds PlacementContext "so the policy excludes this profile" (its own
comment), but FixedPlacementPolicy.select never read ctx.unreachable(). Under
the default `fixed` placement policy (used when `placement` is unset or set
to `fixed`), every retry re-picked the same dead default and a second,
healthy, configured profile was never tried. This also covers the wiring-bug
branch (a candidate profile with no owning adapter), which hit the exact same
symptom for the same reason.

Not live on this fleet: fleetd.yaml sets placement: weighted, which already
consults ctx.unreachable() via PlacementPolicyUtil.available(). This is live
only for a deployment that leaves placement unset or sets it to fixed.

Fix is in FixedPlacementPolicy: consult ctx.unreachable() in the same two
places it already consults quarantined/coolingOff (the default check and the
fallback walk over candidates()), and add a fourth reason to the "no
candidate remains" exception. Considered fixing this in
CompositePeerLauncher's retry loop instead (break when select() returns an
already-unreachable profile), but that only fails faster on the same dead
profile — it cannot make the loop advance to a different candidate, because
only the policy decides which candidate is next. The defect is that one
policy implementation does not honor the loop's stated contract, so the fix
belongs in that policy, matching how weighted/round-robin already behave.

Also fixed: the "no reachable worker profile" exception message said
"trying N candidate(s)" where N was unreachable.size(), a count of DISTINCT
profiles (a HashSet dedupes a profile added twice), under wording that reads
as a count of attempts. Reworded to "N distinct candidate(s)" so the count
matches what is measured and the profile list that follows it.

Tests: two new failover tests next to the three existing ones in
CompositePeerLauncherTest (which all use PlacementPolicies.weighted(), which
is why this had no coverage) — one pinned to PlacementPolicies.fixed() for
the unreachable-default case, one for the wiring-bug (no adapter) case.
Mutation-proofed: reverted FixedPlacementPolicy.java, both new tests failed
with the exact bug ("no reachable worker profile available after trying 1
distinct candidate(s): a" / "...c"), then restored the fix.
2026-09-04 13:52:29 +07:00
179 changed files with 30850 additions and 1703 deletions
+4 -4
View File
@@ -1,15 +1,15 @@
{
"name": "claude-bridge",
"name": "fleetd",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the fleetd MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "claude-bridge",
"name": "fleet",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"description": "Mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.2.0",
"author": {
"name": "LTMS"
}
+23
View File
@@ -0,0 +1,23 @@
---
name: hunter
description: Sweep one assigned scope for defects and report ranked findings without changes.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You sweep the assigned package or scope for real defects. Read the full assigned scope before you
judge it. Report several ranked findings when the evidence supports them. Change nothing: do not
edit code, commit, push, or open a pull request.
You may run the build or tests to check a finding. Read the complete output and report the real
result. Do not hide failures with a pipe. State only checks you actually ran. The primary's IDE
tools are not yours. A mounted forge tool may use a blocked credential and fail by design.
Do only the assigned scope. Note anything outside it in one line and do not investigate it further.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an unclear requirement
or two defensible fixes. Do not ask about something you can decide by reading more code.
Your handoff must name the files you read, each ranked finding or `NO FINDINGS`, the checks you ran,
and any caveat for review.
The launcher provides the required bridge reply instructions for every member.
+226
View File
@@ -0,0 +1,226 @@
---
name: handover
description: Procedure for an outgoing lead to write the handover file that a fresh lead session inherits. Load this when your context is filling up and you are about to be replaced, whether you hand off by hand or fleetd does it for you. The file is the new lead's only inheritance — follow it exactly.
---
# Handover — write the file the next lead depends on
A lead session fills up its context and has to be replaced by a fresh one. The outgoing lead
writes a handover file, and the new session reads that file and carries on.
**There are two ways to hand off, and the file is the same either way.**
- **By hand.** You write the file, then tell the operator where it is. The operator starts the new
session and points it at the file. This always works.
- **With `fleet_handover`** (fleetd #480, merged 2026-09-11). You ask fleetd to do the swap: it
checks the file, clears your pane, and tells the fresh session to read it. This needs
`leadRollover:` in `fleetd.yaml`; without it every action answers a clean refusal naming
`NOT_CONFIGURED`, and you fall back to the manual path. Section 11 below is the procedure.
Nothing else in this skill changes between the two. Only who performs the swap changes.
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
through at the end of a session.
This skill is the procedure for writing it. Every rule below earned its place because a past
handover got it wrong.
## 1. Confirm you are the right session to write this
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
## 2. Every number needs a command, run in this turn
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
you write one, run the command that produces it — now, in this turn, against the live state.
Never take a number from:
- earlier in your own conversation — the state has moved since then,
- a peer lead's report — that is their measurement, not yours,
- your own memory of an earlier session.
Put the command, or its real output, next to the number. That lets the next lead re-run it and
check it still matches. If you cannot measure something yourself, say so instead of guessing:
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
## 3. Say what you measured and what you did not
Mark every claim as one of two things:
- **"I checked this myself, in the code or on this host, at `<time>`."**
- **"I did not check this myself; `<who>` reported it."**
Never present someone else's measurement as your own. This matters most for cross-host claims —
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
re-verify from here.
## 4. Record open decisions, and who owns them
List three things:
- what the operator actually asked for, in their own words where you have them,
- what is still unanswered,
- any question you decided yourself instead of asking, with your reason.
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
reopens a closed question or assumes the operator chose something they never did.
## 5. Record live hazards
List anything that will break if the next lead does the obvious thing next. This includes:
- unpushed commits or unmerged branches,
- a build, a spawn, or a redeploy still running,
- code merged to `main` but not yet redeployed to the live daemon,
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
"tricky."
## 6. Record what is explicitly not owed
List work that is finished, and work that another party has said they do not want touched. Name
who said so and when. Without this line, the next lead re-does closed work or reopens a question
a peer already declined to revisit.
## 7. Open the file with three re-measurement commands
The file's own first section must give the next lead three concrete commands to run before
acting on anything else in the file:
1. confirm role — for example `fleet_whoami`,
2. confirm the state of the working tree — for example `git status` and
`git rev-list origin/main..HEAD`,
3. read the live fleet — for example `fleet_list`.
Record what each command answered when you wrote the file, and tell the reader to run it again
rather than trust your answer. The point of this section is that the reader checks live state
before acting on any claim in the rest of the file, including yours.
## 8. Stamp the file with time and commit
Near the top of the file, write:
- the date and time you wrote it,
- the commit the tree was on (`git rev-parse HEAD`),
- whether the tree was clean (`git status`).
Without this, nobody can tell how old the file is, or which code it describes.
## 9. State plainly that the file goes stale fast
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
of one moment, not a live fact.
## 10. What to leave out
Do not include:
- narration of how the session felt, or how hard something was,
- anything the repo already records — code structure, git history, or a rule already written in
`CLAUDE.md`. Point at it instead of repeating it,
- advice that is only true for the session that is ending — a half-open terminal, a local
variable, a train of thought with no state behind it.
A handover file is a record of state and decisions. It is not a diary.
## 11. Using `fleet_handover` (only if `leadRollover:` is configured)
**Run the three steps in this order. The order is not a style choice — the wrong order is
refused.**
1. **`fleet_handover{action: "open", reason: "<why now>"}`.** It returns a `token` and the
`handoverPath` you must write to. Nothing has happened to your pane yet.
**Write to exactly that path, and do not resolve it yourself.** It is always absolute, even when
the operator configured a relative `handoverPath`: fleetd resolves a relative one against your
own workspace before it hands it to you. The daemon and your pane can run in different
directories, so a path you resolve yourself can point at a different file from the one the daemon
will check.
**Check that the path is ignored by git before you write to it (#491).** A relative
`handoverPath` resolves inside YOUR workspace, which is usually a repository — and usually not
the `fleetd` one, so an ignore rule added to `fleetd` does not protect it. Run
`grep -n handover <your workspace>/.gitignore`. No output means the file you are about to write
will show up as untracked content in that repo. The file is a snapshot of live state and must
never be committed, so tell the operator rather than committing it or silently editing their
`.gitignore`.
2. **Write the handover file at that path**, following sections 1–10 above.
3. **Ask the operator, then `fleet_handover{action: "confirm", token, operatorConfirmed: true}`.**
Why that order: `confirm` refuses with `HANDOVER_STALE` unless the file was modified **after** the
`open` request. That check stops a leftover file from an earlier session being accepted as this
one's handover. So writing the file first and then calling `open` — the obvious order — always
fails.
`{action: "cancel", token}` drops a pending request without rolling.
**Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
get another one.
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
from your connection, so you can only ever roll yourself.
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
defaults to `true` and this is the only thing standing between a judgement call and a wiped
session.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
the operator so before you confirm.** The recovery is the same either way: the file is already
written, so the operator starts a session and points it at the file. That is why you write the
file before you confirm, and never the other way round.
## Writing style
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
class, method, file, flag, and config key exactly as it appears in the code — replacing a
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
first time you use it.
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
brackets, colons, or slashes.
## Template
```markdown
# Handover — <fleet name> lead session, <date and time>
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
spawns, or restarts.
## 0. Do these three things first
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
## 1. What the operator asked for
<the live instructions, in their words where you have them; what is still open; any decision
you made yourself, and why>
## 2. Open decisions, and who owns them
<one line per decision: who owns it, what is unanswered>
## 3. Live hazards
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
## 4. What is not owed
<finished work, and work another party has declined; name who said so and when>
```
+102
View File
@@ -0,0 +1,102 @@
---
name: hunter
description: Defect-hunt procedure for a fleetd worker — sweep an assigned package for real bugs and report several ranked findings without fixing anything. Load this when the lead asks you to hunt or audit a scope rather than review one diff. Do NOT load `reviewer` for this; the two want different output.
---
# Hunter worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies.
**This skill is not `reviewer`.** `reviewer` judges one diff and reports the *single* most
important issue in about 90 words. A hunt sweeps a whole package and reports *several* findings
in a long structured form. Loading both gives you two contradictory output contracts, and the
usual result is a worker that writes a good report into its terminal and ends the turn without
sending it. Load exactly one.
## 0. Read this before you read code: how the report gets home
Your terminal reaches nobody. The lead sees **only** the text inside your `fleet_reply` call.
A long report is exactly the case where this goes wrong, so plan for it:
- **Write the report into the `fleet_reply` argument itself.** Do not compose it in your terminal
and then summarise it into the call.
- If the report is long, **send it anyway** — one `fleet_reply` with everything.
- If you end the turn without replying, the bridge scrapes your pane instead. That scrape carries
at most the last 4000 characters, and on a hunt it usually captures the tail of the lead's own
brief rather than your findings. The lead then has nothing and has to ask you again.
## 1. Change nothing
A hunt is read-only. Do not edit a production file, do not "quickly fix" what you find, and do
not run a formatter. You may run the build and tests to *check* a claim, and you should say so
when you did.
## 2. Read the whole scope first
Read every file in the assigned package before you judge any of it. A defect that a caller
elsewhere in the same package makes unreachable is not a defect, and you cannot know that from
one file.
Stay inside the scope. If a defect there depends on a class outside it, read that class to
confirm — but the defect itself must live in the scope you were given.
## 3. The bar — this matters more than the count
**Name the path into the bad state.** Say which caller, in which state, reaches it. A defect on
paper is not a reachable defect. If you cannot name that path, keep the finding but mark it
`unproven` and say exactly what you could not check. Do not drop it, and do not dress it up.
**Say which direction the harm goes.** Data loss, privilege escalation and silent wrong answers
are worth reporting even when the window is narrow. A finding whose worst outcome is a worse log
line is not worth a block.
Two workers once ran the same scope: the one that applied the direction-of-harm filter found ten
real defects, the one that did not found none. Fewer findings the lead can act on beat many the
lead has to triage.
## 4. Shapes that have produced real merged fixes here
Read for these first:
1. **A one-way gate.** A guard added after an incident closes only the direction that incident
came from. Do not only ask what closes the gate — ask **which states still open it**.
2. **A value read once, then used later to authorise something destructive**, after something
else has had a chance to change it.
3. **A failure downgraded to a value that looks like a legitimate result** — `-1`, `null`, an
empty list, `false` — which a caller then trusts.
4. **A lock held for one half of a read-modify-write and not the other**, or two collections
updated under different locks.
5. **A comment or javadoc stating an invariant the code no longer keeps.** Comments are
load-bearing in this repo; a stale one has already caused a bug.
## 5. What you cannot check, and must not claim you did
- `fleetd/fleetd.yaml` is gitignored and **absent from your worktree**. You cannot read it. If a
finding depends on live configuration, name the key and say you could not check it.
- `.mcp.json`, `opencode.json` and `.autoenv` in your worktree are neutralised stubs, not the
repo's real files.
- The `wiki/` submodule pointer is months old. Do not cite it.
Reporting a fact you took from the lead's brief as something you measured yourself is a false
report, even when the fact is correct. Say where each fact came from.
## 6. The report — what goes in `fleet_reply`
One block per finding, most severe first:
```
FINDING N — <one line>
file:line
Path in: <which caller, in which state, reaches this>
Direction: <data loss | escalation | silent wrong answer | outage | ...>
Window/trigger: <when it actually happens>
Confidence: <confirmed by reading | unproven — say what you could not check>
Why nothing else catches it: <the guard or test you checked, and why it misses>
```
End with one line naming every file you read, so the lead knows the denominator.
**Nothing clears the bar?** Reply `NO FINDINGS`, name the files you read, and say what you ruled
out. A clean sweep is a valid result; an invented defect is worse than none.
+72
View File
@@ -0,0 +1,72 @@
---
name: redeploy-fleetd
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
---
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
+5
View File
@@ -10,6 +10,11 @@ never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and alrea
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
**Wrong skill for a sweep.** This one reviews *one* diff or scope and reports the *single* most
important issue. If the lead asked you to hunt or audit a whole package for several defects, load
`hunter` instead and ignore this file — the two want different output, and following both is how a
worker ends its turn with a good report that never gets sent.
## 1. Read the whole scope before you judge
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
+31 -4
View File
@@ -39,6 +39,10 @@ jobs:
# that runs does so against the fake UDS herdr and fake ccs/claude stubs.
run: mvn -B clean install
- name: javadoc reference lint
working-directory: fleetd
run: mvn -B -DskipTests javadoc:javadoc -Ddoclint=reference
# Deliberately NOT actions/upload-artifact: this Gitea instance presents as GHES, and
# @actions/artifact v2+ (i.e. upload-artifact@v4) refuses to run there —
# "GHESNotSupportedError ... not currently supported on GHES", which red-Xes an otherwise
@@ -54,6 +58,23 @@ jobs:
done
exit 0
# fleetd #550 — nothing ran scripts/test-redeploy-fleetd.sh in CI before this, on any platform,
# so it had run only on macOS by hand and two Linux-only bugs (this issue's items 1 and 2)
# survived undetected: shasum is a macOS-only tool (it ships with Perl; GNU coreutils, i.e. every
# mainstream Linux distro including this runner's ubuntu-latest, does not have it and ships
# sha256sum instead). The gate here is the step's own exit code, nothing else: a `run:` step in
# Gitea/GitHub Actions already fails the job on a non-zero exit with no extra scripting needed,
# so this deliberately does NOT grep the output for a `FAIL:` count. That is the #550 item-2
# lesson one level up — a suite that dies before it runs a single test prints zero FAIL lines,
# which is exactly what a clean pass also prints, so counting FAIL lines can never be the gate.
shell-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: redeploy-fleetd.sh shell suite
run: bash scripts/test-redeploy-fleetd.sh
# CB-521 — actually run the AMQP contract test in CI, against a REAL broker. The broker is a
# RabbitMQ SERVICE CONTAINER, not Testcontainers-with-Docker: the runner image has no Docker, so
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
@@ -87,12 +108,18 @@ jobs:
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
# The `contract` profile clears the default-excludes group, so `-Dgroups=contract` runs every
# @Tag("contract") test and nothing from the unit suite the `build` job already covered — a
# tag selects the whole group, so a test added to it later runs here automatically. A prior
# version of this step pinned `-Dtest=AmqpReplyInboxContractTest` by class name instead: that
# silently excluded every other contract test (including the herdr ones) from CI, and nobody
# noticed until the herdr protocol drifted out from under a test that never ran here
# (fleetd #449). If this runner has no herdr socket, the herdr-backed tests in the group
# skip on their own `assumeTrue` and only the broker-backed ones actually run — check the
# step output rather than assuming which.
- name: Contract tests
working-directory: fleetd
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
run: mvn -B -Pcontract test -Dgroups=contract
- name: Failing test output
if: failure()
+6
View File
@@ -19,3 +19,9 @@
fleetd.out
fleetd/fleetd.out
logs/
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
# has no business in git history.
.handover/
+135 -73
View File
@@ -7,6 +7,14 @@
> wiki ([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
>
> **Anything you measure in an addendum is perishable.** Date it, give the command that
> re-measures it and what each outcome means, and tell the reader to delete the section once
> it stops reproducing. The four parts work together: deciding what would falsify a claim is
> the expensive step, and a reader in the middle of another task will not pay it, so a bare
> "verify before relying on this" costs the same space and does nothing. The case this is for
> is a note that goes stale as a live restriction — it will tell a future session it cannot do
> the thing at the moment doing it becomes the job.
If no `fleet_*` MCP tools are mounted in this session, this section does not apply — skip it.
@@ -50,8 +58,9 @@ and the sender silently receives nothing. Fail toward the recoverable error.
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
### Primary (lead) — run this on every task, in order
@@ -71,8 +80,11 @@ below are the procedure — run them in order, every task, not only the big ones
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `fleet_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model, cost and
LIVENESS, not in tier, so the default is rarely what you want. The default is whatever the
daemon reports, and on a host where it sits on an exhausted or withdrawn credential every
unqualified spawn fails — sometimes loudly, sometimes as a member that spawns fine and then
produces nothing. `fleet_profiles` reports the default; check it once per session.
4. **Then send them all** — `fleet_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
@@ -83,9 +95,21 @@ below are the procedure — run them in order, every task, not only the big ones
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
**A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and
returns a ticket, and is then never delivered — measured here three times in one session, and the
member was released still executing a brief that had been retracted twice. The receipt is true and
it is a fact about the *mailbox*; what you needed was a fact about the *pane*. **A push delivery
needs the recipient free at send time; a pull channel needs only that they look before acting.** So
put every correction on the **ticket**, which they can read whenever they look, and send as the
notification. That obliges you, not them: **all corrections go to the ticket, and the brief is
write-once.** The member cannot check which source is newer — it just always prefers the ticket —
so the day you revise a brief in place instead of commenting, it obeys your rule and does the wrong
thing. A *first* brief for a unit not yet running is not a correction, and may be the whole spec.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge tools it appears to have hold a blocked credential and fail, and a piped command
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
forge MCP server it appears to have holds a blocked credential and fails every call, and a piped
command (`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a
fact. Its injected repo-scoped `GITEA_TOKEN` is a different credential and does work, so a worker
reporting that it opened its own PR is reporting something it really can do.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -94,6 +118,18 @@ below are the procedure — run them in order, every task, not only the big ones
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `fleet_stop{paneId}`.
**If the forge refuses you the merge** — a protected branch, a token without the grant — the
adjudication is still yours. Read the diff, decide, and hand the operator a merge-ready queue
with the refusal quoted. Never report a PR as merged, and never call one "ready to merge"
without having read the diff yourself. A refusal is exactly when that shortcut is tempting,
because no action is left that forces you to look, and taking it turns this step into
forwarding a reviewer's verdict — which is delegating the merge by proxy, two lines above.
**Test a refusal; do not read it off a permissions field.** A protected branch holds its merge
rights separately from the repository permissions, so that field can say yes while the merge is
refused, and still say no after a grant makes it work. Probe instead, with a request that cannot
succeed on its merits, so a rejection can only mean the refusal. Treat a transport failure as a
third answer that proves nothing: a timeout, a DNS error or a bad URL is not a refusal, and
counting it as one makes you sure of something you never measured.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
@@ -103,20 +139,31 @@ prefer `wait:false` + `fleet_poll` for anything non-trivial: a blocking `fleet_s
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
the merge — and merging on a reviewer's word is delegating it by proxy.
**When a decision blocks you, consult architects — not the operator.** Spawn one or more architect
members, give them the question and the evidence you have, and act on what they agree. They are
authorized to settle it, not only to advise. If two of them still disagree after two rounds, they
return both positions and you decide. Go to the operator only for something outside the fleet's
authority: money, credentials, or a promise made to someone else. **Then write the decision on the
ticket.** Taking the operator out of the loop also removes the signal they used to get, because
that signal was the block itself — work stopped, so they found out. A ticket comment replaces it,
and it reaches them whether or not they are at a terminal when you decide.
| Intent | Tool |
|---|---|
| Confirm your own role | `fleet_whoami` |
| See backends available | `fleet_profiles` |
| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `fleet_status{sessionId}` |
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_reply{content}` — the one case a lead replies |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after* `open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
### Lead ↔ lead — coordinate, never delegate
@@ -140,10 +187,21 @@ The traffic between leads is coordination and nothing else:
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
**N observations are N data points only if they differ in the axis you are trusting.** This cuts
both ways. N *failures* blamed on one cause are one data point when the cases share what you are
not varying. N *agreeing measurements* are also one data point when they share an instrument —
two hosts, two operators and the same formula is one formula, not two confirmations.
4. **Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
author is the worst reader of their own qualifier placement — measured here, one addendum carried
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
sentence goes false first, and would a reader reach the caveat before acting?"
Being messaged by a peer does not make you its worker: answer with `fleet_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
Being messaged by a peer does not make you its worker: answer the way you would open —
`fleet_send{coordId}` for another daemon, `fleet_send{sessionId}` on this host — and push back on
the substance if it is wrong. `fleet_reply` resolves a member's blocked `fleet_send`; a peer's
coord-id message is durable and non-blocking, so there is nothing for it to resolve. A peer that
simply complies has thrown away the reason there are two of you.
### Member (worker or architect) — the turn contract
@@ -152,20 +210,26 @@ you.
3. **`fleet_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
4. **Re-read the ticket before you act on anything you were told earlier**, and again before you
commit. A message reaches you only while you are free to receive it; the ticket is there whenever
you look. **If a ticket comment contradicts your brief, the ticket comment is newer and it wins.**
5. **End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `fleet_reply` ⇒ the sender gets nothing and the exchange stalls.
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
5. **Report honestly.** State only what you actually ran and its real output, including failures,
6. **Report honestly.** State only what you actually ran and its real output, including failures,
and never claim the result of a check you had no way to run. **Measure your own tools; do not
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
find there holds a deliberately blocked credential and fails every call, by design.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
yours, whatever you see. And **a mounted tool is not a working tool** — the forge MCP server you
may find there holds a deliberately blocked credential and fails every call, by design. That is
not your only forge route, and the two must not be confused: the repo-scoped `GITEA_TOKEN` the
daemon injects into your environment does work, and using it to open your own PR is part of the
job. A blocked MCP tool is never a reason to skip that step.
7. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
### Where each rule lives (don't duplicate — extend the right layer)
@@ -188,6 +252,19 @@ must obey belongs in the charter, not here.
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Herdr socket tests (measured 2026-09-10).** In this repo, herdr is a subject under test. A
worker assigned to herdr code, and the lead, may let a test open the herdr socket directly in a
throwaway workspace that the test tears down. This only covers
`fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java`, and
`fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java`. It is not a
general licence. Using the herdr CLI or socket to move a real fleet session, pane, or peer stays
banned. That is the control plane that invariant 5 protects. Re-measure with
`grep -rl 'UnixSocketHerdrClient.connect()' fleetd/src/test/java --include='*.java'`. A non-empty
result means tests still open the socket and this note still applies. An empty result means nobody
does this any more; delete this section. Canonical invariant 5 restatement is tracked in #458 and
is not part of this change.
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
@@ -200,11 +277,32 @@ must obey belongs in the charter, not here.
adapter, with a message naming the credential and the remaining seconds ("cooling off after
repeated backend errors") — distinct wording from a quarantine refusal, so don't conflate the
two when reading a spawn failure.
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR),
`reviewer` (one diff → one structured finding) and `hunter` (sweep a package → several ranked
findings, change nothing). Name exactly one in every delegation. **`reviewer` and `hunter` are
not interchangeable** — `reviewer` caps the answer at one finding in about 90 words, so naming
it for a multi-finding sweep hands the worker two contradictory output contracts. That has
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
Spawn `implementer` with role `dev`, `reviewer` with role `reviewer`, and `hunter` with role
`hunter`.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
`fleets-status` (report every fleet that shares one LavinMQ instance).
`port-to-opencode` (make an OpenCode session a participant in this workspace),
`fleets-status` (report every fleet that shares one LavinMQ instance),
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
fleetd #480).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
unmentioned by every instruction file, so it drifted and a later session planned it from scratch
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
worktree; and a plugin cannot deliver anything to members at all, because
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it,
so a member never reads the operator's plugin store. **The plugin is the lead-side surface;
member-facing assets travel in the worktree.**
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
@@ -227,60 +325,14 @@ must obey belongs in the charter, not here.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
worker has no business restarting the daemon it is talking through, and stopping it kills the
worker's own channel mid-turn.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
and its flags, the drain step, the operator's permission grant, and the five checks that have each
gone wrong here before. Do not hand-roll the steps from memory.
### The prompt is part of the product — update it with the code (mandatory)
@@ -328,6 +380,16 @@ print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
PY
```
**Only the lead can run that check (measured 2026-09-10).** A member's provisioned worktree has
`wiki/` uninitialized, so the script dies with `FileNotFoundError: wiki/7-Use-Cases.md`. Measured
in three worker worktrees: `git submodule status` printed a leading `-` and `wiki/` held 0
entries; the primary's own clone printed a leading `+` and the file was there. So never make this
check a member's acceptance criterion — it is unsatisfiable for them, and a brief that asks for it
is asking a worker to invent a pass. A member told to check it must say it could not run it, and
must never report it as passed. The lead runs it in the main clone before merging. Re-measure with
`git submodule status` in a member's worktree: a leading `-` means this still applies; once it
prints a commit with no `-`, delete this paragraph.
## IDE MCP tools & validation workflow (enforced)
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
+54 -46
View File
@@ -1,72 +1,80 @@
# CB-504 — systemd unit for fleetd (Linux).
#
# The macOS launchd agent (deploy/dev.ltms.fleetd.plist) is the supervision target for the
# current single-host deployment. This unit exists for the per-host gateways CB-308 introduces,
# which will run on Linux.
# fleetd #360: the previous version of this file started clean and broke the daemon in three ways
# that nothing logs (see the DO NOT block and the ExecStart/PrivateTmp comments below for what and
# why). The unit below, plus its companion deploy/herdr.service, is the version that has actually
# run on fleet01 without those failures. Do not "improve" it back toward the old shape without
# re-reading why each line is the way it is.
#
# Install (user service — fleetd drives the user's herdr, not a system daemon):
# mkdir -p ~/.config/systemd/user
# cp deploy/fleetd.service ~/.config/systemd/user/
# # edit ExecStart / WorkingDirectory / Environment below, then:
# cp deploy/fleetd.service deploy/herdr.service ~/.config/systemd/user/
# # edit WorkingDirectory / ExecStart below for your host's paths and java location
# systemctl --user daemon-reload
# systemctl --user enable --now fleetd
# systemctl --user enable --now herdr fleetd
# loginctl enable-linger $USER # REQUIRED -- see below
# journalctl --user -u fleetd -f
#
# `loginctl enable-linger` is not optional and is easy to miss, because leaving it out looks like
# success: `systemctl --user enable` reports "enabled" and both units run for as long as you stay
# logged in. A user manager without lingering starts at your first login and stops at your last
# logout, so the fleet simply does not come back after a reboot -- which is the whole reason to
# use systemd here rather than the setsid scripts these units replaced. Check it with
# `loginctl show-user $USER -p Linger`; the answer must be `Linger=yes`.
#
# Secrets (AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI, ...) are not set
# here and need no systemd drop-in: ExecStart runs a login shell, so they come from wherever your
# login shell already sources them (this host: ~/.fleet/secrets.sh via ~/.zprofile). If a token is
# missing there, fleetd still starts — the daemon reports every secret a configured profile
# references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)"; a missing one logs
# "startup secret NAME: MISSING" at WARN and the daemon starts anyway — the first visible symptom
# is a member that cannot open a pull request, hours later and in a different component.
[Unit]
Description=fleetd — claude-bridge message server
Description=fleetd — fleet message server
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
# Ordering only: herdr is a user process and its socket may appear after us. This is advisory —
# fleetd retries the herdr socket rather than exiting, which is what actually makes a late
# socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down with it.
# Ordering only. fleetd retries the herdr socket rather than exiting, which is what actually makes
# a late socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down too.
After=herdr.service
Wants=herdr.service
[Service]
Type=simple
WorkingDirectory=%h/src/claude-bridge/fleetd
ExecStart=/usr/lib/jvm/temurin-25-jdk/bin/java -jar target/fleetd.jar fleetd.yaml
WorkingDirectory=%h/LTMS/fleetd/fleetd
Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
# PATH matters more than it looks (CB-511): fleetd propagates its own PATH to every worker it
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
# Secrets are NOT set here — this file is committed. Put ALL three tokens in a private drop-in
# that systemd reads with restrictive permissions. In `systemctl --user edit fleetd`, add:
# [Service]
# Environment=FLEETD_API_TOKEN=...
# Environment=WORKER_GITEA_TOKEN=...
# Environment=AI_GATEWAY_TOKEN=...
# FLEETD_API_TOKEN protects fleetd's API. WORKER_GITEA_TOKEN lets members open pull requests; if
# it is missing, fleetd still starts, but a member fails when it later tries to open a pull request.
# AI_GATEWAY_TOKEN authenticates gateway profiles; if it is missing, fleetd still starts, but a
# gateway profile later returns HTTP 401. Or, put the same three variables in a 0600 file and add:
# EnvironmentFile=%h/.config/fleetd/env
# After starting, check which of them actually resolved. The daemon reports every secret a
# configured profile references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)". A missing one logs
# "startup secret NAME: MISSING" at WARN — and the daemon starts anyway, which is the whole
# problem: without this grep the first sign is a member that cannot open a pull request, hours
# later and in a different component.
# Note what the report can and cannot tell you. It lists only names some profile actually
# references (tokenEnv, gitTokenEnv, and the broker uriEnv). A secret nothing references is never
# reported, because nothing needs it.
# A LOGIN shell, not java directly. Every secret this daemon needs (AI_GATEWAY_TOKEN,
# WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in ~/.fleet/secrets.sh, which only
# ~/.zprofile sources. systemd runs no login shell. Started any other way the daemon boots fine
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
# unit) has to read them. A private /tmp turns the credential scrub into a silent no-op.
PrivateTmp=false
Restart=on-failure
RestartSec=10s
# A bad config (e.g. a non-loopback bind without token auth) makes fleetd fail fast by design.
# Give up rather than restart-loop on a permanent error.
# A bad config makes fleetd fail fast by design. Give up rather than restart-loop forever.
StartLimitBurst=5
StartLimitIntervalSec=120
# The daemon reads the repo, writes worktrees, and talks to a Unix socket — it needs no more.
# DO NOT add ProtectSystem=, ProtectHome=, ProtectKernelTunables= or ProtectControlGroups=.
# Measured on fleet01 2026-09-05: each of those gives the unit its own mount namespace, and
# fleetd resolves a caller role by running lsof to find the loopback peer PID
# (mcp/LsofPeerPidLookup). Inside such a namespace lsof returns nothing, every caller falls back
# to ANONYMOUS, and the primary is refused every orchestration call with
# "unauthenticated: anonymous may not SPAWN".
# The daemon still starts, healthz still returns ok and the secrets still resolve - the only
# symptom is that the fleet cannot be driven at all. Verified by bisecting the directives:
# no sandbox 3 lsof lines | ProtectSystem=strict 0 | ProtectHome=read-only 0
# ProtectKernelTunables 0 | ProtectControlGroups 0 | RestrictSUIDSGID 3 | NoNewPrivileges 3
# The two below add no mount namespace and are safe.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-write
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true
StandardOutput=journal
+19
View File
@@ -0,0 +1,19 @@
#!/bin/zsh
# fleetd #360 — template for the script deploy/herdr.service's ExecStart wraps in a pty.
#
# `script -qfec <this> /dev/null` needs a real command to run, and that command has to be a LOGIN
# shell script: herdr itself needs the same secrets fleetd.service's login shell picks up (this
# host: ~/.fleet/secrets.sh via ~/.zprofile), because members it spawns inherit its environment.
# systemd's own Environment= lines in herdr.service are not enough for that -- they set TERM and a
# bare PATH so the pty starts at all, nothing more.
#
# Copy this file to the path deploy/herdr.service's ExecStart names
# (%h/LTMS/fleetd/fleetd-run/herdr-inner.sh by default) and `chmod +x` it. Not committed under
# that path itself because the session name below is host-specific.
# A 0x0 pty makes every pane spawn fail with "ghostty error -2" (see herdr-multi-instance-facts /
# fleet01-headless-herdr-standup) -- give it a real size before herdr ever touches it.
stty rows 50 cols 200
# -l: login shell, so herdr and everything it spawns gets the real secrets and PATH.
exec zsh -lc 'exec herdr --session <name>'
+40
View File
@@ -0,0 +1,40 @@
# fleetd #360 — systemd unit for herdr (Linux), the terminal multiplexer fleetd drives.
#
# This is fleetd.service's companion: fleetd.service's After=/Wants=herdr.service assumes this
# unit exists. Before this ticket it did not, so on a fresh host fleetd started against a herdr
# that systemd never supervised at all.
#
# Install: see deploy/fleetd.service's header comment (both units install the same way).
#
# ExecStart below runs deploy/herdr-inner.sh (copy the template of that name from this directory
# to the path in ExecStart, or point ExecStart at wherever you keep it, and make it executable).
# It is a separate file rather than an inline command because it must itself be a login shell (see
# its own header for why) and systemd's ExecStart does not run one.
[Unit]
Description=herdr terminal multiplexer (fleet session)
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
[Service]
Type=simple
# script(1) gives herdr a real pty. Without it the client reports a 0x0 window and every pane
# spawn fails with "ghostty error -2" -- which surfaces as a fleetd spawn failure, not a herdr one.
ExecStart=/usr/bin/script -qfec %h/LTMS/fleetd/fleetd-run/herdr-inner.sh /dev/null
StandardInput=null
Environment=TERM=xterm-256color
Environment=PATH=%h/.local/bin:/usr/local/bin:/usr/bin:/bin
# PrivateTmp MUST stay false. fleetd writes the member ZDOTDIR scrub dir and the opencode config
# dir under its own java.io.tmpdir, and the member pane -- a child of THIS process -- has to read
# them. A private /tmp here silently breaks the credential scrub instead of failing loudly.
PrivateTmp=false
Restart=on-failure
RestartSec=5s
StandardOutput=journal
StandardError=journal
SyslogIdentifier=herdr
[Install]
WantedBy=default.target
+19 -1
View File
@@ -5,6 +5,21 @@
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
with no terminator, and our third-party members cannot detect it (§7.2).
> **Superseded in part — 2026-09-13.** Two claims on this page are no longer true of the live fleet.
> I measured both on this host today.
>
> 1. **The model is named `acoder` now, not `deepseek-v4-flash`.** `acoder` is a stable alias, and
> the model behind it changed on 2026-08-28: it is Qwen3.8-27B, not DeepSeek. The old name is
> still served, so nothing broke — the gateway answers it and reports `"model": "acoder"` in the
> reply, which is how you can see for yourself that it is an alias. `fleetd.yaml` moved to
> `acoder` on 2026-09-13. Do not guess behaviour from the name; ask the gateway's own manifest,
> `GET https://llm.ltms.dev/v1/deployment`, and read its `generation` field.
> 2. **`local` sits at `weight: 0`, not 100.** Only `gx` is auto-selected today.
>
> §2 and §3 below are the plan as written in August. They are the record of the migration, so they
> stay as they are. If this note stops matching `fleetd.yaml`, re-measure and rewrite the note.
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
@@ -380,7 +395,10 @@ one turn; this one costs the whole task and is indistinguishable from a slow wor
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear.
- `/v1/models` returned exactly `["deepseek-v4-flash"]` **on 2026-08-15**, so trap 3 was clear
then. It returns 6 ids now — `acoder`, `qwen3.8-27b-nvfp4`, `deepseek-v4-flash` and three
embedding names — measured on this host 2026-09-13. The exact-name rule still holds; the
one-item list does not.
- **Reasoning survives both surfaces** — see §3b above.
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
key rather than the `fleetd-local-noauth` placeholder.
+1 -1
View File
@@ -192,7 +192,7 @@ Deliberately small; every one maps to a failure mode we have actually hit.
| `fleet_send_duration_seconds` | histogram | delegated turn latency |
| `fleet_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
| `fleet_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ delivered\|exhausted — a rising `exhausted` means the primary is not draining |
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ sent\|exhausted — a rising `exhausted` means the primary is not draining. `sent` was called `delivered` until fleetd #365; it counts the herdr paste-and-submit call returning, never a confirmation the pane read it |
| `fleet_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
| `fleet_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
| `fleet_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
+145 -17
View File
@@ -88,6 +88,47 @@ bind:
# backoffMs: 60000
# quietNudgeCap: 3
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
# own pane and bootstrap a fresh session against that file.
#
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
#
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
# reads must live. No default (an operator-specific path); a present block with no
# handoverPath refuses to start. May be relative: it then resolves against the
# CALLING lead's own fleet.leaders.<name>.cwd (falling back to the daemon's own
# working directory when that lead has none configured) — never against whatever
# directory the daemon process happens to have been started in. An absolute path is
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
# share a working directory (fleetd #480 follow-up).
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
# # operatorConfirmed: true
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
# # lead's own turn to end (its pane to report injectable again)
# # before sending /clear at all. If this elapses, /clear is NEVER
# # sent — a lead that never goes idle is still doing real work.
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
# # again AFTER /clear before giving up (never sends bootstrapText
# # if this elapses). A separate, second wait from turnSettleSeconds.
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
# # the lead once its pane settles after /clear
# leadRollover:
# handoverPath: /path/to/handover.md
# requireOperatorConfirm: true
# maxDocAgeSeconds: 3600
# turnSettleSeconds: 20
# clearSettleSeconds: 20
# bootstrapText: "Fresh lead session: read the handover file and carry on."
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
# per tick.
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
@@ -110,6 +151,14 @@ bind:
# notifications:
# mode: disabled
# Idle-sleep guard: while at least one member is live, hold an OS-level assertion against idle
# sleep (macOS only — a `caffeinate -i` child; a no-op elsewhere or if caffeinate is missing), so
# an unattended host does not idle-sleep out from under a member's long turn. Unlike health/
# configReload above, this is ON BY DEFAULT — omitting the block entirely leaves it enabled, the
# same as `enabled: true`. Uncomment only to turn it off:
# idleSleepGuard:
# enabled: false
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -199,8 +248,12 @@ herdrSocket: ~/.config/herdr/herdr.sock
# omit and this profile's completion fallback behaves exactly as before.
# Every backend words its refusal differently, so this is config, never a
# vendor string baked into fleetd itself.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart, same as this profile's model/baseUrl/argv.
# HOT (fleetd #446): read live, cached by profile name, at every
# completion-fallback check AND by fleet_profiles' exhaustionDetectionArmed —
# editing it and reloading arms or disarms usage-limit detection for this
# profile with no daemon restart. (Before fleetd #446 this was DEFERRED,
# compiled once into a startup pattern map like model/baseUrl/argv still are —
# see errorPattern below, which is still deferred that way on purpose.)
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
# credentialId share one quarantine — the case this exists for is two models
@@ -219,8 +272,9 @@ herdrSocket: ~/.config/herdr/herdr.sock
# happens, just without a profile-specific match; every backend words its
# failure differently, so a hardcoded sentence would only ever match one
# of them.
# DEFERRED: compiled once into a startup pattern map, same as exhaustedPattern
# — editing it needs a daemon restart.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart. Unlike exhaustedPattern above (made hot by fleetd #446),
# errorPattern was scoped out of that ticket on purpose and stays deferred.
# # errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage
#
# What happens once a match fires (BackendOutagePolicy, credentialId-keyed,
@@ -439,6 +493,15 @@ placement: weighted
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
#
# fleetd #466: this is now only the BASE of an escalating backoff, not a flat retry rate. A
# credential quarantined again within one base cooldown of the previous quarantine ending (still
# reporting exhausted — e.g. a weekly subscription limit that hasn't reset) backs off further:
# cooldown doubles each such time, capped at 12x this value (~6 hours at the 1800s default). A
# quarantine that starts after a base-cooldown's worth of quiet resets back to this value. Not
# configurable per se — the multiplier and ceiling are constants in BackendQuarantine, not new
# YAML keys; see its class doc for the exact formula and why there is no automatic probe to clear
# it early (the operator's own design constraint — a probe spends the quota it's measuring).
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
@@ -460,9 +523,10 @@ placement: weighted
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
# / credentialId. Those are hot because the placement policy (and, for credentialId,
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
# not by itself enough to make a key hot.
# / credentialId / exhaustedPattern. Those are hot because the placement policy (and,
# for credentialId, the CB-578 stage B quarantine check; for exhaustedPattern, fleetd
# #446's LiveExhaustedPatterns) reads them through a supplier — being config is not by
# itself enough to make a key hot.
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
@@ -474,10 +538,10 @@ placement: weighted
# stage B — baked once into the quarantine tracker built at startup), ADDING or
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
# configDir, mcpUrl, tabLabel, exhaustedPattern, errorPattern (fleetd #201 / #227 —
# compiled once into a startup pattern map the same way exhaustedPattern is). The
# launcher takes a copy of `profiles:` at startup and resolves every spawn out of
# that copy, so those never
# configDir, mcpUrl, tabLabel, errorPattern (fleetd #201 / #227 — compiled once into
# a startup pattern map; exhaustedPattern used to be compiled the same way until
# fleetd #446 made it hot — see above). The launcher takes a copy of `profiles:` at
# startup and resolves every spawn out of that copy, so those never
# reach a launch until you restart. The reload logs them by name rather than
# pretending they applied.
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
@@ -496,7 +560,7 @@ placement: weighted
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
#
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
# role — which contract: architect, dev, hunter or reviewer. It picks the launch charter, the role
# file, the playbook skill and the authz row.
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
@@ -508,13 +572,13 @@ placement: weighted
#
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
# with being listed here; the entry key just names the entry.
# the candidates, in definition order. A dev, hunter and reviewer staying anonymous is exactly
# compatible with being listed here; the entry key just names the entry.
fleet:
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
# is deliberately not supported.
# hunter, reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put
# secrets here: a later launch step writes this text to a world-readable temp file, and ${ENV}
# interpolation is deliberately not supported.
charters:
architect: |-
You are an architect in this fleet. You refine work before anyone builds it:
@@ -525,6 +589,9 @@ fleet:
dev: |-
You implement the one unit you were given, and nothing else. You test it,
commit it, and open your own pull request. You never merge.
hunter: |-
You sweep the assigned scope for real defects. You may run the build or tests
to check a finding. You change nothing, and report several ranked findings.
reviewer: |-
You review the diff you were given. You report bugs, risks and missing tests.
You do not change code.
@@ -602,6 +669,9 @@ fleet:
developers:
gx10:
profile: gx10
# hunters:
# gx10:
# profile: gx10 # a hunt may run checks, but never changes code
# reviewers:
# gx10:
# profile: gx10 # the same backend may serve two roles; that is the point
@@ -774,6 +844,28 @@ guard:
# so fleetd falls back to the weaker CB-596 sentinel overlay instead (a WARN names the gap).
# worktreeGroup: fleet-workers
# fleetd #362: a directory of skill folders (each a subdirectory holding a SKILL.md, the same
# shape as this repo's own .claude/skills/) copied into every PROVISIONED worktree's
# .claude/skills/, so a member spawned against ANY repo — not only one that already ships its own
# copy — can load a bridge skill (e.g. implementer). Unset (the default): no worktree is touched
# beyond today's behaviour. A skill folder the target repo already carries under
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Best-effort like
# worktreeGroup above: a missing/unreadable directory here is logged and skipped, never a failed
# spawn. Every non-hidden subdirectory of this directory is copied wholesale, with no per-file
# allowlist — don't park scratch files or drafts alongside the real skill folders, they will be
# copied into every provisioned worktree too.
#
# fleetd #393: which member KINDS actually consume this once it is copied. kind: claude-code —
# the Claude Code CLI discovers .claude/skills/ on its own; nothing else is needed. kind: opencode
# — opencode has no such discovery, so OpenCodeLauncher reads whatever landed under
# .claude/skills/ and appends each seeded skill's SKILL.md to the generated instructions[] file
# (opencode's only channel for static guidance text; unlike Claude Code's Skill tool, the content
# is always part of the system prompt, not loaded on demand). Both kinds are covered as of #393 —
# earlier builds copied the files for every kind but only claude-code could read them, and the
# seeding log said "N of M" regardless. Check the per-spawn launcher log (not just the seeding
# log) to see what a given member actually got.
# memberSkills: /path/to/fleetd/checkout/.claude/skills
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
@@ -821,10 +913,17 @@ guard:
# across every daemon sharing this vhost.
# prefetch → consumer basicQos, capping how many unacked messages the mailbox holds in-heap.
# Default 32 when omitted.
# peers → fleetd #361: the coord-ids of the OTHER daemons on this vhost, declared by the
# operator (the daemon never guesses). fleet_list reports each one's live reachability
# (a passive queue check, never a presence protocol) alongside this daemon's own
# mailbox state. Omit, or leave empty, for a daemon with no known peers yet — an
# undeclared peer can still reach you and be reached by fleet_send, it just will not
# show up as a row in fleet_list.
# coordinator:
# uriEnv: LEAD_COORD_URI
# selfId: mac-opus
# prefetch: 32
# peers: [fleet01-lead]
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open fleet_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
@@ -849,3 +948,32 @@ guard:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
# Central allow-list of models any profiles: entry may name. Nothing checked a profile's model:
# value before this block existed — it was a free-form string handed straight to the backend
# adapter, and a withdrawn or misspelled name failed silently instead of at config load (opencode
# falls back to a default model rather than erroring on an unknown -m).
#
# Absent, or present with an empty allow:, is OFF: no profile's model: is checked, exactly like
# before this block existed. fleetd.yaml is gitignored on every host, so an upgrade must not force
# every operator to enumerate their models before the daemon will start.
#
# The list is the authority; profiles: is checked against it, never the reverse — adding or
# editing a profiles: entry cannot, by itself, widen what is permitted here.
#
# Enforcement is at CONFIG LOAD only (a bad model: fails the daemon at startup, naming both the
# model and the profile). There is no spawn-time enforcement, no runtime on/off switch, and no
# interaction with BackendQuarantine — those are separate, later units.
#
# allow → the permitted models. Each entry is its own block (not a bare string) so a later unit
# can add an on/off state or a load limit per model without changing this shape.
# model → the model id exactly as a profiles: entry's model: field would write it. One flat,
# opaque-string namespace: a bare Claude id (claude-sonnet-5) and an opencode
# provider-prefixed id (openai/gpt-5.6-terra) both fit here unchanged — the check is a
# plain string match, never a parse of the provider prefix or a branch on kind:.
# models:
# allow:
# - model: claude-sonnet-5
# - model: claude-opus-5
# - model: openai/gpt-5.6-terra
# - model: amazon.nova-pro-v1:0
File diff suppressed because it is too large Load Diff
@@ -30,8 +30,27 @@ public final class Authz {
DRAIN,
/** Read-only observation: status, roster, profiles, task polling. */
READ,
/**
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
*
* <p>Deliberately <strong>not</strong> folded into {@link #READ}. {@code READ}'s grant
* rests on "the roster carries no secrets" (see its case below) — a lead-to-lead body is
* not the roster; it is where leads discuss host shapes, credentials and unmerged work.
* Mapping this to {@code READ} would let any worker read every peer lead's mail in full
* and would silently falsify that comment for every other {@code READ} caller.
*/
COORD_READ,
/** Scrape the metrics endpoint. */
METRICS
METRICS,
/**
* Drive the lead-rollover executor ({@code fleet_handover}: open/confirm/cancel a
* self-replace, fleetd #480 Unit C). Primary-only, same as {@link #SPAWN}/{@link #STOP}/
* {@link #DRAIN} — and, unlike those, the terminal it acts on is never even an argument:
* {@code LeadRollover#open}/{@code #confirm} are always called with the CALLER's own
* connection-resolved terminal (see {@code dev.ltms.fleet.lead.LeadRollover}'s class
* javadoc, fleetd #480 correction 2), so a primary can only ever roll itself.
*/
HANDOVER
}
/**
@@ -50,7 +69,7 @@ public final class Authz {
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN -> caller.isPrimary();
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
@@ -68,6 +87,11 @@ public final class Authz {
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
// is coordination between leads, not observation of the roster.
case COORD_READ -> caller.isPrimary();
};
}
@@ -221,7 +221,7 @@ public final class CallerResolver {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
// escalating a dev, hunter, or reviewer into an architect. Checked before the worker fallback.
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
@@ -244,7 +244,15 @@ public final class CallerResolver {
// already names what happens if that case is handed the primary role: a worker→primary
// escalation. So an unresolved caller is refused (ANONYMOUS — the same clean, already-tested
// "authenticated as nothing" outcome used everywhere else in this method), never promoted.
return isLoopback(remoteAddr) && c.resolved() ? Principal.primary(c.pid()) : Principal.anonymous();
//
// fleetd #505: the OTHER way a real pid can wrongly reach here with a null terminal — not a
// failed lsof lookup, but a herdr error partway through PaneLocator's pane scan. c.resolved()
// says nothing about that; it only tests the lsof sentinel (by design — see
// ConnectionIdentity.Caller#resolved). c.scanComplete() is the separate signal: a scan that
// could not check every pane must not be read as "checked everywhere, no match" — the pane it
// could not check might have been the caller's own. So both must hold before this promotes.
return isLoopback(remoteAddr) && c.resolved() && c.scanComplete()
? Principal.primary(c.pid()) : Principal.anonymous();
}
private boolean presentedTokenMatches(String authorizationHeader) {
@@ -42,7 +42,7 @@ public interface MemberLifecycle {
* Try to bind a newly spawned {@code terminal} into the role it was granted.
*
* @return the role this session actually holds: {@code role} unchanged for a role with no
* live slot-binding semantics (dev, reviewer), or when the bind succeeded; a fallback
* live slot-binding semantics (dev, hunter, reviewer), or when the bind succeeded; a fallback
* role — never {@code role} — when a slot-bound role (architect) could not be bound.
* Callers must record THIS value on the session, never the requested {@code role}, so
* a later roster read never reports a role the session does not hold (CB-619). In
@@ -11,6 +11,7 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.function.Supplier;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
@@ -19,16 +20,44 @@ import java.util.Objects;
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/{@code hunters}/
* {@code reviewers}
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
* re-reads {@code fleet:} on every call, through a supplier the same shape as
* {@code CompositePeerLauncher}'s (see {@code ConfigRef}'s class doc) — so a config reload
* that removes or adds an architect slot governs the <em>next</em> spawn with no restart.
* Only {@link #MemberRegistry(FleetConfig.Fleet)} freezes the pool at construction, and that
* constructor exists for tests and for the (rare) case of wiring a fixed, code-built config.</li>
* <li><b>terminal bindings</b> — owned by this registry, initially <em>empty</em>, and
* <strong>never</strong> touched by a reload. Config declares no architect terminal, so at
* startup every slot is idle and nothing resolves to an architect; a session only becomes one
* when the spawn lifecycle {@linkplain #bind(String, String) binds} its terminal to a slot.
* {@link CallerResolver} reads this through {@link #snapshot()} to turn a pane into an
* {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p><strong>The binding rule (fleetd #424): config governs what a bound slot still grants, as
* well as what may be bound next.</strong> Removing a slot from config revokes it — that is the
* ticket's entire point ("Revoking an architect slot does not revoke it"). Revoking it means an
* architect already bound to that slot loses the ARCHITECT privilege on its very next request:
* {@link #roleForSlot} and {@link #nameForSlot} read {@link #slots()} directly, with no cache, so
* the moment a slot drops out of config, {@link CallerResolver#resolve} (which calls both on every
* request from a bound pane, {@code CallerResolver.java:220}) can no longer confirm the pane's slot
* is an architect slot, and the pane falls through to {@code Principal.worker(...)}. What does
* <em>not</em> change is the {@code terminalToSlot} <em>occupancy</em> — the binding created by
* {@link #bind} is untouched by a reload, on purpose: unbinding it here would double-book the slot
* key (a second terminal could then bind to the "freed" key while the first is still the terminal
* the operator actually meant to demote) and would silently break {@link #unbind}'s compare-safe
* contract, which needs the original {@code terminal → slot} pair intact to remove it cleanly. So
* the demoted session keeps occupying its slot — {@link #slotForTerminal} and {@link #snapshot()}
* still name it — it just no longer resolves as an architect through that occupancy, and a fresh
* spawn still cannot bind to the same key while it is occupied ({@link #reserve}/
* {@link #requireSlotFor} refuse it anyway, since it is gone from {@link #slots()}). The demoted
* session's own turn is unaffected: {@code fleet_reply}'s authorization
* ({@code Authz.Action.REPLY}) is {@code caller.ownsSession(targetSession)} — identity by terminal,
* not by role — so a demoted architect can still end its own turn normally.
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
@@ -55,14 +84,42 @@ public final class MemberRegistry implements MemberLifecycle {
}
}
private final Map<String, Entry> slots;
private final Supplier<FleetConfig.Fleet> fleet;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code terminalToSlot}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Slot keys held between reservation and the terminal binding. Guarded by terminalToSlot. */
private final java.util.Set<String> reservedSlots = new java.util.HashSet<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
/**
* Freeze the pool at construction — for tests, and for the rare case of wiring a fixed,
* code-built config. Production wiring should prefer {@link #live}, which re-reads
* {@code fleet:} on every call.
*/
public MemberRegistry(FleetConfig.Fleet fleet) {
this(() -> fleet);
}
private MemberRegistry(Supplier<FleetConfig.Fleet> fleet) {
this.fleet = fleet;
}
/**
* Live variant (fleetd #424): {@code fleet} is read fresh on every {@link #slots()} call — pass
* {@code () -> config.get().fleet()}, the same supplier shape {@code CompositePeerLauncher}
* already uses for placement — so a reload that adds or removes an architect slot governs the
* next spawn's {@link #reserve}/{@link #requireSlotFor} check with no restart. A separate,
* private constructor rather than a same-arity public overload of
* {@link #MemberRegistry(FleetConfig.Fleet)}: a {@code FleetConfig.Fleet} and a
* {@code Supplier<FleetConfig.Fleet>} overload are ambiguous for a literal {@code null} — the
* same reason {@code CallerResolver.withLeads} is a static factory rather than a fourth
* constructor overload.
*/
public static MemberRegistry live(Supplier<FleetConfig.Fleet> fleet) {
return new MemberRegistry(Objects.requireNonNull(fleet, "fleet"));
}
/** Flatten every role pool in {@code fleet} into one map. Leaders are not members. */
private static Map<String, Entry> flatten(FleetConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
@@ -74,18 +131,22 @@ public final class MemberRegistry implements MemberLifecycle {
});
}
}
this.slots = Collections.unmodifiableMap(flat);
return Collections.unmodifiableMap(flat);
}
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
/**
* The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot of
* {@code fleet:} <em>as of this call</em> — see the class doc for which constructor makes that
* live versus frozen.
*/
public Map<String, Entry> slots() {
return slots;
return flatten(fleet.get());
}
/** The slots belonging to {@code role}, in definition order. */
/** The slots belonging to {@code role}, in definition order, as of this call. */
public Map<String, Entry> slotsFor(MemberRole role) {
Map<String, Entry> out = new LinkedHashMap<>();
slots.forEach((key, e) -> {
slots().forEach((key, e) -> {
if (e.role() == role) {
out.put(key, e);
}
@@ -117,31 +178,49 @@ public final class MemberRegistry implements MemberLifecycle {
}
/**
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
* The strong-model profile a slot runs under, as of this call.
*
* <p>Nothing in {@code src/main} calls this (fleetd #431 — grepped both the {@code
* .profileForSlot(} and the {@code ::profileForSlot} form). This javadoc used to say "what the
* spawn lifecycle reads", and that seam does not exist: the spawn lifecycle takes its profile
* from the {@link MemberLifecycle.SlotReservation} that {@code reserve} returns, never from
* here. Kept and pinned rather than deleted because it is the natural accessor for that seam
* if one is added; live for the same reason as {@link #roleForSlot}, so a reload cannot leave
* it answering for the old config.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return (e == null || e.profile() == null) ? null : e.profile();
}
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
/**
* The role a qualified slot key belongs to, or {@code null} when the key is not currently
* configured. Deliberately live, with no cache (fleetd #424, see the class doc's binding rule):
* removing a slot from config must make {@link CallerResolver#resolve} stop granting the
* ARCHITECT role for it on the very next request from a terminal that was bound to it, which is
* the ticket's whole point — revoking a slot must actually revoke it, not just refuse the next
* spawn.
*/
public MemberRole roleForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return e == null ? null : e.role();
}
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
/**
* The unqualified configured name for a slot, or {@code null} if it is not currently configured.
* Live for the same reason as {@link #roleForSlot} — see the class doc's binding rule.
*/
public String nameForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
return slots().containsKey(slotName);
}
/**
@@ -244,7 +323,7 @@ public final class MemberRegistry implements MemberLifecycle {
* CB-619 / fleetd #123: refuse an architect acquire before anything spawns when no configured
* slot carries {@code profile} — the config-gap case from the original defect report (a spawn
* asked for {@code role=architect, profile=sonnet}, and {@code fleet.architects} carried only
* {@code opus} and {@code sol}). A dev/reviewer acquire is always a no-op: those pools are
* {@code opus} and {@code sol}). A dev/hunter/reviewer acquire is always a no-op: those pools are
* placement candidates only (see {@code CompositePeerLauncher}), never a live identity binding,
* so there is nothing here to refuse — an explicit profile outside the pool for those roles is a
* documented operator override, not a defect.
@@ -11,6 +11,7 @@ import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Consumer;
import java.util.function.Supplier;
/**
@@ -22,65 +23,296 @@ import java.util.function.Supplier;
* choice rather than an accident of where the field was initialised.
*
* <h2>Not every key can change under a running daemon</h2>
* Keys fall into three classes, and the difference is about what already exists when the reload
* Keys fall into four classes, and the difference is about what already exists when the reload
* happens — not about how important the key is.
*
* <ul>
* <li><strong>Hot</strong> — re-read per use, so a reload takes effect on the next spawn:
* {@code fleet:} (every role pool, {@code charters}, and {@code tabLabel}),
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
* them hot — not the fact that they are config. <strong>This does NOT include
* {@code fleet.leaders}</strong>: {@code Fleetd.main} reads {@code cfg.fleet().leaders()}
* once at startup to build the {@code LeadTabScanner} and the {@code LeadLauncher}, and
* neither is reconstructed on reload — so a lead added, removed, or re-{@code tab}'d under
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Both are
* read through a supplier on {@code CompositePeerLauncher}, which is what makes them hot —
* not the fact that they are config. Most of {@code fleet:} — every role pool
* ({@code architects}/{@code developers}/{@code hunters}/{@code reviewers}), {@code charters}, and
* {@code tabLabel} — is read the same live way, through the same supplier
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
* reads it live for placement (which profile an unqualified architect spawn may land on), and
* {@code MemberRegistry} separately reads it live, through its own instance of the same
* supplier shape, for identity — both which slot a spawn may bind to <em>and</em> what a slot
* already bound still grants. Removing an architect slot from config therefore revokes the
* {@link dev.ltms.fleet.auth.Role#ARCHITECT} role on the bound pane's very next request; only
* the slot <em>occupancy</em> survives, so the demoted session still holds its slot key until
* it unbinds. See {@code MemberRegistry}'s class doc for that binding rule.
* <strong>But {@code fleet:} as a whole is NOT in this class</strong>: {@code fleet.leaders}
* inside the same key is frozen, which is exactly what makes {@code fleet:} split rather than
* hot — see below. {@code models:} (fleetd #422) joined this class whole: {@link
* FleetConfig#validateModels()} re-runs fully against the fresh config on every {@link
* #reload()} (via {@link FleetConfig#validateAll()}), refusing a bad edit outright rather than
* caching a stale copy anywhere, and the on/off half added by fleetd #422 is read live both by
* {@code CompositePeerLauncher}'s spawn gate ({@code enforceModelEnabled} and its candidate
* filter) and by {@code fleet_profiles}/{@code GET /profiles} (via
* {@code PeerLauncher.disabledModels()}). Nothing about {@code models:} is baked into an
* object built at startup, so — unlike the deferred keys below — there is no frozen half left
* to report; it moved here from deferred rather than joining split. An existing profile's
* {@code exhaustedPattern} (fleetd #446) joined this class the same way: it used to be
* compiled once into {@code Fleetd.main}'s startup pattern map (see the Deferred bullet's old
* wording, and {@code LiveExhaustedPatterns}'s class doc for the history), and is now read
* live, cached by profile name, by both {@code CompletionResolver}'s classification (via
* {@code LiveExhaustedPatterns.patternFor}) and {@code fleet_profiles}'s {@code
* exhaustionDetectionArmed} (via {@code LiveExhaustedPatterns.armed}) — the one live object
* both read, so a reload that arms or disarms a profile's usage-limit detection takes effect
* on the next check with no restart. {@code errorPattern}, {@code exhaustedPattern}'s sibling
* key for backend-error (not usage-limit) classification, was deliberately left OUT of this
* fleetd #446 change and stays deferred below — the ticket scoped it out explicitly.
* {@code leadRollover:} (fleetd #480) joined this class whole, the same shape as
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
* than capturing them into fields at construction — unlike its closest
* structural cousin {@code leadHeartbeat:}, whose {@code LeadHeartbeatLoop} bakes
* {@code idleAfterNanos}/{@code backoffMs}/{@code quietNudgeCap} into final fields. The one
* restart-only edge is structural, not a stale value: {@code Fleetd.java} decides whether to
* construct the {@code LeadRollover} object at all off the startup snapshot (the same
* presence gate {@code leadHeartbeat:} uses), so a block ADDED where it was absent at boot
* needs a restart before anything exists to call — the same fact already true of adding a
* brand-new {@code profiles:} entry.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code idleSleepGuard:} ({@code Fleetd.java} reads it once, at startup, to decide whether
* to construct an {@code IdleSleepGuard} and wire {@code SessionManager}'s
* {@code onAcquire}/{@code onRelease} hooks to it — neither is rebuilt on reload, so a
* running daemon keeps whatever this was at startup regardless of a later edit),
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code guard:}, {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* {@code guard:}, {@code worktreeRoot:}, {@code worktreeGroup:} and {@code memberSkills:}
* (all three of the latter baked once into the {@code GitWorktrees} built at
* {@code Fleetd.java:251} and never rebuilt — fleetd #323 instance 2 found
* {@code worktreeGroup} missing from this list and from {@link #changedDeferredKeys};
* {@code memberSkills} (fleetd #362) followed the same shape), {@code primary:} (fleetd #326 — {@code Fleetd.java:506, 519,
* 520} read {@code cfg.primary()} only off the startup snapshot to build {@code
* PrimaryRegistry} and size {@code ReplyPushLoop}'s reminder cap/backoff, and neither is
* rebuilt on reload. Say the consequence exactly: {@code primary.terminal} is DEPRECATED
* (CB-532, and {@code Fleetd.java:511} warns about it at startup) — a lead's identity comes
* from {@code leaders:}/{@code leadScan:}, so changing this pin does not demote or promote a
* lead that uses those. What a changed pin still does not take effect on until a restart is
* the fallback nudge destination the pin remains, the deprecated identity path for an operator
* who still relies on it, and {@code pushReminders}/{@code pushBackoffMs}), {@code configReload:} (fleetd #326 — {@code
* Fleetd.java:679-680} read it only at startup to decide whether to build a {@code
* ConfigWatcher} at all and with what interval; the watcher that would apply a later change is
* itself built once, so a running watcher keeps polling on its original enabled flag and
* interval regardless of what a reload changes it to, the same shape as {@code lifecycle} —
* not cold, because no already-open resource goes inconsistent with the new value, the watcher
* (if any) simply keeps its old settings), adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Fleetd.main}'s
* pattern map at startup), {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
* {@code Fleetd.main}'s backend-error pattern map at startup, the same way), and the rest.
* {@code credentialId} (CB-578 stage B) is NOT on
* this list — it is read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
* {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
* {@code Fleetd.main}'s backend-error pattern map at startup; deliberately NOT made hot
* alongside {@code exhaustedPattern} by fleetd #446 — that ticket scoped {@code errorPattern}
* and cooling-off out explicitly),
* {@code ideProjectDir} / {@code ideOpenCommand} / {@code autoCompactWindow} (fleetd #323
* instance 1 — all three are read at spawn off the same frozen profile map and were missing
* from {@link #sameLaunchSettings}), and the rest of {@link #sameLaunchSettings}.
* {@code credentialId} (CB-578 stage B) and {@code exhaustedPattern} (fleetd #446) are NOT on
* this list — both are read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so both are hot instead
* (see the Hot bullet above for {@code exhaustedPattern}'s history).
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
* reload logs these rather than pretending they applied.</li>
* <li><strong>Cold</strong> — cannot change at all under a running daemon: {@code bind:},
* {@code herdrSocket:}, {@code broker:} and {@code auth:}. The socket is bound, the broker
* <li><strong>Split</strong> (fleetd #330; extended to a third key by fleetd #333) — read
* <em>both</em> ways at different sites, so the key does not fit any class above as a whole:
* {@code health:}, {@code coordinator:} and {@code fleet:}. Each is read off the startup
* snapshot to build a long-lived object, and read live off {@link #get()} at a different,
* unrelated site — so half of a reload's effect already applies while the other half waits
* for a restart, and a bare "config reloaded" would under-claim by exactly that half.
* <ul>
* <li>{@code health:} — the monitor itself ({@code enabled}, {@code intervalSeconds},
* {@code workingSuspectAfterSeconds}) is built once at {@code Fleetd.java:556-563}
* and never rebuilt, so a changed value needs a restart to actually start, stop, or
* retime it. The coverage string {@code fleet_profiles} reports
* ({@code Fleetd.java:648-650}) is read live off {@link #get()} on every call, so it
* already reflects the new value.</li>
* <li>{@code coordinator:} — the {@code LeadMailbox} connection ({@code uri},
* {@code uriEnv}, {@code selfId}, {@code prefetch}) is opened once at
* {@code Fleetd.java:502} and never reopened, so a changed value needs a restart —
* {@code selfId} in particular names this daemon's own AMQP inbox queue, and a peer
* lead that learned the old name would not discover a new one on its own. The broker
* URI env-var <em>name</em> that {@code MemberEnvAllowList} keeps out of a member's
* environment is read live off {@link #get()} on every spawn
* ({@code HerdrPeerLauncher.java:1530}), so it already applies.</li>
* <li>{@code fleet:} (fleetd #333) — {@code fleet.leaders} is the frozen half:
* {@code Fleetd.java:281} reads {@code cfg.fleet().leaders()} off the startup snapshot
* for two long-lived objects built right after it and never rebuilt — the
* {@code LeadTabScanner}'s {@code tab label → lead name} map ({@code Fleetd.java:301},
* wired into {@code CallerResolver.withLeadsAndMembers} at {@code Fleetd.java:620/624},
* which is how a caller's pane is recognised as a lead at all) and, when herdr answered
* at startup, {@code LeadLauncher(...).ensureLeads()} ({@code Fleetd.java:315}), which
* auto-launches each declared lead up to its {@code instances} count. So a lead added,
* removed, or given a new {@code tab:} label under {@code fleet.leaders} needs a
* restart — until then it is invisible to identity resolution, and this is exactly the
* scenario fleetd #333 named: an operator edits a lead's {@code tab:} to match a
* renamed pane, sees "config reloaded", and the pane keeps resolving as a worker,
* because {@code CallerResolver} is still matching against the old label. The live
* half is the rest of {@code fleet:} — {@code architects}/{@code developers}/
* {@code reviewers}, {@code charters}, {@code tabLabel} — read live through the same
* {@code CompositePeerLauncher} supplier the Hot bullet above names, so a reload that
* only touches those already applies with nothing to report. Because the hot and frozen
* halves of {@code fleet:} are disjoint sub-fields rather than the same fields read two
* ways (contrast {@code coordinator.uriEnv} above), {@link #changedSplitKeys} compares
* {@code fleet.leaders} alone, not the whole {@code Fleet} record — comparing the whole
* record would report "split" for a {@code tabLabel}-only change that is actually fully
* hot, over-claiming in exactly the direction this class exists to avoid under-claiming
* in.</li>
* </ul>
* A split change is still accepted — {@link Outcome#applied()} stays {@code true}, the same
* as a deferred change — because the live half genuinely took effect; refusing the whole
* reload would leave the operator worse off than today. {@link Outcome#split()} names the
* key and says which half is which each time, rather than trying to score "how changed" a
* mixed key is or handle "both halves changed in one reload" as a special case.</li>
* <li><strong>Cold</strong> — cannot change at all under a running daemon. All five of
* {@link #COLD_KEYS}: {@code bind:}, {@code herdrSocket:}, {@code memberHerdrSocket:},
* {@code broker:} and {@code auth:}. The sockets are already connected, the broker
* connection is open, and the auth mode decides who may reach the port that is already
* listening.</li>
* listening. This bullet omitted {@code memberHerdrSocket:} until fleetd #333 — say "all
* five of COLD_KEYS" rather than re-listing them, so prose and set cannot drift again.</li>
* </ul>
*
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333);
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, again after
* {@code models:} was added as deferred, again for fleetd #422, which moved {@code models:}
* from deferred to hot-excluded once its on/off half was read live everywhere, and again after
* {@code leadRollover:} was added (fleetd #480).</strong>
* {@code FleetConfig} has 26 top-level record components: 5 cold, 13 deferred, 3 split, 5
* hot-excluded. Five of them are named nowhere in this file, and the reason is the same for all
* five: {@code placement}, {@code memberCredentials}, {@code memberLoginShell}, {@code models} and
* {@code leadRollover} are <strong>hot</strong> and correctly absent — all five are read live off
* {@code config.get()} (placement through the {@code CompositePeerLauncher} supplier the Hot bullet
* names; {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198,
* 205, 729} and {@code HerdrPeerLauncher#configuredMemberLoginShell}; {@code models} the same way,
* through the Hot bullet's {@code models:} paragraph; {@code leadRollover} through the Hot bullet's
* {@code leadRollover:} paragraph), so a reload takes effect on the next spawn (or, for
* {@code models}, the next reported status; for {@code leadRollover}, the next {@code open()}/
* {@code confirm()} call) with no entry needed here.
* {@code health} and {@code coordinator} used to be a third kind — <strong>undecided</strong>, not
* hot — until fleetd #330 added the <strong>split</strong> class above and gave them a home. A
* reload touching either used to report a bare "config reloaded", which under-claimed; now it names
* the key and says which half is which. {@code fleet} was the same story in reverse: fleetd #330's
* own fact-find named {@code health}/{@code coordinator} as "the complete set of split-shaped keys"
* and filed {@code fleet.leaders}'s restart requirement as a documented caveat sitting in the
* <strong>hot-excluded</strong> escape hatch instead — correctly documented, but in the one bucket
* this file's own coverage test cannot check the truth of (see that test's javadoc). fleetd #333
* moved it into <strong>split</strong>, where {@link #changedSplitKeys} actually reports it.
* <p>The point of writing the count down: "not mentioned in this file" looks identical for a key
* that is correctly hot and for a key nobody triaged. Three times now — {@code worktreeGroup} (#323),
* {@code primary}/{@code configReload} (#326), and {@code fleet.leaders} sitting in the escape hatch
* (#333) — the second kind hid among the first. A top-level coverage checker in the
* {@code ConfigRefProfileCoverageTest} shape (one level up, over {@code FleetConfig} itself rather
* than {@code FleetConfig.Profile}) proves this file's four classes exhaust the record's components
* — see {@code ConfigRefTopLevelCoverageTest}. That test proves the record's <em>shape</em> is fully
* triaged; it does NOT prove a {@code SPLIT_KEYS}/{@code COLD_KEYS}/{@code DEFERRED_KEYS} member has
* any reporting code behind it at all — {@code ConfigRefTopLevelReportingCoverageTest} is what
* fleetd #333 added for that, after measuring that a {@code SPLIT_KEYS} entry with its reporting
* branch deleted passes both this file's own "kept in step" assert and
* {@code ConfigRefTopLevelCoverageTest} unchanged. fleetd #337 extended it to {@code DEFERRED_KEYS}
* after measuring the same one-way gap there directly: dropping {@code guard}'s branch out of
* {@link #changedDeferredKeys} while {@code "guard"} stayed in the set left the whole suite green.
*
* <p><strong>A cold change refuses the whole reload.</strong> Not the hot half applied and the cold
* half warned about: that would leave the running daemon in a state matching no file on disk, which
* is the worst thing a reload can do to an operator debugging one. Refusing keeps the invariant that
* the live config is always some version of the file, and the message names the keys that must
* change through a restart.
* change through a restart. A split change does <em>not</em> refuse, for a different reason than a
* deferred change does not: its live half genuinely took effect, so refusing would throw that away
* and leave the operator worse off than the partial-but-honest report {@link Outcome#split()} gives.
*
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*
* <p><strong>fleetd #474</strong> — {@link FleetConfig#validateAll()} is not the only gate startup
* runs before a config takes effect: {@code Fleetd.main} also calls {@code
* dev.ltms.fleet.mcp.CharterToolSurface#assertChartersNameOnlyRegisteredTools}, right after {@code
* cfg.validateAll()}, to refuse a charter that names an MCP tool the server does not register. That
* check cannot live inside {@link FleetConfig} — {@code CharterToolSurface} lives in the {@code mcp}
* package because the canonical tool set ({@code FleetTool}) does, and config is loaded before the
* MCP server exists, so {@code FleetConfig} must not gain a dependency on {@code mcp}. {@link
* #reload} cannot import {@code mcp} either, for the same reason applied one layer up: {@code
* dev.ltms.fleet.config} is loaded before {@code dev.ltms.fleet.mcp} exists, same as {@code
* FleetConfig}. So this class accepts the check as a {@code Consumer<FleetConfig>} —
* {@link #extraValidation} — supplied by whichever caller already sits at the seam that holds both
* a loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}. It is invoked
* inside the same try/catch as {@code fresh.validateAll()}, so a charter that would have refused to
* boot refuses a reload too, and keeps the running config exactly like any other {@code
* validateAll()} failure. A ref built through the two-argument constructor (every test fixture that
* does not care about this check, and {@link #fixed}) gets a no-op consumer, so nothing outside
* {@code Fleetd.main} needs to know this hook exists.
*/
public final class ConfigRef implements Supplier<FleetConfig> {
private static final Logger log = LoggerFactory.getLogger(ConfigRef.class);
/** Keys that cannot change under a running daemon — see the class doc. */
private static final Set<String> COLD_KEYS =
/**
* Keys that cannot change under a running daemon — see the class doc.
*
* <p>Package-private (not {@code private}) so {@code ConfigRefTopLevelCoverageTest} can fold it
* into the top-level triage it checks, the same way it reads {@link #SPLIT_KEYS}.
*/
static final Set<String> COLD_KEYS =
Set.of("bind", "herdrSocket", "memberHerdrSocket", "broker", "auth");
/**
* Keys read BOTH off the startup snapshot and live off {@link #get()} at different sites, so
* neither the hot, deferred nor cold class fits them as a whole — see the class doc's Split
* bullet (fleetd #330). A changed split key is accepted ({@link Outcome#applied()} stays
* {@code true}) and reported by name, with a message naming which half is live and which needs
* a restart.
*/
static final Set<String> SPLIT_KEYS = Set.of("health", "coordinator", "fleet");
/**
* Top-level keys {@link #changedDeferredKeys} compares — see the class doc's Deferred bullet.
* Promoted here from a test-side copy in {@code ConfigRefTopLevelCoverageTest} by fleetd #337,
* the same reason {@link #COLD_KEYS} and {@link #SPLIT_KEYS} live here rather than in a test: a
* second, hand-maintained copy of this set is exactly the kind of thing that silently drifts
* from the method it is supposed to describe. {@code spawnReadyTimeoutMs} and
* {@code spawnReadyPollMs} are compared together in one branch and reported under the combined
* label {@code "spawnReady*"}; {@code profiles} is compared twice over (added/removed names,
* then an existing profile's launch settings) — see {@link #changedDeferredKeys}.
*
* <p>Package-private (not {@code private}) so {@code ConfigRefTopLevelCoverageTest} and
* {@code ConfigRefTopLevelReportingCoverageTest} can both read it, the same way they already
* read {@link #COLD_KEYS} and {@link #SPLIT_KEYS}.
*/
static final Set<String> DEFERRED_KEYS = Set.of(
"guard", "worktreeRoot", "worktreeGroup", "memberSkills", "primary", "configReload",
"leadHeartbeat", "lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs",
"quarantineCooldownSeconds", "profiles", "idleSleepGuard");
private final Path path;
private final AtomicReference<FleetConfig> current;
private final Consumer<FleetConfig> extraValidation;
/** Equivalent to the three-argument constructor with a no-op {@code extraValidation}. */
public ConfigRef(Path path, FleetConfig initial) {
this(path, initial, cfg -> { });
}
/**
* @param extraValidation run on every {@link #reload} candidate, inside the same try/catch as
* {@code fresh.validateAll()} — see the class doc's fleetd #474 note.
* {@code Fleetd.main} passes {@code
* Fleetd::assertChartersNameOnlyRegisteredTools} (a package-private
* {@code FleetConfig -> void} adapter over {@code
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), so a reload
* runs the same gate startup does without this class depending on the
* {@code mcp} package.
*/
public ConfigRef(Path path, FleetConfig initial, Consumer<FleetConfig> extraValidation) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
this.extraValidation = Objects.requireNonNull(extraValidation, "extraValidation");
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
@@ -102,25 +334,37 @@ public final class ConfigRef implements Supplier<FleetConfig> {
/**
* What a reload attempt did.
*
* <p>{@code split} is a separate field from {@code deferred} rather than a differently-worded
* entry inside it, because the two carry different guarantees for any caller that branches on
* them rather than just printing {@link #summary()}: every {@code deferred} entry means "this
* key's whole change waits for a restart", while every {@code split} entry means "part of this
* key's change already applied, and the message says which part" — collapsing them would force
* a caller to re-parse the message to tell those apart. See the class doc's Split bullet
* (fleetd #330) for why the key needs this at all.
*
* @param applied true when the new config is now live
* @param coldKeys cold keys whose value changed, which is why an unapplied reload was refused
* @param deferred keys that changed and were accepted, but whose effect waits for a restart
* @param deferred keys that changed and were accepted, but whose effect waits entirely on a
* restart
* @param split split keys that changed and were accepted, each named with which half of it
* is already live and which half waits for a restart
* @param error the parse or validation failure that refused the reload, else {@code null}
*/
public record Outcome(boolean applied, List<String> coldKeys, List<String> deferred,
String error) {
List<String> split, String error) {
public Outcome {
coldKeys = List.copyOf(coldKeys);
deferred = List.copyOf(deferred);
split = List.copyOf(split);
}
static Outcome refusedCold(List<String> keys) {
return new Outcome(false, keys, List.of(), null);
return new Outcome(false, keys, List.of(), List.of(), null);
}
static Outcome failed(String error) {
return new Outcome(false, List.of(), List.of(), error);
return new Outcome(false, List.of(), List.of(), List.of(), error);
}
/** A one-line summary for the operator — the reason, not just the verdict. */
@@ -132,11 +376,18 @@ public final class ConfigRef implements Supplier<FleetConfig> {
return "config reload refused — these keys cannot change under a running daemon: "
+ String.join(", ", coldKeys) + ". Restart fleetd to apply them.";
}
if (!deferred.isEmpty()) {
return "config reloaded; these changes need a restart to take effect: "
+ String.join(", ", deferred);
if (deferred.isEmpty() && split.isEmpty()) {
return "config reloaded";
}
return "config reloaded";
StringBuilder out = new StringBuilder("config reloaded");
if (!deferred.isEmpty()) {
out.append("; these changes need a restart to take effect: ")
.append(String.join(", ", deferred));
}
if (!split.isEmpty()) {
out.append("; partially live — ").append(String.join(" | ", split));
}
return out.toString();
}
}
@@ -157,12 +408,18 @@ public final class ConfigRef implements Supplier<FleetConfig> {
fresh = FleetConfig.load(path);
// The same gate startup runs. A config that would have refused to boot must not be able
// to slip in through a reload — that is how a daemon ends up in a state it could never
// have started in, which is the hardest kind to debug.
fresh.validateAuthExposure();
fresh.validateLeadTabPrefixes();
fresh.validateSubscriptionProfiles();
fresh.validateCharters();
fresh.validateMembers();
// have started in, which is the hardest kind to debug. fleetd ticket "central allow-list
// of usable models" follow-up: this used to be six individual validateXxx() calls, and
// mutation testing found two of the six unpinned here even though startup pinned nothing
// at all — see FleetConfig#validateAll's javadoc for why the fix is one reflective call,
// not a longer hand-maintained list.
fresh.validateAll();
// fleetd #474: validateAll() does not cover everything startup refuses on — the charter
// tool-surface check (Fleetd.main, right after cfg.validateAll()) lives outside
// FleetConfig on purpose (see this class's doc) and is supplied here as extraValidation.
// Same try/catch as validateAll() above, on purpose: either failure must refuse the whole
// reload and keep the running config the same way.
extraValidation.accept(fresh);
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
@@ -177,14 +434,21 @@ public final class ConfigRef implements Supplier<FleetConfig> {
}
List<String> deferred = changedDeferredKeys(old, fresh);
List<String> split = changedSplitKeys(old, fresh);
current.set(fresh);
Outcome out = new Outcome(true, List.of(), deferred, null);
Outcome out = new Outcome(true, List.of(), deferred, split, null);
log.info(out.summary());
return out;
}
/** Cold keys whose value differs between the running config and the candidate. */
private static List<String> changedColdKeys(FleetConfig old, FleetConfig fresh) {
/**
* Cold keys whose value differs between the running config and the candidate.
*
* <p>Package-private (not {@code private}) so {@code ConfigRefTopLevelReportingCoverageTest}
* can call it directly with a reflection-built {@code FleetConfig} pair, the same reason
* {@link #sameLaunchSettings} is package-private — see that test's class doc (fleetd #333).
*/
static List<String> changedColdKeys(FleetConfig old, FleetConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.bind(), fresh.bind())) {
changed.add("bind");
@@ -206,8 +470,16 @@ public final class ConfigRef implements Supplier<FleetConfig> {
return changed;
}
/** Changed keys that were accepted but whose effect waits for a restart. */
private static List<String> changedDeferredKeys(FleetConfig old, FleetConfig fresh) {
/**
* Changed keys that were accepted but whose effect waits for a restart.
*
* <p>Package-private (not {@code private}) so {@code ConfigRefTopLevelReportingCoverageTest}
* can call it directly with a reflection-built {@code FleetConfig} pair, the same reason
* {@link #changedColdKeys} and {@link #changedSplitKeys} already are (fleetd #333, extended to
* this method by fleetd #337 — membership in {@link #DEFERRED_KEYS} proved nothing about this
* method on its own until then; see that test's class doc).
*/
static List<String> changedDeferredKeys(FleetConfig old, FleetConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.lifecycle(), fresh.lifecycle())) {
changed.add("lifecycle");
@@ -221,6 +493,44 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.worktreeRoot(), fresh.worktreeRoot())) {
changed.add("worktreeRoot");
}
// Baked into the same GitWorktrees as worktreeRoot (Fleetd.java:251) and never rebuilt
// either — see the class doc. Missing this check was fleetd #323 instance 2: a reload
// that changed only worktreeGroup reported "config reloaded" with nothing deferred, and
// newly provisioned worktrees kept the old sharing behaviour.
if (!Objects.equals(old.worktreeGroup(), fresh.worktreeGroup())) {
changed.add("worktreeGroup");
}
// fleetd #362: baked into the same GitWorktrees as worktreeRoot/worktreeGroup
// (Fleetd.java:251) and never rebuilt either — a reload that changes only memberSkills
// must be reported the same way, or a newly provisioned worktree keeps seeding from (or
// skipping) the old source directory with nothing telling the operator why.
if (!Objects.equals(old.memberSkills(), fresh.memberSkills())) {
changed.add("memberSkills");
}
// fleetd #326: Fleetd.java:506, 519, 520 read cfg.primary() only off the startup snapshot
// (PrimaryRegistry's pinned terminal, ReplyPushLoop's reminder cap and backoff) — neither is
// rebuilt on reload, so a changed value needs a restart. Note what it does NOT mean:
// primary.terminal is deprecated (CB-532), identity comes from leaders:/leadScan:, so a lead
// using those is unaffected by this pin either way. See the class doc for the exact scope.
if (!Objects.equals(old.primary(), fresh.primary())) {
changed.add("primary");
}
// fleetd #326: Fleetd.java:679-680 read cfg.configReload() only at startup to decide whether
// to build a ConfigWatcher at all and with what interval — the watcher that would apply a
// later change is itself built once, so a running watcher keeps its original enabled flag and
// interval regardless of what a reload changes it to. Not cold: no already-open resource goes
// inconsistent with the new value, a watcher (if any) simply keeps polling on the old settings.
if (!Objects.equals(old.configReload(), fresh.configReload())) {
changed.add("configReload");
}
// Fleetd.java reads cfg.idleSleepGuard() once, at startup, to decide whether to construct
// an IdleSleepGuard at all and wire SessionManager's onAcquire/onRelease hooks to it —
// neither is rebuilt on reload, so a running daemon keeps whatever this was at startup
// (armed or not) regardless of a later edit here. Not cold: nothing already-open goes
// inconsistent with the new value, an armed-or-not guard just keeps its original answer.
if (!Objects.equals(old.idleSleepGuard(), fresh.idleSleepGuard())) {
changed.add("idleSleepGuard");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
@@ -265,14 +575,114 @@ public final class ConfigRef implements Supplier<FleetConfig> {
}
/**
* Whether two versions of a profile would launch a peer identically. Compares every component
* the launcher reads at spawn; {@code weight}, {@code maxLoad} and {@code credentialId} are
* excluded because those are read live (by the placement policy and, for credentialId, by
* {@code CompositePeerLauncher}/the CB-578 stage B exhaustion sink) and really do take effect on
* the next spawn.
* Split keys whose value differs between the running config and the candidate — see the class
* doc's Split bullet (fleetd #330, extended for {@code fleet:} by fleetd #333). Unlike
* {@link #changedDeferredKeys}, this does not try to tell which sub-field moved for {@code
* health:} or {@code coordinator:}: any change to either gets the same fixed message, because
* the message already names both halves every time, so there is no "which half changed"
* question left for the caller to answer. {@code fleet:} is different on purpose — see below.
*
* <p>Package-private (not {@code private}) so {@code ConfigRefTopLevelReportingCoverageTest}
* can call it directly with a reflection-built {@code FleetConfig} pair, the same reason
* {@link #sameLaunchSettings} is package-private — see that test's class doc (fleetd #333). That
* test exists because membership in {@link #SPLIT_KEYS} proves nothing about this method on its
* own: fleetd #333 measured that dropping the {@code coordinator} branch out of this method
* while leaving {@code "coordinator"} in {@code SPLIT_KEYS} left the whole suite green except a
* hand-written {@code ConfigRefTest} case — neither {@code ConfigRefTopLevelCoverageTest} (it
* only reads the set) nor the "kept in step" assert below (it only checks the reported keys are
* a SUBSET of {@code SPLIT_KEYS}, never that every {@code SPLIT_KEYS} member has a branch here)
* would have caught it.
*/
private static boolean sameLaunchSettings(FleetConfig.Profile a, FleetConfig.Profile b) {
return Objects.equals(a.baseUrl(), b.baseUrl())
static List<String> changedSplitKeys(FleetConfig old, FleetConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.health(), fresh.health())) {
changed.add("health: the monitor itself (enabled, interval, workingSuspectAfter) is "
+ "frozen at startup and needs a restart; the coverage status fleet_profiles "
+ "reports is read live and already applied");
}
if (!Objects.equals(old.coordinator(), fresh.coordinator())) {
changed.add("coordinator: the LeadMailbox connection (uri, uriEnv, selfId, prefetch) is "
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
+ "member's environment is read live on every spawn and already applied");
}
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, hunters, reviewers,
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
// its consumers, not just the one this comment used to name: CompositePeerLauncher reads it
// live for PLACEMENT through the () -> config.get().fleet() supplier named in the class doc's
// Hot bullet, and MemberRegistry separately reads it live for IDENTITY (which slot a spawn
// may bind to, AND what a slot already bound still grants) through its own instance of that
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
// LeadLauncher(...).ensureLeads() at Fleetd.java:315, which auto-launches each lead up to its
// `instances` count; neither is rebuilt on reload). So this compares fleet.leaders alone, not
// the whole Fleet record: comparing the whole record would report "split" for a tabLabel-only
// or architects-only change that is actually fully hot, which is the over-claim mirror of the
// under-claim bug this class exists to prevent.
if (!Objects.equals(leadersOf(old), leadersOf(fresh))) {
changed.add("fleet: fleet.leaders (each lead's tab, workspace, cwd, profile and "
+ "instances count) is read once at startup to build the LeadTabScanner's "
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
+ "lead; the rest of fleet: (developers, hunters, reviewers, charters, tabLabel) is read "
+ "live through the supplier on CompositePeerLauncher, and architects is read "
+ "live through that same supplier for placement AND through a separate supplier "
+ "on MemberRegistry for spawn-time identity — both already applied");
}
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
// every message here must be traceable to one of the split keys the class doc documents.
// NOTE what this does NOT prove, per the javadoc above: it does not catch a SPLIT_KEYS
// member with no branch above at all, only a branch whose message is mis-worded relative to
// the set. ConfigRefTopLevelReportingCoverageTest is what proves the former.
assert changed.stream().allMatch(m -> SPLIT_KEYS.stream().anyMatch(k -> m.startsWith(k + ":")))
: "a split entry was reported that does not start with a SPLIT_KEYS name: " + changed;
return changed;
}
/** {@code cfg.fleet().leaders()}, defensively, in case a caller hands in a non-defaulted config. */
private static Map<String, FleetConfig.Leader> leadersOf(FleetConfig cfg) {
return cfg.fleet() == null ? Map.of() : cfg.fleet().leaders();
}
/**
* {@link FleetConfig.Profile} record components deliberately left out of
* {@link #sameLaunchSettings} because they are read <em>live</em>, not baked in at spawn — see
* the class doc's <em>Hot</em> bullet. {@code weight} and {@code maxLoad} are read live by the
* placement policy on every spawn; {@code credentialId} is read live by
* {@code CompositePeerLauncher} and the CB-578 stage B exhaustion sink; {@code exhaustedPattern}
* (fleetd #446) is read live, cached by profile name, by {@code LiveExhaustedPatterns} — see
* that class's doc and the class doc's <em>Hot</em> bullet for the history (it used to be
* compared here, deferred, like its sibling {@code errorPattern} still is). Nothing else is
* excluded — see {@code sameLaunchSettingsComparesEveryProfileComponentOrExcludesIt} in
* {@code ConfigRefProfileCoverageTest}, which enumerates every {@code Profile} record component
* by reflection and fails the build if one is neither compared below nor named here.
*/
static final Set<String> LAUNCH_SETTINGS_EXCLUDED =
Set.of("weight", "maxLoad", "credentialId", "exhaustedPattern");
/**
* Whether two versions of a profile would launch a peer identically.
*
* <p>This must compare every {@link FleetConfig.Profile} record component except the three in
* {@link #LAUNCH_SETTINGS_EXCLUDED}. That is not a claim this javadoc can make good on by
* itself — a javadoc saying "compares every component" is exactly what fleetd #323 found to be
* false for three fields (and a sibling method's field list, for a fourth). The actual
* guarantee comes from {@code ConfigRefProfileCoverageTest}: it enumerates every record
* component of {@code FleetConfig.Profile} by reflection, mutates each one not in
* {@code LAUNCH_SETTINGS_EXCLUDED} on a base profile, and asserts this method reports a
* difference — so a new component that is neither compared here nor added to
* {@code LAUNCH_SETTINGS_EXCLUDED} (with a reason) fails that test by name, rather than
* silently reporting "config reloaded" for a value the daemon never picked up.
*/
static boolean sameLaunchSettings(FleetConfig.Profile a, FleetConfig.Profile b) {
return Objects.equals(a.profile(), b.profile())
&& Objects.equals(a.baseUrl(), b.baseUrl())
&& Objects.equals(a.model(), b.model())
&& Objects.equals(a.configDir(), b.configDir())
&& Objects.equals(a.tokenEnv(), b.tokenEnv())
@@ -291,13 +701,25 @@ public final class ConfigRef implements Supplier<FleetConfig> {
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription())
// CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern())
// fleetd #446: exhaustedPattern moved to LAUNCH_SETTINGS_EXCLUDED — it is now read
// live, cached by profile name, through LiveExhaustedPatterns (see that class's doc
// and ConfigRef's class doc Hot bullet), so it must NOT be compared here any more: a
// reload that changes only exhaustedPattern must report "config reloaded", not
// "these changes need a restart".
// fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error
// pattern map at startup (see BackendErrorPatternLookup wiring), the same way
// exhaustedPattern is — a reload never re-reads it either.
&& Objects.equals(a.errorPattern(), b.errorPattern());
// pattern map at startup (see BackendErrorPatternLookup wiring) — unlike its sibling
// exhaustedPattern above (fleetd #446), a reload still never re-reads it; fleetd #446
// scoped errorPattern out on purpose (see the class doc's Hot bullet).
&& Objects.equals(a.errorPattern(), b.errorPattern())
// fleetd #323 instance 1: ideProjectDir and ideOpenCommand are read at spawn off the
// same frozen profile map as ideMcpUrl above (ClaudeCodeLauncher.java:267/269,
// OpenCodeLauncher.java:474/480/486) and were missing from this comparison.
&& Objects.equals(a.ideProjectDir(), b.ideProjectDir())
&& Objects.equals(a.ideOpenCommand(), b.ideOpenCommand())
// fleetd #323 instance 1: autoCompactWindow is read at spawn the same way
// (ClaudeCodeLauncher.java:926, OpenCodeLauncher.java:650). Comparing it here only
// makes the reload REPORT that a restart is needed — it deliberately does not make
// autoCompactWindow take effect live, which is a separate, larger change.
&& Objects.equals(a.autoCompactWindow(), b.autoCompactWindow());
}
}
@@ -16,11 +16,15 @@ import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.lang.reflect.InvocationTargetException;
import java.lang.reflect.Method;
import java.lang.reflect.Modifier;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Collections;
import java.util.Comparator;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
@@ -58,7 +62,7 @@ import java.util.regex.PatternSyntaxException;
* @param fleet who the daemon may run and under which role (CB-557). One block replacing
* the former {@code leaders:}, {@code members:}, {@code leadScan:} and
* {@code defaultProfile:}. Role is the containing key — {@code leaders},
* {@code architects}, {@code developers}, {@code reviewers} — and each entry
* {@code architects}, {@code developers}, {@code hunters}, {@code reviewers} — and each entry
* names the {@code profiles:} backend it runs on. See {@link Fleet}
* @param leadHeartbeat opt-in idle-lead heartbeat (CB-551); {@code null} ⇒ off, and an upgraded
* daemon never nudges an idle lead on its own initiative
@@ -66,12 +70,14 @@ import java.util.regex.PatternSyntaxException;
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
* @param quarantineCooldownSeconds the BASE cooldown a credential is quarantined for after a
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
* needs a restart, and a quarantine already running keeps whatever cooldown was
* live when it started.
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Since fleetd #466 this is
* only the first occurrence's length — a credential quarantined again shortly
* after this cooldown ends backs off further, up to a ceiling; see {@code
* BackendQuarantine}'s class doc. Baked once into the {@code BackendQuarantine}
* built at startup, so it is DEFERRED: changing it needs a restart, and a
* quarantine already running keeps whatever cooldown was live when it started.
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
* credentials a spawned member's pane inherits. {@code null} (the block
* omitted) blocks nothing — see {@link MemberCredentials}.
@@ -106,6 +112,36 @@ import java.util.regex.PatternSyntaxException;
* When {@code memberHerdrSocket} is NOT configured this field is never
* consulted at all; fleetd keeps reading its own {@code $SHELL}, exactly as
* before this field existed.
* @param memberSkills fleetd #362: nullable directory of skill folders (each a subdirectory
* holding a {@code SKILL.md}, the same shape as this repo's own {@code
* .claude/skills/}) copied into every provisioned worktree's {@code
* .claude/skills/}, so a member spawned against ANY repo — not only one that
* already ships its own copy — can load a bridge skill such as {@code
* implementer}. {@code null}/blank ⇒ off: no worktree is touched beyond
* today's behaviour. A skill folder the target repo already carries is never
* overwritten — see {@link dev.ltms.fleet.session.GitWorktrees}. Claude Code
* members only; an opencode member's equivalent lives under a different path
* ({@code .opencode/agent}) and is not covered by this key. Every non-hidden
* subdirectory of this directory is copied wholesale, with no per-file
* allowlist — do not park scratch files or drafts alongside the real skill
* folders, they will be copied into every provisioned worktree too.
* @param idleSleepGuard opt-in-by-default: hold an OS-level assertion against idle sleep while at
* least one member is live, so an unattended host does not sleep out from
* under a member's long turn. {@code null} (the block omitted) behaves the
* same as an explicit {@code enabled: true}; set {@code enabled: false} to
* turn it off. See {@link dev.ltms.fleet.power.IdleSleepGuard}.
* @param models central allow-list of models any {@code profiles:} entry may name. {@code
* null} or an empty {@code allow:} ⇒ off: {@link #validateModels()} checks
* nothing and every existing config keeps working exactly as it does today.
* When non-empty, a profile whose {@code model:} is not one of {@link
* Models#ids()} fails config load, naming both the model and the profile —
* see {@link #validateModels()}. This block decides what may be CONFIGURED;
* fleetd #422 added the separate on/off question — whether a configured model
* may be spawned onto RIGHT NOW ({@link Models.ModelEntry#enabled}) — enforced
* live at spawn by {@code CompositePeerLauncher}, not here. See {@link Models}.
* @param leadRollover opt-in lead rollover (fleetd #480): {@code null} ⇒ off, and no
* {@code dev.ltms.fleet.lead.LeadRollover} is constructed at all — an upgraded
* daemon never clears a lead's pane on its own initiative. See {@link LeadRollover}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
@@ -130,7 +166,68 @@ public record FleetConfig(
MemberCredentials memberCredentials,
Coordinator coordinator,
String worktreeGroup,
String memberLoginShell) {
String memberLoginShell,
String memberSkills,
IdleSleepGuard idleSleepGuard,
Models models,
LeadRollover leadRollover) {
/** Back-compat form before the {@code leadRollover:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard,
Models models) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, idleSleepGuard, models, null);
}
/** Back-compat form before the {@code models:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, idleSleepGuard, null);
}
/** Back-compat form before the {@code idleSleepGuard:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, null);
}
/** Back-compat form before the {@code memberSkills} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, null, null);
}
/** Back-compat form before the {@code memberLoginShell} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
@@ -141,7 +238,7 @@ public record FleetConfig(
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup, null);
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup, null, null);
}
/** Back-compat form before the {@code worktreeGroup} key was added. */
@@ -348,6 +445,12 @@ public record FleetConfig(
* {@code null}/blank ⇒ the classification never fires for this profile and
* today's completion-fallback behaviour is unchanged. Every backend words
* its refusal differently, so this is config, never a vendor string in code.
* Read live off the current config through {@code LiveExhaustedPatterns},
* cached by profile name (fleetd #446), so it is HOT: an operator can arm
* or disarm this profile's usage-limit detection by editing this key and
* reloading, with no restart. Before fleetd #446 this was compiled once at
* daemon startup, the same deferred shape its sibling {@code errorPattern}
* (below) still has.
* @param credentialId the shared account this profile authenticates as (CB-578 stage B). Two
* or more profiles setting the <em>same</em> non-blank value are quarantined
* together by one {@code exhaustedPattern} classification on any one of
@@ -888,12 +991,18 @@ public record FleetConfig(
* {@code null} ⇒ kept as {@code null} (no self id configured).
* @param prefetch the consumer's {@code basicQos} prefetch count. {@code null}/non-positive ⇒
* {@link LeadMailbox#DEFAULT_PREFETCH}.
* @param peers fleetd #361: the coord-ids the operator declares as this daemon's peers — the
* daemon never guesses who else exists. {@code fleet_list} reports each one's
* live reachability. Blank entries are dropped; {@code null} ⇒ an empty list, so
* a config written before this field existed still parses unchanged.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Coordinator(String uri, String uriEnv, String selfId, Integer prefetch) {
public record Coordinator(String uri, String uriEnv, String selfId, Integer prefetch, List<String> peers) {
public Coordinator {
selfId = (selfId == null || selfId.isBlank()) ? null : selfId;
peers = peers == null ? List.of()
: peers.stream().filter(p -> p != null && !p.isBlank()).toList();
}
/** True when a {@code uriEnv} is configured by name, whether or not its variable resolves. */
@@ -1091,6 +1200,7 @@ public record FleetConfig(
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
* @param architects profiles the {@code architect} role may run on
* @param developers profiles the {@code dev} role may run on
* @param hunters profiles the {@code hunter} role may run on
* @param reviewers profiles the {@code reviewer} role may run on
* @param charters optional launch-charter text keyed by singular role wire name
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
@@ -1101,6 +1211,7 @@ public record FleetConfig(
public record Fleet(Map<String, Leader> leaders,
Map<String, Slot> architects,
Map<String, Slot> developers,
Map<String, Slot> hunters,
Map<String, Slot> reviewers,
Map<String, String> charters,
String tabLabel) {
@@ -1117,6 +1228,7 @@ public record FleetConfig(
leaders = unmodifiableOrEmpty(leaders);
architects = unmodifiableOrEmpty(architects);
developers = unmodifiableOrEmpty(developers);
hunters = unmodifiableOrEmpty(hunters);
reviewers = unmodifiableOrEmpty(reviewers);
charters = unmodifiableOrEmpty(charters);
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? DEFAULT_TAB_LABEL : tabLabel;
@@ -1131,9 +1243,15 @@ public record FleetConfig(
* constructor: the launcher reads {@code fleet.charters()} from the live config. Jackson
* binds the canonical constructor, so this one cannot swallow an operator's YAML.
*/
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
Map<String, Slot> developers, Map<String, Slot> reviewers,
Map<String, String> charters, String tabLabel) {
this(leaders, architects, developers, null, reviewers, charters, tabLabel);
}
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
Map<String, Slot> developers, Map<String, Slot> reviewers, String tabLabel) {
this(leaders, architects, developers, reviewers, null, tabLabel);
this(leaders, architects, developers, null, reviewers, null, tabLabel);
}
/**
@@ -1155,6 +1273,7 @@ public record FleetConfig(
return switch (role) {
case ARCHITECT -> architects;
case DEV -> developers;
case HUNTER -> hunters;
case REVIEWER -> reviewers;
};
}
@@ -1220,6 +1339,104 @@ public record FleetConfig(
}
}
/**
* Opt-in lead rollover (fleetd #480): a lead that decides it is ready to be replaced writes a
* handover file, then asks fleetd to clear its own pane and bootstrap a fresh session against
* that file. Config + a pure decision/verification layer only — see
* {@code dev.ltms.fleet.lead.LeadRollover} for the executor this block feeds, and
* {@code fleet_*} tool wiring is a later ticket.
*
* <p>Deliberately opt-in ({@code null} ⇒ off, exactly like {@code leadHeartbeat:}): absent this
* block, {@code Fleetd.java} never constructs a {@code LeadRollover} object at all, so an
* upgraded daemon cannot silently acquire the ability to clear the lead's own pane. Even once
* present, nothing but an explicit {@code confirm()} call can ever roll a pane — there is no
* timer, heartbeat or timeout anywhere in this feature that fires one on its own; see that
* class's javadoc.
*
* <p><strong>Hot, not deferred</strong> (see {@code ConfigRef}'s class doc): every field below
* is read live, through a {@code Supplier<LeadRollover>} the same {@code () -> config.get().x()}
* shape {@code fleet}/{@code placement}/{@code models} already use, so an edit to any of the
* six fields takes effect on the very next {@code open()}/{@code confirm()} call once the block
* has been present since startup — nothing here is captured into a frozen field the way {@code
* leadHeartbeat}'s {@code idleAfterNanos}/{@code backoffMs}/{@code quietNudgeCap} are. The one
* restart-only edge left is structural, not a stale value: {@code Fleetd.java} decides whether
* to construct the {@code LeadRollover} object at all off the startup snapshot, the same gate
* {@code leadHeartbeat:} uses, so a block ADDED where it was absent at boot needs a restart
* before anything exists to call — the same fact already true of adding a brand-new
* {@code profiles:} entry.
*
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
* {@code confirm()} validates every gate and schedules the roll. {@code
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
* that SAME pane to report an injectable state again — i.e. for the calling turn to actually
* end — before it sends {@code /clear} at all. If that wait times out, no {@code /clear} is
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
* bounds the SECOND wait, for the pane to re-settle AFTER {@code /clear} has already gone out.
*
* @param handoverPath required when this block is present — where the handover file a fresh
* lead session reads must live. There is no sane non-null default for an
* operator-specific path, so a present block with a {@code null}/blank
* {@code handoverPath} is refused at config load; see
* {@link #validateLeadRollover()}. May be relative: {@code
* dev.ltms.fleet.lead.LeadRollover#open} resolves a relative path against
* the CALLING lead's configured {@code fleet.leaders.<name>.cwd}, falling
* back to the daemon's own working directory ({@code
* System.getProperty("user.dir")}) when that lead has none configured —
* never against whatever directory the daemon process happens to have been
* started in for its own sake. An absolute path is used unchanged.
* @param requireOperatorConfirm default {@code true} — {@code confirm()} refuses unless the
* caller also passes {@code operatorConfirmed: true}. Set {@code false} to
* let the three handover-file checks alone gate the roll.
* @param maxDocAgeSeconds default 3600 — refuse a handover file whose modified time is older
* than this many seconds, so a stale leftover from an earlier rollover
* attempt can never be mistaken for a fresh one.
* @param turnSettleSeconds default 20 — bound on how long the deferred roll waits for the
* CALLING lead's own turn to end (its pane to report injectable again)
* before sending {@code /clear} at all. See the paragraph above.
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
* report an injectable state again after {@code /clear} before giving up. A
* roll that times out here never sends {@code bootstrapText}.
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
* lead's pane once it settles after {@code /clear}, telling the fresh
* session where to read the handover and carry on. Left {@code null} here
* when the operator configures none: the default sentence cannot be built
* at construction time because it must name the path AFTER {@code
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code
* handoverPath} against the calling lead's workspace, which this record has
* no way to know — see {@link #bootstrapTextFor(String)}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
Integer maxDocAgeSeconds, Integer turnSettleSeconds,
Integer clearSettleSeconds, String bootstrapText) {
public LeadRollover {
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 20 : turnSettleSeconds;
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
}
/**
* The text actually sent to the lead's pane once it settles after {@code /clear}: the
* operator's configured {@link #bootstrapText} when one is set, otherwise the default
* sentence built from {@code resolvedHandoverPath}.
*
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
* #open} already resolved — never the raw configured {@link
* #handoverPath}, which may still be relative and would name a
* directory the fresh lead session's own pane cannot resolve
*/
public String bootstrapTextFor(String resolvedHandoverPath) {
return bootstrapText != null
? bootstrapText
: "Fresh lead session: read the handover file at " + resolvedHandoverPath
+ " and carry on from there.";
}
}
/**
* Watch {@code fleetd.yaml} and re-read it when it changes (CB-559).
*
@@ -1245,6 +1462,145 @@ public record FleetConfig(
}
}
/**
* Hold an OS-level assertion against idle sleep while at least one member is live (see
* {@link dev.ltms.fleet.power.IdleSleepGuard}).
*
* <p>Unlike most opt-in blocks in this file, this one defaults to <em>on</em>: an unattended
* host idle-sleeping mid-turn is a correctness problem (a dropped AMQP link, a frozen member),
* not a convenience, so the safer default is armed. An operator who wants the previous
* behaviour (no assertion held, ever) sets {@code enabled: false} explicitly.
*
* @param enabled {@code false} turns the guard off; {@code null} (the block omitted
* entirely) or {@code true} leaves it on
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record IdleSleepGuard(Boolean enabled) {
public boolean isEnabled() {
return !Boolean.FALSE.equals(enabled);
}
}
/**
* Central allow-list of models any {@code profiles:} entry may name (fleetd ticket: "a central
* allow-list of usable models"). Nothing before this block checked a profile's {@code model:}
* against anything — it was a free-form string handed straight to the backend adapter, and a
* withdrawn or misspelled name failed silently (opencode falls back to a default model rather
* than erroring on an unknown {@code -m}) rather than at config load, where a mistake is cheap.
*
* <p><b>Absent or empty {@code allow:} is "off"</b>, on purpose: this is a large deployment
* with gitignored {@code fleetd.yaml} on more than one host, and a change that forced every
* operator to enumerate their models before the daemon would start would break every one of
* them on upgrade. See {@link FleetConfig#validateModels()}, which is where the allow-list is
* actually enforced, at config load.
*
* <p><b>The list is the authority; a {@code profiles:} entry is checked against it, never the
* other way around.</b> Adding or editing a {@code profiles:} entry cannot, by itself, widen
* the set of permitted models — only editing {@code models.allow:} itself can. This is the
* invariant the ticket asked for: the two blocks are validated in one direction only.
*
* <p><b>Spawn-time enforcement (fleetd #422, units 2+3) lives outside this record</b> —
* {@code CompositePeerLauncher.enforceModelEnabled} and its candidate-set filter read this
* block LIVE (through the same kind of supplier {@code weight}/{@code maxLoad} already use),
* so the on/off state below is hot: no restart needed. This block itself still only decides
* what may be CONFIGURED (membership in {@link #allow}); {@link ModelEntry#enabled} decides
* whether a member of that list is currently spawnable. The two questions are deliberately
* separate — see {@link ModelEntry}'s javadoc for why turning a model off must never mean
* removing it from {@link #allow}.
*
* @param allow the permitted models, each its own {@link ModelEntry} rather than a bare
* string — see that record's javadoc for why. {@code null}/empty ⇒ the block is
* treated as absent: {@link FleetConfig#validateModels()} checks nothing.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Models(List<ModelEntry> allow) {
public Models {
allow = (allow == null) ? List.of() : List.copyOf(allow);
}
/**
* One permitted model, named as a record rather than a bare string on purpose: this ticket
* (fleetd #422) is the "later unit" the original comment here predicted — it hangs an on/off
* state ({@link #enabled}) off each entry, and a bare {@code List<String>} could not have
* grown that field without changing the YAML shape underneath every operator who already
* wrote one. {@link #model()} is intentionally a single flat, opaque-string namespace — a
* bare Claude id ({@code claude-sonnet-5}) and an opencode provider-prefixed id
* ({@code openai/gpt-5.6-terra}) both fit it unchanged, because {@link
* FleetConfig#validateModels()} only ever compares a profile's {@code model:} value against
* this string for exact equality; it never parses a provider prefix or branches on a
* profile's {@code kind:}.
*
* @param model the model id exactly as a {@code profiles:} entry's {@code model:} field
* would name it
* @param enabled {@code false} turns spawning onto this model off; {@code null} (the field
* omitted — every config written before fleetd #422 is this shape) or
* {@code true} leaves it on. Turning a model off must NEVER remove it from
* {@link Models#allow} — {@link FleetConfig#validateModels()} checks
* <em>membership</em> only, never the on/off state, so an off entry stays a
* valid thing for a {@code profiles:} entry to name; only
* {@code CompositePeerLauncher}'s spawn-time gate reads {@link #enabled}.
* Collapsing the two — turning a model off by deleting its {@code allow:}
* entry — would make {@link FleetConfig#validateModels()} refuse the whole
* config reload the moment a still-configured profile names it, which is
* exactly the restart-to-flip-a-switch problem this field exists to avoid.
* <p>A model id can be named by more than one {@code profiles:} entry (e.g.
* {@code deepseek-v4-flash} backs both {@code local} and {@code
* local-direct} in the live config) — turning it off disables every profile
* that names it, on purpose: the model is what a subscription's rate limit
* actually constrains, not any one profile alias for it.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record ModelEntry(String model, Boolean enabled) {
public ModelEntry {
model = (model == null || model.isBlank()) ? null : model.trim();
}
/**
* Back-compat form before {@link #enabled} was added (fleetd #422) — the model is
* unconditionally on, exactly as every {@code ModelEntry} behaved before this field
* existed. Keeps pre-#422 call sites (and any YAML that omits {@code enabled:})
* compiling and behaving identically.
*/
public ModelEntry(String model) {
this(model, null);
}
/** {@code true} unless {@link #enabled} is explicitly {@code false} — absent means on. */
public boolean isEnabled() {
return !Boolean.FALSE.equals(enabled);
}
}
/** {@link #allow}'s model ids, as a set for membership checks. Blank/null entries are dropped. */
public Set<String> ids() {
Set<String> ids = new java.util.LinkedHashSet<>();
for (ModelEntry e : allow) {
if (e != null && e.model() != null) {
ids.add(e.model());
}
}
return Collections.unmodifiableSet(ids);
}
/**
* Model ids currently turned off (fleetd #422: {@link ModelEntry#isEnabled()} {@code
* false}). Read live by {@code CompositePeerLauncher}'s spawn gate and by {@code
* fleet_profiles}/{@code GET /profiles} — both must read this same accessor off the same
* live config so the two surfaces cannot disagree about which model is off (the fleetd
* #404 lesson: a status field must read the source the behaviour reads).
*/
public Set<String> offIds() {
Set<String> off = new java.util.LinkedHashSet<>();
for (ModelEntry e : allow) {
if (e != null && e.model() != null && !e.isEnabled()) {
off.add(e.model());
}
}
return Collections.unmodifiableSet(off);
}
}
/**
* The terminal → lead-name map seeded from the legacy singular {@code primary:} pin (CB-530).
*
@@ -1324,7 +1680,7 @@ public record FleetConfig(
* <p>A herdr pane runs a login shell that re-sources the operator's own secret store, so a
* member inherits every credential the operator's shell holds — measured at 31 names on this
* host, of which only one ({@code GITEA_ACCESS_TOKEN}) used to be blocked, and that block was a
* single name hardcoded in {@link HerdrPeerLauncher} rather than driven by config (gitea issue
* single name hardcoded in {@link dev.ltms.fleet.member.HerdrPeerLauncher} rather than driven by config (gitea issue
* #82). This record replaces that hardcoded shadow with a config-driven one.
*
* <p><b>deny-by-default, not a deny-list.</b> A deny-list (block these specific names, let
@@ -1495,7 +1851,8 @@ public record FleetConfig(
"bind", "herdrSocket", "memberHerdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell");
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills",
"idleSleepGuard", "models", "leadRollover");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
@@ -1525,7 +1882,7 @@ public record FleetConfig(
/** The {@code fleet:} child blocks whose direct children are slot names. */
private static final Set<String> FLEET_POOL_KEYS =
Set.of("leaders", "architects", "developers", "reviewers");
Set.of("leaders", "architects", "developers", "hunters", "reviewers");
/**
* Reject a {@code fleet:} role pool whose slot names repeat (CB-548, re-homed by CB-557).
@@ -1535,7 +1892,7 @@ public record FleetConfig(
* daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by
* default, so duplicates are caught here, at parse time, before the map is built.
*
* <p>Only the four pools <em>directly under the top-level {@code fleet:}</em> are considered,
* <p>Only the five pools <em>directly under the top-level {@code fleet:}</em> are considered,
* and only their direct child keys (the slot names). A nested field elsewhere, even one also
* named {@code developers:}, is ignored, so parsing of the rest of the config is unaffected.
*
@@ -1703,7 +2060,7 @@ public record FleetConfig(
+ " and that role's pool supplies the candidate profiles",
"architects", "'fleet.architects'",
"members", "a role pool under 'fleet:' — 'fleet.architects', 'fleet.developers' or"
+ " 'fleet.reviewers'; the role is the containing key, not a 'role:' field",
+ " 'fleet.hunters' or 'fleet.reviewers'; the role is the containing key, not a 'role:' field",
"leaders", "'fleet.leaders'",
"leadScan", "'fleet.leaders.<name>.tabPrefix' and '.scanIntervalSeconds' — lead"
+ " discovery is now configured on the lead it discovers");
@@ -2172,9 +2529,27 @@ public record FleetConfig(
// memberLoginShell is left as-is (fleetd #213), like worktreeGroup: null/blank is "not
// configured", and there is no sane non-null default — a member's login shell is
// operator-specific and only meaningful when memberHerdrSocket is also set.
// memberSkills is left as-is (fleetd #362), like worktreeGroup/memberLoginShell: null/blank
// is "off", and there is no sane non-null default — the daemon may not even run from a
// checkout that ships its own .claude/skills/.
// idleSleepGuard is left as-is, like leadHeartbeat/configReload above, but for the opposite
// reason: it is on by default already (its own isEnabled() treats null the same as
// enabled: true — see its javadoc), so defaulting the block here would change nothing a
// reader observes and would only obscure that "block omitted" and "block present and
// enabled" are deliberately the same outcome.
// models is left as-is, like broker/primary/coordinator above: null/empty is "off", and an
// absent block must validate nothing (see Models's javadoc) — defaulting it here to an
// empty Models would be a no-op for validateModels() either way, since an empty allow-list
// already means "check nothing", so there is nothing to gain and one more null check to
// avoid by leaving it exactly as configured.
// leadRollover is left as-is, like leadHeartbeat above: null is "off", and LeadRollover's
// own compact constructor defaults the fields of a block that IS present. Defaulting it
// here would construct a LeadRollover object (via Fleetd.java's presence gate) for every
// config that never mentioned it.
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell);
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills,
idleSleepGuard, models, leadRollover);
}
/**
@@ -2257,6 +2632,26 @@ public record FleetConfig(
+ "lead tabs cannot be confused.");
}
/**
* Reject a present {@code leadRollover:} block with no (or a blank) {@code handoverPath}
* (fleetd #480). There is no sane non-null default for an operator-specific file path, unlike
* every other field on {@link LeadRollover}, which {@link LeadRollover}'s own compact
* constructor already defaults — so this is the one field that must be refused at load rather
* than silently defaulted to something that would never match a real handover file.
*
* @throws IllegalStateException when {@code leadRollover} is present but {@code handoverPath}
* is {@code null} or blank
*/
public void validateLeadRollover() {
if (leadRollover == null) {
return;
}
if (leadRollover.handoverPath() == null || leadRollover.handoverPath().isBlank()) {
throw new IllegalStateException("refusing to start: leadRollover.handoverPath is "
+ "required when leadRollover: is present.");
}
}
/** Case-insensitive prefix test that tolerates a null/blank label. */
private static boolean startsWithIgnoreCase(String label, String prefix) {
if (label == null || prefix == null || prefix.isBlank()) {
@@ -2394,6 +2789,118 @@ public record FleetConfig(
}
}
/**
* Reject a {@code profiles:} entry whose {@code model:} is not on the configured {@link
* #models} allow-list.
*
* <p>Absent or empty {@code models.allow:} validates nothing — see {@link Models}'s javadoc:
* an existing config with no such block must keep working exactly as it does today. Once the
* operator declares at least one entry, every profile's {@code model:} (when set — a profile
* may legitimately leave it {@code null}, e.g. a {@code subscription: true} profile relying on
* the account's own default) must equal one of {@link Models#ids()} exactly. The comparison is
* a flat string match: a bare Claude id and an opencode {@code provider/model} id are both
* just opaque strings here, so nothing here needs to know which {@code kind:} a profile runs.
*
* <p>The check runs one direction only, by construction: it reads {@link #profiles} and
* {@link #models}, and only ever adds to {@code bad} when a profile's model is missing from
* the allow-list. Nothing here can be satisfied by widening a {@code profiles:} entry — only
* editing {@code models.allow:} itself changes what passes. That is the invariant the ticket
* asked for: the list is the authority, profiles are checked against it.
*
* @throws IllegalStateException when any profile names a model outside the configured
* allow-list, naming both the model and the profile that wanted it
*/
public void validateModels() {
if (models == null || models.allow().isEmpty()) {
return;
}
Set<String> allowed = models.ids();
List<String> bad = new ArrayList<>();
profiles.forEach((name, p) -> {
String model = p.model();
if (model != null && !model.isBlank() && !allowed.contains(model)) {
bad.add("profile '" + name + "' names model '" + model + "', which is not in "
+ "models.allow: (have: " + allowed + ").");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
/**
* Runs every validator this class declares — found by reflection, not by name.
*
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
* that although each of the six validators above was well pinned on its own, nothing proved
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
* per caller; a hand-maintained list of six names here would have the exact same defect its
* own javadoc would warn against: the seventh validator someone adds next month has no reason
* to be added to it. So this method does not name any validator. It sweeps {@link
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
* "validate"} (excluding itself) and invokes every one it finds, via {@link
* #invokeAllValidators}. A new {@code validateXxx()} method is therefore wired into both
* callers the moment it is written — there is no second step to forget, and so no state in
* which it silently never runs.
*
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
* the six individually — see the comments at those two call sites for why
* each must run it.
*
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
* always names the same one first, on every run.
*
* @throws IllegalStateException (or whatever unchecked exception a validator itself throws),
* propagated unchanged from the first validator, in that order,
* that finds a problem
*/
public void validateAll() {
invokeAllValidators(this);
}
/**
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
* seventh validator to this class just to exercise that claim. See {@code
* FleetConfigValidateAllTest} for that proof.
*
* @param target an object whose public, no-argument, {@code void} methods named {@code
* validateXxx} (any name starting with {@code "validate"}, excluding {@code
* validateAll} itself) should all run, in alphabetical-by-name order
*/
static void invokeAllValidators(Object target) {
List<Method> methods = new ArrayList<>();
for (Method m : target.getClass().getMethods()) {
if (Modifier.isPublic(m.getModifiers())
&& m.getParameterCount() == 0
&& m.getReturnType() == void.class
&& m.getName().startsWith("validate")
&& !m.getName().equals("validateAll")) {
methods.add(m);
}
}
methods.sort(Comparator.comparing(Method::getName));
for (Method m : methods) {
try {
m.invoke(target);
} catch (InvocationTargetException e) {
Throwable cause = e.getCause();
if (cause instanceof RuntimeException re) {
throw re;
}
if (cause instanceof Error err) {
throw err;
}
throw new IllegalStateException("validator " + m.getName() + " failed", cause);
} catch (IllegalAccessException e) {
throw new IllegalStateException("cannot invoke validator " + m.getName(), e);
}
}
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
@@ -38,15 +38,46 @@ public final class FleetHealthMonitor {
*/
static final long ASK_LAPSE_RECHECK_DELAY_SECONDS = 120;
/**
* fleetd #386: {@code System.nanoTime()} (or whatever {@link #clock} is) does not advance while
* macOS sleeps, so a raw {@code nowNanos - lastActivityAtNanos} comparison freezes with the
* host and can never cross {@link #workingSuspectAfterNanos}. This is a second, wall-clock
* source used ONLY inside the stall check ({@link #stallElapsedNanos}) to detect and correct
* for that freeze. Nothing else in this class reads it — every other decision (readiness grace,
* the fault classification itself) stays exactly on {@link #clock}, as the ticket requires.
*/
private static final LongSupplier DEFAULT_REALTIME_CLOCK =
() -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis());
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
private final MessageService messages;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final LongSupplier realtimeClock;
private final long intervalSeconds;
private final long tickIntervalNanos;
private final long workingSuspectAfterNanos;
private final BiConsumer<String, String> failTarget;
private final Map<String, HealthPrior> priors = new HashMap<>();
/**
* fleetd #386 clock-drift bookkeeping. {@code haveClockBaseline}/{@code lastTickMonoNanos}/
* {@code lastTickRealNanos} track the previous tick's pair of readings so each new tick can
* measure how far the two clocks moved apart since then. {@code accumulatedDriftNanos} is the
* running total of every such divergence observed since this monitor started (never decreases —
* the monotonic clock can only lag real time, never lead it). {@code busyDriftBaselineNanos}/
* {@code busyBaselineActivityNanos} record, per target, the value of {@code accumulatedDriftNanos}
* at the moment this monitor first saw that target's CURRENT {@code lastActivityAtNanos} while
* BUSY — so {@link #stallElapsedNanos} adds back only the drift observed DURING this BUSY span,
* never drift from a sleep that happened before the member went busy. All five fields are touched
* only from {@code tick()}, like {@link #priors}.
*/
private boolean haveClockBaseline = false;
private long lastTickMonoNanos;
private long lastTickRealNanos;
private long accumulatedDriftNanos = 0;
private final Map<String, Long> busyDriftBaselineNanos = new HashMap<>();
private final Map<String, Long> busyBaselineActivityNanos = new HashMap<>();
/**
* The live classification per member, and the only one of this class's three maps that more
* than one scheduler task touches. {@code tick} writes it (and prunes it to the roster);
@@ -88,12 +119,29 @@ public final class FleetHealthMonitor {
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
long workingSuspectAfterSeconds, BiConsumer<String, String> failTarget) {
this(agents, roster, messages, scheduler, clock, DEFAULT_REALTIME_CLOCK, intervalSeconds,
workingSuspectAfterSeconds, failTarget);
}
/**
* @param realtimeClock fleetd #386: a wall-clock nanosecond source (e.g.
* {@code System.currentTimeMillis()} converted to nanos) that keeps
* advancing while {@code clock} is frozen by a host sleep. Used only to
* correct the stall check — see the class-level javadoc on the
* clock-drift fields.
*/
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, LongSupplier realtimeClock,
long intervalSeconds, long workingSuspectAfterSeconds,
BiConsumer<String, String> failTarget) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.realtimeClock = Objects.requireNonNull(realtimeClock, "realtimeClock");
this.intervalSeconds = intervalSeconds;
this.tickIntervalNanos = TimeUnit.SECONDS.toNanos(intervalSeconds);
this.workingSuspectAfterNanos = TimeUnit.SECONDS.toNanos(workingSuspectAfterSeconds);
this.failTarget = Objects.requireNonNull(failTarget, "failTarget");
}
@@ -123,6 +171,7 @@ public final class FleetHealthMonitor {
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
long nowNanos = clock.getAsLong();
long driftBeforeThisTick = observeClockDrift(nowNanos);
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
@@ -133,7 +182,7 @@ public final class FleetHealthMonitor {
&& session.state() != MemberSession.State.SPAWNING;
boolean readinessGraceElapsed = nowNanos - session.spawnedAtNanos() >= READINESS_GRACE_NANOS;
boolean stalled = session.state() == MemberSession.State.BUSY
&& nowNanos - session.lastActivityAtNanos() >= workingSuspectAfterNanos;
&& stallElapsedNanos(session, nowNanos, driftBeforeThisTick) >= workingSuspectAfterNanos;
// CB-643: the three message-layer facts CB-640 published. Read them here rather than
// leaving them false — that constant is what made 8 of the 9 fault states dead.
boolean queuedDelivery = messages.hasQueuedDelivery(session.terminalId());
@@ -151,6 +200,8 @@ public final class FleetHealthMonitor {
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
orphanStreaks.keySet().retainAll(current);
busyDriftBaselineNanos.keySet().retainAll(current);
busyBaselineActivityNanos.keySet().retainAll(current);
} catch (Throwable error) {
// Any unclassified collection failure must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
@@ -175,6 +226,65 @@ public final class FleetHealthMonitor {
return streak >= ORPHAN_CONFIRM_TICKS;
}
/**
* fleetd #386: compare this tick's monotonic and real-time readings against the previous
* tick's, and fold any positive divergence into {@link #accumulatedDriftNanos} (a ratchet — it
* never decreases, since the monotonic clock can only fall behind real time, never ahead of
* it). Logs once, at WARN, when that single tick's divergence exceeds one full tick interval —
* the signature of a host that slept between the two ticks (a tick literally cannot run while
* the process itself is suspended, so the whole sleep duration lands inside one tick's gap).
*
* @return {@link #accumulatedDriftNanos} as it stood BEFORE this tick's divergence was folded
* in — the baseline {@link #stallElapsedNanos} needs when a target is observed BUSY
* for the first time this tick, so a sleep that happened before this member went busy
* is not attributed to it.
*/
private long observeClockDrift(long nowNanos) {
long nowRealNanos = realtimeClock.getAsLong();
long driftBeforeThisTick = accumulatedDriftNanos;
if (haveClockBaseline) {
long monoDelta = nowNanos - lastTickMonoNanos;
long realDelta = nowRealNanos - lastTickRealNanos;
long tickDrift = realDelta - monoDelta;
if (tickDrift > tickIntervalNanos) {
log.warn("fleet health: the monotonic clock did not advance for about {}s that the "
+ "real clock did since the last tick (host likely slept); the stall "
+ "detector could not see that time", TimeUnit.NANOSECONDS.toSeconds(tickDrift));
}
if (tickDrift > 0) {
accumulatedDriftNanos = driftBeforeThisTick + tickDrift;
}
}
lastTickMonoNanos = nowNanos;
lastTickRealNanos = nowRealNanos;
haveClockBaseline = true;
return driftBeforeThisTick;
}
/**
* fleetd #386: {@code nowNanos - lastActivityAtNanos} alone freezes across a host sleep, since
* both come from the monotonic {@link #clock}. This adds back the real-time drift observed
* since this BUSY span started — not the monitor's whole lifetime, so a sleep that happened
* before this member went busy never leaks into its stall reading (see the class-level javadoc
* on the drift fields). The baseline resets whenever {@code lastActivityAtNanos} changes (a new
* turn) or the member is not currently BUSY.
*/
private long stallElapsedNanos(MemberSession session, long nowNanos, long driftBeforeThisTick) {
String target = session.terminalId();
if (session.state() != MemberSession.State.BUSY) {
busyDriftBaselineNanos.remove(target);
busyBaselineActivityNanos.remove(target);
return nowNanos - session.lastActivityAtNanos();
}
Long baselineActivity = busyBaselineActivityNanos.get(target);
if (baselineActivity == null || baselineActivity != session.lastActivityAtNanos()) {
busyBaselineActivityNanos.put(target, session.lastActivityAtNanos());
busyDriftBaselineNanos.put(target, driftBeforeThisTick);
}
long driftSinceBusyStart = accumulatedDriftNanos - busyDriftBaselineNanos.get(target);
return (nowNanos - session.lastActivityAtNanos()) + driftSinceBusyStart;
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
@@ -3,7 +3,7 @@ package dev.ltms.fleet.herdr;
import com.fasterxml.jackson.databind.JsonNode;
/**
* Client face onto the herdr daemon (protocol 14, herdr 0.7.0).
* Client face onto the herdr daemon (protocol 19, herdr 0.8.0).
*
* <p>This is the ONLY thing in {@code fleetd} that speaks to herdr. Every method
* maps to a herdr JSON-RPC call over its Unix domain socket. Requests are
@@ -8,7 +8,7 @@ import com.fasterxml.jackson.databind.node.ObjectNode;
import java.nio.charset.StandardCharsets;
/**
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 14).
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 19).
*
* <p>Split out from the socket so the framing rules — the ones that actually bit us
* during the spike (id MUST be a string; response carries {@code result} or
@@ -5,6 +5,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.Locale;
import java.util.Map;
@@ -53,17 +54,40 @@ import java.util.function.Supplier;
* daemon and found again by this scan. The trust direction above is unaffected — fleetd writing a
* name for a lead it just started is not a pane promoting itself — but <em>staleness</em> becomes
* real: a label left behind by a session that has since died would read as a live lead forever.
* This scanner does not solve that (its job is naming, and a stale name costs nothing here); the
* launcher does, by requiring a running agent in the tab before it counts the lead as live. If you
* ever make a decision that <em>removes</em> something based on this map, add the same check.
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code FleetConfig.validateLeadTabPrefixes} rather than documented here.
*
* <p><strong>fleetd #359 — the staleness check this class used to skip.</strong> This used to say
* "a stale name costs nothing here" and leave liveness to {@code LeadLauncher}, on the theory that
* naming and removing are different decisions. That was wrong: {@code dev.ltms.fleet.msg.LeadCoordLoop}
* makes exactly the kind of removal decision the old javadoc warned about, by reading this map to
* pick which pane a peer message goes into — and a stale entry there is not free. On a host where
* {@code fleetd} had restarted more than once, a labelled-but-dead tab from a previous life was
* reported right alongside the live one; {@code LeadCoordLoop.resolveLocalLead()} saw more than one
* candidate and refused to guess (safe), but the fix the daemon's own WARN suggests — name a lead
* after {@code coordinator.selfId} — stops being safe once two tabs can share a label: step 1 of
* that resolution picks whichever matching entry it finds first, which can be the dead one, and
* typing a peer's message into a dead shell does not fail — it is silently gone instead of merely
* held. {@link #scan()} now cross-checks every labelled tab against {@code agent.list} (the same
* signal {@code LeadLauncher.countLeads} already trusts for the same purpose) and drops any tab
* with no agent running in it, so a dead tab is never in the map for a caller to pick at all.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
* the previous answer instead of emptying it — a herdr hiccup must not silently demote a live lead
* mid-session.
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan (herdr
* throws) keeps the previous answer instead of emptying it — a herdr hiccup must not silently
* demote a live lead mid-session.
*
* <p><strong>fleetd #359 review, finding 2 — a successful-but-wrong scan is the same hazard.</strong>
* The catch above only fires when a call throws. It does nothing for a call that returns 200 with an
* incomplete answer — exactly what the ticket's own evidence showed {@code agent.list} can do. Once
* this class started trusting that signal, an empty read would otherwise get cached as fact and
* silently drop a lead {@code CallerResolver} had, until then, correctly resolved — turning it into a
* {@code Role.WORKER}, which refuses every orchestration call. So a terminal this class already
* reported as live is not dropped the first time {@code agent.list} loses it: {@link #scan()} grants
* it one grace scan (see {@code gracedTerminals}) and only drops it if a <em>later</em> scan still
* finds no agent. A terminal never reported live before gets no grace — that would weaken the
* original #359 fix itself, which this class's own test suite already pins.
*/
public final class LeadTabScanner implements Supplier<Map<String, String>> {
@@ -79,6 +103,15 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
private long scannedAtNanos;
private boolean everScanned;
/**
* Terminals currently on their one grace scan: {@code cached} reported them live, the most
* recent {@link #scan()} found no agent for them, and they were re-included anyway. Cleared for
* a terminal the instant it is seen live again; a terminal still here on the <em>next</em> scan
* is finally dropped. Scoped separately from {@link #cached} so a graced terminal cannot renew
* its own grace forever just by staying in the exposed map (fleetd #359 review, finding 2).
*/
private Set<String> gracedTerminals = Set.of();
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
@@ -141,7 +174,7 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
return cached;
}
/** One full pass: labelled tabs → their panes → those panes' terminals. */
/** One full pass: labelled tabs → live agents in them → those panes' terminals. */
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
@@ -158,17 +191,47 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
if (!nameByTab.isEmpty()) {
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String name = nameByTab.get(p.path("tab_id").asText(null));
String terminal = p.path("terminal_id").asText(null);
if (name != null && terminal != null && !terminal.isBlank()) {
byTerminal.put(terminal, name);
}
if (nameByTab.isEmpty()) {
gracedTerminals = Set.of();
return Map.of();
}
// fleetd #359: a labelled tab is only a lead when herdr also reports a running agent in
// it — the same liveness signal LeadLauncher.countLeads trusts for the identical purpose.
// Without this, a tab left behind by a session that has since died reads as live forever.
Set<String> tabsWithAgent = new HashSet<>();
for (JsonNode a : herdr.call("agent.list").path("agents")) {
String tabId = a.path("tab_id").asText(null);
if (tabId != null) {
tabsWithAgent.add(tabId);
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
Set<String> stillGraced = new HashSet<>();
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String tabId = p.path("tab_id").asText(null);
String name = nameByTab.get(tabId);
String terminal = p.path("terminal_id").asText(null);
if (name == null || terminal == null || terminal.isBlank()) {
continue;
}
if (tabsWithAgent.contains(tabId)) {
byTerminal.put(terminal, name);
continue;
}
// No agent reported for this tab, but its tab/pane are still here — this is the
// ambiguous case review finding 2 named: a successful agent.list that came back short
// does not prove the lead is dead. Grant one grace scan to a terminal we had already
// reported as live; a terminal we never reported live gets none, so the original #359
// fix (a genuinely dead tab is never reported) is unaffected for the common case.
if (cached.containsKey(terminal) && !gracedTerminals.contains(terminal)) {
byTerminal.put(terminal, name);
stillGraced.add(terminal);
}
}
gracedTerminals = stillGraced;
return Collections.unmodifiableMap(byTerminal);
}
@@ -177,12 +240,14 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
*
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix.
* configured lead just because it shares a prefix. The match strips a trailing
* {@link PendingCloseMarker} first, so a tab {@code LeadLauncher} has flagged as maybe-dead but
* not yet closed keeps resolving normally while that reconcile is pending.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
return tabToName.get(label.strip().toLowerCase(Locale.ROOT));
return tabToName.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
}
}
@@ -1,6 +1,8 @@
package dev.ltms.fleet.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.LinkedHashSet;
import java.util.List;
@@ -36,6 +38,8 @@ import java.util.Set;
*/
public final class PaneLocator {
private static final Logger log = LoggerFactory.getLogger(PaneLocator.class);
/**
* Bound on how many ancestor generations {@link #ancestorsOf} walks. This runs on every MCP
* call, so a cycle or a pathologically deep process tree must not hang identity resolution;
@@ -73,22 +77,46 @@ public final class PaneLocator {
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane on any searched daemon owns it (e.g. the caller is the
* primary, or off-host).
* The outcome of a {@link #terminalForPid} scan: the {@code terminal_id} of the agent pane
* whose process tree contains the pid ({@link #terminal} is {@code null} if none matched),
* and whether the scan that produced that answer ran to completion on every daemon searched.
*
* <p>{@link #complete} is {@code false} exactly when some {@code pane.process_info} call
* failed and, despite that, no pane was ever found to own the pid. In that case a {@code null}
* {@link #terminal} means "could not tell", not "definitely not a worker" — fleetd #505: a
* transient herdr error on the very pane that <em>does</em> own the caller's pid must not read
* as a clean negative and fall through to {@code Principal.primary}, the same way #317's
* {@code Caller.resolved()} already guards a failed lsof lookup. Callers ({@code
* ConnectionIdentity}, {@code CallerResolver}) must refuse rather than promote on an incomplete
* scan.
*
* <p>When a pane genuinely owns the pid, {@link #complete} is {@code true} regardless of
* whether some other, unrelated pane failed to answer earlier in the same scan — a positive
* match is definitive and does not need every pane to have been checked (a pane that "vanished
* mid-scan" but was never the match is still a clean, complete result).
*/
public String terminalForPid(long pid) {
public record Lookup(String terminal, boolean complete) {
private static final Lookup NOT_FOUND = new Lookup(null, true);
}
/**
* Resolve {@code pid} to the agent pane whose process tree contains it, across every searched
* herdr daemon. See {@link Lookup} for how to read a {@code null} terminal.
*/
public Lookup terminalForPid(long pid) {
if (pid <= 0) {
return null;
return Lookup.NOT_FOUND;
}
Set<Long> ancestry = ancestorsOf(pid);
for (HerdrClient herdr : herdrs) {
String terminal = terminalForPid(herdr, ancestry);
if (terminal != null) {
return terminal;
boolean complete = true;
for (int i = 0; i < herdrs.size(); i++) {
Lookup outcome = scan(herdrs.get(i), i, herdrs.size(), ancestry);
if (outcome.terminal() != null) {
return outcome; // a definite match — no need to finish checking other clients
}
complete = complete && outcome.complete();
}
return null;
return new Lookup(null, complete);
}
/**
@@ -117,31 +145,51 @@ public final class PaneLocator {
return ancestry;
}
private static String terminalForPid(HerdrClient herdr, Set<Long> ancestry) {
/** Whether a pane owns one of the scanned pid's ancestors, or the check of it failed outright. */
private enum Ownership { OWNS, DOES_NOT_OWN, UNKNOWN }
private static Lookup scan(HerdrClient herdr, int clientIndex, int clientCount, Set<Long> ancestry) {
boolean complete = true;
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsAnyOf(herdr, paneId, ancestry)) {
return pane.path("terminal_id").asText(null);
if (paneId == null) {
continue;
}
Ownership owns = paneOwnsAnyOf(herdr, clientIndex, clientCount, paneId, ancestry);
if (owns == Ownership.OWNS) {
return new Lookup(pane.path("terminal_id").asText(null), true);
}
if (owns == Ownership.UNKNOWN) {
complete = false;
}
}
return null;
return new Lookup(null, complete);
}
private static boolean paneOwnsAnyOf(HerdrClient herdr, String paneId, Set<Long> ancestry) {
private static Ownership paneOwnsAnyOf(HerdrClient herdr, int clientIndex, int clientCount,
String paneId, Set<Long> ancestry) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
// fleetd #505: this used to be read as a clean "does not own it" (the pane vanished
// mid-scan, just skip it) — one boolean carrying two different facts. It is UNKNOWN
// now: if THIS pane is the one that owns the pid, the caller must not be told "no pane
// owns it", because that reads as a real primary and is promoted under loopback-trust.
log.warn("pane.process_info failed for pane {} on herdr client {} of {} during a "
+ "pid-owner scan — treating it as \"could not tell\", not a clean "
+ "negative (fleetd #505): {}",
paneId, clientIndex + 1, clientCount, e.getMessage());
return Ownership.UNKNOWN;
}
if (ancestry.contains(info.path("shell_pid").asLong(-1))) {
return true;
return Ownership.OWNS;
}
for (JsonNode p : info.path("foreground_processes")) {
if (ancestry.contains(p.path("pid").asLong(-1))) {
return true;
return Ownership.OWNS;
}
}
return false;
return Ownership.DOES_NOT_OWN;
}
}
@@ -0,0 +1,42 @@
package dev.ltms.fleet.herdr;
/**
* The suffix {@code dev.ltms.fleet.lead.LeadLauncher} appends to a lead tab's label the first time a
* reconcile finds no running agent in it, before it is sure enough to close the tab outright.
*
* <p><strong>fleetd #359 review, finding 1.</strong> The daemon's own evidence showed
* {@code agent.list} can read "no agent" for a tab that genuinely has one running — so a single such
* reading must never be treated as proof a tab is dead. {@code LeadLauncher} now writes this marker
* on the first miss, and only closes the tab if a <em>later</em>, independent reconcile still finds
* it dead while the marker is still there. Two consecutive misses, one restart apart, is a much
* stronger claim than one.
*
* <p>{@link LeadTabScanner} strips the same suffix before matching a label against a configured
* lead's {@code tab}, so a flagged-but-actually-still-live tab keeps resolving normally — the marker
* changes nothing about which pane {@code LeadCoordLoop} can reach while the flag is pending. Both
* classes must use exactly this suffix, which is why it lives here rather than as a private constant
* on either.
*/
public final class PendingCloseMarker {
public static final String SUFFIX = " [fleetd:pending-close]";
private PendingCloseMarker() {
}
/** The label with any trailing pending-close marker removed, for name matching. */
public static String strip(String label) {
if (label == null) {
return null;
}
String stripped = label.strip();
return stripped.endsWith(SUFFIX)
? stripped.substring(0, stripped.length() - SUFFIX.length()).strip()
: stripped;
}
/** Whether a label currently carries the marker. */
public static boolean isFlagged(String label) {
return label != null && label.strip().endsWith(SUFFIX);
}
}
@@ -1,10 +1,12 @@
package dev.ltms.fleet.inject;
/**
* Notified when {@link CompletionResolver} actually delivers a typed backend-error classification
* to a waiting send (fleetd#201 / #227) — never on a race that lost. {@link CompletionResolver}
* calls this only after {@code Rendezvous.resolveFailure} returns {@code true} for that exact
* waiter, mirroring the win-only race rule {@link ExhaustionSink} already uses.
* Notified when {@link CompletionResolver} has a backend-error match at the start of a pane line,
* or both a match and its too-fast crash signature, for a waiting send (fleetd#201 / #227). A text
* match inside ordinary pane prose can be a member's report about an error, so it fails the send
* without notifying this sink.
* {@link CompletionResolver} calls this only after {@code Rendezvous.resolveFailure} returns
* {@code true} for that exact waiter, mirroring the win-only race rule {@link ExhaustionSink} uses.
*
* <p>The public send result is unchanged by this classification — it is still a failed send
* ({@code Rendezvous.Kind#FAILED}); this sink is the internal seam a later stage (fleetd#201 Unit
@@ -50,7 +50,7 @@ import java.util.regex.Pattern;
* so the scrape's herdr round-trip never stalls the status poller. The captured waiter is read on the
* poller thread (before any next-turn delivery can overwrite it) and passed into the virtual thread.
*/
public final class CompletionResolver implements TurnListener {
public final class CompletionResolver implements TurnListener, TurnRegistrar {
private static final Logger log = LoggerFactory.getLogger(CompletionResolver.class);
@@ -235,6 +235,26 @@ public final class CompletionResolver implements TurnListener {
this.worktreeBranches = Objects.requireNonNull(worktreeBranches, "worktreeBranches");
}
/**
* fleetd #556: {@link TurnRegistrar}'s structural half of what {@link #captureBaseline} used to
* do alone — record the waiter, no I/O. Called directly and unconditionally by the {@link
* Injector} as part of delivering a turn, before {@link #onDelivered} ever runs, so this
* registration cannot be skipped by a {@link TurnListener} throwing (from {@code onDelivered} or
* any other callback). {@link #captureBaseline} still performs this exact check-and-put itself
* as well — harmless and idempotent when it runs right after this — so a caller that only wires
* the {@link TurnListener} path (e.g. an existing test that never mentions {@link TurnRegistrar})
* keeps working unchanged.
*/
@Override
public void register(String target, TurnToken token) {
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
}
inFlight.put(target, new InFlight(waiter, null, nowNanos.getAsLong(), token.injectedText()));
}
@Override
public void onDelivered(String target, TurnToken token) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
@@ -244,7 +264,14 @@ public final class CompletionResolver implements TurnListener {
captureBaseline(target, token);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
/**
* Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link
* #onDelivered}). fleetd #556: this still does its own full waiter check and {@code inFlight.put}
* — the same registration {@link #register} performs — so it keeps working standalone (as every
* test calling it directly already does) even though the {@link Injector} now also calls {@link
* #register} on its own, earlier and unconditionally. The two writes are idempotent with each
* other; only the second (this one) carries a real baseline, since only this one pays for a scrape.
*/
void captureBaseline(String target, TurnToken token) {
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
if (waiter == null) {
@@ -380,16 +407,19 @@ public final class CompletionResolver implements TurnListener {
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
if (startsWithExhaustion(matchedLine, exhausted)) {
exhaustionSink.onExhausted(target, reason);
}
}
return;
}
// fleetd#164 (part 2) / fleetd#201: a scrape that read cleanly and produced content still
// isn't a real reply when that content is the backend's own rejection (e.g. an HTTP 400
// before the worker did any work). Classify it as a failure naming the member, rather than
// handing the caller a scrape that reads like a completed answer, and — only on the
// resolution that actually wins the race, mirroring the exhaustion sink above — notify the
// typed backend-error sink so a later stage can act on repeated failures.
// handing the caller a scrape that reads like a completed answer. A text match alone is not
// enough to notify the typed backend-error sink: this assistant block can be a member's
// normal prose about an error. A line that starts with the error match is stronger evidence;
// the too-fast path below also has its crash signature before it records a credential failure.
String backendError = firstMatchingLine(assistantBlock, backendErrorPatternOrFallback(target));
if (backendError != null) {
// Carry the whole scrape, not just the matched line. The pattern is a heuristic: a member
@@ -401,9 +431,9 @@ public final class CompletionResolver implements TurnListener {
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
// fleetd#201 Unit 1: only on the resolution that actually won the race — a late
// duplicate must never double-count one backend failure.
backendErrorSink.onBackendError(target, backendError, reason);
if (startsWithBackendError(backendError, backendErrorPatternOrFallback(target))) {
backendErrorSink.onBackendError(target, backendError, reason);
}
}
return;
}
@@ -477,7 +507,9 @@ public final class CompletionResolver implements TurnListener {
+ "usable assistant block; no fleet_reply): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
if (startsWithExhaustion(matchedLine, exhausted)) {
exhaustionSink.onExhausted(target, reason);
}
}
return true;
}
@@ -492,8 +524,9 @@ public final class CompletionResolver implements TurnListener {
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback from the raw scrape: {}", target, reason);
// fleetd#201 Unit 1: only on the resolution that actually won the race.
backendErrorSink.onBackendError(target, backendError, reason);
if (startsWithBackendError(backendError, backendErrorPatternOrFallback(target))) {
backendErrorSink.onBackendError(target, backendError, reason);
}
}
return true;
}
@@ -540,10 +573,21 @@ public final class CompletionResolver implements TurnListener {
* {@link #MIN_TURN_NANOS} — a crash signature (e.g. a backend HTTP 400 before the worker did
* anything) that a bare {@code BUSY -> DONE} transition cannot be told apart from a genuinely
* fast completion. Runs the same backend-error classification the normal and raw-scrape paths
* apply, against whatever is on screen right now: a match is a typed failure that notifies
* {@link #backendErrorSink} (only on the resolution that wins the race); a non-match stays the
* original generic too-fast failure, naming the member and both timings, with whatever the pane
* shows appended so the caller sees the cause, not just "it failed".
* apply, against whatever is on screen right now: a match together with the too-fast crash
* signature notifies {@link #backendErrorSink} (only on the resolution that wins the race). A
* non-match stays the generic too-fast failure, naming the member and both timings, with
* whatever the pane shows appended so the caller sees the evidence, not just "it failed".
*
* <p>fleetd#376: <strong>this path must never resolve a completion.</strong> A fast backend can
* genuinely answer inside the floor, so the failure is sometimes wrong — but it is wrong in the
* loud direction, and the fix for that is honest wording, not a guess at the pane's meaning.
* Reclassifying from the scrape was tried and rejected: there is no reliable positive marker for
* "this is a real reply" across backends. {@link #lastAssistantBlock} falls back to the entire
* pane when it finds no {@code ⏺} marker, so on a crash the candidate "reply" is the whole
* screen; and {@code ⏺} itself is a Claude Code marker that an opencode pane never carries — the
* very backend whose speed raised this ticket. Any weaker test (non-blank, or "contains sentence
* punctuation") passes on almost every crash, because a pane holding a file path, a version
* number or a hostname contains a full stop. That trades a loud wrong answer for a silent one.
*/
private void failTooFast(String target, InFlight turn, CompletableFuture<Rendezvous.Resolution> waiter,
long elapsedNanos) {
@@ -554,13 +598,13 @@ public final class CompletionResolver implements TurnListener {
scrape = "";
}
String clippedScrape = clip(scrape);
String baseReason = String.format(
"member %s went BUSY -> DONE in %dms (floor %dms) — too fast to be real work, most "
+ "likely a backend error before any work started",
String timing = String.format(
"member %s went BUSY -> DONE in %dms (floor %dms)",
target, elapsedNanos / 1_000_000, MIN_TURN_NANOS / 1_000_000);
String backendError = firstMatchingLine(scrape, backendErrorPatternOrFallback(target));
if (backendError != null) {
String reason = baseReason + ": " + clippedScrape;
String reason = timing + " — too fast to be real work, and the pane carries a backend "
+ "error: " + clippedScrape;
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
@@ -569,7 +613,14 @@ public final class CompletionResolver implements TurnListener {
}
return;
}
fail(target, turn, clippedScrape.isBlank() ? baseReason : baseReason + ": " + clippedScrape);
// fleetd#376: no pattern matched, so the cause is genuinely unknown. Say that, rather than
// asserting a backend error the way this message used to — a fast backend really can finish
// inside the floor, and a reader who trusts a wrong cause stops looking at the pane.
String reason = timing + " — inside the floor. That is usually a backend error before any "
+ "work started, but a fast backend can answer inside it too, and nothing here can "
+ "tell those apart, so the turn is reported failed rather than guessed. Read the "
+ "pane below before deciding which it was";
fail(target, turn, clippedScrape.isBlank() ? reason : reason + ": " + clippedScrape);
}
/**
@@ -612,17 +663,119 @@ public final class CompletionResolver implements TurnListener {
}
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* True when nothing before the match on this pane line ends a sentence — that is, the match is
* still inside the line's first sentence rather than inside prose a member wrote about it.
* Used to decide whether an exhaustion match may quarantine a credential (fleetd #348).
*
* <p><strong>Why this is looser than {@link #startsWithBackendError}.</strong> An
* {@code exhaustedPattern} is written per profile and may name only the decisive words of a
* provider message — {@code "usage limit has been reached"} without its leading {@code "The"}.
* A start-of-line check would then reject the genuine refusal. That is the false negative
* fleetd #348's invariant 1 calls the worse direction: an unrecorded exhaustion leaves the
* fleet spawning into a credential with no capacity, and a quarantine runs 1800s against the
* backend-error cooldown's fixed 60s.
*
* <p>This rule accepts a superset of what a start-of-line check accepts: if the match begins
* right after the chrome, there is nothing in front of it, so there is no sentence ending
* either. So moving to it cannot add a false negative.
*
* <p><strong>No chrome skipping here, deliberately.</strong> The first version of this method
* copied {@code startsWithBackendError}'s leading-chrome loop. Measured on merge: deleting that
* loop left all 1369 tests green, and it must — the scan only looks for {@code . ! ?}, and no
* terminal chrome character is one of those. A step that cannot change the result is worse than
* no step, because the next reader takes it as evidence that chrome was handled.
*
* <p>It stays a heuristic. Prose whose <em>first</em> sentence carries the pattern still
* notifies the sink, and a genuine refusal behind an earlier full stop (a hostname, a version
* number) still does not. Both are known and neither is fixed here.
*/
private static boolean startsWithExhaustion(String line, Pattern pattern) {
var matcher = pattern.matcher(line);
if (!matcher.find()) {
return false;
}
for (int prefix = 0; prefix < matcher.start(); prefix++) {
if (".!?".indexOf(line.charAt(prefix)) >= 0) {
return false;
}
}
return true;
}
/**
* True when the error pattern begins the matched pane line, rather than appearing in prose.
*
* <p>Leading terminal chrome is skipped first — box-drawing characters, bullets, gutter bars and
* spaces. #339 introduced this check with a bare {@code lookingAt}, and that rejected a genuine
* error line rendered as {@code "| 503 Service Unavailable: ..."}: the send still failed, but the
* credential outage was never recorded. That is the false negative #339's own invariant 3 called
* worse than the false positive it set out to fix — measured with a throwaway probe on the
* raw-scrape path, which is exactly the path whose comment says to expect leading chrome.
*
* <p>Skipping only a leading run of non-letter, non-digit characters keeps the fix's intent. A
* member's prose ({@code "I checked the retry path. An API Error: makes it back off."}) still
* does not match, because there the pattern sits after words, not after chrome.
*/
private static boolean startsWithBackendError(String line, Pattern pattern) {
int i = 0;
while (i < line.length() && !Character.isLetterOrDigit(line.charAt(i))) {
i++;
}
return pattern.matcher(line.substring(i)).lookingAt();
}
/**
* What an unset pattern key means for the classification it configures (fleetd#415).
* {@code coverage()} cannot infer this from the key's name — the two keys it currently
* describes disagree on it, and a string comparison on the name would just move the same bug
* to a new spot — so every caller must state it explicitly.
*
* <p><strong>This alone does not prove a caller passes the right one for its key.</strong> A
* test that calls {@code coverage()} directly and supplies the meaning itself only proves this
* enum is worded correctly, never that {@code Fleetd}'s two call sites pair each key with its
* true meaning — that pairing is #415's actual defect. Measured on review: swapping the two
* {@code UnsetMeaning} arguments at those call sites (giving {@code exhaustedPattern} the
* built-in-default wording and {@code errorPattern} the off wording — #415's exact defect with
* the keys exchanged) compiled with 0 errors and left all 1506 existing tests green. See
* {@code dev.ltms.fleet.Fleetd#exhaustedPatternCoverageLine}/{@code #errorPatternCoverageLine}
* and {@code FleetdPatternCoverageLineTest}, which exists specifically to catch that swap.
*/
public enum UnsetMeaning {
/** No fallback exists: a profile with no configured pattern truly has this classification off. */
OFF,
/** A built-in pattern applies when unset: the classification still runs for that profile. */
BUILT_IN_DEFAULT
}
/**
* Coverage summary for a fleetd#201/CB-578-style pattern-key classification, logged at startup
* the way {@link dev.ltms.fleet.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* <p>fleetd#415: this method measures <em>pattern coverage</em> — how many profiles set the
* key — which is not the same thing as <em>feature state</em> for a key with a fallback. For
* {@code errorPattern}, an empty {@code configuredProfiles} still runs the classification
* against {@code CompletionResolver}'s built-in compatibility pattern ({@link #BACKEND_ERROR}
* at line ~84); for {@code exhaustedPattern} there is no fallback, so empty really does mean
* off. {@code unsetMeaning} is the single, required source of that fact — see
* {@link dev.ltms.fleet.config.FleetConfig#rejectMalformedProfilePatterns} lines ~2029-2032 for
* where it is documented for config authors. It is a required parameter, not a defaulted
* overload: a third pattern key added later must supply one to compile at all, rather than
* silently inheriting whichever wording this method happened to default to.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
* @param configuredProfiles the subset of {@code allProfiles} that carry the pattern
*/
public static String coverage(String patternKey, Set<String> allProfiles, Set<String> configuredProfiles) {
public static String coverage(String patternKey, UnsetMeaning unsetMeaning, Set<String> allProfiles,
Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return "off (no profile has an " + patternKey + " configured; profiles: " + sorted(allProfiles) + ")";
return switch (unsetMeaning) {
case OFF -> "off (no profile has an " + patternKey + " configured; profiles: "
+ sorted(allProfiles) + ")";
case BUILT_IN_DEFAULT -> "built-in default for all profiles (no profile customises "
+ patternKey + "; profiles: " + sorted(allProfiles) + ")";
};
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
@@ -2,6 +2,7 @@ package dev.ltms.fleet.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.msg.TurnToken;
import org.slf4j.Logger;
@@ -11,10 +12,12 @@ import java.util.ArrayDeque;
import java.util.ArrayList;
import java.util.Deque;
import java.util.List;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
import java.util.function.Predicate;
import java.util.stream.Collectors;
@@ -44,6 +47,17 @@ import java.util.stream.Collectors;
* turn — the pickup-grace path (a turn too fast to sample) unwedges the queue but does not fire
* completion, since without a sampled {@code working} there is no trustworthy "the worker just
* finished the task" signal to act on.
*
* <p><strong>Delivery honesty (fleetd #551).</strong> The one external send call
* ({@link AgentControl#send}) is irreversible, and its own response can still fail after the text
* has already reached the worker's pane — herdr replied (or may have), so the request was already
* processed, but the reply itself then failed to parse or carried an error. A queued entry is
* polled off the queue and marked {@code ATTEMPTED} <em>before</em> that call is made, not after,
* so a failure from the call itself can never be recorded as a confident {@code NOT_DELIVERED} for
* text that may already be sitting in the pane. {@code NOT_DELIVERED} stays reserved for the cases
* where nothing was ever attempted — the readiness grace expiring, {@link #drop}, or a herdr
* {@code *_not_found} error, which the rest of this codebase already treats as a confirmed absence
* rather than a merely inconclusive failure (see {@code StatusPoller}, {@code AgentControl}).
*/
public final class Injector {
@@ -92,8 +106,21 @@ public final class Injector {
private final AgentControl agents;
private final HerdrRouter router;
private final TurnListener turnListener;
/**
* fleetd #556: the Injector's own registration invariant, called directly and unconditionally —
* never folded into {@link #turnListener}'s fan-out. See {@link TurnRegistrar}.
*/
private final TurnRegistrar registrar;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
/**
* Wall-clock source for the readiness-grace elapsed time logged in {@link #onStatus} (fleetd
* #501). Production constructors default this to {@code System::currentTimeMillis}; the
* package-private constructors below take it explicitly so a test can supply a stub whose
* advance does not track {@link #POLL_INTERVAL_MILLIS} — copying the shape {@code LeadRollover}
* already uses for the same purpose.
*/
private final LongSupplier nowMillis;
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
@@ -125,28 +152,170 @@ public final class Injector {
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
// fleetd #556: no explicit registrar named here, so fall back to turnListener itself when it
// happens to also implement TurnRegistrar (true for CompletionResolver, the only production
// TurnListener that needs CB-106 registration) — every existing caller of this overload keeps
// working unchanged. A turnListener that does NOT implement TurnRegistrar (a fan-out object,
// or a bare test lambda) gets TurnRegistrar.NOOP here, same as before this ticket.
this(agents, turnListener, ready, forget,
turnListener instanceof TurnRegistrar r ? r : TurnRegistrar.NOOP);
}
/**
* fleetd #556: explicit-registrar overload. Use this whenever {@code turnListener} does not
* itself implement {@link TurnRegistrar} — e.g. a fan-out object that also notifies unrelated
* observers — so registration is wired directly to the one component that owns it, rather than
* relying on {@code turnListener} happening to implement both interfaces.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, TurnRegistrar registrar) {
this(agents, turnListener, ready, forget, registrar, System::currentTimeMillis);
}
/**
* Package-private constructor — for tests: an injectable wall-clock supplier (fleetd #501),
* with no explicit registrar (fleetd #556 fallback, same as the public 4-arg overload above).
*/
Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, LongSupplier nowMillis) {
this(agents, turnListener, ready, forget,
turnListener instanceof TurnRegistrar r ? r : TurnRegistrar.NOOP, nowMillis);
}
/** Full constructor — for tests: an explicit registrar and an injectable wall-clock supplier. */
Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, TurnRegistrar registrar, LongSupplier nowMillis) {
this.agents = agents;
this.router = null;
this.turnListener = turnListener;
this.registrar = Objects.requireNonNull(registrar, "registrar");
this.ready = ready;
this.forget = forget;
this.nowMillis = nowMillis;
}
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
// fleetd #556: see the AgentControl overload above for the fallback rationale.
this(router, turnListener, ready, forget,
turnListener instanceof TurnRegistrar r ? r : TurnRegistrar.NOOP);
}
/**
* fleetd #556: explicit-registrar overload — production wiring ({@code Fleetd}) uses this, since
* its {@code turnListener} is a fan-out object that also notifies {@code SessionManager} and does
* not itself implement {@link TurnRegistrar}.
*/
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, TurnRegistrar registrar) {
this(router, turnListener, ready, forget, registrar, System::currentTimeMillis);
}
/** Full constructor — for tests: an injectable wall-clock supplier (fleetd #501). */
Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, TurnRegistrar registrar, LongSupplier nowMillis) {
this.agents = null;
this.router = router;
this.turnListener = turnListener;
this.registrar = Objects.requireNonNull(registrar, "registrar");
this.ready = ready;
this.forget = forget;
this.nowMillis = nowMillis;
}
private AgentControl agentsFor(String target) {
return router != null ? router.agentsFor(target) : agents;
}
/** The result of trying to remove an undelivered message from the injector. */
public enum Cancellation {
/** The message was still queued and this call removed it; the target saw nothing. */
CANCELLED,
/** The message reached the worker's pane and is confirmed delivered. */
DELIVERED,
/**
* fleetd #551: the injector called {@link AgentControl#send} for this message and that call
* threw before its outcome was known — the text may or may not have reached the pane. A
* caller reporting this to an operator must say "uncertain", not "definitely not
* delivered": reading it as a confident negative invites a resend of text that may already
* be sitting in the pane (a double delivery), which is worse than the ambiguity itself.
*/
ATTEMPTED,
/**
* The message never reached the worker's pane — nothing was ever attempted for it (the
* readiness grace expired, the target was dropped, or send failed with a herdr
* {@code *_not_found} error, which is a confirmed absence, not merely inconclusive).
*/
NOT_DELIVERED
}
/**
* An identity handle for one queued delivery. It is the only value accepted by
* {@link #cancel(Delivery)}, so a caller cannot cancel a different message with the same target
* or text.
*/
public static final class Delivery {
private final Pending pending;
private Delivery(Pending pending) {
this.pending = pending;
}
public CompletableFuture<Void> completion() {
return pending.delivered;
}
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, TurnToken token, CompletableFuture<Void> delivered) {
private static final class Pending {
enum State {
/** Still in the target's queue, not yet attempted. */
QUEUED,
/**
* fleetd #551: {@link AgentControl#send} has been called for this entry and its outcome
* is not yet known — recorded BEFORE the call (see {@link #onStatus}), so a Throwable
* from send() can never leave the entry either still QUEUED or falsely marked
* NOT_DELIVERED. Upgraded to DELIVERED on success; left as ATTEMPTED on an ordinary
* failure, since reaching the catch does not prove the text never reached the pane.
*/
ATTEMPTED,
/** {@code send()} returned normally: the text is confirmed to have reached the pane. */
DELIVERED,
/**
* Confirmed — not merely inconclusive — that nothing was ever sent for this entry: the
* worker never became ready ({@link #onStatus}'s readiness-grace expiry), the target was
* dropped ({@link #drop}), or send failed with a herdr {@code *_not_found} error. Never
* written for an entry whose send() outcome is unknown; see ATTEMPTED.
*/
NOT_DELIVERED,
/** Removed from the queue by {@link #cancel} before it was ever attempted. */
CANCELLED
}
final String target;
final String text;
final TurnToken token;
final CompletableFuture<Void> delivered;
volatile State state = State.QUEUED; // written under the owning Target monitor
Pending(String target, String text, TurnToken token, CompletableFuture<Void> delivered) {
this.target = target;
this.text = text;
this.token = token;
this.delivered = delivered;
}
String text() {
return text;
}
TurnToken token() {
return token;
}
CompletableFuture<Void> delivered() {
return delivered;
}
}
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
@@ -159,6 +328,7 @@ public final class Injector {
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int unknownSincePostTurn; // the same, for the post-turn housekeeping phase (fleetd #306)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
long notReadySinceMillis; // wall-clock time of the FIRST non-ready sample in the current notReadySincePoll streak (fleetd #501); reset alongside it
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
@@ -177,15 +347,54 @@ public final class Injector {
* <p>Uses an atomic map update so a concurrent {@link #drop} cannot slip between "find the
* target" and "queue the message" and orphan it in a target it just removed.
*/
public CompletableFuture<Void> enqueue(String target, String text, TurnToken token) {
public Delivery enqueue(String target, String text, TurnToken token) {
CompletableFuture<Void> delivered = new CompletableFuture<>();
Pending p = new Pending(text, token, delivered);
Pending p = new Pending(target, text, token, delivered);
targets.compute(target, (_, existing) -> {
Target t = (existing != null) ? existing : new Target();
t.add(p); // synchronized on the Target monitor — atomic with a concurrent drop
return t;
});
return delivered;
return new Delivery(p);
}
/**
* Cancel this exact queued delivery. The target monitor serializes this operation with
* {@link #onStatus}: if delivery wins that race, this returns {@link Cancellation#DELIVERED}
* rather than claiming the message remained queued.
*/
public Cancellation cancel(Delivery delivery) {
Pending p = delivery.pending;
Target t = targets.get(p.target);
if (t == null) {
return cancellationOf(p);
}
synchronized (t) {
if (p.state != Pending.State.QUEUED || !t.queue.remove(p)) {
return cancellationOf(p);
}
p.state = Pending.State.CANCELLED;
if (isQuiescent(t)) {
targets.remove(p.target, t);
}
return Cancellation.CANCELLED;
}
}
private static Cancellation cancellationOf(Pending p) {
// fleetd #551 (comment 17037): a two-way split on a now-three-way question folded ATTEMPTED
// into NOT_DELIVERED with no compiler error and no failing test — the exact defect this
// ticket exists to fix, one layer up. ATTEMPTED gets its own answer instead.
return switch (p.state) {
case DELIVERED -> Cancellation.DELIVERED;
case ATTEMPTED -> Cancellation.ATTEMPTED;
case QUEUED, NOT_DELIVERED, CANCELLED -> Cancellation.NOT_DELIVERED;
};
}
private static boolean isQuiescent(Target t) {
return t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion
&& !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved;
}
/**
@@ -200,7 +409,7 @@ public final class Injector {
if (t == null) return;
Pending sent = null;
RuntimeException sendError = null;
Throwable sendError = null;
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
@@ -220,6 +429,7 @@ public final class Injector {
t.unknownSinceTurn = 0;
t.unknownSincePostTurn = 0;
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
@@ -271,35 +481,93 @@ public final class Injector {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
// fleetd #551: poll and record BEFORE the irreversible send, not after.
// The entry comes off the queue and its state is set to ATTEMPTED here,
// unconditionally — so a Throwable escaping the send call below (caught or
// not) can never leave the entry QUEUED at the head of t.queue (the fleetd
// #546 hazard, since peek() alone would let the next onStatus round re-enter
// this block and send the same text again), and no path can write a
// confident DELIVERED or NOT_DELIVERED before we actually know which one
// happened.
t.queue.poll();
p.state = Pending.State.ATTEMPTED;
try {
agentsFor(target).send(target, p.text());
t.queue.poll();
p.state = Pending.State.DELIVERED;
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
} catch (Throwable e) {
// fleetd #551: leave p.state == ATTEMPTED (recorded above, before the
// call) rather than downgrading it to NOT_DELIVERED here — reaching this
// catch does not prove the text never reached the pane. Three of the
// four HerdrException throw sites in HerdrCodec fire only after herdr
// has already replied (so it processed the request), and the fourth (a
// transport IOException) leaves it genuinely unknown whether herdr even
// received the bytes — see #551 comment 16867. The one exception is a
// herdr `*_not_found` error: that family is already read as "definitely
// absent, not merely inconclusive" everywhere else in this codebase
// (StatusPoller, AgentControl's own retry, WorkspaceControl,
// HerdrPeerLauncher, FleetApp, ReplyPushLoop) because it means the
// target pane/agent does not exist at all, so nothing could have been
// pasted anywhere — #551 keeps the new state consistent with that
// existing vocabulary rather than inventing a second one.
if (e instanceof HerdrException he && he.code() != null
&& he.code().endsWith("_not_found")) {
p.state = Pending.State.NOT_DELIVERED;
}
sent = p;
sendError = e;
}
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
+ "became deliverable, so failing {} queued message(s) that never "
+ "reached its pane",
target, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
} else if (p != null) {
// fleetd #501: stamp the wall-clock time of the FIRST non-ready sample in
// this streak, so the expiry log below can print how long the target
// actually sat non-ready — not just how many polls that took.
if (t.notReadySincePoll == 0) {
t.notReadySinceMillis = nowMillis.getAsLong();
}
if (++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
for (Pending pending : notReady) {
pending.state = Pending.State.NOT_DELIVERED;
}
// fleetd #501: t.notReadySincePoll — the loop's own counter, already in
// scope — is printed here instead of the READINESS_GRACE_POLLS constant.
// On this branch the counter has JUST reached the threshold, so the two
// agree by construction and no test can tell them apart. Printed anyway:
// it gives this line one source of truth instead of two, so a later
// change to the loop above cannot leave this message reporting a number
// the loop no longer produces.
//
// elapsedMillis is a different case: it is NOT equal-by-construction to
// the truth. notReadySincePoll only increments on a sample that reaches
// this branch (p != null, not ready) — a poll that misses that condition
// advances real time without advancing the counter — and this loop's real
// period is not guaranteed to equal POLL_INTERVAL_MILLIS (load, or a host
// sleep, can widen the real gap far past it). READINESS_GRACE_POLLS *
// POLL_INTERVAL_MILLIS / 1000 is arithmetic on two constants, not a
// measurement, so it stays here only as the labelled CONFIGURED budget,
// never presented as elapsed time.
long elapsedMillis = nowMillis.getAsLong() - t.notReadySinceMillis;
log.warn("readiness grace for {} expired after {} polls (configured={} "
+ "polls/{}s elapsed={}ms): target never became "
+ "deliverable, so failing {} queued message(s) that "
+ "never reached its pane",
target, t.notReadySincePoll, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, elapsedMillis,
notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
}
}
}
}
@@ -340,62 +608,189 @@ public final class Injector {
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion
&& !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved) {
if (isQuiescent(t)) {
targets.remove(target, t);
}
}
// Fire listeners / herdr calls after releasing the monitor so nothing runs on the poller
// thread while it holds the target lock.
if (resubmit) {
try {
agentsFor(target).submit(target); // nudge a raced Enter so the pending paste submits
} catch (RuntimeException e) {
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
//
// fleetd #553: wrapped in try/finally. By this point, if `sent != null`, the delivery has
// already happened inside the monitor above — off the queue, p.state == DELIVERED, text
// typed into the target's pane — so `sent`'s future MUST be completed one way or another,
// in every path out of this region, or the caller (a blocking fleet_send, or an async
// ticket) waits forever on a message it actually received. But the finally must not
// swallow whatever escaped: a listener that throws is a defect in THAT listener, and it
// must still reach StatusPoller's catch (Throwable) so log.error fires — converting a
// loud listener bug into a silently orphaned future would be worse than the bug itself.
//
// There are TWO futures at stake here, not one: `sent.delivered()` (the delivery future)
// and `sent.token().waiter()` (the rendezvous waiter a blocking fleet_send actually waits
// on for the worker's ANSWER). fleetd #556: the waiter is registered by `registrar.register`
// now — called directly and unconditionally, structural rather than delegated to a listener
// callback — never by turnListener.onDelivered() alone (before this ticket, that WAS the
// only path, and any TurnListener wired ahead of the one that mattered could throw and skip
// it; #553 could only make the reachable listener behave, not remove that dependency).
// Completing `sent.delivered()` without also calling `registrar.register` leaves the waiter
// unregistered forever — a hang with a success receipt, which is worse than the plain hang
// #553 was about. So the `finally` below does not just complete the future: on the path
// where the `if (sent != null)` block never ran, it does that block's whole job —
// register, then onDelivered (notification only, may throw), then complete.
//
// `sentHandled` is NOT allowed to move earlier than this: an earlier comment on #553
// proposed running `if (sent != null)` first, before turnCompleted/turnFailed, and that
// was withdrawn — `registrar.register` WRITES CompletionResolver's inFlight record for this
// turn, onTurnComplete READS it for the PREVIOUS turn, and running register first makes
// onTurnComplete resolve the NEW turn's waiter with the PREVIOUS turn's stale output
// (CB-116). The `if (sent != null)` block stays last; the `finally` is a backstop for it,
// not a replacement. fleetd #556 keeps this ordering unchanged — it only splits what used
// to be one listener call (register + notify) into two calls in the same position.
boolean sentHandled = false;
try {
if (resubmit) {
try {
agentsFor(target).submit(target); // nudge a raced Enter so the pending paste submits
} catch (Throwable e) {
// fleetd #553: widened from RuntimeException (same reasoning as #549 at :382) —
// this is the FIRST block after the monitor, so an Error escaping it used to skip
// every block below it, including the `sent` completion.
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
}
}
}
if (notReady != null) {
// Worker never became available: forget its (never-set) readiness, unblock every queued
// caller, and route the awaiting send through the same failure path as a stalled turn so
// a blocking or async waiter resolves WORKER_FAILED rather than riding out the timeout.
forget.accept(target);
RuntimeException cause = new IllegalStateException(
target + " never became available (no bridge MCP connection within the boot window)");
for (Pending p : notReady) {
p.delivered().completeExceptionally(cause);
if (notReady != null) {
// Worker never became available: unblock every queued caller FIRST — these messages'
// fate (NOT_DELIVERED, off the queue) was already decided inside the monitor above, so
// fleetd #553 completes every one of them before calling forget.accept or
// turnListener.onTurnFailed below. Either of those is a listener/consumer callback and
// can throw (an ordinary RuntimeException is enough — the same reasoning as the rest of
// this ticket): completing the futures first means such a throw can no longer leave any
// of them permanently pending, regardless of which one throws or in which order.
RuntimeException cause = new IllegalStateException(
target + " never became available (no bridge MCP connection within the boot window)");
for (Pending p : notReady) {
p.delivered().completeExceptionally(cause);
}
// route the awaiting send through the same failure path as a stalled turn so a blocking
// or async waiter resolves WORKER_FAILED rather than riding out the timeout.
forget.accept(target);
turnListener.onTurnFailed(target);
}
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
if (startPostTurn) {
boolean started = turnListener.onTurnCompleteWithPostAction(target);
synchronized (t) {
t.postTurnPending = false;
if (started) {
t.awaitingPostTurnPickup = true;
t.injectableSincePostTurnPickup = 0;
if (turnCompleted) {
if (startPostTurn) {
// fleetd #553: the listener call is wrapped so `t.postTurnPending` (set true inside
// the monitor above, before this call) is always reset. Before this wrapping, a
// RuntimeException from onTurnCompleteWithPostAction skipped the reset below,
// permanently wedging the target: postTurnPending stayed true forever, so this
// target's delivery guard (:376) would never again pass and no further message to it
// would ever be delivered — a target-level lockup, not just one skipped future.
// `started` defaults to false so an exception is treated as "the post-turn action did
// not start" rather than falsely arming the post-turn pickup latch for an action that
// never ran.
boolean started = false;
try {
started = turnListener.onTurnCompleteWithPostAction(target);
} finally {
synchronized (t) {
t.postTurnPending = false;
if (started) {
t.awaitingPostTurnPickup = true;
t.injectableSincePostTurnPickup = 0;
}
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
targets.remove(target, t);
}
}
}
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
targets.remove(target, t);
} else {
turnListener.onTurnComplete(target);
}
}
if (turnFailed) {
turnListener.onTurnFailed(target);
}
if (sent != null) {
// fleetd #553: record the INTENT to handle before any side effect, not the fact of
// having completed it. If register()/onDelivered() throws part way through — e.g.
// after CompletionResolver's captureBaseline has already done its inFlight.put — a
// flag set only after this block would still read false, and the `finally` below
// would re-run this block's job a SECOND time. A second captureBaseline runs later,
// once onTurnComplete has already thrown and possibly scraped the pane, and can
// snapshot a pane that already absorbed this turn's output — which means the pane's
// tail never differs from that baseline again and CompletionResolver.resolve's own
// suppression at :364-369 (`return` keeping the in-flight record) drops every future
// completion for this turn, permanently. So this flag must go true FIRST.
sentHandled = true;
if (sendError != null) {
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
sent.delivered().completeExceptionally(sendError);
} else {
// fleetd #556: register the waiter FIRST and unconditionally — this is the
// Injector's own invariant, kept structural rather than delegated. Even if
// turnListener.onDelivered below throws (a bug in an unrelated observer, e.g.
// SessionManager's bookkeeping, or a fan-out wired in a different order), the
// waiter is already registered and resolvable — a throwing TurnListener can no
// longer take this invariant down with it.
registrar.register(target, sent.token());
// Notification only, from here down: baseline the pane's pre-turn content so a
// misattributed completion (no new output) can't resolve this send with the
// previous turn's stale answer (CB-115). Allowed to throw and to fail (the herdr
// scrape inside already fails open with baseline == null on its own).
turnListener.onDelivered(target, sent.token());
sent.delivered().complete(null);
}
}
} finally {
// Backstop (fleetd #553, extended #556): if an earlier block in this try threw before the
// `if (sent != null)` block above ran, `sentHandled` is still false here, and this does
// that block's WHOLE job — not just the future completion. Skipping registrar.register()
// here would leave sent.token().waiter() never registered with CompletionResolver, so a
// blocking fleet_send would be told its message was delivered and then wait out its full
// timeout for an answer that can never resolve — worse than the plain hang, because it now
// looks like success. This block runs only when sendError == null: nothing was delivered
// on the error path, so there is nothing to register.
//
// fleetd #553 (lead review, comment 16916): `sentHandled` guards ONLY the register/notify
// re-call below, never the future completion outside this inner try. One boolean cannot
// carry both meanings — "the turn was handled" and "the future has been dealt with" —
// because they come apart exactly when onDelivered throws PART WAY THROUGH: sentHandled
// is already true (set before the call, correctly — see the `if (sent != null)` block
// above), so a guard on the outer `if` here would skip this whole recovery, including the
// completion, and leave sent.delivered() pending forever even though the message really
// was typed into the pane and taken off the queue. So `sent != null` alone gates whether
// this target has anything to finish; `!sentHandled` gates only the register/notify
// re-call inside. CompletableFuture.complete/completeExceptionally are idempotent — on the
// ordinary path (sentHandled == true, no throw) the `if (sent != null)` block above has
// already completed this future, so the calls below are a no-op returning false.
//
// This recovery is wrapped in its own try/catch(Throwable) that swallows only ITS OWN
// throwable and logs at WARN — the future is still completed either way — while the
// ORIGINAL throwable from the try block above is left alone to keep unwinding out of
// this method to StatusPoller's catch (Throwable), so a listener bug stays loud.
if (sent != null) {
try {
if (!sentHandled && sendError == null) {
// fleetd #556: register first here too, same as the ordinary path above — this
// recovery path is reached only when an EARLIER block (resubmit/turnCompleted/
// turnFailed) threw before the ordinary path ever ran, so nothing has
// registered this turn yet.
registrar.register(target, sent.token());
turnListener.onDelivered(target, sent.token());
}
} catch (Throwable recoveryError) {
log.warn("fleetd #553 backstop: onDelivered failed for {} while recovering from "
+ "an earlier listener failure; completing its delivery future "
+ "anyway: {}",
target, recoveryError.getMessage());
} finally {
// No-op (returns false) on the ordinary path, where the `if (sent != null)` block
// above already completed this future — see the comment above this block.
if (sendError != null) {
sent.delivered().completeExceptionally(sendError);
} else {
sent.delivered().complete(null);
}
}
} else {
turnListener.onTurnComplete(target);
}
}
if (turnFailed) {
turnListener.onTurnFailed(target);
}
if (sent != null) {
if (sendError != null) {
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
sent.delivered().completeExceptionally(sendError);
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target, sent.token());
sent.delivered().complete(null);
}
}
}
@@ -432,6 +827,9 @@ public final class Injector {
boolean hadDeliveredTurn;
synchronized (t) {
pending = new ArrayList<>(t.queue);
for (Pending p : pending) {
p.state = Pending.State.NOT_DELIVERED;
}
t.queue.clear();
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
@@ -0,0 +1,91 @@
package dev.ltms.fleet.inject;
import dev.ltms.fleet.config.FleetConfig;
import java.util.Map;
import java.util.Objects;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Supplier;
import java.util.regex.Pattern;
/**
* fleetd #446: the single live source for "is {@code exhaustedPattern} configured for this
* profile, and what does it compile to". Read fresh off the config supplier on every call — the
* same reason {@code CompositePeerLauncher#models0} is a live supplier read rather than a value
* captured at construction (see that class's doc) — so an operator can arm or disarm usage-limit
* detection for a profile by editing {@code exhaustedPattern} and reloading, with no restart.
*
* <p>Before this class, {@code exhaustedPattern} was compiled once into a {@code Map<String,
* Pattern>} built inside {@code Fleetd.main} at startup ({@code ConfigRef}'s class doc used to
* list it under <em>Deferred</em>, CB-578 stage A) — the asymmetry fleetd #446 exists to close:
* an operator could turn a model off at runtime ({@code models.allow}'s {@code enabled: false},
* hot since fleetd #422) but could not arm the detector that would tell them to, without a
* restart. That was backwards for a feature whose whole point is to react while the fleet runs.
*
* <p>{@link #patternFor} is what {@link CompletionResolver} enforces on (via the {@link
* ExhaustedPatternLookup} production wiring in {@code Fleetd.main}); {@link #armed} is what
* {@code fleet_profiles}/{@code GET /profiles} report as {@code exhaustionDetectionArmed}. Both
* read this ONE object, so the report can never disagree with the behaviour — the same rule
* {@code CompositePeerLauncher#modelGateState()}'s javadoc states for the model gate's own
* armed/off pair (fleetd #404): "armed" and "which models are off" must come from one read of the
* same accessor the gate enforces on.
*
* <h2>Cache eviction</h2>
* Cached by profile NAME, not by pattern text. A {@code Map<String, Pattern>} keyed by the
* pattern STRING would grow by one entry per distinct regex ever typed for any profile across the
* daemon's uptime — unbounded in practice, because tuning a regex to match a backend's exact
* wording is exactly the kind of edit an operator makes several times while getting it right, and
* every edit-and-reload cycle would leave the previous attempt's compiled {@link Pattern} behind
* forever. Keyed by profile name instead, this cache holds at most one entry per profile name that
* has ever been looked up — and that key space is bounded by the (small, human-authored) set of
* configured profiles, which changes only on a restart: adding or removing a profile is itself a
* deferred key (a new backend needs its own launcher, built once — see {@code ConfigRef}'s class
* doc), so profile names do not churn the way pattern text does. Re-editing an EXISTING profile's
* {@code exhaustedPattern} — the case this class exists to make hot — simply overwrites that
* profile's one cache entry; it never adds a new one.
*/
public final class LiveExhaustedPatterns {
/** A compiled pattern paired with the source string it was compiled from, for change detection. */
private record Cached(String source, Pattern compiled) {
}
private final Supplier<Map<String, FleetConfig.Profile>> profiles;
private final ConcurrentHashMap<String, Cached> cache = new ConcurrentHashMap<>();
public LiveExhaustedPatterns(Supplier<Map<String, FleetConfig.Profile>> profiles) {
this.profiles = Objects.requireNonNull(profiles, "profiles");
}
/**
* The compiled {@code exhaustedPattern} currently configured for {@code profileName}, or
* {@code null} when that profile is unknown or has none configured. Recompiles only when the
* live pattern text differs from what is cached for this profile name; {@code
* FleetConfig#rejectMalformedProfilePatterns} already refuses a config (at load and at reload)
* whose {@code exhaustedPattern} does not compile, so this is not expected to throw in
* production — it is not defended against here for that reason, the same trust
* {@code CompositePeerLauncher#models0} places in config validation having already run.
*/
public Pattern patternFor(String profileName) {
if (profileName == null) {
return null;
}
FleetConfig.Profile profile = profiles.get().get(profileName);
if (profile == null || !profile.hasExhaustedPattern()) {
return null;
}
String source = profile.exhaustedPattern();
Cached cached = cache.get(profileName);
if (cached != null && cached.source().equals(source)) {
return cached.compiled();
}
Cached fresh = new Cached(source, Pattern.compile(source));
cache.put(profileName, fresh);
return fresh.compiled();
}
/** Whether {@code profileName} currently has a usage-limit pattern configured (live). */
public boolean armed(String profileName) {
return patternFor(profileName) != null;
}
}
@@ -0,0 +1,91 @@
package dev.ltms.fleet.inject;
import java.util.function.LongSupplier;
/**
* A progress watchdog for a singleton background loop (fleetd #544): {@link StatusPoller} and
* {@code dev.ltms.fleet.session.SessionReaper} each hold one. It tracks the monotonic timestamp
* of the loop's last <em>completed</em> round and classifies health from that timestamp alone —
* never from thread liveness. That choice is deliberate: measured on {@code main} at
* {@code 7611b69}, {@code UnixSocketHerdrClient.call()} reads with no deadline anywhere in
* {@code herdr/}, so a herdrd that accepts a connection and never answers parks the calling
* loop's thread forever — the thread stays alive and {@code Thread.isAlive()} keeps reporting
* {@code true} the whole time. A last-round timestamp that stops advancing catches that parked
* case exactly the same way it catches a thread that died outright: either way, nothing marks a
* new round complete.
*
* <p>Lives in {@code inject} (rather than {@code health}, which would be the more obvious home
* for a health-reporting fact) because {@code health} already depends on {@code session}
* ({@code HealthSnapshot}/{@code FleetHealthMonitor} read {@code MemberSession}), and
* {@code session} already depends on {@code inject} ({@code SessionManager} implements
* {@code TurnListener} and extends {@code MemberPresence}). Putting this class in {@code health}
* would have made {@code SessionReaper} (in {@code session}) depend back on {@code health},
* closing a cycle through {@code health -> session -> health} — {@code PackageCyclesTest} catches
* exactly this. {@code inject} has no dependency on {@code health}, so {@code session} depending
* on {@code inject} for this stays a one-way edge, same direction it already depends in.
*
* <p>Three states, not two (the fleetd #512 shape — one flag carrying two conditions that need
* opposite handling). {@code stop()} also makes the loop go quiet, exactly like a crash or a
* hang does, so a caller must say which one happened: {@link #markStoppedByCaller()} records
* "this halt was on purpose," and {@link #state()} reports {@link State#STOPPED} for it
* regardless of how stale the last round looks — an intentional stop is never an alarm.
*/
public final class LoopWatchdog {
/**
* {@code RUNNING} — the loop is going and its last round finished recently.
* {@code STALLED} — the loop was never stopped on purpose, and its last-completed-round
* timestamp is older than the threshold. Covers both a dead loop and a parked one.
* {@code STOPPED} — {@link #markStoppedByCaller()} was called. Not an alarm.
*/
public enum State { RUNNING, STALLED, STOPPED }
private final LongSupplier nowNanos;
private final long staleAfterNanos;
private volatile long lastRoundNanos;
private volatile boolean stoppedByCaller;
/**
* @param nowNanos a monotonic elapsed-time clock, e.g. {@code System::nanoTime} live, a
* controllable stub in tests.
* @param staleAfterNanos how long a last-completed-round timestamp may age before {@link #state()}
* reports {@link State#STALLED}. Derive this from the loop's own poll
* interval with generous headroom — see the caller's own derivation comment.
*/
public LoopWatchdog(LongSupplier nowNanos, long staleAfterNanos) {
this.nowNanos = nowNanos;
this.staleAfterNanos = staleAfterNanos;
this.lastRoundNanos = nowNanos.getAsLong();
}
/** Call once per completed round, from inside the loop. */
public void recordRoundComplete() {
lastRoundNanos = nowNanos.getAsLong();
}
/**
* Call from {@code start()} (and any future restart): clears a previous stop's mark and
* resets the clock, so the loop gets a fresh grace period before its first round completes
* rather than being judged against a timestamp left over from before it was (re)started.
*/
public void reset() {
stoppedByCaller = false;
lastRoundNanos = nowNanos.getAsLong();
}
/**
* Call from {@code stop()}: marks the halt as intentional, so {@link #state()} reports
* {@link State#STOPPED} rather than {@link State#STALLED} no matter how stale the last round is.
*/
public void markStoppedByCaller() {
stoppedByCaller = true;
}
/** The current fact about this loop, per the three-state contract above. */
public State state() {
if (stoppedByCaller) {
return State.STOPPED;
}
return (nowNanos.getAsLong() - lastRoundNanos >= staleAfterNanos) ? State.STALLED : State.RUNNING;
}
}
@@ -8,6 +8,8 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
/**
* Drives the {@link Injector} by sampling each active worker's {@code agent_status} and
@@ -22,11 +24,25 @@ public final class StatusPoller {
private static final Logger log = LoggerFactory.getLogger(StatusPoller.class);
/**
* How many multiples of {@code intervalMillis} the last-completed round may age before
* {@link #health()} reports {@link LoopWatchdog.State#STALLED} (fleetd #544). A round this
* loop runs is normally a small fraction of one interval — it only takes longer when a herdr
* call is genuinely wedged (herdr's {@code UnixSocketHerdrClient} has no read deadline, so a
* herdrd that accepts and never answers parks the thread forever) — so a wide multiplier
* avoids false alarms from an ordinarily slow round. At the production interval
* ({@link Injector#POLL_INTERVAL_MILLIS} = 250ms) this puts the threshold at 10s: long enough
* that a transient slow poll never trips it, short enough that a stuck poller is visible well
* before anything else downstream (delivery, `/healthz` consumers) would show a symptom.
*/
private static final long STALE_AFTER_INTERVAL_MULTIPLIER = 40;
private final AgentControl agents;
private final HerdrRouter router;
private final Injector injector;
private final StatusRefiner refiner;
private final long intervalMillis;
private final LoopWatchdog watchdog;
private volatile boolean running;
private Thread thread;
@@ -36,14 +52,26 @@ public final class StatusPoller {
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis) {
this(agents, injector, refiner, intervalMillis, System::nanoTime);
}
/** Full constructor — for tests: an injectable monotonic clock (fleetd #544). */
StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis, LongSupplier nowNanos) {
this.agents = agents;
this.router = null;
this.injector = injector;
this.refiner = refiner;
this.intervalMillis = intervalMillis;
this.watchdog = new LoopWatchdog(nowNanos, staleAfterNanos(intervalMillis));
}
public StatusPoller(HerdrRouter router, Injector injector, long intervalMillis) {
this(router, injector, intervalMillis, System::nanoTime);
}
/** Full constructor — for tests: an injectable monotonic clock (fleetd #544). */
StatusPoller(HerdrRouter router, Injector injector, long intervalMillis, LongSupplier nowNanos) {
this.agents = null;
this.router = router;
this.injector = injector;
@@ -53,45 +81,70 @@ public final class StatusPoller {
// against the LEAD daemon even though this field points at the member one.
this.refiner = new StatusRefiner(router.memberAgents());
this.intervalMillis = intervalMillis;
this.watchdog = new LoopWatchdog(nowNanos, staleAfterNanos(intervalMillis));
}
private static long staleAfterNanos(long intervalMillis) {
return TimeUnit.MILLISECONDS.toNanos(Math.max(intervalMillis, 1)) * STALE_AFTER_INTERVAL_MULTIPLIER;
}
/**
* This loop's progress fact (fleetd #544): {@code RUNNING}, {@code STALLED} (dead or parked —
* indistinguishable from outside, and never observed on purpose), or {@code STOPPED}
* ({@link #stop()} was called). See {@link LoopWatchdog} for why staleness — not thread
* liveness — is the signal.
*/
public LoopWatchdog.State health() {
return watchdog.state();
}
/** Start the polling loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
watchdog.reset(); // fleetd #544: a fresh grace period, not last run's stale mark/timestamp.
thread = Thread.ofVirtual().name("status-poller").start(this::loop);
log.info("status poller started (interval {}ms)", intervalMillis);
}
private void loop() {
while (running) {
Set<String> active = injector.activeTargets();
for (String target : active) {
if (!running) return;
try {
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
// CB-185: refine THROUGH the same control the raw status came from — a router
// splits lead/member targets across two herdr daemons, and reading a lead's pane
// through the (fixed) member refiner never finds it, wedging that lead at UNKNOWN.
AgentControl control = router != null ? router.agentsFor(target) : agents;
AgentStatus status = refiner.refine(target, control.status(target), control);
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
if (e.code() != null && e.code().endsWith("_not_found")) {
log.debug("target {} gone; dropping its queue", target);
injector.drop(target, e);
} else {
log.debug("status poll for {} failed (will retry): {}", target, e.getMessage());
try {
while (running) {
Set<String> active = injector.activeTargets();
for (String target : active) {
if (!running) return;
try {
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
// CB-185: refine THROUGH the same control the raw status came from — a router
// splits lead/member targets across two herdr daemons, and reading a lead's pane
// through the (fixed) member refiner never finds it, wedging that lead at UNKNOWN.
AgentControl control = router != null ? router.agentsFor(target) : agents;
AgentStatus status = refiner.refine(target, control.status(target), control);
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
if (e.code() != null && e.code().endsWith("_not_found")) {
log.debug("target {} gone; dropping its queue", target);
injector.drop(target, e);
} else {
log.debug("status poll for {} failed (will retry): {}", target, e.getMessage());
}
} catch (Throwable e) {
log.error("unexpected failure polling {}; skipping this round", target, e);
}
} catch (RuntimeException e) {
// Never let one target's unexpected error (e.g. an odd agent.get shape) kill
// the single poller thread and stall injection for every worker.
log.warn("unexpected error polling {}; skipping this round", target, e);
}
// fleetd #544: a round is "complete" only once every active target has been polled —
// a target parked mid-for-loop (control.status(target) never returning) means this
// line is never reached, so the watchdog goes stale exactly like a dead loop would.
watchdog.recordRoundComplete();
sleep();
}
sleep();
} finally {
if (running) {
log.error("status poller loop exited unexpectedly; it can be restarted");
}
running = false;
}
}
@@ -107,6 +160,7 @@ public final class StatusPoller {
/** Stop the polling loop. Idempotent. */
public synchronized void stop() {
running = false;
watchdog.markStoppedByCaller(); // fleetd #544: this halt is on purpose — health() must say STOPPED, not STALLED.
if (thread != null) thread.interrupt();
}
}
@@ -0,0 +1,29 @@
package dev.ltms.fleet.inject;
import dev.ltms.fleet.msg.TurnToken;
/**
* fleetd #556: the {@link Injector}'s own invariant — every delivered turn has a registered
* waiter — kept structural rather than delegated to a {@link TurnListener} callback that can
* throw. The {@link Injector} calls {@link #register} directly and unconditionally as part of
* delivering a turn, before it ever calls {@code turnListener.onDelivered}. A {@link TurnListener}
* that throws from every callback (a bug in an unrelated observer, e.g. session bookkeeping, or a
* fan-out object wired in the wrong order) can therefore never leave a delivered turn unregistered
* — the shape #553 could not close, since that fix could only make the reachable listener behave,
* never remove the Injector's dependency on a listener behaving at all.
*
* <p>Deliberately narrow: registration only, no I/O, nothing that scrapes a pane or can reasonably
* fail. The CB-115 staleness baseline (a herdr round-trip, allowed to fail) stays a
* {@link TurnListener#onDelivered} concern — fired by the {@link Injector} only after this call has
* already run, so its own failure cannot un-register anything.
*/
@FunctionalInterface
public interface TurnRegistrar {
/** Record {@code token}'s waiter as the turn currently in flight for {@code target}. */
void register(String target, TurnToken token);
/** No-op registrar for callers that don't need CB-106 rendezvous tracking. */
TurnRegistrar NOOP = (_, _) -> {
};
}
@@ -4,6 +4,7 @@ import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.PendingCloseMarker;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -12,9 +13,11 @@ import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import dev.ltms.fleet.peer.PeerLauncher;
/**
@@ -44,8 +47,33 @@ import dev.ltms.fleet.peer.PeerLauncher;
* privilege escalation (the tab label never granted anything a pane could take for itself; see that
* class's javadoc), but <em>staleness</em>: a label left behind by a crashed session would otherwise
* read as a live lead forever, and the lead would never be relaunched. So a lead counts as live only
* when herdr also reports a <em>running agent</em> in that tab — see {@link #liveLeads}. A labelled
* when herdr also reports a <em>running agent</em> in that tab — see {@link #countLeads}. A labelled
* tab with no agent in it is not a lead.
*
* <p><strong>fleetd #359 — a stale label used to pile up, not just mislead.</strong> Finding "not
* live" here used to mean only one thing: launch another. The old label was left exactly where it
* was, so a daemon that restarted enough times — or hit one herdr read that missed a genuinely
* running agent — accumulated one more identically-labelled dead tab per occurrence, and
* {@code LeadTabScanner} (before its own #359 fix) reported every one of them as a lead.
*
* <p><strong>Review finding 1 — closing on one reading is worse than the bug.</strong> The first
* version of this fix closed a name's dead tabs the moment a single {@link #countLeads} reading
* called them dead. The ticket's own live evidence rules that out: on a real host, {@code
* agent.list} was seen reporting "0 live" for a tab that a plain {@code ps} confirmed was running a
* real session. Closing on that reading would have destroyed the operator's actual lead — a worse
* failure than the extra tab it replaces. So {@link #ensureLeads()} now needs the same dead reading
* <em>twice</em>, one restart apart, before it closes anything: the first time a labelled tab reads
* dead, it is only flagged ({@link PendingCloseMarker}), left running, untouched; it is closed only
* if a <em>later</em>, independently-connected reconcile still finds it dead while the flag is
* still there. A transient miss self-heals — the next reconcile sees the agent again and clears the
* flag (see {@code toUnflag} below) — so the worst case for a single bad reading is one extra tab
* surviving one more restart, never a live session destroyed. Two alternatives were considered and
* rejected: corroborating {@code agent.list} against a second, truly independent signal was dropped
* because nothing else herdr exposes proves "is a process attached to this pane" any better — a
* second call to the same unreliable source is not independent evidence; capping the close to "all
* but the most recent dead tab" was dropped because "most recent" has no reliable ordering across
* tab ids and would leave the true failure mode (a name that is <em>never</em> reconfirmed) growing
* by one tab per bad reading forever, which is the exact defect this ticket exists to fix.
*/
public final class LeadLauncher {
@@ -79,9 +107,9 @@ public final class LeadLauncher {
return 0;
}
Map<String, Integer> live;
Map<String, LeadCount> live;
try {
live = liveLeads(leaders);
live = countLeads(leaders);
} catch (HerdrException e) {
// Counting is the whole safety mechanism against double-spawning. If we cannot count, we
// must not guess — spawning a second orchestrator is worse than starting none.
@@ -93,9 +121,53 @@ public final class LeadLauncher {
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
String name = e.getKey();
FleetConfig.Leader lead = e.getValue();
int running = live.getOrDefault(name, 0);
LeadCount state = live.getOrDefault(name, LeadCount.NONE);
int running = state.running();
int wanted = lead.instances();
// fleetd #359 review finding 1: a tab already flagged pending-close, still labelled for
// this lead, and STILL hosting no agent on this separate reconcile — two independent
// readings agree, so close it. A tab found dead for the first time is only flagged below,
// never closed on the spot.
for (String tabId : state.toClose()) {
log.info("lead '{}': closing tab {} — flagged pending-close on a previous reconcile "
+ "and still no agent running in it", name, tabId);
try {
spaces.closeTab(tabId);
} catch (RuntimeException cleanup) {
log.warn("could not close stale tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
// A tab labelled for this lead, with no agent running in it, seen dead for the first
// time — flag it rather than closing it. One reading of `agent.list` is not enough
// evidence to destroy a tab that might genuinely be live (see the class javadoc).
for (String tabId : state.toFlag()) {
String flagged = lead.tabLabel() + PendingCloseMarker.SUFFIX;
log.info("lead '{}': tab {} has no agent running in it this reconcile — flagging it "
+ "'{}' rather than closing; it is only closed if a later reconcile still "
+ "finds it dead", name, tabId, flagged);
try {
spaces.renameTab(tabId, flagged);
} catch (RuntimeException cleanup) {
log.warn("could not flag stale tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
// A previously-flagged tab that is running an agent again — the miss that flagged it was
// transient. Clear the flag so a future, unrelated miss starts its own two-reading count
// rather than closing on the strength of this one's already-spent flag.
for (String tabId : state.toUnflag()) {
log.info("lead '{}': tab {} is running an agent again — clearing its pending-close flag",
name, tabId);
try {
spaces.renameTab(tabId, lead.tabLabel());
} catch (RuntimeException cleanup) {
log.warn("could not clear the pending-close flag on tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
if (running >= wanted) {
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
continue;
@@ -126,18 +198,29 @@ public final class LeadLauncher {
}
/**
* How many live leads exist per configured name: a running agent in a tab labelled with that
* lead's exact {@code tab} (CB-579). Member workspaces are excluded, exactly as the scanner
* excludes them: a member must not be counted as a lead because it happens to sit in a matching
* tab.
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
* not be counted as a lead because it happens to sit in a matching tab.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
* pin is retired: {@code tab} is now the only field identity depends on, and {@link Agent}
* already carries {@link Agent#tabId()} directly, so a hand-opened lead is found the same way an
* auto-launched one is — by labelling its tab to match.
*
* <p>fleetd #359 review finding 1: a labelled tab with nothing running in it is split into
* {@code toClose} (already flagged pending-close by a previous reconcile, and still dead — two
* independent readings agree) and {@code toFlag} (dead for the first time — not enough evidence
* to close yet). {@code toUnflag} is the reverse: a tab flagged pending-close that is running an
* agent again, so the flag it carries no longer means anything and {@link #ensureLeads()} clears
* it.
*/
private Map<String, Integer> liveLeads(Map<String, FleetConfig.Leader> leaders) {
private record LeadCount(int running, List<String> toClose, List<String> toFlag, List<String> toUnflag) {
static final LeadCount NONE = new LeadCount(0, List.of(), List.of(), List.of());
}
private Map<String, LeadCount> countLeads(Map<String, FleetConfig.Leader> leaders) {
// A lead and the members share ONE workspace now (the operator asked for a single "session"
// with many tabs), so a workspace can no longer be excluded wholesale — the lead lives in the
// member workspace by design. The sole discriminator is the exact tab label: a lead carries
@@ -145,6 +228,7 @@ public final class LeadLauncher {
// profile's `worker: {profile} #{n}` template. These never collide, so an exact-label match
// separates them without needing to know which workspace anyone is in.
Map<String, String> nameByTab = new LinkedHashMap<>();
Set<String> flaggedTabIds = new LinkedHashSet<>();
for (Workspace ws : spaces.listWorkspaces()) {
if (ws.workspaceId() == null) {
continue;
@@ -153,31 +237,65 @@ public final class LeadLauncher {
String declared = leadNameOf(tab.label(), leaders);
if (declared != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), declared);
if (PendingCloseMarker.isFlagged(tab.label())) {
flaggedTabIds.add(tab.tabId());
}
}
}
}
Set<String> liveTabIds = new LinkedHashSet<>();
Map<String, Integer> counts = new LinkedHashMap<>();
for (Agent a : agents.list()) {
String name = nameByTab.get(a.tabId());
if (name != null) {
counts.merge(name, 1, Integer::sum);
liveTabIds.add(a.tabId());
}
}
return counts;
Map<String, List<String>> toCloseByName = new LinkedHashMap<>();
Map<String, List<String>> toFlagByName = new LinkedHashMap<>();
Map<String, List<String>> toUnflagByName = new LinkedHashMap<>();
nameByTab.forEach((tabId, name) -> {
boolean live = liveTabIds.contains(tabId);
boolean flagged = flaggedTabIds.contains(tabId);
if (live) {
if (flagged) {
toUnflagByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
}
return;
}
if (flagged) {
toCloseByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
} else {
toFlagByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
}
});
Map<String, LeadCount> out = new LinkedHashMap<>();
for (String name : leaders.keySet()) {
out.put(name, new LeadCount(counts.getOrDefault(name, 0),
toCloseByName.getOrDefault(name, List.of()),
toFlagByName.getOrDefault(name, List.of()),
toUnflagByName.getOrDefault(name, List.of())));
}
return out;
}
/**
* The configured lead a tab label names, or {@code null} for a label that names none.
*
* <p>Matched exactly (case-insensitively) against each lead's configured {@code tab}, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead. A
* trailing {@link PendingCloseMarker} is stripped first, so a tab this class flagged on a
* previous reconcile is still recognised as the same lead's tab on this one.
*/
private String leadNameOf(String label, Map<String, FleetConfig.Leader> leaders) {
if (label == null) {
return null;
}
String l = label.strip();
String l = PendingCloseMarker.strip(label);
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
String tab = e.getValue().tabLabel();
if (tab != null && l.equalsIgnoreCase(tab.strip())) {
@@ -0,0 +1,649 @@
package dev.ltms.fleet.lead;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.function.Consumer;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
* clears the lead's own pane and bootstraps a fresh session against that file.
*
* <p>This is the executor only. Nothing in this ticket wires an MCP tool onto {@link #open}/
* {@link #confirm}/{@link #cancel} — that is a separate, later unit; until it lands, nothing calls
* this class at all.
*
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING}
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()}
* itself returns. The poll always timed out — but only after the {@code /clear} had already been
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
* ever started, and a refusal return value that claimed nothing had happened.
*
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
* at all — it only records that the request is approved and hands a one-shot continuation to
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
* once the calling turn has ended, in this order:
* <ol>
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
* away live context — exactly the failure this correction exists to prevent.</li>
* <li>{@code agents.send(lead, "/clear")}</li>
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
* see {@link #waitForClearPickupAndSettle})</li>
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
* </ol>
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may
* still be mid-turn, possibly for a long time, when the caller gets that answer back.
*
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket
* that first defined this class required "no timer, no scheduler, no background thread" so that
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
* continuationRunner} launches a single-shot task that exists only because one specific,
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at
* construction time or on any schedule, and no two invocations of it ever share state. A recurring
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work
* slightly later than the method return, as a continuation of the same approved request, rather
* than entirely inside the method body.</strong>
*
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
* correction.</strong> The first version resolved the pane to clear via {@code
* PrimaryRegistry#primaryTerminal()}. That is correct for a background loop with no caller (see
* {@code dev.ltms.fleet.msg.LeadHeartbeatLoop}), but wrong here and a violation of this project's
* own charter invariant 3 — "identity comes from the connection, never an argument." This daemon
* can hold more than one labelled lead tab (see {@code LeadLauncher}'s fleetd #359 two-reading
* dead-tab cleanup), so a single-slot lookup lets lead X's {@link #confirm} clear lead Y's pane: an
* unrecoverable loss of someone else's context, and a different lead's context at that. Both
* {@link #open} and {@link #confirm} now take the lead's terminal id as a parameter instead —
* {@link #open} stores it on the {@link PendingRollover}, and {@link #confirm} refuses with {@link
* RefusalReason#NOT_YOUR_ROLLOVER} unless the caller's terminal matches the one {@link #open}
* recorded. <strong>The terminal id passed to both methods must come from the MCP layer's own
* connection-based caller resolution — the same source {@code auth/CallerResolver#resolve} uses to
* build a {@code Principal.leader(...)} (see its use of {@code ConnectionIdentity.Caller#terminal})
* — never a value the client supplies or chooses.</strong> The later MCP-tool unit that wires
* {@link #open}/{@link #confirm} must pass the resolved caller terminal, not a request field.
*
* <p><strong>A relative {@code handoverPath} resolves against the CALLING lead's workspace, never
* the daemon's own cwd — a fleetd #480 follow-up.</strong> The daemon and a lead's own pane can
* have different working directories (this repo nests {@code fleetd/} inside its own root, so the
* daemon's cwd and {@code fleet.leaders.<name>.cwd} already differ on this host). {@link #open}
* resolves {@code cfg.handoverPath()} to an ABSOLUTE path exactly once — against {@code
* leadWorkspace.apply(leadTerminal)} when that lookup returns a non-null, non-blank workspace, and
* against {@code System.getProperty("user.dir")} otherwise (the same fallback {@code
* LeadLauncher#launch} already uses for a lead with no configured {@code cwd}) — and stores only
* that absolute path on {@link PendingRollover}. Every later read of {@code
* PendingRollover#handoverPath()} (the freshness/exists/empty checks in {@link #checkHandover},
* the value {@code FleetMcp} hands back to the lead in the {@code open} response so it knows where
* to WRITE the file, and {@link FleetConfig.LeadRollover#bootstrapTextFor} which names it in the
* text typed into the fresh session) therefore already sees the resolved absolute form and never
* needs to resolve anything itself.
*/
public final class LeadRollover {
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
static final long SETTLE_POLL_MS = 250;
/**
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
* before releasing rather than wedging the roll — the same constant and the same
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
* nudging again — so 8 polls produce 7 nudges, not 8.
*/
static final int PICKUP_GRACE_POLLS = 8;
/**
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
*
* @param leadTerminal the lead pane that opened this request — the only terminal that may
* later {@link #confirm} it (see {@link RefusalReason#NOT_YOUR_ROLLOVER})
* @param handoverPath the ABSOLUTE, resolved handover path — never the raw configured value,
* which may have been relative. {@link #open} resolves it once, against the
* calling lead's workspace, before storing it here; see this class's
* javadoc. This is the value the MCP layer hands back to the lead as
* "write your file here", so callers may rely on it always being absolute.
*/
public record PendingRollover(String token, String leadTerminal, String handoverPath,
long requestedAtMillis) {}
/** Which check refused a {@link #confirm} call, named so a caller can act on it. */
public enum RefusalReason {
/** {@code leadRollover:} is not configured — absent at construction, or removed since. */
NOT_CONFIGURED,
/** {@code token} names no pending request: never opened, already confirmed, or cancelled. */
UNKNOWN_TOKEN,
/**
* The caller's terminal does not match the terminal that {@link #open} recorded for this
* token. Only the lead that opened a request may confirm it (fleetd #480 correction 2).
*/
NOT_YOUR_ROLLOVER,
/** {@code requireOperatorConfirm: true} and the caller passed {@code operatorConfirmed: false}. */
OPERATOR_NOT_CONFIRMED,
/** The handover file does not exist. */
HANDOVER_MISSING,
/** The handover file exists but is empty. */
HANDOVER_EMPTY,
/**
* The handover file's modified time is not after {@link #open}'s request timestamp, or is
* older than {@code maxDocAgeSeconds}.
*/
HANDOVER_STALE
}
/**
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been
* cleared; the continuation may still be waiting for the calling turn to end when this returns.
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
*/
public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
static RollDecision approved() {
return new RollDecision(true, null, "confirmed; the roll will run once the calling turn ends");
}
static RollDecision refused(RefusalReason reason, String detail) {
return new RollDecision(false, reason, detail);
}
}
private final AgentControl agents;
private final Supplier<FleetConfig.LeadRollover> configSupplier;
/**
* Terminal id → that lead's configured workspace directory (their {@code
* fleet.leaders.<name>.cwd}), or {@code null} when the terminal names no currently-recognised
* lead. {@link #open} calls this to resolve a relative {@code handoverPath} — see this class's
* javadoc. Required: there is no sane default that would not silently reintroduce the
* daemon-cwd bug this parameter exists to fix.
*/
private final Function<String, String> leadWorkspace;
private final LongSupplier nowMillis;
private final Runnable settleSleeper;
/**
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
* thread per confirmed request — see this class's javadoc for why that is a single-shot task,
* not a background scheduler. Tests inject {@code Runnable::run} to make the continuation run
* synchronously and deterministically on the calling thread.
*/
private final Consumer<Runnable> continuationRunner;
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace) {
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis,
() -> sleepUninterruptibly(SETTLE_POLL_MS),
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
}
/**
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
* against a file's modified time, which only a wall clock is comparable to, and {@code
* nanoTime} freezes while the host sleeps (fleetd #386).
*/
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace, LongSupplier nowMillis,
Runnable settleSleeper, Consumer<Runnable> continuationRunner) {
this.agents = agents;
this.configSupplier = configSupplier;
this.leadWorkspace = leadWorkspace;
this.nowMillis = nowMillis;
this.settleSleeper = settleSleeper;
this.continuationRunner = continuationRunner;
}
private static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — this poll loop should not be aborted by an
// interrupt that was not meant for it.
}
}
/**
* The lead says it is ready to be replaced. Generates a token and records the resolved
* handover path, this moment's wall-clock timestamp (the baseline {@link #confirm} checks the
* handover file's modified time against), and {@code leadTerminal} — only that exact terminal
* may later {@link #confirm} this token.
*
* @param leadTerminal the calling lead's terminal id, resolved by the MCP layer from the
* connection (see this class's javadoc) — never a client-supplied value
* @param reason free-text audit note (logged only; not otherwise interpreted or stored)
* @throws IllegalStateException if {@code leadRollover:} is not configured
* @throws IllegalArgumentException if {@code leadTerminal} is null or blank
*/
public PendingRollover open(String leadTerminal, String reason) {
FleetConfig.LeadRollover cfg = configSupplier.get();
if (cfg == null) {
throw new IllegalStateException("leadRollover: is not configured");
}
if (leadTerminal == null || leadTerminal.isBlank()) {
throw new IllegalArgumentException("leadTerminal is required — it must be resolved from "
+ "the caller's connection, never accepted as a client-chosen argument");
}
String token = UUID.randomUUID().toString();
long requestedAt = nowMillis.getAsLong();
String resolvedPath = resolveHandoverPath(cfg.handoverPath(), leadTerminal);
PendingRollover p = new PendingRollover(token, leadTerminal, resolvedPath, requestedAt);
pending.put(token, p);
if (resolvedPath.equals(cfg.handoverPath())) {
log.info("lead-rollover: open token={} lead={} handoverPath={} reason={}",
token, leadTerminal, resolvedPath, reason);
} else {
// The configured value was relative (or merely un-normalized) and resolved to a
// different string — log both, so an operator reading this line can see which
// directory the daemon actually looked in, not just the value it was given.
log.info("lead-rollover: open token={} lead={} configuredHandoverPath={} "
+ "resolvedHandoverPath={} reason={}",
token, leadTerminal, cfg.handoverPath(), resolvedPath, reason);
}
return p;
}
/**
* Resolve {@code configured} to an absolute path exactly once, here, so nothing downstream
* ({@link #checkHandover}, the MCP layer's {@code open} response, {@link
* FleetConfig.LeadRollover#bootstrapTextFor}) ever has to resolve — or worse, silently
* mis-resolve — a relative path again.
*
* <p><strong>The return value is GUARANTEED absolute, not merely usually absolute.</strong>
* {@code leadWorkspace.apply(leadTerminal)} returns an operator-configured {@code
* fleet.leaders.<name>.cwd} string, and nothing forces an operator to write an absolute one —
* a relative {@code cwd} resolved with plain {@link Path#resolve} would still yield a relative
* result, silently reopening the exact bug this class exists to fix (every later reader back to
* interpreting an ambiguous string against ITS OWN working directory). {@link
* Path#toAbsolutePath()} closes that: it resolves any remaining relative path against {@code
* user.dir} (the JVM's own cwd), which is the correct base for an operator-written path the
* daemon process itself is meant to interpret, exactly like the {@code user.dir} fallback used
* below. Applying it unconditionally on both branches means the ALREADY-absolute branch stays a
* no-op (an absolute path is unaffected by {@code toAbsolutePath()}) while the relative-{@code
* cwd} branch above is closed the same way.
*
* <ul>
* <li>already absolute → returned unchanged (normalized)</li>
* <li>relative → resolved against {@code leadWorkspace.apply(leadTerminal)} when that is
* non-null and non-blank; otherwise against {@code System.getProperty("user.dir")} — the
* same fallback {@code LeadLauncher#launch} uses for a lead with no configured {@code
* cwd}. If {@code leadWorkspace}'s own answer is itself relative (an operator wrote a
* relative {@code cwd:}), the result is finished off against the daemon's own
* {@code user.dir} — see the paragraph above.</li>
* </ul>
*/
private String resolveHandoverPath(String configured, String leadTerminal) {
Path path = Path.of(configured);
if (path.isAbsolute()) {
// toAbsolutePath() is a no-op for an already-absolute path — kept here anyway so both
// branches call the exact same guarantee, rather than one branch relying on
// isAbsolute() alone to already imply what toAbsolutePath() enforces.
return path.toAbsolutePath().normalize().toString();
}
String workspace = leadWorkspace.apply(leadTerminal);
Path base = (workspace == null || workspace.isBlank())
? Path.of(System.getProperty("user.dir"))
: Path.of(workspace);
return base.resolve(path).toAbsolutePath().normalize().toString();
}
/**
* Validate every gate, then — if and only if all of them pass — hand a one-shot continuation
* that performs the actual roll to {@code continuationRunner} and return. <strong>This method
* never itself sends anything to the lead's pane</strong> — see this class's javadoc for why
* (it is called FROM the lead's own turn, so the pane cannot possibly be injectable yet).
*
* <p>Order: token lookup, then ownership ({@code callerTerminal} must match the terminal
* {@link #open} recorded — {@link RefusalReason#NOT_YOUR_ROLLOVER}), then the
* operator-confirmation gate, then the three handover-file checks (exists, not empty, fresh —
* see {@link #checkHandover}). The first failing check is returned and {@code token} stays
* pending (so a caller can fix the problem — e.g. rewrite the handover file — and retry with
* the same token); it is consumed only once every gate passes and the continuation is launched.
*
* @param callerTerminal the CALLING lead's terminal id, resolved by the MCP layer from the
* connection — never a client-supplied value (see this class's javadoc)
* @param token the token {@link #open} returned
* @param operatorConfirmed the caller's answer to "has an operator confirmed this roll" —
* consulted only when the live config's {@code requireOperatorConfirm}
* is true
*/
public RollDecision confirm(String callerTerminal, String token, boolean operatorConfirmed) {
FleetConfig.LeadRollover cfg = configSupplier.get();
if (cfg == null) {
return RollDecision.refused(RefusalReason.NOT_CONFIGURED, "leadRollover: is not configured");
}
PendingRollover p = pending.get(token);
if (p == null) {
return RollDecision.refused(RefusalReason.UNKNOWN_TOKEN,
"token " + token + " names no pending rollover request");
}
if (!p.leadTerminal().equals(callerTerminal)) {
return RollDecision.refused(RefusalReason.NOT_YOUR_ROLLOVER,
"token " + token + " was opened by a different lead terminal");
}
if (cfg.requireOperatorConfirm() && !operatorConfirmed) {
return RollDecision.refused(RefusalReason.OPERATOR_NOT_CONFIRMED,
"requireOperatorConfirm is true and operatorConfirmed was false");
}
RollDecision docCheck = checkHandover(p, cfg);
if (docCheck != null) {
return docCheck;
}
pending.remove(token);
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
token, callerTerminal);
continuationRunner.accept(() -> runRollover(p, cfg));
return RollDecision.approved();
}
/**
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
* four-step order. There is no result to return to by this point, so every outcome is logged
* only.
*/
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
String lead = p.leadTerminal();
long rollStartMillis = nowMillis.getAsLong();
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
if (!turnResult.settled()) {
// fleetd #494 follow-up: this line had the SAME defect as the /clear-timeout line below
// — cfg.turnSettleSeconds() is the CONFIGURED budget, not how long this wait actually
// ran. Print the measured elapsed time alongside it, labelled, exactly like the /clear
// path already does.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "confirm() — refusing to send /clear at all; the calling lead's own "
+ "turn is still live and clearing it now would destroy live context "
+ "(token={}, configured={}s elapsed={}ms)",
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
return;
}
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext:
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
// pane forever (see this class's javadoc).
agents.send(lead, "/clear");
ClearSettleResult clearResult = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
if (!clearResult.settled()) {
// fleetd #494: cfg.clearSettleSeconds() is the CONFIGURED budget, not how long the wait
// actually ran — an operator reading only that number wrongly believes it is a measured
// duration. Print the measured elapsed time and nudge count alongside it, each labelled,
// so the two can be compared at a glance.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "/clear — NOT sending bootstrapText (token={}, configured={}s "
+ "elapsed={}ms nudges={})",
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
clearResult.nudges());
return;
}
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
}
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
public boolean cancel(String token) {
return pending.remove(token) != null;
}
/**
* The three handover-file checks, in order: exists, not empty, fresh (modified after
* {@link #open}'s timestamp and not older than {@code maxDocAgeSeconds}). Stats {@code
* p.handoverPath()} directly — {@link #open} already resolved it to an absolute path, so this
* never has to guess which directory it means.
*
* @return the first failing check's refusal, or {@code null} when all three pass
*/
private RollDecision checkHandover(PendingRollover p, FleetConfig.LeadRollover cfg) {
Path path = Path.of(p.handoverPath());
if (!Files.exists(path)) {
return RollDecision.refused(RefusalReason.HANDOVER_MISSING,
"handover file " + p.handoverPath() + " does not exist");
}
long size;
long mtimeMillis;
try {
size = Files.size(path);
mtimeMillis = Files.getLastModifiedTime(path).toMillis();
} catch (IOException e) {
throw new UncheckedIOException("failed to stat handover file " + p.handoverPath(), e);
}
if (size == 0) {
return RollDecision.refused(RefusalReason.HANDOVER_EMPTY,
"handover file " + p.handoverPath() + " is empty");
}
if (mtimeMillis <= p.requestedAtMillis()) {
return RollDecision.refused(RefusalReason.HANDOVER_STALE,
"handover file " + p.handoverPath() + " was not modified after the open() "
+ "request (mtime=" + mtimeMillis + "ms, requestedAt=" + p.requestedAtMillis() + "ms)");
}
long ageMillis = nowMillis.getAsLong() - mtimeMillis;
long maxAgeMillis = TimeUnit.SECONDS.toMillis(cfg.maxDocAgeSeconds());
if (ageMillis > maxAgeMillis) {
return RollDecision.refused(RefusalReason.HANDOVER_STALE,
"handover file " + p.handoverPath() + " is " + TimeUnit.MILLISECONDS.toSeconds(ageMillis)
+ "s old, older than maxDocAgeSeconds=" + cfg.maxDocAgeSeconds());
}
return null;
}
/**
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
*
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
* turn" — and it accepts {@link AgentStatus#BLOCKED} for that purpose, because a pane paused on
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
* reason, on the second wait.)
*
* @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real
* boundary was observed, {@code false} if {@code settleSeconds} elapses first.
* {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis}
* clock, never the configured {@code settleSeconds} budget (fleetd #494 follow-up).
*/
private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle: {}",
target, e.toString());
status = null;
}
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis);
}
settleSleeper.run();
}
return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis);
}
/**
* The measured outcome of {@link #waitUntilAtTurnBoundary} — fleetd #494 follow-up. The sibling
* of {@link ClearSettleResult} for the FIRST wait, which never nudges, so it carries no nudge
* count.
*/
private record TurnSettleResult(boolean settled, long elapsedMillis) {}
/**
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
* own javadoc already records that the submit accompanying a delivery "can race the paste —
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
* it, landing as one concatenated line.
*
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
* <ul>
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
* turn;</li>
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
* Claude Code prompt is a no-op, so repeating it is safe;</li>
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
* observed, releases rather than wedges the roll instead of nudging again — the same
* choice {@code Injector} makes — and returns {@code settled() == true} anyway, logged at
* {@code warn} with the measured elapsed time (fleetd #494) so an operator can see which
* path ran and how long it actually took;</li>
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
* </ul>
*
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
* a real boundary is reached or {@code settleSeconds} runs out.
*
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
* nudge must not abort the roll.
*
* @return a {@link ClearSettleResult} whose {@code settled()} is {@code true} once {@code
* /clear} has settled, or once the nudge budget was exhausted with no pickup ever
* observed (released rather than wedged); {@code false} if {@code settleSeconds} elapses
* first — the caller must NOT send {@code bootstrapText} in that case, exactly as before
* this fix. {@code elapsedMillis()} and {@code nudges()} are MEASURED values (from the
* injected {@link #nowMillis} clock and an actual nudge count), never the configured
* {@code settleSeconds} budget (fleetd #494).
*/
private ClearSettleResult waitForClearPickupAndSettle(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
int idlePollsAwaitingPickup = 0;
int nudges = 0;
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
+ "/clear: {}", target, e.toString());
status = null;
}
if (status == AgentStatus.WORKING) {
pickedUp = true;
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
if (pickedUp) {
// a real WORKING -> IDLE/DONE completion boundary
return new ClearSettleResult(true, nowMillis.getAsLong() - startMillis, nudges);
}
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
long elapsedMillis = nowMillis.getAsLong() - startMillis;
// fleetd #494: this release trades a possibly-unsubmitted /clear for progress
// instead of wedging the roll — that trade is deliberate and stays. But it is
// also exactly the case that reported false success in the real incident (the
// whole roll "succeeded" after 438ms of a 20s budget), so raise it to WARN and
// print the MEASURED elapsed time next to the target pane, not just the count.
//
// fleetd #494 follow-up (2nd pass): BOTH numbers in this line must come from
// the loop's own counters, never from the PICKUP_GRACE_POLLS constant.
// `idlePollsAwaitingPickup` and `nudges` each have exactly one write site in
// this loop, on the same branch, so on this branch they cannot differ from
// PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 today — no test can prove the
// difference on this line, and printing the counters does not change that.
// What it does buy: one source of truth instead of two, so a later change to
// the loop (an early return, a second increment site, a different exit
// condition) cannot leave this message reporting a number the loop no longer
// produces. The place where `nudges` genuinely varies with the run — and is
// covered by a test that can tell it apart from a constant — is the
// /clear-timeout warn in runRollover, which prints clearResult.nudges().
log.warn("lead-rollover: /clear on {} was never observed as WORKING after {} "
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
+ "releasing rather than wedging the roll (elapsed={}ms)",
target, idlePollsAwaitingPickup, nudges, elapsedMillis);
return new ClearSettleResult(true, elapsedMillis, nudges);
}
try {
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
} catch (RuntimeException e) {
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
target, e.getMessage());
} finally {
nudges++; // an attempted nudge, whether or not the submit call itself threw
}
}
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
// signal nor a boundary — keep polling without nudging or releasing.
settleSleeper.run();
}
return new ClearSettleResult(false, nowMillis.getAsLong() - startMillis, nudges);
}
/**
* The measured outcome of {@link #waitForClearPickupAndSettle} — fleetd #494. Carries the
* MEASURED elapsed time (from the injected {@link #nowMillis} clock) and nudge count alongside
* the settle/timeout decision, so callers can log them instead of the configured budget, which
* is not how long the wait actually ran.
*/
private record ClearSettleResult(boolean settled, long elapsedMillis, int nudges) {}
}
@@ -0,0 +1,81 @@
package dev.ltms.fleet.mcp;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* fleetd #469: a launch charter that names an MCP tool the server does not register must stop the
* daemon at startup, not wait for a member to discover the gap by calling something that is not
* there.
*
* <p>Deliberately its own class outside {@code dev.ltms.fleet.config}, not a case in {@link
* dev.ltms.fleet.config.FleetConfig#validateCharters()}. The canonical tool surface ({@link
* FleetTool}) lives in the {@code mcp} package; config is loaded before the MCP server exists and
* must not gain a dependency on it. So this check belongs at the seam that already holds both a
* loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}, called right after
* {@code cfg.validateAll()} and before anything opens a socket or spawns a member.
*
* <p>{@code #464}'s {@code CharterToolSurfaceTest} proved the same comparison against a charter
* fixture it wrote itself into a {@code @TempDir}, which meant nothing anyone wrote into the live
* {@code fleetd.yaml} could ever fail it. This class is what a real charter is actually checked
* against at boot; {@code FleetdStartupValidationTest} exercises it through {@code Fleetd.main}
* itself, the same way it proves every other {@code validateXxx()} still runs there.
*
* <p><strong>fleetd #474</strong> — startup was not the only door: {@code
* dev.ltms.fleet.config.ConfigRef#reload()} used to run {@code FleetConfig#validateAll()} alone,
* which does not look at what a charter's text names, so a charter naming an unregistered tool
* that could not have booted the daemon could still be installed into a running one through a
* reload. This class still knows nothing about {@code ConfigRef} — {@code Fleetd.main} wires a
* small {@code FleetConfig -> void} adapter over {@link #assertChartersNameOnlyRegisteredTools}
* ({@code Fleetd::assertChartersNameOnlyRegisteredTools}) into {@code ConfigRef}'s constructor as
* its {@code Consumer<FleetConfig>} {@code extraValidation}, run inside {@code reload()}'s same
* try/catch as {@code validateAll()}, so both call sites — {@code Fleetd.main} at startup and
* {@code ConfigRef#reload()} afterwards — go through this one method and can never check different
* things. {@code dev.ltms.fleet.config.ConfigRefTest} and {@code
* FleetdConfigRefCharterToolSurfaceWiringTest} are what prove the reload call site, the same way
* {@code FleetdStartupValidationTest} proves the startup one.
*/
public final class CharterToolSurface {
/** A {@code fleet_…} (current) or {@code bridge_…} (pre-CB-634) tool-shaped token in prose. */
private static final Pattern TOOL_REFERENCE = Pattern.compile("(fleet_[a-z_]+|bridge_[a-z_]+)");
private CharterToolSurface() {
}
/**
* @param charters the configured {@code fleet.charters:} map (role wire name → charter text);
* {@code null} or empty is a no-op, same as an absent {@code fleet:} block
* @throws IllegalStateException naming the charter key and every tool it names that {@link
* FleetTool} does not list, when any charter does so
*/
public static void assertChartersNameOnlyRegisteredTools(Map<String, String> charters) {
if (charters == null || charters.isEmpty()) {
return;
}
Set<String> registered = FleetTool.wireNames();
List<String> bad = new ArrayList<>();
charters.forEach((key, text) -> {
if (text == null) {
return;
}
Set<String> named = new LinkedHashSet<>();
Matcher m = TOOL_REFERENCE.matcher(text);
while (m.find()) {
named.add(m.group(1));
}
named.stream()
.filter(t -> !registered.contains(t))
.forEach(unknown -> bad.add("fleet.charters." + key + " names '" + unknown
+ "', which the server does not register (registered: " + registered + ")."));
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
}
@@ -33,9 +33,10 @@ public final class ConnectionIdentity {
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
* the pane scan behind {@code terminal} ran to completion ({@link #scanComplete}).
*/
public record Caller(String terminal, long pid) {
public record Caller(String terminal, long pid, boolean scanComplete) {
/**
* Whether the OS peer-PID lookup actually succeeded — {@code false} means {@code pid} is
@@ -51,6 +52,10 @@ public final class ConnectionIdentity {
* {@link ConnectionIdentity#isLoopback} is centralised rather than left for each caller to
* reimplement: a raw {@code pid > 0} check duplicated at every call site is precisely the
* "one rule, two copies" shape that let #305 drift.
*
* <p>This method is deliberately NOT widened for fleetd #505's failure (a herdr error
* during the pane scan, not a failed lsof lookup) — it still tests only the sentinel it is
* named for. #505 is a different axis, carried separately in {@link #scanComplete}.
*/
public boolean resolved() {
return pid > 0;
@@ -60,10 +65,11 @@ public final class ConnectionIdentity {
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
public Caller resolve(String remoteAddr, int remotePort) {
if (!isLoopback(remoteAddr)) {
return new Caller(null, -1); // only same-host callers can be workers
return new Caller(null, -1, true); // only same-host callers can be workers
}
long pid = pids.pidForLocalPort(remotePort);
return new Caller(panes.terminalForPid(pid), pid);
PaneLocator.Lookup lookup = panes.terminalForPid(pid);
return new Caller(lookup.terminal(), pid, lookup.complete());
}
/**
@@ -90,13 +96,27 @@ public final class ConnectionIdentity {
* {@code 127.0.0.1:8765} with a source address of {@code 127.0.0.2} — measured on the Linux
* fleet host, where binding that source succeeds.
*
* <p><strong>Being strict here does not make the daemon safer; it makes it unsafe.</strong>
* That reads backwards, so it is worth stating plainly. This predicate does not decide whether
* a caller is trusted — it decides whether the caller's identity is <em>resolved at all</em>.
* Returning false means {@link #resolve} answers "no terminal", and downstream a caller with no
* terminal is treated as the primary under loopback-trust. So every address excluded here is an
* address on which a worker silently becomes the lead. Widening a check normally weakens it;
* widening this one is what closes the hole.
* <p><strong>What excluding an address costs, stated as it is today.</strong> This paragraph
* used to say that narrowing this range turned a worker into the lead, and that widening the
* check was what closed the hole. That was true only while there were <em>two</em> definitions
* that disagreed: {@code ConnectionIdentity} skipped the identity lookup for {@code 127.0.0.2}
* while {@code CallerResolver} read the same address as loopback and granted the primary role.
* #305 removed the second copy, and with one shared definition the old sentence no longer holds.
*
* <p>Measured on 2026-09-04 by narrowing this method back to exactly {@code 127.0.0.1} and
* running {@code CallerResolverTest} and {@code ConnectionIdentityTest}: a caller from
* {@code 127.0.0.2} then resolves to {@code ANONYMOUS}, not {@code PRIMARY} — for a worker
* ({@code aWorkerOnAnyLoopbackSourceAddressIsStillAWorkerNotThePrimary}) and for a non-worker
* ({@code aNonWorkerOnAnyLoopbackSourceAddressIsStillThePrimary}) alike. Excluding an address
* now <em>refuses</em> its caller; it does not promote one.
*
* <p>So keep the whole range, but for the plain reason: a genuine worker or primary that
* connects from {@code 127.0.0.2} must be identifiable at all, and narrowing this predicate
* locks it out. That is an outage, and an outage is the direction to fail in — which is exactly
* why the range must not be narrowed casually and also why doing so is no longer a security
* hole. This predicate still does not decide whether a caller is trusted; it decides whether the
* caller's identity is <em>resolved at all</em>. What makes an unresolved caller safe is
* {@link Caller#resolved()} (#317), not this method.
*/
public static boolean isLoopback(String addr) {
if (addr == null) {
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,83 @@
package dev.ltms.fleet.mcp;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
/**
* The one canonical set of MCP tool names this daemon registers (fleetd #469, follow-up to #464).
*
* <p>Before this enum, the tool surface was written twice with nothing tying the copies together:
* once as the literal {@code "fleet_…"} string passed to each tool-schema builder in
* {@link FleetMcp}, and again as the case labels of {@link FleetMcp}'s authorization switch. A
* reader that needed "what does this server register" — a charter check, in particular — had no
* source to ask except scraping {@code FleetMcp.java}'s source text for {@code tool("…")} calls: a
* third copy of the same list, and the weakest of the three forms.
*
* <p>Every reader that needs the registered tool surface now asks this enum instead:
*
* <ul>
* <li>the tool-schema builders in {@code FleetMcp} pass {@code wireName()} rather than a literal;
* <li>{@code FleetMcp}'s constructor asserts, at startup, that the set of tool names it actually
* registers with the MCP SDK equals {@link #wireNames()} exactly — a canonical entry that is
* never registered, or a registration with no canonical entry backing it, fails the daemon's
* own boot rather than only a test's;
* <li>{@code FleetMcp.toolAction}'s dispatch onto {@code Authz.Action} switches on the enum
* (not the raw string) with no {@code default}, so adding a tool here without pinning its
* action is a compile error, not a run-time throw;
* <li>{@link CharterToolSurface} asks {@link #wireNames()} to check a configured launch charter
* against the live tool surface, instead of scraping source text a third time.
* </ul>
*/
public enum FleetTool {
SEND("fleet_send"),
REPLY("fleet_reply"),
ASK("fleet_ask"),
STATUS("fleet_status"),
POLL("fleet_poll"),
ACK("fleet_ack"),
SPAWN("fleet_spawn"),
LIST("fleet_list"),
STOP("fleet_stop"),
PROFILES("fleet_profiles"),
WHOAMI("fleet_whoami"),
HANDOVER("fleet_handover");
private final String wireName;
FleetTool(String wireName) {
this.wireName = wireName;
}
/** The name this tool is registered under, and called by, on the wire ({@code "fleet_send"}, …). */
public String wireName() {
return wireName;
}
private static final Map<String, FleetTool> BY_WIRE_NAME;
private static final Set<String> WIRE_NAMES;
static {
Map<String, FleetTool> byName = new LinkedHashMap<>();
Set<String> names = new LinkedHashSet<>();
for (FleetTool tool : values()) {
byName.put(tool.wireName, tool);
names.add(tool.wireName);
}
BY_WIRE_NAME = Map.copyOf(byName);
WIRE_NAMES = Set.copyOf(names);
}
/** The tool named {@code wireName}, or empty when this daemon registers no such tool. */
public static Optional<FleetTool> byWireName(String wireName) {
return Optional.ofNullable(BY_WIRE_NAME.get(wireName));
}
/** Every wire name this daemon registers — the canonical tool surface. */
public static Set<String> wireNames() {
return WIRE_NAMES;
}
}
@@ -14,6 +14,7 @@ import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementCandidate;
import dev.ltms.fleet.placement.PlacementContext;
import dev.ltms.fleet.placement.PlacementDecision;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.placement.PlacementPolicies;
import dev.ltms.fleet.placement.PlacementPolicy;
@@ -90,6 +91,19 @@ public final class CompositePeerLauncher implements PeerLauncher {
private final Supplier<Map<String, FleetConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/**
* fleetd #422: the central model allow-list's on/off state, read live per spawn — same reason
* {@link #profileConfigs} is a supplier rather than a captured map (see the class doc above and
* {@link #enforceModelEnabled}). A caller with no {@code models:} block to read from (the
* simpler, map-based constructors used throughout this class's own tests) wires this to a
* constant {@code null}, which {@link #models0} treats as "nothing configured, gate never
* fires" — the pre-#422 behaviour.
*/
private final Supplier<FleetConfig.Models> models;
/** The value {@link #models0} normalizes a {@code null} supplier result to. */
private static final FleetConfig.Models NO_MODELS_CONFIGURED = new FleetConfig.Models(List.of());
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
private final BackendQuarantine quarantine;
@@ -205,12 +219,44 @@ public final class CompositePeerLauncher implements PeerLauncher {
FleetConfig.Fleet fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy) {
// No models: block to read from a plain profiles Map — fleetd #422's gate is wired to a
// constant null (see enforceModelEnabled/models0), the pre-#422 behaviour for every caller
// of this overload. Use the 9-arg overload below to test the gate against a LIVE supplier.
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, quarantine,
outagePolicy, constant(null));
}
/**
* As above, plus a LIVE model-gate source (fleetd #422). The 8-arg overload above wires
* {@code models} to a constant {@code null} because it has only a static {@code Map<String,
* Profile>}, never a full {@code FleetConfig}, to read one from; this overload exists so a test
* can prove {@link #enforceModelEnabled} and its candidate filter re-read {@link
* FleetConfig.Models} on every call rather than a value captured once at construction — the
* exact distinction fleetd #422 exists to get right (see the class doc's CB-559 note on {@link
* #profileConfigs}, which this follows). Production wiring uses the
* {@code Supplier<FleetConfig>} constructor below instead, which already threads a live
* {@code config.get().models()} through.
*
* @param models required — pass a supplier returning {@code null} for a caller that has no
* {@code models:} block to gate against, never a defaulting overload (the same
* "explicit opt-out, never a silent default" rule {@code quarantine}/
* {@code outagePolicy} already follow).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, FleetConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
FleetConfig.Fleet fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy,
Supplier<FleetConfig.Models> models) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy);
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy, models);
}
/**
@@ -250,7 +296,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
liveCount,
() -> config.get().fleet(),
quarantine,
outagePolicy);
outagePolicy,
// fleetd #422: read live, same as profiles/placement/fleet above — a models.allow
// edit (on/off or otherwise) is visible to the very next spawn, no restart needed.
() -> config.get().models());
}
/** The all-suppliers form every other constructor funnels into. */
@@ -261,7 +310,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
Function<String, Integer> liveCount,
Supplier<FleetConfig.Fleet> fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy) {
BackendOutagePolicy outagePolicy,
Supplier<FleetConfig.Models> models) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.outagePolicy = Objects.requireNonNull(outagePolicy, "outagePolicy");
@@ -273,6 +323,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
this.profileConfigs = profileConfigs;
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
this.models = Objects.requireNonNull(models, "models");
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
@@ -303,6 +354,16 @@ public final class CompositePeerLauncher implements PeerLauncher {
return m == null ? Map.of() : m;
}
/**
* The currently-configured {@code models:} block, never null (fleetd #422). Read fresh on
* every call, the same reason {@link #profiles0} is — a config reload's on/off edit must reach
* the very next spawn.
*/
private FleetConfig.Models models0() {
FleetConfig.Models m = models.get();
return m == null ? NO_MODELS_CONFIGURED : m;
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
@@ -332,6 +393,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
enforceNotQuarantined(requestedProfile);
enforceNotCoolingOff(requestedProfile);
enforceMaxLoad(requestedProfile);
enforceModelEnabled(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
@@ -341,18 +403,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
// operator overriding, and refusing it would break `fleet_spawn{profile:"opus"}`, which
// carries no role and so would be judged against the dev pool it was never meant for.
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
Set<String> unreachable = new HashSet<>();
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
Set<String> coolingOff = coolingOffProfiles(candidates);
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff);
PlacementContext ctx = placementContextFor(req.role(), unreachable);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
int maxAttempts = ctx.candidates().isEmpty() ? 1 : ctx.candidates().size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
// policy already throws a clear message. Catching it to rethrow a generic
@@ -369,8 +423,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
// CB-547a: route the chosen profile but keep the caller's session identity — dropping it
// here would silently sever the resume handle on every policy-routed spawn. CB-557: the
// role rides along for the same reason, or a routed spawn would be labelled as a dev.
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId(), req.role());
SpawnRequest routedReq = req.withProfile(chosen.profile());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
@@ -380,14 +433,16 @@ public final class CompositePeerLauncher implements PeerLauncher {
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff);
ctx = placementContextFor(req.role(), unreachable);
}
}
// unreachable.size() counts DISTINCT profiles, not attempts (a HashSet dedupes a profile
// added twice) — say "distinct" so the count matches the sentence and the profile list that
// follows, rather than reading as a count of attempts made (fleetd #315).
throw new PeerUnreachableException(
"no reachable worker profile available after trying " + unreachable.size()
+ " candidate(s): " + String.join(", ", unreachable));
+ " distinct candidate(s): " + String.join(", ", unreachable));
}
/**
@@ -490,6 +545,49 @@ public final class CompositePeerLauncher implements PeerLauncher {
}
}
/**
* Refuse an explicit-profile spawn whose {@code model:} the operator has turned off in the
* central {@code models.allow:} list (fleetd #422). Deliberately worded apart from {@link
* #enforceNotQuarantined} and {@link #enforceNotCoolingOff}: those two report a BACKEND-reported
* outage (exhaustion, repeated errors); this one reports an OPERATOR decision, so the message
* says "turned off" and names the model, never "quarantined" or "cooling off". A FOURTH,
* independent reason to refuse a spawn — never layered onto {@code BackendQuarantine} or {@code
* BackendOutagePolicy}, which would misattribute an operator's own choice to the backend.
*
* <p>Reads {@link #models0()} fresh on every call — the same liveness {@link #profiles0()}
* already has — so flipping {@code enabled: false} and reloading takes effect on the very next
* spawn, no restart (criterion 4). A profile naming no model, or a model absent from {@code
* models.allow:} entirely (nothing to gate against), is never refused here.
*
* @throws PlacementException naming the model and the profile, distinct from quarantine/cool-off
*/
private void enforceModelEnabled(String profile) {
FleetConfig.Profile cfg = profiles0().get(profile);
String model = (cfg == null) ? null : cfg.model();
if (model == null) {
return;
}
if (models0().offIds().contains(model)) {
throw new PlacementException("worker profile '" + profile + "' names model '" + model
+ "', which the operator has turned off in models.allow — refusing spawn");
}
}
/** The subset of {@code candidates} whose {@code model:} is currently turned off (fleetd #422). */
private Set<String> modelOffProfiles(List<PlacementCandidate> candidates) {
Set<String> off = models0().offIds();
if (off.isEmpty()) {
return Set.of();
}
return candidates.stream()
.map(PlacementCandidate::profile)
.filter(p -> {
FleetConfig.Profile cfg = profiles0().get(p);
return cfg != null && cfg.model() != null && off.contains(cfg.model());
})
.collect(Collectors.toSet());
}
/**
* The profile names {@code role} may be placed on, in definition order.
*
@@ -506,8 +604,23 @@ public final class CompositePeerLauncher implements PeerLauncher {
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
}
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
private String defaultProfileFor(MemberRole role) {
/**
* {@inheritDoc}
*
* <p>Live: reads {@link #poolFor}, which reads {@link #profileConfigs} and {@link #fleet} fresh
* on every call, so a config reload is visible without a restart (fleetd #425) — unlike {@link
* #defaultProfile}, the field captured once at construction, which this falls back to only when
* {@link #poolFor} has nothing to offer at all (no profiles configured for this composite).
*
* <p>Exact only under the {@code fixed} placement policy — the one that reads this value
* ({@code FixedPlacementPolicy}, package-private, hence not linked) as its first, preferred
* candidate. {@code weighted}/{@code round-robin} placement can choose a different candidate
* from {@code role}'s pool even on the very first spawn; this method does not simulate that
* choice, matching what the {@code defaultProfile:}-derived reporting this replaces has always
* done.
*/
@Override
public String defaultProfileFor(MemberRole role) {
List<String> pool = poolFor(role);
return pool.isEmpty() ? defaultProfile : pool.getFirst();
}
@@ -524,6 +637,142 @@ public final class CompositePeerLauncher implements PeerLauncher {
return out;
}
/**
* Build the {@link PlacementContext} an unqualified spawn of {@code role} would be judged
* against right now — the single source both {@link #spawn} and {@link #place} read, so the two
* can never disagree about which conditions (quarantine, cool-off, model-off) apply to which
* candidate (fleetd #425 rework: round 1 duplicated this into a second, blind resolver —
* {@link #defaultProfileFor} — which is why it regressed; round 2 found that even a single
* shared resolver is not enough on its own if the CALLER re-resolves through an explicit
* profile afterwards — see {@link PlacementDecision}).
*
* @param unreachable the caller's mutable unreachable set; {@link #spawn} grows this across
* retries and rebuilds the context from it, {@link #place} passes a fresh
* empty one since it never retries
*/
private PlacementContext placementContextFor(MemberRole role, Set<String> unreachable) {
List<PlacementCandidate> candidates = candidates(role);
String roleDefault = defaultProfileFor(role);
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
Set<String> coolingOff = coolingOffProfiles(candidates);
// fleetd #422: read live per spawn, same as quarantined/coolingOff above — a config reload
// that flips a model's enabled state is visible to the very next unqualified spawn.
Set<String> modelOff = modelOffProfiles(candidates);
return new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff, modelOff);
}
/**
* {@inheritDoc}
*
* <p>fleetd #425 rework, round 2: runs the exact same selection {@link #spawn} uses for a
* blank-profile request — {@link #placementContextFor} plus one {@link PlacementPolicy#select}
* — rather than {@link #defaultProfileFor}'s blind "pool's first entry", so a quarantined,
* cooling-off, or model-off pool-first candidate is routed around here exactly as it would be
* by a real spawn. Unlike {@link #spawn}, this never retries on {@link
* PeerUnreachableException}: there is no spawn attempt to fail, so "unreachable" never grows
* past the empty set it starts with, and a single {@link PlacementPolicy#select} call already
* reflects the live quarantine/cool-off/model-off state.
*
* <p>Deliberately does <em>not</em> apply {@link #enforceMaxLoad} (or any of the other three
* {@code enforce*} checks): those belong to {@link #spawn}'s EXPLICIT-profile branch, the
* operator-override path, and this method answers a different question — "where would an
* UNQUALIFIED spawn land". That is not the same as {@code select} ignoring these conditions —
* every condition {@code select} filters on (quarantine, cooling off, {@code maxLoad} under
* every placement policy including the default {@code fixed}, since fleetd #435, model-off,
* unreachable, weight-0) is already reflected in the {@link PlacementDecision} this method
* returns, because {@code select} walked past every excluded candidate to find it. What this
* method's caller must not do is take that resolved name and hand it back to {@link
* #spawn(SpawnRequest)} as an explicit profile: the explicit-profile branch treats the same
* exclusion conditions as a reason to REFUSE, where {@code select} had already treated them as
* a reason to fall through — round 1 of this fix did exactly that, turning a fall-through this
* method had already resolved around into a refusal one call later. Round 2 fixes that at the
* caller: {@link #spawn(SpawnRequest, PlacementDecision)} carries this exact decision to the
* spawn without re-resolving or re-checking it, through the same routing path {@code select}
* itself was consulted from.
*
* @throws PlacementException if no candidate in {@code role}'s pool is currently placeable
* (mirrors what an actual unqualified spawn would throw)
*/
@Override
public PlacementDecision place(MemberRole role) {
PlacementContext ctx = placementContextFor(role, new HashSet<>());
return new PlacementDecision(placementPolicy.get().select(ctx).profile());
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #place}, so the two can never disagree about the answer for the same
* {@code role} at the same instant — kept as a convenience for a caller that only wants the
* resolved name (a status report, a log line), never for a caller that will act on it by
* spawning: that caller must hold the {@link PlacementDecision} itself and pass it to {@link
* #spawn(SpawnRequest, PlacementDecision)} — see {@link PlacementDecision}'s javadoc for why
* resolving here and spawning separately, with the name fed back in as an explicit profile,
* regressed fleetd #425 twice.
*/
@Override
public String routedProfileFor(MemberRole role) {
return place(role).profile();
}
/**
* {@inheritDoc}
*
* <p>Routes {@code decision.profile()} directly to its owning delegate — the identical
* {@code d.spawn(routedReq)} call {@link #spawn(SpawnRequest)}'s blank-profile branch makes for
* its first pick — WITHOUT re-running {@link #enforceNotQuarantined}, {@link
* #enforceNotCoolingOff}, {@link #enforceMaxLoad}, or {@link #enforceModelEnabled}: those are
* the EXPLICIT-profile branch's checks, and {@code decision} did not come from an operator
* naming a profile — it came from {@link #place}, which already applied whichever of these
* conditions {@link PlacementPolicy#select} actually filters on (fleetd #425 rework, round 2).
*
* <p>The two branches disagree on purpose about what an excluded profile means, and that
* disagreement is not what this method removes. The blank-profile routing branch (and
* {@link #place}) treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0 profile
* as a reason to fall through to the next candidate; the EXPLICIT-profile branch treats naming
* that same profile as a reason to refuse outright — someone who names a profile should get a
* refusal, not a silent substitution onto a different backend. That is still correct after
* fleetd #435. What round 1 got wrong, and what this method exists to stop happening again, is
* turning a fall-through into a refusal by accident: resolving a name via {@link #place} and
* then handing that same name back to {@link #spawn(SpawnRequest)} as an explicit profile takes
* the refusing branch on a decision the routing branch had already approved by falling through
* past everything else.
*
* <p>Before fleetd #435, this exact accident was reachable through {@code maxLoad} specifically:
* {@code FixedPlacementPolicy} — the default policy — did not evaluate {@code maxLoad} at all
* for automatic selection, so {@link #place} could approve an at-cap profile that {@link
* #enforceMaxLoad} would then refuse one call later. fleetd #435 closed that: {@code
* FixedPlacementPolicy} now walks past an at-cap candidate exactly like {@code weighted}/
* {@code round-robin} already did, so {@link #place} can no longer return one, and this specific
* failure — an approved placement dying at {@code enforceMaxLoad} — cannot happen any more.
* What this method still buys, now that {@code maxLoad} can no longer cause it: it never
* re-evaluates a condition {@link #place} already decided, and it closes the window between
* that decision and the spawn in which the underlying state (another spawn landing on the same
* profile, a config reload) could otherwise move and make a stale explicit re-check wrong.
*
* <p>Deliberately does not retry on {@link PeerUnreachableException} across candidates the way
* {@link #spawn(SpawnRequest)}'s blank-profile branch does: retrying here would silently
* re-place the caller onto a different profile than the one {@code decision} named, behind the
* back of a caller that may already have provisioned something (a worktree's {@code repoRoot},
* parity overlay) specifically for that name. A caller that wants the composite's own failover
* should call {@link #spawn(SpawnRequest)} with a blank profile directly, not resolve through
* {@link #place} first. Losing that retry on a resolve-then-spawn path is an accepted, unrelated
* cost — see {@code SessionManager.acquireWithWorktree}'s own comment on it — never widened by
* this round to include {@code maxLoad}, which is what round 1 actually lost.
*/
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
HerdrPeerLauncher d = route(decision.profile());
SpawnRequest routedReq = req.withProfile(decision.profile());
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
@@ -534,6 +783,32 @@ public final class CompositePeerLauncher implements PeerLauncher {
return route(profileName).parityOverlay(profileName);
}
/**
* fleetd #422 follow-up: the single live read that answers both "is the models.allow: gate
* armed" and "which models are off", off the exact same accessor ({@link #models0()}) {@link
* #enforceModelEnabled} and {@link #modelOffProfiles} read — so {@code fleet_profiles}/{@code
* GET /profiles} (via {@link PeerLauncher#disabledModels()}, which now delegates here) can
* never report a different answer than the gate enforces (the fleetd #404 lesson), and the
* startup log line built from this can never disagree with either.
*
* <p>{@link #models0()} itself normalizes a {@code null} {@link #models} read to the shared
* {@link #NO_MODELS_CONFIGURED} sentinel — deliberately the one object no config-supplied
* {@code Models} instance can ever be identical to, since it is private to this class — so
* comparing by reference here recovers exactly the fact {@code models0()}'s normalization
* would otherwise erase: whether the live source was {@code null} (no {@code models:} block,
* armed = false) or a real, config-supplied block (armed = true, even one whose {@code allow:}
* is itself empty or absent — {@link FleetConfig.Models}'s "absent or empty allow: is off"
* wording governs config-load validation, a distinct question from whether this gate is armed
* for reporting).
*/
@Override
public PeerLauncher.ModelGateState modelGateState() {
FleetConfig.Models m = models0();
return m == NO_MODELS_CONFIGURED
? PeerLauncher.ModelGateState.notConfigured()
: PeerLauncher.ModelGateState.armed(m.offIds());
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.get(id);
@@ -645,9 +920,21 @@ public final class CompositePeerLauncher implements PeerLauncher {
return byProfile.keySet();
}
/**
* {@inheritDoc}
*
* <p>fleetd #425: reports the <em>live</em> {@code dev} pool's first entry — the same value
* {@link #defaultProfileFor} computes for {@link MemberRole#DEV} — not the {@link
* #defaultProfile} field captured at construction. An unqualified {@code fleet_spawn} defaults
* to {@code MemberRole#DEV} (see {@link dev.ltms.fleet.peer.SpawnRequest}), so "the dev pool's
* live first entry" is exactly the profile such a spawn actually lands on right now — the
* question {@code fleet_profiles}' {@code "default"} field exists to answer. The frozen field is
* a role-agnostic fallback used only when {@link #poolFor} has nothing to report at all (no
* profiles configured), which {@link #defaultProfileFor} already handles.
*/
@Override
public String defaultProfile() {
return defaultProfile;
return defaultProfileFor(MemberRole.DEV);
}
/**
@@ -13,6 +13,7 @@ import java.nio.file.attribute.PosixFilePermissions;
import java.time.Duration;
import java.time.Instant;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Set;
import java.util.stream.Stream;
@@ -28,33 +29,85 @@ import java.util.stream.Stream;
* control lacked: herdr applies that overlay BEFORE the shell starts, so any sourced file can undo
* it — and did.
*
* <p><b>Which file is last depends on the platform, so the scrub runs from two of them.</b> zsh
* <p><b>zsh reads its four startup files under three different conditions, so no single file is
* guaranteed to run — the scrub has to cover the gap between them, not just the platforms.</b> zsh
* reads {@code .zshenv} always, {@code .zprofile} and {@code .zlogin} only for a LOGIN shell, and
* {@code .zshrc} only for an INTERACTIVE one. herdr does not open the same kind of shell
* everywhere — measured on herdr 0.8.0: macOS panes run {@code -zsh} (login, so {@code .zlogin}
* runs), Linux panes run a plain {@code /usr/bin/zsh} (interactive but NOT login, so
* {@code .zlogin} never runs at all). A scrub in {@code .zlogin} alone is therefore a control that
* silently does nothing on Linux — the exact failure this class exists to remove, one platform
* over.
* {@code .zshrc} only for an INTERACTIVE one. A pane shell that is at least one of login or
* interactive is covered by sourcing the scrub from {@code .zshrc} and {@code .zlogin} (below), but
* a pane shell that is NEITHER reads only {@code .zshenv} and stops — fleetd #388, measured: a herdr
* pane can be neither login nor interactive, and such a pane read {@code .zshenv}, never reached
* {@code scrub.zsh}, and left no report at all. A bare {@code argv[0]} of {@code /usr/bin/zsh}
* proves the shell is NOT a login shell; it says nothing about whether it is interactive, so it
* must never be read as "therefore interactive" — that wrong inference is what let #388 ship.
*
* <p>So both {@code .zshrc} and {@code .zlogin} source the same generated {@code scrub.zsh} after
* sourcing their {@code $HOME} counterpart. On Linux only the first fires; on macOS both do, and
* the second pass is deliberate rather than merely harmless — it re-scrubs anything the operator's
* own {@code ~/.zlogin} exported after {@code .zshrc} had finished. Re-running is idempotent: a
* name already blank is blanked again, and the report is rewritten with the same counts.
* <p>So {@code .zshenv} carries a THIRD pass, guarded by the exact condition that defines the gap:
* {@code [[ ! -o login && ! -o interactive ]]}. That guard is why this pass cannot double-scrub a
* pane that {@code .zshrc} or {@code .zlogin} will also cover — one of {@code -o login}/
* {@code -o interactive} is always true there, so the {@code .zshenv} pass never fires for them, and
* their own unconditional sourcing is untouched. The guard also carries a sentinel
* ({@value #SCRUB_SENTINEL}) so it fires once per PANE and not once per PROCESS: {@code .zshenv} is
* read by every zsh a member's own tooling forks (a plain {@code zsh -c '...'} for a single
* command is itself neither login nor interactive), and those children inherit variables their
* parent deliberately set for them (git hooks get {@code GIT_DIR}, a venv gets
* {@code VIRTUAL_ENV}, a build tool gets {@code NODE_OPTIONS} or {@code JAVA_TOOL_OPTIONS}).
* Re-scrubbing every such child would blank all of that, and would also make the pane's own
* {@code scrub-report.txt} — rewritten on every pass — describe whichever child exited last
* instead of the pane. The sentinel is exported only AFTER {@code scrub.zsh} runs, so the pass
* that sets it never sees it and cannot blank it; it must also be on the scrub's own allow-list
* (see {@link #generate(Path, Set)}) so a later pass, in the same pane, cannot blank it back to
* empty — an exported-but-empty sentinel reads as unset to the {@code -z} guard and would silently
* re-enable scrubbing for every subsequent child of that pane.
*
* <p>So all three of {@code .zshenv} (gap only, guarded), {@code .zshrc}, and {@code .zlogin}
* source the same generated {@code scrub.zsh} after sourcing their {@code $HOME} counterpart. A
* login-and-interactive pane runs the {@code .zshrc} and {@code .zlogin} passes, and the second is
* deliberate rather than merely harmless — it re-scrubs anything the operator's own
* {@code ~/.zlogin} exported after {@code .zshrc} had finished. A pane that is neither runs only the
* {@code .zshenv} pass. Re-running is idempotent: a name already blank is blanked again, and the
* report is rewritten with the same counts.
*
* <p>Each generated file sources its {@code $HOME} counterpart FIRST, so {@code PATH} and every
* toolchain binary still resolve exactly as the operator configured them; only afterwards does
* {@code .zlogin} run the scrub: every EXPORTED variable not on the derived allow-list is re-exported
* toolchain binary still resolve exactly as the operator configured them; only afterwards does the
* scrub run: every EXPORTED variable not on the derived allow-list is re-exported
* blank. Blank, not credential-shaped-pattern-filtered: a pattern list ({@code *TOKEN*}, …) is an
* enumeration and misses what it did not think of — a username is the other half of a credential and
* is shaped like none. Credential-SHAPED names among the blanked set go to the WARN log only,
* never to the control.
*
* <p>The scrub also writes {@code scrub-report.txt} into its own directory: one {@code allowed N of
* M} line (N = exports left untouched, M = exports present when the scrub ran), then the blanked
* NAMES — never values. The launcher reads this back at teardown and logs it, because a blocked
* count next to an unknown denominator is not a finding.
* M failed F} line (N = exports left untouched, M = exports present when the scrub ran, F = names
* the scrub attempted to blank but could not), then the NAMES — blanked ones bare, unblankable ones
* {@code !}-prefixed — never values. The launcher reads this back at teardown and logs it, because a
* blocked count next to an unknown denominator is not a finding.
*
* <p><b>fleetd #394:</b> plain {@code export "$n="} is a FATAL error for a zsh read-only or special
* parameter (for example {@code UID}) — it aborts the whole sourced file, so every name still to
* come is never blanked and the report above is never written at all. The blanking loop instead
* routes each attempt through {@code eval}, which contains that error to the single iteration: the
* loop always finishes, and a name that could not be blanked is counted as {@code failed} and
* listed {@code !}-prefixed rather than silently disappearing. This is deliberately not a skip-list
* of known-bad names — every enumerated name is still attempted, so a name nobody has thought of
* yet still gets tried and, if it fails, still gets counted.
*
* <p>The blanking loop also re-asserts, on its own, the same {@code [A-Za-z_][A-Za-z0-9_]*} shape
* check the enumeration loop already applied. Before {@code eval} was introduced a non-conforming
* name reaching {@code export "$n="} was harmless either way — the quoting made it inert. With
* {@code eval}, the name is spliced into a string and interpreted as shell syntax, so the enumeration
* loop's check is no longer sufficient on its own to keep that call site safe — it is a guard on a
* different loop, and the two must not silently drift apart. Re-checking right before the
* {@code eval} keeps that call site safe by its own reading, independent of whatever the enumeration
* loop does or stops doing in a later change.
*
* <p><b>fleetd #400:</b> {@code eval}'s exit status is not proof that the blank actually happened.
* zsh coerces a bare {@code NAME=} assignment on an integer special parameter (measured on macOS zsh
* 5.9: {@code SECONDS}, {@code RANDOM}, {@code SHLVL}, {@code HISTSIZE}, {@code COLUMNS},
* {@code LINES}, {@code USERNAME}) to a number instead of failing — {@code eval} returns success,
* the value is untouched, and a status-based classification reports it as blanked when it was not.
* The fix classifies on the observed effect instead: after the attempt, the name's value is read
* back with the {@code (P)} indirection flag and the decision is made from whether that is now
* empty. This one check covers all three shapes a name can take at this point — a genuine blank, a
* fatal read-only error {@code eval} merely contained, and this silent no-op — and the exit status
* plays no part in the decision at all.
*/
public final class EnvAllowListScrub {
@@ -63,7 +116,11 @@ public final class EnvAllowListScrub {
/** Name of the report file written into the generated directory by the scrub itself. */
static final String REPORT_FILE = "scrub-report.txt";
/** The scrub body, generated once and sourced from both {@code .zshrc} and {@code .zlogin}. */
/**
* The scrub body, generated once and sourced from {@code .zshrc} and {@code .zlogin}
* unconditionally, and from {@code .zshenv} when the pane shell is neither login nor
* interactive (fleetd #388) — see the class javadoc.
*/
static final String SCRUB_FILE = "scrub.zsh";
/** Prefix of every generated directory — also what {@link #reapOrphans} matches on. */
@@ -79,14 +136,42 @@ public final class EnvAllowListScrub {
private static final String SOURCE_SCRUB =
"source \"$ZDOTDIR/" + SCRUB_FILE + "\"\n";
/**
* fleetd #388: marks a pane, not a process, as already scrubbed. Set only by the guarded
* {@code .zshenv} pass (see {@link #NEITHER_LOGIN_NOR_INTERACTIVE_SCRUB}) after
* {@code scrub.zsh} has run, so it must also be folded into that pass's own allow-list — see
* the class javadoc's "must also be on the scrub's own allow-list" paragraph.
*/
static final String SCRUB_SENTINEL = "_CB633_SCRUBBED";
/**
* Appended to {@code .zshenv}, after its {@code $HOME} source: the third pass, guarded on the
* exact condition that defines the gap {@code .zshrc}/{@code .zlogin} do not cover — a shell
* that is neither login nor interactive. The sentinel export happens only once the scrub has
* already run, and only for as long as the current pane's environment has not been rebuilt from
* scratch (a fresh {@code env -i} child would not inherit it — that is out of scope here, since
* such a child is no longer running under the pane's own environment at all).
*/
private static final String NEITHER_LOGIN_NOR_INTERACTIVE_SCRUB =
"if [[ ! -o login && ! -o interactive && -z \"${" + SCRUB_SENTINEL + ":-}\" ]]; then\n"
+ " " + SOURCE_SCRUB
+ " export " + SCRUB_SENTINEL + "=1\n"
+ "fi\n";
private EnvAllowListScrub() {
}
/**
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran,
* how many were left untouched (allowed), and the NAMES that were blanked. Values never appear.
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran
* ({@code total}), how many were left untouched ({@code allowed}), how many the scrub attempted
* to blank but could not ({@code failed} — fleetd #394: a zsh read-only/special parameter such
* as {@code UID} fatally errors on plain {@code export NAME=}, so those attempts go through
* {@code eval} instead so the loop keeps going and the failure is counted rather than left
* invisible), and the NAMES in each of the latter two categories. {@code allowed +
* blanked.size() + unblankable.size() == total}, and {@code unblankable.size() == failed}.
* Values never appear.
*/
record ScrubReport(int allowed, int total, List<String> blanked) {
record ScrubReport(int allowed, int total, int failed, List<String> blanked, List<String> unblankable) {
}
/**
@@ -105,11 +190,16 @@ public final class EnvAllowListScrub {
reapOrphans(parentDir);
Path dir = Files.createTempDirectory(parentDir, DIR_PREFIX);
dir.toFile().deleteOnExit();
// The report is written by zsh, after these hooks are registered, so register its path
// too — otherwise the directory is non-empty at JVM exit and cannot be removed at all.
dir.resolve(REPORT_FILE).toFile().deleteOnExit();
write(dir, SCRUB_FILE, scrubScript(allowedNames));
write(dir, ".zshenv", homeSourcingFile(".zshenv"));
// zsh truncates this pre-created receipt after these hooks are registered. Register its
// path too — otherwise the directory is non-empty at JVM exit and cannot be removed.
Files.createFile(dir.resolve(REPORT_FILE)).toFile().deleteOnExit();
// fleetd #388: scrub.zsh's OWN allow-list must also keep SCRUB_SENTINEL, or a later
// pass in the same pane blanks it back to empty and the .zshenv guard below thinks it
// was never scrubbed — see the class javadoc.
Set<String> namesForScrubScript = new HashSet<>(allowedNames);
namesForScrubScript.add(SCRUB_SENTINEL);
write(dir, SCRUB_FILE, scrubScript(namesForScrubScript));
write(dir, ".zshenv", homeSourcingFile(".zshenv") + NEITHER_LOGIN_NOR_INTERACTIVE_SCRUB);
write(dir, ".zprofile", homeSourcingFile(".zprofile"));
write(dir, ".zshrc", homeSourcingFile(".zshrc") + SOURCE_SCRUB);
write(dir, ".zlogin", homeSourcingFile(".zlogin") + SOURCE_SCRUB);
@@ -147,12 +237,11 @@ public final class EnvAllowListScrub {
* into it: owner keeps full access, {@code group} gets traverse+read on the directory ({@code
* rwxr-x---}, so a member process — a login shell reading it via {@code ZDOTDIR}, or another
* process simply opening a file under it — running under that group can find and read the
* files) and read-only on each file ({@code rw-r-----}) — deliberately no group WRITE anywhere,
* since a member never needs to add or change what fleetd generated. (For the ZDOTDIR scrub
* specifically, this also means the scrub script's own report write inside the pane fails
* closed rather than open — see {@code scrub.zsh}'s trailing {@code 2>/dev/null} — which {@link
* dev.ltms.fleet.member.HerdrPeerLauncher#releaseZdotdir} already treats as "cannot be
* confirmed to have run" rather than success.)
* files) and read-only on each file ({@code rw-r-----}), except the pre-created ZDOTDIR
* {@code scrub-report.txt}. That receipt gets group write ({@code rw-rw----}), so
* {@code scrub.zsh} can truncate and write it without granting group write on the directory.
* If its optional permission change fails, the member cannot write a receipt and the launcher
* keeps its existing WARN rather than failing the spawn.
*
* <p>Package-private and named generically on purpose: fleetd #213 built this for the ZDOTDIR
* scrub directory, and fleetd #219 reuses it verbatim for {@link
@@ -167,7 +256,15 @@ public final class EnvAllowListScrub {
setGroupAndPermissions(dir, principal, "rwxr-x---");
try (Stream<Path> entries = Files.list(dir)) {
for (Path file : entries.toList()) {
setGroupAndPermissions(file, principal, "rw-r-----");
if (REPORT_FILE.equals(file.getFileName().toString())) {
try {
setGroupAndPermissions(file, principal, "rw-rw----");
} catch (IOException | UnsupportedOperationException ignored) {
// The receipt is optional. Its absence keeps the existing WARN path.
}
} else {
setGroupAndPermissions(file, principal, "rw-r-----");
}
}
}
} catch (IOException e) {
@@ -245,11 +342,13 @@ public final class EnvAllowListScrub {
}
return """
# generated by fleetd (CB-633 memberCredentials policy=allow-list) — do not edit.
# Sourced from .zshrc and again from .zlogin, each time AFTER that file has sourced
# its $HOME counterpart — so this runs after everything the operator sourced, on a
# login shell (macOS panes) and on a plain interactive one (Linux panes) alike.
# Running twice is idempotent and deliberate: the second pass catches anything
# ~/.zlogin exported after ~/.zshrc had finished.
# Sourced unconditionally from .zshrc and again from .zlogin, each time AFTER that
# file has sourced its $HOME counterpart — so this runs after everything the
# operator sourced, on any pane that is login and/or interactive. Running twice is
# idempotent and deliberate: the second pass catches anything ~/.zlogin exported
# after ~/.zshrc had finished. Also sourced, once, from a guarded pass in .zshenv
# (fleetd #388) when the pane shell is NEITHER login nor interactive — the one gap
# those two files do not cover.
typeset -A _cb633_allowed
for _cb633_n in %s; do _cb633_allowed[$_cb633_n]=1; done
@@ -271,15 +370,57 @@ public final class EnvAllowListScrub {
_cb633_blank+=("$_cb633_n")
done
{ for _cb633_n in "${_cb633_blank[@]}"; do export "$_cb633_n="; done; } 2>/dev/null
# fleetd #394: plain `export "$n="` is FATAL for a zsh read-only/special parameter
# (e.g. UID) and aborts this whole sourced file — every name still to come is never
# blanked, and the report below is never written, silently. `eval` contains that
# error to the single iteration instead: it still fails for that one name, but the
# loop continues and we can tell allowed / blanked / unblankable apart afterwards.
# This is not a skip-list of known-bad names (that would miss the next one nobody
# thought of) — every name in _cb633_blank is still attempted, unconditionally.
# Every name reaching this loop already passed the identical identifier check in the
# enumeration loop above — but that guard is 20 lines away in a different loop, and
# this line is about to splice the name into a string handed to `eval`. Before this
# fix the name only ever reached `export` quoted ("$n="), which is inert on a
# non-identifier string either way; `eval` makes THIS line the only thing standing
# between such a string and code execution in the member's pane, so it re-asserts the
# same check on its own rather than trusting a guard it does not own. Under normal
# operation this can never fire (the enumeration guard already filtered everything
# reaching _cb633_blank), so a name caught here is counted as unblankable rather than
# silently dropped — it is real evidence that the upstream guard was bypassed.
#
# fleetd #400: the attempt's own exit status is NOT proof of its effect. zsh coerces
# a bare `NAME=` assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
# HISTSIZE, COLUMNS, LINES, USERNAME on this host) to a number instead of failing —
# `eval` returns 0, the value is untouched, and the old exit-status check reported it
# as blanked when it was not. Classify on the observed effect instead: attempt the
# export, then read the name's value back with the `(P)` indirection flag and decide
# from whether it is now empty. One check then covers all three shapes a name can
# take here — a genuine blank, a fatal read-only error `eval` merely contained, and
# this silent no-op — without the exit status entering the decision at all.
typeset -a _cb633_ok _cb633_unblankable
_cb633_ok=()
_cb633_unblankable=()
for _cb633_n in "${_cb633_blank[@]}"; do
if [[ ! "$_cb633_n" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
_cb633_unblankable+=("$_cb633_n")
continue
fi
eval "export ${_cb633_n}=" 2>/dev/null
if [[ -z "${(P)_cb633_n}" ]]; then
_cb633_ok+=("$_cb633_n")
else
_cb633_unblankable+=("$_cb633_n")
fi
done
integer _cb633_kept=$(( _cb633_total - ${#_cb633_blank} ))
{
print -r -- "allowed $_cb633_kept of $_cb633_total"
for _cb633_n in "${_cb633_blank[@]}"; do print -r -- "$_cb633_n"; done
print -r -- "allowed $_cb633_kept of $_cb633_total failed ${#_cb633_unblankable}"
for _cb633_n in "${_cb633_ok[@]}"; do print -r -- "$_cb633_n"; done
for _cb633_n in "${_cb633_unblankable[@]}"; do print -r -- "!$_cb633_n"; done
} > "$ZDOTDIR/%s" 2>/dev/null
unset _cb633_allowed _cb633_names _cb633_blank _cb633_n _cb633_total _cb633_kept
unset _cb633_allowed _cb633_names _cb633_blank _cb633_ok _cb633_unblankable _cb633_n _cb633_total _cb633_kept
""".formatted(names, MemberEnvAllowList.zshCasePattern(), REPORT_FILE);
}
@@ -293,6 +434,11 @@ public final class EnvAllowListScrub {
* Read and parse {@link #REPORT_FILE} out of a generated ZDOTDIR directory. Returns {@code null}
* when absent or unreadable (the pane may have been torn down before its login shell ever got to
* the scrub) — callers treat that as "no measurement available", never as success.
*
* <p>First line is {@code "allowed <N> of <M> failed <F>"} (fleetd #394 added the trailing
* {@code failed <F>} — a count of names the scrub attempted to blank but could not, e.g. a zsh
* read-only/special parameter). Every following non-blank line is a name: a bare name was
* blanked, a {@code !}-prefixed name was attempted and failed. Values never appear on either.
*/
static ScrubReport readReport(Path zdotdir) {
Path report = zdotdir.resolve(REPORT_FILE);
@@ -305,17 +451,24 @@ public final class EnvAllowListScrub {
return null;
}
String[] parts = lines.getFirst().substring("allowed ".length()).trim().split("\\s+");
if (parts.length != 3 || !"of".equals(parts[1])) {
if (parts.length != 5 || !"of".equals(parts[1]) || !"failed".equals(parts[3])) {
return null;
}
List<String> blanked = new ArrayList<>();
List<String> unblankable = new ArrayList<>();
for (int i = 1; i < lines.size(); i++) {
if (!lines.get(i).isBlank()) {
blanked.add(lines.get(i));
String line = lines.get(i);
if (line.isBlank()) {
continue;
}
if (line.startsWith("!")) {
unblankable.add(line.substring(1));
} else {
blanked.add(line);
}
}
return new ScrubReport(Integer.parseInt(parts[0]), Integer.parseInt(parts[2]),
List.copyOf(blanked));
Integer.parseInt(parts[4]), List.copyOf(blanked), List.copyOf(unblankable));
} catch (IOException | NumberFormatException e) {
return null;
}
@@ -17,6 +17,7 @@ import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -594,6 +595,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
}
/**
* {@inheritDoc}
*
* <p>fleetd #450: re-enters {@link #spawn(SpawnRequest)} with {@code decision}'s profile named
* explicitly. This is the re-entering form the interface javadoc describes for a launcher with
* no placement concept of its own — an instance of this class spawns a single adapter's own
* profile set by explicit name only ({@link #place}/{@link #defaultProfileFor} are unoverridden
* here and just wrap {@link #defaultProfile()}); it does no quarantine/cool-off/maxLoad/model-off
* filtering of its own to re-apply. That filtering lives one layer up, in {@code
* CompositePeerLauncher}, which is the launcher that routes across more than one profile and
* therefore overrides this method with the routing form instead.
*/
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
return spawn(req.withProfile(decision.profile()));
}
/** The herdr daemon that owns this launcher's pane coordinates. */
public HerdrClient herdr() {
return agents.herdr();
@@ -937,14 +955,20 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
@Override
public void stop(String idOrPane) {
// Teardown knows only the pane, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
// Teardown knows only the pane, not which profile spawned it — and, when this call is
// routed here through CompositePeerLauncher's single-daemon stop() shortcut (fleetd #342),
// not even which adapter's config actually governed the spawn: the shortcut can hand the
// pane to a delegate that never spawned it, whose own profiles say nothing about how THIS
// pane was placed. So the decision to look for a tab to clean up is made from the pane's
// actual state, not from this delegate's static profile config: resolve the tab
// unconditionally and let {@link WorkspaceControl#locatePane} tolerate "not found" (it
// returns null rather than throwing); the single-occupant check below is what actually
// protects the user's shared tabs, exactly as it always has.
String paneId = paneByAgentId.remove(idOrPane);
if (paneId == null) {
paneId = idOrPane; // raw-pane fallback (reap, gate timeout, pane-addressed callers)
}
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
WorkspaceControl.PaneLocation loc = spaces.locatePane(paneId);
try {
agents.close(paneId);
} catch (HerdrException e) {
@@ -985,11 +1009,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(FleetConfig.Profile::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
@@ -1642,11 +1661,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
log.warn("memberCredentials allow-list: pane {} left no scrub report in {} — the "
+ "environment scrub cannot be confirmed to have run. Either the pane ended "
+ "before its shell finished starting, or its shell never read our generated "
+ "startup files, in which case that member saw the full host environment.",
+ "startup files. Either way, we cannot tell from here whether the scrub ran, "
+ "so we do not know what that member's environment contained.",
paneId, dir);
} else {
log.info("memberCredentials allow-list: pane {} allowed {} of {} environment variables",
paneId, report.allowed(), report.total());
if (report.failed() > 0) {
log.warn("memberCredentials allow-list: pane {} could not blank {} environment "
+ "variable(s) — {} (likely a zsh read-only/special parameter) — those "
+ "names were left in the member's environment. Confirm none of them is a "
+ "credential.",
paneId, report.failed(), report.unblankable());
}
List<String> shaped = report.blanked().stream()
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
.toList();
@@ -1665,30 +1692,59 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/**
* Guards the {@code effectiveAllowed == null} branch of {@link #logCredentialGap} — the
* genuinely-unprotected report (deny-by-default, and the allow-list non-zsh fallback) — to one
* WARN per launcher instance, not one per spawn.
* genuinely-unprotected report (deny-by-default, and the allow-list non-zsh fallback) — AND
* the {@code effectiveAllowed != null} / {@code keptByDerivedList} branch, the allow-list case
* where a name is in the gap but the derived allow-list keeps it anyway. Both branches log the
* same severity (WARN) about the same fact — a name genuinely reaching a member pane
* unprotected — so they share this one guard, keyed per NAME rather than per launcher instance:
* each credential-shaped name that is ever reported unprotected gets exactly one WARN, however
* many spawns see it and whichever of the two branches first reports it.
*
* <p>fleetd #341: {@code memberCredentials} is a live, re-read-per-spawn supplier, so the
* policy — and so the gap's actual member names — can change between two spawns on the same
* launcher. Before this fix the guard was a single {@code AtomicBoolean} tripped by either
* branch: spawn 1 could warn about name A and trip the flag, and a later spawn's gap containing
* a different name B would never be reported, even though B is just as unprotected as A was.
* {@code AtomicBoolean} could not express "once per distinct name" at all — only "once, ever,
* for whichever name got there first" — so this is a {@code Set<String>} guard instead, the
* same shape {@link OpenCodeLauncher#modelCheckSkippedWarned} already uses for its own
* once-per-distinct-thing WARN. {@link #add}'s return value (true only the first time a name is
* added) is what turns "log the whole gap" into "log only the names never warned about before".
*
* <p>Bounded by construction: every name added here first passed {@link
* #CREDENTIAL_SHAPED_NAME}'s filter over {@link #hostEnvNames}, i.e. it is an actual
* environment variable name from the daemon's own process — a small, OS-bounded set (the host
* environment has, in practice, tens to a few hundred entries), not an attacker- or
* request-controlled input. So this set cannot grow past "however many distinct credential-
* shaped names this host's environment has ever held across this launcher's lifetime," which is
* effectively fixed for the life of one daemon process — no separate cap is needed.
*
* <p>CB-633 follow-up (#192): kept SEPARATE from {@link #allowListGapLogged} on purpose.
* {@code memberCredentials} is a live, re-read-per-spawn supplier, so the policy can change
* between two spawns on the same launcher. A single shared flag would let a harmless allow-list
* INFO on spawn 1 permanently suppress the real deny-by-default WARN a later spawn deserves —
* the report that matters most getting hidden by the report that doesn't. Two flags mean each
* report kind fires exactly once, independent of what the other kind already logged.
* report kind (WARN vs. INFO) fires independently of what the other kind already logged; within
* the WARN kind itself, the set above further separates by name, for the same reason.
*/
private final AtomicBoolean unprotectedGapLogged = new AtomicBoolean();
private final Set<String> unprotectedGapNamesWarned = ConcurrentHashMap.newKeySet();
/**
* Guards the {@code effectiveAllowed != null} branch of {@link #logCredentialGap} — the
* allow-list-scrub-covered report — to one INFO per launcher instance. See {@link
* #unprotectedGapLogged}'s javadoc for why this is a separate flag rather than a shared one.
* #unprotectedGapNamesWarned}'s javadoc for why this is a separate flag rather than a shared
* one; unlike that guard it stays a per-instance {@code AtomicBoolean}, not a per-name set —
* fleetd #341 fixed the WARN-vs-WARN suppression, not this INFO's own one-shot shape, which was
* not reported as broken and is out of that ticket's scope.
*/
private final AtomicBoolean allowListGapLogged = new AtomicBoolean();
/**
* fleetd #185 stage 2: guards {@link #warnUnknownMemberEnvironment} to one WARN per launcher
* instance, not one per spawn — the same one-per-instance shape as {@link #unprotectedGapLogged}
* and {@link #allowListGapLogged}, kept as its own flag for the same reason those two are split:
* this mode is orthogonal to which of the other two branches would otherwise have fired.
* instance, not one per spawn — the same one-shot shape {@link #unprotectedGapNamesWarned} and
* {@link #allowListGapLogged} guard their own branches with, kept as its own flag for the same
* reason those two are split: this mode is orthogonal to which of the other two branches would
* otherwise have fired.
*/
private final AtomicBoolean unknownMemberEnvironmentWarned = new AtomicBoolean();
@@ -1816,14 +1872,21 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
List<String> blankedByScrub = gap.stream()
.filter(name -> !MemberEnvAllowList.keeps(effectiveAllowed, name))
.toList();
if (!keptByDerivedList.isEmpty() && unprotectedGapLogged.compareAndSet(false, true)) {
// fleetd #341: filter to names this guard has never warned about before — not just
// "isEmpty" on the whole branch — so a name this spawn's gap shares with an EARLIER
// spawn's (already-warned) gap does not re-print, while a name unique to THIS gap still
// does, whichever of the two WARN branches reported it first.
List<String> newlyUnprotected = keptByDerivedList.stream()
.filter(unprotectedGapNamesWarned::add)
.toList();
if (!newlyUnprotected.isEmpty()) {
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — the derived allow-list keeps them anyway (a profile's "
+ "gitTokenEnv/gitHostEnv/tokenEnv/env: names one, or this spawn injects it), "
+ "so every member pane inherits them UNBLOCKED — {}. Add each to "
+ "memberCredentials.known (or .allow if a member legitimately needs it), or "
+ "remove it from whatever profile setting derives it in.",
keptByDerivedList.size(), keptByDerivedList);
newlyUnprotected.size(), newlyUnprotected);
}
if (!blankedByScrub.isEmpty() && allowListGapLogged.compareAndSet(false, true)) {
log.info("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
@@ -1836,12 +1899,18 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/** The deny-by-default (and allow-list non-zsh fallback) WARN — unchanged byte-for-byte by #192. */
private void warnGapUnprotected(List<String> gap) {
if (unprotectedGapLogged.compareAndSet(false, true)) {
// fleetd #341: same "only the names never warned before" filter as the sibling branch in
// logCredentialGap above — see unprotectedGapNamesWarned's javadoc. Both branches share
// this one guard because both report the exact same fact (a name reaching a member pane
// unprotected) at the exact same severity; keying it by name is what lets a later spawn's
// DIFFERENT name still get its own WARN after an earlier spawn's already fired.
List<String> newlyUnprotected = gap.stream().filter(unprotectedGapNamesWarned::add).toList();
if (!newlyUnprotected.isEmpty()) {
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — every member pane inherits them UNBLOCKED — {}. "
+ "Add each to memberCredentials.known (blocked by default) or .allow "
+ "(if a member legitimately needs it).",
gap.size(), gap);
newlyUnprotected.size(), newlyUnprotected);
}
}
@@ -331,11 +331,22 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
+ "form — opencode's per-model context limit could not be applied for this profile",
cfg.profile(), cfg.model());
}
// fleetd #393: memberSkills seeding (GitWorktrees#seedSkills) copies skill folders into
// EVERY provisioned worktree's .claude/skills/ regardless of which kind ultimately spawns
// into it — that copy step cannot know the kind, only the caller of GitWorktrees#add does
// (see that method's own javadoc). .claude/skills/ is a Claude Code CLI convention the CLI
// discovers on its own; opencode has no such discovery, so without this, a seeded skill
// never reaches an opencode member even though GitWorktrees logged it as seeded. Read
// whatever landed under <cwd>/.claude/skills/ here — the one place in this launcher that
// knows both the kind (opencode, by construction: this IS OpenCodeLauncher) and the cwd.
List<Path> skillInstructionFiles = skillInstructionFiles(spec.cwd());
// A config file is needed for the bridge MCP mount, a member charter, the IDE MCP (+ its
// guidance overlay, CB-634), a pinned endpoint (CB-508), or a resolvable autoCompactWindow.
// guidance overlay, CB-634), a pinned endpoint (CB-508), a resolvable autoCompactWindow, or
// at least one seeded skill to deliver via instructions[] (fleetd #393).
if (cfg.hasMcp() || cfg.hasIdeMcp() || spec.charter() != null || hasCustomProvider(cfg)
|| wantsContextLimit) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter(), spec.cwd()).toString());
|| wantsContextLimit || !skillInstructionFiles.isEmpty()) {
workerEnv.put("OPENCODE_CONFIG",
writeConfig(cfg, spec.charter(), spec.cwd(), skillInstructionFiles).toString());
}
applyGitToken(workerEnv, cfg);
List<String> argv = argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId());
@@ -358,6 +369,65 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return withAgent;
}
/**
* fleetd #393: the {@code SKILL.md} paths under {@code <cwd>/.claude/skills/} this launcher can
* turn into {@code instructions[]} entries, plus the honest log this ticket asks for — emitted
* here, at the one point a skill's fate for THIS spawn is actually known, rather than trusting
* {@code GitWorktrees#seedSkills}'s kind-blind "N of M" line to mean "and it will be read."
*
* <p>Every non-hidden subdirectory of {@code .claude/skills/} is a candidate, whether it got
* there via {@code memberSkills:} seeding or because the target repo ships its own copy — this
* launcher does not care which; it only cares what it can find at spawn time. A candidate with
* a {@code SKILL.md} at its top level (the same shape {@link #writeConfig} already requires for
* the charter and IDE-rules instructions entries) is delivered; anything else is a directory
* this launcher cannot turn into a flat instructions entry, named explicitly in the log rather
* than silently dropped, so a caller sees a real "cannot consume" reason and not just a smaller
* number than {@code GitWorktrees}' own count.
*
* <p>No candidates at all (directory absent or empty) logs nothing — the same
* no-log-when-nothing-to-say shape {@link #hasCustomProvider} and friends already follow, and
* the shape {@code GitWorktrees#seedSkills} itself uses when {@code memberSkills:} is unset.
* A failure to even list the directory is logged and treated as "nothing delivered" — best
* effort, must never fail the spawn, matching {@code GitWorktrees#seedSkills}'s own contract.
*/
private List<Path> skillInstructionFiles(String cwd) {
if (cwd == null || cwd.isBlank()) {
return List.of();
}
Path skillsDir = Path.of(cwd, ".claude", "skills");
if (!Files.isDirectory(skillsDir)) {
return List.of();
}
List<Path> candidates;
try (var listing = Files.list(skillsDir)) {
candidates = listing.filter(Files::isDirectory)
.filter(p -> !p.getFileName().toString().startsWith("."))
.sorted()
.toList();
} catch (IOException e) {
log.warn("could not scan {} for skill folders to deliver to this opencode member: {}",
skillsDir, e.getMessage());
return List.of();
}
if (candidates.isEmpty()) {
return List.of();
}
List<Path> delivered = candidates.stream()
.map(dir -> dir.resolve("SKILL.md"))
.filter(Files::isRegularFile)
.toList();
List<String> undeliverable = candidates.stream()
.filter(dir -> !Files.isRegularFile(dir.resolve("SKILL.md")))
.map(dir -> dir.getFileName().toString())
.toList();
log.info("skill delivery: {} of {} skill folder(s) under {} reached this opencode member via "
+ "instructions[] (opencode does not read .claude/skills/ natively, unlike "
+ "Claude Code){}",
delivered.size(), candidates.size(), skillsDir,
undeliverable.isEmpty() ? "" : "; no SKILL.md, could not be delivered: " + undeliverable);
return delivered;
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
@@ -418,8 +488,15 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*
* @param skillInstructionFiles fleetd #393: absolute {@code SKILL.md} paths from
* {@link #skillInstructionFiles(String)}, appended to
* {@code instructions[]} so a {@code memberSkills:}-seeded skill
* reaches this opencode member the same way the charter and IDE
* rules already do.
*/
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd,
List<Path> skillInstructionFiles) {
try {
Path dir = Files.createTempDirectory(configParentDir(), "fleetd-opencode-");
dir.toFile().deleteOnExit();
@@ -447,7 +524,39 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Files.writeString(charter, charterText);
charter.toFile().deleteOnExit();
root.putArray("instructions").add(charter.toAbsolutePath().toString());
// fleetd #393 follow-up: withArray, not putArray. putArray REPLACES whatever node
// is already at "instructions"; withArray gets-or-creates. All three writers on
// this array now use withArray so that write order stops being load-bearing for
// whoever adds a fourth.
//
// Be precise about what this particular line is worth, because an earlier version
// of this comment was wrong and claimed too much. THIS call is the one place where
// the two idioms are equivalent, and no test can tell them apart: it runs first,
// against a still-empty root, so there is never an existing node for putArray to
// replace. Measured on the #393 merge: flipping this one call back to putArray
// leaves the whole suite green (1618 tests), and always will. The edit is a
// readability and future-proofing change with no test behind it, and that is not a
// gap anyone can close.
//
// The ordering hazard is real, just not here. It is the LATER writers that can
// destroy earlier entries. Measured on the same merge: flipping the skills writer
// below to putArray deletes this charter entry and fails
// OpenCodeLauncherTest.instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder;
// flipping the IDE-rules writer fails that test and
// instructionsArrayHoldsCharterThenIdeRulesInOrder. Deleting this line altogether
// fails instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt — the
// assertion that a role contract reaches an opencode member at all, which is the
// hole that predates #393 and is what actually let the mutation hide.
root.withArray("instructions").add(charter.toAbsolutePath().toString());
}
// fleetd #393: each seeded skill's SKILL.md, delivered as a plain instructions[] entry
// — the only mechanism opencode has for static guidance text. Unlike Claude Code's
// Skill tool, opencode cannot load one of these on demand by name; the content is just
// always part of the system prompt from spawn. That is a real difference in HOW the
// content reaches the member, not a reason to withhold it.
for (Path skillFile : skillInstructionFiles) {
root.withArray("instructions").add(skillFile.toAbsolutePath().toString());
}
if (cfg.hasMcp() || cfg.hasIdeMcp()) {
@@ -56,10 +56,13 @@ public final class FleetMetrics {
m.describe(REPLIES, "counter",
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
"CB-307 push-loop nudges to the primary (sent|exhausted). fleetd #365: \"sent\" means "
+ "the herdr paste-and-submit call succeeded, not that the pane read it — this "
+ "layer has no read-receipt concept.");
m.describe(HEARTBEAT_NUDGES, "counter",
"CB-551 idle-lead heartbeat nudges (delivered|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down.");
"CB-551 idle-lead heartbeat nudges (sent|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down. "
+ "fleetd #365: \"sent\" means the herdr call succeeded, not that the lead read it.");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
@@ -8,6 +8,7 @@ import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import com.rabbitmq.client.Return;
import com.rabbitmq.client.impl.DefaultExceptionHandler;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -22,6 +23,7 @@ import java.util.concurrent.ConcurrentSkipListMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicReference;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
@@ -93,8 +95,23 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private final Channel channel;
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
/**
* target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself.
*
* <p><strong>CB-318 tombstone.</strong> The value {@link #RELEASED} is a reserved sentinel: it
* marks a target whose {@link #release} has already run, so {@link #deliverCallback} can tell a
* delivery landing after release() apart from a fresh target it has never seen. See both methods'
* javadoc for why a plain {@code held.remove(target)} is not enough.
*/
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/**
* CB-318 sentinel stored in {@link #held} for a target whose {@link #release} has already run.
* Never mutated — every read site compares it by reference ({@code ==}) before touching it as a
* map, because it is a single object shared across every released target and calling a mutator on
* it would corrupt state for all of them.
*/
private static final LinkedHashMap<String, Held> RELEASED = new LinkedHashMap<>();
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
@@ -137,17 +154,22 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
/** As {@link #open(String)}, with an explicit consumer prefetch (CB-527: caps the held backlog per target). */
public static AmqpReplyInbox open(String uri, int prefetch) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("fleetd-reply-inbox"), prefetch);
return new AmqpReplyInbox(connectionFactory(uri).newConnection(AmqpConnectionFailureLogger.REPLY_INBOX), prefetch);
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
static ConnectionFactory connectionFactory(String uri) throws Exception {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
factory.setExceptionHandler(new AmqpConnectionFailureLogger(AmqpConnectionFailureLogger.REPLY_INBOX, log));
return factory;
}
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this(connection, DEFAULT_PREFETCH);
@@ -201,6 +223,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
String tag = channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
consumerTags.put(target, tag);
// CB-318: drop a stale RELEASED tombstone from a prior ownership of this same target
// string, so a delivery under this fresh consumer is held normally instead of being
// nacked forever by deliverCallback's RELEASED check. Safe to do here, still under
// channelLock: no delivery for the consumer tag just registered above can reach
// deliverCallback before this basicConsume call returns.
held.remove(target, RELEASED);
log.debug("AMQP inbox owns queue {} for target {}", queue, target);
} catch (IOException e) {
throw new IllegalStateException("cannot own queue " + queue, e);
@@ -218,7 +246,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
* neither. So before this method existed with a requeue step, it dropped {@link #held}'s entries
* for {@code target} while the broker still considered them outstanding: never acked, never
* nacked, never requeued, and no longer reachable by {@link #peek} — permanently invisible. This
* is unlike {@link #handleRecovery} and {@link #close()}, whose bare {@code held.clear()} is
* is unlike {@link RecoveryListener#handleRecovery(Recoverable)} and {@link #close()}, whose bare {@code held.clear()} is
* correct because each has already made the broker requeue (a real connection drop, or
* {@code channel.close()} respectively) before clearing local state.
*
@@ -243,6 +271,32 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
* later connection drop even though this release did not manage to requeue it immediately. A
* failed {@code basicCancel} still throws, unchanged from before this fix — that failure means
* the consumer may still be attached, so best-effort requeue is not attempted underneath it.
*
* <p><strong>CB-318: {@code held.remove(target)} alone leaves a second window open.</strong> The
* bullet above already explains why cancelling first does not save a tag from going stale — but
* that only accounts for a delivery landing before this method starts touching {@link #held}.
* {@code basicCancel} stops <em>new</em> dispatches; it does not flush one already handed to the
* consumer work pool. So a delivery can still land on that pool's thread and reach
* {@link #deliverCallback} at any point during, or after, this method's body — and a plain
* {@code held.remove(target)} does nothing to stop it: {@code deliverCallback}'s
* {@code computeIfAbsent} finds the key gone and happily creates a brand-new map under it, which
* this method — already past its {@code remove} — never looks at again. That entry then sits
* delivered-but-unacked on {@link #channel} until the whole inbox closes: never requeued, never
* redelivered, and {@link #peek} is never called again for a target nothing owns any more.
*
* <p>The fix is {@link #held}{@code .compute(target, ...)} instead of {@code remove}: it takes
* whatever was held (to nack, same as before) and, in the same atomic step, leaves the
* {@link #RELEASED} tombstone behind instead of an absent key. {@code computeIfAbsent} and
* {@code compute} calls for the same key are mutually exclusive in {@link ConcurrentHashMap} —
* whichever of this call and a concurrent {@code deliverCallback} runs first is fully visible to
* the other, with no gap between them. So a delivery that loses the race sees a real map here and
* gets nacked by the loop below, same as always; a delivery that wins the race (runs first) is
* itself nacked by that same loop, once it settles into {@code held}. A delivery that arrives once
* this method has stored {@link #RELEASED} finds it via {@code computeIfAbsent} and refuses itself
* — see {@link #deliverCallback}. Either way nothing is silently retained forever, satisfying the
* ticket's invariant against dropping a message. This closes the window rather than merely
* narrowing it — correctness does not depend on how much time elapses between the swap and this
* method returning.
*/
@Override
public void release(String target) {
@@ -255,8 +309,13 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
}
var perTarget = held.remove(target);
if (perTarget != null) {
AtomicReference<LinkedHashMap<String, Held>> previouslyHeld = new AtomicReference<>();
held.compute(target, (_, v) -> {
previouslyHeld.set(v);
return RELEASED;
});
var perTarget = previouslyHeld.get();
if (perTarget != null && perTarget != RELEASED) {
synchronized (perTarget) {
for (Held h : perTarget.values()) {
try {
@@ -326,7 +385,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public List<InboxMessage> peek(String target) {
var perTarget = held.get(target);
if (perTarget == null) {
if (perTarget == null || perTarget == RELEASED) {
return List.of();
}
synchronized (perTarget) {
@@ -335,17 +394,17 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
@Override
public void ack(String target, String msgId) {
public boolean ack(String target, String msgId) {
var perTarget = held.get(target);
if (perTarget == null) {
return;
if (perTarget == null || perTarget == RELEASED) {
return false;
}
Held h;
synchronized (perTarget) {
h = perTarget.remove(msgId);
}
if (h == null) {
return; // never held (or already acked) — no-op
return false; // never held (or already acked) — no-op
}
try {
synchronized (channelLock) {
@@ -359,6 +418,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
}
return true;
}
private DeliverCallback deliverCallback(String target) {
@@ -370,6 +430,20 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
String content = new String(delivery.getBody(), StandardCharsets.UTF_8);
var perTarget = held.computeIfAbsent(target, _ -> new LinkedHashMap<>());
if (perTarget == RELEASED) {
// CB-318: release() already ran for this target and left the RELEASED tombstone in
// held (see release()'s javadoc) — computeIfAbsent() is guaranteed to see it rather
// than recreate a fresh map, because ConcurrentHashMap serializes compute/
// computeIfAbsent calls for the same key against each other. Refuse the delivery
// instead of holding it somewhere release() will never look at again: requeue it, the
// same way release() nacks its own held entries, so a later owner (or a connection
// drop) can still recover it. This does not need channelLock across a broker round
// trip — basicNack, like the duplicate-ack case just below, does not wait for one.
synchronized (channelLock) {
channel.basicNack(tag, false, true);
}
return;
}
boolean duplicate;
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
@@ -504,3 +578,44 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
}
}
/**
* Keeps RabbitMQ's forgiving exception behaviour while adding the connection identity that its
* default logger drops. Package-private so both AMQP connections use the same two names.
*/
final class AmqpConnectionFailureLogger extends DefaultExceptionHandler {
static final String REPLY_INBOX = "fleetd-reply-inbox";
static final String LEAD_MAILBOX = "fleetd-lead-mailbox";
private final String connectionName;
private final Logger logger;
AmqpConnectionFailureLogger(String connectionName, Logger logger) {
this.connectionName = connectionName;
this.logger = logger;
}
String connectionName() {
return connectionName;
}
@Override
protected void log(String message, Throwable cause) {
if (isSocketClosedOrConnectionReset(cause)) {
logger.warn("AMQP connection {}: {} (Exception message: {})", connectionName, message, cause.getMessage());
} else {
logger.error("AMQP connection {}: {}", connectionName, message, cause);
}
}
private static boolean isSocketClosedOrConnectionReset(Throwable cause) {
// Deliberate copy of ForgivingExceptionHandler's private static helper; check it on amqp-client upgrades.
if (!(cause instanceof IOException)) {
return false;
}
return "Connection reset".equals(cause.getMessage())
|| "Socket closed".equals(cause.getMessage())
|| "Connection reset by peer".equals(cause.getMessage());
}
}
@@ -59,16 +59,17 @@ public final class InMemoryReplyInbox implements ReplyInbox {
}
@Override
public void ack(String target, String msgId) {
public boolean ack(String target, String msgId) {
if (!owned.contains(target)) {
return;
return false;
}
var perTarget = store.get(target);
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.remove(msgId);
}
if (perTarget == null) {
return false;
}
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
return perTarget.remove(msgId) != null;
}
}
}
@@ -4,8 +4,8 @@ import java.util.List;
/**
* The lead-to-lead message channel this daemon speaks, as its callers need it — one lead's own
* mailbox: publish to a peer's coord-id, look at what has arrived for me, and ack what I have
* delivered.
* mailbox: publish to a peer's coord-id, look at what has arrived for me, ack what I have
* delivered, and (fleetd #361) inspect any coord-id's mailbox from the outside without owning it.
*
* <p>Extracted from {@link LeadMailbox} purely as a seam. {@code LeadMailbox} is the one production
* implementation and owns a live AMQP connection, so a test that wanted to exercise the routing in
@@ -32,9 +32,99 @@ public interface LeadChannel {
/** Non-destructive FIFO snapshot of the messages held for this daemon's own coord-id. */
List<LeadMessage> peek();
/** Drop {@code msgId} from the held set and ack it on the broker. A no-op if it is not held. */
/**
* Drop {@code msgId} from the held set and ack it on the broker. A repeated ack that this
* connection already completed may be a no-op. Any other unknown msgId must throw rather than
* report an ack that did not reach the broker.
*/
void ack(String msgId);
/** This daemon's own lead coordination id — the mailbox it owns, and the {@code from} it sends as. */
String selfCoordId();
/**
* Whether a message sitting in {@link #peek}'s held set (fetched but not yet {@link #ack}ed) is
* still safe if this daemon crashes or restarts right now — the conclusion of two independent
* facts about how this channel owns its own queue: the queue was declared <em>durable</em>, and
* the consumer that filled {@code held} uses <em>manual ack</em>, so an unacked delivery is still
* owned by the broker rather than only in this process's memory. Both must hold for {@code true};
* an implementation must derive this from what it actually did when it declared and consumed its
* queue, never return a literal — fleetd #440 found {@code FleetMcp}'s {@code heldDurable} field
* doing exactly that, unable to ever report {@code false} even after the fact stopped being true.
*/
boolean heldDurable();
/**
* A non-destructive look at {@code coordId}'s mailbox — does it exist, how many messages are
* waiting on it, and how many consumers are attached — without owning, consuming, or otherwise
* changing it. {@code consumers == 0} on an existing mailbox is the observable form of "nobody
* is reading this right now": a publish to it will sit queued rather than reach a pane.
*
* <p><strong>Never throws</strong> — this is a best-effort fact-finding call, not an operation a
* caller must handle failing. But it must never turn "I could not check" into a false negative:
* {@link MailboxState#absent(String)} means the broker positively confirmed there is no such
* queue, and {@link MailboxState#unknown(String)} — a distinct value — means the look could not
* be completed at all (broker unreachable, timed out, connection closed). A caller that
* collapses those two into one, as fleetd #361 initially did, cannot tell "that peer is down"
* from "I could not check", and a reader of {@code pending}/{@code consumers} cannot tell a
* measured zero from a zero standing in for "not measured".
*
* <p><strong>Must never share fate with {@link #publish} or {@link #peek}/{@link #ack}.</strong>
* fleetd #361: in AMQP 0-9-1 a passive queue declare of a queue that does not exist closes the
* channel it was declared on with a 404. An implementation backed by a real broker connection
* must inspect on a channel it can afford to lose — never the channel {@link #publish} or the
* consume loop depends on — so that looking at a peer that happens to be down can never break
* this daemon's own send or receive path.
*/
MailboxState inspect(String coordId);
/**
* The result of {@link #inspect}. {@code presence} tells apart three states a caller must not
* conflate: a confirmed-existing mailbox ({@link Presence#EXISTS}, the only case where
* {@code pending}/{@code consumers} are measured facts), a confirmed-absent one
* ({@link Presence#ABSENT} — the broker positively said "no such queue"), and one this call
* simply could not determine ({@link Presence#UNKNOWN} — broker unreachable, timed out,
* connection closed). {@code pending}/{@code consumers} are always {@code 0} and meaningless
* outside {@link Presence#EXISTS}; a renderer must gate on {@link #exists()} (or {@code
* presence} directly), never present them as measured otherwise.
*
* @param coordId the coord-id inspected
* @param presence whether the mailbox is confirmed to exist, confirmed absent, or unknown
* @param pending messages ready for delivery but not yet in a consumer's hands (0 unless EXISTS)
* @param consumers how many consumers are attached (0 unless EXISTS)
*/
record MailboxState(String coordId, Presence presence, int pending, int consumers) {
/** Whether {@link #inspect} was able to reach a definite answer, of either kind. */
public enum Presence { EXISTS, ABSENT, UNKNOWN }
/** {@code true} only when the broker confirmed this exact queue is currently declared. */
public boolean exists() {
return presence == Presence.EXISTS;
}
/**
* {@code true} when {@link #inspect} reached a definite answer (exists or confirmed
* absent); {@code false} when it could not determine either way. A caller must never treat
* {@code !known()} the same as a confirmed absence — the mailbox may well exist.
*/
public boolean known() {
return presence != Presence.UNKNOWN;
}
/** The broker confirmed this queue exists, with these measured counts. */
public static MailboxState exists(String coordId, int pending, int consumers) {
return new MailboxState(coordId, Presence.EXISTS, pending, consumers);
}
/** The broker positively confirmed there is no such queue (e.g. a 404 on passive declare). */
public static MailboxState absent(String coordId) {
return new MailboxState(coordId, Presence.ABSENT, 0, 0);
}
/** The look could not be completed — broker unreachable, timed out, or connection closed. */
public static MailboxState unknown(String coordId) {
return new MailboxState(coordId, Presence.UNKNOWN, 0, 0);
}
}
}
@@ -5,6 +5,7 @@ import dev.ltms.fleet.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ScheduledExecutorService;
@@ -43,6 +44,9 @@ public final class LeadCoordLoop {
private static final Logger log = LoggerFactory.getLogger(LeadCoordLoop.class);
/** A bounded window is enough: redelivery can only follow a recent failed ack or recovery. */
private static final int RECENT_DELIVERY_LIMIT = 1_024;
/** How an arriving peer message is rendered into the lead's pane — the sender's coord-id, then its text. */
static final String DELIVERY_FORMAT = "[lead %s] %s";
@@ -51,6 +55,8 @@ public final class LeadCoordLoop {
private final Supplier<Map<String, String>> leads;
private final ScheduledExecutorService scheduler;
private final long intervalMs;
/** msgIds already written to the pane, so recovery redelivery is acked without another pane write. */
private final LinkedHashMap<String, Boolean> delivered = new LinkedHashMap<>();
private volatile boolean running;
@@ -117,6 +123,13 @@ public final class LeadCoordLoop {
if (held.isEmpty()) {
return;
}
LeadMessage msg = held.getFirst();
if (wasDelivered(msg.msgId())) {
// This lives here, rather than in LeadMailbox, because only this loop knows a pane write
// happened. The mailbox only knows broker delivery tags and must still redeliver after a crash.
ackDelivered(msg);
return;
}
String lead = resolveLocalLead();
if (lead == null) {
// Left unacked on purpose: the broker keeps holding it until a lead pane exists.
@@ -136,7 +149,6 @@ public final class LeadCoordLoop {
lead, status, held.size());
return;
}
LeadMessage msg = held.getFirst();
try {
agents.send(lead, DELIVERY_FORMAT.formatted(msg.from(), msg.content()));
} catch (RuntimeException e) {
@@ -145,6 +157,13 @@ public final class LeadCoordLoop {
msg.msgId(), msg.from(), lead, e.toString());
return;
}
rememberDelivered(msg.msgId());
if (ackDelivered(msg)) {
log.debug("lead coordination: delivered message {} from {} to lead {}", msg.msgId(), msg.from(), lead);
}
}
private boolean ackDelivered(LeadMessage msg) {
try {
channel.ack(msg.msgId());
} catch (RuntimeException e) {
@@ -152,9 +171,24 @@ public final class LeadCoordLoop {
// deliberate direction of this trade.
log.warn("lead coordination: delivered message {} but could not ack it: {}",
msg.msgId(), e.toString());
return;
return false;
}
return true;
}
private boolean wasDelivered(String msgId) {
synchronized (delivered) {
return delivered.containsKey(msgId);
}
}
private void rememberDelivered(String msgId) {
synchronized (delivered) {
delivered.put(msgId, Boolean.TRUE);
if (delivered.size() > RECENT_DELIVERY_LIMIT) {
delivered.remove(delivered.keySet().iterator().next());
}
}
log.debug("lead coordination: delivered message {} from {} to lead {}", msg.msgId(), msg.from(), lead);
}
/**
@@ -96,7 +96,14 @@ public final class LeadHeartbeatLoop {
this.metrics = metrics;
}
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
/**
* Count one nudge outcome when a registry is wired; a no-op in unit tests.
*
* <p>fleetd #365: the {@code "sent"} outcome (renamed from {@code "delivered"}) records only
* that {@link #injectNudge} — a one-way herdr {@code agent.prompt} paste-and-submit — returned
* without throwing, not that the lead's pane actually read or acted on the text. This layer has
* no read-receipt concept, so "sent" is the honest word for what this call can ever establish.
*/
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(FleetMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
@@ -245,7 +252,7 @@ public final class LeadHeartbeatLoop {
agents.send(leadTerminal, fleet.nudgeText());
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
leadTerminal, quietCount);
countNudge("delivered");
countNudge("sent");
} catch (RuntimeException e) {
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
countNudge("failed");
@@ -10,6 +10,7 @@ import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import com.rabbitmq.client.Return;
import com.rabbitmq.client.ShutdownSignalException;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -58,7 +59,8 @@ import java.util.concurrent.TimeoutException;
* <p><strong>Recovery.</strong> The connection is opened with automatic + topology recovery
* enabled, mirroring {@code AmqpReplyInbox}: on reconnect the broker hands out fresh delivery tags,
* so the held snapshot is cleared (dedup by {@code msgId} still prevents any double-queue on
* redelivery) and any publish still awaiting its confirm is failed rather than left to idle out
* redelivery). {@link LeadCoordLoop} separately deduplicates pane writes, since it alone knows
* which messages reached a lead. Any publish still awaiting its confirm is failed rather than left to idle out
* the confirm timeout against a sequence number that means nothing on the new channel.
*/
public final class LeadMailbox implements LeadChannel, AutoCloseable {
@@ -84,6 +86,16 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
private final Object channelLock = new Object();
/** msgId → held delivery, for this mailbox's own queue only (there is exactly one). */
private final LinkedHashMap<String, Held> held = new LinkedHashMap<>();
/**
* fleetd #440: the answer to {@link #heldDurable()}, set once by {@link #own()} from the exact
* booleans it passed to {@code queueDeclare}/{@code basicConsume} — never a separate literal that
* could drift from what those calls actually did.
*/
private boolean heldDurable;
/** Successful broker acks on this connection, retained only to make a repeated caller ack quiet. */
private final LinkedHashMap<String, Boolean> recentlyAcked = new LinkedHashMap<>();
/** Bounds {@link #recentlyAcked}: it is only an idempotency aid, never delivery state. */
private static final int RECENT_ACK_LIMIT = 1_024;
/**
* A dedicated channel for {@link #publish}, kept separate from {@link #channel} (consume + ack)
@@ -122,17 +134,22 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
/** As {@link #open(String, String)}, with an explicit consumer prefetch. */
public static LeadMailbox open(String uri, String selfCoordId, int prefetch) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares the queue and re-attaches the consumer.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new LeadMailbox(factory.newConnection("fleetd-lead-mailbox"), selfCoordId, prefetch);
return new LeadMailbox(connectionFactory(uri).newConnection(AmqpConnectionFailureLogger.LEAD_MAILBOX), selfCoordId, prefetch);
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP coordination broker at " + uri, e);
}
}
static ConnectionFactory connectionFactory(String uri) throws Exception {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
factory.setExceptionHandler(new AmqpConnectionFailureLogger(AmqpConnectionFailureLogger.LEAD_MAILBOX, log));
return factory;
}
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for tests). */
LeadMailbox(Connection connection, String selfCoordId) {
this(connection, selfCoordId, DEFAULT_PREFETCH);
@@ -156,16 +173,15 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue). Any publish
// repopulates it with valid tags (dedup by msgId still prevents any double-queue). LeadCoordLoop
// remembers successful pane writes separately, so that redelivery cannot write a pane twice. Any publish
// confirm still in flight when the connection dropped is equally stale — fail it now rather
// than let it silently ride out CONFIRM_TIMEOUT_MS.
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
synchronized (held) {
held.clear();
}
clearHeldForRecovery();
failPendingPublishesOnRecovery();
log.info("AMQP lead mailbox connection recovered; cleared held messages for fresh redelivery");
}
@@ -181,13 +197,22 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
/** Declare + consume this daemon's own {@code lead.<selfCoordId>.inbox}. Called once, at construction. */
private void own() throws IOException {
String queue = queueName(selfCoordId);
boolean durableQueue = true; // durable, non-exclusive, keep on idle
boolean autoAck = false; // manual ack
synchronized (channelLock) {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
channel.basicConsume(queue, false, deliverCallback(), _ -> { }); // autoAck=false: manual ack
channel.queueDeclare(queue, durableQueue, false, false, null);
channel.basicConsume(queue, autoAck, deliverCallback(), _ -> { });
}
// fleetd #440: held mail is durable only while both hold — a durable queue AND manual ack.
this.heldDurable = durableQueue && !autoAck;
log.debug("lead mailbox owns queue {} for coord-id {}", queue, selfCoordId);
}
@Override
public boolean heldDurable() {
return heldDurable;
}
/**
* Publish {@code msg} to {@code toCoordId}'s mailbox and block until the broker's publisher
* confirm for it lands. Does <em>not</em> imply owning or consuming {@code toCoordId}'s queue.
@@ -254,6 +279,89 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
}
}
/**
* fleetd #361: look at {@code coordId}'s mailbox on a fresh, immediately-closed throwaway
* channel — never {@link #channel} (consume/ack) or {@link #publishChannel} (publish). A
* passive queue declare of a queue that does not exist closes the channel it was declared on
* with a 404; using a disposable probe channel means that closure can never touch either
* long-lived channel this instance depends on for {@link #publish} or the consume loop.
*
* <p>Classifies failures rather than collapsing them, both measured against a real broker in
* {@code LeadMailboxTest} rather than assumed from the AMQP 0-9-1 spec text:
* <ul>
* <li>a genuine 404 — an {@link IOException} wrapping a {@link ShutdownSignalException} whose
* {@link AMQP.Channel.Close#getReplyCode()} is {@code 404} — reports
* {@link MailboxState#absent}; every other declare failure reports
* {@link MailboxState#unknown} instead of quietly becoming the same "absent" value;
* <li>{@code catch (RuntimeException e)} on both attempts matters as much as the checked
* catches: a connection that is already closed makes {@link Connection#createChannel()}
* throw {@link com.rabbitmq.client.AlreadyClosedException} (a {@link RuntimeException},
* not an {@link IOException}) — an {@code inspect} that only caught {@code IOException}
* would let that escape, breaking the "never throws" contract this method promises.
* </ul>
*
* <p><strong>Honesty about which catch is measured and which is defensive:</strong> the
* {@code createChannel()} catch above is exercised end-to-end against a real broker by
* {@code LeadMailboxTest.inspectReportsUnknownRatherThanThrowingWhenTheConnectionIsAlreadyClosed}.
* The second {@code catch (RuntimeException e)}, around the passive declare itself — for the
* narrower race where the connection drops <em>between</em> {@code createChannel()} succeeding
* and the declare landing — has no such test; reaching it needs a connection that dies at that
* exact instant, which is not a scenario this suite drives on purpose. It stays purely
* defensive: correct by the same reasoning as the first catch, but unproven the way the first
* one is proven.
*/
@Override
public MailboxState inspect(String coordId) {
String queue = queueName(coordId);
Channel probe;
try {
probe = connection.createChannel();
} catch (IOException | RuntimeException e) {
log.debug("lead mailbox inspect: cannot open a probe channel for {}: {}", coordId, e.toString());
return MailboxState.unknown(coordId);
}
try {
AMQP.Queue.DeclareOk declared = probe.queueDeclarePassive(queue);
return MailboxState.exists(coordId, declared.getMessageCount(), declared.getConsumerCount());
} catch (IOException e) {
// The broker (or the client library) has already closed `probe` for us either way; only
// a confirmed 404 means "no such queue" — anything else (a different declare failure) is
// "could not determine", never silently reported as the same value as a genuine absence.
return isMissingQueue(e) ? MailboxState.absent(coordId) : MailboxState.unknown(coordId);
} catch (RuntimeException e) {
// E.g. the connection dropped between createChannel() and the declare landing.
log.debug("lead mailbox inspect: declare failed unexpectedly for {}: {}", coordId, e.toString());
return MailboxState.unknown(coordId);
} finally {
try {
if (probe.isOpen()) {
probe.close();
}
} catch (Exception e) {
log.debug("lead mailbox inspect: probe channel close for {}: {}", coordId, e.toString());
}
}
}
/**
* {@code true} only for the specific shape a missing-queue passive declare actually produces —
* measured against a real broker, not assumed from the spec text (see {@code
* LeadMailboxTest.passiveDeclareOfAMissingQueueThrowsAnIOExceptionWrappingA404ShutdownSignal}):
* an {@link IOException} whose cause is a {@link ShutdownSignalException} carrying an
* {@link AMQP.Channel.Close} reason with {@code replyCode == 404}. Any other shape (a different
* reply code, a {@code ShutdownSignalException} cause whose reason is not a
* {@code Channel.Close}, or no cause at all) is a declare failure of some other kind and must
* not be read as "confirmed absent" — pinned hermetically, with no broker needed, by
* {@code LeadMailboxIsMissingQueueTest} for exactly those three false shapes. Package-private
* (not {@code private}) so that test can call it directly.
*/
static boolean isMissingQueue(IOException e) {
if (!(e.getCause() instanceof ShutdownSignalException sse)) {
return false;
}
return sse.getReason() instanceof AMQP.Channel.Close close && close.getReplyCode() == AMQP.NOT_FOUND;
}
/** Convenience: {@link #peek} the current snapshot, then {@link #ack} every message in it. */
public List<LeadMessage> drain() {
List<LeadMessage> snapshot = peek();
@@ -261,15 +369,26 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
return snapshot;
}
/** Remove the held message {@code msgId} and ack it on the broker. No-op if not held. */
/**
* Remove the held message {@code msgId} and ack it on the broker.
*
* <p>A repeated ack that this connection already completed is a no-op, tracked in the bounded
* {@link #recentlyAcked} set. Any other missing entry throws: recovery clears {@link #held} while
* the broker still owns the unacked delivery, and quiet success there would hide a required retry.
* The set is bounded because it only distinguishes a recent duplicate caller ack from an unknown
* delivery; it is not a substitute for broker state across a reconnect.
*/
@Override
public void ack(String msgId) {
Held h;
synchronized (held) {
h = held.remove(msgId);
if (h == null && recentlyAcked.containsKey(msgId)) {
return;
}
}
if (h == null) {
return; // never held (or already acked) — no-op
throw new IllegalStateException("cannot ack lead message " + msgId + ": it is not held");
}
try {
synchronized (channelLock) {
@@ -283,6 +402,19 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
}
throw new IllegalStateException("cannot ack lead message " + msgId, e);
}
synchronized (held) {
recentlyAcked.put(msgId, Boolean.TRUE);
if (recentlyAcked.size() > RECENT_ACK_LIMIT) {
recentlyAcked.remove(recentlyAcked.keySet().iterator().next());
}
}
}
/** Clear stale delivery tags after recovery; package-private so the recovery contract test drives this exact path. */
void clearHeldForRecovery() {
synchronized (held) {
held.clear();
}
}
private DeliverCallback deliverCallback() {
@@ -92,8 +92,31 @@ public final class MessageService {
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
TIMED_OUT_WORKING,
/** Timed out before delivery — the message is still queued for the worker. */
/**
* Timed out with no confirmed delivery, and the target saw nothing — the message will not
* arrive later, so a caller may resend. {@link #send} reaches this outcome through {@link
* Injector#cancel} reporting one of two routes: {@link Injector.Cancellation#CANCELLED}
* means the message was still queued and this call removed it; {@link
* Injector.Cancellation#NOT_DELIVERED} means nothing was ever sent — the queue was cleared
* because the target never became ready or was abandoned, or the injector's call to the
* target's terminal ({@link dev.ltms.fleet.herdr.AgentControl#send}) failed with a herdr
* error that this codebase already treats as a confirmed absence. A third route,
* {@link Injector.Cancellation#ATTEMPTED}, used to be folded into this same outcome
* (fleetd #571) — it no longer is; see {@link #TIMED_OUT_UNCONFIRMED}.
*/
TIMED_OUT_QUEUED,
/**
* Timed out with delivery unknown. {@link #send} reaches this outcome when {@link
* Injector#cancel} reports {@link Injector.Cancellation#ATTEMPTED} (fleetd #551): the call
* to the target's terminal ({@link dev.ltms.fleet.herdr.AgentControl#send}) was made, but
* this caller never observed whether it reached the pane. {@code agent.prompt} pastes
* <em>and submits</em> in one call, so the target may already hold a complete, submitted
* turn and be working on it right now — the same reality as {@link #TIMED_OUT_WORKING},
* just not confirmed. The message may or may not have arrived. Treat this as neither a
* confirmed delivery nor a confirmed absence: a caller that resends on this outcome risks a
* double delivery — the same brief typed into the pane twice (fleetd #571).
*/
TIMED_OUT_UNCONFIRMED,
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
@@ -138,6 +161,53 @@ public final class MessageService {
public record AskResult(AskOutcome outcome, String answer) {
}
/**
* How a worker's {@code fleet_reply} ({@link #reply(String, String)}) actually landed
* (fleetd #365) — the two doors that expose it, {@code fleet_reply} and {@code POST
* /sessions/{id}/reply}, both used to report the single word "delivered" whichever of these
* happened, so a caller could not tell an active handoff from a reply merely held for later
* drain. Both are successes; they are not the same fact.
*/
public enum ReplyOutcome {
/** Resolved a {@code fleet_send}/{@code fleet_ask} that was actively waiting on this reply. */
RESOLVED_SEND("resolved_send", true,
"delivered — resolved the fleet_send that was waiting for it"),
/**
* No live waiter was open, but the reply completed a parked async ticket directly
* ({@link #askAnsweredAsyncTasks}) — a {@code fleet_poll} caller sees it immediately.
*/
RESOLVED_ASYNC_TICKET("resolved_async_ticket", true,
"delivered — resolved a pending async ticket (visible to fleet_poll)"),
/** Nothing was waiting; the reply was queued in the inbox for a later drain (CB-307). */
QUEUED("queued", false,
"queued — no send or ticket was waiting; held in the inbox for a later drain");
private final String wireName;
private final boolean delivered;
private final String description;
ReplyOutcome(String wireName, boolean delivered, String description) {
this.wireName = wireName;
this.delivered = delivered;
this.description = description;
}
/** Stable machine-readable name for a JSON/metrics label (REST's {@code outcome} field). */
public String wireName() {
return wireName;
}
/** Whether something was actively waiting and received this reply right now. */
public boolean delivered() {
return delivered;
}
/** Shared human-readable text — the one place both {@code fleet_reply} and REST word this. */
public String description() {
return description;
}
}
/** Lifecycle phase of an async delegation ticket. */
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
@@ -252,12 +322,21 @@ public final class MessageService {
*/
private final ConcurrentHashMap<String, Boolean> strandedReplies = new ConcurrentHashMap<>();
/**
* Targets whose last send timed out with {@link Outcome#TIMED_OUT_QUEUED} (CB-640) — the
* message never reached the {@link Injector} delivery window before the caller's deadline, so
* it is still sitting in the injector's own per-target queue. Set where {@link #send} already
* computes {@code wasDelivered} for that outcome; no new queue is kept here, only the fact.
* Cleared the same way as {@link #strandedReplies}: the next accepted delivery for the target
* ({@link #send} opening a fresh waiter) or a teardown ({@link #abandon}).
* Targets whose last send timed out with no confirmed delivery (CB-640) — {@link #send} called
* {@link Injector#cancel} and got back something other than {@code DELIVERED}. That covers
* three histories, not one: {@link Injector.Cancellation#CANCELLED} — the message was still
* queued and {@code cancel} removed it right there; {@link Injector.Cancellation#NOT_DELIVERED}
* — nothing was ever sent, because the target never became ready, was torn down, or the call to
* its terminal failed with a herdr error this codebase already treats as a confirmed absence; or
* {@link Injector.Cancellation#ATTEMPTED} (fleetd #551) — the call to the target's terminal was
* made and its outcome is unknown, so the target may already hold a complete, submitted turn.
* Only the first two mean the message will not arrive later and the target saw nothing; on the
* third it may already have arrived in full — and the caller sees a different outcome for it
* ({@link Outcome#TIMED_OUT_UNCONFIRMED}, fleetd #571) than for the first two ({@link
* Outcome#TIMED_OUT_QUEUED}). Set where {@link #send} already computes {@code wasDelivered} for
* that outcome; no queue is kept here, only the fact that the send ended with no confirmed
* delivery. Cleared the same way as {@link #strandedReplies}: the next accepted delivery for the
* target ({@link #send} opening a fresh waiter) or a teardown ({@link #abandon}).
*/
private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
@@ -344,12 +423,22 @@ public final class MessageService {
/**
* Read-only delegation fact for fleet views (CB-640): {@code target}'s last send timed out
* before the {@link Injector} ever delivered it — the caller saw
* {@link Outcome#TIMED_OUT_QUEUED} (see the {@code TimeoutException} branch of {@link #send}),
* and the message is still sitting in the injector's per-target queue waiting for the worker
* to go idle. Distinct from {@link Outcome#TIMED_OUT_WORKING}, where delivery already happened
* and only the reply is outstanding. Cleared the next time this target's delivery is accepted
* or the target is abandoned — see {@link #queuedDeliveries}.
* with no confirmed delivery — the caller saw {@link Outcome#TIMED_OUT_QUEUED} or {@link
* Outcome#TIMED_OUT_UNCONFIRMED} (fleetd #571; see the {@code TimeoutException} branch of
* {@link #send}). Despite the method's name, this is not proof that a message is sitting in a
* queue: {@link Injector#cancel} reports this outcome through three routes. {@link
* Injector.Cancellation#CANCELLED} means the message was still queued and got removed right
* there. {@link Injector.Cancellation#NOT_DELIVERED} means nothing was ever sent — the target
* never became ready, was torn down, or the call to its terminal failed with a herdr error this
* codebase already treats as a confirmed absence. Only these two routes mean the message will
* not arrive later, and both report {@code TIMED_OUT_QUEUED}. {@link
* Injector.Cancellation#ATTEMPTED} (fleetd #551) means the call to the target's terminal was
* made and its outcome is unknown: {@code agent.prompt} pastes <em>and submits</em> in one
* call, so on this route the target may already hold a complete, submitted turn and be
* working on it right now — it does NOT follow that the target saw nothing, and this route
* reports {@code TIMED_OUT_UNCONFIRMED} instead. Distinct from {@link Outcome#TIMED_OUT_WORKING},
* where delivery already happened and only the reply is outstanding. Cleared the next time this
* target's delivery is accepted or the target is abandoned — see {@link #queuedDeliveries}.
*/
public boolean hasQueuedDelivery(String target) {
return target != null && queuedDeliveries.containsKey(target);
@@ -369,15 +458,15 @@ public final class MessageService {
/**
* Read-only delegation fact for fleet views (CB-640): an async ticket is still
* {@link Phase#PENDING} against {@code target}, yet nothing is actually in flight for it — no
* open rendezvous waiter ({@link #hasAcceptedDelivery}) and no message still sitting in the
* injector's queue ({@link #hasQueuedDelivery}). A healthy PENDING ticket can briefly look this
* way while its virtual thread has not yet been scheduled or is blocked on the session lock
* behind another send to the same target, so this is a snapshot fact for the health classifier
* to weigh across ticks, not proof on its own that the ticket is stuck. It also genuinely
* persists — not just as a passing race — once an async {@code fleet_ask} lapses unanswered:
* {@link #ask} clears the ticket's question and returns it to {@code PENDING}, but {@link #send}
* already closed the forward waiter the instant the question surfaced, so the target has
* neither an accepted nor a queued delivery left to show for it.
* open rendezvous waiter ({@link #hasAcceptedDelivery}) and no record of a send that ended with
* no confirmed delivery ({@link #hasQueuedDelivery}). A healthy PENDING ticket can briefly look
* this way while its virtual thread has not yet been scheduled or is blocked on the session
* lock behind another send to the same target, so this is a snapshot fact for the health
* classifier to weigh across ticks, not proof on its own that the ticket is stuck. It also
* genuinely persists — not just as a passing race — once an async {@code fleet_ask} lapses
* unanswered: {@link #ask} clears the ticket's question and returns it to {@code PENDING}, but
* {@link #send} already closed the forward waiter the instant the question surfaced, so the
* target has neither an accepted nor a queued delivery left to show for it.
*
* <p><strong>Deliberately still {@code question == null} only (fleetd #275).</strong> This
* method must not also report a still-{@link Phase#ASKING} task as orphaned: the worker may
@@ -439,16 +528,15 @@ public final class MessageService {
* @throws IllegalArgumentException if {@code content} is {@code null} or blank — the caller must
* report this as a client error (REST: 400 {@code bad_request}) rather than resolve
* anything
* @return always {@code true} — the reply resolved a live send, completed a parked ticket, or
* was queued
* @return which of the three ways (fleetd #365) the reply actually landed — never {@code null}
*/
public boolean reply(String session, String content) {
public ReplyOutcome reply(String session, String content) {
if (content == null || content.isBlank()) {
throw new IllegalArgumentException("content is required");
}
if (rendezvous.resolve(session, content)) {
count(FleetMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
return ReplyOutcome.RESOLVED_SEND; // a live send took it — unchanged fast path
}
// #137/fleetd #307: no live rendezvous waiter, but this may be the worker's real fleet_reply resuming
// a turn that either answer() (#137) or ask() (fleetd #307) already gave up waiting on:
@@ -466,11 +554,21 @@ public final class MessageService {
if (candidates.size() == 1) {
Task orphan = candidates.get(0);
if (orphan.future.complete(new Reply(Outcome.REPLIED, content))) {
if (orphan.turnId != null) {
asyncTasksByTurn.remove(orphan.turnId, orphan);
// fleetd #329 (F3): read orphan.turnId once. It used to be read twice, under no
// lock — once for this null check, once as the removal key — the same double-read
// shape fleetd #324 fixed in finishAsyncTask. clearAsyncQuestion's unlocked
// forgetTurn=true path (ask()'s timeout cleanup) can null the field between the two
// reads; capturing it once removes the torn read here too.
String turnId = orphan.turnId;
if (turnId != null) {
if (replyOrphanTurnIdRaceHookForTest != null) {
// Test-only (fleetd #329, F3): see the field's own javadoc.
replyOrphanTurnIdRaceHookForTest.run();
}
asyncTasksByTurn.remove(turnId, orphan);
}
count(FleetMetrics.REPLIES, "path", "async-recovered");
return true; // the ticket itself took it — no inbox stranding at all
return ReplyOutcome.RESOLVED_ASYNC_TICKET; // the ticket itself took it — no inbox stranding
}
} else if (candidates.size() > 1) {
List<String> tickets = candidates.stream().map(t -> t.ticket).toList();
@@ -488,7 +586,7 @@ public final class MessageService {
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
return ReplyOutcome.QUEUED; // held, not lost — but not delivered either
}
/**
@@ -555,7 +653,7 @@ public final class MessageService {
return switch (o) {
case REPLIED -> "replied";
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, TIMED_OUT_UNCONFIRMED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case STALE_TURN, QUESTION -> null; // not a completed delegation
@@ -657,7 +755,7 @@ public final class MessageService {
* failure.
*
* <p><strong>Without {@code sweepAsking} on the release path, a target torn down while
* genuinely {@code ASKING} was unrecoverable.</strong> {@link #resolveQuestion} had already
* genuinely {@code ASKING} was unrecoverable.</strong> {@link Rendezvous#resolveQuestion(String, String, String)} had already
* closed the forward waiter the instant the question surfaced (so the {@code waiter} branch
* below finds nothing to fail), the {@code question == null} guard excluded the task from
* {@code matching} (so the loop below skipped it too), and the worker's own {@code fleet_ask}
@@ -699,27 +797,47 @@ public final class MessageService {
boolean isRecovery = task == recoveryTask && recovered != null;
Reply outcome = isRecovery ? recovered : new Reply(Outcome.WORKER_FAILED, reason);
String turnId = task.turnId;
if (task.future.complete(outcome)) {
if (outcome.outcome() == Outcome.WORKER_FAILED) {
asyncFailed = true;
// fleetd #335: completing THIS task's future must not depend on any other task's
// cleanup succeeding — every task in `matching` is owed its own outcome regardless of
// what happens below, so decide and record that before doing anything that can throw.
boolean completedHere = task.future.complete(outcome);
if (completedHere && outcome.outcome() == Outcome.WORKER_FAILED) {
asyncFailed = true;
}
try {
if (abandonCleanupHookForTest != null) {
// Test-only (fleetd #335 site 1): see the field's own javadoc.
abandonCleanupHookForTest.run();
}
if (turnId != null) {
// #275: whether this task was swept out of ASKING or was already answered and
// only waiting on its resumed turn's real reply (#137), nothing will ever
// complete this turnId now — drop it from this class's own bookkeeping AND the
// reverse-rendezvous itself, so hasAsyncQuestion(target) stops reporting a turn
// that is actually done, and a late answer() sees it as lapsed rather than
// resolving a question nothing is listening for any more.
asyncTasksByTurn.remove(turnId, task);
rendezvous.closeAsk(turnId);
if (completedHere) {
if (turnId != null) {
// #275: whether this task was swept out of ASKING or was already answered and
// only waiting on its resumed turn's real reply (#137), nothing will ever
// complete this turnId now — drop it from this class's own bookkeeping AND the
// reverse-rendezvous itself, so hasAsyncQuestion(target) stops reporting a turn
// that is actually done, and a late answer() sees it as lapsed rather than
// resolving a question nothing is listening for any more.
asyncTasksByTurn.remove(turnId, task);
rendezvous.closeAsk(turnId);
}
} else if (isRecovery) {
// The recovered reply was already drained out of the inbox, but this task resolved
// through another path (e.g. a concurrent reply() or a second abandon() racing this
// one) between us choosing it and completing it here. Put the reply back rather than
// lose it silently — it may still belong to some other still-open task, or the next
// caller that drains this target's inbox.
inbox.publish(target, UUID.randomUUID().toString(), recovered.text());
}
} else if (isRecovery) {
// The recovered reply was already drained out of the inbox, but this task resolved
// through another path (e.g. a concurrent reply() or a second abandon() racing this
// one) between us choosing it and completing it here. Put the reply back rather than
// lose it silently — it may still belong to some other still-open task, or the next
// caller that drains this target's inbox.
inbox.publish(target, UUID.randomUUID().toString(), recovered.text());
} catch (RuntimeException e) {
// fleetd #335: inbox.publish reaches a broker (AmqpReplyInbox throws
// IllegalStateException on an unroutable/unconfirmed/interrupted publish) and this
// loop has no other teardown path — a caller on the release path, or the health
// monitor's GONE/NEVER_READY sweep. Losing this exception uncaught would abort the
// loop and leave every task still to come in `matching` PENDING forever (fleetd
// #335 site 1). completedHere is already recorded above, so only this task's
// best-effort bookkeeping is lost — log it and let the loop reach the rest.
log.error("abandon: per-task cleanup failed for ticket {} (target {}, turnId {})",
task.ticket, target, turnId, e);
}
}
if (failed) {
@@ -750,9 +868,13 @@ public final class MessageService {
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
*
* @return {@code true} if an entry was actually removed, {@code false} if {@code msgId} was not
* in {@code target}'s inbox (wrong id, wrong target, or already acked). The caller —
* {@link dev.ltms.fleet.mcp.FleetMcp#ack} — must not report success on {@code false}.
*/
public void ackReply(String target, String msgId) {
inbox.ack(target, msgId);
public boolean ackReply(String target, String msgId) {
return inbox.ack(target, msgId);
}
/**
@@ -853,20 +975,42 @@ public final class MessageService {
if (onAccepted != null) {
onAccepted.run();
}
CompletableFuture<Void> delivered = injector.enqueue(target, content, token);
Injector.Delivery delivery = injector.enqueue(target, content, token);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
boolean wasDelivered = delivery.completion().isDone()
&& !delivery.completion().isCompletedExceptionally();
Injector.Cancellation cancellation = null;
if (!wasDelivered) {
// CB-640: still sitting in the injector's queue, waiting for the member to
// go idle — record the fact for fleet health (see queuedDeliveries).
queuedDeliveries.put(target, Boolean.TRUE);
if (timeoutCancellationRaceHookForTest != null) {
// Test-only (fleetd #345): see the field's own javadoc.
timeoutCancellationRaceHookForTest.run();
}
// The target monitor makes cancellation atomic with onStatus picking this
// Pending up. If pickup won, report TIMED_OUT_WORKING because the text landed.
cancellation = injector.cancel(delivery);
wasDelivered = cancellation == Injector.Cancellation.DELIVERED;
}
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
Outcome outcome;
if (wasDelivered) {
outcome = Outcome.TIMED_OUT_WORKING;
} else if (cancellation == Injector.Cancellation.ATTEMPTED) {
// fleetd #571: the call to the target's terminal was made and its outcome is
// unknown — the message may already have arrived in full, so this must not
// be reported as TIMED_OUT_QUEUED, which promises it never will.
outcome = Outcome.TIMED_OUT_UNCONFIRMED;
} else {
// CB-640: record that delivery is not confirmed, for fleet health (see
// queuedDeliveries). cancellation is CANCELLED (this call removed a
// still-queued Pending) or NOT_DELIVERED (an earlier attempt already failed
// with a confirmed absence) — both mean the target saw nothing.
queuedDeliveries.put(target, Boolean.TRUE);
outcome = Outcome.TIMED_OUT_QUEUED;
}
return recorded(new Reply(outcome, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
@@ -925,13 +1069,41 @@ public final class MessageService {
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("fleet_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
// fleetd #307: mark the task BEFORE clearAsyncQuestion(forgetTurn=true) below drops it out of
// asyncTasksByTurn and nulls its turnId — that forgetting is deliberate and stays (it is
// what keeps the target from staying BUSY forever), but it would otherwise also erase
// askAnsweredAsyncTasks' only signal that the worker's eventual real fleet_reply still
// belongs to this task, stranding it in the inbox with a false "never replied" verdict.
markAskTimedOut(ticket.turnId());
clearAsyncQuestion(ticket.turnId(), true);
// Only the fresh owner tears down the shared turn (mirrors the finally block below).
// A duplicate's own timeoutMillis says nothing about whether the SHARED ask is actually
// done — it must leave the close/forget bookkeeping to the fresh owner, exactly as it
// already leaves closeAsk to it.
if (ticket.fresh()) {
// fleetd #334: close the ask turn BEFORE forgetting this task's turnId mapping below.
// Before this fix the order was reversed — the mapping was forgotten here first, and
// rendezvous.closeAsk only ran afterward, in the shared finally. A primary's answer()
// call racing this exact timeout could then find rendezvous.askSession(turnId) still
// non-null (the ask still "answerable") after the Task mapping was already gone:
// answer()'s own asyncTasksByTurn lookup returned null, its task != null guard skipped
// the completion, and the async ticket sat at PENDING forever even though answer()
// itself reported the worker's real reply. Closing here first removes that window:
// any answer() call that still observes askSession(turnId) != null is necessarily
// racing a point BEFORE the forgetting below runs (both happen on this one thread, in
// this order, with nothing that yields in between), so the Task mapping is still there
// for it to find; any call that observes askSession(turnId) == null now correctly
// bails out STALE_TURN (see answer()'s own top check) before ever reaching
// asyncTasksByTurn. rendezvous.closeAsk is idempotent — a no-op once the turn is
// already removed, see its own javadoc — so the shared finally below re-running it
// for this same fresh call is harmless.
rendezvous.closeAsk(ticket.turnId());
if (askTimeoutRaceHookForTest != null) {
// Test-only (fleetd #334): see the field's own javadoc.
askTimeoutRaceHookForTest.run();
}
// fleetd #307: mark the task BEFORE clearAsyncQuestion(forgetTurn=true) below drops it
// out of asyncTasksByTurn and nulls its turnId — that forgetting is deliberate and
// stays (it is what keeps the target from staying BUSY forever), but it would
// otherwise also erase askAnsweredAsyncTasks' only signal that the worker's eventual
// real fleet_reply still belongs to this task, stranding it in the inbox with a false
// "never replied" verdict.
markAskTimedOut(ticket.turnId());
clearAsyncQuestion(ticket.turnId(), true);
}
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
@@ -978,21 +1150,33 @@ public final class MessageService {
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
// #282: mirror send()'s registration (:802) so a SECOND fleet_ask inside this same
// resumed turn can re-associate the async ticket with its new turnId via
// markAsyncQuestion — without this, that second ask has no Task to attach to, and
// markAsyncQuestion silently returns null.
Task task = asyncTasksByTurn.get(turnId);
if (task != null) {
asyncTasksByWaiter.put(reply, task);
}
if (!rendezvous.answerAsk(turnId, content)) {
asyncTasksByWaiter.remove(reply);
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
clearAsyncQuestion(turnId, false);
// fleetd #575: this try used to open below, AFTER the Task lookup/registration and the
// STALE_TURN early return that follows it — so that return was covered only by a
// hand-rolled copy of the finally's own cleanup pair, not the finally itself. Widening the
// try up to wrap the registration closes the gap structurally: every exit from here on,
// STALE_TURN included, now runs through the one finally below exactly once, and the
// duplicated pair is gone. This did not fix a live leak — see the ticket: neither
// rendezvous.answerAsk nor clearAsyncQuestion(turnId, false) can throw, so nothing ever
// actually left through the old gap uncovered — but #572 found this exact drift (one
// finally asserted, an identical sibling not) on this same file, and the hand-rolled copy
// was the wrong shape to keep regardless of whether it was ever exercised.
try {
// #282: mirror send()'s registration (:802) so a SECOND fleet_ask inside this same
// resumed turn can re-associate the async ticket with its new turnId via
// markAsyncQuestion — without this, that second ask has no Task to attach to, and
// markAsyncQuestion silently returns null.
Task task = asyncTasksByTurn.get(turnId);
if (task != null) {
asyncTasksByWaiter.put(reply, task);
}
if (answerAskLapseRaceHookForTest != null) {
// Test-only (fleetd #575): see the field's own javadoc.
answerAskLapseRaceHookForTest.run();
}
if (!rendezvous.answerAsk(turnId, content)) {
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
clearAsyncQuestion(turnId, false);
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
Reply result = new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
// #282: this waiter can resolve with a FRESH question rather than a terminal reply —
@@ -1002,13 +1186,58 @@ public final class MessageService {
// Measured when #282 was merged: this guard is DEFENCE IN DEPTH, not the thing
// that makes the chained ask work. ask() calls markAsyncQuestion (:860) before
// resolveQuestion (:861), so by the time this thread wakes, the task has already
// moved to the new turnId and finishAsyncTask(oldTurnId, ...) finds nothing. Removing
// this guard alone leaves the test green. Keep it anyway: it mirrors sendAsync's
// sibling guard, and that sibling's own comment (:1017) warns the two orderings are
// not something to rely on. Do NOT delete it as dead code without re-checking that
// ordering, and do not treat it as the sole protection either.
if (result.outcome() != Outcome.QUESTION) {
finishAsyncTask(turnId, result);
// moved to the new turnId — this outcome check, not a Task lookup, is what tells the
// two cases apart (see fleetd #329 below). Removing this guard alone leaves the test
// green. Keep it anyway: it mirrors sendAsync's sibling guard, and that sibling's own
// comment (:1017) warns the two orderings are not something to rely on. Do NOT delete
// it as dead code without re-checking that ordering, and do not treat it as the sole
// protection either.
//
// fleetd #329 (F1): complete the SAME Task object this method already looked up at
// :991, rather than re-resolving it from turnId a second time. The old
// finishAsyncTask(turnId, result) did its own asyncTasksByTurn.get(turnId) here, and
// that second lookup races ask()'s own timeout path: ask()'s ticket.answer().get(...)
// can time out (or lose that exact race) at essentially the same instant this method's
// rendezvous.answerAsk(turnId, ...) above already succeeded, running
// markAskTimedOut + clearAsyncQuestion(turnId, true) with no lock at all — which
// forgets turnId (removes it from asyncTasksByTurn, nulls Task.turnId) before this
// thread ever gets here. The worker's real reply then arrived, this method's own wait
// woke up with it, and the by-turnId lookup found nothing: the async ticket's future
// was never completed, so fleet_poll{ticket} stayed PENDING forever even though
// answer() itself correctly returned REPLIED. Reusing the reference captured at :991
// — before that race window opens — sidesteps the second lookup entirely: it is the
// identical Task whatever asyncTasksByTurn or Task.turnId say by the time we reach
// this point, so finishAsyncTask(Task, Reply) can still complete its future and detach
// it using whatever turnId it now reads. The #282 chained-ask case is unaffected
// because it is still gated purely by result.outcome() == QUESTION above, which does
// not depend on this lookup — widening what "task" means here cannot complete a ticket
// the chained ask deliberately left open.
//
// A null task is NOT only "this was never an async ticket". That reading was in this
// comment when #329 merged and it is wrong; it is still not the whole story after
// #334. A genuine async ticket can still land here with task == null — a blocking
// (wait:true) send's fleet_ask never has a Task at all, so that case is expected and
// fine. What #334 fixed was a SECOND, unintended way to get here with task == null:
// ask()'s timeout path used to run clearAsyncQuestion(turnId, true) — which drops the
// asyncTasksByTurn entry — in its catch block, while rendezvous.closeAsk(turnId) ran
// later, in its finally. Between those two the ask was still answerable but the map
// entry was already gone, so the lookup at :1053 returned null and this ticket was
// never completed. Measured on 2026-09-04: a probe firing only that first half before
// answer() ran printed "answer=REPLIED phase=PENDING reply=null" — the same stranded
// ticket #329 set out to fix, one step earlier in the same race; the probe used
// forgetTurnForTest, which omits ask()'s markAskTimedOut, and that omission does not
// change the outcome, because askTimedOut is read only by askAnsweredAsyncTasks, and
// reply() never reaches it while this method's own waiter is live. #334's fix
// reorders ask()'s timeout catch to run closeAsk before the forgetting (see the
// fresh-owner block there), which removes this path entirely rather than narrowing
// it further: once closeAsk has run, rendezvous.askSession(turnId) is null and
// answer() returns STALE_TURN from its own top check, before it ever reaches this
// lookup — see aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket in
// MessageServiceTest, which pins the exact window with askTimeoutRaceHookForTest.
// So by the time this line runs, task == null means only the ordinary blocking-send
// case (or #329's own already-fixed race elsewhere) — not this one.
if (result.outcome() != Outcome.QUESTION && task != null) {
finishAsyncTask(task, result);
}
return result;
} catch (TimeoutException e) {
@@ -1059,14 +1288,25 @@ public final class MessageService {
// this fires exactly once, from whichever path completes it: finishAsyncTask(task, result)
// below on any non-QUESTION outcome of send() — a worker's fleet_reply, the CB-106
// completion fallback, a CB-109 wedge, TIMED_OUT, BUSY, or BACKEND_EXHAUSTED — the same
// finishAsyncTask reached via answer()'s finishAsyncTask(turnId, result) once a QUESTION
// finishAsyncTask reached via answer()'s own finishAsyncTask(task, result) once a QUESTION
// is resolved, completeExceptionally(t) just below when send() itself throws, or a CB-516
// abandon() on teardown. Without this, MessageService.reply's rendezvous fast path (the
// one an async ticket always takes) never told the push loop anything happened — see the
// class javadoc on sendAsync/CB-107.
task.future.whenComplete((reply, ex) -> {
boolean failed = ex != null || reply == null || !reply.completed();
pushLoop.onTicketTerminal(ticket, target, failed);
try {
pushLoop.onTicketTerminal(ticket, target, failed);
} catch (Throwable t) {
// fleetd #335 (site 2): this stage's own CompletableFuture is discarded, so an
// uncaught throw here (e.g. a RejectedExecutionException from
// ReplyPushLoop's scheduler, already shut down while this in-flight send's
// whenComplete fires during the daemon's own shutdown sequence — messages.close()
// only stops accepting NEW async work, it does not cancel a delivery already
// running) vanishes with no log line and no metric, and the push loop never learns
// the ticket went terminal — the exact thing this hook exists to tell it.
log.error("push loop failed to learn ticket {} (target {}) went terminal", ticket, target, t);
}
});
}
asyncExecutor.submit(() -> {
@@ -1079,7 +1319,22 @@ public final class MessageService {
finishAsyncTask(task, result);
}
} catch (Throwable t) {
task.future.completeExceptionally(t);
// fleetd #329 (F2): finishAsyncTask above already completes task.future — on its very
// first line — before doing anything else, so anything that throws afterward (inside
// finishAsyncTask's own cleanup, or from a future addition to this try block) lands
// here with the future already resolved. completeExceptionally on an already-completed
// future is a silent no-op: it returns false and does nothing, so without the check
// below the exception simply vanished — no log, no metric, nothing. Measured (see the
// ticket): temporarily reintroducing the fleetd #324 NPE reproduced 19 real exceptions
// on the ordinary path, with 74/74 tests staying green and not one log line produced.
// Do not stop completing the future first — finishAsyncTask completing before it
// cleans up is what makes a late failure harmless to the ticket's own result — only
// add the missing visibility for the case where that step, or whatever ran after it,
// has already lost the race to report through the future.
if (!task.future.completeExceptionally(t)) {
log.error("async send {} -> {} threw after its ticket was already resolved",
ticket, target, t);
}
}
});
pruneTerminalTickets();
@@ -1132,6 +1387,27 @@ public final class MessageService {
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
}
/**
* Test seam only — carries no production behaviour, and nothing in this class calls it;
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
*
* <p>Reports whether {@code ticket}'s completion hook (the {@code whenComplete} registered in
* {@link Task}'s constructor) has actually run yet. Exists because {@link #poll} can report
* {@link Phase#DONE} for a ticket before that hook fires: {@code CompletableFuture.complete()}
* publishes its result and only afterwards runs dependent actions such as {@code whenComplete}
* (fleetd #399), so a caller that observes the future done via {@link #poll} is not thereby
* guaranteed to also observe {@link Task#completedNanos} stamped. A test that must order both
* events — e.g. before advancing an injected clock past the TTL, to avoid stamping the
* *advanced* time and masking a real eviction bug — waits on this instead of on
* {@link Phase#DONE}.
*
* @return {@code false} for an unknown ticket or one whose completion hook has not run yet
*/
boolean isCompletionStampedForTest(String ticket) {
Task task = tasks.get(ticket);
return task != null && task.completedNanos != null;
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
@@ -1230,20 +1506,198 @@ public final class MessageService {
}
}
/** Complete and detach an async ticket after its worker's actual terminal reply. */
/**
* Complete and detach an async ticket after its worker's actual terminal reply.
*
* <p><strong>fleetd #324.</strong> {@code task.turnId} is read into {@code turnId} exactly once.
* It used to be read twice — once for the null check, once as the removal key — and {@code
* volatile} makes each of those reads individually fresh but does not make the pair atomic.
* {@link #answer} calls this while holding {@code sessionLocks} for the target; {@link #ask}'s
* own timeout path calls {@link #clearAsyncQuestion} (which nulls {@link Task#turnId}) under no
* lock at all. When that unlocked null-out landed between the two reads here, the second read saw
* {@code null} and {@code asyncTasksByTurn.remove(null, task)} threw {@code NullPointerException}
* on the lead's own {@code answer()} call — even though {@code task.future.complete(result)} on
* the line above had already run, so the answer was in fact delivered. Capturing the field once
* removes the torn read; see the ticket for why the wider asymmetry between the locked and
* unlocked sides is not fixed by this alone.
*/
private void finishAsyncTask(Task task, Reply result) {
task.future.complete(result);
if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
if (afterFinishAsyncTaskCompleteHookForTest != null) {
// Test-only (fleetd #329, F2): see the field's own javadoc.
afterFinishAsyncTaskCompleteHookForTest.run();
}
String turnId = task.turnId;
if (turnId != null) {
if (finishAsyncTaskRaceHook != null) {
// Test-only (fleetd #324): see the field's own javadoc.
finishAsyncTaskRaceHook.run();
}
asyncTasksByTurn.remove(turnId, task);
}
}
/** Complete the async ticket correlated to a specific answered turn. */
private void finishAsyncTask(String turnId, Reply result) {
Task task = asyncTasksByTurn.get(turnId);
if (task != null) {
finishAsyncTask(task, result);
}
/**
* Null in production; test seam for fleetd #324 — invoked from {@link #finishAsyncTask(Task,
* Reply)} right after {@code task.turnId}'s null-check passes and before the (now-local) value is
* used for the removal. A test installs this to force, deterministically, the exact interleaving
* that a real race between this method and {@link #ask}'s unlocked timeout cleanup can otherwise
* only produce by chance: firing it here reproduces "the field went null between the check and the
* use" against the pre-fix code, and demonstrates the fix tolerates it (the captured local is used
* unconditionally, so a hook that nulls the field afterward cannot affect this call).
*/
private volatile Runnable finishAsyncTaskRaceHook;
/**
* Test-only (fleetd #324): install {@link #finishAsyncTaskRaceHook}. Package-private so the test,
* in the same package, can reach it without widening any production API.
*/
void setFinishAsyncTaskRaceHookForTest(Runnable hook) {
this.finishAsyncTaskRaceHook = hook;
}
/**
* Null in production; test seam for fleetd #329 (F2) — invoked from {@link #finishAsyncTask(Task,
* Reply)} unconditionally, immediately after {@code task.future.complete(result)} runs (before
* {@link Task#turnId} is even read, so it fires regardless of whether this task was ever asked).
* A test installs this to force an exception into exactly the shape fleetd #329 identified:
* something throws inside {@link #sendAsync}'s executor task after the async ticket's future is
* already resolved, so the surrounding {@code catch (Throwable t)} can only report failure through
* {@code completeExceptionally} — a silent no-op on an already-completed future. Engineering a
* real exception to land in that exact post-completion window is what the ticket itself had to do
* by temporarily deleting a production guard (fleetd #324's {@code turnId != null} check); this
* hook drives the identical shape deterministically instead.
*/
private volatile Runnable afterFinishAsyncTaskCompleteHookForTest;
/**
* Test-only (fleetd #329, F2): install {@link #afterFinishAsyncTaskCompleteHookForTest}.
* Package-private so the test, in the same package, can reach it without widening any production
* API.
*/
void setAfterFinishAsyncTaskCompleteHookForTest(Runnable hook) {
this.afterFinishAsyncTaskCompleteHookForTest = hook;
}
/**
* Null in production; test seam for fleetd #345 — invoked in {@link #send}'s timeout path after
* {@link Injector.Delivery#completion()} reports incomplete and before {@link Injector#cancel}
* takes the target monitor. A test installs this to make {@code onStatus} pick the exact queued
* delivery up in that window, so {@code cancel} returns {@link Injector.Cancellation#DELIVERED}.
* This deterministically covers the caller's need to use that result rather than relying on a
* timing-sensitive real race.
*/
private volatile Runnable timeoutCancellationRaceHookForTest;
/**
* Test-only (fleetd #345): install {@link #timeoutCancellationRaceHookForTest}. Package-private
* so the test, in the same package, can reach it without widening any production API.
*/
void setTimeoutCancellationRaceHookForTest(Runnable hook) {
this.timeoutCancellationRaceHookForTest = hook;
}
/**
* Null in production; test seam for fleetd #329 (F3) — invoked from {@link #reply} right after
* the single local read of {@code orphan.turnId} passes its null-check and before that (now-local)
* value is used as the {@code asyncTasksByTurn} removal key. Mirrors {@link
* #finishAsyncTaskRaceHook} exactly, for the structurally identical double-read fleetd #324 fixed
* in {@link #finishAsyncTask}: a test installs this to force, deterministically, {@code
* clearAsyncQuestion}'s unlocked {@code forgetTurn=true} path nulling {@link Task#turnId} in that
* exact window, and to confirm the single-read fix tolerates it (the captured local is used
* unconditionally, so a hook that nulls the field afterward cannot affect this call).
*/
private volatile Runnable replyOrphanTurnIdRaceHookForTest;
/**
* Test-only (fleetd #329, F3): install {@link #replyOrphanTurnIdRaceHookForTest}. Package-private
* so the test, in the same package, can reach it without widening any production API.
*/
void setReplyOrphanTurnIdRaceHookForTest(Runnable hook) {
this.replyOrphanTurnIdRaceHookForTest = hook;
}
/**
* Test-only (fleetd #324): run the exact production cleanup {@link #ask}'s own timeout path runs
* unlocked — {@link #clearAsyncQuestion(String, boolean)} with {@code forgetTurn=true} — so a test
* can reproduce that specific mutation instead of hand-rolling an approximation of it.
*/
void forgetTurnForTest(String turnId) {
clearAsyncQuestion(turnId, true);
}
/**
* Null in production; test seam for fleetd #335 (site 1) — invoked from {@link #abandon(String,
* String, boolean)}'s per-task loop, once per task, right before that task's own cleanup
* (turnId bookkeeping, or the stranded-reply put-back) runs. A test installs this to inject a
* throw at that exact point deterministically.
*
* <p>The one production call there that can really throw is {@code inbox.publish} in the
* put-back branch — {@link AmqpReplyInbox#publish} reaches a broker and throws {@link
* IllegalStateException} on an unroutable, unconfirmed, or interrupted publish — but reaching
* that branch requires a second completion of the very task {@code abandon} is about to
* complete to win the race first (see the branch's own comment), and the #137 follow-up
* investigation above already found the combination this needs (a stranded reply coinciding
* with an open matching task) unreachable through the public API, not merely hard to time.
* This hook reproduces the resulting shape — a per-task cleanup throw — directly, the same
* technique {@link #finishAsyncTaskRaceHook} and {@link
* #afterFinishAsyncTaskCompleteHookForTest} already use for their own hard-to-time races.
*/
private volatile Runnable abandonCleanupHookForTest;
/**
* Test-only (fleetd #335, site 1): install {@link #abandonCleanupHookForTest}. Package-private
* so the test, in the same package, can reach it without widening any production API.
*/
void setAbandonCleanupHookForTest(Runnable hook) {
this.abandonCleanupHookForTest = hook;
}
/**
* Null in production; test seam for fleetd #334 — invoked from {@link #ask}'s {@code
* TimeoutException} catch, only for the fresh owner, right after {@code rendezvous.closeAsk}
* has run and before {@link #markAskTimedOut} / {@link #clearAsyncQuestion} forget this task's
* turnId mapping. A test installs this to call {@link #answer} for the very same {@code turnId}
* synchronously from inside that exact window, deterministically reproducing the race a real
* concurrent {@code answer()} call could otherwise only win by timing luck: with the ask already
* closed, that call must see {@code rendezvous.askSession(turnId) == null} and return {@link
* Outcome#STALE_TURN} immediately, never reaching {@code asyncTasksByTurn} at all — proving the
* window fleetd #334 describes (mapping forgotten while the ask was still "answerable") is
* closed, rather than merely narrowed the way fleetd #329 narrowed the sibling race in {@link
* #finishAsyncTask}.
*/
private volatile Runnable askTimeoutRaceHookForTest;
/**
* Test-only (fleetd #334): install {@link #askTimeoutRaceHookForTest}. Package-private so the
* test, in the same package, can reach it without widening any production API.
*/
void setAskTimeoutRaceHookForTest(Runnable hook) {
this.askTimeoutRaceHookForTest = hook;
}
/**
* Null in production; test seam for fleetd #575 — invoked from {@link #answer}, right after this
* call's own {@code Task} registration and right before its {@code rendezvous.answerAsk(turnId,
* content)} call. A test installs this to complete the SAME turnId's ask directly via {@link
* Rendezvous#answerAsk} from inside that exact window, deterministically reproducing what a
* second, concurrent {@code answer()} call racing to unblock the same ask can otherwise only
* win by timing luck: this call's own {@code askSession(turnId)} lookup at the top already saw
* the ask as open, but by the time it reaches {@code rendezvous.answerAsk} here, the other call
* already completed it (or the worker's own {@code ask()} teardown already closed it) — so this
* call must see {@code false} and return {@link Outcome#STALE_TURN}, exactly the "lapsed between
* the lookup and the unblock" case named at that call site. Proves the #575 fix (widening this
* method's try so a single finally covers this exit) does not change that outcome and still
* cleans this call's own {@code reply} up exactly once.
*/
private volatile Runnable answerAskLapseRaceHookForTest;
/**
* Test-only (fleetd #575): install {@link #answerAskLapseRaceHookForTest}. Package-private so the
* test, in the same package, can reach it without widening any production API.
*/
void setAnswerAskLapseRaceHookForTest(Runnable hook) {
this.answerAskLapseRaceHookForTest = hook;
}
/** A new send must not open a waiter while an async ticket owns this worker's paused turn. */
@@ -46,6 +46,13 @@ public interface ReplyInbox {
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
List<InboxMessage> peek(String target);
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
void ack(String target, String msgId);
/**
* Remove the reply {@code msgId} for {@code target} once the primary has taken it.
*
* @return {@code true} if an entry was actually removed, {@code false} if there was nothing to
* remove (unknown {@code target}, unowned {@code target}, or a {@code msgId} not held for
* it). A {@code false} is not an error — acking a {@code target} this daemon does not own is
* part of the normal contract, not a failure.
*/
boolean ack(String target, String msgId);
}
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -13,6 +14,7 @@ import java.util.Collection;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
@@ -42,7 +44,7 @@ import java.util.stream.Collectors;
* still holds an unacked message ({@link #pendingReplies}), tickets not yet collected
* ({@link #pendingTickets}), and open questions not yet answered or lapsed
* ({@link #pendingQuestions}) — and sends at most one combined nudge per tick
* ({@link #injectNudge(String, int, int, int)}). Work that arrives while the lead is busy is
* ({@link #injectNudge(String, int, int, int, int, int)}). Work that arrives while the lead is busy is
* never lost: it is re-read fresh on every tick until the lead is injectable or its own reminder
* cap ({@link #maxReminders}) is reached — each source spends from its own budget, so one source
* exhausting its cap does not stop nudges about the others (post-CB-590 regression fix; see
@@ -130,7 +132,15 @@ public final class ReplyPushLoop {
this.metrics = metrics;
}
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
/**
* Count one nudge outcome when a registry is wired; a no-op in unit tests.
*
* <p>fleetd #365: the {@code "sent"} outcome (renamed from {@code "delivered"}) records only
* that {@code agents.send} — a one-way herdr {@code agent.prompt} paste-and-submit — returned
* without throwing. Nothing in this loop, or anywhere downstream of it, confirms the pane
* actually read or acted on the text; there is no read-receipt concept at this layer. "Sent"
* says exactly that; "delivered" claimed more than this call can ever establish.
*/
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(FleetMetrics.PUSH_NUDGES, "outcome", outcome);
@@ -220,6 +230,26 @@ public final class ReplyPushLoop {
.collect(Collectors.toUnmodifiableSet());
}
/**
* Test seam only (fleetd #418): carries no production behaviour, and nothing in this class
* calls it. Exposes {@link #pendingQuestionTurnIdsFor} — the exact state {@link #decide} reads
* to decide whether a question keeps a lead's schedule alive.
*
* <p>{@code MessageService.ask()} does three things in order before a question is fully open to
* this loop: it flips the ticket's {@code poll()} phase to {@code Phase.ASKING}, then resolves
* the reverse-rendezvous waiter, then calls {@link #onQuestionOpened}, which is what actually
* populates {@link #pendingQuestions}. A test that barriers on {@code Phase.ASKING} observes only
* the first of those three steps — under load the asker thread can be descheduled between steps
* one and three, so the barrier releases before this method's underlying map is populated, and
* {@link #decide} correctly reports nothing pending yet. A test that must order itself after the
* state {@link #decide} actually reads waits on this instead of on the phase.
*
* @return an unmodifiable snapshot; empty for a lead with no open questions
*/
Set<String> pendingQuestionTurnIdsForTest(String lead) {
return pendingQuestionTurnIdsFor(lead);
}
private List<PendingIncident> pendingIncidentsFor(String lead) {
return pendingIncidents.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
}
@@ -375,6 +405,80 @@ public final class ReplyPushLoop {
return Action.WAIT_BUSY;
}
/**
* Resolve who to nudge about {@code target}, the way every public entry point below wants it:
* {@link PrimaryRegistry#nudgeTargetFor}, but only after checking the delegating lead it names
* is still actually there (fleetd #368).
*
* <p><strong>The bug this closes.</strong> {@code PrimaryRegistry.forgetDelegation} is wired to
* exactly one event — a worker's release — because that is the only teardown the daemon already
* observes for a session in this map. Nothing removes a binding when the LEAD half goes away: a
* lead that is closed, crashes, or is relaunched leaves {@code leadByTarget} entries pointing at
* a terminal that no longer exists. {@code nudgeTargetFor} falls back to the single known
* primary only when the map holds nothing for {@code target} — a stale non-null entry beats the
* fallback every time, which is exactly backwards: the fallback's own javadoc argues it is safe
* precisely in the case a stale entry now hides.
*
* <p><strong>The fix.</strong> Before trusting a recorded delegation, probe the lead the same
* way {@link #decide} already does every tick ({@code agents.status}) — cheap, since it is a
* local herdr round-trip, and it is the same signal {@code AgentControl.paneByTerminal} already
* trusts to tell a genuinely dead target from a live one. A lead that fails the probe is treated
* as if it had never been recorded: the stale entry is forgotten (self-healing, exactly like
* {@code AgentControl.paneByTerminal} already does on {@code agent_not_found}) and resolution is
* retried, which now reaches the fallback {@code nudgeTargetFor} was built to reach — the same
* empty-map state its javadoc already argues is correct.
*
* <p><strong>fleetd #368 review — only a positive "gone" reading forgets the binding.</strong>
* The first version of this method treated <em>any</em> {@code RuntimeException} from the probe
* as death, which is the #359 mistake repeated: a transient socket blip or a codec error on a
* perfectly live lead would silently and permanently unbind it, with no re-record ever coming.
* That is destructive on one bad reading, exactly what #359 shipped a two-reading guard to avoid
* for the analogous lead-tab-liveness question. {@link #isLive} now matches
* {@code AgentControl.agentCall}'s own narrower rule (see its {@code agent_not_found} check): only
* that specific, affirmative "herdr has no such agent" signal counts as gone. Every other failure
* — timeout, transport error, a decode error — is treated as still live and the binding is left
* alone, because guessing wrong here is unrecoverable while guessing "live" merely costs one more
* retry on the next tick, which {@link #decide} already tolerates.
*/
private Optional<String> resolveLiveLead(String target) {
Optional<String> lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty() || isLive(lead.get())) {
return lead;
}
log.debug("push: lead {} delegated to for {} is no longer live, forgetting the stale binding "
+ "and falling back", lead.get(), target);
primaryRegistry.forgetDelegation(target);
return primaryRegistry.nudgeTargetFor(target);
}
/**
* Whether {@code lead} should still be trusted: {@code false} only when herdr affirmatively
* reports the terminal gone ({@code agent_not_found}), never on a merely inconclusive failure.
*
* <p>fleetd #368 review: an earlier version returned {@code false} for any {@code RuntimeException},
* which made a transient herdr hiccup on a live lead indistinguishable from the lead actually
* being dead — and the caller's response to {@code false} ({@code forgetDelegation}) is
* destructive and permanent. Narrowed to the one code {@code AgentControl.agentCall} itself
* already treats as a genuine, resolvable absence (see its {@code agent_not_found} handling) —
* every other {@code RuntimeException} is treated as "still live" and the binding survives to be
* probed again next time, which costs nothing worse than one more retry.
*/
private boolean isLive(String lead) {
try {
agents.status(lead);
return true;
} catch (RuntimeException e) {
boolean gone = e instanceof HerdrException he && "agent_not_found".equals(he.code());
if (gone) {
log.debug("push: lead {} no longer exists ({})", lead, e.toString());
} else {
log.debug("push: liveness check for lead {} was inconclusive ({}); treating as live "
+ "rather than risk destroying a live binding", lead, e.toString());
}
return !gone;
}
}
// --- public entrypoints ----------------------------------------------------------------------
/**
@@ -385,7 +489,7 @@ public final class ReplyPushLoop {
* backstop until a lead is recorded.
*/
public void onReplyQueued(String target) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
return;
@@ -413,7 +517,7 @@ public final class ReplyPushLoop {
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
*/
public void onTicketTerminal(String ticket, String target, boolean failed) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
ticket, target);
@@ -448,7 +552,7 @@ public final class ReplyPushLoop {
* @param question the question text
*/
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}'s question (turnId {}), skipping nudge",
target, turnId);
@@ -476,7 +580,7 @@ public final class ReplyPushLoop {
Collection<String> profiles, int remainingCoolOffSeconds) {
Map<String, List<String>> targetsByLead = new ConcurrentHashMap<>();
for (String target : targets) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend incident {} has no known lead for target {}", incidentId, target);
continue;
@@ -498,7 +602,7 @@ public final class ReplyPushLoop {
* Without an owning lead, emit a warning because no control can act on the target.
*/
public void onBackendTargetUnmapped(String target, String reason) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend target {} could not map to a credential: {}", target, reason);
return;
@@ -654,7 +758,7 @@ public final class ReplyPushLoop {
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
questionReminderCount + 1, maxReminders,
replyTargets.size(), tickets.size(), questions.size());
countNudge("delivered");
countNudge("sent");
for (PendingIncident incident : incidents) {
if (pendingIncidents.remove(incident.key(), incident)) {
deliveredIncidents.add(incident.key());
@@ -29,8 +29,8 @@ public enum MemberRole {
* <p>Reads the repo and writes analysis. Never commits code and never opens a pull request —
* an architect that starts implementing has stopped doing the job that makes it useful.
*
* <p>Architects are the one member kind declared in config, because a lead addresses the same
* slots across many tickets and needs a stable name for them.
* <p>Architects are the one member kind with live slot binding, because a lead addresses the
* same slots across many tickets and needs a stable name for them.
*/
ARCHITECT,
@@ -43,6 +43,14 @@ public enum MemberRole {
*/
DEV,
/**
* Sweeps an assigned package for defects and reports several ranked findings.
*
* <p>Never changes code, commits, or opens a pull request. A hunt gathers evidence, which can
* include running the build, but leaves every fix to a later implementation unit.
*/
HUNTER,
/**
* Reviews a diff it did not write and reports one structured finding.
*
@@ -59,7 +67,7 @@ public enum MemberRole {
/**
* The {@code fleet:} block that holds this role's pool — {@code architects},
* {@code developers}, {@code reviewers}.
* {@code developers}, {@code hunters}, {@code reviewers}.
*
* <p>Plural, and not always the wire name: the pool of things a {@code dev} may run on reads
* naturally as {@code developers:}. The wire name stays the singular {@code dev}, because that
@@ -69,6 +77,7 @@ public enum MemberRole {
return switch (this) {
case ARCHITECT -> "architects";
case DEV -> "developers";
case HUNTER -> "hunters";
case REVIEWER -> "reviewers";
};
}
@@ -1,5 +1,7 @@
package dev.ltms.fleet.peer;
import dev.ltms.fleet.placement.PlacementDecision;
import java.nio.file.Path;
import java.util.List;
import java.util.Set;
@@ -143,9 +145,160 @@ public interface PeerLauncher {
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*
* <p>fleetd #425: for an implementation with role pools (a no-argument spawn is read as {@link
* MemberRole#DEV}, see {@link SpawnRequest}), this must be the profile a live spawn of that role
* would actually be placed on right now, not a value captured once at startup — a caller such as
* {@code fleet_profiles} relies on this to report a live, not frozen, fact.
*/
String defaultProfile();
/**
* The profile an unqualified spawn of {@code role} would resolve to right now — the role-aware,
* live counterpart of {@link #defaultProfile()} (fleetd #425).
*
* <p>A caller that must provision something profile-specific (working directory, parity overlay
* files) <em>before</em> the actual spawn — {@code SessionManager.acquireWithWorktree} is the one
* that exists today — needs the exact profile that spawn will use, for the caller's real role,
* not a role-agnostic guess. Calling {@link #defaultProfile()} for that purpose reads {@code
* MemberRole#DEV}'s answer regardless of the caller's actual role, which is wrong for any other
* role and can provision for a profile the spawn never lands on.
*
* <p>Default implementation returns {@link #defaultProfile()}, ignoring {@code role} — the right
* answer for a launcher with no role-pool concept of its own (e.g. a single {@code
* HerdrPeerLauncher} adapter, which is never reached this way in production: {@code
* CompositePeerLauncher} always fronts it and resolves roles itself).
*
* <p>fleetd #453: this default is deliberately <em>not</em> abstract — unlike {@link
* #spawn(SpawnRequest, PlacementDecision)} (fleetd #450), there is no live defect in inheriting
* it today, and the only current single-adapter implementer ({@code HerdrPeerLauncher}) is
* correct to do so. But it stays correct only as long as that holds: <strong>if a launcher ever
* routes more than one profile per role, it MUST override this method</strong>, or every role
* silently resolves to {@link #defaultProfile()} with no error and no log line. {@code
* HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)} — the override in {@code
* dev.ltms.fleet.member}, not the declaration below — names this method and {@link #place}
* explicitly as "unoverridden here" for exactly this reason. Read it before adding role-pool
* routing to any {@code HerdrPeerLauncher} subclass.
*
* <p>Who is forced to read which paragraph, because it is not symmetric. A new class that
* implements this interface directly must write a body for {@link #spawn(SpawnRequest,
* PlacementDecision)}, which is abstract here, so it lands on this javadoc. A subclass of
* {@code HerdrPeerLauncher} does not: that class already implements the method, and the
* subclass inherits the body. So for a subclass this paragraph is advice, not a gate.
*/
default String defaultProfileFor(MemberRole role) {
return defaultProfile();
}
/**
* The profile an <em>unqualified</em> spawn of {@code role} would actually be routed to right
* now — the same candidate list, the same {@code quarantined}/{@code coolingOff}/{@code
* modelOff} filtering, and the same {@code PlacementPolicy} that {@link #spawn} itself
* consults for a blank-profile request (fleetd #425 rework).
*
* <p>This is <em>not</em> {@link #defaultProfileFor}: that method answers "what is first in
* {@code role}'s pool", blind to quarantine, cool-off, and the model on/off gate — the right
* answer for a role-agnostic, best-effort report ({@code fleet_profiles}' {@code "default"}
* field), but the wrong one for a caller that needs the profile a spawn will actually land on.
* A quarantined or model-off pool-first profile makes {@link #defaultProfileFor} return a name
* an unqualified spawn will never be routed to.
*
* <p>Just the resolved name, not the full {@link PlacementDecision} — a caller that only wants
* to know the answer (a status report, a log line) can call this; a caller that will later
* <em>act</em> on the answer by spawning — provisioning a worktree for a specific profile
* before the peer exists is the one that matters — must call {@link #place} and carry the
* {@link PlacementDecision} itself through to {@link #spawn(SpawnRequest, PlacementDecision)}
* instead of calling this method and feeding the string back in as an explicit profile. Doing
* that re-enters {@link #spawn(SpawnRequest)}'s explicit-profile branch, which disagrees with
* the routing branch on purpose about what an excluded profile means: the routing branch (and
* {@link #place}) falls through a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
* profile to the next candidate, while the explicit branch refuses outright — correct for an
* operator who named that profile on purpose, wrong for a name that only ever came from placement
* itself. That accidental refusal is exactly the regression fleetd #425 rework round 2 fixes:
* the default implementation below delegates to {@link #place}, so the two can never drift apart,
* but a caller that resolves through this method alone and spawns separately can still recreate
* the round-1 defect for itself. (Before fleetd #435, this accident was also reachable through
* {@code maxLoad} specifically, because {@code FixedPlacementPolicy} — the default policy — did
* not evaluate it at all for automatic selection; #435 closed that gap, so a placement decision
* can no longer be at cap in the first place. The refusal-vs-fall-through disagreement above is
* the part that was never about {@code maxLoad} and is still real.)
*
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
* candidate in {@code role}'s pool is currently placeable
*/
default String routedProfileFor(MemberRole role) {
return place(role).profile();
}
/**
* Resolve, <em>without spawning</em>, the {@link PlacementDecision} an unqualified spawn of
* {@code role} would make right now — the same candidate list, the same {@code
* quarantined}/{@code coolingOff}/{@code modelOff} filtering, and the same {@code
* PlacementPolicy} {@link #spawn(SpawnRequest)}'s blank-profile branch itself consults (fleetd
* #425 rework).
*
* <p>Pair this with {@link #spawn(SpawnRequest, PlacementDecision)}, never with {@link
* #spawn(SpawnRequest)} fed the decision's profile as an explicit name — see {@link
* PlacementDecision}'s own javadoc for why that second form regressed.
*
* <p>Default implementation wraps {@link #defaultProfile()}, ignoring {@code role} and every
* placement condition — the right answer for a launcher with no pool or placement-policy
* concept of its own, matching {@link #defaultProfileFor}'s own default.
*
* <p>fleetd #453: same reasoning as {@link #defaultProfileFor}'s own #453 note — this default
* is deliberately not abstract (no live defect today, correct for the sole single-adapter
* implementer), but <strong>a launcher that ever routes more than one profile per role MUST
* override this method too</strong>, or placement silently ignores {@code role} for it. See
* {@code HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)}'s javadoc, which names this
* method as "unoverridden here" and why that is correct only for a single-profile adapter.
*
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
* candidate in {@code role}'s pool is currently placeable
*/
default PlacementDecision place(MemberRole role) {
return new PlacementDecision(defaultProfile());
}
/**
* Spawn against an already-resolved {@link PlacementDecision} from {@link #place}, honoring it
* completely: none of the conditions {@link #place} already applied — quarantine, cooling off,
* {@code maxLoad} (evaluated by every placement policy including the default {@code fixed},
* since fleetd #435), model-off — are re-evaluated here; {@code decision} already reflects them.
* This is not skipping a check {@code place} left undone; it is not repeating one {@code place}
* already did, and not re-opening the window between that decision and this spawn in which the
* underlying state could otherwise move. This is what lets a resolve-then-spawn caller
* ({@code SessionManager.acquireWithWorktree}, which must know the profile before it can
* provision a worktree for it) and a plain blank-profile {@link #spawn(SpawnRequest)} caller
* land on the exact same outcome for the exact same placement state (fleetd #425 rework,
* round 2).
*
* <p>{@code req}'s own {@link SpawnRequest#profileName()} is ignored in favor of {@code
* decision.profile()} — the caller is expected to have built {@code req} with a blank or
* matching profile; passing a request that names a <em>different</em>, explicit profile than
* the decision it is paired with is a caller bug this method does not attempt to detect.
*
* <p>No default implementation (fleetd #450): the two correct bodies disagree on purpose, so an
* implementer must choose one rather than silently inherit whichever this interface happened to
* provide. An implementer with no placement concept of its own — spawns a single profile, e.g.
* {@code HerdrPeerLauncher} — should delegate to {@link #spawn(SpawnRequest)} with the decision's
* profile named explicitly, since there is no separate routing path to honor there: the
* explicit-profile branch it re-enters and the routing branch {@link #place} would have used are
* the same thing. <strong>A launcher that routes across more than one profile — the way {@code
* CompositePeerLauncher} routes across every configured adapter — MUST NOT re-enter {@link
* #spawn(SpawnRequest)}.</strong> Doing so re-applies that single-argument method's
* explicit-profile checks ({@code enforceNotQuarantined}, {@code enforceNotCoolingOff}, {@code
* enforceMaxLoad}, {@code enforceModelEnabled} in {@code CompositePeerLauncher}), which can
* refuse the very profile {@link #place} just chose, if the underlying placement state moved in
* the window between the {@link #place} call and this one — the exact window this method and
* {@link PlacementDecision} exist to close (fleetd #444). Before #450 this was a {@code default}
* method that only {@code CompositePeerLauncher} overrode; a future placement-doing launcher
* could have inherited the re-entering body silently and never known. Making it abstract turns
* that silent inheritance into a compile error.
*
* @throws IllegalArgumentException if the decision names an unknown profile
*/
PeerHandle spawn(SpawnRequest req, PlacementDecision decision);
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
@@ -192,4 +345,66 @@ public interface PeerLauncher {
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
/**
* Model ids the operator has currently turned off in the central {@code models.allow:} list
* (fleetd #422) — empty for a launcher with nothing to gate against. {@code fleet_profiles}/
* {@code GET /profiles} (via {@code FleetMcp.profilesView}) call this to report which models
* are off, and MUST read this exact accessor rather than deriving their own answer: the fleetd
* #404 lesson is that a status field reading a different source than the behaviour it describes
* can drift from what the gate ({@code CompositePeerLauncher.enforceModelEnabled} and its
* candidate filter) actually enforces. A default of {@code Set.of()} keeps every other {@link
* PeerLauncher} implementer (the herdr adapters, and the two test-fake implementers) unchanged.
*
* <p>fleetd #422 follow-up: this alone cannot tell "no {@code models:} block at all" from "a
* {@code models:} block where nothing is currently off" — both report an empty set here. Delegates
* to {@link #modelGateState()} so the two facts always come from the one read {@link
* #modelGateState()}'s implementer makes; do not override this method separately from that one.
*/
default Set<String> disabledModels() {
return modelGateState().off();
}
/**
* Whether the central {@code models.allow:} gate (fleetd #422) is armed at all, together with
* which model ids are currently off — fleetd #422 follow-up. {@link #disabledModels()} alone
* cannot distinguish two states that both report an empty set: a host with no {@code models:}
* block (nothing is gated, and nothing can be) and a host WITH a {@code models:} block where
* nothing is currently turned off (the gate is armed and reporting zero). This method exists so
* a caller — the startup log, {@code fleet_profiles}/{@code GET /profiles} — can tell the two
* apart, the same reason {@code CompletionResolver.UnsetMeaning} exists: an accessor that can
* legitimately report "empty" must never let a caller guess why.
*
* <p>Default {@link ModelGateState#notConfigured()} — every launcher without a {@code models:}
* block to read from (the herdr adapters, and the two test-fake implementers), matching {@link
* #disabledModels()}'s own default of an empty set.
*/
default ModelGateState modelGateState() {
return ModelGateState.notConfigured();
}
/**
* fleetd #422 follow-up: the result of {@link #modelGateState()} — see that method's javadoc
* for why "armed" and "off" must be reported together from one read rather than as two
* separately-derived facts that a reload landing between them could make disagree.
*
* @param configured {@code true} when a {@code models:} block exists at all (armed), regardless
* of whether anything in it is currently turned off; {@code false} when there
* is no block to gate against
* @param off the model ids currently turned off; always empty when {@code configured} is
* {@code false}
*/
record ModelGateState(boolean configured, Set<String> off) {
public ModelGateState {
off = Set.copyOf(off);
}
public static ModelGateState notConfigured() {
return new ModelGateState(false, Set.of());
}
public static ModelGateState armed(Set<String> off) {
return new ModelGateState(true, off);
}
}
}
@@ -29,4 +29,9 @@ public record SpawnRequest(String profileName, String requestedCwd, String calle
String sessionName, String resumeSessionId) {
this(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId, null);
}
/** Return a copy of this request with {@code profileName} replaced by {@code profile}. */
public SpawnRequest withProfile(String profile) {
return new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, role);
}
}
@@ -3,6 +3,7 @@ package dev.ltms.fleet.placement;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
import java.util.Optional;
import java.util.OptionalLong;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
@@ -21,46 +22,158 @@ import java.util.function.LongSupplier;
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
* real sleep.
*
* <h2>Escalation (fleetd #466)</h2>
* A flat cooldown does not fit every exhaustion. A backend that reports "out of capacity for the
* rest of the hour" recovers in one cooldown; a weekly subscription limit does not — it keeps
* reporting exhausted on every attempt made before the window resets, so a flat 30-minute cooldown
* (the default {@code cooldownNanos}) means roughly 336 pointless spawn attempts across a week, one
* every cooldown.
*
* <p><strong>This class only ever sees the exhaustion signal.</strong> Its only production caller is
* {@code Fleetd.exhaustionSink}, wired to fire on a {@code BACKEND_EXHAUSTED} classification alone.
* The daemon's other outage state — a credential "cooling off" after repeated non-exhaustion
* backend errors (an HTTP 5xx storm, say) — is a separate mechanism, {@code BackendOutagePolicy},
* with its own short fixed 60s cooldown and no repeat tracking. The two are never merged: escalating
* on a cooling-off signal would turn a transient 5xx storm into a multi-hour backoff, which is
* exactly the failure this ticket is not asking for. Confirmed by reading every call site of
* {@link #quarantine} — {@code BackendOutagePolicy} has its own {@code coolOff} method and never
* calls this one.
*
* <p><strong>Mechanism</strong> — the {@link #withEscalation} constructors track, per credential, how
* many times in a row {@link #quarantine} has been called without an intervening "quiet" gap.
* Each call computes {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at
* {@code maxCooldownNanos}. A call counts as a continuation of the same streak — {@code repeatCount}
* increments — when it arrives no more than one base {@code cooldownNanos} after the previous
* quarantine's deadline (this covers both "still quarantined" and "quarantine just expired and it
* was exhausted again immediately"); otherwise the streak resets and this call is treated as a fresh
* first occurrence at the base cooldown.
*
* <p><strong>Reset, honestly stated.</strong> The ideal reset signal is "the cooldown expired and the
* next attempt succeeded" — but nothing in this codebase reports a spawn success back to this class
* (checked: {@code SessionManager} and {@code CompositePeerLauncher} never call any method here
* except {@link #quarantine}/{@link #isQuarantined}/{@link #remainingSeconds}, none of which is a
* success hook). Lacking that signal, the reset used here is a time-based proxy: a base-cooldown's
* worth of quiet — no exhaustion report for that credential — since the last quarantine ended. It is
* not proof the credential started working again, only the best available evidence without adding an
* active probe, which is out of scope by the operator's own design constraint (no automatic probing
* of a limited backend).
*
* <p><strong>Ceiling.</strong> {@code maxCooldownNanos} bounds the growth — an unbounded backoff is a
* permanent, unrecoverable-without-a-restart outage, which would be worse than the flat-rate bug this
* escalation fixes. {@link #withEscalation(LongSupplier, long)} defaults the ceiling to
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown (12x the 1800s default ≈ 6 hours), so a
* chronically exhausted credential still gets re-tried roughly every 6 hours instead of every 30
* minutes — about a dozen attempts a week instead of ~336.
*
* <p><strong>Backward compatibility.</strong> The original two-argument {@link #BackendQuarantine(
* LongSupplier, long)} constructor is unchanged in behaviour: it is exactly {@code
* withEscalation}'s mechanism with {@code backoffMultiplier = 1.0} and {@code maxCooldownNanos =
* cooldownNanos}, which collapses the formula back to the original flat {@code now + cooldownNanos}
* on every call regardless of history. Every existing call site (roughly 20 across the test suite,
* plus {@link #none()}) keeps its current shape and behaviour unchanged.
*/
public final class BackendQuarantine {
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
/** Default growth per consecutive exhaustion streak — see the class doc's Mechanism section. */
static final double DEFAULT_BACKOFF_MULTIPLIER = 2.0;
/** Default ceiling, expressed as a multiple of the base cooldown — see the class doc's Ceiling section. */
static final long DEFAULT_MAX_COOLDOWN_MULTIPLE = 12;
private final ConcurrentHashMap<String, QuarantineState> quarantines = new ConcurrentHashMap<>();
private final LongSupplier nowNanos;
private final long cooldownNanos;
private final double backoffMultiplier;
private final long maxCooldownNanos;
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
private final boolean inert;
/** How many consecutive exhaustion reports a credential is on, and when the resulting cooldown ends. */
private record QuarantineState(int repeatCount, long deadlineNanos) {
}
/**
* Flat cooldown, unchanged from before fleetd #466 — every {@link #quarantine} call blocks the
* credential for exactly {@code cooldownNanos}, regardless of how many times it was called
* before. Equivalent to {@link #withEscalation} with no growth ({@code backoffMultiplier = 1.0})
* and a ceiling equal to the base cooldown, so it degrades to the identical {@code now +
* cooldownNanos} formula every call. Kept for the existing call sites that want a fixed cooldown
* (and for tests exercising the fixed-cooldown shape in isolation); production wiring uses
* {@link #withEscalation} instead.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
* must be positive
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
this(nowNanos, cooldownNanos, false);
this(nowNanos, cooldownNanos, 1.0, cooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
/**
* Escalating cooldown (fleetd #466) — see the class doc's Mechanism/Reset/Ceiling sections.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be
* positive
* @param backoffMultiplier growth per consecutive exhaustion; must be {@code >= 1.0} ({@code 1.0}
* disables growth and is exactly the flat two-argument constructor)
* @param maxCooldownNanos ceiling on the escalated cooldown; must be {@code >= cooldownNanos}
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
long maxCooldownNanos) {
this(nowNanos, cooldownNanos, backoffMultiplier, maxCooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
long maxCooldownNanos, boolean inert) {
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
if (cooldownNanos <= 0) {
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
}
if (backoffMultiplier < 1.0) {
throw new IllegalArgumentException("backoffMultiplier must be >= 1.0: " + backoffMultiplier);
}
if (maxCooldownNanos < cooldownNanos) {
throw new IllegalArgumentException(
"maxCooldownNanos must be >= cooldownNanos: " + maxCooldownNanos + " < " + cooldownNanos);
}
this.cooldownNanos = cooldownNanos;
this.backoffMultiplier = backoffMultiplier;
this.maxCooldownNanos = maxCooldownNanos;
this.inert = inert;
}
/**
* Escalating cooldown with the fleetd #466 default shape: cooldown doubles
* ({@value #DEFAULT_BACKOFF_MULTIPLIER}x) per consecutive exhaustion streak, capped at
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown. This is what production wiring
* ({@code Fleetd.main}) uses.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be positive
*/
public static BackendQuarantine withEscalation(LongSupplier nowNanos, long cooldownNanos) {
return new BackendQuarantine(nowNanos, cooldownNanos, DEFAULT_BACKOFF_MULTIPLIER,
cooldownNanos * DEFAULT_MAX_COOLDOWN_MULTIPLE, false);
}
/**
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
*/
public static BackendQuarantine none() {
return new BackendQuarantine(() -> 0L, 1, true);
return new BackendQuarantine(() -> 0L, 1, 1.0, 1, true);
}
/**
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
* Quarantine {@code credentialId} starting now. On a flat instance (the two-argument
* constructor) this always blocks for exactly {@code cooldownNanos}, restarting the cooldown at
* full length on every call — a fresh refusal is fresh evidence the account is still exhausted,
* not a reason to let an earlier, shorter wait stand. On an escalating instance ({@link
* #withEscalation}) the cooldown grows with each call that arrives within one base cooldown of
* the previous deadline, and resets to the base cooldown once a call arrives after a longer gap
* — see the class doc.
*
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
* so recording a deadline would produce a quarantine that never expires — a credential locked out
@@ -73,7 +186,13 @@ public final class BackendQuarantine {
if (inert) {
return;
}
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
long now = nowNanos.getAsLong();
quarantines.compute(credentialId, (id, prev) -> {
int repeatCount = (prev == null || now - prev.deadlineNanos() > cooldownNanos)
? 1
: prev.repeatCount() + 1;
return new QuarantineState(repeatCount, now + escalatedCooldownNanos(repeatCount));
});
}
/** Whether {@code credentialId} is quarantined right now. */
@@ -87,6 +206,49 @@ public final class BackendQuarantine {
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
}
/**
* Remaining seconds together with which consecutive exhaustion this is (fleetd #466 scope item
* 2) — {@code repeatCount} 1 for a first occurrence, 2 for the second in a row, and so on; see
* {@link #quarantine}'s class-doc Mechanism section for exactly when a call continues a streak
* versus starts a fresh one.
*
* <p><strong>Read together, off the one {@link QuarantineState} entry {@link #quarantine} itself
* wrote</strong> — a single {@code quarantines.get(credentialId)}, never a separate lookup or a
* value re-derived from {@code remainingSeconds} (e.g. inverting {@link
* #escalatedCooldownNanos}). That inversion is not just extra work to avoid: once a streak has
* hit {@code maxCooldownNanos}, every further consecutive exhaustion reports the identical
* cooldown, so a derivation that starts from the cooldown value cannot tell the 4th repeat from
* the 9th — only the stored {@code repeatCount} can. This is the same rule {@code
* CompositePeerLauncher.modelGateState()} documents for its own gate/report pair: the report
* reads the exact accessor the behaviour reads, so it can never disagree with what actually
* happened (the fleetd #404/#422 lesson). {@code fleet_profiles}/{@code fleet_list}/{@code GET
* /profiles} all call this — never {@link #remainingSeconds} plus a second, independent count —
* for exactly that reason.
*
* @return empty when {@code credentialId} is not currently quarantined (including on {@link
* #none()}, which quarantines nothing)
*/
public Optional<Status> status(String credentialId) {
QuarantineState state = quarantines.get(credentialId);
if (state == null) {
return Optional.empty();
}
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
return remaining > 0
? Optional.of(new Status(toSecondsRoundedUp(remaining), state.repeatCount()))
: Optional.empty();
}
/**
* @param remainingSeconds seconds left on the quarantine, identical to {@link
* #remainingSeconds(String)}'s answer for the same credential at the
* same instant
* @param repeatCount 1 for a first occurrence, 2 for the second consecutive one, etc. —
* see {@link #status(String)}
*/
public record Status(long remainingSeconds, int repeatCount) {
}
/**
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
@@ -95,8 +257,8 @@ public final class BackendQuarantine {
*/
public Map<String, Long> activeRemainingSeconds() {
Map<String, Long> out = new LinkedHashMap<>();
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
long remaining = deadline - nowNanos.getAsLong();
quarantines.forEach((credentialId, state) -> {
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
if (remaining > 0) {
out.put(credentialId, toSecondsRoundedUp(remaining));
}
@@ -105,8 +267,14 @@ public final class BackendQuarantine {
}
private long remainingNanos(String credentialId) {
Long deadline = quarantinedUntilNanos.get(credentialId);
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
QuarantineState state = quarantines.get(credentialId);
return state == null ? 0L : state.deadlineNanos() - nowNanos.getAsLong();
}
/** {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at {@code maxCooldownNanos}. */
private long escalatedCooldownNanos(int repeatCount) {
double raw = cooldownNanos * Math.pow(backoffMultiplier, repeatCount - 1);
return raw >= (double) maxCooldownNanos ? maxCooldownNanos : (long) raw;
}
private static long toSecondsRoundedUp(long nanos) {
@@ -5,10 +5,12 @@ import java.util.List;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. Reachability is a narrower
* exception (fleetd #315, below): a profile is never checked for reachability up front, only
* skipped once it has already failed in <em>this same</em> spawn call's retry loop — see the
* unreachable case below.
*
* <p>Three exceptions walk past the default instead of returning it unconditionally:
* <p>Six exceptions walk past the default instead of returning it unconditionally:
* <ul>
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
* a usage limit, not a transient capacity or reachability concern.
@@ -16,13 +18,40 @@ import java.util.List;
* ({@code BackendOutagePolicy}) — a separate, shorter-lived source from quarantine. When a
* profile is both quarantined and cooling off, only the quarantine reason is reported
* (exhaustion takes priority), matching {@code CompositePeerLauncher}'s explicit-spawn order.
* <li>At cap (fleetd #435): a profile whose live count has reached its {@code maxLoad}
* ({@link PlacementPolicyUtil#atCap}) — a documented, unconditional capacity limit (see
* {@code FleetConfig.Profile#maxLoad}), so {@code fixed} must gate on it exactly as {@code
* weighted}/{@code round-robin} already do via {@link PlacementPolicyUtil#available}. Before
* this fix {@code fixed} built its own {@link PlacementCandidate} for the default with {@code
* maxLoad} forced to {@code null}, so a capped default was chosen anyway on every unqualified
* spawn — the cap was advisory, not enforced, for the one placement policy every config uses
* by default. Reported only when quarantine and cooling off are both absent, matching {@code
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
* load, then model-off).
* <li>Model off (fleetd #422): a profile whose {@code model:} the operator has turned off in
* {@code models.allow:} — an operator decision, never a backend-reported outage, so it is a
* fifth, independent source from quarantine, cooling off, and at-cap (never merged with any
* of them), exactly as {@code CompositePeerLauncher.enforceModelEnabled} and {@link
* PlacementPolicyUtil#available} treat it. When a profile is model-off <em>and</em> quarantined,
* cooling off, or at cap, only the higher-priority reason is reported, matching {@code
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
* load, then model-off).
* <li>Unreachable (fleetd #315): {@code CompositePeerLauncher.spawn} retries a failed candidate
* on the next one and rebuilds the {@link PlacementContext} so {@code ctx.unreachable()}
* names every profile that already failed with {@code PeerUnreachableException} in this same
* call. Without this check {@code select} kept handing back the same dead default forever —
* the retry loop's own comment says "so the policy excludes this profile", and this is what
* makes that true for {@code fixed} too, matching {@code weighted}/{@code round-robin}
* (both filter on {@code ctx.unreachable()} via {@link PlacementPolicyUtil#available}).
* <li>Weight 0 (CB-554): {@code fixed} is still automatic selection, so a profile the operator
* marked "never auto-select me" ({@code weight <= 0}) must be skipped here exactly as
* {@code weighted}/{@code round-robin} skip it — an explicit {@code fleet_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined, cooling off, or weight-0 never exercises any of these
* paths, so today's behaviour is unchanged.
* A fleet where nothing is ever quarantined, cooling off, at cap, model-off, unreachable, or
* weight-0 never exercises any of these paths, so today's behaviour is unchanged — in particular,
* the very first selection of a spawn call always sees an empty {@code unreachable} set, so the
* first choice is untouched.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@@ -30,12 +59,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !ctx.coolingOff().contains(d)
&& !weightExcluded(ctx, d)) {
&& !ctx.modelOff().contains(d) && !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)
&& !capExcluded(ctx, d)) {
return new PlacementCandidate(d, null, 1.0f, null);
}
for (PlacementCandidate c : ctx.candidates()) {
if (!ctx.quarantined().contains(c.profile()) && !ctx.coolingOff().contains(c.profile())
&& !c.excluded()) {
&& !ctx.modelOff().contains(c.profile()) && !ctx.unreachable().contains(c.profile())
&& !c.excluded() && !PlacementPolicyUtil.atCap(ctx, c)) {
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
}
}
@@ -44,8 +75,18 @@ final class FixedPlacementPolicy implements PlacementPolicy {
// Exhaustion quarantine takes priority: reported only when quarantine is absent, so the
// message never claims "cooling off" for a profile that is really backend-exhausted.
boolean dCoolingOff = !dQuarantined && ctx.coolingOff().contains(d);
// fleetd #435: at-cap sits between cooling off and model-off, matching
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
// then model-off) — reported only when quarantine/cooling-off are both absent.
boolean dAtCap = !dQuarantined && !dCoolingOff && capExcluded(ctx, d);
// fleetd #422: model-off is a fifth, independent source (an operator decision) — but
// quarantine/cooling-off/at-cap still take priority when more than one applies, matching
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
// then model-off).
boolean dModelOff = !dQuarantined && !dCoolingOff && !dAtCap && ctx.modelOff().contains(d);
boolean dUnreachable = ctx.unreachable().contains(d);
boolean dWeightExcluded = weightExcluded(ctx, d);
if (dQuarantined || dCoolingOff || dWeightExcluded) {
if (dQuarantined || dCoolingOff || dAtCap || dModelOff || dUnreachable || dWeightExcluded) {
List<String> reasons = new ArrayList<>();
if (dQuarantined) {
reasons.add("is quarantined (backend exhausted)");
@@ -53,6 +94,17 @@ final class FixedPlacementPolicy implements PlacementPolicy {
if (dCoolingOff) {
reasons.add("is cooling off after repeated backend errors");
}
if (dAtCap) {
PlacementCandidate c = candidateFor(ctx, d);
int live = ctx.liveCount().apply(d);
reasons.add("is at maxLoad (" + live + " live >= " + c.maxLoad() + " cap)");
}
if (dModelOff) {
reasons.add("names a model the operator has turned off in models.allow");
}
if (dUnreachable) {
reasons.add("is unreachable");
}
if (dWeightExcluded) {
reasons.add("has weight 0 (excluded from automatic selection)");
}
@@ -62,18 +114,38 @@ final class FixedPlacementPolicy implements PlacementPolicy {
}
if (!ctx.candidates().isEmpty()) {
throw new PlacementException("all worker profiles are excluded from automatic "
+ "selection (quarantined, cooling off, or weight-0)");
+ "selection (quarantined, cooling off, at cap, model-off, unreachable, or weight-0)");
}
throw new PlacementException("no worker profiles configured");
}
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
private static boolean weightExcluded(PlacementContext ctx, String profile) {
PlacementCandidate c = candidateFor(ctx, profile);
return c != null && c.excluded();
}
/**
* Whether {@code profile} has reached its {@code maxLoad} cap (fleetd #435), using the shared
* {@link PlacementPolicyUtil#atCap} definition — the same one {@code weighted}/{@code
* round-robin} already consult via {@link PlacementPolicyUtil#available}. Looked up by name,
* the same way {@link #weightExcluded} is: the default fast path above builds its own {@link
* PlacementCandidate} with {@code maxLoad} forced to {@code null} (it carries no cap of its
* own), so the candidate actually configured for {@code profile} has to be found in {@code
* ctx.candidates()} first.
*/
private static boolean capExcluded(PlacementContext ctx, String profile) {
PlacementCandidate c = candidateFor(ctx, profile);
return c != null && PlacementPolicyUtil.atCap(ctx, c);
}
/** The configured candidate named {@code profile} in {@code ctx}, or {@code null} if none. */
private static PlacementCandidate candidateFor(PlacementContext ctx, String profile) {
for (PlacementCandidate c : ctx.candidates()) {
if (c.profile().equals(profile)) {
return c.excluded();
return c;
}
}
return false;
return null;
}
}
@@ -22,11 +22,29 @@ import java.util.function.Function;
* the two apart so its refusal message says "cooling off", not "exhausted",
* when only this one is active. A profile can be in both sets at once; when it
* is, exhaustion quarantine is reported (it takes priority).
* @param modelOff profiles whose {@code model:} is currently turned off in {@code
* models.allow:} (fleetd #422) — an operator decision, not a backend-reported
* outage, so a SEPARATE, independent source from both {@code quarantined} and
* {@code coolingOff}. A profile can be in this set together with either (or
* both) of the others; {@link PlacementPolicyUtil} counts it into its own
* bucket rather than merging it into theirs, the same reason
* {@code coolingOff} is kept apart from {@code quarantined}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable,
Set<String> quarantined,
Set<String> coolingOff) {
Set<String> coolingOff,
Set<String> modelOff) {
/**
* Back-compat form before the fleetd #422 model on/off gate was added — no candidate's model
* is off. Keeps pre-#422 call sites (tests included) compiling and behaving identically.
*/
public PlacementContext(String defaultProfile, List<PlacementCandidate> candidates,
Function<String, Integer> liveCount, Set<String> unreachable,
Set<String> quarantined, Set<String> coolingOff) {
this(defaultProfile, candidates, liveCount, unreachable, quarantined, coolingOff, Set.of());
}
}
@@ -0,0 +1,49 @@
package dev.ltms.fleet.placement;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
/**
* An already-completed placement choice — the outcome of one {@link PeerLauncher#place} call,
* carried forward so a later {@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} can honor
* it directly instead of re-resolving the profile a second time (fleetd #425 rework, round 2).
*
* <p>The problem this exists to close: a caller that must know the profile <em>before</em> it can
* spawn — {@code SessionManager.acquireWithWorktree} provisions a worktree's {@code repoRoot} and
* parity overlay for a specific profile before the peer process exists — used to resolve that name
* with {@code PeerLauncher.routedProfileFor(role)} and then hand the SAME string back to {@link
* PeerLauncher#spawn(SpawnRequest)} as an EXPLICIT profile. That re-resolution is not free: naming
* a profile explicitly makes {@code CompositePeerLauncher.spawn} take its THROWING branch
* ({@code enforceNotQuarantined}/{@code enforceNotCoolingOff}/{@code enforceMaxLoad}/{@code
* enforceModelEnabled}), while an unqualified spawn's ROUTING branch never runs those checks at
* all — it instead FALLS THROUGH to the next candidate on exactly the same conditions the throwing
* branch refuses on. That disagreement is deliberate: an operator who names a profile should get a
* refusal, not a silent substitution. The bug is turning the fall-through into a refusal by
* accident — resolving a name through the routing side and then re-entering the refusing side with
* it, for a decision the routing side had already approved by walking past everything else.
* Before fleetd #435, this accident was also reachable through {@code maxLoad} specifically: the
* default {@code fixed} placement policy did not evaluate {@code maxLoad} at all for automatic
* selection, so a profile placement itself just approved could still die at {@code enforceMaxLoad}
* one call later, purely because the caller's route to the spawn passed through an explicit
* profile name instead of the routing branch — a failure a worktree-less unqualified spawn would
* never hit. fleetd #435 closed that specific gap ({@code fixed} now evaluates {@code maxLoad}
* exactly like every other placement policy), so a {@link PlacementDecision} can no longer be
* at-cap in the first place — but the refusal-vs-fall-through disagreement above was never about
* {@code maxLoad}, and resolving a name and re-entering the refusing branch with it is still wrong
* for every OTHER condition placement filters on.
*
* <p>{@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} closes that by spawning through
* the identical code path the routing branch itself uses, keyed off the SAME decision {@link
* PeerLauncher#place} returned — no re-checking of any condition placement already evaluated. A
* resolve-then-spawn caller and a blank-profile {@link PeerLauncher#spawn(SpawnRequest)} caller can
* then never disagree about which conditions apply to the same placement state, and neither one
* re-opens the window between the placement decision and the spawn in which the underlying state
* could otherwise move.
*
* @param profile the profile this decision resolved to (may be {@code null} only when no profile is
* configured at all — the same corner case {@link PeerLauncher#defaultProfile()}
* already tolerates)
*/
public record PlacementDecision(String profile) {
}
@@ -11,28 +11,40 @@ final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* True when {@code c} has reached its {@code maxLoad} cap: {@code liveCount(c.profile()) >=
* c.maxLoad()}. A {@code null} maxLoad means unlimited, so it is never at cap.
*
* <p>Extracted as the single shared definition of "at cap" (fleetd #435): before this fix it
* was computed inline in both {@link #available} and {@link #emptyException}, and {@code
* FixedPlacementPolicy} — not a caller of either — quietly kept its own {@code select} free of
* any cap check at all, so a capped default profile was chosen anyway under the default
* placement policy. Every automatic policy must call this, not re-derive it.
*/
static boolean atCap(PlacementContext ctx, PlacementCandidate c) {
Integer cap = c.maxLoad();
return cap != null && ctx.liveCount().apply(c.profile()) >= cap;
}
/**
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
* first because it is a static config choice rather than transient state), not
* known-unreachable, not quarantined (CB-578 stage B), not cooling off after repeated backend
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), and have not
* reached their maxLoad. A {@code null} maxLoad means unlimited.
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), not naming a
* model the operator has turned off (fleetd #422 — a third, independent source: an operator
* decision, never a backend-reported outage), and have not reached their maxLoad (see {@link
* #atCap}). A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (c.excluded() || ctx.unreachable().contains(c.profile())
|| ctx.quarantined().contains(c.profile())
|| ctx.coolingOff().contains(c.profile())) {
|| ctx.coolingOff().contains(c.profile())
|| ctx.modelOff().contains(c.profile())
|| atCap(ctx, c)) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
@@ -40,12 +52,13 @@ final class PlacementPolicyUtil {
/**
* Build a clear exception describing why every candidate was dropped: all weight-0, all
* quarantined, all cooling off, all at capacity, all unreachable, or a mix. Each candidate is
* counted into exactly one bucket (weight-excluded first, then quarantined, then cooling off)
* so a candidate excluded for more than one reason is never double-counted — a candidate that is
* both quarantined (CB-578 stage B, backend exhausted) and cooling off (fleetd #201 Unit 5,
* repeated backend errors) counts only as quarantined, matching {@code CompositePeerLauncher}'s
* explicit-spawn ordering: exhaustion quarantine takes priority when both are active.
* quarantined, all cooling off, all model-off, all at capacity, all unreachable, or a mix. Each
* candidate is counted into exactly one bucket (weight-excluded first, then quarantined, then
* cooling off, then model-off) so a candidate excluded for more than one reason is never
* double-counted — a candidate that is both quarantined (CB-578 stage B, backend exhausted) and
* cooling off (fleetd #201 Unit 5, repeated backend errors) counts only as quarantined, matching
* {@code CompositePeerLauncher}'s explicit-spawn ordering: exhaustion quarantine takes priority
* when more than one applies.
*/
static PlacementException emptyException(PlacementContext ctx) {
int weightExcluded = 0;
@@ -53,17 +66,19 @@ final class PlacementPolicyUtil {
int unreachable = 0;
int quarantined = 0;
int coolingOff = 0;
int modelOff = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (c.excluded()) {
weightExcluded++;
} else if (ctx.quarantined().contains(c.profile())) {
quarantined++;
} else if (ctx.coolingOff().contains(c.profile())) {
coolingOff++;
} else if (ctx.modelOff().contains(c.profile())) {
modelOff++;
} else if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
} else if (atCap(ctx, c)) {
atCap++;
}
}
@@ -83,6 +98,10 @@ final class PlacementPolicyUtil {
return new PlacementException(
"all worker profiles are cooling off after repeated backend errors");
}
if (modelOff == total) {
return new PlacementException(
"all worker profiles name a model the operator has turned off");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
@@ -92,8 +111,9 @@ final class PlacementPolicyUtil {
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + quarantined + " quarantined, "
+ coolingOff + " cooling off, "
+ modelOff + " model-off, "
+ weightExcluded + " weight-0, "
+ (total - atCap - unreachable - quarantined - coolingOff - weightExcluded)
+ (total - atCap - unreachable - quarantined - coolingOff - modelOff - weightExcluded)
+ " remaining");
}
}
@@ -0,0 +1,93 @@
package dev.ltms.fleet.power;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.util.Locale;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicBoolean;
/**
* Holds macOS idle sleep off by keeping a {@code caffeinate -i} child process alive for the life
* of the returned {@link SleepAssertion}.
*
* <p>{@code -i} asserts only against <em>idle</em> sleep — it does not stop the lid closing or an
* operator-requested sleep from taking effect. That is deliberate: this class exists to stop an
* unattended host from sleeping out from under a member's long turn, never to override the
* operator. {@code -s}/{@code -d} (which also block system/display sleep on demand) are
* intentionally not used here.
*
* <p>{@link #acquire()} never throws. It returns {@code null} — a no-op — off macOS, and again if
* starting the {@code caffeinate} child fails for any reason (binary missing, process table full,
* …); either case is logged once at INFO, not on every occurrence, so a daemon that runs for
* weeks with the tool unavailable does not fill its log.
*/
public final class CaffeinateSleepAssertionMechanism implements SleepAssertionMechanism {
private static final Logger log = LoggerFactory.getLogger(CaffeinateSleepAssertionMechanism.class);
private final AtomicBoolean loggedOnce = new AtomicBoolean(false);
/** {@code true} when running on macOS, the only platform {@code caffeinate} ships on. */
public static boolean isSupportedPlatform() {
return isSupportedPlatform(System.getProperty("os.name"));
}
/** Package-visible so a test can drive the platform check without touching a real property. */
static boolean isSupportedPlatform(String osName) {
return osName != null && osName.toLowerCase(Locale.ROOT).contains("mac");
}
@Override
public SleepAssertion acquire() {
if (!isSupportedPlatform()) {
logOnce("not running on macOS (os.name={}); the idle-sleep guard is a no-op on this platform",
System.getProperty("os.name"));
return null;
}
try {
Process process = new ProcessBuilder("caffeinate", "-i")
.redirectOutput(ProcessBuilder.Redirect.DISCARD)
.redirectError(ProcessBuilder.Redirect.DISCARD)
.start();
return new CaffeinateAssertion(process);
} catch (IOException | RuntimeException e) {
logOnce("could not start 'caffeinate -i' ({}); the host may idle-sleep while members are live",
e.toString());
return null;
}
}
private void logOnce(String format, Object arg) {
if (loggedOnce.compareAndSet(false, true)) {
log.info("idle-sleep guard: " + format, arg);
}
}
/** Wraps the live {@code caffeinate} child; {@link #close} force-destroys it, idempotently. */
private static final class CaffeinateAssertion implements SleepAssertion {
private final Process process;
CaffeinateAssertion(Process process) {
this.process = process;
}
@Override
public void close() {
if (!process.isAlive()) {
return;
}
process.destroy();
try {
if (!process.waitFor(2, TimeUnit.SECONDS)) {
process.destroyForcibly();
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
process.destroyForcibly();
}
}
}
}
@@ -0,0 +1,105 @@
package dev.ltms.fleet.power;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.function.IntSupplier;
/**
* Holds an OS-level assertion against idle sleep for exactly as long as at least one fleet
* member is live.
*
* <p><strong>Why this exists:</strong> a fleetd host was measured idle-sleeping after as little
* as one minute of inactivity (its {@code pmset -g custom} reports {@code sleep 1} on battery).
* Overnight the daemon's AMQP link to the broker dropped 13 times, and cross-checking every drop
* minute against {@code pmset -g log} found a sleep or wake event in the same minute or the one
* before, every time. The AMQP churn is only the visible symptom — the real problem is that a
* member mid-turn freezes with the host, and a long turn with nobody typing is exactly the case
* that goes idle.
*
* <p><strong>How it tracks "live":</strong> this is driven by {@code SessionManager}'s existing
* {@code onAcquire}/{@code onRelease} lifecycle hooks (added for CB-520/CB-516, previously wired
* to nothing but the reply inbox) rather than a second member count kept in parallel. Wire it as:
* <pre>{@code
* IdleSleepGuard guard = new IdleSleepGuard(mechanism, sessions::size);
* sessions.onAcquire(_ -> guard.recheck());
* sessions.onRelease(_ -> guard.recheck());
* }</pre>
* Every acquire/release event re-reads {@code SessionManager#size()} — the same registry {@code
* fleet_list}'s live/capacity numbers are themselves computed from — and only an actual 0→1 or
* 1→0 crossing touches the OS. A listener exception is already caught and logged by {@code
* SessionManager} itself (it must never let a listener failure block the acquire/release it is
* reacting to), so {@link #recheck()} does not need its own top-level try/catch to honor that.
*
* <p><strong>Failure posture:</strong> every method here is safe to call whether or not {@link
* SleepAssertionMechanism#acquire()} actually works. A mechanism that returns {@code null} (wrong
* platform, missing tool, spawn failure) simply means this guard never holds anything — it never
* throws and never blocks a spawn, a release, or shutdown.
*/
public final class IdleSleepGuard implements AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(IdleSleepGuard.class);
private final SleepAssertionMechanism mechanism;
private final IntSupplier liveCount;
private final Object lock = new Object();
private SleepAssertion held;
public IdleSleepGuard(SleepAssertionMechanism mechanism, IntSupplier liveCount) {
this.mechanism = mechanism;
this.liveCount = liveCount;
}
/**
* Re-read the live count and acquire or release the held assertion to match: nothing held and
* at least one member live ⇒ acquire; something held and no member live ⇒ release. A steady
* count (still zero, still positive) is a no-op either way, so a single spawn or release only
* ever touches the OS on the crossing, not on every call.
*/
public void recheck() {
synchronized (lock) {
int live = liveCount.getAsInt();
if (live > 0 && held == null) {
held = mechanism.acquire();
if (held != null) {
log.debug("idle-sleep guard armed: {} live member(s)", live);
}
} else if (live == 0 && held != null) {
releaseHeldLocked();
}
}
}
/** {@code true} while an assertion is actually held. Exposed for tests. */
boolean isHeld() {
synchronized (lock) {
return held != null;
}
}
/**
* Release whatever is held, if anything. Idempotent and safe to call at any time, including
* repeatedly — a daemon shutdown hook calls this unconditionally so no assertion (and no
* {@code caffeinate} child) survives the process, even if the drain that would otherwise have
* driven the live count to zero was itself interrupted or threw.
*/
@Override
public void close() {
synchronized (lock) {
if (held != null) {
releaseHeldLocked();
}
}
}
/** Caller must hold {@link #lock}. */
private void releaseHeldLocked() {
try {
held.close();
} catch (RuntimeException e) {
log.warn("idle-sleep guard: failed to release its assertion cleanly: {}", e.toString());
} finally {
held = null;
}
}
}
@@ -0,0 +1,11 @@
package dev.ltms.fleet.power;
/**
* A held OS-level assertion against idle sleep. {@link #close} must be idempotent — safe to call
* more than once — and must never throw, matching {@link IdleSleepGuard}'s "never break the
* fleet" contract.
*/
public interface SleepAssertion extends AutoCloseable {
@Override
void close();
}
@@ -0,0 +1,21 @@
package dev.ltms.fleet.power;
/**
* The OS mechanism {@link IdleSleepGuard} uses to hold and release an idle-sleep assertion. This
* is the seam a test exercises instead of the real effect (a live {@code caffeinate} child) — see
* {@code IdleSleepGuardTest}.
*
* <p>Implementations must never throw. Every failure — wrong platform, missing tool, a spawn
* error — must show up as {@link #acquire()} returning {@code null}, so a caller can treat "no
* assertion held" and "the mechanism could not be used" identically and the fleet keeps running
* either way.
*/
public interface SleepAssertionMechanism {
/**
* Acquire a fresh assertion against idle sleep, or {@code null} when this mechanism is not
* usable right now (wrong platform, the tool is missing, the child process could not start).
* Never throws.
*/
SleepAssertion acquire();
}
@@ -94,6 +94,7 @@ public final class FleetApp {
// .none() (the honest "feature not wired" view) for every constructor that does not pass one.
private final FleetMcp.QuarantineSource quarantine;
private final FleetMcp.OutageSource outage;
private final FleetMcp.LoopHealthSource loopHealth;
private final ObjectMapper mapper = new ObjectMapper();
/**
@@ -155,7 +156,8 @@ public final class FleetApp {
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials) {
this(herdr, memberHerdr, workers, sessions, messages, presence, mcpServlet, auth, metrics,
deliverable, memberCredentials, FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none());
deliverable, memberCredentials, FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LoopHealthSource.none());
}
/**
@@ -169,8 +171,18 @@ public final class FleetApp {
public FleetApp(HerdrClient herdr, HerdrClient memberHerdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials,
FleetMcp.QuarantineSource quarantine, FleetMcp.OutageSource outage) {
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials,
FleetMcp.QuarantineSource quarantine, FleetMcp.OutageSource outage) {
this(herdr, memberHerdr, workers, sessions, messages, presence, mcpServlet, auth, metrics, deliverable,
memberCredentials, quarantine, outage, FleetMcp.LoopHealthSource.none());
}
public FleetApp(HerdrClient herdr, HerdrClient memberHerdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials,
FleetMcp.QuarantineSource quarantine, FleetMcp.OutageSource outage,
FleetMcp.LoopHealthSource loopHealth) {
this.herdr = herdr;
this.memberHerdr = memberHerdr != null ? memberHerdr : herdr;
this.workers = workers;
@@ -183,6 +195,7 @@ public final class FleetApp {
this.memberCredentials = memberCredentials != null ? memberCredentials : MemberCredentialPolicyView::absent;
this.quarantine = quarantine != null ? quarantine : FleetMcp.QuarantineSource.none();
this.outage = outage != null ? outage : FleetMcp.OutageSource.none();
this.loopHealth = loopHealth != null ? loopHealth : FleetMcp.LoopHealthSource.none();
}
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
@@ -291,15 +304,19 @@ public final class FleetApp {
* spotted by comparing two numbers by eye.
*/
private void healthz(Context ctx) {
HealthzResponse response = healthzResponse(herdr, memberHerdr, loopHealth);
ctx.status(response.status()).json(response.body());
}
record HealthzResponse(int status, Map<String, Object> body) { }
static HealthzResponse healthzResponse(HerdrClient herdr, HerdrClient memberHerdr,
FleetMcp.LoopHealthSource loopHealth) {
JsonNode pong;
try {
pong = herdr.call("ping");
} catch (HerdrException e) {
ctx.status(503).json(Map.of(
"status", "degraded",
"herdr", "unreachable",
"detail", e.getMessage()));
return;
return degradedResponse("unreachable", e.getMessage(), loopHealth);
}
Map<String, Object> body = new LinkedHashMap<>();
body.put("status", "ok");
@@ -311,11 +328,11 @@ public final class FleetApp {
try {
memberPong = memberHerdr.call("ping");
} catch (HerdrException e) {
ctx.status(503).json(Map.of(
return new HealthzResponse(503, Map.of(
"status", "degraded",
"herdr", "member unreachable",
"detail", e.getMessage()));
return;
"detail", e.getMessage(),
"loopHealth", loopHealthView(loopHealth)));
}
int leadProtocol = pong.path("protocol").asInt();
int memberProtocol = memberPong.path("protocol").asInt();
@@ -326,7 +343,20 @@ public final class FleetApp {
body.put("protocolMismatch", true);
}
}
ctx.status(200).json(body);
body.put("loopHealth", loopHealthView(loopHealth));
return new HealthzResponse(200, body);
}
private static Map<String, String> loopHealthView(FleetMcp.LoopHealthSource loopHealth) {
return Map.of("statusPoller", loopHealth.statusPoller().get().name(),
"sessionReaper", loopHealth.sessionReaper().get().name());
}
private static HealthzResponse degradedResponse(String herdr, String detail,
FleetMcp.LoopHealthSource loopHealth) {
return new HealthzResponse(503, Map.of(
"status", "degraded", "herdr", herdr, "detail", detail,
"loopHealth", loopHealthView(loopHealth)));
}
/**
@@ -640,15 +670,31 @@ public final class FleetApp {
}
default -> ctx.status(202).json(Map.of(
"sessionId", id,
// fleetd #571 (ticket comment 17126): no `default` here on purpose. This switch
// is an expression, so the compiler already demands every Outcome constant have
// an arm — adding an 11th constant to Outcome is a compile error here, not a
// silent fall-through. That is exactly the bug this ticket exists to fix:
// `default -> "done"` used to sit here and would have told a REST caller the
// delegation completed for TIMED_OUT_UNCONFIRMED, the one outcome where delivery
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION and STALE_TURN can never
// actually reach this inner switch — the outer switch above always dispatches
// them first — but they still need an arm to keep this switch exhaustive.
"status", switch (reply.outcome()) {
case TIMED_OUT_WORKING -> "working";
case TIMED_OUT_QUEUED -> "queued";
// Delivery here is unknown, not merely still queued — see
// Outcome#TIMED_OUT_UNCONFIRMED's own javadoc.
case TIMED_OUT_UNCONFIRMED -> "unconfirmed";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
default -> "done"; // unreachable (terminal outcomes handled above)
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN -> "done"; // unreachable
},
"detail", (reply.outcome() == MessageService.Outcome.WORKER_FAILED
"detail", reply.outcome() == MessageService.Outcome.TIMED_OUT_UNCONFIRMED
? "no reply within " + timeout + "ms; delivery is unconfirmed — the "
+ "message may already have reached the worker, so a resend "
+ "risks sending it twice; poll status first"
: (reply.outcome() == MessageService.Outcome.WORKER_FAILED
|| reply.outcome() == MessageService.Outcome.BACKEND_EXHAUSTED)
&& reply.text() != null
? reply.text()
@@ -696,6 +742,10 @@ public final class FleetApp {
/**
* The worker's structured reply ({@code fleet_reply}) — resolves the blocking send awaiting
* on this session, or queues the reply in the inbox when no send is open (CB-307).
*
* <p>fleetd #365: the response body's {@code delivered} field used to be unconditionally
* {@code true} for either case; it now reports whether a send/ticket was actually resolved,
* with {@code outcome} naming which (see {@link MessageService.ReplyOutcome}).
*/
private void replyMessage(Context ctx) {
String id = ctx.pathParam("id");
@@ -719,13 +769,19 @@ public final class FleetApp {
// a WRONG value instead of failing loudly. The check lives in MessageService.reply so both
// this door and FleetMcp.reply inherit the same rule; this catch only translates it into the
// {error, detail} envelope this file uses everywhere else.
MessageService.ReplyOutcome outcome;
try {
messages.reply(id, content);
outcome = messages.reply(id, content);
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", e.getMessage()));
return;
}
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
// fleetd #365: "delivered": true used to be unconditional here, whether the reply resolved
// a waiting send or was merely queued in the inbox for a later drain — the same gap
// FleetMcp.reply had over MCP. `delivered` now reflects which actually happened, and
// `outcome` names the specific case (see MessageService.ReplyOutcome).
ctx.status(200).json(Map.of("sessionId", id, "delivered", outcome.delivered(),
"outcome", outcome.wireName()));
}
/**
@@ -92,6 +92,15 @@ public final class GitWorktrees implements Worktrees {
private final String configuredRoot;
/** OS group name for {@link #shareWithGroup} (fleetd #185 stage 3); {@code null} ⇒ feature off. */
private final String group;
/** Source directory of skill folders for {@link #seedSkills} (fleetd #362, {@code memberSkills:}
* in config); {@code null} ⇒ feature off, no worktree is touched beyond today's behaviour. */
private final String memberSkillsSource;
/** Extra environment merged into every {@code git} subprocess this instance runs. Always {@code
* Map.of()} from every production constructor. Test seam only (fleetd #362 review fix): lets
* {@code GitWorktreesTest} point {@code GIT_CONFIG_GLOBAL} at an isolated temp file so it can
* drive the real {@link #add} path against a controlled "operator's global git config" and
* prove the excludesFile composition below without ever touching the real machine's config. */
private final Map<String, String> gitEnv;
private final Consumer<String> afterWorktreeAdded;
/** How the initial {@code git worktree add} command runs. Package-private test seam for an
* interrupted command after Git has made worktree state. */
@@ -125,7 +134,20 @@ public final class GitWorktrees implements Worktrees {
* config); null/blank ⇒ {@link #shareWithGroup} is a no-op.
*/
public GitWorktrees(String configuredRoot, String group) {
this(configuredRoot, group, _ -> {});
this(configuredRoot, group, (String) null);
}
/**
* @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of
* the repo root.
* @param group optional OS group name (fleetd #185 stage 3, {@code worktreeGroup:}
* in config); null/blank ⇒ {@link #shareWithGroup} is a no-op.
* @param memberSkillsSource fleetd #362: optional directory of skill folders ({@code
* memberSkills:} in config) copied into every provisioned worktree's
* {@code .claude/skills/}; null/blank ⇒ {@link #seedSkills} is a no-op.
*/
public GitWorktrees(String configuredRoot, String group, String memberSkillsSource) {
this(configuredRoot, group, _ -> {}, null, null, memberSkillsSource);
}
/** Test seam for changing a real worktree between its creation and its security check. */
@@ -135,13 +157,20 @@ public final class GitWorktrees implements Worktrees {
/** Test seam combining a configurable {@code group} with {@link #afterWorktreeAdded}. */
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded) {
this(configuredRoot, group, afterWorktreeAdded, null, null);
this(configuredRoot, group, afterWorktreeAdded, null, null, null);
}
/** Test seam for changing how {@link #shareWithGroup}'s processes run. */
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner) {
this(configuredRoot, group, afterWorktreeAdded, shareGroupRunner, null);
this(configuredRoot, group, afterWorktreeAdded, shareGroupRunner, null, null);
}
/** Test seam for changing how the initial {@code git worktree add} command runs, with no
* {@code memberSkillsSource} configured. */
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner, Function<String[], String> worktreeAddRunner) {
this(configuredRoot, group, afterWorktreeAdded, shareGroupRunner, worktreeAddRunner, null);
}
/**
@@ -151,11 +180,30 @@ public final class GitWorktrees implements Worktrees {
*
* @param shareGroupRunner {@code null} ⇒ the real {@link #exec(String...)}.
* @param worktreeAddRunner {@code null} ⇒ the real {@link #exec(String...)}.
* @param memberSkillsSource {@code null}/blank ⇒ {@link #seedSkills} is a no-op.
*/
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner, Function<String[], String> worktreeAddRunner) {
Function<String[], String> shareGroupRunner, Function<String[], String> worktreeAddRunner,
String memberSkillsSource) {
this(configuredRoot, group, afterWorktreeAdded, shareGroupRunner, worktreeAddRunner,
memberSkillsSource, Map.of());
}
/**
* Full test seam, plus {@code gitEnv} (fleetd #362 review fix, verification only): extra
* environment merged into every {@code git} subprocess this instance runs, so a test can isolate
* something like {@code GIT_CONFIG_GLOBAL} from the real machine while still driving the real
* {@link #add} path end to end. Every production constructor above delegates here with {@code
* Map.of()}.
*/
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner, Function<String[], String> worktreeAddRunner,
String memberSkillsSource, Map<String, String> gitEnv) {
this.configuredRoot = configuredRoot;
this.group = (group == null || group.isBlank()) ? null : group;
this.memberSkillsSource = (memberSkillsSource == null || memberSkillsSource.isBlank())
? null : memberSkillsSource;
this.gitEnv = gitEnv == null ? Map.of() : gitEnv;
this.afterWorktreeAdded = afterWorktreeAdded == null ? _ -> {} : afterWorktreeAdded;
this.shareGroupRunner = shareGroupRunner != null ? shareGroupRunner : this::exec;
this.worktreeAddRunner = worktreeAddRunner != null ? worktreeAddRunner : this::exec;
@@ -184,6 +232,7 @@ public final class GitWorktrees implements Worktrees {
configureEnvironmentCredentialHelper(repoRoot, wt);
configureHttpsUrlRewriteForSshOrigin(repoRoot, wt);
isolateToolSurface(wt);
seedSkills(wt);
} catch (RuntimeException e) {
cleanupAfterAddFailure(repoRoot, wt, branch, e);
throw e;
@@ -574,6 +623,322 @@ public final class GitWorktrees implements Worktrees {
return true;
}
/**
* fleetd #362: copy each skill folder from the configured {@link #memberSkillsSource} directory
* into {@code <worktreePath>/.claude/skills/}, so a member spawned against ANY repo — not only
* one that already ships its own {@code .claude/skills/} — can load a bridge skill such as
* {@code implementer}. Every brief this fleet sends starts with {@code "Load the <name>
* skill."}; outside a repo carrying its own copy that line was previously a no-op.
*
* <p><b>No-op — nothing read, nothing written, nothing logged</b> — when {@link
* #memberSkillsSource} is null/blank (today's default), the same off-switch shape as
* {@link #shareWithGroup}.
*
* <p><b>Invariant 1 — a repo's own skill wins.</b> A skill folder already present at
* {@code <worktreePath>/.claude/skills/<name>} — because the just-checked-out branch commits its
* own copy — is left completely untouched: never overwritten, and never even opened.
*
* <p><b>Invariant 2 — a seeded skill can never end up in a worker's commit.</b> Every path this
* writes is untracked in the target repo (that is the whole reason it is being seeded), so
* {@code git status} would otherwise show each one as a new, addable, committable path. The
* repo-wide {@code .git/info/exclude} is NOT used for this: measured against a real linked
* worktree, that file resolves to the repository's COMMON git dir even from a worktree (the
* same file {@link dev.ltms.fleet.member.ClaudeCodeLauncher#writeIdeOverlay writeIdeOverlay}
* appends {@code CLAUDE.local.md} to), so an entry written there would hide the seeded skill
* from {@code git status} in the PRIMARY's own checkout and every sibling worktree too — not
* only this one. Instead, {@link #excludeSeededSkillsFromGitStatus} points {@code
* core.excludesFile} at a file scoped {@code --worktree} (the same {@code
* extensions.worktreeConfig} mechanism {@link #configureEnvironmentCredentialHelper} already
* relies on) that itself lives under this worktree's own private git dir ({@code
* .git/worktrees/<nonce>/}, OUTSIDE the working tree) — invisible to this worktree's {@code git
* status} and structurally impossible for this worktree to commit, with no effect on any other
* worktree or the primary checkout. Proven with a real {@code git status --porcelain} in
* {@code GitWorktreesTest}, not by reasoning.
*
* <p><b>Compose, don't replace.</b> {@code core.excludesFile} is single-valued: the first cut of
* this method pointed it at fleetd's own file with {@code --replace-all}, which SHADOWS whatever
* the operator's own (global, or repo-local) {@code core.excludesFile} was already resolving to
* inside this worktree, rather than adding to it. Measured concretely: this repo's own {@code
* .gitignore} does not ignore {@code target/} — only an operator's global excludesFile does — so
* every worker's {@code mvn clean install} would otherwise make {@code target/} appear as
* untracked, and {@link #hasUncommitted}'s deliberately-untracked-inclusive {@code git status
* --porcelain} (CB-576) would then read every such worktree as dirty forever, so it is never
* cleaned up. {@link #excludeSeededSkillsFromGitStatus} now reads whatever {@code
* core.excludesFile} resolves to BEFORE writing anything (falling back to git's own documented
* default, {@code $XDG_CONFIG_HOME/git/ignore} or {@code $HOME/.config/git/ignore}, when the key
* is unset entirely — see {@code gitignore(5)}), and writes that content into its OWN exclude
* file ahead of the seeded skill patterns, so every operator-configured pattern keeps applying
* inside the seeded worktree exactly as it did before seeding ran.
*
* <p>Instead of using worktree-scoped-config as an add-then-append (a second key does not exist
* for {@code core.excludesFile} — it takes exactly one value), an actual second exclude source
* was ruled out because git resolves only ONE {@code core.excludesFile}; concatenating the prior
* content into fleetd's own file is what "compose" reduces to for a single-valued key.
*
* <p>Proven the same way as invariant 2's own leak check: {@code
* GitWorktreesTest#seedSkillsComposesWithAnAlreadyEffectiveGlobalExcludesFile} isolates a
* synthetic "operator's global config" via {@code GIT_CONFIG_GLOBAL} (never the real machine's),
* seeds a skill, and asserts {@code git status --porcelain} is still empty for a file matching
* that global config's own ignore pattern.
*
* <p><b>Invariant 3 — best-effort.</b> A missing/unreadable {@link #memberSkillsSource}, or a
* copy/exclude failure, is logged and skipped — it must never fail the spawn, the same contract
* {@link #overlayParity} and {@link #isolateToolSurface} already hold.
*
* <p>Recorded for the worker itself the same way fleetd #134 records {@code
* fleet.neutralizedConfig}: {@code fleet.seededSkills} (one value per seeded skill folder) and
* {@code fleet.seededSkillsNote}, readable with {@code git config --worktree --get-all
* fleet.seededSkills}.
*
* <p><b>Kind-blind by construction, not by a backend check here — this used to be a real gap
* (fleetd #393).</b> This method only copies files; like {@link #isolateToolSurface} — which
* neutralizes BOTH {@code .mcp.json} and {@code opencode.json} unconditionally — it runs the
* same for every worktree regardless of which backend ultimately spawns into it, because no
* caller of {@link #add} hands this class a kind to consult. Before fleetd #393, that made the
* log line below a false claim of success for a {@code kind: opencode} member: opencode has no
* built-in discovery of {@code .claude/skills/}, unlike the Claude Code CLI, so a seeded skill
* never reached one. It now does — {@code OpenCodeLauncher#skillInstructionFiles} reads
* whatever this method copied into {@code .claude/skills/} and appends each {@code SKILL.md} to
* the generated {@code instructions[]} — but that delivery, and the log line that honestly
* claims it (kind-aware, unlike this one), happens at the launcher, once the kind is actually
* known, not here.
*/
private void seedSkills(String worktreePath) {
if (memberSkillsSource == null) {
return;
}
Path source = Path.of(memberSkillsSource).toAbsolutePath().normalize();
if (!Files.isDirectory(source)) {
log.warn("memberSkills source '{}' is not a directory — skipping skill seeding for worktree {}",
source, worktreePath);
return;
}
Path skillsRoot = Path.of(worktreePath).resolve(".claude").resolve("skills");
List<String> seeded = new ArrayList<>();
List<String> kept = new ArrayList<>();
try (var candidates = Files.list(source)) {
for (Path candidate : candidates
.filter(Files::isDirectory)
.filter(p -> !p.getFileName().toString().startsWith("."))
.sorted()
.toList()) {
String name = candidate.getFileName().toString();
Path dst = skillsRoot.resolve(name);
if (Files.exists(dst)) {
kept.add(name);
continue;
}
copySkillDirectory(candidate, dst);
seeded.add(name);
}
} catch (IOException | RuntimeException e) {
log.warn("failed to seed skills into worktree {} from memberSkills source '{}': {}",
worktreePath, source, e.getMessage());
return;
}
String detail = seeded.isEmpty() ? "" : "seeded: " + String.join(", ", seeded);
if (!kept.isEmpty()) {
detail += (detail.isEmpty() ? "" : "; ") + "kept the repo's own copy of: " + String.join(", ", kept);
}
if (detail.isEmpty()) {
detail = "no skill folders found under " + source;
}
// fleetd #393: this only claims the copy step, deliberately — it cannot know the member
// kind that will spawn into this worktree (see this method's own javadoc), so it must not
// read as "and the member will act on it." Whether that is true depends on the kind: the
// Claude Code CLI discovers .claude/skills/ on its own; OpenCodeLauncher logs its own
// "skill delivery" line, once the kind is known, naming what it could and could not turn
// into instructions[].
log.info("skill seeding: {} of {} candidate(s) from {} into {}/.claude/skills — {} "
+ "(whether the spawned member can act on this depends on its kind — see "
+ "the launcher's own log for that)",
seeded.size(), seeded.size() + kept.size(), source, worktreePath, detail);
if (seeded.isEmpty()) {
return;
}
try {
excludeSeededSkillsFromGitStatus(worktreePath, seeded);
recordSeededSkillsForWorker(worktreePath, seeded);
} catch (RuntimeException e) {
log.warn("seeded skill(s) {} into {} but could not hide them from git status: {} — "
+ "they may show as untracked; never commit them", seeded, worktreePath, e.getMessage());
}
}
/** Recursively copy a skill folder ({@code src}) into a fresh destination ({@code dst}) that
* {@link #seedSkills} has already confirmed does not exist, preserving the directory structure
* (e.g. {@code implementer/SKILL.md}, {@code implementer/references/...}). */
private static void copySkillDirectory(Path src, Path dst) {
try (var walk = Files.walk(src)) {
for (Path path : walk.sorted().toList()) {
Path target = dst.resolve(src.relativize(path).toString());
if (Files.isDirectory(path)) {
Files.createDirectories(target);
} else {
Files.createDirectories(target.getParent());
Files.copy(path, target, StandardCopyOption.COPY_ATTRIBUTES);
}
}
} catch (IOException e) {
throw new WorktreeException("cannot copy skill directory " + src + " -> " + dst + ": "
+ e.getMessage(), e);
}
}
/**
* Make every path in {@code seededSkillNames} (each a name under {@code .claude/skills/})
* invisible to {@code git status} in THIS worktree only — see the invariant-2 discussion on
* {@link #seedSkills}. Sets {@code core.excludesFile} scoped {@code --worktree} to a file
* written under this worktree's own private git dir ({@code git rev-parse
* --absolute-git-dir}), which lives outside the working tree, so the exclude file itself can
* never be committed either.
*
* <p><b>Compose, don't replace.</b> {@code core.excludesFile} is single-valued, so pointing it at
* fleetd's own file would otherwise SHADOW whatever excludesFile this worktree was already
* resolving (an operator's global config, most commonly) rather than add to it — see the
* "Compose, don't replace" discussion on {@link #seedSkills}. {@link
* #previouslyEffectiveExcludesFileContent} is read BEFORE this method's own {@code --worktree}
* write below, so it still sees whatever was effective beforehand; that content is written into
* fleetd's own exclude file ahead of the seeded skill patterns, and the worktree-scoped override
* then points at that combined file — so every pattern the operator's own configuration already
* applied keeps applying, plus the seeded skill paths.
*
* <p><b>Assumes a fresh worktree — not idempotent.</b> {@link #seedSkills} only ever calls this
* from {@link #add}, which always creates a brand-new worktree, so {@code core.excludesFile} is
* never already worktree-scoped-set to fleetd's own file when this runs. A hypothetical second
* call on the SAME worktree would read fleetd's own already-composed file back as "previously
* effective" (worktree scope now wins) and append the seeded patterns a second time — harmless
* to {@code git status} (duplicate exclude lines are a no-op), but not something to rely on. No
* guard is added for this because the path does not exist today; if a future caller ever seeds
* the same worktree twice, it will need one.
*/
private void excludeSeededSkillsFromGitStatus(String worktreePath, List<String> seededSkillNames) {
exec("git", "-C", worktreePath, "config", "extensions.worktreeConfig", "true");
String previouslyEffective = previouslyEffectiveExcludesFileContent(worktreePath);
String gitDir = exec("git", "-C", worktreePath, "rev-parse", "--absolute-git-dir").trim();
Path excludeFile = Path.of(gitDir, "fleet-seeded-skills-exclude");
StringBuilder patterns = new StringBuilder();
if (!previouslyEffective.isEmpty()) {
patterns.append(previouslyEffective);
}
for (String name : seededSkillNames) {
patterns.append("/.claude/skills/").append(name).append('/').append(System.lineSeparator());
}
try {
Files.writeString(excludeFile, patterns.toString());
} catch (IOException e) {
throw new WorktreeException("cannot write skills exclude file " + excludeFile + ": "
+ e.getMessage(), e);
}
exec("git", "-C", worktreePath, "config", "--worktree", "--replace-all", "core.excludesFile",
excludeFile.toString());
}
/**
* The content of whatever {@code core.excludesFile} resolves to for {@code worktreePath} right
* now — BEFORE {@link #excludeSeededSkillsFromGitStatus} points that key at fleetd's own file —
* so it can be carried forward instead of shadowed. {@code --type=path} makes git itself perform
* {@code ~}/{@code ~user} expansion the same way it would when actually reading the key to build
* exclude rules, rather than handing back a raw, unexpanded config string.
*
* <p>When the key is unset entirely (exit code non-zero), falls back to git's own documented
* default excludes file — {@code $XDG_CONFIG_HOME/git/ignore}, or {@code
* $HOME/.config/git/ignore} when that variable is unset — per {@code gitignore(5)}: git applies
* that file even with no {@code core.excludesFile} configured at all, so skipping it here would
* silently drop patterns an operator never had to configure to get.
*
* <p>Never throws: a missing, unreadable, or unresolvable file is treated as "nothing to carry
* forward" (empty string) — this is a best-effort read in service of {@link #seedSkills}'s own
* invariant 3, not a new way for skill seeding to fail a spawn.
*
* <p><b>Review fix, finding 2.</b> The XDG-fallback branch below does not go through {@code git}
* at all, so a first cut of it read {@code XDG_CONFIG_HOME}/{@code HOME} straight from the JVM's
* own environment ({@link System#getenv} / {@code user.home}) — unlike every other value this
* class resolves, which goes through a {@code git} subprocess and therefore already honours
* {@link #gitEnv}. That meant no test could make this branch hermetic, and on any machine
* carrying a real {@code ~/.config/git/ignore} (this repo's own dev machine does), every
* skill-seeding test silently composed with that real file — correct in production, but
* machine-dependent in the test suite, and a future broader pattern in that real file could
* silently change what a seeded worktree's {@code git status} reports depending on whose home
* directory ran the test. {@link #resolveEnv} now checks {@link #gitEnv} first for both
* variables, falling back to the JVM's real environment only when the seam does not supply
* them — production behaviour (empty {@link #gitEnv}) is unchanged, and a test can now isolate
* this branch exactly as it already isolates every {@code git} subprocess call.
*
* <p><b>Snapshot, not a reference.</b> The content below is read once, at seeding time, and
* copied into fleetd's own exclude file. If the operator edits their global excludesFile
* afterward, an already-seeded worktree keeps the old copy — acceptable for a worktree's
* expected lifetime, but worth knowing before reading a stale pattern as a bug.
*
* @return the file's content, trailing-newline-normalized, or {@code ""} when there is nothing
* to compose with.
*/
private String previouslyEffectiveExcludesFileContent(String worktreePath) {
String resolvedPath;
if (exitCode("git", "-C", worktreePath, "config", "--get", "--type=path", "core.excludesFile") == 0) {
resolvedPath = exec("git", "-C", worktreePath, "config", "--get", "--type=path",
"core.excludesFile").trim();
} else {
String xdgConfigHome = resolveEnv("XDG_CONFIG_HOME");
Path fallback = (xdgConfigHome != null && !xdgConfigHome.isBlank())
? Path.of(xdgConfigHome, "git", "ignore")
: Path.of(resolveHome(), ".config", "git", "ignore");
resolvedPath = fallback.toString();
}
if (resolvedPath.isBlank()) {
return "";
}
Path file = Path.of(resolvedPath);
if (!Files.isRegularFile(file) || !Files.isReadable(file)) {
return "";
}
try {
String content = Files.readString(file);
return content.isBlank() ? "" : content.stripTrailing() + System.lineSeparator();
} catch (IOException e) {
log.warn("could not read previously-effective excludesFile {} while seeding skills into "
+ "{}: {} — its patterns will not carry forward into the seeded worktree",
file, worktreePath, e.getMessage());
return "";
}
}
/**
* Resolve environment variable {@code name} for {@link #previouslyEffectiveExcludesFileContent}'s
* XDG fallback, checking {@link #gitEnv} FIRST so a test can isolate this the same way it
* already isolates every {@code git} subprocess this class runs, and falling back to the JVM's
* real environment only when the seam does not supply it (always the case in production, where
* {@link #gitEnv} is {@code Map.of()}).
*/
private String resolveEnv(String name) {
String fromSeam = gitEnv.get(name);
return fromSeam != null ? fromSeam : System.getenv(name);
}
/** Same as {@link #resolveEnv(String)}, for {@code HOME} — falls back to {@code user.home}
* (rather than {@code System.getenv("HOME")}) when the seam does not supply it, matching this
* class's pre-existing behaviour for every other home-directory resolution. */
private String resolveHome() {
String fromSeam = gitEnv.get("HOME");
return fromSeam != null ? fromSeam : System.getProperty("user.home");
}
/**
* The worker-readable half of fleetd #362, mirroring {@link #recordNeutralizedConfigForWorker}:
* record which skill folders were seeded where the worker itself can read it, without a
* working-tree file that would show up in {@code git status}.
*/
private void recordSeededSkillsForWorker(String worktreePath, List<String> seeded) {
exec("git", "-C", worktreePath, "config", "extensions.worktreeConfig", "true");
for (String name : seeded) {
exec("git", "-C", worktreePath, "config", "--worktree", "--add", "fleet.seededSkills", name);
}
exec("git", "-C", worktreePath, "config", "--worktree", "fleet.seededSkillsNote",
"each fleet.seededSkills value names a skill folder fleetd copied into "
+ ".claude/skills/ because this repo did not already ship it; it is excluded "
+ "from git status (core.excludesFile, worktree-scoped) and must never be committed");
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
@@ -1084,6 +1449,9 @@ public final class GitWorktrees implements Worktrees {
Process p;
try {
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
if (!gitEnv.isEmpty()) {
pb.environment().putAll(gitEnv);
}
if (extraEnv != null && !extraEnv.isEmpty()) {
pb.environment().putAll(extraEnv);
}
@@ -1119,7 +1487,11 @@ public final class GitWorktrees implements Worktrees {
private int exitCode(String... command) {
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
if (!gitEnv.isEmpty()) {
pb.environment().putAll(gitEnv);
}
p = pb.start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
@@ -10,6 +10,7 @@ import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -301,9 +302,10 @@ public final class SessionManager implements TurnListener {
* with no copy and no error. Do NOT fuse these back together; the cost of an orphaned worktree
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
*/
private void release(String paneId, ReleaseCause cause) {
private MemberSession release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
releaseRemoved(paneId, removed, handles.remove(paneId), cause);
return removed;
}
/**
@@ -386,6 +388,31 @@ public final class SessionManager implements TurnListener {
// from the registry with no pane stop is an orphaned pane — a live terminal burning a fleet
// slot that no longer appears in the roster and can never be reclaimed.
launcher.stop(paneId);
if (removed != null && !preserveWorktree && removed.worktree() != null) {
// fleetd #316: the `dirty` read above ran while the worker could still write to this
// worktree, so a stale `false` must not be trusted to authorise the --force removal
// below. Re-read the worktree's state one more time, right here — immediately before
// the one step that would destroy it, and only on the path that is actually about to
// do that (invariant 4: no second unconditional `git status` on a release that already
// decided to preserve). By now `launcher.stop` has returned, so this read reflects
// whatever the worker managed to write up to and including its teardown, not whatever
// it had written at release-start time.
if (dirtyImmediatelyBeforeRemoval(removed)) {
preserveWorktree = true;
// The pre-stop snapshot above never ran for this session (the pre-stop read said
// clean), so this is the only chance to get the newly-discovered work into
// refs/wip/* rather than leaving the on-disk preserve as the sole copy. Best-effort,
// like every other snapshot attempt — trySnapshot logs and swallows its own failure.
String lateSnapshotRef = trySnapshot(removed, cause);
log.warn("release {} preserves worktree {} for pane={} terminal={}: it reported "
+ "clean before the pane stopped but dirty immediately before removal — the "
+ "worker wrote to it during teardown, and --force removing it now would "
+ "have destroyed that work{}",
cause, removed.worktree(), paneId, removed.terminalId(),
lateSnapshotRef == null ? "" : " (snapshotted to refs/wip/" + removed.branch()
+ " commit=" + lateSnapshotRef + ")");
}
}
if (removed != null && !preserveWorktree && removed.worktree() != null) {
// fleetd #283: this is the one cleanup step in this method that used to be bare. By the
// time it runs, the registry entry, the retained handle, and the pane are all already
@@ -404,6 +431,24 @@ public final class SessionManager implements TurnListener {
}
}
/**
* fleetd #316: the read that actually authorises {@code worktrees.remove}, taken with the
* worker's pane already stopped. Fails toward preserving (returns {@code true}) on any
* exception — the same rule the pre-stop check applies (CB-581): once we can no longer tell
* whether the worktree is dirty, preserving costs disk while deleting on a guess can destroy
* work that has no other copy.
*/
private boolean dirtyImmediatelyBeforeRemoval(MemberSession removed) {
try {
return worktrees.hasUncommitted(removed.worktree());
} catch (RuntimeException e) {
log.warn("release could not re-check worktree {} for pane={} terminal={} immediately "
+ "before removal; preserving it rather than risk destroying unsaved work: {}",
removed.worktree(), removed.paneId(), removed.terminalId(), e.toString());
return true;
}
}
/**
* Best-effort snapshot of a dirty worktree into {@code refs/wip/<branch>} (CB-578 stage C). A
* failure here must never escalate: the caller has already decided to preserve the worktree
@@ -541,8 +586,65 @@ public final class SessionManager implements TurnListener {
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId,
MemberLifecycle.SlotReservation reservation) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// fleetd #425 rework (round 2): resolved through launcher.place(memberRole) — the same
// candidate list, quarantine/cool-off/model-off filtering, and PlacementPolicy an unqualified
// spawn of this role is actually judged against right now — never launcher.defaultProfile()
// (MemberRole.DEV only, wrong for any other role) and never launcher.defaultProfileFor()
// (the role's pool FIRST entry, blind to quarantine/cool-off/model-off: a first-round fix
// used exactly this and regressed fleetd #429's "the fleet keeps working when a model is
// turned off" guarantee — a quarantined or model-off pool-first profile made this throw
// instead of routing around it, which an unqualified spawn is supposed to do). This same
// resolved decision is reused below for repoRoot, parityOverlay, AND the spawn itself so the
// worktree is always provisioned for the profile the member actually runs on — the two could
// disagree before fleetd #425: this name picked repoRoot/overlay, but the spawn below passed
// the ORIGINAL (blank) profile through to placement, which re-resolves live and can pick a
// different profile if the pool changed between the two reads, or a genuinely different one
// under weighted/round-robin placement.
//
// Round 1 of this rework fed the resolved name back into launcher.spawn(SpawnRequest) as an
// EXPLICIT profile. That was a mistake this round corrects, and the mistake is not that the
// two branches apply different checks — they are SUPPOSED to disagree: the routing branch a
// blank spawn takes treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
// profile as a reason to fall through to the next candidate, while CompositePeerLauncher's
// THROWING branch (enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/
// enforceModelEnabled) treats naming that same profile explicitly as a reason to refuse
// outright. That is correct: an operator who names a profile should get a refusal, not a
// silent substitution onto a different backend. The mistake was turning a fall-through into
// a refusal by accident — resolving a name via the routing side and then re-entering the
// refusing side with it, for a placement the routing side had already approved by walking
// past everything else.
//
// Before fleetd #435, this accident was reachable through maxLoad specifically: the default
// `fixed` placement policy did not evaluate maxLoad at all for automatic selection, so an
// at-cap pool-first profile that placement itself would have picked for a plain unqualified
// spawn could die at enforceMaxLoad one call later, purely because this method's route to
// the spawn passed through an explicit profile name — a failure a worktree-less unqualified
// spawn never hit. fleetd #435 closed that gap (`fixed` now evaluates maxLoad exactly like
// every other placement policy), so that specific failure can no longer happen — a
// PlacementDecision this method resolves can no longer be at-cap in the first place. What
// this round's fix still buys, now that maxLoad can no longer cause the accident: it keeps
// the PlacementDecision from place() and hands it to launcher.spawn(SpawnRequest,
// PlacementDecision) for an unqualified request, which spawns through the SAME routing
// branch a blank spawn uses — no enforce* check is newly applied, and the window between the
// placement decision and the spawn (in which the pool, a config reload, or another spawn
// landing on the same profile could otherwise move the state) never reopens. An
// explicitly-named profile still goes through launcher.spawn(SpawnRequest) and its throwing
// branch, unchanged — that caller asked for one profile by name and still gets everything
// enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/enforceModelEnabled decide about
// it, refusal included.
//
// The one cost that remains, unchanged from round 1: an unqualified worktree-provisioned
// spawn does not get CompositePeerLauncher's cross-candidate retry on a live
// PeerUnreachableException raised by the backend itself at spawn time (a transport-level
// failure placement cannot see in advance) — spawn(req, decision) commits to the one profile
// place() already chose, the same way an explicit-profile spawn commits to its one name. That
// trade is deliberate: a worktree provisioned for the wrong backend (the #425 hazard) is worse
// than a spawn that fails cleanly and can be retried by the caller. Nothing else is lost:
// maxLoad, quarantine, cool-off and model-off all behave identically whether or not a
// worktree was requested — that agreement is the invariant this rework exists to hold.
boolean unqualifiedProfile = profile == null || profile.isBlank();
PlacementDecision decision = unqualifiedProfile ? launcher.place(memberRole) : new PlacementDecision(profile);
String preResolvedProfile = decision.profile();
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
@@ -565,7 +667,15 @@ public final class SessionManager implements TurnListener {
// copies more files into the worktree after add() returns, so sharing the group any earlier
// leaves those overlay files operator-owned and read-only for a different-uid member.
worktrees.shareWithGroup(repoRoot, path);
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
// fleetd #425: preResolvedProfile, not the original (possibly blank) profile — see the
// comment above where it is resolved. The overlay/repoRoot above and the spawn here must
// name the same profile. An unqualified request stays unqualified here and is honored via
// the PlacementDecision already captured above (spawn(req, decision) — the routing branch,
// no enforce* re-check); an explicitly-named profile still goes through the single-arg
// spawn(req) and its throwing branch, exactly as before this rework.
SpawnRequest spawnReq = new SpawnRequest(unqualifiedProfile ? null : preResolvedProfile,
path, callerCwd, sessionName, resumeSessionId, memberRole);
handle = unqualifiedProfile ? launcher.spawn(spawnReq, decision) : launcher.spawn(spawnReq);
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
preResolvedProfile, memberRole, branch, path, e.getMessage());
@@ -593,7 +703,7 @@ public final class SessionManager implements TurnListener {
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String resolvedProfile = resolveProfile(handle, preResolvedProfile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
// CB-619: see the no-worktree path above — bind before recording, and store the returned
@@ -955,16 +1065,62 @@ public final class SessionManager implements TurnListener {
* drain (see above), and a straggler must not buy the drain more time than the flag it lost the
* race against would have. In the ordinary case the sweep finds nothing and costs one empty
* {@link #roster()} call.
*
* <p>fleetd #512: a drain that releases every session cleanly used to log nothing at all — the
* only log calls in this method and {@link #drainSnapshot} sit on abnormal paths, so "nothing
* logged" was indistinguishable from "died on the first session". The {@code log.info} at the
* end below is a positive assertion that the drain actually finished, on the normal path,
* every time — including the all-zero case, which is a common and legitimate outcome (no
* members were live) and must still produce the line. Both {@link #drainSnapshot} passes (the
* main snapshot and the straggler sweep) are folded into the one line: a caller reading two
* lines could not tell a two-pass drain from two separate drains.
*
* <p><strong>Non-goal: this line must never move into a {@code finally} block, and this method
* must never grow one around it.</strong> "Every time" above means every time the drain
* <em>finishes</em>, not every time this method exits. The absence of the line is the signal
* that the drain died, so a {@code finally} would destroy the signal and print confident
* partial counts in the same edit — the line would appear after a drain that threw, carrying
* whatever {@code tally} it had reached. Both halves of the value are lost at once. The line
* has to be the last statement of the successful path and reachable only from it.
*
* <p>This is written down because it is the obvious review comment ("shouldn't we always log
* the drain result?"), it sounds like thoroughness, and the paragraph above reads as an
* invitation to it. Raised by the fleet01 lead on 2026-09-12, from their 2026-09-10 incident:
* we only know that drain died on that host because it <em>threw</em>, and a
* {@code NoClassDefFoundError} reached the JVM's uncaught handler. A drain that hung on one
* session, or returned early on a condition rather than an exception, would leave no stack
* trace, no {@code ERROR} token and no priority — only a missing line. That makes the loud
* variant the one we have seen and the quiet variants the ones this line exists to catch.
*
* <p>Related: {@code released} and {@code abandoned} are counted incrementally inside {@link
* #drainSnapshot}'s loop and folded with {@link DrainTally#plus}, rather than derived from a
* collection read at the end, for the same reason. If a partial report is ever wanted it must
* be a different line with a different verb. One line must not serve both, or a reader cannot
* tell a finished drain from an interrupted one by its wording.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
draining.set(true);
drainSnapshot(roster(), deadline);
DrainTally tally = drainSnapshot(roster(), deadline);
List<MemberSession> stragglers = roster();
if (!stragglers.isEmpty()) {
log.warn("drain sweep found {} session(s) registered after the drain snapshot was "
+ "taken (raced past the shutdown guard); draining them too", stragglers.size());
drainSnapshot(stragglers, deadline);
tally = tally.plus(drainSnapshot(stragglers, deadline));
}
log.info("drain complete: released={} abandoned={} (still BUSY at the shutdown deadline)",
tally.released(), tally.abandoned());
}
/**
* Running count for one {@link #drainAll} invocation, folded across both {@link #drainSnapshot}
* passes (fleetd #512). {@code abandoned} counts sessions that were still {@code BUSY} at the
* moment they were released — i.e. the whole-drain deadline passed before they left {@code BUSY}
* on their own (see {@link #drainSnapshot}) — a subset of {@code released}, not additional to it.
*/
private record DrainTally(int released, int abandoned) {
private DrainTally plus(DrainTally other) {
return new DrainTally(released + other.released, abandoned + other.abandoned);
}
}
@@ -972,8 +1128,12 @@ public final class SessionManager implements TurnListener {
* Drain exactly the sessions in {@code snapshot}, waiting out a {@code BUSY} one against the
* shared whole-drain {@code deadline} before releasing it. Shared by {@link #drainAll}'s main
* pass and its post-loop straggler sweep (fleetd #308) so both honor the same one budget.
* Returns how many sessions this pass released, and how many of those were still {@code BUSY}
* (abandoned mid-turn) at the moment of release.
*/
private void drainSnapshot(List<MemberSession> snapshot, long deadline) {
private DrainTally drainSnapshot(List<MemberSession> snapshot, long deadline) {
int released = 0;
int abandoned = 0;
for (MemberSession s : snapshot) {
try {
if (s.state() == MemberSession.State.BUSY) {
@@ -991,11 +1151,16 @@ public final class SessionManager implements TurnListener {
}
}
}
release(s.paneId(), ReleaseCause.SHUTDOWN);
MemberSession removed = release(s.paneId(), ReleaseCause.SHUTDOWN);
released++;
if (removed != null && removed.state() == MemberSession.State.BUSY) {
abandoned++;
}
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
}
return new DrainTally(released, abandoned);
}
/**
@@ -1,9 +1,11 @@
package dev.ltms.fleet.session;
import dev.ltms.fleet.inject.LoopWatchdog;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
/**
* Periodic virtual-thread reaper that tears down {@code READY}/{@code DONE} sessions which have
@@ -14,6 +16,16 @@ public final class SessionReaper {
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
/**
* How many multiples of {@code intervalMillis} the last-completed round may age before
* {@link #health()} reports {@link LoopWatchdog.State#STALLED} (fleetd #544). One round here
* is {@code sessions.reapIdle} plus the (rare, best-effort) WIP sweep — both normally finish
* in a small fraction of one interval. At the default 5s interval this puts the threshold at
* 60s: generous enough that an occasional slow git call in the WIP sweep never trips it, short
* enough that a genuinely wedged reap round (mirroring the herdr-read-with-no-deadline hang
* {@code StatusPoller} guards against — see fleetd #544) is caught inside about a minute.
*/
private static final long STALE_AFTER_INTERVAL_MULTIPLIER = 12;
/** CB-586: the refs/wip age floor — never sweep a snapshot younger than 24h (the CB-586 rule). */
private static final long WIP_MIN_AGE_MILLIS = TimeUnit.HOURS.toMillis(24);
/**
@@ -25,6 +37,7 @@ public final class SessionReaper {
private final SessionManager sessions;
private final long idleTtlNanos;
private final long intervalMillis;
private final LoopWatchdog watchdog;
private volatile boolean running;
private Thread thread;
/**
@@ -45,29 +58,61 @@ public final class SessionReaper {
/** Construct a reaper with an explicit polling interval (useful for tests). */
public SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis) {
this(sessions, idleTtlSeconds, intervalMillis, System::nanoTime);
}
/** Full constructor — for tests: an injectable monotonic clock (fleetd #544). */
SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis, LongSupplier nowNanos) {
this.sessions = sessions;
this.idleTtlNanos = TimeUnit.SECONDS.toNanos(idleTtlSeconds);
this.intervalMillis = intervalMillis;
this.watchdog = new LoopWatchdog(nowNanos, staleAfterNanos(intervalMillis));
}
private static long staleAfterNanos(long intervalMillis) {
return TimeUnit.MILLISECONDS.toNanos(Math.max(intervalMillis, 1)) * STALE_AFTER_INTERVAL_MULTIPLIER;
}
/**
* This loop's progress fact (fleetd #544): {@code RUNNING}, {@code STALLED} (dead or parked —
* indistinguishable from outside, and never observed on purpose), or {@code STOPPED}
* ({@link #stop()} was called). See {@link LoopWatchdog} for why staleness — not thread
* liveness — is the signal.
*/
public LoopWatchdog.State health() {
return watchdog.state();
}
/** Start the reaper loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
watchdog.reset(); // fleetd #544: a fresh grace period, not last run's stale mark/timestamp.
thread = Thread.ofVirtual().name("session-reaper").start(this::loop);
log.info("session reaper started (idle ttl {}s, interval {}ms)",
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos), intervalMillis);
}
private void loop() {
while (running) {
try {
sessions.reapIdle(idleTtlNanos);
} catch (RuntimeException e) {
log.warn("session reaper iteration failed; continuing", e);
try {
while (running) {
try {
sessions.reapIdle(idleTtlNanos);
} catch (Throwable e) {
log.error("session reaper iteration failed; continuing", e);
}
maybeSweepWipRefs();
// fleetd #544: a round is "complete" only after reapIdle AND the WIP-sweep gate have
// both returned — a hang in either (e.g. a wedged git call) means this line is never
// reached, so the watchdog goes stale exactly like a dead loop would.
watchdog.recordRoundComplete();
sleep();
}
maybeSweepWipRefs();
sleep();
} finally {
if (running) {
log.error("session reaper loop exited unexpectedly; it can be restarted");
}
running = false;
}
}
@@ -90,8 +135,8 @@ public final class SessionReaper {
log.info("refs/wip retention sweep deleted {} snapshot ref(s) older than 24h whose "
+ "content was already reachable from main", deleted);
}
} catch (RuntimeException e) {
log.warn("refs/wip retention sweep failed; continuing", e);
} catch (Throwable e) {
log.error("refs/wip retention sweep failed; continuing", e);
}
// Set even when the sweep threw, so a broken repo is retried on the slow cadence rather
// than hammering git on every 5-second iteration.
@@ -111,6 +156,7 @@ public final class SessionReaper {
/** Stop the reaper loop. Idempotent. */
public synchronized void stop() {
running = false;
watchdog.markStoppedByCaller(); // fleetd #544: this halt is on purpose — health() must say STOPPED, not STALLED.
if (thread != null) thread.interrupt();
}
}
@@ -0,0 +1,150 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
* deliberately opt-in — an unset one leaves usage-limit detection silently OFF for that profile,
* and nothing quarantines its credential. {@link Fleetd#reportExhaustedPatternGap} must say so at
* startup, naming every unarmed profile, and must never fire when every profile is armed. Mirrors
* {@link MemberCredentialsGapReportTest}'s pattern, capturing the real log via a
* {@link ListAppender}.
*
* <p>The 8-profile shape in {@link #theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles} is
* the live {@code fleetd.yaml} shape measured 2026-09-10 (fleetd #395's own ticket): 6 of 8
* profiles unarmed, 2 of those 6 ({@code opus}, {@code sonnet}) running on the operator's Claude
* subscription. {@code fleetd.yaml} itself is gitignored and unavailable to this test, so the
* shape is reproduced as a throwaway config in a {@code @TempDir} rather than read off disk.
*/
class ExhaustedPatternGapReportTest {
private static FleetConfig load(Path dir, String yaml) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml);
return FleetConfig.load(f);
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
}
/** Every profile name mentioned by a WARN-level log line, across every WARN this call produced. */
private static List<String> warnMessages(ListAppender<ILoggingEvent> appender) {
return appender.list.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.toList();
}
@Test
void theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles(@TempDir Path dir) throws Exception {
// Reproduces the live shape measured 2026-09-10: 8 profiles, 2 armed (sol, terra), 6
// unarmed (local, local-direct, gx, opus, sonnet, xf) — 2 of the unarmed 6 (opus, sonnet)
// are subscription: true.
FleetConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
local-direct:
baseUrl: http://gx01.gw:8000
gx:
kind: opencode
baseUrl: https://llm.ltms.dev/v1
opus:
subscription: true
model: claude-opus-5
sonnet:
subscription: true
model: claude-sonnet-5
sol:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
terra:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
xf:
baseUrl: https://llm.ltms.dev/v1
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportExhaustedPatternGap(cfg);
} finally {
detach(appender);
}
List<String> warns = warnMessages(appender);
assertFalse(warns.isEmpty(), "6 of 8 profiles are unarmed — at least one WARN must fire");
// Exactly one WARN aggregates every unarmed profile, naming all 6 and none of the 2 armed.
String aggregate = warns.stream()
.filter(m -> m.contains("no exhaustedPattern configured"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected an aggregate unarmed-profiles WARN: " + warns));
for (String unarmed : List.of("local", "local-direct", "gx", "opus", "sonnet", "xf")) {
assertTrue(aggregate.contains(unarmed), "aggregate WARN must name '" + unarmed + "': " + aggregate);
}
for (String armed : List.of("sol", "terra")) {
assertFalse(aggregate.contains(armed), "aggregate WARN must NOT name armed profile '" + armed + "': " + aggregate);
}
// A second, louder WARN calls out the subscription profiles specifically.
String subscriptionWarn = warns.stream()
.filter(m -> m.contains("metered Claude plan"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a subscription-specific WARN: " + warns));
assertTrue(subscriptionWarn.contains("opus"), subscriptionWarn);
assertTrue(subscriptionWarn.contains("sonnet"), subscriptionWarn);
assertFalse(subscriptionWarn.contains("local-direct"),
"the subscription WARN must not name a non-subscription profile: " + subscriptionWarn);
}
@Test
void everyProfileArmedProducesNoWarningAtAll(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
sol:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
terra:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
opus:
subscription: true
model: claude-opus-5
exhaustedPattern: "5-hour limit reached"
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportExhaustedPatternGap(cfg);
} finally {
detach(appender);
}
assertTrue(warnMessages(appender).isEmpty(),
"every profile is armed — a checker that warns anyway always fires: " + warnMessages(appender));
}
}
@@ -0,0 +1,192 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import com.fasterxml.jackson.databind.JsonNode;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.fail;
/**
* fleetd #498: {@code Fleetd.awaitHerdr} used to return a bare {@code boolean}, collapsing "the
* configured wait budget genuinely ran out" and "the waiting thread was interrupted, possibly
* milliseconds in" onto the same {@code false} — and the caller's log line printed only the
* configured budget, never how long the wait actually ran. This class covers both halves of the
* fix:
* <ul>
* <li>the seam — {@link Fleetd#awaitHerdr} itself, driven with an injected clock and a stub
* {@link HerdrClient}, one test per {@link Fleetd.HerdrWaitResult};</li>
* <li>the call site — {@link Fleetd#logHerdrWaitOutcomeAndShouldReap}, the exact decision {@code
* main} calls (extracted here because {@code main} itself boots the whole daemon and cannot
* be driven from a unit test), pinning the three distinct log messages it emits.</li>
* </ul>
* Every expected message below is a plain literal, not built from {@code HERDR_WAIT_SECONDS} or
* any other production constant — a test that derives its expectation the way the code does
* cannot see a change to either (fleetd #496's identical trap).
*/
class FleetdAwaitHerdrTest {
// ---- the seam: Fleetd.awaitHerdr ----------------------------------------------------------
@Test
void answeredReturnsImmediatelyWithZeroElapsedAndNeverPolls() {
HerdrStub herdr = new HerdrStub(0); // succeeds on the very first call
LongSupplier clock = fixedClock(1_000L);
AtomicBoolean polled = new AtomicBoolean(false);
Runnable poller = () -> polled.set(true);
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.ANSWERED, outcome.result());
assertEquals(0L, outcome.elapsedNanos(), "a fixed clock must measure zero elapsed time");
assertFalse(polled.get(), "herdr answering on the first try must never poll");
}
@Test
void deadlinePassedIsMeasuredNotAssumed() {
HerdrStub herdr = new HerdrStub(-1); // never succeeds
// call order inside awaitHerdr: start, then per failed attempt: deadline-check, elapsed-calc
ScriptedClock clock = new ScriptedClock(0L, 30_500_000_000L, 30_500_000_000L);
Runnable poller = () -> fail("the deadline was already exceeded on the first attempt — must not poll");
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.DEADLINE_PASSED, outcome.result());
assertEquals(30_500_000_000L, outcome.elapsedNanos(),
"elapsed must be the MEASURED clock delta, not the configured budget");
}
@Test
void interruptedIsDistinctFromDeadlinePassedAndPreservesTheInterruptFlag() {
HerdrStub herdr = new HerdrStub(-1); // never succeeds
// start=0, deadline-check returns 500ms (well under the 30s budget) -> not deadline-passed,
// then the poller interrupts, and the elapsed-calc call returns 750ms.
ScriptedClock clock = new ScriptedClock(0L, 500_000_000L, 750_000_000L);
Runnable poller = () -> Thread.currentThread().interrupt();
try {
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.INTERRUPTED, outcome.result());
assertEquals(750_000_000L, outcome.elapsedNanos(),
"elapsed must be measured even when the wait ends via interruption, not the deadline");
assertTrue(Thread.currentThread().isInterrupted(),
"the interrupt flag the old code re-set must still be set on return");
} finally {
Thread.interrupted(); // clear it so it cannot leak into another test on this thread
}
}
// ---- the call site: Fleetd.logHerdrWaitOutcomeAndShouldReap -------------------------------
@Test
void answeredLogsNothingAndSaysReap() {
try (CapturedLog log = attach()) {
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.ANSWERED, 0L));
assertTrue(shouldReap, "only ANSWERED should tell main to reap orphan workers");
assertEquals(0, log.events().size(), "the answered path logs nothing itself");
}
}
@Test
void deadlinePassedLogsConfiguredAndMeasuredElapsedTogether() {
try (CapturedLog log = attach()) {
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.DEADLINE_PASSED, 30_500_000_000L));
assertFalse(shouldReap, "a deadline-passed wait must not tell main to reap");
assertEquals(1, log.events().size());
ILoggingEvent event = log.events().getFirst();
assertEquals(Level.WARN, event.getLevel());
assertEquals("herdr did not answer within the configured wait (configured=30s "
+ "elapsed=30500ms) — starting anyway; /healthz will report degraded until it "
+ "comes up. Orphaned worker panes (if any) were NOT reaped.",
event.getFormattedMessage());
}
}
@Test
void interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed() {
try (CapturedLog log = attach()) {
// 3ms: the ticket's own example of "a few milliseconds in", not the 30s budget.
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.INTERRUPTED, 3_000_000L));
assertFalse(shouldReap, "an interrupted wait must not tell main to reap");
assertEquals(1, log.events().size());
ILoggingEvent event = log.events().getFirst();
assertEquals(Level.WARN, event.getLevel());
String message = event.getFormattedMessage();
assertEquals("herdr wait was interrupted before the configured wait ran out "
+ "(configured=30s elapsed=3ms) — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
message);
assertFalse(message.contains("did not answer"),
"an interrupted wait must not be reported as if herdr failed to answer within the budget");
}
}
// ---- fixtures --------------------------------------------------------------------------
/** Always returns the same value, i.e. a clock that measures zero elapsed time. */
private static LongSupplier fixedClock(long value) {
return () -> value;
}
/** Returns each value in order, then repeats the last one for any call beyond the list. */
private static final class ScriptedClock implements LongSupplier {
private final long[] values;
private int index;
ScriptedClock(long... values) {
this.values = values;
}
@Override
public long getAsLong() {
long v = values[Math.min(index, values.length - 1)];
if (index < values.length - 1) {
index++;
}
return v;
}
}
/** Fails {@code failuresBeforeSuccess} times, then succeeds forever; {@code -1} never succeeds. */
private static final class HerdrStub implements HerdrClient {
private final int failuresBeforeSuccess;
private int calls;
HerdrStub(int failuresBeforeSuccess) {
this.failuresBeforeSuccess = failuresBeforeSuccess;
}
@Override
public JsonNode call(String method, Object params) throws HerdrException {
calls++;
if (failuresBeforeSuccess < 0 || calls <= failuresBeforeSuccess) {
throw new HerdrException("herdr not up yet");
}
return null;
}
@Override
public void close() {
}
}
private static CapturedLog attach() {
return CapturedLog.at(Fleetd.class, Level.DEBUG);
}
}
@@ -20,6 +20,7 @@ import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementDecision;
import dev.ltms.fleet.placement.PlacementPolicies;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
@@ -123,6 +124,11 @@ class FleetdBackendErrorSinkTest {
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
}
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
}
@Override
public Set<String> profiles() {
return Set.of();
@@ -0,0 +1,65 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
* it says nothing about which one {@code main} actually calls.
*
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
* every one of them builds its own instance directly. That silent regression is exactly the shape
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
* factory choice, following their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
* constructor. It does not prove that call actually executes at startup (no test here starts the
* daemon), and it does not prove the escalation reaches a real backend or credential — only
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
* the wiring runs.
*/
class FleetdBackendQuarantineWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"Fleetd.main's BackendQuarantine local must still be built from "
+ "BackendQuarantine.withEscalation(System::nanoTime, "
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
+ "proves the escalation itself works — this source check is what must go red instead. "
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
+ "flat ~30-minute cooldown, about 336 times across the week.");
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
assertFalse(source.contains(
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
}
}
@@ -0,0 +1,155 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #416: {@code fleet_list}'s {@code CapacitySource.configuredProfiles} must enumerate the
* <em>startup</em> profile set, not the live, hot-reloaded one.
*
* <p>{@code profiles} as a whole is a {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}):
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} once at construction, so a profile
* only added to the hot-reloaded map can never actually be spawned. Before this fix, {@code
* Fleetd.main} built {@code CapacitySource} with {@code () -> config.get().profiles().keySet()} —
* the live map — so {@code fleet_list} would report a freshly hot-reloaded profile as available
* ({@code free > 0}) while {@code fleet_spawn} on that same profile failed with
* {@code unknown worker profile}. Measured on another host: adding a throwaway profile and letting
* it hot-reload gave {@code fleet_list} -> {@code free: 3} and {@code fleet_spawn} ->
* {@code error: unknown worker profile}.
*
* <p>This test needs a reload, the same reason {@link FleetdExhaustionDetectionArmedWiringTest}
* does: at startup the two snapshots agree, so a test of only a newly started daemon would not
* detect a live {@code config.get()} lookup for the set.
*
* <p><b>maxLoad must stay hot.</b> It is read off {@code config.get()} exactly like
* {@code credentialId} ({@link ConfigRef} documents both as hot, "read live off the config
* supplier ... exactly like weight/maxLoad"), so a reload that only changes an existing profile's
* {@code maxLoad} — no add/remove — must still change what {@code fleet_list} reports without a
* restart. A fix that freezes the whole {@code CapacitySource} against {@code cfg} (rather than
* only its {@code configuredProfiles} set) would trade this bug for its mirror image and is pinned
* wrong by {@link #reloadedMaxLoadStillChangesWhatFleetListReports}.
*/
class FleetdCapacitySourceWiringTest {
private static final String STARTUP = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 3
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_NEW_PROFILE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 3
ghost404:
baseUrl: http://gx00.gw:8001
model: ghost404
maxLoad: 3
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_CHANGED_MAX_LOAD = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 9
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
@DisplayName("a profile present only in the live (hot-reloaded) config is NOT listed")
void liveOnlyProfileIsNotListed(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
Files.writeString(file, WITH_NEW_PROFILE);
assertTrue(config.reload().applied());
// The live snapshot now has the new profile — proves the reload really happened and this
// test is not accidentally passing because nothing changed.
assertTrue(config.get().profiles().containsKey("ghost404"));
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
assertFalse(source.configuredProfiles().get().contains("ghost404"),
"a profile added only to the hot-reloaded config must not be listed by fleet_list — "
+ "HerdrPeerLauncher never learns about it until a restart, so fleet_spawn on it "
+ "would fail with 'unknown worker profile' while fleet_list claimed it free");
}
@Test
@DisplayName("a profile present in the startup set IS listed")
void startupProfileIsListed(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
// fleetd #416, both-directions requirement: a test that only ever passes an empty/absent
// startup set (the case above) cannot tell a correct lookup from one that is permanently
// empty (e.g. a mutation replacing the supplier with Set::of). This is the direction that
// fails if the fix regresses to reporting nothing at all.
assertTrue(source.configuredProfiles().get().contains("terra"),
"a profile present in the startup snapshot must still be listed by fleet_list");
}
@Test
@DisplayName("a hot maxLoad edit still changes what fleet_list reports")
void reloadedMaxLoadStillChangesWhatFleetListReports(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
assertEquals(3, source.maxLoad().apply("terra"),
"sanity: maxLoad reads 3 from the startup config before any reload");
Files.writeString(file, WITH_CHANGED_MAX_LOAD);
assertTrue(config.reload().applied());
assertEquals(9, source.maxLoad().apply("terra"),
"maxLoad must stay hot — the SAME CapacitySource instance must reflect a reloaded "
+ "maxLoad without a restart, exactly like credentialId. Freezing the whole "
+ "CapacitySource against the startup snapshot (rather than only its "
+ "configuredProfiles set) would trade fleetd #416 for its mirror image.");
}
}
@@ -0,0 +1,106 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #474: proves the exact wiring {@code Fleetd.main} uses to construct its live {@code
* ConfigRef} — {@code new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools)}
* — actually makes {@link ConfigRef#reload()} refuse a charter that names an MCP tool the server
* does not register, the same way {@code Fleetd.main} itself refuses one at startup (see {@code
* FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool}).
*
* <p>{@code dev.ltms.fleet.config.ConfigRefTest} pins the same behaviour through a locally-built
* {@code Consumer<FleetConfig>} adapter that calls the same production {@code CharterToolSurface}
* method, because that test lives in {@code dev.ltms.fleet.config} and cannot see {@code
* Fleetd#assertChartersNameOnlyRegisteredTools} (package-private to {@code dev.ltms.fleet}). This
* class is the companion proof that lives where the real method reference is visible, so the literal
* expression {@code Fleetd::assertChartersNameOnlyRegisteredTools} — not just an equivalent — is
* what gets exercised. {@code Fleetd.main} itself cannot be driven this far in a unit test: every
* fixture in {@code FleetdStartupValidationTest} is deliberately invalid so {@code main} throws
* before opening a socket, binding Javalin, or doing anything else with a real side effect, so a
* test cannot get {@code main} far enough to hold a running daemon it could then reload — this test
* builds the {@code ConfigRef} the same way {@code main} does and drives {@link ConfigRef#reload()}
* directly instead, the same shape {@code FleetdExhaustionDetectionArmedWiringTest} and its
* siblings already use for the rest of {@code Fleetd.main}'s wiring.
*/
class FleetdConfigRefCharterToolSurfaceWiringTest {
private static final String BASE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
void reloadRefusesACharterNamingAnUnregisteredToolThroughFleetdsOwnWiring(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
""");
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
Fleetd::assertChartersNameOnlyRegisteredTools);
FleetConfig before = config.get();
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
""");
ConfigRef.Outcome out = config.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.error().contains("dev"), out.error());
assertTrue(out.error().contains("bridge_send"), out.error());
assertSame(before, config.get());
}
@Test
void reloadAcceptsACharterNamingOnlyRegisteredToolsThroughFleetdsOwnWiring(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, BASE + """
fleet:
charters:
dev: old charter
""");
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
Fleetd::assertChartersNameOnlyRegisteredTools);
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
""");
ConfigRef.Outcome out = config.reload();
assertTrue(out.applied());
assertEquals("config reloaded", out.summary());
}
}
@@ -0,0 +1,80 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #474 follow-up: {@code Fleetd.main} builds its live {@code ConfigRef} from the
* three-argument constructor, {@code new ConfigRef(configPath, cfg,
* Fleetd::assertChartersNameOnlyRegisteredTools)}, so a reload runs the same charter tool-surface
* gate startup does (see {@link dev.ltms.fleet.config.ConfigRef}'s class doc, "fleetd #474" bullet).
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
* three-argument constructor and the {@code Fleetd.assertChartersNameOnlyRegisteredTools} adapter
* work correctly together — both build their OWN {@code ConfigRef} with that constructor, so neither
* proves {@code main} still CHOOSES the three-argument form over the plain two-argument {@code new
* ConfigRef(configPath, cfg)}.
*
* <p>Measured directly: reverting {@code Fleetd.java}'s {@code config} local to the two-argument
* constructor compiles with 0 errors and leaves the entire 1633-test suite green — including every
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
* instance directly, wired with the check by hand. That silent regression is exactly the shape
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
* their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* ConfigRef} and never runs {@code main} — a green result here proves only that the exact text
* {@code main} contains is the three-argument construction with {@code
* Fleetd::assertChartersNameOnlyRegisteredTools}. It does not prove that call actually executes at
* startup (no test here starts the daemon), and it does not prove the reload gate itself works —
* only {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
* behaviour; only a live daemon proves the wiring runs.
*/
class FleetdConfigRefWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main still builds config from the three-argument ConfigRef constructor")
void mainStillWiresTheThreeArgumentConfigRefConstructor() throws Exception {
String source = fleetdSource();
// A broken read (wrong working directory, wrong path, a file that came back empty) would
// make the assertFalse below pass vacuously — a "clean" negative check that actually
// checked nothing. Guard against that first, with an anchor that has nothing to do with
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
assertTrue(source.contains("public final class Fleetd"),
"fleetdSource() did not read anything usable — src/main/java/dev/ltms/fleet/Fleetd.java "
+ "did not come back containing its own class declaration. The assertFalse below "
+ "would pass vacuously on a broken read; fix the read before trusting this test.");
assertTrue(source.contains(
"ConfigRef config = new ConfigRef(configPath, cfg, "
+ "Fleetd::assertChartersNameOnlyRegisteredTools);"),
"Fleetd.main's config local must still be built from the three-argument ConfigRef "
+ "constructor, with Fleetd::assertChartersNameOnlyRegisteredTools as "
+ "extraValidation. Reverting to the plain two-argument constructor (fleetd #474's "
+ "measured M2 regression) compiles with 0 errors and leaves the whole suite green — "
+ "including ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest, because "
+ "neither builds its ConfigRef through main — this source check is what must go "
+ "red instead. A reverted daemon would accept, through a reload with no restart, "
+ "exactly the charter that #469/#474 already refuse at startup.");
// Negative form of the same check: the pre-#474 two-argument call, if it ever reappears at
// this declaration, must not be mistaken for the three-argument one by a looser
// positive-only check — this is the M2 mutation this test exists to kill.
assertFalse(source.contains("ConfigRef config = new ConfigRef(configPath, cfg);"),
"main's config local must never regress to the plain two-argument ConfigRef "
+ "constructor — that drops the reload-path charter check (fleetd #474's measured "
+ "M2 mutation) with no other test catching it");
}
}
@@ -0,0 +1,137 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.placement.BackendQuarantine;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #446: {@code exhaustionDetectionArmed} must describe the LIVE config, not a startup
* snapshot — the opposite of what fleetd #404 asked for. #404 pinned {@code exhaustedPattern} as
* compiled once at startup (matching {@code CompletionResolver}'s then-frozen behaviour); #446's
* whole point is that both sides — the classification {@code CompletionResolver} enforces and the
* {@code exhaustionDetectionArmed} field this reports — now read the SAME live {@link
* LiveExhaustedPatterns} instance, so a reload arms or disarms detection with no restart, and the
* report can never disagree with the behaviour (the fleetd #404 rule, now upheld for real instead
* of by freezing both sides).
*
* <p>This test needs a reload. At startup the two snapshots agree, so a test of only a newly
* started daemon would not detect a live {@code config.get()} lookup in the report field.
*
* <p><b>The hotness proof criterion 1 demands:</b> run this against the fixed code (green) and
* against the pre-#446 compile site — {@code Fleetd.main}'s {@code exhaustedPatternsByProfile}
* compiled once into a {@code Map<String, Pattern>} at startup, handed to {@code quarantineSource}
* — and show it fails there. See the class doc history above: that shape is exactly what
* {@code reloadedPatternArmsDetectionWithNoRestart} below is written to catch, and it is the test
* that used to assert the opposite (see git history for this file's pre-#446 version, which
* asserted {@code assertFalse(...)} on the identical scenario this now asserts {@code assertTrue}
* on).
*/
class FleetdExhaustionDetectionArmedWiringTest {
private static final String NO_PATTERN = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_PATTERN = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
exhaustedPattern: "usage limit"
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
@DisplayName("reloading a profile's exhaustedPattern IN arms exhaustionDetectionArmed, no restart")
void reloadedPatternArmsDetectionWithNoRestart(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, NO_PATTERN);
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
FleetMcp.QuarantineSource before = Fleetd.quarantineSource(config, BackendQuarantine.none(),
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
assertFalse(before.exhaustedPatternArmed().apply("terra"),
"no exhaustedPattern configured yet — must report unarmed");
Files.writeString(file, WITH_PATTERN);
assertTrue(config.reload().applied());
assertTrue(config.get().profiles().get("terra").hasExhaustedPattern());
// Re-read exhaustedPatternArmed off the SAME QuarantineSource built BEFORE the reload — no
// new object, no new wiring — to prove the field itself is live, not merely that a freshly
// built source would be. This is the exact axis the pre-#446 code fails on: the same
// Function<String,Boolean> lambda, called again after a reload, sees the new answer only if
// it re-reads config.get() on every call rather than a value captured earlier.
assertTrue(before.exhaustedPatternArmed().apply("terra"),
"exhaustionDetectionArmed must read the LIVE config on every call — fleetd #446 made "
+ "this hot; it must reflect a reload with no restart and no new QuarantineSource");
}
@Test
@DisplayName("reloading exhaustedPattern OUT disarms exhaustionDetectionArmed, no restart")
void reloadedPatternRemovalDisarmsDetectionWithNoRestart(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, WITH_PATTERN);
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
assertTrue(source.exhaustedPatternArmed().apply("terra"),
"exhaustedPattern is configured from the start — must report armed");
Files.writeString(file, NO_PATTERN);
assertTrue(config.reload().applied());
assertFalse(source.exhaustedPatternArmed().apply("terra"),
"removing exhaustedPattern and reloading must disarm detection with no restart — an "
+ "operator turning detection off (e.g. while debugging a false positive) "
+ "must see that reflected immediately, exactly like arming it is");
}
@Test
@DisplayName("a profile with no exhaustedPattern configured is reported unarmed")
void aProfileWithNoPatternIsUnarmed(@TempDir Path dir) throws Exception {
// fleetd #404's second direction, carried forward: this test alone fails if
// exhaustedPatternArmed degenerates to a constant `profile -> true` — it needs at least one
// profile that is genuinely unarmed to catch that.
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, WITH_PATTERN);
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
assertTrue(source.exhaustedPatternArmed().apply("terra"),
"a profile whose exhaustedPattern is configured must report armed");
assertFalse(source.exhaustedPatternArmed().apply("sonnet"),
"a profile absent from config entirely must not report armed");
}
}
@@ -0,0 +1,175 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.session.SessionManager;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #446 follow-up (round 3): round 2's {@link FleetdUsageLimitFixWarningTest} calls {@link
* Fleetd#usageLimitFixWarning}/{@link Fleetd#usageLimitFixWarningNoModel} directly, which pins
* the two methods' TEXT but cannot see whether {@link Fleetd#exhaustionSink}'s {@code log.warn}
* call actually invokes either one, or invokes the right one. A mutation battery run against
* round 2's merge (1592 green tests) proved both gaps real:
*
* <ul>
* <li><b>Cell A</b> — replacing the whole ternary result with a literal string
* ({@code "usage-limit fix: MUTANT"}) left every test green;</li>
* <li><b>Cell B</b> — swapping the ternary's two branches, so a profile WITH a {@code model:}
* gets the no-model fallback message and vice versa, was measured unmeasured by the lead
* but the existing test file shows no call that would notice it either.</li>
* </ul>
*
* <p>This class drives {@link Fleetd#exhaustionSink} — the actual factory {@code main} wires,
* since round 3's extraction pulled it out of the inline lambda for exactly this reason — and
* asserts on the REAL text a {@link ListAppender} attached to {@code Fleetd}'s own logger
* captures, mirroring {@code ExhaustedPatternGapReportTest}'s idiom. Chosen over a source-text
* assertion (the {@code FleetMcpAuthzTest} {@code everyRegisteredToolHasItsHandlerActionPinned}
* idiom) because that route can prove Cell A (the ternary is still there, still built from the
* two named methods) but not Cell B (which branch a given profile actually reaches) — this route
* answers both from one mechanism, since it reads what the sink actually logged for each shape of
* profile.
*
* <p>Each test filters for the WARN line starting {@code "usage-limit fix:"} specifically — {@link
* Fleetd#exhaustionSink} also logs a separate {@code "credential '...' quarantined for..."} WARN
* on every call, and asserting against the wrong one would pass or fail for the wrong reason. That
* prefix survives even under the M5 mutation above (the mutant keeps the tag, only the body
* becomes {@code "MUTANT"}), so the filter itself is not what either mutation defeats.
*/
class FleetdExhaustionSinkWarningTest {
private static final String YAML = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: claude-opus-5
gx:
baseUrl: http://gx00.gw:8000
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static SessionManager emptyRosterSessions() {
FakeHerdr h = new FakeHerdr();
FleetConfig.Profile dummy = new FleetConfig.Profile(
"dummy", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(dummy.profile(), dummy), dummy.profile(), _ -> "tok");
// Never acquires a session — exhaustionSink's roster lookup is expected to miss and fall
// back to the profileHint argument, exactly like OpenCodeLauncher's real call site does
// (fleetd #234) — so the sink under test never needs a populated roster.
return new SessionManager(launcher);
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
}
/** The one WARN line {@code exhaustionSink} builds from the ternary under test, or {@code null}. */
private static String fixWarning(ListAppender<ILoggingEvent> appender) {
List<String> warns = appender.list.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.startsWith("usage-limit fix:"))
.toList();
assertTrue(warns.size() == 1,
"expected exactly one 'usage-limit fix:' WARN per onExhausted call, got " + warns.size()
+ ": " + warns);
return warns.getFirst();
}
@Test
@DisplayName("a profile WITH a configured model gets the actionable models.allow fix, not the fallback")
void profileWithModelGetsTheActionableFix(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = ConfigRef.fixed(cfg);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
Map<String, String> reasonByCredential = new HashMap<>();
ExhaustionSink sink = Fleetd.exhaustionSink(emptyRosterSessions(), config, quarantine,
reasonByCredential, cfg);
ListAppender<ILoggingEvent> appender = attach();
try {
sink.onExhausted("term_x", "The usage limit has been reached", "terra");
} finally {
detach(appender);
}
String warning = fixWarning(appender);
assertTrue(warning.contains("terra"), "must name the profile: " + warning);
assertTrue(warning.contains("claude-opus-5"), "must name the model: " + warning);
assertTrue(warning.contains("enabled: false"), "must name the action: " + warning);
assertTrue(warning.contains("models.allow"), "must name where the action goes: " + warning);
assertFalse(warning.contains("has no model: configured"),
"a profile WITH a model must not get the no-model fallback text: " + warning);
}
@Test
@DisplayName("a profile with NO configured model gets the fallback, never a fix that does not exist")
void profileWithNoModelGetsTheFallbackNotAFix(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = ConfigRef.fixed(cfg);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
Map<String, String> reasonByCredential = new HashMap<>();
ExhaustionSink sink = Fleetd.exhaustionSink(emptyRosterSessions(), config, quarantine,
reasonByCredential, cfg);
ListAppender<ILoggingEvent> appender = attach();
try {
sink.onExhausted("term_y", "The usage limit has been reached", "gx");
} finally {
detach(appender);
}
String warning = fixWarning(appender);
assertTrue(warning.contains("gx"), "must name the profile: " + warning);
assertTrue(warning.contains("has no model: configured"),
"must say there is no model to gate by name: " + warning);
assertFalse(warning.contains("enabled: false"),
"a profile with no model: configured has no models.allow entry to flip — must "
+ "not claim one exists: " + warning);
}
}
@@ -0,0 +1,160 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #426: {@code FleetHealthMonitor.coverage} had zero references anywhere in the test
* tree — not the method, not either output string, not the field it populates. {@code
* FleetHealthMonitorCoverageTest} (package {@code dev.ltms.fleet.health}) pins the three-branch
* method itself; that is the easy half.
*
* <p>The half that actually matters is this one: {@link Fleetd#healthCoverageSource} is the exact
* call site {@code Fleetd.main} wires into {@code FleetMcp}'s constructor, and it is what feeds
* {@code fleet_list}'s {@code healthCoverage} field (see {@code FleetMcp#listFleet}'s {@code
* result.put("healthCoverage", healthCoverage.value().get())}). Measured precedent on fleetd #423
* (for #415): swapping two arguments at a call site like this one — recreating #415's defect with
* the keys exchanged — compiled with 0 errors and ran the ENTIRE suite (1506 tests) green. A
* thoroughly-tested method proves nothing about whether the call site pairs its arguments correctly;
* only a test that drives the call site itself can catch that.
*
* <p><strong>Why this is not driven through a real {@code Fleetd.main} the way #407 drives its five
* reporters</strong> (keep the config invalid, assert on the log line emitted before {@code
* cfg.validateAll()} throws): all five of #407's reporters run in {@code main} before {@code
* validateAll()} (line ~171). {@link Fleetd#healthCoverageSource} is built during {@code FleetMcp}
* construction, which happens only after {@code UnixSocketHerdrClient.connect} has already opened a
* real herdr socket (line ~188) and after {@code sessions}/{@code workers} are constructed. Reaching
* this call site by actually running {@code main} would require a real socket connect — banned by
* this ticket's hard constraints — so #407's option 1 does not apply here. Instead {@link
* Fleetd#healthCoverageSource} is extracted to a directly-callable package-private factory, the same
* shape {@link Fleetd#capacitySource} and {@link Fleetd#quarantineSource} already use for the same
* reason (see {@code FleetdCapacitySourceWiringTest}, the direct precedent this test follows).
*
* <p>Uses a real {@link FleetConfig#load} + {@link ConfigRef} (no socket, no port bind, no spawn,
* nothing written outside {@code @TempDir}) so the fixture goes through the actual YAML parser and
* {@code FleetConfig.Health}/{@code Notifications} records, not a hand-built stand-in that could
* silently drift from what the parser actually produces.
*/
class FleetdHealthCoverageSourceWiringTest {
private static final String BASE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
""";
private static final String HEALTH_DISABLED_WITH_WEBHOOK = BASE + """
health:
enabled: false
notifications:
mode: webhook
""";
private static final String HEALTH_DETECTION_ONLY = BASE + """
health:
enabled: true
""";
private static final String HEALTH_FULL = BASE + """
health:
enabled: true
notifications:
mode: webhook
""";
private static final String HEALTH_ABSENT = BASE;
@Test
@DisplayName("enabled: false reports off, even with a webhook configured")
void disabledHealthReportsOff(@TempDir Path dir) throws Exception {
FleetMcp.HealthCoverageSource source = sourceFor(dir, HEALTH_DISABLED_WITH_WEBHOOK);
assertEquals("off", source.value().get(),
"health.enabled: false must report 'off' regardless of notifications — flipping "
+ "the 'enabled' argument at the HealthCoverageSource call site would report "
+ "'full' here instead");
}
@Test
@DisplayName("enabled with no notifications reports detection-only")
void enabledWithoutNotificationsReportsDetectionOnly(@TempDir Path dir) throws Exception {
FleetMcp.HealthCoverageSource source = sourceFor(dir, HEALTH_DETECTION_ONLY);
assertEquals("detection-only", source.value().get());
}
@Test
@DisplayName("enabled with a webhook configured reports full")
void enabledWithNotificationsReportsFull(@TempDir Path dir) throws Exception {
FleetMcp.HealthCoverageSource source = sourceFor(dir, HEALTH_FULL);
assertEquals("full", source.value().get(),
"health.enabled: true with notifications.mode: webhook must report 'full' — "
+ "swapping 'full' and 'detection-only' at the call site, or breaking the "
+ "enabled/notificationConfigured argument pairing, would report "
+ "'detection-only' here instead");
}
@Test
@DisplayName("an absent health: block reports off")
void absentHealthBlockReportsOff(@TempDir Path dir) throws Exception {
FleetMcp.HealthCoverageSource source = sourceFor(dir, HEALTH_ABSENT);
assertEquals("off", source.value().get());
}
/**
* fleetd #426, the live-wiring half: {@code health.notifications} is a {@code
* ConfigRef.SPLIT_KEYS} entry, and {@link Fleetd#healthCoverageSource} reads {@code
* config.get().health()} live (not the frozen startup {@code cfg}) — exactly like {@link
* Fleetd#capacitySource}'s {@code maxLoad} ({@code FleetdCapacitySourceWiringTest}'s {@code
* reloadedMaxLoadStillChangesWhatFleetListReports}). A hot-reloaded notifications block must
* change what {@code fleet_list} reports without a restart; a fix that froze the whole source
* against the startup snapshot would silently break that and every other test above would stay
* green, since none of them reload.
*/
@Test
@DisplayName("a hot notifications reload still changes what fleet_list reports")
void reloadedNotificationsStillChangeWhatFleetListReports(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, HEALTH_DETECTION_ONLY);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.HealthCoverageSource source = Fleetd.healthCoverageSource(config);
assertEquals("detection-only", source.value().get(),
"sanity: detection-only before any reload");
Files.writeString(file, HEALTH_FULL);
assertTrue(config.reload().applied());
// The live snapshot now carries the webhook — proves the reload really happened and this
// test is not accidentally passing because nothing changed.
assertTrue(config.get().health().notifications() != null
&& config.get().health().notifications().configured(),
"sanity: the reloaded config really carries a configured webhook");
assertEquals("full", source.value().get(),
"the SAME HealthCoverageSource instance must reflect a reloaded notifications "
+ "block without a restart — health.notifications is read live off "
+ "config.get(), exactly like capacitySource's maxLoad");
}
private static FleetMcp.HealthCoverageSource sourceFor(Path dir, String yaml) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, yaml);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
return Fleetd.healthCoverageSource(config);
}
}
@@ -1,15 +1,12 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.msg.LeadMailbox;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
@@ -52,35 +49,28 @@ class FleetdLeadMailboxSelectionTest {
}
}
private static ListAppender<ILoggingEvent> captureFleetdLogs() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static String joined(ListAppender<ILoggingEvent> appender, Level level) {
return appender.list.stream().filter(e -> e.getLevel() == level)
private static String joined(CapturedLog captured, Level level) {
return captured.events().stream().filter(e -> e.getLevel() == level)
.map(ILoggingEvent::getFormattedMessage).reduce("", (a, b) -> a + "\n" + b);
}
@Test
void noCoordinatorBlockLeavesTheFeatureOffSilently() {
var appender = captureFleetdLogs();
var opener = new RecordingOpener();
try (var captured = CapturedLog.of(Fleetd.class)) {
var opener = new RecordingOpener();
assertNull(Fleetd.openLeadMailbox(null, Map.of(), opener));
assertNull(Fleetd.openLeadMailbox(null, Map.of(), opener));
assertNull(opener.offeredUri, "nothing configured means nothing is opened");
assertEquals("", joined(appender, Level.WARN),
"an opt-in feature nobody asked for must not warn on every boot");
assertNull(opener.offeredUri, "nothing configured means nothing is opened");
assertEquals("", joined(captured, Level.WARN),
"an opt-in feature nobody asked for must not warn on every boot");
}
}
@Test
void opensTheMailboxWhenAUriAndSelfIdAreConfigured() {
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null);
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null, null);
Fleetd.openLeadMailbox(coordinator, Map.of(), opener);
@@ -94,7 +84,7 @@ class FleetdLeadMailboxSelectionTest {
void honoursUriEnvOverALiteralUri() {
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator("amqp://stale:stale@old:5672/x", "COORD_URI",
"mac-opus", 8);
"mac-opus", 8, null);
Fleetd.openLeadMailbox(coordinator, Map.of("COORD_URI", RESOLVED_URI), opener);
@@ -105,7 +95,7 @@ class FleetdLeadMailboxSelectionTest {
@Test
void turnsOffWhenUriEnvDoesNotResolve() {
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(null, "COORD_URI", "mac-opus", null);
var coordinator = new FleetConfig.Coordinator(null, "COORD_URI", "mac-opus", null, null);
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
@@ -114,31 +104,33 @@ class FleetdLeadMailboxSelectionTest {
@Test
void warnsAndStaysOffWhenSelfIdIsMissing() {
var appender = captureFleetdLogs();
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, null, null);
try (var captured = CapturedLog.of(Fleetd.class)) {
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, null, null, null);
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
assertNull(opener.offeredUri, "a mailbox with no owning coord-id has no queue to declare");
String warns = joined(appender, Level.WARN);
assertTrue(warns.contains("coordinator.selfId"), () -> "say which key is missing: " + warns);
assertFalse(warns.contains(SECRET), () -> "the URI's password must never be logged: " + warns);
assertNull(opener.offeredUri, "a mailbox with no owning coord-id has no queue to declare");
String warns = joined(captured, Level.WARN);
assertTrue(warns.contains("coordinator.selfId"), () -> "say which key is missing: " + warns);
assertFalse(warns.contains(SECRET), () -> "the URI's password must never be logged: " + warns);
}
}
@Test
void warnsAndStaysOffWhenTheBrokerIsUnreachableAtBoot() {
var appender = captureFleetdLogs();
var opener = new RecordingOpener();
opener.unreachable = true;
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null);
try (var captured = CapturedLog.of(Fleetd.class)) {
var opener = new RecordingOpener();
opener.unreachable = true;
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null, null);
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener),
"a down coordination broker turns the feature off; it must never take the daemon down");
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener),
"a down coordination broker turns the feature off; it must never take the daemon down");
String warns = joined(appender, Level.WARN);
assertTrue(warns.contains("coord.example"), () -> "name the host that failed: " + warns);
assertFalse(warns.contains(SECRET), () -> "with credentials stripped: " + warns);
assertTrue(warns.contains("Connection refused"), () -> "and the real reason: " + warns);
String warns = joined(captured, Level.WARN);
assertTrue(warns.contains("coord.example"), () -> "name the host that failed: " + warns);
assertFalse(warns.contains(SECRET), () -> "with credentials stripped: " + warns);
assertTrue(warns.contains("Connection refused"), () -> "and the real reason: " + warns);
}
}
}
@@ -0,0 +1,89 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #480 Unit A, hard requirement 6: pin {@code Fleetd.main}'s construction of {@link
* dev.ltms.fleet.lead.LeadRollover} with a source-text assertion, mirroring {@code
* FleetdCompletionResolverWiringTest}'s pattern — five log-only reporters in {@code Fleetd.main}
* already survived mutation batteries this exact way (fleetd #415's extraction antidote note).
*
* <p>What this class still covers, and what it never claimed to. {@code LeadRolloverTest}
* constructs its own {@code LeadRollover} directly (as every prior test of an extracted factory
* does) with a hand-built lookup, so a mutation that deletes the {@code leadRollover(...)} call
* from {@code main} — or replaces one of its arguments with something that still compiles, e.g.
* {@code router.leadAgents()} swapped for {@code null}, or the whole assignment swapped for a bare
* {@code null} literal — leaves every behavioural test green. This is a plain string read, guarded
* by an unrelated anchor assertion so a broken or empty file read cannot pass as a real change.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* LeadRollover} and never runs {@code main}. It pins the {@code leadRollover(...)} CALL SITE's
* argument list — that {@code main} still passes {@code leads} at all — never what the factory
* DOES with that argument once inside its own body.
*
* <p><b>Correction (fleetd #480 relative-handover-path follow-up): that gap used to be real, and
* now is not — but not here.</b> This class's javadoc previously claimed "no behavioural test can
* catch this wiring dropping out" for the whole factory, including the lambda {@code
* leadRollover(...)} builds internally (terminal → lead name → {@code Leader.cwd()}). That claim
* was proven true at the time — mutating that lambda's body to {@code String leadName = null;}
* (always "no lead found", which silently reintroduces the daemon-cwd bug this ticket fixes) left
* the full suite green, {@code Tests run: 1669, Failures: 0}. It is no longer true: {@code
* FleetdLeadRolloverWorkspaceLookupTest} now calls {@code Fleetd.leadRollover(...)} directly with a
* real {@link dev.ltms.fleet.config.ConfigRef} built from a temp {@code fleetd.yaml}, and fails
* against that exact one-line mutation. So: THIS class still covers only the call site's argument
* list; {@code FleetdLeadRolloverWorkspaceLookupTest} is what now covers the lambda's body. Neither
* one subsumes the other — keep both.
*/
class FleetdLeadRolloverWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] unrelated anchor: Fleetd.java still declares the Fleetd class")
void unrelatedAnchorStillPresent() throws Exception {
// Guards the two assertions below: without this, a bad read (empty string, wrong file,
// truncated file) could vacuously fail to contain the leadRollover(...) call too, and a
// test that only asserts "contains X" would report a false pass for the wrong reason if X
// happened to match. Asserting an unrelated, structurally distant string first proves the
// read actually pulled real file content.
String source = fleetdSource();
assertTrue(source.contains("public final class Fleetd"),
"sanity anchor failed — the file read did not return real Fleetd.java source; the "
+ "leadRollover(...) wiring assertions below cannot be trusted until this passes");
}
@Test
@DisplayName("[SOURCE TEXT] main still constructs LeadRollover via the leadRollover(...) factory, exactly as heartbeat is constructed")
void mainStillCallsTheLeadRolloverFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);"),
"Fleetd.main must still assign `LeadRollover leadRollover = leadRollover(cfg, "
+ "router.leadAgents(), config, leads);`. Dropping this call, or swapping one of "
+ "its arguments for something that still compiles (e.g. null in place of "
+ "router.leadAgents()), leaves every behavioural test green — this source check is "
+ "what must go red instead. fleetd #480 correction 2 deliberately dropped "
+ "primaryRegistry from this call — see LeadRollover's class javadoc for why a "
+ "single-slot lookup was wrong here. The fleetd #480 relative-handover-path "
+ "follow-up added `leads` (terminal → lead name) so the factory can resolve a "
+ "relative handoverPath against the calling lead's own workspace.");
}
@Test
@DisplayName("[SOURCE TEXT] the leadRollover(...) factory itself gates construction on cfg.leadRollover() != null")
void factoryGatesOnConfigPresence() throws Exception {
String source = fleetdSource();
assertTrue(source.contains("if (cfg.leadRollover() == null) {"),
"Fleetd.leadRollover(...) must refuse to construct a LeadRollover when the "
+ "leadRollover: block is absent — an upgraded daemon must never silently acquire "
+ "the ability to clear the lead's own pane. See LeadHeartbeatLoop's construction "
+ "gate (cfg.leadHeartbeat() != null) for the pattern this mirrors.");
}
}
@@ -0,0 +1,163 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.lead.LeadRollover;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashMap;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
/**
* fleetd #480 relative-handover-path follow-up, correction round: a BEHAVIOURAL test of {@link
* Fleetd#leadRollover}'s own body — the terminal → lead-name → {@code Leader.cwd()} lookup it
* builds — not another source-text pin.
*
* <p>{@code FleetdLeadRolloverWiringTest} (a plain string read) still earns its keep: it pins the
* call site's ARGUMENT LIST, so a mutation that drops {@code leads} back out of the call, or
* swaps it for {@code Map::of}, still goes red there. But nothing before this class exercised the
* LAMBDA BODY {@code leadRollover(...)} builds — the terminal-to-workspace {@code
* Function<String,String>} that {@link LeadRollover#open} actually calls. Every prior test either
* exercised {@code LeadRollover} directly with a hand-built lookup ({@code LeadRolloverTest}), or
* read source text without ever calling the factory ({@code FleetdLeadRolloverWiringTest}) — so a
* mutation that breaks the lookup ITSELF (e.g. always resolving no lead, or reading a snapshot
* instead of the live supplier) left every existing test green while the daemon's real wiring
* silently reintroduced the exact bug this whole ticket fixes: a relative {@code handoverPath}
* resolving against the daemon's own {@code cwd} instead of the calling lead's.
*
* <p>{@code Fleetd.leadRollover(...)} is package-private and {@code static}, so this test — living
* in the same {@code dev.ltms.fleet} package — calls it directly, exactly the way {@code
* FleetdExhaustionDetectionArmedWiringTest} and {@code FleetdCapacitySourceWiringTest} already call
* other package-private startup factories with a real {@link ConfigRef} built from a temp {@code
* fleetd.yaml} (never the gitignored live one).
*
* <p><b>Proved against the mutation it exists to catch.</b> Before this class was added, mutating
* {@code leadRollover(...)}'s lambda body to {@code String leadName = null;} (always "no lead
* found", which forces every relative {@code handoverPath} onto the {@code user.dir} fallback —
* i.e. the original bug) left the full suite green: {@code Tests run: 1669, Failures: 0}. With
* {@link #relativeHandoverPathResolvesAgainstTheLeadsConfiguredCwd()} added, the same one-line
* mutation now fails that test (it asserts the resolved path equals the configured lead's {@code
* cwd}, which the mutant can never produce) — see this ticket's fleet_reply history for both runs.
*/
class FleetdLeadRolloverWorkspaceLookupTest {
private static AgentControl fakeAgents() {
return new AgentControl(new FakeHerdr());
}
@Test
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
+ "the CALLING lead's configured cwd, not the daemon's own working directory")
void relativeHandoverPathResolvesAgainstTheLeadsConfiguredCwd(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
leadRollover:
handoverPath: handover.md
""".formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
() -> Map.of("term_opus", "opus"));
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
+ "must construct an object");
LeadRollover.PendingRollover pending = rollover.open("term_opus", "test");
String expected = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expected, pending.handoverPath(),
"a relative handoverPath must resolve against the CALLING lead's configured cwd "
+ "through the REAL Fleetd.leadRollover(...) wiring — not the daemon's own "
+ "working directory. This is the exact axis that was proven uncovered: "
+ "mutating the factory's lambda body to always report \"no lead found\" "
+ "left every prior test green.");
}
@Test
@DisplayName("[BEHAVIOURAL] a terminal not present in the live lead-terminal map falls back "
+ "to the daemon's own user.dir")
void terminalNotInLiveMapFallsBackToUserDir(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
leadRollover:
handoverPath: handover.md
""");
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
// has not yet scanned, or one with no fleet.leaders entry at all.
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of);
assertNotNull(rollover);
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
String expected = Path.of(System.getProperty("user.dir")).resolve("handover.md")
.normalize().toString();
assertEquals(expected, pending.handoverPath(),
"a lead not present in the live terminal→name map must fall back to the daemon's "
+ "own user.dir — the same fallback LeadLauncher#launch already uses for a "
+ "lead with no configured cwd. This is deliberate, pinned behaviour, not an "
+ "accident of the null-check chain.");
}
@Test
@DisplayName("[BEHAVIOURAL] the terminal→lead-name lookup is read LIVE on every open() call, "
+ "never snapshotted at Fleetd.leadRollover(...) construction time")
void workspaceLookupIsReadLiveNotSnapshotted(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
leadRollover:
handoverPath: handover.md
""".formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
// EMPTY at the moment leadRollover(...) is called. A lookup captured (snapshotted) here
// instead of read live through the supplier on every call would never see the entry added
// below — exactly the natural mistake to make, since leads are discovered by a live tab
// scan that runs AFTER this factory is constructed at startup.
Map<String, String> liveLeadTerminals = new HashMap<>();
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
() -> liveLeadTerminals);
assertNotNull(rollover);
// The lead is "discovered" only now — mutate the SAME backing map the supplier reads from.
liveLeadTerminals.put("term_opus", "opus");
LeadRollover.PendingRollover pending = rollover.open("term_opus", "test");
String expected = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expected, pending.handoverPath(),
"the terminal→lead-name lookup must be read LIVE on every open() call — a lead "
+ "discovered by the tab scan AFTER Fleetd.leadRollover(...) was constructed "
+ "must still resolve correctly, not only one that was already live at "
+ "construction time");
}
}
@@ -0,0 +1,174 @@
package dev.ltms.fleet;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.LoopWatchdog;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.SessionReaper;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #562 follow-up (issue comment "HOLD on PR #579"): {@code Fleetd.main}'s {@code loopHealth}
* local used to be a bare {@code new FleetMcp.LoopHealthSource(poller::health, ...)} built inline,
* with nothing a test could call directly. Measured on that shape: replacing {@code
* poller::health} with a constant {@code () -> LoopWatchdog.State.RUNNING} at the call site
* compiled with 0 errors and left all 1771 existing tests green — the daemon could be changed to
* always report the {@link StatusPoller} as {@code RUNNING}, so the watchdog could never fire and
* a stalled poller would be invisible, while every test stayed green. That is exactly the false
* negative this ticket exists to prevent.
*
* <p>The five tests PR #579 added ({@code FleetMcpTest}, {@code FleetAppTest}) all build their own
* {@link FleetMcp.LoopHealthSource} directly with fixed lambdas — they prove the seam ({@code
* LoopHealthSource} reports what it is given) and nothing about what {@code Fleetd.main} actually
* gives it. This is the same hand-built-vs-config-wired shape as fleetd #561/#248/#426.
*
* <p>The fix extracts the inline {@code new} into {@link Fleetd#loopHealthSource}, a package-private
* factory in the same style as {@link Fleetd#capacitySource} and {@link Fleetd#healthCoverageSource}
* — which is exactly what makes it directly callable here. This test calls that factory with real
* {@link StatusPoller}/{@link SessionReaper} instances (never started, so no herdr or git I/O
* happens) and pins each half separately, plus the {@code reaper == null} branch: one invariant
* wired at three places needs three assertions, not one combined check whose non-zero total could
* hide a gap at any single place.
*/
class FleetdLoopHealthSourceWiringTest {
@Test
@DisplayName("the statusPoller half reports the real poller's health, not a hardcoded state")
void statusPollerHalfReflectsThePollersRealHealth() {
// Stopped without ever being started — stop() still marks the watchdog STOPPED. A poller
// that has never reported RUNNING is the discriminating case: if Fleetd.loopHealthSource
// ever hardcoded RUNNING (the exact mutation this test exists to catch), this would fail.
StatusPoller stoppedPoller = freshPoller();
stoppedPoller.stop();
SessionReaper unusedReaper = freshReaper(); // present only to satisfy the signature
FleetMcp.LoopHealthSource source = Fleetd.loopHealthSource(stoppedPoller, unusedReaper);
assertEquals(LoopWatchdog.State.STOPPED, source.statusPoller().get(),
"the statusPoller supplier must delegate to the real poller's health() — "
+ "replacing poller::health with a constant () -> RUNNING at the "
+ "Fleetd.loopHealthSource call site must fail this assertion");
}
@Test
@DisplayName("the sessionReaper half reports the real reaper's health, not a hardcoded state")
void sessionReaperHalfReflectsTheReapersRealHealth() {
StatusPoller unusedPoller = freshPoller(); // present only to satisfy the signature
SessionReaper stoppedReaper = freshReaper();
stoppedReaper.stop();
FleetMcp.LoopHealthSource source = Fleetd.loopHealthSource(unusedPoller, stoppedReaper);
assertEquals(LoopWatchdog.State.STOPPED, source.sessionReaper().get(),
"the sessionReaper supplier must delegate to the real reaper's health() — "
+ "replacing reaper.health() with a constant at the "
+ "Fleetd.loopHealthSource call site must fail this assertion");
}
@Test
@DisplayName("a null reaper (idle ttl not configured) still reports STOPPED, not a crash")
void nullReaperStillReportsStopped() {
// SessionReaper is only constructed when lifecycle.idleTtlSeconds is configured (see the
// `reaper` local in Fleetd.main) — a real deployment routinely passes null here. That null
// check is real behaviour, not a simplification to delete: it must keep reporting STOPPED
// rather than throwing a NullPointerException on the first fleet_list/healthz call.
StatusPoller runningPoller = freshPoller();
FleetMcp.LoopHealthSource source = Fleetd.loopHealthSource(runningPoller, null);
assertEquals(LoopWatchdog.State.STOPPED, source.sessionReaper().get(),
"reaper == null must still report STOPPED, exactly like an intentionally-stopped "
+ "reaper would — do not delete this null check to simplify the wiring");
}
/** Never started, so no herdr call is ever made; freshly constructed reports RUNNING. */
private static StatusPoller freshPoller() {
AgentControl agents = new AgentControl(new FakeHerdr());
return new StatusPoller(agents, new Injector(agents), 1000);
}
/** Never started, so no git/session I/O is ever made; freshly constructed reports RUNNING. */
private static SessionReaper freshReaper() {
return new SessionReaper(new SessionManager(new NeverSpawnsLauncher()), 60, 1000);
}
/**
* Same minimal shape as {@code FleetdBackendErrorSinkTest.NeverSpawnsLauncher} — every method
* throws or returns an empty/no-op value, since a {@link SessionReaper} that is only ever
* constructed and then stopped (never started) never calls any of them.
*/
private static final class NeverSpawnsLauncher implements PeerLauncher {
@Override
public Set<Capability> capabilities() {
return Set.of();
}
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return Set.of();
}
@Override
public PeerHandle spawn(SpawnRequest req) {
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
}
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
}
@Override
public Set<String> profiles() {
return Set.of();
}
@Override
public String defaultProfile() {
return null;
}
@Override
public String effectiveCwd(SpawnRequest req) {
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
}
@Override
public List<String> parityOverlay(String profileName) {
return List.of();
}
@Override
public List<?> list() {
return List.of();
}
@Override
public int reapOrphanWorkers() {
return 0;
}
@Override
public void stop(String id) {
}
@Override
public boolean clearContext(String id) {
return false;
}
}
}
@@ -0,0 +1,62 @@
package dev.ltms.fleet;
import dev.ltms.fleet.peer.PeerLauncher;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
/**
* fleetd #422 follow-up: {@code Fleetd.modelGateCoverageLine} is the startup-log counterpart of
* {@code exhaustedPatternCoverageLine}/{@code errorPatternCoverageLine} — see {@code
* FleetdPatternCoverageLineTest} for the identical shape this follows — except here there is no
* {@code UnsetMeaning} choice for a caller to get backwards: {@link
* PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether an empty
* {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate with at all"
* or "a block armed and currently reporting zero off". This class proves {@code
* modelGateCoverageLine} words those two states — plus the third, N off — distinctly, so a
* mutation that made it ignore {@code configured()} either way is caught here.
*/
class FleetdModelGateCoverageLineTest {
@Test
@DisplayName("no models: block reports not configured, distinct from armed-with-zero")
void noModelsBlockReportsNotConfigured() {
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
assertEquals("not configured (no models: block — nothing is gated, and nothing can be)", line);
}
@Test
@DisplayName("a models: block armed with nothing off reports armed, distinct from not configured")
void armedWithNothingOffReportsArmed() {
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
assertEquals("armed (models: block present; 0 models currently turned off)", line);
}
@Test
@DisplayName("a models: block with N off names the off models")
void armedWithModelsOffNamesThem() {
String line = Fleetd.modelGateCoverageLine(
PeerLauncher.ModelGateState.armed(Set.of("deepseek-v4-flash", "claude-opus-9000")));
assertEquals("armed (2 model(s) turned off: [claude-opus-9000, deepseek-v4-flash])", line);
}
@Test
@DisplayName("the three states produce pairwise-distinct wording for the same empty-looking input")
void theThreeStatesProduceDistinctWording() {
String notConfigured = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
String armedZero = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
String armedOne = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of("x")));
// Pinned individually above; restated here so this test alone still catches a regression
// even if one of the three tests above were ever deleted — the exact FleetdPatternCoverageLineTest
// pattern, adapted from "two keys" to "three states of one gate".
assertNotEquals(notConfigured, armedZero,
"collapsing 'no models: block' into 'armed, zero off' is fleetd #422 follow-up's exact defect");
assertNotEquals(armedZero, armedOne);
assertNotEquals(notConfigured, armedOne);
}
}
@@ -0,0 +1,67 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
/**
* fleetd #415 (review follow-up): {@code CompletionResolverTest} proves {@code coverage()} words
* {@code UnsetMeaning.OFF} and {@code UnsetMeaning.BUILT_IN_DEFAULT} correctly — but every one of
* those tests supplies the meaning itself. That proves the enum's wording, never that {@code
* Fleetd} pairs the right meaning with the right pattern key. That pairing is #415's actual
* defect: {@code coverage()} had no way to know what unset meant for its key, so the fix moved
* the fact to the caller — and nothing yet proved the caller states it correctly.
*
* <p><b>Measured or it didn't happen:</b> swapping the two {@code UnsetMeaning} arguments at
* {@code Fleetd}'s two coverage call sites — giving {@code exhaustedPattern} the built-in-default
* wording and {@code errorPattern} the off wording, #415's exact defect with the keys exchanged —
* compiled with 0 errors and left all 1506 existing tests green. This class exists to turn that
* swap red.
*
* <p>It calls {@link Fleetd#exhaustedPatternCoverageLine} and {@link Fleetd#errorPatternCoverageLine}
* directly rather than reading {@code Fleetd.java} as source text (the shape {@code
* FleetdCompletionResolverWiringTest} uses for a different wiring gap): those two methods are the
* extracted call sites {@code main} actually invokes, following the same {@code static} factory +
* dedicated-test pattern as {@link Fleetd#capacitySource} and {@link Fleetd#worktreeBranchLookup}.
*/
class FleetdPatternCoverageLineTest {
private static final Set<String> PROFILES = Set.of("terra", "gx10");
@Test
@DisplayName("exhaustedPatternCoverageLine says off when no profile configures exhaustedPattern")
void exhaustedPatternCoverageLineSaysOffWhenNoProfileConfiguresIt() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of()));
}
@Test
@DisplayName("errorPatternCoverageLine says built-in default when no profile configures errorPattern")
void errorPatternCoverageLineSaysBuiltInDefaultWhenNoProfileConfiguresIt() {
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])",
Fleetd.errorPatternCoverageLine(PROFILES, Set.of()));
}
@Test
@DisplayName("the two keys produce different wording for the identical empty-coverage input")
void theTwoKeysProduceDifferentWordingForTheSameEmptyInput() {
String exhaustedLine = Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of());
String errorLine = Fleetd.errorPatternCoverageLine(PROFILES, Set.of());
// Pinned individually above; restated here so this test alone still catches a swap even
// if one of the two tests above were ever deleted.
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
exhaustedLine);
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])", errorLine);
assertNotEquals(exhaustedLine, errorLine,
"swapping which UnsetMeaning pairs with which pattern key at Fleetd's call sites "
+ "must be caught here — that pairing, not coverage()'s own wording in isolation, "
+ "is fleetd #415's actual defect");
}
}
@@ -1,15 +1,13 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.msg.AmqpReplyInbox;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.net.ServerSocket;
import java.util.List;
@@ -50,43 +48,34 @@ class FleetdReplyInboxSelectionTest {
}
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
private static CapturedLog attach() {
// logback-test.xml pins dev.ltms.fleet to WARN; raise it so INFO selection lines are captured.
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
return CapturedLog.at(Fleetd.class, Level.INFO);
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
}
private static void assertNoLogContains(ListAppender<ILoggingEvent> appender, String secret) {
assertTrue(appender.list.stream().noneMatch(e -> e.getFormattedMessage().contains(secret)),
private static void assertNoLogContains(List<ILoggingEvent> events, String secret) {
assertTrue(events.stream().noneMatch(e -> e.getFormattedMessage().contains(secret)),
"no log line may contain the resolved URI's password");
}
@Test
void uriEnvSetAndPresentSelectsAmqpWithTheResolvedUri() {
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, appender, inbox) -> {
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, events, inbox) -> {
assertEquals(RESOLVED_URI, opener.offeredUri,
"the daemon must connect with the value resolved from uriEnv — selection, not just parse");
assertEquals(opener.inbox, inbox, "the AMQP opener's inbox is what is selected");
assertNoLogContains(appender, SECRET);
assertNoLogContains(events, SECRET);
});
}
@Test
void uriEnvSetButVariableMissingFallsBackToInMemoryAndWarns() {
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
recording(broker, Map.of(), false, (opener, appender, inbox) -> {
recording(broker, Map.of(), false, (opener, events, inbox) -> {
assertInstanceOf(InMemoryReplyInbox.class, inbox);
assertNull(opener.offeredUri, "AMQP must never be attempted when the variable is missing");
assertTrue(hasWarnContaining(appender, "LAVINMQ_URI") && hasWarnContaining(appender, "DISABLED"),
assertTrue(hasWarnContaining(events, "LAVINMQ_URI") && hasWarnContaining(events, "DISABLED"),
"a missing uriEnv variable must warn loudly, not fail silently");
});
}
@@ -94,10 +83,10 @@ class FleetdReplyInboxSelectionTest {
@Test
void uriEnvSetButVariableBlankFallsBackToInMemoryAndWarns() {
FleetConfig.Broker broker = new FleetConfig.Broker("amqp://user:lame@old:5672/", "LAVINMQ_URI", null);
recording(broker, Map.of("LAVINMQ_URI", " "), false, (opener, appender, inbox) -> {
recording(broker, Map.of("LAVINMQ_URI", " "), false, (opener, events, inbox) -> {
assertInstanceOf(InMemoryReplyInbox.class, inbox);
assertNull(opener.offeredUri, "a blank env value must not select AMQP, not even via the literal uri");
assertTrue(hasWarnContaining(appender, "LAVINMQ_URI"),
assertTrue(hasWarnContaining(events, "LAVINMQ_URI"),
"a blank uriEnv value must warn, and must not fall back to the literal uri");
});
}
@@ -106,22 +95,22 @@ class FleetdReplyInboxSelectionTest {
void bothUriAndUriEnvSetUriEnvWinsDeterministically() {
FleetConfig.Broker broker
= new FleetConfig.Broker("amqp://user:oldpw@old.example:5672/", "LAVINMQ_URI", null);
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, appender, inbox) -> {
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, events, inbox) -> {
assertEquals(RESOLVED_URI, opener.offeredUri,
"uriEnv must win over uri, deterministically, every run");
assertTrue(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("broker.uri is ignored")),
assertTrue(events.stream().anyMatch(e -> e.getFormattedMessage().contains("broker.uri is ignored")),
"must log that the literal uri is ignored when uriEnv is set");
assertNoLogContains(appender, SECRET);
assertNoLogContains(appender, "oldpw");
assertNoLogContains(events, SECRET);
assertNoLogContains(events, "oldpw");
});
}
@Test
void unreachableBrokerStartsDaemonWithInMemoryInboxAndLoudWarning() {
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), true, (opener, appender, inbox) -> {
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), true, (opener, events, inbox) -> {
assertInstanceOf(InMemoryReplyInbox.class, inbox, "an unreachable broker must NOT stop the daemon");
String warn = appender.list.stream()
String warn = events.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.reduce("", (a, b) -> a + "\n" + b)
@@ -129,16 +118,16 @@ class FleetdReplyInboxSelectionTest {
assertTrue(warn.contains("durable") && warn.contains("soft-state"),
"the warning must say exactly what was lost: durable delivery off, replies soft-state");
assertTrue(!warn.contains(SECRET), "the failing URI must be logged with credentials stripped");
assertNoLogContains(appender, SECRET);
assertNoLogContains(events, SECRET);
});
}
@Test
void noBrokerConfiguredStaysQuietInMemory() {
FleetConfig.Broker broker = null;
recording(broker, Map.of(), false, (opener, appender, inbox) -> {
recording(broker, Map.of(), false, (opener, events, inbox) -> {
assertInstanceOf(InMemoryReplyInbox.class, inbox);
assertTrue(appender.list.stream().noneMatch(e -> e.getLevel() == Level.WARN),
assertTrue(events.stream().noneMatch(e -> e.getLevel() == Level.WARN),
"no broker configured must keep the existing QUIET in-memory path — no warning");
});
}
@@ -152,21 +141,18 @@ class FleetdReplyInboxSelectionTest {
closedPort = s.getLocalPort();
}
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
ListAppender<ILoggingEvent> appender = attach();
try {
try (CapturedLog log = attach()) {
ReplyInbox inbox = Fleetd.selectReplyInbox(
broker, Map.of("LAVINMQ_URI", "amqp://user:" + SECRET + "@127.0.0.1:" + closedPort + "/vh"),
AmqpReplyInbox::open);
assertInstanceOf(InMemoryReplyInbox.class, inbox,
"a genuinely unreachable broker (real AmqpReplyInbox::open) must fall back to in-memory");
} finally {
detach(appender);
assertNoLogContains(log.events(), SECRET);
}
assertNoLogContains(appender, SECRET);
}
private boolean hasWarnContaining(ListAppender<ILoggingEvent> appender, String fragment) {
return appender.list.stream().anyMatch(e ->
private boolean hasWarnContaining(List<ILoggingEvent> events, String fragment) {
return events.stream().anyMatch(e ->
e.getLevel() == Level.WARN && e.getFormattedMessage().contains(fragment));
}
@@ -174,18 +160,14 @@ class FleetdReplyInboxSelectionTest {
private void recording(FleetConfig.Broker broker, Map<String, String> env, boolean unreachable, Check check) {
RecordingAmqp opener = new RecordingAmqp();
opener.unreachable = unreachable;
ListAppender<ILoggingEvent> appender = attach();
ReplyInbox inbox;
try {
inbox = Fleetd.selectReplyInbox(broker, env, opener);
} finally {
detach(appender);
try (CapturedLog log = attach()) {
ReplyInbox inbox = Fleetd.selectReplyInbox(broker, env, opener);
check.run(opener, log.events(), inbox);
}
check.run(opener, appender, inbox);
}
@FunctionalInterface
private interface Check {
void run(RecordingAmqp opener, ListAppender<ILoggingEvent> appender, ReplyInbox inbox);
void run(RecordingAmqp opener, List<ILoggingEvent> events, ReplyInbox inbox);
}
}
@@ -0,0 +1,79 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Proves {@link Fleetd#main(String[])} calls every startup report before validation aborts startup.
* The invalid non-loopback bind makes {@link FleetConfig#validateAll()} throw before {@code main}
* can open the herdr socket or bind a port. The fixture also triggers every report, so removing any
* one call from {@code main} leaves its expected log line absent.
*/
class FleetdStartupReportTest {
private static Level originalLevel;
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
originalLevel = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
logger.detachAppender(appender);
logger.setLevel(originalLevel);
}
private static boolean contains(ListAppender<ILoggingEvent> appender, String fragment) {
return appender.list.stream()
.map(ILoggingEvent::getFormattedMessage)
.anyMatch(message -> message.contains(fragment));
}
@Test
void mainReportsEveryStartupGapBeforeValidationAborts(@TempDir Path dir) throws Exception {
Path config = dir.resolve("fleetd.yaml");
Files.writeString(config, """
bind:
host: 0.0.0.0
port: 8765
profiles:
worker:
baseUrl: https://llm.ltms.dev/v1
gitTokenEnv: GITEA_TOKEN
""");
ListAppender<ILoggingEvent> appender = attach();
try {
assertThrows(IllegalStateException.class, () -> Fleetd.main(new String[]{config.toString()}));
} finally {
detach(appender);
}
assertTrue(contains(appender, "startup git host GITEA_HOST:"),
"Fleetd.main must report the git host shape");
assertTrue(contains(appender, "member trust model: members run as the same OS user"),
"Fleetd.main must report the member trust model");
assertTrue(contains(appender, "memberCredentials: absent or empty"),
"Fleetd.main must report an absent memberCredentials policy");
assertTrue(contains(appender, "exhaustedPattern: profile(s) [worker] have no exhaustedPattern configured"),
"Fleetd.main must report profiles without exhaustedPattern");
}
}
@@ -0,0 +1,154 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Proves the one thing no test proved before this ticket's follow-up: that {@code Fleetd.main}
* ITSELF — not a copy of its logic, not the validator called directly — still refuses to start on
* a bad config. Mutation testing found that deleting {@code cfg.validateAll();} (née six
* individual {@code cfg.validateXxx();} calls) from {@code Fleetd.main} left the full 1478-test
* suite green; every existing test called a validator directly and none exercised {@code
* Fleetd.main} as the caller. See {@code FleetConfigValidateAllTest} for why the fix collapses
* those six calls into one reflective {@link FleetConfig#validateAll()}, and {@code
* ConfigRefTest} for the equivalent proof on the {@link
* dev.ltms.fleet.config.ConfigRef#reload()} path.
*
* <p>This is deliberately the real {@code static void main(String[] args)} — package-private, so
* only a test in this package can call it, which is exactly what makes this proof strong: it is
* not a helper extracted for testability, it is the literal method {@code java -jar fleetd.jar}
* invokes. Every fixture below is otherwise-valid and fails exactly one validator, and — because
* {@link FleetConfig#validateAll()} runs immediately after {@code SubscriptionGuard.
* assertPrimaryClean}, before {@code main} opens the herdr socket, binds Javalin, or touches
* anything else with a real side effect — calling {@code Fleetd.main} with one of these configs is
* safe: it is guaranteed to throw before reaching any of that, precisely because the config is
* deliberately invalid.
*/
class FleetdStartupValidationTest {
private static void assertMainRefuses(Path dir, String fileName, String yaml,
String mustContain) throws Exception {
Path f = dir.resolve(fileName);
Files.writeString(f, yaml);
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> Fleetd.main(new String[]{f.toString()}),
fileName + ": Fleetd.main must refuse this config before doing anything else");
assertTrue(e.getMessage().contains(mustContain),
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
+ e.getMessage());
}
@Test
void mainRefusesANonLoopbackBindWithoutTokenMode(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "auth-exposure.yaml", """
bind:
host: 0.0.0.0
port: 8765
""", "auth.mode: token");
}
@Test
void mainRefusesALeadTabPrefixCollision(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "lead-tab-prefixes.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
tabLabel: "lead: {role} {profile}"
leaders:
opus:
tab: "lead: opus"
""", "fleet.tabLabel");
}
@Test
void mainRefusesASubscriptionProfileThatReseatsAnthropicBaseUrl(@TempDir Path dir)
throws Exception {
assertMainRefuses(dir, "subscription-profiles.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
""", "ANTHROPIC_BASE_URL");
}
@Test
void mainRefusesAnUnknownCharterKey(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "charters.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
architetc: text
""", "architetc");
}
/**
* fleetd #469: closes the gap left by #464's {@code CharterToolSurfaceTest}, which wrote its
* own charter into a {@code @TempDir} fixture and so could never fail on anything anyone wrote
* into the live {@code fleetd.yaml}. {@code bridge_send} is CB-634's own motivating example — a
* tool name the rename removed — and the charter key ({@code dev}) is a real role wire name, so
* this fixture passes {@code cfg.validateAll()}'s charter check (key valid, text non-blank) and
* is refused only by the new {@code CharterToolSurface} call right after it. The failure message
* must name both the charter key and the unknown tool.
*/
@Test
void mainRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "charter-tool-surface.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
""", "bridge_send");
}
@Test
void mainRefusesAnArchitectSlotNamingAnUnconfiguredProfile(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "members.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
baseUrl: http://gx10.gw:8000
fleet:
architects:
lead-designer:
profile: sonnet
""", "lead-designer");
}
@Test
void mainRefusesAProfileNamingAModelOutsideTheAllowList(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "models.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
rogue:
baseUrl: http://gx11.gw:8000
model: claude-opus-9000
models:
allow:
- model: claude-sonnet-5
""", "rogue");
}
}
@@ -0,0 +1,286 @@
package dev.ltms.fleet;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.inject.TurnListener;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import org.junit.jupiter.api.Test;
import java.util.LinkedHashSet;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #561: {@code Fleetd.turnListener} composes the completion resolver and the session
* manager into one {@link TurnListener}. Each of the four callbacks below used to be two bare,
* unguarded statements (completion first, then sessions) — a throw from the session half used to
* skip nothing <em>after</em> it (there was nothing after it), but nothing enforced that the
* completion half had to come first either, beyond call order in the source. {@code onDelivered}
* had the identical shape and was fixed by #556 (moving its registration off the fan-out
* entirely); these four callbacks cannot be fixed that way, so the fan-out itself is hardened
* instead (see {@link Fleetd#turnListener} and {@code Fleetd.bothMustRun}/{@code
* bothMustRunKeepingSecondResult}).
*
* <p>The invariant under test: <b>a throwing session-listener must not prevent the completion
* resolver from being told the turn ended.</b> Each test below builds the exact production
* composition ({@link Fleetd#turnListener}) from a real {@link CompletionResolver} and a fake
* {@code sessions} half that throws, then asserts the completion half's effect (the captured
* {@link Rendezvous} waiter resolving) happened anyway — never by inspecting call order directly.
*/
class FleetdTurnListenerCompositionTest {
/** Records which callbacks ran and can be told to throw from a chosen one. */
private static final class RecordingSessions implements TurnListener {
final Set<String> called = new LinkedHashSet<>();
private final Set<String> throwing;
RecordingSessions(String... throwingMethods) {
this.throwing = Set.of(throwingMethods);
}
private void maybeThrow(String method) {
called.add(method);
if (throwing.contains(method)) {
throw new IllegalStateException("boom: sessions." + method);
}
}
@Override
public void onTurnComplete(String target) {
maybeThrow("onTurnComplete");
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
maybeThrow("onTurnCompleteWithPostAction");
return true;
}
@Override
public void onTurnFailed(String target) {
maybeThrow("onTurnFailed");
}
@Override
public void onTurnFailed(String target, String reason) {
maybeThrow("onTurnFailedWithReason");
}
}
private static CompletionResolver newResolver(FakeHerdr herdr, Rendezvous rendezvous) {
return new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none());
}
/**
* A resolver with a controllable clock, and the clock itself, so a test can register a turn
* "delivered" at time 0 and then jump the clock past {@link CompletionResolver#MIN_TURN_NANOS}
* before resolving it — otherwise {@code onTurnComplete}/{@code onTurnCompleteWithPostAction}
* resolve within microseconds of registering in-test, well inside the fleetd#164 floor, and
* get classified as a too-fast crash rather than a real completion. Mirrors the {@code
* LongSupplier} clock-injection pattern fleetd#164's own tests use.
*/
private static CompletionResolver newResolverPastTheFloor(FakeHerdr herdr, Rendezvous rendezvous,
AtomicLong clock) {
LongSupplier nowNanos = clock::get;
return new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), nowNanos);
}
@Test
void onTurnCompleteResolvesTheWaiterEvenWhenTheSessionHalfThrows() throws Exception {
FakeHerdr herdr = new FakeHerdr().readText("⏺ real answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
AtomicLong clock = new AtomicLong(0);
CompletionResolver completion = newResolverPastTheFloor(herdr, rendezvous, clock);
var waiter = rendezvous.open("term_a");
completion.register("term_a", new TurnToken("term_a", waiter)); // delivered at clock=0
clock.set(CompletionResolver.MIN_TURN_NANOS + 1_000_000); // past the too-fast floor
RecordingSessions sessions = new RecordingSessions("onTurnComplete");
TurnListener composed = Fleetd.turnListener(completion, sessions);
IllegalStateException thrown = assertThrows(IllegalStateException.class,
() -> composed.onTurnComplete("term_a"),
"the session half's throw must still escape the composed listener");
assertEquals("boom: sessions.onTurnComplete", thrown.getMessage());
assertTrue(sessions.called.contains("onTurnComplete"), "the session half must have run");
// onTurnComplete resolves off-thread (a virtual thread) — wait on the waiter itself,
// exactly like MessageServiceTest.completionFallbackIsNeverQueued does.
Rendezvous.Resolution resolution = waiter.get(5, TimeUnit.SECONDS);
assertEquals(Rendezvous.Kind.COMPLETION, resolution.kind(),
"the completion resolver must still have resolved the send, despite the session "
+ "half throwing");
assertEquals("real answer", resolution.text());
}
@Test
void onTurnFailedResolvesTheWaiterAsFailedEvenWhenTheSessionHalfThrows() throws Exception {
FakeHerdr herdr = new FakeHerdr().readText("⏺ crash context\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = newResolver(herdr, rendezvous);
var waiter = rendezvous.open("term_b");
completion.register("term_b", new TurnToken("term_b", waiter));
RecordingSessions sessions = new RecordingSessions("onTurnFailed");
TurnListener composed = Fleetd.turnListener(completion, sessions);
IllegalStateException thrown = assertThrows(IllegalStateException.class,
() -> composed.onTurnFailed("term_b"),
"the session half's throw must still escape the composed listener");
assertEquals("boom: sessions.onTurnFailed", thrown.getMessage());
assertTrue(sessions.called.contains("onTurnFailed"), "the session half must have run");
Rendezvous.Resolution resolution = waiter.get(5, TimeUnit.SECONDS);
assertEquals(Rendezvous.Kind.FAILED, resolution.kind(),
"the completion resolver must still have failed the send, despite the session "
+ "half throwing");
}
@Test
void onTurnFailedWithReasonResolvesTheWaiterEvenWhenTheSessionHalfThrows() throws Exception {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = newResolver(herdr, rendezvous);
var waiter = rendezvous.open("term_c");
completion.register("term_c", new TurnToken("term_c", waiter));
// fleetd #561: the composed listener calls sessions.onTurnFailed(target) — the ONE-arg
// overload — for this reason-carrying event too (matching the pre-existing production
// behaviour: SessionManager never overrides the two-arg overload either), so the
// throwing key here is "onTurnFailed", not a distinct "...WithReason" one.
RecordingSessions sessions = new RecordingSessions("onTurnFailed");
TurnListener composed = Fleetd.turnListener(completion, sessions);
IllegalStateException thrown = assertThrows(IllegalStateException.class,
() -> composed.onTurnFailed("term_c", "worker unreachable"),
"the session half's throw must still escape the composed listener");
assertEquals("boom: sessions.onTurnFailed", thrown.getMessage());
assertTrue(sessions.called.contains("onTurnFailed"), "the session half must have run");
Rendezvous.Resolution resolution = waiter.get(5, TimeUnit.SECONDS);
assertEquals(Rendezvous.Kind.FAILED, resolution.kind(),
"the completion resolver must still have failed the send, despite the session "
+ "half throwing");
assertEquals("worker unreachable", resolution.text(),
"the explicit reason must still reach the resolved send");
}
@Test
void onTurnCompleteWithPostActionResolvesTheWaiterEvenWhenTheSessionHalfThrows() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ answer before reset\n❯ ");
Rendezvous rendezvous = new Rendezvous();
AtomicLong clock = new AtomicLong(0);
CompletionResolver completion = newResolverPastTheFloor(herdr, rendezvous, clock);
var waiter = rendezvous.open("term_d");
completion.register("term_d", new TurnToken("term_d", waiter)); // delivered at clock=0
clock.set(CompletionResolver.MIN_TURN_NANOS + 1_000_000); // past the too-fast floor
RecordingSessions sessions = new RecordingSessions("onTurnCompleteWithPostAction");
TurnListener composed = Fleetd.turnListener(completion, sessions);
IllegalStateException thrown = assertThrows(IllegalStateException.class,
() -> composed.onTurnCompleteWithPostAction("term_d"),
"the session half's throw must still escape the composed listener");
assertEquals("boom: sessions.onTurnCompleteWithPostAction", thrown.getMessage());
assertTrue(sessions.called.contains("onTurnCompleteWithPostAction"), "the session half must have run");
// resolveBeforePostAction is synchronous by design (it must run before the context reset
// can erase the pane) — the waiter is already resolved by the time the throw propagates.
assertTrue(waiter.isDone(), "resolveBeforePostAction is synchronous — the send must "
+ "already be resolved once the composed call returns (by throwing)");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("answer before reset", waiter.getNow(null).text());
}
/**
* The mirror case (comment 17041's acceptance item 2): the completion half throws, and the
* session half still ran. The production {@link CompletionResolver} is deliberately defensive
* (its scrape reads are wrapped in {@code catch (RuntimeException)}, by the same fail-open
* design {@link CompletionResolver#captureBaseline} documents) so it essentially never throws
* synchronously in normal operation — forcing it to do so needs a genuinely-real trigger, not
* a fabricated one. {@code ConcurrentHashMap.get(null)} is that trigger: passing a {@code null}
* target makes {@code onTurnComplete}'s {@code inFlight.get(target)} throw a
* {@link NullPointerException} before it ever starts its resolving thread — a real code path,
* not a contrived one.
*/
@Test
void sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronously() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = newResolver(herdr, rendezvous);
RecordingSessions sessions = new RecordingSessions(); // throws from nothing
TurnListener composed = Fleetd.turnListener(completion, sessions);
assertThrows(NullPointerException.class, () -> composed.onTurnComplete(null),
"ConcurrentHashMap.get(null) inside CompletionResolver.onTurnComplete must still "
+ "escape the composed listener");
assertTrue(sessions.called.contains("onTurnComplete"),
"the session half must still have run even though the completion half threw first");
}
/**
* The mirror case above only exercises {@code onTurnComplete}/{@code bothMustRun}. {@code
* onTurnCompleteWithPostAction} is composed through the OTHER helper,
* {@code bothMustRunKeepingSecondResult}, and nothing previously asserted that its session half
* still runs when its completion half throws — that helper could be reverted to the pre-#561
* broken shape (run the session half only if the completion half did not throw) and the suite
* would still stay green. Uses the same real, non-fabricated trigger as the test above: a
* {@code null} target makes {@code resolveBeforePostAction}'s {@code inFlight.get(target)}
* throw a {@link NullPointerException} before {@code resolve} is ever entered.
*/
@Test
void sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronouslyForPostAction() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = newResolver(herdr, rendezvous);
RecordingSessions sessions = new RecordingSessions(); // throws from nothing
TurnListener composed = Fleetd.turnListener(completion, sessions);
assertThrows(NullPointerException.class, () -> composed.onTurnCompleteWithPostAction(null),
"ConcurrentHashMap.get(null) inside CompletionResolver.resolveBeforePostAction must "
+ "still escape the composed listener");
assertTrue(sessions.called.contains("onTurnCompleteWithPostAction"),
"the session half must still have run even though the completion half threw first");
}
/**
* {@code bothMustRun} must not just run both halves — it must not DROP a second failure when
* both halves throw. Forces the completion half to throw (the same real {@code
* inFlight.get(null)} NullPointerException trigger used above) while the session half throws a
* distinct {@link IllegalStateException}, and asserts the completion half's throwable is what
* escapes while the session half's throwable survives as a suppressed exception rather than
* being silently discarded.
*/
@Test
void bothFailuresEscapeWhenBothHalvesThrowDistinctExceptions() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = newResolver(herdr, rendezvous);
RecordingSessions sessions = new RecordingSessions("onTurnComplete");
TurnListener composed = Fleetd.turnListener(completion, sessions);
NullPointerException thrown = assertThrows(NullPointerException.class,
() -> composed.onTurnComplete(null),
"the completion half's throw (NPE from inFlight.get(null)) must be what escapes");
assertTrue(sessions.called.contains("onTurnComplete"), "the session half must still have run");
assertEquals(1, thrown.getSuppressed().length,
"the session half's distinct failure must be recorded as suppressed, not dropped");
assertEquals(IllegalStateException.class, thrown.getSuppressed()[0].getClass());
assertEquals("boom: sessions.onTurnComplete", thrown.getSuppressed()[0].getMessage());
}
}
@@ -0,0 +1,75 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #446 follow-up (criterion 2): {@code exhaustionSink} used to build the "name the fix"
* WARNING inline with two SLF4J {@code {}}-placeholder {@code log.warn} calls. Nothing pinned
* that text — a mutation battery run against the merged PR #457 proved it: renaming the leading
* {@code "usage-limit fix:"} tag to {@code "MUTANT no fix named:"} at both call sites left all
* 1586 existing tests green. This class is what makes that mutation red.
*
* <p>{@link Fleetd#usageLimitFixWarning} and {@link Fleetd#usageLimitFixWarningNoModel} are the
* extracted call sites {@code exhaustionSink} actually invokes — {@code log.warn} is called with
* each method's return value as a single, already-formatted argument, so the string these tests
* assert on is byte-identical to what {@code fleetd.out} receives. Same extracted-static-method +
* dedicated-test idiom as {@link Fleetd#exhaustedPatternCoverageLine} /
* {@link FleetdPatternCoverageLineTest}, chosen over the {@link ch.qos.logback.core.read.ListAppender}
* idiom {@code ExhaustedPatternGapReportTest} uses — this warning is built once, at one call site,
* from a single method with no branching inside the message itself, so there is no real log
* plumbing left to prove; capturing an appender would only add setup/teardown around the same
* assertion.
*
* <p>Each assertion below checks for a specific substring — the leading {@code "usage-limit fix:"}
* tag an operator would grep the log for, the profile name, the model name, and the literal
* action ({@code enabled: false} under {@code models.allow}) — rather than the whole sentence.
* Criterion 2 promised those facts, not a particular wording of the prose that carries them; a
* whole-sentence {@code assertEquals} breaks on every copy-edit of that surrounding prose and
* teaches the next person to delete the test rather than read why it failed. The leading tag is
* pinned on purpose, unlike the rest of the sentence: it is the identifying label the fleetd #446
* follow-up's own mutation battery renamed ({@code "usage-limit fix:"} → {@code "MUTANT no fix
* named:"}) to prove this class did not yet catch a tag rename — only the three-fact substrings
* above would have missed it, since none of them mention the tag itself.
*/
class FleetdUsageLimitFixWarningTest {
@Test
@DisplayName("usageLimitFixWarning names the profile, the model, and the models.allow fix")
void usageLimitFixWarningNamesProfileModelAndFix() {
String warning = Fleetd.usageLimitFixWarning("terra", "claude-opus-5");
assertTrue(warning.startsWith("usage-limit fix:"),
"must carry the grep-able identifying tag: " + warning);
assertTrue(warning.contains("terra"), "must name the profile: " + warning);
assertTrue(warning.contains("claude-opus-5"), "must name the model: " + warning);
assertTrue(warning.contains("enabled: false"), "must name the action: " + warning);
assertTrue(warning.contains("models.allow"), "must name where the action goes: " + warning);
}
@Test
@DisplayName("usageLimitFixWarning does not cross-name a different profile or model")
void usageLimitFixWarningDoesNotCrossName() {
String warning = Fleetd.usageLimitFixWarning("terra", "claude-opus-5");
// Guards against the mutation of swapping the two profile/model arguments at the call
// site — a test that only checks "some name appears" cannot catch that.
assertTrue(!warning.contains("sonnet"), "must not name an unrelated model: " + warning);
}
@Test
@DisplayName("usageLimitFixWarningNoModel names the profile but no fix, since none exists")
void usageLimitFixWarningNoModelNamesProfileNotAFix() {
String warning = Fleetd.usageLimitFixWarningNoModel("gx", 1800);
assertTrue(warning.startsWith("usage-limit fix:"),
"must carry the grep-able identifying tag: " + warning);
assertTrue(warning.contains("gx"), "must name the profile: " + warning);
assertTrue(warning.contains("1800"), "must name the quarantine duration it falls back to: " + warning);
assertTrue(!warning.contains("enabled: false"),
"a profile with no model: configured has no models.allow entry to flip — the "
+ "fallback message must not claim one exists: " + warning);
}
}
@@ -0,0 +1,241 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #377: the startup git host line must report the <em>shape</em> of the value a
* member receives as GITEA_HOST — set or unset, length, scheme, trailing slash — and the
* value itself must never reach the log.
*/
class GitHostShapeReportTest {
/** A profile that opts in to the git-forge token, so GITEA_HOST is the var that matters. */
private static final String GIT_TOKEN_CONFIG = """
profiles:
local:
baseUrl: http://gx00.gw:8000
gitTokenEnv: WORKER_GITEA_TOKEN
""";
private static FleetConfig load(Path dir, String yaml) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml);
return FleetConfig.load(f);
}
/**
* The level this logger had before {@link #attach()} raised it, so {@link #detach} can put it
* back. {@code null} is a real value here — it means "inherit from the parent" — and that is
* exactly the state this logger starts in, so it must be restored as {@code null} rather than
* as some concrete level.
*/
private static Level originalLevel;
/**
* fleetd #377: the shape lines are logged at INFO, and {@code logback-test.xml} sets
* {@code dev.ltms.fleet} to WARN — so INFO events are dropped by the level check BEFORE any
* appender sees them. Attaching an appender is therefore not enough: without raising the level
* the list stays empty and every assertion below fails against correct production code. The
* sibling report tests do the same thing at each call site (see
* {@code MemberTrustModelReportTest}); doing it here keeps it in one place.
*/
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
originalLevel = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
logger.detachAppender(appender);
logger.setLevel(originalLevel);
}
private static List<String> messages(ListAppender<ILoggingEvent> appender) {
return appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList();
}
@Test
void setLineReportsLengthSchemeAndTrailingSlash(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
String value = "https://git.example.test/";
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("GITEA_HOST", value));
} finally {
detach(appender);
}
String expected = "startup git host GITEA_HOST: set (profile 'local' gitHostEnv) — "
+ "length=" + value.length() + ", startsWithScheme=true, trailingSlash=true";
assertTrue(messages(appender).contains(expected),
"a value with a scheme and a trailing slash must be reported by shape only: "
+ "its length, startsWithScheme=true, trailingSlash=true");
}
@Test
void bareHostLineReportsNoSchemeNoTrailingSlash(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
String value = "git.example.test";
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("GITEA_HOST", value));
} finally {
detach(appender);
}
String expected = "startup git host GITEA_HOST: set (profile 'local' gitHostEnv) — "
+ "length=" + value.length() + ", startsWithScheme=false, trailingSlash=false";
assertTrue(messages(appender).contains(expected),
"a bare host name must report startsWithScheme=false, trailingSlash=false");
}
@Test
void unsetVariableIsLoggedAtInfoAndDoesNotThrow(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
ListAppender<ILoggingEvent> appender = attach();
try {
// GITEA_HOST absent from the map entirely — the unset case must not throw.
Fleetd.reportGitHostShape(cfg, Map.of());
} finally {
detach(appender);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getLevel() == Level.INFO
&& e.getFormattedMessage()
.startsWith("startup git host GITEA_HOST: unset (profile 'local' gitHostEnv)")),
"an unset git host is useful information, not an error — say it at INFO level");
}
/**
* The important test: the value must NEVER appear in the log output. This test fails if the
* line is ever changed to include the value, because it drives the real logging path with a
* value that carries a marker no shape field could contain.
*/
@Test
void theValueNeverAppearsInLogOutput(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
String marker = "never-logged-host-shape-377";
String value = "https://" + marker + "/";
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("GITEA_HOST", value));
} finally {
detach(appender);
}
List<String> msgs = messages(appender);
assertFalse(msgs.stream().anyMatch(m -> m.contains(value)),
"the full GITEA_HOST value must never reach the log");
assertFalse(msgs.stream().anyMatch(m -> m.contains(marker)),
"no fragment of the value may reach the log — a line that embeds the value"
+ " would leak at least this marker");
assertTrue(msgs.stream().anyMatch(m -> m.contains("startsWithScheme=true")
&& m.contains("trailingSlash=true")),
"the shape must still be reported although the value is not");
}
@Test
void unsetReportAlsoCarriesNoValue(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
// A blank value is not injected either (putIfPresent skips it), so it must report unset
// without echoing anything of it.
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("GITEA_HOST", " "));
} finally {
detach(appender);
}
assertTrue(messages(appender).stream().anyMatch(m ->
m.startsWith("startup git host GITEA_HOST: unset")),
"a blank value is skipped by the launcher, so the line reports unset");
}
@Test
void defaultGitHostEnvIsGITEA_HOST(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
assertEquals(Map.of("GITEA_HOST", List.of("profile 'local' gitHostEnv")),
Fleetd.gitHostEnvVars(cfg));
}
@Test
void explicitGitHostEnvReportsUnderItsOwnName(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
gitTokenEnv: WORKER_GITEA_TOKEN
gitHostEnv: MY_FORGE_HOST
""");
String value = "https://forge.example.test/";
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("MY_FORGE_HOST", value));
} finally {
detach(appender);
}
assertTrue(messages(appender).contains(
"startup git host MY_FORGE_HOST: set (profile 'local' gitHostEnv) — "
+ "length=" + value.length()
+ ", startsWithScheme=true, trailingSlash=true"),
"an explicitly named gitHostEnv is reported under that name");
}
@Test
void noGitTokenMeansNoHostLineIsNeeded(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of());
} finally {
detach(appender);
}
assertEquals(List.of("startup git host: no profile sets a gitTokenEnv — nothing to check"),
messages(appender));
}
@Test
void aHostWithAPortIsNotTreatedAsAScheme() {
assertFalse(Fleetd.startsWithScheme("git.example.test"));
assertFalse(Fleetd.startsWithScheme("git.example.test:3000"));
assertFalse(Fleetd.startsWithScheme("git.example.test:3000/"));
assertTrue(Fleetd.startsWithScheme("https://git.example.test"));
assertTrue(Fleetd.startsWithScheme("https://git.example.test:3000/"));
assertTrue(Fleetd.startsWithScheme("ssh://git.example.test"));
}
}
@@ -1,15 +1,12 @@
package dev.ltms.fleet.auth;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import static org.junit.jupiter.api.Assertions.*;
@@ -24,28 +21,21 @@ import static org.junit.jupiter.api.Assertions.*;
class AuditLogTest {
private final ObjectMapper mapper = new ObjectMapper();
private ListAppender<ILoggingEvent> appender;
private ch.qos.logback.classic.Logger auditLogger;
private CapturedLog auditLog;
@BeforeEach
void attach() {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
auditLogger = ctx.getLogger("audit");
appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
auditLogger.addAppender(appender);
auditLogger.setLevel(Level.INFO);
auditLog = CapturedLog.at("audit", Level.INFO);
}
@AfterEach
void detach() {
auditLogger.detachAppender(appender);
auditLog.close();
}
private JsonNode onlyRecord() throws Exception {
assertEquals(1, appender.list.size(), "exactly one audit line expected");
String line = appender.list.getFirst().getFormattedMessage();
assertEquals(1, auditLog.events().size(), "exactly one audit line expected");
String line = auditLog.events().getFirst().getFormattedMessage();
return mapper.readTree(line); // throws if the line is not valid JSON
}
@@ -113,6 +113,20 @@ class AuthzTest {
assertTrue(Authz.permits(WORKER_A, METRICS, null));
}
/**
* fleetd #421 — unlike READ (the case above), COORD_READ (a lead's own held peer mail) is the
* primary's alone. A worker or an architect reading it would disclose lead-to-lead
* coordination bodies, not the secret-free roster READ is open about.
*/
@Test
void coordReadIsThePrimarysAloneNotAWidenedRead() {
assertTrue(Authz.permits(PRIMARY, COORD_READ, null));
assertFalse(Authz.permits(WORKER_A, COORD_READ, null),
"a worker must not read held lead-to-lead mail");
assertFalse(Authz.permits(ARCH_DESIGN, COORD_READ, null),
"an architect holds READ today, but not-primary must mean not-architect here too");
}
@Test
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
@@ -186,6 +186,43 @@ class CallerResolverTest {
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "BEARER s3cret").role());
}
// ── fleetd #505: a herdr error DURING THE SCAN must not be conflated with "not a worker" ──────
// #317 (above) covers a failed lsof lookup. This is the other input to the same decision: the
// lsof lookup succeeds (a real pid), but PaneLocator's own pane scan hits a herdr error on the
// pane that owns that pid — so c.resolved() is true and c.terminal() is null, exactly like a
// real primary. c.scanComplete() is what tells them apart.
/**
* The discriminating case named in the ticket: the error must land on the pane that DOES own
* the caller's pid, or the test proves nothing (any other pane's failure is invisible to the
* scan's outcome, since a match found elsewhere is definitive regardless).
*/
@Test
void aHerdrErrorOnTheOwningPaneDuringTheScanIsRefusedNotPromotedToPrimary() {
FakeHerdr failing = new FakeHerdr().processInfoFailsForPane("w2:p7", "transient");
ConnectionIdentity incomplete = new ConnectionIdentity(new PaneLocator(failing), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(incomplete).resolve("127.0.0.1", 55555, null);
assertEquals(Role.ANONYMOUS, p.role(),
"an incomplete pane scan must never be read as a clean negative and promoted to primary");
}
/**
* The companion invariant: a herdr error on a DIFFERENT, non-owning pane must not turn every
* mid-scan teardown into a refusal — the real match is still found and resolves as a worker.
*/
@Test
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealWorker() {
FakeHerdr vanishedElsewhere = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(vanishedElsewhere), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(id).resolve("127.0.0.1", 55555, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
@Test
void aNonLoopbackCallerIsNeverThePrimaryUnderLoopbackTrust() {
// Defence in depth: startup already refuses this pairing (validateAuthExposure), but if a
@@ -0,0 +1,360 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.peer.MemberRole;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* fleetd #424 — revoking (or granting) an architect slot must take effect on the next spawn with
* no restart. A session already bound to a slot keeps its <em>binding</em> (the {@code
* terminalToSlot} occupancy) even after that slot drops out of config, but NOT the ARCHITECT
* <em>privilege</em> the slot used to grant — that is revoked on the bound session's very next
* request. See {@link MemberRegistry}'s class doc for the exact rule: "config governs what a
* bound slot still grants, as well as what may be bound next."
*
* <p>Every test here drives a REAL {@link ConfigRef#reload()} against a {@code @TempDir} file and
* asserts {@link ConfigRef.Outcome#applied()}, rather than comparing two frozen
* {@code MemberRegistry} instances in memory — the defect this ticket fixes is specifically that
* {@link MemberRegistry} used to ignore a live reload, so a test that never reloads cannot tell the
* fixed registry from the broken one. {@code requireSlotFor} and {@code reserve} are pinned in
* separate tests, in both directions (removed and added), so a registry that simply refuses (or
* simply allows) everything cannot pass by accident — see {@link MemberRegistryTest} for the
* registry's other invariants (bind/unbind cardinality, thread-safety), which are unaffected by
* this ticket and still exercised against the frozen constructor.
*/
class MemberRegistryLiveTest {
private static String yaml(String fleetBlock) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
opus:
baseUrl: http://gx00.gw:8001
model: opus
guard:
offSubscriptionHosts:
- gx00.gw
""" + fleetBlock;
}
private static final String WITH_SONNET_SLOT = """
fleet:
architects:
designer:
profile: sonnet
""";
/** No architect pool at all — developers is unrelated dead data for this registry (#424 out of scope). */
private static final String WITHOUT_ARCHITECT_SLOTS = """
fleet:
developers:
dev1:
profile: sonnet
""";
/** Same slot name ({@code designer}) as {@link #WITH_SONNET_SLOT}, repointed to a different profile. */
private static final String WITH_OPUS_SLOT = """
fleet:
architects:
designer:
profile: opus
""";
/** Same profile ({@code sonnet}) as {@link #WITH_SONNET_SLOT}, but the pool key is renamed. */
private static final String WITH_RENAMED_SLOT = """
fleet:
architects:
architect-lead:
profile: sonnet
""";
private static ConfigRef refFor(Path f) {
return new ConfigRef(f, FleetConfig.load(f));
}
// ── requireSlotFor is live (criteria 1, 2, 4) ──────────────────────────────────────────────
@Test
void requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"the slot is configured before the reload");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertThrows(IllegalArgumentException.class,
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"revoking the slot must refuse the NEXT spawn that names it");
}
@Test
void requireSlotForAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertThrows(IllegalArgumentException.class,
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"a slot added by reload must be usable with no restart");
}
// ── reserve is live too — tested separately from requireSlotFor (criteria 1, 2, 4) ────────
@Test
void reserveRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation before = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertEquals("architect:designer", before.slot());
registry.release(before); // free it back up so the reload-side reserve below starts clean
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"revoking the slot must refuse the NEXT reservation for it");
}
@Test
void reserveAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
MemberLifecycle.SlotReservation after = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertEquals("architect:designer", after.slot(),
"a slot added by reload must be reservable with no restart");
}
// ── profileForSlot, isSlot and nameForSlot are live too (fleetd #431) ─────────────────────────
// #424 pinned roleForSlot against a live reload but left these three untested — proved by
// mutating each to read a snapshot flattened once at construction: the full suite stayed green
// for all three.
//
// The three differ in how much production behaviour depends on them, and the ticket first got
// this ranking wrong. nameForSlot is wired: CallerResolver passes members::nameForSlot, next to
// members::roleForSlot. isSlot is reached through bind, which calls it to refuse an unknown
// slot. profileForSlot has NO caller in src/main at all — grepped both the ".profileForSlot("
// and the "::profileForSlot" form — so there is no seam to drive its test through beyond the
// accessor itself, and its own javadoc ("what the spawn lifecycle reads") describes a caller
// that does not exist. These tests pin the accessors as they are; whether profileForSlot should
// be wired or deleted is a separate question.
@Test
void profileForSlotReflectsAProfileChangedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertEquals("sonnet", registry.profileForSlot("architect:designer"),
"the slot's profile before the reload");
Files.writeString(f, yaml(WITH_OPUS_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertEquals("opus", registry.profileForSlot("architect:designer"),
"repointing the slot to a different profile must take effect with no restart");
}
@Test
void isSlotStopsReportingASlotRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertTrue(registry.isSlot("architect:designer"), "the slot is configured before the reload");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertFalse(registry.isSlot("architect:designer"),
"removing the slot from config must make isSlot say so on the very next call, "
+ "with no restart");
}
@Test
void isSlotStartsReportingASlotAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertFalse(registry.isSlot("architect:designer"), "no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertTrue(registry.isSlot("architect:designer"),
"a slot added by reload must be visible to isSlot with no restart");
}
@Test
void nameForSlotReflectsANameChangedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertEquals("designer", registry.nameForSlot("architect:designer"),
"the configured name before the reload");
Files.writeString(f, yaml(WITH_RENAMED_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertNull(registry.nameForSlot("architect:designer"),
"the old key no longer names a configured slot — it was renamed away by the reload");
assertEquals("architect-lead", registry.nameForSlot("architect:architect-lead"),
"the new name must be visible under its new qualified key with no restart");
}
// ── a bound architect is demoted, but the binding itself is not touched (fleetd #424) ───────
// The lead's corrected ruling: the PRIVILEGE a slot grants is revoked on the bound session's
// very next request, but the terminalToSlot BINDING itself is untouched by a reload — dropping
// it would double-book the slot key and break unbind's compare-safe contract. See the class
// doc's binding rule.
/** A caller identity resolving the one canned pane (terminal {@code term_a}) in {@link FakeHerdr}. */
private static ConnectionIdentity boundPaneIdentity() {
return new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> FakeHerdr.WORKER_PID);
}
@Test
void anArchitectAlreadyBoundToASlotIsDemotedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
assertEquals("architect:designer", registry.slotForTerminal("term_a"));
// Drive the real caller path, not the roleForSlot seam directly: CallerResolver.resolve is
// what a live request actually goes through (CallerResolver.java:220), and a resolver that
// ignored roleForSlot entirely would still pass a test that only checked the seam.
CallerResolver resolver = CallerResolver.withLeadsAndMembers(
boundPaneIdentity(), false, null, Map::of, registry);
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, before.role(), "sanity check: the harness binds term_a as an architect");
assertEquals("designer", before.name());
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
Principal after = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, after.role(),
"removing the slot from config must demote the bound session to worker on its "
+ "NEXT request — this is the ticket's whole point");
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
}
@Test
void theOriginalBindingStillOccupiesTheRemovedSlotSoASecondTerminalCannotClaimIt(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome removed = ref.reload();
assertTrue(removed.applied(), "the reload must actually take effect: " + removed.summary());
// The binding survives the removal untouched.
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
"a live binding must never be retroactively unbound by a config edit");
assertEquals(Map.of("term_a", "architect:designer"), registry.snapshot());
// Bring the slot back into config. If the binding had been silently dropped by the removal
// (rather than merely losing the privilege it grants), a second terminal could now claim
// the "freed" key — the exact double-booking the class doc's binding rule rules out.
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome restored = ref.reload();
assertTrue(restored.applied(), "the reload must actually take effect: " + restored.summary());
assertFalse(registry.bind("architect:designer", "term_b"),
"the slot is still occupied by term_a — a second terminal must not bind to it");
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"the slot is still occupied by term_a — a fresh reservation must not find it free");
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
"the original binding is unchanged throughout");
}
@Test
void unbindStillSucceedsForTheOriginalTerminalAfterItsSlotIsRemoved(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertTrue(registry.unbind("architect:designer", "term_a"),
"unbind must still work for a slot that config has since removed, or a session "
+ "that outlives its slot's removal could never release it");
assertNull(registry.slotForTerminal("term_a"));
assertEquals(Map.of(), registry.snapshot());
}
}
@@ -0,0 +1,210 @@
package dev.ltms.fleet.config;
import org.junit.jupiter.api.Test;
import java.lang.reflect.Constructor;
import java.lang.reflect.RecordComponent;
import java.util.Arrays;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.TreeSet;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #323: {@code ConfigRef.sameLaunchSettings} javadoc used to claim it "compares every
* component the launcher reads at spawn". It did not — {@code ideProjectDir}, {@code
* ideOpenCommand} and {@code autoCompactWindow} were all baked in at daemon startup (see
* {@code ClaudeCodeLauncher}/{@code OpenCodeLauncher}) and missing from the comparison, so a reload
* that changed only one of them reported "config reloaded" and the running daemon kept the old
* value.
*
* <p>This class is the mechanism the issue asked for: it enumerates every record component of
* {@link FleetConfig.Profile} by reflection and proves — by actually mutating a base profile one
* field at a time and calling the real method — that each component is either compared by
* {@link ConfigRef#sameLaunchSettings} or named in {@link ConfigRef#LAUNCH_SETTINGS_EXCLUDED} with
* a reason. A new profile field that is neither fails this test by name, not a hand-maintained list
* going stale.
*/
class ConfigRefProfileCoverageTest {
private static final RecordComponent[] COMPONENTS = FleetConfig.Profile.class.getRecordComponents();
/**
* One valid, non-blank value per record component — "the a value". None of these trip any
* defaulting/normalization in {@code Profile}'s compact constructor (see {@code
* FleetConfig.java}), so what goes in is what {@code sameLaunchSettings} sees back out.
*/
private static final Map<String, Object> BASE = baseValues();
/** The same shape, each value distinct from {@link #BASE} — "the b value". */
private static final Map<String, Object> ALT = altValues();
private static Map<String, Object> baseValues() {
Map<String, Object> v = new LinkedHashMap<>();
v.put("profile", "sonnet");
v.put("baseUrl", "http://gx00.gw:8000");
v.put("model", "sonnet");
v.put("configDir", "/config/a");
v.put("tokenEnv", "TOKEN_A");
v.put("argv", List.of("claude", "--flag-a"));
v.put("placement", "tab");
v.put("workspace", "workspace-a");
v.put("tabLabel", "label-a");
v.put("mcpUrl", "http://mcp-a");
v.put("cwd", "/cwd/a");
v.put("parityOverlay", List.of(".env", ".env.a"));
v.put("gitTokenEnv", "GIT_TOKEN_A");
v.put("gitHostEnv", "GITEA_HOST_A");
v.put("kind", "claude-code");
v.put("env", Map.of("K", "A"));
v.put("weight", 1.0f);
v.put("maxLoad", 5);
v.put("subscription", Boolean.TRUE);
v.put("exhaustedPattern", "usage limit a");
v.put("credentialId", "cred-a");
v.put("ideMcpUrl", "http://ide-mcp-a");
v.put("ideProjectDir", "modules/a");
v.put("ideOpenCommand", "open-cmd-a {dir}");
v.put("autoCompactWindow", 150000);
v.put("errorPattern", "error a");
assertNamesMatchComponents(v);
return v;
}
private static Map<String, Object> altValues() {
Map<String, Object> v = new LinkedHashMap<>();
v.put("profile", "sonnet-b");
v.put("baseUrl", "http://gx01.gw:8000");
v.put("model", "haiku");
v.put("configDir", "/config/b");
v.put("tokenEnv", "TOKEN_B");
v.put("argv", List.of("claude", "--flag-b"));
v.put("placement", "weighted");
v.put("workspace", "workspace-b");
v.put("tabLabel", "label-b");
v.put("mcpUrl", "http://mcp-b");
v.put("cwd", "/cwd/b");
v.put("parityOverlay", List.of(".env", ".env.b"));
v.put("gitTokenEnv", "GIT_TOKEN_B");
v.put("gitHostEnv", "GITEA_HOST_B");
v.put("kind", "opencode");
v.put("env", Map.of("K", "B"));
v.put("weight", 2.0f);
v.put("maxLoad", 9);
v.put("subscription", Boolean.FALSE);
v.put("exhaustedPattern", "usage limit b");
v.put("credentialId", "cred-b");
v.put("ideMcpUrl", "http://ide-mcp-b");
v.put("ideProjectDir", "modules/b");
v.put("ideOpenCommand", "open-cmd-b {dir}");
v.put("autoCompactWindow", 250000);
v.put("errorPattern", "error b");
assertNamesMatchComponents(v);
return v;
}
private static void assertNamesMatchComponents(Map<String, Object> values) {
Set<String> componentNames = new TreeSet<>();
for (RecordComponent rc : COMPONENTS) {
componentNames.add(rc.getName());
}
assertEquals(componentNames, new TreeSet<>(values.keySet()),
"this test's value map has drifted from FleetConfig.Profile's actual components — "
+ "update BASE/ALT alongside the record");
}
private static FleetConfig.Profile profileOf(Map<String, Object> values) throws ReflectiveOperationException {
Class<?>[] types = Arrays.stream(COMPONENTS).map(RecordComponent::getType).toArray(Class<?>[]::new);
Object[] args = Arrays.stream(COMPONENTS)
.map(rc -> values.get(rc.getName()))
.toArray();
Constructor<FleetConfig.Profile> ctor = FleetConfig.Profile.class.getDeclaredConstructor(types);
return ctor.newInstance(args);
}
/** {@code BASE} with exactly one named component swapped for its {@code ALT} value. */
private static FleetConfig.Profile mutate(String componentName) throws ReflectiveOperationException {
Map<String, Object> values = new LinkedHashMap<>(BASE);
values.put(componentName, ALT.get(componentName));
return profileOf(values);
}
/**
* The mechanism fleetd #323 asked for: enumerate {@link FleetConfig.Profile}'s record
* components, mutate each non-excluded one, and prove {@code sameLaunchSettings} actually
* notices — not just that some hand-maintained list claims it does. Prints the denominator
* (total / compared / excluded) the issue required: a checker that cannot state its own
* denominator is the failure this repo keeps hitting.
*/
@Test
void sameLaunchSettingsComparesEveryProfileComponentOrExcludesIt() throws ReflectiveOperationException {
int total = COMPONENTS.length;
Set<String> excluded = ConfigRef.LAUNCH_SETTINGS_EXCLUDED;
Set<String> allNames = new TreeSet<>();
for (RecordComponent rc : COMPONENTS) {
allNames.add(rc.getName());
}
assertTrue(allNames.containsAll(excluded),
"ConfigRef.LAUNCH_SETTINGS_EXCLUDED names a component that does not exist on "
+ "FleetConfig.Profile — check for a typo: " + excluded);
// The exclusion set is this mechanism's own escape hatch, so it has to be pinned too.
// Found by mutation while verifying fleetd #323: moving autoCompactWindow and
// ideOpenCommand OUT of the comparison and INTO the exclusion set left the whole suite
// green — the loop below simply skips them, and the denominator assertion still balances.
// That is exactly the lazy move a failing coverage test invites, and it silently restores
// the #323 bug. Only ideProjectDir and worktreeGroup were saved by a behavioural test in
// ConfigRefTest; the other two had none. So: growing this set now requires editing this
// line as well, which is a visible, deliberate diff rather than a quiet one.
//
// fleetd #446: "exhaustedPattern" joined the set legitimately — it moved from compiled
// once at Fleetd.main startup to read live, cached by profile name, through
// LiveExhaustedPatterns (see that class's doc and ConfigRef's class doc). Unlike the #323
// hazard this comment warns about, it is pinned by its OWN behavioural tests:
// FleetdExhaustionDetectionArmedWiringTest proves the reload takes effect with no restart,
// and the same test rerun against the pre-#446 compile site (see that ticket's report) is
// what proves the assertion would have failed before the fix.
assertEquals(Set.of("weight", "maxLoad", "credentialId", "exhaustedPattern"), excluded,
"ConfigRef.LAUNCH_SETTINGS_EXCLUDED changed. A component belongs in it ONLY if it "
+ "is read live off the config supplier, not baked into a launcher at "
+ "startup. If you are adding one to silence this test, that is fleetd #323 "
+ "happening again: compare it in sameLaunchSettings instead. If it really "
+ "is read live, name where it is read and update this assertion.");
FleetConfig.Profile base = profileOf(BASE);
List<String> uncovered = new java.util.ArrayList<>();
int compared = 0;
for (RecordComponent rc : COMPONENTS) {
String name = rc.getName();
if (excluded.contains(name)) {
continue;
}
FleetConfig.Profile mutated = mutate(name);
if (ConfigRef.sameLaunchSettings(base, mutated)) {
uncovered.add(name);
} else {
compared++;
}
}
System.out.printf(
"ConfigRef.sameLaunchSettings coverage — %d Profile components total, %d compared, "
+ "%d excluded (%s)%n",
total, compared, excluded.size(), excluded);
assertEquals(List.of(), uncovered,
"these FleetConfig.Profile components changed but ConfigRef.sameLaunchSettings "
+ "reported no difference — add each one to the comparison (it is read at "
+ "spawn and baked in until a restart) or to ConfigRef.LAUNCH_SETTINGS_EXCLUDED "
+ "with a reason it is genuinely read live: " + uncovered);
assertEquals(total, compared + excluded.size(),
"every FleetConfig.Profile record component must be either compared or excluded — "
+ total + " components, " + compared + " compared, " + excluded.size()
+ " excluded");
}
}
@@ -1,10 +1,13 @@
package dev.ltms.fleet.config;
import dev.ltms.fleet.mcp.CharterToolSurface;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.function.Consumer;
import static org.junit.jupiter.api.Assertions.*;
@@ -38,6 +41,23 @@ class ConfigRefTest {
return new ConfigRef(f, FleetConfig.load(f));
}
/**
* The exact {@code Consumer<FleetConfig>} {@code Fleetd.main} wires into {@code ConfigRef}'s
* constructor as {@code extraValidation} (fleetd #474) — an adapter from {@code FleetConfig} to
* the raw charter map {@link CharterToolSurface#assertChartersNameOnlyRegisteredTools} takes.
* Built here rather than referencing {@code dev.ltms.fleet.Fleetd} directly, because that method
* is package-private to {@code dev.ltms.fleet} and this test lives in {@code
* dev.ltms.fleet.config} — but it calls the SAME production {@link CharterToolSurface} method
* {@code Fleetd} calls, so this proves the real check runs on reload, not a stand-in for it.
*/
private static final Consumer<FleetConfig> CHARTER_TOOL_SURFACE = cfg ->
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
private static ConfigRef refForWithCharterToolSurface(Path f) {
return new ConfigRef(f, FleetConfig.load(f), CHARTER_TOOL_SURFACE);
}
@Test
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
@@ -124,6 +144,87 @@ class ConfigRefTest {
assertSame(before, ref.get());
}
/**
* fleetd #474: {@code validateAll()} (via {@code validateCharters()}) only checks that a
* charter's KEY is a role wire name and its text is non-blank — it never looks at what the text
* names, so a charter naming {@code bridge_send} (the pre-CB-634 name, removed from the tool
* surface — #469's own motivating example) passes {@code validateAll()} and used to be applied
* on reload with nothing refusing it, even though the identical charter refuses {@code
* Fleetd.main} at startup ({@code FleetdStartupValidationTest
* #mainRefusesACharterNamingAnUnregisteredTool}). This is the reload-path proof: it drives
* {@link ConfigRef#reload()} itself (not a direct call to {@link
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), through the exact {@code
* extraValidation} wiring {@code Fleetd.main} uses, and the failure message must name both the
* charter key and the unknown tool — the same information the startup failure gives (ticket
* acceptance criterion 1).
*/
@Test
void aReloadRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
"""));
ConfigRef ref = refForWithCharterToolSurface(f);
FleetConfig before = ref.get();
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
"""));
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.error().contains("dev"),
"expected the charter key 'dev' in the refusal, got: " + out.error());
assertTrue(out.error().contains("bridge_send"),
"expected the unknown tool 'bridge_send' in the refusal, got: " + out.error());
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
// The running config must not move at all — half-applying this would leave the daemon in a
// state that could never have booted, exactly the outcome ConfigRef.java's class doc warns
// a cold-key refusal must avoid, and this check must avoid the same way.
assertSame(before, ref.get());
assertEquals("Send the final handoff through fleet_reply.\n",
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
}
/**
* fleetd #474 acceptance criterion 3, the positive case: a reload whose charter names only
* registered tools must still be ACCEPTED. An inverted filter (one that refuses every charter,
* or refuses on any {@code fleet_*}/{@code bridge_*} token regardless of registration) would pass
* the refusal test above alone — #469's own M5 mutation cell showed exactly that shape surviving
* a negative-only suite. This is the test that catches it.
*/
@Test
void aReloadAcceptsACharterNamingOnlyRegisteredTools(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
dev: old charter, no tool names
"""));
ConfigRef ref = refForWithCharterToolSurface(f);
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through fleet_reply, using fleet_send to delegate.
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals("config reloaded", out.summary());
assertEquals("Send the final handoff through fleet_reply, using fleet_send to delegate.\n",
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
}
/**
* The point of the whole class: a consumer holding the ref sees the new value without being
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
@@ -339,12 +440,18 @@ class ConfigRefTest {
}
/**
* CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map at startup
* (see ExhaustedPatternLookup), so a reload never re-reads it — a changed pattern must be
* reported deferred exactly like model/baseUrl, not silently claimed as applied.
* fleetd #446: exhaustedPattern moved from deferred to hot — {@code LiveExhaustedPatterns}
* reads {@code config.get()} fresh on every lookup, cached by profile name (not baked into a
* frozen map inside {@code Fleetd.main} any more, the way it was before this ticket, and the
* way its sibling {@code errorPattern} still is — see
* {@code changingAProfilesErrorPatternIsReportedAsDeferred} right below for the contrast).
* So a reload that changes ONLY exhaustedPattern must report a plain "config reloaded", with
* nothing deferred — the opposite of what this test asserted before fleetd #446, when
* ConfigRef.sameLaunchSettings still compared exhaustedPattern and this reload reported it as
* one of the changes needing a restart.
*/
@Test
void changingAProfilesExhaustedPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
void changingAProfilesExhaustedPatternTakesEffectWithNoRestart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
@@ -379,17 +486,21 @@ class ConfigRefTest {
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size(), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals(0, out.deferred().size(),
"exhaustedPattern is hot since fleetd #446 — changing only this key must not be "
+ "reported as needing a restart: " + out.deferred());
assertEquals(0, out.split().size(), out.split().toString());
assertEquals("config reloaded", out.summary());
// The snapshot carries the new value immediately — this IS the live value now, not
// merely what a restart would eventually pick up.
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
}
/**
* fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error pattern
* map at startup (see BackendErrorPatternLookup), the same way exhaustedPattern is above — a
* reload never re-reads it either, so a changed value must be reported deferred.
* map at startup (see BackendErrorPatternLookup) — unlike its sibling exhaustedPattern above,
* which fleetd #446 made hot, errorPattern was scoped OUT of that ticket on purpose and stays
* deferred: a reload still never re-reads it, so a changed value must be reported deferred.
*/
@Test
void changingAProfilesErrorPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
@@ -434,6 +545,364 @@ class ConfigRefTest {
assertEquals("provider 5xx", ref.get().profiles().get("sonnet").errorPattern());
}
/**
* fleetd #323 instance 1: {@code ideProjectDir} is read at spawn off the frozen profile map
* (see {@code ClaudeCodeLauncher}/{@code OpenCodeLauncher}) exactly like {@code model}, but was
* missing from {@code sameLaunchSettings} — a reload changing only this field used to report a
* bare "config reloaded" and the running daemon kept launching with the old value.
*/
@Test
void changingAProfilesIdeProjectDirIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
ideProjectDir: fleetd
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
ideProjectDir: fleetd-renamed
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size(), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals("fleetd-renamed", ref.get().profiles().get("sonnet").ideProjectDir());
}
/**
* fleetd #323 instance 2: {@code worktreeGroup} is baked into the same {@code GitWorktrees}
* as {@code worktreeRoot} (Fleetd.java:251) and never rebuilt, but only {@code worktreeRoot}
* was on {@code changedDeferredKeys} — a reload changing only the group reported a bare
* "config reloaded" and newly provisioned worktrees kept the old sharing behaviour.
*/
@Test
void changingWorktreeGroupIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("worktreeGroup: devgroup\n"));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("worktreeGroup: devgroup2\n"));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(java.util.List.of("worktreeGroup"), out.deferred());
assertTrue(out.summary().contains("needs a restart") || out.summary().contains("need a restart"),
out.summary());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals("devgroup2", ref.get().worktreeGroup());
}
/**
* fleetd #326: {@code primary} is read only off the startup snapshot — {@code Fleetd.java:506,
* 519, 520} feed {@code PrimaryRegistry} and {@code ReplyPushLoop} at construction and neither is
* rebuilt on reload — but it was missing from {@link ConfigRef#changedDeferredKeys}, so a reload
* that only changed the pinned primary terminal reported a bare "config reloaded" while a lead
* whose tab no longer matched stayed demoted to worker.
*/
@Test
void changingPrimaryIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
primary:
terminal: term-a
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
primary:
terminal: term-b
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(java.util.List.of("primary"), out.deferred());
assertTrue(out.summary().contains("need") && out.summary().contains("restart"), out.summary());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals("term-b", ref.get().primary().terminal());
}
/**
* fleetd #326: {@code configReload} itself is read only at startup ({@code Fleetd.java:679-680})
* to decide whether to build a {@code ConfigWatcher} at all, and with what interval — the watcher
* that would apply a later change is itself built once, so it is deferred rather than cold (see
* {@link ConfigRef}'s class doc: cold means an already-open resource would go inconsistent with
* the new value, and there is no such resource here — a running watcher just keeps polling on its
* original enabled/interval until a restart, exactly like {@code lifecycle} or {@code guard}).
* Before this fix, turning reload off (or changing its interval) through a reload reported a bare
* "config reloaded" — the obvious joke the issue names.
*/
@Test
void changingConfigReloadIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
configReload:
enabled: true
intervalSeconds: 10
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
configReload:
enabled: false
intervalSeconds: 30
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(java.util.List.of("configReload"), out.deferred());
assertTrue(out.summary().contains("need") && out.summary().contains("restart"), out.summary());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertFalse(ref.get().configReload().isEnabled());
assertEquals(30, ref.get().configReload().intervalSeconds());
}
/**
* fleetd #330: {@code health:} is read both ways — {@code Fleetd.java:556-563} builds the
* monitor off the startup snapshot and never rebuilds it, but {@code Fleetd.java:648-650} reads
* {@code config.get().health()} live on every {@code fleet_profiles} call. A changed value is
* neither purely hot nor purely deferred, so it gets its own {@code split} report naming both
* halves rather than a bare "config reloaded" (which would hide the frozen half) or a plain
* {@code deferred} entry (which would hide that the coverage string already applied).
*/
@Test
void changingHealthIsReportedAsSplit(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
health:
enabled: true
intervalSeconds: 30
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
health:
enabled: true
intervalSeconds: 90
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals(1, out.split().size(), out.split().toString());
assertTrue(out.split().getFirst().startsWith("health:"), out.split().toString());
assertTrue(out.split().getFirst().contains("restart"), out.split().toString());
assertTrue(out.split().getFirst().contains("live"), out.split().toString());
assertTrue(out.summary().contains("partially live"), out.summary());
// The snapshot still carries the new value — the monitor itself is what waits for a restart.
assertEquals(90, ref.get().health().intervalSeconds());
}
/**
* fleetd #330: {@code coordinator:} is the other split key — {@code Fleetd.java:502} opens the
* {@code LeadMailbox} off the startup snapshot and never reopens it, but
* {@code HerdrPeerLauncher.java:1530} reads {@code config.get().coordinator()} live on every
* spawn to keep the broker URI env-var name out of a member's environment.
*/
@Test
void changingCoordinatorIsReportedAsSplit(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
coordinator:
selfId: mac-a
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
coordinator:
selfId: mac-b
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals(1, out.split().size(), out.split().toString());
assertTrue(out.split().getFirst().startsWith("coordinator:"), out.split().toString());
assertTrue(out.split().getFirst().contains("restart"), out.split().toString());
assertTrue(out.split().getFirst().contains("live"), out.split().toString());
// The snapshot still carries the new value — the LeadMailbox connection is what waits for a
// restart; selfId names this daemon's own inbox queue and a peer cannot discover a rename.
assertEquals("mac-b", ref.get().coordinator().selfId());
}
/**
* fleetd #333: {@code fleet:} is split too — {@code fleet.leaders} is read only at
* {@code Fleetd.java:281} to build the {@code LeadTabScanner}'s identity map (fed into
* {@code CallerResolver}) and to auto-launch leads via {@code LeadLauncher.ensureLeads()}, and
* neither is rebuilt on reload, while the rest of {@code fleet:} (role pools, charters,
* tabLabel) is read live through the {@code CompositePeerLauncher} supplier. Before this fix
* {@code fleet} sat in the coverage checker's hot-excluded escape hatch, so a reload that
* changed only {@code fleet.leaders} reported a bare "config reloaded" — exactly the
* under-claim the split class exists to prevent for {@code health}/{@code coordinator}.
*/
@Test
void changingFleetLeadersIsReportedAsSplit(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
leaders:
opus:
tab: "lead: opus-a"
profile: sonnet
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
fleet:
leaders:
opus:
tab: "lead: opus-b"
profile: sonnet
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals(1, out.split().size(), out.split().toString());
assertTrue(out.split().getFirst().startsWith("fleet:"), out.split().toString());
assertTrue(out.split().getFirst().contains("restart"), out.split().toString());
assertTrue(out.split().getFirst().contains("live"), out.split().toString());
assertTrue(out.summary().contains("partially live"), out.summary());
// The snapshot still carries the new value — the LeadTabScanner's identity map and
// LeadLauncher's auto-launch are what wait for a restart; a reload rebuilds neither.
assertEquals("lead: opus-b", ref.get().fleet().leaders().get("opus").tab());
}
/**
* fleetd #333: {@code fleet.leaders} is the ONLY frozen part of {@code fleet:}. A reload that
* changes {@code tabLabel} (or charters, or a role pool) without touching {@code fleet.leaders}
* must stay fully hot with nothing reported — proving {@link ConfigRef#changedSplitKeys}
* compares {@code fleet.leaders} specifically rather than the whole {@code Fleet} record, which
* would over-claim "needs a restart" for a change that is genuinely all live (the mirror
* mistake of the under-claim this class exists to prevent).
*/
@Test
void changingFleetTabLabelWithoutLeadersStaysFullyHot(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
tabLabel: "{role}: {profile} #{n}"
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
fleet:
tabLabel: "[{profile}] {role}"
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertTrue(out.split().isEmpty(), out.split().toString());
assertEquals("config reloaded", out.summary());
assertEquals("[{profile}] {role}", ref.get().fleet().tabLabel());
}
/**
* A split change must not refuse the reload (invariant 2 of fleetd #330) and
* {@code Outcome.applied()} must stay {@code true} (invariant 3) — unlike a cold change, the
* live half of a split key genuinely took effect, so refusing would throw that away.
*/
@Test
void aSplitChangeDoesNotRefuseTheReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("coordinator:\n selfId: mac-a\n"));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("coordinator:\n selfId: mac-b\n"));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.error() == null);
assertTrue(out.coldKeys().isEmpty());
}
/**
* Both split keys can change in one reload — the report names both, and a caller reading
* {@code split} does not have to guess which half of which key already applied.
*/
@Test
void changingBothSplitKeysReportsBoth(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
health:
enabled: true
coordinator:
selfId: mac-a
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
health:
enabled: false
coordinator:
selfId: mac-b
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(2, out.split().size(), out.split().toString());
assertTrue(out.split().stream().anyMatch(s -> s.startsWith("health:")), out.split().toString());
assertTrue(out.split().stream().anyMatch(s -> s.startsWith("coordinator:")), out.split().toString());
}
/**
* A split key and a deferred key changing in the same reload must both show up, each in its own
* list — proving the two fields do not step on each other and {@link ConfigRef.Outcome#summary()}
* reports both halves of the message.
*/
@Test
void aSplitChangeAndADeferredChangeCoexist(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
coordinator:
selfId: mac-a
lifecycle:
drainTimeoutSeconds: 30
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
coordinator:
selfId: mac-b
lifecycle:
drainTimeoutSeconds: 60
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(java.util.List.of("lifecycle"), out.deferred());
assertEquals(1, out.split().size(), out.split().toString());
assertTrue(out.split().getFirst().startsWith("coordinator:"), out.split().toString());
assertTrue(out.summary().contains("need a restart") || out.summary().contains("needs a restart"),
out.summary());
assertTrue(out.summary().contains("partially live"), out.summary());
}
@Test
void aFixedRefHasNoFileAndRefusesToReload() {
FleetConfig cfg = new FleetConfig(null, null, null, null, null, null,
@@ -0,0 +1,190 @@
package dev.ltms.fleet.config;
import org.junit.jupiter.api.Test;
import java.lang.reflect.RecordComponent;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Set;
import java.util.TreeSet;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #330: a new top-level {@link FleetConfig} record component must be triaged into a reload
* class before it ships, or it repeats fleetd #323 ({@code worktreeGroup} missing from {@code
* ConfigRef.changedDeferredKeys}) and fleetd #326 ({@code primary}/{@code configReload} missing the
* same way) — a key silently absent from {@link ConfigRef}'s reload machinery, so a reload changing
* only that key reports a bare "config reloaded" for a change the running daemon never picked up.
*
* <p>This is the top-level counterpart of {@link ConfigRefProfileCoverageTest}: instead of
* enumerating {@link FleetConfig.Profile}'s record components, it enumerates {@link FleetConfig}'s
* own — {@code bind}, {@code health}, {@code memberCredentials}, and so on — and requires each to
* fall into exactly one of four homes: {@link ConfigRef#COLD_KEYS}, {@link #DEFERRED_TOP_LEVEL_KEYS}
* (compared in {@code ConfigRef.changedDeferredKeys}), {@link ConfigRef#SPLIT_KEYS}, or
* {@link #HOT_EXCLUDED_TOP_LEVEL_KEYS} (the escape hatch: read live off the config supplier, so no
* reload bookkeeping is needed for it at all).
*
* <h2>What this checker can and cannot prove</h2>
* It proves the record's <em>shape</em> is fully triaged: every one of {@code FleetConfig}'s
* components sits in exactly one of the four sets, none sits in two, and the escape hatch
* ({@link #HOT_EXCLUDED_TOP_LEVEL_KEYS}) cannot silently grow without a visible diff to this file.
* That is what "a new component cannot be added without someone triaging it" means in practice.
*
* <p>It CANNOT prove that any of the citations are <em>true</em>. "Compared in {@code
* changedDeferredKeys}" and "read live off {@code config.get()}" are facts about {@code
* ConfigRef.java}, {@code Fleetd.java} and {@code HerdrPeerLauncher.java} that a reflection-only
* test over {@code FleetConfig}'s shape has no way to inspect — this test would pass identically
* whether or not the cited line still does what the comment next to it says. Trust the citation
* because a person read the source (the exact call sites are named next to each set below), not
* because this test is green. {@link ConfigRefTest} is what behaviourally proves the deferred and
* split keys it covers actually get reported; {@link ConfigRefProfileCoverageTest} does the same,
* behaviourally, for {@code FleetConfig.Profile}'s own fields.
*/
class ConfigRefTopLevelCoverageTest {
private static final RecordComponent[] COMPONENTS = FleetConfig.class.getRecordComponents();
/**
* Top-level components whose change {@code ConfigRef.changedDeferredKeys} reads and reports on.
* Fleetd #337 promoted this out of a hand-maintained copy here into {@link ConfigRef#DEFERRED_KEYS}
* itself, the same reason {@link ConfigRef#COLD_KEYS} and {@link ConfigRef#SPLIT_KEYS} are
* production constants rather than test-side copies: two lists that are supposed to describe the
* same method are exactly the shape that silently drifts apart. {@code spawnReadyTimeoutMs}/
* {@code spawnReadyPollMs} are compared together and reported under one combined label
* ({@code "spawnReady*"}); {@code profiles} is compared twice over — once for added/removed
* profile names, once for an existing profile's launch settings — and that second comparison
* excludes {@code weight}/{@code maxLoad}/{@code credentialId} as hot sub-fields, which is what
* {@link ConfigRefProfileCoverageTest} exists to keep honest at the sub-field level. {@code
* profiles} itself still belongs here, not in the hot-exclusion set below: most of a profile's
* fields are NOT read live, so citing "read live off the config supplier" for the whole
* top-level key would be false.
*/
private static final Set<String> DEFERRED_TOP_LEVEL_KEYS = ConfigRef.DEFERRED_KEYS;
/**
* The escape hatch: top-level components with no reload bookkeeping at all, because every read
* of them goes live through {@link ConfigRef#get()} rather than off a startup snapshot. A
* component belongs here ONLY if that is true — never because adding it here makes this test
* pass. fleetd #323 is the cautionary tale for exactly this pattern: the identical hatch on
* {@code ConfigRef.LAUNCH_SETTINGS_EXCLUDED} let two profile fields be silently re-broken with
* the whole suite green, and it was only caught by mutating the checker itself (see this class's
* own mutation test below, and {@code ConfigRefProfileCoverageTest}'s equivalent).
*
* <ul>
* <li>{@code placement} — read live by the placement policy on every spawn (class doc, Hot
* bullet; {@code ConfigRefTest.aConsumerHoldingTheRefSeesTheNewValue} proves it
* behaviourally).</li>
* <li>{@code memberCredentials} — read live at {@code Fleetd.java:198, 205, 729}.</li>
* <li>{@code memberLoginShell} — read live at
* {@code HerdrPeerLauncher#configuredMemberLoginShell}.</li>
* <li>{@code models} (fleetd #422) — membership is re-validated in full against the fresh
* config on every {@code ConfigRef.reload()} (via {@code FleetConfig#validateAll()}), so
* a bad edit is refused, never cached stale; the on/off half is read live by {@code
* CompositePeerLauncher.enforceModelEnabled} and its candidate filter, and by {@code
* fleet_profiles}/{@code GET /profiles} through {@code PeerLauncher.disabledModels()}.
* Unlike {@code fleet} below, nothing about {@code models} is baked into a startup-built
* object anywhere — there is no frozen half, so it belongs here whole rather than in
* {@code SPLIT_KEYS}.</li>
* <li>{@code leadRollover} (fleetd #480) — {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} and reads every field fresh on each
* {@code open()}/{@code confirm()} call, the same {@code () -> config.get().x()} shape
* {@code fleet}/{@code placement}/{@code models} use — see {@code ConfigRef}'s class doc
* Hot bullet. The only restart-only fact is structural, not a value going stale:
* {@code Fleetd.java} decides whether to construct the {@code LeadRollover} object at
* all off the startup snapshot (the same presence gate {@code leadHeartbeat:} uses), so
* a block ADDED where it was absent at boot needs a restart before anything exists to
* call — the same fact already true of adding a brand-new {@code profiles:} entry, which
* does not stop {@code profiles}' own hot sub-fields (weight/maxLoad/credentialId/
* exhaustedPattern) from being genuinely hot.</li>
* </ul>
*
* <p>{@code fleet} used to sit here too, on the strength of most of it (role pools, charters,
* tabLabel) being read the same live way — but {@code fleet.leaders} inside the same key is
* read only at startup and genuinely needs a restart, which is a real gap this escape hatch
* cannot represent: it excuses a whole top-level key from reload bookkeeping, and {@code fleet}
* needed exactly half of it excused. fleetd #333 moved it into {@link ConfigRef#SPLIT_KEYS}
* instead, where {@code changedSplitKeys} reports the frozen half by name. That is the
* cautionary tale this test's own javadoc already told: this checker proves the record's shape
* is triaged, never that a bucket a key sits in is the right one — a person has to read the
* source, which is exactly how fleetd #333 found {@code fleet} sitting in the wrong bucket
* while this test stayed green throughout.</p>
*/
private static final Set<String> HOT_EXCLUDED_TOP_LEVEL_KEYS =
Set.of("placement", "memberCredentials", "memberLoginShell", "models", "leadRollover");
@Test
void everyTopLevelComponentIsAccountedForInExactlyOneClass() {
Set<String> allNames = new TreeSet<>();
for (RecordComponent rc : COMPONENTS) {
allNames.add(rc.getName());
}
Set<String> cold = ConfigRef.COLD_KEYS;
Set<String> split = ConfigRef.SPLIT_KEYS;
Set<String> deferred = DEFERRED_TOP_LEVEL_KEYS;
Set<String> hot = HOT_EXCLUDED_TOP_LEVEL_KEYS;
// Typo guard on each set — the same check ConfigRefProfileCoverageTest runs on
// LAUNCH_SETTINGS_EXCLUDED. A name that does not exist on FleetConfig is a silent no-op.
assertTrue(allNames.containsAll(cold),
"ConfigRef.COLD_KEYS names a component that does not exist on FleetConfig: " + cold);
assertTrue(allNames.containsAll(split),
"ConfigRef.SPLIT_KEYS names a component that does not exist on FleetConfig: " + split);
assertTrue(allNames.containsAll(deferred),
"DEFERRED_TOP_LEVEL_KEYS names a component that does not exist on FleetConfig: " + deferred);
assertTrue(allNames.containsAll(hot),
"HOT_EXCLUDED_TOP_LEVEL_KEYS names a component that does not exist on FleetConfig: " + hot);
// The escape hatch is pinned. Growing it requires editing this line — a visible, deliberate
// diff, not a quiet one. See the field javadoc above for what "belongs here" actually means.
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell", "models", "leadRollover"), hot,
"HOT_EXCLUDED_TOP_LEVEL_KEYS changed. A component belongs here ONLY if it is read "
+ "live off the config supplier, never because adding it makes this test "
+ "pass. If you are adding one to silence this test, that is fleetd #323 "
+ "happening again: account for it in ConfigRef.changedDeferredKeys (or "
+ "COLD_KEYS/SPLIT_KEYS) instead. If it really is read live, name where and "
+ "update this assertion and the field javadoc together.");
// No component may sit in two buckets at once — the denominator check below could not catch
// that on its own (two buckets double-booking one key still sums to the right total if
// another key is simultaneously missing), so check every pair directly and name the culprit.
record Bucket(String name, Set<String> keys) {}
List<Bucket> buckets = List.of(
new Bucket("COLD_KEYS", cold), new Bucket("SPLIT_KEYS", split),
new Bucket("DEFERRED_TOP_LEVEL_KEYS", deferred), new Bucket("HOT_EXCLUDED_TOP_LEVEL_KEYS", hot));
for (int i = 0; i < buckets.size(); i++) {
for (int j = i + 1; j < buckets.size(); j++) {
Set<String> overlap = new LinkedHashSet<>(buckets.get(i).keys());
overlap.retainAll(buckets.get(j).keys());
assertEquals(Set.of(), overlap, "a component is in both " + buckets.get(i).name()
+ " and " + buckets.get(j).name() + ": " + overlap);
}
}
Set<String> union = new TreeSet<>();
union.addAll(cold);
union.addAll(split);
union.addAll(deferred);
union.addAll(hot);
System.out.printf(
"FleetConfig top-level coverage — %d components total: %d cold %s, %d deferred %s, "
+ "%d split %s, %d hot-excluded %s%n",
allNames.size(), cold.size(), cold, deferred.size(), deferred, split.size(), split,
hot.size(), hot);
Set<String> missing = new TreeSet<>(allNames);
missing.removeAll(union);
assertEquals(Set.of(), missing,
"these FleetConfig components are in none of COLD_KEYS, DEFERRED_TOP_LEVEL_KEYS, "
+ "SPLIT_KEYS or HOT_EXCLUDED_TOP_LEVEL_KEYS — triage each one into whichever "
+ "actually describes it: " + missing);
assertEquals(allNames.size(), cold.size() + deferred.size() + split.size() + hot.size(),
"counts don't sum to the component total even though every component was found in "
+ "the union — " + allNames.size() + " components, " + cold.size()
+ " cold + " + deferred.size() + " deferred + " + split.size() + " split + "
+ hot.size() + " hot-excluded");
}
}

Some files were not shown because too many files have changed in this diff Show More