Compare commits

...

230 Commits

Author SHA1 Message Date
Dai Ha 6e23bf8309 fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m52s
The drain-gate abort message told the operator a rerun "with or without
--no-build" would finish the restart. That is wrong: by the time this
message can fire, stage_built_jar has already moved the jar off $JAR, so
--no-build hits require_no_build_jar's own refusal. Reworded to say the
rerun must NOT use --no-build, and why: the built jar is no longer at the
live path that --no-build requires.

Also added a test pinning jar_id()'s no-argument default (reports $JAR,
the live path) and its explicit-argument behavior (reports that path
instead), per fleetd #511 item 2. Not adding a test for the JAR_STAGED rm
-f at line 578 (fleetd #511 documents it as an equivalent mutant — mvn
clean install deletes target/ on the next line regardless).
2026-09-12 10:56:32 +07:00
ltms aa4c0b84c3 Merge #510: never build into the path a running daemon holds (fleetd #493)
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m4s
2026-09-12 05:44:16 +02:00
ltms 40c593cd09 Merge #508: a herdr error during the pane scan is refused, not promoted to primary (fleetd #505)
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m38s
Closes the worker->primary escalation that fleetd #317 left open on the other side of the same
ternary. #317 stopped an UNRESOLVABLE pid being promoted. This stops a RESOLVED pid whose pane scan
failed being promoted: the scan now reports completeness, and an incomplete scan resolves to
anonymous.

Shape chosen: PaneLocator.terminalForPid returns Lookup(terminal, complete); paneOwnsAnyOf returns a
private Ownership enum OWNS/DOES_NOT_OWN/UNKNOWN, so a HerdrException is UNKNOWN rather than a clean
negative. ConnectionIdentity.Caller gains scanComplete as a third field -- resolved() was NOT
widened, correctly: it is documented as testing the lsof sentinel and #505 is a different axis. A
definite match still short-circuits, so a genuinely vanished non-owning pane does not become a
refusal.

Verified by me, not taken from the report.

Trial-merged onto main (136312f) and built the merge:
  mvn clean install -> Tests run: 1694, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS

Harness proof, by re-running the worker's own mutation rather than trusting it (UNKNOWN ->
DOES_NOT_OWN): Tests run: 63, Failures: 3 -- exactly the three tests meant to pin it, matching their
report line for line:
  CallerResolverTest.aHerdrErrorOnTheOwningPaneDuringTheScanIsRefusedNotPromotedToPrimary
  ConnectionIdentityTest.scanIsIncompleteWhenHerdrErrorsOnThePaneThatOwnsThePid
  PaneLocatorTest.anErrorOnThePaneThatOwnsThePidMakesTheScanIncompleteNotAClearNegative
So the cells can go red, and their numbers are honest. Restored, shasum 81b797a6... byte-identical,
control 63/0.

MY OWN MUTATION, on the half they did not touch, FOUND A GAP. I mutated the two-client completeness
fold in terminalForPid -- `complete = complete && outcome.complete()` -> `complete =
outcome.complete()`, which forgets an earlier client's failure. Two greps proved it applied. Result:
Tests run: 63, Failures: 0. NOT pinned.

It is not an equivalent mutant: with CB-185's two herdr daemons, a first client that scans
incompletely followed by a second that scans cleanly gives `false` originally and `true` mutated, so
an incomplete scan would be reported complete and promoted. It is the #393 shape instead -- no test
assembles that combination, because the single-daemon tests collapse lead and member to one object.
That axis has cost us before: four routing defects passed the suite when one client served both roles
with memberHerdrSocket unset.

Merging anyway, because the production behaviour is correct and the ticket's defect is pinned three
ways -- what is missing is a test for a dimension the PR description claims in prose. Filed as a
follow-up with the numbers above.

Second item for the follow-up, latent not live: FleetMcp.legacyPrincipal still does
`c.terminal() != null ? worker : primary` and drops scanComplete, so it would re-open exactly this
hole. Measured unreachable in production -- Fleetd.java:696 always passes a real CallerResolver, and
the contextExtractor only calls legacyPrincipal when `callers == null`. A defect on paper is not a
reachable defect, so it is not a blocker; it is a trap for whoever next touches that constructor.

That second item is also the fleet01 lead's structural point, and it is the sharper framing: #415's
antidote (make the decision total) is ORTHOGONAL to this defect. Authz.permits is already a
default-less switch over Action and cannot help, because it says nothing about whether the caller was
resolved correctly, and Principal carries no record of how it was resolved. #415 guards a decision
nobody wrote; this was a decision written correctly and fed a bad input. The durable fix direction is
that an unresolved scan should be unrepresentable as a principal rather than checked for at each
consumer.
2026-09-12 05:35:41 +02:00
Dai Ha 979adf82eb fleetd #493: never build into the path a running daemon holds
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 1m37s
redeploy-fleetd.sh's build step wrote straight into fleetd/target/fleetd.jar
via `mvn clean install` while the OLD daemon was still running from that
exact path. A JVM loads classes lazily, so a class the daemon had not
touched yet could be read from a jar already replaced or removed -- the
failure landed on the shutdown drain (NoClassDefFoundError, exit 143,
looks clean).

Stage the freshly built jar at target/fleetd-new.jar (stage_built_jar),
confirm the OLD pid has actually exited (wait_for_daemon_exit, extracted
from the existing wait loop), and only then swap it into the live path
(swap_staged_jar) -- strictly after the wait, strictly before start. A
failed swap dies without starting. --no-build and --check keep their
existing, truthful behavior (require_no_build_jar; jar_id still reads
the live path by default). A leftover staged jar from an interrupted
run is wiped before the next build. The build still runs before
anything is stopped, so a failed build still never takes the fleet down.

Also fixed: the drain-gate abort message ("aborted -- nothing changed")
now names the staged jar when one exists, since staging already moves
the freshly built jar off the live path before that prompt runs.

Adds unit tests for stage_built_jar, swap_staged_jar, require_no_build_jar,
wait_for_daemon_exit, and a source-order test proving swap sits after the
wait and before start (sourcing stops before the main flow ever runs, so
the ordering itself can only be checked by reading the script's own call
sites).
2026-09-12 10:35:19 +07:00
Dai Ha 36870836aa fleetd #505: a herdr error during the pane scan must not read as a clean negative
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m29s
A transient herdr error on pane.process_info during PaneLocator's pid→pane scan used
to be swallowed into a plain "does not own it", so a real worker whose owning pane
errored mid-scan resolved with a null terminal but a resolved (real) pid — exactly
what CallerResolver's loopback-trust fallback reads as the primary. That is a
worker→primary privilege escalation through the door fleetd #317 did not close: #317
guards a failed lsof lookup (c.resolved()), not a failed herdr pane scan.

Fix: add a third state to the scan instead of widening Caller.resolved() (which stays
centralised next to the lsof sentinel it tests, per #505's explicit instruction not
to reopen that decision). PaneLocator.terminalForPid now returns a Lookup(terminal,
complete) record: a HerdrException on one pane marks that pane's ownership UNKNOWN,
not DOES_NOT_OWN, and the scan is complete only if every pane was either matched or
confirmed not to own the pid. A definite match found elsewhere in the same scan
still short-circuits as complete — a pane that genuinely vanished mid-scan without
being the caller's own does not turn into a refusal.

ConnectionIdentity.Caller carries the new scanComplete flag alongside the unchanged
resolved(). CallerResolver's loopback-trust fallback now requires both resolved()
and scanComplete() before promoting to Principal.primary(); an incomplete scan
resolves anonymous, which fails toward the recoverable error (a refused primary
retries loudly; a promoted worker would not).

Logs a warning naming the pane and which herdr client (of how many) failed, so the
incomplete-scan path is diagnosable rather than silent (fleetd #317's own lesson).
2026-09-12 10:27:42 +07:00
ltms 136312fb11 Merge #503: Injector's readiness-grace warn prints measured elapsed time, never arithmetic on constants (fleetd #501)
CI / build (push) Successful in 1m28s
CI / contract (push) Successful in 1m27s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (4f28da6) and built the merge, because a clean auto-merge is not a compiling
merge:
  mvn clean install -> Tests run: 1689, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  InjectorTest -> Tests run: 34, Failures: 0
  Gitea CI on ac351ee: success

The worker mutated the two log arguments. I mutated the half it did not: the STAMP SITE. Changing
the first-sample guard so the clock is re-stamped on every non-ready poll gave 1 failure,
InjectorTest.readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants, on a clock-read
counter assertion: "the readiness-not-ready branch should read the clock exactly twice — once to
stamp the first non-ready sample, once at grace expiry". Two greps proved the mutant applied;
restored, shasum byte-identical, control 34/0 green.

Behaviour preserved: the restructure from `else if (p != null && ++t.notReadySincePoll >= N)` to a
nested `else if (p != null) { ... }` keeps the short-circuit, so the counter still increments only
when p != null. All three reset sites now clear notReadySinceMillis alongside notReadySincePoll.

Two notes for the record, neither a blocker.

The poll-count half is an equivalent mutant and the worker said so instead of reporting a kill it
did not get. That is the right call and the code comment states the limit honestly.

The clock is injected through a package-private constructor overload while the public constructors
default it to System::currentTimeMillis. My brief prescribed that shape. With exactly one production
construction site (Fleetd.java:531) a required parameter — fleetd #415's antidote — would have been
just as cheap and would match LeadRollover and the new Fleetd.awaitHerdr. The silent-survivor risk
#415 names is not present here, because the field is final and both public constructors delegate, so
a future constructor cannot compile without supplying it. Recording the choice so the next person
does not read it as an oversight.
2026-09-12 05:07:35 +02:00
ltms 4f28da62a3 Merge #499: redeploy-fleetd.sh gains an "unclear" supervisor state that cannot reach kill, and the detail survives the subshell (fleetd #492, #497)
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m32s
Two commits, b17f37a and 599419f. Verified by the lead, not taken from the worker's report.

b17f37a adds the fifth answer "unclear" for installed-but-not-loaded and for a probe that could not
answer at all, and makes require_drivable_supervisor die on it.

599419f fixes a defect I found in b17f37a: SUPERVISOR_UNCLEAR_DETAIL was set inside
detect_supervisor, which is called as $(detect_supervisor). A subshell only returns stdout, so the
detail never reached the caller, and a ${VAR:-generic} fallback then printed a message naming no
supervisor and no reason. Measured before the fix: plain call -> detail length 142; via $( ) ->
length 0.

Lead verification of the fix, re-running my own original measurement across every branch:

  launchd installed, not loaded   kind=[unclear]   detail_len=142
  systemd installed, not active   kind=[unclear]   detail_len=146
  launchd loaded (normal)         kind=[launchd]   detail_len=0
  systemd loaded (normal)         kind=[systemd]   detail_len=0
  neither -> none                 kind=[none]      detail_len=0
  both loaded -> ambiguous        kind=[ambiguous] detail_len=0

The separator is always emitted, so the ${RAW#*SEP} unpack cannot silently fall back to the whole
string on the four branches that carry no detail.

Harness proof, run by me: making detect_supervisor print the kind alone (two greps proved the mutant
applied) turned the suite red, exit 1, "FAIL: detail does not name the systemd unit it found
installed-but-not-loaded". Restored, shasum byte-identical, control exit 0, bash -n clean on both
files. 23 test functions defined, 23 invoked.

The three case "$SUPERVISOR_KIND" switches now have final arms, chosen by what each caller does with
the value: the reporting switch warns and continues (a diagnostic that aborts goes silent exactly
when the state is novel), while stop and start die.

Not fixed here, filed separately: the "was already not running" branches at :603 and :610 swallow a
failed launchctl unload / systemctl stop with || true and then print "ok" unconditionally.
2026-09-12 05:04:04 +02:00
ltms 708f1795ad Merge #502: awaitHerdr reports three outcomes with measured elapsed time, not one boolean (fleetd #498)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m49s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (8f59019) locally and built the merge, because a clean auto-merge is not a
compiling merge:
  mvn clean install -> Tests run: 1687, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  FleetdAwaitHerdrTest -> Tests run: 6, Failures: 0
  Gitea CI on 274afaf: success

The worker mutated the three return values. I mutated the half it did not: the reap DECISION at the
call site. Changing logHerdrWaitOutcomeAndShouldReap to return true for INTERRUPTED gave 1 failure,
FleetdAwaitHerdrTest.interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed:135 ("an interrupted
wait must not tell main to reap"). Two greps proved the mutant applied; restored, shasum
byte-identical, control 6/0 green.

Behaviour preserved: herdrUp is still true only for ANSWERED, and both its readers (the orphan reap
at :276 and the lead auto-launch gate at :375) see exactly what they saw before.

One correction to the worker's report, which changes nothing in the code. It wrote that a real
interrupt-detection regression "would also hang the daemon's startup thread forever". It would not:
in production `nanos` is System::nanoTime, so the deadline check still fires. The hang it hit was a
test-fixture property — a frozen injected clock with no iteration bound. That is fleetd #486's
shape, now seen in a second class, and it is recorded there.
2026-09-12 05:01:02 +02:00
Dai Ha 599419f9e6 fleetd #492 follow-up: carry the unclear detail across detect_supervisor's subshell boundary
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m53s
b17f37a set SUPERVISOR_UNCLEAR_DETAIL as a global inside detect_supervisor, but the real call
site invokes it as $(detect_supervisor) — a subshell — so that global died with the subshell and
the die() message's ${VAR:-fallback} silently masked the loss with generic text.

- detect_supervisor now packs kind and detail onto its one stdout line (joined by the ASCII unit
  separator byte, $SUPERVISOR_DETAIL_SEP), the only channel that survives $( ). The real call site
  unpacks both with in-shell parameter expansion — no extra subshell.
- Dropped the ${SUPERVISOR_UNCLEAR_DETAIL:-...} fallback at the die() message: under set -u, a
  missing detail now fails loudly instead of silently defaulting (same defect class as #497).
- Added a constraints comment block above detect_supervisor for future callers: stdout-only,
  no ${VAR:-default} papering over a lost value, and every case on the return value needs an
  explicit *) arm.
- Added *) arms to the three `case "$SUPERVISOR_KIND"` switches (report/stop/start): report warns
  and continues (display-only), stop/start die naming the value (they act on it).
- Rewrote the "unclear" test to go through the real call-site shape ($(detect_supervisor) then
  the same split), not a hand-constructed value, and tightened its final assertion to check for
  the actual detail text rather than $SYSTEMD_UNIT alone (the die() boilerplate names the unit
  either way, so that check could pass on a lost value).
2026-09-12 09:58:45 +07:00
Dai Ha ac351ee1de fleetd #501: readiness-grace expiry logs measured elapsed time and the loop's own poll counter, never the configured budget
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m2s
The line at Injector.java:383-387 printed two numbers that read as measurements
but were both compile-time constants: READINESS_GRACE_POLLS for the poll count
(the loop's own Target.notReadySincePoll counter was in scope at the same call
site), and READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000 for the elapsed
time — arithmetic on two constants, never a measurement, and wrong in the
direction that says everything ran on schedule.

Fix, copying the LongSupplier-clock shape LeadRollover already uses:
- print t.notReadySincePoll instead of the constant for the poll count (the two
  agree by construction on this branch, so no test can tell them apart — the
  comment says so honestly).
- add Target.notReadySinceMillis, stamped at the first non-ready sample and
  reset at all three sites notReadySincePoll already resets (:304, :355, :389
  pre-fix line numbers), to compute a real elapsed time at expiry.
- inject a LongSupplier nowMillis (defaulting to System::currentTimeMillis)
  through new package-private constructor overloads so a test can supply a
  clock whose advance does not track POLL_INTERVAL_MILLIS.

Tests use a ListAppender to assert on the log message contents, per
LeadRolloverTest's pattern. The elapsed-time test drives the loop with a stub
clock returning two literal, non-derived values so it can fail if the fix
regresses to the constant-arithmetic line — proved by mutation: reverting the
elapsed calculation to READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS turns that
one test red (1 failure); reverting the poll-count print to the constant is an
equivalent mutant (0 failures), because the counter and the constant are
identical at that exact call site by construction.

fleetd clean install: Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0.
2026-09-12 09:57:57 +07:00
Dai Ha 274afafde6 fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m51s
- awaitHerdr now returns a HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos) instead of a bare
  boolean, so 'the wait budget genuinely ran out' and 'the waiting thread was interrupted' are
  two distinct, named states instead of the same false (fleetd #497's shape).
- awaitHerdr takes the clock (LongSupplier) and the per-poll sleep (Runnable) as required
  parameters, with no defaulted overload (fleetd #415), so a test can drive it.
- The startup call site is extracted into logHerdrWaitOutcomeAndShouldReap, since main() itself
  cannot be driven from a unit test; it logs a distinct message per outcome, always printing the
  measured elapsed time next to the configured budget, never the budget alone.
- Adds FleetdAwaitHerdrTest covering the seam (all three outcomes, plus the preserved interrupt
  flag) and the call site (the three distinct log messages), using ListAppender.
2026-09-12 09:54:59 +07:00
ltms 8f59019305 Merge #496: lead-rollover logs measured elapsed time and measured counts, never the configured budget (fleetd #494)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 2m1s
Three commits. All verified by the lead in a detached worktree, not taken from the worker's report.

3fb3311 — the /clear-settle and success paths return measured elapsed time and a real nudge count
c87cc25 — print the measured nudge count; fix the same defect on the sibling turn-settle timeout line
e966cba — the poll count on the same line was still a constant; the test could not tell the difference

Lead verification at e966cba:
  mvn clean install -> Tests run: 1681, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  LeadRolloverTest  -> Tests run: 33, Failures: 0, Errors: 0, Skipped: 0
  Gitea CI on e966cba: success

Mutation M2 (PICKUP_GRACE_POLLS 8 -> 5), run by the lead: 2 failures,
clearGraceReleaseLogIsWarnWithMeasuredElapsed and successLogPrintsMeasuredElapsedForTheWholeRoll.
The emitted line under the mutant read "after 5 consecutive IDLE/DONE polls (4 of those were
nudged)", so both numbers now follow the loop. Restored, shasum byte-identical, control 33/0 green.

Known and deliberate limit, stated in the code comment: on the grace-release branch both counters
equal the constants by construction (one write site each), so no test can prove the difference on
that line. The change buys one source of truth, not provable coverage. The place nudges genuinely
varies with the run is the /clear-timeout warn, which is covered by a discriminating test.
2026-09-12 04:44:10 +02:00
Dai Ha e966cbadf9 fleetd #494 follow-up (2nd pass): the grace-release line still had one constant, and its test could not tell the difference
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m55s
Two more fixes on the same line, LeadRollover.java:600-624:

1. The poll-count argument (second, was PICKUP_GRACE_POLLS) now prints the
   loop's own idlePollsAwaitingPickup counter instead of the constant. Same
   defect shape as the nudges fix from c87cc25, one argument over.

2. LeadRolloverTest's clearGraceReleaseLogIsWarnWithMeasuredElapsed asserted
   its nudge-count expectation as `LeadRollover.PICKUP_GRACE_POLLS - 1` — the
   same expression the production code used to build the log line from, so it
   could not discriminate a reverted fix. Rewritten to plain literals
   ("after 8 consecutive IDLE/DONE polls (7 of those were nudged)"), proven to
   trip when PICKUP_GRACE_POLLS's value changes (2 failures under a
   PICKUP_GRACE_POLLS=5 mutation: this test and successLogPrintsMeasured...).

Also corrected the comment above the log line: idlePollsAwaitingPickup and
nudges each have exactly one write site on this loop's release branch, so they
cannot differ from PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 at this call
site — confirmed by re-deriving the loop's control flow, and by reverting just
the nudges argument (M1) and observing 33/0 stayed green even after the test
rewrite. That is an equivalent mutant on this line, not a gap the test rewrite
could close; the honest value of printing the counters is one source of truth
for the loop, not a provable-by-test difference here. The place nudges truly
varies with the run — and is covered by a test that can tell it apart from a
constant — is the /clear-timeout warn's clearResult.nudges() in runRollover.

mvn -Dtest=LeadRolloverTest test: 33/0. mvn clean install: 1681/0, BUILD SUCCESS.

Same-shape sweep (found, not fixed, per instructions):
Injector.java:383-387 — the readiness-grace warn prints the constant
READINESS_GRACE_POLLS where the measured per-target counter
t.notReadySincePoll is in scope (single increment site, just reached the
threshold at this call site — same equivalent-mutant situation as this fix).
2026-09-12 09:39:39 +07:00
Dai Ha b17f37a683 fleetd #492 follow-up: detect_supervisor must never read "could not tell" as "none"
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m32s
Two situations were silently landing in the "none" answer, which
require_drivable_supervisor accepts and the script then falls back to a raw
kill + nohup — exactly the wrong move when a supervisor actually IS present:

- installed-but-not-loaded, on either supervisor. `systemctl --user is-active`
  answers "no" for activating/deactivating/failed and while an auto-restart is
  pending too, and every one of those is a host that IS under systemd (or
  launchd) and about to act again. `*_installed` already knew this; it was
  only ever consulted for a warning line, never by the decision itself.
- a systemd probe that could not answer at all (e.g. systemctl cannot reach
  the user bus over a non-lingering ssh session) looked identical to a clean
  negative, because both probes redirected stderr straight to /dev/null.

detect_supervisor now returns a fifth answer, "unclear", for both cases.
systemd_loaded/systemd_installed capture systemctl's exit status and stderr
separately and set their own *_ERRORED flag only on a real tool failure
(non-zero exit WITH stderr), never on a clean negative. "none" now means only:
neither supervisor installed, neither loaded, neither probe errored.
require_drivable_supervisor die()s on "unclear" exactly like it already does
on "ambiguous", naming the specific supervisor and reason via the new
SUPERVISOR_UNCLEAR_DETAIL global.

Tests: 4 new cases (systemd/launchd installed-but-not-loaded, a real
systemd_loaded run through a systemctl stub that errors on stderr, and the
die() refusal for "unclear" naming the unit). All 3 new guards were verified
by mutation: each was removed from the real script, the suite caught it (a
new FAIL line naming the exact broken assertion), then the file was restored
byte-identically and the suite went green again.
2026-09-12 09:23:21 +07:00
Dai Ha c87cc25aa6 fleetd #494 follow-up: print the measured nudge count, and fix the sibling turn-settle timeout line
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m30s
- waitForClearPickupAndSettle's grace-release warn now prints the measured 'nudges' counter
  instead of the constant PICKUP_GRACE_POLLS - 1. The two happen to agree today, but the
  constant expression was wrong once before (printed PICKUP_GRACE_POLLS itself, claiming 8
  nudges where 7 went out) and a code read did not catch it — only a mutation test did.
  Printing the counter cannot drift from the loop's real behaviour.
- waitUntilAtTurnBoundary (the FIRST wait, ~line 386-391) had the identical 'configured value
  printed as if measured' defect as the three lines fixed in the original #494 commit, but was
  out of scope because the brief named specific lines instead of the shape. Fixed the same way:
  it now returns a TurnSettleResult(settled, elapsedMillis) instead of a bare boolean, and the
  timeout warn prints 'configured={}s elapsed={}ms' instead of presenting cfg.turnSettleSeconds()
  as the measured wait.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Test additions: clearGraceReleaseLogIsWarnWithMeasuredElapsed now also asserts the measured
  nudge count; successLogPrintsMeasuredElapsedForTheWholeRoll's expected elapsed value is
  updated (7000ms, not 6500ms) to account for waitUntilAtTurnBoundary's own new clock read; a
  new turnTimeoutLogPrintsMeasuredElapsedNotJustConfigured test pins the sibling line.
- Proved the nudges fix with a temporary mutation: set PICKUP_GRACE_POLLS to 5, confirmed via
  two greps that the mutant applied and the original constant was gone, ran the grace-release
  test and read the actual log line — nudge count followed to 4 (= 5 - 1), then restored to 8
  and reran the full LeadRolloverTest suite as a control (33/33 green).
2026-09-12 09:22:19 +07:00
Dai Ha 3fb331145a fleetd #494: log measured elapsed time, never the configured budget, on a lead-rollover failure/success
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m35s
- runRollover's /clear-timeout warn now prints configured/elapsed/nudges, each labelled,
  instead of presenting cfg.clearSettleSeconds() as if it were the measured wait.
- waitForClearPickupAndSettle's pickup-grace release is now log.warn (was log.info) and
  prints the measured elapsed time next to the target pane — this is the exact path that
  reported a false-success roll in the real incident (438ms of a 20s budget).
- The success line ('lead-rollover: rolled') now prints the measured elapsed time for the
  whole roll.
- waitForClearPickupAndSettle now returns a ClearSettleResult(settled, elapsedMillis, nudges)
  instead of a bare boolean, so callers can log the measured values instead of the config.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Adds 3 tests to LeadRolloverTest pinning the content of each changed log line, using a
  self-advancing fake clock so the measured elapsed/nudge values are deterministic and
  provably distinct from the configured budget.
2026-09-12 09:08:04 +07:00
Dai Ha dcd505286f fleetd #492: teach redeploy-fleetd.sh systemd --user as a third supervisor
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
launchd, systemd --user, and unsupervised are three different answers, not two.
Refuse (die) rather than fall through to kill+nohup when a supervisor is
detected that this script cannot drive (e.g. both signals fire at once), and
add a post-restart check that fails the run if more than one fleetd process
is alive. detect_supervisor()/require_drivable_supervisor()/
count_daemon_pids()/assert_single_daemon() are pure, overridable functions so
scripts/test-redeploy-fleetd.sh can exercise them without a real launchd or
systemd.
2026-09-12 09:00:59 +07:00
Dai Ha 60fa958da2 docs: a relative handoverPath lands in the LEAD's repo, not fleetd's (#491)
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
On this host the lead's cwd IS the fleetd checkout, so fleetd/.gitignore
protects the handover file and the distinction is invisible. On fleet01 the
lead works in /home/ltms/LTMS/kb while fleetd sits in a different directory,
and that repo has no .handover rule — measured 2026-09-12.

Tell the lead to check its own workspace's .gitignore before writing, and to
report it rather than committing the file or editing someone's .gitignore.

Refs fleetd #491, #487, #480.
2026-09-12 08:46:17 +07:00
Dai Ha 9950361bc9 docs: #489 is fixed and deployed, but criterion 6 is still unmet
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m33s
The skill said the roll 'does nothing until #489 is merged and redeployed'.
Both have happened, so that sentence now reads as a green light. It is not
one: no roll has bootstrapped a fresh session end to end yet. Say that
plainly, and tell the lead to warn the operator before confirming.

Refs fleetd #489, #480.
2026-09-12 07:55:41 +07:00
ltms 9f74b6619a Merge #490: nudge the /clear submit keystroke before bootstrapText (fleetd #489)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m30s
Fixes the live defect measured on 2026-09-12: the roll joined /clear and
bootstrapText into one line and Claude Code refused it as
"Unknown command: /clearFresh".

Verified by the lead before merge:
- mvn clean install: Tests run: 1677, Failures: 0, BUILD SUCCESS
- LeadRolloverTest baseline: 29 tests, 0 failures
- mutation PICKUP_GRACE_POLLS 8 -> 1: 1 failure (the regression test)
- mutation deleting the !clearSettled guard: 3 failures, 2 pre-existing tests
- mutation replacing waitForClearPickupAndSettle with "return true": 5 failures,
  including the strengthened pickup test
- positive control after each restore: 29 tests, 0 failures
- CI run 1738 on f687046: success

A separate mutation of the FIRST gate hung the suite instead of failing it.
That is fleetd #486, not a regression here; the reproduction is recorded there.
2026-09-12 02:52:52 +02:00
Dai Ha f687046450 fleetd #489 follow-up: fix stale class javadoc, off-by-one nudge count, weak test
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m33s
Three review corrections on top of the previous commit:

1. The class javadoc's four-step continuation list (lines 47-58) was stale.
   Step 1 said "report an injectable state", but waitUntilAtTurnBoundary's
   own javadoc excludes BLOCKED - fixed to say IDLE or DONE. Step 3 still
   described the old plain re-check ("the original, pre-correction wait...
   still here") - fixed to describe what waitForClearPickupAndSettle
   actually does: nudge while unpicked-up, then wait for a real WORKING ->
   IDLE/DONE boundary, releasing rather than wedging if WORKING never shows.

2. PICKUP_GRACE_POLLS=8 bounds the number of consecutive not-yet-picked-up
   polls, not the number of nudges - the 8th poll releases instead of
   nudging again, so 8 polls produce 7 nudges. The log.info in the release
   branch and two javadoc spots said "8 nudges"; fixed all three to state
   the poll count and the nudge count separately and correctly. Behavior
   and the constant are unchanged.

3. pickupSeenStopsNudgingAndBootstrapTextIsSent asserted only promptCallCount
   and sendKeysCallCount, both of which a return-true stub also satisfies.
   Added an assertion on the already-tracked postClearGetCalls counter
   (>= 2), which only a real post-/clear poll loop can produce - this is
   what makes the test fail against a return-true mutant.
2026-09-12 07:45:36 +07:00
Dai Ha c6058652be fleetd #489: nudge the /clear submit keystroke before bootstrapText
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m36s
LeadRollover.runRollover's second wait (after /clear) was a no-op: it polled
for IDLE/DONE, which /clear itself never leaves since it starts no real turn,
so it always returned true on the first poll. Combined with a direct
agents.send bypassing Injector (deliberate, to avoid wedging the pane), the
submit Enter that accompanies /clear could race the paste and leave it
unsubmitted — bootstrapText then landed concatenated onto the same input
line, exactly as measured live on 2026-09-12.

Replace that second wait with waitForClearPickupAndSettle, which copies the
pickup-nudge pattern Injector already ships for its own post-turn /clear
housekeeping (fleetd #306): nudge agents.submit while the pane hasn't
reported WORKING yet, release after PICKUP_GRACE_POLLS=8 nudges rather than
wedge, and require a real WORKING -> IDLE/DONE boundary once a pickup is
observed. BLOCKED stays excluded from both the nudge and the boundary check,
same as the (unchanged) first wait — a paused live turn is not settled, and
nudging Enter into an open prompt could wrongly answer it.

Adds four tests to LeadRolloverTest covering the paste-race regression
(nudge ordered between /clear and bootstrapText), a confirmed pickup, a
deadline expiry with no boundary ever reached, and a throwing submit().
2026-09-12 07:29:26 +07:00
Dai Ha 2a95da2ff4 docs: the handover skill must warn that the automatic roll is broken (#489)
CI / contract (push) Successful in 1m23s
CI / build (push) Successful in 1m53s
Section 11 said the bootstrap prompt landing was 'not yet proven
end-to-end'. It is now measured failing: the first real roll joined
/clear and bootstrapText into one line. Nothing was cleared, so the
failure is safe, but a lead that reads the old wording would reach for
fleet_handover expecting it to work.

Refs fleetd #489, #480.
2026-09-12 07:21:38 +07:00
Dai Ha 7f9137fcb6 fleetd #480: ignore the handover file, and note the absolute-path guarantee in the handover skill
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m52s
The lead rollover handover file now lives inside the workspace, at the relative
path fleetd.yaml's leadRollover.handoverPath names. It is a snapshot of one
moment's live state, so it must never enter git history.

The handover skill also now says the handoverPath fleetd hands back is always
absolute, even when the configured value is relative — a lead that resolves it
itself can pick a different file from the one the daemon checks.
2026-09-12 05:46:34 +07:00
Dai Ha 3bf3968bc7 Merge #487: resolve a relative leadRollover.handoverPath against the calling lead's workspace (fleetd #480 follow-up) 2026-09-12 05:43:12 +07:00
Dai Ha 261aa056f9 fleetd #480 follow-up correction 2: guarantee resolveHandoverPath is always absolute
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m32s
LeadRollover.resolveHandoverPath's relative branch resolved the configured
handoverPath against leadWorkspace.apply(...) (fleet.leaders.<name>.cwd) but
never forced the result absolute. If an operator writes a RELATIVE cwd, the
returned path stays relative, silently breaking the "always absolute"
contract documented on PendingRollover.

Fix: call toAbsolutePath() unconditionally on both branches (the
already-absolute input branch, where it is a no-op, and the relative
branch), so neither branch trusts isAbsolute() alone to already imply what
toAbsolutePath() enforces. Method javadoc now states the absolute result is
guaranteed, not merely usual.

Added a test: a lead with a RELATIVE cwd and a relative handoverPath still
yields an absolute PendingRollover.handoverPath. Asserts both isAbsolute()
and the exact resolved value, since isAbsolute() alone would also pass for a
path resolved against the wrong base.

Proved the test discriminates: reverting the toAbsolutePath() calls (keeping
the test) made it fail with an AssertionFailedError ("expected: <true> but
was: <false>"); restoring the fix made it pass again.

Note: FleetConfig has no validation on fleet.leaders.<name>.cwd at config
load (grep across every validate* method: 0 matches for .cwd()) — a relative
cwd is silently accepted. Not adding validation here per instruction; that
is a separate ticket.
2026-09-12 05:41:40 +07:00
Dai Ha 042b8c99dd fleetd #480 follow-up correction: cover Fleetd.leadRollover(...)'s own wiring behaviourally
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m51s
Add FleetdLeadRolloverWorkspaceLookupTest, calling the package-private Fleetd.leadRollover(...)
factory directly (with a real ConfigRef built from a temp fleetd.yaml, never the gitignored live
one) to prove the terminal -> lead-name -> Leader.cwd() lookup it builds actually works: a relative
handoverPath resolves against the calling lead's configured cwd; a terminal absent from the live
lead-terminal map falls back to user.dir; and the lookup is read live, not snapshotted at
construction time (a lead discovered by the tab scan after leadRollover(...) was built still
resolves correctly).

Proved this closes the gap: mutating the factory's lambda body (String leadName = null;, always
"no lead found", which forces the daemon-cwd fallback this ticket exists to fix) left the full
1669-test suite green before this commit. With the new test added, the same one-line mutation now
fails 2 of its 3 cases; reverting it goes green again (3/3). Mutation applied/reverted only during
verification and is not part of this commit (git diff on Fleetd.java is empty).

FleetdLeadRolloverWiringTest's class javadoc corrected: it previously claimed no behavioural test
could catch this wiring dropping out, which was true only before this commit and only covered the
factory's own body, not its call site. Restated what each test class actually covers: the source-
text pin covers the call site's argument list; the new behavioural test covers the lambda's body.
2026-09-12 05:32:25 +07:00
Dai Ha 4bfab6b718 fleetd #480 follow-up: resolve a relative leadRollover.handoverPath against the calling lead's workspace
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m4s
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the
calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that
lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores
only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed
back to the lead in the fleet_handover open response, and the default bootstrapText sentence
all see the same absolute location instead of a value resolved against whatever directory the
daemon process happened to start in.

FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it
would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor
(resolvedHandoverPath) method builds the default sentence from the resolved path instead.

Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to
lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on
every call, never off a startup snapshot.
2026-09-12 05:19:49 +07:00
Dai Ha 008a457557 docs: fleet_handover is shipped — update the charter table and the handover skill
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m38s
fleetd #480 merged in #483, #484 and #485, so the instruction surface has to
catch up. Per CLAUDE.md's own rule, a code change that silently invalidates the
canonical block is an incomplete change.

CLAUDE.md + wiki/7-Use-Cases.md: one new intent-to-tool row for fleet_handover.
Both copies edited identically; the byte-identical sync check prints True.
The row carries the trap rather than just the call: open FIRST, then write the
file, then confirm — because confirm refuses the file as stale unless its
modified time is later than the open request, so the obvious order fails.

.claude/skills/handover/SKILL.md: the skill said "do not call a fleet_handover
tool: it does not exist." That was true when it was written this morning and is
now false, which is the worst state for an instruction file to be in. Replaced
with the two real paths (by hand, or with the tool when leadRollover: is
configured), plus a new section 11 giving the three-step order and the five
things that surprise a caller — chiefly that `accepted` does not mean the pane
has been cleared, and that operatorConfirmed is a report of what a human said,
not a confidence level.

Section 11 also states plainly that the bootstrap prompt landing in a freshly
cleared pane is not yet proven end-to-end, and that the recovery is the manual
path — which is why the file is written before confirm, never after.
2026-09-11 07:29:59 +07:00
ltms 9494a6b99a Merge #485: fleetd #480 Unit C — the fleet_handover MCP tool
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m58s
fleet_handover{action: "open"|"confirm"|"cancel"} drives LeadRollover, which #483 and
#484 landed with nothing calling it. Primary-only via a new Authz.Action.HANDOVER, on
the same case line as SPAWN/STOP/DRAIN.

The tool has NO terminal, session or leadTerminal parameter of any kind — the pane is
always callerTerminal(exchange), resolved from the connection. A lead can therefore only
ever roll itself, never another lead. That is charter invariant 3, and it is the second
of the two corrections recorded in LeadRollover's class javadoc.

Registered unconditionally, so the tool surface does not vary with config: with
leadRollover: absent, every action returns a clean NOT_CONFIGURED refusal instead of
failing, and open()'s IllegalStateException (config removed by a hot reload after
construction) is caught and turned into the same refusal. A config-dependent tool set
would have collided with #474's charter tool-surface gate and McpContractDocTest.

Correction round applied before merge, and it is the reason this took two passes.
The unit first shipped with a defaulted 15-argument FleetMcp constructor delegating to
the new 16-argument one with leadRollover = null. I mutated the wiring rather than
reasoning about it: deleting just the leadRollover argument from Fleetd.main's FleetMcp
call compiled with 0 errors and passed all 1659 tests, BUILD SUCCESS — while the live
daemon would have answered NOT_CONFIGURED to every fleet_handover call for ever.
Neither FleetMcpHandoverTest (it builds its own FleetMcp) nor FleetdLeadRolloverWiringTest
(it pins that LeadRollover is constructed, not that it is passed on) could see it.

The worker then found the defect was wider than I had named: all five shorter
constructors (11/12/13/14/15-arg) formed one defaulting chain into the 16-arg one, each
silently supplying another feature's "off" value — leadChannel, outage, leadSeats, peers,
and finally leadRollover. All five are deleted. FleetMcp now has exactly one public
constructor, so every one of those features is compile-enforced at its call site, not
just this one.

Verified by the lead before merge, on PR head merged with current main (0176378):
- CI run 1728 green on eb05576.
- The identical mutation re-run by me: control against the original -> 1; mutant present
  (MUT-DROP-ARG2) -> 1; original gone -> 0; `mvn -o -q compile` now FAILS —
  "constructor FleetMcp ... cannot be applied to given types; reason: actual and formal
  argument lists differ in length" at Fleetd.java:[698,24]. The antidote holds: required
  parameter = compile error, defaulted overload = silent survivor a green suite vouches for.
- Restored clean (0 changed files), then mvn clean install on the merged tree:
  BUILD SUCCESS 1, BUILD FAILURE 0, Tests run: 1662, Failures: 0, Errors: 0, Skipped: 0.
- Three call sites of `new FleetMcp(` measured, all accounted for: Fleetd.java,
  FleetMcpAuthzTest.java, FleetMcpHandoverTest.java.
2026-09-11 02:27:56 +02:00
Dai Ha eb0557621e fleetd #480 correction round: collapse FleetMcp to one required constructor
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 1m32s
FleetMcp had a defaulted 15-argument constructor that delegated to the new
16-argument one with an implicit null for leadRollover. Dropping the
leadRollover argument from Fleetd.main's FleetMcp(...) call fell back to that
shorter overload, compiled fine, and left all 1659 tests green — the live
daemon would then answer NOT_CONFIGURED to fleet_handover forever with
nothing going red.

Delete every overload that could reach the 16-arg constructor with a
silently-defaulted leadRollover (11/12/13/14/15-arg forms all chained to it),
leaving the 16-arg constructor as FleetMcp's sole public constructor. Update
FleetMcpAuthzTest's call site to pass every parameter explicitly (leadChannel
null, OutageSource.none(), LeadSeatSource.none(), List.of(), leadRollover
null) — Fleetd.java and FleetMcpHandoverTest already called the full form.
Proved with mvn -o -q compile: removing the leadRollover argument from
Fleetd.main now fails to compile instead of silently defaulting.

No behaviour changes — NOT_CONFIGURED refusals are unchanged.
2026-09-11 07:25:08 +07:00
ltms a0eed6f01b Merge #484: fleetd #480 Unit E — a BLOCKED pane is not a settled pane
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m35s
LeadRollover's two settle waits used AgentStatus#injectable(), which is
IDLE || BLOCKED || DONE. That is the right rule for Injector ("may I deliver a message
without stepping on a live turn") and the wrong one here ("has the turn actually
ended"), because BLOCKED is a live turn that is merely paused — a pane sitting on an
approval prompt.

Path in: the lead calls confirm(); its turn carries on and hits anything needing
approval; herdr reports blocked; within 250ms the deferred continuation reads that as
settled; /clear is typed into an open prompt. That destroys the lead's live context
mid-turn, which is exactly what the turnSettleSeconds gate added in #483 exists to
prevent. The window is turnSettleSeconds, default 20s.

Both waits now require a real turn boundary — IDLE or DONE. AgentStatus#injectable() is
untouched: it is correct for Injector, LeadHeartbeatLoop, LeadCoordLoop, ReplyPushLoop
and HerdrPeerLauncher's two readiness checks, all of which are delivery gates.

Verified by the lead before merge:
- CI run 1726 green on head e2a91e8.
- Mutation on the half the worker did NOT mutate: dropped the DONE arm, leaving
  `if (status == AgentStatus.IDLE)`. Mutant proven applied with two unrelated proofs
  using different search strings (MUT-DROP-DONE present -> 1; 'IDLE || status' gone
  -> 0). Same single test, same 100s cap: unmutated passes in 0.172s; mutated is still
  running at 100s and does not pass. So doneStatusStillCompletesTheFullRoll really does
  pin the DONE arm, and the fix did not over-tighten to IDLE-only.

Why the earlier #483 battery could not have caught this: it proved the guard FIRES, not
that the guard's predicate was tight enough. LeadRolloverTest drove the fake status with
only "idle" and "working"; neither value discriminates this defect, only "blocked" does.

Correction round applied before merge: both warning lines said "never went idle" and
"did not become injectable". Retiring that word is the whole point of the change, so a
future reader debugging from those lines would have been handed back the wrong mental
model. Both now say "did not reach a turn boundary (IDLE or DONE)".

Filed separately as #486, deliberately not fixed here: with the DONE arm removed the
test hangs rather than fails, because waitUntilAtTurnBoundary is bounded only by an
injected clock that the tests hold fixed. Not a production bug — nowMillis is
System::currentTimeMillis there — but it turns a future regression into a CI hang
instead of a red test.
2026-09-11 02:21:06 +02:00
Dai Ha e2a91e883e fleetd #480 Unit E correction: retire "injectable" wording from the log lines
CI / build (pull_request) Successful in 1m28s
CI / contract (pull_request) Successful in 1m27s
Both waitUntilAtTurnBoundary guard messages still said "never went idle" /
"did not become injectable" — the old mental model the rename was meant to
retire. Made both say what the code now actually waits for: a turn boundary
(IDLE or DONE).
2026-09-11 07:18:51 +07:00
Dai Ha 62646957ea fleetd #480 Unit C: wire fleet_handover MCP tool onto LeadRollover
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m56s
Adds the fleet_handover tool (open/confirm/cancel) as a thin adapter over
LeadRollover, registered unconditionally so the charter tool-surface gate
sees a stable set regardless of whether leadRollover: is configured. With a
null LeadRollover every action degrades to a clean NOT_CONFIGURED refusal
instead of throwing. Gated on a new Authz.Action.HANDOVER (primary-only,
same as SPAWN/STOP/DRAIN). The caller's own connection-resolved terminal is
the only lead identity ever used — the tool's input schema carries no
terminal/session/leadTerminal parameter, so a lead can only ever roll
itself. Fleetd.main now passes its existing leadRollover local into FleetMcp
via a new trailing constructor parameter.
2026-09-11 07:07:24 +07:00
Dai Ha 802c0ab701 fleetd #480 Unit E: BLOCKED is not a settled turn boundary
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m55s
LeadRollover's waitUntilInjectable used AgentStatus#injectable(), which
accepts BLOCKED. A BLOCKED pane is paused mid-turn on a prompt, not
settled — reusing injectable() let /clear (or the bootstrap text after
it) fire into an open approval prompt within the 20s settle window,
destroying the lead's live context.

Renamed the helper to waitUntilAtTurnBoundary and restricted both waits
to IDLE or DONE only, with a comment explaining why this class does not
reuse injectable() (it answers "may I deliver", not "has the turn
ended"). Added tests for BLOCKED-forever on both waits (zero sends /
exactly one send) and for DONE still completing the full roll.
2026-09-11 07:01:28 +07:00
ltms 4bab23e241 Merge #483: fleetd #480 Unit A — lead rollover core (config + executor)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m37s
Two corrections were applied to the original unit before this merge, and the PR
description above still describes the pre-correction shape:

1. confirm() no longer rolls inline. It validates every gate, then hands a one-shot
   continuation to continuationRunner and returns. confirm() is called BY the lead FROM
   its own turn, so the pane is WORKING and cannot report injectable until confirm()
   returns; the old code sent /clear first and then timed out waiting, which destroyed
   the lead's context and started no fresh session. The continuation waits for the
   calling turn to settle FIRST (turnSettleSeconds, a new knob), and if that wait times
   out it sends no /clear at all.
2. The pane to roll comes from the caller's terminal, not PrimaryRegistry. A single-slot
   lookup let lead X's confirm() clear lead Y's pane. open() records the caller's
   terminal; confirm() refuses with NOT_YOUR_ROLLOVER on a mismatch.

Verified by the lead before merge:
- CI run 1722 green on head a942622.
- Mutation battery on the two safety-critical branches, in a scratch worktree at a942622.
  Controls measured against the ORIGINAL first; each mutant proven applied with two
  unrelated proofs using different search strings.
  M1 `if (!turnSettled)` -> `if (false)`: turnThatNeverSettlesSendsNoClearAtAll FAILS
     (LeadRolloverTest:247, expected 0 sends but was 1).
  M2 ownership check -> `if (false)`: aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover
     FAILS (LeadRolloverTest:266). Unmutated harness-proof run: exit 0 with both guard
     WARN lines logged, so the branches really are exercised.

Nothing calls LeadRollover yet. The fleet_handover MCP tool is Unit C.
2026-09-11 01:51:20 +02:00
Dai Ha a94262271b fleetd #480 correction round: defer the roll, and gate it on caller identity
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m3s
Two defects found after the fact, both from the original brief, both fixed here.

1. confirm() is called FROM the calling lead's own turn, so its pane is still
   WORKING and can never report injectable inside that same call. The old
   confirm() sent /clear before polling for that — the poll always timed out,
   but only after /clear had already fired and queued, destroying the lead's
   context with no fresh session ever started and a refusal return that lied
   about what had happened.

   Fix: confirm() now only validates and, if every gate passes, hands a
   one-shot continuation to a new continuationRunner (a real virtual thread in
   production, Runnable::run in tests) and returns RollDecision.approved()
   immediately - "scheduled", not "rolled". The continuation itself does the
   actual work, once the calling turn has ended: wait for the SAME pane to
   report injectable again (new turnSettleSeconds config key, default 20) -
   if this never happens, /clear is NEVER sent, at all - then /clear, then
   wait again (clearSettleSeconds, as before), then bootstrapText. The "no
   timer/scheduler, only confirm() can roll" invariant is restated precisely
   in LeadRollover's class javadoc: it is about initiative, not synchronicity
   - a single-shot continuation of an already-approved confirm() call still
   satisfies it; a recurring background loop would not.

2. confirm() resolved the pane to clear via PrimaryRegistry.primaryTerminal(),
   a single-slot lookup that is correct for a background loop with no caller
   but wrong here: on a daemon with more than one labelled lead tab, lead X's
   confirm() could clear lead Y's pane, violating the charter's "identity
   comes from the connection, never an argument" invariant.

   Fix: open() and confirm() now take the caller's terminal id as a parameter
   (resolved by the MCP layer from the connection - the later MCP-tool unit
   must pass it in, never accept it as a request field). confirm() refuses
   with a new NOT_YOUR_ROLLOVER reason unless it matches the terminal open()
   recorded. LeadRollover no longer depends on PrimaryRegistry at all.

Also: renamed RollResult to RollDecision (rolled -> accepted) to reflect the
new meaning - approved and scheduled, not necessarily cleared yet. Added
turnSettleSeconds to the leadRollover: config block (documented in
fleetd.example.yaml alongside the existing keys) and updated Fleetd.java's
leadRollover(...) factory to drop the primaryRegistry parameter, with
FleetdLeadRolloverWiringTest's source-text pin updated to match.

New tests: turnThatNeverSettlesSendsNoClearAtAll (the branch that matters
most - a turn that never ends means /clear is never sent) and
aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover (NOT_YOUR_ROLLOVER),
plus a settle-after-clear timeout test and an open() input-validation test.
LeadRolloverTest: 11 -> 14 tests.
2026-09-11 06:44:58 +07:00
Dai Ha 5c12865c25 fleetd #480 Unit A: lead rollover core (config block + executor)
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m58s
Adds the opt-in leadRollover: config block and LeadRollover, the executor a
later unit's MCP tool will call. A lead writes a handover file, then open()
records a token and confirm() verifies it (exists, non-empty, fresh) and an
operator confirmation before clearing the lead's own pane via /clear (sent
directly through AgentControl, bypassing Injector, same as
ClaudeCodeLauncher#clearContext) and bootstrapping a fresh session. Nothing
but an explicit confirm() call can ever roll a pane - no timer, no heartbeat,
no background thread anywhere in this class.

Wired into Fleetd.java exactly like LeadHeartbeatLoop: constructed only when
leadRollover: is present at startup, and nothing calls it yet - the MCP tool
is a separate, later unit.

Classified leadRollover: as HOT in ConfigRef (joins placement/
memberCredentials/memberLoginShell/models): the executor holds
Supplier<FleetConfig.LeadRollover> and reads every field fresh per call,
unlike LeadHeartbeatLoop's frozen final fields. The one caveat: the object's
construction is still gated on presence in the startup config snapshot, so a
freshly-added block needs a restart before anything exists to call.

Tests: LeadRolloverTest (14 cases covering the 6 hard requirements - no
object without the config block, only confirm() can roll, missing/empty/
stale handover file each refuse by name, requireOperatorConfirm gating, and
the injected wall-clock supplier) and FleetdLeadRolloverWiringTest (source-
text pin on Fleetd.main's construction call, mirroring
FleetdCompletionResolverWiringTest). Also updated the existing
FleetConfigValidateAllTest, FleetConfigWithDefaultsPreservesEveryComponentTest,
ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest to
account for the new record component.
2026-09-11 06:33:01 +07:00
Dai Ha 3f8c956325 docs: the handover skill must not promise a tool that does not exist
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m30s
The skill shipped in #481 described fleetd #480's automation in the present
tense: "fleetd ticket #480 lets a lead session hand off to a fresh one". The
tool is not built. A session loading the skill would look for a fleet_handover
tool it cannot call, and this repo already knows what a confidently wrong
instruction costs — a wrong comment stops the reader investigating with a false
conclusion, which is worse than no comment.

Now says plainly: today the handoff is manual, #480 will automate the same
cycle, and do not call fleet_handover because it does not exist. The nine
procedure rules are unchanged — only who performs the swap changes when #480
ships, not what the file must contain.

My wording to fix, not the worker's: the brief handed them the present tense.
2026-09-11 06:17:11 +07:00
ltms 2051ceeb26 Merge #481: the handover skill (fleetd #480 Unit B)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m27s
Adds .claude/skills/handover/SKILL.md — the outgoing lead's procedure for writing
the file a fresh lead session inherits.

Reviewed by the lead against the eight requirements in the brief: every number
carries its command, measured-vs-reported is marked, open decisions name their
owner, live hazards, what is not owed, a first section of three re-measurement
commands, a time-and-commit stamp, and a plain statement that the file goes
stale. All eight are present, plus a leave-out section and a template.

Verified by the lead, not taken from the worker's report:
- CLAUDE.md diff is the addendum's skill list only, in the primary-side group.
- The canonical block is byte-identical with the wiki template (sync check True).
- CI run 1718 green on 325d0771, the same commit reviewed.

Follow-up owed by the lead: the skill describes #480's automation in the present
tense, but the tool does not exist yet. The lead will soften that wording so the
skill does not promise a fleet_handover tool a session cannot call.
2026-09-11 01:16:40 +02:00
Dai Ha 325d0771a4 fleetd #480 Unit B: add handover skill for lead session handoff
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m54s
Adds .claude/skills/handover/SKILL.md, the procedure an outgoing lead
follows to write the handover file a fresh lead session inherits when
fleetd clears the pane. Registers the new skill in CLAUDE.md's
primary-side skills list; no other change to CLAUDE.md.
2026-09-11 06:13:49 +07:00
Dai Ha 7b97aae85b docs: move the redeploy procedure out of CLAUDE.md into a skill
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m34s
CLAUDE.md loads into every session. The daemon-redeploy procedure is
needed only when someone redeploys, so it paid for context it did not
use: 58 lines, about 918 est. tokens, every session.

The procedure now lives in .claude/skills/redeploy-fleetd/SKILL.md,
which loads only when invoked. The moved text is byte-identical to
what was removed, plus one new paragraph documenting --no-build
(scripts/redeploy-fleetd.sh:34,62,266) — the script and the wiki
already had that flag, CLAUDE.md never did.

CLAUDE.md keeps an 11-line pointer, because two rules must stay
resident: a merge is not a deployment, and workers must never
redeploy. A session learns it needs the procedure before it needs
the skill, then the next line names the skill.

Also updates the addendum's primary-side skill list, as this file's
own rule for .claude/skills/** changes requires.

CLAUDE.md: 34,442 -> 31,591 chars.
Canonical block untouched — the wiki sync check still prints True.
2026-09-11 05:48:06 +07:00
Dai Ha 49a404ddf3 Merge #474 follow-up: pin main's ConfigRef wiring against the surviving mutation
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m37s
My battery on the #474 merge found one survivor: reverting Fleetd.java:154
from the three-argument ConfigRef constructor to the plain two-argument
one turns the live reload gate off and leaves all 1633 tests green. Both
new #474 tests build their own ConfigRef with the method reference, so
neither reads what main chose.

This adds FleetdConfigRefWiringTest, following the three source-text
precedents already in the tree (FleetdBackendQuarantineWiringTest,
FleetdLeadSeatWiringTest, FleetdCompletionResolverWiringTest) rather
than the weaker sibling pattern that builds the wiring itself. No
production change.

The worker branched fresh off 4466ee0 rather than continuing its old
branch off the stranded 435e022 base. That was its own call and it was
the right one: one merge base, a clean 80-line diff, and no cherry-pick
needed this time.

Its vacuity guard is worth keeping in mind for the next source-text
test: it asserts the file it read contains 'public final class Fleetd'
before asserting anything about the mutation, with a message saying the
assertFalse below would pass vacuously on a broken read. An unrelated
anchor is the right choice there, because a guard that shares the
mutation's text cannot tell a bad read from a real change.
2026-09-10 20:50:13 +07:00
Dai Ha 72d6a6878b fleetd #474 follow-up: pin main's config wiring against M2
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m32s
Fleetd.main's own choice of the three-argument ConfigRef constructor
(with Fleetd::assertChartersNameOnlyRegisteredTools as extraValidation)
was unpinned. Reverting Fleetd.java:154 to the plain two-argument
constructor compiled with 0 errors and left the whole suite green,
because ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest
each build their own ConfigRef directly rather than through main.

Adds FleetdConfigRefWiringTest, a source-text check on Fleetd.java
following the FleetdBackendQuarantineWiringTest/FleetdLeadSeatWiringTest/
FleetdCompletionResolverWiringTest precedent: asserts the exact
three-argument construction is present, asserts the plain two-argument
form is absent, and guards against a vacuous pass on a broken/empty
source read by first asserting an unrelated anchor is present.
2026-09-10 20:48:05 +07:00
Dai Ha 4466ee0ef2 fleetd #474: ConfigRef.reload() runs the charter tool-surface gate too
CI / build (push) Successful in 1m37s
CI / contract (push) Successful in 1m44s
A charter naming an MCP tool the server does not register refused Fleetd.main
at startup but slipped through ConfigRef.reload(), because reload() only ran
FleetConfig.validateAll(), which never looks at what a charter's text names.

CharterToolSurface stays in the mcp package (config must not depend on it), so
ConfigRef now accepts the check as a Consumer<FleetConfig> extraValidation,
run inside reload()'s same try/catch as validateAll(). Fleetd.main wires a new
package-private adapter, Fleetd.assertChartersNameOnlyRegisteredTools, into
both the startup call site and ConfigRef's constructor, so the two call sites
can never check different things.

Tests: ConfigRefTest (reload refuses/accepts, via a locally-built equivalent
consumer since Fleetd's method is package-private to dev.ltms.fleet) and the
new FleetdConfigRefCharterToolSurfaceWiringTest (same proof through the exact
Fleetd::assertChartersNameOnlyRegisteredTools reference production uses).
Verified deleting the new extraValidation.accept(fresh) call site fails both
new "refuses" tests by name.

(cherry picked from commit 97e4c1d658)

Lead review. Cherry-picked, not merged: the worker branched from 435e022,
which is not an ancestor of main (I had reset and rewritten that commit's
message while the worker was already on it). Merging its branch would have
added a second merge commit for work already on main. Same content, no
duplicated history. Rule learned: once a worker is spawned, its base
commit is published.

My own mutation battery, five cells, each a full `mvn -B clean test` on
this commit's own tree:

  CONTROL 1, unmutated       1633 tests, 0 failures, BUILD SUCCESS
  M1 delete accept(fresh)    KILLED  2 by name
  M3 3-arg ctor stores no-op KILLED  2 by name
  M4 set(fresh) before valid KILLED  6
  M5 accept outside try      KILLED  2 errors
  M2 main uses the 2-arg ctor SURVIVED

M2 is a real gap and it is a test gap, not a defect: the production code
here is correct. Reverting Fleetd.java:154 to `new ConfigRef(configPath,
cfg)` turns the live reload gate off and leaves all 1633 tests green,
because both new tests build their own ConfigRef with the method
reference rather than reading what main wires.

That is the fourth Fleetd.main call site to survive a battery (#446 M5
and M7, #466 M1, now this). Extracting a check into a well-tested helper
moves the untested surface UP, into the line that chooses to call it. The
repo already has the answer in three source-text wiring tests
(FleetdBackendQuarantineWiringTest, FleetdLeadSeatWiringTest,
FleetdCompletionResolverWiringTest); the worker followed the other
sibling pattern, which builds the wiring itself and so cannot pin main.
Delegated as a follow-up, with the surviving mutation as its acceptance
criterion.

Also mine to correct: my M3 cell's annotation said the no-op store "must
be 2" and measured 1. The 2-arg constructor delegates rather than
assigning, so my pattern only ever matched the mutant. The cell still
stands on its other proof (requireNonNull left = 0). Third battery in a
row with a wrong CONTROL annotation of my own.
2026-09-10 20:40:49 +07:00
Dai Ha 25ba7f16bb Merge #473: fleet_profiles and fleet_list report which attempt a quarantine is on (fleetd #466 item 2)
CI / contract (push) Successful in 1m20s
CI / build (push) Successful in 1m24s
The escalating cooldown landed in 789b6a8, but the only thing an operator
could see was quarantinedForSeconds. A long number does not say whether
this is the first exhaustion or the fifth, and after escalation those look
the same from outside: 3600 seconds could be a big base cooldown or a
credential on its fourth strike.

Both reports now carry quarantineAttempt beside quarantinedForSeconds.

The part that matters is HOW, not that the field exists.
BackendQuarantine#status(credentialId) does ONE quarantines.get() and
returns both values from the same QuarantineState. It is not two accessors
that a caller pairs up. That is the pattern
CompositePeerLauncher.modelGateState() already documents: a report that
reads a different source from the behaviour it describes will eventually
disagree with it, and the disagreement is invisible because both halves
look right on their own.

This is reporting only. Escalation, the ceiling and the quiet-gap reset are
untouched. The flat two-argument constructor still reports a real growing
attempt count even though its cooldown stays flat - which is the honest
answer, since the streak is real whether or not the cooldown uses it.

The worker's report was lost: its fleet_reply never arrived and its inbox
was empty. Nothing was lost with it, because the brief required pushing
first and putting the report in the PR body. That is the second time that
practice has saved a turn.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:21:14 +07:00
Dai Ha 5ba69c9cf2 fleetd #393: correct a comment that claimed two tests pin a call they cannot
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m51s
The comment above the charter writer said flipping that call back to
putArray was "proven load-bearing" and pointed at two named tests. I
mutated exactly that, on the merge commit, and it SURVIVED at 1618 green -
so neither named test covers it, and neither can.

The reason is structural, not a missing test: that writer runs first
against an empty array, so putArray has nothing to replace and the two
idioms are equivalent there. No test can distinguish them. The worker
measured the same thing independently and said so in their report; the
comment was left over from the pre-fix state, where the mutation being
described was a DIFFERENT one (a later writer destroying the charter).

A comment that names tests which do not cover the line is worse than no
comment. The next person mutates the line, sees green, and concludes the
tests are broken.

The corrected version states what each cell actually measures: this line
is unpinnable and why, the later writers ARE pinned and by which test
names, and deleting this line entirely fails the charter-reaches-the-member
assertion - the hole that predates #393.
2026-09-10 20:17:21 +07:00
Dai Ha 17052bb515 Merge #471: seeded skills reach an opencode member, and instructions[] stops depending on write order (fleetd #393)
memberSkills: copied skill folders into every provisioned worktree's
.claude/skills/ and stopped there. Claude Code reads that directory
natively; opencode never does. So the feature was INERT for opencode
members rather than broken: the copy succeeded, the files were correct,
and nothing ever read them. No test failed because there was nothing to
fail - the feature worked at the only layer it implemented.

An opencode member's only channel for static guidance is the instructions[]
array in its generated config. OpenCodeLauncher now adds each seeded
skill's SKILL.md there. A folder with no SKILL.md is never delivered and
the log names it.

The second half is an ordering hazard fleet01 found by reading, and that I
then measured. Three writers append to instructions[]: the role charter,
the seeded skills, and the IDE rules. The charter used putArray
(CREATE-OR-REPLACE) while the other two used withArray (get-or-create).
That was safe only because the charter ran first against an empty array -
an undeclared constraint that nothing tested. Measured on the earlier
merge: making the skills writer use putArray left 1603 tests green while
silently deleting the charter entry, so an opencode member would launch
with no role contract at all. Worse than the bug being fixed, and
invisible.

All three writers now use withArray, and three tests pin the array's
CONTENTS (never its size - a size assertion passes when putArray swaps two
entries for two others) across the combinations that matter:
charter-only, charter+ide, charter+skills+ide.

WHY NOTHING CAUGHT IT, MEASURED RATHER THAN ASSUMED. fleet01 first said
the missing axis was the COMBINATION of writers, then revised that to a
stronger claim: that nothing asserted the charter reaches an opencode
member at all. I checked the second claim on this merge and it is FALSE.
Deleting the charter writer outright fails 6 tests, and 3 of those existed
before #393: writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet,
roleCharterWithoutMcpOrCustomProviderStillWritesAConfig, and
aPinnedEndpointAndTheFleetMcpCoexistInOneConfig. The charter reaching a
member was already pinned.

So their FIRST diagnosis was the right one. The surviving mutation did not
delete the charter write; it made a LATER writer replace the whole array.
Every pre-existing test had exactly one writer active, and with one writer
putArray and withArray are indistinguishable. The gap was never "is the
charter delivered" - it was "are two writers ever active at once", which is
the combination axis. Recording this because the stronger claim is the more
quotable one and it would have sent the next reader looking for a hole that
is not there.

One honest residue: the charter writer's own idiom cannot be pinned.
Flipping it back to putArray leaves the suite green, and always will,
because it runs first against an empty array where the two idioms are
equivalent. The edit removes an undeclared constraint for the next person
to add a writer; it is not a change any test can detect. The comment in
the source claimed two named tests cover it - that was wrong, and the
commit after this one corrects it rather than leaving a false claim beside
the code.

What this does NOT do: opencode has no equivalent of Claude Code's Skill
tool, so the content is static system-prompt text present from spawn, not
something a member can invoke by name. This closes the DELIVERY gap and
cannot close the ACTIVATION gap. That residue is opencode's design.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:17:21 +07:00
Dai Ha e95ed99bf7 fleetd #466 scope item 2: report the quarantine repeat count, not only the seconds
CI / build (pull_request) Successful in 1m31s
CI / contract (pull_request) Successful in 1m33s
Add BackendQuarantine#status(credentialId) -> Optional<Status>, a single
QuarantineState read that answers both remainingSeconds and repeatCount
together -- the same "one accessor" pattern CompositePeerLauncher.
modelGateState() already uses, so the two facts can never disagree.

fleet_profiles/REST GET /profiles and fleet_list's capacity rows
(FleetMcp.profilesView/capacityView) now call status() instead of
remainingSeconds() and add a "quarantineAttempt" field beside
"quarantinedForSeconds": 1 for a first occurrence, 2 for the second
in a row, and so on.

No change to the escalation, ceiling, or reset logic itself -- this
unit is reporting only.
2026-09-10 20:09:13 +07:00
Dai Ha 1477e4358a Merge #472: one canonical tool-name set, and charters are checked against it (fleetd #469)
CI / contract (push) Successful in 57s
CI / build (push) Successful in 2m0s
A role charter is free text in config that tells a member which tools to
call, and nothing checked that those tools exist. A charter naming
bridge_send - a name CB-634 removed - started the daemon cleanly, and the
member found out at run time by calling something that was not there.

#464 shipped a test for this, but it wrote its own charter into a @TempDir,
so nothing anyone put in the real config could fail it. That was a defect
in my acceptance criteria, not in that work.

Now: FleetTool is one enum of the 11 registered wire names, and every
reader goes through it.

- FleetMcp's schema builders pass FleetTool.X.wireName() instead of a
  literal.
- FleetMcp's constructor asserts at startup that what it registers with the
  SDK equals FleetTool.wireNames() exactly, in both directions.
- toolAction(String, Map) resolves arbitrary wire input against
  FleetTool.byWireName() and keeps its run-time throw, which is necessary -
  network input has no closed compile-time form. It then hands off to
  authzAction(FleetTool, Map), a switch over the enum with NO default, so a
  new tool is a compile error at that layer.
- CharterToolSurface lives in mcp, not config, and Fleetd.main calls it
  right after validateAll(). Config must not depend on the MCP server:
  config loads before the server exists.

My brief undercounted the problem and the worker corrected it. I said there
were two tool-name inventories plus a test fixture. There were FIVE: the
registrations, the authz switch, and three separate source-text scrapes of
FleetMcp.java in CharterToolSurfaceTest, FleetMcpAuthzTest and
McpContractDocTest - none of which the ticket mentioned. Fixing those three
was required, not scope creep: once the literals moved into FleetTool their
regexes matched zero names, so one would have failed on its vacuity guard
and the other two would have gone quietly vacuous. Inventories after: one.

Verified independently on origin/main before accepting the wider diff:
three test files did read FleetMcp.java as source text, with a control file
at zero to prove the search discriminated.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:59:03 +07:00
Dai Ha 9e4e423ad6 fleetd #393 follow-up: remove the instructions[] writer-ordering hazard
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 2m11s
OpenCodeLauncher.writeConfig has three writers into the instructions[]
array (charter, seeded skills, IDE rules). The charter writer used
putArray (create-or-REPLACE) instead of withArray (get-or-create), which
"worked" only because it happened to run first against a still-empty
array — an undeclared ordering dependency nothing tested. Found by the
fleet01 lead and verified on this branch's merge: flipping the skills
writer to putArray left the full 1603-test suite green while silently
deleting the charter entry, which would launch an opencode member with
no role contract at all.

Fix: charter's putArray -> withArray (one-word change, behavior-identical
today). Add three tests asserting instructions[] CONTENT as an exact
ordered list (not size) across writer combinations: charter only,
charter + IDE rules, and charter + IDE rules + seeded skills. Mutation
testing (see PR body) shows the skills and IDE-rules writers are each
independently detectable by name; the charter writer's own mutation is
not detectable by any test, because it structurally always runs first
against an empty array, so putArray and withArray are equivalent there.
2026-09-10 19:58:25 +07:00
Dai Ha 6c2d6e93cb fleetd #469: one canonical FleetTool set backs registration, authz and charter checks
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m37s
FleetConfig.validateCharters() only checked that a charter key is a role
wire name and its text is non-blank; #464's CharterToolSurfaceTest compared
charter text against the registered tool surface, but wrote its own charter
into a @TempDir fixture, so nothing anyone wrote into the live fleetd.yaml
could ever fail it.

Add FleetTool, an enum in dev.ltms.fleet.mcp holding the one canonical set
of registered tool wire names. FleetMcp's tool schemas now derive their
names from it, its constructor asserts at startup that what it actually
registers with the SDK equals FleetTool.wireNames() exactly, and its authz
dispatch (toolAction/authzAction) resolves the wire string against FleetTool
before switching on the enum itself with no default -- adding a tool without
pinning its Authz.Action is now a compile error, not just a test gap.

Add CharterToolSurface (mcp package, not config -- config loads before the
MCP server exists) and call it from Fleetd.main right after
cfg.validateAll(), so a charter naming a tool the server does not register
refuses the daemon's startup, naming both the charter key and the unknown
tool. FleetdStartupValidationTest proves this through Fleetd.main itself
against a live-shaped config fixture (bridge_send, CB-634's own removed
name).

CharterToolSurfaceTest, FleetMcpAuthzTest and McpContractDocTest each kept
an independent regex scrape of FleetMcp.java's source for the registered
side of their own comparison -- three more copies of the same list nothing
tied together. All three now read FleetTool.wireNames() instead.

Proved canonical by removal: deleting FleetTool.ACK while ackTool() still
referenced it broke mvn compile in two places (FleetMcp.java:940,:1898);
registering a schema under a literal not backed by FleetTool
("fleet_ack_v2") failed FleetMcp's new startup assertion in every test that
constructs it (12 errors, IllegalStateException at FleetMcp.<init>). Both
reverted before this commit.
2026-09-10 19:55:02 +07:00
Dai Ha 789b6a8716 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
CI / contract (push) Successful in 1m56s
CI / build (push) Successful in 2m4s
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things on the record because they are judgement calls, not facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this codebase
  reports a spawn success back to this class, so "it started working again"
  cannot be observed here. A base cooldown of quiet is the best available
  evidence, and the class doc says that plainly instead of implying the
  stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - the same
formula with multiplier 1.0 and a ceiling equal to the base, which collapses
to the original flat "now + cooldown". Every existing call site keeps its
shape.

The second commit exists because my own mutation on the first merge found
the wiring unpinned: putting Fleetd.main back on the flat constructor left
all 1608 tests green, so the factory was pinned and the decision to use it
was not. FleetdBackendQuarantineWiringTest closes that, following the five
existing *WiringTest files rather than inventing an idiom. It checks source
TEXT, and its class doc says so: it does not prove the call executes, and it
cannot tell "wrong factory" apart from "renamed the anchor" - both fail the
same assertion. That limit is real and recorded rather than papered over.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:51:00 +07:00
Dai Ha 01462c9695 fleetd #466 follow-up: pin main's choice of the escalating quarantine factory
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m25s
BackendQuarantineTest proves BackendQuarantine.withEscalation itself
escalates, is ceilinged, and resets. Nothing proved Fleetd.main still
calls it: reverting main to the flat two-argument constructor compiled
with 0 errors and left the full 1608-test suite green, because every
other test builds its own BackendQuarantine directly.

Adds FleetdBackendQuarantineWiringTest, a source-text assertion on
Fleetd.java (same idiom as FleetdLeadSeatWiringTest and
FleetdCompletionResolverWiringTest) pinning that main's declaration is
built from BackendQuarantine.withEscalation(...), not `new
BackendQuarantine(...)`. Measured: reverting to the flat constructor
fails it (expected true, was false); renaming the anchored local
variable while keeping behaviour identical also fails it loudly, not
silently. The class doc states plainly this checks source text only,
not that the call executes or that escalation reaches a real backend.
2026-09-10 19:47:59 +07:00
Dai Ha 8b4ff78546 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things I want on the record because they are judgement calls, not
facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this
  codebase reports a spawn success back to this class, so "it started
  working again" cannot be observed here. A base cooldown of quiet is the
  best available evidence. The class doc says this plainly rather than
  implying the stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred classification question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and is deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - it is the
same formula with multiplier 1.0 and a ceiling equal to the base, which
collapses to the original flat "now + cooldown". Every existing call site
keeps its shape.

Build number and my own mutation results are reported in the PR and on the
ticket, measured on this merge commit rather than on the branch.
2026-09-10 19:40:15 +07:00
Dai Ha d4a2cd720c fleetd #393: deliver memberSkills to opencode members, and stop overclaiming seeding success
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m46s
GitWorktrees.seedSkills copies memberSkills:-seeded skill folders into every
provisioned worktree's .claude/skills/ and logged "skill seeding: N of M" as
if that were success — but .claude/skills/ is a Claude Code CLI convention.
opencode has no such discovery, so a kind: opencode member never actually
read a seeded skill even though the log said N of M succeeded.

Two changes, both required:

1. Deliver it. OpenCodeLauncher.skillInstructionFiles scans
   <cwd>/.claude/skills/*/SKILL.md at spawn time (the one point the launcher
   knows both the kind and the cwd) and appends each to the generated
   opencode.json's instructions[] array, the same channel already used for
   the member charter and IDE rules. A skill folder with no SKILL.md is
   named and skipped rather than silently dropped.

2. Stop claiming it where the claim can't be verified. GitWorktrees.seedSkills'
   log now says explicitly that consumption depends on the member's kind and
   points at the launcher's own log; OpenCodeLauncher logs its own kind-aware
   "skill delivery: M of N ..." line once the kind is actually known, naming
   any folder it could not turn into an instructions[] entry.

fleetd.example.yaml's memberSkills: doc previously claimed "Claude Code
members only; an opencode member reads a different path (.opencode/agent)
this key does not touch" — false as of this fix, corrected to name both
kinds and how each consumes it.

Tests: OpenCodeLauncherTest gains two cases driving the real
GitWorktrees#add seeding path (not a hand-built fixture) into an
opencode-kind spawn — one asserting a seeded skill's SKILL.md lands in
instructions[] plus the honest log line, one covering a skill folder
without SKILL.md (delivered skills still flow, the malformed one is named
in the log and excluded from instructions[]). ClaudeCodeLauncher is
untouched — its native .claude/skills/ discovery already worked and is out
of scope.

mvn -B clean test: Tests run: 1603, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
2026-09-10 19:37:00 +07:00
Dai Ha 5a467e1f8b fleetd #466: escalate BackendQuarantine's cooldown on repeated exhaustion
CI / build (pull_request) Successful in 1m34s
CI / contract (pull_request) Successful in 1m36s
A flat 30-minute quarantine retries a weekly subscription limit about
336 times before the window resets. BackendQuarantine now doubles the
cooldown on each consecutive exhaustion of the same credential (no
more than one base cooldown after the previous quarantine's deadline),
capped at 12x the base cooldown (~6h at the 1800s default), and resets
back to the base cooldown once a base-cooldown's worth of quiet has
passed with no further exhaustion.

The flat two-argument constructor is unchanged (equivalent to
multiplier 1.0 / ceiling == base), so all ~20 existing call sites keep
their current shape and behaviour. Production wiring (Fleetd.main)
switches to the new BackendQuarantine.withEscalation factory.

This only touches the exhaustion path (BackendQuarantine's one
production caller is Fleetd.exhaustionSink, fired on BACKEND_EXHAUSTED
alone) and never the separate, unescalated cooling-off mechanism
(BackendOutagePolicy, fixed 60s) that guards against a transient
backend-error storm.
2026-09-10 19:33:09 +07:00
Dai Ha 1fb6176783 Merge #457: exhaustedPattern goes hot, and the warning names the fix (fleetd #446)
CI / contract (push) Successful in 1m31s
CI / build (push) Successful in 1m33s
Three rounds. Round 1 made exhaustedPattern a hot config key, made the
quarantine warning name the fix instead of only the fact, and added model/reason
to fleet_profiles' quarantined rows. Round 2 extracted the warning text into
usageLimitFixWarning/usageLimitFixWarningNoModel and pinned both. Round 3
extracted the sink itself into a static exhaustionSink(...) factory and pinned
what it actually logs, using a ListAppender on this class's own logger.

Where the mutation numbers below come from, stated exactly. The battery ran on
merge commit 3d2d521, tree 1dcca23: origin/main at 235644c plus this branch.
This merge commit's tree is 953ce11, which is that tree plus one file --
CharterToolSurfaceTest, from the #464 merge (49df792) that landed on main while
the battery was running. So the battery did not run on this exact tree. That one
added file was verified green on its own merge. The build on THIS tree, tree
953ce11, is: Tests run: 1601, Failures: 0, Errors: 0, Skipped: 0, BUILD
SUCCESS, 0 compile-error blocks. That is the battery tree's 1600 plus that one
test, which is the arithmetic the two trees predict. Saying this rather than
implying one tree.

On the battery tree: 1600 tests green, 0 compile-error blocks. FleetMcp.java
auto-merged there against #463's change to the same file, so that build was also
the gate on the auto-merge -- a clean auto-merge is not a compiling merge.

- M5, the round-2 survivor reproduced verbatim: the call site stops using either
  pinned method and logs a literal instead. Round 2 left this green across 1592
  tests. Now KILLED by
  FleetdExhaustionSinkWarningTest.profileWithModelGetsTheActionableFix and
  .profileWithNoModelGetsTheFallbackNotAFix.
- M6, the ternary's two branches swapped, so a profile WITH a model gets the
  no-model text and vice versa. Both extracted methods and both their unit tests
  untouched. KILLED by the same two tests, independently.
- M7, THE HALF THIS ROUND DID NOT PIN, and it SURVIVED. main()'s call to the
  factory replaced with an inert lambda: the factory and its test stay perfect,
  the daemon quarantines nothing and logs nothing. 1600 green.

M7 is the same defect shape as round 2, moved one level up, and it is worth
naming plainly. Extracting a thing in order to pin it CREATES the seam the test
then lands on. Round 2 extracted the warning TEXT, pinned the two methods, and
left the call site that selects between them unpinned -- M5. Round 3 extracted
the SINK, pinned the factory's behaviour, and left main()'s wiring of it
unpinned -- M7. The question to ask on any such fix is which of three things a
test now reaches: the value, the call site, or the selection between values.
Extraction only ever answers the first.

M7 is not a reason to hold this merge. Proving main() wires this factory means
driving daemon startup, which nothing here does -- that is fleetd #460. A
cheaper option exists and is recorded there: a source-reading assertion of the
kind #439 used, which would kill M7 without starting the daemon.

Round 3's own caveat, disclosed by the worker and worth keeping: its first Cell A
run was corrupted by a second concurrent mvn against the same module directory.
It killed that run, checked for stray ForkedBooter processes, and re-ran. The
numbers above are mine, from this battery, not that run.
2026-09-10 19:16:24 +07:00
Dai Ha 49df79203c Merge #464: a test that charter text names only registered tools (fleetd #464)
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m56s
CharterToolSurfaceTest extracts every fleet_* / bridge_* token from configured
launch charters and every tool("fleet_...") FleetMcp registers, then asserts the
first set is a subset of the second.

Verified on the merge commit. Its three acceptance criteria are met:

- Catches the ticket's own example. Fixture charter naming bridge_send, the tool
  CB-634 renamed away: KILLED.
- Fails loudly with no charter text. Fixture stripped of every tool name: KILLED
  by its named.isEmpty() guard, not a silent pass.
- Fails loudly with no registered tools. The tool("...") scrape broken so it
  matches nothing: KILLED by its registered.isEmpty() guard.
- And one cell of my own: the server stops registering fleet_reply, which the
  fixture names. KILLED. This is what proves the 'registered' half reads real
  production source and is not a second fixture.

The scrape finds 11 registered tools: ack, ask, list, poll, profiles, reply,
send, spawn, status, stop, whoami. An independent count of every "fleet_x"
literal in FleetMcp.java is also 11.

WHAT THIS DOES NOT CLOSE, and it is the ticket's actual gap. The charter half is
a @TempDir fixture the test writes itself, so no charter text anyone writes can
make this test fail. Measured: the test mentions fleetd.yaml 0 times, and the
commit changes 0 production files -- FleetConfig.validateCharters() still never
reads charter text (0 lines of its body mention a tool name). So this pins the
comparison logic and acts as a rename tripwire for the two tools the fixture
names. It does not check the live config. That needs a production-side check and
is filed as a follow-up.

That residue is my ticket's fault, not the worker's: the three criteria I wrote
are exactly the three it met.

A note on my own battery, because it nearly published four false kills. The
first run showed rc=1 on all four mutation cells and I would have read that as
four kills. It was zsh: unquoted parameters are NOT word-split, so
'mvn -B $scope test' passed '-Dtest=X -DfailIfNoTests=false' as ONE argument and
surefire ran zero tests. The tell was a missing 'Tests run:' line. The rerun
proves the harness first -- selector alone must report 'Tests run: 1' -- and
every cell now prints its surefire summary count so a void cell cannot pass for
a kill.
2026-09-10 19:10:56 +07:00
Dai Ha 235644c0f0 Merge #467: listFleet's callerIsPrimary default fails closed (fleetd #463)
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m38s
A wrapper overload that is called with no callerIsPrimary argument used to
default it to true, so a forgotten argument silently handed out the lead's
coordination state. It now defaults to false: a missing identity fails closed.

Verified on the merge commit, not the branch:

- The funnel is real. FleetMcp has 7 listFleet declarations and 7 real calls
  (an 8th 'listFleet(' match is a javadoc {@link}). Exactly one call writes a
  literal for the new boolean, and it writes false; exactly one writes the real
  predicate, coordinatorVisibleTo(principal(exchange)) in the MCP handler. No
  call writes true.
- Control battery on the merge: 1584 tests green unmutated.
- M1, the fix reverted at the one line that writes the default (false -> true):
  KILLED by listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey.
- M2, the half this round weakened. The worker rewrote
  listIsByteForByteUnchangedForThePrimaryCaller and dropped its byte-for-byte
  equality assertion, which was correct because that comparison ran against the
  implicit-default overload -- now the path #463 closes. So: does anything still
  notice if the primary's coordinator row silently loses a field? Dropped
  heldCount: KILLED by two tests, that same test and
  listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero.
- M3, a regression check on #439's own gate, 'if (callerIsPrimary)' -> 'if (true)':
  KILLED by three tests.

What is no longer pinned, stated plainly: the primary's output is now checked
field by field, not as a whole string. A field that no test names could
disappear without failing anything. Every field the tests do name is pinned,
proven by M2. The whole-output answer belongs to fleetd #460.

One control in my own battery was wrong and is worth recording: I labelled
', true);' in FleetMcp.java 'must be 0' and it is 2 -- row.put("configured",
true) and m.put("self", true), neither a listFleet delegation. The pattern was
too wide. The count above comes from a walk over each declaration and call
instead.
2026-09-10 19:04:51 +07:00
Dai Ha 7772b41993 fleetd #446 round 3: extract exhaustionSink and pin its caller (Cell A/B)
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 2m13s
Round 2 pinned usageLimitFixWarning/usageLimitFixWarningNoModel's TEXT via
FleetdUsageLimitFixWarningTest, but a mutation battery against the merged PR
proved two gaps in the caller that builds main()'s real quarantine
ExhaustionSink: nothing proved the sink's log.warn actually invokes either
method (Cell A), and nothing proved it picks the right one for a profile
with vs without a configured model: (Cell B).

Extract the inline lambda into a new static Fleetd.exhaustionSink(...)
factory (same refactor-for-testability class the lead approved in round 2
for the two static warning methods), behaviourally unchanged from the
lambda it replaces. FleetdExhaustionSinkWarningTest drives this factory's
return value directly and asserts on the real text a ListAppender attached
to Fleetd's own logger captures, covering both cells from one mechanism:

- Cell A (ternary result replaced by a literal string): confirmed red,
  1594 run / 1 failure, restored, confirmed green (1594/0).
- Cell B (ternary's two branches swapped): confirmed red, 1594 run /
  2 failures (both test methods independently caught it), restored,
  confirmed green (1594/0).
2026-09-10 19:00:07 +07:00
lead eccd0548ce Merge #465: state canonical invariant 5 as a purpose, not a banned tool (fleetd #458)
CI / contract (push) Successful in 52s
CI / build (push) Successful in 2m1s
Verified by me on a local merge of 29e7a06 onto f5e02fe:

- One file, one hunk. 5 changed lines inside the canonical block, all of them
  invariant 5; a line-by-line diff of the block against origin/main shows
  nothing else moved.
- The addendum is untouched: both 'Herdr socket tests (measured 2026-09-10)'
  and 'Only the lead can run that check' appear once on each side.
- Block size: origin/main 17467 chars / 17626 bytes, the merge 17557 chars /
  17718 bytes. The worker's numbers were bytes and were right; I am naming both
  units because len(str) in python is characters and this block is full of em
  dashes, which is a mistake I have made before.
- mvn -B clean test: Tests run: 1583, Failures: 0, Errors: 0, Skipped: 0.
  BUILD SUCCESS, rc=0.

The three things I reserved from the worker, because a provisioned worktree
cannot do them:

- wiki/7-Use-Cases.md updated to match, in the wiki repo's own commit.
- The sync script run in this clone: in sync: True.
- Other projects carrying the block: NONE. 29 CLAUDE.md files under the
  operator's project roots, 18 of them carry the block and 17 of those are this
  repo's own worktrees, which inherit it from git. Only
  LTMS/claude-bridge/CLAUDE.md is a real second copy, and it is this file.

On the worker's criterion-4 question: the addendum note stays. The restated
invariant removes the unsatisfiability, which was the sharp problem, but the
note still names the exact four files the carve-out covers and records why
(fleetd #449, where a fake and the real herdr disagreed for weeks under a green
suite). A rule that is merely satisfiable is not the same as a rule a worker can
apply without re-deriving it.
2026-09-10 18:57:24 +07:00
Dai Ha c4e23eebad fleetd #464: guard charter tool names
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m37s
2026-09-10 18:55:16 +07:00
Dai Ha 7df503dfc2 fleetd #463: default listFleet's callerIsPrimary to false, fail closed
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 2m9s
The compat overload at FleetMcp.java:1373 defaulted callerIsPrimary to a
literal true, so a caller that forgot the argument silently got the
coordinator row (this daemon's coord-id, mailbox state, held-mail previews,
peer reachability) -- lead-to-lead state fleetd #439 just gated. Flip the
default to false: a forgotten argument now yields a missing row instead of
a leaked one.

Six FleetMcpTest methods relied on the implicit true to see the coordinator
row at all; they now pass true explicitly through the canonical overload.
listIsByteForByteUnchangedForThePrimaryCaller's own premise (comparing the
implicit-default path against an explicit-true path) was the shape of the
bug, so it now only exercises the explicit-true path.

Added listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey
to pin the new default: a compat overload called with no callerIsPrimary
argument, against a fully-configured lead channel, must produce a result
with the coordinator key absent -- not empty, not redacted, absent.
2026-09-10 18:55:13 +07:00
Dai Ha f5e02fedd6 plans: commit the fleet01 move plan, with today's state measured on the host
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m58s
This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.

Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:

- Phases 1 to 3 are done. Both systemd user units exist, are active and
  enabled, linger is on, one java process (so the old restart.sh
  double-daemon problem is gone), the live fleetd.yaml is on the host, and
  the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at 4887731, 77 commits behind
  origin/main, 0 ahead. So fleet01's daemon runs code from before this
  week's merges.
- Phases 6 and 7 are not done. The Mac still runs the daemon this fleet
  uses, and fleet01 has 2 weeks of uptime, so no reboot proof exists.
- Section 9 is out of date: claude is installed on fleet01 and a fleet01
  lead is live on the coordination channel, so the headless-login blocker
  it names is solved.

Also records fleet01's four profiles with their weights, because
placement: weighted plus gx at 100 and xf at 80 means an unqualified spawn
there almost never lands on a claude-code profile.

And one search that found the opposite of what I expected, written down so
the next session does not repeat it: opencode IS installed on fleet01, at
~/.opencode/bin/opencode, but only on the INTERACTIVE PATH. zsh -ic finds
it, zsh -lc does not. A herdr pane is a non-login interactive zsh and
fleetd types the launch command into the pane, so gx and xf are not broken
by this. The daemon's own login-shell ExecStart cannot see it, which is the
mirror of the trap that ExecStart exists to fix: credentials live on the
login side, ~/.opencode/bin on the interactive side.
2026-09-10 18:53:11 +07:00
Dai Ha 29e7a06c49 fleetd #458: restate invariant 5 by purpose, not mechanism
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Failing after 1m56s
Invariant 5 banned 'driving the terminal multiplexer directly', naming
herdr CLI and socket as the banned tool. That bans a mechanism. What it
protects is the control plane: nobody may move a fleet session, pane or
peer by a route that skips the bridge's policy checks.

In this repo herdr is itself the subject under test, so four contract
tests must open its socket on purpose (see the addendum's 'Herdr socket
tests' note). Under the old wording, a worker assigned to that code
reads invariant 5 and finds its only path to finish the task banned.

Restate the invariant by purpose: never move fleet state except through
the bridge. herdr's CLI and socket stay as the named example of the
banned route, not the definition of it.
2026-09-10 18:50:41 +07:00
Dai Ha 92a96fcbd8 Merge #462: gate the coordinator row on a named predicate the handler must consult (fleetd #439)
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 1m51s
Verified by me on a local merge of c1ca627 onto 1348287:

- mvn -B clean test: see the totals below. Ran in fleetd/, redirected to a file,
  exit code captured on its own line.
- Mutation battery on merge 8402923 (same two parents, earlier base 5f1b260),
  4 cells, each with a proof gate on the occurrence count:
    * CONTROL, unmutated: 1583 tests, 0 failures, rc=0.
    * M2r - the handler stops asking who called: coordinatorVisibleTo(principal(exchange))
      replaced by a literal true. KILLED by
      FleetMcpAuthzTest.theFleetListHandlerActuallyConsultsCoordinatorVisibleTo.
      This is the mutant that survived round 1, so the gap the follow-up was for is closed.
    * M3 - is the new source-reading detector vacuous? Renamed its anchor
      (listHandler -> listHandlerX, 2 sites, behaviour identical). The detector FAILED,
      rc=1, as a source-reading test must: it cannot silently pass on an empty scrape.
    * M4 - the predicate itself always says yes: return caller.isPrimary() replaced by
      return true. KILLED by FleetMcpAuthzTest.onlyThePrimaryMaySeeTheCoordinatorRow.
  Tree verified clean before the battery and restored after each cell.
- ANON is pinned too, not only worker and architect: FleetMcpAuthzTest asserts
  coordinatorVisibleTo(ANON) is false.
- listFleet( appears in exactly one main file, mcp/FleetMcp.java, and no REST class builds
  the coordinator row, so the detector's single-file scope covers every live call site
  today. Control for that sweep: 109 main .java files matched a string they all contain.

Not in this PR, and my call, not the worker's: the six compat overloads of listFleet still
default callerIsPrimary = true, which fails open. Safe today because the one production
call site passes the computed value. Filed separately.
2026-09-10 18:40:48 +07:00
Dai Ha 13482872bb contract test: keep the loaded run's passes, discard only its failure
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 2m6s
The javadoc threw away the whole loaded run as "not clean evidence". The
fixed cell order makes that too strong in one direction. A cell that runs
last on a climbing load has a free explanation for FAILING. It has no free
explanation for PASSING: surviving a worse condition than a fair order
would have given it is evidence in the safe direction. So the old version's
0 of 3 is still discarded, and the three passes are kept with the load each
one ran at.

Also names the cell that tests the swallow explanation head-on, which the
old text left as "nobody has managed that yet". SHELL_READY_TIMEOUT_MS at 0
types input at once — the worst case for "typed before the prompt" — and it
passed 3 of 3 at load 18.42 to 23.65. The wider read window cannot explain
that away, because a swallowed keystroke is lost, not late: the command
never runs, so no amount of polling makes its output appear.

And a warning not to carry the raw load average to another host. Load
average counts differently per core and per operating system, so only load
per core compares. I broke that rule myself when comparing this run with
another host's numbers.

Both points came from the fleet01 lead reviewing 2af13ab.
2026-09-10 18:37:41 +07:00
Dai Ha c1ca6273fc fleetd #439: pin the caller at the fleet_list call site, not just the gate
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m33s
Review of PR #462 found M2: the coordinatorVisibleTo gate (then an inline
principal(exchange).isPrimary() check) could survive a mutation that
replaced the argument with a literal true at the one production call
site, because every existing test drove listFleet directly and supplied
the boolean itself -- nothing exercised the handler's own call.

- Name the decision: FleetMcp.coordinatorVisibleTo(Principal), a small
  package-private predicate next to denyFor/recordPrimarySingleton. The
  fleet_list handler now calls coordinatorVisibleTo(principal(exchange))
  instead of inlining .isPrimary().
- Pin the predicate's role table in FleetMcpAuthzTest
  (onlyThePrimaryMaySeeTheCoordinatorRow), covering primary/worker/
  architect and, newly, anonymous.
- Add a source-reading detector at the boundary
  (theFleetListHandlerActuallyConsultsCoordinatorVisibleTo), same idiom as
  toolsTheServerRegisters/everyRegisteredToolHasItsHandlerActionPinned: it
  reads FleetMcp.java, isolates the listHandler block, asserts (as a
  control) that the block actually contains a listFleet( call, then
  asserts the call's trailing boolean argument is exactly
  coordinatorVisibleTo(principal(exchange)) -- not a literal true/false.

Both mutations from the review were reproduced and killed by these tests,
then reverted; see the PR body for the full break-and-restore transcript.
2026-09-10 18:28:16 +07:00
Dai Ha 5f1b260c81 addendum: the canonical sync check is the lead's, not a member's
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 1m45s
I put "the sync script must print in sync: True" in a worker's acceptance
criteria for #455. The worker could not run it and said so, honestly, instead
of inventing a pass. My brief was the defect.

A member's provisioned worktree has wiki/ uninitialized, so the script dies
with FileNotFoundError. Measured in three worker worktrees: git submodule
status printed a leading '-' and wiki/ held 0 entries. The primary's own clone
printed a leading '+' and the file was there.

The note is dated, gives the re-measure command, says what each outcome means,
and says to delete it once it stops reproducing - as this file requires of any
measurement in an addendum.

Canonical block untouched: the sync script itself reports in sync: True.
2026-09-10 18:25:32 +07:00
Dai Ha 30d6872779 fleetd #446 follow-up: pin the WARNING text and the fleet_profiles model/reason fields
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m4s
A mutation battery run against merged PR #457 (d02dd1b) proved criteria 2 and 3
shipped without a test that could catch them breaking: deleting either
row.put("model", model) or row.put("reason", reason) in FleetMcp.profilesView
left all 1586 tests green (rc=0), and renaming the "usage-limit fix:" log tag
to something meaningless did too. Criterion 1's own mutation (LiveExhaustedPatterns
snapshotting instead of reading live) was correctly killed by the existing
FleetdExhaustionDetectionArmedWiringTest/LiveExhaustedPatternsTest — only 2 and 3
were unguarded.

- Extracted the exhaustionSink WARNING text out of two inline SLF4J {}-placeholder
  log.warn calls into two static methods, Fleetd.usageLimitFixWarning(profile, model)
  and Fleetd.usageLimitFixWarningNoModel(profile, quarantineCooldownSeconds), the
  same extracted-static-method + dedicated-test idiom as
  exhaustedPatternCoverageLine/errorPatternCoverageLine. log.warn is now called with
  each method's return value as a single already-formatted argument, so the string a
  test asserts on is byte-identical to what fleetd.out receives. No behaviour change:
  same text, same two branches, same call site.
- New FleetdUsageLimitFixWarningTest pins the leading "usage-limit fix:" grep tag and
  three specific facts (profile name, model name, "enabled: false" under
  models.allow) rather than the whole sentence, plus the no-model fallback's profile
  name and cooldown-seconds substitution and its explicit absence of "enabled: false".
- New FleetProfilesQuarantineModelReasonFieldsTest exercises FleetMcp.QuarantineSource
  with a modelFor/reasonFor that actually return values (every existing test used
  QuarantineSource.none() or a 3-arg form defaulting both to null), asserting the
  quarantined row's model/reason are present when supplied and absent (not null, not
  blank) when modelFor/reasonFor return null or a blank string.
- Each of the three: implemented, broken by hand (row.put deleted / tag renamed),
  confirmed the new test goes red, restored, confirmed green again. See PR body and
  this ticket's fleet_reply for the verbatim failure output of all three.
2026-09-10 18:22:51 +07:00
ltms 9d1306d442 Merge #461: project addendum for the herdr socket carve-out (fleetd #455)
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m56s
Doc-only, one file, 13 added lines, inside §Project addendum. Verified by me:

- Canonical block byte-identical: 17468 chars on main and on the branch.
- The sync script in CLAUDE.md prints "in sync: True" both before and after.
- The note's own re-measure command works. Run verbatim, it returns exactly the 4 files
  the note names, and no others:
    fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java
  Control that the search reaches the tree: 120 test java files, 63 of them mention herdr.
  That control matters here — a zero-match re-measure command would tell a future session to
  delete a live restriction.
- All four required elements are present: the carve-out, what stays banned (the control plane),
  who it applies to, and the perishable half (dated 2026-09-10, the command, what each outcome
  means, and delete-when-stale).
- Cross-references #458 for the canonical restatement, which is deliberately not in this change.

Pre-send check applied to the note itself: a worker assigned to AgentControlContractTest can now
name one legal action that finishes its task — let the test open the herdr socket in a throwaway
workspace it tears down.
2026-09-10 13:18:17 +02:00
Dai Ha e54e3d87ea fleetd #439: omit fleet_list's coordinator key for non-primary callers
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
The coordinator row is lead-to-lead coordination state (coord-ids, mailbox
facts, held-message previews). fleet_list returned it to every caller,
including a worker or an architect, because coordinatorView() had no way to
know who was asking.

Gate at the call site inside listFleet: a new overload takes
callerIsPrimary and only assembles/attaches the coordinator row when it is
true, so the key is absent (not empty) for a worker or an architect. The
MCP handler now passes principal(exchange).isPrimary(); every other
listFleet overload keeps passing true, so callers with no caller identity
(existing unit tests, the no-op wrappers) are unaffected -- confirmed by a
byte-for-byte comparison test against the pre-fix overload.

Authz's READ case is untouched: it stays shared by fleet_status,
fleet_profiles and fleet_whoami, and the gate here is purely inside
fleet_list's own result assembly.
2026-09-10 18:16:48 +07:00
Dai Ha bf027f10b9 #455: document herdr test socket exception
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m50s
2026-09-10 18:13:50 +07:00
Dai Ha 3c5873dfe2 fleetd #453: point the new javadoc at the right javadoc
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m52s
#456's first paragraph said "HerdrPeerLauncher's own {@link
#spawn(SpawnRequest, PlacementDecision)} javadoc". That link resolves to this
interface's own abstract declaration, not to HerdrPeerLauncher's override, so a
reader who follows it lands on the wrong text. The second paragraph already
used the plain {@code HerdrPeerLauncher.spawn(...)} form; both now match.

Also says who is actually forced to read the paragraph, because #456's
reasoning rests on it and the two cases differ. A class that implements this
interface directly must write a body for spawn(SpawnRequest,
PlacementDecision) - it is abstract here - so it reads this javadoc. A
subclass of HerdrPeerLauncher does not: HerdrPeerLauncher already implements
that method (member/HerdrPeerLauncher.java:611) and the subclass inherits the
body. For a subclass the paragraph is advice, not a gate.

javadoc -Ddoclint=reference: 5 "reference not found", the same 5 in the same 5
untouched files as origin/main at 2af13ab, and none in PeerLauncher.java. Those
5 are ticket #459. mvn compile rc=0.
2026-09-10 18:05:07 +07:00
ltms e29227d5f4 Merge #456: document the override obligation on PeerLauncher.defaultProfileFor/place (fleetd #453)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m34s
Doc-only, one file, 17 added lines. Verified by me on a local merge of 0788d84
onto 2af13ab (merge commit 5b46538):

- mvn -B clean test: Tests run: 1578, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS, rc=0.
- javadoc reference lint (mvn javadoc:javadoc -Ddoclint=reference): the merge has 5
  "reference not found" errors; origin/main at 2af13ab has the same 5, in the same 5
  untouched files. So this PR adds no broken link. Both runs rc=1 for that pre-existing
  reason, which is now ticket #459.
- The quoted phrase is real: "are unoverridden here and just wrap {@link #defaultProfile()}"
  is at member/HerdrPeerLauncher.java:604. place() is not overloaded, so {@link #place}
  is unambiguous.

One inaccuracy I am fixing in a follow-up commit rather than sending the PR back: the first
new paragraph writes "HerdrPeerLauncher's own {@link #spawn(SpawnRequest, PlacementDecision)}
javadoc", but that link resolves to PeerLauncher's own abstract declaration, not to
HerdrPeerLauncher's override. The second paragraph already gets this right with plain
{@code HerdrPeerLauncher.spawn(...)}.
2026-09-10 13:04:09 +02:00
Dai Ha 45aca9eb3e fleetd #446: make exhaustedPattern hot, name the fix in the warning, report it in fleet_profiles
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 1m34s
The model gate can be turned off at runtime (models.allow[].enabled: false, hot
since fleetd #422), but the usage-limit detector it's meant to react to was
compiled once at Fleetd.main startup into a frozen Map<String,Pattern> — arming
or disarming exhaustedPattern needed a daemon restart. Backwards for a feature
meant to react live.

- New LiveExhaustedPatterns: reads exhaustedPattern off the live config supplier
  per lookup (matching CompositePeerLauncher#models0's live-supplier pattern),
  caching compiled Pattern objects by PROFILE NAME (not pattern text — pattern
  text would grow unboundedly as an operator tunes a regex across reloads;
  profile names are bounded by the small, restart-gated set of configured
  profiles). patternFor() backs CompletionResolver's classification; armed()
  backs fleet_profiles' exhaustionDetectionArmed — both read the same object,
  the fleetd #404 single-accessor rule CompositePeerLauncher.modelGateState()
  established for the model gate.
- Fleetd.java: on BACKEND_EXHAUSTED, log a WARNING naming the profile's model
  and the exact fix (enabled: false under models.allow, hot, no restart; remove
  it again once the window resets) — or, when the profile has no model:
  configured, say quarantine is the only thing keeping spawns off it.
- FleetMcp.QuarantineSource gains modelFor/reasonFor; fleet_profiles'
  quarantined rows gain model/reason fields so a lead can see why without
  reading the daemon log. capacityView (fleet_list) intentionally untouched —
  scoped to fleet_profiles only.
- ConfigRef/FleetConfig docs + fleetd.example.yaml updated: exhaustedPattern
  moves from Deferred to Hot. errorPattern stays deferred on purpose (out of
  scope for this ticket).
- Tests: LiveExhaustedPatternsTest (new, unit-level hotness/caching proof),
  FleetdExhaustionDetectionArmedWiringTest (rewritten — same QuarantineSource
  object, read before and after a reload, asserts the answer flips with no
  restart), ConfigRefTest/ConfigRefProfileCoverageTest updated for the new
  Hot/Deferred classification.
2026-09-10 18:02:00 +07:00
Dai Ha 2af13ab1ff fleetd #449: name the decisive cell and the load the claim was measured at
CI / contract (push) Successful in 54s
CI / build (push) Successful in 2m8s
The previous comment named a cause with no cell behind it. fleet01's review
made the point: a confident wrong mechanism gets copied, and a confident
under-determined one gets copied the same way.

So the comment now names the cell that settles it. Hold the old 1000ms write
sleep and change only the read - 800ms fixed sleep becomes a 5s poll - and the
test goes 0 of 3 to 3 of 3. The read deadline was the whole story.

It also says where: a 12-core macOS host near idle (load 2.6 to 5.9). The
loaded run agreed, but its load climbed from 7 to 50 while the cells ran and
the old version ran last, so it is not clean evidence and the comment says so.

Comment only. No test or production code changed.
2026-09-10 17:59:52 +07:00
Dai Ha 0788d84be8 fleetd #453: document the override obligation on PeerLauncher.defaultProfileFor/place
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m34s
Decision: leave both as default methods (option 1), not abstract. Neither
default is a live defect today — HerdrPeerLauncher is the sole single-profile
implementer and the degenerate answer (ignore role, always defaultProfile())
is correct for it. Making them abstract would force ~10 boilerplate one-line
overrides across 5 unrelated PeerLauncher test doubles (NeverSpawnsLauncher,
RaceLauncher, NoResumeLauncher, ClearContextSpyLauncher, LazyIdLauncher) that
never call either method, for a risk that is speculative (no multi-profile
HerdrPeerLauncher subclass exists or is planned).

Strengthens both javadocs with an explicit MUST-override warning and cross-
references HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)'s existing
#450 javadoc, which already names both methods as "unoverridden here" and
ties that to being a single-profile adapter -- the concrete place a future
multi-profile launcher author would read, since #450 made that method
abstract and any subclass must write its body.
2026-09-10 17:49:37 +07:00
Dai Ha 20c1094cbf fleetd #449: say what the timing fix actually proved, not what it assumed
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m33s
The polling fix that landed in #452 is right, but its javadoc named a
mechanism nobody measured: that input typed before the shell's prompt was
swallowed by the shell's own startup.

I mutated the settle poll away — SHELL_READY_TIMEOUT_MS = 0, so input is
typed at once with no wait — and the test passed 3 of 3. So waitForText is
the load-bearing half, and the proven cause is the old 800ms READ deadline,
not the 1000ms write delay.

The direction is the point: typing at 0ms works where typing at 1000ms
failed. If early input were swallowed, 0ms would be worse than 1000ms. It is
better, so the swallow explanation is unsupported.

waitUntilSettled stays as cheap insurance, now labelled as insurance rather
than as the fix. Comment-only; AgentControlContractTest still green.
2026-09-10 17:36:22 +07:00
ltms 9011c59b9f Merge #452: run the contract tag in CI, fix the stale protocol 14 assertion (fleetd #449)
CI / build (push) Successful in 1m28s
CI / contract (push) Successful in 1m33s
Verified on a locally built merge onto c11ad71 (the PR was branched from 822327e, before #448):
1578 unit tests green, then `-Dgroups=contract` gives 30 tests / 0 failures / 0 skipped here.

CI's own run of the tag on d4f93a7: 29 tests, 0 failures, 6 skipped — all 6 herdr tests skip on
the runner, and LeadMailboxTest's 15 tests run there for the first time.

Two mutations killed: restoring `assertEquals(14` fails the test, and pointing the expected env
value at a wrong URL fails it with the real pane text in the message.
2026-09-10 12:35:10 +02:00
ltms c11ad71ed0 Merge #448: fleet_ack errors instead of claiming success on a miss (fleetd #437)
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m33s
Verified by the lead on head 5289eb5, and then again on the MERGE.

Battery on the branch tip:

  FULL BUILD  Tests run: 1577, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0
  CONTROL     contract test, real broker: Tests run: 9, Failures: 0, Skipped: 0

  M3a  AmqpReplyInbox.ack's `h == null` branch reports true
       -> KILLED  AmqpReplyInboxContractTest.ackReportsHitVsMissAgainstARealBroker
  M3b  AmqpReplyInbox.ack's `perTarget == null` branch reports true
       -> KILLED  same method

M3a is the gap I found in round 1: reporting true for something never held
survived both the default suite and -Pcontract against a real broker. It is now
closed in the adapter this daemon actually runs.

Both branches die, and they die through DIFFERENT assertions, which the worker
worked out and I confirmed by reading the code. held.get(target) is populated by
the deliver callback, not by own(), so after own() with nothing ever delivered
the map entry is still null — the "never held" assertion therefore exercises
`perTarget == null`, and only the double-ack assertion reaches `h == null`. Both
assertions are load-bearing; neither is redundant.

This PR was pushed before #451 landed, so the battery above tested the branch,
not the result. Merging main (bdcf285) in gave 0 conflicts and #448 adds no
PeerLauncher implementer, but a clean auto-merge is not a compiling merge, so I
built the merged tree 0829542:

  FULL BUILD               Tests run: 1578, Failures: 0  BUILD SUCCESS, 0 compile errors
  contract, real broker    Tests run: 9,    Failures: 0
  CompositePeerLauncherTest (#447's guarantee)  Tests run: 76, Failures: 0

The worker also declined to point the contract test at the shared local LavinMQ
broker, because this adapter never deletes queues and a run would leave orphaned
durable queues on the instance backing the live fleet. It started a disposable
rabbitmq:3.13-management container on a throwaway port instead, then removed it.
That was its own judgment and it was correct.

Held peer mail is deliberately NOT ackable: fleet_ack against a coord-id errors
and names fleet_poll{coordId}. See #437 for why refusing is the right answer
today, and why that is a policy choice rather than a structural one.
2026-09-10 12:20:49 +02:00
ltms bdcf285265 Merge #451: make PeerLauncher.spawn(req, decision) abstract (fleetd #450)
CI / contract (push) Successful in 1m17s
CI / build (push) Successful in 1m33s
Verified by the lead on head cfebc57 (base 822327e is current main).

  FULL BUILD  Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0

  M1  delete HerdrPeerLauncher's new override
      -> BUILD FAILURE, 1 COMPILATION ERROR block, naming exactly
         ClaudeCodeLauncher.java:[49,14] and OpenCodeLauncher.java:[59,14]

  M2  CONTROL for M1: put the interface method back to a `default` AND delete
      the override
      -> BUILD SUCCESS, 0 compile errors
      This is the row that makes M1 mean something. Without it, M1 only shows
      that the build broke; with it, the break is attributable to the method
      being abstract rather than to anything else the edit disturbed.

  M3  CompositePeerLauncher's override reverts to the re-entering form
      -> KILLED, Errors: 1
      CompositePeerLauncherTest
        .spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace
      So #447's guarantee survives this refactor of the interface it rests on.

Tree restored clean after each mutation (git status --porcelain empty).

The worker corrected my ticket, and it was right. My #450 body listed five
src/main implementers of PeerLauncher and quoted `grep -rln 'implements
PeerLauncher'` as the source; that command returns two files. I had run a wider
pattern that also matched a comment in ConfigRef and the `extends
HerdrPeerLauncher` line in two subclasses, then quoted the narrow command beside
the wide command's output. Ground truth: ConfigRef implements
Supplier<FleetConfig>; ClaudeCodeLauncher and OpenCodeLauncher extend
HerdrPeerLauncher. So one override in that parent serves both, which is what the
worker built. The ticket body is corrected.

Out-of-scope note carried forward from the worker: defaultProfileFor(MemberRole)
and place(MemberRole) are two more default methods with the same shape. Noted,
not fixed here.
2026-09-10 12:16:47 +02:00
Dai Ha d4f93a7b13 fleetd #449: fix stale herdr protocol 14 javadocs/assertion, diagnose and fix the timing-raced AgentControlContractTest, select contract tests by tag in CI
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m9s
- HerdrClient.java, HerdrCodec.java, HerdrContractTest.java: the herdr port to
  protocol 19 (CB-521) left the client javadoc and the contract test's own
  assertion still saying protocol 14 / herdr 0.7.0. Updated to 19 / 0.8.0 and
  renamed pingReturnsProtocol14 -> pingReturnsProtocol19. Verified the
  assertion is real by temporarily changing the expected value to 20 (fails),
  then restoring 19 (passes).

- AgentControlContractTest.java: tabCreateInjectsEnvIntoTheSeedShell was
  failing, not skipping, on a host with a live herdr socket. Diagnosed with a
  temporary instrumented run (not committed) that polled the pane every
  200ms before and after sending input: the seed shell reliably takes ~2.5s
  to reach its prompt (measured 3x), while the test's fixed 1000ms sleep
  raced that startup. Input typed too early was swallowed by the shell's own
  startup, leaving the typed line followed by the "Restored session" banner
  and no command output — indistinguishable at a glance from the env map
  never reaching the shell. Once the shell was actually ready, the injected
  env value showed up in ~200ms, ruling out an env-seam defect. Replaced both
  fixed sleeps with bounded polling on the actual conditions (pane text
  settling, then the expected output appearing). Ran the fixed test 3x
  standalone, all green.

- .gitea/workflows/ci.yml: the "Contract tests" step ran exactly one class by
  name (-Dtest=AmqpReplyInboxContractTest), silently excluding every other
  @Tag("contract") test from CI including the herdr ones above -- which is
  how the stale protocol 14 assertion went unnoticed. Changed to
  -Dgroups=contract, which selects the whole tagged group and picks up
  future contract tests automatically.
2026-09-10 17:14:09 +07:00
Dai Ha cfebc575ea fleetd #450: make PeerLauncher.spawn(SpawnRequest, PlacementDecision) abstract
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m5s
The default re-entered the single-argument spawn(SpawnRequest), which re-runs
checks that can refuse the profile place() just chose (#444's window). Only
CompositePeerLauncher overrode it; a future placement-doing launcher could
have inherited the wrong body silently.

Give every current implementer an explicit override, chosen by what it does:
- HerdrPeerLauncher (base of ClaudeCodeLauncher/OpenCodeLauncher, neither of
  which overrides spawn(req) or place()) does no placement filtering of its
  own, so it gets the re-entering form.
- CompositePeerLauncher's routing-form override is untouched.
- 5 test-fake PeerLauncher implementers (SessionManagerTest, FleetdBackendErrorSinkTest)
  get overrides matching their existing spawn(SpawnRequest) shape: delegating
  wrappers delegate, unreachable stubs throw, the single-profile fake re-enters.

ConfigRef does not implement PeerLauncher at all (confirmed in this tree at
822327e) despite the ticket listing it as a src/main implementer.
2026-09-10 17:10:44 +07:00
ltms 822327eed5 Merge #447: pin the place()-to-spawn() window PlacementDecision closes (fleetd #444)
CI / contract (push) Successful in 1m33s
CI / build (push) Successful in 1m36s
Verified by the lead on the exact tree that lands (head e4c703a, base 82fae94 is
an ancestor, so this is the tree I measured):

  FULL BUILD  Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0
  CONTROL     CompositePeerLauncherTest  Tests run: 76, Failures: 0  -> GREEN

  M1  the 2-arg spawn re-enters the 1-arg spawn (the inherited default this
      ticket forbids for a multi-profile launcher)
      -> KILLED  Errors: 1
      CompositePeerLauncherTest
        .spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace

  M2  drop the stamping: route to the decided profile but do not carry it
      (SpawnRequest routedReq = req)
      -> KILLED  Failures: 1
      same test method

Both mutations proved applied by printing the mutated method, and the tree was
restored clean after each (git status --porcelain empty).

Round 1 of this PR had a fixture weakness I found by mutation: StubLauncher's own
fallback default was "sol", the same profile place() decides, so an unstamped
request landed on spawnCount("sol") by coincidence and M2 survived. e4c703a gives
the adapter "b" as its fallback instead. One fixture now kills both mutations.

src/main is javadoc-only in this PR: 12 added lines, 0 added code lines, measured
by filtering the main-side diff.
2026-09-10 12:00:59 +02:00
Dai Ha 5289eb509f fleetd #437: pin the ack hit/miss contract in the AMQP contract test
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Successful in 1m34s
AmqpReplyInboxContractTest is the one contract-group class CI actually
runs, and it never asserted on ack()'s return value at all — so the
exact defect this ticket fixes (reporting success for an ack that
removed nothing) was unpinned in the adapter fleetd runs live.

Add ackReportsHitVsMissAgainstARealBroker: a msgId never held for an
owned target returns false without throwing, a real held reply returns
true and is removed, and acking the same msgId again returns false.
Ran against both broker modes the class supports: Testcontainers
(AMQP_URI unset) and an external broker via AMQP_URI (the CI shape,
using a disposable container — not the shared local LavinMQ instance).
2026-09-10 16:58:55 +07:00
Dai Ha e4c703a51a fleetd #444: separate the adapter's fallback default from the decided profile
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m56s
Review found the fixture's StubLauncher fell back to 'sol' too — the
same profile place() decides — so an UNSTAMPED request could land on
spawnCount('sol') by coincidence, and the assertion's claim that the
request 'actually carried sol' was unproven. Dropping the stamping
(SpawnRequest routedReq = req) while keeping the routing survived the
test unchanged.

Fix: give the adapter 'b' as its own fallback default instead, so an
unstamped request counts against 'b', not 'sol'. Verified both
mutations against the single test in isolation:
  - drop-stamping (routedReq = req): RED, expected <sol> but was <b>
  - re-entering (return spawn(req.withProfile(decision.profile())))
    i.e. M1 from the first round: still RED, PlacementException
    naming the now-quarantined 'sol'
Restored both; full suite green at 1575 tests.

No changes to src/main — PeerLauncher's javadoc from the first round
is unchanged.
2026-09-10 16:50:39 +07:00
Dai Ha 703a05db41 fleetd #437: fleet_ack errors instead of claiming success on a miss
CI / contract (pull_request) Successful in 1m25s
CI / build (pull_request) Successful in 1m37s
ReplyInbox.ack now returns boolean (true = removed, false = nothing to
remove) instead of void, so FleetMcp.ack can finally tell a hit from a
miss. FleetMcp.ack returns an error when the boolean is false, naming
fleet_poll{coordId} for held peer mail, which has no route through this
call. MessageService.ackReply propagates the boolean; drainReplies keeps
ignoring it (its own javadoc already documents that loss window as
deliberate). Updated the tool schema's target description to match.

Rewrote FleetMcpTest's ack tests to publish a real message before
asserting success, and added tests for a never-queued id and a coord-id
target, both now erroring. Added boolean assertions to
InMemoryReplyInboxTest's existing ack cases.
2026-09-10 16:44:08 +07:00
Dai Ha 3f036b2a62 fleetd #444: pin the place()-to-spawn() window PlacementDecision closes
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m53s
Add a test that resolves place(role) while nothing is quarantined,
then quarantines the resolved profile's credential BEFORE spawning
against the held PlacementDecision. CompositePeerLauncher.spawn(req,
decision) must still honor the decision and land on the quarantined
profile, since it never re-runs the explicit-profile enforce* checks.

Verified the test kills the regression: with the override's body
replaced by the re-entering spawn(req.withProfile(...)) form, this
exact test goes RED with a PlacementException naming the now-
quarantined profile; restored, the full suite is green (1575 tests).

Also documents on PeerLauncher's default spawn(req, decision) that a
launcher routing across more than one profile MUST override it,
naming the four enforce* checks the default's re-entry re-applies.
2026-09-10 16:41:52 +07:00
ltms 82fae94c55 Merge #445: pin every startup report call in Fleetd.main (fleetd #442)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m59s
Test written by a worker whose backend died before it could report; evidence
re-run by the lead against the merged tree.

Verified: merge of current main clean (0 conflicts); full build 1574 tests,
0 failures, BUILD SUCCESS, 0 compile errors; control green; deleting each of
reportGitHostShape, reportMemberTrustModel, reportMemberCredentialsGap and
reportExhaustedPatternGap from main() is KILLED by
mainReportsEveryStartupGapBeforeValidationAborts.
2026-09-10 11:30:31 +02:00
Dai Ha b1f34c2e6b fleetd #442: drop the unused java.util.List import
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m8s
The new test never names List — only ListAppender, which has its own
import. An unused import is an IDE warning, and this repo treats
warnings as gates. No behaviour change: FleetdStartupReportTest still
runs 1 test, 0 failures, BUILD SUCCESS, 0 compile errors.
2026-09-10 16:29:56 +07:00
ltms 3f807d9f1b Merge #443: derive coordinator.heldDurable from queue durability + ack mode (fleetd #440)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m35s
Found by the fleet01 lead reviewing #438 after I had merged it. Verified
independently before merging.

The implementation choice is the load-bearing part: LeadMailbox.own() now
assigns queueDeclare's durable flag and basicConsume's autoAck flag to named
locals, passes those SAME locals into the two real AMQP calls (:203/:204), and
derives heldDurable from them (:207). So the reported fact cannot drift from a
duplicate constant - a mutation to either call's argument moves the behaviour
and the report together. LeadChannel.heldDurable() is abstract, so a future
implementer gets a compile error rather than a silent default.

My own battery, merged tree, control green, tree restored clean:
- M1 revert to the literal true -> KILLED by
  FleetMcpTest.listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable
- M3 heldCount forced to 0 (a half the worker did not touch) -> KILLED by
  FleetMcpTest.listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero
- M2 break the derivation itself -> SURVIVED under plain clean install, exactly
  as the worker reported. LeadMailbox needs a real broker, so the only test that
  reaches the real queueDeclare/basicConsume is @Tag("contract"), excluded from
  the default build. The worker ran that arm with -Pcontract and got 3 reds
  including its own new test. Pre-existing structural limit of this class, not
  introduced here, and the worker flagged it rather than hiding it.

Full build, my own run: Tests run: 1573, Failures: 0, Errors: 0, Skipped: 0 -
BUILD SUCCESS, 0 compile errors.

Not blocking, noted for a possible follow-up: one boolean over two independent
facts cannot say WHICH fact was lost. The merged field is still strictly better
than the literal it replaces, because it can now go false at all.
2026-09-10 11:21:41 +02:00
Dai Ha e70263062c fleetd #442: pin startup report calls
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m36s
2026-09-10 14:41:10 +07:00
ltms 1a1e586b62 Merge #433: carry the PlacementDecision instead of re-resolving the profile (fleetd #425)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m47s
Round 4 pins the fix. Verified independently before merging.

My own battery in the worker's tree (control green, tree restored clean):
- M1 revert SessionManager:677 to launcher.spawn(spawnReq) -> KILLED by
  SessionManagerTest.acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedThe
  WorktreeForUnderARotatingPolicy. This is the ticket's own deliverable and it
  survived 186 tests in round 3.
- M3 drop the withProfile stamping in the 2-arg spawn -> KILLED by 3 tests.
- M2 make the 2-arg spawn re-enter the refusing branch -> SURVIVED, but it is a
  near-equivalent mutant, not a gap in this work. Post-#435 all four enforce*
  conditions are ones the routing branch already filtered on, so the two paths
  differ only if placement state moves between place() and spawn(). Filed
  separately.

Full build, my own run this turn, whole log redirected and grepped:
Tests run: 1572, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, 0 compile
errors. Branch already contains current main.

Read the src/main diff. The new test uses roundRobin() (stateful) and asserts
agreement between the overlay profile and the spawned profile, rather than a
hardcoded name, which is the right shape - select() is stateful, so two calls
disagree by design.
2026-09-10 09:35:02 +02:00
Dai Ha c16d118f09 fleetd #440: derive coordinator.heldDurable from queue durability + ack mode
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m36s
FleetMcp.coordinatorView wrote heldDurable as a literal true, so a change
that broke either the durable queue declare or the manual-ack consume in
LeadMailbox would leave the field, and the full suite, green.

- LeadChannel gets a new heldDurable() method: the conclusion of a durable
  queue declare AND a manual-ack consumer, derived by the implementation
  from what it actually did, never asserted.
- LeadMailbox.own() captures the exact booleans it passes to
  queueDeclare/basicConsume and stores their conjunction; heldDurable()
  returns it.
- FleetMcp.coordinatorView now reads channel.heldDurable() instead of a
  literal; updated the javadoc to say where the fact comes from.
- FakeLeadChannel gets a heldDurable field (default true) + withHeldDurable
  setter so FleetMcpTest can prove the field goes false.
- FleetMcpTest: new test asserts heldDurable:false when the channel says so.
- LeadMailboxTest (contract, real broker): new test asserts heldDurable()
  true against a real LeadMailbox. Verified by hand that flipping own()'s
  autoAck local to true turns this test (and two pre-existing redelivery
  tests) red, and restoring it turns them green again.
2026-09-10 14:31:44 +07:00
Dai Ha 4b10d02207 fleetd #425 rework round 4: mutation-pinning test for the dropped PlacementDecision
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m3s
SessionManager.acquireWithWorktree's unqualified branch must carry the
PlacementDecision it already resolved via launcher.place() into
launcher.spawn(spawnReq, decision) rather than re-deriving it through a
blank-profile launcher.spawn(spawnReq). Every existing test in this file uses
PlacementPolicies.fixed(), which answers select() the same way on every call,
so dropping the decision (handle = launcher.spawn(spawnReq);) was invisible:
186 tests stayed green under that mutation.

acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy
uses PlacementPolicies.roundRobin() instead — deterministic AND stateful, so
two select() calls on the same policy instance disagree (index 0 then index 1
across a two-profile pool). It asserts AGREEMENT between the profile the
worktree's parity overlay was provisioned for and the profile the member
actually spawned on, never a hardcoded expected profile name.

Verified as a real mutation, not a no-op: applying the exact mutation
(handle = launcher.spawn(spawnReq);) turns it red — expected [b.mcp.json]
but was [a.mcp.json] — and reverting turns it green again. Full build:
1572 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-09-10 14:24:14 +07:00
Dai Ha c5fbfdbf4a Merge main into #425 rework branch (brings #438 held-peer-mail read) 2026-09-10 14:11:37 +07:00
ltms 12cff28abb Merge #438: let a lead read its own held peer mail, primary-only (fleetd #421)
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 2m6s
fleet_poll{coordId} peeks this daemon's own held lead-to-lead mail and
returns full bodies without acking. New Authz.Action COORD_READ, primary
only — not the architect, which holds READ today. pollAction is now
argument-derived over both target and coordId.

heldView and HELD_PREVIEW_MAX_CHARS are untouched: fleet_list stays a
cheap always-safe scan, and the full read is a separately authorized call.

Verified by me, not taken from the report: merged tree builds 1562 green
(main 1555 + 7 new methods), 0 compile errors, no merge conflicts.
Mutation battery on lines the worker did NOT mutate, control 111 green:
  - COORD_READ widened to the architect  KILLED (AuthzTest + FleetMcpAuthzTest)
  - self-coord-id guard removed          KILLED (FleetMcpTest)
  - full body swapped for the preview    KILLED (FleetMcpTest)

The third mutation targets the ticket's own deliverable, and it is pinned.
Authz.permits has no default, so a new action is a compile error rather
than a silently unhandled case.

Follow-up filed as #439: fleet_list's coordinator row is still READ-gated,
so a worker sees peer coord-ids and 80-char previews of lead-to-lead
bodies. Pre-existing; the implementer flagged it and left it alone.
2026-09-10 09:08:17 +02:00
Dai Ha 77a6a7142e Merge main into #421 branch (brings #434 model-gate observability and #436 fixed-placement cap)
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m48s
2026-09-10 14:06:34 +07:00
Dai Ha 9f3671b801 fleetd #425 rework round 3: rewrite prose after #435 made fixed honour maxLoad
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m56s
fleetd #435 (merged to main) made FixedPlacementPolicy evaluate maxLoad during automatic
selection, the same way weighted/round-robin already did. Six comments across
CompositePeerLauncher.java, PeerLauncher.java, PlacementDecision.java, and SessionManager.java
justified round 2's place()/spawn(req, decision) mechanism by saying fixed "deliberately never
evaluates maxLoad" — that claim is now false, and needed restating, not just deleting.

The honest case after #435: the two-path shape (routing branch falls through an excluded
candidate; explicit-profile branch refuses on it) is still real and still deliberate — an
operator who names a profile should get a refusal, not a silent substitution. What round 1 got
wrong, and what round 2 still needs to prevent, is turning a fall-through into a refusal by
accident: resolving a name via place() and then feeding it back to spawn(SpawnRequest) as an
explicit profile. Before #435 that accident was reachable through maxLoad specifically, because
fixed never evaluated it; #435 closed that specific gap, so a PlacementDecision can no longer be
at-cap in the first place. What survives as the justification for spawn(req, decision): it never
re-evaluates a condition place() already decided, and it closes the window between that decision
and the spawn in which the underlying state could otherwise move — not a failure #435 already
prevents.

Re-measured the sibling paths this round exists to keep in agreement (one profile at maxLoad: 1,
liveCount pinned at 1, PlacementPolicies.fixed(), unqualified spawn): both the with-worktree and
without-worktree paths now throw the identical PlacementException — "worker profile 'a' is at
maxLoad (1 live >= 1 cap), and no available candidate remains" — closed upstream by #435, at
place()/select(), before either path ever reaches a spawn call. The observable asymmetry this PR
was filed to fix is gone; what remains is the structural argument above.

No behavior change: place()/PlacementDecision/spawn(req, decision) are untouched, and
FixedPlacementPolicy/PlacementPolicyUtil are taken wholesale from main's merge.
2026-09-10 14:06:10 +07:00
Dai Ha 84034b34d1 Merge remote-tracking branch 'origin/main' into worker/425-rework-placement-resolve-c58ba1-9 2026-09-10 13:56:54 +07:00
Dai Ha 1e9b2c9b7e fleetd #421: let a lead peek its own held peer mail, primary-only
CI / contract (pull_request) Successful in 1m27s
CI / build (pull_request) Successful in 1m36s
fleet_list truncated held lead-to-lead messages to an 80-char preview with
no way to read the full body, and fleet_poll{target} drained the wrong
inbox (a worker's reply queue, not the coordinator mailbox) -- it silently
returned []. fleet_ack would have destroyed the message unread.

Add a non-destructive read: fleet_poll{coordId} peeks (never acks) this
daemon's own held mail via LeadChannel.peek(). The coordId must equal the
caller's own selfCoordId -- passing a peer's id is refused with a reason,
instead of repeating the original silent-[] confusion.

This is authorization-sensitive: mapping it to the existing READ action
would let any worker read every peer lead's mail in full. READ's openness
rests on "the roster carries no secrets" (Authz.java), which does not hold
for lead-to-lead coordination bodies. Added Authz.Action.COORD_READ,
primary-only (not even the architect, which holds READ today), and made
pollAction's signature depend on both target and coordId so every call
site states explicitly what it passes.

Also fixes fleet_list's "pending: 0" trap: mailbox.pending only counts
broker-ready messages, so a healthy held mailbox reads as empty. Added
heldCount/heldDurable beside held[] so the durability fact isn't implied
only by reading the code.

Mutation-tested: pollAction's COORD_READ->READ mapping, the peek->ack
substitution, the 80-char preview cap widened to 81, and the Authz case
widened to include caller.isWorker() -- each breaks exactly its matching
test and nothing else. The first attempt at the preview-cap test used a
homogeneous "x"*200 body, which a widened cap slipped through unnoticed
(contains() found a shifted match); replaced with a sentinel character at
index 80 to actually pin the boundary.

Updates CLAUDE.md's intent->tool table for fleet_poll's new coordId
semantics, per this repo's own "prompt is part of the product" rule.
wiki/ is a submodule and not committable from a worker's worktree --
wiki-bound content is in the PR body instead.
2026-09-10 13:54:50 +07:00
ltms 5d422f85fa Merge #436: make fixed placement honour maxLoad (fleetd #435)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m37s
fixed was the one automatic policy that ignored maxLoad, and it is the
default for an absent placement: key. It now gates on the same shared
atCap predicate weighted and round-robin use, and falls through to the
next candidate rather than refusing.

Verified here: 1555 tests green (main was 1548, +7 new methods), 0
compile errors. Mutation battery, control 113 green:
  - atCap boundary >= -> >        KILLED (15 tests, across all 3 policies)
  - drop cap check in fixed walk  KILLED (3 tests)
  - weightExcluded -> false       KILLED (2 pre-existing tests)

Neither live host is affected today: both run placement: weighted (Mac
fleetd.yaml:210, fleet01 fleetd.yaml:172), and weighted already skipped
at-cap candidates via PlacementPolicyUtil.available. This aligns fixed
with the other policies and with maxLoad's documented contract.
2026-09-10 08:54:18 +02:00
Dai Ha ed2027b202 fleetd #435: make FixedPlacementPolicy honor maxLoad
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m54s
FixedPlacementPolicy (the default placement policy) never consulted
maxLoad, so an at-cap default was chosen anyway on every unqualified
spawn -- the cap was advisory, not enforced, for the one policy every
config uses by default. weighted/round-robin already gated on it via
PlacementPolicyUtil.available().

Extract the "at cap" predicate into PlacementPolicyUtil.atCap(ctx, c)
so all three policies share one definition, and consult it at both of
FixedPlacementPolicy's filter sites (the default fast path and the
candidate walk), mirroring the existing weightExcluded pattern. An
at-cap default now falls through to the next candidate instead of
refusing the spawn -- only when every candidate is unusable does the
policy still throw, naming the cap in the message. Update the class
javadoc (five exceptions -> six) and the reason-priority comments to
match CompositePeerLauncher's explicit-spawn order (quarantine,
cooling off, max load, model-off).
2026-09-10 13:46:06 +07:00
Dai Ha 6b0a99b2b7 fleetd #425 rework round 2: stop routedProfileFor's caller re-entering the throwing branch
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 1m34s
Round 1 closed quarantine/cool-off/model-off routing for acquireWithWorktree by resolving the
profile through routedProfileFor(role) and handing that name back to launcher.spawn(SpawnRequest)
as an EXPLICIT profile. That re-resolution has a cost the lead measured directly: naming a
profile explicitly makes CompositePeerLauncher.spawn take its THROWING branch (enforceMaxLoad
included), while the routing branch a blank spawn takes never calls enforceMaxLoad at all, and
FixedPlacementPolicy (the default) deliberately never evaluates maxLoad during automatic
selection. So an at-cap pool-first profile that placement itself would have picked for a plain
unqualified spawn could die at enforceMaxLoad one call later, purely because the worktree path's
route to the spawn passed through an explicit profile name — a new failure a worktree-less
unqualified spawn never hits.

This closes the two-path shape instead of moving it: PeerLauncher gains place(role), returning an
opaque PlacementDecision, and spawn(req, decision), which honors that decision through the SAME
routing branch a blank spawn uses — no enforce* check is newly applied. SessionManager.
acquireWithWorktree now keeps the PlacementDecision from place() and hands it to
spawn(req, decision) for an unqualified request, instead of re-resolving through an explicit
profile name. An explicitly-named profile is unaffected: it still goes through spawn(req) and its
throwing branch, exactly as before.

Also corrects the acquireWithWorktree comment's false claim that round 1 "loses nothing else" —
maxLoad was lost too, as a new hard failure, not a retry. The comment now names it explicitly.

Kept the four round-1 tests (still pass — routedProfileFor now just delegates to place()). Added
one class asserting the invariant itself: an unqualified spawn on a maxLoad-capped profile must
land the same outcome with and without a worktree, asserting on the pair rather than a hardcoded
direction, so it stays correct however fleetd #435 (not this ticket) resolves whether maxLoad
should gate an unqualified spawn at all.
2026-09-10 13:42:06 +07:00
ltms 6b7caba248 Merge #434: make the model gate's own state observable
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m49s
Verified in the worker's tree at 7fd914d: 1544 tests green (base 1535 + 9),
0 compile errors, unpiped mvn clean install. merge-tree against f8b0d42 reports
no conflicts and the two sides share no files.

Read all four production diffs. The design is right: one ModelGateState record
carrying both "armed" and "off" from a single models0() read, so the startup
log line, fleet_profiles' modelGateArmed and the spawn gate cannot disagree —
the fleetd #404 lesson applied properly. disabledModels() now delegates to it
rather than being a second independent read.

Mutated the two subtlest lines, with a control in the same script and the
changed line echoed back with its number:

- modelGateState(): sentinel identity check replaced by the naive
  "m.offIds().isEmpty()" -> 2 failures. FleetProfilesModelGateStateTest
  .modelsBlockWithNothingOffReportsGateArmedAndZeroOff:86 and
  CompositePeerLauncherTest.modelGateStateIsHotReloadedThroughARealConfigRef
  :1519, both "expected: <true> but was: <false>". So the one state this
  ticket exists to expose — a block present with nothing off — is pinned.
- notConfigured() returning a non-empty off set, breaking the invariant the
  record's javadoc states but does not enforce -> 3 failures, including
  disabledModelsIsEmptyWithNoModelsConfigured:1581. So the invariant is
  observable even though the constructor does not check it.

Control run unmutated: 1544 green.

Follow-up on me, not a merge blocker: fleet_profiles gains an operator-visible
field, so this needs a wiki/11-Features.md entry. Workers cannot commit the
wiki submodule, so I am adding it.
2026-09-10 08:32:54 +02:00
Dai Ha f8b0d42a5c fleetd #431 follow-up: profileForSlot's javadoc named a caller that does not exist
CI / contract (push) Successful in 1m28s
CI / build (push) Failing after 1m53s
The javadoc said "what the spawn lifecycle reads". Nothing in src/main calls
profileForSlot at all, in either the ".profileForSlot(" or the
"::profileForSlot" form. My own #431 ticket text repeated that sentence as a
fact and ranked the three accessors by it, and the #432 worker copied it into
the test file's comment and one assertion message. Corrected in all three
places; the ticket correction is posted on #431.

What the spawn lifecycle actually reads for an architect's profile is the
SlotReservation that reserve() returns — SessionManager.java:225,
"reservation == null ? profile : reservation.profile()".

The corrected ranking, measured rather than read off the javadoc:
- nameForSlot is wired, at CallerResolver.java:137 (method reference, which is
  why a ".nameForSlot(" grep missed it)
- isSlot is reached through bind, called at MemberRegistry.java:235 and :376
- profileForSlot has no caller at all

Prose only. No behaviour change.
2026-09-10 13:25:19 +07:00
ltms d1e7d71eee Merge #432: pin profileForSlot, isSlot and nameForSlot against a live reload
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
Verified in my own tree at a196d34: 1539 tests green, 0 compile errors, 0 files
changed under src/main (test-only, as reported).

Mutated two halves the worker's own proof did not cover, with a control in the
same script and the changed line echoed back each time:

- roleForSlot returning ARCHITECT for any configured slot (flatten is
  role-blind) -> 2 failures, incl. CallerResolverTest
  .aBoundNonArchitectSlotStillResolvesAsAWorker:454. So the role-blind flatten
  cannot grant ARCHITECT through a non-architect pool; that was already pinned.
- nameForSlot parsing the key suffix instead of reading config -> 1 failure,
  MemberRegistryLiveTest.nameForSlotReflectsANameChangedByReload:255. The new
  test has teeth beyond the freeze the worker ran.

Control run unmutated: green.

The ticket's severity ranking was wrong and I corrected it on #431. profileForSlot
has no caller in src/main in either the "." or "::" form, so its javadoc ("what
the spawn lifecycle reads") names a caller that does not exist; nameForSlot is
wired at CallerResolver.java:137; isSlot is reached through bind, at :235 and
:376. A follow-up commit fixes the test prose that repeated my claim.
2026-09-10 08:24:26 +02:00
Dai Ha 7fd914df1a fleetd #422 follow-up: make the model gate's own state observable
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m53s
PeerLauncher.disabledModels() reported an empty set both when no
models: block exists and when a block exists with nothing off, so
fleet_profiles/GET /profiles and the startup log could not tell an
inert gate from an armed one reporting zero. Add
PeerLauncher.ModelGateState (configured + off), a modelGateState()
default method disabledModels() now delegates to, and a
CompositePeerLauncher override that reads models0() once and
distinguishes the null-supplier case (no models: block) from a real,
config-supplied block via identity against the NO_MODELS_CONFIGURED
sentinel — reusing the exact accessor the spawn gate itself reads, per
the fleetd #404 lesson.

Wires the state into a new startup log line (Fleetd.modelGateCoverageLine)
and a new modelGateArmed field in FleetMcp.profilesView, reported
unconditionally alongside the existing modelsOff set.
2026-09-10 13:18:52 +07:00
Dai Ha b066eb1903 fleetd #425 rework: resolve acquireWithWorktree through real placement, not a blind pool-first read
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m13s
e1d7dde (PR #430) kept two good fixes and one regressed one. Kept: (1)
CompositePeerLauncher.defaultProfile() delegating to
defaultProfileFor(MemberRole.DEV) so fleet_profiles' "default" tracks a live
reload, and (2) PeerLauncher.defaultProfileFor(MemberRole). Redone:
acquireWithWorktree's profile pre-resolution.

The regression: acquireWithWorktree pre-resolved via
launcher.defaultProfileFor(memberRole), which just returns the role pool's
FIRST entry, blind to quarantine/cool-off/model-off. That name was then
passed to launcher.spawn as an EXPLICIT profile, which takes
CompositePeerLauncher.spawn's THROWING branch (enforceNotQuarantined /
enforceMaxLoad / enforceModelEnabled) instead of the ROUTING branch a blank
profile gets. So a quarantined or model-off pool-first profile turned a
routine unqualified spawn into a hard PlacementException -- undermining
fleetd #429's "the fleet keeps working when a model is turned off"
guarantee for every worktree spawn.

Fix: add PeerLauncher.routedProfileFor(MemberRole), the profile an
unqualified spawn of that role would actually be routed to right now --
same candidate list, same quarantined/coolingOff/modelOff filtering, same
PlacementPolicy spawn() itself consults. CompositePeerLauncher implements it
by extracting spawn()'s context-building into a shared private
placementContextFor(role, unreachable), so spawn() and routedProfileFor()
can never disagree about which conditions apply to which candidate.
acquireWithWorktree now calls routedProfileFor once and reuses that name for
repoRoot, parityOverlay, and the spawn -- the fleetd #425 defect (the three
disagreeing) stays fixed, now on the routed path instead of the blind one.

An explicit profile named by the caller is untouched -- it still hits the
throwing branch, which is correct for an operator override.

Trade-off carried over from e1d7dde, now precisely scoped: an unqualified
worktree spawn still loses CompositePeerLauncher's cross-candidate retry on
a live PeerUnreachableException (a transport failure at spawn time, which
placement cannot see in advance) -- but NOT the quarantine/cool-off/
model-off routing, which routedProfileFor already resolved before spawn
ever runs. Accepted: a worktree provisioned for the wrong backend is worse
than a spawn that fails cleanly and can be retried by the caller.

Tests: CompositePeerLauncherTest gains
routedProfileForSkipsAQuarantinedPoolFirstProfileUnderFixedPolicy and
...ModelOff..., both under PlacementPolicies.fixed() (the default policy,
not weighted() -- the previous round's tests all used weighted() and never
exercised FixedPlacementPolicy's own inline filter, which is exactly what
regressed). SessionManagerTest gains
acquireWithWorktreeRoutesAroundAQuarantinedPoolFirstProfile, proving
repoRoot/parityOverlay/spawn agree on the ROUTED profile, not just the
pool-reordered-by-reload one the existing #425 tests already covered.
2026-09-10 13:10:30 +07:00
Dai Ha a196d34455 fleetd #431: pin profileForSlot, isSlot and nameForSlot against a live reload
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m7s
#424 made MemberRegistry.slots() re-read fleet: on every call, but only
roleForSlot was tested against a real reload. profileForSlot, isSlot and
nameForSlot all have the same live-read line and none was pinned — proved by
freezing each to a construction-time snapshot and watching the full suite
stay green.

Adds 4 tests to MemberRegistryLiveTest, each driving a real ConfigRef.reload()
against a @TempDir config file (never two frozen registries compared in
memory, which would test the constructor instead of the reload):
- profileForSlotReflectsAProfileChangedByReload
- isSlotStopsReportingASlotRemovedByReload / isSlotStartsReportingASlotAddedByReload
- nameForSlotReflectsANameChangedByReload

No production change. profileForSlot has no call site anywhere in src/main
yet, so there is no spawn-lifecycle seam to drive the test through beyond the
accessor itself.
2026-09-10 13:06:56 +07:00
Dai Ha 051d320ea0 fleetd #425: fleet_profiles' default and worktree provisioning must read live placement
fleet_profiles' "default" was CompositePeerLauncher.defaultProfile, a value
frozen at construction from cfg.effectiveDefaultProfile(). An unqualified
fleet_spawn instead resolves the dev pool live via defaultProfileFor(DEV) on
every call, so reordering fleet.developers and reloading changed where a
spawn landed without ever changing what fleet_profiles reported.

- CompositePeerLauncher.defaultProfile() now delegates to
  defaultProfileFor(MemberRole.DEV) -- the same live, reload-aware pool read
  placement already uses -- falling back to the frozen field only when no
  profiles are configured at all.
- PeerLauncher gains a default defaultProfileFor(MemberRole) method so a
  generic PeerLauncher reference can ask for a role's live default; the
  default implementation delegates to defaultProfile() for launchers with no
  pool concept of their own.
- SessionManager.acquireWithWorktree resolved a profile via
  launcher.defaultProfile() (DEV-only) to provision repoRoot/parityOverlay,
  then spawned with the original (possibly blank) profile, which re-resolves
  independently through placement -- for any non-DEV role, or across a config
  reload between the two reads, the two resolutions could disagree and
  provision a worktree for a profile the member never runs on. Fixed by
  resolving once, through defaultProfileFor(the caller's actual role), and
  reusing that same resolved name for repoRoot, parityOverlay, and the spawn
  itself. Trade-off: this path now spawns with an explicit profile rather
  than a blank one, so it loses CompositePeerLauncher's cross-candidate retry
  on PeerUnreachableException -- accepted because a worktree provisioned for
  the wrong backend is worse than a spawn that fails cleanly and can be
  retried.

Tests: CompositePeerLauncherTest (live dev-pool reorder + empty-pool
fallback), FleetProfilesLiveDefaultTest (drives FleetMcp.profilesView
directly), SessionManagerTest (worktree overlay follows a reorder, and a
non-DEV role's worktree spawn uses that role's pool, not DEV's).
2026-09-10 12:58:20 +07:00
Dai Ha 766772763f Merge #428: revoke the ARCHITECT privilege on reload, not just future spawns
CI / contract (push) Successful in 1m27s
CI / build (push) Successful in 1m32s
fleetd #424. MemberRegistry.slots() now re-reads fleet: through a supplier,
so removing an architect slot demotes the bound pane on its very next request.
The boundEntries cache and entryFor() fallback from the first round are gone:
the slot OCCUPANCY (terminalToSlot) survives a reload, the ARCHITECT role does
not. That split was my own ticket wording's fault -- I asked for a test that a
bound architect "survives the rebuild", which conflated the binding with the
privilege.

Conflict resolved by hand in ConfigRef.java: #422 (models:) and #424
(architects) both rewrote the same Hot bullet. Kept both.

Also corrected two claims #424's own second commit left stale -- 7f672f0
reversed the behaviour but never touched ConfigRef, whose whole job is to tell
the operator what a reload does:
  - the Hot bullet said MemberRegistry's rule "keeps a live session's identity
    even after its slot is removed from config"
  - the reload-report comment said "only a NEW bind is refused"
Both now say what the code does: removal revokes ARCHITECT on the next
request, and only the slot occupancy survives.

Verified by the lead: 1535 tests, 0 failures, 0 compile errors, BUILD SUCCESS
on the merged tree.

Mutation of three halves the worker's own proof did not cover -- profileForSlot,
nameForSlot and isSlot each pointed at a frozen snapshot taken at construction
(live readers 5 -> 4, each mutation naming its method and line). All three
PASSED at 1535. The ticket's own fix is well pinned; these three sibling live
reads are not. Follow-up filed.
2026-09-10 12:54:35 +07:00
ltms eab8185d7b Merge #429: enforce the models.allow on/off gate at spawn, hot — including under fixed placement
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m33s
fleetd #422 + follow-up. Verified by the lead: 1524 tests green, 0 compile errors.

Mutation proof of three halves the worker's own proof did not cover:
- modelOffProfiles() -> Set.of() (the feed into ctx.modelOff): kills 3, incl. fixedPlacementSkipsAnOffModelProfileToo
- enforceModelEnabled() -> no-op (the explicit-spawn gate): kills 3, incl. the real-ConfigRef hot-reload test
- disabledModels() -> Set.of() (the reporting accessor): kills 1, alone
2026-09-10 07:44:35 +02:00
Dai Ha 7f672f0fb8 fleetd #424: revoke the ARCHITECT privilege on reload, not just future spawns
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 2m3s
Correction to the #424 fix in PR #428: the ticket asked to revoke a removed
architect slot, but the previous change (boundEntries) kept BOTH the binding
and the ARCHITECT privilege alive for an already-bound session after its slot
left config. That left the ticket's actual headline defect half-open.

The corrected rule: config governs both what may be bound next AND what a
bound slot still grants. Removing a slot now demotes its bound session to
worker on the very next request (roleForSlot/nameForSlot read slots() with no
cache, so CallerResolver.resolve falls through to Principal.worker(...)). The
terminalToSlot binding itself is untouched by a reload, on purpose: dropping
it would double-book the slot key and break unbind's compare-safe contract.

- Delete boundEntries and entryFor; profileForSlot/roleForSlot/nameForSlot/
  isSlot all read slots() directly, live, with no cache.
- Rewrite the class doc's binding rule for the corrected semantic.
- Replace the old "survives removal" test with anArchitectAlreadyBoundToASlotIsDemotedByReload,
  asserted through a real CallerResolver.resolve (not the roleForSlot seam),
  plus two tests for what must NOT change: the binding still occupies the
  slot after removal (a second terminal cannot claim it, even once the slot
  returns to config), and unbind still succeeds for the original terminal.

Verified snapshot()/CallerResolver.members() need no change: snapshot() only
ever reported raw terminalToSlot occupancy, and CallerResolver.members() has
no production caller.
2026-09-10 12:29:35 +07:00
Dai Ha e2801b9bbc fleetd #422 follow-up: gate FixedPlacementPolicy on modelOff too
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
FixedPlacementPolicy is the DEFAULT placement policy (PlacementPolicies.fromName
returns it for an absent/blank name) and it built its own inline candidate
filter instead of calling PlacementPolicyUtil.available(). That filter checked
quarantined/coolingOff/unreachable/excluded() but never modelOff(), so an
unqualified fleet_spawn on any fleet without an explicit placement: policy
could still land on a profile whose model the operator turned off.

- Add ctx.modelOff() to both filter sites: the default-profile fast path and
  the fallback walk over ctx.candidates().
- Add a modelOff refusal reason to the default-profile reasons list, worded as
  an operator decision ("turned off in models.allow"), matching
  enforceModelEnabled. Quarantine and cooling off still take priority when a
  profile is also model-off, matching CompositePeerLauncher's explicit-spawn
  check order.
- Update the two stale "excluded from automatic selection" messages to name
  model-off, consistent with PlacementPolicyUtil.emptyException.
- Update the class javadoc: four exceptions -> five, with a new bullet for
  model-off (fleetd #422).

Tests: PlacementPolicyTest gains fixedSkipsModelOffDefault (fast-path),
fixedFallbackWalkSkipsModelOffCandidate (fallback walk),
fixedThrowsWhenDefaultAndEveryCandidateModelOff (all-off refusal wording), and
fixedReportsQuarantineNotModelOffWhenBothApply (priority). CompositePeerLauncherTest
gains fixedPlacementSkipsAnOffModelProfileToo, an integration-level mirror of
the existing placementSkipsAnOffModelProfileAndRoutesToAnotherOne but under
PlacementPolicies.fixed(). The two existing weighted()-based tests are
untouched.
2026-09-10 12:25:31 +07:00
Dai Ha ea02c7b248 fleetd #422: enforce the model allow-list on/off state at spawn, hot
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 1m30s
Ships the two halves left out of the earlier allow-list ticket in one PR,
since apart they are inert: a gate with no flag always allows, and a flag
nothing reads does nothing.

- FleetConfig.Models.ModelEntry gains `enabled` (default on; absent/true =
  on, false = off). Turning a model off never removes it from `allow:` —
  validateModels() checks membership only, so an off model stays valid
  config and a still-configured profile naming it does not refuse reload.
  Models.offIds() is the one live accessor both the gate and the status
  report read.
- CompositePeerLauncher.enforceModelEnabled is a FOURTH, independent
  spawn-refusal reason (operator intent) — never layered onto
  BackendQuarantine/BackendOutagePolicy, which are backend-reported outage.
  Wired into the explicit-profile branch. modelOffProfiles() feeds the same
  off-model exclusion into PlacementContext for unqualified spawns via
  PlacementPolicyUtil (a new modelOff set, counted into its own bucket in
  emptyException so "all off" is named as the cause, not generic).
  Both read models0(), a live Supplier<FleetConfig.Models>, so a reload
  reaches the very next spawn — no restart.
- PeerLauncher.disabledModels() (default empty) lets fleet_profiles/
  GET /profiles report off models by reading the exact same accessor the
  gate reads (the fleetd #404 lesson: a status field must read the source
  the behaviour reads).
- ConfigRef: `models` reclassified from deferred to hot-excluded — nothing
  about it is baked into a startup-built object anymore; membership is
  re-validated in full on every reload via validateAll(), and the on/off
  half is read live everywhere. Tally: 5 cold, 13 deferred, 3 split, 4
  hot-excluded (25 total). ConfigRefTopLevelCoverageTest and
  ConfigRefTopLevelReportingCoverageTest updated with no new exclusion
  added just to force green.

Tests: FleetConfigTest (old-style fixture stays on; an off model is still
valid config; one model disables every profile naming it),
CompositePeerLauncherTest (explicit refusal wording distinct from
quarantine/cool-off; unqualified spawn skips an off candidate and names
model-off when every candidate is off; a real ConfigRef.reload() proves
the hot path; disabledModels() matches the gate).
2026-09-10 12:08:18 +07:00
Dai Ha ce74e164c6 fleetd #424: make architect-slot identity checks read fleet.architects live
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m32s
MemberRegistry used to flatten fleet.architects into an unmodifiable map at
construction, so removing (revoking) an architect slot from config never
took effect: reserve()/requireSlotFor() kept granting spawns against the
frozen snapshot forever, while ConfigRef told the operator "already
applied" for the wrong consumer.

- MemberRegistry gains a live constructor (MemberRegistry.live(Supplier))
  that re-flattens fleet.architects/developers/reviewers on every
  slots()/slotsFor() call, so reserve() and requireSlotFor() (which both
  read through slotsFor) govern the NEXT spawn with no restart. The frozen
  single-arg constructor is kept for tests and fixed/code-built configs.
- Binding rule: config governs what may be bound next; it never
  retroactively unbinds a live session. A slot removed from config while a
  terminal is bound to it keeps that binding. To keep the bound terminal's
  IDENTITY too (CallerResolver.resolve reads roleForSlot/nameForSlot on
  every request), every successful bind now caches the slot's Entry into a
  new boundEntries map; roleForSlot/nameForSlot/profileForSlot/isSlot fall
  back to it when the slot is no longer live, and unbind clears it in the
  same critical section it clears the binding.
- Fleetd.java now wires MemberRegistry.live(() -> config.get().fleet())
  instead of the frozen constructor.
- ConfigRef: corrected the fleet.leaders split-key message and the class
  doc's Hot bullet — architects is now hot for two independent consumers
  (CompositePeerLauncher for placement, MemberRegistry for identity), not
  only the one the message used to name. An architects-only edit still
  reports nothing beyond "config reloaded", which is now honest since the
  key really is fully hot for both consumers.

Added MemberRegistryLiveTest: real ConfigRef.reload() against a @TempDir
file, both directions (slot removed / slot added) for requireSlotFor and
reserve tested separately, plus a bound-architect-survives-removal test
that checks the binding AND the identity (roleForSlot/nameForSlot).

Mutation-tested: reverting requireSlotFor to a frozen snapshot fails
requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload and its mirror;
reverting reserve the same way fails the two reserve tests; dropping the
boundEntries fallback fails the survives-removal test's roleForSlot
assertion. All three restored before commit.

mvn clean install: Tests run: 1510, Failures: 0, Errors: 0, Skipped: 0 —
BUILD SUCCESS.
2026-09-10 12:05:18 +07:00
ltms e60f892efd Merge pull request 'fleetd #415: split coverage() feature-state wording by pattern fallback semantics' (#423) from worker/415-coverage-wording-2cbf9c-5 into main
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m37s
2026-09-10 06:57:33 +02:00
Dai Ha ce05886831 fleetd #415: pin which UnsetMeaning Fleetd pairs with which pattern key
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m54s
Review found a gap: the earlier tests all called CompletionResolver.coverage()
directly, supplying the UnsetMeaning themselves — proving the enum's wording,
never that Fleetd's two call sites pair the right meaning with the right key.
Swapping the two UnsetMeaning arguments at those call sites (recreating #415's
defect with exhaustedPattern and errorPattern exchanged) compiled with 0 errors
and left all 1506 tests green.

Extract the two coverage-line call sites out of main() into package-private
static factories (Fleetd.exhaustedPatternCoverageLine /
errorPatternCoverageLine), the same pattern already used for capacitySource
and worktreeBranchLookup. Add FleetdPatternCoverageLineTest, which calls both
factories directly and asserts the actual wording each produces for the same
empty-coverage input, including that the two differ.

Also recorded the swap-mutation measurement (0 errors, 1506 green) in
UnsetMeaning's javadoc so a future reader does not delete the new test as
redundant with CompletionResolverTest.
2026-09-10 11:54:02 +07:00
Dai Ha be123d0ac7 fleetd #415: split coverage() feature-state wording by pattern-key fallback semantics
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m4s
CompletionResolver.coverage() measured pattern coverage (how many profiles set
a key) but its 'off' wording read as feature state. That is false for
errorPattern: an unset errorPattern still runs the classification against the
built-in BACKEND_ERROR pattern (CompletionResolver.java:84), so the empty case
is not off.

Add CompletionResolver.UnsetMeaning (OFF / BUILT_IN_DEFAULT), a required
parameter every coverage() call must supply — no defaulted overload, so a
future third pattern key cannot compile without stating what unset means for
it. Fleetd.java now passes UnsetMeaning.OFF for exhaustedPattern (no fallback
exists) and UnsetMeaning.BUILT_IN_DEFAULT for errorPattern.

Tests: updated the three existing empty/full/partial cases to pass the new
parameter, corrected the one test that pinned the old (wrong) errorPattern
wording, and added a test that asserts the same empty-coverage input produces
different wording for the two keys.
2026-09-10 11:40:32 +07:00
ltms 2d09c8b027 Merge pull request 'fleetd #418: barrier the throw-path push-loop test on state decide() reads' (#419) from worker/418-588283-3 into main
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m39s
2026-09-10 06:29:23 +02:00
ltms ed54f0224e Merge pull request 'fleetd #416: fleet_list must enumerate the STARTUP profile set' (#420) from worker/416-3ad1da-1 into main
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m34s
2026-09-10 06:27:36 +02:00
Dai Ha 8d5bc3ee89 fleetd #416: fleet_list must enumerate the STARTUP profile set
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m52s
`profiles` is a DEFERRED key: HerdrPeerLauncher takes Map.copyOf(profiles) once at
construction, so a profile added only to the hot-reloaded map can never be spawned.
CapacitySource was built with `() -> config.get().profiles().keySet()` — the live map —
so fleet_list reported a hot-added profile as free while fleet_spawn on that same
profile failed with "unknown worker profile". fleet_list's contract for `free` says it
runs "the same check the spawn gate itself runs"; it did not.

The set now comes from the startup snapshot, the same shape as the coordinator.peers
wiring three lines below, whose comment already stated the rule. maxLoad stays live on
purpose — ConfigRef documents it as hot, like credentialId and weight — so a maxLoad
edit still takes effect without a restart.

Found by the fleet01 lead, who proved it with a live ghost profile rather than an
argument. Implemented by a sonnet member; its reply was lost to an empty scrape, so the
work was salvaged uncommitted from its worktree and both proof steps were run by the
lead instead:

  full build                          Tests run: 1505, Failures: 0, 0 compile errors
  mutation A: live keySet restored    liveOnlyProfileIsNotListed FAILS (alone)
  mutation B: permanently empty set   startupProfileIsListed FAILS (alone)

Mutation B is the point of the second direction: per #404, a test that only ever checks
the absent case cannot tell a correct lookup from one that returns nothing at all.
2026-09-10 11:25:56 +07:00
Dai Ha ab0cc71aa4 fleetd #418: barrier the throw-path push-loop test on state decide() reads
CI / contract (pull_request) Successful in 1m27s
CI / build (pull_request) Successful in 1m50s
anAskThatLeavesByThrowingStillClosesItsQuestion barriered on Phase.ASKING,
which markAsyncQuestion sets in ask()'s FIRST step. The assertion right
after it depends on ask()'s THIRD step (pushLoop.onQuestionOpened), which
is what actually populates ReplyPushLoop's pendingQuestions map. Under
load the asker thread can be descheduled between those two steps, so the
barrier released before decide() had anything to see, and it correctly
returned STOP instead of the expected INJECT.

Add ReplyPushLoop#pendingQuestionTurnIdsForTest, a package-private test
seam (modeled on MessageService#isCompletionStampedForTest) exposing the
private pendingQuestionTurnIdsFor. The test now waits for its own turnId
to appear there before asserting on decide() — not for decide() itself to
return INJECT, which would make the barrier assert nothing.

Checked every other awaitTicketPhaseOn(..., Phase.ASKING) in the file
(two, in the CB-582 nudge tests): both are followed by a real awaitNudge()
that waits for an actual agent.prompt push-loop call before any assertion
depends on push-loop state, so they are not exposed to this race.

No production behaviour changed.
2026-09-10 11:24:41 +07:00
ltms 4cd9046353 Merge pull request 'fleetd #409: deterministic test for the #399 completion-stamp ordering race' (#414) from worker/deterministic-stamp-race-409-3cb7b6-10 into main
CI / contract (push) Successful in 59s
CI / build (push) Successful in 1m36s
2026-09-10 04:48:44 +02:00
Dai Ha 4e98a74047 fleetd #409: deterministic test for the #399 completion-stamp ordering race
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 2m5s
Widens the completed-hook's real race window (normally instructions-wide, needing
~2x-core host load to hit by chance per #399) by injecting a bounded sleep into the
test clock's completion-stamp read. This makes the ordering invariant — a test must
wait for isCompletionStampedForTest, not just DONE, before advancing the clock past
the TTL — fail deterministically on the first run when the barrier is removed, and
pass deterministically with it present. No production code changed.
2026-09-10 09:39:08 +07:00
ltms a502ba53e0 Merge #404: the armed field reads the startup pattern map, both directions pinned
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m0s
The armed lookup now uses the same compiled map detection uses. Second test added during review after a mutation proved the true direction was unpinned.
2026-09-10 04:20:54 +02:00
Dai Ha 446cc11d1d fleetd #404: pin the armed field's TRUE direction too
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m51s
Mutation-tested during review: replacing the armed lambda with
`profile -> false` left the whole suite green at 1475 tests, because the
only existing test passes an EMPTY startup map. That mutation would make
#395's visibility feature silently dead.

With this test the same mutation fails, and it is the only test that
fails, so nothing else covers this direction.
2026-09-10 09:20:43 +07:00
ltms ddd81fe174 Merge #391: fleet_reply refuses a lead, and the role argument is now mandatory
CI / contract (push) Successful in 46s
CI / build (push) Successful in 2m1s
The permissive 3-arg overload is deleted, so a handler that forgets the role no longer compiles.
2026-09-10 04:18:36 +02:00
ltms ae74cc081f Merge #399: wait on a real completion stamp, not on DONE
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m29s
Adds a volatile completedNanos stamp and a bounded test barrier. Mutation-tested: removing the production stamp fails 3 tests.
2026-09-10 04:14:29 +02:00
Dai Ha cf8da1d5fa fleetd #391: require reply caller role
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 1m45s
2026-09-10 09:13:16 +07:00
Dai Ha b8aedeafcb fleetd #404: test armed startup map behavior
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m30s
2026-09-10 09:13:07 +07:00
Dai Ha 3d61af6f6f Merge #400: the scrub receipt now measures the blank, not the attempt
CI / contract (push) Successful in 1m8s
CI / build (push) Successful in 1m34s
The blanking loop classified a name as blanked from eval's exit status.
zsh coerces a bare NAME= assignment on an integer special parameter
(SECONDS, RANDOM, SHLVL, HISTSIZE, COLUMNS, LINES, USERNAME) to a number
instead of failing, so eval returned 0 with the value untouched. Measured
7 false receipts in 10 names. This is a security receipt, so a count that
overstates the scrub is worse than no count.

The loop now runs the eval unconditionally and decides from the observed
value, read back with the (P) indirection flag. One check covers all
three shapes a name can take: a real blank, a fatal error eval merely
contained, and this silent no-op. Exit status plays no part.

Lead review: mutation removing the '!' unblankable report line is CAUGHT
(2 failures in EnvAllowListScrubTest, both asserting the name is reported
rather than silently dropped). The eval-site identifier guard is
untouched.

Unplanted evidence the fix works: the #394 test
unblankableNameInTheMiddleDoesNotAbortNamesAfterIt began failing under
the fix, because zsh auto-exports SHLVL and the old exit-status bug had
been miscounting it as blanked all along. Its exact-count assertion had
only ever passed because of the bug beside it; it is now a presence
check, since which names a zsh version auto-exports is not this test's
to pin.

Still not pinned, tracked in #394's follow-up: the eval-site identifier
guard has no test behind it.
2026-09-10 09:05:30 +07:00
Dai Ha ed99c209ac Merge #398: a central allow-list of usable models, with the startup validators pinned
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m25s
models: allow: is a single place that names every model the fleet may
use. Absent or empty keeps today's behaviour, so this ships inert until
configured. Once set, a profile naming a model outside the list refuses
to start, and refuses a reload, rather than reaching a backend adapter
as a free-form string.

The allow-list cannot be checked against a provider catalogue: for an
opencode profile fleetd SYNTHESIZES the provider from provider/model
plus baseUrl (OpenCodeLauncher:562-583), so a valid fleetd model id
appears in no published catalogue. An operator-owned list is therefore
the only workable gate.

Includes the #398 follow-up: FleetConfig.validateAll() reflectively
sweeps this class's validateXxx() methods, and both real call sites
(Fleetd.main and ConfigRef.reload) call that one method. Before this,
deleting a validateXxx() call from either caller left the whole suite
green. FleetdStartupValidationTest now drives the real Fleetd.main.

Lead review: mutation on the reload call site is caught (ConfigRefTest,
2 failures). Mutation replacing the reflective sweep with a hardcoded
list is NOT caught (1491 green) — so the sweep is a convenience and the
denominator test is the real guarantee; two false statements in the test
javadoc were corrected to say so (af4c88d).

Recovered work: the worker's agent died mid-turn with the follow-up
uncommitted and the startup call left disabled as
'// MUTATION-TEST-TEMP: cfg.validateAll();'. I restored it before
committing.

NOT covered: the five log-only reportXxx(cfg) calls in main are still
unpinned — filed as #407.
2026-09-10 09:02:18 +07:00
Dai Ha af4c88d54b t398: correct two false statements in FleetConfigValidateAllTest
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m58s
Measured at review: reverting validateAll() to a hardcoded list of
today's six calls leaves the suite green (1491 tests, 0 failures). The
class javadoc claimed that mutation fails a test. It does not — claim 1
pins the generic helper on an unrelated class, claim 2 pins today's six,
and a hardcoded list satisfies both.

The interaction was the real hazard. The denominator assertion IS a
tripwire (declaring a seventh validator fails it), but its failure
message said the sweep reaches new validators 'by construction' and told
the author to just update the expected set. If the sweep were ever
replaced by a name list, the one assertion that fires would hand back a
false all-clear at the moment it fired.

Javadoc now states the measurement, and names the denominator test as
the actual guarantee. The assertion message now says to confirm
validateAll() still delegates to invokeAllValidators(this) BEFORE
updating the expected set.
2026-09-10 09:01:49 +07:00
Dai Ha 44c735f6f5 fleetd #404: report armed detection from startup map
CI / build (pull_request) Successful in 1m25s
CI / contract (pull_request) Successful in 1m20s
2026-09-10 08:56:39 +07:00
Dai Ha b540a1744b fleetd #398 follow-up: pin the startup validators with a reflective validateAll()
CI / contract (pull_request) Successful in 1m9s
CI / build (pull_request) Successful in 1m28s
Mutation testing found that deleting a cfg.validateXxx() call from
Fleetd.main left the full suite green: every test called a validator
directly and none exercised main as the caller.

FleetConfig.validateAll() sweeps this class's own public no-arg void
validateXxx() methods by reflection and invokes each in alphabetical
order, so a newly written validator is wired into both callers
(Fleetd.main and ConfigRef.reload) with no second step to forget.
FleetdStartupValidationTest calls the real Fleetd.main with six configs,
each failing exactly one validator.

Recovered by the lead: the worker's agent died mid-turn with this work
uncommitted, and had left the startup call commented out as
'// MUTATION-TEST-TEMP: cfg.validateAll();' from its own mutation run.
I restored the call before committing. Build after restoring:
Tests run: 1491, Failures: 0, BUILD SUCCESS.

NOT covered, and not claimed to be: the five log-only reporters in
main (reportRequiredSecrets, reportGitHostShape, reportMemberTrustModel,
reportMemberCredentialsGap, and reportExhaustedPatternGap on current
main) are not validateXxx() methods, so the sweep does not reach them
and their call sites stay unpinned.
2026-09-10 08:56:20 +07:00
Dai Ha 0f51d53098 fleetd #399: fix TTL test race by waiting on a real completion stamp, not DONE
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m2s
poll() can report Phase.DONE for a ticket before the whenComplete hook that
stamps Task.completedNanos has run — CompletableFuture.complete() publishes
its result and only then runs dependents. The TTL tests advanced an injected
clock right after observing DONE, so on a host where the hook runs late it
stamps the ADVANCED time and the eviction never happens (fails on Linux,
passes on macOS).

Add a package-private test seam, MessageService.isCompletionStampedForTest,
that reports whether completedNanos is stamped. Both TTL tests now wait on
that (a real volatile read/write happens-before edge) before advancing the
clock, instead of on Phase.DONE. The prune condition in pruneTerminalTickets
is untouched.
2026-09-10 08:54:38 +07:00
Dai Ha c670792ffe #400: classify the blanking loop's result on the observed value, not eval's exit status
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m1s
eval "export NAME=" can return success even when zsh coerces the bare
assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
HISTSIZE, COLUMNS, LINES, USERNAME) instead of failing, leaving the
value unchanged. The old exit-status check then reported the name as
blanked when it was not -- a false receipt.

Classify on the observed effect instead: attempt the export, then read
the name's value back with the (P) indirection flag and decide from
whether it is now empty. One check now covers all three shapes a name
can take here -- a genuine blank, a fatal read-only error eval merely
contains, and this silent no-op -- with the exit status playing no
part in the decision.

Adds a test driving all three shapes through the real scrubScript in
one run (a normal name, LINENO for the fatal case, SECONDS for the
silent no-op), with the parent environment explicitly carrying those
names since a cleared ProcessBuilder parent does not expose them on
its own. Also corrects the previously-merged
unblankableNameInTheMiddleDoesNotAbortNamesAfterIt test, whose "exactly
one failed name" assertion turned out to only pass by accident: zsh
itself auto-exports SHLVL on every shell start, and the old exit-status
bug was silently miscounting it as blanked. The fixed classification
now correctly reports it unblankable too, so the test asserts presence
rather than an exact count.
2026-09-10 08:46:27 +07:00
Dai Ha 3982ace544 fleetd #391: refuse lead fleet replies
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m31s
2026-09-10 08:43:54 +07:00
Dai Ha 7180b1aad0 Merge #395: warn when exhaustedPattern usage-limit detection is off
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m27s
A profile with no exhaustedPattern has usage-limit detection silently
disabled. Startup now reports every unarmed profile (a louder warning for
subscription profiles, which are the ones a limit actually stops), and
fleet_profiles / GET /profiles carry exhaustionDetectionArmed per profile.

Reviewed by the lead: all three requested mutations fail a test, and the
back-compat QuarantineSource ctor defaults to 'not armed' when the source
is unknown. KNOWN GAP, not fixed here: deleting the
reportExhaustedPatternGap(cfg) call at Fleetd.java:140 leaves the suite
green (Tests run: 1472, Failures: 0). That is the same unpinned-startup-
call shape as PR #398's six validators, and fleetd #398's ticket owns it.
2026-09-10 08:43:09 +07:00
Dai Ha c26f695402 fleetd #395: warn when exhaustedPattern detection is silently off
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m25s
exhaustedPattern is opt-in per profile: unset means a usage-limit
refusal on that profile is never classified BACKEND_EXHAUSTED and
never quarantines its credential, with nothing telling the operator.
Add a startup WARN naming every unarmed profile (a louder, separate
WARN for a subscription: true profile, since that is the operator's
own metered plan). Surface the same fact per profile in fleet_profiles
as exhaustionDetectionArmed, so an operator can tell "healthy" from
"can never be caught" without reading fleetd.yaml.
2026-09-10 08:34:33 +07:00
Dai Ha 48877315ca Merge #394: contain the fatal export that aborted the credential scrub
CI / contract (push) Successful in 50s
CI / build (push) Failing after 1m31s
The CB-633 allow-list scrub has been dying mid-loop on every fleet01 member
pane and saying nothing. `export UID=` in zsh is not a failed command — it is
a fatal parameter error that terminates the whole sourced file. The blanking
loop is wrapped in `{ ... } 2>/dev/null`, so the message was swallowed and the
report block after the loop never ran.

Root cause found by the fleet01 lead, with xtrace on a live pane's own ZDOTDIR:

    +scrub.zsh:28> _cb633_n=UID
    +scrub.zsh:28> export 'UID='
    +zsh:1> rc=126        <- file aborted

The severity is the INVERSION, and this is their finding, quoted:

  "env lists inherited names first and the names a startup file exports last.
   So the loop blanks the harmless inherited half and dies immediately before
   the operator's own exports — exactly the credentials the policy exists to
   remove. The selection is inverted, not merely partial."

Measured there: UID is name 42 of 57, and a ~/.zshrc decoy at 58 survived on
8 of 8 spawns. "Partial scrub" reads as "we got most of it"; it got precisely
the wrong half.

Fixed with `eval "export ${n}=" 2>/dev/null` rather than a skip-list of the
known-fatal names (UID EUID GID EGID PPID LINENO). A skip-list has to be
complete forever and this is a security control; eval needs no list. Measured:
plain export dies at UID and every later name keeps its value, while the eval
form completes the loop and blanks all of them. PR #396 proposed the skip-list
and is closed in favour of this; its claim that the abort happens "however the
assignment is wrapped" holds for a direct `if ! export` but not for eval,
which reparses in a nested context.

The report now carries `failed N` and `!`-prefixed unblankable names, and
HerdrPeerLauncher WARNs when any name could not be blanked. The old "no report"
WARN no longer claims the daemon knows what the member saw.

Why the suite stayed green: EnvAllowListScrubTest starts zsh from
pb.environment().clear(), and under a cleared parent UID is not an exported
name at all, so the abort could not reproduce in that harness.

Reviewed by mutation, which found a second gap now also closed: the eval is
only safe because names are filtered to ^[A-Za-z_][A-Za-z0-9_]*$. Replacing
that pattern with .* left the class green, so the line the security property
rests on was unpinned. The guard is now re-asserted at the eval site and pinned
by a test. The reachable vector is a VALUE with an embedded newline, not a
hostile name — measured: zsh strips non-identifier env names outright, while
MULTI=$'keep\njunk.fragment' forges 'junk.fragment' as a candidate name out of
its own value.

Closes #394. Refs #396, #388.
2026-09-10 08:30:15 +07:00
Dai Ha 5c08054533 #394 follow-up: re-assert the identifier guard at the eval call site
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m48s
EnvAllowListScrub's blanking loop splices each name into a string
handed to eval ("export ${n}="). That is only safe because every name
reaching _cb633_blank already passed an identifier check in the
enumeration loop -- 20 lines away, in a different loop. Before eval
was introduced a non-conforming name reaching plain `export "$n="`
was inert either way (the quoting neutralized it); eval removed that
safety net, so the enumeration loop's guard became the ONLY thing
standing between a non-identifier string and code execution in the
member's pane, with nothing at the eval site itself defending that
property.

Re-assert the same [A-Za-z_][A-Za-z0-9_]* check immediately before
the eval call, independent of the enumeration loop's own guard (left
untouched, not moved). A name that fails it is counted unblankable
rather than silently dropped, so a bypass of the upstream guard would
leave real evidence in the report.

New test exploits the "junk from multi-line values" gap the
enumeration loop's own comment already documents: a value with an
embedded newline makes `command env`'s text output split into a
spurious extra "name" line that was never a real variable. Runs the
real generated scrubScript() end-to-end under zsh and asserts the
non-conforming fragment is neither blanked nor counted unblankable.
The fragment used is merely non-conforming (contains a dot) --
never command-shaped.

Mutation-verified both guards. Weakening the enumeration guard alone
DOES break the new test (the fragment then reaches the new eval-site
guard and gets counted unblankable, failing the "not unblankable"
assertion). Removing the new eval-site guard alone, with the
enumeration guard intact, does NOT break it: _cb633_blank has exactly
one producer (the enumeration loop), so nothing can reach the eval
site without already having passed the identical check there. That is
expected given the single-source architecture, and it is exactly why
the eval-site guard is defense-in-depth against a future change that
adds a second path into _cb633_blank or decouples the two loops --
not a currently independently-observable divergence.
2026-09-10 08:24:46 +07:00
Dai Ha b6db9c31f5 charter: a peer lead is answered with fleet_send, not fleet_reply
CI / contract (push) Successful in 46s
CI / build (push) Successful in 2m0s
`fleet_reply` has no route to a peer lead. `AmqpReplyInbox` publishes to
`agent.<target>.inbox`, mandatory, and a lead's own terminal has no such
queue, so the publish is refused. `MessageService.reply()` has no peer
branch at all — `grep -c 'coord\|LeadMailbox'` on it returns 0. The charter
told every lead to use a tool that cannot work, and both leads here hit it.

Three edits to the canonical block, byte-identical with the wiki template
(pushed as 803726a; the in-sync check in this file reports True):

- the intent->tool row now says `fleet_send{coordId}`, or `{sessionId}` for
  a peer on the same host, and says plainly that `fleet_reply` is refused
- the prose says WHY: `fleet_reply` resolves a member's blocked `fleet_send`,
  while a peer's coord-id message is durable and non-blocking, so there is
  nothing for it to resolve
- lead<->lead item 3 gains the data-point rule: N observations are N data
  points only if they differ in the axis you are trusting

Wording for all three drafted by the fleet01 lead, who verified the missing
queue namespace independently in its own tree. The data-point rule has now
caught three separate errors in a day, in both directions: one cause blamed
for N failures, and N agreeing measurements that shared a single instrument.

The refusal message itself is still wrong — it says "queue not declared or
owned", which sends the reader to the broker instead of to this file. That
half stays open on #391.

Tracked as fleetd #391.
2026-09-10 08:22:35 +07:00
Dai Ha e7b33fe3a0 fleetd: central allow-list of usable models (models: + validateModels())
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m34s
Add an optional top-level `models:` block (Models{allow: List<ModelEntry>})
naming the models any profiles: entry may use. Absent/empty allow: keeps
today's behaviour exactly (no check, no warning). When configured,
FleetConfig.validateModels() fails config load (and reload, via ConfigRef)
naming both the model and the profile, if any profile's model: is outside
the list. The check is one-way: editing profiles: alone can never widen
what is permitted, only models.allow: can.

Wired into Fleetd.main() alongside the other validateXxx() calls, and into
ConfigRef.reload()/DEFERRED_KEYS so a bad edit can't slip in through a
reload either. Each ModelEntry is its own record (not a bare string) so a
later unit can add per-model on/off or load-limit state without changing
the YAML shape. One flat string namespace covers both a bare Claude id and
an opencode provider-prefixed id.
2026-09-10 08:09:27 +07:00
Dai Ha e3e403e5c8 #394: contain a fatal export error instead of letting it abort the scrub
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m52s
EnvAllowListScrub's blanking loop used a plain `export "$n="` on every
name not on the allow-list. For a zsh read-only/special parameter (e.g.
UID) that is a FATAL parameter error, and since the loop runs inside the
sourced startup file, the error aborts the whole file: every name still
to come is never blanked, and scrub-report.txt is never written at all
-- silently, because 2>/dev/null on the group swallows it.

Route each blanking attempt through `eval` instead, which contains the
error to that one iteration. The loop always finishes; a name it could
not blank is now counted separately ("failed" on the report's first
line) and listed !-prefixed rather than disappearing. No skip-list of
known-bad names is added -- every enumerated name is still attempted,
so a name nobody has thought of is still tried and, if it fails, still
counted.

HerdrPeerLauncher: log a WARN when a pane's report carries a nonzero
failed count, and reword the "no report at all" WARN so it no longer
claims the daemon knows the member "saw the full host environment" --
a partial vs. a missing scrub are different situations and only the
first is now distinguishable from the report alone.
2026-09-10 08:08:41 +07:00
Dai Ha 799014e99d Merge #388: scrub a pane shell that is neither login nor interactive
CI / contract (push) Successful in 1m8s
CI / build (push) Successful in 1m38s
2026-09-10 07:21:48 +07:00
Dai Ha 69e09b10fa #388: scrub a pane shell that is neither login nor interactive
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m32s
EnvAllowListScrub generated four zsh startup files but only .zshrc and
.zlogin sourced the scrub — .zshenv (the one file zsh always reads) did
not. A pane shell that is neither login nor interactive reads only
.zshenv and stops, so it was never scrubbed at all (measured on fleet01,
issue #388).

Adding an unguarded scrub to .zshenv (the ticket's own suggested fix) is
wrong: .zshenv is read by every zsh, including a short-lived `zsh -c`
a member's own tooling forks for a single command. Those children are
also neither login nor interactive, so they would scrub the environment
their parent deliberately set for them (GIT_DIR, VIRTUAL_ENV, ...), and
the rewritten scrub-report.txt would describe the last child to exit
instead of the pane.

Fix (per comment 15387, measured): keep .zshrc/.zlogin unconditional,
and add to .zshenv a pass guarded on the exact condition that defines
the gap (neither login nor interactive), plus a per-pane sentinel
(_CB633_SCRUBBED) so it runs once per pane, not once per process. The
sentinel is exported only after the scrub runs, and is folded into the
scrub's own allow-list so a later pass in the same pane cannot blank it
back to empty.

Also corrects the class javadoc's wrong premise (a bare argv[0] proves
NOT login, not "therefore interactive") and its now-stale two-file
walkthrough.

Tests: two new real-zsh tests in EnvAllowListScrubTest run actual
non-login/non-interactive zsh processes (never string-match the
generated files) to prove: a neither-shell pane is scrubbed; a child
that pane forks keeps variables the pane deliberately set for it; the
child does not re-scrub; and scrub-report.txt still describes the pane
after the child exits. Both fail without the production fix (verified
by reverting it and re-running: AssertionFailedError on the sentinel
and on the decoy secret surviving).
2026-09-10 07:21:08 +07:00
Dai Ha b9d09e044e t386: pin the per-member drift baseline the fix's own tests left open
CI / contract (push) Successful in 37s
CI / build (push) Successful in 1m57s
The two tests merged with #386 both start with the member already BUSY, so a
single global drift baseline passes them. This one sleeps the host while nothing
is busy and only then starts a turn, which fails without the per-member map.
2026-09-10 06:59:42 +07:00
Dai Ha fd8650cda4 Merge #386: correct the stall check for a monotonic clock frozen by host sleep 2026-09-10 06:56:37 +07:00
Dai Ha 11050e24ed t384: fix javadoc indentation on the merged shareWithGroup lines
CI / contract (push) Successful in 1m7s
CI / build (push) Successful in 1m29s
2026-09-10 06:54:51 +07:00
Dai Ha 769f282408 #386: give the stall detector a real-time clock, log the divergence
CI / contract (pull_request) Successful in 1m18s
CI / build (pull_request) Successful in 1m26s
FleetHealthMonitor.tick's stalled check compared two monotonic-clock
readings (System.nanoTime(), which macOS freezes across a host sleep),
so a member BUSY for 101 real minutes was never flagged.

The monitor now also takes a wall-clock LongSupplier (realtimeClock),
used only inside the stall check. Each tick measures how far the two
clocks moved apart since the previous tick and folds any positive
divergence into a running total; when a single tick's divergence
exceeds one tick interval (the signature of a sleep, since a tick
cannot run while the process itself is suspended) it logs one WARN
naming how long the detector could not see. The correction is applied
per member, keyed to when that member's current lastActivityAtNanos
was first observed BUSY - not since monitor start - so a sleep that
happened before a member went busy is never charged to it.

Every other use of the monitor's clock (readiness grace, snapshot
timestamp) is unchanged. Backend quarantine/cool-off, the lead tab
scan, the completion resolver, the session reaper and the message
service TTLs are untouched, per the ticket's decision.

Existing FleetHealthMonitor/FleetHealth tests pass unmodified (none of
them ticks a BUSY session more than once, so the drift path never
engages for them). Two new tests: a frozen monotonic clock past the
real-time threshold produces STALL_SUSPECTED, and a single sleep gap
logs the divergence exactly once, not once per tick.
2026-09-10 06:53:30 +07:00
Dai Ha 6f71f40047 Merge #384: pre-create the scrub receipt and give it group write 2026-09-10 06:49:46 +07:00
Dai Ha fb36c5238f Merge #382: SpawnRequest.withProfile() replaces the six-accessor rebuild 2026-09-10 06:49:46 +07:00
Dai Ha cc9cdc938b #384: write shared scrub receipts 2026-09-10 06:45:30 +07:00
Dai Ha 5fede82468 #382: give SpawnRequest a withProfile wither, guard it against the arity trap
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m49s
CompositePeerLauncher:372 rebuilt a routed SpawnRequest from a literal
new SpawnRequest(...) call listing six of the original request's own
accessors. That call is only correct because it happens to match the
canonical 6-arg constructor today; add a 7th component plus the
established back-compat constructor at the old (now-shorter) arity and
this call would silently rebind to it, dropping the new field on every
profile-routed spawn with no compile error — the same defect shape
already guarded on FleetConfig.withDefaults() (#357) and MemberSession
(#358).

Add SpawnRequest.withProfile(String), modeled on
MemberSession.withState/withActivity, and use it at the call site
instead. Add a guard test that resolves the true canonical constructor
by exact component types (never by argument count), gives every
component a distinctive value, and asserts every component but
profileName survives withProfile() unchanged.

Proved the guard against the real mechanism: temporarily dropped the
last (role) argument from withProfile()'s constructor call so it bound
to the 5-arg back-compat constructor — it still compiled, and the new
test failed, catching the silently-defaulted role. Restored the fix
and reconfirmed green.
2026-09-10 06:44:15 +07:00
Dai Ha 1515025804 charter: a profile differs in liveness, not just model and cost
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m31s
The old line said profiles differ in model and cost. That is true and it is
not the reason the default hurts. Measured on two hosts: a default sitting on
an exhausted or withdrawn credential either fails the spawn loudly or, worse,
produces a member that starts fine and then returns nothing.

Wording proposed by the fleet01 lead; merged with the existing 'not in tier'
clause, which is still right. Applied byte-identically to the wiki template.
2026-09-10 06:28:51 +07:00
Dai Ha 7754f53662 Merge t385 follow-up: pin the mid-turn redelivery ack 2026-09-10 06:28:08 +07:00
Dai Ha a507f7b31b t385: pin that a redelivery is acked even while the lead is mid-turn
Verification found the fix's dedup check could be moved behind the
injectable-status gate and every test still passed. That placement matters:
this lead is mid-turn most of the time, so gating the ack on an idle pane
leaves the redelivered message held, and the next recovery delivers it again.

The new test fails on that mutant and passes on the fix.
2026-09-10 06:28:03 +07:00
Dai Ha 71c322f104 Merge #385: a redelivered lead message must not write the pane twice 2026-09-10 06:24:40 +07:00
Dai Ha fde2c15627 t385: a redelivered lead message must not write the pane twice
An AMQP recovery clears the held delivery tags, the broker redelivers with
fresh ones, and the coordination loop wrote the same peer message into the
lead pane again. Measured on the live daemon: one msgId reached the pane 12
times in 9 hours, across 19 recovery events.

LeadCoordLoop now remembers the msgIds it has written to a pane (bounded at
1024) and acks a redelivery without a second write. LeadMailbox.ack no longer
returns quietly for an unknown msgId: it throws, so an ack that never reached
the broker is reported instead of hidden. A repeat ack that this connection
already completed stays quiet, tracked in a bounded set.
2026-09-10 06:24:35 +07:00
Dai Ha 2830735644 t377: raise the logger level in the test, or it asserts on an empty list
CI / contract (push) Successful in 46s
CI / build (push) Failing after 1m32s
The 9 tests the member wrote failed 7 of 9 on first build. The production
code was correct; the test harness was not. logback-test.xml sets
dev.ltms.fleet to WARN, so the INFO shape lines were dropped by the level
check before any appender saw them.

The member copied attach()/detach() from MemberTrustModelReportTest but not
the setLevel(INFO) those siblings do at each call site. Doing it inside
attach()/detach() covers all nine at once, and restores the original level
(null, meaning inherit) rather than a concrete one.

Proven by mutation: making the code log the value fails 4 tests, including
theValueNeverAppearsInLogOutput — 'the full GITEA_HOST value must never
reach the log'. Full build 1459 tests green.
2026-09-09 07:52:43 +07:00
Dai Ha 4e3ac91a22 Merge worker/t377-7f587b-7: startup reports the git host value's shape, never its value 2026-09-09 07:49:09 +07:00
Dai Ha 29d3f0b41f t377: report the git host value's SHAPE at startup, never its value
A member receives the git host as GITEA_HOST and often gets a full URL
(scheme, trailing slash) where it expects a bare host, so it builds
https://https://... and the request never leaves the machine.

Logs set/unset, length, startsWithScheme and trailingSlash next to the
existing secret report. The value itself is never logged, and the value
is still passed to members unchanged — rewriting it here would change
what works on one host and breaks on another.

Written by the gx member on branch worker/t377-7f587b-7; it could not
build, commit or push because the host command classifier refused every
shell command (see #381). Build and verification are mine.
2026-09-09 07:49:02 +07:00
Dai Ha cbc732444f fleetd #365: the push-nudge metric outcome is 'sent', not 'delivered'
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m31s
The label was renamed in the #365 merge. This table still named the old one.
'sent' counts the herdr paste-and-submit call returning, never a confirmation
that the pane read it.
2026-09-09 07:40:57 +07:00
Dai Ha 6f828b8c38 Merge worker/t365-3920c5-3: t365: fleet_reply/REST reply distinguish resolved vs queued; rename nudge metric outcome
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m48s
2026-09-09 07:35:00 +07:00
Dai Ha 145a8c8862 Merge worker/t373-336973-2: t373: pin the production XDG-excludes seam GitWorktreesTest.seedingGitWorktrees builds 2026-09-09 07:35:00 +07:00
Dai Ha 7057291739 Merge worker/t358-6e989b-1: t358: guard MemberSession's 5 rebuild sites and Profile.withProfile() against the back-compat-arity trap 2026-09-09 07:35:00 +07:00
Dai Ha 2302b3bc11 #376: the too-fast failure stops asserting a cause it cannot know
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m31s
A turn that settles inside MIN_TURN_NANOS is still reported FAILED. Only the claim
about WHY is withdrawn, and the pane is still carried.

Measured on fleet01 on 2026-09-08 UTC: an opencode member on mimo-v2.5-free answered
a real question in 1575ms, below the 2000ms floor. The daemon reported the turn as
failed with 'most likely a backend error before any work started'. The answer was
right there in the scrape. With no errorPattern configured — the live state on both
hosts, which both log as 'backend-error classification: off' — that sentence is a
guess, and a reader who believes it stops looking at the pane.

WHAT I REJECTED, because the next person will try it. A worker implemented the
ticket's first suggested direction: inspect the pane inside the floor and resolve a
COMPLETION when the text looks like a real reply. Its test for 'looks like a real
reply' was non-blank plus a '.', '!' or '?' anywhere in the text. That is unsafe
twice over. lastAssistantBlock falls back to the WHOLE pane when it finds no U+23FA
marker, so on a crash the candidate reply is the entire screen; and a crash pane
almost always contains a full stop, in a file path, a version or a hostname. I ran
that implementation against the new guard test and it resolved

    Error: connection reset while loading src/main/java/Foo.java v1.2.3

as a COMPLETION — expected: <FAILED> but was: <COMPLETION>. A loud wrong answer
became a silent one, which is the trade the ticket brief forbade.

The obvious repair does not work either. Requiring the U+23FA marker as positive
evidence would be safe, but that marker is Claude Code chrome and an opencode pane
never carries it — and an opencode member is what raised this ticket. There is no
reliable cross-backend marker for 'this is a real reply', so this path must not try
to judge one. That reasoning is now in the failTooFast javadoc.

Two guard tests. aPlausibleLookingReplyInsideTheFloorStillFails pins the safety
property against exactly the rejected approach; it passes today and fails against
that implementation, which is how it was verified rather than assumed.
theTooFastFailureDoesNotAssertACauseItCannotKnow pins the wording.

One existing test changed. aNonMatchInsideTheFloorStaysGenericAndNeverNotifiesTheSink
asserted the phrase 'too fast to be real work', which carried the withdrawn claim. It
now asserts what it was really guarding: the floor alone fails the turn, the reason
stays generic, the pane is carried, and the typed sink is never notified.

MIN_TURN_NANOS is unchanged at 2000ms. mvn clean install: 1441 tests, 0 failures.
2026-09-09 07:28:04 +07:00
Dai Ha 3a004dc1b3 t373: pin the production XDG-excludes seam GitWorktreesTest.seedingGitWorktrees builds
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m32s
fleetd #362 review finding 2 protects GitWorktrees#previouslyEffectiveExcludesFileContent's
Java-side XDG_CONFIG_HOME/HOME read (it never goes through a git subprocess, so no
GIT_CONFIG_GLOBAL/GIT_CONFIG_SYSTEM isolation reaches it) with a gitEnv constructor seam. A
mutation run during the #372/#369 merge found that seam unpinned: stripping hermeticGitEnv(tmp)
from seedingGitWorktrees left every test green, poisoned XDG_CONFIG_HOME or not.

Adds seedingGitWorktreesResolvesTheExcludesFileFallbackInsideItsThrowawayDirectory, which asserts
the property directly (a GitWorktrees built for seeding resolves the fallback inside its own
throwaway directory) using a self-contained marker instead of relying on an externally poisoned
env var. Refactors hermeticGitEnv/seedingGitWorktrees into two-argument overloads (one taking an
explicit XDG_CONFIG_HOME / gitEnv) so the new test can pre-populate the marker before construction
while still going through the same production construction every other seeding test uses; no
behavior change for the 4 existing call sites.
2026-09-09 07:24:22 +07:00
Dai Ha a9a3c12232 t365: fleet_reply/REST reply distinguish resolved vs queued; rename nudge metric outcome
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m51s
fleetd #365. fleet_reply always returned the literal "delivered" and
POST /sessions/{id}/reply always returned {"delivered": true}, whether
the reply resolved a live waiting send/ticket or was merely queued in
the inbox for a later drain (CB-307) — both are successes, but not the
same fact.

MessageService.reply() now returns a ReplyOutcome (RESOLVED_SEND,
RESOLVED_ASYNC_TICKET, or QUEUED) instead of an always-true boolean.
FleetMcp.reply and FleetApp.replyMessage both read it: the MCP tool
result names which happened, and the REST body's "delivered" field is
now accurate, with an added "outcome" field.

Also renames the heartbeat/push-loop nudge metric's "delivered" outcome
to "sent" (LeadHeartbeatLoop, ReplyPushLoop, FleetMetrics): it only
records that the herdr agent.prompt paste-and-submit call succeeded,
never that the lead's pane actually read it — there is no read-receipt
concept at that layer, so "delivered" overclaimed there too.

Tests: MessageServiceTest/FleetMcpTest/FleetAppTest strengthened to
assert the specific outcome per case (a resolved send, a resolved async
ticket, and a queued reply); ReplyPushLoopTest updated for the outcome
rename.
2026-09-09 07:23:41 +07:00
Dai Ha 94ec77a1bc t358: guard MemberSession's 5 rebuild sites and Profile.withProfile() against the back-compat-arity trap
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m53s
Follow-up to #357 (FleetConfig.withDefaults()). Reflectively enumerate each
record's own components, resolve the canonical constructor by exact
component types, build a real non-null value per component, run each
rebuild site, and assert every component survives (except the one it is
documented to change). Exclusion lists pinned at 0 for both.

Re-counted the Profile back-compat ladder directly against the source:
8 constructors (arities 25, 24, 22, 20, 18, 15, 14, 12) against a
canonical arity of 26 — the ticket's own number was explicitly untrusted.
2026-09-09 07:19:12 +07:00
Dai Ha 49f285cfda charter: a measured fact in an addendum must carry its own deletion trigger
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m36s
The operator chose this and its size; the reasoning below is the fleet01 lead's.

Placed in the canonical block's boundary paragraph rather than the orchestration
body. That paragraph already talks about the addendum layer instead of protocol,
every project that mounts the bridge inherits it, and it sits about 3800 characters
before the primary's step list, so it does not dilute the steps a lead reads while
working. Perishability is structurally an addendum problem: the block is
byte-identical across projects by construction, so a dated local measurement in the
block body would already be a layering violation.

What happened. The fleet01 lead's kb addendum held a dated merge-refusal section
that carried an instruction to delete itself once it stopped reproducing. On
2026-09-08 UTC the operator granted merge rights on akb/kb, the lead re-ran the
probe, got 409 'head out of date' where the identical request had returned 405
'User not allowed to merge PR' on 2026-09-06, and deleted the section as instructed.

Why four parts and not one. The lead's finding is that the banner did not work
because it was emphatic. It worked because the falsification condition was
executable: it carried the exact probe, the reason for the all-zeroes
head_commit_id, and what each response code meant. The lead did not have to
reconstruct the experiment or decide what would count as refutation, and just ran
it. A banner saying 'this may be out of date, verify before relying on it' costs
the same space and does nothing, because deciding what would falsify a claim is the
expensive step and a reader in the middle of another task will not pay it. So: the
date, the command, what each outcome means, and the instruction to delete. The
fourth without the second is decoration.

The closing clause is the justification for the machinery. Most stale notes are
merely wrong. This one went stale in the dangerous direction: it would have told a
future lead it could not merge at the exact moment merging became its job, silently
and with confidence. A note that goes harmlessly stale does not need this.

Note what is NOT centralized here. The banner text itself cannot be. What fired for
the lead was a specific instruction sitting on top of the specific stale fact, which
it could not read past on its way to acting. A rule elsewhere saying 'date your
measurements' would not have fired, because nobody reads that rule at the moment
they re-measure. This sentence sets the convention; the trigger still has to live
next to the fact it governs.

Propagated to the wiki template in the same turn, wiki 8c4f152 on main; the sync
check in this file's addendum reports 'in sync: True'.
2026-09-09 04:24:17 +07:00
Dai Ha 5a12ae7930 charter: test a refusal, and do not count a transport failure as one
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m38s
Step 8 gained a refusal paragraph in 2f71a30, which said what a lead does once the
forge refuses a merge. It did not say how a lead establishes that it was refused.
Both halves of this amendment come from the fleet01 lead, measured on akb/kb on
2026-09-08 UTC, and both are ways to be wrong about a permission you never tested.

Do not read a refusal off a permissions field. After the operator granted merge
rights, the lead re-ran its probe: POST .../pulls/53/merge with an all-zeroes
head_commit_id, chosen so the request cannot succeed on its merits and a rejection
can only mean the refusal. It returned 409 'head out of date' where the identical
request returned 405 'User not allowed to merge PR' on 2026-09-06. A 409 is payload
validation and sits after the permission gate, so the grant took. The lead reports
the repository permissions object did not change across that flip -- still
admin:false, push:true, pull:true. I did not read that object myself; my forge token
is a different identity and would return a different one, so this stays the lead's
measurement and not mine. Merge rights on a protected branch live in branch
protection, so a permissions field can be wrong in both directions.

Do not count a transport failure as a refusal. The lead's first attempt returned
HTTP 000, because GITEA_HOST already carries a scheme and a trailing slash and the
URL came out as https://https://git.ltms.dev//api/... Under a 'not 200' test that is
indistinguishable from being refused. A probe exists to separate a refusal from
everything else, so an error that never reached the gate has to be a third answer
that concludes nothing.

Propagated to the wiki template in the same turn, wiki 8c2ef96 on main; the sync
check in this file's addendum reports 'in sync: True'.
2026-09-09 04:09:08 +07:00
Dai Ha 127e6832a9 Merge #374: fleetd holds off idle sleep while any member is live
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 2m3s
Lands PR #355 (fleetd #354's sibling), rebased onto current main by a worker
after 39 commits of drift left it unmergeable.

The problem, measured on the original branch: a fleetd host idle-slept after as
little as one minute (pmset -g custom reported 'sleep 1' on battery). Overnight
the daemon's AMQP link dropped 13 times, and every drop minute had a sleep or
wake event in pmset -g log in the same minute or the one before. The AMQP churn
is the visible symptom; the real cost is a member mid-turn freezing with the
host, and a long turn with nobody typing is exactly the case that goes idle.

IdleSleepGuard holds an OS-level assertion for as long as at least one member is
live. It is driven by SessionManager's existing onAcquire/onRelease hooks rather
than a second member count kept in parallel, so it reads the same registry
fleet_list's numbers come from, and only a real 0->1 or 1->0 crossing touches the
OS. It fails safe: a mechanism that cannot acquire means nothing is ever held,
and it never throws, never blocks a spawn, a release, or shutdown.

Conflict resolution was the whole job, and all three were in config plumbing:
ConfigRef, FleetConfig and ConfigRefTopLevelReportingCoverageTest. The power
package is byte-identical to the original branch commit.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1439, Failures: 0, Errors: 0, BUILD SUCCESS
  (1425 on main + 14 new: 4 caffeinate, 5 guard, 1 wiring, 4 config)

The denominator recount, which the worker flagged as its own weakest number
because this file's count has drifted three times before (#330/#333/#337). I
counted it mechanically rather than reading it: FleetConfig has 24 canonical
record components; COLD_KEYS 5, DEFERRED_KEYS 13, SPLIT_KEYS 3, plus the 3 the
javadoc names as hot-excluded (placement, memberCredentials, memberLoginShell).
5+13+3+3 = 24. The javadoc's '24 components: 5 cold, 13 deferred, 3 split, 3
hot-excluded' is correct. The worker's prose called idleSleepGuard the 25th
constructor argument; it is the 24th. The code is right, the report was off by
one.

Mutation run on merge, on the half the worker verified by READING rather than by
proving -- it said it had checked that withDefaults()'s final call binds the true
canonical constructor. I dropped the trailing idleSleepGuard argument so the call
silently binds the 23-arg back-compat overload. It compiles, which is the whole
hazard. Caught: 1 failure, 3 errors, BUILD FAILURE, and
FleetConfigWithDefaultsPreservesEveryComponentTest names the dropped component
and prints its own denominator -- '24 components, 24 checked, 0 excluded, 23
survived'. That test was added on main after this exact defect happened live when
idleSleepGuard was added on a sibling branch; the worker had to add the missing
entry to it, and doing so is what makes the guard cover this component at all.
2026-09-07 20:31:41 +07:00
Dai Ha 6442a583ae t355: give the withDefaults() coverage test a value for idleSleepGuard
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m6s
FleetConfigWithDefaultsPreservesEveryComponentTest was added on main after
the sleep-guard commit was cherry-picked, specifically to catch this exact
rebase hazard (a withDefaults() call silently rebinding to a stale-arity
back-compat constructor). Its baseValues() map didn't know about the new
idleSleepGuard component yet, so the coverage test itself failed the
name-drift check. Add a real, non-null value for it, consistent with how
the sibling ConfigRefTopLevelReportingCoverageTest already covers it.
2026-09-07 20:25:33 +07:00
Dai Ha 24b96d29ae fleetd must not let the host idle-sleep while members are live
Adds a small IdleSleepGuard (dev.ltms.fleet.power) that holds a macOS
caffeinate -i child while at least one fleet member is live, and
releases it once none are. It hangs off SessionManager's existing
onAcquire/onRelease hooks and SessionManager#size() rather than
tracking members a second way. New idleSleepGuard: config block,
on by default, following the FleetConfig.Health/ConfigReload pattern.
2026-09-07 20:22:11 +07:00
Dai Ha 4721771052 charter: a peer reads your project addendum
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m47s
Fourth item under Lead <-> lead. The addendum is instruction surface -- every
future session on that host obeys it, and a wrong one is obeyed as faithfully as
a right one -- but nothing in the block said to have anyone check it.

Evidence is one addendum, written on fleet01 this week, and it carried two
defects that its author did not see and a non-author did. First: it described a
forge permission wall as if it were policy, so a token regrade would have made
it tell a session NOT to merge at the moment merging became its job. Second,
found only because the first was raised: it DID carry the 'this is a dated
measurement' caveat, at the bottom, after the prohibition -- so a session that
read the instruction and stopped had already taken it as policy. The caveat sat
downstream of the thing it qualified.

Neither is a writing slip. Both are the author being unable to see their own
qualifier placement, which is what a second reader is for.

Deliberately conditional: not every operator runs two leads, so the rule ends
with what a lone lead can still do -- ask which sentence goes false first, and
whether a reader reaches the caveat before acting.

Weaker evidence than the step 8 amendment (n=1 addendum, 2 defects, versus a
measured 405). Recorded as such so it can be dropped if it does not earn the
context it costs.
2026-09-07 05:03:58 +07:00
Dai Ha 2f71a30bd7 charter: say what step 8 means when the forge refuses the merge
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m35s
Step 8 said 'then merge' and nothing about a refusal. fleet01 is the first host
to hit that: on akb/kb its lead gets 405 'User not allowed to merge PR' from the
API, and main is protected, so merging locally and pushing is refused too. Both
routes shut. The charter was telling a lead to do something the forge would not
let it do, and that gap was latent in every copy of the block.

The amendment keeps the rule that the merge decision is never delegated, while
admitting the mechanical merge may not be the lead's to make.

The second sentence is the one that earns its keep, and it came from the fleet01
lead rather than from me: never call a PR 'ready to merge' without having read
the diff. A refusal is exactly when that shortcut is tempting, because no action
is left that forces the lead to look. Without it, a refused merge quietly turns
step 8 from 'read it yourself, then merge' into 'forward the reviewer's verdict'
— the proxy-delegation the same step forbids two lines earlier.

Propagated to the wiki template in the same turn; the sync check in this file's
addendum reports 'in sync: True'.
2026-09-07 05:01:31 +07:00
Dai Ha 22cdebbdbe fleetd #369: state hermeticGitEnv's real scope, measured
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m32s
The comment claimed no test in this class can reach the real machine's home
directory. Measured on the merge: 53 of the 58 new GitWorktrees(...)
constructions here pass no env override, and stripping the override from
seedingGitWorktrees leaves the class green under the poison command that the
same comment cites as proof. Say what it covers and what it does not.
2026-09-06 20:32:21 +07:00
Dai Ha 154971c2b8 Merge #372: GitWorktreesTest no longer reads the operator's real git config
fleetd #369. Every git subprocess the test class starts now goes through one
gitProcessBuilder factory that applies the hermetic environment. Before this,
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus set nothing at all -- so the tests inherited the JVM's
whole real environment, including the operator's default excludes file. That
file applies with no core.excludesFile configured at all, and /dev/null for the
global config does not stop it.

Round 1 pinned this with a call-site count: exactly 2 literal
new ProcessBuilder( occurrences. That catches a NEW helper built the old way,
but it is a proxy, not the property. I measured the gap -- deleting
pb.environment().putAll(hermeticEnv()) from inside the factory left every call
site unchanged, the count stayed 2, and 1413 tests stayed green. Round 2 added
gitProcessBuilderCarriesTheFullHermeticEnvironment, which asserts on what the
factory actually hands to ProcessBuilder#start(). Both checks are kept: they
catch different regressions.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1414, Failures: 0, Errors: 0, BUILD SUCCESS
  poison control on the PRE-FIX file:
    XDG_CONFIG_HOME=<dir with a '*' git/ignore> mvn test -Dtest=GitWorktreesTest
    -> tests=59 failures=56, so the poison genuinely reaches these tests

Mutation run on merge, on a half neither the worker nor the reviewer touched --
made seedingGitWorktrees pass null instead of hermeticGitEnv(tmp), which strips
the hermetic environment from the PRODUCTION GitWorktrees instances rather than
from the test's own subprocesses: tests=61 failures=0 unpoisoned AND poisoned.
That half is unpinned. It is fleetd #362's protection, not this ticket's, and it
guards a different path -- the Java-side XDG read in
previouslyEffectiveExcludesFileContent, which no assertion observes. Out of
scope here; filed as a follow-up rather than held against this PR.

One javadoc sentence corrected in the merge: hermeticGitEnv claimed "no test in
this class can reach the real machine's home directory". 53 of the 58
new GitWorktrees(...) constructions in this file pass no env override at all, so
the claim is true of the 5 seeding sites and of every test-started subprocess,
not of the class.
2026-09-06 20:31:01 +07:00
Dai Ha fef287c346 Merge #371: a stale lead binding no longer swallows a worker's reply nudge
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m55s
fleetd #368. PrimaryRegistry.forgetDelegation fires only when a WORKER is
released, never when the delegating LEAD goes away. The map is keyed by the
worker, so a closed, crashed or relaunched lead left its bindings behind. A
stale non-null entry then beat nudgeTargetFor's single-primary fallback every
time -- and that fallback's own javadoc argues it is correct precisely in the
case the stale entry was hiding.

ReplyPushLoop now probes the recorded lead with agents.status before trusting
it, at all five entry points, and a lead that is really gone is forgotten so
resolution reaches the fallback.

Round 1 caught any RuntimeException and treated it as death, and death here
calls forgetDelegation -- destructive and permanent on ONE reading. That is the
#359 mistake repeated two days later: a socket blip or a decode error on a live
lead would silently unbind it forever. I measured the breadth was unpinned
(narrowing it left 1414 green), and round 2 narrowed it to the one affirmative
signal AgentControl.agentCall itself uses, agent_not_found. Every other failure
is treated as live, because guessing wrong costs one extra retry next tick while
guessing wrong the other way is unrecoverable.

The fallback it lands on is live, not stale: PrimaryRegistry.record overwrites
the single slot on every orchestration-side MCP call, and neither host pins
primary.terminal, so a relaunched lead re-registers on its first tool call.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1415, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge, on a half the worker did not touch -- dropped the
forgetDelegation call while still returning the fallback, so behaviour on the
first nudge is identical and only the self-healing is lost: 2 failures,
BUILD FAILURE. The cleanup is pinned, not just the fallback.
2026-09-06 20:13:07 +07:00
Dai Ha a37acd5ee3 fleetd #369 review round 2: pin the factory's behaviour, not its call count
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m26s
everyGitSubprocessGoesThroughTheHermeticFactory counts ProcessBuilder("git",
...) call sites, so it catches a new helper built the old way, but a
reviewer proved it does not catch gitProcessBuilder itself being gutted:
removing pb.environment().putAll(hermeticEnv()) from inside the factory
leaves every call site unchanged, the count stays 2, and the whole
unpoisoned suite stays green.

Add gitProcessBuilderCarriesTheFullHermeticEnvironment, which inspects what
the factory actually hands to ProcessBuilder#start(): every hermetic key
present with the isolating value, and XDG_CONFIG_HOME pointed inside the
class's own throwaway directory rather than left unset or pointing at the
operator's real one. This fails the moment the hermetic environment stops
being applied, on any machine, with no poison needed. Keep the call-site
count check too — the two catch different regressions.
2026-09-06 20:09:32 +07:00
Dai Ha d6ef0c8013 fleetd #368 review: only agent_not_found may forget a lead binding
CI / build (pull_request) Successful in 1m22s
CI / contract (pull_request) Successful in 1m22s
isLive treated any RuntimeException from the liveness probe as "the lead is
gone", which forgetDelegation then acted on destructively and permanently.
That made a transient herdr hiccup (socket blip, decode error) on a perfectly
live lead indistinguishable from the lead actually being dead — the same
one-bad-reading mistake #359 shipped a guard against for lead-tab liveness.

Narrow isLive to match AgentControl.agentCall's own rule: only an affirmative
HerdrException("agent_not_found") counts as gone. Every other failure is
treated as still live and the binding is left alone.

Adds aTransientLivenessFailureMustNotForgetABindingToAStillLiveLead, which
fails with the bare RuntimeException catch and passes with the narrowed one.
2026-09-06 20:08:18 +07:00
Dai Ha dfb70871b4 fleetd #360 follow-up: the install block must say loginctl enable-linger
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m54s
PR #370 shipped both units but its install block stopped at
`systemctl --user enable --now`. Without lingering a user manager starts at
your first login and stops at your last logout, so the units do not come back
after a reboot -- which is the whole reason this ticket moved fleet01 off the
setsid scripts.

It is easy to miss because leaving it out looks like success: `systemctl --user
enable` reports "enabled" and both units run while you stay logged in. The
issue named this and the PR did not carry it over.

fleet01 itself is fine -- measured `Linger=yes`, both units `enabled`. This is
about the next host that follows these instructions.

Comment only; SystemdUnitSafetyTest still 8 green (a commented line is not an
active directive).
2026-09-06 20:08:05 +07:00
Dai Ha 4ee7b16929 Merge #370: ship systemd units that do not silently disable the daemon
CI / contract (push) Successful in 48s
CI / build (push) Successful in 2m1s
fleetd #360. deploy/fleetd.service shipped four mount-namespacing directives
(ProtectSystem, ProtectHome, ProtectKernelTunables, ProtectControlGroups),
PrivateTmp=true, and an ExecStart that ran java directly. Each one starts green
and breaks the daemon in a way nothing logs: lsof goes blind so every caller is
resolved ANONYMOUS and refused; the member ZDOTDIR scrub becomes a no-op; every
credential is empty. deploy/herdr.service did not exist at all, though
fleetd.service's After=/Wants= already named it.

Both units are now the ones running on fleet01, comments included -- the
bisected lsof counts and the reasons live in the files, because the next person
to 'harden' this needs the reason, not the rule.

A unit file has no compile step, so SystemdUnitSafetyTest reads both units plus
deploy/herdr-inner.sh and fails on an active forbidden directive, on
PrivateTmp=true, on an ExecStart that skips the login shell, on a herdr-inner.sh
that does not exec a login shell, and on one that does not set a non-zero pty
size. Each message names the consequence. A vacuity guard pins that all three
files exist and that the DO-NOT-add comment block still mentions every forbidden
directive, so 'comment survives, directive does not' is actually exercised.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1420, Failures: 0, Errors: 0, BUILD SUCCESS

Round 1 shipped herdr-inner.sh untested; I mutated it (dropped both the login
shell and the stty sizing) and got 6 green. Round 2 added those two checks.

Mutation run on merge, on a half the worker never touched -- ProtectHome=read-only
added to herdr.service, the unit it only ever mutated fleetd.service for:
Tests run: 8, Failures: 1, BUILD FAILURE. The parameterisation really covers
both files.
2026-09-06 20:04:48 +07:00
Dai Ha 8557289dc0 fleetd #360 review round 2: extend SystemdUnitSafetyTest to herdr-inner.sh
CI / build (pull_request) Successful in 1m27s
CI / contract (pull_request) Successful in 1m30s
herdr.service's ExecStart only names deploy/herdr-inner.sh, so that script was the only
place herdr's login-shell and pty-size properties lived, and nothing was reading it -- the
same silent-at-startup shape the ticket was about, one file further down the chain. A
mutation dropping both the login shell and the stty sizing left SystemdUnitSafetyTest green.

Add herdrInnerScriptUsesALoginShell and herdrInnerScriptSetsANonZeroPtySize, extend the
vacuity guard to cover herdr-inner.sh too, and tighten execStartUsesALoginShell's check from
a bare contains("-lc") substring match to the same login-shell-invocation regex the new
checks use.
2026-09-06 19:54:36 +07:00
Dai Ha dd2efd8541 fleetd #369: make GitWorktreesTest hermetic against the machine's real git config
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Failing after 1m28s
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus (plus every other raw git subprocess in this class) set no
isolation at all — inheriting the JVM's real environment, including the
operator's real ~/.gitconfig and default excludes file. Measured: with a
poisoned XDG_CONFIG_HOME, 56 of 59 tests failed.

Centralize every git subprocess this test starts through one factory,
gitProcessBuilder, which always applies the existing hermeticGitEnv isolation
(extended with XDG_CONFIG_HOME, the same fix #366 already applied to the
production-instance seam). Add a self-check test that counts direct
ProcessBuilder("git", ...) constructions in this file's own source and fails
if a future helper bypasses the factory, so the omission that caused this
ticket is caught by name instead of rediscovered on a poisoned machine.
2026-09-06 19:54:35 +07:00
Dai Ha 5af786d135 fleetd #368: a stale lead delegation binding must not shadow the primary fallback
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m55s
PrimaryRegistry.forgetDelegation only fires when a WORKER is released, never when
the delegating LEAD terminal itself disappears (closed, crashed, or relaunched).
A stale, non-null leadByTarget entry always beat nudgeTargetFor's single-primary
fallback, so a dead lead silently swallowed every reply nudge for its workers.

Fix: ReplyPushLoop now verifies (via the same agents.status check decide() already
uses every tick) that a recorded delegating lead is actually live before trusting
it. A dead lead is treated as if never recorded — self-healing the binding
(mirroring AgentControl.paneByTerminal's self-heal on agent_not_found) and falling
through to PrimaryRegistry's existing fallback.
2026-09-06 19:54:26 +07:00
Dai Ha 380eb63277 fleetd #360: fix systemd units that silently disabled caller identity and the credential scrub
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 2m12s
deploy/fleetd.service started clean on fleet01 but broke the daemon in three ways nothing
logs: ProtectSystem/ProtectHome/ProtectKernelTunables/ProtectControlGroups each put the unit
in its own mount namespace, which blinds fleetd's lsof-based caller lookup and falls every
caller back to ANONYMOUS; PrivateTmp=true silently no-ops the credential scrub the member
pane depends on; and running java directly from ExecStart skips the login shell that sources
the daemon's secrets, so it boots with empty credentials.

Replace the unit with the version verified working on fleet01 for a day, and add the
deploy/herdr.service companion unit it was already depending on via After=/Wants= but which
did not exist in the repo. Add deploy/herdr-inner.sh as the login-shell template
herdr.service's ExecStart wraps in a pty.

Add SystemdUnitSafetyTest (fleetd/src/test/java/dev/ltms/fleet/deploy) to read both unit
files from disk and fail if a forbidden mount-namespacing directive is active, PrivateTmp is
true, or fleetd.service's ExecStart does not go through a login shell -- the only guard
possible for a unit file with no compile step.
2026-09-06 19:45:23 +07:00
Dai Ha 6c61355f8f Merge #367: dead lead tabs stop accumulating, without destroying a live one
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m7s
fleetd #359. LeadTabScanner used to join labelled tabs straight to terminals
with no liveness check, and its javadoc excused that ("a stale name costs
nothing here"). It cost plenty: LeadCoordLoop reads that map to pick which
pane a peer message goes into, so a dead tab was a candidate it could pick.
LeadLauncher, meanwhile, had no cleanup path at all -- every reconcile that
found 0 live created another tab and left the old one.

Both now cross-check agent.list, and neither trusts a single reading of it.
That matters because the daemon's own evidence on fleet01 was agent.list
reporting 0 live while ps showed one real claude. A first cut of this fix
closed tabs on that single reading, which would have closed the operator's
live lead instead of leaving a spare tab. So: a dead tab is flagged, not
closed, and only closed when a later reconcile still finds it dead; and the
scanner grants one grace scan to a terminal it already knew was live.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1412, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge (PendingCloseMarker.strip made identity, so a flagged
tab stops matching its configured label): 4 failures, BUILD FAILURE.

Known and accepted: ensureLeads() runs at startup, so the second reading
arrives at the next restart. A tab that dies mid-session stays flagged and
open until then. Deliberate -- the bug is about repeated restarts, and one
leftover tab is cheaper than closing a live session on unverified evidence.
2026-09-05 16:00:04 +07:00
Dai Ha 96c406b968 #359 review: require two independent readings before destroying a lead's tab
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m33s
Finding 1 (LeadLauncher): closing a labelled tab on a single agent.list
miss could destroy a live lead's session — the ticket's own evidence
showed that exact signal missing a genuinely running agent. A dead
reading now only flags the tab (PendingCloseMarker); it is closed only
if a later, independently-connected reconcile still finds it dead
while flagged. A tab found live again has its flag cleared instead.

Finding 2 (LeadTabScanner): the new agent.list cross-check in scan()
was not covered by get()'s "keep the cache on a failed scan" contract,
which only fires on a thrown HerdrException. A successful-but-short
agent.list could silently drop a lead CallerResolver had already
resolved, demoting it to Role.WORKER. A terminal already reported live
now gets one grace scan before being dropped; a terminal never
reported live gets none, so the original #359 exclusion is unaffected.

Both mechanisms were mutation-tested: reverting either change turns
exactly its own new tests red and nothing else.
2026-09-05 15:54:52 +07:00
Dai Ha a5d81c3f70 Merge remote-tracking branch 'origin/main' into worker/359-dead-lead-tabs-f1253b-4 2026-09-05 15:28:49 +07:00
Dai Ha 92c0f164f1 Merge #366: seed the bridge's worker skills into every provisioned worktree
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m46s
fleetd #362 item 3. A member spawned against a repo that does not ship its
own .claude/skills/ could not load implementer, reviewer or hunter at all.
Every brief starts with "Load the <name> skill", and outside this repo that
line was silently a no-op. memberSkills: <dir> now copies those folders into
each provisioned worktree, skipping any name the target repo already ships.

Two review rounds, both about the same hazard: core.excludesFile is
single-valued, so pointing it at fleetd's own file would SHADOW the
operator's. It now composes instead of replacing, and the XDG default
excludes file is carried forward too.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1402, Failures: 0, Errors: 0, BUILD SUCCESS

Two mutations run on merge:
  drop the XDG fallback          -> 1 failure, BUILD FAILURE
  remove the composition itself  -> 2 failures, BUILD FAILURE
Both directions are pinned.
2026-09-05 13:32:16 +07:00
Dai Ha 9e813ec179 fleetd #362 review fix 2: route the XDG excludesFile fallback through gitEnv too
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m24s
Finding 1 (lead): the XDG fallback branch of previouslyEffectiveExcludesFileContent was
unpinned — deleting it left the suite green (Tests run: 1386, Failures: 0). Added
seedSkillsComposesWithTheXdgDefaultExcludesFileWhenNoneIsConfigured to pin it: isolates
XDG_CONFIG_HOME via the gitEnv seam at a temp dir carrying a synthetic git/ignore, points
GIT_CONFIG_GLOBAL at an empty file so core.excludesFile is genuinely unset (forcing the
fallback branch), seeds a skill, and asserts a file matching the XDG-default pattern still
reads as clean. Reverting the fix (mutating the fallback to resolve to "") turns this test
red with a real pasted failure (see PR body): "expected: <> but was: <?? xdg-fallback-marker>".

Finding 2 (lead, the one that actually needed a code fix): the fallback read XDG_CONFIG_HOME
and HOME straight from the JVM's own environment, not through the gitEnv seam every git
subprocess in this class already honours — so no test could isolate it, and on a machine
carrying a real ~/.config/git/ignore (this dev machine does), every seeding test silently
composed with that real file. Added resolveEnv/resolveHome, which check gitEnv first and
fall back to the JVM's real environment only when the seam doesn't supply a value (production
behaviour, where gitEnv is always Map.of(), is unchanged). Added a hermeticGitEnv() test
helper and routed every seeding test in GitWorktreesTest through it, so no test in the class
can reach the real machine's home directory for this fallback.

Also documents two non-defects the lead asked for one javadoc line each on: the composed
excludesFile is a snapshot taken at seed time, not a live reference to the operator's file;
and excludeSeededSkillsFromGitStatus assumes a fresh worktree (not idempotent, but the
double-seed path does not exist today, so no guard was added for it).
2026-09-05 13:23:39 +07:00
Dai Ha 395b3b5c46 #359: LeadTabScanner drops dead lead tabs; LeadLauncher closes them on relaunch
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m49s
LeadTabScanner.scan() joined labelled tabs to panes with no liveness check at
all, so a tab left behind by a crashed/relaunched lead read as a live lead
forever -- exactly the hazard its own javadoc predicted but excused. It now
cross-checks agent.list, the same signal LeadLauncher already trusted, and
drops any labelled tab with no agent running in it. That alone removes the
duplicate candidates LeadCoordLoop.resolveLocalLead() could pick from,
including the dead one its own WARN's advice (name a lead after
coordinator.selfId) could land a message in.

LeadLauncher never actually stopped the accumulation: relaunching on "0 live"
always created a brand new tab and left the old dead-labelled one right where
it was, so any restart that found 0 live for any reason (a real crash, or a
herdr read that missed a still-running agent) added one more dead tab,
forever. ensureLeads() now closes every dead-labelled tab for a lead as part
of the same reconcile that decides to relaunch, so at most one tab survives
per configured lead once a restart's reconcile has run.
2026-09-05 13:20:37 +07:00
Dai Ha 6e9e464d62 Merge #364: fleet_list reports lead-coordination state instead of guessing at it
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m56s
fleetd #361. fleet_list's coordination block now carries this daemon's own
mailbox state and one row per operator-declared peer, each as a tri-state
status (exists / absent / unknown) rather than a boolean. pending and
consumers appear only when status is "exists", so an unresolved probe can
never render as a measured zero.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1394, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge (isMissingQueue always returns true, which restores
the exact defect the ticket fixes): 4 failures, BUILD FAILURE. The
discriminator's false branch is pinned.
2026-09-05 13:19:46 +07:00
Dai Ha 105c065615 Merge origin/main into worker/362-worktree-skills-c03e51-3 (picks up #363) 2026-09-05 13:18:32 +07:00
Dai Ha 01492059d4 fleetd #361 review round 2: pin isMissingQueue's false branch
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m34s
The reviewer's mutation (isMissingQueue always returns true) restored
the exact overstatement fleetd #361 exists to fix -- every declare
failure reading as a confirmed absence -- and still left mvn clean
install green (1389/1389), because no test drove a non-404 shape
through inspect(). The false branch was the whole discriminator
between MailboxState.absent() and MailboxState.unknown(), unpinned.

Widened LeadMailbox.isMissingQueue from private to package-private and
added LeadMailboxIsMissingQueueTest: five hermetic tests (no broker)
covering the true case and all three false shapes isMissingQueue's own
javadoc lists -- a different reply code, a ShutdownSignalException
whose reason isn't a Channel.Close, and an IOException with no such
cause at all (plus an IOException wrapping an unrelated exception
type). Re-ran the reviewer's exact mutation locally: 4 of 5 new tests
went red with the expected assertion messages; reverted, and mvn clean
install is green again at 1394/1394 (1389 + 5 new).

Also added a one-line javadoc note on LeadMailbox.inspect being honest
about which of its two RuntimeException catches is proven by a test
(the createChannel() one, end-to-end against a real broker) and which
stays purely defensive (the declare-site one, for a connection-drops-
mid-call race no test drives on purpose).
2026-09-05 13:12:24 +07:00
Dai Ha f84824ee29 fleetd #362 review fix: compose skill-seeding excludes with the operator's own excludesFile
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m54s
core.excludesFile is single-valued, so pointing it at fleetd's own seeded-skill exclude file
with --replace-all at worktree scope was SHADOWING whatever excludesFile the worktree already
resolved (an operator's global config, most commonly) instead of adding to it. This repo's own
.gitignore does not ignore target/ — only an operator's global excludesFile does — so every
worker's `mvn clean install` would make target/ show up as untracked, and CB-576's deliberately
untracked-inclusive hasUncommitted would then read every such worktree as dirty forever, so it
is never cleaned up.

excludeSeededSkillsFromGitStatus now reads whatever core.excludesFile resolves to BEFORE writing
anything (falling back to git's own $XDG_CONFIG_HOME/git/ignore default when the key is unset
entirely, per gitignore(5)), and writes that content into fleetd's own exclude file ahead of the
seeded skill patterns, so every operator-configured pattern keeps applying inside the seeded
worktree. Proven with a new test, seedSkillsComposesWithAnAlreadyEffectiveGlobalExcludesFile,
which isolates a synthetic "operator's global config" via a new gitEnv test seam on GitWorktrees
(GIT_CONFIG_GLOBAL pointed at a throwaway temp file, never the real machine's config) and drives
the real add() path end to end.

Also documents (FleetConfig javadoc + fleetd.example.yaml) that memberSkills copies every
non-hidden subdirectory of its source wholesale, with no per-file allowlist.
2026-09-05 13:10:27 +07:00
Dai Ha c4d40fbc2b fleetd #361 review: fix false-negative absent, uncancelled probes, and a throw contract gap
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m45s
Three findings from review of #364, fixed on the same branch:

1. LeadChannel.MailboxState.absent() was returned both for a genuinely
   absent mailbox AND for "the probe could not determine anything"
   (timeout, unreachable broker, other declare failure) -- exactly the
   overstatement #361 exists to fix, one level down. MailboxState now
   carries a Presence enum (EXISTS/ABSENT/UNKNOWN) with exists()/known()
   accessors; LeadMailbox.inspect classifies a real AMQP 404 (measured
   against a live broker, not assumed: an IOException wrapping a
   ShutdownSignalException whose Channel.Close reply code is 404) as
   ABSENT and everything else as UNKNOWN. fleet_list's mailbox/peer rows
   now render a "status" of exists/absent/unknown and only include
   pending/consumers when status is "exists", so an unresolved self- or
   peer-probe can never render as a measured zero.

2. FleetMcp.probe's get(timeoutMs) left a timed-out inspect() task
   running forever on its own virtual thread, holding the AMQP channel
   it had already opened -- against a hung (not down) broker this would
   orphan one channel per fleet_list call until the connection's
   channel-max was exhausted, breaking publish() too. probe() now holds
   the Future and calls cancel(true) on timeout/failure so the orphaned
   task is interrupted instead of abandoned, and now returns
   MailboxState.unknown() (never absent()) on timeout/exception.

3. LeadMailbox.inspect only caught IOException, but createChannel() on
   an already-closed connection throws AlreadyClosedException, an
   unchecked RuntimeException (measured against a live broker) -- so it
   could escape the "never throws" contract. Both places in inspect now
   also catch RuntimeException and report unknown().

Tests: MailboxState.exists()/absent()/unknown() call sites updated
across FleetMcpTest/FleetMcpLeadCoordTest; new hermetic tests cover the
tri-state fleet_list rendering (self-probe unknown, a peer that's
absent vs. one that's unknown) and probe cancellation (a LeadChannel
fake that blocks until interrupted, proving probe() doesn't just give
up on it); new @Tag("contract") LeadMailboxTest cases pin the real
exception shapes for both the 404 and the already-closed-connection
paths and prove inspect() reports unknown (never throws) when the
connection is already closed.
2026-09-05 13:01:48 +07:00
Dai Ha 7c684e40d3 fleetd #362 (item 3): seed .claude/skills/ into provisioned worktrees
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 1m46s
Add memberSkills: <dir> to FleetConfig. GitWorktrees#add copies each
skill folder from that directory into <worktree>/.claude/skills/ so a
member spawned against ANY repo — not only one that already ships its
own skills — can load a bridge skill (e.g. implementer). A skill the
target repo already carries is never overwritten.

Every seeded path is hidden from `git status` in that worktree ONLY,
via a --worktree-scoped core.excludesFile pointing at a file under the
worktree's own private git dir (outside the working tree, so it can
never be committed) — not the shared .git/info/exclude, which a linked
worktree resolves to the repo's common git dir and would otherwise leak
visibility changes into the primary checkout and every sibling
worktree. Proven with a real `git status --porcelain` in
GitWorktreesTest, not by reasoning.

Seeding is best-effort like the existing overlayParity/isolateToolSurface
steps: a missing/unreadable source or a copy/exclude failure is logged
and skipped, never fails the spawn. memberSkills is triaged as a
DEFERRED config key in ConfigRef (baked once into GitWorktrees at
startup, like worktreeGroup), with its own changedDeferredKeys branch
and coverage-test entries.
2026-09-05 12:54:01 +07:00
ltms 6938f52155 Merge #363: make the plugin visible, and fix the drift that made it unusable
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m28s
2026-09-05 07:50:19 +02:00
Dai Ha 0df34f3220 fleetd #361: close the lead-coordination visibility gap
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m34s
Lead-to-lead AMQP coordination had a send half with tools and a receive
half without. This closes three blind spots:

- LeadChannel gains inspect(coordId) -> MailboxState(exists, pending,
  consumers), implemented in LeadMailbox with a throwaway probe channel
  (never the long-lived publish/consume channels) so a passive-declare
  404 on a missing queue can never take down publish() on the same
  instance.
- FleetConfig.Coordinator gains peers: List<String> (defaults to empty,
  blank entries dropped) so a daemon can declare which peer coord-ids
  it expects to reach.
- fleet_list reports coordination state via a new CoordinationSource
  (own coord-id, own mailbox state, held messages as msgId/from/preview
  only, and one row per configured peer with reachability/pending/
  consumers), following the existing OutageSource/QuarantineSource
  "Source record with none()" idiom instead of growing listFleet's
  overload chain by another positional parameter. Every peer probe is
  bounded by a 1.5s timeout on a virtual-thread pool and degrades to
  absent rather than ever slowing or failing fleet_list.
- fleet_send{coordId}'s success text now says "durably confirmed by the
  broker" instead of "delivered", and warns (while still reporting
  success) when the target mailbox has zero consumers attached.

Tests: hermetic unit tests for exists/absent/zero-consumer/old-config-
no-peers-key/new fleet_list shape using FakeLeadChannel, plus a
@Tag("contract") LeadMailboxTest.inspectingAMissingMailboxNeverBreaks
PublishOnTheSameInstance proving the invariant against a real broker.
2026-09-05 12:44:29 +07:00
Dai Ha 457458437f #362: make the plugin visible, and fix the drift that made it unusable
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m31s
CB-527 shipped a Claude Code plugin and a marketplace in this repo. Nothing in
CLAUDE.md or docs/ ever named it, so a later session planned the same feature
from scratch. The wiki Features entry existed and was correct, but wiki/ is a
submodule whose pointer is never advanced, so no session reads it.

Visibility:
- CLAUDE.md addendum now names plugin/ and both structural limits, so every
  session sees it. This is the change that stops the rebuild happening again.
- wiki/11-Features.md records the rename and why the entry alone was not enough.

Drift (each measured against the code, not assumed):
- mount name fleetd -> fleet, matching PeerLauncher.MCP_MOUNT_NAME. The old name
  gave a lead with both a project .mcp.json and the plugin two mounts of one
  daemon and a duplicated fleet_* tool set.
- url is now ${FLEETD_MCP_URL} instead of a hardcoded address, so one plugin can
  serve hosts running the daemon on different ports. Plain ${VAR}, the form
  kb-alms proves works here; ${VAR:-default} is untested and not used.
- plugin claude-bridge -> fleet, marketplace claude-bridge -> fleetd, version
  0.2.0. Breaking for a 0.1.0 install: mcp__fleetd__* becomes mcp__fleet__*.
- README install path ltms/claude-bridge -> the fleet/fleetd remote.
- the setup skill's §5 told operators to pin primary.terminal:. CB-579 replaced
  that with fleet.leaders.*.tab. Replaced, with the duplicate-tab warning (#359).

Scope: the plugin is lead-side only, and cannot be otherwise. The launcher adds
--agent only when <worktree>/.claude/agents/<role>.md exists in the member's own
tree (ClaudeCodeLauncher.java:371,391), and a member's CLAUDE_CONFIG_DIR points
at its profile's config dir (ClaudeCodeLauncher.java:285), so a member never
reads the operator's plugin store. On this Mac all four Claude profiles set
configDir, and the four ccs instances hold four separate copies of the plugin
store -- same md5, different inodes. Seeding member skills through the worktree
is #362 scope item 3, implemented separately.

Note for anyone verifying a plugin: `claude plugin validate` does NOT read
.mcp.json. Replacing it with `{ this is not json at all` still passes, exit 0.

Refs #362, #359
2026-09-05 12:42:20 +07:00
Dai Ha 3759c41f99 Merge #354: the redeploy health gate classifies AMQP errors instead of counting them
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m36s
The gate counted ERROR lines since RESTART_MARK. On a laptop that idle-sleeps after one
minute on battery that meant 6 ERROR lines for an AMQP link that recovered every time,
and a gate that cries wolf is a gate nobody reads.

It now reports three states: no errors; only errors proven to have recovered (quiet, and
the gate passes); anything else (the old warning, unchanged). Attribution is per
connection, using the names #356 put into the log -- a lead-mailbox recovery can no
longer clear an unrecovered reply-inbox reset. A candidate carrying neither name is
unattributable and stays LOUD.

Two earlier rounds were rejected. Round 1 was inert: it matched nothing in the real log,
because the layout abbreviates the logger and 'Connection reset' sits in the stack trace,
not on the ERROR line -- my brief had pointed the worker at fleetd.out, which is untracked
and so absent from its worktree. Round 2 was correct and honest but could not attribute
anything, which is what motivated #356.

Verified on merge beyond the worker's own mutations:
 - ran the classifier against the REAL log, which is still in the pre-#356 format: 6 total,
   0 recovered, 6 unexplained. Old-format lines carry no connection name, so they stay loud
   -- the safe direction, on genuine data rather than a fixture.
 - adversarial fixture the worker did not write: a lead recovery BEFORE any failure banks
   no credit; 2 inbox resets with 1 recovery leaves 1 unexplained; a non-AMQP ERROR stays
   loud. total=3 recovered=1 unexplained=2, as intended.
 - RESTART_MARK still anchors the scanned region.

Caveat carried from the PR: the patterns are source-derived. The daemon has not been
redeployed, so they are not yet confirmed against a live log.
2026-09-05 06:09:29 +07:00
Dai Ha 09159f2857 Classify named AMQP recovery errors
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m16s
2026-09-05 06:05:52 +07:00
Dai Ha 29cd1194c2 Merge remote-tracking branch 'origin/main' into worker/errscan-bed2ca-2 2026-09-05 06:02:29 +07:00
Dai Ha 815e8f8b23 Merge #356: name the AMQP connection in its own log lines
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m38s
Both connections were already named at newConnection() -- 'fleetd-reply-inbox' and
'fleetd-lead-mailbox' -- and neither name ever reached the log: 0 occurrences in
fleetd.out, and both connections logged under the same thread name
'[AMQP Connection 10.10.20.13:5672]'. So when one of the two died and never came
back, the log could not say which.

AmqpConnectionFailureLogger extends DefaultExceptionHandler and overrides only the
protected log(String, Throwable) sink that every handle* method calls virtually, so
the identity is added without changing any handler action.

My brief caused a defect here and the correction is the interesting part. I told the
worker the client 'currently uses ForgivingExceptionHandler', read off the log line
c.r.c.i.ForgivingExceptionHandler -- which names where the LOGGER FIELD is declared,
not the instance's class. javap on the jar shows ConnectionFactory's constructor does
'new DefaultExceptionHandler', and DefaultExceptionHandler extends StrictExceptionHandler
extends ForgivingExceptionHandler. The first version therefore extended the base and
silently dropped strict channel-closing on four listener/consumer paths. Now pinned by
a type assertion on both factories plus a behavioural test that handleConsumerException
still closes the channel once.

Verified on merge with a mutation the worker did not run: it mutated the parent class,
so I mutated the copied private-static isSocketClosedOrConnectionReset in the DANGEROUS
direction (always true => every failure logs at WARN and vanishes from the redeploy
gate's ERROR count). Caught: 'inbox failure line ==> expected: <ERROR> but was: <WARN>'.

Merged main in first; the auto-merge compiled. 1379 green, unpiped.
2026-09-05 06:01:30 +07:00
Dai Ha 1e60ac0745 merge main for verification 2026-09-05 05:59:22 +07:00
Dai Ha 650a4c146b Merge #357: a FleetConfig component dropped by withDefaults() now fails the build
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m44s
Test-only. FleetConfig.java itself is unchanged.

The hazard is the back-compat constructor ladder (21/20/18/17/16/15/14 alongside the
22-arg canonical). Add a component and leave withDefaults()'s call at the old arity and
it binds to a back-compat constructor: it compiles, the suite passes, and the new key is
silently defaulted away on every load().

Verified on merge with a mutation the worker did not run: I made withDefaults() issue a
21-arg call, reproducing the real binding rather than an explicit null. It compiled, and
the guard failed by name -- 'memberLoginShell: ... a component silently dropped by
withDefaults(), the shape of the defect this test exists to catch'.

Exclusion list is empty and its size is pinned, so a future exemption must touch an
assertion rather than grow quietly.
2026-09-05 05:55:22 +07:00
Dai Ha 23f299e105 Preserve strict AMQP exception handling
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m46s
2026-09-05 05:53:44 +07:00
Dai Ha dbf6fef0e9 config: guard FleetConfig.withDefaults() against silently dropping a component
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m53s
Adding a component to FleetConfig follows an established pattern: the
record grows by one arg, and a back-compat constructor is added at the
OLD arity so existing callers keep compiling. That back-compat
constructor also silently captures withDefaults()'s own literal-arity
'return new FleetConfig(...)' call the next time this happens, since
that call is now a legal overload match too. It compiles, every other
test passes, and the new component is defaulted away on every load().
This is not hypothetical - it happened live while building the (now
parked) idle-sleep-guard PR, caught only because that branch's own new
tests asserted on the new field.

Add a reflective test that builds a FleetConfig through the true
canonical constructor (resolved by record-component types, not arg
count - the same pattern ConfigRefTopLevelReportingCoverageTest already
uses in this file) with a real, non-null value in every component, runs
the real withDefaults(), and asserts every value survives unchanged.
Never hardcodes the arity - it enumerates
FleetConfig.class.getRecordComponents() - so it keeps working as the
record grows. No back-compat constructor is touched or removed.
2026-09-05 05:51:59 +07:00
Dai Ha d292522d00 Name AMQP connection failure logs
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 1m29s
2026-09-05 05:45:37 +07:00
Dai Ha 0241e0d3a8 Keep unattributed AMQP errors loud
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m30s
2026-09-05 05:36:53 +07:00
Dai Ha e4973eb8a4 Classify recovered AMQP redeploy errors
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m51s
2026-09-05 05:29:53 +07:00
Dai Ha b6b88c5f1c #334: pin the fresh-owner gate on ask()'s timeout teardown
CI / build (push) Successful in 2m7s
CI / contract (push) Successful in 34m37s
2026-09-04 17:12:49 +07:00
Dai Ha 86dddfe240 Merge #334: ask()'s timeout closes the turn before it forgets the task mapping 2026-09-04 17:08:40 +07:00
Dai Ha 0d5944af63 fleetd #334: close ask()'s turn before forgetting its Task, closing the last stranding window
CI / build (pull_request) Successful in 2m14s
CI / contract (pull_request) Successful in 19m14s
ask()'s TimeoutException catch used to run clearAsyncQuestion(turnId, true) -- forgetting the
Task's asyncTasksByTurn mapping -- before rendezvous.closeAsk(turnId) ran in the shared finally.
Between those two calls the ask was still "answerable" (askSession(turnId) non-null) but the Task
mapping was already gone, so a racing answer() call found task == null, skipped
finishAsyncTask, and stranded the async ticket at PENDING even though answer() itself reported a
result. #329 fixed one step of this same race; this closes the remaining one.

The fix reorders the fresh owner's teardown: closeAsk runs first, then markAskTimedOut and
clearAsyncQuestion. A racing answer() call now either sees the ask still open (and the Task
mapping guaranteed intact) or sees it already closed (STALE_TURN, before it ever reaches
asyncTasksByTurn). It also gates the whole block by ticket.fresh(), matching the invariant the
finally block already states ("only the fresh owner tears down the shared turn") -- a duplicate
coalesced ask() timing out no longer forgets bookkeeping the fresh owner still needs.

Adds aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket, which pins the exact window
with a new test-only hook (askTimeoutRaceHookForTest) and proves both invariants: a late answer()
racing the timeout sees STALE_TURN, and the async ticket still resolves DONE from the worker's
real reply. Reverting the reorder (verified locally, not committed) makes this test fail with
"expected STALE_TURN but was TIMED_OUT_WORKING".
2026-09-04 17:06:15 +07:00
Dai Ha 4a5030a5c6 #348: drop a chrome skip that cannot fire, and pin the live pattern shape
CI / contract (push) Successful in 2m10s
CI / build (push) Successful in 2m12s
2026-09-04 16:57:36 +07:00
Dai Ha c1c8794c48 Merge #348: a member's prose about a usage limit no longer quarantines a credential 2026-09-04 16:53:19 +07:00
Dai Ha 65a78932c1 Merge #335: a per-task cleanup throw in abandon() no longer strands the tasks behind it
CI / contract (push) Successful in 1m24s
CI / build (push) Successful in 2m8s
2026-09-04 16:44:04 +07:00
Dai Ha f429ca1a50 Avoid exhaustion cooldown for member prose 2026-09-04 16:43:16 +07:00
Dai Ha 73aab3f83e Merge #342: teardown resolves the pane's real tab instead of trusting the delegate's placement config
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m47s
2026-09-04 16:38:31 +07:00
Dai Ha 3fd23ecafa fleetd #342: base tab-cleanup teardown on the pane's real placement, not a delegate's static config
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Failing after 1m38s
HerdrPeerLauncher.stop() used to gate spaces.locatePane() on usesTabPlacement(),
which reads the delegate's OWN configured profiles. When CompositePeerLauncher's
single-daemon stop() shortcut hands a pane to a delegate that never spawned it
(spawnedBy empty after a daemon restart, herdrDaemonCount()==1), that delegate's
placement config says nothing true about how the pane was actually placed, and a
dedicated tab could be skipped and leaked.

Resolve the tab unconditionally instead — WorkspaceControl#locatePane already
tolerates a missing pane by returning null — and let the existing single-occupant
check (tabPaneCount()==1) be the only thing that decides whether to close it, same
as it already protects a shared tab regardless of declared placement.

Adds a mixed-placement CompositePeerLauncherTest (every existing stop-fallback test
configured both adapters as tab placement, so the mis-routing never showed) and
updates FleetAppTest#stopWorkerInPanePlacementClosesOnlyThePane, whose old
assertion (no pane.get on pane placement) documented exactly the skip this fix
removes.
2026-09-04 16:36:11 +07:00
Dai Ha f379847942 Merge #345: the timeout path's use of Cancellation.DELIVERED is now pinned
CI / contract (push) Successful in 47s
CI / build (push) Successful in 2m8s
2026-09-04 16:32:45 +07:00
Dai Ha ea12107497 fleetd #345: test timeout cancellation race
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 2m3s
2026-09-04 16:30:46 +07:00
145 changed files with 20720 additions and 1030 deletions
+4 -4
View File
@@ -1,15 +1,15 @@
{
"name": "claude-bridge",
"name": "fleetd",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the fleetd MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "claude-bridge",
"name": "fleet",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"description": "Mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.2.0",
"author": {
"name": "LTMS"
}
+226
View File
@@ -0,0 +1,226 @@
---
name: handover
description: Procedure for an outgoing lead to write the handover file that a fresh lead session inherits. Load this when your context is filling up and you are about to be replaced, whether you hand off by hand or fleetd does it for you. The file is the new lead's only inheritance — follow it exactly.
---
# Handover — write the file the next lead depends on
A lead session fills up its context and has to be replaced by a fresh one. The outgoing lead
writes a handover file, and the new session reads that file and carries on.
**There are two ways to hand off, and the file is the same either way.**
- **By hand.** You write the file, then tell the operator where it is. The operator starts the new
session and points it at the file. This always works.
- **With `fleet_handover`** (fleetd #480, merged 2026-09-11). You ask fleetd to do the swap: it
checks the file, clears your pane, and tells the fresh session to read it. This needs
`leadRollover:` in `fleetd.yaml`; without it every action answers a clean refusal naming
`NOT_CONFIGURED`, and you fall back to the manual path. Section 11 below is the procedure.
Nothing else in this skill changes between the two. Only who performs the swap changes.
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
through at the end of a session.
This skill is the procedure for writing it. Every rule below earned its place because a past
handover got it wrong.
## 1. Confirm you are the right session to write this
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
## 2. Every number needs a command, run in this turn
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
you write one, run the command that produces it — now, in this turn, against the live state.
Never take a number from:
- earlier in your own conversation — the state has moved since then,
- a peer lead's report — that is their measurement, not yours,
- your own memory of an earlier session.
Put the command, or its real output, next to the number. That lets the next lead re-run it and
check it still matches. If you cannot measure something yourself, say so instead of guessing:
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
## 3. Say what you measured and what you did not
Mark every claim as one of two things:
- **"I checked this myself, in the code or on this host, at `<time>`."**
- **"I did not check this myself; `<who>` reported it."**
Never present someone else's measurement as your own. This matters most for cross-host claims —
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
re-verify from here.
## 4. Record open decisions, and who owns them
List three things:
- what the operator actually asked for, in their own words where you have them,
- what is still unanswered,
- any question you decided yourself instead of asking, with your reason.
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
reopens a closed question or assumes the operator chose something they never did.
## 5. Record live hazards
List anything that will break if the next lead does the obvious thing next. This includes:
- unpushed commits or unmerged branches,
- a build, a spawn, or a redeploy still running,
- code merged to `main` but not yet redeployed to the live daemon,
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
"tricky."
## 6. Record what is explicitly not owed
List work that is finished, and work that another party has said they do not want touched. Name
who said so and when. Without this line, the next lead re-does closed work or reopens a question
a peer already declined to revisit.
## 7. Open the file with three re-measurement commands
The file's own first section must give the next lead three concrete commands to run before
acting on anything else in the file:
1. confirm role — for example `fleet_whoami`,
2. confirm the state of the working tree — for example `git status` and
`git rev-list origin/main..HEAD`,
3. read the live fleet — for example `fleet_list`.
Record what each command answered when you wrote the file, and tell the reader to run it again
rather than trust your answer. The point of this section is that the reader checks live state
before acting on any claim in the rest of the file, including yours.
## 8. Stamp the file with time and commit
Near the top of the file, write:
- the date and time you wrote it,
- the commit the tree was on (`git rev-parse HEAD`),
- whether the tree was clean (`git status`).
Without this, nobody can tell how old the file is, or which code it describes.
## 9. State plainly that the file goes stale fast
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
of one moment, not a live fact.
## 10. What to leave out
Do not include:
- narration of how the session felt, or how hard something was,
- anything the repo already records — code structure, git history, or a rule already written in
`CLAUDE.md`. Point at it instead of repeating it,
- advice that is only true for the session that is ending — a half-open terminal, a local
variable, a train of thought with no state behind it.
A handover file is a record of state and decisions. It is not a diary.
## 11. Using `fleet_handover` (only if `leadRollover:` is configured)
**Run the three steps in this order. The order is not a style choice — the wrong order is
refused.**
1. **`fleet_handover{action: "open", reason: "<why now>"}`.** It returns a `token` and the
`handoverPath` you must write to. Nothing has happened to your pane yet.
**Write to exactly that path, and do not resolve it yourself.** It is always absolute, even when
the operator configured a relative `handoverPath`: fleetd resolves a relative one against your
own workspace before it hands it to you. The daemon and your pane can run in different
directories, so a path you resolve yourself can point at a different file from the one the daemon
will check.
**Check that the path is ignored by git before you write to it (#491).** A relative
`handoverPath` resolves inside YOUR workspace, which is usually a repository — and usually not
the `fleetd` one, so an ignore rule added to `fleetd` does not protect it. Run
`grep -n handover <your workspace>/.gitignore`. No output means the file you are about to write
will show up as untracked content in that repo. The file is a snapshot of live state and must
never be committed, so tell the operator rather than committing it or silently editing their
`.gitignore`.
2. **Write the handover file at that path**, following sections 1–10 above.
3. **Ask the operator, then `fleet_handover{action: "confirm", token, operatorConfirmed: true}`.**
Why that order: `confirm` refuses with `HANDOVER_STALE` unless the file was modified **after** the
`open` request. That check stops a leftover file from an earlier session being accepted as this
one's handover. So writing the file first and then calling `open` — the obvious order — always
fails.
`{action: "cancel", token}` drops a pending request without rolling.
**Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
get another one.
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
from your connection, so you can only ever roll yourself.
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
defaults to `true` and this is the only thing standing between a judgement call and a wiped
session.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
the operator so before you confirm.** The recovery is the same either way: the file is already
written, so the operator starts a session and points it at the file. That is why you write the
file before you confirm, and never the other way round.
## Writing style
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
class, method, file, flag, and config key exactly as it appears in the code — replacing a
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
first time you use it.
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
brackets, colons, or slashes.
## Template
```markdown
# Handover — <fleet name> lead session, <date and time>
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
spawns, or restarts.
## 0. Do these three things first
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
## 1. What the operator asked for
<the live instructions, in their words where you have them; what is still open; any decision
you made yourself, and why>
## 2. Open decisions, and who owns them
<one line per decision: who owns it, what is unanswered>
## 3. Live hazards
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
## 4. What is not owed
<finished work, and work another party has declined; name who said so and when>
```
+72
View File
@@ -0,0 +1,72 @@
---
name: redeploy-fleetd
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
---
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
+10 -4
View File
@@ -87,12 +87,18 @@ jobs:
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
# The `contract` profile clears the default-excludes group, so `-Dgroups=contract` runs every
# @Tag("contract") test and nothing from the unit suite the `build` job already covered — a
# tag selects the whole group, so a test added to it later runs here automatically. A prior
# version of this step pinned `-Dtest=AmqpReplyInboxContractTest` by class name instead: that
# silently excluded every other contract test (including the herdr ones) from CI, and nobody
# noticed until the herdr protocol drifted out from under a test that never ran here
# (fleetd #449). If this runner has no herdr socket, the herdr-backed tests in the group
# skip on their own `assumeTrue` and only the broker-backed ones actually run — check the
# step output rather than assuming which.
- name: Contract tests
working-directory: fleetd
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
run: mvn -B -Pcontract test -Dgroups=contract
- name: Failing test output
if: failure()
+6
View File
@@ -19,3 +19,9 @@
fleetd.out
fleetd/fleetd.out
logs/
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
# has no business in git history.
.handover/
+91 -63
View File
@@ -7,6 +7,14 @@
> wiki ([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
>
> **Anything you measure in an addendum is perishable.** Date it, give the command that
> re-measures it and what each outcome means, and tell the reader to delete the section once
> it stops reproducing. The four parts work together: deciding what would falsify a claim is
> the expensive step, and a reader in the middle of another task will not pay it, so a bare
> "verify before relying on this" costs the same space and does nothing. The case this is for
> is a note that goes stale as a live restriction — it will tell a future session it cannot do
> the thing at the moment doing it becomes the job.
If no `fleet_*` MCP tools are mounted in this session, this section does not apply — skip it.
@@ -50,8 +58,9 @@ and the sender silently receives nothing. Fail toward the recoverable error.
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
### Primary (lead) — run this on every task, in order
@@ -71,8 +80,11 @@ below are the procedure — run them in order, every task, not only the big ones
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `fleet_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model, cost and
LIVENESS, not in tier, so the default is rarely what you want. The default is whatever the
daemon reports, and on a host where it sits on an exhausted or withdrawn credential every
unqualified spawn fails — sometimes loudly, sometimes as a member that spawns fine and then
produces nothing. `fleet_profiles` reports the default; check it once per session.
4. **Then send them all** — `fleet_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
@@ -94,6 +106,18 @@ below are the procedure — run them in order, every task, not only the big ones
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `fleet_stop{paneId}`.
**If the forge refuses you the merge** — a protected branch, a token without the grant — the
adjudication is still yours. Read the diff, decide, and hand the operator a merge-ready queue
with the refusal quoted. Never report a PR as merged, and never call one "ready to merge"
without having read the diff yourself. A refusal is exactly when that shortcut is tempting,
because no action is left that forces you to look, and taking it turns this step into
forwarding a reviewer's verdict — which is delegating the merge by proxy, two lines above.
**Test a refusal; do not read it off a permissions field.** A protected branch holds its merge
rights separately from the repository permissions, so that field can say yes while the merge is
refused, and still say no after a grant makes it work. Probe instead, with a request that cannot
succeed on its merits, so a rejection can only mean the refusal. Treat a transport failure as a
third answer that proves nothing: a timeout, a DNS error or a bad URL is not a refusal, and
counting it as one makes you sure of something you never measured.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
@@ -114,9 +138,11 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_reply{content}` — the one case a lead replies |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after* `open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
### Lead ↔ lead — coordinate, never delegate
@@ -140,10 +166,21 @@ The traffic between leads is coordination and nothing else:
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
**N observations are N data points only if they differ in the axis you are trusting.** This cuts
both ways. N *failures* blamed on one cause are one data point when the cases share what you are
not varying. N *agreeing measurements* are also one data point when they share an instrument —
two hosts, two operators and the same formula is one formula, not two confirmations.
4. **Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
author is the worst reader of their own qualifier placement — measured here, one addendum carried
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
sentence goes false first, and would a reader reach the caveat before acting?"
Being messaged by a peer does not make you its worker: answer with `fleet_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
Being messaged by a peer does not make you its worker: answer the way you would open —
`fleet_send{coordId}` for another daemon, `fleet_send{sessionId}` on this host — and push back on
the substance if it is wrong. `fleet_reply` resolves a member's blocked `fleet_send`; a peer's
coord-id message is durable and non-blocking, so there is nothing for it to resolve. A peer that
simply complies has thrown away the reason there are two of you.
### Member (worker or architect) — the turn contract
@@ -188,6 +225,19 @@ must obey belongs in the charter, not here.
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Herdr socket tests (measured 2026-09-10).** In this repo, herdr is a subject under test. A
worker assigned to herdr code, and the lead, may let a test open the herdr socket directly in a
throwaway workspace that the test tears down. This only covers
`fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java`, and
`fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java`. It is not a
general licence. Using the herdr CLI or socket to move a real fleet session, pane, or peer stays
banned. That is the control plane that invariant 5 protects. Re-measure with
`grep -rl 'UnixSocketHerdrClient.connect()' fleetd/src/test/java --include='*.java'`. A non-empty
result means tests still open the socket and this note still applies. An empty result means nobody
does this any more; delete this section. Canonical invariant 5 restatement is tracked in #458 and
is not part of this change.
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
@@ -208,8 +258,22 @@ must obey belongs in the charter, not here.
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
`fleets-status` (report every fleet that shares one LavinMQ instance).
`port-to-opencode` (make an OpenCode session a participant in this workspace),
`fleets-status` (report every fleet that shares one LavinMQ instance),
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
fleetd #480).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
unmentioned by every instruction file, so it drifted and a later session planned it from scratch
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
worktree; and a plugin cannot deliver anything to members at all, because
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it,
so a member never reads the operator's plugin store. **The plugin is the lead-side surface;
member-facing assets travel in the worktree.**
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
@@ -232,60 +296,14 @@ must obey belongs in the charter, not here.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
worker has no business restarting the daemon it is talking through, and stopping it kills the
worker's own channel mid-turn.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
and its flags, the drain step, the operator's permission grant, and the five checks that have each
gone wrong here before. Do not hand-roll the steps from memory.
### The prompt is part of the product — update it with the code (mandatory)
@@ -333,6 +351,16 @@ print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
PY
```
**Only the lead can run that check (measured 2026-09-10).** A member's provisioned worktree has
`wiki/` uninitialized, so the script dies with `FileNotFoundError: wiki/7-Use-Cases.md`. Measured
in three worker worktrees: `git submodule status` printed a leading `-` and `wiki/` held 0
entries; the primary's own clone printed a leading `+` and the file was there. So never make this
check a member's acceptance criterion — it is unsatisfiable for them, and a brief that asks for it
is asking a worker to invent a pass. A member told to check it must say it could not run it, and
must never report it as passed. The lead runs it in the main clone before merging. Re-measure with
`git submodule status` in a member's worktree: a leading `-` means this still applies; once it
prints a commit with no `-`, delete this paragraph.
## IDE MCP tools & validation workflow (enforced)
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
+54 -46
View File
@@ -1,72 +1,80 @@
# CB-504 — systemd unit for fleetd (Linux).
#
# The macOS launchd agent (deploy/dev.ltms.fleetd.plist) is the supervision target for the
# current single-host deployment. This unit exists for the per-host gateways CB-308 introduces,
# which will run on Linux.
# fleetd #360: the previous version of this file started clean and broke the daemon in three ways
# that nothing logs (see the DO NOT block and the ExecStart/PrivateTmp comments below for what and
# why). The unit below, plus its companion deploy/herdr.service, is the version that has actually
# run on fleet01 without those failures. Do not "improve" it back toward the old shape without
# re-reading why each line is the way it is.
#
# Install (user service — fleetd drives the user's herdr, not a system daemon):
# mkdir -p ~/.config/systemd/user
# cp deploy/fleetd.service ~/.config/systemd/user/
# # edit ExecStart / WorkingDirectory / Environment below, then:
# cp deploy/fleetd.service deploy/herdr.service ~/.config/systemd/user/
# # edit WorkingDirectory / ExecStart below for your host's paths and java location
# systemctl --user daemon-reload
# systemctl --user enable --now fleetd
# systemctl --user enable --now herdr fleetd
# loginctl enable-linger $USER # REQUIRED -- see below
# journalctl --user -u fleetd -f
#
# `loginctl enable-linger` is not optional and is easy to miss, because leaving it out looks like
# success: `systemctl --user enable` reports "enabled" and both units run for as long as you stay
# logged in. A user manager without lingering starts at your first login and stops at your last
# logout, so the fleet simply does not come back after a reboot -- which is the whole reason to
# use systemd here rather than the setsid scripts these units replaced. Check it with
# `loginctl show-user $USER -p Linger`; the answer must be `Linger=yes`.
#
# Secrets (AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI, ...) are not set
# here and need no systemd drop-in: ExecStart runs a login shell, so they come from wherever your
# login shell already sources them (this host: ~/.fleet/secrets.sh via ~/.zprofile). If a token is
# missing there, fleetd still starts — the daemon reports every secret a configured profile
# references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)"; a missing one logs
# "startup secret NAME: MISSING" at WARN and the daemon starts anyway — the first visible symptom
# is a member that cannot open a pull request, hours later and in a different component.
[Unit]
Description=fleetd — claude-bridge message server
Description=fleetd — fleet message server
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
# Ordering only: herdr is a user process and its socket may appear after us. This is advisory —
# fleetd retries the herdr socket rather than exiting, which is what actually makes a late
# socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down with it.
# Ordering only. fleetd retries the herdr socket rather than exiting, which is what actually makes
# a late socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down too.
After=herdr.service
Wants=herdr.service
[Service]
Type=simple
WorkingDirectory=%h/src/claude-bridge/fleetd
ExecStart=/usr/lib/jvm/temurin-25-jdk/bin/java -jar target/fleetd.jar fleetd.yaml
WorkingDirectory=%h/LTMS/fleetd/fleetd
Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
# PATH matters more than it looks (CB-511): fleetd propagates its own PATH to every worker it
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
# Secrets are NOT set here — this file is committed. Put ALL three tokens in a private drop-in
# that systemd reads with restrictive permissions. In `systemctl --user edit fleetd`, add:
# [Service]
# Environment=FLEETD_API_TOKEN=...
# Environment=WORKER_GITEA_TOKEN=...
# Environment=AI_GATEWAY_TOKEN=...
# FLEETD_API_TOKEN protects fleetd's API. WORKER_GITEA_TOKEN lets members open pull requests; if
# it is missing, fleetd still starts, but a member fails when it later tries to open a pull request.
# AI_GATEWAY_TOKEN authenticates gateway profiles; if it is missing, fleetd still starts, but a
# gateway profile later returns HTTP 401. Or, put the same three variables in a 0600 file and add:
# EnvironmentFile=%h/.config/fleetd/env
# After starting, check which of them actually resolved. The daemon reports every secret a
# configured profile references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)". A missing one logs
# "startup secret NAME: MISSING" at WARN — and the daemon starts anyway, which is the whole
# problem: without this grep the first sign is a member that cannot open a pull request, hours
# later and in a different component.
# Note what the report can and cannot tell you. It lists only names some profile actually
# references (tokenEnv, gitTokenEnv, and the broker uriEnv). A secret nothing references is never
# reported, because nothing needs it.
# A LOGIN shell, not java directly. Every secret this daemon needs (AI_GATEWAY_TOKEN,
# WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in ~/.fleet/secrets.sh, which only
# ~/.zprofile sources. systemd runs no login shell. Started any other way the daemon boots fine
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
# unit) has to read them. A private /tmp turns the credential scrub into a silent no-op.
PrivateTmp=false
Restart=on-failure
RestartSec=10s
# A bad config (e.g. a non-loopback bind without token auth) makes fleetd fail fast by design.
# Give up rather than restart-loop on a permanent error.
# A bad config makes fleetd fail fast by design. Give up rather than restart-loop forever.
StartLimitBurst=5
StartLimitIntervalSec=120
# The daemon reads the repo, writes worktrees, and talks to a Unix socket — it needs no more.
# DO NOT add ProtectSystem=, ProtectHome=, ProtectKernelTunables= or ProtectControlGroups=.
# Measured on fleet01 2026-09-05: each of those gives the unit its own mount namespace, and
# fleetd resolves a caller role by running lsof to find the loopback peer PID
# (mcp/LsofPeerPidLookup). Inside such a namespace lsof returns nothing, every caller falls back
# to ANONYMOUS, and the primary is refused every orchestration call with
# "unauthenticated: anonymous may not SPAWN".
# The daemon still starts, healthz still returns ok and the secrets still resolve - the only
# symptom is that the fleet cannot be driven at all. Verified by bisecting the directives:
# no sandbox 3 lsof lines | ProtectSystem=strict 0 | ProtectHome=read-only 0
# ProtectKernelTunables 0 | ProtectControlGroups 0 | RestrictSUIDSGID 3 | NoNewPrivileges 3
# The two below add no mount namespace and are safe.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-write
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true
StandardOutput=journal
+19
View File
@@ -0,0 +1,19 @@
#!/bin/zsh
# fleetd #360 — template for the script deploy/herdr.service's ExecStart wraps in a pty.
#
# `script -qfec <this> /dev/null` needs a real command to run, and that command has to be a LOGIN
# shell script: herdr itself needs the same secrets fleetd.service's login shell picks up (this
# host: ~/.fleet/secrets.sh via ~/.zprofile), because members it spawns inherit its environment.
# systemd's own Environment= lines in herdr.service are not enough for that -- they set TERM and a
# bare PATH so the pty starts at all, nothing more.
#
# Copy this file to the path deploy/herdr.service's ExecStart names
# (%h/LTMS/fleetd/fleetd-run/herdr-inner.sh by default) and `chmod +x` it. Not committed under
# that path itself because the session name below is host-specific.
# A 0x0 pty makes every pane spawn fail with "ghostty error -2" (see herdr-multi-instance-facts /
# fleet01-headless-herdr-standup) -- give it a real size before herdr ever touches it.
stty rows 50 cols 200
# -l: login shell, so herdr and everything it spawns gets the real secrets and PATH.
exec zsh -lc 'exec herdr --session <name>'
+40
View File
@@ -0,0 +1,40 @@
# fleetd #360 — systemd unit for herdr (Linux), the terminal multiplexer fleetd drives.
#
# This is fleetd.service's companion: fleetd.service's After=/Wants=herdr.service assumes this
# unit exists. Before this ticket it did not, so on a fresh host fleetd started against a herdr
# that systemd never supervised at all.
#
# Install: see deploy/fleetd.service's header comment (both units install the same way).
#
# ExecStart below runs deploy/herdr-inner.sh (copy the template of that name from this directory
# to the path in ExecStart, or point ExecStart at wherever you keep it, and make it executable).
# It is a separate file rather than an inline command because it must itself be a login shell (see
# its own header for why) and systemd's ExecStart does not run one.
[Unit]
Description=herdr terminal multiplexer (fleet session)
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
[Service]
Type=simple
# script(1) gives herdr a real pty. Without it the client reports a 0x0 window and every pane
# spawn fails with "ghostty error -2" -- which surfaces as a fleetd spawn failure, not a herdr one.
ExecStart=/usr/bin/script -qfec %h/LTMS/fleetd/fleetd-run/herdr-inner.sh /dev/null
StandardInput=null
Environment=TERM=xterm-256color
Environment=PATH=%h/.local/bin:/usr/local/bin:/usr/bin:/bin
# PrivateTmp MUST stay false. fleetd writes the member ZDOTDIR scrub dir and the opencode config
# dir under its own java.io.tmpdir, and the member pane -- a child of THIS process -- has to read
# them. A private /tmp here silently breaks the credential scrub instead of failing loudly.
PrivateTmp=false
Restart=on-failure
RestartSec=5s
StandardOutput=journal
StandardError=journal
SyslogIdentifier=herdr
[Install]
WantedBy=default.target
+1 -1
View File
@@ -192,7 +192,7 @@ Deliberately small; every one maps to a failure mode we have actually hit.
| `fleet_send_duration_seconds` | histogram | delegated turn latency |
| `fleet_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
| `fleet_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ delivered\|exhausted — a rising `exhausted` means the primary is not draining |
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ sent\|exhausted — a rising `exhausted` means the primary is not draining. `sent` was called `delivered` until fleetd #365; it counts the herdr paste-and-submit call returning, never a confirmation the pane read it |
| `fleet_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
| `fleet_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
| `fleet_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
+133 -11
View File
@@ -88,6 +88,47 @@ bind:
# backoffMs: 60000
# quietNudgeCap: 3
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
# own pane and bootstrap a fresh session against that file.
#
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
#
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
# reads must live. No default (an operator-specific path); a present block with no
# handoverPath refuses to start. May be relative: it then resolves against the
# CALLING lead's own fleet.leaders.<name>.cwd (falling back to the daemon's own
# working directory when that lead has none configured) — never against whatever
# directory the daemon process happens to have been started in. An absolute path is
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
# share a working directory (fleetd #480 follow-up).
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
# # operatorConfirmed: true
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
# # lead's own turn to end (its pane to report injectable again)
# # before sending /clear at all. If this elapses, /clear is NEVER
# # sent — a lead that never goes idle is still doing real work.
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
# # again AFTER /clear before giving up (never sends bootstrapText
# # if this elapses). A separate, second wait from turnSettleSeconds.
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
# # the lead once its pane settles after /clear
# leadRollover:
# handoverPath: /path/to/handover.md
# requireOperatorConfirm: true
# maxDocAgeSeconds: 3600
# turnSettleSeconds: 20
# clearSettleSeconds: 20
# bootstrapText: "Fresh lead session: read the handover file and carry on."
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
# per tick.
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
@@ -110,6 +151,14 @@ bind:
# notifications:
# mode: disabled
# Idle-sleep guard: while at least one member is live, hold an OS-level assertion against idle
# sleep (macOS only — a `caffeinate -i` child; a no-op elsewhere or if caffeinate is missing), so
# an unattended host does not idle-sleep out from under a member's long turn. Unlike health/
# configReload above, this is ON BY DEFAULT — omitting the block entirely leaves it enabled, the
# same as `enabled: true`. Uncomment only to turn it off:
# idleSleepGuard:
# enabled: false
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -199,8 +248,12 @@ herdrSocket: ~/.config/herdr/herdr.sock
# omit and this profile's completion fallback behaves exactly as before.
# Every backend words its refusal differently, so this is config, never a
# vendor string baked into fleetd itself.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart, same as this profile's model/baseUrl/argv.
# HOT (fleetd #446): read live, cached by profile name, at every
# completion-fallback check AND by fleet_profiles' exhaustionDetectionArmed —
# editing it and reloading arms or disarms usage-limit detection for this
# profile with no daemon restart. (Before fleetd #446 this was DEFERRED,
# compiled once into a startup pattern map like model/baseUrl/argv still are —
# see errorPattern below, which is still deferred that way on purpose.)
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
# credentialId share one quarantine — the case this exists for is two models
@@ -219,8 +272,9 @@ herdrSocket: ~/.config/herdr/herdr.sock
# happens, just without a profile-specific match; every backend words its
# failure differently, so a hardcoded sentence would only ever match one
# of them.
# DEFERRED: compiled once into a startup pattern map, same as exhaustedPattern
# — editing it needs a daemon restart.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart. Unlike exhaustedPattern above (made hot by fleetd #446),
# errorPattern was scoped out of that ticket on purpose and stays deferred.
# # errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage
#
# What happens once a match fires (BackendOutagePolicy, credentialId-keyed,
@@ -439,6 +493,15 @@ placement: weighted
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
#
# fleetd #466: this is now only the BASE of an escalating backoff, not a flat retry rate. A
# credential quarantined again within one base cooldown of the previous quarantine ending (still
# reporting exhausted — e.g. a weekly subscription limit that hasn't reset) backs off further:
# cooldown doubles each such time, capped at 12x this value (~6 hours at the 1800s default). A
# quarantine that starts after a base-cooldown's worth of quiet resets back to this value. Not
# configurable per se — the multiplier and ceiling are constants in BackendQuarantine, not new
# YAML keys; see its class doc for the exact formula and why there is no automatic probe to clear
# it early (the operator's own design constraint — a probe spends the quota it's measuring).
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
@@ -460,9 +523,10 @@ placement: weighted
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
# / credentialId. Those are hot because the placement policy (and, for credentialId,
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
# not by itself enough to make a key hot.
# / credentialId / exhaustedPattern. Those are hot because the placement policy (and,
# for credentialId, the CB-578 stage B quarantine check; for exhaustedPattern, fleetd
# #446's LiveExhaustedPatterns) reads them through a supplier — being config is not by
# itself enough to make a key hot.
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
@@ -474,10 +538,10 @@ placement: weighted
# stage B — baked once into the quarantine tracker built at startup), ADDING or
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
# configDir, mcpUrl, tabLabel, exhaustedPattern, errorPattern (fleetd #201 / #227 —
# compiled once into a startup pattern map the same way exhaustedPattern is). The
# launcher takes a copy of `profiles:` at startup and resolves every spawn out of
# that copy, so those never
# configDir, mcpUrl, tabLabel, errorPattern (fleetd #201 / #227 — compiled once into
# a startup pattern map; exhaustedPattern used to be compiled the same way until
# fleetd #446 made it hot — see above). The launcher takes a copy of `profiles:` at
# startup and resolves every spawn out of that copy, so those never
# reach a launch until you restart. The reload logs them by name rather than
# pretending they applied.
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
@@ -774,6 +838,28 @@ guard:
# so fleetd falls back to the weaker CB-596 sentinel overlay instead (a WARN names the gap).
# worktreeGroup: fleet-workers
# fleetd #362: a directory of skill folders (each a subdirectory holding a SKILL.md, the same
# shape as this repo's own .claude/skills/) copied into every PROVISIONED worktree's
# .claude/skills/, so a member spawned against ANY repo — not only one that already ships its own
# copy — can load a bridge skill (e.g. implementer). Unset (the default): no worktree is touched
# beyond today's behaviour. A skill folder the target repo already carries under
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Best-effort like
# worktreeGroup above: a missing/unreadable directory here is logged and skipped, never a failed
# spawn. Every non-hidden subdirectory of this directory is copied wholesale, with no per-file
# allowlist — don't park scratch files or drafts alongside the real skill folders, they will be
# copied into every provisioned worktree too.
#
# fleetd #393: which member KINDS actually consume this once it is copied. kind: claude-code —
# the Claude Code CLI discovers .claude/skills/ on its own; nothing else is needed. kind: opencode
# — opencode has no such discovery, so OpenCodeLauncher reads whatever landed under
# .claude/skills/ and appends each seeded skill's SKILL.md to the generated instructions[] file
# (opencode's only channel for static guidance text; unlike Claude Code's Skill tool, the content
# is always part of the system prompt, not loaded on demand). Both kinds are covered as of #393 —
# earlier builds copied the files for every kind but only claude-code could read them, and the
# seeding log said "N of M" regardless. Check the per-spawn launcher log (not just the seeding
# log) to see what a given member actually got.
# memberSkills: /path/to/fleetd/checkout/.claude/skills
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
@@ -821,10 +907,17 @@ guard:
# across every daemon sharing this vhost.
# prefetch → consumer basicQos, capping how many unacked messages the mailbox holds in-heap.
# Default 32 when omitted.
# peers → fleetd #361: the coord-ids of the OTHER daemons on this vhost, declared by the
# operator (the daemon never guesses). fleet_list reports each one's live reachability
# (a passive queue check, never a presence protocol) alongside this daemon's own
# mailbox state. Omit, or leave empty, for a daemon with no known peers yet — an
# undeclared peer can still reach you and be reached by fleet_send, it just will not
# show up as a row in fleet_list.
# coordinator:
# uriEnv: LEAD_COORD_URI
# selfId: mac-opus
# prefetch: 32
# peers: [fleet01-lead]
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open fleet_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
@@ -849,3 +942,32 @@ guard:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
# Central allow-list of models any profiles: entry may name. Nothing checked a profile's model:
# value before this block existed — it was a free-form string handed straight to the backend
# adapter, and a withdrawn or misspelled name failed silently instead of at config load (opencode
# falls back to a default model rather than erroring on an unknown -m).
#
# Absent, or present with an empty allow:, is OFF: no profile's model: is checked, exactly like
# before this block existed. fleetd.yaml is gitignored on every host, so an upgrade must not force
# every operator to enumerate their models before the daemon will start.
#
# The list is the authority; profiles: is checked against it, never the reverse — adding or
# editing a profiles: entry cannot, by itself, widen what is permitted here.
#
# Enforcement is at CONFIG LOAD only (a bad model: fails the daemon at startup, naming both the
# model and the profile). There is no spawn-time enforcement, no runtime on/off switch, and no
# interaction with BackendQuarantine — those are separate, later units.
#
# allow → the permitted models. Each entry is its own block (not a bare string) so a later unit
# can add an on/off state or a load limit per model without changing this shape.
# model → the model id exactly as a profiles: entry's model: field would write it. One flat,
# opaque-string namespace: a bare Claude id (claude-sonnet-5) and an opencode
# provider-prefixed id (openai/gpt-5.6-terra) both fit here unchanged — the check is a
# plain string match, never a parse of the provider prefix or a branch on kind:.
# models:
# allow:
# - model: claude-sonnet-5
# - model: claude-opus-5
# - model: openai/gpt-5.6-terra
# - model: amazon.nova-pro-v1:0
+675 -103
View File
@@ -10,6 +10,7 @@ import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -18,6 +19,7 @@ import dev.ltms.fleet.inject.BackendErrorSink;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.inject.TurnListener;
@@ -25,6 +27,7 @@ import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.mcp.CharterToolSurface;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -55,6 +58,8 @@ import dev.ltms.fleet.member.MemberCredentialPolicyView;
import dev.ltms.fleet.member.OpenCodeLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.power.CaffeinateSleepAssertionMechanism;
import dev.ltms.fleet.power.IdleSleepGuard;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -63,18 +68,23 @@ import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
/**
* {@code fleetd} entry point. Wires the real herdr socket client to the REST app and
@@ -121,35 +131,55 @@ public final class Fleetd {
// else can fail on a silently-empty one. A daemon started without a login shell (launchd)
// boots fine either way — this is the only thing that says so out loud.
reportRequiredSecrets(cfg);
// fleetd #377: the git host value a member receives as GITEA_HOST is often a full URL
// (scheme and trailing slash), not a host name. A member that assumes a bare host then
// builds https://https://... and the request never leaves the machine. Report the shape
// next to the secret report — shape only, never the value.
reportGitHostShape(cfg);
reportMemberTrustModel(cfg);
// CB-596: an absent (or empty) memberCredentials: block blocks NOTHING — no credential
// name is hardcoded any more to fall back on. Say so loudly, the same way a missing
// secret is reported above, so upgrading past this commit never silently drops CB-592's
// protection.
reportMemberCredentialsGap(cfg);
// fleetd #395: an unset exhaustedPattern is a silent opt-out of usage-limit detection for
// that profile — say so loudly, the same way the two reports above do, rather than let an
// operator discover it only when a limit goes undetected.
reportExhaustedPatternGap(cfg);
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
// contract; adding a reader here does not make a key reloadable by itself.
ConfigRef config = new ConfigRef(configPath, cfg);
// fleetd #474: pass the charter/tool-surface check in as ConfigRef's extraValidation, so
// ConfigRef#reload() runs the same gate main() runs below, without dev.ltms.fleet.config
// gaining a dependency on dev.ltms.fleet.mcp — Fleetd is the seam that already holds both.
ConfigRef config = new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools);
// The primary/host env that launched fleetd must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
cfg.validateLeadTabPrefixes();
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
cfg.validateCharters();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateMembers();
// Every FleetConfig.validateXxx() the operator's config can fail — CB-501's auth-exposure
// check, CB-531's lead-tab-prefix check, CB-542's subscription-profile check, the charter
// and member-slot checks, and the "central allow-list of usable models" check — must run
// here, at load, before anything below opens a socket or spawns a member. fleetd ticket
// "central allow-list of usable models" follow-up: mutation testing found six individual
// calls here with nothing proving any of them still ran (deleting one left the full suite
// green). validateAll() replaces them with the one call that FleetConfigValidateAllTest
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
// why a name-by-name list here would have the same defect it replaces.
cfg.validateAll();
// fleetd #469, follow-up to #464: validateAll() (and validateCharters() inside it) only
// checks that a charter's KEY is a role wire name and its text is non-blank — it never
// looks at what the text actually names. This is the separate check that does: it asks
// dev.ltms.fleet.mcp.FleetTool (the canonical registered-tool set) whether every fleet_*/
// bridge_* token a charter names is a tool this server actually registers. It cannot live
// inside FleetConfig#validateCharters() — config loads before the MCP server exists, and
// must not gain a dependency on the mcp package — so it runs here instead, at the one seam
// that already holds both a loaded FleetConfig and the mcp package, before anything below
// opens a socket or spawns a member. fleetd #474: the same check is also wired into `config`
// above as ConfigRef's extraValidation, so a reload refuses what this line refuses at startup.
assertChartersNameOnlyRegisteredTools(cfg);
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
@@ -209,7 +239,13 @@ public final class Fleetd {
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
// fleetd #466: escalating, not flat — a credential that keeps reporting exhaustion (e.g. a
// weekly subscription limit, which would otherwise be retried on every ~30-minute cooldown,
// about 336 times across the week) backs off further each consecutive time, capped at
// BackendQuarantine.DEFAULT_MAX_COOLDOWN_MULTIPLE x the base cooldown. See BackendQuarantine's
// class doc for the mechanism, why this never fires on cooling-off (a separate, unescalated
// mechanism — BackendOutagePolicy below), and the reset.
BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
@@ -223,20 +259,24 @@ public final class Fleetd {
profileName -> liveCountRef.get().apply(profileName),
quarantine,
outagePolicy);
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into —
// no models: block at all, a block armed with nothing off, or a block with N off — the same
// way exhaustedPatternCoverageLine/errorPatternCoverageLine report CB-578 stage A/fleetd
// #201 Unit 5 coverage just below. Read from workers.modelGateState() (never a separate
// config.get().models() here) so this line and fleet_profiles' modelGateArmed can never
// disagree about what CompositePeerLauncher's spawn gate actually enforces.
log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState()));
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
boolean herdrUp = awaitHerdr(herdr);
HerdrAwaitOutcome herdrOutcome = awaitHerdr(herdr, System::nanoTime, Fleetd::sleepHerdrPoll);
boolean herdrUp = logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
@@ -248,10 +288,31 @@ public final class Fleetd {
contextCap = cfg.lifecycle().contextCap();
}
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup()),
SessionManager sessions = new SessionManager(workers,
new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup(), cfg.memberSkills()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> liveSessionCount(sessions.roster(), profileName));
// Idle-sleep guard: hold an OS-level assertion against idle sleep while at least one
// member is live, so an unattended host does not idle-sleep out from under a member's
// long turn (see FleetConfig.IdleSleepGuard / dev.ltms.fleet.power.IdleSleepGuard for the
// measurement that motivated this). Opt-out via idleSleepGuard.enabled: false; on by
// default. Hangs off SessionManager's own onAcquire/onRelease hooks (CB-520/CB-516,
// previously wired only to the reply inbox) and SessionManager#size() — the exact registry
// fleet_list's live/capacity numbers are themselves computed from — rather than tracking
// members a second way. No-op (never constructed) off macOS or when idleSleepGuard.enabled
// is explicitly false; the mechanism itself is additionally a no-op if 'caffeinate' cannot
// be started, so this can never fail a spawn, a release, or startup.
boolean idleSleepGuardEnabled = cfg.idleSleepGuard() == null || cfg.idleSleepGuard().isEnabled();
final IdleSleepGuard idleSleepGuard;
if (idleSleepGuardEnabled) {
idleSleepGuard = new IdleSleepGuard(new CaffeinateSleepAssertionMechanism(), sessions::size);
sessions.onAcquire(_ -> idleSleepGuard.recheck());
sessions.onRelease(_ -> idleSleepGuard.recheck());
} else {
idleSleepGuard = null;
}
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
@@ -323,7 +384,11 @@ public final class Fleetd {
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
MemberRegistry members = new MemberRegistry(cfg.fleet());
// fleetd #424: MemberRegistry.live re-reads fleet.architects through `config` on every
// reserve/requireSlotFor call, so a reload that removes or adds an architect slot governs
// the next spawn with no restart — the frozen `new MemberRegistry(cfg.fleet())` this used
// to be let a "revoked" slot keep granting new architect spawns forever.
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
@@ -337,29 +402,34 @@ public final class Fleetd {
Rendezvous rendezvous = new Rendezvous();
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
// Compiled once at startup, keyed by profile name; a profile with no exhaustedPattern is
// simply absent here, so its workers keep today's completion-fallback behaviour unchanged.
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasExhaustedPattern()) {
exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));
}
});
// fleetd #446: read live off `config` per lookup, cached by profile name — see
// LiveExhaustedPatterns's class doc for why this replaces the old compiled-once-at-startup
// map. A profile with no exhaustedPattern simply returns null here, so its workers keep
// today's completion-fallback behaviour unchanged.
LiveExhaustedPatterns liveExhaustedPatterns = new LiveExhaustedPatterns(() -> config.get().profiles());
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
.map(session -> liveExhaustedPatterns.patternFor(session.profile()))
.orElse(null);
// The startup coverage line still reports the boot-time snapshot only — it is printed once,
// here, and a reload no longer needs to change what it said; exhaustionDetectionArmed (via
// liveExhaustedPatterns.armed, wired into quarantineSource below) is what stays live.
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
.filter(e -> e.getValue().hasExhaustedPattern())
.map(Map.Entry::getKey)
.collect(Collectors.toCollection(LinkedHashSet::new));
log.info("backend-exhausted classification (CB-578 stage A): {}",
CompletionResolver.coverage("exhaustedPattern", cfg.profiles().keySet(),
exhaustedPatternsByProfile.keySet()));
exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
// rather than handing it back as a real answer. Compiled once at startup, keyed by profile
// name, mirroring exhaustedPatternsByProfile above — a profile with no configured
// errorPattern is simply absent here, so CompletionResolver falls back to its built-in
// narrow {@code (?i)\bAPI Error\s*:} compatibility pattern for that profile's targets
// (BackendErrorPatternLookup#legacy's contract — see backendErrorPatterns below).
// rather than handing it back as a real answer. Still compiled once at startup, keyed by
// profile name — unlike exhaustedPattern above (fleetd #446), errorPattern was left DEFERRED
// on purpose: the ticket that made exhaustedPattern hot scoped errorPattern/cooling-off out
// explicitly. A profile with no configured errorPattern is simply absent here, so
// CompletionResolver falls back to its built-in narrow {@code (?i)\bAPI Error\s*:}
// compatibility pattern for that profile's targets (BackendErrorPatternLookup#legacy's
// contract — see backendErrorPatterns below).
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasErrorPattern()) {
@@ -372,52 +442,25 @@ public final class Fleetd {
BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,
errorPatternsByProfile);
log.info("backend-error classification (fleetd #201 Unit 5): {}",
CompletionResolver.coverage("errorPattern", cfg.profiles().keySet(),
errorPatternsByProfile.keySet()));
errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
//
// fleetd #175: this sink is now also the quarantine target for OpenCodeLauncher's
// model-mismatch check — it needs nothing profile-specific from the caller beyond `target`
// (a herdr terminal id) and `reason`, so reusing it here is exactly "the existing
// ExhaustionSink path", not a new mechanism.
//
// fleetd #234: that check fires from SessionAwareHandle.agentSessionId(), which runs during
// SessionManager.acquire() BEFORE this session is registered in sessions.roster() — so the
// roster-only lookup below used to find nothing, .ifPresent silently no-op'd, and the
// ERROR the check had just logged ("quarantining this profile's credential") was a lie:
// nothing was quarantined, and nothing said so. Two changes: (1) OpenCodeLauncher now
// passes its OWN profile name via ExhaustionSink's 3-arg overload — it already has the
// FleetConfig.Profile in hand and does not need the roster at all — used here as a
// fallback whenever the roster lookup misses; (2) if a profile still cannot be resolved
// (neither the roster nor the hint names a configured one), this logs loudly at ERROR
// instead of silently doing nothing — a control that cannot act must say so.
// fleetd #234, round 4: the 3-arg overload is now ExhaustionSink's single abstract method,
// so this is safely a lambda — there is no separate 2-arg overload left for it to bind to
// instead and silently drop profileHint (that was rounds 1-3's whole hazard).
ExhaustionSink exhaustionSink = (target, reason, profileHint) -> {
String profileName = sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.orElse(profileHint);
FleetConfig.Profile profile = profileName == null ? null : config.get().profiles().get(profileName);
if (profile == null) {
log.error("quarantine requested for target '{}' ({}) but no profile could be "
+ "resolved — the target is not (yet) in the roster, and {} — "
+ "credential NOT quarantined (fleetd #234)",
target, reason,
profileHint == null ? "no profile hint was given"
: "the hinted profile '" + profileHint + "' is not configured");
return;
}
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
};
// fleetd #446 criterion 3: the backend text that triggered the most recent quarantine of
// each credential, so fleet_profiles can report WHY a limit was hit, not only that it was.
// Keyed by credential id — the same key BackendQuarantine's own remainingSeconds uses —
// and written at the one call site that actually quarantines (inside exhaustionSink below),
// so a reason can never be reported for a quarantine that never happened. Bounded the same
// way BackendQuarantine's own internal map is documented to be: by the number of distinct
// credentials ever exhausted, not by how often the config is edited.
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
// fleetd #446 follow-up (round 3): extracted into exhaustionSink(...) below — see that
// method's javadoc for the full fleetd #175/#234/#446 history this used to carry inline —
// so a dedicated test can drive the exact ExhaustionSink main() builds, not a hand-rebuilt
// copy of its shape.
ExhaustionSink exhaustionSink = exhaustionSink(sessions, config, quarantine,
quarantineReasonByCredential, cfg);
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
// now that `sessions` exists to resolve target -> session -> profile.
exhaustionSinkRef.set(exhaustionSink);
@@ -546,6 +589,16 @@ public final class Fleetd {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed, so
// an upgraded daemon cannot silently acquire the ability to clear the lead's own pane.
// Unlike heartbeat above, this has no recurring scheduler of its own — nothing but an
// explicit confirm() call (wired to an MCP tool by a later ticket; nothing calls it yet)
// that passes every gate can ever schedule a roll. It does not take primaryRegistry: the
// lead terminal to roll comes from the caller of open()/confirm() (resolved by the MCP
// layer from the connection, the same way auth/CallerResolver#resolve builds a
// Principal.leader(...)), never from a single-slot lookup — see LeadRollover's class
// javadoc, fleetd #480 correction 2.
LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
pushLoop, metrics);
@@ -556,8 +609,12 @@ public final class Fleetd {
if (cfg.health() != null && cfg.health().isEnabled()) {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
// through the same idempotent target-wide operation CB-516 already uses on release.
// fleetd #386: System::nanoTime freezes across a macOS sleep, so the stall check also
// gets a wall-clock source to detect and correct for that freeze. Every other decision
// in FleetHealthMonitor stays on the monotonic clock, unchanged.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, cfg.health().intervalOrDefault(),
System::nanoTime, () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis()),
cfg.health().intervalOrDefault(),
cfg.health().workingSuspectAfterOrDefault(), messages::abandon);
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
@@ -629,21 +686,16 @@ public final class Fleetd {
// built a second time — two independently-constructed sources reading the SAME BackendQuarantine
// / BackendOutagePolicy would still be able to drift (e.g. a future edit to the credentialIdFor
// closure in only one of the two places), exactly the shape #284 was.
FleetMcp.QuarantineSource quarantineSource = new FleetMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine);
FleetMcp.QuarantineSource quarantineSource = quarantineSource(config, quarantine,
liveExhaustedPatterns, quarantineReasonByCredential);
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, outagePolicy);
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics, new FleetMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime),
primaryRegistry, callers, metrics,
capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
new FleetMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
@@ -652,7 +704,15 @@ public final class Fleetd {
quarantineSource,
leadMailbox,
outageSource,
new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), leaders, leads)));
new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), leaders, leads)),
// fleetd #361: the operator-declared peers this daemon's fleet_list should try to
// reach. Read from the SAME snapshot leadMailbox itself opened from (cfg.coordinator()),
// not the live config.get() — coordinator wiring is already boot-time-fixed (see
// leadMailbox above), so peers follows the same rule rather than half hot-reloading.
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
// fleetd #480 Unit C: the executor behind fleet_handover — null whenever
// leadRollover: is not configured (see the leadRollover local above).
leadRollover);
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
@@ -698,6 +758,11 @@ public final class Fleetd {
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
// Idle-sleep guard: release unconditionally, even though sessions.close() above already
// drained every session (and each release already drove the live count to 0, which
// releases the guard's assertion on its own) — this is the backstop for a drain that was
// itself interrupted or threw, so no caffeinate child ever outlives the daemon.
if (idleSleepGuard != null) idleSleepGuard.close();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
@@ -756,6 +821,223 @@ public final class Fleetd {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
* fleetd #446 follow-up (round 3): the quarantine {@link ExhaustionSink} main() actually wires
* — extracted out of {@code main} for the same reason {@link #quarantineSource} below and
* {@link #usageLimitFixWarning}/{@link #usageLimitFixWarningNoModel} were. Round 2 pinned the
* WARNING text by having {@code FleetdUsageLimitFixWarningTest} call those two methods
* directly, but a mutation battery proved that left two gaps: nothing proved this sink's
* {@code log.warn} call actually uses either method (replacing the whole ternary with a
* literal string), and nothing proved it picks the right one for a profile WITH a configured
* {@code model:} versus one WITHOUT (swapping the ternary's two branches) — both mutations
* left round 2's full 1592-test suite green. {@code FleetdExhaustionSinkWarningTest} drives
* THIS factory's return value directly and asserts on the real text a {@code ListAppender}
* attached to this class's own logger captures — the only way to prove the caller, not just
* the two callees in isolation.
*
* <p>Behaviourally unchanged from the inline lambda this replaces:
* <ul>
* <li>fleetd #175: also the quarantine target for {@code OpenCodeLauncher}'s model-mismatch
* check — it needs nothing profile-specific beyond {@code target}/{@code reason}, so
* reusing this sink is the existing {@code ExhaustionSink} path, not a new mechanism;</li>
* <li>fleetd #234 (round 4): resolves {@code target} to a profile via the live session
* roster first, falling back to {@code profileHint} — {@code
* SessionAwareHandle.agentSessionId()} fires this sink during {@code
* SessionManager.acquire()}, <em>before</em> that session is registered, so the
* roster-only lookup alone used to silently miss it. A profile resolved neither way
* logs loudly at ERROR instead of doing nothing;</li>
* <li>fleetd #446 criterion 3: writes {@code quarantineReasonByCredential} at the one call
* site that actually quarantines, keyed by credential id — see the field's declaration
* in {@code main} for the bound on its size.</li>
* </ul>
*
* <p>{@code cfg} is the startup {@link FleetConfig} snapshot, read only for {@code
* quarantineCooldownSeconds()} in the log text — a Cold key (see {@code ConfigRef}'s class
* doc), so reading it off the snapshot rather than {@code config.get()} makes no observable
* difference and matches what the inline version already did.
*/
static ExhaustionSink exhaustionSink(SessionManager sessions, ConfigRef config, BackendQuarantine quarantine,
Map<String, String> quarantineReasonByCredential, FleetConfig cfg) {
return (target, reason, profileHint) -> {
String profileName = sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.orElse(profileHint);
FleetConfig.Profile profile = profileName == null ? null : config.get().profiles().get(profileName);
if (profile == null) {
log.error("quarantine requested for target '{}' ({}) but no profile could be "
+ "resolved — the target is not (yet) in the roster, and {} — "
+ "credential NOT quarantined (fleetd #234)",
target, reason,
profileHint == null ? "no profile hint was given"
: "the hinted profile '" + profileHint + "' is not configured");
return;
}
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
quarantineReasonByCredential.put(credentialId, reason);
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
// fleetd #446 criterion 2: name the fix, not just the fact — an operator reading this
// should not have to work out which of several configured models to touch.
String model = profile.model();
log.warn(model != null && !model.isBlank()
? usageLimitFixWarning(profile.profile(), model)
: usageLimitFixWarningNoModel(profile.profile(), cfg.quarantineCooldownSeconds()));
};
}
/**
* fleetd #404, superseded by fleetd #446: production source for quarantine reporting.
* Credential IDs were already hot; {@code exhaustedPattern} used to be compiled once at
* startup for {@link CompletionResolver}, so the armed field had to read that same frozen map
* until restart — the exact asymmetry #446 exists to close. Both {@code exhaustedPatternArmed}
* here and {@link CompletionResolver}'s own classification now read the ONE live {@link
* LiveExhaustedPatterns} instance ({@code exhaustedPatterns.armed}/{@code .patternFor}), so a
* reload that arms or disarms a profile's detection is visible to both at once — never two
* independently-updated copies that could disagree, the fleetd #404 rule this method used to
* violate on purpose and now upholds.
*
* <p>{@code modelFor} and {@code reasonFor} back fleetd #446 criterion 3 — {@code
* fleet_profiles}'s quarantined row naming which model a quarantined profile runs
* ({@code modelFor}, read live off {@code config} the same way {@code credentialIdFor} already
* is) and the backend text that triggered the most recent quarantine of that credential
* ({@code reasonFor}, backed by {@code reasonByCredential} — see its call site in {@code main}
* for where that map is written, at the one place a credential is actually quarantined).
*/
static FleetMcp.QuarantineSource quarantineSource(ConfigRef config, BackendQuarantine quarantine,
LiveExhaustedPatterns exhaustedPatterns, Map<String, String> reasonByCredential) {
return new FleetMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine, exhaustedPatterns::armed, profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.model();
}, reasonByCredential::get);
}
/**
* fleetd #415 (review follow-up): package-private factory for the CB-578 stage A {@code
* exhaustedPattern} startup coverage line, paired explicitly with {@link
* CompletionResolver.UnsetMeaning#OFF} — {@code exhaustedPattern} has no fallback, so a
* profile with none configured really does have the classification off.
*
* <p>Extracted out of {@code main} for the same reason {@link #capacitySource} and {@link
* #worktreeBranchLookup} were: {@code coverage()}'s own tests ({@code CompletionResolverTest})
* prove it words {@code OFF} and {@link CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT}
* correctly when a test supplies the meaning itself — they cannot prove {@code main} pairs the
* right meaning with the right key, which is the actual fleetd #415 defect. <b>Measured:</b>
* swapping the {@code UnsetMeaning} arguments between this method and {@link
* #errorPatternCoverageLine} — recreating #415's defect with the two keys exchanged — compiled
* with 0 errors and left all 1506 existing tests green before {@code
* FleetdPatternCoverageLineTest} was added to catch exactly that swap.
*/
/**
* fleetd #446 follow-up: the criterion-2 WARNING text — "name the fix, not just the fact" —
* for a profile whose {@code model:} is configured. Extracted out of the {@code
* exhaustionSink} lambda the same way {@link #exhaustedPatternCoverageLine} was extracted out
* of {@code main}: the inline version compiled, ran, and read correctly, but nothing pinned
* its text, so a later edit could silently stop naming the fix and every test would stay
* green (measured: renaming the leading {@code "usage-limit fix:"} tag left all 1586
* pre-follow-up tests passing). {@link FleetdUsageLimitFixWarningTest} asserts on the return
* value of this method directly, which is exactly what the caller logs — {@code log.warn} is
* called with this method's result as a single, already-formatted argument, so the string a
* test sees here is byte-identical to what {@code fleetd.out} receives.
*
* <p>Deliberately asserts on three substrings, not the whole sentence (the profile name, the
* model name, and the literal {@code enabled: false} / {@code models.allow} action) — see that
* test's class doc for why a whole-sentence pin is the wrong granularity here.
*/
static String usageLimitFixWarning(String profile, String model) {
return "usage-limit fix: profile '" + profile + "' runs model '" + model + "' — set "
+ "`enabled: false` on that model's entry under models.allow in fleetd.yaml to "
+ "stop new spawns landing on it (models: is hot, no restart needed); remove the "
+ "line again once the subscription window resets";
}
/**
* fleetd #446 follow-up: the {@link #usageLimitFixWarning} counterpart for a profile with no
* {@code model:} configured — {@code models.allow} gates by model name, so there is nothing
* for an operator to flip, and this says so instead of naming a fix that does not exist.
*/
static String usageLimitFixWarningNoModel(String profile, Integer quarantineCooldownSeconds) {
return "usage-limit fix: profile '" + profile + "' has no model: configured, so "
+ "models.allow cannot gate it by name — the " + quarantineCooldownSeconds
+ "s quarantine above is the only thing keeping new spawns off it for now";
}
static String exhaustedPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
return CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
allProfiles, configuredProfiles);
}
/**
* fleetd #415 (review follow-up): the {@code errorPattern} counterpart of {@link
* #exhaustedPatternCoverageLine}, paired explicitly with {@link
* CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT} — an unset {@code errorPattern} still runs
* backend-error classification against {@code CompletionResolver}'s built-in {@code
* BACKEND_ERROR} pattern, so the empty case is not "off". See {@link
* #exhaustedPatternCoverageLine}'s javadoc for the measured swap mutation this pairing guards
* against.
*/
static String errorPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
return CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
allProfiles, configuredProfiles);
}
/**
* fleetd #422 follow-up: package-private factory for the startup line reporting which of the
* three central {@code models.allow:} gate states the daemon booted into. Extracted the same
* way {@link #exhaustedPatternCoverageLine}/{@link #errorPatternCoverageLine} are, so a
* dedicated test can call it directly rather than parsing log output, and so {@code main}'s
* only source for this line is {@link PeerLauncher#modelGateState()} — never a second,
* independently-derived read of {@code cfg.models()} that could disagree with what {@code
* CompositePeerLauncher}'s spawn gate actually enforces (the fleetd #404 lesson).
*
* <p>Unlike the two pattern-key lines above, there is no {@code UnsetMeaning} choice to make
* here: {@link PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether
* an empty {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate
* with" or "a block armed and currently reporting zero off" — the exact two states a bare
* {@code disabledModels()} read could not tell apart before this ticket.
*/
static String modelGateCoverageLine(PeerLauncher.ModelGateState state) {
if (!state.configured()) {
return "not configured (no models: block — nothing is gated, and nothing can be)";
}
Set<String> off = state.off();
return off.isEmpty()
? "armed (models: block present; 0 models currently turned off)"
: "armed (" + off.size() + " model(s) turned off: " + new TreeSet<>(off) + ")";
}
/**
* fleetd #416: production source for {@code fleet_list}'s per-profile capacity facts.
*
* <p>The profile <em>set</em> ({@code configuredProfiles}) must come from {@code cfg} — the
* startup snapshot — not the live {@code config.get()}. {@code profiles} as a whole is a
* {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}): {@code HerdrPeerLauncher} takes
* {@code Map.copyOf(profiles)} once at construction and a profile only added to the
* hot-reloaded map can never actually be spawned, so enumerating it live made {@code fleet_list}
* report a profile as available when {@code fleet_spawn} on that same profile fails with
* {@code unknown worker profile}. {@code fleet_list}'s own contract for {@code free} is "the
* same check the spawn gate itself runs" — the set the spawn gate can see is the startup one,
* so this must enumerate that one too, the same shape as {@code coordinator.peers} above.
*
* <p>{@code maxLoad} stays live on purpose: it is read off {@code config.get()} exactly like
* {@code credentialId} ({@link ConfigRef} documents both as hot), so an existing profile's
* {@code maxLoad} edit must still change what {@code fleet_list} reports without a restart.
*/
static FleetMcp.CapacitySource capacitySource(ConfigRef config, FleetConfig cfg,
Function<String, Integer> liveCount) {
return new FleetMcp.CapacitySource(liveCount,
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
},
cfg.profiles()::keySet, System::nanoTime);
}
/**
* fleetd #248: package-private factory for the member worktree/branch lookup {@link
* CompletionResolver} uses to name a fallback report's worktree and branch (fleetd#241).
@@ -783,6 +1065,68 @@ public final class Fleetd {
.orElse(null);
}
/**
* fleetd #480: construct the {@link LeadRollover} executor only when {@code leadRollover:} is
* present at startup — the same presence gate {@code leadHeartbeat:} uses just above this
* call site in {@code main}. Extracted to its own factory, the same reason
* {@link #worktreeBranchLookup} and {@link #exhaustionSink} are: a unit test can call this
* directly with a fabricated {@link FleetConfig} (absent block ⇒ {@code null}, present block ⇒
* constructed) without needing the whole of {@code main}, and a source-text test on the real
* call site proves {@code main} still calls this factory rather than inlining a copy that could
* silently diverge.
*
* <p>The returned object reads every {@code leadRollover:} field fresh on each {@code open()}/
* {@code confirm()} call through {@code () -> config.get().leadRollover()} — see {@code
* ConfigRef}'s class doc Hot bullet and {@link FleetConfig.LeadRollover}'s javadoc for why that
* makes the block's fields HOT despite this presence gate being evaluated once, at startup.
*
* <p>Deliberately does NOT take {@link PrimaryRegistry}: fleetd #480 correction 2 found that a
* single-slot lookup lets one lead's {@code confirm()} clear a DIFFERENT lead's pane on a
* daemon with more than one labelled lead tab. The lead terminal to roll instead comes from
* whoever calls {@code open()}/{@code confirm()} — the later MCP-tool unit must resolve it from
* the connection and pass it in, never take it as a request field. See {@link LeadRollover}'s
* class javadoc.
*
* <p><strong>fleetd #480 follow-up:</strong> also builds the terminal → lead-workspace lookup
* {@link LeadRollover#open} needs to resolve a relative {@code handoverPath} against the
* CALLING lead's own {@code cwd} rather than the daemon's — the daemon and a lead's pane can
* have different working directories (this repo nests {@code fleetd/} inside its own root, so
* they already differ on this host). The lookup is terminal → lead name (via {@code
* liveLeadTerminals}) → that lead's {@code cwd} (via {@code config.get().fleet().leaders()}),
* and both hops are read LIVE on every call, never off a snapshot taken here: leads are
* discovered by a live tab scan ({@code LeadTabScanner}), so a map captured at construction
* time could be empty (no lead has been scanned yet) or stale (a lead added since).
*
* @param cfg the startup config snapshot — read ONCE here, only to decide whether
* to construct the object at all, exactly like {@code
* cfg.leadHeartbeat()}
* @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not
* {@code memberAgents}), normally {@code router.leadAgents()}
* @param config the live {@link ConfigRef}, captured only inside the returned
* supplier and the workspace lookup — never dereferenced here
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
* the same {@code leads} supplier {@code main} already builds for
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
* absent from the startup config
*/
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents, ConfigRef config,
Supplier<Map<String, String>> liveLeadTerminals) {
if (cfg.leadRollover() == null) {
return null;
}
Function<String, String> leadWorkspace = terminal -> {
String leadName = liveLeadTerminals.get().get(terminal);
if (leadName == null) {
return null;
}
FleetConfig.Fleet fleet = config.get().fleet();
FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName);
return leader == null ? null : leader.cwd();
};
return new LeadRollover(leadAgents, () -> config.get().leadRollover(), leadWorkspace);
}
/**
* fleetd #176: per-profile factory for {@link FleetMcp.LeadSeatSource} — how many seats a
* profile's own live LEAD session(s) hold on the same Claude subscription.
@@ -1162,6 +1506,76 @@ public final class Fleetd {
});
}
/**
* fleetd #377: the env var names holding the git host value that members receive as
* {@code GITEA_HOST}. {@code HerdrPeerLauncher.applyGitToken} injects {@code GITEA_HOST}
* only for profiles that opted in via {@code gitTokenEnv} (CB-302), reading the value from
* that profile's {@code gitHostEnv} (default {@code GITEA_HOST}), so the shape only matters
* where a git token is opted in. A var used by more than one profile is one entry naming
* every profile that reads it, the same shape as {@link #requiredSecretEnvVars}.
*
* <p>Package-private and pure (no I/O, no logging) so the derivation is unit-testable
* without capturing log output; {@link #reportGitHostShape(FleetConfig)} is the logging caller.
*/
static Map<String, List<String>> gitHostEnvVars(FleetConfig cfg) {
Map<String, List<String>> hostsBy = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasGitToken()) {
hostsBy.computeIfAbsent(profile.gitHostEnv(), _ -> new ArrayList<>())
.add("profile '" + name + "' gitHostEnv");
}
});
return hostsBy;
}
/**
* True when the value already starts with a URI scheme ({@code https://...}, {@code
* http://...}). A bare host name and a host:port must both report {@code false} — the shape
* this ticket exists for is a value that <em>looks like</em> a host but is a full URL, and
* confusing those in the report would move the failure to the log instead of the network.
*/
static boolean startsWithScheme(String value) {
return value.matches("[A-Za-z][A-Za-z0-9+.-]*://.*");
}
/**
* fleetd #377: log, on the same startup path as {@link #reportRequiredSecrets}, the SHAPE of
* each git host value that members receive as {@code GITEA_HOST}: set or unset, its length,
* whether it starts with a scheme, whether it ends with a slash. Never the value itself — the
* same discipline as {@link #reportRequiredSecrets}, which logs by name only. Either shape is
* legitimate: the value is passed to members unchanged, and a line that quietly rewrites it
* would change what works on one host and breaks on another. The shape line only tells the
* operator which URL form to expect from a member that builds on {@code GITEA_HOST}. An
* unset variable is logged at INFO — useful information, not an error — and the daemon
* starts on either way.
*
* <p>The env read lives in the overload below so a test can drive the line with a known
* value and prove that value never reaches the log.
*/
static void reportGitHostShape(FleetConfig cfg) {
reportGitHostShape(cfg, System.getenv());
}
static void reportGitHostShape(FleetConfig cfg, Map<String, String> env) {
Map<String, List<String>> hostsBy = gitHostEnvVars(cfg);
if (hostsBy.isEmpty()) {
log.info("startup git host: no profile sets a gitTokenEnv — nothing to check");
return;
}
hostsBy.forEach((varName, sources) -> {
String value = env.get(varName);
if (value == null || value.isBlank()) {
log.info("startup git host {}: unset ({}) — a member gets GITEA_TOKEN but no "
+ "GITEA_HOST value", varName, String.join(", ", sources));
} else {
log.info("startup git host {}: set ({}) — length={}, startsWithScheme={}, "
+ "trailingSlash={}",
varName, String.join(", ", sources),
value.length(), startsWithScheme(value), value.endsWith("/"));
}
});
}
/**
* fleetd #184: state the member trust model at startup. Environment controls and worktrees do
* not make a sandbox when fleetd and its members use the same OS user. A separate herdr may
@@ -1212,12 +1626,138 @@ public final class Fleetd {
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
* deliberately opt-in — {@code null}/blank means a backend refusal on that profile is never
* classified as {@code BACKEND_EXHAUSTED}, so its credential is never quarantined. That is a
* legitimate choice (guessing the vendor's wording would be worse), but an operator who never
* opted a profile in should not discover the gap only when a usage limit silently goes
* undetected. Warn once at startup, naming every unarmed profile, exactly like {@link
* #reportMemberCredentialsGap} — never refuse to start over it.
*
* @return true if herdr answered, false if it never did
* <p>A {@code subscription: true} profile that is unarmed gets a SECOND, louder WARN of its
* own: it bills the operator's metered Claude plan, the case where an undetected usage limit
* costs the most.
*
* <p>Package-private so a test can capture the real log via a {@link
* ch.qos.logback.core.read.ListAppender}, the same pattern {@link
* #reportMemberCredentialsGap}'s own test uses.
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
static void reportExhaustedPatternGap(FleetConfig cfg) {
List<String> unarmedSubscription = new ArrayList<>();
List<String> unarmedOther = new ArrayList<>();
cfg.profiles().forEach((name, profile) -> {
if (!profile.hasExhaustedPattern()) {
(profile.isSubscription() ? unarmedSubscription : unarmedOther).add(name);
}
});
if (unarmedSubscription.isEmpty() && unarmedOther.isEmpty()) {
log.info("exhaustedPattern: every configured profile has usage-limit detection armed");
return;
}
List<String> allUnarmed = new ArrayList<>(unarmedSubscription);
allUnarmed.addAll(unarmedOther);
allUnarmed = allUnarmed.stream().sorted().toList();
log.warn("exhaustedPattern: profile(s) {} have no exhaustedPattern configured — a "
+ "usage-limit refusal on any of them is never detected and never "
+ "quarantines its credential. Set exhaustedPattern (see "
+ "fleetd.example.yaml) to arm detection for a profile.",
allUnarmed);
if (!unarmedSubscription.isEmpty()) {
List<String> sortedSubscription = unarmedSubscription.stream().sorted().toList();
log.warn("exhaustedPattern: subscription profile(s) {} run on the operator's metered "
+ "Claude plan and have NO usage-limit detection armed — this is the "
+ "case where a missed usage limit costs the most.",
sortedSubscription);
}
}
/**
* fleetd #474: the one place both the startup call (right after {@code cfg.validateAll()} in
* {@link #main}) and the reload call (wired into {@code config}'s {@code extraValidation} above,
* via a method reference to this method) go through, so the two can never drift into checking
* different things. Extracted only to give {@link ConfigRef}'s {@code Consumer<FleetConfig>}
* hook a {@code FleetConfig -> void} shape to bind to — {@link CharterToolSurface} itself still
* takes the raw charter map and knows nothing about {@code ConfigRef} or {@code Fleetd}.
*
* <p>Package-private so a test can call it directly the same way the other startup-report
* helpers above are tested, without needing to drive {@link #main} for a unit-level check;
* {@code FleetdStartupValidationTest} proves the startup call site, and {@code
* FleetdConfigRefCharterToolSurfaceWiringTest} — by constructing {@code ConfigRef} with this
* exact method reference, the same way {@code main} does above — proves the reload call site.
* {@code dev.ltms.fleet.config.ConfigRefTest} pins the same reload behaviour too, through an
* equivalent {@code Consumer<FleetConfig>} it builds locally (it cannot see this package-private
* method from {@code dev.ltms.fleet.config}).
*/
static void assertChartersNameOnlyRegisteredTools(FleetConfig cfg) {
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
}
/**
* How {@link #awaitHerdr} ended (fleetd #498). The old code returned a bare {@code boolean},
* which collapsed two different facts onto the same {@code false}: the configured wait budget
* genuinely running out, and the waiting thread being interrupted possibly milliseconds in.
* Those need different operator messages — see {@link #logHerdrWaitOutcomeAndShouldReap} — so
* this is a third state, not a better number (the same shape fleetd #497 named). Never treat
* {@link #INTERRUPTED} as if it were {@link #DEADLINE_PASSED}: only the latter means herdr was
* actually given the full {@link #HERDR_WAIT_SECONDS} and still failed to answer.
*/
enum HerdrWaitResult {
/** herdr answered {@code ping} before the deadline. */
ANSWERED,
/** the configured {@link #HERDR_WAIT_SECONDS} budget elapsed with no answer. */
DEADLINE_PASSED,
/**
* the waiting thread was interrupted before the budget ran out — a different event from
* {@link #DEADLINE_PASSED} and must never be reported as "did not answer within Ns".
*/
INTERRUPTED
}
/**
* The outcome of one {@link #awaitHerdr} call, carrying the MEASURED elapsed wait time
* alongside {@link #result}. {@code elapsedNanos} is always measured against the {@code nanos}
* supplier passed to {@link #awaitHerdr} — never assume it equals the configured budget, the
* same defect fleetd #494 already fixed once in {@code LeadRollover}.
*/
record HerdrAwaitOutcome(HerdrWaitResult result, long elapsedNanos) {}
/**
* The real per-poll wait {@link #main} passes to {@link #awaitHerdr}: sleep
* {@link #HERDR_WAIT_POLL_MILLIS}, and on interruption re-set the thread's interrupt flag
* rather than throwing — {@link #awaitHerdr} detects an interruption by checking {@link
* Thread#isInterrupted()} right after this returns, so a poller that swallowed the flag
* instead of restoring it would make that check silently miss the interruption.
*/
private static void sleepHerdrPoll() {
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
}
}
/**
* Poll herdr's {@code ping} until it answers, the configured {@link #HERDR_WAIT_SECONDS}
* budget elapses, or the waiting thread is interrupted (CB-504, fleetd #498).
*
* <p>{@code nanos} and {@code poller} are required parameters with no defaulted overload
* (fleetd #415's shape: a defaulted overload is a silent survivor a green suite would vouch
* for) — the previous version read {@link System#nanoTime()} and called {@link Thread#sleep}
* directly, so nothing could drive it from a test. The one production call site in {@link
* #main} passes {@code System::nanoTime} and {@link #sleepHerdrPoll}.
*
* @param nanos a monotonic elapsed-time clock, e.g. {@code System::nanoTime} — never a
* wall-clock source, since only elapsed time (not a timestamp) is measured here
* @param poller called once per failed ping while the budget remains; must, on an
* {@link InterruptedException}, re-set the thread's interrupt flag rather than
* throw or swallow it — this method's interruption check reads that flag right
* after {@code poller.run()} returns
* @return the outcome and the measured elapsed wait time — see {@link HerdrAwaitOutcome}
*/
static HerdrAwaitOutcome awaitHerdr(HerdrClient herdr, LongSupplier nanos, Runnable poller) {
long start = nanos.getAsLong();
long deadline = start + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
@@ -1225,25 +1765,57 @@ public final class Fleetd {
if (waited) {
log.info("herdr is up");
}
return true;
return new HerdrAwaitOutcome(HerdrWaitResult.ANSWERED, nanos.getAsLong() - start);
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
if (nanos.getAsLong() >= deadline) {
return new HerdrAwaitOutcome(HerdrWaitResult.DEADLINE_PASSED, nanos.getAsLong() - start);
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
poller.run();
if (Thread.currentThread().isInterrupted()) {
return new HerdrAwaitOutcome(HerdrWaitResult.INTERRUPTED, nanos.getAsLong() - start);
}
}
}
}
/**
* Log the right message for {@code outcome} — never the configured {@link #HERDR_WAIT_SECONDS}
* budget alone, always the measured elapsed time next to it — and say whether {@link #main}
* should now reap orphan worker panes (fleetd #498).
*
* <p>Extracted out of {@link #main} so this decision is drivable from a test: {@link #main}
* boots the whole daemon and cannot itself be run in a unit test, but this is the exact,
* unmodified code {@link #main} calls for the decision, not a re-derivation of it.
*
* @return true only for {@link HerdrWaitResult#ANSWERED} — orphan workers are reaped only
* then, exactly as before this ticket
*/
static boolean logHerdrWaitOutcomeAndShouldReap(HerdrAwaitOutcome outcome) {
long elapsedMillis = TimeUnit.NANOSECONDS.toMillis(outcome.elapsedNanos());
if (outcome.result() == HerdrWaitResult.ANSWERED) {
return true;
}
if (outcome.result() == HerdrWaitResult.DEADLINE_PASSED) {
log.warn("herdr did not answer within the configured wait (configured={}s elapsed={}ms) "
+ "— starting anyway; /healthz will report degraded until it comes up. Orphaned "
+ "worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS, elapsedMillis);
return false;
}
// HerdrWaitResult.INTERRUPTED — a different fact from DEADLINE_PASSED (fleetd #498): the
// wait was cut short, not exhausted, and must never be reported as "did not answer within
// Ns" — that claim would be false and would send an operator to debug herdr for nothing.
log.warn("herdr wait was interrupted before the configured wait ran out (configured={}s "
+ "elapsed={}ms) — starting anyway; /healthz will report degraded until it comes "
+ "up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS, elapsedMillis);
return false;
}
private Fleetd() {
}
}
@@ -30,8 +30,27 @@ public final class Authz {
DRAIN,
/** Read-only observation: status, roster, profiles, task polling. */
READ,
/**
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
*
* <p>Deliberately <strong>not</strong> folded into {@link #READ}. {@code READ}'s grant
* rests on "the roster carries no secrets" (see its case below) — a lead-to-lead body is
* not the roster; it is where leads discuss host shapes, credentials and unmerged work.
* Mapping this to {@code READ} would let any worker read every peer lead's mail in full
* and would silently falsify that comment for every other {@code READ} caller.
*/
COORD_READ,
/** Scrape the metrics endpoint. */
METRICS
METRICS,
/**
* Drive the lead-rollover executor ({@code fleet_handover}: open/confirm/cancel a
* self-replace, fleetd #480 Unit C). Primary-only, same as {@link #SPAWN}/{@link #STOP}/
* {@link #DRAIN} — and, unlike those, the terminal it acts on is never even an argument:
* {@code LeadRollover#open}/{@code #confirm} are always called with the CALLER's own
* connection-resolved terminal (see {@code dev.ltms.fleet.lead.LeadRollover}'s class
* javadoc, fleetd #480 correction 2), so a primary can only ever roll itself.
*/
HANDOVER
}
/**
@@ -50,7 +69,7 @@ public final class Authz {
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN -> caller.isPrimary();
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
@@ -68,6 +87,11 @@ public final class Authz {
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
// is coordination between leads, not observation of the roster.
case COORD_READ -> caller.isPrimary();
};
}
@@ -244,7 +244,15 @@ public final class CallerResolver {
// already names what happens if that case is handed the primary role: a worker→primary
// escalation. So an unresolved caller is refused (ANONYMOUS — the same clean, already-tested
// "authenticated as nothing" outcome used everywhere else in this method), never promoted.
return isLoopback(remoteAddr) && c.resolved() ? Principal.primary(c.pid()) : Principal.anonymous();
//
// fleetd #505: the OTHER way a real pid can wrongly reach here with a null terminal — not a
// failed lsof lookup, but a herdr error partway through PaneLocator's pane scan. c.resolved()
// says nothing about that; it only tests the lsof sentinel (by design — see
// ConnectionIdentity.Caller#resolved). c.scanComplete() is the separate signal: a scan that
// could not check every pane must not be read as "checked everywhere, no match" — the pane it
// could not check might have been the caller's own. So both must hold before this promotes.
return isLoopback(remoteAddr) && c.resolved() && c.scanComplete()
? Principal.primary(c.pid()) : Principal.anonymous();
}
private boolean presentedTokenMatches(String authorizationHeader) {
@@ -11,6 +11,7 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.function.Supplier;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
@@ -19,16 +20,43 @@ import java.util.Objects;
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/{@code reviewers}
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
* re-reads {@code fleet:} on every call, through a supplier the same shape as
* {@code CompositePeerLauncher}'s (see {@code ConfigRef}'s class doc) — so a config reload
* that removes or adds an architect slot governs the <em>next</em> spawn with no restart.
* Only {@link #MemberRegistry(FleetConfig.Fleet)} freezes the pool at construction, and that
* constructor exists for tests and for the (rare) case of wiring a fixed, code-built config.</li>
* <li><b>terminal bindings</b> — owned by this registry, initially <em>empty</em>, and
* <strong>never</strong> touched by a reload. Config declares no architect terminal, so at
* startup every slot is idle and nothing resolves to an architect; a session only becomes one
* when the spawn lifecycle {@linkplain #bind(String, String) binds} its terminal to a slot.
* {@link CallerResolver} reads this through {@link #snapshot()} to turn a pane into an
* {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p><strong>The binding rule (fleetd #424): config governs what a bound slot still grants, as
* well as what may be bound next.</strong> Removing a slot from config revokes it — that is the
* ticket's entire point ("Revoking an architect slot does not revoke it"). Revoking it means an
* architect already bound to that slot loses the ARCHITECT privilege on its very next request:
* {@link #roleForSlot} and {@link #nameForSlot} read {@link #slots()} directly, with no cache, so
* the moment a slot drops out of config, {@link CallerResolver#resolve} (which calls both on every
* request from a bound pane, {@code CallerResolver.java:220}) can no longer confirm the pane's slot
* is an architect slot, and the pane falls through to {@code Principal.worker(...)}. What does
* <em>not</em> change is the {@code terminalToSlot} <em>occupancy</em> — the binding created by
* {@link #bind} is untouched by a reload, on purpose: unbinding it here would double-book the slot
* key (a second terminal could then bind to the "freed" key while the first is still the terminal
* the operator actually meant to demote) and would silently break {@link #unbind}'s compare-safe
* contract, which needs the original {@code terminal → slot} pair intact to remove it cleanly. So
* the demoted session keeps occupying its slot — {@link #slotForTerminal} and {@link #snapshot()}
* still name it — it just no longer resolves as an architect through that occupancy, and a fresh
* spawn still cannot bind to the same key while it is occupied ({@link #reserve}/
* {@link #requireSlotFor} refuse it anyway, since it is gone from {@link #slots()}). The demoted
* session's own turn is unaffected: {@code fleet_reply}'s authorization
* ({@code Authz.Action.REPLY}) is {@code caller.ownsSession(targetSession)} — identity by terminal,
* not by role — so a demoted architect can still end its own turn normally.
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
@@ -55,14 +83,42 @@ public final class MemberRegistry implements MemberLifecycle {
}
}
private final Map<String, Entry> slots;
private final Supplier<FleetConfig.Fleet> fleet;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code terminalToSlot}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Slot keys held between reservation and the terminal binding. Guarded by terminalToSlot. */
private final java.util.Set<String> reservedSlots = new java.util.HashSet<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
/**
* Freeze the pool at construction — for tests, and for the rare case of wiring a fixed,
* code-built config. Production wiring should prefer {@link #live}, which re-reads
* {@code fleet:} on every call.
*/
public MemberRegistry(FleetConfig.Fleet fleet) {
this(() -> fleet);
}
private MemberRegistry(Supplier<FleetConfig.Fleet> fleet) {
this.fleet = fleet;
}
/**
* Live variant (fleetd #424): {@code fleet} is read fresh on every {@link #slots()} call — pass
* {@code () -> config.get().fleet()}, the same supplier shape {@code CompositePeerLauncher}
* already uses for placement — so a reload that adds or removes an architect slot governs the
* next spawn's {@link #reserve}/{@link #requireSlotFor} check with no restart. A separate,
* private constructor rather than a same-arity public overload of
* {@link #MemberRegistry(FleetConfig.Fleet)}: a {@code FleetConfig.Fleet} and a
* {@code Supplier<FleetConfig.Fleet>} overload are ambiguous for a literal {@code null} — the
* same reason {@code CallerResolver.withLeads} is a static factory rather than a fourth
* constructor overload.
*/
public static MemberRegistry live(Supplier<FleetConfig.Fleet> fleet) {
return new MemberRegistry(Objects.requireNonNull(fleet, "fleet"));
}
/** Flatten every role pool in {@code fleet} into one map. Leaders are not members. */
private static Map<String, Entry> flatten(FleetConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
@@ -74,18 +130,22 @@ public final class MemberRegistry implements MemberLifecycle {
});
}
}
this.slots = Collections.unmodifiableMap(flat);
return Collections.unmodifiableMap(flat);
}
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
/**
* The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot of
* {@code fleet:} <em>as of this call</em> — see the class doc for which constructor makes that
* live versus frozen.
*/
public Map<String, Entry> slots() {
return slots;
return flatten(fleet.get());
}
/** The slots belonging to {@code role}, in definition order. */
/** The slots belonging to {@code role}, in definition order, as of this call. */
public Map<String, Entry> slotsFor(MemberRole role) {
Map<String, Entry> out = new LinkedHashMap<>();
slots.forEach((key, e) -> {
slots().forEach((key, e) -> {
if (e.role() == role) {
out.put(key, e);
}
@@ -117,31 +177,49 @@ public final class MemberRegistry implements MemberLifecycle {
}
/**
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
* The strong-model profile a slot runs under, as of this call.
*
* <p>Nothing in {@code src/main} calls this (fleetd #431 — grepped both the {@code
* .profileForSlot(} and the {@code ::profileForSlot} form). This javadoc used to say "what the
* spawn lifecycle reads", and that seam does not exist: the spawn lifecycle takes its profile
* from the {@link MemberLifecycle.SlotReservation} that {@code reserve} returns, never from
* here. Kept and pinned rather than deleted because it is the natural accessor for that seam
* if one is added; live for the same reason as {@link #roleForSlot}, so a reload cannot leave
* it answering for the old config.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return (e == null || e.profile() == null) ? null : e.profile();
}
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
/**
* The role a qualified slot key belongs to, or {@code null} when the key is not currently
* configured. Deliberately live, with no cache (fleetd #424, see the class doc's binding rule):
* removing a slot from config must make {@link CallerResolver#resolve} stop granting the
* ARCHITECT role for it on the very next request from a terminal that was bound to it, which is
* the ticket's whole point — revoking a slot must actually revoke it, not just refuse the next
* spawn.
*/
public MemberRole roleForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return e == null ? null : e.role();
}
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
/**
* The unqualified configured name for a slot, or {@code null} if it is not currently configured.
* Live for the same reason as {@link #roleForSlot} — see the class doc's binding rule.
*/
public String nameForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
return slots().containsKey(slotName);
}
/**
@@ -11,6 +11,7 @@ import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Consumer;
import java.util.function.Supplier;
/**
@@ -32,17 +33,64 @@ import java.util.function.Supplier;
* not the fact that they are config. Most of {@code fleet:} — every role pool
* ({@code architects}/{@code developers}/{@code reviewers}), {@code charters}, and
* {@code tabLabel} — is read the same live way, through the same supplier
* ({@code () -> config.get().fleet()}). <strong>But {@code fleet:} as a whole is NOT in this
* class</strong>: {@code fleet.leaders} inside the same key is frozen, which is exactly what
* makes {@code fleet:} split rather than hot — see below.</li>
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
* reads it live for placement (which profile an unqualified architect spawn may land on), and
* {@code MemberRegistry} separately reads it live, through its own instance of the same
* supplier shape, for identity — both which slot a spawn may bind to <em>and</em> what a slot
* already bound still grants. Removing an architect slot from config therefore revokes the
* {@link dev.ltms.fleet.auth.Role#ARCHITECT} role on the bound pane's very next request; only
* the slot <em>occupancy</em> survives, so the demoted session still holds its slot key until
* it unbinds. See {@code MemberRegistry}'s class doc for that binding rule.
* <strong>But {@code fleet:} as a whole is NOT in this class</strong>: {@code fleet.leaders}
* inside the same key is frozen, which is exactly what makes {@code fleet:} split rather than
* hot — see below. {@code models:} (fleetd #422) joined this class whole: {@link
* FleetConfig#validateModels()} re-runs fully against the fresh config on every {@link
* #reload()} (via {@link FleetConfig#validateAll()}), refusing a bad edit outright rather than
* caching a stale copy anywhere, and the on/off half added by fleetd #422 is read live both by
* {@code CompositePeerLauncher}'s spawn gate ({@code enforceModelEnabled} and its candidate
* filter) and by {@code fleet_profiles}/{@code GET /profiles} (via
* {@code PeerLauncher.disabledModels()}). Nothing about {@code models:} is baked into an
* object built at startup, so — unlike the deferred keys below — there is no frozen half left
* to report; it moved here from deferred rather than joining split. An existing profile's
* {@code exhaustedPattern} (fleetd #446) joined this class the same way: it used to be
* compiled once into {@code Fleetd.main}'s startup pattern map (see the Deferred bullet's old
* wording, and {@code LiveExhaustedPatterns}'s class doc for the history), and is now read
* live, cached by profile name, by both {@code CompletionResolver}'s classification (via
* {@code LiveExhaustedPatterns.patternFor}) and {@code fleet_profiles}'s {@code
* exhaustionDetectionArmed} (via {@code LiveExhaustedPatterns.armed}) — the one live object
* both read, so a reload that arms or disarms a profile's usage-limit detection takes effect
* on the next check with no restart. {@code errorPattern}, {@code exhaustedPattern}'s sibling
* key for backend-error (not usage-limit) classification, was deliberately left OUT of this
* fleetd #446 change and stays deferred below — the ticket scoped it out explicitly.
* {@code leadRollover:} (fleetd #480) joined this class whole, the same shape as
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
* than capturing them into fields at construction — unlike its closest
* structural cousin {@code leadHeartbeat:}, whose {@code LeadHeartbeatLoop} bakes
* {@code idleAfterNanos}/{@code backoffMs}/{@code quietNudgeCap} into final fields. The one
* restart-only edge is structural, not a stale value: {@code Fleetd.java} decides whether to
* construct the {@code LeadRollover} object at all off the startup snapshot (the same
* presence gate {@code leadHeartbeat:} uses), so a block ADDED where it was absent at boot
* needs a restart before anything exists to call — the same fact already true of adding a
* brand-new {@code profiles:} entry.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code idleSleepGuard:} ({@code Fleetd.java} reads it once, at startup, to decide whether
* to construct an {@code IdleSleepGuard} and wire {@code SessionManager}'s
* {@code onAcquire}/{@code onRelease} hooks to it — neither is rebuilt on reload, so a
* running daemon keeps whatever this was at startup regardless of a later edit),
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code guard:}, {@code worktreeRoot:} and {@code worktreeGroup:} (both baked once into the
* {@code GitWorktrees} built at {@code Fleetd.java:251} and never rebuilt — fleetd #323
* instance 2 found {@code worktreeGroup} missing from this list and from
* {@link #changedDeferredKeys}), {@code primary:} (fleetd #326 — {@code Fleetd.java:506, 519,
* {@code guard:}, {@code worktreeRoot:}, {@code worktreeGroup:} and {@code memberSkills:}
* (all three of the latter baked once into the {@code GitWorktrees} built at
* {@code Fleetd.java:251} and never rebuilt — fleetd #323 instance 2 found
* {@code worktreeGroup} missing from this list and from {@link #changedDeferredKeys};
* {@code memberSkills} (fleetd #362) followed the same shape), {@code primary:} (fleetd #326 — {@code Fleetd.java:506, 519,
* 520} read {@code cfg.primary()} only off the startup snapshot to build {@code
* PrimaryRegistry} and size {@code ReplyPushLoop}'s reminder cap/backoff, and neither is
* rebuilt on reload. Say the consequence exactly: {@code primary.terminal} is DEPRECATED
@@ -59,15 +107,17 @@ import java.util.function.Supplier;
* (if any) simply keeps its old settings), adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Fleetd.main}'s
* pattern map at startup), {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
* {@code Fleetd.main}'s backend-error pattern map at startup, the same way),
* {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
* {@code Fleetd.main}'s backend-error pattern map at startup; deliberately NOT made hot
* alongside {@code exhaustedPattern} by fleetd #446 — that ticket scoped {@code errorPattern}
* and cooling-off out explicitly),
* {@code ideProjectDir} / {@code ideOpenCommand} / {@code autoCompactWindow} (fleetd #323
* instance 1 — all three are read at spawn off the same frozen profile map and were missing
* from {@link #sameLaunchSettings}), and the rest of {@link #sameLaunchSettings}.
* {@code credentialId} (CB-578 stage B) is NOT on
* this list — it is read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
* {@code credentialId} (CB-578 stage B) and {@code exhaustedPattern} (fleetd #446) are NOT on
* this list — both are read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so both are hot instead
* (see the Hot bullet above for {@code exhaustedPattern}'s history).
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
* reload logs these rather than pretending they applied.</li>
@@ -129,15 +179,22 @@ import java.util.function.Supplier;
* five of COLD_KEYS" rather than re-listing them, so prose and set cannot drift again.</li>
* </ul>
*
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333).</strong>
* {@code FleetConfig} has 22 top-level record components: 5 cold, 11 deferred, 3 split, 3
* hot-excluded. Three of them are named nowhere in this file, and the reason is the same for all
* three: {@code placement}, {@code memberCredentials} and {@code memberLoginShell} are
* <strong>hot</strong> and correctly absent — all three are read live off {@code config.get()}
* (placement through the {@code CompositePeerLauncher} supplier the Hot bullet names;
* {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198, 205, 729}
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}), so a reload takes effect on the next
* spawn with no entry needed here.
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333);
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, again after
* {@code models:} was added as deferred, again for fleetd #422, which moved {@code models:}
* from deferred to hot-excluded once its on/off half was read live everywhere, and again after
* {@code leadRollover:} was added (fleetd #480).</strong>
* {@code FleetConfig} has 26 top-level record components: 5 cold, 13 deferred, 3 split, 5
* hot-excluded. Five of them are named nowhere in this file, and the reason is the same for all
* five: {@code placement}, {@code memberCredentials}, {@code memberLoginShell}, {@code models} and
* {@code leadRollover} are <strong>hot</strong> and correctly absent — all five are read live off
* {@code config.get()} (placement through the {@code CompositePeerLauncher} supplier the Hot bullet
* names; {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198,
* 205, 729} and {@code HerdrPeerLauncher#configuredMemberLoginShell}; {@code models} the same way,
* through the Hot bullet's {@code models:} paragraph; {@code leadRollover} through the Hot bullet's
* {@code leadRollover:} paragraph), so a reload takes effect on the next spawn (or, for
* {@code models}, the next reported status; for {@code leadRollover}, the next {@code open()}/
* {@code confirm()} call) with no entry needed here.
* {@code health} and {@code coordinator} used to be a third kind — <strong>undecided</strong>, not
* hot — until fleetd #330 added the <strong>split</strong> class above and gave them a home. A
* reload touching either used to report a bare "config reloaded", which under-claimed; now it names
@@ -173,6 +230,24 @@ import java.util.function.Supplier;
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*
* <p><strong>fleetd #474</strong> — {@link FleetConfig#validateAll()} is not the only gate startup
* runs before a config takes effect: {@code Fleetd.main} also calls {@code
* dev.ltms.fleet.mcp.CharterToolSurface#assertChartersNameOnlyRegisteredTools}, right after {@code
* cfg.validateAll()}, to refuse a charter that names an MCP tool the server does not register. That
* check cannot live inside {@link FleetConfig} — {@code CharterToolSurface} lives in the {@code mcp}
* package because the canonical tool set ({@code FleetTool}) does, and config is loaded before the
* MCP server exists, so {@code FleetConfig} must not gain a dependency on {@code mcp}. {@link
* #reload} cannot import {@code mcp} either, for the same reason applied one layer up: {@code
* dev.ltms.fleet.config} is loaded before {@code dev.ltms.fleet.mcp} exists, same as {@code
* FleetConfig}. So this class accepts the check as a {@code Consumer<FleetConfig>} —
* {@link #extraValidation} — supplied by whichever caller already sits at the seam that holds both
* a loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}. It is invoked
* inside the same try/catch as {@code fresh.validateAll()}, so a charter that would have refused to
* boot refuses a reload too, and keeps the running config exactly like any other {@code
* validateAll()} failure. A ref built through the two-argument constructor (every test fixture that
* does not care about this check, and {@link #fixed}) gets a no-op consumer, so nothing outside
* {@code Fleetd.main} needs to know this hook exists.
*/
public final class ConfigRef implements Supplier<FleetConfig> {
@@ -211,16 +286,33 @@ public final class ConfigRef implements Supplier<FleetConfig> {
* read {@link #COLD_KEYS} and {@link #SPLIT_KEYS}.
*/
static final Set<String> DEFERRED_KEYS = Set.of(
"guard", "worktreeRoot", "worktreeGroup", "primary", "configReload",
"guard", "worktreeRoot", "worktreeGroup", "memberSkills", "primary", "configReload",
"leadHeartbeat", "lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs",
"quarantineCooldownSeconds", "profiles");
"quarantineCooldownSeconds", "profiles", "idleSleepGuard");
private final Path path;
private final AtomicReference<FleetConfig> current;
private final Consumer<FleetConfig> extraValidation;
/** Equivalent to the three-argument constructor with a no-op {@code extraValidation}. */
public ConfigRef(Path path, FleetConfig initial) {
this(path, initial, cfg -> { });
}
/**
* @param extraValidation run on every {@link #reload} candidate, inside the same try/catch as
* {@code fresh.validateAll()} — see the class doc's fleetd #474 note.
* {@code Fleetd.main} passes {@code
* Fleetd::assertChartersNameOnlyRegisteredTools} (a package-private
* {@code FleetConfig -> void} adapter over {@code
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), so a reload
* runs the same gate startup does without this class depending on the
* {@code mcp} package.
*/
public ConfigRef(Path path, FleetConfig initial, Consumer<FleetConfig> extraValidation) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
this.extraValidation = Objects.requireNonNull(extraValidation, "extraValidation");
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
@@ -316,12 +408,18 @@ public final class ConfigRef implements Supplier<FleetConfig> {
fresh = FleetConfig.load(path);
// The same gate startup runs. A config that would have refused to boot must not be able
// to slip in through a reload — that is how a daemon ends up in a state it could never
// have started in, which is the hardest kind to debug.
fresh.validateAuthExposure();
fresh.validateLeadTabPrefixes();
fresh.validateSubscriptionProfiles();
fresh.validateCharters();
fresh.validateMembers();
// have started in, which is the hardest kind to debug. fleetd ticket "central allow-list
// of usable models" follow-up: this used to be six individual validateXxx() calls, and
// mutation testing found two of the six unpinned here even though startup pinned nothing
// at all — see FleetConfig#validateAll's javadoc for why the fix is one reflective call,
// not a longer hand-maintained list.
fresh.validateAll();
// fleetd #474: validateAll() does not cover everything startup refuses on — the charter
// tool-surface check (Fleetd.main, right after cfg.validateAll()) lives outside
// FleetConfig on purpose (see this class's doc) and is supplied here as extraValidation.
// Same try/catch as validateAll() above, on purpose: either failure must refuse the whole
// reload and keep the running config the same way.
extraValidation.accept(fresh);
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
@@ -402,6 +500,13 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.worktreeGroup(), fresh.worktreeGroup())) {
changed.add("worktreeGroup");
}
// fleetd #362: baked into the same GitWorktrees as worktreeRoot/worktreeGroup
// (Fleetd.java:251) and never rebuilt either — a reload that changes only memberSkills
// must be reported the same way, or a newly provisioned worktree keeps seeding from (or
// skipping) the old source directory with nothing telling the operator why.
if (!Objects.equals(old.memberSkills(), fresh.memberSkills())) {
changed.add("memberSkills");
}
// fleetd #326: Fleetd.java:506, 519, 520 read cfg.primary() only off the startup snapshot
// (PrimaryRegistry's pinned terminal, ReplyPushLoop's reminder cap and backoff) — neither is
// rebuilt on reload, so a changed value needs a restart. Note what it does NOT mean:
@@ -418,6 +523,14 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.configReload(), fresh.configReload())) {
changed.add("configReload");
}
// Fleetd.java reads cfg.idleSleepGuard() once, at startup, to decide whether to construct
// an IdleSleepGuard at all and wire SessionManager's onAcquire/onRelease hooks to it —
// neither is rebuilt on reload, so a running daemon keeps whatever this was at startup
// (armed or not) regardless of a later edit here. Not cold: nothing already-open goes
// inconsistent with the new value, an armed-or-not guard just keeps its original answer.
if (!Objects.equals(old.idleSleepGuard(), fresh.idleSleepGuard())) {
changed.add("idleSleepGuard");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
@@ -492,26 +605,35 @@ public final class ConfigRef implements Supplier<FleetConfig> {
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
+ "member's environment is read live on every spawn and already applied");
}
// fleetd #333: unlike health/coordinator above, most of `fleet:` (architects, developers,
// reviewers, charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, reviewers,
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
// with no restart note. Only fleet.leaders is frozen (Fleetd.java:281 reads
// cfg.fleet().leaders() off the startup snapshot to build both the LeadTabScanner's
// tab-label-to-name map, wired into CallerResolver.withLeadsAndMembers at Fleetd.java:620/624,
// and — when herdr answered — LeadLauncher(...).ensureLeads() at Fleetd.java:315, which
// auto-launches each lead up to its `instances` count; neither is rebuilt on reload). So this
// compares fleet.leaders alone, not the whole Fleet record: comparing the whole record would
// report "split" for a tabLabel-only or charters-only change that is actually fully hot,
// which is the over-claim mirror of the under-claim bug this class exists to prevent.
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
// its consumers, not just the one this comment used to name: CompositePeerLauncher reads it
// live for PLACEMENT through the () -> config.get().fleet() supplier named in the class doc's
// Hot bullet, and MemberRegistry separately reads it live for IDENTITY (which slot a spawn
// may bind to, AND what a slot already bound still grants) through its own instance of that
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
// LeadLauncher(...).ensureLeads() at Fleetd.java:315, which auto-launches each lead up to its
// `instances` count; neither is rebuilt on reload). So this compares fleet.leaders alone, not
// the whole Fleet record: comparing the whole record would report "split" for a tabLabel-only
// or architects-only change that is actually fully hot, which is the over-claim mirror of the
// under-claim bug this class exists to prevent.
if (!Objects.equals(leadersOf(old), leadersOf(fresh))) {
changed.add("fleet: fleet.leaders (each lead's tab, workspace, cwd, profile and "
+ "instances count) is read once at startup to build the LeadTabScanner's "
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
+ "lead; the rest of fleet: (architects, developers, reviewers, charters, "
+ "tabLabel) is read live through the supplier on CompositePeerLauncher and "
+ "already applied");
+ "lead; the rest of fleet: (developers, reviewers, charters, tabLabel) is read "
+ "live through the supplier on CompositePeerLauncher, and architects is read "
+ "live through that same supplier for placement AND through a separate supplier "
+ "on MemberRegistry for spawn-time identity — both already applied");
}
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
// every message here must be traceable to one of the split keys the class doc documents.
@@ -533,12 +655,16 @@ public final class ConfigRef implements Supplier<FleetConfig> {
* {@link #sameLaunchSettings} because they are read <em>live</em>, not baked in at spawn — see
* the class doc's <em>Hot</em> bullet. {@code weight} and {@code maxLoad} are read live by the
* placement policy on every spawn; {@code credentialId} is read live by
* {@code CompositePeerLauncher} and the CB-578 stage B exhaustion sink. Nothing else is
* {@code CompositePeerLauncher} and the CB-578 stage B exhaustion sink; {@code exhaustedPattern}
* (fleetd #446) is read live, cached by profile name, by {@code LiveExhaustedPatterns} — see
* that class's doc and the class doc's <em>Hot</em> bullet for the history (it used to be
* compared here, deferred, like its sibling {@code errorPattern} still is). Nothing else is
* excluded — see {@code sameLaunchSettingsComparesEveryProfileComponentOrExcludesIt} in
* {@code ConfigRefProfileCoverageTest}, which enumerates every {@code Profile} record component
* by reflection and fails the build if one is neither compared below nor named here.
*/
static final Set<String> LAUNCH_SETTINGS_EXCLUDED = Set.of("weight", "maxLoad", "credentialId");
static final Set<String> LAUNCH_SETTINGS_EXCLUDED =
Set.of("weight", "maxLoad", "credentialId", "exhaustedPattern");
/**
* Whether two versions of a profile would launch a peer identically.
@@ -575,13 +701,15 @@ public final class ConfigRef implements Supplier<FleetConfig> {
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription())
// CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern())
// fleetd #446: exhaustedPattern moved to LAUNCH_SETTINGS_EXCLUDED — it is now read
// live, cached by profile name, through LiveExhaustedPatterns (see that class's doc
// and ConfigRef's class doc Hot bullet), so it must NOT be compared here any more: a
// reload that changes only exhaustedPattern must report "config reloaded", not
// "these changes need a restart".
// fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error
// pattern map at startup (see BackendErrorPatternLookup wiring), the same way
// exhaustedPattern is — a reload never re-reads it either.
// pattern map at startup (see BackendErrorPatternLookup wiring) — unlike its sibling
// exhaustedPattern above (fleetd #446), a reload still never re-reads it; fleetd #446
// scoped errorPattern out on purpose (see the class doc's Hot bullet).
&& Objects.equals(a.errorPattern(), b.errorPattern())
// fleetd #323 instance 1: ideProjectDir and ideOpenCommand are read at spawn off the
// same frozen profile map as ideMcpUrl above (ClaudeCodeLauncher.java:267/269,
@@ -16,11 +16,15 @@ import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.lang.reflect.InvocationTargetException;
import java.lang.reflect.Method;
import java.lang.reflect.Modifier;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Collections;
import java.util.Comparator;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
@@ -66,12 +70,14 @@ import java.util.regex.PatternSyntaxException;
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
* @param quarantineCooldownSeconds the BASE cooldown a credential is quarantined for after a
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
* needs a restart, and a quarantine already running keeps whatever cooldown was
* live when it started.
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Since fleetd #466 this is
* only the first occurrence's length — a credential quarantined again shortly
* after this cooldown ends backs off further, up to a ceiling; see {@code
* BackendQuarantine}'s class doc. Baked once into the {@code BackendQuarantine}
* built at startup, so it is DEFERRED: changing it needs a restart, and a
* quarantine already running keeps whatever cooldown was live when it started.
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
* credentials a spawned member's pane inherits. {@code null} (the block
* omitted) blocks nothing — see {@link MemberCredentials}.
@@ -106,6 +112,36 @@ import java.util.regex.PatternSyntaxException;
* When {@code memberHerdrSocket} is NOT configured this field is never
* consulted at all; fleetd keeps reading its own {@code $SHELL}, exactly as
* before this field existed.
* @param memberSkills fleetd #362: nullable directory of skill folders (each a subdirectory
* holding a {@code SKILL.md}, the same shape as this repo's own {@code
* .claude/skills/}) copied into every provisioned worktree's {@code
* .claude/skills/}, so a member spawned against ANY repo — not only one that
* already ships its own copy — can load a bridge skill such as {@code
* implementer}. {@code null}/blank ⇒ off: no worktree is touched beyond
* today's behaviour. A skill folder the target repo already carries is never
* overwritten — see {@link dev.ltms.fleet.session.GitWorktrees}. Claude Code
* members only; an opencode member's equivalent lives under a different path
* ({@code .opencode/agent}) and is not covered by this key. Every non-hidden
* subdirectory of this directory is copied wholesale, with no per-file
* allowlist — do not park scratch files or drafts alongside the real skill
* folders, they will be copied into every provisioned worktree too.
* @param idleSleepGuard opt-in-by-default: hold an OS-level assertion against idle sleep while at
* least one member is live, so an unattended host does not sleep out from
* under a member's long turn. {@code null} (the block omitted) behaves the
* same as an explicit {@code enabled: true}; set {@code enabled: false} to
* turn it off. See {@link dev.ltms.fleet.power.IdleSleepGuard}.
* @param models central allow-list of models any {@code profiles:} entry may name. {@code
* null} or an empty {@code allow:} ⇒ off: {@link #validateModels()} checks
* nothing and every existing config keeps working exactly as it does today.
* When non-empty, a profile whose {@code model:} is not one of {@link
* Models#ids()} fails config load, naming both the model and the profile —
* see {@link #validateModels()}. This block decides what may be CONFIGURED;
* fleetd #422 added the separate on/off question — whether a configured model
* may be spawned onto RIGHT NOW ({@link Models.ModelEntry#enabled}) — enforced
* live at spawn by {@code CompositePeerLauncher}, not here. See {@link Models}.
* @param leadRollover opt-in lead rollover (fleetd #480): {@code null} ⇒ off, and no
* {@code dev.ltms.fleet.lead.LeadRollover} is constructed at all — an upgraded
* daemon never clears a lead's pane on its own initiative. See {@link LeadRollover}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
@@ -130,7 +166,68 @@ public record FleetConfig(
MemberCredentials memberCredentials,
Coordinator coordinator,
String worktreeGroup,
String memberLoginShell) {
String memberLoginShell,
String memberSkills,
IdleSleepGuard idleSleepGuard,
Models models,
LeadRollover leadRollover) {
/** Back-compat form before the {@code leadRollover:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard,
Models models) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, idleSleepGuard, models, null);
}
/** Back-compat form before the {@code models:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, idleSleepGuard, null);
}
/** Back-compat form before the {@code idleSleepGuard:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, null);
}
/** Back-compat form before the {@code memberSkills} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, null, null);
}
/** Back-compat form before the {@code memberLoginShell} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
@@ -141,7 +238,7 @@ public record FleetConfig(
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup, null);
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup, null, null);
}
/** Back-compat form before the {@code worktreeGroup} key was added. */
@@ -348,6 +445,12 @@ public record FleetConfig(
* {@code null}/blank ⇒ the classification never fires for this profile and
* today's completion-fallback behaviour is unchanged. Every backend words
* its refusal differently, so this is config, never a vendor string in code.
* Read live off the current config through {@code LiveExhaustedPatterns},
* cached by profile name (fleetd #446), so it is HOT: an operator can arm
* or disarm this profile's usage-limit detection by editing this key and
* reloading, with no restart. Before fleetd #446 this was compiled once at
* daemon startup, the same deferred shape its sibling {@code errorPattern}
* (below) still has.
* @param credentialId the shared account this profile authenticates as (CB-578 stage B). Two
* or more profiles setting the <em>same</em> non-blank value are quarantined
* together by one {@code exhaustedPattern} classification on any one of
@@ -888,12 +991,18 @@ public record FleetConfig(
* {@code null} ⇒ kept as {@code null} (no self id configured).
* @param prefetch the consumer's {@code basicQos} prefetch count. {@code null}/non-positive ⇒
* {@link LeadMailbox#DEFAULT_PREFETCH}.
* @param peers fleetd #361: the coord-ids the operator declares as this daemon's peers — the
* daemon never guesses who else exists. {@code fleet_list} reports each one's
* live reachability. Blank entries are dropped; {@code null} ⇒ an empty list, so
* a config written before this field existed still parses unchanged.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Coordinator(String uri, String uriEnv, String selfId, Integer prefetch) {
public record Coordinator(String uri, String uriEnv, String selfId, Integer prefetch, List<String> peers) {
public Coordinator {
selfId = (selfId == null || selfId.isBlank()) ? null : selfId;
peers = peers == null ? List.of()
: peers.stream().filter(p -> p != null && !p.isBlank()).toList();
}
/** True when a {@code uriEnv} is configured by name, whether or not its variable resolves. */
@@ -1220,6 +1329,104 @@ public record FleetConfig(
}
}
/**
* Opt-in lead rollover (fleetd #480): a lead that decides it is ready to be replaced writes a
* handover file, then asks fleetd to clear its own pane and bootstrap a fresh session against
* that file. Config + a pure decision/verification layer only — see
* {@code dev.ltms.fleet.lead.LeadRollover} for the executor this block feeds, and
* {@code fleet_*} tool wiring is a later ticket.
*
* <p>Deliberately opt-in ({@code null} ⇒ off, exactly like {@code leadHeartbeat:}): absent this
* block, {@code Fleetd.java} never constructs a {@code LeadRollover} object at all, so an
* upgraded daemon cannot silently acquire the ability to clear the lead's own pane. Even once
* present, nothing but an explicit {@code confirm()} call can ever roll a pane — there is no
* timer, heartbeat or timeout anywhere in this feature that fires one on its own; see that
* class's javadoc.
*
* <p><strong>Hot, not deferred</strong> (see {@code ConfigRef}'s class doc): every field below
* is read live, through a {@code Supplier<LeadRollover>} the same {@code () -> config.get().x()}
* shape {@code fleet}/{@code placement}/{@code models} already use, so an edit to any of the
* six fields takes effect on the very next {@code open()}/{@code confirm()} call once the block
* has been present since startup — nothing here is captured into a frozen field the way {@code
* leadHeartbeat}'s {@code idleAfterNanos}/{@code backoffMs}/{@code quietNudgeCap} are. The one
* restart-only edge left is structural, not a stale value: {@code Fleetd.java} decides whether
* to construct the {@code LeadRollover} object at all off the startup snapshot, the same gate
* {@code leadHeartbeat:} uses, so a block ADDED where it was absent at boot needs a restart
* before anything exists to call — the same fact already true of adding a brand-new
* {@code profiles:} entry.
*
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
* {@code confirm()} validates every gate and schedules the roll. {@code
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
* that SAME pane to report an injectable state again — i.e. for the calling turn to actually
* end — before it sends {@code /clear} at all. If that wait times out, no {@code /clear} is
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
* bounds the SECOND wait, for the pane to re-settle AFTER {@code /clear} has already gone out.
*
* @param handoverPath required when this block is present — where the handover file a fresh
* lead session reads must live. There is no sane non-null default for an
* operator-specific path, so a present block with a {@code null}/blank
* {@code handoverPath} is refused at config load; see
* {@link #validateLeadRollover()}. May be relative: {@code
* dev.ltms.fleet.lead.LeadRollover#open} resolves a relative path against
* the CALLING lead's configured {@code fleet.leaders.<name>.cwd}, falling
* back to the daemon's own working directory ({@code
* System.getProperty("user.dir")}) when that lead has none configured —
* never against whatever directory the daemon process happens to have been
* started in for its own sake. An absolute path is used unchanged.
* @param requireOperatorConfirm default {@code true} — {@code confirm()} refuses unless the
* caller also passes {@code operatorConfirmed: true}. Set {@code false} to
* let the three handover-file checks alone gate the roll.
* @param maxDocAgeSeconds default 3600 — refuse a handover file whose modified time is older
* than this many seconds, so a stale leftover from an earlier rollover
* attempt can never be mistaken for a fresh one.
* @param turnSettleSeconds default 20 — bound on how long the deferred roll waits for the
* CALLING lead's own turn to end (its pane to report injectable again)
* before sending {@code /clear} at all. See the paragraph above.
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
* report an injectable state again after {@code /clear} before giving up. A
* roll that times out here never sends {@code bootstrapText}.
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
* lead's pane once it settles after {@code /clear}, telling the fresh
* session where to read the handover and carry on. Left {@code null} here
* when the operator configures none: the default sentence cannot be built
* at construction time because it must name the path AFTER {@code
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code
* handoverPath} against the calling lead's workspace, which this record has
* no way to know — see {@link #bootstrapTextFor(String)}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
Integer maxDocAgeSeconds, Integer turnSettleSeconds,
Integer clearSettleSeconds, String bootstrapText) {
public LeadRollover {
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 20 : turnSettleSeconds;
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
}
/**
* The text actually sent to the lead's pane once it settles after {@code /clear}: the
* operator's configured {@link #bootstrapText} when one is set, otherwise the default
* sentence built from {@code resolvedHandoverPath}.
*
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
* #open} already resolved — never the raw configured {@link
* #handoverPath}, which may still be relative and would name a
* directory the fresh lead session's own pane cannot resolve
*/
public String bootstrapTextFor(String resolvedHandoverPath) {
return bootstrapText != null
? bootstrapText
: "Fresh lead session: read the handover file at " + resolvedHandoverPath
+ " and carry on from there.";
}
}
/**
* Watch {@code fleetd.yaml} and re-read it when it changes (CB-559).
*
@@ -1245,6 +1452,145 @@ public record FleetConfig(
}
}
/**
* Hold an OS-level assertion against idle sleep while at least one member is live (see
* {@link dev.ltms.fleet.power.IdleSleepGuard}).
*
* <p>Unlike most opt-in blocks in this file, this one defaults to <em>on</em>: an unattended
* host idle-sleeping mid-turn is a correctness problem (a dropped AMQP link, a frozen member),
* not a convenience, so the safer default is armed. An operator who wants the previous
* behaviour (no assertion held, ever) sets {@code enabled: false} explicitly.
*
* @param enabled {@code false} turns the guard off; {@code null} (the block omitted
* entirely) or {@code true} leaves it on
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record IdleSleepGuard(Boolean enabled) {
public boolean isEnabled() {
return !Boolean.FALSE.equals(enabled);
}
}
/**
* Central allow-list of models any {@code profiles:} entry may name (fleetd ticket: "a central
* allow-list of usable models"). Nothing before this block checked a profile's {@code model:}
* against anything — it was a free-form string handed straight to the backend adapter, and a
* withdrawn or misspelled name failed silently (opencode falls back to a default model rather
* than erroring on an unknown {@code -m}) rather than at config load, where a mistake is cheap.
*
* <p><b>Absent or empty {@code allow:} is "off"</b>, on purpose: this is a large deployment
* with gitignored {@code fleetd.yaml} on more than one host, and a change that forced every
* operator to enumerate their models before the daemon would start would break every one of
* them on upgrade. See {@link FleetConfig#validateModels()}, which is where the allow-list is
* actually enforced, at config load.
*
* <p><b>The list is the authority; a {@code profiles:} entry is checked against it, never the
* other way around.</b> Adding or editing a {@code profiles:} entry cannot, by itself, widen
* the set of permitted models — only editing {@code models.allow:} itself can. This is the
* invariant the ticket asked for: the two blocks are validated in one direction only.
*
* <p><b>Spawn-time enforcement (fleetd #422, units 2+3) lives outside this record</b> —
* {@code CompositePeerLauncher.enforceModelEnabled} and its candidate-set filter read this
* block LIVE (through the same kind of supplier {@code weight}/{@code maxLoad} already use),
* so the on/off state below is hot: no restart needed. This block itself still only decides
* what may be CONFIGURED (membership in {@link #allow}); {@link ModelEntry#enabled} decides
* whether a member of that list is currently spawnable. The two questions are deliberately
* separate — see {@link ModelEntry}'s javadoc for why turning a model off must never mean
* removing it from {@link #allow}.
*
* @param allow the permitted models, each its own {@link ModelEntry} rather than a bare
* string — see that record's javadoc for why. {@code null}/empty ⇒ the block is
* treated as absent: {@link FleetConfig#validateModels()} checks nothing.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Models(List<ModelEntry> allow) {
public Models {
allow = (allow == null) ? List.of() : List.copyOf(allow);
}
/**
* One permitted model, named as a record rather than a bare string on purpose: this ticket
* (fleetd #422) is the "later unit" the original comment here predicted — it hangs an on/off
* state ({@link #enabled}) off each entry, and a bare {@code List<String>} could not have
* grown that field without changing the YAML shape underneath every operator who already
* wrote one. {@link #model()} is intentionally a single flat, opaque-string namespace — a
* bare Claude id ({@code claude-sonnet-5}) and an opencode provider-prefixed id
* ({@code openai/gpt-5.6-terra}) both fit it unchanged, because {@link
* FleetConfig#validateModels()} only ever compares a profile's {@code model:} value against
* this string for exact equality; it never parses a provider prefix or branches on a
* profile's {@code kind:}.
*
* @param model the model id exactly as a {@code profiles:} entry's {@code model:} field
* would name it
* @param enabled {@code false} turns spawning onto this model off; {@code null} (the field
* omitted — every config written before fleetd #422 is this shape) or
* {@code true} leaves it on. Turning a model off must NEVER remove it from
* {@link Models#allow} — {@link FleetConfig#validateModels()} checks
* <em>membership</em> only, never the on/off state, so an off entry stays a
* valid thing for a {@code profiles:} entry to name; only
* {@code CompositePeerLauncher}'s spawn-time gate reads {@link #enabled}.
* Collapsing the two — turning a model off by deleting its {@code allow:}
* entry — would make {@link FleetConfig#validateModels()} refuse the whole
* config reload the moment a still-configured profile names it, which is
* exactly the restart-to-flip-a-switch problem this field exists to avoid.
* <p>A model id can be named by more than one {@code profiles:} entry (e.g.
* {@code deepseek-v4-flash} backs both {@code local} and {@code
* local-direct} in the live config) — turning it off disables every profile
* that names it, on purpose: the model is what a subscription's rate limit
* actually constrains, not any one profile alias for it.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record ModelEntry(String model, Boolean enabled) {
public ModelEntry {
model = (model == null || model.isBlank()) ? null : model.trim();
}
/**
* Back-compat form before {@link #enabled} was added (fleetd #422) — the model is
* unconditionally on, exactly as every {@code ModelEntry} behaved before this field
* existed. Keeps pre-#422 call sites (and any YAML that omits {@code enabled:})
* compiling and behaving identically.
*/
public ModelEntry(String model) {
this(model, null);
}
/** {@code true} unless {@link #enabled} is explicitly {@code false} — absent means on. */
public boolean isEnabled() {
return !Boolean.FALSE.equals(enabled);
}
}
/** {@link #allow}'s model ids, as a set for membership checks. Blank/null entries are dropped. */
public Set<String> ids() {
Set<String> ids = new java.util.LinkedHashSet<>();
for (ModelEntry e : allow) {
if (e != null && e.model() != null) {
ids.add(e.model());
}
}
return Collections.unmodifiableSet(ids);
}
/**
* Model ids currently turned off (fleetd #422: {@link ModelEntry#isEnabled()} {@code
* false}). Read live by {@code CompositePeerLauncher}'s spawn gate and by {@code
* fleet_profiles}/{@code GET /profiles} — both must read this same accessor off the same
* live config so the two surfaces cannot disagree about which model is off (the fleetd
* #404 lesson: a status field must read the source the behaviour reads).
*/
public Set<String> offIds() {
Set<String> off = new java.util.LinkedHashSet<>();
for (ModelEntry e : allow) {
if (e != null && e.model() != null && !e.isEnabled()) {
off.add(e.model());
}
}
return Collections.unmodifiableSet(off);
}
}
/**
* The terminal → lead-name map seeded from the legacy singular {@code primary:} pin (CB-530).
*
@@ -1495,7 +1841,8 @@ public record FleetConfig(
"bind", "herdrSocket", "memberHerdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell");
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills",
"idleSleepGuard", "models", "leadRollover");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
@@ -2172,9 +2519,27 @@ public record FleetConfig(
// memberLoginShell is left as-is (fleetd #213), like worktreeGroup: null/blank is "not
// configured", and there is no sane non-null default — a member's login shell is
// operator-specific and only meaningful when memberHerdrSocket is also set.
// memberSkills is left as-is (fleetd #362), like worktreeGroup/memberLoginShell: null/blank
// is "off", and there is no sane non-null default — the daemon may not even run from a
// checkout that ships its own .claude/skills/.
// idleSleepGuard is left as-is, like leadHeartbeat/configReload above, but for the opposite
// reason: it is on by default already (its own isEnabled() treats null the same as
// enabled: true — see its javadoc), so defaulting the block here would change nothing a
// reader observes and would only obscure that "block omitted" and "block present and
// enabled" are deliberately the same outcome.
// models is left as-is, like broker/primary/coordinator above: null/empty is "off", and an
// absent block must validate nothing (see Models's javadoc) — defaulting it here to an
// empty Models would be a no-op for validateModels() either way, since an empty allow-list
// already means "check nothing", so there is nothing to gain and one more null check to
// avoid by leaving it exactly as configured.
// leadRollover is left as-is, like leadHeartbeat above: null is "off", and LeadRollover's
// own compact constructor defaults the fields of a block that IS present. Defaulting it
// here would construct a LeadRollover object (via Fleetd.java's presence gate) for every
// config that never mentioned it.
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell);
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills,
idleSleepGuard, models, leadRollover);
}
/**
@@ -2257,6 +2622,26 @@ public record FleetConfig(
+ "lead tabs cannot be confused.");
}
/**
* Reject a present {@code leadRollover:} block with no (or a blank) {@code handoverPath}
* (fleetd #480). There is no sane non-null default for an operator-specific file path, unlike
* every other field on {@link LeadRollover}, which {@link LeadRollover}'s own compact
* constructor already defaults — so this is the one field that must be refused at load rather
* than silently defaulted to something that would never match a real handover file.
*
* @throws IllegalStateException when {@code leadRollover} is present but {@code handoverPath}
* is {@code null} or blank
*/
public void validateLeadRollover() {
if (leadRollover == null) {
return;
}
if (leadRollover.handoverPath() == null || leadRollover.handoverPath().isBlank()) {
throw new IllegalStateException("refusing to start: leadRollover.handoverPath is "
+ "required when leadRollover: is present.");
}
}
/** Case-insensitive prefix test that tolerates a null/blank label. */
private static boolean startsWithIgnoreCase(String label, String prefix) {
if (label == null || prefix == null || prefix.isBlank()) {
@@ -2394,6 +2779,118 @@ public record FleetConfig(
}
}
/**
* Reject a {@code profiles:} entry whose {@code model:} is not on the configured {@link
* #models} allow-list.
*
* <p>Absent or empty {@code models.allow:} validates nothing — see {@link Models}'s javadoc:
* an existing config with no such block must keep working exactly as it does today. Once the
* operator declares at least one entry, every profile's {@code model:} (when set — a profile
* may legitimately leave it {@code null}, e.g. a {@code subscription: true} profile relying on
* the account's own default) must equal one of {@link Models#ids()} exactly. The comparison is
* a flat string match: a bare Claude id and an opencode {@code provider/model} id are both
* just opaque strings here, so nothing here needs to know which {@code kind:} a profile runs.
*
* <p>The check runs one direction only, by construction: it reads {@link #profiles} and
* {@link #models}, and only ever adds to {@code bad} when a profile's model is missing from
* the allow-list. Nothing here can be satisfied by widening a {@code profiles:} entry — only
* editing {@code models.allow:} itself changes what passes. That is the invariant the ticket
* asked for: the list is the authority, profiles are checked against it.
*
* @throws IllegalStateException when any profile names a model outside the configured
* allow-list, naming both the model and the profile that wanted it
*/
public void validateModels() {
if (models == null || models.allow().isEmpty()) {
return;
}
Set<String> allowed = models.ids();
List<String> bad = new ArrayList<>();
profiles.forEach((name, p) -> {
String model = p.model();
if (model != null && !model.isBlank() && !allowed.contains(model)) {
bad.add("profile '" + name + "' names model '" + model + "', which is not in "
+ "models.allow: (have: " + allowed + ").");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
/**
* Runs every validator this class declares — found by reflection, not by name.
*
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
* that although each of the six validators above was well pinned on its own, nothing proved
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
* per caller; a hand-maintained list of six names here would have the exact same defect its
* own javadoc would warn against: the seventh validator someone adds next month has no reason
* to be added to it. So this method does not name any validator. It sweeps {@link
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
* "validate"} (excluding itself) and invokes every one it finds, via {@link
* #invokeAllValidators}. A new {@code validateXxx()} method is therefore wired into both
* callers the moment it is written — there is no second step to forget, and so no state in
* which it silently never runs.
*
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
* the six individually — see the comments at those two call sites for why
* each must run it.
*
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
* always names the same one first, on every run.
*
* @throws IllegalStateException (or whatever unchecked exception a validator itself throws),
* propagated unchanged from the first validator, in that order,
* that finds a problem
*/
public void validateAll() {
invokeAllValidators(this);
}
/**
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
* seventh validator to this class just to exercise that claim. See {@code
* FleetConfigValidateAllTest} for that proof.
*
* @param target an object whose public, no-argument, {@code void} methods named {@code
* validateXxx} (any name starting with {@code "validate"}, excluding {@code
* validateAll} itself) should all run, in alphabetical-by-name order
*/
static void invokeAllValidators(Object target) {
List<Method> methods = new ArrayList<>();
for (Method m : target.getClass().getMethods()) {
if (Modifier.isPublic(m.getModifiers())
&& m.getParameterCount() == 0
&& m.getReturnType() == void.class
&& m.getName().startsWith("validate")
&& !m.getName().equals("validateAll")) {
methods.add(m);
}
}
methods.sort(Comparator.comparing(Method::getName));
for (Method m : methods) {
try {
m.invoke(target);
} catch (InvocationTargetException e) {
Throwable cause = e.getCause();
if (cause instanceof RuntimeException re) {
throw re;
}
if (cause instanceof Error err) {
throw err;
}
throw new IllegalStateException("validator " + m.getName() + " failed", cause);
} catch (IllegalAccessException e) {
throw new IllegalStateException("cannot invoke validator " + m.getName(), e);
}
}
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
@@ -38,15 +38,46 @@ public final class FleetHealthMonitor {
*/
static final long ASK_LAPSE_RECHECK_DELAY_SECONDS = 120;
/**
* fleetd #386: {@code System.nanoTime()} (or whatever {@link #clock} is) does not advance while
* macOS sleeps, so a raw {@code nowNanos - lastActivityAtNanos} comparison freezes with the
* host and can never cross {@link #workingSuspectAfterNanos}. This is a second, wall-clock
* source used ONLY inside the stall check ({@link #stallElapsedNanos}) to detect and correct
* for that freeze. Nothing else in this class reads it — every other decision (readiness grace,
* the fault classification itself) stays exactly on {@link #clock}, as the ticket requires.
*/
private static final LongSupplier DEFAULT_REALTIME_CLOCK =
() -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis());
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
private final MessageService messages;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final LongSupplier realtimeClock;
private final long intervalSeconds;
private final long tickIntervalNanos;
private final long workingSuspectAfterNanos;
private final BiConsumer<String, String> failTarget;
private final Map<String, HealthPrior> priors = new HashMap<>();
/**
* fleetd #386 clock-drift bookkeeping. {@code haveClockBaseline}/{@code lastTickMonoNanos}/
* {@code lastTickRealNanos} track the previous tick's pair of readings so each new tick can
* measure how far the two clocks moved apart since then. {@code accumulatedDriftNanos} is the
* running total of every such divergence observed since this monitor started (never decreases —
* the monotonic clock can only lag real time, never lead it). {@code busyDriftBaselineNanos}/
* {@code busyBaselineActivityNanos} record, per target, the value of {@code accumulatedDriftNanos}
* at the moment this monitor first saw that target's CURRENT {@code lastActivityAtNanos} while
* BUSY — so {@link #stallElapsedNanos} adds back only the drift observed DURING this BUSY span,
* never drift from a sleep that happened before the member went busy. All five fields are touched
* only from {@code tick()}, like {@link #priors}.
*/
private boolean haveClockBaseline = false;
private long lastTickMonoNanos;
private long lastTickRealNanos;
private long accumulatedDriftNanos = 0;
private final Map<String, Long> busyDriftBaselineNanos = new HashMap<>();
private final Map<String, Long> busyBaselineActivityNanos = new HashMap<>();
/**
* The live classification per member, and the only one of this class's three maps that more
* than one scheduler task touches. {@code tick} writes it (and prunes it to the roster);
@@ -88,12 +119,29 @@ public final class FleetHealthMonitor {
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
long workingSuspectAfterSeconds, BiConsumer<String, String> failTarget) {
this(agents, roster, messages, scheduler, clock, DEFAULT_REALTIME_CLOCK, intervalSeconds,
workingSuspectAfterSeconds, failTarget);
}
/**
* @param realtimeClock fleetd #386: a wall-clock nanosecond source (e.g.
* {@code System.currentTimeMillis()} converted to nanos) that keeps
* advancing while {@code clock} is frozen by a host sleep. Used only to
* correct the stall check — see the class-level javadoc on the
* clock-drift fields.
*/
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, LongSupplier realtimeClock,
long intervalSeconds, long workingSuspectAfterSeconds,
BiConsumer<String, String> failTarget) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.realtimeClock = Objects.requireNonNull(realtimeClock, "realtimeClock");
this.intervalSeconds = intervalSeconds;
this.tickIntervalNanos = TimeUnit.SECONDS.toNanos(intervalSeconds);
this.workingSuspectAfterNanos = TimeUnit.SECONDS.toNanos(workingSuspectAfterSeconds);
this.failTarget = Objects.requireNonNull(failTarget, "failTarget");
}
@@ -123,6 +171,7 @@ public final class FleetHealthMonitor {
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
long nowNanos = clock.getAsLong();
long driftBeforeThisTick = observeClockDrift(nowNanos);
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
@@ -133,7 +182,7 @@ public final class FleetHealthMonitor {
&& session.state() != MemberSession.State.SPAWNING;
boolean readinessGraceElapsed = nowNanos - session.spawnedAtNanos() >= READINESS_GRACE_NANOS;
boolean stalled = session.state() == MemberSession.State.BUSY
&& nowNanos - session.lastActivityAtNanos() >= workingSuspectAfterNanos;
&& stallElapsedNanos(session, nowNanos, driftBeforeThisTick) >= workingSuspectAfterNanos;
// CB-643: the three message-layer facts CB-640 published. Read them here rather than
// leaving them false — that constant is what made 8 of the 9 fault states dead.
boolean queuedDelivery = messages.hasQueuedDelivery(session.terminalId());
@@ -151,6 +200,8 @@ public final class FleetHealthMonitor {
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
orphanStreaks.keySet().retainAll(current);
busyDriftBaselineNanos.keySet().retainAll(current);
busyBaselineActivityNanos.keySet().retainAll(current);
} catch (Throwable error) {
// Any unclassified collection failure must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
@@ -175,6 +226,65 @@ public final class FleetHealthMonitor {
return streak >= ORPHAN_CONFIRM_TICKS;
}
/**
* fleetd #386: compare this tick's monotonic and real-time readings against the previous
* tick's, and fold any positive divergence into {@link #accumulatedDriftNanos} (a ratchet — it
* never decreases, since the monotonic clock can only fall behind real time, never ahead of
* it). Logs once, at WARN, when that single tick's divergence exceeds one full tick interval —
* the signature of a host that slept between the two ticks (a tick literally cannot run while
* the process itself is suspended, so the whole sleep duration lands inside one tick's gap).
*
* @return {@link #accumulatedDriftNanos} as it stood BEFORE this tick's divergence was folded
* in — the baseline {@link #stallElapsedNanos} needs when a target is observed BUSY
* for the first time this tick, so a sleep that happened before this member went busy
* is not attributed to it.
*/
private long observeClockDrift(long nowNanos) {
long nowRealNanos = realtimeClock.getAsLong();
long driftBeforeThisTick = accumulatedDriftNanos;
if (haveClockBaseline) {
long monoDelta = nowNanos - lastTickMonoNanos;
long realDelta = nowRealNanos - lastTickRealNanos;
long tickDrift = realDelta - monoDelta;
if (tickDrift > tickIntervalNanos) {
log.warn("fleet health: the monotonic clock did not advance for about {}s that the "
+ "real clock did since the last tick (host likely slept); the stall "
+ "detector could not see that time", TimeUnit.NANOSECONDS.toSeconds(tickDrift));
}
if (tickDrift > 0) {
accumulatedDriftNanos = driftBeforeThisTick + tickDrift;
}
}
lastTickMonoNanos = nowNanos;
lastTickRealNanos = nowRealNanos;
haveClockBaseline = true;
return driftBeforeThisTick;
}
/**
* fleetd #386: {@code nowNanos - lastActivityAtNanos} alone freezes across a host sleep, since
* both come from the monotonic {@link #clock}. This adds back the real-time drift observed
* since this BUSY span started — not the monitor's whole lifetime, so a sleep that happened
* before this member went busy never leaks into its stall reading (see the class-level javadoc
* on the drift fields). The baseline resets whenever {@code lastActivityAtNanos} changes (a new
* turn) or the member is not currently BUSY.
*/
private long stallElapsedNanos(MemberSession session, long nowNanos, long driftBeforeThisTick) {
String target = session.terminalId();
if (session.state() != MemberSession.State.BUSY) {
busyDriftBaselineNanos.remove(target);
busyBaselineActivityNanos.remove(target);
return nowNanos - session.lastActivityAtNanos();
}
Long baselineActivity = busyBaselineActivityNanos.get(target);
if (baselineActivity == null || baselineActivity != session.lastActivityAtNanos()) {
busyBaselineActivityNanos.put(target, session.lastActivityAtNanos());
busyDriftBaselineNanos.put(target, driftBeforeThisTick);
}
long driftSinceBusyStart = accumulatedDriftNanos - busyDriftBaselineNanos.get(target);
return (nowNanos - session.lastActivityAtNanos()) + driftSinceBusyStart;
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
@@ -3,7 +3,7 @@ package dev.ltms.fleet.herdr;
import com.fasterxml.jackson.databind.JsonNode;
/**
* Client face onto the herdr daemon (protocol 14, herdr 0.7.0).
* Client face onto the herdr daemon (protocol 19, herdr 0.8.0).
*
* <p>This is the ONLY thing in {@code fleetd} that speaks to herdr. Every method
* maps to a herdr JSON-RPC call over its Unix domain socket. Requests are
@@ -8,7 +8,7 @@ import com.fasterxml.jackson.databind.node.ObjectNode;
import java.nio.charset.StandardCharsets;
/**
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 14).
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 19).
*
* <p>Split out from the socket so the framing rules — the ones that actually bit us
* during the spike (id MUST be a string; response carries {@code result} or
@@ -5,6 +5,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.Locale;
import java.util.Map;
@@ -53,17 +54,40 @@ import java.util.function.Supplier;
* daemon and found again by this scan. The trust direction above is unaffected — fleetd writing a
* name for a lead it just started is not a pane promoting itself — but <em>staleness</em> becomes
* real: a label left behind by a session that has since died would read as a live lead forever.
* This scanner does not solve that (its job is naming, and a stale name costs nothing here); the
* launcher does, by requiring a running agent in the tab before it counts the lead as live. If you
* ever make a decision that <em>removes</em> something based on this map, add the same check.
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code FleetConfig.validateLeadTabPrefixes} rather than documented here.
*
* <p><strong>fleetd #359 — the staleness check this class used to skip.</strong> This used to say
* "a stale name costs nothing here" and leave liveness to {@code LeadLauncher}, on the theory that
* naming and removing are different decisions. That was wrong: {@code dev.ltms.fleet.msg.LeadCoordLoop}
* makes exactly the kind of removal decision the old javadoc warned about, by reading this map to
* pick which pane a peer message goes into — and a stale entry there is not free. On a host where
* {@code fleetd} had restarted more than once, a labelled-but-dead tab from a previous life was
* reported right alongside the live one; {@code LeadCoordLoop.resolveLocalLead()} saw more than one
* candidate and refused to guess (safe), but the fix the daemon's own WARN suggests — name a lead
* after {@code coordinator.selfId} — stops being safe once two tabs can share a label: step 1 of
* that resolution picks whichever matching entry it finds first, which can be the dead one, and
* typing a peer's message into a dead shell does not fail — it is silently gone instead of merely
* held. {@link #scan()} now cross-checks every labelled tab against {@code agent.list} (the same
* signal {@code LeadLauncher.countLeads} already trusts for the same purpose) and drops any tab
* with no agent running in it, so a dead tab is never in the map for a caller to pick at all.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
* the previous answer instead of emptying it — a herdr hiccup must not silently demote a live lead
* mid-session.
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan (herdr
* throws) keeps the previous answer instead of emptying it — a herdr hiccup must not silently
* demote a live lead mid-session.
*
* <p><strong>fleetd #359 review, finding 2 — a successful-but-wrong scan is the same hazard.</strong>
* The catch above only fires when a call throws. It does nothing for a call that returns 200 with an
* incomplete answer — exactly what the ticket's own evidence showed {@code agent.list} can do. Once
* this class started trusting that signal, an empty read would otherwise get cached as fact and
* silently drop a lead {@code CallerResolver} had, until then, correctly resolved — turning it into a
* {@code Role.WORKER}, which refuses every orchestration call. So a terminal this class already
* reported as live is not dropped the first time {@code agent.list} loses it: {@link #scan()} grants
* it one grace scan (see {@code gracedTerminals}) and only drops it if a <em>later</em> scan still
* finds no agent. A terminal never reported live before gets no grace — that would weaken the
* original #359 fix itself, which this class's own test suite already pins.
*/
public final class LeadTabScanner implements Supplier<Map<String, String>> {
@@ -79,6 +103,15 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
private long scannedAtNanos;
private boolean everScanned;
/**
* Terminals currently on their one grace scan: {@code cached} reported them live, the most
* recent {@link #scan()} found no agent for them, and they were re-included anyway. Cleared for
* a terminal the instant it is seen live again; a terminal still here on the <em>next</em> scan
* is finally dropped. Scoped separately from {@link #cached} so a graced terminal cannot renew
* its own grace forever just by staying in the exposed map (fleetd #359 review, finding 2).
*/
private Set<String> gracedTerminals = Set.of();
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
@@ -141,7 +174,7 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
return cached;
}
/** One full pass: labelled tabs → their panes → those panes' terminals. */
/** One full pass: labelled tabs → live agents in them → those panes' terminals. */
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
@@ -158,17 +191,47 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
if (!nameByTab.isEmpty()) {
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String name = nameByTab.get(p.path("tab_id").asText(null));
String terminal = p.path("terminal_id").asText(null);
if (name != null && terminal != null && !terminal.isBlank()) {
byTerminal.put(terminal, name);
}
if (nameByTab.isEmpty()) {
gracedTerminals = Set.of();
return Map.of();
}
// fleetd #359: a labelled tab is only a lead when herdr also reports a running agent in
// it — the same liveness signal LeadLauncher.countLeads trusts for the identical purpose.
// Without this, a tab left behind by a session that has since died reads as live forever.
Set<String> tabsWithAgent = new HashSet<>();
for (JsonNode a : herdr.call("agent.list").path("agents")) {
String tabId = a.path("tab_id").asText(null);
if (tabId != null) {
tabsWithAgent.add(tabId);
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
Set<String> stillGraced = new HashSet<>();
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String tabId = p.path("tab_id").asText(null);
String name = nameByTab.get(tabId);
String terminal = p.path("terminal_id").asText(null);
if (name == null || terminal == null || terminal.isBlank()) {
continue;
}
if (tabsWithAgent.contains(tabId)) {
byTerminal.put(terminal, name);
continue;
}
// No agent reported for this tab, but its tab/pane are still here — this is the
// ambiguous case review finding 2 named: a successful agent.list that came back short
// does not prove the lead is dead. Grant one grace scan to a terminal we had already
// reported as live; a terminal we never reported live gets none, so the original #359
// fix (a genuinely dead tab is never reported) is unaffected for the common case.
if (cached.containsKey(terminal) && !gracedTerminals.contains(terminal)) {
byTerminal.put(terminal, name);
stillGraced.add(terminal);
}
}
gracedTerminals = stillGraced;
return Collections.unmodifiableMap(byTerminal);
}
@@ -177,12 +240,14 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
*
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix.
* configured lead just because it shares a prefix. The match strips a trailing
* {@link PendingCloseMarker} first, so a tab {@code LeadLauncher} has flagged as maybe-dead but
* not yet closed keeps resolving normally while that reconcile is pending.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
return tabToName.get(label.strip().toLowerCase(Locale.ROOT));
return tabToName.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
}
}
@@ -1,6 +1,8 @@
package dev.ltms.fleet.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.LinkedHashSet;
import java.util.List;
@@ -36,6 +38,8 @@ import java.util.Set;
*/
public final class PaneLocator {
private static final Logger log = LoggerFactory.getLogger(PaneLocator.class);
/**
* Bound on how many ancestor generations {@link #ancestorsOf} walks. This runs on every MCP
* call, so a cycle or a pathologically deep process tree must not hang identity resolution;
@@ -73,22 +77,46 @@ public final class PaneLocator {
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane on any searched daemon owns it (e.g. the caller is the
* primary, or off-host).
* The outcome of a {@link #terminalForPid} scan: the {@code terminal_id} of the agent pane
* whose process tree contains the pid ({@link #terminal} is {@code null} if none matched),
* and whether the scan that produced that answer ran to completion on every daemon searched.
*
* <p>{@link #complete} is {@code false} exactly when some {@code pane.process_info} call
* failed and, despite that, no pane was ever found to own the pid. In that case a {@code null}
* {@link #terminal} means "could not tell", not "definitely not a worker" — fleetd #505: a
* transient herdr error on the very pane that <em>does</em> own the caller's pid must not read
* as a clean negative and fall through to {@code Principal.primary}, the same way #317's
* {@code Caller.resolved()} already guards a failed lsof lookup. Callers ({@code
* ConnectionIdentity}, {@code CallerResolver}) must refuse rather than promote on an incomplete
* scan.
*
* <p>When a pane genuinely owns the pid, {@link #complete} is {@code true} regardless of
* whether some other, unrelated pane failed to answer earlier in the same scan — a positive
* match is definitive and does not need every pane to have been checked (a pane that "vanished
* mid-scan" but was never the match is still a clean, complete result).
*/
public String terminalForPid(long pid) {
public record Lookup(String terminal, boolean complete) {
private static final Lookup NOT_FOUND = new Lookup(null, true);
}
/**
* Resolve {@code pid} to the agent pane whose process tree contains it, across every searched
* herdr daemon. See {@link Lookup} for how to read a {@code null} terminal.
*/
public Lookup terminalForPid(long pid) {
if (pid <= 0) {
return null;
return Lookup.NOT_FOUND;
}
Set<Long> ancestry = ancestorsOf(pid);
for (HerdrClient herdr : herdrs) {
String terminal = terminalForPid(herdr, ancestry);
if (terminal != null) {
return terminal;
boolean complete = true;
for (int i = 0; i < herdrs.size(); i++) {
Lookup outcome = scan(herdrs.get(i), i, herdrs.size(), ancestry);
if (outcome.terminal() != null) {
return outcome; // a definite match — no need to finish checking other clients
}
complete = complete && outcome.complete();
}
return null;
return new Lookup(null, complete);
}
/**
@@ -117,31 +145,51 @@ public final class PaneLocator {
return ancestry;
}
private static String terminalForPid(HerdrClient herdr, Set<Long> ancestry) {
/** Whether a pane owns one of the scanned pid's ancestors, or the check of it failed outright. */
private enum Ownership { OWNS, DOES_NOT_OWN, UNKNOWN }
private static Lookup scan(HerdrClient herdr, int clientIndex, int clientCount, Set<Long> ancestry) {
boolean complete = true;
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsAnyOf(herdr, paneId, ancestry)) {
return pane.path("terminal_id").asText(null);
if (paneId == null) {
continue;
}
Ownership owns = paneOwnsAnyOf(herdr, clientIndex, clientCount, paneId, ancestry);
if (owns == Ownership.OWNS) {
return new Lookup(pane.path("terminal_id").asText(null), true);
}
if (owns == Ownership.UNKNOWN) {
complete = false;
}
}
return null;
return new Lookup(null, complete);
}
private static boolean paneOwnsAnyOf(HerdrClient herdr, String paneId, Set<Long> ancestry) {
private static Ownership paneOwnsAnyOf(HerdrClient herdr, int clientIndex, int clientCount,
String paneId, Set<Long> ancestry) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
// fleetd #505: this used to be read as a clean "does not own it" (the pane vanished
// mid-scan, just skip it) — one boolean carrying two different facts. It is UNKNOWN
// now: if THIS pane is the one that owns the pid, the caller must not be told "no pane
// owns it", because that reads as a real primary and is promoted under loopback-trust.
log.warn("pane.process_info failed for pane {} on herdr client {} of {} during a "
+ "pid-owner scan — treating it as \"could not tell\", not a clean "
+ "negative (fleetd #505): {}",
paneId, clientIndex + 1, clientCount, e.getMessage());
return Ownership.UNKNOWN;
}
if (ancestry.contains(info.path("shell_pid").asLong(-1))) {
return true;
return Ownership.OWNS;
}
for (JsonNode p : info.path("foreground_processes")) {
if (ancestry.contains(p.path("pid").asLong(-1))) {
return true;
return Ownership.OWNS;
}
}
return false;
return Ownership.DOES_NOT_OWN;
}
}
@@ -0,0 +1,42 @@
package dev.ltms.fleet.herdr;
/**
* The suffix {@code dev.ltms.fleet.lead.LeadLauncher} appends to a lead tab's label the first time a
* reconcile finds no running agent in it, before it is sure enough to close the tab outright.
*
* <p><strong>fleetd #359 review, finding 1.</strong> The daemon's own evidence showed
* {@code agent.list} can read "no agent" for a tab that genuinely has one running — so a single such
* reading must never be treated as proof a tab is dead. {@code LeadLauncher} now writes this marker
* on the first miss, and only closes the tab if a <em>later</em>, independent reconcile still finds
* it dead while the marker is still there. Two consecutive misses, one restart apart, is a much
* stronger claim than one.
*
* <p>{@link LeadTabScanner} strips the same suffix before matching a label against a configured
* lead's {@code tab}, so a flagged-but-actually-still-live tab keeps resolving normally — the marker
* changes nothing about which pane {@code LeadCoordLoop} can reach while the flag is pending. Both
* classes must use exactly this suffix, which is why it lives here rather than as a private constant
* on either.
*/
public final class PendingCloseMarker {
public static final String SUFFIX = " [fleetd:pending-close]";
private PendingCloseMarker() {
}
/** The label with any trailing pending-close marker removed, for name matching. */
public static String strip(String label) {
if (label == null) {
return null;
}
String stripped = label.strip();
return stripped.endsWith(SUFFIX)
? stripped.substring(0, stripped.length() - SUFFIX.length()).strip()
: stripped;
}
/** Whether a label currently carries the marker. */
public static boolean isFlagged(String label) {
return label != null && label.strip().endsWith(SUFFIX);
}
}
@@ -380,7 +380,9 @@ public final class CompletionResolver implements TurnListener {
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
if (startsWithExhaustion(matchedLine, exhausted)) {
exhaustionSink.onExhausted(target, reason);
}
}
return;
}
@@ -478,7 +480,9 @@ public final class CompletionResolver implements TurnListener {
+ "usable assistant block; no fleet_reply): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
if (startsWithExhaustion(matchedLine, exhausted)) {
exhaustionSink.onExhausted(target, reason);
}
}
return true;
}
@@ -544,8 +548,19 @@ public final class CompletionResolver implements TurnListener {
* fast completion. Runs the same backend-error classification the normal and raw-scrape paths
* apply, against whatever is on screen right now: a match together with the too-fast crash
* signature notifies {@link #backendErrorSink} (only on the resolution that wins the race). A
* non-match stays the original generic too-fast failure, naming the member and both timings,
* with whatever the pane shows appended so the caller sees the cause, not just "it failed".
* non-match stays the generic too-fast failure, naming the member and both timings, with
* whatever the pane shows appended so the caller sees the evidence, not just "it failed".
*
* <p>fleetd#376: <strong>this path must never resolve a completion.</strong> A fast backend can
* genuinely answer inside the floor, so the failure is sometimes wrong — but it is wrong in the
* loud direction, and the fix for that is honest wording, not a guess at the pane's meaning.
* Reclassifying from the scrape was tried and rejected: there is no reliable positive marker for
* "this is a real reply" across backends. {@link #lastAssistantBlock} falls back to the entire
* pane when it finds no {@code ⏺} marker, so on a crash the candidate "reply" is the whole
* screen; and {@code ⏺} itself is a Claude Code marker that an opencode pane never carries — the
* very backend whose speed raised this ticket. Any weaker test (non-blank, or "contains sentence
* punctuation") passes on almost every crash, because a pane holding a file path, a version
* number or a hostname contains a full stop. That trades a loud wrong answer for a silent one.
*/
private void failTooFast(String target, InFlight turn, CompletableFuture<Rendezvous.Resolution> waiter,
long elapsedNanos) {
@@ -556,13 +571,13 @@ public final class CompletionResolver implements TurnListener {
scrape = "";
}
String clippedScrape = clip(scrape);
String baseReason = String.format(
"member %s went BUSY -> DONE in %dms (floor %dms) — too fast to be real work, most "
+ "likely a backend error before any work started",
String timing = String.format(
"member %s went BUSY -> DONE in %dms (floor %dms)",
target, elapsedNanos / 1_000_000, MIN_TURN_NANOS / 1_000_000);
String backendError = firstMatchingLine(scrape, backendErrorPatternOrFallback(target));
if (backendError != null) {
String reason = baseReason + ": " + clippedScrape;
String reason = timing + " — too fast to be real work, and the pane carries a backend "
+ "error: " + clippedScrape;
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
@@ -571,7 +586,14 @@ public final class CompletionResolver implements TurnListener {
}
return;
}
fail(target, turn, clippedScrape.isBlank() ? baseReason : baseReason + ": " + clippedScrape);
// fleetd#376: no pattern matched, so the cause is genuinely unknown. Say that, rather than
// asserting a backend error the way this message used to — a fast backend really can finish
// inside the floor, and a reader who trusts a wrong cause stops looking at the pane.
String reason = timing + " — inside the floor. That is usually a backend error before any "
+ "work started, but a fast backend can answer inside it too, and nothing here can "
+ "tell those apart, so the turn is reported failed rather than guessed. Read the "
+ "pane below before deciding which it was";
fail(target, turn, clippedScrape.isBlank() ? reason : reason + ": " + clippedScrape);
}
/**
@@ -613,6 +635,46 @@ public final class CompletionResolver implements TurnListener {
return null;
}
/**
* True when nothing before the match on this pane line ends a sentence — that is, the match is
* still inside the line's first sentence rather than inside prose a member wrote about it.
* Used to decide whether an exhaustion match may quarantine a credential (fleetd #348).
*
* <p><strong>Why this is looser than {@link #startsWithBackendError}.</strong> An
* {@code exhaustedPattern} is written per profile and may name only the decisive words of a
* provider message — {@code "usage limit has been reached"} without its leading {@code "The"}.
* A start-of-line check would then reject the genuine refusal. That is the false negative
* fleetd #348's invariant 1 calls the worse direction: an unrecorded exhaustion leaves the
* fleet spawning into a credential with no capacity, and a quarantine runs 1800s against the
* backend-error cooldown's fixed 60s.
*
* <p>This rule accepts a superset of what a start-of-line check accepts: if the match begins
* right after the chrome, there is nothing in front of it, so there is no sentence ending
* either. So moving to it cannot add a false negative.
*
* <p><strong>No chrome skipping here, deliberately.</strong> The first version of this method
* copied {@code startsWithBackendError}'s leading-chrome loop. Measured on merge: deleting that
* loop left all 1369 tests green, and it must — the scan only looks for {@code . ! ?}, and no
* terminal chrome character is one of those. A step that cannot change the result is worse than
* no step, because the next reader takes it as evidence that chrome was handled.
*
* <p>It stays a heuristic. Prose whose <em>first</em> sentence carries the pattern still
* notifies the sink, and a genuine refusal behind an earlier full stop (a hostname, a version
* number) still does not. Both are known and neither is fixed here.
*/
private static boolean startsWithExhaustion(String line, Pattern pattern) {
var matcher = pattern.matcher(line);
if (!matcher.find()) {
return false;
}
for (int prefix = 0; prefix < matcher.start(); prefix++) {
if (".!?".indexOf(line.charAt(prefix)) >= 0) {
return false;
}
}
return true;
}
/**
* True when the error pattern begins the matched pane line, rather than appearing in prose.
*
@@ -636,17 +698,57 @@ public final class CompletionResolver implements TurnListener {
}
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* What an unset pattern key means for the classification it configures (fleetd#415).
* {@code coverage()} cannot infer this from the key's name — the two keys it currently
* describes disagree on it, and a string comparison on the name would just move the same bug
* to a new spot — so every caller must state it explicitly.
*
* <p><strong>This alone does not prove a caller passes the right one for its key.</strong> A
* test that calls {@code coverage()} directly and supplies the meaning itself only proves this
* enum is worded correctly, never that {@code Fleetd}'s two call sites pair each key with its
* true meaning — that pairing is #415's actual defect. Measured on review: swapping the two
* {@code UnsetMeaning} arguments at those call sites (giving {@code exhaustedPattern} the
* built-in-default wording and {@code errorPattern} the off wording — #415's exact defect with
* the keys exchanged) compiled with 0 errors and left all 1506 existing tests green. See
* {@code dev.ltms.fleet.Fleetd#exhaustedPatternCoverageLine}/{@code #errorPatternCoverageLine}
* and {@code FleetdPatternCoverageLineTest}, which exists specifically to catch that swap.
*/
public enum UnsetMeaning {
/** No fallback exists: a profile with no configured pattern truly has this classification off. */
OFF,
/** A built-in pattern applies when unset: the classification still runs for that profile. */
BUILT_IN_DEFAULT
}
/**
* Coverage summary for a fleetd#201/CB-578-style pattern-key classification, logged at startup
* the way {@link dev.ltms.fleet.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* <p>fleetd#415: this method measures <em>pattern coverage</em> — how many profiles set the
* key — which is not the same thing as <em>feature state</em> for a key with a fallback. For
* {@code errorPattern}, an empty {@code configuredProfiles} still runs the classification
* against {@code CompletionResolver}'s built-in compatibility pattern ({@link #BACKEND_ERROR}
* at line ~84); for {@code exhaustedPattern} there is no fallback, so empty really does mean
* off. {@code unsetMeaning} is the single, required source of that fact — see
* {@link dev.ltms.fleet.config.FleetConfig#rejectMalformedProfilePatterns} lines ~2029-2032 for
* where it is documented for config authors. It is a required parameter, not a defaulted
* overload: a third pattern key added later must supply one to compile at all, rather than
* silently inheriting whichever wording this method happened to default to.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
* @param configuredProfiles the subset of {@code allProfiles} that carry the pattern
*/
public static String coverage(String patternKey, Set<String> allProfiles, Set<String> configuredProfiles) {
public static String coverage(String patternKey, UnsetMeaning unsetMeaning, Set<String> allProfiles,
Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return "off (no profile has an " + patternKey + " configured; profiles: " + sorted(allProfiles) + ")";
return switch (unsetMeaning) {
case OFF -> "off (no profile has an " + patternKey + " configured; profiles: "
+ sorted(allProfiles) + ")";
case BUILT_IN_DEFAULT -> "built-in default for all profiles (no profile customises "
+ patternKey + "; profiles: " + sorted(allProfiles) + ")";
};
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
@@ -15,6 +15,7 @@ import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
import java.util.function.Predicate;
import java.util.stream.Collectors;
@@ -94,6 +95,14 @@ public final class Injector {
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
/**
* Wall-clock source for the readiness-grace elapsed time logged in {@link #onStatus} (fleetd
* #501). Production constructors default this to {@code System::currentTimeMillis}; the
* package-private constructors below take it explicitly so a test can supply a stub whose
* advance does not track {@link #POLL_INTERVAL_MILLIS} — copying the shape {@code LeadRollover}
* already uses for the same purpose.
*/
private final LongSupplier nowMillis;
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
@@ -125,20 +134,34 @@ public final class Injector {
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this(agents, turnListener, ready, forget, System::currentTimeMillis);
}
/** Full constructor — for tests: an injectable wall-clock supplier (fleetd #501). */
Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, LongSupplier nowMillis) {
this.agents = agents;
this.router = null;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
this.nowMillis = nowMillis;
}
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this(router, turnListener, ready, forget, System::currentTimeMillis);
}
/** Full constructor — for tests: an injectable wall-clock supplier (fleetd #501). */
Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, LongSupplier nowMillis) {
this.agents = null;
this.router = router;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
this.nowMillis = nowMillis;
}
private AgentControl agentsFor(String target) {
@@ -209,6 +232,7 @@ public final class Injector {
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int unknownSincePostTurn; // the same, for the post-turn housekeeping phase (fleetd #306)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
long notReadySinceMillis; // wall-clock time of the FIRST non-ready sample in the current notReadySincePoll streak (fleetd #501); reset alongside it
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
@@ -302,6 +326,7 @@ public final class Injector {
t.unknownSinceTurn = 0;
t.unknownSincePostTurn = 0;
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
@@ -353,6 +378,7 @@ public final class Injector {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
try {
agentsFor(target).send(target, p.text());
t.queue.poll();
@@ -370,23 +396,52 @@ public final class Injector {
sent = p;
sendError = e;
}
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
for (Pending pending : notReady) {
pending.state = Pending.State.NOT_DELIVERED;
} else if (p != null) {
// fleetd #501: stamp the wall-clock time of the FIRST non-ready sample in
// this streak, so the expiry log below can print how long the target
// actually sat non-ready — not just how many polls that took.
if (t.notReadySincePoll == 0) {
t.notReadySinceMillis = nowMillis.getAsLong();
}
if (++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
for (Pending pending : notReady) {
pending.state = Pending.State.NOT_DELIVERED;
}
// fleetd #501: t.notReadySincePoll — the loop's own counter, already in
// scope — is printed here instead of the READINESS_GRACE_POLLS constant.
// On this branch the counter has JUST reached the threshold, so the two
// agree by construction and no test can tell them apart. Printed anyway:
// it gives this line one source of truth instead of two, so a later
// change to the loop above cannot leave this message reporting a number
// the loop no longer produces.
//
// elapsedMillis is a different case: it is NOT equal-by-construction to
// the truth. notReadySincePoll only increments on a sample that reaches
// this branch (p != null, not ready) — a poll that misses that condition
// advances real time without advancing the counter — and this loop's real
// period is not guaranteed to equal POLL_INTERVAL_MILLIS (load, or a host
// sleep, can widen the real gap far past it). READINESS_GRACE_POLLS *
// POLL_INTERVAL_MILLIS / 1000 is arithmetic on two constants, not a
// measurement, so it stays here only as the labelled CONFIGURED budget,
// never presented as elapsed time.
long elapsedMillis = nowMillis.getAsLong() - t.notReadySinceMillis;
log.warn("readiness grace for {} expired after {} polls (configured={} "
+ "polls/{}s elapsed={}ms): target never became "
+ "deliverable, so failing {} queued message(s) that "
+ "never reached its pane",
target, t.notReadySincePoll, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, elapsedMillis,
notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
}
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
+ "became deliverable, so failing {} queued message(s) that never "
+ "reached its pane",
target, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
}
}
}
@@ -0,0 +1,91 @@
package dev.ltms.fleet.inject;
import dev.ltms.fleet.config.FleetConfig;
import java.util.Map;
import java.util.Objects;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Supplier;
import java.util.regex.Pattern;
/**
* fleetd #446: the single live source for "is {@code exhaustedPattern} configured for this
* profile, and what does it compile to". Read fresh off the config supplier on every call — the
* same reason {@code CompositePeerLauncher#models0} is a live supplier read rather than a value
* captured at construction (see that class's doc) — so an operator can arm or disarm usage-limit
* detection for a profile by editing {@code exhaustedPattern} and reloading, with no restart.
*
* <p>Before this class, {@code exhaustedPattern} was compiled once into a {@code Map<String,
* Pattern>} built inside {@code Fleetd.main} at startup ({@code ConfigRef}'s class doc used to
* list it under <em>Deferred</em>, CB-578 stage A) — the asymmetry fleetd #446 exists to close:
* an operator could turn a model off at runtime ({@code models.allow}'s {@code enabled: false},
* hot since fleetd #422) but could not arm the detector that would tell them to, without a
* restart. That was backwards for a feature whose whole point is to react while the fleet runs.
*
* <p>{@link #patternFor} is what {@link CompletionResolver} enforces on (via the {@link
* ExhaustedPatternLookup} production wiring in {@code Fleetd.main}); {@link #armed} is what
* {@code fleet_profiles}/{@code GET /profiles} report as {@code exhaustionDetectionArmed}. Both
* read this ONE object, so the report can never disagree with the behaviour — the same rule
* {@code CompositePeerLauncher#modelGateState()}'s javadoc states for the model gate's own
* armed/off pair (fleetd #404): "armed" and "which models are off" must come from one read of the
* same accessor the gate enforces on.
*
* <h2>Cache eviction</h2>
* Cached by profile NAME, not by pattern text. A {@code Map<String, Pattern>} keyed by the
* pattern STRING would grow by one entry per distinct regex ever typed for any profile across the
* daemon's uptime — unbounded in practice, because tuning a regex to match a backend's exact
* wording is exactly the kind of edit an operator makes several times while getting it right, and
* every edit-and-reload cycle would leave the previous attempt's compiled {@link Pattern} behind
* forever. Keyed by profile name instead, this cache holds at most one entry per profile name that
* has ever been looked up — and that key space is bounded by the (small, human-authored) set of
* configured profiles, which changes only on a restart: adding or removing a profile is itself a
* deferred key (a new backend needs its own launcher, built once — see {@code ConfigRef}'s class
* doc), so profile names do not churn the way pattern text does. Re-editing an EXISTING profile's
* {@code exhaustedPattern} — the case this class exists to make hot — simply overwrites that
* profile's one cache entry; it never adds a new one.
*/
public final class LiveExhaustedPatterns {
/** A compiled pattern paired with the source string it was compiled from, for change detection. */
private record Cached(String source, Pattern compiled) {
}
private final Supplier<Map<String, FleetConfig.Profile>> profiles;
private final ConcurrentHashMap<String, Cached> cache = new ConcurrentHashMap<>();
public LiveExhaustedPatterns(Supplier<Map<String, FleetConfig.Profile>> profiles) {
this.profiles = Objects.requireNonNull(profiles, "profiles");
}
/**
* The compiled {@code exhaustedPattern} currently configured for {@code profileName}, or
* {@code null} when that profile is unknown or has none configured. Recompiles only when the
* live pattern text differs from what is cached for this profile name; {@code
* FleetConfig#rejectMalformedProfilePatterns} already refuses a config (at load and at reload)
* whose {@code exhaustedPattern} does not compile, so this is not expected to throw in
* production — it is not defended against here for that reason, the same trust
* {@code CompositePeerLauncher#models0} places in config validation having already run.
*/
public Pattern patternFor(String profileName) {
if (profileName == null) {
return null;
}
FleetConfig.Profile profile = profiles.get().get(profileName);
if (profile == null || !profile.hasExhaustedPattern()) {
return null;
}
String source = profile.exhaustedPattern();
Cached cached = cache.get(profileName);
if (cached != null && cached.source().equals(source)) {
return cached.compiled();
}
Cached fresh = new Cached(source, Pattern.compile(source));
cache.put(profileName, fresh);
return fresh.compiled();
}
/** Whether {@code profileName} currently has a usage-limit pattern configured (live). */
public boolean armed(String profileName) {
return patternFor(profileName) != null;
}
}
@@ -4,6 +4,7 @@ import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.PendingCloseMarker;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -12,9 +13,11 @@ import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import dev.ltms.fleet.peer.PeerLauncher;
/**
@@ -44,8 +47,33 @@ import dev.ltms.fleet.peer.PeerLauncher;
* privilege escalation (the tab label never granted anything a pane could take for itself; see that
* class's javadoc), but <em>staleness</em>: a label left behind by a crashed session would otherwise
* read as a live lead forever, and the lead would never be relaunched. So a lead counts as live only
* when herdr also reports a <em>running agent</em> in that tab — see {@link #liveLeads}. A labelled
* when herdr also reports a <em>running agent</em> in that tab — see {@link #countLeads}. A labelled
* tab with no agent in it is not a lead.
*
* <p><strong>fleetd #359 — a stale label used to pile up, not just mislead.</strong> Finding "not
* live" here used to mean only one thing: launch another. The old label was left exactly where it
* was, so a daemon that restarted enough times — or hit one herdr read that missed a genuinely
* running agent — accumulated one more identically-labelled dead tab per occurrence, and
* {@code LeadTabScanner} (before its own #359 fix) reported every one of them as a lead.
*
* <p><strong>Review finding 1 — closing on one reading is worse than the bug.</strong> The first
* version of this fix closed a name's dead tabs the moment a single {@link #countLeads} reading
* called them dead. The ticket's own live evidence rules that out: on a real host, {@code
* agent.list} was seen reporting "0 live" for a tab that a plain {@code ps} confirmed was running a
* real session. Closing on that reading would have destroyed the operator's actual lead — a worse
* failure than the extra tab it replaces. So {@link #ensureLeads()} now needs the same dead reading
* <em>twice</em>, one restart apart, before it closes anything: the first time a labelled tab reads
* dead, it is only flagged ({@link PendingCloseMarker}), left running, untouched; it is closed only
* if a <em>later</em>, independently-connected reconcile still finds it dead while the flag is
* still there. A transient miss self-heals — the next reconcile sees the agent again and clears the
* flag (see {@code toUnflag} below) — so the worst case for a single bad reading is one extra tab
* surviving one more restart, never a live session destroyed. Two alternatives were considered and
* rejected: corroborating {@code agent.list} against a second, truly independent signal was dropped
* because nothing else herdr exposes proves "is a process attached to this pane" any better — a
* second call to the same unreliable source is not independent evidence; capping the close to "all
* but the most recent dead tab" was dropped because "most recent" has no reliable ordering across
* tab ids and would leave the true failure mode (a name that is <em>never</em> reconfirmed) growing
* by one tab per bad reading forever, which is the exact defect this ticket exists to fix.
*/
public final class LeadLauncher {
@@ -79,9 +107,9 @@ public final class LeadLauncher {
return 0;
}
Map<String, Integer> live;
Map<String, LeadCount> live;
try {
live = liveLeads(leaders);
live = countLeads(leaders);
} catch (HerdrException e) {
// Counting is the whole safety mechanism against double-spawning. If we cannot count, we
// must not guess — spawning a second orchestrator is worse than starting none.
@@ -93,9 +121,53 @@ public final class LeadLauncher {
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
String name = e.getKey();
FleetConfig.Leader lead = e.getValue();
int running = live.getOrDefault(name, 0);
LeadCount state = live.getOrDefault(name, LeadCount.NONE);
int running = state.running();
int wanted = lead.instances();
// fleetd #359 review finding 1: a tab already flagged pending-close, still labelled for
// this lead, and STILL hosting no agent on this separate reconcile — two independent
// readings agree, so close it. A tab found dead for the first time is only flagged below,
// never closed on the spot.
for (String tabId : state.toClose()) {
log.info("lead '{}': closing tab {} — flagged pending-close on a previous reconcile "
+ "and still no agent running in it", name, tabId);
try {
spaces.closeTab(tabId);
} catch (RuntimeException cleanup) {
log.warn("could not close stale tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
// A tab labelled for this lead, with no agent running in it, seen dead for the first
// time — flag it rather than closing it. One reading of `agent.list` is not enough
// evidence to destroy a tab that might genuinely be live (see the class javadoc).
for (String tabId : state.toFlag()) {
String flagged = lead.tabLabel() + PendingCloseMarker.SUFFIX;
log.info("lead '{}': tab {} has no agent running in it this reconcile — flagging it "
+ "'{}' rather than closing; it is only closed if a later reconcile still "
+ "finds it dead", name, tabId, flagged);
try {
spaces.renameTab(tabId, flagged);
} catch (RuntimeException cleanup) {
log.warn("could not flag stale tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
// A previously-flagged tab that is running an agent again — the miss that flagged it was
// transient. Clear the flag so a future, unrelated miss starts its own two-reading count
// rather than closing on the strength of this one's already-spent flag.
for (String tabId : state.toUnflag()) {
log.info("lead '{}': tab {} is running an agent again — clearing its pending-close flag",
name, tabId);
try {
spaces.renameTab(tabId, lead.tabLabel());
} catch (RuntimeException cleanup) {
log.warn("could not clear the pending-close flag on tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
if (running >= wanted) {
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
continue;
@@ -126,18 +198,29 @@ public final class LeadLauncher {
}
/**
* How many live leads exist per configured name: a running agent in a tab labelled with that
* lead's exact {@code tab} (CB-579). Member workspaces are excluded, exactly as the scanner
* excludes them: a member must not be counted as a lead because it happens to sit in a matching
* tab.
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
* not be counted as a lead because it happens to sit in a matching tab.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
* pin is retired: {@code tab} is now the only field identity depends on, and {@link Agent}
* already carries {@link Agent#tabId()} directly, so a hand-opened lead is found the same way an
* auto-launched one is — by labelling its tab to match.
*
* <p>fleetd #359 review finding 1: a labelled tab with nothing running in it is split into
* {@code toClose} (already flagged pending-close by a previous reconcile, and still dead — two
* independent readings agree) and {@code toFlag} (dead for the first time — not enough evidence
* to close yet). {@code toUnflag} is the reverse: a tab flagged pending-close that is running an
* agent again, so the flag it carries no longer means anything and {@link #ensureLeads()} clears
* it.
*/
private Map<String, Integer> liveLeads(Map<String, FleetConfig.Leader> leaders) {
private record LeadCount(int running, List<String> toClose, List<String> toFlag, List<String> toUnflag) {
static final LeadCount NONE = new LeadCount(0, List.of(), List.of(), List.of());
}
private Map<String, LeadCount> countLeads(Map<String, FleetConfig.Leader> leaders) {
// A lead and the members share ONE workspace now (the operator asked for a single "session"
// with many tabs), so a workspace can no longer be excluded wholesale — the lead lives in the
// member workspace by design. The sole discriminator is the exact tab label: a lead carries
@@ -145,6 +228,7 @@ public final class LeadLauncher {
// profile's `worker: {profile} #{n}` template. These never collide, so an exact-label match
// separates them without needing to know which workspace anyone is in.
Map<String, String> nameByTab = new LinkedHashMap<>();
Set<String> flaggedTabIds = new LinkedHashSet<>();
for (Workspace ws : spaces.listWorkspaces()) {
if (ws.workspaceId() == null) {
continue;
@@ -153,31 +237,65 @@ public final class LeadLauncher {
String declared = leadNameOf(tab.label(), leaders);
if (declared != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), declared);
if (PendingCloseMarker.isFlagged(tab.label())) {
flaggedTabIds.add(tab.tabId());
}
}
}
}
Set<String> liveTabIds = new LinkedHashSet<>();
Map<String, Integer> counts = new LinkedHashMap<>();
for (Agent a : agents.list()) {
String name = nameByTab.get(a.tabId());
if (name != null) {
counts.merge(name, 1, Integer::sum);
liveTabIds.add(a.tabId());
}
}
return counts;
Map<String, List<String>> toCloseByName = new LinkedHashMap<>();
Map<String, List<String>> toFlagByName = new LinkedHashMap<>();
Map<String, List<String>> toUnflagByName = new LinkedHashMap<>();
nameByTab.forEach((tabId, name) -> {
boolean live = liveTabIds.contains(tabId);
boolean flagged = flaggedTabIds.contains(tabId);
if (live) {
if (flagged) {
toUnflagByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
}
return;
}
if (flagged) {
toCloseByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
} else {
toFlagByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
}
});
Map<String, LeadCount> out = new LinkedHashMap<>();
for (String name : leaders.keySet()) {
out.put(name, new LeadCount(counts.getOrDefault(name, 0),
toCloseByName.getOrDefault(name, List.of()),
toFlagByName.getOrDefault(name, List.of()),
toUnflagByName.getOrDefault(name, List.of())));
}
return out;
}
/**
* The configured lead a tab label names, or {@code null} for a label that names none.
*
* <p>Matched exactly (case-insensitively) against each lead's configured {@code tab}, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead. A
* trailing {@link PendingCloseMarker} is stripped first, so a tab this class flagged on a
* previous reconcile is still recognised as the same lead's tab on this one.
*/
private String leadNameOf(String label, Map<String, FleetConfig.Leader> leaders) {
if (label == null) {
return null;
}
String l = label.strip();
String l = PendingCloseMarker.strip(label);
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
String tab = e.getValue().tabLabel();
if (tab != null && l.equalsIgnoreCase(tab.strip())) {
@@ -0,0 +1,649 @@
package dev.ltms.fleet.lead;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.function.Consumer;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
* clears the lead's own pane and bootstraps a fresh session against that file.
*
* <p>This is the executor only. Nothing in this ticket wires an MCP tool onto {@link #open}/
* {@link #confirm}/{@link #cancel} — that is a separate, later unit; until it lands, nothing calls
* this class at all.
*
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING}
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()}
* itself returns. The poll always timed out — but only after the {@code /clear} had already been
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
* ever started, and a refusal return value that claimed nothing had happened.
*
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
* at all — it only records that the request is approved and hands a one-shot continuation to
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
* once the calling turn has ended, in this order:
* <ol>
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
* away live context — exactly the failure this correction exists to prevent.</li>
* <li>{@code agents.send(lead, "/clear")}</li>
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
* see {@link #waitForClearPickupAndSettle})</li>
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
* </ol>
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may
* still be mid-turn, possibly for a long time, when the caller gets that answer back.
*
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket
* that first defined this class required "no timer, no scheduler, no background thread" so that
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
* continuationRunner} launches a single-shot task that exists only because one specific,
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at
* construction time or on any schedule, and no two invocations of it ever share state. A recurring
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work
* slightly later than the method return, as a continuation of the same approved request, rather
* than entirely inside the method body.</strong>
*
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
* correction.</strong> The first version resolved the pane to clear via {@code
* PrimaryRegistry#primaryTerminal()}. That is correct for a background loop with no caller (see
* {@code dev.ltms.fleet.msg.LeadHeartbeatLoop}), but wrong here and a violation of this project's
* own charter invariant 3 — "identity comes from the connection, never an argument." This daemon
* can hold more than one labelled lead tab (see {@code LeadLauncher}'s fleetd #359 two-reading
* dead-tab cleanup), so a single-slot lookup lets lead X's {@link #confirm} clear lead Y's pane: an
* unrecoverable loss of someone else's context, and a different lead's context at that. Both
* {@link #open} and {@link #confirm} now take the lead's terminal id as a parameter instead —
* {@link #open} stores it on the {@link PendingRollover}, and {@link #confirm} refuses with {@link
* RefusalReason#NOT_YOUR_ROLLOVER} unless the caller's terminal matches the one {@link #open}
* recorded. <strong>The terminal id passed to both methods must come from the MCP layer's own
* connection-based caller resolution — the same source {@code auth/CallerResolver#resolve} uses to
* build a {@code Principal.leader(...)} (see its use of {@code ConnectionIdentity.Caller#terminal})
* — never a value the client supplies or chooses.</strong> The later MCP-tool unit that wires
* {@link #open}/{@link #confirm} must pass the resolved caller terminal, not a request field.
*
* <p><strong>A relative {@code handoverPath} resolves against the CALLING lead's workspace, never
* the daemon's own cwd — a fleetd #480 follow-up.</strong> The daemon and a lead's own pane can
* have different working directories (this repo nests {@code fleetd/} inside its own root, so the
* daemon's cwd and {@code fleet.leaders.<name>.cwd} already differ on this host). {@link #open}
* resolves {@code cfg.handoverPath()} to an ABSOLUTE path exactly once — against {@code
* leadWorkspace.apply(leadTerminal)} when that lookup returns a non-null, non-blank workspace, and
* against {@code System.getProperty("user.dir")} otherwise (the same fallback {@code
* LeadLauncher#launch} already uses for a lead with no configured {@code cwd}) — and stores only
* that absolute path on {@link PendingRollover}. Every later read of {@code
* PendingRollover#handoverPath()} (the freshness/exists/empty checks in {@link #checkHandover},
* the value {@code FleetMcp} hands back to the lead in the {@code open} response so it knows where
* to WRITE the file, and {@link FleetConfig.LeadRollover#bootstrapTextFor} which names it in the
* text typed into the fresh session) therefore already sees the resolved absolute form and never
* needs to resolve anything itself.
*/
public final class LeadRollover {
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
static final long SETTLE_POLL_MS = 250;
/**
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
* before releasing rather than wedging the roll — the same constant and the same
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
* nudging again — so 8 polls produce 7 nudges, not 8.
*/
static final int PICKUP_GRACE_POLLS = 8;
/**
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
*
* @param leadTerminal the lead pane that opened this request — the only terminal that may
* later {@link #confirm} it (see {@link RefusalReason#NOT_YOUR_ROLLOVER})
* @param handoverPath the ABSOLUTE, resolved handover path — never the raw configured value,
* which may have been relative. {@link #open} resolves it once, against the
* calling lead's workspace, before storing it here; see this class's
* javadoc. This is the value the MCP layer hands back to the lead as
* "write your file here", so callers may rely on it always being absolute.
*/
public record PendingRollover(String token, String leadTerminal, String handoverPath,
long requestedAtMillis) {}
/** Which check refused a {@link #confirm} call, named so a caller can act on it. */
public enum RefusalReason {
/** {@code leadRollover:} is not configured — absent at construction, or removed since. */
NOT_CONFIGURED,
/** {@code token} names no pending request: never opened, already confirmed, or cancelled. */
UNKNOWN_TOKEN,
/**
* The caller's terminal does not match the terminal that {@link #open} recorded for this
* token. Only the lead that opened a request may confirm it (fleetd #480 correction 2).
*/
NOT_YOUR_ROLLOVER,
/** {@code requireOperatorConfirm: true} and the caller passed {@code operatorConfirmed: false}. */
OPERATOR_NOT_CONFIRMED,
/** The handover file does not exist. */
HANDOVER_MISSING,
/** The handover file exists but is empty. */
HANDOVER_EMPTY,
/**
* The handover file's modified time is not after {@link #open}'s request timestamp, or is
* older than {@code maxDocAgeSeconds}.
*/
HANDOVER_STALE
}
/**
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been
* cleared; the continuation may still be waiting for the calling turn to end when this returns.
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
*/
public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
static RollDecision approved() {
return new RollDecision(true, null, "confirmed; the roll will run once the calling turn ends");
}
static RollDecision refused(RefusalReason reason, String detail) {
return new RollDecision(false, reason, detail);
}
}
private final AgentControl agents;
private final Supplier<FleetConfig.LeadRollover> configSupplier;
/**
* Terminal id → that lead's configured workspace directory (their {@code
* fleet.leaders.<name>.cwd}), or {@code null} when the terminal names no currently-recognised
* lead. {@link #open} calls this to resolve a relative {@code handoverPath} — see this class's
* javadoc. Required: there is no sane default that would not silently reintroduce the
* daemon-cwd bug this parameter exists to fix.
*/
private final Function<String, String> leadWorkspace;
private final LongSupplier nowMillis;
private final Runnable settleSleeper;
/**
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
* thread per confirmed request — see this class's javadoc for why that is a single-shot task,
* not a background scheduler. Tests inject {@code Runnable::run} to make the continuation run
* synchronously and deterministically on the calling thread.
*/
private final Consumer<Runnable> continuationRunner;
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace) {
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis,
() -> sleepUninterruptibly(SETTLE_POLL_MS),
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
}
/**
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
* against a file's modified time, which only a wall clock is comparable to, and {@code
* nanoTime} freezes while the host sleeps (fleetd #386).
*/
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace, LongSupplier nowMillis,
Runnable settleSleeper, Consumer<Runnable> continuationRunner) {
this.agents = agents;
this.configSupplier = configSupplier;
this.leadWorkspace = leadWorkspace;
this.nowMillis = nowMillis;
this.settleSleeper = settleSleeper;
this.continuationRunner = continuationRunner;
}
private static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — this poll loop should not be aborted by an
// interrupt that was not meant for it.
}
}
/**
* The lead says it is ready to be replaced. Generates a token and records the resolved
* handover path, this moment's wall-clock timestamp (the baseline {@link #confirm} checks the
* handover file's modified time against), and {@code leadTerminal} — only that exact terminal
* may later {@link #confirm} this token.
*
* @param leadTerminal the calling lead's terminal id, resolved by the MCP layer from the
* connection (see this class's javadoc) — never a client-supplied value
* @param reason free-text audit note (logged only; not otherwise interpreted or stored)
* @throws IllegalStateException if {@code leadRollover:} is not configured
* @throws IllegalArgumentException if {@code leadTerminal} is null or blank
*/
public PendingRollover open(String leadTerminal, String reason) {
FleetConfig.LeadRollover cfg = configSupplier.get();
if (cfg == null) {
throw new IllegalStateException("leadRollover: is not configured");
}
if (leadTerminal == null || leadTerminal.isBlank()) {
throw new IllegalArgumentException("leadTerminal is required — it must be resolved from "
+ "the caller's connection, never accepted as a client-chosen argument");
}
String token = UUID.randomUUID().toString();
long requestedAt = nowMillis.getAsLong();
String resolvedPath = resolveHandoverPath(cfg.handoverPath(), leadTerminal);
PendingRollover p = new PendingRollover(token, leadTerminal, resolvedPath, requestedAt);
pending.put(token, p);
if (resolvedPath.equals(cfg.handoverPath())) {
log.info("lead-rollover: open token={} lead={} handoverPath={} reason={}",
token, leadTerminal, resolvedPath, reason);
} else {
// The configured value was relative (or merely un-normalized) and resolved to a
// different string — log both, so an operator reading this line can see which
// directory the daemon actually looked in, not just the value it was given.
log.info("lead-rollover: open token={} lead={} configuredHandoverPath={} "
+ "resolvedHandoverPath={} reason={}",
token, leadTerminal, cfg.handoverPath(), resolvedPath, reason);
}
return p;
}
/**
* Resolve {@code configured} to an absolute path exactly once, here, so nothing downstream
* ({@link #checkHandover}, the MCP layer's {@code open} response, {@link
* FleetConfig.LeadRollover#bootstrapTextFor}) ever has to resolve — or worse, silently
* mis-resolve — a relative path again.
*
* <p><strong>The return value is GUARANTEED absolute, not merely usually absolute.</strong>
* {@code leadWorkspace.apply(leadTerminal)} returns an operator-configured {@code
* fleet.leaders.<name>.cwd} string, and nothing forces an operator to write an absolute one —
* a relative {@code cwd} resolved with plain {@link Path#resolve} would still yield a relative
* result, silently reopening the exact bug this class exists to fix (every later reader back to
* interpreting an ambiguous string against ITS OWN working directory). {@link
* Path#toAbsolutePath()} closes that: it resolves any remaining relative path against {@code
* user.dir} (the JVM's own cwd), which is the correct base for an operator-written path the
* daemon process itself is meant to interpret, exactly like the {@code user.dir} fallback used
* below. Applying it unconditionally on both branches means the ALREADY-absolute branch stays a
* no-op (an absolute path is unaffected by {@code toAbsolutePath()}) while the relative-{@code
* cwd} branch above is closed the same way.
*
* <ul>
* <li>already absolute → returned unchanged (normalized)</li>
* <li>relative → resolved against {@code leadWorkspace.apply(leadTerminal)} when that is
* non-null and non-blank; otherwise against {@code System.getProperty("user.dir")} — the
* same fallback {@code LeadLauncher#launch} uses for a lead with no configured {@code
* cwd}. If {@code leadWorkspace}'s own answer is itself relative (an operator wrote a
* relative {@code cwd:}), the result is finished off against the daemon's own
* {@code user.dir} — see the paragraph above.</li>
* </ul>
*/
private String resolveHandoverPath(String configured, String leadTerminal) {
Path path = Path.of(configured);
if (path.isAbsolute()) {
// toAbsolutePath() is a no-op for an already-absolute path — kept here anyway so both
// branches call the exact same guarantee, rather than one branch relying on
// isAbsolute() alone to already imply what toAbsolutePath() enforces.
return path.toAbsolutePath().normalize().toString();
}
String workspace = leadWorkspace.apply(leadTerminal);
Path base = (workspace == null || workspace.isBlank())
? Path.of(System.getProperty("user.dir"))
: Path.of(workspace);
return base.resolve(path).toAbsolutePath().normalize().toString();
}
/**
* Validate every gate, then — if and only if all of them pass — hand a one-shot continuation
* that performs the actual roll to {@code continuationRunner} and return. <strong>This method
* never itself sends anything to the lead's pane</strong> — see this class's javadoc for why
* (it is called FROM the lead's own turn, so the pane cannot possibly be injectable yet).
*
* <p>Order: token lookup, then ownership ({@code callerTerminal} must match the terminal
* {@link #open} recorded — {@link RefusalReason#NOT_YOUR_ROLLOVER}), then the
* operator-confirmation gate, then the three handover-file checks (exists, not empty, fresh —
* see {@link #checkHandover}). The first failing check is returned and {@code token} stays
* pending (so a caller can fix the problem — e.g. rewrite the handover file — and retry with
* the same token); it is consumed only once every gate passes and the continuation is launched.
*
* @param callerTerminal the CALLING lead's terminal id, resolved by the MCP layer from the
* connection — never a client-supplied value (see this class's javadoc)
* @param token the token {@link #open} returned
* @param operatorConfirmed the caller's answer to "has an operator confirmed this roll" —
* consulted only when the live config's {@code requireOperatorConfirm}
* is true
*/
public RollDecision confirm(String callerTerminal, String token, boolean operatorConfirmed) {
FleetConfig.LeadRollover cfg = configSupplier.get();
if (cfg == null) {
return RollDecision.refused(RefusalReason.NOT_CONFIGURED, "leadRollover: is not configured");
}
PendingRollover p = pending.get(token);
if (p == null) {
return RollDecision.refused(RefusalReason.UNKNOWN_TOKEN,
"token " + token + " names no pending rollover request");
}
if (!p.leadTerminal().equals(callerTerminal)) {
return RollDecision.refused(RefusalReason.NOT_YOUR_ROLLOVER,
"token " + token + " was opened by a different lead terminal");
}
if (cfg.requireOperatorConfirm() && !operatorConfirmed) {
return RollDecision.refused(RefusalReason.OPERATOR_NOT_CONFIRMED,
"requireOperatorConfirm is true and operatorConfirmed was false");
}
RollDecision docCheck = checkHandover(p, cfg);
if (docCheck != null) {
return docCheck;
}
pending.remove(token);
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
token, callerTerminal);
continuationRunner.accept(() -> runRollover(p, cfg));
return RollDecision.approved();
}
/**
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
* four-step order. There is no result to return to by this point, so every outcome is logged
* only.
*/
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
String lead = p.leadTerminal();
long rollStartMillis = nowMillis.getAsLong();
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
if (!turnResult.settled()) {
// fleetd #494 follow-up: this line had the SAME defect as the /clear-timeout line below
// — cfg.turnSettleSeconds() is the CONFIGURED budget, not how long this wait actually
// ran. Print the measured elapsed time alongside it, labelled, exactly like the /clear
// path already does.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "confirm() — refusing to send /clear at all; the calling lead's own "
+ "turn is still live and clearing it now would destroy live context "
+ "(token={}, configured={}s elapsed={}ms)",
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
return;
}
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext:
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
// pane forever (see this class's javadoc).
agents.send(lead, "/clear");
ClearSettleResult clearResult = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
if (!clearResult.settled()) {
// fleetd #494: cfg.clearSettleSeconds() is the CONFIGURED budget, not how long the wait
// actually ran — an operator reading only that number wrongly believes it is a measured
// duration. Print the measured elapsed time and nudge count alongside it, each labelled,
// so the two can be compared at a glance.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "/clear — NOT sending bootstrapText (token={}, configured={}s "
+ "elapsed={}ms nudges={})",
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
clearResult.nudges());
return;
}
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
}
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
public boolean cancel(String token) {
return pending.remove(token) != null;
}
/**
* The three handover-file checks, in order: exists, not empty, fresh (modified after
* {@link #open}'s timestamp and not older than {@code maxDocAgeSeconds}). Stats {@code
* p.handoverPath()} directly — {@link #open} already resolved it to an absolute path, so this
* never has to guess which directory it means.
*
* @return the first failing check's refusal, or {@code null} when all three pass
*/
private RollDecision checkHandover(PendingRollover p, FleetConfig.LeadRollover cfg) {
Path path = Path.of(p.handoverPath());
if (!Files.exists(path)) {
return RollDecision.refused(RefusalReason.HANDOVER_MISSING,
"handover file " + p.handoverPath() + " does not exist");
}
long size;
long mtimeMillis;
try {
size = Files.size(path);
mtimeMillis = Files.getLastModifiedTime(path).toMillis();
} catch (IOException e) {
throw new UncheckedIOException("failed to stat handover file " + p.handoverPath(), e);
}
if (size == 0) {
return RollDecision.refused(RefusalReason.HANDOVER_EMPTY,
"handover file " + p.handoverPath() + " is empty");
}
if (mtimeMillis <= p.requestedAtMillis()) {
return RollDecision.refused(RefusalReason.HANDOVER_STALE,
"handover file " + p.handoverPath() + " was not modified after the open() "
+ "request (mtime=" + mtimeMillis + "ms, requestedAt=" + p.requestedAtMillis() + "ms)");
}
long ageMillis = nowMillis.getAsLong() - mtimeMillis;
long maxAgeMillis = TimeUnit.SECONDS.toMillis(cfg.maxDocAgeSeconds());
if (ageMillis > maxAgeMillis) {
return RollDecision.refused(RefusalReason.HANDOVER_STALE,
"handover file " + p.handoverPath() + " is " + TimeUnit.MILLISECONDS.toSeconds(ageMillis)
+ "s old, older than maxDocAgeSeconds=" + cfg.maxDocAgeSeconds());
}
return null;
}
/**
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
*
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
* turn" — and it accepts {@link AgentStatus#BLOCKED} for that purpose, because a pane paused on
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
* reason, on the second wait.)
*
* @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real
* boundary was observed, {@code false} if {@code settleSeconds} elapses first.
* {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis}
* clock, never the configured {@code settleSeconds} budget (fleetd #494 follow-up).
*/
private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle: {}",
target, e.toString());
status = null;
}
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis);
}
settleSleeper.run();
}
return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis);
}
/**
* The measured outcome of {@link #waitUntilAtTurnBoundary} — fleetd #494 follow-up. The sibling
* of {@link ClearSettleResult} for the FIRST wait, which never nudges, so it carries no nudge
* count.
*/
private record TurnSettleResult(boolean settled, long elapsedMillis) {}
/**
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
* own javadoc already records that the submit accompanying a delivery "can race the paste —
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
* it, landing as one concatenated line.
*
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
* <ul>
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
* turn;</li>
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
* Claude Code prompt is a no-op, so repeating it is safe;</li>
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
* observed, releases rather than wedges the roll instead of nudging again — the same
* choice {@code Injector} makes — and returns {@code settled() == true} anyway, logged at
* {@code warn} with the measured elapsed time (fleetd #494) so an operator can see which
* path ran and how long it actually took;</li>
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
* </ul>
*
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
* a real boundary is reached or {@code settleSeconds} runs out.
*
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
* nudge must not abort the roll.
*
* @return a {@link ClearSettleResult} whose {@code settled()} is {@code true} once {@code
* /clear} has settled, or once the nudge budget was exhausted with no pickup ever
* observed (released rather than wedged); {@code false} if {@code settleSeconds} elapses
* first — the caller must NOT send {@code bootstrapText} in that case, exactly as before
* this fix. {@code elapsedMillis()} and {@code nudges()} are MEASURED values (from the
* injected {@link #nowMillis} clock and an actual nudge count), never the configured
* {@code settleSeconds} budget (fleetd #494).
*/
private ClearSettleResult waitForClearPickupAndSettle(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
int idlePollsAwaitingPickup = 0;
int nudges = 0;
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
+ "/clear: {}", target, e.toString());
status = null;
}
if (status == AgentStatus.WORKING) {
pickedUp = true;
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
if (pickedUp) {
// a real WORKING -> IDLE/DONE completion boundary
return new ClearSettleResult(true, nowMillis.getAsLong() - startMillis, nudges);
}
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
long elapsedMillis = nowMillis.getAsLong() - startMillis;
// fleetd #494: this release trades a possibly-unsubmitted /clear for progress
// instead of wedging the roll — that trade is deliberate and stays. But it is
// also exactly the case that reported false success in the real incident (the
// whole roll "succeeded" after 438ms of a 20s budget), so raise it to WARN and
// print the MEASURED elapsed time next to the target pane, not just the count.
//
// fleetd #494 follow-up (2nd pass): BOTH numbers in this line must come from
// the loop's own counters, never from the PICKUP_GRACE_POLLS constant.
// `idlePollsAwaitingPickup` and `nudges` each have exactly one write site in
// this loop, on the same branch, so on this branch they cannot differ from
// PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 today — no test can prove the
// difference on this line, and printing the counters does not change that.
// What it does buy: one source of truth instead of two, so a later change to
// the loop (an early return, a second increment site, a different exit
// condition) cannot leave this message reporting a number the loop no longer
// produces. The place where `nudges` genuinely varies with the run — and is
// covered by a test that can tell it apart from a constant — is the
// /clear-timeout warn in runRollover, which prints clearResult.nudges().
log.warn("lead-rollover: /clear on {} was never observed as WORKING after {} "
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
+ "releasing rather than wedging the roll (elapsed={}ms)",
target, idlePollsAwaitingPickup, nudges, elapsedMillis);
return new ClearSettleResult(true, elapsedMillis, nudges);
}
try {
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
} catch (RuntimeException e) {
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
target, e.getMessage());
} finally {
nudges++; // an attempted nudge, whether or not the submit call itself threw
}
}
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
// signal nor a boundary — keep polling without nudging or releasing.
settleSleeper.run();
}
return new ClearSettleResult(false, nowMillis.getAsLong() - startMillis, nudges);
}
/**
* The measured outcome of {@link #waitForClearPickupAndSettle} — fleetd #494. Carries the
* MEASURED elapsed time (from the injected {@link #nowMillis} clock) and nudge count alongside
* the settle/timeout decision, so callers can log them instead of the configured budget, which
* is not how long the wait actually ran.
*/
private record ClearSettleResult(boolean settled, long elapsedMillis, int nudges) {}
}
@@ -0,0 +1,81 @@
package dev.ltms.fleet.mcp;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* fleetd #469: a launch charter that names an MCP tool the server does not register must stop the
* daemon at startup, not wait for a member to discover the gap by calling something that is not
* there.
*
* <p>Deliberately its own class outside {@code dev.ltms.fleet.config}, not a case in {@link
* dev.ltms.fleet.config.FleetConfig#validateCharters()}. The canonical tool surface ({@link
* FleetTool}) lives in the {@code mcp} package; config is loaded before the MCP server exists and
* must not gain a dependency on it. So this check belongs at the seam that already holds both a
* loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}, called right after
* {@code cfg.validateAll()} and before anything opens a socket or spawns a member.
*
* <p>{@code #464}'s {@code CharterToolSurfaceTest} proved the same comparison against a charter
* fixture it wrote itself into a {@code @TempDir}, which meant nothing anyone wrote into the live
* {@code fleetd.yaml} could ever fail it. This class is what a real charter is actually checked
* against at boot; {@code FleetdStartupValidationTest} exercises it through {@code Fleetd.main}
* itself, the same way it proves every other {@code validateXxx()} still runs there.
*
* <p><strong>fleetd #474</strong> — startup was not the only door: {@code
* dev.ltms.fleet.config.ConfigRef#reload()} used to run {@code FleetConfig#validateAll()} alone,
* which does not look at what a charter's text names, so a charter naming an unregistered tool
* that could not have booted the daemon could still be installed into a running one through a
* reload. This class still knows nothing about {@code ConfigRef} — {@code Fleetd.main} wires a
* small {@code FleetConfig -> void} adapter over {@link #assertChartersNameOnlyRegisteredTools}
* ({@code Fleetd::assertChartersNameOnlyRegisteredTools}) into {@code ConfigRef}'s constructor as
* its {@code Consumer<FleetConfig>} {@code extraValidation}, run inside {@code reload()}'s same
* try/catch as {@code validateAll()}, so both call sites — {@code Fleetd.main} at startup and
* {@code ConfigRef#reload()} afterwards — go through this one method and can never check different
* things. {@code dev.ltms.fleet.config.ConfigRefTest} and {@code
* FleetdConfigRefCharterToolSurfaceWiringTest} are what prove the reload call site, the same way
* {@code FleetdStartupValidationTest} proves the startup one.
*/
public final class CharterToolSurface {
/** A {@code fleet_…} (current) or {@code bridge_…} (pre-CB-634) tool-shaped token in prose. */
private static final Pattern TOOL_REFERENCE = Pattern.compile("(fleet_[a-z_]+|bridge_[a-z_]+)");
private CharterToolSurface() {
}
/**
* @param charters the configured {@code fleet.charters:} map (role wire name → charter text);
* {@code null} or empty is a no-op, same as an absent {@code fleet:} block
* @throws IllegalStateException naming the charter key and every tool it names that {@link
* FleetTool} does not list, when any charter does so
*/
public static void assertChartersNameOnlyRegisteredTools(Map<String, String> charters) {
if (charters == null || charters.isEmpty()) {
return;
}
Set<String> registered = FleetTool.wireNames();
List<String> bad = new ArrayList<>();
charters.forEach((key, text) -> {
if (text == null) {
return;
}
Set<String> named = new LinkedHashSet<>();
Matcher m = TOOL_REFERENCE.matcher(text);
while (m.find()) {
named.add(m.group(1));
}
named.stream()
.filter(t -> !registered.contains(t))
.forEach(unknown -> bad.add("fleet.charters." + key + " names '" + unknown
+ "', which the server does not register (registered: " + registered + ")."));
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
}
@@ -33,9 +33,10 @@ public final class ConnectionIdentity {
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
* the pane scan behind {@code terminal} ran to completion ({@link #scanComplete}).
*/
public record Caller(String terminal, long pid) {
public record Caller(String terminal, long pid, boolean scanComplete) {
/**
* Whether the OS peer-PID lookup actually succeeded — {@code false} means {@code pid} is
@@ -51,6 +52,10 @@ public final class ConnectionIdentity {
* {@link ConnectionIdentity#isLoopback} is centralised rather than left for each caller to
* reimplement: a raw {@code pid > 0} check duplicated at every call site is precisely the
* "one rule, two copies" shape that let #305 drift.
*
* <p>This method is deliberately NOT widened for fleetd #505's failure (a herdr error
* during the pane scan, not a failed lsof lookup) — it still tests only the sentinel it is
* named for. #505 is a different axis, carried separately in {@link #scanComplete}.
*/
public boolean resolved() {
return pid > 0;
@@ -60,10 +65,11 @@ public final class ConnectionIdentity {
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
public Caller resolve(String remoteAddr, int remotePort) {
if (!isLoopback(remoteAddr)) {
return new Caller(null, -1); // only same-host callers can be workers
return new Caller(null, -1, true); // only same-host callers can be workers
}
long pid = pids.pidForLocalPort(remotePort);
return new Caller(panes.terminalForPid(pid), pid);
PaneLocator.Lookup lookup = panes.terminalForPid(pid);
return new Caller(lookup.terminal(), pid, lookup.complete());
}
/**
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,83 @@
package dev.ltms.fleet.mcp;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
/**
* The one canonical set of MCP tool names this daemon registers (fleetd #469, follow-up to #464).
*
* <p>Before this enum, the tool surface was written twice with nothing tying the copies together:
* once as the literal {@code "fleet_…"} string passed to each tool-schema builder in
* {@link FleetMcp}, and again as the case labels of {@link FleetMcp}'s authorization switch. A
* reader that needed "what does this server register" — a charter check, in particular — had no
* source to ask except scraping {@code FleetMcp.java}'s source text for {@code tool("…")} calls: a
* third copy of the same list, and the weakest of the three forms.
*
* <p>Every reader that needs the registered tool surface now asks this enum instead:
*
* <ul>
* <li>the tool-schema builders in {@code FleetMcp} pass {@code wireName()} rather than a literal;
* <li>{@code FleetMcp}'s constructor asserts, at startup, that the set of tool names it actually
* registers with the MCP SDK equals {@link #wireNames()} exactly — a canonical entry that is
* never registered, or a registration with no canonical entry backing it, fails the daemon's
* own boot rather than only a test's;
* <li>{@code FleetMcp.toolAction}'s dispatch onto {@code Authz.Action} switches on the enum
* (not the raw string) with no {@code default}, so adding a tool here without pinning its
* action is a compile error, not a run-time throw;
* <li>{@link CharterToolSurface} asks {@link #wireNames()} to check a configured launch charter
* against the live tool surface, instead of scraping source text a third time.
* </ul>
*/
public enum FleetTool {
SEND("fleet_send"),
REPLY("fleet_reply"),
ASK("fleet_ask"),
STATUS("fleet_status"),
POLL("fleet_poll"),
ACK("fleet_ack"),
SPAWN("fleet_spawn"),
LIST("fleet_list"),
STOP("fleet_stop"),
PROFILES("fleet_profiles"),
WHOAMI("fleet_whoami"),
HANDOVER("fleet_handover");
private final String wireName;
FleetTool(String wireName) {
this.wireName = wireName;
}
/** The name this tool is registered under, and called by, on the wire ({@code "fleet_send"}, …). */
public String wireName() {
return wireName;
}
private static final Map<String, FleetTool> BY_WIRE_NAME;
private static final Set<String> WIRE_NAMES;
static {
Map<String, FleetTool> byName = new LinkedHashMap<>();
Set<String> names = new LinkedHashSet<>();
for (FleetTool tool : values()) {
byName.put(tool.wireName, tool);
names.add(tool.wireName);
}
BY_WIRE_NAME = Map.copyOf(byName);
WIRE_NAMES = Set.copyOf(names);
}
/** The tool named {@code wireName}, or empty when this daemon registers no such tool. */
public static Optional<FleetTool> byWireName(String wireName) {
return Optional.ofNullable(BY_WIRE_NAME.get(wireName));
}
/** Every wire name this daemon registers — the canonical tool surface. */
public static Set<String> wireNames() {
return WIRE_NAMES;
}
}
@@ -14,6 +14,7 @@ import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementCandidate;
import dev.ltms.fleet.placement.PlacementContext;
import dev.ltms.fleet.placement.PlacementDecision;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.placement.PlacementPolicies;
import dev.ltms.fleet.placement.PlacementPolicy;
@@ -90,6 +91,19 @@ public final class CompositePeerLauncher implements PeerLauncher {
private final Supplier<Map<String, FleetConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/**
* fleetd #422: the central model allow-list's on/off state, read live per spawn — same reason
* {@link #profileConfigs} is a supplier rather than a captured map (see the class doc above and
* {@link #enforceModelEnabled}). A caller with no {@code models:} block to read from (the
* simpler, map-based constructors used throughout this class's own tests) wires this to a
* constant {@code null}, which {@link #models0} treats as "nothing configured, gate never
* fires" — the pre-#422 behaviour.
*/
private final Supplier<FleetConfig.Models> models;
/** The value {@link #models0} normalizes a {@code null} supplier result to. */
private static final FleetConfig.Models NO_MODELS_CONFIGURED = new FleetConfig.Models(List.of());
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
private final BackendQuarantine quarantine;
@@ -205,12 +219,44 @@ public final class CompositePeerLauncher implements PeerLauncher {
FleetConfig.Fleet fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy) {
// No models: block to read from a plain profiles Map — fleetd #422's gate is wired to a
// constant null (see enforceModelEnabled/models0), the pre-#422 behaviour for every caller
// of this overload. Use the 9-arg overload below to test the gate against a LIVE supplier.
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, quarantine,
outagePolicy, constant(null));
}
/**
* As above, plus a LIVE model-gate source (fleetd #422). The 8-arg overload above wires
* {@code models} to a constant {@code null} because it has only a static {@code Map<String,
* Profile>}, never a full {@code FleetConfig}, to read one from; this overload exists so a test
* can prove {@link #enforceModelEnabled} and its candidate filter re-read {@link
* FleetConfig.Models} on every call rather than a value captured once at construction — the
* exact distinction fleetd #422 exists to get right (see the class doc's CB-559 note on {@link
* #profileConfigs}, which this follows). Production wiring uses the
* {@code Supplier<FleetConfig>} constructor below instead, which already threads a live
* {@code config.get().models()} through.
*
* @param models required — pass a supplier returning {@code null} for a caller that has no
* {@code models:} block to gate against, never a defaulting overload (the same
* "explicit opt-out, never a silent default" rule {@code quarantine}/
* {@code outagePolicy} already follow).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, FleetConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
FleetConfig.Fleet fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy,
Supplier<FleetConfig.Models> models) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy);
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy, models);
}
/**
@@ -250,7 +296,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
liveCount,
() -> config.get().fleet(),
quarantine,
outagePolicy);
outagePolicy,
// fleetd #422: read live, same as profiles/placement/fleet above — a models.allow
// edit (on/off or otherwise) is visible to the very next spawn, no restart needed.
() -> config.get().models());
}
/** The all-suppliers form every other constructor funnels into. */
@@ -261,7 +310,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
Function<String, Integer> liveCount,
Supplier<FleetConfig.Fleet> fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy) {
BackendOutagePolicy outagePolicy,
Supplier<FleetConfig.Models> models) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.outagePolicy = Objects.requireNonNull(outagePolicy, "outagePolicy");
@@ -273,6 +323,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
this.profileConfigs = profileConfigs;
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
this.models = Objects.requireNonNull(models, "models");
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
@@ -303,6 +354,16 @@ public final class CompositePeerLauncher implements PeerLauncher {
return m == null ? Map.of() : m;
}
/**
* The currently-configured {@code models:} block, never null (fleetd #422). Read fresh on
* every call, the same reason {@link #profiles0} is — a config reload's on/off edit must reach
* the very next spawn.
*/
private FleetConfig.Models models0() {
FleetConfig.Models m = models.get();
return m == null ? NO_MODELS_CONFIGURED : m;
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
@@ -332,6 +393,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
enforceNotQuarantined(requestedProfile);
enforceNotCoolingOff(requestedProfile);
enforceMaxLoad(requestedProfile);
enforceModelEnabled(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
@@ -341,18 +403,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
// operator overriding, and refusing it would break `fleet_spawn{profile:"opus"}`, which
// carries no role and so would be judged against the dev pool it was never meant for.
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
Set<String> unreachable = new HashSet<>();
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
Set<String> coolingOff = coolingOffProfiles(candidates);
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff);
PlacementContext ctx = placementContextFor(req.role(), unreachable);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
int maxAttempts = ctx.candidates().isEmpty() ? 1 : ctx.candidates().size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
// policy already throws a clear message. Catching it to rethrow a generic
@@ -369,8 +423,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
// CB-547a: route the chosen profile but keep the caller's session identity — dropping it
// here would silently sever the resume handle on every policy-routed spawn. CB-557: the
// role rides along for the same reason, or a routed spawn would be labelled as a dev.
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId(), req.role());
SpawnRequest routedReq = req.withProfile(chosen.profile());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
@@ -380,8 +433,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff);
ctx = placementContextFor(req.role(), unreachable);
}
}
@@ -493,6 +545,49 @@ public final class CompositePeerLauncher implements PeerLauncher {
}
}
/**
* Refuse an explicit-profile spawn whose {@code model:} the operator has turned off in the
* central {@code models.allow:} list (fleetd #422). Deliberately worded apart from {@link
* #enforceNotQuarantined} and {@link #enforceNotCoolingOff}: those two report a BACKEND-reported
* outage (exhaustion, repeated errors); this one reports an OPERATOR decision, so the message
* says "turned off" and names the model, never "quarantined" or "cooling off". A FOURTH,
* independent reason to refuse a spawn — never layered onto {@code BackendQuarantine} or {@code
* BackendOutagePolicy}, which would misattribute an operator's own choice to the backend.
*
* <p>Reads {@link #models0()} fresh on every call — the same liveness {@link #profiles0()}
* already has — so flipping {@code enabled: false} and reloading takes effect on the very next
* spawn, no restart (criterion 4). A profile naming no model, or a model absent from {@code
* models.allow:} entirely (nothing to gate against), is never refused here.
*
* @throws PlacementException naming the model and the profile, distinct from quarantine/cool-off
*/
private void enforceModelEnabled(String profile) {
FleetConfig.Profile cfg = profiles0().get(profile);
String model = (cfg == null) ? null : cfg.model();
if (model == null) {
return;
}
if (models0().offIds().contains(model)) {
throw new PlacementException("worker profile '" + profile + "' names model '" + model
+ "', which the operator has turned off in models.allow — refusing spawn");
}
}
/** The subset of {@code candidates} whose {@code model:} is currently turned off (fleetd #422). */
private Set<String> modelOffProfiles(List<PlacementCandidate> candidates) {
Set<String> off = models0().offIds();
if (off.isEmpty()) {
return Set.of();
}
return candidates.stream()
.map(PlacementCandidate::profile)
.filter(p -> {
FleetConfig.Profile cfg = profiles0().get(p);
return cfg != null && cfg.model() != null && off.contains(cfg.model());
})
.collect(Collectors.toSet());
}
/**
* The profile names {@code role} may be placed on, in definition order.
*
@@ -509,8 +604,23 @@ public final class CompositePeerLauncher implements PeerLauncher {
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
}
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
private String defaultProfileFor(MemberRole role) {
/**
* {@inheritDoc}
*
* <p>Live: reads {@link #poolFor}, which reads {@link #profileConfigs} and {@link #fleet} fresh
* on every call, so a config reload is visible without a restart (fleetd #425) — unlike {@link
* #defaultProfile}, the field captured once at construction, which this falls back to only when
* {@link #poolFor} has nothing to offer at all (no profiles configured for this composite).
*
* <p>Exact only under the {@code fixed} placement policy — the one that reads this value
* ({@code FixedPlacementPolicy}, package-private, hence not linked) as its first, preferred
* candidate. {@code weighted}/{@code round-robin} placement can choose a different candidate
* from {@code role}'s pool even on the very first spawn; this method does not simulate that
* choice, matching what the {@code defaultProfile:}-derived reporting this replaces has always
* done.
*/
@Override
public String defaultProfileFor(MemberRole role) {
List<String> pool = poolFor(role);
return pool.isEmpty() ? defaultProfile : pool.getFirst();
}
@@ -527,6 +637,142 @@ public final class CompositePeerLauncher implements PeerLauncher {
return out;
}
/**
* Build the {@link PlacementContext} an unqualified spawn of {@code role} would be judged
* against right now — the single source both {@link #spawn} and {@link #place} read, so the two
* can never disagree about which conditions (quarantine, cool-off, model-off) apply to which
* candidate (fleetd #425 rework: round 1 duplicated this into a second, blind resolver —
* {@link #defaultProfileFor} — which is why it regressed; round 2 found that even a single
* shared resolver is not enough on its own if the CALLER re-resolves through an explicit
* profile afterwards — see {@link PlacementDecision}).
*
* @param unreachable the caller's mutable unreachable set; {@link #spawn} grows this across
* retries and rebuilds the context from it, {@link #place} passes a fresh
* empty one since it never retries
*/
private PlacementContext placementContextFor(MemberRole role, Set<String> unreachable) {
List<PlacementCandidate> candidates = candidates(role);
String roleDefault = defaultProfileFor(role);
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
Set<String> coolingOff = coolingOffProfiles(candidates);
// fleetd #422: read live per spawn, same as quarantined/coolingOff above — a config reload
// that flips a model's enabled state is visible to the very next unqualified spawn.
Set<String> modelOff = modelOffProfiles(candidates);
return new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff, modelOff);
}
/**
* {@inheritDoc}
*
* <p>fleetd #425 rework, round 2: runs the exact same selection {@link #spawn} uses for a
* blank-profile request — {@link #placementContextFor} plus one {@link PlacementPolicy#select}
* — rather than {@link #defaultProfileFor}'s blind "pool's first entry", so a quarantined,
* cooling-off, or model-off pool-first candidate is routed around here exactly as it would be
* by a real spawn. Unlike {@link #spawn}, this never retries on {@link
* PeerUnreachableException}: there is no spawn attempt to fail, so "unreachable" never grows
* past the empty set it starts with, and a single {@link PlacementPolicy#select} call already
* reflects the live quarantine/cool-off/model-off state.
*
* <p>Deliberately does <em>not</em> apply {@link #enforceMaxLoad} (or any of the other three
* {@code enforce*} checks): those belong to {@link #spawn}'s EXPLICIT-profile branch, the
* operator-override path, and this method answers a different question — "where would an
* UNQUALIFIED spawn land". That is not the same as {@code select} ignoring these conditions —
* every condition {@code select} filters on (quarantine, cooling off, {@code maxLoad} under
* every placement policy including the default {@code fixed}, since fleetd #435, model-off,
* unreachable, weight-0) is already reflected in the {@link PlacementDecision} this method
* returns, because {@code select} walked past every excluded candidate to find it. What this
* method's caller must not do is take that resolved name and hand it back to {@link
* #spawn(SpawnRequest)} as an explicit profile: the explicit-profile branch treats the same
* exclusion conditions as a reason to REFUSE, where {@code select} had already treated them as
* a reason to fall through — round 1 of this fix did exactly that, turning a fall-through this
* method had already resolved around into a refusal one call later. Round 2 fixes that at the
* caller: {@link #spawn(SpawnRequest, PlacementDecision)} carries this exact decision to the
* spawn without re-resolving or re-checking it, through the same routing path {@code select}
* itself was consulted from.
*
* @throws PlacementException if no candidate in {@code role}'s pool is currently placeable
* (mirrors what an actual unqualified spawn would throw)
*/
@Override
public PlacementDecision place(MemberRole role) {
PlacementContext ctx = placementContextFor(role, new HashSet<>());
return new PlacementDecision(placementPolicy.get().select(ctx).profile());
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #place}, so the two can never disagree about the answer for the same
* {@code role} at the same instant — kept as a convenience for a caller that only wants the
* resolved name (a status report, a log line), never for a caller that will act on it by
* spawning: that caller must hold the {@link PlacementDecision} itself and pass it to {@link
* #spawn(SpawnRequest, PlacementDecision)} — see {@link PlacementDecision}'s javadoc for why
* resolving here and spawning separately, with the name fed back in as an explicit profile,
* regressed fleetd #425 twice.
*/
@Override
public String routedProfileFor(MemberRole role) {
return place(role).profile();
}
/**
* {@inheritDoc}
*
* <p>Routes {@code decision.profile()} directly to its owning delegate — the identical
* {@code d.spawn(routedReq)} call {@link #spawn(SpawnRequest)}'s blank-profile branch makes for
* its first pick — WITHOUT re-running {@link #enforceNotQuarantined}, {@link
* #enforceNotCoolingOff}, {@link #enforceMaxLoad}, or {@link #enforceModelEnabled}: those are
* the EXPLICIT-profile branch's checks, and {@code decision} did not come from an operator
* naming a profile — it came from {@link #place}, which already applied whichever of these
* conditions {@link PlacementPolicy#select} actually filters on (fleetd #425 rework, round 2).
*
* <p>The two branches disagree on purpose about what an excluded profile means, and that
* disagreement is not what this method removes. The blank-profile routing branch (and
* {@link #place}) treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0 profile
* as a reason to fall through to the next candidate; the EXPLICIT-profile branch treats naming
* that same profile as a reason to refuse outright — someone who names a profile should get a
* refusal, not a silent substitution onto a different backend. That is still correct after
* fleetd #435. What round 1 got wrong, and what this method exists to stop happening again, is
* turning a fall-through into a refusal by accident: resolving a name via {@link #place} and
* then handing that same name back to {@link #spawn(SpawnRequest)} as an explicit profile takes
* the refusing branch on a decision the routing branch had already approved by falling through
* past everything else.
*
* <p>Before fleetd #435, this exact accident was reachable through {@code maxLoad} specifically:
* {@code FixedPlacementPolicy} — the default policy — did not evaluate {@code maxLoad} at all
* for automatic selection, so {@link #place} could approve an at-cap profile that {@link
* #enforceMaxLoad} would then refuse one call later. fleetd #435 closed that: {@code
* FixedPlacementPolicy} now walks past an at-cap candidate exactly like {@code weighted}/
* {@code round-robin} already did, so {@link #place} can no longer return one, and this specific
* failure — an approved placement dying at {@code enforceMaxLoad} — cannot happen any more.
* What this method still buys, now that {@code maxLoad} can no longer cause it: it never
* re-evaluates a condition {@link #place} already decided, and it closes the window between
* that decision and the spawn in which the underlying state (another spawn landing on the same
* profile, a config reload) could otherwise move and make a stale explicit re-check wrong.
*
* <p>Deliberately does not retry on {@link PeerUnreachableException} across candidates the way
* {@link #spawn(SpawnRequest)}'s blank-profile branch does: retrying here would silently
* re-place the caller onto a different profile than the one {@code decision} named, behind the
* back of a caller that may already have provisioned something (a worktree's {@code repoRoot},
* parity overlay) specifically for that name. A caller that wants the composite's own failover
* should call {@link #spawn(SpawnRequest)} with a blank profile directly, not resolve through
* {@link #place} first. Losing that retry on a resolve-then-spawn path is an accepted, unrelated
* cost — see {@code SessionManager.acquireWithWorktree}'s own comment on it — never widened by
* this round to include {@code maxLoad}, which is what round 1 actually lost.
*/
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
HerdrPeerLauncher d = route(decision.profile());
SpawnRequest routedReq = req.withProfile(decision.profile());
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
@@ -537,6 +783,32 @@ public final class CompositePeerLauncher implements PeerLauncher {
return route(profileName).parityOverlay(profileName);
}
/**
* fleetd #422 follow-up: the single live read that answers both "is the models.allow: gate
* armed" and "which models are off", off the exact same accessor ({@link #models0()}) {@link
* #enforceModelEnabled} and {@link #modelOffProfiles} read — so {@code fleet_profiles}/{@code
* GET /profiles} (via {@link PeerLauncher#disabledModels()}, which now delegates here) can
* never report a different answer than the gate enforces (the fleetd #404 lesson), and the
* startup log line built from this can never disagree with either.
*
* <p>{@link #models0()} itself normalizes a {@code null} {@link #models} read to the shared
* {@link #NO_MODELS_CONFIGURED} sentinel — deliberately the one object no config-supplied
* {@code Models} instance can ever be identical to, since it is private to this class — so
* comparing by reference here recovers exactly the fact {@code models0()}'s normalization
* would otherwise erase: whether the live source was {@code null} (no {@code models:} block,
* armed = false) or a real, config-supplied block (armed = true, even one whose {@code allow:}
* is itself empty or absent — {@link FleetConfig.Models}'s "absent or empty allow: is off"
* wording governs config-load validation, a distinct question from whether this gate is armed
* for reporting).
*/
@Override
public PeerLauncher.ModelGateState modelGateState() {
FleetConfig.Models m = models0();
return m == NO_MODELS_CONFIGURED
? PeerLauncher.ModelGateState.notConfigured()
: PeerLauncher.ModelGateState.armed(m.offIds());
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.get(id);
@@ -648,9 +920,21 @@ public final class CompositePeerLauncher implements PeerLauncher {
return byProfile.keySet();
}
/**
* {@inheritDoc}
*
* <p>fleetd #425: reports the <em>live</em> {@code dev} pool's first entry — the same value
* {@link #defaultProfileFor} computes for {@link MemberRole#DEV} — not the {@link
* #defaultProfile} field captured at construction. An unqualified {@code fleet_spawn} defaults
* to {@code MemberRole#DEV} (see {@link dev.ltms.fleet.peer.SpawnRequest}), so "the dev pool's
* live first entry" is exactly the profile such a spawn actually lands on right now — the
* question {@code fleet_profiles}' {@code "default"} field exists to answer. The frozen field is
* a role-agnostic fallback used only when {@link #poolFor} has nothing to report at all (no
* profiles configured), which {@link #defaultProfileFor} already handles.
*/
@Override
public String defaultProfile() {
return defaultProfile;
return defaultProfileFor(MemberRole.DEV);
}
/**
@@ -13,6 +13,7 @@ import java.nio.file.attribute.PosixFilePermissions;
import java.time.Duration;
import java.time.Instant;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Set;
import java.util.stream.Stream;
@@ -28,33 +29,85 @@ import java.util.stream.Stream;
* control lacked: herdr applies that overlay BEFORE the shell starts, so any sourced file can undo
* it — and did.
*
* <p><b>Which file is last depends on the platform, so the scrub runs from two of them.</b> zsh
* <p><b>zsh reads its four startup files under three different conditions, so no single file is
* guaranteed to run — the scrub has to cover the gap between them, not just the platforms.</b> zsh
* reads {@code .zshenv} always, {@code .zprofile} and {@code .zlogin} only for a LOGIN shell, and
* {@code .zshrc} only for an INTERACTIVE one. herdr does not open the same kind of shell
* everywhere — measured on herdr 0.8.0: macOS panes run {@code -zsh} (login, so {@code .zlogin}
* runs), Linux panes run a plain {@code /usr/bin/zsh} (interactive but NOT login, so
* {@code .zlogin} never runs at all). A scrub in {@code .zlogin} alone is therefore a control that
* silently does nothing on Linux — the exact failure this class exists to remove, one platform
* over.
* {@code .zshrc} only for an INTERACTIVE one. A pane shell that is at least one of login or
* interactive is covered by sourcing the scrub from {@code .zshrc} and {@code .zlogin} (below), but
* a pane shell that is NEITHER reads only {@code .zshenv} and stops — fleetd #388, measured: a herdr
* pane can be neither login nor interactive, and such a pane read {@code .zshenv}, never reached
* {@code scrub.zsh}, and left no report at all. A bare {@code argv[0]} of {@code /usr/bin/zsh}
* proves the shell is NOT a login shell; it says nothing about whether it is interactive, so it
* must never be read as "therefore interactive" — that wrong inference is what let #388 ship.
*
* <p>So both {@code .zshrc} and {@code .zlogin} source the same generated {@code scrub.zsh} after
* sourcing their {@code $HOME} counterpart. On Linux only the first fires; on macOS both do, and
* the second pass is deliberate rather than merely harmless — it re-scrubs anything the operator's
* own {@code ~/.zlogin} exported after {@code .zshrc} had finished. Re-running is idempotent: a
* name already blank is blanked again, and the report is rewritten with the same counts.
* <p>So {@code .zshenv} carries a THIRD pass, guarded by the exact condition that defines the gap:
* {@code [[ ! -o login && ! -o interactive ]]}. That guard is why this pass cannot double-scrub a
* pane that {@code .zshrc} or {@code .zlogin} will also cover — one of {@code -o login}/
* {@code -o interactive} is always true there, so the {@code .zshenv} pass never fires for them, and
* their own unconditional sourcing is untouched. The guard also carries a sentinel
* ({@value #SCRUB_SENTINEL}) so it fires once per PANE and not once per PROCESS: {@code .zshenv} is
* read by every zsh a member's own tooling forks (a plain {@code zsh -c '...'} for a single
* command is itself neither login nor interactive), and those children inherit variables their
* parent deliberately set for them (git hooks get {@code GIT_DIR}, a venv gets
* {@code VIRTUAL_ENV}, a build tool gets {@code NODE_OPTIONS} or {@code JAVA_TOOL_OPTIONS}).
* Re-scrubbing every such child would blank all of that, and would also make the pane's own
* {@code scrub-report.txt} — rewritten on every pass — describe whichever child exited last
* instead of the pane. The sentinel is exported only AFTER {@code scrub.zsh} runs, so the pass
* that sets it never sees it and cannot blank it; it must also be on the scrub's own allow-list
* (see {@link #generate(Path, Set)}) so a later pass, in the same pane, cannot blank it back to
* empty — an exported-but-empty sentinel reads as unset to the {@code -z} guard and would silently
* re-enable scrubbing for every subsequent child of that pane.
*
* <p>So all three of {@code .zshenv} (gap only, guarded), {@code .zshrc}, and {@code .zlogin}
* source the same generated {@code scrub.zsh} after sourcing their {@code $HOME} counterpart. A
* login-and-interactive pane runs the {@code .zshrc} and {@code .zlogin} passes, and the second is
* deliberate rather than merely harmless — it re-scrubs anything the operator's own
* {@code ~/.zlogin} exported after {@code .zshrc} had finished. A pane that is neither runs only the
* {@code .zshenv} pass. Re-running is idempotent: a name already blank is blanked again, and the
* report is rewritten with the same counts.
*
* <p>Each generated file sources its {@code $HOME} counterpart FIRST, so {@code PATH} and every
* toolchain binary still resolve exactly as the operator configured them; only afterwards does
* {@code .zlogin} run the scrub: every EXPORTED variable not on the derived allow-list is re-exported
* toolchain binary still resolve exactly as the operator configured them; only afterwards does the
* scrub run: every EXPORTED variable not on the derived allow-list is re-exported
* blank. Blank, not credential-shaped-pattern-filtered: a pattern list ({@code *TOKEN*}, …) is an
* enumeration and misses what it did not think of — a username is the other half of a credential and
* is shaped like none. Credential-SHAPED names among the blanked set go to the WARN log only,
* never to the control.
*
* <p>The scrub also writes {@code scrub-report.txt} into its own directory: one {@code allowed N of
* M} line (N = exports left untouched, M = exports present when the scrub ran), then the blanked
* NAMES — never values. The launcher reads this back at teardown and logs it, because a blocked
* count next to an unknown denominator is not a finding.
* M failed F} line (N = exports left untouched, M = exports present when the scrub ran, F = names
* the scrub attempted to blank but could not), then the NAMES — blanked ones bare, unblankable ones
* {@code !}-prefixed — never values. The launcher reads this back at teardown and logs it, because a
* blocked count next to an unknown denominator is not a finding.
*
* <p><b>fleetd #394:</b> plain {@code export "$n="} is a FATAL error for a zsh read-only or special
* parameter (for example {@code UID}) — it aborts the whole sourced file, so every name still to
* come is never blanked and the report above is never written at all. The blanking loop instead
* routes each attempt through {@code eval}, which contains that error to the single iteration: the
* loop always finishes, and a name that could not be blanked is counted as {@code failed} and
* listed {@code !}-prefixed rather than silently disappearing. This is deliberately not a skip-list
* of known-bad names — every enumerated name is still attempted, so a name nobody has thought of
* yet still gets tried and, if it fails, still gets counted.
*
* <p>The blanking loop also re-asserts, on its own, the same {@code [A-Za-z_][A-Za-z0-9_]*} shape
* check the enumeration loop already applied. Before {@code eval} was introduced a non-conforming
* name reaching {@code export "$n="} was harmless either way — the quoting made it inert. With
* {@code eval}, the name is spliced into a string and interpreted as shell syntax, so the enumeration
* loop's check is no longer sufficient on its own to keep that call site safe — it is a guard on a
* different loop, and the two must not silently drift apart. Re-checking right before the
* {@code eval} keeps that call site safe by its own reading, independent of whatever the enumeration
* loop does or stops doing in a later change.
*
* <p><b>fleetd #400:</b> {@code eval}'s exit status is not proof that the blank actually happened.
* zsh coerces a bare {@code NAME=} assignment on an integer special parameter (measured on macOS zsh
* 5.9: {@code SECONDS}, {@code RANDOM}, {@code SHLVL}, {@code HISTSIZE}, {@code COLUMNS},
* {@code LINES}, {@code USERNAME}) to a number instead of failing — {@code eval} returns success,
* the value is untouched, and a status-based classification reports it as blanked when it was not.
* The fix classifies on the observed effect instead: after the attempt, the name's value is read
* back with the {@code (P)} indirection flag and the decision is made from whether that is now
* empty. This one check covers all three shapes a name can take at this point — a genuine blank, a
* fatal read-only error {@code eval} merely contained, and this silent no-op — and the exit status
* plays no part in the decision at all.
*/
public final class EnvAllowListScrub {
@@ -63,7 +116,11 @@ public final class EnvAllowListScrub {
/** Name of the report file written into the generated directory by the scrub itself. */
static final String REPORT_FILE = "scrub-report.txt";
/** The scrub body, generated once and sourced from both {@code .zshrc} and {@code .zlogin}. */
/**
* The scrub body, generated once and sourced from {@code .zshrc} and {@code .zlogin}
* unconditionally, and from {@code .zshenv} when the pane shell is neither login nor
* interactive (fleetd #388) — see the class javadoc.
*/
static final String SCRUB_FILE = "scrub.zsh";
/** Prefix of every generated directory — also what {@link #reapOrphans} matches on. */
@@ -79,14 +136,42 @@ public final class EnvAllowListScrub {
private static final String SOURCE_SCRUB =
"source \"$ZDOTDIR/" + SCRUB_FILE + "\"\n";
/**
* fleetd #388: marks a pane, not a process, as already scrubbed. Set only by the guarded
* {@code .zshenv} pass (see {@link #NEITHER_LOGIN_NOR_INTERACTIVE_SCRUB}) after
* {@code scrub.zsh} has run, so it must also be folded into that pass's own allow-list — see
* the class javadoc's "must also be on the scrub's own allow-list" paragraph.
*/
static final String SCRUB_SENTINEL = "_CB633_SCRUBBED";
/**
* Appended to {@code .zshenv}, after its {@code $HOME} source: the third pass, guarded on the
* exact condition that defines the gap {@code .zshrc}/{@code .zlogin} do not cover — a shell
* that is neither login nor interactive. The sentinel export happens only once the scrub has
* already run, and only for as long as the current pane's environment has not been rebuilt from
* scratch (a fresh {@code env -i} child would not inherit it — that is out of scope here, since
* such a child is no longer running under the pane's own environment at all).
*/
private static final String NEITHER_LOGIN_NOR_INTERACTIVE_SCRUB =
"if [[ ! -o login && ! -o interactive && -z \"${" + SCRUB_SENTINEL + ":-}\" ]]; then\n"
+ " " + SOURCE_SCRUB
+ " export " + SCRUB_SENTINEL + "=1\n"
+ "fi\n";
private EnvAllowListScrub() {
}
/**
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran,
* how many were left untouched (allowed), and the NAMES that were blanked. Values never appear.
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran
* ({@code total}), how many were left untouched ({@code allowed}), how many the scrub attempted
* to blank but could not ({@code failed} — fleetd #394: a zsh read-only/special parameter such
* as {@code UID} fatally errors on plain {@code export NAME=}, so those attempts go through
* {@code eval} instead so the loop keeps going and the failure is counted rather than left
* invisible), and the NAMES in each of the latter two categories. {@code allowed +
* blanked.size() + unblankable.size() == total}, and {@code unblankable.size() == failed}.
* Values never appear.
*/
record ScrubReport(int allowed, int total, List<String> blanked) {
record ScrubReport(int allowed, int total, int failed, List<String> blanked, List<String> unblankable) {
}
/**
@@ -105,11 +190,16 @@ public final class EnvAllowListScrub {
reapOrphans(parentDir);
Path dir = Files.createTempDirectory(parentDir, DIR_PREFIX);
dir.toFile().deleteOnExit();
// The report is written by zsh, after these hooks are registered, so register its path
// too — otherwise the directory is non-empty at JVM exit and cannot be removed at all.
dir.resolve(REPORT_FILE).toFile().deleteOnExit();
write(dir, SCRUB_FILE, scrubScript(allowedNames));
write(dir, ".zshenv", homeSourcingFile(".zshenv"));
// zsh truncates this pre-created receipt after these hooks are registered. Register its
// path too — otherwise the directory is non-empty at JVM exit and cannot be removed.
Files.createFile(dir.resolve(REPORT_FILE)).toFile().deleteOnExit();
// fleetd #388: scrub.zsh's OWN allow-list must also keep SCRUB_SENTINEL, or a later
// pass in the same pane blanks it back to empty and the .zshenv guard below thinks it
// was never scrubbed — see the class javadoc.
Set<String> namesForScrubScript = new HashSet<>(allowedNames);
namesForScrubScript.add(SCRUB_SENTINEL);
write(dir, SCRUB_FILE, scrubScript(namesForScrubScript));
write(dir, ".zshenv", homeSourcingFile(".zshenv") + NEITHER_LOGIN_NOR_INTERACTIVE_SCRUB);
write(dir, ".zprofile", homeSourcingFile(".zprofile"));
write(dir, ".zshrc", homeSourcingFile(".zshrc") + SOURCE_SCRUB);
write(dir, ".zlogin", homeSourcingFile(".zlogin") + SOURCE_SCRUB);
@@ -147,12 +237,11 @@ public final class EnvAllowListScrub {
* into it: owner keeps full access, {@code group} gets traverse+read on the directory ({@code
* rwxr-x---}, so a member process — a login shell reading it via {@code ZDOTDIR}, or another
* process simply opening a file under it — running under that group can find and read the
* files) and read-only on each file ({@code rw-r-----}) — deliberately no group WRITE anywhere,
* since a member never needs to add or change what fleetd generated. (For the ZDOTDIR scrub
* specifically, this also means the scrub script's own report write inside the pane fails
* closed rather than open — see {@code scrub.zsh}'s trailing {@code 2>/dev/null} — which {@link
* dev.ltms.fleet.member.HerdrPeerLauncher#releaseZdotdir} already treats as "cannot be
* confirmed to have run" rather than success.)
* files) and read-only on each file ({@code rw-r-----}), except the pre-created ZDOTDIR
* {@code scrub-report.txt}. That receipt gets group write ({@code rw-rw----}), so
* {@code scrub.zsh} can truncate and write it without granting group write on the directory.
* If its optional permission change fails, the member cannot write a receipt and the launcher
* keeps its existing WARN rather than failing the spawn.
*
* <p>Package-private and named generically on purpose: fleetd #213 built this for the ZDOTDIR
* scrub directory, and fleetd #219 reuses it verbatim for {@link
@@ -167,7 +256,15 @@ public final class EnvAllowListScrub {
setGroupAndPermissions(dir, principal, "rwxr-x---");
try (Stream<Path> entries = Files.list(dir)) {
for (Path file : entries.toList()) {
setGroupAndPermissions(file, principal, "rw-r-----");
if (REPORT_FILE.equals(file.getFileName().toString())) {
try {
setGroupAndPermissions(file, principal, "rw-rw----");
} catch (IOException | UnsupportedOperationException ignored) {
// The receipt is optional. Its absence keeps the existing WARN path.
}
} else {
setGroupAndPermissions(file, principal, "rw-r-----");
}
}
}
} catch (IOException e) {
@@ -245,11 +342,13 @@ public final class EnvAllowListScrub {
}
return """
# generated by fleetd (CB-633 memberCredentials policy=allow-list) — do not edit.
# Sourced from .zshrc and again from .zlogin, each time AFTER that file has sourced
# its $HOME counterpart — so this runs after everything the operator sourced, on a
# login shell (macOS panes) and on a plain interactive one (Linux panes) alike.
# Running twice is idempotent and deliberate: the second pass catches anything
# ~/.zlogin exported after ~/.zshrc had finished.
# Sourced unconditionally from .zshrc and again from .zlogin, each time AFTER that
# file has sourced its $HOME counterpart — so this runs after everything the
# operator sourced, on any pane that is login and/or interactive. Running twice is
# idempotent and deliberate: the second pass catches anything ~/.zlogin exported
# after ~/.zshrc had finished. Also sourced, once, from a guarded pass in .zshenv
# (fleetd #388) when the pane shell is NEITHER login nor interactive — the one gap
# those two files do not cover.
typeset -A _cb633_allowed
for _cb633_n in %s; do _cb633_allowed[$_cb633_n]=1; done
@@ -271,15 +370,57 @@ public final class EnvAllowListScrub {
_cb633_blank+=("$_cb633_n")
done
{ for _cb633_n in "${_cb633_blank[@]}"; do export "$_cb633_n="; done; } 2>/dev/null
# fleetd #394: plain `export "$n="` is FATAL for a zsh read-only/special parameter
# (e.g. UID) and aborts this whole sourced file — every name still to come is never
# blanked, and the report below is never written, silently. `eval` contains that
# error to the single iteration instead: it still fails for that one name, but the
# loop continues and we can tell allowed / blanked / unblankable apart afterwards.
# This is not a skip-list of known-bad names (that would miss the next one nobody
# thought of) — every name in _cb633_blank is still attempted, unconditionally.
# Every name reaching this loop already passed the identical identifier check in the
# enumeration loop above — but that guard is 20 lines away in a different loop, and
# this line is about to splice the name into a string handed to `eval`. Before this
# fix the name only ever reached `export` quoted ("$n="), which is inert on a
# non-identifier string either way; `eval` makes THIS line the only thing standing
# between such a string and code execution in the member's pane, so it re-asserts the
# same check on its own rather than trusting a guard it does not own. Under normal
# operation this can never fire (the enumeration guard already filtered everything
# reaching _cb633_blank), so a name caught here is counted as unblankable rather than
# silently dropped — it is real evidence that the upstream guard was bypassed.
#
# fleetd #400: the attempt's own exit status is NOT proof of its effect. zsh coerces
# a bare `NAME=` assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
# HISTSIZE, COLUMNS, LINES, USERNAME on this host) to a number instead of failing —
# `eval` returns 0, the value is untouched, and the old exit-status check reported it
# as blanked when it was not. Classify on the observed effect instead: attempt the
# export, then read the name's value back with the `(P)` indirection flag and decide
# from whether it is now empty. One check then covers all three shapes a name can
# take here — a genuine blank, a fatal read-only error `eval` merely contained, and
# this silent no-op — without the exit status entering the decision at all.
typeset -a _cb633_ok _cb633_unblankable
_cb633_ok=()
_cb633_unblankable=()
for _cb633_n in "${_cb633_blank[@]}"; do
if [[ ! "$_cb633_n" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
_cb633_unblankable+=("$_cb633_n")
continue
fi
eval "export ${_cb633_n}=" 2>/dev/null
if [[ -z "${(P)_cb633_n}" ]]; then
_cb633_ok+=("$_cb633_n")
else
_cb633_unblankable+=("$_cb633_n")
fi
done
integer _cb633_kept=$(( _cb633_total - ${#_cb633_blank} ))
{
print -r -- "allowed $_cb633_kept of $_cb633_total"
for _cb633_n in "${_cb633_blank[@]}"; do print -r -- "$_cb633_n"; done
print -r -- "allowed $_cb633_kept of $_cb633_total failed ${#_cb633_unblankable}"
for _cb633_n in "${_cb633_ok[@]}"; do print -r -- "$_cb633_n"; done
for _cb633_n in "${_cb633_unblankable[@]}"; do print -r -- "!$_cb633_n"; done
} > "$ZDOTDIR/%s" 2>/dev/null
unset _cb633_allowed _cb633_names _cb633_blank _cb633_n _cb633_total _cb633_kept
unset _cb633_allowed _cb633_names _cb633_blank _cb633_ok _cb633_unblankable _cb633_n _cb633_total _cb633_kept
""".formatted(names, MemberEnvAllowList.zshCasePattern(), REPORT_FILE);
}
@@ -293,6 +434,11 @@ public final class EnvAllowListScrub {
* Read and parse {@link #REPORT_FILE} out of a generated ZDOTDIR directory. Returns {@code null}
* when absent or unreadable (the pane may have been torn down before its login shell ever got to
* the scrub) — callers treat that as "no measurement available", never as success.
*
* <p>First line is {@code "allowed <N> of <M> failed <F>"} (fleetd #394 added the trailing
* {@code failed <F>} — a count of names the scrub attempted to blank but could not, e.g. a zsh
* read-only/special parameter). Every following non-blank line is a name: a bare name was
* blanked, a {@code !}-prefixed name was attempted and failed. Values never appear on either.
*/
static ScrubReport readReport(Path zdotdir) {
Path report = zdotdir.resolve(REPORT_FILE);
@@ -305,17 +451,24 @@ public final class EnvAllowListScrub {
return null;
}
String[] parts = lines.getFirst().substring("allowed ".length()).trim().split("\\s+");
if (parts.length != 3 || !"of".equals(parts[1])) {
if (parts.length != 5 || !"of".equals(parts[1]) || !"failed".equals(parts[3])) {
return null;
}
List<String> blanked = new ArrayList<>();
List<String> unblankable = new ArrayList<>();
for (int i = 1; i < lines.size(); i++) {
if (!lines.get(i).isBlank()) {
blanked.add(lines.get(i));
String line = lines.get(i);
if (line.isBlank()) {
continue;
}
if (line.startsWith("!")) {
unblankable.add(line.substring(1));
} else {
blanked.add(line);
}
}
return new ScrubReport(Integer.parseInt(parts[0]), Integer.parseInt(parts[2]),
List.copyOf(blanked));
Integer.parseInt(parts[4]), List.copyOf(blanked), List.copyOf(unblankable));
} catch (IOException | NumberFormatException e) {
return null;
}
@@ -17,6 +17,7 @@ import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -594,6 +595,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
}
/**
* {@inheritDoc}
*
* <p>fleetd #450: re-enters {@link #spawn(SpawnRequest)} with {@code decision}'s profile named
* explicitly. This is the re-entering form the interface javadoc describes for a launcher with
* no placement concept of its own — an instance of this class spawns a single adapter's own
* profile set by explicit name only ({@link #place}/{@link #defaultProfileFor} are unoverridden
* here and just wrap {@link #defaultProfile()}); it does no quarantine/cool-off/maxLoad/model-off
* filtering of its own to re-apply. That filtering lives one layer up, in {@code
* CompositePeerLauncher}, which is the launcher that routes across more than one profile and
* therefore overrides this method with the routing form instead.
*/
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
return spawn(req.withProfile(decision.profile()));
}
/** The herdr daemon that owns this launcher's pane coordinates. */
public HerdrClient herdr() {
return agents.herdr();
@@ -937,14 +955,20 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
@Override
public void stop(String idOrPane) {
// Teardown knows only the pane, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
// Teardown knows only the pane, not which profile spawned it — and, when this call is
// routed here through CompositePeerLauncher's single-daemon stop() shortcut (fleetd #342),
// not even which adapter's config actually governed the spawn: the shortcut can hand the
// pane to a delegate that never spawned it, whose own profiles say nothing about how THIS
// pane was placed. So the decision to look for a tab to clean up is made from the pane's
// actual state, not from this delegate's static profile config: resolve the tab
// unconditionally and let {@link WorkspaceControl#locatePane} tolerate "not found" (it
// returns null rather than throwing); the single-occupant check below is what actually
// protects the user's shared tabs, exactly as it always has.
String paneId = paneByAgentId.remove(idOrPane);
if (paneId == null) {
paneId = idOrPane; // raw-pane fallback (reap, gate timeout, pane-addressed callers)
}
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
WorkspaceControl.PaneLocation loc = spaces.locatePane(paneId);
try {
agents.close(paneId);
} catch (HerdrException e) {
@@ -985,11 +1009,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(FleetConfig.Profile::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
@@ -1642,11 +1661,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
log.warn("memberCredentials allow-list: pane {} left no scrub report in {} — the "
+ "environment scrub cannot be confirmed to have run. Either the pane ended "
+ "before its shell finished starting, or its shell never read our generated "
+ "startup files, in which case that member saw the full host environment.",
+ "startup files. Either way, we cannot tell from here whether the scrub ran, "
+ "so we do not know what that member's environment contained.",
paneId, dir);
} else {
log.info("memberCredentials allow-list: pane {} allowed {} of {} environment variables",
paneId, report.allowed(), report.total());
if (report.failed() > 0) {
log.warn("memberCredentials allow-list: pane {} could not blank {} environment "
+ "variable(s) — {} (likely a zsh read-only/special parameter) — those "
+ "names were left in the member's environment. Confirm none of them is a "
+ "credential.",
paneId, report.failed(), report.unblankable());
}
List<String> shaped = report.blanked().stream()
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
.toList();
@@ -331,11 +331,22 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
+ "form — opencode's per-model context limit could not be applied for this profile",
cfg.profile(), cfg.model());
}
// fleetd #393: memberSkills seeding (GitWorktrees#seedSkills) copies skill folders into
// EVERY provisioned worktree's .claude/skills/ regardless of which kind ultimately spawns
// into it — that copy step cannot know the kind, only the caller of GitWorktrees#add does
// (see that method's own javadoc). .claude/skills/ is a Claude Code CLI convention the CLI
// discovers on its own; opencode has no such discovery, so without this, a seeded skill
// never reaches an opencode member even though GitWorktrees logged it as seeded. Read
// whatever landed under <cwd>/.claude/skills/ here — the one place in this launcher that
// knows both the kind (opencode, by construction: this IS OpenCodeLauncher) and the cwd.
List<Path> skillInstructionFiles = skillInstructionFiles(spec.cwd());
// A config file is needed for the bridge MCP mount, a member charter, the IDE MCP (+ its
// guidance overlay, CB-634), a pinned endpoint (CB-508), or a resolvable autoCompactWindow.
// guidance overlay, CB-634), a pinned endpoint (CB-508), a resolvable autoCompactWindow, or
// at least one seeded skill to deliver via instructions[] (fleetd #393).
if (cfg.hasMcp() || cfg.hasIdeMcp() || spec.charter() != null || hasCustomProvider(cfg)
|| wantsContextLimit) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter(), spec.cwd()).toString());
|| wantsContextLimit || !skillInstructionFiles.isEmpty()) {
workerEnv.put("OPENCODE_CONFIG",
writeConfig(cfg, spec.charter(), spec.cwd(), skillInstructionFiles).toString());
}
applyGitToken(workerEnv, cfg);
List<String> argv = argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId());
@@ -358,6 +369,65 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return withAgent;
}
/**
* fleetd #393: the {@code SKILL.md} paths under {@code <cwd>/.claude/skills/} this launcher can
* turn into {@code instructions[]} entries, plus the honest log this ticket asks for — emitted
* here, at the one point a skill's fate for THIS spawn is actually known, rather than trusting
* {@code GitWorktrees#seedSkills}'s kind-blind "N of M" line to mean "and it will be read."
*
* <p>Every non-hidden subdirectory of {@code .claude/skills/} is a candidate, whether it got
* there via {@code memberSkills:} seeding or because the target repo ships its own copy — this
* launcher does not care which; it only cares what it can find at spawn time. A candidate with
* a {@code SKILL.md} at its top level (the same shape {@link #writeConfig} already requires for
* the charter and IDE-rules instructions entries) is delivered; anything else is a directory
* this launcher cannot turn into a flat instructions entry, named explicitly in the log rather
* than silently dropped, so a caller sees a real "cannot consume" reason and not just a smaller
* number than {@code GitWorktrees}' own count.
*
* <p>No candidates at all (directory absent or empty) logs nothing — the same
* no-log-when-nothing-to-say shape {@link #hasCustomProvider} and friends already follow, and
* the shape {@code GitWorktrees#seedSkills} itself uses when {@code memberSkills:} is unset.
* A failure to even list the directory is logged and treated as "nothing delivered" — best
* effort, must never fail the spawn, matching {@code GitWorktrees#seedSkills}'s own contract.
*/
private List<Path> skillInstructionFiles(String cwd) {
if (cwd == null || cwd.isBlank()) {
return List.of();
}
Path skillsDir = Path.of(cwd, ".claude", "skills");
if (!Files.isDirectory(skillsDir)) {
return List.of();
}
List<Path> candidates;
try (var listing = Files.list(skillsDir)) {
candidates = listing.filter(Files::isDirectory)
.filter(p -> !p.getFileName().toString().startsWith("."))
.sorted()
.toList();
} catch (IOException e) {
log.warn("could not scan {} for skill folders to deliver to this opencode member: {}",
skillsDir, e.getMessage());
return List.of();
}
if (candidates.isEmpty()) {
return List.of();
}
List<Path> delivered = candidates.stream()
.map(dir -> dir.resolve("SKILL.md"))
.filter(Files::isRegularFile)
.toList();
List<String> undeliverable = candidates.stream()
.filter(dir -> !Files.isRegularFile(dir.resolve("SKILL.md")))
.map(dir -> dir.getFileName().toString())
.toList();
log.info("skill delivery: {} of {} skill folder(s) under {} reached this opencode member via "
+ "instructions[] (opencode does not read .claude/skills/ natively, unlike "
+ "Claude Code){}",
delivered.size(), candidates.size(), skillsDir,
undeliverable.isEmpty() ? "" : "; no SKILL.md, could not be delivered: " + undeliverable);
return delivered;
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
@@ -418,8 +488,15 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*
* @param skillInstructionFiles fleetd #393: absolute {@code SKILL.md} paths from
* {@link #skillInstructionFiles(String)}, appended to
* {@code instructions[]} so a {@code memberSkills:}-seeded skill
* reaches this opencode member the same way the charter and IDE
* rules already do.
*/
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd,
List<Path> skillInstructionFiles) {
try {
Path dir = Files.createTempDirectory(configParentDir(), "fleetd-opencode-");
dir.toFile().deleteOnExit();
@@ -447,7 +524,39 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Files.writeString(charter, charterText);
charter.toFile().deleteOnExit();
root.putArray("instructions").add(charter.toAbsolutePath().toString());
// fleetd #393 follow-up: withArray, not putArray. putArray REPLACES whatever node
// is already at "instructions"; withArray gets-or-creates. All three writers on
// this array now use withArray so that write order stops being load-bearing for
// whoever adds a fourth.
//
// Be precise about what this particular line is worth, because an earlier version
// of this comment was wrong and claimed too much. THIS call is the one place where
// the two idioms are equivalent, and no test can tell them apart: it runs first,
// against a still-empty root, so there is never an existing node for putArray to
// replace. Measured on the #393 merge: flipping this one call back to putArray
// leaves the whole suite green (1618 tests), and always will. The edit is a
// readability and future-proofing change with no test behind it, and that is not a
// gap anyone can close.
//
// The ordering hazard is real, just not here. It is the LATER writers that can
// destroy earlier entries. Measured on the same merge: flipping the skills writer
// below to putArray deletes this charter entry and fails
// OpenCodeLauncherTest.instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder;
// flipping the IDE-rules writer fails that test and
// instructionsArrayHoldsCharterThenIdeRulesInOrder. Deleting this line altogether
// fails instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt — the
// assertion that a role contract reaches an opencode member at all, which is the
// hole that predates #393 and is what actually let the mutation hide.
root.withArray("instructions").add(charter.toAbsolutePath().toString());
}
// fleetd #393: each seeded skill's SKILL.md, delivered as a plain instructions[] entry
// — the only mechanism opencode has for static guidance text. Unlike Claude Code's
// Skill tool, opencode cannot load one of these on demand by name; the content is just
// always part of the system prompt from spawn. That is a real difference in HOW the
// content reaches the member, not a reason to withhold it.
for (Path skillFile : skillInstructionFiles) {
root.withArray("instructions").add(skillFile.toAbsolutePath().toString());
}
if (cfg.hasMcp() || cfg.hasIdeMcp()) {
@@ -56,10 +56,13 @@ public final class FleetMetrics {
m.describe(REPLIES, "counter",
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
"CB-307 push-loop nudges to the primary (sent|exhausted). fleetd #365: \"sent\" means "
+ "the herdr paste-and-submit call succeeded, not that the pane read it — this "
+ "layer has no read-receipt concept.");
m.describe(HEARTBEAT_NUDGES, "counter",
"CB-551 idle-lead heartbeat nudges (delivered|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down.");
"CB-551 idle-lead heartbeat nudges (sent|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down. "
+ "fleetd #365: \"sent\" means the herdr call succeeded, not that the lead read it.");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
@@ -8,6 +8,7 @@ import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import com.rabbitmq.client.Return;
import com.rabbitmq.client.impl.DefaultExceptionHandler;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -153,17 +154,22 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
/** As {@link #open(String)}, with an explicit consumer prefetch (CB-527: caps the held backlog per target). */
public static AmqpReplyInbox open(String uri, int prefetch) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("fleetd-reply-inbox"), prefetch);
return new AmqpReplyInbox(connectionFactory(uri).newConnection(AmqpConnectionFailureLogger.REPLY_INBOX), prefetch);
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
static ConnectionFactory connectionFactory(String uri) throws Exception {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
factory.setExceptionHandler(new AmqpConnectionFailureLogger(AmqpConnectionFailureLogger.REPLY_INBOX, log));
return factory;
}
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this(connection, DEFAULT_PREFETCH);
@@ -388,17 +394,17 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
@Override
public void ack(String target, String msgId) {
public boolean ack(String target, String msgId) {
var perTarget = held.get(target);
if (perTarget == null || perTarget == RELEASED) {
return;
return false;
}
Held h;
synchronized (perTarget) {
h = perTarget.remove(msgId);
}
if (h == null) {
return; // never held (or already acked) — no-op
return false; // never held (or already acked) — no-op
}
try {
synchronized (channelLock) {
@@ -412,6 +418,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
}
return true;
}
private DeliverCallback deliverCallback(String target) {
@@ -571,3 +578,44 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
}
}
/**
* Keeps RabbitMQ's forgiving exception behaviour while adding the connection identity that its
* default logger drops. Package-private so both AMQP connections use the same two names.
*/
final class AmqpConnectionFailureLogger extends DefaultExceptionHandler {
static final String REPLY_INBOX = "fleetd-reply-inbox";
static final String LEAD_MAILBOX = "fleetd-lead-mailbox";
private final String connectionName;
private final Logger logger;
AmqpConnectionFailureLogger(String connectionName, Logger logger) {
this.connectionName = connectionName;
this.logger = logger;
}
String connectionName() {
return connectionName;
}
@Override
protected void log(String message, Throwable cause) {
if (isSocketClosedOrConnectionReset(cause)) {
logger.warn("AMQP connection {}: {} (Exception message: {})", connectionName, message, cause.getMessage());
} else {
logger.error("AMQP connection {}: {}", connectionName, message, cause);
}
}
private static boolean isSocketClosedOrConnectionReset(Throwable cause) {
// Deliberate copy of ForgivingExceptionHandler's private static helper; check it on amqp-client upgrades.
if (!(cause instanceof IOException)) {
return false;
}
return "Connection reset".equals(cause.getMessage())
|| "Socket closed".equals(cause.getMessage())
|| "Connection reset by peer".equals(cause.getMessage());
}
}
@@ -59,16 +59,17 @@ public final class InMemoryReplyInbox implements ReplyInbox {
}
@Override
public void ack(String target, String msgId) {
public boolean ack(String target, String msgId) {
if (!owned.contains(target)) {
return;
return false;
}
var perTarget = store.get(target);
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.remove(msgId);
}
if (perTarget == null) {
return false;
}
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
return perTarget.remove(msgId) != null;
}
}
}
@@ -4,8 +4,8 @@ import java.util.List;
/**
* The lead-to-lead message channel this daemon speaks, as its callers need it — one lead's own
* mailbox: publish to a peer's coord-id, look at what has arrived for me, and ack what I have
* delivered.
* mailbox: publish to a peer's coord-id, look at what has arrived for me, ack what I have
* delivered, and (fleetd #361) inspect any coord-id's mailbox from the outside without owning it.
*
* <p>Extracted from {@link LeadMailbox} purely as a seam. {@code LeadMailbox} is the one production
* implementation and owns a live AMQP connection, so a test that wanted to exercise the routing in
@@ -32,9 +32,99 @@ public interface LeadChannel {
/** Non-destructive FIFO snapshot of the messages held for this daemon's own coord-id. */
List<LeadMessage> peek();
/** Drop {@code msgId} from the held set and ack it on the broker. A no-op if it is not held. */
/**
* Drop {@code msgId} from the held set and ack it on the broker. A repeated ack that this
* connection already completed may be a no-op. Any other unknown msgId must throw rather than
* report an ack that did not reach the broker.
*/
void ack(String msgId);
/** This daemon's own lead coordination id — the mailbox it owns, and the {@code from} it sends as. */
String selfCoordId();
/**
* Whether a message sitting in {@link #peek}'s held set (fetched but not yet {@link #ack}ed) is
* still safe if this daemon crashes or restarts right now — the conclusion of two independent
* facts about how this channel owns its own queue: the queue was declared <em>durable</em>, and
* the consumer that filled {@code held} uses <em>manual ack</em>, so an unacked delivery is still
* owned by the broker rather than only in this process's memory. Both must hold for {@code true};
* an implementation must derive this from what it actually did when it declared and consumed its
* queue, never return a literal — fleetd #440 found {@code FleetMcp}'s {@code heldDurable} field
* doing exactly that, unable to ever report {@code false} even after the fact stopped being true.
*/
boolean heldDurable();
/**
* A non-destructive look at {@code coordId}'s mailbox — does it exist, how many messages are
* waiting on it, and how many consumers are attached — without owning, consuming, or otherwise
* changing it. {@code consumers == 0} on an existing mailbox is the observable form of "nobody
* is reading this right now": a publish to it will sit queued rather than reach a pane.
*
* <p><strong>Never throws</strong> — this is a best-effort fact-finding call, not an operation a
* caller must handle failing. But it must never turn "I could not check" into a false negative:
* {@link MailboxState#absent(String)} means the broker positively confirmed there is no such
* queue, and {@link MailboxState#unknown(String)} — a distinct value — means the look could not
* be completed at all (broker unreachable, timed out, connection closed). A caller that
* collapses those two into one, as fleetd #361 initially did, cannot tell "that peer is down"
* from "I could not check", and a reader of {@code pending}/{@code consumers} cannot tell a
* measured zero from a zero standing in for "not measured".
*
* <p><strong>Must never share fate with {@link #publish} or {@link #peek}/{@link #ack}.</strong>
* fleetd #361: in AMQP 0-9-1 a passive queue declare of a queue that does not exist closes the
* channel it was declared on with a 404. An implementation backed by a real broker connection
* must inspect on a channel it can afford to lose — never the channel {@link #publish} or the
* consume loop depends on — so that looking at a peer that happens to be down can never break
* this daemon's own send or receive path.
*/
MailboxState inspect(String coordId);
/**
* The result of {@link #inspect}. {@code presence} tells apart three states a caller must not
* conflate: a confirmed-existing mailbox ({@link Presence#EXISTS}, the only case where
* {@code pending}/{@code consumers} are measured facts), a confirmed-absent one
* ({@link Presence#ABSENT} — the broker positively said "no such queue"), and one this call
* simply could not determine ({@link Presence#UNKNOWN} — broker unreachable, timed out,
* connection closed). {@code pending}/{@code consumers} are always {@code 0} and meaningless
* outside {@link Presence#EXISTS}; a renderer must gate on {@link #exists()} (or {@code
* presence} directly), never present them as measured otherwise.
*
* @param coordId the coord-id inspected
* @param presence whether the mailbox is confirmed to exist, confirmed absent, or unknown
* @param pending messages ready for delivery but not yet in a consumer's hands (0 unless EXISTS)
* @param consumers how many consumers are attached (0 unless EXISTS)
*/
record MailboxState(String coordId, Presence presence, int pending, int consumers) {
/** Whether {@link #inspect} was able to reach a definite answer, of either kind. */
public enum Presence { EXISTS, ABSENT, UNKNOWN }
/** {@code true} only when the broker confirmed this exact queue is currently declared. */
public boolean exists() {
return presence == Presence.EXISTS;
}
/**
* {@code true} when {@link #inspect} reached a definite answer (exists or confirmed
* absent); {@code false} when it could not determine either way. A caller must never treat
* {@code !known()} the same as a confirmed absence — the mailbox may well exist.
*/
public boolean known() {
return presence != Presence.UNKNOWN;
}
/** The broker confirmed this queue exists, with these measured counts. */
public static MailboxState exists(String coordId, int pending, int consumers) {
return new MailboxState(coordId, Presence.EXISTS, pending, consumers);
}
/** The broker positively confirmed there is no such queue (e.g. a 404 on passive declare). */
public static MailboxState absent(String coordId) {
return new MailboxState(coordId, Presence.ABSENT, 0, 0);
}
/** The look could not be completed — broker unreachable, timed out, or connection closed. */
public static MailboxState unknown(String coordId) {
return new MailboxState(coordId, Presence.UNKNOWN, 0, 0);
}
}
}
@@ -5,6 +5,7 @@ import dev.ltms.fleet.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ScheduledExecutorService;
@@ -43,6 +44,9 @@ public final class LeadCoordLoop {
private static final Logger log = LoggerFactory.getLogger(LeadCoordLoop.class);
/** A bounded window is enough: redelivery can only follow a recent failed ack or recovery. */
private static final int RECENT_DELIVERY_LIMIT = 1_024;
/** How an arriving peer message is rendered into the lead's pane — the sender's coord-id, then its text. */
static final String DELIVERY_FORMAT = "[lead %s] %s";
@@ -51,6 +55,8 @@ public final class LeadCoordLoop {
private final Supplier<Map<String, String>> leads;
private final ScheduledExecutorService scheduler;
private final long intervalMs;
/** msgIds already written to the pane, so recovery redelivery is acked without another pane write. */
private final LinkedHashMap<String, Boolean> delivered = new LinkedHashMap<>();
private volatile boolean running;
@@ -117,6 +123,13 @@ public final class LeadCoordLoop {
if (held.isEmpty()) {
return;
}
LeadMessage msg = held.getFirst();
if (wasDelivered(msg.msgId())) {
// This lives here, rather than in LeadMailbox, because only this loop knows a pane write
// happened. The mailbox only knows broker delivery tags and must still redeliver after a crash.
ackDelivered(msg);
return;
}
String lead = resolveLocalLead();
if (lead == null) {
// Left unacked on purpose: the broker keeps holding it until a lead pane exists.
@@ -136,7 +149,6 @@ public final class LeadCoordLoop {
lead, status, held.size());
return;
}
LeadMessage msg = held.getFirst();
try {
agents.send(lead, DELIVERY_FORMAT.formatted(msg.from(), msg.content()));
} catch (RuntimeException e) {
@@ -145,6 +157,13 @@ public final class LeadCoordLoop {
msg.msgId(), msg.from(), lead, e.toString());
return;
}
rememberDelivered(msg.msgId());
if (ackDelivered(msg)) {
log.debug("lead coordination: delivered message {} from {} to lead {}", msg.msgId(), msg.from(), lead);
}
}
private boolean ackDelivered(LeadMessage msg) {
try {
channel.ack(msg.msgId());
} catch (RuntimeException e) {
@@ -152,9 +171,24 @@ public final class LeadCoordLoop {
// deliberate direction of this trade.
log.warn("lead coordination: delivered message {} but could not ack it: {}",
msg.msgId(), e.toString());
return;
return false;
}
return true;
}
private boolean wasDelivered(String msgId) {
synchronized (delivered) {
return delivered.containsKey(msgId);
}
}
private void rememberDelivered(String msgId) {
synchronized (delivered) {
delivered.put(msgId, Boolean.TRUE);
if (delivered.size() > RECENT_DELIVERY_LIMIT) {
delivered.remove(delivered.keySet().iterator().next());
}
}
log.debug("lead coordination: delivered message {} from {} to lead {}", msg.msgId(), msg.from(), lead);
}
/**
@@ -96,7 +96,14 @@ public final class LeadHeartbeatLoop {
this.metrics = metrics;
}
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
/**
* Count one nudge outcome when a registry is wired; a no-op in unit tests.
*
* <p>fleetd #365: the {@code "sent"} outcome (renamed from {@code "delivered"}) records only
* that {@link #injectNudge} — a one-way herdr {@code agent.prompt} paste-and-submit — returned
* without throwing, not that the lead's pane actually read or acted on the text. This layer has
* no read-receipt concept, so "sent" is the honest word for what this call can ever establish.
*/
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(FleetMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
@@ -245,7 +252,7 @@ public final class LeadHeartbeatLoop {
agents.send(leadTerminal, fleet.nudgeText());
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
leadTerminal, quietCount);
countNudge("delivered");
countNudge("sent");
} catch (RuntimeException e) {
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
countNudge("failed");
@@ -10,6 +10,7 @@ import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import com.rabbitmq.client.Return;
import com.rabbitmq.client.ShutdownSignalException;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -58,7 +59,8 @@ import java.util.concurrent.TimeoutException;
* <p><strong>Recovery.</strong> The connection is opened with automatic + topology recovery
* enabled, mirroring {@code AmqpReplyInbox}: on reconnect the broker hands out fresh delivery tags,
* so the held snapshot is cleared (dedup by {@code msgId} still prevents any double-queue on
* redelivery) and any publish still awaiting its confirm is failed rather than left to idle out
* redelivery). {@link LeadCoordLoop} separately deduplicates pane writes, since it alone knows
* which messages reached a lead. Any publish still awaiting its confirm is failed rather than left to idle out
* the confirm timeout against a sequence number that means nothing on the new channel.
*/
public final class LeadMailbox implements LeadChannel, AutoCloseable {
@@ -84,6 +86,16 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
private final Object channelLock = new Object();
/** msgId → held delivery, for this mailbox's own queue only (there is exactly one). */
private final LinkedHashMap<String, Held> held = new LinkedHashMap<>();
/**
* fleetd #440: the answer to {@link #heldDurable()}, set once by {@link #own()} from the exact
* booleans it passed to {@code queueDeclare}/{@code basicConsume} — never a separate literal that
* could drift from what those calls actually did.
*/
private boolean heldDurable;
/** Successful broker acks on this connection, retained only to make a repeated caller ack quiet. */
private final LinkedHashMap<String, Boolean> recentlyAcked = new LinkedHashMap<>();
/** Bounds {@link #recentlyAcked}: it is only an idempotency aid, never delivery state. */
private static final int RECENT_ACK_LIMIT = 1_024;
/**
* A dedicated channel for {@link #publish}, kept separate from {@link #channel} (consume + ack)
@@ -122,17 +134,22 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
/** As {@link #open(String, String)}, with an explicit consumer prefetch. */
public static LeadMailbox open(String uri, String selfCoordId, int prefetch) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares the queue and re-attaches the consumer.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new LeadMailbox(factory.newConnection("fleetd-lead-mailbox"), selfCoordId, prefetch);
return new LeadMailbox(connectionFactory(uri).newConnection(AmqpConnectionFailureLogger.LEAD_MAILBOX), selfCoordId, prefetch);
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP coordination broker at " + uri, e);
}
}
static ConnectionFactory connectionFactory(String uri) throws Exception {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
factory.setExceptionHandler(new AmqpConnectionFailureLogger(AmqpConnectionFailureLogger.LEAD_MAILBOX, log));
return factory;
}
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for tests). */
LeadMailbox(Connection connection, String selfCoordId) {
this(connection, selfCoordId, DEFAULT_PREFETCH);
@@ -156,16 +173,15 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue). Any publish
// repopulates it with valid tags (dedup by msgId still prevents any double-queue). LeadCoordLoop
// remembers successful pane writes separately, so that redelivery cannot write a pane twice. Any publish
// confirm still in flight when the connection dropped is equally stale — fail it now rather
// than let it silently ride out CONFIRM_TIMEOUT_MS.
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
synchronized (held) {
held.clear();
}
clearHeldForRecovery();
failPendingPublishesOnRecovery();
log.info("AMQP lead mailbox connection recovered; cleared held messages for fresh redelivery");
}
@@ -181,13 +197,22 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
/** Declare + consume this daemon's own {@code lead.<selfCoordId>.inbox}. Called once, at construction. */
private void own() throws IOException {
String queue = queueName(selfCoordId);
boolean durableQueue = true; // durable, non-exclusive, keep on idle
boolean autoAck = false; // manual ack
synchronized (channelLock) {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
channel.basicConsume(queue, false, deliverCallback(), _ -> { }); // autoAck=false: manual ack
channel.queueDeclare(queue, durableQueue, false, false, null);
channel.basicConsume(queue, autoAck, deliverCallback(), _ -> { });
}
// fleetd #440: held mail is durable only while both hold — a durable queue AND manual ack.
this.heldDurable = durableQueue && !autoAck;
log.debug("lead mailbox owns queue {} for coord-id {}", queue, selfCoordId);
}
@Override
public boolean heldDurable() {
return heldDurable;
}
/**
* Publish {@code msg} to {@code toCoordId}'s mailbox and block until the broker's publisher
* confirm for it lands. Does <em>not</em> imply owning or consuming {@code toCoordId}'s queue.
@@ -254,6 +279,89 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
}
}
/**
* fleetd #361: look at {@code coordId}'s mailbox on a fresh, immediately-closed throwaway
* channel — never {@link #channel} (consume/ack) or {@link #publishChannel} (publish). A
* passive queue declare of a queue that does not exist closes the channel it was declared on
* with a 404; using a disposable probe channel means that closure can never touch either
* long-lived channel this instance depends on for {@link #publish} or the consume loop.
*
* <p>Classifies failures rather than collapsing them, both measured against a real broker in
* {@code LeadMailboxTest} rather than assumed from the AMQP 0-9-1 spec text:
* <ul>
* <li>a genuine 404 — an {@link IOException} wrapping a {@link ShutdownSignalException} whose
* {@link AMQP.Channel.Close#getReplyCode()} is {@code 404} — reports
* {@link MailboxState#absent}; every other declare failure reports
* {@link MailboxState#unknown} instead of quietly becoming the same "absent" value;
* <li>{@code catch (RuntimeException e)} on both attempts matters as much as the checked
* catches: a connection that is already closed makes {@link Connection#createChannel()}
* throw {@link com.rabbitmq.client.AlreadyClosedException} (a {@link RuntimeException},
* not an {@link IOException}) — an {@code inspect} that only caught {@code IOException}
* would let that escape, breaking the "never throws" contract this method promises.
* </ul>
*
* <p><strong>Honesty about which catch is measured and which is defensive:</strong> the
* {@code createChannel()} catch above is exercised end-to-end against a real broker by
* {@code LeadMailboxTest.inspectReportsUnknownRatherThanThrowingWhenTheConnectionIsAlreadyClosed}.
* The second {@code catch (RuntimeException e)}, around the passive declare itself — for the
* narrower race where the connection drops <em>between</em> {@code createChannel()} succeeding
* and the declare landing — has no such test; reaching it needs a connection that dies at that
* exact instant, which is not a scenario this suite drives on purpose. It stays purely
* defensive: correct by the same reasoning as the first catch, but unproven the way the first
* one is proven.
*/
@Override
public MailboxState inspect(String coordId) {
String queue = queueName(coordId);
Channel probe;
try {
probe = connection.createChannel();
} catch (IOException | RuntimeException e) {
log.debug("lead mailbox inspect: cannot open a probe channel for {}: {}", coordId, e.toString());
return MailboxState.unknown(coordId);
}
try {
AMQP.Queue.DeclareOk declared = probe.queueDeclarePassive(queue);
return MailboxState.exists(coordId, declared.getMessageCount(), declared.getConsumerCount());
} catch (IOException e) {
// The broker (or the client library) has already closed `probe` for us either way; only
// a confirmed 404 means "no such queue" — anything else (a different declare failure) is
// "could not determine", never silently reported as the same value as a genuine absence.
return isMissingQueue(e) ? MailboxState.absent(coordId) : MailboxState.unknown(coordId);
} catch (RuntimeException e) {
// E.g. the connection dropped between createChannel() and the declare landing.
log.debug("lead mailbox inspect: declare failed unexpectedly for {}: {}", coordId, e.toString());
return MailboxState.unknown(coordId);
} finally {
try {
if (probe.isOpen()) {
probe.close();
}
} catch (Exception e) {
log.debug("lead mailbox inspect: probe channel close for {}: {}", coordId, e.toString());
}
}
}
/**
* {@code true} only for the specific shape a missing-queue passive declare actually produces —
* measured against a real broker, not assumed from the spec text (see {@code
* LeadMailboxTest.passiveDeclareOfAMissingQueueThrowsAnIOExceptionWrappingA404ShutdownSignal}):
* an {@link IOException} whose cause is a {@link ShutdownSignalException} carrying an
* {@link AMQP.Channel.Close} reason with {@code replyCode == 404}. Any other shape (a different
* reply code, a {@code ShutdownSignalException} cause whose reason is not a
* {@code Channel.Close}, or no cause at all) is a declare failure of some other kind and must
* not be read as "confirmed absent" — pinned hermetically, with no broker needed, by
* {@code LeadMailboxIsMissingQueueTest} for exactly those three false shapes. Package-private
* (not {@code private}) so that test can call it directly.
*/
static boolean isMissingQueue(IOException e) {
if (!(e.getCause() instanceof ShutdownSignalException sse)) {
return false;
}
return sse.getReason() instanceof AMQP.Channel.Close close && close.getReplyCode() == AMQP.NOT_FOUND;
}
/** Convenience: {@link #peek} the current snapshot, then {@link #ack} every message in it. */
public List<LeadMessage> drain() {
List<LeadMessage> snapshot = peek();
@@ -261,15 +369,26 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
return snapshot;
}
/** Remove the held message {@code msgId} and ack it on the broker. No-op if not held. */
/**
* Remove the held message {@code msgId} and ack it on the broker.
*
* <p>A repeated ack that this connection already completed is a no-op, tracked in the bounded
* {@link #recentlyAcked} set. Any other missing entry throws: recovery clears {@link #held} while
* the broker still owns the unacked delivery, and quiet success there would hide a required retry.
* The set is bounded because it only distinguishes a recent duplicate caller ack from an unknown
* delivery; it is not a substitute for broker state across a reconnect.
*/
@Override
public void ack(String msgId) {
Held h;
synchronized (held) {
h = held.remove(msgId);
if (h == null && recentlyAcked.containsKey(msgId)) {
return;
}
}
if (h == null) {
return; // never held (or already acked) — no-op
throw new IllegalStateException("cannot ack lead message " + msgId + ": it is not held");
}
try {
synchronized (channelLock) {
@@ -283,6 +402,19 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
}
throw new IllegalStateException("cannot ack lead message " + msgId, e);
}
synchronized (held) {
recentlyAcked.put(msgId, Boolean.TRUE);
if (recentlyAcked.size() > RECENT_ACK_LIMIT) {
recentlyAcked.remove(recentlyAcked.keySet().iterator().next());
}
}
}
/** Clear stale delivery tags after recovery; package-private so the recovery contract test drives this exact path. */
void clearHeldForRecovery() {
synchronized (held) {
held.clear();
}
}
private DeliverCallback deliverCallback() {
@@ -138,6 +138,53 @@ public final class MessageService {
public record AskResult(AskOutcome outcome, String answer) {
}
/**
* How a worker's {@code fleet_reply} ({@link #reply(String, String)}) actually landed
* (fleetd #365) — the two doors that expose it, {@code fleet_reply} and {@code POST
* /sessions/{id}/reply}, both used to report the single word "delivered" whichever of these
* happened, so a caller could not tell an active handoff from a reply merely held for later
* drain. Both are successes; they are not the same fact.
*/
public enum ReplyOutcome {
/** Resolved a {@code fleet_send}/{@code fleet_ask} that was actively waiting on this reply. */
RESOLVED_SEND("resolved_send", true,
"delivered — resolved the fleet_send that was waiting for it"),
/**
* No live waiter was open, but the reply completed a parked async ticket directly
* ({@link #askAnsweredAsyncTasks}) — a {@code fleet_poll} caller sees it immediately.
*/
RESOLVED_ASYNC_TICKET("resolved_async_ticket", true,
"delivered — resolved a pending async ticket (visible to fleet_poll)"),
/** Nothing was waiting; the reply was queued in the inbox for a later drain (CB-307). */
QUEUED("queued", false,
"queued — no send or ticket was waiting; held in the inbox for a later drain");
private final String wireName;
private final boolean delivered;
private final String description;
ReplyOutcome(String wireName, boolean delivered, String description) {
this.wireName = wireName;
this.delivered = delivered;
this.description = description;
}
/** Stable machine-readable name for a JSON/metrics label (REST's {@code outcome} field). */
public String wireName() {
return wireName;
}
/** Whether something was actively waiting and received this reply right now. */
public boolean delivered() {
return delivered;
}
/** Shared human-readable text — the one place both {@code fleet_reply} and REST word this. */
public String description() {
return description;
}
}
/** Lifecycle phase of an async delegation ticket. */
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
@@ -439,16 +486,15 @@ public final class MessageService {
* @throws IllegalArgumentException if {@code content} is {@code null} or blank — the caller must
* report this as a client error (REST: 400 {@code bad_request}) rather than resolve
* anything
* @return always {@code true} — the reply resolved a live send, completed a parked ticket, or
* was queued
* @return which of the three ways (fleetd #365) the reply actually landed — never {@code null}
*/
public boolean reply(String session, String content) {
public ReplyOutcome reply(String session, String content) {
if (content == null || content.isBlank()) {
throw new IllegalArgumentException("content is required");
}
if (rendezvous.resolve(session, content)) {
count(FleetMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
return ReplyOutcome.RESOLVED_SEND; // a live send took it — unchanged fast path
}
// #137/fleetd #307: no live rendezvous waiter, but this may be the worker's real fleet_reply resuming
// a turn that either answer() (#137) or ask() (fleetd #307) already gave up waiting on:
@@ -480,7 +526,7 @@ public final class MessageService {
asyncTasksByTurn.remove(turnId, orphan);
}
count(FleetMetrics.REPLIES, "path", "async-recovered");
return true; // the ticket itself took it — no inbox stranding at all
return ReplyOutcome.RESOLVED_ASYNC_TICKET; // the ticket itself took it — no inbox stranding
}
} else if (candidates.size() > 1) {
List<String> tickets = candidates.stream().map(t -> t.ticket).toList();
@@ -498,7 +544,7 @@ public final class MessageService {
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
return ReplyOutcome.QUEUED; // held, not lost — but not delivered either
}
/**
@@ -780,9 +826,13 @@ public final class MessageService {
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
*
* @return {@code true} if an entry was actually removed, {@code false} if {@code msgId} was not
* in {@code target}'s inbox (wrong id, wrong target, or already acked). The caller —
* {@link dev.ltms.fleet.mcp.FleetMcp#ack} — must not report success on {@code false}.
*/
public void ackReply(String target, String msgId) {
inbox.ack(target, msgId);
public boolean ackReply(String target, String msgId) {
return inbox.ack(target, msgId);
}
/**
@@ -891,6 +941,10 @@ public final class MessageService {
boolean wasDelivered = delivery.completion().isDone()
&& !delivery.completion().isCompletedExceptionally();
if (!wasDelivered) {
if (timeoutCancellationRaceHookForTest != null) {
// Test-only (fleetd #345): see the field's own javadoc.
timeoutCancellationRaceHookForTest.run();
}
// The target monitor makes cancellation atomic with onStatus picking this
// Pending up. If pickup won, report TIMED_OUT_WORKING because the text landed.
wasDelivered = injector.cancel(delivery) == Injector.Cancellation.DELIVERED;
@@ -961,13 +1015,41 @@ public final class MessageService {
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("fleet_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
// fleetd #307: mark the task BEFORE clearAsyncQuestion(forgetTurn=true) below drops it out of
// asyncTasksByTurn and nulls its turnId — that forgetting is deliberate and stays (it is
// what keeps the target from staying BUSY forever), but it would otherwise also erase
// askAnsweredAsyncTasks' only signal that the worker's eventual real fleet_reply still
// belongs to this task, stranding it in the inbox with a false "never replied" verdict.
markAskTimedOut(ticket.turnId());
clearAsyncQuestion(ticket.turnId(), true);
// Only the fresh owner tears down the shared turn (mirrors the finally block below).
// A duplicate's own timeoutMillis says nothing about whether the SHARED ask is actually
// done — it must leave the close/forget bookkeeping to the fresh owner, exactly as it
// already leaves closeAsk to it.
if (ticket.fresh()) {
// fleetd #334: close the ask turn BEFORE forgetting this task's turnId mapping below.
// Before this fix the order was reversed — the mapping was forgotten here first, and
// rendezvous.closeAsk only ran afterward, in the shared finally. A primary's answer()
// call racing this exact timeout could then find rendezvous.askSession(turnId) still
// non-null (the ask still "answerable") after the Task mapping was already gone:
// answer()'s own asyncTasksByTurn lookup returned null, its task != null guard skipped
// the completion, and the async ticket sat at PENDING forever even though answer()
// itself reported the worker's real reply. Closing here first removes that window:
// any answer() call that still observes askSession(turnId) != null is necessarily
// racing a point BEFORE the forgetting below runs (both happen on this one thread, in
// this order, with nothing that yields in between), so the Task mapping is still there
// for it to find; any call that observes askSession(turnId) == null now correctly
// bails out STALE_TURN (see answer()'s own top check) before ever reaching
// asyncTasksByTurn. rendezvous.closeAsk is idempotent — a no-op once the turn is
// already removed, see its own javadoc — so the shared finally below re-running it
// for this same fresh call is harmless.
rendezvous.closeAsk(ticket.turnId());
if (askTimeoutRaceHookForTest != null) {
// Test-only (fleetd #334): see the field's own javadoc.
askTimeoutRaceHookForTest.run();
}
// fleetd #307: mark the task BEFORE clearAsyncQuestion(forgetTurn=true) below drops it
// out of asyncTasksByTurn and nulls its turnId — that forgetting is deliberate and
// stays (it is what keeps the target from staying BUSY forever), but it would
// otherwise also erase askAnsweredAsyncTasks' only signal that the worker's eventual
// real fleet_reply still belongs to this task, stranding it in the inbox with a false
// "never replied" verdict.
markAskTimedOut(ticket.turnId());
clearAsyncQuestion(ticket.turnId(), true);
}
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
@@ -1066,19 +1148,28 @@ public final class MessageService {
// the chained ask deliberately left open.
//
// A null task is NOT only "this was never an async ticket". That reading was in this
// comment when #329 merged and it is wrong. A genuine async ticket also lands here
// with task == null, because ask()'s timeout path runs clearAsyncQuestion(turnId,
// true) — which drops the asyncTasksByTurn entry — in its catch block, while
// rendezvous.closeAsk(turnId) runs later, in its finally. Between those two the ask
// is still answerable but the map entry is already gone, so the lookup at :991
// returns null and this ticket is never completed. Measured on 2026-09-04: a probe
// firing only that first half before answer() runs printed
// "answer=REPLIED phase=PENDING reply=null" — the same stranded ticket #329 set out
// to fix, one step earlier in the same race. The probe used forgetTurnForTest, which
// omits ask()'s markAskTimedOut; that cannot change the outcome, because askTimedOut
// is read only by askAnsweredAsyncTasks, and reply() never reaches it while this
// method's own waiter is live. So #329 narrows this window rather than closing it.
// Open as fleetd #334 — do not read this guard as complete.
// comment when #329 merged and it is wrong; it is still not the whole story after
// #334. A genuine async ticket can still land here with task == null — a blocking
// (wait:true) send's fleet_ask never has a Task at all, so that case is expected and
// fine. What #334 fixed was a SECOND, unintended way to get here with task == null:
// ask()'s timeout path used to run clearAsyncQuestion(turnId, true) — which drops the
// asyncTasksByTurn entry — in its catch block, while rendezvous.closeAsk(turnId) ran
// later, in its finally. Between those two the ask was still answerable but the map
// entry was already gone, so the lookup at :1053 returned null and this ticket was
// never completed. Measured on 2026-09-04: a probe firing only that first half before
// answer() ran printed "answer=REPLIED phase=PENDING reply=null" — the same stranded
// ticket #329 set out to fix, one step earlier in the same race; the probe used
// forgetTurnForTest, which omits ask()'s markAskTimedOut, and that omission does not
// change the outcome, because askTimedOut is read only by askAnsweredAsyncTasks, and
// reply() never reaches it while this method's own waiter is live. #334's fix
// reorders ask()'s timeout catch to run closeAsk before the forgetting (see the
// fresh-owner block there), which removes this path entirely rather than narrowing
// it further: once closeAsk has run, rendezvous.askSession(turnId) is null and
// answer() returns STALE_TURN from its own top check, before it ever reaches this
// lookup — see aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket in
// MessageServiceTest, which pins the exact window with askTimeoutRaceHookForTest.
// So by the time this line runs, task == null means only the ordinary blocking-send
// case (or #329's own already-fixed race elsewhere) — not this one.
if (result.outcome() != Outcome.QUESTION && task != null) {
finishAsyncTask(task, result);
}
@@ -1230,6 +1321,27 @@ public final class MessageService {
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
}
/**
* Test seam only — carries no production behaviour, and nothing in this class calls it;
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
*
* <p>Reports whether {@code ticket}'s completion hook (the {@code whenComplete} registered in
* {@link Task}'s constructor) has actually run yet. Exists because {@link #poll} can report
* {@link Phase#DONE} for a ticket before that hook fires: {@code CompletableFuture.complete()}
* publishes its result and only afterwards runs dependent actions such as {@code whenComplete}
* (fleetd #399), so a caller that observes the future done via {@link #poll} is not thereby
* guaranteed to also observe {@link Task#completedNanos} stamped. A test that must order both
* events — e.g. before advancing an injected clock past the TTL, to avoid stamping the
* *advanced* time and masking a real eviction bug — waits on this instead of on
* {@link Phase#DONE}.
*
* @return {@code false} for an unknown ticket or one whose completion hook has not run yet
*/
boolean isCompletionStampedForTest(String ticket) {
Task task = tasks.get(ticket);
return task != null && task.completedNanos != null;
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
@@ -1401,6 +1513,24 @@ public final class MessageService {
this.afterFinishAsyncTaskCompleteHookForTest = hook;
}
/**
* Null in production; test seam for fleetd #345 — invoked in {@link #send}'s timeout path after
* {@link Injector.Delivery#completion()} reports incomplete and before {@link Injector#cancel}
* takes the target monitor. A test installs this to make {@code onStatus} pick the exact queued
* delivery up in that window, so {@code cancel} returns {@link Injector.Cancellation#DELIVERED}.
* This deterministically covers the caller's need to use that result rather than relying on a
* timing-sensitive real race.
*/
private volatile Runnable timeoutCancellationRaceHookForTest;
/**
* Test-only (fleetd #345): install {@link #timeoutCancellationRaceHookForTest}. Package-private
* so the test, in the same package, can reach it without widening any production API.
*/
void setTimeoutCancellationRaceHookForTest(Runnable hook) {
this.timeoutCancellationRaceHookForTest = hook;
}
/**
* Null in production; test seam for fleetd #329 (F3) — invoked from {@link #reply} right after
* the single local read of {@code orphan.turnId} passes its null-check and before that (now-local)
@@ -1457,6 +1587,29 @@ public final class MessageService {
this.abandonCleanupHookForTest = hook;
}
/**
* Null in production; test seam for fleetd #334 — invoked from {@link #ask}'s {@code
* TimeoutException} catch, only for the fresh owner, right after {@code rendezvous.closeAsk}
* has run and before {@link #markAskTimedOut} / {@link #clearAsyncQuestion} forget this task's
* turnId mapping. A test installs this to call {@link #answer} for the very same {@code turnId}
* synchronously from inside that exact window, deterministically reproducing the race a real
* concurrent {@code answer()} call could otherwise only win by timing luck: with the ask already
* closed, that call must see {@code rendezvous.askSession(turnId) == null} and return {@link
* Outcome#STALE_TURN} immediately, never reaching {@code asyncTasksByTurn} at all — proving the
* window fleetd #334 describes (mapping forgotten while the ask was still "answerable") is
* closed, rather than merely narrowed the way fleetd #329 narrowed the sibling race in {@link
* #finishAsyncTask}.
*/
private volatile Runnable askTimeoutRaceHookForTest;
/**
* Test-only (fleetd #334): install {@link #askTimeoutRaceHookForTest}. Package-private so the
* test, in the same package, can reach it without widening any production API.
*/
void setAskTimeoutRaceHookForTest(Runnable hook) {
this.askTimeoutRaceHookForTest = hook;
}
/** A new send must not open a waiter while an async ticket owns this worker's paused turn. */
private boolean hasAsyncQuestion(String target) {
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
@@ -46,6 +46,13 @@ public interface ReplyInbox {
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
List<InboxMessage> peek(String target);
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
void ack(String target, String msgId);
/**
* Remove the reply {@code msgId} for {@code target} once the primary has taken it.
*
* @return {@code true} if an entry was actually removed, {@code false} if there was nothing to
* remove (unknown {@code target}, unowned {@code target}, or a {@code msgId} not held for
* it). A {@code false} is not an error — acking a {@code target} this daemon does not own is
* part of the normal contract, not a failure.
*/
boolean ack(String target, String msgId);
}
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -13,6 +14,7 @@ import java.util.Collection;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
@@ -130,7 +132,15 @@ public final class ReplyPushLoop {
this.metrics = metrics;
}
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
/**
* Count one nudge outcome when a registry is wired; a no-op in unit tests.
*
* <p>fleetd #365: the {@code "sent"} outcome (renamed from {@code "delivered"}) records only
* that {@code agents.send} — a one-way herdr {@code agent.prompt} paste-and-submit — returned
* without throwing. Nothing in this loop, or anywhere downstream of it, confirms the pane
* actually read or acted on the text; there is no read-receipt concept at this layer. "Sent"
* says exactly that; "delivered" claimed more than this call can ever establish.
*/
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(FleetMetrics.PUSH_NUDGES, "outcome", outcome);
@@ -220,6 +230,26 @@ public final class ReplyPushLoop {
.collect(Collectors.toUnmodifiableSet());
}
/**
* Test seam only (fleetd #418): carries no production behaviour, and nothing in this class
* calls it. Exposes {@link #pendingQuestionTurnIdsFor} — the exact state {@link #decide} reads
* to decide whether a question keeps a lead's schedule alive.
*
* <p>{@code MessageService.ask()} does three things in order before a question is fully open to
* this loop: it flips the ticket's {@code poll()} phase to {@code Phase.ASKING}, then resolves
* the reverse-rendezvous waiter, then calls {@link #onQuestionOpened}, which is what actually
* populates {@link #pendingQuestions}. A test that barriers on {@code Phase.ASKING} observes only
* the first of those three steps — under load the asker thread can be descheduled between steps
* one and three, so the barrier releases before this method's underlying map is populated, and
* {@link #decide} correctly reports nothing pending yet. A test that must order itself after the
* state {@link #decide} actually reads waits on this instead of on the phase.
*
* @return an unmodifiable snapshot; empty for a lead with no open questions
*/
Set<String> pendingQuestionTurnIdsForTest(String lead) {
return pendingQuestionTurnIdsFor(lead);
}
private List<PendingIncident> pendingIncidentsFor(String lead) {
return pendingIncidents.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
}
@@ -375,6 +405,80 @@ public final class ReplyPushLoop {
return Action.WAIT_BUSY;
}
/**
* Resolve who to nudge about {@code target}, the way every public entry point below wants it:
* {@link PrimaryRegistry#nudgeTargetFor}, but only after checking the delegating lead it names
* is still actually there (fleetd #368).
*
* <p><strong>The bug this closes.</strong> {@code PrimaryRegistry.forgetDelegation} is wired to
* exactly one event — a worker's release — because that is the only teardown the daemon already
* observes for a session in this map. Nothing removes a binding when the LEAD half goes away: a
* lead that is closed, crashes, or is relaunched leaves {@code leadByTarget} entries pointing at
* a terminal that no longer exists. {@code nudgeTargetFor} falls back to the single known
* primary only when the map holds nothing for {@code target} — a stale non-null entry beats the
* fallback every time, which is exactly backwards: the fallback's own javadoc argues it is safe
* precisely in the case a stale entry now hides.
*
* <p><strong>The fix.</strong> Before trusting a recorded delegation, probe the lead the same
* way {@link #decide} already does every tick ({@code agents.status}) — cheap, since it is a
* local herdr round-trip, and it is the same signal {@code AgentControl.paneByTerminal} already
* trusts to tell a genuinely dead target from a live one. A lead that fails the probe is treated
* as if it had never been recorded: the stale entry is forgotten (self-healing, exactly like
* {@code AgentControl.paneByTerminal} already does on {@code agent_not_found}) and resolution is
* retried, which now reaches the fallback {@code nudgeTargetFor} was built to reach — the same
* empty-map state its javadoc already argues is correct.
*
* <p><strong>fleetd #368 review — only a positive "gone" reading forgets the binding.</strong>
* The first version of this method treated <em>any</em> {@code RuntimeException} from the probe
* as death, which is the #359 mistake repeated: a transient socket blip or a codec error on a
* perfectly live lead would silently and permanently unbind it, with no re-record ever coming.
* That is destructive on one bad reading, exactly what #359 shipped a two-reading guard to avoid
* for the analogous lead-tab-liveness question. {@link #isLive} now matches
* {@code AgentControl.agentCall}'s own narrower rule (see its {@code agent_not_found} check): only
* that specific, affirmative "herdr has no such agent" signal counts as gone. Every other failure
* — timeout, transport error, a decode error — is treated as still live and the binding is left
* alone, because guessing wrong here is unrecoverable while guessing "live" merely costs one more
* retry on the next tick, which {@link #decide} already tolerates.
*/
private Optional<String> resolveLiveLead(String target) {
Optional<String> lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty() || isLive(lead.get())) {
return lead;
}
log.debug("push: lead {} delegated to for {} is no longer live, forgetting the stale binding "
+ "and falling back", lead.get(), target);
primaryRegistry.forgetDelegation(target);
return primaryRegistry.nudgeTargetFor(target);
}
/**
* Whether {@code lead} should still be trusted: {@code false} only when herdr affirmatively
* reports the terminal gone ({@code agent_not_found}), never on a merely inconclusive failure.
*
* <p>fleetd #368 review: an earlier version returned {@code false} for any {@code RuntimeException},
* which made a transient herdr hiccup on a live lead indistinguishable from the lead actually
* being dead — and the caller's response to {@code false} ({@code forgetDelegation}) is
* destructive and permanent. Narrowed to the one code {@code AgentControl.agentCall} itself
* already treats as a genuine, resolvable absence (see its {@code agent_not_found} handling) —
* every other {@code RuntimeException} is treated as "still live" and the binding survives to be
* probed again next time, which costs nothing worse than one more retry.
*/
private boolean isLive(String lead) {
try {
agents.status(lead);
return true;
} catch (RuntimeException e) {
boolean gone = e instanceof HerdrException he && "agent_not_found".equals(he.code());
if (gone) {
log.debug("push: lead {} no longer exists ({})", lead, e.toString());
} else {
log.debug("push: liveness check for lead {} was inconclusive ({}); treating as live "
+ "rather than risk destroying a live binding", lead, e.toString());
}
return !gone;
}
}
// --- public entrypoints ----------------------------------------------------------------------
/**
@@ -385,7 +489,7 @@ public final class ReplyPushLoop {
* backstop until a lead is recorded.
*/
public void onReplyQueued(String target) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
return;
@@ -413,7 +517,7 @@ public final class ReplyPushLoop {
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
*/
public void onTicketTerminal(String ticket, String target, boolean failed) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
ticket, target);
@@ -448,7 +552,7 @@ public final class ReplyPushLoop {
* @param question the question text
*/
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}'s question (turnId {}), skipping nudge",
target, turnId);
@@ -476,7 +580,7 @@ public final class ReplyPushLoop {
Collection<String> profiles, int remainingCoolOffSeconds) {
Map<String, List<String>> targetsByLead = new ConcurrentHashMap<>();
for (String target : targets) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend incident {} has no known lead for target {}", incidentId, target);
continue;
@@ -498,7 +602,7 @@ public final class ReplyPushLoop {
* Without an owning lead, emit a warning because no control can act on the target.
*/
public void onBackendTargetUnmapped(String target, String reason) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend target {} could not map to a credential: {}", target, reason);
return;
@@ -654,7 +758,7 @@ public final class ReplyPushLoop {
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
questionReminderCount + 1, maxReminders,
replyTargets.size(), tickets.size(), questions.size());
countNudge("delivered");
countNudge("sent");
for (PendingIncident incident : incidents) {
if (pendingIncidents.remove(incident.key(), incident)) {
deliveredIncidents.add(incident.key());
@@ -1,5 +1,7 @@
package dev.ltms.fleet.peer;
import dev.ltms.fleet.placement.PlacementDecision;
import java.nio.file.Path;
import java.util.List;
import java.util.Set;
@@ -143,9 +145,160 @@ public interface PeerLauncher {
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*
* <p>fleetd #425: for an implementation with role pools (a no-argument spawn is read as {@link
* MemberRole#DEV}, see {@link SpawnRequest}), this must be the profile a live spawn of that role
* would actually be placed on right now, not a value captured once at startup — a caller such as
* {@code fleet_profiles} relies on this to report a live, not frozen, fact.
*/
String defaultProfile();
/**
* The profile an unqualified spawn of {@code role} would resolve to right now — the role-aware,
* live counterpart of {@link #defaultProfile()} (fleetd #425).
*
* <p>A caller that must provision something profile-specific (working directory, parity overlay
* files) <em>before</em> the actual spawn — {@code SessionManager.acquireWithWorktree} is the one
* that exists today — needs the exact profile that spawn will use, for the caller's real role,
* not a role-agnostic guess. Calling {@link #defaultProfile()} for that purpose reads {@code
* MemberRole#DEV}'s answer regardless of the caller's actual role, which is wrong for any other
* role and can provision for a profile the spawn never lands on.
*
* <p>Default implementation returns {@link #defaultProfile()}, ignoring {@code role} — the right
* answer for a launcher with no role-pool concept of its own (e.g. a single {@code
* HerdrPeerLauncher} adapter, which is never reached this way in production: {@code
* CompositePeerLauncher} always fronts it and resolves roles itself).
*
* <p>fleetd #453: this default is deliberately <em>not</em> abstract — unlike {@link
* #spawn(SpawnRequest, PlacementDecision)} (fleetd #450), there is no live defect in inheriting
* it today, and the only current single-adapter implementer ({@code HerdrPeerLauncher}) is
* correct to do so. But it stays correct only as long as that holds: <strong>if a launcher ever
* routes more than one profile per role, it MUST override this method</strong>, or every role
* silently resolves to {@link #defaultProfile()} with no error and no log line. {@code
* HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)} — the override in {@code
* dev.ltms.fleet.member}, not the declaration below — names this method and {@link #place}
* explicitly as "unoverridden here" for exactly this reason. Read it before adding role-pool
* routing to any {@code HerdrPeerLauncher} subclass.
*
* <p>Who is forced to read which paragraph, because it is not symmetric. A new class that
* implements this interface directly must write a body for {@link #spawn(SpawnRequest,
* PlacementDecision)}, which is abstract here, so it lands on this javadoc. A subclass of
* {@code HerdrPeerLauncher} does not: that class already implements the method, and the
* subclass inherits the body. So for a subclass this paragraph is advice, not a gate.
*/
default String defaultProfileFor(MemberRole role) {
return defaultProfile();
}
/**
* The profile an <em>unqualified</em> spawn of {@code role} would actually be routed to right
* now — the same candidate list, the same {@code quarantined}/{@code coolingOff}/{@code
* modelOff} filtering, and the same {@code PlacementPolicy} that {@link #spawn} itself
* consults for a blank-profile request (fleetd #425 rework).
*
* <p>This is <em>not</em> {@link #defaultProfileFor}: that method answers "what is first in
* {@code role}'s pool", blind to quarantine, cool-off, and the model on/off gate — the right
* answer for a role-agnostic, best-effort report ({@code fleet_profiles}' {@code "default"}
* field), but the wrong one for a caller that needs the profile a spawn will actually land on.
* A quarantined or model-off pool-first profile makes {@link #defaultProfileFor} return a name
* an unqualified spawn will never be routed to.
*
* <p>Just the resolved name, not the full {@link PlacementDecision} — a caller that only wants
* to know the answer (a status report, a log line) can call this; a caller that will later
* <em>act</em> on the answer by spawning — provisioning a worktree for a specific profile
* before the peer exists is the one that matters — must call {@link #place} and carry the
* {@link PlacementDecision} itself through to {@link #spawn(SpawnRequest, PlacementDecision)}
* instead of calling this method and feeding the string back in as an explicit profile. Doing
* that re-enters {@link #spawn(SpawnRequest)}'s explicit-profile branch, which disagrees with
* the routing branch on purpose about what an excluded profile means: the routing branch (and
* {@link #place}) falls through a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
* profile to the next candidate, while the explicit branch refuses outright — correct for an
* operator who named that profile on purpose, wrong for a name that only ever came from placement
* itself. That accidental refusal is exactly the regression fleetd #425 rework round 2 fixes:
* the default implementation below delegates to {@link #place}, so the two can never drift apart,
* but a caller that resolves through this method alone and spawns separately can still recreate
* the round-1 defect for itself. (Before fleetd #435, this accident was also reachable through
* {@code maxLoad} specifically, because {@code FixedPlacementPolicy} — the default policy — did
* not evaluate it at all for automatic selection; #435 closed that gap, so a placement decision
* can no longer be at cap in the first place. The refusal-vs-fall-through disagreement above is
* the part that was never about {@code maxLoad} and is still real.)
*
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
* candidate in {@code role}'s pool is currently placeable
*/
default String routedProfileFor(MemberRole role) {
return place(role).profile();
}
/**
* Resolve, <em>without spawning</em>, the {@link PlacementDecision} an unqualified spawn of
* {@code role} would make right now — the same candidate list, the same {@code
* quarantined}/{@code coolingOff}/{@code modelOff} filtering, and the same {@code
* PlacementPolicy} {@link #spawn(SpawnRequest)}'s blank-profile branch itself consults (fleetd
* #425 rework).
*
* <p>Pair this with {@link #spawn(SpawnRequest, PlacementDecision)}, never with {@link
* #spawn(SpawnRequest)} fed the decision's profile as an explicit name — see {@link
* PlacementDecision}'s own javadoc for why that second form regressed.
*
* <p>Default implementation wraps {@link #defaultProfile()}, ignoring {@code role} and every
* placement condition — the right answer for a launcher with no pool or placement-policy
* concept of its own, matching {@link #defaultProfileFor}'s own default.
*
* <p>fleetd #453: same reasoning as {@link #defaultProfileFor}'s own #453 note — this default
* is deliberately not abstract (no live defect today, correct for the sole single-adapter
* implementer), but <strong>a launcher that ever routes more than one profile per role MUST
* override this method too</strong>, or placement silently ignores {@code role} for it. See
* {@code HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)}'s javadoc, which names this
* method as "unoverridden here" and why that is correct only for a single-profile adapter.
*
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
* candidate in {@code role}'s pool is currently placeable
*/
default PlacementDecision place(MemberRole role) {
return new PlacementDecision(defaultProfile());
}
/**
* Spawn against an already-resolved {@link PlacementDecision} from {@link #place}, honoring it
* completely: none of the conditions {@link #place} already applied — quarantine, cooling off,
* {@code maxLoad} (evaluated by every placement policy including the default {@code fixed},
* since fleetd #435), model-off — are re-evaluated here; {@code decision} already reflects them.
* This is not skipping a check {@code place} left undone; it is not repeating one {@code place}
* already did, and not re-opening the window between that decision and this spawn in which the
* underlying state could otherwise move. This is what lets a resolve-then-spawn caller
* ({@code SessionManager.acquireWithWorktree}, which must know the profile before it can
* provision a worktree for it) and a plain blank-profile {@link #spawn(SpawnRequest)} caller
* land on the exact same outcome for the exact same placement state (fleetd #425 rework,
* round 2).
*
* <p>{@code req}'s own {@link SpawnRequest#profileName()} is ignored in favor of {@code
* decision.profile()} — the caller is expected to have built {@code req} with a blank or
* matching profile; passing a request that names a <em>different</em>, explicit profile than
* the decision it is paired with is a caller bug this method does not attempt to detect.
*
* <p>No default implementation (fleetd #450): the two correct bodies disagree on purpose, so an
* implementer must choose one rather than silently inherit whichever this interface happened to
* provide. An implementer with no placement concept of its own — spawns a single profile, e.g.
* {@code HerdrPeerLauncher} — should delegate to {@link #spawn(SpawnRequest)} with the decision's
* profile named explicitly, since there is no separate routing path to honor there: the
* explicit-profile branch it re-enters and the routing branch {@link #place} would have used are
* the same thing. <strong>A launcher that routes across more than one profile — the way {@code
* CompositePeerLauncher} routes across every configured adapter — MUST NOT re-enter {@link
* #spawn(SpawnRequest)}.</strong> Doing so re-applies that single-argument method's
* explicit-profile checks ({@code enforceNotQuarantined}, {@code enforceNotCoolingOff}, {@code
* enforceMaxLoad}, {@code enforceModelEnabled} in {@code CompositePeerLauncher}), which can
* refuse the very profile {@link #place} just chose, if the underlying placement state moved in
* the window between the {@link #place} call and this one — the exact window this method and
* {@link PlacementDecision} exist to close (fleetd #444). Before #450 this was a {@code default}
* method that only {@code CompositePeerLauncher} overrode; a future placement-doing launcher
* could have inherited the re-entering body silently and never known. Making it abstract turns
* that silent inheritance into a compile error.
*
* @throws IllegalArgumentException if the decision names an unknown profile
*/
PeerHandle spawn(SpawnRequest req, PlacementDecision decision);
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
@@ -192,4 +345,66 @@ public interface PeerLauncher {
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
/**
* Model ids the operator has currently turned off in the central {@code models.allow:} list
* (fleetd #422) — empty for a launcher with nothing to gate against. {@code fleet_profiles}/
* {@code GET /profiles} (via {@code FleetMcp.profilesView}) call this to report which models
* are off, and MUST read this exact accessor rather than deriving their own answer: the fleetd
* #404 lesson is that a status field reading a different source than the behaviour it describes
* can drift from what the gate ({@code CompositePeerLauncher.enforceModelEnabled} and its
* candidate filter) actually enforces. A default of {@code Set.of()} keeps every other {@link
* PeerLauncher} implementer (the herdr adapters, and the two test-fake implementers) unchanged.
*
* <p>fleetd #422 follow-up: this alone cannot tell "no {@code models:} block at all" from "a
* {@code models:} block where nothing is currently off" — both report an empty set here. Delegates
* to {@link #modelGateState()} so the two facts always come from the one read {@link
* #modelGateState()}'s implementer makes; do not override this method separately from that one.
*/
default Set<String> disabledModels() {
return modelGateState().off();
}
/**
* Whether the central {@code models.allow:} gate (fleetd #422) is armed at all, together with
* which model ids are currently off — fleetd #422 follow-up. {@link #disabledModels()} alone
* cannot distinguish two states that both report an empty set: a host with no {@code models:}
* block (nothing is gated, and nothing can be) and a host WITH a {@code models:} block where
* nothing is currently turned off (the gate is armed and reporting zero). This method exists so
* a caller — the startup log, {@code fleet_profiles}/{@code GET /profiles} — can tell the two
* apart, the same reason {@code CompletionResolver.UnsetMeaning} exists: an accessor that can
* legitimately report "empty" must never let a caller guess why.
*
* <p>Default {@link ModelGateState#notConfigured()} — every launcher without a {@code models:}
* block to read from (the herdr adapters, and the two test-fake implementers), matching {@link
* #disabledModels()}'s own default of an empty set.
*/
default ModelGateState modelGateState() {
return ModelGateState.notConfigured();
}
/**
* fleetd #422 follow-up: the result of {@link #modelGateState()} — see that method's javadoc
* for why "armed" and "off" must be reported together from one read rather than as two
* separately-derived facts that a reload landing between them could make disagree.
*
* @param configured {@code true} when a {@code models:} block exists at all (armed), regardless
* of whether anything in it is currently turned off; {@code false} when there
* is no block to gate against
* @param off the model ids currently turned off; always empty when {@code configured} is
* {@code false}
*/
record ModelGateState(boolean configured, Set<String> off) {
public ModelGateState {
off = Set.copyOf(off);
}
public static ModelGateState notConfigured() {
return new ModelGateState(false, Set.of());
}
public static ModelGateState armed(Set<String> off) {
return new ModelGateState(true, off);
}
}
}
@@ -29,4 +29,9 @@ public record SpawnRequest(String profileName, String requestedCwd, String calle
String sessionName, String resumeSessionId) {
this(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId, null);
}
/** Return a copy of this request with {@code profileName} replaced by {@code profile}. */
public SpawnRequest withProfile(String profile) {
return new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, role);
}
}
@@ -3,6 +3,7 @@ package dev.ltms.fleet.placement;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
import java.util.Optional;
import java.util.OptionalLong;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
@@ -21,46 +22,158 @@ import java.util.function.LongSupplier;
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
* real sleep.
*
* <h2>Escalation (fleetd #466)</h2>
* A flat cooldown does not fit every exhaustion. A backend that reports "out of capacity for the
* rest of the hour" recovers in one cooldown; a weekly subscription limit does not — it keeps
* reporting exhausted on every attempt made before the window resets, so a flat 30-minute cooldown
* (the default {@code cooldownNanos}) means roughly 336 pointless spawn attempts across a week, one
* every cooldown.
*
* <p><strong>This class only ever sees the exhaustion signal.</strong> Its only production caller is
* {@code Fleetd.exhaustionSink}, wired to fire on a {@code BACKEND_EXHAUSTED} classification alone.
* The daemon's other outage state — a credential "cooling off" after repeated non-exhaustion
* backend errors (an HTTP 5xx storm, say) — is a separate mechanism, {@code BackendOutagePolicy},
* with its own short fixed 60s cooldown and no repeat tracking. The two are never merged: escalating
* on a cooling-off signal would turn a transient 5xx storm into a multi-hour backoff, which is
* exactly the failure this ticket is not asking for. Confirmed by reading every call site of
* {@link #quarantine} — {@code BackendOutagePolicy} has its own {@code coolOff} method and never
* calls this one.
*
* <p><strong>Mechanism</strong> — the {@link #withEscalation} constructors track, per credential, how
* many times in a row {@link #quarantine} has been called without an intervening "quiet" gap.
* Each call computes {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at
* {@code maxCooldownNanos}. A call counts as a continuation of the same streak — {@code repeatCount}
* increments — when it arrives no more than one base {@code cooldownNanos} after the previous
* quarantine's deadline (this covers both "still quarantined" and "quarantine just expired and it
* was exhausted again immediately"); otherwise the streak resets and this call is treated as a fresh
* first occurrence at the base cooldown.
*
* <p><strong>Reset, honestly stated.</strong> The ideal reset signal is "the cooldown expired and the
* next attempt succeeded" — but nothing in this codebase reports a spawn success back to this class
* (checked: {@code SessionManager} and {@code CompositePeerLauncher} never call any method here
* except {@link #quarantine}/{@link #isQuarantined}/{@link #remainingSeconds}, none of which is a
* success hook). Lacking that signal, the reset used here is a time-based proxy: a base-cooldown's
* worth of quiet — no exhaustion report for that credential — since the last quarantine ended. It is
* not proof the credential started working again, only the best available evidence without adding an
* active probe, which is out of scope by the operator's own design constraint (no automatic probing
* of a limited backend).
*
* <p><strong>Ceiling.</strong> {@code maxCooldownNanos} bounds the growth — an unbounded backoff is a
* permanent, unrecoverable-without-a-restart outage, which would be worse than the flat-rate bug this
* escalation fixes. {@link #withEscalation(LongSupplier, long)} defaults the ceiling to
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown (12x the 1800s default ≈ 6 hours), so a
* chronically exhausted credential still gets re-tried roughly every 6 hours instead of every 30
* minutes — about a dozen attempts a week instead of ~336.
*
* <p><strong>Backward compatibility.</strong> The original two-argument {@link #BackendQuarantine(
* LongSupplier, long)} constructor is unchanged in behaviour: it is exactly {@code
* withEscalation}'s mechanism with {@code backoffMultiplier = 1.0} and {@code maxCooldownNanos =
* cooldownNanos}, which collapses the formula back to the original flat {@code now + cooldownNanos}
* on every call regardless of history. Every existing call site (roughly 20 across the test suite,
* plus {@link #none()}) keeps its current shape and behaviour unchanged.
*/
public final class BackendQuarantine {
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
/** Default growth per consecutive exhaustion streak — see the class doc's Mechanism section. */
static final double DEFAULT_BACKOFF_MULTIPLIER = 2.0;
/** Default ceiling, expressed as a multiple of the base cooldown — see the class doc's Ceiling section. */
static final long DEFAULT_MAX_COOLDOWN_MULTIPLE = 12;
private final ConcurrentHashMap<String, QuarantineState> quarantines = new ConcurrentHashMap<>();
private final LongSupplier nowNanos;
private final long cooldownNanos;
private final double backoffMultiplier;
private final long maxCooldownNanos;
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
private final boolean inert;
/** How many consecutive exhaustion reports a credential is on, and when the resulting cooldown ends. */
private record QuarantineState(int repeatCount, long deadlineNanos) {
}
/**
* Flat cooldown, unchanged from before fleetd #466 — every {@link #quarantine} call blocks the
* credential for exactly {@code cooldownNanos}, regardless of how many times it was called
* before. Equivalent to {@link #withEscalation} with no growth ({@code backoffMultiplier = 1.0})
* and a ceiling equal to the base cooldown, so it degrades to the identical {@code now +
* cooldownNanos} formula every call. Kept for the existing call sites that want a fixed cooldown
* (and for tests exercising the fixed-cooldown shape in isolation); production wiring uses
* {@link #withEscalation} instead.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
* must be positive
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
this(nowNanos, cooldownNanos, false);
this(nowNanos, cooldownNanos, 1.0, cooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
/**
* Escalating cooldown (fleetd #466) — see the class doc's Mechanism/Reset/Ceiling sections.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be
* positive
* @param backoffMultiplier growth per consecutive exhaustion; must be {@code >= 1.0} ({@code 1.0}
* disables growth and is exactly the flat two-argument constructor)
* @param maxCooldownNanos ceiling on the escalated cooldown; must be {@code >= cooldownNanos}
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
long maxCooldownNanos) {
this(nowNanos, cooldownNanos, backoffMultiplier, maxCooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
long maxCooldownNanos, boolean inert) {
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
if (cooldownNanos <= 0) {
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
}
if (backoffMultiplier < 1.0) {
throw new IllegalArgumentException("backoffMultiplier must be >= 1.0: " + backoffMultiplier);
}
if (maxCooldownNanos < cooldownNanos) {
throw new IllegalArgumentException(
"maxCooldownNanos must be >= cooldownNanos: " + maxCooldownNanos + " < " + cooldownNanos);
}
this.cooldownNanos = cooldownNanos;
this.backoffMultiplier = backoffMultiplier;
this.maxCooldownNanos = maxCooldownNanos;
this.inert = inert;
}
/**
* Escalating cooldown with the fleetd #466 default shape: cooldown doubles
* ({@value #DEFAULT_BACKOFF_MULTIPLIER}x) per consecutive exhaustion streak, capped at
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown. This is what production wiring
* ({@code Fleetd.main}) uses.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be positive
*/
public static BackendQuarantine withEscalation(LongSupplier nowNanos, long cooldownNanos) {
return new BackendQuarantine(nowNanos, cooldownNanos, DEFAULT_BACKOFF_MULTIPLIER,
cooldownNanos * DEFAULT_MAX_COOLDOWN_MULTIPLE, false);
}
/**
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
*/
public static BackendQuarantine none() {
return new BackendQuarantine(() -> 0L, 1, true);
return new BackendQuarantine(() -> 0L, 1, 1.0, 1, true);
}
/**
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
* Quarantine {@code credentialId} starting now. On a flat instance (the two-argument
* constructor) this always blocks for exactly {@code cooldownNanos}, restarting the cooldown at
* full length on every call — a fresh refusal is fresh evidence the account is still exhausted,
* not a reason to let an earlier, shorter wait stand. On an escalating instance ({@link
* #withEscalation}) the cooldown grows with each call that arrives within one base cooldown of
* the previous deadline, and resets to the base cooldown once a call arrives after a longer gap
* — see the class doc.
*
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
* so recording a deadline would produce a quarantine that never expires — a credential locked out
@@ -73,7 +186,13 @@ public final class BackendQuarantine {
if (inert) {
return;
}
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
long now = nowNanos.getAsLong();
quarantines.compute(credentialId, (id, prev) -> {
int repeatCount = (prev == null || now - prev.deadlineNanos() > cooldownNanos)
? 1
: prev.repeatCount() + 1;
return new QuarantineState(repeatCount, now + escalatedCooldownNanos(repeatCount));
});
}
/** Whether {@code credentialId} is quarantined right now. */
@@ -87,6 +206,49 @@ public final class BackendQuarantine {
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
}
/**
* Remaining seconds together with which consecutive exhaustion this is (fleetd #466 scope item
* 2) — {@code repeatCount} 1 for a first occurrence, 2 for the second in a row, and so on; see
* {@link #quarantine}'s class-doc Mechanism section for exactly when a call continues a streak
* versus starts a fresh one.
*
* <p><strong>Read together, off the one {@link QuarantineState} entry {@link #quarantine} itself
* wrote</strong> — a single {@code quarantines.get(credentialId)}, never a separate lookup or a
* value re-derived from {@code remainingSeconds} (e.g. inverting {@link
* #escalatedCooldownNanos}). That inversion is not just extra work to avoid: once a streak has
* hit {@code maxCooldownNanos}, every further consecutive exhaustion reports the identical
* cooldown, so a derivation that starts from the cooldown value cannot tell the 4th repeat from
* the 9th — only the stored {@code repeatCount} can. This is the same rule {@code
* CompositePeerLauncher.modelGateState()} documents for its own gate/report pair: the report
* reads the exact accessor the behaviour reads, so it can never disagree with what actually
* happened (the fleetd #404/#422 lesson). {@code fleet_profiles}/{@code fleet_list}/{@code GET
* /profiles} all call this — never {@link #remainingSeconds} plus a second, independent count —
* for exactly that reason.
*
* @return empty when {@code credentialId} is not currently quarantined (including on {@link
* #none()}, which quarantines nothing)
*/
public Optional<Status> status(String credentialId) {
QuarantineState state = quarantines.get(credentialId);
if (state == null) {
return Optional.empty();
}
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
return remaining > 0
? Optional.of(new Status(toSecondsRoundedUp(remaining), state.repeatCount()))
: Optional.empty();
}
/**
* @param remainingSeconds seconds left on the quarantine, identical to {@link
* #remainingSeconds(String)}'s answer for the same credential at the
* same instant
* @param repeatCount 1 for a first occurrence, 2 for the second consecutive one, etc. —
* see {@link #status(String)}
*/
public record Status(long remainingSeconds, int repeatCount) {
}
/**
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
@@ -95,8 +257,8 @@ public final class BackendQuarantine {
*/
public Map<String, Long> activeRemainingSeconds() {
Map<String, Long> out = new LinkedHashMap<>();
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
long remaining = deadline - nowNanos.getAsLong();
quarantines.forEach((credentialId, state) -> {
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
if (remaining > 0) {
out.put(credentialId, toSecondsRoundedUp(remaining));
}
@@ -105,8 +267,14 @@ public final class BackendQuarantine {
}
private long remainingNanos(String credentialId) {
Long deadline = quarantinedUntilNanos.get(credentialId);
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
QuarantineState state = quarantines.get(credentialId);
return state == null ? 0L : state.deadlineNanos() - nowNanos.getAsLong();
}
/** {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at {@code maxCooldownNanos}. */
private long escalatedCooldownNanos(int repeatCount) {
double raw = cooldownNanos * Math.pow(backoffMultiplier, repeatCount - 1);
return raw >= (double) maxCooldownNanos ? maxCooldownNanos : (long) raw;
}
private static long toSecondsRoundedUp(long nanos) {
@@ -5,14 +5,12 @@ import java.util.List;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps
* ({@code maxLoad}) so that a pre-existing config behaves identically after upgrade — capacity
* gating for automatic placement is deliberately out of scope for {@code fixed}, exactly as it
* always has been. Reachability is a narrower exception (fleetd #315, below): a profile is never
* checked for reachability up front, only skipped once it has already failed in <em>this same</em>
* spawn call's retry loop — see the unreachable case below.
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. Reachability is a narrower
* exception (fleetd #315, below): a profile is never checked for reachability up front, only
* skipped once it has already failed in <em>this same</em> spawn call's retry loop — see the
* unreachable case below.
*
* <p>Four exceptions walk past the default instead of returning it unconditionally:
* <p>Six exceptions walk past the default instead of returning it unconditionally:
* <ul>
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
* a usage limit, not a transient capacity or reachability concern.
@@ -20,6 +18,24 @@ import java.util.List;
* ({@code BackendOutagePolicy}) — a separate, shorter-lived source from quarantine. When a
* profile is both quarantined and cooling off, only the quarantine reason is reported
* (exhaustion takes priority), matching {@code CompositePeerLauncher}'s explicit-spawn order.
* <li>At cap (fleetd #435): a profile whose live count has reached its {@code maxLoad}
* ({@link PlacementPolicyUtil#atCap}) — a documented, unconditional capacity limit (see
* {@code FleetConfig.Profile#maxLoad}), so {@code fixed} must gate on it exactly as {@code
* weighted}/{@code round-robin} already do via {@link PlacementPolicyUtil#available}. Before
* this fix {@code fixed} built its own {@link PlacementCandidate} for the default with {@code
* maxLoad} forced to {@code null}, so a capped default was chosen anyway on every unqualified
* spawn — the cap was advisory, not enforced, for the one placement policy every config uses
* by default. Reported only when quarantine and cooling off are both absent, matching {@code
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
* load, then model-off).
* <li>Model off (fleetd #422): a profile whose {@code model:} the operator has turned off in
* {@code models.allow:} — an operator decision, never a backend-reported outage, so it is a
* fifth, independent source from quarantine, cooling off, and at-cap (never merged with any
* of them), exactly as {@code CompositePeerLauncher.enforceModelEnabled} and {@link
* PlacementPolicyUtil#available} treat it. When a profile is model-off <em>and</em> quarantined,
* cooling off, or at cap, only the higher-priority reason is reported, matching {@code
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
* load, then model-off).
* <li>Unreachable (fleetd #315): {@code CompositePeerLauncher.spawn} retries a failed candidate
* on the next one and rebuilds the {@link PlacementContext} so {@code ctx.unreachable()}
* names every profile that already failed with {@code PeerUnreachableException} in this same
@@ -32,9 +48,10 @@ import java.util.List;
* {@code weighted}/{@code round-robin} skip it — an explicit {@code fleet_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined, cooling off, unreachable, or weight-0 never exercises
* any of these paths, so today's behaviour is unchanged — in particular, the very first selection
* of a spawn call always sees an empty {@code unreachable} set, so the first choice is untouched.
* A fleet where nothing is ever quarantined, cooling off, at cap, model-off, unreachable, or
* weight-0 never exercises any of these paths, so today's behaviour is unchanged — in particular,
* the very first selection of a spawn call always sees an empty {@code unreachable} set, so the
* first choice is untouched.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@@ -42,12 +59,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !ctx.coolingOff().contains(d)
&& !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)) {
&& !ctx.modelOff().contains(d) && !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)
&& !capExcluded(ctx, d)) {
return new PlacementCandidate(d, null, 1.0f, null);
}
for (PlacementCandidate c : ctx.candidates()) {
if (!ctx.quarantined().contains(c.profile()) && !ctx.coolingOff().contains(c.profile())
&& !ctx.unreachable().contains(c.profile()) && !c.excluded()) {
&& !ctx.modelOff().contains(c.profile()) && !ctx.unreachable().contains(c.profile())
&& !c.excluded() && !PlacementPolicyUtil.atCap(ctx, c)) {
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
}
}
@@ -56,9 +75,18 @@ final class FixedPlacementPolicy implements PlacementPolicy {
// Exhaustion quarantine takes priority: reported only when quarantine is absent, so the
// message never claims "cooling off" for a profile that is really backend-exhausted.
boolean dCoolingOff = !dQuarantined && ctx.coolingOff().contains(d);
// fleetd #435: at-cap sits between cooling off and model-off, matching
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
// then model-off) — reported only when quarantine/cooling-off are both absent.
boolean dAtCap = !dQuarantined && !dCoolingOff && capExcluded(ctx, d);
// fleetd #422: model-off is a fifth, independent source (an operator decision) — but
// quarantine/cooling-off/at-cap still take priority when more than one applies, matching
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
// then model-off).
boolean dModelOff = !dQuarantined && !dCoolingOff && !dAtCap && ctx.modelOff().contains(d);
boolean dUnreachable = ctx.unreachable().contains(d);
boolean dWeightExcluded = weightExcluded(ctx, d);
if (dQuarantined || dCoolingOff || dUnreachable || dWeightExcluded) {
if (dQuarantined || dCoolingOff || dAtCap || dModelOff || dUnreachable || dWeightExcluded) {
List<String> reasons = new ArrayList<>();
if (dQuarantined) {
reasons.add("is quarantined (backend exhausted)");
@@ -66,6 +94,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
if (dCoolingOff) {
reasons.add("is cooling off after repeated backend errors");
}
if (dAtCap) {
PlacementCandidate c = candidateFor(ctx, d);
int live = ctx.liveCount().apply(d);
reasons.add("is at maxLoad (" + live + " live >= " + c.maxLoad() + " cap)");
}
if (dModelOff) {
reasons.add("names a model the operator has turned off in models.allow");
}
if (dUnreachable) {
reasons.add("is unreachable");
}
@@ -78,18 +114,38 @@ final class FixedPlacementPolicy implements PlacementPolicy {
}
if (!ctx.candidates().isEmpty()) {
throw new PlacementException("all worker profiles are excluded from automatic "
+ "selection (quarantined, cooling off, unreachable, or weight-0)");
+ "selection (quarantined, cooling off, at cap, model-off, unreachable, or weight-0)");
}
throw new PlacementException("no worker profiles configured");
}
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
private static boolean weightExcluded(PlacementContext ctx, String profile) {
PlacementCandidate c = candidateFor(ctx, profile);
return c != null && c.excluded();
}
/**
* Whether {@code profile} has reached its {@code maxLoad} cap (fleetd #435), using the shared
* {@link PlacementPolicyUtil#atCap} definition — the same one {@code weighted}/{@code
* round-robin} already consult via {@link PlacementPolicyUtil#available}. Looked up by name,
* the same way {@link #weightExcluded} is: the default fast path above builds its own {@link
* PlacementCandidate} with {@code maxLoad} forced to {@code null} (it carries no cap of its
* own), so the candidate actually configured for {@code profile} has to be found in {@code
* ctx.candidates()} first.
*/
private static boolean capExcluded(PlacementContext ctx, String profile) {
PlacementCandidate c = candidateFor(ctx, profile);
return c != null && PlacementPolicyUtil.atCap(ctx, c);
}
/** The configured candidate named {@code profile} in {@code ctx}, or {@code null} if none. */
private static PlacementCandidate candidateFor(PlacementContext ctx, String profile) {
for (PlacementCandidate c : ctx.candidates()) {
if (c.profile().equals(profile)) {
return c.excluded();
return c;
}
}
return false;
return null;
}
}
@@ -22,11 +22,29 @@ import java.util.function.Function;
* the two apart so its refusal message says "cooling off", not "exhausted",
* when only this one is active. A profile can be in both sets at once; when it
* is, exhaustion quarantine is reported (it takes priority).
* @param modelOff profiles whose {@code model:} is currently turned off in {@code
* models.allow:} (fleetd #422) — an operator decision, not a backend-reported
* outage, so a SEPARATE, independent source from both {@code quarantined} and
* {@code coolingOff}. A profile can be in this set together with either (or
* both) of the others; {@link PlacementPolicyUtil} counts it into its own
* bucket rather than merging it into theirs, the same reason
* {@code coolingOff} is kept apart from {@code quarantined}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable,
Set<String> quarantined,
Set<String> coolingOff) {
Set<String> coolingOff,
Set<String> modelOff) {
/**
* Back-compat form before the fleetd #422 model on/off gate was added — no candidate's model
* is off. Keeps pre-#422 call sites (tests included) compiling and behaving identically.
*/
public PlacementContext(String defaultProfile, List<PlacementCandidate> candidates,
Function<String, Integer> liveCount, Set<String> unreachable,
Set<String> quarantined, Set<String> coolingOff) {
this(defaultProfile, candidates, liveCount, unreachable, quarantined, coolingOff, Set.of());
}
}
@@ -0,0 +1,49 @@
package dev.ltms.fleet.placement;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
/**
* An already-completed placement choice — the outcome of one {@link PeerLauncher#place} call,
* carried forward so a later {@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} can honor
* it directly instead of re-resolving the profile a second time (fleetd #425 rework, round 2).
*
* <p>The problem this exists to close: a caller that must know the profile <em>before</em> it can
* spawn — {@code SessionManager.acquireWithWorktree} provisions a worktree's {@code repoRoot} and
* parity overlay for a specific profile before the peer process exists — used to resolve that name
* with {@code PeerLauncher.routedProfileFor(role)} and then hand the SAME string back to {@link
* PeerLauncher#spawn(SpawnRequest)} as an EXPLICIT profile. That re-resolution is not free: naming
* a profile explicitly makes {@code CompositePeerLauncher.spawn} take its THROWING branch
* ({@code enforceNotQuarantined}/{@code enforceNotCoolingOff}/{@code enforceMaxLoad}/{@code
* enforceModelEnabled}), while an unqualified spawn's ROUTING branch never runs those checks at
* all — it instead FALLS THROUGH to the next candidate on exactly the same conditions the throwing
* branch refuses on. That disagreement is deliberate: an operator who names a profile should get a
* refusal, not a silent substitution. The bug is turning the fall-through into a refusal by
* accident — resolving a name through the routing side and then re-entering the refusing side with
* it, for a decision the routing side had already approved by walking past everything else.
* Before fleetd #435, this accident was also reachable through {@code maxLoad} specifically: the
* default {@code fixed} placement policy did not evaluate {@code maxLoad} at all for automatic
* selection, so a profile placement itself just approved could still die at {@code enforceMaxLoad}
* one call later, purely because the caller's route to the spawn passed through an explicit
* profile name instead of the routing branch — a failure a worktree-less unqualified spawn would
* never hit. fleetd #435 closed that specific gap ({@code fixed} now evaluates {@code maxLoad}
* exactly like every other placement policy), so a {@link PlacementDecision} can no longer be
* at-cap in the first place — but the refusal-vs-fall-through disagreement above was never about
* {@code maxLoad}, and resolving a name and re-entering the refusing branch with it is still wrong
* for every OTHER condition placement filters on.
*
* <p>{@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} closes that by spawning through
* the identical code path the routing branch itself uses, keyed off the SAME decision {@link
* PeerLauncher#place} returned — no re-checking of any condition placement already evaluated. A
* resolve-then-spawn caller and a blank-profile {@link PeerLauncher#spawn(SpawnRequest)} caller can
* then never disagree about which conditions apply to the same placement state, and neither one
* re-opens the window between the placement decision and the spawn in which the underlying state
* could otherwise move.
*
* @param profile the profile this decision resolved to (may be {@code null} only when no profile is
* configured at all — the same corner case {@link PeerLauncher#defaultProfile()}
* already tolerates)
*/
public record PlacementDecision(String profile) {
}
@@ -11,28 +11,40 @@ final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* True when {@code c} has reached its {@code maxLoad} cap: {@code liveCount(c.profile()) >=
* c.maxLoad()}. A {@code null} maxLoad means unlimited, so it is never at cap.
*
* <p>Extracted as the single shared definition of "at cap" (fleetd #435): before this fix it
* was computed inline in both {@link #available} and {@link #emptyException}, and {@code
* FixedPlacementPolicy} — not a caller of either — quietly kept its own {@code select} free of
* any cap check at all, so a capped default profile was chosen anyway under the default
* placement policy. Every automatic policy must call this, not re-derive it.
*/
static boolean atCap(PlacementContext ctx, PlacementCandidate c) {
Integer cap = c.maxLoad();
return cap != null && ctx.liveCount().apply(c.profile()) >= cap;
}
/**
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
* first because it is a static config choice rather than transient state), not
* known-unreachable, not quarantined (CB-578 stage B), not cooling off after repeated backend
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), and have not
* reached their maxLoad. A {@code null} maxLoad means unlimited.
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), not naming a
* model the operator has turned off (fleetd #422 — a third, independent source: an operator
* decision, never a backend-reported outage), and have not reached their maxLoad (see {@link
* #atCap}). A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (c.excluded() || ctx.unreachable().contains(c.profile())
|| ctx.quarantined().contains(c.profile())
|| ctx.coolingOff().contains(c.profile())) {
|| ctx.coolingOff().contains(c.profile())
|| ctx.modelOff().contains(c.profile())
|| atCap(ctx, c)) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
@@ -40,12 +52,13 @@ final class PlacementPolicyUtil {
/**
* Build a clear exception describing why every candidate was dropped: all weight-0, all
* quarantined, all cooling off, all at capacity, all unreachable, or a mix. Each candidate is
* counted into exactly one bucket (weight-excluded first, then quarantined, then cooling off)
* so a candidate excluded for more than one reason is never double-counted — a candidate that is
* both quarantined (CB-578 stage B, backend exhausted) and cooling off (fleetd #201 Unit 5,
* repeated backend errors) counts only as quarantined, matching {@code CompositePeerLauncher}'s
* explicit-spawn ordering: exhaustion quarantine takes priority when both are active.
* quarantined, all cooling off, all model-off, all at capacity, all unreachable, or a mix. Each
* candidate is counted into exactly one bucket (weight-excluded first, then quarantined, then
* cooling off, then model-off) so a candidate excluded for more than one reason is never
* double-counted — a candidate that is both quarantined (CB-578 stage B, backend exhausted) and
* cooling off (fleetd #201 Unit 5, repeated backend errors) counts only as quarantined, matching
* {@code CompositePeerLauncher}'s explicit-spawn ordering: exhaustion quarantine takes priority
* when more than one applies.
*/
static PlacementException emptyException(PlacementContext ctx) {
int weightExcluded = 0;
@@ -53,17 +66,19 @@ final class PlacementPolicyUtil {
int unreachable = 0;
int quarantined = 0;
int coolingOff = 0;
int modelOff = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (c.excluded()) {
weightExcluded++;
} else if (ctx.quarantined().contains(c.profile())) {
quarantined++;
} else if (ctx.coolingOff().contains(c.profile())) {
coolingOff++;
} else if (ctx.modelOff().contains(c.profile())) {
modelOff++;
} else if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
} else if (atCap(ctx, c)) {
atCap++;
}
}
@@ -83,6 +98,10 @@ final class PlacementPolicyUtil {
return new PlacementException(
"all worker profiles are cooling off after repeated backend errors");
}
if (modelOff == total) {
return new PlacementException(
"all worker profiles name a model the operator has turned off");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
@@ -92,8 +111,9 @@ final class PlacementPolicyUtil {
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + quarantined + " quarantined, "
+ coolingOff + " cooling off, "
+ modelOff + " model-off, "
+ weightExcluded + " weight-0, "
+ (total - atCap - unreachable - quarantined - coolingOff - weightExcluded)
+ (total - atCap - unreachable - quarantined - coolingOff - modelOff - weightExcluded)
+ " remaining");
}
}
@@ -0,0 +1,93 @@
package dev.ltms.fleet.power;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.util.Locale;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicBoolean;
/**
* Holds macOS idle sleep off by keeping a {@code caffeinate -i} child process alive for the life
* of the returned {@link SleepAssertion}.
*
* <p>{@code -i} asserts only against <em>idle</em> sleep — it does not stop the lid closing or an
* operator-requested sleep from taking effect. That is deliberate: this class exists to stop an
* unattended host from sleeping out from under a member's long turn, never to override the
* operator. {@code -s}/{@code -d} (which also block system/display sleep on demand) are
* intentionally not used here.
*
* <p>{@link #acquire()} never throws. It returns {@code null} — a no-op — off macOS, and again if
* starting the {@code caffeinate} child fails for any reason (binary missing, process table full,
* …); either case is logged once at INFO, not on every occurrence, so a daemon that runs for
* weeks with the tool unavailable does not fill its log.
*/
public final class CaffeinateSleepAssertionMechanism implements SleepAssertionMechanism {
private static final Logger log = LoggerFactory.getLogger(CaffeinateSleepAssertionMechanism.class);
private final AtomicBoolean loggedOnce = new AtomicBoolean(false);
/** {@code true} when running on macOS, the only platform {@code caffeinate} ships on. */
public static boolean isSupportedPlatform() {
return isSupportedPlatform(System.getProperty("os.name"));
}
/** Package-visible so a test can drive the platform check without touching a real property. */
static boolean isSupportedPlatform(String osName) {
return osName != null && osName.toLowerCase(Locale.ROOT).contains("mac");
}
@Override
public SleepAssertion acquire() {
if (!isSupportedPlatform()) {
logOnce("not running on macOS (os.name={}); the idle-sleep guard is a no-op on this platform",
System.getProperty("os.name"));
return null;
}
try {
Process process = new ProcessBuilder("caffeinate", "-i")
.redirectOutput(ProcessBuilder.Redirect.DISCARD)
.redirectError(ProcessBuilder.Redirect.DISCARD)
.start();
return new CaffeinateAssertion(process);
} catch (IOException | RuntimeException e) {
logOnce("could not start 'caffeinate -i' ({}); the host may idle-sleep while members are live",
e.toString());
return null;
}
}
private void logOnce(String format, Object arg) {
if (loggedOnce.compareAndSet(false, true)) {
log.info("idle-sleep guard: " + format, arg);
}
}
/** Wraps the live {@code caffeinate} child; {@link #close} force-destroys it, idempotently. */
private static final class CaffeinateAssertion implements SleepAssertion {
private final Process process;
CaffeinateAssertion(Process process) {
this.process = process;
}
@Override
public void close() {
if (!process.isAlive()) {
return;
}
process.destroy();
try {
if (!process.waitFor(2, TimeUnit.SECONDS)) {
process.destroyForcibly();
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
process.destroyForcibly();
}
}
}
}
@@ -0,0 +1,105 @@
package dev.ltms.fleet.power;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.function.IntSupplier;
/**
* Holds an OS-level assertion against idle sleep for exactly as long as at least one fleet
* member is live.
*
* <p><strong>Why this exists:</strong> a fleetd host was measured idle-sleeping after as little
* as one minute of inactivity (its {@code pmset -g custom} reports {@code sleep 1} on battery).
* Overnight the daemon's AMQP link to the broker dropped 13 times, and cross-checking every drop
* minute against {@code pmset -g log} found a sleep or wake event in the same minute or the one
* before, every time. The AMQP churn is only the visible symptom — the real problem is that a
* member mid-turn freezes with the host, and a long turn with nobody typing is exactly the case
* that goes idle.
*
* <p><strong>How it tracks "live":</strong> this is driven by {@code SessionManager}'s existing
* {@code onAcquire}/{@code onRelease} lifecycle hooks (added for CB-520/CB-516, previously wired
* to nothing but the reply inbox) rather than a second member count kept in parallel. Wire it as:
* <pre>{@code
* IdleSleepGuard guard = new IdleSleepGuard(mechanism, sessions::size);
* sessions.onAcquire(_ -> guard.recheck());
* sessions.onRelease(_ -> guard.recheck());
* }</pre>
* Every acquire/release event re-reads {@code SessionManager#size()} — the same registry {@code
* fleet_list}'s live/capacity numbers are themselves computed from — and only an actual 0→1 or
* 1→0 crossing touches the OS. A listener exception is already caught and logged by {@code
* SessionManager} itself (it must never let a listener failure block the acquire/release it is
* reacting to), so {@link #recheck()} does not need its own top-level try/catch to honor that.
*
* <p><strong>Failure posture:</strong> every method here is safe to call whether or not {@link
* SleepAssertionMechanism#acquire()} actually works. A mechanism that returns {@code null} (wrong
* platform, missing tool, spawn failure) simply means this guard never holds anything — it never
* throws and never blocks a spawn, a release, or shutdown.
*/
public final class IdleSleepGuard implements AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(IdleSleepGuard.class);
private final SleepAssertionMechanism mechanism;
private final IntSupplier liveCount;
private final Object lock = new Object();
private SleepAssertion held;
public IdleSleepGuard(SleepAssertionMechanism mechanism, IntSupplier liveCount) {
this.mechanism = mechanism;
this.liveCount = liveCount;
}
/**
* Re-read the live count and acquire or release the held assertion to match: nothing held and
* at least one member live ⇒ acquire; something held and no member live ⇒ release. A steady
* count (still zero, still positive) is a no-op either way, so a single spawn or release only
* ever touches the OS on the crossing, not on every call.
*/
public void recheck() {
synchronized (lock) {
int live = liveCount.getAsInt();
if (live > 0 && held == null) {
held = mechanism.acquire();
if (held != null) {
log.debug("idle-sleep guard armed: {} live member(s)", live);
}
} else if (live == 0 && held != null) {
releaseHeldLocked();
}
}
}
/** {@code true} while an assertion is actually held. Exposed for tests. */
boolean isHeld() {
synchronized (lock) {
return held != null;
}
}
/**
* Release whatever is held, if anything. Idempotent and safe to call at any time, including
* repeatedly — a daemon shutdown hook calls this unconditionally so no assertion (and no
* {@code caffeinate} child) survives the process, even if the drain that would otherwise have
* driven the live count to zero was itself interrupted or threw.
*/
@Override
public void close() {
synchronized (lock) {
if (held != null) {
releaseHeldLocked();
}
}
}
/** Caller must hold {@link #lock}. */
private void releaseHeldLocked() {
try {
held.close();
} catch (RuntimeException e) {
log.warn("idle-sleep guard: failed to release its assertion cleanly: {}", e.toString());
} finally {
held = null;
}
}
}
@@ -0,0 +1,11 @@
package dev.ltms.fleet.power;
/**
* A held OS-level assertion against idle sleep. {@link #close} must be idempotent — safe to call
* more than once — and must never throw, matching {@link IdleSleepGuard}'s "never break the
* fleet" contract.
*/
public interface SleepAssertion extends AutoCloseable {
@Override
void close();
}
@@ -0,0 +1,21 @@
package dev.ltms.fleet.power;
/**
* The OS mechanism {@link IdleSleepGuard} uses to hold and release an idle-sleep assertion. This
* is the seam a test exercises instead of the real effect (a live {@code caffeinate} child) — see
* {@code IdleSleepGuardTest}.
*
* <p>Implementations must never throw. Every failure — wrong platform, missing tool, a spawn
* error — must show up as {@link #acquire()} returning {@code null}, so a caller can treat "no
* assertion held" and "the mechanism could not be used" identically and the fleet keeps running
* either way.
*/
public interface SleepAssertionMechanism {
/**
* Acquire a fresh assertion against idle sleep, or {@code null} when this mechanism is not
* usable right now (wrong platform, the tool is missing, the child process could not start).
* Never throws.
*/
SleepAssertion acquire();
}
@@ -696,6 +696,10 @@ public final class FleetApp {
/**
* The worker's structured reply ({@code fleet_reply}) — resolves the blocking send awaiting
* on this session, or queues the reply in the inbox when no send is open (CB-307).
*
* <p>fleetd #365: the response body's {@code delivered} field used to be unconditionally
* {@code true} for either case; it now reports whether a send/ticket was actually resolved,
* with {@code outcome} naming which (see {@link MessageService.ReplyOutcome}).
*/
private void replyMessage(Context ctx) {
String id = ctx.pathParam("id");
@@ -719,13 +723,19 @@ public final class FleetApp {
// a WRONG value instead of failing loudly. The check lives in MessageService.reply so both
// this door and FleetMcp.reply inherit the same rule; this catch only translates it into the
// {error, detail} envelope this file uses everywhere else.
MessageService.ReplyOutcome outcome;
try {
messages.reply(id, content);
outcome = messages.reply(id, content);
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", e.getMessage()));
return;
}
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
// fleetd #365: "delivered": true used to be unconditional here, whether the reply resolved
// a waiting send or was merely queued in the inbox for a later drain — the same gap
// FleetMcp.reply had over MCP. `delivered` now reflects which actually happened, and
// `outcome` names the specific case (see MessageService.ReplyOutcome).
ctx.status(200).json(Map.of("sessionId", id, "delivered", outcome.delivered(),
"outcome", outcome.wireName()));
}
/**
@@ -92,6 +92,15 @@ public final class GitWorktrees implements Worktrees {
private final String configuredRoot;
/** OS group name for {@link #shareWithGroup} (fleetd #185 stage 3); {@code null} ⇒ feature off. */
private final String group;
/** Source directory of skill folders for {@link #seedSkills} (fleetd #362, {@code memberSkills:}
* in config); {@code null} ⇒ feature off, no worktree is touched beyond today's behaviour. */
private final String memberSkillsSource;
/** Extra environment merged into every {@code git} subprocess this instance runs. Always {@code
* Map.of()} from every production constructor. Test seam only (fleetd #362 review fix): lets
* {@code GitWorktreesTest} point {@code GIT_CONFIG_GLOBAL} at an isolated temp file so it can
* drive the real {@link #add} path against a controlled "operator's global git config" and
* prove the excludesFile composition below without ever touching the real machine's config. */
private final Map<String, String> gitEnv;
private final Consumer<String> afterWorktreeAdded;
/** How the initial {@code git worktree add} command runs. Package-private test seam for an
* interrupted command after Git has made worktree state. */
@@ -125,7 +134,20 @@ public final class GitWorktrees implements Worktrees {
* config); null/blank ⇒ {@link #shareWithGroup} is a no-op.
*/
public GitWorktrees(String configuredRoot, String group) {
this(configuredRoot, group, _ -> {});
this(configuredRoot, group, (String) null);
}
/**
* @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of
* the repo root.
* @param group optional OS group name (fleetd #185 stage 3, {@code worktreeGroup:}
* in config); null/blank ⇒ {@link #shareWithGroup} is a no-op.
* @param memberSkillsSource fleetd #362: optional directory of skill folders ({@code
* memberSkills:} in config) copied into every provisioned worktree's
* {@code .claude/skills/}; null/blank ⇒ {@link #seedSkills} is a no-op.
*/
public GitWorktrees(String configuredRoot, String group, String memberSkillsSource) {
this(configuredRoot, group, _ -> {}, null, null, memberSkillsSource);
}
/** Test seam for changing a real worktree between its creation and its security check. */
@@ -135,13 +157,20 @@ public final class GitWorktrees implements Worktrees {
/** Test seam combining a configurable {@code group} with {@link #afterWorktreeAdded}. */
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded) {
this(configuredRoot, group, afterWorktreeAdded, null, null);
this(configuredRoot, group, afterWorktreeAdded, null, null, null);
}
/** Test seam for changing how {@link #shareWithGroup}'s processes run. */
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner) {
this(configuredRoot, group, afterWorktreeAdded, shareGroupRunner, null);
this(configuredRoot, group, afterWorktreeAdded, shareGroupRunner, null, null);
}
/** Test seam for changing how the initial {@code git worktree add} command runs, with no
* {@code memberSkillsSource} configured. */
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner, Function<String[], String> worktreeAddRunner) {
this(configuredRoot, group, afterWorktreeAdded, shareGroupRunner, worktreeAddRunner, null);
}
/**
@@ -151,11 +180,30 @@ public final class GitWorktrees implements Worktrees {
*
* @param shareGroupRunner {@code null} ⇒ the real {@link #exec(String...)}.
* @param worktreeAddRunner {@code null} ⇒ the real {@link #exec(String...)}.
* @param memberSkillsSource {@code null}/blank ⇒ {@link #seedSkills} is a no-op.
*/
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner, Function<String[], String> worktreeAddRunner) {
Function<String[], String> shareGroupRunner, Function<String[], String> worktreeAddRunner,
String memberSkillsSource) {
this(configuredRoot, group, afterWorktreeAdded, shareGroupRunner, worktreeAddRunner,
memberSkillsSource, Map.of());
}
/**
* Full test seam, plus {@code gitEnv} (fleetd #362 review fix, verification only): extra
* environment merged into every {@code git} subprocess this instance runs, so a test can isolate
* something like {@code GIT_CONFIG_GLOBAL} from the real machine while still driving the real
* {@link #add} path end to end. Every production constructor above delegates here with {@code
* Map.of()}.
*/
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner, Function<String[], String> worktreeAddRunner,
String memberSkillsSource, Map<String, String> gitEnv) {
this.configuredRoot = configuredRoot;
this.group = (group == null || group.isBlank()) ? null : group;
this.memberSkillsSource = (memberSkillsSource == null || memberSkillsSource.isBlank())
? null : memberSkillsSource;
this.gitEnv = gitEnv == null ? Map.of() : gitEnv;
this.afterWorktreeAdded = afterWorktreeAdded == null ? _ -> {} : afterWorktreeAdded;
this.shareGroupRunner = shareGroupRunner != null ? shareGroupRunner : this::exec;
this.worktreeAddRunner = worktreeAddRunner != null ? worktreeAddRunner : this::exec;
@@ -184,6 +232,7 @@ public final class GitWorktrees implements Worktrees {
configureEnvironmentCredentialHelper(repoRoot, wt);
configureHttpsUrlRewriteForSshOrigin(repoRoot, wt);
isolateToolSurface(wt);
seedSkills(wt);
} catch (RuntimeException e) {
cleanupAfterAddFailure(repoRoot, wt, branch, e);
throw e;
@@ -574,6 +623,322 @@ public final class GitWorktrees implements Worktrees {
return true;
}
/**
* fleetd #362: copy each skill folder from the configured {@link #memberSkillsSource} directory
* into {@code <worktreePath>/.claude/skills/}, so a member spawned against ANY repo — not only
* one that already ships its own {@code .claude/skills/} — can load a bridge skill such as
* {@code implementer}. Every brief this fleet sends starts with {@code "Load the <name>
* skill."}; outside a repo carrying its own copy that line was previously a no-op.
*
* <p><b>No-op — nothing read, nothing written, nothing logged</b> — when {@link
* #memberSkillsSource} is null/blank (today's default), the same off-switch shape as
* {@link #shareWithGroup}.
*
* <p><b>Invariant 1 — a repo's own skill wins.</b> A skill folder already present at
* {@code <worktreePath>/.claude/skills/<name>} — because the just-checked-out branch commits its
* own copy — is left completely untouched: never overwritten, and never even opened.
*
* <p><b>Invariant 2 — a seeded skill can never end up in a worker's commit.</b> Every path this
* writes is untracked in the target repo (that is the whole reason it is being seeded), so
* {@code git status} would otherwise show each one as a new, addable, committable path. The
* repo-wide {@code .git/info/exclude} is NOT used for this: measured against a real linked
* worktree, that file resolves to the repository's COMMON git dir even from a worktree (the
* same file {@link dev.ltms.fleet.member.ClaudeCodeLauncher#writeIdeOverlay writeIdeOverlay}
* appends {@code CLAUDE.local.md} to), so an entry written there would hide the seeded skill
* from {@code git status} in the PRIMARY's own checkout and every sibling worktree too — not
* only this one. Instead, {@link #excludeSeededSkillsFromGitStatus} points {@code
* core.excludesFile} at a file scoped {@code --worktree} (the same {@code
* extensions.worktreeConfig} mechanism {@link #configureEnvironmentCredentialHelper} already
* relies on) that itself lives under this worktree's own private git dir ({@code
* .git/worktrees/<nonce>/}, OUTSIDE the working tree) — invisible to this worktree's {@code git
* status} and structurally impossible for this worktree to commit, with no effect on any other
* worktree or the primary checkout. Proven with a real {@code git status --porcelain} in
* {@code GitWorktreesTest}, not by reasoning.
*
* <p><b>Compose, don't replace.</b> {@code core.excludesFile} is single-valued: the first cut of
* this method pointed it at fleetd's own file with {@code --replace-all}, which SHADOWS whatever
* the operator's own (global, or repo-local) {@code core.excludesFile} was already resolving to
* inside this worktree, rather than adding to it. Measured concretely: this repo's own {@code
* .gitignore} does not ignore {@code target/} — only an operator's global excludesFile does — so
* every worker's {@code mvn clean install} would otherwise make {@code target/} appear as
* untracked, and {@link #hasUncommitted}'s deliberately-untracked-inclusive {@code git status
* --porcelain} (CB-576) would then read every such worktree as dirty forever, so it is never
* cleaned up. {@link #excludeSeededSkillsFromGitStatus} now reads whatever {@code
* core.excludesFile} resolves to BEFORE writing anything (falling back to git's own documented
* default, {@code $XDG_CONFIG_HOME/git/ignore} or {@code $HOME/.config/git/ignore}, when the key
* is unset entirely — see {@code gitignore(5)}), and writes that content into its OWN exclude
* file ahead of the seeded skill patterns, so every operator-configured pattern keeps applying
* inside the seeded worktree exactly as it did before seeding ran.
*
* <p>Instead of using worktree-scoped-config as an add-then-append (a second key does not exist
* for {@code core.excludesFile} — it takes exactly one value), an actual second exclude source
* was ruled out because git resolves only ONE {@code core.excludesFile}; concatenating the prior
* content into fleetd's own file is what "compose" reduces to for a single-valued key.
*
* <p>Proven the same way as invariant 2's own leak check: {@code
* GitWorktreesTest#seedSkillsComposesWithAnAlreadyEffectiveGlobalExcludesFile} isolates a
* synthetic "operator's global config" via {@code GIT_CONFIG_GLOBAL} (never the real machine's),
* seeds a skill, and asserts {@code git status --porcelain} is still empty for a file matching
* that global config's own ignore pattern.
*
* <p><b>Invariant 3 — best-effort.</b> A missing/unreadable {@link #memberSkillsSource}, or a
* copy/exclude failure, is logged and skipped — it must never fail the spawn, the same contract
* {@link #overlayParity} and {@link #isolateToolSurface} already hold.
*
* <p>Recorded for the worker itself the same way fleetd #134 records {@code
* fleet.neutralizedConfig}: {@code fleet.seededSkills} (one value per seeded skill folder) and
* {@code fleet.seededSkillsNote}, readable with {@code git config --worktree --get-all
* fleet.seededSkills}.
*
* <p><b>Kind-blind by construction, not by a backend check here — this used to be a real gap
* (fleetd #393).</b> This method only copies files; like {@link #isolateToolSurface} — which
* neutralizes BOTH {@code .mcp.json} and {@code opencode.json} unconditionally — it runs the
* same for every worktree regardless of which backend ultimately spawns into it, because no
* caller of {@link #add} hands this class a kind to consult. Before fleetd #393, that made the
* log line below a false claim of success for a {@code kind: opencode} member: opencode has no
* built-in discovery of {@code .claude/skills/}, unlike the Claude Code CLI, so a seeded skill
* never reached one. It now does — {@code OpenCodeLauncher#skillInstructionFiles} reads
* whatever this method copied into {@code .claude/skills/} and appends each {@code SKILL.md} to
* the generated {@code instructions[]} — but that delivery, and the log line that honestly
* claims it (kind-aware, unlike this one), happens at the launcher, once the kind is actually
* known, not here.
*/
private void seedSkills(String worktreePath) {
if (memberSkillsSource == null) {
return;
}
Path source = Path.of(memberSkillsSource).toAbsolutePath().normalize();
if (!Files.isDirectory(source)) {
log.warn("memberSkills source '{}' is not a directory — skipping skill seeding for worktree {}",
source, worktreePath);
return;
}
Path skillsRoot = Path.of(worktreePath).resolve(".claude").resolve("skills");
List<String> seeded = new ArrayList<>();
List<String> kept = new ArrayList<>();
try (var candidates = Files.list(source)) {
for (Path candidate : candidates
.filter(Files::isDirectory)
.filter(p -> !p.getFileName().toString().startsWith("."))
.sorted()
.toList()) {
String name = candidate.getFileName().toString();
Path dst = skillsRoot.resolve(name);
if (Files.exists(dst)) {
kept.add(name);
continue;
}
copySkillDirectory(candidate, dst);
seeded.add(name);
}
} catch (IOException | RuntimeException e) {
log.warn("failed to seed skills into worktree {} from memberSkills source '{}': {}",
worktreePath, source, e.getMessage());
return;
}
String detail = seeded.isEmpty() ? "" : "seeded: " + String.join(", ", seeded);
if (!kept.isEmpty()) {
detail += (detail.isEmpty() ? "" : "; ") + "kept the repo's own copy of: " + String.join(", ", kept);
}
if (detail.isEmpty()) {
detail = "no skill folders found under " + source;
}
// fleetd #393: this only claims the copy step, deliberately — it cannot know the member
// kind that will spawn into this worktree (see this method's own javadoc), so it must not
// read as "and the member will act on it." Whether that is true depends on the kind: the
// Claude Code CLI discovers .claude/skills/ on its own; OpenCodeLauncher logs its own
// "skill delivery" line, once the kind is known, naming what it could and could not turn
// into instructions[].
log.info("skill seeding: {} of {} candidate(s) from {} into {}/.claude/skills — {} "
+ "(whether the spawned member can act on this depends on its kind — see "
+ "the launcher's own log for that)",
seeded.size(), seeded.size() + kept.size(), source, worktreePath, detail);
if (seeded.isEmpty()) {
return;
}
try {
excludeSeededSkillsFromGitStatus(worktreePath, seeded);
recordSeededSkillsForWorker(worktreePath, seeded);
} catch (RuntimeException e) {
log.warn("seeded skill(s) {} into {} but could not hide them from git status: {} — "
+ "they may show as untracked; never commit them", seeded, worktreePath, e.getMessage());
}
}
/** Recursively copy a skill folder ({@code src}) into a fresh destination ({@code dst}) that
* {@link #seedSkills} has already confirmed does not exist, preserving the directory structure
* (e.g. {@code implementer/SKILL.md}, {@code implementer/references/...}). */
private static void copySkillDirectory(Path src, Path dst) {
try (var walk = Files.walk(src)) {
for (Path path : walk.sorted().toList()) {
Path target = dst.resolve(src.relativize(path).toString());
if (Files.isDirectory(path)) {
Files.createDirectories(target);
} else {
Files.createDirectories(target.getParent());
Files.copy(path, target, StandardCopyOption.COPY_ATTRIBUTES);
}
}
} catch (IOException e) {
throw new WorktreeException("cannot copy skill directory " + src + " -> " + dst + ": "
+ e.getMessage(), e);
}
}
/**
* Make every path in {@code seededSkillNames} (each a name under {@code .claude/skills/})
* invisible to {@code git status} in THIS worktree only — see the invariant-2 discussion on
* {@link #seedSkills}. Sets {@code core.excludesFile} scoped {@code --worktree} to a file
* written under this worktree's own private git dir ({@code git rev-parse
* --absolute-git-dir}), which lives outside the working tree, so the exclude file itself can
* never be committed either.
*
* <p><b>Compose, don't replace.</b> {@code core.excludesFile} is single-valued, so pointing it at
* fleetd's own file would otherwise SHADOW whatever excludesFile this worktree was already
* resolving (an operator's global config, most commonly) rather than add to it — see the
* "Compose, don't replace" discussion on {@link #seedSkills}. {@link
* #previouslyEffectiveExcludesFileContent} is read BEFORE this method's own {@code --worktree}
* write below, so it still sees whatever was effective beforehand; that content is written into
* fleetd's own exclude file ahead of the seeded skill patterns, and the worktree-scoped override
* then points at that combined file — so every pattern the operator's own configuration already
* applied keeps applying, plus the seeded skill paths.
*
* <p><b>Assumes a fresh worktree — not idempotent.</b> {@link #seedSkills} only ever calls this
* from {@link #add}, which always creates a brand-new worktree, so {@code core.excludesFile} is
* never already worktree-scoped-set to fleetd's own file when this runs. A hypothetical second
* call on the SAME worktree would read fleetd's own already-composed file back as "previously
* effective" (worktree scope now wins) and append the seeded patterns a second time — harmless
* to {@code git status} (duplicate exclude lines are a no-op), but not something to rely on. No
* guard is added for this because the path does not exist today; if a future caller ever seeds
* the same worktree twice, it will need one.
*/
private void excludeSeededSkillsFromGitStatus(String worktreePath, List<String> seededSkillNames) {
exec("git", "-C", worktreePath, "config", "extensions.worktreeConfig", "true");
String previouslyEffective = previouslyEffectiveExcludesFileContent(worktreePath);
String gitDir = exec("git", "-C", worktreePath, "rev-parse", "--absolute-git-dir").trim();
Path excludeFile = Path.of(gitDir, "fleet-seeded-skills-exclude");
StringBuilder patterns = new StringBuilder();
if (!previouslyEffective.isEmpty()) {
patterns.append(previouslyEffective);
}
for (String name : seededSkillNames) {
patterns.append("/.claude/skills/").append(name).append('/').append(System.lineSeparator());
}
try {
Files.writeString(excludeFile, patterns.toString());
} catch (IOException e) {
throw new WorktreeException("cannot write skills exclude file " + excludeFile + ": "
+ e.getMessage(), e);
}
exec("git", "-C", worktreePath, "config", "--worktree", "--replace-all", "core.excludesFile",
excludeFile.toString());
}
/**
* The content of whatever {@code core.excludesFile} resolves to for {@code worktreePath} right
* now — BEFORE {@link #excludeSeededSkillsFromGitStatus} points that key at fleetd's own file —
* so it can be carried forward instead of shadowed. {@code --type=path} makes git itself perform
* {@code ~}/{@code ~user} expansion the same way it would when actually reading the key to build
* exclude rules, rather than handing back a raw, unexpanded config string.
*
* <p>When the key is unset entirely (exit code non-zero), falls back to git's own documented
* default excludes file — {@code $XDG_CONFIG_HOME/git/ignore}, or {@code
* $HOME/.config/git/ignore} when that variable is unset — per {@code gitignore(5)}: git applies
* that file even with no {@code core.excludesFile} configured at all, so skipping it here would
* silently drop patterns an operator never had to configure to get.
*
* <p>Never throws: a missing, unreadable, or unresolvable file is treated as "nothing to carry
* forward" (empty string) — this is a best-effort read in service of {@link #seedSkills}'s own
* invariant 3, not a new way for skill seeding to fail a spawn.
*
* <p><b>Review fix, finding 2.</b> The XDG-fallback branch below does not go through {@code git}
* at all, so a first cut of it read {@code XDG_CONFIG_HOME}/{@code HOME} straight from the JVM's
* own environment ({@link System#getenv} / {@code user.home}) — unlike every other value this
* class resolves, which goes through a {@code git} subprocess and therefore already honours
* {@link #gitEnv}. That meant no test could make this branch hermetic, and on any machine
* carrying a real {@code ~/.config/git/ignore} (this repo's own dev machine does), every
* skill-seeding test silently composed with that real file — correct in production, but
* machine-dependent in the test suite, and a future broader pattern in that real file could
* silently change what a seeded worktree's {@code git status} reports depending on whose home
* directory ran the test. {@link #resolveEnv} now checks {@link #gitEnv} first for both
* variables, falling back to the JVM's real environment only when the seam does not supply
* them — production behaviour (empty {@link #gitEnv}) is unchanged, and a test can now isolate
* this branch exactly as it already isolates every {@code git} subprocess call.
*
* <p><b>Snapshot, not a reference.</b> The content below is read once, at seeding time, and
* copied into fleetd's own exclude file. If the operator edits their global excludesFile
* afterward, an already-seeded worktree keeps the old copy — acceptable for a worktree's
* expected lifetime, but worth knowing before reading a stale pattern as a bug.
*
* @return the file's content, trailing-newline-normalized, or {@code ""} when there is nothing
* to compose with.
*/
private String previouslyEffectiveExcludesFileContent(String worktreePath) {
String resolvedPath;
if (exitCode("git", "-C", worktreePath, "config", "--get", "--type=path", "core.excludesFile") == 0) {
resolvedPath = exec("git", "-C", worktreePath, "config", "--get", "--type=path",
"core.excludesFile").trim();
} else {
String xdgConfigHome = resolveEnv("XDG_CONFIG_HOME");
Path fallback = (xdgConfigHome != null && !xdgConfigHome.isBlank())
? Path.of(xdgConfigHome, "git", "ignore")
: Path.of(resolveHome(), ".config", "git", "ignore");
resolvedPath = fallback.toString();
}
if (resolvedPath.isBlank()) {
return "";
}
Path file = Path.of(resolvedPath);
if (!Files.isRegularFile(file) || !Files.isReadable(file)) {
return "";
}
try {
String content = Files.readString(file);
return content.isBlank() ? "" : content.stripTrailing() + System.lineSeparator();
} catch (IOException e) {
log.warn("could not read previously-effective excludesFile {} while seeding skills into "
+ "{}: {} — its patterns will not carry forward into the seeded worktree",
file, worktreePath, e.getMessage());
return "";
}
}
/**
* Resolve environment variable {@code name} for {@link #previouslyEffectiveExcludesFileContent}'s
* XDG fallback, checking {@link #gitEnv} FIRST so a test can isolate this the same way it
* already isolates every {@code git} subprocess this class runs, and falling back to the JVM's
* real environment only when the seam does not supply it (always the case in production, where
* {@link #gitEnv} is {@code Map.of()}).
*/
private String resolveEnv(String name) {
String fromSeam = gitEnv.get(name);
return fromSeam != null ? fromSeam : System.getenv(name);
}
/** Same as {@link #resolveEnv(String)}, for {@code HOME} — falls back to {@code user.home}
* (rather than {@code System.getenv("HOME")}) when the seam does not supply it, matching this
* class's pre-existing behaviour for every other home-directory resolution. */
private String resolveHome() {
String fromSeam = gitEnv.get("HOME");
return fromSeam != null ? fromSeam : System.getProperty("user.home");
}
/**
* The worker-readable half of fleetd #362, mirroring {@link #recordNeutralizedConfigForWorker}:
* record which skill folders were seeded where the worker itself can read it, without a
* working-tree file that would show up in {@code git status}.
*/
private void recordSeededSkillsForWorker(String worktreePath, List<String> seeded) {
exec("git", "-C", worktreePath, "config", "extensions.worktreeConfig", "true");
for (String name : seeded) {
exec("git", "-C", worktreePath, "config", "--worktree", "--add", "fleet.seededSkills", name);
}
exec("git", "-C", worktreePath, "config", "--worktree", "fleet.seededSkillsNote",
"each fleet.seededSkills value names a skill folder fleetd copied into "
+ ".claude/skills/ because this repo did not already ship it; it is excluded "
+ "from git status (core.excludesFile, worktree-scoped) and must never be committed");
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
@@ -1084,6 +1449,9 @@ public final class GitWorktrees implements Worktrees {
Process p;
try {
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
if (!gitEnv.isEmpty()) {
pb.environment().putAll(gitEnv);
}
if (extraEnv != null && !extraEnv.isEmpty()) {
pb.environment().putAll(extraEnv);
}
@@ -1119,7 +1487,11 @@ public final class GitWorktrees implements Worktrees {
private int exitCode(String... command) {
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
if (!gitEnv.isEmpty()) {
pb.environment().putAll(gitEnv);
}
p = pb.start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
@@ -10,6 +10,7 @@ import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -584,8 +585,65 @@ public final class SessionManager implements TurnListener {
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId,
MemberLifecycle.SlotReservation reservation) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// fleetd #425 rework (round 2): resolved through launcher.place(memberRole) — the same
// candidate list, quarantine/cool-off/model-off filtering, and PlacementPolicy an unqualified
// spawn of this role is actually judged against right now — never launcher.defaultProfile()
// (MemberRole.DEV only, wrong for any other role) and never launcher.defaultProfileFor()
// (the role's pool FIRST entry, blind to quarantine/cool-off/model-off: a first-round fix
// used exactly this and regressed fleetd #429's "the fleet keeps working when a model is
// turned off" guarantee — a quarantined or model-off pool-first profile made this throw
// instead of routing around it, which an unqualified spawn is supposed to do). This same
// resolved decision is reused below for repoRoot, parityOverlay, AND the spawn itself so the
// worktree is always provisioned for the profile the member actually runs on — the two could
// disagree before fleetd #425: this name picked repoRoot/overlay, but the spawn below passed
// the ORIGINAL (blank) profile through to placement, which re-resolves live and can pick a
// different profile if the pool changed between the two reads, or a genuinely different one
// under weighted/round-robin placement.
//
// Round 1 of this rework fed the resolved name back into launcher.spawn(SpawnRequest) as an
// EXPLICIT profile. That was a mistake this round corrects, and the mistake is not that the
// two branches apply different checks — they are SUPPOSED to disagree: the routing branch a
// blank spawn takes treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
// profile as a reason to fall through to the next candidate, while CompositePeerLauncher's
// THROWING branch (enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/
// enforceModelEnabled) treats naming that same profile explicitly as a reason to refuse
// outright. That is correct: an operator who names a profile should get a refusal, not a
// silent substitution onto a different backend. The mistake was turning a fall-through into
// a refusal by accident — resolving a name via the routing side and then re-entering the
// refusing side with it, for a placement the routing side had already approved by walking
// past everything else.
//
// Before fleetd #435, this accident was reachable through maxLoad specifically: the default
// `fixed` placement policy did not evaluate maxLoad at all for automatic selection, so an
// at-cap pool-first profile that placement itself would have picked for a plain unqualified
// spawn could die at enforceMaxLoad one call later, purely because this method's route to
// the spawn passed through an explicit profile name — a failure a worktree-less unqualified
// spawn never hit. fleetd #435 closed that gap (`fixed` now evaluates maxLoad exactly like
// every other placement policy), so that specific failure can no longer happen — a
// PlacementDecision this method resolves can no longer be at-cap in the first place. What
// this round's fix still buys, now that maxLoad can no longer cause the accident: it keeps
// the PlacementDecision from place() and hands it to launcher.spawn(SpawnRequest,
// PlacementDecision) for an unqualified request, which spawns through the SAME routing
// branch a blank spawn uses — no enforce* check is newly applied, and the window between the
// placement decision and the spawn (in which the pool, a config reload, or another spawn
// landing on the same profile could otherwise move the state) never reopens. An
// explicitly-named profile still goes through launcher.spawn(SpawnRequest) and its throwing
// branch, unchanged — that caller asked for one profile by name and still gets everything
// enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/enforceModelEnabled decide about
// it, refusal included.
//
// The one cost that remains, unchanged from round 1: an unqualified worktree-provisioned
// spawn does not get CompositePeerLauncher's cross-candidate retry on a live
// PeerUnreachableException raised by the backend itself at spawn time (a transport-level
// failure placement cannot see in advance) — spawn(req, decision) commits to the one profile
// place() already chose, the same way an explicit-profile spawn commits to its one name. That
// trade is deliberate: a worktree provisioned for the wrong backend (the #425 hazard) is worse
// than a spawn that fails cleanly and can be retried by the caller. Nothing else is lost:
// maxLoad, quarantine, cool-off and model-off all behave identically whether or not a
// worktree was requested — that agreement is the invariant this rework exists to hold.
boolean unqualifiedProfile = profile == null || profile.isBlank();
PlacementDecision decision = unqualifiedProfile ? launcher.place(memberRole) : new PlacementDecision(profile);
String preResolvedProfile = decision.profile();
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
@@ -608,7 +666,15 @@ public final class SessionManager implements TurnListener {
// copies more files into the worktree after add() returns, so sharing the group any earlier
// leaves those overlay files operator-owned and read-only for a different-uid member.
worktrees.shareWithGroup(repoRoot, path);
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
// fleetd #425: preResolvedProfile, not the original (possibly blank) profile — see the
// comment above where it is resolved. The overlay/repoRoot above and the spawn here must
// name the same profile. An unqualified request stays unqualified here and is honored via
// the PlacementDecision already captured above (spawn(req, decision) — the routing branch,
// no enforce* re-check); an explicitly-named profile still goes through the single-arg
// spawn(req) and its throwing branch, exactly as before this rework.
SpawnRequest spawnReq = new SpawnRequest(unqualifiedProfile ? null : preResolvedProfile,
path, callerCwd, sessionName, resumeSessionId, memberRole);
handle = unqualifiedProfile ? launcher.spawn(spawnReq, decision) : launcher.spawn(spawnReq);
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
preResolvedProfile, memberRole, branch, path, e.getMessage());
@@ -636,7 +702,7 @@ public final class SessionManager implements TurnListener {
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String resolvedProfile = resolveProfile(handle, preResolvedProfile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
// CB-619: see the no-worktree path above — bind before recording, and store the returned
@@ -0,0 +1,150 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
* deliberately opt-in — an unset one leaves usage-limit detection silently OFF for that profile,
* and nothing quarantines its credential. {@link Fleetd#reportExhaustedPatternGap} must say so at
* startup, naming every unarmed profile, and must never fire when every profile is armed. Mirrors
* {@link MemberCredentialsGapReportTest}'s pattern, capturing the real log via a
* {@link ListAppender}.
*
* <p>The 8-profile shape in {@link #theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles} is
* the live {@code fleetd.yaml} shape measured 2026-09-10 (fleetd #395's own ticket): 6 of 8
* profiles unarmed, 2 of those 6 ({@code opus}, {@code sonnet}) running on the operator's Claude
* subscription. {@code fleetd.yaml} itself is gitignored and unavailable to this test, so the
* shape is reproduced as a throwaway config in a {@code @TempDir} rather than read off disk.
*/
class ExhaustedPatternGapReportTest {
private static FleetConfig load(Path dir, String yaml) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml);
return FleetConfig.load(f);
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
}
/** Every profile name mentioned by a WARN-level log line, across every WARN this call produced. */
private static List<String> warnMessages(ListAppender<ILoggingEvent> appender) {
return appender.list.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.toList();
}
@Test
void theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles(@TempDir Path dir) throws Exception {
// Reproduces the live shape measured 2026-09-10: 8 profiles, 2 armed (sol, terra), 6
// unarmed (local, local-direct, gx, opus, sonnet, xf) — 2 of the unarmed 6 (opus, sonnet)
// are subscription: true.
FleetConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
local-direct:
baseUrl: http://gx01.gw:8000
gx:
kind: opencode
baseUrl: https://llm.ltms.dev/v1
opus:
subscription: true
model: claude-opus-5
sonnet:
subscription: true
model: claude-sonnet-5
sol:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
terra:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
xf:
baseUrl: https://llm.ltms.dev/v1
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportExhaustedPatternGap(cfg);
} finally {
detach(appender);
}
List<String> warns = warnMessages(appender);
assertFalse(warns.isEmpty(), "6 of 8 profiles are unarmed — at least one WARN must fire");
// Exactly one WARN aggregates every unarmed profile, naming all 6 and none of the 2 armed.
String aggregate = warns.stream()
.filter(m -> m.contains("no exhaustedPattern configured"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected an aggregate unarmed-profiles WARN: " + warns));
for (String unarmed : List.of("local", "local-direct", "gx", "opus", "sonnet", "xf")) {
assertTrue(aggregate.contains(unarmed), "aggregate WARN must name '" + unarmed + "': " + aggregate);
}
for (String armed : List.of("sol", "terra")) {
assertFalse(aggregate.contains(armed), "aggregate WARN must NOT name armed profile '" + armed + "': " + aggregate);
}
// A second, louder WARN calls out the subscription profiles specifically.
String subscriptionWarn = warns.stream()
.filter(m -> m.contains("metered Claude plan"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a subscription-specific WARN: " + warns));
assertTrue(subscriptionWarn.contains("opus"), subscriptionWarn);
assertTrue(subscriptionWarn.contains("sonnet"), subscriptionWarn);
assertFalse(subscriptionWarn.contains("local-direct"),
"the subscription WARN must not name a non-subscription profile: " + subscriptionWarn);
}
@Test
void everyProfileArmedProducesNoWarningAtAll(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
sol:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
terra:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
opus:
subscription: true
model: claude-opus-5
exhaustedPattern: "5-hour limit reached"
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportExhaustedPatternGap(cfg);
} finally {
detach(appender);
}
assertTrue(warnMessages(appender).isEmpty(),
"every profile is armed — a checker that warns anyway always fires: " + warnMessages(appender));
}
}
@@ -0,0 +1,212 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.fail;
/**
* fleetd #498: {@code Fleetd.awaitHerdr} used to return a bare {@code boolean}, collapsing "the
* configured wait budget genuinely ran out" and "the waiting thread was interrupted, possibly
* milliseconds in" onto the same {@code false} — and the caller's log line printed only the
* configured budget, never how long the wait actually ran. This class covers both halves of the
* fix:
* <ul>
* <li>the seam — {@link Fleetd#awaitHerdr} itself, driven with an injected clock and a stub
* {@link HerdrClient}, one test per {@link Fleetd.HerdrWaitResult};</li>
* <li>the call site — {@link Fleetd#logHerdrWaitOutcomeAndShouldReap}, the exact decision {@code
* main} calls (extracted here because {@code main} itself boots the whole daemon and cannot
* be driven from a unit test), pinning the three distinct log messages it emits.</li>
* </ul>
* Every expected message below is a plain literal, not built from {@code HERDR_WAIT_SECONDS} or
* any other production constant — a test that derives its expectation the way the code does
* cannot see a change to either (fleetd #496's identical trap).
*/
class FleetdAwaitHerdrTest {
// ---- the seam: Fleetd.awaitHerdr ----------------------------------------------------------
@Test
void answeredReturnsImmediatelyWithZeroElapsedAndNeverPolls() {
HerdrStub herdr = new HerdrStub(0); // succeeds on the very first call
LongSupplier clock = fixedClock(1_000L);
AtomicBoolean polled = new AtomicBoolean(false);
Runnable poller = () -> polled.set(true);
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.ANSWERED, outcome.result());
assertEquals(0L, outcome.elapsedNanos(), "a fixed clock must measure zero elapsed time");
assertFalse(polled.get(), "herdr answering on the first try must never poll");
}
@Test
void deadlinePassedIsMeasuredNotAssumed() {
HerdrStub herdr = new HerdrStub(-1); // never succeeds
// call order inside awaitHerdr: start, then per failed attempt: deadline-check, elapsed-calc
ScriptedClock clock = new ScriptedClock(0L, 30_500_000_000L, 30_500_000_000L);
Runnable poller = () -> fail("the deadline was already exceeded on the first attempt — must not poll");
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.DEADLINE_PASSED, outcome.result());
assertEquals(30_500_000_000L, outcome.elapsedNanos(),
"elapsed must be the MEASURED clock delta, not the configured budget");
}
@Test
void interruptedIsDistinctFromDeadlinePassedAndPreservesTheInterruptFlag() {
HerdrStub herdr = new HerdrStub(-1); // never succeeds
// start=0, deadline-check returns 500ms (well under the 30s budget) -> not deadline-passed,
// then the poller interrupts, and the elapsed-calc call returns 750ms.
ScriptedClock clock = new ScriptedClock(0L, 500_000_000L, 750_000_000L);
Runnable poller = () -> Thread.currentThread().interrupt();
try {
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.INTERRUPTED, outcome.result());
assertEquals(750_000_000L, outcome.elapsedNanos(),
"elapsed must be measured even when the wait ends via interruption, not the deadline");
assertTrue(Thread.currentThread().isInterrupted(),
"the interrupt flag the old code re-set must still be set on return");
} finally {
Thread.interrupted(); // clear it so it cannot leak into another test on this thread
}
}
// ---- the call site: Fleetd.logHerdrWaitOutcomeAndShouldReap -------------------------------
@Test
void answeredLogsNothingAndSaysReap() {
ListAppender<ILoggingEvent> events = attach();
try {
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.ANSWERED, 0L));
assertTrue(shouldReap, "only ANSWERED should tell main to reap orphan workers");
assertEquals(0, events.list.size(), "the answered path logs nothing itself");
} finally {
detach(events);
}
}
@Test
void deadlinePassedLogsConfiguredAndMeasuredElapsedTogether() {
ListAppender<ILoggingEvent> events = attach();
try {
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.DEADLINE_PASSED, 30_500_000_000L));
assertFalse(shouldReap, "a deadline-passed wait must not tell main to reap");
assertEquals(1, events.list.size());
ILoggingEvent event = events.list.getFirst();
assertEquals(Level.WARN, event.getLevel());
assertEquals("herdr did not answer within the configured wait (configured=30s "
+ "elapsed=30500ms) — starting anyway; /healthz will report degraded until it "
+ "comes up. Orphaned worker panes (if any) were NOT reaped.",
event.getFormattedMessage());
} finally {
detach(events);
}
}
@Test
void interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed() {
ListAppender<ILoggingEvent> events = attach();
try {
// 3ms: the ticket's own example of "a few milliseconds in", not the 30s budget.
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.INTERRUPTED, 3_000_000L));
assertFalse(shouldReap, "an interrupted wait must not tell main to reap");
assertEquals(1, events.list.size());
ILoggingEvent event = events.list.getFirst();
assertEquals(Level.WARN, event.getLevel());
String message = event.getFormattedMessage();
assertEquals("herdr wait was interrupted before the configured wait ran out "
+ "(configured=30s elapsed=3ms) — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
message);
assertFalse(message.contains("did not answer"),
"an interrupted wait must not be reported as if herdr failed to answer within the budget");
} finally {
detach(events);
}
}
// ---- fixtures --------------------------------------------------------------------------
/** Always returns the same value, i.e. a clock that measures zero elapsed time. */
private static LongSupplier fixedClock(long value) {
return () -> value;
}
/** Returns each value in order, then repeats the last one for any call beyond the list. */
private static final class ScriptedClock implements LongSupplier {
private final long[] values;
private int index;
ScriptedClock(long... values) {
this.values = values;
}
@Override
public long getAsLong() {
long v = values[Math.min(index, values.length - 1)];
if (index < values.length - 1) {
index++;
}
return v;
}
}
/** Fails {@code failuresBeforeSuccess} times, then succeeds forever; {@code -1} never succeeds. */
private static final class HerdrStub implements HerdrClient {
private final int failuresBeforeSuccess;
private int calls;
HerdrStub(int failuresBeforeSuccess) {
this.failuresBeforeSuccess = failuresBeforeSuccess;
}
@Override
public JsonNode call(String method, Object params) throws HerdrException {
calls++;
if (failuresBeforeSuccess < 0 || calls <= failuresBeforeSuccess) {
throw new HerdrException("herdr not up yet");
}
return null;
}
@Override
public void close() {
}
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
logger.setLevel(Level.DEBUG);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
}
}
@@ -20,6 +20,7 @@ import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementDecision;
import dev.ltms.fleet.placement.PlacementPolicies;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
@@ -123,6 +124,11 @@ class FleetdBackendErrorSinkTest {
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
}
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
}
@Override
public Set<String> profiles() {
return Set.of();
@@ -0,0 +1,65 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
* it says nothing about which one {@code main} actually calls.
*
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
* every one of them builds its own instance directly. That silent regression is exactly the shape
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
* factory choice, following their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
* constructor. It does not prove that call actually executes at startup (no test here starts the
* daemon), and it does not prove the escalation reaches a real backend or credential — only
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
* the wiring runs.
*/
class FleetdBackendQuarantineWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"Fleetd.main's BackendQuarantine local must still be built from "
+ "BackendQuarantine.withEscalation(System::nanoTime, "
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
+ "proves the escalation itself works — this source check is what must go red instead. "
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
+ "flat ~30-minute cooldown, about 336 times across the week.");
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
assertFalse(source.contains(
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
}
}
@@ -0,0 +1,155 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #416: {@code fleet_list}'s {@code CapacitySource.configuredProfiles} must enumerate the
* <em>startup</em> profile set, not the live, hot-reloaded one.
*
* <p>{@code profiles} as a whole is a {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}):
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} once at construction, so a profile
* only added to the hot-reloaded map can never actually be spawned. Before this fix, {@code
* Fleetd.main} built {@code CapacitySource} with {@code () -> config.get().profiles().keySet()} —
* the live map — so {@code fleet_list} would report a freshly hot-reloaded profile as available
* ({@code free > 0}) while {@code fleet_spawn} on that same profile failed with
* {@code unknown worker profile}. Measured on another host: adding a throwaway profile and letting
* it hot-reload gave {@code fleet_list} -> {@code free: 3} and {@code fleet_spawn} ->
* {@code error: unknown worker profile}.
*
* <p>This test needs a reload, the same reason {@link FleetdExhaustionDetectionArmedWiringTest}
* does: at startup the two snapshots agree, so a test of only a newly started daemon would not
* detect a live {@code config.get()} lookup for the set.
*
* <p><b>maxLoad must stay hot.</b> It is read off {@code config.get()} exactly like
* {@code credentialId} ({@link ConfigRef} documents both as hot, "read live off the config
* supplier ... exactly like weight/maxLoad"), so a reload that only changes an existing profile's
* {@code maxLoad} — no add/remove — must still change what {@code fleet_list} reports without a
* restart. A fix that freezes the whole {@code CapacitySource} against {@code cfg} (rather than
* only its {@code configuredProfiles} set) would trade this bug for its mirror image and is pinned
* wrong by {@link #reloadedMaxLoadStillChangesWhatFleetListReports}.
*/
class FleetdCapacitySourceWiringTest {
private static final String STARTUP = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 3
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_NEW_PROFILE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 3
ghost404:
baseUrl: http://gx00.gw:8001
model: ghost404
maxLoad: 3
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_CHANGED_MAX_LOAD = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 9
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
@DisplayName("a profile present only in the live (hot-reloaded) config is NOT listed")
void liveOnlyProfileIsNotListed(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
Files.writeString(file, WITH_NEW_PROFILE);
assertTrue(config.reload().applied());
// The live snapshot now has the new profile — proves the reload really happened and this
// test is not accidentally passing because nothing changed.
assertTrue(config.get().profiles().containsKey("ghost404"));
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
assertFalse(source.configuredProfiles().get().contains("ghost404"),
"a profile added only to the hot-reloaded config must not be listed by fleet_list — "
+ "HerdrPeerLauncher never learns about it until a restart, so fleet_spawn on it "
+ "would fail with 'unknown worker profile' while fleet_list claimed it free");
}
@Test
@DisplayName("a profile present in the startup set IS listed")
void startupProfileIsListed(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
// fleetd #416, both-directions requirement: a test that only ever passes an empty/absent
// startup set (the case above) cannot tell a correct lookup from one that is permanently
// empty (e.g. a mutation replacing the supplier with Set::of). This is the direction that
// fails if the fix regresses to reporting nothing at all.
assertTrue(source.configuredProfiles().get().contains("terra"),
"a profile present in the startup snapshot must still be listed by fleet_list");
}
@Test
@DisplayName("a hot maxLoad edit still changes what fleet_list reports")
void reloadedMaxLoadStillChangesWhatFleetListReports(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
assertEquals(3, source.maxLoad().apply("terra"),
"sanity: maxLoad reads 3 from the startup config before any reload");
Files.writeString(file, WITH_CHANGED_MAX_LOAD);
assertTrue(config.reload().applied());
assertEquals(9, source.maxLoad().apply("terra"),
"maxLoad must stay hot — the SAME CapacitySource instance must reflect a reloaded "
+ "maxLoad without a restart, exactly like credentialId. Freezing the whole "
+ "CapacitySource against the startup snapshot (rather than only its "
+ "configuredProfiles set) would trade fleetd #416 for its mirror image.");
}
}
@@ -0,0 +1,106 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #474: proves the exact wiring {@code Fleetd.main} uses to construct its live {@code
* ConfigRef} — {@code new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools)}
* — actually makes {@link ConfigRef#reload()} refuse a charter that names an MCP tool the server
* does not register, the same way {@code Fleetd.main} itself refuses one at startup (see {@code
* FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool}).
*
* <p>{@code dev.ltms.fleet.config.ConfigRefTest} pins the same behaviour through a locally-built
* {@code Consumer<FleetConfig>} adapter that calls the same production {@code CharterToolSurface}
* method, because that test lives in {@code dev.ltms.fleet.config} and cannot see {@code
* Fleetd#assertChartersNameOnlyRegisteredTools} (package-private to {@code dev.ltms.fleet}). This
* class is the companion proof that lives where the real method reference is visible, so the literal
* expression {@code Fleetd::assertChartersNameOnlyRegisteredTools} — not just an equivalent — is
* what gets exercised. {@code Fleetd.main} itself cannot be driven this far in a unit test: every
* fixture in {@code FleetdStartupValidationTest} is deliberately invalid so {@code main} throws
* before opening a socket, binding Javalin, or doing anything else with a real side effect, so a
* test cannot get {@code main} far enough to hold a running daemon it could then reload — this test
* builds the {@code ConfigRef} the same way {@code main} does and drives {@link ConfigRef#reload()}
* directly instead, the same shape {@code FleetdExhaustionDetectionArmedWiringTest} and its
* siblings already use for the rest of {@code Fleetd.main}'s wiring.
*/
class FleetdConfigRefCharterToolSurfaceWiringTest {
private static final String BASE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
void reloadRefusesACharterNamingAnUnregisteredToolThroughFleetdsOwnWiring(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
""");
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
Fleetd::assertChartersNameOnlyRegisteredTools);
FleetConfig before = config.get();
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
""");
ConfigRef.Outcome out = config.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.error().contains("dev"), out.error());
assertTrue(out.error().contains("bridge_send"), out.error());
assertSame(before, config.get());
}
@Test
void reloadAcceptsACharterNamingOnlyRegisteredToolsThroughFleetdsOwnWiring(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, BASE + """
fleet:
charters:
dev: old charter
""");
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
Fleetd::assertChartersNameOnlyRegisteredTools);
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
""");
ConfigRef.Outcome out = config.reload();
assertTrue(out.applied());
assertEquals("config reloaded", out.summary());
}
}
@@ -0,0 +1,80 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #474 follow-up: {@code Fleetd.main} builds its live {@code ConfigRef} from the
* three-argument constructor, {@code new ConfigRef(configPath, cfg,
* Fleetd::assertChartersNameOnlyRegisteredTools)}, so a reload runs the same charter tool-surface
* gate startup does (see {@link dev.ltms.fleet.config.ConfigRef}'s class doc, "fleetd #474" bullet).
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
* three-argument constructor and the {@code Fleetd.assertChartersNameOnlyRegisteredTools} adapter
* work correctly together — both build their OWN {@code ConfigRef} with that constructor, so neither
* proves {@code main} still CHOOSES the three-argument form over the plain two-argument {@code new
* ConfigRef(configPath, cfg)}.
*
* <p>Measured directly: reverting {@code Fleetd.java}'s {@code config} local to the two-argument
* constructor compiles with 0 errors and leaves the entire 1633-test suite green — including every
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
* instance directly, wired with the check by hand. That silent regression is exactly the shape
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
* their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* ConfigRef} and never runs {@code main} — a green result here proves only that the exact text
* {@code main} contains is the three-argument construction with {@code
* Fleetd::assertChartersNameOnlyRegisteredTools}. It does not prove that call actually executes at
* startup (no test here starts the daemon), and it does not prove the reload gate itself works —
* only {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
* behaviour; only a live daemon proves the wiring runs.
*/
class FleetdConfigRefWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main still builds config from the three-argument ConfigRef constructor")
void mainStillWiresTheThreeArgumentConfigRefConstructor() throws Exception {
String source = fleetdSource();
// A broken read (wrong working directory, wrong path, a file that came back empty) would
// make the assertFalse below pass vacuously — a "clean" negative check that actually
// checked nothing. Guard against that first, with an anchor that has nothing to do with
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
assertTrue(source.contains("public final class Fleetd"),
"fleetdSource() did not read anything usable — src/main/java/dev/ltms/fleet/Fleetd.java "
+ "did not come back containing its own class declaration. The assertFalse below "
+ "would pass vacuously on a broken read; fix the read before trusting this test.");
assertTrue(source.contains(
"ConfigRef config = new ConfigRef(configPath, cfg, "
+ "Fleetd::assertChartersNameOnlyRegisteredTools);"),
"Fleetd.main's config local must still be built from the three-argument ConfigRef "
+ "constructor, with Fleetd::assertChartersNameOnlyRegisteredTools as "
+ "extraValidation. Reverting to the plain two-argument constructor (fleetd #474's "
+ "measured M2 regression) compiles with 0 errors and leaves the whole suite green — "
+ "including ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest, because "
+ "neither builds its ConfigRef through main — this source check is what must go "
+ "red instead. A reverted daemon would accept, through a reload with no restart, "
+ "exactly the charter that #469/#474 already refuse at startup.");
// Negative form of the same check: the pre-#474 two-argument call, if it ever reappears at
// this declaration, must not be mistaken for the three-argument one by a looser
// positive-only check — this is the M2 mutation this test exists to kill.
assertFalse(source.contains("ConfigRef config = new ConfigRef(configPath, cfg);"),
"main's config local must never regress to the plain two-argument ConfigRef "
+ "constructor — that drops the reload-path charter check (fleetd #474's measured "
+ "M2 mutation) with no other test catching it");
}
}
@@ -0,0 +1,137 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.placement.BackendQuarantine;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #446: {@code exhaustionDetectionArmed} must describe the LIVE config, not a startup
* snapshot — the opposite of what fleetd #404 asked for. #404 pinned {@code exhaustedPattern} as
* compiled once at startup (matching {@code CompletionResolver}'s then-frozen behaviour); #446's
* whole point is that both sides — the classification {@code CompletionResolver} enforces and the
* {@code exhaustionDetectionArmed} field this reports — now read the SAME live {@link
* LiveExhaustedPatterns} instance, so a reload arms or disarms detection with no restart, and the
* report can never disagree with the behaviour (the fleetd #404 rule, now upheld for real instead
* of by freezing both sides).
*
* <p>This test needs a reload. At startup the two snapshots agree, so a test of only a newly
* started daemon would not detect a live {@code config.get()} lookup in the report field.
*
* <p><b>The hotness proof criterion 1 demands:</b> run this against the fixed code (green) and
* against the pre-#446 compile site — {@code Fleetd.main}'s {@code exhaustedPatternsByProfile}
* compiled once into a {@code Map<String, Pattern>} at startup, handed to {@code quarantineSource}
* — and show it fails there. See the class doc history above: that shape is exactly what
* {@code reloadedPatternArmsDetectionWithNoRestart} below is written to catch, and it is the test
* that used to assert the opposite (see git history for this file's pre-#446 version, which
* asserted {@code assertFalse(...)} on the identical scenario this now asserts {@code assertTrue}
* on).
*/
class FleetdExhaustionDetectionArmedWiringTest {
private static final String NO_PATTERN = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_PATTERN = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
exhaustedPattern: "usage limit"
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
@DisplayName("reloading a profile's exhaustedPattern IN arms exhaustionDetectionArmed, no restart")
void reloadedPatternArmsDetectionWithNoRestart(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, NO_PATTERN);
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
FleetMcp.QuarantineSource before = Fleetd.quarantineSource(config, BackendQuarantine.none(),
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
assertFalse(before.exhaustedPatternArmed().apply("terra"),
"no exhaustedPattern configured yet — must report unarmed");
Files.writeString(file, WITH_PATTERN);
assertTrue(config.reload().applied());
assertTrue(config.get().profiles().get("terra").hasExhaustedPattern());
// Re-read exhaustedPatternArmed off the SAME QuarantineSource built BEFORE the reload — no
// new object, no new wiring — to prove the field itself is live, not merely that a freshly
// built source would be. This is the exact axis the pre-#446 code fails on: the same
// Function<String,Boolean> lambda, called again after a reload, sees the new answer only if
// it re-reads config.get() on every call rather than a value captured earlier.
assertTrue(before.exhaustedPatternArmed().apply("terra"),
"exhaustionDetectionArmed must read the LIVE config on every call — fleetd #446 made "
+ "this hot; it must reflect a reload with no restart and no new QuarantineSource");
}
@Test
@DisplayName("reloading exhaustedPattern OUT disarms exhaustionDetectionArmed, no restart")
void reloadedPatternRemovalDisarmsDetectionWithNoRestart(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, WITH_PATTERN);
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
assertTrue(source.exhaustedPatternArmed().apply("terra"),
"exhaustedPattern is configured from the start — must report armed");
Files.writeString(file, NO_PATTERN);
assertTrue(config.reload().applied());
assertFalse(source.exhaustedPatternArmed().apply("terra"),
"removing exhaustedPattern and reloading must disarm detection with no restart — an "
+ "operator turning detection off (e.g. while debugging a false positive) "
+ "must see that reflected immediately, exactly like arming it is");
}
@Test
@DisplayName("a profile with no exhaustedPattern configured is reported unarmed")
void aProfileWithNoPatternIsUnarmed(@TempDir Path dir) throws Exception {
// fleetd #404's second direction, carried forward: this test alone fails if
// exhaustedPatternArmed degenerates to a constant `profile -> true` — it needs at least one
// profile that is genuinely unarmed to catch that.
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, WITH_PATTERN);
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
assertTrue(source.exhaustedPatternArmed().apply("terra"),
"a profile whose exhaustedPattern is configured must report armed");
assertFalse(source.exhaustedPatternArmed().apply("sonnet"),
"a profile absent from config entirely must not report armed");
}
}
@@ -0,0 +1,175 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.session.SessionManager;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #446 follow-up (round 3): round 2's {@link FleetdUsageLimitFixWarningTest} calls {@link
* Fleetd#usageLimitFixWarning}/{@link Fleetd#usageLimitFixWarningNoModel} directly, which pins
* the two methods' TEXT but cannot see whether {@link Fleetd#exhaustionSink}'s {@code log.warn}
* call actually invokes either one, or invokes the right one. A mutation battery run against
* round 2's merge (1592 green tests) proved both gaps real:
*
* <ul>
* <li><b>Cell A</b> — replacing the whole ternary result with a literal string
* ({@code "usage-limit fix: MUTANT"}) left every test green;</li>
* <li><b>Cell B</b> — swapping the ternary's two branches, so a profile WITH a {@code model:}
* gets the no-model fallback message and vice versa, was measured unmeasured by the lead
* but the existing test file shows no call that would notice it either.</li>
* </ul>
*
* <p>This class drives {@link Fleetd#exhaustionSink} — the actual factory {@code main} wires,
* since round 3's extraction pulled it out of the inline lambda for exactly this reason — and
* asserts on the REAL text a {@link ListAppender} attached to {@code Fleetd}'s own logger
* captures, mirroring {@code ExhaustedPatternGapReportTest}'s idiom. Chosen over a source-text
* assertion (the {@code FleetMcpAuthzTest} {@code everyRegisteredToolHasItsHandlerActionPinned}
* idiom) because that route can prove Cell A (the ternary is still there, still built from the
* two named methods) but not Cell B (which branch a given profile actually reaches) — this route
* answers both from one mechanism, since it reads what the sink actually logged for each shape of
* profile.
*
* <p>Each test filters for the WARN line starting {@code "usage-limit fix:"} specifically — {@link
* Fleetd#exhaustionSink} also logs a separate {@code "credential '...' quarantined for..."} WARN
* on every call, and asserting against the wrong one would pass or fail for the wrong reason. That
* prefix survives even under the M5 mutation above (the mutant keeps the tag, only the body
* becomes {@code "MUTANT"}), so the filter itself is not what either mutation defeats.
*/
class FleetdExhaustionSinkWarningTest {
private static final String YAML = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: claude-opus-5
gx:
baseUrl: http://gx00.gw:8000
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static SessionManager emptyRosterSessions() {
FakeHerdr h = new FakeHerdr();
FleetConfig.Profile dummy = new FleetConfig.Profile(
"dummy", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(dummy.profile(), dummy), dummy.profile(), _ -> "tok");
// Never acquires a session — exhaustionSink's roster lookup is expected to miss and fall
// back to the profileHint argument, exactly like OpenCodeLauncher's real call site does
// (fleetd #234) — so the sink under test never needs a populated roster.
return new SessionManager(launcher);
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
}
/** The one WARN line {@code exhaustionSink} builds from the ternary under test, or {@code null}. */
private static String fixWarning(ListAppender<ILoggingEvent> appender) {
List<String> warns = appender.list.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.startsWith("usage-limit fix:"))
.toList();
assertTrue(warns.size() == 1,
"expected exactly one 'usage-limit fix:' WARN per onExhausted call, got " + warns.size()
+ ": " + warns);
return warns.getFirst();
}
@Test
@DisplayName("a profile WITH a configured model gets the actionable models.allow fix, not the fallback")
void profileWithModelGetsTheActionableFix(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = ConfigRef.fixed(cfg);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
Map<String, String> reasonByCredential = new HashMap<>();
ExhaustionSink sink = Fleetd.exhaustionSink(emptyRosterSessions(), config, quarantine,
reasonByCredential, cfg);
ListAppender<ILoggingEvent> appender = attach();
try {
sink.onExhausted("term_x", "The usage limit has been reached", "terra");
} finally {
detach(appender);
}
String warning = fixWarning(appender);
assertTrue(warning.contains("terra"), "must name the profile: " + warning);
assertTrue(warning.contains("claude-opus-5"), "must name the model: " + warning);
assertTrue(warning.contains("enabled: false"), "must name the action: " + warning);
assertTrue(warning.contains("models.allow"), "must name where the action goes: " + warning);
assertFalse(warning.contains("has no model: configured"),
"a profile WITH a model must not get the no-model fallback text: " + warning);
}
@Test
@DisplayName("a profile with NO configured model gets the fallback, never a fix that does not exist")
void profileWithNoModelGetsTheFallbackNotAFix(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = ConfigRef.fixed(cfg);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
Map<String, String> reasonByCredential = new HashMap<>();
ExhaustionSink sink = Fleetd.exhaustionSink(emptyRosterSessions(), config, quarantine,
reasonByCredential, cfg);
ListAppender<ILoggingEvent> appender = attach();
try {
sink.onExhausted("term_y", "The usage limit has been reached", "gx");
} finally {
detach(appender);
}
String warning = fixWarning(appender);
assertTrue(warning.contains("gx"), "must name the profile: " + warning);
assertTrue(warning.contains("has no model: configured"),
"must say there is no model to gate by name: " + warning);
assertFalse(warning.contains("enabled: false"),
"a profile with no model: configured has no models.allow entry to flip — must "
+ "not claim one exists: " + warning);
}
}
@@ -80,7 +80,7 @@ class FleetdLeadMailboxSelectionTest {
@Test
void opensTheMailboxWhenAUriAndSelfIdAreConfigured() {
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null);
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null, null);
Fleetd.openLeadMailbox(coordinator, Map.of(), opener);
@@ -94,7 +94,7 @@ class FleetdLeadMailboxSelectionTest {
void honoursUriEnvOverALiteralUri() {
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator("amqp://stale:stale@old:5672/x", "COORD_URI",
"mac-opus", 8);
"mac-opus", 8, null);
Fleetd.openLeadMailbox(coordinator, Map.of("COORD_URI", RESOLVED_URI), opener);
@@ -105,7 +105,7 @@ class FleetdLeadMailboxSelectionTest {
@Test
void turnsOffWhenUriEnvDoesNotResolve() {
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(null, "COORD_URI", "mac-opus", null);
var coordinator = new FleetConfig.Coordinator(null, "COORD_URI", "mac-opus", null, null);
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
@@ -116,7 +116,7 @@ class FleetdLeadMailboxSelectionTest {
void warnsAndStaysOffWhenSelfIdIsMissing() {
var appender = captureFleetdLogs();
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, null, null);
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, null, null, null);
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
@@ -131,7 +131,7 @@ class FleetdLeadMailboxSelectionTest {
var appender = captureFleetdLogs();
var opener = new RecordingOpener();
opener.unreachable = true;
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null);
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null, null);
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener),
"a down coordination broker turns the feature off; it must never take the daemon down");
@@ -0,0 +1,89 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #480 Unit A, hard requirement 6: pin {@code Fleetd.main}'s construction of {@link
* dev.ltms.fleet.lead.LeadRollover} with a source-text assertion, mirroring {@code
* FleetdCompletionResolverWiringTest}'s pattern — five log-only reporters in {@code Fleetd.main}
* already survived mutation batteries this exact way (fleetd #415's extraction antidote note).
*
* <p>What this class still covers, and what it never claimed to. {@code LeadRolloverTest}
* constructs its own {@code LeadRollover} directly (as every prior test of an extracted factory
* does) with a hand-built lookup, so a mutation that deletes the {@code leadRollover(...)} call
* from {@code main} — or replaces one of its arguments with something that still compiles, e.g.
* {@code router.leadAgents()} swapped for {@code null}, or the whole assignment swapped for a bare
* {@code null} literal — leaves every behavioural test green. This is a plain string read, guarded
* by an unrelated anchor assertion so a broken or empty file read cannot pass as a real change.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* LeadRollover} and never runs {@code main}. It pins the {@code leadRollover(...)} CALL SITE's
* argument list — that {@code main} still passes {@code leads} at all — never what the factory
* DOES with that argument once inside its own body.
*
* <p><b>Correction (fleetd #480 relative-handover-path follow-up): that gap used to be real, and
* now is not — but not here.</b> This class's javadoc previously claimed "no behavioural test can
* catch this wiring dropping out" for the whole factory, including the lambda {@code
* leadRollover(...)} builds internally (terminal → lead name → {@code Leader.cwd()}). That claim
* was proven true at the time — mutating that lambda's body to {@code String leadName = null;}
* (always "no lead found", which silently reintroduces the daemon-cwd bug this ticket fixes) left
* the full suite green, {@code Tests run: 1669, Failures: 0}. It is no longer true: {@code
* FleetdLeadRolloverWorkspaceLookupTest} now calls {@code Fleetd.leadRollover(...)} directly with a
* real {@link dev.ltms.fleet.config.ConfigRef} built from a temp {@code fleetd.yaml}, and fails
* against that exact one-line mutation. So: THIS class still covers only the call site's argument
* list; {@code FleetdLeadRolloverWorkspaceLookupTest} is what now covers the lambda's body. Neither
* one subsumes the other — keep both.
*/
class FleetdLeadRolloverWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] unrelated anchor: Fleetd.java still declares the Fleetd class")
void unrelatedAnchorStillPresent() throws Exception {
// Guards the two assertions below: without this, a bad read (empty string, wrong file,
// truncated file) could vacuously fail to contain the leadRollover(...) call too, and a
// test that only asserts "contains X" would report a false pass for the wrong reason if X
// happened to match. Asserting an unrelated, structurally distant string first proves the
// read actually pulled real file content.
String source = fleetdSource();
assertTrue(source.contains("public final class Fleetd"),
"sanity anchor failed — the file read did not return real Fleetd.java source; the "
+ "leadRollover(...) wiring assertions below cannot be trusted until this passes");
}
@Test
@DisplayName("[SOURCE TEXT] main still constructs LeadRollover via the leadRollover(...) factory, exactly as heartbeat is constructed")
void mainStillCallsTheLeadRolloverFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);"),
"Fleetd.main must still assign `LeadRollover leadRollover = leadRollover(cfg, "
+ "router.leadAgents(), config, leads);`. Dropping this call, or swapping one of "
+ "its arguments for something that still compiles (e.g. null in place of "
+ "router.leadAgents()), leaves every behavioural test green — this source check is "
+ "what must go red instead. fleetd #480 correction 2 deliberately dropped "
+ "primaryRegistry from this call — see LeadRollover's class javadoc for why a "
+ "single-slot lookup was wrong here. The fleetd #480 relative-handover-path "
+ "follow-up added `leads` (terminal → lead name) so the factory can resolve a "
+ "relative handoverPath against the calling lead's own workspace.");
}
@Test
@DisplayName("[SOURCE TEXT] the leadRollover(...) factory itself gates construction on cfg.leadRollover() != null")
void factoryGatesOnConfigPresence() throws Exception {
String source = fleetdSource();
assertTrue(source.contains("if (cfg.leadRollover() == null) {"),
"Fleetd.leadRollover(...) must refuse to construct a LeadRollover when the "
+ "leadRollover: block is absent — an upgraded daemon must never silently acquire "
+ "the ability to clear the lead's own pane. See LeadHeartbeatLoop's construction "
+ "gate (cfg.leadHeartbeat() != null) for the pattern this mirrors.");
}
}
@@ -0,0 +1,163 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.lead.LeadRollover;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashMap;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
/**
* fleetd #480 relative-handover-path follow-up, correction round: a BEHAVIOURAL test of {@link
* Fleetd#leadRollover}'s own body — the terminal → lead-name → {@code Leader.cwd()} lookup it
* builds — not another source-text pin.
*
* <p>{@code FleetdLeadRolloverWiringTest} (a plain string read) still earns its keep: it pins the
* call site's ARGUMENT LIST, so a mutation that drops {@code leads} back out of the call, or
* swaps it for {@code Map::of}, still goes red there. But nothing before this class exercised the
* LAMBDA BODY {@code leadRollover(...)} builds — the terminal-to-workspace {@code
* Function<String,String>} that {@link LeadRollover#open} actually calls. Every prior test either
* exercised {@code LeadRollover} directly with a hand-built lookup ({@code LeadRolloverTest}), or
* read source text without ever calling the factory ({@code FleetdLeadRolloverWiringTest}) — so a
* mutation that breaks the lookup ITSELF (e.g. always resolving no lead, or reading a snapshot
* instead of the live supplier) left every existing test green while the daemon's real wiring
* silently reintroduced the exact bug this whole ticket fixes: a relative {@code handoverPath}
* resolving against the daemon's own {@code cwd} instead of the calling lead's.
*
* <p>{@code Fleetd.leadRollover(...)} is package-private and {@code static}, so this test — living
* in the same {@code dev.ltms.fleet} package — calls it directly, exactly the way {@code
* FleetdExhaustionDetectionArmedWiringTest} and {@code FleetdCapacitySourceWiringTest} already call
* other package-private startup factories with a real {@link ConfigRef} built from a temp {@code
* fleetd.yaml} (never the gitignored live one).
*
* <p><b>Proved against the mutation it exists to catch.</b> Before this class was added, mutating
* {@code leadRollover(...)}'s lambda body to {@code String leadName = null;} (always "no lead
* found", which forces every relative {@code handoverPath} onto the {@code user.dir} fallback —
* i.e. the original bug) left the full suite green: {@code Tests run: 1669, Failures: 0}. With
* {@link #relativeHandoverPathResolvesAgainstTheLeadsConfiguredCwd()} added, the same one-line
* mutation now fails that test (it asserts the resolved path equals the configured lead's {@code
* cwd}, which the mutant can never produce) — see this ticket's fleet_reply history for both runs.
*/
class FleetdLeadRolloverWorkspaceLookupTest {
private static AgentControl fakeAgents() {
return new AgentControl(new FakeHerdr());
}
@Test
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
+ "the CALLING lead's configured cwd, not the daemon's own working directory")
void relativeHandoverPathResolvesAgainstTheLeadsConfiguredCwd(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
leadRollover:
handoverPath: handover.md
""".formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
() -> Map.of("term_opus", "opus"));
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
+ "must construct an object");
LeadRollover.PendingRollover pending = rollover.open("term_opus", "test");
String expected = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expected, pending.handoverPath(),
"a relative handoverPath must resolve against the CALLING lead's configured cwd "
+ "through the REAL Fleetd.leadRollover(...) wiring — not the daemon's own "
+ "working directory. This is the exact axis that was proven uncovered: "
+ "mutating the factory's lambda body to always report \"no lead found\" "
+ "left every prior test green.");
}
@Test
@DisplayName("[BEHAVIOURAL] a terminal not present in the live lead-terminal map falls back "
+ "to the daemon's own user.dir")
void terminalNotInLiveMapFallsBackToUserDir(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
leadRollover:
handoverPath: handover.md
""");
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
// has not yet scanned, or one with no fleet.leaders entry at all.
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of);
assertNotNull(rollover);
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
String expected = Path.of(System.getProperty("user.dir")).resolve("handover.md")
.normalize().toString();
assertEquals(expected, pending.handoverPath(),
"a lead not present in the live terminal→name map must fall back to the daemon's "
+ "own user.dir — the same fallback LeadLauncher#launch already uses for a "
+ "lead with no configured cwd. This is deliberate, pinned behaviour, not an "
+ "accident of the null-check chain.");
}
@Test
@DisplayName("[BEHAVIOURAL] the terminal→lead-name lookup is read LIVE on every open() call, "
+ "never snapshotted at Fleetd.leadRollover(...) construction time")
void workspaceLookupIsReadLiveNotSnapshotted(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
leadRollover:
handoverPath: handover.md
""".formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
// EMPTY at the moment leadRollover(...) is called. A lookup captured (snapshotted) here
// instead of read live through the supplier on every call would never see the entry added
// below — exactly the natural mistake to make, since leads are discovered by a live tab
// scan that runs AFTER this factory is constructed at startup.
Map<String, String> liveLeadTerminals = new HashMap<>();
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
() -> liveLeadTerminals);
assertNotNull(rollover);
// The lead is "discovered" only now — mutate the SAME backing map the supplier reads from.
liveLeadTerminals.put("term_opus", "opus");
LeadRollover.PendingRollover pending = rollover.open("term_opus", "test");
String expected = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expected, pending.handoverPath(),
"the terminal→lead-name lookup must be read LIVE on every open() call — a lead "
+ "discovered by the tab scan AFTER Fleetd.leadRollover(...) was constructed "
+ "must still resolve correctly, not only one that was already live at "
+ "construction time");
}
}
@@ -0,0 +1,62 @@
package dev.ltms.fleet;
import dev.ltms.fleet.peer.PeerLauncher;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
/**
* fleetd #422 follow-up: {@code Fleetd.modelGateCoverageLine} is the startup-log counterpart of
* {@code exhaustedPatternCoverageLine}/{@code errorPatternCoverageLine} — see {@code
* FleetdPatternCoverageLineTest} for the identical shape this follows — except here there is no
* {@code UnsetMeaning} choice for a caller to get backwards: {@link
* PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether an empty
* {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate with at all"
* or "a block armed and currently reporting zero off". This class proves {@code
* modelGateCoverageLine} words those two states — plus the third, N off — distinctly, so a
* mutation that made it ignore {@code configured()} either way is caught here.
*/
class FleetdModelGateCoverageLineTest {
@Test
@DisplayName("no models: block reports not configured, distinct from armed-with-zero")
void noModelsBlockReportsNotConfigured() {
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
assertEquals("not configured (no models: block — nothing is gated, and nothing can be)", line);
}
@Test
@DisplayName("a models: block armed with nothing off reports armed, distinct from not configured")
void armedWithNothingOffReportsArmed() {
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
assertEquals("armed (models: block present; 0 models currently turned off)", line);
}
@Test
@DisplayName("a models: block with N off names the off models")
void armedWithModelsOffNamesThem() {
String line = Fleetd.modelGateCoverageLine(
PeerLauncher.ModelGateState.armed(Set.of("deepseek-v4-flash", "claude-opus-9000")));
assertEquals("armed (2 model(s) turned off: [claude-opus-9000, deepseek-v4-flash])", line);
}
@Test
@DisplayName("the three states produce pairwise-distinct wording for the same empty-looking input")
void theThreeStatesProduceDistinctWording() {
String notConfigured = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
String armedZero = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
String armedOne = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of("x")));
// Pinned individually above; restated here so this test alone still catches a regression
// even if one of the three tests above were ever deleted — the exact FleetdPatternCoverageLineTest
// pattern, adapted from "two keys" to "three states of one gate".
assertNotEquals(notConfigured, armedZero,
"collapsing 'no models: block' into 'armed, zero off' is fleetd #422 follow-up's exact defect");
assertNotEquals(armedZero, armedOne);
assertNotEquals(notConfigured, armedOne);
}
}
@@ -0,0 +1,67 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
/**
* fleetd #415 (review follow-up): {@code CompletionResolverTest} proves {@code coverage()} words
* {@code UnsetMeaning.OFF} and {@code UnsetMeaning.BUILT_IN_DEFAULT} correctly — but every one of
* those tests supplies the meaning itself. That proves the enum's wording, never that {@code
* Fleetd} pairs the right meaning with the right pattern key. That pairing is #415's actual
* defect: {@code coverage()} had no way to know what unset meant for its key, so the fix moved
* the fact to the caller — and nothing yet proved the caller states it correctly.
*
* <p><b>Measured or it didn't happen:</b> swapping the two {@code UnsetMeaning} arguments at
* {@code Fleetd}'s two coverage call sites — giving {@code exhaustedPattern} the built-in-default
* wording and {@code errorPattern} the off wording, #415's exact defect with the keys exchanged —
* compiled with 0 errors and left all 1506 existing tests green. This class exists to turn that
* swap red.
*
* <p>It calls {@link Fleetd#exhaustedPatternCoverageLine} and {@link Fleetd#errorPatternCoverageLine}
* directly rather than reading {@code Fleetd.java} as source text (the shape {@code
* FleetdCompletionResolverWiringTest} uses for a different wiring gap): those two methods are the
* extracted call sites {@code main} actually invokes, following the same {@code static} factory +
* dedicated-test pattern as {@link Fleetd#capacitySource} and {@link Fleetd#worktreeBranchLookup}.
*/
class FleetdPatternCoverageLineTest {
private static final Set<String> PROFILES = Set.of("terra", "gx10");
@Test
@DisplayName("exhaustedPatternCoverageLine says off when no profile configures exhaustedPattern")
void exhaustedPatternCoverageLineSaysOffWhenNoProfileConfiguresIt() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of()));
}
@Test
@DisplayName("errorPatternCoverageLine says built-in default when no profile configures errorPattern")
void errorPatternCoverageLineSaysBuiltInDefaultWhenNoProfileConfiguresIt() {
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])",
Fleetd.errorPatternCoverageLine(PROFILES, Set.of()));
}
@Test
@DisplayName("the two keys produce different wording for the identical empty-coverage input")
void theTwoKeysProduceDifferentWordingForTheSameEmptyInput() {
String exhaustedLine = Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of());
String errorLine = Fleetd.errorPatternCoverageLine(PROFILES, Set.of());
// Pinned individually above; restated here so this test alone still catches a swap even
// if one of the two tests above were ever deleted.
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
exhaustedLine);
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])", errorLine);
assertNotEquals(exhaustedLine, errorLine,
"swapping which UnsetMeaning pairs with which pattern key at Fleetd's call sites "
+ "must be caught here — that pairing, not coverage()'s own wording in isolation, "
+ "is fleetd #415's actual defect");
}
}
@@ -0,0 +1,79 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Proves {@link Fleetd#main(String[])} calls every startup report before validation aborts startup.
* The invalid non-loopback bind makes {@link FleetConfig#validateAll()} throw before {@code main}
* can open the herdr socket or bind a port. The fixture also triggers every report, so removing any
* one call from {@code main} leaves its expected log line absent.
*/
class FleetdStartupReportTest {
private static Level originalLevel;
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
originalLevel = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
logger.detachAppender(appender);
logger.setLevel(originalLevel);
}
private static boolean contains(ListAppender<ILoggingEvent> appender, String fragment) {
return appender.list.stream()
.map(ILoggingEvent::getFormattedMessage)
.anyMatch(message -> message.contains(fragment));
}
@Test
void mainReportsEveryStartupGapBeforeValidationAborts(@TempDir Path dir) throws Exception {
Path config = dir.resolve("fleetd.yaml");
Files.writeString(config, """
bind:
host: 0.0.0.0
port: 8765
profiles:
worker:
baseUrl: https://llm.ltms.dev/v1
gitTokenEnv: GITEA_TOKEN
""");
ListAppender<ILoggingEvent> appender = attach();
try {
assertThrows(IllegalStateException.class, () -> Fleetd.main(new String[]{config.toString()}));
} finally {
detach(appender);
}
assertTrue(contains(appender, "startup git host GITEA_HOST:"),
"Fleetd.main must report the git host shape");
assertTrue(contains(appender, "member trust model: members run as the same OS user"),
"Fleetd.main must report the member trust model");
assertTrue(contains(appender, "memberCredentials: absent or empty"),
"Fleetd.main must report an absent memberCredentials policy");
assertTrue(contains(appender, "exhaustedPattern: profile(s) [worker] have no exhaustedPattern configured"),
"Fleetd.main must report profiles without exhaustedPattern");
}
}
@@ -0,0 +1,154 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Proves the one thing no test proved before this ticket's follow-up: that {@code Fleetd.main}
* ITSELF — not a copy of its logic, not the validator called directly — still refuses to start on
* a bad config. Mutation testing found that deleting {@code cfg.validateAll();} (née six
* individual {@code cfg.validateXxx();} calls) from {@code Fleetd.main} left the full 1478-test
* suite green; every existing test called a validator directly and none exercised {@code
* Fleetd.main} as the caller. See {@code FleetConfigValidateAllTest} for why the fix collapses
* those six calls into one reflective {@link FleetConfig#validateAll()}, and {@code
* ConfigRefTest} for the equivalent proof on the {@link
* dev.ltms.fleet.config.ConfigRef#reload()} path.
*
* <p>This is deliberately the real {@code static void main(String[] args)} — package-private, so
* only a test in this package can call it, which is exactly what makes this proof strong: it is
* not a helper extracted for testability, it is the literal method {@code java -jar fleetd.jar}
* invokes. Every fixture below is otherwise-valid and fails exactly one validator, and — because
* {@link FleetConfig#validateAll()} runs immediately after {@code SubscriptionGuard.
* assertPrimaryClean}, before {@code main} opens the herdr socket, binds Javalin, or touches
* anything else with a real side effect — calling {@code Fleetd.main} with one of these configs is
* safe: it is guaranteed to throw before reaching any of that, precisely because the config is
* deliberately invalid.
*/
class FleetdStartupValidationTest {
private static void assertMainRefuses(Path dir, String fileName, String yaml,
String mustContain) throws Exception {
Path f = dir.resolve(fileName);
Files.writeString(f, yaml);
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> Fleetd.main(new String[]{f.toString()}),
fileName + ": Fleetd.main must refuse this config before doing anything else");
assertTrue(e.getMessage().contains(mustContain),
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
+ e.getMessage());
}
@Test
void mainRefusesANonLoopbackBindWithoutTokenMode(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "auth-exposure.yaml", """
bind:
host: 0.0.0.0
port: 8765
""", "auth.mode: token");
}
@Test
void mainRefusesALeadTabPrefixCollision(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "lead-tab-prefixes.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
tabLabel: "lead: {role} {profile}"
leaders:
opus:
tab: "lead: opus"
""", "fleet.tabLabel");
}
@Test
void mainRefusesASubscriptionProfileThatReseatsAnthropicBaseUrl(@TempDir Path dir)
throws Exception {
assertMainRefuses(dir, "subscription-profiles.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
""", "ANTHROPIC_BASE_URL");
}
@Test
void mainRefusesAnUnknownCharterKey(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "charters.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
architetc: text
""", "architetc");
}
/**
* fleetd #469: closes the gap left by #464's {@code CharterToolSurfaceTest}, which wrote its
* own charter into a {@code @TempDir} fixture and so could never fail on anything anyone wrote
* into the live {@code fleetd.yaml}. {@code bridge_send} is CB-634's own motivating example — a
* tool name the rename removed — and the charter key ({@code dev}) is a real role wire name, so
* this fixture passes {@code cfg.validateAll()}'s charter check (key valid, text non-blank) and
* is refused only by the new {@code CharterToolSurface} call right after it. The failure message
* must name both the charter key and the unknown tool.
*/
@Test
void mainRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "charter-tool-surface.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
""", "bridge_send");
}
@Test
void mainRefusesAnArchitectSlotNamingAnUnconfiguredProfile(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "members.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
baseUrl: http://gx10.gw:8000
fleet:
architects:
lead-designer:
profile: sonnet
""", "lead-designer");
}
@Test
void mainRefusesAProfileNamingAModelOutsideTheAllowList(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "models.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
rogue:
baseUrl: http://gx11.gw:8000
model: claude-opus-9000
models:
allow:
- model: claude-sonnet-5
""", "rogue");
}
}
@@ -0,0 +1,75 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #446 follow-up (criterion 2): {@code exhaustionSink} used to build the "name the fix"
* WARNING inline with two SLF4J {@code {}}-placeholder {@code log.warn} calls. Nothing pinned
* that text — a mutation battery run against the merged PR #457 proved it: renaming the leading
* {@code "usage-limit fix:"} tag to {@code "MUTANT no fix named:"} at both call sites left all
* 1586 existing tests green. This class is what makes that mutation red.
*
* <p>{@link Fleetd#usageLimitFixWarning} and {@link Fleetd#usageLimitFixWarningNoModel} are the
* extracted call sites {@code exhaustionSink} actually invokes — {@code log.warn} is called with
* each method's return value as a single, already-formatted argument, so the string these tests
* assert on is byte-identical to what {@code fleetd.out} receives. Same extracted-static-method +
* dedicated-test idiom as {@link Fleetd#exhaustedPatternCoverageLine} /
* {@link FleetdPatternCoverageLineTest}, chosen over the {@link ch.qos.logback.core.read.ListAppender}
* idiom {@code ExhaustedPatternGapReportTest} uses — this warning is built once, at one call site,
* from a single method with no branching inside the message itself, so there is no real log
* plumbing left to prove; capturing an appender would only add setup/teardown around the same
* assertion.
*
* <p>Each assertion below checks for a specific substring — the leading {@code "usage-limit fix:"}
* tag an operator would grep the log for, the profile name, the model name, and the literal
* action ({@code enabled: false} under {@code models.allow}) — rather than the whole sentence.
* Criterion 2 promised those facts, not a particular wording of the prose that carries them; a
* whole-sentence {@code assertEquals} breaks on every copy-edit of that surrounding prose and
* teaches the next person to delete the test rather than read why it failed. The leading tag is
* pinned on purpose, unlike the rest of the sentence: it is the identifying label the fleetd #446
* follow-up's own mutation battery renamed ({@code "usage-limit fix:"} → {@code "MUTANT no fix
* named:"}) to prove this class did not yet catch a tag rename — only the three-fact substrings
* above would have missed it, since none of them mention the tag itself.
*/
class FleetdUsageLimitFixWarningTest {
@Test
@DisplayName("usageLimitFixWarning names the profile, the model, and the models.allow fix")
void usageLimitFixWarningNamesProfileModelAndFix() {
String warning = Fleetd.usageLimitFixWarning("terra", "claude-opus-5");
assertTrue(warning.startsWith("usage-limit fix:"),
"must carry the grep-able identifying tag: " + warning);
assertTrue(warning.contains("terra"), "must name the profile: " + warning);
assertTrue(warning.contains("claude-opus-5"), "must name the model: " + warning);
assertTrue(warning.contains("enabled: false"), "must name the action: " + warning);
assertTrue(warning.contains("models.allow"), "must name where the action goes: " + warning);
}
@Test
@DisplayName("usageLimitFixWarning does not cross-name a different profile or model")
void usageLimitFixWarningDoesNotCrossName() {
String warning = Fleetd.usageLimitFixWarning("terra", "claude-opus-5");
// Guards against the mutation of swapping the two profile/model arguments at the call
// site — a test that only checks "some name appears" cannot catch that.
assertTrue(!warning.contains("sonnet"), "must not name an unrelated model: " + warning);
}
@Test
@DisplayName("usageLimitFixWarningNoModel names the profile but no fix, since none exists")
void usageLimitFixWarningNoModelNamesProfileNotAFix() {
String warning = Fleetd.usageLimitFixWarningNoModel("gx", 1800);
assertTrue(warning.startsWith("usage-limit fix:"),
"must carry the grep-able identifying tag: " + warning);
assertTrue(warning.contains("gx"), "must name the profile: " + warning);
assertTrue(warning.contains("1800"), "must name the quarantine duration it falls back to: " + warning);
assertTrue(!warning.contains("enabled: false"),
"a profile with no model: configured has no models.allow entry to flip — the "
+ "fallback message must not claim one exists: " + warning);
}
}
@@ -0,0 +1,241 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #377: the startup git host line must report the <em>shape</em> of the value a
* member receives as GITEA_HOST — set or unset, length, scheme, trailing slash — and the
* value itself must never reach the log.
*/
class GitHostShapeReportTest {
/** A profile that opts in to the git-forge token, so GITEA_HOST is the var that matters. */
private static final String GIT_TOKEN_CONFIG = """
profiles:
local:
baseUrl: http://gx00.gw:8000
gitTokenEnv: WORKER_GITEA_TOKEN
""";
private static FleetConfig load(Path dir, String yaml) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml);
return FleetConfig.load(f);
}
/**
* The level this logger had before {@link #attach()} raised it, so {@link #detach} can put it
* back. {@code null} is a real value here — it means "inherit from the parent" — and that is
* exactly the state this logger starts in, so it must be restored as {@code null} rather than
* as some concrete level.
*/
private static Level originalLevel;
/**
* fleetd #377: the shape lines are logged at INFO, and {@code logback-test.xml} sets
* {@code dev.ltms.fleet} to WARN — so INFO events are dropped by the level check BEFORE any
* appender sees them. Attaching an appender is therefore not enough: without raising the level
* the list stays empty and every assertion below fails against correct production code. The
* sibling report tests do the same thing at each call site (see
* {@code MemberTrustModelReportTest}); doing it here keeps it in one place.
*/
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
originalLevel = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
logger.detachAppender(appender);
logger.setLevel(originalLevel);
}
private static List<String> messages(ListAppender<ILoggingEvent> appender) {
return appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList();
}
@Test
void setLineReportsLengthSchemeAndTrailingSlash(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
String value = "https://git.example.test/";
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("GITEA_HOST", value));
} finally {
detach(appender);
}
String expected = "startup git host GITEA_HOST: set (profile 'local' gitHostEnv) — "
+ "length=" + value.length() + ", startsWithScheme=true, trailingSlash=true";
assertTrue(messages(appender).contains(expected),
"a value with a scheme and a trailing slash must be reported by shape only: "
+ "its length, startsWithScheme=true, trailingSlash=true");
}
@Test
void bareHostLineReportsNoSchemeNoTrailingSlash(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
String value = "git.example.test";
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("GITEA_HOST", value));
} finally {
detach(appender);
}
String expected = "startup git host GITEA_HOST: set (profile 'local' gitHostEnv) — "
+ "length=" + value.length() + ", startsWithScheme=false, trailingSlash=false";
assertTrue(messages(appender).contains(expected),
"a bare host name must report startsWithScheme=false, trailingSlash=false");
}
@Test
void unsetVariableIsLoggedAtInfoAndDoesNotThrow(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
ListAppender<ILoggingEvent> appender = attach();
try {
// GITEA_HOST absent from the map entirely — the unset case must not throw.
Fleetd.reportGitHostShape(cfg, Map.of());
} finally {
detach(appender);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getLevel() == Level.INFO
&& e.getFormattedMessage()
.startsWith("startup git host GITEA_HOST: unset (profile 'local' gitHostEnv)")),
"an unset git host is useful information, not an error — say it at INFO level");
}
/**
* The important test: the value must NEVER appear in the log output. This test fails if the
* line is ever changed to include the value, because it drives the real logging path with a
* value that carries a marker no shape field could contain.
*/
@Test
void theValueNeverAppearsInLogOutput(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
String marker = "never-logged-host-shape-377";
String value = "https://" + marker + "/";
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("GITEA_HOST", value));
} finally {
detach(appender);
}
List<String> msgs = messages(appender);
assertFalse(msgs.stream().anyMatch(m -> m.contains(value)),
"the full GITEA_HOST value must never reach the log");
assertFalse(msgs.stream().anyMatch(m -> m.contains(marker)),
"no fragment of the value may reach the log — a line that embeds the value"
+ " would leak at least this marker");
assertTrue(msgs.stream().anyMatch(m -> m.contains("startsWithScheme=true")
&& m.contains("trailingSlash=true")),
"the shape must still be reported although the value is not");
}
@Test
void unsetReportAlsoCarriesNoValue(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
// A blank value is not injected either (putIfPresent skips it), so it must report unset
// without echoing anything of it.
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("GITEA_HOST", " "));
} finally {
detach(appender);
}
assertTrue(messages(appender).stream().anyMatch(m ->
m.startsWith("startup git host GITEA_HOST: unset")),
"a blank value is skipped by the launcher, so the line reports unset");
}
@Test
void defaultGitHostEnvIsGITEA_HOST(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, GIT_TOKEN_CONFIG);
assertEquals(Map.of("GITEA_HOST", List.of("profile 'local' gitHostEnv")),
Fleetd.gitHostEnvVars(cfg));
}
@Test
void explicitGitHostEnvReportsUnderItsOwnName(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
gitTokenEnv: WORKER_GITEA_TOKEN
gitHostEnv: MY_FORGE_HOST
""");
String value = "https://forge.example.test/";
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of("MY_FORGE_HOST", value));
} finally {
detach(appender);
}
assertTrue(messages(appender).contains(
"startup git host MY_FORGE_HOST: set (profile 'local' gitHostEnv) — "
+ "length=" + value.length()
+ ", startsWithScheme=true, trailingSlash=true"),
"an explicitly named gitHostEnv is reported under that name");
}
@Test
void noGitTokenMeansNoHostLineIsNeeded(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportGitHostShape(cfg, Map.of());
} finally {
detach(appender);
}
assertEquals(List.of("startup git host: no profile sets a gitTokenEnv — nothing to check"),
messages(appender));
}
@Test
void aHostWithAPortIsNotTreatedAsAScheme() {
assertFalse(Fleetd.startsWithScheme("git.example.test"));
assertFalse(Fleetd.startsWithScheme("git.example.test:3000"));
assertFalse(Fleetd.startsWithScheme("git.example.test:3000/"));
assertTrue(Fleetd.startsWithScheme("https://git.example.test"));
assertTrue(Fleetd.startsWithScheme("https://git.example.test:3000/"));
assertTrue(Fleetd.startsWithScheme("ssh://git.example.test"));
}
}
@@ -113,6 +113,20 @@ class AuthzTest {
assertTrue(Authz.permits(WORKER_A, METRICS, null));
}
/**
* fleetd #421 — unlike READ (the case above), COORD_READ (a lead's own held peer mail) is the
* primary's alone. A worker or an architect reading it would disclose lead-to-lead
* coordination bodies, not the secret-free roster READ is open about.
*/
@Test
void coordReadIsThePrimarysAloneNotAWidenedRead() {
assertTrue(Authz.permits(PRIMARY, COORD_READ, null));
assertFalse(Authz.permits(WORKER_A, COORD_READ, null),
"a worker must not read held lead-to-lead mail");
assertFalse(Authz.permits(ARCH_DESIGN, COORD_READ, null),
"an architect holds READ today, but not-primary must mean not-architect here too");
}
@Test
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
@@ -186,6 +186,43 @@ class CallerResolverTest {
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "BEARER s3cret").role());
}
// ── fleetd #505: a herdr error DURING THE SCAN must not be conflated with "not a worker" ──────
// #317 (above) covers a failed lsof lookup. This is the other input to the same decision: the
// lsof lookup succeeds (a real pid), but PaneLocator's own pane scan hits a herdr error on the
// pane that owns that pid — so c.resolved() is true and c.terminal() is null, exactly like a
// real primary. c.scanComplete() is what tells them apart.
/**
* The discriminating case named in the ticket: the error must land on the pane that DOES own
* the caller's pid, or the test proves nothing (any other pane's failure is invisible to the
* scan's outcome, since a match found elsewhere is definitive regardless).
*/
@Test
void aHerdrErrorOnTheOwningPaneDuringTheScanIsRefusedNotPromotedToPrimary() {
FakeHerdr failing = new FakeHerdr().processInfoFailsForPane("w2:p7", "transient");
ConnectionIdentity incomplete = new ConnectionIdentity(new PaneLocator(failing), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(incomplete).resolve("127.0.0.1", 55555, null);
assertEquals(Role.ANONYMOUS, p.role(),
"an incomplete pane scan must never be read as a clean negative and promoted to primary");
}
/**
* The companion invariant: a herdr error on a DIFFERENT, non-owning pane must not turn every
* mid-scan teardown into a refusal — the real match is still found and resolves as a worker.
*/
@Test
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealWorker() {
FakeHerdr vanishedElsewhere = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(vanishedElsewhere), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(id).resolve("127.0.0.1", 55555, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
@Test
void aNonLoopbackCallerIsNeverThePrimaryUnderLoopbackTrust() {
// Defence in depth: startup already refuses this pairing (validateAuthExposure), but if a
@@ -0,0 +1,360 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.peer.MemberRole;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* fleetd #424 — revoking (or granting) an architect slot must take effect on the next spawn with
* no restart. A session already bound to a slot keeps its <em>binding</em> (the {@code
* terminalToSlot} occupancy) even after that slot drops out of config, but NOT the ARCHITECT
* <em>privilege</em> the slot used to grant — that is revoked on the bound session's very next
* request. See {@link MemberRegistry}'s class doc for the exact rule: "config governs what a
* bound slot still grants, as well as what may be bound next."
*
* <p>Every test here drives a REAL {@link ConfigRef#reload()} against a {@code @TempDir} file and
* asserts {@link ConfigRef.Outcome#applied()}, rather than comparing two frozen
* {@code MemberRegistry} instances in memory — the defect this ticket fixes is specifically that
* {@link MemberRegistry} used to ignore a live reload, so a test that never reloads cannot tell the
* fixed registry from the broken one. {@code requireSlotFor} and {@code reserve} are pinned in
* separate tests, in both directions (removed and added), so a registry that simply refuses (or
* simply allows) everything cannot pass by accident — see {@link MemberRegistryTest} for the
* registry's other invariants (bind/unbind cardinality, thread-safety), which are unaffected by
* this ticket and still exercised against the frozen constructor.
*/
class MemberRegistryLiveTest {
private static String yaml(String fleetBlock) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
opus:
baseUrl: http://gx00.gw:8001
model: opus
guard:
offSubscriptionHosts:
- gx00.gw
""" + fleetBlock;
}
private static final String WITH_SONNET_SLOT = """
fleet:
architects:
designer:
profile: sonnet
""";
/** No architect pool at all — developers is unrelated dead data for this registry (#424 out of scope). */
private static final String WITHOUT_ARCHITECT_SLOTS = """
fleet:
developers:
dev1:
profile: sonnet
""";
/** Same slot name ({@code designer}) as {@link #WITH_SONNET_SLOT}, repointed to a different profile. */
private static final String WITH_OPUS_SLOT = """
fleet:
architects:
designer:
profile: opus
""";
/** Same profile ({@code sonnet}) as {@link #WITH_SONNET_SLOT}, but the pool key is renamed. */
private static final String WITH_RENAMED_SLOT = """
fleet:
architects:
architect-lead:
profile: sonnet
""";
private static ConfigRef refFor(Path f) {
return new ConfigRef(f, FleetConfig.load(f));
}
// ── requireSlotFor is live (criteria 1, 2, 4) ──────────────────────────────────────────────
@Test
void requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"the slot is configured before the reload");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertThrows(IllegalArgumentException.class,
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"revoking the slot must refuse the NEXT spawn that names it");
}
@Test
void requireSlotForAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertThrows(IllegalArgumentException.class,
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"a slot added by reload must be usable with no restart");
}
// ── reserve is live too — tested separately from requireSlotFor (criteria 1, 2, 4) ────────
@Test
void reserveRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation before = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertEquals("architect:designer", before.slot());
registry.release(before); // free it back up so the reload-side reserve below starts clean
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"revoking the slot must refuse the NEXT reservation for it");
}
@Test
void reserveAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
MemberLifecycle.SlotReservation after = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertEquals("architect:designer", after.slot(),
"a slot added by reload must be reservable with no restart");
}
// ── profileForSlot, isSlot and nameForSlot are live too (fleetd #431) ─────────────────────────
// #424 pinned roleForSlot against a live reload but left these three untested — proved by
// mutating each to read a snapshot flattened once at construction: the full suite stayed green
// for all three.
//
// The three differ in how much production behaviour depends on them, and the ticket first got
// this ranking wrong. nameForSlot is wired: CallerResolver passes members::nameForSlot, next to
// members::roleForSlot. isSlot is reached through bind, which calls it to refuse an unknown
// slot. profileForSlot has NO caller in src/main at all — grepped both the ".profileForSlot("
// and the "::profileForSlot" form — so there is no seam to drive its test through beyond the
// accessor itself, and its own javadoc ("what the spawn lifecycle reads") describes a caller
// that does not exist. These tests pin the accessors as they are; whether profileForSlot should
// be wired or deleted is a separate question.
@Test
void profileForSlotReflectsAProfileChangedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertEquals("sonnet", registry.profileForSlot("architect:designer"),
"the slot's profile before the reload");
Files.writeString(f, yaml(WITH_OPUS_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertEquals("opus", registry.profileForSlot("architect:designer"),
"repointing the slot to a different profile must take effect with no restart");
}
@Test
void isSlotStopsReportingASlotRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertTrue(registry.isSlot("architect:designer"), "the slot is configured before the reload");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertFalse(registry.isSlot("architect:designer"),
"removing the slot from config must make isSlot say so on the very next call, "
+ "with no restart");
}
@Test
void isSlotStartsReportingASlotAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertFalse(registry.isSlot("architect:designer"), "no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertTrue(registry.isSlot("architect:designer"),
"a slot added by reload must be visible to isSlot with no restart");
}
@Test
void nameForSlotReflectsANameChangedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertEquals("designer", registry.nameForSlot("architect:designer"),
"the configured name before the reload");
Files.writeString(f, yaml(WITH_RENAMED_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertNull(registry.nameForSlot("architect:designer"),
"the old key no longer names a configured slot — it was renamed away by the reload");
assertEquals("architect-lead", registry.nameForSlot("architect:architect-lead"),
"the new name must be visible under its new qualified key with no restart");
}
// ── a bound architect is demoted, but the binding itself is not touched (fleetd #424) ───────
// The lead's corrected ruling: the PRIVILEGE a slot grants is revoked on the bound session's
// very next request, but the terminalToSlot BINDING itself is untouched by a reload — dropping
// it would double-book the slot key and break unbind's compare-safe contract. See the class
// doc's binding rule.
/** A caller identity resolving the one canned pane (terminal {@code term_a}) in {@link FakeHerdr}. */
private static ConnectionIdentity boundPaneIdentity() {
return new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> FakeHerdr.WORKER_PID);
}
@Test
void anArchitectAlreadyBoundToASlotIsDemotedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
assertEquals("architect:designer", registry.slotForTerminal("term_a"));
// Drive the real caller path, not the roleForSlot seam directly: CallerResolver.resolve is
// what a live request actually goes through (CallerResolver.java:220), and a resolver that
// ignored roleForSlot entirely would still pass a test that only checked the seam.
CallerResolver resolver = CallerResolver.withLeadsAndMembers(
boundPaneIdentity(), false, null, Map::of, registry);
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, before.role(), "sanity check: the harness binds term_a as an architect");
assertEquals("designer", before.name());
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
Principal after = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, after.role(),
"removing the slot from config must demote the bound session to worker on its "
+ "NEXT request — this is the ticket's whole point");
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
}
@Test
void theOriginalBindingStillOccupiesTheRemovedSlotSoASecondTerminalCannotClaimIt(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome removed = ref.reload();
assertTrue(removed.applied(), "the reload must actually take effect: " + removed.summary());
// The binding survives the removal untouched.
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
"a live binding must never be retroactively unbound by a config edit");
assertEquals(Map.of("term_a", "architect:designer"), registry.snapshot());
// Bring the slot back into config. If the binding had been silently dropped by the removal
// (rather than merely losing the privilege it grants), a second terminal could now claim
// the "freed" key — the exact double-booking the class doc's binding rule rules out.
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome restored = ref.reload();
assertTrue(restored.applied(), "the reload must actually take effect: " + restored.summary());
assertFalse(registry.bind("architect:designer", "term_b"),
"the slot is still occupied by term_a — a second terminal must not bind to it");
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"the slot is still occupied by term_a — a fresh reservation must not find it free");
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
"the original binding is unchanged throughout");
}
@Test
void unbindStillSucceedsForTheOriginalTerminalAfterItsSlotIsRemoved(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertTrue(registry.unbind("architect:designer", "term_a"),
"unbind must still work for a slot that config has since removed, or a session "
+ "that outlives its slot's removal could never release it");
assertNull(registry.slotForTerminal("term_a"));
assertEquals(Map.of(), registry.snapshot());
}
}
@@ -161,7 +161,15 @@ class ConfigRefProfileCoverageTest {
// the #323 bug. Only ideProjectDir and worktreeGroup were saved by a behavioural test in
// ConfigRefTest; the other two had none. So: growing this set now requires editing this
// line as well, which is a visible, deliberate diff rather than a quiet one.
assertEquals(Set.of("weight", "maxLoad", "credentialId"), excluded,
//
// fleetd #446: "exhaustedPattern" joined the set legitimately — it moved from compiled
// once at Fleetd.main startup to read live, cached by profile name, through
// LiveExhaustedPatterns (see that class's doc and ConfigRef's class doc). Unlike the #323
// hazard this comment warns about, it is pinned by its OWN behavioural tests:
// FleetdExhaustionDetectionArmedWiringTest proves the reload takes effect with no restart,
// and the same test rerun against the pre-#446 compile site (see that ticket's report) is
// what proves the assertion would have failed before the fix.
assertEquals(Set.of("weight", "maxLoad", "credentialId", "exhaustedPattern"), excluded,
"ConfigRef.LAUNCH_SETTINGS_EXCLUDED changed. A component belongs in it ONLY if it "
+ "is read live off the config supplier, not baked into a launcher at "
+ "startup. If you are adding one to silence this test, that is fleetd #323 "
@@ -1,10 +1,13 @@
package dev.ltms.fleet.config;
import dev.ltms.fleet.mcp.CharterToolSurface;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.function.Consumer;
import static org.junit.jupiter.api.Assertions.*;
@@ -38,6 +41,23 @@ class ConfigRefTest {
return new ConfigRef(f, FleetConfig.load(f));
}
/**
* The exact {@code Consumer<FleetConfig>} {@code Fleetd.main} wires into {@code ConfigRef}'s
* constructor as {@code extraValidation} (fleetd #474) — an adapter from {@code FleetConfig} to
* the raw charter map {@link CharterToolSurface#assertChartersNameOnlyRegisteredTools} takes.
* Built here rather than referencing {@code dev.ltms.fleet.Fleetd} directly, because that method
* is package-private to {@code dev.ltms.fleet} and this test lives in {@code
* dev.ltms.fleet.config} — but it calls the SAME production {@link CharterToolSurface} method
* {@code Fleetd} calls, so this proves the real check runs on reload, not a stand-in for it.
*/
private static final Consumer<FleetConfig> CHARTER_TOOL_SURFACE = cfg ->
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
private static ConfigRef refForWithCharterToolSurface(Path f) {
return new ConfigRef(f, FleetConfig.load(f), CHARTER_TOOL_SURFACE);
}
@Test
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
@@ -124,6 +144,87 @@ class ConfigRefTest {
assertSame(before, ref.get());
}
/**
* fleetd #474: {@code validateAll()} (via {@code validateCharters()}) only checks that a
* charter's KEY is a role wire name and its text is non-blank — it never looks at what the text
* names, so a charter naming {@code bridge_send} (the pre-CB-634 name, removed from the tool
* surface — #469's own motivating example) passes {@code validateAll()} and used to be applied
* on reload with nothing refusing it, even though the identical charter refuses {@code
* Fleetd.main} at startup ({@code FleetdStartupValidationTest
* #mainRefusesACharterNamingAnUnregisteredTool}). This is the reload-path proof: it drives
* {@link ConfigRef#reload()} itself (not a direct call to {@link
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), through the exact {@code
* extraValidation} wiring {@code Fleetd.main} uses, and the failure message must name both the
* charter key and the unknown tool — the same information the startup failure gives (ticket
* acceptance criterion 1).
*/
@Test
void aReloadRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
"""));
ConfigRef ref = refForWithCharterToolSurface(f);
FleetConfig before = ref.get();
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
"""));
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.error().contains("dev"),
"expected the charter key 'dev' in the refusal, got: " + out.error());
assertTrue(out.error().contains("bridge_send"),
"expected the unknown tool 'bridge_send' in the refusal, got: " + out.error());
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
// The running config must not move at all — half-applying this would leave the daemon in a
// state that could never have booted, exactly the outcome ConfigRef.java's class doc warns
// a cold-key refusal must avoid, and this check must avoid the same way.
assertSame(before, ref.get());
assertEquals("Send the final handoff through fleet_reply.\n",
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
}
/**
* fleetd #474 acceptance criterion 3, the positive case: a reload whose charter names only
* registered tools must still be ACCEPTED. An inverted filter (one that refuses every charter,
* or refuses on any {@code fleet_*}/{@code bridge_*} token regardless of registration) would pass
* the refusal test above alone — #469's own M5 mutation cell showed exactly that shape surviving
* a negative-only suite. This is the test that catches it.
*/
@Test
void aReloadAcceptsACharterNamingOnlyRegisteredTools(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
dev: old charter, no tool names
"""));
ConfigRef ref = refForWithCharterToolSurface(f);
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through fleet_reply, using fleet_send to delegate.
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals("config reloaded", out.summary());
assertEquals("Send the final handoff through fleet_reply, using fleet_send to delegate.\n",
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
}
/**
* The point of the whole class: a consumer holding the ref sees the new value without being
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
@@ -339,12 +440,18 @@ class ConfigRefTest {
}
/**
* CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map at startup
* (see ExhaustedPatternLookup), so a reload never re-reads it — a changed pattern must be
* reported deferred exactly like model/baseUrl, not silently claimed as applied.
* fleetd #446: exhaustedPattern moved from deferred to hot — {@code LiveExhaustedPatterns}
* reads {@code config.get()} fresh on every lookup, cached by profile name (not baked into a
* frozen map inside {@code Fleetd.main} any more, the way it was before this ticket, and the
* way its sibling {@code errorPattern} still is — see
* {@code changingAProfilesErrorPatternIsReportedAsDeferred} right below for the contrast).
* So a reload that changes ONLY exhaustedPattern must report a plain "config reloaded", with
* nothing deferred — the opposite of what this test asserted before fleetd #446, when
* ConfigRef.sameLaunchSettings still compared exhaustedPattern and this reload reported it as
* one of the changes needing a restart.
*/
@Test
void changingAProfilesExhaustedPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
void changingAProfilesExhaustedPatternTakesEffectWithNoRestart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
@@ -379,17 +486,21 @@ class ConfigRefTest {
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size(), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals(0, out.deferred().size(),
"exhaustedPattern is hot since fleetd #446 — changing only this key must not be "
+ "reported as needing a restart: " + out.deferred());
assertEquals(0, out.split().size(), out.split().toString());
assertEquals("config reloaded", out.summary());
// The snapshot carries the new value immediately — this IS the live value now, not
// merely what a restart would eventually pick up.
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
}
/**
* fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error pattern
* map at startup (see BackendErrorPatternLookup), the same way exhaustedPattern is above — a
* reload never re-reads it either, so a changed value must be reported deferred.
* map at startup (see BackendErrorPatternLookup) — unlike its sibling exhaustedPattern above,
* which fleetd #446 made hot, errorPattern was scoped OUT of that ticket on purpose and stays
* deferred: a reload still never re-reads it, so a changed value must be reported deferred.
*/
@Test
void changingAProfilesErrorPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
@@ -79,6 +79,25 @@ class ConfigRefTopLevelCoverageTest {
* <li>{@code memberCredentials} — read live at {@code Fleetd.java:198, 205, 729}.</li>
* <li>{@code memberLoginShell} — read live at
* {@code HerdrPeerLauncher#configuredMemberLoginShell}.</li>
* <li>{@code models} (fleetd #422) — membership is re-validated in full against the fresh
* config on every {@code ConfigRef.reload()} (via {@code FleetConfig#validateAll()}), so
* a bad edit is refused, never cached stale; the on/off half is read live by {@code
* CompositePeerLauncher.enforceModelEnabled} and its candidate filter, and by {@code
* fleet_profiles}/{@code GET /profiles} through {@code PeerLauncher.disabledModels()}.
* Unlike {@code fleet} below, nothing about {@code models} is baked into a startup-built
* object anywhere — there is no frozen half, so it belongs here whole rather than in
* {@code SPLIT_KEYS}.</li>
* <li>{@code leadRollover} (fleetd #480) — {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} and reads every field fresh on each
* {@code open()}/{@code confirm()} call, the same {@code () -> config.get().x()} shape
* {@code fleet}/{@code placement}/{@code models} use — see {@code ConfigRef}'s class doc
* Hot bullet. The only restart-only fact is structural, not a value going stale:
* {@code Fleetd.java} decides whether to construct the {@code LeadRollover} object at
* all off the startup snapshot (the same presence gate {@code leadHeartbeat:} uses), so
* a block ADDED where it was absent at boot needs a restart before anything exists to
* call — the same fact already true of adding a brand-new {@code profiles:} entry, which
* does not stop {@code profiles}' own hot sub-fields (weight/maxLoad/credentialId/
* exhaustedPattern) from being genuinely hot.</li>
* </ul>
*
* <p>{@code fleet} used to sit here too, on the strength of most of it (role pools, charters,
@@ -93,7 +112,7 @@ class ConfigRefTopLevelCoverageTest {
* while this test stayed green throughout.</p>
*/
private static final Set<String> HOT_EXCLUDED_TOP_LEVEL_KEYS =
Set.of("placement", "memberCredentials", "memberLoginShell");
Set.of("placement", "memberCredentials", "memberLoginShell", "models", "leadRollover");
@Test
void everyTopLevelComponentIsAccountedForInExactlyOneClass() {
@@ -120,7 +139,7 @@ class ConfigRefTopLevelCoverageTest {
// The escape hatch is pinned. Growing it requires editing this line — a visible, deliberate
// diff, not a quiet one. See the field javadoc above for what "belongs here" actually means.
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell"), hot,
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell", "models", "leadRollover"), hot,
"HOT_EXCLUDED_TOP_LEVEL_KEYS changed. A component belongs here ONLY if it is read "
+ "live off the config supplier, never because adding it makes this test "
+ "pass. If you are adding one to silence this test, that is fleetd #323 "
@@ -104,9 +104,16 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("configReload", new FleetConfig.ConfigReload(true, 10));
v.put("quarantineCooldownSeconds", 1800);
v.put("memberCredentials", null);
v.put("coordinator", new FleetConfig.Coordinator("amqp://coord-a", null, "self-a", 1));
v.put("coordinator", new FleetConfig.Coordinator("amqp://coord-a", null, "self-a", 1, null));
v.put("worktreeGroup", "group-a");
v.put("memberLoginShell", null);
v.put("memberSkills", "/skills/a");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-a"))));
// fleetd #480: leadRollover joins placement/memberCredentials/memberLoginShell as
// hot-excluded — never compared by any changed*Keys method, so it stays null like them.
v.put("leadRollover", null);
assertNamesMatchComponents(v);
return v;
}
@@ -144,9 +151,14 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("configReload", new FleetConfig.ConfigReload(false, 20));
v.put("quarantineCooldownSeconds", 3600);
v.put("memberCredentials", null);
v.put("coordinator", new FleetConfig.Coordinator("amqp://coord-b", null, "self-b", 2));
v.put("coordinator", new FleetConfig.Coordinator("amqp://coord-b", null, "self-b", 2, null));
v.put("worktreeGroup", "group-b");
v.put("memberLoginShell", null);
v.put("memberSkills", "/skills/b");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(false));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-b"))));
v.put("leadRollover", null);
assertNamesMatchComponents(v);
return v;
}
@@ -0,0 +1,172 @@
package dev.ltms.fleet.config;
import org.junit.jupiter.api.Test;
import java.lang.reflect.Constructor;
import java.lang.reflect.RecordComponent;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* Fleetd #358, the same "defect factory" #357 guarded on {@code FleetConfig.withDefaults()}
* (see {@code FleetConfigWithDefaultsPreservesEveryComponentTest}), reproduced here on
* {@link FleetConfig.Profile}. {@code Profile} carries a long back-compat constructor ladder — 8
* constructors, re-counted directly against the source rather than trusted from the ticket, at
* arities 25, 24, 22, 20, 18, 15, 14 and 12, against a canonical arity of 26 — and exactly ONE
* rebuild site, {@link FleetConfig.Profile#withProfile(String)}, whose own
* {@code return new Profile(...)} call is written at a literal 26-arg count. Add a 27th component
* and its established back-compat constructor at the old (26-arg) arity, and {@code withProfile}'s
* own call becomes a legal match for that new overload — silently dropping the new component every
* time a profile's name is defaulted from its {@code workers:} key.
*
* <p>Builds one {@link FleetConfig.Profile} through the TRUE canonical constructor — resolved by
* the record's own component types via {@code getDeclaredConstructor}, never by argument count —
* with a real, distinctive, non-null value in every component, calls {@link
* FleetConfig.Profile#withProfile(String)}, and asserts every component except {@code profile}
* itself survives unchanged, while {@code profile} comes back as the new name it was given.
*
* <p>Every value here is chosen so {@code Profile}'s own compact constructor (which normalizes
* several components — defaults {@code argv}/{@code kind}/{@code placement}/{@code workspace}/
* {@code gitHostEnv}, nulls a handful of blank-checked strings, clamps {@code weight}, coerces
* {@code subscription}) leaves it unchanged: every String is non-blank and already in the shape the
* compact constructor would otherwise coerce it to (e.g. {@code placement} is already lowercase),
* and every collection is non-empty. That is what makes "must survive unchanged" a valid assertion
* for every component below, the same reasoning {@code FleetConfigWithDefaultsPreservesEveryComponentTest}
* documents for {@code withDefaults()}.
*
* <p>{@link #EXCLUDED_FROM_SURVIVAL_CHECK} is kept deliberately empty and size-pinned by
* {@link #exclusionListSizeIsPinned()} — a checker whose escape hatch can grow to silence a failure
* is not a checker. Every one of {@code Profile}'s 26 current components has a real, non-null,
* non-blank value here and none is excluded.
*/
class FleetConfigProfileWithProfilePreservesEveryComponentTest {
private static final RecordComponent[] COMPONENTS = FleetConfig.Profile.class.getRecordComponents();
/** Deliberately empty today; grow it only with a matching justification, and re-pin the size. */
private static final Set<String> EXCLUDED_FROM_SURVIVAL_CHECK = Set.of();
/** One real, distinctive, non-null value per component, chosen to survive the compact ctor. */
private static Map<String, Object> baseValues() {
Map<String, Object> v = new LinkedHashMap<>();
v.put("profile", "profile-guard");
v.put("baseUrl", "https://guard.example/base");
v.put("model", "model-guard");
v.put("configDir", "/config/guard");
v.put("tokenEnv", "GUARD_TOKEN");
v.put("argv", List.of("guard-cmd"));
v.put("placement", "guard-placement");
v.put("workspace", "workspace-guard");
v.put("tabLabel", "tab-guard");
v.put("mcpUrl", "https://mcp.guard/");
v.put("cwd", "/cwd/guard");
v.put("parityOverlay", List.of(".guardrc"));
v.put("gitTokenEnv", "GUARD_GIT_TOKEN");
v.put("gitHostEnv", "GUARD_GIT_HOST");
v.put("kind", "claude-code");
v.put("env", Map.of("GUARD_ENV", "1"));
v.put("weight", 2.5f);
v.put("maxLoad", 4);
v.put("subscription", Boolean.TRUE);
v.put("exhaustedPattern", "pattern-guard");
v.put("credentialId", "cred-guard");
v.put("ideMcpUrl", "https://ide.guard/");
v.put("ideProjectDir", "ide-project-guard");
v.put("ideOpenCommand", "open-guard {dir}");
v.put("autoCompactWindow", 150_000);
v.put("errorPattern", "error-pattern-guard");
assertNamesMatchComponents(v);
return v;
}
/**
* Guards {@link #baseValues()} itself against drifting from the record's real shape — forgetting
* to add a new component here fails this assertion by name, rather than silently checking one
* component fewer than the record has.
*/
private static void assertNamesMatchComponents(Map<String, Object> values) {
Set<String> names = new TreeSet<>();
for (RecordComponent rc : COMPONENTS) {
names.add(rc.getName());
}
assertEquals(names, new TreeSet<>(values.keySet()),
"this test's value map has drifted from FleetConfig.Profile's actual components — "
+ "update baseValues() alongside the record");
}
/**
* Builds a {@link FleetConfig.Profile} through the TRUE canonical constructor — resolved by the
* record's own component types, not by argument count — so this never accidentally exercises a
* back-compat overload the way a literal {@code new Profile(...)} call risks doing.
*/
private static FleetConfig.Profile profileOf(Map<String, Object> values) throws ReflectiveOperationException {
Class<?>[] types = Arrays.stream(COMPONENTS).map(RecordComponent::getType).toArray(Class<?>[]::new);
Object[] args = Arrays.stream(COMPONENTS).map(rc -> values.get(rc.getName())).toArray();
Constructor<FleetConfig.Profile> ctor = FleetConfig.Profile.class.getDeclaredConstructor(types);
return ctor.newInstance(args);
}
@Test
void exclusionListSizeIsPinned() {
assertEquals(0, EXCLUDED_FROM_SURVIVAL_CHECK.size(),
"EXCLUDED_FROM_SURVIVAL_CHECK grew from 0 — every entry needs a justification in "
+ "this test class's javadoc AND this assertion re-pinned to the new size; a "
+ "growing exclusion list that silences failures on its own is not a guard");
}
/**
* The mutation this is built to catch: make {@code withProfile(String)}'s final constructor call
* literal at some arg count, add one more component to the record with a new back-compat
* constructor at the old arity, and the stale call silently rebinds. Every component here is real
* and non-null/non-blank, so none of it should be replaced by {@code withProfile}, except
* {@code profile} itself, which the method is documented to replace.
*/
@Test
void withProfilePreservesEveryOtherComponent() throws ReflectiveOperationException {
Map<String, Object> base = baseValues();
FleetConfig.Profile profile = profileOf(base);
FleetConfig.Profile renamed = profile.withProfile("renamed-profile-guard");
List<String> dropped = new ArrayList<>();
int checked = 0;
for (RecordComponent rc : COMPONENTS) {
String name = rc.getName();
if (EXCLUDED_FROM_SURVIVAL_CHECK.contains(name)) {
continue;
}
checked++;
Object expected = "profile".equals(name) ? "renamed-profile-guard" : base.get(name);
Object actual;
try {
actual = rc.getAccessor().invoke(renamed);
} catch (ReflectiveOperationException e) {
throw new RuntimeException("failed to read FleetConfig.Profile." + name + "()", e);
}
if (!Objects.equals(expected, actual)) {
dropped.add(String.format(Locale.ROOT,
"%s: withProfile() was expected to carry (%s) for '%s' but returned %s — a "
+ "component silently dropped by withProfile(), the shape of the "
+ "defect this test exists to catch (its final \"return new "
+ "Profile(...)\" call binding to a back-compat constructor instead "
+ "of the true canonical one)",
name, expected, name, actual));
}
}
System.out.printf(Locale.ROOT,
"FleetConfig.Profile.withProfile() component-survival coverage — %d components, %d "
+ "checked, %d excluded, %d survived%n",
COMPONENTS.length, checked, EXCLUDED_FROM_SURVIVAL_CHECK.size(), checked - dropped.size());
assertEquals(List.of(), dropped,
"withProfile() silently dropped these components: " + dropped);
}
}
@@ -1189,7 +1189,7 @@ class FleetConfigTest {
@Test
void coordinatorEffectiveUriHonorsUriEnv() {
FleetConfig.Coordinator withEnv = new FleetConfig.Coordinator(
"amqp://stale-clear-text@127.0.0.1:5672/coord", "LEAD_COORD_URI", "fleet01-lead", null);
"amqp://stale-clear-text@127.0.0.1:5672/coord", "LEAD_COORD_URI", "fleet01-lead", null, null);
assertEquals("amqp://from-env@127.0.0.1:5672/coord",
withEnv.effectiveUri(Map.of("LEAD_COORD_URI", "amqp://from-env@127.0.0.1:5672/coord")),
@@ -1200,12 +1200,61 @@ class FleetConfigTest {
"a blank uriEnv variable must not fall back to the literal uri");
FleetConfig.Coordinator noEnv = new FleetConfig.Coordinator(
"amqp://guest:guest@127.0.0.1:5672/coord", null, null, null);
"amqp://guest:guest@127.0.0.1:5672/coord", null, null, null, null);
assertEquals("amqp://guest:guest@127.0.0.1:5672/coord", noEnv.effectiveUri(Map.of()),
"the literal uri is used when no uriEnv is configured");
assertEquals(LeadMailbox.DEFAULT_PREFETCH, noEnv.prefetchOrDefault());
}
/**
* fleetd #361: the live block ships as just {@code uriEnv} + {@code selfId} (see
* {@code Fleetd.example.yaml} / the operator's real {@code fleetd.yaml}, gitignored). That exact
* shape, with no {@code peers:} key at all, must keep parsing unchanged after this field is added.
*/
@Test
void coordinatorBlockWithNoPeersKeyStillParses(@TempDir Path dir) throws Exception {
Path f = dir.resolve("coordinator-no-peers.yaml");
Files.writeString(f, """
bind:
port: 8080
coordinator:
uriEnv: COORD_AMQP_URI
selfId: mac
""");
FleetConfig cfg = FleetConfig.load(f);
assertNotNull(cfg.coordinator());
assertEquals("mac", cfg.coordinator().selfId());
assertEquals(List.of(), cfg.coordinator().peers(), "no peers: key means no configured peers, never null");
}
@Test
void coordinatorPeersParsesAndDropsBlankEntries(@TempDir Path dir) throws Exception {
Path f = dir.resolve("coordinator-peers.yaml");
Files.writeString(f, """
bind:
port: 8080
coordinator:
selfId: mac
uri: amqp://guest:guest@127.0.0.1:5672/coord
peers:
- fleet01
- ""
- fleet02
""");
FleetConfig cfg = FleetConfig.load(f);
assertEquals(List.of("fleet01", "fleet02"), cfg.coordinator().peers(),
"a blank peer entry must be dropped, never kept as an empty coord-id");
}
@Test
void coordinatorPeersDefaultsToEmptyWhenConstructedWithNull() {
FleetConfig.Coordinator c = new FleetConfig.Coordinator(
"amqp://guest:guest@127.0.0.1:5672/coord", null, "mac", null, null);
assertEquals(List.of(), c.peers(), "a null peers list must default to empty, never NPE downstream");
}
@Test
void absentWorktreeGroupLeavesItNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-worktree-group.yaml");
@@ -2580,4 +2629,391 @@ class FleetConfigTest {
"with no pool to choose from, every configured profile is a candidate and the "
+ "first one wins");
}
// ── idle-sleep guard: default-on config block ───────────────────────────────────────────────
@Test
void idleSleepGuardIsOnByDefaultWhenTheBlockIsEntirelyAbsent(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8080
""");
FleetConfig cfg = FleetConfig.load(f);
assertNull(cfg.idleSleepGuard(), "an absent block parses to null, unlike most other blocks here");
// The block itself is absent, but the FEATURE stays on: Fleetd treats a null block the
// same as enabled: true (see FleetConfig.idleSleepGuard's javadoc) — this test only pins
// the parse result, the on-by-default behaviour is Fleetd's own null check.
}
@Test
void idleSleepGuardExplicitlyEnabledIsOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
idleSleepGuard:
enabled: true
""");
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.idleSleepGuard().isEnabled());
}
@Test
void idleSleepGuardExplicitlyDisabledIsOff(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
idleSleepGuard:
enabled: false
""");
FleetConfig cfg = FleetConfig.load(f);
assertFalse(cfg.idleSleepGuard().isEnabled());
}
@Test
void idleSleepGuardBlockPresentButEmptyDefaultsToEnabled(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
idleSleepGuard: {}
""");
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.idleSleepGuard().isEnabled(),
"unlike ConfigReload/Health, this block defaults to ON even when present but empty");
}
// --- models: central allow-list of usable models -----------------------------------------
/**
* An absent {@code models:} block is today's behaviour exactly: no profile's {@code model:} is
* checked against anything, whatever it says. This is the "existing config keeps working"
* invariant — an operator on a gitignored {@code fleetd.yaml} that predates this feature must
* not be broken by upgrading the daemon.
*/
@Test
void absentModelsBlockValidatesNothing(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
baseUrl: http://gx10.gw:8000
model: totally-unheard-of-model-xyz
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/**
* {@code models:} present but with an empty (or absent) {@code allow:} must behave exactly like
* the block being absent — an operator adding the block for the first time with nothing in it
* yet must not be surprised by every profile suddenly refusing to start.
*/
@Test
void modelsBlockPresentButEmptyValidatesNothing(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
baseUrl: http://gx10.gw:8000
model: whatever-the-operator-typed
models:
allow: []
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/**
* The core of the ticket: a profile naming a model outside the configured allow-list fails
* config load, naming both the model and the profile that wanted it.
*/
@Test
void aProfileNamingAModelOutsideTheAllowListRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
rogue:
baseUrl: http://gx11.gw:8000
model: claude-opus-9000
models:
allow:
- model: claude-sonnet-5
- model: claude-haiku-5
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
assertTrue(e.getMessage().contains("rogue"), "the refusal names the offending profile");
assertTrue(e.getMessage().contains("claude-opus-9000"), "the refusal names the offending model");
assertFalse(e.getMessage().contains("'sonnet'"),
"the profile whose model IS allowed must not be reported");
}
/** A profile whose {@code model:} is on the allow-list loads fine. */
@Test
void aProfileNamingAModelOnTheAllowListLoadsFine(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/**
* A profile that never sets {@code model:} (a {@code subscription: true} profile relying on the
* account's own default is the live example) must not be refused just because an allow-list is
* active — there is nothing to check it against.
*/
@Test
void aProfileWithNoModelPassesEvenWithAnActiveAllowList(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
opus:
subscription: true
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/**
* Reproduces the live config shape this ticket measured: a mix of {@code amazon-bedrock},
* {@code opencode} and {@code openai}-backed profiles, some naming a bare Claude id and some a
* provider-prefixed opencode id, in ONE allow-list. Both forms are just opaque strings compared
* for exact equality — proves the shape decision holds against the shape actually seen live,
* not just against a synthetic single-provider example.
*/
@Test
void aBareClaudeIdAndAnOpencodeProviderPrefixedIdBothFitOneAllowList(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
terra:
kind: opencode
model: openai/gpt-5.6-terra
nova:
kind: opencode
model: amazon-bedrock/amazon.nova-pro-v1:0
models:
allow:
- model: claude-sonnet-5
- model: openai/gpt-5.6-terra
- model: amazon-bedrock/amazon.nova-pro-v1:0
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/** The same live shape, but one opencode profile's model is missing from the allow-list. */
@Test
void anUnlistedOpencodeProviderPrefixedModelRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
terra:
kind: opencode
model: openai/gpt-5.6-terra-withdrawn
models:
allow:
- model: claude-sonnet-5
- model: openai/gpt-5.6-terra
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
assertTrue(e.getMessage().contains("terra"), "the refusal names the offending profile");
assertTrue(e.getMessage().contains("openai/gpt-5.6-terra-withdrawn"),
"the refusal names the offending model, with its provider prefix intact");
}
/**
* Editing a {@code profiles:} entry alone must never be able to widen what is permitted — only
* editing {@code models.allow:} itself can. This is the invariant the ticket states explicitly;
* this test pins it by giving a profile a plausible-looking model that was never added to the
* allow-list and confirming it is still refused.
*/
@Test
void addingAProfileCannotWidenTheAllowListByItself(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
brand-new:
baseUrl: http://gx12.gw:8000
model: claude-sonnet-6-preview
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
assertTrue(e.getMessage().contains("brand-new"));
assertTrue(e.getMessage().contains("claude-sonnet-6-preview"));
}
/** {@code models} is a brand-new top-level key and must be recognized, not WARN-ed as unknown. */
@Test
void modelsIsAKnownTopLevelKey() {
assertTrue(FleetConfig.KNOWN_TOP_LEVEL_KEYS.contains("models"));
}
// --- fleetd #422: model on/off (spawn-time gate + hot reload) -----------------------------
/**
* Criterion 3: an {@code allow:} entry written before {@code enabled:} existed — no such key in
* the YAML at all — must behave exactly as it always did: on, and absent from {@link
* FleetConfig.Models#offIds()}. This is the old-style fixture the ticket asks for, proven
* through a real YAML load rather than only through the {@code ModelEntry(String)} back-compat
* constructor.
*/
@Test
void anOldStyleAllowEntryWithNoEnabledFieldStaysOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
assertTrue(cfg.models().ids().contains("claude-sonnet-5"), "membership is unaffected");
assertTrue(cfg.models().offIds().isEmpty(), "no enabled: field ⇒ nothing is off");
}
/**
* Criterion 7 (the whole point of the ticket, mirroring criterion 3's shape): a profile naming a
* model that is turned off ({@code enabled: false}) must still be VALID config —
* {@code validateModels()} checks membership only, never the on/off state, so it must not throw.
* Collapsing "off" into "removed from allow:" would make this test fail, which is exactly the
* design trap the ticket calls out.
*/
@Test
void anOffModelIsStillValidConfigForValidateModels(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels,
"an off model must stay a valid allow-list member — only the spawn gate reads enabled");
assertTrue(cfg.models().ids().contains("deepseek-v4-flash"));
assertTrue(cfg.models().offIds().contains("deepseek-v4-flash"));
}
/**
* Criterion 8: one {@code enabled: false} entry names a model, not a profile — every profile
* naming that model is off, on purpose (the live example: {@code deepseek-v4-flash} backs both
* {@code local} and {@code local-direct}). Proven here at the {@code Models}/config-load level;
* {@code CompositePeerLauncherTest} proves the spawn-time consequence for both profiles.
*/
@Test
void turningOffOneModelIsIndependentOfHowManyProfilesNameIt(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
local-direct:
baseUrl: http://local.gw:8001
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
// One allow-list entry, one off id — the fan-out to both profiles happens at the reader
// (CompositePeerLauncher), not by duplicating the entry per profile.
assertEquals(Set.of("deepseek-v4-flash"), cfg.models().offIds());
assertEquals("deepseek-v4-flash", cfg.profiles().get("local").model());
assertEquals("deepseek-v4-flash", cfg.profiles().get("local-direct").model());
}
/** An {@code enabled: true} entry (explicit, not just absent) also stays on — not just null. */
@Test
void explicitlyEnabledTrueStaysOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
models:
allow:
- model: claude-sonnet-5
enabled: true
""");
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.models().offIds().isEmpty());
}
}
@@ -0,0 +1,358 @@
package dev.ltms.fleet.config;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Method;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.Set;
import java.util.TreeSet;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The gap this class exists to close: mutation testing on the fleetd ticket "central allow-list
* of usable models" found that although {@link FleetConfig#validateModels()}'s own logic was well
* pinned, nothing proved either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()})
* still invoked it — deleting the call site left the full suite green (1478/0/0/0). A follow-up
* measurement (same technique — remove one call site, run the suite, not read the code) found the
* SAME gap for all five of {@link FleetConfig}'s other validators at startup, and for four of the
* six inside {@link ConfigRef#reload()}. This is a class of gap, not one line's mistake: every one
* of those thirteen tests called the validator itself directly, never the real caller that was
* supposed to.
*
* <p>The fix replaces the six individual {@code cfg.validateXxx()} calls at each of the two real
* call sites with one {@link FleetConfig#validateAll()}, which reaches every validator by
* reflection rather than by a hand-maintained list of names. A hand-maintained list of six names
* would have exactly the defect it replaces: the seventh validator someone adds next month has no
* reason to be added to it, and nothing would say so. This class proves TWO separate claims, and
* keeps them separate on purpose:
*
* <ol>
* <li>{@link #theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames()} and its neighbours
* prove the reflective sweep itself ({@link FleetConfig#invokeAllValidators}) is a general
* mechanism — it runs whatever public, no-arg, void {@code validateXxx()} methods a class
* happens to declare today, including a class with more of them than {@link FleetConfig}
* has right now. This is the proof that a future, real seventh validator on {@link
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
* seventh validator just to exercise the claim.</li>
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
* reaches each of today's six real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, so a
* single call to {@code validateAll()} is shown to reproduce every one of those six
* failures.</li>
* </ol>
*
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
* FleetdStartupValidationTest} (which prove the real startup call site still calls {@code
* validateAll()}) and the {@code ConfigRefTest} reload tests (which prove the same for {@link
* ConfigRef#reload()}), removing {@code cfg.validateAll();} from either real call site now fails
* a test in this module.
*
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
* is declared, which forces whoever adds it to look at this file.
*/
class FleetConfigValidateAllTest {
// ── Claim 1: the reflective sweep is a general mechanism, not six names in disguise ──────────
/**
* A throwaway fixture class, unrelated to {@link FleetConfig} in every way except shape: three
* public, no-arg, void methods named {@code validateXxx}. Proves the sweep works on ANY class
* with this shape, not on something special-cased to {@link FleetConfig}.
*/
static class ThreeValidators {
final List<String> ran = new ArrayList<>();
public void validateAlpha() {
ran.add("validateAlpha");
}
public void validateBeta() {
ran.add("validateBeta");
}
public void validateGamma() {
ran.add("validateGamma");
}
}
@Test
void theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames() {
ThreeValidators target = new ThreeValidators();
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateAlpha", "validateBeta", "validateGamma"), target.ran,
"every validateXxx() method on this unrelated class must run, in alphabetical "
+ "order — the sweep reads the class's own shape, not a name FleetConfig "
+ "happens to know about");
}
/**
* The core of the "self-maintaining" requirement: the exact same class shape as {@link
* ThreeValidators}, plus one more method — standing in for "a developer adds a validator next
* month". Nothing about the sweep changes to pick it up; the new method is invoked purely
* because it exists and matches the shape. This is what makes adding a seventh real validator
* to {@link FleetConfig} safe without touching {@link FleetConfig#validateAll()} or either
* call site — there is no "wire it in" step left to forget.
*/
static class FourValidators {
final List<String> ran = new ArrayList<>();
public void validateAlpha() {
ran.add("validateAlpha");
}
public void validateBeta() {
ran.add("validateBeta");
}
public void validateGamma() {
ran.add("validateGamma");
}
public void validateDelta() {
ran.add("validateDelta");
}
}
@Test
void addingAFourthValidatorMethodGetsSweptWithNoOtherChange() {
FourValidators target = new FourValidators();
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateAlpha", "validateBeta", "validateDelta", "validateGamma"),
sorted(target.ran),
"the fourth method must be reached automatically — proving a class can grow the "
+ "set of things it validates with no change to the sweep itself");
}
private static List<String> sorted(List<String> in) {
List<String> copy = new ArrayList<>(in);
copy.sort(String::compareTo);
return copy;
}
@Test
void aFailingValidatorStopsTheSweepAndPropagatesUnchanged() {
class OneFails {
@SuppressWarnings("unused")
public void validateOk() {
// passes
}
public void validateBoom() {
throw new IllegalStateException("refusing to start: boom");
}
}
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> FleetConfig.invokeAllValidators(new OneFails()));
assertEquals("refusing to start: boom", e.getMessage(),
"the real exception must propagate unchanged, not be wrapped or swallowed");
}
/**
* Every rule the sweep's filter applies, proven independently: only public, no-arg, void
* methods whose name starts with {@code "validate"} run, {@code validateAll} itself is
* excluded (so a class that declares one of its own — as {@link FleetConfig} does — cannot
* recurse into itself), and a same-shaped-but-wrongly-named or wrongly-shaped method never
* runs. A reader who "simplifies" the filter in {@link FleetConfig#invokeAllValidators} in a
* way that widens or narrows it breaks one of these.
*/
static class FilterEdgeCases {
final List<String> ran = new ArrayList<>();
public void validateReal() {
ran.add("validateReal");
}
/** Wrong name — must not run. */
public void checkSomething() {
ran.add("checkSomething");
}
/** Wrong shape — takes an argument. */
public void validateWithArg(String ignored) {
ran.add("validateWithArg");
}
/** Wrong shape — returns something. */
public boolean validateReturnsBoolean() {
ran.add("validateReturnsBoolean");
return true;
}
/** Excluded by name on purpose, so the sweep cannot call itself. */
public void validateAll() {
ran.add("validateAll");
}
}
@Test
void onlyPublicNoArgVoidValidateNamedMethodsRun() {
FilterEdgeCases target = new FilterEdgeCases();
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateReal"), target.ran,
"checkSomething (wrong name), validateWithArg (wrong shape), "
+ "validateReturnsBoolean (wrong shape), and validateAll (excluded by "
+ "name) must all be skipped");
}
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
/**
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
* above, on an unrelated class) — it is a visible denominator: today there are six, named
* here, so a reader adding a seventh sees this assertion name the new count rather than a
* silent pass at the old one.
*/
@Test
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
Set<String> names = new TreeSet<>();
for (Method m : FleetConfig.class.getMethods()) {
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
&& m.getParameterCount() == 0
&& m.getReturnType() == void.class
&& m.getName().startsWith("validate")
&& !m.getName().equals("validateAll")) {
names.add(m.getName());
}
}
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
"validateModels", "validateLeadRollover")), names,
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
+ "order. First confirm validateAll() still delegates to "
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
+ "test in this class, so this assertion is the only place that will ever "
+ "make you check. Only then update the expected set to match.");
}
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
private static String minimalValidYaml() {
return """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
""";
}
@Test
void aFullyValidConfigPassesValidateAll(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, minimalValidYaml());
assertDoesNotThrow(() -> FleetConfig.load(f).validateAll());
}
/**
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
* filter), exactly one of these six would start passing when it must not.
*/
@Test
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
// validateAuthExposure: a non-loopback bind without token mode.
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
bind:
host: 0.0.0.0
port: 8765
""", "auth.mode: token");
// validateLeadTabPrefixes: a fleet-wide tabLabel that starts with a lead's own tabPrefix.
assertValidateAllRefuses(dir, "lead-tab-prefixes.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
tabLabel: "lead: {role} {profile}"
leaders:
opus:
tab: "lead: opus"
""", "fleet.tabLabel");
// validateSubscriptionProfiles: subscription: true with env: reseating ANTHROPIC_BASE_URL.
assertValidateAllRefuses(dir, "subscription-profiles.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
""", "ANTHROPIC_BASE_URL");
// validateCharters: a charter key that is not a role wire name.
assertValidateAllRefuses(dir, "charters.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
architetc: text
""", "architetc");
// validateMembers: an architect slot naming an unconfigured profile.
assertValidateAllRefuses(dir, "members.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
baseUrl: http://gx10.gw:8000
fleet:
architects:
lead-designer:
profile: sonnet
""", "lead-designer");
// validateModels: a profile naming a model outside the configured allow-list.
assertValidateAllRefuses(dir, "models.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
rogue:
baseUrl: http://gx11.gw:8000
model: claude-opus-9000
models:
allow:
- model: claude-sonnet-5
""", "rogue");
}
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
String mustContain) throws Exception {
Path f = dir.resolve(fileName);
Files.writeString(f, yaml);
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll,
fileName + ": validateAll() must refuse this config");
assertTrue(e.getMessage().contains(mustContain),
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
+ e.getMessage());
}
}
@@ -0,0 +1,201 @@
package dev.ltms.fleet.config;
import org.junit.jupiter.api.Test;
import java.lang.reflect.Constructor;
import java.lang.reflect.RecordComponent;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* Guards against a "defect factory" built into this file's own established pattern, found live
* while building the (parked) idle-sleep-guard PR: every time a component is added to
* {@link FleetConfig}, the record grows by one arg AND a new back-compat constructor is added at
* the OLD arity, so existing callers keep compiling. That is correct and required — see the
* constructor ladder just below the record header. But {@link #withDefaults()}'s own {@code return
* new FleetConfig(...)} call sits in this same file, written at a literal argument count. The very
* next time a component is added, the freshly-added back-compat constructor at the OLD arity
* silently captures that stale call, because it is now a legal overload at that arg count too. It
* compiles. Every other test passes, because nothing else exercises the new field. The new
* component is defaulted away — {@code null}, or whatever that back-compat overload defaults it to
* — on every {@link FleetConfig#load}. Measured, not theoretical: this exact sequence happened
* live when the {@code idleSleepGuard} component was added on a sibling branch; it was caught only
* because that branch's own new tests happened to assert on the new field's value.
*
* <p>This test proves the opposite property, and does it in a way that survives the next field
* being added without being rewritten: reflectively enumerate {@link FleetConfig}'s own record
* components (never a hardcoded count — the arity is exactly what changes over time), build one
* config through the true canonical constructor with a real, distinctive, non-null value in EVERY
* component (reusing the exact reflective-construction pattern
* {@link ConfigRefTopLevelReportingCoverageTest} already established for this file:
* {@code getDeclaredConstructor(exact record-component types)}, which resolves the canonical
* constructor by its true shape, not by binding to whichever overload happens to match arg count —
* the same way Jackson resolves it), call the real {@link FleetConfig#withDefaults()}, and assert
* every one of those values survives unchanged.
*
* <p>Why this is a valid check for every component, not just some: {@link #withDefaults()}'s own
* comments document that it only ever REPLACES a component when the incoming value is {@code null}
* (or blank, for {@code placement}) — {@code broker}/{@code primary}/{@code leadHeartbeat}/
* {@code configReload}/{@code coordinator}/{@code worktreeGroup}/{@code memberLoginShell}/
* {@code memberSkills}/{@code idleSleepGuard} are left as-is unconditionally, and {@code bind}/{@code guard}/{@code lifecycle}/{@code auth}/
* {@code fleet}/{@code quarantineCooldownSeconds}/{@code memberCredentials}/{@code placement} are
* replaced only on null/blank input. A value that is never null or blank going in must therefore
* never change coming out, for every current component. No exclusion is needed today.
*
* <p>{@link #EXCLUDED_FROM_SURVIVAL_CHECK} exists anyway, kept deliberately empty and size-pinned
* by {@link #exclusionListSizeIsPinned()}: a future component that {@code withDefaults()} is
* <em>documented</em> to transform unconditionally (unlike every field today) would legitimately
* need one. Pinning the size at 0 means growing that set to make a failure go away is itself a
* visible diff to this test, not a silent one — a checker that can be silenced by adding to its
* own escape hatch is not a checker.
*/
class FleetConfigWithDefaultsPreservesEveryComponentTest {
private static final RecordComponent[] COMPONENTS = FleetConfig.class.getRecordComponents();
/** See the class javadoc — deliberately empty today; grow it only with a matching justification. */
private static final Set<String> EXCLUDED_FROM_SURVIVAL_CHECK = Set.of();
/** One real, distinctive, non-null (non-blank where blankness would mean "unset") value per component. */
private static Map<String, Object> baseValues() {
Map<String, Object> v = new LinkedHashMap<>();
v.put("bind", new FleetConfig.Bind("127.0.0.1", 8765));
v.put("herdrSocket", "~/.config/herdr/guard.sock");
v.put("memberHerdrSocket", "~/.config/herdr/member-guard.sock");
v.put("profiles", Map.of("sonnet", minimalProfile("sonnet")));
v.put("guard", new FleetConfig.Guard(List.of("host-guard")));
v.put("worktreeRoot", "/wt/guard");
v.put("lifecycle", new FleetConfig.Lifecycle(300, 5, 30, true));
v.put("spawnReadyTimeoutMs", 12_345);
v.put("spawnReadyPollMs", 234);
v.put("broker", new FleetConfig.Broker("amqp://guard", null, 7));
v.put("primary", new FleetConfig.Primary("term-guard", 4, 4000));
v.put("fleet", new FleetConfig.Fleet(
Map.of("opus", new FleetConfig.Leader("sonnet", "lead: opus-guard", 1, null, 10,
"claude", null, null, null)),
Map.of(), Map.of(), Map.of(), Map.of(), "{role}: {profile} #{n}"));
v.put("leadHeartbeat", new FleetConfig.LeadHeartbeat(301, 61_000L, 4));
v.put("health", new FleetConfig.Health(true, 31, 601, 61, null));
v.put("placement", "round-robin");
v.put("auth", new FleetConfig.Auth("loopback-trust", null));
v.put("configReload", new FleetConfig.ConfigReload(true, 11));
v.put("quarantineCooldownSeconds", 1801);
v.put("memberCredentials", new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT,
List.of("git"), List.of("git", "ssh"), null));
v.put("coordinator", new FleetConfig.Coordinator("amqp://coord-guard", null, "self-guard", 3, null));
v.put("worktreeGroup", "group-guard");
v.put("memberLoginShell", "/bin/zsh");
v.put("memberSkills", "/skills/guard");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-guard"))));
// fleetd #480: leadRollover is left as-is unconditionally by withDefaults() (see its
// comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value
// here proves it, rather than leaving it null and proving nothing.
v.put("leadRollover", new FleetConfig.LeadRollover(
"/handover/guard.md", true, 3600, 20, 20, "read the handover file"));
assertNamesMatchComponents(v);
return v;
}
/** A minimal, otherwise-null {@link FleetConfig.Profile} — just enough to name one in a map. */
private static FleetConfig.Profile minimalProfile(String name) {
return new FleetConfig.Profile(name, null, null, null, null, null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null, null, null, null,
null, null);
}
/**
* Guards {@link #baseValues()} itself against drifting from the record's real shape — the same
* assurance {@link ConfigRefTopLevelReportingCoverageTest} already relies on. This is what makes
* "no hardcoded arity" true in practice: forgetting to add a new component here fails this
* assertion by name, rather than silently checking one component fewer than the record has.
*/
private static void assertNamesMatchComponents(Map<String, Object> values) {
Set<String> names = new TreeSet<>();
for (RecordComponent rc : COMPONENTS) {
names.add(rc.getName());
}
assertEquals(names, new TreeSet<>(values.keySet()),
"this test's value map has drifted from FleetConfig's actual top-level components — "
+ "update baseValues() alongside the record");
}
/**
* Builds a {@link FleetConfig} through the TRUE canonical constructor — resolved by the record's
* own component types, not by argument count — so this never accidentally exercises a
* back-compat overload the way a literal {@code new FleetConfig(...)} call risks doing.
*/
private static FleetConfig configOf(Map<String, Object> values) throws ReflectiveOperationException {
Class<?>[] types = Arrays.stream(COMPONENTS).map(RecordComponent::getType).toArray(Class<?>[]::new);
Object[] args = Arrays.stream(COMPONENTS).map(rc -> values.get(rc.getName())).toArray();
Constructor<FleetConfig> ctor = FleetConfig.class.getDeclaredConstructor(types);
return ctor.newInstance(args);
}
@Test
void exclusionListSizeIsPinned() {
assertEquals(0, EXCLUDED_FROM_SURVIVAL_CHECK.size(),
"EXCLUDED_FROM_SURVIVAL_CHECK grew from 0 — every entry needs a justification in "
+ "this test class's javadoc AND this assertion re-pinned to the new size; a "
+ "growing exclusion list that silences failures on its own is not a guard");
}
/**
* The mutation this is built to catch: make {@code withDefaults()}'s final constructor call
* literal at some arg count, add one more component to the record with a new back-compat
* constructor at the old arity, and the stale call silently rebinds. Every component here is
* real and non-null (non-blank for the one String — {@code placement} — where blank has
* meaning), so none of it should be replaced by {@code withDefaults()}; any component that
* comes back different was silently dropped.
*/
@Test
void everyComponentGivenARealValueSurvivesWithDefaults() throws ReflectiveOperationException {
Map<String, Object> base = baseValues();
FleetConfig config = configOf(base);
FleetConfig defaulted = config.withDefaults();
List<String> dropped = new ArrayList<>();
int checked = 0;
for (RecordComponent rc : COMPONENTS) {
String name = rc.getName();
if (EXCLUDED_FROM_SURVIVAL_CHECK.contains(name)) {
continue;
}
checked++;
Object expected = base.get(name);
Object actual;
try {
actual = rc.getAccessor().invoke(defaulted);
} catch (ReflectiveOperationException e) {
throw new RuntimeException("failed to read FleetConfig." + name + "()", e);
}
if (!Objects.equals(expected, actual)) {
dropped.add(String.format(Locale.ROOT,
"%s: withDefaults() was given a real, non-null value (%s) for '%s' but "
+ "returned %s — a component silently dropped by withDefaults(), the "
+ "shape of the defect this test exists to catch (its final "
+ "\"return new FleetConfig(...)\" call binding to a back-compat "
+ "constructor instead of the true canonical one)",
name, expected, name, actual));
}
}
System.out.printf(Locale.ROOT,
"FleetConfig.withDefaults() component-survival coverage — %d components, %d checked, "
+ "%d excluded, %d survived%n",
COMPONENTS.length, checked, EXCLUDED_FROM_SURVIVAL_CHECK.size(), checked - dropped.size());
assertEquals(List.of(), dropped,
"withDefaults() silently dropped these real, given components: " + dropped);
}
}
@@ -0,0 +1,232 @@
package dev.ltms.fleet.deploy;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Pattern;
import java.util.stream.Stream;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.params.ParameterizedTest;
import org.junit.jupiter.params.provider.MethodSource;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #360: {@code deploy/fleetd.service} started clean on fleet01 and broke the daemon in
* three ways that nothing logs at startup.
*
* <ul>
* <li>{@code ProtectSystem=}, {@code ProtectHome=}, {@code ProtectKernelTunables=} and
* {@code ProtectControlGroups=} each give the unit its own mount namespace. fleetd resolves
* a caller's role by running {@code lsof} to find the loopback peer PID (see
* {@code mcp/LsofPeerPidLookup}). Inside such a namespace {@code lsof} returns nothing,
* every caller falls back to ANONYMOUS, and the primary is refused every orchestration call
* with "unauthenticated: anonymous may not SPAWN" -- while healthz still reports ok.
* Measured by bisection on fleet01 2026-09-05 (lsof line count): no sandbox 3,
* {@code ProtectSystem=strict} 0, {@code ProtectHome=read-only} 0,
* {@code ProtectKernelTunables} 0, {@code ProtectControlGroups} 0,
* {@code RestrictSUIDSGID} 3, {@code NoNewPrivileges} 3 -- so only the first four are
* forbidden.
* <li>{@code PrivateTmp=true} gives the unit its own {@code /tmp}. fleetd writes the member
* ZDOTDIR credential-scrub directory and the opencode config directory under
* {@code java.io.tmpdir}, and the member pane -- a child of a *different* unit
* (herdr.service) -- has to read them back. A private {@code /tmp} on either unit turns the
* credential scrub into a silent no-op.
* <li>Running {@code java} directly from {@code ExecStart} skips the login shell. Every secret
* fleetd needs (AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in
* a file only the login shell sources; systemd runs no login shell on its own. Started that
* way the daemon boots fine with empty credentials, and the failure appears hours later as a
* member that cannot open a pull request.
* </ul>
*
* <p>Review round 2 on fleetd #360 found the same shape one file further down the chain:
* {@code herdr.service}'s {@code ExecStart} only names {@code deploy/herdr-inner.sh}, so that
* script is the ONLY place two more of these properties live, and nothing was reading it.
*
* <ul>
* <li>If the script does not exec a login shell, herdr -- and every member pane it spawns as a
* child -- starts with none of the secrets only a login shell sources, the same silent
* empty-credentials failure as fleetd.service's {@code ExecStart}, one process further away.
* <li>If the script does not set a real, non-zero pty size before starting herdr (with
* {@code stty}), herdr reports a 0x0 window and every pane spawn fails with
* "ghostty error -2" -- which surfaces as a fleetd spawn failure, not a herdr one.
* </ul>
*
* <p>None of these five failures makes the unit (or the script) fail to start, or makes
* {@code /healthz} report unhealthy, so nothing short of reading the files catches a regression.
* This test is that read.
*
* <p><b>It checks source text, not behaviour</b> -- it cannot start systemd or fork an mount
* namespace in a build sandbox. It parses the same two files an operator would install and fails
* if a forbidden directive is active, exactly the way {@code McpContractDocTest} guards
* {@code docs/MCP-Contract.md} against naming a tool that does not exist.
*/
class SystemdUnitSafetyTest {
/** Tests run with the module directory (fleetd/) as cwd; deploy/ is the repo-root sibling. */
private static final Path FLEETD_SERVICE = Path.of("../deploy/fleetd.service");
private static final Path HERDR_SERVICE = Path.of("../deploy/herdr.service");
private static final Path HERDR_INNER_SCRIPT = Path.of("../deploy/herdr-inner.sh");
private static final List<String> NAMESPACING_DIRECTIVES = List.of(
"ProtectSystem", "ProtectHome", "ProtectKernelTunables", "ProtectControlGroups");
/**
* A shell invoked with a login flag, e.g. {@code zsh -lc '...'} or {@code /bin/sh -l}. The
* property under test is the {@code -l}, not the specific shell or the rest of its flags, so
* this matches a token ending in "sh" followed by a flag cluster containing "l" -- tighter
* than a bare {@code contains("-lc")}, which a "-lc" anywhere in the line would also satisfy.
*/
private static final Pattern LOGIN_SHELL_INVOCATION =
Pattern.compile("(?:^|\\s)\\S*sh\\s+-[A-Za-z]*l[A-Za-z]*\\b");
private static final Pattern POSITIVE_ROWS = Pattern.compile("\\brows\\s+([1-9]\\d*)\\b");
private static final Pattern POSITIVE_COLS = Pattern.compile("\\bcols\\s+([1-9]\\d*)\\b");
static Stream<Path> bothUnits() {
return Stream.of(FLEETD_SERVICE, HERDR_SERVICE);
}
private static boolean invokesALoginShell(List<String> activeLines) {
return activeLines.stream().anyMatch(l -> LOGIN_SHELL_INVOCATION.matcher(l).find());
}
/** An active {@code stty} line naming both a positive row count and a positive column count. */
private static boolean setsANonZeroPtySize(List<String> activeLines) {
return activeLines.stream()
.filter(l -> l.startsWith("stty"))
.anyMatch(l -> POSITIVE_ROWS.matcher(l).find() && POSITIVE_COLS.matcher(l).find());
}
/** Lines that are actually in force: comments and blank lines don't count. */
private static List<String> activeLines(Path unit) throws Exception {
List<String> active = new ArrayList<>();
for (String line : Files.readAllLines(unit)) {
String stripped = line.strip();
if (!stripped.isEmpty() && !stripped.startsWith("#")) {
active.add(stripped);
}
}
return active;
}
@ParameterizedTest
@MethodSource("bothUnits")
@DisplayName("[SOURCE TEXT] no active mount-namespacing directive -- it blinds lsof and turns every caller ANONYMOUS")
void doesNotActivateAMountNamespace(Path unit) throws Exception {
List<String> active = activeLines(unit);
for (String directive : NAMESPACING_DIRECTIVES) {
Pattern activeDirective = Pattern.compile("^" + Pattern.quote(directive) + "\\s*=");
List<String> hits = active.stream().filter(l -> activeDirective.matcher(l).find()).toList();
assertTrue(hits.isEmpty(),
unit + " sets " + directive + " (" + hits + "). That directive gives the unit its "
+ "own mount namespace; inside it, fleetd's lsof-based caller lookup "
+ "(mcp/LsofPeerPidLookup) returns nothing, so every MCP caller falls back to "
+ "ANONYMOUS and the primary is refused every orchestration call with "
+ "\"unauthenticated: anonymous may not SPAWN\" -- while the daemon still "
+ "starts and /healthz still reports ok. Measured on fleet01 2026-09-05 "
+ "(fleetd #360). A commented-out mention in the file's own DO-NOT-add block "
+ "is fine; an active directive is not.");
}
}
@ParameterizedTest
@MethodSource("bothUnits")
@DisplayName("[SOURCE TEXT] PrivateTmp is not true -- it silently no-ops the credential scrub")
void privateTmpIsNotTrue(Path unit) throws Exception {
List<String> active = activeLines(unit);
Pattern privateTmpTrue = Pattern.compile("^PrivateTmp\\s*=\\s*true\\b");
boolean hasPrivateTmpTrue = active.stream().anyMatch(l -> privateTmpTrue.matcher(l).find());
assertFalse(hasPrivateTmpTrue,
unit + " sets PrivateTmp=true. fleetd writes the member ZDOTDIR credential-scrub "
+ "directory and the opencode config directory under java.io.tmpdir, and the "
+ "member pane -- a child of a DIFFERENT unit -- has to read them back. A "
+ "private /tmp on either unit turns the credential scrub into a silent no-op: "
+ "no error, no log line, the scrub just never happens (fleetd #360).");
}
@Test
@DisplayName("[SOURCE TEXT] fleetd.service's ExecStart goes through a login shell -- otherwise every secret is empty")
void execStartUsesALoginShell() throws Exception {
List<String> active = activeLines(FLEETD_SERVICE);
List<String> execStartLines = active.stream().filter(l -> l.startsWith("ExecStart=")).toList();
assertTrue(execStartLines.size() == 1,
FLEETD_SERVICE + " must have exactly one active ExecStart= line; found "
+ execStartLines.size() + ": " + execStartLines);
String execStart = execStartLines.get(0);
assertTrue(LOGIN_SHELL_INVOCATION.matcher(execStart).find(),
FLEETD_SERVICE + "'s ExecStart (" + execStart + ") does not run a login shell (a shell "
+ "invoked with a \"-l\" flag, e.g. \"zsh -lc\"). Every secret fleetd needs "
+ "(AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in a "
+ "file only the login shell sources; systemd runs no login shell on its own. "
+ "Running java directly boots fine with every credential empty, and the failure "
+ "surfaces hours later as a member that cannot open a pull request (fleetd #360).");
}
@Test
@DisplayName("[SOURCE TEXT] herdr-inner.sh execs a login shell -- otherwise herdr and every member it spawns start with empty credentials")
void herdrInnerScriptUsesALoginShell() throws Exception {
List<String> active = activeLines(HERDR_INNER_SCRIPT);
assertTrue(invokesALoginShell(active),
HERDR_INNER_SCRIPT + " does not exec a login shell (a shell invoked with a \"-l\" flag, "
+ "e.g. \"zsh -lc\"). herdr.service's ExecStart only names this script, so this "
+ "is the ONLY place herdr's login-shell property lives. Without it, herdr -- and "
+ "every member pane it spawns as a child of herdr -- starts with none of the "
+ "secrets that only a login shell sources (this host: ~/.fleet/secrets.sh via "
+ "~/.zprofile), and the failure surfaces hours later as a member with no "
+ "credentials at all, not just fleetd (fleetd #360).");
}
@Test
@DisplayName("[SOURCE TEXT] herdr-inner.sh sets a non-zero pty size before starting herdr -- otherwise every pane spawn fails with \"ghostty error -2\"")
void herdrInnerScriptSetsANonZeroPtySize() throws Exception {
List<String> active = activeLines(HERDR_INNER_SCRIPT);
assertTrue(setsANonZeroPtySize(active),
HERDR_INNER_SCRIPT + " does not run \"stty rows N cols M\" with both N and M positive "
+ "before starting herdr. Without a real pty size, herdr reports a 0x0 window and "
+ "every pane spawn fails with \"ghostty error -2\" -- which surfaces as a fleetd "
+ "spawn failure, not a herdr one, and gives no hint that the actual cause is this "
+ "script (fleetd #360).");
}
/**
* The denominator guard, same shape as {@code McpContractDocTest}'s: a check that scans for a
* forbidden pattern passes trivially if it is handed nothing to scan. Pin that all three files
* exist, are non-trivial, and that the DO-NOT block's commented mentions are still there -- so
* the "comment survives, directive doesn't" distinction above is actually being exercised.
*/
@Test
@DisplayName("[SOURCE TEXT] the namespacing check is not vacuous -- all three files exist and the DO-NOT block still mentions every forbidden directive in a comment")
void theCheckActuallyHasSomethingToCheck() throws Exception {
assertTrue(Files.size(FLEETD_SERVICE) > 200,
FLEETD_SERVICE + " is missing or unexpectedly small -- the checks above would pass "
+ "vacuously against an empty or absent file");
assertTrue(Files.size(HERDR_SERVICE) > 200,
HERDR_SERVICE + " is missing or unexpectedly small -- the checks above would pass "
+ "vacuously against an empty or absent file");
assertTrue(Files.size(HERDR_INNER_SCRIPT) > 50,
HERDR_INNER_SCRIPT + " is missing or unexpectedly small -- the login-shell and "
+ "pty-size checks above would either pass vacuously or fail with an opaque "
+ "\"file not found\" instead of naming the actual consequence (empty "
+ "credentials / \"ghostty error -2\") against an empty or absent file");
String fleetdService = Files.readString(FLEETD_SERVICE);
for (String directive : NAMESPACING_DIRECTIVES) {
assertTrue(fleetdService.contains(directive),
FLEETD_SERVICE + " no longer mentions " + directive + " anywhere, not even in the "
+ "DO-NOT-add comment block that explains why it must stay out. That comment "
+ "is the whole point of fleetd #360 -- it is what stops the next edit from "
+ "re-adding the directive without knowing why.");
}
}
}
@@ -30,6 +30,7 @@ import java.util.concurrent.ScheduledFuture;
import java.util.concurrent.ScheduledThreadPoolExecutor;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.BiConsumer;
import java.util.function.LongSupplier;
@@ -206,6 +207,139 @@ class FleetHealthMonitorTest {
.workingSuspectAfterOrDefault());
}
// --- fleetd #386: a stall detector whose only clock freezes with a sleeping host is worse
// than a silent one — it reports "quiet" for a member that was genuinely busy for hours.
@Test void monotonicClockFrozenPastThresholdOnRealClockStillReportsStallSuspected() {
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
AtomicLong mono = new AtomicLong(0);
AtomicLong real = new AtomicLong(0);
FakeHerdr herdr = new FakeHerdr().withAgent("busy", "term_busy", "pane_busy", "tab_busy");
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = monitorWithClocks(herdr,
List.of(member("term_busy", MemberSession.State.BUSY, 0, 0)), scheduler,
mono::get, real::get, 60, 600, (_, _) -> { });
monitor.tick(); // establishes the clock baseline; nothing has diverged yet
assertEquals(0, appender.list.stream().filter(event -> event.getFormattedMessage()
.contains("state=STALL_SUSPECTED")).count());
// The host "sleeps": the monotonic clock stands completely still while the real clock
// keeps moving, past the 600s stall threshold.
real.set(TimeUnit.SECONDS.toNanos(700));
monitor.tick();
monitor.stop();
assertTrue(appender.list.stream().anyMatch(event -> event.getFormattedMessage()
.contains("member=term_busy state=STALL_SUSPECTED")),
"the real clock crossed the stall threshold even though the monotonic clock never moved");
} finally {
logger.detachAppender(appender);
}
}
@Test void clockDivergenceIsLoggedOnceNotOncePerTick() {
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
AtomicLong mono = new AtomicLong(0);
AtomicLong real = new AtomicLong(0);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = monitorWithClocks(new FakeHerdr(), List.of(), scheduler,
mono::get, real::get, 60, 600, (_, _) -> { });
monitor.tick(); // baseline: no divergence possible yet
// One sleep gap: the monotonic clock is frozen while the real clock jumps far past one
// tick interval (60s).
real.set(TimeUnit.SECONDS.toNanos(700));
monitor.tick();
// The host is awake again: both clocks advance together from here, so no more divergence.
mono.set(TimeUnit.SECONDS.toNanos(10));
real.set(TimeUnit.SECONDS.toNanos(710));
monitor.tick();
mono.set(TimeUnit.SECONDS.toNanos(20));
real.set(TimeUnit.SECONDS.toNanos(720));
monitor.tick();
monitor.stop();
assertEquals(1, appender.list.stream().filter(event -> event.getFormattedMessage()
.contains("the monotonic clock did not advance"))
.count(), "one sleep gap must produce exactly one divergence line, not one per tick");
} finally {
logger.detachAppender(appender);
}
}
/**
* fleetd #386 follow-up, added on merge. The fix carries a PER-MEMBER drift baseline, so drift
* from a sleep that happened BEFORE a member went busy is never charged to that member. The
* two tests shipped with the fix both start with the member already BUSY, so a single global
* baseline passes them — this one fails without the per-member map.
*
* <p>Order matters: the host sleeps while nothing is busy, and only then does a member take a
* turn. Its stall clock must start at zero.
*/
@Test void driftFromASleepBeforeAMemberWentBusyIsNotChargedToThatMember() {
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
AtomicLong mono = new AtomicLong(0);
AtomicLong real = new AtomicLong(0);
AtomicReference<List<MemberSession>> roster = new AtomicReference<>(List.of());
FakeHerdr herdr = new FakeHerdr().withAgent("busy", "term_busy", "pane_busy", "tab_busy");
var scheduler = Executors.newSingleThreadScheduledExecutor();
AgentControl agents = new AgentControl(herdr);
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, roster::get,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, mono::get, real::get, 60, 600, (_, _) -> { });
monitor.tick(); // baseline, no members yet
// The host sleeps for 700s with nobody busy: the monotonic clock stands still.
real.set(TimeUnit.SECONDS.toNanos(700));
monitor.tick();
// Awake again. Only NOW does a member start a turn, with a fresh activity stamp taken
// from the monotonic clock. Both clocks advance together from here.
mono.set(TimeUnit.SECONDS.toNanos(10));
real.set(TimeUnit.SECONDS.toNanos(710));
roster.set(List.of(member("term_busy", MemberSession.State.BUSY,
0, TimeUnit.SECONDS.toNanos(10))));
monitor.tick();
mono.set(TimeUnit.SECONDS.toNanos(20));
real.set(TimeUnit.SECONDS.toNanos(720));
monitor.tick();
monitor.stop();
assertEquals(0, appender.list.stream().filter(event -> event.getFormattedMessage()
.contains("member=term_busy state=STALL_SUSPECTED")).count(),
"the member has been busy for 10s, not 710s — the earlier sleep is not its stall");
} finally {
logger.detachAppender(appender);
}
}
private static FleetHealthMonitor monitorWithClocks(FakeHerdr herdr, List<MemberSession> roster,
java.util.concurrent.ScheduledExecutorService scheduler, LongSupplier clock,
LongSupplier realtimeClock, long intervalSeconds, long workingSuspectAfterSeconds,
BiConsumer<String, String> failTarget) {
AgentControl agents = new AgentControl(herdr);
return new FleetHealthMonitor(agents, () -> roster,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, clock, realtimeClock, intervalSeconds, workingSuspectAfterSeconds, failTarget);
}
@Test void goneMemberRecoveryLogsOnceWithoutRefiringTargetFailure() {
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
Level previousLevel = logger.getLevel();
@@ -17,15 +17,105 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
* the throwaway space down.
*
* <p>The seed shell's own startup (restoring its session, printing its banner) is asynchronous
* and its length is not a fleetd contract — measured here at ~2.5s on one host (fleetd #449).
* The old version used two fixed sleeps: 1000ms before typing, then 800ms before reading. It
* failed, and the pane showed the typed line followed by the startup banner with no command
* output at all — which looks, at a glance, exactly like the env map never reaching the shell.
* So this polls for a real signal instead of guessing a sleep length.
*
* <p><strong>What was measured, and what was not.</strong> Polling fixes it: 5 standalone runs
* green. The load-bearing half is {@link #waitForText}, and one cell proves it. Keep the old
* 1000ms write sleep and change only the read — the 800ms fixed sleep becomes a 5s poll — and
* the test goes from 0 of 3 passing to 3 of 3. Removing the write wait instead
* ({@link #SHELL_READY_TIMEOUT_MS} set to 0, so input is typed at once) also passes 3 of 3. So
* the cause is the 800ms READ deadline, not the 1000ms write delay. The old version fails every
* time, not sometimes, so "race" is the wrong word for it. The earlier explanation — that input
* typed before the prompt is swallowed by the shell's startup — is not supported by anything
* measured here. Please do not repeat it: if it were true, typing at 0ms would be worse than
* typing at 1000ms, and it is not.
*
* <p><strong>Where this was measured.</strong> A 12-core macOS host, load average 2.6 to 5.9,
* on commit 20c1094. The same four cells were also run under load, with 8 spinners on 12 cores.
* The failure and the passes from that run are not worth the same. The cells ran in a fixed
* order while the load climbed from 7 to 50. The old version ran last, at the top of that climb,
* so it has a free explanation for failing and its 0 of 3 is discarded. A pass has no such free
* explanation: a cell that survives a worse condition than a fair order would have given it is
* evidence in the safe direction. So keep the three passes, each with the load it ran at: the
* fixed version 3 of 3 at load 7.42 to 18.42, the 0ms-write cell 3 of 3 at 18.42 to 23.65, the
* 1000ms-write cell 3 of 3 at 23.65 to 46.40. Above about load 20 everything here is slow for
* reasons that have nothing to do with this seam, so read the positive claim — that the read
* deadline was the whole cause — as "measured near idle on a 12-core host", and nothing
* stronger. Do not carry that raw load average to another host either: load average counts
* differently per core and per operating system, so only load per core compares. If this test
* fails on a smaller or busier machine, raise {@link #OUTPUT_TIMEOUT_MS} before you suspect the
* seam.
*
* <p>{@link #waitUntilSettled} is kept as cheap insurance against that swallow case, not because
* anyone showed it was needed. The cell that tests the swallow case head-on is the one with
* {@link #SHELL_READY_TIMEOUT_MS} at 0: input is typed at once, which is the worst case for
* "typed before the prompt is ready". It passed 3 of 3 at load 18.42 to 23.65. The wider read
* window cannot explain that pass away, because a swallowed keystroke is LOST, not late — the
* command never runs, so no amount of polling makes its output appear. So the swallow mechanism
* was tested and did not show up. If you want to delete this call, that is the cell to re-run.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@Tag("contract")
class AgentControlContractTest {
private static final long POLL_INTERVAL_MS = 150;
/** Bound for the seed shell to settle: observed ~2.5s three times running; this leaves headroom. */
private static final long SHELL_READY_TIMEOUT_MS = 8_000;
/** Bound for the typed command's output to appear once the shell is ready: observed ~0.2s. */
private static final long OUTPUT_TIMEOUT_MS = 5_000;
private boolean noSocket() {
return !Files.exists(UnixSocketHerdrClient.defaultSocketPath());
}
private static String readPane(UnixSocketHerdrClient herdr, String paneId) {
return herdr.call("pane.read", Map.of("pane_id", paneId, "source", "visible"))
.path("read").path("text").asText("");
}
/**
* Poll {@code pane.read} until two consecutive reads come back identical — the shell's own
* startup output (restore banner, prompt) has stopped changing — or {@code timeoutMs} elapses.
* Never asserts by itself; the caller's own assertion is what actually verifies the outcome.
* Setting {@code timeoutMs} to 0 skips the wait entirely and the test still passes here, so
* treat this as insurance rather than as the fix — see the class javadoc.
*/
private static String waitUntilSettled(UnixSocketHerdrClient herdr, String paneId, long timeoutMs)
throws InterruptedException {
long deadline = System.currentTimeMillis() + timeoutMs;
String previous = null;
while (System.currentTimeMillis() < deadline) {
Thread.sleep(POLL_INTERVAL_MS);
String current = readPane(herdr, paneId);
if (current.equals(previous) && !current.isBlank()) {
return current;
}
previous = current;
}
return previous == null ? "" : previous;
}
/** Poll {@code pane.read} until {@code needle} appears or {@code timeoutMs} elapses. */
private static String waitForText(UnixSocketHerdrClient herdr, String paneId, String needle, long timeoutMs)
throws InterruptedException {
long deadline = System.currentTimeMillis() + timeoutMs;
String last = "";
while (System.currentTimeMillis() < deadline) {
last = readPane(herdr, paneId);
if (last.contains(needle)) {
return last;
}
Thread.sleep(POLL_INTERVAL_MS);
}
return last;
}
@Test
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
@@ -36,15 +126,13 @@ class AgentControlContractTest {
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
try {
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
Thread.sleep(1000); // let the seed shell reach its prompt
waitUntilSettled(herdr, tab.rootPaneId(), SHELL_READY_TIMEOUT_MS);
herdr.call("pane.send_input", Map.of(
"pane_id", tab.rootPaneId(),
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
"keys", List.of("enter")));
Thread.sleep(800);
String visible = herdr.call("pane.read",
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
.path("read").path("text").asText("");
String visible = waitForText(herdr, tab.rootPaneId(),
"PROBE_BASE=[http://gx00.gw:8000]", OUTPUT_TIMEOUT_MS);
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
"env map must reach the seed shell; saw: " + visible);
} finally {
@@ -44,6 +44,7 @@ public final class FakeHerdr implements HerdrClient {
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private final Map<String, String> paneCloseErrorCodeFor = new ConcurrentHashMap<>();
private final Map<String, String> processInfoErrorCodeFor = new ConcurrentHashMap<>();
private String tabCloseErrorCode = null;
private final Map<String, String> tabCloseErrorCodeFor = new ConcurrentHashMap<>();
private String agentSendErrorCode = null;
@@ -144,6 +145,18 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Make {@code pane.process_info} fail with this herdr error code, but only for the given
* {@code pane_id} — every other pane's {@code pane.process_info} still succeeds. Models a
* transient herdr failure partway through a {@link PaneLocator} pid→pane scan (fleetd #505):
* the scan must be able to tell "this pane does not own the pid" apart from "the scan could
* not check this pane at all", instead of collapsing both into one {@code false}.
*/
public FakeHerdr processInfoFailsForPane(String paneId, String code) {
this.processInfoErrorCodeFor.put(paneId, code);
return this;
}
/** Set the {@code agent_status} that {@code agent.get} reports (drives the injector). */
public FakeHerdr agentStatus(String status) {
this.agentStatus = status;
@@ -407,6 +420,13 @@ public final class FakeHerdr implements HerdrClient {
{"pane_id":"w2:p9","terminal_id":"term_shell","workspace_id":"w2","tab_id":"w2:t8"}]}""");
case "pane.process_info" -> {
Object paneId = params instanceof java.util.Map<?, ?> m ? m.get("pane_id") : null;
String failCode = paneId == null ? null
: processInfoErrorCodeFor.get(String.valueOf(paneId));
if (failCode != null) {
throw new HerdrException(
"herdr error [" + failCode + "]: pane.process_info failed",
failCode, null);
}
yield "w2:p7".equals(paneId)
? mapper.readTree(("""
{"type":"pane_process_info","process_info":{"pane_id":"w2:p7","shell_pid":%d,
@@ -14,7 +14,7 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
* Contract test against a REAL running herdr. Tagged {@code contract} so it is
* excluded from {@code mvn test}; run it with {@code mvn test -Pcontract}. It fails
* loudly if herdr drifts from the protocol {@code fleetd} was built against
* (0.7.0, protocol 14) — catching breakage that unit tests with canned frames cannot.
* (0.8.0, protocol 19) — catching breakage that unit tests with canned frames cannot.
*/
@Tag("contract")
class HerdrContractTest {
@@ -24,13 +24,13 @@ class HerdrContractTest {
}
@Test
void pingReturnsProtocol14() {
void pingReturnsProtocol19() {
assumeTrue(Files.exists(socket()), "no herdr socket at " + socket() + " — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
JsonNode pong = herdr.call("ping");
assertEquals("pong", pong.get("type").asText());
assertEquals(14, pong.get("protocol").asInt(),
"fleetd is built against herdr protocol 14");
assertEquals(19, pong.get("protocol").asInt(),
"fleetd is built against herdr protocol 19");
assertFalse(pong.get("version").asText().isBlank());
}
}
@@ -6,6 +6,7 @@ import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
@@ -35,6 +36,12 @@ class LeadTabScannerTest {
final Map<String, String[]> tabs = new LinkedHashMap<>();
/** pane_id → [tab_id, terminal_id]. */
final Map<String, String[]> panes = new LinkedHashMap<>();
/**
* tab_id → whether herdr reports a running agent there. Defaults to {@code true} for every
* tab that has a pane, so every existing fixture keeps meaning "a live lead" unless a test
* says otherwise via {@link #deadAgent}.
*/
final Set<String> deadTabs = new LinkedHashSet<>();
int calls;
boolean failing;
@@ -53,6 +60,18 @@ class LeadTabScannerTest {
return this;
}
/** Mark {@code tabId} as labelled but agent-less — a dead lead's leftover tab (fleetd #359). */
TopologyHerdr deadAgent(String tabId) {
deadTabs.add(tabId);
return this;
}
/** Undo {@link #deadAgent} — models {@code agent.list} reporting the tab live again. */
TopologyHerdr reviveAgent(String tabId) {
deadTabs.remove(tabId);
return this;
}
@Override
public JsonNode call(String method, Object params) {
calls++;
@@ -83,6 +102,19 @@ class LeadTabScannerTest {
.formatted(id, p[0], p[1])));
return read("{\"panes\":[%s]}".formatted(String.join(",", items)));
}
case "agent.list" -> {
// One agent per distinct tab that has a pane and isn't marked dead — mirrors
// AgentControl.list()'s "agents" shape closely enough for the scanner's join,
// which only reads tab_id off each entry.
Set<String> seen = new LinkedHashSet<>();
panes.forEach((paneId, p) -> {
String tabId = p[0];
if (!deadTabs.contains(tabId) && seen.add(tabId)) {
items.add("{\"tab_id\":\"%s\"}".formatted(tabId));
}
});
return read("{\"agents\":[%s]}".formatted(String.join(",", items)));
}
default -> throw new AssertionError("unexpected herdr call: " + method);
}
}
@@ -208,6 +240,114 @@ class LeadTabScannerTest {
scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get().get("term_opus_split"));
}
// ── fleetd #359: a labelled tab is only a lead when something is running in it ─────────────────
/**
* The core invariant this ticket restores: {@code get()} must never report a terminal for a
* lead whose pane no longer runs an agent. Before this fix the scanner joined labelled tabs to
* panes with no liveness check at all, so a tab left behind by a crashed/relaunched lead (see
* {@code LeadLauncher}'s own staleness handling) was reported as live forever — which is exactly
* what let duplicate lead tabs make {@code LeadCoordLoop.resolveLocalLead()} permanently unable
* to pick one. Mutate this away (drop the {@code agent.list} cross-check in {@link
* LeadTabScanner#scan()}) and this test must fail.
*/
@Test
void aLabelledTabWithNoRunningAgentIsNotReported() {
TopologyHerdr herdr = twoLeads().deadAgent("w1:t1"); // opus-5.0's tab is labelled but dead
Map<String, String> leads = scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get();
assertFalse(leads.containsKey("term_opus"),
"a labelled tab with no running agent must never be reported as a live lead");
assertEquals("gpt-sol-5.6", leads.get("term_gpt"),
"the other, genuinely live lead must be unaffected");
}
/**
* The other direction, pinned separately so a fix cannot satisfy the test above by simply
* returning nothing: a labelled tab that DOES have a running agent must still be reported. A
* scanner that always comes back empty is worse than the bug it fixes.
*/
@Test
void aLabelledTabWithARunningAgentIsStillReported() {
Map<String, String> leads = scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong()).get();
assertEquals("opus-5.0", leads.get("term_opus"));
assertEquals("gpt-sol-5.6", leads.get("term_gpt"));
}
/**
* fleetd #359 review, finding 2 — the exact scenario the ticket's own evidence showed:
* {@code agent.list} can come back successfully but short, without the herdr call ever throwing.
* A lead this class already reported as live must not be dropped on the strength of one such
* read: {@link LeadTabScanner#get()}'s "keep the cache on failure" contract only fires on an
* exception, so without a fix a single short {@code agent.list} silently empties the cached map —
* which would resolve that lead's pane as {@code Role.WORKER} downstream, refusing every
* orchestration call. Mutate this away (drop the one-scan grace in {@link
* LeadTabScanner#scan()}) and this test must fail.
*/
@Test
void aTransientAgentListMissDoesNotDemoteALeadAlreadyKnownLive() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertTrue(s.get().containsKey("term_opus"), "opus must be known live before the miss");
// One scan where agent.list comes back without opus's tab, even though the tab and pane are
// completely unchanged — the tab/pane are still there, only the liveness read is short.
herdr.deadAgent("w1:t1");
clock.addAndGet(TTL);
assertTrue(s.get().containsKey("term_opus"),
"a single missed detection must not empty the cached map for a lead already known live");
assertEquals("gpt-sol-5.6", s.get().get("term_gpt"), "the unaffected lead is unchanged");
}
/**
* The other half of finding 2, so the grace above cannot be mistaken for permanent amnesty: the
* original #359 invariant (a genuinely dead tab is not reported forever) must still hold once a
* SECOND, independent scan agrees the agent is gone.
*/
@Test
void aLeadMissingFromAgentListOnTwoConsecutiveScansIsFinallyDropped() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertTrue(s.get().containsKey("term_opus"));
herdr.deadAgent("w1:t1");
clock.addAndGet(TTL);
assertTrue(s.get().containsKey("term_opus"), "first miss is a grace period, not a verdict");
clock.addAndGet(TTL); // a second, independent scan — still no agent
assertFalse(s.get().containsKey("term_opus"),
"a second consecutive miss for the same terminal must finally drop it");
}
/** A lead that recovers between the two misses keeps its grace spent, not renewed for free. */
@Test
void aLeadThatRecoversBetweenMissesIsReportedNormallyAndResetsItsGrace() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertTrue(s.get().containsKey("term_opus"));
herdr.deadAgent("w1:t1");
clock.addAndGet(TTL);
assertTrue(s.get().containsKey("term_opus"), "graced on the first miss");
herdr.reviveAgent("w1:t1"); // the miss really was transient
clock.addAndGet(TTL);
assertTrue(s.get().containsKey("term_opus"), "found live again — reported normally");
// A later, unrelated miss must get its own fresh grace scan rather than being dropped
// immediately because the earlier miss had already "used up" a slot for this terminal.
herdr.deadAgent("w1:t1");
clock.addAndGet(TTL);
assertTrue(s.get().containsKey("term_opus"),
"a miss after a genuine recovery is a new event and deserves its own grace scan");
}
/**
* CB-579 acceptance (6): this is the bug the ticket closes. A stale pin used to be merged back
* over every scan and never expire; now a scan is the whole answer, so a lead whose tab is gone
@@ -37,7 +37,7 @@ class PaneLocatorContractTest {
.path("pane").path("terminal_id").asText(null);
assertNotNull(terminalId, "seed pane should carry a terminal_id");
assertEquals(terminalId, new PaneLocator(herdr).terminalForPid(shellPid),
assertEquals(terminalId, new PaneLocator(herdr).terminalForPid(shellPid).terminal(),
"a real PID must resolve back to its own pane's terminal_id");
} finally {
spaces.closeTab(tab.tab().tabId());
@@ -15,18 +15,20 @@ class PaneLocatorTest {
@Test
void resolvesTerminalForAForegroundPid() {
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID));
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID).terminal());
}
@Test
void nullForAPidInNoPane() {
assertNull(loc.terminalForPid(999_999));
PaneLocator.Lookup outcome = loc.terminalForPid(999_999);
assertNull(outcome.terminal());
assertTrue(outcome.complete(), "a full, error-free scan that finds no match is complete");
}
@Test
void nullForNonPositivePid() {
assertNull(loc.terminalForPid(0));
assertNull(loc.terminalForPid(-1));
assertNull(loc.terminalForPid(0).terminal());
assertNull(loc.terminalForPid(-1).terminal());
}
// --- two-daemon fallback (CB-185) -----------------------------------------
@@ -38,7 +40,7 @@ class PaneLocatorTest {
HerdrClient lead = new FakeHerdr().withNoPanes();
HerdrClient member = new FakeHerdr();
PaneLocator two = new PaneLocator(lead, member);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID).terminal());
}
@Test
@@ -48,13 +50,13 @@ class PaneLocatorTest {
HerdrClient lead = new FakeHerdr();
HerdrClient member = new FakeHerdr().withNoPanes();
PaneLocator two = new PaneLocator(lead, member);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID).terminal());
}
@Test
void nullWhenNeitherClientHasTheMatch() {
PaneLocator two = new PaneLocator(new FakeHerdr().withNoPanes(), new FakeHerdr().withNoPanes());
assertNull(two.terminalForPid(FakeHerdr.WORKER_PID));
assertNull(two.terminalForPid(FakeHerdr.WORKER_PID).terminal());
}
@Test
@@ -63,7 +65,7 @@ class PaneLocatorTest {
// must behave exactly like the one-arg constructor, including making only one herdr call.
FakeHerdr shared = new FakeHerdr();
PaneLocator two = new PaneLocator(shared, shared);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID).terminal());
long paneListCalls = shared.calls.stream().filter(c -> c.method().equals("pane.list")).count();
assertEquals(1, paneListCalls, "same-object lead/member must scan exactly once, not twice");
}
@@ -75,7 +77,7 @@ class PaneLocatorTest {
// Regression: a pid with no parent chain at all — no ancestry walk is needed to match it.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
PaneLocator loc = new PaneLocator(pane, new FakeParentResolver());
assertEquals("term_x", loc.terminalForPid(5000));
assertEquals("term_x", loc.terminalForPid(5000).terminal());
}
@Test
@@ -83,7 +85,7 @@ class PaneLocatorTest {
// Regression: same as above, but matching via the foreground-processes list.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
PaneLocator loc = new PaneLocator(pane, new FakeParentResolver());
assertEquals("term_x", loc.terminalForPid(6000));
assertEquals("term_x", loc.terminalForPid(6000).terminal());
}
@Test
@@ -97,7 +99,7 @@ class PaneLocatorTest {
.parent(7002, 7001) // grandchild -> child
.parent(7001, 5000); // child -> shell (the pane's shell_pid)
PaneLocator loc = new PaneLocator(pane, parents);
assertEquals("term_x", loc.terminalForPid(7002));
assertEquals("term_x", loc.terminalForPid(7002).terminal());
}
@Test
@@ -110,7 +112,7 @@ class PaneLocatorTest {
.parent(9002, 9001)
.parent(9001, 9000); // chain never reaches 5000 or 6000
PaneLocator loc = new PaneLocator(pane, parents);
assertNull(loc.terminalForPid(9002));
assertNull(loc.terminalForPid(9002).terminal());
}
@Test
@@ -122,7 +124,7 @@ class PaneLocatorTest {
.parent(100, 101)
.parent(101, 100); // cycle, never reaches the pane's pids
PaneLocator loc = new PaneLocator(pane, parents);
assertNull(loc.terminalForPid(100));
assertNull(loc.terminalForPid(100).terminal());
}
@Test
@@ -141,11 +143,43 @@ class PaneLocatorTest {
};
HerdrClient noPanes = new FakeHerdr().withNoPanes();
PaneLocator two = new PaneLocator(noPanes, pane, counting);
assertEquals("term_x", two.terminalForPid(7002));
assertEquals("term_x", two.terminalForPid(7002).terminal());
assertEquals(3, calls.get(), "ancestry must be walked once (3 lookups: 7002, 7001, 5000), "
+ "not re-walked per herdr client");
}
// --- fleetd #505: a herdr error during the scan must not read as a clean negative ---------
@Test
void anErrorOnThePaneThatOwnsThePidMakesTheScanIncompleteNotAClearNegative() {
// The discriminating case: pane.process_info fails for exactly the pane that DOES own the
// caller's pid ("w2:p7", term_a). Before the fix, that failure was swallowed into a plain
// "does not own it" and the scan finished with a clean-looking null — indistinguishable
// from a real primary. It must now report incomplete, not a definite null.
FakeHerdr herdr = new FakeHerdr().processInfoFailsForPane("w2:p7", "transient");
PaneLocator loc = new PaneLocator(herdr);
PaneLocator.Lookup outcome = loc.terminalForPid(FakeHerdr.WORKER_PID);
assertNull(outcome.terminal(), "the failing pane's ownership could not be confirmed");
assertFalse(outcome.complete(),
"a scan that could not check the owning pane must not report as complete");
}
@Test
void aVanishedPaneThatIsNotTheMatchLeavesAnOtherwiseSuccessfulScanComplete() {
// The companion invariant: a DIFFERENT pane (not the caller's own) failing mid-scan must
// not turn every mid-scan teardown into a refusal — the real match is still found, and the
// scan is still reported complete.
FakeHerdr herdr = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
PaneLocator loc = new PaneLocator(herdr);
PaneLocator.Lookup outcome = loc.terminalForPid(FakeHerdr.WORKER_PID);
assertEquals("term_a", outcome.terminal());
assertTrue(outcome.complete(), "a positive match elsewhere in the scan is definitive");
}
/** Minimal single-pane {@link HerdrClient} fake, purpose-built for the ancestry tests above. */
private static final class OnePaneHerdr implements HerdrClient {
private final ObjectMapper mapper = new ObjectMapper();
@@ -20,6 +20,7 @@ import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
@@ -357,6 +358,54 @@ class CompletionResolverTest {
"the failure carries whatever was on screen: " + waiter.getNow(null).text());
}
@Test
void aPlausibleLookingReplyInsideTheFloorStillFails() {
// fleetd#376 guard. A fix was attempted that inspected the pane inside the floor and resolved
// a COMPLETION when the text "looked like a real reply". Every cheap test for that is unsafe:
// lastAssistantBlock falls back to the WHOLE pane when there is no ⏺ marker, and a crash pane
// almost always contains sentence punctuation — in a file path, a version, or a hostname.
// This pane is the trap: it reads like a finished answer and it is a backend failure.
FakeHerdr herdr = new FakeHerdr().readText(
"Error: connection reset while loading src/main/java/Foo.java v1.2.3\n❯ ");
Rendezvous rendezvous = new Rendezvous();
long[] clock = {10_000_000_000L};
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), () -> clock[0]);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null, clock[0]);
clock[0] += CompletionResolver.MIN_TURN_NANOS - 1; // inside the floor
resolver.resolve("term_a", turn);
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"inside the floor the verdict is always FAILED — never guess a completion from pane text");
}
@Test
void theTooFastFailureDoesNotAssertACauseItCannotKnow() {
// fleetd#376: the message used to say "most likely a backend error before any work started".
// When no error pattern matches, that cause is a guess, and a reader who believes it stops
// looking at the pane. The verdict stays FAILED; only the claim about WHY is withdrawn.
FakeHerdr herdr = new FakeHerdr().readText("I am running on opencode/mimo-v2.5-free.\n❯ ");
Rendezvous rendezvous = new Rendezvous();
long[] clock = {10_000_000_000L};
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), () -> clock[0]);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null, clock[0]);
clock[0] += CompletionResolver.MIN_TURN_NANOS - 1; // inside the floor
resolver.resolve("term_a", turn);
String text = waiter.getNow(null).text();
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(), "still fails, still loud");
assertFalse(text.contains("most likely a backend error"),
"an unmatched fast turn must not assert a backend error: " + text);
assertTrue(text.contains("mimo-v2.5-free"), "the pane is still carried: " + text);
}
@Test
void aBusyToDoneTransitionJustOutsideTheFloorResolvesNormally() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ a real, if quick, answer\n❯ ");
@@ -484,6 +533,27 @@ class CompletionResolverTest {
// --- CB-578 stage A: backend-exhausted classification ---------------------------------
@Test
void aNormalMemberReportMentioningTheExhaustionPatternDoesNotNotifyTheSink() {
String block = "⏺ I reviewed capacity handling. The usage limit has been reached means no more work can start.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"a matching report still fails the send as exhausted");
assertTrue(waiter.getNow(null).text().contains("I reviewed capacity handling."),
"the exhausted result keeps the whole matched pane line");
assertTrue(notified.isEmpty(),
"a normal report mentioning an exhaustion pattern must not quarantine a credential");
}
@Test
void classifiesAMatchingScrapeAsBackendExhaustedInsteadOfACompletedReply() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
@@ -534,6 +604,53 @@ class CompletionResolverTest {
"the sink is told the matched reason: " + notified.get(0));
}
@Test
void aRealExhaustionBehindTerminalChromeStillNotifiesTheSink() {
String block = "⏺ │ The usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"a real exhaustion must still fail the send as exhausted");
assertEquals(1, notified.size(),
"a real exhaustion behind terminal chrome must reach the sink");
}
/**
* The live fleet configures {@code exhaustedPattern: "The usage limit has been reached"} — with
* the leading {@code "The"}. Every other test here uses a pattern without it, which is the shape
* that made fleetd #348 need a looser rule than a start-of-line check. This pins the deployed
* shape as well, so a later tightening of {@link CompletionResolver} cannot silently stop
* recording the exhaustion this fleet actually reports.
*
* <p>What it does not prove: that this is the only pattern shape an operator will write.
*/
@Test
void anExhaustionPatternCarryingItsLeadingWordsStillNotifiesTheSink() {
String block = "⏺ │ The usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("The usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"the live pattern shape must still fail the send as exhausted");
assertEquals(1, notified.size(),
"the live pattern shape must still reach the sink");
}
@Test
void aLosingBackendExhaustedClassificationNeverNotifiesTheExhaustionSink() {
// The waiter was already resolved (e.g. by the worker's own reply) before this scrape landed —
@@ -768,19 +885,22 @@ class CompletionResolverTest {
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
CompletionResolver.coverage("exhaustedPattern", Set.of("terra"), Set.of()));
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra"), Set.of()));
}
@Test
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
assertEquals("full (all profiles configured: [gx10, terra])",
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra", "gx10")));
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra", "gx10"), Set.of("terra", "gx10")));
}
@Test
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
assertEquals("partial (configured: [terra]; not configured: [gx10])",
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra")));
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra", "gx10"), Set.of("terra")));
}
/**
@@ -794,11 +914,42 @@ class CompletionResolverTest {
*
* <p>Every earlier test here passed the exhaustion case only, so none of them could see it. This
* one pins that the message names the key the caller actually meant.
*
* <p>fleetd#415: the expected wording changed here too. {@code errorPattern} has a built-in
* fallback ({@link CompletionResolver#BACKEND_ERROR}), so an empty {@code configuredProfiles}
* for it is not "off" — see {@link #coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput}
* for the test built specifically to pin that distinction.
*/
@Test
void coverageNamesTheConfigKeyItsCallerMeansRatherThanAlwaysSayingExhaustedPattern() {
assertEquals("off (no profile has an errorPattern configured; profiles: [gx10, terra])",
CompletionResolver.coverage("errorPattern", Set.of("terra", "gx10"), Set.of()));
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])",
CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
Set.of("terra", "gx10"), Set.of()));
}
/**
* fleetd#415: {@code coverage()} measures pattern coverage (how many profiles set the key), but
* for {@code errorPattern} the empty case is not the feature-off state — a profile with no
* configured {@code errorPattern} still runs the classification against
* {@link CompletionResolver#BACKEND_ERROR}. For {@code exhaustedPattern} there is no fallback,
* so empty really is off. Same shape of input (empty {@code configuredProfiles}, one profile),
* different {@link CompletionResolver.UnsetMeaning} — the wording must differ, or this method is
* back to conflating pattern coverage with feature state for the one key where they disagree.
*/
@Test
void coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput() {
String exhaustedLine = CompletionResolver.coverage("exhaustedPattern",
CompletionResolver.UnsetMeaning.OFF, Set.of("gx10", "terra"), Set.of());
String errorLine = CompletionResolver.coverage("errorPattern",
CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT, Set.of("gx10", "terra"), Set.of());
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
exhaustedLine);
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])", errorLine);
assertNotEquals(exhaustedLine, errorLine,
"the same empty-coverage input must not read as the same feature state for both keys");
}
// --- fleetd#201 Unit 1: target-keyed backend-error pattern + typed sink ----------------------
@@ -1049,8 +1200,16 @@ class CompletionResolverTest {
assertTrue(waiter.isDone());
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"still a failure — the floor itself, not the pattern, is why");
assertTrue(waiter.getNow(null).text().contains("too fast to be real work"),
"a non-match inside the floor stays the generic too-fast reason: " + waiter.getNow(null).text());
// fleetd#376: this used to assert the phrase "too fast to be real work", which carried the
// claim "most likely a backend error before any work started". With no pattern matched that
// cause is a guess, so the wording was withdrawn. What this test really guards is unchanged:
// the floor alone still fails the turn, it stays generic, and it never notifies the sink.
String reason = waiter.getNow(null).text();
assertTrue(reason.contains("inside the floor"),
"a non-match inside the floor stays the generic floor reason: " + reason);
assertFalse(reason.contains("most likely a backend error"),
"a non-match must not assert a cause it did not establish: " + reason);
assertTrue(reason.contains("still starting up"), "the pane is still carried: " + reason);
assertTrue(notified.isEmpty(), "a non-match must never notify the typed sink");
}
}
@@ -19,6 +19,8 @@ import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.*;
@@ -521,6 +523,97 @@ class InjectorTest {
}
}
@Test
void readinessGraceExpiryLogsTheMeasuredPollCountNextToTheConfiguredBudget() {
// fleetd #501, defect 1: READINESS_GRACE_POLLS (240) used to be printed twice — once as
// "after {} polls" and once inside the parenthesised budget — even though the loop's own
// counter (Target.notReadySincePoll) was in scope at the same call site. On THIS branch the
// counter has just reached the threshold, so it equals the constant by construction and this
// test cannot tell the two apart — it only pins that the message still carries both a poll
// count and a labelled configured budget, using literal numbers (240, 60), never
// READINESS_GRACE_POLLS or POLL_INTERVAL_MILLIS, so the assertion can't silently track a
// constant change instead of catching a real regression.
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
});
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no grace-expiry WARN logged");
assertTrue(warn.contains("after 240 polls"),
"must print the measured poll count as a plain number: " + warn);
assertTrue(warn.contains("configured=240 polls/60s"),
"must print the configured budget, clearly labelled: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
@Test
void readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants() {
// fleetd #501, defect 2: the old line computed "({}s)" as READINESS_GRACE_POLLS *
// POLL_INTERVAL_MILLIS / 1000 — arithmetic on two constants, never a measurement, and wrong
// in the direction that says everything ran on schedule. This stub clock returns two FIXED
// values (1_000ms at the first non-ready sample, 318_412ms at the poll that trips the grace)
// whose difference — 317_412ms — does NOT equal 240 * POLL_INTERVAL_MILLIS (=60_000ms).
// Asserting on that literal, non-derived number is what makes this test able to fail if the
// production code goes back to printing the constant-arithmetic value instead of the
// injected clock's measurement.
long[] readings = {1_000L, 318_412L};
AtomicInteger call = new AtomicInteger(0);
LongSupplier stubClock = () -> {
int i = call.getAndIncrement();
if (i >= readings.length) {
throw new AssertionError("nowMillis read more times than this fixture expects (" + i
+ "); the readiness-not-ready branch should read the clock exactly twice — "
+ "once to stamp the first non-ready sample, once at grace expiry");
}
return readings[i];
};
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
}, stubClock);
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no grace-expiry WARN logged");
assertTrue(warn.contains("elapsed=317412ms"), "must print the MEASURED elapsed time from "
+ "the injected clock (318412 - 1000 = 317412), not an arithmetic value: " + warn);
assertFalse(warn.contains("elapsed=60000ms"), "must not print "
+ "READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS (240 * 250 = 60000ms) as if it "
+ "were the measured elapsed time: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
@Test
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
@@ -0,0 +1,138 @@
package dev.ltms.fleet.inject;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotSame;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #446: {@code LiveExhaustedPatterns} is the mechanism that makes {@code exhaustedPattern}
* hot — read fresh off a config supplier per lookup, cached by profile name. This is the direct,
* minimal unit test of that mechanism: read the pattern once, mutate the backing "config", read
* again, and assert the second read saw the new value — the exact hotness proof shape.
*
* <p>This class alone does not prove {@code Fleetd.main} actually wires the live source into
* production — {@link dev.ltms.fleet.FleetdExhaustionDetectionArmedWiringTest} and {@code
* CompletionResolverTest}'s {@code ExhaustedPatternLookup} tests do that at the wiring level. What
* this class proves is that the mechanism itself is genuinely live and genuinely cached, not that
* it is used.
*/
class LiveExhaustedPatternsTest {
private static FleetConfig.Profile profileWithPattern(String pattern) {
return new FleetConfig.Profile("terra", "http://gx00.gw:8000", "terra",
null, null, null, "tab", "fleetd-workers", "w #{n}", null, null, null,
null, null, null, null, null, null, null, pattern);
}
/** Simple mutable holder standing in for {@code ConfigRef} — swapped, never mutated in place. */
private static final class MutableProfiles {
private volatile Map<String, FleetConfig.Profile> profiles;
MutableProfiles(Map<String, FleetConfig.Profile> initial) {
this.profiles = initial;
}
Map<String, FleetConfig.Profile> get() {
return profiles;
}
void set(Map<String, FleetConfig.Profile> fresh) {
this.profiles = fresh;
}
}
@Test
@DisplayName("patternFor reads the live config: read, change, read again, see the change — no restart")
void patternForIsHotAcrossAConfigChange() {
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
Pattern before = patterns.patternFor("terra");
assertTrue(before.matcher("your usage limit has been reached").find());
// Change the "config" — the exact thing a reload does to ConfigRef in production.
live.set(Map.of("terra", profileWithPattern("rate limit exceeded")));
Pattern after = patterns.patternFor("terra");
assertTrue(after.matcher("429: rate limit exceeded, try later").find(),
"the second read must see the NEW pattern text with no restart");
assertFalse(after.matcher("your usage limit has been reached").find(),
"the second read must have actually recompiled — not just reused a stale match");
}
@Test
@DisplayName("armed follows the same live read: on, then off, after a config change — no restart")
void armedIsHotAcrossAConfigChange() {
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
assertTrue(patterns.armed("terra"));
live.set(Map.of("terra", profileWithPattern(null)));
assertFalse(patterns.armed("terra"),
"removing exhaustedPattern and 'reloading' must disarm detection with no restart");
}
@Test
@DisplayName("an unchanged pattern string is not recompiled — the cache is reused")
void unchangedPatternTextReusesTheCompiledInstance() {
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
Pattern first = patterns.patternFor("terra");
// A NEW Profile object, same pattern text — a reload of an unrelated key produces a fresh
// FleetConfig even though this profile's own text did not change.
live.set(Map.of("terra", profileWithPattern("usage limit")));
Pattern second = patterns.patternFor("terra");
assertSame(first, second,
"identical pattern text must reuse the cached compiled Pattern, not recompile it — "
+ "recompiling a regex on every check is exactly the cost the cache exists to avoid");
}
@Test
@DisplayName("a changed pattern string is recompiled, not silently reused from the cache")
void changedPatternTextIsRecompiled() {
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
Pattern first = patterns.patternFor("terra");
live.set(Map.of("terra", profileWithPattern("rate limit")));
Pattern second = patterns.patternFor("terra");
assertNotSame(first, second, "changed pattern text must produce a freshly compiled Pattern");
}
@Test
void unknownProfileReturnsNullAndUnarmed() {
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(Map::of);
assertNull(patterns.patternFor("nope"));
assertFalse(patterns.armed("nope"));
}
@Test
void nullProfileNameReturnsNullAndUnarmed() {
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(
() -> Map.of("terra", profileWithPattern("usage limit")));
assertNull(patterns.patternFor(null));
assertFalse(patterns.armed(null));
}
@Test
void profileWithNoPatternConfiguredReturnsNull() {
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(
() -> Map.of("terra", profileWithPattern(null)));
assertNull(patterns.patternFor("terra"));
assertFalse(patterns.armed("terra"));
}
}

Some files were not shown because too many files have changed in this diff Show More