Compare commits

..

47 Commits

Author SHA1 Message Date
ltms 26f198675c Merge pull request 'fleetd #612 Unit A: extract main's boot composition into FleetdAssembly/FleetdRuntime' (#620) from worker/fleetd-612-unita-87807e-1 into main
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 2m30s
2026-09-22 07:43:39 +02:00
lead 72f46d7c0c fleetd #612: merge main into Unit A, carrying #621's requireOperatorConfirm into the assembly
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 2m59s
Resolves the one conflict in Fleetd.java. main's side is the inline boot block
that Unit A had already moved into FleetdAssembly.assembleAndStart, so the
resolution keeps Unit A's single assembly call.

That resolution is not purely mechanical. #622 (fleetd #621) landed on main
AFTER Unit A forked, and it added

    boolean requireOperatorConfirm = cfg.leadRollover() == null
            || cfg.leadRollover().requireOperatorConfirm();

plus a 14th argument to the LeadHeartbeatLoop constructor, inside the very block
Unit A moved. Taking Unit A's side alone would have dropped both and silently
reverted the operator's #621 fix: the 13-argument overload still exists and
delegates with `true`, so the daemon would go back to telling every lead to ask
the operator before a context roll. Both are carried into FleetdAssembly here.

Measured: with the carried line removed, the full suite is
`Tests run: 1883, Failures: 0, Errors: 0` — nothing pins it. That is the fleetd
#612 defect shape applied to #621's own wiring, and it is filed separately
rather than fixed here, because this commit is a merge resolution and must not
also introduce new tests.

Full suite on this resolved tree: Tests run: 1883, Failures: 0, Errors: 0.
2026-09-22 12:43:22 +07:00
ltms 640f4d5f23 Merge pull request 'fleetd #612 B3: behavioural replacements for lead-seat, quarantine, lead-rollover guards' (#628) from worker/612-b3-mcpwirings-da2b58-3 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 12s
CI / contract (pull_request) Successful in 1m33s
CI / build (pull_request) Failing after 2m25s
2026-09-22 07:36:05 +02:00
Dai Ha 6edeb70bc4 fleetd #612 B3 correction: distinguish lead vs member herdr in FleetdLeadRolloverAssemblyTest
Ticket comment 17553 on fleetd #612 found that the test's single shared
FakeHerdr made router.leadAgents() and router.memberAgents() collapse to
the identical client (FleetdAssembly.java:140-142's no-distinct-socket
fallback), so a mutation swapping leadAgents() for memberAgents() at the
FleetdAssembly.java:408 call site was invisible to this test even though
the two are genuinely different daemons in production.

Configure two distinct herdr sockets and two distinct FakeHerdr instances
(the same TwoHerdrResourcePorts shape B2's FleetdAssemblyConnectionIdentityTest
uses) and assert the roll's /clear + bootstrap sends land on the LEAD fake
and never on the MEMBER one.

Proven red against the router.memberAgents() mutation, reverted, touched,
and re-run green — both outputs recorded in the PR.
2026-09-22 12:31:39 +07:00
Dai Ha bc49d87cb8 Merge remote-tracking branch 'origin/worker/fleetd-612-unita-87807e-1' into worker/612-b3-mcpwirings-da2b58-3 2026-09-22 12:27:47 +07:00
ltms 14b169c410 Merge pull request 'fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair' (#626) from worker/612-b2-cb185-176d3a-2 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 2m6s
2026-09-22 07:19:28 +02:00
Dai Ha 2b52324d9a fleetd #612 step 2 unit B3: behavioural replacements for lead-seat, quarantine
and lead-rollover source-text guards

Replaces three FleetdAssembly.java source-text guards (each scraped
Fleetd.java for a call site that fleetd #612 Unit A moved into
FleetdAssembly.java) with tests that drive the real assembled objects
through FleetdAssembly.assembleAndStart(...) -> FleetdRuntime.mcp(),
per the step-2 B-unit split (issue #612 comment 17513).

- Deleted FleetdLeadSeatWiringTest (fleetd #176): pinned that
  FleetMcp's LeadSeatSource construction still wires
  Fleetd.leadSeatLookup(...) by scraping the constructor call's text.
  Replaced by FleetdLeadSeatAssemblyTest, which seeds one FakeHerdr
  tab labelled to match a configured fleet.leaders.opus.tab and
  asserts the REAL assembled LeadSeatSource (via
  runtime.mcp().leadSeatSource()) reports the live lead's seat against
  its own subscription profile -- 1, not the 0 LeadSeatSource.none()
  (the inert stand-in) could ever report.

- Deleted FleetdBackendQuarantineWiringTest (fleetd #466): pinned that
  the escalating BackendQuarantine.withEscalation(...) text was
  present and the flat two-argument constructor's text was absent.
  Replaced by FleetdBackendQuarantineAssemblyTest, which quarantines
  the same credential twice through the REAL assembled
  BackendQuarantine (via runtime.mcp().quarantineSource().quarantine())
  at controlled fake-clock offsets and asserts the second cooldown
  doubles (200s vs 100s) -- the one behavioural difference escalation
  and the flat constructor actually produce.

- Deleted FleetdLeadRolloverWiringTest (fleetd #480), all three
  methods: unrelatedAnchorStillPresent was a scaffold anchor with no
  independent claim, needing no replacement.
  mainStillCallsTheLeadRolloverFactory pinned the leadRollover
  assignment's call-site text. factoryGatesOnConfigPresence pinned
  that an absent leadRollover: config yields no LeadRollover.
  Replaced by FleetdLeadRolloverAssemblyTest's two tests:
  assembledLeadRolloverRunsTheRealClearAndBootstrapSequence drives the
  REAL assembled LeadRollover (via runtime.mcp().leadRollover())
  through open()/confirm() end to end and asserts /clear then
  bootstrapText were actually sent through the real herdr router,
  reaching ROLLED. absentLeadRolloverConfigMeansNoRolloverIsBuilt
  calls Fleetd.leadRollover(...) directly with no leadRollover: block
  and asserts null -- this claim was found uncovered elsewhere
  (LeadRolloverTest's only related assertion is vacuous, assertNull
  (null), and never calls the real factory).

Each of the three FleetdAssembly.java call sites (quarantine
line 179-180, leadRollover line 408, lead seats line 479) was mutated
to its named inert variant, run against ONLY its new test (RED),
reverted, touch'd (Maven mtime trap) and re-run (GREEN) -- six proven
runs, pasted in the PR body.

FleetMcp.java: adds three accessors (quarantineSource(),
leadSeatSource(), leadRollover()) alongside the existing
registeredTools() -- but public, not package-private, and this is a
deliberate deviation from that precedent, not an oversight: these new
assembly tests cannot live in package dev.ltms.fleet.mcp the way
registeredTools()'s callers do, because they also build the
ResourcePorts FleetdAssembly.assembleAndStart(...) needs, and
ResourcePorts' methods return Fleetd-nested types visible only from
package dev.ltms.fleet. Package-private would compile but be
unreachable from there.

Full mvn -o test in fleetd/: Tests run: 1879, Failures: 6 (down from
the branch baseline's 1880/9 by exactly the 3 guards this unit
deletes) -- the remaining 6 are FleetdCompletionResolverWiringTest (4)
and FleetdConnectionIdentityConstructionTest /
FleetdFleetAppConstructionTest (1 each), all out of this unit's scope
(B1/B2).
2026-09-22 12:18:48 +07:00
ltms b6147a39f6 Merge pull request 'fleetd #612 Unit B1: CompletionResolver behavioural test (replaces source-text guard)' (#627) from worker/612-b1-completion-457459-1 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 1m37s
2026-09-22 07:14:25 +02:00
Dai Ha dbfc34cb6d fleetd #612 B2 fixup: cover the symmetric FleetApp daemon-drop mutation
Per ticket comment 17525: my second independent mutation on
FleetdAssembly.java:518 survived — new FleetApp(memberHerdr, memberHerdr, ...)
(dropping the LEAD client instead of the member one) left both existing
FleetdAssemblyFleetAppTest cases green. That is the symmetric form of the
CB-185 defect (/healthz green while the LEAD daemon is down), and the guard
this PR deletes would have caught it: its positive assertion required the
exact pair "new FleetApp(herdr, memberHerdr, workers,", which does not
survive either daemon being dropped.

Adds healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp, symmetric
to the existing member-down case.

Proven red today: reverting FleetdAssembly.java:518 to
"new FleetApp(memberHerdr, memberHerdr, workers, ..." and running only
FleetdAssemblyFleetAppTest gives Tests run: 3, Failures: 1 — the new case
fails ("expected: <503> but was: <200>", body has no "member" key); the other
two cases stay green. Reverted, git diff --stat empty, file touched, re-ran:
Tests run: 3, Failures: 0.

Full mvn -o test: Tests run: 1884, Failures: 7 (same 7 B1/B3-scope failures
as before this fixup; 1884 = 1883 + 1 new case).
2026-09-22 12:11:28 +07:00
Dai Ha a20cb96730 fleetd #612 Unit B1: replace CompletionResolver source-text guards with behavioural tests
FleetdCompletionResolverWiringTest read Fleetd.java's literal source text and
asserted the CompletionResolver construction call still named the right
arguments — proof of spelling, not behaviour. Unit A (FleetdAssembly) moved
that call site out of Fleetd.main, breaking all four of its tests on a
harmless relocation.

FleetdCompletionResolverAssemblyTest replaces it, driving the real
FleetdAssembly.assembleAndStart(...) and reading FleetdRuntime.completion() —
the exact CompletionResolver instance production uses — through its public
onDelivered/resolveBeforePostAction API, with a controllable ResourcePorts
nanoClock in place of real sleeps.

Deleted test -> what it pinned -> replacement:
- worktreeBranchLookupIsStillPassedAtTheCallSite (8th constructor arg) ->
  assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport: a
  real git-worktree-provisioned MemberSession's branch must appear in a
  noReportMessage fallback (fleetd #241).
- backendErrorArgumentsAreStillNamedAtTheCallSite,
  backendErrorPatternsComesFromTheFactory, backendErrorSinkComesFromTheFactory
  (5th/6th args) -> assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern:
  a configured errorPattern the built-in fallback never matches must classify
  as FAILED (not COMPLETION), transition the session to BACKEND_ERROR, and
  cool the credential off after two distinct targets within the window
  (fleetd #201 Unit 5).

Both behaviours were proven red today by mutating FleetdAssembly.java's real
call site to its inert variant (_ -> null; BackendErrorPatternLookup.legacy();
BackendErrorSink.none()), confirming the new test failed with the expected
message, then reverting (touching the file to defeat Maven's stale-mtime
skip) and confirming green again. FleetdAssembly.java itself is unchanged in
this commit.

Full suite: Tests run: 1878, Failures: 5 (down from the baseline 9 — the
remaining 5 are the other in-flight workers' own guard files:
FleetdBackendQuarantineWiringTest, FleetdConnectionIdentityConstructionTest,
FleetdFleetAppConstructionTest, FleetdLeadRolloverWiringTest,
FleetdLeadSeatWiringTest), Errors: 0, Skipped: 0.
2026-09-22 12:10:38 +07:00
Dai Ha cda1a6a917 fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair
Deletes the two source-text guards fleetd #612 Unit A broke by moving their
scraped call sites from Fleetd.java into FleetdAssembly.java, replacing each
with a behavioural test that drives the real assembled graph instead.

- FleetdConnectionIdentityConstructionTest pinned that Fleetd.java contained
  "new PaneLocator(herdr, memberHerdr)". Replaced by
  FleetdAssemblyConnectionIdentityTest, which drives the real PaneLocator a
  real FleetdAssembly.assembleAndStart(...) built (reached via
  runtime.mcp().identity().panes(), never a copy) with two distinct FakeHerdr
  daemons, and proves it finds a pane that exists on only one of them —
  first the lead-only case (the CB-185 bug: a lead's own connection going
  unresolvable), then the member-only case, plus a no-match control.

- FleetdFleetAppConstructionTest pinned that Fleetd.java contained
  "new FleetApp(herdr, memberHerdr, workers,". Replaced by
  FleetdAssemblyFleetAppTest, which binds the real Javalin app
  FleetdAssembly built (runtime.app()) to a real ephemeral port and proves
  GET /healthz goes 503 when only the member daemon is down — the consequence
  named in the deleted test's javadoc (a down member daemon invisible behind
  a healthy lead). The GET /sessions merge half of that javadoc could not be
  driven the same way: it requires Authz.Action.READ, which needs a real
  positive pid from the real assembly's hardcoded LsofPeerPidLookup, and an
  in-process test's HTTP client and server share one JVM pid so that pid is
  always -1 (fleetd #317's fail-closed rule then refuses the request before
  the route, and its merge, is ever reached) — documented in the new test's
  javadoc; FleetAppTwoDaemonTest remains the full behavioural proof that
  FleetApp itself merges /sessions correctly given two clients.

Both replacements were proven red today: reverting FleetdAssembly.java:443-444
to "new PaneLocator(memberHerdr)" turns the identity test red (expected
"term_a", got null); reverting :518 to "new FleetApp(herdr, herdr, ..." turns
the app test red (expected 503, got 200 with no "member" key). Both mutations
reverted, tree confirmed clean, and the mutated file touched afterward so
Maven does not skip recompiling a stale .class.

Adds two small production accessors needed to reach the real objects rather
than a copy, since FleetdRuntime may not gain a field (three workers touch
that file): ConnectionIdentity#panes() exposes the PaneLocator it resolves
against, and FleetMcp#identity() exposes the ConnectionIdentity it was built
with (now also stored as a field).

Baseline on 608e449: 1880 tests, 9 failures (the known source-text set).
After this change: 1883 tests, 7 failures — the remaining #248/#176/#466/#480
guards, which are B1's and B3's scope, not this one's.
2026-09-22 12:03:38 +07:00
ltms 608e4496be Merge pull request 'fleetd #612 A-gaps: coordinator path + reportRoleFallbackGaps inside the assembly boundary' (#624) from worker/612-agaps-73a926-2 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m34s
CI / build (pull_request) Failing after 1m42s
2026-09-22 06:41:42 +02:00
ltms fa61dc587c Merge pull request '#608 replace MessageService timing sleeps' (#623) from worker/608-sleeps-3a64ff-3 into main
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 3h14m55s
2026-09-22 06:38:29 +02:00
Dai Ha 526e3b7459 fleetd #612 A-gaps: exercise the coordinator path and pull reportRoleFallbackGaps inside the assembly boundary
Gap 1: generalise Fleetd.LeadMailboxOpener (and FleetdRuntime's field) from the concrete
LeadMailbox to a new closeable LeadChannelHandle (LeadChannel + AutoCloseable), so a test can
fake the configured-coordinator path without a real broker. FleetdAssemblyCoordinatorLifecycleTest
drives FleetdAssembly.assembleAndStart with a coordinator: block and a fake channel, proving the
assembly builds it and shutdown closes it.

Gap 2: move reportRoleFallbackGaps(cfg) and assertChartersNameOnlyRegisteredTools(cfg) out of
Fleetd.main and into FleetdAssembly.assembleAndStart, immediately after cfg.validateAll(), so
both run inside the tested assembly boundary before any I/O. FleetdAssemblyRoleFallbackBoundaryTest
pins the moved call site behaviourally (captured log output), not by reading source text.

Both new tests were verified red under a targeted mutation (a non-closing/null coordinator for
gap 1; deleting the moved call for gap 2) and restored.
2026-09-22 11:38:19 +07:00
Dai Ha b6b006c651 #608 replace MessageService timing sleeps
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Failing after 2m27s
2026-09-22 11:33:48 +07:00
ltms 63eec8a0da Merge pull request 'fleetd #621: make the context-roll notice obey requireOperatorConfirm' (#622) from worker/621-b4520b-1 into main
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m11s
CI / build (push) Failing after 1m52s
2026-09-22 06:31:02 +02:00
Dai Ha cbe872b538 fleetd #621: make the context-roll notice obey requireOperatorConfirm
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m0s
CI / build (pull_request) Failing after 1m53s
contextNotice hardcoded 'ask the operator' and 'Only the operator can
approve the roll', so setting leadRollover.requireOperatorConfirm to
false stopped the daemon refusing the roll but never stopped the lead
being told to ask. Thread the effective config value into
contextNotice: when true the text stays byte-identical, when false it
tells the lead to confirm on its own judgement against the three
handover-file checks instead.

LeadRollover.confirm's own enforcement is untouched — this is the
message only.
2026-09-22 11:27:17 +07:00
Dai Ha 7d9a807243 handover skill: the rollover bootstrap is proven, and requireOperatorConfirm is per-host
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 2m33s
Two bullets in the handover skill were telling every outgoing lead something
that is no longer true.

1. The skill said the bootstrap prompt "has never yet landed, and the fix is
   unproven (fleetd #489)", and told the lead to warn the operator it may fail.
   Measured today from fleetd/fleetd.out:

     grep -c "lead-rollover: rolled" -> 4
     grep -c "lead-rollover:"        -> 16   (positive control)
     grep -c "Unknown command"       -> 0

   Three of the four rolls ran on 2026-09-22 (10:01:43, 10:38:28, 11:15:47).
   Each cleared the old lead and bootstrapped a fresh one against the handover
   file. The old "Unknown command: /clearFresh" failure does not appear at all.
   The paragraph now carries the measured result and the three re-measure
   commands, including the control line, because a broken grep pattern returns
   a clean 0 that reads like good news.

   It also records that the "/clear was never observed as WORKING ... releasing
   rather than wedging the roll" WARN accompanies every successful roll. That
   is the safe branch, not a failure, and it was being misread as one.

2. The skill said requireOperatorConfirm "defaults to true and this is the only
   thing standing between a judgement call and a wiped session", which reads as
   if asking is always required. The default is still true
   (FleetConfig.java:1426), but this host set it to false on 2026-09-22 on the
   operator's explicit grant. The bullet now says to read the live value rather
   than assume, and notes the key is deferred, not hot.

   It also warns that until fleetd #621 merges, LeadHeartbeatLoop.contextNotice()
   still hardcodes "ask the operator" and takes no config, so the nudge text and
   the config disagree. Trust the config. That warning names the ticket that
   removes it.

Documentation only. No code or test changes.
2026-09-22 11:24:03 +07:00
ltms 8915e40c7d Merge pull request 'fleetd #618: state the measured auto-compact precedence' (#619) from worker/618-b83894-2 into main
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 2m12s
2026-09-22 05:53:35 +02:00
Dai Ha 3f7bc3815e fleetd #612 Unit A: extract Fleetd.main's boot composition into FleetdAssembly/FleetdRuntime
CI / shell-tests (pull_request) Failing after 12s
CI / build (pull_request) Failing after 1m29s
CI / contract (pull_request) Successful in 1m28s
Fleetd.main kept config loading, startup reports and validation. Everything from
the herdr socket connect onward moved verbatim, same order, into
FleetdAssembly.assembleAndStart(AssemblyInputs, ResourcePorts), which returns a
FleetdRuntime owning the real objects (package-private accessors, never a copy)
and their single ordered close(). ResourcePorts/SystemResourcePorts abstract every
boot-time side effect (env, herdr connect, broker openers, clocks, schedulers,
shutdown-hook registration, HTTP start) with no inert production variant, per the
architect proposal on the ticket.

sleepHerdrPoll widened from private to package-private so FleetdAssembly can pass
a method reference to it; no other signature changed.

FleetdAssemblyLifecycleTest drives the real assembly with FakeHerdr, a temp
FleetConfig and a fake ResourcePorts recording a start/close ledger, asserting it
against the order recorded from the pre-move main() and shutdown hook, and proving
every resource the ledger can observe (herdr client, three schedulers, the AMQP
reply inbox) closes via FleetdRuntime.close(). FakeHerdr gained a closed flag for
this.
2026-09-22 10:52:35 +07:00
Dai Ha 6cb31a10e4 fleetd #618: fix the third stale spot the brief missed
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Failing after 2m6s
The method-level javadoc on FleetConfig.warnConflictingAutoCompactWindows
(above the log.warn call) still claimed the autoCompactWindow vs
CLAUDE_CODE_AUTO_COMPACT_WINDOW precedence was 'intentionally not
asserted' and cited fleetd.yaml's now-corrected comment as evidence the
question was open. Replace it with the measured answer from #618: the
env var wins, so autoCompactWindow is inert on a profile that sets both.
Kept the WARN-not-throw rationale paragraph above it untouched (#601)
and kept the ClaudeCodeArguments cross-reference, which now points to an
agreeing claim instead of a contradicting one. No behaviour change.
2026-09-22 10:50:59 +07:00
Dai Ha 8368a274a0 fleetd #618: state the measured auto-compact precedence, not 'unverified'
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Failing after 1m57s
ClaudeCodeArguments.withAutoCompactWindow's javadoc and FleetConfig's
warnConflictingAutoCompactWindows WARN text both used to say the
precedence between --autocompact and CLAUDE_CODE_AUTO_COMPACT_WINDOW was
not verified. fleetd #618 measured it: the env var wins, so the flag has
no effect when both are set. Update both texts to say so, name #618, and
warn that deleting the env var to resolve the conflict LOWERS the live
window rather than fixing anything. No behaviour change; the WARN still
fires on the same condition and stays a WARN (per #601).
2026-09-22 10:46:31 +07:00
ltms 17127efb88 Merge #601: pass auto-compact window to leads; warn instead of refusing on a conflict (CB-617)
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m16s
CI / build (push) Failing after 2m33s
2026-09-22 05:22:56 +02:00
ltms 203f034528 Merge #617: write FAILED instead of leaving a dead roll stuck at IN_PROGRESS (fleetd #615)
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m14s
CI / build (push) Failing after 1m47s
2026-09-22 05:21:21 +02:00
ltms 9ee16f5b85 Merge #616: report role-fallback gaps at boot, name contextHighNudge (fleetd #613)
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m15s
CI / build (push) Failing after 1m43s
2026-09-22 05:17:52 +02:00
Dai Ha 388ef5a3c3 fleetd #615: write FAILED instead of leaving status(token) stuck at IN_PROGRESS
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 2m25s
LeadRollover.runRollover made two unwrapped agents.send calls. HerdrException
is unchecked, and the production continuationRunner is a bare virtual thread
with no uncaught-exception handler, so a throw from either send call killed
the continuation silently — confirm() had already written IN_PROGRESS into
outcomes before scheduling it, and nothing ever overwrote that entry with a
terminal state.

Wrap the whole continuation body in one try/catch(RuntimeException), matching
the local convention already used around agents.status in
waitUntilAtTurnBoundary. On a throw, write a new terminal RollState.FAILED
entry naming the exception, in the same diagnostic style as
TURN_NEVER_SETTLED and CLEAR_NEVER_SETTLED.

Two new tests make send() throw on the /clear call and on the bootstrap-text
call respectively, each asserting status(token) reports FAILED, not
IN_PROGRESS. Reverting only the production catch (keeping the tests) turns
both red; restoring it turns them green again.
2026-09-22 10:17:12 +07:00
Dai Ha be6c45ff78 CB-617 review: warn instead of refuse on conflicting autoCompactWindow
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m9s
rejectConflictingAutoCompactWindows threw and stopped fleetd from starting when a
Claude Code profile's autoCompactWindow flag and CLAUDE_CODE_AUTO_COMPACT_WINDOW
env var disagreed. Under launchd that is a restart loop, and the config that
would fix it (fleetd.yaml) is gitignored, so the cause is invisible on the host
where it bites (measured live: 4 profiles on this host trip it, including the
lead's own profile and the one every worker spawns on).

Renamed to warnConflictingAutoCompactWindows: it now logs a WARN naming each
offending profile with BOTH values (autoCompactWindow=... and
env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=...) instead of throwing, so the daemon
starts and an operator can fix the config without reading the source. Equal
values still load silently.

Also reworded ClaudeCodeArguments' javadoc, which stated as fact that the env
var takes precedence over the flag. That was never measured, and this host's
own fleetd.yaml comment asserts the opposite — the javadoc no longer picks a
side.
2026-09-22 10:13:14 +07:00
Dai Ha e99cb70a8b CB-617: pass auto-compact window to leads 2026-09-22 10:12:56 +07:00
Dai Ha 987ccef4c7 fleetd #613: log role-fallback gaps at boot, name contextHighNudge in the heartbeat line
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Failing after 1m48s
- reportRoleFallbackGaps(cfg), called right after cfg.validateAll() in Fleetd.main, logs every
  MemberRole with no fleet.<role>s: pool (naming the profile count and the resolved
  defaultProfileFor(role) first choice) and, separately, every role with no
  fleet.charters.<role>: entry. Log only — the deliberate 'unconstrained' fallback in
  FleetConfig#candidateProfiles / CompositePeerLauncher#poolFor is unchanged, and a config with
  profiles: and no fleet: block still starts and still spawns.
- LeadHeartbeatLoop#start()'s boot line now also names contextHighNudge (fleetd #609), alongside
  the three settings it already logged.
- RoleFallbackGapReportTest (new) and two new LeadHeartbeatLoopTest cases pin both lines' content
  via a ListAppender, raising the dev.ltms.fleet logger past logback-test.xml's WARN override for
  the INFO-level lines.
2026-09-22 10:12:09 +07:00
ltms 076cc43f7b Merge #614: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 2m20s
CI has been red on main itself since #602/#606, on this one test, so the build has been giving no second opinion on any PR. Cause: the CI job runs in a container as root. setReadable(false) really does clear the read bit, so the test's own setup guard passes, but root opens the file anyway and the gauge correctly returns OK. The test was asserting on a condition the environment never created.

The fix adds an assumeFalse(Files.isReadable(file), ...) after the chmod and before the gauge is built, inside the existing try, so the finally still restores the bit on a skip.

Verified by me on a scratch worktree merging this onto 955b9ea:
- 1864 tests, 0 failures, 0 errors, 0 skipped, 149 surefire reports, mvn exit 0. The suite-wide skipped=0 is the point: the fix did not quietly turn the test into a permanent skip.
- LeadContextGaugeTest on this non-root Mac: 9 tests, 0 skipped, and unreadableFileIsUnknown present in the report. The assumption does not fire here, so developers keep the coverage.
- Mutation: made the IOException path return OK instead of UNKNOWN. unreadableFileIsUnknown failed with "expected: <UNKNOWN> but was: <OK>". The test still has teeth. Production file reverted, git diff clean before merge.

Known trade, recorded rather than hidden: under root this case is now covered by nothing at all. A skip is honest about that, which an assertion on an unreachable state was not. The durable fix is to run the CI build as a non-root user; that is a CI configuration change and out of scope here.
2026-09-20 12:43:43 +02:00
Dai Ha bad47a8444 fleetd CI: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m27s
The test set the file's read bit off via setReadable(false), but on the
Gitea CI runner (root inside the container) the OS ignores that bit and
opens the file anyway, so the test asserted on a condition it never
actually created (LeadContextGaugeTest.java:142 UNKNOWN vs OK, CI run
1887 job 3104, commit fa62e99 on main).

Add Files.isReadable(file) after setReadable(false) and before the gauge
runs, and assumeFalse on it: a skip means "could not set up the case",
never "the behaviour is fine". Restores the read bit either way so
@TempDir cleanup still works.
2026-09-20 17:40:47 +07:00
ltms 955b9ea013 Merge #610: nudge an idle lead to hand over when its own context reads HIGH (fleetd #609)
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m51s
Closes fleetd #609. Completes the second half of the context work: #602/#606 could detect a full lead context, and nothing acted on it. LeadHeartbeatLoop now offers a handover when the lead's own gauge reads HIGH.

Never rolls a pane by itself. The nudge is text only; the lead still has to call fleet_handover, and that still needs operatorConfirmed.

Verified by me on a scratch worktree merging 89cb8ff onto fa62e99:
- 1864 tests, 0 failures, 0 errors, 149 surefire reports, mvn exit 0.
- Four mutations, all killed: the two the implementer ran (latch back in applyDecision: 4 failures; call site drops the latch argument: 2 failures), one of my own at the line the logic moved TO (latch set regardless of send outcome: 1 failure), and a control on an untouched line (quiet-cap boundary < to <=: 4 failures). The control is what makes the other kills evidence.
- The new tests assert on herdr.sentTexts() — what actually reached the fake pane — not on source text. That is the right observable for a defect whose essence was "the latch says told, the pane got nothing".

The review blocker from the first round is fixed: the latch used to be committed by applyDecision before injectNudge tried to send, and injectNudge swallows its own RuntimeException. On quietNudgeCap: 0, which is this host's configuration, that was the normal path and not an edge case. The latch is now set only when a notice was actually included and the send returned.

Known and deliberately not blocked: Fleetd.main's own one-line call to leadContextSource is not pinned by a test. That is pre-existing and class-wide, tracked in #612.
2026-09-20 12:37:43 +02:00
Dai Ha 89cb8ff79b fleetd #609 review: the context latch must mean the notice reached the pane
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m6s
Fixes the PR #610 review blocker: LeadHeartbeatLoop committed contextNotified
before injectNudge attempted the send, so a transient herdr failure marked the
lead as told when nothing reached its pane, and contextNotice() carried no
latch at all, so a pending-driven INJECT re-appended the notice on every tick
while the context stayed HIGH.

- injectNudge now reports whether agents.send succeeded and persists
  contextNotified only when a notice was actually included in the text and
  the send did not throw. The latch is split out of applyDecision (kept for
  idleSinceNanos/quietCount, applied unconditionally as before) so it is
  written on the success path only, once per branch in tick().
- contextNotice gained an overloaded 3-arg form gated on the latch as it
  stood before the tick's decision; the existing 2-arg form delegates to it
  with alreadyNotified=false, so all pre-existing callers/tests are unchanged.
- tick() is now package-private (mirrors ReplyPushLoop#tick(String)) so tests
  can drive the real send path with a fake AgentControl instead of only the
  pure decide() function.
- Added tests I-L covering: a failed send does not consume the notice and
  retries; a successful send does; the text is gated when the latch is
  already set; and the notice appears exactly once across three differently
  driven INJECTs.

Both required mutations verified red and reverted:
1. Setting the latch from the Decision regardless of send outcome -> test I
   (iAFailedSendDoesNotConsumeTheNotice) fails.
2. Dropping the latch argument at the contextNotice call site -> tests K
   (kAPendingDrivenInjectWithTheLatchAlreadySetSendsNoNotice) and L
   (lTheNoticeAppearsExactlyOnceAcrossThreeDifferentlyDrivenInjects) fail.
2026-09-20 17:34:05 +07:00
ltms fa62e9906d Merge #611: make the collected-ticket nudge test deterministic (fleetd #608)
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m48s
Replaces a wall-clock bet with a manually-driven scheduler, so the tick runs
only when the test runs it. Also closes a second, smaller race the brief did
not name: waiting on Phase.DONE is not enough, because complete() can publish
isDone() before every whenComplete dependent has run.

Verified by the lead: full suite 1841/1841 green in a clean worktree, and an
independent mutation (hasTicketWork forced true) diagnosed as an equivalent
mutant — ReplyPushLoop.injectNudge re-reads the pending collections and returns
early, so that line cannot reach agent.prompt. The worker's own mutation
(ticketCollected made a no-op) is the one that reaches the observable, and it
killed.
2026-09-20 12:26:16 +02:00
Dai Ha d7390ccd37 fleetd #609 review: repair a garbled comment carried over from the brief
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 1m41s
The brief's sentence about a null token count at HIGH was broken, and the
worker copied it into the source verbatim. The code was already right; only
the comment was unreadable.

Says what is actually true: a HIGH reading always carries a non-null token
count today, because LeadContextGauge only reaches HIGH by comparing a number
against HIGH_THRESHOLD_TOKENS. That invariant lives in another class and
nothing asserts it, so the branch stays.
2026-09-20 17:20:48 +07:00
Dai Ha aa517ae0ec fleetd #608: make anAlreadyCollectedTicketProducesNoNudge deterministic
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Failing after 1m34s
Replace the real ScheduledExecutorService backing ReplyPushLoop in this one
test with ManualScheduler, a fake that only runs a tick when the test calls
runDueTasks(). The old test bet a 300ms backoff was wide enough that
collecting the ticket always won the race against the scheduler's own timer
- true on an idle machine, false under a loaded full-suite run, which is
exactly the flake reported.

The rewritten test also waits on setAfterFinishAsyncTaskCompleteHookForTest
(already used elsewhere in this file) instead of polling Phase.DONE, so it
does not race CompletableFuture.complete()'s own publish-then-run-dependents
gap (fleetd #399) while proving ReplyPushLoop.onTicketTerminal really ran
before the ticket is collected.

Verified: backoff=1 (the most hostile value) still passes; mutating
ReplyPushLoop.ticketCollected to a no-op turns the test red with the same
assertion message the original flake reported; three consecutive full-suite
runs are green (1841/1841 each).
2026-09-20 17:18:45 +07:00
Dai Ha 60496831c2 fleetd #609: nudge an idle lead to hand over when its own context reads HIGH
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 2m5s
LeadHeartbeatLoop can now append a text-only notice to its nudge when the lead's
own LeadContextGauge reading is HIGH and leadHeartbeat.contextHighNudge is on.
Fires once per HIGH stretch (a latch, cleared only by a later OK reading; UNKNOWN
neither sets nor clears it), never spends the quietNudgeCap budget, and never
rolls a pane itself — only the operator can approve a handover.

- LeadContextGauge.Reading.unknown() widened to public for LeadContextSource.none()
- FleetConfig.LeadHeartbeat gains contextHighNudge (null/false = off, unchanged default)
- LeadHeartbeatLoop.decide gains context/contextNotified; Fleetd wires a new
  leadContextLookup/leadContextSource factory pair (LeadHeartbeatLoop.LeadContextSource)
- fleetd.example.yaml documents the new key
2026-09-20 17:11:54 +07:00
ltms 9a992d0f70 Merge #602: report a lead's live context usage in fleet_list
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m14s
CI / build (push) Failing after 1m33s
LeadContextGauge reports OK / HIGH / UNKNOWN for each lead in fleet_list, read
from the transcript Claude Code writes for itself — never from the lead's pane,
so it does not touch the control plane invariant 5 protects.

Why: a lead on this host auto-compacted 30 times in one session, discarding
roughly 250,000 tokens each time, and nothing could see it coming. The handover
feature already existed; what was missing was any way to know when to use it.

Three design points that earned their place:
  - finds the transcript by NAME under <configDir>/projects/, never by deriving
    the project slug, which is an undocumented Claude Code internal
  - three states, not two. Every path that cannot establish a token count says
    UNKNOWN with no number, so a lead is never told it is fine when the honest
    answer is "I could not look"
  - bounded twice: TAIL_BYTES caps bytes read, a 5s cache TTL caps how often

Includes #606, which fixed the two defects this PR's body recorded before it
was ever merged: the null configDir that made the gauge inert on this fleet,
and the torn final line that would have made it flap.

The authoring member died mid-task on a DNS error with the work uncommitted;
the lead recovered it, verified it, and committed it with that stated.

Verified by the lead on the combined state with current main merged in:
mvn -o clean install exit 0, 147 reports, 1841 tests, 0 failures, 0 errors,
0 skipped. LeadContextGaugeTest 9/9, FleetMcpLeadContextGaugeWiringTest 2/2,
FleetdLeadConfigDirLookupTest 6/6, FleetdLeadConfigDirSourceWiringTest 3/3.

Not deployed by this merge. A merge is not a deployment; the running daemon
still holds its old jar until redeploy.
2026-09-20 11:38:31 +02:00
ltms 3762aca307 Merge #606: wire the lead context gauge to the real configDir, and stop flapping on a torn line
CI / shell-tests (pull_request) Failing after 6s
CI / build (pull_request) Failing after 1m21s
CI / contract (pull_request) Successful in 1m50s
Fixes the two defects recorded in #602's own body.

1. FleetMcp.contextView passed configDir=null, so the gauge read <user.home>/
   .claude while this host's lead profile sets an override. Measured before the
   fix: 19 transcripts under the real directory, 0 under the fallback. The gauge
   would have deployed green and reported UNKNOWN forever, for every lead.
   Now threaded via FleetMcp.LeadConfigDirSource, built by Fleetd
   .leadConfigDirSource, following fleet.leaders.<name>.profile to that
   profile's configDir and reading config.get() live inside the lambda.

2. A torn final line no longer means UNKNOWN. fleet_list reads a transcript
   Claude Code may be mid-write on, so the last line can be cut. The old code
   treated that as fatal, which would make the gauge flap at random. The stated
   reason ("a format change should show as UNKNOWN") does not hold: a real
   format change makes EVERY line unparseable, and that case is still caught.

Verified by the lead on the combined state with current main merged in, not on
the branch alone: mvn -o clean install exit 0, 147 reports, 1841 tests, 0
failures, 0 errors, 0 skipped.

Mutation-checked independently by the lead:
  Fleetd.leadConfigDirSource body -> none()   -> KILLED (1 failure)
  FleetMcp.contextView configDir -> null      -> KILLED (worker-measured)

KNOWN RESIDUAL, documented rather than overclaimed. main's own one-line call to
leadConfigDirSource could be swapped for none() and the suite stays green. Every
member of this wiring-test family (loopHealthSource, capacitySource,
healthCoverageSource) has the identical gap — no test runs Fleetd.main far enough
to observe which factory it called. Filed separately as a class-wide problem
rather than patched here. The empirical close is the dogfood check after redeploy.
2026-09-20 11:38:22 +02:00
Dai Ha 4e27bde2d7 fleetd #602 gauge-wiring follow-up: pin Fleetd.main's LeadConfigDirSource wiring
Extract the inline new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(...))
construction in Fleetd.main into a package-private factory,
Fleetd.leadConfigDirSource, mirroring loopHealthSource/capacitySource/
healthCoverageSource. Add FleetdLeadConfigDirSourceWiringTest, which calls the
factory directly with real Profile/Leader fixtures and asserts the returned
source resolves a real configDir -- a property that is false if the factory's
body is mutated to return LeadConfigDirSource.none().

Neither FleetMcpLeadContextGaugeWiringTest nor FleetdLeadConfigDirLookupTest
could catch main losing this wiring: each builds its own instance instead of
calling what main calls. This closes that gap at the factory level, matching
the standard already accepted for loopHealthSource's own wiring test.
2026-09-20 16:34:49 +07:00
ltms b85d9b0e46 Merge #607: redeploy-fleetd.sh no longer fails a deploy that worked
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m49s
fleetd #603. The pid poll had its own fixed 10s budget and then hard-died,
while the health check right after it was allowed 60s for the same daemon.
Under launchd the java process does not exist yet when launchctl load returns,
so the script reported FAIL on a fully successful deploy — and a false FAIL in
that direction invites the hand-rolled stop/start this script exists to replace.

The pid poll now shares HEALTH_WAIT, and a miss falls through to the health
check rather than killing the run. A genuine failure still dies and still
prints the log tail. Worst-case time-to-fail roughly doubles; that cost lands
only on real failures and is the right trade.

Verified by the lead, not taken from the report. Suite exit 0, unpiped.
Two mutations run independently:
  budget back to a hardcoded 10          -> KILLED (suite exit 1)
  warn back to die on a pid miss         -> KILLED (suite exit 1) after 856dfc6
The second survived on the first submission, which is why that commit exists:
nothing covered "pid never appears but healthz answers, so do not die" — the
one case the fall-through is for, and reachable because running_pid() only
recognises a plain java -jar.
2026-09-20 11:30:10 +02:00
Dai Ha d345b14e43 fleetd #602 gauge-wiring: thread a lead's configured configDir into the context gauge
FleetMcp.contextView hardcoded LeadContextGauge.read(null, ...), so a lead whose
profile sets its own CLAUDE_CONFIG_DIR always read the wrong transcript directory
and reported UNKNOWN forever, with no error anywhere.

- Add FleetMcp.LeadConfigDirSource (same idiom as LeadSeatSource) and thread it
  through the constructor / listFleet overload chain / leadView / contextView.
- Add Fleetd.leadConfigDirLookup, wired at construction, following the same
  fleet.leaders.<name>.profile link leadSeatLookup already uses, one step
  further to that profile's own configDir.
- LeadContextGauge.parse: a single unparseable line (typically the final one,
  torn by a write this read raced) is now skipped rather than forcing UNKNOWN;
  only when every line in the read window fails to parse does it report
  UNKNOWN, which is the real format-change signal.
- Tests: FleetdLeadConfigDirLookupTest (lookup logic), FleetMcpLeadContextGaugeWiringTest
  (end-to-end: config naming directory A vs B decides which is read; a lead with
  no configured dir degrades without throwing), and two replacement properties in
  LeadContextGaugeTest for the torn-line fix plus its all-unparseable control.
2026-09-20 16:19:37 +07:00
ltms 9640deeffc Merge #605: fleet_list reports charterBytes alongside charterSha256
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m34s
#604 item 1. A digest answers "same or different" and cannot say how much. The
byte count is the second signal, needed exactly when two hosts find they differ.

Verified by the lead before merging, from fleetd/: mvn -o clean install exit 0,
143 reports, 1821 tests, 0 failures, 0 errors, 0 skipped; SessionManagerTest 74/74.

Includes ad3d819, correcting an invariant the added comment claimed but that
CharterReceipt.compose() does not hold: digestOf() returns null for blank text
while getBytes().length does not, so a whitespace-only charter pairs a null
digest with a non-zero size. No behaviour change — the digest gate already
omits both on that path, which is the right answer.
2026-09-20 11:18:17 +02:00
Dai Ha ad3d81941f Correct the invariant claimed in the charterBytes comment
CI / shell-tests (pull_request) Failing after 7s
CI / build (pull_request) Successful in 1m36s
CI / contract (pull_request) Successful in 1m42s
The comment said CharterReceipt never pairs a null digest with a non-zero
byte count. It can. compose() derives the digest with digestOf(), which
returns null for blank text, while the byte count is getBytes().length,
which does not. A whitespace-only role charter on a profile with no MCP
produces exactly that pair.

No behaviour change. The gate already omits both fields on that path, which
is the right answer — a size with no digest would describe an artifact we
cannot fingerprint. Only the stated reason was wrong, and a false invariant
in a comment is worse than no comment, because the next reader will widen
the gate on the strength of it.
2026-09-20 16:16:30 +07:00
Dai Ha 8b986a52e0 #604 item 1: fleet_list reports charterBytes alongside charterSha256
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 2m26s
CharterReceipt carries a byte count next to its digest, but the roster
projection in SessionManager.rosterView only ever copied the digest
across. A digest tells a lead whether two members' charters match; it
cannot say how far apart they are when they don't. Report charterBytes
too, nested in the same conditional as charterSha256 so the two travel
together: the receipt's own contract only ever pairs a non-null digest
with a real byte count, and a member with no composed charter reports
charterSource alone, unchanged from before.

Tests: the existing charter-receipt roster test now asserts charterBytes
against the receipt's own value (not a literal), plus two new cases —
no charter composed (source "none", no digest, no size) and the receipt
itself absent (no charter keys at all).
2026-09-20 16:14:45 +07:00
Dai Ha f336bcef39 CLAUDE.md: rewrap the long line my last edit left in the architects paragraph
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m48s
The fleet01 lead found this. My edit in f5c6a0e left one prose line at 128
columns inside a paragraph wrapped at about 96. It is invisible to any check
that normalises by paragraph, and visible in any raw digest of the block.

No wording changed. Block stays 20937 bytes. The wiki template gets the same
rewrap in its own commit, so the two stay byte-identical.
2026-09-20 16:13:23 +07:00
Dai Ha 3ed7bfca67 fleetd: report a lead's live context usage in fleet_list
CI / shell-tests (pull_request) Failing after 7s
CI / build (pull_request) Failing after 1m37s
CI / contract (pull_request) Successful in 2m7s
fleetd had no way to see how full a lead's context window is. On this host a
lead auto-compacted 30 times in one session, discarding roughly 250,000 tokens
and costing 46s to 3m16s each time, and nothing could see it coming.

LeadContextGauge reads the transcript Claude Code itself writes, never the
lead's pane. It finds <sessionId>.jsonl by NAME under <configDir>/projects/
rather than deriving the project slug, which is an undocumented internal.

Three states, not two: OK, HIGH, UNKNOWN. Every path that cannot positively
establish a token count reports UNKNOWN with no number, so a lead is never
told it is fine when the honest answer is "I could not look".

Bounded two ways: TAIL_BYTES caps bytes read off disk, and a 5s cache TTL caps
how often that read happens, because fleet_list is polled constantly.

Recovered by the lead: the authoring member ended on a backend error (DNS
ENOTFOUND) with this work uncommitted and unpushed in its worktree. Verified
before committing: mvn -o clean install exit 0, 1813 tests, 0 failures,
0 errors, 0 skipped, 144 reports; LeadContextGaugeTest 8/8.

KNOWN INCOMPLETE - see the PR. The fleet_list call site passes configDir=null,
which falls back to ~/.claude, but this host's lead profile sets configDir to
an override. Measured: 19 transcripts under the real configDir, 0 under the
fallback. The gauge is therefore INERT on this fleet until that is wired.
2026-09-20 16:05:15 +07:00
52 changed files with 6284 additions and 1116 deletions
+51 -11
View File
@@ -167,19 +167,59 @@ fails.
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
from your connection, so you can only ever roll yourself.
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
defaults to `true` and this is the only thing standing between a judgement call and a wiped
session.
are confident. Ask, wait for the answer, then pass what they said.
- **Whether you must ask at all depends on `leadRollover.requireOperatorConfirm`. Check it; do not
assume.** The default is `true` (`FleetConfig.java:1426`), and then `confirm` refuses unless you
also pass `operatorConfirmed: true`. **This host set it to `false` on 2026-09-22**, on the
operator's explicit grant, because they do not want to approve routine context rolls. Where it is
`false`, the three handover-file checks are the whole gate: the file must exist, be fresher than
`maxDocAgeSeconds`, and have been modified after the open request.
Read the live value rather than trusting this line:
```bash
grep -A1 'requireOperatorConfirm' fleetd/fleetd.yaml
```
No match means the key is unset, so the default `true` applies and you must ask. The key is
**deferred, not hot** — it is read once at boot, so an edit does nothing until the daemon is
redeployed.
**Until fleetd #621 merges, the nudge text will tell you to ask the operator even where the
daemon no longer requires it.** `LeadHeartbeatLoop.contextNotice()` hardcodes "ask the operator"
and takes no config, so it cannot know. Trust the config value over the nudge text. Once #621 is
merged and deployed, the nudge matches the config and this warning can be deleted.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
the operator so before you confirm.** The recovery is the same either way: the file is already
written, so the operator starts a session and points it at the file. That is why you write the
file before you confirm, and never the other way round.
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43,
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the
handover file, with the configured `bootstrapText` arriving as its first message. No context was
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log
at all. Re-measure both numbers with:
```bash
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
grep -c "lead-rollover:" fleetd/fleetd.out # positive control: must be larger
grep -c "Unknown command" fleetd/fleetd.out # the old failure: expect 0
```
Run the control line too. A broken pattern returns a clean `0` that reads exactly like good news.
If the first number stops growing across rolls, or `Unknown command` returns anything above 0,
the bootstrap has regressed and this paragraph is stale again.
**You still write the file before you confirm, and never the other way round.** That order is not
about the bootstrap being unreliable. It is what the daemon checks: the handover file must have
been modified *after* the open request, or `confirm` refuses it as stale.
- **One warning in the log is normal and is not a failure.** Every one of the three rolls above also
logged `/clear on term_… was never observed as WORKING after 8 consecutive IDLE/DONE polls —
releasing rather than wedging the roll`. The daemon could not see the pane go WORKING after
`/clear`, so it released instead of hanging. The roll then succeeded anyway. That is the safe
branch behaving correctly. Do not report it as a broken roll.
## Writing style
+4 -3
View File
@@ -145,9 +145,10 @@ authorized to settle it, not only to advise. Architects first form independent p
compare them. If they still disagree after that comparison, they return both positions and their
checked evidence; the lead decides. Go to the operator only for an action the fleet has no
authority to take, such as spending money, granting access, or making a promise to someone else.
**Then write the decision on the ticket.** Taking the operator out of the loop also removes the signal they used to get, because
that signal was the block itself — work stopped, so they found out. A ticket comment replaces it,
and it reaches them whether or not they are at a terminal when you decide.
**Then write the decision on the ticket.** Taking the operator out of the loop also removes the
signal they used to get, because that signal was the block itself — work stopped, so they found
out. A ticket comment replaces it, and it reaches them whether or not they are at a terminal when
you decide.
| Intent | Tool |
|---|---|
+15 -6
View File
@@ -77,16 +77,25 @@ bind:
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
# block = feature off, exactly as before.
#
# Three knobs, each with a default that errs on the side of not burning context:
# Four knobs, each with a default that errs on the side of not burning context:
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
# # until real state appears (default 3 — never nag an empty fleet forever)
# contextHighNudge: false # fleetd #609 — when true, an idle lead whose OWN Claude Code context
# # reads HIGH (see LeadContextGauge; fleet_list's context row) gets a text
# # notice appended to its nudge telling it to consider fleet_handover. Text
# # only — it never rolls a pane by itself, and only the operator can approve
# # a roll. Fires once per HIGH stretch (a later OK reading re-arms it), and
# # never spends the quietNudgeCap budget. Default false/absent = off, same
# # as every other knob here — an upgraded daemon must not start telling
# # leads to hand over on its own.
# leadHeartbeat:
# idleAfterSeconds: 300
# backoffMs: 60000
# quietNudgeCap: 3
# contextHighNudge: false
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
@@ -200,13 +209,13 @@ herdrSocket: ~/.config/herdr/herdr.sock
# e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never
# fails the spawn. Omit to open the member's module by hand. There is no close
# half yet — an opened module stays open until the operator closes it.
# autoCompactWindow → opt-in, default off. A bounded token window that forces a spawned member to
# compact its context instead of running on the backend's own default and dying
# mid-turn (losing its fleet_reply — the whole point of the turn — with it).
# autoCompactWindow → opt-in, default off. A bounded token window that forces a launched Claude Code
# lead or member to compact its context instead of running on the backend's own
# default. A member that runs out of context can die mid-turn and lose its fleet_reply.
# Validated at config load to [100000, 1000000] — the band Claude Code's own
# --autocompact flag accepts.
# CROSS-BACKEND SEMANTICS DIFFER: on claude-code this is a launch-time
# `--autocompact <tokens>` flag — the member compacts AT this window. opencode
# `--autocompact <tokens>` flag — the Claude Code session compacts AT this window. opencode
# has no equivalent flag (it only forces `compaction.auto: true`, unconditionally,
# already), so this is instead applied as the model's `limit.context` in the
# generated opencode.json — the member compacts WITHIN this window, not exactly
@@ -406,7 +415,7 @@ profiles:
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
# ideProjectDir: fleetd # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in fleetd/)
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
# autoCompactWindow: 250000 # opt-in: bound member context; claude-code compacts AT this, opencode within it (model limit.context)
# autoCompactWindow: 250000 # opt-in: bound Claude Code lead/member context; claude-code compacts AT this, opencode within it (model limit.context)
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
@@ -0,0 +1,22 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
/**
* fleetd #612 Unit A: everything {@code Fleetd.main} has ready once config is loaded, reported
* and validated — the exact point {@code main} used to keep going straight into socket and broker
* work. Building this (and handing it to {@link FleetdAssembly#assembleAndStart}) is the seam a
* test now has to drive the real boot composition without being {@code main} itself.
*
* @param cfg the boot-time {@link FleetConfig} snapshot every one-time wiring decision reads —
* see {@code Fleetd.main}'s own comment on why this must never be swapped for a live
* reference once loaded
* @param config the live {@link ConfigRef} the hot-reloadable paths read per use
* @param guard the same {@link SubscriptionGuard} {@code Fleetd.main} already used to assert the
* launching environment is clean, reused rather than rebuilt so the assembled
* launchers see the identical instance {@code main} already validated with
*/
record AssemblyInputs(FleetConfig cfg, ConfigRef config, SubscriptionGuard guard) {
}
+199 -602
View File
@@ -4,11 +4,13 @@ import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.ConfigWatcher;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.herdr.PaneLocator;
@@ -40,6 +42,7 @@ import dev.ltms.fleet.mcp.LsofProcessCwdLookup;
import dev.ltms.fleet.msg.AmqpReplyInbox;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.LeadChannel;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadMailbox;
import dev.ltms.fleet.msg.MessageService;
@@ -51,6 +54,7 @@ import dev.ltms.fleet.rest.FleetApp;
import dev.ltms.fleet.session.GitWorktrees;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.session.SessionReaper;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
@@ -174,595 +178,18 @@ public final class Fleetd {
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
// why a name-by-name list here would have the same defect it replaces.
cfg.validateAll();
// fleetd #469, follow-up to #464: validateAll() (and validateCharters() inside it) only
// checks that a charter's KEY is a role wire name and its text is non-blank — it never
// looks at what the text actually names. This is the separate check that does: it asks
// dev.ltms.fleet.mcp.FleetTool (the canonical registered-tool set) whether every fleet_*/
// bridge_* token a charter names is a tool this server actually registers. It cannot live
// inside FleetConfig#validateCharters() — config loads before the MCP server exists, and
// must not gain a dependency on the mcp package — so it runs here instead, at the one seam
// that already holds both a loaded FleetConfig and the mcp package, before anything below
// opens a socket or spawns a member. fleetd #474: the same check is also wired into `config`
// above as ConfigRef's extraValidation, so a reload refuses what this line refuses at startup.
assertChartersNameOnlyRegisteredTools(cfg);
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
UnixSocketHerdrClient memberHerdr = cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank()
? UnixSocketHerdrClient.connect(Path.of(cfg.memberHerdrSocket()), new com.fasterxml.jackson.databind.ObjectMapper())
: herdr;
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
target -> leadsRef.get().get().containsKey(target));
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, FleetConfig.Profile> claudeProfiles = new LinkedHashMap<>();
Map<String, FleetConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
cfg.profiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// fleetd #175: the daemon's real ExhaustionSink can only be built once `sessions` exists
// (below), but `sessions` needs `workers`, which needs the adapters built right here — a
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
// pointed at the real one once it exists.
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
// fleetd #234, round 4: routed through the shared ExhaustionSink.forwardingTo factory
// rather than written inline here — not because a lambda at this call site is unsafe
// anymore (it is not: the 3-arg overload is now the interface's single abstract method, so
// there is no 2-arg overload left for any lambda to silently bind to instead), but so a
// test can call the exact same object this line builds, instead of asserting a copy of its
// shape (round 3's lesson).
// fleetd #589 Group 1: extracted to forwardingExhaustionSink(...) below (see that method's
// javadoc) so a dedicated test can prove this factory keeps reading the reference live,
// rather than a rebuilt copy of its shape.
ExhaustionSink forwardingExhaustionSink = forwardingExhaustionSink(exhaustionSinkRef);
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
// fleetd #589 Group 2: extracted to claudeCodeLauncher(...)/openCodeLauncher(...) below (see
// those methods' javadoc) so a dedicated test can prove the CB-596 memberCredentials policy
// supplier is actually wired to each adapter, not silently replaced with `() -> null`.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(claudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
claudeProfiles, cfg, config));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(openCodeLauncher(router.memberAgents(), router.memberSpaces(),
opencodeProfiles, cfg, config, forwardingExhaustionSink));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
// fleetd #466: escalating, not flat — a credential that keeps reporting exhaustion (e.g. a
// weekly subscription limit, which would otherwise be retried on every ~30-minute cooldown,
// about 336 times across the week) backs off further each consecutive time, capped at
// BackendQuarantine.DEFAULT_MAX_COOLDOWN_MULTIPLE x the base cooldown. See BackendQuarantine's
// class doc for the mechanism, why this never fires on cooling-off (a separate, unescalated
// mechanism — BackendOutagePolicy below), and the reset.
BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
// below (written on a classified backend error). A SEPARATE, shorter-lived mechanism from
// `quarantine` — see BackendOutagePolicy's class doc — never merged with it.
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(System::nanoTime);
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.effectiveDefaultProfile(),
config,
profileName -> liveCountRef.get().apply(profileName),
quarantine,
outagePolicy);
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into —
// no models: block at all, a block armed with nothing off, or a block with N off — the same
// way exhaustedPatternCoverageLine/errorPatternCoverageLine report CB-578 stage A/fleetd
// #201 Unit 5 coverage just below. Read from workers.modelGateState() (never a separate
// config.get().models() here) so this line and fleet_profiles' modelGateArmed can never
// disagree about what CompositePeerLauncher's spawn gate actually enforces.
log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState()));
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
HerdrAwaitOutcome herdrOutcome = awaitHerdr(herdr, System::nanoTime, Fleetd::sleepHerdrPoll);
boolean herdrUp = logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers,
new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup(), cfg.memberSkills()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> liveSessionCount(sessions.roster(), profileName));
// Idle-sleep guard: hold an OS-level assertion against idle sleep while at least one
// member is live, so an unattended host does not idle-sleep out from under a member's
// long turn (see FleetConfig.IdleSleepGuard / dev.ltms.fleet.power.IdleSleepGuard for the
// measurement that motivated this). Opt-out via idleSleepGuard.enabled: false; on by
// default. Hangs off SessionManager's own onAcquire/onRelease hooks (CB-520/CB-516,
// previously wired only to the reply inbox) and SessionManager#size() — the exact registry
// fleet_list's live/capacity numbers are themselves computed from — rather than tracking
// members a second way. No-op (never constructed) off macOS or when idleSleepGuard.enabled
// is explicitly false; the mechanism itself is additionally a no-op if 'caffeinate' cannot
// be started, so this can never fail a spawn, a release, or startup.
boolean idleSleepGuardEnabled = cfg.idleSleepGuard() == null || cfg.idleSleepGuard().isEnabled();
final IdleSleepGuard idleSleepGuard;
if (idleSleepGuardEnabled) {
idleSleepGuard = new IdleSleepGuard(new CaffeinateSleepAssertionMechanism(), sessions::size);
sessions.onAcquire(_ -> idleSleepGuard.recheck());
sessions.onRelease(_ -> idleSleepGuard.recheck());
} else {
idleSleepGuard = null;
}
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
// loop's nudges, which need a single destination — so it keeps the legacy pin.
Map<String, String> leadTerminals = cfg.leaderTerminals();
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the legacy primary.terminal pin, discover leads by the tab labels the
// operator writes. CB-557 moved the settings onto the lead they describe, so scanning is on
// whenever a `fleet.leaders:` entry exists — with no leads configured the supplier is a
// constant and never touches herdr, exactly as a missing `leadScan:` block used to behave.
// CB-579: each lead now names its own exact `tab:` label, so one scanner discovers every
// configured lead regardless of how differently their tabs are labelled — the old
// single-shared-tabPrefix limitation (and its warning) is gone.
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
// The lead and the members now share ONE workspace (the operator asked for a single
// "session" with many tabs), so no workspace can be excluded — the lead lives in the
// members' space by design. A lead is told from a member by its exact tab label alone:
// a lead carries its configured `tab`, a member its `worker: {profile} #{n}` template,
// and the two never collide. (The scanner still supports an exclusion set for a split
// layout; the fleet's policy here is simply not to use one.)
//
// One shared rescan cadence: still taken from the first entry, as before — it is an
// operational cadence, not identity, so there is no correctness reason to give every
// lead its own scanner.
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
tabToName.keySet(), scanIntervalSeconds);
} else {
leads = () -> leadTerminals;
}
leadsRef.set(leads);
// CB-558: start any declared lead that is not already running. After the scanner is built,
// because both read the same tab labels and the ordering makes that dependency visible; and
// only when herdr answered, because the launcher's whole safety property is that it can
// count live leads first — it must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
}
// CB-548: config-declared architect slots. Config supplies only the stable name → profile
// map; the terminal → slot binding is owned by the registry and is empty at startup, so no
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
// fleetd #424: MemberRegistry.live re-reads fleet.architects through `config` on every
// reserve/requireSlotFor call, so a reload that removes or adds an architect slot governs
// the next spawn with no restart — the frozen `new MemberRegistry(cfg.fleet())` this used
// to be let a "revoked" slot keep granting new architect spawns forever.
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
+ "spawn lifecycle binds a live terminal to it)",
members.slots().size(), members.slots().keySet());
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
// fleetd #446: read live off `config` per lookup, cached by profile name — see
// LiveExhaustedPatterns's class doc for why this replaces the old compiled-once-at-startup
// map. A profile with no exhaustedPattern simply returns null here, so its workers keep
// today's completion-fallback behaviour unchanged.
// fleetd #589 Group 1: both extracted to liveExhaustedPatterns(...)/
// exhaustedPatternLookup(...) below (see those methods' javadoc) — this is the worst
// consequence in the whole #589 sweep: silently losing either wiring means a genuine
// usage-limit refusal is handed back as a real completion instead of BACKEND_EXHAUSTED.
LiveExhaustedPatterns liveExhaustedPatterns = liveExhaustedPatterns(config);
ExhaustedPatternLookup exhaustedPatterns = exhaustedPatternLookup(sessions::roster, liveExhaustedPatterns);
// The startup coverage line still reports the boot-time snapshot only — it is printed once,
// here, and a reload no longer needs to change what it said; exhaustionDetectionArmed (via
// liveExhaustedPatterns.armed, wired into quarantineSource below) is what stays live.
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
.filter(e -> e.getValue().hasExhaustedPattern())
.map(Map.Entry::getKey)
.collect(Collectors.toCollection(LinkedHashSet::new));
log.info("backend-exhausted classification (CB-578 stage A): {}",
exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
// rather than handing it back as a real answer. Still compiled once at startup, keyed by
// profile name — unlike exhaustedPattern above (fleetd #446), errorPattern was left DEFERRED
// on purpose: the ticket that made exhaustedPattern hot scoped errorPattern/cooling-off out
// explicitly. A profile with no configured errorPattern is simply absent here, so
// CompletionResolver falls back to its built-in narrow {@code (?i)\bAPI Error\s*:}
// compatibility pattern for that profile's targets (BackendErrorPatternLookup#legacy's
// contract — see backendErrorPatterns below).
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasErrorPattern()) {
errorPatternsByProfile.put(name, Pattern.compile(profile.errorPattern()));
}
});
// fleetd #248: extracted to a static factory (see backendErrorPatternLookup below) so a
// test can prove main() actually PASSES this into CompletionResolver, not only that the
// lookup itself behaves correctly — the exact gap fleetd #248 exists to close.
BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,
errorPatternsByProfile);
log.info("backend-error classification (fleetd #201 Unit 5): {}",
errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
// fleetd #446 criterion 3: the backend text that triggered the most recent quarantine of
// each credential, so fleet_profiles can report WHY a limit was hit, not only that it was.
// Keyed by credential id — the same key BackendQuarantine's own remainingSeconds uses —
// and written at the one call site that actually quarantines (inside exhaustionSink below),
// so a reason can never be reported for a quarantine that never happened. Bounded the same
// way BackendQuarantine's own internal map is documented to be: by the number of distinct
// credentials ever exhausted, not by how often the config is edited.
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
// fleetd #446 follow-up (round 3): extracted into exhaustionSink(...) below — see that
// method's javadoc for the full fleetd #175/#234/#446 history this used to carry inline —
// so a dedicated test can drive the exact ExhaustionSink main() builds, not a hand-rebuilt
// copy of its shape.
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
// now that `sessions` exists to resolve target -> session -> profile.
// fleetd #589 Group 1: both statements (build + set) folded into publishExhaustionSink(...)
// below (see that method's javadoc), so a test can prove the reference is actually
// repointed at the real sink, not silently left at ExhaustionSink.none().
ExhaustionSink exhaustionSink = publishExhaustionSink(exhaustionSinkRef, sessions, config,
quarantine, quarantineReasonByCredential, cfg);
// fleetd #201 Unit 5: the production BackendErrorSink needs `pushLoop` (built further below,
// after `sessions`) to tell a lead about an incident or an unmapped target — the same
// construction-order cycle `exhaustionSinkRef` breaks above, broken the same way: a mutable
// holder set once `pushLoop` exists, read lazily from inside the lambda built here.
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>();
// fleetd #248: extracted to a static factory (see backendErrorSink below), public rather
// than package-private like the other two factories here, so
// dev.ltms.fleet.inject.BackendOutageFlowTest can exercise the REAL production sink
// directly instead of a hand-mirrored copy of this lambda — that copy was precisely the
// gap fleetd #248 exists to close (see that test's class doc for the history).
BackendErrorSink backendErrorSink = backendErrorSink(sessions, () -> config.get().profiles(),
outagePolicy, pushLoopRef::get);
AgentControl agents = router.memberAgents();
// Both fleetd#201 Unit 5 (backend-error patterns + sink) and fleetd#241 (the worktree/branch
// lookup the fallback report names) land on this one call. The full constructor takes both,
// so neither feature is dropped; nowNanos must be passed explicitly to reach it.
//
// fleetd #248: every argument built specifically for this call (backendErrorPatterns and
// backendErrorSink above, and the worktree/branch lookup right here) now comes from a
// static factory tested on its own; FleetdCompletionResolverWiringTest source-asserts that
// THIS call actually passes them, which is the coverage that was missing before.
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns,
exhaustionSink, backendErrorPatterns, backendErrorSink, System::nanoTime,
worktreeBranchLookup(sessions::roster));
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
MemberPresence presence = sessions.asPresence();
// fleetd #561: extracted to a static factory (see turnListener below) — the anonymous class
// this replaced had two bare, unguarded statements per callback, and nothing enforced that
// the completion half went first beyond call order in the source.
TurnListener turnListener = turnListener(completion, sessions);
Predicate<String> deliverable = deliverableTo(presence, leads);
// fleetd #556: registration is wired directly to `completion`, not folded into the
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
// listener) throwing, regardless of call order. See TurnRegistrar's javadoc.
// fleetd #589 Group 3 (:505): extracted to turnRegistrar(...) — see FleetdTurnRegistrarWiringTest.
Injector injector = new Injector(router, turnListener, deliverable,
presence::forget, turnRegistrar(completion));
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
// unusable), fleetd stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
// fleetd #589 Group 3 (:512): opener extracted to replyInboxOpener() — see
// FleetdReplyInboxOpenerWiringTest.
final ReplyInbox replyInbox = selectReplyInbox(cfg.broker(), System.getenv(), replyInboxOpener());
// CB-637: this daemon's lead-to-lead mailbox on the SHARED coordination vhost — a separate
// broker from the reply inbox by design (see FleetConfig.Coordinator). Absent a coordinator:
// block this is null and every lead path below is simply not wired, which is exactly the
// behaviour before this ticket. It owns a broker connection, so keep the reference for the
// ordered shutdown hook.
// fleetd #589 Group 3 (:518): opener extracted to leadMailboxOpener() — see
// FleetdLeadMailboxOpenerWiringTest.
final LeadMailbox leadMailbox = openLeadMailbox(cfg.coordinator(), System.getenv(), leadMailboxOpener());
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs —
// identity comes from `leaders:`/`leadScan:`, and reply nudges now follow the delegating
// lead. Say so once at startup rather than leaving a redundant pin to look load-bearing.
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
+ "It still works, and is still the fallback nudge destination when a restart "
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open fleet_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-push-").unstarted(r));
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
// fleetd #201 Unit 5: point the forwarding holder captured by the backendErrorSink lambda
// above at the real push loop, now that it exists.
pushLoopRef.set(pushLoop);
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
// It has its own single-thread scheduler and holds its own scheduler shutdown via close().
final LeadHeartbeatLoop heartbeat;
var heartbeatScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, System::nanoTime,
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics);
heartbeat.start();
} else {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed, so
// an upgraded daemon cannot silently acquire the ability to clear the lead's own pane.
// Unlike heartbeat above, this has no recurring scheduler of its own — nothing but an
// explicit confirm() call (wired to an MCP tool by a later ticket; nothing calls it yet)
// that passes every gate can ever schedule a roll. It does not take primaryRegistry: the
// lead terminal to roll comes from the caller of open()/confirm() (resolved by the MCP
// layer from the connection, the same way auth/CallerResolver#resolve builds a
// Principal.leader(...)), never from a single-slot lookup — see LeadRollover's class
// javadoc, fleetd #480 correction 2.
LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
final FleetHealthMonitor healthMonitor;
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-health-").unstarted(r));
if (cfg.health() != null && cfg.health().isEnabled()) {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
// through the same idempotent target-wide operation CB-516 already uses on release.
// fleetd #386: System::nanoTime freezes across a macOS sleep, so the stall check also
// gets a wall-clock source to detect and correct for that freeze. Every other decision
// in FleetHealthMonitor stays on the monotonic clock, unchanged.
// fleetd #589 Group 3 (:591): failTarget extracted to healthFailTarget(...) — see
// FleetdHealthFailTargetWiringTest.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis()),
cfg.health().intervalOrDefault(),
cfg.health().workingSuspectAfterOrDefault(), healthFailTarget(messages));
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
log.warn("fleet health: {} (no notification sink configured)", coverage);
} else {
log.info("fleet health: {}", coverage);
}
healthMonitor.start();
} else {
healthMonitor = null;
healthScheduler.shutdownNow();
}
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
// fleetd #589 Group 3 (:611-631): the whole cleanup lambda extracted to releaseCleanup(...)
// — see FleetdReleaseCleanupWiringTest, and MessageService.abandon's javadoc for the
// documented incident (a torn-down worker's rendezvous waiter left open) this lambda exists
// to prevent.
sessions.onRelease(releaseCleanup(messages, replyInbox, primaryRegistry));
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
// CB-185: a caller's pane can live on either daemon (a lead's on the lead daemon, a
// member's on the member daemon) — search both, lead first. Collapses to one scan when
// memberHerdrSocket is unset (herdr == memberHerdr).
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting fleetd");
}
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
// fleetd #297: named once and reused verbatim below for FleetApp's GET /profiles, rather than
// built a second time — two independently-constructed sources reading the SAME BackendQuarantine
// / BackendOutagePolicy would still be able to drift (e.g. a future edit to the credentialIdFor
// closure in only one of the two places), exactly the shape #284 was.
FleetMcp.QuarantineSource quarantineSource = quarantineSource(config, quarantine,
liveExhaustedPatterns, quarantineReasonByCredential);
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, outagePolicy);
FleetMcp.LoopHealthSource loopHealth = loopHealthSource(poller, reaper);
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, FleetMcp.AuthorizationMode.ENFORCED, metrics,
capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
healthCoverageSource(config),
loopHealth,
quarantineSource,
leadMailbox,
outageSource,
new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), leaders, leads)),
// fleetd #361: the operator-declared peers this daemon's fleet_list should try to
// reach. Read from the SAME snapshot leadMailbox itself opened from (cfg.coordinator()),
// not the live config.get() — coordinator wiring is already boot-time-fixed (see
// leadMailbox above), so peers follows the same rule rather than half hot-reloading.
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
// fleetd #480 Unit C: the executor behind fleet_handover — null whenever
// leadRollover: is not configured (see the leadRollover local above).
leadRollover);
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
// created and no thread runs. It reads the SAME live lead supplier the injector's
// deliverability gate does, so a lead found by the tab scan after startup is reachable
// without a restart.
final LeadCoordLoop leadCoordLoop;
final ScheduledExecutorService leadCoordSchedulerRef;
if (leadMailbox != null) {
var leadCoordScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-leadcoord-").unstarted(r));
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
LEAD_COORD_INTERVAL_MS);
leadCoordLoop.start();
leadCoordSchedulerRef = leadCoordScheduler;
} else {
leadCoordLoop = null;
leadCoordSchedulerRef = null;
}
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
final ConfigWatcher configWatcher;
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
configWatcher.start();
} else {
configWatcher = null;
}
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (leadCoordLoop != null) leadCoordLoop.close(); // CB-637: stop delivering peer-lead messages
if (leadCoordSchedulerRef != null) leadCoordSchedulerRef.shutdownNow();
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
// Idle-sleep guard: release unconditionally, even though sessions.close() above already
// drained every session (and each release already drove the live count to 0, which
// releases the guard's assertion on its own) — this is the backstop for a drain that was
// itself interrupted or threw, so no caffeinate child ever outlives the daemon.
if (idleSleepGuard != null) idleSleepGuard.close();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
// CB-637: the coordination connection goes with it — after the loop that reads it has
// stopped, so no tick can be mid-ack against a closed channel.
if (leadMailbox != null) {
try {
leadMailbox.close();
} catch (Exception e) {
log.debug("lead mailbox close: {}", e.toString());
}
}
router.close();
}));
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
// fleetd #111: live (re-read-per-request) memberCredentials view for GET /member-credentials —
// same hot-reload shape as the memberCredentials supplier passed to ClaudeCodeLauncher above.
// fleetd #297: quarantineSource/outageSource are the SAME instances passed to FleetMcp above —
// GET /profiles must report the identical quarantine/cool-off facts as fleet_profiles.
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics, deliverable,
() -> MemberCredentialPolicyView.of(config.get().memberCredentials()),
quarantineSource, outageSource, loopHealth).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("fleetd listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
// fleetd #612 A-gaps (gap 2): everything from here on — including the two post-validation
// reports that used to run inline right here (reportRoleFallbackGaps,
// assertChartersNameOnlyRegisteredTools) — now lives in FleetdAssembly.assembleAndStart,
// built against a real ResourcePorts. Unit A originally moved only the socket/broker/HTTP
// composition and left those two calls here, between validateAll() and the assembly call —
// outside the boundary FleetdAssemblyLifecycleTest drives, so deleting either call
// compiled clean and left the whole suite green. Moving the boundary to start immediately
// after validateAll() (this line) puts both back under test, in the same relative order,
// before either one does any I/O — see FleetdAssembly's javadoc for the full boot-order
// contract this preserves exactly.
FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ResourcePorts.system());
}
/**
@@ -1576,6 +1003,109 @@ public final class Fleetd {
};
}
/**
* fleetd #602 gauge-wiring: per-lead-name factory for {@link FleetMcp.LeadConfigDirSource} — the
* {@code CLAUDE_CONFIG_DIR} the named lead's own profile runs on, so {@code LeadContextGauge}
* reads the transcript directory that lead's Claude Code process actually writes to, not always
* the built-in {@code <user.home>/.claude} default.
*
* <p>Follows the same {@code fleet.leaders.<name>.profile} link {@link #leadSeatLookup} already
* uses to find a lead's profile, one step further to that profile's own {@code configDir:}. A
* lead entry that names no {@code profile:}, or whose named profile is not configured, or whose
* profile sets no {@code configDir:} override, returns {@code null} — {@code LeadContextGauge}
* then falls back to its own default, exactly as before this ticket.
*
* @param profiles the live profile map, normally {@code () -> config.get().profiles()} in
* {@code main} — read live, like every other {@code configDir} lookup, so a
* config reload takes effect on the very next {@code fleet_list} call
* @param leaders {@code fleet.leaders}, read once at startup like {@link #leadSeatLookup}'s own
* {@code leaders} parameter — passed as a plain map, never re-read from
* {@code config.get()}
*/
static Function<String, String> leadConfigDirLookup(Supplier<Map<String, FleetConfig.Profile>> profiles,
Map<String, FleetConfig.Leader> leaders) {
return leadName -> {
FleetConfig.Leader lead = leaders.get(leadName);
if (lead == null || lead.profile() == null || lead.profile().isBlank()) {
return null;
}
FleetConfig.Profile leadProfile = profiles.get().get(lead.profile());
return leadProfile == null ? null : leadProfile.configDir();
};
}
/**
* fleetd #602 gauge-wiring follow-up (PR #606 review comment 17353): {@code main} used to build
* {@code new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(...))} inline, with nothing a test
* could call directly. Measured on that shape: replacing the whole expression with {@code
* FleetMcp.LeadConfigDirSource.none()} at the call site compiled with 0 errors and left the full
* 1822-test suite green — the daemon could be changed to always report every lead's context as
* {@code UNKNOWN}, forever, while every test stayed green. That is the same hand-built-vs-wired
* shape as fleetd #561/#248/#426/#562 ({@link #loopHealthSource}).
*
* <p>The fix extracts the inline {@code new} into this factory, in the same style as {@link
* #loopHealthSource}/{@link #capacitySource}/{@link #healthCoverageSource} — which is exactly
* what makes it directly callable from {@code FleetdLeadConfigDirSourceWiringTest}. That test
* calls this factory with real {@link FleetConfig.Profile}/{@link FleetConfig.Leader} fixtures and
* asserts the returned source resolves a real {@code configDir} — a property that would be false
* if this method's body were mutated to {@code return FleetMcp.LeadConfigDirSource.none();}.
*/
static FleetMcp.LeadConfigDirSource leadConfigDirSource(Supplier<Map<String, FleetConfig.Profile>> profiles,
Map<String, FleetConfig.Leader> leaders) {
return new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(profiles, leaders));
}
/**
* fleetd #609: per-terminal factory for {@link LeadHeartbeatLoop.LeadContextSource} — the lead's
* own {@link LeadContextGauge} reading, so the heartbeat loop can tell an idle, HIGH-context lead
* to consider a handover.
*
* <p>Three hops, each degrading to {@link LeadContextGauge.Reading#unknown()} rather than
* throwing, since a herdr hiccup or an unrecognised terminal must never kill the heartbeat's own
* tick: {@code liveLeadTerminals} (terminal id → lead name, the same live supplier {@link
* #leadSeatLookup} and {@code LeadCoordLoop} already read) → {@code configDirForLeadName} (that
* lead's {@code configDir}, normally {@link #leadConfigDirLookup}'s return) → {@code agents.get}
* for the live {@link Agent#sessionId()}/{@link Agent#agentType()} the gauge itself needs.
*
* @param gauge the {@link LeadContextGauge} instance to read through — shares its
* cache across every call this factory's function makes
* @param agents the {@link AgentControl} instance that reaches the LEAD's pane
* (not {@code memberAgents}), normally {@code router.leadAgents()}
* @param liveLeadTerminals terminal id → lead name for every CURRENTLY recognised lead
* @param configDirForLeadName lead name → {@code configDir}, normally {@link
* #leadConfigDirLookup}'s return
*/
static Function<String, LeadContextGauge.Reading> leadContextLookup(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
return terminal -> {
String leadName = liveLeadTerminals.get().get(terminal);
if (leadName == null) {
return LeadContextGauge.Reading.unknown();
}
String configDir = configDirForLeadName.apply(leadName);
Agent live;
try {
live = agents.get(terminal);
} catch (RuntimeException e) {
return LeadContextGauge.Reading.unknown();
}
return gauge.read(configDir, live.sessionId(), live.agentType());
};
}
/**
* fleetd #609: wraps {@link #leadContextLookup} into a {@link LeadHeartbeatLoop.LeadContextSource}
* — the same hand-built-vs-wired shape as {@link #leadConfigDirSource}/{@link #loopHealthSource}/
* {@link #capacitySource}/{@link #healthCoverageSource}. Extracted so a test can call the exact
* factory {@code main} calls, rather than only a lookup nothing in {@code main} is proven to use
* (see {@code FleetdLeadConfigDirSourceWiringTest}'s javadoc for the measured gap this shape closes).
*/
static LeadHeartbeatLoop.LeadContextSource leadContextSource(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
return new LeadHeartbeatLoop.LeadContextSource(
leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName));
}
/**
* fleetd #248 / fleetd#201 Unit 5: package-private factory for the per-target backend-error
* pattern lookup {@link CompletionResolver} classifies a pane scrape against. Closes over the
@@ -1685,10 +1215,18 @@ public final class Fleetd {
ReplyInbox open(String uri, int prefetch);
}
/** Injection seam for {@link #openLeadMailbox}: production binds {@link LeadMailbox#open}. */
/**
* Injection seam for {@link #openLeadMailbox}: production binds {@link LeadMailbox#open}.
*
* <p>fleetd #612 A-gaps (gap 1): returns {@link LeadChannelHandle}, not the concrete {@link
* LeadMailbox}, so a test can supply a fake closeable channel instead of a real broker
* connection — see {@link LeadChannelHandle}'s own javadoc for why the narrower type existed
* and what widening it to add {@code close()} costs (nothing: {@code LeadMailbox} already
* implements it).
*/
@FunctionalInterface
interface LeadMailboxOpener {
LeadMailbox open(String uri, String selfCoordId, int prefetch);
LeadChannelHandle open(String uri, String selfCoordId, int prefetch);
}
/**
@@ -1711,7 +1249,7 @@ public final class Fleetd {
* fleet still works exactly as it did before this feature existed.</li>
* </ul>
*/
static LeadMailbox openLeadMailbox(FleetConfig.Coordinator coordinator, Map<String, String> env,
static LeadChannelHandle openLeadMailbox(FleetConfig.Coordinator coordinator, Map<String, String> env,
LeadMailboxOpener opener) {
if (coordinator == null) {
return null; // opt-in: nothing configured, nothing to say
@@ -1731,7 +1269,7 @@ public final class Fleetd {
return null;
}
try {
LeadMailbox mailbox = opener.open(uri, coordinator.selfId(), coordinator.prefetchOrDefault());
LeadChannelHandle mailbox = opener.open(uri, coordinator.selfId(), coordinator.prefetchOrDefault());
log.info("lead coordination: ON as coord-id {} (prefetch={})",
coordinator.selfId(), coordinator.prefetchOrDefault());
return mailbox;
@@ -2059,16 +1597,72 @@ public final class Fleetd {
}
/**
* fleetd #474: the one place both the startup call (right after {@code cfg.validateAll()} in
* {@link #main}) and the reload call (wired into {@code config}'s {@code extraValidation} above,
* via a method reference to this method) go through, so the two can never drift into checking
* different things. Extracted only to give {@link ConfigRef}'s {@code Consumer<FleetConfig>}
* hook a {@code FleetConfig -> void} shape to bind to — {@link CharterToolSurface} itself still
* takes the raw charter map and knows nothing about {@code ConfigRef} or {@code Fleetd}.
* fleetd #613: {@code FleetConfig.candidateProfiles(MemberRole)} (FleetConfig.java:1827) and
* {@code CompositePeerLauncher.poolFor} (CompositePeerLauncher.java:598-603) both fall back to
* <em>every</em> configured profile when a role has no {@code fleet.<role>s:} pool — a
* deliberate "unconstrained" behaviour, kept unchanged here, that lets a config with only
* {@code profiles:} and no {@code fleet:} block still spawn. That fallback is silent, and on
* the host that opened this ticket it widened an unqualified {@code hunter} spawn to all 8
* configured profiles and picked {@code local} as the resolved first choice — a profile every
* other pool on that same config gives weight 0 to. Report it once at boot instead, naming both
* how many profiles the gap opens onto and the exact first choice, since the first choice (not
* the pool size) is what actually surprised the operator.
*
* <p>A missing {@code fleet.charters.<role>:} entry is reported separately: the member still
* runs, but with only the launcher's own reply charter and no role contract. Unlike the pool
* gap this already has a per-spawn instrument ({@code HerdrPeerLauncher.logCharterReceipt},
* {@code SessionManager}'s {@code charterSource} roster field) — this boot line is the same
* information surfaced once, up front, rather than discovered per member later.
*
* <p>Never refuses to start over either gap — both are legitimate configurations, and this is a
* report, not a validation. Package-private so a test can call it directly and capture the log
* via a {@link ch.qos.logback.core.read.ListAppender}, the same pattern {@link
* #reportExhaustedPatternGap} and {@link #reportMemberCredentialsGap} already use.
*/
static void reportRoleFallbackGaps(FleetConfig cfg) {
List<String> poolGaps = new ArrayList<>();
List<String> charterGaps = new ArrayList<>();
int profileCount = cfg.profiles().size();
for (MemberRole role : MemberRole.values()) {
boolean hasPool = cfg.fleet() != null && !cfg.fleet().profilesFor(role).isEmpty();
if (!hasPool) {
String firstChoice = cfg.defaultProfileFor(role);
poolGaps.add(role.wireName() + " (may land on any of " + profileCount
+ " profile(s), first choice "
+ (firstChoice == null ? "none — no profiles configured" : "'" + firstChoice + "'")
+ ")");
}
String charter = cfg.fleet() == null ? null : cfg.fleet().charterFor(role);
if (charter == null || charter.isBlank()) {
charterGaps.add(role.wireName());
}
}
if (!poolGaps.isEmpty()) {
log.info("role fallback: no fleet.<role>s: pool for {} — an unqualified spawn of that "
+ "role falls back to every configured profile (deliberate; see "
+ "FleetConfig#candidateProfiles)", poolGaps);
}
if (!charterGaps.isEmpty()) {
log.info("role fallback: no fleet.charters: entry for {} — that role runs with only "
+ "the launcher's reply charter, no role contract", charterGaps);
}
}
/**
* fleetd #474: the one place both the startup call (fleetd #612 A-gaps gap 2: right after
* {@code cfg.validateAll()} in {@link FleetdAssembly#assembleAndStart}, immediately after
* {@code main} hands off to it — see that method's javadoc) and the reload call (wired into
* {@code config}'s {@code extraValidation} above, via a method reference to this method) go
* through, so the two can never drift into checking different things. Extracted only to give
* {@link ConfigRef}'s {@code Consumer<FleetConfig>} hook a {@code FleetConfig -> void} shape to
* bind to — {@link CharterToolSurface} itself still takes the raw charter map and knows nothing
* about {@code ConfigRef} or {@code Fleetd}.
*
* <p>Package-private so a test can call it directly the same way the other startup-report
* helpers above are tested, without needing to drive {@link #main} for a unit-level check;
* {@code FleetdStartupValidationTest} proves the startup call site, and {@code
* {@code FleetdStartupValidationTest} proves the startup call site by driving {@code main}
* itself end to end (the throw still happens before {@code main} reaches any real socket or
* broker work, since the assembly runs this before either), and {@code
* FleetdConfigRefCharterToolSurfaceWiringTest} — by constructing {@code ConfigRef} with this
* exact method reference, the same way {@code main} does above — proves the reload call site.
* {@code dev.ltms.fleet.config.ConfigRefTest} pins the same reload behaviour too, through an
@@ -2110,13 +1704,16 @@ public final class Fleetd {
record HerdrAwaitOutcome(HerdrWaitResult result, long elapsedNanos) {}
/**
* The real per-poll wait {@link #main} passes to {@link #awaitHerdr}: sleep
* {@link #HERDR_WAIT_POLL_MILLIS}, and on interruption re-set the thread's interrupt flag
* The real per-poll wait {@link FleetdAssembly#assembleAndStart} passes to {@link #awaitHerdr}:
* sleep {@link #HERDR_WAIT_POLL_MILLIS}, and on interruption re-set the thread's interrupt flag
* rather than throwing — {@link #awaitHerdr} detects an interruption by checking {@link
* Thread#isInterrupted()} right after this returns, so a poller that swallowed the flag
* instead of restoring it would make that check silently miss the interruption.
*
* <p>Package-private (fleetd #612 Unit A), not {@code private}: {@code FleetdAssembly}, a
* different class in this package, needs a {@code Runnable} reference to this exact method.
*/
private static void sleepHerdrPoll() {
static void sleepHerdrPoll() {
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
@@ -0,0 +1,536 @@
package dev.ltms.fleet;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.ConfigWatcher;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.health.FleetHealthMonitor;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
import dev.ltms.fleet.inject.BackendErrorPatternLookup;
import dev.ltms.fleet.inject.BackendErrorSink;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.inject.TurnListener;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.mcp.LsofPeerPidLookup;
import dev.ltms.fleet.mcp.LsofProcessCwdLookup;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.member.HerdrPeerLauncher;
import dev.ltms.fleet.member.MemberCredentialPolicyView;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.msg.ReplyPushLoop;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.power.CaffeinateSleepAssertionMechanism;
import dev.ltms.fleet.power.IdleSleepGuard;
import dev.ltms.fleet.rest.FleetApp;
import dev.ltms.fleet.session.GitWorktrees;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.SessionReaper;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
/**
* fleetd #612 Unit A: the real boot assembly, extracted out of {@code Fleetd.main} so a test can
* drive it directly. {@link #assembleAndStart} is <em>the same statements {@code main} used to run
* inline</em>, in the same order, against a real {@link ResourcePorts} in production and a fake one
* in a test — see {@code FleetdAssemblyLifecycleTest}. {@code Fleetd.main} keeps config loading and
* {@code cfg.validateAll()}; everything from immediately after that call onward moved here —
* including, since fleetd #612 A-gaps (gap 2), the two post-validation reports ({@code
* reportRoleFallbackGaps}, {@code assertChartersNameOnlyRegisteredTools}) that Unit A originally
* left behind in {@code main}. Those two calls do no I/O themselves, but leaving them outside this
* method meant deleting either one compiled clean and left the whole suite green — nothing drove
* {@code main} itself, so nothing could notice. They run first here, in the same relative order,
* before the herdr socket or anything else that touches the outside world.
*
* <p><strong>Construction and start order is preserved exactly, on purpose.</strong> This is not
* rebuilt into "construct everything, then start everything" — that would change boot timing. The
* order recorded before any code moved (see the ticket and {@code FleetdAssemblyLifecycleTest}):
* {@code SessionReaper} starts first (if {@code lifecycle.idleTtlSeconds} is configured), then
* {@link StatusPoller}, then the optional {@link LeadHeartbeatLoop} and {@link FleetHealthMonitor},
* then the optional {@link LeadCoordLoop} and {@link ConfigWatcher}, and the HTTP server starts
* last of all. The close order (see {@link FleetdRuntime#close()}) is the mirror the original
* shutdown hook always used.
*
* <p><strong>One statement could not move without reordering startup.</strong> {@code Fleetd.main}
* registered its shutdown hook <em>before</em> building the Javalin {@code FleetApp} — the hook
* itself never touched {@code app} (it still doesn't; see {@link FleetdRuntime#close()}), but the
* hook needs a {@link FleetdRuntime} to close over, and the ticket asks that runtime to also own
* {@code FleetApp}. Building the runtime before the app exists and mutating it afterward (via
* {@link FleetdRuntime#attachApp}) preserves the exact original order — hook registered, then app
* built, then HTTP started — without moving the app's construction earlier or the hook's
* registration later. That is the one seam this ticket did not get to pin any other way.
*
* <p><strong>No inert variant.</strong> Deliberately, there is no overload of this method that
* accepts a smaller/optional {@link ResourcePorts} or defaults one internally. A future edit that
* wants to skip {@code FleetdAssembly} entirely and build its own graph is still possible — no
* static analysis stops that — but it cannot do so by quietly swapping this call for an inert
* substitute that still compiles, because none exists.
*/
final class FleetdAssembly {
private static final Logger log = LoggerFactory.getLogger(Fleetd.class);
/** CB-637: how often the lead coordination loop looks for peer messages — see {@code Fleetd}'s own constant. */
private static final long LEAD_COORD_INTERVAL_MS = 3_000L;
private FleetdAssembly() {
}
static FleetdRuntime assembleAndStart(AssemblyInputs inputs, ResourcePorts ports) {
FleetConfig cfg = inputs.cfg();
ConfigRef config = inputs.config();
SubscriptionGuard guard = inputs.guard();
// fleetd #612 A-gaps (gap 2): moved in from Fleetd.main, immediately after cfg.validateAll()
// there — the exact point main used to call these two, and still the first thing this
// method does, before any socket or broker work below. See this class's javadoc and each
// method's own for why they run here rather than in FleetConfig#validateAll() itself.
Fleetd.reportRoleFallbackGaps(cfg);
Fleetd.assertChartersNameOnlyRegisteredTools(cfg);
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
HerdrClient herdr = ports.connectHerdr(socket);
HerdrClient memberHerdr = cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank()
? ports.connectHerdr(Path.of(cfg.memberHerdrSocket()))
: herdr;
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
target -> leadsRef.get().get().containsKey(target));
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, FleetConfig.Profile> claudeProfiles = new LinkedHashMap<>();
Map<String, FleetConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
cfg.profiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// fleetd #175: the daemon's real ExhaustionSink can only be built once `sessions` exists
// (below), but `sessions` needs `workers`, which needs the adapters built right here — a
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
// pointed at the real one once it exists.
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
ExhaustionSink forwardingExhaustionSink = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(Fleetd.claudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
claudeProfiles, cfg, config));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(Fleetd.openCodeLauncher(router.memberAgents(), router.memberSpaces(),
opencodeProfiles, cfg, config, forwardingExhaustionSink));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
BackendQuarantine quarantine = BackendQuarantine.withEscalation(ports.nanoClock(),
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
// below (written on a classified backend error).
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(ports.nanoClock());
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.effectiveDefaultProfile(),
config,
profileName -> liveCountRef.get().apply(profileName),
quarantine,
outagePolicy);
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into.
log.info("model gate (fleetd #422): {}", Fleetd.modelGateCoverageLine(workers.modelGateState()));
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
// exists. Wait, then degrade rather than die: serving with /healthz reporting "degraded" is
// strictly more useful than exiting.
Fleetd.HerdrAwaitOutcome herdrOutcome = Fleetd.awaitHerdr(herdr, ports.nanoClock(), Fleetd::sleepHerdrPoll);
boolean herdrUp = Fleetd.logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers,
new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup(), cfg.memberSkills()),
ports.nanoClock(), contextCap, clearAfterTurn);
liveCountRef.set(profileName -> Fleetd.liveSessionCount(sessions.roster(), profileName));
// Idle-sleep guard: hold an OS-level assertion against idle sleep while at least one
// member is live. No-op (never constructed) off macOS or when idleSleepGuard.enabled is
// explicitly false; the mechanism itself is additionally a no-op if 'caffeinate' cannot be
// started, so this can never fail a spawn, a release, or startup.
boolean idleSleepGuardEnabled = cfg.idleSleepGuard() == null || cfg.idleSleepGuard().isEnabled();
final IdleSleepGuard idleSleepGuard;
if (idleSleepGuardEnabled) {
idleSleepGuard = new IdleSleepGuard(new CaffeinateSleepAssertionMechanism(), sessions::size);
sessions.onAcquire(_ -> idleSleepGuard.recheck());
sessions.onRelease(_ -> idleSleepGuard.recheck());
} else {
idleSleepGuard = null;
}
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled. FIRST of the
// recurring background loops to start — see this class's own javadoc for the full order.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
// loop's nudges, which need a single destination — so it keeps the legacy pin.
Map<String, String> leadTerminals = cfg.leaderTerminals();
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531/CB-579: discover leads by the tab labels the operator writes, one scanner per
// configured lead's own exact `tab:` label.
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
tabToName.keySet(), scanIntervalSeconds);
} else {
leads = () -> leadTerminals;
}
leadsRef.set(leads);
// CB-558: start any declared lead that is not already running. After the scanner is built,
// and only when herdr answered — the launcher's whole safety property is that it can count
// live leads first, and must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
}
// CB-548: config-declared architect slots. Nothing here spawns a slot; the terminal → slot
// binding is owned by the registry and empty at startup.
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
+ "spawn lifecycle binds a live terminal to it)",
members.slots().size(), members.slots().keySet());
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
LiveExhaustedPatterns liveExhaustedPatterns = Fleetd.liveExhaustedPatterns(config);
ExhaustedPatternLookup exhaustedPatterns = Fleetd.exhaustedPatternLookup(sessions::roster, liveExhaustedPatterns);
// The startup coverage line still reports the boot-time snapshot only.
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
.filter(e -> e.getValue().hasExhaustedPattern())
.map(Map.Entry::getKey)
.collect(Collectors.toCollection(LinkedHashSet::new));
log.info("backend-exhausted classification (CB-578 stage A): {}",
Fleetd.exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
// rather than handing it back as a real answer.
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasErrorPattern()) {
errorPatternsByProfile.put(name, Pattern.compile(profile.errorPattern()));
}
});
BackendErrorPatternLookup backendErrorPatterns = Fleetd.backendErrorPatternLookup(sessions::roster,
errorPatternsByProfile);
log.info("backend-error classification (fleetd #201 Unit 5): {}",
Fleetd.errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL, so a profile sharing that credential is refused too, not just the one that
// happened to report it.
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
// now that `sessions` exists to resolve target -> session -> profile.
ExhaustionSink exhaustionSink = Fleetd.publishExhaustionSink(exhaustionSinkRef, sessions, config,
quarantine, quarantineReasonByCredential, cfg);
// fleetd #201 Unit 5: the production BackendErrorSink needs `pushLoop` (built further below,
// after `sessions`) to tell a lead about an incident or an unmapped target — the same
// construction-order cycle `exhaustionSinkRef` breaks above, broken the same way: a mutable
// holder set once `pushLoop` exists, read lazily from inside the lambda built here.
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>();
BackendErrorSink backendErrorSink = Fleetd.backendErrorSink(sessions, () -> config.get().profiles(),
outagePolicy, pushLoopRef::get);
AgentControl agents = router.memberAgents();
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns,
exhaustionSink, backendErrorPatterns, backendErrorSink, ports.nanoClock(),
Fleetd.worktreeBranchLookup(sessions::roster));
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
MemberPresence presence = sessions.asPresence();
TurnListener turnListener = Fleetd.turnListener(completion, sessions);
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads);
// fleetd #556: registration is wired directly to `completion`, not folded into the
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
// listener) throwing, regardless of call order.
Injector injector = new Injector(router, turnListener, deliverable,
presence::forget, Fleetd.turnRegistrar(completion));
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start(); // SECOND of the recurring background loops to start, after the reaper.
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
// unusable), fleetd stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox = Fleetd.selectReplyInbox(cfg.broker(), ports.environment(),
ports.replyInboxOpener());
// CB-637: this daemon's lead-to-lead mailbox on the SHARED coordination vhost — a separate
// broker from the reply inbox by design. Absent a coordinator: block this is null and every
// lead path below is simply not wired, exactly the behaviour before this ticket. It owns a
// broker connection, so keep the reference for the ordered shutdown hook.
final LeadChannelHandle leadMailbox = Fleetd.openLeadMailbox(cfg.coordinator(), ports.environment(),
ports.leadMailboxOpener());
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs.
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
+ "It still works, and is still the fallback nudge destination when a restart "
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open fleet_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = ports.newScheduler("bridge-push-");
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
// fleetd #201 Unit 5: point the forwarding holder captured by the backendErrorSink lambda
// above at the real push loop, now that it exists.
pushLoopRef.set(pushLoop);
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
// THIRD of the recurring background loops to start (optional).
final LeadHeartbeatLoop heartbeat;
var heartbeatScheduler = ports.newScheduler("bridge-heartbeat-");
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
var leadContextGauge = new LeadContextGauge();
// fleetd #621: the context-high notice's own wording must track this same effective
// value — LeadRollover.confirm(...) already gates the roll on it (LeadRollover.java:480),
// and absent `leadRollover:` entirely the roll is unusable regardless (NOT_CONFIGURED),
// so `true` (the FleetConfig.LeadRollover default) is the safe, byte-identical fallback.
// Carried in from Fleetd.main when #612 Unit A merged main: #622 added this line to the
// block Unit A had already moved here, so the merge would otherwise have silently
// dropped it — with a fully green suite, because nothing pins it (see the follow-up issue).
boolean requireOperatorConfirm = cfg.leadRollover() == null || cfg.leadRollover().requireOperatorConfirm();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, ports.nanoClock(),
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics,
Fleetd.leadContextSource(leadContextGauge, router.leadAgents(), leads,
Fleetd.leadConfigDirLookup(() -> config.get().profiles(), leaders)),
Boolean.TRUE.equals(hb.contextHighNudge()), requireOperatorConfirm);
heartbeat.start();
} else {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads);
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
// FOURTH of the recurring background loops to start (optional).
final FleetHealthMonitor healthMonitor;
var healthScheduler = ports.newScheduler("bridge-health-");
if (cfg.health() != null && cfg.health().isEnabled()) {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
ports.nanoClock(), ports.wallClockNanos(),
cfg.health().intervalOrDefault(),
cfg.health().workingSuspectAfterOrDefault(), Fleetd.healthFailTarget(messages));
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
log.warn("fleet health: {} (no notification sink configured)", coverage);
} else {
log.info("fleet health: {}", coverage);
}
healthMonitor.start();
} else {
healthMonitor = null;
healthScheduler.shutdownNow();
}
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it.
sessions.onRelease(Fleetd.releaseCleanup(messages, replyInbox, primaryRegistry));
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = ports.environment().get(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting fleetd");
}
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
FleetMcp.QuarantineSource quarantineSource = Fleetd.quarantineSource(config, quarantine,
liveExhaustedPatterns, quarantineReasonByCredential);
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, outagePolicy);
FleetMcp.LoopHealthSource loopHealth = Fleetd.loopHealthSource(poller, reaper);
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, FleetMcp.AuthorizationMode.ENFORCED, metrics,
Fleetd.capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
Fleetd.healthCoverageSource(config),
loopHealth,
quarantineSource,
leadMailbox,
outageSource,
new FleetMcp.LeadSeatSource(Fleetd.leadSeatLookup(() -> config.get().profiles(), leaders, leads)),
Fleetd.leadConfigDirSource(() -> config.get().profiles(), leaders),
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
leadRollover);
// CB-637: the receive half. Only constructed when a lead mailbox actually opened.
final LeadCoordLoop leadCoordLoop;
final ScheduledExecutorService leadCoordSchedulerRef;
if (leadMailbox != null) {
var leadCoordScheduler = ports.newScheduler("bridge-leadcoord-");
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
LEAD_COORD_INTERVAL_MS);
leadCoordLoop.start(); // FIFTH of the recurring background loops to start (optional).
leadCoordSchedulerRef = leadCoordScheduler;
} else {
leadCoordLoop = null;
leadCoordSchedulerRef = null;
}
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed.
// SIXTH of the recurring background loops to start (optional).
final ConfigWatcher configWatcher;
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
configWatcher.start();
} else {
configWatcher = null;
}
// CB-303 part 3: single ordered shutdown hook — see FleetdRuntime#close() for the statements
// this used to be. Registered here, at the exact point `main` used to register it: after
// configWatcher, before the Javalin app exists (see this class's own javadoc for why).
FleetdRuntime runtime = new FleetdRuntime(cfg, sessions, router, poller, messages, pushLoop, heartbeat,
leadCoordLoop, leadCoordSchedulerRef, healthMonitor, configWatcher, mcp, reaper, idleSleepGuard,
replyInbox, leadMailbox, completion, injector);
ports.addShutdownHook(runtime::close);
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics, deliverable,
() -> MemberCredentialPolicyView.of(config.get().memberCredentials()),
quarantineSource, outageSource, loopHealth).build();
runtime.attachApp(app);
ports.startHttp(app, cfg.bind().host(), cfg.bind().port()); // HTTP starts LAST, always.
log.info("fleetd listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
return runtime;
}
}
@@ -0,0 +1,166 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigWatcher;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.health.FleetHealthMonitor;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.msg.ReplyPushLoop;
import dev.ltms.fleet.power.IdleSleepGuard;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.SessionReaper;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.ScheduledExecutorService;
/**
* fleetd #612 Unit A: the assembled daemon. {@link FleetdAssembly#assembleAndStart} builds exactly
* one of these and hands it to {@code ResourcePorts.addShutdownHook}; {@link #close()} is the
* single ordered shutdown, moved verbatim out of {@code Fleetd.main}'s old shutdown-hook
* {@code Thread} body — same statements, same order, see that method's javadoc.
*
* <p><strong>Owns the real production objects, never a copy.</strong> Every package-private
* accessor below returns the identical instance the running daemon is using. That is the entire
* point of this class existing (see fleetd #612's problem statement): a test that inspected a
* snapshot built alongside the real objects could pass while production silently received
* something else — the exact shape of the #602/#606 defect this ticket exists to stop from
* recurring one call site at a time. Nothing here is rebuilt or copied for a test's benefit.
*/
final class FleetdRuntime implements AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(FleetdRuntime.class);
private final FleetConfig cfg;
private final SessionManager sessions;
private final HerdrRouter router;
private final StatusPoller poller;
private final MessageService messages;
private final ReplyPushLoop pushLoop;
private final LeadHeartbeatLoop heartbeat; // nullable — leadHeartbeat: opt-in
private final LeadCoordLoop leadCoordLoop; // nullable — coordinator: opt-in
private final ScheduledExecutorService leadCoordScheduler; // nullable, paired with leadCoordLoop
private final FleetHealthMonitor healthMonitor; // nullable — health.enabled opt-in
private final ConfigWatcher configWatcher; // nullable — configReload.enabled opt-in
private final FleetMcp mcp;
private final SessionReaper reaper; // nullable — lifecycle.idleTtlSeconds opt-in
private final IdleSleepGuard idleSleepGuard; // nullable — idleSleepGuard.enabled: false
private final ReplyInbox replyInbox;
private final LeadChannelHandle leadMailbox; // nullable — coordinator: opt-in
private final CompletionResolver completion;
private final Injector injector;
/**
* Not final: {@code Fleetd.main}'s shutdown hook was registered <em>before</em> the Javalin
* {@code FleetApp} was built and started — see {@link FleetdAssembly#assembleAndStart}'s javadoc
* for why that order could not be preserved AND have this constructor take {@code app}. {@link
* #attachApp} is called immediately after the real app is built, still before HTTP starts
* listening, so this is set long before any test or caller could observe it unset.
*/
private Javalin app;
FleetdRuntime(FleetConfig cfg, SessionManager sessions, HerdrRouter router, StatusPoller poller,
MessageService messages, ReplyPushLoop pushLoop, LeadHeartbeatLoop heartbeat,
LeadCoordLoop leadCoordLoop, ScheduledExecutorService leadCoordScheduler,
FleetHealthMonitor healthMonitor, ConfigWatcher configWatcher, FleetMcp mcp,
SessionReaper reaper, IdleSleepGuard idleSleepGuard, ReplyInbox replyInbox,
LeadChannelHandle leadMailbox, CompletionResolver completion, Injector injector) {
this.cfg = cfg;
this.sessions = sessions;
this.router = router;
this.poller = poller;
this.messages = messages;
this.pushLoop = pushLoop;
this.heartbeat = heartbeat;
this.leadCoordLoop = leadCoordLoop;
this.leadCoordScheduler = leadCoordScheduler;
this.healthMonitor = healthMonitor;
this.configWatcher = configWatcher;
this.mcp = mcp;
this.reaper = reaper;
this.idleSleepGuard = idleSleepGuard;
this.replyInbox = replyInbox;
this.leadMailbox = leadMailbox;
this.completion = completion;
this.injector = injector;
}
/** See the {@link #app} field doc for why this is a late-bound setter rather than a constructor arg. */
void attachApp(Javalin app) {
this.app = app;
}
// --- package-private accessors: the SAME instances this runtime owns, never a copy ---------
SessionManager sessions() { return sessions; }
HerdrRouter router() { return router; }
StatusPoller poller() { return poller; }
MessageService messages() { return messages; }
ReplyPushLoop pushLoop() { return pushLoop; }
LeadHeartbeatLoop heartbeat() { return heartbeat; }
LeadCoordLoop leadCoordLoop() { return leadCoordLoop; }
FleetHealthMonitor healthMonitor() { return healthMonitor; }
ConfigWatcher configWatcher() { return configWatcher; }
FleetMcp mcp() { return mcp; }
SessionReaper reaper() { return reaper; }
IdleSleepGuard idleSleepGuard() { return idleSleepGuard; }
ReplyInbox replyInbox() { return replyInbox; }
LeadChannelHandle leadMailbox() { return leadMailbox; }
CompletionResolver completion() { return completion; }
Injector injector() { return injector; }
Javalin app() { return app; }
/**
* CB-303 part 3: the single ordered shutdown. Moved verbatim out of {@code Fleetd.main}'s
* shutdown-hook {@code Thread} body (fleetd #612 Unit A) — drain sessions first while herdr is
* still open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close
* herdr last, exactly as before. {@code Fleetd.main} never calls this directly; it hands the
* reference to {@code ResourcePorts.addShutdownHook} the moment this runtime exists, the same
* point it used to register the hook {@code Thread} itself.
*/
@Override
public void close() {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (leadCoordLoop != null) leadCoordLoop.close(); // CB-637: stop delivering peer-lead messages
if (leadCoordScheduler != null) leadCoordScheduler.shutdownNow();
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
// Idle-sleep guard: release unconditionally, even though sessions.close() above already
// drained every session (and each release already drove the live count to 0, which
// releases the guard's assertion on its own) — this is the backstop for a drain that was
// itself interrupted or threw, so no caffeinate child ever outlives the daemon.
if (idleSleepGuard != null) idleSleepGuard.close();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
// CB-637: the coordination connection goes with it — after the loop that reads it has
// stopped, so no tick can be mid-ack against a closed channel.
if (leadMailbox != null) {
try {
leadMailbox.close();
} catch (Exception e) {
log.debug("lead mailbox close: {}", e.toString());
}
}
router.close();
}
}
@@ -0,0 +1,67 @@
package dev.ltms.fleet;
import dev.ltms.fleet.herdr.HerdrClient;
import io.javalin.Javalin;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
/**
* fleetd #612 Unit A: every boot-time side effect {@link FleetdAssembly#assembleAndStart} performs
* that a real daemon must do for real, and a test must not — read the process environment, connect
* a herdr client, open a broker (the reply inbox, the lead mailbox), read a clock, start a
* background scheduler, register the JVM shutdown hook, and bind the HTTP server.
*
* <p>{@link #system()} is the one production implementation ({@code SystemResourcePorts}), wired
* verbatim from what {@code Fleetd.main} used to call directly at each of these call sites. A test
* builds its own implementation instead of receiving an inert default from this interface —
* deliberately, there is no {@code ResourcePorts.none()}. fleetd #612's whole problem is a call
* site quietly swapped for an inert variant that still compiles; adding one here, even for tests,
* would hand a future edit to {@code FleetdAssembly} the exact compiling substitute this ticket
* exists to rule out. A test that wants an inert resource writes its own fake and owns that
* decision explicitly.
*/
public interface ResourcePorts {
/** The process environment. Production: {@link System#getenv()}. */
Map<String, String> environment();
/**
* Connect a herdr client bound to {@code socketPath}. Production returns a real
* {@code UnixSocketHerdrClient} — connection-per-call, so this itself never touches the socket.
*/
HerdrClient connectHerdr(Path socketPath);
/** The reply-inbox AMQP opener (CB-307). Production: {@link Fleetd#replyInboxOpener()}. */
Fleetd.AmqpOpener replyInboxOpener();
/** The lead-mailbox AMQP opener (CB-637). Production: {@link Fleetd#leadMailboxOpener()}. */
Fleetd.LeadMailboxOpener leadMailboxOpener();
/** A monotonic elapsed-time clock. Production: {@link System#nanoTime()}. */
LongSupplier nanoClock();
/**
* A wall-clock reading, in nanoseconds. Production: {@code System.currentTimeMillis()}
* converted to nanoseconds. Kept separate from {@link #nanoClock()} because {@link
* dev.ltms.fleet.health.FleetHealthMonitor} needs both — one monotonic clock for elapsed-time
* decisions, one wall clock to detect and correct for a macOS sleep freezing the monotonic one.
*/
LongSupplier wallClockNanos();
/** A dedicated single-thread scheduler; production names its (virtual) thread {@code purpose}. */
ScheduledExecutorService newScheduler(String purpose);
/** Register a JVM shutdown hook that runs {@code hook} on JVM exit. */
void addShutdownHook(Runnable hook);
/** Bind and start the HTTP server. */
void startHttp(Javalin app, String host, int port);
/** The real production ports: a live herdr socket, a live broker, real threads, a real bind. */
static ResourcePorts system() {
return new SystemResourcePorts();
}
}
@@ -0,0 +1,66 @@
package dev.ltms.fleet;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
import io.javalin.Javalin;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
/**
* fleetd #612 Unit A: the one production {@link ResourcePorts} — every method here is the exact
* call {@code Fleetd.main} used to make directly at each of these sites before this ticket.
* Package-private: obtained only through {@link ResourcePorts#system()}.
*/
final class SystemResourcePorts implements ResourcePorts {
@Override
public Map<String, String> environment() {
return System.getenv();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return UnixSocketHerdrClient.connect(socketPath, new ObjectMapper());
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return Fleetd.replyInboxOpener();
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return Fleetd.leadMailboxOpener();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis());
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor(r -> Thread.ofVirtual().name(purpose).unstarted(r));
}
@Override
public void addShutdownHook(Runnable hook) {
Runtime.getRuntime().addShutdownHook(new Thread(hook));
}
@Override
public void startHttp(Javalin app, String host, int port) {
app.start(host, port);
}
}
@@ -462,8 +462,8 @@ public record FleetConfig(
* profile that does not opt in. Read live off the current config, so it is
* HOT: a change takes effect on the next exhaustion classification / spawn,
* no restart needed.
* @param autoCompactWindow opt-in per-profile token window that forces a spawned member to
* auto-compact its context at (Claude Code) or within (opencode) a bound the
* @param autoCompactWindow opt-in per-profile token window that forces a launched Claude Code
* session to auto-compact its context at, or an opencode session within, a bound the
* operator chooses, instead of the backend's own default. {@code null} (the
* default) leaves today's behaviour exactly — opencode already forces
* {@code compaction.auto: true} unconditionally (CB-523) but has no absolute
@@ -1329,9 +1329,20 @@ public record FleetConfig(
* each such nudge costs the lead a turn just to read "nothing pending";
* three is enough to tell it it may stand down without nagging forever, and
* it is the bound that stops an idle fleet from being a subscription burner.
* @param contextHighNudge fleetd #609: when {@code true}, an idle lead whose own {@code
* LeadContextGauge} reading is {@code HIGH} gets a text notice telling it
* to consider a handover, appended to whatever heartbeat nudge the loop
* already sends. Default {@code false} ({@code null} also means off) — an
* upgraded daemon must not silently start telling leads to hand over.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap) {
public record LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap,
Boolean contextHighNudge) {
/** Convenience constructor for every call site that predates fleetd #609: no context notice. */
public LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap) {
this(idleAfterSeconds, backoffMs, quietNudgeCap, null);
}
public LeadHeartbeat {
idleAfterSeconds = (idleAfterSeconds == null || idleAfterSeconds <= 0) ? 300 : idleAfterSeconds;
backoffMs = (backoffMs == null || backoffMs <= 0) ? 60_000L : backoffMs;
@@ -1864,6 +1875,7 @@ public record FleetConfig(
rejectDuplicateMemberSlots(yaml);
rejectNegativeMaxLoad(yaml);
rejectAutoCompactWindowOutOfRange(yaml);
warnConflictingAutoCompactWindows(yaml);
rejectMalformedProfilePatterns(yaml);
rejectUnknownKind(yaml);
rejectUnknownAuthMode(yaml);
@@ -2214,6 +2226,73 @@ public record FleetConfig(
}
}
/**
* Warn (never refuse to start) about a Claude Code profile whose auto-compaction flag and
* environment setting disagree.
*
* <p>Renamed from {@code rejectConflictingAutoCompactWindows} (fleetd #601 review, measured
* 2026-09-22): that method threw {@link IllegalStateException}, so {@link #load(Path)} refused
* to start on a config carrying this conflict. On this host, four profiles trip it, including
* the lead's own profile and the one every worker spawns on — so the throw is not a rare edge
* case. Under launchd, a throw inside {@code load()} is a restart loop, not an error an operator
* reads once, and the config that would fix it ({@code fleetd/fleetd.yaml}) is gitignored, so
* the cause is invisible on the host where it bites. A WARN gives the operator the same
* information — which profiles, and now both values, so they can fix it without reading the
* source — without ever taking the fleet down.
*
* <p>fleetd #618 measured which of the two inputs Claude Code actually follows when they
* disagree: the environment variable wins, so {@code autoCompactWindow} is inert on a profile
* that also sets the env var. This method only detects and reports the disagreement — it does
* not correct it — see {@link dev.ltms.fleet.launch.ClaudeCodeArguments} for the full measured
* precedence.
*
* <p>Equal values never warn: either input then produces the same session window, so there is
* nothing to reconcile.
*/
static void warnConflictingAutoCompactWindows(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return;
}
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
return;
}
List<String> names = new ArrayList<>();
List<String> detail = new ArrayList<>();
for (Map.Entry<?, ?> entry : profiles.entrySet()) {
if (!(entry.getValue() instanceof Map<?, ?> profile)
|| !(profile.get("autoCompactWindow") instanceof Number window)
|| !(profile.get("env") instanceof Map<?, ?> env)
|| !env.containsKey("CLAUDE_CODE_AUTO_COMPACT_WINDOW")) {
continue;
}
Object kind = profile.get("kind");
boolean claudeCode = kind == null || String.valueOf(kind).isBlank()
|| Profile.KIND_CLAUDE_CODE.equalsIgnoreCase(String.valueOf(kind));
Object envValue = env.get("CLAUDE_CODE_AUTO_COMPACT_WINDOW");
if (claudeCode && !String.valueOf(window).equals(String.valueOf(envValue))) {
String name = String.valueOf(entry.getKey());
names.add(name);
detail.add(name + " (autoCompactWindow=" + window
+ ", env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=" + envValue + ")");
}
}
if (names.isEmpty()) {
return;
}
names.sort(String::compareTo);
detail.sort(String::compareTo);
log.warn("Claude Code profile(s) {} set disagreeing autoCompactWindow and env."
+ "CLAUDE_CODE_AUTO_COMPACT_WINDOW — the daemon starts anyway: {}. fleetd "
+ "#618 measured that CLAUDE_CODE_AUTO_COMPACT_WINDOW wins, so "
+ "autoCompactWindow is inert on these profiles. Set equal values on each "
+ "to resolve this — do not just delete the env var, since that LOWERS the "
+ "live window to autoCompactWindow's value rather than fixing anything.",
names, String.join(", ", detail));
}
/**
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
@@ -0,0 +1,37 @@
package dev.ltms.fleet.launch;
import dev.ltms.fleet.config.FleetConfig;
import java.util.ArrayList;
import java.util.List;
/** Arguments shared by every fleetd path that starts Claude Code. */
public final class ClaudeCodeArguments {
private ClaudeCodeArguments() {
}
/**
* Append the configured Claude Code auto-compaction window when the profile opts in.
*
* <p>This flag and the environment variable {@code CLAUDE_CODE_AUTO_COMPACT_WINDOW} can
* disagree, and fleetd #618 measured which one Claude Code actually follows: the environment
* variable wins, ahead of this {@code --autocompact} flag, ahead of the settings file, ahead of
* clientdata, the experiment, and the model default. So when a profile sets both, the flag this
* method appends has NO effect — Claude Code reads {@code CLAUDE_CODE_AUTO_COMPACT_WINDOW}
* first and never consults the flag. {@link FleetConfig#load(java.nio.file.Path)} only WARNS
* when a Claude Code profile sets both to different values (see {@code
* FleetConfig.warnConflictingAutoCompactWindows}) — it does not stop the daemon from starting,
* and the launched session honours the env var, not this flag. Measured against Claude Code
* 2.1.278 (fleetd #618) — a later version could reorder this precedence.
*/
public static List<String> withAutoCompactWindow(List<String> argv, FleetConfig.Profile profile) {
if (profile.autoCompactWindow() == null) {
return argv;
}
List<String> withAutoCompact = new ArrayList<>(argv);
withAutoCompact.add("--autocompact");
withAutoCompact.add(String.valueOf(profile.autoCompactWindow()));
return withAutoCompact;
}
}
@@ -0,0 +1,313 @@
package dev.ltms.fleet.lead;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
import java.io.RandomAccessFile;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.function.LongSupplier;
/**
* Reads how full a lead's own Claude Code context window is, from the transcript Claude Code
* itself writes — never from the lead's pane (fleetd has {@code AgentControl.read} for that, and
* this must not use it: a pane holds terminal text, not the structured usage numbers a transcript
* carries, and scraping it would also race the lead's own rendering).
*
* <p><strong>Why this exists.</strong> A lead auto-compacts when its context fills — on the host
* this was built for, that happened 30 times in one session, discarding roughly 250,000 tokens
* and costing 46 seconds to 3 minutes each time, and fleetd had no way to see it coming. This
* class is the first thing that looks.
*
* <p><strong>The route.</strong> Claude Code appends one JSON object per line to
* {@code <configDir>/projects/<slug>/<sessionId>.jsonl}. {@code <slug>} is an undocumented,
* internal encoding of the working directory — this class never derives it. Instead it lists the
* one-level-deep subdirectories of {@code <configDir>/projects/} and looks for
* {@code <sessionId>.jsonl} by name, so the slug rule can change without breaking this reader.
*
* <ul>
* <li><strong>Live context</strong> is read off the last record in the read window that carries
* a {@code message.usage} object: {@code input_tokens + cache_read_input_tokens +
* cache_creation_input_tokens}. This is what actually fills the window — a plain
* {@code input_tokens} count alone understates it once the conversation has any cached
* prefix, which on a long-lived lead is always.</li>
* <li><strong>Compaction history</strong> is a count of {@code subtype: "compact_boundary"}
* records seen in the same read window — see {@link Reading#compactions()}. It is a count
* within the window this reader actually looked at, not a lifetime total: a session with
* more compactions than fit in {@link #TAIL_BYTES} of transcript will undercount. That
* trade-off is deliberate — see {@link #TAIL_BYTES}.</li>
* </ul>
*
* <p><strong>Three states, not two (OK / HIGH / UNKNOWN).</strong> Every path that cannot
* positively establish the live token count — a missing file, an unreadable one, a peer that
* is not a Claude backend, or every line in the read window failing to parse as JSON — returns
* {@link State#UNKNOWN} with no token number, never a default "0" or "ok" that would read as
* "this lead is fine" when the honest answer is "I could not look".
*
* <p><strong>A torn final line does not mean UNKNOWN.</strong> {@code fleet_list} reads this
* transcript while Claude Code may be mid-write on it, so the last line in the window can be cut
* off mid-flush — that is an ordinary, expected race, not a sign the format has changed. Earlier
* this class treated ANY unparseable last line as UNKNOWN, on the theory that "if the format
* changes, we should see UNKNOWN". That reasoning does not hold: a real format change makes
* <em>every</em> line in the window unparseable, not only the last one written. So a single
* malformed line (most often the final, torn one, but the check is not position-specific) is
* simply skipped rather than treated as fatal, and the reading is built from whatever lines in the
* window did parse. Only when <em>none</em> of them parse — the real format-change signal — does
* this return {@link State#UNKNOWN}, still with no stale number standing in for "I could not
* tell".
*
* <p><strong>Bounded cost.</strong> {@code fleet_list} is polled constantly, so every read is
* capped two ways: {@link #TAIL_BYTES} bounds how much of the transcript is ever read from disk
* (never the whole 52 MB a long-lived transcript reaches on the host this was measured on), and
* {@link #DEFAULT_CACHE_TTL_MILLIS} bounds how often that bounded read actually happens — a burst
* of {@code fleet_list} calls inside one TTL window reads the file once. One instance's cache is
* keyed by {@code (configDir, sessionId)}, so it is safe to share across every lead a single
* {@code fleet_list} call reports on.
*/
public final class LeadContextGauge {
private static final String PROJECTS_DIR = "projects";
private static final ObjectMapper MAPPER = new ObjectMapper();
/**
* How many trailing bytes of a transcript a single read ever pulls off disk. Chosen so one
* read comfortably spans many recent turns — each usage or compact_boundary record is at most
* a few KB — while staying nowhere near the 52 MB a long session's real transcript reaches on
* the host this was built for; reading that whole file on every {@code fleet_list} call is
* exactly the cost this bound exists to avoid. 2 MiB holds on the order of hundreds of recent
* lines even when a turn's tool output is unusually large, which is far more than needed to
* find the most recent usage record and any recent compaction.
*/
static final int TAIL_BYTES = 2 * 1024 * 1024;
/**
* How long a {@link Reading} is served from cache before the file is read again.
* {@code fleet_list} is called constantly (by design — it is the fleet's own status probe), so
* without a TTL a burst of calls would re-read the transcript tail once per call. 5 seconds is
* short enough that a caller watching for a state change never waits long, and long enough that
* a poll loop calling every second or two only touches disk once per window.
*/
static final long DEFAULT_CACHE_TTL_MILLIS = 5_000;
/**
* Live tokens at or above this count report {@link State#HIGH}. On the host this was measured
* on, auto-compaction actually fires around 267,000–270,000 tokens, but the point of a HIGH
* state is to warn before that happens, not at it — 200,000 is the standard Claude context
* window size and a sensible built-in default: no config key is required to pick it, and a
* lead crossing it is already deep enough into its window that a compaction is foreseeable.
*/
static final long HIGH_THRESHOLD_TOKENS = 200_000;
/** The only peer kind this reader understands ({@code Agent.agentType()}'s wire value). */
private static final String CLAUDE_AGENT_TYPE = "claude";
public enum State { OK, HIGH, UNKNOWN }
/**
* @param state {@link State#UNKNOWN} whenever {@code tokens} could not be established
* @param tokens live context tokens, or {@code null} exactly when {@code state} is
* {@link State#UNKNOWN}
* @param compactions {@code compact_boundary} records seen in the read window (see class
* javadoc) — {@code 0} both for "genuinely none seen" and for "unknown",
* since a caller that already sees {@code state: UNKNOWN} has no reason to
* trust this number either way
*/
public record Reading(State state, Long tokens, int compactions) {
/**
* fleetd #609: widened from package-private to public so {@code
* dev.ltms.fleet.msg.LeadHeartbeatLoop.LeadContextSource.none()} (a different package) can
* return the same inert "I could not look" reading the gauge itself uses, without inventing
* a parallel unknown-reading constant. Behaviour of this class is otherwise unchanged.
*/
public static Reading unknown() {
return new Reading(State.UNKNOWN, null, 0);
}
}
private record CacheEntry(Reading reading, long readAtMillis) {
}
private final LongSupplier clock;
private final long ttlMillis;
private final Map<String, CacheEntry> cache = new ConcurrentHashMap<>();
/** Test seam only (package-private) — counts real disk reads, i.e. cache misses. */
private final AtomicInteger diskReads = new AtomicInteger();
public LeadContextGauge() {
this(System::currentTimeMillis, DEFAULT_CACHE_TTL_MILLIS);
}
/** Test seam: an injectable clock and TTL so cache expiry is provable without sleeping. */
LeadContextGauge(LongSupplier clock, long ttlMillis) {
this.clock = clock;
this.ttlMillis = ttlMillis;
}
/** How many times this instance has actually read a transcript off disk — test seam only. */
int diskReadCount() {
return diskReads.get();
}
/**
* @param configDir the lead's {@code CLAUDE_CONFIG_DIR}, or {@code null}/blank to use the
* default {@code <user.home>/.claude} — the right answer for the common case
* where the lead's profile sets no {@code configDir} override
* @param sessionId the lead's own Claude session id ({@code Agent.sessionId()}), or
* {@code null} when herdr has not resolved one yet
* @param agentType the detected peer kind ({@code Agent.agentType()}); anything other than
* {@code "claude"} (including {@code null}, meaning undetected) reports
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
* transcript format
*/
public Reading read(String configDir, String sessionId, String agentType) {
if (sessionId == null || sessionId.isBlank()) {
return Reading.unknown();
}
if (!CLAUDE_AGENT_TYPE.equalsIgnoreCase(agentType)) {
return Reading.unknown();
}
String base = (configDir == null || configDir.isBlank())
? System.getProperty("user.home") + "/.claude"
: configDir;
String cacheKey = base + '\u0000' + sessionId;
long now = clock.getAsLong();
CacheEntry cached = cache.get(cacheKey);
if (cached != null && now - cached.readAtMillis() < ttlMillis) {
return cached.reading();
}
Reading fresh = readUncached(base, sessionId);
cache.put(cacheKey, new CacheEntry(fresh, now));
return fresh;
}
private Reading readUncached(String base, String sessionId) {
diskReads.incrementAndGet();
Path file = findTranscript(base, sessionId);
if (file == null) {
return Reading.unknown();
}
TailRead tail;
try {
tail = tailBytes(file, TAIL_BYTES);
} catch (IOException e) {
return Reading.unknown();
}
return parse(tail);
}
/**
* Finds {@code <sessionId>.jsonl} under {@code <base>/projects/}, one level deep — never by
* deriving the slug directory from a working directory (see class javadoc). Bounded to a
* single {@code list()} of {@code projects/} itself: it never recurses further, so the cost is
* the number of project directories, not the size of any transcript inside them.
*/
private Path findTranscript(String base, String sessionId) {
Path projectsDir = Path.of(base, PROJECTS_DIR);
if (!Files.isDirectory(projectsDir)) {
return null;
}
String filename = sessionId + ".jsonl";
Path direct = projectsDir.resolve(filename);
if (Files.isRegularFile(direct)) {
return direct;
}
try (var children = Files.list(projectsDir)) {
return children.filter(Files::isDirectory)
.map(dir -> dir.resolve(filename))
.filter(Files::isRegularFile)
.findFirst()
.orElse(null);
} catch (IOException e) {
return null;
}
}
/** Package-private (not {@code private}): {@link #tailBytes} is a test seam, see its javadoc. */
record TailRead(byte[] bytes, boolean fromStart) {
}
/**
* Reads at most {@code maxBytes} trailing bytes of {@code file}. Package-private (not
* {@code private}) so a test can assert directly on the returned array's length — "bytes
* actually read", not on any parsed answer — without needing a file anywhere near
* {@link #TAIL_BYTES} in size to prove the cap holds.
*/
static TailRead tailBytes(Path file, int maxBytes) throws IOException {
try (RandomAccessFile raf = new RandomAccessFile(file.toFile(), "r")) {
long length = raf.length();
long start = Math.max(0, length - maxBytes);
raf.seek(start);
byte[] buf = new byte[(int) (length - start)];
raf.readFully(buf);
return new TailRead(buf, start == 0);
}
}
/**
* Parses the tail into a {@link Reading}. The first line is dropped unconditionally whenever
* the tail is not the whole file (it starts mid-line, cut by {@link #TAIL_BYTES} — an expected
* artefact of the bound, not a data problem). Every remaining line is then parsed on a
* best-effort basis: a line that fails to parse (most often the last one, torn by a write this
* read raced — see the "torn final line" section of the class javadoc) is skipped, not fatal.
* Only when none of the remaining lines parse does this report {@link State#UNKNOWN}.
*/
private Reading parse(TailRead tail) {
String text = new String(tail.bytes(), StandardCharsets.UTF_8);
List<String> lines = new ArrayList<>(List.of(text.split("\n", -1)));
if (!lines.isEmpty() && lines.get(lines.size() - 1).isEmpty()) {
lines.remove(lines.size() - 1); // trailing newline leaves a phantom empty last element
}
if (!tail.fromStart() && !lines.isEmpty()) {
lines.remove(0); // first line is a fragment cut by our own tail bound, not real data
}
if (lines.isEmpty()) {
return Reading.unknown();
}
Long tokens = null;
int compactions = 0;
boolean anyLineParsed = false;
for (String line : lines) {
JsonNode node = tryParse(line);
if (node == null) {
// A malformed line — typically the last one, cut mid-flush by a write this read
// raced — is skipped rather than treated as fatal. See the class javadoc's "torn
// final line" section for why: a real format change makes EVERY line unparseable,
// not only this one, and that case is still caught below by anyLineParsed.
continue;
}
anyLineParsed = true;
JsonNode usage = node.path("message").path("usage");
if (usage.isObject()) {
tokens = usage.path("input_tokens").asLong(0)
+ usage.path("cache_read_input_tokens").asLong(0)
+ usage.path("cache_creation_input_tokens").asLong(0);
}
if ("compact_boundary".equals(node.path("subtype").asText(null))) {
compactions++;
}
}
if (!anyLineParsed) {
return Reading.unknown();
}
if (tokens == null) {
return new Reading(State.UNKNOWN, null, compactions);
}
State state = tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK;
return new Reading(state, tokens, compactions);
}
private JsonNode tryParse(String line) {
try {
return MAPPER.readTree(line);
} catch (IOException e) {
return null;
}
}
}
@@ -8,6 +8,7 @@ import dev.ltms.fleet.herdr.PendingCloseMarker;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.launch.ClaudeCodeArguments;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -358,7 +359,8 @@ public final class LeadLauncher {
}
/**
* The lead's argv: the profile's own command, the model pin, and the bridge MCP mount.
* The lead's argv: the profile's own command, the model and auto-compaction pins, and the bridge
* MCP mount.
*
* <p>No {@code --append-system-prompt}. That flag carries the worker reply charter, and a lead
* is not a worker — it reads its orchestration rules from the project's {@code CLAUDE.md} like
@@ -380,7 +382,7 @@ public final class LeadLauncher {
argv.add("--model");
argv.add(profile.model());
}
return argv;
return profile.isOpenCode() ? argv : ClaudeCodeArguments.withAutoCompactWindow(argv, profile);
}
/**
@@ -231,7 +231,9 @@ public final class LeadRollover {
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
* that is, in fact, actively running. This is not sticky: the deferred continuation
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
* #TURN_NEVER_SETTLED}, or {@link #CLEAR_NEVER_SETTLED}) once it finishes.
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into
* {@link #FAILED} instead of leaving this entry stuck forever.
*/
IN_PROGRESS,
/**
@@ -253,6 +255,19 @@ public final class LeadRollover {
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent.
*/
CLEAR_NEVER_SETTLED,
/**
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it.
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS}
* forever, because the production {@code continuationRunner} is a bare virtual thread with
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a
* terminal outcome. {@code detail} names the exception, so a reader has something to act on
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck
* lead must {@link #open} a fresh request.
*/
FAILED,
/**
* {@code token} names nothing this instance currently knows about: never issued by {@link
* #open}, dropped by {@link #cancel}, or aged out of {@link #outcomes}'s bounded history.
@@ -493,8 +508,42 @@ public final class LeadRollover {
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
* four-step order. There is no result to return to by this point, so every outcome is logged
* only.
*
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} →
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual
* thread with no uncaught-exception handler (see this class's public constructor). Before this
* fix, either throw killed the continuation thread silently, leaving the {@link
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch}
* below is scoped to the method body rather than to each {@code send} call individually, so it
* also covers anything else added to this continuation later, not just today's two call sites —
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
* branches below) rather than inside the helpers that detect them.</p>
*
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
* broader {@link Exception} or {@link Throwable}, which would also swallow something like an
* {@link OutOfMemoryError} this continuation has no business handling.</p>
*/
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
try {
runRolloverUnguarded(p, cfg);
} catch (RuntimeException e) {
log.warn("lead-rollover: continuation for token={} lead={} threw {} — the roll is dead; "
+ "no further step in this continuation will run",
p.token(), p.leadTerminal(), e.toString(), e);
outcomes.put(p.token(), new RollStatus(RollState.FAILED,
"the roll's continuation threw " + e.toString() + " — the roll is dead and will "
+ "not retry itself; check the daemon log for the stack trace, then open() "
+ "a fresh rollover request"));
}
}
/** The actual body of {@link #runRollover}, unwrapped — see that method's javadoc for the catch. */
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
String lead = p.leadTerminal();
long rollStartMillis = nowMillis.getAsLong();
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
@@ -31,6 +31,20 @@ public final class ConnectionIdentity {
this.cwds = cwds;
}
/**
* The {@link PaneLocator} this identity resolves callers against — fleetd #612 CB-185: lets a
* test drive the exact {@link PaneLocator} a real assembly wired up (e.g. {@code
* FleetdAssembly}'s {@code new ConnectionIdentity(new PaneLocator(herdr, memberHerdr), ...)})
* directly with a chosen pid, bypassing the OS-dependent {@link PeerPidLookup} that {@link
* #resolve} otherwise goes through. A full HTTP round trip cannot exercise this: {@code
* LsofPeerPidLookup} excludes its own pid, and an in-process test client and server share one
* JVM pid, so {@code pidForLocalPort} always returns {@code -1} and {@link PaneLocator} never
* gets called at all.
*/
public PaneLocator panes() {
return panes;
}
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
@@ -12,6 +12,7 @@ import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.msg.LeadChannel;
import dev.ltms.fleet.msg.LeadMessage;
@@ -106,6 +107,13 @@ public final class FleetMcp {
* untested identity heuristic.
*/
private final boolean authorizationEnforced;
/**
* fleetd #612 CB-185: kept as a field (rather than only captured by the {@code
* contextExtractor} closure built in the constructor) so a test can reach the exact {@link
* ConnectionIdentity} — and, through {@link ConnectionIdentity#panes()}, the exact {@link
* dev.ltms.fleet.herdr.PaneLocator} — that a real assembly wired up. See {@link #identity()}.
*/
private final ConnectionIdentity identity;
private final Metrics metrics; // CB-502: null → auth failures not counted
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
@@ -115,6 +123,8 @@ public final class FleetMcp {
private final OutageSource outage;
/** fleetd #176: SEPARATE from both of the above — see {@link LeadSeatSource}'s doc. */
private final LeadSeatSource leadSeats;
/** fleetd #602 gauge-wiring: see {@link LeadConfigDirSource}. */
private final LeadConfigDirSource leadConfigDirs;
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
private final LeadChannel leadChannel;
/** fleetd #361: {@code coordinator.peers} — see {@link CoordinationSource}. Empty when unset. */
@@ -126,6 +136,17 @@ public final class FleetMcp {
* clean {@code NOT_CONFIGURED} refusal rather than throwing. See {@link #handover}.
*/
private final LeadRollover leadRollover;
/**
* "Lead context gauge": how full each lead's own Claude Code context window is, reported on
* {@code fleet_list}'s {@code leads} rows (see {@link #contextView}). Built unconditionally, in
* the field initializer rather than a constructor parameter — this is not a togglable feature
* with an on/off config knob the way {@link OutageSource}/{@link LeadSeatSource} are: it needs
* no config at all (see {@link LeadContextGauge}'s own javadoc for the built-in defaults), so
* there is no "off" value to thread through every existing constructor call site. One instance
* per daemon so its read cache (keyed by session, TTL'd) is actually shared across
* {@code fleet_list} calls rather than rebuilt — and therefore useless — on every call.
*/
private final LeadContextGauge leadContextGauge = new LeadContextGauge();
/** Capacity facts used by {@code fleet_list}; production must supply the placement live count. */
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
@@ -246,6 +267,29 @@ public final class FleetMcp {
public static LeadSeatSource none() { return new LeadSeatSource(_ -> 0); }
}
/**
* fleetd #602 gauge-wiring: a lead name's configured {@code CLAUDE_CONFIG_DIR} override, fed to
* {@link LeadContextGauge#read} so {@code fleet_list}'s {@code context} row reads the transcript
* directory the lead's own profile actually writes to, not always the built-in
* {@code <user.home>/.claude} default.
*
* <p>The same idiom as {@link LeadSeatSource} — a {@code FleetMcp} constructor field, not a
* lookup {@code contextView} performs itself, because {@code FleetMcp} holds no
* {@link dev.ltms.fleet.config.FleetConfig} and {@code contextView} is {@code static}. See
* {@code Fleetd.leadConfigDirLookup} for the derivation: the same
* {@code fleet.leaders.<name>.profile} link {@link LeadSeatSource} already follows, resolved to
* that profile's own {@code configDir:}.
*
* @param configDirFor lead name → {@code configDir}, or {@code null} when the lead's entry names
* no profile, or that profile sets no {@code configDir} override — either
* way {@link LeadContextGauge} then falls back to its own built-in default,
* exactly as before this ticket
*/
public record LeadConfigDirSource(Function<String, String> configDirFor) {
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in default {@code configDir}. */
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null); }
}
/**
* fleetd #361: peer-visibility facts for {@code fleet_list}'s {@code coordinator} row — this
* daemon's own {@link LeadChannel} (for its self mailbox state and held messages) plus the
@@ -346,7 +390,7 @@ public final class FleetMcp {
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, authorizationMode, metrics,
capacity, healthCoverage, LoopHealthSource.none(), quarantine, leadChannel, outage, leadSeats,
peers, leadRollover);
LeadConfigDirSource.none(), peers, leadRollover);
}
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
@@ -354,16 +398,19 @@ public final class FleetMcp {
CallerResolver callers, AuthorizationMode authorizationMode, Metrics metrics,
CapacitySource capacity, HealthCoverageSource healthCoverage, LoopHealthSource loopHealth,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
LeadSeatSource leadSeats, LeadConfigDirSource leadConfigDirs, List<String> peers,
LeadRollover leadRollover) {
Objects.requireNonNull(callers, "callers");
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
== AuthorizationMode.ENFORCED;
this.identity = identity;
this.leadChannel = leadChannel;
this.peers = peers == null ? List.of() : List.copyOf(peers);
this.capacity = capacity;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.outage = Objects.requireNonNull(outage, "outage");
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
this.leadConfigDirs = Objects.requireNonNull(leadConfigDirs, "leadConfigDirs");
this.healthCoverage = healthCoverage;
this.loopHealth = Objects.requireNonNull(loopHealth, "loopHealth");
this.leadRollover = leadRollover;
@@ -498,7 +545,7 @@ public final class FleetMcp {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_list", Map.of()), null);
if (denied != null) return denied;
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
leadSeats, callers.leads(),
leadSeats, leadContextGauge, leadConfigDirs, callers.leads(),
callerTerminal(exchange),
new CoordinationSource(leadChannel, peers),
coordinatorVisibleTo(principal(exchange)));
@@ -716,6 +763,15 @@ public final class FleetMcp {
return transport;
}
/**
* The {@link ConnectionIdentity} this server resolves every caller against — fleetd #612
* CB-185: lets a test reach the exact {@link dev.ltms.fleet.herdr.PaneLocator} a real assembly
* wired up (via {@link ConnectionIdentity#panes()}), rather than a copy built for the test.
*/
public ConnectionIdentity identity() {
return identity;
}
/** Mark a connected spawned member available for the injector readiness gate. */
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember()) {
@@ -741,6 +797,46 @@ public final class FleetMcp {
return server.listTools();
}
/**
* fleetd #612 B3 — same reason as {@link #registeredTools()}: a test that must drive the REAL
* {@link QuarantineSource} (and the real {@link BackendQuarantine} it wraps) this daemon was
* assembled with, rather than scraping {@code FleetdAssembly.java}'s source text for the
* constructor call that built it. Unlike {@link #registeredTools()}'s callers, that test cannot
* live in this package: it also builds the {@code ResourcePorts} that drives
* {@code FleetdAssembly.assembleAndStart}, and {@code ResourcePorts}' methods return
* {@code Fleetd}-nested types that are only visible from package {@code dev.ltms.fleet} — so
* this accessor is {@code public}, not package-private, to stay reachable from there. {@code
* FleetdBackendQuarantineAssemblyTest} quarantines a credential twice through this exact
* instance and checks the second cooldown is longer than the first — the one behavioural
* difference {@link BackendQuarantine#withEscalation} and the flat two-argument constructor
* actually produce.
*/
public QuarantineSource quarantineSource() {
return quarantine;
}
/**
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
* reason, for the real {@link LeadSeatSource} this daemon was assembled with. {@code
* FleetdLeadSeatAssemblyTest} calls {@code seatsFor} on this exact instance and checks it
* reports a live lead's seat, which {@link LeadSeatSource#none()} can never do (it is a
* constant-zero function regardless of input).
*/
public LeadSeatSource leadSeatSource() {
return leadSeats;
}
/**
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText}
* through the real {@code router.leadAgents()}.
*/
public LeadRollover leadRollover() {
return leadRollover;
}
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/**
@@ -1626,7 +1722,7 @@ public final class FleetMcp {
Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
OutageSource.none(), LeadSeatSource.none(), leads, selfTerm, coordination, false);
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1635,7 +1731,7 @@ public final class FleetMcp {
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), leads, selfTerm, CoordinationSource.none(), false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, CoordinationSource.none(), false);
}
/**
@@ -1652,7 +1748,7 @@ public final class FleetMcp {
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
LeadSeatSource.none(), leads, selfTerm, coordination, false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1661,7 +1757,7 @@ public final class FleetMcp {
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), leads, selfTerm, coordination, false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/**
@@ -1683,7 +1779,7 @@ public final class FleetMcp {
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
leadSeats, leads, selfTerm, coordination, false);
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/**
@@ -1707,14 +1803,24 @@ public final class FleetMcp {
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
outage, leadSeats, leads, selfTerm, coordination, callerIsPrimary);
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, callerIsPrimary);
}
/**
* The canonical implementation. {@code contextGauge} is the "lead context gauge" (see
* {@link LeadContextGauge}) — every wrapper overload above passes a freshly constructed one,
* which is correct for them (none of them exercise repeated calls where a shared cache would
* matter); the one caller that matters for caching, {@code fleet_list}'s MCP handler, passes
* its own single long-lived instance instead (see {@code FleetMcp}'s {@code leadContextGauge}
* field).
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
LoopHealthSource loopHealth,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
LeadSeatSource leadSeats, LeadContextGauge contextGauge,
LeadConfigDirSource leadConfigDirs,
Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary) {
try {
Map<String, Agent> live = workers.list().stream()
@@ -1723,7 +1829,8 @@ public final class FleetMcp {
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
leadConfigDirs))
.toList();
// fleetd #209: this is the caller-driven fleet_list read that actually reports
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
@@ -2018,15 +2125,25 @@ public final class FleetMcp {
}
/**
* One lead's row: its address, its name, and whether it can be reached right now.
* One lead's row: its address, its name, whether it can be reached right now, and how full its
* own Claude Code context window is.
*
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
* see is a lead a {@code fleet_send} cannot be typed into. It is reported rather than hidden,
* because a peer that has gone unreachable is exactly what the sender needs to know.
*
* <p>{@code context} is the "lead context gauge" (fleetd's context-usage visibility ticket):
* {@code {state: "ok"|"high"|"unknown", tokens?: number, compactions: number}}, read from the
* lead's transcript — see {@link LeadContextGauge}. Kept small on purpose (a token count, a
* state, and a compaction count) rather than echoing the whole reading history: this is a
* roster row a caller glances at, not a diagnostics dump. {@code tokens} is present only when
* {@code state} is not {@code "unknown"} — never a stale or default number standing in for "I
* could not tell".
*/
private static Map<String, Object> leadView(String terminal, String name, Agent live,
String selfTerm) {
String selfTerm, LeadContextGauge contextGauge,
LeadConfigDirSource leadConfigDirs) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", terminal);
m.put("name", name);
@@ -2035,9 +2152,32 @@ public final class FleetMcp {
if (terminal.equals(selfTerm)) {
m.put("self", true);
}
String configDir = leadConfigDirs.configDirFor().apply(name);
m.put("context", contextView(contextGauge, live, configDir));
return m;
}
/**
* Reads the lead context gauge for one lead. {@code configDir} is this lead's configured
* {@code CLAUDE_CONFIG_DIR} override (see {@link LeadConfigDirSource}), derived from
* {@code fleet.leaders.<name>.profile} → that profile's own {@code configDir:} — or {@code null}
* when the lead's entry names no profile, or that profile sets no override, in which case
* {@link LeadContextGauge#read} falls back to its own built-in default
* ({@code <user.home>/.claude}).
*/
private static Map<String, Object> contextView(LeadContextGauge contextGauge, Agent live, String configDir) {
String sessionId = live == null ? null : live.sessionId();
String agentType = live == null ? null : live.agentType();
LeadContextGauge.Reading reading = contextGauge.read(configDir, sessionId, agentType);
Map<String, Object> c = new LinkedHashMap<>();
c.put("state", reading.state().name().toLowerCase());
if (reading.tokens() != null) {
c.put("tokens", reading.tokens());
}
c.put("compactions", reading.compactions());
return c;
}
/** {@code fleet_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
@@ -8,6 +8,7 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.launch.ClaudeCodeArguments;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.PeerLauncher;
import org.slf4j.Logger;
@@ -298,7 +299,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
// has neither MCP nor a charter — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithFleet(cfg, spec));
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
return new Launch(workerEnv, argvWithAutoCompact(argvWithModel(argv, cfg), cfg), agentSessionId);
return new Launch(workerEnv, ClaudeCodeArguments.withAutoCompactWindow(argvWithModel(argv, cfg), cfg), agentSessionId);
}
/**
@@ -908,30 +909,6 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
return withModel;
}
/**
* Pin a bounded auto-compaction window on the command line via {@code --autocompact <tokens>},
* opt-in per profile (CB-634's sibling ticket: a member that runs out of context dies mid-turn
* and its {@code fleet_reply} — the whole point of the turn — is lost with it; opencode already
* forces {@code compaction.auto: true} unconditionally, CB-523, but Claude Code has no equivalent
* and runs at the backend's own default window).
*
* <p>Mirrors {@link #argvWithModel}: appended after it, so it survives the {@code ccs <profile>}
* wrapper the same way {@code --model} does, and outranks env/settings and the operator's own
* {@code argv}. Verified: {@code claude 2.1.241 --help} lists {@code --autocompact <auto|tokens>}
* (either the literal {@code auto}, or an integer 100k–1M) — {@link FleetConfig#load} rejects a
* configured value outside that band before this ever runs, so the flag Claude Code receives here
* is always in range.
*/
private static List<String> argvWithAutoCompact(List<String> argv, FleetConfig.Profile cfg) {
if (cfg.autoCompactWindow() == null) {
return argv;
}
List<String> withAutoCompact = mutableArgv(argv);
withAutoCompact.add("--autocompact");
withAutoCompact.add(String.valueOf(cfg.autoCompactWindow()));
return withAutoCompact;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
@@ -0,0 +1,24 @@
package dev.ltms.fleet.msg;
/**
* fleetd #612 A-gaps (gap 1): a {@link LeadChannel} that its owner can also close.
*
* <p>{@link LeadChannel}'s own javadoc says plainly that {@code close()} is deliberately left out
* of that interface — draining is a caller convenience nobody uses, and closing is the
* <em>owner's</em> job. This interface is that owner's own, wider view: whoever opens the
* coordination mailbox (the assembly that builds the daemon) also needs to close it from the
* shutdown path, and a test standing in for a real broker connection needs a fake it can mark
* closed, without ever holding a live connection. Every ordinary consumer ({@code FleetMcp},
* {@link LeadCoordLoop}) keeps taking the narrower {@link LeadChannel} exactly as before — only
* the owner speaks this wider one.
*
* <p>{@link LeadMailbox} is still the only production implementation. This only generalises the
* TYPE its owner holds it as (previously the concrete class), so a test can substitute a fake
* closeable channel instead of a real AMQP connection.
*/
public interface LeadChannelHandle extends LeadChannel, AutoCloseable {
/** Release the underlying connection. Declared with no checked exception, unlike the plain {@link AutoCloseable#close()}. */
@Override
void close();
}
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -13,6 +14,7 @@ import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
@@ -44,6 +46,15 @@ import java.util.function.Supplier;
* loop stands down, so two competing injections never start two turns in the same pane
* (constraint 6).</li>
* </ol>
*
* <p><b>fleetd #609 — context-high notice.</b> Optionally ({@code contextHighNudge}, opt-in like the
* loop itself), a tick that finds the lead's own {@link LeadContextGauge} reading at {@link
* LeadContextGauge.State#HIGH} appends a text notice to whatever nudge it sends, telling the lead to
* consider {@code fleet_handover}. This is text only — it never rolls a pane itself. It fires once per
* HIGH stretch (a latch, cleared only by a later {@code OK} reading — {@code UNKNOWN} neither sets nor
* clears it, since "I could not look" must not be read as "it got better"), and it never spends the
* quiet-nudge budget: an idle, quiet, HIGH-context lead is exactly the case {@link Action#QUIET_DONE}
* would otherwise swallow, and it is the one case most worth interrupting the quiet cap for.
*/
public final class LeadHeartbeatLoop {
@@ -63,11 +74,16 @@ public final class LeadHeartbeatLoop {
private final long backoffMs;
private final int quietNudgeCap;
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
private final LeadContextSource contextSource; // fleetd #609
private final boolean contextHighNudge; // fleetd #609: opt-in, like the loop itself
private final boolean requireOperatorConfirm; // fleetd #621: mirrors leadRollover.requireOperatorConfirm
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
private long idleSinceNanos = NOT_IDLE;
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
private int quietCount = 0;
/** fleetd #609: latched "already told this lead about this HIGH stretch". Single scheduler thread only. */
private boolean contextNotified = false;
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
@@ -75,7 +91,7 @@ public final class LeadHeartbeatLoop {
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, null);
idleAfterNanos, backoffMs, quietNudgeCap, null, LeadContextSource.none(), false);
}
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
@@ -83,6 +99,40 @@ public final class LeadHeartbeatLoop {
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, metrics, LeadContextSource.none(), false);
}
/**
* fleetd #609: as above, plus the lead's own context source and whether a HIGH reading should
* append a hand-over notice to the loop's nudge. Pass {@link LeadContextSource#none()} and
* {@code false} to keep the pre-#609 behaviour exactly (both existing public constructors do).
*
* <p>fleetd #621: delegates to the full constructor with {@code requireOperatorConfirm=true} —
* the pre-#621 wording ("ask the operator ... only the operator can approve the roll") assumed
* the config default, so every caller of this overload keeps that text byte-identical.
*/
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
LeadContextSource contextSource, boolean contextHighNudge) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, metrics, contextSource, contextHighNudge, true);
}
/**
* fleetd #621: as above, plus the daemon's effective {@code leadRollover.requireOperatorConfirm}
* value — threaded into {@link #contextNotice(boolean, LeadContextGauge.Reading, boolean, boolean)}
* so the notice's wording tracks the config the daemon actually enforces (see {@code
* LeadRollover.confirm}) instead of always asserting the operator gate is on.
*/
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
LeadContextSource contextSource, boolean contextHighNudge,
boolean requireOperatorConfirm) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
@@ -94,6 +144,20 @@ public final class LeadHeartbeatLoop {
this.backoffMs = backoffMs;
this.quietNudgeCap = quietNudgeCap;
this.metrics = metrics;
this.contextSource = contextSource;
this.contextHighNudge = contextHighNudge;
this.requireOperatorConfirm = requireOperatorConfirm;
}
/**
* fleetd #609: one lead's own context reading, keyed by its terminal id — the same injected-source
* idiom {@code FleetMcp.LeadSeatSource}/{@code FleetMcp.LeadConfigDirSource} already use.
*/
public record LeadContextSource(Function<String, LeadContextGauge.Reading> readingFor) {
/** Inert source — every lead reads UNKNOWN, so the context notice can never fire. */
public static LeadContextSource none() {
return new LeadContextSource(_ -> LeadContextGauge.Reading.unknown());
}
}
/**
@@ -126,8 +190,11 @@ public final class LeadHeartbeatLoop {
STAND_DOWN
}
/** The outcome of one decision: the action plus the state to persist for the next tick. */
record Decision(Action action, Long idleSinceNanos, int quietCount) {}
/**
* The outcome of one decision: the action, the state to persist for the next tick, and (fleetd
* #609) whether the lead has now been told about the current HIGH context stretch.
*/
record Decision(Action action, Long idleSinceNanos, int quietCount, boolean contextNotified) {}
/**
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
@@ -142,34 +209,47 @@ public final class LeadHeartbeatLoop {
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
* @param leadKnown whether a lead terminal is known to nudge at all
* @param fleet a snapshot of the pending fleet state (constraint 5)
* @param context fleetd #609: the lead's own {@link LeadContextGauge} reading for this tick
* @param contextNotified fleetd #609: whether the lead has already been told about the current HIGH
* stretch — a latch, carried forward by {@link #applyDecision}
* @return the action to take and the state to persist
*/
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
boolean pushLoopActive, boolean leadKnown, FleetState fleet) {
boolean pushLoopActive, boolean leadKnown, FleetState fleet,
LeadContextGauge.State context, boolean contextNotified) {
// fleetd #609: re-arm the latch only on a positive OK reading. UNKNOWN means "I could not
// look", not "it got better" — re-arming on UNKNOWN would let a flapping gauge (a transcript
// read that misses one tick) nudge a full lead again on every recovery, defeating the "once
// per HIGH stretch" promise. Computed once, up front, so every gate below carries it forward
// unchanged unless it is the gate that actually discharges it.
boolean latch = context == LeadContextGauge.State.OK ? false : contextNotified;
boolean contextHigh = contextHighNudge && context == LeadContextGauge.State.HIGH;
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
// competing prompt into the same pane would start a second turn — racing loops multiply
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
// quiet counter), because the reply that drove it is exactly the kind of new state that
// should reset the cap.
// should reset the cap. The context latch is untouched: standing down must not spend the
// one notice this stretch gets.
if (pushLoopActive) {
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0);
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0, latch);
}
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
// status (read failure, or the agent is gone) is safest treated the same way — never inject
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
if (status == null || !status.injectable()) {
return new Decision(Action.LEAD_BUSY, null, 0);
return new Decision(Action.LEAD_BUSY, null, 0, latch);
}
if (idleSinceNanos == null) {
// The lead just became injectable — record the start of an idle stretch and wait out the
// debounce quiet period before ever nudging (constraint 3).
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount);
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount, latch);
}
if (nowNanos - idleSinceNanos < idleAfterNanos) {
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
// and must not be re-prompted into every natural pause.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
}
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
// gating facts decide whether and how:
@@ -177,23 +257,33 @@ public final class LeadHeartbeatLoop {
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
// quiet period.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
}
if (fleet.hasPending()) {
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it.
return new Decision(Action.INJECT, idleSinceNanos, 0);
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it. The
// nudge text carries the context notice too when contextHigh — see injectNudge/contextNotice
// — so this route discharges the same duty and must set the latch.
return new Decision(Action.INJECT, idleSinceNanos, 0, latch || contextHigh);
}
if (contextHigh && !latch) {
// fleetd #609: the lead is idle, its context is full, and nothing is pending. This is the
// one case the quiet cap would otherwise swallow, and it is exactly when the lead most
// needs to hear it. Fire once per HIGH stretch, and do NOT spend the quiet budget on it:
// this is an event notice, not a "are you still there" nudge.
return new Decision(Action.INJECT, idleSinceNanos, quietCount, true);
}
if (quietCount < quietNudgeCap) {
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
// exactly that nothing is waiting so it can choose to stand down rather than hunt
// (constraint 5). Count it toward the consecutive-quiet cap.
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1);
// (constraint 5). Count it toward the consecutive-quiet cap. This nudge also carries the
// context notice when contextHigh (already latched above, or being latched now).
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1, latch || contextHigh);
}
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
// re-arms it — QUIET_DONE stops injection, not observation.
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount);
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount, latch);
}
// --- loop ----------------------------------------------------------------------------------
@@ -203,62 +293,195 @@ public final class LeadHeartbeatLoop {
* does not evaluate the lead's idle state before the fleet has settled.
*/
public void start() {
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {})",
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap);
// fleetd #613: contextHighNudge added alongside the three settings already here — an
// operator otherwise cannot tell from the boot log whether the #609 handover notice is
// armed, and had to load the deployed jar's config to confirm it.
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {}, "
+ "context-high nudge {})",
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap,
contextHighNudge);
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
}
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. */
private void tick() {
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. Package-private
* (mirroring {@link ReplyPushLoop#tick(String)}) so tests can drive it directly with a fake clock and a
* fake {@link AgentControl} instead of racing the scheduler thread. */
void tick() {
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
FleetState fleet = snapshot(inbox, roster);
AgentStatus status = AgentStatus.UNKNOWN;
LeadContextGauge.Reading reading = LeadContextGauge.Reading.unknown();
if (leadKnown) {
String leadTerminal = primaryRegistry.primaryTerminal().orElseThrow();
try {
status = agents.status(primaryRegistry.primaryTerminal().orElseThrow());
status = agents.status(leadTerminal);
} catch (RuntimeException e) {
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
// and never injects into a state it cannot read. Retry on the next backoff.
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
}
// fleetd #609: read the lead's own context regardless of status — decide() still gates on
// status first (constraint 2), so this is harmless work on a WORKING lead and lets the
// latch state stay accurate for whenever the lead does go idle.
reading = contextSource.readingFor().apply(leadTerminal);
}
Decision d = decide(clock.getAsLong(),
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
quietCount, status, pushLoop.isActive(), leadKnown, fleet);
quietCount, status, pushLoop.isActive(), leadKnown, fleet,
reading.state(), contextNotified);
applyDecision(d);
switch (d.action()) {
case INJECT -> injectNudge(fleet);
case QUIET_DONE -> countNudge("exhausted");
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> { /* nothing to inject, nothing to count */ }
case INJECT -> injectNudge(d, fleet, reading);
case QUIET_DONE -> {
countNudge("exhausted");
contextNotified = d.contextNotified();
}
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> contextNotified = d.contextNotified();
}
scheduleNext();
}
/** Persist the state a decision returned, so the next tick starts from it. */
/**
* Persist the idle/quiet state a decision returned, so the next tick starts from it.
*
* <p>fleetd #609 review: the context latch ({@link #contextNotified}) is deliberately <em>not</em>
* set here any more. Setting it from the decision unconditionally — before {@link #injectNudge} even
* tries to send — is exactly the review's blocker: a decision to notify is not the same fact as "the
* notice reached the pane". Every branch of {@link #tick} now assigns {@link #contextNotified} itself,
* once it knows whether a send happened and whether it carried the notice (see {@link #injectNudge}).
*/
private void applyDecision(Decision d) {
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
quietCount = d.quietCount();
}
/** Send the nudge to the known lead. */
private void injectNudge(FleetState fleet) {
/**
* Send the nudge to the known lead, with the fleetd #609 context notice appended when it applies, and
* persist the context latch based on what actually happened this tick — not merely what {@code d}
* chose to attempt.
*/
private void injectNudge(Decision d, FleetState fleet, LeadContextGauge.Reading reading) {
// fleetd #609 review: build the notice from the latch as it stood BEFORE this tick's decision —
// d.contextNotified() is the value to persist once delivery is confirmed, not the value the text
// itself should be built from. Otherwise a HIGH stretch that is still latched would never see the
// notice at all, defeating the very check this fixes.
String notice = contextNotice(contextHighNudge, reading, contextNotified, requireOperatorConfirm);
var lead = primaryRegistry.primaryTerminal();
if (lead.isEmpty()) {
return; // the lead disappeared between the decision and the injection
}
String leadTerminal = lead.get();
boolean sent = lead.isPresent() && trySend(lead.get(), fleet.nudgeText() + notice, notice);
// The latch becomes true only when all three hold: decide() chose to notify, a notice was
// actually included in the text, and the send reached the pane without throwing. Whenever no
// notice was attempted (disabled, not HIGH, or already latched), nothing was promised to the lead
// this tick, so apply the decision's own carried-forward value unconditionally — that is how the
// OK-only re-arm rule and STAND_DOWN's "don't burn the notice" rule keep working through this path
// too. A lead that disappeared between the decision and the send (lead.isEmpty()) is treated the
// same as a failed send: nothing reached the pane, so the latch must not be set.
contextNotified = notice.isEmpty() ? d.contextNotified() : sent;
}
/**
* Attempt one herdr send and count its outcome. Returns whether {@code agents.send} returned without
* throwing — the caller ({@link #injectNudge}) needs this to decide whether the fleetd #609 context
* latch may be persisted as set.
*/
private boolean trySend(String leadTerminal, String text, String notice) {
try {
agents.send(leadTerminal, fleet.nudgeText());
agents.send(leadTerminal, text);
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
leadTerminal, quietCount);
countNudge("sent");
// fleetd #609: a nudge that carries the context notice is counted under its own outcome so
// it is visible in /metrics — one count per nudge either way, never two.
countNudge(notice.isEmpty() ? "sent" : "sent_context");
return true;
} catch (RuntimeException e) {
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
countNudge("failed");
return false;
}
}
/**
* fleetd #609: the text appended to a nudge when the lead's own context is full — {@code ""}
* whenever the notice does not apply, so callers can unconditionally append this without an extra
* branch. Wording stays plain (CEFR B1) and honest about who actually gates the roll — see the
* {@code requireOperatorConfirm} overload (fleetd #621) for which check that is. This loop only
* ever prints text, it never calls {@code fleet_handover} itself.
*
* @param enabled the {@code leadHeartbeat.contextHighNudge} config flag
* @param reading the lead's current {@link LeadContextGauge} reading
* @return the notice text (starting with a leading space, to append directly after {@link
* FleetState#nudgeText()}), or {@code ""} when disabled or the state is not {@code HIGH}
*/
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading) {
return contextNotice(enabled, reading, false);
}
/**
* fleetd #609 review: as {@link #contextNotice(boolean, LeadContextGauge.Reading)}, but also gated on
* {@code alreadyNotified} — the context latch as it stood <em>before</em> the current tick's decision.
* Without this gate, every pending-driven {@code INJECT} that lands while the context stays {@code
* HIGH} would re-append the full notice on top of an already-latched stretch, making the notice's own
* closing sentence ("You will not be told again until your context reads ok.") false. {@link
* #injectNudge} is the only caller that passes a non-default {@code alreadyNotified}.
*
* <p>fleetd #621: delegates with {@code requireOperatorConfirm=true} — the pre-#621 default and the
* value every existing caller of this overload (including every test written before #621) already
* assumed, so the text this overload returns stays byte-identical.
*
* @param alreadyNotified whether the lead has already been told about the current HIGH stretch
*/
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified) {
return contextNotice(enabled, reading, alreadyNotified, true);
}
/**
* fleetd #621: as {@link #contextNotice(boolean, LeadContextGauge.Reading, boolean)}, but the closing
* instructions also track the daemon's effective {@code leadRollover.requireOperatorConfirm} value,
* instead of always asserting that only the operator can approve the roll.
*
* <p>{@code LeadRollover.confirm(...)} already honours this flag: when it is {@code false}, the daemon
* itself gates the roll on the three handover-file checks alone (exists, modified after the {@code
* open()} request, and no older than {@code maxDocAgeSeconds}) and never consults {@code
* operatorConfirmed}. Before this parameter existed, this notice told the lead to ask the operator
* regardless — so a lead that followed its own instructions asked anyway, and setting the config knob
* to {@code false} stopped the daemon refusing the roll without stopping the operator being
* interrupted. This parameter is how the text is kept honest about which gate is actually live.
*
* @param requireOperatorConfirm the effective {@code leadRollover.requireOperatorConfirm} value
*/
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified,
boolean requireOperatorConfirm) {
if (!enabled || alreadyNotified || reading.state() != LeadContextGauge.State.HIGH) {
return "";
}
StringBuilder sb = new StringBuilder(" Your own context is nearly full");
String compactionWord = reading.compactions() == 1 ? "compaction" : "compactions";
if (reading.tokens() != null) {
sb.append(": ").append(reading.tokens()).append(" tokens used, ")
.append(reading.compactions()).append(' ').append(compactionWord).append(" so far.");
} else {
// A HIGH reading always carries a non-null token count today: LeadContextGauge only
// reaches HIGH by comparing a number against HIGH_THRESHOLD_TOKENS. That invariant
// lives in another class and nothing asserts it, so this branch does not rely on it —
// it drops the token clause rather than printing "null tokens".
sb.append(" (").append(reading.compactions()).append(' ').append(compactionWord)
.append(" so far).");
}
if (requireOperatorConfirm) {
sb.append(" A fresh session would work better. To hand over: call fleet_handover(action=\"open\"), "
+ "write the file it names, ask the operator, then call fleet_handover(action=\"confirm\", "
+ "token, operatorConfirmed). Only the operator can approve the roll. You will not be told "
+ "again until your context reads ok.");
} else {
sb.append(" A fresh session would work better. To hand over: call fleet_handover(action=\"open\"), "
+ "write the file it names, then call fleet_handover(action=\"confirm\", token). Decide for "
+ "yourself when to confirm: the roll goes through if the handover file exists, was "
+ "changed after you opened it, and is not older than maxDocAgeSeconds. You will not be "
+ "told again until your context reads ok.");
}
return sb.toString();
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext() {
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
@@ -63,7 +63,7 @@ import java.util.concurrent.TimeoutException;
* which messages reached a lead. Any publish still awaiting its confirm is failed rather than left to idle out
* the confirm timeout against a sequence number that means nothing on the new channel.
*/
public final class LeadMailbox implements LeadChannel, AutoCloseable {
public final class LeadMailbox implements LeadChannelHandle {
private static final Logger log = LoggerFactory.getLogger(LeadMailbox.class);
@@ -871,10 +871,23 @@ public final class SessionManager implements TurnListener {
// digest lets a lead tell at a glance whether all members got the same charter; the source
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
// charter was composed ("none").
// #604: charterBytes rides along with charterSha256, not with charterSource — it is only
// meaningful as the digest's companion (a length turns "they differ" into "by how much").
// A member with no composed charter reports charterSource and nothing else, as before.
//
// Gating on the DIGEST rather than on the receipt is deliberate, and the two are not
// always null together. CharterReceipt.compose() derives the digest with digestOf(), which
// returns null for BLANK text, while the byte count is composed.getBytes().length, which
// does not. So a whitespace-only charter (a blank fleet.charters.<role> on a profile with
// no MCP, so no reply charter is appended) yields a null digest beside a non-zero size.
// Reporting a size with no digest would say "they differ by N bytes" about an artifact we
// cannot fingerprint, so this gate omits both. Never widen it to the receipt-level null
// check without deciding what that case should report.
if (session.charterReceipt() != null) {
m.put("charterSource", session.charterReceipt().charterSource());
if (session.charterReceipt().charterSha256() != null) {
m.put("charterSha256", session.charterReceipt().charterSha256());
m.put("charterBytes", session.charterReceipt().charterBytes());
}
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
@@ -0,0 +1,222 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.PaneLocator;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* fleetd #612 step 2, unit B2 (CB-185, identity half). Replaces the deleted
* {@code FleetdConnectionIdentityConstructionTest}, which pinned this claim by reading {@code
* Fleetd.java}'s source text for {@code "new PaneLocator(herdr, memberHerdr)"}. That claim moved
* to {@code FleetdAssembly.java} (fleetd #612 Unit A) and is pinned here instead, by driving the
* real {@link ConnectionIdentity} — via {@code runtime.mcp().identity()}, not a copy — that {@link
* FleetdAssembly#assembleAndStart} built.
*
* <p><strong>What this guards against</strong> (from the deleted test's own javadoc): pinning
* {@code PaneLocator} to {@code memberHerdr} alone leaves every LEAD's own MCP connection
* unresolvable ({@code callerTerminal == null}) the moment {@code memberHerdrSocket} names a
* second daemon, which breaks {@code fleet_reply}/{@code fleet_ask}/{@code fleet_whoami} for a
* lead. {@code PaneLocatorTest} already proves {@link PaneLocator} itself can search two clients
* given two — the gap this pins is that the assembly actually passes it two, and in the right
* order (lead first).
*
* <p><strong>Why this cannot be driven through a real MCP/HTTP round trip.</strong> The natural
* way to observe {@code ConnectionIdentity} would be a real {@code fleet_whoami} call over the
* built {@code FleetMcp}, the way {@code FleetMcpContextExtractorTest} drives its own
* hand-built one. That does not work for the REAL assembly, because {@code FleetdAssembly} wires
* {@code ConnectionIdentity} with a hardcoded {@code new LsofPeerPidLookup()} (see {@code
* FleetdAssembly.java:444}), and {@code LsofPeerPidLookup} explicitly excludes its own PID — see
* its javadoc: "we exclude our own PID and take the other end". In a JUnit test the HTTP client
* and the daemon under test run in the very same JVM, so the "client" and "server" ends of the
* loopback connection ARE the same PID, and {@code pidForLocalPort} always returns {@code -1}
* before {@link PaneLocator} is ever reached — proving nothing about which daemon(s) got searched.
* This test instead reaches the real {@link PaneLocator} the assembly built (through {@link
* ConnectionIdentity#panes()}, added for exactly this) and drives it with a chosen pid directly,
* bypassing the OS-dependent PID lookup entirely — a legitimate substitute, since the pid lookup
* is not what CB-185 is about.
*/
class FleetdAssemblyConnectionIdentityTest {
/** Same shape as {@code FleetdAssemblyLifecycleTest}'s fake, but keys {@code connectHerdr} by
* socket path so the lead and member daemons can be two DIFFERENT {@link FakeHerdr}s. */
private static final class TwoHerdrResourcePorts implements ResourcePorts {
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
HerdrClient client = herdrsBySocket.get(socketPath);
if (client == null) {
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
}
return client;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
schedulers.add(scheduler);
return scheduler;
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind — this test never issues a real HTTP request.
}
}
private FleetdRuntime runtime;
private TwoHerdrResourcePorts ports;
@AfterEach
void tearDown() {
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
private static FleetConfig writeConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: "%s"
memberHerdrSocket: "%s"
lifecycle:
idleTtlSeconds: 600
health:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
return FleetConfig.load(f);
}
private FleetdRuntime assemble(Path dir, FakeHerdr lead, FakeHerdr member) throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
ports = new TwoHerdrResourcePorts();
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
return runtime;
}
/**
* The pin. {@code lead} carries the one pane {@link FakeHerdr}'s canned {@code
* pane.process_info} ties to {@link FakeHerdr#WORKER_PID} (pane {@code w2:p7}); {@code member}
* reports NO panes at all ({@link FakeHerdr#withNoPanes()}) — modelling a second daemon that
* simply does not host the caller's pane, exactly the CB-185 javadoc's scenario for a lead's
* own connection. If {@code PaneLocator} only ever searches the member daemon (the bug), this
* pid resolves to nothing, because the pane that owns it lives on the LEAD daemon the bug
* skips.
*/
@Test
void connectionIdentitySearchesTheLeadDaemonNotJustTheMemberOne(@TempDir Path dir) throws Exception {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr().withNoPanes();
assemble(dir, lead, member);
PaneLocator panes = runtime.mcp().identity().panes();
PaneLocator.Lookup lookup = panes.terminalForPid(FakeHerdr.WORKER_PID);
assertEquals("term_a", lookup.terminal(),
"the pane owning WORKER_PID lives on the LEAD daemon only (the member fake reports "
+ "no panes) — PaneLocator must still find it, which is only possible if it "
+ "searches the lead client and not just the member one");
}
/**
* The mirror control: when the pane instead lives ONLY on the member daemon (the lead reports
* no panes), the lookup must still find it — proving the member client is genuinely searched
* too, not merely tolerated as a second, always-losing argument.
*/
@Test
void connectionIdentityAlsoSearchesTheMemberDaemon(@TempDir Path dir) throws Exception {
FakeHerdr lead = new FakeHerdr().withNoPanes();
FakeHerdr member = new FakeHerdr();
assemble(dir, lead, member);
PaneLocator panes = runtime.mcp().identity().panes();
PaneLocator.Lookup lookup = panes.terminalForPid(FakeHerdr.WORKER_PID);
assertEquals("term_a", lookup.terminal(),
"the pane owning WORKER_PID lives on the MEMBER daemon only — PaneLocator must "
+ "find it there too");
}
/** Sanity control: a pid nobody owns resolves to nothing on either daemon. */
@Test
void aPidNoPaneOwnsResolvesToNoTerminalOnEitherDaemon(@TempDir Path dir) throws Exception {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr();
assemble(dir, lead, member);
PaneLocator panes = runtime.mcp().identity().panes();
PaneLocator.Lookup lookup = panes.terminalForPid(999_999L);
assertNull(lookup.terminal());
}
}
@@ -0,0 +1,211 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.LeadChannel;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 A-gaps (gap 1): {@code FleetdAssemblyLifecycleTest}'s own class javadoc says plainly
* that it leaves {@code coordinator:} unset, so {@code leadMailbox} and {@code leadCoordLoop} stay
* {@code null} throughout — the configured-coordinator path is never exercised by Unit A's own
* test. This class drives that path instead: a real {@code coordinator:} block, a fake {@link
* Fleetd.LeadMailboxOpener} returning a fake closeable channel (never a real broker connection),
* and proof that {@link FleetdAssembly#assembleAndStart} both builds it and, on shutdown, closes it.
*
* <p>Made possible by generalising {@code Fleetd.LeadMailboxOpener}'s return type (and {@code
* FleetdRuntime}'s field) from the concrete {@code LeadMailbox} to {@link LeadChannelHandle} — a
* {@link LeadChannel} its owner can also close. {@code FleetMcp} and {@code LeadCoordLoop} already
* consumed the narrower {@link LeadChannel}; this only widens the one seam that owns and closes it.
*/
class FleetdAssemblyCoordinatorLifecycleTest {
private static final String SELF_COORD_ID = "test-lead";
/** A fake {@link LeadChannelHandle}: never touches a broker, and records whether it was closed. */
private static final class FakeLeadChannel implements LeadChannelHandle {
volatile boolean closed = false;
@Override
public void publish(String toCoordId, LeadMessage m) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return SELF_COORD_ID;
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
closed = true;
}
}
/** Minimal fake {@link ResourcePorts}: a real herdr fake, a fake reply inbox, and a real, offered fake lead channel. */
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
final FakeLeadChannel leadChannel = new FakeLeadChannel();
String offeredUri;
String offeredSelfId;
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
offeredUri = uri;
offeredSelfId = selfCoordId;
return leadChannel;
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// No real HTTP bind in a unit test.
}
}
private static final class SentinelReplyInbox implements ReplyInbox {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
coordinator:
uri: "amqp://fake-lead-broker/vh"
selfId: "test-lead"
""");
return FleetConfig.load(f);
}
@Test
void configuredCoordinatorIsBuiltByTheAssemblyAndClosedOnShutdown(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
// --- the assembly actually calls the configured opener and builds the coordinator path ---
assertEquals("amqp://fake-lead-broker/vh", ports.offeredUri,
"the assembly must open the mailbox at the configured broker uri");
assertEquals(SELF_COORD_ID, ports.offeredSelfId,
"the assembly must open the mailbox under the configured selfId");
assertSame(ports.leadChannel, runtime.leadMailbox(),
"FleetdRuntime must own the exact LeadChannelHandle the opener returned, not a copy");
assertNotNull(runtime.leadCoordLoop(),
"a configured coordinator: block must build the receiving LeadCoordLoop too");
assertFalse(ports.leadChannel.closed, "the channel must still be open while the daemon is running");
// --- shutting the assembly down closes it -------------------------------------------------
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
ports.shutdownHook.run();
assertTrue(ports.leadChannel.closed,
"FleetdRuntime.close() must close the configured LeadChannelHandle");
}
}
@@ -0,0 +1,237 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #612 step 2, unit B2 (CB-185, {@code FleetApp} half). Replaces the deleted {@code
* FleetdFleetAppConstructionTest}, which pinned this claim by reading {@code Fleetd.java}'s
* source text for {@code "new FleetApp(herdr, memberHerdr, workers,"}. That claim moved to {@code
* FleetdAssembly.java} (fleetd #612 Unit A) and is pinned here instead, by driving the real {@code
* Javalin} app — via {@code runtime.app()}, not a copy — that {@link
* FleetdAssembly#assembleAndStart} built and handed to {@link FleetdRuntime}.
*
* <p><strong>What this guards against</strong> (from the deleted test's own javadoc): constructing
* {@code FleetApp} with the lead-only {@code herdr} client (dropping {@code memberHerdr}) makes
* {@code GET /healthz} report green while the MEMBER daemon is down — so every spawn fails
* invisibly — and silently drops every member workspace from {@code GET /sessions}. {@code
* FleetAppTwoDaemonTest} already proves {@code FleetApp} itself merges/gates correctly given two
* clients; the gap this pins is that the assembly actually passes it two.
*
* <p><strong>Both directions, not just one</strong> (fleetd #612 issue comment 17525): the deleted
* guard's positive assertion required the exact pair {@code "new FleetApp(herdr, memberHerdr,
* workers,"}, which does not survive EITHER daemon being dropped. An earlier version of this class
* only proved the member-dropped direction, which left {@code new FleetApp(memberHerdr,
* memberHerdr, ...)} — the symmetric bug, {@code /healthz} green while the LEAD daemon is down —
* an undetected regression. {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}
* closes that.
*
* <p>Unlike the {@code ConnectionIdentity} half of CB-185 ({@code
* FleetdAssemblyConnectionIdentityTest}), {@code /healthz} needs no caller identity at all, so
* this test can bind {@link FleetdRuntime#app()} to a REAL ephemeral port (exactly {@code
* FleetAppTwoDaemonTest} does for its own hand-built {@code FleetApp}) and drive it with a real
* {@code HttpClient} — no accessor needed for this half.
*
* <p><strong>{@code GET /sessions} could not be driven the same way</strong>, so this class does
* not pin the merge half of the deleted test's javadoc. {@code /sessions} requires
* {@code Authz.Action.READ}, which — through the REAL assembly's real {@code
* CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded {@code
* new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid from
* {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a JUnit
* test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is always
* {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's fail-closed
* rule) before the route handler — and its {@code memberHerdr} merge — is ever reached. Verified
* directly: driving {@code GET /sessions} here returns {@code 401 unauthenticated}, not the
* merged body. {@code FleetAppTwoDaemonTest} avoids this because it builds {@code FleetApp} with
* {@code callers: null}, which is not what the real assembly passes. The {@code /healthz} pin
* below is what this class relies on for CB-185's {@code FleetApp} half; {@code
* FleetAppTwoDaemonTest} remains the full behavioural proof that {@code FleetApp} itself merges
* {@code /sessions} correctly once handed two clients.
*/
class FleetdAssemblyFleetAppTest {
private static final class TwoHerdrResourcePorts implements ResourcePorts {
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
HerdrClient client = herdrsBySocket.get(socketPath);
if (client == null) {
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
}
return client;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
schedulers.add(scheduler);
return scheduler;
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind here — this test binds runtime.app() itself, for real, below.
}
}
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
private final HttpClient http = HttpClient.newHttpClient();
private FleetdRuntime runtime;
private TwoHerdrResourcePorts ports;
private Javalin boundApp;
@AfterEach
void tearDown() {
if (boundApp != null) {
boundApp.stop();
}
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: "%s"
memberHerdrSocket: "%s"
lifecycle:
idleTtlSeconds: 600
health:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
return FleetConfig.load(f);
}
/** Assembles the real graph, then binds the real {@code Javalin app} to an ephemeral port. */
private int assembleAndBind(Path dir, FakeHerdr lead, FakeHerdr member) throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
ports = new TwoHerdrResourcePorts();
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
boundApp = runtime.app().start("127.0.0.1", 0);
return boundApp.port();
}
private HttpResponse<String> get(int port, String path) throws Exception {
HttpRequest req = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path)).GET().build();
return http.send(req, HttpResponse.BodyHandlers.ofString());
}
/**
* The pin. The MEMBER daemon is down; the LEAD daemon is healthy. If the assembly built
* {@code FleetApp} with only the lead client (the bug: passing {@code herdr} where {@code
* memberHerdr} is expected), the down member is invisible and {@code /healthz} stays 200.
*/
@Test
void healthzGoesRedWhenTheMemberDaemonIsDownEvenThoughTheLeadIsUp(@TempDir Path dir) throws Exception {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr().healthy(false);
int port = assembleAndBind(dir, lead, member);
HttpResponse<String> res = get(port, "/healthz");
assertEquals(503, res.statusCode(),
"a down MEMBER daemon must not be masked by a healthy lead: " + res.body());
}
/** Sanity control: both daemons healthy must still be green through the real assembly. */
@Test
void healthzIsGreenWhenBothDaemonsAreUp(@TempDir Path dir) throws Exception {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr();
int port = assembleAndBind(dir, lead, member);
assertEquals(200, get(port, "/healthz").statusCode());
}
/**
* The symmetric pin (fleetd #612 issue comment 17525): the LEAD daemon is down; the MEMBER
* daemon is healthy. If the assembly built {@code FleetApp} with only the member client
* (dropping {@code herdr} — the mirror of the bug above, {@code new FleetApp(memberHerdr,
* memberHerdr, ...)}), the down LEAD is invisible and {@code /healthz} stays 200. Without this
* case the pair above is one-directional and does not cover the deleted guard's positive
* assertion (it required BOTH {@code herdr,} and {@code memberHerdr,} in that order).
*/
@Test
void healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp(@TempDir Path dir) throws Exception {
FakeHerdr lead = new FakeHerdr().healthy(false);
FakeHerdr member = new FakeHerdr();
int port = assembleAndBind(dir, lead, member);
HttpResponse<String> res = get(port, "/healthz");
assertEquals(503, res.statusCode(),
"a down LEAD daemon must not be masked by a healthy member: " + res.body());
}
}
@@ -0,0 +1,258 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.LoopWatchdog;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 Unit A: proves {@link FleetdAssembly#assembleAndStart} — not a copy of its logic —
* against a real {@link FleetConfig}, a {@link FakeHerdr} and a fake {@link ResourcePorts}, with no
* real herdr socket, no real broker, and no real HTTP bind.
*
* <p><strong>The canonical order this test asserts against was recorded from {@code Fleetd.java}
* BEFORE any code moved</strong> (fleetd #612 Unit A's mandated order of work), by reading the
* original {@code main}'s body and its shutdown-hook {@code Thread}:
*
* <p>Start order: {@code SessionReaper.start()} → {@code StatusPoller.start()} →
* {@code LeadHeartbeatLoop.start()} (opt-in) → {@code FleetHealthMonitor.start()} (opt-in) →
* {@code LeadCoordLoop.start()} (opt-in) → {@code ConfigWatcher.start()} (opt-in) →
* {@code app.start()} (HTTP), always last.
*
* <p>Close order (from the original shutdown hook body): {@code sessions.close(drainTimeoutSeconds)}
* → {@code poller.stop()} → {@code messages.close()} → {@code pushLoop.close()} →
* {@code heartbeat.close()} (if present) → {@code leadCoordLoop.close()} (if present) →
* {@code leadCoordScheduler.shutdownNow()} (if present) → {@code healthMonitor.stop()} (if present)
* → {@code configWatcher.stop()} (if present) → {@code mcp.close()} → {@code reaper.stop()} (if
* present) → {@code idleSleepGuard.close()} (if present) → {@code replyInbox.close()} (if
* {@code AutoCloseable}) → {@code leadMailbox.close()} (if present) → {@code router.close()}.
*
* <p>This test's config deliberately leaves {@code coordinator:} unset, so {@code leadMailbox} and
* {@code leadCoordLoop} stay {@code null} throughout — the lead-mailbox resource-ledger criterion is
* NOT exercised here; see the class-level caveat in the implementer's hand-off. {@code
* idleSleepGuard.enabled: false} is set for the same kind of reason: it would otherwise try to spawn
* a real {@code caffeinate} subprocess, which is not one of the resources the ticket's acceptance
* criteria names (scheduler/inbox/mailbox/loop/MCP server/herdr router).
*/
class FleetdAssemblyLifecycleTest {
/** The one {@link HerdrClient} both {@code herdrSocket} and {@code memberHerdrSocket} resolve to. */
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
final List<String> ledger = new CopyOnWriteArrayList<>();
final List<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
final List<String> schedulerPurposes = new ArrayList<>();
Runnable shutdownHook;
Javalin startedApp;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
ledger.add("connectHerdr");
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
ledger.add("replyInboxOpener");
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
// Never invoked: this test's config has no `coordinator:` block, so
// Fleetd.openLeadMailbox returns null before calling the opener at all.
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
ledger.add("newScheduler:" + purpose);
schedulerPurposes.add(purpose);
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
schedulers.add(scheduler);
return scheduler;
}
@Override
public void addShutdownHook(Runnable hook) {
ledger.add("addShutdownHook");
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Deliberately never call app.start(host, port): no real HTTP bind in a unit test.
ledger.add("startHttp");
this.startedApp = app;
}
}
/** A fake {@link ReplyInbox} that is also {@link AutoCloseable}, so the ledger can prove it closes. */
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
volatile boolean closed = false;
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
closed = true;
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
lifecycle:
idleTtlSeconds: 600
health:
enabled: true
intervalSeconds: 30
leadHeartbeat:
idleAfterSeconds: 600
backoffMs: 15000
quietNudgeCap: 5
configReload:
enabled: true
intervalSeconds: 30
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
""");
return FleetConfig.load(f);
}
@Test
void assemblesTheRealBootGraphWithoutTouchingAnyRealSocketBrokerOrPort(@TempDir Path dir)
throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
// --- start order: the herdr connect, then the recurring background loops, in the recorded
// order, then the shutdown hook is registered, then (finally) HTTP "starts" -------------
assertTrue(ports.ledger.indexOf("connectHerdr") < ports.ledger.indexOf("newScheduler:bridge-push-"),
"herdr must connect before the push scheduler is created: " + ports.ledger);
assertEquals(List.of("bridge-push-", "bridge-heartbeat-", "bridge-health-"), ports.schedulerPurposes,
"the three always-created schedulers must be requested in exactly this order: "
+ ports.schedulerPurposes);
assertTrue(ports.ledger.indexOf("addShutdownHook") < ports.ledger.indexOf("startHttp"),
"the shutdown hook must be registered before HTTP starts — the one statement that "
+ "could not be reordered without changing FleetdRuntime's constructor shape, "
+ "see FleetdAssembly's javadoc: " + ports.ledger);
assertEquals(ports.ledger.size() - 1, ports.ledger.indexOf("startHttp"),
"HTTP must start LAST of everything this fake observes: " + ports.ledger);
assertNotNull(ports.startedApp, "FleetdAssembly must have built and handed off a real FleetApp");
// No real HTTP bind and no real herdr socket: this call returning at all, plus the ledger
// above, is the proof — a real bind or a real UnixSocketHerdrClient.connect would have
// thrown or hung against the sockets/ports this test never opened.
assertNotNull(runtime.app(), "FleetdRuntime must own the same Javalin app that was built");
assertEquals(ports.startedApp, runtime.app(), "attachApp must hand FleetdRuntime the SAME instance");
// --- the runtime owns the real, live objects — not a copy --------------------------------
assertEquals(LoopWatchdog.State.RUNNING, runtime.reaper().health(),
"SessionReaper must be running: lifecycle.idleTtlSeconds is configured");
assertEquals(LoopWatchdog.State.RUNNING, runtime.poller().health(), "StatusPoller must be running");
assertNotNull(runtime.heartbeat(), "leadHeartbeat: is configured, so the loop must be built and started");
assertNotNull(runtime.healthMonitor(), "health.enabled: true, so the monitor must be built and started");
assertNotNull(runtime.configWatcher(), "configReload.enabled: true, so the watcher must be built and started");
// Documented gap (see class javadoc): no coordinator: block, so these stay null.
assertEquals(null, runtime.leadCoordLoop(), "no coordinator: block — leadCoordLoop must stay unbuilt");
assertEquals(null, runtime.leadMailbox(), "no coordinator: block — leadMailbox must stay unbuilt");
assertEquals(runtime.replyInbox(), ports.replyInbox,
"FleetdRuntime must own the exact ReplyInbox instance this fake's AmqpOpener returned");
assertFalse(ports.herdr.closed, "herdr must still be open while the daemon is running");
assertFalse(ports.replyInbox.closed, "the reply inbox must still be open while the daemon is running");
for (ScheduledExecutorService scheduler : ports.schedulers) {
assertFalse(scheduler.isShutdown(), "a scheduler must still be running while the daemon is up");
}
// --- close order: invoke the captured shutdown-hook Runnable directly (no real JVM shutdown
// happens in a unit test) and prove every resource this fake can observe is released --------
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
ports.shutdownHook.run();
assertEquals(LoopWatchdog.State.STOPPED, runtime.reaper().health(), "SessionReaper must stop on close");
assertEquals(LoopWatchdog.State.STOPPED, runtime.poller().health(), "StatusPoller must stop on close");
assertTrue(ports.herdr.closed, "router.close() must close the herdr client last");
assertTrue(ports.replyInbox.closed, "the AutoCloseable reply inbox must be closed");
for (int i = 0; i < ports.schedulers.size(); i++) {
assertTrue(ports.schedulers.get(i).isShutdown(),
"scheduler for '" + ports.schedulerPurposes.get(i) + "' must be shut down by close(): "
+ "LeadHeartbeatLoop/FleetHealthMonitor/ReplyPushLoop each call "
+ "scheduler.shutdownNow() on the exact instance ports.newScheduler(...) handed them");
}
// Calling the captured hook a second time must never happen for a real JVM shutdown hook,
// but nothing above should have thrown — that already proves every accessed field's close()
// tolerated running once, in the recorded order, without an exception escaping.
}
}
@@ -0,0 +1,165 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.testing.CapturedLog;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 A-gaps (gap 2): {@code Fleetd.java:185} used to call {@code
* reportRoleFallbackGaps(cfg)} <em>outside</em> the boundary {@code FleetdAssembly.assembleAndStart}
* — the assembly call itself sat at line 201, after it — so nothing that drives the assembly (the
* seam {@code FleetdAssemblyLifecycleTest} exercises) could ever notice the call being deleted.
* {@code Fleetd.main} still refuses a bad config end to end (see {@code
* FleetdStartupValidationTest}), but that only pins {@code validateAll()} and {@code
* assertChartersNameOnlyRegisteredTools} — a validator that <em>throws</em>. {@code
* reportRoleFallbackGaps} only logs; nothing about {@code main} throwing or not throwing can
* observe whether that particular call ran.
*
* <p>This drives {@link FleetdAssembly#assembleAndStart} directly — never a copy of its logic —
* with a config that has no {@code fleet:} pools or charters configured for any role, so every role
* trips both of {@code reportRoleFallbackGaps}' log branches, and asserts on the real log line a
* {@code ListAppender} attached to the shared {@code Fleetd}/{@code FleetdAssembly} logger
* captures. Deleting the call from {@code FleetdAssembly} (verified by hand, see the ticket) turns
* this test red; deleting it from {@code Fleetd.main} instead (its old location) would not, which
* is exactly the gap this test closes.
*/
class FleetdAssemblyRoleFallbackBoundaryTest {
/** Minimal fake {@link ResourcePorts}: enough for {@code assembleAndStart} to run with no real I/O. */
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// No real HTTP bind in a unit test.
}
}
private static final class SentinelReplyInbox implements ReplyInbox {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
// Deliberately no `fleet:` block at all: every MemberRole has neither a pool nor a
// charter, so reportRoleFallbackGaps' "role fallback: no fleet.<role>s: pool for ..."
// branch is guaranteed to log something to assert on.
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
""");
return FleetConfig.load(f);
}
@Test
void assembleAndStartItselfReportsTheRoleFallbackGaps(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
try (var captured = CapturedLog.at(Fleetd.class, Level.INFO)) {
FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
String infoLines = captured.events().stream()
.filter(e -> e.getLevel() == Level.INFO)
.map(ILoggingEvent::getFormattedMessage)
.reduce("", (a, b) -> a + "\n" + b);
assertTrue(infoLines.contains("role fallback: no fleet.<role>s: pool for"),
() -> "FleetdAssembly.assembleAndStart itself must call reportRoleFallbackGaps "
+ "(fleetd #612 A-gaps gap 2) — captured INFO lines: " + infoLines);
} finally {
if (ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
}
}
@@ -0,0 +1,199 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.placement.BackendQuarantine;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.OptionalLong;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 B3 — replaces {@code FleetdBackendQuarantineWiringTest} (fleetd #466), a source-text
* test that scraped {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by fleetd #612
* Unit A) for the {@code BackendQuarantine.withEscalation(...)} call, and separately asserted the
* flat two-argument constructor's text was ABSENT. That proves the right method NAME appears in
* source; it proves nothing about what the constructed object actually DOES.
*
* <p>This test instead drives the REAL {@link BackendQuarantine} the real {@link
* FleetdAssembly#assembleAndStart} builds — reached through {@link
* dev.ltms.fleet.mcp.FleetMcp#quarantineSource()} on the real, live {@code FleetMcp} {@code
* FleetdRuntime} owns — and asserts the ONE behavioural difference {@code withEscalation} and the
* flat constructor actually produce (see {@link BackendQuarantine}'s own class doc, "Mechanism"):
* quarantining the same credential twice in a row, within one base cooldown of the first deadline,
* must escalate the second cooldown past the first. A flat instance reports the identical cooldown
* both times.
*/
class FleetdBackendQuarantineAssemblyTest {
/** Base cooldown used throughout — long enough that rounding never blurs the 2x escalation. */
private static final int COOLDOWN_SECONDS = 100;
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
final AtomicLong clockNanos = new AtomicLong(0L);
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
// Controllable: the SAME LongSupplier instance BackendQuarantine.withEscalation(...) is
// built with, so advancing clockNanos after assembly moves the quarantine tracker's own
// clock, with no real sleep needed to observe escalation.
return clockNanos::get;
}
@Override
public LongSupplier wallClockNanos() {
return clockNanos::get;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
quarantineCooldownSeconds: %d
""".formatted(COOLDOWN_SECONDS));
return FleetConfig.load(f);
}
@Test
@DisplayName("[BEHAVIOURAL] the real assembled BackendQuarantine escalates a repeated exhaustion, "
+ "which the flat two-argument constructor can never do")
void assembledQuarantineEscalatesOnARepeatedExhaustion(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
// First exhaustion, at clock=0: a fresh occurrence, blocked for exactly the base cooldown.
quarantine.quarantine("cred-x");
BackendQuarantine.Status first = quarantine.status("cred-x").orElseThrow(
() -> new AssertionError("credential must be quarantined immediately after quarantine()"));
assertEquals(1, first.repeatCount(), "the first call is repeat #1");
assertEquals(COOLDOWN_SECONDS, first.remainingSeconds(),
"a fresh quarantine blocks for exactly the base cooldown");
// Second exhaustion, arriving just after the first deadline — well within one base cooldown
// of it, so this is a CONTINUATION of the same streak (repeat #2), not a fresh occurrence.
long firstDeadlineNanos = COOLDOWN_SECONDS * 1_000_000_000L;
ports.clockNanos.set(firstDeadlineNanos + 1);
quarantine.quarantine("cred-x");
BackendQuarantine.Status second = quarantine.status("cred-x").orElseThrow(
() -> new AssertionError("credential must be quarantined immediately after the second "
+ "quarantine() call"));
assertEquals(2, second.repeatCount(), "the second call, arriving within one base cooldown of "
+ "the first deadline, continues the streak as repeat #2");
// The one behavioural difference: withEscalation doubles the cooldown on repeat #2 (capped
// well above this at 12x base), the flat two-argument constructor never grows past the base
// cooldown no matter how many times quarantine() is called in a row.
assertEquals(2 * COOLDOWN_SECONDS, second.remainingSeconds(),
"withEscalation's default backoff doubles the cooldown on the second consecutive "
+ "exhaustion — this is the exact call FleetdAssembly.java makes at the "
+ "BackendQuarantine.withEscalation(...) call site");
assertTrue(second.remainingSeconds() > first.remainingSeconds(),
"the flat two-argument BackendQuarantine constructor would report the SAME remaining "
+ "seconds both times — this inequality is what a mutation to the flat "
+ "constructor at that call site must fail");
// Also confirm isQuarantined/remainingSeconds agree, exercising the accessors a real caller
// (fleet_profiles / fleet_list, per BackendQuarantine's own class doc) actually reads.
assertTrue(quarantine.isQuarantined("cred-x"));
OptionalLong remaining = quarantine.remainingSeconds("cred-x");
assertTrue(remaining.isPresent());
assertEquals(2 * COOLDOWN_SECONDS, remaining.getAsLong());
}
}
@@ -1,65 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
* it says nothing about which one {@code main} actually calls.
*
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
* every one of them builds its own instance directly. That silent regression is exactly the shape
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
* factory choice, following their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
* constructor. It does not prove that call actually executes at startup (no test here starts the
* daemon), and it does not prove the escalation reaches a real backend or credential — only
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
* the wiring runs.
*/
class FleetdBackendQuarantineWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"Fleetd.main's BackendQuarantine local must still be built from "
+ "BackendQuarantine.withEscalation(System::nanoTime, "
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
+ "proves the escalation itself works — this source check is what must go red instead. "
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
+ "flat ~30-minute cooldown, about 336 times across the week.");
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
assertFalse(source.contains(
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
}
}
@@ -0,0 +1,307 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.WorktreeRequest;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 step 2, Unit B1 — replaces {@code FleetdCompletionResolverWiringTest} (deleted in
* this same commit), whose four tests read {@code Fleetd.java}'s source text and asserted the
* {@code CompletionResolver} construction call still named the right arguments. That proved the
* call site's spelling, never that the assembled resolver actually behaves differently when an
* argument is dropped.
*
* <p>These tests drive {@link FleetdAssembly#assembleAndStart} — the real boot composition,
* fleetd #612 Unit A — and read {@link FleetdRuntime#completion()}: the exact {@link
* CompletionResolver} instance the assembled daemon uses, never a copy built alongside it for the
* test's benefit. Two behaviours are pinned, matching the ticket's own two measured mutations:
*
* <ul>
* <li>the 8th constructor argument ({@code Fleetd.worktreeBranchLookup(sessions::roster)}) —
* {@link #assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport}; and</li>
* <li>the 5th/6th arguments ({@code backendErrorPatterns}, {@code backendErrorSink}, both
* assigned from {@code Fleetd}'s extracted factories rather than an inline lambda) —
* {@link #assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern}.</li>
* </ul>
*
* <p>Both tests bypass {@link dev.ltms.fleet.inject.StatusPoller} and drive {@link
* CompletionResolver#onDelivered} / {@link CompletionResolver#resolveBeforePostAction} directly —
* the same public, synchronous entry points {@code CompletionResolverTest} uses — with a
* hand-built {@link CompletableFuture} waiter, so no real poller loop or herdr status poll is
* needed. The pane scrape comes from {@link FakeHerdr#readText}; the elapsed-time floor
* ({@code CompletionResolver.MIN_TURN_NANOS}) is controlled via a fake, advanceable {@link
* ResourcePorts#nanoClock()} rather than a real sleep.
*/
class FleetdCompletionResolverAssemblyTest {
/** Same shape as {@code FleetdAssemblyLifecycleTest}'s fake, plus a nanoClock this test can advance. */
private static final class ControllableResourcePorts implements ResourcePorts {
final FakeHerdr herdr;
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
Runnable shutdownHook;
ControllableResourcePorts(FakeHerdr herdr) {
this.herdr = herdr;
}
void advanceSeconds(long seconds) {
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
}
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
// Never invoked: this test's config has no `broker:` block, so Fleetd.selectReplyInbox
// returns the in-memory inbox before calling the opener at all.
return (uri, prefetch) -> {
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
// Never invoked: no `coordinator:` block configured either.
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
};
}
@Override
public LongSupplier nanoClock() {
return nowNanos::get;
}
@Override
public LongSupplier wallClockNanos() {
return nowNanos::get;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind a real port.
}
}
private static FleetConfig writeConfig(Path dir, String profilesYaml, String extraGuardHost,
String worktreeRootYamlLine) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
%s
profiles:
%s
guard:
offSubscriptionHosts:
- %s
""".formatted(worktreeRootYamlLine == null ? "" : worktreeRootYamlLine, profilesYaml, extraGuardHost));
return FleetConfig.load(f);
}
private static void gitQuiet(Path cwd, String... args) throws Exception {
List<String> cmd = new java.util.ArrayList<>(List.of("git"));
cmd.addAll(List.of(args));
Process p = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: git " + String.join(" ", args));
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
}
private static Path initRepo(Path dir) throws Exception {
Files.createDirectories(dir);
gitQuiet(dir, "init", "-q", "-b", "main");
gitQuiet(dir, "config", "user.email", "test@example.invalid");
gitQuiet(dir, "config", "user.name", "Test");
Files.writeString(dir.resolve("README.md"), "seed\n");
gitQuiet(dir, "add", "README.md");
gitQuiet(dir, "commit", "-q", "-m", "seed");
return dir;
}
/**
* fleetd #248's first measured mutation: replacing {@code CompletionResolver}'s 8th constructor
* argument with the inert {@code _ -> null} compiles clean and leaves every existing test green
* — it silently drops fleetd #241's fallback-report location. This drives the real assembled
* resolver through a member echoing its own injected brief back (no {@code fleet_reply}), which
* resolves via {@code noReportMessage(target)}, and proves the real member's {@code branch} —
* only obtainable via {@code Fleetd.worktreeBranchLookup(sessions::roster)} reading the real,
* worktree-provisioned {@link MemberSession} — appears in the reported text.
*/
@Test
void assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport(@TempDir Path dir) throws Exception {
Path repo = initRepo(dir.resolve("repo"));
FleetConfig cfg = writeConfig(dir, """
wtprofile:
baseUrl: http://wthost.local:8000
model: sonnet
""", "wthost.local", "worktreeRoot: " + dir.resolve("wts"));
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
MemberSession session = runtime.sessions().acquire("wtprofile", repo.toString(), repo.toString(),
null, new WorktreeRequest("fleetd-612-b1", null));
String target = session.terminalId();
String branch = session.branch();
assertTrue(branch != null && branch.startsWith("worker/"),
"sanity: a worktree-provisioned session must carry a real branch, got: " + branch);
CompletionResolver completion = runtime.completion();
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
String echoedBrief = "z".repeat(450); // >= CompletionResolver.ECHO_MIN_CHARS normalised chars
ports.herdr.readText("idle, nothing yet");
completion.onDelivered(target, new TurnToken(target, waiter, echoedBrief));
ports.herdr.readText(echoedBrief); // the pane just echoes the injected brief back — no real report
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s) without a real sleep
completion.resolveBeforePostAction(target);
Rendezvous.Resolution resolution = waiter.getNow(null);
assertTrue(resolution != null, "the waiter must have resolved synchronously");
assertEquals(Rendezvous.Kind.COMPLETION, resolution.kind());
assertTrue(resolution.text().contains(CompletionResolver.NO_REPORT_PREFIX),
"sanity: must have gone down the noReportMessage sub-path: " + resolution.text());
assertTrue(resolution.text().contains("branch=" + branch),
"the assembled resolver must report the member's real branch (fleetd #241 via "
+ "fleetd #248's worktreeBranchLookup wiring); got: " + resolution.text());
} finally {
if (ports.shutdownHook != null) ports.shutdownHook.run();
}
}
/**
* fleetd #248's second measured mutation, and fleetd#201 Unit 5's own gap: replacing {@code
* backendErrorPatterns}/{@code backendErrorSink} with {@code BackendErrorPatternLookup.legacy()}
* / {@code BackendErrorSink.none()} compiles clean and leaves every existing behavioural test
* green.
*
* <p>Classification proof: this test's profile configures {@code errorPattern: "credential
* outage"} — text the built-in {@code (?i)\bAPI Error\s*:} fallback ({@code legacy()}'s only
* behaviour) never matches. So a real {@code Fleetd.backendErrorPatternLookup(...)} wiring
* classifies the send as {@code FAILED}; {@code legacy()} would fall through to the plain
* completion path instead ({@code Kind.COMPLETION}).
*
* <p>Cool-off proof: two distinct targets on the same profile/credential each classified as a
* backend error inside the 60s window must cool the credential off ({@link
* dev.ltms.fleet.placement.BackendOutagePolicy}, fleetd#201 Unit 5) — observable two ways: (1)
* the real {@code Fleetd.backendErrorSink(...)} marks each session {@code BACKEND_ERROR} (only
* the real sink calls {@code sessions.onBackendError}; {@code BackendErrorSink.none()} never
* does), and (2) a third explicit-profile spawn attempt is refused with a {@link
* PlacementException} naming the cool-off — only reachable because the real sink's {@code
* outagePolicy.record(...)} call actually ran.
*/
@Test
void assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir, """
coolprofile:
baseUrl: http://coolhost.local:8000
model: sonnet
errorPattern: "credential outage"
""", "coolhost.local", null);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
MemberSession session1 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
MemberSession session2 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
String target1 = session1.terminalId();
String target2 = session2.terminalId();
assertTrue(!target1.equals(target2), "sanity: the two spawns must be distinct targets");
CompletionResolver completion = runtime.completion();
CompletableFuture<Rendezvous.Resolution> waiter1 = new CompletableFuture<>();
ports.herdr.readText("idle 1");
completion.onDelivered(target1, new TurnToken(target1, waiter1, null));
ports.herdr.readText("credential outage: upstream 503");
ports.advanceSeconds(3);
completion.resolveBeforePostAction(target1);
Rendezvous.Resolution resolution1 = waiter1.getNow(null);
assertTrue(resolution1 != null, "target1's waiter must have resolved synchronously");
assertEquals(Rendezvous.Kind.FAILED, resolution1.kind(),
"a configured errorPattern the built-in fallback never matches must classify as "
+ "a backend error, not a plain completion; got: " + resolution1);
assertTrue(resolution1.text().contains("credential outage: upstream 503"), resolution1.text());
CompletableFuture<Rendezvous.Resolution> waiter2 = new CompletableFuture<>();
ports.herdr.readText("idle 2");
completion.onDelivered(target2, new TurnToken(target2, waiter2, null));
ports.herdr.readText("credential outage: upstream 503 again");
ports.advanceSeconds(3);
completion.resolveBeforePostAction(target2);
Rendezvous.Resolution resolution2 = waiter2.getNow(null);
assertTrue(resolution2 != null, "target2's waiter must have resolved synchronously");
assertEquals(Rendezvous.Kind.FAILED, resolution2.kind());
List<MemberSession> roster = runtime.sessions().roster();
assertTrue(roster.stream().anyMatch(s -> target1.equals(s.terminalId())
&& s.state() == MemberSession.State.BACKEND_ERROR),
"the real backendErrorSink must have transitioned target1 to BACKEND_ERROR: " + roster);
assertTrue(roster.stream().anyMatch(s -> target2.equals(s.terminalId())
&& s.state() == MemberSession.State.BACKEND_ERROR),
"the real backendErrorSink must have transitioned target2 to BACKEND_ERROR: " + roster);
PlacementException coolOff = assertThrows(PlacementException.class,
() -> runtime.sessions().acquire("coolprofile", null, dir.toString(), null),
"two distinct targets classified within the 60s window must cool the credential "
+ "off (BackendOutagePolicy), refusing a third explicit-profile spawn");
assertTrue(coolOff.getMessage().contains("cooling off"), coolOff.getMessage());
} finally {
if (ports.shutdownHook != null) ports.shutdownHook.run();
}
}
}
@@ -1,89 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #248: this is the test that was actually missing. {@code Fleetd.main} builds its {@code
* CompletionResolver} from an 8-argument constructor, and the ticket's own measurement proved two
* ways to silently unwire it — both compiled with 0 errors and left every existing test green:
*
* <ul>
* <li>replacing the worktree/branch argument (the 8th) with {@code _ -> null} — drops
* fleetd#241's fallback-report location entirely;</li>
* <li>replacing {@code backendErrorPatterns, backendErrorSink} (5th/6th) with {@code
* BackendErrorPatternLookup.legacy(), BackendErrorSink.none()} — drops fleetd#201 Unit 5's
* backend-error classification and cool-off entirely.</li>
* </ul>
*
* <p>Neither mutation could be caught by any test that constructs its own {@code
* CompletionResolver} (every test before this one did exactly that) or by a test of {@link
* Fleetd#worktreeBranchLookup}, {@link Fleetd#backendErrorPatternLookup}, or {@link
* Fleetd#backendErrorSink} in isolation (see {@code FleetdWorktreeBranchLookupTest}, {@code
* FleetdBackendErrorPatternLookupTest}, {@code FleetdBackendErrorSinkTest}) — those prove the
* factories work, never that {@code main} still calls them. This class is a plain source-text
* assertion on {@code Fleetd.java} — crude, but honest about what it checks, and it turns red the
* instant the wiring is dropped, mirroring the same fallback shape {@link
* FleetdFleetAppConstructionTest} already uses for a different constructor argument.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* CompletionResolver} and never runs {@code main}.
*/
class FleetdCompletionResolverWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still names backendErrorPatterns and backendErrorSink")
void backendErrorArgumentsAreStillNamedAtTheCallSite() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"exhaustionSink, backendErrorPatterns, backendErrorSink, System::nanoTime,"),
"CompletionResolver's construction call must still pass backendErrorPatterns and "
+ "backendErrorSink as its 5th/6th arguments. Replacing them with "
+ "BackendErrorPatternLookup.legacy()/BackendErrorSink.none() (fleetd #248's measured "
+ "mutation) compiles with 0 errors and leaves every behavioural test green — this "
+ "source check is what must go red instead.");
}
@Test
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still passes worktreeBranchLookup(sessions::roster)")
void worktreeBranchLookupIsStillPassedAtTheCallSite() throws Exception {
String source = fleetdSource();
assertTrue(source.contains("worktreeBranchLookup(sessions::roster)"),
"CompletionResolver's construction call must still pass worktreeBranchLookup(sessions::roster) "
+ "as its 8th (last) argument. Replacing it with the inert `_ -> null` (fleetd #248's "
+ "other measured mutation) compiles with 0 errors and leaves every behavioural test "
+ "green — this source check is what must go red instead.");
assertFalse(source.contains("System::nanoTime,\n _ -> null"),
"the worktree/branch argument must never regress to the inert `_ -> null` literal");
}
@Test
@DisplayName("[SOURCE TEXT] backendErrorPatterns is assigned from the extracted backendErrorPatternLookup(...) factory")
void backendErrorPatternsComesFromTheFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,"),
"backendErrorPatterns must be assigned from Fleetd.backendErrorPatternLookup(...), not an "
+ "inline lambda that a source check on the CompletionResolver call alone cannot see "
+ "through");
}
@Test
@DisplayName("[SOURCE TEXT] backendErrorSink is assigned from the extracted backendErrorSink(...) factory")
void backendErrorSinkComesFromTheFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"BackendErrorSink backendErrorSink = backendErrorSink(sessions, () -> config.get().profiles(),"),
"backendErrorSink must be assigned from Fleetd.backendErrorSink(...), not an inline lambda "
+ "that a source check on the CompletionResolver call alone cannot see through");
}
}
@@ -1,31 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-185: {@code ConnectionIdentity} must resolve a caller's pane on EITHER herdr daemon (a
* lead's MCP connection resolves against the lead daemon; a member's against the member daemon).
* Pinning {@code PaneLocator} to {@code memberHerdr} alone — the bug this guards against — leaves
* every lead's own connection unresolvable ({@code callerTerminal == null}) the moment
* {@code memberHerdrSocket} names a second daemon, which breaks {@code fleet_reply}/{@code
* fleet_ask} and {@code fleet_whoami} for a lead. A unit test on {@link
* dev.ltms.fleet.herdr.PaneLocator} alone (see {@code PaneLocatorTest}) proves the class CAN
* search two clients, but not that {@code Fleetd.main} actually wires it that way — hence this
* source-level assertion, the same technique {@code FleetdHerdrControlConstructionTest} uses.
*/
class FleetdConnectionIdentityConstructionTest {
@Test
void connectionIdentitySearchesBothDaemonsNotJustTheMemberOne() throws Exception {
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
assertFalse(source.contains("new PaneLocator(memberHerdr)"),
"PaneLocator must not be pinned to the member daemon alone — a lead's own "
+ "connection resolves against the LEAD daemon and would never be found");
assertTrue(source.contains("new PaneLocator(herdr, memberHerdr)"),
"PaneLocator must search the lead daemon first, then the member daemon");
}
}
@@ -1,29 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-185: {@code FleetApp} must be constructed with BOTH herdr clients (the lead's and the
* member's), never the raw lead-only {@code herdr}. Passing only {@code herdr} — the bug this
* guards against — makes {@code GET /healthz} green while the member daemon is down (so every
* spawn fails invisibly) and silently drops every member workspace from {@code GET /sessions}.
* A behavioural test on {@code FleetApp} alone (see {@code FleetAppTwoDaemonTest}) proves the
* class merges/gates correctly when given two clients, but not that {@code Fleetd.main} actually
* passes it two — hence this source-level assertion, mirroring
* {@code FleetdHerdrControlConstructionTest}.
*/
class FleetdFleetAppConstructionTest {
@Test
void fleetAppIsConstructedWithBothHerdrDaemons() throws Exception {
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
assertFalse(source.contains("new FleetApp(herdr, workers,"),
"FleetApp must not be constructed with the lead-only herdr client");
assertTrue(source.contains("new FleetApp(herdr, memberHerdr, workers,"),
"FleetApp must be constructed with both the lead and the member herdr client");
}
}
@@ -0,0 +1,100 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #602 gauge-wiring: {@link Fleetd#leadConfigDirLookup} is the factory {@code Fleetd.main}
* wires into {@code FleetMcp.LeadConfigDirSource} so {@code fleet_list}'s {@code context} row reads
* the transcript directory a lead's OWN profile actually writes to, instead of always falling back
* to the built-in {@code <user.home>/.claude} default (see {@code LeadContextGauge}).
*
* <p>{@code FleetMcpLeadContextGaugeWiringTest} proves the directory this factory returns is what
* actually gets read; this class proves the factory's own matching logic — the same
* {@code fleet.leaders.<name>.profile} link {@code Fleetd.leadSeatLookup} already follows (see
* {@code FleetdLeadSeatLookupTest}), one step further to that profile's own {@code configDir:}.
*/
class FleetdLeadConfigDirLookupTest {
private static FleetConfig.Profile profileWithConfigDir(String name, String configDir) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", configDir, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, null, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("a lead on a profile that sets configDir resolves to that directory")
void leadOnAProfileWithConfigDirResolvesToIt() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
assertEquals("/mnt/opus-claude", lookup.apply("primary"));
}
@Test
@DisplayName("changing the config from directory A to directory B changes what the lookup reports")
void configChangeFromDirectoryAToDirectoryBChangesTheAnswer() {
java.util.concurrent.atomic.AtomicReference<Map<String, FleetConfig.Profile>> profilesRef =
new java.util.concurrent.atomic.AtomicReference<>(
Map.of("opus", profileWithConfigDir("opus", "/mnt/dir-a")));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, String> lookup = Fleetd.leadConfigDirLookup(profilesRef::get, leaders);
assertEquals("/mnt/dir-a", lookup.apply("primary"), "must read directory A before the config changes");
profilesRef.set(Map.of("opus", profileWithConfigDir("opus", "/mnt/dir-b")));
assertEquals("/mnt/dir-b", lookup.apply("primary"), "must read directory B once the LIVE config changes — "
+ "a lookup that snapshotted the profile map at construction would still answer directory A here");
}
@Test
@DisplayName("a lead entry with no `profile:` (recognise-only) resolves to null, not a thrown exception")
void recogniseOnlyLeadWithNoProfileResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
"claude", "claude-sonnet-5");
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("a lead naming a profile that is not configured resolves to null, not a thrown exception")
void leadOnAnUnconfiguredProfileResolvesToNull() {
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("ghost-profile"));
Function<String, String> lookup = Fleetd.leadConfigDirLookup(Map::of, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("a lead on a profile that sets no configDir override resolves to null")
void leadOnAProfileWithNoConfigDirResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
void unrecognisedLeadNameResolvesToNull() {
Function<String, String> lookup = Fleetd.leadConfigDirLookup(Map::of, Map.of());
assertNull(lookup.apply("ghost-lead"));
}
}
@@ -0,0 +1,98 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* fleetd #602 gauge-wiring follow-up (PR #606 review comment 17353): {@code Fleetd.main}'s {@code
* LeadConfigDirSource} local used to be a bare {@code new FleetMcp.LeadConfigDirSource(
* leadConfigDirLookup(...))} built inline, with nothing a test could call directly. Measured on
* that shape: replacing the whole expression with {@code FleetMcp.LeadConfigDirSource.none()} at
* the call site compiled with 0 errors and left the full 1822-test suite green — the daemon could
* be changed to always report every lead's context as {@code UNKNOWN}, forever, and no test would
* notice. That is the same hand-built-vs-config-wired shape as fleetd #561/#248/#426/#562
* ({@code FleetdLoopHealthSourceWiringTest}).
*
* <p>{@code FleetMcpLeadContextGaugeWiringTest} and {@code FleetdLeadConfigDirLookupTest} both
* predate this class and are both still correct — but neither can catch the mutation above. One
* builds its own {@code FleetMcp} and hands it its own {@code LeadConfigDirSource}; the other
* builds its own lookup and calls {@link Fleetd#leadConfigDirLookup} directly. Neither one ever
* calls the thing {@code Fleetd.main} actually calls.
*
* <p>The fix extracts the inline {@code new} into {@link Fleetd#leadConfigDirSource}, a
* package-private factory in the same style as {@link Fleetd#loopHealthSource}/{@link
* Fleetd#capacitySource}/{@link Fleetd#healthCoverageSource} — which is exactly what makes it
* directly callable here. This test calls that factory with real {@link FleetConfig.Profile}/
* {@link FleetConfig.Leader} fixtures (the same shapes {@code FleetdLeadConfigDirLookupTest}
* already uses) and asserts the returned source resolves a real {@code configDir} — a property
* that would be false if {@link Fleetd#leadConfigDirSource} were mutated to {@code return
* FleetMcp.LeadConfigDirSource.none();}. Measured: mutating exactly that line makes
* {@link #resolvesTheRealConfiguredConfigDir()} fail ({@code expected: </mnt/opus-claude> but was:
* <null>}); restoring it makes the whole suite green again.
*
* <p><b>What this class does not and cannot cover.</b> {@code main}'s own line —
* {@code leadConfigDirSource(() -> config.get().profiles(), leaders)} — could itself be swapped
* for a bare {@code FleetMcp.LeadConfigDirSource.none()}, bypassing this factory entirely. Measured:
* that exact mutation compiles with 0 errors and leaves every test in this file, and the full
* 1825-test suite, green. {@link Fleetd#loopHealthSource}'s own wiring test has the identical gap
* for its own one-line call in {@code main} — no test in this codebase calls {@code Fleetd.main}
* far enough to observe which factory call it made. This class narrows the gap from "nothing tests
* the wiring" (the pre-extraction state this ticket found) to "the factory's own logic is pinned,
* and main's call to it is a one-line, visually-verifiable delegation" — the same standard already
* accepted for {@code loopHealthSource}/{@code capacitySource}/{@code healthCoverageSource}.
*/
class FleetdLeadConfigDirSourceWiringTest {
private static FleetConfig.Profile profileWithConfigDir(String name, String configDir) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", configDir, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, null, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("the returned source resolves the lead's REAL configured configDir, not a hardcoded null")
void resolvesTheRealConfiguredConfigDir() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertEquals("/mnt/opus-claude", source.configDirFor().apply("primary"),
"the configDirFor function must delegate to the real leadConfigDirLookup — mutating "
+ "Fleetd.leadConfigDirSource's own body to `return FleetMcp.LeadConfigDirSource.none();` "
+ "must fail this assertion (measured: it does — see this test's class javadoc for the "
+ "companion measurement on main's one-line call to this factory, which this assertion "
+ "does not and structurally cannot cover)");
}
@Test
@DisplayName("a lead on a profile with no configDir override still resolves to null, not a crash")
void leadWithNoConfigDirOverrideResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertNull(source.configDirFor().apply("primary"),
"no configDir: override configured ⇒ null, so LeadContextGauge falls back to its own default");
}
@Test
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
void unrecognisedLeadNameResolvesToNull() {
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(Map::of, Map.of());
assertNull(source.configDirFor().apply("ghost-lead"));
}
}
@@ -0,0 +1,156 @@
package dev.ltms.fleet;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.lead.LeadContextGauge;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* fleetd #609: {@link Fleetd#leadContextLookup} is the factory {@code Fleetd.main} wires into
* {@code LeadHeartbeatLoop.LeadContextSource} so the idle-lead heartbeat can read a lead's own
* {@link LeadContextGauge} reading by its terminal id. Three hops: terminal → lead name (unknown ⇒
* UNKNOWN), lead name → configDir, and {@code agents.get(terminal)} for the live session id and
* agent type the gauge itself needs — a herdr failure on that last hop must degrade to UNKNOWN, not
* throw and kill the heartbeat's own tick.
*
* <p>{@code AgentControl.agentCall} resolves a {@code term_}-prefixed target's pane id via a first
* {@code agent.list} round trip (see its own javadoc); this class's terminal id deliberately does
* NOT start with {@code term_} so the stub {@link HerdrClient} below only needs to answer
* {@code agent.get} — the one call this factory actually depends on.
*/
class FleetdLeadContextLookupTest {
private static final String LEAD_TERMINAL = "leadpane1";
private static final String LEAD_NAME = "opus";
private static final String SESSION_ID = "sess-609-happy-path";
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
}
/** An {@link AgentControl} whose every {@code agent.get} answers with the given session/type/status. */
private static AgentControl agentControlStub(String sessionId, String agentType, String status) {
HerdrClient client = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
if (!"agent.get".equals(method)) {
throw new HerdrException("stub has no canned response for " + method);
}
try {
String sessionField = sessionId == null ? ""
: ",\"agent_session\":{\"kind\":\"id\",\"value\":\"" + sessionId + "\"}";
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
return new ObjectMapper().readTree(("""
{"type":"agent_info","agent":{"terminal_id":"%s","agent":%s,
"agent_status":"%s"%s}}""")
.formatted(LEAD_TERMINAL, agentField, status, sessionField));
} catch (Exception e) {
throw new HerdrException("stub decode failed", e);
}
}
@Override
public void close() {
}
};
return new AgentControl(client);
}
/** An {@link AgentControl} whose every herdr call fails — models a herdr hiccup mid-tick. */
private static AgentControl throwingAgentControl() {
HerdrClient client = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
throw new HerdrException("herdr unreachable (stub)");
}
@Override
public void close() {
}
};
return new AgentControl(client);
}
@Test
@DisplayName("an unrecognised terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply("ghost-terminal"));
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
}
@Test
@DisplayName("agents.get throwing degrades to UNKNOWN, no exception escapes")
void agentsGetThrowingDegradesToUnknown() {
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply(LEAD_TERMINAL));
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"a herdr failure resolving the live agent must degrade to UNKNOWN, never kill the heartbeat's tick");
}
@Test
@DisplayName("the happy path resolves the configDir, sessionId and agentType through to the gauge")
void happyPathPassesResolvedFactsThroughToTheGauge(@TempDir Path tmp) throws IOException {
writeTranscript(tmp, SESSION_ID, 12_345);
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
AgentControl agents = agentControlStub(SESSION_ID, "claude", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.OK, reading.state(),
"the resolved configDir + the live agent's own sessionId/agentType must reach the gauge — "
+ "a wrong hop anywhere in the chain would read no transcript and report UNKNOWN instead");
assertEquals(12_345L, reading.tokens());
}
@Test
@DisplayName("a lead whose agent type is not claude still resolves to UNKNOWN, never a crash")
void nonClaudeAgentTypeResolvesToUnknown(@TempDir Path tmp) throws IOException {
writeTranscript(tmp, SESSION_ID, 12_345);
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
AgentControl agents = agentControlStub(SESSION_ID, "opencode", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"the agentType hop must reach the gauge too — a non-claude peer must not be misread as claude");
}
}
@@ -0,0 +1,108 @@
package dev.ltms.fleet;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #609: {@code Fleetd.main}'s {@code LeadHeartbeatLoop.LeadContextSource} local could be
* swapped for a bare {@code LeadHeartbeatLoop.LeadContextSource.none()} at the call site — compiling
* with 0 errors and leaving every pre-existing test green — exactly the shape #602/#606 already found
* for {@code LeadConfigDirSource} (see {@code FleetdLeadConfigDirSourceWiringTest}'s own javadoc for
* the measured version of that gap).
*
* <p>The fix follows the same pattern: {@link Fleetd#leadContextSource} is the extracted,
* directly-callable factory {@code main} calls to build the source it hands {@code
* LeadHeartbeatLoop}'s constructor. This test calls that exact factory and asserts it resolves a
* REAL reading off a real transcript file — a property that would be false if {@link
* Fleetd#leadContextSource} were mutated to {@code return LeadHeartbeatLoop.LeadContextSource.none();}.
*
* <p>What this class does not and cannot cover: {@code main}'s own one-line call to this factory
* could itself be swapped for {@code LeadHeartbeatLoop.LeadContextSource.none()}, bypassing this
* factory entirely — the same structural gap {@code FleetdLeadConfigDirSourceWiringTest} names for
* its own factory, and for the same reason (no test in this codebase calls {@code Fleetd.main} far
* enough to observe which factory call it made).
*/
class FleetdLeadContextSourceWiringTest {
/** Deliberately not {@code term_}-prefixed — see {@code FleetdLeadContextLookupTest}'s class doc. */
private static final String LEAD_TERMINAL = "leadpane1";
private static final String LEAD_NAME = "opus";
private static final String SESSION_ID = "sess-609-wiring";
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
}
private static AgentControl agentControlStub() {
HerdrClient client = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
if (!"agent.get".equals(method)) {
throw new HerdrException("stub has no canned response for " + method);
}
try {
return new ObjectMapper().readTree(("""
{"type":"agent_info","agent":{"terminal_id":"%s","agent":"claude",
"agent_status":"idle","agent_session":{"kind":"id","value":"%s"}}}""")
.formatted(LEAD_TERMINAL, SESSION_ID));
} catch (Exception e) {
throw new HerdrException("stub decode failed", e);
}
}
@Override
public void close() {
}
};
return new AgentControl(client);
}
@Test
@DisplayName("main's factory resolves a REAL reading, not the inert none() answer")
void resolvesARealReadingNotTheInertNoneAnswer(@TempDir Path tmp) throws IOException {
writeTranscript(tmp, SESSION_ID, 54_321);
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), () -> liveLeadTerminals, name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.OK, reading.state(),
"mutating Fleetd.leadContextSource's own body to `return LeadHeartbeatLoop.LeadContextSource.none();` "
+ "must fail this assertion");
assertEquals(54_321L, reading.tokens());
}
@Test
@DisplayName("an unrecognised lead terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), Map::of, name -> null);
assertEquals(LeadContextGauge.State.UNKNOWN, source.readingFor().apply("ghost-terminal").state());
}
}
@@ -3,6 +3,7 @@ package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadMailbox;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
@@ -35,7 +36,7 @@ class FleetdLeadMailboxSelectionTest {
boolean unreachable;
@Override
public LeadMailbox open(String uri, String selfCoordId, int prefetch) {
public LeadChannelHandle open(String uri, String selfCoordId, int prefetch) {
this.offeredUri = uri;
this.offeredSelfId = selfCoordId;
this.offeredPrefetch = prefetch;
@@ -0,0 +1,285 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.fail;
/**
* fleetd #612 B3 — replaces {@code FleetdLeadRolloverWiringTest} (fleetd #480). That class was a
* source-text test scraping {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
* not an independent claim — needs no replacement of its own), {@code
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
* vacuous, {@code assertNull(null)}, and never calls the real factory).
*
* <p><strong>fleetd #612 B3 correction (ticket comment 17553):</strong> the first version of this
* test configured a single shared {@link FakeHerdr} for both the lead and member herdr sockets.
* {@code FleetdAssembly.java:140-142} falls back to {@code memberHerdr = herdr} whenever no
* distinct {@code memberHerdrSocket} is configured, so with one fake, {@code
* router.leadAgents()} and {@code router.memberAgents()} wrapped the identical client — a
* mutation swapping {@code Fleetd.leadRollover(cfg, router.leadAgents(), config, leads)} for
* {@code ..., router.memberAgents(), ...} at {@code FleetdAssembly.java:408} was therefore
* invisible to this test, even though the two are genuinely different daemons in production. This
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake
* and never on the MEMBER one.
*/
class FleetdLeadRolloverAssemblyTest {
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
/** Keys {@code connectHerdr} by socket path so the lead and member daemons can be two
* DIFFERENT {@link FakeHerdr}s — same shape as B2's {@code FleetdAssemblyConnectionIdentityTest
* .TwoHerdrResourcePorts}. */
private static final class RecordingResourcePorts implements ResourcePorts {
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
HerdrClient client = herdrsBySocket.get(socketPath);
if (client == null) {
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
}
return client;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
}
}
private static FleetConfig writeConfig(Path dir, Path leadCwd) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: "%s"
memberHerdrSocket: "%s"
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
leadRollover:
handoverPath: handover.md
requireOperatorConfirm: false
""".formatted(LEAD_SOCKET, MEMBER_SOCKET, leadCwd.toString()));
return FleetConfig.load(f);
}
@SuppressWarnings("unchecked")
@Test
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation "
+ "sequence — /clear, then bootstrapText — through the real herdr router")
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
FleetConfig cfg = writeConfig(dir, leadCwd);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
// Two DISTINCT fakes — one per configured socket — so leadAgents()/memberAgents() wrap
// genuinely different clients, exactly like production when memberHerdrSocket is set.
FakeHerdr lead = new FakeHerdr();
lead.withTab("w2", "w2:t7", "lead: opus");
FakeHerdr member = new FakeHerdr();
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
LeadRollover rollover = runtime.mcp().leadRollover();
assertNotNull(rollover, "leadRollover: is present in this test's config, so "
+ "FleetdAssembly.assembleAndStart must have built a real LeadRollover through the "
+ "Fleetd.leadRollover(...) call site — a mutation to `LeadRollover leadRollover = "
+ "null;` at that call site can never pass this");
LeadRollover.PendingRollover pending = rollover.open("term_a", "fleetd #612 B3 test");
String expectedHandoverPath = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expectedHandoverPath, pending.handoverPath());
// Ensure the handover file's mtime lands strictly AFTER open()'s requestedAtMillis —
// LeadRollover.checkHandover refuses on mtime <= requestedAt (HANDOVER_STALE).
Thread.sleep(50);
Files.writeString(Path.of(pending.handoverPath()), "handover content for fleetd #612 B3");
LeadRollover.RollDecision decision = rollover.confirm("term_a", pending.token(), true);
assertTrue(decision.accepted(), "confirm() must approve: requireOperatorConfirm is false, "
+ "the caller terminal matches open()'s, and the handover file exists, is non-empty "
+ "and fresh — got: " + decision);
// The production LeadRollover constructor always runs the post-confirm continuation on a
// real virtual thread (see Fleetd.leadRollover, which never passes the package-private test
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
// until the real continuation finishes.
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
+ "the turn-boundary wait settles immediately and the post-/clear wait "
+ "releases via its pickup-grace path — detail: " + status.detail());
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD
// pane — this is the one thing a source-text pin on the call site could never show.
List<FakeHerdr.Call> prompts = lead.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.toList();
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send "
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts);
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
"the first send must be the literal /clear housekeeping command");
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
"the second send must be the default bootstrapText naming the resolved handover "
+ "path, got: " + secondText);
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
// swapping router.leadAgents() for router.memberAgents() at the real call site would move
// both sends above onto `member` instead, which this assertion catches — the thing the
// single-fake version of this test could never see, because both wrapped the same client.
List<FakeHerdr.Call> memberPrompts = member.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.toList();
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: "
+ memberPrompts);
}
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
throws InterruptedException {
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10);
while (System.nanoTime() < deadline) {
LeadRollover.RollStatus status = rollover.status(token);
if (status.state() != LeadRollover.RollState.PENDING
&& status.state() != LeadRollover.RollState.IN_PROGRESS) {
return status;
}
Thread.sleep(50);
}
fail("the real continuation did not reach a terminal state within 10s — last status: "
+ rollover.status(token));
throw new AssertionError("unreachable");
}
@Test
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) returns null when leadRollover: is absent "
+ "from config — the opt-in gate FleetdLeadRolloverWiringTest's "
+ "factoryGatesOnConfigPresence pinned by source text alone")
void absentLeadRolloverConfigMeansNoRolloverIsBuilt(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
host: 127.0.0.1
port: 8765
""");
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
AgentControl agents = new AgentControl(new FakeHerdr());
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of);
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
+ "LeadRollover must be constructed at all");
}
}
@@ -1,89 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #480 Unit A, hard requirement 6: pin {@code Fleetd.main}'s construction of {@link
* dev.ltms.fleet.lead.LeadRollover} with a source-text assertion, mirroring {@code
* FleetdCompletionResolverWiringTest}'s pattern — five log-only reporters in {@code Fleetd.main}
* already survived mutation batteries this exact way (fleetd #415's extraction antidote note).
*
* <p>What this class still covers, and what it never claimed to. {@code LeadRolloverTest}
* constructs its own {@code LeadRollover} directly (as every prior test of an extracted factory
* does) with a hand-built lookup, so a mutation that deletes the {@code leadRollover(...)} call
* from {@code main} — or replaces one of its arguments with something that still compiles, e.g.
* {@code router.leadAgents()} swapped for {@code null}, or the whole assignment swapped for a bare
* {@code null} literal — leaves every behavioural test green. This is a plain string read, guarded
* by an unrelated anchor assertion so a broken or empty file read cannot pass as a real change.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* LeadRollover} and never runs {@code main}. It pins the {@code leadRollover(...)} CALL SITE's
* argument list — that {@code main} still passes {@code leads} at all — never what the factory
* DOES with that argument once inside its own body.
*
* <p><b>Correction (fleetd #480 relative-handover-path follow-up): that gap used to be real, and
* now is not — but not here.</b> This class's javadoc previously claimed "no behavioural test can
* catch this wiring dropping out" for the whole factory, including the lambda {@code
* leadRollover(...)} builds internally (terminal → lead name → {@code Leader.cwd()}). That claim
* was proven true at the time — mutating that lambda's body to {@code String leadName = null;}
* (always "no lead found", which silently reintroduces the daemon-cwd bug this ticket fixes) left
* the full suite green, {@code Tests run: 1669, Failures: 0}. It is no longer true: {@code
* FleetdLeadRolloverWorkspaceLookupTest} now calls {@code Fleetd.leadRollover(...)} directly with a
* real {@link dev.ltms.fleet.config.ConfigRef} built from a temp {@code fleetd.yaml}, and fails
* against that exact one-line mutation. So: THIS class still covers only the call site's argument
* list; {@code FleetdLeadRolloverWorkspaceLookupTest} is what now covers the lambda's body. Neither
* one subsumes the other — keep both.
*/
class FleetdLeadRolloverWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] unrelated anchor: Fleetd.java still declares the Fleetd class")
void unrelatedAnchorStillPresent() throws Exception {
// Guards the two assertions below: without this, a bad read (empty string, wrong file,
// truncated file) could vacuously fail to contain the leadRollover(...) call too, and a
// test that only asserts "contains X" would report a false pass for the wrong reason if X
// happened to match. Asserting an unrelated, structurally distant string first proves the
// read actually pulled real file content.
String source = fleetdSource();
assertTrue(source.contains("public final class Fleetd"),
"sanity anchor failed — the file read did not return real Fleetd.java source; the "
+ "leadRollover(...) wiring assertions below cannot be trusted until this passes");
}
@Test
@DisplayName("[SOURCE TEXT] main still constructs LeadRollover via the leadRollover(...) factory, exactly as heartbeat is constructed")
void mainStillCallsTheLeadRolloverFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);"),
"Fleetd.main must still assign `LeadRollover leadRollover = leadRollover(cfg, "
+ "router.leadAgents(), config, leads);`. Dropping this call, or swapping one of "
+ "its arguments for something that still compiles (e.g. null in place of "
+ "router.leadAgents()), leaves every behavioural test green — this source check is "
+ "what must go red instead. fleetd #480 correction 2 deliberately dropped "
+ "primaryRegistry from this call — see LeadRollover's class javadoc for why a "
+ "single-slot lookup was wrong here. The fleetd #480 relative-handover-path "
+ "follow-up added `leads` (terminal → lead name) so the factory can resolve a "
+ "relative handoverPath against the calling lead's own workspace.");
}
@Test
@DisplayName("[SOURCE TEXT] the leadRollover(...) factory itself gates construction on cfg.leadRollover() != null")
void factoryGatesOnConfigPresence() throws Exception {
String source = fleetdSource();
assertTrue(source.contains("if (cfg.leadRollover() == null) {"),
"Fleetd.leadRollover(...) must refuse to construct a LeadRollover when the "
+ "leadRollover: block is absent — an upgraded daemon must never silently acquire "
+ "the ability to clear the lead's own pane. See LeadHeartbeatLoop's construction "
+ "gate (cfg.leadHeartbeat() != null) for the pattern this mirrors.");
}
}
@@ -0,0 +1,172 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #612 B3 — replaces {@code FleetdLeadSeatWiringTest} (fleetd #176), a source-text test that
* scraped {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by fleetd #612 Unit A)
* for the exact {@code new FleetMcp.LeadSeatSource(Fleetd.leadSeatLookup(...))} constructor-call
* text. That proves the right symbols appear in source; it proves nothing about what the daemon's
* live {@code fleet_list} actually reports.
*
* <p>This test instead drives the REAL {@link FleetMcp.LeadSeatSource} the real {@link
* FleetdAssembly#assembleAndStart} builds — including the REAL {@code LeadTabScanner} it wires
* {@code Fleetd.leadSeatLookup} through — reached via {@link FleetMcp#leadSeatSource()} on the
* live {@code FleetMcp} {@code FleetdRuntime} owns. It seeds one FakeHerdr tab labelled to match a
* configured {@code fleet.leaders.opus.tab}, with a live agent already in it (FakeHerdr's own
* default {@code agent.list}/{@code pane.list} entries for {@code term_a}/{@code w2:p7}/{@code
* w2:t7} — no FakeHerdr change needed), and asserts the assembled seat source reports exactly the
* seat {@link FleetMcp.LeadSeatSource#none()} (the inert stand-in) could never produce: 1, not 0.
*/
class FleetdLeadSeatAssemblyTest {
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
fleet:
leaders:
opus:
tab: "lead: opus"
profile: sonnet
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""");
return FleetConfig.load(f);
}
@Test
@DisplayName("[BEHAVIOURAL] the real assembled LeadSeatSource, backed by the real LeadTabScanner, "
+ "reports a live lead's seat against its own subscription profile")
void assembledLeadSeatSourceReportsALiveLeadsSeat(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
// agent) to match fleet.leaders.opus.tab exactly — no FakeHerdr change needed at all.
ports.herdr.withTab("w2", "w2:t7", "lead: opus");
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
FleetMcp.LeadSeatSource seatSource = runtime.mcp().leadSeatSource();
assertEquals(1, seatSource.seatsFor().apply("sonnet"),
"the real LeadTabScanner recognises the labelled tab as a live 'opus' lead on "
+ "profile 'sonnet' (same credential, matched by Fleetd.leadSeatLookup), so "
+ "subscription profile 'sonnet' must be charged one seat — "
+ "FleetMcp.LeadSeatSource.none() (the inert stand-in this test's mutation "
+ "swaps the call site for) always reports 0, whatever the input");
// A profile no lead is running on gets no seat charged — the same seat source, applied to
// an input that must stay at the inert answer even on the real, non-inert instance.
assertEquals(0, seatSource.seatsFor().apply("no-such-profile"));
}
}
@@ -1,43 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #176: {@code Fleetd.main} builds its {@code FleetMcp} from a 14-argument constructor whose
* last argument is a {@code FleetMcp.LeadSeatSource} wrapping {@link Fleetd#leadSeatLookup}. That
* argument is exactly the kind of wiring fleetd #248 warned about: dropping it (or swapping it for
* the inert {@code FleetMcp.LeadSeatSource.none()}) compiles with 0 errors and leaves every test
* that builds its own {@code FleetMcp}/{@code CapacitySource} directly — every test that predates
* this ticket — green, because none of them go through {@code main} at all.
*
* <p>{@link FleetdLeadSeatLookupTest} proves the factory's own matching logic; this class is the
* plain source-text assertion that proves {@code main} still passes its result in, mirroring
* {@code FleetdCompletionResolverWiringTest}'s approach for the same class of gap.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a
* {@code FleetMcp} and never runs {@code main}.
*/
class FleetdLeadSeatWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] FleetMcp's construction call still passes a LeadSeatSource built from leadSeatLookup(...)")
void fleetMcpConstructionStillWiresLeadSeatLookup() throws Exception {
String source = fleetdSource();
assertTrue(source.contains("new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), "
+ "leaders, leads))"),
"FleetMcp's construction call must still pass a LeadSeatSource built from "
+ "Fleetd.leadSeatLookup(...). Dropping it or swapping in "
+ "FleetMcp.LeadSeatSource.none() (fleetd #176's would-be silent regression, the same "
+ "shape as fleetd #248's measured mutations) compiles with 0 errors and leaves every "
+ "existing behavioural test green — this source check is what must go red instead.");
}
}
@@ -0,0 +1,252 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #613: a {@code MemberRole} with no {@code fleet.<role>s:} pool falls back to
* <em>every</em> configured profile ({@code FleetConfig#candidateProfiles}), and one with no
* {@code fleet.charters.<role>:} entry runs with only the launcher's reply charter. Both are
* deliberate, legitimate states — neither is refused by {@code validateMembers()} — but both were
* silent at boot. On the host that opened this ticket, an unqualified {@code hunter} spawn silently
* widened to all 8 configured profiles and its resolved first choice was {@code local}, a profile
* every other pool on that same config gave weight 0 to.
*
* <p>{@link Fleetd#reportRoleFallbackGaps} must name every gapped role, and for a pool gap, the
* exact resolved first-choice profile — that number, not the pool size, is what actually surprised
* the operator. Mirrors {@link ExhaustedPatternGapReportTest}'s pattern: capture the real log via a
* {@link ListAppender} rather than asserting on the call site's source text.
*/
class RoleFallbackGapReportTest {
private static FleetConfig load(Path dir, String yaml) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml);
return FleetConfig.load(f);
}
/**
* The level this logger had before {@link #attach()} raised it, so {@link #detach} can put it
* back. {@code null} means "inherit from the parent" — the state this logger starts in.
*/
private static Level originalLevel;
/**
* {@code reportRoleFallbackGaps} logs at INFO, and {@code logback-test.xml} sets
* {@code dev.ltms.fleet} to WARN — so INFO events are dropped by the level check before any
* appender sees them. Raise the level for the duration of the test, exactly like {@code
* GitHostShapeReportTest#attach}.
*/
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
originalLevel = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
logger.detachAppender(appender);
logger.setLevel(originalLevel);
}
private static List<String> infoMessages(ListAppender<ILoggingEvent> appender) {
return appender.list.stream()
.filter(e -> e.getLevel() == Level.INFO)
.map(ILoggingEvent::getFormattedMessage)
.toList();
}
/**
* Reproduces the shape measured in the ticket: {@code dev}, {@code reviewer} and
* {@code architect} each have a pool and a charter; {@code hunter} has neither. The pool-gap
* line must name {@code hunter}, the profile count (3), and the resolved first choice
* ({@code local}, the first profile in definition order) — and must not name the three healthy
* roles. The charter-gap line must separately name only {@code hunter}.
*/
@Test
void hunterWithNoPoolOrCharterIsNamedWithItsResolvedFirstChoice(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
sonnet:
baseUrl: https://llm.ltms.dev/v1
terra:
baseUrl: https://llm.ltms.dev/v1
fleet:
developers:
a:
profile: sonnet
reviewers:
b:
profile: terra
architects:
c:
profile: sonnet
charters:
dev: "dev charter text"
reviewer: "reviewer charter text"
architect: "architect charter text"
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportRoleFallbackGaps(cfg);
} finally {
detach(appender);
}
List<String> infos = infoMessages(appender);
String poolLine = infos.stream()
.filter(m -> m.contains("no fleet.<role>s: pool"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a pool-gap INFO line: " + infos));
assertTrue(poolLine.contains("hunter"), poolLine);
assertTrue(poolLine.contains("3"), "must name the profile count the gap opens onto: " + poolLine);
assertTrue(poolLine.contains("'local'"),
"must name the resolved first-choice profile, the number that actually surprised "
+ "the operator: " + poolLine);
for (String healthy : List.of("dev", "reviewer", "architect")) {
assertFalse(poolLine.contains(healthy),
"pool-gap line must not name a role that has a pool: " + poolLine);
}
String charterLine = infos.stream()
.filter(m -> m.contains("no fleet.charters: entry"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a charter-gap INFO line: " + infos));
assertTrue(charterLine.contains("hunter"), charterLine);
for (String healthy : List.of("dev", "reviewer", "architect")) {
assertFalse(charterLine.contains(healthy),
"charter-gap line must not name a role that has a charter: " + charterLine);
}
}
/**
* The resolved first choice must be genuinely computed from definition order, not hardcoded —
* reordering {@code profiles:} so a different entry comes first changes the reported choice.
*/
@Test
void theResolvedFirstChoiceFollowsProfileDefinitionOrder(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
sonnet:
baseUrl: https://llm.ltms.dev/v1
local:
baseUrl: http://gx00.gw:8000
fleet:
developers:
a:
profile: sonnet
charters:
dev: "dev charter text"
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportRoleFallbackGaps(cfg);
} finally {
detach(appender);
}
String poolLine = infoMessages(appender).stream()
.filter(m -> m.contains("no fleet.<role>s: pool"))
.findFirst()
.orElseThrow();
// hunter, reviewer and architect all lack a pool here; each falls back to the full 2-profile
// set and the first-choice is 'sonnet' because it is first in profiles: definition order.
assertTrue(poolLine.contains("'sonnet'"), poolLine);
assertFalse(poolLine.contains("'local'"), poolLine);
}
/** A config with a pool and a charter for every role produces no role-fallback log at all. */
@Test
void everyRoleWithAPoolAndACharterProducesNoLogAtAll(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
sonnet:
baseUrl: https://llm.ltms.dev/v1
fleet:
developers:
a:
profile: sonnet
reviewers:
b:
profile: sonnet
hunters:
c:
profile: sonnet
architects:
d:
profile: sonnet
charters:
dev: "dev charter text"
reviewer: "reviewer charter text"
hunter: "hunter charter text"
architect: "architect charter text"
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportRoleFallbackGaps(cfg);
} finally {
detach(appender);
}
assertTrue(appender.list.isEmpty(),
"a config with no gaps must not print a per-role block: " + infoMessages(appender));
}
/**
* A config with only {@code profiles:} and no {@code fleet:} block at all must still be
* reported (every role is gapped, both pool and charter) rather than throwing — this is the
* exact shape {@code candidateProfiles}' fallback exists to keep starting.
*/
@Test
void aConfigWithNoFleetBlockAtAllReportsEveryRoleGapped(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
sonnet:
baseUrl: https://llm.ltms.dev/v1
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportRoleFallbackGaps(cfg);
} finally {
detach(appender);
}
List<String> infos = infoMessages(appender);
String poolLine = infos.stream()
.filter(m -> m.contains("no fleet.<role>s: pool"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a pool-gap INFO line: " + infos));
String charterLine = infos.stream()
.filter(m -> m.contains("no fleet.charters: entry"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a charter-gap INFO line: " + infos));
for (String role : List.of("dev", "hunter", "reviewer", "architect")) {
assertTrue(poolLine.contains(role), poolLine);
assertTrue(charterLine.contains(role), charterLine);
}
}
}
@@ -1,8 +1,11 @@
package dev.ltms.fleet.config;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.msg.LeadMailbox;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
@@ -93,6 +96,72 @@ class FleetConfigTest {
"unset means off — today's behaviour, unchanged");
}
/**
* fleetd #601 review (measured 2026-09-22): this guard used to throw {@link
* IllegalStateException} and refuse to start. On a host with the conflict configured, that
* turned into a launchd restart loop with no readable cause, since {@code fleetd.yaml} is
* gitignored. It must now WARN and let the daemon start, and the warning must carry both values
* so an operator can fix the config without reading the source. Pinning the log line itself (via
* {@link CapturedLog}) rather than a snippet of production source text — the latter is the
* anti-pattern this repo avoids; the former is the actual observable behaviour a reader (or an
* alert on the log) depends on.
*/
@Test
void aClaudeProfileWithConflictingAutoCompactFlagAndEnvironmentWindowLoadsAndWarnsWithBothValues(
@TempDir Path dir) throws Exception {
Path f = dir.resolve("conflicting-auto-compact-window.yaml");
Files.writeString(f, """
profiles:
claude-profile:
autoCompactWindow: 250000
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "300000"
""");
FleetConfig cfg;
List<String> warnings;
try (CapturedLog log = CapturedLog.at(FleetConfig.class, Level.WARN)) {
cfg = FleetConfig.load(f);
warnings = log.events().stream().map(ILoggingEvent::getFormattedMessage).toList();
}
assertEquals(250_000, cfg.profiles().get("claude-profile").autoCompactWindow(),
"the disagreement is reported, not corrected — the flag value still loads as-is");
assertEquals(1, warnings.size(), "exactly one warning for the one conflicting profile: " + warnings);
String warning = warnings.get(0);
assertTrue(warning.contains("claude-profile"), "names the offending profile: " + warning);
assertTrue(warning.contains("autoCompactWindow=250000"), "names the flag value: " + warning);
assertTrue(warning.contains("CLAUDE_CODE_AUTO_COMPACT_WINDOW=300000"), "names the env value: " + warning);
}
/**
* The negative probe paired with the test above (per fleetd #601 review): a warning that fires
* on every load and a warning that never fires read the same from a single test, so both must be
* checked. No conflict here — the flag and the env value agree — so no warning should be logged.
*/
@Test
void aClaudeProfileWithEqualAutoCompactFlagAndEnvironmentWindowLoadsWithNoWarning(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("equal-auto-compact-window.yaml");
Files.writeString(f, """
profiles:
claude-profile:
autoCompactWindow: 250000
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "250000"
""");
FleetConfig cfg;
List<ILoggingEvent> events;
try (CapturedLog log = CapturedLog.at(FleetConfig.class, Level.WARN)) {
cfg = FleetConfig.load(f);
events = log.events();
}
assertEquals(250_000, cfg.profiles().get("claude-profile").autoCompactWindow());
assertTrue(events.isEmpty(), "equal values must not warn: " + events);
}
// ── fleetd #201 Unit 5: errorPattern ────────────────────────────────────────────────────────
@Test
@@ -456,7 +456,11 @@ public final class FakeHerdr implements HerdrClient {
}
}
/** fleetd #612 Unit A: set by {@link #close()} so a test can prove a router/client actually closed this. */
public volatile boolean closed = false;
@Override
public void close() {
closed = true;
}
}
@@ -0,0 +1,248 @@
package dev.ltms.fleet.lead;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.io.RandomAccessFile;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assumptions.assumeFalse;
/**
* Ticket "lead context gauge" — fleetd could not see how full a lead's own Claude Code context
* window was, so a lead auto-compacting was always a surprise. {@link LeadContextGauge} reads that
* off the lead's own transcript. Five acceptance properties from the ticket, one test class each
* (plus the three separate UNKNOWN cases the ticket calls out by name):
* <ol>
* <li>{@link #tokenCountTracksTheLastUsageRecordAndChangesWithIt()}</li>
* <li>{@link #compactionCountTracksCompactBoundaryRecordsAndChangesWithIt()}</li>
* <li>{@link #missingFileIsUnknown()}, {@link #unreadableFileIsUnknown()}</li>
* <li>{@link #tornFinalLineFallsBackToTheLastGoodReading()},
* {@link #everyLineUnparseableIsUnknown()} (fleetd #602 gauge-wiring, finding 2)</li>
* <li>{@link #readNeverExceedsTheTailBound()}</li>
* <li>{@link #secondReadInsideTtlDoesNotTouchDiskAgain()}</li>
* </ol>
*
* <p>Every fixture lives under {@code @TempDir} — never the operator's real config directory (see
* the ticket's hard constraint on this).
*
* <p><strong>fleetd #602 gauge-wiring, finding 2.</strong> The property above used to be "a
* truncated or invalid LAST line reports UNKNOWN" — on the theory that a bad last line signals a
* format change. That reasoning did not hold: {@code fleet_list} reads this transcript while Claude
* Code may be mid-write on it, so a torn LAST line is an ordinary race, not a format change, and a
* real format change makes EVERY line unparseable, not only the last one written. So the property
* is now split in two: {@link #tornFinalLineFallsBackToTheLastGoodReading} (the torn-line case must
* NOT destroy a good earlier reading) and its control, {@link #everyLineUnparseableIsUnknown} (only
* when NOTHING in the window parses does this report UNKNOWN).
*/
class LeadContextGaugeTest {
private static final String SESSION_ID = "11111111-1111-1111-1111-111111111111";
/** One line Claude Code would write for a turn with the given live-context total. */
private static String usageLine(long inputTokens, long cacheRead, long cacheCreation) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + inputTokens + ","
+ "\"cache_read_input_tokens\":" + cacheRead + ","
+ "\"cache_creation_input_tokens\":" + cacheCreation + "}}}";
}
/** One line Claude Code writes when an auto-compaction happens. */
private static String compactionLine() {
return "{\"type\":\"system\",\"subtype\":\"compact_boundary\","
+ "\"compactMetadata\":{\"preTokens\":250000,\"postTokens\":5000,"
+ "\"trigger\":\"auto\",\"durationMs\":54000}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} and writes {@code lines}. */
private static String writeTranscript(Path configDir, String sessionId, String... lines) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Path file = projectDir.resolve(sessionId + ".jsonl");
StringBuilder sb = new StringBuilder();
for (String line : lines) {
sb.append(line).append('\n');
}
Files.writeString(file, sb.toString(), StandardCharsets.UTF_8);
return configDir.toString();
}
// --- property 1: live token count tracks the last usage record -----------------------------
@Test
@DisplayName("reported tokens equal the LAST usage record's total, and change when it does")
void tokenCountTracksTheLastUsageRecordAndChangesWithIt(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID,
usageLine(1_000, 0, 0),
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
assertEquals(LeadContextGauge.State.OK, first.state());
// Change N in the fixture (a fresh session id avoids the cache) -- the reported number
// must change with it, not stay pinned to the first fixture's total.
String otherSession = "22222222-2222-2222-2222-222222222222";
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
assertEquals(200_000L, second.tokens());
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
}
// --- property 2: compaction count tracks compact_boundary records --------------------------
@Test
@DisplayName("reported compaction count equals K compact_boundary records, and changes with K")
void compactionCountTracksCompactBoundaryRecordsAndChangesWithIt(@TempDir Path tmp) throws IOException {
String sessionTwoCompactions = "33333333-3333-3333-3333-333333333333";
writeTranscript(tmp, sessionTwoCompactions,
usageLine(1_000, 0, 0),
compactionLine(),
usageLine(2_000, 0, 0),
compactionLine(),
usageLine(3_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
assertEquals(2, twoCompactions.compactions());
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
}
// --- property 3: two separate UNKNOWN cases -------------------------------------------------
@Test
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
void missingFileIsUnknown(@TempDir Path tmp) {
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
}
@Test
@DisplayName("an unreadable transcript file reports UNKNOWN with no token number")
void unreadableFileIsUnknown(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
Path file = tmp.resolve("projects").resolve("some-project-slug").resolve(SESSION_ID + ".jsonl");
assertTrue(file.toFile().setReadable(false), "test setup: must be able to revoke read permission");
try {
// setReadable(false) really did clear the read bit (asserted above), but that alone
// does not prove the file is UNREADABLE: running as root (e.g. a CI container) ignores
// the read bit and opens the file anyway. Files.isReadable checks what actually happens
// on open, not the bit. When it still reports readable, this test cannot create the
// condition it needs on this machine, so it skips honestly instead of asserting on a
// state that was never reached. A skip here means "I could not set up the case", NOT
// "the UNKNOWN behaviour is fine" -- it is not evidence either way.
assumeFalse(Files.isReadable(file),
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
} finally {
file.toFile().setReadable(true); // so @TempDir cleanup can delete it
}
}
@Test
@DisplayName("fleetd #602 finding 2: a torn final line does not destroy a good earlier reading")
void tornFinalLineFallsBackToTheLastGoodReading(@TempDir Path tmp) throws IOException {
// Deliberately NOT via writeTranscript: that helper appends a trailing newline after every
// line, including the last one, which would not model what a write caught mid-flush looks
// like on disk. Claude Code appends and flushes one line at a time, so the torn line here has
// no trailing newline at all -- exactly the shape a read racing an in-progress write sees.
Path projectDir = tmp.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Path file = projectDir.resolve(SESSION_ID + ".jsonl");
String lastCompleteLine = usageLine(1_000, 2_000, 3_000); // last COMPLETE record: 6,000
String tornLine = "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{\"input_tok";
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.OK, reading.state(),
"a torn final line must not turn a good earlier reading into UNKNOWN");
assertEquals(6_000L, reading.tokens(),
"must report the last COMPLETE line's total, not fail the whole read");
}
@Test
@DisplayName("fleetd #602 finding 2 control: when EVERY line is unparseable, the reading IS UNKNOWN")
void everyLineUnparseableIsUnknown(@TempDir Path tmp) throws IOException {
// The real signal a format change gives: not one bad line, but ALL of them. Without this
// control, code that never reports UNKNOWN at all would still pass the torn-line property.
String configDir = writeTranscript(tmp, SESSION_ID,
"{this is not json at all",
"neither is this{{{");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"every line unparseable is the real format-change signal and must still report UNKNOWN");
assertNull(reading.tokens());
}
// --- property 4: the tail bound is actually enforced ----------------------------------------
@Test
@DisplayName("reading a transcript far larger than the tail bound reads no more than the bound")
void readNeverExceedsTheTailBound(@TempDir Path tmp) throws IOException {
int smallBound = 4_096; // exercise the mechanism without writing a multi-MB fixture
Path file = tmp.resolve("big.jsonl");
// One line far bigger than smallBound, repeated, so the whole file is many times the bound.
String line = usageLine(42, 0, 0) + " ".repeat(500);
StringBuilder sb = new StringBuilder();
for (int i = 0; i < 50; i++) {
sb.append(line).append('\n');
}
Files.writeString(file, sb.toString(), StandardCharsets.UTF_8);
long fileLength = Files.size(file);
assertTrue(fileLength > (long) smallBound * 5, "test setup: fixture must genuinely dwarf the bound");
var tail = LeadContextGauge.tailBytes(file, smallBound);
assertEquals(smallBound, tail.bytes().length,
"must read exactly the bound, not the whole " + fileLength + "-byte file — assert on bytes actually read");
}
// --- property 5: the cache TTL means a burst of calls reads the file once -------------------
@Test
@DisplayName("the same (configDir, sessionId) read twice inside the TTL touches disk once")
void secondReadInsideTtlDoesNotTouchDiskAgain(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
AtomicLong now = new AtomicLong(0);
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
gauge.read(configDir, SESSION_ID, "claude");
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
now.set(6_000); // past the TTL
gauge.read(configDir, SESSION_ID, "claude");
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
}
// --- non-Claude peer / unresolved session -----------------------------------------------------
@Test
@DisplayName("a non-claude agent type or a null session id is UNKNOWN, never OK")
void nonClaudePeerOrUnresolvedSessionIsUnknown(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
}
}
@@ -32,13 +32,26 @@ class LeadLauncherTest {
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
}
private static FleetConfig.Profile profileWithAutoCompactWindow(String kind) {
return new FleetConfig.Profile(
"opus", null, "claude-opus-5", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms"), "tab", "fleet", null,
"http://127.0.0.1:8765/mcp", null, null,
null, null, kind, Map.of(), null, null, true, null, null,
null, null, null, 250_000);
}
private static FleetConfig configWith(FleetConfig.Leader lead) {
return configWith(lead, opusProfile());
}
private static FleetConfig configWith(FleetConfig.Leader lead, FleetConfig.Profile profile) {
Map<String, FleetConfig.Leader> leaders = new LinkedHashMap<>();
leaders.put("opus", lead);
FleetConfig.Fleet fleet =
new FleetConfig.Fleet(leaders, Map.of(), Map.of(), Map.of(), null);
return new FleetConfig(
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
null, null, Map.of("opus", profile), null, null, null, null, null,
null, null, fleet, null, "fixed", null).withDefaults();
}
@@ -330,6 +343,36 @@ class LeadLauncherTest {
"--model is appended last so it outranks the ccs wrapper (CB-533)");
}
@Test
void aClaudeLeadPassesItsConfiguredAutoCompactWindowToClaudeCode() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile profile = profileWithAutoCompactWindow("claude-code");
launcher(herdr, configWith(lead("opus", "lead: opus", 1), profile)).ensureLeads();
List<String> args = startedArgs(herdr);
assertEquals("250000", args.get(args.indexOf("--autocompact") + 1));
}
@Test
void aClaudeLeadWithNoAutoCompactWindowGetsNoAutoCompactFlag() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
assertFalse(startedArgs(herdr).contains("--autocompact"));
}
@Test
void aNonClaudeLeadDoesNotGetAnAutoCompactFlag() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile profile = profileWithAutoCompactWindow("opencode");
launcher(herdr, configWith(lead("opus", "lead: opus", 1), profile)).ensureLeads();
assertFalse(startedArgs(herdr).contains("--autocompact"));
}
/** A lead runs on the operator's subscription. Nothing may move it off. */
@Test
void theLeadEnvCarriesNoAnthropicBinding() {
@@ -1359,4 +1359,92 @@ class LeadRolloverTest {
+ "it were the measured wait duration: " + message);
}
}
// ---- fleetd #615: a HerdrException out of either unwrapped agents.send call must leave a ----
// ---- TERMINAL FAILED outcome, never a stuck IN_PROGRESS ---------------------------------------
@Test
@DisplayName("[fleetd #615 — 1] send() throwing on the /clear call leaves status(token) "
+ "reporting FAILED, not stuck at IN_PROGRESS")
void sendThrowingOnClearLeavesStatusReportingFailed() throws IOException {
FakeHerdr fake = new FakeHerdr(); // default idle — the turn-settle wait passes immediately
HerdrClient throwsOnClear = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) throws HerdrException {
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
throw new HerdrException("simulated herdr transport failure sending /clear");
}
return fake.call(method, params);
}
@Override
public void close() {
fake.close();
}
};
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(throwsOnClear, cfg(handover.toString()), fixedClock(clock));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
+ "inside the deferred continuation, which this test's synchronous runner has "
+ "already run to completion by the time confirm() returns");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.FAILED, status.state(),
"a HerdrException out of the /clear send must leave a TERMINAL FAILED outcome — "
+ "before fleetd #615's fix, the continuation thread died silently and "
+ "status() was stuck reporting the IN_PROGRESS confirm() wrote at hand-off, "
+ "forever: got " + status.state() + " / " + status.detail());
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
+ "so an operator reading status() has something to act on: " + status.detail());
}
@Test
@DisplayName("[fleetd #615 — 2] send() throwing on the bootstrap-text call (after /clear "
+ "succeeded and the pane settled) also leaves status(token) reporting FAILED — a "
+ "DIFFERENT exit from the /clear-throw case above")
void sendThrowingOnBootstrapTextLeavesStatusReportingFailed() throws IOException {
FakeHerdr fake = new FakeHerdr(); // default idle throughout — both settle waits pass promptly
HerdrClient throwsOnBootstrapText = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) throws HerdrException {
if ("agent.prompt".equals(method) && String.valueOf(params).contains("read the handover file")) {
throw new HerdrException("simulated herdr transport failure sending bootstrapText");
}
return fake.call(method, params);
}
@Override
public void close() {
fake.close();
}
};
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(throwsOnBootstrapText, cfg(handover.toString()), fixedClock(clock));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
+ "inside the deferred continuation, which this test's synchronous runner has "
+ "already run to completion by the time confirm() returns");
assertEquals(1, promptCallCount(fake), "sanity: /clear was sent and settled — only the "
+ "SECOND agent.prompt call (bootstrapText) threw");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.FAILED, status.state(),
"a HerdrException out of the bootstrapText send — a DIFFERENT exit from the /clear "
+ "throw, reached only after /clear already succeeded and the pane already "
+ "settled — must also leave a TERMINAL FAILED outcome, not a stuck "
+ "IN_PROGRESS: got " + status.state() + " / " + status.detail());
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
+ "so an operator reading status() has something to act on: " + status.detail());
}
}
@@ -0,0 +1,126 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.session.SessionManager;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #602 gauge-wiring: {@code FleetMcp}'s {@code fleet_list} handler shipped calling
* {@code LeadContextGauge.read(null, ...)} unconditionally, so every lead whose profile sets a
* {@code configDir:} override read the wrong transcript directory forever, with no error anywhere —
* {@link LeadContextGauge}'s own 8 tests all passed because none of them exercised THIS wiring; each
* one hands the gauge its own directory directly.
*
* <p>This class proves the directory {@code fleet.leaders.<name>.profile}'s own {@code configDir:}
* names is the one {@code fleet_list}'s {@code context} row actually reads — not merely that SOME
* directory got passed. If the wiring in {@code FleetMcp.leadView}/{@code contextView} is ever
* reverted to a hardcoded {@code null}, {@link #configuredDirectoryDecidesWhichTranscriptIsRead()}
* must go red: both directories below hold a REAL, DIFFERENT token count for the SAME session id,
* so a hardcoded {@code null} (always reading the built-in default, where neither transcript lives)
* would report {@code UNKNOWN} both times instead of the two distinct numbers this test asserts.
*/
class FleetMcpLeadContextGaugeWiringTest {
private static final String LEAD_TERMINAL = "term_lead";
private static final String LEAD_NAME = "opus";
/** {@code FakeHerdr.withAgent} always projects {@code agent_session.value} as {@code sess-<terminalId>}. */
private static final String LEAD_SESSION_ID = "sess-" + LEAD_TERMINAL;
private static ClaudeCodeLauncher workerService(FakeHerdr h) {
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
/** One line Claude Code would write for a turn with the given live-context total. */
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
private static String writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
return configDir.toString();
}
@Test
@DisplayName("the CONFIGURED directory decides which transcript is read, and follows a config change")
void configuredDirectoryDecidesWhichTranscriptIsRead(@TempDir Path tmp) throws IOException {
Path dirA = Files.createDirectory(tmp.resolve("dir-a"));
Path dirB = Files.createDirectory(tmp.resolve("dir-b"));
writeTranscript(dirA, LEAD_SESSION_ID, 11_000);
writeTranscript(dirB, LEAD_SESSION_ID, 22_000);
FakeHerdr herdr = new FakeHerdr().withAgent("lead-opus", LEAD_TERMINAL, "wL:p1", "wL:t1");
SessionManager sessions = new SessionManager(workerService(herdr));
LeadContextGauge contextGauge = new LeadContextGauge();
AtomicReference<String> configuredDir = new AtomicReference<>(dirA.toString());
FleetMcp.LeadConfigDirSource source = new FleetMcp.LeadConfigDirSource(name ->
LEAD_NAME.equals(name) ? configuredDir.get() : null);
String firstRead = textOf(listFleet(herdr, sessions, contextGauge, source));
assertTrue(firstRead.contains("\"tokens\":11000"),
"config names directory A ⇒ fleet_list must report A's token count: " + firstRead);
// The config changes to name directory B instead -- the very next read must follow it.
configuredDir.set(dirB.toString());
String secondRead = textOf(listFleet(herdr, sessions, contextGauge, source));
assertTrue(secondRead.contains("\"tokens\":22000"),
"config now names directory B ⇒ fleet_list must report B's token count, not A's stale one: "
+ secondRead);
}
@Test
@DisplayName("a lead whose config names no directory still degrades to the built-in default, never throws")
void aLeadWhoseConfigNamesNoDirectoryStillFallsBackWithoutThrowing() {
FakeHerdr herdr = new FakeHerdr().withAgent("lead-opus", LEAD_TERMINAL, "wL:p1", "wL:t1");
SessionManager sessions = new SessionManager(workerService(herdr));
LeadContextGauge contextGauge = new LeadContextGauge();
McpSchema.CallToolResult res = listFleet(herdr, sessions, contextGauge, FleetMcp.LeadConfigDirSource.none());
assertNotEquals(Boolean.TRUE, res.isError(), "a lead with no configured configDir must degrade, never throw");
String out = textOf(res);
assertTrue(out.contains("\"state\":\"unknown\""), "no transcript under the built-in default ⇒ UNKNOWN: " + out);
assertFalse(out.contains("\"tokens\""), "UNKNOWN must never carry a stale/default token number: " + out);
}
private static McpSchema.CallToolResult listFleet(FakeHerdr herdr, SessionManager sessions,
LeadContextGauge contextGauge, FleetMcp.LeadConfigDirSource leadConfigDirs) {
return FleetMcp.listFleet(workerService(herdr), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs,
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, FleetMcp.CoordinationSource.none(), false);
}
}
@@ -1,27 +1,46 @@
package dev.ltms.fleet.msg;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.MemberSession;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for the CB-551 idle-lead heartbeat: the pure {@link LeadHeartbeatLoop#decide} decision
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot.
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot. Also covers
* fleetd #609's context-high notice: the {@code context}/{@code contextNotified} parameters {@code
* decide} gained, and the standalone {@link LeadHeartbeatLoop#contextNotice} text builder.
*
* <p>All decision tests call {@code decide} directly with explicit nanoTime values from an injected
* clock — no sleeping, no scheduler races. This mirrors how {@code ReplyPushLoopTest} pins the pure
* decision before exercising the loop.
*
* <p>Every pre-#609 test below passes {@code LeadContextGauge.State.UNKNOWN, false} for the two new
* {@code decide} parameters — the same as a lead whose context could not be read and has never been
* notified — so each one still pins exactly the behaviour it pinned before this ticket.
*/
class LeadHeartbeatLoopTest {
@@ -56,12 +75,18 @@ class LeadHeartbeatLoopTest {
return new LeadHeartbeatLoop.FleetState(2, 1, 3, List.of("term_a", "term_b"));
}
/** The loop under test; the scheduler is never invoked on the pure decide path. */
/** The loop under test, context notice off; the scheduler is never invoked on the pure decide path. */
private static LeadHeartbeatLoop loop(int quietCap) {
return loop(quietCap, false);
}
/** As above, with the fleetd #609 {@code contextHighNudge} flag set explicitly. */
private static LeadHeartbeatLoop loop(int quietCap, boolean contextHighNudge) {
return new LeadHeartbeatLoop(
new PrimaryRegistry("term_lead"), null /*agents — unused on the decide path*/,
null /*inbox*/, List::of, null /*pushLoop*/, null /*scheduler*/, () -> 0L,
IDLE_AFTER_NANOS, 1_000L, quietCap);
IDLE_AFTER_NANOS, 1_000L, quietCap, null,
LeadHeartbeatLoop.LeadContextSource.none(), contextHighNudge);
}
// ── (a) a working lead is never injected ───────────────────────────────────────────────────
@@ -69,7 +94,8 @@ class LeadHeartbeatLoopTest {
@Test
void aWorkingLeadIsNeverInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet());
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
"a WORKING lead is making progress and must not be touched");
assertNull(d.idleSinceNanos(), "a busy lead resets the idle window");
@@ -80,17 +106,71 @@ class LeadHeartbeatLoopTest {
void anUnreadableStatusIsNeverInjectedEither() {
// A failed status read (or a gone agent) must degrade to "do not inject", never hammer the pane.
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet());
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
"never inject into a state the loop cannot read");
}
// ── fleetd #613: the boot line names all four heartbeat settings ──────────────────────────
/**
* fleetd #613: {@link LeadHeartbeatLoop#start()}'s boot line named only 3 of the 4 constructor
* settings — {@code contextHighNudge} (fleetd #609) was missing, so an operator could not tell
* from the log whether the handover notice was armed. Captures the real log via a
* {@link ListAppender}, raising the logger's level past {@code logback-test.xml}'s
* {@code dev.ltms.fleet -> WARN} override for the duration of the call — the same seam {@code
* GitHostShapeReportTest#attach} uses for its own INFO-level boot line.
*/
private static String heartbeatBootLine(boolean contextHighNudge, ScheduledExecutorService scheduler) {
Logger logger = (Logger) LoggerFactory.getLogger(LeadHeartbeatLoop.class);
Level originalLevel = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
LeadHeartbeatLoop l = new LeadHeartbeatLoop(
new PrimaryRegistry("term_lead"), null /*agents*/, null /*inbox*/, List::of,
null /*pushLoop*/, scheduler, () -> 0L, IDLE_AFTER_NANOS, 1_000L, 3, null,
LeadHeartbeatLoop.LeadContextSource.none(), contextHighNudge);
l.start();
} finally {
logger.detachAppender(appender);
logger.setLevel(originalLevel);
}
return appender.list.stream()
.filter(e -> e.getLevel() == Level.INFO)
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.startsWith("idle-lead heartbeat:"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected the heartbeat boot line to be logged"));
}
@Test
void theBootLineNamesContextHighNudgeWhenArmed() {
String line = heartbeatBootLine(true, scheduler);
assertTrue(line.contains("300s idle"), line);
assertTrue(line.contains("recheck 1000ms"), line);
assertTrue(line.contains("quiet cap 3"), line);
assertTrue(line.contains("context-high nudge true"),
"the boot line must name the 4th setting, contextHighNudge, alongside the other "
+ "three: " + line);
}
@Test
void theBootLineNamesContextHighNudgeWhenOff() {
String line = heartbeatBootLine(false, scheduler);
assertTrue(line.contains("context-high nudge false"), line);
}
// ── (b) an idle lead within the quiet period is not yet injected ───────────────────────────
@Test
void justBecameIdleStartsTheDebounceWindow() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet());
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"the first injectable tick only records the start of the idle stretch");
assertEquals(NOW, d.idleSinceNanos(), "the idle window opens at the moment the lead became injectable");
@@ -99,7 +179,8 @@ class LeadHeartbeatLoopTest {
@Test
void idleWithinQuietPeriodIsNotInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet());
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"a lead idle for 10s (< 300s) has just finished a turn — do not re-prompt it");
}
@@ -109,7 +190,8 @@ class LeadHeartbeatLoopTest {
@Test
void idlePastQuietPeriodIsInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet());
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
"a lead continuously idle past the quiet period is the reason to nudge");
}
@@ -117,14 +199,17 @@ class LeadHeartbeatLoopTest {
@Test
void blockedAndDoneAreInjectableViewsOfIdle() {
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet()).action());
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false).action());
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet()).action());
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false).action());
}
@Test
void nothingIsInjectedWhenNoLeadIsKnown() {
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet());
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"with no known lead there is nobody to nudge — keep waiting until one is discovered");
assertEquals(IDLE_PAST, d.idleSinceNanos(), "the idle window stays open so discovery re-arms it");
@@ -135,7 +220,8 @@ class LeadHeartbeatLoopTest {
@Test
void quietNudgeCapStopsTheLoopWhenNothingIsPending() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet());
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
"3 consecutive nothing-pending nudges have already happened — stop nagging an empty fleet");
}
@@ -145,7 +231,8 @@ class LeadHeartbeatLoopTest {
@Test
void newPendingStateResetsTheQuietCap() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet());
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
"real state appearing re-arms the loop past an exhausted cap");
assertEquals(0, d.quietCount(), "the pending state resets the consecutive-quiet counter");
@@ -156,7 +243,8 @@ class LeadHeartbeatLoopTest {
@Test
void standsDownWhileReplyPushLoopIsActive() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet());
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(),
"a second injection would start a competing turn — stand aside instead");
assertEquals(0, d.quietCount(), "the active push is real state, so it re-arms the cap");
@@ -213,4 +301,398 @@ class LeadHeartbeatLoopTest {
assertFalse(fs.hasPending());
assertEquals(0, fs.liveWorkers());
}
// ── fleetd #609: context-high notice ───────────────────────────────────────────────────────
/** One gate scenario, keyed by name, over which the off-state/context-independence property is checked. */
private record Gate(String name, Long idleSince, int quietCount, AgentStatus status,
boolean pushLoopActive, boolean leadKnown, LeadHeartbeatLoop.FleetState fleet) {}
private static List<Gate> allEightGates() {
return List.of(
new Gate("1-standDown", IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet()),
new Gate("2-leadBusy", IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet()),
new Gate("3-justBecameIdle", null, 0, AgentStatus.IDLE, false, true, quietFleet()),
new Gate("4-withinQuietPeriod", IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet()),
new Gate("5-noLeadKnown", IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet()),
new Gate("6-pending", IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet()),
new Gate("7-quietNotExhausted", IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet()),
new Gate("8-quietExhausted", IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet()));
}
@Test
void aOffFlagIgnoresContextAcrossAllEightGates() {
// Property A: with contextHighNudge off, decide() must not depend on the context reading at
// all — not just "usually agrees", but identical Action/idleSinceNanos/quietCount whatever
// context state is passed, and contextNotified must stay false throughout. This is the proof
// that fleetd #609 is opt-in: a daemon upgraded to carry this code, but never configuring
// `contextHighNudge: true`, behaves exactly as it did before this ticket for every one of the
// 8 gates the class javadoc numbers.
LeadHeartbeatLoop offLoop = loop(3, false);
for (Gate g : allEightGates()) {
var withHigh = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.HIGH, false);
var withUnknown = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.UNKNOWN, false);
var withOk = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.OK, false);
assertEquals(withUnknown.action(), withHigh.action(), g.name() + ": action must not depend on context");
assertEquals(withUnknown.idleSinceNanos(), withHigh.idleSinceNanos(), g.name());
assertEquals(withUnknown.quietCount(), withHigh.quietCount(), g.name());
assertEquals(withUnknown.action(), withOk.action(), g.name() + ": nor on an OK reading");
assertFalse(withHigh.contextNotified(), g.name() + ": contextNotified must stay false when the flag is off");
}
}
@Test
void bHighContextFiresEvenPastAnExhaustedQuietCapWithoutSpendingIt() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.HIGH, false);
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
"an idle, quiet, HIGH-context lead is exactly the case the exhausted cap must not swallow");
assertEquals(3, d.quietCount(), "the context notice is an event notice, not a quiet nudge — it must not "
+ "spend or grow the consecutive-quiet budget");
assertTrue(d.contextNotified(), "firing the notice sets the latch");
}
@Test
void cASecondTickWithTheLatchAlreadySetDoesNotTellTheLeadAgain() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.HIGH, true);
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
"the lead was already told once about this HIGH stretch — telling it every tick would be nagging, "
+ "not a notice");
}
@Test
void dFlappingBetweenHighAndUnknownNeverReInjectsWhileLatched() {
LeadHeartbeatLoop on = loop(3, true);
// The lead was already told once (latch = true from a prior HIGH tick).
LeadHeartbeatLoop.Decision afterUnknown = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, true);
assertTrue(afterUnknown.contextNotified(),
"UNKNOWN means 'I could not look', not 'it got better' — it must not clear the latch");
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, afterUnknown.action(),
"a still-latched, still-quiet tick must not inject just because the reading is UNKNOWN");
LeadHeartbeatLoop.Decision afterHighAgain = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.HIGH, afterUnknown.contextNotified());
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, afterHighAgain.action(),
"HIGH -> UNKNOWN -> HIGH with the latch already set must never inject again — this is the "
+ "flapping case that would otherwise cost a full lead its remaining turns");
}
@Test
void eAnOkReadingClearsTheLatchSoALaterHighInjectsAgain() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision afterOk = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.OK, true);
assertFalse(afterOk.contextNotified(), "an OK reading clears the latch — the lead's context recovered");
LeadHeartbeatLoop.Decision afterHighAgain = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.HIGH, afterOk.contextNotified());
assertEquals(LeadHeartbeatLoop.Action.INJECT, afterHighAgain.action(),
"the cleared latch lets a genuinely new HIGH stretch notify again");
}
@Test
void fStandingDownDoesNotBurnTheOneContextNotice() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet(),
LeadContextGauge.State.HIGH, false);
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(), "ReplyPushLoop is active — stand aside");
assertFalse(d.contextNotified(),
"standing down must not spend the one notice this HIGH stretch gets — the latch stays clear so a "
+ "later tick can still fire it");
}
@Test
void gAWorkingLeadWithHighContextIsStillLeadBusy() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet(),
LeadContextGauge.State.HIGH, false);
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(), "the status gate wins over the context notice");
}
@Test
void hAPendingDrivenInjectWhileHighSetsTheLatchTooSoTheLeadIsNotToldTwice() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet(),
LeadContextGauge.State.HIGH, false);
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action());
assertTrue(d.contextNotified(),
"the pending-driven nudge carries the context notice too (see contextNotice), so it must set the "
+ "latch — otherwise the lead could be told twice by two different routes");
}
// ── contextNotice text builder ─────────────────────────────────────────────────────────────
@Test
void contextNoticeIsEmptyWhenDisabled() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
assertEquals("", LeadHeartbeatLoop.contextNotice(false, reading));
}
@Test
void contextNoticeIsEmptyWhenStateIsOk() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.OK, 1_000L, 0);
assertEquals("", LeadHeartbeatLoop.contextNotice(true, reading));
}
@Test
void contextNoticeIsEmptyWhenStateIsUnknown() {
assertEquals("", LeadHeartbeatLoop.contextNotice(true, LeadContextGauge.Reading.unknown()));
}
@Test
void contextNoticeNamesTheTokenCountWhenHighAndEnabled() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
String notice = LeadHeartbeatLoop.contextNotice(true, reading);
assertFalse(notice.isEmpty());
assertTrue(notice.contains("260771"), notice);
assertTrue(notice.contains("1 compaction"), notice);
assertTrue(notice.contains("fleet_handover"), notice);
assertTrue(notice.contains("operator"), notice);
}
@Test
void contextNoticeOmitsTheTokenClauseRatherThanPrintingNull() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, null, 2);
String notice = LeadHeartbeatLoop.contextNotice(true, reading);
assertFalse(notice.isEmpty());
assertFalse(notice.toLowerCase().contains("null"), notice);
assertTrue(notice.contains("2 compactions"), notice);
}
// ── fleetd #621: the notice must track the effective requireOperatorConfirm value ─────────────
@Test
void contextNoticeKeepsAskingTheOperatorWhenRequireOperatorConfirmIsTrue() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
String notice = LeadHeartbeatLoop.contextNotice(true, reading, false, true);
assertTrue(notice.contains("ask the operator"), notice);
assertTrue(notice.contains("Only the operator can approve the roll"), notice);
}
@Test
void contextNoticeDropsTheOperatorAskWhenRequireOperatorConfirmIsFalse() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
String notice = LeadHeartbeatLoop.contextNotice(true, reading, false, false);
assertFalse(notice.contains("ask the operator"), notice);
assertFalse(notice.contains("Only the operator can approve the roll"), notice);
assertTrue(notice.contains("fleet_handover"), notice);
assertTrue(notice.contains("maxDocAgeSeconds"), notice);
}
// ── fleetd #609 review: the latch must mean "the notice reached the pane" ────────────────────
//
// These four drive LeadHeartbeatLoop.tick() directly (package-private, same reasoning as
// ReplyPushLoop#tick(String) being directly testable) against a real AgentControl wrapping a
// FailableHerdrClient, so the send path (agents.send -> herdr -> possible throw) is exercised
// for real rather than assumed from decide()'s Decision alone.
private static final String LEAD = "term_lead";
private static final String WORKER = "term_w1";
/** A HIGH reading with a fixed token/compaction count, for the four tests below. */
private static LeadContextGauge.Reading highReading() {
return new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
}
/**
* Builds a real {@link LeadHeartbeatLoop} wired to {@code herdr} via a real {@link AgentControl},
* a mutable fake clock, and a mutable roster so a test can change fleet state between ticks. The
* lead is always reported IDLE by {@code herdr}, so every tick's outcome is governed only by the
* idle-window/quiet-cap/context gates under test.
*/
private static LeadHeartbeatLoop tickableLoop(FailableHerdrClient herdr, AtomicLong now,
List<MemberSession>[] rosterBox, InMemoryReplyInbox inbox,
int quietNudgeCap, ScheduledExecutorService scheduler) {
AgentControl agents = new AgentControl(herdr);
PrimaryRegistry registry = new PrimaryRegistry(LEAD);
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, agents, inbox, scheduler, 5, 100_000);
return new LeadHeartbeatLoop(registry, agents, inbox, () -> rosterBox[0], pushLoop, scheduler,
now::get, IDLE_AFTER_NANOS, 100_000L, quietNudgeCap, null,
new LeadHeartbeatLoop.LeadContextSource(t -> highReading()), true);
}
@Test
void iAFailedSendDoesNotConsumeTheNotice() {
var herdr = new FailableHerdrClient(LEAD);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()}; // quiet: nothing pending
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
loop.tick(); // first injectable tick: only opens the idle window (WAIT_IDLE)
now.addAndGet(TimeUnit.SECONDS.toNanos(400)); // now clearly past the quiet period
herdr.throwOnNextSend();
loop.tick(); // quiet fleet, quiet cap exhausted (0), context HIGH, latch clear -> INJECT, send throws
assertEquals(0, herdr.sentTexts().size(), "the failed send must not have recorded any text");
loop.tick(); // same inputs — the latch must still be clear, so this must INJECT and send again
assertEquals(1, herdr.sentTexts().size(),
"a retried tick with the latch still clear must attempt the send again");
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"),
"the retried, successful send must carry the notice: " + herdr.sentTexts().get(0));
}
@Test
void jASuccessfulSendDoesConsumeIt() {
var herdr = new FailableHerdrClient(LEAD);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
loop.tick(); // opens the idle window
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
loop.tick(); // INJECT, send succeeds -> latch set
assertEquals(1, herdr.sentTexts().size());
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"), herdr.sentTexts().get(0));
loop.tick(); // same inputs — the lead was already told this stretch
assertEquals(1, herdr.sentTexts().size(),
"the next tick with the same inputs must not send a second notice");
}
@Test
void kAPendingDrivenInjectWithTheLatchAlreadySetSendsNoNotice() {
var herdr = new FailableHerdrClient(LEAD);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
loop.tick();
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
loop.tick(); // latches the notice (quiet, HIGH, cap exhausted -> the forced context INJECT)
assertEquals(1, herdr.sentTexts().size());
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"));
// Now make the fleet have real pending state, so the NEXT INJECT is pending-driven, not the
// forced-context route — with the latch already set from the tick above.
inbox.own(WORKER);
inbox.publish(WORKER, "m1", "hello");
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null));
loop.tick();
assertEquals(2, herdr.sentTexts().size(), "the pending-driven tick must still send a nudge");
assertFalse(herdr.sentTexts().get(1).contains("Your own context is nearly full"),
"a pending-driven INJECT while the latch is already set must carry no context notice: "
+ herdr.sentTexts().get(1));
}
@Test
void lTheNoticeAppearsExactlyOnceAcrossThreeDifferentlyDrivenInjects() {
var herdr = new FailableHerdrClient(LEAD);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
// quietNudgeCap=1 so a still-not-exhausted quiet nudge is available as the "forced" route below,
// distinct from both the pending-driven route and the exhausted-cap forced-context route.
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 1, scheduler);
loop.tick(); // opens the idle window
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
// 1) pending-driven INJECT: real fleet state present. Sets the latch and carries the notice.
inbox.own(WORKER);
inbox.publish(WORKER, "m1", "hello");
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null));
loop.tick();
assertEquals(1, herdr.sentTexts().size());
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"), herdr.sentTexts().get(0));
// 2) "forced" INJECT: nothing pending, but the quiet cap (1) is not yet exhausted, so decide()
// nudges anyway. The latch is already set, so no notice.
inbox.ack(WORKER, "m1");
rosterBox[0] = List.of();
loop.tick();
assertEquals(2, herdr.sentTexts().size(), "the quiet-cap-not-yet-exhausted nudge must still fire");
assertFalse(herdr.sentTexts().get(1).contains("Your own context is nearly full"), herdr.sentTexts().get(1));
// 3) pending-driven INJECT again. Still latched, still no notice.
inbox.publish(WORKER, "m2", "hello again");
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null));
loop.tick();
assertEquals(3, herdr.sentTexts().size());
assertFalse(herdr.sentTexts().get(2).contains("Your own context is nearly full"), herdr.sentTexts().get(2));
long noticeCount = herdr.sentTexts().stream()
.filter(t -> t.contains("Your own context is nearly full")).count();
assertEquals(1, noticeCount,
"the notice text must appear exactly once across all three sends: " + herdr.sentTexts());
}
/**
* Fake herdr client for the four tests above: always reports {@code lead} as IDLE, records the
* {@code text} of every {@code agent.prompt} call, and can be told to throw on the very next
* {@code agent.prompt} call — standing in for one transient herdr send failure.
*/
private static final class FailableHerdrClient implements HerdrClient {
private static final ObjectMapper MAPPER = new ObjectMapper();
private final String lead;
private final List<String> sentTexts = new ArrayList<>();
private boolean throwOnNextSend = false;
FailableHerdrClient(String lead) {
this.lead = lead;
}
void throwOnNextSend() {
throwOnNextSend = true;
}
List<String> sentTexts() {
return List.copyOf(sentTexts);
}
@Override
@SuppressWarnings("unchecked")
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", lead)
.put("agent_status", "idle"));
}
if ("agent.prompt".equals(method)) {
if (throwOnNextSend) {
throwOnNextSend = false;
throw new RuntimeException("simulated transient herdr send failure");
}
Map<String, Object> p = params instanceof Map ? (Map<String, Object>) params : Map.of();
sentTexts.add(String.valueOf(p.get("text")));
}
return MAPPER.createObjectNode();
}
@Override
public void close() {
}
}
}
@@ -0,0 +1,168 @@
package dev.ltms.fleet.msg;
import java.util.ArrayDeque;
import java.util.ArrayList;
import java.util.Collection;
import java.util.Deque;
import java.util.List;
import java.util.concurrent.Callable;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.Future;
import java.util.concurrent.RejectedExecutionException;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.ScheduledFuture;
import java.util.concurrent.TimeUnit;
/**
* A {@link ScheduledExecutorService} for tests that never runs a task on a timer — it records what
* {@link ReplyPushLoop} schedules and only ever runs it when the test itself calls
* {@link #runDueTasks()} (fleetd #608).
*
* <p>The test this exists for ({@code anAlreadyCollectedTicketProducesNoNudge}) used to wire
* {@link ReplyPushLoop} to a real {@code Executors.newSingleThreadScheduledExecutor()} and bet a
* 300ms backoff was "wide" enough that the test's own work (collecting the ticket) always won the
* race against the scheduler's own timer firing the next tick. On an idle machine that held; under
* a loaded full-suite run it did not, and the test went red on perfectly correct code. Replacing
* the timer with this fake removes the race rather than widening it: nothing here ever runs on a
* schedule of its own, so no backoff value — however small — can make the tick fire before the test
* is ready for it.
*
* <p><strong>Deliberately narrow.</strong> {@link ReplyPushLoop} calls exactly two methods on its
* {@link ScheduledExecutorService} field — {@link #schedule(Runnable, long, TimeUnit)} (every tick,
* including the very first one {@code startOrCoalesce} kicks off) and {@link #shutdownNow()} (on
* {@code ReplyPushLoop.stop()}/{@code close()}) — verified by reading every {@code scheduler.} call
* site in that class. Every other {@link ScheduledExecutorService} method throws
* {@link UnsupportedOperationException} rather than silently doing the wrong thing, so a future
* change to {@code ReplyPushLoop} that starts calling one of them fails this fake loudly, on the
* very first test that exercises it, instead of being quietly mishandled.
*/
final class ManualScheduler implements ScheduledExecutorService {
private final Deque<Runnable> pending = new ArrayDeque<>();
private boolean shutdown = false;
/**
* Run every task pending as of the START of this call — a snapshot taken before anything runs.
* {@link ReplyPushLoop#tick} always reschedules its own next tick before returning (see
* {@code scheduleNext} in its {@code INJECT}/{@code WAIT_BUSY} branches), so without the
* snapshot a single call here would recurse forever. Taking it up front means one call is
* exactly one tick, deterministically, no matter what that tick itself goes on to schedule.
*
* @return how many tasks actually ran
*/
synchronized int runDueTasks() {
List<Runnable> due = new ArrayList<>(pending);
pending.clear();
for (Runnable task : due) {
task.run();
}
return due.size();
}
/** How many tasks are currently queued, without running any of them. */
synchronized int pendingCount() {
return pending.size();
}
@Override
public synchronized ScheduledFuture<?> schedule(Runnable command, long delay, TimeUnit unit) {
if (shutdown) {
throw new RejectedExecutionException("ManualScheduler is shut down");
}
pending.add(command);
return null; // ReplyPushLoop discards the return value of every schedule() call it makes.
}
@Override
public synchronized List<Runnable> shutdownNow() {
shutdown = true;
List<Runnable> left = new ArrayList<>(pending);
pending.clear();
return left;
}
@Override
public synchronized boolean isShutdown() {
return shutdown;
}
// --- everything below: not called by ReplyPushLoop, and not supported by this fake (see the
// class javadoc's "deliberately narrow" note) -------------------------------------------------
@Override
public ScheduledFuture<?> scheduleAtFixedRate(Runnable command, long initialDelay, long period, TimeUnit unit) {
throw unsupported();
}
@Override
public ScheduledFuture<?> scheduleWithFixedDelay(Runnable command, long initialDelay, long delay, TimeUnit unit) {
throw unsupported();
}
@Override
public <V> ScheduledFuture<V> schedule(Callable<V> callable, long delay, TimeUnit unit) {
throw unsupported();
}
@Override
public void shutdown() {
throw unsupported();
}
@Override
public boolean isTerminated() {
throw unsupported();
}
@Override
public boolean awaitTermination(long timeout, TimeUnit unit) {
throw unsupported();
}
@Override
public <T> Future<T> submit(Callable<T> task) {
throw unsupported();
}
@Override
public <T> Future<T> submit(Runnable task, T result) {
throw unsupported();
}
@Override
public Future<?> submit(Runnable task) {
throw unsupported();
}
@Override
public <T> List<Future<T>> invokeAll(Collection<? extends Callable<T>> tasks) {
throw unsupported();
}
@Override
public <T> List<Future<T>> invokeAll(Collection<? extends Callable<T>> tasks, long timeout, TimeUnit unit) {
throw unsupported();
}
@Override
public <T> T invokeAny(Collection<? extends Callable<T>> tasks) throws ExecutionException {
throw unsupported();
}
@Override
public <T> T invokeAny(Collection<? extends Callable<T>> tasks, long timeout, TimeUnit unit)
throws ExecutionException {
throw unsupported();
}
@Override
public void execute(Runnable command) {
throw unsupported();
}
private static UnsupportedOperationException unsupported() {
return new UnsupportedOperationException(
"ManualScheduler only supports schedule(Runnable, long, TimeUnit), shutdownNow() and isShutdown() "
+ "— see the class javadoc");
}
}
@@ -1841,6 +1841,43 @@ class MessageServiceTest {
return new PushWiring(service, leadHerdr, scheduler);
}
/**
* A {@link MessageService} wired to a real {@link ReplyPushLoop}, like {@link #wireWithPushLoop},
* but backed by {@link ManualScheduler} instead of a real timer (fleetd #608): its tick never
* fires on its own — a test drives it explicitly via {@link ManualScheduler#runDueTasks()}.
*/
private record ManualPushWiring(MessageService service, FakeHerdr leadHerdr, ManualScheduler scheduler,
ReplyPushLoop pushLoop)
implements AutoCloseable {
@Override
public void close() {
scheduler.shutdownNow();
}
}
/**
* As {@link #wireWithPushLoop}, but the schedule never runs on its own. {@code backoffMs} is
* still threaded through to {@link ReplyPushLoop}'s constructor (it takes one), but nothing here
* ever waits it out, so its value cannot affect anything a test built on this observes — see
* {@code anAlreadyCollectedTicketProducesNoNudge}, which sets it to 1 to prove exactly that.
*/
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs) {
return wireWithManualScheduler(maxReminders, backoffMs, System::nanoTime);
}
/** As above, with an injectable clock for tests that exercise terminal-ticket pruning. */
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs,
java.util.function.LongSupplier nowNanos) {
PrimaryRegistry registry = new PrimaryRegistry(null);
registry.recordDelegation(T, LEAD);
FakeHerdr leadHerdr = new FakeHerdr();
AgentControl leadAgents = new AgentControl(leadHerdr);
ManualScheduler scheduler = new ManualScheduler();
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, nowNanos);
return new ManualPushWiring(service, leadHerdr, scheduler, pushLoop);
}
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
long deadline = System.currentTimeMillis() + 3000;
while (!leadHerdr.called("agent.prompt") && System.currentTimeMillis() < deadline) {
@@ -1882,16 +1919,46 @@ class MessageServiceTest {
@Test
void anAlreadyCollectedTicketProducesNoNudge() throws Exception {
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires
String ticket = wiring.service().sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
// fleetd #608: backoff is 1ms — the most hostile value there is, the tick due immediately —
// and the test still must pass, because with ManualScheduler the tick never runs on a timer
// at all; it only runs when this test calls runDueTasks() below. The old version bet a 300ms
// backoff was "wide" enough that collecting the ticket always won the race against a real
// scheduler's own timer; that held on an idle machine and failed under a loaded full-suite
// run — a false red on correct code, since nothing here forced that ordering, it only made
// it likely.
try (var wiring = wireWithManualScheduler(1, 1)) {
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
// finishAsyncTask's task.future.complete(...) runs every whenComplete registered on that
// same future — including sendAsync's own hook that calls ReplyPushLoop.onTicketTerminal
// — synchronously, before complete() returns (see finishAsyncTask's javadoc and fleetd
// #399). This test-only hook fires right after that complete() call, on the very same
// (async executor) thread, so waiting for it guarantees onTicketTerminal has already run
// and the ticket is really sitting in the push loop's pending set — unlike waiting for
// Phase.DONE via poll(), which CompletableFuture.complete() can make visible to another
// thread before every whenComplete dependent has actually finished running (the same
// publish-then-run-dependents gap awaitCompletionStamped exists to close elsewhere).
String ticket;
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(terminalReached::countDown);
try {
ticket = wiring.service().sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
assertTrue(terminalReached.await(5, TimeUnit.SECONDS),
"the ticket never reached its terminal phase");
} finally {
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
}
// Collect it — the exact action the nudge must never follow.
MessageService.TaskView view = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.DONE);
assertEquals(MessageService.Phase.DONE, view.phase());
Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
// Now run the pending tick explicitly. onTicketTerminal scheduled exactly one (the ticket
// is the only thing this lead has ever had pending); it must find nothing left pending —
// the collection above already removed it — and send no nudge.
assertEquals(1, wiring.scheduler().runDueTasks(),
"expected exactly the one tick onTicketTerminal scheduled");
assertFalse(wiring.leadHerdr().called("agent.prompt"),
"a ticket the lead already polled must never be nudged");
}
@@ -1899,24 +1966,35 @@ class MessageServiceTest {
@Test
void severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: both tickets land before the tick fires
String first = wiring.service().sendAsync(T, "first task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "first done"));
// Settle without polling: poll() itself marks a ticket collected (that's the point of
// anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion
// would collect the ticket before the coalescing this test checks ever gets a chance.
Thread.sleep(100);
// A 1ms backoff is due immediately. ManualScheduler still cannot run it until this test
// explicitly calls runDueTasks(), so both terminal tickets join one scheduled tick.
try (var wiring = wireWithManualScheduler(1, 1)) {
String first;
String second;
java.util.concurrent.CountDownLatch firstTerminalReached = new java.util.concurrent.CountDownLatch(1);
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(firstTerminalReached::countDown);
try {
first = wiring.service().sendAsync(T, "first task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "first done"));
assertTrue(firstTerminalReached.await(5, TimeUnit.SECONDS),
"the first ticket never reached its terminal phase");
String second = wiring.service().sendAsync(T, "second task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "second done"));
Thread.sleep(100);
java.util.concurrent.CountDownLatch secondTerminalReached = new java.util.concurrent.CountDownLatch(1);
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(secondTerminalReached::countDown);
second = wiring.service().sendAsync(T, "second task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "second done"));
assertTrue(secondTerminalReached.await(5, TimeUnit.SECONDS),
"the second ticket never reached its terminal phase");
} finally {
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
}
awaitNudge(wiring.leadHerdr());
Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge
assertEquals(1, wiring.scheduler().runDueTasks(),
"both terminal tickets must coalesce onto one scheduled tick");
long nudgeCount = wiring.leadHerdr().calls.stream()
.filter(c -> c.method().equals("agent.prompt")).count();
assertEquals(1, nudgeCount, "two tickets finishing together must produce ONE nudge, not two");
@@ -1995,7 +2073,8 @@ class MessageServiceTest {
@Test
void answeringAQuestionStopsFurtherNudgesAboutIt() throws Exception {
try (var wiring = wireWithPushLoop(5, 50)) {
// A 1ms backoff is due immediately, but ManualScheduler only ticks when this test asks it to.
try (var wiring = wireWithManualScheduler(5, 1)) {
String ticket = wiring.service().sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
@@ -2003,8 +2082,9 @@ class MessageServiceTest {
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> wiring.service().ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
awaitQuestionPendingOn(wiring.pushLoop(), LEAD, asking.turnId());
awaitNudge(wiring.leadHerdr());
assertEquals(1, wiring.scheduler().runDueTasks(), "the open question must have one scheduled tick");
long callsBeforeAnswer = wiring.leadHerdr().calls.stream()
.filter(c -> c.method().equals("agent.prompt")).count();
@@ -2012,12 +2092,22 @@ class MessageServiceTest {
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(terminalReached::countDown);
assertTrue(rendezvous.resolve(T, "done"));
try {
assertTrue(terminalReached.await(5, TimeUnit.SECONDS),
"the answered ticket never reached its terminal phase");
} finally {
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
}
answer.get(5, TimeUnit.SECONDS);
// Let several more ticks (and the ticket's own now-legitimate terminal nudge) fire —
// none of them may still name the question's turnId, which is closed.
Thread.sleep(300);
// Run the ticket's legitimate terminal nudge and every later scheduled tick through its
// reminder cap. None may still name the closed question.
for (int tick = 0; tick < 6; tick++) {
assertEquals(1, wiring.scheduler().runDueTasks(), "expected one scheduled reminder tick");
}
boolean anyNamesClosedQuestion = wiring.leadHerdr().calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.skip(callsBeforeAnswer)
@@ -2100,35 +2190,45 @@ class MessageServiceTest {
// decideTickets hits the cap and STOPs — activeLeads drops the lead, but (before the fix)
// pendingTickets never drops the ticket. That is the exact "cap already STOPped" branch of
// the bug report, reached deterministically rather than by timing it against a live tick.
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
String stale = wiring.service().sendAsync(T, "first task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "stale result"));
// A 1ms backoff is due immediately, but ManualScheduler runs only the ticks below.
try (var wiring = wireWithManualScheduler(1, 1, clock::get)) {
String stale;
String fresh;
java.util.concurrent.CountDownLatch staleTerminalReached = new java.util.concurrent.CountDownLatch(1);
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(staleTerminalReached::countDown);
try {
stale = wiring.service().sendAsync(T, "first task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "stale result"));
assertTrue(staleTerminalReached.await(5, TimeUnit.SECONDS),
"the stale ticket never reached its terminal phase");
// Let the reminder loop fire its one nudge and hit the cap (STOP removes it from
// activeLeads; pendingTickets is untouched either way — that asymmetry is the bug).
awaitNudge(wiring.leadHerdr());
Thread.sleep(300);
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
"sanity: the stale ticket's own reminder must have fired first");
// Fire the stale ticket's one nudge, then its cap tick (STOP removes it from activeLeads;
// pendingTickets is untouched either way — that asymmetry is the bug).
assertEquals(1, wiring.scheduler().runDueTasks(), "the stale ticket must have one nudge tick");
assertEquals(1, wiring.scheduler().runDueTasks(), "the stale ticket must have one cap tick");
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
"sanity: the stale ticket's own reminder must have fired first");
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
String fresh = wiring.service().sendAsync(T, "second task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "fresh result"));
// The fresh ticket restarts the (now-dormant) reminder loop with its own nudge.
long before = wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count();
long deadline = System.currentTimeMillis() + 3000;
while (wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count() <= before
&& System.currentTimeMillis() < deadline) {
Thread.sleep(10);
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
java.util.concurrent.CountDownLatch freshTerminalReached = new java.util.concurrent.CountDownLatch(1);
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(freshTerminalReached::countDown);
fresh = wiring.service().sendAsync(T, "second task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "fresh result"));
assertTrue(freshTerminalReached.await(5, TimeUnit.SECONDS),
"the fresh ticket never reached its terminal phase");
} finally {
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
}
// The fresh ticket restarts the now-dormant reminder loop with its own nudge.
assertEquals(1, wiring.scheduler().runDueTasks(), "the fresh ticket must have one nudge tick");
String latestNudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(latestNudge.contains(fresh), "the fresh ticket's nudge must still arrive: " + latestNudge);
assertFalse(latestNudge.contains(stale),
@@ -328,9 +328,10 @@ class SessionManagerTest {
void rosterViewExposesTheCharterReceiptButNeverTheCharterText() {
// The roster (fleet_list and GET /members both render through rosterView) must let a lead
// see which charter a member got, without ever carrying the charter prose itself (CB-571).
String composed = "role charter\n\nreply";
CharterReceipt receipt = CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", composed);
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null,
CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", "role charter\n\nreply"), null);
0, 0, 0, MemberSession.State.READY, null, null, receipt, null);
Map<String, Object> view = SessionManager.rosterView(s, null);
@@ -338,10 +339,49 @@ class SessionManagerTest {
"the config key that supplied the role charter is reported");
assertEquals(CharterReceipt.digestOf("role charter\n\nreply"), view.get("charterSha256"),
"the digest of the exact composed charter bytes is reported");
// #604: the byte count rides alongside the digest, and must match what the receipt itself
// carries (not a hardcoded literal) so a bug that reads the wrong field is caught.
assertEquals(receipt.charterBytes(), view.get("charterBytes"),
"the exact composed byte count is reported, read from the receipt");
assertEquals(composed.getBytes(java.nio.charset.StandardCharsets.UTF_8).length, view.get("charterBytes"),
"the byte count is the real UTF-8 length of the composed charter");
assertFalse(view.values().toString().contains("role charter"),
"the roster row must not embed the charter text itself");
}
@Test
void rosterViewOmitsCharterBytesAndDigestWhenNoCharterWasComposed() {
// #604: a member with no role charter and no reply charter (composed == null) still reports
// charterSource ("none"), but neither a digest nor a size — the digest is absent, and the
// size only ever accompanies a real digest. This must not be satisfiable by code that always
// writes charterBytes.
CharterReceipt receipt = CharterReceipt.compose(MemberRole.DEV, "prof", null, null);
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null, receipt, null);
Map<String, Object> view = SessionManager.rosterView(s, null);
assertEquals(CharterReceipt.NO_SOURCE, view.get("charterSource"),
"no configured role or reply charter reports the explicit \"none\" source");
assertFalse(view.containsKey("charterSha256"), "no digest is reported when no charter was composed");
assertFalse(view.containsKey("charterBytes"), "no byte count is reported when no charter was composed");
}
@Test
void rosterViewOmitsCharterFieldsEntirelyWhenTheReceiptItselfIsAbsent() {
// #604 acceptance criterion 3: charterReceipt() can be null on its own (a session recorded
// before CB-571, or a launcher that never composed one) — the outer null-guard must still
// suppress charterSource, charterSha256 AND charterBytes together.
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null, null, null);
Map<String, Object> view = SessionManager.rosterView(s, null);
assertFalse(view.containsKey("charterSource"), "no charter fields at all when the receipt is null");
assertFalse(view.containsKey("charterSha256"), "no charter fields at all when the receipt is null");
assertFalse(view.containsKey("charterBytes"), "no charter fields at all when the receipt is null");
}
@Test
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
// The primary resolves to a Principal with no terminal, and FleetMcp's context extractor