Compare commits

...

30 Commits

Author SHA1 Message Date
Dai Ha 388ef5a3c3 fleetd #615: write FAILED instead of leaving status(token) stuck at IN_PROGRESS
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 2m25s
LeadRollover.runRollover made two unwrapped agents.send calls. HerdrException
is unchecked, and the production continuationRunner is a bare virtual thread
with no uncaught-exception handler, so a throw from either send call killed
the continuation silently — confirm() had already written IN_PROGRESS into
outcomes before scheduling it, and nothing ever overwrote that entry with a
terminal state.

Wrap the whole continuation body in one try/catch(RuntimeException), matching
the local convention already used around agents.status in
waitUntilAtTurnBoundary. On a throw, write a new terminal RollState.FAILED
entry naming the exception, in the same diagnostic style as
TURN_NEVER_SETTLED and CLEAR_NEVER_SETTLED.

Two new tests make send() throw on the /clear call and on the bootstrap-text
call respectively, each asserting status(token) reports FAILED, not
IN_PROGRESS. Reverting only the production catch (keeping the tests) turns
both red; restoring it turns them green again.
2026-09-22 10:17:12 +07:00
ltms 076cc43f7b Merge #614: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 2m20s
CI has been red on main itself since #602/#606, on this one test, so the build has been giving no second opinion on any PR. Cause: the CI job runs in a container as root. setReadable(false) really does clear the read bit, so the test's own setup guard passes, but root opens the file anyway and the gauge correctly returns OK. The test was asserting on a condition the environment never created.

The fix adds an assumeFalse(Files.isReadable(file), ...) after the chmod and before the gauge is built, inside the existing try, so the finally still restores the bit on a skip.

Verified by me on a scratch worktree merging this onto 955b9ea:
- 1864 tests, 0 failures, 0 errors, 0 skipped, 149 surefire reports, mvn exit 0. The suite-wide skipped=0 is the point: the fix did not quietly turn the test into a permanent skip.
- LeadContextGaugeTest on this non-root Mac: 9 tests, 0 skipped, and unreadableFileIsUnknown present in the report. The assumption does not fire here, so developers keep the coverage.
- Mutation: made the IOException path return OK instead of UNKNOWN. unreadableFileIsUnknown failed with "expected: <UNKNOWN> but was: <OK>". The test still has teeth. Production file reverted, git diff clean before merge.

Known trade, recorded rather than hidden: under root this case is now covered by nothing at all. A skip is honest about that, which an assertion on an unreachable state was not. The durable fix is to run the CI build as a non-root user; that is a CI configuration change and out of scope here.
2026-09-20 12:43:43 +02:00
Dai Ha bad47a8444 fleetd CI: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m27s
The test set the file's read bit off via setReadable(false), but on the
Gitea CI runner (root inside the container) the OS ignores that bit and
opens the file anyway, so the test asserted on a condition it never
actually created (LeadContextGaugeTest.java:142 UNKNOWN vs OK, CI run
1887 job 3104, commit fa62e99 on main).

Add Files.isReadable(file) after setReadable(false) and before the gauge
runs, and assumeFalse on it: a skip means "could not set up the case",
never "the behaviour is fine". Restores the read bit either way so
@TempDir cleanup still works.
2026-09-20 17:40:47 +07:00
ltms 955b9ea013 Merge #610: nudge an idle lead to hand over when its own context reads HIGH (fleetd #609)
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m51s
Closes fleetd #609. Completes the second half of the context work: #602/#606 could detect a full lead context, and nothing acted on it. LeadHeartbeatLoop now offers a handover when the lead's own gauge reads HIGH.

Never rolls a pane by itself. The nudge is text only; the lead still has to call fleet_handover, and that still needs operatorConfirmed.

Verified by me on a scratch worktree merging 89cb8ff onto fa62e99:
- 1864 tests, 0 failures, 0 errors, 149 surefire reports, mvn exit 0.
- Four mutations, all killed: the two the implementer ran (latch back in applyDecision: 4 failures; call site drops the latch argument: 2 failures), one of my own at the line the logic moved TO (latch set regardless of send outcome: 1 failure), and a control on an untouched line (quiet-cap boundary < to <=: 4 failures). The control is what makes the other kills evidence.
- The new tests assert on herdr.sentTexts() — what actually reached the fake pane — not on source text. That is the right observable for a defect whose essence was "the latch says told, the pane got nothing".

The review blocker from the first round is fixed: the latch used to be committed by applyDecision before injectNudge tried to send, and injectNudge swallows its own RuntimeException. On quietNudgeCap: 0, which is this host's configuration, that was the normal path and not an edge case. The latch is now set only when a notice was actually included and the send returned.

Known and deliberately not blocked: Fleetd.main's own one-line call to leadContextSource is not pinned by a test. That is pre-existing and class-wide, tracked in #612.
2026-09-20 12:37:43 +02:00
Dai Ha 89cb8ff79b fleetd #609 review: the context latch must mean the notice reached the pane
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m6s
Fixes the PR #610 review blocker: LeadHeartbeatLoop committed contextNotified
before injectNudge attempted the send, so a transient herdr failure marked the
lead as told when nothing reached its pane, and contextNotice() carried no
latch at all, so a pending-driven INJECT re-appended the notice on every tick
while the context stayed HIGH.

- injectNudge now reports whether agents.send succeeded and persists
  contextNotified only when a notice was actually included in the text and
  the send did not throw. The latch is split out of applyDecision (kept for
  idleSinceNanos/quietCount, applied unconditionally as before) so it is
  written on the success path only, once per branch in tick().
- contextNotice gained an overloaded 3-arg form gated on the latch as it
  stood before the tick's decision; the existing 2-arg form delegates to it
  with alreadyNotified=false, so all pre-existing callers/tests are unchanged.
- tick() is now package-private (mirrors ReplyPushLoop#tick(String)) so tests
  can drive the real send path with a fake AgentControl instead of only the
  pure decide() function.
- Added tests I-L covering: a failed send does not consume the notice and
  retries; a successful send does; the text is gated when the latch is
  already set; and the notice appears exactly once across three differently
  driven INJECTs.

Both required mutations verified red and reverted:
1. Setting the latch from the Decision regardless of send outcome -> test I
   (iAFailedSendDoesNotConsumeTheNotice) fails.
2. Dropping the latch argument at the contextNotice call site -> tests K
   (kAPendingDrivenInjectWithTheLatchAlreadySetSendsNoNotice) and L
   (lTheNoticeAppearsExactlyOnceAcrossThreeDifferentlyDrivenInjects) fail.
2026-09-20 17:34:05 +07:00
ltms fa62e9906d Merge #611: make the collected-ticket nudge test deterministic (fleetd #608)
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m48s
Replaces a wall-clock bet with a manually-driven scheduler, so the tick runs
only when the test runs it. Also closes a second, smaller race the brief did
not name: waiting on Phase.DONE is not enough, because complete() can publish
isDone() before every whenComplete dependent has run.

Verified by the lead: full suite 1841/1841 green in a clean worktree, and an
independent mutation (hasTicketWork forced true) diagnosed as an equivalent
mutant — ReplyPushLoop.injectNudge re-reads the pending collections and returns
early, so that line cannot reach agent.prompt. The worker's own mutation
(ticketCollected made a no-op) is the one that reaches the observable, and it
killed.
2026-09-20 12:26:16 +02:00
Dai Ha d7390ccd37 fleetd #609 review: repair a garbled comment carried over from the brief
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 1m41s
The brief's sentence about a null token count at HIGH was broken, and the
worker copied it into the source verbatim. The code was already right; only
the comment was unreadable.

Says what is actually true: a HIGH reading always carries a non-null token
count today, because LeadContextGauge only reaches HIGH by comparing a number
against HIGH_THRESHOLD_TOKENS. That invariant lives in another class and
nothing asserts it, so the branch stays.
2026-09-20 17:20:48 +07:00
Dai Ha aa517ae0ec fleetd #608: make anAlreadyCollectedTicketProducesNoNudge deterministic
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Failing after 1m34s
Replace the real ScheduledExecutorService backing ReplyPushLoop in this one
test with ManualScheduler, a fake that only runs a tick when the test calls
runDueTasks(). The old test bet a 300ms backoff was wide enough that
collecting the ticket always won the race against the scheduler's own timer
- true on an idle machine, false under a loaded full-suite run, which is
exactly the flake reported.

The rewritten test also waits on setAfterFinishAsyncTaskCompleteHookForTest
(already used elsewhere in this file) instead of polling Phase.DONE, so it
does not race CompletableFuture.complete()'s own publish-then-run-dependents
gap (fleetd #399) while proving ReplyPushLoop.onTicketTerminal really ran
before the ticket is collected.

Verified: backoff=1 (the most hostile value) still passes; mutating
ReplyPushLoop.ticketCollected to a no-op turns the test red with the same
assertion message the original flake reported; three consecutive full-suite
runs are green (1841/1841 each).
2026-09-20 17:18:45 +07:00
Dai Ha 60496831c2 fleetd #609: nudge an idle lead to hand over when its own context reads HIGH
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 2m5s
LeadHeartbeatLoop can now append a text-only notice to its nudge when the lead's
own LeadContextGauge reading is HIGH and leadHeartbeat.contextHighNudge is on.
Fires once per HIGH stretch (a latch, cleared only by a later OK reading; UNKNOWN
neither sets nor clears it), never spends the quietNudgeCap budget, and never
rolls a pane itself — only the operator can approve a handover.

- LeadContextGauge.Reading.unknown() widened to public for LeadContextSource.none()
- FleetConfig.LeadHeartbeat gains contextHighNudge (null/false = off, unchanged default)
- LeadHeartbeatLoop.decide gains context/contextNotified; Fleetd wires a new
  leadContextLookup/leadContextSource factory pair (LeadHeartbeatLoop.LeadContextSource)
- fleetd.example.yaml documents the new key
2026-09-20 17:11:54 +07:00
ltms 9a992d0f70 Merge #602: report a lead's live context usage in fleet_list
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m14s
CI / build (push) Failing after 1m33s
LeadContextGauge reports OK / HIGH / UNKNOWN for each lead in fleet_list, read
from the transcript Claude Code writes for itself — never from the lead's pane,
so it does not touch the control plane invariant 5 protects.

Why: a lead on this host auto-compacted 30 times in one session, discarding
roughly 250,000 tokens each time, and nothing could see it coming. The handover
feature already existed; what was missing was any way to know when to use it.

Three design points that earned their place:
  - finds the transcript by NAME under <configDir>/projects/, never by deriving
    the project slug, which is an undocumented Claude Code internal
  - three states, not two. Every path that cannot establish a token count says
    UNKNOWN with no number, so a lead is never told it is fine when the honest
    answer is "I could not look"
  - bounded twice: TAIL_BYTES caps bytes read, a 5s cache TTL caps how often

Includes #606, which fixed the two defects this PR's body recorded before it
was ever merged: the null configDir that made the gauge inert on this fleet,
and the torn final line that would have made it flap.

The authoring member died mid-task on a DNS error with the work uncommitted;
the lead recovered it, verified it, and committed it with that stated.

Verified by the lead on the combined state with current main merged in:
mvn -o clean install exit 0, 147 reports, 1841 tests, 0 failures, 0 errors,
0 skipped. LeadContextGaugeTest 9/9, FleetMcpLeadContextGaugeWiringTest 2/2,
FleetdLeadConfigDirLookupTest 6/6, FleetdLeadConfigDirSourceWiringTest 3/3.

Not deployed by this merge. A merge is not a deployment; the running daemon
still holds its old jar until redeploy.
2026-09-20 11:38:31 +02:00
ltms 3762aca307 Merge #606: wire the lead context gauge to the real configDir, and stop flapping on a torn line
CI / shell-tests (pull_request) Failing after 6s
CI / build (pull_request) Failing after 1m21s
CI / contract (pull_request) Successful in 1m50s
Fixes the two defects recorded in #602's own body.

1. FleetMcp.contextView passed configDir=null, so the gauge read <user.home>/
   .claude while this host's lead profile sets an override. Measured before the
   fix: 19 transcripts under the real directory, 0 under the fallback. The gauge
   would have deployed green and reported UNKNOWN forever, for every lead.
   Now threaded via FleetMcp.LeadConfigDirSource, built by Fleetd
   .leadConfigDirSource, following fleet.leaders.<name>.profile to that
   profile's configDir and reading config.get() live inside the lambda.

2. A torn final line no longer means UNKNOWN. fleet_list reads a transcript
   Claude Code may be mid-write on, so the last line can be cut. The old code
   treated that as fatal, which would make the gauge flap at random. The stated
   reason ("a format change should show as UNKNOWN") does not hold: a real
   format change makes EVERY line unparseable, and that case is still caught.

Verified by the lead on the combined state with current main merged in, not on
the branch alone: mvn -o clean install exit 0, 147 reports, 1841 tests, 0
failures, 0 errors, 0 skipped.

Mutation-checked independently by the lead:
  Fleetd.leadConfigDirSource body -> none()   -> KILLED (1 failure)
  FleetMcp.contextView configDir -> null      -> KILLED (worker-measured)

KNOWN RESIDUAL, documented rather than overclaimed. main's own one-line call to
leadConfigDirSource could be swapped for none() and the suite stays green. Every
member of this wiring-test family (loopHealthSource, capacitySource,
healthCoverageSource) has the identical gap — no test runs Fleetd.main far enough
to observe which factory it called. Filed separately as a class-wide problem
rather than patched here. The empirical close is the dogfood check after redeploy.
2026-09-20 11:38:22 +02:00
Dai Ha 4e27bde2d7 fleetd #602 gauge-wiring follow-up: pin Fleetd.main's LeadConfigDirSource wiring
Extract the inline new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(...))
construction in Fleetd.main into a package-private factory,
Fleetd.leadConfigDirSource, mirroring loopHealthSource/capacitySource/
healthCoverageSource. Add FleetdLeadConfigDirSourceWiringTest, which calls the
factory directly with real Profile/Leader fixtures and asserts the returned
source resolves a real configDir -- a property that is false if the factory's
body is mutated to return LeadConfigDirSource.none().

Neither FleetMcpLeadContextGaugeWiringTest nor FleetdLeadConfigDirLookupTest
could catch main losing this wiring: each builds its own instance instead of
calling what main calls. This closes that gap at the factory level, matching
the standard already accepted for loopHealthSource's own wiring test.
2026-09-20 16:34:49 +07:00
ltms b85d9b0e46 Merge #607: redeploy-fleetd.sh no longer fails a deploy that worked
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m49s
fleetd #603. The pid poll had its own fixed 10s budget and then hard-died,
while the health check right after it was allowed 60s for the same daemon.
Under launchd the java process does not exist yet when launchctl load returns,
so the script reported FAIL on a fully successful deploy — and a false FAIL in
that direction invites the hand-rolled stop/start this script exists to replace.

The pid poll now shares HEALTH_WAIT, and a miss falls through to the health
check rather than killing the run. A genuine failure still dies and still
prints the log tail. Worst-case time-to-fail roughly doubles; that cost lands
only on real failures and is the right trade.

Verified by the lead, not taken from the report. Suite exit 0, unpiped.
Two mutations run independently:
  budget back to a hardcoded 10          -> KILLED (suite exit 1)
  warn back to die on a pid miss         -> KILLED (suite exit 1) after 856dfc6
The second survived on the first submission, which is why that commit exists:
nothing covered "pid never appears but healthz answers, so do not die" — the
one case the fall-through is for, and reachable because running_pid() only
recognises a plain java -jar.
2026-09-20 11:30:10 +02:00
Dai Ha 856dfc6318 fleetd #603 review: close the untested fall-through path
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m15s
PR review (comment 17358) found a real gap via mutation testing: replacing
the "pid never found" warn with die "no process appeared" left the whole
suite green, because neither existing test drove the case the fall-through
exists for -- running_pid() never finds anything (as its own doc comment
says it eventually will) while /healthz answers anyway.

Adds test_await_daemon_started_pid_never_found_but_healthy_warns_and_survives:
running_pid always empty, poll_health_body succeeds. Asserts DIED_CALLED=0,
the warn line is emitted, and NEW_PID stays empty (the honest "could not
establish this" answer, never a guessed pid).

Verified both halves myself: reverting the warn to die "no process appeared"
turns this one test red (FAIL: await_daemon_started must not die...);
restoring it returns the suite to green.
2026-09-20 16:28:53 +07:00
Dai Ha 81c1d8e91c fleetd #603: share HEALTH_WAIT between the pid poll and the health check
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m18s
CI / build (pull_request) Successful in 1m47s
The start step gave the new process its own short, fixed 10s budget before a
hard die, while the health check right after it waits a full HEALTH_WAIT
(60s) for the same daemon. Under launchd, launchctl load returns before the
java process exists, and on a slow host that took longer than 10s -- so the
script reported "no process appeared" on a deploy that had fully succeeded.

wait_for_new_pid/await_daemon_started fold the pid poll and the health check
into one decision: the pid poll now shares HEALTH_WAIT instead of its own
shorter budget, and a miss there falls through to the health check (direct
proof the daemon is up) instead of killing the run. A genuine failure still
dies, and still prints the log tail.

Adds behavioural tests for both acceptance criteria (slow start succeeds,
genuine failure still fails and prints the log tail) in one suite run, plus
unit tests for wait_for_new_pid and a call-site test for the new function.
2026-09-20 16:21:52 +07:00
Dai Ha d345b14e43 fleetd #602 gauge-wiring: thread a lead's configured configDir into the context gauge
FleetMcp.contextView hardcoded LeadContextGauge.read(null, ...), so a lead whose
profile sets its own CLAUDE_CONFIG_DIR always read the wrong transcript directory
and reported UNKNOWN forever, with no error anywhere.

- Add FleetMcp.LeadConfigDirSource (same idiom as LeadSeatSource) and thread it
  through the constructor / listFleet overload chain / leadView / contextView.
- Add Fleetd.leadConfigDirLookup, wired at construction, following the same
  fleet.leaders.<name>.profile link leadSeatLookup already uses, one step
  further to that profile's own configDir.
- LeadContextGauge.parse: a single unparseable line (typically the final one,
  torn by a write this read raced) is now skipped rather than forcing UNKNOWN;
  only when every line in the read window fails to parse does it report
  UNKNOWN, which is the real format-change signal.
- Tests: FleetdLeadConfigDirLookupTest (lookup logic), FleetMcpLeadContextGaugeWiringTest
  (end-to-end: config naming directory A vs B decides which is read; a lead with
  no configured dir degrades without throwing), and two replacement properties in
  LeadContextGaugeTest for the torn-line fix plus its all-unparseable control.
2026-09-20 16:19:37 +07:00
ltms 9640deeffc Merge #605: fleet_list reports charterBytes alongside charterSha256
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m34s
#604 item 1. A digest answers "same or different" and cannot say how much. The
byte count is the second signal, needed exactly when two hosts find they differ.

Verified by the lead before merging, from fleetd/: mvn -o clean install exit 0,
143 reports, 1821 tests, 0 failures, 0 errors, 0 skipped; SessionManagerTest 74/74.

Includes ad3d819, correcting an invariant the added comment claimed but that
CharterReceipt.compose() does not hold: digestOf() returns null for blank text
while getBytes().length does not, so a whitespace-only charter pairs a null
digest with a non-zero size. No behaviour change — the digest gate already
omits both on that path, which is the right answer.
2026-09-20 11:18:17 +02:00
Dai Ha ad3d81941f Correct the invariant claimed in the charterBytes comment
CI / shell-tests (pull_request) Failing after 7s
CI / build (pull_request) Successful in 1m36s
CI / contract (pull_request) Successful in 1m42s
The comment said CharterReceipt never pairs a null digest with a non-zero
byte count. It can. compose() derives the digest with digestOf(), which
returns null for blank text, while the byte count is getBytes().length,
which does not. A whitespace-only role charter on a profile with no MCP
produces exactly that pair.

No behaviour change. The gate already omits both fields on that path, which
is the right answer — a size with no digest would describe an artifact we
cannot fingerprint. Only the stated reason was wrong, and a false invariant
in a comment is worse than no comment, because the next reader will widen
the gate on the strength of it.
2026-09-20 16:16:30 +07:00
Dai Ha 8b986a52e0 #604 item 1: fleet_list reports charterBytes alongside charterSha256
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 2m26s
CharterReceipt carries a byte count next to its digest, but the roster
projection in SessionManager.rosterView only ever copied the digest
across. A digest tells a lead whether two members' charters match; it
cannot say how far apart they are when they don't. Report charterBytes
too, nested in the same conditional as charterSha256 so the two travel
together: the receipt's own contract only ever pairs a non-null digest
with a real byte count, and a member with no composed charter reports
charterSource alone, unchanged from before.

Tests: the existing charter-receipt roster test now asserts charterBytes
against the receipt's own value (not a literal), plus two new cases —
no charter composed (source "none", no digest, no size) and the receipt
itself absent (no charter keys at all).
2026-09-20 16:14:45 +07:00
Dai Ha f336bcef39 CLAUDE.md: rewrap the long line my last edit left in the architects paragraph
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m48s
The fleet01 lead found this. My edit in f5c6a0e left one prose line at 128
columns inside a paragraph wrapped at about 96. It is invisible to any check
that normalises by paragraph, and visible in any raw digest of the block.

No wording changed. Block stays 20937 bytes. The wiki template gets the same
rewrap in its own commit, so the two stay byte-identical.
2026-09-20 16:13:23 +07:00
Dai Ha 3ed7bfca67 fleetd: report a lead's live context usage in fleet_list
CI / shell-tests (pull_request) Failing after 7s
CI / build (pull_request) Failing after 1m37s
CI / contract (pull_request) Successful in 2m7s
fleetd had no way to see how full a lead's context window is. On this host a
lead auto-compacted 30 times in one session, discarding roughly 250,000 tokens
and costing 46s to 3m16s each time, and nothing could see it coming.

LeadContextGauge reads the transcript Claude Code itself writes, never the
lead's pane. It finds <sessionId>.jsonl by NAME under <configDir>/projects/
rather than deriving the project slug, which is an undocumented internal.

Three states, not two: OK, HIGH, UNKNOWN. Every path that cannot positively
establish a token count reports UNKNOWN with no number, so a lead is never
told it is fine when the honest answer is "I could not look".

Bounded two ways: TAIL_BYTES caps bytes read off disk, and a 5s cache TTL caps
how often that read happens, because fleet_list is polled constantly.

Recovered by the lead: the authoring member ended on a backend error (DNS
ENOTFOUND) with this work uncommitted and unpushed in its worktree. Verified
before committing: mvn -o clean install exit 0, 1813 tests, 0 failures,
0 errors, 0 skipped, 144 reports; LeadContextGaugeTest 8/8.

KNOWN INCOMPLETE - see the PR. The fleet_list call site passes configDir=null,
which falls back to ~/.claude, but this host's lead profile sets configDir to
an override. Measured: 19 transcripts under the real configDir, 0 under the
fallback. The gauge is therefore INERT on this fleet until that is wired.
2026-09-20 16:05:15 +07:00
Dai Ha f5c6a0e4fc CLAUDE.md: architects settled the two invented specifics at line 144
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 47s
CI / build (push) Successful in 3m0s
Both specifics in the "consult architects" paragraph were mine, not the
operator's. The operator declined twice to rule on them and directed the lead
to consult architects instead, so two architects on different models settled
them over two rounds.

"after two rounds" is gone. It was a ceiling nobody had evidence for, and it
implied a counter fleetd does not have - nothing in the daemon counts rounds.
The bound is now expressed as a shape: form independent positions, then
compare. That is a floor of two without naming a number.

The three-item operator list read as complete, so a lead hitting anything not
on it would conclude it must not ask. It is now explicitly examples, and
"granting access" replaces "credentials" - the case that motivated this was a
forge merge refusal on a protected branch, which "credentials" covers only
awkwardly.

Canonical block and the wiki template updated together; sync check passes.
2026-09-19 23:32:41 +07:00
Dai Ha a7aee5b982 Merge #600: fleetd lead-rollover outcomes readable after confirm()
Adds LeadRollover.status() and a 'status' action on fleet_handover, so a lead
can find out what happened to its own roll. Every failure past confirm() was a
log.warn the lead cannot read.

Five states. IN_PROGRESS is written at the confirm hand-off, BEFORE the token
leaves 'pending', and status() reads 'outcomes' first — so there is no window
in which an in-flight roll reports UNKNOWN.

Gated by the lead: mvn -o clean install exit 0, 1819 tests, 0 failures,
0 errors, 0 skipped, 143 reports; LeadRolloverTest 43, FleetMcpHandoverTest 12.
Diff read in full. 0 source-text assertions in the new tests.
2026-09-19 23:32:32 +07:00
Dai Ha bf895616a5 fleetd: distinguish an in-flight roll from an unknown token
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Successful in 2m8s
confirm() removed the token from pending before handing the roll to the
continuation, and outcomes was only written when runRollover reached an exit.
For the whole duration of the roll the token was in neither map, so status()
answered UNKNOWN - documented as "never issued, cancelled, or aged out". A lead
polling right after its own confirm was told the roll had never been requested.

Adds RollState.IN_PROGRESS, written at the confirm hand-off rather than at the
roll's end, so there is no gap. status() now reads outcomes before pending, so
the hand-off write cannot race the removal.

Also corrects the class javadoc, which still claimed nothing calls this class.

Recovered by the lead: the authoring member ended on a backend error (the host
slept mid-response) with this work uncommitted in its worktree. Verified before
committing: mvn -o clean install exit 0, 1819 tests, 0 failures, 0 errors,
0 skipped, 143 reports; LeadRolloverTest 43 (was 39).
2026-09-19 22:31:47 +07:00
Dai Ha b874afb0af fleetd: make lead-rollover outcomes readable after confirm()
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m31s
CI / build (pull_request) Successful in 2m14s
LeadRollover previously logged every post-confirm() failure only — a lead
has no way to read the daemon log, so a roll that timed out because its
own turn never settled (or /clear never re-settled) was invisible; the
lead would carry on believing a fresh session was coming.

Add a bounded (cap=200) token -> outcome record, written at each of the
three exits in runRollover (ROLLED, TURN_NEVER_SETTLED, CLEAR_NEVER_SETTLED),
and a read-only LeadRollover#status(token) accessor. The TURN_NEVER_SETTLED
detail names turnSettleSeconds explicitly so a reader knows what to raise.

Wire a "status" action onto the fleet_handover MCP tool (handler + schema);
it never schedules, cancels, or retries anything — confirm() remains the
only path that can ever cause a /clear.

Extends LeadRolloverTest (33 -> 39 tests) covering the six acceptance
properties, and FleetMcpHandoverTest (8 -> 12) for the new tool action.
2026-09-19 16:17:39 +07:00
ltms 6eb34a654f Merge #599: fleetd #589 groups 1+2 — wiring-test 6 sites in Fleetd.main()
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m43s
Extracts 6 inline constructions in Fleetd.main() (:214-:468) to
package-private factories, each pinned by a behavioural wiring test.
Production behaviour unchanged.

Gate: test-merged onto current main (already carrying #598, which rewrote
117 lines of the same file). No conflict — the two units insert at different
anchors, #599 after capacitySource and #598 after loopHealthSource, as their
briefs specified. mvn exit=0, 1805 tests / 0 failures / 0 errors from 143
surefire reports; all 6 new classes confirmed to have run with their
expected counts. Diff confirmed extraction-only.

No source-text assertions in any of the 6 new files; that zero carries a
positive control (the same pattern finds 18 such files elsewhere in the
repo). Worker self-reported fixing a mutation that threw NullPointerException
rather than failing an assertion — the assertNotNull guard is present ahead
of the matcher call, confirmed.
2026-09-19 10:39:47 +02:00
ltms 7084d99b89 Merge #597: fleetd #593 — running_pid() counts only the daemon
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m39s
CI / build (push) Successful in 1m50s
running_pid() filters pgrep -f hits by `comm = java` (allowlist) instead of
denying a fixed list of shell names. Closes two false-positive holes: an
exited pid (empty comm matched no denied name) and any non-shell wrapper
(ssh, perl, python3, ruby) carrying the pattern in its own argv.

Gate: merged onto current main, bash suite exit 0 and structurally identical
to the baseline run on main (the one "Unattributable mutation" line is
pre-existing, confirmed by running the suite on origin/main). All four
running_pid tests confirmed defined AND invoked. Mutation check run by me:
neutering the allowlist makes the suite exit 1 with a named failure; restore
is byte-identical to baseline by git hash-object and green again.

Also checked and cleared: the new die-message advice `ps -eo pid,comm,args`
does NOT expose process environments on macOS — `-e` with `-o` selects all
processes, it does not imply `-E`. Verified with an isolated two-phase probe
and a positive control, after three earlier probes gave false positives by
self-matching (the grep's own argv, and the probe script's own text).
2026-09-19 10:37:50 +02:00
Dai Ha ae7845c375 fleetd #589 (Groups 1 & 2): pin 6 main() wiring sites with named factories
CI / shell-tests (pull_request) Successful in 9s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m52s
Extracts 6 inline wiring expressions from Fleetd.main() into named,
directly-testable package-private static factories, following the
FleetdLoopHealthSourceWiringTest (#584) shape, and adds one wiring test
per factory:

Group 1 (exhaustion/quarantine):
- forwardingExhaustionSink(exhaustionSinkRef) — was inline
  ExhaustionSink.forwardingTo(exhaustionSinkRef::get)
- publishExhaustionSink(...) — was two untested statements building the
  real sink and .set()-ing it into exhaustionSinkRef
- liveExhaustedPatterns(config) — was inline
  new LiveExhaustedPatterns(() -> config.get().profiles())
- exhaustedPatternLookup(roster, liveExhaustedPatterns) — was an inline
  lambda resolving a herdr target to its profile's live pattern; silently
  losing this is the worst regression in the sweep, since a real
  usage-limit refusal would stop being classified as BACKEND_EXHAUSTED

Group 2 (CB-596 credential policy):
- claudeCodeLauncher(...) — was an inline `new ClaudeCodeLauncher(...)`
  whose memberCredentials supplier argument was untestable wiring
- openCodeLauncher(...) — same, for OpenCodeLauncher

Each new test pins its factory behaviorally (never via source-text
assertions): built and confirmed RED by name against the named inert
mutation, then confirmed GREEN again after restoring, and separately
confirmed GREEN after a behavior-preserving reformat/local-variable
extraction of the same call, to rule out a disguised source-text test.

Suite: 1789 -> 1799 tests (+10, matching the 10 tests added), 0
failures, mvn -o clean install BUILD SUCCESS.

Scope strictly limited to main()'s :214-:468 range per the ticket split
with the concurrent worker handling Group 3 at line 500+.
2026-09-19 15:35:32 +07:00
ltms 61115f6f61 Merge #598: fleetd #589 group 3 — wiring-test 5 sites in Fleetd.main()
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 50s
CI / build (push) Successful in 2m9s
Extracts 5 inline lambdas/method-refs in Fleetd.main() to package-private
factories and pins each with a wiring test. Production behaviour unchanged;
the releaseCleanup body moved verbatim.

Gate: test-merged onto current main in a scratch worktree, mvn exit=0,
1795 tests / 0 failures / 0 errors from 137 surefire reports. Diff read in
full. Worker's self-disclosed bare `git stash push` verified as recovered —
all 3 surviving stash entries predate today, so no other worktree lost work.
2026-09-19 10:33:07 +02:00
Dai Ha 6cccd458d4 #589 Group 3: wiring-test the 5 sites below line 500 in Fleetd.main()
CI / shell-tests (pull_request) Successful in 11s
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 2m11s
Extracts the inline lambdas/method references at the 5 assigned wiring
sites into named package-private factories on Fleetd, following the
FleetdLoopHealthSourceWiringTest pattern from #584:

- turnRegistrar(CompletionResolver) — was completion::register (Injector)
- healthFailTarget(MessageService) — was messages::abandon (FleetHealthMonitor)
- releaseCleanup(MessageService, ReplyInbox, PrimaryRegistry) — was the
  inline sessions.onRelease(detail -> {...}) cleanup lambda
- replyInboxOpener() — was AmqpReplyInbox::open passed to selectReplyInbox
- leadMailboxOpener() — was LeadMailbox::open passed to openLeadMailbox

Each factory has a new runtime test (not source-text) that drives real
collaborators through public APIs: MessageService.poll(ticket).phase(),
InMemoryReplyInbox.peek(), PrimaryRegistry.nudgeTargetFor(), and the
opener tests connect to a guaranteed-closed local port to prove a real
network attempt vs. an inert stub.

releaseCleanup was done first per the brief: MessageService.abandon's
javadoc documents that losing this cleanup leaves a torn-down worker's
rendezvous waiter open forever.

Tests: 1789 -> 1794 (+5), 0 failures, 0 errors. mvn -q -o test exit 0,
no BUILD FAILURE, no piped exit status. Each new test verified RED on
the inert form named in the ticket, and GREEN after reformatting the
call across lines and extracting the argument into a local/factory.
2026-09-19 15:24:11 +07:00
34 changed files with 4327 additions and 156 deletions
+8 -6
View File
@@ -141,12 +141,14 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
**When a decision blocks you, consult architects — not the operator.** Spawn one or more architect
members, give them the question and the evidence you have, and act on what they agree. They are
authorized to settle it, not only to advise. If two of them still disagree after two rounds, they
return both positions and you decide. Go to the operator only for something outside the fleet's
authority: money, credentials, or a promise made to someone else. **Then write the decision on the
ticket.** Taking the operator out of the loop also removes the signal they used to get, because
that signal was the block itself — work stopped, so they found out. A ticket comment replaces it,
and it reaches them whether or not they are at a terminal when you decide.
authorized to settle it, not only to advise. Architects first form independent positions, then
compare them. If they still disagree after that comparison, they return both positions and their
checked evidence; the lead decides. Go to the operator only for an action the fleet has no
authority to take, such as spending money, granting access, or making a promise to someone else.
**Then write the decision on the ticket.** Taking the operator out of the loop also removes the
signal they used to get, because that signal was the block itself — work stopped, so they found
out. A ticket comment replaces it, and it reaches them whether or not they are at a terminal when
you decide.
| Intent | Tool |
|---|---|
+10 -1
View File
@@ -77,16 +77,25 @@ bind:
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
# block = feature off, exactly as before.
#
# Three knobs, each with a default that errs on the side of not burning context:
# Four knobs, each with a default that errs on the side of not burning context:
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
# # until real state appears (default 3 — never nag an empty fleet forever)
# contextHighNudge: false # fleetd #609 — when true, an idle lead whose OWN Claude Code context
# # reads HIGH (see LeadContextGauge; fleet_list's context row) gets a text
# # notice appended to its nudge telling it to consider fleet_handover. Text
# # only — it never rolls a pane by itself, and only the operator can approve
# # a roll. Fires once per HIGH stretch (a later OK reading re-arms it), and
# # never spends the quietNudgeCap budget. Default false/absent = off, same
# # as every other knob here — an upgraded daemon must not start telling
# # leads to hand over on its own.
# leadHeartbeat:
# idleAfterSeconds: 300
# backoffMs: 60000
# quietNudgeCap: 3
# contextHighNudge: false
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
+371 -46
View File
@@ -4,11 +4,13 @@ import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.ConfigWatcher;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.herdr.PaneLocator;
@@ -24,6 +26,7 @@ import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.LoopWatchdog;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.inject.TurnListener;
import dev.ltms.fleet.inject.TurnRegistrar;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.CallerResolver;
@@ -80,7 +83,9 @@ import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.BiConsumer;
import java.util.function.BooleanSupplier;
import java.util.function.Consumer;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Predicate;
@@ -218,23 +223,23 @@ public final class Fleetd {
// there is no 2-arg overload left for any lambda to silently bind to instead), but so a
// test can call the exact same object this line builds, instead of asserting a copy of its
// shape (round 3's lesson).
ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);
// fleetd #589 Group 1: extracted to forwardingExhaustionSink(...) below (see that method's
// javadoc) so a dedicated test can prove this factory keeps reading the reference live,
// rather than a rebuilt copy of its shape.
ExhaustionSink forwardingExhaustionSink = forwardingExhaustionSink(exhaustionSinkRef);
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
// fleetd #589 Group 2: extracted to claudeCodeLauncher(...)/openCodeLauncher(...) below (see
// those methods' javadoc) so a dedicated test can prove the CB-596 memberCredentials policy
// supplier is actually wired to each adapter, not silently replaced with `() -> null`.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials(), null, config::get));
adapters.add(claudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
claudeProfiles, cfg, config));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(router.memberAgents(), router.memberSpaces(),
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials(), config::get, forwardingExhaustionSink));
adapters.add(openCodeLauncher(router.memberAgents(), router.memberSpaces(),
opencodeProfiles, cfg, config, forwardingExhaustionSink));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
@@ -408,12 +413,12 @@ public final class Fleetd {
// LiveExhaustedPatterns's class doc for why this replaces the old compiled-once-at-startup
// map. A profile with no exhaustedPattern simply returns null here, so its workers keep
// today's completion-fallback behaviour unchanged.
LiveExhaustedPatterns liveExhaustedPatterns = new LiveExhaustedPatterns(() -> config.get().profiles());
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(session -> liveExhaustedPatterns.patternFor(session.profile()))
.orElse(null);
// fleetd #589 Group 1: both extracted to liveExhaustedPatterns(...)/
// exhaustedPatternLookup(...) below (see those methods' javadoc) — this is the worst
// consequence in the whole #589 sweep: silently losing either wiring means a genuine
// usage-limit refusal is handed back as a real completion instead of BACKEND_EXHAUSTED.
LiveExhaustedPatterns liveExhaustedPatterns = liveExhaustedPatterns(config);
ExhaustedPatternLookup exhaustedPatterns = exhaustedPatternLookup(sessions::roster, liveExhaustedPatterns);
// The startup coverage line still reports the boot-time snapshot only — it is printed once,
// here, and a reload no longer needs to change what it said; exhaustionDetectionArmed (via
// liveExhaustedPatterns.armed, wired into quarantineSource below) is what stays live.
@@ -461,11 +466,13 @@ public final class Fleetd {
// method's javadoc for the full fleetd #175/#234/#446 history this used to carry inline —
// so a dedicated test can drive the exact ExhaustionSink main() builds, not a hand-rebuilt
// copy of its shape.
ExhaustionSink exhaustionSink = exhaustionSink(sessions, config, quarantine,
quarantineReasonByCredential, cfg);
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
// now that `sessions` exists to resolve target -> session -> profile.
exhaustionSinkRef.set(exhaustionSink);
// fleetd #589 Group 1: both statements (build + set) folded into publishExhaustionSink(...)
// below (see that method's javadoc), so a test can prove the reference is actually
// repointed at the real sink, not silently left at ExhaustionSink.none().
ExhaustionSink exhaustionSink = publishExhaustionSink(exhaustionSinkRef, sessions, config,
quarantine, quarantineReasonByCredential, cfg);
// fleetd #201 Unit 5: the production BackendErrorSink needs `pushLoop` (built further below,
// after `sessions`) to tell a lead about an incident or an unmapped target — the same
// construction-order cycle `exhaustionSinkRef` breaks above, broken the same way: a mutable
@@ -501,21 +508,26 @@ public final class Fleetd {
// fleetd #556: registration is wired directly to `completion`, not folded into the
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
// listener) throwing, regardless of call order. See TurnRegistrar's javadoc.
// fleetd #589 Group 3 (:505): extracted to turnRegistrar(...) — see FleetdTurnRegistrarWiringTest.
Injector injector = new Injector(router, turnListener, deliverable,
presence::forget, completion::register);
presence::forget, turnRegistrar(completion));
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
// unusable), fleetd stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox = selectReplyInbox(cfg.broker(), System.getenv(), AmqpReplyInbox::open);
// fleetd #589 Group 3 (:512): opener extracted to replyInboxOpener() — see
// FleetdReplyInboxOpenerWiringTest.
final ReplyInbox replyInbox = selectReplyInbox(cfg.broker(), System.getenv(), replyInboxOpener());
// CB-637: this daemon's lead-to-lead mailbox on the SHARED coordination vhost — a separate
// broker from the reply inbox by design (see FleetConfig.Coordinator). Absent a coordinator:
// block this is null and every lead path below is simply not wired, which is exactly the
// behaviour before this ticket. It owns a broker connection, so keep the reference for the
// ordered shutdown hook.
final LeadMailbox leadMailbox = openLeadMailbox(cfg.coordinator(), System.getenv(), LeadMailbox::open);
// fleetd #589 Group 3 (:518): opener extracted to leadMailboxOpener() — see
// FleetdLeadMailboxOpenerWiringTest.
final LeadMailbox leadMailbox = openLeadMailbox(cfg.coordinator(), System.getenv(), leadMailboxOpener());
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
@@ -553,10 +565,18 @@ public final class Fleetd {
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
// fleetd #609: own LeadContextGauge instance for the heartbeat loop — separate from the
// one FleetMcp builds internally for fleet_list's context row. Each caches independently
// (keyed by configDir+sessionId), so this costs at most one extra bounded tail read per
// TTL window, never a shared-mutable-state hazard between the two callers.
var leadContextGauge = new LeadContextGauge();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, System::nanoTime,
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics);
metrics,
leadContextSource(leadContextGauge, router.leadAgents(), leads,
leadConfigDirLookup(() -> config.get().profiles(), leaders)),
Boolean.TRUE.equals(hb.contextHighNudge()));
heartbeat.start();
} else {
heartbeat = null;
@@ -585,10 +605,12 @@ public final class Fleetd {
// fleetd #386: System::nanoTime freezes across a macOS sleep, so the stall check also
// gets a wall-clock source to detect and correct for that freeze. Every other decision
// in FleetHealthMonitor stays on the monotonic clock, unchanged.
// fleetd #589 Group 3 (:591): failTarget extracted to healthFailTarget(...) — see
// FleetdHealthFailTargetWiringTest.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis()),
cfg.health().intervalOrDefault(),
cfg.health().workingSuspectAfterOrDefault(), messages::abandon);
cfg.health().workingSuspectAfterOrDefault(), healthFailTarget(messages));
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
@@ -608,27 +630,11 @@ public final class Fleetd {
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(detail -> {
// CB-578 stage C, acceptance criterion 10: a failed ticket's detail should tell a lead
// where to re-dispatch onto the same tree, not just that the worker vanished.
String reason = "the worker session was released before it replied";
if (detail.worktreePath() != null) {
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
}
// CB-584 (issue #65 criterion 5): also name the agent session, so a lead can resume the
// member's conversation instead of only re-dispatching a fresh one onto the same files.
if (detail.agentSessionId() != null) {
reason += " agentSessionId=" + detail.agentSessionId();
}
// fleetd #275: this is an explicit teardown (fleet_stop, or the idle reaper) — the
// worker's pane is being stopped right now, so an open fleet_ask has no turn left to
// resume into. Sweep it too, unlike FleetHealthMonitor's health-classification call
// (see MessageService.abandon's javadoc for why those two must differ).
messages.abandon(detail.terminalId(), reason, true);
replyInbox.release(detail.terminalId());
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
});
// fleetd #589 Group 3 (:611-631): the whole cleanup lambda extracted to releaseCleanup(...)
// — see FleetdReleaseCleanupWiringTest, and MessageService.abandon's javadoc for the
// documented incident (a torn-down worker's rendezvous waiter left open) this lambda exists
// to prevent.
sessions.onRelease(releaseCleanup(messages, replyInbox, primaryRegistry));
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
@@ -676,6 +682,10 @@ public final class Fleetd {
leadMailbox,
outageSource,
new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), leaders, leads)),
// fleetd #602 gauge-wiring: threads each lead's configured configDir into the
// context gauge — see leadConfigDirSource's own doc for why this, not a hardcoded
// null, is what fleet_list's context row now reads.
leadConfigDirSource(() -> config.get().profiles(), leaders),
// fleetd #361: the operator-declared peers this daemon's fleet_list should try to
// reach. Read from the SAME snapshot leadMailbox itself opened from (cfg.coordinator()),
// not the live config.get() — coordinator wiring is already boot-time-fixed (see
@@ -1009,6 +1019,120 @@ public final class Fleetd {
cfg.profiles()::keySet, System::nanoTime);
}
/**
* fleetd #589 Group 1: the forwarding {@link ExhaustionSink} handed to the adapters built
* before {@code sessions} exists (see the {@code exhaustionSinkRef}/{@code
* forwardingExhaustionSink} locals in {@code main}, just above {@link #capacitySource}'s call
* site). Before this ticket, {@code ExhaustionSink.forwardingTo(exhaustionSinkRef::get)} was
* built inline — nothing a test could call directly, so a mutation swapping the supplier for a
* hardcoded {@code () -> ExhaustionSink.none()} compiled clean and left the suite green: the
* forwarder would silently stop reading the reference at all, and {@link
* #publishExhaustionSink} repointing that reference later would have no effect.
*
* <p>Extracted the same way {@link #capacitySource}/{@link #loopHealthSource} were, so {@code
* FleetdExhaustionSinkForwardingWiringTest} can call this factory directly with a real {@link
* AtomicReference}, mutate the reference AFTER the forwarder is built, and prove the forwarder
* still reads it live rather than a fixed target captured at construction time.
*/
static ExhaustionSink forwardingExhaustionSink(AtomicReference<ExhaustionSink> exhaustionSinkRef) {
return ExhaustionSink.forwardingTo(exhaustionSinkRef::get);
}
/**
* fleetd #589 Group 1: publish the real {@link ExhaustionSink} — built the same way {@link
* #exhaustionSink} always was — into the forwarding reference {@link #forwardingExhaustionSink}
* built above, replacing {@code main}'s previously untested two-statement sequence ({@code
* ExhaustionSink exhaustionSink = exhaustionSink(...); exhaustionSinkRef.set(exhaustionSink);}).
* Before this ticket, nothing proved the {@code .set(...)} call actually received the real sink
* rather than a hardcoded {@code ExhaustionSink.none()} — the whole point of {@code
* exhaustionSinkRef} existing (fleetd #175) is that {@link
* dev.ltms.fleet.member.OpenCodeLauncher}'s model-mismatch check, built before {@code sessions}
* exists, keeps working once this line runs; silently keeping the reference at {@code none()}
* would mean that check permanently does nothing, with the full suite still green because no
* existing test drives this exact call site.
*
* <p>Returns the built sink so {@code main} can still pass it to {@link CompletionResolver}'s
* constructor at the same call site it already does, without building it twice.
*/
static ExhaustionSink publishExhaustionSink(AtomicReference<ExhaustionSink> exhaustionSinkRef,
SessionManager sessions, ConfigRef config, BackendQuarantine quarantine,
Map<String, String> quarantineReasonByCredential, FleetConfig cfg) {
ExhaustionSink sink = exhaustionSink(sessions, config, quarantine, quarantineReasonByCredential, cfg);
exhaustionSinkRef.set(sink);
return sink;
}
/**
* fleetd #589 Group 1: the live {@code exhaustedPattern} source (fleetd #446) {@link
* CompletionResolver} enforces on, extracted out of {@code main} for the same reason {@link
* #capacitySource} was. Before this ticket {@code new LiveExhaustedPatterns(() ->
* config.get().profiles())} was built inline; replacing the supplier with a hardcoded {@code ()
* -> Map.of()} compiled clean and left the suite green, meaning every profile's {@code
* exhaustedPattern} would silently stop being recognised and a genuine usage-limit refusal
* would be handed back as a real completion instead of {@code BACKEND_EXHAUSTED}.
*/
static LiveExhaustedPatterns liveExhaustedPatterns(ConfigRef config) {
return new LiveExhaustedPatterns(() -> config.get().profiles());
}
/**
* fleetd #589 Group 1: the {@link ExhaustedPatternLookup} {@link CompletionResolver} enforces
* on, resolving a herdr {@code target} to its session's profile and then to that profile's live
* {@link LiveExhaustedPatterns#patternFor}. Extracted out of {@code main} the same way {@link
* #worktreeBranchLookup} was — same {@code Supplier<List<MemberSession>>} roster shape, same
* reason: before this ticket the lambda was built inline, and replacing it with {@code target ->
* null} (the exact shape of {@link ExhaustedPatternLookup#none()}) compiled clean and left the
* suite green. This is the worst consequence in the whole #589 sweep (see the ticket): a
* genuine usage-limit refusal would stop being classified as {@code BACKEND_EXHAUSTED} and
* would be handed back to a waiting {@code fleet_send} as if it were real completed work.
*
* @param roster the live member roster, normally {@code sessions::roster}
*/
static ExhaustedPatternLookup exhaustedPatternLookup(Supplier<List<MemberSession>> roster,
LiveExhaustedPatterns liveExhaustedPatterns) {
return target -> roster.get().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(session -> liveExhaustedPatterns.patternFor(session.profile()))
.orElse(null);
}
/**
* fleetd #589 Group 2: the production {@link ClaudeCodeLauncher} adapter, extracted out of
* {@code main} the same way {@link #capacitySource} was. Before this ticket the constructor
* call (11 arguments, including the CB-596 {@code memberCredentials} policy supplier) was built
* inline; replacing the {@code () -> config.get().memberCredentials()} argument with {@code ()
* -> null} compiled clean and left the suite green — {@code memberCredentials} is not {@code
* null} itself (a lambda is never {@code null}), so {@link
* dev.ltms.fleet.member.HerdrPeerLauncher#applyMemberCredentialPolicy} sees {@code
* memberCredentials.get() == null} and silently shadows nothing, reopening the exact CB-592
* exposure gap CB-596's policy closed. {@code FleetdClaudeCodeLauncherCredentialWiringTest}
* calls this factory with a real {@link ConfigRef} carrying a {@code memberCredentials:} block
* and proves a known-but-not-allowed name is actually shadowed on {@code spawn()}.
*/
static ClaudeCodeLauncher claudeCodeLauncher(AgentControl agents, WorkspaceControl spaces,
SubscriptionGuard guard, Map<String, FleetConfig.Profile> claudeProfiles, FleetConfig cfg,
ConfigRef config) {
return new ClaudeCodeLauncher(agents, spaces, guard, claudeProfiles, cfg.effectiveDefaultProfile(),
System::getenv, cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(), () -> config.get().memberCredentials(), null, config::get);
}
/**
* fleetd #589 Group 2: the production {@link OpenCodeLauncher} adapter, the {@code opencode}
* counterpart to {@link #claudeCodeLauncher} above and extracted for the identical reason: the
* same {@code () -> config.get().memberCredentials()} argument, reopening the same CB-592
* exposure gap if silently replaced with {@code () -> null}.
*/
static OpenCodeLauncher openCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> opencodeProfiles, FleetConfig cfg, ConfigRef config,
ExhaustionSink forwardingExhaustionSink) {
return new OpenCodeLauncher(agents, spaces, opencodeProfiles, cfg.effectiveDefaultProfile(),
System::getenv, cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(), () -> config.get().memberCredentials(), config::get,
forwardingExhaustionSink);
}
/**
* fleetd #426: package-private factory for {@code fleet_list}'s {@code healthCoverage} source,
* extracted out of {@code main} for the same reason {@link #capacitySource} and {@link
@@ -1066,6 +1190,104 @@ public final class Fleetd {
() -> reaper == null ? LoopWatchdog.State.STOPPED : reaper.health());
}
/**
* fleetd #589 Group 3 (site {@code :505}): package-private factory for the {@link Injector}'s
* {@link TurnRegistrar}, extracted out of {@code main} for the same reason {@link
* #loopHealthSource} was — before this ticket {@code completion::register} was an inline
* argument to {@code new Injector(...)}, so nothing could pin it directly. Replacing it with
* {@link TurnRegistrar#NOOP} compiles clean and leaves every existing test green: {@code
* onDelivered}'s own {@code captureBaseline} does the identical {@code inFlight} check-and-put a
* moment later on the ordinary path, so the two are indistinguishable once a turn's delivery
* finishes normally. The gap {@link TurnRegistrar}'s own javadoc (fleetd #556) exists to close is
* a {@code turnListener} callback throwing between the two — {@link
* FleetdTurnRegistrarWiringTest} pins that {@code register} itself (not just {@code
* captureBaseline}) makes a delivered turn's waiter resolvable.
*/
static TurnRegistrar turnRegistrar(CompletionResolver completion) {
return completion::register;
}
/**
* fleetd #589 Group 3 (site {@code :591}): package-private factory for {@link
* FleetHealthMonitor}'s {@code failTarget} callback, extracted out of {@code main} for the same
* reason {@link #loopHealthSource} was. Before this ticket {@code messages::abandon} was an
* inline argument to {@code new FleetHealthMonitor(...)}; replacing it with a no-op {@code
* BiConsumer} compiles clean and leaves every existing test green, and in production it means a
* member found {@code GONE}/{@code NEVER_READY} never fails the ticket waiting on it — the
* caller reports {@code PENDING} for the full 30-minute async timeout instead of the immediate,
* accurate failure CB-580 exists to give it. {@link FleetdHealthFailTargetWiringTest} pins that
* the returned callback actually reaches the real {@link MessageService#abandon}.
*/
static BiConsumer<String, String> healthFailTarget(MessageService messages) {
return messages::abandon;
}
/**
* fleetd #589 Group 3 (site {@code :611-631}): package-private factory for the whole {@link
* SessionManager#onRelease} cleanup callback, extracted out of {@code main} for the same reason
* {@link #loopHealthSource} was. Before this ticket this was an inline lambda built directly
* inside {@code main}; replacing its body with a no-op {@code detail -> { }} compiles clean and
* leaves every existing test green, and in production it is the exact incident {@link
* MessageService#abandon}'s own javadoc documents: a torn-down worker's rendezvous waiter is
* left open, so a blocking {@code fleet_send} keeps blocking and an async one reports {@code
* PENDING} for a hardcoded thirty minutes on every {@code fleet_stop} and every idle-reap.
*
* <p>{@link FleetdReleaseCleanupWiringTest} pins all three collaborator calls this lambda makes
* — {@code messages.abandon}, {@code replyInbox.release}, and {@code
* primaryRegistry.forgetDelegation} — each already tested on its own ({@code MessageServiceTest},
* {@code PrimaryRegistryTest}), but never before proven to actually be reached from here.
*/
static Consumer<SessionManager.ReleaseDetail> releaseCleanup(MessageService messages, ReplyInbox replyInbox,
PrimaryRegistry primaryRegistry) {
return detail -> {
// CB-578 stage C, acceptance criterion 10: a failed ticket's detail should tell a lead
// where to re-dispatch onto the same tree, not just that the worker vanished.
String reason = "the worker session was released before it replied";
if (detail.worktreePath() != null) {
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
}
// CB-584 (issue #65 criterion 5): also name the agent session, so a lead can resume the
// member's conversation instead of only re-dispatching a fresh one onto the same files.
if (detail.agentSessionId() != null) {
reason += " agentSessionId=" + detail.agentSessionId();
}
// fleetd #275: this is an explicit teardown (fleet_stop, or the idle reaper) — the
// worker's pane is being stopped right now, so an open fleet_ask has no turn left to
// resume into. Sweep it too, unlike FleetHealthMonitor's health-classification call
// (see MessageService.abandon's javadoc for why those two must differ).
messages.abandon(detail.terminalId(), reason, true);
replyInbox.release(detail.terminalId());
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
};
}
/**
* fleetd #589 Group 3 (site {@code :512}): package-private factory for {@code selectReplyInbox}'s
* production {@link AmqpOpener}, extracted out of {@code main} for the same reason {@link
* #loopHealthSource} was. Before this ticket {@code AmqpReplyInbox::open} was an inline argument
* to the {@code selectReplyInbox(...)} call; replacing it with {@code (uri, prefetch) -> new
* InMemoryReplyInbox()} compiles clean and leaves every existing test green — {@link
* FleetdReplyInboxSelectionTest} drives {@code selectReplyInbox} with its own injected opener and
* never sees what {@code main} actually passes. {@link FleetdReplyInboxOpenerWiringTest} pins
* that this returns the real opener by pointing it at a guaranteed-closed local port and
* asserting the real network attempt throws — the inert stub never attempts a connection at all.
*/
static AmqpOpener replyInboxOpener() {
return AmqpReplyInbox::open;
}
/**
* fleetd #589 Group 3 (site {@code :518}): package-private factory for {@code
* openLeadMailbox}'s production {@link LeadMailboxOpener}, extracted out of {@code main} for the
* same reason {@link #replyInboxOpener} was — same gap, same fix, the lead-coordination mailbox
* instead of the reply inbox. {@link FleetdLeadMailboxOpenerWiringTest} pins that this returns
* the real opener the same way.
*/
static LeadMailboxOpener leadMailboxOpener() {
return LeadMailbox::open;
}
/**
* fleetd #248: package-private factory for the member worktree/branch lookup {@link
* CompletionResolver} uses to name a fallback report's worktree and branch (fleetd#241).
@@ -1368,6 +1590,109 @@ public final class Fleetd {
};
}
/**
* fleetd #602 gauge-wiring: per-lead-name factory for {@link FleetMcp.LeadConfigDirSource} — the
* {@code CLAUDE_CONFIG_DIR} the named lead's own profile runs on, so {@code LeadContextGauge}
* reads the transcript directory that lead's Claude Code process actually writes to, not always
* the built-in {@code <user.home>/.claude} default.
*
* <p>Follows the same {@code fleet.leaders.<name>.profile} link {@link #leadSeatLookup} already
* uses to find a lead's profile, one step further to that profile's own {@code configDir:}. A
* lead entry that names no {@code profile:}, or whose named profile is not configured, or whose
* profile sets no {@code configDir:} override, returns {@code null} — {@code LeadContextGauge}
* then falls back to its own default, exactly as before this ticket.
*
* @param profiles the live profile map, normally {@code () -> config.get().profiles()} in
* {@code main} — read live, like every other {@code configDir} lookup, so a
* config reload takes effect on the very next {@code fleet_list} call
* @param leaders {@code fleet.leaders}, read once at startup like {@link #leadSeatLookup}'s own
* {@code leaders} parameter — passed as a plain map, never re-read from
* {@code config.get()}
*/
static Function<String, String> leadConfigDirLookup(Supplier<Map<String, FleetConfig.Profile>> profiles,
Map<String, FleetConfig.Leader> leaders) {
return leadName -> {
FleetConfig.Leader lead = leaders.get(leadName);
if (lead == null || lead.profile() == null || lead.profile().isBlank()) {
return null;
}
FleetConfig.Profile leadProfile = profiles.get().get(lead.profile());
return leadProfile == null ? null : leadProfile.configDir();
};
}
/**
* fleetd #602 gauge-wiring follow-up (PR #606 review comment 17353): {@code main} used to build
* {@code new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(...))} inline, with nothing a test
* could call directly. Measured on that shape: replacing the whole expression with {@code
* FleetMcp.LeadConfigDirSource.none()} at the call site compiled with 0 errors and left the full
* 1822-test suite green — the daemon could be changed to always report every lead's context as
* {@code UNKNOWN}, forever, while every test stayed green. That is the same hand-built-vs-wired
* shape as fleetd #561/#248/#426/#562 ({@link #loopHealthSource}).
*
* <p>The fix extracts the inline {@code new} into this factory, in the same style as {@link
* #loopHealthSource}/{@link #capacitySource}/{@link #healthCoverageSource} — which is exactly
* what makes it directly callable from {@code FleetdLeadConfigDirSourceWiringTest}. That test
* calls this factory with real {@link FleetConfig.Profile}/{@link FleetConfig.Leader} fixtures and
* asserts the returned source resolves a real {@code configDir} — a property that would be false
* if this method's body were mutated to {@code return FleetMcp.LeadConfigDirSource.none();}.
*/
static FleetMcp.LeadConfigDirSource leadConfigDirSource(Supplier<Map<String, FleetConfig.Profile>> profiles,
Map<String, FleetConfig.Leader> leaders) {
return new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(profiles, leaders));
}
/**
* fleetd #609: per-terminal factory for {@link LeadHeartbeatLoop.LeadContextSource} — the lead's
* own {@link LeadContextGauge} reading, so the heartbeat loop can tell an idle, HIGH-context lead
* to consider a handover.
*
* <p>Three hops, each degrading to {@link LeadContextGauge.Reading#unknown()} rather than
* throwing, since a herdr hiccup or an unrecognised terminal must never kill the heartbeat's own
* tick: {@code liveLeadTerminals} (terminal id → lead name, the same live supplier {@link
* #leadSeatLookup} and {@code LeadCoordLoop} already read) → {@code configDirForLeadName} (that
* lead's {@code configDir}, normally {@link #leadConfigDirLookup}'s return) → {@code agents.get}
* for the live {@link Agent#sessionId()}/{@link Agent#agentType()} the gauge itself needs.
*
* @param gauge the {@link LeadContextGauge} instance to read through — shares its
* cache across every call this factory's function makes
* @param agents the {@link AgentControl} instance that reaches the LEAD's pane
* (not {@code memberAgents}), normally {@code router.leadAgents()}
* @param liveLeadTerminals terminal id → lead name for every CURRENTLY recognised lead
* @param configDirForLeadName lead name → {@code configDir}, normally {@link
* #leadConfigDirLookup}'s return
*/
static Function<String, LeadContextGauge.Reading> leadContextLookup(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
return terminal -> {
String leadName = liveLeadTerminals.get().get(terminal);
if (leadName == null) {
return LeadContextGauge.Reading.unknown();
}
String configDir = configDirForLeadName.apply(leadName);
Agent live;
try {
live = agents.get(terminal);
} catch (RuntimeException e) {
return LeadContextGauge.Reading.unknown();
}
return gauge.read(configDir, live.sessionId(), live.agentType());
};
}
/**
* fleetd #609: wraps {@link #leadContextLookup} into a {@link LeadHeartbeatLoop.LeadContextSource}
* — the same hand-built-vs-wired shape as {@link #leadConfigDirSource}/{@link #loopHealthSource}/
* {@link #capacitySource}/{@link #healthCoverageSource}. Extracted so a test can call the exact
* factory {@code main} calls, rather than only a lookup nothing in {@code main} is proven to use
* (see {@code FleetdLeadConfigDirSourceWiringTest}'s javadoc for the measured gap this shape closes).
*/
static LeadHeartbeatLoop.LeadContextSource leadContextSource(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
return new LeadHeartbeatLoop.LeadContextSource(
leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName));
}
/**
* fleetd #248 / fleetd#201 Unit 5: package-private factory for the per-target backend-error
* pattern lookup {@link CompletionResolver} classifies a pane scrape against. Closes over the
@@ -1329,9 +1329,20 @@ public record FleetConfig(
* each such nudge costs the lead a turn just to read "nothing pending";
* three is enough to tell it it may stand down without nagging forever, and
* it is the bound that stops an idle fleet from being a subscription burner.
* @param contextHighNudge fleetd #609: when {@code true}, an idle lead whose own {@code
* LeadContextGauge} reading is {@code HIGH} gets a text notice telling it
* to consider a handover, appended to whatever heartbeat nudge the loop
* already sends. Default {@code false} ({@code null} also means off) — an
* upgraded daemon must not silently start telling leads to hand over.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap) {
public record LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap,
Boolean contextHighNudge) {
/** Convenience constructor for every call site that predates fleetd #609: no context notice. */
public LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap) {
this(idleAfterSeconds, backoffMs, quietNudgeCap, null);
}
public LeadHeartbeat {
idleAfterSeconds = (idleAfterSeconds == null || idleAfterSeconds <= 0) ? 300 : idleAfterSeconds;
backoffMs = (backoffMs == null || backoffMs <= 0) ? 60_000L : backoffMs;
@@ -0,0 +1,313 @@
package dev.ltms.fleet.lead;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
import java.io.RandomAccessFile;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.function.LongSupplier;
/**
* Reads how full a lead's own Claude Code context window is, from the transcript Claude Code
* itself writes — never from the lead's pane (fleetd has {@code AgentControl.read} for that, and
* this must not use it: a pane holds terminal text, not the structured usage numbers a transcript
* carries, and scraping it would also race the lead's own rendering).
*
* <p><strong>Why this exists.</strong> A lead auto-compacts when its context fills — on the host
* this was built for, that happened 30 times in one session, discarding roughly 250,000 tokens
* and costing 46 seconds to 3 minutes each time, and fleetd had no way to see it coming. This
* class is the first thing that looks.
*
* <p><strong>The route.</strong> Claude Code appends one JSON object per line to
* {@code <configDir>/projects/<slug>/<sessionId>.jsonl}. {@code <slug>} is an undocumented,
* internal encoding of the working directory — this class never derives it. Instead it lists the
* one-level-deep subdirectories of {@code <configDir>/projects/} and looks for
* {@code <sessionId>.jsonl} by name, so the slug rule can change without breaking this reader.
*
* <ul>
* <li><strong>Live context</strong> is read off the last record in the read window that carries
* a {@code message.usage} object: {@code input_tokens + cache_read_input_tokens +
* cache_creation_input_tokens}. This is what actually fills the window — a plain
* {@code input_tokens} count alone understates it once the conversation has any cached
* prefix, which on a long-lived lead is always.</li>
* <li><strong>Compaction history</strong> is a count of {@code subtype: "compact_boundary"}
* records seen in the same read window — see {@link Reading#compactions()}. It is a count
* within the window this reader actually looked at, not a lifetime total: a session with
* more compactions than fit in {@link #TAIL_BYTES} of transcript will undercount. That
* trade-off is deliberate — see {@link #TAIL_BYTES}.</li>
* </ul>
*
* <p><strong>Three states, not two (OK / HIGH / UNKNOWN).</strong> Every path that cannot
* positively establish the live token count — a missing file, an unreadable one, a peer that
* is not a Claude backend, or every line in the read window failing to parse as JSON — returns
* {@link State#UNKNOWN} with no token number, never a default "0" or "ok" that would read as
* "this lead is fine" when the honest answer is "I could not look".
*
* <p><strong>A torn final line does not mean UNKNOWN.</strong> {@code fleet_list} reads this
* transcript while Claude Code may be mid-write on it, so the last line in the window can be cut
* off mid-flush — that is an ordinary, expected race, not a sign the format has changed. Earlier
* this class treated ANY unparseable last line as UNKNOWN, on the theory that "if the format
* changes, we should see UNKNOWN". That reasoning does not hold: a real format change makes
* <em>every</em> line in the window unparseable, not only the last one written. So a single
* malformed line (most often the final, torn one, but the check is not position-specific) is
* simply skipped rather than treated as fatal, and the reading is built from whatever lines in the
* window did parse. Only when <em>none</em> of them parse — the real format-change signal — does
* this return {@link State#UNKNOWN}, still with no stale number standing in for "I could not
* tell".
*
* <p><strong>Bounded cost.</strong> {@code fleet_list} is polled constantly, so every read is
* capped two ways: {@link #TAIL_BYTES} bounds how much of the transcript is ever read from disk
* (never the whole 52 MB a long-lived transcript reaches on the host this was measured on), and
* {@link #DEFAULT_CACHE_TTL_MILLIS} bounds how often that bounded read actually happens — a burst
* of {@code fleet_list} calls inside one TTL window reads the file once. One instance's cache is
* keyed by {@code (configDir, sessionId)}, so it is safe to share across every lead a single
* {@code fleet_list} call reports on.
*/
public final class LeadContextGauge {
private static final String PROJECTS_DIR = "projects";
private static final ObjectMapper MAPPER = new ObjectMapper();
/**
* How many trailing bytes of a transcript a single read ever pulls off disk. Chosen so one
* read comfortably spans many recent turns — each usage or compact_boundary record is at most
* a few KB — while staying nowhere near the 52 MB a long session's real transcript reaches on
* the host this was built for; reading that whole file on every {@code fleet_list} call is
* exactly the cost this bound exists to avoid. 2 MiB holds on the order of hundreds of recent
* lines even when a turn's tool output is unusually large, which is far more than needed to
* find the most recent usage record and any recent compaction.
*/
static final int TAIL_BYTES = 2 * 1024 * 1024;
/**
* How long a {@link Reading} is served from cache before the file is read again.
* {@code fleet_list} is called constantly (by design — it is the fleet's own status probe), so
* without a TTL a burst of calls would re-read the transcript tail once per call. 5 seconds is
* short enough that a caller watching for a state change never waits long, and long enough that
* a poll loop calling every second or two only touches disk once per window.
*/
static final long DEFAULT_CACHE_TTL_MILLIS = 5_000;
/**
* Live tokens at or above this count report {@link State#HIGH}. On the host this was measured
* on, auto-compaction actually fires around 267,000–270,000 tokens, but the point of a HIGH
* state is to warn before that happens, not at it — 200,000 is the standard Claude context
* window size and a sensible built-in default: no config key is required to pick it, and a
* lead crossing it is already deep enough into its window that a compaction is foreseeable.
*/
static final long HIGH_THRESHOLD_TOKENS = 200_000;
/** The only peer kind this reader understands ({@code Agent.agentType()}'s wire value). */
private static final String CLAUDE_AGENT_TYPE = "claude";
public enum State { OK, HIGH, UNKNOWN }
/**
* @param state {@link State#UNKNOWN} whenever {@code tokens} could not be established
* @param tokens live context tokens, or {@code null} exactly when {@code state} is
* {@link State#UNKNOWN}
* @param compactions {@code compact_boundary} records seen in the read window (see class
* javadoc) — {@code 0} both for "genuinely none seen" and for "unknown",
* since a caller that already sees {@code state: UNKNOWN} has no reason to
* trust this number either way
*/
public record Reading(State state, Long tokens, int compactions) {
/**
* fleetd #609: widened from package-private to public so {@code
* dev.ltms.fleet.msg.LeadHeartbeatLoop.LeadContextSource.none()} (a different package) can
* return the same inert "I could not look" reading the gauge itself uses, without inventing
* a parallel unknown-reading constant. Behaviour of this class is otherwise unchanged.
*/
public static Reading unknown() {
return new Reading(State.UNKNOWN, null, 0);
}
}
private record CacheEntry(Reading reading, long readAtMillis) {
}
private final LongSupplier clock;
private final long ttlMillis;
private final Map<String, CacheEntry> cache = new ConcurrentHashMap<>();
/** Test seam only (package-private) — counts real disk reads, i.e. cache misses. */
private final AtomicInteger diskReads = new AtomicInteger();
public LeadContextGauge() {
this(System::currentTimeMillis, DEFAULT_CACHE_TTL_MILLIS);
}
/** Test seam: an injectable clock and TTL so cache expiry is provable without sleeping. */
LeadContextGauge(LongSupplier clock, long ttlMillis) {
this.clock = clock;
this.ttlMillis = ttlMillis;
}
/** How many times this instance has actually read a transcript off disk — test seam only. */
int diskReadCount() {
return diskReads.get();
}
/**
* @param configDir the lead's {@code CLAUDE_CONFIG_DIR}, or {@code null}/blank to use the
* default {@code <user.home>/.claude} — the right answer for the common case
* where the lead's profile sets no {@code configDir} override
* @param sessionId the lead's own Claude session id ({@code Agent.sessionId()}), or
* {@code null} when herdr has not resolved one yet
* @param agentType the detected peer kind ({@code Agent.agentType()}); anything other than
* {@code "claude"} (including {@code null}, meaning undetected) reports
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
* transcript format
*/
public Reading read(String configDir, String sessionId, String agentType) {
if (sessionId == null || sessionId.isBlank()) {
return Reading.unknown();
}
if (!CLAUDE_AGENT_TYPE.equalsIgnoreCase(agentType)) {
return Reading.unknown();
}
String base = (configDir == null || configDir.isBlank())
? System.getProperty("user.home") + "/.claude"
: configDir;
String cacheKey = base + '\u0000' + sessionId;
long now = clock.getAsLong();
CacheEntry cached = cache.get(cacheKey);
if (cached != null && now - cached.readAtMillis() < ttlMillis) {
return cached.reading();
}
Reading fresh = readUncached(base, sessionId);
cache.put(cacheKey, new CacheEntry(fresh, now));
return fresh;
}
private Reading readUncached(String base, String sessionId) {
diskReads.incrementAndGet();
Path file = findTranscript(base, sessionId);
if (file == null) {
return Reading.unknown();
}
TailRead tail;
try {
tail = tailBytes(file, TAIL_BYTES);
} catch (IOException e) {
return Reading.unknown();
}
return parse(tail);
}
/**
* Finds {@code <sessionId>.jsonl} under {@code <base>/projects/}, one level deep — never by
* deriving the slug directory from a working directory (see class javadoc). Bounded to a
* single {@code list()} of {@code projects/} itself: it never recurses further, so the cost is
* the number of project directories, not the size of any transcript inside them.
*/
private Path findTranscript(String base, String sessionId) {
Path projectsDir = Path.of(base, PROJECTS_DIR);
if (!Files.isDirectory(projectsDir)) {
return null;
}
String filename = sessionId + ".jsonl";
Path direct = projectsDir.resolve(filename);
if (Files.isRegularFile(direct)) {
return direct;
}
try (var children = Files.list(projectsDir)) {
return children.filter(Files::isDirectory)
.map(dir -> dir.resolve(filename))
.filter(Files::isRegularFile)
.findFirst()
.orElse(null);
} catch (IOException e) {
return null;
}
}
/** Package-private (not {@code private}): {@link #tailBytes} is a test seam, see its javadoc. */
record TailRead(byte[] bytes, boolean fromStart) {
}
/**
* Reads at most {@code maxBytes} trailing bytes of {@code file}. Package-private (not
* {@code private}) so a test can assert directly on the returned array's length — "bytes
* actually read", not on any parsed answer — without needing a file anywhere near
* {@link #TAIL_BYTES} in size to prove the cap holds.
*/
static TailRead tailBytes(Path file, int maxBytes) throws IOException {
try (RandomAccessFile raf = new RandomAccessFile(file.toFile(), "r")) {
long length = raf.length();
long start = Math.max(0, length - maxBytes);
raf.seek(start);
byte[] buf = new byte[(int) (length - start)];
raf.readFully(buf);
return new TailRead(buf, start == 0);
}
}
/**
* Parses the tail into a {@link Reading}. The first line is dropped unconditionally whenever
* the tail is not the whole file (it starts mid-line, cut by {@link #TAIL_BYTES} — an expected
* artefact of the bound, not a data problem). Every remaining line is then parsed on a
* best-effort basis: a line that fails to parse (most often the last one, torn by a write this
* read raced — see the "torn final line" section of the class javadoc) is skipped, not fatal.
* Only when none of the remaining lines parse does this report {@link State#UNKNOWN}.
*/
private Reading parse(TailRead tail) {
String text = new String(tail.bytes(), StandardCharsets.UTF_8);
List<String> lines = new ArrayList<>(List.of(text.split("\n", -1)));
if (!lines.isEmpty() && lines.get(lines.size() - 1).isEmpty()) {
lines.remove(lines.size() - 1); // trailing newline leaves a phantom empty last element
}
if (!tail.fromStart() && !lines.isEmpty()) {
lines.remove(0); // first line is a fragment cut by our own tail bound, not real data
}
if (lines.isEmpty()) {
return Reading.unknown();
}
Long tokens = null;
int compactions = 0;
boolean anyLineParsed = false;
for (String line : lines) {
JsonNode node = tryParse(line);
if (node == null) {
// A malformed line — typically the last one, cut mid-flush by a write this read
// raced — is skipped rather than treated as fatal. See the class javadoc's "torn
// final line" section for why: a real format change makes EVERY line unparseable,
// not only this one, and that case is still caught below by anyLineParsed.
continue;
}
anyLineParsed = true;
JsonNode usage = node.path("message").path("usage");
if (usage.isObject()) {
tokens = usage.path("input_tokens").asLong(0)
+ usage.path("cache_read_input_tokens").asLong(0)
+ usage.path("cache_creation_input_tokens").asLong(0);
}
if ("compact_boundary".equals(node.path("subtype").asText(null))) {
compactions++;
}
}
if (!anyLineParsed) {
return Reading.unknown();
}
if (tokens == null) {
return new Reading(State.UNKNOWN, null, compactions);
}
State state = tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK;
return new Reading(state, tokens, compactions);
}
private JsonNode tryParse(String line) {
try {
return MAPPER.readTree(line);
} catch (IOException e) {
return null;
}
}
}
@@ -10,6 +10,8 @@ import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
@@ -25,9 +27,11 @@ import java.util.function.Supplier;
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
* clears the lead's own pane and bootstraps a fresh session against that file.
*
* <p>This is the executor only. Nothing in this ticket wires an MCP tool onto {@link #open}/
* {@link #confirm}/{@link #cancel} — that is a separate, later unit; until it lands, nothing calls
* this class at all.
* <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code
* dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link
* #cancel}, and {@link #status} from a tool call — wired in fleetd #480 Unit C. <strong>An earlier
* version of this paragraph said nothing called this class at all; that stopped being true once
* that unit landed, and this correction exists so the javadoc does not go on claiming it.</strong>
*
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
@@ -181,6 +185,106 @@ public final class LeadRollover {
}
}
/**
* How many tokens {@link #outcomes} remembers before it starts evicting the oldest — bounded
* so a long-running daemon never grows this map without limit. Chosen generously rather than
* tightly: production rolls are rare (this class's own ticket found exactly ONE completed roll
* ever logged on this host), and each entry is a handful of short strings, so even a full cap
* costs a few tens of kilobytes — nowhere near a reason to make it configurable. 200 entries
* comfortably outlasts any operator's own memory of "did that roll I asked for actually
* happen", which is the whole reason {@link #status} exists.
*
* <p><strong>This cap counts {@link RollState#IN_PROGRESS} entries exactly the same as
* finished ones.</strong> There is only the one bounded map: {@link #confirm} writes an {@link
* RollState#IN_PROGRESS} entry into {@link #outcomes} at hand-off, and the deferred
* continuation later overwrites that SAME key with a terminal state — it never inserts a
* second entry. An approved roll therefore occupies one slot in this map for its entire
* lifetime, from the moment {@link #confirm} hands off, not only once it finishes; a
* confirmed-but-not-yet-finished roll counts against the cap exactly like a finished one. The
* alternative (a separate, uncapped in-flight map) would let a burst of confirmed-but-stuck
* rolls grow without bound — the exact failure this cap exists to prevent — so it was rejected.
*/
static final int OUTCOME_HISTORY_CAP = 200;
/**
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
* three terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* that has been approved but has not finished yet, and two answers for a token that names no
* active work at all: still pending confirmation, or nothing known about this token at all.
*/
public enum RollState {
/**
* {@code token} is still open: either {@link #open} was called and {@link #confirm} has not
* been (or not successfully) yet, or a {@link #confirm} call failed one of its gate checks
* and left the token pending for a retry — see {@link #confirm}'s javadoc ("token stays
* pending"). Indistinguishable from a genuinely fresh request; a caller wanting to know
* WHICH gate most recently refused should read the {@link RollDecision} that {@link
* #confirm} itself returned, not this status. <strong>Never the state of an APPROVED
* roll</strong> — see {@link #IN_PROGRESS}, which {@link #confirm} records at the moment it
* hands off, before this token is even removed from the pending set.
*/
PENDING,
/**
* {@link #confirm} approved this roll and handed it to the deferred continuation, which has
* not finished yet. Recorded by {@link #confirm} itself, at hand-off — <strong>before</strong>
* {@code token} is removed from the pending set — so there is never a gap in which {@link
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
* that is, in fact, actively running. This is not sticky: the deferred continuation
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into
* {@link #FAILED} instead of leaving this entry stuck forever.
*/
IN_PROGRESS,
/**
* {@link #confirm} was approved and the deferred continuation completed the entire roll:
* the calling lead's turn settled, {@code /clear} was sent and settled, and {@code
* bootstrapText} was sent.
*/
ROLLED,
/**
* {@link #confirm} was approved, but the calling lead's own turn never reached a boundary
* (IDLE or DONE) within {@code turnSettleSeconds} — no {@code /clear} was ever sent, at
* all. This is the branch the fleetd #480 correction exists to make safe, and the one this
* status exists to make VISIBLE: before this, a lead that hit this case had no way to find
* out, and would carry on believing it was about to be replaced. See this class's javadoc.
*/
TURN_NEVER_SETTLED,
/**
* {@link #confirm} was approved and {@code /clear} was sent, but the pane never re-settled
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent.
*/
CLEAR_NEVER_SETTLED,
/**
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it.
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS}
* forever, because the production {@code continuationRunner} is a bare virtual thread with
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a
* terminal outcome. {@code detail} names the exception, so a reader has something to act on
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck
* lead must {@link #open} a fresh request.
*/
FAILED,
/**
* {@code token} names nothing this instance currently knows about: never issued by {@link
* #open}, dropped by {@link #cancel}, or aged out of {@link #outcomes}'s bounded history.
* These three causes are not distinguished — all of them mean "there is nothing to tell
* you", which is the entire content of a clean answer here.
*/
UNKNOWN
}
/**
* The answer {@link #status} gives for one token: a {@link RollState} and a human-readable
* {@code detail}. For {@link RollState#TURN_NEVER_SETTLED}, {@code detail} names {@code
* turnSettleSeconds} and its configured value explicitly, so a reader who sees this knows what
* to raise.
*/
public record RollStatus(RollState state, String detail) {}
private final AgentControl agents;
private final Supplier<FleetConfig.LeadRollover> configSupplier;
/**
@@ -201,6 +305,23 @@ public final class LeadRollover {
*/
private final Consumer<Runnable> continuationRunner;
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
/**
* Finished tokens → what actually happened, for {@link #status}. Bounded by {@link
* #OUTCOME_HISTORY_CAP}, oldest evicted first ({@code removeEldestEntry} on an insertion-order
* {@link LinkedHashMap}). Wrapped in {@link Collections#synchronizedMap} because entries are
* written from whatever thread {@code continuationRunner} runs the roll on (a fresh virtual
* thread in production, the calling test thread under {@code Runnable::run}) and read from
* whatever thread calls {@link #status} (the MCP handler thread) — a plain {@code
* LinkedHashMap} is not safe for that, and {@code removeEldestEntry} additionally requires
* external synchronization even for a thread-safe map that merely wraps it.
*/
private final Map<String, RollStatus> outcomes = Collections.synchronizedMap(
new LinkedHashMap<>(16, 0.75f, false) {
@Override
protected boolean removeEldestEntry(Map.Entry<String, RollStatus> eldest) {
return size() > OUTCOME_HISTORY_CAP;
}
});
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
@@ -366,6 +487,15 @@ public final class LeadRollover {
return docCheck;
}
// Record IN_PROGRESS BEFORE removing from `pending` — see RollState#IN_PROGRESS and
// OUTCOME_HISTORY_CAP's javadoc. This ordering means `token` is written into `outcomes`
// while it is STILL present in `pending`; status() checks `outcomes` first (see that
// method), so it reports IN_PROGRESS immediately, not the brief-but-real gap a
// remove-then-put ordering would leave in which the token is in neither map.
outcomes.put(token, new RollStatus(RollState.IN_PROGRESS,
"confirm() approved this roll and handed it to the deferred continuation; it has "
+ "not finished yet — still waiting for the calling turn to settle, for "
+ "/clear to be sent and settle, or for bootstrapText to be sent"));
pending.remove(token);
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
token, callerTerminal);
@@ -378,8 +508,42 @@ public final class LeadRollover {
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
* four-step order. There is no result to return to by this point, so every outcome is logged
* only.
*
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} →
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual
* thread with no uncaught-exception handler (see this class's public constructor). Before this
* fix, either throw killed the continuation thread silently, leaving the {@link
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch}
* below is scoped to the method body rather than to each {@code send} call individually, so it
* also covers anything else added to this continuation later, not just today's two call sites —
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
* branches below) rather than inside the helpers that detect them.</p>
*
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
* broader {@link Exception} or {@link Throwable}, which would also swallow something like an
* {@link OutOfMemoryError} this continuation has no business handling.</p>
*/
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
try {
runRolloverUnguarded(p, cfg);
} catch (RuntimeException e) {
log.warn("lead-rollover: continuation for token={} lead={} threw {} — the roll is dead; "
+ "no further step in this continuation will run",
p.token(), p.leadTerminal(), e.toString(), e);
outcomes.put(p.token(), new RollStatus(RollState.FAILED,
"the roll's continuation threw " + e.toString() + " — the roll is dead and will "
+ "not retry itself; check the daemon log for the stack trace, then open() "
+ "a fresh rollover request"));
}
}
/** The actual body of {@link #runRollover}, unwrapped — see that method's javadoc for the catch. */
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
String lead = p.leadTerminal();
long rollStartMillis = nowMillis.getAsLong();
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
@@ -393,6 +557,11 @@ public final class LeadRollover {
+ "turn is still live and clearing it now would destroy live context "
+ "(token={}, configured={}s elapsed={}ms)",
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED,
"the calling lead's own turn never reached a boundary (IDLE or DONE) within "
+ "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed="
+ turnResult.elapsedMillis() + "ms) — no /clear was ever sent. If this "
+ "keeps happening, raise turnSettleSeconds in fleetd.yaml"));
return;
}
@@ -411,11 +580,17 @@ public final class LeadRollover {
+ "elapsed={}ms nudges={})",
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
clearResult.nudges());
outcomes.put(p.token(), new RollStatus(RollState.CLEAR_NEVER_SETTLED,
"/clear was sent, but the pane never re-settled within clearSettleSeconds="
+ cfg.clearSettleSeconds() + "s (measured elapsed=" + clearResult.elapsedMillis()
+ "ms, nudges=" + clearResult.nudges() + ") — bootstrapText was never sent"));
return;
}
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
outcomes.put(p.token(), new RollStatus(RollState.ROLLED,
"rolled successfully in " + rollElapsedMillis + "ms"));
}
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
@@ -423,6 +598,47 @@ public final class LeadRollover {
return pending.remove(token) != null;
}
/**
* Read-only: what is currently known about {@code token}. <strong>Never sends anything, never
* schedules, cancels, or retries a roll</strong> — a caller may poll this as often as it likes
* with no side effect at all, which is exactly why it exists: every failure past {@link
* #confirm} used to be a {@code log.warn} a lead can never read (see this class's javadoc), and
* this is the only route back.
*
* @param token the token {@link #open} returned; {@code null} or blank is a clean {@link
* RollState#UNKNOWN}, never a {@link NullPointerException} — {@link #pending} is a
* {@link ConcurrentHashMap}, which throws on a {@code null} key lookup, so this
* short-circuits before ever reaching it
* @return {@link RollState#IN_PROGRESS} for an approved roll whose continuation has not
* finished yet, or a terminal state once it has (both read from {@link #outcomes} —
* checked FIRST, see below); {@link RollState#PENDING} while {@code token} is still
* open and has not yet been approved (including one left pending by a {@link #confirm}
* gate refusal — see that method's javadoc); or {@link RollState#UNKNOWN} for a token
* never issued, cancelled, or aged out of the bounded history
*/
public RollStatus status(String token) {
if (token == null || token.isBlank()) {
return new RollStatus(RollState.UNKNOWN, "no token given");
}
// `outcomes` is checked BEFORE `pending`, deliberately: `confirm` writes an IN_PROGRESS
// entry into `outcomes` before it removes `token` from `pending` (see `confirm`'s own
// comment at that call site), so for the brief window where a token is present in BOTH
// maps, this order reports the more accurate answer (IN_PROGRESS, already approved) rather
// than the stale one (PENDING, not yet approved) a pending-first check would give.
RollStatus recorded = outcomes.get(token);
if (recorded != null) {
return recorded;
}
if (pending.containsKey(token)) {
return new RollStatus(RollState.PENDING, "open() has been called for this token and "
+ "it has not yet been confirmed — or a confirm() gate check failed and left it "
+ "pending, so the same token may be retried once the problem is fixed");
}
return new RollStatus(RollState.UNKNOWN, "token names no pending or finished rollover "
+ "request known to this instance — never issued, cancelled, or aged out of the "
+ "bounded history (cap=" + OUTCOME_HISTORY_CAP + ")");
}
/**
* The three handover-file checks, in order: exists, not empty, fresh (modified after
* {@link #open}'s timestamp and not older than {@code maxDocAgeSeconds}). Stats {@code
@@ -12,6 +12,7 @@ import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.msg.LeadChannel;
import dev.ltms.fleet.msg.LeadMessage;
@@ -115,6 +116,8 @@ public final class FleetMcp {
private final OutageSource outage;
/** fleetd #176: SEPARATE from both of the above — see {@link LeadSeatSource}'s doc. */
private final LeadSeatSource leadSeats;
/** fleetd #602 gauge-wiring: see {@link LeadConfigDirSource}. */
private final LeadConfigDirSource leadConfigDirs;
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
private final LeadChannel leadChannel;
/** fleetd #361: {@code coordinator.peers} — see {@link CoordinationSource}. Empty when unset. */
@@ -126,6 +129,17 @@ public final class FleetMcp {
* clean {@code NOT_CONFIGURED} refusal rather than throwing. See {@link #handover}.
*/
private final LeadRollover leadRollover;
/**
* "Lead context gauge": how full each lead's own Claude Code context window is, reported on
* {@code fleet_list}'s {@code leads} rows (see {@link #contextView}). Built unconditionally, in
* the field initializer rather than a constructor parameter — this is not a togglable feature
* with an on/off config knob the way {@link OutageSource}/{@link LeadSeatSource} are: it needs
* no config at all (see {@link LeadContextGauge}'s own javadoc for the built-in defaults), so
* there is no "off" value to thread through every existing constructor call site. One instance
* per daemon so its read cache (keyed by session, TTL'd) is actually shared across
* {@code fleet_list} calls rather than rebuilt — and therefore useless — on every call.
*/
private final LeadContextGauge leadContextGauge = new LeadContextGauge();
/** Capacity facts used by {@code fleet_list}; production must supply the placement live count. */
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
@@ -246,6 +260,29 @@ public final class FleetMcp {
public static LeadSeatSource none() { return new LeadSeatSource(_ -> 0); }
}
/**
* fleetd #602 gauge-wiring: a lead name's configured {@code CLAUDE_CONFIG_DIR} override, fed to
* {@link LeadContextGauge#read} so {@code fleet_list}'s {@code context} row reads the transcript
* directory the lead's own profile actually writes to, not always the built-in
* {@code <user.home>/.claude} default.
*
* <p>The same idiom as {@link LeadSeatSource} — a {@code FleetMcp} constructor field, not a
* lookup {@code contextView} performs itself, because {@code FleetMcp} holds no
* {@link dev.ltms.fleet.config.FleetConfig} and {@code contextView} is {@code static}. See
* {@code Fleetd.leadConfigDirLookup} for the derivation: the same
* {@code fleet.leaders.<name>.profile} link {@link LeadSeatSource} already follows, resolved to
* that profile's own {@code configDir:}.
*
* @param configDirFor lead name → {@code configDir}, or {@code null} when the lead's entry names
* no profile, or that profile sets no {@code configDir} override — either
* way {@link LeadContextGauge} then falls back to its own built-in default,
* exactly as before this ticket
*/
public record LeadConfigDirSource(Function<String, String> configDirFor) {
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in default {@code configDir}. */
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null); }
}
/**
* fleetd #361: peer-visibility facts for {@code fleet_list}'s {@code coordinator} row — this
* daemon's own {@link LeadChannel} (for its self mailbox state and held messages) plus the
@@ -346,7 +383,7 @@ public final class FleetMcp {
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, authorizationMode, metrics,
capacity, healthCoverage, LoopHealthSource.none(), quarantine, leadChannel, outage, leadSeats,
peers, leadRollover);
LeadConfigDirSource.none(), peers, leadRollover);
}
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
@@ -354,7 +391,8 @@ public final class FleetMcp {
CallerResolver callers, AuthorizationMode authorizationMode, Metrics metrics,
CapacitySource capacity, HealthCoverageSource healthCoverage, LoopHealthSource loopHealth,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
LeadSeatSource leadSeats, LeadConfigDirSource leadConfigDirs, List<String> peers,
LeadRollover leadRollover) {
Objects.requireNonNull(callers, "callers");
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
== AuthorizationMode.ENFORCED;
@@ -364,6 +402,7 @@ public final class FleetMcp {
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.outage = Objects.requireNonNull(outage, "outage");
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
this.leadConfigDirs = Objects.requireNonNull(leadConfigDirs, "leadConfigDirs");
this.healthCoverage = healthCoverage;
this.loopHealth = Objects.requireNonNull(loopHealth, "loopHealth");
this.leadRollover = leadRollover;
@@ -498,7 +537,7 @@ public final class FleetMcp {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_list", Map.of()), null);
if (denied != null) return denied;
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
leadSeats, callers.leads(),
leadSeats, leadContextGauge, leadConfigDirs, callers.leads(),
callerTerminal(exchange),
new CoordinationSource(leadChannel, peers),
coordinatorVisibleTo(principal(exchange)));
@@ -1248,14 +1287,16 @@ public final class FleetMcp {
Map<String, Object> args) {
String action = str(args, "action");
if (isBlank(action)) {
return error("action is required: \"open\", \"confirm\" or \"cancel\"");
return error("action is required: \"open\", \"confirm\", \"cancel\" or \"status\"");
}
return switch (action) {
case "open" -> handoverOpen(leadRollover, callerTerminal, str(args, "reason"));
case "confirm" -> handoverConfirm(leadRollover, callerTerminal, str(args, "token"),
truthy(args, "operatorConfirmed"));
case "cancel" -> handoverCancel(leadRollover, str(args, "token"));
default -> error("unknown action \"" + action + "\" — must be \"open\", \"confirm\" or \"cancel\"");
case "status" -> handoverStatus(leadRollover, str(args, "token"));
default -> error("unknown action \"" + action
+ "\" — must be \"open\", \"confirm\", \"cancel\" or \"status\"");
};
}
@@ -1325,6 +1366,29 @@ public final class FleetMcp {
return text(json(m));
}
/**
* {@code action: "status"}. Read-only — see {@link LeadRollover#status}: never schedules,
* cancels, or retries anything, and it is the only way for a lead to find out what happened to
* a token past {@code confirm()}, since every outcome after that point is otherwise logged only
* (see {@link LeadRollover}'s class javadoc).
*/
private static McpSchema.CallToolResult handoverStatus(LeadRollover leadRollover, String token) {
if (leadRollover == null) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("state", "NOT_CONFIGURED");
m.put("detail", "leadRollover: is not configured");
return text(json(m));
}
if (isBlank(token)) {
return error("token is required for action \"status\"");
}
LeadRollover.RollStatus s = leadRollover.status(token);
Map<String, Object> m = new LinkedHashMap<>();
m.put("state", s.state().name());
m.put("detail", s.detail());
return text(json(m));
}
/** The one shared {@code NOT_CONFIGURED} refusal shape for {@code open}/{@code confirm}. */
private static McpSchema.CallToolResult notConfigured() {
return refusalJson(false, "NOT_CONFIGURED", "leadRollover: is not configured");
@@ -1601,7 +1665,7 @@ public final class FleetMcp {
Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
OutageSource.none(), LeadSeatSource.none(), leads, selfTerm, coordination, false);
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1610,7 +1674,7 @@ public final class FleetMcp {
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), leads, selfTerm, CoordinationSource.none(), false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, CoordinationSource.none(), false);
}
/**
@@ -1627,7 +1691,7 @@ public final class FleetMcp {
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
LeadSeatSource.none(), leads, selfTerm, coordination, false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1636,7 +1700,7 @@ public final class FleetMcp {
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), leads, selfTerm, coordination, false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/**
@@ -1658,7 +1722,7 @@ public final class FleetMcp {
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
leadSeats, leads, selfTerm, coordination, false);
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/**
@@ -1682,14 +1746,24 @@ public final class FleetMcp {
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
outage, leadSeats, leads, selfTerm, coordination, callerIsPrimary);
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, callerIsPrimary);
}
/**
* The canonical implementation. {@code contextGauge} is the "lead context gauge" (see
* {@link LeadContextGauge}) — every wrapper overload above passes a freshly constructed one,
* which is correct for them (none of them exercise repeated calls where a shared cache would
* matter); the one caller that matters for caching, {@code fleet_list}'s MCP handler, passes
* its own single long-lived instance instead (see {@code FleetMcp}'s {@code leadContextGauge}
* field).
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
LoopHealthSource loopHealth,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
LeadSeatSource leadSeats, LeadContextGauge contextGauge,
LeadConfigDirSource leadConfigDirs,
Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary) {
try {
Map<String, Agent> live = workers.list().stream()
@@ -1698,7 +1772,8 @@ public final class FleetMcp {
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
leadConfigDirs))
.toList();
// fleetd #209: this is the caller-driven fleet_list read that actually reports
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
@@ -1993,15 +2068,25 @@ public final class FleetMcp {
}
/**
* One lead's row: its address, its name, and whether it can be reached right now.
* One lead's row: its address, its name, whether it can be reached right now, and how full its
* own Claude Code context window is.
*
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
* see is a lead a {@code fleet_send} cannot be typed into. It is reported rather than hidden,
* because a peer that has gone unreachable is exactly what the sender needs to know.
*
* <p>{@code context} is the "lead context gauge" (fleetd's context-usage visibility ticket):
* {@code {state: "ok"|"high"|"unknown", tokens?: number, compactions: number}}, read from the
* lead's transcript — see {@link LeadContextGauge}. Kept small on purpose (a token count, a
* state, and a compaction count) rather than echoing the whole reading history: this is a
* roster row a caller glances at, not a diagnostics dump. {@code tokens} is present only when
* {@code state} is not {@code "unknown"} — never a stale or default number standing in for "I
* could not tell".
*/
private static Map<String, Object> leadView(String terminal, String name, Agent live,
String selfTerm) {
String selfTerm, LeadContextGauge contextGauge,
LeadConfigDirSource leadConfigDirs) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", terminal);
m.put("name", name);
@@ -2010,9 +2095,32 @@ public final class FleetMcp {
if (terminal.equals(selfTerm)) {
m.put("self", true);
}
String configDir = leadConfigDirs.configDirFor().apply(name);
m.put("context", contextView(contextGauge, live, configDir));
return m;
}
/**
* Reads the lead context gauge for one lead. {@code configDir} is this lead's configured
* {@code CLAUDE_CONFIG_DIR} override (see {@link LeadConfigDirSource}), derived from
* {@code fleet.leaders.<name>.profile} → that profile's own {@code configDir:} — or {@code null}
* when the lead's entry names no profile, or that profile sets no override, in which case
* {@link LeadContextGauge#read} falls back to its own built-in default
* ({@code <user.home>/.claude}).
*/
private static Map<String, Object> contextView(LeadContextGauge contextGauge, Agent live, String configDir) {
String sessionId = live == null ? null : live.sessionId();
String agentType = live == null ? null : live.agentType();
LeadContextGauge.Reading reading = contextGauge.read(configDir, sessionId, agentType);
Map<String, Object> c = new LinkedHashMap<>();
c.put("state", reading.state().name().toLowerCase());
if (reading.tokens() != null) {
c.put("tokens", reading.tokens());
}
c.put("compactions", reading.compactions());
return c;
}
/** {@code fleet_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
@@ -2265,20 +2373,26 @@ public final class FleetMcp {
return tool(FleetTool.HANDOVER.wireName(),
"Replace your OWN lead session once its context is full: write a handover file, "
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
+ "session against it. Three actions: 'open' (requests a token and the "
+ "session against it. Four actions: 'open' (requests a token and the "
+ "handoverPath you must write the handover file to before confirming), "
+ "'confirm' (validates every gate and — only if every one passes — schedules "
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
+ "own turn ends), and 'cancel' (drops a pending request without rolling). "
+ "Primary-only. There is deliberately no terminal/session/leadTerminal "
+ "parameter: the pane to roll is always resolved from YOUR OWN connection, "
+ "never a value you pass, so you can only ever roll yourself — never another "
+ "lead. Requires leadRollover: to be configured; when it is not, every action "
+ "returns a clean refusal naming NOT_CONFIGURED instead of failing.",
+ "own turn ends), 'cancel' (drops a pending request without rolling), and "
+ "'status' (read-only: what happened to a token after 'confirm' — still "
+ "running (approved but not finished yet), the roll completed, the calling "
+ "turn never settled within turnSettleSeconds so no /clear was ever sent, or "
+ "/clear itself never settled so bootstrapText was never sent; never "
+ "schedules, cancels or retries anything). Primary-only. "
+ "There is deliberately no terminal/session/leadTerminal parameter: the pane "
+ "to roll is always resolved from YOUR OWN connection, never a value you "
+ "pass, so you can only ever roll yourself — never another lead. Requires "
+ "leadRollover: to be configured; when it is not, every action returns a "
+ "clean refusal naming NOT_CONFIGURED instead of failing.",
objectSchema(Map.of(
"action", stringProp("\"open\", \"confirm\" or \"cancel\""),
"action", stringProp("\"open\", \"confirm\", \"cancel\" or \"status\""),
"reason", stringProp("Free-text audit note for \"open\" (optional, logged only)"),
"token", stringProp("The token \"open\" returned — required for \"confirm\" and \"cancel\""),
"token", stringProp("The token \"open\" returned — required for \"confirm\", "
+ "\"cancel\" and \"status\""),
"operatorConfirmed", Map.of("type", "boolean",
"description", "For \"confirm\": your answer to \"has the human operator "
+ "confirmed this wipe\" (default false; only consulted when "
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -13,6 +14,7 @@ import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
@@ -44,6 +46,15 @@ import java.util.function.Supplier;
* loop stands down, so two competing injections never start two turns in the same pane
* (constraint 6).</li>
* </ol>
*
* <p><b>fleetd #609 — context-high notice.</b> Optionally ({@code contextHighNudge}, opt-in like the
* loop itself), a tick that finds the lead's own {@link LeadContextGauge} reading at {@link
* LeadContextGauge.State#HIGH} appends a text notice to whatever nudge it sends, telling the lead to
* consider {@code fleet_handover}. This is text only — it never rolls a pane itself. It fires once per
* HIGH stretch (a latch, cleared only by a later {@code OK} reading — {@code UNKNOWN} neither sets nor
* clears it, since "I could not look" must not be read as "it got better"), and it never spends the
* quiet-nudge budget: an idle, quiet, HIGH-context lead is exactly the case {@link Action#QUIET_DONE}
* would otherwise swallow, and it is the one case most worth interrupting the quiet cap for.
*/
public final class LeadHeartbeatLoop {
@@ -63,11 +74,15 @@ public final class LeadHeartbeatLoop {
private final long backoffMs;
private final int quietNudgeCap;
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
private final LeadContextSource contextSource; // fleetd #609
private final boolean contextHighNudge; // fleetd #609: opt-in, like the loop itself
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
private long idleSinceNanos = NOT_IDLE;
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
private int quietCount = 0;
/** fleetd #609: latched "already told this lead about this HIGH stretch". Single scheduler thread only. */
private boolean contextNotified = false;
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
@@ -75,7 +90,7 @@ public final class LeadHeartbeatLoop {
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, null);
idleAfterNanos, backoffMs, quietNudgeCap, null, LeadContextSource.none(), false);
}
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
@@ -83,6 +98,20 @@ public final class LeadHeartbeatLoop {
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, metrics, LeadContextSource.none(), false);
}
/**
* fleetd #609: as above, plus the lead's own context source and whether a HIGH reading should
* append a hand-over notice to the loop's nudge. Pass {@link LeadContextSource#none()} and
* {@code false} to keep the pre-#609 behaviour exactly (both existing public constructors do).
*/
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
LeadContextSource contextSource, boolean contextHighNudge) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
@@ -94,6 +123,19 @@ public final class LeadHeartbeatLoop {
this.backoffMs = backoffMs;
this.quietNudgeCap = quietNudgeCap;
this.metrics = metrics;
this.contextSource = contextSource;
this.contextHighNudge = contextHighNudge;
}
/**
* fleetd #609: one lead's own context reading, keyed by its terminal id — the same injected-source
* idiom {@code FleetMcp.LeadSeatSource}/{@code FleetMcp.LeadConfigDirSource} already use.
*/
public record LeadContextSource(Function<String, LeadContextGauge.Reading> readingFor) {
/** Inert source — every lead reads UNKNOWN, so the context notice can never fire. */
public static LeadContextSource none() {
return new LeadContextSource(_ -> LeadContextGauge.Reading.unknown());
}
}
/**
@@ -126,8 +168,11 @@ public final class LeadHeartbeatLoop {
STAND_DOWN
}
/** The outcome of one decision: the action plus the state to persist for the next tick. */
record Decision(Action action, Long idleSinceNanos, int quietCount) {}
/**
* The outcome of one decision: the action, the state to persist for the next tick, and (fleetd
* #609) whether the lead has now been told about the current HIGH context stretch.
*/
record Decision(Action action, Long idleSinceNanos, int quietCount, boolean contextNotified) {}
/**
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
@@ -142,34 +187,47 @@ public final class LeadHeartbeatLoop {
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
* @param leadKnown whether a lead terminal is known to nudge at all
* @param fleet a snapshot of the pending fleet state (constraint 5)
* @param context fleetd #609: the lead's own {@link LeadContextGauge} reading for this tick
* @param contextNotified fleetd #609: whether the lead has already been told about the current HIGH
* stretch — a latch, carried forward by {@link #applyDecision}
* @return the action to take and the state to persist
*/
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
boolean pushLoopActive, boolean leadKnown, FleetState fleet) {
boolean pushLoopActive, boolean leadKnown, FleetState fleet,
LeadContextGauge.State context, boolean contextNotified) {
// fleetd #609: re-arm the latch only on a positive OK reading. UNKNOWN means "I could not
// look", not "it got better" — re-arming on UNKNOWN would let a flapping gauge (a transcript
// read that misses one tick) nudge a full lead again on every recovery, defeating the "once
// per HIGH stretch" promise. Computed once, up front, so every gate below carries it forward
// unchanged unless it is the gate that actually discharges it.
boolean latch = context == LeadContextGauge.State.OK ? false : contextNotified;
boolean contextHigh = contextHighNudge && context == LeadContextGauge.State.HIGH;
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
// competing prompt into the same pane would start a second turn — racing loops multiply
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
// quiet counter), because the reply that drove it is exactly the kind of new state that
// should reset the cap.
// should reset the cap. The context latch is untouched: standing down must not spend the
// one notice this stretch gets.
if (pushLoopActive) {
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0);
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0, latch);
}
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
// status (read failure, or the agent is gone) is safest treated the same way — never inject
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
if (status == null || !status.injectable()) {
return new Decision(Action.LEAD_BUSY, null, 0);
return new Decision(Action.LEAD_BUSY, null, 0, latch);
}
if (idleSinceNanos == null) {
// The lead just became injectable — record the start of an idle stretch and wait out the
// debounce quiet period before ever nudging (constraint 3).
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount);
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount, latch);
}
if (nowNanos - idleSinceNanos < idleAfterNanos) {
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
// and must not be re-prompted into every natural pause.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
}
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
// gating facts decide whether and how:
@@ -177,23 +235,33 @@ public final class LeadHeartbeatLoop {
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
// quiet period.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
}
if (fleet.hasPending()) {
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it.
return new Decision(Action.INJECT, idleSinceNanos, 0);
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it. The
// nudge text carries the context notice too when contextHigh — see injectNudge/contextNotice
// — so this route discharges the same duty and must set the latch.
return new Decision(Action.INJECT, idleSinceNanos, 0, latch || contextHigh);
}
if (contextHigh && !latch) {
// fleetd #609: the lead is idle, its context is full, and nothing is pending. This is the
// one case the quiet cap would otherwise swallow, and it is exactly when the lead most
// needs to hear it. Fire once per HIGH stretch, and do NOT spend the quiet budget on it:
// this is an event notice, not a "are you still there" nudge.
return new Decision(Action.INJECT, idleSinceNanos, quietCount, true);
}
if (quietCount < quietNudgeCap) {
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
// exactly that nothing is waiting so it can choose to stand down rather than hunt
// (constraint 5). Count it toward the consecutive-quiet cap.
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1);
// (constraint 5). Count it toward the consecutive-quiet cap. This nudge also carries the
// context notice when contextHigh (already latched above, or being latched now).
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1, latch || contextHigh);
}
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
// re-arms it — QUIET_DONE stops injection, not observation.
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount);
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount, latch);
}
// --- loop ----------------------------------------------------------------------------------
@@ -208,57 +276,152 @@ public final class LeadHeartbeatLoop {
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
}
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. */
private void tick() {
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. Package-private
* (mirroring {@link ReplyPushLoop#tick(String)}) so tests can drive it directly with a fake clock and a
* fake {@link AgentControl} instead of racing the scheduler thread. */
void tick() {
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
FleetState fleet = snapshot(inbox, roster);
AgentStatus status = AgentStatus.UNKNOWN;
LeadContextGauge.Reading reading = LeadContextGauge.Reading.unknown();
if (leadKnown) {
String leadTerminal = primaryRegistry.primaryTerminal().orElseThrow();
try {
status = agents.status(primaryRegistry.primaryTerminal().orElseThrow());
status = agents.status(leadTerminal);
} catch (RuntimeException e) {
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
// and never injects into a state it cannot read. Retry on the next backoff.
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
}
// fleetd #609: read the lead's own context regardless of status — decide() still gates on
// status first (constraint 2), so this is harmless work on a WORKING lead and lets the
// latch state stay accurate for whenever the lead does go idle.
reading = contextSource.readingFor().apply(leadTerminal);
}
Decision d = decide(clock.getAsLong(),
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
quietCount, status, pushLoop.isActive(), leadKnown, fleet);
quietCount, status, pushLoop.isActive(), leadKnown, fleet,
reading.state(), contextNotified);
applyDecision(d);
switch (d.action()) {
case INJECT -> injectNudge(fleet);
case QUIET_DONE -> countNudge("exhausted");
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> { /* nothing to inject, nothing to count */ }
case INJECT -> injectNudge(d, fleet, reading);
case QUIET_DONE -> {
countNudge("exhausted");
contextNotified = d.contextNotified();
}
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> contextNotified = d.contextNotified();
}
scheduleNext();
}
/** Persist the state a decision returned, so the next tick starts from it. */
/**
* Persist the idle/quiet state a decision returned, so the next tick starts from it.
*
* <p>fleetd #609 review: the context latch ({@link #contextNotified}) is deliberately <em>not</em>
* set here any more. Setting it from the decision unconditionally — before {@link #injectNudge} even
* tries to send — is exactly the review's blocker: a decision to notify is not the same fact as "the
* notice reached the pane". Every branch of {@link #tick} now assigns {@link #contextNotified} itself,
* once it knows whether a send happened and whether it carried the notice (see {@link #injectNudge}).
*/
private void applyDecision(Decision d) {
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
quietCount = d.quietCount();
}
/** Send the nudge to the known lead. */
private void injectNudge(FleetState fleet) {
/**
* Send the nudge to the known lead, with the fleetd #609 context notice appended when it applies, and
* persist the context latch based on what actually happened this tick — not merely what {@code d}
* chose to attempt.
*/
private void injectNudge(Decision d, FleetState fleet, LeadContextGauge.Reading reading) {
// fleetd #609 review: build the notice from the latch as it stood BEFORE this tick's decision —
// d.contextNotified() is the value to persist once delivery is confirmed, not the value the text
// itself should be built from. Otherwise a HIGH stretch that is still latched would never see the
// notice at all, defeating the very check this fixes.
String notice = contextNotice(contextHighNudge, reading, contextNotified);
var lead = primaryRegistry.primaryTerminal();
if (lead.isEmpty()) {
return; // the lead disappeared between the decision and the injection
}
String leadTerminal = lead.get();
boolean sent = lead.isPresent() && trySend(lead.get(), fleet.nudgeText() + notice, notice);
// The latch becomes true only when all three hold: decide() chose to notify, a notice was
// actually included in the text, and the send reached the pane without throwing. Whenever no
// notice was attempted (disabled, not HIGH, or already latched), nothing was promised to the lead
// this tick, so apply the decision's own carried-forward value unconditionally — that is how the
// OK-only re-arm rule and STAND_DOWN's "don't burn the notice" rule keep working through this path
// too. A lead that disappeared between the decision and the send (lead.isEmpty()) is treated the
// same as a failed send: nothing reached the pane, so the latch must not be set.
contextNotified = notice.isEmpty() ? d.contextNotified() : sent;
}
/**
* Attempt one herdr send and count its outcome. Returns whether {@code agents.send} returned without
* throwing — the caller ({@link #injectNudge}) needs this to decide whether the fleetd #609 context
* latch may be persisted as set.
*/
private boolean trySend(String leadTerminal, String text, String notice) {
try {
agents.send(leadTerminal, fleet.nudgeText());
agents.send(leadTerminal, text);
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
leadTerminal, quietCount);
countNudge("sent");
// fleetd #609: a nudge that carries the context notice is counted under its own outcome so
// it is visible in /metrics — one count per nudge either way, never two.
countNudge(notice.isEmpty() ? "sent" : "sent_context");
return true;
} catch (RuntimeException e) {
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
countNudge("failed");
return false;
}
}
/**
* fleetd #609: the text appended to a nudge when the lead's own context is full — {@code ""}
* whenever the notice does not apply, so callers can unconditionally append this without an extra
* branch. Wording stays plain (CEFR B1) and honest that only the operator approves a roll — this
* loop only ever prints text, it never calls {@code fleet_handover} itself.
*
* @param enabled the {@code leadHeartbeat.contextHighNudge} config flag
* @param reading the lead's current {@link LeadContextGauge} reading
* @return the notice text (starting with a leading space, to append directly after {@link
* FleetState#nudgeText()}), or {@code ""} when disabled or the state is not {@code HIGH}
*/
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading) {
return contextNotice(enabled, reading, false);
}
/**
* fleetd #609 review: as {@link #contextNotice(boolean, LeadContextGauge.Reading)}, but also gated on
* {@code alreadyNotified} — the context latch as it stood <em>before</em> the current tick's decision.
* Without this gate, every pending-driven {@code INJECT} that lands while the context stays {@code
* HIGH} would re-append the full notice on top of an already-latched stretch, making the notice's own
* closing sentence ("You will not be told again until your context reads ok.") false. {@link
* #injectNudge} is the only caller that passes a non-default {@code alreadyNotified}.
*
* @param alreadyNotified whether the lead has already been told about the current HIGH stretch
*/
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified) {
if (!enabled || alreadyNotified || reading.state() != LeadContextGauge.State.HIGH) {
return "";
}
StringBuilder sb = new StringBuilder(" Your own context is nearly full");
String compactionWord = reading.compactions() == 1 ? "compaction" : "compactions";
if (reading.tokens() != null) {
sb.append(": ").append(reading.tokens()).append(" tokens used, ")
.append(reading.compactions()).append(' ').append(compactionWord).append(" so far.");
} else {
// A HIGH reading always carries a non-null token count today: LeadContextGauge only
// reaches HIGH by comparing a number against HIGH_THRESHOLD_TOKENS. That invariant
// lives in another class and nothing asserts it, so this branch does not rely on it —
// it drops the token clause rather than printing "null tokens".
sb.append(" (").append(reading.compactions()).append(' ').append(compactionWord)
.append(" so far).");
}
sb.append(" A fresh session would work better. To hand over: call fleet_handover(action=\"open\"), "
+ "write the file it names, ask the operator, then call fleet_handover(action=\"confirm\", "
+ "token, operatorConfirmed). Only the operator can approve the roll. You will not be told "
+ "again until your context reads ok.");
return sb.toString();
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext() {
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
@@ -871,10 +871,23 @@ public final class SessionManager implements TurnListener {
// digest lets a lead tell at a glance whether all members got the same charter; the source
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
// charter was composed ("none").
// #604: charterBytes rides along with charterSha256, not with charterSource — it is only
// meaningful as the digest's companion (a length turns "they differ" into "by how much").
// A member with no composed charter reports charterSource and nothing else, as before.
//
// Gating on the DIGEST rather than on the receipt is deliberate, and the two are not
// always null together. CharterReceipt.compose() derives the digest with digestOf(), which
// returns null for BLANK text, while the byte count is composed.getBytes().length, which
// does not. So a whitespace-only charter (a blank fleet.charters.<role> on a profile with
// no MCP, so no reply charter is appended) yields a null digest beside a non-zero size.
// Reporting a size with no digest would say "they differ by N bytes" about an artifact we
// cannot fingerprint, so this gate omits both. Never widen it to the receipt-level null
// check without deciding what that case should report.
if (session.charterReceipt() != null) {
m.put("charterSource", session.charterReceipt().charterSource());
if (session.charterReceipt().charterSha256() != null) {
m.put("charterSha256", session.charterReceipt().charterSha256());
m.put("charterBytes", session.charterReceipt().charterBytes());
}
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
@@ -0,0 +1,84 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
/**
* fleetd #589 Group 2: {@link Fleetd#claudeCodeLauncher} is the factory that replaced {@code
* main}'s inline {@code new ClaudeCodeLauncher(...)} call, whose 10th argument is the CB-596 {@code
* memberCredentials} policy supplier ({@code () -> config.get().memberCredentials()}). Before this
* ticket that argument was untestable wiring: replacing it with {@code () -> null} compiled with 0
* errors and left every existing test green, since no existing test builds the exact object {@code
* main} wires and then spawns it. {@code memberCredentials} being a lambda is never itself {@code
* null}, so {@link dev.ltms.fleet.member.HerdrPeerLauncher#applyMemberCredentialPolicy} sees {@code
* memberCredentials.get() == null} and silently shadows nothing — reopening the exact CB-592
* exposure gap CB-596's policy closed (gitea issue #82).
*
* <p>This test drives the factory with a real {@link ConfigRef} carrying a {@code
* memberCredentials:} block, spawns through the resulting launcher, and inspects what {@code
* tab.create} actually carried — the same observable surface {@code ClaudeCodeLauncherTest}'s
* {@code everyKnownNameNotAllowedIsShadowedWithTheSentinel} uses for the launcher's own credential
* policy, applied here to prove {@code main}'s wiring reaches it.
*/
class FleetdClaudeCodeLauncherCredentialWiringTest {
private static final String YAML = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
ltms-local:
baseUrl: http://gx00.gw:8000
model: coder
memberCredentials:
policy: deny-by-default
known:
- GITEA_ACCESS_TOKEN
""";
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
}
@Test
@DisplayName("main's memberCredentials wiring reaches ClaudeCodeLauncher: a known-but-not-allowed "
+ "name is shadowed on spawn")
void memberCredentialsWiringReachesClaudeCodeLauncher(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = ConfigRef.fixed(cfg);
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher launcher = Fleetd.claudeCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
cfg.profiles(), cfg, config);
launcher.spawn();
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
assertNotNull(shadowed,
"GITEA_ACCESS_TOKEN is 'known' but not 'allow'-ed in the loaded config — it must be "
+ "explicitly shadowed on spawn; replacing the memberCredentials supplier with "
+ "() -> null at the Fleetd.claudeCodeLauncher call site must fail this "
+ "assertion, since a null policy shadows nothing");
assertFalse(shadowed.isBlank(), "the overlay value must be non-blank");
}
}
@@ -0,0 +1,86 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.MemberSession;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 1: {@link Fleetd#exhaustedPatternLookup} is the factory that replaced {@code
* main}'s inline lambda — resolve a herdr {@code target} to its session's profile, then to that
* profile's live {@link LiveExhaustedPatterns#patternFor}. Same shape as {@link
* Fleetd#worktreeBranchLookup} (which {@code FleetdWorktreeBranchLookupTest} pins the same way).
*
* <p>Before this ticket the lambda was built inline in {@code main} and untestable: replacing it
* with {@code target -> null} — the exact shape of {@link ExhaustedPatternLookup#none()} — compiled
* with 0 errors and left every existing test green. Per the ticket, this is the worst consequence
* in the whole #589 sweep: a genuine usage-limit refusal would stop being classified as {@code
* BACKEND_EXHAUSTED} and would be handed back to a waiting {@code fleet_send} as if it were real
* completed work.
*/
class FleetdExhaustedPatternLookupWiringTest {
private static final String YAML = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: claude-opus-5
exhaustedPattern: "usage limit"
""";
private static MemberSession session(String terminal, String profile) {
return new MemberSession("pane-" + terminal, terminal, profile, MemberRole.DEV,
"/cwd", null, 0L, 0L, 0, MemberSession.State.READY, null, null);
}
private static LiveExhaustedPatterns liveExhaustedPatterns(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
return new LiveExhaustedPatterns(() -> ConfigRef.fixed(cfg).get().profiles());
}
@Test
@DisplayName("a known target resolves through its session's profile to that profile's live pattern")
void knownTargetResolvesThroughItsProfile(@TempDir Path dir) throws Exception {
LiveExhaustedPatterns patterns = liveExhaustedPatterns(dir);
ExhaustedPatternLookup lookup = Fleetd.exhaustedPatternLookup(
() -> List.of(session("term1", "terra")), patterns);
Pattern resolved = lookup.patternFor("term1");
assertNotNull(resolved,
"the lookup must resolve term1 -> profile 'terra' -> LiveExhaustedPatterns.patternFor("
+ "'terra') — replacing the lambda body with 'target -> null' at the "
+ "Fleetd.exhaustedPatternLookup call site must fail this assertion");
assertTrue(resolved.matcher("the usage limit has been reached").find());
}
@Test
@DisplayName("an unknown target resolves to null, not a thrown exception")
void unknownTargetResolvesToNull(@TempDir Path dir) throws Exception {
LiveExhaustedPatterns patterns = liveExhaustedPatterns(dir);
ExhaustedPatternLookup lookup = Fleetd.exhaustedPatternLookup(
() -> List.of(session("term1", "terra")), patterns);
assertNull(lookup.patternFor("term_stranger"));
}
}
@@ -0,0 +1,62 @@
package dev.ltms.fleet;
import dev.ltms.fleet.inject.ExhaustionSink;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 1: {@link Fleetd#forwardingExhaustionSink} is the factory that replaced
* {@code main}'s inline {@code ExhaustionSink.forwardingTo(exhaustionSinkRef::get)} (fleetd #175's
* construction-order break: the adapters need a sink before {@code sessions} exists to build the
* real one). Before this ticket that call site was untestable wiring: replacing the supplier
* argument with a hardcoded {@code () -> ExhaustionSink.none()} compiled with 0 errors and left
* every existing test green, because no test builds the object {@code main} actually wires and
* then mutates the reference afterward — every existing {@code ExhaustionSink.forwardingTo} caller
* in this codebase reads and writes the SAME reference within one test, so a hardcoded-none supplier
* and a correctly-forwarding one are indistinguishable to them.
*
* <p>This test builds the reference, builds the forwarder from it, and only THEN repoints the
* reference at a spy sink — the discriminating order fleetd #175's whole design depends on
* ({@code exhaustionSinkRef} starts at {@code none()} and is repointed once {@code sessions}
* exists). A forwarder that captured a fixed target at construction time (the inert form) can never
* see that later repoint.
*/
class FleetdExhaustionSinkForwardingWiringTest {
@Test
@DisplayName("the forwarder reads the reference live: repointing it AFTER construction is honoured")
void forwarderReadsTheReferenceLiveNotAFixedTargetCapturedAtConstruction() {
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
ExhaustionSink forwarder = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
AtomicBoolean spyCalled = new AtomicBoolean(false);
exhaustionSinkRef.set((target, reason, profile) -> spyCalled.set(true));
forwarder.onExhausted("term_x", "usage limit reached", "terra");
assertTrue(spyCalled.get(),
"forwardingExhaustionSink must delegate to whatever exhaustionSinkRef currently "
+ "holds — hardcoding the supplier to () -> ExhaustionSink.none() at the "
+ "Fleetd.forwardingExhaustionSink call site must fail this assertion, "
+ "since the spy set into the reference after construction would never run");
}
@Test
@DisplayName("before any repoint, the forwarder is inert — it starts at none(), not a crash")
void beforeAnyRepointTheForwarderIsInert() {
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
ExhaustionSink forwarder = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
AtomicBoolean spyCalled = new AtomicBoolean(false);
forwarder.onExhausted("term_x", "usage limit reached", "terra");
assertFalse(spyCalled.get(), "nothing was ever wired to be called here — this only pins "
+ "that the factory does not throw before a real sink is published");
}
}
@@ -0,0 +1,110 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.session.SessionManager;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashMap;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 1: {@link Fleetd#publishExhaustionSink} is the factory that replaced {@code
* main}'s previously untested two-statement sequence — build the real {@link
* Fleetd#exhaustionSink}, then {@code exhaustionSinkRef.set(exhaustionSink)}. {@link
* Fleetd#exhaustionSink} itself is already pinned by {@code FleetdExhaustionSinkWarningTest} (its
* log text) — what was NEVER pinned is the {@code .set(...)} call: {@code main} could replace it
* with {@code exhaustionSinkRef.set(ExhaustionSink.none())} and compile with 0 errors, leaving
* every existing test green, because {@link Fleetd#exhaustionSink}'s own tests build and call the
* sink directly, never through the reference {@code main} publishes it into.
*
* <p>This test proves the PUBLISHED reference — not a freshly rebuilt sink — is the one that
* actually quarantines a credential, by reading {@link BackendQuarantine#isQuarantined} after
* calling {@code exhaustionSinkRef.get().onExhausted(...)}, the same object {@link
* Fleetd#forwardingExhaustionSink} forwards to in production.
*/
class FleetdExhaustionSinkPublishWiringTest {
private static final String YAML = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: claude-opus-5
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static SessionManager emptyRosterSessions() {
FakeHerdr h = new FakeHerdr();
FleetConfig.Profile dummy = new FleetConfig.Profile(
"dummy", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(dummy.profile(), dummy), dummy.profile(), _ -> "tok");
// Never acquires a session — publishExhaustionSink's built sink resolves target -> profile
// via the profileHint fallback (fleetd #234), exactly like OpenCodeLauncher's real call
// site does, so this never needs a populated roster.
return new SessionManager(launcher);
}
@Test
@DisplayName("the published reference actually quarantines — not a rebuilt-but-never-set sink")
void publishedReferenceActuallyQuarantines(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = ConfigRef.fixed(cfg);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
Map<String, String> reasonByCredential = new HashMap<>();
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
Fleetd.publishExhaustionSink(exhaustionSinkRef, emptyRosterSessions(), config, quarantine,
reasonByCredential, cfg);
exhaustionSinkRef.get().onExhausted("term_x", "The usage limit has been reached", "terra");
assertTrue(quarantine.isQuarantined("terra"),
"publishExhaustionSink must repoint exhaustionSinkRef at the REAL sink — "
+ "replacing the .set(...) call with exhaustionSinkRef.set(ExhaustionSink.none()) "
+ "at the Fleetd.publishExhaustionSink call site must fail this assertion, "
+ "since none()'s onExhausted does nothing");
}
@Test
@DisplayName("before publishing, the reference is still inert — no quarantine, no crash")
void beforePublishingTheReferenceIsStillInert(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = ConfigRef.fixed(cfg);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
exhaustionSinkRef.get().onExhausted("term_x", "The usage limit has been reached", "terra");
assertFalse(quarantine.isQuarantined("terra"),
"nothing was published yet — this only pins the starting state the other test's "
+ "assertion actually distinguishes from");
}
}
@@ -0,0 +1,80 @@
package dev.ltms.fleet;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.function.BiConsumer;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 3, site {@code :591}. {@code Fleetd.main} wires {@link
* dev.ltms.fleet.health.FleetHealthMonitor}'s {@code failTarget} callback with {@code
* messages::abandon} — before this ticket that was an inline argument to {@code new
* FleetHealthMonitor(...)}. Measured: replacing it with a no-op {@code BiConsumer} at the call
* site compiles with 0 errors and leaves the full suite green, because nothing else in the tree
* ever drives that specific constructor argument. In production it means a member the monitor
* classifies {@code GONE}/{@code NEVER_READY} never has its pending ticket failed — the caller
* keeps reporting {@code PENDING} for the full 30-minute async timeout instead of the immediate,
* accurate failure CB-580 exists to give it.
*
* <p>This test calls {@link Fleetd#healthFailTarget} directly — never {@code FleetHealthMonitor}
* or {@code main} — against a real {@link MessageService}, using the same {@code sendAsync} +
* {@code poll} observable {@link MessageServiceTest} already relies on to pin {@code
* MessageService.abandon} itself.
*/
class FleetdHealthFailTargetWiringTest {
private static final String T = "term_a";
@Test
@DisplayName("Fleetd.healthFailTarget delegates to the real MessageService.abandon, not a no-op")
void healthFailTargetDelegatesToMessagesAbandon() throws Exception {
FakeHerdr herdr = new FakeHerdr().readText("$ prompt");
AgentControl agents = new AgentControl(herdr);
Rendezvous rendezvous = new Rendezvous();
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, rendezvous);
BiConsumer<String, String> failTarget = Fleetd.healthFailTarget(messages);
String ticket = messages.sendAsync(T, "long task");
awaitWaiting(rendezvous);
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase());
failTarget.accept(T, "member unreachable (health monitor)");
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 3000;
while (System.currentTimeMillis() < deadline) {
view = messages.poll(ticket);
if (view.phase() != MessageService.Phase.PENDING) {
break;
}
//noinspection BusyWait
Thread.sleep(10);
}
assertEquals(MessageService.Phase.FAILED, view == null ? null : view.phase(),
"Fleetd.healthFailTarget(messages) must return messages::abandon — replacing it "
+ "with a no-op BiConsumer at the Fleetd.healthFailTarget call site means "
+ "this ticket is never failed and keeps polling as PENDING");
assertTrue(view.detail() != null && view.detail().contains("member unreachable"),
"the failure reason passed to failTarget.accept must reach MessageService.abandon "
+ "and end up in the ticket's detail");
}
private static void awaitWaiting(Rendezvous rendezvous) throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
}
}
@@ -0,0 +1,100 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #602 gauge-wiring: {@link Fleetd#leadConfigDirLookup} is the factory {@code Fleetd.main}
* wires into {@code FleetMcp.LeadConfigDirSource} so {@code fleet_list}'s {@code context} row reads
* the transcript directory a lead's OWN profile actually writes to, instead of always falling back
* to the built-in {@code <user.home>/.claude} default (see {@code LeadContextGauge}).
*
* <p>{@code FleetMcpLeadContextGaugeWiringTest} proves the directory this factory returns is what
* actually gets read; this class proves the factory's own matching logic — the same
* {@code fleet.leaders.<name>.profile} link {@code Fleetd.leadSeatLookup} already follows (see
* {@code FleetdLeadSeatLookupTest}), one step further to that profile's own {@code configDir:}.
*/
class FleetdLeadConfigDirLookupTest {
private static FleetConfig.Profile profileWithConfigDir(String name, String configDir) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", configDir, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, null, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("a lead on a profile that sets configDir resolves to that directory")
void leadOnAProfileWithConfigDirResolvesToIt() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
assertEquals("/mnt/opus-claude", lookup.apply("primary"));
}
@Test
@DisplayName("changing the config from directory A to directory B changes what the lookup reports")
void configChangeFromDirectoryAToDirectoryBChangesTheAnswer() {
java.util.concurrent.atomic.AtomicReference<Map<String, FleetConfig.Profile>> profilesRef =
new java.util.concurrent.atomic.AtomicReference<>(
Map.of("opus", profileWithConfigDir("opus", "/mnt/dir-a")));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, String> lookup = Fleetd.leadConfigDirLookup(profilesRef::get, leaders);
assertEquals("/mnt/dir-a", lookup.apply("primary"), "must read directory A before the config changes");
profilesRef.set(Map.of("opus", profileWithConfigDir("opus", "/mnt/dir-b")));
assertEquals("/mnt/dir-b", lookup.apply("primary"), "must read directory B once the LIVE config changes — "
+ "a lookup that snapshotted the profile map at construction would still answer directory A here");
}
@Test
@DisplayName("a lead entry with no `profile:` (recognise-only) resolves to null, not a thrown exception")
void recogniseOnlyLeadWithNoProfileResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
"claude", "claude-sonnet-5");
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("a lead naming a profile that is not configured resolves to null, not a thrown exception")
void leadOnAnUnconfiguredProfileResolvesToNull() {
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("ghost-profile"));
Function<String, String> lookup = Fleetd.leadConfigDirLookup(Map::of, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("a lead on a profile that sets no configDir override resolves to null")
void leadOnAProfileWithNoConfigDirResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
void unrecognisedLeadNameResolvesToNull() {
Function<String, String> lookup = Fleetd.leadConfigDirLookup(Map::of, Map.of());
assertNull(lookup.apply("ghost-lead"));
}
}
@@ -0,0 +1,98 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* fleetd #602 gauge-wiring follow-up (PR #606 review comment 17353): {@code Fleetd.main}'s {@code
* LeadConfigDirSource} local used to be a bare {@code new FleetMcp.LeadConfigDirSource(
* leadConfigDirLookup(...))} built inline, with nothing a test could call directly. Measured on
* that shape: replacing the whole expression with {@code FleetMcp.LeadConfigDirSource.none()} at
* the call site compiled with 0 errors and left the full 1822-test suite green — the daemon could
* be changed to always report every lead's context as {@code UNKNOWN}, forever, and no test would
* notice. That is the same hand-built-vs-config-wired shape as fleetd #561/#248/#426/#562
* ({@code FleetdLoopHealthSourceWiringTest}).
*
* <p>{@code FleetMcpLeadContextGaugeWiringTest} and {@code FleetdLeadConfigDirLookupTest} both
* predate this class and are both still correct — but neither can catch the mutation above. One
* builds its own {@code FleetMcp} and hands it its own {@code LeadConfigDirSource}; the other
* builds its own lookup and calls {@link Fleetd#leadConfigDirLookup} directly. Neither one ever
* calls the thing {@code Fleetd.main} actually calls.
*
* <p>The fix extracts the inline {@code new} into {@link Fleetd#leadConfigDirSource}, a
* package-private factory in the same style as {@link Fleetd#loopHealthSource}/{@link
* Fleetd#capacitySource}/{@link Fleetd#healthCoverageSource} — which is exactly what makes it
* directly callable here. This test calls that factory with real {@link FleetConfig.Profile}/
* {@link FleetConfig.Leader} fixtures (the same shapes {@code FleetdLeadConfigDirLookupTest}
* already uses) and asserts the returned source resolves a real {@code configDir} — a property
* that would be false if {@link Fleetd#leadConfigDirSource} were mutated to {@code return
* FleetMcp.LeadConfigDirSource.none();}. Measured: mutating exactly that line makes
* {@link #resolvesTheRealConfiguredConfigDir()} fail ({@code expected: </mnt/opus-claude> but was:
* <null>}); restoring it makes the whole suite green again.
*
* <p><b>What this class does not and cannot cover.</b> {@code main}'s own line —
* {@code leadConfigDirSource(() -> config.get().profiles(), leaders)} — could itself be swapped
* for a bare {@code FleetMcp.LeadConfigDirSource.none()}, bypassing this factory entirely. Measured:
* that exact mutation compiles with 0 errors and leaves every test in this file, and the full
* 1825-test suite, green. {@link Fleetd#loopHealthSource}'s own wiring test has the identical gap
* for its own one-line call in {@code main} — no test in this codebase calls {@code Fleetd.main}
* far enough to observe which factory call it made. This class narrows the gap from "nothing tests
* the wiring" (the pre-extraction state this ticket found) to "the factory's own logic is pinned,
* and main's call to it is a one-line, visually-verifiable delegation" — the same standard already
* accepted for {@code loopHealthSource}/{@code capacitySource}/{@code healthCoverageSource}.
*/
class FleetdLeadConfigDirSourceWiringTest {
private static FleetConfig.Profile profileWithConfigDir(String name, String configDir) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", configDir, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, null, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("the returned source resolves the lead's REAL configured configDir, not a hardcoded null")
void resolvesTheRealConfiguredConfigDir() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertEquals("/mnt/opus-claude", source.configDirFor().apply("primary"),
"the configDirFor function must delegate to the real leadConfigDirLookup — mutating "
+ "Fleetd.leadConfigDirSource's own body to `return FleetMcp.LeadConfigDirSource.none();` "
+ "must fail this assertion (measured: it does — see this test's class javadoc for the "
+ "companion measurement on main's one-line call to this factory, which this assertion "
+ "does not and structurally cannot cover)");
}
@Test
@DisplayName("a lead on a profile with no configDir override still resolves to null, not a crash")
void leadWithNoConfigDirOverrideResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertNull(source.configDirFor().apply("primary"),
"no configDir: override configured ⇒ null, so LeadContextGauge falls back to its own default");
}
@Test
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
void unrecognisedLeadNameResolvesToNull() {
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(Map::of, Map.of());
assertNull(source.configDirFor().apply("ghost-lead"));
}
}
@@ -0,0 +1,156 @@
package dev.ltms.fleet;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.lead.LeadContextGauge;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* fleetd #609: {@link Fleetd#leadContextLookup} is the factory {@code Fleetd.main} wires into
* {@code LeadHeartbeatLoop.LeadContextSource} so the idle-lead heartbeat can read a lead's own
* {@link LeadContextGauge} reading by its terminal id. Three hops: terminal → lead name (unknown ⇒
* UNKNOWN), lead name → configDir, and {@code agents.get(terminal)} for the live session id and
* agent type the gauge itself needs — a herdr failure on that last hop must degrade to UNKNOWN, not
* throw and kill the heartbeat's own tick.
*
* <p>{@code AgentControl.agentCall} resolves a {@code term_}-prefixed target's pane id via a first
* {@code agent.list} round trip (see its own javadoc); this class's terminal id deliberately does
* NOT start with {@code term_} so the stub {@link HerdrClient} below only needs to answer
* {@code agent.get} — the one call this factory actually depends on.
*/
class FleetdLeadContextLookupTest {
private static final String LEAD_TERMINAL = "leadpane1";
private static final String LEAD_NAME = "opus";
private static final String SESSION_ID = "sess-609-happy-path";
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
}
/** An {@link AgentControl} whose every {@code agent.get} answers with the given session/type/status. */
private static AgentControl agentControlStub(String sessionId, String agentType, String status) {
HerdrClient client = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
if (!"agent.get".equals(method)) {
throw new HerdrException("stub has no canned response for " + method);
}
try {
String sessionField = sessionId == null ? ""
: ",\"agent_session\":{\"kind\":\"id\",\"value\":\"" + sessionId + "\"}";
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
return new ObjectMapper().readTree(("""
{"type":"agent_info","agent":{"terminal_id":"%s","agent":%s,
"agent_status":"%s"%s}}""")
.formatted(LEAD_TERMINAL, agentField, status, sessionField));
} catch (Exception e) {
throw new HerdrException("stub decode failed", e);
}
}
@Override
public void close() {
}
};
return new AgentControl(client);
}
/** An {@link AgentControl} whose every herdr call fails — models a herdr hiccup mid-tick. */
private static AgentControl throwingAgentControl() {
HerdrClient client = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
throw new HerdrException("herdr unreachable (stub)");
}
@Override
public void close() {
}
};
return new AgentControl(client);
}
@Test
@DisplayName("an unrecognised terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply("ghost-terminal"));
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
}
@Test
@DisplayName("agents.get throwing degrades to UNKNOWN, no exception escapes")
void agentsGetThrowingDegradesToUnknown() {
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply(LEAD_TERMINAL));
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"a herdr failure resolving the live agent must degrade to UNKNOWN, never kill the heartbeat's tick");
}
@Test
@DisplayName("the happy path resolves the configDir, sessionId and agentType through to the gauge")
void happyPathPassesResolvedFactsThroughToTheGauge(@TempDir Path tmp) throws IOException {
writeTranscript(tmp, SESSION_ID, 12_345);
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
AgentControl agents = agentControlStub(SESSION_ID, "claude", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.OK, reading.state(),
"the resolved configDir + the live agent's own sessionId/agentType must reach the gauge — "
+ "a wrong hop anywhere in the chain would read no transcript and report UNKNOWN instead");
assertEquals(12_345L, reading.tokens());
}
@Test
@DisplayName("a lead whose agent type is not claude still resolves to UNKNOWN, never a crash")
void nonClaudeAgentTypeResolvesToUnknown(@TempDir Path tmp) throws IOException {
writeTranscript(tmp, SESSION_ID, 12_345);
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
AgentControl agents = agentControlStub(SESSION_ID, "opencode", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"the agentType hop must reach the gauge too — a non-claude peer must not be misread as claude");
}
}
@@ -0,0 +1,108 @@
package dev.ltms.fleet;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #609: {@code Fleetd.main}'s {@code LeadHeartbeatLoop.LeadContextSource} local could be
* swapped for a bare {@code LeadHeartbeatLoop.LeadContextSource.none()} at the call site — compiling
* with 0 errors and leaving every pre-existing test green — exactly the shape #602/#606 already found
* for {@code LeadConfigDirSource} (see {@code FleetdLeadConfigDirSourceWiringTest}'s own javadoc for
* the measured version of that gap).
*
* <p>The fix follows the same pattern: {@link Fleetd#leadContextSource} is the extracted,
* directly-callable factory {@code main} calls to build the source it hands {@code
* LeadHeartbeatLoop}'s constructor. This test calls that exact factory and asserts it resolves a
* REAL reading off a real transcript file — a property that would be false if {@link
* Fleetd#leadContextSource} were mutated to {@code return LeadHeartbeatLoop.LeadContextSource.none();}.
*
* <p>What this class does not and cannot cover: {@code main}'s own one-line call to this factory
* could itself be swapped for {@code LeadHeartbeatLoop.LeadContextSource.none()}, bypassing this
* factory entirely — the same structural gap {@code FleetdLeadConfigDirSourceWiringTest} names for
* its own factory, and for the same reason (no test in this codebase calls {@code Fleetd.main} far
* enough to observe which factory call it made).
*/
class FleetdLeadContextSourceWiringTest {
/** Deliberately not {@code term_}-prefixed — see {@code FleetdLeadContextLookupTest}'s class doc. */
private static final String LEAD_TERMINAL = "leadpane1";
private static final String LEAD_NAME = "opus";
private static final String SESSION_ID = "sess-609-wiring";
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
}
private static AgentControl agentControlStub() {
HerdrClient client = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
if (!"agent.get".equals(method)) {
throw new HerdrException("stub has no canned response for " + method);
}
try {
return new ObjectMapper().readTree(("""
{"type":"agent_info","agent":{"terminal_id":"%s","agent":"claude",
"agent_status":"idle","agent_session":{"kind":"id","value":"%s"}}}""")
.formatted(LEAD_TERMINAL, SESSION_ID));
} catch (Exception e) {
throw new HerdrException("stub decode failed", e);
}
}
@Override
public void close() {
}
};
return new AgentControl(client);
}
@Test
@DisplayName("main's factory resolves a REAL reading, not the inert none() answer")
void resolvesARealReadingNotTheInertNoneAnswer(@TempDir Path tmp) throws IOException {
writeTranscript(tmp, SESSION_ID, 54_321);
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), () -> liveLeadTerminals, name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.OK, reading.state(),
"mutating Fleetd.leadContextSource's own body to `return LeadHeartbeatLoop.LeadContextSource.none();` "
+ "must fail this assertion");
assertEquals(54_321L, reading.tokens());
}
@Test
@DisplayName("an unrecognised lead terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), Map::of, name -> null);
assertEquals(LeadContextGauge.State.UNKNOWN, source.readingFor().apply("ghost-terminal").state());
}
}
@@ -0,0 +1,48 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.net.ServerSocket;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 3, site {@code :518}. {@code Fleetd.main} passes {@code LeadMailbox::open} as
* the {@link Fleetd.LeadMailboxOpener} argument to {@code openLeadMailbox(...)} — before this
* ticket that method reference was inline at the call site. Measured: replacing it with the inert
* {@code (uri, selfCoordId, prefetch) -> null} compiles with 0 errors and leaves the full suite
* green — {@code FleetdLeadMailboxSelectionTest} drives {@code openLeadMailbox} with its own
* injected opener and never observes what {@code main} itself actually passes.
*
* <p>This test calls {@link Fleetd#leadMailboxOpener} directly and proves it is the real,
* network-attempting opener rather than a stub: pointed at a guaranteed-closed local port, it must
* throw, exactly mirroring {@link FleetdReplyInboxOpenerWiringTest} for the reply-inbox opener.
* The inert form never attempts a connection and returns {@code null} without throwing, so it
* fails this assertion silently.
*/
class FleetdLeadMailboxOpenerWiringTest {
@Test
@DisplayName("Fleetd.leadMailboxOpener is the real LeadMailbox::open, not a stub that never connects")
void leadMailboxOpenerAttemptsARealConnection() throws Exception {
int closedPort;
try (ServerSocket socket = new ServerSocket(0)) {
closedPort = socket.getLocalPort();
} // released immediately — connecting to it now is a guaranteed refusal, not a fluke
Fleetd.LeadMailboxOpener opener = Fleetd.leadMailboxOpener();
IllegalStateException thrown = assertThrows(IllegalStateException.class,
() -> opener.open("amqp://user:pw@127.0.0.1:" + closedPort + "/coord", "coord-1", 50),
"Fleetd.leadMailboxOpener() must be LeadMailbox::open — a real network attempt "
+ "against a genuinely unreachable broker must throw. The inert form "
+ "(uri, selfCoordId, prefetch) -> null never attempts a connection and "
+ "returns null instead of throwing, so it would fail this assertion "
+ "silently.");
assertTrue(thrown.getMessage().contains("cannot connect to AMQP coordination broker"),
"must be LeadMailbox.open's own real failure message, not a different exception "
+ "shape standing in for it");
}
}
@@ -0,0 +1,72 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 1: {@link Fleetd#liveExhaustedPatterns} is the factory that replaced {@code
* main}'s inline {@code new LiveExhaustedPatterns(() -> config.get().profiles())}. Before this
* ticket, that supplier argument was untestable wiring: replacing it with a hardcoded {@code () ->
* Map.of()} compiled with 0 errors and left every existing test green, because {@code
* LiveExhaustedPatternsTest} builds its own instance directly with a hand-supplied map and never
* goes through {@code main}'s call site.
*
* <p>Silently losing this wiring means every profile's {@code exhaustedPattern} stops being
* recognised — {@link Fleetd#exhaustedPatternLookup} would never see a match, and a genuine
* usage-limit refusal would be handed back to a waiting {@code fleet_send} as real completed work.
*/
class FleetdLiveExhaustedPatternsWiringTest {
private static final String YAML = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: claude-opus-5
exhaustedPattern: "usage limit"
gx:
baseUrl: http://gx00.gw:8000
""";
private static ConfigRef loadConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
return ConfigRef.fixed(cfg);
}
@Test
@DisplayName("a profile with a configured exhaustedPattern is armed, with a compiled matcher")
void configuredProfileIsArmed(@TempDir Path dir) throws Exception {
LiveExhaustedPatterns patterns = Fleetd.liveExhaustedPatterns(loadConfig(dir));
assertTrue(patterns.armed("terra"),
"the config's live profiles() supplier must reach LiveExhaustedPatterns — hardcoding "
+ "the supplier to () -> Map.of() at the Fleetd.liveExhaustedPatterns call "
+ "site must fail this assertion");
assertTrue(patterns.patternFor("terra").matcher("the usage limit has been reached").find());
}
@Test
@DisplayName("a profile with no configured exhaustedPattern is not armed, but is still resolvable")
void unconfiguredProfileIsNotArmed(@TempDir Path dir) throws Exception {
LiveExhaustedPatterns patterns = Fleetd.liveExhaustedPatterns(loadConfig(dir));
assertFalse(patterns.armed("gx"), "'gx' has no exhaustedPattern configured");
assertNull(patterns.patternFor("gx"));
}
}
@@ -0,0 +1,77 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.member.OpenCodeLauncher;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
/**
* fleetd #589 Group 2: {@link Fleetd#openCodeLauncher} is the factory that replaced {@code main}'s
* inline {@code new OpenCodeLauncher(...)} call — the {@code opencode} counterpart to {@link
* Fleetd#claudeCodeLauncher}, extracted for the identical reason. Its {@code memberCredentials}
* argument is the same {@code () -> config.get().memberCredentials()} supplier; replacing it with
* {@code () -> null} compiled with 0 errors and left every existing test green before this ticket,
* reopening the same CB-592 exposure gap CB-596's policy closed.
*
* <p>Same observable surface as {@code OpenCodeLauncherTest}'s own {@code memberCredentials} tests:
* spawn through the launcher {@code main} actually wires and inspect what {@code tab.create}
* carried.
*/
class FleetdOpenCodeLauncherCredentialWiringTest {
private static final String YAML = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
gemini:
kind: opencode
model: google/gemini-2.5-pro
memberCredentials:
policy: deny-by-default
known:
- GITEA_ACCESS_TOKEN
""";
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
}
@Test
@DisplayName("main's memberCredentials wiring reaches OpenCodeLauncher: a known-but-not-allowed "
+ "name is shadowed on spawn")
void memberCredentialsWiringReachesOpenCodeLauncher(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, YAML);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = ConfigRef.fixed(cfg);
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = Fleetd.openCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), cfg.profiles(), cfg, config, ExhaustionSink.none());
launcher.spawn();
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
assertNotNull(shadowed,
"GITEA_ACCESS_TOKEN is 'known' but not 'allow'-ed in the loaded config — it must be "
+ "explicitly shadowed on spawn; replacing the memberCredentials supplier with "
+ "() -> null at the Fleetd.openCodeLauncher call site must fail this "
+ "assertion, since a null policy shadows nothing");
assertFalse(shadowed.isBlank(), "the overlay value must be non-blank");
}
}
@@ -0,0 +1,107 @@
package dev.ltms.fleet;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.session.SessionManager;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.function.Consumer;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 3, site {@code :611-631}. {@code Fleetd.main} wires {@code
* sessions.onRelease(...)} with a lambda that calls three collaborators — {@code
* messages.abandon}, {@code replyInbox.release}, and {@code primaryRegistry.forgetDelegation} —
* before this ticket built inline inside {@code main}. Measured: replacing the whole lambda body
* with {@code detail -> { }} compiles with 0 errors and leaves the full suite green, because each
* collaborator is separately tested in isolation ({@code MessageServiceTest}, {@code
* PrimaryRegistryTest}) but nothing before this ticket drove the lambda that calls all three from
* {@code main}.
*
* <p>{@code MessageService.abandon}'s own javadoc documents the consequence: without this,
* tearing a worker down leaves its rendezvous waiter open, so a blocking {@code fleet_send} keeps
* blocking and an async one reports {@code PENDING} for a hardcoded thirty minutes on every {@code
* fleet_stop} and every idle-reap.
*
* <p>This test calls {@link Fleetd#releaseCleanup} directly — never {@code SessionManager} or
* {@code main} — against real {@link MessageService}, {@link InMemoryReplyInbox}, and {@link
* PrimaryRegistry} instances, and asserts each collaborator's own observable effect: the pending
* async ticket transitions to {@code FAILED} (abandon), the inbox no longer owns the target's
* queue (release), and the recorded delegation is forgotten (forgetDelegation).
*/
class FleetdReleaseCleanupWiringTest {
private static final String T = "term_a";
@Test
@DisplayName("Fleetd.releaseCleanup reaches messages.abandon, replyInbox.release, and primaryRegistry.forgetDelegation")
void releaseCleanupReachesAllThreeCollaborators() throws Exception {
FakeHerdr herdr = new FakeHerdr().readText("$ prompt");
AgentControl agents = new AgentControl(herdr);
Rendezvous rendezvous = new Rendezvous();
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, rendezvous);
InMemoryReplyInbox replyInbox = new InMemoryReplyInbox();
PrimaryRegistry primaryRegistry = new PrimaryRegistry(null);
// Set up the "before" state each collaborator's own effect is measured against.
String ticket = messages.sendAsync(T, "long task");
awaitWaiting(rendezvous);
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase(),
"sanity: the async ticket is pending before cleanup runs");
replyInbox.own(T);
replyInbox.publish(T, "msg-1", "hello");
assertEquals(1, replyInbox.peek(T).size(),
"sanity: the inbox owns T and holds one message before cleanup runs");
primaryRegistry.recordDelegation(T, "lead-1");
assertEquals("lead-1", primaryRegistry.nudgeTargetFor(T).orElse(null),
"sanity: the delegation is recorded before cleanup runs");
Consumer<SessionManager.ReleaseDetail> cleanup =
Fleetd.releaseCleanup(messages, replyInbox, primaryRegistry);
cleanup.accept(new SessionManager.ReleaseDetail(T, null, null, null, null));
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 3000;
while (System.currentTimeMillis() < deadline) {
view = messages.poll(ticket);
if (view.phase() != MessageService.Phase.PENDING) {
break;
}
//noinspection BusyWait
Thread.sleep(10);
}
assertEquals(MessageService.Phase.FAILED, view == null ? null : view.phase(),
"releaseCleanup must call messages.abandon(...) — an inert detail -> { } lambda "
+ "leaves this ticket PENDING forever");
assertTrue(view.detail() != null && view.detail().contains("released"),
"the abandon reason must say the worker session was released");
assertTrue(replyInbox.peek(T).isEmpty(),
"releaseCleanup must call replyInbox.release(...) — an inert lambda leaves the "
+ "inbox still owning T with its message");
assertTrue(primaryRegistry.nudgeTargetFor(T).isEmpty(),
"releaseCleanup must call primaryRegistry.forgetDelegation(...) — an inert lambda "
+ "leaves the stale delegation in place");
}
private static void awaitWaiting(Rendezvous rendezvous) throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
}
}
@@ -0,0 +1,49 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.net.ServerSocket;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 3, site {@code :512}. {@code Fleetd.main} passes {@code AmqpReplyInbox::open}
* as the {@link Fleetd.AmqpOpener} argument to {@code selectReplyInbox(...)} — before this ticket
* that method reference was inline at the call site. Measured: replacing it with the inert {@code
* (uri, prefetch) -> new InMemoryReplyInbox()} compiles with 0 errors and leaves the full suite
* green — {@code FleetdReplyInboxSelectionTest} drives {@code selectReplyInbox} with its own
* injected opener (including one test that passes the real {@code AmqpReplyInbox::open}
* explicitly) and never observes what {@code main} itself actually passes.
*
* <p>This test calls {@link Fleetd#replyInboxOpener} directly and proves it is the real, network-
* attempting opener rather than a stub: pointed at a guaranteed-closed local port, it must throw —
* the same shape {@code FleetdReplyInboxSelectionTest.aRealUnreachableBrokerFallsBackViaTheRealOpener}
* already relies on for {@code AmqpReplyInbox.open} itself. The inert form never attempts a
* connection and never throws, so it fails this assertion silently (by returning normally).
*/
class FleetdReplyInboxOpenerWiringTest {
@Test
@DisplayName("Fleetd.replyInboxOpener is the real AmqpReplyInbox::open, not a stub that never connects")
void replyInboxOpenerAttemptsARealConnection() throws Exception {
int closedPort;
try (ServerSocket socket = new ServerSocket(0)) {
closedPort = socket.getLocalPort();
} // released immediately — connecting to it now is a guaranteed refusal, not a fluke
Fleetd.AmqpOpener opener = Fleetd.replyInboxOpener();
IllegalStateException thrown = assertThrows(IllegalStateException.class,
() -> opener.open("amqp://user:pw@127.0.0.1:" + closedPort + "/vh", 50),
"Fleetd.replyInboxOpener() must be AmqpReplyInbox::open — a real network attempt "
+ "against a genuinely unreachable broker must throw. The inert form "
+ "(uri, prefetch) -> new InMemoryReplyInbox() never attempts a connection "
+ "and never throws, so it would return normally here and fail this "
+ "assertion silently.");
assertTrue(thrown.getMessage().contains("cannot connect to AMQP broker"),
"must be AmqpReplyInbox.open's own real failure message, not a different exception "
+ "shape standing in for it");
}
}
@@ -0,0 +1,72 @@
package dev.ltms.fleet;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.inject.TurnRegistrar;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #589 Group 3, site {@code :505}. {@code Fleetd.main} wires the {@link
* dev.ltms.fleet.inject.Injector}'s {@link TurnRegistrar} with {@code completion::register} —
* before this ticket that was an inline argument to {@code new Injector(...)}, so nothing could
* pin it directly. Measured: replacing it with {@link TurnRegistrar#NOOP} at the call site
* compiles with 0 errors and leaves the full suite green, because {@code onDelivered}'s own {@code
* captureBaseline} performs the identical {@code inFlight} check-and-put a moment later on the
* ordinary delivery path — the two are indistinguishable unless something reads the resolver
* between {@code register} and {@code onDelivered}, or {@code onDelivered} never runs at all (the
* gap fleetd #556 introduced {@link TurnRegistrar} to close).
*
* <p>This test calls {@link Fleetd#turnRegistrar} directly — never {@code Injector} or {@code
* main} — and drives {@link CompletionResolver} entirely through its public API: {@link
* TurnRegistrar#register} followed by {@link CompletionResolver#resolveBeforePostAction}, which
* looks up the same {@code inFlight} entry {@code onTurnComplete} would. With the real registrar,
* that entry exists and the waiter opened by {@link Rendezvous#open} resolves; with {@link
* TurnRegistrar#NOOP} nothing was ever registered, {@code resolveBeforePostAction} finds no
* in-flight turn, and the waiter is left exactly as it started — never done.
*/
class FleetdTurnRegistrarWiringTest {
private static final String T = "term_a";
@Test
@DisplayName("Fleetd.turnRegistrar delegates to the real CompletionResolver, not a no-op")
void turnRegistrarDelegatesToCompletionRegister() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ BUILD GREEN: 391 files\n❯ ");
AgentControl agents = new AgentControl(herdr);
Rendezvous rendezvous = new Rendezvous();
// fleetd#164: an ever-advancing fake clock stands in for the real time a turn would take
// between delivery and resolution, so the MIN_TURN_NANOS "too fast" floor never trips here —
// see MessageServiceTest's resolverClock for the same technique.
AtomicLong clock = new AtomicLong();
CompletionResolver completion = new CompletionResolver(agents, rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(),
() -> clock.addAndGet(CompletionResolver.MIN_TURN_NANOS + 1));
TurnRegistrar registrar = Fleetd.turnRegistrar(completion);
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(T);
registrar.register(T, new TurnToken(T, waiter));
// Mirrors what Injector.onStatus's confirmed working->idle boundary would trigger via
// CompletionResolver.onTurnComplete — resolveBeforePostAction is the public, synchronous
// twin of that path and reads the exact same inFlight entry register() must have written.
completion.resolveBeforePostAction(T);
assertTrue(waiter.isDone(),
"Fleetd.turnRegistrar(completion) must return completion::register — replacing it "
+ "with TurnRegistrar.NOOP at the Fleetd.turnRegistrar call site means this "
+ "turn is never registered with CompletionResolver, so resolveBeforePostAction "
+ "finds no in-flight turn and this waiter is never resolved");
}
}
@@ -0,0 +1,248 @@
package dev.ltms.fleet.lead;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.io.RandomAccessFile;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assumptions.assumeFalse;
/**
* Ticket "lead context gauge" — fleetd could not see how full a lead's own Claude Code context
* window was, so a lead auto-compacting was always a surprise. {@link LeadContextGauge} reads that
* off the lead's own transcript. Five acceptance properties from the ticket, one test class each
* (plus the three separate UNKNOWN cases the ticket calls out by name):
* <ol>
* <li>{@link #tokenCountTracksTheLastUsageRecordAndChangesWithIt()}</li>
* <li>{@link #compactionCountTracksCompactBoundaryRecordsAndChangesWithIt()}</li>
* <li>{@link #missingFileIsUnknown()}, {@link #unreadableFileIsUnknown()}</li>
* <li>{@link #tornFinalLineFallsBackToTheLastGoodReading()},
* {@link #everyLineUnparseableIsUnknown()} (fleetd #602 gauge-wiring, finding 2)</li>
* <li>{@link #readNeverExceedsTheTailBound()}</li>
* <li>{@link #secondReadInsideTtlDoesNotTouchDiskAgain()}</li>
* </ol>
*
* <p>Every fixture lives under {@code @TempDir} — never the operator's real config directory (see
* the ticket's hard constraint on this).
*
* <p><strong>fleetd #602 gauge-wiring, finding 2.</strong> The property above used to be "a
* truncated or invalid LAST line reports UNKNOWN" — on the theory that a bad last line signals a
* format change. That reasoning did not hold: {@code fleet_list} reads this transcript while Claude
* Code may be mid-write on it, so a torn LAST line is an ordinary race, not a format change, and a
* real format change makes EVERY line unparseable, not only the last one written. So the property
* is now split in two: {@link #tornFinalLineFallsBackToTheLastGoodReading} (the torn-line case must
* NOT destroy a good earlier reading) and its control, {@link #everyLineUnparseableIsUnknown} (only
* when NOTHING in the window parses does this report UNKNOWN).
*/
class LeadContextGaugeTest {
private static final String SESSION_ID = "11111111-1111-1111-1111-111111111111";
/** One line Claude Code would write for a turn with the given live-context total. */
private static String usageLine(long inputTokens, long cacheRead, long cacheCreation) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + inputTokens + ","
+ "\"cache_read_input_tokens\":" + cacheRead + ","
+ "\"cache_creation_input_tokens\":" + cacheCreation + "}}}";
}
/** One line Claude Code writes when an auto-compaction happens. */
private static String compactionLine() {
return "{\"type\":\"system\",\"subtype\":\"compact_boundary\","
+ "\"compactMetadata\":{\"preTokens\":250000,\"postTokens\":5000,"
+ "\"trigger\":\"auto\",\"durationMs\":54000}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} and writes {@code lines}. */
private static String writeTranscript(Path configDir, String sessionId, String... lines) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Path file = projectDir.resolve(sessionId + ".jsonl");
StringBuilder sb = new StringBuilder();
for (String line : lines) {
sb.append(line).append('\n');
}
Files.writeString(file, sb.toString(), StandardCharsets.UTF_8);
return configDir.toString();
}
// --- property 1: live token count tracks the last usage record -----------------------------
@Test
@DisplayName("reported tokens equal the LAST usage record's total, and change when it does")
void tokenCountTracksTheLastUsageRecordAndChangesWithIt(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID,
usageLine(1_000, 0, 0),
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
assertEquals(LeadContextGauge.State.OK, first.state());
// Change N in the fixture (a fresh session id avoids the cache) -- the reported number
// must change with it, not stay pinned to the first fixture's total.
String otherSession = "22222222-2222-2222-2222-222222222222";
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
assertEquals(200_000L, second.tokens());
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
}
// --- property 2: compaction count tracks compact_boundary records --------------------------
@Test
@DisplayName("reported compaction count equals K compact_boundary records, and changes with K")
void compactionCountTracksCompactBoundaryRecordsAndChangesWithIt(@TempDir Path tmp) throws IOException {
String sessionTwoCompactions = "33333333-3333-3333-3333-333333333333";
writeTranscript(tmp, sessionTwoCompactions,
usageLine(1_000, 0, 0),
compactionLine(),
usageLine(2_000, 0, 0),
compactionLine(),
usageLine(3_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
assertEquals(2, twoCompactions.compactions());
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
}
// --- property 3: two separate UNKNOWN cases -------------------------------------------------
@Test
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
void missingFileIsUnknown(@TempDir Path tmp) {
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
}
@Test
@DisplayName("an unreadable transcript file reports UNKNOWN with no token number")
void unreadableFileIsUnknown(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
Path file = tmp.resolve("projects").resolve("some-project-slug").resolve(SESSION_ID + ".jsonl");
assertTrue(file.toFile().setReadable(false), "test setup: must be able to revoke read permission");
try {
// setReadable(false) really did clear the read bit (asserted above), but that alone
// does not prove the file is UNREADABLE: running as root (e.g. a CI container) ignores
// the read bit and opens the file anyway. Files.isReadable checks what actually happens
// on open, not the bit. When it still reports readable, this test cannot create the
// condition it needs on this machine, so it skips honestly instead of asserting on a
// state that was never reached. A skip here means "I could not set up the case", NOT
// "the UNKNOWN behaviour is fine" -- it is not evidence either way.
assumeFalse(Files.isReadable(file),
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
} finally {
file.toFile().setReadable(true); // so @TempDir cleanup can delete it
}
}
@Test
@DisplayName("fleetd #602 finding 2: a torn final line does not destroy a good earlier reading")
void tornFinalLineFallsBackToTheLastGoodReading(@TempDir Path tmp) throws IOException {
// Deliberately NOT via writeTranscript: that helper appends a trailing newline after every
// line, including the last one, which would not model what a write caught mid-flush looks
// like on disk. Claude Code appends and flushes one line at a time, so the torn line here has
// no trailing newline at all -- exactly the shape a read racing an in-progress write sees.
Path projectDir = tmp.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Path file = projectDir.resolve(SESSION_ID + ".jsonl");
String lastCompleteLine = usageLine(1_000, 2_000, 3_000); // last COMPLETE record: 6,000
String tornLine = "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{\"input_tok";
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.OK, reading.state(),
"a torn final line must not turn a good earlier reading into UNKNOWN");
assertEquals(6_000L, reading.tokens(),
"must report the last COMPLETE line's total, not fail the whole read");
}
@Test
@DisplayName("fleetd #602 finding 2 control: when EVERY line is unparseable, the reading IS UNKNOWN")
void everyLineUnparseableIsUnknown(@TempDir Path tmp) throws IOException {
// The real signal a format change gives: not one bad line, but ALL of them. Without this
// control, code that never reports UNKNOWN at all would still pass the torn-line property.
String configDir = writeTranscript(tmp, SESSION_ID,
"{this is not json at all",
"neither is this{{{");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"every line unparseable is the real format-change signal and must still report UNKNOWN");
assertNull(reading.tokens());
}
// --- property 4: the tail bound is actually enforced ----------------------------------------
@Test
@DisplayName("reading a transcript far larger than the tail bound reads no more than the bound")
void readNeverExceedsTheTailBound(@TempDir Path tmp) throws IOException {
int smallBound = 4_096; // exercise the mechanism without writing a multi-MB fixture
Path file = tmp.resolve("big.jsonl");
// One line far bigger than smallBound, repeated, so the whole file is many times the bound.
String line = usageLine(42, 0, 0) + " ".repeat(500);
StringBuilder sb = new StringBuilder();
for (int i = 0; i < 50; i++) {
sb.append(line).append('\n');
}
Files.writeString(file, sb.toString(), StandardCharsets.UTF_8);
long fileLength = Files.size(file);
assertTrue(fileLength > (long) smallBound * 5, "test setup: fixture must genuinely dwarf the bound");
var tail = LeadContextGauge.tailBytes(file, smallBound);
assertEquals(smallBound, tail.bytes().length,
"must read exactly the bound, not the whole " + fileLength + "-byte file — assert on bytes actually read");
}
// --- property 5: the cache TTL means a burst of calls reads the file once -------------------
@Test
@DisplayName("the same (configDir, sessionId) read twice inside the TTL touches disk once")
void secondReadInsideTtlDoesNotTouchDiskAgain(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
AtomicLong now = new AtomicLong(0);
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
gauge.read(configDir, SESSION_ID, "claude");
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
now.set(6_000); // past the TTL
gauge.read(configDir, SESSION_ID, "claude");
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
}
// --- non-Claude peer / unresolved session -----------------------------------------------------
@Test
@DisplayName("a non-claude agent type or a null session id is UNKNOWN, never OK")
void nonClaudePeerOrUnresolvedSessionIsUnknown(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
}
}
@@ -16,6 +16,7 @@ import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
@@ -1004,6 +1005,324 @@ class LeadRolloverTest {
}
}
// ---- CB-... : LeadRollover#status makes the outcome of a confirmed roll readable -----------
@Test
@DisplayName("[STATUS 1] after the calling turn never settles, status() reports "
+ "TURN_NEVER_SETTLED for that token — this assertion could not even be written before "
+ "status() existed")
void statusReportsTurnNeverSettledAfterTheRollIsAbandoned() throws IOException {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("working"); // the calling lead's own pane — never goes idle in this test
Path handover = writeHandover("handover contents");
FleetConfig.LeadRollover config =
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 1 /*turnSettleSeconds*/, 20, "text");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(herdr, config, () -> clock.addAndGet(500));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "every synchronous gate should pass; the refusal happens "
+ "only inside the deferred continuation, which this test's synchronous runner has "
+ "already run to completion by the time confirm() returns");
assertEquals(0, promptCallCount(herdr), "sanity: /clear was never sent");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.TURN_NEVER_SETTLED, status.state());
assertTrue(status.detail().contains("turnSettleSeconds"), "the detail must name the knob a "
+ "reader needs to raise: " + status.detail());
assertTrue(status.detail().toLowerCase().contains("no /clear"), "the detail must say plainly "
+ "that no /clear was ever sent: " + status.detail());
}
@Test
@DisplayName("[STATUS 2] a roll that completes reports ROLLED for its token")
void statusReportsRolledForACompletedRoll() throws IOException {
FakeHerdr herdr = new FakeHerdr(); // default idle — a full successful roll
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
assertEquals(2, promptCallCount(herdr), "sanity: /clear then bootstrapText were both sent");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.ROLLED, status.state());
assertNotNull(status.detail());
}
@Test
@DisplayName("[STATUS 3] a roll where /clear never settles reports CLEAR_NEVER_SETTLED — "
+ "distinct from both ROLLED and TURN_NEVER_SETTLED")
void statusReportsClearNeverSettledDistinctFromTheOtherTwoStates() throws IOException {
// Idle until /clear is sent, then permanently working — the SECOND wait never settles.
FakeHerdr fake = new FakeHerdr();
HerdrClient flipsAfterClear = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) throws HerdrException {
JsonNode result = fake.call(method, params);
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
fake.agentStatus("working");
}
return result;
}
@Override
public void close() {
fake.close();
}
};
Path handover = writeHandover("handover contents");
FleetConfig.LeadRollover config =
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*clearSettleSeconds*/, "boot text");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(flipsAfterClear, config, () -> clock.addAndGet(500));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged only, "
+ "deep inside the deferred continuation");
assertEquals(1, promptCallCount(fake), "sanity: only /clear was sent, never bootstrapText");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.CLEAR_NEVER_SETTLED, status.state());
assertNotEquals(LeadRollover.RollState.ROLLED, status.state());
assertNotEquals(LeadRollover.RollState.TURN_NEVER_SETTLED, status.state());
assertTrue(status.detail().contains("clearSettleSeconds"), status.detail());
}
@Test
@DisplayName("[STATUS 4] a token that was never issued, or was cancelled, gives a clean "
+ "UNKNOWN answer rather than an exception or a false ROLLED")
void statusOnUnissuedOrCancelledTokenIsCleanNotAnExceptionOrFalseSuccess() throws IOException {
FakeHerdr herdr = new FakeHerdr();
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
// never issued at all
LeadRollover.RollStatus neverIssued = assertDoesNotThrow(() -> rollover.status("no-such-token"));
assertEquals(LeadRollover.RollState.UNKNOWN, neverIssued.state());
assertNotEquals(LeadRollover.RollState.ROLLED, neverIssued.state());
// null / blank must not throw either — pending is a ConcurrentHashMap, which throws on a
// null-key lookup unless status() guards it first
assertEquals(LeadRollover.RollState.UNKNOWN, assertDoesNotThrow(() -> rollover.status(null)).state());
assertEquals(LeadRollover.RollState.UNKNOWN, assertDoesNotThrow(() -> rollover.status(" ")).state());
// opened, then cancelled
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
assertTrue(rollover.cancel(pending.token()));
LeadRollover.RollStatus cancelled = assertDoesNotThrow(() -> rollover.status(pending.token()));
assertEquals(LeadRollover.RollState.UNKNOWN, cancelled.state());
assertNotEquals(LeadRollover.RollState.ROLLED, cancelled.state());
}
@Test
@DisplayName("[STATUS 5] the bounded outcome history never grows past its cap")
void statusHistoryDoesNotGrowPastItsCap() throws IOException {
FakeHerdr herdr = new FakeHerdr(); // default idle — every roll completes
Path handover = writeHandover("handover contents"); // one file, reused by every roll below —
// checkHandover only compares its mtime against each open()'s OWN requestedAtMillis (the
// fake clock, in the low thousands), and the file's real (wall-clock) mtime is always far
// larger than that, so freshness passes on every iteration without rewriting the file.
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), () -> clock.addAndGet(1));
int rolls = LeadRollover.OUTCOME_HISTORY_CAP + 50;
String[] tokens = new String[rolls];
for (int i = 0; i < rolls; i++) {
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full #" + i);
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "roll #" + i + " should have been approved: " + decision.detail());
tokens[i] = pending.token();
}
assertEquals(LeadRollover.RollState.ROLLED, rollover.status(tokens[rolls - 1]).state(),
"the most recently finished roll's outcome must still be in the bounded history");
// There is no direct size accessor for the bounded history, so boundedness is asserted
// indirectly and behaviourally: the OLDEST finished roll's outcome must have been evicted
// (reads back as a clean UNKNOWN, exactly like a token that was never issued) once more than
// OUTCOME_HISTORY_CAP rolls have gone through this instance. If the cap were not enforced,
// tokens[0] would still read back ROLLED here, and this assertion would fail.
assertEquals(LeadRollover.RollState.UNKNOWN, rollover.status(tokens[0]).state(),
"the oldest finished roll's outcome must have been evicted once the cap was "
+ "exceeded — otherwise the bounded history is not actually bounded");
}
@Test
@DisplayName("[STATUS 6] a status() call against a pending roll is read-only — no herdr call is "
+ "made and the pending request is left untouched")
void statusCallAgainstAPendingRollIsReadOnlyAndTouchesNothing() throws IOException {
FakeHerdr herdr = new FakeHerdr();
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.PENDING, status.state());
assertEquals(0, herdr.calls.size(), "status() must never make any herdr call at all — not "
+ "just no agent.prompt — since it must never schedule, cancel, or retry anything");
// the pending request must be left exactly as it was: the SAME token can still be confirmed
// afterwards, as if status() had never been called.
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "status() must not have consumed or otherwise disturbed the "
+ "pending request: " + decision.reason() + " / " + decision.detail());
}
// ---- PR #600 review round 2: IN_PROGRESS — the gap between confirm() handing off and the ----
// ---- continuation finishing must never read back as UNKNOWN ("nothing was ever requested") --
/**
* A {@code continuationRunner} that CAPTURES the roll instead of running it, so a test can
* observe {@link LeadRollover#status} in the window between {@link LeadRollover#confirm}
* handing off and the roll actually finishing — the window the synchronous {@code
* Runnable::run} runner used everywhere else in this class collapses to nothing. Call {@link
* #runNext()} to finish exactly one held roll, once the test is done observing the in-flight
* state.
*/
private static final class HoldingRunner implements java.util.function.Consumer<Runnable> {
private final List<Runnable> held = new ArrayList<>();
@Override
public void accept(Runnable runnable) {
held.add(runnable);
}
int heldCount() {
return held.size();
}
/** Runs (and removes) the oldest held roll — FIFO, matching confirm() call order. */
void runNext() {
held.remove(0).run();
}
}
private static LeadRollover newRolloverWithHoldingRunner(HerdrClient herdr,
FleetConfig.LeadRollover config, LongSupplier nowMillis, HoldingRunner runner) {
AgentControl agents = new AgentControl(herdr);
return new LeadRollover(agents, () -> config, _ -> null, nowMillis, () -> { }, runner);
}
@Test
@DisplayName("[IN-PROGRESS 1] between an approved confirm() and the continuation finishing, "
+ "status() reports IN_PROGRESS — not UNKNOWN, and not PENDING")
void statusReportsInProgressBetweenConfirmAndTheContinuationFinishing() throws IOException {
FakeHerdr herdr = new FakeHerdr(); // default idle — the held roll WOULD complete once run
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
HoldingRunner runner = new HoldingRunner();
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
fixedClock(clock), runner);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
assertEquals(1, runner.heldCount(), "sanity: the roll must have been handed to the "
+ "continuation runner and held there, not run yet");
assertEquals(0, promptCallCount(herdr), "sanity: the held continuation has not run, so "
+ "nothing has been sent to the pane yet");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.IN_PROGRESS, status.state(),
"a lead calling status() right after confirm() returned approved, while the roll "
+ "is still running, must be told IN_PROGRESS — not UNKNOWN (\"nothing was "
+ "ever requested\", which would wrongly invite it to call open() again "
+ "mid-roll) and not PENDING (\"not yet approved\", which is simply false "
+ "here): got " + status.state() + " / " + status.detail());
}
@Test
@DisplayName("[IN-PROGRESS 2] once the continuation finishes, the same token reports its "
+ "terminal state — IN_PROGRESS is not sticky")
void inProgressStateIsNotStickyOnceTheContinuationFinishes() throws IOException {
FakeHerdr herdr = new FakeHerdr(); // default idle — the held roll completes once run
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
HoldingRunner runner = new HoldingRunner();
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
fixedClock(clock), runner);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
assertEquals(LeadRollover.RollState.IN_PROGRESS, rollover.status(pending.token()).state(),
"sanity: must be IN_PROGRESS before the held continuation is run");
runner.runNext(); // finish the held roll now
assertEquals(2, promptCallCount(herdr), "sanity: the roll actually ran to completion "
+ "once released — /clear then bootstrapText");
LeadRollover.RollStatus after = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.ROLLED, after.state(),
"the SAME token must now report its terminal state — IN_PROGRESS must not still be "
+ "reported once the roll has actually finished");
}
@Test
@DisplayName("[IN-PROGRESS 4] IN_PROGRESS is distinct from every other RollState")
void inProgressStateIsDistinctFromAllOtherStates() throws IOException {
FakeHerdr herdr = new FakeHerdr();
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
HoldingRunner runner = new HoldingRunner();
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
fixedClock(clock), runner);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
LeadRollover.RollState inProgress = rollover.status(pending.token()).state();
assertEquals(LeadRollover.RollState.IN_PROGRESS, inProgress);
for (LeadRollover.RollState other : LeadRollover.RollState.values()) {
if (other == LeadRollover.RollState.IN_PROGRESS) {
continue;
}
assertNotEquals(other, inProgress, "IN_PROGRESS must be distinct from " + other);
}
}
@Test
@DisplayName("[IN-PROGRESS 5] eviction counts IN_PROGRESS entries toward the cap exactly like "
+ "finished ones — a burst of confirmed-but-not-yet-finished rolls still ages out")
void evictionCountsInProgressEntriesTowardTheCap() throws IOException {
FakeHerdr herdr = new FakeHerdr();
Path handover = writeHandover("handover contents"); // reused by every roll — see
// statusHistoryDoesNotGrowPastItsCap for why one shared file is enough for freshness.
AtomicLong clock = new AtomicLong(1_000);
HoldingRunner runner = new HoldingRunner(); // nothing run below — every roll stays IN_PROGRESS
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
() -> clock.addAndGet(1), runner);
int rolls = LeadRollover.OUTCOME_HISTORY_CAP + 50;
String[] tokens = new String[rolls];
for (int i = 0; i < rolls; i++) {
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full #" + i);
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "roll #" + i + " should have been approved: " + decision.detail());
tokens[i] = pending.token();
}
assertEquals(rolls, runner.heldCount(), "sanity: none of these rolls have been run — every "
+ "one of them is sitting in outcomes as IN_PROGRESS, not a separate uncapped map");
assertEquals(LeadRollover.RollState.IN_PROGRESS, rollover.status(tokens[rolls - 1]).state(),
"the most recently confirmed (still in-flight) roll must still be in the bounded "
+ "history");
assertEquals(LeadRollover.RollState.UNKNOWN, rollover.status(tokens[0]).state(),
"the oldest confirmed roll's IN_PROGRESS entry must have been evicted once the cap "
+ "was exceeded, exactly like a finished entry would be — proving IN_PROGRESS "
+ "entries share the SAME bounded map and count against the SAME cap, rather "
+ "than living in a second, uncapped in-flight map");
}
@Test
@DisplayName("[fleetd #494 follow-up] the turn-settle timeout warn line prints the MEASURED "
+ "elapsed time next to the configured budget, never the configured value alone")
@@ -1040,4 +1359,92 @@ class LeadRolloverTest {
+ "it were the measured wait duration: " + message);
}
}
// ---- fleetd #615: a HerdrException out of either unwrapped agents.send call must leave a ----
// ---- TERMINAL FAILED outcome, never a stuck IN_PROGRESS ---------------------------------------
@Test
@DisplayName("[fleetd #615 — 1] send() throwing on the /clear call leaves status(token) "
+ "reporting FAILED, not stuck at IN_PROGRESS")
void sendThrowingOnClearLeavesStatusReportingFailed() throws IOException {
FakeHerdr fake = new FakeHerdr(); // default idle — the turn-settle wait passes immediately
HerdrClient throwsOnClear = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) throws HerdrException {
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
throw new HerdrException("simulated herdr transport failure sending /clear");
}
return fake.call(method, params);
}
@Override
public void close() {
fake.close();
}
};
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(throwsOnClear, cfg(handover.toString()), fixedClock(clock));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
+ "inside the deferred continuation, which this test's synchronous runner has "
+ "already run to completion by the time confirm() returns");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.FAILED, status.state(),
"a HerdrException out of the /clear send must leave a TERMINAL FAILED outcome — "
+ "before fleetd #615's fix, the continuation thread died silently and "
+ "status() was stuck reporting the IN_PROGRESS confirm() wrote at hand-off, "
+ "forever: got " + status.state() + " / " + status.detail());
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
+ "so an operator reading status() has something to act on: " + status.detail());
}
@Test
@DisplayName("[fleetd #615 — 2] send() throwing on the bootstrap-text call (after /clear "
+ "succeeded and the pane settled) also leaves status(token) reporting FAILED — a "
+ "DIFFERENT exit from the /clear-throw case above")
void sendThrowingOnBootstrapTextLeavesStatusReportingFailed() throws IOException {
FakeHerdr fake = new FakeHerdr(); // default idle throughout — both settle waits pass promptly
HerdrClient throwsOnBootstrapText = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) throws HerdrException {
if ("agent.prompt".equals(method) && String.valueOf(params).contains("read the handover file")) {
throw new HerdrException("simulated herdr transport failure sending bootstrapText");
}
return fake.call(method, params);
}
@Override
public void close() {
fake.close();
}
};
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(throwsOnBootstrapText, cfg(handover.toString()), fixedClock(clock));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
+ "inside the deferred continuation, which this test's synchronous runner has "
+ "already run to completion by the time confirm() returns");
assertEquals(1, promptCallCount(fake), "sanity: /clear was sent and settled — only the "
+ "SECOND agent.prompt call (bootstrapText) threw");
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.FAILED, status.state(),
"a HerdrException out of the bootstrapText send — a DIFFERENT exit from the /clear "
+ "throw, reached only after /clear already succeeded and the pane already "
+ "settled — must also leave a TERMINAL FAILED outcome, not a stuck "
+ "IN_PROGRESS: got " + status.state() + " / " + status.detail());
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
+ "so an operator reading status() has something to act on: " + status.detail());
}
}
@@ -258,5 +258,51 @@ class FleetMcpHandoverTest {
assertTrue(textOf(r).contains("\"cancelled\":true"), textOf(r));
}
// --- new: action "status" — makes the outcome of a confirm() readable through the tool ----
@Test
@DisplayName("status on a null LeadRollover is a clean NOT_CONFIGURED refusal, never a throw")
void statusWithNullLeadRolloverRefusesCleanly() {
McpSchema.CallToolResult r = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD,
Map.of("action", "status", "token", "whatever")));
assertFalse(r.isError());
assertTrue(textOf(r).contains("NOT_CONFIGURED"), textOf(r));
}
@Test
@DisplayName("status on a token that was never opened reports UNKNOWN")
void statusOnUnknownTokenReportsUnknown() {
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
Map.of("action", "status", "token", "does-not-exist"));
assertFalse(r.isError());
assertTrue(textOf(r).contains("\"state\":\"UNKNOWN\""), textOf(r));
}
@Test
@DisplayName("status on a token that is still pending (opened, not confirmed) reports PENDING")
void statusOnPendingTokenReportsPending() {
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
String token = extractToken(textOf(FleetMcp.handover(rollover, LEAD, Map.of("action", "open"))));
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
Map.of("action", "status", "token", token));
assertFalse(r.isError());
assertTrue(textOf(r).contains("\"state\":\"PENDING\""), textOf(r));
}
@Test
@DisplayName("status is registered on the tool's schema and the schema still names no caller-identity parameter")
void statusActionIsAdvertisedOnTheSchema() {
FleetMcp m = mcp(null);
McpSchema.Tool tool = m.registeredTools().stream()
.filter(t -> "fleet_handover".equals(t.name()))
.findFirst()
.orElseThrow(() -> new AssertionError("fleet_handover was not registered"));
assertTrue(tool.description().contains("'status'"),
"the tool's own description must advertise the 'status' action: " + tool.description());
}
// --- acceptance 7 (wiring) is covered by FleetdLeadRolloverWiringTest, unchanged -----------
}
@@ -0,0 +1,126 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.session.SessionManager;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #602 gauge-wiring: {@code FleetMcp}'s {@code fleet_list} handler shipped calling
* {@code LeadContextGauge.read(null, ...)} unconditionally, so every lead whose profile sets a
* {@code configDir:} override read the wrong transcript directory forever, with no error anywhere —
* {@link LeadContextGauge}'s own 8 tests all passed because none of them exercised THIS wiring; each
* one hands the gauge its own directory directly.
*
* <p>This class proves the directory {@code fleet.leaders.<name>.profile}'s own {@code configDir:}
* names is the one {@code fleet_list}'s {@code context} row actually reads — not merely that SOME
* directory got passed. If the wiring in {@code FleetMcp.leadView}/{@code contextView} is ever
* reverted to a hardcoded {@code null}, {@link #configuredDirectoryDecidesWhichTranscriptIsRead()}
* must go red: both directories below hold a REAL, DIFFERENT token count for the SAME session id,
* so a hardcoded {@code null} (always reading the built-in default, where neither transcript lives)
* would report {@code UNKNOWN} both times instead of the two distinct numbers this test asserts.
*/
class FleetMcpLeadContextGaugeWiringTest {
private static final String LEAD_TERMINAL = "term_lead";
private static final String LEAD_NAME = "opus";
/** {@code FakeHerdr.withAgent} always projects {@code agent_session.value} as {@code sess-<terminalId>}. */
private static final String LEAD_SESSION_ID = "sess-" + LEAD_TERMINAL;
private static ClaudeCodeLauncher workerService(FakeHerdr h) {
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
/** One line Claude Code would write for a turn with the given live-context total. */
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
private static String writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
return configDir.toString();
}
@Test
@DisplayName("the CONFIGURED directory decides which transcript is read, and follows a config change")
void configuredDirectoryDecidesWhichTranscriptIsRead(@TempDir Path tmp) throws IOException {
Path dirA = Files.createDirectory(tmp.resolve("dir-a"));
Path dirB = Files.createDirectory(tmp.resolve("dir-b"));
writeTranscript(dirA, LEAD_SESSION_ID, 11_000);
writeTranscript(dirB, LEAD_SESSION_ID, 22_000);
FakeHerdr herdr = new FakeHerdr().withAgent("lead-opus", LEAD_TERMINAL, "wL:p1", "wL:t1");
SessionManager sessions = new SessionManager(workerService(herdr));
LeadContextGauge contextGauge = new LeadContextGauge();
AtomicReference<String> configuredDir = new AtomicReference<>(dirA.toString());
FleetMcp.LeadConfigDirSource source = new FleetMcp.LeadConfigDirSource(name ->
LEAD_NAME.equals(name) ? configuredDir.get() : null);
String firstRead = textOf(listFleet(herdr, sessions, contextGauge, source));
assertTrue(firstRead.contains("\"tokens\":11000"),
"config names directory A ⇒ fleet_list must report A's token count: " + firstRead);
// The config changes to name directory B instead -- the very next read must follow it.
configuredDir.set(dirB.toString());
String secondRead = textOf(listFleet(herdr, sessions, contextGauge, source));
assertTrue(secondRead.contains("\"tokens\":22000"),
"config now names directory B ⇒ fleet_list must report B's token count, not A's stale one: "
+ secondRead);
}
@Test
@DisplayName("a lead whose config names no directory still degrades to the built-in default, never throws")
void aLeadWhoseConfigNamesNoDirectoryStillFallsBackWithoutThrowing() {
FakeHerdr herdr = new FakeHerdr().withAgent("lead-opus", LEAD_TERMINAL, "wL:p1", "wL:t1");
SessionManager sessions = new SessionManager(workerService(herdr));
LeadContextGauge contextGauge = new LeadContextGauge();
McpSchema.CallToolResult res = listFleet(herdr, sessions, contextGauge, FleetMcp.LeadConfigDirSource.none());
assertNotEquals(Boolean.TRUE, res.isError(), "a lead with no configured configDir must degrade, never throw");
String out = textOf(res);
assertTrue(out.contains("\"state\":\"unknown\""), "no transcript under the built-in default ⇒ UNKNOWN: " + out);
assertFalse(out.contains("\"tokens\""), "UNKNOWN must never carry a stale/default token number: " + out);
}
private static McpSchema.CallToolResult listFleet(FakeHerdr herdr, SessionManager sessions,
LeadContextGauge contextGauge, FleetMcp.LeadConfigDirSource leadConfigDirs) {
return FleetMcp.listFleet(workerService(herdr), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs,
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, FleetMcp.CoordinationSource.none(), false);
}
}
@@ -1,6 +1,11 @@
package dev.ltms.fleet.msg;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.MemberSession;
@@ -8,20 +13,29 @@ import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for the CB-551 idle-lead heartbeat: the pure {@link LeadHeartbeatLoop#decide} decision
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot.
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot. Also covers
* fleetd #609's context-high notice: the {@code context}/{@code contextNotified} parameters {@code
* decide} gained, and the standalone {@link LeadHeartbeatLoop#contextNotice} text builder.
*
* <p>All decision tests call {@code decide} directly with explicit nanoTime values from an injected
* clock — no sleeping, no scheduler races. This mirrors how {@code ReplyPushLoopTest} pins the pure
* decision before exercising the loop.
*
* <p>Every pre-#609 test below passes {@code LeadContextGauge.State.UNKNOWN, false} for the two new
* {@code decide} parameters — the same as a lead whose context could not be read and has never been
* notified — so each one still pins exactly the behaviour it pinned before this ticket.
*/
class LeadHeartbeatLoopTest {
@@ -56,12 +70,18 @@ class LeadHeartbeatLoopTest {
return new LeadHeartbeatLoop.FleetState(2, 1, 3, List.of("term_a", "term_b"));
}
/** The loop under test; the scheduler is never invoked on the pure decide path. */
/** The loop under test, context notice off; the scheduler is never invoked on the pure decide path. */
private static LeadHeartbeatLoop loop(int quietCap) {
return loop(quietCap, false);
}
/** As above, with the fleetd #609 {@code contextHighNudge} flag set explicitly. */
private static LeadHeartbeatLoop loop(int quietCap, boolean contextHighNudge) {
return new LeadHeartbeatLoop(
new PrimaryRegistry("term_lead"), null /*agents — unused on the decide path*/,
null /*inbox*/, List::of, null /*pushLoop*/, null /*scheduler*/, () -> 0L,
IDLE_AFTER_NANOS, 1_000L, quietCap);
IDLE_AFTER_NANOS, 1_000L, quietCap, null,
LeadHeartbeatLoop.LeadContextSource.none(), contextHighNudge);
}
// ── (a) a working lead is never injected ───────────────────────────────────────────────────
@@ -69,7 +89,8 @@ class LeadHeartbeatLoopTest {
@Test
void aWorkingLeadIsNeverInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet());
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
"a WORKING lead is making progress and must not be touched");
assertNull(d.idleSinceNanos(), "a busy lead resets the idle window");
@@ -80,7 +101,8 @@ class LeadHeartbeatLoopTest {
void anUnreadableStatusIsNeverInjectedEither() {
// A failed status read (or a gone agent) must degrade to "do not inject", never hammer the pane.
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet());
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
"never inject into a state the loop cannot read");
}
@@ -90,7 +112,8 @@ class LeadHeartbeatLoopTest {
@Test
void justBecameIdleStartsTheDebounceWindow() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet());
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"the first injectable tick only records the start of the idle stretch");
assertEquals(NOW, d.idleSinceNanos(), "the idle window opens at the moment the lead became injectable");
@@ -99,7 +122,8 @@ class LeadHeartbeatLoopTest {
@Test
void idleWithinQuietPeriodIsNotInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet());
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"a lead idle for 10s (< 300s) has just finished a turn — do not re-prompt it");
}
@@ -109,7 +133,8 @@ class LeadHeartbeatLoopTest {
@Test
void idlePastQuietPeriodIsInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet());
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
"a lead continuously idle past the quiet period is the reason to nudge");
}
@@ -117,14 +142,17 @@ class LeadHeartbeatLoopTest {
@Test
void blockedAndDoneAreInjectableViewsOfIdle() {
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet()).action());
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false).action());
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet()).action());
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false).action());
}
@Test
void nothingIsInjectedWhenNoLeadIsKnown() {
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet());
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"with no known lead there is nobody to nudge — keep waiting until one is discovered");
assertEquals(IDLE_PAST, d.idleSinceNanos(), "the idle window stays open so discovery re-arms it");
@@ -135,7 +163,8 @@ class LeadHeartbeatLoopTest {
@Test
void quietNudgeCapStopsTheLoopWhenNothingIsPending() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet());
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
"3 consecutive nothing-pending nudges have already happened — stop nagging an empty fleet");
}
@@ -145,7 +174,8 @@ class LeadHeartbeatLoopTest {
@Test
void newPendingStateResetsTheQuietCap() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet());
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
"real state appearing re-arms the loop past an exhausted cap");
assertEquals(0, d.quietCount(), "the pending state resets the consecutive-quiet counter");
@@ -156,7 +186,8 @@ class LeadHeartbeatLoopTest {
@Test
void standsDownWhileReplyPushLoopIsActive() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet());
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, false);
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(),
"a second injection would start a competing turn — stand aside instead");
assertEquals(0, d.quietCount(), "the active push is real state, so it re-arms the cap");
@@ -213,4 +244,378 @@ class LeadHeartbeatLoopTest {
assertFalse(fs.hasPending());
assertEquals(0, fs.liveWorkers());
}
// ── fleetd #609: context-high notice ───────────────────────────────────────────────────────
/** One gate scenario, keyed by name, over which the off-state/context-independence property is checked. */
private record Gate(String name, Long idleSince, int quietCount, AgentStatus status,
boolean pushLoopActive, boolean leadKnown, LeadHeartbeatLoop.FleetState fleet) {}
private static List<Gate> allEightGates() {
return List.of(
new Gate("1-standDown", IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet()),
new Gate("2-leadBusy", IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet()),
new Gate("3-justBecameIdle", null, 0, AgentStatus.IDLE, false, true, quietFleet()),
new Gate("4-withinQuietPeriod", IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet()),
new Gate("5-noLeadKnown", IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet()),
new Gate("6-pending", IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet()),
new Gate("7-quietNotExhausted", IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet()),
new Gate("8-quietExhausted", IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet()));
}
@Test
void aOffFlagIgnoresContextAcrossAllEightGates() {
// Property A: with contextHighNudge off, decide() must not depend on the context reading at
// all — not just "usually agrees", but identical Action/idleSinceNanos/quietCount whatever
// context state is passed, and contextNotified must stay false throughout. This is the proof
// that fleetd #609 is opt-in: a daemon upgraded to carry this code, but never configuring
// `contextHighNudge: true`, behaves exactly as it did before this ticket for every one of the
// 8 gates the class javadoc numbers.
LeadHeartbeatLoop offLoop = loop(3, false);
for (Gate g : allEightGates()) {
var withHigh = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.HIGH, false);
var withUnknown = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.UNKNOWN, false);
var withOk = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.OK, false);
assertEquals(withUnknown.action(), withHigh.action(), g.name() + ": action must not depend on context");
assertEquals(withUnknown.idleSinceNanos(), withHigh.idleSinceNanos(), g.name());
assertEquals(withUnknown.quietCount(), withHigh.quietCount(), g.name());
assertEquals(withUnknown.action(), withOk.action(), g.name() + ": nor on an OK reading");
assertFalse(withHigh.contextNotified(), g.name() + ": contextNotified must stay false when the flag is off");
}
}
@Test
void bHighContextFiresEvenPastAnExhaustedQuietCapWithoutSpendingIt() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.HIGH, false);
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
"an idle, quiet, HIGH-context lead is exactly the case the exhausted cap must not swallow");
assertEquals(3, d.quietCount(), "the context notice is an event notice, not a quiet nudge — it must not "
+ "spend or grow the consecutive-quiet budget");
assertTrue(d.contextNotified(), "firing the notice sets the latch");
}
@Test
void cASecondTickWithTheLatchAlreadySetDoesNotTellTheLeadAgain() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.HIGH, true);
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
"the lead was already told once about this HIGH stretch — telling it every tick would be nagging, "
+ "not a notice");
}
@Test
void dFlappingBetweenHighAndUnknownNeverReInjectsWhileLatched() {
LeadHeartbeatLoop on = loop(3, true);
// The lead was already told once (latch = true from a prior HIGH tick).
LeadHeartbeatLoop.Decision afterUnknown = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.UNKNOWN, true);
assertTrue(afterUnknown.contextNotified(),
"UNKNOWN means 'I could not look', not 'it got better' — it must not clear the latch");
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, afterUnknown.action(),
"a still-latched, still-quiet tick must not inject just because the reading is UNKNOWN");
LeadHeartbeatLoop.Decision afterHighAgain = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.HIGH, afterUnknown.contextNotified());
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, afterHighAgain.action(),
"HIGH -> UNKNOWN -> HIGH with the latch already set must never inject again — this is the "
+ "flapping case that would otherwise cost a full lead its remaining turns");
}
@Test
void eAnOkReadingClearsTheLatchSoALaterHighInjectsAgain() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision afterOk = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.OK, true);
assertFalse(afterOk.contextNotified(), "an OK reading clears the latch — the lead's context recovered");
LeadHeartbeatLoop.Decision afterHighAgain = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
LeadContextGauge.State.HIGH, afterOk.contextNotified());
assertEquals(LeadHeartbeatLoop.Action.INJECT, afterHighAgain.action(),
"the cleared latch lets a genuinely new HIGH stretch notify again");
}
@Test
void fStandingDownDoesNotBurnTheOneContextNotice() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet(),
LeadContextGauge.State.HIGH, false);
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(), "ReplyPushLoop is active — stand aside");
assertFalse(d.contextNotified(),
"standing down must not spend the one notice this HIGH stretch gets — the latch stays clear so a "
+ "later tick can still fire it");
}
@Test
void gAWorkingLeadWithHighContextIsStillLeadBusy() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet(),
LeadContextGauge.State.HIGH, false);
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(), "the status gate wins over the context notice");
}
@Test
void hAPendingDrivenInjectWhileHighSetsTheLatchTooSoTheLeadIsNotToldTwice() {
LeadHeartbeatLoop on = loop(3, true);
LeadHeartbeatLoop.Decision d = on.decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet(),
LeadContextGauge.State.HIGH, false);
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action());
assertTrue(d.contextNotified(),
"the pending-driven nudge carries the context notice too (see contextNotice), so it must set the "
+ "latch — otherwise the lead could be told twice by two different routes");
}
// ── contextNotice text builder ─────────────────────────────────────────────────────────────
@Test
void contextNoticeIsEmptyWhenDisabled() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
assertEquals("", LeadHeartbeatLoop.contextNotice(false, reading));
}
@Test
void contextNoticeIsEmptyWhenStateIsOk() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.OK, 1_000L, 0);
assertEquals("", LeadHeartbeatLoop.contextNotice(true, reading));
}
@Test
void contextNoticeIsEmptyWhenStateIsUnknown() {
assertEquals("", LeadHeartbeatLoop.contextNotice(true, LeadContextGauge.Reading.unknown()));
}
@Test
void contextNoticeNamesTheTokenCountWhenHighAndEnabled() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
String notice = LeadHeartbeatLoop.contextNotice(true, reading);
assertFalse(notice.isEmpty());
assertTrue(notice.contains("260771"), notice);
assertTrue(notice.contains("1 compaction"), notice);
assertTrue(notice.contains("fleet_handover"), notice);
assertTrue(notice.contains("operator"), notice);
}
@Test
void contextNoticeOmitsTheTokenClauseRatherThanPrintingNull() {
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, null, 2);
String notice = LeadHeartbeatLoop.contextNotice(true, reading);
assertFalse(notice.isEmpty());
assertFalse(notice.toLowerCase().contains("null"), notice);
assertTrue(notice.contains("2 compactions"), notice);
}
// ── fleetd #609 review: the latch must mean "the notice reached the pane" ────────────────────
//
// These four drive LeadHeartbeatLoop.tick() directly (package-private, same reasoning as
// ReplyPushLoop#tick(String) being directly testable) against a real AgentControl wrapping a
// FailableHerdrClient, so the send path (agents.send -> herdr -> possible throw) is exercised
// for real rather than assumed from decide()'s Decision alone.
private static final String LEAD = "term_lead";
private static final String WORKER = "term_w1";
/** A HIGH reading with a fixed token/compaction count, for the four tests below. */
private static LeadContextGauge.Reading highReading() {
return new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
}
/**
* Builds a real {@link LeadHeartbeatLoop} wired to {@code herdr} via a real {@link AgentControl},
* a mutable fake clock, and a mutable roster so a test can change fleet state between ticks. The
* lead is always reported IDLE by {@code herdr}, so every tick's outcome is governed only by the
* idle-window/quiet-cap/context gates under test.
*/
private static LeadHeartbeatLoop tickableLoop(FailableHerdrClient herdr, AtomicLong now,
List<MemberSession>[] rosterBox, InMemoryReplyInbox inbox,
int quietNudgeCap, ScheduledExecutorService scheduler) {
AgentControl agents = new AgentControl(herdr);
PrimaryRegistry registry = new PrimaryRegistry(LEAD);
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, agents, inbox, scheduler, 5, 100_000);
return new LeadHeartbeatLoop(registry, agents, inbox, () -> rosterBox[0], pushLoop, scheduler,
now::get, IDLE_AFTER_NANOS, 100_000L, quietNudgeCap, null,
new LeadHeartbeatLoop.LeadContextSource(t -> highReading()), true);
}
@Test
void iAFailedSendDoesNotConsumeTheNotice() {
var herdr = new FailableHerdrClient(LEAD);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()}; // quiet: nothing pending
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
loop.tick(); // first injectable tick: only opens the idle window (WAIT_IDLE)
now.addAndGet(TimeUnit.SECONDS.toNanos(400)); // now clearly past the quiet period
herdr.throwOnNextSend();
loop.tick(); // quiet fleet, quiet cap exhausted (0), context HIGH, latch clear -> INJECT, send throws
assertEquals(0, herdr.sentTexts().size(), "the failed send must not have recorded any text");
loop.tick(); // same inputs — the latch must still be clear, so this must INJECT and send again
assertEquals(1, herdr.sentTexts().size(),
"a retried tick with the latch still clear must attempt the send again");
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"),
"the retried, successful send must carry the notice: " + herdr.sentTexts().get(0));
}
@Test
void jASuccessfulSendDoesConsumeIt() {
var herdr = new FailableHerdrClient(LEAD);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
loop.tick(); // opens the idle window
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
loop.tick(); // INJECT, send succeeds -> latch set
assertEquals(1, herdr.sentTexts().size());
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"), herdr.sentTexts().get(0));
loop.tick(); // same inputs — the lead was already told this stretch
assertEquals(1, herdr.sentTexts().size(),
"the next tick with the same inputs must not send a second notice");
}
@Test
void kAPendingDrivenInjectWithTheLatchAlreadySetSendsNoNotice() {
var herdr = new FailableHerdrClient(LEAD);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
loop.tick();
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
loop.tick(); // latches the notice (quiet, HIGH, cap exhausted -> the forced context INJECT)
assertEquals(1, herdr.sentTexts().size());
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"));
// Now make the fleet have real pending state, so the NEXT INJECT is pending-driven, not the
// forced-context route — with the latch already set from the tick above.
inbox.own(WORKER);
inbox.publish(WORKER, "m1", "hello");
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null));
loop.tick();
assertEquals(2, herdr.sentTexts().size(), "the pending-driven tick must still send a nudge");
assertFalse(herdr.sentTexts().get(1).contains("Your own context is nearly full"),
"a pending-driven INJECT while the latch is already set must carry no context notice: "
+ herdr.sentTexts().get(1));
}
@Test
void lTheNoticeAppearsExactlyOnceAcrossThreeDifferentlyDrivenInjects() {
var herdr = new FailableHerdrClient(LEAD);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
// quietNudgeCap=1 so a still-not-exhausted quiet nudge is available as the "forced" route below,
// distinct from both the pending-driven route and the exhausted-cap forced-context route.
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 1, scheduler);
loop.tick(); // opens the idle window
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
// 1) pending-driven INJECT: real fleet state present. Sets the latch and carries the notice.
inbox.own(WORKER);
inbox.publish(WORKER, "m1", "hello");
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null));
loop.tick();
assertEquals(1, herdr.sentTexts().size());
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"), herdr.sentTexts().get(0));
// 2) "forced" INJECT: nothing pending, but the quiet cap (1) is not yet exhausted, so decide()
// nudges anyway. The latch is already set, so no notice.
inbox.ack(WORKER, "m1");
rosterBox[0] = List.of();
loop.tick();
assertEquals(2, herdr.sentTexts().size(), "the quiet-cap-not-yet-exhausted nudge must still fire");
assertFalse(herdr.sentTexts().get(1).contains("Your own context is nearly full"), herdr.sentTexts().get(1));
// 3) pending-driven INJECT again. Still latched, still no notice.
inbox.publish(WORKER, "m2", "hello again");
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null));
loop.tick();
assertEquals(3, herdr.sentTexts().size());
assertFalse(herdr.sentTexts().get(2).contains("Your own context is nearly full"), herdr.sentTexts().get(2));
long noticeCount = herdr.sentTexts().stream()
.filter(t -> t.contains("Your own context is nearly full")).count();
assertEquals(1, noticeCount,
"the notice text must appear exactly once across all three sends: " + herdr.sentTexts());
}
/**
* Fake herdr client for the four tests above: always reports {@code lead} as IDLE, records the
* {@code text} of every {@code agent.prompt} call, and can be told to throw on the very next
* {@code agent.prompt} call — standing in for one transient herdr send failure.
*/
private static final class FailableHerdrClient implements HerdrClient {
private static final ObjectMapper MAPPER = new ObjectMapper();
private final String lead;
private final List<String> sentTexts = new ArrayList<>();
private boolean throwOnNextSend = false;
FailableHerdrClient(String lead) {
this.lead = lead;
}
void throwOnNextSend() {
throwOnNextSend = true;
}
List<String> sentTexts() {
return List.copyOf(sentTexts);
}
@Override
@SuppressWarnings("unchecked")
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", lead)
.put("agent_status", "idle"));
}
if ("agent.prompt".equals(method)) {
if (throwOnNextSend) {
throwOnNextSend = false;
throw new RuntimeException("simulated transient herdr send failure");
}
Map<String, Object> p = params instanceof Map ? (Map<String, Object>) params : Map.of();
sentTexts.add(String.valueOf(p.get("text")));
}
return MAPPER.createObjectNode();
}
@Override
public void close() {
}
}
}
@@ -0,0 +1,168 @@
package dev.ltms.fleet.msg;
import java.util.ArrayDeque;
import java.util.ArrayList;
import java.util.Collection;
import java.util.Deque;
import java.util.List;
import java.util.concurrent.Callable;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.Future;
import java.util.concurrent.RejectedExecutionException;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.ScheduledFuture;
import java.util.concurrent.TimeUnit;
/**
* A {@link ScheduledExecutorService} for tests that never runs a task on a timer — it records what
* {@link ReplyPushLoop} schedules and only ever runs it when the test itself calls
* {@link #runDueTasks()} (fleetd #608).
*
* <p>The test this exists for ({@code anAlreadyCollectedTicketProducesNoNudge}) used to wire
* {@link ReplyPushLoop} to a real {@code Executors.newSingleThreadScheduledExecutor()} and bet a
* 300ms backoff was "wide" enough that the test's own work (collecting the ticket) always won the
* race against the scheduler's own timer firing the next tick. On an idle machine that held; under
* a loaded full-suite run it did not, and the test went red on perfectly correct code. Replacing
* the timer with this fake removes the race rather than widening it: nothing here ever runs on a
* schedule of its own, so no backoff value — however small — can make the tick fire before the test
* is ready for it.
*
* <p><strong>Deliberately narrow.</strong> {@link ReplyPushLoop} calls exactly two methods on its
* {@link ScheduledExecutorService} field — {@link #schedule(Runnable, long, TimeUnit)} (every tick,
* including the very first one {@code startOrCoalesce} kicks off) and {@link #shutdownNow()} (on
* {@code ReplyPushLoop.stop()}/{@code close()}) — verified by reading every {@code scheduler.} call
* site in that class. Every other {@link ScheduledExecutorService} method throws
* {@link UnsupportedOperationException} rather than silently doing the wrong thing, so a future
* change to {@code ReplyPushLoop} that starts calling one of them fails this fake loudly, on the
* very first test that exercises it, instead of being quietly mishandled.
*/
final class ManualScheduler implements ScheduledExecutorService {
private final Deque<Runnable> pending = new ArrayDeque<>();
private boolean shutdown = false;
/**
* Run every task pending as of the START of this call — a snapshot taken before anything runs.
* {@link ReplyPushLoop#tick} always reschedules its own next tick before returning (see
* {@code scheduleNext} in its {@code INJECT}/{@code WAIT_BUSY} branches), so without the
* snapshot a single call here would recurse forever. Taking it up front means one call is
* exactly one tick, deterministically, no matter what that tick itself goes on to schedule.
*
* @return how many tasks actually ran
*/
synchronized int runDueTasks() {
List<Runnable> due = new ArrayList<>(pending);
pending.clear();
for (Runnable task : due) {
task.run();
}
return due.size();
}
/** How many tasks are currently queued, without running any of them. */
synchronized int pendingCount() {
return pending.size();
}
@Override
public synchronized ScheduledFuture<?> schedule(Runnable command, long delay, TimeUnit unit) {
if (shutdown) {
throw new RejectedExecutionException("ManualScheduler is shut down");
}
pending.add(command);
return null; // ReplyPushLoop discards the return value of every schedule() call it makes.
}
@Override
public synchronized List<Runnable> shutdownNow() {
shutdown = true;
List<Runnable> left = new ArrayList<>(pending);
pending.clear();
return left;
}
@Override
public synchronized boolean isShutdown() {
return shutdown;
}
// --- everything below: not called by ReplyPushLoop, and not supported by this fake (see the
// class javadoc's "deliberately narrow" note) -------------------------------------------------
@Override
public ScheduledFuture<?> scheduleAtFixedRate(Runnable command, long initialDelay, long period, TimeUnit unit) {
throw unsupported();
}
@Override
public ScheduledFuture<?> scheduleWithFixedDelay(Runnable command, long initialDelay, long delay, TimeUnit unit) {
throw unsupported();
}
@Override
public <V> ScheduledFuture<V> schedule(Callable<V> callable, long delay, TimeUnit unit) {
throw unsupported();
}
@Override
public void shutdown() {
throw unsupported();
}
@Override
public boolean isTerminated() {
throw unsupported();
}
@Override
public boolean awaitTermination(long timeout, TimeUnit unit) {
throw unsupported();
}
@Override
public <T> Future<T> submit(Callable<T> task) {
throw unsupported();
}
@Override
public <T> Future<T> submit(Runnable task, T result) {
throw unsupported();
}
@Override
public Future<?> submit(Runnable task) {
throw unsupported();
}
@Override
public <T> List<Future<T>> invokeAll(Collection<? extends Callable<T>> tasks) {
throw unsupported();
}
@Override
public <T> List<Future<T>> invokeAll(Collection<? extends Callable<T>> tasks, long timeout, TimeUnit unit) {
throw unsupported();
}
@Override
public <T> T invokeAny(Collection<? extends Callable<T>> tasks) throws ExecutionException {
throw unsupported();
}
@Override
public <T> T invokeAny(Collection<? extends Callable<T>> tasks, long timeout, TimeUnit unit)
throws ExecutionException {
throw unsupported();
}
@Override
public void execute(Runnable command) {
throw unsupported();
}
private static UnsupportedOperationException unsupported() {
return new UnsupportedOperationException(
"ManualScheduler only supports schedule(Runnable, long, TimeUnit), shutdownNow() and isShutdown() "
+ "— see the class javadoc");
}
}
@@ -1841,6 +1841,36 @@ class MessageServiceTest {
return new PushWiring(service, leadHerdr, scheduler);
}
/**
* A {@link MessageService} wired to a real {@link ReplyPushLoop}, like {@link #wireWithPushLoop},
* but backed by {@link ManualScheduler} instead of a real timer (fleetd #608): its tick never
* fires on its own — a test drives it explicitly via {@link ManualScheduler#runDueTasks()}.
*/
private record ManualPushWiring(MessageService service, FakeHerdr leadHerdr, ManualScheduler scheduler)
implements AutoCloseable {
@Override
public void close() {
scheduler.shutdownNow();
}
}
/**
* As {@link #wireWithPushLoop}, but the schedule never runs on its own. {@code backoffMs} is
* still threaded through to {@link ReplyPushLoop}'s constructor (it takes one), but nothing here
* ever waits it out, so its value cannot affect anything a test built on this observes — see
* {@code anAlreadyCollectedTicketProducesNoNudge}, which sets it to 1 to prove exactly that.
*/
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs) {
PrimaryRegistry registry = new PrimaryRegistry(null);
registry.recordDelegation(T, LEAD);
FakeHerdr leadHerdr = new FakeHerdr();
AgentControl leadAgents = new AgentControl(leadHerdr);
ManualScheduler scheduler = new ManualScheduler();
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, System::nanoTime);
return new ManualPushWiring(service, leadHerdr, scheduler);
}
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
long deadline = System.currentTimeMillis() + 3000;
while (!leadHerdr.called("agent.prompt") && System.currentTimeMillis() < deadline) {
@@ -1882,16 +1912,46 @@ class MessageServiceTest {
@Test
void anAlreadyCollectedTicketProducesNoNudge() throws Exception {
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires
String ticket = wiring.service().sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
// fleetd #608: backoff is 1ms — the most hostile value there is, the tick due immediately —
// and the test still must pass, because with ManualScheduler the tick never runs on a timer
// at all; it only runs when this test calls runDueTasks() below. The old version bet a 300ms
// backoff was "wide" enough that collecting the ticket always won the race against a real
// scheduler's own timer; that held on an idle machine and failed under a loaded full-suite
// run — a false red on correct code, since nothing here forced that ordering, it only made
// it likely.
try (var wiring = wireWithManualScheduler(1, 1)) {
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
// finishAsyncTask's task.future.complete(...) runs every whenComplete registered on that
// same future — including sendAsync's own hook that calls ReplyPushLoop.onTicketTerminal
// — synchronously, before complete() returns (see finishAsyncTask's javadoc and fleetd
// #399). This test-only hook fires right after that complete() call, on the very same
// (async executor) thread, so waiting for it guarantees onTicketTerminal has already run
// and the ticket is really sitting in the push loop's pending set — unlike waiting for
// Phase.DONE via poll(), which CompletableFuture.complete() can make visible to another
// thread before every whenComplete dependent has actually finished running (the same
// publish-then-run-dependents gap awaitCompletionStamped exists to close elsewhere).
String ticket;
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(terminalReached::countDown);
try {
ticket = wiring.service().sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
assertTrue(terminalReached.await(5, TimeUnit.SECONDS),
"the ticket never reached its terminal phase");
} finally {
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
}
// Collect it — the exact action the nudge must never follow.
MessageService.TaskView view = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.DONE);
assertEquals(MessageService.Phase.DONE, view.phase());
Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
// Now run the pending tick explicitly. onTicketTerminal scheduled exactly one (the ticket
// is the only thing this lead has ever had pending); it must find nothing left pending —
// the collection above already removed it — and send no nudge.
assertEquals(1, wiring.scheduler().runDueTasks(),
"expected exactly the one tick onTicketTerminal scheduled");
assertFalse(wiring.leadHerdr().called("agent.prompt"),
"a ticket the lead already polled must never be nudged");
}
@@ -328,9 +328,10 @@ class SessionManagerTest {
void rosterViewExposesTheCharterReceiptButNeverTheCharterText() {
// The roster (fleet_list and GET /members both render through rosterView) must let a lead
// see which charter a member got, without ever carrying the charter prose itself (CB-571).
String composed = "role charter\n\nreply";
CharterReceipt receipt = CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", composed);
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null,
CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", "role charter\n\nreply"), null);
0, 0, 0, MemberSession.State.READY, null, null, receipt, null);
Map<String, Object> view = SessionManager.rosterView(s, null);
@@ -338,10 +339,49 @@ class SessionManagerTest {
"the config key that supplied the role charter is reported");
assertEquals(CharterReceipt.digestOf("role charter\n\nreply"), view.get("charterSha256"),
"the digest of the exact composed charter bytes is reported");
// #604: the byte count rides alongside the digest, and must match what the receipt itself
// carries (not a hardcoded literal) so a bug that reads the wrong field is caught.
assertEquals(receipt.charterBytes(), view.get("charterBytes"),
"the exact composed byte count is reported, read from the receipt");
assertEquals(composed.getBytes(java.nio.charset.StandardCharsets.UTF_8).length, view.get("charterBytes"),
"the byte count is the real UTF-8 length of the composed charter");
assertFalse(view.values().toString().contains("role charter"),
"the roster row must not embed the charter text itself");
}
@Test
void rosterViewOmitsCharterBytesAndDigestWhenNoCharterWasComposed() {
// #604: a member with no role charter and no reply charter (composed == null) still reports
// charterSource ("none"), but neither a digest nor a size — the digest is absent, and the
// size only ever accompanies a real digest. This must not be satisfiable by code that always
// writes charterBytes.
CharterReceipt receipt = CharterReceipt.compose(MemberRole.DEV, "prof", null, null);
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null, receipt, null);
Map<String, Object> view = SessionManager.rosterView(s, null);
assertEquals(CharterReceipt.NO_SOURCE, view.get("charterSource"),
"no configured role or reply charter reports the explicit \"none\" source");
assertFalse(view.containsKey("charterSha256"), "no digest is reported when no charter was composed");
assertFalse(view.containsKey("charterBytes"), "no byte count is reported when no charter was composed");
}
@Test
void rosterViewOmitsCharterFieldsEntirelyWhenTheReceiptItselfIsAbsent() {
// #604 acceptance criterion 3: charterReceipt() can be null on its own (a session recorded
// before CB-571, or a launcher that never composed one) — the outer null-guard must still
// suppress charterSource, charterSha256 AND charterBytes together.
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null, null, null);
Map<String, Object> view = SessionManager.rosterView(s, null);
assertFalse(view.containsKey("charterSource"), "no charter fields at all when the receipt is null");
assertFalse(view.containsKey("charterSha256"), "no charter fields at all when the receipt is null");
assertFalse(view.containsKey("charterBytes"), "no charter fields at all when the receipt is null");
}
@Test
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
// The primary resolves to a Principal with no terminal, and FleetMcp's context extractor
+77 -20
View File
@@ -50,6 +50,14 @@
# absence is the real signal, because a drain that dies on its first session prints nothing
# else either. Warns loudly; never fails the redeploy, because by the time this is detectable
# the new daemon is already up and healthy.
# 10. fleetd #603 — the same shape as trap 3 above, through a different door: the step that waited
# for the NEW process to appear gave it its own short, fixed 10s budget, then hard-`die`d,
# while the health check right after it waits a full $HEALTH_WAIT (60s) for the same daemon to
# answer. Under launchd, `launchctl load` returns as soon as launchd accepts the job, before the
# java process exists, and on a slow host that took longer than 10s — so the script died with
# "no process appeared" on a deploy that had fully succeeded. The pid poll now shares
# $HEALTH_WAIT instead of a separate, shorter budget, and a miss there falls through to the
# health check (the truer signal: is it actually answering?) instead of killing the run.
#
# Usage:
# scripts/redeploy-fleetd.sh # build, confirm, restart, verify
@@ -79,7 +87,9 @@ OUT="$MODULE/fleetd.out"
PATTERN='target/fleetd.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
# poll budget below (wait_for_new_pid/await_daemon_started), so the two checks
# share one named budget instead of the pid poll holding its own shorter one
# CB-594: the launchd agent this script must not fight with (see trap 6 above).
LAUNCHD_LABEL='dev.ltms.fleetd'
@@ -1067,6 +1077,67 @@ $(tail -30 "$out_file" 2>/dev/null)"
fi
}
# fleetd #603 — the dual of wait_for_daemon_exit above: waits for a pid to APPEAR instead of
# disappear. Used to be a bare `for _ in $(seq 10)` sitting directly in the main flow, with its own
# short, fixed budget that had nothing to do with $HEALTH_WAIT (60s) — the budget the health check
# right after it gets for the very same daemon. Under launchd, `launchctl load` returns as soon as
# launchd accepts the job, before the java process exists, and on a slow host that took longer than
# 10s — so the script died with "no process appeared" on a deploy that had fully succeeded (the
# operator confirmed pid, healthz, and a fresh log line all present, by hand, right afterwards).
# Sharing $HEALTH_WAIT here removes the extra, shorter magic number without inventing a new one.
wait_for_new_pid() {
local timeout="$1" _i
for _i in $(seq "$timeout"); do
[ -n "$(running_pid)" ] && return 0
sleep 1
done
[ -n "$(running_pid)" ]
}
# fleetd #603 — the pid-appeared check and the healthz check, folded into one decision. Same shape
# as swap_if_built/refuse_drain_gate/report_shutdown_drain above (#521/#528/#512): the main flow
# calls this ONE function unconditionally, so there is no bare guard left for a future edit to
# invert independently of it. That matters more here than for most of those: a plain grep of this
# script's source cannot tell "a pid miss falls through to the health check" from "a pid miss still
# dies" apart, because both read as the same two lines of text with only the runtime branch
# changed — a source-text test could pass on either behavior. A real behavioural test on this
# function is the only thing that can actually tell them apart, which is why one exists below.
#
# wait_for_new_pid's miss is no longer fatal by itself: it falls through to the health check, which
# is direct proof the new daemon is up (/healthz answers 200) rather than a proxy for it (a process
# merely existing under a name running_pid() recognises). A genuine failure still dies here: it
# misses the pid poll AND the health poll, and report_health's own die() still prints the tail of
# $out_file, exactly as before this fix.
#
# Sets NEW_PID (global — the caller's "pid N, jar ..." result line reads it afterwards) and
# HEALTH_BODY/HEALTH_CODE (globals, the same reason report_health already needs them handed back).
# Only ever called as a bare statement in the main flow below, never from inside a `$( )`: a die()
# reached from inside a command substitution only kills that subshell, not the whole script, which
# would silently turn a genuine failure back into a false "succeeded" exit (see poll_health_body's
# own `|| true` idiom for the same hazard from the other direction).
await_daemon_started() {
local health_wait="$1" old_pid="$2" health_url="$3" out_file="$4"
NEW_PID=""
if wait_for_new_pid "$health_wait"; then
NEW_PID="$(running_pid)"
[ "$NEW_PID" != "${old_pid:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
ok "started, pid $NEW_PID"
else
warn "no process matched $PATTERN within ${health_wait}s of starting — falling through to the health check, which is the more truthful signal"
fi
HEALTH_BODY="$(poll_health_body "$health_url" "$health_wait")" || true
HEALTH_CODE="000"
[ -n "$HEALTH_BODY" ] || HEALTH_CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$health_url" 2>/dev/null || echo 000)"
report_health "$HEALTH_BODY" "$HEALTH_CODE" "$out_file" "$health_wait"
# By now /healthz has answered (report_health above would have died otherwise), so the daemon is
# confirmed up even if the pid poll never matched it — see running_pid()'s own comment on
# under-counting if the launch method ever stops being a plain `java -jar`. Fill NEW_PID in for the
# result line rather than leave it blank on an otherwise fully successful redeploy.
[ -n "$NEW_PID" ] || NEW_PID="$(running_pid)"
}
# Ticket item 8 — `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1`. Measured safe under `set -e`
# at both bash 3.2.57 and 5.x (see the header comment trap 9 discussion in the ticket) — not a `set
# -e` hazard, but still an untested computation feeding report_shutdown_drain's own four-way
@@ -1285,28 +1356,14 @@ say "start"
# there now, so this line is the only thing left in the main flow to get wrong.
dispatch_start "$SUPERVISOR_KIND"
for _ in $(seq 10); do
NEW_PID="$(running_pid)"
[ -n "$NEW_PID" ] && break
sleep 1
done
[ -n "${NEW_PID:-}" ] || die "no process appeared. Last lines of $OUT:
$(tail -20 "$OUT" 2>/dev/null)"
[ "$NEW_PID" != "${OLD_PID:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
ok "started, pid $NEW_PID"
# ------------------------------------------------------------------ verify
say "verify"
# fleetd #555: poll_health_body/health_is_up/report_health above. HEALTH_CODE is only ever
# consulted by report_health when the body came back empty; `|| true` on both assignments is the
# same "an absent/failing command substitution must not kill the script under set -e" idiom the
# swap/drain helpers already rely on (see poll_health_body's own comment).
HEALTH_BODY="$(poll_health_body "$HEALTH" "$HEALTH_WAIT")" || true
HEALTH_CODE="000"
[ -n "$HEALTH_BODY" ] || HEALTH_CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$HEALTH" 2>/dev/null || echo 000)"
report_health "$HEALTH_BODY" "$HEALTH_CODE" "$OUT" "$HEALTH_WAIT"
# fleetd #603: await_daemon_started above folds the pid-appeared check and the healthz check into
# one decision — see its own comment for why a bare `if` here, split back into the old two pieces,
# would put back an untestable branch this ticket exists to close.
await_daemon_started "$HEALTH_WAIT" "$OLD_PID" "$HEALTH" "$OUT"
# A fresh listening line, strictly after the restart mark. An old daemon that never died would
# otherwise let an old line pass for a new one.
@@ -1361,7 +1418,7 @@ report_shutdown_drain "$FRESH_LOG" "$HAD_OLD_PID"
assert_single_daemon "$(running_pid)"
say "result"
ok "pid $NEW_PID, jar $(jar_id)"
ok "pid ${NEW_PID:-unknown}, jar $(jar_id)"
if [ "$REDEPLOY_AMQP_CHECK_SKIPPED" -eq 1 ]; then
# fleetd #552: the fourth reader of the fresh-log region. Without this branch REDEPLOY_ERROR_COUNT
# stays at its untouched 0 (classify_amqp_connection_errors never got a file to read) and
+141 -2
View File
@@ -770,6 +770,137 @@ test_wait_for_daemon_exit_times_out_if_pid_never_clears() {
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
}
# fleetd #603 — wait_for_new_pid is the dual of wait_for_daemon_exit above: it must not report
# success while running_pid() still answers empty, and must report success the moment a pid
# appears. Same counter-file idiom as test_wait_for_daemon_exit_returns_true_once_pid_clears above,
# for the same reason (running_pid() runs inside a `$(...)` subshell on every call).
test_wait_for_new_pid_returns_true_once_pid_appears() {
local counter_file="$TMP/wait-new-pid-calls" final_calls
printf '0' > "$counter_file"
running_pid() {
local n
n="$(cat "$counter_file")"
n=$((n + 1))
printf '%s' "$n" > "$counter_file"
if [ "$n" -lt 3 ]; then printf ''; else printf '4242'; fi
}
sleep() { :; }
wait_for_new_pid 10 || fail "wait_for_new_pid did not report success once the pid appeared"
final_calls="$(cat "$counter_file")"
[ "$final_calls" -ge 3 ] || fail "wait_for_new_pid returned before actually re-checking running_pid"
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
}
test_wait_for_new_pid_times_out_if_pid_never_appears() {
local rc=0
running_pid() { printf ''; }
sleep() { :; }
wait_for_new_pid 3 || rc=$?
[ "$rc" -ne 0 ] || fail "wait_for_new_pid reported success while the pid never appeared"
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
}
# fleetd #603 — await_daemon_started folds the pid-appeared check and the healthz check into one
# decision (see its own comment in redeploy-fleetd.sh for why a source-text grep cannot tell the two
# possible behaviors apart here). These two tests are the ticket's own acceptance criteria, run
# together in this one suite invocation so neither can be satisfied by code that never fails at all:
#
# 1. a slow start must still succeed — running_pid mimics a process that does not appear until
# well after the OLD, buggy 10-second budget, but does appear, and healthz answers.
# 2. a genuine failure must still fail, and still print the log tail — running_pid and the health
# check both report nothing at all, ever.
test_await_daemon_started_slow_pid_then_healthy_succeeds() {
source "$ROOT/scripts/redeploy-fleetd.sh"
stub_die_recorder
local counter_file="$TMP/await-slow-pid-calls" output
printf '0' > "$counter_file"
running_pid() {
local n
n="$(cat "$counter_file")"
n=$((n + 1))
printf '%s' "$n" > "$counter_file"
# Stays empty well past the old 10-second budget, then appears — the exact shape #603 reports.
if [ "$n" -lt 12 ]; then printf ''; else printf '4242'; fi
}
sleep() { :; }
poll_health_body() { printf '{"status":"ok"}'; return 0; }
# NOT `output="$(await_daemon_started ...)"`: that would run the call in a subshell, and
# NEW_PID — a plain global assignment inside the function, by design (see its own comment) — would
# die with that subshell instead of reaching this test's own shell. Redirect to a file instead, the
# same hazard await_daemon_started's own comment warns die() itself is subject to.
await_daemon_started 60 "" "http://ignored/healthz" "$TMP/await-slow-pid.out" \
> "$TMP/await-slow-pid-output.log" 2>&1
output="$(cat "$TMP/await-slow-pid-output.log")"
[ "$DIED_CALLED" = 0 ] \
|| fail "await_daemon_started must not die on a slow-but-real start: $DIED_MESSAGE"
assert_equals "4242" "$NEW_PID" "await_daemon_started NEW_PID after a slow-but-real start"
printf '%s' "$output" | grep -qF 'started, pid 4242' \
|| fail "await_daemon_started did not report the pid once it finally appeared"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
test_await_daemon_started_never_appears_dies_with_log_tail() {
source "$ROOT/scripts/redeploy-fleetd.sh"
stub_die_recorder
local out_file="$TMP/await-never-appears.out"
printf 'boot line one\nboot line two\n' > "$out_file"
running_pid() { printf ''; }
sleep() { :; }
poll_health_body() { return 1; }
await_daemon_started 2 "" "http://127.0.0.1:1/healthz" "$out_file" > /dev/null 2>&1
[ "$DIED_CALLED" = 1 ] \
|| fail "await_daemon_started must die when the daemon never appears and never becomes healthy"
printf '%s' "$DIED_MESSAGE" | grep -qF 'never answered' \
|| fail "await_daemon_started die message does not say healthz never answered"
printf '%s' "$DIED_MESSAGE" | grep -qF 'boot line two' \
|| fail "await_daemon_started die message does not include the log tail"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
# fleetd #603 review — the path the fall-through actually exists for, and the one gap a lead
# mutation found in the first version of this test file: running_pid() NEVER finds anything (as its
# own doc comment says it eventually will, once the daemon stops being launched as a plain
# `java -jar` its allowlist recognises), while /healthz answers anyway. Neither of the two tests
# above drives this: the slow-pid test has the pid appear, so the `else` branch never runs, and the
# never-appears test fails BOTH checks, so it dies either way and cannot tell which branch fired.
# This must not die, must warn (so the operator is told the pid could not be identified), and must
# leave NEW_PID empty — the honest "could not establish this" answer, never a guessed pid, which is
# what the final result line's `${NEW_PID:-unknown}` fallback exists to print truthfully.
#
# Proof this actually pins the behavior, not just the source text (paste from a real run, not
# claimed): reverting the `warn` below back to `die "no process appeared"` (the old fleetd #603
# defect, reintroduced) turns this one test red —
# FAIL: await_daemon_started must not die when the pid is never found but healthz answers
# — and restoring `warn` turns the whole suite green again. Both halves observed, not asserted.
test_await_daemon_started_pid_never_found_but_healthy_warns_and_survives() {
source "$ROOT/scripts/redeploy-fleetd.sh"
stub_die_recorder
local out_file="$TMP/await-pid-never-found.out" output
printf 'boot line\n' > "$out_file"
running_pid() { printf ''; }
sleep() { :; }
poll_health_body() { printf '{"status":"ok"}'; return 0; }
# NOT `output="$(await_daemon_started ...)"` — see the slow-pid test above for why that would
# drop NEW_PID's assignment in a subshell instead of reaching this test's own shell.
await_daemon_started 3 "" "http://ignored/healthz" "$out_file" \
> "$TMP/await-pid-never-found-output.log" 2>&1
output="$(cat "$TMP/await-pid-never-found-output.log")"
[ "$DIED_CALLED" = 0 ] \
|| fail "await_daemon_started must not die when the pid is never found but healthz answers: $DIED_MESSAGE"
printf '%s' "$output" | grep -qF 'falling through to the health check' \
|| fail "await_daemon_started did not warn that the pid could not be identified"
assert_equals "" "$NEW_PID" \
"await_daemon_started NEW_PID when the pid is never found but healthz answers — must stay empty, never a guessed pid"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
test_await_daemon_started_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'await_daemon_started "$HEALTH_WAIT" "$OLD_PID" "$HEALTH" "$OUT"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's await_daemon_started call site in redeploy-fleetd.sh"
}
# fleetd #521 — the swap step's guard, at two levels.
#
# The first two tests call the predicate should_swap() directly. They pin its logic, and that is all
@@ -1350,9 +1481,11 @@ test_poll_health_body_returns_nonzero_when_unreachable() {
test_report_health_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'report_health "$HEALTH_BODY" "$HEALTH_CODE" "$OUT" "$HEALTH_WAIT"' "$src" | head -1 | cut -d: -f1 || true)"
# fleetd #603 — moved from a literal main-flow call into await_daemon_started (see its own
# comment above for why); this now finds the call inside that function instead.
call_line="$(grep -Fn 'report_health "$HEALTH_BODY" "$HEALTH_CODE" "$out_file" "$health_wait"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's report_health call site in redeploy-fleetd.sh"
|| fail "could not find await_daemon_started's report_health call site in redeploy-fleetd.sh"
}
# fleetd #555 item 8 — `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1`. Measured safe under
@@ -2041,6 +2174,12 @@ test_require_no_build_jar_dies_when_absent
test_require_no_build_jar_accepts_present_jar
test_wait_for_daemon_exit_returns_true_once_pid_clears
test_wait_for_daemon_exit_times_out_if_pid_never_clears
test_wait_for_new_pid_returns_true_once_pid_appears
test_wait_for_new_pid_times_out_if_pid_never_appears
test_await_daemon_started_slow_pid_then_healthy_succeeds
test_await_daemon_started_never_appears_dies_with_log_tail
test_await_daemon_started_pid_never_found_but_healthy_warns_and_survives
test_await_daemon_started_call_site_present
test_swap_ordered_after_wait_and_before_start
test_drain_gate_abort_message_says_no_no_build
test_drain_gate_refusal_build_ran_staged_present